The attacker didn't exploit anything clever. They typed a request into a chat box and the agent answered.
That's the core of the first incident METR, the AI evaluation nonprofit, disclosed in a security update on August 31, 2026. In March, someone found an agent orchestration dashboard running on a researcher's personal EC2 instance, "prompted an agent directly to reveal its model provider API key," added an SSH key for persistence, and consumed about $600,000 in inference credits over three weeks. METR's own post gives that figure, and The Hacker News and Dark Reading both reported it from the same disclosure. (One secondary write-up circulating on DEV Community says $60,000. It dropped a zero. METR's number is $600,000.)
Nobody paid cash for it. The model developer had granted those credits to METR for free and absorbed the loss. That's the only reason this story reads as an embarrassment instead of a budget crisis, and it's why I think detection teams should read it more closely than they will. Swap "donated research credits" for your production OpenAI or Bedrock account and the same three weeks would have landed on your invoice.
I've spent a lot of this year writing about keys leaking out of dev tooling. This one is different in a way that matters for how you design detection, so let me walk the attacker's path step by step and point at the signal each step left behind.
Step one: the attacker read the certificate logs
METR says it suspects the attacker "found the instance by looking through recently-registered websites (e.g. in certificate transparency lists)."
If you've never watched this happen live, it's worth doing once. Certificate Transparency logs publish every publicly trusted TLS certificate, and the moment a researcher spins up agent-dashboard.something.dev behind Let's Encrypt, that hostname is in a public, searchable, streaming feed. Scanners tail those feeds and hit new hostnames within minutes. When I've stood up throwaway hosts for testing, the first unsolicited requests usually beat me to my own browser.
The METR attacker apparently went a step further and filtered for LLM and agent keywords. That's a smart filter. A hostname with "agent" or "eval" in it, freshly issued, sitting on a personal cloud account, is very likely to be a hand-built tool with a model provider key somewhere inside it.
The signal: a request to a hostname that has never been shared with anyone. A brand new host has no legitimate traffic except its owner. Everything else that shows up is someone reading CT logs.
Step two: the login screen quietly stopped being a login screen
The instance was meant to sit behind Google authentication. METR describes the orchestration tool, which Dark Reading notes was vibe-coded, as containing "a fail-open vulnerability that silently disabled authentication." For several days the dashboard was exposed to the public internet with no auth at all.
Fail-open bugs are the most boring vulnerability class there is, and the one I'd bet on in any internally built AI tool. The pattern is almost always the same: the auth middleware hits an error (a missing env var, an expired OAuth client secret, a config file that didn't load) and a try/except around it logs the problem and lets the request through. Generated code is especially prone to this because the model is optimizing for "the app runs," and an app that crashes on a missing secret doesn't run.
Here's the part that should bother you. Nothing about this looks like an attack in your logs. There's no failed login, no brute force, no impossible travel. The identity provider never saw a request at all, because the app stopped asking it. Every identity-based detection you own is blind to an authentication layer that has switched itself off.
The signal: none, from the identity side. The only way to see this is from the application side or from the thing being protected.
Step three: the agent was the credential store
This is where the METR incident stops being a generic "exposed dashboard" story.
The attacker didn't need to find a .env file, read process memory, or pivot to the instance metadata service. They asked the agent. An agent that can run tools, read its own environment, and describe its configuration is, functionally, an interactive API for its own secrets. Whatever is in its context window or its reachable environment is one polite prompt away from whoever controls the chat box.
I'd put it this way: if an agent can use a credential, assume anyone who can talk to the agent can read it. System prompts that say "never reveal your API key" are not access control. They're a suggestion to a text generator, and people have been talking models out of those suggestions since 2023.
This changes where secrets should live. In the traditional app model we argued about env vars versus a secrets manager. With agents the question is different: does the agent process ever hold a long-lived provider key at all? If it does, the key's exposure boundary is no longer the server, it's the conversation. METR's post-incident list includes spend alerts and "monitoring for unusual API key usage," which is correct, but the architectural fix is to keep the real key out of the agent's reach entirely. A local proxy or broker that holds the provider key and hands the agent a short-lived, scoped, rate-limited token turns "reveal your key" into "reveal a token that dies in an hour."
We covered a cousin of this problem in 294 LiteLLM gateways still answering to the default sk-1234 key. The gateway pattern is the right idea. It only helps if the gateway itself isn't the next exposed dashboard.
The signal: the key being used from somewhere other than the agent's host. That's detectable on day one, if you're watching the right thing.
Step four: an SSH key, then three quiet weeks
The attacker added an SSH key for persistence. Then they spent.
Why did it take three weeks to notice? METR is candid: high baseline token usage from large-scale evaluations, incomplete monitoring dashboards, and "no way to put a spending limit on keys like this one." Read that list as a detection engineer and you'll recognize every item.
- High baseline. Anomaly detection on volume fails when legitimate volume is already huge and spiky. An eval lab that routinely burns through millions of tokens in a burst is the worst possible place to set a threshold. Plenty of enterprise AI platform teams now look the same.
- Incomplete dashboards. Provider usage consoles are billing tools, not security tools. They'll tell you what you spent. They mostly won't tell you which source IP spent it.
- No spend cap. Some provider keys support hard limits, some don't, and grant-funded or enterprise-negotiated keys are often the ones without guardrails.
The SSH key is the detail I'd want every cloud team to take away. An authorized_keys change on a personal research box doesn't trip anything in most organizations, because personal research boxes aren't in the asset inventory to begin with. That's the shadow infrastructure problem in one line.
The signal: a new key in authorized_keys, and a login from an IP that has never touched the account.
The May incident shows where this is heading
METR's second incident is less dramatic and more instructive. In May it watched attackers "systematically probing our publicly accessible infrastructure, with heavy use of agents to automate vulnerability discovery, including by credential stuffing authentication providers, attempting OAuth token grants, scanning newly deployed services," and phishing staff. Separately, an independent researcher reported that METR's public transcript viewer exposed a read-only SQL query mechanism that could reach unpublished evaluation data. METR took that API offline and says it has no evidence the attackers used it.
Notice the list. Credential stuffing, OAuth grants, fresh-deployment scanning, phishing. Every item is an identity attack, and the attacker ran them in parallel with agents. That lines up with what Unit 42 reported in its 2026 Global Incident Response Report: identity weaknesses played a material role in nearly 90% of its roughly 750 investigations, and about 65% of initial access came from identity-based techniques. The same report put the fastest quarter of intrusions at 72 minutes from access to exfiltration.
It also lines up with Google Threat Intelligence Group's September 8 reporting on TeamPCP, a financially motivated group that used an autonomous multi-agent framework to harvest thousands of third-party credentials in under six hours. We wrote about a similar pipeline in the Zerofot scanner, which pulled 2,975 keys in 48 days. The scanning side of credential theft is automated end to end now. What METR adds is the other half: attackers are using agents to find targets, and they're finding agents as targets.
What I'd actually deploy after reading this
Each of the four steps above left a signal. The trouble is that three of them are signals you normally can't see, because the thing emitting them (a personal EC2 box, a vibe-coded dashboard, a donated API key) isn't on anyone's watch list. You can't tune a SIEM rule for infrastructure you don't know exists.
That's why I keep coming back to deception for this class of problem. A decoy doesn't need a baseline, and it doesn't need an inventory. It needs one property: nobody legitimate ever touches it.
Put a canary key where the agent will find it first. If your agents need provider access, give them a brokered token for real work and put a honeytoken key in the environment, named the way a real one would be (OPENAI_API_KEY, ANTHROPIC_API_KEY_PROD). An attacker who asks the agent for "its key" gets the canary, and the first API call with it fires an alert with the source IP. Compare that with METR's three weeks. There's an edge case worth planning for: if the agent can dump its whole environment, the attacker gets both values. That's fine. They'll almost always try the one that looks most like production, and you only need them to try one.
Plant a decoy on a freshly certificated hostname. Stand up a host with an agent-sounding name, get it a real certificate so it lands in CT logs, and put a decoy service behind it. Every connection is a scanner that's filtering CT feeds the way METR's attacker did. You learn who's hunting for AI tooling in your namespace, and their source ranges, before they find the real one. A MineField decoy service is built for exactly this: it listens, looks plausible, and has zero legitimate users, so there's nothing to tune.
Seed decoy credentials in the places a dashboard owner would stash them. A fake SSH private key in the home directory, a fake cloud credentials file, a fake .env next to the real app. If someone gets shell on the box, as METR's attacker did, they'll go looking. We showed how quickly attackers pull .env files from exposed dev servers in our analysis of the Vite CVE-2026-39364 harvesting wave. The same reflex works in your favor when the file is a trap.
Test your auth for fail-open, on purpose. Break the OAuth client secret on a staging copy of every internal AI tool and see whether it denies or waves you through. It's a ten-minute test. I'd bet a good share of teams running internally built agent dashboards will find at least one that fails open.
None of this replaces the fundamentals METR listed: review before anything goes public, spend alerts, short-lived keys, no organizational credentials in personal cloud accounts. Those are right. But fundamentals depend on people following process, and the METR incident started with a researcher doing something reasonable and a tool doing something silent. Decoys cover the gap between the policy and the box nobody told you about.
METR deserves credit for publishing this in detail. Most organizations would have quietly rotated the key and moved on, and the rest of us would have learned nothing. If you want to see what a canary key in an agent environment or a CT-log decoy looks like in your own setup, book a Mine2 demo and we'll build one with you.
Arjun
Lead Detection Engineer, Mine2
Arjun builds detection logic at Mine2, focusing on the blind spots EDR and SIEM leave behind and how honeytokens close them.
Recent Articles
32,000 Requests for Your .env: How a Vite Dev Server Bug Became a Cloud Key Harvester
2,975 Keys in 48 Days: The Zerofot Scanner Proves Credential Theft Has Been Automated End to End
Your API Keys Don't Have MFA: Why Non-Human Identities Are the Biggest Blind Spot in Enterprise Security
Need Security Help?
Protect your organization with MINE2's cyber deception platform.
