The most interesting file on the attacker's server wasn't a payload. It was a plain markdown document called AGENTS.md.
Google Threat Intelligence Group put that detail in its AI Threat Tracker report, published September 9, 2026. GTIG found an exposed command-and-control server running a framework it calls "Recon." The first directory listing showed agent configuration and knowledge files, AGENTS.md, KNOWLEDGE.md, and agentic_vuln_research.md, sitting next to a .openclaw/ folder and a memory/ folder. Shortly after GTIG spotted it, that directory turned into a live dashboard built to "organize, validate, and manage over 23,800 harvested secrets in real time, including API keys for cloud and AI services."
In a separate case in the same report, Mandiant watched a financially motivated actor compromise an organization's cloud infrastructure and use "an AI coding chatbot, a prompt, and a set of agent instructions to plan, build, and execute a mass credential harvesting campaign in less than six hours." That campaign grabbed thousands of third-party credentials. The Hacker News, GBHackers and Biometric Update all reported both cases from GTIG's write-up, and the figures match across them.
Most of the coverage framed this as a speed story. Six hours. No human approving each step. I read it differently. I've built credential triage during red-team work, and the word GTIG keeps using, validate, points at the moment an attacker is loudest. An operation that checks every secret it collects is an operation that trips every trap you leave in the pile.
Why a stolen-secret pile needs sorting at all
When you pull secrets off a compromised host, most of them are dead weight. Expired tokens. Rotated passwords. Test keys from a staging account that got torn down a year ago. Placeholder strings someone pasted into a .env.example and never removed. The raw haul from one build server might be a few hundred strings, and only a small fraction will still open anything.
So the harvester's whole value is the sort. A dashboard that holds 23,800 secrets is useless as a flat list. What makes it a weapon is the column next to each entry: still works, or doesn't. Which account. Which cloud. How much that account can spend. GTIG describes exactly this evolution, a move away from "conventional infostealers that passively search infected endpoints for passwords and browser data" toward a system that actively assesses what it holds and feeds the live ones into the next stage.
That column has to be populated somehow. The only way to know a key still works is to present it to the service it belongs to and see if the service answers. There's no offline check. You cannot tell from the shape of a token whether the account behind it was closed yesterday. You have to knock.
And that is the hinge this entire report turns on for defenders. The agent knocks on every door, including the ones that were never real.
There's a second detail worth pausing on. GTIG notes the Recon dashboard didn't just mark keys live or dead, it organized them, which means it was reasoning about value. A key that opens a cloud account with a fat billing limit is worth more than one that opens a read-only logging bucket, and a harvester holding 23,800 secrets has to prioritize or it drowns. That prioritization is itself a tell. It implies the operator cares which accounts spend money, which maps cleanly onto the financially motivated attribution. It also means a planted key that advertises itself as high value, a decoy named like a production billing credential or a root service account, isn't just one more row to test. It's the row the agent rushes to first.
A honeytoken is a door that only rings an alarm
Here's the asymmetry I want red and blue teams to sit with.
A real credential, when an attacker checks it, does its job silently. The provider answers, the key works, nothing fires. That success is invisible to you unless you've instrumented every provider you use, which almost nobody has done well.
A honeytoken is the inverse. It's a credential that looks exactly like a live one, lives where a real one would live, and has exactly one job: when anyone anywhere tries to use it, it tells you. No legitimate process ever touches it, so there is no baseline of normal activity to tune out. The first time it's used, that use is the attacker. That's the zero-false-positive property we build Mine2 around, and it's not a marketing line here, it's a structural fact about bait that nothing real references.
Now layer the two cases together. The Recon operator built a machine whose core function is to try thousands of keys against live services as fast as possible. If even one of those 23,800 entries was a planted token, the machine announced the intrusion the instant it reached that row. The attacker's own efficiency, the thing every headline admired, is what rings the bell. The faster the agent sorts, the sooner you know.
I don't say this to pretend a honeytoken would have unwound the whole campaign. It wouldn't. But detection is about getting a clean, high-confidence signal early, and a harvester that validates in bulk is the most cooperative adversary a deception grid will ever meet. It volunteers to test your bait.
The six-hour campaign makes the window smaller, not the problem harder
The Mandiant case is the one that unsettles people, because the time from cloud compromise to mass harvesting was under six hours, and the agent handled the tedious parts itself: it managed the scanning pipeline, troubleshot its own errors, and rotated IP addresses so the traffic came from legitimate-looking cloud ranges. GTIG is blunt that this "significantly reduced the human-in-the-loop latency."
Reduced latency cuts both ways, and I think defenders are reading only one side of it. Yes, the attacker moves faster. But the attacker also removed the human pause that used to protect them. A careful operator might eyeball a batch of keys, notice that one belongs to an account named billing-prod-canary, and skip it. An agent chewing through a scan queue at machine speed with a markdown playbook telling it to validate everything does not get suspicious. It has no instinct to protect. It checks the canary with the same indifference it checks a real key.
So the compression that makes agentic attacks scary is also what makes decoy-based detection work better against them, not worse. Six hours of autonomous validation is six hours of an attacker interrogating every credential it can find, with no judgment filtering out the ones that bite back. We saw the money side of this gap last week when METR's own agent handed over its API key and an attacker burned about $600,000 in credits over three weeks. The difference between that story and this one is a planted key would have screamed on first use instead of three weeks later on the invoice.
This is the same lesson the supply-chain half of the report teaches
The GTIG report doesn't stop at harvesting. It also details DUSTMAKER, malware from the group it tracks as UNC6780, which hunts for OIDC tokens in the memory of GitHub Actions runners, then uses them to publish poisoned packages with valid signed attestations so automated trust checks pass. It drops malicious config into hidden AI-assistant directories like .claude/ and .vscode/, and in one genuinely strange touch, it embeds extreme adversarial prompts in its loaders, nonsense about weapons, specifically to make LLM security scanners refuse or skip analysis of the code underneath.
Every one of those tricks is built to defeat checks that reason about content. Is this package signed? Yes. Does this token authenticate? Yes. Does this file look like a normal workspace config? Yes. The attacker has learned to make the real signals lie. That's the same arms race we wrote about when the BigBear 2.0 phishing kit started deleting canary tokens from stolen sessions and when a forged checksum broke Chrome's extension trust model. When adversaries can forge legitimacy, the defenses that hold up are the ones that don't depend on a judgment call.
A decoy doesn't judge. It doesn't ask whether the caller is authorized, whether the signature is valid, or whether the prompt is adversarial. It asserts one fact: I am fake, nothing should ever touch me, and something just did. An agent that validates 23,800 secrets cannot route around that, because the trap isn't a check it can pass. It's a landmine that only exists to be stepped on.
What I'd actually do this week
If you run cloud infrastructure or a CI pipeline, two moves cost you almost nothing and change what these campaigns see.
First, salt the places harvesters read. The Recon framework and the six-hour campaign both work by scraping secrets out of environments, files, and memory, then testing them. Put planted AWS keys, fake provider tokens, and decoy service credentials exactly where a scanner would expect to find the real ones: in environment variables, config files, CI secret stores, the .env a careless commit would leak. Our honeytoken product is built to make those indistinguishable from live secrets and to alert the moment one is used, with nothing to tune and no false alarms to chase.
Second, stand up decoy services that have no business receiving traffic, so that automated scanning lights them up before it reaches anything real. A harvesting agent maps before it validates, and a MineField decoy turns that reconnaissance into a dated, attributable alert instead of silence.
GTIG's report is being read as a warning that attackers now move at machine speed. Fair. But machine speed with no human judgment is an adversary that will test every key you own, including the ones you planted to catch it. If you want to see how fast a decoy credential turns an autonomous harvester into a confirmed intrusion, book a Mine2 demo and watch one trip on the first key it tries.
Monty
Offensive Security Lead, Mine2
Monty leads offensive security research at Mine2, breaking down how attackers turn vulnerabilities into footholds — and where deception trips them up first.
Recent Articles
27 Seconds: What CrowdStrike's 2026 Threat Report Really Means for Your Detection Stack
Your EDR Is Dead — Now What? Why Deception Is the Detection Layer That Survives EDR Killers
Ransomware's Invisible Kill Chain: Why Lateral Movement Is the Phase Your EDR Can't See
Need Security Help?
Protect your organization with MINE2's cyber deception platform.
