Attackers Started Using AI Before Companies Did. 12,000 Inboxes, 10,000 Organisations.

Security & AI Governance

· 9 min read

Microsoft disrupted EvilTokens, the first cybercrime service where AI ran across the attack chain. What the case reveals about phishing, non-human identities and AI agent governance.

There is an irony in the 2026 AI landscape worth naming directly.

While many companies are still debating whether AI justifies the investment, attackers finished that debate long ago. They implemented, scaled and sold the result as a subscription service.

On 22 September 2026, Microsoft's Digital Crimes Unit obtained authorisation from the US District Court for the Eastern District of Virginia and seized 50 websites and more than 150 domains operating EvilTokens — a phishing-as-a-service platform in which AI ran the entire attack chain.

The toll: more than 12,000 compromised Microsoft 365 inboxes across more than 10,000 organisations worldwide. Victims concentrated in the United States, Canada, the UK, France, Australia and India.

It is Microsoft DCU's 40th court-authorised disruption — and its first against an end-to-end AI-enabled cybercrime service.

How It Actually Worked

The technical detail matters, because it explains why conventional defences did not stop it.

The entry: device-code phishing. The attacker initiates the sign-in themselves and sends the resulting code to the victim in a phishing lure. The victim enters it on the genuine microsoft.com/devicelogin page, and the attacker's session enters the account without any password changing hands — which also gets past multi-factor authentication.

Note the detail: no password stolen, 2FA bypassed. And according to the analysis, access survived password resets, because sessions and tokens remained valid.

The lures: 44 templates adapted by AI. Not generic emails with grammatical mistakes. Lures automatically tailored to each target's professional role.

The exploitation: a chatbot that read the inbox for the attacker. This is the new part. Once inside the account, a chatbot analysed the victim's messages: surfacing payment-related threads, charting who held which role in the organisation, flagging trusted relationships and suggesting where fraud was most likely to succeed.

Steven Masada, General Manager of Microsoft's DCU, summarised the shift in a single sentence: once an inbox falls, criminals may understand its contents in minutes, not days.

The platform sold on Telegram as a monthly subscription. Microsoft noted that it consolidated functions that previously required separate expertise in identity exploitation, cloud environments, social engineering and financial fraud — delivered through a single commercial interface.

A Coinbase investigator involved in the case described the effect bluntly: it completely obliterates the barrier to entry for phishing-as-a-service, and can be operated at industrial scale.

This Is Not an Isolated Case

In the same week, Cisco Talos documented CLOSEDQUORUM — described as the first fully autonomous command-and-control implant built on multiple AI models. Cisco simultaneously released CAIRN, a tool to help security teams detect AI-integrated malware. Cisco Talos noted there was no confirmation of active in-the-wild deployment.

And Akamai published a brief arguing that enterprise risk is shifting toward governing autonomous non-human identities.

That, in fact, is September's most important conclusion. And it is not only about attackers.

The Part Being Discussed Too Little: Your Agents Are Identities Too

If you are deploying AI agents in your company — and more companies are — you have created non-human identities with access to your systems.

An agent that reads from your ERP, writes to your CRM and sends emails is, from a security standpoint, a user. One that does not sleep, does not tire and does not pause to wonder whether a request looks strange.

The market recognised this in September. On 24 September, Dataiku launched Agent Management — a standalone product that inventories AI agents across the organisation, tracks their performance and tiers them by risk. Noma released governance and protection tooling for agents at the endpoint level. Alibaba unveiled AgentCore, an enterprise platform for building and managing agents across critical business systems.

Three announcements, one theme: AI agents have become an identity category that must be inventoried, monitored and governed.

Companies that deployed agents without thinking about this now have users with broad permissions, no audit, no inventory and no clear answer to "who authorised this?"

What You Can Do This Week

Concrete recommendations, most at no cost:

Block the device-code flow wherever you do not need it. This is Microsoft's explicit recommendation following the investigation. Few organisations use it legitimately; many have it enabled by default.

If you suspect a compromised account, temporarily disable it — do not stop at a password reset. Microsoft warns that revoking sessions can leave access tokens valid for up to an hour. A new password does not close an already-stolen session.

Take an inventory of non-human identities. Service accounts, integrations, API keys, AI agents, automations. For each: what it accesses, who authorised it, when it was last reviewed.

Check what your AI agents can do, not what they should be able to do. The difference between policy and architecture is exactly the difference between a document and actual protection.

Train your team on this specific pattern. A request to enter a code on a legitimate Microsoft page does not look like classic phishing. That is precisely why it worked on 12,000 inboxes.

How We Build, and Why It Matters Now

Our approach at Visual AI Labs has always been that governance belongs in the architecture, not as a layer added after delivery. September's events confirm why.

Architectural access limits, not declarative ones. If an agent does not need access to the HR system, that access does not technically exist — it is not merely forbidden in a document.

Immutable audit trails for every action. What the agent accessed, what decision it made, what it executed, when. If an incident occurs, you can reconstruct exactly what happened.

A distinct identity per agent. Not one shared service account used by three different automations. Each agent is a separate identity, independently revocable.

A stop mechanism known to at least two people in the organisation. Practically, not theoretically.

Data in the EU, with GDPR and AI Act compliance designed in from the start — not bolted on when the first audit arrives.

We deliver fast, without replacing the systems that already work for you.

If You Have AI Agents in Production and No Inventory of Them

It is worth a conversation. Not because your system is certainly compromised — but because you cannot answer "what can this agent do right now?" without checking.

Write to us on the contact page with a short description of what you have deployed. We will come back with an honest assessment of what we would check first.

Write to us on the contact page →

Verified sources: Microsoft Security Blog — https://www.microsoft.com/en-us/security/blog/2026/09/22/unmasking-eviltokens-getting-to-the-root-of-device-code-phishing/; Microsoft On the Issues — https://blogs.microsoft.com/on-the-issues/2026/09/22/disrupting-eviltokens-the-ai-chatbot-built-for-cybercrime/; — *Disrupting EvilTokens* (22 Sept. 2026); Axios (22 Sept. 2026); Fortune / Coinbase (22 Sept. 2026); The Hacker News (Sept. 2026); Security Boulevard (Sept. 2026); Cyber Kendra (Sept. 2026); Enterprise Times — *Security and AI news, week of 21 Sept. 2026*; AI Agent Store — *AI Agents News, week of 25 Sept. 2026*

Contact