The AI offence-defence debate is asking the wrong question

Image: Giovanni Battista Tiepolo, 1696-1770, Illustration for a Book: Soldiers Storming a City. Courtesy Metropolitan Museum of Art

Image: Giovanni Battista Tiepolo, 1696-1770, Illustration for a Book: Soldiers Storming a City 

Courtesy Metropolitan Museum of Art

07 July 2026

On 7 April 2026, Anthropic announced that an unreleased frontier model, Claude Mythos Preview, had found thousands of zero-day vulnerabilities – on its own, across every major operating system and web browser. Some of the flaws had survived decades of human review, including a 27-year-old bug in the OpenBSD operating system. In one case, the model built a working browser exploit by chaining four of them together. 

Anthropic judged the model too dangerous to release. Instead, it launched Project Glasswing, giving a vetted set of technology companies and infrastructure operators gated access so they could patch their own code first. It has since extended the programme to roughly 150 more organisations across more than 15 countries.

In the two months since then, the announcement has revived one of the oldest questions in security studies: whether new technology favours the attacker or the defender, and has drawn two confident and opposite answers. One camp says the balance has broken decisively toward offence: AI-scale vulnerability discovery collapses the window between finding a flaw and exploiting it, leaves defenders structurally outmatched, and marks a genuine inflexion point. 

The other camp is optimistic. Anthropic itself describes Glasswing as a bid to secure ‘a permanent advantage for defenders.’ The most rigorous version of that case is academic and predates Mythos: because offense in cyberspace depends on creative deception that AI cannot supply, while defence mostly needs detection at scale, automation should systematically favour the defender. The argument rests on the steep, multi-year decline in attacker ‘dwell time’ – the amount of time the intruder has access to a system – recorded in a 2025 report from Mandiant (the 2026 report showed it rising from 11 to 14 days).

Both camps are wrong in the same way. They treat the offence-defence balance as something that can be settled by technology, a property baked into the model’s capabilities. The evidence since April points the other way. AI’s effect on the balance is conditional. It turns on who is using it, what kind of operation they are running, and whether the institutions that have to convert a technical finding into a deployed defence can actually do so.

Attacker intent matters more than the model

When pessimists and optimists argue about AI and cyber conflict, they tend to picture the same things: the exquisite sabotage campaign, like Stuxnet and the attacks on Ukraine’s power grid, which takes years to prepare and turns on the kind of creative deception AI still struggles to produce. From that shared image, they draw opposite conclusions. The optimists believe AI adds little to serious offense; the pessimists think it is about to make such attacks easy for anyone.

But that shared image is much of the problem. It is the rarest, hardest-to-combat end of the threat spectrum, and the place where AI changes the least. The pressure is at the other end: the high-volume criminal operations where AI’s advantages, like scale and automation, do the most work. This is where the real offence-defence answer can be found. 

Most of what advanced persistent threat groups do falls outside of the realm of sabotage, in two other categories. One is espionage: sustained access campaigns aimed at political and economic intelligence. The other, revenue generation, covers financial theft, cryptocurrency schemes, and IT worker fraud – North Korea’s signature. And scale is what AI supplies.

In November 2025, Anthropic reported that a group it assessed with high confidence as sponsored by the Chinese state had manipulated its coding agent, Claude Code, into conducting espionage intrusions against roughly 30 organisations. The AI did 80 to 90% of the tactical work, and humans stepped in at only a handful of decision points. 

It also overstated its findings and sometimes invented results outright. AI is not yet a flawless attacker, but flawless was never the bar for spying and theft. Cheap and good enough is the bar, and for low-stakes operations it has already been cleared. For sabotage, it has not. The technology is identical in both cases; what changes the answer is the attacker.

Discovery is not defence

The defender’s side is where Mythos exposes a gap the optimists tend to omit: finding a vulnerability and fixing it are not the same task. 

For an attacker, finding one exploitable vulnerability is almost the whole game: a single unpatched flaw in a deployed system will do. 

Defenders face something larger and differently shaped: they have to determine which identified bugs are actually exploitable, weigh the severity of each, and close every gap that matters. For a defender, an AI-discovered flaw is where the work begins: triage, prioritisation, getting code owners on board, patching, testing, and rolling the fix out across every affected system. That pipeline runs on budgets, change windows, and office politics. AI has dramatically compressed the discovery half of it. There is little sign it has compressed the rest of the process. Anthropic claims Glasswing surfaced more than 10,000 high- or critical-severity vulnerabilities in its first two months, a figure that counts as a defensive win only if remediation capacity grows to match it.

Attackers need to be right once, but defenders need to be right every time. As a result, a capability that appears balanced in controlled environments often leads to a highly unequal advantage in real-world situations. 

If the model by itself handed defenders the advantage, Anthropic could have shipped it. Instead, the company has had to build an institutional workaround: vetting organisations for their capacity to absorb the findings, sequencing the disclosures, and coordinating with vendors and governments. (This was inhibited by the since-rescinded June bans by the US government on foreign access to Fable 5 and Mythos 5 models, for national security reasons.) The defender’s advantage, in other words, is being manufactured through institutional design, not conferred by the model. 

Independent testing points in the same way. The UK’s AI Security Institute reportedly found that Mythos cannot reliably carry out autonomous attacks against well-hardened targets. Where institutions are strong, the optimists are right. Where they are weak, like trailing-edge firms that can’t patch at machine speed and legacy systems that can’t be patched at all, the pessimists are. 

We only count what defenders catch

That brings us to the evidence itself. The dwell-time decline that anchors the optimistic case is assembled from detected incidents: operations that left sufficient telemetry or forensic evidence to be added to an incident-response firm’s caseload. That is not a random sample of cyber operations. It is the subset that failed at staying hidden. 

As my coauthors and I argue in recent work on the ‘missingness’ problem in cyber-conflict data, the observable record is the tip of an iceberg whose hidden mass is shaped by attackers’ own decisions about stealth. An AI-assisted operation whose main advance is improved evasion would, by definition, never appear in the data at all. Falling dwell times are consistent with defenders winning. They are just as consistent with attackers generating an ever-larger pool of operations that are simply never found.

The same blind spot limits our understanding of AI-enabled offence. Anthropic could see the 2025 espionage campaign for one reason: the attackers ran Anthropic’s model on Anthropic’s infrastructure. An operator running its own models on its own servers leaves no such trace. The public record of AI-enabled attacks is not a census; it captures only the attackers who happened to operate within view of the handful of firms that publish threat intelligence. Neither camp should be confident in a balance we measure poorly.

Reading the balance as conditional points to two priorities. The first is to resist both fatalism and complacency: ‘offense has won’ and ‘defenders now hold a durable advantage’ are each half-true, and each licenses bad policy: by inviting us to abandon hardening as futile, or by treating it as unnecessary. The second is to disaggregate the threat model. Sabotage, espionage, and state-sponsored cybercrime have different tolerances for detection, and therefore very different exposure to AI advances. A posture calibrated to the hardest case, the meticulous sabotage operation, will systematically misjudge the far more common ones.