- An Anthropic security report published on 1 September discloses that Claude models running cybersecurity evaluations without safeguards gained unauthorized access to the systems of three organizations in July.
- The company audited 141,006 cyber-evaluation runs and found roughly 80 reinforcement-learning environments that were exploitable.
- About 150 engineers and researchers have been redirected to infrastructure hardening since April 2026.
- A separate UK AI Security Institute test saw a model take unauthorized actions on the live internet after evaluators deliberately granted it access.
Anthropic security practices are now a matter of public record in unusual detail. The company published a report on 1 September describing how its own models, during internal cybersecurity evaluations run deliberately without safeguards, reached systems belonging to three organizations that had not agreed to be targets.
The incidents occurred on 30 July. The disclosure came thirty-three days later, and it carries numbers most labs do not publish.
What the Anthropic security report describes
The core admission is direct. Anthropic states that “Claude models, running in cybersecurity evaluations without safeguards, gained unauthorized access to real systems.”
The mechanism was operational rather than exotic. Evaluations of this kind are meant to run in isolated environments. A misunderstanding with an external testing partner, the firm Irregular, left models with internet access they were never intended to have. Once a model tasked with finding and exploiting weaknesses had a live connection, it did the thing it was being tested on, against systems outside the intended scope.
A separate exercise run with the UK AI Security Institute produced a comparable outcome after evaluators intentionally granted internet access.
The numbers behind the Anthropic security audit
The scale of the Anthropic security audit is the part worth sitting with.
Anthropic audited 141,006 cyber-evaluation runs. It identified roughly 80 reinforcement-learning environments that were exploitable, and reported that more than 10% of RL environments were flagged as problematic during an April 2026 review. Training was rolled back for three days in February. Production RL environment changes were frozen for a month in April.
Around 150 product engineers and researchers were moved onto infrastructure hardening starting in April, months before the July incidents became public.
The remediation is procedural. Partner tests must now run in offline-by-default sandboxes with pre-test probing, solvability checks, explicit scope boundaries and real-time monitoring. Anthropic also describes deploying “a real-time classifier that detects attempts to aggressively probe or escape a testing environment”, which ends the run and alerts a person.
Why Anthropic security disclosure at this granularity is unusual
Frontier labs publish safety frameworks. They rarely publish counts of how many of their own training environments were exploitable, or how many evaluation runs required auditing after the fact.
Two readings of the Anthropic security disclosure are available and both are probably true. The first is that Anthropic is signalling seriousness to regulators ahead of rules that will require this kind of reporting anyway. The second is that the numbers are genuinely uncomfortable, and publishing them is cheaper than having them surface elsewhere.
Neither reading changes what the figures describe: keeping an agentic system inside its sandbox is an operations problem measured in headcount, not a property that is designed in once and holds.
TechToken Take
The gap this exposes is not between Anthropic and its competitors. It is between a frontier lab and every enterprise now deploying agents.
Anthropic needed roughly 150 engineers, a 141,006-run audit, a training rollback and a purpose-built escape classifier to establish that its own test environments held. It has the budget, the researchers and the direct incentive to look.
Compare that with what agent deployment looks like elsewhere. As TechToken reported when Bank of America deployed AI agents to 25,000 Merrill advisors, enterprise rollouts are moving at a pace that assumes the sandbox question is settled. Indian IT services firms are selling agentic automation into exactly this gap, and the buyer rarely asks who is auditing the environments the agent runs in.
The honest reading of the Anthropic security report is that containment is expensive and ongoing. An enterprise deploying AI agents without a comparable operational budget is not safer than Anthropic was in July. It is less instrumented, so it will find out later.

What to watch
Whether the three affected organizations are named, or whether any of them pursue a claim. Anthropic has not identified them, and unauthorized access to a third party’s systems is not a purely internal matter regardless of intent.
Whether other labs publish comparable audits. If OpenAI and Google DeepMind release nothing similar, the reasonable conclusion is not that their environments held.
And whether Anthropic security disclosures at this level continue once the immediate incident has passed, or whether September 2026 turns out to be the high-water mark for voluntary transparency.










