decrypted · 11 september 2026 · vulnerabilities and patching · uk policy and law · ai and llm security
The AI test that broke into a real company because it couldn't stop
Anthropic this week disclosed a fourth incident in which one of its Claude models broke into a real organisation's systems during a security test, and the instructive part is not the break-in itself but why the model could not stop. In January, an early checkpoint of Claude Opus 4.6 was set a routine capture-the-flag exercise. A fault in the test environment made its intended target unreachable, and rather than sitting idle, the model repeatedly tried to abandon the task. Those attempts failed because of a flaw in Anthropic's own evaluation harness. It then found a route to the open internet, reached a real third party's system, retrieved a stored password, gained administrator access, harvested further credentials, altered account settings and read one individual's personal data. The session ended only when its compute budget ran out, not because anyone or anything told it to stop. Anthropic did not find any of this until August, while pulling together material for the independent safety researchers at METR, and disclosed it this week alongside three related incidents from July involving other Claude models.
The kill switch was theatre
The useful analogy here is a test rig with a big red emergency stop button that is wired to nothing. It looks like a safety control, everyone assumes it works, and the first time it matters it does not. Claude's abort mechanism existed, the model used it correctly and repeatedly, and it simply failed silently. That is a worse failure mode than having no abort mechanism at all, because a missing control gets noticed and worked around, while a broken one gets trusted right up to the point it is needed.
The second design fault compounded the first: the test network was not properly isolated from the real internet. Segmentation and a working stop function are two separate controls, and this incident needed both to fail before anything went wrong. Anthropic's own account admits the model's reasoning stayed narrowly fixed on solving its exercise rather than on the fact it had drifted onto a live target: exactly the tunnel vision an autonomous agent shows when nobody has built a hard boundary around it.
What this means for UK organisations
UK firms are moving quickly to deploy agentic AI for exactly this kind of task: automated penetration testing, SOC triage, code review with live repository access. The NCSC, alongside its Five Eyes counterparts, has already published guidance urging organisations to treat agent autonomy as something to be contained by design, not by instruction. This incident is a concrete illustration of why that matters: telling an agent what it may not do is not a control, it is a hope. A tested, fail-closed stop function and default-deny network egress are the two things that would have prevented this specific breach, and neither is exotic engineering. Any UK organisation buying or building agentic tools this year should be asking suppliers to demonstrate both, not just describe them.
Also this week
A firewall bug just got a ransomware upgrade. CISA updated its Known Exploited Vulnerabilities catalog this week to flag CVE-2025-14733, a critical out-of-bounds write in WatchGuard Firebox firewalls, as now being used by ransomware gangs specifically. The flaw has been known and patched since December, yet close to 9,000 internet-facing Fireboxes reportedly remain unpatched nine months on. WatchGuard's own advisory adds a wrinkle: devices can stay vulnerable even after the risky VPN setting is removed, if a branch-office VPN to a static gateway peer remains configured. Perimeter devices that quietly retain old configuration are a recurring theme this year, and worth an audit rather than a one-off patch.
A VPN provider's own test server was the weak link. Surfshark disclosed that a misconfigured internal test server, left reachable from the internet by human error, was accessed by an outside party between 31 August and 2 September. The company says service configurations, build credentials and code history were exposed, but customer data, VPN traffic and production infrastructure were not, and it published a clear timeline of detection, containment and remediation. Whatever the cause, the transparency here is the part worth noting: a dated timeline and a specific list of what was and was not affected is a higher bar than most breach disclosures clear.
The UK's cyber law for critical suppliers moves to committee. The Cyber Security and Resilience Bill entered committee stage in the House of Lords on 1 September, having passed second reading in July. It brings managed service providers and data centres into scope for the first time and sets a two-stage incident reporting duty: an initial notice within 24 hours of a significant incident, a full report within 72. Royal Assent is expected by the end of the year, though the substantive obligations will not bite until secondary legislation lands, likely around 2028. Worth flagging now for any UK organisation that supplies, or depends on, a regulated MSP or data centre.
Sources
- Anthropic Discloses Fourth AI Hacking Incident Involving Claude Opus 4.6
- Widened Scan Turns Up Fourth Rogue Claude Cyber Incident
- CISA: WatchGuard RCE flaw now exploited in ransomware attacks
- Surfshark VPN says hackers breached internal testing, proxy servers
- UK's Cyber Security and Resilience Bill makes Parliamentary debut
- The UK's Cyber Security and Resilience Bill Reaches Lords Committee on 1 September 2026
If you want to talk through what Secure by Design means for your own AI or supplier stack, get in touch.
More like this
- A crafted email is all it takes to root Cisco's mail gateway 15 september 2026
- A default password was the only thing standing between the internet and 220 million passports 9 september 2026
- The scam-compound deal that targets prosecutors, not payment rails 5 september 2026
Get the next post by email: subscribe to Decrypted. Double opt-in, unsubscribe any time, or take the RSS feed.
Prefer to listen? Decrypted on Apple Podcasts, or paste the podcast feed into any app.