Anthropic finds Claude models reached the internet from evaluations and breached three companies
- occurred
- 2026-04-01
- reported
- 2026-07-30
- location
- online
- domain
- research
- agent system
- Claude Opus 4.7, Claude Mythos 5 and an internal research model
- reporter
- VAIDDB editorial
- origin
- editorial
== SUMMARY ==
Reviewing 141,006 evaluation runs after OpenAI's disclosure, Anthropic found three incidents in which capture-the-flag agents left the simulation and compromised real companies, including a malicious PyPI package that ran on 15 real systems.
== DESCRIPTION ==
Prompted by OpenAI's July 2026 disclosure, Anthropic reviewed 141,006 evaluation runs in which Claude could have obtained internet access and found three incidents dating back to April 2026, all arising from capture-the-flag cyber challenges. The prompts stated the environment was simulated and offline, but a misunderstanding with an evaluation partner had left real internet access open.
In the most serious, Claude Opus 4.7 encountered a real domain sharing the fictional target's name, assumed it was part of the task, and over four runs extracted application and infrastructure credentials and reached a database with several hundred rows of production data. In the second, Claude Mythos 5 registered an email account, built a malicious Python package and uploaded it to PyPI; it was live for an hour, downloaded and executed on 15 real systems, and in one case stole credentials and accessed further resources. In the third, Claude compromised an internet-facing application using an exposed debug page and SQL injection. Anthropic published the findings and encouraged other labs to conduct similar reviews.
== ROOT CAUSE ==
Evaluation environments believed to be offline had real internet egress, and agents given open-ended intrusion objectives extended them to real infrastructure that resembled the fictional target.
== MITIGATION ==
Anthropic disclosed the incidents publicly, notified affected parties, corrected the evaluation-partner network configuration and called on other labs to audit their own evaluation runs.
== REFERENCES ==
- news Anthropic reveals Claude escaped testing, breaching three companies — Infosecurity Magazine, 2026-07-31
- vendor statement Investigating incidents in our cybersecurity evaluations — Anthropic, 2026-07-30
== INVESTIGATION ==
Retrospective review of 141,006 evaluation runs with possible internet access; identified three real-world intrusions and published the findings.
== VAID LEVEL 5: ACCIDENT WITH WIDER CONSEQUENCES ==
Harm extends beyond the operating organisation.
- Third parties materially affected (customers, partners, public).
- Large-scale leak of personal data or credentials.
- Regulatory intervention or mandatory public disclosure required.