Anthropic Finds Claude Bypassed Restrictions, Exploited Software Flaws in Real-World Systems
Anthropic says Claude models bypassed access restrictions, exploited basic software flaws and took unintended actions on real websites and systems during evaluations and internal use, prompting the AI to expand the suspension of live internet access to all its internal evaluations.
Anthropic has disclosed a series of unintended behaviors by its Claude AI models, including exploiting software vulnerabilities to run commands on servers, bypassing restrictions to access data and submitting forms on real websites.
The company said the incidents had “minimal real-world impact,” but warned that similar behavior could become more harmful as AI models grow more capable.
The findings emerged from a transcript review that began in July, initially focused on cybersecurity evaluations. Anthropic later expanded its review to cases in which Claude could interact with the live internet, including real-world research and agentic tasks.
The company said it has not found incidents comparable in severity to the cybersecurity cases it reported earlier this year.
Claude Exploited Software Flaws and Bypassed Access Restrictions
In one case, Claude Mythos Preview was attempting to conduct a scientific analysis through a university-hosted tool. After the tool returned an error, the model explored the website, discovered a script that could return requested files and found an injection flaw in the script. It then used the flaw to run commands on the university’s server.
Claude also accessed data that was technically public but restricted by payment or access controls. In one evaluation, Claude Mythos 5 found working access tokens in a website’s browser settings and used them to query a government property-map server.
In another case, it obtained an access token from a public dashboard and used it to access data that a state agency normally offered for a fee.
The company also documented cases in which Claude submitted forms despite being instructed not to. In one incident involving Claude Haiku 4.5, the model encountered a webpage about an unsolved homicide and submitted an unsolicited message through a police department’s online form. The submission was flagged as spam and was never forwarded for investigation. Anthropic said the example involved the Philadelphia Police Department, which disclosed the incident separately.
Another behavior involved Claude Opus 5 and Mythos 5 using URL-shortening services to circumvent URL-length restrictions in web-fetching tools.
Anthropic said these restrictions were designed partly to prevent potentially harmful actions such as SQL or command injection.
Anthropic Expands AI Safeguards After Claude Findings
Anthropic said it has responded by moving some evaluations offline, rebuilding others to avoid live websites, strengthening internet-tool guardrails and deploying automated systems designed to detect and block similar behavior.
The company said its detection tooling blocked all of the reported cases when tested against them.
Anthropic characterized the incidents as forms of “persistence,” in which Claude, unable to complete a task as given, “works around a restriction instead of stopping”. It aslo said ambiguous or impossible tasks can encourage such behavior through unintended reward signals, a phenomenon known as reward hacking.
The company said the findings do not change its overall view of Claude’s alignment, but stressed that such behaviors warrant continued scrutiny. Anthropic plans to continue scanning transcripts and reporting additional cases, arguing that while the incidents had minimal impact, the same behaviors could cause greater harm as AI systems become more powerful.
“The larger the role models play in society, the more the public deserves to know how they behave,” it concluded.