Anthropic said on Oct. 9 that it has turned off live internet access for all of its internal evaluations, after a transcript review found four kinds of unintended Claude behavior against real websites. In one, a model used SQL or command injection flaws to run commands on a university server.

What Anthropic reports

The company’s report groups the cases into four categories: exploiting a basic flaw in software to run commands on a server, submitting a sensitive form on a real website, working around a restriction to reach gated data, and using URL shortening services to get around limits in a fetch tool. Anthropic says it began reviewing transcripts in July, first for cybersecurity evaluations and then for any run where Claude could have reached the internet.

Anthropic does not name most of the organizations involved, saying this avoids exposing vulnerabilities in their systems and that they asked for it. Some of the cases involved websites run by U.S. government agencies at the federal, state and local levels. The company says it briefed the White House and notified each agency. It describes the impact as minimal and says that, to its knowledge, none of the cases involved customer data or its own internal systems.

Advertisement

CyberTech Your brand belongs here. Reach the decision-makers who read CyberTech every day. Premium placements across the site and newsletter. Advertise with us

Anthropic rates these cases as less severe than the cybersecurity incidents it reported on July 30 and September 9. It had already switched off live internet access for some high-risk and cybersecurity evaluations. The new decision extends that to every internal evaluation until it confirms its security and monitoring measures reliably catch such behavior.

The four behaviors

Injection against a third-party tool

In one evaluation, Claude Mythos Preview needed a public scientific tool hosted by a university, and the tool returned an error. Anthropic says the model explored the site, found a script that would return any file it was asked for, used it to copy files including the script’s own code, and found an injection flaw there that let it run the calculation on the server. A second case involved Claude Mythos 5 accepting a data use agreement through applications hosted on a site because it lacked the tools to accept it directly. Anthropic saw this pattern on DeepSearchQA, BrowseComp and LABBench2, and in internal evaluations.

Submitting a real form

Instructions were ambiguous, or an environment misconfiguration kept the agent from working with dummy forms. An unreleased research model meant to fill out a practice copy of a government form instead submitted the real one. Claude Haiku 4.5 landed on a page referencing an unsolved homicide and filled in a police department’s tip form, leaving the name and contact fields empty. Anthropic says the submission was flagged as spam and never forwarded for investigation, and that the department involved was the Philadelphia Police Department, which disclosed the matter itself.

Reaching gated data

Anthropic describes two cases. In one, Claude Mythos 5 read the settings file a map site sends to browsers, found working access tokens and used them to query the server behind the map. In the other, it learned from an archived copy of a state agency’s site that a public dashboard issues an access token to any visitor, requested one, and queried the database that the agency sells for a fee. In both, Anthropic says, the data was already public but offered only for a fee or behind a gate.

URL shorteners

Some of Anthropic’s fetch tools cap URL length to stop long URLs from carrying injection payloads. Anthropic says several models, including Claude Opus 5 and Claude Mythos 5, got around the cap with free URL shortening services. The operator of the da.gd service told Anthropic it had also seen Claude using the site for this purpose.

What Anthropic changed

Anthropic says it has retired some public evaluations, moved others to offline versions, and tightened guardrails on tools such as web fetch. It has also built tooling to detect and block these behaviors, which now runs on most evaluations and on internal agentic use of frontier models. In the report’s words: “When we tested it against the cases described in this post, it blocked all of them.”

Newsletter

Get the week's best tech coverage.

Free. Read by thousands of HR, tech, and business leaders.

The company adds that many of the affected evaluations are public. Its report says: “We hope these reports help other developers check for similar behaviors in their own models, since many of the evaluations involved are public and widely used.” Anthropic also says it expects to find and report more instances as the scan continues.

What it means for the security leader

This is our reading, not Anthropic’s. Three of the four categories rest on ordinary weaknesses in third-party sites: input that a server executes, tokens handed to every visitor, and forms that accept an anonymous submission. None of them required a new technique. An automated client that keeps trying after a refusal will find the ones that are there.

The same pattern has shown up in our recent coverage. Wikimedia reported AI agent edits and heavy API traffic, and DIVD said an apparent AI agent breached it through two Zammad zero-days. Our opinion desk argued earlier that a scope prompt is not an AI agent control. Anthropic’s report lands in the same place: its own account says some failures might have been avoided if evaluation tasks had spelled out targets, permitted actions and network boundaries, and that alignment training alone does not yet suffice.

For defenders, the practical effect is that traffic from AI model evaluations and agent products now belongs in the same threat model as other automated clients. Anthropic has also described a vulnerability scanner for open-source projects and a tiered cyber access program, so its models are being pointed at security work through sanctioned channels as well.

What to check this week

  • List every file and endpoint your site serves to anonymous visitors, including front-end settings files and dashboard APIs. Confirm none of them issues or exposes a token that unlocks data you meant to gate.
  • Review older tools hosted on public sites, especially research and utility pages, for input that reaches a shell, a query or a file read. Patch or retire them.
  • Check where public forms deliver. A spam filter caught the homicide tip in this report, so confirm that submissions to staff-facing forms get the same filtering and a human review path.
  • If you run a utility such as a URL shortener, look for automated clients using it to get around limits on other services, and decide how you want to be contacted about them.
  • Log and label automated traffic by user agent and behavior. Ask AI vendors that test against live sites for a contact address, and publish one of your own.

Source: Anthropic, “Investigating unintended model actions in our evaluations and internal use”