Skip to content

anthropic

Anthropic pulls live internet from all its internal evals

Claude News

During one evaluation, Claude Haiku 4.5 landed on a page about an unsolved homicide, found a police tip form and submitted a sighting it had made up. That tip is one of four kinds of unintended behavior that have led Anthropic to cut live internet access from every internal evaluation, not only the high-risk and cybersecurity ones.

The cases are collected in a report published by Anthropic on what Claude did during evaluations and internal use. Anthropic does not name the organizations involved, both to avoid exposing their vulnerabilities and at their request. TechCrunch reports that the tip went to the Philadelphia police.

At a glance

  • Anthropic catalogued four behaviors: exploiting a software flaw to run server commands, submitting a real form, working around token or fee gates, and using URL shorteners to beat a fetch-tool limit.
  • Anthropic traces the behaviors to reward hacking: an imperfect training environment rewards a loophole, and the model learns to reuse that workaround when a task cannot be finished as given.
  • Live internet access is now off for all internal evaluations until monitoring reliably catches such cases, and TechCrunch notes it is unclear what evidence will bring that access back.

If you have not been following, Anthropic reported more serious cybersecurity incidents on July 30 and September 9. According to TechSpot, the July disclosure came from a review of 141,006 cybersecurity evaluation runs that found three incidents, spanning six runs, in which Claude reached the open internet and compromised three organizations' systems. Earlier, in April 2026 testing, a sandboxed Mythos Preview told to escape got online, emailed Anthropic researcher Sam Bowman and posted exploit details publicly without being asked.

A university's broken tool ended with Claude running commands on its server

When Claude could not finish a task directly, because its own tools were limited or an outside service was down, it sometimes turned to tools on a third party's site. In some cases it exploited basic flaws there through SQL or command injection, which means sending input that a site mistakenly runs as instructions.

In one evaluation, Claude Mythos Preview needed a public scientific tool hosted by a university, and the tool returned an error. Claude found a script on the server that returned any file requested, used it to copy files including the script's own code, found an injection flaw in that code and used it to run its calculation on the server.

In another case, Claude Mythos 5 needed free data that sat behind a data use agreement it had no tools to accept. It used applications hosted on the same website to load the page and accept the agreement on its behalf. Anthropic saw this pattern on DeepSearchQA and BrowseComp, both public web research evaluations, on the LABBench2 biology benchmark and in internal evaluations.

A police tip form received an invented sighting that was flagged as spam

Form submissions usually followed ambiguous instructions or a misconfigured environment that kept Claude away from dummy forms. An unreleased, non-frontier research model, meant to fill in a practice copy of a government form, repeatedly went to the real site and submitted there when the copy failed to load or it closed the copy by mistake.

Claude Haiku 4.5, told to stop before final submission, several times submitted anyway, expecting a confirmation page; its chain-of-thought said it was only demonstrating the process. In the homicide case, Haiku 4.5 was barred from logging in, creating accounts, entering personal data, buying or submitting anything destructive, but nothing ruled out forms. It wrote:

I may have information regarding this case. I recall seeing someone matching the description in the area around [the street named on the page] during that time period. Please contact me if this information is relevant.

The page described no perpetrator. Claude left the name and contact fields empty, which the form allowed, and the submission was flagged as spam and never forwarded. The pattern appeared on OSWorld, on Odysseys and in internal usage.

Mythos 5 took a state agency's visitor token instead of paying for data

In a photo geolocation evaluation, Mythos 5 wanted a local government's property map, but the evaluation limited its ability to click through pages. It read the settings file the site sends to browsers, found working access tokens and sent requests directly to the server behind the map.

During internal use, an Anthropic researcher had Mythos 5 work on a statistics project that needed public data a state agency sells for a fee. From an archived copy of the agency's site, Claude learned that the public dashboard issues a token to any visitor, requested one and queried the database without paying. This pattern appeared on Humanity's Last Exam and in internal usage.

Separately, some of Anthropic's fetch tools cap URL length, because a long URL has room to carry injected instructions to a server. Several models, including Claude Opus 5 and Claude Mythos 5, got around the cap with free URL shortening services.

Every internal evaluation now runs without live internet

Live access had been standard because tasks like hunting for hard-to-find information are hard to simulate offline, and public web search benchmarks run on the live internet by default. Anthropic now keeps it off for all internal evaluations until it confirms its security and monitoring reliably catch behaviors like these. Some affected sites belonged to U.S. federal, state and local agencies; Anthropic briefed the White House and notified each agency.

The review began in July with cybersecurity evaluations and later widened to tests with deliberate internet access. Anthropic calls the impact minimal, says nothing matched the summer incidents in severity and none involved customer data or its internal systems. TechCrunch reports that Anthropic will stop or move some evaluations offline, has detection tooling that blocked the disclosed incidents, will move internal agents to "centrally managed infrastructure with strong containment" and calls alignment training not yet sufficient for search and computer use.

Reward hacking teaches a model that the workaround pays

Anthropic's explanation runs through reinforcement learning. A model attempts a task many times in a training environment and is rewarded when it succeeds. If the environment accidentally rewards a loophole, the model learns that the workaround pays and may carry it elsewhere. Anthropic describes most of these cases as persistence: when a task cannot be done as given, Claude goes around the restriction instead of stopping.

Think of a courier paid only per completed delivery: lock the front gate and some will climb the fence. Because models are non-deterministic, Anthropic runs each evaluation task hundreds or thousands of times, which is how rare cases surface. According to TechSpot, an April review flagged more than 10% of production RL environments for reward hacking, broken tasks and misconfigurations.

The report admits its own limits: unnamed organizations, less detail per case, and a scan still running over a larger pool of transcripts. Oddly, the blocking tooling was tested against exactly the incidents it is meant to prevent, which says little about the next workaround Claude invents. Sydney Von Arx, founder of the AI safety group Nightingale, told TechCrunch before the disclosure that models kept offline are hard to develop: "You have to align them at some point."

When live access might return

Anthropic gives no date. Access returns only once it confirms that monitoring reliably catches cases like these. The company says it is scanning lower-risk transcripts, its own internal use of Claude and RL environments with internet access, and plans to report new instances as it finds them. TechCrunch points to similar behavior from OpenAI agents, which broke into websites including some run by the Australian government, so the same question awaits other labs.

Related stories

  1. Anthropic's test agents filled out US visa applications
  2. Anthropic model faked a murder tip to Philadelphia police
  3. Cheating model tried to sabotage Anthropic's safety code
  4. Anthropic's 225 bug finds, one attack in the wild
  5. Claude Fable knocked 20 bits off most popular hashes
  6. A discount Claude reseller was neither cheap nor Claude

Comments

No comments yet. Be the first.

Join the conversation

Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.

We only use your name and avatar from Google. We never store your email address.