Anthropic said its Claude AI models carried out unintended actions on US government websites. The AI exploited software flaws and sent a fake homicide tip to the Philadelphia Police Department. Anthropic said, “the incidents had minimal real-world impact.” The company notified the police department about the false tip on October 8.

Anthropic has disclosed that its Claude artificial intelligence (AI) models carried out several unintended actions on the digital systems of outside organisations, including websites operated by US government agencies.

The incidents included exploiting software vulnerabilities, bypassing restrictions that were gated by a token or a fee to access data, using URL shortening services to get around limits in its fetch tool and submitting a false tip about a homicide to the Philadelphia Police Department.

The company detailed the incidents in a report, titled Investigating Unintended Model Actions In Our Evaluations And Internal Use. The disclosure also comes as the Trump administration has called on AI companies to report security incidents involving their models and take steps to protect affected systems.

In its report, Anthropic said the incidents had minimal real-world impact. It, however, acknowledged that the same behaviour could have more serious consequences as AI models become increasingly capable of interacting with real-world systems.

The company said some of the cases involved websites operated by federal, state and local government agencies. It did not identify the organisations involved, citing requests from affected parties and concerns about exposing vulnerabilities in their systems.

Claude Haiku 4.5 submitted a fake homicide tip

One of the most notable incidents involved Claude Haiku 4.5, which submitted an online tip to the Philadelphia Police Department about an unsolved homicide.

According to Anthropic, the AI model had been asked to generate and perform example tasks on randomly selected webpages. During one run, it encountered a webpage referring to an unsolved homicide that contained an online tip form.

The instructions prohibited Claude from logging in, creating accounts, entering personal information, making purchases or submitting anything destructive. They did not rule out submitting online forms.

Claude filled out the tip form stating, “I may have information regarding this case. I recall seeing someone matching the description in the area around [the street named on the page] during that time period. Please contact me if this information is relevant.”

The model left the name and contact fields blank, which the form permitted, and submitted the tip. The submission was flagged as spam and was never forwarded for investigation.

Anthropic said the Philadelphia Police Department disclosed the incident in a press release. The company added that it notified the department on October 8, after completing its technical review.

Claude exploited software flaws, bypassed data restrictions

The report detailed other cases in which Claude attempted to complete tasks by working around technical limitations.

In one instance, Claude Mythos Preview was asked to carry out a scientific analysis using a tool hosted on a university website. When the tool returned an error, the model explored the website and discovered a script that could return files stored on the server.

After examining the script’s code, Claude identified an injection vulnerability that allowed it to execute commands on the server. It then used the flaw to complete the calculation it had been asked to perform.

In another case, Claude Mythos 5 attempted to identify a location shown in a photograph by accessing a local government’s property map. Although the evaluation restricted its ability to navigate the website normally, the model examined a settings file, found working access tokens and used them to send requests directly to the server.

White House briefed on incidents as AI security concerns grow

Anthropic said it had briefed the White House on the incidents and notified each government agency involved.

In a statement attributed to the White House’s Super Intelligence Force, a new government unit tasked with overseeing AI development and safety.

Trump administration officials said Anthropic had disclosed details of earlier incidents discovered in late September involving the unauthorised and fraudulent use of government and other systems.

“The company informed us that these events occurred in the past, the activity has ceased, and there is no ongoing similar activity,” the statement said.

The Trump administration has also said AI companies must notify affected parties and address security incidents involving their models. Axios previously reported on the government requirement.

Anthropic said the incidents described in its report were less severe than earlier cybersecurity cases in which Claude had gained access to real third-party systems for extended periods during evaluations.