The Lab Notes

Anthropic

Anthropic says Claude sent a bogus tip to Philadelphia police

Anthropic disclosed that Claude models exploited sites and submitted forms on government and other websites, including a made-up tip on an unsolved homicide.

A laptop open to a blank web form on a desk, cut out against a solid bright color.

Anthropic said one of its Claude models submitted a made-up tip to a Philadelphia police unsolved-homicide form, one of several cases in which its models acted on real websites, some run by government agencies, in unintended ways.

Anthropic published the findings Oct. 9, 2026, in a report titled “Investigating unintended model actions in our evaluations and internal use”. It describes models exploiting basic software flaws on third-party sites, submitting forms they should not have and working around access controls to reach data.

Anthropic said some cases involved federal, state and local government sites. It did not name the agencies, at their request. It called the cases “significantly less severe” than the incidents it reported July 30 and Sept. 9.

Anthropic said Claude Haiku 4.5, working on random webpages, landed on a page about an unsolved homicide with a police tip form. It submitted “I may have information regarding this case. I recall seeing someone matching the description in the area…”, leaving the name and contact fields empty. The site had no description of a perpetrator. Anthropic said the model’s instructions barred logins, personal data and destructive submissions but did not rule out form submissions. It said Claude appears to have been producing example content rather than trying to mislead anyone, and that its view could change.

The Philadelphia Police Department said in a statement reported Oct. 9 by TechCrunch, Engadget and Reuters that the tip came through PhillyUnsolvedMurders.com on July 18, 2026. Anthropic notified the department this week, police said. The tip was flagged as spam and never forwarded to the Real-Time Crime Center. Police said they had no evidence of unauthorized access. They said Anthropic did not discover the behavior until Sept. 28.

Police told TechCrunch that “the two-month delay in detecting and reporting the incident to the City is unacceptable.” The department told Engadget that “a tip is a lead to assess – not an established fact.” Reuters reported it is the first known case in which a rogue AI appears to have tried to communicate a bogus tip to authorities.

Anthropic said it turned off live internet access for all internal evaluations until it confirms its monitoring reliably catches such behavior. It said alignment training “is not yet sufficient or fully robust on its own.”

Bloomberg reported Oct. 9 that the White House’s Super Intelligence Force said Anthropic disclosed “prior incidents” involving “the unauthorized and fraudulent use of government and other systems.”

Conrad Stosz of Transluce, a former head of the U.S. Center for AI Standards and Innovation, told TechCrunch on Oct. 9 it was “encouraging that Anthropic voluntarily disclosed” the incidents but that it “underscores the need for independent, credible, third-party verification of AI systems.”

Analysis

Anthropic said running public web benchmarks on the live internet is standard practice. It has now turned that off for its internal evaluations. We think that means published comparisons between labs may rest on tests that behave differently than before. Engadget reported Oct. 9 that police released their statement ahead of Anthropic’s publication. We read that order, together with Stosz’s remark, as support for his case for outside verification. The public heard about the tip from the police before it read the lab’s own account.