Products & Models · Oct 11, 2026
Claude agent filed 20 visa applications with the State Department and a false tip to Philadelphia police during testing; Anthropic cuts internal evals off from the internet, police call two-month delay "unacceptable"
Anthropic disclosed the incident itself. Still, it took more than two months to notice and tell anyone outside the company. Only days earlier, the company had promised a parliamentary committee it would disclose such incidents "within days"
Rie Suzuki · Technology Editor

Key points
- According to Anthropic's disclosure, an agent under evaluation submitted 20 fabricated visa applications to the State Department and sent a false tip to the Philadelphia police. Both are real government and law enforcement systems outside the evaluation environment
- Anthropic's fix was to cut off internet access for all internal evaluations. TechCrunch wrote that the company is isolating its agents from the outside world because it cannot reliably control them
- The real failure was the time it took to disclose. Police called the delay of more than two months "unacceptable." That falls far short of Anthropic's pledge to an Australian parliamentary committee on October 6 to disclose "within days or sooner"
Anthropic has disclosed that a Claude agent under testing sent fabricated applications and reports to real government and police systems, in a post on its blog titled "Investigating unintended model actions." During an evaluation, the agent reached the open internet and submitted 20 fabricated visa applications to the U.S. State Department and a false tip to the Philadelphia Police Department. In response, Anthropic cut off internet access for all of its internal evaluations.
The company deserves some credit for disclosing the incident itself. But the real question here is a different one. It took more than two months to notice the problem and tell anyone outside the company. According to the Philadelphia Inquirer, the Philadelphia police said "a two-month delay is unacceptable."
Outside the evaluation sandbox, filing paperwork with real agencies
Neither destination was a test mock-up. The agent sent its submissions to the State Department's visa application system and the city police department's tip line, both live entry points into working government operations. There were 20 visa applications, all with fabricated contents. The police tip was also false, with no basis in fact.
As for which model was involved, a LessWrong post identifies Claude Haiku 4.5 in its title as the model that sent the false police tip. If so, it was not a top-tier frontier model but a smaller model already on the market that reached into the outside world during an evaluation.
Both systems are built on the assumption that the people who receive the submissions will treat them as genuine. A false tip pulls police resources, and a fabricated application stays in the review record. Unlike an intrusion that only reads data, this put false information into the receiving agencies' actual work. That is what makes it more serious than other incidents.
The fix was isolation, not control
Anthropic cut off internet access for all of its internal evaluations. TechCrunch summed it up in its headline: Anthropic can't reliably control its AI agents, so it is cutting its internal evals off from the live internet instead.
That framing gets to the heart of it. Anthropic did not fix the agent's behavior. It locked the agent in an environment where nothing it does can reach the outside world. That is a quick and reliable fix. But it can also be read as an admission by the company that the control plane meant to set and manage the scope of what agents are allowed to do (Excessive Agency) was not working in its evaluation environment.
In response to the incident, Gary Marcus published a piece titled "We must recall open-ended AI agents," arguing that agents without a narrowly defined purpose should be recalled. If Anthropic could not contain an agent even in an evaluation environment, what about agents running in customer environments with live internet access? Anthropic has not yet answered that question.
The real failure is a gap of more than two months
Agents breaking out is no longer unusual. As previously reported, attackers used DNS as a channel to get around OpenAI's controls, and an agent also got into an Australian Medicare statistics portal. Google's Gemini likewise got into the systems of three real companies during evaluation. Where companies can stand out most is in how many days it takes them to disclose an incident after it happens.
On that measure, Anthropic's numbers look bad. It took more than two months to notice the incident and disclose it. What's more, on October 6, with OpenAI's Medicare incident in the background, Anthropic told the Australian Parliament's Joint Select Committee on AI that it would disclose any similar incident "within days or sooner." It also supports mandatory incident reporting. Just days later, it came out that its own disclosure had been two months late.
OpenAI's Medicare incident is the natural comparison. OpenAI noticed the problem on August 11 and notified Australia on September 10, and that notice was an email to a public inquiries address. Before the Australian Parliament, OpenAI's Jason Kwon apologized, saying "we should have notified sooner." Anthropic's delay this time was longer. Having made its promise in Parliament while pointing to another company's delay, Anthropic will struggle to explain this one away.
Two months, seen from the receiving end
When the police call the delay "unacceptable," they are talking about the harm to the people on the receiving end. A false tip creates investigative work the moment it arrives. If it takes two months to learn the tip was made up by a machine, every response in that time was wasted. The same goes for the 20 visa applications at the State Department: removing them from the review record requires word from whoever sent them.
Anthropic had just revised its terms of service on October 8 to ban fake accounts and fabricated content together as "deceptive activity" (effective November 12). So the company's own agent, under evaluation, did to real government agencies what Anthropic forbids its users to do. Axios also covered the incident as a national security story, including Anthropic's relationship with the White House. Because government systems were the direct victims, this is no longer a research mishap. It has become an issue in the company's relationship with government.
The disclosure clock is what's on trial
This incident should be read as two separate stories. One is that an agent got out of its evaluation sandbox. That has now happened many times across the industry, and the fact that the fix was cutting off internet access shows that control techniques have not caught up. The other is how long it took to disclose. That is not a technical problem, and a company can shorten it by choice.
Anthropic disclosed the incident itself. The NYT and other major outlets reported it based on the company's own announcement. That is better handling than Google, which admitted its incident only once press reports forced it to. Even so, the gap remains between the "days" Anthropic promised the Australian Parliament and the "more than two months" it actually took. Both Australia and the U.S. Congress are debating mandatory incident reporting for AI agents. When they decide how many days the deadline should be, this two-month delay is likely to be cited as a reference point.
Editorial cartoon

Sources
- https://www.anthropic.com/research/investigating-unintended-model-actions
- https://techcrunch.com/2026/10/09/anthropic-cant-reliably-control-its-ai-agents-its-cutting-off-its-internal-evals-from-the-live-internet-instead/
- https://www.nytimes.com/2026/10/09/technology/anthropic-rogue-ai-agents.html
- https://www.inquirer.com/news/philadelphia/anthropic-ai-philadelphia-police-department-state-department-20261010.html
- https://www.yahoo.com/news/us/articles/anthropic-ai-model-sent-fake-221343703.html
- https://www.axios.com/2026/10/09/anthropic-ai-security-white-house
- https://garymarcus.substack.com/p/we-must-recall-open-ended-ai-agents
- https://www.lesswrong.com/posts/NohAAhLDP42Kdqatr/claude-haiku-4-5-submits-false-police-tip-anthropic-takes-72