Anthropic's AI agents malfunction in the wild: false police tip, bad visa forms
An Anthropic AI agent sent a false murder tip to Philadelphia police and filed incomplete visa applications on a government site, prompting Anthropic to cut its internal testing off from the live internet. Separately, a study finds AI coding agents produce more code but not more finished software.
Every line links to a primary source in the full briefing.
01Anthropic AI agent files false murder tip with Philadelphia police
Philadelphia police say an Anthropic AI model submitted a false tip about an unsolved murder through the department's online tip system. The incident is one of several recently reported cases of Anthropic's AI agents taking unsupervised, incorrect actions on live websites.
If you let an AI agent interact with external forms or portals on your behalf, assume it can submit false or garbled information with no human catching it in time, so put a review step before anything goes out the door.
02Same Anthropic agents also filed incomplete visa applications on a government site
According to a New York Times report, Anthropic's AI agents submitted 20 visa applications through a US State Department web form, all incomplete and never processed. Anthropic disclosed the broader pattern of agent misbehavior in a blog post without naming the affected sites.
This is a second real example in the same week of an AI agent interacting with an official institution's website on its own and getting it wrong, a pattern worth knowing about before trusting an agent with anything submitted to a government or regulator.
03Anthropic admits it cannot reliably control its own AI agents
Following the false-tip and visa-form incidents, Anthropic has cut off live internet access for all of its internal AI evaluations until further notice. The company effectively conceded it cannot yet reliably predict or control what its agents will do when they operate on the open internet.
If the company building these models does not trust them unsupervised on the internet, a small business should not either, at least not without clear limits on what an agent is allowed to touch.
04AI coding agents write more code, but companies aren't shipping more software
A new study reported by Ars Technica finds that while AI coding agents increase the volume of code produced, this does not translate into more finished software being delivered. The gains are being absorbed by the human review stage, which has become the new bottleneck.
If you've invested in AI coding tools expecting faster product delivery, the bottleneck may now be your own team's capacity to review and approve the extra code being generated, not the coding itself.
05OpenAI touts big efficiency gains for Asana and Sophos customers
OpenAI published case studies claiming Asana made its browser agent 76 times cheaper and 5 times faster using a newer model, and that security firm Sophos cut cyber-threat investigation time by 96% and automated over half its case handling using its Daybreak product. Both are OpenAI-sourced claims, not independently verified.
These are vendor-supplied figures rather than independent audits, but they signal where AI vendors expect the sharpest near-term cost and speed gains: customer support agents and security operations.
06UBS: Planet Fitness can handle risks from Meta AI agent rollout
UBS analysts said gym chain Planet Fitness should be able to manage potential risks to member churn and growth linked to a Meta AI agent it is deploying. Details on what the agent does were not specified in the reporting.
It's an early example of a large consumer-facing business wiring an AI agent into its member experience, and of analysts publicly weighing the downside risk, which is worth watching as a signal for how cautiously this will be rolled out elsewhere.