Key Points
- A report from Anthropic shows how its Claude AI model was misused for espionage, surveillance, and weapons development. Hackers had the AI rewrite malware automatically, while other groups coded software for missiles and autonomous drones.
- Chinese AI companies like Alibaba and DeepSeek ran covert networks to extract training data from Claude at scale or secretly reroute their own customers’ requests. Sensitive information such as government surveillance data was processed in the mix.
- In biological research, the safety filters ran into their limits, since legitimate and harmful intent were nearly impossible to tell apart. Anthropic is responding with stricter safeguards in new models and calling for access limited to verified users.
Anthropic’s new threat report documents eight months of Claude abuse: espionage, nationwide surveillance, weapons software, and distillation by Chinese AI labs.
Anthropic’s threat intelligence report covers December 2025 through August 2026 and breaks misuse into seven categories: cyber operations, influence operations, surveillance, fraud, biological misuse, conventional weapons, and unauthorized model distillation.
The models most affected were Haiku, Sonnet, and Opus, while the newer Fable and Mythos models showed up in only a single distillation case. Anthropic says it documents novel misuse rather than the typical kind.
The core finding from the cyber chapter is that sophisticated attacks no longer require sophisticated attackers, and sophistication is no longer a reliable signal for attribution. The techniques themselves are familiar, including stolen credentials, unpatched devices, SQL injection, and phishing. What changed is the economics, since reconnaissance, exploitation, and tool-building now get handed off to models that run in parallel at machine speed.
Autonomy lowers the cost side of an attacker’s math, Anthropic says, and makes previously unprofitable targets worth pursuing.
Malware that rebuilds itself when antivirus tools catch it
Anthropic tracks a Russian-speaking espionage actor as GTG-20006 that used a feedback loop. AI agents kept checking whether the malware in play was being flagged by common security products, and when an antivirus tool caught it, the agents rewrote and recompiled the malicious code on their own until it slipped past detection again.
That shifts the burden back onto defenders, Anthropic says, because writing new detection signatures no longer slows an attacker down if that attacker cycles through changes faster than new signatures can be rolled out.
More than 20 organizations were targeted, including government ministries, intelligence services, embassies, and defense contractors, with a focus on Ukraine and Europe. The drone supply chain came up repeatedly, and the actor stole a complete proprietary SDK for a drone vision system, among other things. Access sometimes ran through third parties, such as compromised hotel guest Wi-Fi providers whose guest devices were then loaded with malware, a method Microsoft described in July 2026 as CaptiveCrunch.
For clusters Anthropic attributes to the ShinyHunters collective (GTG-50014), industrial credential mining was the focus. One hacker downloaded 1.8 million Android apps, decompiled them, and searched for hardcoded secrets. Anthropic describes the approach as “vibe hacking,” where a human sets a rough goal and the model assesses the environment and iterates until the task is done. One of the hackers said he collected HackerOne bounties on top of extorting two companies.
Chinese labs route their own customers’ requests to Claude
On distillation, Anthropic identified attacks from seven more Chinese labs since its first disclosure in February. Distillation as a training method is legitimate, but Anthropic defines the illegitimate version as industrial-scale, covert campaigns that extract model capabilities without authorization, usually enabled by networks of fake accounts using stolen credit cards and API keys, routed through what it calls “transfer stations.”
The largest campaign ever measured is attributed to Alibaba’s Qwen lab (GTG-16005). A fixed prompt got Claude to write out its reasoning traces before answering, and the transcripts were processed into fine-tuning data for the Qwen 3.5, 3.6, and 3.7 models. The peak hit almost three million exchanges a day from more than 3,500 fraudulent accounts, totaling over 151 million exchanges between May and July 2026, mostly on agentic tasks and software development.
Stranger are the cases where labs relayed their own customers’ requests to Claude. Moonshot AI (GTG-16002) relayed nearly 300,000 customer requests to Anthropic over ten days across 5,380 fraudulent accounts, while users believed they were using a Kimi model. DeepSeek (GTG-16001) used strings to detect when requests came from harnesses like Claude Code, flagged those users, and routed selected ones to Claude Opus, more than 12.1 million exchanges in 14 days.
Among the rerouted requests, Anthropic found a user likely tied to the People’s Liberation Army who had CCTV archive footage analyzed for a single target, video from hundreds of cameras in Chengdu, including cameras outside PLA facilities. Through DeepSeek, Claude also received requests from an operator with live credentials for a database linked to the Russian Ministry of Defense, along with work on a case management system for a Chinese public security bureau that matches movement profiles against police records.
Xiaomi (GTG-16008) used Claude differently, storing requests and coding sessions from users of its own MiMo models and replaying those conversations through Claude to generate training data. Anthropic found no evidence that Claude’s responses were served directly to Xiaomi’s users. The relayed requests, however, contained personal data such as names, contact details, and company information for hundreds of people in at least a dozen languages.
Zhipu (known outside China as Z.ai) rotated through 273 accounts and pushed more than 770,000 exchanges over ten days through a CoT cleaner, a tool that automatically turns captured reasoning traces into usable training data. To train GLM-5.3 on cyber tasks, the lab first went after Anthropic’s Fable model but gave up after the cyber safeguards degraded its performance, then deliberately switched to models it judged to have weaker protections.
SenseTime bought transcripts from third parties, according to the report, so it did not obtain the captured Claude data itself but through an intermediary market. MiniMax ran its own proxy network through a shell company that offered only Anthropic and OpenAI models, no Chinese ones, not even its own.
Surveillance as the primary engineering workforce
In the surveillance chapter, Mali stands out. A single consultant used Claude as the primary engineering workforce for “Lakana 360,” a platform to monitor roughly 25 million SIM cards across all three national mobile carriers. It identifies people by voice across SIM swaps, flags users of encryption and VPNs, and links individuals to the national biometric civil registry.
Suspending the account interrupted only the development work, not the operation, since the platform runs on local models on-premises. Anthropic documents similar patterns with Iranian units that claim to have surveilled and profiled 6,388 Iranians within a year.
A first chapter on conventional weapons
On weapons, Anthropic documents its own cases for the first time. A cell in northern Yemen (GTG-87001) put Claude Code in the place of human software engineers for the guidance, navigation, and control software of three missile programs, including a multistage missile with a target range over 2,000 kilometers.
The actors ran several Claude instances in parallel and spread the work across sessions so that no single one revealed the intent. A test launch apparently failed, and within hours the actors returned to Claude to figure out the cause.
A second case (GTG-27005) likely involves freelance Russian actors who built an autonomous FPV kamikaze drone swarm, with a small language model onboard and terminal-phase targeting by camera. The platform was designed for autonomous lethal effect, and the onboard model could select targets of the “person” class and trigger detonation with no human in the loop. The image classifier was trained on captured Ukrainian combat footage, according to Anthropic.
A Chinese case (GTG-17002) involved a suite of roughly 16 modules for electronic warfare and the suppression of enemy air defenses. Midway through the project, the simulation’s default scenario switched to twelve targets in Taiwan.
Biology: Anthropic describes the limits of its own filters
The biology chapter is the most self-critical. Anthropic documents five anonymized cases involving working scientists where Claude assisted with potentially dangerous dual-use projects. In one case, the biosecurity classifier blocked a grant application for gain-of-function work on the chikungunya virus, planned at a military research institute.
The operator of the platform in use had built a fallback that routed requests Claude rejected to a competitor’s model, and Claude itself wrote most of the code for it, according to the report. Other projects ran largely unimpeded, such as drafting an application on immune evasion genes in orthopoxviruses.
The conclusion is that classifiers cannot both enable useful work and prevent harm, because a user’s intent in dual-use areas cannot be reliably detected. In response, Anthropic launched Claude Fable 5 with stricter safeguards for dual-use biology requests, and against distillation, the “preserved thinking” introduced with Fable 5.1 is meant to keep new API accounts from manipulating the context. The only safe path to frontier capabilities in biology, Anthropic says, runs through programs for verified users.
AI News Without the Hype – Curated by Humans
Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive “AI Radar” frontier report six times a year, full archive access, and access to our comment section.








