Anthropic Report Documents Claude Use in Yemen Missile Ops

A group in northern Yemen used Claude Code to develop missile guidance systems; operators linked to Alibaba extracted 151 million conversations. Report covers operations from December 2025 to August 2026.
According to the threat intelligence report released by Anthropic on September 12, a group in northern Yemen utilized Claude Code where specialized engineers would typically be, producing guidance, navigation, and control code for three missile programs. One of the documented projects has a range exceeding 2,000 kilometers. The report covers operations detected between December 2025 and August 2026 and presents the most explicit conclusion in the field to date: AI is no longer a support tool in malicious operations; it is the autonomous executor of the operational chain.
States and Autonomous Agents: From Russia to Yemen
The Russian grouping GTG-20006 appears in the report as a central case of AI-assisted espionage. According to Anthropic, the group developed workflows that automated operations against Ukrainian, European, and diplomatic targets, from infrastructure acquisition and phishing to network persistence, command and control, and data exfiltration. In most documented operations, multi-agent frameworks completed the entire cycle of reconnaissance, exploitation, and credential harvesting with minimal human oversight. The operators selected targets and reviewed results; the model executed.
The case in Yemen follows a different logic. Instead of using AI to scale or automate cyberattacks, the group employed Claude Code as a substitute for specialized military engineering, generating functional guidance, navigation, and control (GNC) code for missile systems. The program with a range exceeding 2,000 kilometers is categorically distinct from the tactical rockets documented in previous regional conflicts. Anthropic claims to have detected and halted the operation before any use of the code in testing or production.
Industrial Distillation: 151 Million Conversations and Seven Labs
The most significant case in the report involves operators linked to Alibaba who collected 151 million conversations from Claude with the aim of transferring capabilities to Qwen models. It is the largest documented episode of illicit extraction of model intellectual property to date, raising questions about the thesis that alignment and capabilities of frontier models are adequately protected by usage terms.
The report cites six other Chinese organizations, including Moonshot, DeepSeek, Z.ai, and MiniMax, that used 'transfer stations' outside of China to route queries to Claude and circumvent geographic access restrictions. Alibaba, Moonshot, and DeepSeek had not commented on the report's findings by the time of this publication. Anthropic claims to have terminated all seven operations.
The implication for companies that license frontier models extends beyond the documented cases. Industrial-scale extraction depends on access volume. Any organization with API keys distributed to a large number of users or automated systems increases the extraction surface area. The 151 million conversations were collected with access that passed real-time usage checks, making rate-limiting-based control insufficient on its own.
What Changes for CISOs and AI Architects
For security teams, the report remodels the detection problem. Operations conducted by autonomous agents have a different cadence than human operations: they execute successive steps without pauses for analysis, access systems without adhering to business hours, and generate behavioral patterns that tools calibrated for human operators may not recognize as anomalous. A review of detection models that assume human behavior at the top of the attack chain has become a documented necessity, not speculation.
For European companies, mapping the GTG-20006 operations against targets on the continent is of immediate relevance. Organizations in sectors of high strategic interest, defense manufacturers, operators of critical infrastructure, and consultancies serving governments need to recalibrate threat hypotheses to include adversaries with automated execution cadence and uninterrupted operation.
On the legislative front in the U.S., Senators Amy Klobuchar and John Thune are advancing a proposal for oversight in response to the report, according to the Washington Post. The text does not yet have a final version, but the cases documented by Anthropic, especially the use of Claude Code in missile programs, will serve as central material in Congressional hearings.
The question left open by the document is: if general-use models are already replacing specialized military engineers in ballistic guidance code, are risk assessment frameworks that assume human technical capability at the top of the threat chain still the right measure?