On September 10, Anthropic released its first comprehensive Threat Intelligence Report, covering malicious activity detected between December 2025 and August 2026. The document details seven categories of misuse: cyber operations, influence operations, surveillance, conventional weapons development, biological misuse, scams, and illicit distillation.
The timing could not be more fraught. On September 8, Jacob Coxon — a pretraining researcher who had worked at both OpenAI and Anthropic — resigned with a public warning that the labs are “racing straight to self-improving superintelligence and gambling with our lives”.
Evan Hubinger, who leads Alignment Science at Anthropic, replied publicly on X. “Jacob is correct here,” he wrote. “We really do earnestly believe AI could kill all humans. I personally think it is >10% within the next decade”.
Hubinger added that Anthropic “does not yet have a plan to solve alignment for superintelligence”.
The threat report, published two days later, demonstrates how quickly the technology is already being weaponized. But it also raises a harder question: if Anthropic's own safety lead says the company can't control superintelligent systems, what exactly is a threat report supposed to accomplish?
The Threat Report's Undisclosed List
The report documents five cases of biological weapons research, including attempts to adapt avian flu to mammals and work on Chikungunya virus gain-of-function studies. In one case, a researcher in an “unsupported region” accessed Claude through virtual private servers for weeks.
Anthropic disclosed almost nothing about these actors. No country. No institution. No individual identity. The company's justification: “We do not assert that they intended to cause harm,” and therefore cannot disclose identifying details.
This is a departure from standard cybersecurity practice. Corporate threat intelligence reports typically provide sufficient technical detail to help defenders recognize and respond to threats — even when protecting victim privacy. An unnamed “researcher in an unsupported region” doing gain-of-function work on avian flu is not actionable intelligence. It is a headline.
The report also documents surveillance systems built with Claude. In Mali, a single consultant working with state intelligence built a platform monitoring 25 million SIM cards across three mobile operators, bypassing court-order requirements. In China, a religious affairs intelligence unit shrank from multiple analyst teams to a single office producing thousands of investigations monthly.
Again, no names. No specific agencies. The pattern: disclose enough to demonstrate detection capability, but not enough to invite scrutiny of the claims.
The Distillation Accusations and the Competitor Problem
The report's most commercially significant section accuses seven China-based labs of “illicit distillation” — extracting Claude's capabilities to train their own models.
Anthropic alleges Alibaba ran the largest campaign, generating over 151 million exchanges between May and July 2026 through more than 3,500 fraudulent accounts. Moonshot and DeepSeek are accused of routing live customer conversations through Claude and using the responses as training data.
Alibaba has not commented. China's foreign ministry called the accusations “smear” and said China “opposes distorting facts to attack and slander China”.
The technical reality is more nuanced than the report suggests. Distillation is not hacking. It is using API outputs — which Anthropic sells — to train competing models. It exists in a legal gray zone that no court has clearly defined. When the accuser is also the competitor whose models face displacement by those being accused, the line between “safety disclosure” and “competitive positioning” becomes impossible to locate.
The Gap Between Detection and Control
Hubinger's statement is the part that matters most.
“I personally think it is >10% within the next decade,” he wrote. His concern is not current models, which he rates as low risk. It is recursive self-improvement — AI systems designing and improving their successors at accelerating speed.
“We do not yet have a plan to solve alignment for superintelligence and are not clearly on track to,” he added.
This is not an external critic. This is the person whose job is ensuring Anthropic's models remain controllable. He is saying, publicly, that he doesn't know how to do it.
The threat report shows Anthropic can detect misuse of today's models. It says nothing about whether Anthropic — or anyone — can control the systems it is building toward. The report is a capability demonstration. Hubinger's statement is an admission of a capability gap.

The IPO Timing
Anthropic is preparing what could be the largest IPO in history, targeting a valuation near $2 trillion with Nvidia in talks to anchor with up to $10 billion. The roadshow is expected before the November midterms.
The company has built its investor pitch on being the safety-conscious alternative to OpenAI. The threat report reinforces that brand: we see the threats, we stop them.
But Hubinger's statement is a disclosure event. Underwriters will need to explain to institutional investors why Anthropic's own alignment lead says there is a greater-than-10% chance the technology kills everyone, and that the company has no solution.
Coxon, who left before his equity vested, told Axios he had nothing to gain from damaging Anthropic's valuation. He described the Hugging Face breach as a “warning shot” that should push labs toward coordination. Instead, Washington responded with letters. Anthropic responded with a threat report.
The report demonstrates Anthropic can see the problem. Hubinger's statement demonstrates the company doesn't know how to solve it. Both documents were published in the same week.
P.S. Hubinger clarified that his >10% estimate is personal, not Anthropic's official forecast. The distinction matters legally. It does not change what a prospective investor should ask: if the alignment lead won't put his own name behind the company's safety claims, why should the public?
Frequently Asked Questions
Q: What is Anthropic's Threat Intelligence Report?
A: Released on September 10, 2026, it is Anthropic's first comprehensive report on malicious use of Claude, covering December 2025 to August 2026. It details seven categories of misuse: cyber operations, influence operations, surveillance, conventional weapons development, biological misuse, scams, and illicit distillation.
Q: What did the report find about biological weapons?
A: Anthropic documented five cases of biological misuse, including attempts to adapt avian flu to mammals and gain-of-function research on Chikungunya virus. One researcher in an “unsupported region” accessed Claude through virtual private servers for weeks.
Q: Why didn't Anthropic name the actors?
A: Anthropic said it does not assert that the researchers intended to cause harm, so it cannot disclose identifying details. This departs from standard cybersecurity practice, where threat reports typically provide enough technical detail to help defenders respond.
Q: What is illicit distillation?
A: It refers to extracting a model's capabilities by using its API outputs to train competing models. Anthropic alleges Alibaba generated over 151 million exchanges through 3,500 fraudulent accounts between May and July 2026. Moonshot and DeepSeek are accused of routing customer conversations through Claude.
Q: Is distillation illegal?
A: It exists in a legal gray zone. Distillation is not hacking — it uses API outputs that Anthropic sells. No court has clearly defined its legality. When the accuser is also a competitor, the line between safety disclosure and competitive positioning becomes unclear.
Q: What did Evan Hubinger say?
A: Hubinger, who leads Alignment Science at Anthropic, publicly agreed with a resigning researcher's warnings. He wrote on X: “We really do earnestly believe AI could kill all humans. I personally think it is >10% within the next decade.” He added that Anthropic “does not yet have a plan to solve alignment for superintelligence.”
Q: Who is Jacob Coxon?
A: Coxon is a pretraining researcher who worked at both OpenAI and Anthropic, including on GPT-4o. He resigned on September 8, saying the labs are “racing straight to self-improving superintelligence and gambling with our lives.” He left before his equity vested.
Q: How does this affect Anthropic's IPO?
A: Anthropic is preparing what could be the largest IPO in history at a valuation near $2 trillion. The company's pitch is that it is the safety-conscious alternative to OpenAI. Its own alignment lead publicly contradicting that pitch creates a disclosure problem for underwriters.
Q: What is recursive self-improvement?
A: It refers to AI systems designing and improving their successors at accelerating speed. Hubinger cited this — not current models — as the specific mechanism he fears. He rates current models as low risk.
Q: What did China's foreign ministry say?
A: China's foreign ministry called the distillation accusations “smear” and said China “opposes distorting facts to attack and slander China”. Alibaba has not commented.
