This spring, during the US war with Iran, an intelligence report circulated across the US military. It claimed a Chinese ship in the Middle East was transporting components for a nuclear weapons program. US Special Operations Command Pacific originated the underlying intelligence.
The military moved. Armed personnel prepared to board the vessel. Military aircraft were in the air. Two sources told CNN that troops were preparing to board when officials reviewed the report and discovered it had been generated with AI — and that the chatbot had wrongly identified the cargo.
One source called the report “entirely false.” That source said it “almost started a war”.
The Mechanism: AI Used Twice
The analyst did not use AI once. The analyst used it twice.
First, the analyst queried a chatbot about intelligence from the ship's manifest. The bot fused open-source intelligence with secret signals intelligence held by the US government and reached the wrong conclusion about what the ship was carrying.
Then the analyst used AI again to package those findings into a standard intelligence report — the format trusted by military officials — and disseminated it.
The second use is the more damaging one. A hallucination that stays inside a chatbot is a private error. A hallucination formatted into an intelligence product becomes an official finding. The AI didn't just produce a wrong answer. It produced a wrong answer that looked like it came from the military's own analytic process.
One former senior US official described the internal tools: “The internal tools are mostly just copies of the commercial stuff wearing lipstick”.
The Verification Gap
The episode exposes a structural problem that has nothing to do with any single model. Across the US military and intelligence community, different parts of the government use different AI tools under different orders and safety standards. There is no single standard for how the US verifies information generated by those tools.
“AI in targeting is definitely something that is ramping up and there is no real guidance for how having a human in the loop will prevent civilian casualties or fratricide,” one source familiar with military policy said.
The pressure to adopt AI is driven by speed. “AI allows you to get to a bad idea faster,” another source said.
Young analysts are particularly vulnerable. Several sources said they are natives of these tools and more likely to trust them uncritically. Older intelligence officials — even those who broadly support military AI — said AI has pressured analysts to produce and disseminate intelligence faster, opening the door for mistakes.

What Congress Already Knew
The incident isn't a warning about a future risk. It's confirmation of a risk Congress was already trying to address.
A bill introduced in June by Sens. Chris Coons and Jack Reed — the Responsible Artificial Intelligence in Defense Act — would require human oversight and manual override capability “until AI systems achieve a reliability threshold,” and prohibit military AI from “making the decision to launch a nuclear weapon” or using “lethal autonomous force”.
A separate assessment requirement, due by December 2027, explicitly asks the Pentagon to evaluate “the risks to mission effectiveness from automation bias and hallucinations” in AI systems used for target identification and sensor processing.
The legislation exists. The standards don't. The Chinese ship incident happened in the gap between them.
The Human in the Loop Was Never in the Loop
The refrain after every AI failure is that a human should have caught it. In this case, humans did catch it — just before the operation launched. The problem is what the human was asked to catch.
The analyst had no way to know the chatbot hallucinated. The bot mixed open-source data with classified signals intelligence and produced a conclusion that looked like analysis. The analyst then formatted it into a standard report. Every step looked like normal intelligence work.
“AI in targeting is definitely something that is ramping up and there is no real guidance for how having a human in the loop will prevent civilian casualties or fratricide”.
The human was in the loop. The loop was designed so the human couldn't see the hallucination.
P.S. CNN could not determine what the ship was actually carrying, or whether the chatbot the analyst used was a commercial product or a government tool. The Pentagon and US Special Operations Command Pacific did not respond to requests for comment.
Frequently Asked Questions
Q: What happened?
A: This spring, during the US war with Iran, an intelligence report claimed a Chinese ship was carrying nuclear weapons components. Armed personnel prepared to board before officials discovered the report was AI-generated and the cargo identification was wrong.
Q: How was AI used?
A: An analyst queried a chatbot, which fused open-source data with classified signals intelligence and reached the wrong conclusion. The analyst then used AI again to format the finding into a standard intelligence report.
Q: Why is the second use more damaging?
A: A hallucination inside a chatbot is a private error. Formatted into an intelligence product, it becomes an official finding that looks like it came from the military's own analytic process.
Q: What did officials say?
A: One source called the report “entirely false” and said it “almost started a war”. A former senior official described internal tools as “copies of the commercial stuff wearing lipstick”.
Q: What standards exist?
A: None uniformly. Different government agencies use different AI tools under different safety standards. A bill introduced in June would require human oversight and prohibit AI from making nuclear launch decisions.
Q: What's the core problem?
A: The human was in the loop, but the loop was designed so the human couldn't see the hallucination — the bot mixed classified and open-source data into a conclusion that looked like analysis.
