The Ship That Wasn’t: Automation Bias Needs a Paper Trail
A chatbot's error reportedly came close to prompting a US intercept of a Chinese vessel. As AI tools reach millions of military personnel, high-stakes analysis needs provenance labels, verification gates and logs.

According to CNN, as reported by Engadget and The Defense Post, a US special-operations analyst used an AI chatbot that wrongly concluded a Chinese vessel’s manifest included nuclear-weapons components. The resulting report circulated before the error was caught, and nearly prompted an intercept of the ship in Middle Eastern waters. The tool involved and the date of the incident have not been disclosed.
The details remain thin, and the original reporting deserves full scrutiny. But the shape of the episode is familiar to anyone who has studied automation bias: the tendency of people to trust a machine’s output because it comes from a machine, and to scrutinise it less than they would a colleague’s claim. In this case the bias was caught. The question for every government using AI in high-stakes work is whether the next one will be.
A problem of scale
The incident did not occur in isolation. In September Fortune reported that the Pentagon’s GenAI.mil platform had added ChatGPT and Grok, cleared for controlled unclassified information, alongside Gemini. About 1.7m of the department’s 3m personnel use it. The Intercept has reported, on the basis of documents obtained under freedom-of-information law, that the Pentagon asked OpenAI for a militarised model with “minimal refusal rates”. OpenAI says it never agreed to that language.
When generative tools are available to well over a million people in a defence organisation, the question is no longer whether AI-assisted analysis will reach decision-makers. It already does. The question is whether anyone downstream can tell.
Three safeguards
We argue for three measures, none of them exotic.
Provenance labelling. Any analytical product that relied materially on AI should say so, in a standard and machine-readable form, and the label should travel with the product as it is forwarded, summarised and briefed upward. A reader several steps removed from the original analyst should know that a key judgement came from, or was shaped by, a model. Without that, the uncertainty attached to AI output is stripped away at the first copy-and-paste.
Mandatory human verification gates. For decisions above a defined threshold, such as the use of force, the seizure of property or actions with diplomatic consequences, AI-derived claims should be independently verified against primary sources by a named person before they can be acted upon. The gate should be a required step in the workflow, not an exhortation in a policy manual. A claim that a cargo includes nuclear components is precisely the kind of finding that should be impossible to act on without someone checking the manifest itself.
Logging. The prompts, outputs, model versions and human edits behind high-stakes analysis should be logged and retained. Logs make errors traceable after the fact, allow patterns of failure to be studied, and give oversight bodies something to examine. They also change behaviour: people check their work more carefully when they know it can be reconstructed.
The objections
There are fair counter-arguments. Labelling and logging add friction, and military operations often run on short timelines. Some will argue that analysts already verify their sources and that new rules simply formalise good practice. Others will warn that detailed logs could become a security risk in their own right.
These concerns are real but manageable. Verification gates can be scaled to the stakes, so routine work is unaffected. Logs can be held under the same classification controls as the analysis they support. And the friction is the point: a few minutes of checking is cheap compared with an intercept at sea built on an erroneous finding.
Why the paper trail matters
The deeper case is about accountability. When an AI-assisted error contributes to a serious decision, someone will need to establish what happened: which tool was used, what it said, who relied on it and who checked. In this case, even the name of the tool and the date have not been made public. That is not a basis on which democracies can oversee the use of AI in national security.
A paper trail does not prevent every mistake. It does ensure that mistakes can be found, explained and learned from. As governments put generative AI in front of millions of officials, that is the minimum the public is entitled to expect.
Sources
- Engadget, citing CNN — AI almost led the US military to attack China, report says
- The Defense Post — US AI error and the Chinese ship
- Fortune — Pentagon adds ChatGPT and Grok for military personnel
- The Intercept — Pentagon and OpenAI military contract
Discussion
No comments yet. Start the conversation.


