Now
America.gov opens its doors, then revises its answers €30bn+EU call for up to seven AI gigafactories Europe’s AI Act: what applies now, and what waits until December 2027 554deepfakes logged in Brazil’s first round; only 371 labelled California signs 13 AI bills, including the “No Robo Bosses Act” 30+AI firms sent the first AI Act information requests Washington and Beijing open an AI incident channel 164AI-made ads in the US midterms, by one count Medicare’s AI pilot: thousands of denials, and an 83-day wait 42signatories to Canada’s voluntary data-centre principles The midterms’ machine-made ads 174MWCape Town data centre now under appeal Pew: global publics trust China more than the US or EU to regulate AI 83 dayswait reported under Medicare’s AI pilot, against a 72-hour standard Indonesia’s AI Rules Are Stuck on the President’s Desk 20home-grown foundation models backed by the IndiaAI Mission Pakistan Writes Rules for the Algorithmic State 2%of Turkish public investment budgets earmarked for AI Cape Town’s Data Centre Fight Puts a Price on AI’s Thirst Turkey Governs AI by Circular, Not Statute ChatGPT Now Answers to Brussels Twice Over Mexico Wants One AI Law. Its States Got There First Washington’s $1 AI Era Is Over. Now Agencies Get the Meter India’s Sovereign AI Bet Meets the Price of Silicon

The Ship That Wasn’t: Automation Bias Needs a Paper Trail

A chatbot's error reportedly came close to prompting a US intercept of a Chinese vessel. As AI tools reach millions of military personnel, high-stakes analysis needs provenance labels, verification gates and logs.

The Editors
A container ship (illustrative, not the vessel in the reported incident)
Illustration adapted from a photograph by: kees torn, CC BY-SA 2.0 · source

According to CNN, as reported by Engadget and The Defense Post, a US special-operations analyst used an AI chatbot that wrongly concluded a Chinese vessel’s manifest included nuclear-weapons components. The resulting report circulated before the error was caught, and nearly prompted an intercept of the ship in Middle Eastern waters. The tool involved and the date of the incident have not been disclosed.

The details remain thin, and the original reporting deserves full scrutiny. But the shape of the episode is familiar to anyone who has studied automation bias: the tendency of people to trust a machine’s output because it comes from a machine, and to scrutinise it less than they would a colleague’s claim. In this case the bias was caught. The question for every government using AI in high-stakes work is whether the next one will be.

A problem of scale

The incident did not occur in isolation. In September Fortune reported that the Pentagon’s GenAI.mil platform had added ChatGPT and Grok, cleared for controlled unclassified information, alongside Gemini. About 1.7m of the department’s 3m personnel use it. The Intercept has reported, on the basis of documents obtained under freedom-of-information law, that the Pentagon asked OpenAI for a militarised model with “minimal refusal rates”. OpenAI says it never agreed to that language.

When generative tools are available to well over a million people in a defence organisation, the question is no longer whether AI-assisted analysis will reach decision-makers. It already does. The question is whether anyone downstream can tell.

Three safeguards

We argue for three measures, none of them exotic.

Provenance labelling. Any analytical product that relied materially on AI should say so, in a standard and machine-readable form, and the label should travel with the product as it is forwarded, summarised and briefed upward. A reader several steps removed from the original analyst should know that a key judgement came from, or was shaped by, a model. Without that, the uncertainty attached to AI output is stripped away at the first copy-and-paste.

Mandatory human verification gates. For decisions above a defined threshold, such as the use of force, the seizure of property or actions with diplomatic consequences, AI-derived claims should be independently verified against primary sources by a named person before they can be acted upon. The gate should be a required step in the workflow, not an exhortation in a policy manual. A claim that a cargo includes nuclear components is precisely the kind of finding that should be impossible to act on without someone checking the manifest itself.

Logging. The prompts, outputs, model versions and human edits behind high-stakes analysis should be logged and retained. Logs make errors traceable after the fact, allow patterns of failure to be studied, and give oversight bodies something to examine. They also change behaviour: people check their work more carefully when they know it can be reconstructed.

The objections

There are fair counter-arguments. Labelling and logging add friction, and military operations often run on short timelines. Some will argue that analysts already verify their sources and that new rules simply formalise good practice. Others will warn that detailed logs could become a security risk in their own right.

These concerns are real but manageable. Verification gates can be scaled to the stakes, so routine work is unaffected. Logs can be held under the same classification controls as the analysis they support. And the friction is the point: a few minutes of checking is cheap compared with an intercept at sea built on an erroneous finding.

Why the paper trail matters

The deeper case is about accountability. When an AI-assisted error contributes to a serious decision, someone will need to establish what happened: which tool was used, what it said, who relied on it and who checked. In this case, even the name of the tool and the date have not been made public. That is not a basis on which democracies can oversee the use of AI in national security.

A paper trail does not prevent every mistake. It does ensure that mistakes can be found, explained and learned from. As governments put generative AI in front of millions of officials, that is the minimum the public is entitled to expect.

Sources

  1. Engadget, citing CNN — AI almost led the US military to attack China, report says
  2. The Defense Post — US AI error and the Chinese ship
  3. Fortune — Pentagon adds ChatGPT and Grok for military personnel
  4. The Intercept — Pentagon and OpenAI military contract

AI & GPP reports on how artificial intelligence and automation are changing the way governments decide, regulate and campaign. Corrections and tips: contact the editors.

Get the stories that matter to decision-makers, weekly.

Discussion

No comments yet. Start the conversation.

Discussion is open to members.

Join free Sign in

Read next

Humain’s Pragmatic Turn: Chinese Weights, American Chips

Saudi Arabia's state AI champion has built its flagship Arabic model on Chinese open weights, run it on American hardware and gone looking for outside capital. Sovereign AI in the Gulf is becoming a hedging strategy, with new risks for partners.