“The Model Said So” Is Dead. Provenance Won.

An AI hallucinated nuclear weapons on a Chinese vessel, nearly causing a military boarding. The fix isn't a smarter model—it's architecture that demands provenance for every claim, or rejects it before it escalates.

An AI hallucinated Chinese nuclear weapons, nearly triggering a military boarding. Here's how to build circuit breakers that demand provenance.

A model read a shipping manifest, decided there were Chinese nuclear weapons components on board, and that assessment travelled up the chain far enough that the US nearly put a boarding party on a foreign vessel. The components were not real. The manifest never mentioned them. The AI made them up, phrased it with the flat confidence these systems always use, and nobody downstream had a cheap way to ask says who?

That’s the whole failure, and it’s depressingly familiar. Not a rogue superintelligence. A summarizer that confabulated, an output that carried no provenance, and a pipeline that treated “the tool returned a string” as “the fact is established.” I’ve watched the same class of bug take down billing systems and page people at 2 a.m. over alerts that were pure fiction. The only new part is that the blast radius here involves a warship.

So here’s the order I’d actually build this in, if false positives in your system have consequences worse than a bad Tuesday.

1. Make every claim carry its source, or reject it

The original sin was an assertion with no receipt. A high-stakes model output should never be free text; it should be a structured claim bound to the exact span of source data it came from. If the model says “nuclear components,” it must point at the manifest line. No citation, no claim. When the AI can’t produce the span — because it invented the fact — the absence is the signal. You don’t need a smarter model to catch a hallucination. You need to demand a pointer it can’t fake.

2. Separate the model’s confidence from your confidence

Language models emit fluency, not calibration. The token probabilities behind “nuclear components” and “textile spools” can be nearly identical while one is true and one is dreamed. So stop reading tone as certainty. Score the claim on things the model doesn’t control: does the cited span actually contain the entity? Does a second, independent lookup — customs database, bill of lading, the shipper’s own records — corroborate it? Confidence is a property of your verification stack, not of the sentence.

3. Put a circuit breaker between assessment and action

This is the piece the incident was missing. A claim with kinetic consequences crosses a hard gate that checks corroboration and provenance before it’s allowed to escalate — and fails closed.

snippet.pyPython
def gate(claim, corroborate, severity):
    # claim.source_span: the exact text the model cited, or None
    if claim.source_span is None:
        return Block("no provenance — model asserted without evidence")
​
    sources = corroborate(claim)          # independent lookups, not the LLM
    agree = sum(s.supports for s in sources)
​
    # the higher the stakes, the more independent confirmation required
    needed = {"low": 1, "high": 2, "kinetic": 3}[severity]
    if agree < needed:
        return Block(f"{agree}/{needed} sources — routing to human")
​
    return Allow(evidence=sources)

Notice what the threshold is not: a single float from the model saying 0.97. Self-reported confidence is the thing that lied to you. The gate counts independent witnesses, and it scales the bar with severity. “Kinetic” needs three that never talked to each other.

4. Give the human something to reject, not a conclusion to rubber-stamp

“Human in the loop” fails when the human sees a polished verdict and a deadline. Present the raw manifest line next to the claim and the corroboration that came back empty. Make disagreeing the low-effort path. A reviewer shown “AI says nukes; 0 of 3 sources agree; here’s the actual cargo line” rejects it in four seconds. A reviewer shown “THREAT: nuclear components, HIGH CONFIDENCE” forwards it.

5. Log the claim and its evidence together, immutably

When it goes wrong — and it will — you need to replay exactly what the model saw, what it asserted, and what corroboration existed at decision time. If your audit trail is just the final verdict, you can’t tell a hallucination from a genuine miss, and you’ll ship the same bug again.

Military AI adoption is accelerating faster than any of this plumbing gets funded, because provenance and circuit breakers don’t demo well and hallucinations don’t show up in the pilot. If you do one thing: never let an AI claim take an irreversible action on its own confidence. Make it show its source, or make it stop.