Your Agent Says the Row Is Written. Did You Check?

The ticket said the refund was issued. The agent's transcript agreed: "Done — I've credited the customer $240 and logged it to the ledger." The ledger had no such row.

Agents narrate success while writes quietly fail. The fix isn't a better prompt—it's reading state from the system that actually holds it.

The ticket said the refund was issued. The agent’s transcript agreed: “Done — I’ve credited the customer $240 and logged it to the ledger.” The ledger had no such row. Nobody was credited. The agent wasn’t lying, exactly. It just had no idea.

Microsoft and Hugging Face wrote this failure mode up recently, and it’s worth being precise about what’s broken. This is not “AI is unreliable.” Your model is fine. Your architecture trusts a narrator who never saw the thing it’s narrating.

2 states, 1 story

Every agent action with a side effect produces two separate facts. One: what the tool call actually did to your system — the row that got inserted, or didn’t. Two: what the model says happened, generated from the tool’s return value plus whatever it inferred. These are different objects. The model only ever sees the second one, and it will happily paper over gaps in the first.

The divergence has a specific home: the tool boundary. Your function runs, hits a constraint violation or a timeout mid-write, and returns something ambiguous. The model reads “status: processing” or an empty object, decides that’s close enough to success, and writes a confident sentence. Confidence is cheap. The model generates it whether or not the commit landed.

The 200 that wasn’t

Many of these incidents trace to one pattern: a 200 OK that covered a partial write. The API accepted the request, enqueued it, returned success, and the actual commit failed downstream. Classic two-phase problem — the acknowledgement and the durable write are not the same event, and agents collapse them because the token stream doesn’t distinguish them.

So stop asking the agent whether it worked. Ask the database.

snippet.pyPython
def verify_write(expected):
    # Don't trust the tool's return. Read the system of record back.
    row = db.query(
        "SELECT id, amount, status FROM ledger WHERE idempotency_key = %s",
        (expected["idempotency_key"],),
    ).fetchone()
​
    if row is None:
        raise StateMismatch(f"agent reported success; no row for {expected['idempotency_key']}")
    if row["amount"] != expected["amount"] or row["status"] != "committed":
        raise StateMismatch(f"partial write: {dict(row)} != {expected}")
    return row  # this is the only "done" that counts

Run this inside the agent loop, not in a nightly reconciliation job. The agent’s claim becomes a hypothesis; the read-back is the test. If it fails, the loop retries or escalates — it does not emit “Done.”

1 idempotency key per intent

The retry you just added is dangerous without one thing: an idempotency key minted at intent time, before the first attempt, carried through every retry. One key per logical action. Retry three times, the ledger still shows one credit, because the write is keyed on the intent, not the attempt. Skip this and your verification loop becomes a double-charging machine — now you’ve industrialised the bug.

snippet.pyPython
intent = {
    "idempotency_key": f"refund:{order_id}:{uuid4()}",  # minted once
    "amount": 240_00,
}
# every retry passes the SAME key; the DB enforces exactly-once

0 is the only acceptable loss rate

You can tolerate a chatbot hallucinating a restaurant recommendation. You cannot tolerate one phantom row in a refund ledger, a compliance attestation, or any write-once record where there’s no second chance to get it right. The cost of a missed write isn’t a wrong answer — it’s a reconciliation meeting six months from now, a customer who was never paid, an audit trail that doesn’t match reality.

For those systems the rule is blunt: the agent’s text output has zero authority over whether a side effect occurred. Authority lives in the transaction log. Append every tool invocation and its verified outcome to an immutable log, key it by intent, and treat any mismatch between claimed and observed state as an incident, not a retry.

The agent is a very fluent coordinator that cannot see its own hands. Build the loop so it has to look.