Every week another team wires an AI agent into something real: a refund queue, a CRM, a deploy pipeline, a spreadsheet that finance actually trusts. And every week the same quiet failure shows up. The agent replies "Done, updated 42 records" and nobody can tell, a day later, whether that sentence was true.
The chat transcript is not evidence. It is the agent describing its own work. If you would not accept "trust me" from a junior engineer touching production, you should not accept it from a model either.
The gap: claim vs state
An agent run produces two different things:
- A claim: the text it shows you ("I closed 3 tickets and emailed the customer").
- A state change: rows, files, API calls, messages that actually happened.
Most agent setups log the first one beautifully and the second one barely. When something goes wrong, you end up reading a friendly paragraph and guessing.
A 5-field receipt per action
You do not need a blockchain or a research lab for this. Emit one small record for every side effect the agent causes:
| Field | What it holds | Why it matters |
|---|---|---|
who |
agent id, model version, human who approved | accountability when the model or prompt changes |
what |
tool name plus normalized arguments | lets you replay or diff the exact call |
before |
hash or snapshot of the target state | proves what the agent started from |
after |
hash or snapshot after the call | proves the change really landed |
link |
hash of the previous receipt | makes silent deletion or reordering obvious |
That last field turns a plain log into an append-only chain. If someone (or something) edits receipt #17, receipts #18 onward stop matching. You get tamper evidence with a few lines of code and a SHA-256 call.
What this buys you
-
Debugging in minutes, not meetings. "The agent said it refunded the order" becomes a lookup: is there a receipt with
what = refund(order_id)and anafterstate showing the refund? -
Safer autonomy. You can let agents act on low-risk tools automatically and require a human signature in
whofor high-risk ones. The receipt shows which path each action took. - Honest metrics. Count actions with matching before/after states, not messages that contain the word "done". The gap between those two numbers is your real error rate.
- Privacy by design. Store hashes and field-level diffs instead of raw personal data. You can prove a change happened without copying the customer record into yet another log.
Start small
Pick one tool your agent calls in production. Wrap it so every call writes a receipt before returning. Add a tiny checker that walks the chain and flags breaks or missing after states. Run it nightly.
That is it. No new platform, no vendor. Just a habit: an agent action is not finished until there is a receipt a third party could verify.
The models will keep getting smarter. That does not make their self-reports more trustworthy, it just makes them more convincing. Receipts are how you keep the two apart.
What is the first tool in your stack you would wrap with a receipt? I am curious which side effects people worry about most.
I wrote a longer, free paper on this idea of verifiable claims for public and AI systems, if you want the deeper version: Proof, not promises.

Top comments (2)
The receipt is the part most teams skip. If the agent says done, someone still has to name the commit, the test that ran, and who would notice if it rolled back. "Done" without those three is just a sentence.
Agreed.
The chat transcript is just the agent describing its own work.
Your who, what, before, after table is very usable.
We write a hash-chained receipt like that on every refund and payout decision, before the money moves.
If a free shadow look at your recent refund queue would help, send us a DM.