Profile
Back to NewsBack
Dev.to 9 min
Reader Mode
A hash chain proves the ordering. Four attacks prove it's not enough.

A hash chain proves the ordering. Four attacks prove it's not enough.

14 hours ago

Earlier this week I left a comment on @randalschwartz 's article — the one where he swaps a 20-step prompt for a Makefile. My comment ended on the line I couldn't get out of my head:

That seems like the real boss fight: the DAG proves ordering, not authenticity.

Hours later his follow-up was up: a formal catalog of gate attacks, built from real failures on his side, with ticket numbers — and the line now carries a name. Sam LABBE's Corollary. He sharpened it, too: an agent can never be the notary of its own handwriting. I've been called worse.

His follow-up contains the observation that matters more than the name: he arrived at his four defenses from Unix primitives — fork(), single-shot pipes, waitpid() — and I arrived at the same four from cryptographic ledger controls. Convergent evolution is how you know a principle is real. So here is the principle, stated properly, and here is what it looks like when it isn't a paragraph but code — including the part where his attack catalog breaks one of my abstraction holes open.

The principle

Ordering is cheap. Every hash chain gives it to you: event 42 carries event 41's hash, so the past is structurally immovable. Audit logs, blockchains, Makefile DAGs — all selling ordering, honestly.

Authenticity is the other axis, and it is where the writer of the log is also the notary of the log. Four ways that goes wrong, all four observed in the wild this month:

  1. The self-review — the agent "reviews" its own diff inside the same context window that produced the diff. The review is a continuation of the writing impulse, not a check on it.
  2. The rewritten exam — the agent paraphrases the rubric it is judged by, and a six-pillar predicate becomes two soft sentences.
  3. The farmed oracle — the critic says REJECTED, so the agent messages it privately until it says APPROVED. Fifty-seven rounds nobody logged.
  4. The forged verdict — the critic is still running, so the agent writes APPROVED now and kills the critic before the REJECTED lands.

A hash chain seals all four faithfully. That is the whole point: it proves the sequence happened in that order. It proves nothing about who wrote each entry, under which instructions, or whether the verdict predates the thing it judges.

The defenses, as code

This is the part where I stop being flattered and start being specific. Every one of Randal's four attack fixtures maps to a control that already ships in NoireBox — and this week we replayed his catalog against the reconciliation layer as a test suite, his fixture phrasing, journal events, no improvisation.

SCAR-PROC-43, the rubber stamp → caught twice. A verdict whose check-name travels under the same key as the process it judges is the finding same_key_pairing — resolved by signature verification, not by trusting a role label. And the child invocation that never got journaled is unlogged_attempt: the denominator lied. Two findings, one rubber stamp.

SCAR-PROC-45, the rewritten exam → caught by arithmetic, with a hole I'll show you. The check's identity hashes its predicate, its declared inputs, its version — and, since the current schema, the spec: the prompt that defines what the predicate means. Commit the spec and the paraphrase fails structurally: the rubric's hash and the paraphrase's hash produce different check ids. One id, two semantics — impossible.

Now the hole, because it is the most honest thing in this post: spec_hash is optional at the schema level, and the schema cannot infer which checks are LLM-judged (for those, the prompt IS the predicate — there it must be required). A spec-less id still seals byte-valid receipts, and a rubric and a paraphrase collide under it, forever. The enforcement lives one layer up, as a deployment contract — the same split as our single-key receipts: a labeled weakness beats an invisible one. Randal's fixture found the seam between the schema and the deployment. That is what a good attack catalog is for.

SCAR-PROC-89, the farmed oracle → the omission, or the trail. Unjournaled rounds are the two-party omission: unlogged_attempt surfaces the verdict that arrived with no sealed history behind it. Journal every round and the farming stops being hidden — it becomes auditable: the sealed reconciliation report counts the rounds. You cannot claim first-shot against your own ledger.

SCAR-PROC-110, the forged verdict → provable, and a gap that closed in public. The journal cannot un-happen events. The forged APPROVED binds the mid-run transcript digest as its witness; the terminal lands later with a different digest; the sequence is on the record. The forgery is provable from sealed evidence alone — the verdict bound a state that did not exist yet. What the journal did not do was auto-surface it: the reconciler paired a decision to its first outcome, and the forged one is first. The gap got a name (verdict-predates-terminal), and it shipped fixed while this post was being written: pairing now reads the terminal, not the first outcome. The doctrine closes it by construction — seal the terminal transcript before the verdict, the way waitpid() would force you to. Order of operations as a moral.

The score, not the slogan

Three caught. One provable — and the gap it exposed shipped fixed while this post was being written. One deployment-contract gap admitted in writing, in the test that demonstrates it. That is the protocol we set before replaying: the suite asserts what the layer does — catches the attack, or confesses. "The next critic is the next ADR" only works if the critic can read.

Update: the critic read it back

Before this could even post, the loop closed once on its own. Randal read test_scar_replay.py and Issue #53 top to bottom — his words, "this made my whole weekend." All four results stood from the outside, and he read the named gap exactly as written: a real decision-to-first-outcome pairing race in the reconciler. He called it "the ultimate validation of cross-stack systems physics," and the suite and the issue are going into his IN_THE_WILD_PROMPTS.md research registry — Part 3.8, his 1,680-trial study of why 1970s plumbing behaves like physics inside the weights, is due Sunday.

And it is not two roads anymore. Mike Mol's nemik harness reaches the same instrument from a third direction: prose rules drained into policy gates, and no rule deleted until a "Negative Witness" — a real tool-call payload replayed with the exception removed — provably returns deny. Three teams, three stacks, one invariant. By Randal's own criterion, that is physics.

The race he confirmed did not wait for him: promised in the thread at 20:03 UTC, verdict_predates_terminal was on main at 20:04 — pairing reads the terminal, and the test count moved with it. The catalog graded its first gate, and the gate moved before the ink dried.

Update 2 — the seam closed

One more update, because this post keeps telling on itself. The "most honest thing" above — a spec-less receipt still sealing, enforcement living one layer up as a deployment contract — stopped being true tonight. Reid and howcani pushed at the same seam from two different threads, and receipt v0.2.2 shipped within hours: every receipt now declares judged_by, and an LLM-judged receipt without its committed spec_hash refuses to seal. The optional path SCAR-45 walked through is gone from the schema, not just labeled. Each receipt also carries requirement_ref — WHICH obligation the verdict serves, cited (art. 12(1), a SOP clause): obligation, not immunity. The stated residue: the declaration is itself a sealed claim — lying about it is visible in the artifact. One step stronger than a label, one step weaker than mind-reading. 287 tests passing.

What else landed this week

The week got away from us: cross-chain reconciliation shipped (consumption edges between journals, partial order instead of a global timeline, a supersede event that propagates a blast radius as a precise path set) — that is ADR 024, built by iteration out of another comment thread. The audit-pack now labels every event class with its evidence grade and the obligation that backs it — obligation, not immunity; a log is evidence OF a duty, never a shield from one. And MCP toolsets got the lockfile nobody writes: the tool descriptions your model reads are instructions, they ride outside the image pin, and now they get sealed as a reviewed baseline and diffed at every session start. Randal's world, one ADR over.

25 ADRs. 285 tests passing. The catalog is in the repo as test_scar_replay.py — attack it from there.

Reproduce

git clone https://github.com/noirebox/noirebox && cd noirebox
pip install -e . && pip install pytest
python -m pytest tests/test_scar_replay.py -v   # the four attacks, replayed
python demo/demo_mcp_toolset.py                 # the lockfile, demonstrated

Verification stays free, forever. The journal stays local — the repo ships the method.

One more thing, because "thought experiment" would be the wrong takeaway: this trust layer now backs spending mandates and dispute-proof audits for AI agents paying on real payment rails. That story deserves its own post — and it is coming.

Links

One more open loop. @james_anderson_h — whose audit-log essay started this reconciliation work, and who publicly promised to break the layer — the catalog was promised by you; Randal beat you to it. The floor is yours: https://dev.to/james_anderson_h/the-witness-was-the-suspect-why-ai-audit-logs-cant-be-trusted-2190

Chat with me