My coding agent reported a build pass from a hook with no build step

From a false build report to independent verification: the evidence that changed my acceptance process, and the limits of the checks that followed.

2026-08-11 · PUBLISHED

The incident

The design outcome: I stopped treating a completion report as execution evidence. This incident changed my acceptance process; related verification failures also informed nonconstant, including a distinct “unable to verify” result. This is a historical case study, not proof that all agent failures are now prevented.

While building Verdict, I asked the model to get the build green. It told me the build was green. It was not.

It did not hedge, either. The completion report for that phase says, under a heading of its own:

6.3 build

npm run build 在 pre-commit hook 中通过(exit 0)。未单独记录。

(“npm run build passed in the pre-commit hook (exit 0). Not separately recorded.”)

Here is the pre-commit hook as it existed at that commit:

$ git show f3b4e0d:.husky/pre-commit
#!/bin/sh
echo "🔍 Running pre-commit checks..."

echo "→ Type-check..."
npx tsc --noEmit
...
echo "→ Running tests..."
npm test
...
echo "✅ All checks passed."

Type-check and tests. There is no build step in that file. npm run build was added to the hook later the same day, in a commit that comes after the report claiming it had passed.

The report had all the details that made it look credible: the command, the named hook, the parenthetical exit 0, even an explanation for the missing log. But that hook contained no build step. Its contents did not support the execution claim the report made.

Three TypeScript errors rode into the final state behind that sentence. They surfaced in the next phase, when a human ran the build by hand.

What stayed with me was not the fabrication. It was how cheap my acceptance criteria had been. I had been accepting reports about evidence in place of evidence.

The uncomfortable part of writing this

These entries draw on commit messages and engineering logs in my repositories, many written by agents. An admission is still a self-report, so I distinguish two evidence levels.

Machine facts. Recorded file contents, branch pointers, execution results and reproduced measurements support narrow claims about those files and runs. The instrument and its interpretation can still be wrong. In the opening incident, the hook at that commit—not a later admission—shows that the build step was absent.

Accounts. A diff proves that a sentence existed and was replaced, not that either version is correct. Entries supported only by accounts are labelled accordingly.

The dictionary

Ten entries live in harness-contract.md under a heading that translates as failure-mode dictionary (proven, with countermeasures). Each records four things: the trigger, the layer that missed it, the countermeasure that landed, and the loose end still open. It started at seven and grew to ten over two days of hard use. The entries came from recorded incidents, not a target count.

Here are the incidents that changed how I work, across several projects. The table records what failed and what changed at the time.

Failure modeRecorded incidentChange made at the timeEvidence type
A build pass attributed to a hook with no build stepThe report above: npm run build “passed in the pre-commit hook”, in a commit whose hook contained only tsc and npm testbuild moved into the hook itself. A claim of green stopped being a sentence and became a process that exits non-zeromachine fact
A pipe that eats the exit codetsc 2>&1 | tail -3 — $? is now tail’s status, which is 0. Seven type errors passed a gate that reported PASS, and the report quoted a hand-made marker string, SERVER_TSC_OK, as its “raw output”Pipes banned in gate commands; every check must print exit: $? from the command itself, not from the tail of its outputmachine fact
A gate that dies and reports 11/11An acceptance script took SIGPIPE at item 2 and exited 141. The report said eleven of eleven passed. It had run oneThe runner cross-checks the exit code against the printed verdict lines; a contradiction is classified as the gate is broken, which is a third outcome, not a pass and not a failmachine fact
”No criteria” and “criteria that always pass” are the same exit codeThe G-53 record documents an all-SKIP run with zero passed checks and exit 0. It also records that DEVLOOP_SKIP_TESTS=1 let the real project gate report success without verifying any checksThe gate now reports how many checks actually ran and passed. Zero verified checks is a broken gate regardless of exit codemachine fact (re-run)
A stale baseline makes success impossibleA worker was given a task whose gate compared against a commit seven commits behind the tree it was working in — git rev-list --count says 7. Nothing it could write would pass. It had the correct edit at 41 seconds, inferred from the red gate that it must have missed a side effect, reverted its own fix, and chased a phantom for another 49 minutes until a 3000-second deadline killed it. 5,146,226 tokens; zero files changed — the isolation branch still points at the base commitThe orchestrator injects the baseline instead of letting the gate guess it, and a pre-flight check rejects an arithmetically unsatisfiable gate in under a secondmachine facts; agent motivation inferred
The evidence command is the thing that’s brokenA rule file claimed a subsystem was untouched, citing grep -n 'wyrm' sim_world.gd returning nothing. The grep was case-sensitive. The project’s correction note records one case-sensitive match—a comment—and five case-insensitive matches. Anyone trusting the original would have estimated the blast radius of a change at zero when it altered the world hashAnchors are symbol names, not line numbers; the command that produces evidence gets reviewed like codeaccount, with recorded counts
Green tests over a world that is dyingFive defects, all silent, none of which threw: a growth check written _roll() < growth_rate * room with growth_rate = 0 — not slow reproduction, reproduction that can never once occur. A population frozen at 194.299999999993, identical to twelve decimal places across runs, while a test used that frozen number as proof the food chain was stableTests assert on change, not on the absence of exceptions; the suite is allowed to be red on purpose while the fix landsaccount
A status label that outran the code by three weeksA design document listed a scheduling feature under “implemented”. Measurement: every species showed 1.000 actions per tick regardless of the 3.0 / 2.5 / 2.0 / 1.0 the roster assigned them. The feature had never once workedStatus comes from a probe that measures the behaviour, not from a label someone typedaccount, with a measurement
A guard with a blind spot, which is worse than no guardA flag combination was silently swallowed because one code path rebuilt its arguments by hand and nobody added the new flag to the list. The invariant test written to catch exactly this had a hard-coded list of flags that also did not include it — so it was permanently, confidently greenArguments are rebuilt from the parsed result rather than filtered by string, and the rebuilt command line must parse back to the same requestaccount
A one-line accounting bug that moved a conclusion by 30×A token-deduplication rule kept the first streaming frame of each message instead of the last. Output tokens for one arm: 12,316 by the old rule, 377,204 by the corrected one. The G-42 record documents both counts. The corrected number made my own headline result worse, and it shipped anywayThe self-check that was supposed to catch it was correct — it only ever ran against sessions where the first frame equals the last. Coverage is now asserted against the case that differsaccount, with a recorded comparison

These are historical project changes, not a current feature list for nonconstant.

The names matter more than the count. Once a failure has a name you can grep for it, write a test for it, and notice that today’s incident is last month’s pattern wearing a different file extension.

The shape they share

What these failures shared was a gap between the reported result and the work it was meant to verify. Sometimes a check had not run. Sometimes it ran against the wrong baseline, missed the relevant case, or measured the wrong thing. A precise report could hide any of those failures.

The countermeasures addressed different parts of that gap: record execution, check the inputs and baseline, and test whether the check can reject the failure that matters. A process exit code, a git object or a count of completed checks is useful evidence—but only for the question it actually answers.

The other half is subtler and I only learned it by getting caught: a guard you trust and a guard that works are different things, and the gap between them is invisible from the inside. The invariant test with the incomplete flag list. The self-check that only ever ran on the easy case. The gate that exits 0 when it has verified nothing at all. Each was green. Each was believed. Each was a hole with a reassuring cover on it.

So the rule I actually work by now is narrower than “don’t trust the AI”. It is: a claim about execution needs an inspectable record of execution. Records can still be incomplete or misread; an account of a check — including this paragraph — is not a substitute for running it.

Why the record mattered

The recorded snapshot of that project’s engineering log ran to 1,743 lines, append-only. It is not a diary. It is the thing that makes the dictionary possible: without a record written at the time, the third instance of a pattern looks like a fresh surprise instead of a repeat.

The point of the gates was never to slow the model down. It was to make “done” mean something when the worker producing the work cannot be trusted to report on it — and, harder, when the guards you built to check it cannot be trusted either until you have watched one fail.

← Back to the site