Twelve false diagnoses: why my specs now start by checking every claim
Agents write confident diagnoses. Twelve of mine turned out to be false once someone opened the file. Checking once per session was not enough; it has to be claim by claim.
I keep product specs for Freem, the grocery assistant I am building, as PRDs: one document per feature, written so that a fresh session can implement the feature by reading only the PRD and its links. Most of these documents are drafted with an agent and then reviewed by me.
In August I counted the diagnoses in those documents that turned out to be false. Four were already documented by mid-August. Three days later, while building a catalog of open work, I found eight more.
They were not subtle. A spec said price history was missing; it already existed. Another said the privacy policy was not published; an endpoint was serving it. Each claim was written with total confidence, cited in later documents, and wrong.
Checking once is not enough
My first fix was obvious: verify the premises at the start of every session. It did not hold. A cleanup plan verified its premises when it started, and a premise added later in the same plan still turned out to be false.
So the rule became stricter. Every PRD now begins with a Phase 0, and it is applied claim by claim, not session by session. Before writing a single line about the current state of the system, every claim inherited from a ticket, a roadmap or an earlier plan goes into a table:
| Claim | Command run | Evidence | Verdict |
|---|---|---|---|
| "Price history is missing" | rg -n "price_history" src/ |
the file and line, quoted | False |
If a premise falls, the PRD changes scope or is not written at all.
Quote the line, or it did not happen
The "current state" section of a spec is where invention is easiest and where it gets cited most. So every path:line reference must come with the literal text of that line, copied from the file, not from another document. A reference that cannot be quoted is a guess.
The same goes for causes. If a root cause has not been measured, the spec has to call it a hypothesis, and the technical design has to start with a measurement step whose result can cancel the rest of the document.
Criteria that can fail
The other half of the problem is "done". Acceptance criteria now have to name the command that verifies them and the output that counts as failure. If I cannot write that command, the criterion is not falsifiable, so it gets rewritten or deleted.
Then comes a sabotage test: name what you would have to break for the checks to turn red. This one earned its place. A gap in how a database connection pool was set up had been declared closed. In reality it had only moved. Rebuilding the pool inline in the app's startup code left the whole test suite green, because nothing tested the startup path. An instrument that cannot fail is not evidence, and that applies most of all to the instruments you write to watch your own fixes.
Numbers follow the same rule. Every figure in a spec carries the date it was measured, the command and the environment. Before that rule, the repository's documentation held six different numbers for test coverage.
Why this matters more with agents
None of these rules are new. Engineers have always written confident wrong diagnoses. What changed is the volume. An agent can produce a well-structured, plausible spec in minutes, and the structure itself makes it look verified. The cost of writing a claim dropped to almost zero; the cost of checking it did not.
Phase 0 moves that cost back to where it belongs: to the moment before anyone builds on the claim. It is slower for the first hour and much faster for the rest of the week.