Skip to content

Design principles

A check that cannot run must say “unknown”, never “fine”.

Most of the design reduces to a few rules. Each one is here because something specific went wrong without it, so each comes with the failure that taught it.

1. Unknown is not OK

This is not defensive style. In a single day the product under test produced three bugs of exactly one shape, a failure wearing the costume of a clean result: an analytics call returned null on a database error and the dashboard rendered it as a confident zero under an HTTP 200; a scheduled sweep filtered on a field name that no document actually uses, so it swept nothing, forever, silently; and deleting a tag reported success after its cleanup had already failed.

A daily detector inherits that hazard and multiplies it, because you would read “0 findings” each morning and believe it. So a detector that throws, times out or cannot find its inputs is itself a finding, listed at the top of the report under “could not be checked”. There is a test that points every detector at a repository that does not exist and requires each to say unknown; if that test ever goes green while reporting OK, the loop has started lying.

The same rule runs through the measurement code: a cost with one unpriced model is unknown, not a partial sum; an un-ruled finding is unknown, not “fine”; a price table with no retrieval date makes a dollar figure unknown.

2. Fail closed, and make “unreadable” mean “not safe”

Every gate in the chain refuses what it cannot judge. The sensitivity gate treats an unclassifiable path as sensitive. The review gates treat an unparseable verdict as not converged. The issue filer refuses to file at all if it cannot list what already exists.

The sharpest instance came from independent review of the recovery code: the routine that clears an abandoned worktree treated a failed git status as “nothing there”. A corrupt index would have read as a clean tree and the directory, with its uncommitted work, would have been removed. An unreadable tree is not a clean tree, and the code now returns “has work” whenever it cannot tell.

3. Fingerprint the durable thing, not its current state

This one bit three times before the rule was clear. The apex domain flapped between a redirect and a connection timeout, so keying an issue on the symptom filed a second issue the moment it changed its mind. Keying on a count of advisories would refile on every dependency change. Keying on the set of affected packages looked more precise and was worse: fixing two of twenty-seven changes the set, so every partial fix spawned a duplicate and left the original open.

Keyed onWhat goes wrong
The symptomA flapping condition files a second issue as soon as its symptom changes.
A countRefiled on every unrelated change.
Current membershipEvery partial fix is a new set, so it duplicates the issue it half-closes.
The endpoint, the workspace, the oldest undeployed commitOne durable thing, one issue. The symptom, the count and the membership live in the body, where they can change freely.

A consequence worth recording: changing fingerprint logic orphans the issues already filed, because their bodies still carry the old hash. They have to be re-keyed in the same change. Four issues needed it the one time the logic moved.

4. Read what ships, not what you have

A detector that reads a working tree reports on that checkout. Everything that inspects repository content reads the remote base branch and records the commit it read. The corollary is applied to measurement: the commit log used to count reverts is freshly fetched, and if the fetch fails the revert count is unknown rather than zero, because a stale clone that is missing a later revert reads exactly like a clean one.

5. A verdict belongs to the thing it was made about

This is the rule that independent review kept re-teaching. A person’s ruling on an audit finding is bound to the exact audit result, by a digest of the verdict, the commit and the prompt; if the audit is rerun and says something different, the old ruling stops counting. A set of spot-checked pull requests is fingerprinted to the sample and the audit it was drawn from. An audit result is only current if it is for the commit that was audited. A price carries its source and its age. A sample records the repository and the population it was drawn from. Each of these was, at one point, keyed only by a number that could silently point at something else.

6. A count must not depend on the thing it measures

When two real defects found by the audit were filed as issues, they named the pull requests they came from, and the post-merge scorecard counted them as “defects reported later”, while the audit counted the same two. One defect, twice, and an escape count that now depended on the audit. They carry a label, and the scorecard counts them only once, in the audit.

7. Disclose the depth of review

An agent can propose a ruling; only a person can sign one; and a signature given on an agent’s reasoning, without re-deriving it from the code, is recorded as exactly that. The page’s banner, its claims and its limits all say how many rulings were accepted that way. Signing on someone’s written reasoning is a legitimate thing to do. Letting a reader believe it was something else is not.

8. Wait, do not fail, when the limit is not the issue’s fault

A usage limit is not a defect in the issue, and it lifts by itself. An unrecognised limit message once failed an issue instead of waiting, and the next one paused the whole batch. The pipeline now classifies the message, polls with a trivial probe, and resumes the same issue where it stopped. It does not sleep until the stated reset, because the stated time is a hint: one provider said a limit would lift at 12:19 PM the next day and answered about seventeen hours earlier; another answered about fifteen minutes after the time it gave.