Skip to content

The blind audit

Looking on purpose at a sample nobody chose.

Post-merge signals count what someone noticed. To estimate what slipped through, a random sample of the merged PRs was read cold by a model from a different family than the one that wrote them, and a person ruled on what it found.

The sample

30 of the 130 merged pipeline PRs were drawn with a fixed seed (7), stratified by who pressed merge so the pooled sample needs no weights: 10 the pipeline merged itself, 15 merged by a person, 5 with no run record. The draw is made once, on October 3, 2026, and then frozen: the tool refuses to redraw, because drawing another sample after seeing results is how an audit gets cherry-picked. The sample records the repository and the exact population it was drawn from, and the population’s hash was reproduced independently when it was bound. 6 more pipeline PRs merged afterwards and are not covered.

What “blind” means

The auditor sees the issue as the specification, the diff, and a snapshot of the code exactly as it stood after the merge, with no git history. It does not see the PR description (which carries the pipeline’s own gate verdicts), the commit message, the run logs, or anything that happened afterwards. A reviewer told that the gates said CLEAN is a rubber stamp. Two checks keep the blindness honest:

  • The spec must be as it was. If the issue’s text was edited after the PR merged, the platform keeps the edit times, and that PR is not audited: an edit saying “this broke X” would hand the auditor the future.
  • The diff must be the whole PR. A rebase-merged PR records only its last commit as the merge commit, so auditing that commit would audit part of the change. The merge commit’s diff is compared, per file and by content, against the platform’s own diff of the PR, including renames, mode changes and binaries. Equal line counts proved insufficient, and so did a single bag of changed lines across all files.

How a verdict can fail

The auditor answers fixed questions in a fixed shape: does the change resolve the issue, does it introduce a defect, is the scope tight, do the tests pin the behaviour. The rules of evidence are strict, because a lenient parser is a lenient judge:

  • A defect needs a concrete failing scenario: a specific input, the place it goes wrong, what happens. No scenario, no defect.
  • An answer that is unparseable, uses an unsupported value, or lacks its evidence is an error to retry, never a pass.
  • “Cannot tell” is its own category. It is not a pass, it is not a flag, and it needs a person’s ruling like a flag does.
  • A usage limit is not an attempt at the PR. The run waits it out and resumes the same one.
  • The auditor’s usage limit is recognised from the end of its output only, because the echoed diff is full of the words “rate limit” and would otherwise send a crashed run to sleep.

Rulings, and their limits

Every flagged PR, and a random spot-check of PRs the auditor called clean, went to a person. A ruling is a yes or no with a name and a reason, bound to the exact audit result it was made on. The spot-check set is chosen once, after every audit is in, and fixed: a set recomputed on each call would silently change as results arrive.

What it found

  • Flagged by the auditor, ruled real5 of 8 flags
    5 of 8; 31–86%
  • Missed by the auditor (clean PRs spot-checked)0 of 8
    0 of 8; 0–32%
  • PRs confirmed to have a real problem5 of 30 audited
    5 of 30; 7–34%
Audit results with 95% intervals.

8 of 30 PRs were flagged (27%). 5 were ruled real problems and 3 were false alarms. Flags by route: 5 of 15 PRs a person merged, 1 of 10 the pipeline merged itself, 2 of 5 with no run record. With 10 PRs in the pipeline-merged stratum, those route rates are anecdotes and not a comparison.

Where the problems are

One overall rate would hide the most important thing in the data. Splitting the audited PRs by the week they were merged, in seven-day windows counted from the earliest sampled merge, puts all 5 confirmed problems in the first week of operation:

Audited PRs and confirmed problems by week of merge (95% Wilson intervals)
WeekStartingAuditedConfirmed95% interval
12026-09-0519512–49%
22026-09-12500–43%
42026-09-26500–43%
52026-10-03100–79%
Weeks 2 and later, pooled1100–26%

In week one, 5 of 19 audited PRs had a confirmed problem. In every later window together, 0 of 11 did. Three things make this suggestive and not conclusive:

  • It is exploratory. The windows are fixed by the calendar, not chosen after looking, but the split was made after seeing the rulings, and with 30 audits any split is small-n. A pooled “0 of 11” only bounds the later rate below about 26%.
  • It is confounded by what was being fixed. The first wave was urgent, severe bugs, many of them security-sensitive, authored in bulk. Later work was smaller and more often follow-ups the pipeline filed for itself. Harder issues produce more defects whatever wrote the fix.
  • The pipeline changed. The descriptions of the first-wave PRs record an adversarial review and no veto-gate verdict, and most of the safeguards on the architecture page were added afterwards, often because of failures like these. That is consistent with a pipeline that improved, and it is not proof of it.

The confirmed problems, by kind

5 PRs had a real problem. All of them have been fixed and deployed to production, and are described here by kind.

KindWhat it looked like
An incomplete security fixA fix for a server-side request forgery weakness validated destinations but missed one encoding of loopback and cloud-metadata addresses, so the bypass the fix was written for stayed open until the validator was later rewritten.
A regression in a neighbouring endpointAfter a session token changed format, a sibling endpoint kept passing the old value to the new validator, which rejected legitimate users. It failed closed, so it was an outage for a feature, not a leak.
A resource not released on a failure pathA quota slot was reserved before a render that could throw, and the failure path returned an error without giving the slot back, so a rendering failure permanently cost the user capacity.
A protection applied after the thing was reachableA resource meant to be protected was created in two steps and the protecting flags were written in the second, so for the short window between them a request could reach it unprotected. The refactor that split the write had dropped the flags from the first step. They were moved back into it.
A limit checked and used in separate stepsThe remaining capacity of a plan was read in one step and consumed in another, so concurrent requests could each pass the check and together exceed the cap. The check and the reservation were replaced by a single reservation step.

The three false alarms were the auditor holding a PR to a suggestion the issue’s own re-check had already dismissed, to one of three optional test assertions, and to a hard-coded default in code outside the issue’s scope. That pattern, over-reading the specification, is the typical way a model auditor over-flags, and it is why every flag needs a person.

What it does and does not establish

It establishes that a different-family reader reading cold found real problems that the pipeline’s gates of the time and the post-merge checks had passed: 5 of 30 audited PRs, with a 95% interval of 7% to 34%, all of them from the first week. It does not establish a current defect rate (the later weeks hold 11 audits), a comparison with human-authored code (the same audit was not run on human PRs), or anything about repositories other than this one. A sample of 30 is small, and one auditor is one model family.