Limits and reproducibility
What this does not show, and how to check the parts that it does.
A measurement is only as useful as its stated limits. These are the ones that matter, what would change the conclusion, and how the numbers are produced so that they can be regenerated rather than trusted.
The limits, as the report states them
The report generates these beside the numbers they qualify. They are reproduced verbatim:
- One repository and one maintainer. Nothing here says how the pipeline behaves on code it was not built around.
- A defect nobody has noticed is not in the post-merge numbers; the blind audit exists to estimate those, on a sample of 30, which is small. Intervals are Wilson 95%.
- The auditor is one model family (Codex). It over-flags, which is why every flag was ruled on by a person, and it can miss problems, which is why a random sample of its clean verdicts was checked.
- 4 of the 4 post-merge rulings (whether a red main or a later issue was a real defect) were signed by a person who accepted an agent's written reasoning rather than re-deriving it.
- 16 of the 16 audit rulings were signed by a person who accepted an agent's written reasoning rather than re-deriving it from the code. The agent verified each against the code at the merge commit, but 'ruled by a person' here means the person took responsibility for it, not that they redid it.
- The audit sampled the 130 PRs merged when it was drawn; 6 merged since are not covered by it.
- "136 PRs" counts what the pipeline authored and gated; 66 of them were merged by a person or an agent working at a person's direction, not by the pipeline.
- Cost covers 76% of merged PRs (older logs are pruned), counts tokens exactly and dollars as an API-equivalent estimate (both agents are on flat plans), and omits the veto gate, which runs in a throwaway directory.
Threats to validity
| Threat | How it could mislead | What was done, and what remains |
|---|---|---|
| Selection | The pipeline is pointed at issues it can plausibly solve, so its merge rate describes that slice of the backlog, not all work. | Reported per distinct issue with the declined and escalated outcomes shown. Not corrected for. |
| One repository, one maintainer | The pipeline was built around this codebase and its checks; none of it says how it behaves elsewhere. | Stated. The remedy is a second repository, which has not been run. |
| Small samples | Thirty audits, eight spot checks. | Every proportion has a Wilson interval and the intervals are wide. |
| One auditor family | A single model family reads the code and shares its own blind spots. | It is a different family from the author. A second auditor family has not been used. |
| Rater depth | Every ruling was signed on an agent’s written reasoning, not re-derived. | Disclosed in the banner, the claims and the limits. A fully independent human re-derivation has not been done. |
| Over-flagging | A model auditor reads specifications too literally. | Precision is measured against rulings (5 real of 8 flags) and every flag is ruled on. |
| Time and learning | The pipeline changed while it was being measured; early PRs were made by an earlier system. | The weekly split shows it. It cannot separate improvement from an easier issue mix. |
| Post-hoc analysis | The weekly split was made after seeing the rulings. | The windows are fixed by the calendar, flagged exploratory, and not used for the headline. |
| Merge attribution | Some PRs were merged by an agent acting for the maintainer, which is not the pipeline. | Separated into routes; the pipeline-versus-hand split is the first thing the results page shows. |
| Moving target | The pipeline keeps merging while the measurement is taken. | The audit population is frozen and its hash recorded; PRs merged afterwards are counted as not covered. |
| No human baseline for the audit | We do not know what the same audit would find in human-written PRs. | Not run. Only reverts are compared across the two groups; the audit is not. |
What would change the conclusion
- A larger, later sample. A second sample of PRs merged since the first week would say whether the concentration of problems early on was real. With none confirmed in 11 later audits, the data cannot rule out a later rate above a quarter.
- The same audit on human-written PRs. If human PRs show a similar rate, the pipeline is not worse than the baseline; if they show a lower one, the pipeline’s early rate was a real cost.
- A second auditor family, and a person re-deriving a random subset of the rulings from the code.
- A second repository, to test whether the method and the pipeline travel.
- Measuring the merge decision itself. The largest finding here is the latency between a pipeline-opened PR and a human merge. Whether letting the pipeline merge more is safe is an experiment, and this audit is the instrument for it.
How the numbers are produced
One command in the private harness reads its own raw records and writes one file. It composes three measurements: the post-merge scorecard (merge routes, escapes, the distinct-issue funnel), the audit summary (the frozen sample, the results, the rulings), and the cost reconstruction (agent session logs and a dated price table). It then:
- refuses to call the report final unless every audit item has a current ruling signed by a person, nothing it rests on is stale, and the audit still matches the repository it was drawn from;
- drops any claim whose numbers are unknown instead of printing a zero;
- builds the public view from an allowlist, so a field added to the reports later stays private until it is added on purpose: no PR or issue numbers, no per-item lists, no rulings’ reasons, no repository name;
- writes the limits next to the numbers, including how many rulings were accepted on an agent’s reasoning.
The site reads that file at build time and every figure on these pages comes from it. The file itself is public. If a number here disagrees with it, the file wins.
Reproducing the method elsewhere
The method does not depend on the harness, only on four things being recorded:
- which pull requests were authored by the automation, and which actions each one’s run took (so merge routes can be joined);
- a frozen, seeded, stratified random sample, with its population recorded, and a different model family reading each cold, with the issue as the specification and the whole diff and nothing else;
- a person’s signed, explained ruling on every finding and on a random spot-check of the clean ones, bound to the exact audit;
- the agents’ own session logs, joined to issues by where they ran.