Results
Every figure, with its denominator and its interval.
The measurements are snapshots of one repository over 28 days. Proportions carry a Wilson 95% interval, because the interesting cases here are the extreme ones: “none in 136” is a statement about how rare something is, not a claim that it is absent.
Who pressed merge
The pipeline authored and gated 136 pull requests that were merged. The headline “136 PRs” overstates what it did alone, so each merge is joined back to the pipeline’s own run records to see who actually pressed the button.
| Route | PRs | Share | Median open → merge |
|---|---|---|---|
| The pipeline itself | 50 | 37% | 0 minutes |
| A person or an agent, after the pipeline opened it | 66 | 49% | 21 hours |
| No run record | 20 | 15% | — |
What was caught after merge
- Reverted0 of 1360 of 136; 95% interval 0.0–2.7%
- Confirmed to have a real problem (audit sample)5 of 305 of 30; 95% interval 7–34%
| Signal after merge | Count | How it is judged |
|---|---|---|
| Reverted | 0 of 136 | Mechanical: a later commit says it reverts the PR’s merge commit. |
| Named by a later issue | 2 candidates, 0 confirmed | Candidates only: a person rules whether the issue was a defect in that PR. |
| The post-merge gate turned red | 2 events: 1 real, 1 flake | A person rules real or flake. Not pinned on one PR: it names a commit range. |
The same measurements on the repository’s human-authored PRs over the same history: 66 merged, 0 reverted. The pipeline’s PRs are larger (median 193 changed lines against 92), which is expected of a system that is handed well-specified bug reports.
The funnel, by distinct issue
The pipeline’s own run totals count every attempt, so an issue retried six times because the host was short of memory counts six times. Here each issue counts once, by its best outcome across every run. 10 issues never started because of the host or an outage; that says something about the machine, not the pipeline, so they are excluded from the 211 actionable ones.
“Declined” means the pipeline looked and correctly did nothing: the bug was already fixed, or the change belonged to a part of the system it is not allowed to touch. The median merged issue needed 1 attempt and took 5 hours from first drive to merge.
The blind audit
30 of the 130 merged PRs (seed 7) were read cold by a different model family, without the pipeline’s verdicts, and a person ruled on every finding. The method is on the audit page; the results are:
| Question | Answers |
|---|---|
| Resolves the issue? | yes 25 · partial 5 |
| Introduces a defect? | none 25 · minor 2 · major 3 |
| Scope | tight 30 |
| Tests pin the behaviour? | meaningful 29 · weak 1 |
| Flagged (a defect, or not fully resolved) | 8 of 30 |
- Precision of the auditor’s flags5 of 8 flags were real5 of 8; 95% interval 31–86%
- Problems the auditor missed in PRs it called clean0 of 8 spot-checked0 of 8; 95% interval 0–32%
Agreement between the auditor and the person who ruled, beyond chance, was κ = 0.63 over 16 PRs. Of the 8 flags, 5 were ruled real and 3 were false alarms. The auditor over-flags, which is why every flag was ruled on. A spot check of 8 of its clean verdicts found nothing it had missed, though with n = 8 that only bounds the miss rate below about a third.
The overall rate hides where the problems are: all 5 came from PRs merged in the first week of operation, and none from the 11 audited PRs merged later. The audit page has the weekly table and the reasons not to over-read it.
What it costs
Cost is rebuilt after the fact from the agents’ own session logs, with no change to the pipeline. Tokens are measured; dollars are an API-equivalent estimate at list prices retrieved 2026-10-03, because both agents run on flat plans.
| Merged PRs with usage on disk | 104 of 136 (76%) |
| Median tokens per merged PR | 1,470,655 |
| Median API-equivalent cost per merged PR | $1.50 (mean $2.98) |
| All spend, API-equivalent | $445.76 across 167 issues |
| Tokens spent on issues that merged | 69% |
| Tokens spent on issues that did not | 31% |
The veto gate runs in a throwaway directory and is not in these figures, and logs older than the retention window are gone, which is why coverage is 76% and not 100%.