Skip to content

Results

Every figure, with its denominator and its interval.

The measurements are snapshots of one repository over 28 days. Proportions carry a Wilson 95% interval, because the interesting cases here are the extreme ones: “none in 136” is a statement about how rare something is, not a claim that it is absent.

Who pressed merge

The pipeline authored and gated 136 pull requests that were merged. The headline “136 PRs” overstates what it did alone, so each merge is joined back to the pipeline’s own run records to see who actually pressed the button.

Merged by the pipeline: 50 (37%)Opened by the pipeline, merged by a person or an agent at their direction: 66 (49%)No run record names the PR: 20 (15%)
Of 136 pipeline-authored PRs merged on the repository, 50 (37%, 95% interval 29–45%) were merged by the pipeline itself.
Merge route and time from PR opened to merged (median)
RoutePRsShareMedian open → merge
The pipeline itself5037%0 minutes
A person or an agent, after the pipeline opened it6649%21 hours
No run record2015%—

What was caught after merge

  • Reverted0 of 136
    0 of 136; 95% interval 0.0–2.7%
  • Confirmed to have a real problem (audit sample)5 of 30
    5 of 30; 95% interval 7–34%
Share of merged pipeline PRs with each outcome, on one 0–100% axis. The bar is the 95% interval.
Signal after mergeCountHow it is judged
Reverted0 of 136Mechanical: a later commit says it reverts the PR’s merge commit.
Named by a later issue2 candidates, 0 confirmedCandidates only: a person rules whether the issue was a defect in that PR.
The post-merge gate turned red2 events: 1 real, 1 flakeA person rules real or flake. Not pinned on one PR: it names a commit range.

The same measurements on the repository’s human-authored PRs over the same history: 66 merged, 0 reverted. The pipeline’s PRs are larger (median 193 changed lines against 92), which is expected of a system that is handed well-specified bug reports.

The funnel, by distinct issue

The pipeline’s own run totals count every attempt, so an issue retried six times because the host was short of memory counts six times. Here each issue counts once, by its best outcome across every run. 10 issues never started because of the host or an outage; that says something about the machine, not the pipeline, so they are excluded from the 211 actionable ones.

  • Merged50
  • PR open54
  • Needs a human42
  • Declined57
  • Failed7
  • Never started10
221 distinct issues attempted. The merged rate is 23.7% (95% interval 18–30%); 104 reached a pull request.

“Declined” means the pipeline looked and correctly did nothing: the bug was already fixed, or the change belonged to a part of the system it is not allowed to touch. The median merged issue needed 1 attempt and took 5 hours from first drive to merge.

The blind audit

30 of the 130 merged PRs (seed 7) were read cold by a different model family, without the pipeline’s verdicts, and a person ruled on every finding. The method is on the audit page; the results are:

What the auditor said about the 30 sampled PRs
QuestionAnswers
Resolves the issue?yes 25 · partial 5
Introduces a defect?none 25 · minor 2 · major 3
Scopetight 30
Tests pin the behaviour?meaningful 29 · weak 1
Flagged (a defect, or not fully resolved)8 of 30
  • Precision of the auditor’s flags5 of 8 flags were real
    5 of 8; 95% interval 31–86%
  • Problems the auditor missed in PRs it called clean0 of 8 spot-checked
    0 of 8; 95% interval 0–32%
The auditor’s own error rates, measured against a person’s rulings.

Agreement between the auditor and the person who ruled, beyond chance, was κ = 0.63 over 16 PRs. Of the 8 flags, 5 were ruled real and 3 were false alarms. The auditor over-flags, which is why every flag was ruled on. A spot check of 8 of its clean verdicts found nothing it had missed, though with n = 8 that only bounds the miss rate below about a third.

The overall rate hides where the problems are: all 5 came from PRs merged in the first week of operation, and none from the 11 audited PRs merged later. The audit page has the weekly table and the reasons not to over-read it.

What it costs

Cost is rebuilt after the fact from the agents’ own session logs, with no change to the pipeline. Tokens are measured; dollars are an API-equivalent estimate at list prices retrieved 2026-10-03, because both agents run on flat plans.

Merged PRs with usage on disk104 of 136 (76%)
Median tokens per merged PR1,470,655
Median API-equivalent cost per merged PR$1.50 (mean $2.98)
All spend, API-equivalent$445.76 across 167 issues
Tokens spent on issues that merged69%
Tokens spent on issues that did not31%

The veto gate runs in a throwaway directory and is not in these figures, and logs older than the retention window are gone, which is why coverage is 76% and not 100%.