Case study
An AI pipeline that merges its own code, and the evidence about it.
Shipper takes a bug report, has one model write the fix, has a second model from a different vendor try to break it, lets a third veto it, and merges only when every gate is green. This case study is about the question that matters more than whether it works: how good are the changes it merges, and how much of the work does it really do on its own?
What the numbers say, in four sentences
- The pipeline writes and gates well, but a person is still the bottleneck. It merged 50 PRs on its own; another 66 it opened and gated, then waited a median of 21 hours for a human to press merge.
- “Zero reverts” is not “zero defects.” A blind audit of 30 randomly chosen merged PRs found 5 with a real problem. Revert counts only see what someone noticed; the audit looks on purpose.
- The problems are concentrated in the first week. All 5 confirmed problems were in PRs merged in the first seven days of operation; none of the 11 audited PRs merged later had one. That is suggestive of a pipeline that improved and is not proof: the sample is small and the early work was harder.
- Most of the effort is not waste, but a third of it is. 31% of all tokens were spent on issues that never reached a merge, which is the price of a pipeline that is willing to give up.
Read it in the order that suits you
Status of this snapshot
Generated 2026-10-04 02:45 UTC from records spanning September 5, 2026 to October 4, 2026. Status: FINAL. “Final” means every audit item has a ruling signed by a person; it does not mean the rulings were re-derived from the code. 16 of 16 audit rulings and 4 of 4 post-merge rulings were signed on an agent’s written reasoning, and the limits page says so too.
The pipeline’s source code is private. The method is documented here in enough detail to be criticised and reproduced. The numbers are in one JSON file.