Skip to content

Questions and objections

The objections worth taking seriously.

These are the questions a sceptical reader should ask, answered as plainly as the evidence allows.

Is this just a wrapper around a coding model?

The model writes the fix. Nearly everything else is what happens around it: refusing what it cannot judge, reviewing with a model from another vendor, a veto gate that cannot see the author’s account, recovering from limits and outages without deleting work, never running two batches at once, and a measurement layer that is willing to say “unknown”. The engineering is in that surrounding system, and so are most of the failures in the engineering log.

Why not use an existing benchmark?

Benchmarks measure whether a model can solve a curated task. This measures a deployed loop on a real, unfiltered backlog: how many issues it closed, how many a person had to finish, how long merges waited, what it cost, and what a later reader finds wrong. No benchmark can see who pressed merge. The cost of that realism is that the numbers describe one repository and cannot be compared with anyone else’s.

Why trust a model to audit a model?

Don’t, without measuring it. The auditor is a different model family from the author, it sees only the specification and the change, and its answers are held to strict rules of evidence. Then its error rates are measured against a person’s rulings: of its 8 flags, 5 were real (31–86% interval), and in 8 spot-checked clean PRs it missed 0. It over-flags, which is why every flag needs a ruling.

The rulings were accepted from an agent. Isn’t that circular?

Partly, and the page says so. The code was written by one model family, the findings came from a second, and the agent that pre-read each finding and proposed the ruling is from the same family as the author. The person who signed took responsibility without re-deriving them. 16 of 16 audit rulings carry that disclosure.

What limits the damage: the findings themselves were produced by a different family, five of the eight went against the pipeline, and the rulings that favour it (the three false alarms and the clean spot-checks) are the ones to doubt most. The remedy is a person, or a third family, re-deriving a random subset from the code, which has not been done.

Does “0 reverts” mean it is safe?

No. Reverts only count what somebody noticed. The audit found 5 of 30 sampled PRs with a real problem and none had been reverted.

How does it compare with a human developer?

That was not measured, beyond one number. The same repository’s 66 human-authored PRs had 0 reverts, but the blind audit was only run on the pipeline’s PRs, so there is no human baseline for the thing the audit measures. Running it on a sample of human PRs is the single most informative next experiment.

What does it cost?

A median of $1.50 per merged PR at list prices, from the agents’ own logs, with 31% of tokens spent on issues that never merged. Both agents run on flat plans, so that is an estimate of what the same usage would cost through the API, not what was paid.

Is the pipeline autonomous?

Less than the headline “136 PRs” suggests. It merged 37% of them itself. The rest it authored and gated and then waited, a median of 21 hours, for a person. It runs unattended and files its own work, and the final decision to merge was, for most of its history, a human one.

Is the sample representative?

It is a seeded, stratified random sample of every PR the pipeline had merged on the day it was drawn, so it is representative of that population. It is not representative of the pipeline today: the confirmed problems were all in the first week and the sample holds few audits from the weeks after. The audit page has the weekly split and the reasons not to over-read it.

What was the hardest part?

Not the model. It was making the measurement honest: refusing to quote a number whose inputs could have gone stale, telling candidates from conclusions, and keeping the account of who did what accurate when an agent, a pipeline and a person all touched the same pull request. The same bug, a verdict attached to something that later changed, was found six different ways by independent review.

What would you do next?

  1. Run the same audit on a sample of human-written PRs for a baseline.
  2. Audit a second, later sample to test whether the early concentration was real.
  3. Have a person re-derive a random subset of the rulings from the code, and add a second auditor family.
  4. Run the pipeline and the method on a second repository.
  5. Treat letting the pipeline merge more as an experiment, measured with this audit.

Can I see the code?

Not today. The pipeline’s source is private. The method, the definitions and every aggregate number are published so they can be criticised, and so the approach can be reproduced on another repository.

Are the bugs the audit found still open?

No. All five confirmed problems have been fixed and deployed. Two of them were still live in production when the audit found them. The pipeline wrote both fixes. It merged one itself; for the other its veto gate objected, so it did not merge it, and that fix was merged separately outside the pipeline.