Method
Counting activity is easy. Counting value takes definitions.
The pipeline already counted what it did: attempts, reviews, refusals. None of that says whether a merged change was any good. This page defines what was measured instead, so the numbers can be argued with.
Cohorts
Pipeline-authored pull requests are told apart from human ones on the same repository by the branch they come from, so the comparison is between two cohorts on one codebase with one maintainer, not between two projects. A PR is in the pipeline cohort only if it merged and its branch is one the pipeline creates.
Who pressed merge
“The pipeline merged N PRs” is not one fact. Each merged PR is joined to the pipeline’s own run records and put in one of three routes:
| Route | Meaning |
|---|---|
| Pipeline | A run recorded this PR with the result MERGED: every gate was green and auto-merge was on, so nobody pressed merge. |
| Hand | A run opened this PR and stopped, and it was merged afterwards by a person, or by an agent working at a person’s direction. The pipeline did the authoring and the gating, not the final decision. |
| Unrecorded | No run record names the PR (it predates the records, or was made by hand on a pipeline-style branch). Counted separately, never folded into “pipeline”. |
This join is what keeps the headline honest. Of 136 PRs, 37% were merged by the pipeline itself, and reporting the full count as autonomy would have been the most flattering and least true sentence available.
Distinct issues, not attempts
An issue retried six times because the host was short of memory is six attempts and one issue. Everything about the funnel counts each issue once, by its best outcome across all runs, ranked from merged, to a PR opened, to a hand-off to a person, to a correct decline, to a failure. Runs the host or an outage stopped before the issue started say something about the machine, so they are reported on their own and excluded from the actionable denominator.
There is a guard against drift here: the pipeline emits a fixed vocabulary of result codes, and a test fails if the pipeline ever emits one the scorecard does not know how to classify. Two such codes had been silently falling into an “other” bucket until independent review noticed.
Escapes: what is found after a merge
| Signal | Definition | Judged by |
|---|---|---|
| Reverted | A later commit says “This reverts commit” the PR’s merge commit, or is titled as a revert naming the PR number. A revert’s own number is not counted as a revert. | Mechanical. |
| Named by a later issue | A later, non-automatic issue that names the PR. Follow-ups the author filed on purpose, the deploy-drift alarm, and issues filed because the audit found a defect are excluded. | A person rules each candidate real or not. An un-ruled one is unknown. |
| Post-merge gate red | The post-merge gate found the base branch failing its own blocking checks and filed an event. The event names a commit range, not a culprit. | A person rules real or flake. Never attributed to one PR. |
Reverts are fetched fresh before they are counted. If the repository cannot be brought up to date or its log cannot be read, the report stops: a stale clone that is missing a later revert reads exactly like a clean one.
Intervals, and why “none” is a measurement
Every proportion carries a 95% Wilson score interval. The normal approximation collapses at the extremes: zero reverts in 136 would get an interval of exactly [0, 0] and claim a certainty the data cannot support. The Wilson interval gives an honest upper bound instead (0.0–2.7% here), which is a statement about how rare reverts are and not a claim that there are none. The auditor’s precision and miss rate come with the same kind of interval, and are wide because the samples are small.
Rulings: who may say a finding is real
A script can find candidates. It cannot decide that one is a defect. So a ruling is a yes or no with a name and a reason, recorded in a file the operator owns. A ruling counts only if it is signed by a person (not an agent), is explained, and, for the audit, was made on the audit result that is still current. An agent may attach a provisional ruling with its evidence, and that is shown beside the numbers and is in none of them.
Cost, rebuilt after the fact
The pipeline could only price a run whose console log kept the model CLI’s reported cost, and the author calls never printed one, so per-issue cost was unknown for almost everything. But both agents keep a transcript of every session with token counts and the directory it ran in, and each issue runs in its own worktree. So cost was recovered without touching the pipeline: the author’s sessions by the worktree they ran in, the reviewer’s by its working directory.
| Trap | What it would have done |
|---|---|
| A model message repeated across transcript lines | Summing lines instead of messages counted the first issue checked at 22 lines for 14 messages: a 1.5× overcount. Usage is taken once per message. |
| Cumulative token totals | The reviewer’s totals are cumulative within a session and its input figure already contains the cached tokens; adding either again would overcount. |
| Lossy directory names | The author’s sessions are filed under a path with every non-alphanumeric character turned into a dash, so two repositories can share a prefix. Each session’s own recorded start directory decides, and a repository reached through a symlink is matched under both spellings. |
| A path anyone can use | A worktree name is a path anyone could have used, so only issues the pipeline’s own run records show it drove are counted. |
| Prices that age | A dollar figure carries the date and source of the price table, and a table with no date or older than ninety days makes the figure unknown. |
Dollars are an input, not a fact: tokens are measured, a price per million tokens is a number somebody has to supply, and with both agents on flat plans the result is an API-equivalent estimate and is labelled as one.