Reference
Every figure on these pages is in one file.
shipper-evidence.json is generated from the pipeline’s raw records and read by these pages at build time. This page documents it. Generated 2026-10-04 02:45 UTC; status FINAL.
Top-level fields
| Field | What it holds |
|---|---|
| generated_at, status | When it was generated and whether the report was FINAL or DRAFT. |
| window | First and last merge of a pipeline PR, and the days between. |
| cohorts | Merged counts, lines added and deleted, and median size for the pipeline’s PRs and the human PRs on the same repository. |
| autonomy | The three merge routes (pipeline, hand, unrecorded): PR count, share with its 95% interval, and median hours from opened to merged. |
| baseline | The human cohort’s PR count and reverts, measured the same way. |
| escapes | Reverts with their interval; later-issue candidates and how many were confirmed; post-merge gate events split real and flake; how many of those rulings were accepted on an agent’s reasoning. |
| funnel | Distinct issues by best outcome, the merged rate and interval, attempts, and time from first drive to merge. |
| audit | The frozen sample (size, seed, population, strata), the auditor’s answers, flags and rulings as counts, precision, miss rate, agreement, confirmed problems, and the weekly split. |
| cost | Coverage, tokens and API-equivalent dollars per merged PR, total spend, and the share of tokens on work that did or did not merge. |
| claims, limits | The generated summary sentences and the generated limits, verbatim. |
Conventions
- Intervals are 95% Wilson score intervals, as
[low, high]on 0 to 1. - Shares and rates are fractions on 0 to 1; the pages format them as percentages.
- null means the figure could not be computed from the evidence. It is never zero.
- Hours are medians, which is why a pipeline merge shows as a fraction of an hour.
Glossary
| Term | Meaning |
|---|---|
| Author | The model that writes the fix, working in its own worktree. |
| Adversarial reviewer | A model from another vendor, read-only, that tries to break the diff. |
| Veto gate | A third model that sees only the issue and the diff and may refuse. |
| Sensitivity gate | A manifest of paths that always need a person, including the gate code itself. |
| Worktree | An isolated checkout of the repository in which one issue is worked. |
| Distinct issue | An issue counted once, by its best outcome across all runs. |
| Merge route | Who pressed merge: the pipeline, a person or an agent at their direction, or unrecorded. |
| Escape | A problem found after a merge. |
| Candidate | A possible escape that a script found and a person has not yet ruled on. |
| Ruling | A signed, explained yes or no on whether a finding is a real problem, bound to the audit it was made on. |
| Provisional ruling | An agent’s proposed ruling, shown beside the numbers and counted in none of them. |
| Accepted ruling | A person’s signature on an agent’s written reasoning, without re-deriving it. Always disclosed. |
| Spot-check | A seeded random pick of PRs the auditor called clean, checked to measure what it missed. |
| Wilson interval | A confidence interval for a proportion that stays honest at the extremes, such as zero in 136. |
| DRAFT / FINAL | FINAL only when every item has a current, signed ruling and nothing the report rests on is stale. |