Skip to content

Engineering log

What broke, what review caught, and what changed.

The most useful thing a measurement project produces is a list of the ways it was nearly wrong. These are the entries that changed the design, each as symptom, cause, fix, and what it taught.

The veto gate ran out of credit and every batch ended with nothing

Symptom. Batches began ending with no pull request. Cause. The model behind the last gate started answering that it required usage credits, its fallback was at its limit too, and nothing checked the last gate before the first expensive step, so each issue was authored, verified and reviewed and then parked. A credit lapse does not clear by waiting, so the next batch wasted the same work. Fix. The preflight now probes the last gate before any spend, mirroring what the gate does call for call, with the same fallback rule, the same verdict contract and a timeout no longer than the gate’s. What review added. The first version of the probe was more lenient than the gate: it fell back on an error payload with a success exit code, and accepted any non-empty answer from the fallback. A probe that accepts more than the gate passes exactly where the gate fails.

It taught: probe every dependency of the last step before the first expensive one, and make the probe no more forgiving than the thing it guards.

A usage limit failed an issue instead of waiting

Symptom. One issue failed on a message about a session limit, the next paused the batch, and five more sat queued. Two leftover empty branches were treated as “taken” forever, and one worktree held real uncommitted work that the existing clean-up would have deleted. Cause. The limit message matched nothing the classifier knew. Fix. Recognise it, wait it out by polling, and resume the same issue in place. Because the stated reset time proved unreliable in both directions, the wait probes on its own schedule and resumes the moment the provider answers.

Seven review rounds on that one feature; the first six each found a real defect
What review foundThe failing scenario
A substring match decided what counts as “dependencies, not work”A real file whose name contained “node_modules” was treated as an install artefact, and the clean-up deleted its worktree. The special case was then found to hide work three different ways and was removed outright.
The recovery probe checked the wrong modelThe wait probed the author’s default model while the limited thing was a different model with a different fallback, so it resumed straight into the same pause.
A remote-only leftover measured as emptyA branch full of someone’s commits, known only from the remote, counted as zero commits ahead and would have been healed away. In a single-branch clone even the fix failed until it used an explicit refspec.
Unclamped configurationA poll interval of zero probed the provider in a tight loop; a negative resume count made the loop’s range empty, so the issue silently never ran.
A failed status read as cleanA corrupt index made the work check print nothing, and nothing was read as no work.
PackagingThe new module was missing from the package’s module list, so a built wheel would have failed to import. A guard test now fails if any top-level module is not packaged.

It taught: recovery code is the most dangerous code in a system that deletes things, and it needs the most adversarial review.

The audit’s first reading hid where the problems were

Symptom. “Five of thirty” read as a single rate. Cause. Nothing in the report asked whether the problems were spread out. Fix. Splitting the same rulings by merge week showed all five in the first week and none in the 11 audited PRs after it. That does not prove the pipeline improved (the sample is small and the early work was harder), but one overall number would have misdescribed both stories.

It taught: after computing a headline, cut the data in the ways most likely to embarrass it, and publish what the cuts show, flagged as exploratory.

The same stale-state defect, six different ways

Symptom. Review of the measurement stack kept finding the same family of defect in different places. Cause. Each artefact was keyed by something that could silently point at something else.

ArtefactKeyed byWhat could go wrongNow bound to
A person’s rulingThe PR numberA re-audit says something different and the old ruling still counts.A digest of the exact verdict, commit and prompt.
The spot-check setA fileA redrawn sample reuses a set drawn from another one.A fingerprint of the sample and the audit.
An audit resultThe PR numberA re-merged PR or an older prompt is reported as current.The merge commit, the prompt version and a content check.
The sampleA fileOne repository’s audit published as another’s.The repository and the exact population list.
A dollar figureA price tableMonths-old prices give a precise wrong number.The table’s source and retrieval date.
A cost attributionA directory nameTwo repositories share a sanitised prefix, or one is reached through a symlink.The session’s own recorded start directory, under either spelling, and only issues the pipeline drove.

It taught: whenever a number or a decision is stored, ask what it would take for the thing it refers to to change underneath it, and bind it to that.

A diff that was not the whole change

Symptom. Review pointed out that a rebase-merged pull request records only its last commit as its merge commit, so the audit would read a fraction of the change and could pass a PR it had barely seen. Fix, in stages. First compare the changed-line count with the platform’s record of the PR’s size; review showed equal counts prove nothing (one commit changing a line from one to two and another from two to three has the same size as the net change one to three). Then compare the content of the changed lines against the platform’s own diff; review showed that a rename with no content change is invisible to that, so file-level facts joined the comparison; then that a single bag of lines lets changes swap between files, so the comparison became per file. All 30 audited results were re-verified against each stricter version.

It taught: a check that passes is only as good as the thing it compares; keep asking what a partial input would look like and whether it would pass.

Cost counted one and a half times

Symptom. The first issue checked read twenty-two usage lines for fourteen model messages. Cause. The agent’s transcript repeats a message across lines, one per content block, with the same usage. Fix. Usage is taken once per message identifier, with a regression test built from the real proportions. A reviewer’s totals, by contrast, are cumulative within a session and already include their cached input, so the last count is the total and the cached part is split out, not added.

Merging is not shipping

Symptom. The pipeline merged for days while production ran an older build: the first time anyone looked, twenty-one undeployed commits, including five security fixes, sat on the base branch. Cause. Deploying was a separate manual step that nothing watched. Fix. A drift detector compares the running version with the base branch on the paths that deploy, keyed on the oldest undeployed commit, and files an issue once it is old enough to matter. The deploy hook itself stays blank by default, so merging and deploying remain separate decisions.

One thing that is still open

  • The post-merge gate can cry wolf on a loaded machine. A “base branch is red” issue was filed from a single failing run of a frontend suite; re-running the same commit on an idle machine passed 727 of 727 tests in 54 seconds, against 152 seconds when it failed. The gate should re-run a failure once before filing, and does not yet: it happened again days later, when a test-runner worker timed out while the machine was short of memory, and a later run of the same commit passed.