Architecture
A loop that finds problems, and a chain of gates that decides what ships.
Two things are easy to confuse. The daily pass finds work and files it. The gate chain turns one issue into at most one merged change, and every link in it is allowed to say no. This page describes both, and the properties that stop the whole thing from hurting its host.
The daily pass
An unattended scheduler runs the same four steps several times a day: detect, triage, crunch, report. Detect looks for problems and emits findings. Triage turns each finding into a tracked issue, exactly once. Crunch drives selected issues through the gate chain below. Report writes down what happened, including what could not be checked.
| Detector | Finds | The design choice |
|---|---|---|
| types | The type checker failing on the base branch | If main does not build, every later signal is suspect. |
| secrets | Credential material tracked by git | Reuses the project’s own checker instead of a second opinion. |
| prod_health | Production endpoints not answering | Keyed on the endpoint, not the symptom. |
| deploy_drift | Production running an older commit than the base branch | Keyed on the oldest undeployed commit, so closing the issue cannot silence the next drift. |
| deps | High and critical advisories in dependencies | Keyed on the package set; the count lives in the body. |
| stale_baselines | Known-failure entries pinning issues that are closed | Reads the remote base branch, not the checkout. |
| host | Orphaned test stacks and low disk | The pipeline watching its own footing. |
Every repository-content detector reads through the remote base branch rather than the working tree, and the report records which commit it inspected. A detector that reads a checkout reports on the checkout, not on what ships. The first run flagged a baseline entry the remote had already dropped because the checkout was one commit behind.
The gate chain for one issue
Each issue is worked in its own git worktree, in an isolated copy of the repository, by a stage that can refuse it. The stages run in this order, and a refusal anywhere stops the issue with a named reason rather than a generic failure.
- PreflightThe host
Before any model is called, the pipeline checks free memory and disk, probes that each model it will need can actually answer, and takes the one-batch-at-a-time lock.
Can refuse: a host short of resources, a model that is offline or out of credit, a second batch already running.
- AuthorClaude Code, in a worktree
The author reads the issue, confirms the defect by reading the code, writes the fix and its test, and commits. It is told to look at the real code rather than trust the title.
Can refuse: an issue that is already fixed (recorded as “no change”), or one that would need a file it is not allowed to touch.
- Sensitivity gateA path manifest
The committed diff is checked against a manifest of paths and patterns that need a person’s eyes: the gate code itself, credentials, deployment configuration. It is fail-closed: a path it cannot classify is sensitive.
Can refuse: any diff that touches the machinery that judges it.
- VerificationThe project’s own checks
Changed files are routed through a map of the repository’s package graph to the same blocking checks CI would run. The pipeline never pushes a red branch.
Can refuse: type errors, failing tests, a regression in the end-to-end ratchet.
- Adversarial review loopCodex, read-only
A model from a different vendor, with no ability to edit, attacks the diff. Findings go back to the author, who fixes or rebuts them; the loop repeats until the reviewer’s verdict is CLEAN, it stops making progress, or the time budget runs out. An unparseable verdict is treated as not converged.
Can refuse: a diff the reviewer can still break, or one the author and reviewer cannot agree on.
- Veto gateA third model, diff only
A final reviewer sees only the issue and the diff, in an empty directory with no tools. It may veto. A veto that rests on evidence it could not possibly have is re-checked once, diff-only, and flagged as low confidence rather than silently trusted.
Can refuse: anything that is not a well-formed CLEAN. An offline reviewer is an outage, not a verdict.
- Pull requestThe pipeline
The PR body records provenance: which model authored, which reviewer actually answered each round (a stand-in is labelled as one), which gate answered, and the local checks that ran. Auto-merge is off unless the operator turned it on.
Can refuse: nothing further: this is where a person, or the merge gate, takes over.
- Post-merge gateThe merged base branch
After merges land, the same blocking checks run against the base branch itself. Two branches that are each green can still break the base together, and nothing in the per-issue loop can see that.
Can refuse: a base branch that fails its own checks, which files an issue instead of advancing the verified watermark.
Why three models
The author and the reviewers are deliberately different. A model reviewing its own family’s code shares its blind spots, so the adversarial reviewer comes from another vendor, and the veto gate is a third reading with less context on purpose: it cannot be persuaded by the author’s account of the change because it never sees one. When one provider is offline the pipeline falls back to a stand-in and says so in the PR body, because a single-provider review reported as a cross-vendor one is the plausible-but-false answer this design exists to refuse.
Safety properties around it
| Property | How it holds |
|---|---|
| One batch at a time | A host-wide lock using directory creation, not flock: macOS ships no flock(1), so a flock-based guard there fails open and does nothing, silently. Stale locks are cleared; live ones are respected. |
| Stops before it hurts | Refuses to start below a memory and disk floor, and a watchdog pauses in flight. A run that exhausts the host cannot even record why it stopped. |
| Never deletes what it did not create | Worktrees are registered with an ownership fingerprint. A directory is removed only if the registry proves this tool made it, and an unreadable working tree counts as containing work, not as clean. |
| Cannot spam | A cap on new issues per run, and if it cannot list existing issues it refuses to file at all rather than filing blind duplicates. |
| Deploys only if configured to | Merging and deploying are separate. The post-merge deploy hook is blank by default; a deploy failure files its own finding and is never confused with “the base is red”. |
| Waits out usage limits | A usage limit is not a defect in the issue. The pipeline polls with a trivial probe and resumes the same issue in place, because a stated reset time is a hint, not a promise. |
| Cannot edit its own gate | The sensitivity manifest refuses any issue that touches the veto and sensitivity code, so the machinery that judges changes can only be changed by a person. |