Repository navigation
feat(release): T-1 deep, dogfood and models start together and join before readiness - #4887
Conversation
…nd preflight
The autopilot created a public release right after the tag push, ahead of
the cleanroom, assets and preflight steps. binary-release.yml fired on
`release: published`, so the assets could only be built by publishing first:
the order was inverted by design.
Now:
tag step gh release create --draft, then binary-release.yml is
dispatched on the tag (workflow_dispatch, input tag, the tag's
own workflow file). Every upload step already finds the
release by listing, which returns drafts.
assets step waits for that dispatch run (--event workflow_dispatch).
preflight writes "PASS <tag> <commit>" on success, truncated first.
publish (new step, between preflight and dryrun) publish_release()
re-reads each fact itself and only then runs
gh release edit --draft=false:
clean-room (aprender) job = success on the recorded run;
a preflight PASS naming this tag and commit;
check_release_assets.sh <tag> = 0 (1 and 2 both refuse);
the release is still a draft.
Publishing then fires `release: published` once more; that run finds every
asset present and rebuilds nothing (#4286). No token is widened and no check
is renamed; binary-release.yml is unchanged.
Case table: scripts/release/check_release_draft_gated.sh extracts the
autopilot's own tag..publish step bodies and runs them against a stub gh
that records the call order (gh api writes, failing edits, a missing
release and --jq filters modelled; rows may resume over shared state),
plus 4 structural rows over the whole file: only publish_release() sets
draft to false, every create is a --draft, publish_release is called once
from the publish step, and publish follows its gates in STEPS and file.
before (origin/main 316dee2 autopilot.sh):
bash scripts/release/check_release_draft_gated.sh <main copy>
17/17 rows WRONG + 3/4 structural WRONG; every run row reads
public-EARLY (created public at the tag step) or, for publish-alone
rows, no refusal at all
after (this commit):
bash scripts/release/check_release_draft_gated.sh
17/17 rows + 4/4 structural ok, mutants 19/19 killed, 18 s
It lives in scripts/release/, outside guard_tree's universe: report-only
until three green nights (L31).
Contract: tag-step-milestone-gate-v1 1.0.0 -> 1.1.0, equation
release_is_draft_until_gated, TSMG-INV-004, FALSIFY-TSMG-007.
ont-delta: none (extends an existing pattern contract; no new kind, class or edge)
Agent: aprender-78
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Review round 2 found three ways past the table: a publish in a step it does not run, spelled through a variable (--draft=$F); publish_release called through a variable; and a clean-room check weakened to `!= failure`, which only success/failure fixtures could not see. - structure: outside publish_release(), with `\` continuations joined, any draft set by value (--draft=X, -f/-F draft=X, a JSON "draft": X) and any `gh api` write to a release (-X/--method/-f/-F/--field/--input) is RED, whatever X is; a plain `--draft` stays allowed, as does a notes edit. - structure: any mention of publish_release, not only a call followed by a space, counts toward "called once, from the publish step". - rows: a cancelled clean-room (full run and publish alone) and an empty conclusion (publish alone). - mutants: draft-by-variable, publish-by-variable, api-field-variable, api-input-body, cleanroom-not-failure. before (origin/main 316dee2 autopilot.sh): bash scripts/release/check_release_draft_gated.sh <main copy> 20/20 rows WRONG + 3/4 structural WRONG after (this commit): bash scripts/release/check_release_draft_gated.sh 20/20 rows + 4/4 structural ok, mutants 24/24 killed ont-delta: none (prose of FALSIFY-TSMG-007's prediction; no new kind, class or edge) Agent: aprender-78 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-78 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ity guard can judge release scripts (#4690) check_release_scripts_derive_identity.sh reads every script under scripts/release/ and refuses a literal train identity (R2) or a literal count (R3). check_release_draft_gated.sh builds a fixture train with a fixed version and a fixed row count, so it was RED there on two lines. It is a guard, not a release script: it moves to scripts/ beside the other check_*.sh guards and reads autopilot.sh from scripts/release/. No rule of the identity guard changes. Measured on this tree: bash scripts/check_release_draft_gated.sh rc 0, mutants 24/24 bash scripts/check_release_scripts_derive_identity.sh rc 0 (was R2 + R3 FAIL) Agent: aprender-78 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
autopilot ran `rc_publish_gate.sh --verify` (every publishable tarball
built against the local overlay, = cargo publish --dry-run for the whole
cascade) in the `dryrun` step, after the tag was pushed and the draft and
asset build existed. A tarball that does not compile was found only after
the irreversible steps.
Now the tag step runs it first, on the worktree at the release commit the
tag will name, and dies on red before cut_tag: no carry, no tag, no draft.
A green run writes publish-dryrun-commit = the release commit. The
`dryrun` step keeps cascade-publish.sh --check and the T-4 receipt, and
requires that receipt to name exactly the release commit; a missing one
or another commit's is red, never a skip. STEPS and their order do not
change, so every anchor that reads them stays put.
Guards:
- rc_publish_gate.sh --self-test: the wiring check now requires --verify
and its die inside `if run_step tag`, before cut_tag, and a fixture with
the old order (dry run after the tag) is refused by the same predicate.
- release-ready list: RR-P07 moves to RR-T16 (tag). RR-T17 and RR-P37
list the asset dispatch in `tag` and the asset check in `publish`, which
this branch had left unlisted. The header's step list names `publish`.
- check_release_draft_gated.sh: its harness stubs rc_publish_gate.sh and
sets WT, as it already stubs tag_coverage_gate.sh. No case or mutant
changes.
Measured on this tree:
bash scripts/release/rc_publish_gate.sh --self-test PASS (both wiring rows ok)
bash scripts/release/release_ready.sh unlisted=0 orphaned=0 contradictions=1
(was unlisted=2 contradictions=2; the one left is publish.executor,
also on origin/main 682dab1)
bash scripts/check_release_draft_gated.sh rc 0, mutants 24/24
bash scripts/check_tag_coverage_gated.sh (+ --self-test) rc 0
bash scripts/check_release_scripts_derive_identity.sh rc 0
Agent: aprender-78
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…efore readiness The three T-1 steps were run in series and the first red ended the pass. They are three independent measurements of the same commit, so they now start together as background lanes and join before readiness. The rule is unchanged: a red in any lane stops the pass. The first red stops the lanes still running (each lane is its own process group, so its cargo and ssh children go with it) and nothing after the join runs. Each lane keeps its own log as before and gets one row in $AP/t1-steps.tsv (step, start, end, seconds, verdict: GO, RED, or STOPPED when a sibling was red first), plus one STEP line in STATUS. All three build release/apr with different features. Run in parallel, one shared target dir would be a race over which apr each lane measures. So deep and dogfood build into their own dirs under it, and models keeps CARGO_TARGET_DIR: readiness still reads the apr that models built, as before. The step bodies and every die line are byte-identical. scripts/check_release_t1_lanes_joined.sh runs the join block against stub lanes (all green; one red stops a running sibling and its child; a single step; the lanes overlap) and kills 4 mutants (no final die, no kill, no rows, serial start). It is RED on the serial autopilot. Agent: aprender-wprodmodels Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…new guard Since it moved to scripts/, guard_tree runs it on every merge, and it had no report mode: a new check would have blocked from its first night. It now follows the row-order guard's shape. A wrong row or a surviving mutant prints a REPORT line and exits 0. RELEASE_DRAFT_GATED_ENFORCE=1 makes it exit 1. ENV stays rc 2 in both modes. The contract's falsification test runs it under ENFORCE, so the contract still fails on a broken table. The header no longer says guard_tree skips it. Measured on this tree: rc 0 by default and under ENFORCE. The create-public mutant on a copy of the autopilot gives 8 WRONG rows, rc 0 by default and rc 1 under ENFORCE. A missing autopilot is rc 2. Refs #4690 Agent: aprender-78 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e guard pins the target dirs Review of the first commit found three things the parallel lanes changed: - The remote leg of models is an ssh with no pty. Killing the local lane left the remote build and ladder running in the release dir the next pass reuses. A red now stops only deep and dogfood, which run only here; models runs to its own verdict, and every red lane is named in the STOP line. - dogfood's C14 parity and the model ladder now run on the same GPU at once. The ladder takes the fleet GPU lock per apr call; parity did not. Each parity run now takes the same lock (a lock not free in time is exit 75, read as FAIL). - The guard stubbed the lanes, so deleting the per-lane target dirs stayed green. It now checks that deep and dogfood build into their own dirs and that models keeps CARGO_TARGET_DIR, with two mutants for it. The lanes no longer share the autopilot's process group, so INT or TERM to the autopilot now stops every lane (a trap held for the join). The guard is now 7 cases and 9 mutants: no-die, no-kill, stop-models, no-rows, serial, first-red, no-trap, shared-deep, shared-dogfood. Agent: aprender-wprodmodels Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…se PR moved) Agent: aprender-wprodmodels Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The receipt and findings for 951a030: the draft guard green at head under enforce, the publish gate's self-test green, and the guard's create-public mutant turning 8 rows wrong. The cross-vendor reviewer could not be consulted (the budget probe exited 126), so the verdict is DEGRADED, not PASS. Refs #4690 Agent: aprender-78 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-wprodmodels Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…r STOPPED, parity lock is pinned From the C129 review of 19af0aa: - Two lanes red in the same second: the second was recorded STOPPED and left out of the STOP line. STOPPED now needs the lane to have died of the TERM we sent (exit 128 + TERM, derived with kill -l); any other non-zero exit is RED. New guard case self-red, and mutant stopped-any-rc. - Parity's GPU lock was untested (a flock -> env mutant survived). The guard now checks that every apr parity in check_model_parity.sh takes the ladder's lock, and mutant parity-unlocked is killed. - A lock not free in time (exit 75) now writes its reason, so the FAIL names the lock instead of an empty reason. - Parity runs under choom 1000, as the ladder's apr calls do. - The comment now says INT/TERM stops every lane, models included. Guard: 8 cases + target dirs + parity lock, 11 mutants killed. Agent: aprender-wprodmodels Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…se PR moved) Agent: aprender-wprodmodels Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-wprodmodels Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
|
Review of #4887 at d4ad3f2: NO. There is one blocking finding. The release-path change itself reads correct. Finding:
What I checked and found sound
Non-blocking
|
…lease guard guard_tree runs every scripts/check_*.sh, and this one ran bare. Its red, overlap and interrupt cases read the wall clock. A new check blocks only after three green nights (L31), so it now prints a REPORT line and exits 0 by default. RELEASE_T1_LANES_ENFORCE=1 makes it exit 1, and ENV stays rc 2. Measured: on the serial autopilot, report mode gives rc 0 with REPORT and ENFORCE=1 gives rc 1. On the real tree, ENFORCE=1 passes (8 cases, 11 mutants). Agent: aprender-wprodmodels Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-wprodmodels Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…draft-release guard does From the C129 review of 4f55ddf: a failed baseline case called finish at once, so report mode printed no mutant lines. The baseline result is now held and the verdict comes after the mutants. The success path ends in finish 0. Measured on the serial autopilot: report mode gives rc 0 with 13 mutant lines, and ENFORCE=1 gives rc 1. On the real tree, ENFORCE=1 passes. Agent: aprender-wprodmodels Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-wprodmodels Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
|
Re-review of #4887 at 472096b: YES. My earlier finding is closed.
Non-author sign-off round, bound to 472096b: AGREED 2/2 (claude-sonnet-5-5, gemini-3.1-pro-high). Both lanes said YES to the join block as written and NO to a planted decoy that records a red lane and continues to readiness. Both cited |
|
Second non-author review, head 472096b: YES Scope: Measured on a detached checkout of 472096b:
Read and checked:
Advisory (asserted, not blocking):
|
…ain) Agent: aprender-wprodmodels Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
§13.11 rung 1 — quorum shadow verdict Shadow mode: this records a verdict and merges nothing. A refusal |
Agent: aprender-wprodmodels Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
|
#4887: my YES and the non-author sign-off round (AGREED 2/2) from 472096b now bind to a0e49c2.
|
…02, SC2242) x86-main guard-cargo was red on #4887 for two real findings in the scripts this PR adds: - m16 (bashrs SEC/DET/IDEM): `date +%s` timers at autopilot.sh:188/196 and check_release_t1_lanes_joined.sh:71/76 are DET002. Durations now come from bash's $SECONDS, as dogfood.sh already does. The t1-steps.tsv start and end columns are seconds after the T-1 launch; STATUS's "T-1 LANES started together" line carries the wall-clock anchor. No disable-line, no baseline change. - m12 (bashrs error ratchet, 6 -> 7): the pinned bashrs 7.4.1 read `continue` inside the stop-models mutant's sed string as a loop keyword (SC2242). The regex now spells it cont[i]nue, which matches the same text; the mutant still applies and is killed. Measured on this tree: check_bashrs_gate.sh PASS (0 SEC/DET/IDEM errors over 517 files), check_shell_lint_ratchet.sh PASS (6, baseline 6), the T-1 guard passes in report and enforce mode (8 cases, 11/11 mutants killed), and the other autopilot guards are unchanged. Agent: aprender-wprodmodels Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-wprodmodels Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
|
Re-review of #4887 at 4672271: YES.
Fresh non-author sign-off round, bound to 4672271 (the judged block changed, so this is a new round, not a re-bind): AGREED 2/2 (claude-sonnet-5-5, gemini-3.1-pro-high). Both lanes said YES to the join block as written and NO to a planted decoy that records a red lane and continues to readiness. Both cited |
|
Re-review, head 4672271 (fix 1b6603e): YES My review covers What I measured on a detached checkout of 4672271:
The ordering and verdict logic of the join is unchanged from 472096b, so my earlier reasoning still holds. The two advisories from my first review still apply. |
What
In
scripts/release/autopilot.sh, the T-1 steps deep, dogfood and models now start together and join before readiness. Before this they ran in series.Stacked on #4883 (the draft-release change). Base is its branch; review only the top commit.
The rule is unchanged
A red in any lane still stops the pass. The script dies naming every red lane. Nothing after the join runs, so nothing is tagged.
What a red stops:
set -m), so its cargo children are stopped with it.The GPU lock
In parallel, dogfood's C14 parity and the models ladder can run on the same GPU at once. The ladder already takes the fleet GPU lock for every
aprcall; parity did not.Now each
apr parityinscripts/check_model_parity.shtakes the same lock, with the same wait. A lock that is not free in time exits 75, which parity reports as a FAIL naming the lock, never as a pass. Parity also runs underchoom 1000, as the ladder does. The parity self-test is 26/26.The per-pass table
Each lane still writes its own log, as before. It also gets:
$AP/t1-steps.tsv, with the columns step, start, end, seconds and verdict;STEP <name> <verdict> rc= seconds=line in STATUS.The verdict is one of:
Target dirs
All three lanes build
release/apr, each with different features. In parallel, a shared target dir would be a race over whichapreach lane measures. So:target/t1-deep.target/t1-dogfood.CARGO_TARGET_DIR. Readiness still reads theaprthat models built, exactly as before.Everything else is unchanged:
dieline are byte-identical.from=to) runs that lane alone.Proof
New guard:
scripts/check_release_t1_lanes_joined.sh. It runs the join block against stub lanes in eight cases:STOPPED means the lane died of the TERM the join sent (exit 128 + TERM). Any other non-zero exit is RED.
It also checks the scripts themselves: deep and dogfood build into their own target dirs, models keeps
CARGO_TARGET_DIR, and everyapr paritytakes the GPU lock.It kills 11 mutants: no-die, no-kill, stop-models, no-rows, serial, first-red, no-trap, shared-deep, shared-dogfood, stopped-any-rc, parity-unlocked. It is RED on the serial autopilot.
Every existing autopilot guard passes on this head:
check_release_models_t1(41 rows)check_release_autopilot_dogfood_close(17)check_tag_step_gatedcheck_release_draft_gated(20 + 4)check_dogfood_matrix_is_visitedcheck_tag_coverage_gatedcheck_release_host_receipts(94)check_milestone_cut(31)check_release_scripts_derive_identitybashrs lintreports 1 error onautopilot.sh. It is SC2135 at line 33, and the base has the same error.keep-open: #4883 is the PR this one stacks on, not an issue this PR resolves.
🤖 Generated with Claude Code