Skip to content

feat(release): T-1 deep, dogfood and models start together and join before readiness - #4887

Merged
noahgift merged 22 commits into
mainfrom
feat/c316-2d-autopilot-parallel
Oct 7, 2026
Merged

noahgift merged 22 commits into
mainfrom
feat/c316-2d-autopilot-parallel

Conversation

@noahgift

@noahgift noahgift commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

What

In scripts/release/autopilot.sh, the T-1 steps deep, dogfood and models now start together and join before readiness. Before this they ran in series.

Stacked on #4883 (the draft-release change). Base is its branch; review only the top commit.

The rule is unchanged

A red in any lane still stops the pass. The script dies naming every red lane. Nothing after the join runs, so nothing is tagged.

What a red stops:

  • deep and dogfood run only on the release host. The first red stops whichever of them are still running. Each lane is its own process group (set -m), so its cargo children are stopped with it.
  • models is never stopped. Its remote leg is an ssh with no pty, so killing the local ssh would leave the remote build and ladder running in the release dir that the next pass reuses. models runs to its own verdict.
  • An INT or TERM to the autopilot stops every lane. A trap is held for the join, because the lanes no longer share the autopilot's process group.

The GPU lock

In parallel, dogfood's C14 parity and the models ladder can run on the same GPU at once. The ladder already takes the fleet GPU lock for every apr call; parity did not.

Now each apr parity in scripts/check_model_parity.sh takes the same lock, with the same wait. A lock that is not free in time exits 75, which parity reports as a FAIL naming the lock, never as a pass. Parity also runs under choom 1000, as the ladder does. The parity self-test is 26/26.

The per-pass table

Each lane still writes its own log, as before. It also gets:

  • one row in $AP/t1-steps.tsv, with the columns step, start, end, seconds and verdict;
  • one STEP <name> <verdict> rc= seconds= line in STATUS.

The verdict is one of:

  • GO: the lane exited 0.
  • RED: the lane exited non-zero on its own.
  • STOPPED: a sibling went red first and this lane was stopped.

Target dirs

All three lanes build release/apr, each with different features. In parallel, a shared target dir would be a race over which apr each lane measures. So:

  • deep builds into target/t1-deep.
  • dogfood builds into target/t1-dogfood.
  • models keeps CARGO_TARGET_DIR. Readiness still reads the apr that models built, exactly as before.

Everything else is unchanged:

  • The step bodies and every die line are byte-identical.
  • A single-step run (from = to) runs that lane alone.
  • Trade-off: the two extra target dirs mean two more full builds' worth of disk on the release host.

Proof

New guard: scripts/check_release_t1_lanes_joined.sh. It runs the join block against stub lanes in eight cases:

  • all green;
  • dogfood red: deep is stopped with its child, and models runs to GO;
  • two reds, both named;
  • one lane red on its own;
  • a single step;
  • overlap: three 2 s lanes finish in under 5 s;
  • TERM to the run, which stops every lane;
  • self-red: deep red while dogfood fails by itself; dogfood is RED, not STOPPED, and both are named.

STOPPED means the lane died of the TERM the join sent (exit 128 + TERM). Any other non-zero exit is RED.

It also checks the scripts themselves: deep and dogfood build into their own target dirs, models keeps CARGO_TARGET_DIR, and every apr parity takes the GPU lock.

It kills 11 mutants: no-die, no-kill, stop-models, no-rows, serial, first-red, no-trap, shared-deep, shared-dogfood, stopped-any-rc, parity-unlocked. It is RED on the serial autopilot.

Every existing autopilot guard passes on this head:

  • check_release_models_t1 (41 rows)
  • check_release_autopilot_dogfood_close (17)
  • check_tag_step_gated
  • check_release_draft_gated (20 + 4)
  • check_dogfood_matrix_is_visited
  • check_tag_coverage_gated
  • check_release_host_receipts (94)
  • check_milestone_cut (31)
  • check_release_scripts_derive_identity

bashrs lint reports 1 error on autopilot.sh. It is SC2135 at line 33, and the base has the same error.

keep-open: #4883 is the PR this one stacks on, not an issue this PR resolves.

🤖 Generated with Claude Code

noahgift and others added 14 commits October 4, 2026 13:32
…nd preflight

The autopilot created a public release right after the tag push, ahead of
the cleanroom, assets and preflight steps. binary-release.yml fired on
`release: published`, so the assets could only be built by publishing first:
the order was inverted by design.

Now:
  tag step      gh release create --draft, then binary-release.yml is
                dispatched on the tag (workflow_dispatch, input tag, the tag's
                own workflow file). Every upload step already finds the
                release by listing, which returns drafts.
  assets step   waits for that dispatch run (--event workflow_dispatch).
  preflight     writes "PASS <tag> <commit>" on success, truncated first.
  publish (new step, between preflight and dryrun) publish_release()
                re-reads each fact itself and only then runs
                gh release edit --draft=false:
                  clean-room (aprender) job = success on the recorded run;
                  a preflight PASS naming this tag and commit;
                  check_release_assets.sh <tag> = 0 (1 and 2 both refuse);
                  the release is still a draft.
Publishing then fires `release: published` once more; that run finds every
asset present and rebuilds nothing (#4286). No token is widened and no check
is renamed; binary-release.yml is unchanged.

Case table: scripts/release/check_release_draft_gated.sh extracts the
autopilot's own tag..publish step bodies and runs them against a stub gh
that records the call order (gh api writes, failing edits, a missing
release and --jq filters modelled; rows may resume over shared state),
plus 4 structural rows over the whole file: only publish_release() sets
draft to false, every create is a --draft, publish_release is called once
from the publish step, and publish follows its gates in STEPS and file.
  before (origin/main 316dee2 autopilot.sh):
    bash scripts/release/check_release_draft_gated.sh <main copy>
    17/17 rows WRONG + 3/4 structural WRONG; every run row reads
    public-EARLY (created public at the tag step) or, for publish-alone
    rows, no refusal at all
  after (this commit):
    bash scripts/release/check_release_draft_gated.sh
    17/17 rows + 4/4 structural ok, mutants 19/19 killed, 18 s
It lives in scripts/release/, outside guard_tree's universe: report-only
until three green nights (L31).

Contract: tag-step-milestone-gate-v1 1.0.0 -> 1.1.0, equation
release_is_draft_until_gated, TSMG-INV-004, FALSIFY-TSMG-007.

ont-delta: none (extends an existing pattern contract; no new kind, class or edge)

Agent: aprender-78
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Review round 2 found three ways past the table: a publish in a step it does
not run, spelled through a variable (--draft=$F); publish_release called
through a variable; and a clean-room check weakened to `!= failure`, which
only success/failure fixtures could not see.
- structure: outside publish_release(), with `\` continuations joined, any
  draft set by value (--draft=X, -f/-F draft=X, a JSON "draft": X) and any
  `gh api` write to a release (-X/--method/-f/-F/--field/--input) is RED,
  whatever X is; a plain `--draft` stays allowed, as does a notes edit.
- structure: any mention of publish_release, not only a call followed by a
  space, counts toward "called once, from the publish step".
- rows: a cancelled clean-room (full run and publish alone) and an empty
  conclusion (publish alone).
- mutants: draft-by-variable, publish-by-variable, api-field-variable,
  api-input-body, cleanroom-not-failure.
  before (origin/main 316dee2 autopilot.sh):
    bash scripts/release/check_release_draft_gated.sh <main copy>
    20/20 rows WRONG + 3/4 structural WRONG
  after (this commit):
    bash scripts/release/check_release_draft_gated.sh
    20/20 rows + 4/4 structural ok, mutants 24/24 killed

ont-delta: none (prose of FALSIFY-TSMG-007's prediction; no new kind, class or edge)

Agent: aprender-78
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-78
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ity guard can judge release scripts (#4690)

check_release_scripts_derive_identity.sh reads every script under
scripts/release/ and refuses a literal train identity (R2) or a literal
count (R3). check_release_draft_gated.sh builds a fixture train with a
fixed version and a fixed row count, so it was RED there on two lines.

It is a guard, not a release script: it moves to scripts/ beside the
other check_*.sh guards and reads autopilot.sh from scripts/release/.
No rule of the identity guard changes.

Measured on this tree:
  bash scripts/check_release_draft_gated.sh            rc 0, mutants 24/24
  bash scripts/check_release_scripts_derive_identity.sh rc 0 (was R2 + R3 FAIL)

Agent: aprender-78
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
autopilot ran `rc_publish_gate.sh --verify` (every publishable tarball
built against the local overlay, = cargo publish --dry-run for the whole
cascade) in the `dryrun` step, after the tag was pushed and the draft and
asset build existed. A tarball that does not compile was found only after
the irreversible steps.

Now the tag step runs it first, on the worktree at the release commit the
tag will name, and dies on red before cut_tag: no carry, no tag, no draft.
A green run writes publish-dryrun-commit = the release commit. The
`dryrun` step keeps cascade-publish.sh --check and the T-4 receipt, and
requires that receipt to name exactly the release commit; a missing one
or another commit's is red, never a skip. STEPS and their order do not
change, so every anchor that reads them stays put.

Guards:
- rc_publish_gate.sh --self-test: the wiring check now requires --verify
  and its die inside `if run_step tag`, before cut_tag, and a fixture with
  the old order (dry run after the tag) is refused by the same predicate.
- release-ready list: RR-P07 moves to RR-T16 (tag). RR-T17 and RR-P37
  list the asset dispatch in `tag` and the asset check in `publish`, which
  this branch had left unlisted. The header's step list names `publish`.
- check_release_draft_gated.sh: its harness stubs rc_publish_gate.sh and
  sets WT, as it already stubs tag_coverage_gate.sh. No case or mutant
  changes.

Measured on this tree:
  bash scripts/release/rc_publish_gate.sh --self-test  PASS (both wiring rows ok)
  bash scripts/release/release_ready.sh  unlisted=0 orphaned=0 contradictions=1
    (was unlisted=2 contradictions=2; the one left is publish.executor,
    also on origin/main 682dab1)
  bash scripts/check_release_draft_gated.sh  rc 0, mutants 24/24
  bash scripts/check_tag_coverage_gated.sh (+ --self-test)  rc 0
  bash scripts/check_release_scripts_derive_identity.sh  rc 0

Agent: aprender-78
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…efore readiness

The three T-1 steps were run in series and the first red ended the pass. They are
three independent measurements of the same commit, so they now start together as
background lanes and join before readiness. The rule is unchanged: a red in any
lane stops the pass. The first red stops the lanes still running (each lane is its
own process group, so its cargo and ssh children go with it) and nothing after the
join runs.

Each lane keeps its own log as before and gets one row in $AP/t1-steps.tsv
(step, start, end, seconds, verdict: GO, RED, or STOPPED when a sibling was red
first), plus one STEP line in STATUS.

All three build release/apr with different features. Run in parallel, one shared
target dir would be a race over which apr each lane measures. So deep and dogfood
build into their own dirs under it, and models keeps CARGO_TARGET_DIR: readiness
still reads the apr that models built, as before. The step bodies and every die
line are byte-identical.

scripts/check_release_t1_lanes_joined.sh runs the join block against stub
lanes (all green; one red stops a running sibling and its child; a single step;
the lanes overlap) and kills 4 mutants (no final die, no kill, no rows, serial
start). It is RED on the serial autopilot.

Agent: aprender-wprodmodels
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…new guard

Since it moved to scripts/, guard_tree runs it on every merge, and it had no
report mode: a new check would have blocked from its first night. It now
follows the row-order guard's shape. A wrong row or a surviving mutant prints
a REPORT line and exits 0. RELEASE_DRAFT_GATED_ENFORCE=1 makes it exit 1.
ENV stays rc 2 in both modes. The contract's falsification test runs it
under ENFORCE, so the contract still fails on a broken table. The header no
longer says guard_tree skips it.

Measured on this tree: rc 0 by default and under ENFORCE. The create-public
mutant on a copy of the autopilot gives 8 WRONG rows, rc 0 by default and
rc 1 under ENFORCE. A missing autopilot is rc 2.

Refs #4690

Agent: aprender-78
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e guard pins the target dirs

Review of the first commit found three things the parallel lanes changed:

- The remote leg of models is an ssh with no pty. Killing the local lane left the
  remote build and ladder running in the release dir the next pass reuses. A red
  now stops only deep and dogfood, which run only here; models runs to its own
  verdict, and every red lane is named in the STOP line.
- dogfood's C14 parity and the model ladder now run on the same GPU at once. The
  ladder takes the fleet GPU lock per apr call; parity did not. Each parity run
  now takes the same lock (a lock not free in time is exit 75, read as FAIL).
- The guard stubbed the lanes, so deleting the per-lane target dirs stayed green.
  It now checks that deep and dogfood build into their own dirs and that models
  keeps CARGO_TARGET_DIR, with two mutants for it.

The lanes no longer share the autopilot's process group, so INT or TERM to the
autopilot now stops every lane (a trap held for the join). The guard is now 7 cases
and 9 mutants: no-die, no-kill, stop-models, no-rows, serial, first-red, no-trap,
shared-deep, shared-dogfood.

Agent: aprender-wprodmodels
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…se PR moved)

Agent: aprender-wprodmodels
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The receipt and findings for 951a030: the draft guard green at head under
enforce, the publish gate's self-test green, and the guard's create-public
mutant turning 8 rows wrong. The cross-vendor reviewer could not be consulted
(the budget probe exited 126), so the verdict is DEGRADED, not PASS.

Refs #4690

Agent: aprender-78
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-wprodmodels
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…r STOPPED, parity lock is pinned

From the C129 review of 19af0aa:
- Two lanes red in the same second: the second was recorded STOPPED and left
  out of the STOP line. STOPPED now needs the lane to have died of the TERM we
  sent (exit 128 + TERM, derived with kill -l); any other non-zero exit is RED.
  New guard case self-red, and mutant stopped-any-rc.
- Parity's GPU lock was untested (a flock -> env mutant survived). The guard
  now checks that every apr parity in check_model_parity.sh takes the ladder's
  lock, and mutant parity-unlocked is killed.
- A lock not free in time (exit 75) now writes its reason, so the FAIL names
  the lock instead of an empty reason.
- Parity runs under choom 1000, as the ladder's apr calls do.
- The comment now says INT/TERM stops every lane, models included.

Guard: 8 cases + target dirs + parity lock, 11 mutants killed.

Agent: aprender-wprodmodels
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…se PR moved)

Agent: aprender-wprodmodels
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-wprodmodels
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@noahgift

noahgift commented Oct 6, 2026

Copy link
Copy Markdown
Contributor Author

Review of #4887 at d4ad3f2: NO. There is one blocking finding. The release-path change itself reads correct.

Finding: scripts/check_release_t1_lanes_joined.sh blocks every merge as soon as it lands.

  • scripts/guard_tree.sh runs every scripts/check_*.sh in the merge-path guard step. guard_tree.sh --dry-run at this head prints run: scripts/check_release_t1_lanes_joined.sh.
  • The guard has no report mode. A failed case or a surviving mutant exits 1 (lines 133 and 158). So this PR wires a new blocking check, where a new check should report first and block only after three green nights.
  • Its overlap case asserts wall-clock time: three 2 s lanes must finish in under 5 s. Its red case asserts that a lane was STOPPED "well before it would have ended". On a loaded runner both can go red with no defect, which on the merge path is a flake.
  • Fix, as R11 (#4690): no public GitHub release before its assets, clean-room and preflight #4883 did for check_release_draft_gated.sh: report by default (a failure prints a REPORT line and exits 0), exit 1 only under an ..._ENFORCE=1 env, rc 2 stays rc 2, and the contract's falsification test runs it under ENFORCE. Say so in the header.

What I checked and found sound

  • The join: a lane is GO only on rc 0. STOPPED applies only to a lane this script signalled, and only when its rc is 143. Anything else is RED, and any RED dies before readiness. A wait -n that returns no lane dies. The end state is the same as running in series: no tag unless all three lanes are GO.
  • A lane's die exits only its own subshell. The join sees that as RED. Nothing a lane sets is read after the join except CARGO_TARGET_DIR, which keeps the parent's value, so readiness still reads the binary the models lane built.
  • run_step is pure, and the autopilot had no INT/TERM trap of its own, so trap - INT TERM drops nothing.
  • Dogfood's parity check (check_model_parity.sh) takes the same GPU lock as the model ladder (same default path, same env var, choom -n 1000). A lock timeout is exit 75, a FAIL that names the lock.
  • The guard at this head: rc 0, PASS 8 case(s) + target dirs + parity lock, 11 mutant(s) killed.

Non-blocking

  • Three release builds now run at once: deep, dogfood with its own C14 cuda build, and models with cuda. Peak memory and disk on the release host rise. A red from contention fails closed, but check free space on the build volume before the first parallel pass.
  • After a red in deep or dogfood, the pass still waits for the models lane to finish before it stops. That is by design, and it is in the comment above the join.

noahgift and others added 4 commits October 6, 2026 20:32
…lease guard

guard_tree runs every scripts/check_*.sh, and this one ran bare. Its red, overlap
and interrupt cases read the wall clock. A new check blocks only after three
green nights (L31), so it now prints a REPORT line and exits 0 by default.
RELEASE_T1_LANES_ENFORCE=1 makes it exit 1, and ENV stays rc 2.

Measured: on the serial autopilot, report mode gives rc 0 with REPORT and
ENFORCE=1 gives rc 1. On the real tree, ENFORCE=1 passes (8 cases, 11 mutants).

Agent: aprender-wprodmodels
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-wprodmodels
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…draft-release guard does

From the C129 review of 4f55ddf: a failed baseline case called finish at once,
so report mode printed no mutant lines. The baseline result is now held and the
verdict comes after the mutants. The success path ends in finish 0.

Measured on the serial autopilot: report mode gives rc 0 with 13 mutant lines,
and ENFORCE=1 gives rc 1. On the real tree, ENFORCE=1 passes.

Agent: aprender-wprodmodels
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-wprodmodels
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@noahgift

noahgift commented Oct 6, 2026

Copy link
Copy Markdown
Contributor Author

Re-review of #4887 at 472096b: YES. My earlier finding is closed.

  • scripts/check_release_t1_lanes_joined.sh now reports by default. A failed case or a surviving mutant prints FAIL lines and a REPORT line and exits 0. RELEASE_T1_LANES_ENFORCE=1 exits 1, and an environment problem exits 2 in both modes. Since 04fbad0, report mode runs every mutant before it gives a verdict.
  • Measured at this head:
    • the pre-PR serial autopilot: report mode rc 0, enforce mode rc 1;
    • this tree under enforce: rc 0, PASS, 8 cases, 11/11 mutants killed;
    • a missing autopilot: rc 2.
  • The guard runs bare in the merge-path guard step and nothing else references it, so it adds no new blocking check.
  • 472096b adds only the review receipt on top of 04fbad0.

Non-author sign-off round, bound to 472096b: AGREED 2/2 (claude-sonnet-5-5, gemini-3.1-pro-high). Both lanes said YES to the join block as written and NO to a planted decoy that records a red lane and continues to readiness. Both cited scripts/release/autopilot.sh:209.

@noahgift

noahgift commented Oct 6, 2026

Copy link
Copy Markdown
Contributor Author

Second non-author review, head 472096b: YES

Scope: 637d4aca05..472096b649, covering the T-1 join in scripts/release/autopilot.sh, the parity GPU lock in scripts/check_model_parity.sh, and the new guard scripts/check_release_t1_lanes_joined.sh.

Measured on a detached checkout of 472096b:

  • bash scripts/check_release_t1_lanes_joined.sh exits 0. All 8 cases hold and 11/11 of the guard's own mutants are killed. Wall time is 110 s.

  • RELEASE_T1_LANES_ENFORCE=1 bash scripts/check_release_t1_lanes_joined.sh exits 0 (PASS).

  • I planted two mutants that are not in the guard's list and passed them as $1 under RELEASE_T1_LANES_ENFORCE=1. Both exit 1:

    • Delete set -m. The lanes then share the autopilot's process group, so kill -- -pid reaches nothing. The guard catches it in the red case: deep runs 20 s to GO instead of STOPPED.
    • Count models as GO whatever its rc. The two-reds case catches it (models=GO).
  • bashrs lint, never-worse against the base:

    file errors base → head warnings base → head
    autopilot.sh 1 → 1 186 → 206
    check_model_parity.sh 0 → 0 31 → 34
    the new guard 0 errors —
  • scripts/check_sourced_libs_option_neutral.sh exits 0.

Read and checked:

  • Readiness still reads the models binary. The lanes export CARGO_TARGET_DIR inside their own subshells, so the parent keeps $REPO_ROOT/target. Readiness resolves $TD/release/apr from the parent and checks it against $MC, so the binary it reads is the one models built.
  • The join. wait -n -p takes every lane. A missing pid dies. Any non-zero rc that is not a TERM we sent counts as RED. A RED stops deep and dogfood by process group and leaves models running to its own verdict. After the join, any RED dies before readiness. INT and TERM stop every lane and die. This fails closed.
  • The parity lock. It uses the same lock file, the same wait variable and the same -E 75 as model_ladder.sh's apr_locked. A lock timeout writes an error line naming the lock and stays a FAIL. A missing flock or choom gives a non-zero prc and counts as a refusal, never a pass.

Advisory (asserted, not blocking):

  1. guard_tree adds about 110 s per run. The red, overlap and interrupt cases read the wall clock, so a heavily loaded runner could miss a timing bound. Report mode covers this for now. Before enforcing, watch for flaky results over the three nights.
  2. During T-1, three cargo builds and the model ladder now share one host. Only the apr parity and ladder calls run under choom. If the OOM killer hits a build, that lane goes RED, which fails closed but costs a pass. This is worth recording in the first pass's step table.

Base automatically changed from build-kaizen/0702-r11-draft-release to main October 6, 2026 20:01
…ain)

Agent: aprender-wprodmodels
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

§13.11 rung 1 — quorum shadow verdict

S13-SHADOW pr=4887 head=46722719a536bfa99ee1227cb40c4f9bbc796f2c verdict=REFUSE class=Q1 arm_rc=1

Shadow mode: this records a verdict and merges nothing. A refusal
to arm is not a block (§13 adds zero rows to §7) — the pull request is
exactly as green as it was.

Agent: aprender-wprodmodels
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@noahgift

noahgift commented Oct 6, 2026

Copy link
Copy Markdown
Contributor Author

#4887: my YES and the non-author sign-off round (AGREED 2/2) from 472096b now bind to a0e49c2.

  • Against main, this PR still changes only scripts/release/autopilot.sh, scripts/check_model_parity.sh, scripts/check_release_t1_lanes_joined.sh and the review evidence. All three scripts are byte-identical to 472096b, so autopilot.sh:209, the line the round judged, is unchanged.
  • a0e49c2 contains main, including the new dogfood.sh row order, which the dogfood lane runs as a whole. Re-measured on the merged tree: RELEASE_T1_LANES_ENFORCE=1 bash scripts/check_release_t1_lanes_joined.sh exits 0 with PASS, 8 cases and 11/11 mutants killed. The guard still reports by default in the merge-path guard step.

noahgift and others added 2 commits October 7, 2026 00:45
…02, SC2242)

x86-main guard-cargo was red on #4887 for two real findings in the scripts this PR adds:

- m16 (bashrs SEC/DET/IDEM): `date +%s` timers at autopilot.sh:188/196 and
  check_release_t1_lanes_joined.sh:71/76 are DET002. Durations now come from bash's
  $SECONDS, as dogfood.sh already does. The t1-steps.tsv start and end columns are
  seconds after the T-1 launch; STATUS's "T-1 LANES started together" line carries
  the wall-clock anchor. No disable-line, no baseline change.
- m12 (bashrs error ratchet, 6 -> 7): the pinned bashrs 7.4.1 read `continue` inside
  the stop-models mutant's sed string as a loop keyword (SC2242). The regex now spells
  it cont[i]nue, which matches the same text; the mutant still applies and is killed.

Measured on this tree: check_bashrs_gate.sh PASS (0 SEC/DET/IDEM errors over 517
files), check_shell_lint_ratchet.sh PASS (6, baseline 6), the T-1 guard passes in
report and enforce mode (8 cases, 11/11 mutants killed), and the other autopilot
guards are unchanged.

Agent: aprender-wprodmodels
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-wprodmodels
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@noahgift

noahgift commented Oct 6, 2026

Copy link
Copy Markdown
Contributor Author

Re-review of #4887 at 4672271: YES.

  • 1b6603e changes no verdict logic. The join's timers and the guard's case timer read $SECONDS instead of date +%s, and t1-steps.tsv start/end are now seconds after the T-1 launch. Nothing reads that file except the guard, and the guard reads only its verdict column. The stop-models mutant's sed spells cont[i]nue, which matches the same line.
  • The die after the join is unchanged and now sits at autopilot.sh:210.
  • Measured at this head: the guard exits 0 in default mode, and under RELEASE_T1_LANES_ENFORCE=1 it passes 8 cases with 11/11 mutants killed. The pre-PR serial autopilot still exits 1 under enforce. bashrs reports one error, SC1086 inside a message string, which main already has (line 390 there). It is not from this PR.

Fresh non-author sign-off round, bound to 4672271 (the judged block changed, so this is a new round, not a re-bind): AGREED 2/2 (claude-sonnet-5-5, gemini-3.1-pro-high). Both lanes said YES to the join block as written and NO to a planted decoy that records a red lane and continues to readiness. Both cited scripts/release/autopilot.sh:210.

@noahgift

noahgift commented Oct 6, 2026

Copy link
Copy Markdown
Contributor Author

Re-review, head 4672271 (fix 1b6603e): YES

My review covers a0e49c24a1..1b6603ee5c, a 7-line change. The lane timers and the guard's case timers now read $SECONDS instead of date +%s. The tsv start and end columns are now seconds after the launch. The stop-models mutant pattern now spells cont[i]nue, so the shell lint no longer reads it as a continue.

What I measured on a detached checkout of 4672271:

  • RELEASE_T1_LANES_ENFORCE=1 bash scripts/check_release_t1_lanes_joined.sh exits 0. All 8 cases pass and the guard kills all 11 of its own mutants. stop-models still applies and is still killed.
  • I planted one mutant of my own: I removed set -m from the new launch line and passed that file as $1 under ENFORCE. The guard exits 1, so it is still caught.
  • bashrs lint scripts/release/autopilot.sh reports 1 error, the same count as the base, and none of them are DET002 or SC2242. The new guard reports 0 errors.
  • I searched for anything else that reads t1-steps.tsv and found nothing outside these two files, so the column change to relative seconds has no other consumer.

The ordering and verdict logic of the join is unchanged from 472096b, so my earlier reasoning still holds. The two advisories from my first review still apply.

@noahgift
noahgift added this pull request to the merge queue Oct 7, 2026
Merged via the queue into main with commit 4098007 Oct 7, 2026
20 of 21 checks passed
@noahgift
noahgift deleted the feat/c316-2d-autopilot-parallel branch October 7, 2026 02:19
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant