fold(pr-review-ci): #4503 + #4533 + #4517 — diff classifier/docs tier, FLOW-003 v2.1, fork attest - #4605
Closed
noahgift wants to merge 43 commits into
Closed
fold(pr-review-ci): #4503 + #4533 + #4517 — diff classifier/docs tier, FLOW-003 v2.1, fork attest#4605noahgift wants to merge 43 commits into
noahgift wants to merge 43 commits into
Conversation
…4415) On the 0.69.4 RC PR #4318, every red guard job reported exactly one failure. GitHub Actions skips a job's later steps once one fails, so guard-cargo skipped 15-71 of its 77 steps (median 55). Each further red was found only after another 35-50 min CI cycle: 02d4618 -> cef1b12 -> 696576e -> 6281d9f. - ci.yml: the last setup step of guard-cargo and guard-tree gets `id: guard-setup`. Every guard step after it gets `if: ${{ !cancelled() && steps.guard-setup.outcome == 'success' }}`. All guards run and the job stays red if any fails. If setup failed, the guards skip instead of painting dozens of meaningless red rows. The diff is if/id keys only; that was verified by a YAML-level comparison of every job. - scripts/lib/ci_guard_steps.py: reads a guard job's steps from ci.yml. `check-run-all` counts fail-fast guard steps, `run` executes the same run: blocks locally without stopping at the first red (the #4416 v0 seed), and `list` prints the steps. - scripts/check_guard_steps_run_all.sh: a shrink-only ratchet at 0 (scripts/guard_fail_fast_baseline.txt), so a new guard step cannot be fail-fast. 11-row case table; 5/5 mutants killed. It is wired through guard_tree.sh --no-cargo. - scripts/ci_guards_local.sh: the local entry point (#4416, handed to aprender-cb). Its first local run caught a missing tool_version header in this PR's own baseline file. Operator ruling Y1 (relayed by the cop, 2026-09-25): guards use !cancelled() and the job stays red. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… script, CI and local (#4415, #4416) GitHub stops a job at its first red step. On RC PR #4318 every red guard job reported ONE failure; guard-cargo skipped 15-71 of its 77 steps. Operator amendment A4: guard-cargo and guard-tree each run ONE step, `bash scripts/ci_guards.sh <job>`, and `make guards-local` + the pre-push hook run the same file. Both print `ci_guards: sha256 <runner> manifest <steps>`, so "CI and local ran the same guards" is a line comparison. - The guard steps stay VERBATIM in ci.yml, moved into `guard-cargo-steps` / `guard-tree-steps` manifest jobs (`if: false`, never run by GitHub), so every guard that greps ci.yml for an invocation still finds it. Transform verified by YAML equality: moved steps == originals minus `if:`, every other job and the top level unchanged. - scripts/lib/ci_guard_steps.py runs the manifest: bash -eo pipefail per step, GITHUB_PATH/GITHUB_ENV propagated between steps, github.token / runner.temp resolved (anything else refused), per-step timeout with a process-group kill, ::error per failure, step-summary table, exit 2 on a vacuous run. - check_guard_steps_run_all.sh: manifest must be `if: false`, steps are name/run/env only, only resolvable ${{ }}, the job must call the runner and have `id: guard-setup` (else the runner's if: skips it and the job is green having run nothing), no orphan manifest. 25/25 rows; 10/10 non-equivalent mutants of the runner killed. - guard_tree_test / guard_tree_job_test read the manifest too; new mutant (cargo planted in guard-tree-steps) is RED under the new leg and was GREEN under the old one. - Absorbed from aprender-cb's add2fd4 (#4416): check-coverage + ack file (2 rows stale on main, removed), Makefile guards-local, pre-push block. scripts/ci_guards_local.sh -> scripts/ci_guards.sh. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
#4415) Quorum lane b on 15883bb: - usage() printed lines 2-31, so --help never said "self-test", and guard_tree.sh's advertises_self_test never dispatched the case table. usage() now prints the whole header, up to `set -uo pipefail`. Verified: the old file is NO, the new one YES. - ci_guards.sh --check-coverage ran only in make guards-local and the pre-push hook. The main path now also runs check-coverage, so a guard step outside a manifest turns guard-tree red. Mutant: removing one ack row gives rc=1. - A new self-test row: a literal ${{ }} inside run: turns it RED (26/26). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…in check, rc 2 in CI run (#4415) Lane b (sonnet-5) reproduced it: with the guard-tree-steps job deleted, the leftover steps behind the runner counted as 'ran', so ran==0 never tripped. There is a new fixture variant, no_manifest, and a mutant dropping the check turns it RED. 27/27. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ader refuses what it cannot read (#4415) The self-test now calls scripts/lib/ci_guard_steps.sh (not yet written) and adds 14 reader rows: 10 refusals (anchor, alias, tag, folded, multi-line plain, unclosed quote, duplicate key, tab indent, flow map, colon in a plain scalar), each beside a passing control, plus quoted-scalar decoding, a flow-list gate needs:, and no python on the lib or runner path. Pmat-Ticket: PMAT-4415 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…leaves the runner path (#4415) scripts/lib/ci_guard_steps.py -> scripts/lib/ci_guard_steps.sh, with the YAML reader in ci_guard_yaml.awk and the jq defs in ci_guard_steps.jq (both hashed into the `ci_guards: sha256` line). yq is not on the clean-room fleet, so the reader is awk: it reads the block-YAML subset ci.yml is written in and REFUSES (rc 2) anchors, aliases, tags, folded >, multi-line scalars, flow maps, duplicate keys and tab indentation. Measured on this tree: - case table 41/41 (was 1 ok / 40 FAIL at RED), under gawk AND mawk - reader vs PyYAML on ci.yml .jobs: byte-identical JSON (18 jobs), gawk + mawk - list / check-coverage / check-run-all: byte-identical to the .py at 3fb4514 - 3 reader mutants (dup-key, backslash escape, tab) each turn a row RED - bashrs lint: 0 errors (warnings are jq/awk text inside strings) ci.yml is unchanged: it never installed PyYAML and never named the .py. Pmat-Ticket: PMAT-4415 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…in an agy lane sandbox (#4415) The self-test killed a hung step and then asked `kill -0 child` 0.2 s later. kill -0 answers for a zombie, so a killed child that its reaper has not yet collected read as "still alive", and the row went RED in a quorum lane's sandbox while passing 5/5 here. Read the child's state with ps instead (gone or Z = dead) and give the reaper up to 2 s. The mutant with a really escaped child (setsid sleep) still goes RED; table 41/41. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… rc 0 The run loop read jq's step records through <(...), whose exit status nothing checks: a jq error after the first record silently dropped every later step and the run still reported success. Records now go to a scratch file first, and a nonzero jq stops the run with rc 2. Case row 42 injects a jq error on step 2; it is RED on the old form and GREEN now (42/42). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ner-attest)
A fork PR's own runs get no secrets, so it could never carry a receipt signed
with PR_REVIEW_SIGNING_KEY_B64 and was unmergeable by construction.
pr-review-fork-attest.yml (pull_request_target, labeled) authorizes the
labeler server-side (admin/maintain/write; a triage labeler reads as `read`),
refuses a self-label and a same-repo head, reads the head as git objects only,
builds an L2-maintainer-attest receipt (verdict DEGRADED, no consultations),
signs it with the base secret, and publishes it fast-forward to the
base-owned branch pr-review-fork-receipts. Arm 4 reads L2 receipts only from
that branch (PR_REVIEW_ATTEST_ROOT), prints "DEGRADED: maintainer-attest by
<login>", and the signed diff patch-id voids the attest on any later push.
RED proofs (cop conditions): author self-label, triage-only labeler, stale
patch-id, receipt on a non-base branch — each a case-table row that goes RED,
each mutant verified killed. Guard fix found on the way: jq `inside()` on a
string is substring containment ("" and "rite" passed); now IN(...).
Refs #4462
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…4472) Docs-only PRs paid full CI and a full quorum (evidence: #4467, README + two book pages). This is PR 1 of 2 and touches no workflow file. - scripts/ci/diff_class.sh: ONE classifier (class=docs|code|empty, subclass=ledger|prose). Docs = root *.md, book/**/*.md, docs/**, each present at head; a deleted doc, crate README, .claude/, contracts/, evidence/ and any config is code. --no-renames. Self-test 18/18, all 5 mutants RED. - scripts/ci_test_tier.sh: docs_only() now asks the classifier (subclass=ledger, same #3658 set as before); stale self-test row repointed at ci/sections.yml, where #4471 moved the touched-list line. - check_pr_review_receipt.sh S3.E docs tier: antigravity may be not-triggered only when class=docs AND docs/BEATS.md untouched AND no added line states a comparative ratio AND trigger_reason names the docs tier. The classifier is looked up lazily and fails closed. - Fixtures: row 34 moved to a code diff; rows 44-47 new (SB1 head added after K1, no existing SHA moved). reject-53-drop survived until row 47 existed. Mutation set 233 -> 241; all 10 new mutants killed. bats 170/170. - Contract F-PRREV-016 + formula, skill S3.E, spec count amended. PR 2 (workflow gating of mutants/determinism/sov on class=docs) follows a2's #4482. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ls closed agy on 253d2e4: docs/specifications/pp-066-dag.yaml is a CI input, so docs/** now means docs/**/*.md (plus the #3658 ledger dirs). The ratio scan read the added lines through < <(...), so a failed git diff passed; it now captures them first and rejects B1. The new probe (blob removed) goes RED. Mutation set 243. WIP: mutants reject-51..56 not yet re-run (stopped for the operator budget hold). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…red killed alone The derived set (mutate-guard.sh --list) is 239. Only the six added by the L2-attest guard (reject-01..03, drop+flip) were run, each alone against the full bats oracle: 6/6 killed. The spec's PENDING row says exactly that and keeps the full-sweep kill count pending on Arm 3. Refs #4462 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ing, #4517) The attest-root loop matched on patch-id only. A receipt signed for PR A, copied under PR B's directory with an identical diff, would arm B. The signed .predicate.pr must now equal PR_NUMBER. Self-test row added; mutant (check removed) turns it RED. Refs #4462 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… false positive) Refs #4462 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
# Conflicts: # .github/workflows/pr-review-quorum.yml
…ng v1.1 fixes, a Parts II–III review (R2-1..R2-14), Prop. 19 and pv contracts v2.0 said "infra-5a applies its remaining v1.1 fixes on top of this version". This commit applies them: R-3..R-12 on Part I. It also reviews Parts II and III. Two soundness findings are among them: the input set goes stale between the nightly and the base (R2-8), and an openat trace misses directory listings and ENOENT lookups (R2-9). Doctests drop out of nextest archives (R2-10), and Prop. 18 measured only what full review missed (R2-12). New: Prop. 19 (tier routing lowers rho_HOL), §11 (ci-input-set-v1 written out in full), and §12 (the review record). R2-14: the QM-08 number collides with aprender#4519, flagged for the cop to re-map (not renumbered here). Refs #4514 (PMAT-4514) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…-4514 to v2.1 Quorum lane 2 found that v2.0 cut QM-00's mutant list from v1.1's four to three with no record: a gate weakened silently. Restored as the union (the 1/(1-phi) retry mutant, the Prop. 10 rescued-flake mutant, and v2.0's double-count mutant) plus one for Prop. 20. The ont:falsifier minCount goes from 3 to 6. Recorded as R2-15 in §12.2. Quorum lane 3 found that PMAT-4514's title still says v1.1 only. The cop put v2.1 under PMAT-4514, so the title in the fragment and the aggregate now names both. Refs #4514 (PMAT-4514) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ve-diff guard (PMAT-980) forbids a title edit; the v2.1 scope rides in the PR body and issue #4514 Refs PMAT-4514 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Resolution table: | path | conflict | resolution | |-----------------------------|--------------------------------------------|------------| | docs/roadmaps/roadmap.yaml | both sides appended an entry at the tail: ours PMAT-4472, main's PMAT-4514 | union: both entries kept verbatim, PMAT-4472 then PMAT-4514; YAML parses | | ci/sections.yml | auto-merged | no manual edit | Checked on the merged tree: - diff_class.sh --self-test: 20/20 - ci_test_tier.sh --self-test: 93 ok, rc=0 - check_pr_review_wiring.sh: rc=0 - bats tests/pr-review.bats: 171 ok, rc=0 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-59 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-59 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Sep 28, 2026
|
§13.11 rung 1 — quorum shadow verdict Shadow mode: this records a verdict and merges nothing. A refusal |
…r onto the sections shape Main moved every CI job body into ci/sections.yml (#4441) and ci.yml's gate now requires fat jobs, not guard-*. A text merge could not carry #4457 across: - ci.yml is main's, unchanged. #4457's guard-tree / guard-cargo edits (guard-setup, one `bash scripts/ci_guards.sh <job>` step, the `<job>-steps` manifests) are applied to ci/sections.yml, with main's `setsid --wait` ported onto the moved steps. All guard-setup / ci_guards.sh lines are present (1076-1094, 1663-1702). - scripts/lib/ci_guard_steps.sh reads ci/sections.yml (its `jobs:` mapping only; the header carries flow maps the reader refuses). With no gate job in the file, guard jobs are the guard-* jobs that are not `-steps` manifests; a ci.yml-shaped file with a gate still uses the gate's needs. New self-test row for the no-gate case; repo fixtures follow the path (43/43). - guard_tree_job_test.sh: main's CI_YML+SECT_YML test plus #4457's legs; the no-cargo leg reads guard-tree-steps too, and mutant 2b plants cargo in the manifest (8/8, yq and awk modes). - ci_guards_uncovered.txt: two `gate` rows removed as stale on main (the check says so; check_receipt_gate_base_owned.sh now runs inside guard-tree-steps). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- pr-review-quorum.yml: union. The fold's PR_REVIEW_ATTEST_ROOT (#4517) and main's B1 ARM4_PR_HEAD_SHA + GH_TOKEN both feed check_pr_review_arm4.sh. - roadmap.yaml: main now aggregates from entries/, so #4503's inline PMAT-4472 entry moves verbatim to docs/roadmaps/entries/PMAT-4472.yaml, and roadmap.yaml is regenerated (aggregate --check: ok, idempotent). Agent: aprender-59 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… 5 s (FLAKE-0) (#4613) assertion_failure_reports_nonzero_and_traceback gave exit_code None under a full `cargo test -p apr-cli --tests` on intel (2/2 runs) and passed alone (3/3): `assert 1 == 2` was killed by the 5 s deadline while python3 was still starting. The deadline in these tests guards a hang; a program that exits returns at once, so the terminating-program tests now share TERMINATING_DEADLINE_SECS = 120, and the success/exit-code asserts print timed_out so the next failure names its cause. Agent: aprender-a2 Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…ALE (PMAT-4594) (#4598) * fix(silicon): a short listing that agrees with itself scored two green axes STALE Run 36410817478 (2026-09-28 11:25Z, #4512) failed check_silicon_coverage.sh with x86_64-cpu and aarch64-cuda-sm121 "STALE 9d" while Silicon Nightly 36364433328 had carried both 10 hours earlier. The created>= listing of silicon-nightly.yml answered "21 of 21": total_count agreed with the page, and the newest 9 runs were not on it. The runs before and after that one all PASS, and the same guard run live now reads 30 of 30. The reader cross-checked only an EMPTY listing against the unfiltered newest run (PMAT-3337 §8). A short one went straight to the verdict. Now every listing is cross-checked: if the unfiltered newest run is inside the lookback and newer than anything the filtered listing held, the listing is read once more, and if it is still short the axis is NO-GO (unknown, rc 2), never STALE. The lookback, the stale window and every limit are unchanged. Self-test rows on the real policy lines, the real silicon-nightly.yml and run 36364433328's four jobs as the API returns them: g-real-green fresh nightlies -> ok, rc 0 h-real-stale newest nightly 9d old, whole list -> STALE, rc 1 i-short-once 11:25Z shape on the first read -> ok, rc 0 j-short-always 11:25Z shape on every read -> NO-GO, rc 2 Against origin/main's reader, i and j go RED (rc 1, STALE): the defect. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Agent: aprender-57 * roadmap(PMAT-4594): entry for the silicon short-listing fix Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Agent: aprender-57 * evidence(PMAT-4594): quorum verdict (AGREED 3/3) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Agent: aprender-57 * evidence(PMAT-4594): pr-review receipt for 9aa1595 (reviewer sonnet-5, FINDINGS advisory-only, mutation 3/3) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Agent: aprender-57 * roadmap(PMAT-4594): entry arrives through its fragment (check_roadmap_fragment_required) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Agent: aprender-57 * evidence(PMAT-4594): pr-review receipt for 4c7e5c3 (reviewer sonnet-5, DEGRADED: agy lane unreachable) The fragment commit moved the diff patch-id, so the 9aa1595 receipt no longer binds. Re-reviewed at the new head by an independent sonnet-5 subagent; self-test 12/12, roadmap guards rc 0. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Agent: aprender-57 * fix(ci): pr-review-sign got an empty key — ci.yml fed FAT_SECRET_..._KEY_B from a secret that does not exist #4441 wired `FAT_SECRET_PR_REVIEW_SIGNING_KEY_B: ${{ secrets.PR_REVIEW_SIGNING_KEY_B }}`. The repo secret is PR_REVIEW_SIGNING_KEY_B64 and the pr-review-sign section reads secrets.PR_REVIEW_SIGNING_KEY_B64, so the driver always handed it "" and every same-repo PR that commits a pr-review receipt failed x86-main ("carries an UNSIGNED receipt and PR_REVIEW_SIGNING_KEY_B64 is empty"): #4598 job 108952599966, #4587 job 109019972501. Mechanism: fat_driver self-test gains secret_wiring_gaps() — every secrets.X a section (ci/sections.yml, vendored sov) reads must arrive as FAT_SECRET_X fed from secrets.X. Red/green: with the old ci.yml line the self-test exits 1 (row "secret wiring: every secrets.X ..." BAD); with this fix 52/52. Refs PMAT-4594 Agent: aprender-e6 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * chore(pr-review): sign this PR's receipt (PR-REVIEW-SKILL-002 v2 §4.3 CI signer) * evidence(PMAT-4594): pr-review receipt for a0e0e9c (reviewer sonnet-5, FINDINGS advisory-only) Fresh quorum after #4512 changed the patch-id lib (receipt rebind, cop DECIDED-UNATTENDED). diff_patch_id c860ecfc40 under the in-tree lib. Second vendor agy gemini-3.1-pro-high consulted. Agent: aprender-e6 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * evidence(PMAT-4594): §3.C CRUX consulted on the a0e0e9c receipt Arm 4 rejected consultations.crux=not-triggered: the S3.C surface regex matches 'ToolDefinition' inside a prior receipt's prose in evidence/pr-review/4598/ (false positive; no CLI/HTTP/MCP/config surface changes). Consulted and recorded; patch-id unchanged (c860ecfc). Agent: aprender-e6 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * evidence(PMAT-4594): a0e0e9c receipt — SARIF driver names in the closed vocabulary Arm 4 A4 rejected the receipt [B1]: runs named pr-review-mutation / pr-review-reviewer / pr-review-crux are outside { pmat, nvidia-cuda-docs, crux, cargo-mutants, antigravity }. Renamed per the prior 9aa1595 receipt's convention: primary-reviewer results under pmat, mutation -> cargo-mutants, crux -> crux, plus the (empty) antigravity run that consultation records. The two results gain grounding=asserted / precision_class=advisory (as their own text states) and a failure_scenario. findings_ref.sha256 rebound. No verdict, finding or consultation changed. check_pr_review_receipt.sh ACCEPTs it (verified under a throwaway key; the CI signer signs for real). Agent: aprender-e6 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * evidence(PMAT-4594): commit the CI signer's signature for the a0e0e9c receipt The B1 check-run route cannot go green on a rerun: the signature check-run lands in present's own check suite, so a re-attempt never reads it (inbox 21:09Z). Per COP 21:10Z 4598-ARM, arm via the committed route. The .minisig is verbatim from pr-review-signature check-run 109135289610 (CI signer, repository key). minisign -V under .github/pr-review.pub is OK, and check_pr_review_arm4.sh passes locally (A2 binds patch-id c860ecfc, A3 positive control fired, A4 ACCEPT). Agent: aprender-e6 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…context_length (#4599) (#4601) * fix(code): apr code -p never silently drops an over-budget prompt (#4599) The sliding window skipped the newest message whole when it did not fit, and apr code hard-coded a 32K window, so a >~105 KB prompt reached the model as an empty conversation: rc 0, tokens_in 19 (#3715 pilot, 8 cells). - serve/context.rs: the newest message is never dropped; when it cannot fit on its own the window is refused (ExceedsLimit -> context_overflow, nonzero). - agent/code.rs: the window is the model's GGUF context_length (bounded header read), an explicit manifest value still wins, 32K only as fallback. - Falsifiers: the 8 failing cells fit at 262144 and are refused (never dropped) at 32K; a >105 KB prompt is refused at 32K; the refusal is named and nonzero; the window comes from the model. Refs #4599 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Agent: aprender-3d * test(#4599): over-budget boundary at 128 KiB; tiny-window test gets an input budget Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Agent: aprender-3d * test(#4599): MockDriver default window 32768, above the default 4096 output reserve At 4096 every mock-driven agent test had zero input budget; the old sliding window silently sent the model an empty conversation, which is #4599 itself. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Agent: aprender-3d * test(#4599): retry-test drivers get a window above the output reserve Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Agent: aprender-3d * test(#4599): 001 asserts the whole prompt, tail included, reached the model Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Agent: aprender-3d * fix(#4599): a tool result never evicts its turn's prompt; one tested driver-window resolver Quorum (sonnet-5) on #4601: in a tool loop the sliding window kept the newest tool result and could drop the current prompt, rc 0. truncate_messages now refuses (ContextOverflow) when the last User message would be evicted; older turns may still be. Both drivers launch through code_driver_window, pinned by FALSIFY-4599-005. New: FALSIFY-4599-006/007. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Agent: aprender-3d * test(#4599): 007 asserts the current turn is whole and the old prompt evicted Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Agent: aprender-3d * test(#4599): echoed input token count == expected; planted 1-byte truncation is RED (FALSIFY-4599-008) Operator R2: 212 KB and 510 KB prompts must echo the exact input token count. The recording driver echoes a byte-level count (the 4 B/token estimate cannot see one lost byte); cell_verdict checks count, needle and full content. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Agent: aprender-3d * test(#4599): agent_integration drivers had zero input budget under the 4096 reserve test_routing_driver_fallback_integration's FailPrimary reported a 4096 window and test_context_truncation_integration a 300 window, both at or under the default 4096 output reserve. Before #4599 the runtime sent them an empty conversation; it now refuses with context_overflow. Raise FailPrimary to 32768 and give the tiny-window test max_tokens 64, as the lib tests already do. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Agent: aprender-3d * test(#4599): kill the 4 missed and 2 timed-out mutants on the diff - context.rs `>`→`>=` on the newest-message check: a newest message that exactly fills the window must be kept (boundary case added). - model_context_length: one fn with cfg blocks, so the whole-fn mutant lands on the compiled body; a 0 context_length falls through to 32K (case added). - build_default_manifest: drop the explicit `context_window: None` (the default); deleting it was an equivalent mutant. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Agent: aprender-3d * ci(#4601): re-run on a body that carries keep-open for #3715 guard-tree reads the PR body from the event payload, so a rerun of run 36454753912 re-reads the old body and cannot pass. The body now has a "keep-open: #3715 <reason>" line; this empty commit makes a new synchronize event that carries it. No code change. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Agent: aprender-3d * review(#4601): pr-review receipt at e3d9361 — DEGRADED (mutation not run on lambda by rule), no blocking class Reviewer: independent Sonnet 5 session; agy gemini-3.1-pro-high advisory: pre-existing orphan tool_use eviction in truncate_sliding_window, not introduced here. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Agent: aprender-3d * review(#4601): stamp predicate.diff_patch_id into the receipt so Arm 4 A2 can bind it (#4421) Computed with scripts/lib/pr_review_patch_id.sh (git-patch-id-verbatim/pinned-diff-v1) over base_sha..head_sha; the CI signer recomputes and refuses a mismatch. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Agent: aprender-3d * review(#4601): commit the CI signer's .minisig so Arm 4 verifies it unchanged (cop ruling 21:23Z, same route as #4598) Taken verbatim from the pr-review-signature check run on b87d6b7; minisign -V against .github/pr-review.pub rc=0; check_pr_review_receipt.sh ACCEPT. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Agent: aprender-3d --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…rds its backend (#4609) (#4612) * fix(apr chat): the Qwen3.5 GPU turn records the backend that answered (#4609) The qwen35 session branch returned without setting generated_on_gpu; the moe and dense branches set it. So every CUDA Qwen3.5 chat turn printed backend {ran: cpu, fell_back: true} while stderr said 'Backend: GPU (CUDA ...)' and apr run on the same model said ran: gpu. With #4609's grader (fell_back => fail) every qwen35 chat cell of #3715 would be RED on a real GPU. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Refs: #4609 Agent: aprender-ec (cherry picked from commit baec88e) * fix(apr chat): a --gpu turn the CPU answered is exit 14, never a silent pass (#4609) `apr chat --gpu` whose turn ran on the CPU printed the answer, reported `fell_back: true` only in the --json epilogue, and exited 0, so a #3715 cell that asked for the GPU graded PASS. `apr run --gpu` already refuses that case (R-0b, exit 14). Chat now calls the same `registry::after_generation` per turn: it prints no answer and ends the session with BackendUnavailable (14), after printing the --json epilogue. `generated_on_gpu` is reset at the start of each turn, so a CPU turn cannot inherit an earlier GPU turn's flag. A default request (no --gpu) is unchanged: the epilogue still says fell_back. Tests: pmat4609_forced_gpu_turn_check covers both polarities (--gpu+CPU → 14; --gpu+GPU, no flag, --no-gpu → Ok). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Refs: #4609 Agent: aprender-3d * review(#4612): pr-review receipt at c61482f — DEGRADED (mutation unreachable: include!() files), no blocking class Reviewer: independent Sonnet 5 session; agy gemini-3.1-pro-high advisory, no findings. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Agent: aprender-3d * review(#4612): stamp predicate.diff_patch_id into the receipt so Arm 4 A2 can bind it (#4421) Computed with scripts/lib/pr_review_patch_id.sh (git-patch-id-verbatim/pinned-diff-v1) over base_sha..head_sha; the CI signer recomputes and refuses a mismatch. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Agent: aprender-3d * review(#4612): SARIF driver crux-consultation -> crux (guard's allowed set); findings_ref.sha256 recomputed by the reviewer Found only once the receipt was signed: the unsigned rejection masked it. diff_patch_id unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Agent: aprender-3d * review(#4612): restore the receipt's trailing newline (the signer requires exactly one line) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Agent: aprender-3d * review(#4612): commit the CI signature as .minisig so present verifies with Arm 4 unchanged Taken verbatim from the pr-review-signature check run on 51bb553 (x86-main). minisign -V rc=0; check_pr_review_receipt ACCEPT. Route per cop 21:23Z ruling. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Agent: aprender-3d --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…uld never go green (#4618) * fix(arm4): B1 reads every attempt's signature — a rerun of present could never go green The pr-review-signature check run lands in the pr-review-quorum run's own check suite. Arm 4 read check-runs with the API default filter (latest) and took the newest run, so a re-attempt of `present` read nothing and printed UNSIGNED beside a verifying signature on the same head. #4598: sig 109130365825 success at 20:52:12Z; run 36481105318 attempt 2 UNSIGNED at ~21:04Z; the signer then posted another sig after present failed. 2/94 recent present runs passed, both via a committed .minisig. - the listing uses filter=all - the selector takes the latest successful run on the head that CARRIES this receipt's key, not the latest run Nothing loosens: the pick must still verify under the repository key and name exactly pr, head and pid. Case table 33 -> 37 rows: check-sig-prior-attempt GREEN (sig in an earlier attempt, newer unsigned + queued) check-sig-prior-wrong-pid RED check-sig-prior-wrong-head RED check-sig-prior-other-sha RED Red proof: reverting the selector alone turns check-sig-prior-attempt RED (rc=1, wanted 0); the three controls stay RED. NOT covered by the table: the filter=all URL itself, because injection bypasses the API. Agent: aprender-e6 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(arm4): prove filter=all through the API, not only through a file Review of 4c12e25 (BLOCK): dropping &filter=all from the URL left the self-test 37/37 GREEN — every prior-attempt row injected a file, so the URL was never exercised. New row check-sig-api-all-attempts drives the real fetch through a gh shim that returns the earlier attempt only when the query carries filter=all. Mutation: drop filter=all → that row RED. Also: a newer run whose output.text is valid JSON but not an object (e.g. "[]") crashed jq's .[$k] and read as UNSIGNED. Guarded; row check-sig-prior-array-text; mutation: drop the guard → that row RED. Self-test 39/39. Agent: aprender-e6 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * evidence(pr-review): non-author quorum receipt for #4618 at 78eb06a Reviewer claude-sonnet-5 (author claude-opus-5-5): PASS, 0 findings. Measured by the reviewer: self-test 39/39; drop &filter=all → only check-sig-api-all-attempts RED. pid 3aa28558bc (evidence excluded). check_pr_review_receipt.sh ACCEPT under a throwaway key. Agent: aprender-e6 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * evidence(pr-review): commit the CI signer's .minisig for #4618's receipt Verbatim from the pr-review-signature check run on ee080ab (trusted comment pr=4618 pid=3aa28558bc). minisign -V under .github/pr-review.pub passes; local Arm 4: A3 fired, A4 ACCEPT, PASS. Same route as #4598 (present ran before the CI signer posted). Agent: aprender-e6 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(counts): the Arm 4 step name states 39 rows, not 33 x86-main on 500fb1f: check_pr_review_counts.sh derived arm4_rows=39 and found "Arm 4 case table: 39 rows" 0 times in pr-review-quorum.yml. Label-only change; no gate logic touched. Local: counts run + --self-test green. Agent: aprender-e6 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * evidence(pr-review): fresh non-author receipt for #4618 at 437037f The count-label fix changed the patch-id (3aa28558bc → 253965ff05), so the 78eb06a receipt no longer binds; this is a fresh review, not a patched one. claude-sonnet-5, PASS; counts guard run + self-test green as measured by the reviewer. ACCEPT under a throwaway key. Agent: aprender-e6 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * evidence(pr-review): commit the CI signer's .minisig for #4618's 437037f receipt Verbatim from the pr-review-signature check run on 17efd38 (pid 253965ff05). minisign -V under .github/pr-review.pub passes; local Arm 4 PASS. Agent: aprender-e6 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Fragment copied byte-identical from cop-inbox/handoff/PMAT-4621.yaml; roadmap.yaml regenerated (aggregate --check ok, +22 -0). Agent: aprender-59 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(deps): wasmtime 47.0.4 -> 48.0.3 for RUSTSEC-2026-0315/0316 cargo audit on origin/main's Cargo.lock reports RUSTSEC-2026-0315 and RUSTSEC-2026-0316 against wasmtime 47.0.4 (fixed >=48.0.3, <49 or >=49.0.1), which turns the security job red on main and on every PR. A version bump only; no audit ignore is added, so no gate is relaxed. wasmtime stays optional behind aprender-test-lib's runtime feature, which nothing enables: cargo tree -i wasmtime --workspace finds no active path, so the 1.95 MSRV of the 48 line never fires on the 1.93 pin (comment updated). Receipts (intel/devont, fresh advisory DB): cargo audit rc=0 (was: 2 vulnerabilities), cargo deny check rc=0 (advisories, bans, licenses, sources ok), cargo fmt --check rc=0, cargo metadata --locked rc=0. Agent: aprender-00 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(roadmap): PMAT-4633 — wasmtime RUSTSEC-2026-0315/0316 bump Agent: aprender-00 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(audit): PMAT-4633 — review input for the wasmtime 48.0.3 bump The review lanes get the Cargo.toml diff, a lock-delta table generated from Cargo.lock (crate, old -> new, checksum), the RUSTSEC-2026-0315/0316 text, and cargo audit / cargo deny / cargo metadata --locked on base and head. The raw lockfile stays with the machine checks. cargo audit goes from rc=1 (both advisories on 47.0.4) to rc=0. cargo deny passes on both sides because wasmtime is reachable only through aprender-test-lib's optional `runtime` feature, so it does not tell them apart. Refs #4633 Agent: aprender-00 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * review(PMAT-4633): full cargo deny evidence; acceptance names the C45 review doc Q0 round on e9728ba: both Gemini lanes failed on (1) no evidence for deny bans/licenses/sources, and (2) the acceptance criterion's file list omitting the audit doc that C45 requires in place of the raw Cargo.lock. Full cargo deny check (cargo-deny 0.19.0) was run on base 2817c6d and on this head: rc=0 on both, all four checks ok. Pmat-Ticket: PMAT-4633 Agent: aprender-07 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * review(PMAT-4633): pr-review receipt for #4632 @ 7af51f3 (receipt-only) Independent reviewer agent:claude-sonnet-5-5/pr-4632-review; verdict FINDINGS (7 advisory, 0 blocking); agy arm gemini-3.8-flash-high after gpt-oss-120b-medium 503x2 (operator fallback ruling). predicate.diff_patch_id 74b701ae170949b97b6ea5f75126a5588c716438; the signature is posted by CI as the pr-review-signature check run. Operator C72 exception to C69. Pmat-Ticket: PMAT-4633 Agent: aprender-07 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…ine (PMAT-4646) (#4647) * ci(mutants): ratchet mode — report-only in ci / gate, shrink-only debt baseline (PMAT-4646) Operator rulings C187(a)/C188(a), 2026-10-01: a new gate is adopted as a ratchet, never as an instant absolute bar. No ruleset change. - ci / gate: mutation results (the mutants section and, on PRs that add it, the survivor table) are printed and no longer fail the gate. Every check before it is unchanged: workspace-test, guard-tree, guard-cargo, sov.gate, determinism-compare. - ci/mutants-debt.tsv: 590 known survivors (MissedMutant + Timeout), with file, mutant, sha and run id, from completed shard artifacts of #4587 run 36725778557 and #4588 run 36703131059. - scripts/check_mutants_debt_ratchet.sh (auto-run by guard-tree): the debt file is shrink-only against the merge base. A grown file, a swapped row, a missing file, a malformed or duplicate row, or an unresolvable base is RED. Self-test: 10 cases; 7/7 planted mutants RED. The RESTORE PR deletes the MUTANTS-RATCHET block after crates.io 0.70. Agent: aprender-ca Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(roadmap): PMAT-4646 entry (MUT-RATCHET acceptance criteria for the quorum) Refs #4646 Agent: aprender-ca Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(ratchet): base via scripts/lib/resolve_base.sh (CI depth-1 checkout); gate echo names only defined vars check_mutants_debt_ratchet.sh [run] was RED in run 36829781758: merge-base(HEAD, origin/main) is unresolvable on a depth-1 checkout. It now uses resolve_base, the base every differential guard uses, which refuses rather than judge HEAD against HEAD (new case R11, plus R12 for the default origin/main path; 12/12; a fail-open mutant of the new branch is RED). check_workflow_env_defined: the gate echo interpolated IN_SCOPE/SCOPE/TABLE, which main's gate job does not define; dropped. Refs #4646 Agent: aprender-ca Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(roadmap): PMAT-4646 entry states the 12-case self-test (R1-R12), per q-4647c finding Pmat-Ticket: PMAT-4646 Agent: aprender-ca Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * ci(gate): the ratchet comment says where the baseline lands (#4647) and that the mutants rules are off on purpose until #4648 Quorum lane finding (q-4587c, sonnet): the comment named ci/mutants-debt.tsv and scripts/check_mutants_debt_ratchet.sh as the blocking rule, but on a branch without #4647 neither exists. Same block text in #4647/#4587/#4588 so the three still merge without conflict. No check changed. Pmat-Ticket: PMAT-4646 Agent: aprender-ca Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * evidence(pr-review): non-author receipt for #4647 at 368f05f Sonnet seat (agent:claude-sonnet-5-5/pr-4647-review), base 00052c0. Verdict FINDINGS (2 warnings, 5 notes; none blocking): PRREV-RATCHET-001 (new-survivor detection from CI results is follow-up, report-only by C188(a)); bashrs 0 errors. Pmat-Ticket: PMAT-4646 Agent: aprender-ca Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(ratchet): bashrs-clean guard; roadmap ACs match the diff; new-survivor detection is #4649 q-4647d (3 lanes) findings, all fixed: - scripts/check_mutants_debt_ratchet.sh: bashrs lint 0 warnings (quoted assignments, EXIT trap beside RETURN, false positives tagged); 12/12 self-test. - AC3 named a survivor-table section the gate does not have: now 'the mutants section result'. - AC4 now lists the PR's own review artifacts (evidence/, quorum json). - New-survivor RED (C187) is out of scope here, filed as #4649. Pmat-Ticket: PMAT-4646 Agent: aprender-ca Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(ratchet): drop the stale 368f05f receipt; AC4 separates the reviewed diff from review outputs q-4647f (gemini, haiku): the in-tree receipt for 368f05f still carried the pre-fix bashrs finding, so lanes judged the current script by it; and AC4 listed roadmap.yaml and the quorum json, which the quorum brief does not contain. The receipt for this head is committed next; the stale one binds nothing. Pmat-Ticket: PMAT-4646 Agent: aprender-ca Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * evidence(pr-review): non-author receipt for #4647 at 1f2e359 Sonnet seat (agent:claude-sonnet-5-5/pr-4647-review), base 00052c0, verdict FINDINGS, none blocking. Quorum q-4647g AGREED 3/3 (committed after the signer). Pmat-Ticket: PMAT-4646 Agent: aprender-ca Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * spec(pp-llama-001): re-date §12 rows 13, 15, 18, 19, 21 to 2026-10-15 — operator C197 None of the five was discharged by its 2026-09-30 expiry, and the D6 andon (scripts/lib/spec_conformance.py) turned every PR and main RED on 2026-10-01. §12 allows exactly one way to move an expiry: an amendment recording who moved it and why. This is that amendment: - Moved by the operator (Noah), ruling C197. Reason: "not delivered for 0.70; carried to 0.71". - Rows 13, 15, 19 and 21 carry the typed date. Row 18 derives it from row 15. - An Appendix D row records the move. - derived_expiries.json was regenerated with --write; exactly the five rows changed. The andon is unchanged. It passes today, and with SPEC_CONFORMANCE_TODAY=2026-10-16 it is RED again (D6). The release notes carry §12's consequence: "NO SPEED DELIVERED: PP-LLAMA-001 rows 13, 15, 18, 19, 21 (re-dated to 2026-10-15)". Refs #4651. Advance warning before expiry: #4652 (post-0.70). Agent: aprender-ca Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(ratchet): a debt file shrunk to zero rows read as malformed — the ratchet could never reach clean Quorum on 65d3fb6 (lanes claude-sonnet-5 + gemini-3.1-pro, independently): check_rows fed awk `<<<""` when the debt file had only its header, so the single empty line hit `NF != 5` and the guard went RED on the final shrink. PMAT-4646 requires GREEN on a shrink. - awk skips NF == 0 (rows() already strips blank lines, so no real row is empty) - self-test R13: header-only file vs a 2-row base is GREEN "2 -> 0". RED before the fix (rc=1), GREEN after; deleting the NF==0 line reds R13 only. - PMAT-4646 AC: 13/13 (R1-R13). - Non-author receipt for 65d3fb6 (verdict FINDINGS) committed as evidence. Refs #4646 Agent: aprender-ca Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(ratchet): self-test kills the 3 surviving guard mutants; ci.yml comment states what actually blocks Review of a683b4a (haiku lane FAIL, non-author seat FINDINGS): - 3 sed mutants survived the self-test (PRREV-MUT-002): empty run id accepted, file-name check unanchored, NF<5 instead of NF!=5. New cases R14 (empty run id), R15 (file named mid-string), R16 (6 fields): each mutant now reds exactly its own case; 16/16. - ci.yml comment said "the blocking rule is the shrink-only survivor baseline". It blocks only GROWING the file; new survivors are #4649 and line-shifted rows are #4653 (new issue). Comment and script header say so. Comment-only; no gate changed. - PMAT-4651 AC: the "NO SPEED DELIVERED" line belongs to the v0.70.0 release notes at tag time, not to this PR. - Not changed: the lane's "R13 is RED" claim. Measured at this head: R13 rc=0, self-test 13/13 then 16/16. - Non-author receipt for a683b4a committed as evidence. Refs #4646 #4651 #4653 Agent: aprender-ca Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(roadmap): quote the PMAT-4646 AC; its "R1-R16: ..." parsed as a YAML map 0711f26 broke roadmap-valid (sov.security + sov.gate, CI 36873128678): "roadmap[1098].acceptance_criteria[1]: invalid type: map". The ": " inside the unquoted AC made it a mapping. Quoted it; pmat work validate rc 0, pmat work status PMAT-4646 rc 0, roadmap sorted rc 0. The cop committed without checking validate's rc; it is checked explicitly now. Refs #4646 Agent: aprender-ca Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * evidence(4647): round-4 review receipt + quorum AGREED (PASS/PASS/PASS) at bd85154 Agent: aprender-ca Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * evidence(4647): stamp diff_patch_id 41b11c0a into the round-4 receipt (Arm 4 A2 binds on-disk) Agent: aprender-ca Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(queue): #4647 was dequeued twice by two queue-only defects — R11 read the CI event, Arm 4 read the squash 1. check_mutants_debt_ratchet.sh self-test R11/R12 ran the script with the job's GITHUB_EVENT_NAME. Under merge_group, resolve_base accepts a single parent as the base, so "no nameable base" returned rc=0 and R11 failed in the queue only (runs 36888741624, 36890959980). The temp repo is not a CI checkout: the subprocess now runs with -u GITHUB_EVENT_NAME. Proof: GITHUB_EVENT_NAME=merge_group --self-test rc 1 before, rc 0 after; pull_request and unset rc 0 both. 2. pr-review-quorum.yml passed the merge_group head (the queue squash) as ARM4_PR_HEAD_SHA. The signer posts pr-review-signature on the PR head, so every queued PR read UNSIGNED (#4632 run 36602609576; #4647 runs 36888741583, 36890960012). ARM4_PR_HEAD_SHA is now the pull_request event's head (empty on merge_group), and Arm 4 reads the PR head from the pulls API; the job gains pull-requests: read. Proof on queue sha 79594c8 against the live API: old wiring rc 1 UNSIGNED; new wiring rc 0, signature on 163c660 verifies, binds pr=4647 pid=41b11c0a, A3 positive control still rejects a corrupt sig. No gate loosened: same key, same trusted-comment binding, same A2 diff bind. Pmat-Ticket: PMAT-4646 Agent: aprender-27 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(roadmap): PMAT-4646 AC names the pr-review-quorum.yml queue-only fix (round-5 quorum failed it as out-of-ticket) Operator 2026-09-28 14:52Z pre-approves gate-keeping workflow edits with red/green proof; cop decision c213 (option a). Cites runs 36888741583/36890960012 and queue sha 79594c8. Pmat-Ticket: PMAT-4646 Agent: aprender-r5 Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * evidence(4647): round-6 review receipt + quorum AGREED (PASS/PASS/PASS) at ce8036e diff_patch_id c75016cb stamped. Round 5 failed as out-of-ticket; ce8036e amended the PMAT-4646 AC to name the pr-review-quorum.yml queue-only fix. Pmat-Ticket: PMAT-4646 Agent: aprender-r5 Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…AT-937 (C221 fat_driver fix) (#4655) * fix(release): models_t1 readiness rows ran without the release apr (#3715) Quorum lanes 1 and 3 (sonnet, haiku) measured 4 of 39 rows RED: 997a34381f made the T-1 readiness step re-prove $REPO_ROOT/target/release/apr against MC and take its surface, and the harness never put an apr there. ready() now plants a stub apr at the path autopilot pins; r_goes also requires the wrapper to receive --surface $AP/surface-t1.json from that binary. New row readiness-stale-apr-stops (an apr from another commit stops the step before the wrapper runs) and mutant readiness-any-apr (the check removed -> RED). 41/41 rows; floor 41. Lane 2's minor: the R8 fixture comment still called report the committed mode. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Agent: aprender-0d (cherry picked from commit 57a94972f37968111062b0d7c8dbb06ef69aa4ad) * fix(#3715 R10): delete the report-only arm -- enforce is the only mode, and a WARN row stops T-1 and T-4 Operator R10 (2026-09-28): "report-only is not an option". L19: a silent report-only check = hard stop. - release_readiness.sh: resolve_mode prints enforce or fails; `report` as the env value OR as a committed DEFAULT_MODE is a caller error (3). Both WARN R8 REPORT-ONLY arms (the "not a stop until #3712" line) are gone. Selftest row mode_has_no_report_arm. - autopilot.sh T-1: a WARN R8 row from the wrapper is a die, even at rc 0. - check_publish_preflight.sh T-4 R8: rc 0 is not a pass without "#3715 ENFORCE PASS" for this version and HEAD; a WARN R8 row refuses. - Rows flipped: readiness-warn-continues -> readiness-warn-stops; r8_report_mode_warn_passes -> r8_report_mode_warn_refuses; new r8_rc0_without_enforce_pass_refuses. Each killed by its mutant. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Agent: aprender-0d * G2 (#4590): withdraw every GB10 dense claim through data — qwen2@gx10 de-claimed with qwen3 Operator ruling 2026-09-28 16:15Z widens R1 (Qwen3 only) to every dense model on GB10: serial prefill is the sm_12x default for every non-hybrid arch (select_prefill_path). - cells.declaimed: + qwen2 @ gx10 (issue 4590, until 0.70.1); rung qwen2-1.5b-q4km hosts: [lambda] - README + CHANGELOG: "GB10 dense models: prompt processing unbatched (11.4 tok/s); fixed in 0.70.1 (#4590)" - enforce-gate proof (ont_release_readiness.rs): withdrawn on gx10 owes nothing; a planted gx10 claim of the dense model is RED; lambda dropping one dense row is RED (withdrawn != waived) Agent: aprender-e6 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(ci #4587): lib-test container trusts /workspace so the pinned origin/main resolves The workspace-lib docker run is root over a host-owned checkout; git refused it (dubious ownership), rev-parse origin/main failed, the refinement gate skipped and lint_passes_on_real_contracts asserted under CI (run 36449404284 job 109019972318). Reproduced on intel with sovereign-ci:stable: rc=128 without, rc=0 with the env. Agent: aprender-88 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * release(pv #4604): C2 bump — pv set at 0.69.4, update-check + build-sha off by default macros, aprender-common, aprender-contracts and aprender-contracts-cli pin 0.69.4 (v0.69.3 is taken; a version string is never reused). The cli's aprender-update and aprender-build-sha deps become optional features so the crates.io pv builds only from crates live there (B2a). Without build-sha, build.rs stamps APR_GIT_SHA from APR_GIT_SHA_OVERRIDE or v<ver>+no-git. pv_bin.sh and the Makefile pass --features update-check,build-sha, so in-tree pv keeps its update check and sha. binary-release.yml is a separate branch, pending operator OK. Agent: aprender-88 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> (cherry picked from commit 34ff1f7432a20f643161250d869cd8e0c9729828) * release(pv #4604): 6-crate set — aprender-build-sha + aprender-update at 0.69.4 cargo publish needs optional deps live on crates.io too (0d C1 dry-run). Both are new crate names there; publishing them waits on the operator ruling. Fallback without them: release/pv-0.69.4-fallback. Agent: aprender-88 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> (cherry picked from commit 1e5f8a84849003421e8e8640346a048bc87d595c) * ci(binary-release #4604): release pv with --features update-check,build-sha The C2 bump makes both features off by default so crates.io pv builds from live crates. Release binaries keep the update check and build sha. Held: no merge until operator OK (workflow edit). Agent: aprender-88 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> (cherry picked from commit 372c980391143706bd811473bd43b3bc3385712b) * fix(shapes #4587): the #3610 exemption moves off the shape so the fleet-pinned pv parses it pv 0.69.1 (the fleet pin, intel-clean-room) refuses the shape key allowEmpty (rc=3: shape release-readiness-v1.refusal uses unsupported allowEmpty), so check_fleet_pv_shapes_gate.sh reported pv-no-verdict (run 36449404284 job 109019972501). Deleting the key would make every green release decline under --shape release-readiness-v1 (the refusal shape is empty by design). The reason now lives in a contract-level allow_empty map keyed by shape id, which 0.69.1 ignores and HEAD applies with the same semantics; the in-shape key still parses. Agent: aprender-88 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(#4587): x86-main reds from run 36462668621 — contracts recipe and facade guard vs the scoped 0.69.4 pv set 1. `make contracts` called scripts/contracts_gate.sh, which PR-1 does not carry (it lands with PR-2). Restore main's step-list recipe; PR-2 brings the gate script and its recipe together. make_contracts_propagates.sh: 17/17 rows, mutants included. 2. make_contracts_propagates.sh faked pv at the ROOT workspace version, but pv_bin.sh declares the CLI crate's own version first. The fixture now reads it the same way. 3. check_facade_compat.sh CURRENCY and PUBLISH ORDER compared against the root workspace version; R3 already used the fronted crate's own. All three now use the version this tree publishes for the fronted crate (identical when unscoped; an unreadable version is a FAIL). Facade pins -> 0.69.4, crates/facades/Cargo.lock regen. Agent: aprender-88 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * feat(ladder): a de-claimed (host, arch) the ladder still claims is RED — planted RED, qwen3moe@gx10 (#3715) STAGED, branch-only (cop 20:00Z/20:26Z; operator rules 07:30Z). Not for merge. cells.declaimed withdrew a (host, arch) from the owed cells, but nothing failed when a rung's host set (what released pv reads) or a long-rung representative still claimed it. declaim_still_claimed() in the judge (enforce: judge rc 1 -> check exits 1): - a rung whose hosts (or, absent, every required host) name a de-claimed (host, its arch) -> FAIL - a long_rungs_for representative for an arch de-claimed on every required host -> FAIL Case table: red-cells-declaimed-still-claimed (new), green-cells-declaimed narrowed to hosts: [lambda] (it claimed the de-claimed pair); cmutant declaim-claimed killed. Self-test 153/153. Data: cells.declaimed + (gx10, qwen3moe), issue 3715, until 0.70.1 (0d H4 19:54Z: arch-level DATA is clean for qwen3moe only). Real-contract proof in evidence/3715-declaim-planted-red/: real = 0, planted qwen3moe rung = 1, narrowed = 0, orphan representative = 1. Agent: aprender-e6 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(serve): a turn past the device KV cache is a 400 "context exceeds N", not a 500 (D5) Operator ruling D5 (ruling-3715-2010): "a 500 is a crash-class lie". A 27155-token turn on a dense CUDA model reached BorrowedCudaForward::reserve, failed with "the device KV cache holds 4096", and came back as HTTP 500 (h4-depcheck-0d, lambda + gx10 serve/code cells). - serving_context(): the model's context capped by the device KV cache; the borrowed serve session reports it, so max_tokens clamps to what is left. - AppState caches it at construction (PMAT-073: no read lock in the handler). - serve_context_refusal(): the pre-flight 400 in the CUDA chat, CUDA /v1/completions and /generate/stream routes, before any path takes the model, so a streaming client gets a 400, not a mid-stream error. - The direct paths map ContextLimitExceeded through generation_error_status. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(#3715): release-note warning — qwen3moe on GB10 is not claimed for 0.70 Operator D1 (relayed by the cop, 2026-09-28 20:10Z): the qwen3moe-gx10 de-claim lands with the planted RED and a release-note warning. CHANGELOG Known issues + README known-issue note, same shape as the #4590 entries. Agent: aprender-e6 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(D1): the qwen3moe de-claim said lambda does not hold qwen3moe — it holds both files lambda holds Qwen3-30B-A3B-Instruct-2507 and Qwen3-Coder-30B-A3B (~/models, lambda H4 queue), so the entry now says what withdrawn-not-waived means there: lambda still owes every qwen3moe cell. Data text only; the judge keys on (host, arch) and is unchanged. Falsifier re-run on the real contract: entry -> withdrawn rc 0, drop -> owed again, bare -> rc 1. Refs #3715 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Agent: aprender-0d * fix(facade guard #4587): exact scoped allowlist — only the declared #4604 set at 0.69.4 may leave the workspace version Quorum q-pr1-r81cd2 (gemini FAIL): comparing a crate against its own version let a coordinated off-workspace drift pass. pub_ver now admits the workspace version, or the EXACT version scoped_ver declares for that crate; anything else (unlisted crate, other version, unreadable) is RED. The only input the old workspace-only rows refused and these admit is the declared set itself. Agent: aprender-88 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(D1): CHANGELOG said no other host holds qwen3moe — lambda holds both files and still owes them Same false claim 0d fixed in the contract (e59329fe67); the release note carried it too. Agent: aprender-e6 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(make #4587): no PR-1 recipe calls a script only PR-2 carries Four Makefile recipes in the split named scripts that live in PR-2 alone: census (check_census_derived.sh), guards-local (ci_guards.sh), oracle-owl (tests/oracle/owl) and coverage-serve (the .py -> .sh rename would have broken `make coverage`). Those hunks go back to main and travel with their scripts in PR-2. lint-ratchet stays: PR-1 carries EV-11, whose contract names scripts/lint_ratchet.sh as its writer, so the script comes with it (--self-test 12/12). Agent: aprender-88 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * roadmap: PMAT-4616 (D5): the ticket #4614 implements, not #3715 #3715 is the pv SHACL release-readiness shape. D5 is its serve × consumer-context child: a turn longer than the device KV returned 500. The quorum judged this diff against #3715's acceptance and failed it 3/3 on the mismatch, so the serve fix now has its own ticket and acceptance. Refs #4616, #3715 Agent: aprender-2e-serve Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(dogfood #4587): re-derive per_binary/band/cluster baselines for PR-1's 22 ledger rows run 36473794463 guard-tree step 11: the split moved the 22 surface_audit rows (11 apr, 11 pv) and the rows/ratio totals into PR-1, but not the per_binary / per_band / per_cluster entries. Re-derived with scripts/dogfood_baseline.py; --check PASSES. Agent: aprender-88 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(guards #4587): mutation sets for the three guards PR-1 adds, and a CI caller for the two lean ones Review B3 (PR-REVIEW-SKILL-002 v2 S3.D): check_facade_compat.sh (scoped allowlist), check_leanchecker_scoped.sh and check_lean_modules_reachable.sh had only self-authored self-tests, and the two lean guards had no CI caller. - mutate_{facade_compat,leanchecker_scoped,lean_modules_reachable}_guard.sh: 17/17, 31/31, 18/18 killed. Each patch fails closed (INVALID, never a survivor), the M0 baseline must pass, and mutants run against copies. - First runs killed 0/17, 27/31 and 16/18. The self-test rows that kill the survivors were added. No check was removed or loosened. scoped_ver/pub_ver were moved, unchanged, above the self-test so it can reach them. - ci/sections.yml guard-cargo runs the mutation sets, plus both lean guards (self-test and the real check). Full facade guard on intel: PASS. Agent: aprender-88 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(version #4587): a facade pin must name the crate it re-exports, not the workspace beside it guard-tree step 17 (bump-version.sh --check) failed on run 36485785730. The check compared each facade pin with the ROOT workspace version (0.69.3). But PR-1's scoped pv set moves aprender-contracts{,-macros} to 0.69.4, so the pins on those two crates (0.69.4) were the correct values, and the check called them wrong. --check now compares each pin with the version its `upstream` path crate actually carries. A crate on version.workspace = true carries the workspace version. Case table rows 7-10: one pass and one RED for a literal-versioned upstream, and one pass and one RED for a workspace-versioned upstream. Each RED row must stay RED. The self-test passes 10/10, and --check passes on this tree. Agent: aprender-88 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * evidence(PMAT-4616): pr-review receipt for 167bd5e126 (reviewer sonnet-5, DEGRADED, advisory finding) Non-author reviewer (Sonnet 5), per COP 22:16Z applying the 21:10Z committed- .minisig route. diff_patch_id cf52ad666e. Verdict DEGRADED: cuda (docs MCP OAuth-gated) and mutation (no cargo test on lambda) unreachable; pmat and agy gemini-3.1-pro-high consulted; crux not triggered. Guard rejects only B1 (unsigned) until the CI signer's .minisig is committed. Advisory finding (measured): /generate and /batch/generate still map every CUDA error to 500 (batch_processing.rs:492, batch.rs:515). Follow-up, not folded here: changing the diff would void this receipt. Agent: aprender-2e-serve Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * evidence(PMAT-4616): commit the CI signer's signature for the 167bd5e126 receipt Per COP 22:16Z (21:10Z committed-.minisig route). The .minisig is verbatim from pr-review-signature check-run 109170961010 (CI signer, repository key). minisign -V under .github/pr-review.pub: verified, trusted comment pr=4614 pid=cf52ad666e. check_pr_review_receipt.sh: ACCEPT (positive controls fired). Agent: aprender-2e-serve Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * evidence(pr-review #4587): Sonnet review receipt for head ba09f72af5, bound to patch-id 1817f9d2 The reviewer is claude-sonnet-5; the author is claude-opus-5-5. The verdict is BLOCK, and its only reason is the PR1-MUTANTS HARD-STOP (3555 mutants in the diff, cap 60, operator ruling at 07:30Z). B3 (guard mutation sets) is resolved by c1e4d195c5. The receipt is unsigned; the ci pr-review-sign job signs it. Agent: aprender-88 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(ci #4587): the lean-guards step runs each guard under setsid --wait check_guard_steps_isolated.sh was RED on the local guard-tree run: the six invocations the step added in c1e4d195c5 shared the runner's process group (#4133). Now green: every direct guard invocation in 25 workflows is isolated. Agent: aprender-88 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(cuda): chunk the dense batched prefill; QK-norm models stop taking the serial prefill (#3715) The #3715 lambda cells timed out or failed at the 20k rung on two causes, both measured on the RTX 4090 with apr 0.70.0 (8d021f61e): - Qwen3-1.7B (QK-norm): `[PREFILL-PATH] serial cc=89 (qk-norm model)`. One forward_gpu_resident per prompt token: 7955 tokens took 109 s, 26131 did not finish in 300 s at GPU util 99-100 %. Not #4609: the backend was the GPU throughout. - Qwen2.5-0.5B IQ3_M (batched): CUDA_ERROR_OUT_OF_MEMORY in prefill_all_layers_gpu at 26148 tokens. One pass allocates a num_heads x S x S f32 score matrix (38 GB), and its u32 row offsets wrap past 2^32. The batched prefill now runs in chunks whose score matrix fits the 1 GiB budget the Qwen3.5 prefill already uses; each chunk appends to the KV cache through the existing cache_len > 0 attention path. A prompt that fits one chunk runs exactly as before. QK-norm models keep FP8 off and now take the FP16 batched prefill, which #3413 measured passing CPU parity; serial was chosen only because the unchunked prefill OOMed 8B. Refs: #3715 Agent: aprender-ec Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(chat): size the chat device KV to the context when VRAM allows (#3715) apr chat built its CUDA model with a flat 2048-position KV. The dense session refuses a longer turn on the device, so a 26k-token chat turn on Qwen3-1.7B ran on the CPU and hit the 600 s cell timeout (GPU util 0%, measured on lambda). apr run already sizes its KV to the turn (#4268); chat loads before it has read a turn, so it now takes the model context, capped by free VRAM after the weights, the FP16/FP8 prefill cache, the GH-178 reserve and the chunked-prefill score budget. When nothing is left (Qwen3-8B on 24 GB), it keeps the 2048 floor. Refs: #3715 Agent: aprender-ec Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(cuda): Qwen3.5 prefill dequantizes F16 and IQ4_XS on the device (#3715) The Qwen3.5 batched prefill dequantizes each weight to F32 for one SGEMM, and had kernels for Q4_K/Q5_K/Q6_K/Q8_0 only. Qwen3.5-4B-UD-Q4_K_XL (F16 ssm_alpha/beta, IQ4_XS FFN blocks) and Qwen3.5-0.8B-IQ4_XS therefore refused the CUDA prefill, the F2 guard rejected the GPU, and every lambda cell ran on the CPU (GPU 0%) into the 600 s timeout. Adds F16DequantKernel and Iq4XsDequantKernel. The values match the per-token GEMV weights bit for bit (device tests against the ggml reader). Refs #3715 Agent: aprender-ec Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(cuda): a QK-norm model keeps the serial prefill when its FP16 batched prefill does not fit (#3715) 3c2a354b7d moved QK-norm models from the serial prefill to the chunked FP16 batched one. The FP16 weight cache is warmed on either path, but only batched also needs the prefill workspace. Qwen3-8B on the 24 GB 4090 has 4.7 GB of weights and a 15.1 GB FP16 cache, which leaves 0.1 GB, so init_prefill_workspace failed and a 2k chat turn fell back to the CPU (382 s, rg5 C-q38b-chat-2k). The path is now decided against the device before the cache is warmed: weights + cache + KV + prefill reserve over free VRAM keeps serial, as on the base. Qwen3-1.7B (about 9 GB of 23) stays batched. BATCHED_PREFILL still overrides. The resident-bytes estimate is shared with session_kv_len. Refs: #3715 Agent: aprender-ec Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * evidence(models #4587): re-derive supported.yaml after the README edit derive_model_manifest.sh --check was RED on the local guard-tree run (step 28): PR-1 moved README.md lines, so every README citation line in the derived manifest drifted. Re-derived with the script; line numbers only, 24 models. Agent: aprender-88 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * evidence(pr-review): commit the CI signer's .minisig for #4587's ba09f72af5 receipt Verbatim from the pr-review-signature check run 109176017607 on 7eb331468c (trusted comment pr=4587 pid=1817f9d24d). minisign -V under .github/pr-review.pub passes. Same route as #4618 500fb1fb10 and #4612. Agent: aprender-88 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * evidence(pr-review): #4620 non-author receipt for fdc165f50d (DEGRADED: mutation unreachable) Reviewer agent:claude-sonnet-5/pr-review-4620; author claude-opus-5-5/aprender-ec. Written without diff_patch_id or .minisig. The CI signer stamps both. Refs: #3715 Agent: aprender-ec Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * evidence(pr-review #4587): Sonnet review receipt for head 7218961ef9, bound to patch-id 9d93fbed Reviews the delta ba09f72af5..7218961ef9 in full (setsid --wait on the lean guards step; supported.yaml re-derived) plus the committed CI .minisig. Verdict BLOCK: PR1-MUTANTS only (3555 > cap 60), the operator HARD-STOP. Agent: aprender-88 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * evidence(pr-review #4587): cuda consultation is unreachable, not consulted-with-no-queries Arm 4 on 293425f2d8 (pr-review-quorum 36504222978) rejected the receipt: B1, cuda consulted with queries: []. The S3.B path trigger fires on two files and the CUDA docs tool is not authenticated in the reviewer's session or the cuda-docs-reviewer's, so no query ran. Recorded as unreachable (allowed with a non-PASS verdict); the device claims were not checked against docs. Verdict stays BLOCK. check_pr_review_receipt.sh ACCEPTs a signed copy. Agent: aprender-88 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test(cuda): the fit-check no-op cases assert the path of the profile they called (#4621) The last loop in fit_check_leaves_non_qk_norm_fp8_and_forced_paths_alone built two fresh profiles and asserted they were batched, so a mutant that switched the path to Serial while still returning false passed it. Each case now asserts the path of the profile it called. Red on that mutant, green on the code (intel). Quorum finding, PMAT-4621 round 1. Agent: aprender-ec Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(PMAT-4616): D5 gate reds — the cap is CUDA-free and tested, try_cuda_backend no longer grows x86-main on 099d07de7d was red on this diff, not on the runner: - complexity: try_cuda_backend grew (cog 43->46, cyc 14->15). The chat pre-flight now rides the tokenize result through fit_serving_context, so the function gains no branch. - mutants: serving_context lived in the cuda-gated dense_session_borrowed, where no CPU lib test can reach it. The rule is now dense_session::cap_context(context_length, device_kv), CUDA-free, with lib tests on every arm; fit_serving_context has lib tests for pass, refuse (400) and a kept tokenize error. - roadmap.yaml regenerated for the PMAT-4616 fragment. - surface_audit.csv: the /v1/completions row's cited line moved 398 -> 405. Refs #4616 Agent: aprender-2e-serve Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * evidence(pr-review): #4620 non-author receipt for 65370adc9a (DEGRADED: mutation sweep unreachable; targeted mutant killed) Agent: aprender-ec Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * evidence(pr-review): drop the superseded fdc165f50d receipt from #4620's tree Its findings.sarif carries a placeholder excerpt_sha256 ("n/a-short-code-idiom-not-hashed") that S1.1 would reject once signed, and the receipt binds that sarif by sha256, so the author cannot correct it without editing the reviewer's record. It is superseded by the 65370adc9a receipt (the head's code) and stays in history at 18b6c8341f. Quorum finding, PMAT-4621 round 2 (gemini lane). Agent: aprender-ec Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(audits): quorum PMAT-4621 AGREED 3/3 (sonnet-5, gemini-3.1-pro-high, haiku-4-5) at 2b8d79acd7 Agent: aprender-ec Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * evidence(pr-review): stamp #4620's receipt with the diff patch-id Arm 4 computes (b3f1fd942a) Arm 4 A2 reads predicate.diff_patch_id from the COMMITTED receipt; the CI signer stamps only its runner copy, so an unstamped receipt is legacy and RED. Computed with scripts/lib/pr_review_patch_id.sh prpid_compute over 722fd1621f..HEAD (evidence and quorum json excluded), equal to the value Arm 4 printed in run 36510383864. Agent: aprender-ec Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(PMAT-4616): CUDA /generate and /batch/generate map ContextLimitExceeded to 400 Quorum lane 2 (gemini-3.1-pro-high) at 926edb2135 found two GPU-resident generate paths that still hardcoded 500 for every generation error, bypassing generation_error_status(): batch_processing.rs try_cuda_generate and batch.rs try_cuda_batch_generate. Both now route the status through it, so an over-length request is a 400 on the CUDA path as on the CPU path. cuda clippy on intel: only the pre-existing residual.rs:314/343 borrow_deref_ref. Agent: aprender-2e-serve Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * evidence(pr-review): #4620 receipt passes the whole guard, not only the signature arm Arm 4 (run 36511265641) verified the signature, then rejected the content: the antigravity divergence counted 1 against 0 findings. The reviewer corrected it to the agy output it actually recorded (findings: []), added the pmat duplication fields the guard requires, renamed two SARIF drivers to the allowed set, and re-bound findings_ref. The guard ACCEPTs it under a throwaway key. diff_patch_id is unchanged. Agent: aprender-ec Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(PMAT-4616): every serve generate/logprobs/perplexity error routes through generation_error_status Quorum at 1ead14b2d7 (sonnet-5 + gemini-3.1-pro-high FAIL) found more handlers that hardcoded 500 on a generation error, so ContextLimitExceeded still leaked as 500: SafeTensors CUDA chat, GPU /generate, logprobs, perplexity, the APR/demo batch + stream paths and the openai GPU fallback. Swept all 12 generate-error sites in api/ (decode errors stay 500: those are server faults). generation_error_status keeps every non-context error at 500, so only the over-length case changes. intel: clippy --lib default rc0; --features cuda only the pre-existing residual.rs:314/343 borrow_deref_ref; 15 targeted lib tests pass. Local mutants not_measured (unmutated baseline timed out at 600s). Agent: aprender-2e-serve Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(guard): C13 read a present release asset as MISSING (printf into grep -q under pipefail) scripts/check_release_assets.sh:113, under the file's own `set -uo pipefail`: if printf '%s\n' "$have" | grep -qxF "$want"; then grep -q exits on its first match; printf then takes SIGPIPE (141), pipefail reports the pipeline as 141 and the `if` takes the MISSING branch though grep matched. This turned guard-tree RED in CI run 36510076745 (step 26, release_criteria.sh --self-test row 6, "C13 is credited when the release carries all sixteen assets", rc=1 wanted 0) on a code delta that was evidence-only; the identical merge passed 10/10 three times locally. Fix: a here-string (`grep -qxF -- "$want" <<< "$have"`), no pipeline. Same class as 0c6932fd54. The C0 leg-1 check in release_criteria.sh:85 had the inverse form (a matched '✗' could read as no failure) and gets the same fix; that one could only ever PASS wrongly. Regression row (check_release_assets.sh --selftest row 7): a complete release that also carries 20000 other assets, so $have outgrows the 64 KiB pipe buffer and the race becomes deterministic. Old script: rc=1 (x3). New script: rc=0 (x3). selftest 12/12; release_criteria --self-test 10/10. Agent: aprender-88 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * evidence(pr-review #4587): Sonnet receipt for code head 92aac9b4e3 (BLOCK: PR1-MUTANTS, unchanged) Supersedes the 7218961ef9 receipt. Reviews the one-commit delta 92aac9b4e3 (printf|grep -q under pipefail -> here-string in check_release_assets.sh:113 and release_criteria.sh:85, plus the 20016-asset regression row). Reviewer reproduced old rc=1 x3 / new rc=0 x3; selftests 12/12 and 10/10. diff_patch_id 5a638e1100a63e65b81061815d2f8581ce08644d (base ef72c05..92aac9b4e3). Signature comes from CI pr-review-sign on the pushed head. Agent: aprender-88 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * split(#4621): #4620 carries roots 1+5 only (chunked dense prefill, F16/IQ4_XS device dequant) The diff held 97 mutants against the gate's cap of 60. The QK-norm fit check (root 2) and the chat session KV (root 3) move to their own PR, per the cop's split ruling; they are restored here to the merge base, unchanged in content. Refs #3715 Agent: aprender-ec Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * evidence(PMAT-4616): pr-review receipt for 7008062734 (reviewer sonnet-5, DEGRADED, advisory findings) Non-author reviewer (Sonnet 5). Verdict DEGRADED: cuda (docs MCP OAuth-gated) and mutation (no cargo test on lambda) unreachable. Confirms the 167bd5e126 receipt's finding is fixed (try_cuda_generate / try_cuda_batch_generate now route through generation_error_status). Guard rejects only B1 (unsigned) until the CI signer's .minisig is committed. Advisory follow-ups, not folded (changing the diff voids this receipt): batch_processing.rs:343 batch_generate_gpu still 500; /generate and /batch/generate have no serve_context_refusal pre-flight; the effective_max_tokens `>` vs serve_context_refusal `>=` boundary. Agent: aprender-2e-serve Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * evidence(PMAT-4616): bind the 7008062734 receipt to diff patch-id 2fe3407e2f The reviewer (sonnet-5) had left a placeholder in predicate.diff_patch_id; it now carries the pinned prpid value, equal to the one CI's Arm 4 measured (base 2817c6d97b). Guard: only B1 (unsigned) until the signer's .minisig. Agent: aprender-2e-serve Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * evidence(PMAT-4616): 7008062734 receipt — crux block and citations made guard-complete The signature exposed a B1 the unsigned state had hidden: crux was `consulted` without surfaces/comparative_claims. Reviewer (sonnet-5) filled the S3.C arrays honestly and added the missing citation fields; verified ACCEPT under a throwaway key. patch-id binding unchanged (2fe3407e2f). Agent: aprender-2e-serve Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(#4621): a chunked prefill's later chunks never replay the prefill graph PREFILL_GRAPH=1 (opt-in, off by default) replays a graph that bakes positions 0..S-1 and resets the KV lengths to 0. With the chunked dense prefill, chunk 2+ of the same length would have replayed it over a filled KV cache, silently corrupting it. Only a chunk starting at position 0 may take the graph now; the rest run eager. Found by the non-author Arm-4 review of #4620. Refs #4621, Refs #3715 Agent: aprender-ec Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * evidence(pr-review): drop #4620's pre-split receipt (it reviewed roots 2+3, now in the stacked PRs) Agent: aprender-ec Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * evidence(PMAT-4616): commit the CI signer's signature for the 7008062734 receipt Agent: aprender-2e-serve Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * evidence(pr-review): #4620 Arm-4 receipt for split head 31cd3f6c40 (PMAT-4621) Non-author Sonnet 5 review of part A (chunked dense prefill, F16/IQ4_XS device dequant, PREFILL_GRAPH first-chunk-only). Verdict DEGRADED; diff_patch_id stamped. CI signer adds the .minisig. Refs #4621 Agent: aprender-ec Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(serve): CUDA /generate and /batch/generate refuse over-context prompts with 400 before the model lock (PMAT-4616) Quorum @0c2928642d (gemini, haiku) found two remaining 500 paths: - try_cuda_generate / try_cuda_batch_generate had no pre-flight, so an over-context prompt reached the engine. They now call preflight_serving_context, the same check chat and completions run through fit_serving_context: a 400 "context exceeds N" before the lock. - gpu_batch_completions_handler's batch_generate_gpu Err arm hard-coded 500; it now maps through generation_error_status (ContextLimitExceeded -> 400). Two tests: no cap -> Ok; cap 3 -> 2 Ok, 3 refused with 400. Roadmap entry: title and criterion 3 name every covered route; criterion 1 states that the --context-length 32768 -> 200 case is #4603, not this ticket. Complexity vs main: cyclomatic and cognitive unchanged on all touched fns. intel: clippy default rc0; cuda only pre-existing residual.rs; d5_/cap_context 13 passed. Agent: aprender-2e-serve Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(serve): /batch/generate pre-flights every prompt before taking the CUDA write lock (PMAT-4616) Quorum @7bcdde94bd (sonnet) FAIL: try_cuda_batch_generate ran preflight_serving_context inside the generate loop, after cuda_model_lock.write(), so the doc/commit claim "before the model lock" was false for /batch/generate. Tokenize, empty-check and pre-flight all prompts first, then lock and generate: an over-context prompt anywhere in the batch is a 400 with no lock taken and no GPU work spent on earlier prompts. intel: clippy default rc0; cuda only pre-existing residual.rs:314/343; d5_/cap_context 13 passed. Agent: aprender-2e-serve Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * evidence(pr-review): #4620 post-split quorum (3/3 AGREED, scope A) + S4.2-compliant receipt SARIF (PMAT-4621) The quorum json is replaced with the round at head 6562916bee, which covers part A only (the pre-split file claimed root-3 chat KV, now C's). The receipt SARIF had its runs moved from driver 'pr-review-primary' to one run per consulted consultation, and findings_ref is rehashed. Same subject and patch-id. Refs #4621 Agent: aprender-ec Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(serve): /v1/batch/completions pre-flights the serving context (PMAT-4616) Non-author review @30b29b3798 (blocking, measured): gpu_batch_completions_handler had no serving-context pre-flight, and its GPU callee batch_generate_gpu has no context guard at all, so mapping its Err through generation_error_status could not turn an over-device-KV prompt into a 400. Pre-flight every encoded prompt right after tokenization, before either the GPU or the CPU branch generates. intel: clippy default rc0; cuda only pre-existing residual.rs:314/343; d5_/cap_context/batch_completions 32 passed. Agent: aprender-2e-serve Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * evidence(pr-review #4587): cuda consultation run for real (consulted, 5 queries, 0 contradicted) The nvidia-cuda-docs MCP was signed in 2026-09-29 06:04Z. A headless claude-sonnet-5 lane (not the author's model) ran 5 named queries on the FP8 E4M3/E5M2 claims in fp8-interchange-v1.lean and the cuda_simd_scores fixture: 4 found, 1 no-authority-found, 0 contradicted. Every excerpt_sha256 verified. Replaces the "unreachable" block; PRREV-CUDA-001 updated; findings_ref re-hashed. Verdict unchanged: BLOCK (PR1-MUTANTS). Agent: aprender-88 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(serve): /v1/batch/completions pre-flights against the cached model's own context (PMAT-4616) Non-author review @ad102790a4 (blocking, measured): the pre-flight read AppState::serving_context, which every cached-model constructor sets to None, so on the only states this handler serves it was a runtime no-op. And batch_generate_gpu sizes KV caches from prompt + max_tokens with no context check. Split preflight_context(Option<usize>, n) out of preflight_serving_context; the handler refuses against the device serving context when set, else cached_model.model().config.context_length. +1 test for the explicit-context arm. intel: clippy default rc0; cuda only pre-existing residual.rs:314/343; d5_/cap_context/batch_completions 33 passed. Agent: aprender-2e-serve Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(mutants-gate): a mutant in code the build never compiles is NOT_MEASURED, never MISSED (PMAT-4621) cargo-mutants tests a mutant with `cargo test -p <pkg> --lib` at default features, so a mutant in #[cfg(feature = "cuda")] code is never compiled and its run is the baseline. On #4620 (run 36521053810, artifact 11015244648) all 7 mutants were in cfg(cuda) code. 0 cuda:: tests ran of 16149, and the MISSED/TIMEOUT split was timing only (292 s vs 300.1 s). The compiled set is now measured from dep-info (cargo check -p <pkg> --lib in a private target dir). A survivor outside it is NOT_MEASURED and stays RED unless evidence/mutants-local/*.txt carries 'CAUGHT <line>' from a build that compiles it. Compiled survivors stay RED as before. A failed check or empty dep-info is RED. No limit changed. Red/green: - Self-test: 5 new rows, and 5 rule mutants are each killed. - Real artifact, intel, real cargo: 7/7 NOT_MEASURED OPEN, rc=1. The dep-info names 488 aprender-serve files and 0 under src/cuda/. Refs #4621 Agent: aprender-ec Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * ci(mutants): A' survivor table -- CI matrix shards bound to head sha + diff hash, verified before trusted Operator ruling 2026-09-29 06:36Z (via the cop): CI generates the survivor table itself; any gap is RED. Scope PR-1 and PR-2 of 0.70 only (ci/mutants-table-refs.txt). - mutants-table-scope: decides in_scope from ci/mutants-table-refs.txt at the head sha and re-proves the gate rule + checker case tables on every PR. - mutants-shard: 16 matrix jobs on the PR HEAD sha, cap 0, each writes the full --list universe, its slice's outcomes.json and meta (head sha, diff sha256, run). - mutants-table: shard matrix must be success, then scripts/ci/mutants_survivor_table.py checks every listed id exactly once, killed/equivalent/survived, equivalents need a proof + a quorum receipt naming the id, survivors <= MUTANTS_MAX_MISSED, every shard's sha and diff hash match. - gate: GATE-MUTANTS-TABLE-RULE requires mutants-table success for a listed ref and a scope decision on every pull_request (case table: scripts/check_ci_gate_mutants_table_rule.sh, 14 rows, 3 planted weak rules RED). - mutants section: hands over for a listed ref only; the 60 cap is unchanged for every other PR. Checker case table 18/18; each of 10 mutations of the checker turns a row RED. Agent: aprender-88 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(mutants-table): an equivalent's receipt must be a pr-review statement, not any file naming the id Sonnet review F1 on a79535bd51: the equivalents row was accepted when ANY repo file contained the mutant id, so a stray dump could reclassify a real survivor. The receipt must now sit at evidence/pr-review/<pr>/<sha>/receipt.intoto.jsonl and parse as an in-toto pr-review Statement. Two new RED rows (stray file, right path but not a statement); disabling either check turns them RED. Agent: aprender-88 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(PMAT-4616): scope criterion 3 to dense GGUF CUDA; APR Q4K split to #4625 Quorum @1531ead607 (sonnet) FAIL on three scheduler-channel sites: - cuda_chat_backend continuous-batch arm and gpu_completions_handler's token_rx loop map scheduler errors to 500. Both run only after the D5 pre-flight (fit_serving_context at cuda_chat_backend.rs:126, the serve_context_refusal check in the completions handler), and the channel carries String, not RealizarError, so there is no ContextLimitExceeded to map. Comment added at the chat arm, matching the completions one. - APR Q4K chat/completions backends have no pre-flight: a separate engine bounded by max_position_embeddings that its AppState constructors never see. Filed #4625; criterion 3 now names the dense GGUF CUDA routes and "typed" generate-error sites, and points at #4625. Agent: aprender-2e-serve Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(mutants-table): an equivalent needs a CI-signed receipt, verified with the BASE branch key Sonnet re-review F2 on 0a5c1325c7: a shape-checked receipt could be hand-written by the PR author in the same PR. An equivalents row now also needs reviewer_actor != author_actor and <receipt>.minisig verifying with minisign against --pubkey; mutants-table passes the BASE branch's .github/pr-review.pub (git show origin/<base>:...), which the PR cannot edit, and only the CI signer holds the secret half. No --pubkey or no minisign with a row is RED. minisign comes from the pinned installer in both jobs. Four new RED rows (self-review, unsigned, other key, no --pubkey); disabling each check turns the table RED. Agent: aprender-88 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * evidence(pr-review #4587): Sonnet receipt for code head b132bf3c34 (A-prime survivor table; F1+F2 fixed) Three Sonnet lanes over the 3-commit delta since d7b36adff0: a79535bd51 APPROVE (F1-F3 minor/nit), 0a5c1325c7 BLOCK (F2), b132bf3c34 APPROVE. Verdict stays BLOCK only on PR1-MUTANTS, now judged by the required mutants-table job. The cuda consultation is carried unchanged (no device code in the delta). Agent: aprender-88 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test(cli_roles): kill the 14 M1 shard-4 survivors (PR-1 #4587) Under --features cli-roles, 14 of 26 PR-1 mutants in cli_roles.rs survived: the role and sampling spellings, SamplingKind::id, SamplingArg::groups and kind_of, strings and path_bufs had no test that checked their values. The four tests pin each spelling and id, the one group per kind, the kind of each sampling argument (and none for a plain one), and the order-preserving conversions. cargo-mutants 27.1.0, --features cli-roles: 25 caught, 1 unviable, 0 missed (was 11 caught, 14 missed). Without the feature the module is not compiled and every mutant reads MISSED; that is the gate's feature gap, not these tests. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Agent: aprender-3d * fix(mutants-table): feature-gated files run with their feature, not as false MISSED (#4587 A') cargo-mutants tests with default features, so every mutant in a module behind #[cfg(feature = "x")] never compiled and read MISSED whatever its tests said (M1 FALSE-RED: cli_roles.rs 26/26 missed; main.rs under update-check). A workspace-wide --features pkg/feat is refused by cargo for any selected package that does not own the feature, so it cannot go on the one run. ci/mutants-feature-groups.txt names each gated file with its package, features and test args. mutants_table_shard.sh excludes those files from the default run, runs each row as its own `cargo mutants -p PKG --features F --file FILE --shard K-1/N`, and merges the outcomes (perl JSON::PP; the sovereign-ci image has no python3). list.txt, the universe, is unchanged: a row that tests nothing leaves its mutants MISSING, so it is RED, never a skip. cuda rows are refused. Red/green proof, scripts/check_mutants_feature_groups.sh, wired next to the gate-rule self-test in ci.yml: the REAL shard script + REAL checker over a fake cargo -> all caught GREEN; planted survivor in a gated file RED; planted survivor in a default file RED; group tests nothing RED; no group rows (the old behaviour) RED; group writes no outcomes DEAD; cuda row DEAD. Mutating the merge to drop group outcomes turns the GREEN case RED. Measured on intel with the kill tests (prev commit, from aprender-3d): aprender-common --features cli-roles: 26 tested, 24 caught, 2 unviable, 0 missed. Agent: aprender-88 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * evidence(pr-review #4587): Sonnet receipt for code head 05c9e6f2d9 (feature groups + cli_roles kill tests; BLOCK: PR1-MUTANTS until the table runs) Agent: aprender-88 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test(ontology): kill CellsError::file -> String::new() survivor (M1 shard 5) Both refusal tests now assert the file the error names, so the mutant that blanks it is caught. Agent: aprender-27 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test(ontology): kill version_key || -> && and v_star > -> >= survivors (M1 shard 5) "0.+69" is refused only by the digit check, since u64::from_str takes a leading +. Two spellings of one version key keep the first as V*, and cells are read at that spelling. Also rustfmt the file() assertion. Agent: aprender-27 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * evidence(pr-review #4587): Sonnet receipt for code head 38cabdb8af (+ M1 shard-5 kill tests; BLOCK: PR1-MUTANTS until the table runs) Agent: aprender-88 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(release): publish-order.txt omits aprender-build-sha and aprender-update (PR-1 #4587) PR-1 adds two publishable crates, but publish-order.txt does not list them. publish_strict.sh requires the order to equal cascade_universe.py's set, so the pv 0.69.4 cascade would stop before uploading anything: "in the universe but not in the publish order, so never uploaded: aprender-build-sha aprender-update". Neither crate depends on a workspace crate, and contracts-cli, apr-cli and 24 others depend on them, so they go first. Re-proved at this tree: the set difference with the universe is 0, and the publish_strict dependency-order check reports 0 violations. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Agent: aprender-3d * evidence(pr-review #4587): Sonnet receipt for code head 1028fdf9f9 (+ publish-order leaves; BLOCK: PR1-MUTANTS until the table runs) Agent: aprender-88 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(build): watch only paths that exist, so cargo stops rebuilding every invocation (PMAT-4616) cargo-mutants on #4614 (run 36520343459) reported 9 TIMEOUTs on the D5 functions. The test phase was compiling rather than hanging: 4m18s of the 300s budget went to rebuilding aprender-compute, aprender-serve, aprender-present-core and aprender-present-terminal. The baseline suite takes 32s. Each of those build scripts emits `rerun-if-changed` on a path that does not exist: - the pre-monorepo sibling ../../../provable-contracts/... for compute, present-core, and serve's binding, architecture-requirements and tensor-names reads; - `.git-sha` in build_sha::emit, which no crate in the tree carries. Cargo reruns a build script whose watched path is missing on every invocation, so each `cargo test` after the build rebuilt those crates and everything above them. Each of these sites now emits the watch only when the path exists. The cost: a sibling file that appears later is not noticed until the script reruns for another reason. Before, the script reran on every invocation. The in-tree reads (arch-constraints, contracts/) and the git HEAD triggers are unchanged. Acceptance criterion 4 of PMAT-4616 records this. #4369 moves the bindings in-tree but leaves serve's two sibling reads and `.git-sha`, so it does not fix this alone. Agent: aprender-2e-serve Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(mutants-table): main.rs group runs the pv binary tests, not --lib (TIMEOUT = survived) main.rs is bin code; --lib tests cannot observe it, and the --lib suite ran past the 900 s mutant budget on intel, so 'replace main with ()' read TIMEOUT, which the checker counts as a survivor. binary:: in cli_integration execs pv: both main.rs mutants caught on intel. Timeout unchanged. Agent: aprender-88 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(mutants-gate): NOT_MEASURED is RED with no local-receipt bypass (operator E1) (PMAT-4621) A committed "CAUGHT <line>" text file could turn a mutant CI never built green (e6 requorum finding). Operator ruling E1 (2026-09-29 07:55Z): cuda-gated mutants are measured only by a CI shard on a CUDA runner, and not_measured > 0 on changed gated code is RED. The gate now lists each NOT_MEASURED mutant, prints the count and fails; the self-test plants a receipt and a mutant that honours it is killed. Agent: aprender-ec Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(mutants-table): scratch copies keep .git (--copy-vcs true); every baseline failed without it Run 36538406766 on 8a3abba725: every default-run baseline FAILED on lint::tests::lint_passes_on_real_contracts ('no BASE to compare with: neither merge-base(HEAD, origin/main) nor origin/main resolves'): cargo-mutants copies the tree without .git, so 3527 mutants were never tested (MISSING -> RED). --copy-vcs true on the default and feature-group runs gives the copy the checkout's .git (fetch-depth 0 + the fetched base), as workspace-test has. Intel: aprender-contracts baseline 'ok' with it. Nothing is skipped or relaxed. Proof: the fake cargo in check_mutants_feature_groups.sh dies on any run without --copy-vcs; removing it from either call turns every case DEAD (checked both). Agent: aprender-88 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * evidence(pr-review): #4614 receipt for 8467f2bcb7 (PMAT-4616) Non-author review receipt for the head carrying the build.rs fix; quorum r5 3/3 PASS. Agent: aprender-2e-serve Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * ci(mutants): cuda-gated mutants are NOT_MEASURED on the CPU gate and measured on yoga (#4621, operator E1) cargo-mutants never compiles cfg(feature="cuda") code at default features, so those mutants were invisible. The CPU mutants section now defers the sorted NOT_MEASURED set (mutants_diff_gate.sh --defer-unbuilt); the new mutants-cuda section in the yoga job mutates the diff's feature-gated files at the head sha (--features aprender-serve/cuda --gated-only); the gate job is RED unless the two set shas match (GATE-MUTANTS-CUDA-RULE). Missing/failed/skipped = RED. Guard: scripts/check_ci_gate_mutants_cuda_rule.sh (13-row table, 5 planted weaker rules all RED). mutants_diff_gate.sh self-test PASS. Quorum r3 fixes: PMAT-4621.yaml cites the mutants gate (cop ruling 08:32Z); the review SARIF no longer claims graph_may_run cannot be mutated on intel (host-side, 4/4 CAUGHT with --features aprender-serve/cuda); findings_ref sha256 rebound. Refs #4621 Agent: aprender-ec Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test(cuda): model-free kill tests for qwen35_dequant_f32 and prefill_all_layers_gpu (#4621) The mutants-cuda shard's GPU runner has no TinyLlama, and these two fns were killed only by model-backed suites, so their mutants would survive there. - qwen35_dequant_f32: an F32 weight is returned as it is (kills Ok(0)/Ok(1)); an F16 [2x64] weight dequantizes on the device bit-exact. - prefill_all_layers_gpu: a wrong embedding length and an uninitialized workspace are refused; S=0 is a no-op (kills the Ok(()) body). Without a device they skip; under mutation a skip survives, so the shard is RED, never vacuously green (L25). Built on intel with --features cuda. Refs #4621 Agent: aprender-ec Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * chore(regen): single-writer regen on PR-2 — README contract_count 1889, ont-ratchet, contracts.nt, roadmap Generated paths only (batch_fold.sh regen recipe + make ont-ratchet): - README.md frontmatter/blocks: 1837 -> 1889 (= census.json n_files; PV-ONT-012) - contracts/lint-baseline.json verified_commands re-measured (F-33 / PV-ONT-013) - contracts/contracts.nt via pv extract contracts - docs/roadmaps/roadmap.yaml via make roadmap-aggregate Fixed point: readme_sync --check, pv extract contracts --check, roadmap-aggregate-check all 0; pv lint contracts: 0 errors. Pmat-Ticket: ONT-10 Agent: aprender-98 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(mutants): the feature shard passed --lib to cargo test twice (#4621) yoga run 36547274574: `cargo test ... --lib --lib -- --exact ...` is a cargo usage error, so mutants-cuda tested 0 of the 7 listed mutants. The gate failed closed ("tested 0 mutant(s) but listed 7"), but no mutant was measured. --lib already reaches cargo test through the test args; drop the --cargo-arg=--lib copy. Self-test: gated-scope-and-measured-set now also requires --lib at most once per invocation; planted mutant lib-passed-twice (the old line) is killed. bashrs: 0 errors. Refs #4621 Agent: aprender-ec Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test(ont4b): name the nine ONT-10 rest shapes — shapes_n 57 -> 66 PR-2 adds eight binary-<bin>-surface shapes (apr, apr-qa, aprender-data, aprender-present, aprender-ptop, aprender-simulate, aprender-train-lora, aprender-train-shell) and binary-aprender-apr-http-mcp. Measured: pv extract contracts --check reports shapes_n 66 on this branch. Pmat-Ticket: ONT-10 Agent: aprender-07 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test(ont4c5): follow #4590's de-claim — B names qwen3-8b@lambda, not qwen3-1.7b@gx10 #4590 (ca2daad7db) de-claims Qwen3 on gx10 through data, so the measured domain D no longer holds qwen3-1.7b-q4km@gx10 and gains qwen3-8b-q4km@lambda. The probe's owed set B and the deleted-row mutation still named the de-claimed cell: the_rows_probe_is_green_on_this_repository went red (B ⊄ D) and a_deleted_current_release_row_rejects_naming_the_cell got exit 0. B now equals the measured D; the mutation deletes a gx10 row that is still owed (qwen2-1.5b-q4km). Pmat-Ticket: ONT-10 Agent: aprender-07 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test(merge 2aa46a19df): undo two union-merge defects in contracts-cli tests - pv_surface_gate: 483cc76d4c and 7c418eb75c each added a VS-COUNT-001 Case; the merge kept both, so CASES has 27 rows over 26 distinct rules and gate_self_test_rule_output_depends_on_input (:1136) can never reach 27. Keep the first row. - pvl_discharge_leanchecker: 7c418eb75c made `check()` unscoped while e85da1a039 added check_unscoped() plus a row that needs the SCOPED path; the merge kept both, so scoped_without_user_systemd_declines_and_never_runs_the_checker ran the stub unscoped. That row now calls check_scoped(). Neither test file's defect exists on main (the lean test is absent there; main's pv_surface_gate has no VS-COUNT-001 row). Pmat-Ticket: ONT-10 Agent: aprender-07 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(ci): PR-1 run 36543136454 guard-tree + pr-review-sign reds (PR-introduced) guard-tree, three guards over code this PR adds: - check_guard_steps_isolated: ci.yml's survivor-table step ran three guards outside `setsid --wait` (#4133). Prefixed; check_mutants_feature_groups.sh moves to its own step so it stays bare-wired and guard_tree still runs it. - check_ci_gate_mutants_table_rule [run]: the checker's case table needs minisign, which the sovereign-ci image (guard-tree) lacks -> ENV -> rc=1. The ci.yml step installs minisign first; it now names its input (env GATE_RULE_WORKFLOW, which the script reads), so guard_tree rule 1 leaves the guard to that step. Without minisign it is still rc=1 there (measured: PATH=/usr/bin:/bin -> rc=1). - check_no_silent_truncation: mutants_survivor_table.py:360 printed errs[:2]; it prints every error now. pr-review-sign: three unsigned receipts (heads 05c9e6f2d9, 38cabdb8af, 1028fdf9f9; all verdict BLOCK) carried a hand-supplied diff_patch_id the signer refuses by design (bc7b4296/8b8c5880/4accd62e vs computed d53bb0c1/789241f1/198c753d over the same base..head). No tool writes that field; the field is dropped so the signer stamps the computed one. All 7 unsigned receipts sign+verify locally under the test key. A BLOCK receipt approves nothing, so re-binding it cannot launder an approval. Local: guard_steps_isolated, no_silent_truncation, guards_are_wired, ci_gate_mutants_table_rule (+ --self-test via the step), mutants_feature_groups all rc=0; guard_tree --dry-run: 139 run / 35 skipped, rule guard "wired-with-args in ci.yml". Pmat-Ticket: PMAT-4502 Agent: aprender-07 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(mutants): --lib must reach the feature shard's build, once (#4621) yoga run 36553822552: 451938d114 dropped --cargo-arg=--lib, so --lib no longer reached cargo-mutants' baseline (`cargo test --no-run`, which takes only --cargo-arg). Every cuda-gated integration target was compiled, and tests/driver_cuda_gguf.rs does not build under --features cuda on main (GGUFConfig initializers lack query_pre_attn_scalar, E0063 x10). Baseline FAILED, 0/7 tested, and the gate failed closed. Under --features, pass --lib once, as --cargo-arg (it reaches the build and the test), and drop it from the test args (a second one is the usage error of run 36547274574). Self-test: gated-scope-and-measured-set also requires --cargo-arg=--lib. Mutants lib-passed-twice and lib-missing-from-build (the 451938d114 shape) are both killed. bashrs: 0 errors. Refs #4621 Agent: aprender-ec Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(mutants,roadmap): --cargo-arg --lib spelling for bashrs 7.4.1; regenerate roadmap aggregate (#4621) x86-main on f3077017b1 (job 109390857552), two reds from this branch: 1. check_shell_lint_ratchet grew 6 -> 8 on the runner's bashrs 7.4.1, which reads `--cargo-arg=--lib` as a `--` prefix operator (SC2210, lines 56 and 309; 7.4.2 is clean). Spelled `--cargo-arg --lib`: 0 errors on both. Red/green (cargo-mutants 27.1.0, a crate with a non-compiling integration test): no arg -> "FAILED Unmutated baseline"; both spellings -> baseline `cargo test --no-run --lib`, 4 mutants tested. 2. roadmap.yaml was not regenerated after adding the PMAT-4621 fragment (make roadmap-aggregate; --check idempotent). Self-test PASS, including lib-passed-twice and lib-missing-from-build. Refs #4621 Agent: aprender-ec Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * roadmap(#4502 PR-1): carry the entries for PR-1's own rows The split left PR-1's roadmap entries in PR-2, so a Q0 round on PR-1 had no ticket to judge its rows against and every lane failed it on scope. Bytes are identical to PR-2's copies, so PR-2's merge of PR-1 adds nothing. Pmat-Ticket: PMAT-4502 Agent: aprender-07 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test(ontology): a glob re-export cycle reports the real miss, not the cycle Kills the MISSED mutant from #4588 run 36572203357 shard 9: code.rs:628:27, guard e.reason.contains("re-export cycle") -> false. Under the mutant the cycle reason overwrites the last real miss and the new test goes RED. Pmat-Ticket: PMAT-4502 Agent: aprender-07 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * roadmap(#4502 PR-1): rows for folded #4604 and #4369; regenerate the aggregate C78(2): PR-1 folds the pv 0.69.4 release set (#4604) and the build_helper::env_key migration (#4369); each now has its entry. fad4fb0453 added six entries without regenerating roadmap.yaml (check_roadmap_fragment_required DRIFT); regenerated here. Pmat-Ticket: PMAT-4502 Agent: aprender-07 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * roadmap(#4502 PR-1): AC rows for the #4587 A' survivor table and the folded-ticket list (C78(2)) Answers the Q0 gemini objection ci.yml:51 (survivor table not asked for by PMAT-4502): it is the operator's 06:36Z ruling on this PR, now an AC row. Makefile label/lint ratchets are PMAT-4139/4166, carried with their own entries. Pmat-Ticket: PMAT-4502 Agent: aprender-07 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * roadmap: rows for the #4083 (EV-9) and #4189 (G5) folds in PR-2 (C78, cop ruling a) Pmat-Ticket: PMAT-4502 Agent: aprender-07 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(#3715 B1): receipt §8 describes THIS diff; README/CHANGELOG say 'planned for 0.70.1', not 'fixed' (quorum round 2) Agent: apr-0d-b1 Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * census: drop the batch-regen census edit; README contract_count back to 1837 (cop ruling a, #3569) census.json has one writer (the release train); it is regenerated once after the batch merges. contracts.nt re-extracted (readme:contractCount 1837); pv extract --check fixed point holds; check_census_derived.sh PASS. Pmat-Ticket: PMAT-4502 Agent: aprender-07 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * evidence(#3715 B1): Q0 quorum AGREED at 33b032506a (opus-5-5, haiku-4-5, gemini-3.1-pro-high; base split/ont10-rest-0.70) Agent: apr-0d-b1 Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * revert(#4588 split): #4429 x86-main intel pin and #3668 docs-only skip move to PR-2b Cop ruling 19:35Z (split). Gate-graph walk from rc.1 (rc-cut.yml: ci/gate + workspace-test; gate.needs closure) and FINAL (paiml-ontology @60179ae2:714, ONT-10 depends_on): neither row is needed, directly or transitively. - x86-main: runs-on back to main's [self-hosted, Linux, X64, clean-room]. The gate needs the job, not the runner label. - #3668: the `changes` job, guard-cargo's needs/if, the gate's GATE-DOCS-ONLY-RULE, ci_docs_only.sh, check_ci_gate_docs_only_rule.sh, fat_driver's outputs pass-through and the PMAT-3668 roadmap entry are removed. The three guard steps return to guard-cargo. The gate again reads guard-cargo and determinism-compare strictly, as main does, so no gate is loosened versus main. Receipts: fat_driver self-test 59/59; check_ci_gate_mutants_rule, check_guards_are_wired, check_ci_fat_secrets_plumbed, check_ci_gate_mutants_table_rule rc 0; actionlint 2 findings = HEAD's 2; make roadmap-aggregate-check ok. Pmat-Ticket: PMAT-4502 Agent: aprender-07 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * audit(PMAT-4616): Q0 quorum for a8e97a4711 — 3/3 PASS (sonnet-5, gemini-3.1-pro-high, haiku-4-5) Agent: apr-2e-d5 Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * docs(#4588): receipts for 4502/4073/3712/4445/4083; PMAT-4083 scoped to part 1 (advisory) Cargo-test rows are NOT_MEASURED (intel only) and are not counted as passes. Pmat-Ticket: PMAT-4502 Agent: aprender-07 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(ci): the fat driver's results output dropped section outputs — the NOT_MEASURED gate rule read "" and could never be RED (#4621) emit_results_output wrote {result, continue_on_error} only, so the gate job's GATE-MUTANTS-CUDA-RULE (which reads .<section>.outputs.not_measured / measured_sha) saw an empty string on every run and took the "no deferred set" branch: a fail-open gate. Found by the non-author pr-review of d07e187b66 (F1); its own case table passed because its fixture JSON hand-wrote `outputs`. The results output now carries each section's outputs, a fat_driver self-test row pins it, and check_ci_gate_mutants_cuda_rule.sh gains a row that feeds the rule the driver's REAL emission (RED with 3 deferred and no shard). Mutating the driver back turns both rows RED (52/53 and row 14 FAIL). Refs #4621 Agent: aprender-ec Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * evidence(pr-review): #4614 receipt for 12c0e28a00 (PMAT-4616), DEGRADED (mutation unreachable), unsigned pending CI signer Agent: apr-2e-d5 Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * docs(#4588): cells-producer receipt row measured (rc 0, 25 ok) Pmat-Ticket: PMAT-4502 Agent: aprender-07 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(mutants): --defer-unbuilt file's directory is made before it is written (#4621) x86-main went red on run 36607873893: 'mutants-gate/not_measured.txt: No such file or directory'. The gate truncated the deferred file before mkdir -p of --out, and CI hands it a path in a directory a fresh checkout does not have. Red/green: a new self-test row (defer into a missing dir) and a mutant that drops the mkdir; the mutant is killed, the self-test passes 58 ok, bashrs 0 errors. Agent: aprender-ec Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * chore(review): #4620 quorum artifact (r4, 98d7ffb1c7) + pr-review receipt for the reviewed head Quorum r4: 3/3 PASS but degraded same-family (agy quota exhausted, 429); r3 on fc88268134 was full Q0 3/3 (gemini-3.1-pro-high + sonnet-5 + haiku-4-5). Receipt is unsigned; CI signs it. Agent: aprender-ec Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * revert(cb200): drop the 599->617 re-baseline row from #4588 (cop ruling 20:31Z) A raised stored limit is a gate waiver. Head vs base on the same scanner and run measures 621 vs 622, so there are no new findings to fix; the row is only the limit raise and goes. Pmat-Ticket: PMAT-4502 Agent: aprender-07 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(CB-200): head-vs-base release gate + clear the 4 findings over v0.69.3 dogfood.sh judged CB-200 against a stored baseline (599); main measured 622. Now: on Fail, compare head to the newest final tag (v0.69.3 = 618), one pmat, one run, baseline neutralised (scripts/check_cb200_head_vs_base.sh). Refactored the 4 added definitions (check_valid_under, coverage_serve_shards.main, get_one_patchid, nightly_manifest.main) with behaviour-identical splits; golden/differential proofs in receipt. Agent: apr-0d-b1 Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * fix(CB-200): split get_one_patchid state into helpers (behaviour-identical) Agent: apr-0d-b1 Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * test(contracts): kill the MISSED mutants CI's survivor table found on #4587 mutants-shard 5/9/10/12 reported MISSED in bindings_gate, evidence_gate, ratchet_gates, refines_gate, shapes_gate, extract/binary, extract/code, witness_small and scoring/codebase. Each test pins the behaviour a surviving mutant changed. aprender-contracts…
… on main (bed17c8) (#4656) * chore(contracts): discharge-summary.json regenerated by pv discharge run at db86acf (C268#2a) In-tree pv (cargo build --locked -p aprender-contracts-cli --bin pv), run as `pv discharge run crates/aprender-contracts-staging/lean` under systemd-run --user --scope -p MemoryMax=24G, on the ladder's lean cache link (bfbabfa35e3e25aa): rc 0, 423 s, memory.peak 25769803776 (= the cap), oom_kill 0. Only tree_sha moves (a31602f -> bed17c8); byte-identical (cmp) to 14a3b1b3a4's file. Agent: aprender-27 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * evidence(pr-review): #4656 review receipt for 492dcd3 Agent: aprender-27 Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * evidence(pr-review): #4656 receipt — index_commit/ancestry re-derived for this PR (A4 B1) Agent: aprender-27 Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * fix(ci): quick-tier test steps never saw the pinned origin/main — git refused the root-owned mount (#4656) Step 8 pins the event's base as refs/remotes/origin/main, and #4655 gave the full-tier "Workspace lib tests" container safe.directory=/workspace so git, as root over the host-owned checkout, would read it. The two quick-tier test containers (selected crates; tree readers) were left without it. A quick-tier PR (#4656, run 37052925185, step 21) therefore ran lint_passes_on_real_contracts with git refusing /workspace ("dubious ownership"): origin/main did not resolve, the refinement gate skipped, and the CI assert failed 3/3. #4655 passed only because it ran the full tier. Red/green (depth-1 clone, base pinned exactly as step 8 does, git as root in a container over the host-owned mount): without the env: rc=128 "detected dubious ownership in repository at '/workspace'" with the env: 989cb01, rc=0 Test level, same clone, localhost:5000/sovereign-ci:stable as root, CI=true, cargo test -p aprender-contracts --lib lint::tests::lint_passes_on_real_contracts: without the env: rc=101, "the refinement gate declined ... Skipped { reason: \"no BASE to compare with ...\" }" (the exact CI failure) with the env: rc=0, 1 passed cargo test -p aprender-contracts --lib (intel): 2289 passed, 0 failed. Tightens: the refinement gate now RUNS on quick-tier PRs instead of being skipped. No gate is removed or loosened. Agent: aprender-27 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * evidence(pr-review): #4656 review receipt for bd66e7c (quick-tier safe.directory) Agent: aprender-27 Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * evidence(pr-review): #4656 receipt — S3.E antigravity consulted (A4 B1) Agent: aprender-27 Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * fix(deps): bump wasmtime 48.0.3 -> 48.0.4 (RUSTSEC-2026-0325/0326/0327) cargo audit in x86-main sov.security (CI 37061601053, #4656 @ ab68432) went red on three wasmtime advisories published 2026-10-02. Fixed in >=48.0.4,<49.0.0. Tool-regenerated: `cargo update -p wasmtime --precise 48.0.4`, Cargo.lock only, no hand edits (C265#10, C269#3). Receipts (intel, scratch clone at ab68432 + this lock, byte-identical sha256 5a8e7b32fcac4fec on lambda and intel): - cargo audit: rc=0, 0 hits for RUSTSEC-2026-032[567] - cargo deny check advisories: rc=0, "advisories ok" - wasmtime is reached only via the optional aprender-test-lib/runtime feature; that feature already fails on the 1.93.0 pin at base (48.0.3 -> cranelift requires rustc 1.95.0), so unchanged by this bump. Agent: aprender-27 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * evidence(pr-review): #4656 receipt for 98e7256 (wasmtime 48.0.4, RUSTSEC-2026-0325/6/7) Agent: aprender-27 Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * evidence(pr-review): #4656 receipt 98e7256 — stamp diff_patch_id (A2 binding) diff_patch_id was null. Arm 4 computes the branch diff 989cb01..848a584 (evidence/pr-review/4656 excluded) as 4ac8c02f02a1fbb5a10c8f3257c7cf70fbca74c6; prpid_compute over 989cb01..98e7256 gives the same id. Only that field changes. Agent: aprender-27 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…se got binaries (#4657) (#4658) * fix(binary-release): assets pre-check died under bash -e on missing assets, so no release got binaries The `assets` job ran `check_release_assets.sh "$TAG"; rc=$?` in the default `bash -e` step shell. Exit 1 (assets missing: every fresh release) killed the step before the `case`, the job failed, and every build lane was skipped. v0.70.0-rc.1 (run 37034899163) shipped zero assets, so install.sh row 12 (`--channel rc`) 404s (run 37040341799). The final v0.70.0 would hit the same. Now `rc=0; ... || rc=$?`. Case table under bash -e with a stub script: exit 0: before present=true after present=true exit 1: before STEP FAILS after present=false (builds run) exit 2: before step fails after step fails (::error:: refusal kept) No gate weakened: exit-2 refusal and verify-apr-assets unchanged. actionlint findings identical head vs base (2 pre-existing SC2016 info). Agent: aprender-27 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * review(4658): receipt for df432eb Agent: aprender-27 Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…rip found no apr (#4659) (#4660) * fix(binary-release): mini lane built into a redirected target dir, so strip found no apr (#4659) Run 37084403773 failed on both attempts: cargo finished into the runner's redirected CARGO_TARGET_DIR, and strip/version/package all read ./target. Pin CARGO_TARGET_DIR to the workspace target for the build step. No gate changed. Agent: aprender-27 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * review(4660): receipt for 1b61c42 Agent: aprender-27 Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Guard-step files (#4415) take main's side: main carries the later ci/sections.yml form of the same work, and this branch's late fixes (jq stream death, zombie child, no-manifest step) are all present there. ci/sections.yml takes main's live-API PR-body reads (#3298) and keeps its run-all guard step. tests/pr-review.bats keeps both new blocks (L2 attest and the #3594 jq rows). roadmap.yaml is the union. The mutation set is now 253 (this branch's 249 plus main's four); its kill count stays pending until the receipt job's Arm 3 runs. Refs #4472 #4503 #4517 #4533 Agent: aprender-78 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Contributor
Author
|
Closed for the open-PR cap: this changes what a gate accepts and waits for the operator's sign-off batch (0.70.2 ledger). Branch kept; reopens on a yes. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fold of like-kind PR-review / CI-guard PRs (PRCAP, k=4, fold = MOVE). Built on
origin/main15f1b2a withgit merge --no-ffper source, then one merge of main. Every source branch is kept.c208d41e504cfe2485e43200765b19712e58dd67Proof:
git merge-base --is-ancestor <head> 6227d1fd6aexits 0 for all four.Merge notes
scripts/mutate-guard.sh --list= 249..github/workflows/pr-review-quorum.yml: the Arm 4 step env block is the union of both sides. The fold side gavePR_REVIEW_ATTEST_ROOT: ${{ steps.attest.outputs.root }}and main (fix(ci): plumb secrets.PR_REVIEW_SIGNING_KEY_B64 under its real name; guard section secrets vs plumbing (#4510) #4512/B1) gaveARM4_PR_HEAD_SHA+GH_TOKEN. Count of those lines: main 2, fold side 1, HEAD 3.docs/roadmaps/roadmap.yaml: main aggregates it fromdocs/roadmaps/entries/. feat(ci): one diff classifier + a docs tier for the pr-review quorum (#4472) #4503's inline PMAT-4472 entry was moved verbatim intodocs/roadmaps/entries/PMAT-4472.yaml, then the file was regenerated withscripts/lib/roadmap_fragments.py aggregate --write.--checkis ok and the output is idempotent.ci/sections.yml, so it no longer touches ci.yml. 11/11 of its files are byte-identical to its head.Agent:trailer (L18). It is pushed and will not be rewritten. It is authored by aprender-59.Red/green evidence (workflow edit)
bash scripts/check_pr_review_arm4.sh --self-testbash scripts/check_guard_steps_run_all.shbash scripts/check_pr_review_wiring.shbash scripts/check_workflow_env_defined.shbash scripts/tests/guard_tree_job_test.shcargo fmt --all -- --checkCoverage gap, stated plainly: I mutated the workflow by deleting the
PR_REVIEW_ATTEST_ROOTline, then separately theARM4_PR_HEAD_SHAline. Bothcheck_pr_review_wiring.shandcheck_workflow_env_defined.shstayed green (rc=0). The script reads both with${VAR:-}defaults, and its self-test sets them itself. So the behaviour has a red/green (the 33-row table), but the wiring of these two env lines is not guarded by anything. That was true on main before this fold. It is reported as a finding and is not fixed here, because fold = MOVE.Touches .github/workflows → needs a 3/3 non-Claude quorum, or a labelled claude-only quorum with seat ≠ opus. Per the cop ruling, do not arm until PR-1 (#4587) is in the merge queue.
Agent: aprender-59
🤖 Generated with Claude Code