Skip to content

fold(pr-review-ci): #4503 + #4533 + #4517 — diff classifier/docs tier, FLOW-003 v2.1, fork attest - #4605

Closed
noahgift wants to merge 43 commits into
mainfrom
fold/pr-review-ci
Closed

noahgift wants to merge 43 commits into
mainfrom
fold/pr-review-ci

Conversation

@noahgift

@noahgift noahgift commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Fold of like-kind PR-review / CI-guard PRs (PRCAP, k=4, fold = MOVE). Built on origin/main 15f1b2a with git merge --no-ff per source, then one merge of main. Every source branch is kept.

Source PR Head Ancestor of fold head
#4503 feat(ci): one diff classifier + a docs tier for the pr-review quorum (#4472) c208d41e50 yes
#4533 FLOW-003 v2.1: fleet hardware + tiered testing, Parts II–III review, Props 19–20, pv contracts (docs only) 4cfe2485e4 yes
#4517 feat(pr-review): pr-review:attest label path for fork PRs (#4462) 3200765b19 yes
#4457 fix(ci): run every guard step and report every failure (#4415) 712e58dd67 yes

Proof: git merge-base --is-ancestor <head> 6227d1fd6a exits 0 for all four.

Merge notes

Red/green evidence (workflow edit)

Check on head 6227d1f Result
bash scripts/check_pr_review_arm4.sh --self-test 33/33 rows + 2 fingerprint checks, both polarities
bash scripts/check_guard_steps_run_all.sh ok
bash scripts/check_pr_review_wiring.sh PASS
bash scripts/check_workflow_env_defined.sh OK, 24 workflows
bash scripts/tests/guard_tree_job_test.sh 8 checks, 0 failed
cargo fmt --all -- --check ok

Coverage gap, stated plainly: I mutated the workflow by deleting the PR_REVIEW_ATTEST_ROOT line, then separately the ARM4_PR_HEAD_SHA line. Both check_pr_review_wiring.sh and check_workflow_env_defined.sh stayed green (rc=0). The script reads both with ${VAR:-} defaults, and its self-test sets them itself. So the behaviour has a red/green (the 33-row table), but the wiring of these two env lines is not guarded by anything. That was true on main before this fold. It is reported as a finding and is not fixed here, because fold = MOVE.

Touches .github/workflows → needs a 3/3 non-Claude quorum, or a labelled claude-only quorum with seat ≠ opus. Per the cop ruling, do not arm until PR-1 (#4587) is in the merge queue.

Agent: aprender-59

🤖 Generated with Claude Code

noahgift and others added 27 commits September 25, 2026 14:56
…4415)

On the 0.69.4 RC PR #4318, every red guard job reported exactly one failure.
GitHub Actions skips a job's later steps once one fails, so guard-cargo skipped
15-71 of its 77 steps (median 55). Each further red was found only after
another 35-50 min CI cycle: 02d4618 -> cef1b12 -> 696576e -> 6281d9f.

- ci.yml: the last setup step of guard-cargo and guard-tree gets
  `id: guard-setup`. Every guard step after it gets
  `if: ${{ !cancelled() && steps.guard-setup.outcome == 'success' }}`.
  All guards run and the job stays red if any fails. If setup failed, the
  guards skip instead of painting dozens of meaningless red rows. The diff
  is if/id keys only; that was verified by a YAML-level comparison of every job.
- scripts/lib/ci_guard_steps.py: reads a guard job's steps from ci.yml.
  `check-run-all` counts fail-fast guard steps, `run` executes the same run:
  blocks locally without stopping at the first red (the #4416 v0 seed),
  and `list` prints the steps.
- scripts/check_guard_steps_run_all.sh: a shrink-only ratchet at 0
  (scripts/guard_fail_fast_baseline.txt), so a new guard step cannot be
  fail-fast. 11-row case table; 5/5 mutants killed. It is wired through
  guard_tree.sh --no-cargo.
- scripts/ci_guards_local.sh: the local entry point (#4416, handed to
  aprender-cb). Its first local run caught a missing tool_version header
  in this PR's own baseline file.

Operator ruling Y1 (relayed by the cop, 2026-09-25): guards use !cancelled()
and the job stays red.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… script, CI and local (#4415, #4416)

GitHub stops a job at its first red step. On RC PR #4318 every red guard
job reported ONE failure; guard-cargo skipped 15-71 of its 77 steps.

Operator amendment A4: guard-cargo and guard-tree each run ONE step,
`bash scripts/ci_guards.sh <job>`, and `make guards-local` + the pre-push
hook run the same file. Both print `ci_guards: sha256 <runner> manifest
<steps>`, so "CI and local ran the same guards" is a line comparison.

- The guard steps stay VERBATIM in ci.yml, moved into `guard-cargo-steps`
  / `guard-tree-steps` manifest jobs (`if: false`, never run by GitHub),
  so every guard that greps ci.yml for an invocation still finds it.
  Transform verified by YAML equality: moved steps == originals minus
  `if:`, every other job and the top level unchanged.
- scripts/lib/ci_guard_steps.py runs the manifest: bash -eo pipefail per
  step, GITHUB_PATH/GITHUB_ENV propagated between steps, github.token /
  runner.temp resolved (anything else refused), per-step timeout with a
  process-group kill, ::error per failure, step-summary table, exit 2 on
  a vacuous run.
- check_guard_steps_run_all.sh: manifest must be `if: false`, steps are
  name/run/env only, only resolvable ${{ }}, the job must call the runner
  and have `id: guard-setup` (else the runner's if: skips it and the job
  is green having run nothing), no orphan manifest. 25/25 rows; 10/10
  non-equivalent mutants of the runner killed.
- guard_tree_test / guard_tree_job_test read the manifest too; new mutant
  (cargo planted in guard-tree-steps) is RED under the new leg and was
  GREEN under the old one.
- Absorbed from aprender-cb's add2fd4 (#4416): check-coverage + ack
  file (2 rows stale on main, removed), Makefile guards-local, pre-push
  block. scripts/ci_guards_local.sh -> scripts/ci_guards.sh.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
#4415)

Quorum lane b on 15883bb:
- usage() printed lines 2-31, so --help never said "self-test", and
  guard_tree.sh's advertises_self_test never dispatched the case table.
  usage() now prints the whole header, up to `set -uo pipefail`.
  Verified: the old file is NO, the new one YES.
- ci_guards.sh --check-coverage ran only in make guards-local and the
  pre-push hook. The main path now also runs check-coverage, so a guard
  step outside a manifest turns guard-tree red. Mutant: removing one
  ack row gives rc=1.
- A new self-test row: a literal ${{ }} inside run: turns it RED (26/26).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…in check, rc 2 in CI run (#4415)

Lane b (sonnet-5) reproduced it: with the guard-tree-steps job deleted, the
leftover steps behind the runner counted as 'ran', so ran==0 never tripped.
There is a new fixture variant, no_manifest, and a mutant dropping the check
turns it RED. 27/27.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ader refuses what it cannot read (#4415)

The self-test now calls scripts/lib/ci_guard_steps.sh (not yet written) and
adds 14 reader rows: 10 refusals (anchor, alias, tag, folded, multi-line plain,
unclosed quote, duplicate key, tab indent, flow map, colon in a plain scalar),
each beside a passing control, plus quoted-scalar decoding, a flow-list gate
needs:, and no python on the lib or runner path.

Pmat-Ticket: PMAT-4415

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…leaves the runner path (#4415)

scripts/lib/ci_guard_steps.py -> scripts/lib/ci_guard_steps.sh, with the
YAML reader in ci_guard_yaml.awk and the jq defs in ci_guard_steps.jq (both
hashed into the `ci_guards: sha256` line). yq is not on the clean-room fleet,
so the reader is awk: it reads the block-YAML subset ci.yml is written in and
REFUSES (rc 2) anchors, aliases, tags, folded >, multi-line scalars, flow maps,
duplicate keys and tab indentation.

Measured on this tree:
- case table 41/41 (was 1 ok / 40 FAIL at RED), under gawk AND mawk
- reader vs PyYAML on ci.yml .jobs: byte-identical JSON (18 jobs), gawk + mawk
- list / check-coverage / check-run-all: byte-identical to the .py at 3fb4514
- 3 reader mutants (dup-key, backslash escape, tab) each turn a row RED
- bashrs lint: 0 errors (warnings are jq/awk text inside strings)

ci.yml is unchanged: it never installed PyYAML and never named the .py.

Pmat-Ticket: PMAT-4415

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…in an agy lane sandbox (#4415)

The self-test killed a hung step and then asked `kill -0 child` 0.2 s later.
kill -0 answers for a zombie, so a killed child that its reaper has not yet
collected read as "still alive", and the row went RED in a quorum lane's
sandbox while passing 5/5 here. Read the child's state with ps instead
(gone or Z = dead) and give the reaper up to 2 s. The mutant with a
really escaped child (setsid sleep) still goes RED; table 41/41.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… rc 0

The run loop read jq's step records through <(...), whose exit status nothing
checks: a jq error after the first record silently dropped every later step and
the run still reported success. Records now go to a scratch file first, and a
nonzero jq stops the run with rc 2. Case row 42 injects a jq error on step 2;
it is RED on the old form and GREEN now (42/42).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ner-attest)

A fork PR's own runs get no secrets, so it could never carry a receipt signed
with PR_REVIEW_SIGNING_KEY_B64 and was unmergeable by construction.

pr-review-fork-attest.yml (pull_request_target, labeled) authorizes the
labeler server-side (admin/maintain/write; a triage labeler reads as `read`),
refuses a self-label and a same-repo head, reads the head as git objects only,
builds an L2-maintainer-attest receipt (verdict DEGRADED, no consultations),
signs it with the base secret, and publishes it fast-forward to the
base-owned branch pr-review-fork-receipts. Arm 4 reads L2 receipts only from
that branch (PR_REVIEW_ATTEST_ROOT), prints "DEGRADED: maintainer-attest by
<login>", and the signed diff patch-id voids the attest on any later push.

RED proofs (cop conditions): author self-label, triage-only labeler, stale
patch-id, receipt on a non-base branch — each a case-table row that goes RED,
each mutant verified killed. Guard fix found on the way: jq `inside()` on a
string is substring containment ("" and "rite" passed); now IN(...).

Refs #4462

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…4472)

Docs-only PRs paid full CI and a full quorum (evidence: #4467, README + two
book pages). This is PR 1 of 2 and touches no workflow file.

- scripts/ci/diff_class.sh: ONE classifier (class=docs|code|empty,
  subclass=ledger|prose). Docs = root *.md, book/**/*.md, docs/**, each present
  at head; a deleted doc, crate README, .claude/, contracts/, evidence/ and any
  config is code. --no-renames. Self-test 18/18, all 5 mutants RED.
- scripts/ci_test_tier.sh: docs_only() now asks the classifier (subclass=ledger,
  same #3658 set as before); stale self-test row repointed at ci/sections.yml,
  where #4471 moved the touched-list line.
- check_pr_review_receipt.sh S3.E docs tier: antigravity may be not-triggered
  only when class=docs AND docs/BEATS.md untouched AND no added line states a
  comparative ratio AND trigger_reason names the docs tier. The classifier is
  looked up lazily and fails closed.
- Fixtures: row 34 moved to a code diff; rows 44-47 new (SB1 head added after
  K1, no existing SHA moved). reject-53-drop survived until row 47 existed.
  Mutation set 233 -> 241; all 10 new mutants killed. bats 170/170.
- Contract F-PRREV-016 + formula, skill S3.E, spec count amended.

PR 2 (workflow gating of mutants/determinism/sov on class=docs) follows a2's
#4482.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ls closed

agy on 253d2e4: docs/specifications/pp-066-dag.yaml is a CI input, so docs/**
now means docs/**/*.md (plus the #3658 ledger dirs). The ratio scan read the added
lines through < <(...), so a failed git diff passed; it now captures them first and
rejects B1. The new probe (blob removed) goes RED. Mutation set 243. WIP: mutants
reject-51..56 not yet re-run (stopped for the operator budget hold).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…red killed alone

The derived set (mutate-guard.sh --list) is 239. Only the six added by the
L2-attest guard (reject-01..03, drop+flip) were run, each alone against the
full bats oracle: 6/6 killed. The spec's PENDING row says exactly that and
keeps the full-sweep kill count pending on Arm 3.

Refs #4462

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ing, #4517)

The attest-root loop matched on patch-id only. A receipt signed for PR A,
copied under PR B's directory with an identical diff, would arm B. The
signed .predicate.pr must now equal PR_NUMBER. Self-test row added;
mutant (check removed) turns it RED.

Refs #4462

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… false positive)

Refs #4462

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
# Conflicts:
#	.github/workflows/pr-review-quorum.yml
…ng v1.1 fixes, a Parts II–III review (R2-1..R2-14), Prop. 19 and pv contracts

v2.0 said "infra-5a applies its remaining v1.1 fixes on top of this
version". This commit applies them: R-3..R-12 on Part I. It also reviews
Parts II and III. Two soundness findings are among them: the input set
goes stale between the nightly and the base (R2-8), and an openat trace
misses directory listings and ENOENT lookups (R2-9). Doctests drop out of
nextest archives (R2-10), and Prop. 18 measured only what full review
missed (R2-12). New: Prop. 19 (tier routing lowers rho_HOL), §11
(ci-input-set-v1 written out in full), and §12 (the review record).
R2-14: the QM-08 number collides with aprender#4519, flagged for the cop
to re-map (not renumbered here).

Refs #4514 (PMAT-4514)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…-4514 to v2.1

Quorum lane 2 found that v2.0 cut QM-00's mutant list from v1.1's four to
three with no record: a gate weakened silently. Restored as the union (the
1/(1-phi) retry mutant, the Prop. 10 rescued-flake mutant, and v2.0's
double-count mutant) plus one for Prop. 20. The ont:falsifier minCount goes
from 3 to 6. Recorded as R2-15 in §12.2.

Quorum lane 3 found that PMAT-4514's title still says v1.1 only. The cop
put v2.1 under PMAT-4514, so the title in the fragment and the aggregate
now names both.

Refs #4514 (PMAT-4514)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ve-diff guard (PMAT-980) forbids a title edit; the v2.1 scope rides in the PR body and issue #4514

Refs PMAT-4514

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Resolution table:

| path                        | conflict                                   | resolution |
|-----------------------------|--------------------------------------------|------------|
| docs/roadmaps/roadmap.yaml  | both sides appended an entry at the tail: ours PMAT-4472, main's PMAT-4514 | union: both entries kept verbatim, PMAT-4472 then PMAT-4514; YAML parses |
| ci/sections.yml             | auto-merged                                | no manual edit |

Checked on the merged tree:
- diff_class.sh --self-test: 20/20
- ci_test_tier.sh --self-test: 93 ok, rc=0
- check_pr_review_wiring.sh: rc=0
- bats tests/pr-review.bats: 171 ok, rc=0

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-59
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-59
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Conflict resolved: the mutation-set count moves to 249/249, which is
243 from #4503 plus #4517's six L2-attest mutants. Derived with
`scripts/mutate-guard.sh --list` on the merged tree, which prints 249 rows.

Agent: aprender-59
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

§13.11 rung 1 — quorum shadow verdict

S13-SHADOW pr=4605 head=b9476d24d7d2e3e1f03969c2b2858fbf12940ca2 verdict=REFUSE class=Q1 arm_rc=1

Shadow mode: this records a verdict and merges nothing. A refusal
to arm is not a block (§13 adds zero rows to §7) — the pull request is
exactly as green as it was.

noahgift and others added 3 commits September 28, 2026 18:24
…r onto the sections shape

Main moved every CI job body into ci/sections.yml (#4441) and ci.yml's gate now
requires fat jobs, not guard-*. A text merge could not carry #4457 across:

- ci.yml is main's, unchanged. #4457's guard-tree / guard-cargo edits (guard-setup,
  one `bash scripts/ci_guards.sh <job>` step, the `<job>-steps` manifests) are
  applied to ci/sections.yml, with main's `setsid --wait` ported onto the moved
  steps. All guard-setup / ci_guards.sh lines are present (1076-1094, 1663-1702).
- scripts/lib/ci_guard_steps.sh reads ci/sections.yml (its `jobs:` mapping only;
  the header carries flow maps the reader refuses). With no gate job in the file,
  guard jobs are the guard-* jobs that are not `-steps` manifests; a ci.yml-shaped
  file with a gate still uses the gate's needs. New self-test row for the no-gate
  case; repo fixtures follow the path (43/43).
- guard_tree_job_test.sh: main's CI_YML+SECT_YML test plus #4457's legs; the
  no-cargo leg reads guard-tree-steps too, and mutant 2b plants cargo in the
  manifest (8/8, yq and awk modes).
- ci_guards_uncovered.txt: two `gate` rows removed as stale on main (the check
  says so; check_receipt_gate_base_owned.sh now runs inside guard-tree-steps).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- pr-review-quorum.yml: union. The fold's PR_REVIEW_ATTEST_ROOT (#4517) and
  main's B1 ARM4_PR_HEAD_SHA + GH_TOKEN both feed check_pr_review_arm4.sh.
- roadmap.yaml: main now aggregates from entries/, so #4503's inline
  PMAT-4472 entry moves verbatim to docs/roadmaps/entries/PMAT-4472.yaml,
  and roadmap.yaml is regenerated (aggregate --check: ok, idempotent).

Agent: aprender-59
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
noahgift and others added 13 commits September 28, 2026 19:52
… 5 s (FLAKE-0) (#4613)

assertion_failure_reports_nonzero_and_traceback gave exit_code None under
a full `cargo test -p apr-cli --tests` on intel (2/2 runs) and passed
alone (3/3): `assert 1 == 2` was killed by the 5 s deadline while python3
was still starting. The deadline in these tests guards a hang; a program
that exits returns at once, so the terminating-program tests now share
TERMINATING_DEADLINE_SECS = 120, and the success/exit-code asserts print
timed_out so the next failure names its cause.

Agent: aprender-a2

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…ALE (PMAT-4594) (#4598)

* fix(silicon): a short listing that agrees with itself scored two green axes STALE

Run 36410817478 (2026-09-28 11:25Z, #4512) failed check_silicon_coverage.sh
with x86_64-cpu and aarch64-cuda-sm121 "STALE 9d" while Silicon Nightly
36364433328 had carried both 10 hours earlier. The created>= listing of
silicon-nightly.yml answered "21 of 21": total_count agreed with the page,
and the newest 9 runs were not on it. The runs before and after that one
all PASS, and the same guard run live now reads 30 of 30.

The reader cross-checked only an EMPTY listing against the unfiltered
newest run (PMAT-3337 §8). A short one went straight to the verdict. Now
every listing is cross-checked: if the unfiltered newest run is inside the
lookback and newer than anything the filtered listing held, the listing is
read once more, and if it is still short the axis is NO-GO (unknown, rc 2),
never STALE. The lookback, the stale window and every limit are unchanged.

Self-test rows on the real policy lines, the real silicon-nightly.yml and
run 36364433328's four jobs as the API returns them:
  g-real-green    fresh nightlies                    -> ok, rc 0
  h-real-stale    newest nightly 9d old, whole list  -> STALE, rc 1
  i-short-once    11:25Z shape on the first read     -> ok, rc 0
  j-short-always  11:25Z shape on every read         -> NO-GO, rc 2
Against origin/main's reader, i and j go RED (rc 1, STALE): the defect.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-57

* roadmap(PMAT-4594): entry for the silicon short-listing fix

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-57

* evidence(PMAT-4594): quorum verdict (AGREED 3/3)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-57

* evidence(PMAT-4594): pr-review receipt for 9aa1595 (reviewer sonnet-5, FINDINGS advisory-only, mutation 3/3)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-57

* roadmap(PMAT-4594): entry arrives through its fragment (check_roadmap_fragment_required)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-57

* evidence(PMAT-4594): pr-review receipt for 4c7e5c3 (reviewer sonnet-5, DEGRADED: agy lane unreachable)

The fragment commit moved the diff patch-id, so the 9aa1595 receipt no
longer binds. Re-reviewed at the new head by an independent sonnet-5
subagent; self-test 12/12, roadmap guards rc 0.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-57

* fix(ci): pr-review-sign got an empty key — ci.yml fed FAT_SECRET_..._KEY_B from a secret that does not exist

#4441 wired `FAT_SECRET_PR_REVIEW_SIGNING_KEY_B: ${{ secrets.PR_REVIEW_SIGNING_KEY_B }}`.
The repo secret is PR_REVIEW_SIGNING_KEY_B64 and the pr-review-sign section reads
secrets.PR_REVIEW_SIGNING_KEY_B64, so the driver always handed it "" and every
same-repo PR that commits a pr-review receipt failed x86-main ("carries an UNSIGNED
receipt and PR_REVIEW_SIGNING_KEY_B64 is empty"): #4598 job 108952599966, #4587 job
109019972501.

Mechanism: fat_driver self-test gains secret_wiring_gaps() — every secrets.X a
section (ci/sections.yml, vendored sov) reads must arrive as FAT_SECRET_X fed from
secrets.X. Red/green: with the old ci.yml line the self-test exits 1 (row
"secret wiring: every secrets.X ..." BAD); with this fix 52/52.

Refs PMAT-4594

Agent: aprender-e6
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore(pr-review): sign this PR's receipt (PR-REVIEW-SKILL-002 v2 §4.3 CI signer)

* evidence(PMAT-4594): pr-review receipt for a0e0e9c (reviewer sonnet-5, FINDINGS advisory-only)

Fresh quorum after #4512 changed the patch-id lib (receipt rebind, cop
DECIDED-UNATTENDED). diff_patch_id c860ecfc40 under the in-tree lib.
Second vendor agy gemini-3.1-pro-high consulted.

Agent: aprender-e6
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* evidence(PMAT-4594): §3.C CRUX consulted on the a0e0e9c receipt

Arm 4 rejected consultations.crux=not-triggered: the S3.C surface regex
matches 'ToolDefinition' inside a prior receipt's prose in
evidence/pr-review/4598/ (false positive; no CLI/HTTP/MCP/config surface
changes). Consulted and recorded; patch-id unchanged (c860ecfc).

Agent: aprender-e6
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* evidence(PMAT-4594): a0e0e9c receipt — SARIF driver names in the closed vocabulary

Arm 4 A4 rejected the receipt [B1]: runs named pr-review-mutation /
pr-review-reviewer / pr-review-crux are outside { pmat, nvidia-cuda-docs,
crux, cargo-mutants, antigravity }. Renamed per the prior 9aa1595 receipt's
convention: primary-reviewer results under pmat, mutation -> cargo-mutants,
crux -> crux, plus the (empty) antigravity run that consultation records.
The two results gain grounding=asserted / precision_class=advisory (as their
own text states) and a failure_scenario. findings_ref.sha256 rebound.
No verdict, finding or consultation changed. check_pr_review_receipt.sh
ACCEPTs it (verified under a throwaway key; the CI signer signs for real).

Agent: aprender-e6
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* evidence(PMAT-4594): commit the CI signer's signature for the a0e0e9c receipt

The B1 check-run route cannot go green on a rerun: the signature check-run
lands in present's own check suite, so a re-attempt never reads it
(inbox 21:09Z). Per COP 21:10Z 4598-ARM, arm via the committed route.
The .minisig is verbatim from pr-review-signature check-run 109135289610
(CI signer, repository key). minisign -V under .github/pr-review.pub is OK,
and check_pr_review_arm4.sh passes locally (A2 binds patch-id c860ecfc, A3
positive control fired, A4 ACCEPT).

Agent: aprender-e6
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…context_length (#4599) (#4601)

* fix(code): apr code -p never silently drops an over-budget prompt (#4599)

The sliding window skipped the newest message whole when it did not fit, and
apr code hard-coded a 32K window, so a >~105 KB prompt reached the model as an
empty conversation: rc 0, tokens_in 19 (#3715 pilot, 8 cells).

- serve/context.rs: the newest message is never dropped; when it cannot fit
  on its own the window is refused (ExceedsLimit -> context_overflow, nonzero).
- agent/code.rs: the window is the model's GGUF context_length (bounded
  header read), an explicit manifest value still wins, 32K only as fallback.
- Falsifiers: the 8 failing cells fit at 262144 and are refused (never
  dropped) at 32K; a >105 KB prompt is refused at 32K; the refusal is named
  and nonzero; the window comes from the model.

Refs #4599

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-3d

* test(#4599): over-budget boundary at 128 KiB; tiny-window test gets an input budget

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-3d

* test(#4599): MockDriver default window 32768, above the default 4096 output reserve

At 4096 every mock-driven agent test had zero input budget; the old sliding
window silently sent the model an empty conversation, which is #4599 itself.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-3d

* test(#4599): retry-test drivers get a window above the output reserve

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-3d

* test(#4599): 001 asserts the whole prompt, tail included, reached the model

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-3d

* fix(#4599): a tool result never evicts its turn's prompt; one tested driver-window resolver

Quorum (sonnet-5) on #4601: in a tool loop the sliding window kept the newest
tool result and could drop the current prompt, rc 0. truncate_messages now
refuses (ContextOverflow) when the last User message would be evicted; older
turns may still be. Both drivers launch through code_driver_window, pinned by
FALSIFY-4599-005. New: FALSIFY-4599-006/007.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-3d

* test(#4599): 007 asserts the current turn is whole and the old prompt evicted

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-3d

* test(#4599): echoed input token count == expected; planted 1-byte truncation is RED (FALSIFY-4599-008)

Operator R2: 212 KB and 510 KB prompts must echo the exact input token count.
The recording driver echoes a byte-level count (the 4 B/token estimate cannot
see one lost byte); cell_verdict checks count, needle and full content.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-3d

* test(#4599): agent_integration drivers had zero input budget under the 4096 reserve

test_routing_driver_fallback_integration's FailPrimary reported a 4096 window
and test_context_truncation_integration a 300 window, both at or under the
default 4096 output reserve. Before #4599 the runtime sent them an empty
conversation; it now refuses with context_overflow. Raise FailPrimary to 32768
and give the tiny-window test max_tokens 64, as the lib tests already do.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-3d

* test(#4599): kill the 4 missed and 2 timed-out mutants on the diff

- context.rs `>`→`>=` on the newest-message check: a newest message that exactly
  fills the window must be kept (boundary case added).
- model_context_length: one fn with cfg blocks, so the whole-fn mutant lands on the
  compiled body; a 0 context_length falls through to 32K (case added).
- build_default_manifest: drop the explicit `context_window: None` (the default);
  deleting it was an equivalent mutant.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-3d

* ci(#4601): re-run on a body that carries keep-open for #3715

guard-tree reads the PR body from the event payload, so a rerun of run
36454753912 re-reads the old body and cannot pass. The body now has a
"keep-open: #3715 <reason>" line; this empty commit makes a new
synchronize event that carries it. No code change.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-3d

* review(#4601): pr-review receipt at e3d9361 — DEGRADED (mutation not run on lambda by rule), no blocking class

Reviewer: independent Sonnet 5 session; agy gemini-3.1-pro-high advisory: pre-existing orphan tool_use eviction in truncate_sliding_window, not introduced here.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-3d

* review(#4601): stamp predicate.diff_patch_id into the receipt so Arm 4 A2 can bind it (#4421)

Computed with scripts/lib/pr_review_patch_id.sh (git-patch-id-verbatim/pinned-diff-v1) over base_sha..head_sha; the CI signer recomputes and refuses a mismatch.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-3d

* review(#4601): commit the CI signer's .minisig so Arm 4 verifies it unchanged (cop ruling 21:23Z, same route as #4598)

Taken verbatim from the pr-review-signature check run on b87d6b7; minisign -V against .github/pr-review.pub rc=0; check_pr_review_receipt.sh ACCEPT.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-3d

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…rds its backend (#4609) (#4612)

* fix(apr chat): the Qwen3.5 GPU turn records the backend that answered (#4609)

The qwen35 session branch returned without setting generated_on_gpu; the moe and dense branches
set it. So every CUDA Qwen3.5 chat turn printed backend {ran: cpu, fell_back: true} while stderr
said 'Backend: GPU (CUDA ...)' and apr run on the same model said ran: gpu. With #4609's grader
(fell_back => fail) every qwen35 chat cell of #3715 would be RED on a real GPU.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Refs: #4609
Agent: aprender-ec
(cherry picked from commit baec88e)

* fix(apr chat): a --gpu turn the CPU answered is exit 14, never a silent pass (#4609)

`apr chat --gpu` whose turn ran on the CPU printed the answer, reported
`fell_back: true` only in the --json epilogue, and exited 0, so a #3715
cell that asked for the GPU graded PASS. `apr run --gpu` already refuses
that case (R-0b, exit 14). Chat now calls the same
`registry::after_generation` per turn: it prints no answer and ends the
session with BackendUnavailable (14), after printing the --json epilogue.

`generated_on_gpu` is reset at the start of each turn, so a CPU turn
cannot inherit an earlier GPU turn's flag. A default request (no --gpu)
is unchanged: the epilogue still says fell_back.

Tests: pmat4609_forced_gpu_turn_check covers both polarities (--gpu+CPU
→ 14; --gpu+GPU, no flag, --no-gpu → Ok).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Refs: #4609
Agent: aprender-3d

* review(#4612): pr-review receipt at c61482f — DEGRADED (mutation unreachable: include!() files), no blocking class

Reviewer: independent Sonnet 5 session; agy gemini-3.1-pro-high advisory, no findings.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-3d

* review(#4612): stamp predicate.diff_patch_id into the receipt so Arm 4 A2 can bind it (#4421)

Computed with scripts/lib/pr_review_patch_id.sh (git-patch-id-verbatim/pinned-diff-v1) over base_sha..head_sha; the CI signer recomputes and refuses a mismatch.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-3d

* review(#4612): SARIF driver crux-consultation -> crux (guard's allowed set); findings_ref.sha256 recomputed by the reviewer

Found only once the receipt was signed: the unsigned rejection masked it. diff_patch_id unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-3d

* review(#4612): restore the receipt's trailing newline (the signer requires exactly one line)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-3d

* review(#4612): commit the CI signature as .minisig so present verifies with Arm 4 unchanged

Taken verbatim from the pr-review-signature check run on 51bb553 (x86-main).
minisign -V rc=0; check_pr_review_receipt ACCEPT. Route per cop 21:23Z ruling.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-3d

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…uld never go green (#4618)

* fix(arm4): B1 reads every attempt's signature — a rerun of present could never go green

The pr-review-signature check run lands in the pr-review-quorum run's own
check suite. Arm 4 read check-runs with the API default filter (latest)
and took the newest run, so a re-attempt of `present` read nothing and
printed UNSIGNED beside a verifying signature on the same head. #4598: sig
109130365825 success at 20:52:12Z; run 36481105318 attempt 2 UNSIGNED at
~21:04Z; the signer then posted another sig after present failed.
2/94 recent present runs passed, both via a committed .minisig.

- the listing uses filter=all
- the selector takes the latest successful run on the head that CARRIES this
  receipt's key, not the latest run
Nothing loosens: the pick must still verify under the repository key and
name exactly pr, head and pid.

Case table 33 -> 37 rows:
  check-sig-prior-attempt    GREEN (sig in an earlier attempt, newer unsigned + queued)
  check-sig-prior-wrong-pid  RED
  check-sig-prior-wrong-head RED
  check-sig-prior-other-sha  RED
Red proof: reverting the selector alone turns check-sig-prior-attempt RED
(rc=1, wanted 0); the three controls stay RED.
NOT covered by the table: the filter=all URL itself, because injection
bypasses the API.

Agent: aprender-e6
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(arm4): prove filter=all through the API, not only through a file

Review of 4c12e25 (BLOCK): dropping &filter=all from the URL left the
self-test 37/37 GREEN — every prior-attempt row injected a file, so the
URL was never exercised. New row check-sig-api-all-attempts drives the
real fetch through a gh shim that returns the earlier attempt only when
the query carries filter=all. Mutation: drop filter=all → that row RED.

Also: a newer run whose output.text is valid JSON but not an object
(e.g. "[]") crashed jq's .[$k] and read as UNSIGNED. Guarded; row
check-sig-prior-array-text; mutation: drop the guard → that row RED.

Self-test 39/39.

Agent: aprender-e6
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* evidence(pr-review): non-author quorum receipt for #4618 at 78eb06a

Reviewer claude-sonnet-5 (author claude-opus-5-5): PASS, 0 findings.
Measured by the reviewer: self-test 39/39; drop &filter=all → only
check-sig-api-all-attempts RED. pid 3aa28558bc (evidence excluded).
check_pr_review_receipt.sh ACCEPT under a throwaway key.

Agent: aprender-e6
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* evidence(pr-review): commit the CI signer's .minisig for #4618's receipt

Verbatim from the pr-review-signature check run on ee080ab (trusted
comment pr=4618 pid=3aa28558bc). minisign -V under .github/pr-review.pub
passes; local Arm 4: A3 fired, A4 ACCEPT, PASS. Same route as #4598
(present ran before the CI signer posted).

Agent: aprender-e6
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(counts): the Arm 4 step name states 39 rows, not 33

x86-main on 500fb1f: check_pr_review_counts.sh derived arm4_rows=39
and found "Arm 4 case table: 39 rows" 0 times in pr-review-quorum.yml.
Label-only change; no gate logic touched. Local: counts run + --self-test
green.

Agent: aprender-e6
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* evidence(pr-review): fresh non-author receipt for #4618 at 437037f

The count-label fix changed the patch-id (3aa28558bc → 253965ff05), so
the 78eb06a receipt no longer binds; this is a fresh review, not a
patched one. claude-sonnet-5, PASS; counts guard run + self-test green
as measured by the reviewer. ACCEPT under a throwaway key.

Agent: aprender-e6
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* evidence(pr-review): commit the CI signer's .minisig for #4618's 437037f receipt

Verbatim from the pr-review-signature check run on 17efd38 (pid
253965ff05). minisign -V under .github/pr-review.pub passes; local Arm 4 PASS.

Agent: aprender-e6
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Fragment copied byte-identical from cop-inbox/handoff/PMAT-4621.yaml;
roadmap.yaml regenerated (aggregate --check ok, +22 -0).

Agent: aprender-59
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* fix(deps): wasmtime 47.0.4 -> 48.0.3 for RUSTSEC-2026-0315/0316

cargo audit on origin/main's Cargo.lock reports RUSTSEC-2026-0315 and
RUSTSEC-2026-0316 against wasmtime 47.0.4 (fixed >=48.0.3, <49 or >=49.0.1),
which turns the security job red on main and on every PR. A version bump
only; no audit ignore is added, so no gate is relaxed.

wasmtime stays optional behind aprender-test-lib's runtime feature, which
nothing enables: cargo tree -i wasmtime --workspace finds no active path, so
the 1.95 MSRV of the 48 line never fires on the 1.93 pin (comment updated).

Receipts (intel/devont, fresh advisory DB): cargo audit rc=0 (was: 2
vulnerabilities), cargo deny check rc=0 (advisories, bans, licenses, sources
ok), cargo fmt --check rc=0, cargo metadata --locked rc=0.

Agent: aprender-00
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(roadmap): PMAT-4633 — wasmtime RUSTSEC-2026-0315/0316 bump

Agent: aprender-00
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(audit): PMAT-4633 — review input for the wasmtime 48.0.3 bump

The review lanes get the Cargo.toml diff, a lock-delta table generated from
Cargo.lock (crate, old -> new, checksum), the RUSTSEC-2026-0315/0316 text, and
cargo audit / cargo deny / cargo metadata --locked on base and head. The raw
lockfile stays with the machine checks. cargo audit goes from rc=1 (both
advisories on 47.0.4) to rc=0. cargo deny passes on both sides because
wasmtime is reachable only through aprender-test-lib's optional `runtime`
feature, so it does not tell them apart.

Refs #4633

Agent: aprender-00
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* review(PMAT-4633): full cargo deny evidence; acceptance names the C45 review doc

Q0 round on e9728ba: both Gemini lanes failed on (1) no evidence for deny bans/licenses/sources, and (2) the acceptance criterion's file list omitting the audit doc that C45 requires in place of the raw Cargo.lock. Full cargo deny check (cargo-deny 0.19.0) was run on base 2817c6d and on this head: rc=0 on both, all four checks ok.

Pmat-Ticket: PMAT-4633

Agent: aprender-07

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* review(PMAT-4633): pr-review receipt for #4632 @ 7af51f3 (receipt-only)

Independent reviewer agent:claude-sonnet-5-5/pr-4632-review; verdict FINDINGS (7 advisory, 0 blocking); agy arm gemini-3.8-flash-high after gpt-oss-120b-medium 503x2 (operator fallback ruling). predicate.diff_patch_id 74b701ae170949b97b6ea5f75126a5588c716438; the signature is posted by CI as the pr-review-signature check run. Operator C72 exception to C69.

Pmat-Ticket: PMAT-4633
Agent: aprender-07
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…ine (PMAT-4646) (#4647)

* ci(mutants): ratchet mode — report-only in ci / gate, shrink-only debt baseline (PMAT-4646)

Operator rulings C187(a)/C188(a), 2026-10-01: a new gate is adopted as a
ratchet, never as an instant absolute bar. No ruleset change.

- ci / gate: mutation results (the mutants section and, on PRs that add
  it, the survivor table) are printed and no longer fail the gate.
  Every check before it is unchanged: workspace-test, guard-tree,
  guard-cargo, sov.gate, determinism-compare.
- ci/mutants-debt.tsv: 590 known survivors (MissedMutant + Timeout),
  with file, mutant, sha and run id, from completed shard artifacts of
  #4587 run 36725778557 and #4588 run 36703131059.
- scripts/check_mutants_debt_ratchet.sh (auto-run by guard-tree): the
  debt file is shrink-only against the merge base. A grown file, a
  swapped row, a missing file, a malformed or duplicate row, or an
  unresolvable base is RED. Self-test: 10 cases; 7/7 planted mutants RED.

The RESTORE PR deletes the MUTANTS-RATCHET block after crates.io 0.70.

Agent: aprender-ca
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(roadmap): PMAT-4646 entry (MUT-RATCHET acceptance criteria for the quorum)

Refs #4646
Agent: aprender-ca
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ratchet): base via scripts/lib/resolve_base.sh (CI depth-1 checkout); gate echo names only defined vars

check_mutants_debt_ratchet.sh [run] was RED in run 36829781758: merge-base(HEAD, origin/main)
is unresolvable on a depth-1 checkout. It now uses resolve_base, the base every differential
guard uses, which refuses rather than judge HEAD against HEAD (new case R11, plus R12 for the
default origin/main path; 12/12; a fail-open mutant of the new branch is RED).
check_workflow_env_defined: the gate echo interpolated IN_SCOPE/SCOPE/TABLE, which main's gate
job does not define; dropped.

Refs #4646
Agent: aprender-ca
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(roadmap): PMAT-4646 entry states the 12-case self-test (R1-R12), per q-4647c finding

Pmat-Ticket: PMAT-4646
Agent: aprender-ca
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* ci(gate): the ratchet comment says where the baseline lands (#4647) and that the mutants rules are off on purpose until #4648

Quorum lane finding (q-4587c, sonnet): the comment named ci/mutants-debt.tsv and
scripts/check_mutants_debt_ratchet.sh as the blocking rule, but on a branch
without #4647 neither exists. Same block text in #4647/#4587/#4588 so the three
still merge without conflict. No check changed.

Pmat-Ticket: PMAT-4646
Agent: aprender-ca
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* evidence(pr-review): non-author receipt for #4647 at 368f05f

Sonnet seat (agent:claude-sonnet-5-5/pr-4647-review), base 00052c0. Verdict
FINDINGS (2 warnings, 5 notes; none blocking): PRREV-RATCHET-001 (new-survivor
detection from CI results is follow-up, report-only by C188(a)); bashrs 0 errors.

Pmat-Ticket: PMAT-4646
Agent: aprender-ca
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ratchet): bashrs-clean guard; roadmap ACs match the diff; new-survivor detection is #4649

q-4647d (3 lanes) findings, all fixed:
- scripts/check_mutants_debt_ratchet.sh: bashrs lint 0 warnings (quoted
  assignments, EXIT trap beside RETURN, false positives tagged); 12/12 self-test.
- AC3 named a survivor-table section the gate does not have: now 'the mutants
  section result'.
- AC4 now lists the PR's own review artifacts (evidence/, quorum json).
- New-survivor RED (C187) is out of scope here, filed as #4649.

Pmat-Ticket: PMAT-4646
Agent: aprender-ca
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ratchet): drop the stale 368f05f receipt; AC4 separates the reviewed diff from review outputs

q-4647f (gemini, haiku): the in-tree receipt for 368f05f still carried the
pre-fix bashrs finding, so lanes judged the current script by it; and AC4 listed
roadmap.yaml and the quorum json, which the quorum brief does not contain. The
receipt for this head is committed next; the stale one binds nothing.

Pmat-Ticket: PMAT-4646
Agent: aprender-ca
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* evidence(pr-review): non-author receipt for #4647 at 1f2e359

Sonnet seat (agent:claude-sonnet-5-5/pr-4647-review), base 00052c0, verdict
FINDINGS, none blocking. Quorum q-4647g AGREED 3/3 (committed after the signer).

Pmat-Ticket: PMAT-4646
Agent: aprender-ca
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* spec(pp-llama-001): re-date §12 rows 13, 15, 18, 19, 21 to 2026-10-15 — operator C197

None of the five was discharged by its 2026-09-30 expiry, and the D6 andon
(scripts/lib/spec_conformance.py) turned every PR and main RED on 2026-10-01.
§12 allows exactly one way to move an expiry: an amendment recording who moved
it and why. This is that amendment:

- Moved by the operator (Noah), ruling C197. Reason: "not delivered for 0.70;
  carried to 0.71".
- Rows 13, 15, 19 and 21 carry the typed date. Row 18 derives it from row 15.
- An Appendix D row records the move.
- derived_expiries.json was regenerated with --write; exactly the five rows
  changed.

The andon is unchanged. It passes today, and with
SPEC_CONFORMANCE_TODAY=2026-10-16 it is RED again (D6). The release notes carry
§12's consequence: "NO SPEED DELIVERED: PP-LLAMA-001 rows 13, 15, 18, 19, 21
(re-dated to 2026-10-15)".

Refs #4651. Advance warning before expiry: #4652 (post-0.70).

Agent: aprender-ca
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ratchet): a debt file shrunk to zero rows read as malformed — the ratchet could never reach clean

Quorum on 65d3fb6 (lanes claude-sonnet-5 + gemini-3.1-pro, independently):
check_rows fed awk `<<<""` when the debt file had only its header, so the
single empty line hit `NF != 5` and the guard went RED on the final shrink.
PMAT-4646 requires GREEN on a shrink.

- awk skips NF == 0 (rows() already strips blank lines, so no real row is empty)
- self-test R13: header-only file vs a 2-row base is GREEN "2 -> 0".
  RED before the fix (rc=1), GREEN after; deleting the NF==0 line reds R13 only.
- PMAT-4646 AC: 13/13 (R1-R13).
- Non-author receipt for 65d3fb6 (verdict FINDINGS) committed as evidence.

Refs #4646

Agent: aprender-ca
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ratchet): self-test kills the 3 surviving guard mutants; ci.yml comment states what actually blocks

Review of a683b4a (haiku lane FAIL, non-author seat FINDINGS):
- 3 sed mutants survived the self-test (PRREV-MUT-002): empty run id
  accepted, file-name check unanchored, NF<5 instead of NF!=5. New cases
  R14 (empty run id), R15 (file named mid-string), R16 (6 fields): each
  mutant now reds exactly its own case; 16/16.
- ci.yml comment said "the blocking rule is the shrink-only survivor
  baseline". It blocks only GROWING the file; new survivors are #4649 and
  line-shifted rows are #4653 (new issue). Comment and script header say so.
  Comment-only; no gate changed.
- PMAT-4651 AC: the "NO SPEED DELIVERED" line belongs to the v0.70.0 release
  notes at tag time, not to this PR.
- Not changed: the lane's "R13 is RED" claim. Measured at this head:
  R13 rc=0, self-test 13/13 then 16/16.
- Non-author receipt for a683b4a committed as evidence.

Refs #4646 #4651 #4653

Agent: aprender-ca
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(roadmap): quote the PMAT-4646 AC; its "R1-R16: ..." parsed as a YAML map

0711f26 broke roadmap-valid (sov.security + sov.gate, CI 36873128678):
"roadmap[1098].acceptance_criteria[1]: invalid type: map". The ": " inside
the unquoted AC made it a mapping. Quoted it; pmat work validate rc 0,
pmat work status PMAT-4646 rc 0, roadmap sorted rc 0. The cop committed
without checking validate's rc; it is checked explicitly now.

Refs #4646

Agent: aprender-ca
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* evidence(4647): round-4 review receipt + quorum AGREED (PASS/PASS/PASS) at bd85154

Agent: aprender-ca
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* evidence(4647): stamp diff_patch_id 41b11c0a into the round-4 receipt (Arm 4 A2 binds on-disk)

Agent: aprender-ca
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(queue): #4647 was dequeued twice by two queue-only defects — R11 read the CI event, Arm 4 read the squash

1. check_mutants_debt_ratchet.sh self-test R11/R12 ran the script with the
   job's GITHUB_EVENT_NAME. Under merge_group, resolve_base accepts a single
   parent as the base, so "no nameable base" returned rc=0 and R11 failed
   in the queue only (runs 36888741624, 36890959980). The temp repo is not a
   CI checkout: the subprocess now runs with -u GITHUB_EVENT_NAME.
   Proof: GITHUB_EVENT_NAME=merge_group --self-test rc 1 before, rc 0 after;
   pull_request and unset rc 0 both.

2. pr-review-quorum.yml passed the merge_group head (the queue squash) as
   ARM4_PR_HEAD_SHA. The signer posts pr-review-signature on the PR head, so
   every queued PR read UNSIGNED (#4632 run 36602609576; #4647 runs
   36888741583, 36890960012). ARM4_PR_HEAD_SHA is now the pull_request
   event's head (empty on merge_group), and Arm 4 reads the PR head from
   the pulls API; the job gains pull-requests: read.
   Proof on queue sha 79594c8 against the live API: old wiring rc 1
   UNSIGNED; new wiring rc 0, signature on 163c660 verifies, binds
   pr=4647 pid=41b11c0a, A3 positive control still rejects a corrupt sig.
   No gate loosened: same key, same trusted-comment binding, same A2 diff bind.

Pmat-Ticket: PMAT-4646
Agent: aprender-27
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(roadmap): PMAT-4646 AC names the pr-review-quorum.yml queue-only fix (round-5 quorum failed it as out-of-ticket)

Operator 2026-09-28 14:52Z pre-approves gate-keeping workflow edits with red/green proof;
cop decision c213 (option a). Cites runs 36888741583/36890960012 and queue sha 79594c8.

Pmat-Ticket: PMAT-4646
Agent: aprender-r5
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* evidence(4647): round-6 review receipt + quorum AGREED (PASS/PASS/PASS) at ce8036e

diff_patch_id c75016cb stamped. Round 5 failed as out-of-ticket; ce8036e amended the
PMAT-4646 AC to name the pr-review-quorum.yml queue-only fix.

Pmat-Ticket: PMAT-4646
Agent: aprender-r5
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…AT-937 (C221 fat_driver fix) (#4655)

* fix(release): models_t1 readiness rows ran without the release apr (#3715)

Quorum lanes 1 and 3 (sonnet, haiku) measured 4 of 39 rows RED: 997a34381f
made the T-1 readiness step re-prove $REPO_ROOT/target/release/apr against
MC and take its surface, and the harness never put an apr there. ready() now
plants a stub apr at the path autopilot pins; r_goes also requires the
wrapper to receive --surface $AP/surface-t1.json from that binary. New row
readiness-stale-apr-stops (an apr from another commit stops the step before
the wrapper runs) and mutant readiness-any-apr (the check removed -> RED).
41/41 rows; floor 41. Lane 2's minor: the R8 fixture comment still called
report the committed mode.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-0d
(cherry picked from commit 57a94972f37968111062b0d7c8dbb06ef69aa4ad)

* fix(#3715 R10): delete the report-only arm -- enforce is the only mode, and a WARN row stops T-1 and T-4

Operator R10 (2026-09-28): "report-only is not an option". L19: a silent report-only check = hard stop.

- release_readiness.sh: resolve_mode prints enforce or fails; `report` as the env value OR as a committed
  DEFAULT_MODE is a caller error (3). Both WARN R8 REPORT-ONLY arms (the "not a stop until #3712" line)
  are gone. Selftest row mode_has_no_report_arm.
- autopilot.sh T-1: a WARN R8 row from the wrapper is a die, even at rc 0.
- check_publish_preflight.sh T-4 R8: rc 0 is not a pass without "#3715 ENFORCE PASS" for this version
  and HEAD; a WARN R8 row refuses.
- Rows flipped: readiness-warn-continues -> readiness-warn-stops; r8_report_mode_warn_passes ->
  r8_report_mode_warn_refuses; new r8_rc0_without_enforce_pass_refuses. Each killed by its mutant.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-0d

* G2 (#4590): withdraw every GB10 dense claim through data — qwen2@gx10 de-claimed with qwen3

Operator ruling 2026-09-28 16:15Z widens R1 (Qwen3 only) to every dense model on GB10:
serial prefill is the sm_12x default for every non-hybrid arch (select_prefill_path).

- cells.declaimed: + qwen2 @ gx10 (issue 4590, until 0.70.1); rung qwen2-1.5b-q4km hosts: [lambda]
- README + CHANGELOG: "GB10 dense models: prompt processing unbatched (11.4 tok/s); fixed in 0.70.1 (#4590)"
- enforce-gate proof (ont_release_readiness.rs): withdrawn on gx10 owes nothing; a planted gx10
  claim of the dense model is RED; lambda dropping one dense row is RED (withdrawn != waived)

Agent: aprender-e6
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ci #4587): lib-test container trusts /workspace so the pinned origin/main resolves

The workspace-lib docker run is root over a host-owned checkout; git refused it
(dubious ownership), rev-parse origin/main failed, the refinement gate skipped and
lint_passes_on_real_contracts asserted under CI (run 36449404284 job 109019972318).
Reproduced on intel with sovereign-ci:stable: rc=128 without, rc=0 with the env.

Agent: aprender-88
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* release(pv #4604): C2 bump — pv set at 0.69.4, update-check + build-sha off by default

macros, aprender-common, aprender-contracts and aprender-contracts-cli pin
0.69.4 (v0.69.3 is taken; a version string is never reused). The cli's
aprender-update and aprender-build-sha deps become optional features so the
crates.io pv builds only from crates live there (B2a). Without build-sha,
build.rs stamps APR_GIT_SHA from APR_GIT_SHA_OVERRIDE or v<ver>+no-git.
pv_bin.sh and the Makefile pass --features update-check,build-sha, so
in-tree pv keeps its update check and sha. binary-release.yml is a
separate branch, pending operator OK.

Agent: aprender-88
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 34ff1f7432a20f643161250d869cd8e0c9729828)

* release(pv #4604): 6-crate set — aprender-build-sha + aprender-update at 0.69.4

cargo publish needs optional deps live on crates.io too (0d C1 dry-run).
Both are new crate names there; publishing them waits on the operator
ruling. Fallback without them: release/pv-0.69.4-fallback.

Agent: aprender-88
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 1e5f8a84849003421e8e8640346a048bc87d595c)

* ci(binary-release #4604): release pv with --features update-check,build-sha

The C2 bump makes both features off by default so crates.io pv builds from
live crates. Release binaries keep the update check and build sha. Held:
no merge until operator OK (workflow edit).

Agent: aprender-88
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 372c980391143706bd811473bd43b3bc3385712b)

* fix(shapes #4587): the #3610 exemption moves off the shape so the fleet-pinned pv parses it

pv 0.69.1 (the fleet pin, intel-clean-room) refuses the shape key allowEmpty
(rc=3: shape release-readiness-v1.refusal uses unsupported allowEmpty), so
check_fleet_pv_shapes_gate.sh reported pv-no-verdict (run 36449404284 job
109019972501). Deleting the key would make every green release decline under
--shape release-readiness-v1 (the refusal shape is empty by design). The reason
now lives in a contract-level allow_empty map keyed by shape id, which 0.69.1
ignores and HEAD applies with the same semantics; the in-shape key still parses.

Agent: aprender-88
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(#4587): x86-main reds from run 36462668621 — contracts recipe and facade guard vs the scoped 0.69.4 pv set

1. `make contracts` called scripts/contracts_gate.sh, which PR-1 does not carry (it
   lands with PR-2). Restore main's step-list recipe; PR-2 brings the gate script and
   its recipe together. make_contracts_propagates.sh: 17/17 rows, mutants included.
2. make_contracts_propagates.sh faked pv at the ROOT workspace version, but pv_bin.sh
   declares the CLI crate's own version first. The fixture now reads it the same way.
3. check_facade_compat.sh CURRENCY and PUBLISH ORDER compared against the root
   workspace version; R3 already used the fronted crate's own. All three now use the
   version this tree publishes for the fronted crate (identical when unscoped; an
   unreadable version is a FAIL). Facade pins -> 0.69.4, crates/facades/Cargo.lock regen.

Agent: aprender-88
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(ladder): a de-claimed (host, arch) the ladder still claims is RED — planted RED, qwen3moe@gx10 (#3715)

STAGED, branch-only (cop 20:00Z/20:26Z; operator rules 07:30Z). Not for merge.

cells.declaimed withdrew a (host, arch) from the owed cells, but nothing
failed when a rung's host set (what released pv reads) or a long-rung
representative still claimed it. declaim_still_claimed() in the judge
(enforce: judge rc 1 -> check exits 1):
- a rung whose hosts (or, absent, every required host) name a de-claimed
  (host, its arch) -> FAIL
- a long_rungs_for representative for an arch de-claimed on every required
  host -> FAIL

Case table: red-cells-declaimed-still-claimed (new), green-cells-declaimed
narrowed to hosts: [lambda] (it claimed the de-claimed pair); cmutant
declaim-claimed killed. Self-test 153/153.

Data: cells.declaimed + (gx10, qwen3moe), issue 3715, until 0.70.1 (0d H4
19:54Z: arch-level DATA is clean for qwen3moe only). Real-contract proof in
evidence/3715-declaim-planted-red/: real = 0, planted qwen3moe rung = 1,
narrowed = 0, orphan representative = 1.

Agent: aprender-e6
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(serve): a turn past the device KV cache is a 400 "context exceeds N", not a 500 (D5)

Operator ruling D5 (ruling-3715-2010): "a 500 is a crash-class lie".
A 27155-token turn on a dense CUDA model reached BorrowedCudaForward::reserve,
failed with "the device KV cache holds 4096", and came back as HTTP 500
(h4-depcheck-0d, lambda + gx10 serve/code cells).

- serving_context(): the model's context capped by the device KV cache;
  the borrowed serve session reports it, so max_tokens clamps to what is left.
- AppState caches it at construction (PMAT-073: no read lock in the handler).
- serve_context_refusal(): the pre-flight 400 in the CUDA chat, CUDA
  /v1/completions and /generate/stream routes, before any path takes the
  model, so a streaming client gets a 400, not a mid-stream error.
- The direct paths map ContextLimitExceeded through generation_error_status.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(#3715): release-note warning — qwen3moe on GB10 is not claimed for 0.70

Operator D1 (relayed by the cop, 2026-09-28 20:10Z): the qwen3moe-gx10 de-claim
lands with the planted RED and a release-note warning. CHANGELOG Known
issues + README known-issue note, same shape as the #4590 entries.

Agent: aprender-e6
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(D1): the qwen3moe de-claim said lambda does not hold qwen3moe — it holds both files

lambda holds Qwen3-30B-A3B-Instruct-2507 and Qwen3-Coder-30B-A3B (~/models, lambda H4 queue), so the
entry now says what withdrawn-not-waived means there: lambda still owes every qwen3moe cell.
Data text only; the judge keys on (host, arch) and is unchanged. Falsifier re-run on the real contract:
entry -> withdrawn rc 0, drop -> owed again, bare -> rc 1.

Refs #3715

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-0d

* fix(facade guard #4587): exact scoped allowlist — only the declared #4604 set at 0.69.4 may leave the workspace version

Quorum q-pr1-r81cd2 (gemini FAIL): comparing a crate against its own version let a
coordinated off-workspace drift pass. pub_ver now admits the workspace version, or the
EXACT version scoped_ver declares for that crate; anything else (unlisted crate, other
version, unreadable) is RED. The only input the old workspace-only rows refused and these
admit is the declared set itself.

Agent: aprender-88
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(D1): CHANGELOG said no other host holds qwen3moe — lambda holds both files and still owes them

Same false claim 0d fixed in the contract (e59329fe67); the release note carried it too.

Agent: aprender-e6
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(make #4587): no PR-1 recipe calls a script only PR-2 carries

Four Makefile recipes in the split named scripts that live in PR-2 alone:
census (check_census_derived.sh), guards-local (ci_guards.sh),
oracle-owl (tests/oracle/owl) and coverage-serve (the .py -> .sh rename
would have broken `make coverage`). Those hunks go back to main and travel
with their scripts in PR-2. lint-ratchet stays: PR-1 carries EV-11, whose
contract names scripts/lint_ratchet.sh as its writer, so the script comes
with it (--self-test 12/12).

Agent: aprender-88
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* roadmap: PMAT-4616 (D5): the ticket #4614 implements, not #3715

#3715 is the pv SHACL release-readiness shape. D5 is its serve × consumer-context
child: a turn longer than the device KV returned 500. The quorum judged this diff
against #3715's acceptance and failed it 3/3 on the mismatch, so the serve fix
now has its own ticket and acceptance.

Refs #4616, #3715

Agent: aprender-2e-serve
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(dogfood #4587): re-derive per_binary/band/cluster baselines for PR-1's 22 ledger rows

run 36473794463 guard-tree step 11: the split moved the 22 surface_audit rows (11 apr, 11 pv)
and the rows/ratio totals into PR-1, but not the per_binary / per_band / per_cluster entries.
Re-derived with scripts/dogfood_baseline.py; --check PASSES.

Agent: aprender-88
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(guards #4587): mutation sets for the three guards PR-1 adds, and a CI caller for the two lean ones

Review B3 (PR-REVIEW-SKILL-002 v2 S3.D): check_facade_compat.sh (scoped
allowlist), check_leanchecker_scoped.sh and check_lean_modules_reachable.sh
had only self-authored self-tests, and the two lean guards had no CI caller.

- mutate_{facade_compat,leanchecker_scoped,lean_modules_reachable}_guard.sh:
  17/17, 31/31, 18/18 killed. Each patch fails closed (INVALID, never a
  survivor), the M0 baseline must pass, and mutants run against copies.
- First runs killed 0/17, 27/31 and 16/18. The self-test rows that kill the
  survivors were added. No check was removed or loosened. scoped_ver/pub_ver
  were moved, unchanged, above the self-test so it can reach them.
- ci/sections.yml guard-cargo runs the mutation sets, plus both lean guards
  (self-test and the real check).

Full facade guard on intel: PASS.

Agent: aprender-88
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(version #4587): a facade pin must name the crate it re-exports, not the workspace beside it

guard-tree step 17 (bump-version.sh --check) failed on run 36485785730. The
check compared each facade pin with the ROOT workspace version (0.69.3). But
PR-1's scoped pv set moves aprender-contracts{,-macros} to 0.69.4, so the pins
on those two crates (0.69.4) were the correct values, and the check called them
wrong.

--check now compares each pin with the version its `upstream` path crate
actually carries. A crate on version.workspace = true carries the workspace
version. Case table rows 7-10: one pass and one RED for a literal-versioned
upstream, and one pass and one RED for a workspace-versioned upstream. Each RED
row must stay RED. The self-test passes 10/10, and --check passes on this tree.

Agent: aprender-88
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* evidence(PMAT-4616): pr-review receipt for 167bd5e126 (reviewer sonnet-5, DEGRADED, advisory finding)

Non-author reviewer (Sonnet 5), per COP 22:16Z applying the 21:10Z committed-
.minisig route. diff_patch_id cf52ad666e. Verdict DEGRADED: cuda (docs MCP
OAuth-gated) and mutation (no cargo test on lambda) unreachable; pmat and agy
gemini-3.1-pro-high consulted; crux not triggered. Guard rejects only B1
(unsigned) until the CI signer's .minisig is committed.

Advisory finding (measured): /generate and /batch/generate still map every
CUDA error to 500 (batch_processing.rs:492, batch.rs:515). Follow-up, not
folded here: changing the diff would void this receipt.

Agent: aprender-2e-serve
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* evidence(PMAT-4616): commit the CI signer's signature for the 167bd5e126 receipt

Per COP 22:16Z (21:10Z committed-.minisig route). The .minisig is verbatim
from pr-review-signature check-run 109170961010 (CI signer, repository key).
minisign -V under .github/pr-review.pub: verified, trusted comment pr=4614
pid=cf52ad666e. check_pr_review_receipt.sh: ACCEPT (positive controls fired).

Agent: aprender-2e-serve
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* evidence(pr-review #4587): Sonnet review receipt for head ba09f72af5, bound to patch-id 1817f9d2

The reviewer is claude-sonnet-5; the author is claude-opus-5-5. The verdict is
BLOCK, and its only reason is the PR1-MUTANTS HARD-STOP (3555 mutants in the
diff, cap 60, operator ruling at 07:30Z). B3 (guard mutation sets) is resolved
by c1e4d195c5. The receipt is unsigned; the ci pr-review-sign job signs it.

Agent: aprender-88
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ci #4587): the lean-guards step runs each guard under setsid --wait

check_guard_steps_isolated.sh was RED on the local guard-tree run: the six
invocations the step added in c1e4d195c5 shared the runner's process group
(#4133). Now green: every direct guard invocation in 25 workflows is isolated.

Agent: aprender-88
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(cuda): chunk the dense batched prefill; QK-norm models stop taking the serial prefill (#3715)

The #3715 lambda cells timed out or failed at the 20k rung on two causes,
both measured on the RTX 4090 with apr 0.70.0 (8d021f61e):

- Qwen3-1.7B (QK-norm): `[PREFILL-PATH] serial cc=89 (qk-norm model)`.
  One forward_gpu_resident per prompt token: 7955 tokens took 109 s,
  26131 did not finish in 300 s at GPU util 99-100 %. Not #4609: the
  backend was the GPU throughout.
- Qwen2.5-0.5B IQ3_M (batched): CUDA_ERROR_OUT_OF_MEMORY in
  prefill_all_layers_gpu at 26148 tokens. One pass allocates a
  num_heads x S x S f32 score matrix (38 GB), and its u32 row offsets
  wrap past 2^32.

The batched prefill now runs in chunks whose score matrix fits the
1 GiB budget the Qwen3.5 prefill already uses; each chunk appends to the
KV cache through the existing cache_len > 0 attention path. A prompt that
fits one chunk runs exactly as before. QK-norm models keep FP8 off and
now take the FP16 batched prefill, which #3413 measured passing CPU
parity; serial was chosen only because the unchunked prefill OOMed 8B.

Refs: #3715
Agent: aprender-ec
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(chat): size the chat device KV to the context when VRAM allows (#3715)

apr chat built its CUDA model with a flat 2048-position KV. The dense session
refuses a longer turn on the device, so a 26k-token chat turn on Qwen3-1.7B ran
on the CPU and hit the 600 s cell timeout (GPU util 0%, measured on lambda).
apr run already sizes its KV to the turn (#4268); chat loads before it has read
a turn, so it now takes the model context, capped by free VRAM after the weights,
the FP16/FP8 prefill cache, the GH-178 reserve and the chunked-prefill score
budget. When nothing is left (Qwen3-8B on 24 GB), it keeps the 2048 floor.

Refs: #3715
Agent: aprender-ec
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(cuda): Qwen3.5 prefill dequantizes F16 and IQ4_XS on the device (#3715)

The Qwen3.5 batched prefill dequantizes each weight to F32 for one SGEMM, and
had kernels for Q4_K/Q5_K/Q6_K/Q8_0 only. Qwen3.5-4B-UD-Q4_K_XL (F16
ssm_alpha/beta, IQ4_XS FFN blocks) and Qwen3.5-0.8B-IQ4_XS therefore refused
the CUDA prefill, the F2 guard rejected the GPU, and every lambda cell ran on
the CPU (GPU 0%) into the 600 s timeout.

Adds F16DequantKernel and Iq4XsDequantKernel. The values match the per-token
GEMV weights bit for bit (device tests against the ggml reader).

Refs #3715

Agent: aprender-ec
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(cuda): a QK-norm model keeps the serial prefill when its FP16 batched prefill does not fit (#3715)

3c2a354b7d moved QK-norm models from the serial prefill to the chunked FP16
batched one. The FP16 weight cache is warmed on either path, but only batched
also needs the prefill workspace. Qwen3-8B on the 24 GB 4090 has 4.7 GB of
weights and a 15.1 GB FP16 cache, which leaves 0.1 GB, so init_prefill_workspace
failed and a 2k chat turn fell back to the CPU (382 s, rg5 C-q38b-chat-2k).

The path is now decided against the device before the cache is warmed:
weights + cache + KV + prefill reserve over free VRAM keeps serial, as on the
base. Qwen3-1.7B (about 9 GB of 23) stays batched. BATCHED_PREFILL still
overrides. The resident-bytes estimate is shared with session_kv_len.

Refs: #3715
Agent: aprender-ec
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* evidence(models #4587): re-derive supported.yaml after the README edit

derive_model_manifest.sh --check was RED on the local guard-tree run (step 28):
PR-1 moved README.md lines, so every README citation line in the derived
manifest drifted. Re-derived with the script; line numbers only, 24 models.

Agent: aprender-88
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* evidence(pr-review): commit the CI signer's .minisig for #4587's ba09f72af5 receipt

Verbatim from the pr-review-signature check run 109176017607 on 7eb331468c
(trusted comment pr=4587 pid=1817f9d24d). minisign -V under
.github/pr-review.pub passes. Same route as #4618 500fb1fb10 and #4612.

Agent: aprender-88
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* evidence(pr-review): #4620 non-author receipt for fdc165f50d (DEGRADED: mutation unreachable)

Reviewer agent:claude-sonnet-5/pr-review-4620; author claude-opus-5-5/aprender-ec.
Written without diff_patch_id or .minisig. The CI signer stamps both.

Refs: #3715
Agent: aprender-ec
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* evidence(pr-review #4587): Sonnet review receipt for head 7218961ef9, bound to patch-id 9d93fbed

Reviews the delta ba09f72af5..7218961ef9 in full (setsid --wait on the lean
guards step; supported.yaml re-derived) plus the committed CI .minisig.
Verdict BLOCK: PR1-MUTANTS only (3555 > cap 60), the operator HARD-STOP.

Agent: aprender-88
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* evidence(pr-review #4587): cuda consultation is unreachable, not consulted-with-no-queries

Arm 4 on 293425f2d8 (pr-review-quorum 36504222978) rejected the receipt: B1,
cuda consulted with queries: []. The S3.B path trigger fires on two files and
the CUDA docs tool is not authenticated in the reviewer's session or the
cuda-docs-reviewer's, so no query ran. Recorded as unreachable (allowed with a
non-PASS verdict); the device claims were not checked against docs. Verdict
stays BLOCK. check_pr_review_receipt.sh ACCEPTs a signed copy.

Agent: aprender-88
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(cuda): the fit-check no-op cases assert the path of the profile they called (#4621)

The last loop in fit_check_leaves_non_qk_norm_fp8_and_forced_paths_alone
built two fresh profiles and asserted they were batched, so a mutant that
switched the path to Serial while still returning false passed it.
Each case now asserts the path of the profile it called. Red on that
mutant, green on the code (intel). Quorum finding, PMAT-4621 round 1.

Agent: aprender-ec
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(PMAT-4616): D5 gate reds — the cap is CUDA-free and tested, try_cuda_backend no longer grows

x86-main on 099d07de7d was red on this diff, not on the runner:
- complexity: try_cuda_backend grew (cog 43->46, cyc 14->15). The chat
  pre-flight now rides the tokenize result through fit_serving_context,
  so the function gains no branch.
- mutants: serving_context lived in the cuda-gated dense_session_borrowed,
  where no CPU lib test can reach it. The rule is now
  dense_session::cap_context(context_length, device_kv), CUDA-free, with
  lib tests on every arm; fit_serving_context has lib tests for pass,
  refuse (400) and a kept tokenize error.
- roadmap.yaml regenerated for the PMAT-4616 fragment.
- surface_audit.csv: the /v1/completions row's cited line moved 398 -> 405.

Refs #4616

Agent: aprender-2e-serve
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* evidence(pr-review): #4620 non-author receipt for 65370adc9a (DEGRADED: mutation sweep unreachable; targeted mutant killed)

Agent: aprender-ec
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* evidence(pr-review): drop the superseded fdc165f50d receipt from #4620's tree

Its findings.sarif carries a placeholder excerpt_sha256 ("n/a-short-code-idiom-not-hashed")
that S1.1 would reject once signed, and the receipt binds that sarif by sha256, so the
author cannot correct it without editing the reviewer's record. It is superseded by the
65370adc9a receipt (the head's code) and stays in history at 18b6c8341f. Quorum finding,
PMAT-4621 round 2 (gemini lane).

Agent: aprender-ec
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(audits): quorum PMAT-4621 AGREED 3/3 (sonnet-5, gemini-3.1-pro-high, haiku-4-5) at 2b8d79acd7

Agent: aprender-ec
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* evidence(pr-review): stamp #4620's receipt with the diff patch-id Arm 4 computes (b3f1fd942a)

Arm 4 A2 reads predicate.diff_patch_id from the COMMITTED receipt; the CI signer stamps
only its runner copy, so an unstamped receipt is legacy and RED. Computed with
scripts/lib/pr_review_patch_id.sh prpid_compute over 722fd1621f..HEAD (evidence and
quorum json excluded), equal to the value Arm 4 printed in run 36510383864.

Agent: aprender-ec
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(PMAT-4616): CUDA /generate and /batch/generate map ContextLimitExceeded to 400

Quorum lane 2 (gemini-3.1-pro-high) at 926edb2135 found two GPU-resident
generate paths that still hardcoded 500 for every generation error,
bypassing generation_error_status(): batch_processing.rs try_cuda_generate
and batch.rs try_cuda_batch_generate. Both now route the status through it,
so an over-length request is a 400 on the CUDA path as on the CPU path.
cuda clippy on intel: only the pre-existing residual.rs:314/343 borrow_deref_ref.

Agent: aprender-2e-serve
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* evidence(pr-review): #4620 receipt passes the whole guard, not only the signature arm

Arm 4 (run 36511265641) verified the signature, then rejected the content: the
antigravity divergence counted 1 against 0 findings. The reviewer corrected it to the
agy output it actually recorded (findings: []), added the pmat duplication fields the
guard requires, renamed two SARIF drivers to the allowed set, and re-bound
findings_ref. The guard ACCEPTs it under a throwaway key. diff_patch_id is unchanged.

Agent: aprender-ec
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(PMAT-4616): every serve generate/logprobs/perplexity error routes through generation_error_status

Quorum at 1ead14b2d7 (sonnet-5 + gemini-3.1-pro-high FAIL) found more
handlers that hardcoded 500 on a generation error, so ContextLimitExceeded
still leaked as 500: SafeTensors CUDA chat, GPU /generate, logprobs,
perplexity, the APR/demo batch + stream paths and the openai GPU fallback.
Swept all 12 generate-error sites in api/ (decode errors stay 500: those are
server faults). generation_error_status keeps every non-context error at
500, so only the over-length case changes.

intel: clippy --lib default rc0; --features cuda only the pre-existing
residual.rs:314/343 borrow_deref_ref; 15 targeted lib tests pass. Local
mutants not_measured (unmutated baseline timed out at 600s).

Agent: aprender-2e-serve
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(guard): C13 read a present release asset as MISSING (printf into grep -q under pipefail)

scripts/check_release_assets.sh:113, under the file's own `set -uo pipefail`:

    if printf '%s\n' "$have" | grep -qxF "$want"; then

grep -q exits on its first match; printf then takes SIGPIPE (141), pipefail
reports the pipeline as 141 and the `if` takes the MISSING branch though grep
matched. This turned guard-tree RED in CI run 36510076745 (step 26,
release_criteria.sh --self-test row 6, "C13 is credited when the release
carries all sixteen assets", rc=1 wanted 0) on a code delta that was
evidence-only; the identical merge passed 10/10 three times locally.

Fix: a here-string (`grep -qxF -- "$want" <<< "$have"`), no pipeline.
Same class as 0c6932fd54. The C0 leg-1 check in release_criteria.sh:85
had the inverse form (a matched '✗' could read as no failure) and gets the
same fix; that one could only ever PASS wrongly.

Regression row (check_release_assets.sh --selftest row 7): a complete
release that also carries 20000 other assets, so $have outgrows the 64 KiB
pipe buffer and the race becomes deterministic. Old script: rc=1 (x3).
New script: rc=0 (x3). selftest 12/12; release_criteria --self-test 10/10.

Agent: aprender-88
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* evidence(pr-review #4587): Sonnet receipt for code head 92aac9b4e3 (BLOCK: PR1-MUTANTS, unchanged)

Supersedes the 7218961ef9 receipt. Reviews the one-commit delta 92aac9b4e3
(printf|grep -q under pipefail -> here-string in check_release_assets.sh:113
and release_criteria.sh:85, plus the 20016-asset regression row). Reviewer
reproduced old rc=1 x3 / new rc=0 x3; selftests 12/12 and 10/10.
diff_patch_id 5a638e1100a63e65b81061815d2f8581ce08644d (base ef72c05..92aac9b4e3).
Signature comes from CI pr-review-sign on the pushed head.

Agent: aprender-88
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* split(#4621): #4620 carries roots 1+5 only (chunked dense prefill, F16/IQ4_XS device dequant)

The diff held 97 mutants against the gate's cap of 60. The QK-norm fit check
(root 2) and the chat session KV (root 3) move to their own PR, per the cop's
split ruling; they are restored here to the merge base, unchanged in content.

Refs #3715
Agent: aprender-ec
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* evidence(PMAT-4616): pr-review receipt for 7008062734 (reviewer sonnet-5, DEGRADED, advisory findings)

Non-author reviewer (Sonnet 5). Verdict DEGRADED: cuda (docs MCP
OAuth-gated) and mutation (no cargo test on lambda) unreachable. Confirms the
167bd5e126 receipt's finding is fixed (try_cuda_generate /
try_cuda_batch_generate now route through generation_error_status).
Guard rejects only B1 (unsigned) until the CI signer's .minisig is committed.

Advisory follow-ups, not folded (changing the diff voids this receipt):
batch_processing.rs:343 batch_generate_gpu still 500; /generate and
/batch/generate have no serve_context_refusal pre-flight; the
effective_max_tokens `>` vs serve_context_refusal `>=` boundary.

Agent: aprender-2e-serve
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* evidence(PMAT-4616): bind the 7008062734 receipt to diff patch-id 2fe3407e2f

The reviewer (sonnet-5) had left a placeholder in predicate.diff_patch_id;
it now carries the pinned prpid value, equal to the one CI's Arm 4 measured
(base 2817c6d97b). Guard: only B1 (unsigned) until the signer's .minisig.

Agent: aprender-2e-serve
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* evidence(PMAT-4616): 7008062734 receipt — crux block and citations made guard-complete

The signature exposed a B1 the unsigned state had hidden: crux was
`consulted` without surfaces/comparative_claims. Reviewer (sonnet-5) filled
the S3.C arrays honestly and added the missing citation fields; verified
ACCEPT under a throwaway key. patch-id binding unchanged (2fe3407e2f).

Agent: aprender-2e-serve
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(#4621): a chunked prefill's later chunks never replay the prefill graph

PREFILL_GRAPH=1 (opt-in, off by default) replays a graph that bakes positions
0..S-1 and resets the KV lengths to 0. With the chunked dense prefill, chunk 2+
of the same length would have replayed it over a filled KV cache, silently
corrupting it. Only a chunk starting at position 0 may take the graph now; the
rest run eager. Found by the non-author Arm-4 review of #4620.

Refs #4621, Refs #3715
Agent: aprender-ec
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* evidence(pr-review): drop #4620's pre-split receipt (it reviewed roots 2+3, now in the stacked PRs)

Agent: aprender-ec
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* evidence(PMAT-4616): commit the CI signer's signature for the 7008062734 receipt

Agent: aprender-2e-serve
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* evidence(pr-review): #4620 Arm-4 receipt for split head 31cd3f6c40 (PMAT-4621)

Non-author Sonnet 5 review of part A (chunked dense prefill, F16/IQ4_XS
device dequant, PREFILL_GRAPH first-chunk-only). Verdict DEGRADED;
diff_patch_id stamped. CI signer adds the .minisig.

Refs #4621

Agent: aprender-ec
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(serve): CUDA /generate and /batch/generate refuse over-context prompts with 400 before the model lock (PMAT-4616)

Quorum @0c2928642d (gemini, haiku) found two remaining 500 paths:
- try_cuda_generate / try_cuda_batch_generate had no pre-flight, so an
  over-context prompt reached the engine. They now call
  preflight_serving_context, the same check chat and completions run
  through fit_serving_context: a 400 "context exceeds N" before the lock.
- gpu_batch_completions_handler's batch_generate_gpu Err arm hard-coded 500;
  it now maps through generation_error_status (ContextLimitExceeded -> 400).

Two tests: no cap -> Ok; cap 3 -> 2 Ok, 3 refused with 400.
Roadmap entry: title and criterion 3 name every covered route; criterion 1
states that the --context-length 32768 -> 200 case is #4603, not this ticket.

Complexity vs main: cyclomatic and cognitive unchanged on all touched fns.
intel: clippy default rc0; cuda only pre-existing residual.rs; d5_/cap_context 13 passed.

Agent: aprender-2e-serve
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(serve): /batch/generate pre-flights every prompt before taking the CUDA write lock (PMAT-4616)

Quorum @7bcdde94bd (sonnet) FAIL: try_cuda_batch_generate ran
preflight_serving_context inside the generate loop, after
cuda_model_lock.write(), so the doc/commit claim "before the model lock" was
false for /batch/generate. Tokenize, empty-check and pre-flight all prompts
first, then lock and generate: an over-context prompt anywhere in the batch is
a 400 with no lock taken and no GPU work spent on earlier prompts.

intel: clippy default rc0; cuda only pre-existing residual.rs:314/343;
d5_/cap_context 13 passed.

Agent: aprender-2e-serve
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* evidence(pr-review): #4620 post-split quorum (3/3 AGREED, scope A) + S4.2-compliant receipt SARIF (PMAT-4621)

The quorum json is replaced with the round at head 6562916bee, which covers
part A only (the pre-split file claimed root-3 chat KV, now C's).
The receipt SARIF had its runs moved from driver 'pr-review-primary' to one
run per consulted consultation, and findings_ref is rehashed. Same subject
and patch-id.

Refs #4621

Agent: aprender-ec
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(serve): /v1/batch/completions pre-flights the serving context (PMAT-4616)

Non-author review @30b29b3798 (blocking, measured): gpu_batch_completions_handler
had no serving-context pre-flight, and its GPU callee batch_generate_gpu has no
context guard at all, so mapping its Err through generation_error_status could
not turn an over-device-KV prompt into a 400. Pre-flight every encoded prompt
right after tokenization, before either the GPU or the CPU branch generates.

intel: clippy default rc0; cuda only pre-existing residual.rs:314/343;
d5_/cap_context/batch_completions 32 passed.

Agent: aprender-2e-serve
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* evidence(pr-review #4587): cuda consultation run for real (consulted, 5 queries, 0 contradicted)

The nvidia-cuda-docs MCP was signed in 2026-09-29 06:04Z. A headless
claude-sonnet-5 lane (not the author's model) ran 5 named queries on the FP8
E4M3/E5M2 claims in fp8-interchange-v1.lean and the cuda_simd_scores fixture:
4 found, 1 no-authority-found, 0 contradicted. Every excerpt_sha256 verified.
Replaces the "unreachable" block; PRREV-CUDA-001 updated; findings_ref re-hashed.
Verdict unchanged: BLOCK (PR1-MUTANTS).

Agent: aprender-88
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(serve): /v1/batch/completions pre-flights against the cached model's own context (PMAT-4616)

Non-author review @ad102790a4 (blocking, measured): the pre-flight read
AppState::serving_context, which every cached-model constructor sets to None,
so on the only states this handler serves it was a runtime no-op. And
batch_generate_gpu sizes KV caches from prompt + max_tokens with no context
check. Split preflight_context(Option<usize>, n) out of
preflight_serving_context; the handler refuses against the device serving
context when set, else cached_model.model().config.context_length.
+1 test for the explicit-context arm.

intel: clippy default rc0; cuda only pre-existing residual.rs:314/343;
d5_/cap_context/batch_completions 33 passed.

Agent: aprender-2e-serve
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(mutants-gate): a mutant in code the build never compiles is NOT_MEASURED, never MISSED (PMAT-4621)

cargo-mutants tests a mutant with `cargo test -p <pkg> --lib` at default
features, so a mutant in #[cfg(feature = "cuda")] code is never compiled and
its run is the baseline. On #4620 (run 36521053810, artifact 11015244648)
all 7 mutants were in cfg(cuda) code. 0 cuda:: tests ran of 16149, and the
MISSED/TIMEOUT split was timing only (292 s vs 300.1 s).

The compiled set is now measured from dep-info (cargo check -p <pkg> --lib in a
private target dir). A survivor outside it is NOT_MEASURED and stays RED unless
evidence/mutants-local/*.txt carries 'CAUGHT <line>' from a build that compiles
it. Compiled survivors stay RED as before. A failed check or empty dep-info is
RED. No limit changed.

Red/green:
- Self-test: 5 new rows, and 5 rule mutants are each killed.
- Real artifact, intel, real cargo: 7/7 NOT_MEASURED OPEN, rc=1. The
  dep-info names 488 aprender-serve files and 0 under src/cuda/.

Refs #4621

Agent: aprender-ec
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* ci(mutants): A' survivor table -- CI matrix shards bound to head sha + diff hash, verified before trusted

Operator ruling 2026-09-29 06:36Z (via the cop): CI generates the survivor table
itself; any gap is RED. Scope PR-1 and PR-2 of 0.70 only (ci/mutants-table-refs.txt).

- mutants-table-scope: decides in_scope from ci/mutants-table-refs.txt at the head
  sha and re-proves the gate rule + checker case tables on every PR.
- mutants-shard: 16 matrix jobs on the PR HEAD sha, cap 0, each writes the full
  --list universe, its slice's outcomes.json and meta (head sha, diff sha256, run).
- mutants-table: shard matrix must be success, then
  scripts/ci/mutants_survivor_table.py checks every listed id exactly once,
  killed/equivalent/survived, equivalents need a proof + a quorum receipt naming
  the id, survivors <= MUTANTS_MAX_MISSED, every shard's sha and diff hash match.
- gate: GATE-MUTANTS-TABLE-RULE requires mutants-table success for a listed ref
  and a scope decision on every pull_request (case table:
  scripts/check_ci_gate_mutants_table_rule.sh, 14 rows, 3 planted weak rules RED).
- mutants section: hands over for a listed ref only; the 60 cap is unchanged for
  every other PR.

Checker case table 18/18; each of 10 mutations of the checker turns a row RED.

Agent: aprender-88
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(mutants-table): an equivalent's receipt must be a pr-review statement, not any file naming the id

Sonnet review F1 on a79535bd51: the equivalents row was accepted when ANY repo
file contained the mutant id, so a stray dump could reclassify a real survivor.
The receipt must now sit at evidence/pr-review/<pr>/<sha>/receipt.intoto.jsonl
and parse as an in-toto pr-review Statement. Two new RED rows (stray file,
right path but not a statement); disabling either check turns them RED.

Agent: aprender-88
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(PMAT-4616): scope criterion 3 to dense GGUF CUDA; APR Q4K split to #4625

Quorum @1531ead607 (sonnet) FAIL on three scheduler-channel sites:
- cuda_chat_backend continuous-batch arm and gpu_completions_handler's
  token_rx loop map scheduler errors to 500. Both run only after the D5
  pre-flight (fit_serving_context at cuda_chat_backend.rs:126, the
  serve_context_refusal check in the completions handler), and the channel
  carries String, not RealizarError, so there is no ContextLimitExceeded to
  map. Comment added at the chat arm, matching the completions one.
- APR Q4K chat/completions backends have no pre-flight: a separate engine
  bounded by max_position_embeddings that its AppState constructors never
  see. Filed #4625; criterion 3 now names the dense GGUF CUDA routes and
  "typed" generate-error sites, and points at #4625.

Agent: aprender-2e-serve
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(mutants-table): an equivalent needs a CI-signed receipt, verified with the BASE branch key

Sonnet re-review F2 on 0a5c1325c7: a shape-checked receipt could be hand-written
by the PR author in the same PR. An equivalents row now also needs
reviewer_actor != author_actor and <receipt>.minisig verifying with minisign
against --pubkey; mutants-table passes the BASE branch's .github/pr-review.pub
(git show origin/<base>:...), which the PR cannot edit, and only the CI signer
holds the secret half. No --pubkey or no minisign with a row is RED. minisign
comes from the pinned installer in both jobs. Four new RED rows (self-review,
unsigned, other key, no --pubkey); disabling each check turns the table RED.

Agent: aprender-88
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* evidence(pr-review #4587): Sonnet receipt for code head b132bf3c34 (A-prime survivor table; F1+F2 fixed)

Three Sonnet lanes over the 3-commit delta since d7b36adff0: a79535bd51 APPROVE
(F1-F3 minor/nit), 0a5c1325c7 BLOCK (F2), b132bf3c34 APPROVE. Verdict stays
BLOCK only on PR1-MUTANTS, now judged by the required mutants-table job. The
cuda consultation is carried unchanged (no device code in the delta).

Agent: aprender-88
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(cli_roles): kill the 14 M1 shard-4 survivors (PR-1 #4587)

Under --features cli-roles, 14 of 26 PR-1 mutants in cli_roles.rs survived:
the role and sampling spellings, SamplingKind::id, SamplingArg::groups and
kind_of, strings and path_bufs had no test that checked their values. The
four tests pin each spelling and id, the one group per kind, the kind of
each sampling argument (and none for a plain one), and the order-preserving
conversions.

cargo-mutants 27.1.0, --features cli-roles: 25 caught, 1 unviable, 0 missed
(was 11 caught, 14 missed). Without the feature the module is not compiled
and every mutant reads MISSED; that is the gate's feature gap, not these tests.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-3d

* fix(mutants-table): feature-gated files run with their feature, not as false MISSED (#4587 A')

cargo-mutants tests with default features, so every mutant in a module behind
#[cfg(feature = "x")] never compiled and read MISSED whatever its tests said
(M1 FALSE-RED: cli_roles.rs 26/26 missed; main.rs under update-check). A
workspace-wide --features pkg/feat is refused by cargo for any selected
package that does not own the feature, so it cannot go on the one run.

ci/mutants-feature-groups.txt names each gated file with its package,
features and test args. mutants_table_shard.sh excludes those files from the
default run, runs each row as its own `cargo mutants -p PKG --features F
--file FILE --shard K-1/N`, and merges the outcomes (perl JSON::PP; the
sovereign-ci image has no python3). list.txt, the universe, is unchanged: a
row that tests nothing leaves its mutants MISSING, so it is RED, never a skip.
cuda rows are refused.

Red/green proof, scripts/check_mutants_feature_groups.sh, wired next to the
gate-rule self-test in ci.yml: the REAL shard script + REAL checker over a
fake cargo -> all caught GREEN; planted survivor in a gated file RED; planted
survivor in a default file RED; group tests nothing RED; no group rows (the
old behaviour) RED; group writes no outcomes DEAD; cuda row DEAD. Mutating
the merge to drop group outcomes turns the GREEN case RED.

Measured on intel with the kill tests (prev commit, from aprender-3d):
aprender-common --features cli-roles: 26 tested, 24 caught, 2 unviable, 0 missed.

Agent: aprender-88
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* evidence(pr-review #4587): Sonnet receipt for code head 05c9e6f2d9 (feature groups + cli_roles kill tests; BLOCK: PR1-MUTANTS until the table runs)

Agent: aprender-88
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(ontology): kill CellsError::file -> String::new() survivor (M1 shard 5)

Both refusal tests now assert the file the error names, so the mutant
that blanks it is caught.

Agent: aprender-27
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(ontology): kill version_key || -> && and v_star > -> >= survivors (M1 shard 5)

"0.+69" is refused only by the digit check, since u64::from_str takes a
leading +. Two spellings of one version key keep the first as V*, and
cells are read at that spelling. Also rustfmt the file() assertion.

Agent: aprender-27
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* evidence(pr-review #4587): Sonnet receipt for code head 38cabdb8af (+ M1 shard-5 kill tests; BLOCK: PR1-MUTANTS until the table runs)

Agent: aprender-88
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(release): publish-order.txt omits aprender-build-sha and aprender-update (PR-1 #4587)

PR-1 adds two publishable crates, but publish-order.txt does not list them.
publish_strict.sh requires the order to equal cascade_universe.py's set,
so the pv 0.69.4 cascade would stop before uploading anything:
"in the universe but not in the publish order, so never uploaded:
aprender-build-sha aprender-update".

Neither crate depends on a workspace crate, and contracts-cli, apr-cli
and 24 others depend on them, so they go first. Re-proved at this tree:
the set difference with the universe is 0, and the publish_strict
dependency-order check reports 0 violations.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-3d

* evidence(pr-review #4587): Sonnet receipt for code head 1028fdf9f9 (+ publish-order leaves; BLOCK: PR1-MUTANTS until the table runs)

Agent: aprender-88
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(build): watch only paths that exist, so cargo stops rebuilding every invocation (PMAT-4616)

cargo-mutants on #4614 (run 36520343459) reported 9 TIMEOUTs on the D5
functions. The test phase was compiling rather than hanging: 4m18s of the
300s budget went to rebuilding aprender-compute, aprender-serve,
aprender-present-core and aprender-present-terminal. The baseline suite
takes 32s. Each of those build scripts emits `rerun-if-changed` on a path
that does not exist:
- the pre-monorepo sibling ../../../provable-contracts/... for compute,
  present-core, and serve's binding, architecture-requirements and
  tensor-names reads;
- `.git-sha` in build_sha::emit, which no crate in the tree carries.
Cargo reruns a build script whose watched path is missing on every
invocation, so each `cargo test` after the build rebuilt those crates and
everything above them.

Each of these sites now emits the watch only when the path exists. The cost:
a sibling file that appears later is not noticed until the script reruns for
another reason. Before, the script reran on every invocation. The in-tree
reads (arch-constraints, contracts/) and the git HEAD triggers are unchanged.
Acceptance criterion 4 of PMAT-4616 records this. #4369 moves the bindings in-tree but leaves serve's two
sibling reads and `.git-sha`, so it does not fix this alone.

Agent: aprender-2e-serve
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(mutants-table): main.rs group runs the pv binary tests, not --lib (TIMEOUT = survived)

main.rs is bin code; --lib tests cannot observe it, and the --lib suite ran
past the 900 s mutant budget on intel, so 'replace main with ()' read TIMEOUT,
which the checker counts as a survivor. binary:: in cli_integration execs pv:
both main.rs mutants caught on intel. Timeout unchanged.

Agent: aprender-88
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(mutants-gate): NOT_MEASURED is RED with no local-receipt bypass (operator E1) (PMAT-4621)

A committed "CAUGHT <line>" text file could turn a mutant CI never built
green (e6 requorum finding). Operator ruling E1 (2026-09-29 07:55Z):
cuda-gated mutants are measured only by a CI shard on a CUDA runner, and
not_measured > 0 on changed gated code is RED. The gate now lists each
NOT_MEASURED mutant, prints the count and fails; the self-test plants a
receipt and a mutant that honours it is killed.

Agent: aprender-ec
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(mutants-table): scratch copies keep .git (--copy-vcs true); every baseline failed without it

Run 36538406766 on 8a3abba725: every default-run baseline FAILED on
lint::tests::lint_passes_on_real_contracts ('no BASE to compare with: neither
merge-base(HEAD, origin/main) nor origin/main resolves'): cargo-mutants copies
the tree without .git, so 3527 mutants were never tested (MISSING -> RED).
--copy-vcs true on the default and feature-group runs gives the copy the
checkout's .git (fetch-depth 0 + the fetched base), as workspace-test has.
Intel: aprender-contracts baseline 'ok' with it. Nothing is skipped or relaxed.

Proof: the fake cargo in check_mutants_feature_groups.sh dies on any run without
--copy-vcs; removing it from either call turns every case DEAD (checked both).

Agent: aprender-88
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* evidence(pr-review): #4614 receipt for 8467f2bcb7 (PMAT-4616)

Non-author review receipt for the head carrying the build.rs fix; quorum r5 3/3 PASS.

Agent: aprender-2e-serve
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* ci(mutants): cuda-gated mutants are NOT_MEASURED on the CPU gate and measured on yoga (#4621, operator E1)

cargo-mutants never compiles cfg(feature="cuda") code at default features, so
those mutants were invisible. The CPU mutants section now defers the sorted
NOT_MEASURED set (mutants_diff_gate.sh --defer-unbuilt); the new mutants-cuda
section in the yoga job mutates the diff's feature-gated files at the head sha
(--features aprender-serve/cuda --gated-only); the gate job is RED unless the
two set shas match (GATE-MUTANTS-CUDA-RULE). Missing/failed/skipped = RED.

Guard: scripts/check_ci_gate_mutants_cuda_rule.sh (13-row table, 5 planted
weaker rules all RED). mutants_diff_gate.sh self-test PASS.

Quorum r3 fixes: PMAT-4621.yaml cites the mutants gate (cop ruling 08:32Z);
the review SARIF no longer claims graph_may_run cannot be mutated on intel
(host-side, 4/4 CAUGHT with --features aprender-serve/cuda); findings_ref
sha256 rebound.

Refs #4621
Agent: aprender-ec
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(cuda): model-free kill tests for qwen35_dequant_f32 and prefill_all_layers_gpu (#4621)

The mutants-cuda shard's GPU runner has no TinyLlama, and these two fns were
killed only by model-backed suites, so their mutants would survive there.

- qwen35_dequant_f32: an F32 weight is returned as it is (kills Ok(0)/Ok(1));
  an F16 [2x64] weight dequantizes on the device bit-exact.
- prefill_all_layers_gpu: a wrong embedding length and an uninitialized
  workspace are refused; S=0 is a no-op (kills the Ok(()) body).

Without a device they skip; under mutation a skip survives, so the shard
is RED, never vacuously green (L25). Built on intel with --features cuda.

Refs #4621
Agent: aprender-ec
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore(regen): single-writer regen on PR-2 — README contract_count 1889, ont-ratchet, contracts.nt, roadmap

Generated paths only (batch_fold.sh regen recipe + make ont-ratchet):
- README.md frontmatter/blocks: 1837 -> 1889 (= census.json n_files; PV-ONT-012)
- contracts/lint-baseline.json verified_commands re-measured (F-33 / PV-ONT-013)
- contracts/contracts.nt via pv extract contracts
- docs/roadmaps/roadmap.yaml via make roadmap-aggregate

Fixed point: readme_sync --check, pv extract contracts --check,
roadmap-aggregate-check all 0; pv lint contracts: 0 errors.

Pmat-Ticket: ONT-10
Agent: aprender-98
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(mutants): the feature shard passed --lib to cargo test twice (#4621)

yoga run 36547274574: `cargo test ... --lib --lib -- --exact ...` is a cargo
usage error, so mutants-cuda tested 0 of the 7 listed mutants. The gate
failed closed ("tested 0 mutant(s) but listed 7"), but no mutant was
measured. --lib already reaches cargo test through the test args; drop
the --cargo-arg=--lib copy.

Self-test: gated-scope-and-measured-set now also requires --lib at most
once per invocation; planted mutant lib-passed-twice (the old line) is
killed. bashrs: 0 errors.

Refs #4621
Agent: aprender-ec
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(ont4b): name the nine ONT-10 rest shapes — shapes_n 57 -> 66

PR-2 adds eight binary-<bin>-surface shapes (apr, apr-qa, aprender-data,
aprender-present, aprender-ptop, aprender-simulate, aprender-train-lora,
aprender-train-shell) and binary-aprender-apr-http-mcp. Measured:
pv extract contracts --check reports shapes_n 66 on this branch.

Pmat-Ticket: ONT-10
Agent: aprender-07
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(ont4c5): follow #4590's de-claim — B names qwen3-8b@lambda, not qwen3-1.7b@gx10

#4590 (ca2daad7db) de-claims Qwen3 on gx10 through data, so the measured
domain D no longer holds qwen3-1.7b-q4km@gx10 and gains qwen3-8b-q4km@lambda.
The probe's owed set B and the deleted-row mutation still named the
de-claimed cell: the_rows_probe_is_green_on_this_repository went red
(B ⊄ D) and a_deleted_current_release_row_rejects_naming_the_cell got
exit 0. B now equals the measured D; the mutation deletes a gx10 row
that is still owed (qwen2-1.5b-q4km).

Pmat-Ticket: ONT-10
Agent: aprender-07
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(merge 2aa46a19df): undo two union-merge defects in contracts-cli tests

- pv_surface_gate: 483cc76d4c and 7c418eb75c each added a VS-COUNT-001
  Case; the merge kept both, so CASES has 27 rows over 26 distinct rules
  and gate_self_test_rule_output_depends_on_input (:1136) can never reach
  27. Keep the first row.
- pvl_discharge_leanchecker: 7c418eb75c made `check()` unscoped while
  e85da1a039 added check_unscoped() plus a row that needs the SCOPED path;
  the merge kept both, so scoped_without_user_systemd_declines_and_never_runs_the_checker
  ran the stub unscoped. That row now calls check_scoped().

Neither test file's defect exists on main (the lean test is absent there;
main's pv_surface_gate has no VS-COUNT-001 row).

Pmat-Ticket: ONT-10
Agent: aprender-07
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ci): PR-1 run 36543136454 guard-tree + pr-review-sign reds (PR-introduced)

guard-tree, three guards over code this PR adds:
- check_guard_steps_isolated: ci.yml's survivor-table step ran three guards
  outside `setsid --wait` (#4133). Prefixed; check_mutants_feature_groups.sh
  moves to its own step so it stays bare-wired and guard_tree still runs it.
- check_ci_gate_mutants_table_rule [run]: the checker's case table needs
  minisign, which the sovereign-ci image (guard-tree) lacks -> ENV -> rc=1.
  The ci.yml step installs minisign first; it now names its input
  (env GATE_RULE_WORKFLOW, which the script reads), so guard_tree rule 1
  leaves the guard to that step. Without minisign it is still rc=1 there
  (measured: PATH=/usr/bin:/bin -> rc=1).
- check_no_silent_truncation: mutants_survivor_table.py:360 printed errs[:2];
  it prints every error now.

pr-review-sign: three unsigned receipts (heads 05c9e6f2d9, 38cabdb8af,
1028fdf9f9; all verdict BLOCK) carried a hand-supplied diff_patch_id the
signer refuses by design (bc7b4296/8b8c5880/4accd62e vs computed
d53bb0c1/789241f1/198c753d over the same base..head). No tool writes that
field; the field is dropped so the signer stamps the computed one. All 7
unsigned receipts sign+verify locally under the test key. A BLOCK receipt
approves nothing, so re-binding it cannot launder an approval.

Local: guard_steps_isolated, no_silent_truncation, guards_are_wired,
ci_gate_mutants_table_rule (+ --self-test via the step), mutants_feature_groups
all rc=0; guard_tree --dry-run: 139 run / 35 skipped, rule guard
"wired-with-args in ci.yml".

Pmat-Ticket: PMAT-4502
Agent: aprender-07
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(mutants): --lib must reach the feature shard's build, once (#4621)

yoga run 36553822552: 451938d114 dropped --cargo-arg=--lib, so --lib no
longer reached cargo-mutants' baseline (`cargo test --no-run`, which takes
only --cargo-arg). Every cuda-gated integration target was compiled, and
tests/driver_cuda_gguf.rs does not build under --features cuda on main
(GGUFConfig initializers lack query_pre_attn_scalar, E0063 x10). Baseline
FAILED, 0/7 tested, and the gate failed closed.

Under --features, pass --lib once, as --cargo-arg (it reaches the build
and the test), and drop it from the test args (a second one is the
usage error of run 36547274574).

Self-test: gated-scope-and-measured-set also requires --cargo-arg=--lib.
Mutants lib-passed-twice and lib-missing-from-build (the 451938d114
shape) are both killed. bashrs: 0 errors.

Refs #4621
Agent: aprender-ec
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(mutants,roadmap): --cargo-arg --lib spelling for bashrs 7.4.1; regenerate roadmap aggregate (#4621)

x86-main on f3077017b1 (job 109390857552), two reds from this branch:

1. check_shell_lint_ratchet grew 6 -> 8 on the runner's bashrs 7.4.1, which
   reads `--cargo-arg=--lib` as a `--` prefix operator (SC2210, lines 56
   and 309; 7.4.2 is clean). Spelled `--cargo-arg --lib`: 0 errors on
   both. Red/green (cargo-mutants 27.1.0, a crate with a non-compiling
   integration test): no arg -> "FAILED Unmutated baseline"; both
   spellings -> baseline `cargo test --no-run --lib`, 4 mutants tested.
2. roadmap.yaml was not regenerated after adding the PMAT-4621 fragment
   (make roadmap-aggregate; --check idempotent).

Self-test PASS, including lib-passed-twice and lib-missing-from-build.

Refs #4621
Agent: aprender-ec
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* roadmap(#4502 PR-1): carry the entries for PR-1's own rows

The split left PR-1's roadmap entries in PR-2, so a Q0 round on PR-1 had no ticket to judge its rows against and every lane failed it on scope. Bytes are identical to PR-2's copies, so PR-2's merge of PR-1 adds nothing.

Pmat-Ticket: PMAT-4502

Agent: aprender-07

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(ontology): a glob re-export cycle reports the real miss, not the cycle

Kills the MISSED mutant from #4588 run 36572203357 shard 9: code.rs:628:27, guard
e.reason.contains("re-export cycle") -> false. Under the mutant the cycle reason
overwrites the last real miss and the new test goes RED.

Pmat-Ticket: PMAT-4502
Agent: aprender-07
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* roadmap(#4502 PR-1): rows for folded #4604 and #4369; regenerate the aggregate

C78(2): PR-1 folds the pv 0.69.4 release set (#4604) and the build_helper::env_key
migration (#4369); each now has its entry. fad4fb0453 added six entries without
regenerating roadmap.yaml (check_roadmap_fragment_required DRIFT); regenerated here.

Pmat-Ticket: PMAT-4502
Agent: aprender-07
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* roadmap(#4502 PR-1): AC rows for the #4587 A' survivor table and the folded-ticket list (C78(2))

Answers the Q0 gemini objection ci.yml:51 (survivor table not asked for by PMAT-4502): it is the operator's 06:36Z ruling on this PR, now an AC row. Makefile label/lint ratchets are PMAT-4139/4166, carried with their own entries.

Pmat-Ticket: PMAT-4502
Agent: aprender-07
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* roadmap: rows for the #4083 (EV-9) and #4189 (G5) folds in PR-2 (C78, cop ruling a)

Pmat-Ticket: PMAT-4502

Agent: aprender-07

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(#3715 B1): receipt §8 describes THIS diff; README/CHANGELOG say 'planned for 0.70.1', not 'fixed' (quorum round 2)

Agent: apr-0d-b1

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* census: drop the batch-regen census edit; README contract_count back to 1837 (cop ruling a, #3569)

census.json has one writer (the release train); it is regenerated once after the batch merges. contracts.nt re-extracted (readme:contractCount 1837); pv extract --check fixed point holds; check_census_derived.sh PASS.

Pmat-Ticket: PMAT-4502

Agent: aprender-07

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* evidence(#3715 B1): Q0 quorum AGREED at 33b032506a (opus-5-5, haiku-4-5, gemini-3.1-pro-high; base split/ont10-rest-0.70)

Agent: apr-0d-b1

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* revert(#4588 split): #4429 x86-main intel pin and #3668 docs-only skip move to PR-2b

Cop ruling 19:35Z (split). Gate-graph walk from rc.1 (rc-cut.yml: ci/gate +
workspace-test; gate.needs closure) and FINAL (paiml-ontology @60179ae2:714,
ONT-10 depends_on): neither row is needed, directly or transitively.

- x86-main: runs-on back to main's [self-hosted, Linux, X64, clean-room].
  The gate needs the job, not the runner label.
- #3668: the `changes` job, guard-cargo's needs/if, the gate's
  GATE-DOCS-ONLY-RULE, ci_docs_only.sh, check_ci_gate_docs_only_rule.sh,
  fat_driver's outputs pass-through and the PMAT-3668 roadmap entry are
  removed. The three guard steps return to guard-cargo. The gate again reads
  guard-cargo and determinism-compare strictly, as main does, so no gate is
  loosened versus main.

Receipts: fat_driver self-test 59/59; check_ci_gate_mutants_rule,
check_guards_are_wired, check_ci_fat_secrets_plumbed,
check_ci_gate_mutants_table_rule rc 0; actionlint 2 findings = HEAD's 2;
make roadmap-aggregate-check ok.

Pmat-Ticket: PMAT-4502
Agent: aprender-07

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* audit(PMAT-4616): Q0 quorum for a8e97a4711 — 3/3 PASS (sonnet-5, gemini-3.1-pro-high, haiku-4-5)

Agent: apr-2e-d5

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* docs(#4588): receipts for 4502/4073/3712/4445/4083; PMAT-4083 scoped to part 1 (advisory)

Cargo-test rows are NOT_MEASURED (intel only) and are not counted as passes.

Pmat-Ticket: PMAT-4502
Agent: aprender-07

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ci): the fat driver's results output dropped section outputs — the NOT_MEASURED gate rule read "" and could never be RED (#4621)

emit_results_output wrote {result, continue_on_error} only, so the gate job's
GATE-MUTANTS-CUDA-RULE (which reads .<section>.outputs.not_measured / measured_sha)
saw an empty string on every run and took the "no deferred set" branch: a fail-open
gate. Found by the non-author pr-review of d07e187b66 (F1); its own case table
passed because its fixture JSON hand-wrote `outputs`.

The results output now carries each section's outputs, a fat_driver self-test row
pins it, and check_ci_gate_mutants_cuda_rule.sh gains a row that feeds the rule the
driver's REAL emission (RED with 3 deferred and no shard). Mutating the driver back
turns both rows RED (52/53 and row 14 FAIL).

Refs #4621

Agent: aprender-ec
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* evidence(pr-review): #4614 receipt for 12c0e28a00 (PMAT-4616), DEGRADED (mutation unreachable), unsigned pending CI signer

Agent: apr-2e-d5

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* docs(#4588): cells-producer receipt row measured (rc 0, 25 ok)

Pmat-Ticket: PMAT-4502
Agent: aprender-07

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(mutants): --defer-unbuilt file's directory is made before it is written (#4621)

x86-main went red on run 36607873893: 'mutants-gate/not_measured.txt: No such
file or directory'. The gate truncated the deferred file before mkdir -p of --out,
and CI hands it a path in a directory a fresh checkout does not have.
Red/green: a new self-test row (defer into a missing dir) and a mutant that
drops the mkdir; the mutant is killed, the self-test passes 58 ok, bashrs 0 errors.

Agent: aprender-ec
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* chore(review): #4620 quorum artifact (r4, 98d7ffb1c7) + pr-review receipt for the reviewed head

Quorum r4: 3/3 PASS but degraded same-family (agy quota exhausted, 429); r3 on fc88268134
was full Q0 3/3 (gemini-3.1-pro-high + sonnet-5 + haiku-4-5). Receipt is unsigned; CI signs it.

Agent: aprender-ec
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* revert(cb200): drop the 599->617 re-baseline row from #4588 (cop ruling 20:31Z)

A raised stored limit is a gate waiver. Head vs base on the same scanner and run measures 621 vs 622, so there are no new findings to fix; the row is only the limit raise and goes.

Pmat-Ticket: PMAT-4502
Agent: aprender-07
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(CB-200): head-vs-base release gate + clear the 4 findings over v0.69.3

dogfood.sh judged CB-200 against a stored baseline (599); main measured 622.
Now: on Fail, compare head to the newest final tag (v0.69.3 = 618), one pmat, one run,
baseline neutralised (scripts/check_cb200_head_vs_base.sh). Refactored the 4 added
definitions (check_valid_under, coverage_serve_shards.main, get_one_patchid,
nightly_manifest.main) with behaviour-identical splits; golden/differential proofs in receipt.

Agent: apr-0d-b1
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* fix(CB-200): split get_one_patchid state into helpers (behaviour-identical)

Agent: apr-0d-b1
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* test(contracts): kill the MISSED mutants CI's survivor table found on #4587

mutants-shard 5/9/10/12 reported MISSED in bindings_gate, evidence_gate,
ratchet_gates, refines_gate, shapes_gate, extract/binary, extract/code,
witness_small and scoring/codebase. Each test pins the behaviour a surviving
mutant changed. aprender-contracts…
… on main (bed17c8) (#4656)

* chore(contracts): discharge-summary.json regenerated by pv discharge run at db86acf (C268#2a)

In-tree pv (cargo build --locked -p aprender-contracts-cli --bin pv), run as
`pv discharge run crates/aprender-contracts-staging/lean` under
systemd-run --user --scope -p MemoryMax=24G, on the ladder's lean cache link
(bfbabfa35e3e25aa): rc 0, 423 s, memory.peak 25769803776 (= the cap), oom_kill 0.
Only tree_sha moves (a31602f -> bed17c8); byte-identical (cmp) to 14a3b1b3a4's file.

Agent: aprender-27
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* evidence(pr-review): #4656 review receipt for 492dcd3

Agent: aprender-27
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* evidence(pr-review): #4656 receipt — index_commit/ancestry re-derived for this PR (A4 B1)

Agent: aprender-27
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* fix(ci): quick-tier test steps never saw the pinned origin/main — git refused the root-owned mount (#4656)

Step 8 pins the event's base as refs/remotes/origin/main, and #4655 gave the
full-tier "Workspace lib tests" container safe.directory=/workspace so git, as
root over the host-owned checkout, would read it. The two quick-tier test
containers (selected crates; tree readers) were left without it. A quick-tier
PR (#4656, run 37052925185, step 21) therefore ran
lint_passes_on_real_contracts with git refusing /workspace ("dubious
ownership"): origin/main did not resolve, the refinement gate skipped, and the
CI assert failed 3/3. #4655 passed only because it ran the full tier.

Red/green (depth-1 clone, base pinned exactly as step 8 does, git as root in a
container over the host-owned mount):
  without the env: rc=128 "detected dubious ownership in repository at '/workspace'"
  with the env:    989cb01, rc=0
Test level, same clone, localhost:5000/sovereign-ci:stable as root, CI=true,
cargo test -p aprender-contracts --lib lint::tests::lint_passes_on_real_contracts:
  without the env: rc=101, "the refinement gate declined ... Skipped { reason:
                   \"no BASE to compare with ...\" }" (the exact CI failure)
  with the env:    rc=0, 1 passed
cargo test -p aprender-contracts --lib (intel): 2289 passed, 0 failed.

Tightens: the refinement gate now RUNS on quick-tier PRs instead of being
skipped. No gate is removed or loosened.

Agent: aprender-27
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* evidence(pr-review): #4656 review receipt for bd66e7c (quick-tier safe.directory)

Agent: aprender-27
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* evidence(pr-review): #4656 receipt — S3.E antigravity consulted (A4 B1)

Agent: aprender-27
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* fix(deps): bump wasmtime 48.0.3 -> 48.0.4 (RUSTSEC-2026-0325/0326/0327)

cargo audit in x86-main sov.security (CI 37061601053, #4656 @ ab68432)
went red on three wasmtime advisories published 2026-10-02. Fixed in
>=48.0.4,<49.0.0. Tool-regenerated: `cargo update -p wasmtime --precise
48.0.4`, Cargo.lock only, no hand edits (C265#10, C269#3).

Receipts (intel, scratch clone at ab68432 + this lock, byte-identical
sha256 5a8e7b32fcac4fec on lambda and intel):
- cargo audit: rc=0, 0 hits for RUSTSEC-2026-032[567]
- cargo deny check advisories: rc=0, "advisories ok"
- wasmtime is reached only via the optional aprender-test-lib/runtime
  feature; that feature already fails on the 1.93.0 pin at base
  (48.0.3 -> cranelift requires rustc 1.95.0), so unchanged by this bump.

Agent: aprender-27
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* evidence(pr-review): #4656 receipt for 98e7256 (wasmtime 48.0.4, RUSTSEC-2026-0325/6/7)

Agent: aprender-27
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* evidence(pr-review): #4656 receipt 98e7256 — stamp diff_patch_id (A2 binding)

diff_patch_id was null. Arm 4 computes the branch diff 989cb01..848a584
(evidence/pr-review/4656 excluded) as 4ac8c02f02a1fbb5a10c8f3257c7cf70fbca74c6;
prpid_compute over 989cb01..98e7256 gives the same id. Only that field changes.

Agent: aprender-27
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…se got binaries (#4657) (#4658)

* fix(binary-release): assets pre-check died under bash -e on missing assets, so no release got binaries

The `assets` job ran `check_release_assets.sh "$TAG"; rc=$?` in the default
`bash -e` step shell. Exit 1 (assets missing: every fresh release) killed the
step before the `case`, the job failed, and every build lane was skipped.
v0.70.0-rc.1 (run 37034899163) shipped zero assets, so install.sh row 12
(`--channel rc`) 404s (run 37040341799). The final v0.70.0 would hit the same.

Now `rc=0; ... || rc=$?`. Case table under bash -e with a stub script:
  exit 0: before present=true      after present=true
  exit 1: before STEP FAILS         after present=false (builds run)
  exit 2: before step fails         after step fails (::error:: refusal kept)
No gate weakened: exit-2 refusal and verify-apr-assets unchanged. actionlint
findings identical head vs base (2 pre-existing SC2016 info).

Agent: aprender-27
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* review(4658): receipt for df432eb

Agent: aprender-27
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…rip found no apr (#4659) (#4660)

* fix(binary-release): mini lane built into a redirected target dir, so strip found no apr (#4659)

Run 37084403773 failed on both attempts: cargo finished into the runner's
redirected CARGO_TARGET_DIR, and strip/version/package all read ./target.
Pin CARGO_TARGET_DIR to the workspace target for the build step. No gate changed.

Agent: aprender-27
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* review(4660): receipt for 1b61c42

Agent: aprender-27
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Guard-step files (#4415) take main's side: main carries the later
ci/sections.yml form of the same work, and this branch's late fixes
(jq stream death, zombie child, no-manifest step) are all present there.
ci/sections.yml takes main's live-API PR-body reads (#3298) and keeps
its run-all guard step. tests/pr-review.bats keeps both new blocks
(L2 attest and the #3594 jq rows). roadmap.yaml is the union.
The mutation set is now 253 (this branch's 249 plus main's four);
its kill count stays pending until the receipt job's Arm 3 runs.

Refs #4472 #4503 #4517 #4533
Agent: aprender-78
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@noahgift

noahgift commented Oct 4, 2026

Copy link
Copy Markdown
Contributor Author

Closed for the open-PR cap: this changes what a gate accepts and waits for the operator's sign-off batch (0.70.2 ledger). Branch kept; reopens on a yes.

@noahgift noahgift closed this Oct 4, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant