Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
75 commits
Select commit Hold shift + click to select a range
a08a6bd
chore(roadmap): PMAT-989 kind:code label (kind-gate)
noahgift Sep 6, 2026
d36c409
docs(PP-066): R-0 design quorum record (3/3 implement-with-changes, 3…
noahgift Sep 6, 2026
1ef2a22
test(R-0a): registry_case_table — RED: trueno::registry does not exis…
noahgift Sep 6, 2026
5795ccc
feat(R-0a): trueno::registry — BackendRegistry::discover(): cpu alway…
noahgift Sep 6, 2026
121fa43
feat(R-0a): apr devices [--json] — the registry printed (every kind a…
noahgift Sep 6, 2026
8df3519
contract(R-0a): apr-backend-registry-v1 (invariants i and iii dischar…
noahgift Sep 6, 2026
ae5e9cb
docs(audits): impl-PMAT-989 receipt (R-0a, status: partial) + estimat…
noahgift Sep 6, 2026
3c6cbea
fix(R-0a): review quorum changes — the CLI contract lists devices onc…
noahgift Sep 6, 2026
1acdc2e
docs(audits): impl-PMAT-989 receipt — review quorum record and adjudi…
noahgift Sep 6, 2026
871fd49
fix(R-0a): two different cards with one name through one API stay two…
noahgift Sep 6, 2026
b83c231
Merge remote-tracking branch 'origin/main' into agent/R-0
noahgift Sep 6, 2026
ad574f2
docs(PP-066): §5.0 re-rendered on the merged tree (G-10 + R-0b, 92 ro…
noahgift Sep 6, 2026
4987923
docs(R-0): book chapter book/src/cli/devices.md + page contract apr-p…
noahgift Sep 6, 2026
020e926
docs(R-0): README contract count 1814 on the third line too
noahgift Sep 6, 2026
bcf1227
docs(R-0): DAG/spec edits are carried by the R-0 amendment PR (#3003)…
noahgift Sep 6, 2026
ea697a3
docs(R-0): the design-quorum record is carried by #3003 too
noahgift Sep 6, 2026
caa0e68
docs(R-0): roadmap entry edits for PMAT-989 are carried by #3003; thi…
noahgift Sep 6, 2026
99f961d
evidence(L0-1): RED-first record on lambda (RTX 4090, apr 0.65.2 cuda…
noahgift Sep 6, 2026
41fd2f4
feat(L0-1): the model manifest is DERIVED (scripts/derive_model_manif…
noahgift Sep 6, 2026
c3ea066
docs(L0-1): receipt PMAT-1065 (partial) — RED-first record, what the …
noahgift Sep 6, 2026
c3d29d2
docs(L0-1): root-cause quorum folded — position 0 is token 785 not BO…
noahgift Sep 6, 2026
006ae7e
docs(L0-1): the row's issue is #2971 (the user report; #3017 closed a…
noahgift Sep 6, 2026
beb2fa4
feat(L0-1a): contract apr-gpu-cpu-parity-v1 (PAR-OB-001..003; the hor…
noahgift Sep 6, 2026
b3736d1
feat(L0-1a): REG-15 model admission — ParityVerdict/Admission (admit:…
noahgift Sep 6, 2026
29b9cef
feat(L0-1a): REG-15 wired — on_cuda_load_error() decides every CUDA l…
noahgift Sep 6, 2026
3778ad1
docs(L0-1a): receipt — REG-15 landed (module, tests, chat wiring, com…
noahgift Sep 6, 2026
f37b1bd
feat(L0-1a): dogfood gets C14 — the must-RED twin as the row's own fa…
noahgift Sep 6, 2026
82e5e6a
test(L0-1a): the PR-time gate is two sentinels in workspace-test — th…
noahgift Sep 6, 2026
75e60cf
evidence(L0-1a): threshold basis from n=5 — the known-good pair (7B@l…
noahgift Sep 6, 2026
499a9cc
feat(L0-1a): GET /v1/effective-config carries parity:{status,cosine,p…
noahgift Sep 6, 2026
baaad3c
fix(L0-1a): OwnedQuantizedModelCuda keeps its doc comment (missing_do…
noahgift Sep 6, 2026
d6aba8a
docs(L0-1a): accept.sh judges the recorded runs and runs the three te…
noahgift Sep 6, 2026
3edfd5b
evidence(L0-1a): manifest re-derived after the receipt relabel (the c…
noahgift Sep 6, 2026
d56d793
mutants(L0-1a) round 1: a hand-typed manifest entry nothing cites (de…
noahgift Sep 6, 2026
0b1e683
fix(L0-1a): review fold — a threshold without a basis is refused (I4,…
noahgift Sep 6, 2026
b4c8d59
fix(L0-1a): diff-benchmark's GPU half extracted into gpu_profile_or_n…
noahgift Sep 6, 2026
d481487
docs(L0-1a): review-lane record and its dispositions (six fixed, two …
noahgift Sep 6, 2026
3e30a62
fix(L0-1a): shell-lint — no single-quote escape idiom in the UNMEASUR…
noahgift Sep 6, 2026
4e90fb4
Merge remote-tracking branch 'origin/main' into agent/R-0
noahgift Sep 6, 2026
27f5ff9
fix(R-0a): delete the two complexity_baseline rows this PR made STALE
noahgift Sep 6, 2026
4a66e20
test(R-0a): MUTANT — drop the pass-1 refusals so the reserve refusal …
noahgift Sep 6, 2026
5901220
chore(L0-1a): delete the three complexity_baseline rows this branch m…
noahgift Sep 6, 2026
665b711
fix(L0-1a): gpu_profile_or_none names its error type — the bare Resul…
noahgift Sep 6, 2026
193d626
test(L0-1a): MUTANT (round 2) — thresholds.yaml min_positions: 1; the…
noahgift Sep 6, 2026
fae3b7f
Revert "test(R-0a): MUTANT" + receipt status: complete with the mutat…
noahgift Sep 6, 2026
879c3e1
audit(R-0a): re-audit the surface ledger for the code this PR moved (…
noahgift Sep 6, 2026
0ef312c
test(L0-1a): revert the round-2 mutant — min_positions back to 64 (ba…
noahgift Sep 6, 2026
2498f40
docs(R-0): receipt records the surface-ledger re-audit as resume item…
noahgift Sep 6, 2026
6fecabe
fix(R-0): the three per-{binary,band,cluster} baselines the new row m…
noahgift Sep 6, 2026
c11ec8c
Merge branch 'main' into agent/L0-1
noahgift Sep 7, 2026
eb9642a
test(L0-1a): the effective-config key set carries parity (REG-15) — R…
noahgift Sep 7, 2026
059b5c4
Merge branch 'agent/L0-1' of github.com:paiml/aprender into agent/L0-1
noahgift Sep 7, 2026
fffdfa3
Merge branch 'main' into agent/R-0
noahgift Sep 7, 2026
dff0225
plan(R-0b): P0 discover + P1 plan and accept.sh (RED, 4 legs) — resol…
noahgift Sep 7, 2026
f3bb33e
Merge remote-tracking branch 'origin/agent/R-0' into agent/R-0b
noahgift Sep 7, 2026
d89db05
Merge remote-tracking branch 'origin/agent/L0-1' into agent/R-0b
noahgift Sep 7, 2026
9610dd5
feat(R-0b): backend resolution module reads the registry — a forced a…
noahgift Sep 7, 2026
278417b
fix(L0-1a): drop a redundant #[must_use] on ParityBlock::not_run — th…
noahgift Sep 7, 2026
4057329
Merge branch 'agent/L0-1' into agent/R-0b
noahgift Sep 7, 2026
99589f9
feat(R-0b): the last two backend cfg! reads read the registry; check_…
noahgift Sep 7, 2026
805342e
feat(R-0b): selected: line on every run/chat/serve, parity: line at t…
noahgift Sep 7, 2026
a0269f3
feat(R-0b): GET /v1/effective-config reports the startup backend reso…
noahgift Sep 7, 2026
5b5116d
fix(R-0b): a forced accelerator that fell to CPU at runtime is refuse…
noahgift Sep 7, 2026
c93f74a
docs(R-0b): receipt — quorum folded, gaps narrowed to the CI pair, fl…
noahgift Sep 7, 2026
860ccfb
Merge origin/main into agent/L0-1 — re-anchor the L0-1a manifest step…
noahgift Sep 8, 2026
176b986
feat(L0-1a): the second required host — gx10/GB10 measured, the thres…
noahgift Sep 8, 2026
68ac4d6
Merge origin/main into agent/R-0b — re-anchor the R-0b and L0-1a guar…
noahgift Sep 8, 2026
f48cd90
fix(L0-1a): the parity records carry no machine path, and the n=5 ser…
noahgift Sep 8, 2026
48053da
docs(L0-1a): the receipt records why the parity evidence carries no m…
noahgift Sep 8, 2026
799760f
Merge remote-tracking branch 'origin/agent/L0-1' into agent/R-0b
noahgift Sep 8, 2026
a8faded
fix(R-0b): regenerate the derived tree-reader registry — this row add…
noahgift Sep 8, 2026
f8a94b3
fix(R-0b): backend_refusal_case_table was wired into NOTHING — this r…
noahgift Sep 8, 2026
d4e1f41
fix(L0-1a): reg15_admission was named by NO workflow — the row's own …
noahgift Sep 8, 2026
6b8bc97
Merge agent/L0-1 into agent/R-0b — the gated test line takes the UNIO…
noahgift Sep 8, 2026
245ca32
fix(R-0b): throughput figures leave rustdoc — this row promoted them …
noahgift Sep 8, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .claude/skills/apr-dogfood/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -1476,7 +1476,7 @@ The nine clusters at zero, largest first: `http-orchestrate-banco` (95),
`test-harness` (49), `rag-eval` (44), `qa-cgp` (37), `simulation` (18),
`orchestrate-pacha-secrets` (17).

**Report both numbers or neither (T2).** "5 of 14 clusters gated (35.7%)" without "143 of 833 features gated (17.2%)" beside it is a proxy masquerading as coverage, and the gate refuses to emit it — *on this line too*.
**Report both numbers or neither (T2).** "5 of 14 clusters gated (35.7%)" without "144 of 834 features gated (17.3%)" beside it is a proxy masquerading as coverage, and the gate refuses to emit it — *on this line too*.

The rule is about the NUMBER, not about a phrasing, and it is enforced on every
surface that can emit one: the gate's own report, the receipt, the output of
Expand Down
19 changes: 18 additions & 1 deletion .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -485,7 +485,7 @@ jobs:
-e CARGO_INCREMENTAL=0 \
-e CARGO_BUILD_JOBS=8 \
"$IMAGE" \
bash -c 'cargo test -p aprender-core --features setfit --lib setfit && cargo test -p aprender-core --features setfit,conformance-fixtures --test setfit_conformance && cargo test -p aprender-core --test monorepo_invariants && cargo test -p aprender-core --test readme_contract && cargo test -p apr-cli --test cli_commands && cargo test -p aprender-train-inspect --test falsify_no_fabricated_metadata_2519 && cargo test -p aprender-train-bench --test falsify_no_fabricated_benchmarks_2519 && cargo test -p aprender-train-shell --test falsify_no_fabricated_fetch_2519 && cargo test -p aprender-core --test beat_sklearn_iris && cargo test -p aprender-core --test beat_sklearn_nmi && cargo test -p aprender-core --test beat_sklearn_metrics_parity && cargo test -p aprender-core --test beat_sklearn_gaussiannb_accuracy && cargo test -p aprender-core --test beat_sklearn_svc_accuracy && cargo test -p aprender-core --test beat_sklearn_pipeline_encoder && cargo test -p aprender-serve --test beat_fail_closed_garbage && cargo test -p aprender-compute --lib beat_nf4_bitsandbytes_equivalence && cargo test -p aprender-core --test beat_pytorch_autograd_grad && cargo test -p aprender-train-lora --lib beat_lora_merge_forward_equivalence && cargo test -p apr-cli --release --test beat_pytorch_deploy_footprint && cargo test -p aprender-serve --test beat_fail_closed_structural && cargo test -p aprender-serve --test ollama_http_compat && cargo test -p apr-cli --test ollama_ndjson_streaming && cargo test -p apr-cli --test falsification_chat_http_cli && cargo test -p aprender-contracts --test apr_serve_api_key_auth_contract && cargo test -p apr-cli --test falsify_auth_001 --test falsify_auth_002 --test falsify_auth_003 --no-fail-fast && cargo test -p apr-cli --test beat_apr_data_alimentar_reach && cargo test -p apr-cli --test beat_apr_sibling_cli_reach && cargo test -p aprender-contracts-cli --test pv_surface_gate && cargo test -p aprender-contracts-cli --test cli_integration && cargo test -p aprender-contracts-cli --test ground_truth && cargo test -p aprender-contracts-cli --test version_identity && cargo build --examples --workspace --keep-going && cargo check --workspace --benches --locked'
bash -c 'cargo test -p aprender-core --features setfit --lib setfit && cargo test -p aprender-core --features setfit,conformance-fixtures --test setfit_conformance && cargo test -p aprender-core --test monorepo_invariants && cargo test -p aprender-core --test readme_contract && cargo test -p apr-cli --test cli_commands && cargo test -p apr-cli --test reg15_admission && cargo test -p apr-cli --test registry_failure_catalogue && cargo test -p apr-cli --test backend_refusal_case_table && cargo test -p aprender-compute --test registry_case_table && cargo test -p aprender-train-inspect --test falsify_no_fabricated_metadata_2519 && cargo test -p aprender-train-bench --test falsify_no_fabricated_benchmarks_2519 && cargo test -p aprender-train-shell --test falsify_no_fabricated_fetch_2519 && cargo test -p aprender-core --test beat_sklearn_iris && cargo test -p aprender-core --test beat_sklearn_nmi && cargo test -p aprender-core --test beat_sklearn_metrics_parity && cargo test -p aprender-core --test beat_sklearn_gaussiannb_accuracy && cargo test -p aprender-core --test beat_sklearn_svc_accuracy && cargo test -p aprender-core --test beat_sklearn_pipeline_encoder && cargo test -p aprender-serve --test beat_fail_closed_garbage && cargo test -p aprender-compute --lib beat_nf4_bitsandbytes_equivalence && cargo test -p aprender-core --test beat_pytorch_autograd_grad && cargo test -p aprender-train-lora --lib beat_lora_merge_forward_equivalence && cargo test -p apr-cli --release --test beat_pytorch_deploy_footprint && cargo test -p aprender-serve --test beat_fail_closed_structural && cargo test -p aprender-serve --test ollama_http_compat && cargo test -p apr-cli --test ollama_ndjson_streaming && cargo test -p apr-cli --test falsification_chat_http_cli && cargo test -p aprender-contracts --test apr_serve_api_key_auth_contract && cargo test -p apr-cli --test falsify_auth_001 --test falsify_auth_002 --test falsify_auth_003 --no-fail-fast && cargo test -p apr-cli --test beat_apr_data_alimentar_reach && cargo test -p apr-cli --test beat_apr_sibling_cli_reach && cargo test -p aprender-contracts-cli --test pv_surface_gate && cargo test -p aprender-contracts-cli --test cli_integration && cargo test -p aprender-contracts-cli --test ground_truth && cargo test -p aprender-contracts-cli --test version_identity && cargo build --examples --workspace --keep-going && cargo check --workspace --benches --locked'
# BSE-17 quick tier, part 1: the selected crates' unit and integration
# targets (default features) under the same image and mounts as the full
# tier. Empty selection (a scripts/docs PR) skips this step; part 2 still runs.
Expand Down Expand Up @@ -770,6 +770,12 @@ jobs:
run: bash scripts/bump-version.sh --self-test
- name: Every workspace is at a consistent version
run: bash scripts/bump-version.sh --check
# R-0b (#3002, PP-066 C11): a backend DECISION in apr-cli reads the registry,
# never cfg!(feature = "cuda"|"wgpu"). Case table first, then the live tree.
- name: "Backend-registry static guard can still go RED (case table)"
run: bash scripts/check_backend_registry.sh --self-test
- name: No cfg!(feature) backend decision in apr-cli (R-0b, #3002)
run: bash scripts/check_backend_registry.sh --static
- name: Every DAG row marked complete has a receipt whose marker says so (C0-7)
run: bash scripts/check_receipt_complete.sh --dag docs/specifications/pp-066-dag.yaml
# APR-PERF-GATE-001's PRIMARY gate — Arms A/B1/B2/C/D/E and
Expand Down Expand Up @@ -834,6 +840,17 @@ jobs:
run: bash scripts/check_dag_invariants.sh docs/specifications/pp-066-dag.yaml --min-slack-days 6
- name: A row PR writes no shared file (G-11, PMAT-1062)
run: bash scripts/check_row_pr_write_set.sh --event "${GITHUB_EVENT_NAME:-pull_request}" --branch "${GITHUB_HEAD_REF:-}"
# L0-1a (#2971, PMAT-1065): the supported-model manifest is DERIVED from what the tree
# names (README, BEATS, book, dogfood receipts, perf-matrix) — a named model absent from
# the manifest, or a typed entry nothing cites, is DRIFT; C14's own case table runs over
# the lambda RED-first records so the gate can be seen to turn RED. Case tables, then live.
# Both scripts are cargo-free, so guard-tree (not guard-cargo) is their job (BSE-02).
- name: Supported-model manifest case table (L0-1a, #2971)
run: bash scripts/derive_model_manifest.sh --self-test
- name: The manifest equals its derivation (a named model is never absent, an entry is never typed)
run: bash scripts/derive_model_manifest.sh --check
- name: C14 model-parity case table (L0-1a, #2971)
run: bash scripts/check_model_parity.sh --self-test
# PP-23's bandwidth probe is release-phase INSTRUMENTATION (needs nvcc and an
# idle GPU) and never runs its measurement here; its case table is GPU-free
# and must, or the number the roofline rests on comes from a probe nobody
Expand Down
22 changes: 22 additions & 0 deletions .pr/L0-1/accept.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,22 @@
#!/usr/bin/env bash
# L0-1a accept.sh — re-runs every A_i in one call (I5). GREEN on a GPU-less host by judging the
# recorded runs; the live manifest (`check_model_parity.sh --manifest`) is the fleet-verify leg on
# lambda and gx10 (card acceptance vi), never faked here.
set -uo pipefail; cd "$(dirname "$0")/../.."; rc=0
run() { printf '== %s\n' "$*"; "$@"; local r=$?; printf 'rc=%s\n' "$r"; [ "$r" = 0 ] || rc=1; }
expect_fail() { printf '== (must FAIL) %s\n' "$*"; if "$@"; then printf 'rc=0 (wanted non-zero)\n'; rc=1; else printf 'rc=%s (as required)\n' "$?"; fi; }
CARGO="${CARGO:-$HOME/.cargo/bin/cargo}"; # never a bare `cargo`: a shell function of that name overrides CARGO_TARGET_DIR (and never a machine path either) export CARGO_TARGET_DIR="${CARGO_TARGET_DIR:-/mnt/nvme-raid0/agent-wt/target-l01}"
run bash scripts/derive_model_manifest.sh --self-test
run bash scripts/derive_model_manifest.sh --check
run bash scripts/check_model_parity.sh --self-test
run bash scripts/check_model_parity.sh --judge evidence/parity/l0-1/lambda/qwen2.5-coder-7b-instruct-q4_k_m.json --model qwen2.5-coder-7b-instruct
expect_fail bash scripts/check_model_parity.sh --judge evidence/parity/l0-1/lambda/qwen2.5-coder-1.5b-instruct-q4_k_m.json --model qwen2.5-coder-1.5b-instruct
run bash scripts/check_model_parity.sh --judge evidence/parity/l0-1/gx10/qwen2.5-coder-7b-instruct-q4_k_m.json --model qwen2.5-coder-7b-instruct
expect_fail bash scripts/check_model_parity.sh --judge evidence/parity/l0-1/gx10/qwen2.5-coder-1.5b-instruct-q4_k_m.json --model qwen2.5-coder-1.5b-instruct
expect_fail bash scripts/check_model_parity.sh --judge tests/fixtures/parity/defective/one-position-at-0.5.json --model qwen2.5-coder-7b-instruct
run env "$CARGO" test -p apr-cli --test reg15_admission
run env "$CARGO" test -p apr-cli --lib sentinel_tests
run env "$CARGO" test -p aprender-serve --lib parity_report_carries
. scripts/pv_bin.sh >/dev/null 2>&1 && run "$PV" validate contracts/apr-gpu-cpu-parity-v1.yaml
if [ -d "${APR_MODELS_DIR:-$HOME/models}" ] && [ -n "${APR_BIN_FOR_C14:-}" ]; then run bash scripts/check_model_parity.sh --manifest --apr "$APR_BIN_FOR_C14"; else printf '== check_model_parity.sh --manifest: not on this host (no models dir / no cuda apr) — the fleet-verify leg on lambda and gx10\n'; fi
exit "$rc"
18 changes: 18 additions & 0 deletions .pr/L0-1/diff_benchmark_report.override.patch
Original file line number Diff line number Diff line change
@@ -0,0 +1,18 @@
diff --git a/crates/apr-cli/src/commands/diff_benchmark_report.rs b/crates/apr-cli/src/commands/diff_benchmark_report.rs
index 74a4a3a43..88e5242f6 100644
--- a/crates/apr-cli/src/commands/diff_benchmark_report.rs
+++ b/crates/apr-cli/src/commands/diff_benchmark_report.rs
@@ -79,7 +79,12 @@ pub(crate) fn run(
// PMAT-203: Skip parity gate for profiling — known false positive on CUDA 13.1 driver.
// The parity gate compares GPU/CPU logits and fails spuriously, but profiling
// only needs generation throughput, not parity verification.
- std::env::set_var("SKIP_PARITY_GATE", "1");
+ // REG-15 (#2971): the load-time parity gate is never disabled silently. If the user set
+ // SKIP_PARITY_GATE themselves it is printed as an override (every receipt of this run is
+ // INVALID-CORRECTNESS); otherwise the gate runs and a failing model is refused by name.
+ if let Some(line) = crate::commands::parity_admission::override_line() {
+ eprintln!("{line}");
+ }

// GH-578: Spawn GPU profiling on a thread with 16MB stack.
// The deep call chain (profile_gpu_generation → generate_gpu_resident →
21 changes: 21 additions & 0 deletions .pr/L0-1/plan.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,21 @@
# L0-1 — P0 discovery (2026-09-06, read-only) and P1 plan

## What exists (cited)
- Load-time gate: `crates/aprender-serve/src/gguf/cuda/mod_parity_gate.rs::parity_gate` — ONE token (BOS, position 0), cosine of CPU vs GPU logits against `PARITY_GATE_COSINE_MIN`, a retry on the unfused FFN path when cosine ∈ [0.90, gate). One position: violates I8 (≥ 64 positions) by construction.
- Sequence check: `crates/aprender-serve/src/gguf/parity.rs::check_parity(tokens)` → per-position `ParityResult` (cosine_similarity …) — the engine behind `apr parity` (`crates/apr-cli/src/commands/parity_03.rs`, cuda-only, `--json` emits cosine per position).
- The downgrade: the gate's failure is a `RealizarError` at CUDA model load; the CLI catches it and prints `[CUDA init failed: …, falling back to CPU]` (`chat_generate_session_02.rs:471`; `run_resolve_tokenizer.rs:132` "Fallback to CPU") — a forced `--gpu` silently becomes CPU. This is the "apr then runs cpu under --gpu".
- `SKIP_PARITY_GATE`: read at `gguf/cuda/mod.rs:349`; SET SILENTLY by `commands/comparison.rs:204` and `commands/diff_benchmark_report.rs:82`; `contracts-staging/gpu-multi-backend-parity-v1.yaml:231` already says "SKIP_PARITY_GATE=1 is forbidden in production" — a contract nothing enforces.
- Model names in the tree: README/BEATS/book name qwen2.5-coder-1.5b(-instruct)(-q4_k_m), qwen2.5-coder-0.5b, qwen3.5-9b, qwen3.5-0.8b; perf-matrix.yaml names qwen2.5-coder-1.5b-instruct; the 0.65.2 dogfood receipts name qwen2.5-coder-1.5b-instruct-q4_k_m.gguf. No 7B is named anywhere the manifest would read — the "7B GREEN" leg of the horizon test needs a manifest entry that a doc actually names, or it is not in the manifest.

## P1 phases (A_i as commands; `.pr/L0-1/accept.sh` re-runs them)
| P | deliverable | A_i |
|---|---|---|
| 1 | `scripts/derive_model_manifest.sh` → `evidence/models/supported.yaml` from README, docs/BEATS.md, book/src, evidence/dogfood/*/*.json, scripts/perf-matrix.yaml (names + where cited); `check_readme_claims.sh`: a model named ∉ manifest → RED (case table) | `bash scripts/derive_model_manifest.sh --check` exit 0; `bash scripts/check_readme_claims.sh --self-test` (new rows) |
| 2 | RED first: `scripts/check_model_parity.sh --manifest` (C14) runs `apr parity --json` per manifest model × host over ≥ 64 positions and compares cosine to the threshold file; recorded 1.5B RED / 7B GREEN on lambda and gx10 BEFORE any kernel edit; must-RED twin fixture (`tests/fixtures/parity/defective/`) | `bash scripts/check_model_parity.sh --self-test`; the recorded RED receipts in `evidence/parity/l0-1/` |
| 3 | five whys by `apr parity` SPC down to the op (N-lane root cause: three model families), then the fix | the horizon test flips to GREEN on 1.5B; revert → RED |
| 4 | REG-15 admission in apr-cli: forced backend never downgrades (code from error.rs + reason), unforced prints `selected: cpu (reason: parity FAILED …)`, effective-config `parity: {status, cosine, positions, threshold, basis}`, `SKIP_PARITY_GATE` prints `override:` + receipts INVALID-CORRECTNESS + asserted unset in dogfood and `ci / gate`; the two silent `set_var` sites removed | `cargo test -p apr-cli --test backend_refusal_case_table` (REG-15 rows); `SKIP_PARITY_GATE=1` exported in the gate step → RED |
| 5 | threshold: `evidence/parity/thresholds.yaml` measured n ≥ 5 per known-good pair, `basis=` (0.98 stays [U] until then) | the file's rows carry n ≥ 5 and a command |
| 6 | contract `contracts/apr-gpu-cpu-parity-v1.yaml`; dogfood P6 falsifier; C14 wired in `apr-dogfood --release`, C4 post-publish, R-8 nightly; every (1.5B, cuda) receipt relabelled INVALID-CORRECTNESS citing #3017 | `pv validate`; `check_guards_are_wired.sh` |

Routing: Fable (N-lane root-cause row); lanes: three model families on agy, async. Hosts: lambda + gx10 via `make fleet-verify ROW=L0-1` (G-11b) — until it exists, ad-hoc ssh to lambda is the one excepted host (I10).
K̂ [U].
45 changes: 45 additions & 0 deletions .pr/L0-1/pr-body.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,45 @@
## Ticket
#2971 (alfredodeza's report; P0; DAG row **L0-1a**, ticket PMAT-1065, epic #2873). #3017 was closed as its duplicate.

## Claim
Claim 2 of 0.66, the bounded half: every model in `evidence/models/supported.yaml` (derived, never typed) is measured over ≥ 64 positions against a threshold that carries its basis, and a model that fails is refused by the GPU with its reason printed — never silently run on the CPU under a forced request. The op that diverges on Qwen2.5-1.5B and its fix are L0-1b.

## RED test
- The RED-first record, before any kernel edit: `c977cad58` — `apr parity` on lambda (RTX 4090, the 0.65.2 cuda install `c642576eecb62daa`), 78 positions: **1.5B min cosine 0.9508 at position 0 (token 785), 7B 0.9986**; n=5 repeats are bit-for-bit identical (`d2a566642`).
- The sentinel gate: `d56f65430` — `sentinel_1p5b_on_lambda_is_red_under_the_horizon_rule` / `sentinel_7b_on_lambda_is_green_under_the_horizon_rule` (lib tests, ride `workspace-test`).
- The must-RED twin: `tests/fixtures/parity/defective/one-position-at-0.5.json` (C14 case-table row 3).

## Acceptance (`.pr/L0-1/accept.sh`, orchestrator's run, `.pr/L0-1/accept.log`)
```
derive_model_manifest.sh --self-test rc=0 (6/6)
derive_model_manifest.sh --check rc=0 (18 models)
check_model_parity.sh --self-test rc=0 (6/6)
--judge 7B record rc=0 PASS min 0.9986
--judge 1.5B record rc=1 as required (RED)
--judge one-position-at-0.5 twin rc=1 as required (RED)
cargo test -p apr-cli --test reg15_admission rc=0 (7 passed)
cargo test -p apr-cli --lib sentinel_tests rc=0 (3 passed)
cargo test -p aprender-serve --lib parity_report rc=0
pv validate contracts/apr-gpu-cpu-parity-v1.yaml rc=0 (valid)
--manifest: the fleet-verify leg on lambda/gx10 (not this host)
```
`cargo check -p aprender-serve -p apr-cli --features cuda` on lambda (CUDA 12.8): clean at 35c330c16.

## Mutation (I3)
| round | commit | expected RED | run |
|---|---|---|---|
| 1 | `a5db15200` — a hand-typed manifest entry + `min_cosine: 0.90` | guard-runner-labels: "The manifest equals its derivation" FAILS; workspace-test: `sentinel_1p5b_on_lambda_is_red` FAILS | _run id after CI_ |
| 2 | (next) `min_positions: 1` | guard-runner-labels: C14 case-table row 4 FAILS | _run id after CI_ |
| GREEN | the reverts | 6/6 · 6/6 · 3 sentinels | _run id after CI_ |

## Contract
`contracts/apr-gpu-cpu-parity-v1.yaml` — kind: pattern; PAR-OB-001..003 ↔ PAR-F-001..003. `pv validate` (via `scripts/pv_bin.sh`): `0 error(s), 0 warning(s) — Contract is valid.`

## Quorum
review-only row: one agy lane on this diff — _verdict recorded in `.pr/L0-1/quorum.md` and the receipt before arming_. The L0-1 root-cause quorum (three lanes, one family — a recorded gap) refuted the fused-FFN hypothesis on default config; L0-1b answers the rest by measurement.

## Receipt
`docs/audits/impl-PMAT-1065-receipt.md` (v6 DONE-IF ledger: (i) admission level ✓, (ii) ✓, (iii) blocked on R-0a, (iv) one pair measured, (v) ✓, (vi) pending G-11b, (vii) ✓).

## Writes
`scripts/derive_model_manifest.sh`, `scripts/check_model_parity.sh`, `scripts/dogfood.sh` (C14 rows), `evidence/models/supported.yaml`, `evidence/parity/{thresholds.yaml,l0-1/**}`, `evidence/dogfood/0.65.2/{lambda,gx10}.json` (validity relabel), `tests/fixtures/parity/defective/**`, `crates/apr-cli/src/commands/{parity_admission.rs,chat_generate_session_02.rs,comparison.rs,mod.rs}`, `crates/apr-cli/src/error.rs`, `crates/apr-cli/tests/reg15_admission.rs`, `crates/aprender-serve/src/gguf/cuda/{mod.rs,mod_parity_gate.rs}`, `crates/aprender-serve/src/api/effective_config.rs`, `contracts/apr-gpu-cpu-parity-v1.yaml`, `.github/workflows/ci.yml` (three guard-runner-labels steps), `.pr/L0-1/accept.sh`, the receipt. No DAG, roadmap, README or spec edit.
Loading
Loading