Skip to content

fold(gpu-perf): #4576 + #4586 + #4532 — GGUF header-only read, dense F2 guard, sm_121 logits budget - #4607

Open
noahgift wants to merge 26 commits into
mainfrom
fold/gpu-perf
Open

noahgift wants to merge 26 commits into
mainfrom
fold/gpu-perf

Conversation

@noahgift

@noahgift noahgift commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

PRCAP fold (cop order 14:02Z, 24/10 CRIT). Fold = MOVE: each source PR's head commit is an ancestor of this branch (merged with --no-ff, so the original commits are kept). Built on origin/main 15f1b2a496. Fold head: 25fdbc4a12d0c2b62b517ad628aa089b95723c11.

Source PR Head Ancestor of fold head
#4576 perf(inspect): GGUF header-only read with a capped growing prefix (#3761) 8d12866154 yes
#4586 fix(#3602): dense F2 guard behind its receipt — --gpu setup 24.7s → 1.5s 2a56299f80 yes
#4532 fix(qwen35): per-arch e2e logits budget from measurement + sm_121 nightly lane (Closes #4523) ef44857b9f yes

Clean merge, fmt clean. Touches .github/workflows/cuda-nightly.yml (via #4532), so it needs a 3/3 non-Claude quorum.

Carried quorum/receipt files move with their commits. A fold changes the head, so this PR needs a new quorum receipt signed with --ticket listing every source ticket. Not armed (workers don't arm).

Workflow gate proof (cop 20:16Z: pre-approval conditions)

  • Diff: git diff origin/main...HEAD -- .github/ touches .github/workflows/cuda-nightly.yml only: +25 −0. It adds one step, qwen35 CUDA parity module on sm_121 (aprender#4523), to the gx10 job. No existing step, check or if: is changed or removed.
  • Permissions: the diff has no permissions: lines at all. No token scope changes.
  • Gates: tightened only. The new step fails on a non-zero cargo rc, on any SKIP/is absent line, on a missing model file, and on zero tests selected.
  • GREEN (CI): gh workflow run cuda-nightly.yml --ref fold/gpu-perf -f silicon=blackwell → run 36477877442 → NOT MEASURED. The run concluded success, but the decide step yielded (GPU busy (0 MiB, 1 procs) — yielding to training), so the new step was skipped. That is not a GREEN. It needs a re-dispatch while gx10 is idle (or force=true if the cop authorizes overriding the training yield).
  • RED (offline, exact step script): the step run: block was extracted verbatim from the YAML and run with a stub cargo over fixture logs. That gives 6/6 rows as expected: green-3-pass rc=0; red-model-absent, red-cargo-fails, red-skip-line, red-zero-selected and red-no-summary each rc=1. A RED CI run was not dispatched. perf-gx10 keeps a single pending run, so a second dispatch could cancel main's nightly.

Agent: aprender-59

🤖 Generated with Claude Code

noahgift and others added 25 commits September 27, 2026 09:23
… read only their magic or header — one bounded-prefix policy, a peak-RSS case row per reader

ONE policy: apr-format's prefix module (16 MiB first read, doubling, 256 MiB cap, refused
past it, never read whole) with the APR v2 header reader. aprender-core's format::prefix
re-exports it and adds the SafeTensors header reader beside its format. apr-cli's
model_header delegates to it.

The 17 sites: bench (8 bytes to detect the format; the GGUF model mapped ONCE and handed
to the CPU, CUDA and MoE paths), gguf_vocab (the header cut at data_offset), eval
(count_safetensors_keys by header; verify_single_file's hash STREAMED), the converter's
five APR readers, lint (SafeTensors, APR), rosetta inspect, is_onnx_file (5 bytes), and
aprender-serve's is_apr_file / format_from_magic (4 bytes).

The case rows found two copies the grep could not see:
- list_tensors_gguf copied its bytes (data.to_vec()): 4.2 GB on a 2 GiB file. A GGUF
  listing without --stats is now its header.
- apr bench mapped the model twice: realizar's map pre-faults every page, so the two
  maps counted the file twice.

Case rows: a real file extended sparsely to 2 GiB, the readers run in a child process,
VmHWM under 256 MiB (the bench: one map + 256 MiB). 22 mutants each put a whole-file read
back into one reader, and every one goes RED at 2.10-4.22 GB.

Ledger: the 17 rows are CONVERTED (#3761). The same git grep now finds 215 hits.

Closes #3761
Refs #3750

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit cba552c)
… the complexity ratchet (cognitive 28 > 25)

read_apr_metadata_json reads the 24-byte header, then only up to the end of the metadata
section it names, under the shared cap, exactly as before. extract_user_metadata keeps the
source_metadata lookup. Behaviour unchanged; the converter's RSS row and mutant are re-run.

Refs #3761

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit 3060414)
)

`tensors` on a truncated GGUF rendered "Failed to parse GGUF: header parse
failed: Invalid model format: Unexpected EOF…". The header-only path passed
GgufReader's error through parse_growing_prefix by Display, which is what
#3661 removed from the whole-file path. Both paths now share one unwrap.

model_file_error_tests_3661 was RED on 4ca017b and is GREEN here.

Refs #3761, #3661

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…vs sm_121 GPU)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…the first over-budget one

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… sm_121 GPU

GPU sm_89 vs sm_121 pos0 logits agree to 6.5e-7; x86 AVX2 vs aarch64 NEON CPU
references differ by 2.6e-2. aarch64 budget = measured worst 7.105e-2 rounded up
(8e-2); x86_64 stays 7e-2. QWEN35_E2E_DUMP now fails after printing, never passes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… step 2

`apr inspect M --json` is the ladder's one header-only verb (#4520 step 2), and on
Qwen3.5-27B-Q4_K_M it read all 16.7 GB and peaked at 32.8 GB RSS: rosetta
`inspect_gguf` went through `load_gguf_raw`, which reads the whole file and then
copies every tensor, to report sizes and metadata the header holds. The 17 readers
converted by #3761 did not include it, and the 2 GiB RSS row inspected only the APR.

`GgufReader::header_from_file` parses a prefix that doubles until the header parses
(the file length bounds it; a header that never parses reads to EOF and that error is
final). `tensor_extents` sizes each tensor from the header with get_tensor_raw's own
arithmetic and refusals, checked against the file length, so a truncated file is
refused with the same message. Metadata rendering is shared with load_gguf_raw.

Measured on lambda, 27B Q4_K_M, `apr inspect M --json` (strace + time -v):
  rc 0.70.0 (817d633):  model read 16,740,812,720 B, peak RSS 32.8 GB, 62.1 s
  this commit:            model read     16,777,232 B, peak RSS 75 MB,   0.24 s
  JSON output byte-identical.

The 2 GiB peak-RSS row now inspects the GGUF too. Mutation: the whole-file read put
back in inspect_gguf makes that row read 2,124,048 KiB (bound 262,144) — RED.

Refs #4520
Closes #3761

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…y gx10) + findings

yoga-eph has no models, so the module SKIPped at PR time and the whole-forward
budgets were only ever read on sm_89/x86_64. The gx10 nightly holds the model;
the step refuses a skip or an empty selection. Verified: 17/17 on GB10.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… a corrupt GGUF header is refused at 256 MiB, never read to EOF

Quorum finding (PMAT-3761 round 1, lane 2, sonnet-5): header_from_file
doubled its prefix up to file_len, so a GGUF whose header never parses
buffered the whole file (17 GB on the 27B) before its error — the blowup
this ticket fixes, triggered by a malformed file. It now goes through
parse_growing_prefix_within (16 MiB first, HEADER_READ_CAP 256 MiB), like
apr_v2_header_prefix and safetensors_header_prefix.

Case row a_header_past_the_cap_is_refused_not_read_whole: 700k-token
header, first 64 KiB, cap 1 MiB -> refused by name. Mutant (cap -> usize::MAX/2):
that row FAILS, 4 others pass. aprender-core gguf/prefix/rosetta 661 pass,
apr-cli inspect 141 pass, clippy -D warnings clean.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…prefix (#3661)

header_from_file_within passed GgufReader::from_bytes straight to
parse_growing_prefix_within, so the FormatError's Display ('Invalid model
format: …') was embedded in the prefix message and re-wrapped.
on a truncated GGUF then read '…header parse failed: Invalid model format:
Unexpected EOF…' and model_file_error_tests_3661 went red on the folded tree.
The closure now contributes the message, as safetensors.rs's listing path
already does (error_message).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
….5s on qwen2.5-coder-0.5b

The cosine-0.4153 rejection no longer reproduces on car/0.70.0 (#3727 made PMAT-084 activation reuse opt-in). What remained of the 2-3x was validate_gpu_first_token itself: a CPU reference forward of the whole prompt on every dense --gpu run. Port the #3604 receipt (model sha256, apr build, device) to the dense path, keyed additionally on the prefill precision that passed, so a receipt never re-enables an FP8 prefill the guard rejected (#3807).

Pmat-Ticket: PMAT-3602

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…()), per quorum lane 1

Pmat-Ticket: PMAT-3602

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Pmat-Ticket: PMAT-3602

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…merge (guard-tree DRIFT)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-59
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-59
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-59
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

§13.11 rung 1 — quorum shadow verdict

S13-SHADOW pr=4607 head=b3faa936083aad4270e9e7f3916bc3ac986edce3 verdict=REFUSE class=Q1 arm_rc=1

Shadow mode: this records a verdict and merges nothing. A refusal
to arm is not a block (§13 adds zero rows to §7) — the pull request is
exactly as green as it was.

@noahgift

Copy link
Copy Markdown
Contributor Author

quorum-review (AD-04): three PASS — agreed (auto_merge: checked=true was_armed=false disarmed=false)

{
 "ticket": "PMAT-3761,PMAT-3602,PMAT-4523",
 "head": "25fdbc4a12d0c2b62b517ad628aa089b95723c11",
 "width": 3,
 "executor": "agy",
 "agreed": true,
 "auto_merge": {
  "checked": true,
  "was_armed": false,
  "disarmed": false,
  "note": "auto-merge not armed"
 },
 "lanes": [
  {
   "lane": 1,
   "verdict": "PASS",
   "findings": 3
  },
  {
   "lane": 2,
   "verdict": "PASS",
   "findings": 0
  },
  {
   "lane": 3,
   "verdict": "PASS",
   "findings": 0
  }
 ]
}

…t-5, gemini-3.1-pro-high, haiku-4-5; mixed, no author seat)

Agent: aprender-59
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

qwen35 CUDA-vs-CPU e2e logits exceed LOGITS_BUDGET 7e-2 on GB10 sm_121 (7.105e-2, 2/2 deterministic; budget measured on 4090 only)

1 participant