fold(gpu-perf): #4576 + #4586 + #4532 — GGUF header-only read, dense F2 guard, sm_121 logits budget - #4607
Open
noahgift wants to merge 26 commits into
Open
fold(gpu-perf): #4576 + #4586 + #4532 — GGUF header-only read, dense F2 guard, sm_121 logits budget#4607noahgift wants to merge 26 commits into
noahgift wants to merge 26 commits into
Conversation
… read only their magic or header — one bounded-prefix policy, a peak-RSS case row per reader ONE policy: apr-format's prefix module (16 MiB first read, doubling, 256 MiB cap, refused past it, never read whole) with the APR v2 header reader. aprender-core's format::prefix re-exports it and adds the SafeTensors header reader beside its format. apr-cli's model_header delegates to it. The 17 sites: bench (8 bytes to detect the format; the GGUF model mapped ONCE and handed to the CPU, CUDA and MoE paths), gguf_vocab (the header cut at data_offset), eval (count_safetensors_keys by header; verify_single_file's hash STREAMED), the converter's five APR readers, lint (SafeTensors, APR), rosetta inspect, is_onnx_file (5 bytes), and aprender-serve's is_apr_file / format_from_magic (4 bytes). The case rows found two copies the grep could not see: - list_tensors_gguf copied its bytes (data.to_vec()): 4.2 GB on a 2 GiB file. A GGUF listing without --stats is now its header. - apr bench mapped the model twice: realizar's map pre-faults every page, so the two maps counted the file twice. Case rows: a real file extended sparsely to 2 GiB, the readers run in a child process, VmHWM under 256 MiB (the bench: one map + 256 MiB). 22 mutants each put a whole-file read back into one reader, and every one goes RED at 2.10-4.22 GB. Ledger: the 17 rows are CONVERTED (#3761). The same git grep now finds 215 hits. Closes #3761 Refs #3750 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> (cherry picked from commit cba552c)
… the complexity ratchet (cognitive 28 > 25) read_apr_metadata_json reads the 24-byte header, then only up to the end of the metadata section it names, under the shared cap, exactly as before. extract_user_metadata keeps the source_metadata lookup. Behaviour unchanged; the converter's RSS row and mutant are re-run. Refs #3761 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> (cherry picked from commit 3060414)
) `tensors` on a truncated GGUF rendered "Failed to parse GGUF: header parse failed: Invalid model format: Unexpected EOF…". The header-only path passed GgufReader's error through parse_growing_prefix by Display, which is what #3661 removed from the whole-file path. Both paths now share one unwrap. model_file_error_tests_3661 was RED on 4ca017b and is GREEN here. Refs #3761, #3661 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…vs sm_121 GPU) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…the first over-budget one Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… sm_121 GPU GPU sm_89 vs sm_121 pos0 logits agree to 6.5e-7; x86 AVX2 vs aarch64 NEON CPU references differ by 2.6e-2. aarch64 budget = measured worst 7.105e-2 rounded up (8e-2); x86_64 stays 7e-2. QWEN35_E2E_DUMP now fails after printing, never passes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… step 2 `apr inspect M --json` is the ladder's one header-only verb (#4520 step 2), and on Qwen3.5-27B-Q4_K_M it read all 16.7 GB and peaked at 32.8 GB RSS: rosetta `inspect_gguf` went through `load_gguf_raw`, which reads the whole file and then copies every tensor, to report sizes and metadata the header holds. The 17 readers converted by #3761 did not include it, and the 2 GiB RSS row inspected only the APR. `GgufReader::header_from_file` parses a prefix that doubles until the header parses (the file length bounds it; a header that never parses reads to EOF and that error is final). `tensor_extents` sizes each tensor from the header with get_tensor_raw's own arithmetic and refusals, checked against the file length, so a truncated file is refused with the same message. Metadata rendering is shared with load_gguf_raw. Measured on lambda, 27B Q4_K_M, `apr inspect M --json` (strace + time -v): rc 0.70.0 (817d633): model read 16,740,812,720 B, peak RSS 32.8 GB, 62.1 s this commit: model read 16,777,232 B, peak RSS 75 MB, 0.24 s JSON output byte-identical. The 2 GiB peak-RSS row now inspects the GGUF too. Mutation: the whole-file read put back in inspect_gguf makes that row read 2,124,048 KiB (bound 262,144) — RED. Refs #4520 Closes #3761 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…y gx10) + findings yoga-eph has no models, so the module SKIPped at PR time and the whole-forward budgets were only ever read on sm_89/x86_64. The gx10 nightly holds the model; the step refuses a skip or an empty selection. Verified: 17/17 on GB10. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… a corrupt GGUF header is refused at 256 MiB, never read to EOF Quorum finding (PMAT-3761 round 1, lane 2, sonnet-5): header_from_file doubled its prefix up to file_len, so a GGUF whose header never parses buffered the whole file (17 GB on the 27B) before its error — the blowup this ticket fixes, triggered by a malformed file. It now goes through parse_growing_prefix_within (16 MiB first, HEADER_READ_CAP 256 MiB), like apr_v2_header_prefix and safetensors_header_prefix. Case row a_header_past_the_cap_is_refused_not_read_whole: 700k-token header, first 64 KiB, cap 1 MiB -> refused by name. Mutant (cap -> usize::MAX/2): that row FAILS, 4 others pass. aprender-core gguf/prefix/rosetta 661 pass, apr-cli inspect 141 pass, clippy -D warnings clean. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…prefix (#3661) header_from_file_within passed GgufReader::from_bytes straight to parse_growing_prefix_within, so the FormatError's Display ('Invalid model format: …') was embedded in the prefix message and re-wrapped. on a truncated GGUF then read '…header parse failed: Invalid model format: Unexpected EOF…' and model_file_error_tests_3661 went red on the folded tree. The closure now contributes the message, as safetensors.rs's listing path already does (error_message). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
….5s on qwen2.5-coder-0.5b The cosine-0.4153 rejection no longer reproduces on car/0.70.0 (#3727 made PMAT-084 activation reuse opt-in). What remained of the 2-3x was validate_gpu_first_token itself: a CPU reference forward of the whole prompt on every dense --gpu run. Port the #3604 receipt (model sha256, apr build, device) to the dense path, keyed additionally on the prefill precision that passed, so a receipt never re-enables an FP8 prefill the guard rejected (#3807). Pmat-Ticket: PMAT-3602 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…()), per quorum lane 1 Pmat-Ticket: PMAT-3602 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Pmat-Ticket: PMAT-3602 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…merge (guard-tree DRIFT) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-59 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-59 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-59 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Sep 28, 2026
|
§13.11 rung 1 — quorum shadow verdict Shadow mode: this records a verdict and merges nothing. A refusal |
Contributor
Author
|
quorum-review (AD-04): three PASS — agreed (auto_merge: checked=true was_armed=false disarmed=false) {
"ticket": "PMAT-3761,PMAT-3602,PMAT-4523",
"head": "25fdbc4a12d0c2b62b517ad628aa089b95723c11",
"width": 3,
"executor": "agy",
"agreed": true,
"auto_merge": {
"checked": true,
"was_armed": false,
"disarmed": false,
"note": "auto-merge not armed"
},
"lanes": [
{
"lane": 1,
"verdict": "PASS",
"findings": 3
},
{
"lane": 2,
"verdict": "PASS",
"findings": 0
},
{
"lane": 3,
"verdict": "PASS",
"findings": 0
}
]
} |
…t-5, gemini-3.1-pro-high, haiku-4-5; mixed, no author seat) Agent: aprender-59 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
PRCAP fold (cop order 14:02Z, 24/10 CRIT). Fold = MOVE: each source PR's head commit is an ancestor of this branch (merged with
--no-ff, so the original commits are kept). Built on origin/main15f1b2a496. Fold head:25fdbc4a12d0c2b62b517ad628aa089b95723c11.8d128661542a56299f80ef44857b9fClean merge, fmt clean. Touches .github/workflows/cuda-nightly.yml (via #4532), so it needs a 3/3 non-Claude quorum.
Carried quorum/receipt files move with their commits. A fold changes the head, so this PR needs a new quorum receipt signed with
--ticketlisting every source ticket. Not armed (workers don't arm).Workflow gate proof (cop 20:16Z: pre-approval conditions)
git diff origin/main...HEAD -- .github/touches.github/workflows/cuda-nightly.ymlonly: +25 −0. It adds one step, qwen35 CUDA parity module on sm_121 (aprender#4523), to the gx10 job. No existing step, check orif:is changed or removed.permissions:lines at all. No token scope changes.SKIP/is absentline, on a missing model file, and on zero tests selected.gh workflow run cuda-nightly.yml --ref fold/gpu-perf -f silicon=blackwell→ run 36477877442 → NOT MEASURED. The run concluded success, but the decide step yielded (GPU busy (0 MiB, 1 procs) — yielding to training), so the new step was skipped. That is not a GREEN. It needs a re-dispatch while gx10 is idle (orforce=trueif the cop authorizes overriding the training yield).run:block was extracted verbatim from the YAML and run with a stubcargoover fixture logs. That gives 6/6 rows as expected: green-3-pass rc=0; red-model-absent, red-cargo-fails, red-skip-line, red-zero-selected and red-no-summary each rc=1. A RED CI run was not dispatched.perf-gx10keeps a single pending run, so a second dispatch could cancel main's nightly.Agent: aprender-59
🤖 Generated with Claude Code