Skip to content

perf(inspect): GGUF header-only read with a capped growing prefix (#3761) - #4576

Closed
noahgift wants to merge 10 commits into
mainfrom
91/3761-on-main
Closed

noahgift wants to merge 10 commits into
mainfrom
91/3761-on-main

Conversation

@noahgift

@noahgift noahgift commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

Closes #3761
Refs #4520, #3750
keep-open: #4520 is the LADDER-HOST-SAFE epic; this PR removes one host-memory hazard (whole-file GGUF reads on header-only verbs) and does not complete the epic.

Header-only verbs: apr inspect on GGUF parses only the header prefix instead of reading the whole file.

  • GgufReader::header_from_file uses a growing prefix: the first read is 16 MiB and it grows to a 256 MiB cap. A header past the cap is refused rather than read whole (quorum round 1 finding). header_from_file_within(path, first, cap) is exposed so tests can pin the cap.
  • Measured on a real 27B GGUF (apr inspect --json): bytes read went from 16.7 GB to 16 MB, peak RSS from 32.8 GB to 75 MB, and wall time from 62 s to 0.24 s. The JSON output is identical.
  • New tests:
    • a_header_longer_than_the_first_prefix_still_parses (700k-token vocab fixture)
    • a_header_past_the_cap_is_refused_not_read_whole. Mutating the cap to usize::MAX/2 makes it fail.
  • Tests: aprender-core format::gguf/prefix/rosetta 661 pass, and apr-cli inspect 141 pass.

Quorum: round 1 Sonnet FAIL on the missing cap, which is now fixed. Round 2 Haiku FAIL was a false positive (it cited metadata as undefined, but the variable is defined and the code compiles). Round 3 AGREED 3/3 (docs/audits/quorum-PMAT-3761.json).

🤖 Generated with Claude Code

noahgift and others added 7 commits September 27, 2026 09:23
… read only their magic or header — one bounded-prefix policy, a peak-RSS case row per reader

ONE policy: apr-format's prefix module (16 MiB first read, doubling, 256 MiB cap, refused
past it, never read whole) with the APR v2 header reader. aprender-core's format::prefix
re-exports it and adds the SafeTensors header reader beside its format. apr-cli's
model_header delegates to it.

The 17 sites: bench (8 bytes to detect the format; the GGUF model mapped ONCE and handed
to the CPU, CUDA and MoE paths), gguf_vocab (the header cut at data_offset), eval
(count_safetensors_keys by header; verify_single_file's hash STREAMED), the converter's
five APR readers, lint (SafeTensors, APR), rosetta inspect, is_onnx_file (5 bytes), and
aprender-serve's is_apr_file / format_from_magic (4 bytes).

The case rows found two copies the grep could not see:
- list_tensors_gguf copied its bytes (data.to_vec()): 4.2 GB on a 2 GiB file. A GGUF
  listing without --stats is now its header.
- apr bench mapped the model twice: realizar's map pre-faults every page, so the two
  maps counted the file twice.

Case rows: a real file extended sparsely to 2 GiB, the readers run in a child process,
VmHWM under 256 MiB (the bench: one map + 256 MiB). 22 mutants each put a whole-file read
back into one reader, and every one goes RED at 2.10-4.22 GB.

Ledger: the 17 rows are CONVERTED (#3761). The same git grep now finds 215 hits.

Closes #3761
Refs #3750

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit cba552c)
… the complexity ratchet (cognitive 28 > 25)

read_apr_metadata_json reads the 24-byte header, then only up to the end of the metadata
section it names, under the shared cap, exactly as before. extract_user_metadata keeps the
source_metadata lookup. Behaviour unchanged; the converter's RSS row and mutant are re-run.

Refs #3761

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit 3060414)
)

`tensors` on a truncated GGUF rendered "Failed to parse GGUF: header parse
failed: Invalid model format: Unexpected EOF…". The header-only path passed
GgufReader's error through parse_growing_prefix by Display, which is what
#3661 removed from the whole-file path. Both paths now share one unwrap.

model_file_error_tests_3661 was RED on 4ca017b and is GREEN here.

Refs #3761, #3661

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… step 2

`apr inspect M --json` is the ladder's one header-only verb (#4520 step 2), and on
Qwen3.5-27B-Q4_K_M it read all 16.7 GB and peaked at 32.8 GB RSS: rosetta
`inspect_gguf` went through `load_gguf_raw`, which reads the whole file and then
copies every tensor, to report sizes and metadata the header holds. The 17 readers
converted by #3761 did not include it, and the 2 GiB RSS row inspected only the APR.

`GgufReader::header_from_file` parses a prefix that doubles until the header parses
(the file length bounds it; a header that never parses reads to EOF and that error is
final). `tensor_extents` sizes each tensor from the header with get_tensor_raw's own
arithmetic and refusals, checked against the file length, so a truncated file is
refused with the same message. Metadata rendering is shared with load_gguf_raw.

Measured on lambda, 27B Q4_K_M, `apr inspect M --json` (strace + time -v):
  rc 0.70.0 (817d633):  model read 16,740,812,720 B, peak RSS 32.8 GB, 62.1 s
  this commit:            model read     16,777,232 B, peak RSS 75 MB,   0.24 s
  JSON output byte-identical.

The 2 GiB peak-RSS row now inspects the GGUF too. Mutation: the whole-file read put
back in inspect_gguf makes that row read 2,124,048 KiB (bound 262,144) — RED.

Refs #4520
Closes #3761

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… a corrupt GGUF header is refused at 256 MiB, never read to EOF

Quorum finding (PMAT-3761 round 1, lane 2, sonnet-5): header_from_file
doubled its prefix up to file_len, so a GGUF whose header never parses
buffered the whole file (17 GB on the 27B) before its error — the blowup
this ticket fixes, triggered by a malformed file. It now goes through
parse_growing_prefix_within (16 MiB first, HEADER_READ_CAP 256 MiB), like
apr_v2_header_prefix and safetensors_header_prefix.

Case row a_header_past_the_cap_is_refused_not_read_whole: 700k-token
header, first 64 KiB, cap 1 MiB -> refused by name. Mutant (cap -> usize::MAX/2):
that row FAILS, 4 others pass. aprender-core gguf/prefix/rosetta 661 pass,
apr-cli inspect 141 pass, clippy -D warnings clean.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@noahgift
noahgift enabled auto-merge September 27, 2026 13:19
@github-actions

github-actions Bot commented Sep 27, 2026 •

Copy link
Copy Markdown

§13.11 rung 1 — quorum shadow verdict

S13-SHADOW pr=4576 head=8d12866154714cc7cebb0dfd415cd591329b85ef verdict=REFUSE class=Q1 arm_rc=1

Shadow mode: this records a verdict and merges nothing. A refusal
to arm is not a block (§13 adds zero rows to §7) — the pull request is
exactly as green as it was.

@noahgift noahgift added the owner:aprender-91 owning session (cop inbox claims) label Sep 27, 2026
noahgift and others added 3 commits September 27, 2026 17:32
…prefix (#3661)

header_from_file_within passed GgufReader::from_bytes straight to
parse_growing_prefix_within, so the FormatError's Display ('Invalid model
format: …') was embedded in the prefix message and re-wrapped.
on a truncated GGUF then read '…header parse failed: Invalid model format:
Unexpected EOF…' and model_file_error_tests_3661 went red on the folded tree.
The closure now contributes the message, as safetensors.rs's listing path
already does (error_message).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…merge (guard-tree DRIFT)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@noahgift

Copy link
Copy Markdown
Contributor Author

Moved into #4607 (PRCAP fold, cop order 14:02Z). Fold = MOVE: head 8d12866154714cc7cebb0dfd415cd591329b85ef is an ancestor of the pushed fold head, so no work is lost. The branch is kept. Agent: aprender-59

@noahgift noahgift closed this Sep 28, 2026
auto-merge was automatically disabled September 28, 2026 14:24

Pull request was closed

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

owner:aprender-91 owning session (cop inbox claims)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

#3750 PR B: the 17 header-only whole-file model reads outside apr qa — magic/header/metadata readers in their format crates

1 participant