Skip to content

feat(qwen36): derive the default expert-cache cap from host RAM instead of a hardcoded 16 - #1747

Open
jtinbergen wants to merge 4 commits into
JustVugg:devfrom
jtinbergen:dense-cap-default-v1
Open

jtinbergen wants to merge 4 commits into
JustVugg:devfrom
jtinbergen:dense-cap-default-v1

Conversation

@jtinbergen

@jtinbergen jtinbergen commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Title: feat(qwen36): derive the default expert-cache cap from host RAM instead of a hardcoded 16

Summary

  • cap (expert-cache slots/layer, argv[1]) no longer defaults to a hardcoded
    16 when omitted. Bare omission and an explicit cap=0 are now the same
    sentinel ("auto-size from host RAM"), matching the cap<=0 convention
    colibri.c/olmoe.c already use -- this PR only extends that convention to
    bare CLI omission, which neither of those engines does either today.
  • Adds qwen36_cap_for_ram(): a pure function (no Model*, no globals, no
    I/O -- same testability contract as coli_resolve_cap/k3_cap_for_ram)
    that derives the cap from resident + mem_available_gb()*0.88 (or an
    explicit RAM_GB override), the int4-vs-int8 per-slot byte cost
    (xf_mode-aware, half the bytes when experts are stored packed int4), and
    the number of active layers. Mirrors olmoe.c's own cap<=0 budget block
    and reuses compat_mem_available_gb() -- the same cross-platform free-RAM
    probe every other engine (colibri.c, olmoe.c, inkling.c, kimi_k3.c,
    deepseek_v4.c) already wraps in a local mem_available_gb(); qwen36.c
    had no such wrapper before this PR. RAM_GB as the override name and
    0.88 as the margin are the existing repo-wide convention, not invented
    here. The resolved budget and per-slot size are returned through two
    outparams rather than recomputed by the caller, so the diagnostic line
    below can never drift from what the function actually decided.
  • Called from model_init_range, inserted right after all dense weights for
    the range finish loading (so rss_gb() reflects a real, all-dense-resident
    footprint) and right before the per-layer expert-cache allocation that
    consumes cap for the first time -- no reordering of main()'s argument
    parsing needed. Prints [qwen36] cache auto-sized: N slots/layer of M experts (...) so the derivation is visible and testable, the same way
    olmoe.c's own cache-line is.
  • qwen36_segment_engine_open's memory_limit_bytes == 0 path also now
    passes the same sentinel through, instead of leaving cap hardcoded at
    16 -- ColiSegmentEngineOptions.memory_limit_bytes's own doc comment
    already promised 0 means "the adapter's ordinary automatic budget"; this
    closes that gap with the same one function, one line changed. The
    memory_limit_bytes != 0 path (an explicit caller-supplied byte ceiling,
    a different concept from measuring host RAM) is unchanged.
  • An explicit cap>0, from either call site, is never touched by any of
    this -- the new code path is only reached when cap<=0.
  • Second commit: dense checkpoints (num_experts == 0, [Feature]: Qwen 3.8 27b #1757's Qwen3.8-27B
    and the rest of the family) short-circuit to cap=1 before the derivation
    runs. Nothing is routed there, so the per-layer cache is never touched by
    moe()/expert_get(), and qwen36_cap_for_ram's n_experts clamp has
    nothing to clamp against -- an unclamped RAM budget could otherwise size a
    cache that will sit empty. Found while rebasing onto current dev: this
    PR was written when validate_cfg still guaranteed n_experts > 0.

Known, pre-existing, out-of-scope bug found while reading this code, not
fixed here
: xf_mode() memoizes its int4-vs-int8 answer in a process-global
static int, not per-Model*. If a single process opens more than one
qwen36_segment_engine_open() model with different on-disk container
formats, the second model's decode kernel (not just this PR's cap sizing)
would silently reuse the first model's answer. This predates this PR and
lives on the hot decode path (matmul_d's dispatch, slot_ensure_allocated);
fixing it is a separate, higher-risk change and is intentionally not part of
this diff. Flagging it here rather than staying quiet about something found
along the way.

Scope

Default-selection only. No numerical behavior changes anywhere: an explicit
cap produces byte-identical cache allocation/behavior to before this PR
(verified directly, see Measured). The auto-derived value itself is a
resource-sizing heuristic, not something with a single correct answer to
verify bit-exactly -- what's verified instead is that it's bounded correctly
(never <1, never >n_experts), responds to its documented inputs
(RAM_GB, int4 vs int8, layer count), and that a model loads and runs
correctly at whatever cap it derives.

Measured / Demonstrated

Not a throughput change, so CONTRIBUTING.md's experiment-manifest
requirement (built around samples.tok_s) doesn't apply here -- same
reasoning as #1716's own "Measured" section: this changes what number gets
chosen before the cache exists, not decode speed at a given cap. Demonstrated
instead with real command transcripts on the same real model and machine as
#1716: qwen36-i4-gs64 (35B), Intel Core Ultra 7 155H (16C/22T, hybrid P/E),
61 GiB RAM, local ext4 SSD, commit 9fe4666e (rebased onto origin/dev
eefa57a3). Single runs, not medians over several -- the
derived cap is a deterministic function of its inputs (resident RSS,
mem_available_gb(), geometry), not a timing measurement, so repeating a
run would reproduce the same number rather than add information; the
benchmark-reporting rigor (median + spread over repeated runs) that
CONTRIBUTING.md asks for applies to noisy throughput claims, which this
PR doesn't make.

Bare omission now auto-sizes instead of defaulting to 16:

cd c
SNAP=~/models/qwen36-i4-gs64 SERVE=1 ./qwen36
== qwen36 Phase-2 engine | cache=auto/layer bits=4 ctx=8192 ... ==
[dense-i8] 291 matrices quantized during load, 7.2 GB f32 freed
[qwen36] cache auto-sized: 256 slots/layer of 256 experts (24.4 GB budget via 88% of available RAM, 4.8 GB dense resident, 2 MB/slot)
resident weights loaded in 15.1s | RSS after load: 4.83 GB

Before vs. after, same model, varying available RAM

RAM_GB stands in for "what's actually free" (real mem_available_gb()
reads whatever the host has at that moment; RAM_GB makes the same code
path reproducible for this table). Same qwen36-i4-gs64, same machine:

available RAM old default (hardcoded) new default (auto)
6 GB 16 slots/layer 18 slots/layer
8 GB 16 slots/layer 50 slots/layer
16 GB 16 slots/layer 176 slots/layer
32 GB 16 slots/layer 256 slots/layer (= every expert)
61 GB (this machine, unset) 16 slots/layer 256 slots/layer (= every expert)
for ram in 6 8 16 32 61; do RAM_GB=$ram SNAP=~/models/qwen36-i4-gs64 SERVE=1 ./qwen36; done

The risk on a low-memory system that motivates this: the old default of
16 knows nothing about the machine it's running on. On this particular
model's geometry (~2 MB/slot) that happens not to be dramatic at 6 GB (16
vs. 18) -- but the old number carries no relationship to available RAM at
all, on any model. For a model with a larger hidden/inter (bigger
per-slot cost), the same hardcoded 16 could just as easily land above
what a tight machine actually has free, with nothing in the old code path
even aware of that -- the cache would still be allowed to grow toward 16
slots/layer regardless of what else is resident, and only a user who already
knew to pass a smaller explicit cap by hand was protected. The new default
is bounded by the same mem_available_gb() measurement on every model, so
it can't hand out a cache ceiling bigger than what's actually free in the
first place -- on a genuinely tight machine it now picks a smaller cap
than 16 automatically instead of a user finding out the hard way. The
trade-off is the same one already inherent to a smaller cap on any engine:
fewer resident slots means more disk streaming/cache misses during
generation, not incorrect output -- this PR does not change that trade-off,
it only chooses a cap that respects the machine's actual headroom instead of
a number that ignores it.

RAM_GB bounds the choice on the same model (tight budget vs. generous):

RAM_GB=6  SNAP=~/models/qwen36-i4-gs64 SERVE=1 ./qwen36   # -> 18 slots/layer
RAM_GB=64 SNAP=~/models/qwen36-i4-gs64 SERVE=1 ./qwen36   # -> 256 slots/layer (clamped at n_experts)

An explicit cap is never re-derived, regression-checked on the same model:

SNAP=~/models/qwen36-i4-gs64 SERVE=1 ./qwen36 32 4
== qwen36 Phase-2 engine | cache=32/layer bits=4 ctx=8192 ... ==   (no "cache auto-sized" line printed)

Token-exact oracle, both as a pure regression check and on a cap the new
auto-derivation itself picked (make_qwen36_tiny.py --ref-mode full +
convert_qwen36.py --ebits 8, COLI_DENSE_I8=0, the same tiny fixture and
flags the existing qwen36-tiny-check CI job uses):

for cap in 1 2 8; do
  COLI_DENSE_I8=0 SNAP=qwen36_tiny_c ./qwen36 "$cap" 8 qwen36_tiny/ref_full.json
done   # 16/16 at each of cap=1, cap=2, cap=8 -- unchanged from before this PR
RAM_GB=8 COLI_DENSE_I8=0 SNAP=qwen36_tiny_c ./qwen36 0 8 qwen36_tiny/ref_full.json
# [qwen36] cache auto-sized: 8 slots/layer of 8 experts ... -> Matching tokens: 16/16

Verification

  • test_qwen36_cap_precedence (new, unit, pure function): RAM_GB override
    wins over the computed budget, and its outparam echoes the override
    exactly; a negative or zero override (garbage/unset RAM_GB) falls back
    to the computed budget rather than propagating a negative one; int4
    derives a cap >= int8 for identical budget/geometry; a budget at or
    below the resident floor still returns cap==1, never 0; a generous
    budget clamps at n_experts; n_active_layers scales the result
    (Segment/Edge partial-model ranges pass layer_end-layer_begin, not
    always the full model's n_layers); degenerate inputs
    (n_active_layers<=0) don't crash.
  • test_qwen36_cap_budget.py (new, integration, black-box, mirrors
    test_olmoe_cap_budget.py's approach for the sibling engine): explicit cap
    is never second-guessed; a generous RAM_GB holds every expert; a budget
    below the floor still runs at cap==1; the budget bounds the choice; no
    RAM_GB still decides (measures the real host); bare CLI omission produces
    the identical cache-line as an explicit cap=0 bits=4 -- the equivalence
    this PR's whole premise rests on, tested directly rather than only argued
    from reading main().
  • test_segment_adapters_real qwen36 <fixture> 0 4 8 8 (existing real-model
    Segment/Edge gate, not part of routine make check -- it needs
    QWEN_SEGMENT_MODEL set to a real converted container, same as every
    other engine's entry in the segment-adapters-real Makefile target): run
    directly against the tiny fixture to exercise
    qwen36_segment_engine_open's memory_limit_bytes==0 path end to end,
    including a genuinely partial layer range ([0,4)/[4,8) alongside the
    full [0,8) engine in the same process) -- this is the one call site the
    unit test above can't reach on its own. qwen36 real Segment range/chaining/snapshot: ok.
  • Dense checkpoint (num_experts == 0): a qwen38-27b-dense tiny fixture
    (make_qwen36_tiny.py --geometry qwen38-27b-dense, the geometry [Feature]: Qwen 3.8 27b #1757
    added) is token-exact 16/16 both at an explicit cap and at the cap<=0
    sentinel, and the sentinel prints no auto-sizing line -- confirming the
    short-circuit runs and no oversized cache is allocated.
  • make check: pass, 1923 tests, 0 new warnings (after make clean, on the
    rebased tree).
  • make test-asan (ASan + UBSan): clean.
  • qwen36 tiny oracle, token-exact, both the existing cap=1/2/8 regression and
    the new auto-derivation path (see Measured above): 16/16 in all four runs.

@JustVugg

JustVugg commented Oct 1, 2026

Copy link
Copy Markdown
Owner

Thanks. With an explicit cap this is fine on the CPU path: the qwen36 oracles are byte-identical to dev, test_qwen36_cap_precedence passes and test_qwen36_cap_budget passes 6/6.

One bug blocks it, on the path this PR is about. When --cap is omitted, main() keeps its own cap at the "size it from RAM" sentinel; model_init_range resolves only its local copy. That unresolved value is what reaches qt_init (c/qwen36.c:4556 at 9fe4666), and qt_init disables the tier whenever cap != n_experts outside fp8-stream mode (c/qwen36_tier.c:536). So COLI_CUDA=1 with an automatic cap switches the VRAM tier off even when auto-sizing resolved to every expert (8 of 8 on the tiny fixture). Passing the resolved cap back to main() before qt_init, for example m.cache[layer_begin].cap, fixes it. This is from reading the code; we have no GPU to run it on.

Also a heads-up: the new Makefile rule lists headers by hand and builds several units in one command. Once #1758 lands, its guard rejects rules like that, so it is easier to write it the way #1758 does from the start.

jtinbergen and others added 4 commits October 1, 2026 13:54
…ad of a hardcoded 16

cap<=0 (explicit 0, or bare CLI omission) now derives the expert-cache
slots/layer from mem_available_gb()*0.88 (or an explicit RAM_GB override),
the int4-vs-int8 per-slot byte cost, and the active layer count, instead of
always returning 16 regardless of the machine. Mirrors olmoe.c's own cap<=0
budget block and colibri.c's 0.88 margin; qwen36.c had no mem_available_gb()
wrapper before this. qwen36_segment_engine_open's memory_limit_bytes==0 path
gets the same sentinel treatment, honoring ColiSegmentEngineOptions'
documented default-budget promise.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
JustVugg#1757 made num_experts 0 a valid config (Qwen3.8-27B and the other dense
checkpoints of the family): the layer MLP loads as an ungated shared expert
and nothing is routed, so the per-layer expert cache is never touched by
moe()/expert_get().

qwen36_cap_for_ram clamps its derived value against n_experts, which a dense
model does not have -- an unclamped RAM budget could size a cache that will
sit empty. Short-circuit to cap=1 instead, before the derivation runs.

Verified token-exact (16/16) on a qwen38-27b-dense tiny fixture at both an
explicit cap and the cap<=0 sentinel; the MoE path is unchanged.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Review on JustVugg#1747: main() keeps its own `cap` at the cap<=0 "auto" sentinel.
model_init_range resolves it into its own local copy and writes the result to
every layer's cache, but main() never reads it back, so the unresolved 0 is
what reaches qt_init.

qt_init refuses the VRAM expert tier for any cap != n_experts outside
fp8-stream mode (qwen36_tier.c:536). Not a regression -- dev's hardcoded 16
is equally != 256 for Qwen3.6-35B-A3B, so a bare invocation never got the
tier and callers passed `--cap 256` for it. But reading the value back makes
the automatic cap work where the old default could not: with enough RAM the
sentinel resolves to n_experts, and the tier comes up without an explicit cap.

qwen36_resolved_cap() keeps the shape qwen36_cap_for_ram() established: no
Model pointer, no globals, no I/O, so test_qwen36_cap_precedence.c pins it
without a container. An explicit cap passes through untouched; a null or
empty cache hands the sentinel on rather than inventing a value.

The guard sits behind COLI_CUDA and the CPU build links the inline qt_init
stub, so no CPU-only test can observe it -- the unit test pins the helper
instead, and fails on the pre-fix body.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ustVugg#1758)

JustVugg#1758 removed hand-written header lists from 202 rules and renamed
QWEN36_TIER_SRC to QWEN36_TIER_OBJ. The rule this PR added predates it and
named five headers plus the old variable, so after the rebase onto dev it no
longer resolved.

Now spelled exactly like its neighbour tests/test_qwen36_slot_int8: one
translation unit, $(QWEN36_CFLAGS) so -MMD -MP writes the .d, no header list.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@jtinbergen
jtinbergen force-pushed the dense-cap-default-v1 branch from 9fe4666 to 617f9de Compare October 1, 2026 13:07
@jtinbergen

Copy link
Copy Markdown
Contributor Author

Thanks. Fixed.

main() now reads the resolved value back through qwen36_resolved_cap(), in the same pure shape as qwen36_cap_for_ram() — no Model*, no globals, no I/O — so test_qwen36_cap_precedence.c pins it without a container. Six assertions: the sentinel resolves, an explicit cap passes through untouched, a null or empty cache hands the sentinel on rather than inventing a value. Fails on the pre-fix body (2 assertions), passes after.

One correction: this was not a regression. dev's hardcoded 16 is equally != 256 for Qwen3.6-35B-A3B, so a bare invocation never got the tier either — callers passed --cap 256, which is what the README's "needs full RAM residency" refers to. The unfixed PR did what dev does. The fix adds what the old default could not: with enough RAM the sentinel resolves to n_experts and the tier comes up without an explicit cap.

Makefile: rebased onto dev and rewritten in #1758's form. The old rule named five headers plus QWEN36_TIER_SRC, which #1758 renamed, so after the rebase it no longer resolved. Now spelled like its neighbour tests/test_qwen36_slot_int8: one translation unit, $(QWEN36_CFLAGS) so -MMD -MP writes the .d, no header list.

Local: make check 1967 tests OK (149 skipped), 0 warnings. The six test_qwen36_cap_budget cases skip here — no torch for the fixture — so CI and your run cover those.

make -C c cuda-test not run: no CUDA host here either. Nothing in the CUDA path changed though — a different integer reaches an existing guard, and that guard is what the new assertions pin.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants