Repository navigation
Conversation
The Markdown subset listed tables among the things it did not cover, on the grounds that an uncovered construct "degrades to the literal text instead of disappearing". For footnotes and raw HTML that holds. For tables it did not: with no table branch, every row fell through to the paragraph branch, which joins its lines with spaces — so a table arrived as a single run-on line with the pipes still in it, which is neither a table nor the literal text. Parse them instead. lib/markdown-table reads a GFM pipe table to plain data and the component maps it to a table element, so the escaping stance is unchanged: cell text goes through the same inline() path to React text nodes and still never reaches innerHTML. Column alignment comes from the delimiter row, the outer pipes are optional, and a backslash-escaped pipe stays content. A table inside a fence is still literal code, and a line that merely contains a pipe is still a paragraph — the delimiter row is what makes a run of pipes a table. Ragged rows are kept rather than truncated to the header width, which is where this departs from GFM: a short row is padded so the grid holds, and a long one keeps its extra cells. Models emit slightly ragged tables often enough that silently dropping their content is the worse failure. The parser is in lib/ so it is testable without a DOM, next to its 15 tests; the component's 9 render tests go through renderToStaticMarkup, which needs no DOM either, so both run in the default vitest environment. One asserts that markup written by the model inside a cell is escaped, since the file's no-sanitiser stance depends on that. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SQvpvgtVuZ71LE5WV8f3ax
…1758) #1758 removed hand-written header lists from 202 rules and renamed QWEN36_TIER_SRC to QWEN36_TIER_OBJ. The rule this PR added predates it and named five headers plus the old variable, so after the rebase onto dev it no longer resolved. Now spelled exactly like its neighbour tests/test_qwen36_slot_int8: one translation unit, $(QWEN36_CFLAGS) so -MMD -MP writes the .d, no header list, and $(VK_OBJ) on the link -- qwen36.c calls coli_vk_chain_decide since the Vulkan chain landed, which test_makefile_vk_obj.py enforces. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Install the Inkling oracle dependencies in both dense-only jobs. Compare warm-start text on stdout separately from streaming diagnostics, and derive the Kimi device-loss injection point from the measured frame count.
…ub.com/JustVugg/colibri into feat/vk-partial-chain
…culation feat(vulkan): integrate expert streaming, memory budgeting and speculative verification
qwen36 (Qwen3.6, Qwen3-Coder, Qwen3.8-27B and Clef's backbone) adopts the shared fit (vkc_fit). q36c_start computes each layer's device bytes from the shapes (its matrices as vk_qw_tensor places them, the DeltaNet b|a and shared gate rows, the DeltaNet state and conv ring, the K/V mirror at the split's floor of three blocks, its share of the parameter arena), the fixed bytes (one prompt chunk's scratch from the reservations themselves) and the head as the tail, before any upload and before the dense-host decision and the tier. The CPU's layers and, unless the fit puts it up, the head are marked vk_off: the per-matrix path never uploads them. N = 0 turns the chain off and releases its blocks. The chain now sets itself up at start, before the tier sizes its budget: the parameter arena, then each layer whole (a layer that fails is freed and N stops before it, vkc_fit_shrink), its host copies dropped once it is there when the dense weights live on the device only, then the head. The forward runs the N layers chunk by chunk and brings every row's residual back; the caller runs layers N..L-1 on the CPU and the head. Every loop over the chain's layers (mirrors, state sync, pushes, watermarks, verify copies, rollback) is bounded by N; a lost device rebuilds the N layers' recurrent state only. With every layer and the head fitting, the uploads are today's; the tier's dense_bytes is 0 because they are placed before it reads the free memory.
olmoe adopts the shared fit (vkc_fit) as qwen36 does. olc_fit_start, in model_init before the dense-host decision and the tier, decides the chain as olm_dho_start already did (silently; olc_start still prints the decision after the tier), brings its pipelines up and computes N from each layer's bytes (q, k, v, o and the router as f32 tensors, the K/V mirror at the split's floor of three blocks, its share of the parameter arena), the fixed bytes (one prompt chunk's scratch, its rows' read-back, PILOT's) and lm_head as the tail. The CPU's layers and, unless the fit puts it up, the head are marked refused, so matmul_res never uploads them. N = 0 turns the chain off and releases its blocks. olc_place sets the chain up there and then, layer by layer: a layer that does not reach the device is freed whole and N stops before it. With the dense weights on the device only each layer's host copies go once the layer is there, and the automatic cache grows by what was given back (only the N layers'). The forward runs the N layers and brings every row's residual back; the caller runs layers N..L-1 and the head on the CPU. The mirrors, pushes and watermarks cover the N layers; PILOT keeps prefetching the model's next layers from the chain's rows. A lost device redoes the step on the CPU, as before.
_VK_CHAIN_LAYOUT gains qwen36 (Qwen3.6, Qwen3-Coder, Qwen3.8-27B, Clef's backbone) and olmoe, so coli plan predicts their N with vkc_fit's rule and credits the first N layers' host copies alone. qwen36: each layer's matrices in the format the engine gives them (f32 with COLI_DENSE_I8=0, f16 with COLI_DENSE_BITS=16, int4 in groups of 64 where COLI_DENSE_BITS=4 and COLI_DENSE_INT4 take them, else int8 rows; int8 and f16 rows read in place under COLI_VK_IMPORT count their scales only), the DeltaNet b|a rows and the shared expert's gate row as f32, the DeltaNet state and conv ring, the K/V mirror at the split's floor, the layer's parameters; the fixed bytes are q36c_bufs at the fit's rows, their read-back and the final norm; lm_head is the tail. Shapes come from the scanned tensors, the scalars from qwen36_meta.json (config.json's without one). olmoe: q, k, v, o and the router as f32, the K/V mirror, the parameters; olc_bufs, the read-back and PILOT's rows; lm_head. tests/test_resource_plan.py: both layouts against the numbers the engines printed for the tiny fixtures on Lavapipe (qwen36 in f32, int8, f16, int4-g64 and with COLI_VK_CHAIN_ROWS; olmoe with and without PILOT), and the credit for every k of a budget aimed at k layers, forced, none, and every layer without and with room for the head.
"[VK] olmoe: N matmuls on the GPU" now says how many dense matrices the device holds and their size, as qwen36's line does: with a partial chain a test can check that the per-matrix path put nothing of the CPU's layers there. The count's format leaves the lines' parsers (the matmul count) as they were.
tests/vulkan_partial_qwen36-olmoe.sh (partial-qwen36-olmoe and partial-qwen36-olmoe-sanitize), qwen36 (the hybrid, Qwen3-Coder, the 27B dense geometry, int8, int4-g64 and f32 rows, an image) and olmoe on Lavapipe: - COLI_VK_CHAIN_LAYERS from 0 to L: the CPU's tokens and logits, N, the matrices on the device after setup the N layers' exactly, and with N below L the same count resident at exit (the per-matrix path put nothing more there); - COLI_VK_DEVICE_CAP_MB aimed at k layers from a probe's numbers: the line's N, the rule's from the run's own numbers, and coli plan's (free, per-layer and fixed bytes too); - a staged upload failing inside layer k's setup: N = k, nothing of layer k left; - with N below L: prompt chunks, the tiled GEMM, the tier off, prompts only, the per-matrix path beside, prompt-lookup verifies against the CPU and byte for byte against plain decoding, the KV split, a lost device mid-decode, serve sessions, the prefix-reuse, dashboard and Brio tests, Clef's oracle, PILOT; - the dense weights on the device only: the N layers' matrices alone dropped, none read back on a healthy run, only theirs after a lost device; - nothing forced: N = L, the head with it, the same logits byte for byte as COLI_VK_CHAIN_LAYERS=L. The sanitize variant runs a forced middle N, a cap, a fault in a layer, a lost device, verifies, the device-only weights and a serve session per engine. ci.yml: the two legs beside the dense-only ones. docs/vulkan.md: what the handoff moves and where each side keeps its state, per engine.
vkc_init allocated the chain's first block (its placeholder buffer) before vkc_fit read the free memory, so the fit saw it held and counted the pools' share on top: under COLI_VK_DEVICE_CAP_MB the free bytes the line printed were not the device's, and coli plan (which sees nothing held yet) predicted from other numbers. Both engines now bring the pipelines up once the fit decided N > 0; N = 0 leaves the chain's pipelines and blocks off the device altogether. Without the fit (PILOT, a geometry the shaders do not take, the CUDA tier) qwen36 initialises the chain as before, just later in main.
What their handoff moves (the residual rows alone) and where a verify's or PILOT's state stays, beside DeepSeek V4's row; the engines' own paragraphs say the rest.
qwen38's chain now takes the layers that fit the device instead of all or none (docs/vulkan.md, "A partial chain"): - q38c_start, at startup before any upload, the dense-host pass and the expert tier: vkc_fit from each layer's bytes (its matrices in the format they go up in, the shared expert's gate, its state at its starting size, its share of the parameter buffer), the scratch of one prompt chunk and the tail (the final mixer, lm_head and the MTP head's matrices). A partial chain is placed there and then, layer by layer; the full chain keeps its setup at the first forward, so with everything fitting the uploads and the tier's dense bytes are today's. N = 0 turns the chain off and gives its device memory back. - A layer that does not fully reach the device is freed with every layer after it and the tail (host copies read back where the dense-host pass dropped them), and the chain keeps the layers before it. From then on the per-matrix path uploads nothing new: the CPU's layers and the head keep their host copies. - The forward runs the N layers chunk by chunk and hands every row's four streams back once per chunk; the CPU runs layers N.. and the head from them. Every loop over the chain's state (mirrors, watermarks, verify copies, rollback, sync, recovery) covers the N layers; a lost device rebuilds only their state. - The dense-host pass drops the N layers' host copies only (and the head's with the tail); with the full chain fitted each layer is dropped once all of it is on the device.
…ir CI legs tests/vulkan_partial_qwen38.sh, sourced by tests/vulkan_engines.sh as partial-qwen38 and partial-qwen38-sanitize: COLI_VK_CHAIN_LAYERS for every k on the tiny fixture in every resident format, the int4-g64 experts and the PLE on either side of the handoff; the cap under which exactly k layers fit (a 256 MiB probe, so its pools' blocks are the run's); an upload refused inside a layer at the first forward, at startup and in the dense-host pass; prompt chunks and streaming, MTP and lookup drafts with the speculative harness's byte gate, serve sessions, the KV split, prompts only and a lost device, all with N < L; the host copies of the N layers only; and the full chain when everything fits. tools/make_qwen38_tiny.py --ple-layer moves the PLE (the default fixture is byte-identical). The chain's setup line now also counts the matrices resident, so a test can tell that nothing went up after a partial chain's setup; with no layer placed the chain's own pools go too.
What the fit counts for a Qwen3.8 layer, where the handoff falls, which state stays on which side (the PLE with its layer, the MTP head on the CPU reading the final streams), the lines on the tiny fixture under a device cap, and what was not measured (no checkpoint, no discrete GPU).
resource_plan._q38_chain_layout mirrors qwen38_chain.h's fit: each layer's matrices in the format q38_vk_fmt gives them (the trunk's int8 rows by Q38_TRUNK_CPU_INT8, Q38_TRUNK_MIN_KB and Q38_TRUNK_SKIP, else bf16 by Q38_NATIVE_BF16, else f32), the shared expert's gate, the DeltaNet state and ring with a verify's first copy, the K/V and index-key mirrors at the KV split's floor, the attention layers' read-back rows, the PLE ring and the parameters; the scratch of one chunk; the tail with the final mixer, lm_head and, under Q38_MTP=1, the MTP head priced from the config (the scan leaves its tensors out). Registered for vk_chain_fit, so coli plan predicts qwen38's N and credits only those layers' host copies. tests/test_resource_plan.py: the layout against the numbers the engine printed for the tiny fixture (bf16, the int8 trunk, f32, the PLE at layer 2, the MTP head), and the credit for 1 to 3 layers, every layer without the head, forced and off. The engine clamps COLI_VK_KV_BLOCK below 1 as the planner does.
… the family The fit now runs before vkc_init, as the pilot's does: nothing of the chain is on the device when it reads the free memory (vkc_init's first buffer took a pool block), so coli plan, which sees the device before the engine starts, predicts the same free bytes. N = 0 never brings the chain's pipelines up. The per-matrix gate goes on after a partial chain's own uploads. partial-qwen38 cross-checks coli plan against the engine's fit line under each cap (free, per-layer, fixed and tail bytes, N), on the int8 trunk with the MTP head too, and the full chain's logits with and without COLI_VK_CHAIN_LAYERS=4. docs/vulkan.md: qwen38's row in the partial chain's table, and where the MTP head runs.
A device the plan describes without a budget, a heap size or a cap (the existing tests' integrated GPU, say) left vk_chain_fit with no room, so qwen38's plan put N at 0 and credited no host copy. Its layout now gives no prediction there unless N is forced: the plan as before, every layer and every host copy, and the existing qwen38 planner tests pass unchanged.
…ne prints them for
GLM-5.2's dense part (9.9 GB) does not fit an 8 or 12 GB card whole. The chain now takes the layers that fit and the CPU runs the rest: - glmc_start runs right after the load, before the pins, cap_for_ram and the tier: the decision, the pipelines, the shapes check (glmc_check: no upload), then vkc_fit with each layer's device bytes (its matrices as glmc_tensor uploads them, its share of the parameter arena, its KV mirror at the first size), the fixed bytes (one prompt chunk's scratch at the planned context) and the tail (the MTP layer, eh_proj, lm_head). COLI_VK_CHAIN_LAYERS forces N. - glmc_setup places layers 0..N-1 one at a time. A layer that does not reach the device is freed whole and vkc_fit_shrink keeps the layers before it. - COLI_VK_DENSE_HOST: glm_dho_start decides on the N layers' bytes, glmc_setup drops a layer's host copies once all of it is placed, glm_dho_finish drops the tail only when the full chain took it. resident_bytes, and with it the pins and cap_for_ram, count only what was dropped. - glmc_forward runs every chunk through the N device layers and returns N. The CPU's layer loop runs from there. When layer N is a shared DSA layer, the device's selection comes back with the rows into the CPU's (dsa_sel, dsa_nsel). - The mirror, its watermarks, the KV split, push and pull cover the N layers only. - With a partial fit the per-matrix path multiplies on the device only what is there (VK_MAY): no lazy upload of a CPU layer or the tail. vk_dense_preload uploads nothing. The tier's dense_bytes adds the chain's first-forward scratch and mirrors. When everything fits, the uploads, the tier's budget and the tokens are as before.
GLM-5.3 Flash's chain takes the layers the device holds and the CPU runs the rest: - g53c_start moves into model_load_range, where the device opens, before expert_cache_init sizes the cache from the free memory and before the tier: the decision, the pipelines, the shapes check (g53c_check: no upload), then vkc_fit with each layer's device bytes (the mHC mixes and its matrices as g53c_setup uploads them, its share of the parameter arena, a KDA layer's state and window, an MLA layer's caches at the first size), the fixed bytes (one prompt chunk's scratch at GLM53_MAXT) and the tail (the head). COLI_VK_CHAIN_LAYERS forces N. - g53c_setup places layers 0..N-1 one at a time. A layer that does not reach the device is freed whole (g53c_layer_free) and vkc_fit_shrink keeps the layers before it. - COLI_VK_DENSE_HOST: g53_dho_start decides on the N layers' bytes, g53c_setup drops a layer's host copies once all of it is placed, g53_dho_finish drops the head only when the full chain took it. - g53c_forward runs the N device layers and returns N. run_layers goes on from layer N on the CPU with the streams, and its KDA state and MLA rows are the host's. - The KDA state's sync, push and the CPU step, the MLA mirror, the watermarks and the KV split cover the N layers only. After a lost device the rebuild runs the device's layers only: the CPU layers ran every position already. - With a partial fit mv and mm multiply on the device only what is there (G53_VK_MAY). The tier's dense_bytes is the chain's first-forward scratch and caches. When everything fits, the uploads, the tier's budget and the tokens are as before.
- tests/vulkan_partial_glm.sh (partial-glm, partial-glm-sanitize): colibri and glm53 with COLI_VK_CHAIN_LAYERS from 0 to L, COLI_VK_DEVICE_CAP_MB aimed at k layers, a staged upload failing inside layer k's setup, every forward path with N < L (prompt chunks, expert streaming, the DSA selection handed over at a shared indexer layer, n-gram and MTP drafts, prompts only, the per-matrix path beside, the tier off, the KV split, a lost device, serve sessions, glm53's image and pin-branch harness), the dense weights on the device only with N < L (the host copies dropped counted from the config), and the fit with everything fitting against N = L asked for. Two fixtures of its own: glm_tiny with shared indexer layers and a six-layer GLM-5.3 whose KDA and MLA layers alternate. - The engines print, with a partial fit, "[VK] <engine> chain: N of L layers held at exit: M B of matrices on the device": the family checks that the per-matrix path put nothing of a CPU layer or the tail on the device. - ci.yml: the partial-glm and partial-glm-sanitize legs beside the dense-only ones. - docs/vulkan.md: the partial chain in the GLM section.
_glm53_chain_layout is g53c_fit_plan transcribed: each layer's device bytes (the mHC mixes in f32, its matrices at GLM53_BITS as quantize_loaded leaves them or an int4 container's as it is, the absorbed kv_b halves, its share of the parameter arena, a KDA layer's state and window, an MLA layer's caches at their first 256 positions), the scratch of one prompt chunk (COLI_VK_CHAIN_ROWS, else 128 rows) at a serve slot's context (GLM53_MAXT, else 8192), and the head. build_plan then credits only the N layers' matrices. tests/test_resource_plan.py: the layout against the numbers glm53 printed for the partial-glm family's six-layer fixture on Lavapipe at GLM53_BITS 32, 8 and 4 and with a chunk of 7 rows at 300 positions, and the credit for N = 1, 3, 5 from the budget and forced, and for no layer. colibri (the glm family) gets no layout: its dense formats come from the command line's dense bits, which the scan does not see, and the plan gives it no device-only credit to scale with N.
…ss-checked colibri and glm53 now decide N before vkc_init, so the fit's free bytes are the device's before the chain holds its first pool blocks (vkc_fit_pools counts them), as coli plan sees it: under COLI_VK_DEVICE_CAP_MB the engine's free bytes are the cap. When the pipelines or the MLA shaders do not come up after the fit, the chain stays off and everything is as before. partial-glm: glm53's cap cases check coli plan's prediction against the engine's fit line (free, per-layer and fixed bytes, N), one of them on the 4-bit trunk with chunks of 7 rows; the device lost after an image aims at a decode step (the partial chain's forwards are one frame each there, so five frames back was the prompt's, before the device held any state); the bit-for-bit comparisons run the tier without its timing-driven balance; a dense-only case per engine with prompts only (kv_b and the indexer kept on the host).
…ned verify, docs - glm53's chain rounds differently with a chunk's rows with every layer on the device too (the binary before this branch: 3.6e-7 of 1.54 between chunks of 1 and 7 rows), so its chunk comparison holds the logits within the bound and the tokens exactly; colibri's stays bit for bit. - colibri with MTP drafts and COLI_EXACT_VERIFY=1: the verify declines the chain and runs every layer on the CPU, the other forwards take the device's two layers. - docs/vulkan.md: colibri and glm53 in the partial chain's table, and coli plan's glm53 layout in the GLM section (colibri gets none: its dense formats are the command line's).
… aligned to 16 bytes (#1908) COLI_VK_DEV was documented in docs/ENVIRONMENT.md and read nowhere: device selection ranked by type only, so with two discrete GPUs it always took the first one (an RTX 3050 beside an RTX 5060 Ti on #1908). It now takes the enumeration index it names (the order vulkaninfo lists), as proposed by ibboucco on #1908. An index out of range or not a number is ignored with a line, and the ranking applies. qmatmul_coop.comp staged each accumulator through shared memory with coopMatStore at a stride of 17 floats (68 bytes, odd to avoid bank conflicts). The stride of a cooperative-matrix load or store must be aligned to the lesser of 16 bytes and a row's natural alignment (VUID-RuntimeSpirv-OpCooperativeMatrixLoadKHR-08986): 16 bytes here. RADV tolerated it. On NVIDIA the default (cooperative) path returned black Qwen-Image pictures that changed from run to run, and COLI_VK_COOP=0 fixed them (#1908). The stride is now 20 floats (80 bytes), named once as VK_COOP_EST in backend_vulkan.c for both shared-memory budget checks. Verified on a Radeon 780M (RADV), the only cooperative-matrix device here: the VK_TEST harness passes before and after, with 64-lane and with 32-lane subgroups (COLI_VK_COOP_SG=32, the width NVIDIA uses), 20 cooperative GEMM cases each. Whether it fixes the NVIDIA cards needs the reporter's run. COLI_VK_DEV checked with Dozen (Iris Xe) and Lavapipe enumerated together.
The scheme of qwen38's: a Q36Seq per cache slot (K and V rows, DeltaNet, the token record, an image turn's rope positions), requests started through q36_serve_start (split out of serve_one) on their slot's state, one forward over a row of each active request a step (q36_step_rows), the attention and DeltaNet on each row's own state, the head once for every row. The KV is the whole context's for every slot, allocated before READY. KV_SLOTS=1 runs serve_one as before. With several slots nothing drafts, Q36_DN_GPU and the dense chain are off; a Clef DECIDE runs at once on its slot. serve_mux_check.py: a STOP may end in ERROR CANCELLED (qwen36 treats it as a cancel alone too), and a request ended early compares by the bytes it sent (the engine flushes the partial UTF-8 it held). test_qwen36_vision_serve: an image turn decoded beside two text turns gives the oracle's tokens, the text turns theirs.
An OlmSeq per cache slot (K and V rows, the token record), requests started through olm_serve_start (split out of serve_one) on their slot's KV, one forward over a row of each active request a step (olm_step_rows), the attention on each row's KV and lm_head once for every row. The KV and the expert cache share one room: kv_room_fit gets the positions every conversation's KV reaches, summed. KV_SLOTS=1 runs serve_one as before; with several slots the dense chain is off. As alone, CANCEL ends a request with DONE and STOP is not read.
A MimoSeq per cache slot (every layer's K and V, a sliding window's ring and its positions, where it stands, the token record), requests started through mimo_serve_start (split out of serve_loop) on their slot's caches, one forward over a row of each active request a step (forward_rows). attention_rows gives each row its own conversation's window or full cache, with attention()'s numbers for a block of that one token; the projections, the experts and lm_head run once for every row. KV_SLOTS=1 serves as before; with several slots the dense chain is off. As alone, STOP ends a request with DONE and CANCEL with ERROR CANCELLED.
An InkSeq per cache slot (K and V, a sliding layer's ring, the four short-convolution banks, the token record), requests started through ink_serve_start (split out of serve_one) on their slot's state, one forward over a row of each active request a step (ink_step_rows). The attention reads each row's cache and its own position as the scratch's first, the short convolutions run each row through its conversation's bank (ink_sconv), the projections, the experts and lm_head once for every row. The state is the whole context's for every slot, before READY. KV_SLOTS=1 serves as before; with several slots the dense chain is off. As alone, CANCEL ends a request with DONE and STOP is not read; each request keeps its repetition-penalty history.
A K3Seq per cache slot (every KDA layer's recurrent state and convolution windows, every MLA layer's latent caches and index keys, their capacity, the token record), requests started through k3_serve_start (split out of serve_one) on their slot's state, one forward over a row of each active request a step (k3_step_rows). The KDA and MLA read and write each row's own state and position; the projections, AttnRes, the experts and lm_head run once for every row. Each request keeps its chat framing: the XTML markers, the thinking close and the tool sideband. Its prefill is not interrupted (commands wait for it). KV_SLOTS=1 serves as before; with several slots the dense chain and Metal are off. As alone, STOP and CANCEL end with DONE.
… answers COLI_VULKAN=1 with no device is a refusal the engine reports with exit 2, and the msys2 shell runs steps with -e: the subshell stopped the verify before the [VK] line was read.
A V41Seq per cache slot: every layer's window ring and positions, compressed KV, index keys and partial compressor group, and the cross-layer state attention_run reads from the Model (the engram history, the index keys and top-k an index source last published, the candidate mask). Requests start through v41_serve_start (split out of serve_loop) on their slot's state; a decode step runs one forward over a row of each active request (forward_rows): the hyper-connection mixes, the routed experts and the head once over all rows, the engram and the attention row by row with that row's conversation bound (pointer swaps, no copies). KV_SLOTS=1 serves as before; with several slots DSpark stays unloaded and the dense chain is off. As alone, STOP ends with DONE and CANCEL with ERROR CANCELLED. serve_mux_check.py: MUX_LONG=k makes the prompts k times longer, past a window, compression groups and the index top-k.
KV_SLOTS=n in serve mode gives the engine n sessions, one a cache slot. A SUBMIT on a free slot is prefilled at once on its slot's session up to its first token (coli_v4_session_generate with prefill_only: the prompt cache, the checkpoints and the numeric channel work as alone), and every step then decodes one token of each active request as a row of one batch (coli_v4_sessions_step): the mixes, the hyper-connections and the routed experts take the rows together (the experts through the union over them, token-exact), each row's window attention runs on its own conversation's state, and the head reads its weights once for the batch. A request's frames are the ones it gets alone, byte for byte on the CPU (tests/serve_mux_check.py on the tiny fixture with a byte tokenizer, 4, 6 and 16 slots, prompts up to eight times longer, and under ASan and UBSan); on Lavapipe with the expert tier and with the dense matrices on the device, within 1e-4. STOP and CANCEL end a request with DONE, as alone. With more than one slot DSpark stays unloaded and the dense chain is off; the device's KV ring is dropped before each conversation's attention. KV_SLOTS=1 is the serve as before (v4_serve_one split into admit, marks and finish, which the multiplexed loop shares).
The prompts were slices of the base string from its start, so MUX_LONG only reached the slots past the sixth. Each prompt is now its slice times k.
qwen36, qwen38, OLMoE, Inkling and Kimi K3 answered "cache slot busy"; colibri, glm53, MiMo and DeepSeek answer SLOT_BUSY, the code docs/serve_protocol.md lists.
release: Vulkan engines and shaders in every archive, Metal on macOS, a Windows CUDA package
…shader paths and aligned frees On an Intel Iris Xe (Windows driver 101.7076) qwen36 answered garbage with the dense chain on (the default on an integrated GPU with the expert tier) and correctly with it off. tests/test_vk_chain on the device: every case of chain_gemv.comp and chain_gemv2.comp's long rows wrong (relative error 0.6 to 1.0). Both size their work by gl_SubgroupSize and gl_SubgroupID; the device compiles compute shaders at 8, 16 or 32 lanes and the chain's pipelines did not say which, so they ran at another width than the one they read. The same pipeline created with a required subgroup size of 8, 16 or 32 passes; a smaller shared-memory staging or one lane a row does not change anything. - backend_vulkan.c turns subgroup size control on (VK_EXT_subgroup_size_control) on a device whose compute subgroups can be more than one size, not only with cooperative matrices, and hands the reported size to the chain (ColiVkCore.pin_sg; the second device too). A device with one size (NVIDIA, Lavapipe) is left as it was. COLI_VK_SUBGROUP=0 keeps the driver's choice. - vk_chain.c's make_pipe requires that size for every pipeline it makes. - The chain found its optional shaders (MLA, KDA, mHC, AttnRes, DeepSeek V4, the split KV cache) by the last '/' of qmatmul.spv's path: on Windows the path has backslashes, so none of them loaded and their ops ran on the CPU. dir_prefix takes either separator. - vk_kvsplit.h freed its host shadows, from posix_memalign, with free(); on Windows that is _aligned_malloc and the heap broke at the first split KV test that ran there. Its own page allocator pairs them. glm53's mirror probe had the same free(). Iris Xe, tests/test_vk_chain: 23 failures before, PASS after (the optional shaders' ops included). qwen36 int4-gs64, "ciao": CPU "Ciao! Come posso aiutarti oggi?", chain on before the fix garbage, after it the same answer. Lavapipe: test_vk_chain and test_vk_tier PASS.
family_registry gives qwen36, qwen38, OLMoE, Inkling, Kimi K3, MiMo, DeepSeek V4 and V4.1 the 16 slots colibri and glm53 had: `coli serve --kv-slots N` (and coli setup, coli plan) take them, and the resource plan already counts each slot's state. api.md, serve_protocol.md, ENVIRONMENT.md and vulkan.md say what several conversations at once do on these engines: one batch a step, each row on its own conversation's state, the frames a request gets alone; no drafting, the dense chain off, the expert tier on.
tests/vulkan_engines.sh mux, mux-sanitize, mux-deepseek and mux-deepseek-sanitize: for qwen36, qwen38, OLMoE, Inkling, Kimi K3, MiMo, DeepSeek V4.1 and V4, tests/serve_mux_check.py on 4 slots and on 3 with prompts three times longer on the CPU (frame for frame), and on 4 with the expert tier on Lavapipe (logprobs within 1e-4); the same builds under ASan and UBSan. Kimi's tier gate runs with K3_IDOT=0: with the default int8 activations an expert's result depends on whether it is resident on the device, which the reference sessions and the shared one reach in another order (six replays of the family's sequence: 0 of 6 differ with it, the first differed without). colibri and glm53 keep theirs in glm-chain.
MiMo's logprobs reach -270, and with the expert tier the device's sums move their last digits by up to 1.5e-4 (6e-7 of the value) between the reference sessions and the shared one: the same tokens, frames 1.5e-4 apart. Its other Vulkan serve gates (vulkan_chain_serve.py in kv-split and mimo-chain) already compare within 1e-3; this one does too (three runs, all within it).
vulkan: the chain's pipelines at the device's subgroup size (Intel Iris Xe decoded garbage)
Signed-off-by: Rudy Celekli <47457359+rudycelekli@users.noreply.github.com>
Several conversations at once on every text engine (KV_SLOTS up to 16)
The version in c/version.py, the six READMEs' banners and the site's "Currently shipping" line; setup_catalog's prebuilt_since for the engines first shipped in this release (Qwen3-Coder, MiMo-V2.6 Flash and Pro, Qwen-Image) from the 1.12.2 that never came out to 2.0.0; the CHANGELOG entry: 166 pull requests since v1.12.1, 87 of them from contributors.
release 2.0.0
In all six READMEs: - CUDA: every release ships the Windows CUDA package built (colibri-<version>-windows-x86_64-cuda.zip: coli_cuda.dll for compute 8.0 and newer, with the colibri, qwen36 and kimi_k3 engines that load it); - Apple Silicon: the macOS archive's colibri, inkling and kimi_k3 have Metal built in, turned on with COLI_METAL=1 (K3_METAL=1 for Kimi K3); the paragraph used to say only "build with METAL=1"; - Use it from other apps: several conversations at once on every text engine (coli serve --kv-slots N, up to 16); - Install by hand: the Linux and Windows engines have Vulkan built in, the macOS ones Metal. The site: the prebuilt engine has Vulkan built in on Linux and Windows, the archives carry the GPU backends, and every language engine serves up to 16 conversations at once.
All six READMEs: coli setup --backend vulkan uses the GPU for any model, --backend cpu (or --no-gpu) keeps everything on the CPU; a Vulkan build uses the GPU only with COLI_VULKAN=1 (the setup sets it when it chose Vulkan), and COLI_VK_CHAIN=0 keeps the expert tier with the dense layers on the CPU. On an integrated GPU, try both: on an Intel Iris Xe laptop (i7-1355U) Qwen3.6 decoded 2.1 tok/s on the CPU, 1.7 to 1.9 with Vulkan and 2.1 with the dense chain off. The site names the two setup options.
…ntract-20261006 fix(setup): require Colibri health success before reusing a port
docs: the READMEs and the site for 2.0, #1959 in the changelog
Contributor
|
Fantastic work everyone lets keep this up! BTW we are looking for input on a new language we are bootstraping from hex this language is built specifically for AI/ML and closes the gaps in memory architecture and cache missing during inference time. We are are just steps away from self compilation expecting self comp in a week's time I'm looking for a fellow system 0 architect to help build the transpiler for Rust and C++ we are working around the clock to kick this off please contact me if you're interested in building complete computational sovereignty at JackKnifeAI@proton.me. The snake eats its tail. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Release 2.0.0: dev to main. These are the notes of CHANGELOG.md's 2.0.0 entry; release.yml takes them for the GitHub Release when the tag v2.0.0 is pushed. The site (site/) goes live from main through site.yml.
169 pull requests since v1.12.1, 88 of them from contributors. Every MoE
engine now runs on any GPU a Vulkan driver can see: the routed experts on a
shared device tier, the dense layers on a device chain (all of them, or the
first N that fit), and a second GPU for more experts. Three model families
arrive (MiMo-V2.6, Qwen3-Coder, Qwen-Image-2.1) and one dense checkpoint
(Qwen3.8-27B). Qwen3.8 Flash Next gets int4 experts and speculative decoding
that is on by default. Every text engine serves up to 16 conversations at once.
The release archives carry the GPU backends. Setup is one step, and there is one
decision API (System One).
Several conversations at once
coli serve --kv-slots N, up to 16), as GLM-5.2 and GLM-5.3 Flash did: qwen36 (Qwen3.6,Qwen3-Coder, Qwen3.8-27B), qwen38, OLMoE, Inkling, Kimi K3, MiMo-V2.6,
DeepSeek V4 and V4.1. Each slot keeps its own conversation's state, and every
decode step takes the next token of each active request as one batch, so the
weights and experts a step reads serve all of them. A request gets the tokens
it gets alone, checked frame for frame on every engine, on the CPU and with
the expert tier. With more than one slot nothing drafts and the dense chain
stays off.
Vulkan on every engine
and 12 added), each configuration checked against the engine's own CPU run on
Lavapipe (
tests/vulkan_engines.sh).cooperative-matrix GEMM where the device has one.
vk_tier.c). It holds acache of experts in device memory, filled from the expert history and adapting
while you chat, and computes the resident experts of each step as one async
batch while the CPU computes the rest. First on qwen38 and qwen36, then on
every MoE engine: GLM-5.2, GLM-5.3 Flash, Inkling, OLMoE, Kimi K3, MiMo-V2.6,
DeepSeek V4 and V4.1.
(
vk_chain.c), a layer's dense part in one submission, on qwen36, qwen38,MiMo-V2.6 (sliding window and sinks on the device), OLMoE, Inkling (and
OLMoE's PILOT race fixed), GLM-5.2 and GLM-5.3 Flash (shared MLA, DSA, KDA and
mHC ops), Kimi K3, DeepSeek V4.1 and DeepSeek V4.
the mapped path failed or spilled to system RAM.
speculative verification on the device.
layers that fit go to the device, the CPU runs the rest, and nothing is
uploaded that the chain will not use, so an 8 GB card with a larger trunk keeps
the chain and its expert tier instead of losing both.
COLI_VK_CHAIN_LAYERSforces N,
coli planpredicts it. DeepSeek V4, qwen36 and olmoe, qwen38,GLM-5.2 and GLM-5.3 Flash.
COLI_VK_DEV2=auto|<index>) onevery MoE engine: the experts after the primary device's, each step one batch
per device, both in flight at once. A failing second device gives its experts
back to the CPU and the tier goes on with the first.
COLI_VK_DEV2, thelayers a partial chain leaves go to the second device instead of the CPU, on
every chain engine (qwen36, qwen38, OLMoE, MiMo-V2.6, Inkling, GLM-5.2,
GLM-5.3 Flash, Kimi K3, DeepSeek V4.1 and V4).
COLI_VK_CHAIN_LAYERS2forceshow many,
COLI_VK_CHAIN_DEV2=0keeps them off. A lost second device leavesthe CPU to run its layers, the recurrent state rebuilt where it was held.
Checked on Lavapipe opened twice; two real GPUs not measured yet.
COLI_VK_KV_COLD=device(opt-in): with the KV cache split past thedevice's budget, the host's part of the attention runs on the device too, from
host memory it reads in place, so a step runs in one frame. The prompt gains
4-6% on a Radeon 780M, decode is slower there (hence off by default); a
dedicated GPU has not been measured.
multiplexed decode (
KV_SLOTS, one row from each active conversation) runs onthe device: the batch's matrices as one, each row's attention over its own
conversation's KV mirror (up to
COLI_VK_CHAIN_MUX, default 16, beside thechain's own) with its own DSA list; two devices and the partial chain too.
Every frame checked against the CPU's on Lavapipe; not measured with a real
checkpoint yet.
the matrix units for prompt attention and int8/int4 prompt GEMMs on the
chain, a step's experts as one grouped GEMM (decode: one grouped GEMV per
phase), qwen36's routing in parallel, and int8 decode GEMVs sized by the
matrix. On a Radeon 780M with Qwen3.6-35B-A3B the first token of a
1000-token prompt comes in 9.3 s instead of 12.6, decode 6.1 tok/s instead
of 5.7, perplexity unchanged.
Xe, the default there). Its pipelines now run at the subgroup size the device
reports; the device compiles at 8 to 32 lanes, and the chain's GEMVs had run
at another width than the one they read. On Windows the chain also finds its
optional shaders (MLA, KDA, mHC, DeepSeek V4, the split KV cache), and the
split cache frees its host memory with the matching call.
COLI_VK_DEV=<index>picks the device on a machine with two GPUs ofthe same kind; the cooperative GEMM's epilogue stride is aligned to 16 bytes, a
validation error on some drivers ([Bug]: qwen-image on vulkan and Nvidia GPU #1908).
default mode.
under
-static, and the Vulkan objects linked into every engine-includingtest rule.
Models
Pro): the engine, the vision tower, tool calling and Brio mode with logprobs,
read as released; Pro verified against Xiaomi's reference on the real
checkpoint, and
coli servenames the loaded variant.qwen3_moe) on the qwen36 engine.engine (Qwen3.8-27B), text and images.
terminal, the web app and the OpenAI images API.
7.5 s instead of 34 (four layers at once, quantized straight from bf16, the
same int8 bytes). With Vulkan every block of a denoising step runs on the
device, its attention included, and the VAE decodes there too; a card that
does not hold the 7.1 GB transformer keeps the blocks that fit and streams
the others each step. On a Radeon 780M a 512x512 image takes 61 s instead of
112 and a 1024x1024 step 38 s instead of 67; every stage checked against the
diffusers reference.
decode.
lossless, +12-14% on the release; with speculative decoding on by default: Qwen3.8's MTP head and gated prompt lookup; the plan prices the MTP head #1917 it is on by default with prompt
lookup on every engine that has it (
Q38_MTP=0,COLI_LOOKUP=0turn themoff), and
coli plancounts the MTP head's memory.writes.
System One: one decision API
POST /v1/systemone, one decision API, and the Laya decisionengine;
/v1/briois removed.GLiNER2's head).
Qwen3.8-27B, with a 16-bit dense mode.
mixes option token counts.
Setup and the CLI
release: Vulkan engines and shaders in every archive, Metal on macOS, a Windows CUDA package #1953: the release archives carry the GPU backends: every Linux and
Windows engine built with Vulkan and its shaders, the macOS engines with
Metal, and a Windows CUDA package (
coli_cuda.dllfor compute 8.0 and newer,with colibri, qwen36 and kimi_k3 built to load it). The Vulkan loader is
opened at run time, so these engines still start on a machine without one,
on the CPU.
coli setupuses them without a Vulkan toolchain.setup: one-step install and start (start-here, coli setup), AI_SETUP.md and an MCP server #1842: one-step install and start (
start-here,coli setup), theAI_SETUP.mdguide and an MCP server.setup: on an integrated GPU, Vulkan only where it was measured faster #1850, setup: CUDA only when the toolkit builds for the card; Vulkan otherwise, and a failed build falls back #1896: setup picks the backend that works on the machine: on an
integrated GPU, Vulkan only where it was measured faster; CUDA only when the
toolkit builds for the card, Vulkan otherwise, and a failed build falls back.
setup: bring an engine already here up to date with make before using it; coli logs says where a foreground server prints (#1852) #1903: an engine already built is brought up to date with
makebefore itis used (a
git pullused to keep running the old binary), andcoli logssays where a foreground server prints ([Feature]: Deb or pkg #1852).
setup: ship a starting expert history for Qwen3.6 and Qwen3.8 #1919: a fresh install of Qwen3.6-35B-A3B or Qwen3.8 Flash Next gets a
starting expert history, so its first run fills the GPU tier from the start:
on a Radeon 780M Qwen3.6's first run went from about 6 to 11 tok/s.
coli chat: a line that starts with a picture's path is a message, not a command #1817: in
coli chat, a line that starts with a picture's path is amessage, not a command.
fix(setup): require Colibri health success before reusing a port #1959 (@rudycelekli):
coli setupcounts a server on its port as a runningcolibri only when
/healthanswers as colibri does; another program's HTTP 200no longer passes for one, and the setup starts its own on the next free port.
Add resumable, hash-verified qpack installers for Hugging Face and static mirrors #1315 (@Avicennasis): resumable, hash-verified qpack installers for
Hugging Face and static mirrors.
fix(planner): honor cgroup memory limits #1316 (@Avicennasis): the planner honours cgroup memory limits.
Fixes from the issue triage: Windows build, Vulkan defaults, serve errors, DeepSeek V4 (#1945, #1900, #1941, #1906) #1949: on Windows every engine has its bare
make <engine>target (sixfell to make's built-in rule and linked without CUDA or Vulkan, [Bug]: Dynamic DLL linking on Windows fail in qwen38.c, but not qwen36.c #1945,
[Performance]: Qwen3.8-Flash-Next FP8 on the Vulkan expert tier, RTX 4070 Laptop 8 GB (Windows 11) #1900);
coli setupprefers an MSYS2 that can build Vulkan over a PATHcompiler that cannot, and installs libgomp;
COLI_VK_DENSE=0keeps the densechain off too (on an 8 GB laptop it had taken the expert tier's budget,
[Performance]: Qwen3.8-Flash-Next FP8 on the Vulkan expert tier, RTX 4070 Laptop 8 GB (Windows 11) #1900); DeepSeek V4 uses the physical cores on an SMT machine (+10% decode on
a Threadripper, Findings on Zen2/Windows with DS4F REAP-150B (v1.12.1): SMT thread default, non-monotonic --ram, cross-build token divergence (+#1136), and a few docs/UX items #1906);
coli planleaves a card below an engine's CUDAfloor out; the server says why its engine stopped (GLM-5.2 int4 on Windows: colibri.exe serve engine stops with "dispatcher stopped" on first request #1941); clearer messages
for a missing model and a CUDA DLL that does not load.
The serve contract
logprobs,top_logprobsandechoon the OpenAIendpoints, aligned by raw-stream span.
preserve_thinkingrendered so standard clients get KV prefix reuse./v1/messagesanswer.tool result's picture is not dropped; the modalities a model card claims are
the ones the engine loaded.
refused; the completions keepalive streams as a text chunk.
text parts, tools, messages and image URLs answer 400 instead of 500.
in detail and logs CANCEL, does not restore a pin over rows another branch
rewrote, handles cancellation during prompt processing, and continues from the
cache instead of prefilling again.
Correctness and security
model (
.coli_kv,.coli_usage.tmp,hot_pinned.bin,.coli_ckpt/) areopened without following a symlink and only as regular files, so a link
planted in a downloaded model directory no longer redirects the engine's
writes. Request grammars are limited to 32767 alternates and symbols per rule
(past that the walker's index wrapped) and are refused with a message.
colibri, kimi_k3 and deepseek_v41 check the indexer's and RoPE's head
dimensions in the config.
(@namespaceMarcello): the published advisory fixes carried to glm53, qwen36 and
deepseek_v41; glm53 refuses f32 tensors shorter than the config reads and
matrices whose shape differs from it; deepseek_v41 keeps engram lookups inside
their table; qwen36 gives the same logits at every expert-cache capacity on FMA
builds and never evicts an expert a MoE run is still reading.
tok.h, so Turkishtokenizes as HF does.
lowest index and its indexer never keeps -inf.
plausibility floor.
node_idno longer kills thecluster's topology and health requests, and a manifest that is not a JSON
object no longer crashes the experiments validator.
digest-bound one.
PROFmeasures the expert wait instead of printing 0and folding it into
attention_s([Feature]: Deb or pkg #1852).Speed on the CPU
prefill.
matmul_fp8, bit-exact, 6 to 10xat prefill shapes.
fmafchains atonce with the same bits on every FMA build, and one vectorized f32 matmul
serves quant.h, olmoe, inkling and qwen36 ([Performance]: hot-path float reductions are scalar and don't auto-vectorize (router matmul + MLA-absorb attention) #442).
token.
(
Q36_DN_GPU=1), expert homes per layer range (QT_HOME=layer), the sharedexpert offered on request.
container through bounded Metal slots, and MLX-affine dense weights and norms.
shared expert.
Build, tests and CI
-MMD -MP(Derive header prerequisites with `-MMD -MP` instead of listing them by hand #1741).
runners, whose CPUs keep bf16 on the CPU.
deterministic.
Docs
models per machine, GPUs, System One) in five languages, with the site.
in its own section, and the weights-from-disk options.
labels.