Skip to content

release 2.0.0: dev to main - #1966

Merged
JustVugg merged 931 commits into
mainfrom
dev
Oct 6, 2026
Merged

JustVugg merged 931 commits into
mainfrom
dev

Conversation

@JustVugg

@JustVugg JustVugg commented Oct 6, 2026

Copy link
Copy Markdown
Owner

Release 2.0.0: dev to main. These are the notes of CHANGELOG.md's 2.0.0 entry; release.yml takes them for the GitHub Release when the tag v2.0.0 is pushed. The site (site/) goes live from main through site.yml.

169 pull requests since v1.12.1, 88 of them from contributors. Every MoE
engine now runs on any GPU a Vulkan driver can see: the routed experts on a
shared device tier, the dense layers on a device chain (all of them, or the
first N that fit), and a second GPU for more experts. Three model families
arrive (MiMo-V2.6, Qwen3-Coder, Qwen-Image-2.1) and one dense checkpoint
(Qwen3.8-27B). Qwen3.8 Flash Next gets int4 experts and speculative decoding
that is on by default. Every text engine serves up to 16 conversations at once.
The release archives carry the GPU backends. Setup is one step, and there is one
decision API (System One).

Several conversations at once

  • Several conversations at once on every text engine (KV_SLOTS up to 16) #1955: every text engine serves several conversations at once (coli serve --kv-slots N, up to 16), as GLM-5.2 and GLM-5.3 Flash did: qwen36 (Qwen3.6,
    Qwen3-Coder, Qwen3.8-27B), qwen38, OLMoE, Inkling, Kimi K3, MiMo-V2.6,
    DeepSeek V4 and V4.1. Each slot keeps its own conversation's state, and every
    decode step takes the next token of each active request as one batch, so the
    weights and experts a step reads serve all of them. A request gets the tokens
    it gets alone, checked frame for frame on every engine, on the CPU and with
    the expert tier. With more than one slot nothing drafts and the dense chain
    stays off.

Vulkan on every engine

Models

System One: one decision API

Setup and the CLI

The serve contract

Correctness and security

Speed on the CPU

Build, tests and CI

Docs

rayhanhanaputra and others added 30 commits October 5, 2026 11:18
The Markdown subset listed tables among the things it did not cover, on the
grounds that an uncovered construct "degrades to the literal text instead of
disappearing". For footnotes and raw HTML that holds. For tables it did not:
with no table branch, every row fell through to the paragraph branch, which
joins its lines with spaces — so a table arrived as a single run-on line with
the pipes still in it, which is neither a table nor the literal text.

Parse them instead. lib/markdown-table reads a GFM pipe table to plain data
and the component maps it to a table element, so the escaping stance is
unchanged: cell text goes through the same inline() path to React text nodes
and still never reaches innerHTML.

Column alignment comes from the delimiter row, the outer pipes are optional,
and a backslash-escaped pipe stays content. A table inside a fence is still
literal code, and a line that merely contains a pipe is still a paragraph —
the delimiter row is what makes a run of pipes a table.

Ragged rows are kept rather than truncated to the header width, which is where
this departs from GFM: a short row is padded so the grid holds, and a long one
keeps its extra cells. Models emit slightly ragged tables often enough that
silently dropping their content is the worse failure.

The parser is in lib/ so it is testable without a DOM, next to its 15 tests;
the component's 9 render tests go through renderToStaticMarkup, which needs no
DOM either, so both run in the default vitest environment. One asserts that
markup written by the model inside a cell is escaped, since the file's
no-sanitiser stance depends on that.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SQvpvgtVuZ71LE5WV8f3ax
…1758)

#1758 removed hand-written header lists from 202 rules and renamed
QWEN36_TIER_SRC to QWEN36_TIER_OBJ. The rule this PR added predates it and
named five headers plus the old variable, so after the rebase onto dev it no
longer resolved.

Now spelled exactly like its neighbour tests/test_qwen36_slot_int8: one
translation unit, $(QWEN36_CFLAGS) so -MMD -MP writes the .d, no header list,
and $(VK_OBJ) on the link -- qwen36.c calls coli_vk_chain_decide since the
Vulkan chain landed, which test_makefile_vk_obj.py enforces.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Install the Inkling oracle dependencies in both dense-only jobs. Compare warm-start text on stdout separately from streaming diagnostics, and derive the Kimi device-loss injection point from the measured frame count.
…culation

feat(vulkan): integrate expert streaming, memory budgeting and speculative verification
qwen36 (Qwen3.6, Qwen3-Coder, Qwen3.8-27B and Clef's backbone) adopts the
shared fit (vkc_fit). q36c_start computes each layer's device bytes from the
shapes (its matrices as vk_qw_tensor places them, the DeltaNet b|a and shared
gate rows, the DeltaNet state and conv ring, the K/V mirror at the split's
floor of three blocks, its share of the parameter arena), the fixed bytes (one
prompt chunk's scratch from the reservations themselves) and the head as the
tail, before any upload and before the dense-host decision and the tier. The
CPU's layers and, unless the fit puts it up, the head are marked vk_off: the
per-matrix path never uploads them. N = 0 turns the chain off and releases
its blocks.

The chain now sets itself up at start, before the tier sizes its budget:
the parameter arena, then each layer whole (a layer that fails is freed and N
stops before it, vkc_fit_shrink), its host copies dropped once it is there
when the dense weights live on the device only, then the head. The forward
runs the N layers chunk by chunk and brings every row's residual back; the
caller runs layers N..L-1 on the CPU and the head. Every loop over the chain's
layers (mirrors, state sync, pushes, watermarks, verify copies, rollback) is
bounded by N; a lost device rebuilds the N layers' recurrent state only.

With every layer and the head fitting, the uploads are today's; the tier's
dense_bytes is 0 because they are placed before it reads the free memory.
olmoe adopts the shared fit (vkc_fit) as qwen36 does. olc_fit_start, in
model_init before the dense-host decision and the tier, decides the chain as
olm_dho_start already did (silently; olc_start still prints the decision after
the tier), brings its pipelines up and computes N from each layer's bytes (q,
k, v, o and the router as f32 tensors, the K/V mirror at the split's floor of
three blocks, its share of the parameter arena), the fixed bytes (one prompt
chunk's scratch, its rows' read-back, PILOT's) and lm_head as the tail. The
CPU's layers and, unless the fit puts it up, the head are marked refused, so
matmul_res never uploads them. N = 0 turns the chain off and releases its
blocks.

olc_place sets the chain up there and then, layer by layer: a layer that does
not reach the device is freed whole and N stops before it. With the dense
weights on the device only each layer's host copies go once the layer is
there, and the automatic cache grows by what was given back (only the N
layers'). The forward runs the N layers and brings every row's residual back;
the caller runs layers N..L-1 and the head on the CPU. The mirrors, pushes and
watermarks cover the N layers; PILOT keeps prefetching the model's next layers
from the chain's rows. A lost device redoes the step on the CPU, as before.
_VK_CHAIN_LAYOUT gains qwen36 (Qwen3.6, Qwen3-Coder, Qwen3.8-27B, Clef's
backbone) and olmoe, so coli plan predicts their N with vkc_fit's rule and
credits the first N layers' host copies alone.

qwen36: each layer's matrices in the format the engine gives them (f32 with
COLI_DENSE_I8=0, f16 with COLI_DENSE_BITS=16, int4 in groups of 64 where
COLI_DENSE_BITS=4 and COLI_DENSE_INT4 take them, else int8 rows; int8 and f16
rows read in place under COLI_VK_IMPORT count their scales only), the DeltaNet
b|a rows and the shared expert's gate row as f32, the DeltaNet state and conv
ring, the K/V mirror at the split's floor, the layer's parameters; the fixed
bytes are q36c_bufs at the fit's rows, their read-back and the final norm;
lm_head is the tail. Shapes come from the scanned tensors, the scalars from
qwen36_meta.json (config.json's without one). olmoe: q, k, v, o and the router
as f32, the K/V mirror, the parameters; olc_bufs, the read-back and PILOT's
rows; lm_head.

tests/test_resource_plan.py: both layouts against the numbers the engines
printed for the tiny fixtures on Lavapipe (qwen36 in f32, int8, f16, int4-g64
and with COLI_VK_CHAIN_ROWS; olmoe with and without PILOT), and the credit for
every k of a budget aimed at k layers, forced, none, and every layer without
and with room for the head.
"[VK] olmoe: N matmuls on the GPU" now says how many dense matrices the device holds
and their size, as qwen36's line does: with a partial chain a test can check that the
per-matrix path put nothing of the CPU's layers there. The count's format leaves the
lines' parsers (the matmul count) as they were.
tests/vulkan_partial_qwen36-olmoe.sh (partial-qwen36-olmoe and
partial-qwen36-olmoe-sanitize), qwen36 (the hybrid, Qwen3-Coder, the 27B dense
geometry, int8, int4-g64 and f32 rows, an image) and olmoe on Lavapipe:

- COLI_VK_CHAIN_LAYERS from 0 to L: the CPU's tokens and logits, N, the matrices
  on the device after setup the N layers' exactly, and with N below L the same
  count resident at exit (the per-matrix path put nothing more there);
- COLI_VK_DEVICE_CAP_MB aimed at k layers from a probe's numbers: the line's N,
  the rule's from the run's own numbers, and coli plan's (free, per-layer and
  fixed bytes too);
- a staged upload failing inside layer k's setup: N = k, nothing of layer k left;
- with N below L: prompt chunks, the tiled GEMM, the tier off, prompts only, the
  per-matrix path beside, prompt-lookup verifies against the CPU and byte for
  byte against plain decoding, the KV split, a lost device mid-decode, serve
  sessions, the prefix-reuse, dashboard and Brio tests, Clef's oracle, PILOT;
- the dense weights on the device only: the N layers' matrices alone dropped,
  none read back on a healthy run, only theirs after a lost device;
- nothing forced: N = L, the head with it, the same logits byte for byte as
  COLI_VK_CHAIN_LAYERS=L.

The sanitize variant runs a forced middle N, a cap, a fault in a layer, a lost
device, verifies, the device-only weights and a serve session per engine. ci.yml:
the two legs beside the dense-only ones. docs/vulkan.md: what the handoff moves
and where each side keeps its state, per engine.
vkc_init allocated the chain's first block (its placeholder buffer) before
vkc_fit read the free memory, so the fit saw it held and counted the pools'
share on top: under COLI_VK_DEVICE_CAP_MB the free bytes the line printed were
not the device's, and coli plan (which sees nothing held yet) predicted from
other numbers. Both engines now bring the pipelines up once the fit decided N
> 0; N = 0 leaves the chain's pipelines and blocks off the device altogether.
Without the fit (PILOT, a geometry the shaders do not take, the CUDA tier)
qwen36 initialises the chain as before, just later in main.
What their handoff moves (the residual rows alone) and where a verify's or PILOT's
state stays, beside DeepSeek V4's row; the engines' own paragraphs say the rest.
qwen38's chain now takes the layers that fit the device instead of all or none
(docs/vulkan.md, "A partial chain"):

- q38c_start, at startup before any upload, the dense-host pass and the expert
  tier: vkc_fit from each layer's bytes (its matrices in the format they go up in,
  the shared expert's gate, its state at its starting size, its share of the
  parameter buffer), the scratch of one prompt chunk and the tail (the final mixer,
  lm_head and the MTP head's matrices). A partial chain is placed there and then,
  layer by layer; the full chain keeps its setup at the first forward, so with
  everything fitting the uploads and the tier's dense bytes are today's. N = 0
  turns the chain off and gives its device memory back.
- A layer that does not fully reach the device is freed with every layer after it
  and the tail (host copies read back where the dense-host pass dropped them), and
  the chain keeps the layers before it. From then on the per-matrix path uploads
  nothing new: the CPU's layers and the head keep their host copies.
- The forward runs the N layers chunk by chunk and hands every row's four streams
  back once per chunk; the CPU runs layers N.. and the head from them. Every loop
  over the chain's state (mirrors, watermarks, verify copies, rollback, sync,
  recovery) covers the N layers; a lost device rebuilds only their state.
- The dense-host pass drops the N layers' host copies only (and the head's with
  the tail); with the full chain fitted each layer is dropped once all of it is on
  the device.
…ir CI legs

tests/vulkan_partial_qwen38.sh, sourced by tests/vulkan_engines.sh as
partial-qwen38 and partial-qwen38-sanitize: COLI_VK_CHAIN_LAYERS for every k on
the tiny fixture in every resident format, the int4-g64 experts and the PLE on
either side of the handoff; the cap under which exactly k layers fit (a 256 MiB
probe, so its pools' blocks are the run's); an upload refused inside a layer at
the first forward, at startup and in the dense-host pass; prompt chunks and
streaming, MTP and lookup drafts with the speculative harness's byte gate, serve
sessions, the KV split, prompts only and a lost device, all with N < L; the host
copies of the N layers only; and the full chain when everything fits.

tools/make_qwen38_tiny.py --ple-layer moves the PLE (the default fixture is
byte-identical). The chain's setup line now also counts the matrices resident, so
a test can tell that nothing went up after a partial chain's setup; with no layer
placed the chain's own pools go too.
What the fit counts for a Qwen3.8 layer, where the handoff falls, which state
stays on which side (the PLE with its layer, the MTP head on the CPU reading the
final streams), the lines on the tiny fixture under a device cap, and what was
not measured (no checkpoint, no discrete GPU).
resource_plan._q38_chain_layout mirrors qwen38_chain.h's fit: each layer's matrices
in the format q38_vk_fmt gives them (the trunk's int8 rows by Q38_TRUNK_CPU_INT8,
Q38_TRUNK_MIN_KB and Q38_TRUNK_SKIP, else bf16 by Q38_NATIVE_BF16, else f32), the
shared expert's gate, the DeltaNet state and ring with a verify's first copy, the
K/V and index-key mirrors at the KV split's floor, the attention layers' read-back
rows, the PLE ring and the parameters; the scratch of one chunk; the tail with the
final mixer, lm_head and, under Q38_MTP=1, the MTP head priced from the config (the
scan leaves its tensors out). Registered for vk_chain_fit, so coli plan predicts
qwen38's N and credits only those layers' host copies.

tests/test_resource_plan.py: the layout against the numbers the engine printed for
the tiny fixture (bf16, the int8 trunk, f32, the PLE at layer 2, the MTP head), and
the credit for 1 to 3 layers, every layer without the head, forced and off. The
engine clamps COLI_VK_KV_BLOCK below 1 as the planner does.
… the family

The fit now runs before vkc_init, as the pilot's does: nothing of the chain is on the
device when it reads the free memory (vkc_init's first buffer took a pool block), so
coli plan, which sees the device before the engine starts, predicts the same free bytes.
N = 0 never brings the chain's pipelines up. The per-matrix gate goes on after a
partial chain's own uploads.

partial-qwen38 cross-checks coli plan against the engine's fit line under each cap
(free, per-layer, fixed and tail bytes, N), on the int8 trunk with the MTP head too,
and the full chain's logits with and without COLI_VK_CHAIN_LAYERS=4. docs/vulkan.md:
qwen38's row in the partial chain's table, and where the MTP head runs.
A device the plan describes without a budget, a heap size or a cap (the existing
tests' integrated GPU, say) left vk_chain_fit with no room, so qwen38's plan put N at
0 and credited no host copy. Its layout now gives no prediction there unless N is
forced: the plan as before, every layer and every host copy, and the existing qwen38
planner tests pass unchanged.
GLM-5.2's dense part (9.9 GB) does not fit an 8 or 12 GB card whole. The chain now
takes the layers that fit and the CPU runs the rest:

- glmc_start runs right after the load, before the pins, cap_for_ram and the tier:
  the decision, the pipelines, the shapes check (glmc_check: no upload), then vkc_fit
  with each layer's device bytes (its matrices as glmc_tensor uploads them, its share
  of the parameter arena, its KV mirror at the first size), the fixed bytes (one
  prompt chunk's scratch at the planned context) and the tail (the MTP layer, eh_proj,
  lm_head). COLI_VK_CHAIN_LAYERS forces N.
- glmc_setup places layers 0..N-1 one at a time. A layer that does not reach the
  device is freed whole and vkc_fit_shrink keeps the layers before it.
- COLI_VK_DENSE_HOST: glm_dho_start decides on the N layers' bytes, glmc_setup drops a
  layer's host copies once all of it is placed, glm_dho_finish drops the tail only
  when the full chain took it. resident_bytes, and with it the pins and cap_for_ram,
  count only what was dropped.
- glmc_forward runs every chunk through the N device layers and returns N. The CPU's
  layer loop runs from there. When layer N is a shared DSA layer, the device's
  selection comes back with the rows into the CPU's (dsa_sel, dsa_nsel).
- The mirror, its watermarks, the KV split, push and pull cover the N layers only.
- With a partial fit the per-matrix path multiplies on the device only what is
  there (VK_MAY): no lazy upload of a CPU layer or the tail. vk_dense_preload uploads
  nothing. The tier's dense_bytes adds the chain's first-forward scratch and mirrors.

When everything fits, the uploads, the tier's budget and the tokens are as before.
GLM-5.3 Flash's chain takes the layers the device holds and the CPU runs the rest:

- g53c_start moves into model_load_range, where the device opens, before
  expert_cache_init sizes the cache from the free memory and before the tier: the
  decision, the pipelines, the shapes check (g53c_check: no upload), then vkc_fit with
  each layer's device bytes (the mHC mixes and its matrices as g53c_setup uploads them,
  its share of the parameter arena, a KDA layer's state and window, an MLA layer's
  caches at the first size), the fixed bytes (one prompt chunk's scratch at
  GLM53_MAXT) and the tail (the head). COLI_VK_CHAIN_LAYERS forces N.
- g53c_setup places layers 0..N-1 one at a time. A layer that does not reach the
  device is freed whole (g53c_layer_free) and vkc_fit_shrink keeps the layers before
  it.
- COLI_VK_DENSE_HOST: g53_dho_start decides on the N layers' bytes, g53c_setup drops
  a layer's host copies once all of it is placed, g53_dho_finish drops the head only
  when the full chain took it.
- g53c_forward runs the N device layers and returns N. run_layers goes on from layer N
  on the CPU with the streams, and its KDA state and MLA rows are the host's.
- The KDA state's sync, push and the CPU step, the MLA mirror, the watermarks and the
  KV split cover the N layers only. After a lost device the rebuild runs the device's
  layers only: the CPU layers ran every position already.
- With a partial fit mv and mm multiply on the device only what is there (G53_VK_MAY).
  The tier's dense_bytes is the chain's first-forward scratch and caches.

When everything fits, the uploads, the tier's budget and the tokens are as before.
- tests/vulkan_partial_glm.sh (partial-glm, partial-glm-sanitize): colibri and glm53
  with COLI_VK_CHAIN_LAYERS from 0 to L, COLI_VK_DEVICE_CAP_MB aimed at k layers, a
  staged upload failing inside layer k's setup, every forward path with N < L (prompt
  chunks, expert streaming, the DSA selection handed over at a shared indexer layer,
  n-gram and MTP drafts, prompts only, the per-matrix path beside, the tier off, the
  KV split, a lost device, serve sessions, glm53's image and pin-branch harness), the
  dense weights on the device only with N < L (the host copies dropped counted from
  the config), and the fit with everything fitting against N = L asked for. Two
  fixtures of its own: glm_tiny with shared indexer layers and a six-layer GLM-5.3
  whose KDA and MLA layers alternate.
- The engines print, with a partial fit, "[VK] <engine> chain: N of L layers held at
  exit: M B of matrices on the device": the family checks that the per-matrix path put
  nothing of a CPU layer or the tail on the device.
- ci.yml: the partial-glm and partial-glm-sanitize legs beside the dense-only ones.
- docs/vulkan.md: the partial chain in the GLM section.
_glm53_chain_layout is g53c_fit_plan transcribed: each layer's device bytes (the mHC
mixes in f32, its matrices at GLM53_BITS as quantize_loaded leaves them or an int4
container's as it is, the absorbed kv_b halves, its share of the parameter arena, a KDA
layer's state and window, an MLA layer's caches at their first 256 positions), the
scratch of one prompt chunk (COLI_VK_CHAIN_ROWS, else 128 rows) at a serve slot's
context (GLM53_MAXT, else 8192), and the head. build_plan then credits only the N
layers' matrices.

tests/test_resource_plan.py: the layout against the numbers glm53 printed for the
partial-glm family's six-layer fixture on Lavapipe at GLM53_BITS 32, 8 and 4 and with a
chunk of 7 rows at 300 positions, and the credit for N = 1, 3, 5 from the budget and
forced, and for no layer.

colibri (the glm family) gets no layout: its dense formats come from the command line's
dense bits, which the scan does not see, and the plan gives it no device-only credit to
scale with N.
…ss-checked

colibri and glm53 now decide N before vkc_init, so the fit's free bytes are the device's
before the chain holds its first pool blocks (vkc_fit_pools counts them), as coli plan
sees it: under COLI_VK_DEVICE_CAP_MB the engine's free bytes are the cap. When the
pipelines or the MLA shaders do not come up after the fit, the chain stays off and
everything is as before.

partial-glm: glm53's cap cases check coli plan's prediction against the engine's fit
line (free, per-layer and fixed bytes, N), one of them on the 4-bit trunk with chunks
of 7 rows; the device lost after an image aims at a decode step (the partial chain's
forwards are one frame each there, so five frames back was the prompt's, before the
device held any state); the bit-for-bit comparisons run the tier without its
timing-driven balance; a dense-only case per engine with prompts only (kv_b and the
indexer kept on the host).
…ned verify, docs

- glm53's chain rounds differently with a chunk's rows with every layer on the device
  too (the binary before this branch: 3.6e-7 of 1.54 between chunks of 1 and 7 rows),
  so its chunk comparison holds the logits within the bound and the tokens exactly;
  colibri's stays bit for bit.
- colibri with MTP drafts and COLI_EXACT_VERIFY=1: the verify declines the chain and
  runs every layer on the CPU, the other forwards take the device's two layers.
- docs/vulkan.md: colibri and glm53 in the partial chain's table, and coli plan's
  glm53 layout in the GLM section (colibri gets none: its dense formats are the
  command line's).
… aligned to 16 bytes (#1908)

COLI_VK_DEV was documented in docs/ENVIRONMENT.md and read nowhere: device
selection ranked by type only, so with two discrete GPUs it always took the
first one (an RTX 3050 beside an RTX 5060 Ti on #1908). It now takes the
enumeration index it names (the order vulkaninfo lists), as proposed by
ibboucco on #1908. An index out of range or not a number is ignored with a
line, and the ranking applies.

qmatmul_coop.comp staged each accumulator through shared memory with
coopMatStore at a stride of 17 floats (68 bytes, odd to avoid bank
conflicts). The stride of a cooperative-matrix load or store must be aligned
to the lesser of 16 bytes and a row's natural alignment
(VUID-RuntimeSpirv-OpCooperativeMatrixLoadKHR-08986): 16 bytes here. RADV
tolerated it. On NVIDIA the default (cooperative) path returned black
Qwen-Image pictures that changed from run to run, and COLI_VK_COOP=0 fixed
them (#1908). The stride is now 20 floats (80 bytes), named once as
VK_COOP_EST in backend_vulkan.c for both shared-memory budget checks.

Verified on a Radeon 780M (RADV), the only cooperative-matrix device here:
the VK_TEST harness passes before and after, with 64-lane and with 32-lane
subgroups (COLI_VK_COOP_SG=32, the width NVIDIA uses), 20 cooperative GEMM
cases each. Whether it fixes the NVIDIA cards needs the reporter's run.
COLI_VK_DEV checked with Dozen (Iris Xe) and Lavapipe enumerated together.
JustVugg and others added 28 commits October 6, 2026 16:21
The scheme of qwen38's: a Q36Seq per cache slot (K and V rows, DeltaNet, the token
record, an image turn's rope positions), requests started through q36_serve_start
(split out of serve_one) on their slot's state, one forward over a row of each
active request a step (q36_step_rows), the attention and DeltaNet on each row's own
state, the head once for every row. The KV is the whole context's for every slot,
allocated before READY. KV_SLOTS=1 runs serve_one as before. With several slots
nothing drafts, Q36_DN_GPU and the dense chain are off; a Clef DECIDE runs at once
on its slot.

serve_mux_check.py: a STOP may end in ERROR CANCELLED (qwen36 treats it as a
cancel alone too), and a request ended early compares by the bytes it sent (the
engine flushes the partial UTF-8 it held). test_qwen36_vision_serve: an image turn
decoded beside two text turns gives the oracle's tokens, the text turns theirs.
An OlmSeq per cache slot (K and V rows, the token record), requests started
through olm_serve_start (split out of serve_one) on their slot's KV, one forward
over a row of each active request a step (olm_step_rows), the attention on each
row's KV and lm_head once for every row. The KV and the expert cache share one
room: kv_room_fit gets the positions every conversation's KV reaches, summed.
KV_SLOTS=1 runs serve_one as before; with several slots the dense chain is off.
As alone, CANCEL ends a request with DONE and STOP is not read.
A MimoSeq per cache slot (every layer's K and V, a sliding window's ring and its
positions, where it stands, the token record), requests started through
mimo_serve_start (split out of serve_loop) on their slot's caches, one forward over
a row of each active request a step (forward_rows). attention_rows gives each row
its own conversation's window or full cache, with attention()'s numbers for a block
of that one token; the projections, the experts and lm_head run once for every row.
KV_SLOTS=1 serves as before; with several slots the dense chain is off. As alone,
STOP ends a request with DONE and CANCEL with ERROR CANCELLED.
An InkSeq per cache slot (K and V, a sliding layer's ring, the four short-convolution
banks, the token record), requests started through ink_serve_start (split out of
serve_one) on their slot's state, one forward over a row of each active request a
step (ink_step_rows). The attention reads each row's cache and its own position as
the scratch's first, the short convolutions run each row through its conversation's
bank (ink_sconv), the projections, the experts and lm_head once for every row. The
state is the whole context's for every slot, before READY. KV_SLOTS=1 serves as
before; with several slots the dense chain is off. As alone, CANCEL ends a request
with DONE and STOP is not read; each request keeps its repetition-penalty history.
A K3Seq per cache slot (every KDA layer's recurrent state and convolution windows,
every MLA layer's latent caches and index keys, their capacity, the token record),
requests started through k3_serve_start (split out of serve_one) on their slot's
state, one forward over a row of each active request a step (k3_step_rows). The KDA
and MLA read and write each row's own state and position; the projections, AttnRes,
the experts and lm_head run once for every row. Each request keeps its chat
framing: the XTML markers, the thinking close and the tool sideband. Its prefill is
not interrupted (commands wait for it). KV_SLOTS=1 serves as before; with several
slots the dense chain and Metal are off. As alone, STOP and CANCEL end with DONE.
… answers

COLI_VULKAN=1 with no device is a refusal the engine reports with exit 2, and the
msys2 shell runs steps with -e: the subshell stopped the verify before the [VK]
line was read.
A V41Seq per cache slot: every layer's window ring and positions, compressed KV,
index keys and partial compressor group, and the cross-layer state attention_run
reads from the Model (the engram history, the index keys and top-k an index source
last published, the candidate mask). Requests start through v41_serve_start (split
out of serve_loop) on their slot's state; a decode step runs one forward over a row
of each active request (forward_rows): the hyper-connection mixes, the routed
experts and the head once over all rows, the engram and the attention row by row
with that row's conversation bound (pointer swaps, no copies). KV_SLOTS=1 serves as
before; with several slots DSpark stays unloaded and the dense chain is off. As
alone, STOP ends with DONE and CANCEL with ERROR CANCELLED.

serve_mux_check.py: MUX_LONG=k makes the prompts k times longer, past a window,
compression groups and the index top-k.
KV_SLOTS=n in serve mode gives the engine n sessions, one a cache slot. A
SUBMIT on a free slot is prefilled at once on its slot's session up to its
first token (coli_v4_session_generate with prefill_only: the prompt cache,
the checkpoints and the numeric channel work as alone), and every step then
decodes one token of each active request as a row of one batch
(coli_v4_sessions_step): the mixes, the hyper-connections and the routed
experts take the rows together (the experts through the union over them,
token-exact), each row's window attention runs on its own conversation's
state, and the head reads its weights once for the batch.

A request's frames are the ones it gets alone, byte for byte on the CPU
(tests/serve_mux_check.py on the tiny fixture with a byte tokenizer, 4, 6
and 16 slots, prompts up to eight times longer, and under ASan and UBSan);
on Lavapipe with the expert tier and with the dense matrices on the device,
within 1e-4. STOP and CANCEL end a request with DONE, as alone. With more
than one slot DSpark stays unloaded and the dense chain is off; the device's
KV ring is dropped before each conversation's attention. KV_SLOTS=1 is the
serve as before (v4_serve_one split into admit, marks and finish, which the
multiplexed loop shares).
The prompts were slices of the base string from its start, so MUX_LONG only
reached the slots past the sixth. Each prompt is now its slice times k.
qwen36, qwen38, OLMoE, Inkling and Kimi K3 answered "cache slot busy";
colibri, glm53, MiMo and DeepSeek answer SLOT_BUSY, the code
docs/serve_protocol.md lists.
release: Vulkan engines and shaders in every archive, Metal on macOS, a Windows CUDA package
…shader paths and aligned frees

On an Intel Iris Xe (Windows driver 101.7076) qwen36 answered garbage with the
dense chain on (the default on an integrated GPU with the expert tier) and
correctly with it off. tests/test_vk_chain on the device: every case of
chain_gemv.comp and chain_gemv2.comp's long rows wrong (relative error 0.6 to
1.0). Both size their work by gl_SubgroupSize and gl_SubgroupID; the device
compiles compute shaders at 8, 16 or 32 lanes and the chain's pipelines did
not say which, so they ran at another width than the one they read. The same
pipeline created with a required subgroup size of 8, 16 or 32 passes; a
smaller shared-memory staging or one lane a row does not change anything.

- backend_vulkan.c turns subgroup size control on (VK_EXT_subgroup_size_control)
  on a device whose compute subgroups can be more than one size, not only
  with cooperative matrices, and hands the reported size to the chain
  (ColiVkCore.pin_sg; the second device too). A device with one size (NVIDIA,
  Lavapipe) is left as it was. COLI_VK_SUBGROUP=0 keeps the driver's choice.
- vk_chain.c's make_pipe requires that size for every pipeline it makes.
- The chain found its optional shaders (MLA, KDA, mHC, AttnRes, DeepSeek V4,
  the split KV cache) by the last '/' of qmatmul.spv's path: on Windows the
  path has backslashes, so none of them loaded and their ops ran on the CPU.
  dir_prefix takes either separator.
- vk_kvsplit.h freed its host shadows, from posix_memalign, with free(); on
  Windows that is _aligned_malloc and the heap broke at the first split KV
  test that ran there. Its own page allocator pairs them. glm53's mirror
  probe had the same free().

Iris Xe, tests/test_vk_chain: 23 failures before, PASS after (the optional
shaders' ops included). qwen36 int4-gs64, "ciao": CPU "Ciao! Come posso
aiutarti oggi?", chain on before the fix garbage, after it the same answer.
Lavapipe: test_vk_chain and test_vk_tier PASS.
family_registry gives qwen36, qwen38, OLMoE, Inkling, Kimi K3, MiMo,
DeepSeek V4 and V4.1 the 16 slots colibri and glm53 had: `coli serve
--kv-slots N` (and coli setup, coli plan) take them, and the resource plan
already counts each slot's state. api.md, serve_protocol.md, ENVIRONMENT.md
and vulkan.md say what several conversations at once do on these engines:
one batch a step, each row on its own conversation's state, the frames a
request gets alone; no drafting, the dense chain off, the expert tier on.
tests/vulkan_engines.sh mux, mux-sanitize, mux-deepseek and
mux-deepseek-sanitize: for qwen36, qwen38, OLMoE, Inkling, Kimi K3, MiMo,
DeepSeek V4.1 and V4, tests/serve_mux_check.py on 4 slots and on 3 with
prompts three times longer on the CPU (frame for frame), and on 4 with the
expert tier on Lavapipe (logprobs within 1e-4); the same builds under ASan
and UBSan. Kimi's tier gate runs with K3_IDOT=0: with the default int8
activations an expert's result depends on whether it is resident on the
device, which the reference sessions and the shared one reach in another
order (six replays of the family's sequence: 0 of 6 differ with it, the
first differed without). colibri and glm53 keep theirs in glm-chain.
MiMo's logprobs reach -270, and with the expert tier the device's sums move
their last digits by up to 1.5e-4 (6e-7 of the value) between the reference
sessions and the shared one: the same tokens, frames 1.5e-4 apart. Its
other Vulkan serve gates (vulkan_chain_serve.py in kv-split and mimo-chain)
already compare within 1e-3; this one does too (three runs, all within it).
vulkan: the chain's pipelines at the device's subgroup size (Intel Iris Xe decoded garbage)
Signed-off-by: Rudy Celekli <47457359+rudycelekli@users.noreply.github.com>
Several conversations at once on every text engine (KV_SLOTS up to 16)
The version in c/version.py, the six READMEs' banners and the site's
"Currently shipping" line; setup_catalog's prebuilt_since for the engines
first shipped in this release (Qwen3-Coder, MiMo-V2.6 Flash and Pro,
Qwen-Image) from the 1.12.2 that never came out to 2.0.0; the CHANGELOG
entry: 166 pull requests since v1.12.1, 87 of them from contributors.
In all six READMEs:
- CUDA: every release ships the Windows CUDA package built
  (colibri-<version>-windows-x86_64-cuda.zip: coli_cuda.dll for compute 8.0
  and newer, with the colibri, qwen36 and kimi_k3 engines that load it);
- Apple Silicon: the macOS archive's colibri, inkling and kimi_k3 have
  Metal built in, turned on with COLI_METAL=1 (K3_METAL=1 for Kimi K3); the
  paragraph used to say only "build with METAL=1";
- Use it from other apps: several conversations at once on every text
  engine (coli serve --kv-slots N, up to 16);
- Install by hand: the Linux and Windows engines have Vulkan built in, the
  macOS ones Metal.

The site: the prebuilt engine has Vulkan built in on Linux and Windows, the
archives carry the GPU backends, and every language engine serves up to 16
conversations at once.
All six READMEs: coli setup --backend vulkan uses the GPU for any model,
--backend cpu (or --no-gpu) keeps everything on the CPU; a Vulkan build uses
the GPU only with COLI_VULKAN=1 (the setup sets it when it chose Vulkan), and
COLI_VK_CHAIN=0 keeps the expert tier with the dense layers on the CPU. On an
integrated GPU, try both: on an Intel Iris Xe laptop (i7-1355U) Qwen3.6
decoded 2.1 tok/s on the CPU, 1.7 to 1.9 with Vulkan and 2.1 with the dense
chain off. The site names the two setup options.
…ntract-20261006

fix(setup): require Colibri health success before reusing a port
169 pull requests since v1.12.1, 88 of them from contributors: the 166
counted before, #1958 (the version bump), #1959 and this one.
docs: the READMEs and the site for 2.0, #1959 in the changelog
@JustVugg
JustVugg merged commit bf24429 into main Oct 6, 2026
119 checks passed
@JackKnifeAI

Copy link
Copy Markdown
Contributor

Fantastic work everyone lets keep this up! BTW we are looking for input on a new language we are bootstraping from hex this language is built specifically for AI/ML and closes the gaps in memory architecture and cache missing during inference time. We are are just steps away from self compilation expecting self comp in a week's time I'm looking for a fellow system 0 architect to help build the transpiler for Rust and C++ we are working around the clock to kick this off please contact me if you're interested in building complete computational sovereignty at JackKnifeAI@proton.me. The snake eats its tail.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants