Skip to content

feat(glm53): handle cancellation between layers of a prompt chunk - #1822

Open
enitimeago wants to merge 1 commit into
JustVugg:devfrom
enitimeago:feat/glm53-prefill-layer-cancel
Open

enitimeago wants to merge 1 commit into
JustVugg:devfrom
enitimeago:feat/glm53-prefill-layer-cancel

Conversation

@enitimeago

Copy link
Copy Markdown
Contributor

Summary

Closes #1748.

Since #1789, glm53 checks CANCEL before each chunk of prompt tokens, so a cancellation that arrives just after a chunk starts waits for that whole chunk. Reducing GLM53_PREFILL_CHUNK shortens the wait but makes every prompt slower, because routed experts are read again for every chunk and every layer (see the alternatives in #1748).

In serve mode, the engine now also checks between layers. A chunk stopped part-way is discarded:

  • forward_prefill copies the KDA recurrent state and convolution windows at the start of each chunk, then arms a halt hook for the duration of that chunk only.
  • run_layers asks the hook before each layer after the first. If it fires, it returns without finishing the chunk, and filled does not advance.
  • forward_prefill restores the KDA state from the copy and returns as a between-chunk cancellation does. The DSA rows the completed layers wrote for the chunk's positions are positional, and the next prefill overwrites them.

The cache is therefore left at the previous chunk boundary, exactly as #1789 leaves it, and a retry reuses it the same way. Decode and other callers of run_layers never see the hook. GLM53_VERBOSE=1 prints HALT layer <i> of <n>, back to <filled>.

Cost. The copy needs a buffer the size of the KDA state, about 149 MiB on GLM-5.3-Flash, outside the GLM53_EXPERT_GB budget. It is one buffer per process, not per slot, allocated on the first serve-mode prefill and kept for reuse. The state is copied once per prompt chunk, never during generation. If the buffer can't be allocated, that chunk falls back to the between-chunk check. CLI runs pass no cancellation hook and allocate nothing.

The GLM53_PREFILL_CHUNK row in docs/ENVIRONMENT.md now says cancellation is checked between layers and that chunk size no longer bounds the wait.

Validation

  • make -C c check
  • CUDA changes were tested with make -C c cuda-test (if applicable) (N/A; no CUDA changes)
  • Performance claims include hardware, commands, and repeatable measurements (N/A; no timing or throughput claims)
  • Performance claims include a validated experiment manifest with raw evidence (N/A; no timing or throughput claims)

c/tests/glm53_serve_harness.py adds a case with GLM53_PREFILL_CHUNK=4096, so the whole prompt is a single chunk:

  • Prime a 10-token cache, then submit a 1000-token prompt and send CANCEL after a delay. Expect a HALT line, ERROR CANCELLED and CANCEL 51 1000 0 10. A retry must reuse the 10 tokens (extend) and produce the same output as a fresh engine.
  • This is the harness's only timed case. If the CANCEL lands before or after the chunk, it retries with another delay (five in all). If none lands inside the chunk, it fails rather than passing without having tested anything.
  • Negative controls: the parent engine never halts mid-chunk, and with the KDA restore removed the retry's output differs from a fresh engine's. The harness fails on both.

The pin-branch, multimodal and vision-serve harnesses and the dashboard and context-exceeded tests also pass.

Compatibility

  • The default CPU build remains dependency-free
  • No model files, generated binaries, or benchmark artifacts are included

CANCEL was checked only between prefill chunks, so a cancel that arrived
early in a chunk held the engine until the chunk finished. Smaller chunks
are not the answer: routed experts are re-read once per chunk per layer,
so every prompt would pay for a shorter wait (measurements in JustVugg#1748).

run_layers now asks a halt hook before each layer. forward_prefill arms it
only for the duration of a chunk, after copying the KDA state (~149 MiB on
GLM-5.3-Flash, once per chunk). If the hook fires, run_layers returns
NULL, forward_span leaves `filled` where it was, and forward_prefill puts
the KDA state back to the start of the chunk and stops. The layers already
run have written their DSA rows for the chunk's positions; those are
positional and the next prefill rewrites them. The cache is left at the
previous chunk boundary, exactly as a between-chunk cancel leaves it, so
a retry reuses it the same way. Decode and the other callers of run_layers
never see the hook. If the copy can't be allocated, that chunk falls back
to the between-chunk check. GLM53_VERBOSE prints `HALT layer <i> of <n>,
back to <filled>`.

The serve harness adds a case: prime 10 tokens, then a 1000-token prompt in
one chunk with CANCEL sent after a delay. The delays are fractions of the
chunk's time, measured first on the same machine without a CANCEL, since
fixed delays miss the chunk on a machine ten times faster. It needs a HALT
line, and retries other delays rather than passing without one. It expects
`CANCEL 51 1000 0 10`, a retry that reuses the 10 tokens, and the same
answer as a fresh engine. The parent engine never halts mid-chunk, and
with the restore removed the answer differs; the harness fails on both. The
pin, multimodal, vision-serve harnesses and the dashboard and
context-exceeded unit tests pass.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@enitimeago
enitimeago force-pushed the feat/glm53-prefill-layer-cancel branch from 98c1076 to 9249d99 Compare October 1, 2026 14:42

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant