feat(glm53): handle cancellation between layers of a prompt chunk - #1822
Open
enitimeago wants to merge 1 commit into
Open
enitimeago wants to merge 1 commit into
enitimeago wants to merge 1 commit into
Conversation
CANCEL was checked only between prefill chunks, so a cancel that arrived early in a chunk held the engine until the chunk finished. Smaller chunks are not the answer: routed experts are re-read once per chunk per layer, so every prompt would pay for a shorter wait (measurements in JustVugg#1748). run_layers now asks a halt hook before each layer. forward_prefill arms it only for the duration of a chunk, after copying the KDA state (~149 MiB on GLM-5.3-Flash, once per chunk). If the hook fires, run_layers returns NULL, forward_span leaves `filled` where it was, and forward_prefill puts the KDA state back to the start of the chunk and stops. The layers already run have written their DSA rows for the chunk's positions; those are positional and the next prefill rewrites them. The cache is left at the previous chunk boundary, exactly as a between-chunk cancel leaves it, so a retry reuses it the same way. Decode and the other callers of run_layers never see the hook. If the copy can't be allocated, that chunk falls back to the between-chunk check. GLM53_VERBOSE prints `HALT layer <i> of <n>, back to <filled>`. The serve harness adds a case: prime 10 tokens, then a 1000-token prompt in one chunk with CANCEL sent after a delay. The delays are fractions of the chunk's time, measured first on the same machine without a CANCEL, since fixed delays miss the chunk on a machine ten times faster. It needs a HALT line, and retries other delays rather than passing without one. It expects `CANCEL 51 1000 0 10`, a retry that reuses the 10 tokens, and the same answer as a fresh engine. The parent engine never halts mid-chunk, and with the restore removed the answer differs; the harness fails on both. The pin, multimodal, vision-serve harnesses and the dashboard and context-exceeded unit tests pass. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
enitimeago
force-pushed
the
feat/glm53-prefill-layer-cancel
branch
from
October 1, 2026 14:42
98c1076 to
9249d99
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Closes #1748.
Since #1789, glm53 checks
CANCELbefore each chunk of prompt tokens, so a cancellation that arrives just after a chunk starts waits for that whole chunk. ReducingGLM53_PREFILL_CHUNKshortens the wait but makes every prompt slower, because routed experts are read again for every chunk and every layer (see the alternatives in #1748).In serve mode, the engine now also checks between layers. A chunk stopped part-way is discarded:
forward_prefillcopies the KDA recurrent state and convolution windows at the start of each chunk, then arms a halt hook for the duration of that chunk only.run_layersasks the hook before each layer after the first. If it fires, it returns without finishing the chunk, andfilleddoes not advance.forward_prefillrestores the KDA state from the copy and returns as a between-chunk cancellation does. The DSA rows the completed layers wrote for the chunk's positions are positional, and the next prefill overwrites them.The cache is therefore left at the previous chunk boundary, exactly as #1789 leaves it, and a retry reuses it the same way. Decode and other callers of
run_layersnever see the hook.GLM53_VERBOSE=1printsHALT layer <i> of <n>, back to <filled>.Cost. The copy needs a buffer the size of the KDA state, about 149 MiB on GLM-5.3-Flash, outside the
GLM53_EXPERT_GBbudget. It is one buffer per process, not per slot, allocated on the first serve-mode prefill and kept for reuse. The state is copied once per prompt chunk, never during generation. If the buffer can't be allocated, that chunk falls back to the between-chunk check. CLI runs pass no cancellation hook and allocate nothing.The
GLM53_PREFILL_CHUNKrow indocs/ENVIRONMENT.mdnow says cancellation is checked between layers and that chunk size no longer bounds the wait.Validation
make -C c checkmake -C c cuda-test(if applicable) (N/A; no CUDA changes)c/tests/glm53_serve_harness.pyadds a case withGLM53_PREFILL_CHUNK=4096, so the whole prompt is a single chunk:CANCELafter a delay. Expect aHALTline,ERROR CANCELLEDandCANCEL 51 1000 0 10. A retry must reuse the 10 tokens (extend) and produce the same output as a fresh engine.CANCELlands before or after the chunk, it retries with another delay (five in all). If none lands inside the chunk, it fails rather than passing without having tested anything.The pin-branch, multimodal and vision-serve harnesses and the dashboard and context-exceeded tests also pass.
Compatibility