Skip to content

feat(glm53): continue from the cache instead of re-prefilling - #1803

Merged
JustVugg merged 2 commits into
JustVugg:devfrom
enitimeago:feat/glm53-continue-rewind
Oct 1, 2026
Merged

JustVugg merged 2 commits into
JustVugg:devfrom
enitimeago:feat/glm53-continue-rewind

Conversation

@enitimeago

@enitimeago enitimeago commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Related to #1748.

Stopping a reply and clicking continue can make the glm53 engine process the entire prompt again, even though it has already cached the work.

The existing reuse path requires at least one new prompt token, hence an exact match misses, and so does a shorter prompt after the client trims trailing whitespace tokens to satisfy the continuation API. (For example, "# Hummingbird\n\n" cannot be continued by sending "# Hummingbird", whilst "# Hummingbird\n\n" is rejected by the gateway.)

This change supports preserving that work for two continuation cases:

  • Exact match (equal): retain the final logits row, which predicts the next token at the cached position. The next request can use that row even though it has no new prompt tokens to process.
  • Trimmed whitespace (shorter): save the KDA recurrent state, convolution windows, and logits before feeding the first whitespace-only token of a generated run. If the next prompt ends at that saved position and matches the cached prefix, restore it. This handles a client trimming the reply before continuation; the gateway rejects trailing whitespace.

Rules for both resume points:

  • They are valid only for the turn right after the one that made them. Each turn decides, clears them, and writes them again at the end.
  • Token IDs are checked against the slot's history before any pin restore.
  • Requests with image embeddings or logprobs use the existing path, and image turns leave no resume points: matching placeholder IDs cannot establish that the image embeddings match.
  • Whitespace snapshots recognize only ASCII whitespace-only tokens. A shorter prompt must match the saved boundary exactly; other cases use the existing prefill and pin logic.

This change does not provide arbitrary cache rewind.

The trimmed whitespace case is opt-in (GLM53_REWIND=0, documented in docs/ENVIRONMENT.md). It requires a buffer the size of the KDA state, about 149 MiB per slot, plus a logits row, on first use; the buffer is retained across slot resets and reused. The state is copied once at the start of each eligible generated whitespace run, including during ordinary generation.

Exact-match resumption stays on, since it costs one logits row.

Validation

  • make -C c check
  • CUDA changes were tested with make -C c cuda-test (if applicable) (N/A; no CUDA changes)
  • Performance claims include hardware, commands, and repeatable measurements (N/A; no timing or throughput claims)
  • Performance claims include a validated experiment manifest with raw evidence (N/A; no timing or throughput claims)

Serve harness checks (glm53-oracle CI job and vision serve harness passed locally):

  • Continue after a mid-turn CANCEL reuses the exact cached prompt (REUSE 20 3 3 3 3 equal) and matches a fresh engine's output.
  • A generated run of two whitespace tokens is tested at two stopping points. Ending the budget at the second whitespace token leaves only the first cached. Allowing one more token of budget feeds both whitespace tokens, exercising the snapshot guard on the second as well. Continuing from the prefix before the run must report shorter, reuse that prefix, and match a fresh engine's output in both cases. These are deterministic cache-state tests using token budgets; the existing CANCEL case covers cancellation.
  • The first whitespace case reuses zero tokens with GLM53_REWIND=0.
  • The vision serve harness ends an image turn after one token, then sends the same prompt without IMAGE. It must reuse nothing; without the image rule it resumed all 22 tokens.

Compatibility

  • The default CPU build remains dependency-free
  • No model files, generated binaries, or benchmark artifacts are included

A continue after a client disconnect resent a prompt the slot already held
and got a full prefill anyway. Measured on a CPU host with GLM53_VERBOSE,
three mid-reply disconnects lost no tokens in flight, and every continue
was one of two shapes:

  REUSE 2 0 68 68 68 equal
  REUSE 4 0 42 43 42 shorter

`equal`: the prompt is exactly the cached history. Reuse needs at least one
new token to prefill, because that is where the next logits come from, so it
was refused. The engine now keeps the logits for the session's final
position when a turn ends (the row it already had, not a copy) and resumes
from them. Keeping them meant freeing `logits` only when a newer row
replaces it, so every exit from the decode loop leaves the row for `filled`.

`shorter`: the reply ended in a whitespace run ("# Hummingbird\n\n"). The
gateway refuses a continued turn that ends in whitespace (the template
strips it), so the client trims it and the continue arrives short of the
cache by the trailing whitespace tokens. The KDA recurrent state can't
rewind, so the engine now snapshots it, with the logits row that chose the
token, just before the first whitespace-only token of each generated run is
fed. That's one copy per run, not per token. A continue that stops exactly
there resumes from the snapshot. Only ASCII whitespace-only tokens trigger
it; anything else str.rstrip() removes retokenizes differently and falls
back to a prefill, which is slow but never wrong. GLM53_REWIND=0 turns the
snapshot off; it costs a buffer the size of the KDA state (~149 MiB).

Both resume points are valid only for the turn right after the one that
made them: each turn decides, clears them, and writes them again at the
end. The id check is against the slot's history, the record of what the
rows hold, as JustVugg#1751 does. They are decided before a pin restore can
change the state, and skipped with an image (ids don't describe it) and
with logprobs (the ECHO reading has to prefill the positions it reports).
A turn that carried an image leaves neither: the next request is checked
by ids only, and a text-only request with the same placeholder ids would
otherwise resume from rows built with an image it never sent.

On the same host the two continues now reuse all 68 and 42 tokens instead
of none. The `equal` text matches an uninterrupted temperature-0 reply, and
the `shorter` one writes the trimmed "\n\n" again byte for byte.

The serve harness now expects the `equal` continue after a mid-turn CANCEL
to reuse all of it (`REUSE 20 3 3 3 3 equal`) and to answer like a freshly
started engine. A new case searches the tiny fixture for a reply with text
and then two whitespace tokens, and runs it twice: stopped at the second
whitespace token, so the cache is one token past the trimmed continue, and
one token later, so both whitespace tokens are in the cache and the second
has passed the snapshot point too. Each expects `shorter` resumed at the
trimmed length and the same answer as a fresh engine; the first also expects
no reuse under GLM53_REWIND=0. The vision-serve harness ends an image turn
after one token and sends the same prompt without IMAGE: it must reuse
nothing (without the image rule it resumed all 22 tokens).

The cases fail without the fix: against the parent engine the serve harness
fails at request 20, and with the snapshot restore removed it reuses the
right count but the answer differs, which the fresh-engine comparison
catches. With a snapshot taken at every whitespace token instead of once per
run, the later stop reuses nothing. The pin, multimodal, vision-serve
harnesses and the dashboard and context-exceeded unit tests pass.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@enitimeago
enitimeago force-pushed the feat/glm53-continue-rewind branch from 9617138 to 15c4b8c Compare September 29, 2026 14:08
@enitimeago enitimeago changed the title feat(glm53): resume continue from the cache instead of re-prefilling feat(glm53): continue from the cache instead of re-prefilling Sep 29, 2026
@JustVugg

JustVugg commented Oct 1, 2026

Copy link
Copy Markdown
Owner

Thanks, this works: the oracles and serve harnesses pass, also under ASan and UBSan, including the new equal and shorter cases compared against a fresh engine.

One change before merging, about the default. The whitespace snapshot costs about 149 MiB per slot outside the GLM53_EXPERT_GB budget, and a full state copy at every generated whitespace run. Every reply has whitespace runs, so with the default on, every user pays both the memory and the copies during ordinary generation, including users who never continue a reply. On the machines GLM-5.3 targets that margin is small. Please make GLM53_REWIND default to 0 (opt-in) and update docs/ENVIRONMENT.md; exact-match resumption stays on, since it costs one logits row.

About #1792 (image-keyed KV prefix): whichever of the two lands second has to fold in the other. The points are:

  • slot_shared and slot_remember work on the keys;
  • keys[total] = next stays in the reworked decode loop;
  • slot_pin_restore(..., common) sits under resume ? 0 :;
  • vision_new, n_new are passed to forward_prefill;
  • the "no resume point after an image turn" rule should test image_here rather than n_vision == 0, since after glm53: keep the KV prefix across image requests by keying the picture #1792 a turn whose picture was fully reused also has n_vision == 0.

The whitespace snapshot costs about 149 MiB per slot outside the
GLM53_EXPERT_GB budget, plus a full KDA state copy at every generated
whitespace run. Every reply has whitespace runs, so with the snapshot on by
default every user paid both, including users who never continue a reply.
GLM53_REWIND now defaults to 0; set it to 1 to rewind a continue that trims
trailing whitespace. Exact-match resumption stays on, since it costs one
logits row.

The serve harness turns GLM53_REWIND=1 on for the rewind cases and checks
that a trimmed continue re-prefills under the default.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@enitimeago

Copy link
Copy Markdown
Contributor Author

@JustVugg Done in 1ba5b09.

GLM53_REWIND now defaults to 0 and is enabled with GLM53_REWIND=1; docs/ENVIRONMENT.md is updated and notes the snapshot sits outside GLM53_EXPERT_GB. Exact-match resumption is unchanged. The serve harness turns rewind on for the shorter cases and now checks that a trimmed continue re-prefills under the default. Harness and tiny oracles pass.

Noted on the image-keyed KV prefix PR, including image_here instead of n_vision == 0 for the no-resume-after-image rule. I'll fold it in if this lands second.

@JustVugg
JustVugg merged commit 1b0dd1b into JustVugg:dev Oct 1, 2026
31 checks passed
@enitimeago
enitimeago deleted the feat/glm53-continue-rewind branch October 1, 2026 11:20
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants