feat(glm53): continue from the cache instead of re-prefilling - #1803
Conversation
A continue after a client disconnect resent a prompt the slot already held
and got a full prefill anyway. Measured on a CPU host with GLM53_VERBOSE,
three mid-reply disconnects lost no tokens in flight, and every continue
was one of two shapes:
REUSE 2 0 68 68 68 equal
REUSE 4 0 42 43 42 shorter
`equal`: the prompt is exactly the cached history. Reuse needs at least one
new token to prefill, because that is where the next logits come from, so it
was refused. The engine now keeps the logits for the session's final
position when a turn ends (the row it already had, not a copy) and resumes
from them. Keeping them meant freeing `logits` only when a newer row
replaces it, so every exit from the decode loop leaves the row for `filled`.
`shorter`: the reply ended in a whitespace run ("# Hummingbird\n\n"). The
gateway refuses a continued turn that ends in whitespace (the template
strips it), so the client trims it and the continue arrives short of the
cache by the trailing whitespace tokens. The KDA recurrent state can't
rewind, so the engine now snapshots it, with the logits row that chose the
token, just before the first whitespace-only token of each generated run is
fed. That's one copy per run, not per token. A continue that stops exactly
there resumes from the snapshot. Only ASCII whitespace-only tokens trigger
it; anything else str.rstrip() removes retokenizes differently and falls
back to a prefill, which is slow but never wrong. GLM53_REWIND=0 turns the
snapshot off; it costs a buffer the size of the KDA state (~149 MiB).
Both resume points are valid only for the turn right after the one that
made them: each turn decides, clears them, and writes them again at the
end. The id check is against the slot's history, the record of what the
rows hold, as JustVugg#1751 does. They are decided before a pin restore can
change the state, and skipped with an image (ids don't describe it) and
with logprobs (the ECHO reading has to prefill the positions it reports).
A turn that carried an image leaves neither: the next request is checked
by ids only, and a text-only request with the same placeholder ids would
otherwise resume from rows built with an image it never sent.
On the same host the two continues now reuse all 68 and 42 tokens instead
of none. The `equal` text matches an uninterrupted temperature-0 reply, and
the `shorter` one writes the trimmed "\n\n" again byte for byte.
The serve harness now expects the `equal` continue after a mid-turn CANCEL
to reuse all of it (`REUSE 20 3 3 3 3 equal`) and to answer like a freshly
started engine. A new case searches the tiny fixture for a reply with text
and then two whitespace tokens, and runs it twice: stopped at the second
whitespace token, so the cache is one token past the trimmed continue, and
one token later, so both whitespace tokens are in the cache and the second
has passed the snapshot point too. Each expects `shorter` resumed at the
trimmed length and the same answer as a fresh engine; the first also expects
no reuse under GLM53_REWIND=0. The vision-serve harness ends an image turn
after one token and sends the same prompt without IMAGE: it must reuse
nothing (without the image rule it resumed all 22 tokens).
The cases fail without the fix: against the parent engine the serve harness
fails at request 20, and with the snapshot restore removed it reuses the
right count but the answer differs, which the fresh-engine comparison
catches. With a snapshot taken at every whitespace token instead of once per
run, the later stop reuses nothing. The pin, multimodal, vision-serve
harnesses and the dashboard and context-exceeded unit tests pass.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
9617138 to
15c4b8c
Compare
|
Thanks, this works: the oracles and serve harnesses pass, also under ASan and UBSan, including the new One change before merging, about the default. The whitespace snapshot costs about 149 MiB per slot outside the About #1792 (image-keyed KV prefix): whichever of the two lands second has to fold in the other. The points are:
|
The whitespace snapshot costs about 149 MiB per slot outside the GLM53_EXPERT_GB budget, plus a full KDA state copy at every generated whitespace run. Every reply has whitespace runs, so with the snapshot on by default every user paid both, including users who never continue a reply. GLM53_REWIND now defaults to 0; set it to 1 to rewind a continue that trims trailing whitespace. Exact-match resumption stays on, since it costs one logits row. The serve harness turns GLM53_REWIND=1 on for the rewind cases and checks that a trimmed continue re-prefills under the default. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Noted on the image-keyed KV prefix PR, including |
Summary
Related to #1748.
Stopping a reply and clicking continue can make the glm53 engine process the entire prompt again, even though it has already cached the work.
The existing reuse path requires at least one new prompt token, hence an exact match misses, and so does a shorter prompt after the client trims trailing whitespace tokens to satisfy the continuation API. (For example,
"# Hummingbird\n\n"cannot be continued by sending"# Hummingbird", whilst"# Hummingbird\n\n"is rejected by the gateway.)This change supports preserving that work for two continuation cases:
equal): retain the final logits row, which predicts the next token at the cached position. The next request can use that row even though it has no new prompt tokens to process.shorter): save the KDA recurrent state, convolution windows, and logits before feeding the first whitespace-only token of a generated run. If the next prompt ends at that saved position and matches the cached prefix, restore it. This handles a client trimming the reply before continuation; the gateway rejects trailing whitespace.Rules for both resume points:
This change does not provide arbitrary cache rewind.
The trimmed whitespace case is opt-in (
GLM53_REWIND=0, documented indocs/ENVIRONMENT.md). It requires a buffer the size of the KDA state, about 149 MiB per slot, plus a logits row, on first use; the buffer is retained across slot resets and reused. The state is copied once at the start of each eligible generated whitespace run, including during ordinary generation.Exact-match resumption stays on, since it costs one logits row.
Validation
make -C c checkmake -C c cuda-test(if applicable) (N/A; no CUDA changes)Serve harness checks (
glm53-oracleCI job and vision serve harness passed locally):CANCELreuses the exact cached prompt (REUSE 20 3 3 3 3 equal) and matches a fresh engine's output.shorter, reuse that prefix, and match a fresh engine's output in both cases. These are deterministic cache-state tests using token budgets; the existingCANCELcase covers cancellation.GLM53_REWIND=0.IMAGE. It must reuse nothing; without the image rule it resumed all 22 tokens.Compatibility