Skip to content

[Feature]: improve glm53 cache reuse after interruption #1748

Description

@enitimeago

Problem

Following the assistant-continuation support in #1402 and the web UI continue action in #1717, I'd like to improve how much completed work the glm53 engine can reuse when continuing or retrying, and how quickly it releases the engine on cancellation. This would significantly improve responsiveness of continuing a manually stopped generation in the web UI.

Specifically, the glm53 engine appears to have an all-or-nothing reuse rule: it reuses the slot only if the new prompt matches every cached token. On a disconnect, the engine keeps decoding for a moment before the cancel lands, and it keeps those extra tokens in the slot. The client never receives them, so a partial sent back is shorter than the cache and nothing is reused.

Measured on v1.12.1 (GLM-5.3-Flash int4, coli serve --gpu none --ram 20 --ctx 32768, one KV slot, about 1 s/token prefill):

Case GLM53_VERBOSE=1 output First token
Reply stopped at max_tokens 60, then Continue REUSE 5 96 97 1.5 s
Client disconnected after 48 tokens, then Continue REUSE 2 0 86 66.9 s
281-token prompt disconnected mid-prefill at 20 s, then identical retry REUSE 7 0 281 206 s

(REUSE <id> <reused> <prompt> respectively: the request id, how many prompt tokens were taken from the slot's KV cache instead of running through the model again, and the prompt's length in tokens. So REUSE 5 96 97 computed one token, and REUSE 2 0 86 computed all 86.)

In the last case, the client disconnected 20 seconds into the request, but /health recorded cancellation only at 211.7 seconds from the request's start. The engine continued processing the prompt while other requests waited.

The existing cache-reuse path requires the new prompt to extend the cached prefix. A continue can instead match the cache exactly, or end before it after the client trims trailing whitespace to satisfy the continuation API. Those cases currently fall back to processing the prompt again. Cancellation is also checked during decoding, but not during prompt processing.

Proposed solution

Allow the glm53 engine to stop prompt processing promptly on cancellation, preserve safely completed work for a retry, and resume supported continue requests from cached state.

I have an implementation split into incremental changes:

  1. feat(glm53): add detailed REUSE information, and log CANCEL #1749: Add verbose diagnostics showing the cache state at cancellation and the reason for each reuse decision. I've opened this PR alongside the issue so the existing behavior and later improvements are easier to inspect.
  2. feat(glm53): handle cancellation during prompt processing #1789: Check cancellation between prompt-processing chunks and retain completed chunks for reuse.
  3. feat(glm53): continue from the cache instead of re-prefilling #1803: Resume from saved logits when a continue exactly matches the cache, and from a recurrent-state snapshot when it matches the start of a trailing run of whitespace-only tokens.
  4. feat(glm53): handle cancellation between layers of a prompt chunk #1822: Check cancellation between layers within a chunk, restoring the state from the chunk's start when interrupted.

The later changes have been exercised locally, and I'd submit them after the initial diagnostics PR is reviewed and merged. I'm open to adjusting the scope or order if a different split would be easier to review.

Alternatives considered

  • Keep a ring of recurrent-state snapshots during decoding, so Continue could rewind past tokens generated after the client stopped reading. This was my first plan, and the diagnostics PR was written partly to size it. In three mid-reply disconnects on the tested host, the cache held exactly the tokens the client had received, so a ring (~149 MiB per snapshot for GLM-5.3-Flash) would pay memory for a case I haven't observed. The diagnostics would show it if it does occur.
  • Reduce the prompt-processing chunk size. Cancellation between chunks waits up to one chunk, so smaller chunks respond sooner: the worst-case wait on the tested host was roughly 92 s at 128, 56 s at 64 and 20 s at 16. But every prompt pays for it, not just cancelled ones: on a 277-token prompt, total processing time rose 22% at 64 and 73% at 16, because experts streamed under --ram 20 are read again for every chunk. Checking between layers brought the wait down to a few seconds at the default chunk size, at the cost of one state copy per chunk.
  • Accept trailing whitespace in continuation requests instead of rewinding. The GLM-5.3-Flash template strips assistant content, so the rendered prompt would still end before the whitespace and the cache would still be one token ahead. Keeping the whitespace would require the prompt to diverge from the reference template, and serve: continue a trailing assistant turn instead of answering in a new one #1402 rejects it with a 400 so that the model doesn't silently resume from different bytes than the client sent.

Scope and compatibility

  • The proposed changes are limited to the glm53 engine, its diagnostics, tests, and documentation. They require no new dependencies, model formats, or client request fields. The default CPU path remains dependency-free.
  • The timings above are from CPU-only runs (--gpu none) with experts streamed from disk; they do not establish performance with Metal or Vulkan offload.
  • Reuse must preserve cache validity. Image requests and scoring requests need separate handling; the implementation does not assume token IDs alone are sufficient for every request.
  • This doesn't provide general rewind support. It covers the exact-match and trailing-whitespace cases; a request that can't safely reuse the cached state still processes the prompt again.
  • Whitespace resumption requires a recurrent-state snapshot, approximately 149 MiB for GLM-5.3-Flash. A new flag GLM53_REWIND is added in the associated draft PR. Cancellation within a chunk also requires saving recurrent state and adds a copy at each chunk's start.
  • The diagnostics PR changes logging only. The behavioral changes would be reviewed in the follow-up PRs.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions