Skip to content

Repository files navigation

KV Cache Capacity Estimator

This tool analyzes KV cache storage capacity requirements by replaying offline requests.

It can also run as a service that receives a live request stream and estimates KV cache capacity requirements in real time.

Features

1. Offline Token ID Replay

This is the recommended mode. It has minimal dependencies and does not rely on engines such as SGLang or vLLM. By comparison, the OpenAI request replay mode described below may encounter issues during tokenization.

This mode depends only on the Python standard library and does not load a tokenizer, model, or chat template. The input is a JSONL file in which each line may use any of the following formats:

[1, 2, 3, 4]
{"input_ids": [1, 2, 3, 4]}
{"token_ids": [1, 2, 3, 4]}

Run it as follows:

python src/token_ids_replay.py token_ids.jsonl \
  --kv-bytes-per-token 61505 \
  --page-size 64

Both replay modes directly calculate the capacity required for consecutive prefix hits for each request. If a request can consecutively hit H prefix pages with unlimited capacity and the target factor is R, the first ceil(H × R) prefix pages must all be hits. The required capacity is the maximum LRU depth of those pages plus one, multiplied by the number of bytes per page.

By default, the target is 99% of the request's theoretical consecutive prefix-page hit rate under unlimited capacity. The results report the mean, p50, p90, p95, p99, and p999 of the exact capacity requirements.

--target-hit-rate-percent 99

--capacities estimates the hit rate at different capacities. The default values are 100GiB,200GiB,400GiB,800GiB,1TiB,2TiB,4TiB,6TiB,8TiB,12TiB,16TiB,24TiB,32TiB,64TiB. Pass --capacities explicitly to override these defaults. Note that the tool determines a capacity requirement for each request, and these per-request results are independent of the --capacities setting.

--warm-up-percent PERCENT accepts a value from 0 to 100 (or a percentage such as 20%). The first floor(total number of requests × PERCENT / 100) requests are used only to establish the LRU/cache state. They are excluded from the request count, page-access count, all hit-rate metrics, and per-request capacity distribution. For example:

--warm-up-percent 20

2. OpenAI-Compatible Request Replay

The input is an OpenAI Chat/Completions-compatible JSONL file. Both direct request objects and the body wrapper format used by the OpenAI Batch API are supported. The tool provides three tokenization backends, and the resulting token IDs are passed through exactly the same analysis pipeline as in Feature 1.

By default, request_replay.py uses four threads for concurrent tokenization. Use --tokenize-workers N to adjust the number of threads, or set it to 1 to restore serial processing. Even if requests finish out of order across threads, their token IDs are always sent to the simulator in the original JSONL input order, so the LRU access trace remains unchanged.

Local SGLang Backend (No GPU Deployment Required, Recommended)

The sglang backend directly uses SGLang's OpenAI request types, tool-call parser, content-format conversion, chat template, and tokenizer. It reads only the model configuration, tokenizer, and chat template; it neither loads model weights nor starts an inference service.

Note: Run this in an SGLang container environment.

The following command processes the GLM-5.2 example included in this repository. --tool-call-parser glm47 validates the parser against SGLang's parser registry and initializes the corresponding parser for requests that include tools:

python src/request_replay.py \
  test_data/LongBench-v2/data_short.openai.jsonl \
  --tokenize-backend sglang \
  --model test_model/GLM-5.2-FP8 \
  --tool-call-parser glm47 \
  --kv-bytes-per-token 61505 \
  --page-size 64 \
  --capacities 40GiB,80GiB,160GiB,200GiB,250GiB,300GiB,512GiB

For completion prompts, this backend uses the tokenizer loaded by SGLang. For chat messages, it reuses SGLang's content-format normalization and apply_chat_template path. This includes converting arguments in historical assistant tool calls from JSON strings to objects and forwarding chat_template_kwargs such as enable_thinking.

Tested models

  • GLM-5.2: The token IDs match those generated by the SGLang engine. With SGLang v0.5.19, an fp8_e4m3 KV cache data type, and MTP disabled, the logs report a KV cache size of 60 KiB per token.

3. Convert OpenAI Requests to Token IDs

openai_to_token_ids.py reuses the three tokenization backends described above to convert an OpenAI Chat/Completions-compatible JSONL file into an offline token-ID JSONL file. Each output line is an array of token IDs that can be used directly as input to token_ids_replay.py. If --output is omitted, the results are written to standard output.

For example, to use the tokenizer from a local SGLang source tree:

python src/openai_to_token_ids.py \
  test_data/LongBench-v2/data_short.openai.jsonl \
  --tokenize-backend sglang \
  --model test_model/GLM-5.2-FP8 \
  --tool-call-parser glm47 \
  --output token_ids.jsonl

4. Online KV Cache Capacity Analysis

Currently, kv_capacity_estimator is integrated into SGLang to enable online KV cache capacity estimation. The adapted SGLang source code is available at:

https://github.com/llc-kc/sglang/tree/kv_capacity_estimator

Usage

1. Start the SGLang container

Pull the official SGLang image:

docker pull lmsysorg/sglang:v0.5.20-cu130

Start a container based on this image.

2. Install kv_capacity_estimator

Follow the “Installation and Wheel Build” part of this readme to install kv_capacity_estimator.

3. Install kv_capacity_estimator SGLang frontend.

Clone and install the SGLang branch with kv_capacity_estimator support:

git clone -b kv_capacity_estimator https://github.com/llc-kc/sglang.git
cd sglang
pip uninstall -y sglang
pip install -e "python"

4. Start the KV capacity estimation service

The following example starts an SGLang service with KV capacity estimation enabled:

sglang serve models/GLM-5.2-FP8 \
  --model-type llm \
  --tool-call-parser glm47 \
  --tokenizer-worker-num 32 \
  --enable-metrics \
  --host 0.0.0.0 \
  --port 30000 \
  --log-level debug \
  --kv-capacity-estimator \
  --kv-capacity-estimator-config '{
    "kv_bytes_per_token": 61505,
    "target_hit_rate_ratio": 0.99,
    "slide_window_size": 3000
  }' 

The main configuration parameters are:

  • kv_bytes_per_token: KV cache memory consumption per token.
  • target_hit_rate_ratio: Target fraction of theoretical hit rate.
  • slide_window_size: Number of recent requests used for online capacity estimation.

5. Send requests to the service

The service provides an OpenAI-compatible API:

curl -s http://localhost:30000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model": "GLM-5.2", "messages": [{"role": "user", "content": "hello?"}]}'

For more details on the KV capacity estimator SGLang front end, please refer to the documentation:

KV Capacity Estimator documentation

Interpreting Replay Results

Example replay output:

    Capacity        Pages    Page hit   Reuse hit   Prefix hit
--------------------------------------------------------------
      100GiB       27,277     40.95%     46.23%      40.95%
      200GiB       54,555     57.13%     64.50%      57.13%
      300GiB       81,833     67.92%     76.68%      67.92%
      512GiB      139,662     81.56%     92.08%      81.56%
        1TiB      279,324     87.99%     99.34%      87.99%
      1.5TiB      418,987     88.58%    100.00%      88.58%
        2TiB      558,649     88.58%    100.00%      88.58%
        3TiB      837,974     88.58%    100.00%      88.58%
    infinite     infinite     88.58%    100.00%      88.58%

Per-request target capacity (99.00% of theoretical prefix hit rate, samples=1500):
  mean=188.24GiB  p50=103.60GiB  p90=463.88GiB  p95=704.22GiB  p99=1023.81GiB  p999=1.33TiB

Metric definitions:

Page hit   = Number of page-access hits / Total number of page accesses
Reuse hit  = Number of reusable-page hits / All non-cold-start page accesses
Prefix hit = Sum of consecutive page hits from the start of each request / Total number of page accesses

The output provides two complementary views: request hit rates at different capacities and the distribution of per-request capacity requirements. Together, they can be used to determine an appropriate cache capacity.

We have found that the p99 estimate of the capacity required for requests to achieve 99% of their theoretical hit rate is generally a reasonable minimum capacity requirement.

Installation and Wheel Build

The project uses a standard pyproject.toml build configuration and requires Python 3.10 or later.

Install from source:

pip install -e ./

Build a wheel:

python -m pip install build
python -m build --wheel  --no-isolation

The build artifact is written to the dist/ directory and can be installed with:

python -m pip install dist/kv_capacity_estimator-0.2.0-py3-none-any.whl

After installation, the following command-line entry points are available. They correspond to the three original scripts under src/:

kv-capacity-token-replay
kv-capacity-request-replay
kv-capacity-openai-to-token-ids

For example, offline Token ID replay can also be run as follows:

kv-capacity-token-replay requests.jsonl \
  --kv-bytes-per-token 61505 \
  --page-size 64 \
  --capacities 10GiB,100GiB,1TiB

The base wheel depends only on the Python standard library. When using the local sglang or vllm tokenization backend, run the tool in the corresponding engine's Docker environment.

Citation

{
  title={The KV Cache Working Set: Online Capacity Planning for LLM Inference Systems},
  author={Luchang Li, Shuaishuai Wang, Zhao Ruan, Dongfang Li, Bozhao Gong},
  year={2026}
}

About

kv_cache_capacity_estimator

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages