Pi extension that shrinks the thinking in the context.
- in some tasks reasoning takes more than half of the context;
- the extension compresses that part by 30–60%.
Old thinking blocks are rewritten by a background call of the same model into short digests: decisions, rejected options, facts and open questions survive, repetition and filler do not. Recent blocks stay raw.
Footer: 💭 -40k/120k +76k — reasoning tokens saved, raw reasoning tokens in
context, tokens spent by the background calls. Spinner while a rewrite runs,
💭 off when disabled.
/compact-thinking— report;/compact-thinking on/off;/compact-thinking dump— raw and compacted blocks next to the session file.
/compact-thinking dump writes two files next to the session JSONL:
thinking-a— the raw blocks as the model wrote them;thinking-b— the same blocks as the model sees them, with digests in place.
Every block is preceded by the same heading in both files
(## entryId:blockIndex timestamp), so a diff aligns block by block. Only
applied digests that are still in context are dumped.
Unified diff:
diff -u thinking-a thinking-bSide by side:
delta --side-by-side thinking-a thinking-b
diff --side-by-side --width=180 thinking-a thinking-b
vimdiff thinking-a thinking-bdelta is the readable default: side-by-side panes, syntax highlight, an
aligned vertical line between the sides. vimdiff (or meld, code --diff)
works when you want to edit or fold.
The dump notification prints the directory it wrote to; the files sit next to
the session JSONL under ~/.pi/agent/sessions/.
pi install git:github.com/NikolayXHD/pi-compact-thinkingSettings: ~/.pi/agent/compact-thinking.json.
| field | default | meaning |
|---|---|---|
enabled |
true |
mode is on at session start |
k |
5 |
recent blocks that stay raw |
minTokens |
512 |
shorter blocks are left alone |
acceptanceRatio |
0.95 |
a longer digest is rejected |
A digest is a separate call of the same model over the same session prefix, so
the provider reads it from the warm cache: no tools, no reasoning, the answer
wrapped in <digest>…</digest>.
Digests are substituted into the context by one function called in context
(before the main model request) and in session_before_compact (before Pi's
own summarizer). Nothing is written to the session, so Pi keeps deriving the
context size and the compaction threshold from the real provider usage.
The background job runs while the agent is busy with tools and right before the session settles; it never competes with the model stream.
- works while the extension is loaded, nothing is stored in the session;
- every block costs a full prefix read: a warm cache matters;
- does not replace Pi compaction, it shrinks thinking only.
node src/probe.mjs for pure rules, src/TESTING.md for session scenarios.