fix(driver-sql): reclaimSpace() returns the freed bytes from the SQLite -wal sidecar too, never waiting on another connection (#20426) - #20463
Conversation
… sidecar too, never waiting on a reader Co-authored-by: Claude <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N8TPEsoJxPsdSdNKGnNGEN
…al sidecar Co-authored-by: Claude <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N8TPEsoJxPsdSdNKGnNGEN
…for its follow-up call Co-authored-by: Claude <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N8TPEsoJxPsdSdNKGnNGEN
…se file plus its -wal sidecar Co-authored-by: Claude <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N8TPEsoJxPsdSdNKGnNGEN
…tion back before cleanup destroys it Co-authored-by: Claude <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N8TPEsoJxPsdSdNKGnNGEN
…20106 entry drops its now-false WAL sentence Co-authored-by: Claude <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N8TPEsoJxPsdSdNKGnNGEN
Co-authored-by: Claude <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N8TPEsoJxPsdSdNKGnNGEN
📓 Docs Drift CheckThis PR changes 1 package(s): 10 hand-written doc(s) NAME something this change touched and may need an implementation-accuracy re-verification:
⛔ 2 release-owned page(s) also name something this change touched. These are read-only:
What this run could not see
Coarse fallback — 11 page(s) merely mention a changed package (the pre-#9192 predicate, kept for the deliberately-wide backstop): Which tree this was computed onThis run read A worktree cut from an older # while this PR is open — GitHub drops the merge commit once it closes
git fetch origin 985984b9d1cf386b7c1c523cb1533d163cbfb81a && git checkout 985984b9d1cf386b7c1c523cb1533d163cbfb81a
# afterwards, rebuild it from the two parents, which stay fetchable
git fetch origin bea6d2ea3b355265e78da1a1fa131842c1e15fca effb34a8a32591d1360386c007eed7611e3222e7 && git checkout -B drift-repro bea6d2ea3b355265e78da1a1fa131842c1e15fca && git merge --no-ff effb34a8a32591d1360386c007eed7611e3222e7
node scripts/docs-audit/affected-docs.mjs --json bea6d2ea3b355265e78da1a1fa131842c1e15fca
|
Contract reviewServed-tier: Inputs, and nothing else: card #20426 (body and all four comments: triage 5868711622, claim 5871282236, os-dev-report 5873136055, seat answer 5873177721); PR #20463 (body, file list of 5 files at +203 / −10, and the net diff against the merge base ① Derived judgmentsContract:
The DELIBERATE CORRECTION — ② Semver level
Changeset sentences (20426): heading — TRUE of the checkpoint and of a reader, overbroad against a held write lock (①.8). Paragraph 1 (the WAL default; the whole freelist returned but the bytes left in the sidecar; the 25,754-page readings; the sweep listing the datasource as reclaimed) — TRUE: card #20426 and H1, and 103,149,568 + 4,255,992 = 107,405,560 checks. Paragraph 2 (chunks of a quarter of the page cache, 1,000 at the default; a PR-body sentences. "What was wrong" — TRUE. "What changed": the four helper steps, ③ Boundary flagsDev deviations (report 5873136055): the file surface widened by
Acceptance-note carriers:
Check-runs on Implemented-by: VERDICT: PASS |
Fixes #20426
Clause-②: no
reclaimSpace()on better-sqlite3 now returns the freed bytes from the-walsidecar as well as the freelist, and it never waits on another connection. Every size below is the database file plus its-walfile, read from the file system while the driver is still open. Every freelist and page count is read from a second connection. Measured head:effb34a8a(the branch after mergingorigin/mainatb28550818, which carries PR #20427).What was wrong
With PR #20425, one
Database.exec('PRAGMA incremental_vacuum')returns the whole freelist in one transaction. In WAL mode, the file-backed default, that transaction's dirty pages outgrow the page cache, so SQLite spills them into the WAL before the commit truncates them away. Nothing afterwards truncates the WAL, so the sidecar keeps its high-water size until the last connection closes.What changed
packages/drivers/driver-sql/src/sql-driver.ts: the better-sqlite3 arm ofSqlDriver.reclaimSpacecalls a module-localreclaimBetterSqlite3(connection). It is module-local, likeformatDuplicateGroups, becauseSqlDriver's.d.tscarries its non-public members and this helper is no entry point. The published types are unchanged; the.d.tsgains one doc-comment sentence onreclaimSpace. The helper:PRAGMA freelist_count, and sends nothing more when it is0;PRAGMA incremental_vacuum(N)in chunks, N being a quarter of this connection's page cache (1,000 pages at better-sqlite3's defaultcache_size = -16000and 4 KiB pages), with aPASSIVEcheckpoint after each chunk;auto_vacuum = NONEfile never shrinks its freelist);PRAGMA wal_checkpoint(TRUNCATE)under a busy timeout of0, and puts the connection's own busy timeout back in afinally.Every statement goes through the binding's
exec()/pragma(), which step to completion. The loop is synchronous, so nothing else runs on the connection between chunks. Every other SQLite client stays onknex.raw, as before.driver-sqlanddriver-turso(below), and.changeset/20426-reclaim-space-wal-sidecar.md(@objectstack/driver-sql:patch)..changeset/20106-reclaim-space-full-freelist.md: one paragraph removed. It said the freed pages pass through the-walfile, "which keeps its size until the last connection closes". This PR makes that false, and that note is still pending release. This keepscheck-empty-changesetred on purpose — see "The one red gate" below.The dispatch's hypotheses
H1 — confirmed on
origin/main8cdbe0c6e, throughSqlDriver(25,754 free pages):-walreclaimSpace()(351 ms)disconnect()The DELETE-journal control on the same tree: 105,631,744 → 16,384 while open, with no
-walfile.H2 — re-measured on this tree, and the picked variant is a fourth one. Each variant ran on
SqlDriver's own pooled connection after the real fill-and-delete path (25,754 free pages, chunk 1,000). The rows show database file +-walafter the call, driver open. This is one run per cell on a shared box, so read the ratios, not the absolute times.execalone (PR fix(driver-sql, driver-turso): reclaimSpace() returns the whole SQLite freelist, not one page per call (#20106) #20425)wal_checkpoint(TRUNCATE)PASSIVEPASSIVE+TRUNCATEat busy timeout 0 (this PR)reclaimSpace()frees ONE freelist page per call, not the freelist —PRAGMA incremental_vacuummeasured 300 → 299 pages on SQLite, so the lifecycle sweep never returns bulk-deleted space (ADR-0057 §3.4) #20106 reading of about 210 KB for chunked +PASSIVEdoes not hold throughSqlDriver.PASSIVEnever shrinks the sidecar: it stays at whatever high-water size the sweep's own deletes left (4,255,992 here). Only aTRUNCATEcheckpoint returns it.TRUNCATEcheckpoint blocks the whole process on this synchronous binding, for up to the connection's busy timeout (5,000 ms; knex's better-sqlite3 client passes notimeout, so it is always better-sqlite3's default). The lifecycle sweep runs in the server process, so the triage's never-wait direction holds.TRUNCATEcheckpoint that cannot wait. It is the only row that both returns the space with no reader and never waits with one.The chunk size, and why. A chunk that outgrows the page cache spills its pages into the WAL, just as one statement does. Frames left in the WAL by the call, with a reader pinning every frame so none is reused:
-16000)-2000)PRAGMA cache_spillreads 3,871), and between 250 and 500 at-2000.cache_sizeandpage_size, and the quarter leaves room for the per-page overhead and the b-tree pages each chunk rewrites. At the default that is 1,000 pages (4 MB).H3 — confirmed.
resolveSqliteJournalMode()answerswalfor a file-backed database unless configured otherwise, and the probe's second connection readsjournal_mode = wal. The DELETE-journal control is unchanged by the fix. Before and after, the file shrinks while the driver is open and no-walfile exists: 105,631,744 → 16,384, 216 ms before and 127 ms after.H4 — confirmed. The local
TursoDriverface uses knex'sbetter-sqlite3client, so it takes this arm throughsuper.reclaimSpace(). Its suite reached the method, but it read only the freelist and the page count. It now has a WAL-size case. The remote route is untouched.H5 — nothing new is thrown, so the sweep logs nothing new. Both checkpoints report "busy" as a result row, not as an error. So a busy checkpoint degrades to "vacuumed, not checkpointed": the call resolves, the pages are off the freelist, and
LifecycleService.sweep()lists the datasource as reclaimed, as before.BEGIN IMMEDIATE) makes the vacuum statement itself wait out the busy timeout and throwSQLITE_BUSY. Measured:main5,021 ms and this PR 5,014 ms, both freelist unchanged, busy timeout 5,000 afterwards.space reclaim on datasource 'X' failed (database is locked)) and does not list the datasource.The fix through
SqlDriverSame 25,754-page fixture:
-walafter the call-walTests
driver-sql/src/sql-driver-sqlite-reclaim-space.test.ts, 7 cases (4 before). Each size is the database file plus the-walfile, read while the driver is open. Each freelist and page count is read from a second connection.{ file: pages × 4096, wal: 0 }while open and again after close.{ file: pages × 4096, wal: 0 }.auto_vacuum = NONEcontrol:-walis no larger than before.{ file: pages × 4096, wal: 0 }while open.driver-turso/src/turso-remote-inherited-members.test.ts: new case "local face: in WAL mode the freed bytes leave the -wal sidecar too, while the driver is still open".Suites on the merged head
effb34a8a, all throughscripts/pm/os-verify-lock.sh, each exit code recorded:pnpm --filter @objectstack/driver-sql test: exit 0, 195 files passed and 11 skipped; 3,179 tests passed and 178 skipped. The count before the merge was 3,227; the merge brought in PR fix(security)!: the RLS write check refuses a field-to-field comparison the read refuses — one comparison class, one answer per policy (#20355) #20427, which removed tests of its own.pnpm --filter @objectstack/driver-turso test: exit 0, 74 files; 1,982 passed and 16 skipped.typecheckfordriver-sqlanddriver-turso: exit 0 each.tsc --listFilesOnlyshows both changed test files are in each package's program.Ablations
Every leg ran on the committed state through
scripts/ablation-replace.mjs. In each, the anchor went from 1 hit to 0, and the restore was proven blob-equal to HEAD with an emptygit diff HEAD. Thedriver-sqlsuite imports./sql-driver.js(source), so those legs needed no build.TRUNCATEcheckpoint removed{16,384 + 1,334,912}vs{16,384 + 0}; the reader case's follow-up{16,384 + 296,672}; the NONE control's pair grew 1,318,384 → 2,555,376. 4 green.TRUNCATEdriver-sql'sdist/, whichdriver-tursoresolvesablation-dist-preflightfound the marker in 2 built files.driver-turso: 1 red (local face{32,768 + 280,192}vs{32,768 + 0}), 82 green. After the restore and a rebuild,--absentfound the marker in none of the 6 built files, and the tree was clean.In the first B–E runs, the red reader case also timed out its cleanup hook: the failed assertion left the reader's transaction open. The fixed case rolls the transaction back first. A rerun of leg B went red in 91 ms with no hook timeout.
Gates
node scripts/pm/dispatch-gates.mjs --commands --repo objectstack-ai/objectstackateffb34a8aderived 63 commands. All 63 ran, and every exit code was recorded before any pipe. 62 exited 0;check-empty-changeset --base origin/mainexited 1 (next section).--ran: 63 derived, 63 run, 0 NOT-MEASURED, 0 UNRUN.--ranpass printed a STALE TREE warning:origin/mainmoved 6 commits after the merge, andscripts/cross-package-test-inputs.mjschanged in that range. Of those 6 commits, only PR fix(driver-turso)!: new TursoDriver refuses syncUrl under a forced remote mode, and sync with no syncUrl (#20200) #20447 touches a driver: it changes thedriver-tursoconstructor, and none of this PR's files. CI reads the merge ref.check:driver-conformance: 50 covered, 0 DEBT, 0 exempt, both before (8cdbe0c6e) and after (effb34a8a).pnpm lintis CI's run. The narrowed run:eslint --no-inline-config --format jsonover the 3 changed.tsfiles reports 3 files, 0 errors and 0 warnings.ESLint.isPathIgnoredanswers false for each, so all three are inpnpm lint's population.eslint.config.mjssets noparserOptions.projectand no typed rule, so this diff cannot move the verdict of an untouched file.The one red gate:
check-empty-changeset(a deliberate correction, for confirmation)This PR edits
.changeset/20106-reclaim-space-full-freelist.md, which exists on the merge base. The gate refuses that by name, and its own text sets out two classes. This is the deliberate correction class, not a collision.-walfile "keeps its size until the last connection closes". After this PR it is truncated at the end of the call unless another connection is reading.skip-changesetis not applied and must not be: this PR publishes apatch.For the seat: please confirm, or choose the other route. The other route is to restore the 20106 file from the merge base. The gate then goes green, but the release would carry that sentence beside this PR's own changeset, which describes the new behaviour.
Acceptance notes
.changeset/20426-*.md, and this PR also edits.changeset/20106-reclaim-space-full-freelist.md(one paragraph removed). It is the same defect, a mechanical removal, a card that has already landed, and the same changeset gate family.LifecycleService.sweep()still lists the datasource as reclaimed. The next sweep that deletes rows returns it, and SQLite's auto-checkpoint or the last close returns it sooner. No producer is left worse off than onmain, where the same reader left 197,580,000 bytes instead of 107,405,560.TursoDriverroute,SqliteWasmDriver,LifecycleServiceandpackages/specare untouched.Generated by Claude Code