Repository navigation
Conversation
809086d to
f757dab
Compare
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
f757dab to
a614c0a
Compare
cevheri
left a comment
There was a problem hiding this comment.
Thanks @koraysrn, the child-process approach fits D227, and timers do keep firing while a long on-disk statement runs. I ran the worker under Node 24.14 and hit the problems below, so changes are requested.
-
Queued requests never settle after a kill.
Cause:failAllinsqlite-worker.tsemptiesqueuewithout rejecting its entries. Cancelling q2 while q1 ran left q2 pending.
Change: reject every queued request infailAll.
Done when: a test cancels a running statement and the queued one rejects. -
The read-only path breaks after one deadline kill.
Cause:enforceQueryOnlyinsqlite.tsusesthis.worker!instead ofgetHandle(), andsend()rejects on!alivewithout resettingbusy, so the next call fails and the one after hangs.
Change: go throughgetHandle()and resetbusyon that rejection.
Done when: two read-only statements succeed after a deadline kill. -
Concurrent restarts leak children.
Cause:getHandlehas no single-flight guard; three queries after a cancel started three children and two outliveddisconnect().
Change: share one pending start.
Done when: parallel queries after a cancel leave exactly one child, and none afterdisconnect(). -
Large results are about 130x slower.
Cause: the parent appends each chunk and scans the whole buffer, and BLOBs cross as JSON number arrays. 100k rows took 32 s against 245 ms in-process, on the parent's event loop.
Change: scan only the new chunk, and send BLOBs compactly (base64).
Done when: 100k rows come back in roughly in-process time. -
Test the worker instead of excluding it.
Cause:tests/setup.tsturns it off andmerge-lcov.mjsdrops it from coverage, which is how 1 to 3 passed CI. The 100% gate cannot be narrowed for this.
Change: test the worker under Node, wire the editor query timeout to the kill (D227 asks for it), and add the/api/healthtest D227 names.
Done when: the exclusions are gone and coverage holds at 100%. -
Copy and docs still describe the old blocking.
Change: updateConsentCard.tsx,AgentRail.tsx,docs/AGENT.mdanddocs/MCP.md. Also give the child a minimalenv, not the parent's secrets.
Done when: nothing says SQLite statements are not interrupted.
5e118c8 to
afaeadc
Compare
|
All six review points are addressed, and the worker is now exercised rather than excluded.
Also fixed along the way: a decimal 64-bit parameter is converted back to a BigInt tag before the wire so a re-read INTEGER still matches, a failed connect releases the file handle before rethrowing, and the in-process and worker paths now agree on null columns. The tests whose subject is the query_only boundary or the provider factory, not the child transport, run the synchronous in-process driver, because Bun 1.4.x on Windows can drop a stdin write to the node:sqlite child under repeated |
…, A1) Run on-disk SQLite statements in a child process the provider can SIGKILL, so cancelQuery and the read-only deadline become preemptive (D227, A1). :memory: stays synchronous because a child process cannot share it. A LIBREDB_SQLITE_WORKER escape hatch restores the synchronous driver, used by the test suite because Bun 1.4.x hangs a SQLite child process a bun test run spawned.
afaeadc to
aade569
Compare
cevheri
left a comment
There was a problem hiding this comment.
Thanks @koraysrn, the six asks from the last round hold: I ran cancel, restart, the read-only deadline and the editor timeout under Node 24 and through the routes, and /api/health answered in under 20 ms during a long statement.
One new problem blocks the merge, and two smaller ones are cheap to fix in the same push.
-
Non-ASCII text is corrupted across the pipe, in both directions.
Cause:buffer += chunkin the parent andstdinBuf += chunkin the child decode every 64 KB chunk on its own, so a multibyte character split across two chunks turns into U+FFFD.
Measured: a SELECT returned 14 of 50,000 rows changed, and a 200 KB Turkish and CJK value was stored in the file with 11 replacement characters, as a literal and as a bound parameter. With the worker off, both are intact.
Change:child.stdout.setEncoding("utf8")in the parent andprocess.stdin.setEncoding("utf8")in the child; with those two lines both probes came back clean.
Done when: a test writes and reads back more than 64 KB of multibyte text through the worker unchanged. -
Cancelling a queued statement kills the running one.
Change: remove a queued id from the queue, and terminate only when it is the statement running.
Done when: cancelling a queued id leaves the running statement alone. -
disconnect()waits for a running statement, becauseclosequeues behind it.
Change: terminate when a statement is in flight.
Done when: disconnect during a long statement returns at once.
Summary
Run on-disk SQLite statements in a child process the provider can
SIGKILL, socancelQueryand the read-only deadline become preemptive. This closes D227 and A1 fromdocs/BACKLOG.md.:memory:stays on the synchronous driver because a child process cannot share an in-memory database.Problem
bun:sqliteandnode:sqliteare both synchronous and expose neithersqlite3_interruptnor a progress handler. A long statement therefore ran on the Studio server's only JavaScript thread and blocked every other request:/api/healthfrom answering for 69.7 s instead of 6 ms.cancelQuerydid not exist for SQLite (supportsQueryCancel: false).statementTimeoutMswas checked only after the statement returned, so an overrunning statement was never preempted (A1).Solution
A new
src/lib/db/providers/sql/sqlite-worker.tsisolates on-disk statements:spawn(process.execPath, ["-e", <bootstrap>])and talks JSON lines over stdin/stdout. No script file is needed, so nothing has to survive the production payload pruning inscripts/lib/prune-standalone-payload.sh.__libredb_sqlite_bigint, BLOBs as__libredb_sqlite_bytes. The parent resolves both throughnormalizeSQLiteBigIntandBuffer, so the value logic stays in exactly one place.terminate()ends the child's stdin and kills it withSIGKILL, which is what makescancelQueryand the read-only deadline preemptive.LIBREDB_SQLITE_WORKER_TIMEOUT_MSoptionally bounds every request with a watchdog that kills the child if no answer arrives.SQLiteProvidernow routes every statement through a smallSQLiteHandleabstraction::memory:uses the synchronous driver as before.cancelQuery(queryId)is implemented and answersfalsefor an unknown or already finished id.queryReadOnlyapplies its time budget preemptively viarunWithStatementTimeout.getCapabilities()reportssupportsQueryCancel: trueandblocksServerWhileRunning: falsefor on-disk connections.Escape hatch
Bun 1.4.x hangs a SQLite child process that a
bun testrun spawned (verified on Windows and Linux). The test suite therefore setsLIBREDB_SQLITE_WORKER=0intests/setup.tsso on-disk tests exercise the synchronous driver. Node (production) never sees that variable and keeps the worker on. With the worker off, on-disk statements run on the server thread again,supportsQueryCancelreadsfalse, andblocksServerWhileRunningreadstrue.Verification
bun tests/run-tests.ts tests/integration/db/sqlite-provider.test.ts: 258 passed, 0 failed (6 skipped are POSIX file-mode tests unavailable on NTFS).tests/unit/lib/db/sqlite-worker-boundary.test.ts: 8 passed, pinning the serialize/deserialize round trip (64-bit integers, BLOBs, nesting, lookalike tags).bun run typecheckpasses.node:sqliteand the worker on, outsidebun test:Files changed
src/lib/db/providers/sql/sqlite-worker.ts(new): child-process transport, wire tags, serialized queue, watchdog, SIGKILL terminate.src/lib/db/providers/sql/sqlite.ts:SQLiteHandleabstraction,cancelQuery, preemptive read-only deadline,shouldUseWorker()degrade gate, handle-based reads for schema/monitoring,buildDbstatSizesshared between the two paths.tests/setup.ts:LIBREDB_SQLITE_WORKER ??= "0"sobun testruns the synchronous fallback.tests/unit/lib/db/sqlite-worker-boundary.test.ts(new): boundary round-trip tests.docs/providers/sqlite.md: section 3.4 rewritten for cancellation via the worker and the escape hatch.docs/BACKLOG.md: D227 and A1 removed (work landed).Notes and known limits
node:sqlite.:memory:remains synchronous and therefore still blocks the server during a long statement, documented indocs/providers/sqlite.md.