Skip to content

Turso v0.8.1 parity: corpus refresh, 0.8 planner and bindings, MVCC and spill fixes - #75

Merged
Marc-André Moreau (mamoreau-devolutions) merged 27 commits into
masterfrom
claude/ahtola-turso-parity-check-b1919f
Oct 1, 2026
Merged

Marc-André Moreau (mamoreau-devolutions) merged 27 commits into
masterfrom
claude/ahtola-turso-parity-check-b1919f

Conversation

@mamoreau-devolutions

Copy link
Copy Markdown
Contributor

Brings Ahtola up to the Turso v0.8.1 pin (turso-src @ 8549c1659): vendors the v0.8.1 sqltest corpus, closes most of the differences it exposed, ports the 0.8 planner and .NET binding additions, and fixes the MVCC and performance bugs found along the way.

Conformance

  • Vendored the v0.8.1 corpus: 33 new and 41 refreshed sqltest files. This exposed 113 new case-level differences.
  • 80 of those are closed, plus two older generated-column integrity markers. The expected-failures ledger goes from 191 entries to 109, and docs/turso-gap-inventory.json is reconciled.
  • The remaining 33 v0.8.1 entries are documented in docs/turso-remaining-gap-plan.md ("v0.8.1 refresh"):
    • 19 planner/EQP shapes that the evaluator route used by the harness does not execute.
    • 14 deliberate policies: :memory: journal_mode, CLI Parse error: prefix, SQLite schema-text quoting, readonly VACUUM message, file-backed ATTACH, COVERING wording.

Correctness

  • MVCC lost update. A BEGIN CONCURRENT writer could overwrite a peer's update committed after its snapshot, or resurrect a row the peer had deleted. This is now a write-write conflict (Turso 0f7f30eac).
  • MVCC explicit rowids. Concurrent inserts rewrote explicit rowids and INTEGER PRIMARY KEY values. Explicit keys are now kept, and a clash between them is a conflict.
  • SQL:
    • CROSS JOIN … ON/USING and a comma join followed by ON now parse.
    • Correlated subqueries after a FULL JOIN now run.
    • Errors are no longer swallowed by unnested NOT EXISTS.
    • Generated columns are computed after REPLACE/DEFAULT.
    • UPSERT honours the target alias, and BEFORE UPDATE triggers now refresh the row.
    • Internal tables are protected in any letter case.
    • The 2,000-column limit is enforced, and parameter-name tokenizing follows SQLite.
    • PRAGMA count_changes is implemented.
  • Functions: fixes to JSON (NULL labels, jsonb(NULL), json_tree rowid), date/time (BLOB arguments, pre-epoch rounding), min/max ties and generate_series with NULL arguments.
  • Error codes: SqliteException now carries SQLite primary and extended result codes, locally and remotely: 19/2067/1299/275/1555/787, 13, 26, 5, 8, …
  • Storage:
    • max_page_count is enforced on file-backed and :memory: databases, and :memory: gets exact page counts.
    • integrity_check reports STRICT and generated-column type violations.
    • Explicit checkpoints are refused while another statement is active on the connection.
    • WAL readers use read marks 1–4 while a checkpoint holds read mark 0.
    • auto_vacuum reports the mode stored in page 1.
    • MVCC ALTER COLUMN rejects changes that need an index rebuild.

Turso 0.8 features

  • Planner: Turso's cost model is ported:
    • unnesting of correlated comparisons
    • hash anti-joins
    • IN-list and OR-implied IN index seeks
    • multi-index OR with AND branches
    • partial-index OR terms
    • costed compiled joins, including hash builds and materialized prefixes
  • EXPLAIN: EQP describes the executed join tree, and FORMAT=JSON reports real estimates.
  • .NET replica/sync API:
    • 13 connection-string replica keys.
    • Pull, Push, Checkpoint, GetSyncStatistics.
    • Automatic-sync status and a change event, plus an opt-in PullOnly mode. The default still pushes and pulls.
    • An auth-token provider hook.
    • Tolerant replication_index handling.

Performance

Workload (default 2 MB budget) Before After
10k-row spilled GROUP BY ~28 s ~0.5 s
groupby/default.sqltest (whole file) ~15 min ~80 s
3k × 3k spilled equi-join ~32 s ~0.4 s
20k-row count(DISTINCT …) ~16 s ~0.1 s
20k-row self-UNION ~42 s ~0.25 s
  • Sorter spill: writes are coalesced and runs read ahead. Both buffers scale with the memory limit and are charged to it.
  • Hash join spill: probe batches can evict cached partitions to keep growing, groups are answered resident-first, partition loads are pre-sized, and scans read ahead.
  • Deduplication: evaluator DISTINCT, compound set operators, DISTINCT aggregates and recursive UNION share a set that buckets rows by a hash consistent with DISTINCT equality.

Behaviour changes worth reviewing

  • Opening a garbage-header file without a key now raises SQLITE_NOTADB (26), "file is not a database". Previously this was an InvalidDataException.
  • SqliteErrorCode is no longer always 1.
  • VDBE "datatype mismatch" now reports 20, where it used to report Constraint.
  • max_page_count is now per-connection, as in SQLite.

Not ported (product decisions)

  • PostgreSQL frontend
  • general typed values / custom types
  • incremental materialized views
  • replica pooling
  • standalone Serverless/Platform client packages
  • Force Logical MVCC Pull=True
  • per-step batch statistics

Validation

  • Full managed suite on net10.0 at the head commit: 8,679 passed, 1 failed, 19 skipped (skipped by platform or corpus policy).
    • The one failure, SqlitePagerPortableLockCoordinatorTests.DeleteModePagerFaultsWhenPeerSwitchesMainFileToWal, was a cross-process worker hitting its 30 s exit timeout. That class passes 22/22 on rerun.
  • Each slice adds differential tests against bundled SQLite (error codes, DISTINCT hashing, max_page_count, schema/DML, functions) plus focused regressions (MVCC conflicts, spill buffers, planner).
  • ./build.ps1 validate-trim passed on the bindings slice.

🤖 Generated with Claude Code

Refresh conformance/sqlite-sqltests from the pinned turso-src v0.8.1
(8549c1659): 33 new files and 41 refreshed files since v0.8.0-pre.7. The
113 newly exposed case-level differences are recorded in the expected-
failures ledger pending closure; the inventory's live count follows.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A BEGIN CONCURRENT transaction could overwrite a peer's update committed
after its snapshot, or resurrect a row the peer deleted: base-row deletion
markers are separate tombstone versions, invisible to the older snapshot,
and the no-visible-version path only checked in-flight peers. Treat any
begin/end stamp committed after the writer's snapshot as a write-write
conflict (Turso 0f7f30eac).

Concurrent inserts also promoted every rowid through the store-global
allocator, silently rewriting explicit rowids and INTEGER PRIMARY KEY
values after a peer delete. Only rowids the statement allocated itself are
promoted now; an explicit key is kept and a clash with a concurrent
writer's live version is a write-write conflict.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rences

Port the Turso v0.8.1 behavior the refreshed sqltest corpus exposed:

- json_group_object/jsonb_group_object skip rows whose label is SQL NULL
  (517ec809f "json: skip NULL labels in object aggregates").
- jsonb(NULL) returns SQL NULL (5775d5337 "json: preserve SQL NULL in jsonb()").
- json_each/json_tree bind a selectable rowid (qualified, unqualified, and in
  INSERT ... SELECT), numbered from zero like SQLite and core/json/vtab.rs;
  generate_series reports the generated value as its rowid (core/series.rs)
  (16a02b139 "core/translate: bind rowid on virtual tables").
- Date/time functions read a BLOB value or modifier as UTF-8 text
  (b2512eb71 "core/functions: read a BLOB date/time argument as text").
- unixepoch() and strftime('%s') floor to whole seconds, so 1 ms before the
  epoch is -1 (535f684f1 "core/functions: fix datetime rounding").
- Scalar min() keeps the last of tied arguments while max() keeps the first,
  matching SQLite's minmaxFunc (78b1ac633, merged as d81560556).
- generate_series yields no rows when any supplied bound is NULL.
- The simple-count INDEXED BY exemption covers COUNT(col) over a provably
  non-NULL column, as Turso's detect_simple_aggregate does.

The integrity_check/memory.sqltest attached-corrupt-database case stays in the
ledger as an intentional managed-ATTACH policy.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…nts, coerce FTS queries

Error-code fidelity (Microsoft.Data.Sqlite parity):
- EmbeddedSqlException.SqliteErrorCode now yields the (extended) SQLite result
  code for every classic failure class: explicit where the raise site knows it
  (primary-key/rowid conflicts 1555/2579, RAISE 1811, datatype mismatch 20,
  sequence exhaustion 13) and otherwise derived from SQLite's own error text
  (UNIQUE 2067, NOT NULL 1299, CHECK 275, FOREIGN KEY 787, STRICT 3091,
  readonly 8, busy 5, full 13, too big 18, corrupt 11, ...). Wrappers that
  re-raise the same failure keep the wrapped code.
- The SQLite facade reports the primary code in SqliteErrorCode and the
  "SQLite Error N" text, and the extended code in SqliteExtendedErrorCode, on
  the embedded, remote (Hrana error.code names such as
  SQLITE_CONSTRAINT_UNIQUE) and SqliteRemoteException paths.
- A garbage plain header (bad magic, invalid page size, too few usable bytes,
  too-short file) opened without a key or codec reports SQLITE_NOTADB (26)
  "file is not a database" like SQLite and upstream 6480cd77e; keyed/codec
  opens keep the encrypted-or-not-a-database phrase.
- SqliteErrorCodeDifferentialTests pins the codes against e_sqlite3 for
  in-memory, file-backed and MVCC databases.

Explicit checkpoints (upstream a95d01284 / 7ad5d9de4): PRAGMA wal_checkpoint
is refused with SQLITE_BUSY "cannot checkpoint while another statement is
active - SQL statements in progress" while another statement on the same
connection is active or suspended; open blob handles do not block it.

FTS (upstream d485aa6d8): integer, real and blob fts_match/fts_score queries
are searched by their text form on the indexed and scalar paths instead of
raising "requires a text query".

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…stats, sync status, token provider

Close the remaining gaps against turso-src v0.8.1 bindings/dotnet:

- Connection strings (AhtolaConnectionStringBuilder and the SQLite facade,
  so EF Core UseAhtola reaches them): Sync Client Name, Sync Long Poll
  Timeout, Bootstrap If Empty, Partial Bootstrap Prefix/Query, Partial Sync
  Segment Size/Prefetch, Remote Encryption Cipher/Key, Push Operations
  Threshold, Pull Bytes Threshold, Force Logical MVCC Pull (only False is
  supported; True fails closed), Sync Experimental Features (validated,
  no effect) and the Ahtola-specific Automatic Sync Mode. They map onto
  AhtolaReplicaOptions and are rejected on local/direct-remote connections,
  matching Turso's HasAdvancedReplicaOptions.
- Explicit sync operations on AhtolaConnection and SqliteConnection:
  Pull/PullAsync, Push/PushAsync (drains every batch pending at the call),
  Checkpoint/CheckpointAsync and GetSyncStatistics/GetSyncStatisticsAsync,
  reusing the managed replica engine. Pull-only and push-only requests
  coalesce per kind in ManagedReplicaSyncRegistry. Sync results now report
  real main/revert WAL sizes; statistics report pending journal changes.
- Automatic sync: AutomaticSyncStatus and AutomaticSyncStatusChanged
  (ported from TursoAutomaticSyncCoordinator) make background failures
  observable. AhtolaAutomaticSyncMode.PullOnly matches Turso's pull-only
  loop; the default stays PushAndPull.
- AuthTokenProvider hook (connection, facade, replica options) resolved per
  HTTP request, per Hrana WebSocket connect, and per replica sync request.
- Hrana replication_index is treated as opaque per PROTOCOL.md 7.2.6: the
  numeric watermark is still tracked and echoed, other values are ignored
  instead of failing the batch.
- Sync client ids may now carry a validated client-name prefix.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…uilds

PRAGMA auto_vacuum (Turso v0.8.1 translate/pragma.rs without the
experimental autovacuum feature, which Ahtola does not have):
- The read form reports the mode page 1 declares (largest-root-page and
  incremental-vacuum header fields), so an auto-vacuum database SQLite
  created reports FULL (1) / INCREMENTAL (2) instead of a hard-coded 0.
- The set form resolves the mode like upstream's requested_mode: names in any
  case, or a numeric literal valued 0..2 (so 00 and 0x0 are NONE). NONE stays
  accepted; FULL, INCREMENTAL and unrecognized values keep upstream's
  --experimental-autovacuum diagnostic.

ALTER COLUMN in MVCC mode (upstream validate_indexes_can_be_rewritten): a
rewrite that changes stored values (becoming generated, toggling VIRTUAL, an
affinity change) or may change a VIRTUAL generated value, and that touches
an index observing the column or a generated column derived from it, is
rejected with 'cannot ALTER COLUMN "x": rebuilding affected indexes is not
supported in MVCC mode'. The two alter_column *-mvcc corpus cases now fail
only on the CLI "Parse error:" prefix, like the other prefix-only entries;
their ledger summaries are updated accordingly.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Ports the Turso v0.8.1 (turso-src 8549c1659) fixes behind 19 ledger
entries recorded by the v0.8.1 corpus adoption, plus non-corpus parity:

- REPLACE defaults now land before virtual generated columns are
  recomputed (Turso ea7b14a0a), so CHECK, UNIQUE keys, NOT NULL, STRICT
  typing and RETURNING see the substituted value; generated-column NOT NULL
  is enforced after the default pass.
- Invalid generated-column clauses report SQLite's messages (be3c173b7):
  a second generation clause, a clause after DEFAULT, or an unknown type
  token fail with `error in generated column "<col>"`.
- UPSERT honors the INSERT target alias (859c91ade): the alias names the
  target row, hides the base name, and shadows `excluded`.
- UPSERT DO UPDATE keeps columns a BEFORE UPDATE trigger rewrote and
  recomputes generated columns (a99d061a5).
- User INSERT/UPDATE/DELETE on the __turso_internal_ namespace fails with
  `table X may not be modified` regardless of case (67262f427); managed
  backup/snapshot replay is exempt like upstream nested statements.
- ALTER TABLE RENAME onto a view reports `there is already another table
  or index with this name` (799e4737c).
- PRAGMA temp.synchronous reads OFF and ignores writes, like SQLite.
- CREATE TABLE / ADD COLUMN enforce SQLITE_MAX_COLUMN 2000 with SQLite's
  messages (a81efd85b).
- Named parameters tokenize with SQLite's CC_VARALPHA rule: bare `$`, `:`,
  `@` (or `$::`, unclosed `$a(b`) are `unrecognized token`, and TCL names
  `$::v`, `$ns::v`, `$arr(e)` are single parameters (22a0a0f6e, a45cd87ff).
- PRAGMA count_changes uses SQLite's column names, reports zero-row
  statements, leaves RETURNING rows alone, and counts only rows an UPSERT
  inserted.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

# Conflicts:
#	src/Ahtola.Tests/Conformance/managed-sqltest-expected-failures.txt
…grity_check

Turso "report STRICT type violations from integrity_check" parity:
- PRAGMA integrity_check/quick_check now checks every non-rowid-alias column
  of a STRICT table against its declared INT/INTEGER/REAL/TEXT/BLOB storage
  class ("non-TEXT value in t.b"), before its NOT NULL check, using SQLite's
  OP_IsType masks (NULL always passes; REAL accepts the integer storage class
  SQLite writes for integral reals). ANY and domain/custom types are exempt.
- Stored rows' VIRTUAL generated columns are recomputed on read with affinity
  only, like a SQLite/Turso column read, instead of enforcing NOT NULL and the
  STRICT storage class at load. A row that violates them stays readable and
  integrity_check reports it ("NULL value in t.b", "non-INT value in t.b")
  rather than the database failing to open.
- The sqltest harness now builds the four database/integrity_strict_*.db
  fixtures like upstream's generate_strict_fixture (managed write path plus a
  same-length first-row serial-type swap), so the parity_strict_* corpus files
  run instead of being skipped, and all pass.

Ledger: removed the two parity_gencol_not_null_violation entries, which now
pass.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Turso 74eb80c83 / SQLite walTryBeginRead: a reader whose WAL is fully
backfilled normally takes read-mark 0 and ignores the WAL, but a checkpoint
holds read-mark 0 exclusively for as long as it runs. Instead of retrying
until the checkpoint finishes, SqliteWalReadSnapshotCoordinator now pins the
committed boundary with an existing mark already at mxFrame or an idle mark
it advances to mxFrame. Every frame is already in the database file, so the
mark only keeps a later checkpoint from restarting the WAL underneath the
reader; the usual confirmation (mark value, boundary, commit frame) still
applies. An empty WAL has no commit frame to pin and still waits.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…er cases

Correctness:
- Parser: ON/USING join constraints on comma joins and CROSS JOIN
  (`a, b ON ...`, `a CROSS JOIN b USING (...)`, NATURAL CROSS JOIN), matching
  SQLite's joinop/on_using grammar. JoinTableSource carries a Cross flag;
  CROSS JOIN is a join-order barrier (Turso JoinInfo::no_reorder) and never
  flips its hash-build orientation, so its left operand stays the outer loop.
- FULL OUTER JOIN with correlated subqueries referencing the joined tables is
  no longer rejected: Turso v0.8.1 unnests these after the FULL JOIN, and the
  managed evaluator already evaluates them per joined row, null-padded rows
  included.
- The transient equality hash probe no longer prunes rows a SQLite scan would
  still evaluate: an equality is only probeable while no conjunct before it
  can raise. A positive single-table EXISTS keeps pruning, matching SQLite's
  existsToJoin automatic-index semantics. Fixes NOT EXISTS hiding
  'malformed JSON' errors.

Planner/EQP (from the executed plan):
- Partial-index implication accepts either disjunct of an OR predicate
  (Turso query_term_implies_predicate), and a covering partial index can be
  scanned without a key constraint.
- A correlated IN the semi-join rewrite declines is described as its real
  per-row CORRELATED LIST SUBQUERY scan.
- FORMAT=JSON reports the ported Turso scan estimate for an unfiltered table
  scan (sqlite_stat1 row count or the 1,000,000-row default).

Ledger: 17 v0.8.1 entries closed; 5 now-parsing join/memory entries record
their new plan-pattern differences.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

# Conflicts:
#	src/Ahtola.Tests/Conformance/managed-sqltest-expected-failures.txt
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e Turso cost model

Port Turso v0.8.1 f7aac1fc5 ("optimizer: allow unnesting correlated
subqueries with non-equality comparisons"): a correlated EXISTS, NOT
EXISTS or positive IN may now move any direct inner/outer comparison
(=, <>, <, <=, >, >=, IS, IS NOT) into the semi/anti join condition
(unnest.rs is_inner_outer_comparison). The fallible/non-deterministic
inner-WHERE gates are unchanged.

The inner access of a semi/anti join is now chosen by a port of the
v0.8.1 join planner (choose_best_btree_candidate, the temporary-index
alternative of find_best_access_method_for_btree, and the LeftAnti hash
join of join_lhs_and_rhs / try_hash_join_access_method with its
should_not_use_hash_join and probe-index gates), costed with exact ports
of cost.rs / estimate_hash_join_cost (TursoCostModel). Execution and
EXPLAIN QUERY PLAN both consume the same SemiAntiInnerAccess:

- declared index searches (equality prefix plus one range),
- automatic (ephemeral) index searches,
- a new left anti hash join that hashes the outer rows, probes the inner
  table once and emits unmatched build rows in order.

The compiled automatic-index semi-join lowering declines when the cost
model prefers another access, so EQP and execution stay one plan.

Closes unnest-correlated.sqltest correlated-comparison-uses-joins,
correlated-in-inequality-uses-indexed-semi-join,
correlated-in-with-only-inequality-uses-indexed-semi-join and
not-exists-can-use-hash-anti-join.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…R, and semi/anti JSON estimates

- Port Turso v0.8.1 in_filters_implied_by_or_branches (70f75beb7,
  lift_common_subexpressions.rs): an OR whose every branch pins a column
  to copyable literals implies column IN (values); when several columns of
  one table qualify only the sole leading-index column is used, and none
  when a compound index starts with two of them. The planner consumes the
  implied filter while choosing an access path; the statement itself is
  unchanged, so the original OR still filters every returned row.
- TryPlanManagedIndexScan now searches an index by an explicit positive
  IN list or an implied IN filter on its plain ascending leading column
  (choose_best_in_seek_candidate), reported as Turso's (col=?) seek.
- The OR-by-union planner replans compound AND branches with the
  compound-seek analysis and may use one index in several branches
  (MULTI-INDEX OR t (txy, txy)), matching consider_multi_index_union.
- IndexCoversSelect understands IN lists and BETWEEN.
- EXPLAIN QUERY PLAN FORMAT=JSON attaches the ported cost model's
  estimates (input/output rows, access and running cost, including
  cost_with_where_work with Turso's extra_steps) to planned semi/anti
  join nodes; search, multi-index and hash-join ops can now carry
  estimates.

Closes multi_index_or_compound.sqltest
compound-or-keeps-correlated-index-columns-together,
nested-or-branches-use-one-index-search,
nested-and-branches-use-one-index-search and
turso/explain-query-plan-json.sqltest exists-is-reported-as-a-semi-join,
not-exists-is-reported-as-an-anti-join.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
PRAGMA page_count now reports the real size for every database, and a
connection's PRAGMA max_page_count ceiling is enforced with SQLite's
SQLITE_FULL ("database or disk is full", result code 13) and its
transaction semantics.

- :memory: databases keep rows as heap objects, so their page count is
  measured by serializing the catalog through the same page builders a
  persisted image (VACUUM INTO / full rewrite) uses, via a private
  in-memory page planner store. Per-tree counts are cached by row-store
  lineage/revision; page_count is a high-water mark with freelist, and
  VACUUM compacts it. Values match bundled SQLite for the tested shapes.
- File-backed databases: SqlitePager refuses a write transaction whose
  target size exceeds both the ceiling and the committed size (Turso
  Pager::allocate_page), before any frame is written. Explicit
  transactions plan the persist (pager PlanCommitsOnly probe) after each
  statement so the statement fails like SQLite's, not the COMMIT.
- max_page_count is per connection, never persisted, clamps to the
  connection-visible page count, and treats 0/negative as a query.
- In an explicit transaction a single-row write without a statement
  journal rolls back the whole transaction (sqlite3VdbeHalt); any other
  statement only rolls back itself.
- MVCC: autocommit writes (which build pages at commit) and checkpoints
  are held to the ceiling; MVCC transaction COMMITs are not (Turso).
- Removes the sql-error-message-prefix database-full ledger entry.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

# Conflicts:
#	src/Ahtola.Core/EmbeddedDatabase.cs
#	src/Ahtola.Tests/Conformance/managed-sqltest-expected-failures.txt
…he executed join tree

Port the per-step costing of optimizer/join.rs join_lhs_and_rhs and access_method.rs
(find_best_access_method_for_btree, the temporary-index fallback, try_hash_join_access_method,
estimate_hash_join_cost, cost_with_where_work) into JoinOrderEnumerator.EvaluateTursoStep:
declared index seeks, INTEGER PRIMARY KEY lookups, the ephemeral index over a full scan,
left-deep hash joins that build the last placed table (materialized behind a longer prefix up to
MAX_MATERIALIZED_BUILD_ROWS), outer-join Poisson output estimates and the cross-product penalty.
Two-table LEFT/FULL base-table joins are planned the same way, and VdbeHashJoinRuntime now
keeps unmatched rows for either build side so LEFT/FULL can hash their preserved input.

EXPLAIN QUERY PLAN for a compiled join now walks the OpenJoinCursor plan it will execute and
renders Turso's EqpDetail rows (SCAN / SEARCH ... USING [COVERING] INDEX (a=? AND b=?) /
SEARCH ... USING INTEGER PRIMARY KEY / HASH JOIN / MATERIALIZE hash build input / LEFT-JOIN /
USE SORTER FOR ORDER BY). Shapes Turso would not produce for a table with INDEXED BY or
NOT INDEXED (temporary index, hash join) keep the previous placeholder description.

Unit tests that pinned the old single-row EQP, the automatic-index choice or the old build-side
heuristic now assert the plan the Turso cost model selects.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…nd seek join indexes on constant key prefixes

FORMAT=JSON nodes now carry Turso v0.8.1's plan_estimate figures (optimizer/mod.rs) wherever
the described access path is one the ported cost model prices exactly:
- filtered full-table scans (estimate_scan_cost + cost_with_where_work, with
  constraint_output_multipliers and the closed-range factor), declining for BETWEEN/LIKE/IN,
  subqueries, OR-implied IN filters, or a constraint Turso could seek instead;
- single-table index searches and ordered index scans (estimate_cost_for_scan_or_seek, with the
  is_index_ordered exemption from the non-covering full-scan penalty);
- the evaluator's LEFT JOIN index plan (the null-supplying seek priced by the same inner-access
  planner as semi/anti joins, rows_after_join's outer-join Poisson term);
- multi-index OR unions (multi_index.rs estimate_multi_index_scan_cost with rowid-only branches);
- uncorrelated IN-subquery seeks (choose_best_in_seek_candidate, list rows capped at
  sqrt(rows)) and the correlated scalar subquery seek priced per outer row;
- compound-select arms, FROM-clause coroutine subqueries (scan plus body cost), and recursive
  CTE scans (1,000,000-row fallback) with their one-row recursive input.

Join planning: a `column = literal` term on a join member can now extend an index seek key
(constraints.rs constant constraints) when the literal needs no affinity conversion and the
column's collation matches the index; the compiled index seek uses the literal verbatim.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…us in AGENTS.md

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…oin costing, EQP estimates

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

# Conflicts:
#	src/Ahtola.Tests/Conformance/managed-sqltest-expected-failures.txt
Record what the v0.8.1 refresh delivered, the 33 remaining v0.8.1
differences (evaluator-route planner shapes and documented policies), what
stays unported, and known follow-ups. The inventory's live expected-failure
count is now 109.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A GROUP BY whose sort spills under the default 2 MB execution budget took
25-40 s at 10k rows: the spill codec issued two or three unbuffered file
operations per value written or read. The sorter's spill file now coalesces
writes in a buffer that absorbs the codec's reserved length headers, and
each run reader reads ahead in blocks. Both buffers scale with the memory
limit (none below 512 bytes, so tiny budgets keep their exact accounting)
and are charged to the spill and merge infrastructure.

The 10k-row GROUP BY drops from ~28 s to ~0.5 s; groupby/default.sqltest
from ~15 minutes to ~80 s, so it no longer hits the 30 s per-case cap.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A 3,000 x 3,000 equi-join whose build side spills under the default 2 MB
budget took ~32 s:

- Resident build partitions filled the budget, so probe-batch admission
  (which keeps a quarter of the budget free) never grew a batch past one
  probe, and each probe reloaded its partition. Batch admission now evicts
  cached partitions to keep growing, and each batch answers groups whose
  partition is already resident first, so cyclic batches stop evicting a
  partition just before reusing it.
- About half of the loads ran out of memory part-way and restarted after
  each eviction. Partitions now record their entries' retention as they are
  written; a load makes room for that lower bound first and skips a load
  that cannot fit.
- Every partition read went through the spill codec one value at a time.
  Sequential partition scans now read through a transient read-ahead block
  charged to the statement's memory.

The join drops to ~0.4 s (10,000 x 10,000: ~2.4 s); output keeps probe
order.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ggregates

Evaluator DISTINCT, UNION/INTERSECT/EXCEPT, DISTINCT aggregates and
recursive-CTE UNION compared every row with every kept row, whatever the
memory budget: count(DISTINCT pad) over 20k rows took ~16 s and a
self-UNION ~42 s. They now share a DistinctRowSet that buckets rows by a
hash consistent with DISTINCT equality (integer/real numerically, BINARY,
NOCASE and RTRIM text by their built-in rules; any other collation hashes
text as a constant, so it stays correct). RowsEqual still decides equality
within a bucket; UNION keeps the later row of an equal group as the pinned
corpus expects. The same queries now take ~0.1-0.25 s.

Differential tests against bundled SQLite cover value classes, collations,
compound operators, DISTINCT aggregates and recursive UNION.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@mamoreau-devolutions
Marc-André Moreau (mamoreau-devolutions) merged commit 2e78ba3 into master Oct 1, 2026
11 checks passed
@mamoreau-devolutions
Marc-André Moreau (mamoreau-devolutions) deleted the claude/ahtola-turso-parity-check-b1919f branch October 1, 2026 20:18
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Development

Successfully merging this pull request may close these issues.

1 participant