Repository navigation
feat(db): asynchronous vector and text index publication - #1156
Merged
Merged
Conversation
|
Preview deployment for your docs. Learn more about Mintlify Previews.
💡 Tip: Enable Automations to automatically generate PRs for you. |
xav-db
marked this pull request as draft
October 1, 2026 14:54
xav-db
force-pushed
the
async-index-queue
branch
from
October 3, 2026 15:13
747e31f to
f02e63a
Compare
xav-db
marked this pull request as ready for review
October 3, 2026 16:13
One merge-backed value per (scope, logical index, generation) holds every committed but unpublished vector or text operation (RecordKind 0x14). Producers append blind insert operands and publication removes exact operation IDs; operands, partial merges and resolved values share one canonical encoding. A row-per-operation layout (RecordKind 0x15) is kept for test and benchmark builds as a layout baseline. Cleanup and cursor validation treat both key kinds as outside their lanes.
…re errors The queue store reads and writes generation queues in the map layout, or one row per operation in the row layout that test and benchmark builds can select. The admission ledger charges every committed but unacknowledged operation against per-index retained-byte and pending-member limits, and writers rebuild it from durable queues before open returns; readers only check the layout. A write that would exceed a limit fails before commit with the retryable index_backpressure code (HTTP 429, gRPC RESOURCE_EXHAUSTED); one that exceeds a limit on its own fails with index_operation_batch_too_large. Nothing enqueues yet: the producer and publisher follow.
xav-db
force-pushed
the
async-index-queue
branch
from
October 3, 2026 18:11
f02e63a to
e4f159a
Compare
Graph writes no longer update vector or text indexes in their own transaction. Each write builds one complete immutable operation per changed entity and touched generation, reserves admission capacity (1 GB / 250K pending members per index by default) and enqueues it with the graph change. A publisher owned by the index worker applies queued work in bounded batches under generation ownership and acknowledges exact operation IDs in the same transaction. It defers Building generations and discards retired ones. Vector entities are planned by the build's plan_and_apply, now generic over a Build or Publication target and sharing vector_build_cache_bytes, and commit through the vector cache fence; text publishes through the Active text epoch path, with a per-document admission limit. Searches overlay unpublished work. Strong searches (the default, and every write transaction) see all committed and local changes and fail with a retryable 429 past 800 superseded results; read requests may choose eventual consistency, which overlays the oldest 800 selected entities and never fails. Builds scan source rows only: writes made during a build queue for the hidden generation and publish after activation, so build deltas, the catch-up stage, ActiveVectorMutationRuntime and the synchronous text runtime are removed. HelixDB reports queue stats and exact-ID publication lag. The scoped search bench and gate retry backpressure and wait for the queue to drain before measuring. Search dispatch boxes the overlaid search futures, keeping them out of the frames of recursive callers such as count cursors.
Publication keeps its VectorBuildSession across attempts in the VectorBuildCache and max-min budget builds use, instead of starting every attempt cold. Retained sessions are keyed by owner: a build operation or a publication target (scope, index, generation). An attempt checks out its target's session only at the exact checkpoint of the target's latest commit (target, index record revision, publisher sequence). Only a successful commit whose session holds no unflushed rows retains it; every other outcome forgets it. Leases carry their own demand, builds and publication targets have separate retention caps of 16, and drained targets keep only budget the fair split leaves free. Warm, cold and built graphs agree over inserts, updates, deletes, tenant moves and partition reclaim, including warm sessions that evict, rebind and resume under a 24 KiB budget. A warm insert reads 8 storage keys against about 200 cold.
Behind the async-index-benchmark feature the server opens as an explicit writer or reader with a selectable queue layout, through the same disk cache claim and storage configuration as the product runner, and prints one cumulative JSON sample per interval: queue backlog, publication lag, merge cost, object-store connector I/O and SlateDB metrics. Readers open with HelixDB::open_reader_for_server so only the server's transport reports their queries. docker-image/build.sh --async-index-benchmark builds the image, and CI lints and tests the feature.
The DB production-coverage gate counts only production integration targets, so queue paths reached only by lib unit tests showed as uncovered. production_index_operation_queue, production_queue_publication and production_text_queue_contracts drive the producer, recovery, publication (budgets, trims, reconciliation, retired and foreign generations, fenced writers, uncertain commits), vector publication planning, text publication, text build limits and search overlays through production entry points. The vector driver's build, cleanup, adoption and allocation-race unit tests move into a support module both targets run.
A text search that overlays pending work may search its splits more than once as it widens past superseded results; only the first physical search of one logical search now records split demand, so one request no longer admits every split to the disk tier. A traversal-restricted vector or text search whose candidates are all superseded now returns only pending results instead of running the physical restricted search or loading the text manifest. Also lock in that index membership after an overlaid vector or text search keeps exactly the rows of the per-row filter, for strong and eventual reads and inside a write batch.
Vector Scan and text ScanSource steps scanned graph rows through their serializable outbox transaction, so any write in the remaining source range failed the step's commit and a backfill under steady writes retried its first step indefinitely. Every graph write committed after a build is created already queues a complete operation for the hidden generation, applied in commit order after activation. Steps therefore read source rows from a snapshot; build-owned rows stay in the serializable read set. A blocker is durable and the queue corrects documents, not blockers, so a step that blocks on a graph row also reads that row through its transaction and retries if the row is repaired. Tests commit writes inside held step windows and run 16 concurrent writers through vector and text builds; once the queue drains, both indexes equal cold builds of the final documents.
- Skip upserts that replay a node's exact indexed state: an entry-candidate row with the same vector bytes at the same deterministic layer returns before any delete, relink or reinsert, so replays of values a build already indexed stage nothing. - Hydrate relink candidates once per deletion layer with the batched item loader and keep each source's nearest M with select_nth_unstable, instead of ~36K cache lookups and ~35 full sorts per delete. A per-candidate oracle pins byte-identical rows. - Order vector planning caches by an append-only recency queue over a foldhash map instead of a BTreeSet with SipHash, bounded to four queue entries per value. Eviction, flush order and the graph are unchanged. A pinned SHA-256 of a drained graph under warm, cold and evicting sessions guards that caches change which rows publication reads, never a byte it writes.
…arness docker-image/tests/index_queue_contracts.py runs the product image through restart and SIGKILL exactness on disk and S3, the member-limit backpressure boundary, and the benchmark image's sample stream. The queue_acceptance_verify example opens a fixture written by the pre-queue release and checks graph data and strong and eventual search against exact oracles across queued writes, publication and restart. scripts/async-index-benchmark provisions, seeds, replays and summarizes map-versus-rows comparisons on EC2 and S3; nothing in it runs without an explicit, cost-approved invocation.
Read requests take search_consistency (strong by default, or eventual) in the Rust, TypeScript, Python and Go SDKs and the OpenAPI schema, with request-parity fixtures for both values. Write requests always search strongly and reject eventual.
A Search consistency guide is the one place strong and eventual search are described, with examples in every SDK and JSON. Prefiltered search explains how it treats superseded candidates; the HTTP API page, error references and troubleshooting cover the retryable 429 index_backpressure, including during an index build; guarantees, limits, per-document text admission and the Learn freshness bullet reflect queued vector/text publication. The llms.txt files are regenerated.
Points the DB vector coverage exclusions and dispositions at the lines they describe after the queue moved vector maintenance into publication.
…exes Review coverage the queue lacked: - An enqueue replayed below its acknowledgement never brings the operation back, in the merge algebra and through real SlateDB flushes and compactions. - Queue work committed only to the WAL survives a fenced writer and a failed WAL upload; two tenant scopes queue, recover, and discard independently; a blocked text build recovers once its row is repaired. - Every edge mutation keeps queued edge vector and text indexes exact against an oracle, with no row naming a dead edge.
…lushes A queued re-embedding in one partition is an upsert. Its delete removed every locator naming the node directly while the rows it edited stayed cached; the reinsertion relinked many rows back to their cached originals, so their flush staged no locator. A later delete then missed those rows and left links to a node without an item: searches failed with "missing simhash" and reinserting the node blocked publication. A cached row's flush is now the only writer of neighbor rows and their locators, staging the row's transition from the value the transaction holds. A delete stages the node's rows absent and directly deletes only locators no cached row's baseline links. Upserts now write only each row's net change. Adds a seeded concurrent soak under compaction and restarts (nightly, in release), random vector write sequences that check every HNSW namespace after each commit, stray-locator and bounded one-off cache runs, and a test pinning links released versions left without a locator.
…limits A vector Scan step that blocked behind admitted rows committed them without advancing its cursor, so every retry after a repair failed with invariant_violation. Such a blocker now ends the step like a full batch. A blocked build's hidden generation publishes nothing until a retry or abort, yet writes kept queueing there; once they filled the member or byte limit, the repair the blocker asks for was refused with a backpressure that never cleared. A saturated blocked build now admits its blocker's first repair and a removal of the blocker's entity beyond the limits, and refuses other writes with the non-retryable index_build_blocked. Writes to entities already pending stay admitted above the member limit; only new members are refused. Covered in unit and production contracts, including a repair racing a retry or abort of the build. Docs describe index_build_blocked.
Queue reads followed the whole backlog: - Acknowledgements piled up as removals in unresolved queue values. An acknowledgement composed with its own enqueue now cancels it inside partial merges, so unresolved values stay bounded by the outstanding backlog. Queue format 0x14 (new on this branch) accepts an empty value as a partial merge result. - Eventual searches decoded every queued operation and failed on corrupt records their budget never selects. They now decode only the entities whose latest operations fit the budget, comparing entities by raw bytes, and show each at its latest state or not at all. - Every publication attempt re-read the whole queue to publish one batch. The writer now retains each target's remainder after a successful commit and continues from it. Retained-byte ceilings whose worst-case queue value outgrows SlateDB's u32 value length are rejected. With a merge operand pending, SlateDB still resolves the whole value; tests pin that as a known limitation.
Strong text overlays analyzed every pending document of a partition with no budget. They are now bounded by one text publication's analysis charge: strong searches fail with retryable index_backpressure past it, eventual searches keep the oldest prefix within it, and a write batch's own documents past it fail with index_operation_batch_too_large. One queued operation that could never fit a publication blocked every later entity of its generation. It now holds its entity back in process memory while the rest publish; a newer write repairs it in rotation from full width, past one acknowledgement when needed, and a generation left with only held entities waits for a write instead of polling. Queues retained across uncommitted attempts share one writer-wide budget, and a stalled retained queue reads storage before waiting. The ledger wakes the index worker whenever an enqueue commit returns, committed, uncertain, or cancelled, so a racing write never waits out the stalled deadline. Health responses report the held-back entity count. No stored format changes.
An eventual search serves unpublished changes beyond its budget as last published, so until publication catches up it can return a node or edge that has since moved to another tenant, changed label, or lost the indexed property, through whole-index and prefiltered searches alike. Strong search never does. Document this on the search consistency guide and on SearchConsistency::Eventual.
Remaps the DB vector coverage exclusions and dispositions to the merged vector lines: main's refreshed entries plus the branch's, each pointing at the source line it describes. The server fingerprint is recorded as a non-root user. Against main the uncovered lines only shift, except the query event's backpressure message arm, which joins its uncovered sibling arms: 102 to 103 lines.
…oss the operation queue Version 5 marks stores that may hold asynchronous index-operation queues (0x14). Binaries that support at most version 4 ignore those queues and collapse their merge operands, so they must refuse an upgraded store instead of silently diverging. Version 5 shares version 4's physical layout. A writer upgrades a version-4 store by rewriting only the storage marker in one transaction; no index is rebuilt. Version 2/3 stores still run the equality-bitmap migration, which now publishes 5 directly. Current readers serve versions 4 and 5, so readers can be upgraded before the writer, and recovery-only managed failover reports WriterMigrationRequired for a version-4 store. The V4 cleanup-marker checks now compare against the equality-bitmap version, so stores written by the last version-4 release stay openable.
xav-db
force-pushed
the
async-index-queue
branch
from
October 4, 2026 10:38
e4f159a to
f27d8e3
Compare
xav-db
added a commit
that referenced
this pull request
Oct 5, 2026
…ing (#1165) Stacked on #1156; the base branch is `async-index-queue`, and it should be retargeted to `main` once #1156 merges. Fixes HEL-956. ## Problem In #1156, if planning one entity's queued index operation failed the same way every time, its whole (scope, index, generation) stopped publishing for good. The error was retried forever, the queue filled, writes to that index got 429, and `blocked_index_entity_count` stayed at 0. A real trigger exists: data written by released versions (HEL-952) can contain a self-link, and the vector insert searches never excluded the inserting node, so planning failed every time with `ContainsOwner`. ## Fix **1. A node is never its own neighbour.** - Both insert searches and entry-point resolution now exclude the inserting node. - A self-link already in storage is dropped in memory when its row loads, instead of failing every mutation. The row is rewritten without it only when a mutation actually changes that row. - Each guard has a test that fails without it, and a property test over damaged graphs checks that no new self-link is ever created. **2. Only the failing entity is held back.** ```text publish attempt fails FailureKind::of(error) // exhaustive match on HelixDbError, no string matching Fatal (closed / fenced) → stop Transient (I/O, conflict, cancel) → retry the generation with backoff (unchanged) Deterministic (corrupt input, invariant) vector → hold the entity at the failed position text → halve the epoch until one entity fails alone, then hold it held entity: never acknowledged or dropped; counted in blocked_index_entity_count retried alone after 60 s, or immediately when a newer write for it arrives every other entity keeps publishing ``` - Held state lives in memory only. After a restart, the entity is held back again on its first failed attempt instead of stalling the generation. - When several entities fail in a row, which suggests the whole generation is damaged, attempts back off instead of spinning. ## Behaviour changes - `blocked_index_entity_count` (in stats and health) now also counts entities held back because their planning failed, not only entities too large to fit a publication. The docs and the proto/OpenAPI descriptions say so. - A held entity that keeps being written still counts toward the index's shared queue budget, so it can eventually cause 429s for the whole index. This is documented, and bounding it is HEL-969. ## Testing (EC2) - fmt and clippy (workspace) are clean. - `cargo test -p db --lib`: 2,058 passed. - Release production targets (queue publication, internal contracts, text queue, index-operation queue, text correctness regressions): 92 passed. - `cargo test -p server` passes. - The release queue soak passes with 5 seeds. - The new isolation tests cover: - a vector entity failing mid-batch; - text epoch splitting; - a transient failure, where nothing is held back; - a restart while an entity is held. They also cover repair writes and the 60 s retry. - The text correctness regressions now expect a damaged root to hold back only its own entity, with the search still failing closed, instead of retrying forever. No stored format changes. The only effect on stored data is that a damaged row loses its self-link when a mutation rewrites it. **Expected CI follow-up:** the DB coverage fingerprint will change. Take the new count and sha from this PR's ARM64 "DB production thresholds" job. <!-- greptile_comment --> <!-- greptile_summary --> <p><a href="https://app.greptile.com/api/retrigger?id=74418662"><picture><source media="(prefers-color-scheme: dark)" srcset="https://greptile-static-assets.s3.amazonaws.com/badges/RetriggerDark.svg?v=2"><source media="(prefers-color-scheme: light)" srcset="https://greptile-static-assets.s3.amazonaws.com/badges/Retrigger.svg?v=2"><img alt="Retrigger" src="https://greptile-static-assets.s3.amazonaws.com/badges/Retrigger.svg?v=2" align="right"></picture></a><br clear="all"></p> <!-- greptile-risk --> The PR should wait for the production coverage baseline update so its required CI gate can pass. <h2>Findings</h2> 1. <img alt="P1" src="https://greptile-static-assets.s3.amazonaws.com/badges/p1.svg?v=9" align="top"> **Coverage baseline remains unchanged** <a href="https://github.com/HelixDB/helix-db/pull/1165#discussion_r4178564885">▶</a> <h3>Summary</h3> The PR makes queued index publication hold back an entity whose planning fails deterministically while continuing to publish other entities. It also prevents HNSW inserts from choosing their own node as a neighbor and tolerates stored self-links during mutation. - Vector failures identify the failing effect; text failures narrow an epoch until one entity remains. - Failed entities remain queued and become eligible for a timed retry or a newer write. - The production coverage baseline still needs updating for the changed source. <details><summary>Diagram</summary> ```mermaid %%{init: {'theme': 'neutral'}}%% flowchart TD A[Select queued entities] --> B[Plan publication] B -->|Success| C[Commit effects and acknowledge] B -->|Transient failure| D[Retry batch with backoff] B -->|Deterministic vector failure| E[Hold identified entity] B -->|Deterministic text failure| F[Narrow epoch] F -->|One entity fails| E E --> G[Publish other entities] E --> H[Retry on newer write or timer] ``` </details> <sub>Reviews (1) · Last reviewed commit: ["docs: count failed-planning holds in blo..."](https://github.com/helixdb/helix-db/commit/8708fe2ec237f7fbe545fd150b7657983a089224)</sub> <!-- /greptile_comment -->
matthewsanetra
added a commit
that referenced
this pull request
Oct 5, 2026
Release PR for **CLI 3.4.3** and **Docker image v0.0.10**, following the same shape as #1161. ```diff crates/cli/Cargo.toml version 3.4.2 → 3.4.3 (and Cargo.lock) crates/cli/src/config.rs DEFAULT_LOCAL_IMAGE_TAG v0.0.9 → v0.0.10 CLI tests, README, CLAUDE.md v0.0.9 → v0.0.10 docs default image, release notes, storage v5 warning, llms-full.txt docker-image/README.md published release v0.0.9 → v0.0.10 CONTRIBUTORS.md CLI image v0.0.9 → v0.0.10 sdks/tests/parity/COVERAGE.md published-image parity example v0.0.5 → v0.0.10 ``` Version-history references stay as written: - "v0.0.9 and earlier refuse to open" in the docker-image README's index storage version 5 section. - "Upgrading S3 deployments from v0.0.8 and earlier". - The "Docker images through v0.0.9" locator comments in `crates/db`. - The v0.0.5 compatibility statements. The TypeScript README's `package-smoke` commands stay on v0.0.5, because they mirror CI's digest-pinned compatibility check in `typescript-sdk.yml`. The `COVERAGE.md` parity pin was bumped together with the CLI default in v0.0.5 (`620fb543`) and missed afterwards. v0.0.5 can't serve the `search_consistency` parity fixtures added in #1156. ## Order 1. Merge this. The `Docker image` workflow's publish job requires `DEFAULT_LOCAL_IMAGE_TAG` on main to equal `release_version`, so the image can only be published after this lands. 2. Dispatch the `Docker image` workflow from main with `release_version=v0.0.10`. It reruns the full suite and publishes only if it passes. 3. Once v0.0.10 is published, run `cli.yml` from main to tag and publish v3.4.3. ## What's in v0.0.10 / 3.4.3 -⚠️ **Index storage version 4 → 5** (#1156). The first writer open rewrites only the version marker. After that, v0.0.9 and earlier refuse the database with `unsupported_index_storage_version`. - Upgrade readers before the writer. - Back up first if a rollback may be needed. - The local server guide now carries this warning next to the "set `tag`" instructions. - **Asynchronous vector and text index publication** (#1156). - Writes commit the graph change plus a queued index operation, and a background worker publishes it. - `search_consistency: strong | eventual`. Strong is the default and exact. - `index_backpressure` (429, retryable) is returned in three cases: past 1 GB or 250K pending entities per index, when a whole-index strong search is behind more than 800 changes, and when strong text analysis would exceed one publication's budget. - `index_operation_batch_too_large` (400) is returned when one write stages more than 8 MiB for one index. - Text admission limits are now per document. - Builds no longer livelock under concurrent writes. - **Failing entries are held back** (#1165). - An entry whose index update fails deterministically no longer stalls its whole generation. It is counted in `blocked_index_entity_count` and retried about once a minute. - Insert searches never pick the inserting node as its own neighbour. - **Fixes the v0.0.9 `missing simhash` known issue** (#1156 C1, #1165). Re-embeds keep their HNSW reverse-link locators. Indexes damaged by earlier versions still need a drop and recreate (HEL-952). - **Range-driven counts apply every filter** (#1169). For example, 149 → 49. - **CLI 3.4.3:** defaults to v0.0.10. There are no other CLI changes since 3.4.2. - **Not in this release:** the SDK changes in #1156, #1163 and #1164 ship with the next SDK releases. The release notes say to send `search_consistency` over HTTP until then. #1166 is a docs dependency bump. ## Testing - `cargo check --workspace --locked` passes, so the hand-edited `Cargo.lock` is consistent. - `cargo test --locked -p helix-cli` passes: lib 241, `e2e_cli` 13, `runtime_commands` 20, `typescript_runtime`, and the other targets. The Docker-only `e2e_runtime` tests are ignored by design, because they need the unpublished v0.0.10 image. - rustfmt 1.9.0 (the pinned 1.97.1 toolchain) `--edition 2024 --check` is clean on the edited Rust files. - Docs: `check-docs`, `generate-llms --check`, `generate-llms-full --check` and `check-openapi` pass. - Grep: every `helixdb/helixdb:v0.0.x` reference outside release notes is v0.0.10, except the CI-pinned v0.0.5 TypeScript smoke. - Workspace clippy is left to CI. The local nix nightly clippy rejects the repo's clippy configuration before linting, independent of this change. <!-- greptile_comment --> <!-- greptile_summary --> <p><a href="https://app.greptile.com/api/retrigger?id=74957602"><picture><source media="(prefers-color-scheme: dark)" srcset="https://greptile-static-assets.s3.amazonaws.com/badges/RetriggerDark.svg?v=2"><source media="(prefers-color-scheme: light)" srcset="https://greptile-static-assets.s3.amazonaws.com/badges/Retrigger.svg?v=2"><img alt="Retrigger" src="https://greptile-static-assets.s3.amazonaws.com/badges/Retrigger.svg?v=2" align="right"></picture></a><br clear="all"></p> <!-- greptile-risk --> The PR appears safe to merge, though the published-image parity instructions should include an image pull step. <h3>Summary</h3> This release PR updates the CLI default to Docker v0.0.10, bumps the CLI to 3.4.3, and updates image references and release guidance. The published-image parity instructions need a pull step before their new image pin can work without a local copy. <sub>Reviews (1) · Last reviewed commit: ["chore(release): CLI 3.4.3 and Docker ima..."](https://github.com/helixdb/helix-db/commit/281197c0c91e5f72a7ce89fb8a248e04176cdee4)</sub> <!-- /greptile_comment --> ## Also: fix main's workspace tests `counts_over_range_intersections_apply_every_filter` (#1169) overflows the default 2 MiB test-thread stack in Linux debug builds. That fails `Workspace tests`, `DB all targets` and `Workspace and tooling quality` on main, and the Docker release's `quality` job runs the same tests. `92c15755` runs it through the file's existing 16 MiB `run_high_stack_contract`, like the other heavy contracts. It passes locally even with `RUST_MIN_STACK=524288`.
xav-db
added a commit
that referenced
this pull request
Oct 6, 2026
…log (#1175) First PR of the queue stack: the HEL-957 PR (streamed backlog load at open), then the HEL-958 PR (admission memory charge), build on this one. Fixes HEL-961. ## Problem Draining a publication queue cost more than the size of its backlog: - Every attempt regrouped the whole retained queue by entity, and every commit rescanned it to remove the acknowledged operations. With 512-entity batches that is quadratic: going from a 100K to a 250K backlog (2.5×) cost 7.6× the bookkeeping. - A hold (an entity too large for a batch, or one whose planning failed) dropped the queue. The next attempt then re-read and decoded the whole backlog: 136 reads for 64 failing entities at 250K. - A retired generation's discard re-read the whole queue for every acknowledgement chunk: 63 reads at 250K with small acknowledgements. That also kept the generation's charges, which a recreated index shares, held for longer. - A target's schedule was removed only when a read found its queue empty, which never happens once a commit drains it. So `schedules` kept one entry for every target and generation ever published. ## Fix All changes are in memory. **1. The queue is grouped once per read** (`StoredQueue` in `queue/storage.rs`). ```text read order 0:a1 1:b1 2:a2 3:c1 4:b2 chains a: 0 -> 2 b: 1 -> 4 c: 3 order {0: a, 1: b, 3: c} // entities by their oldest outstanding op without(a1) a: 2 order {1: b, 2: a, 3: c} ``` - `rotation(after)` yields entities in the same order as the old regroup, at one map step per entity visited. - `without(acked)` pops chain heads in O(acked · log entities), with no rescan. It asserts that every acknowledgement names an entity's oldest operations in order. - `HeldEntity::repair_width` is O(1): a newer operation is queued exactly when the newest one isn't the `through` that is known not to publish. **2. Holds keep the queue.** A size block or a `FailureKind` isolation commits nothing, so the queue is retained instead of being read again. A retained queue whose entities are all `Waiting` or `Failed` is re-read once the target's latest admission has moved past the one seen before the read. So retries that keep falling due still never pin publication to a stale queue (the 61240f7 fix). **3. A discard continues from the retained queue.** It acknowledges whole-entity prefixes and keeps the rest, so a retired backlog is read and decoded once. Charges are released at each commit. **4. A drained target's schedule is dropped.** Only its latest vector commit is kept, in a ring of the 16 most recently drained targets (`MAX_RETAINED_PUBLICATIONS`). This lets its retained planning session still serve its next write. The other #1156 and #1165 invariants are unchanged: - the indexer is the only writer of index rows, and the ack is in the same transaction; - exact-ID enqueue and ack; - held-entity states, the 60 s retry, and `blocked_index_entity_count`. ## Behaviour/config changes - None visible to users or in config. - Fewer queue reads during holds and discards, and a recreated index gets its shared queue budget back sooner after a drop. ## Measurements (release, c7i.8xlarge, in-process, 2-dim vectors, 512-entity batches; main → branch) **Bookkeeping only** (selecting a batch and removing what was acknowledged): | Backlog | main | branch | |---|---|---| | 10K | 0.011 s | 0.003 s | | 100K | 1.23 s | 0.047 s | | 250K | 9.36 s | 0.157 s | | 1M | — | 0.856 s | **Publish drain:** | Backlog | main | branch | |---|---|---| | 10K | 5.47 s | 5.27 s | | 100K | 60.3 s | 59.3 s | | 250K | 176.2 s (1,418 ops/s) | 154.9 s (1,614 ops/s), −12% | HNSW planning takes most of the time. **64 entities that fail to plan, then a drain to Stalled** (time, reads): | Backlog | main | branch | |---|---|---| | 10K | 7.69 s, 67 reads | 7.19 s, 2 reads | | 100K | 73.0 s, 80 reads | 66.1 s, 2 reads | | 250K | 185.0 s, 136 reads | 161.1 s, 2 reads | At 250K, time spent reading falls from 12.4 s to 0.28 s. **Retired discard with a 64 KiB ack budget** (time, reads): | Backlog | main | branch | |---|---|---| | 10K | 0.31 s, 4 reads | 0.31 s, 2 reads | | 100K | 2.57 s, 26 reads | 2.57 s, 2 reads | | 250K | 12.68 s, 63 reads | 6.52 s, 2 reads | With default acks at 250K, the discard goes from 5 reads to 2 and from 0.96 s to 0.79 s. ## Testing (EC2) - fmt, clippy (workspace, all targets) and `cargo check --locked --workspace --all-features` are clean. - `cargo test -p db --lib`: 2,069 passed. `cargo test -p db --doc`: 40 passed. - `cargo test -p server` passes (71 unit, 2 integration, 5 doc tests). `scripts/validate-cargo-target-references.py` passes. - Release production targets, with `production-coverage,index-lifecycle-testing`: 185 passed. These are queue publication, index-operation queue, text queue, text correctness regressions, internal, contracts, migration, and index lifecycle (run single-threaded, as CI runs it). - Product image: `index_queue_contracts.py --scenario all` passes. Indexes are exact after restart and after SIGKILL, on local disk and on S3. Writes get a 429 at 250,000 members and are accepted 1.4 s after activation. - New `grouping_tests.rs`: a property test (1,024 cases × up to 95 steps). It runs random enqueues, rereads, publishes, every kind of hold, releases, discards and cursor moves, and checks the result against a verbatim copy of the old regrouping. Selections, remaining queues, rotation order and retained bytes must all be identical. There are also should-panic tests for out-of-order and duplicate acks. - New `drain_tests.rs` covers: - retiring mid-drain; - recreating an index whose old backlog is exactly at the limit (429 until the first discard commits, then admitted); - schedule cleanup across 24 targets; - restarting mid-drain with updates and deletes; - concurrent writes during a drain; - holds reusing one read. - The #1165 isolation tests pass unchanged. No stored format changes. **Expected CI follow-up:** the DB coverage fingerprint will change. Take the new count and sha from this PR's ARM64 "DB production thresholds" job. <!-- greptile_comment --> <!-- greptile_summary --> <p><a href="https://app.greptile.com/api/retrigger?id=75528033"><picture><source media="(prefers-color-scheme: dark)" srcset="https://greptile-static-assets.s3.amazonaws.com/badges/RetriggerDark.svg?v=2"><source media="(prefers-color-scheme: light)" srcset="https://greptile-static-assets.s3.amazonaws.com/badges/Retrigger.svg?v=2"><img alt="Retrigger" src="https://greptile-static-assets.s3.amazonaws.com/badges/Retrigger.svg?v=2" align="right"></picture></a><br clear="all"></p> <!-- greptile-risk --> The PR appears safe to merge, with two non-blocking memory and planning-cache efficiency issues worth addressing. <h3>Summary</h3> The PR groups each queue once per read, advances retained queues by acknowledged entity prefixes, reuses queues across holds and retired-generation discards, and bounds schedule history for drained targets. It adds property, drain, and production-support coverage. Two non-blocking resource-efficiency concerns remain: retained allocation size during a long drain and unusable commits entering the drained-target ring. <details><summary>Diagram</summary> ```mermaid %%{init: {'theme': 'neutral'}}%% flowchart LR A[Read durable queue] --> B[Group operations by entity] B --> C[Select ordered prefixes] C --> D{Attempt outcome} D -->|Commit| E[Acknowledge IDs and retain remainder] D -->|Hold or definite conflict| F[Retain unchanged queue] D -->|No selectable work| G[Drop queue and wait] E --> C F --> C ``` </details> <sub>Reviews (1) · Last reviewed commit: ["test(db): cover drains across retirement..."](https://github.com/helixdb/helix-db/commit/4288421a163fb46485908c11b6d29277759bb7a3)</sub> <!-- /greptile_comment -->
xav-db
added a commit
that referenced
this pull request
Oct 6, 2026
Fixes HEL-957.
## Problem
At writer open, `load_backlog` fully decoded every queue of a scope
(every payload, including 1536-dimension vectors) into one `Vec` before
charging any of it to the admission ledger. So open's peak memory and
time grew with the whole durable backlog: about 25 s and 4.5 GB for 2.26
GB of queued work. Each scope also cost about 4 scans (scope seek,
record scan, layout probe, queue scan).
## Fix
```text
before: for scope: decode every queue of scope -> Vec -> charge all
after: one forward scan per keyspace (legacy, then tenants)
row -> QueueDiscovery::push -> completed queue? -> validate owner -> charge -> drop
```
- **Streaming discovery.** `QueueStore::discover` is replaced by
`QueueDiscovery`, which takes rows in key order and returns each queue
as soon as it is complete. In the map layout each row is a whole queue.
In the row layout a generation completes when the next generation's
first row arrives, or at `finish()`. Open now holds one queue at a time
plus the ledger.
- **Framing only.** `OperationFrame::decode_queue` / `decode_row` read
each operation's ID, entity and retained bytes. Payloads are validated
in place (`algebra::validate_body`) and never decoded. The values and
rows they reject are exactly the ones `OperationQueue::decode` /
`QueueRow::decode` reject (duplicate IDs, NaN, UTF-8, trailing bytes,
unresolved values). A proptest checks this, and `OperationQueue::decode`
and `OperationFrame::decode_queue` share one record walk
(`decode_unique`), so their duplicate-ID check cannot drift.
- **Fewer scans.**
- One scan covers the legacy scope.
- One forward-only iterator covers the tenant keyspace, with at most 2
seeks per tenant.
- Index records are read only for scopes that have a queue, and are
cached per scope.
- **Layout check folded in.** `discovery_range` spans both layouts'
adjacent record kinds (a const assert checks this). A queue of the other
layout still fails closed with `Config`, during the same scan. The
separate probes (`require_scope_layout`, `require_layout`,
`discover_scopes`) are now compiled only for tests and the benchmark
feature.
- **Ledger only.** Open fills the admission ledger and nothing else:
publication starts with no retained queue, schedule or held entity, and
each target's first attempt reads and groups its own queue
(`StoredQueue`).
- **Removed:** `QueueStore::discover`, `finish_rows` and
`OperationQueue::from_rows`. `OperationQueue::operations()` is now
compiled only for tests and `index-lifecycle-testing`.
- **Unchanged:**
- Charge order (legacy first, then tenants ascending).
- Owner validation.
- The row-layout sequence resume.
- The worker wake.
- The #1156/#1165 invariants: the indexer is the only writer of index
rows, and held-entity state starts empty either way.
## Behaviour changes
- A corrupt `IndexRecord` in a tenant that has **no** queue no longer
fails writer open. That tenant's first catalog read fails instead
(`Encoding`), and the other tenants keep working. A corrupt record in a
scope that has a queue still fails the open.
- No config or API changes.
## Measurements (EC2 c7i.8xlarge, release)
Fixture: 6 legacy vector queues of 256 MiB, 8 tenant queues of 64 MiB
(dropped indexes), and 2000 tenants with rows but no queue. That is 14
queues, 366,930 ops and 2.26 GB retained, about 6 GB on disk. Each run
starts from a fresh copy of the fixture, 4 runs each, and every run
loaded the same totals.
| Phase | base | this PR |
|---|---|---|
| Backlog load only (no compactor) | 15.1–15.3 s, peak 2.17–2.22 GB |
4.63–4.86 s, peak 1.14 GB (flat) |
| Full writer open (compactor running) | 24.4–24.8 s, peak 4.33–4.65 GB
| 4.83–5.01 s, peak 1.22–1.30 GB |
- Open is about 5× faster.
- The base column was measured on main. `linear-queue-publication-drain`
does not touch the open-time load path (`recovery.rs`, queue discovery),
so it applies unchanged.
- Without the compactor, this PR's peak scales with the largest queue:
about 0.85 GB at 128 MiB queues, 1.14 GB at 256 MiB. The base's peak
scales with the whole backlog.
- Open peaks above the load-only peak come from SlateDB's background
compactor running during open, not from the load.
## Testing (EC2)
- `cargo fmt --all --check`, `cargo clippy --workspace --all-targets --
-D warnings` and `cargo check --locked --workspace --all-features` are
clean.
- `cargo test -p db --lib`: 2,082 passed. `cargo test -p db --doc`: 40
passed.
- Release production targets (queue publication, index-operation queue,
text queue, text correctness regressions, internal, contracts,
migration, index lifecycle), one test thread each: 186 passed.
- `cargo test -p server`: 79 passed.
- `scripts/validate-cargo-target-references.py` passes.
- Product image, `docker-image/test.sh` and `index_queue_contracts.py
--scenario all` pass. After SIGKILL and restart, 10,750 docs are exact
on local disk and on S3. The member limit rejects at 250k and accepts
again 1.4 s after the build activates.
- New `recovery_tests` compare every load against a reference loader
that fully decodes each queue the old way. They cover:
- legacy and tenants 0, 7, 8 and `u128::MAX`, with keys on both sides of
each range;
- the map layout across the compacted run, L0 and the memtable, with
partial and full acks;
- the row layout, with interleaved generations and the sequence resume;
- the other layout failing closed in every scope;
- corrupt payloads, duplicate IDs, foreign keys and invalid tenant
envelopes failing closed;
- misowned tenant queues after other scopes' records are cached (this
test fails if the per-scope cache goes stale);
- an SST read failure injected at every point during open, after which
every reopen reloads the same ledger and writes nothing;
- a full writer restart that reloads the ledger exactly, starts with no
queue reads, retained bytes, schedules or blocked entities, and drains
each reloaded target in exactly 2 reads.
- Production contracts add a tenant scope-walk test and duplicate-ID /
corrupt-payload cases.
- `open_measurement_tests` is an ignored, manual release benchmark that
produced the numbers above.
No stored format changes.
**Expected CI follow-up:** the DB coverage fingerprint will change.
Rows-layout-only lines are uncovered in production targets, and line
numbers shift. Take the new count and sha from this PR's ARM64 "DB
production thresholds" job, after `linear-queue-publication-drain`
merges and this PR is retargeted to main.
xav-db
added a commit
that referenced
this pull request
Oct 6, 2026
…ext analyses (#1176) Fixes HEL-960. ## Problem A strong text search overlays every committed but unpublished document of its partition so its BM25 answer stays exact. On main, and unchanged by the base PR: - it was bounded by one text publication's analysis budget (64 MiB, about 5,500 pending 40-word documents). Any ingest burst past that made strong text searches fail with retryable 429 `pending_text_analysis_bytes` until the index worker caught up (106 and 234 rejections in the 20K and 40K burst runs below); - every search analyzed the whole backlog again from scratch and copied it into an in-memory Tantivy index, so cost and memory grew with the backlog and with each concurrent search (about 200 ms and 1 GB RSS at 5K pending with 1 searcher, 1.7 GB with 8). ## Fix ```text strong text search pending set (per request, from #1174) for each pending doc in the partition cached analysis (by queued op id)? reuse it, charged as if fresh else wait for the analysis turn -> analyze on the blocking pool -> cache charge > strong_text_search_max_analysis_bytes -> 429 pending_text_analysis_bytes score pending docs in place (Tantivy Bm25Weight, best k) -> merge with published results ``` **1. Strong text searches have their own bound.** - `IndexOperationQueueTuning::strong_text_search_max_analysis_bytes` (default 512 MiB, about 44K pending 40-word documents), next to the base's strong vector bound. Past it, the search fails with the same retryable 429 `index_backpressure`, resource `pending_text_analysis_bytes`. - Applies to read and write requests. A write's own documents are charged first; alone past the bound they fail the write with `index_operation_batch_too_large`, as before. - Eventual searches are unchanged: they overlay at most one publication budget (64 MiB) and serve the rest as last published. **2. Analyses of pending text are reused (in memory only).** - `PendingTextAnalyses`, per database, keyed by (queue target, partition, queued operation ID). Selection reads the base's per-request decoded queue, where each committed entity now carries its queued operation ID. - It is exact without invalidation: a queued payload never changes and its random 121-bit ID is never reused, so publication, supersession, a rebuild or a drop can only make an entry unreachable. A reused analysis is still compared with the queued text and fails closed on a mismatch. - A reused analysis is charged exactly as a fresh one, so a 429's `requested` value is the same cold or warm. - Capped at the bound; the least recently used partitions give up only the bytes a replacement needs. - Released when the publisher drains or discards a generation, and when a strong search finds its queue empty (the only path on reader handles). - A write transaction's own text is never cached. **3. Pending documents are scored in place.** - No clone and no in-memory index. Scoring uses Tantivy's `Bm25Weight` with the same corpus statistics and fieldnorm quantization, ranked by score then entity ID. - Scoring runs on the blocking pool through the base's `run_blocking`, in steps of 4,096 documents with a request check between steps, keeping only the best k per step, so memory grows with k, not with the number of hits. - A proptest against a real Tantivy index requires bit-exact results for queries of up to 2 terms. For 3 or more terms results must match within rounding, because Tantivy's own union sums in an order that depends on index layout. **4. Only one strong search at a time analyzes uncached text.** - A process-wide first-come-first-served turn. A search that waited reads the cache again and reuses what the previous holder analyzed. - Cold analysis runs through `run_blocking` and checks the request before each fresh document. The turn is an owned guard moved into that blocking work and released only when it stops, so an abandoned request cannot let a second cold analysis overlap. - A pass that stops early (deadline, cancellation, reader retirement) still caches what it analyzed plus the cached entries after it, so the next search resumes instead of starting over. - Fully cached strong searches and eventual searches never wait. - However many searches run at once, analyses of pending text stay at about two bounds: one cached and one being analyzed. **5. Observability.** `IndexOperationQueueStats::strong_text_search_rejections` counts refusals. A warning naming the setting is logged at most once a minute. #1156/#1165 invariants are unchanged: the indexer is still the only index writer, acks stay in the same transaction, strong search stays exact, and held-back entities still count toward the bound. ## Behaviour/config changes - New server env var `HELIX_STRONG_TEXT_SEARCH_MAX_ANALYSIS_BYTES`: a positive byte count, default 512 MiB. It feeds the same tuning builder as the base's vector variable. An invalid value fails startup with `IndexQueueBytes` naming the variable, before any cache directory is created. Benchmark builds keep it. Embedded users set it with `IndexOperationQueueTuning::with_strong_text_search_max_analysis_bytes`. - Strong text 429s now start at 512 MiB of pending analysis per partition instead of 64 MiB. The resource name is unchanged. - While text is pending, the database keeps up to one bound of cached analyses, plus about one more while a cold search analyzes. - New stats field `strong_text_search_rejections`. On reader handles it is the only nonzero field. - Docs updated: search-consistency (one table for both strong bounds), troubleshooting, error-handling, prefiltering, Cloud limits, local-server, docker-image README, OpenAPI 429 description, llms-full.txt. ## Measurements EC2 c7i.8xlarge. Ingest burst, product benchmark image: 100 docs/s plus one burst, 2 strong searchers, 90 s. Latency counts successful searches only, so main's numbers leave out the searches it refused, which were the ones behind the largest backlogs. Measured before the restack onto the base PR. | Burst | Image | Strong-search 429s | p50 / p99 (ms) | Peak container memory | |---|---|---|---|---| | 20K | main | 106 of 1,705 | 50.3 / 129.4 | 1,716 MiB | | 20K | branch | **0** of 1,580 | 58.5 / 139.5 | 1,271 MiB | | 40K | main | 234 of 1,598 | 58.7 / 166.1 | 2,002 MiB | | 40K | branch | **0** of 1,335 | 77.3 / 223.0 | 1,287 MiB | The 40K burst peaked at about 279 MiB of pending analysis. In-process probe, release build: 20K published docs, then N pending with publication paused. Every document and query contains a common term, so no pending document is filtered out, which is the worst case. Values are strong search p50 / p99 (ms) and process RSS. `*` marks values measured before the restack and not re-run. | Pending | Searchers | main | branch p50 / p99 | branch RSS | |---|---|---|---|---| | 0 | 1 | 13.2 / 18.5 | 12.8 / 18.6 * | | | 5K (59 MiB) | 1 | 204 / 217, 1,062 MiB | 27 / 39 | 322 MiB * | | 24K (282 MiB) | 1 | 429 | 113 / 439 * | 689 MiB * | | 40K (470 MiB) | 1 | 429 | 208 / 299 | 813 MiB * | | 5K | 8 | 261 / 293, 1,668 MiB | 46 / 63 | 376 MiB * | | 40K | 8 | 429 | 248 / 401 | 1,529–1,680 MiB peak | | 40K | 32 | 429 | 904 / 1,152 * | 1,860 MiB peak * | - Versus the pre-restack build on the same host in the same session: single-searcher latency is the same; at 8 searchers p50 is up to about 8% higher and p99 the same or lower. No search was rejected. - RSS at 40K with 8 searchers is higher than before the restack (1,118–1,192 MiB): glibc keeps per-thread arenas for the extra blocking-pool threads. With `MALLOC_ARENA_MAX=2` both builds sit at about 790 MiB. The multi-searcher RSS values marked `*` are likely higher for the same reason. - Eventual search at 5K pending with 1 searcher: 184 → 27 ms p50 *, because it also scores in place now. - A warm strong search still costs O(backlog) CPU: one statistics-marker read per pending entity, plus scoring. The statistics pass and the scope/term filter stay on the async workers, one cheap pass over the backlog each. At 32 concurrent searches and 40K pending this saturates the 32 cores. ## Testing (EC2) - `cargo fmt --all --check`, `cargo clippy --workspace --all-targets -- -D warnings` and `cargo check --locked --workspace --all-features` are clean. `validate-cargo-target-references.py` passes. - `cargo test -p db --lib`: 2,097 passed (26 more than the base). `cargo test -p db --doc`: 40 passed. - Release production targets (`production-coverage,index-lifecycle-testing`, `--test-threads=1` as CI runs them): 187 passed: contracts 58, internal contracts 41, migration contracts 27, text correctness regressions 23, index-operation queue 15, text queue contracts 12, index lifecycle contracts 8, and queue publication 3. - `cargo test -p server` passes (72 unit tests plus the integration and doc targets), as does `--features async-index-benchmark` (76 unit tests). - Product image: `docker-image/test.sh` passes, and `index_queue_contracts.py --scenario all` passes. Correctness is exact after restart and SIGKILL on disk and on S3 (10,750 docs each). In the member-limit scenario, the batch past 250,000 members gets a 429 and the retried batch is accepted 2.5 s after activation. - Docs: `check-docs`, `generate-llms --check`, `generate-llms-full --check` and `check-openapi` pass. - New coverage: - Cache: replacement keeps exactly the latest selection; exact-budget and one-byte-over boundaries; partial LRU shedding across partitions; forget per target; a search holding an old snapshot still reads it. - Selection: identical charges cold and warm; refusals report the same `requested` value cold and warm, at the text charge and at a token; refused searches still cache the prefix they analyzed; a write's own text is never cached; a cached text mismatch fails closed; committed and local entities keep their order and source. - Turn: a cold search waits, then reuses the holder's analyses. Cached searches, eventual searches and searches of another partition do not wait. A search over a write's own text does. A holder whose request is dropped, retired or expires mid-analysis keeps the turn until its blocking work stops; it caches its prefix plus the cached tail, and the next search finishes without analyzing any document twice. - Scoring: proptest against a real Tantivy index covering fieldnorm quantization, ties, restricted scopes, empty documents and extra corpus statistics. Stepped scoring keeps the k best across steps and checks the request at each step. - Overlay: - exact past the publication budget, across a restart and after publication; - the exact bound passes and one byte over fails, cold and warm; - a write batch sees its own staged text despite a warm cache; - updates, deletes and re-inserts; - concurrent strong searches racing writes and automatic publication; - the rejection counter. - Release: the publisher keeps the cache while work remains and drops it once drained. Dropping a text index releases its analyses. A strong search over an empty queue releases them on writer and reader handles; with that release removed, both tests fail. - Production contract (`production_text_queue_contracts`): - the default bound stays exact past the publication budget; - a lowered bound gives a retryable 429 in read requests, through `HelixQueryService` (as Backpressure), and in write requests, which roll back; - eventual results are unchanged, and publication clears the 429; - a write's own text past the bound gets `IndexOperationBatchTooLarge`. - Server config: one test covers each strong bound variable and both together. Values 1, 2^31 and u64::MAX are accepted. 0, -1, empty, `512MiB` and u64::MAX+1 are rejected before any cache directory is created. Benchmark builds keep both bounds. No stored format changes. **Expected CI follow-up:** the DB coverage fingerprint will change. Take the new count and sha from this PR's ARM64 "DB production thresholds" job.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Vector and text index maintenance moves off the write path. A write now commits its graph change plus a queued index operation. A single background indexer publishes those operations into the HNSW and text indexes. Searches overlay any work still queued, so results stay exact.
sequenceDiagram participant W as Write participant Q as Index queue (0x14) participant P as Indexer participant I as HNSW / text index participant S as Search W->>Q: commit graph change + queued op (one txn) P->>Q: read pending ops P->>I: plan + apply batch (build planner, retained session) P->>Q: ack exact op IDs (same txn) S->>I: search published index S->>Q: overlay pending ops (strong) / skip within lag (eventual)The indexer, made up of lifecycle builds and the queue publisher, is the only thing that writes index rows. User writes only enqueue operations.
0x14: one merge-backed queue value per (scope, index, generation), holding exact operation IDs, with a merge algebra for enqueue and ack.0x15: a rows layout used only by tests and benchmarks.0x14queues and silently ignore them. They now refuse an upgraded store:Unsupported index storage version 5; this binary supports 4.WriterMigrationRequired { found: 4, target: 5 }and leaves the store for a controlled migration.What changes
429 index_backpressure.plan_and_applyoverVectorPlanTarget::{Build, Publication}, prefix commits, deterministic layers, and planning sessions retained between commits. The vector memory cache commits throughcommit_fenced.search_consistency: strong | eventual. Strong overlays pending work, and fails with a 429 when more than 800 superseded rows are ahead of its answer. Eventual degrades instead of failing.CatchUpandBuildDeltaare removed.Benchmarks (EC2 c7i.8xlarge, 100K existing docs, 100 writes/s during the build)
mainTesting
Everything below ran on EC2 at this head,
f02e63a1b, which sits on mainf7fe7b39d.production-coverage,index-lifecycle-testingsuite (2,565 passed);production-scalesuite;scoped_search_gate;queue/soak_tests.rs). 8 writers and 4 readers run against tenant-partitioned vector and text indexes while builds, an index drop and recreate, and restarts happen. Each seed sees about 2,000 compactions, with queues holding work during them. Every read is compared with an exact oracle, and the HNSW graph invariants are checked throughout. It passed 20 seeds, and it also runs nightly.The only failure was a test from main,
object_store_warm_starts_only_when_the_tier_enables_it. It runs out of open files on a host with a 1,024 soft limit, and it fails the same way on main.Expected CI follow-up: the ARM64 DB coverage fingerprint will change. Its count and sha will be taken from this PR's CI run.
Reviewing
There are 21 commits, ordered bottom-up. The last one bumps the index storage version. Commits 1–13 are the original feature: codec, then the store and ledger, then the core write and publish path, then the performance and fix commits, then tests, SDKs and docs. Commits 14–20 are the final-review fixes, one per bug group. Each commit builds on its own.
Final review fixes (commits 14–20)
A final correctness review found these bugs. Each one was fixed against a test that reproduced it, reviewed adversarially, and then the full validation above was re-run.
missing simhashsearch errors, or a publication that stalls for good.Follow-ups (Linear):
cachePuts.The PR has no established blocking behavioral defect, but its
OnceLockinitializations must satisfy the explicit repository requirement before merging.Findings
Summary
The PR moves vector and text indexing to queued background publication, overlays pending work in searches, and adds lifecycle, cache, benchmark, SDK, and documentation support. The revisions also address blocked builds, oversized publication work, text-overlay limits, and vector locator maintenance.
Diagram
sequenceDiagram participant W as Write participant Q as Durable index queue participant P as Publisher participant I as Published indexes participant S as Search W->>Q: Commit graph change and operation P->>Q: Read pending operations P->>I: Apply bounded publication P->>Q: Acknowledge published IDs S->>I: Read published results S->>Q: Overlay pending work by consistency modeReviews (2) · Last reviewed commit: "ci: remap DB coverage metadata and recor..."