Skip to content

feat(db): asynchronous vector and text index publication - #1156

Merged
xav-db merged 22 commits into
mainfrom
async-index-queue
Oct 5, 2026
Merged

xav-db merged 22 commits into
mainfrom
async-index-queue

Conversation

@xav-db

@xav-db xav-db commented Oct 1, 2026 •

Copy link
Copy Markdown
Member

Summary

Vector and text index maintenance moves off the write path. A write now commits its graph change plus a queued index operation. A single background indexer publishes those operations into the HNSW and text indexes. Searches overlay any work still queued, so results stay exact.

sequenceDiagram
    participant W as Write
    participant Q as Index queue (0x14)
    participant P as Indexer
    participant I as HNSW / text index
    participant S as Search
    W->>Q: commit graph change + queued op (one txn)
    P->>Q: read pending ops
    P->>I: plan + apply batch (build planner, retained session)
    P->>Q: ack exact op IDs (same txn)
    S->>I: search published index
    S->>Q: overlay pending ops (strong) / skip within lag (eventual)
Loading

The indexer, made up of lifecycle builds and the queue publisher, is the only thing that writes index rows. User writes only enqueue operations.

⚠️ Stored format

  • New record kind 0x14: one merge-backed queue value per (scope, index, generation), holding exact operation IDs, with a merge algebra for enqueue and ack.
  • 0x15: a rows layout used only by tests and benchmarks.
  • Index storage version 4 → 5. Version 5 shares version 4's physical layout, so a writer upgrades a v4 store by rewriting only the version marker, in one transaction. No index is rebuilt.
    • Why: v0.0.9 and earlier would otherwise open a store with pending 0x14 queues and silently ignore them. They now refuse an upgraded store: Unsupported index storage version 5; this binary supports 4.
    • Upgrade order: current readers serve both v4 and v5, so upgrade readers before the writer.
    • Managed failover: it reports WriterMigrationRequired { found: 4, target: 5 } and leaves the store for a controlled migration.
    • Older stores: v2/v3 stores still run the equality-bitmap migration, which now publishes 5 directly.
    • Verified end to end: the v0.0.9 image wrote indexed data, and this branch's image opened it with identical search results, accepted new inserts, updates and deletes, and reopened cleanly. v0.0.9 then exited on the upgraded store.

What changes

Area Change
Writes Enqueue only. An admission ledger caps the backlog at 1 GB / 250K members, and writes past the cap get 429 index_backpressure.
Publisher Reuses the build planner: plan_and_apply over VectorPlanTarget::{Build, Publication}, prefix commits, deterministic layers, and planning sessions retained between commits. The vector memory cache commits through commit_fenced.
Search search_consistency: strong | eventual. Strong overlays pending work, and fails with a 429 when more than 800 superseded rows are ahead of its answer. Eventual degrades instead of failing.
Backfill Builds cut at a fixed ID watermark and read source rows from an untracked snapshot, which fixes a livelock under concurrent writes. Writes made during a build are queued for the hidden generation and published after activation. CatchUp and BuildDelta are removed.
Update drain Skips exact replays, hydrates relink candidates once per layer, and makes cache recency O(1).
SDKs / docs Search consistency is in every SDK and the API schema. The docs gain a new section on search consistency and the 429 behaviour.

Benchmarks (EC2 c7i.8xlarge, 100K existing docs, 100 writes/s during the build)

Backfill with writes, until every write is searchable main This PR
vec128, disk livelocked 958 s
vec768, disk livelocked 2,507 s
vec768, S3 + disk cache not run 2,512 s
  • Writes during builds: p99 193–306 ms, with no 429s.
  • Searches during the drain: no rejections.
  • Drain throughput: steady inserts went from 6.2 to 239 ops/s. In-process, updates went from 32 to 63 ops/s and deletes from 43 to 97 ops/s.

Testing

Everything below ran on EC2 at this head, f02e63a1b, which sits on main f7fe7b39d.

  • Lint: fmt, clippy (workspace and benchmark features).
  • Rust tests:
    • workspace (4,770 passed);
    • the release production-coverage,index-lifecycle-testing suite (2,565 passed);
    • lifecycle contracts;
    • the 100K production-scale suite;
    • scoped_search_gate;
    • 50 consecutive passing runs of the vector cache node tests.
  • SDKs: TypeScript, Python, Go, Rust and uniffi, plus request/server/embedded parity.
  • Images: the product and benchmark images build. The image suite, queue contracts on disk and real S3, member-limit and the benchmark contract all pass.
  • Soak: a new release soak (queue/soak_tests.rs). 8 writers and 4 readers run against tenant-partitioned vector and text indexes while builds, an index drop and recreate, and restarts happen. Each seed sees about 2,000 compactions, with queues holding work during them. Every read is compared with an exact oracle, and the HNSW graph invariants are checked throughout. It passed 20 seeds, and it also runs nightly.
  • Server coverage: run as non-root; the fingerprint has been re-recorded.

The only failure was a test from main, object_store_warm_starts_only_when_the_tier_enables_it. It runs out of open files on a host with a 1,024 soft limit, and it fails the same way on main.

Expected CI follow-up: the ARM64 DB coverage fingerprint will change. Its count and sha will be taken from this PR's CI run.

Reviewing

There are 21 commits, ordered bottom-up. The last one bumps the index storage version. Commits 1–13 are the original feature: codec, then the store and ledger, then the core write and publish path, then the performance and fix commits, then tests, SDKs and docs. Commits 14–20 are the final-review fixes, one per bug group. Each commit builds on its own.

Final review fixes (commits 14–20)

A final correctness review found these bugs. Each one was fixed against a test that reproduced it, reviewed adversarially, and then the full validation above was re-run.

Bug Fix
C1 (critical): a vector re-embed dropped the HNSW reverse-edge locators for links that the re-insert restored. That leads to dangling links later, missing simhash search errors, or a publication that stalls for good. Flushing a cached row is now the only writer of neighbour rows and their locators. The flush moves each row from its baseline to its current state. Randomized graph-invariant tests cover the publication and build paths with cosine and euclidean. The published golden digest changes only because 443 locators the bug used to lose are now written.
H1: a build could not recover from a blocker that came after valid rows in the same scan batch. Every retry after the repair hit an invariant violation. The scan step now ends before the blocker and commits the rows admitted so far, with the cursor moved past them.
M2: a blocked build plus ongoing writes could fill the member cap, and then the repair write itself got a 429. A write that repairs or deletes a blocked build's blocker entity is admitted past the queue limits.
H2: acknowledgement removals piled up in partial merge values until a compaction reached the bottom run, so read cost grew without bound (and a value past 4 GiB would hit SlateDB's u32 value-length limit). A partial merge now cancels an enqueue against its own acknowledgement. This relies on SlateDB applying each committed operand exactly once, which HelixDB's counter merge operator already assumes.
M3: every search and every publication attempt decoded the whole queue. Eventual searches decode only their budget's prefix, and publication keeps the decoded queue between commits.
H3: strong text search analyzed the whole pending backlog in memory on every query. Strong text overlays are bounded by one publication's analysis budget. Past that bound, strong search returns 429 and eventual search degrades.
M1: one operation that could never fit a publication stalled its whole generation. That entity alone is held back and every other entity keeps publishing. The health response reports a count of held-back entities, with no identifiers. A superseding write repairs it.
M4: eventual search can return rows that have since moved partition, changed label or lost the indexed property. This is kept as the documented behaviour of eventual search. Strong search never returns such rows.

Follow-ups (Linear):

  • HEL-952: main's synchronous update path has the same C1 bug, so data written by released versions may contain links without locators. A damaged index can be dropped and recreated.
  • HEL-953: a writer-wide cap on retained queues.
  • HEL-954: eventual reads still cost O(backlog) while merge operands are pending, which is negligible at typical write rates.
  • HEL-950: object-store cachePuts.

Retrigger

The PR has no established blocking behavioral defect, but its OnceLock initializations must satisfy the explicit repository requirement before merging.

Findings

  1. P1 Cgroup path escapes root ▶

Summary

The PR moves vector and text indexing to queued background publication, overlays pending work in searches, and adds lifecycle, cache, benchmark, SDK, and documentation support. The revisions also address blocked builds, oversized publication work, text-overlay limits, and vector locator maintenance.

Diagram
sequenceDiagram
    participant W as Write
    participant Q as Durable index queue
    participant P as Publisher
    participant I as Published indexes
    participant S as Search
    W->>Q: Commit graph change and operation
    P->>Q: Read pending operations
    P->>I: Apply bounded publication
    P->>Q: Acknowledge published IDs
    S->>I: Read published results
    S->>Q: Overlay pending work by consistency mode
Loading

Reviews (2) · Last reviewed commit: "ci: remap DB coverage metadata and recor..."

@mintlify

mintlify Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

Preview deployment for your docs. Learn more about Mintlify Previews.

Project Status Preview Updated
helix 🟢 Ready View Preview Oct 4, 2026, 10:38 AM

💡 Tip: Enable Automations to automatically generate PRs for you.

Comment thread scripts/async-index-benchmark/resources.py
@xav-db
xav-db marked this pull request as draft October 1, 2026 14:54
@xav-db
xav-db force-pushed the async-index-queue branch from 747e31f to f02e63a Compare October 3, 2026 15:13
@xav-db
xav-db marked this pull request as ready for review October 3, 2026 16:13
xav-db added 2 commits October 3, 2026 18:33
One merge-backed value per (scope, logical index, generation) holds every
committed but unpublished vector or text operation (RecordKind 0x14).
Producers append blind insert operands and publication removes exact
operation IDs; operands, partial merges and resolved values share one
canonical encoding. A row-per-operation layout (RecordKind 0x15) is kept
for test and benchmark builds as a layout baseline. Cleanup and cursor
validation treat both key kinds as outside their lanes.
…re errors

The queue store reads and writes generation queues in the map layout, or
one row per operation in the row layout that test and benchmark builds can
select. The admission ledger charges every committed but unacknowledged
operation against per-index retained-byte and pending-member limits, and
writers rebuild it from durable queues before open returns; readers only
check the layout.

A write that would exceed a limit fails before commit with the retryable
index_backpressure code (HTTP 429, gRPC RESOURCE_EXHAUSTED); one that
exceeds a limit on its own fails with index_operation_batch_too_large.
Nothing enqueues yet: the producer and publisher follow.
xav-db added 19 commits October 3, 2026 22:02
Graph writes no longer update vector or text indexes in their own
transaction. Each write builds one complete immutable operation per changed
entity and touched generation, reserves admission capacity (1 GB / 250K
pending members per index by default) and enqueues it with the graph
change.

A publisher owned by the index worker applies queued work in bounded
batches under generation ownership and acknowledges exact operation IDs in
the same transaction. It defers Building generations and discards retired
ones. Vector entities are planned by the build's plan_and_apply, now
generic over a Build or Publication target and sharing
vector_build_cache_bytes, and commit through the vector cache fence; text
publishes through the Active text epoch path, with a per-document
admission limit.

Searches overlay unpublished work. Strong searches (the default, and every
write transaction) see all committed and local changes and fail with a
retryable 429 past 800 superseded results; read requests may choose
eventual consistency, which overlays the oldest 800 selected entities and
never fails.

Builds scan source rows only: writes made during a build queue for the
hidden generation and publish after activation, so build deltas, the
catch-up stage, ActiveVectorMutationRuntime and the synchronous text
runtime are removed. HelixDB reports queue stats and exact-ID publication
lag. The scoped search bench and gate retry backpressure and wait for the
queue to drain before measuring.

Search dispatch boxes the overlaid search futures, keeping them out of
the frames of recursive callers such as count cursors.
Publication keeps its VectorBuildSession across attempts in the
VectorBuildCache and max-min budget builds use, instead of starting every
attempt cold. Retained sessions are keyed by owner: a build operation or a
publication target (scope, index, generation).

An attempt checks out its target's session only at the exact checkpoint of
the target's latest commit (target, index record revision, publisher
sequence). Only a successful commit whose session holds no unflushed rows
retains it; every other outcome forgets it. Leases carry their own demand,
builds and publication targets have separate retention caps of 16, and
drained targets keep only budget the fair split leaves free.

Warm, cold and built graphs agree over inserts, updates, deletes, tenant
moves and partition reclaim, including warm sessions that evict, rebind and
resume under a 24 KiB budget. A warm insert reads 8 storage keys against
about 200 cold.
Behind the async-index-benchmark feature the server opens as an explicit
writer or reader with a selectable queue layout, through the same disk
cache claim and storage configuration as the product runner, and prints
one cumulative JSON sample per interval: queue backlog, publication lag,
merge cost, object-store connector I/O and SlateDB metrics. Readers open
with HelixDB::open_reader_for_server so only the server's transport reports
their queries. docker-image/build.sh --async-index-benchmark builds the
image, and CI lints and tests the feature.
The DB production-coverage gate counts only production integration
targets, so queue paths reached only by lib unit tests showed as
uncovered. production_index_operation_queue, production_queue_publication
and production_text_queue_contracts drive the producer, recovery,
publication (budgets, trims, reconciliation, retired and foreign
generations, fenced writers, uncertain commits), vector publication
planning, text publication, text build limits and search overlays through
production entry points. The vector driver's build, cleanup, adoption and
allocation-race unit tests move into a support module both targets run.
A text search that overlays pending work may search its splits more than
once as it widens past superseded results; only the first physical search
of one logical search now records split demand, so one request no longer
admits every split to the disk tier.

A traversal-restricted vector or text search whose candidates are all
superseded now returns only pending results instead of running the
physical restricted search or loading the text manifest.

Also lock in that index membership after an overlaid vector or text
search keeps exactly the rows of the per-row filter, for strong and
eventual reads and inside a write batch.
Vector Scan and text ScanSource steps scanned graph rows through their
serializable outbox transaction, so any write in the remaining source range
failed the step's commit and a backfill under steady writes retried its
first step indefinitely.

Every graph write committed after a build is created already queues a
complete operation for the hidden generation, applied in commit order after
activation. Steps therefore read source rows from a snapshot; build-owned
rows stay in the serializable read set. A blocker is durable and the queue
corrects documents, not blockers, so a step that blocks on a graph row also
reads that row through its transaction and retries if the row is repaired.

Tests commit writes inside held step windows and run 16 concurrent writers
through vector and text builds; once the queue drains, both indexes equal
cold builds of the final documents.
- Skip upserts that replay a node's exact indexed state: an entry-candidate
  row with the same vector bytes at the same deterministic layer returns
  before any delete, relink or reinsert, so replays of values a build
  already indexed stage nothing.
- Hydrate relink candidates once per deletion layer with the batched item
  loader and keep each source's nearest M with select_nth_unstable, instead
  of ~36K cache lookups and ~35 full sorts per delete. A per-candidate
  oracle pins byte-identical rows.
- Order vector planning caches by an append-only recency queue over a
  foldhash map instead of a BTreeSet with SipHash, bounded to four queue
  entries per value. Eviction, flush order and the graph are unchanged.

A pinned SHA-256 of a drained graph under warm, cold and evicting sessions
guards that caches change which rows publication reads, never a byte it
writes.
…arness

docker-image/tests/index_queue_contracts.py runs the product image through
restart and SIGKILL exactness on disk and S3, the member-limit backpressure
boundary, and the benchmark image's sample stream. The
queue_acceptance_verify example opens a fixture written by the pre-queue
release and checks graph data and strong and eventual search against exact
oracles across queued writes, publication and restart.
scripts/async-index-benchmark provisions, seeds, replays and summarizes
map-versus-rows comparisons on EC2 and S3; nothing in it runs without an
explicit, cost-approved invocation.
Read requests take search_consistency (strong by default, or eventual) in
the Rust, TypeScript, Python and Go SDKs and the OpenAPI schema, with
request-parity fixtures for both values. Write requests always search
strongly and reject eventual.
A Search consistency guide is the one place strong and eventual search are
described, with examples in every SDK and JSON. Prefiltered search explains
how it treats superseded candidates; the HTTP API page, error references
and troubleshooting cover the retryable 429 index_backpressure, including
during an index build; guarantees, limits, per-document text admission and
the Learn freshness bullet reflect queued vector/text publication. The
llms.txt files are regenerated.
Points the DB vector coverage exclusions and dispositions at the lines
they describe after the queue moved vector maintenance into publication.
…exes

Review coverage the queue lacked:
- An enqueue replayed below its acknowledgement never brings the
  operation back, in the merge algebra and through real SlateDB flushes
  and compactions.
- Queue work committed only to the WAL survives a fenced writer and a
  failed WAL upload; two tenant scopes queue, recover, and discard
  independently; a blocked text build recovers once its row is repaired.
- Every edge mutation keeps queued edge vector and text indexes exact
  against an oracle, with no row naming a dead edge.
…lushes

A queued re-embedding in one partition is an upsert. Its delete removed
every locator naming the node directly while the rows it edited stayed
cached; the reinsertion relinked many rows back to their cached
originals, so their flush staged no locator. A later delete then missed
those rows and left links to a node without an item: searches failed
with "missing simhash" and reinserting the node blocked publication.

A cached row's flush is now the only writer of neighbor rows and their
locators, staging the row's transition from the value the transaction
holds. A delete stages the node's rows absent and directly deletes only
locators no cached row's baseline links. Upserts now write only each
row's net change.

Adds a seeded concurrent soak under compaction and restarts (nightly, in
release), random vector write sequences that check every HNSW namespace
after each commit, stray-locator and bounded one-off cache runs, and a
test pinning links released versions left without a locator.
…limits

A vector Scan step that blocked behind admitted rows committed them
without advancing its cursor, so every retry after a repair failed with
invariant_violation. Such a blocker now ends the step like a full batch.

A blocked build's hidden generation publishes nothing until a retry or
abort, yet writes kept queueing there; once they filled the member or
byte limit, the repair the blocker asks for was refused with a
backpressure that never cleared. A saturated blocked build now admits
its blocker's first repair and a removal of the blocker's entity beyond
the limits, and refuses other writes with the non-retryable
index_build_blocked. Writes to entities already pending stay admitted
above the member limit; only new members are refused.

Covered in unit and production contracts, including a repair racing a
retry or abort of the build. Docs describe index_build_blocked.
Queue reads followed the whole backlog:
- Acknowledgements piled up as removals in unresolved queue values. An
  acknowledgement composed with its own enqueue now cancels it inside
  partial merges, so unresolved values stay bounded by the outstanding
  backlog. Queue format 0x14 (new on this branch) accepts an empty value
  as a partial merge result.
- Eventual searches decoded every queued operation and failed on corrupt
  records their budget never selects. They now decode only the entities
  whose latest operations fit the budget, comparing entities by raw
  bytes, and show each at its latest state or not at all.
- Every publication attempt re-read the whole queue to publish one batch.
  The writer now retains each target's remainder after a successful
  commit and continues from it.

Retained-byte ceilings whose worst-case queue value outgrows SlateDB's
u32 value length are rejected. With a merge operand pending, SlateDB
still resolves the whole value; tests pin that as a known limitation.
Strong text overlays analyzed every pending document of a partition with
no budget. They are now bounded by one text publication's analysis
charge: strong searches fail with retryable index_backpressure past it,
eventual searches keep the oldest prefix within it, and a write batch's
own documents past it fail with index_operation_batch_too_large.

One queued operation that could never fit a publication blocked every
later entity of its generation. It now holds its entity back in process
memory while the rest publish; a newer write repairs it in rotation from
full width, past one acknowledgement when needed, and a generation left
with only held entities waits for a write instead of polling. Queues
retained across uncommitted attempts share one writer-wide budget, and a
stalled retained queue reads storage before waiting.

The ledger wakes the index worker whenever an enqueue commit returns,
committed, uncertain, or cancelled, so a racing write never waits out
the stalled deadline. Health responses report the held-back entity
count. No stored format changes.
An eventual search serves unpublished changes beyond its budget as last
published, so until publication catches up it can return a node or edge
that has since moved to another tenant, changed label, or lost the indexed
property, through whole-index and prefiltered searches alike. Strong search
never does. Document this on the search consistency guide and on
SearchConsistency::Eventual.
Remaps the DB vector coverage exclusions and dispositions to the merged
vector lines: main's refreshed entries plus the branch's, each pointing
at the source line it describes.

The server fingerprint is recorded as a non-root user. Against main the
uncovered lines only shift, except the query event's backpressure message
arm, which joins its uncovered sibling arms: 102 to 103 lines.
…oss the operation queue

Version 5 marks stores that may hold asynchronous index-operation queues
(0x14). Binaries that support at most version 4 ignore those queues and
collapse their merge operands, so they must refuse an upgraded store instead
of silently diverging.

Version 5 shares version 4's physical layout. A writer upgrades a version-4
store by rewriting only the storage marker in one transaction; no index is
rebuilt. Version 2/3 stores still run the equality-bitmap migration, which now
publishes 5 directly. Current readers serve versions 4 and 5, so readers can be
upgraded before the writer, and recovery-only managed failover reports
WriterMigrationRequired for a version-4 store. The V4 cleanup-marker checks now
compare against the equality-bitmap version, so stores written by the last
version-4 release stay openable.
@xav-db
xav-db merged commit 0e05651 into main Oct 5, 2026
47 of 49 checks passed
xav-db added a commit that referenced this pull request Oct 5, 2026
…ing (#1165)

Stacked on #1156; the base branch is `async-index-queue`, and it should
be retargeted to `main` once #1156 merges. Fixes HEL-956.

## Problem
In #1156, if planning one entity's queued index operation failed the
same way every time, its whole (scope, index, generation) stopped
publishing for good. The error was retried forever, the queue filled,
writes to that index got 429, and `blocked_index_entity_count` stayed at
0.

A real trigger exists: data written by released versions (HEL-952) can
contain a self-link, and the vector insert searches never excluded the
inserting node, so planning failed every time with `ContainsOwner`.

## Fix

**1. A node is never its own neighbour.**
- Both insert searches and entry-point resolution now exclude the
inserting node.
- A self-link already in storage is dropped in memory when its row
loads, instead of failing every mutation. The row is rewritten without
it only when a mutation actually changes that row.
- Each guard has a test that fails without it, and a property test over
damaged graphs checks that no new self-link is ever created.

**2. Only the failing entity is held back.**

```text
publish attempt fails
  FailureKind::of(error)                  // exhaustive match on HelixDbError, no string matching
    Fatal         (closed / fenced)        → stop
    Transient     (I/O, conflict, cancel)  → retry the generation with backoff   (unchanged)
    Deterministic (corrupt input, invariant)
      vector → hold the entity at the failed position
      text   → halve the epoch until one entity fails alone, then hold it
      held entity: never acknowledged or dropped; counted in blocked_index_entity_count
      retried alone after 60 s, or immediately when a newer write for it arrives
      every other entity keeps publishing
```

- Held state lives in memory only. After a restart, the entity is held
back again on its first failed attempt instead of stalling the
generation.
- When several entities fail in a row, which suggests the whole
generation is damaged, attempts back off instead of spinning.

## Behaviour changes
- `blocked_index_entity_count` (in stats and health) now also counts
entities held back because their planning failed, not only entities too
large to fit a publication. The docs and the proto/OpenAPI descriptions
say so.
- A held entity that keeps being written still counts toward the index's
shared queue budget, so it can eventually cause 429s for the whole
index. This is documented, and bounding it is HEL-969.

## Testing (EC2)
- fmt and clippy (workspace) are clean.
- `cargo test -p db --lib`: 2,058 passed.
- Release production targets (queue publication, internal contracts,
text queue, index-operation queue, text correctness regressions): 92
passed.
- `cargo test -p server` passes.
- The release queue soak passes with 5 seeds.
- The new isolation tests cover:
  - a vector entity failing mid-batch;
  - text epoch splitting;
  - a transient failure, where nothing is held back;
  - a restart while an entity is held.

  They also cover repair writes and the 60 s retry.
- The text correctness regressions now expect a damaged root to hold
back only its own entity, with the search still failing closed, instead
of retrying forever.

No stored format changes. The only effect on stored data is that a
damaged row loses its self-link when a mutation rewrites it.

**Expected CI follow-up:** the DB coverage fingerprint will change. Take
the new count and sha from this PR's ARM64 "DB production thresholds"
job.


<!-- greptile_comment -->

<!-- greptile_summary -->

<p><a
href="https://app.greptile.com/api/retrigger?id=74418662"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/RetriggerDark.svg?v=2"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/Retrigger.svg?v=2"><img
alt="Retrigger"
src="https://greptile-static-assets.s3.amazonaws.com/badges/Retrigger.svg?v=2"
align="right"></picture></a><br clear="all"></p>

<!-- greptile-risk -->

The PR should wait for the production coverage baseline update so its
required CI gate can pass.

<h2>Findings</h2>

1. <img alt="P1"
src="https://greptile-static-assets.s3.amazonaws.com/badges/p1.svg?v=9"
align="top">&nbsp;**Coverage baseline remains unchanged** <a
href="https://github.com/HelixDB/helix-db/pull/1165#discussion_r4178564885">▶</a>

<h3>Summary</h3>

The PR makes queued index publication hold back an entity whose planning
fails deterministically while continuing to publish other entities. It
also prevents HNSW inserts from choosing their own node as a neighbor
and tolerates stored self-links during mutation.
- Vector failures identify the failing effect; text failures narrow an
epoch until one entity remains.
- Failed entities remain queued and become eligible for a timed retry or
a newer write.
- The production coverage baseline still needs updating for the changed
source.

<details><summary>Diagram</summary>

```mermaid
%%{init: {'theme': 'neutral'}}%%
flowchart TD
  A[Select queued entities] --> B[Plan publication]
  B -->|Success| C[Commit effects and acknowledge]
  B -->|Transient failure| D[Retry batch with backoff]
  B -->|Deterministic vector failure| E[Hold identified entity]
  B -->|Deterministic text failure| F[Narrow epoch]
  F -->|One entity fails| E
  E --> G[Publish other entities]
  E --> H[Retry on newer write or timer]
```
</details>

<sub>Reviews (1) · Last reviewed commit: ["docs: count failed-planning
holds in
blo..."](https://github.com/helixdb/helix-db/commit/8708fe2ec237f7fbe545fd150b7657983a089224)</sub>

<!-- /greptile_comment -->
matthewsanetra added a commit that referenced this pull request Oct 5, 2026
Release PR for **CLI 3.4.3** and **Docker image v0.0.10**, following the
same shape as #1161.

```diff
 crates/cli/Cargo.toml           version 3.4.2 → 3.4.3 (and Cargo.lock)
 crates/cli/src/config.rs        DEFAULT_LOCAL_IMAGE_TAG v0.0.9 → v0.0.10
 CLI tests, README, CLAUDE.md    v0.0.9 → v0.0.10
 docs                            default image, release notes, storage v5 warning, llms-full.txt
 docker-image/README.md          published release v0.0.9 → v0.0.10
 CONTRIBUTORS.md                 CLI image v0.0.9 → v0.0.10
 sdks/tests/parity/COVERAGE.md   published-image parity example v0.0.5 → v0.0.10
```

Version-history references stay as written:
- "v0.0.9 and earlier refuse to open" in the docker-image README's index
storage version 5 section.
- "Upgrading S3 deployments from v0.0.8 and earlier".
- The "Docker images through v0.0.9" locator comments in `crates/db`.
- The v0.0.5 compatibility statements.

The TypeScript README's `package-smoke` commands stay on v0.0.5, because
they mirror CI's digest-pinned compatibility check in
`typescript-sdk.yml`.

The `COVERAGE.md` parity pin was bumped together with the CLI default in
v0.0.5 (`620fb543`) and missed afterwards. v0.0.5 can't serve the
`search_consistency` parity fixtures added in #1156.

## Order

1. Merge this. The `Docker image` workflow's publish job requires
`DEFAULT_LOCAL_IMAGE_TAG` on main to equal `release_version`, so the
image can only be published after this lands.
2. Dispatch the `Docker image` workflow from main with
`release_version=v0.0.10`. It reruns the full suite and publishes only
if it passes.
3. Once v0.0.10 is published, run `cli.yml` from main to tag and publish
v3.4.3.

## What's in v0.0.10 / 3.4.3

- ⚠️ **Index storage version 4 → 5** (#1156). The first writer open
rewrites only the version marker. After that, v0.0.9 and earlier refuse
the database with `unsupported_index_storage_version`.
  - Upgrade readers before the writer.
  - Back up first if a rollback may be needed.
- The local server guide now carries this warning next to the "set
`tag`" instructions.
- **Asynchronous vector and text index publication** (#1156).
- Writes commit the graph change plus a queued index operation, and a
background worker publishes it.
- `search_consistency: strong | eventual`. Strong is the default and
exact.
- `index_backpressure` (429, retryable) is returned in three cases: past
1 GB or 250K pending entities per index, when a whole-index strong
search is behind more than 800 changes, and when strong text analysis
would exceed one publication's budget.
- `index_operation_batch_too_large` (400) is returned when one write
stages more than 8 MiB for one index.
  - Text admission limits are now per document.
  - Builds no longer livelock under concurrent writes.
- **Failing entries are held back** (#1165).
- An entry whose index update fails deterministically no longer stalls
its whole generation. It is counted in `blocked_index_entity_count` and
retried about once a minute.
  - Insert searches never pick the inserting node as its own neighbour.
- **Fixes the v0.0.9 `missing simhash` known issue** (#1156 C1, #1165).
Re-embeds keep their HNSW reverse-link locators. Indexes damaged by
earlier versions still need a drop and recreate (HEL-952).
- **Range-driven counts apply every filter** (#1169). For example, 149 →
49.
- **CLI 3.4.3:** defaults to v0.0.10. There are no other CLI changes
since 3.4.2.
- **Not in this release:** the SDK changes in #1156, #1163 and #1164
ship with the next SDK releases. The release notes say to send
`search_consistency` over HTTP until then. #1166 is a docs dependency
bump.

## Testing

- `cargo check --workspace --locked` passes, so the hand-edited
`Cargo.lock` is consistent.
- `cargo test --locked -p helix-cli` passes: lib 241, `e2e_cli` 13,
`runtime_commands` 20, `typescript_runtime`, and the other targets. The
Docker-only `e2e_runtime` tests are ignored by design, because they need
the unpublished v0.0.10 image.
- rustfmt 1.9.0 (the pinned 1.97.1 toolchain) `--edition 2024 --check`
is clean on the edited Rust files.
- Docs: `check-docs`, `generate-llms --check`, `generate-llms-full
--check` and `check-openapi` pass.
- Grep: every `helixdb/helixdb:v0.0.x` reference outside release notes
is v0.0.10, except the CI-pinned v0.0.5 TypeScript smoke.
- Workspace clippy is left to CI. The local nix nightly clippy rejects
the repo's clippy configuration before linting, independent of this
change.


<!-- greptile_comment -->

<!-- greptile_summary -->

<p><a
href="https://app.greptile.com/api/retrigger?id=74957602"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/RetriggerDark.svg?v=2"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/Retrigger.svg?v=2"><img
alt="Retrigger"
src="https://greptile-static-assets.s3.amazonaws.com/badges/Retrigger.svg?v=2"
align="right"></picture></a><br clear="all"></p>

<!-- greptile-risk -->

The PR appears safe to merge, though the published-image parity
instructions should include an image pull step.

<h3>Summary</h3>

This release PR updates the CLI default to Docker v0.0.10, bumps the CLI
to 3.4.3, and updates image references and release guidance. The
published-image parity instructions need a pull step before their new
image pin can work without a local copy.

<sub>Reviews (1) · Last reviewed commit: ["chore(release): CLI 3.4.3 and
Docker
ima..."](https://github.com/helixdb/helix-db/commit/281197c0c91e5f72a7ce89fb8a248e04176cdee4)</sub>

<!-- /greptile_comment -->

## Also: fix main's workspace tests

`counts_over_range_intersections_apply_every_filter` (#1169) overflows
the default 2 MiB test-thread stack in Linux debug builds. That fails
`Workspace tests`, `DB all targets` and `Workspace and tooling quality`
on main, and the Docker release's `quality` job runs the same tests.
`92c15755` runs it through the file's existing 16 MiB
`run_high_stack_contract`, like the other heavy contracts. It passes
locally even with `RUST_MIN_STACK=524288`.
xav-db added a commit that referenced this pull request Oct 6, 2026
…log (#1175)

First PR of the queue stack: the HEL-957 PR (streamed backlog load at
open), then the HEL-958 PR (admission memory charge), build on this one.

Fixes HEL-961.

## Problem
Draining a publication queue cost more than the size of its backlog:
- Every attempt regrouped the whole retained queue by entity, and every
commit rescanned it to remove the acknowledged operations. With
512-entity batches that is quadratic: going from a 100K to a 250K
backlog (2.5×) cost 7.6× the bookkeeping.
- A hold (an entity too large for a batch, or one whose planning failed)
dropped the queue. The next attempt then re-read and decoded the whole
backlog: 136 reads for 64 failing entities at 250K.
- A retired generation's discard re-read the whole queue for every
acknowledgement chunk: 63 reads at 250K with small acknowledgements.
That also kept the generation's charges, which a recreated index shares,
held for longer.
- A target's schedule was removed only when a read found its queue
empty, which never happens once a commit drains it. So `schedules` kept
one entry for every target and generation ever published.

## Fix
All changes are in memory.

**1. The queue is grouped once per read** (`StoredQueue` in
`queue/storage.rs`).

```text
read order   0:a1  1:b1  2:a2  3:c1  4:b2
chains       a: 0 -> 2   b: 1 -> 4   c: 3
order        {0: a, 1: b, 3: c}          // entities by their oldest outstanding op
without(a1)  a: 2        order {1: b, 2: a, 3: c}
```

- `rotation(after)` yields entities in the same order as the old
regroup, at one map step per entity visited.
- `without(acked)` pops chain heads in O(acked · log entities), with no
rescan. It asserts that every acknowledgement names an entity's oldest
operations in order.
- `HeldEntity::repair_width` is O(1): a newer operation is queued
exactly when the newest one isn't the `through` that is known not to
publish.

**2. Holds keep the queue.** A size block or a `FailureKind` isolation
commits nothing, so the queue is retained instead of being read again. A
retained queue whose entities are all `Waiting` or `Failed` is re-read
once the target's latest admission has moved past the one seen before
the read. So retries that keep falling due still never pin publication
to a stale queue (the 61240f7 fix).

**3. A discard continues from the retained queue.** It acknowledges
whole-entity prefixes and keeps the rest, so a retired backlog is read
and decoded once. Charges are released at each commit.

**4. A drained target's schedule is dropped.** Only its latest vector
commit is kept, in a ring of the 16 most recently drained targets
(`MAX_RETAINED_PUBLICATIONS`). This lets its retained planning session
still serve its next write.

The other #1156 and #1165 invariants are unchanged:
- the indexer is the only writer of index rows, and the ack is in the
same transaction;
- exact-ID enqueue and ack;
- held-entity states, the 60 s retry, and `blocked_index_entity_count`.

## Behaviour/config changes
- None visible to users or in config.
- Fewer queue reads during holds and discards, and a recreated index
gets its shared queue budget back sooner after a drop.

## Measurements (release, c7i.8xlarge, in-process, 2-dim vectors,
512-entity batches; main → branch)

**Bookkeeping only** (selecting a batch and removing what was
acknowledged):

| Backlog | main | branch |
|---|---|---|
| 10K | 0.011 s | 0.003 s |
| 100K | 1.23 s | 0.047 s |
| 250K | 9.36 s | 0.157 s |
| 1M | — | 0.856 s |

**Publish drain:**

| Backlog | main | branch |
|---|---|---|
| 10K | 5.47 s | 5.27 s |
| 100K | 60.3 s | 59.3 s |
| 250K | 176.2 s (1,418 ops/s) | 154.9 s (1,614 ops/s), −12% |

HNSW planning takes most of the time.

**64 entities that fail to plan, then a drain to Stalled** (time,
reads):

| Backlog | main | branch |
|---|---|---|
| 10K | 7.69 s, 67 reads | 7.19 s, 2 reads |
| 100K | 73.0 s, 80 reads | 66.1 s, 2 reads |
| 250K | 185.0 s, 136 reads | 161.1 s, 2 reads |

At 250K, time spent reading falls from 12.4 s to 0.28 s.

**Retired discard with a 64 KiB ack budget** (time, reads):

| Backlog | main | branch |
|---|---|---|
| 10K | 0.31 s, 4 reads | 0.31 s, 2 reads |
| 100K | 2.57 s, 26 reads | 2.57 s, 2 reads |
| 250K | 12.68 s, 63 reads | 6.52 s, 2 reads |

With default acks at 250K, the discard goes from 5 reads to 2 and from
0.96 s to 0.79 s.

## Testing (EC2)
- fmt, clippy (workspace, all targets) and `cargo check --locked
--workspace --all-features` are clean.
- `cargo test -p db --lib`: 2,069 passed. `cargo test -p db --doc`: 40
passed.
- `cargo test -p server` passes (71 unit, 2 integration, 5 doc tests).
`scripts/validate-cargo-target-references.py` passes.
- Release production targets, with
`production-coverage,index-lifecycle-testing`: 185 passed. These are
queue publication, index-operation queue, text queue, text correctness
regressions, internal, contracts, migration, and index lifecycle (run
single-threaded, as CI runs it).
- Product image: `index_queue_contracts.py --scenario all` passes.
Indexes are exact after restart and after SIGKILL, on local disk and on
S3. Writes get a 429 at 250,000 members and are accepted 1.4 s after
activation.
- New `grouping_tests.rs`: a property test (1,024 cases × up to 95
steps). It runs random enqueues, rereads, publishes, every kind of hold,
releases, discards and cursor moves, and checks the result against a
verbatim copy of the old regrouping. Selections, remaining queues,
rotation order and retained bytes must all be identical. There are also
should-panic tests for out-of-order and duplicate acks.
- New `drain_tests.rs` covers:
  - retiring mid-drain;
- recreating an index whose old backlog is exactly at the limit (429
until the first discard commits, then admitted);
  - schedule cleanup across 24 targets;
  - restarting mid-drain with updates and deletes;
  - concurrent writes during a drain;
  - holds reusing one read.
- The #1165 isolation tests pass unchanged.

No stored format changes.

**Expected CI follow-up:** the DB coverage fingerprint will change. Take
the new count and sha from this PR's ARM64 "DB production thresholds"
job.


<!-- greptile_comment -->

<!-- greptile_summary -->

<p><a
href="https://app.greptile.com/api/retrigger?id=75528033"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/RetriggerDark.svg?v=2"><source
media="(prefers-color-scheme: light)"
srcset="https://greptile-static-assets.s3.amazonaws.com/badges/Retrigger.svg?v=2"><img
alt="Retrigger"
src="https://greptile-static-assets.s3.amazonaws.com/badges/Retrigger.svg?v=2"
align="right"></picture></a><br clear="all"></p>

<!-- greptile-risk -->

The PR appears safe to merge, with two non-blocking memory and
planning-cache efficiency issues worth addressing.

<h3>Summary</h3>

The PR groups each queue once per read, advances retained queues by
acknowledged entity prefixes, reuses queues across holds and
retired-generation discards, and bounds schedule history for drained
targets. It adds property, drain, and production-support coverage. Two
non-blocking resource-efficiency concerns remain: retained allocation
size during a long drain and unusable commits entering the
drained-target ring.

<details><summary>Diagram</summary>

```mermaid
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  A[Read durable queue] --> B[Group operations by entity]
  B --> C[Select ordered prefixes]
  C --> D{Attempt outcome}
  D -->|Commit| E[Acknowledge IDs and retain remainder]
  D -->|Hold or definite conflict| F[Retain unchanged queue]
  D -->|No selectable work| G[Drop queue and wait]
  E --> C
  F --> C
```
</details>

<sub>Reviews (1) · Last reviewed commit: ["test(db): cover drains across
retirement..."](https://github.com/helixdb/helix-db/commit/4288421a163fb46485908c11b6d29277759bb7a3)</sub>

<!-- /greptile_comment -->
xav-db added a commit that referenced this pull request Oct 6, 2026
Fixes HEL-957.

## Problem
At writer open, `load_backlog` fully decoded every queue of a scope
(every payload, including 1536-dimension vectors) into one `Vec` before
charging any of it to the admission ledger. So open's peak memory and
time grew with the whole durable backlog: about 25 s and 4.5 GB for 2.26
GB of queued work. Each scope also cost about 4 scans (scope seek,
record scan, layout probe, queue scan).

## Fix

```text
before: for scope: decode every queue of scope -> Vec -> charge all
after:  one forward scan per keyspace (legacy, then tenants)
          row -> QueueDiscovery::push -> completed queue? -> validate owner -> charge -> drop
```

- **Streaming discovery.** `QueueStore::discover` is replaced by
`QueueDiscovery`, which takes rows in key order and returns each queue
as soon as it is complete. In the map layout each row is a whole queue.
In the row layout a generation completes when the next generation's
first row arrives, or at `finish()`. Open now holds one queue at a time
plus the ledger.
- **Framing only.** `OperationFrame::decode_queue` / `decode_row` read
each operation's ID, entity and retained bytes. Payloads are validated
in place (`algebra::validate_body`) and never decoded. The values and
rows they reject are exactly the ones `OperationQueue::decode` /
`QueueRow::decode` reject (duplicate IDs, NaN, UTF-8, trailing bytes,
unresolved values). A proptest checks this, and `OperationQueue::decode`
and `OperationFrame::decode_queue` share one record walk
(`decode_unique`), so their duplicate-ID check cannot drift.
- **Fewer scans.**
  - One scan covers the legacy scope.
- One forward-only iterator covers the tenant keyspace, with at most 2
seeks per tenant.
- Index records are read only for scopes that have a queue, and are
cached per scope.
- **Layout check folded in.** `discovery_range` spans both layouts'
adjacent record kinds (a const assert checks this). A queue of the other
layout still fails closed with `Config`, during the same scan. The
separate probes (`require_scope_layout`, `require_layout`,
`discover_scopes`) are now compiled only for tests and the benchmark
feature.
- **Ledger only.** Open fills the admission ledger and nothing else:
publication starts with no retained queue, schedule or held entity, and
each target's first attempt reads and groups its own queue
(`StoredQueue`).
- **Removed:** `QueueStore::discover`, `finish_rows` and
`OperationQueue::from_rows`. `OperationQueue::operations()` is now
compiled only for tests and `index-lifecycle-testing`.
- **Unchanged:**
  - Charge order (legacy first, then tenants ascending).
  - Owner validation.
  - The row-layout sequence resume.
  - The worker wake.
- The #1156/#1165 invariants: the indexer is the only writer of index
rows, and held-entity state starts empty either way.

## Behaviour changes
- A corrupt `IndexRecord` in a tenant that has **no** queue no longer
fails writer open. That tenant's first catalog read fails instead
(`Encoding`), and the other tenants keep working. A corrupt record in a
scope that has a queue still fails the open.
- No config or API changes.

## Measurements (EC2 c7i.8xlarge, release)
Fixture: 6 legacy vector queues of 256 MiB, 8 tenant queues of 64 MiB
(dropped indexes), and 2000 tenants with rows but no queue. That is 14
queues, 366,930 ops and 2.26 GB retained, about 6 GB on disk. Each run
starts from a fresh copy of the fixture, 4 runs each, and every run
loaded the same totals.

| Phase | base | this PR |
|---|---|---|
| Backlog load only (no compactor) | 15.1–15.3 s, peak 2.17–2.22 GB |
4.63–4.86 s, peak 1.14 GB (flat) |
| Full writer open (compactor running) | 24.4–24.8 s, peak 4.33–4.65 GB
| 4.83–5.01 s, peak 1.22–1.30 GB |

- Open is about 5× faster.
- The base column was measured on main. `linear-queue-publication-drain`
does not touch the open-time load path (`recovery.rs`, queue discovery),
so it applies unchanged.
- Without the compactor, this PR's peak scales with the largest queue:
about 0.85 GB at 128 MiB queues, 1.14 GB at 256 MiB. The base's peak
scales with the whole backlog.
- Open peaks above the load-only peak come from SlateDB's background
compactor running during open, not from the load.

## Testing (EC2)
- `cargo fmt --all --check`, `cargo clippy --workspace --all-targets --
-D warnings` and `cargo check --locked --workspace --all-features` are
clean.
- `cargo test -p db --lib`: 2,082 passed. `cargo test -p db --doc`: 40
passed.
- Release production targets (queue publication, index-operation queue,
text queue, text correctness regressions, internal, contracts,
migration, index lifecycle), one test thread each: 186 passed.
- `cargo test -p server`: 79 passed.
- `scripts/validate-cargo-target-references.py` passes.
- Product image, `docker-image/test.sh` and `index_queue_contracts.py
--scenario all` pass. After SIGKILL and restart, 10,750 docs are exact
on local disk and on S3. The member limit rejects at 250k and accepts
again 1.4 s after the build activates.
- New `recovery_tests` compare every load against a reference loader
that fully decodes each queue the old way. They cover:
- legacy and tenants 0, 7, 8 and `u128::MAX`, with keys on both sides of
each range;
- the map layout across the compacted run, L0 and the memtable, with
partial and full acks;
- the row layout, with interleaved generations and the sequence resume;
  - the other layout failing closed in every scope;
- corrupt payloads, duplicate IDs, foreign keys and invalid tenant
envelopes failing closed;
- misowned tenant queues after other scopes' records are cached (this
test fails if the per-scope cache goes stale);
- an SST read failure injected at every point during open, after which
every reopen reloads the same ledger and writes nothing;
- a full writer restart that reloads the ledger exactly, starts with no
queue reads, retained bytes, schedules or blocked entities, and drains
each reloaded target in exactly 2 reads.
- Production contracts add a tenant scope-walk test and duplicate-ID /
corrupt-payload cases.
- `open_measurement_tests` is an ignored, manual release benchmark that
produced the numbers above.

No stored format changes.

**Expected CI follow-up:** the DB coverage fingerprint will change.
Rows-layout-only lines are uncovered in production targets, and line
numbers shift. Take the new count and sha from this PR's ARM64 "DB
production thresholds" job, after `linear-queue-publication-drain`
merges and this PR is retargeted to main.
xav-db added a commit that referenced this pull request Oct 6, 2026
…ext analyses (#1176)

Fixes HEL-960.

## Problem
A strong text search overlays every committed but unpublished document
of its partition so its BM25 answer stays exact. On main, and unchanged
by the base PR:
- it was bounded by one text publication's analysis budget (64 MiB,
about 5,500 pending 40-word documents). Any ingest burst past that made
strong text searches fail with retryable 429
`pending_text_analysis_bytes` until the index worker caught up (106 and
234 rejections in the 20K and 40K burst runs below);
- every search analyzed the whole backlog again from scratch and copied
it into an in-memory Tantivy index, so cost and memory grew with the
backlog and with each concurrent search (about 200 ms and 1 GB RSS at 5K
pending with 1 searcher, 1.7 GB with 8).

## Fix

```text
strong text search
  pending set (per request, from #1174)
  for each pending doc in the partition
    cached analysis (by queued op id)?  reuse it, charged as if fresh
    else  wait for the analysis turn -> analyze on the blocking pool -> cache
    charge > strong_text_search_max_analysis_bytes  ->  429 pending_text_analysis_bytes
  score pending docs in place (Tantivy Bm25Weight, best k)  ->  merge with published results
```

**1. Strong text searches have their own bound.**
- `IndexOperationQueueTuning::strong_text_search_max_analysis_bytes`
(default 512 MiB, about 44K pending 40-word documents), next to the
base's strong vector bound. Past it, the search fails with the same
retryable 429 `index_backpressure`, resource
`pending_text_analysis_bytes`.
- Applies to read and write requests. A write's own documents are
charged first; alone past the bound they fail the write with
`index_operation_batch_too_large`, as before.
- Eventual searches are unchanged: they overlay at most one publication
budget (64 MiB) and serve the rest as last published.

**2. Analyses of pending text are reused (in memory only).**
- `PendingTextAnalyses`, per database, keyed by (queue target,
partition, queued operation ID). Selection reads the base's per-request
decoded queue, where each committed entity now carries its queued
operation ID.
- It is exact without invalidation: a queued payload never changes and
its random 121-bit ID is never reused, so publication, supersession, a
rebuild or a drop can only make an entry unreachable. A reused analysis
is still compared with the queued text and fails closed on a mismatch.
- A reused analysis is charged exactly as a fresh one, so a 429's
`requested` value is the same cold or warm.
- Capped at the bound; the least recently used partitions give up only
the bytes a replacement needs.
- Released when the publisher drains or discards a generation, and when
a strong search finds its queue empty (the only path on reader handles).
- A write transaction's own text is never cached.

**3. Pending documents are scored in place.**
- No clone and no in-memory index. Scoring uses Tantivy's `Bm25Weight`
with the same corpus statistics and fieldnorm quantization, ranked by
score then entity ID.
- Scoring runs on the blocking pool through the base's `run_blocking`,
in steps of 4,096 documents with a request check between steps, keeping
only the best k per step, so memory grows with k, not with the number of
hits.
- A proptest against a real Tantivy index requires bit-exact results for
queries of up to 2 terms. For 3 or more terms results must match within
rounding, because Tantivy's own union sums in an order that depends on
index layout.

**4. Only one strong search at a time analyzes uncached text.**
- A process-wide first-come-first-served turn. A search that waited
reads the cache again and reuses what the previous holder analyzed.
- Cold analysis runs through `run_blocking` and checks the request
before each fresh document. The turn is an owned guard moved into that
blocking work and released only when it stops, so an abandoned request
cannot let a second cold analysis overlap.
- A pass that stops early (deadline, cancellation, reader retirement)
still caches what it analyzed plus the cached entries after it, so the
next search resumes instead of starting over.
- Fully cached strong searches and eventual searches never wait.
- However many searches run at once, analyses of pending text stay at
about two bounds: one cached and one being analyzed.

**5. Observability.**
`IndexOperationQueueStats::strong_text_search_rejections` counts
refusals. A warning naming the setting is logged at most once a minute.

#1156/#1165 invariants are unchanged: the indexer is still the only
index writer, acks stay in the same transaction, strong search stays
exact, and held-back entities still count toward the bound.

## Behaviour/config changes
- New server env var `HELIX_STRONG_TEXT_SEARCH_MAX_ANALYSIS_BYTES`: a
positive byte count, default 512 MiB. It feeds the same tuning builder
as the base's vector variable. An invalid value fails startup with
`IndexQueueBytes` naming the variable, before any cache directory is
created. Benchmark builds keep it. Embedded users set it with
`IndexOperationQueueTuning::with_strong_text_search_max_analysis_bytes`.
- Strong text 429s now start at 512 MiB of pending analysis per
partition instead of 64 MiB. The resource name is unchanged.
- While text is pending, the database keeps up to one bound of cached
analyses, plus about one more while a cold search analyzes.
- New stats field `strong_text_search_rejections`. On reader handles it
is the only nonzero field.
- Docs updated: search-consistency (one table for both strong bounds),
troubleshooting, error-handling, prefiltering, Cloud limits,
local-server, docker-image README, OpenAPI 429 description,
llms-full.txt.

## Measurements
EC2 c7i.8xlarge.

Ingest burst, product benchmark image: 100 docs/s plus one burst, 2
strong searchers, 90 s. Latency counts successful searches only, so
main's numbers leave out the searches it refused, which were the ones
behind the largest backlogs. Measured before the restack onto the base
PR.

| Burst | Image | Strong-search 429s | p50 / p99 (ms) | Peak container
memory |
|---|---|---|---|---|
| 20K | main | 106 of 1,705 | 50.3 / 129.4 | 1,716 MiB |
| 20K | branch | **0** of 1,580 | 58.5 / 139.5 | 1,271 MiB |
| 40K | main | 234 of 1,598 | 58.7 / 166.1 | 2,002 MiB |
| 40K | branch | **0** of 1,335 | 77.3 / 223.0 | 1,287 MiB |

The 40K burst peaked at about 279 MiB of pending analysis.

In-process probe, release build: 20K published docs, then N pending with
publication paused. Every document and query contains a common term, so
no pending document is filtered out, which is the worst case. Values are
strong search p50 / p99 (ms) and process RSS. `*` marks values measured
before the restack and not re-run.

| Pending | Searchers | main | branch p50 / p99 | branch RSS |
|---|---|---|---|---|
| 0 | 1 | 13.2 / 18.5 | 12.8 / 18.6 * | |
| 5K (59 MiB) | 1 | 204 / 217, 1,062 MiB | 27 / 39 | 322 MiB * |
| 24K (282 MiB) | 1 | 429 | 113 / 439 * | 689 MiB * |
| 40K (470 MiB) | 1 | 429 | 208 / 299 | 813 MiB * |
| 5K | 8 | 261 / 293, 1,668 MiB | 46 / 63 | 376 MiB * |
| 40K | 8 | 429 | 248 / 401 | 1,529–1,680 MiB peak |
| 40K | 32 | 429 | 904 / 1,152 * | 1,860 MiB peak * |

- Versus the pre-restack build on the same host in the same session:
single-searcher latency is the same; at 8 searchers p50 is up to about
8% higher and p99 the same or lower. No search was rejected.
- RSS at 40K with 8 searchers is higher than before the restack
(1,118–1,192 MiB): glibc keeps per-thread arenas for the extra
blocking-pool threads. With `MALLOC_ARENA_MAX=2` both builds sit at
about 790 MiB. The multi-searcher RSS values marked `*` are likely
higher for the same reason.
- Eventual search at 5K pending with 1 searcher: 184 → 27 ms p50 *,
because it also scores in place now.
- A warm strong search still costs O(backlog) CPU: one statistics-marker
read per pending entity, plus scoring. The statistics pass and the
scope/term filter stay on the async workers, one cheap pass over the
backlog each. At 32 concurrent searches and 40K pending this saturates
the 32 cores.

## Testing (EC2)
- `cargo fmt --all --check`, `cargo clippy --workspace --all-targets --
-D warnings` and `cargo check --locked --workspace --all-features` are
clean. `validate-cargo-target-references.py` passes.
- `cargo test -p db --lib`: 2,097 passed (26 more than the base). `cargo
test -p db --doc`: 40 passed.
- Release production targets
(`production-coverage,index-lifecycle-testing`, `--test-threads=1` as CI
runs them): 187 passed: contracts 58, internal contracts 41, migration
contracts 27, text correctness regressions 23, index-operation queue 15,
text queue contracts 12, index lifecycle contracts 8, and queue
publication 3.
- `cargo test -p server` passes (72 unit tests plus the integration and
doc targets), as does `--features async-index-benchmark` (76 unit
tests).
- Product image: `docker-image/test.sh` passes, and
`index_queue_contracts.py --scenario all` passes. Correctness is exact
after restart and SIGKILL on disk and on S3 (10,750 docs each). In the
member-limit scenario, the batch past 250,000 members gets a 429 and the
retried batch is accepted 2.5 s after activation.
- Docs: `check-docs`, `generate-llms --check`, `generate-llms-full
--check` and `check-openapi` pass.
- New coverage:
- Cache: replacement keeps exactly the latest selection; exact-budget
and one-byte-over boundaries; partial LRU shedding across partitions;
forget per target; a search holding an old snapshot still reads it.
- Selection: identical charges cold and warm; refusals report the same
`requested` value cold and warm, at the text charge and at a token;
refused searches still cache the prefix they analyzed; a write's own
text is never cached; a cached text mismatch fails closed; committed and
local entities keep their order and source.
- Turn: a cold search waits, then reuses the holder's analyses. Cached
searches, eventual searches and searches of another partition do not
wait. A search over a write's own text does. A holder whose request is
dropped, retired or expires mid-analysis keeps the turn until its
blocking work stops; it caches its prefix plus the cached tail, and the
next search finishes without analyzing any document twice.
- Scoring: proptest against a real Tantivy index covering fieldnorm
quantization, ties, restricted scopes, empty documents and extra corpus
statistics. Stepped scoring keeps the k best across steps and checks the
request at each step.
  - Overlay:
- exact past the publication budget, across a restart and after
publication;
    - the exact bound passes and one byte over fails, cold and warm;
    - a write batch sees its own staged text despite a warm cache;
    - updates, deletes and re-inserts;
- concurrent strong searches racing writes and automatic publication;
    - the rejection counter.
- Release: the publisher keeps the cache while work remains and drops it
once drained. Dropping a text index releases its analyses. A strong
search over an empty queue releases them on writer and reader handles;
with that release removed, both tests fail.
  - Production contract (`production_text_queue_contracts`):
    - the default bound stays exact past the publication budget;
- a lowered bound gives a retryable 429 in read requests, through
`HelixQueryService` (as Backpressure), and in write requests, which roll
back;
    - eventual results are unchanged, and publication clears the 429;
- a write's own text past the bound gets `IndexOperationBatchTooLarge`.
- Server config: one test covers each strong bound variable and both
together. Values 1, 2^31 and u64::MAX are accepted. 0, -1, empty,
`512MiB` and u64::MAX+1 are rejected before any cache directory is
created. Benchmark builds keep both bounds.

No stored format changes.

**Expected CI follow-up:** the DB coverage fingerprint will change. Take
the new count and sha from this PR's ARM64 "DB production thresholds"
job.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant