Repository navigation
fix(db): hold back only the index entity whose publication keeps failing - #1165
Merged
Merged
Conversation
A link that released versions left without a reverse locator outlives its target's delete, so re-embedding that node searched back to it through the link and selected it as its own neighbor, failing planning with ContainsOwner on every attempt. Both mutation searches and the beam's entry resolution now skip the inserting node. A neighbor row that already links its own node failed every mutation that loaded it. The deployed-row adapter now drops the self-link in memory; a mutation that changes the row stores it self-free. No codec or stored format changes.
Planning one queued operation could fail the same way on every attempt, for example on damaged graph rows. Publication retried the whole batch forever, so its generation stalled, the queue filled, writes failed with backpressure, and no entity was reported as blocked. Publication failures are now classified by FailureKind. A deterministic failure planning one entity holds that entity back through the existing held-entity state: the vector planner reports the failing effect's position, and a failed text epoch halves its entity ceiling until one entity fails alone. The rest of the generation keeps publishing, the held operations stay queued, and a newer write to the entity, or a restart, plans it again. Transient failures still retry the batch with backoff, and a deterministic failure outside one entity's planning still retries the generation. No stored format changes.
The self-link fix moved lines in the HNSW mutation module, so the line-anchored vector coverage dispositions and exclusions there now name the same source lines at their new positions.
Removing the greedy descent's guard or either beam-root guard left every vector test passing. Stale metadata naming the inserting node now has an end-to-end case, and each guard a direct case that first shows the same traversal reaching the node when another node inserts.
…back Every isolation case planned the failing entity first, so holding back the batch's first entity instead of the failing one went unnoticed. The failing update now follows two inserts that plan in the same batch.
A failure outside an entity's own input, such as missing namespace metadata, held back every queued entity one immediate attempt at a time, and nothing planned them again until a write or a restart, which a held delete never gets. A failure hold now repairs its entity again once MAX_STALLED_WAIT passes, a stalled generation wakes at the first such retry, and from the second hold in a row without a publication the next attempt backs off. The docs now say a held entity's writes keep using the index's shared queue capacity, and llms-full.txt is regenerated.
The health field also counts entities whose planning fails the same way every time, but its OpenAPI, proto, and server descriptions still named only operations that cannot fit a publication.
| @@ -1666,13 +1991,88 @@ async fn load_generation_record( | |||
| #[path = "../../../tests/production_support/queue_publication.rs"] | |||
Contributor
There was a problem hiding this comment.
Coverage baseline remains unchanged
These new publication branches change the uncovered non-vector source lines used by the production coverage job, but this PR does not update the tracked count and SHA-256 in db-production-coverage-baselines.json. The job requires both values to match its generated report, so the coverage gate will fail until the baseline is updated.
A hold after failed planning kept the attempt's queue for the next attempt, against the rule that a blocked outcome drops it. With enough entities whose retries keep falling due, no attempt stalled, storage was never read again, and newer writes, including repairs of held entities, never published. A hold now drops the queue like a size-blocked one; a trimmed text epoch still keeps it.
…window Judge eligibility as of an instant before the attempt rescheduled and bound the retry instant by the attempt's start and return, so a slow host cannot let the 20 ms backoff pass before the check.
|
Preview deployment for your docs. Learn more about Mintlify Previews.
💡 Tip: Enable Automations to automatically generate PRs for you. |
matthewsanetra
added a commit
that referenced
this pull request
Oct 5, 2026
Release PR for **CLI 3.4.3** and **Docker image v0.0.10**, following the same shape as #1161. ```diff crates/cli/Cargo.toml version 3.4.2 → 3.4.3 (and Cargo.lock) crates/cli/src/config.rs DEFAULT_LOCAL_IMAGE_TAG v0.0.9 → v0.0.10 CLI tests, README, CLAUDE.md v0.0.9 → v0.0.10 docs default image, release notes, storage v5 warning, llms-full.txt docker-image/README.md published release v0.0.9 → v0.0.10 CONTRIBUTORS.md CLI image v0.0.9 → v0.0.10 sdks/tests/parity/COVERAGE.md published-image parity example v0.0.5 → v0.0.10 ``` Version-history references stay as written: - "v0.0.9 and earlier refuse to open" in the docker-image README's index storage version 5 section. - "Upgrading S3 deployments from v0.0.8 and earlier". - The "Docker images through v0.0.9" locator comments in `crates/db`. - The v0.0.5 compatibility statements. The TypeScript README's `package-smoke` commands stay on v0.0.5, because they mirror CI's digest-pinned compatibility check in `typescript-sdk.yml`. The `COVERAGE.md` parity pin was bumped together with the CLI default in v0.0.5 (`620fb543`) and missed afterwards. v0.0.5 can't serve the `search_consistency` parity fixtures added in #1156. ## Order 1. Merge this. The `Docker image` workflow's publish job requires `DEFAULT_LOCAL_IMAGE_TAG` on main to equal `release_version`, so the image can only be published after this lands. 2. Dispatch the `Docker image` workflow from main with `release_version=v0.0.10`. It reruns the full suite and publishes only if it passes. 3. Once v0.0.10 is published, run `cli.yml` from main to tag and publish v3.4.3. ## What's in v0.0.10 / 3.4.3 -⚠️ **Index storage version 4 → 5** (#1156). The first writer open rewrites only the version marker. After that, v0.0.9 and earlier refuse the database with `unsupported_index_storage_version`. - Upgrade readers before the writer. - Back up first if a rollback may be needed. - The local server guide now carries this warning next to the "set `tag`" instructions. - **Asynchronous vector and text index publication** (#1156). - Writes commit the graph change plus a queued index operation, and a background worker publishes it. - `search_consistency: strong | eventual`. Strong is the default and exact. - `index_backpressure` (429, retryable) is returned in three cases: past 1 GB or 250K pending entities per index, when a whole-index strong search is behind more than 800 changes, and when strong text analysis would exceed one publication's budget. - `index_operation_batch_too_large` (400) is returned when one write stages more than 8 MiB for one index. - Text admission limits are now per document. - Builds no longer livelock under concurrent writes. - **Failing entries are held back** (#1165). - An entry whose index update fails deterministically no longer stalls its whole generation. It is counted in `blocked_index_entity_count` and retried about once a minute. - Insert searches never pick the inserting node as its own neighbour. - **Fixes the v0.0.9 `missing simhash` known issue** (#1156 C1, #1165). Re-embeds keep their HNSW reverse-link locators. Indexes damaged by earlier versions still need a drop and recreate (HEL-952). - **Range-driven counts apply every filter** (#1169). For example, 149 → 49. - **CLI 3.4.3:** defaults to v0.0.10. There are no other CLI changes since 3.4.2. - **Not in this release:** the SDK changes in #1156, #1163 and #1164 ship with the next SDK releases. The release notes say to send `search_consistency` over HTTP until then. #1166 is a docs dependency bump. ## Testing - `cargo check --workspace --locked` passes, so the hand-edited `Cargo.lock` is consistent. - `cargo test --locked -p helix-cli` passes: lib 241, `e2e_cli` 13, `runtime_commands` 20, `typescript_runtime`, and the other targets. The Docker-only `e2e_runtime` tests are ignored by design, because they need the unpublished v0.0.10 image. - rustfmt 1.9.0 (the pinned 1.97.1 toolchain) `--edition 2024 --check` is clean on the edited Rust files. - Docs: `check-docs`, `generate-llms --check`, `generate-llms-full --check` and `check-openapi` pass. - Grep: every `helixdb/helixdb:v0.0.x` reference outside release notes is v0.0.10, except the CI-pinned v0.0.5 TypeScript smoke. - Workspace clippy is left to CI. The local nix nightly clippy rejects the repo's clippy configuration before linting, independent of this change. <!-- greptile_comment --> <!-- greptile_summary --> <p><a href="https://app.greptile.com/api/retrigger?id=74957602"><picture><source media="(prefers-color-scheme: dark)" srcset="https://greptile-static-assets.s3.amazonaws.com/badges/RetriggerDark.svg?v=2"><source media="(prefers-color-scheme: light)" srcset="https://greptile-static-assets.s3.amazonaws.com/badges/Retrigger.svg?v=2"><img alt="Retrigger" src="https://greptile-static-assets.s3.amazonaws.com/badges/Retrigger.svg?v=2" align="right"></picture></a><br clear="all"></p> <!-- greptile-risk --> The PR appears safe to merge, though the published-image parity instructions should include an image pull step. <h3>Summary</h3> This release PR updates the CLI default to Docker v0.0.10, bumps the CLI to 3.4.3, and updates image references and release guidance. The published-image parity instructions need a pull step before their new image pin can work without a local copy. <sub>Reviews (1) · Last reviewed commit: ["chore(release): CLI 3.4.3 and Docker ima..."](https://github.com/helixdb/helix-db/commit/281197c0c91e5f72a7ce89fb8a248e04176cdee4)</sub> <!-- /greptile_comment --> ## Also: fix main's workspace tests `counts_over_range_intersections_apply_every_filter` (#1169) overflows the default 2 MiB test-thread stack in Linux debug builds. That fails `Workspace tests`, `DB all targets` and `Workspace and tooling quality` on main, and the Docker release's `quality` job runs the same tests. `92c15755` runs it through the file's existing 16 MiB `run_high_stack_contract`, like the other heavy contracts. It passes locally even with `RUST_MIN_STACK=524288`.
xav-db
added a commit
that referenced
this pull request
Oct 6, 2026
…log (#1175) First PR of the queue stack: the HEL-957 PR (streamed backlog load at open), then the HEL-958 PR (admission memory charge), build on this one. Fixes HEL-961. ## Problem Draining a publication queue cost more than the size of its backlog: - Every attempt regrouped the whole retained queue by entity, and every commit rescanned it to remove the acknowledged operations. With 512-entity batches that is quadratic: going from a 100K to a 250K backlog (2.5×) cost 7.6× the bookkeeping. - A hold (an entity too large for a batch, or one whose planning failed) dropped the queue. The next attempt then re-read and decoded the whole backlog: 136 reads for 64 failing entities at 250K. - A retired generation's discard re-read the whole queue for every acknowledgement chunk: 63 reads at 250K with small acknowledgements. That also kept the generation's charges, which a recreated index shares, held for longer. - A target's schedule was removed only when a read found its queue empty, which never happens once a commit drains it. So `schedules` kept one entry for every target and generation ever published. ## Fix All changes are in memory. **1. The queue is grouped once per read** (`StoredQueue` in `queue/storage.rs`). ```text read order 0:a1 1:b1 2:a2 3:c1 4:b2 chains a: 0 -> 2 b: 1 -> 4 c: 3 order {0: a, 1: b, 3: c} // entities by their oldest outstanding op without(a1) a: 2 order {1: b, 2: a, 3: c} ``` - `rotation(after)` yields entities in the same order as the old regroup, at one map step per entity visited. - `without(acked)` pops chain heads in O(acked · log entities), with no rescan. It asserts that every acknowledgement names an entity's oldest operations in order. - `HeldEntity::repair_width` is O(1): a newer operation is queued exactly when the newest one isn't the `through` that is known not to publish. **2. Holds keep the queue.** A size block or a `FailureKind` isolation commits nothing, so the queue is retained instead of being read again. A retained queue whose entities are all `Waiting` or `Failed` is re-read once the target's latest admission has moved past the one seen before the read. So retries that keep falling due still never pin publication to a stale queue (the 61240f7 fix). **3. A discard continues from the retained queue.** It acknowledges whole-entity prefixes and keeps the rest, so a retired backlog is read and decoded once. Charges are released at each commit. **4. A drained target's schedule is dropped.** Only its latest vector commit is kept, in a ring of the 16 most recently drained targets (`MAX_RETAINED_PUBLICATIONS`). This lets its retained planning session still serve its next write. The other #1156 and #1165 invariants are unchanged: - the indexer is the only writer of index rows, and the ack is in the same transaction; - exact-ID enqueue and ack; - held-entity states, the 60 s retry, and `blocked_index_entity_count`. ## Behaviour/config changes - None visible to users or in config. - Fewer queue reads during holds and discards, and a recreated index gets its shared queue budget back sooner after a drop. ## Measurements (release, c7i.8xlarge, in-process, 2-dim vectors, 512-entity batches; main → branch) **Bookkeeping only** (selecting a batch and removing what was acknowledged): | Backlog | main | branch | |---|---|---| | 10K | 0.011 s | 0.003 s | | 100K | 1.23 s | 0.047 s | | 250K | 9.36 s | 0.157 s | | 1M | — | 0.856 s | **Publish drain:** | Backlog | main | branch | |---|---|---| | 10K | 5.47 s | 5.27 s | | 100K | 60.3 s | 59.3 s | | 250K | 176.2 s (1,418 ops/s) | 154.9 s (1,614 ops/s), −12% | HNSW planning takes most of the time. **64 entities that fail to plan, then a drain to Stalled** (time, reads): | Backlog | main | branch | |---|---|---| | 10K | 7.69 s, 67 reads | 7.19 s, 2 reads | | 100K | 73.0 s, 80 reads | 66.1 s, 2 reads | | 250K | 185.0 s, 136 reads | 161.1 s, 2 reads | At 250K, time spent reading falls from 12.4 s to 0.28 s. **Retired discard with a 64 KiB ack budget** (time, reads): | Backlog | main | branch | |---|---|---| | 10K | 0.31 s, 4 reads | 0.31 s, 2 reads | | 100K | 2.57 s, 26 reads | 2.57 s, 2 reads | | 250K | 12.68 s, 63 reads | 6.52 s, 2 reads | With default acks at 250K, the discard goes from 5 reads to 2 and from 0.96 s to 0.79 s. ## Testing (EC2) - fmt, clippy (workspace, all targets) and `cargo check --locked --workspace --all-features` are clean. - `cargo test -p db --lib`: 2,069 passed. `cargo test -p db --doc`: 40 passed. - `cargo test -p server` passes (71 unit, 2 integration, 5 doc tests). `scripts/validate-cargo-target-references.py` passes. - Release production targets, with `production-coverage,index-lifecycle-testing`: 185 passed. These are queue publication, index-operation queue, text queue, text correctness regressions, internal, contracts, migration, and index lifecycle (run single-threaded, as CI runs it). - Product image: `index_queue_contracts.py --scenario all` passes. Indexes are exact after restart and after SIGKILL, on local disk and on S3. Writes get a 429 at 250,000 members and are accepted 1.4 s after activation. - New `grouping_tests.rs`: a property test (1,024 cases × up to 95 steps). It runs random enqueues, rereads, publishes, every kind of hold, releases, discards and cursor moves, and checks the result against a verbatim copy of the old regrouping. Selections, remaining queues, rotation order and retained bytes must all be identical. There are also should-panic tests for out-of-order and duplicate acks. - New `drain_tests.rs` covers: - retiring mid-drain; - recreating an index whose old backlog is exactly at the limit (429 until the first discard commits, then admitted); - schedule cleanup across 24 targets; - restarting mid-drain with updates and deletes; - concurrent writes during a drain; - holds reusing one read. - The #1165 isolation tests pass unchanged. No stored format changes. **Expected CI follow-up:** the DB coverage fingerprint will change. Take the new count and sha from this PR's ARM64 "DB production thresholds" job. <!-- greptile_comment --> <!-- greptile_summary --> <p><a href="https://app.greptile.com/api/retrigger?id=75528033"><picture><source media="(prefers-color-scheme: dark)" srcset="https://greptile-static-assets.s3.amazonaws.com/badges/RetriggerDark.svg?v=2"><source media="(prefers-color-scheme: light)" srcset="https://greptile-static-assets.s3.amazonaws.com/badges/Retrigger.svg?v=2"><img alt="Retrigger" src="https://greptile-static-assets.s3.amazonaws.com/badges/Retrigger.svg?v=2" align="right"></picture></a><br clear="all"></p> <!-- greptile-risk --> The PR appears safe to merge, with two non-blocking memory and planning-cache efficiency issues worth addressing. <h3>Summary</h3> The PR groups each queue once per read, advances retained queues by acknowledged entity prefixes, reuses queues across holds and retired-generation discards, and bounds schedule history for drained targets. It adds property, drain, and production-support coverage. Two non-blocking resource-efficiency concerns remain: retained allocation size during a long drain and unusable commits entering the drained-target ring. <details><summary>Diagram</summary> ```mermaid %%{init: {'theme': 'neutral'}}%% flowchart LR A[Read durable queue] --> B[Group operations by entity] B --> C[Select ordered prefixes] C --> D{Attempt outcome} D -->|Commit| E[Acknowledge IDs and retain remainder] D -->|Hold or definite conflict| F[Retain unchanged queue] D -->|No selectable work| G[Drop queue and wait] E --> C F --> C ``` </details> <sub>Reviews (1) · Last reviewed commit: ["test(db): cover drains across retirement..."](https://github.com/helixdb/helix-db/commit/4288421a163fb46485908c11b6d29277759bb7a3)</sub> <!-- /greptile_comment -->
xav-db
added a commit
that referenced
this pull request
Oct 6, 2026
Fixes HEL-957.
## Problem
At writer open, `load_backlog` fully decoded every queue of a scope
(every payload, including 1536-dimension vectors) into one `Vec` before
charging any of it to the admission ledger. So open's peak memory and
time grew with the whole durable backlog: about 25 s and 4.5 GB for 2.26
GB of queued work. Each scope also cost about 4 scans (scope seek,
record scan, layout probe, queue scan).
## Fix
```text
before: for scope: decode every queue of scope -> Vec -> charge all
after: one forward scan per keyspace (legacy, then tenants)
row -> QueueDiscovery::push -> completed queue? -> validate owner -> charge -> drop
```
- **Streaming discovery.** `QueueStore::discover` is replaced by
`QueueDiscovery`, which takes rows in key order and returns each queue
as soon as it is complete. In the map layout each row is a whole queue.
In the row layout a generation completes when the next generation's
first row arrives, or at `finish()`. Open now holds one queue at a time
plus the ledger.
- **Framing only.** `OperationFrame::decode_queue` / `decode_row` read
each operation's ID, entity and retained bytes. Payloads are validated
in place (`algebra::validate_body`) and never decoded. The values and
rows they reject are exactly the ones `OperationQueue::decode` /
`QueueRow::decode` reject (duplicate IDs, NaN, UTF-8, trailing bytes,
unresolved values). A proptest checks this, and `OperationQueue::decode`
and `OperationFrame::decode_queue` share one record walk
(`decode_unique`), so their duplicate-ID check cannot drift.
- **Fewer scans.**
- One scan covers the legacy scope.
- One forward-only iterator covers the tenant keyspace, with at most 2
seeks per tenant.
- Index records are read only for scopes that have a queue, and are
cached per scope.
- **Layout check folded in.** `discovery_range` spans both layouts'
adjacent record kinds (a const assert checks this). A queue of the other
layout still fails closed with `Config`, during the same scan. The
separate probes (`require_scope_layout`, `require_layout`,
`discover_scopes`) are now compiled only for tests and the benchmark
feature.
- **Ledger only.** Open fills the admission ledger and nothing else:
publication starts with no retained queue, schedule or held entity, and
each target's first attempt reads and groups its own queue
(`StoredQueue`).
- **Removed:** `QueueStore::discover`, `finish_rows` and
`OperationQueue::from_rows`. `OperationQueue::operations()` is now
compiled only for tests and `index-lifecycle-testing`.
- **Unchanged:**
- Charge order (legacy first, then tenants ascending).
- Owner validation.
- The row-layout sequence resume.
- The worker wake.
- The #1156/#1165 invariants: the indexer is the only writer of index
rows, and held-entity state starts empty either way.
## Behaviour changes
- A corrupt `IndexRecord` in a tenant that has **no** queue no longer
fails writer open. That tenant's first catalog read fails instead
(`Encoding`), and the other tenants keep working. A corrupt record in a
scope that has a queue still fails the open.
- No config or API changes.
## Measurements (EC2 c7i.8xlarge, release)
Fixture: 6 legacy vector queues of 256 MiB, 8 tenant queues of 64 MiB
(dropped indexes), and 2000 tenants with rows but no queue. That is 14
queues, 366,930 ops and 2.26 GB retained, about 6 GB on disk. Each run
starts from a fresh copy of the fixture, 4 runs each, and every run
loaded the same totals.
| Phase | base | this PR |
|---|---|---|
| Backlog load only (no compactor) | 15.1–15.3 s, peak 2.17–2.22 GB |
4.63–4.86 s, peak 1.14 GB (flat) |
| Full writer open (compactor running) | 24.4–24.8 s, peak 4.33–4.65 GB
| 4.83–5.01 s, peak 1.22–1.30 GB |
- Open is about 5× faster.
- The base column was measured on main. `linear-queue-publication-drain`
does not touch the open-time load path (`recovery.rs`, queue discovery),
so it applies unchanged.
- Without the compactor, this PR's peak scales with the largest queue:
about 0.85 GB at 128 MiB queues, 1.14 GB at 256 MiB. The base's peak
scales with the whole backlog.
- Open peaks above the load-only peak come from SlateDB's background
compactor running during open, not from the load.
## Testing (EC2)
- `cargo fmt --all --check`, `cargo clippy --workspace --all-targets --
-D warnings` and `cargo check --locked --workspace --all-features` are
clean.
- `cargo test -p db --lib`: 2,082 passed. `cargo test -p db --doc`: 40
passed.
- Release production targets (queue publication, index-operation queue,
text queue, text correctness regressions, internal, contracts,
migration, index lifecycle), one test thread each: 186 passed.
- `cargo test -p server`: 79 passed.
- `scripts/validate-cargo-target-references.py` passes.
- Product image, `docker-image/test.sh` and `index_queue_contracts.py
--scenario all` pass. After SIGKILL and restart, 10,750 docs are exact
on local disk and on S3. The member limit rejects at 250k and accepts
again 1.4 s after the build activates.
- New `recovery_tests` compare every load against a reference loader
that fully decodes each queue the old way. They cover:
- legacy and tenants 0, 7, 8 and `u128::MAX`, with keys on both sides of
each range;
- the map layout across the compacted run, L0 and the memtable, with
partial and full acks;
- the row layout, with interleaved generations and the sequence resume;
- the other layout failing closed in every scope;
- corrupt payloads, duplicate IDs, foreign keys and invalid tenant
envelopes failing closed;
- misowned tenant queues after other scopes' records are cached (this
test fails if the per-scope cache goes stale);
- an SST read failure injected at every point during open, after which
every reopen reloads the same ledger and writes nothing;
- a full writer restart that reloads the ledger exactly, starts with no
queue reads, retained bytes, schedules or blocked entities, and drains
each reloaded target in exactly 2 reads.
- Production contracts add a tenant scope-walk test and duplicate-ID /
corrupt-payload cases.
- `open_measurement_tests` is an ignored, manual release benchmark that
produced the numbers above.
No stored format changes.
**Expected CI follow-up:** the DB coverage fingerprint will change.
Rows-layout-only lines are uncovered in production targets, and line
numbers shift. Take the new count and sha from this PR's ARM64 "DB
production thresholds" job, after `linear-queue-publication-drain`
merges and this PR is retargeted to main.
xav-db
added a commit
that referenced
this pull request
Oct 6, 2026
…ext analyses (#1176) Fixes HEL-960. ## Problem A strong text search overlays every committed but unpublished document of its partition so its BM25 answer stays exact. On main, and unchanged by the base PR: - it was bounded by one text publication's analysis budget (64 MiB, about 5,500 pending 40-word documents). Any ingest burst past that made strong text searches fail with retryable 429 `pending_text_analysis_bytes` until the index worker caught up (106 and 234 rejections in the 20K and 40K burst runs below); - every search analyzed the whole backlog again from scratch and copied it into an in-memory Tantivy index, so cost and memory grew with the backlog and with each concurrent search (about 200 ms and 1 GB RSS at 5K pending with 1 searcher, 1.7 GB with 8). ## Fix ```text strong text search pending set (per request, from #1174) for each pending doc in the partition cached analysis (by queued op id)? reuse it, charged as if fresh else wait for the analysis turn -> analyze on the blocking pool -> cache charge > strong_text_search_max_analysis_bytes -> 429 pending_text_analysis_bytes score pending docs in place (Tantivy Bm25Weight, best k) -> merge with published results ``` **1. Strong text searches have their own bound.** - `IndexOperationQueueTuning::strong_text_search_max_analysis_bytes` (default 512 MiB, about 44K pending 40-word documents), next to the base's strong vector bound. Past it, the search fails with the same retryable 429 `index_backpressure`, resource `pending_text_analysis_bytes`. - Applies to read and write requests. A write's own documents are charged first; alone past the bound they fail the write with `index_operation_batch_too_large`, as before. - Eventual searches are unchanged: they overlay at most one publication budget (64 MiB) and serve the rest as last published. **2. Analyses of pending text are reused (in memory only).** - `PendingTextAnalyses`, per database, keyed by (queue target, partition, queued operation ID). Selection reads the base's per-request decoded queue, where each committed entity now carries its queued operation ID. - It is exact without invalidation: a queued payload never changes and its random 121-bit ID is never reused, so publication, supersession, a rebuild or a drop can only make an entry unreachable. A reused analysis is still compared with the queued text and fails closed on a mismatch. - A reused analysis is charged exactly as a fresh one, so a 429's `requested` value is the same cold or warm. - Capped at the bound; the least recently used partitions give up only the bytes a replacement needs. - Released when the publisher drains or discards a generation, and when a strong search finds its queue empty (the only path on reader handles). - A write transaction's own text is never cached. **3. Pending documents are scored in place.** - No clone and no in-memory index. Scoring uses Tantivy's `Bm25Weight` with the same corpus statistics and fieldnorm quantization, ranked by score then entity ID. - Scoring runs on the blocking pool through the base's `run_blocking`, in steps of 4,096 documents with a request check between steps, keeping only the best k per step, so memory grows with k, not with the number of hits. - A proptest against a real Tantivy index requires bit-exact results for queries of up to 2 terms. For 3 or more terms results must match within rounding, because Tantivy's own union sums in an order that depends on index layout. **4. Only one strong search at a time analyzes uncached text.** - A process-wide first-come-first-served turn. A search that waited reads the cache again and reuses what the previous holder analyzed. - Cold analysis runs through `run_blocking` and checks the request before each fresh document. The turn is an owned guard moved into that blocking work and released only when it stops, so an abandoned request cannot let a second cold analysis overlap. - A pass that stops early (deadline, cancellation, reader retirement) still caches what it analyzed plus the cached entries after it, so the next search resumes instead of starting over. - Fully cached strong searches and eventual searches never wait. - However many searches run at once, analyses of pending text stay at about two bounds: one cached and one being analyzed. **5. Observability.** `IndexOperationQueueStats::strong_text_search_rejections` counts refusals. A warning naming the setting is logged at most once a minute. #1156/#1165 invariants are unchanged: the indexer is still the only index writer, acks stay in the same transaction, strong search stays exact, and held-back entities still count toward the bound. ## Behaviour/config changes - New server env var `HELIX_STRONG_TEXT_SEARCH_MAX_ANALYSIS_BYTES`: a positive byte count, default 512 MiB. It feeds the same tuning builder as the base's vector variable. An invalid value fails startup with `IndexQueueBytes` naming the variable, before any cache directory is created. Benchmark builds keep it. Embedded users set it with `IndexOperationQueueTuning::with_strong_text_search_max_analysis_bytes`. - Strong text 429s now start at 512 MiB of pending analysis per partition instead of 64 MiB. The resource name is unchanged. - While text is pending, the database keeps up to one bound of cached analyses, plus about one more while a cold search analyzes. - New stats field `strong_text_search_rejections`. On reader handles it is the only nonzero field. - Docs updated: search-consistency (one table for both strong bounds), troubleshooting, error-handling, prefiltering, Cloud limits, local-server, docker-image README, OpenAPI 429 description, llms-full.txt. ## Measurements EC2 c7i.8xlarge. Ingest burst, product benchmark image: 100 docs/s plus one burst, 2 strong searchers, 90 s. Latency counts successful searches only, so main's numbers leave out the searches it refused, which were the ones behind the largest backlogs. Measured before the restack onto the base PR. | Burst | Image | Strong-search 429s | p50 / p99 (ms) | Peak container memory | |---|---|---|---|---| | 20K | main | 106 of 1,705 | 50.3 / 129.4 | 1,716 MiB | | 20K | branch | **0** of 1,580 | 58.5 / 139.5 | 1,271 MiB | | 40K | main | 234 of 1,598 | 58.7 / 166.1 | 2,002 MiB | | 40K | branch | **0** of 1,335 | 77.3 / 223.0 | 1,287 MiB | The 40K burst peaked at about 279 MiB of pending analysis. In-process probe, release build: 20K published docs, then N pending with publication paused. Every document and query contains a common term, so no pending document is filtered out, which is the worst case. Values are strong search p50 / p99 (ms) and process RSS. `*` marks values measured before the restack and not re-run. | Pending | Searchers | main | branch p50 / p99 | branch RSS | |---|---|---|---|---| | 0 | 1 | 13.2 / 18.5 | 12.8 / 18.6 * | | | 5K (59 MiB) | 1 | 204 / 217, 1,062 MiB | 27 / 39 | 322 MiB * | | 24K (282 MiB) | 1 | 429 | 113 / 439 * | 689 MiB * | | 40K (470 MiB) | 1 | 429 | 208 / 299 | 813 MiB * | | 5K | 8 | 261 / 293, 1,668 MiB | 46 / 63 | 376 MiB * | | 40K | 8 | 429 | 248 / 401 | 1,529–1,680 MiB peak | | 40K | 32 | 429 | 904 / 1,152 * | 1,860 MiB peak * | - Versus the pre-restack build on the same host in the same session: single-searcher latency is the same; at 8 searchers p50 is up to about 8% higher and p99 the same or lower. No search was rejected. - RSS at 40K with 8 searchers is higher than before the restack (1,118–1,192 MiB): glibc keeps per-thread arenas for the extra blocking-pool threads. With `MALLOC_ARENA_MAX=2` both builds sit at about 790 MiB. The multi-searcher RSS values marked `*` are likely higher for the same reason. - Eventual search at 5K pending with 1 searcher: 184 → 27 ms p50 *, because it also scores in place now. - A warm strong search still costs O(backlog) CPU: one statistics-marker read per pending entity, plus scoring. The statistics pass and the scope/term filter stay on the async workers, one cheap pass over the backlog each. At 32 concurrent searches and 40K pending this saturates the 32 cores. ## Testing (EC2) - `cargo fmt --all --check`, `cargo clippy --workspace --all-targets -- -D warnings` and `cargo check --locked --workspace --all-features` are clean. `validate-cargo-target-references.py` passes. - `cargo test -p db --lib`: 2,097 passed (26 more than the base). `cargo test -p db --doc`: 40 passed. - Release production targets (`production-coverage,index-lifecycle-testing`, `--test-threads=1` as CI runs them): 187 passed: contracts 58, internal contracts 41, migration contracts 27, text correctness regressions 23, index-operation queue 15, text queue contracts 12, index lifecycle contracts 8, and queue publication 3. - `cargo test -p server` passes (72 unit tests plus the integration and doc targets), as does `--features async-index-benchmark` (76 unit tests). - Product image: `docker-image/test.sh` passes, and `index_queue_contracts.py --scenario all` passes. Correctness is exact after restart and SIGKILL on disk and on S3 (10,750 docs each). In the member-limit scenario, the batch past 250,000 members gets a 429 and the retried batch is accepted 2.5 s after activation. - Docs: `check-docs`, `generate-llms --check`, `generate-llms-full --check` and `check-openapi` pass. - New coverage: - Cache: replacement keeps exactly the latest selection; exact-budget and one-byte-over boundaries; partial LRU shedding across partitions; forget per target; a search holding an old snapshot still reads it. - Selection: identical charges cold and warm; refusals report the same `requested` value cold and warm, at the text charge and at a token; refused searches still cache the prefix they analyzed; a write's own text is never cached; a cached text mismatch fails closed; committed and local entities keep their order and source. - Turn: a cold search waits, then reuses the holder's analyses. Cached searches, eventual searches and searches of another partition do not wait. A search over a write's own text does. A holder whose request is dropped, retired or expires mid-analysis keeps the turn until its blocking work stops; it caches its prefix plus the cached tail, and the next search finishes without analyzing any document twice. - Scoring: proptest against a real Tantivy index covering fieldnorm quantization, ties, restricted scopes, empty documents and extra corpus statistics. Stepped scoring keeps the k best across steps and checks the request at each step. - Overlay: - exact past the publication budget, across a restart and after publication; - the exact bound passes and one byte over fails, cold and warm; - a write batch sees its own staged text despite a warm cache; - updates, deletes and re-inserts; - concurrent strong searches racing writes and automatic publication; - the rejection counter. - Release: the publisher keeps the cache while work remains and drops it once drained. Dropping a text index releases its analyses. A strong search over an empty queue releases them on writer and reader handles; with that release removed, both tests fail. - Production contract (`production_text_queue_contracts`): - the default bound stays exact past the publication budget; - a lowered bound gives a retryable 429 in read requests, through `HelixQueryService` (as Backpressure), and in write requests, which roll back; - eventual results are unchanged, and publication clears the 429; - a write's own text past the bound gets `IndexOperationBatchTooLarge`. - Server config: one test covers each strong bound variable and both together. Values 1, 2^31 and u64::MAX are accepted. 0, -1, empty, `512MiB` and u64::MAX+1 are rejected before any cache directory is created. Benchmark builds keep both bounds. No stored format changes. **Expected CI follow-up:** the DB coverage fingerprint will change. Take the new count and sha from this PR's ARM64 "DB production thresholds" job.
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #1156; the base branch is
async-index-queue, and it should be retargeted tomainonce #1156 merges. Fixes HEL-956.Problem
In #1156, if planning one entity's queued index operation failed the same way every time, its whole (scope, index, generation) stopped publishing for good. The error was retried forever, the queue filled, writes to that index got 429, and
blocked_index_entity_countstayed at 0.A real trigger exists: data written by released versions (HEL-952) can contain a self-link, and the vector insert searches never excluded the inserting node, so planning failed every time with
ContainsOwner.Fix
1. A node is never its own neighbour.
2. Only the failing entity is held back.
Behaviour changes
blocked_index_entity_count(in stats and health) now also counts entities held back because their planning failed, not only entities too large to fit a publication. The docs and the proto/OpenAPI descriptions say so.Testing (EC2)
fmt and clippy (workspace) are clean.
cargo test -p db --lib: 2,058 passed.Release production targets (queue publication, internal contracts, text queue, index-operation queue, text correctness regressions): 92 passed.
cargo test -p serverpasses.The release queue soak passes with 5 seeds.
The new isolation tests cover:
They also cover repair writes and the 60 s retry.
The text correctness regressions now expect a damaged root to hold back only its own entity, with the search still failing closed, instead of retrying forever.
No stored format changes. The only effect on stored data is that a damaged row loses its self-link when a mutation rewrites it.
Expected CI follow-up: the DB coverage fingerprint will change. Take the new count and sha from this PR's ARM64 "DB production thresholds" job.
The PR should wait for the production coverage baseline update so its required CI gate can pass.
Findings
Summary
The PR makes queued index publication hold back an entity whose planning fails deterministically while continuing to publish other entities. It also prevents HNSW inserts from choosing their own node as a neighbor and tolerates stored self-links during mutation.
Diagram
%%{init: {'theme': 'neutral'}}%% flowchart TD A[Select queued entities] --> B[Plan publication] B -->|Success| C[Commit effects and acknowledge] B -->|Transient failure| D[Retry batch with backoff] B -->|Deterministic vector failure| E[Hold identified entity] B -->|Deterministic text failure| F[Narrow epoch] F -->|One entity fails| E E --> G[Publish other entities] E --> H[Retry on newer write or timer]Reviews (1) · Last reviewed commit: "docs: count failed-planning holds in blo..."