Skip to content

[BUG] Index pruning silently drops an index it cannot reach, hiding missing shards #5808

Description

@ahkcs

Describe the bug

Request-level time-bounds index pruning (#5766) narrows a wildcard index expression to the indices that can match the query's time range, before the search is issued. It decides what can match by probing each index with _field_caps plus an index_filter and keeping whatever the probe returns.

An index whose primary shard is unassigned — a node is down, or the shard has not been allocated — cannot be probed, and _field_caps reports no failure for it. It is simply absent from the response, which is indistinguishable from an index the probe proved holds nothing in range. Pruning therefore drops it from the expression.

That loses the index's documents with nothing left anywhere to report it. The search that runs afterwards covers every shard it was given, so:

This is strictly worse than allow_partial_search_results=true. With partial results allowed the evidence survives in _shards and a client can surface it. Here the evidence is destroyed before the search exists: pruning does not return a partial answer, it silently makes the question smaller so the answer comes back complete.

To Reproduce

Two-node cluster, a wildcard pattern of 4 shards holding 990 documents, with one index (osdp-broken, 240 docs) pinned to node-2. Send the query with the request-level bounds OpenSearch Dashboards sends, then stop node-2:

curl -XPOST localhost:9200/_plugins/_ppl -H 'Content-Type: application/json' -d '{
  "query": "source=osdp-* | stats count() as n",
  "time_field": "@timestamp",
  "start_time": "2026-09-23 21:18:39",
  "end_time":   "2026-09-24 21:18:40"
}'
result
PPL, node-2 up 990
PPL, node-2 stopped 750, HTTP 200, no warning, no error
PPL, node-2 stopped, search.default_allow_partial_results=false 750, still no error
equivalent DSL search over the same pattern, same setting HTTP 503, Search rejected due to missing shards [[osdp-broken][0]]
PPL, node-2 stopped, without the request bounds (no pruning) 750 + PARTIAL_RESULT_SHARD_FAILURE warning

The last two rows isolate it: the data loss is detectable and reportable until pruning removes the index.

The probe's blindness is directly observable — osdp-broken is missing from indices and there is no failure entry to notice:

curl -XPOST 'localhost:9200/osdp-*/_field_caps?fields=@timestamp' -H 'Content-Type: application/json' -d '{
  "index_filter": {"range": {"@timestamp": {"gte": "...", "lte": "..."}}}
}'
indices:        ['osdp-main', 'osdp-textconflict']
failed_indices: None
failures:       null

Root cause

IndexPruner.prune (opensearch module) treats "absent from the probe's candidate list" as "proven to hold no data in range":

  1. IndexExpression.probeMatching issues FieldCapabilitiesRequest.indexFilter(filter) and reads response.getIndices().
  2. _field_caps with an index_filter runs a can-match per index. An index with no available shard copy is omitted from indices, and — verified on 3.9 — populates neither failed_indices nor failures, so the caller cannot tell it apart from an index that matched nothing.
  3. prune then returns the candidate list whenever candidates.length < resolved.getIndices().size(), so the unreachable index is dropped from the expression the search reads.

Nothing later in the pipeline can recover the fact, because the narrowed search is genuinely complete over the shards it was given.

Expected behavior

Pruning is an optimization, so it must never drop an index it could not prove empty. An index that could not be probed should stay in the expression, so the search includes it and OpenSearch reports it as the missing shard it is — which then feeds the existing partial-result warning, or a rejection under allow_partial_search_results=false.

Proposed fix

Before narrowing the expression, read the routing table and keep any resolved index with a missing primary:

  • Read it from the local cluster state (ClusterStateRequest.clear().routingTable(true).local(true)), so there is no round trip to the cluster manager.
  • Run the check only when pruning would actually narrow the expression, leaving the healthy path unchanged.
  • Skip indices with no routing table (closed indices are not searched either way).
  • When nothing matched the filter, pruning already declines — keep that, so a partial answer does not turn into an all-shards-failed error by narrowing to only unavailable indices.

Being conservative here costs at most a few extra indices in the search while a node is down, and only in patterns that contain one.

Follow-up worth splitting out

With this fixed and search.default_allow_partial_results=false, the query correctly fails, but as HTTP 500 java.sql.SQLException: exception while executing query: Failed to fetch data from the index: the background task failed or interrupted rather than core's HTTP 503 Search rejected due to missing shards [...]. Wrong status class and an unhelpful message; the shard-failure cause should be mapped through rather than wrapped by the background scanner.

Plugins

SQL/PPL (Calcite path). Affects any query carrying request-level time bounds, i.e. every PPL query issued from Dashboards Discover/Explore with a time picker.

Host/Environment

OpenSearch 3.9.0-SNAPSHOT, plugins.query.pruning.enabled at its default (true).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    PPLPiped processing languagebugSomething isn't working

    Type

    No type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions