Skip to content

Find out how to improve Opensearch reliability #5648

Description

@neilmb

Feature/what we're after

We rely critically on our Opensearch cluster for the new Catalog's search capability. We have experienced instability with Opensearch going "red" entirely and stopping working. When this happens, we have to involve Cloud.gov support and wait until they can figure out what's wrong or the Opensearch system heals itself.

We need to get better reliability from Opensearch than we have now. We should research different ways to do that. Here are some things to look into:

  1. Are there things we can do to keep the Opensearch cluster "healthy"? Are there maintenance things that heavy Opensearch users do that we should add?
  2. Does our heavy daily sync job exacerbate this reliability issue? If so, should we consider raising the priority of Implement an incremental harvest-time process for opensearch sync #5607 so we stop doing that sync job?
  3. Could we have a "hot-standby" Opensearch cluster that we keep running and synced up alongside of the one we are using so that if it goes "red" then we can easily bind to the fully synced cluster.
  4. Could we go one step further and use two Opensearch clusters in a blue-green fashion so that we sync against the "inactive" cluster and then bind/undbind to switch to the newly synced cluster? This would give us more "atomic" synchronization if we need to recreate the whole index.

For this ticket, we want a recommendation on what to do next. We might have to try some of these things to make the recommendation, but the point of this ticket is to figure things out.

Anticipated/hypothesized benefits

  • [benefit A]
  • [benefit B]

Measurements/metrics

  • [description of a thing that can be monitored or tested to measure and validate when the outcome has been achieved]

References/background

  • [notes about desirable attributes or context for the epic]
  • [links to previous research, possible routes to take, things that multiple stories will likely need to refer to]

Activity

  1. moved this to 🌈 Catalog UI Project in data.gov team boardon Jan 21, 2026
  2. added theissue type on Jan 21, 2026
  3. FuhuXia commented on Jan 22, 2026

    @FuhuXia
    Member

    OpenSearch went red on 2025-01-21.

    • Incident thread on slack
    • cg support response to the ticket with troubleshooting notes and suggestions. (email from Mark Boyd on 2025-01-21)
    • we bumped the plan from es-medium to es-large.
  4. moved this from 🌈 Catalog UI Project to 📟 Sprint Backlog [7] in data.gov team boardon Jan 27, 2026
  5. moved this from 📟 Sprint Backlog [7] to 🏗 In Progress [8] in data.gov team boardon Jan 28, 2026
  6. self-assigned this
    on Jan 28, 2026
  7. rshewitt commented on Feb 2, 2026

    @rshewitt
    Contributor

    here's the latest reasons for why opensearch sync fails:

    • psycopg.errors.SerializationFailure) terminating connection due to conflict with recovery. this is caused by the read-only replica cancelling a query because the data is out of sync with the read-write (specifically rows have been deleted). this would be avoided in #5607
    • opensearchpy.exceptions.TransportError: TransportError(500, 'search_phase_execution_exception', '[238e0529432848c2a6dfea88e7c5832b][x.x.x.x:9300][indices:data/read/search[phase/query/scroll]] disconnected'). this typically means the manager node can't communicate with the data node so the search fails to execute. apparently this can be due to network issues, node overload (high CPU/JVM), misconfigured certificates, or nodes leaving the cluster.
      • here's a generic one related to this opensearchpy.exceptions.TransportError: TransportError(503, 'search_phase_execution_exception')
    • opensearchpy.exceptions.TransportError: TransportError(429, 'circuit_breaking_exception', '[parent] Data too large. the request has exceeded the configured memory limit. we gather opensearch ids using helpers.scan which is the recommended method for retrieving large volumes of documents. it uses the scroll API for efficient batching.
    • opensearchpy.exceptions.ConnectionTimeout: ConnectionTimeout caused by - ReadTimeout(HTTPSConnectionPool(host='vpc-cg-broker-prd-8qhmo39iv-qg6zpvom2uuyjqxfgckdn6tpum.us-gov-west-1.es.amazonaws.com', port=443): Read timed out. (read timeout=60)). apparently this can be caused by slow queries, long garbage collection pauses in OpenSearch nodes, or network device idle timeouts.

    Max JVM memory pressure for staging and prod in the past 2 weeks

    Image
  8. rshewitt commented on Feb 4, 2026

    @rshewitt
    Contributor

    I think if we want to do this right we should keep sync as an isolated process instead of integrating it into harvesting. i figure we'll always have it isolated anyways for retries similar to db-solr-sync? here's some drawbacks to integrating it into harvesting

    • if harvesting fails then search sync fails
    • if search is slow then harvesting is slow (i suppose we've never cared about the speed of things though?)
    • they can't be scaled independently
    • harvesting occurs whenever there's work to be done which means JVM pressure can occur whenever work is being done instead of, let's say, before business hours (if we're sticking to a write-only approach then maybe this isn't a big deal)

    looking into the matter it seems like a common approach to synchronizing 2 backends like this (without implementing change-data-capture into opensearch which i don't think we can do with cloud.gov postgres) is to use a monotonic version counter or timestamp (this can be problematic because of the potential of duplicates which doesn't affect the version approach). the idea is to keep a latest_version int (global incrementing sequence counting changes over time) or latest_sync datetime as the watermark for when sync occurred and only updating it on successful bulk sync (if 1 record out of the chunk fails then we don't update the watermark and stop syncing). we can store this in its own table as 1 row. here's how sync would work...

    • find all datasets with a last_harvested_date or version (this would be a new field if we chose this approach) greater than our watermark. these records would be sorted based on the datetime or version int
    • bulk sync all of these records. after each successful chunk update the watermark. stop syncing on the first failing chunk. this ensures our measurement of records needing to be synced is accurate. i suppose we could update the watermark per record to avoid unnecessary work.
    • for deletes, we can avoid needing to do a full scan by either adding a "is_deleted" field to the dataset table and soft delete it by labelling it or we can keep a table of deleted ids and update opensearch based on those values.

    sync would be write-only to opensearch. this should improve the reliability of opensearch because sync would no longer read anything from it.

    another option is setting a boolean synced_with_os that we flick off on any change and back on when synced

    • i suppose there's a possibility of data loss if a record is updated in between sync and synced_with_os=true (race condition). i can't imagine that happening often.
    • this option may produce less duplication of indexing because each record in the chunk would be updated to synced_with_os=true. if the last record in the chunk, for example, fails to sync then only that one would still be labelled as synced_with_os=false. retry would only pick that one up. in contrast to the monotonic option which wouldn't update the watermark despite processing all but one record which means the entire chunk needs to be processed again on retry.
  9. moved this from 🏗 In Progress [8] to 👀 Needs Review [2] in data.gov team boardon Feb 4, 2026
  10. rshewitt commented on Mar 2, 2026

    @rshewitt
    Contributor

    despite recent opensearch failures 5745 OS has been relatively stable after upgrading the service plan. if OS continues to act up (referencing 5745) then i'll re-open this ticket.

  11. moved this from 👀 Needs Review to ✔ Done in data.gov team boardon Mar 2, 2026
  12. moved this from ✔ Done to 🗄 Closed in data.gov team boardon Mar 11, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions