Repository navigation
Find out how to improve Opensearch reliability #5648
Description
Activity
OpenSearch went red on 2025-01-21.
- Incident thread on slack
- cg support response to the ticket with troubleshooting notes and suggestions. (email from Mark Boyd on 2025-01-21)
- we bumped the plan from es-medium to es-large.
- moved this from 🌈 Catalog UI Project to 📟 Sprint Backlog [7] in data.gov team board
on Jan 27, 2026 - moved this from 📟 Sprint Backlog [7] to 🏗 In Progress [8] in data.gov team board
on Jan 28, 2026 here's the latest reasons for why opensearch sync fails:
psycopg.errors.SerializationFailure) terminating connection due to conflict with recovery. this is caused by the read-only replica cancelling a query because the data is out of sync with the read-write (specifically rows have been deleted). this would be avoided in #5607opensearchpy.exceptions.TransportError: TransportError(500, 'search_phase_execution_exception', '[238e0529432848c2a6dfea88e7c5832b][x.x.x.x:9300][indices:data/read/search[phase/query/scroll]] disconnected'). this typically means the manager node can't communicate with the data node so the search fails to execute. apparently this can be due to network issues, node overload (high CPU/JVM), misconfigured certificates, or nodes leaving the cluster.- here's a generic one related to this
opensearchpy.exceptions.TransportError: TransportError(503, 'search_phase_execution_exception')
- here's a generic one related to this
opensearchpy.exceptions.TransportError: TransportError(429, 'circuit_breaking_exception', '[parent] Data too large. the request has exceeded the configured memory limit. we gather opensearch ids usinghelpers.scanwhich is the recommended method for retrieving large volumes of documents. it uses thescrollAPI for efficient batching.opensearchpy.exceptions.ConnectionTimeout: ConnectionTimeout caused by - ReadTimeout(HTTPSConnectionPool(host='vpc-cg-broker-prd-8qhmo39iv-qg6zpvom2uuyjqxfgckdn6tpum.us-gov-west-1.es.amazonaws.com', port=443): Read timed out. (read timeout=60)). apparently this can be caused by slow queries, long garbage collection pauses in OpenSearch nodes, or network device idle timeouts.
Max JVM memory pressure for staging and prod in the past 2 weeks

I think if we want to do this right we should keep sync as an isolated process instead of integrating it into harvesting. i figure we'll always have it isolated anyways for retries similar to db-solr-sync? here's some drawbacks to integrating it into harvesting
- if harvesting fails then search sync fails
- if search is slow then harvesting is slow (i suppose we've never cared about the speed of things though?)
- they can't be scaled independently
- harvesting occurs whenever there's work to be done which means JVM pressure can occur whenever work is being done instead of, let's say, before business hours (if we're sticking to a write-only approach then maybe this isn't a big deal)
looking into the matter it seems like a common approach to synchronizing 2 backends like this (without implementing change-data-capture into opensearch which i don't think we can do with cloud.gov postgres) is to use a monotonic version counter or timestamp (this can be problematic because of the potential of duplicates which doesn't affect the version approach). the idea is to keep a
latest_versionint (global incrementing sequence counting changes over time) orlatest_syncdatetime as the watermark for when sync occurred and only updating it on successful bulk sync (if 1 record out of the chunk fails then we don't update the watermark and stop syncing). we can store this in its own table as 1 row. here's how sync would work...- find all datasets with a
last_harvested_dateorversion(this would be a new field if we chose this approach) greater than our watermark. these records would be sorted based on the datetime or version int - bulk sync all of these records. after each successful chunk update the watermark. stop syncing on the first failing chunk. this ensures our measurement of records needing to be synced is accurate. i suppose we could update the watermark per record to avoid unnecessary work.
- for deletes, we can avoid needing to do a full scan by either adding a "is_deleted" field to the
datasettable and soft delete it by labelling it or we can keep a table of deleted ids and update opensearch based on those values.
sync would be write-only to opensearch. this should improve the reliability of opensearch because sync would no longer read anything from it.
another option is setting a boolean
synced_with_osthat we flick off on any change and back on when synced- i suppose there's a possibility of data loss if a record is updated in between sync and
synced_with_os=true(race condition). i can't imagine that happening often. - this option may produce less duplication of indexing because each record in the chunk would be updated to
synced_with_os=true. if the last record in the chunk, for example, fails to sync then only that one would still be labelled assynced_with_os=false. retry would only pick that one up. in contrast to the monotonic option which wouldn't update the watermark despite processing all but one record which means the entire chunk needs to be processed again on retry.
- moved this from 🏗 In Progress [8] to 👀 Needs Review [2] in data.gov team board
on Feb 4, 2026 despite recent opensearch failures 5745 OS has been relatively stable after upgrading the service plan. if OS continues to act up (referencing 5745) then i'll re-open this ticket.
- moved this from 👀 Needs Review to ✔ Done in data.gov team board
on Mar 2, 2026
Metadata
Metadata
Assignees
Labels
Type
Projects
- StatusShow more project fields🗄 Closed
Feature/what we're after
We rely critically on our Opensearch cluster for the new Catalog's search capability. We have experienced instability with Opensearch going "red" entirely and stopping working. When this happens, we have to involve Cloud.gov support and wait until they can figure out what's wrong or the Opensearch system heals itself.
We need to get better reliability from Opensearch than we have now. We should research different ways to do that. Here are some things to look into:
For this ticket, we want a recommendation on what to do next. We might have to try some of these things to make the recommendation, but the point of this ticket is to figure things out.
Anticipated/hypothesized benefits
Measurements/metrics
References/background