shuffle: use blocking reads for gapped-producer replays - #3576
Conversation
A replay reads [offset, end_offset), where end_offset is the begin of a trigger document the main read already observed, so the whole range is known to exist. But a lagging broker may not yet index it: a newly-assigned replica serves reads before it has listed a fragment still being persisted by its prior topology, and reports a write head at the end of the last fragment it knows. A non-blocking replay read routed to such a broker either fails with OFFSET_NOT_YET_AVAILABLE (failing the task), or reads through to the stale write head and ends cleanly short of end_offset, silently skipping the remainder of the replay. A blocking read instead waits for the broker to index the fragment, and still terminates at end_offset.
Strix Security ReviewNo security issues found. Review summaryReviewed the single-file diff in Updated for Reviewed by Strix |
dgreer-dev
left a comment
There was a problem hiding this comment.
LGTM.
Do you think there's any value in a regression test where the replay hits a stale write head, waits for a missing fragment to become available, and then verifies that all replayed documents are delivered exactly once? Should be red with the old non-blocking read and green with this change.
Not especially. I think it would be a lot of test SLOC, and would predominantly be testing Gazette and client behavior -- behavior that's well covered under main-path reads. This change is fundamentally just making replay reads less special and like main-path blocking reads, but with an end offset, and we have existing coverage that the end-offset is respected. |
Summary
Gapped-producer replays read
[offset, end_offset), whereend_offsetis the begin of a trigger document the main read already saw, so the whole range is known to exist. A broker that lags, though, may not have indexed it yet. A newly-assigned replica serves reads before it has listed a fragment still being persisted by its prior topology.Replays used non-blocking reads, which fail against such a broker in one of two ways:
classify_read_failuretreats this as terminal, so the task fails.end_offset: when the broker's write head is exactly where the replay has read to. The gazette client ends a non-blocking read at the write head, so the replay reports "complete" and silently skips[write_head, end_offset).This switches replays to
block: true. The broker waits until it has indexed the fragment, and the read still stops atend_offset.