Repository navigation
runtime-next: hot-path optimizations - #3575
Conversation
…de is set OverrideFilter reported `Interest::sometimes` and a TRACE level hint for every callsite, so the process's max level was TRACE and every disabled `trace!`, `debug!`, and span went through a dynamic `enabled()` check: EnvFilter, then a current-span lookup (sharded_slab CAS) and scope walk. h2 opens spans per frame, and the shuffle Slice and Log actors trace per document, so this ran on the hottest paths of a sidecar. Now a process-wide count tracks live handlers with an override set. While it's zero, OverrideFilter reports `Interest::never` and an INFO hint (which keeps handler `info_span!`s alive), so disabled callsites are skipped statically. Setting the first override, or clearing the last (including by its handler finishing), rebuilds the callsite interest cache. Performance (runtime-lab, mean ± sd of n=3 interleaved runs, vs. parent): materialize-sink, 3 shards: 196.6k ± 0.7k docs/s (+8.4%), 25.0 CPU-s/M docs (-9.1%) capture-fake-postgres, 1 shard: backfill 145.2k ± 1.1k docs/s (+23%), replication 140.3k ± 0.8k docs/s (+76%)
…ocking pool `ChildStdio` was `tokio::fs::File`, which runs every read and write as a blocking-pool task: a thread handoff, a buffer copy, and futex wakeups per operation. It carries the connector protocol, so a materialization shard paid that per ~32KB of C:Store requests and per read of the connector's responses. `ChildStdio` is replaced by `ChildStdin` and `ChildOutput` (stdout and stderr), aliases of tokio's non-blocking pipe Sender and Receiver. They register with the IO driver of the runtime current at `Child::from`, and a direction mismatch is now a type error rather than a runtime EBADF. The crate is now unix-only: the non-unix path never built (it imported `std::os::fd::OwnedHandle`), and Windows is not a target. A pipe Sender's `flush` and `shutdown` are no-ops, so materialize-consistency drops its `shutdown` calls; as before, dropping the pipe is what sends EOF. Performance (runtime-lab, mean ± sd of n=3 interleaved runs, vs. parent): materialize-sink, 3 shards: 208.3k ± 0.9k docs/s (+5.9%), 23.7 CPU-s/M docs (-5.4%), context switches -46% capture-fake-postgres, 1 shard: backfill 142.4k ± 0.9k docs/s (-1.9%), replication 150.3k ± 0.9k docs/s (+7.2%)
Validation visits each document node with every active Frame, and each
pass (items, properties, strings, numbers, containers, nodes, unwind)
scanned all of a Frame's keywords looking for ones it handles. Most
Frames, especially in-place applications like $ref and allOf, have none.
wind_frame now folds a Frame's keywords into KIND_* classes. Frame keeps
the eight classes the per-node passes check in a `u8`, so Frame stays one
cache line, and each pass skips Frames lacking its class. Two further
classes (in-place applications, and unevaluatedItems/Properties) are only
needed while winding. Retained classes are `u8` and wind-only classes `u16`,
so testing a wind-only class against Frame::kinds is a type error.
unwind_frame now unwinds the top Frame in place and then truncates it,
rather than popping it by value: the Frame was typically just written by
wind_frame, and moving it whole stalled on those pending stores.
Criterion medians, before -> after:
citi_rides rides1x 3.36 ms -> 2.81 ms (-16%)
rides4x 14.38 ms -> 11.99 ms (-17%)
github scrape0 202 µs -> 182 µs (-10%)
scrape1 230 µs -> 208 µs (-10%)
scrape2 231 µs -> 209 µs (-9%)
scrape3 220 µs -> 198 µs (-10%)
Performance (runtime-lab, mean ± sd of n=3 interleaved runs, vs. parent):
materialize-sink, 3 shards: 211.3k ± 0.6k docs/s (+1.5%), 23.3 CPU-s/M docs (-1.5%)
capture-fake-postgres, 1 shard: backfill 149.9k ± 1.8k docs/s (+5.3%), replication 158.4k ± 0.4k docs/s (+5.3%)
…ner disk A capture closed its transaction once the combiner had spilled 64MB to disk, so a large backfill transaction closed only after it had already paid to spill: serialization, compression, and IO of documents which are then read back to drain. Backfill transactions ran up to ~110MB of captured JSON. Instead bound transactions at 64MB of captured document bytes, which keeps the combiner within memory and favors small transactions, and leave combiner disk usage at the close policy's default. The bound is evaluated at connector checkpoints, so a transaction may exceed it by a checkpoint's worth of documents. Performance (runtime-lab, mean ± sd of n=3 interleaved runs, vs. parent): capture-fake-postgres, 1 shard: backfill 160.4k ± 1.1k docs/s (+7.1%), replication 157.9k ± 1.6k docs/s (no change) materialize-sink: not applicable (capture-only change)
Strix Security ReviewNo security issues found. Review summaryReviewed all 11 changed files in this performance-optimization PR. The changes are behavior-preserving hot-path optimizations across four crates: static callsite-interest skipping in service-kit tracing, IO-driver-based child stdio in async-process, keyword-class-based Frame skipping in the JSON validator, and a capture transaction bound switch in runtime-next. The security-sensitive area was the JSON validator's new Updated for Reviewed by Strix |
|
There are a bunch of other performance levers to pull, after further investigation using Adding worker threads is not an obvious improvement (for captures, in particular). It requires a deeper dive into how tasks yield, when, and how tasks bounce between tokio worker threads, and also splitting up the monolith capture shard actor so that draining can be parallel to loading a next combiner. |
Summary
Four independent hot-path optimizations. Together they speed up a 3-shard materialization by
16.5%, a capture's backfill by 36%, and its replication by about 2×, while cutting CPU per document.
Changes
service-kit: skip disabled callsites statically unless a trace override is set. Everydisabled
trace!/debug!and span paid a dynamic filter check, per h2 frame and per shuffledocument. Now they're skipped statically until some handler sets a trace override.
async-process: drive child stdio with the IO driver, not the blocking pool. Connectorprotocol reads and writes no longer pay a thread handoff and futex wakeups per operation.
json: skip Frames with no keyword relevant to a validation pass. Most Frames have none.Criterion benchmarks improve 9–17%.
runtime-next: bound capture transactions by captured bytes, not combiner disk. The old64MB bound on spilled bytes closed a large backfill transaction only after it had paid to spill.
Transactions are now bounded at 64MB of captured JSON.
Testing
Covered by existing tests of each crate, including the official JSON Schema suite for
json. A newservice-kittest checks that override changes rebuild cached callsite interest.Performance
Setup. GCP
c4a-standard-8-lssd(8 arm64 vCPUs, local SSD), release builds, 3 runs per cell(mean ± sd), interleaved across commits.
demo/wikipedia/recentchange(4 journals), with each shard on its own 2-core host.
Each row is measured against the row above it:
becomes throughput.
1 shard and scale-out (measured together, before and after this PR):