Skip to content

runtime-next: hot-path optimizations - #3575

Merged
jgraettinger merged 4 commits into
masterfrom
johnny/runtime-hot-paths
Oct 6, 2026
Merged

jgraettinger merged 4 commits into
masterfrom
johnny/runtime-hot-paths

Conversation

@jgraettinger

Copy link
Copy Markdown
Member

Summary

Four independent hot-path optimizations. Together they speed up a 3-shard materialization by
16.5%, a capture's backfill by 36%, and its replication by about 2×, while cutting CPU per document.

Changes

  • service-kit: skip disabled callsites statically unless a trace override is set. Every
    disabled trace!/debug! and span paid a dynamic filter check, per h2 frame and per shuffle
    document. Now they're skipped statically until some handler sets a trace override.
  • async-process: drive child stdio with the IO driver, not the blocking pool. Connector
    protocol reads and writes no longer pay a thread handoff and futex wakeups per operation.
  • json: skip Frames with no keyword relevant to a validation pass. Most Frames have none.
    Criterion benchmarks improve 9–17%.
  • runtime-next: bound capture transactions by captured bytes, not combiner disk. The old
    64MB bound on spilled bytes closed a large backfill transaction only after it had paid to spill.
    Transactions are now bounded at 64MB of captured JSON.

Testing

Covered by existing tests of each crate, including the official JSON Schema suite for json. A new
service-kit test checks that override changes rebuild cached callsite interest.

Performance

Setup. GCP c4a-standard-8-lssd (8 arm64 vCPUs, local SSD), release builds, 3 runs per cell
(mean ± sd), interleaved across commits.

  • materialize-sink: catches up on about 16M docs (22.7 GB) of demo/wikipedia/recentchange
    (4 journals), with each shard on its own 2-core host.
  • capture-fake-postgres: 1 shard, backfilling 7.2M docs and then replicating.

Each row is measured against the row above it:

commit materialize, 3 shards CPU per M docs capture backfill capture replication
(base) 181.4k ± 3.8k docs/s 27.5 CPU-s 117.8k ± 0.6k docs/s 79.8k ± 1.0k docs/s
service-kit 196.6k ± 0.7k (+8.4%) 25.0 (−9.1%) 145.2k ± 1.1k (+23%) 140.3k ± 0.8k (+76%)
async-process 208.3k ± 0.9k (+5.9%) 23.7 (−5.4%) 142.4k ± 0.9k (−1.9%) 150.3k ± 0.9k (+7.2%)
json 211.3k ± 0.6k (+1.5%) 23.3 (−1.5%) 149.9k ± 1.8k (+5.3%) 158.4k ± 0.4k (+5.3%)
capture txn bound n/a 160.4k ± 1.1k (+7.1%) 157.9k ± 1.6k (no change)
cumulative +16.5% −15% +36% +98%
  • The capture gains most because it's bound by one shard thread, so every CPU saving there
    becomes throughput.
  • async-process mostly helps the materialization, halving its context switches (90k → 49k/s).
  • The capture transaction bound trims the largest backfill transactions from 110MB to 77MB.

1 shard and scale-out (measured together, before and after this PR):

before after Δ
materialize, 1 shard 69.3k ± 2.7k docs/s 74.4k ± 2.7k docs/s +7.5%
materialize, 3 shards 180.8k ± 3.6k docs/s 213.1k ± 2.1k docs/s +17.8%
scale-out (3 shards ÷ 3 × 1 shard) 0.87 0.95
  • One shard gains less because it's bound by its runtime thread.
  • Three shards are bound by the host whose Slice reads 2 of the 4 journals, at 1.89 of 2 cores.

…de is set

OverrideFilter reported `Interest::sometimes` and a TRACE level hint for
every callsite, so the process's max level was TRACE and every disabled
`trace!`, `debug!`, and span went through a dynamic `enabled()` check:
EnvFilter, then a current-span lookup (sharded_slab CAS) and scope walk.
h2 opens spans per frame, and the shuffle Slice and Log actors trace per
document, so this ran on the hottest paths of a sidecar.

Now a process-wide count tracks live handlers with an override set. While
it's zero, OverrideFilter reports `Interest::never` and an INFO hint (which
keeps handler `info_span!`s alive), so disabled callsites are skipped
statically. Setting the first override, or clearing the last (including
by its handler finishing), rebuilds the callsite interest cache.

Performance (runtime-lab, mean ± sd of n=3 interleaved runs, vs. parent):
  materialize-sink, 3 shards: 196.6k ± 0.7k docs/s (+8.4%), 25.0 CPU-s/M docs (-9.1%)
  capture-fake-postgres, 1 shard: backfill 145.2k ± 1.1k docs/s (+23%), replication 140.3k ± 0.8k docs/s (+76%)
…ocking pool

`ChildStdio` was `tokio::fs::File`, which runs every read and write as a
blocking-pool task: a thread handoff, a buffer copy, and futex wakeups per
operation. It carries the connector protocol, so a materialization shard
paid that per ~32KB of C:Store requests and per read of the connector's
responses.

`ChildStdio` is replaced by `ChildStdin` and `ChildOutput` (stdout and
stderr), aliases of tokio's non-blocking pipe Sender and Receiver. They
register with the IO driver of the runtime current at `Child::from`, and
a direction mismatch is now a type error rather than a runtime EBADF.

The crate is now unix-only: the non-unix path never built (it imported
`std::os::fd::OwnedHandle`), and Windows is not a target.

A pipe Sender's `flush` and `shutdown` are no-ops, so materialize-consistency
drops its `shutdown` calls; as before, dropping the pipe is what sends EOF.

Performance (runtime-lab, mean ± sd of n=3 interleaved runs, vs. parent):
  materialize-sink, 3 shards: 208.3k ± 0.9k docs/s (+5.9%), 23.7 CPU-s/M docs (-5.4%), context switches -46%
  capture-fake-postgres, 1 shard: backfill 142.4k ± 0.9k docs/s (-1.9%), replication 150.3k ± 0.9k docs/s (+7.2%)
Validation visits each document node with every active Frame, and each
pass (items, properties, strings, numbers, containers, nodes, unwind)
scanned all of a Frame's keywords looking for ones it handles. Most
Frames, especially in-place applications like $ref and allOf, have none.

wind_frame now folds a Frame's keywords into KIND_* classes. Frame keeps
the eight classes the per-node passes check in a `u8`, so Frame stays one
cache line, and each pass skips Frames lacking its class. Two further
classes (in-place applications, and unevaluatedItems/Properties) are only
needed while winding. Retained classes are `u8` and wind-only classes `u16`,
so testing a wind-only class against Frame::kinds is a type error.

unwind_frame now unwinds the top Frame in place and then truncates it,
rather than popping it by value: the Frame was typically just written by
wind_frame, and moving it whole stalled on those pending stores.

Criterion medians, before -> after:
  citi_rides  rides1x   3.36 ms -> 2.81 ms  (-16%)
              rides4x  14.38 ms -> 11.99 ms (-17%)
  github      scrape0   202 µs -> 182 µs  (-10%)
              scrape1   230 µs -> 208 µs  (-10%)
              scrape2   231 µs -> 209 µs  (-9%)
              scrape3   220 µs -> 198 µs  (-10%)

Performance (runtime-lab, mean ± sd of n=3 interleaved runs, vs. parent):
  materialize-sink, 3 shards: 211.3k ± 0.6k docs/s (+1.5%), 23.3 CPU-s/M docs (-1.5%)
  capture-fake-postgres, 1 shard: backfill 149.9k ± 1.8k docs/s (+5.3%), replication 158.4k ± 0.4k docs/s (+5.3%)
…ner disk

A capture closed its transaction once the combiner had spilled 64MB to
disk, so a large backfill transaction closed only after it had already
paid to spill: serialization, compression, and IO of documents which are
then read back to drain. Backfill transactions ran up to ~110MB of
captured JSON.

Instead bound transactions at 64MB of captured document bytes, which
keeps the combiner within memory and favors small transactions, and leave
combiner disk usage at the close policy's default. The bound is evaluated
at connector checkpoints, so a transaction may exceed it by a checkpoint's
worth of documents.

Performance (runtime-lab, mean ± sd of n=3 interleaved runs, vs. parent):
  capture-fake-postgres, 1 shard: backfill 160.4k ± 1.1k docs/s (+7.1%), replication 157.9k ± 1.6k docs/s (no change)
  materialize-sink: not applicable (capture-only change)
@strix-security

strix-security Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

Strix Security Review

No security issues found.

Review summary

Reviewed all 11 changed files in this performance-optimization PR. The changes are behavior-preserving hot-path optimizations across four crates: static callsite-interest skipping in service-kit tracing, IO-driver-based child stdio in async-process, keyword-class-based Frame skipping in the JSON validator, and a capture transaction bound switch in runtime-next. The security-sensitive area was the JSON validator's new keyword_kinds classification, which gates which validation passes examine each Frame; I verified the classification is exhaustive over the Keyword enum and that every variant's class exactly matches the pass that enforces it, so no schema constraint can be silently skipped (the u16-only KIND_IN_PLACE/KIND_UNEVALUATED bits are only tested against the untruncated value, so the kinds as u8 truncation is safe). The async-process stdio switch, the trace-override atomic counting/rebuild logic, and the transaction-bound change were also reviewed and found correct with no new attacker-reachable path. No security vulnerabilities were identified.

Updated for e331fbc.


Reviewed by Strix
Re-run review · Configure security review settings

@jgraettinger
jgraettinger requested a review from a team October 3, 2026 18:09
@jgraettinger

Copy link
Copy Markdown
Member Author

There are a bunch of other performance levers to pull, after further investigation using runtime-lab, but these are the ones that were unambiguous wins in the time I had available, without any downsides that I can see.

Adding worker threads is not an obvious improvement (for captures, in particular). It requires a deeper dive into how tasks yield, when, and how tasks bounce between tokio worker threads, and also splitting up the monolith capture shard actor so that draining can be parallel to loading a next combiner.

@dgreer-dev dgreer-dev left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@jgraettinger
jgraettinger merged commit a65cc98 into master Oct 6, 2026
12 checks passed
@jgraettinger
jgraettinger deleted the johnny/runtime-hot-paths branch October 6, 2026 14:35
@github-actions github-actions Bot added pending:agent Merged, in the control-plane-agent image, and not yet rolled to flow-agent pending:flowctl Merged, changes the flowctl binary, and not in a published release pending:agent-api Merged, ships via Deploy agent-api, and not yet deployed labels Oct 6, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

pending:agent Merged, in the control-plane-agent image, and not yet rolled to flow-agent pending:agent-api Merged, ships via Deploy agent-api, and not yet deployed pending:flowctl Merged, changes the flowctl binary, and not in a published release

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants