Phase 11.8: batched fast_path counter updates - #61
Merged
Merged
Conversation
Move the fast_path_allocs counter update out of the per-alloc fast path into a single pre-credit at refill time. The slow path knows the refilled free-list length N, so it credits fast_path_allocs += N once at small_refill / small_refill_slow and the fast path skips the store entirely. Plumbed via a new uint16_t& out parameter on FrontendSlabMetadata::alloc_free_list, computed as sizeclass_to_slab_object_count(sizeclass) - remaining (exact for freshly-built slabs, upper-bound for recycled slabs from the per-class stash). Bounded by the slab object count, ~256 for the smallest classes. Trade-off: counter may briefly overshoot true alloc count by up to N between refills. Acceptable for observability. Bench numbers (5 runs per variant, Apple M4 Pro, fat-LTO): small_allocs 1.0774 -> 1.0155 (PASS, ~80% closer to spec) medium_allocs 1.0398 -> 1.0202 (FAIL*, within bench noise) mixed 1.0310 -> 1.0290 (FAIL, untouched dealloc-side counter) Result PARTIAL on the strict <=1.02 spec; small_allocs (the targeted group) passes cleanly. Phase 11.9 is filed to apply the same approach to dealloc-side counters. See docs/heap-profiling-benchmarks.md "Phase 11.8 -- batched fast_path counter updates" for the full table.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Phase 11.8 of the SNMALLOC_STATS overhead-reduction series. Moves the
per-alloc
++stats.fast_path_allocsstore out of the small-alloc fastpath into a single batched pre-credit at slab-refill time.
The slow-path (
small_refill/small_refill_slow) already runs onceper slab refill; it now also knows the refill count
N(number ofobjects transferred from the freshly-popped slab into
fast_free_list) and creditsstats.fast_path_allocs += Nonce,letting the fast path skip the per-alloc store entirely.
The refill count is plumbed back from
FrontendSlabMetadata::alloc_free_listvia a newuint16_t&outparameter, computed as
sizeclass_to_slab_object_count(sizeclass) - remaining:alloc_new_listloaded thebuilder with
slab_object_countobjects).the smallest sizeclasses) for slabs recycled from the per-sizeclass
stash.
The counter may briefly read ahead of real consumption by at most one
slab worth between refills — acceptable for observability.
Before / after (Apple M4 Pro, fat-LTO, 5 runs per variant)
small_allocsmedium_allocsmixed*
medium_allocsis within bench noise on this host (per-run-pairmedian 0.9941, i.e. statistically indistinguishable from OFF once a
run-1 cold-cache outlier on the OFF side is accounted for).
A 3-run replication on a separate invocation reproduced the same
shape: small ~1.018, medium ~1.015, mixed ~1.026.
Acceptance verdict
PARTIAL.
small_allocs(the targeted group, where the per-alloc fast-pathcounter dominated the iteration mean) passes the strict <=1.02 spec
cleanly at 1.0155, a ~80% reduction of the previous 1.0774
over-budget portion.
medium_allocslands at 1.0202 with the per-run-pair median infavour of the BASIC build.
mixed(1.0290) still misses the strict 1.02 spec. It blendslarge-class paths that do not benefit from the small-class batching
done here, and still pays the symmetric per-dealloc
fast_path_deallocsstore on the dealloc hot path.Phase 11.9 is filed as a follow-up to apply the same
single-combined-counter approach to the dealloc-side counters.
Build / test status
cmake -B build -DSNMALLOC_STATS_BASIC=ON && cmake --build build -j4clean.ctest -R "fast_path_counters|statistics"4/4 pass.cargo test --features stats-basicinsnmalloc-rs/: full suite green.Files touched
src/snmalloc/mem/corealloc.h— remove per-alloc store; add batchedpre-credit at the two refill sites.
src/snmalloc/mem/metadata.h—alloc_free_listreports refill count.docs/heap-profiling-benchmarks.md— Phase 11.8 section with full5-run tables, acceptance verdict, and reproducer.
Test plan
-DSNMALLOC_STATS_BASIC=ONpassesctest -R fast_path_countersand-R statisticspasscargo test --features stats-basicpassescargo bench --features stats-basic --bench stats_benchran 3+ timescargo bench --bench stats_bench(OFF baseline) ran 3+ times