Skip to content

Phase 11.8: batched fast_path counter updates - #61

Merged
jayakasadev merged 1 commit into
mainfrom
feature/phase-11-8-batched-counters
Jun 12, 2026
Merged

jayakasadev merged 1 commit into
mainfrom
feature/phase-11-8-batched-counters

Conversation

@jayakasadev

Copy link
Copy Markdown
Owner

Summary

Phase 11.8 of the SNMALLOC_STATS overhead-reduction series. Moves the
per-alloc ++stats.fast_path_allocs store out of the small-alloc fast
path into a single batched pre-credit at slab-refill time.

The slow-path (small_refill / small_refill_slow) already runs once
per slab refill; it now also knows the refill count N (number of
objects transferred from the freshly-popped slab into
fast_free_list) and credits stats.fast_path_allocs += N once,
letting the fast path skip the per-alloc store entirely.

The refill count is plumbed back from
FrontendSlabMetadata::alloc_free_list via a new uint16_t& out
parameter, computed as
sizeclass_to_slab_object_count(sizeclass) - remaining:

  • Exact for freshly-built slabs (where alloc_new_list loaded the
    builder with slab_object_count objects).
  • Upper-bound (bounded by the slab object count, at most ~256 for
    the smallest sizeclasses) for slabs recycled from the per-sizeclass
    stash.

The counter may briefly read ahead of real consumption by at most one
slab worth between refills — acceptable for observability.

Before / after (Apple M4 Pro, fat-LTO, 5 runs per variant)

Group 11.6 ratio 11.8 ratio acceptance (<=1.02)
small_allocs 1.0774 1.0155 PASS
medium_allocs 1.0398 1.0202 FAIL*
mixed 1.0310 1.0290 FAIL

* medium_allocs is within bench noise on this host (per-run-pair
median 0.9941, i.e. statistically indistinguishable from OFF once a
run-1 cold-cache outlier on the OFF side is accounted for).

A 3-run replication on a separate invocation reproduced the same
shape: small ~1.018, medium ~1.015, mixed ~1.026.

Acceptance verdict

PARTIAL.

  • small_allocs (the targeted group, where the per-alloc fast-path
    counter dominated the iteration mean) passes the strict <=1.02 spec
    cleanly at 1.0155, a ~80% reduction of the previous 1.0774
    over-budget portion.
  • medium_allocs lands at 1.0202 with the per-run-pair median in
    favour of the BASIC build.
  • mixed (1.0290) still misses the strict 1.02 spec. It blends
    large-class paths that do not benefit from the small-class batching
    done here, and still pays the symmetric per-dealloc
    fast_path_deallocs store on the dealloc hot path.

Phase 11.9 is filed as a follow-up to apply the same
single-combined-counter approach to the dealloc-side counters.

Build / test status

  • cmake -B build -DSNMALLOC_STATS_BASIC=ON && cmake --build build -j4 clean.
  • ctest -R "fast_path_counters|statistics" 4/4 pass.
  • cargo test --features stats-basic in snmalloc-rs/: full suite green.

Files touched

  • src/snmalloc/mem/corealloc.h — remove per-alloc store; add batched
    pre-credit at the two refill sites.
  • src/snmalloc/mem/metadata.h — alloc_free_list reports refill count.
  • docs/heap-profiling-benchmarks.md — Phase 11.8 section with full
    5-run tables, acceptance verdict, and reproducer.

Test plan

  • Build with -DSNMALLOC_STATS_BASIC=ON passes
  • ctest -R fast_path_counters and -R statistics pass
  • cargo test --features stats-basic passes
  • cargo bench --features stats-basic --bench stats_bench ran 3+ times
  • cargo bench --bench stats_bench (OFF baseline) ran 3+ times

Move the fast_path_allocs counter update out of the per-alloc fast path
into a single pre-credit at refill time. The slow path knows the refilled
free-list length N, so it credits fast_path_allocs += N once at
small_refill / small_refill_slow and the fast path skips the store
entirely.

Plumbed via a new uint16_t& out parameter on
FrontendSlabMetadata::alloc_free_list, computed as
sizeclass_to_slab_object_count(sizeclass) - remaining (exact for
freshly-built slabs, upper-bound for recycled slabs from the per-class
stash). Bounded by the slab object count, ~256 for the smallest classes.

Trade-off: counter may briefly overshoot true alloc count by up to N
between refills. Acceptable for observability.

Bench numbers (5 runs per variant, Apple M4 Pro, fat-LTO):
  small_allocs  1.0774 -> 1.0155  (PASS, ~80% closer to spec)
  medium_allocs 1.0398 -> 1.0202  (FAIL*, within bench noise)
  mixed         1.0310 -> 1.0290  (FAIL, untouched dealloc-side counter)

Result PARTIAL on the strict <=1.02 spec; small_allocs (the targeted
group) passes cleanly. Phase 11.9 is filed to apply the same approach
to dealloc-side counters.

See docs/heap-profiling-benchmarks.md "Phase 11.8 -- batched fast_path
counter updates" for the full table.
@jayakasadev
jayakasadev merged commit 8cde584 into main Jun 12, 2026
19 of 211 checks passed
@jayakasadev
jayakasadev deleted the feature/phase-11-8-batched-counters branch June 12, 2026 18:14
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant