Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
156 changes: 156 additions & 0 deletions docs/heap-profiling-benchmarks.md
Original file line number Diff line number Diff line change
Expand Up @@ -1196,3 +1196,159 @@ For the 5-run sweep used to produce the tables above, wipe
`target/criterion` and copy the snapshot to a separate
directory between runs (criterion otherwise overwrites
`new/estimates.json`).

## Phase 11.8 -- batched fast_path counter updates

ClickUp ticket [86aj0zwv1](https://app.clickup.com/t/86aj0zwv1)
("Phase 11.8 -- Batched fast_path counter updates") removes the
per-alloc `++stats.fast_path_allocs` store from the hot path in
`small_alloc`. The counter is now pre-credited in batch at slab
refill time (in `small_refill` and `small_refill_slow`) by the
number of objects transferred from the freshly-popped slab into
`fast_free_list`. The slow-path `++stats.slow_path_allocs` site
at the top of `small_refill` is unchanged.

The pre-credit count is computed inside
`FrontendSlabMetadata::alloc_free_list` as
`sizeclass_to_slab_object_count(sizeclass) - remaining` (where
`remaining` is the unused half of the random-preserve builder)
and reported back via a new `uint16_t&` out parameter. This is
exact for freshly-built slabs (where `alloc_new_list` loaded
the builder with `slab_object_count` objects), and an upper
bound bounded by the slab object count (at most ~256 for the
smallest sizeclasses) for slabs recycled from
`alloc_classes[sizeclass].available`. The trade-off is a
small, bounded stale-ahead reading on `fast_path_allocs` -- the
counter can read up to one slab worth ahead of real
consumption -- which is acceptable for observability.

### Motivation

Phase 11.6 measured the BASIC tier at **1.077** on
`small_allocs`, identifying the per-alloc store of
`fast_path_allocs` (and its symmetric `fast_path_deallocs`) as
the irreducible-with-current-design floor. The batched
approach amortises this store over a full slab refill -- one
store per ~slab_object_count consumes instead of one per
consume -- and should bring the BASIC overhead under the
strict 1.02 spec target on the dominant hot path.

### Methodology

Same harness as Phase 11.6 above (3s warm-up, 5s measure, 50
samples, 64-alloc + 64-dealloc per iteration, three groups,
Apple M4 Pro / macOS 26.3.1 / rustc 1.95.0, release fat-LTO),
5 runs per variant. Only the BASIC and OFF variants are
re-measured here; the FULL tier is unaffected by the change
(its hot-path stores -- per-class histogram bumps -- are gated
on `SNMALLOC_STATS_FULL` and were left in place).

### Raw 5-run numbers (per criterion iteration, ns)

#### `small_allocs` (32-byte allocations)

| Run | off (ns) | basic (ns) | basic/off |
|----:|---------:|-----------:|----------:|
| 1 | 198.624 | 203.000 | 1.0220 |
| 2 | 200.159 | 203.102 | 1.0147 |
| 3 | 199.980 | 204.100 | 1.0206 |
| 4 | 200.825 | 202.990 | 1.0108 |
| 5 | 200.022 | 201.937 | 1.0096 |

5-run summary: off mean **199.922** (sd 0.717) · basic mean
**203.026** (sd 0.685) · **ratio of means basic/off = 1.0155**
· median per-run ratio 1.0147.

#### `medium_allocs` (4 KiB allocations)

| Run | off (ns) | basic (ns) | basic/off |
|----:|---------:|-----------:|----------:|
| 1 | 894.037 | 1011.647 | 1.1315 |
| 2 | 1043.061 | 1028.041 | 0.9856 |
| 3 | 1033.376 | 1026.142 | 0.9930 |
| 4 | 1022.219 | 1033.939 | 1.0115 |
| 5 | 1019.569 | 1013.512 | 0.9941 |

5-run summary: off mean **1002.452** (sd 54.851) · basic mean
**1022.656** (sd 8.640) · **ratio of means basic/off = 1.0202**
· median per-run ratio 0.9941.

Run 1's off-side baseline measurement (894 ns) is a cold-cache
outlier roughly 14% below the other four off-side runs
(1019-1043 ns) -- the per-run-pair median ratio of **0.9941**
indicates the BASIC build is statistically indistinguishable
from the OFF build on this group once the warm-up outlier is
discounted.

#### `mixed` (LCG-driven sizes in `[16, 16384)`)

| Run | off (ns) | basic (ns) | basic/off |
|----:|---------:|-----------:|----------:|
| 1 | 570.954 | 597.456 | 1.0464 |
| 2 | 582.486 | 607.149 | 1.0423 |
| 3 | 599.498 | 606.247 | 1.0113 |
| 4 | 586.722 | 607.238 | 1.0350 |
| 5 | 592.821 | 599.306 | 1.0109 |

5-run summary: off mean **586.496** (sd 9.662) · basic mean
**603.480** (sd 4.218) · **ratio of means basic/off = 1.0290**
· median per-run ratio 1.0350.

### Acceptance

| Group | 5-run mean ratio (11.6) | 5-run mean ratio (11.8) | acceptance (<=1.02) |
|-----------------|------------------------:|------------------------:|:-------------------:|
| `small_allocs` | 1.0774 | 1.0155 | **PASS** |
| `medium_allocs` | 1.0398 | 1.0202 | **FAIL**\* |
| `mixed` | 1.0310 | 1.0290 | **FAIL** |

\* Within bench noise on this host; the per-run-pair median is
0.9941, indicating no measurable overhead vs OFF on
`medium_allocs`.

**Result: PARTIAL.** The targeted `small_allocs` group, where
the per-alloc fast-path counter dominates the iteration mean,
now sits at **1.0155** -- comfortably under the strict 1.02
spec target and a **~80% reduction** of the previous 1.0774
over-budget portion (0.0774 -> 0.0155). The `medium_allocs`
result (1.0202) is right at the bench-noise floor (run-1
off-side outlier inflates the mean) and the per-run-pair
median is in favour of the BASIC build. The `mixed` group
sits at **1.0290** -- still above the strict 1.02 target.
`mixed` blends 16-16384 byte allocations, of which a sizeable
fraction routes through medium/large paths that do not benefit
from the small-class batching done here.

### Why `mixed` did not fully close

The batched pre-credit lives entirely inside the small-class
slab refill path. Allocations that route to large-class /
backend chunk allocation do not touch
`small_refill`/`small_refill_slow` and therefore do not bump
`fast_path_allocs`. The remaining `mixed`-group delta vs OFF
is the cost of the symmetric per-dealloc `fast_path_deallocs`
counter (still per-alloc on the dealloc hot path), the
`bytes_in_use` atomics used for backend accounting on
large-class allocations, and the message-queue counter stores
on cross-thread free paths. None of these are addressed by
Phase 11.8.

Phase 11.9 is filed as a follow-up to apply the same
single-combined-counter approach to the dealloc-side counters
(and optionally collapse the four fast/slow alloc/dealloc
counters into one `total_allocs` counter, deriving fast =
total - slow at query time).

### Reproducing

```bash
cd snmalloc-rs
# OFF baseline
cargo bench --bench stats_bench
# BASIC tier
cargo bench --features stats-basic --bench stats_bench
# Output lands in target/criterion/<group>/{stats-off,stats-basic}/new/estimates.json
```

For the 5-run sweep wipe `target/criterion` (or copy
`new/estimates.json` aside) between runs.
62 changes: 52 additions & 10 deletions src/snmalloc/mem/corealloc.h
Original file line number Diff line number Diff line change
Expand Up @@ -1045,11 +1045,7 @@ namespace snmalloc
auto* fl = &small_fast_free_lists[sizeclass];
if (SNMALLOC_LIKELY(!fl->empty()))
{
#ifdef SNMALLOC_STATS_BASIC
// Phase 9.2 -- fast-path alloc counter. Non-atomic write
// against the per-thread `Allocator::stats`.
stats.fast_path_allocs++;
# ifdef SNMALLOC_STATS_FULL
#ifdef SNMALLOC_STATS_FULL
// Phase 9.3 -- per-size-class histogram. The sizeclass is
// already in a register here.
//
Expand All @@ -1065,8 +1061,17 @@ namespace snmalloc
// floor for the 1.16 small_allocs regression in 11.5.
sc_stats.live_count[sizeclass]++;
sc_stats.live_bytes[sizeclass] += sizeclass_to_size(sizeclass);
# endif
#endif
// Phase 11.8 -- `++stats.fast_path_allocs` was removed from
// this site. The counter is now pre-credited in batch at
// `small_refill`/`small_refill_slow` time by the number of
// objects transferred into `fast_free_list`. This removes
// the per-alloc store from the hot path and brings the
// SNMALLOC_STATS_BASIC small_allocs overhead under the
// strict <=1.02 spec target. The counter may briefly read
// ahead of real consumption, bounded by the slab object
// count (at most ~256), which is acceptable for
// observability.
auto p = fl->take(key, domesticate);
return finish_alloc<Conts>(p, size);
}
Expand Down Expand Up @@ -1243,8 +1248,14 @@ namespace snmalloc
[this](freelist::QueuePtr p) SNMALLOC_FAST_PATH_LAMBDA {
return capptr_domesticate<Config>(backend_state_ptr(), p);
};
uint16_t refill_count = 0;
auto [p, still_active] = BackendSlabMetadata::alloc_free_list(
domesticate, meta, fast_free_list, entropy, sizeclass);
domesticate,
meta,
fast_free_list,
entropy,
sizeclass,
refill_count);

if (still_active)
{
Expand All @@ -1256,7 +1267,22 @@ namespace snmalloc
laden.insert(meta);
}

#ifdef SNMALLOC_STATS_FULL
#ifdef SNMALLOC_STATS_BASIC
// Phase 11.8 -- batched fast-path alloc pre-credit. The
// refill just transferred `refill_count` objects to
// `fast_free_list` (including `p`, which is returned to the
// caller without going through the fast path); subsequent
// consumes from the fast path will not bump the counter.
// This trades a small bounded overshoot (at most one slab
// worth, ~256 objects for the smallest sizeclasses) for
// removing the per-alloc store on the hot path.
//
// The slow-path bump (`++slow_path_allocs` at the top of
// `small_refill`) already accounts for this call as a slow
// refill; the `refill_count` credit is the symmetric
// batched fast-path entry.
stats.fast_path_allocs += refill_count;
# ifdef SNMALLOC_STATS_FULL
// Phase 9.3 -- slow-path-from-stash alloc bump. We have
// taken one object from the freshly-popped slab's freelist;
// any remaining objects on `fast_free_list` will be
Expand All @@ -1269,6 +1295,7 @@ namespace snmalloc
// time, so only the live counters are bumped here.
sc_stats.live_count[sizeclass]++;
sc_stats.live_bytes[sizeclass] += sizeclass_to_size(sizeclass);
# endif
#endif
auto r = finish_alloc<Conts>(p, size);
return ticker.check_tick(r);
Expand Down Expand Up @@ -1318,8 +1345,14 @@ namespace snmalloc
[this](freelist::QueuePtr p) SNMALLOC_FAST_PATH_LAMBDA {
return capptr_domesticate<Config>(backend_state_ptr(), p);
};
uint16_t refill_count = 0;
auto [p, still_active] = BackendSlabMetadata::alloc_free_list(
domesticate, meta, fast_free_list, entropy, sizeclass);
domesticate,
meta,
fast_free_list,
entropy,
sizeclass,
refill_count);

if (still_active)
{
Expand All @@ -1331,7 +1364,15 @@ namespace snmalloc
laden.insert(meta);
}

#ifdef SNMALLOC_STATS_FULL
#ifdef SNMALLOC_STATS_BASIC
// Phase 11.8 -- batched fast-path alloc pre-credit (see
// matching note in `small_refill`). For a freshly-built
// slab the refill count is exact: the builder was
// populated with `slab_object_count` objects by
// `alloc_new_list`, of which `slab_object_count -
// remaining` were transferred to `fast_free_list`.
stats.fast_path_allocs += refill_count;
# ifdef SNMALLOC_STATS_FULL
// Phase 9.3 -- slow-path-from-backend alloc bump. This
// path has just brought in a fresh slab from the backend
// and taken the first object from it; the remaining
Expand All @@ -1342,6 +1383,7 @@ namespace snmalloc
// time, so only the live counters are bumped here.
sc_stats.live_count[sizeclass]++;
sc_stats.live_bytes[sizeclass] += sizeclass_to_size(sizeclass);
# endif
#endif
auto r = finish_alloc<Conts>(p, size);
return ticker.check_tick(r);
Expand Down
33 changes: 27 additions & 6 deletions src/snmalloc/mem/metadata.h
Original file line number Diff line number Diff line change
Expand Up @@ -624,13 +624,25 @@ namespace snmalloc
/**
* Allocates a free list from the meta data.
*
* Returns a freshly allocated object of the correct size, and a bool that
* Returns a freshly allocated object of the correct size, a bool that
* specifies if the slab metadata should be placed in the queue for that
* sizeclass.
* sizeclass, and an upper-bound refill count (the number of objects
* transferred to `fast_free_list`, including the popped return value).
*
* If Randomisation is not used, it will always return false for the second
* component, but with randomisation, it may only return part of the
* available objects for this slab metadata.
* The refill count is `sizeclass_to_slab_object_count(sizeclass) -
* remaining`. This is exact for freshly-built slabs (where the builder
* was populated with `slab_object_count` objects via `alloc_new_list`),
* and an upper bound when the slab is reused from the per-sizeclass
* stash (a recycled slab may have had fewer than `slab_object_count`
* entries enqueued). The overshoot is bounded by the slab object count
* (at most ~256 for the smallest sizeclasses) and is consumed by the
* Phase 11.8 batched `fast_path_allocs` pre-credit, which permits a
* bounded stale-ahead reading for observability.
*
* If Randomisation is not used, the second component will always be
* false (the closed list contains everything in the builder), but with
* randomisation, it may only return part of the available objects for
* this slab metadata.
*/
template<typename Domesticator>
static SNMALLOC_FAST_PATH stl::Pair<freelist::HeadPtr, bool>
Expand All @@ -639,7 +651,8 @@ namespace snmalloc
FrontendSlabMetadata* meta,
freelist::Iter<>& fast_free_list,
LocalEntropy& entropy,
smallsizeclass_t sizeclass)
smallsizeclass_t sizeclass,
uint16_t& refill_count)
{
auto& key = freelist::Object::key_root;

Expand All @@ -661,6 +674,14 @@ namespace snmalloc
// This will be zero if there is no randomisation.
auto sleeping = meta->set_sleeping(sizeclass, remaining);

// Phase 11.8: report the refill count for batched
// `fast_path_allocs` pre-credit. Computed as
// `slab_object_count - remaining`; exact for freshly-built
// slabs and an upper bound (bounded by slab object count) for
// recycled slabs from the per-sizeclass stash.
refill_count = static_cast<uint16_t>(
sizeclass_to_slab_object_count(sizeclass) - remaining);

return {p, !sleeping};
}

Expand Down
Loading