Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
26 changes: 26 additions & 0 deletions snmalloc-rs/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -93,6 +93,32 @@ Then render to SVG:
inferno-flamegraph < heap.folded > heap.svg
```

### When to use snapshot vs streaming

The two profiling modes answer different questions and have different
biases. Pick the one that matches your workload:

| | `SnMalloc::snapshot()` | `ProfilingSession::start` (streaming) |
| - | - | - |
| **What it captures** | Sampled allocations *currently live* in the process at the time of the call. | *Every* sampled event (alloc / dealloc / resize) as it happens. |
| **Best for** | "What is holding memory *right now*?" — heap-state audits, leak triage, before/after diffs across a steady-state. | "Which call site is the highest-rate allocator?" — hot-path optimisation, rate-based attribution, transient-churn analysis. |
| **Bias** | Biased toward long-lived allocations — short-lived churn (allocate-and-free inside a request, scratch buffers in a tight loop) is freed before the snapshot and vanishes from view. | None on the event stream itself, but the consumer pays for storage / aggregation of every event. |
| **Output** | In-process `HeapProfile`; serialise via `write_pprof` / `write_flamegraph`. | Live callback; the application chooses how to persist events (commonly a JSON-Lines log file). |
| **Tooling** | `snmalloc-tools profile-top` for top-N live sites. | `snmalloc-tools rate-report` for per-site alloc/dealloc rate + peak-live-bytes. |

**Rule of thumb.** If the question is "where is my live heap?" use a
snapshot. If the question is "which call site is hottest and how
churny is it?" use streaming. A snapshot will systematically
under-count a hot allocate-and-free site; a streaming log captures
that churn but requires you to keep an event log around.

The `snmalloc-tools` CLI ships dedicated subcommands for each mode:
`profile-top` walks a snapshot, and `rate-report` stream-parses a
streaming event log file without loading the whole log into memory
(safe for multi-million-event traces). See
[`snmalloc-tools/README.md`](../snmalloc-tools/README.md) for the
streaming-log on-disk schema.

### Streaming mode

For long-running services, `ProfilingSession::start` registers a
Expand Down
53 changes: 51 additions & 2 deletions snmalloc-tools/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -27,10 +27,56 @@ snmalloc-tools branch-misses --perf-script <file> --hints <branch_hints.json> [-
Parse `perf script` output and cross-reference with the Phase
10.2 branch-hint inventory. High-miss-rate inverted hints are
candidates for `LIKELY` <-> `UNLIKELY` swap.

snmalloc-tools rate-report --input <streaming-log.jsonl> [--top N] [--pretty]
Stream-parse a snmalloc streaming event log (JSON Lines) and
emit a per-site row: alloc/dealloc counts, peak live bytes,
alloc-rate per second. Output is CSV by default; `--pretty`
emits a fixed-width table. Stream-based — 6M-event logs use
O(distinct sites) memory, not O(events).
```

All subcommands except `rate-report` accept `--json` for structured
output; the default is a plain-text table. `rate-report` emits CSV
by default (the friendliest format for downstream awk/jq/spreadsheet
pipelines) and a fixed-width table under `--pretty`.

## Streaming event-log schema (`rate-report`)

`rate-report` consumes **JSON Lines** (UTF-8, one event object per
line). The producer is typically an application using
[`snmalloc_rs::ProfilingSession`](../snmalloc-rs/src/streaming.rs)
that serialises each callback to a file. Schema:

```jsonl
{"ts_ns": 1000000, "kind": "alloc", "site": "0x55a0c0001000", "size": 4096}
{"ts_ns": 1001000, "kind": "dealloc", "site": "0x55a0c0001000", "size": 4096}
```

All subcommands accept `--json` for structured output; the default is
a plain-text table.
Fields:

- `ts_ns` (u64, optional) — monotonic-clock timestamp in nanoseconds.
Used to compute the alloc-rate denominator; when missing across all
records the rate column is reported as `0.0`.
- `kind` (string, required) — one of `"alloc"`, `"dealloc"`,
`"resize"`. Unknown values are skipped (forward-compat).
- `site` (string, required) — the allocation site key. Typically the
leaf-frame address as `0x` + 16 hex digits, matching the
`site_leaf` field emitted by the other subcommands.
- `size` (u64, optional) — bytes attributable to this event.

Malformed lines are skipped silently — the reader is resilient to
truncated tails and the occasional blank line. See
`tests/fixtures/streaming_log_sample.jsonl` for a worked example.

## Snapshot vs streaming

`profile-top` walks a `HeapProfile::snapshot()` (currently-live
sampled allocations) and is biased toward long-lived state;
`rate-report` walks a streaming log and captures transient churn.
See the "When to use snapshot vs streaming" section in
[`../snmalloc-rs/README.md`](../snmalloc-rs/README.md) for a fuller
treatment of the tradeoff.

## Live-process limitation (important)

Expand Down Expand Up @@ -70,6 +116,9 @@ the branch-hint inventory is a static sidecar.
- `perf_c2c_sample.txt` — two contended cache lines with detail rows.
- `branch_hints_sample.json` — three hint sites matching the schema
in `scripts/dump_branch_hints.py`.
- `streaming_log_sample.jsonl` — eight events across two sites,
exercising alloc, dealloc, resize, and the peak-then-drop pattern
that `rate-report` is built to surface.

The integration tests in `tests/integration.rs` exercise each
parser/joiner against these fixtures.
Expand Down
1 change: 1 addition & 0 deletions snmalloc-tools/src/lib.rs
Original file line number Diff line number Diff line change
Expand Up @@ -7,3 +7,4 @@ pub mod branch_hints;
pub mod joiner;
pub mod perf_c2c;
pub mod perf_script;
pub mod rate_report;
44 changes: 44 additions & 0 deletions snmalloc-tools/src/main.rs
Original file line number Diff line number Diff line change
Expand Up @@ -32,6 +32,7 @@ use snmalloc_tools::branch_hints::{BranchHintIndex, HintKind};
use snmalloc_tools::joiner;
use snmalloc_tools::perf_c2c::{self, C2cLine};
use snmalloc_tools::perf_script;
use snmalloc_tools::rate_report;

/// snmalloc-tools — CLI for joining perf PMU output with snmalloc's
/// in-tree allocation-site lookup and branch-hint inventory.
Expand All @@ -57,6 +58,9 @@ enum Cmd {
/// Cross-reference `perf script` branch-miss samples with the
/// Phase 10.2 branch-hint inventory.
BranchMisses(BranchMissesArgs),
/// Stream-parse a snmalloc streaming event log and emit a per-site
/// rate report (alloc/dealloc counts, peak live bytes, alloc rate).
RateReport(RateReportArgs),
}

#[derive(Args, Debug)]
Expand Down Expand Up @@ -121,6 +125,25 @@ struct C2cArgs {
json: bool,
}

#[derive(Args, Debug)]
struct RateReportArgs {
/// Path to the streaming event log to read. Must be JSON-Lines
/// (one event object per line). See
/// `snmalloc_tools::rate_report` module docs for the schema. The
/// file is stream-parsed -- 6M-event logs are fine.
#[arg(long)]
input: PathBuf,
/// Limit to the top-N highest-alloc-count sites. `0` means "no
/// limit"; the report still arrives sorted by alloc-count desc.
#[arg(long, default_value_t = 0)]
top: usize,
/// Render as a fixed-width pretty table instead of CSV. The
/// default (CSV) is the friendliest format for downstream
/// awk/jq/spreadsheet pipelines.
#[arg(long)]
pretty: bool,
}

#[derive(Args, Debug)]
struct BranchMissesArgs {
/// Path to the `perf script` output to parse.
Expand All @@ -146,6 +169,7 @@ fn main() -> Result<()> {
PmuJoinKind::C2c(c) => run_c2c(c),
},
Cmd::BranchMisses(a) => run_branch_misses(a),
Cmd::RateReport(a) => run_rate_report(a),
}
}

Expand Down Expand Up @@ -375,3 +399,23 @@ fn run_branch_misses(args: BranchMissesArgs) -> Result<()> {
Ok(())
}

// -- rate-report ----------------------------------------------------------

fn run_rate_report(args: RateReportArgs) -> Result<()> {
let mut rows = rate_report::read_path(&args.input)?;
if args.top > 0 && rows.len() > args.top {
rows.truncate(args.top);
}
// Writing to a locked stdout once per run is materially faster
// than repeated `println!` for large reports, and matters when a
// user pipes the output through downstream tools.
let stdout = std::io::stdout();
let mut out = stdout.lock();
if args.pretty {
rate_report::write_pretty(&rows, &mut out)?;
} else {
rate_report::write_csv(&rows, &mut out)?;
}
Ok(())
}

Loading