A command-line benchmark for comparing automatic speech recognition (ASR) models on your own audio.
Built for professors, researchers, and accessibility teams who need to pick an ASR engine for their actual content — not whatever was on a public leaderboard last year. Outputs markdown tables you can paste into a doc.
- What it is — a single-file Python CLI that benchmarks local speech-to-text engines on your audio and your hardware, writing one markdown report per run (WER, speed/RTFx, wall-clock, peak VRAM, disk size).
- Why — model rankings shift with your content and your GPU. Measure them yourself instead of trusting a public leaderboard.
- Engines — four Whisper variants via faster-whisper (the stable core), plus an experimental NVIDIA NIM engine (Riva gRPC) that's implemented but not yet validated against a live server. See NVIDIA NIM engine (optional).
- Headline result — on a 12-lecture, single-speaker corpus (RTX 5090), Large V3 Turbo won on both accuracy (8.9% WER) and speed (~65× realtime) — unusual, and exactly the kind of per-corpus surprise this tool surfaces. See Example output.
- Bottom line — run it on your own audio, or just read the numbers below to get a feel for the accuracy/speed/VRAM tradeoffs.
- License — MIT. Use it freely; keep the copyright notice.
v0.3.5 (on main). Local Whisper variants + optional fusion + WhisperX word alignment and speaker diarization; pip-installable with an asr-bench console command:
small(244M, ~470MB)medium(769M, ~1.5GB)large-v3(1550M, ~3.1GB)large-v3-turbo(809M, ~1.6GB)<size>+whisperx— any of the above paired with WhisperX alignment + diarization (e.g.large-v3-turbo+whisperx)
New in v0.2: MER/WIL metrics, per-clip S/D/I counts, --show-alignment, and the
--fuse post-processing stage (verbatim captions + RAG knowledge base). See
What it measures and Fusion below.
New in v0.3: WhisperX word-level alignment, speaker-labeled VTT, DER scoring, and word-timestamp sidecars. See WhisperX (word alignment + diarization) below.
See SPEC.md for the full roadmap (Canary-Qwen + NVIDIA NeMo, multi-language coverage, hand-corrected reference sets).
| Metric | What it means |
|---|---|
| WER% | Word Error Rate vs your reference transcript |
| CER% | Character Error Rate (jiwer) — finer-grained companion to WER; counts character edits rather than whole-word edits |
| MER% | Match Error Rate — fraction of reference+hypothesis words that are errors (bounded [0,1]) |
| WIL% | Word Information Lost — information-theoretic complement to WER (bounded [0,1]) |
| S/D/I | Per-clip substitution / deletion / insertion counts |
| RTFx | Audio seconds processed per wall-clock second (higher = faster than realtime) |
| RTFx (med) | Median per-clip RTFx — robust to a single slow/locked clip that skews the totals-based aggregate RTFx |
| Wall clock | Total processing time |
| Peak VRAM | NVIDIA GPU memory peak during transcription (requires nvidia-ml-py) |
| Disk size | Model file size after first download |
| Params | Model parameter count |
WER is the load-bearing metric and requires a reference transcript. If your reference is hand-corrected gold standard, the numbers are defensible. If you use auto-generated captions (like Panopto exports) as the reference, treat the WER as a relative divergence rate rather than an absolute accuracy score.
MER and WIL (Morris, Maier & Green 2004) sit alongside WER in every table.
Both are bounded [0, 1] and information-theoretic: WIL in particular captures
how much information was lost rather than just counting token-level edits,
which makes it a better proxy for comprehension quality on lecture content. Use
--show-alignment to print per-clip alignment diffs when you want to see
exactly what was substituted, deleted, or inserted.
# Python 3.10+ recommended
python -m pip install -r requirements.txtOr install it as a package to get the asr-bench console command:
pip install -e . # core (CPU) deps only — torch-free
pip install -e ".[gpu]" # + NVML VRAM tracking + CUDA wheels
pip install -e ".[nim]" # + nvidia-riva-client for NIM runs
asr-bench --versionasr-bench ... is equivalent to python asr_bench.py ... (same for the
compare and prepare-gold subcommands). The nvidia-ml-py dep is optional but
enables peak-VRAM tracking; without it the VRAM column shows n/a. For the fusion
stage, Ollama is an additional optional dependency (only
needed with --llm ollama:...). WhisperX is not a pip extra — it needs a
separate Python ≤ 3.13 venv (see the WhisperX section below).
Put audio files and their reference transcripts in matching pairs under one folder. Three layouts are recognized:
Layout A — flat folder, name-matched pairs:
test-corpus/
lecture-week-3.mp4
lecture-week-3.txt # reference transcript
lecture-week-4.mp4
lecture-week-4.txt
Layout B — Panopto export shape (auto-detected):
test-corpus/
Lecture 12_default.mp4
Lecture 12_Captions_English (United States).txt
Layout C — explicit pairing via manifest.json:
{
"clips": [
{"audio": "wk03.mp4", "reference": "wk03-corrected.txt"},
{"audio": "wk04.mp4", "reference": "wk04-corrected.txt"}
]
}Reference files can be:
- Plain text (one transcript per file)
- SRT-shaped (Panopto exports use this — timestamps + cue numbers stripped automatically)
- WebVTT (
.vtt)
Converting captions to references (prepare-gold): if your references are
.vtt/.srt caption files, turn them into the plain .txt the scorer reads:
python asr_bench.py prepare-gold --corpus ./test-corpus # *.vtt/*.srt -> *.txt
python asr_bench.py prepare-gold lecture.vtt --overwrite # one file
python asr_bench.py prepare-gold ./corpus --dry-run # preview onlyTiming and cue numbers are stripped and cues joined. asr-bench's own
_Captions_*.vtt outputs are skipped (so a model's output can't become its own
reference). Auto-generated / Panopto sources keep a proxy marker so they stay
flagged as not gold — proofread the resulting .txt and delete that header
line to promote it to a true gold reference.
python asr_bench.py --corpus ./test-corpusWith explicit model selection:
python asr_bench.py --corpus ./test-corpus --models small,medium,large-v3Output goes to stdout (markdown) and ./report/<timestamp>.md.
The Whisper variants will use CUDA if available; otherwise CPU. Force CPU
explicitly with --device cpu. Expect large-v3 to take 5-10× the audio
duration on CPU — start it overnight.
A real run over a 12-lecture corpus (~614 min ≈ 10.2 hours of single-speaker lecture audio) on an NVIDIA RTX 5090, default settings. The reference here was auto-generated captions, so these WER numbers are relative divergence between engines, not absolute accuracy — and lecture names are anonymized.
You don't have to run a full benchmark yourself to get a feel for the tradeoffs — here's what these four models did on that ~10 hours of audio:
| Model | Disk | Overall WER% | Speed (RTFx) | Total time | Peak VRAM |
|---|---|---|---|---|---|
| Whisper Small | 1.4 GB | 10.7 | 43.5× | 14.1 min | 372 MB |
| Whisper Medium | 1.4 GB | 11.8 | 29.1× | 21.1 min | 269 MB |
| Whisper Large V3 | 8.6 GB | 14.2 | 14.7× | 41.9 min | 1.2 GB |
| Whisper Large V3 Turbo | 1.5 GB | 8.9 | 64.8× | 9.5 min | 168 MB |
Total time is wall-clock to transcribe the entire 10-hour corpus. RTFx (real-time
factor) is the hardware-portable way to estimate your own runtime:
time ≈ audio length ÷ RTFx. At 64.8× a one-hour lecture takes ~55 s; at 14.7× it
takes ~4 min. A faster GPU pushes RTFx higher; CPU pushes it far lower (CPU large-v3
can run 5–10× slower than realtime). So you can either run the tool on your own
audio, or just read these numbers off the run above.
Per-lecture WER%:
| Lecture | Audio | Small | Medium | Large V3 | Large V3 Turbo |
|---|---|---|---|---|---|
| Lecture 1 | 58.4 min | 12.2 | 14.4 | 13.8 | 8.3 |
| Lecture 2 | 54.5 min | 10.2 | 9.6 | 29.7 | 8.2 |
| Lecture 3 | 54.5 min | 13.8 | 14.4 | 15.9 | 12.7 |
| Lecture 4 | 68.9 min | 9.2 | 13.1 | 15.5 | 8.3 |
| Lecture 5 | 11.4 min | 13.9 | 19.3 | 18.0 | 13.4 |
| Lecture 6 | 56.2 min | 9.3 | 9.3 | 10.2 | 8.7 |
| Lecture 7 | 64.4 min | 9.1 | 12.9 | 10.9 | 9.2 |
| Lecture 8 | 58.3 min | 12.4 | 10.4 | 9.6 | 8.7 |
| Lecture 9 | 56.1 min | 11.6 | 11.7 | 14.7 | 9.8 |
| Lecture 10 | 66.2 min | 10.8 | 6.5 | 11.6 | 5.6 |
| Lecture 11 | 6.8 min | 7.6 | 10.6 | 7.5 | 6.8 |
| Lecture 12 | 58.5 min | 8.8 | 9.8 | 13.3 | 7.2 |
| Overall | 614.4 min | 10.7 | 11.8 | 14.2 | 8.9 |
On this content, Large V3 Turbo had both the lowest overall WER (8.9%) and the highest throughput (~65× realtime) — unusual, since Large V3 normally has the accuracy edge. Lecture 2's Large V3 spike (29.7%) was a decoder lockup on one clip; it's the failure mode the default VAD filter now prevents. This is exactly the kind of per-corpus surprise the tool exists to surface — run it on your audio rather than trusting a generic leaderboard.
WhisperX adds forced wav2vec2 word-level alignment and pyannote speaker
diarization on top of any Whisper size. Use <size>+whisperx model IDs:
python asr_bench.py --models large-v3-turbo+whisperx --diarizeWhat it adds compared to plain Whisper:
- Word-level timestamps —
<base>_Words_<Model>.jsonsidecar next to each audio file. - Speaker-labeled VTT — cues prefixed
SPEAKER_00: text,SPEAKER_01: text, etc. - DER column — Diarization Error Rate appears in the headline table whenever a
<base>.rttmground-truth sidecar is present next to the audio. Without an RTTM file the DER column is omitted and diarization still runs (speaker labels in VTT). - Speakers column — detected speaker count alongside DER.
WhisperX requires a Python ≤ 3.13 environment (PyTorch has no 3.14 wheels). The convenience script sets up the venv asr-bench auto-detects:
./setup_whisperx_venv.ps1 # default CUDA build (cu128, RTX 50xx/Blackwell)
./setup_whisperx_venv.ps1 -CudaIndex cu124 # older CUDA toolkit
./setup_whisperx_venv.ps1 -CudaIndex cpu # CPU-onlyOr by hand:
py -3.12 -m venv .venv-whisperx
# Install the CUDA torch build FIRST — `pip install whisperx` otherwise pulls the
# CPU-only torch wheel on Windows (torch.cuda.is_available() == False, silent CPU).
.venv-whisperx\Scripts\pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu128
.venv-whisperx\Scripts\pip install whisperx
.venv-whisperx\Scripts\python -c "import torch; print(torch.cuda.is_available())" # expect Trueasr-bench auto-detects ./.venv-whisperx and uses it as a subprocess for
WhisperX runs. To use a different path:
python asr_bench.py --models large-v3-turbo+whisperx --whisperx-python path\to\python.exeIf torch is already importable in the running interpreter (e.g. you are running
asr-bench from a 3.12 venv that already has WhisperX installed), asr-bench runs
WhisperX in-process automatically — no subprocess overhead.
Speaker diarization uses the gated pyannote/speaker-diarization-community-1
model from HuggingFace (pyannote-audio 4.x unified on this single self-contained
repo — it bundles segmentation + embedding, so you accept just one). To use it:
- Create a free account at huggingface.co.
- Accept the model terms at huggingface.co/pyannote/speaker-diarization-community-1.
- Generate a read token and pass it via
--hf-tokenor theHF_TOKEN/HUGGINGFACE_TOKENenvironment variable.
On a pyannote 3.x install, pass
--diarize-model pyannote/speaker-diarization-3.1(and accept that repo +pyannote/segmentation-3.0instead). The runner defaults to community-1 for pyannote 4.x.
python asr_bench.py --models large-v3-turbo+whisperx --diarize --hf-token hf_...Missing token: asr-bench warns and falls back to alignment-only (word timestamps
- VTT without speaker labels) — it does not hard-fail. To skip diarization entirely:
python asr_bench.py --models large-v3-turbo+whisperx --no-diarizeDrop a <base>.rttm ground-truth annotation file next to the audio file. asr-bench
picks it up automatically and adds the DER% and Speakers columns to the report.
Without an RTTM file diarization still runs but DER is not computed.
Tip — long recordings: pyannote tends to over-cluster on long, noisy audio
(e.g. a 2-person call estimated as 12 speakers). If you know the speaker count,
pass --min-speakers/--max-speakers — on an 82-min 2-speaker validation clip
this took the detected count 12 → 2 and DER 27.4% → 13.8%.
python asr_bench.py --models large-v3-turbo+whisperx --diarize \
--min-speakers 2 --max-speakers 4--min-speakers / --max-speakers are hints to pyannote, not hard limits.
Every run also writes a machine-readable results/<timestamp>.json (same
timestamp as the markdown report; sibling <output>.json when you pass
--output). It mirrors the full run — run config, per-model and per-clip
metrics, transcripts, and speaker/DER data — for cross-run aggregation. NaN
values (e.g. DER on a non-diarized clip) serialize as null. Secrets
(hf_token, nim_api_key) are never written. Opt out with --no-json.
Every run checks each clip for looping or fabricated output using two
reference-free heuristics and adds a "
repeat_coverage— fraction of words that appear inside a repeated 4-gram. Normal prose is near 0; looped Whisper output is typically > 0.30.compression_ratio— gzip ratio of the hypothesis text. This is the same signal Whisper uses internally to detect hallucination (normal prose ≈ 1.5–2.2; repetitive/looped output compresses more, typically > 2.4).
A clip is flagged when repeat_coverage > 0.30 OR compression_ratio > 2.4.
The report lists every flagged (model, clip) pair with both signal values and an
insertion-rate annotation ("insertion burst") when a reference is available.
Per-model hallucination_rate (fraction of clips flagged) and the per-clip
signals (repeat_coverage, compression_ratio, hallucination_suspect) are
written to the JSON sidecar. Works without a reference and in single-model runs,
unlike cue-density anomaly detection (which is cross-model and structural).
Every run writes a results/<timestamp>.json sidecar. compare reads 2+ of them
and prints a markdown comparison:
# 2 files -> delta view (baseline -> candidate, signed Δ with ✓/✗)
python asr_bench.py compare results/20260605-190913.json results/20260606-101500.json
# 3+ files -> matrix (one column per run)
python asr_bench.py compare results/a.json results/b.json results/c.json
# the two most recent runs, with per-clip detail
python asr_bench.py compare --last 2 --per-clip
# write to a file instead of stdout
python asr_bench.py compare results/a.json results/b.json --output report/compare.mdForce a layout with --delta (exactly 2 files) or --matrix. WER/MER/WIL/CER/DER
are lower-is-better (improvement marked ✓); RTFx is higher-is-better. When any
compared run has hallucination data, a Halluc% column appears (lower is better).
Differing corpus or key config between runs is flagged as a
asr-bench/
├── README.md ← this file
├── SPEC.md ← full roadmap including Canary-Qwen + NeMo
├── requirements.txt
├── asr_bench.py ← the script
├── whisperx_runner.py ← standalone subprocess for WhisperX runs
├── test-corpus/ ← bring-your-own (gitignored except README)
└── report/ ← timestamped markdown outputs (gitignored)
Your reference transcript IS the ground truth. The WER score is only as good as the reference.
Best: hand-correct 5-10 short clips (10-30 sec each) covering the kinds of audio you actually deal with — single speaker, multi-speaker, technical vocab, accents, your idiolect on numbers and dates. One evening of labor for defensible numbers.
Acceptable: use existing auto-generated captions (Panopto, YouTube auto-caps, Zoom transcripts) as reference. Treat the resulting WER as a relative divergence rate — engine vs your existing pipeline. Useful for "does engine X disagree with Panopto more than engine Y?" but not "what is engine X's absolute accuracy?"
Worst: use one ASR model's output as the reference for benchmarking another ASR model. This produces numbers that look real but measure nothing.
The --fuse flag adds a post-benchmark pass that combines each model's VTT output
with the Panopto reference into a consensus transcript. It works by windowing the
timed cues into overlapping chunks (default 25 s windows, 5 s overlap) and asking
a pluggable LLM to fuse each window into one of two profiles:
| Profile | Output | Use for |
|---|---|---|
verbatim |
<base>_Captions_Fused.vtt |
Accessibility captions; optional rescoring reference |
kb |
<base>_KB_Fused.jsonl + .md |
RAG / knowledge-base ingestion |
--profile both (the default) produces both from a single chunked pass.
Two important caveats:
- Only the verbatim profile is ADA/WCAG caption-eligible. The kb profile deliberately rephrases and condenses for retrieval quality — it is explicitly NOT compliant captions and is labeled as such in the report.
--rescore-against-fusedis agreement-biased. The fused verbatim VTT is built from the same model outputs it is then used to score — it measures consensus, not ground truth. The report labels this table clearly.
Pass --llm <backend> to choose how the fusion prompt is served:
ollama:<model>(default:ollama:qwen2.5) — calls a locally-running Ollama server. Fully offline, no API key. Requiresollama serveand the model pulled (ollama pull qwen2.5).cli:<command>— shells out to an authenticated frontier CLI. Uses your existing subscription — no asr-bench API key required. The prompt is piped on stdin by default (e.g.cli:claude -p); if the command contains a{prompt}token it is substituted as an argument instead, which some CLIs require (e.g."cli:gemini -p {prompt}"). The executable is resolved viaPATH(Windows.cmd/.batshims included). Caveat: agentic CLIs likegeminireload their whole harness per invocation (often 10s–minutes per call), so they are impractical for bulk per-window fusion — prefer Ollama for full runs and reservecli:for one-off/small fusions or a fast headless completion CLI.fake— returns deterministic stub output. No LLM required. Good for testing pipeline wiring before you have Ollama set up.
Provide domain context to improve fusion quality:
# Generate a guided template
python asr_bench.py --init-context context.md
# Fill in context.md (course schedule, speaker names, jargon, glossary, …)
# Then pass it to a fusion run:
python asr_bench.py --models large-v3-turbo --fuse --context context.md# Local Ollama — both profiles, with context
python asr_bench.py --models small,medium,large-v3-turbo \
--fuse --profile both --llm ollama:qwen2.5 --context context.md
# Frontier CLI backend — verbatim captions only (prompt on stdin)
python asr_bench.py --models large-v3-turbo \
--fuse --profile verbatim --llm "cli:claude -p" --context context.md
# Frontier CLI that needs the prompt as an argument (e.g. gemini)
python asr_bench.py --models large-v3-turbo \
--fuse --profile verbatim --llm "cli:gemini -p {prompt}" --context context.md
# Re-score all models against the fused verbatim reference
python asr_bench.py --models small,medium,large-v3-turbo \
--fuse --rescore-against-fused --context context.md
# Dry-run with no LLM (FakeLLMBackend — no Ollama required)
python asr_bench.py --models small --fuse --llm fake --limit 1Note: Fusion is fully unit-tested via
FakeLLMBackendbut has not yet been validated end-to-end against a live Ollama orcli:backend on real lecture audio. Expect to tune the drift-guard threshold and review output quality on your first real run.
Status — experimental; not yet validated against a live NIM. asr-bench's core purpose is benchmarking local Whisper variants. The NIM engine (and other extra models) are a nice-to-have, not critical. The engine is implemented and its
riva.clientAPI usage is verified statically againstnvidia-riva-client2.26.0, and the supporting paths — audio decode, engine dispatch, report rendering, and graceful failure — are tested. But it has not been run end-to-end against a live NIM: the intended deployment is a self-hosted NIM container running locally, which is untested here because NIM ships only as a container (nvcr.ioimage, no native binary) and no container runtime was set up on the reference machine. The remote/hosted path (--nim-api-key/--nim-ssl) is implemented but likewise not fully tested. Expect to debug the first real run.
To benchmark a NIM ASR endpoint (e.g. a self-hosted Canary NIM):
pip install nvidia-riva-clientThis is only needed when you request a nim-engine model. The engine reaches a
server two ways. Local self-hosted is the preferred, default path; the remote
hosted endpoint is a flag-gated fallback for users without a local container
runtime (and is never the default — consistent with the "local engines only"
ethos).
# Local (preferred): self-hosted container over gRPC
python asr_bench.py --models large-v3-turbo,canary-nim --nim-url localhost:50051
# Remote (fallback): NVIDIA hosted NVCF endpoint — needs an NGC API key
python asr_bench.py --models large-v3-turbo,canary-nim \
--nim-url <host>:443 --nim-api-key <NGC key> --nim-sslNIM rows report WER, RTFx, and wall-clock like any engine. Because the model
runs behind a gRPC service, VRAM is reported as total GPU memory in use
(marked *) rather than a per-clip delta, and disk size shows n/a. Use
--models nim:<riva-model-name> to benchmark an unregistered NIM model.
MIT — see LICENSE.
Issues + PRs welcome. The model registry in asr_bench.py is the natural place
to add new engines. Each engine needs a wrapper that exposes
transcribe(audio_path) -> str plus static metadata.
- faster-whisper — CTranslate2 port of OpenAI Whisper
- WhisperX — word-level alignment + speaker diarization
- pyannote.audio — speaker diarization model
- jiwer — WER / MER / WIL computation
- nvidia-ml-py — GPU memory tracking
- Ollama — optional local LLM backend for fusion
- Morris, Maier & Green (2004) — MER and WIL metric definitions