Use the public recognize() pipeline to measure what the current version gets right, where it
returns no suggestion, and where it suggests the wrong amount. Start with a small development set,
then add a synthetic regression fixture whenever you fix a failure.
See the 9 September 2026 development baseline for the initial observations and their limits.
From the repository root:
python3 -m venv .venv
.venv/bin/python -m pip install --editable '.[dev]'
./scripts/setup-models.sh
.venv/bin/python -m pytest -m 'not integration'
.venv/bin/python -m pytest -m integrationTesseract must be installed separately (brew install tesseract on macOS). Both pinned English and
Devanagari models are required. macOS uses the default development runtime policy. The reference
conformance runtime is Ubuntu 24.04 / amd64 / Tesseract 5.3.4.
The committed synthetic manifest is ready to run without supplying private receipts:
.venv/bin/python tools/benchmark.py \
--manifest tests/fixtures/capability-cases.json \
--repetitions 1 \
--output output/synthetic-capabilities.jsonThese eight cases cover payable totals and images that should return no suggestion. Inspect the saved report to see each expected answer beside the observed answer, evidence grade, and timing.
.venv/bin/python -m payable_receipt_ocr private-inputs/receipt.png --currency INRInspect result.total, result.evidence_grade, and warnings. An evidence grade of strong
still requires the person to confirm the amount. Use --diagnostics only when you need the raw OCR
text and candidate list to investigate a failure; keep that output with the private images.
Create a local private-inputs/development-cases.json. Label each image before running OCR so the
expected answer is independent of the model's suggestion. For example:
{
"cases": [
{
"id": "receipt-001",
"image": "receipt.png",
"expected_total": "289.86",
"expected_currency": "INR",
"category": "checkout"
},
{
"id": "receipt-002",
"image": "no-total.png",
"expected_total": null,
"expected_currency": null,
"category": "no-total"
}
]
}Image paths resolve relative to the manifest file; absolute paths are also accepted. Every id
must be unique. Amount labels must be nonnegative canonical two-decimal strings, such as 0.00
or 1234.50, with currency INR. Use both values as null when the image should produce no
suggestion. A category is optional. All records and image paths are validated before OCR starts.
Keep real images, manifests, and reports private and out of commits and CI artifacts. The benchmark excludes raw OCR text, but its report contains image paths, amounts, and error messages. The repository's contribution rules permit only synthetic receipt fixtures in committed tests.
.venv/bin/python tools/benchmark.py \
--manifest private-inputs/development-cases.json \
--repetitions 1 \
--output output/development-before.jsonThe report includes expected and observed totals/currencies, exact matches, wrong results marked
strong, warning codes, no-suggestion counts, errors, and accuracy by category. It also records
confirmation-contract violations, degraded processing, deadline exhaustion, and p50/p95/p99 wall
times. Latency statistics include failed calls. wall_duration_ms covers the complete OCR loop.
Exit status is 0 when every labelled run matches and all returned results preserve confirmation
requirements; 1 means an OCR error, labelled mismatch, or confirmation-contract violation; 2
means invalid inputs or a report-write error. A mismatch still produces the complete report.
For the existing unlabelled latency workflow:
.venv/bin/python tools/benchmark.py \
tests/fixtures/clean-screenshot.png tests/fixtures/multiple-totals.png \
--repetitions 3 --output output/synthetic-timing.jsonUnlabelled runs have no accuracy or false-strong rate; ok and success_count only mean the OCR
call returned a result. A returned result can still contain no amount. Confirmation rate is
calculated across returned results; OCR errors are reported separately and always fail the run.
Include clean screenshots, competing totals, discounts/wallet use, small text, wide price columns, dark backgrounds, blur/rotation, and images without a payable total. Include each app you intend to support. Examine a failing image alongside its expected amount and opt-in diagnostics to decide whether text recognition, row reconstruction, amount selection, or evidence grading caused it.
After a change, rerun the same labelled images with the same runtime and repetition count:
.venv/bin/python tools/benchmark.py \
--manifest private-inputs/development-cases.json \
--repetitions 1 \
--output output/development-after.jsonCompare summary.exact_match_rate, summary.false_strong_count, summary.error_count, category
results, and durations.p95_ms. Match individual runs by case_id and repetition to check both
fixed cases and regressions. Repeat runs when investigating timing or inconsistent outcomes; the
reported accuracy then counts every repetition, while case_count remains the unique case count.
Keep the input images and labels unchanged during a comparison.
Add synthetic reproductions under tests/fixtures/ and public-interface assertions in
tests/test_public_interface.py, then run the portable and integration suites again. Changes to
preprocessing, OCR pass scheduling, evidence grading, or the runtime baseline need holdout
requalification before release; see CONTRIBUTING.md.
A development benchmark helps discover failures and compare changes on your chosen images. It does
not measure general receipt accuracy or pass the v1 release gate, even under a conformant runtime.
The report always sets eligible_for_release_gate to false; runtime conformance reporting only
indicates the runtime policy and baseline checks succeeded. Repeated or tuned-on cases are not an
independent holdout.
The release evaluator remains tools/holdout/evaluate.py, with a frozen authorized corpus of at least 300 images and 100 per supported app. See the holdout process and release gate. Current scope is INR payable totals from Blinkit, Swiggy, and Zepto images; PDFs, HEIC, handwriting, line-item extraction, and merchant/date extraction are outside the supported scope.