Tests: the volume oracle — a deterministic 'ambient night' fixture gates triage usability at realistic scale - #104
Merged
Merged
Conversation
…tes triage usability at realistic scale (#50) Adds tests/volume/ as a second regression oracle (ADR-0068), sibling to tests/golden/: golden pins exact scores for 1-12 event scenarios; this oracle pins usability invariants (queue set-equality, the flood tripwire, breach-among-noise anti-suppression, conservation, calm-state reachability, generation determinism) under a realistic ~130-actor/~850- event ambient night, built by a seeded generator driving the REAL Suricata/syslog/CEF normalizers — never a hand-built SecurityEvent. The ADR-0070 (+ Amendment 1) distribution-table personas — the 50/min attacker, the single burst that queues during and fades after, the nightly recidivist, the moderate grinder (endurance at ~24h), the sub-theta_press paced actor (INFORM forever, the designed exclusion), the 10x4min hysteresis pin, and ambient priority-2 Suricata noise — are each a named, individually-failing test through the real normalizer, mirroring (not duplicating) the unit-level pins in test_issue_54_attack_in_progress_campaign.py. A frontend vitest sibling (triageBand.volume.test.ts) feeds the committed derived_threats.json fixture through deriveTriageActors and independently closes the JS `tier: null <= 2` coercion channel. tests/golden/fixtures/expected_scores.json is untouched (sha fe4787643955c920e934e3789c79f741cd8c8cde6b2adbc6540b66ff3743f31f, verified unchanged). No firewatch-core edits — this is pure test/fixture infrastructure exercising the already-shipped ADR-0069/0070 intensity model and severity recalibration. Closes #50
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
tests/volume/— a second regression oracle (ADR-0068), sibling totests/golden/: golden pins exact scores for 1-12 event scenarios; this oracle pins usability invariants (queue set-equality, the flood tripwire, breach-among-noise anti-suppression, conservation, calm-state reachability, generation determinism) under a realistic ~130-actor/~850-event ambient night.generator.py) expanding recorded templates (tests/golden/fixtures/eve_*.json, thepackages/sources/syslog/tests"Failed password" line shape) and driven through the REALfirewatch_suricata/firewatch_syslog/firewatch_syslog_cefnormalizers (harness.py) — no hand-builtSecurityEventanywhere in the scenario path.frontend/src/test/triageBand.volume.test.ts) feeds the committedderived_threats.jsonfixture throughderiveTriageActorsand independently closes the JStier: null <= 2coercion channel.scripts/regen_volume_fixtures.pyis the single regeneration entrypoint; a drift-check test fails loudly if the committed fixture goes stale.firewatch-coreedits — this is pure test/fixture infrastructure exercising the already-shipped ADR-0069/0070 severity recalibration + intensity model.Structure self-check
generator.py(359 lines) — one concern: manifest × templates →RawEvents, no scoring imports.harness.py(108 lines) — one concern:RawEvents → real normalizers → per-actorThreatScore, mirroring (not reimplementing)Pipeline.analyze_ip's decision slice.test_triage_volume.py(483 lines) — the invariants + the persona ledger; at the ~500-line guideline, kept as one file because the invariant fixtures (ambient_scores/breach_scores) are shared session-scoped state across every test class — splitting would either duplicate that fixture setup or force an import-heavy conftest split for no isolation benefit.Test plan
bash scripts/gates-backend.shgreen (sync + ruff + pyright + pytest) — see gate report below.tests/golden/fixtures/expected_scores.jsonsha unchanged:fe4787643955c920e934e3789c79f741cd8c8cde6b2adbc6540b66ff3743f31f.tests/volume/runs in the defaultuv run pytestinvocation (no opt-in marker), full scenario scores in well under 1s (budget: ≤5s).npm run lint && npm run typecheck && npm run test && npm run buildall green (178 test files / 4332 tests).tests/volume/README.md's ledger section).Gate report
Frontend (
frontend/):npm run lintclean,npm run typecheckclean,npm run test→ 178 files / 4332 tests passed,npm run buildsucceeds.Closes #50