Skip to content

Latest commit

 

History

History
2341 lines (2063 loc) · 231 KB

File metadata and controls

2341 lines (2063 loc) · 231 KB

Testing (gbrain repo)

test/local-install-lifecycle.serial.test.ts exercises isolated Bun-link install, keyless memory write/read/search, process reopen and migration/post-upgrade with service-command tripwires. test/e2e/grandfather-projection-postgres.test.ts checks guarded metadata-only grandfathering, duplicate source slugs, preserved valid text projections and refusal to seal previously unsealed rows on Postgres. test/reindex-markdown-persistence.slow.test.ts retains the bounded 3,600-page real-CLI reindex, SIGKILL and resume workload; the diagnostic benchmark launcher is scripts/bench-reindex-markdown.ts.

On-demand reference (see CLAUDE.md Reference map). Current behavior + invariants only.

CI runner capacity

Repository-owned Linux validation jobs use ephemeral Ubicloud runners pinned to Ubuntu 24.04. The Ubicloud Managed Runners GitHub App must have access to this repository and active billing in its connected project; runner labels alone do not grant access. No Ubicloud API token is passed to workflow jobs.

Workload Runner Capacity
Unit shards, slow and eval jobs, BrainBench, admin browser, shared-skills compatibility, persistence soak, reconciliation crashes and read latency, native Linux cells, OpenClaw startup, JSONB parity, selected E2E and Tier 2 ubicloud-standard-4-ubuntu-2404 4 vCPU, 16 GB RAM
Serial pool (PR and nightly coverage), verify, PgBouncer/RLS deployment matrix ubicloud-standard-8-ubuntu-2404 8 vCPU, 32 GB RAM
E2E Tier 1 (its CLI init spawns exceed their timeouts on 4 vCPUs), label-gated and nightly heavy-tests jobs ubicloud-standard-16-ubuntu-2404 16 vCPU, 64 GB RAM
Label-gated heavy test suite ubicloud-standard-30-ubuntu-2404 30 vCPU, 120 GB RAM
Native ARM64 glibc and musl tests ubicloud-standard-4-arm-ubuntu-2404 4 vCPU
Coverage reports and Semgrep ubicloud-standard-4-ubuntu-2404 4 vCPU, 16 GB RAM
Planning, status aggregation, dependency audit, gitleaks, security regressions and actionlint ubicloud-standard-2-ubuntu-2404 2 vCPU, 8 GB RAM

macOS and Windows matrices stay on GitHub-hosted runners. Release building and publishing also stay unchanged. The pinned upstream OSV reusable workflow does not expose a runner override, so its runner remains upstream-owned.

Sizes come from measured CPU use, not guesses. Each Ubicloud project shares one vCPU quota between pull-request CI and agent ci:ubicloud VMs, so an oversized runner makes every other job wait. On matched VMs (September 2026), a unit shard is one Bun process that averaged 1.1-1.5 busy cores and took 361s on 4, 8 and 16 vCPUs (488s on 2); the serial pool and verify took the same time on 8 and 16 vCPUs (serial shard 2: 222s and 224s; 277s on 4); the 2,500-write PGLite soak averaged 1.3-1.8 busy cores and took 523s on 4 vCPUs and 500s on 16. Re-measure with scripts/ubicloud/ubi-runner.sh run -s standard-N before growing a runner: more CPU does not shorten a single-process job. test/scripts/ci-runner-routing.test.ts pins capacity and platform routing; .github/actionlint.yaml declares the exact custom runner labels. Actual GitHub job records and completed checks establish runner availability; local workflow tests do not.

Pull request, master and nightly scope

Every test file runs on every push to master, on the nightly schedule and on manual dispatch. Pull requests run a narrower matrix of the same files:

Lane Pull request Push to master, nightly, manual
Security regressions Linux, macOS and Windows on Bun 1.4.2 Also Bun 1.4.0
Persistence read latency, deployment matrix, soak, reconciliation crashes Bun 1.4.2 Bun 1.4.0 and 1.4.2
Persistence soak size 2,500 writes 10,000 writes
Native writer locks, native paths changed Every target on Bun 1.4.2, musl, both Windows probes, OpenClaw Every target, musl and Windows probe on Bun 1.4.0 and 1.4.2, OpenClaw
Native writer locks, other changes linux-x64-glibc / Bun 1.4.2 smoke cell (full native step list) Same as above
test/export-scale.slow.test.ts 10,001 pages 100,001 pages

The changes job classifies a pull request's changed files with scripts/ci-native-scope.sh: native lock sources, the native toolchain, IPC, persistence, publication, backup, export and sync sources, their native tests, package.json, bun.lock and the workflow files select every target. An unreadable file list selects every target too. Skipped cells never report a failure: test-status needs the native-locks and persistence-validation workflow calls, which succeed when their remaining cells do, so the required check names are unchanged. test/scripts/ci-pr-scope.test.ts pins every scope.

Shared-skill tests distinguish canonical publication, protocol delivery, installed files and native harness use. test/shared-skills-transports.test.ts and test/e2e/shared-skills-transports.test.ts use real HTTP authentication, OAuth issuance and a new stdio process; they do not prove vendor-native activation. test/persistence-skill-bundles.serial.test.ts and test/persistence-skill-crash.slow.test.ts exercise typed file-set CAS and independent-process publication/restoration kills on PGLite in the unit lane and PostgreSQL through test/e2e/persistence-skill-bundles-postgres.test.ts.

Shared persistence suites use test/helpers/test-backends.ts: direct invocation defaults to PGLite, and a safe DATABASE_URL opts into both engines. Their E2E wrappers select PostgreSQL before registering tests, refusing a missing or unsafe database instead of silently running only the local backend. Backend selection is captured at registration so hooks retain it after the import environment restores. Every backend's assertions remain in the shared suites; engine-specific cases run in their owning lane.

Ordinary PostgreSQL setupDB() clears fixture data, operator configuration and source sync identity while retaining config.version and the stored embedding identity, avoiding historical migration replay against an already-current schema. Migration-focused fixtures use setupDB({ replayMigrations: true }); an absent ledger also runs the cold chain. test/e2e/fixture-reset-postgres.test.ts checks both paths, cleanup and vector-shape preservation, including the deliberate legacy-width restoration helper. That helper aligns both the physical columns and stored embedding identity with the legacy test configuration; ordinary resets preserve that identity.

The required shared-skills-compatibility CI job builds the pinned pre-feature executable with scripts/build-shared-skills-baseline.sh and supplies GBRAIN_TEST_OLD_BINARY to test/persistence-skill-old-binary.slow.test.ts. An absent old executable is an explicit skip, never old-writer evidence. test/shared-skills-catalog-performance.test.ts runs the reproducible 10/100/1,000-skill read benchmark when GBRAIN_TEST_SHARED_SKILLS_BENCHMARK=1; its timings and database-call counts are recorded diagnostics, while identity and catalog-size assertions are deterministic. The shared-skills cases under evals/harness-instructions/ test interpretation separately from executed calls and native sessions.

scripts/shared-skills/lifecycle.ts measures real authenticated HTTP enrollment, revision/asset reads, publication, missed-notification polling, acknowledgments, queue/recovery bytes, and concurrent read latency at 10/100/1,000 skills. Run it on a quiet machine with the protocol and fixture boundaries in its README. Its additional body/asset latency comparison is experimental; report each engine's measured result without substituting it for the existing scripts/persistence/performance.ts memory-read gate. The five pure accounting tests run normally; the small runtime smoke is explicitly opt-in and never counts as a full performance pass.

Canonical reconciliation

test/persistence-reconcile-merge.test.ts pins loss-preserving field choices. test/persistence-reconcile.test.ts runs the guarded repair and replay contracts on PGLite and, with an explicit safe DATABASE_URL, isolated PostgreSQL databases. It covers stale preconditions, current/original grants, private facts, retained backups, ordinary mutations after repair, and competing publications. test/reconcile-owner-journey.serial.test.ts drives real CLI requests through HTTP and stdio PGLite owners before and after activation, restarts the owner, and independently reads the newly remembered private fact and provenance.

test/reconcile-crash.slow.test.ts and test/e2e/reconcile-crash*.test.ts kill real processes at all eight publication boundaries with activation off/on. PostgreSQL uses one file per activation state to stay within the unchanged per-file cap. Optional GBRAIN_TEST_RECONCILE_CRASH_MANIFEST_DIR retains executed-case evidence. test/e2e/reconcile-pgbouncer.test.ts requires the transaction-mode pooler when GBRAIN_CI_REQUIRE_PGBOUNCER=1 and proves repair followed by a new private memory write. The durable-persistence workflow runs these contracts on both supported Bun versions and uploads the crash manifests; local CI runs the slow and E2E lanes.

test/docs-navigation.test.ts checks local links and fragments in the primary install/memory guides and all docs/architecture/key-files/ references, requires every subsystem to be linked from KEY_FILES.md, and guards against blanket graph-write and preference-routing claims. The fixture suite test/scripts/check-key-files-current-state.test.ts covers history markers, cross-subsystem duplicate entries, and byte caps for the entry docs and references.

Search reliability has real-planner and transport regressions in test/e2e/vector-candidate-safety-postgres.test.ts, test/e2e/search-query-contract-postgres.test.ts, test/e2e/projection-statistics-postgres.test.ts, and test/search-readiness-http.test.ts. The statistics tests include owner, restricted-reader and FORCE-RLS roles; the candidate tests distinguish natural plans from forced-HNSW controls and prove server cancellation of exact fallback. test/e2e/vector-plan-real-column-postgres.test.ts is the #5824 plan proof on the real embedding column: in a dedicated 64-dim database (it needs CREATEDB and CREATE EXTENSION vector, so it runs under bun run ci:local) it first shows the legacy-guard statement seq-scans on its fixture, then that the emitted statement uses idx_chunks_embedding across a filter matrix incl. RLS scope binding, then the stale-heavy escalation, exact-fallback and short-window cases and every doctor vector_plan outcome. The PGLite side is test/search/vector-freshness.test.ts; the SQL shape and lockstep are test/search/vector-statement.test.ts. The opt-in reporter-scale bench is scripts/bench/vector-plan-5824.ts. test/e2e/projection-recovery-parity.test.ts runs the shared Markdown/code recovery, graph-edge preservation and migration-origin contracts against PostgreSQL; their root suites cover PGLite in the unit lane. PGLite work caps never count a Promise race as cancellation evidence. The recovery parity entry also runs symbol-resolver-projection-race.test.ts: paused resolver/rebuild ordering, atomic rollback, candidate revalidation, and a real PostgreSQL lock-wait receipt before releasing the competing writer.

test/pglite-in-memory-create-retry.serial.test.ts injects create failures while using real PGLite instances and a validated schema snapshot. It pins one cold retry only before an in-memory database has opened, both failure diagnostics, schema replay after snapshot fallback, post-open cleanup and close poisoning, concurrent connect/disconnect ordering, exit-code preservation, and exclusion of the persistent repair path. Run it in its own Bun process because it mocks the PGLite module. Its recovery cases discriminate against the no-retry base; its post-open cleanup cases discriminate against a retry that replaces a live database.

test/e2e/serve-http-oauth.test.ts additionally pins confidential POST/Basic revocation, public-client SDK fallthrough, malformed/mixed authentication rejection, cross-client isolation, unknown-token opacity, metadata auth methods, no-store responses, strict post-revoke 401, and retryable backend 503 semantics. SDK-driven discovery and real owner-approved PKCE also pin read-only bootstrap, explicit writer requests, scope clamping, and DCR delegation refusal. test/oauth-scope-hint.test.ts exercises the actual SDK middleware over HTTP without requiring a database.

test/put-page-persistence.test.ts and test/e2e/put-page-persistence-postgres.test.ts pin durable page acceptance and ordinary-error publication: native contention returns an accepted pending receipt without changing the page, and replay of its original UUID commits exactly once after release. Filesystem or required source-path failure rolls back the database transaction. Embedding failure preserves the canonical receipt; a delayed result superseded by another revision cannot install vectors. The PGLite suite also covers scoped physical file paths, unchanged-content no-ops, legacy hashes, deletion/recreation, and sanitized diagnostics. Actual process-death boundaries belong to the crash suites below.

test/subagent-required-writes.test.ts and test/subagent-put-page-rejection.serial.test.ts distinguish a persisted write from prose-only completion, rejected imports, and historical rejected ledger envelopes across the Anthropic, gateway, and oneshot lanes. Unchanged saves, optional-write jobs, and saved pages with failed enrichment are positive controls. test/cycle/global-freshness-postcondition.serial.test.ts exercises the registered maintenance handler with failed phases, incomplete children, budget deferrals, abort/lock loss, and successful warning-only controls.

Coverage responsibilities before consolidation

Assign ownership to an assertion and its execution boundary, not to a test filename or a shared helper. Record the contract, backend, runtime version, OS/architecture/libc, source-versus-compiled artifact, transport/authentication, process/storage/crash boundary, workload size and required cadence. Shared scenario code across two engines is not duplicate engine coverage: PostgreSQL JSONB, locking and pooler behavior are not established by a PGLite pass.

Responsibility Execution owner What it does not establish
Keyless behavior, structural guards and shared contracts Unit shards and verify in test.yml; process-isolated serial and dedicated slow lanes where required Real PostgreSQL, native activation or compiled behavior
PostgreSQL behavior and engine parity Named and diff-selected jobs in e2e.yml; the complete nightly runner corpus Execution of key-gated or native-door assertions merely because their files were discovered
Durable publication and recovery under sustained load persistence-validation.yml and scripts/persistence/README.md Power-loss safety, production authentication or equivalence to two smaller databases
Native lock ABI and compiled-process exclusion native-locks.yml, compiled smoke and release validation All compiled CLI features or native-harness activation
Browser journeys Required admin-browser job and admin/e2e/*.pw.ts Vendor-native agent behavior
Native agent doors and heavier operational scenarios Explicit jobs in heavy-tests.yml A passing skipped door or generic protocol test is not native activation
Live-provider and optional recipe/eval behavior Their explicitly configured opt-in commands/jobs A missing key, early return or skipped assertion is not live-provider evidence
Line-coverage accounting PR prCorpus and nightly fullCorpus reports Subprocess coverage, all platforms or proof that every discovered case executed

Before removing repeated work, identify the surviving owner for the same contract and every relevant boundary, prove that owner actually executes, and retain its cadence, failure gate and coverage artifacts. A shared fixture can reduce maintenance while keeping both engine arms. Making one crash lane authoritative or collecting LCOV in a named owner requires a separate ownership change; nightly sharding alone makes neither change.

The 2026-09-29 test audit's lane reports, inventories and mutation-probe logs are committed under docs/test-audit/2026-09-29/; cite them for the surviving-owner and probe evidence behind a consolidation.

Recorded ownership changes:

  • test/e2e/reconcile-crash.test.ts and test/e2e/reconcile-crash-unactivated.test.ts: the PR owner is persistence-validation.yml, called from test.yml on every PR on both supported Bun versions against pg16. Its "Require all reconciliation crash boundaries" step runs both files by name and uploads the crash manifests, unchanged. Both files are in E2E_EXCLUSIONS (PERSISTENCE_VALIDATION_OWNED in scripts/e2e-matrix.ts), so PR selected-e2e does not run them a second time; scripts/select-e2e.ts prints excluded: <file> (owned by persistence-validation.yml) on stderr when a mapped source changes. The nightly full-corpus E2E run and the local gates (ci:local, ci:ubicloud, their :diff forms) still run them. Run them locally with the same command the workflow uses, with DATABASE_URL exported for the test database from "E2E test DB lifecycle":

    GBRAIN_TEST_ALLOW_DATABASE_URL=1 \
    GBRAIN_TEST_RECONCILE_CRASH_MANIFEST_DIR=.context/reconcile-crashes \
      bun --no-env-file test --timeout=180000 \
      test/e2e/reconcile-crash.test.ts test/e2e/reconcile-crash-unactivated.test.ts
  • Attendance parity (test/attendance-retrieval.test.ts, test/attendance-repair.test.ts, test/extract-timeline-attendance.test.ts): the unit lane owns the PGLite arm; the test/e2e/*-postgres.test.ts wrappers load the scenarios through registerPostgresTests, so E2E runs only the PostgreSQL arm.

Name the profile when reporting “all tests.” The local fast loop, test:full, ci:local, required PR checks and nightly fullCorpus are not interchangeable supersets. Native matrices, sustained persistence validation, browser tests and optional recipe/eval commands have separate responsibilities. A faster nightly E2E schedule does not shorten a PR critical path dominated by persistence. Report matched executed timings separately from dry-run partition estimates, including setup, queueing and retries; never count skip-only output as coverage.

Scale tier

The gate shape and cadence are defined once, by O-CEO-16 (with O-ENG-16 and O-CEO-9) in the Foundations 1 plan; scripts/scale/gates.ts and .github/workflows/scale-tier.yml implement it. This section says how to run it.

bun run test:scale -- --pages 10000 [--engine pglite|postgres] [--seed 1] [--corpus-dir <dir>] [--import-mode cli|content] [--enforce] [--out <file.json>] (scripts/scale/run.ts) generates a deterministic two-source brain from the seed (scripts/scale/fixture.ts: links, timeline bullets, ## Facts and ## Takes fences, partly overlapping bodies, island pages, a seeded dense vector in 16 dimensions per page), and imports it into a fresh brain under a temporary GBRAIN_HOME: PGLite in a fresh data dir, or Postgres in a fresh database created from DATABASE_URL and dropped afterwards. The default --import-mode cli writes the Markdown corpus once into --corpus-dir (reused while its manifest matches) and runs the real gbrain import per source, timing each file from its progress events; --import-mode content keeps the per-page importFromContent loop. It then extracts links, timeline, facts and takes, writes the vectors onto every chunk so the vector arm runs keylessly through queryEmbedFn, and measures p50 over five runs after a warmup for each op, each with a known-answer check: get_health, list_pages, local and MCP-path search, a source-scoped grant search, hybrid query with an injected vector, traverse_graph, get_backlinks and find_orphans, plus a cold-process first query and two concurrent receipt-bearing put_pages. find_orphans is checked with one call at the op's maximum page size: its rows must hold every fixture island and match total_orphans. It captures every statement each op sends and replays the reads under EXPLAIN ANALYZE for the planner check. With the brain closed, it then runs the large-brain operational ceilings through the real CLI (scripts/scale/f4d.ts), each an enforced data check with its measurement in the report's f4d section: a gbrain sync of a fresh source (1,000 files at 10k and up) past a 1 s progress-aware deadline completes (f4d_sync_deadline); gbrain embed --stale against a local stub embedding endpoint stops at its time budget with exit 11, the remaining count and the resume command (f4d_embed_budget_stop; nothing leaves the machine); gbrain serve answers initialize and a search without hitting its boot deadline (f4d_serve_boot); and at 20,000 pages and up, gbrain sources add registers a 20,000-file checkout on a fresh managed brain (f4d_sources_add_20k). Each phase logs its start and end, so a run stopped by a job timeout shows where it was. The report leads with the headline metric, MCP search p50 at the run's size as shipped (no manual ANALYZE).

Exit codes: 0 when every enforced gate passes, or always without --enforce; 1 when an enforced gate fails (each failure names the gate and op and prints its EXPLAIN; the JSON report and a .explain.txt land next to --out); 2 on a usage error; 3 when the harness itself crashed (not a verdict). Enforced under --enforce: import rate, known answers, no-op re-import, no duplicates across sources, the import phase timer, and the stats-dependent gates (PLANNER_HEALTH_ENFORCED in scripts/scale/gates.ts): the Nested Loop inner-loop gate on the key plans, the budgets phase timer, and planner stats. Planner stats are probed after the first timed op, because F4b analyzes on the first planner-sensitive read, and only hot tables above 500 rows must have pg_stats rows (PGLite only: on Postgres autovacuum owns statistics, so a missing row is report-only). Interactive ceilings and calibrated budgets (scripts/scale/budgets.json, written by --calibrate) stay report-only until bun scripts/scale/trend.ts prints "ceilings stable" over the last five nightly runs; a reviewer then sets the repo variable GBRAIN_SCALE_ENFORCE_CEILINGS=1. The same script picks the nightly sizes. Reproduce any report with the command it prints. The fixture's determinism is pinned by test/scripts/scale-fixture.test.ts, the gate policy by test/scripts/scale-gates.test.ts and scale-trend.test.ts, the find_orphans known answer by test/scripts/scale-orphans-verifier.test.ts, and a 40-page enforced run by test/scripts/scale-harness.slow.test.ts.

Authoring gate

Before adding a test, answer four questions in the PR description or the test header:

  1. What observable behavior or contract does it protect?
  2. What credible regression makes it fail?
  3. Why does existing coverage not already catch it?
  4. Does it need a production seam that no production caller needs?

A regression test must fail when its fix is reverted; prove it with scripts/check-test-discriminates.sh (see CONTRIBUTING.md). If question 1 has no answer, or the answer to question 3 names an existing owner at the same boundary, do not add the test.

Good: a test that runs gbrain remote ping against a fake MCP server returning { status: 'failed' } and asserts exit code 1 with the failure reason in the JSON output. It protects a user-visible contract, fails if the poll loop reads the wrong field, and needs no seam.

Bad: a test that reads src/commands/remote.ts and asserts it contains job.status. It passes when the loop is broken in a way that keeps the token, fails on a harmless rename, and duplicates the behavioral test above.

Retiring a test

Delete or merge a test only with evidence, recorded in the PR body:

  1. Name the contract the test claims to protect and classify the evidence case below.
  2. Probe it: make a behavior-breaking edit to the production code (or, for a vacuous assertion, show that such an edit passes), run the test and the surviving owner, then revert. Behavior-preserving edits that fail the test are useful extra evidence of implementation coupling.
  3. Confirm the surviving owner executes (executed-test counts, not skip output) at the same or a more frequent cadence, with an equal or stronger failure gate, per "Coverage responsibilities before consolidation" above.
  4. Remove the deleted file's entries from scripts/ubicloud/weights.json, scripts/test-weights.json, scripts/serial-weights.json and scripts/e2e-weights.json, grep scripts/, .github/, scripts/e2e-test-map.ts and test/fixtures/e2e-unmapped-baseline.txt for the path, and regenerate scripts/structural-suites.tsv (bun scripts/classify-tests.ts).

Evidence cases:

  • Retained contract: the contract still matters. Evidence is a surviving owner at the same boundary plus an executed mutation that fails it.
  • Intentionally abandoned contract: the behavior is being removed or was never shipped. Evidence is the approved disposition plus reachability proof (no production caller) and a check that no user-facing promise (docs, skills, --help, CHANGELOG) still describes it.
  • Vacuous assertion: the test asserts nothing about product behavior (a constant compared to itself, a copied function, a typeof probe that typecheck already enforces). Evidence is a demonstration that a behavior-breaking edit leaves it passing, or that it imports no product code.

Evidence template:

Deleted test Probe edit Result Surviving owner Owner result
test/x.test.ts › "name" src/y.ts: what changed deleted test passes (blind) test/z.test.ts › "name" fails (N of M)

Test command tiers

The sequential E2E runner gives each test file a fresh HOME and GBRAIN_HOME. Configuration written by a CLI initialization or schema migration remains available within that file, but cannot change a later file's selected schema or harness state. Each file's home is removed after it exits, including failures; the runner's exit trap also cleans up interrupted runs.

Test command tiers, each with a clear scope:

Command What it runs Wallclock When to use
bun run test Parallel unit-test fast loop. Sharded fan-out via scripts/run-unit-parallel.sh (default 4 shards — CPU-detected, clamped to a max of 8; 4 limits local PGLite WASM-init contention; GitHub CI uses 8 unit shards), then a serial pass over *.serial.test.ts. Excludes *.slow.test.ts and test/e2e/*. No pre-checks, no typecheck. Builds/refreshes the PGLite schema snapshot BEFORE the shard fan-out and exports GBRAIN_PGLITE_SNAPSHOT so PGLite-booting files restore a baked schema instead of replaying every migration (~3.5x per booting file; see "PGLite schema snapshot" below). Opt out: GBRAIN_NO_SNAPSHOT=1. Memory-safe by default: total concurrency (shards × intra-shard width) is capped to available memory at GBRAIN_TEST_MEM_PER_FILE_MB (default 1536 — a PGLite WASM instance) per concurrent slot, shedding INTRA-SHARD width first and shards only after it (bun's --max-concurrency bounds only test.concurrent tests — 1 file in the corpus — so intra width is nearly free to shed, while every dropped shard removes a whole bun process of real fan-out; shedding shards first would collapse a 16GB box to a serial 1×4 run, measured 3.25× slower than 4×1 on the same machine). Two phantom-failure classes are automatically re-run serially (the rescue pass): failures carrying the WASM out-of-memory signature, and shards killed externally (SIGTERM/SIGKILL well before the shard timeout — sibling workspaces' process cleanup, memory jetsam). On machines without coreutils timeout, the fallback watchdog drops a .watchdog sentinel before TERMing a shard at the cap so the WEDGED/EXIT-HANG classifier stays reachable there (a bare rc=143 would otherwise read as a plain failure). Phantoms pass serially and the run goes green with an oom_rescued note; real failures fail again serially and stay red. Knobs: GBRAIN_TEST_NO_MEM_ADAPT=1, GBRAIN_TEST_NO_OOM_FALLBACK=1, GBRAIN_TEST_MAX_CONCURRENCY (intra-shard, default 4), GBRAIN_TEST_SHARD_TIMEOUT / GBRAIN_TEST_SHARD_KILL_AFTER, plus --shards N / --max-concurrency N / --dry-run script args. a few minutes on a Mac dev box Inner edit loop. Default.
bun run verify CI's authoritative pre-test gate set, fanned out by scripts/run-verify-parallel.sh through a bounded worker pool (default detect_cpus; override GBRAIN_VERIFY_MAX_PARALLEL) with the heavy checks ordered first (typecheck, the two compile-embed checks, admin build, fuzz bundles, guard self-tests, the PGLite-booting chronicle eval check, whole-tree greps). The battery includes the deterministic check:eval-chronicle eval gate; check:eval-canary is deliberately NOT in the battery (its test-file twin test/eval-canary.test.ts spawns the identical runner in the unit matrix, and CI's verify job and matrix always run together — the package script stays for on-demand runs, so verify-only local callers should know the canary rides the unit lane instead). The CHECKS array in that script is the single source of truth — CI literally calls bun run verify in a dedicated job. ~50s (pool-bounded; longest check dominates) Before pushing; before /ship.
bun run test:full verify && bun run test && bun run test:slow && [smart e2e]. Smart e2e runs only when DATABASE_URL is set and propagates its failure; otherwise it prints a skip notice to stderr. Use ci:local to provision the databases and require PgBouncer execution. ~3-5min depending on slow + e2e Pre-merge sanity, before opening a PR.
bun run ci:local Independent host gitleaks scans, then frozen dependencies, guards/typecheck, the complete serial and slow lanes, and four unit/E2E shards inside Docker. Each E2E shard has its own pgvector database; selected PgBouncer tests must execute against the transaction-mode pooler. Unit, serial, and slow lanes have database URL overrides unset. Any failed stage fails the command. Complete shard logs survive container teardown under .context/ci-local-shards/. ci:local:diff narrows E2E selection; --no-shard runs unit/E2E sequentially. Doc-only diffs still require successful gitleaks scans. Depends on the full corpus Full local gate before shipping.
bun run ci:ubicloud The ci:local lanes (gitleaks, guards/typecheck, serial, slow, unit, all E2E with required PgBouncer execution) fanned out across ephemeral Ubicloud VMs from one heaviest-first work queue; ci:ubicloud:diff narrows E2E like ci:local:diff. Needs UBICLOUD_API_KEY or UBICLOUD_API_TOKEN, no local Docker. See "Ubicloud fan-out" below. ~5 min (floor: the longest single file) Full gate before shipping when a Ubicloud token is available.
bun run test:slow Just the *.slow.test.ts set (intentional cold-path correctness checks). seconds-to-minutes When touching slow-path code.
bun run test:serial Just the *.serial.test.ts set (cross-file-contention quarantine; one bun process per file for true module-registry isolation), run through a POOL of concurrent per-file processes — the isolation is per-process, not per-machine. Dispatch is heaviest-first (LPT) from the advisory scripts/serial-weights.json (seconds; mined from the .context/serial-durations.txt table each run banks; absent/corrupt weights fall back to discovery order, absent keys to the corpus p75 — scheduling only, never correctness; LPT order + the corrupt-weights fail-soft are pinned by test/scripts/run-serial-pool.test.ts). Pool defaults to min(detect_cpus, 4) then memory-adapts (same doctrine as the parallel runner); a small growth-guarded set of files (machine-global state or contention-critical timing — see the justified EXCLUSIVE_FILES list in scripts/run-serial-tests.sh, capped at 3 by test/scripts/serial-files.test.ts) runs on a sequential EXCLUSIVE lane after the pool. Per-test timeout 120s (pooled contention headroom); each pooled file is wall-clock-killed at 300s (timeout -k, exit-hang containment). SHARD=N/M partitions pooled files by duration; the three exclusive files run only on shard 1. Unset runs the complete corpus. Routing variables are cleared before tests start, so nested runners remain independent. Externally-killed files (exit 143/137 or a missing exit sentinel — sibling-workspace cleanup, memory jetsam) get ONE sequential rescue re-run, mirroring the parallel runner's doctrine: phantoms stay green with a rescue note, real failures stay red. Prints per-file PASS lines plus a top-10 slowest-files list. Knobs: GBRAIN_SERIAL_POOL=N (explicit pool width — bypasses the memory clamp; 1 restores fully-sequential), GBRAIN_SERIAL_FILE_TIMEOUT. a few minutes for all ~220 files at pool=4 Debugging quarantined files; CI's serial-tests job.
bun run test:e2e Real Postgres E2E. Requires Docker + DATABASE_URL. Sequential within a shard; SHARD=N/M fans out against separate databases (ci-local runs 4 containers). Activates the PGLite snapshot like every other runner (per-file cold-path opt-outs where the test asserts the path TO post-initSchema state), exporting it as an ABSOLUTE path so CLI children spawned with varying cwd still find it. ~5-10min Pre-ship; nightly.
bun run test:compile-smoke Self-update integrity verify under a REAL bun build --compile binary, offline (sets GBRAIN_SELFUPDATE_COMPILE_SMOKE=1). The unit suite mocks the network seams; this proves the dependency-free crypto/base64/JSON verify path survives compilation — the failure mode sigstore-js would have hit. ~5s (one compile) When touching src/core/binary-self-update.ts; pre-ship on self-update changes.
bun run test:admin Pinned Playwright Chromium tests for the production embedded admin UI, served with an isolated temporary home/cwd and in-memory PGLite. Exercises owner login, OAuth consent, registration, setup, and lifecycle actions. seconds-to-minutes When touching the admin browser flow; required admin-browser CI job.

For the admin browser lane, install frozen dependencies in the repository and admin/, run bunx playwright install --with-deps chromium on Linux, then run bun run build:admin before bun run test:admin. Tests live in admin/e2e/*.pw.ts so Bun's unit-test discovery does not execute them. The browser suite proves the GBrain dashboard journey; it does not establish activation inside a native vendor harness.

There is no check:all script: a second, hand-synced guard registry would drift from verify, leaving checks that never run anywhere. The CHECKS array in scripts/run-verify-parallel.sh is the single execution list (including check:newlines, check:exports-count, check:no-legacy-getconnection). The guard REGISTRY is scripts/guards-manifest.tsv (see "Guard registry and self-test" below).

bun run typecheck uses TypeScript's native incremental analysis in node_modules/.cache/gbrain-typecheck.tsbuildinfo. Every invocation still runs the compiler; source, root-file, configuration and dependency changes invalidate the affected analysis, and cached diagnostics remain failures. The cache is local and ignored by Git; CI does not restore prior typecheck results.

The local Docker runner isolates root and admin node_modules, plus the generated admin bundle, in named volumes. Admin build dependencies, Vite's generated cache and build output stay inside container volumes instead of replacing host files or leaving root-owned directories behind. ci:local --clean removes these volumes too; build the admin app on the host when updating its committed bundle.

Ubicloud fan-out (ci:ubicloud)

scripts/ci-ubicloud.ts runs the ci:local gate on ephemeral Ubicloud VMs instead of one Docker host. It packs the working tree once (tracked files, untracked files that are not ignored, and .git) and streams it to every VM, so uncommitted edits are tested. scripts/ubicloud/ubi-runner.sh creates and destroys the VMs; every VM is destroyed on exit, including Ctrl-C, and any ubirun-* VM older than 12 hours is garbage-collected by the next run.

Each VM runs scripts/ubicloud/setup-ci-vm.sh: the pinned Bun from docker-compose.ci.yml, the runner container's test prerequisites plus Node, frozen dependencies, both PGLite snapshot fixtures, and one pgvector/pgvector:pg16 server fronted by a transaction-mode PgBouncer per slot. Every slot's schema is bootstrapped with setupLegacyEmbeddingDB(), the same step nightly full-corpus E2E workers run, so no E2E file depends on which file reaches a database first. Setup takes 70-90 seconds, including VM boot.

Scheduling is dynamic. Every unit, serial, slow and E2E file is one item in a global queue ordered by weight, heaviest first. Each idle slot on any VM takes the next item, so a slow VM or a mis-weighted file delays only the slot holding it. Light items leave in same-lane batches to amortize SSH round trips, and the batch target shrinks as the queue drains. Items of 60 seconds or more are the run's long poles, so they spread one per VM before any VM takes a second one. The first VM to finish setup runs the machine-level work first: gitleaks, verify, then the serial lane's machine-exclusive files one at a time with nothing else on that VM. After that it joins the pool. Items run through the ci:local wrappers (scripts/ubicloud/ci-item.sh): run-unit-shard.sh, run-serial-tests.sh and run-slow-tests.sh accept explicit file arguments for this purpose, and run-e2e.sh runs each E2E file against its slot's own server and pooler. The unit, serial and slow lanes run with database URLs unset. Tests run natively as a non-root user on Ubuntu 24.04, the same OS as the CI runners, instead of as root in the oven/bun container.

Weights come from, in order: .context/ci-ubicloud/weights.json (merged after every run), the committed scripts/ubicloud/weights.json (refresh it with --record-weights on a green full run), then the lane weight files mined from GitHub CI. Unknown files get their lane's p75. Per-item logs, failure logs and summary.json land in .context/ci-ubicloud/<run>/. The exit status is non-zero when any item fails, an item never produces a result, or no VM becomes ready. An item whose SSH batch dies without a result is retried once on any slot, and a VM with three such infrastructure errors is retired.

Defaults are four standard-16 VMs (64 vCPUs) in eu-central-h1 with 8 slots each, one per two vCPUs (--vms, --size, --slots, --location). The Ubicloud project's vCPU quota (256) is shared with pull-request CI, so the default leaves room for about two concurrent PR runs; the former default of ten VMs took 160 vCPUs and queued PR jobs for up to 28 minutes. A VM that the quota refuses fails to provision and the run continues on the VMs that did start, so a busy project shrinks the fleet instead of failing. Pass --vms 10 only when the quota is idle. Slow-lane items run test/export-scale.slow.test.ts at the pull-request scale (GBRAIN_TEST_EXPORT_SCALE_PAGES=10001). --lanes runs a subset, --keep leaves the VMs up for debugging, and --diff follows ci:local:diff (a doc-only diff runs gitleaks alone). The corpus is roughly 8,000 seconds of test compute at that density, so 80 slots finish everything but the longest files about two minutes after setup; more slots per VM add CPU contention that slows timing-sensitive files without shortening the run. Wall time is bounded by setup plus the longest single file, test/reindex-markdown-persistence.slow.test.ts (one test, about 230 seconds), so adding VMs past the default does not shorten a run.

E2E backend matrix

scripts/e2e-backend-matrix.txt lists the E2E files that must pass on direct Postgres and through a transaction-mode PgBouncer: the E5 executor binding matrix (test/e2e/executor-binding-matrix.test.ts, whose PGLite arm is test/executor-binding-matrix.test.ts) and every test/e2e/*parity* file. When GBRAIN_PGBOUNCER_E2E_URL is set, scripts/run-e2e.sh runs each listed file twice: first against DATABASE_URL with GBRAIN_TEST_BACKEND=postgres-direct, then with DATABASE_URL set to the pooled URL and GBRAIN_TEST_BACKEND=pgbouncer. The PGLite arm inside each parity file runs in both passes. Both passes must execute the same, non-zero number of tests, and the summary prints the per-backend counts. The pooled URL must carry ?prepare=false, because resolvePrepare only auto-detects port 6543 and CI poolers listen elsewhere; the runner refuses a pooled URL without it. With GBRAIN_CI_REQUIRE_PGBOUNCER=1, a listed file fails when no pooled URL is configured.

Instead of a full URL, a lane may set GBRAIN_PGBOUNCER_E2E_DB=<name>: the runner then reaches that database through the pooler in GBRAIN_PGBOUNCER_URL, pins prepare=false itself, and creates the database on first use through GBRAIN_PGBOUNCER_DIRECT_URL (scripts/lib/ensure-e2e-database.ts). ci:ubicloud routes each slot's own pooler at the slot database, ci:local gives each shard a gbrain_pooled_<N>_test database behind its single pooler, and e2e.yml tier1 runs the list against a pgbouncer service. An entry may carry <TAB>pooled-timeout=<seconds> when its pooled pass needs more than the per-file cap; !path<TAB>reason records a parity file deliberately left out. test/scripts/e2e-backend-matrix.test.ts pins the list's completeness, the CI wiring and the runner's count assertion.

Engine-sql

The engine-sql executor (src/core/engine-sql/, refactor wave 1 W1) is pinned by these tests; the *-parity and RLS files run on every backend in the matrix above, and each E2E file keeps a PGLite arm in the unit lane.

  • test/executor-binding-matrix.test.ts / test/e2e/executor-binding-matrix.test.ts: the E5 case table runs twice per backend, through engine.executeRaw and through the dialect adapters (engineSqlExecutor factory), so the adapters bind, count, fail and cancel exactly like master's raw path.
  • test/engine-sql-executor.test.ts: sqlFragment renders the same text and values as the postgres.js tagged template (vendored serializer); Postgres driver options (prepare: true, simple: false for converted statements, master's options for executeRaw / unsafe), gauge bypass, EO1 transaction lane, brand @ts-expect-error fixtures.
  • test/e2e/engine-sql-prepare-parity.test.ts: pg_prepared_statements holds a converted statement on direct Postgres and nothing through PgBouncer; a zero-parameter multi-statement string is rejected on every backend.
  • test/engine-sql-transaction.test.ts / test/e2e/engine-sql-transaction-parity.test.ts: per-domain write-then-throw rollback through engine.transaction() and transactionDirect() (dual pool on Postgres), with a concurrent pool read. Add a case to test/helpers/engine-sql-rollback-cases.ts for every migrated domain write. A mutation that caches the executor on the engine fails it (DISCRIMINATE_BASE=<mutation> bash scripts/check-test-discriminates.sh).
  • test/engine-sql-capabilities.test.ts / test/e2e/engine-sql-capabilities-parity.test.ts: each dialect capability with a boundary-size and a concurrent-write case.
  • test/e2e/engine-sql-normalize-parity.test.ts: every declared column kind of normalize.ts decodes to one shape on each backend.
  • test/e2e/engine-sql-rls-scope.test.ts: ScopedRead reads under a non-owner NOBYPASSRLS role (cross-source denial, concurrent isolation, nested rollback restoration, connection reuse).
  • SQL text: test/engine-sql-sql-text.test.ts goldens must stay byte-identical after a conversion; only sql-text/_driver.json moves (tagged -> runUnsafe).

Native writer locks

bun test test/native-lock.test.ts test/scripts/native-lock-prebuilds.test.ts checks real process exclusion, crash handoff, retained files, cancellation, missing-addon failure and source/binary manifest integrity. Tests use isolated temporary paths and never open an operator datastore. The required native-locks.yml lane rebuilds and executes all eight OS/architecture/libc targets on Bun 1.4.0 and 1.4.2, including native musl Docker userspace. Every pair also runs bun scripts/native/compiled-smoke.ts to prove compiled process locking. Release CI verifies the shipped CLI embeds the matching addon and runs the compiled smoke on its two release platforms. Rebuild instructions and the precise packaging/runtime distinction are in native/locks/README.md. Release compilation uses Bun 1.4.2; strict Darwin codesign verification must pass before publication. The native macOS 26.2 smoke is not macOS 27 certification, and Linux fault injection is not a full native Windows backup create/restore test.

The glibc Linux, macOS and Windows matrix also runs persistence-publication-native.serial.test.ts, persistence-git-publication.test.ts, persistence-sync-origin-native.serial.test.ts and backup-portability-native.serial.test.ts in separate Bun processes. These exercise real Git publication, historical source origins and backup restoration with fresh-process reopen. macOS and Windows set GBRAIN_TEST_REQUIRE_CASE_INSENSITIVE=1; a case-sensitive fixture must fail rather than silently skip the publication regression. Inspect executed test counts before claiming native coverage; a configured lane alone is not evidence.

The publication, sync-origin, processing-option and company-sync suites also run separately with an explicit PostgreSQL URL in the persistence deployment matrix. test/e2e/persistence-publication-parity.test.ts and test/e2e/persistence-sync-{origin,options,company}-parity.test.ts include them in the full local E2E gate; the keyless serial runner alone cannot exercise their PostgreSQL branches.

The OpenClaw 2026.9.4 / Node 24.18.0 native-host fixture proves plugin startup, restarted-turn saved-page pointer retrieval and same-slug source isolation with a deterministic loopback provider:

GBRAIN_TEST_OPENCLAW_BIN=<absolute-installed-cli> \
GBRAIN_TEST_OPENCLAW_DATABASE_URL=<isolated-postgres-test-db> \
bun test test/openclaw-context-engine-native.serial.test.ts

The database user needs CREATEDB; fixtures create/drop unique databases rather than truncating shared rows. Real-provider recall and macOS 27 behavior remain unverified.

Focused safety coverage: test/apply-migrations-safety.serial.test.ts checks force dry-run previews before DB/ledger access and failed-phase partial exit; test/real-home-guard-preload.test.ts pins the test-home fingerprint backstop (detection, not prevention). Managed retry, durable diagnostics, restart and PGLite/Postgres parity are covered by test/persistence-sync-failures.serial.test.ts and test/e2e/managed-sync-failures.test.ts. Backup remote readback and fsync fault cases run in test/backup-verification.serial.test.ts and test/backup-fsync.serial.test.ts; test/e2e/backup-coverage-parity.test.ts covers PGLite/Postgres page/fact/config parity. Output redaction uses test/search/output-redaction.serial.test.ts and test/search/output-redaction.test.ts, including unchanged internal capture.

Managed writer fixtures use isolated PGLite and guarded disposable Postgres: test/e2e/fact-vector-repair-parity.test.ts, test/e2e/fact-embedding-backfill-parity.test.ts, and test/fact-backfill-resident.test.ts cover preserved vectors, bounded NULL-only fact backfill, selected-config refusal and owner-held PGLite IPC; test/ai/google-embed-batch-items.test.ts pins 100-item provider batches. test/persistence-embedding-effects.test.ts, test/persistence-effect-retry.test.ts, and test/embedding-completion-atomic.serial.test.ts cover partial vector completion, exhausted durable attempts and state-bound explicit retry. test/managed-extract-atoms.test.ts, test/managed-facts-backstop.test.ts and their test/e2e/ counterparts exercise admitted atom/fact replay, including fresh-process facts authority. test/persistence-connectors.test.ts covers managed bound/unbound Google/GitHub sources, API pagination and source-scoped deletions. test/persistence-connector-retry.test.ts covers explicit retry, compaction, checkpoint dependency identity, concurrent approval and lost acknowledgements. Each suite creates its own home, engines and lifecycle through test/helpers/connector-fixture.ts; the helper shares no live engine or mutable suite state. Their separate E2E entry points, test/e2e/managed-connector-routing.test.ts and test/e2e/managed-connector-retry.test.ts, retain the runner's default 180-second per-file cap without duplicating the base cases in the retry lane. Linux root runners execute the complete EACCES case in an isolated setpriv child and assert UID 65534 before testing permissions. This needs a readable checkout, not changes to the parent process identity or checkout permissions; the CI runner image supplies setpriv. test/managed-maintenance.test.ts and test/helpers/maintenance-restart.ts cover local synthesize/patterns/ consolidation, restart replay, retired takes and semantic snapshots; test/managed-unsupported-preflight.serial.test.ts checks that the managed facts-family bulk lanes refuse an unaccepted writer before spend, and test/managed-facts-writers.test.ts proves each of them (fence reconcile, phantom redirect, fence writes, loops extraction, bulk conversation facts) publishes through the coordinator on PGLite and Postgres. These use synthetic provider/API transports, not paid model calls or production connectors. PGLite dream/job CLI with an active owner is not proven delegated by the live fact-backfill IPC test.

test/facts-worker-config.test.ts and its PostgreSQL E2E counterpart dispose the original consumer before executing a real facts-absorb job. They verify the worker passes trusted selected configuration, ignores job-supplied configuration and settles the entity-page effect with zero fact or chunk embedding calls when disabled. Fact extraction still captures the generated fact with a NULL embedding.

test/managed-facts-embedding.test.ts and its PostgreSQL counterpart bind retained fact vectors to the selected brain's model and dimensions, including equal-width host/mount mismatches, keyless capture, policy changes and replay without new spend. test/managed-atom-regressions.test.ts and its PostgreSQL counterpart preserve later target edits through explicit retries and honor database-only storage policy without relaxing source authority. test/managed-synthesis-postprocess.test.ts and its E2E wrapper verify that completed quote/provenance work never rewrites a later user edit, while unfinished work resumes against its original revision. The synthesis suite also preserves the existing same-date summary on replay and rebuilds a complete index after partial recovery. test/managed-atom-compaction.test.ts and its PostgreSQL counterpart age and compact real receipts: permanent completion identity still prevents repeated extraction, while expired retry payloads produce an explicit refusal without changing terminal outcomes or compaction accounting. test/managed-facts-compaction.test.ts and its PostgreSQL counterpart cover the same lifetime boundary for explicit and derived fact-batch identities, including failed or partially committed batches and successful replay without new spend. Connector sweep fencing and physical-path normalization have separate parity coverage in test/persistence-connector-fencing.test.ts. Standalone crash/recovery cases live in test/persistence-connector-recovery.test.ts and their own E2E wrapper so they do not share the routing file's wall-clock budget; their original assertions, child watchdogs and per-file timeout are unchanged.

test/managed-atoms-cli.slow.test.ts exercises real disk-backed PGLite CLI recovery with a loopback provider: live-owner refusal, graceful owner stop, malformed extraction, explicit same-input retry, idempotent replay and owner restart. Fresh-process readback checks the private canonical file, searchable chunk, retained failure receipt, committed completion and released leases. test/managed-connector-routing.serial.test.ts pins actual activation and performSync routing for API sources; the maintenance suite also drives runCycle with eligible facts in two sources and proves the other source is unchanged. The E2E wrapper files ensure these optional PostgreSQL arms execute in the database lane rather than only passing their PGLite controls.

The persistence invariant jobs run the complete scripts/persistence/validate.ts gate (10,000-write soak) on pushes to master and manual dispatches. Pull requests run the same schedules and crash boundaries with a 2,500-write soak; the full PGLite soak alone takes 15-28 minutes and would otherwise set every PR's wall time. test/scripts/data-safety-native-workflow.test.ts pins the split. The crash robot (generated sequences of real operations, SIGKILL at every crash seam they reach, process faults, the reference model) runs as its own job beside the soak: 150 s on pull requests, 600 s elsewhere, Postgres through a transaction-mode PgBouncer. That job also replays the shrunk crash-robot regressions (test/persistence-crash-robot.slow.test.ts) and the history fixture test on both engines; see scripts/persistence/README.md.

For platform-only feedback, dispatch gh workflow run test.yml --ref <branch> -f native_only=true. This explicit manual option uses a separate concurrency group so it does not cancel an ongoing full persistence soak. Its native-only-validation-scope artifact records the exact commit and full_ci: false; it never emits the required test-status check for unrun full CI. Omitting the option preserves every normal PR, push and full manual gate.

Datastore shutdown and lease ownership

test/pglite-lock.test.ts proves process pause/crash handoff, metadata damage, legacy migration refusal and stable ownership across datastore replacement. test/pglite-engine-disconnect.serial.test.ts uses actual disk-backed PGLite for concurrent opens, consumer/statement drains, persisted reopen, delayed close and failed close. A close deadline retains the kernel lock; it is never successful shutdown evidence. Watchdog and telemetry regression suites cover loop starvation and background statement teardown.

test/db-lock-concurrency.test.ts proves unique identities even when two acquisitions have identical database timestamps, exact successor-safe cleanup, renewal cancellation/late-completion drain and mandatory loss propagation. test/e2e/db-lock-acquisition-token.test.ts repeats acquisition/cleanup invariants against real Postgres; the E2E map selects it for lease and engine changes. test/engine-control-routing.test.ts pins direct/shared pool routing, nested transaction confinement and the Postgres resident-stop barrier.

Durable persistence schedules and process crashes

test/persistence-consumer-scheduling.test.ts pins completion wake-ups, including a wake-up arriving during an active tick, without lowering the idle poll interval. Per-root deadlines preserve blocked/retryable backoff even while another root keeps committing; expired deadlines permit retries. Shutdown drains active preparation without starting another request. test/persistence-root-refresh.test.ts checks that unchanged root registrations do not replace their durable files while unbound, moved and original bound paths all remain fenced.

test/e2e/persistence-http-liveness.test.ts drives real authenticated HTTP MCP with legacy and source-bound OAuth credentials against an isolated PostgreSQL database. It covers pending page/fact writes, lost acknowledgments, owner interruption, exact canonical readback, 44 pages over three clients, and current receipt authorization. Set GBRAIN_TEST_OLD_BINARY to a retained compatible executable to additionally exercise old/new owner handoffs in both directions for queued, actually claimed and recovery-bearing rows. Interrupted owners are reaped and native exclusion is verified before their successors start; only the abandoned running lease is advanced by the fixture. An omitted executable skips these compatibility cases, not proves them. test/e2e/persistence-phase-liveness.test.ts holds real table/row locks and ordinary/direct pool slots to verify phase cancellation (including capacity marking), tracked queued BEGIN, renewal, expired-head FIFO and shutdown fences. A loopback TCP gate separately delays or rejects cold direct initialization, checking retained work through shutdown and same-ID completion after retry. The memory-mutation tests use a deterministic loopback embedding transport for slow and aborted preparation; they do not contact a paid provider. Receipt contract, MCP parser, IPC and CLI tests pin the same nested allowlist and advisory age policy. Diagnostic tests exercise both fresh/index-upgrade parity and a genuinely interrupted PostgreSQL concurrent index build when a PostgreSQL fixture is supplied.

test/persistence-chaos.slow.test.ts and test/e2e/persistence-chaos.test.ts execute real journal/coordinator schedules and eight SIGKILL publication boundaries, followed by a small multi-process soak. The Postgres test creates and drops fresh test databases, requiring CREATEDB on the explicit test URL. It never truncates the shared E2E database. The reusable persistence-validation.yml gate runs 1,000 schedules and 10,000 writes per engine under Bun 1.4.0 and 1.4.2 and uploads actual executed-case manifests. See scripts/persistence/README.md for workloads, reruns, performance measurements and the process-crash scope.

test/e2e/persistence-runtime-matrix.test.ts additionally requires the real transaction-mode PgBouncer fixture. Its 24 cells exercise direct/pooler connections, enforced RLS under a non-bypass role, ordinary pool sizes 1/2/3, and shared pools or a separate one-connection direct route. It verifies reserved short control capacity while bulk connections remain held, then drains and commits the original request. Ownership cases cover mismatched successor manifests, stale owners, root replacement under a held kernel lock, and actual source deletion/recreation. The reusable persistence lane runs this matrix on both supported Bun versions and uploads its manifest.

The required persistence lane also runs scripts/persistence/performance.ts on both engines and Bun versions. Three independent instances use the existing 500-page/200-query read-latency corpus, with public put_page mutations and actual in-flight interval coverage of at least 90%. Any read or write failure invalidates the sample. CI uses --informational: the 50% loaded-versus-idle p99 threshold is advisory, not a merge blocker. Loaded reads compete with additional write work, so this ratio alone is not evidence of a change regressing the same workload. Manifests retain the original threshold verdict, each sample, admission/commit latency, queue age, RSS, recovery bytes and pool activity. The CLI without --informational still enforces the threshold. The original heavy shell entry invokes this harness; its optional STRICT_LATENCY=1 flag affects only the latency threshold, never validity requirements.

Engine graduation

Graduation (gbrain migrate --to postgres) is tested against two fixtures. test/fixtures/graduation/legacy-brain.ts builds a small brain by hand-written SQL on a fresh schema, and its expected outcomes live in expected.json beside it (test/graduation-legacy-fixture.test.ts proves the build matches and that the target checker discriminates). scripts/persistence/graduation-fixture.ts wraps buildHistoryFixture for the 1k and 10k history brains, builds keylessly in a child process, and caches each build as a tarball keyed by the fixture sources, schema version, seed and size (GBRAIN_GRADUATION_FIXTURE_CACHE, default ~/.cache/gbrain-graduation-fixtures). A restore re-homes paths and owner stamps and marks sources synced, so the source doctor stays green.

The E2E suites drive the real CLI in child processes: graduation-crash (SIGKILL at every run and rollback boundary), graduation-faults (ENOSPC on a tmpfs tablespace, which needs Docker; password rotation; DDL route mismatch), graduation-clients (older releases, respawned and resident serve, stale CLI and MCP configs) and graduation-cli (agent flow, zero-mutation --plan/--status, --force, PgBouncer through GBRAIN_PGBOUNCER_URL, a NOSUPERUSER role, the 1k round trip). Kill and pause points come from graduationBoundary() hooks that only test/helpers/graduation-hooks-preload.ts registers. Older release binaries are built once per tag under GBRAIN_OLDER_RELEASE_DIR. The crash suite takes about 40 seconds per case, so run it with GBRAIN_E2E_FILE_TIMEOUT=3600. scripts/persistence/graduation-ttv.ts records the commands and wall time from the first plan to a green doctor, the run's phase timings and query p50/p95 on both engines. The 1k-page gate is five minutes; tests/heavy/graduation_10k.sh reports the 10k run.

PGLite schema snapshot (default-on)

scripts/build-pglite-snapshot.ts (bun run build:pglite-snapshot) bakes a post-initSchema() PGLite data dir into test/fixtures/pglite-snapshot.tar plus a version file; PGLiteEngine.initSchema() restores the tar instead of replaying the embedded schema + all migrations when the env var GBRAIN_PGLITE_SNAPSHOT points at it. Runners activate it through the shared ensure_pglite_snapshot helper in scripts/lib/test-env.sh (also home of detect_cpus and detect_available_mem_mb), sourced by run-unit-parallel.sh, test-shard.sh, run-slow-tests.sh, run-serial-tests.sh, run-verify-parallel.sh, and run-e2e.sh (which re-exports the path as ABSOLUTE — its tests spawn CLI children with varying cwd); scripts/ci-local.sh calls the builder directly. The helper builds/refreshes the snapshot and exports the env var, no-ops on GBRAIN_NO_SNAPSHOT=1 or an already-inherited path, and is non-fatal on build failure — tests fall back to cold init, with a one-line "active" echo so a silent fallback stays visible in CI logs. Measured effect: ~3.5x per PGLite-booting file (a cold boot replays every migration, ~3.1s each on a CI shard). Properties:

  • Idempotent. A hash short-circuit exits in ~40ms when the snapshot is fresh, and REBUILDS a stale one. The hash covers the raw file bytes of the static import closure of pglite-schema.ts, the schema-migration registry, migrate.ts and the forward-reference bootstrap (engine-sql/bootstrap.ts), plus pglite-engine.ts (src/core/snapshot-schema-inputs.ts computes the list; no hand list). Imported SQL and handler changes invalidate the fixture; coverage instrumentation does not change the hash. test/snapshot-inputs-closure.test.ts checks that list against an independent TS-AST closure, requires every literal dynamic import in the closure to be classified, and discovers all 13 pglite-snapshot-* CI cache keys: identical hashFiles inputs covering every hash input, each profile restoring its own tar. A failure names the missing file and both workflow files to edit.

  • Concurrency-safe. Each profile has its own lock with a PID/token owner and host/process-namespace identity. Only a confirmed dead local owner using the current retirement protocol can be reclaimed. Both normal release and crash recovery retain a nonempty owner tombstone so a delayed observer cannot remove the next builder's lock. Keep those records while builders may run. Live, foreign, ownerless, or older-protocol locks time out without building; callers visibly fall back to cold initialization. Temporary tar/version files are atomically renamed, with the version last.

  • Never authoritative. The loader (tryLoadSnapshot in src/core/pglite-engine.ts) verifies the schema hash AND the embedding shape the snapshot was baked with (dims= / model= lines in the version file) against what this process would create; any mismatch — including a version file without shape lines — warns once and falls through to normal cold init. A wrong fixture can never poison the suite.

  • Opt out. GBRAIN_NO_SNAPSHOT=1 skips the build + env export for a run; the migration-replay canary tests clear the env themselves regardless.

Pinned by test/snapshot-shape-guard.test.ts (hash + shape refusal matrix, imported SQL/handler dependency hash sensitivity).

The builder accepts --profile legacy|default (legacy remains the default). Legacy uses the unit preload's embedding shape. Default uses the CLI's canonical embedding shape and writes pglite-snapshot-default.tar plus its .version. The artifacts, locks, and CI caches are separate. ensure_default_pglite_snapshot exports an absolute GBRAIN_TEST_DEFAULT_SNAPSHOT; BrainBench applies it only to CLI children, including run-all. The parent unit process retains its legacy snapshot. The slow runner and direct BrainBench test invocation prepare the default profile automatically. GBRAIN_NO_SNAPSHOT=1 clears both paths and survives test preloads.

PGLite checkpoint harness (outside PR CI)

scripts/pglite-checkpoint-harness/supervisor.ts reproduces the large-store PGLite freeze. It spawns worker.ts, which imports 40 KB pages through PGLiteEngine.transaction() into one long-lived store. It then watches that process from outside: CPU from /proc, committed pages from the worker's progress file, and WAL and checkpoint activity from the data directory. A worker that burns CPU with no committed page and no checkpoint progress for --stall-sec is reported as wedged. A completed run also asserts that WAL since the last redo point never exceeded the guard threshold plus the largest single transaction. --shared-buffers and --max-wal-size scale the store down with ALTER SYSTEM, so a small machine reaches the same trigger in minutes:

bun scripts/pglite-checkpoint-harness/supervisor.ts --dir /tmp/h --fresh --pages 3000 --stall-sec 600 --timeout-sec 1800
bun scripts/pglite-checkpoint-harness/supervisor.ts --dir /tmp/h --fresh --pages 1500 \
  --shared-buffers 16MB --max-wal-size 160MB --stall-sec 120 --expect-wedge   # baseline check

--min-store-gb <n> also fails a completed run whose data directory is smaller than n GiB; the result reports the store size with and without WAL. On macOS the supervisor samples worker CPU with ps instead of /proc.

macOS 26 validation (macos-validation)

.github/workflows/macos-validation.yml runs on a GitHub-hosted macos-26 runner nightly, on manual dispatch, and on pull requests that carry the macos-validation label. It uses no secrets and has a 90-minute cap. It checks the device-identity re-stamp (#5604) on real APFS (the step first asserts an APFS volume with a non-zero birth time and inode; the device-number change itself stays simulated because a runner never reboots), runs the checkpoint harness above on a store of at least 2 GiB with the WAL-bound assertion (#5449), and verifies the latest published signed darwin-arm64 release binary with codesign --verify --strict and --version against its release tag (#5286). It also runs the bash 3.2 parse guard under the runner's /bin/bash, then bun run verify (#5810), so a script or guard that only works under bash 4 or later fails there. Maintainers with triage or write access apply the label; an outside contributor whose change touches macOS-specific persistence, locking or release code asks for it in the pull request. Scheduled and dispatched runs use the default branch's workflow file, so the label is the way to get this evidence for a change before it merges. test.yml's security matrix, release.yml's darwin build and native-locks.yml's darwin cells pin macos-26 / macos-26-intel rather than macos-latest.

Keeping CI partitions balanced

Required CI runs eight weighted unit workers, four serial workers with bounded per-file pools, and up to four selected E2E workers. E2E selection and exclusions run once before setup; the resulting file lists are frozen and executed against separate Postgres services. An explicit empty selection launches no tests; selection errors, failed workers, cancellations, and unexpected skips fail the existing aggregate checks. Nightly full-corpus E2E uses four independent Postgres jobs with the same weighted partitioner and one fresh Bun process per file, sequential within each job. It does not use the selected-E2E exclusion list: default discovery includes every test/e2e/*.test.ts and test/phantom-redirect-engine-parity.test.ts. Each full-profile worker first initializes its own service schema with the guarded setupLegacyEmbeddingDB() helper, in a temporary home with provider keys stripped and local environment-file loading disabled. No partition relies on a preceding file to create shared tables. Bootstrap is a separate timed CI step and must be included in end-to-end comparisons.

Refresh after a large test wave or when the longest shard repeatedly exceeds the mean shard execution time by 25%:

bun run weights:mine --lane unit --run <successful-test-run>
bun run weights:mine --lane serial --run <successful-test-run>
bun run weights:mine --lane e2e --run <successful-e2e-run>
bun run weights:mine --lane e2e --e2e-profile full --run <successful-full-corpus-run>

The miner accepts --from-file or stdin for timestamped GitHub-format logs and --out for inspection before replacing a checked-in map. Unit timing uses only unit matrix jobs, includes evals/, and closes the final file at the Bun summary. Serial timing uses runner durations, never timestamps of buffered output. E2E uses each file's Bun summary and merges partial selections into known weights. File/stdin imports also merge unobserved entries; only a complete GitHub unit, serial or explicit full-profile E2E run replaces that lane's entire map. Captured artifacts need their final successful completion marker, and GitHub imports verify every expected job. Incomplete or failed inputs leave the existing map intact. Sidecar metadata records the source run/commit, units, and counts. Unit/E2E weights are milliseconds; serial weights remain seconds. New files receive the corpus p75 estimate. Empty maps and zero-cost ties distribute files deterministically; corrupt serial weights warn and retain safe fallback scheduling.

The default E2E miner reads selected jobs and merges partial observations. --e2e-profile full instead requires a successful GitHub run with a source SHA matching checkout HEAD. It pins the run attempt and full-job IDs, reconstructs default discovery using the committed runner and tracked test paths, and requires every source file exactly once across complete successful job logs. Missing/extra files, ambiguous basenames, duplicate execution or failed evidence leave weights and metadata untouched. Basename resolution preserves the outside-directory parity entry. Full mode replaces the complete map and records source SHA, attempt, jobs, log hash and corpus hash; file/stdin imports cannot claim full-profile provenance. These weights schedule work; they are not an observed parallel runtime.

CI retains timestamped unit/E2E logs, frozen E2E selection, and serial attempt records for 14 days. Compare push-to-required-green time including queueing, first failure, rescues/reruns, runner minutes, unique file counts, and coverage completeness. Compare cold and warm caches separately. Snapshot timings and partition estimates are projections until matched workflow runs confirm them; successful test results are never cached.

Refactor wave 1 goldens

Outputs captured on master before refactor wave 1 moves any code live under test/fixtures/goldens/; test/fixtures/goldens/README.md maps each file to its owning test and named normalizer. test/helpers/golden.ts writes and compares them (expectGolden) and proves each normalizer by capturing twice (expectNormalizerStable). Regenerate only deliberately, never in a refactor commit: GBRAIN_TEST_UPDATE_GOLDENS=1 bun test <file> (the switch carries the GBRAIN_TEST_ prefix because the unit preload scrubs other GBRAIN_* overrides). Performance baselines are a bench, not a test: docs/designs/refactor-wave-1/perf-baseline.md.

Doctor check registry

gbrain doctor runs DOCTOR_CHECK_REGISTRY (src/commands/doctor/registry.ts) in order: one { name, emits, run(ctx) } entry per topic block under src/commands/doctor/checks/, each returning its checks or STOP_DOCTOR. test/doctor-registry.test.ts fails with a FAIL / Why / Fix / See block when an entry's name or any emits[] name is missing from src/core/doctor-categories.ts, when emits[] differs from what the entry's run can push (AST walk in test/helpers/doctor-registry-ast.ts), or when a STOP gate moves away from where master's buildChecks returned early. test/doctor-mode-matrix.serial.test.ts wraps every entry and the engine with recorders and asserts, per mode (default, --fast, --fix, --fix --dry-run, no engine, connection failure), which entries ran, where the run stopped, which engine calls happened and which mutations landed (the SKILL.md DRY auto-repair, the dead-holder lock reap). The W0 registry, early-stop and --json goldens pin the output itself.

Move-only verifier

scripts/verify-move-only.ts proves a commit tagged Move-Only: yes moves code without editing it: every top-level statement of every touched TS file on the base side reappears token for token on the head side (tokens from scripts/lib/normalize-tokens.ts, so whitespace and comments are ignored and string/SQL text is exact). Imports, export ... from lines and toggling the export modifier on a moved statement are allowed and counted. Run bun scripts/verify-move-only.ts <commit> (default HEAD~1..HEAD); --wrapper migration inlines export const vNNN: Migration = {...} files into the generated registry array so the W3 split must reproduce the base side's single MIGRATIONS literal entry for entry; --wrapper doctor-entry inlines each run<Topic>(ctx: DoctorContext): Promise<Check[]> body (minus its ctx destructure / connectedEngine / const checks prologue and return checks;) at its checks.push(...(await runX(ctx))); call in buildChecks, drops the const ctx: DoctorContext = {...}; glue and resolves relative import() / require() specifiers to repo paths, so the W4 doctor peel must reproduce the original buildChecks body; and --rename-map <json> applies identifier rewrites for Mechanical-Rename: yes commits. Failures print FAIL: <file:line> with the first differing token. Pinned by test/scripts/verify-move-only.test.ts.

Schema migration registry

Schema migrations live one per file in src/core/schema-migrations/v<NNN>-<name>.ts (NNN zero-padded to 3, name = the slug with - → _, one export const v<NNN>: Migration = {...} per file). bun run new:migration <snake_name> scaffolds the next version; bun run build:schema-migrations regenerates the committed static-import registry registry.generated.ts (regenerate, never hand-merge). The array order is master's historical order (HISTORICAL_ARRAY_ORDER in scripts/build-schema-migrations.ts), then ascending; the runner sorts by version. Two guards run in bun run verify:

  • check:schema-migrations (scripts/check-schema-migrations-fresh.sh) regenerates the registry into a temp file and diffs it; the generator also fails on a filename/version/name mismatch and on a version defined twice, naming both files with the git mv + version: + regenerate recipe.
  • check:schema-migration-order (scripts/check-schema-migration-order.ts) fails when a migration origin/master does not have is numbered at or below origin/master's latest version (it would be skipped forever on current brains) or reuses a version with a different name. Base ref: GBRAIN_MIGRATION_BASE_REF (default origin/master); skipped with a notice when the ref is missing, failed under CI=true.

Collision recovery: an unapplied branch migration is renumbered (git mv, edit version, regenerate); one already applied to a disposable dev DB means rebuilding that DB and replaying; one applied to retained data needs explicit schema_version reconciliation, never just a counter edit. Pinned by test/scripts/build-schema-migrations.test.ts and test/migrations-golden.test.ts.

Schema generator freshness

check:schema-fresh (scripts/check-schema-fresh.sh) runs scripts/build-schema.ts --out-dir <tmp> (fragments -> src/schema.sql regions -> schema-embedded.generated.ts -> pglite-schema.generated.ts) and diffs every output, naming the source to edit. Canonical sources and PGLite capability rules: docs/ENGINES.md#canonical-schema-sources. Pinned by test/scripts/build-schema.test.ts; the end state by the E4 catalog goldens.

Guard registry and self-test

The privacy and test-isolation guards use scripts/lib/guard-candidates.sh to scan fresh file contents in bounded batches before applying their detailed per-file rules. They do not cache passing results. Candidate scanner failures fail the guard, and matching files retain the same allowlists and diagnostics.

scripts/guards-manifest.tsv is THE single registry of scripts/check-* guards (currently 68), each classified scanner (greps/parses repo sources — must eventually carry fixtures), buildfresh, or repostate (build/freshness guards are exempt-with-reason, not fixture-tested). scripts/guard-self-test.sh (bun run check:guard-self-test, wired into bun run verify) proves every selftest=yes scanner CAN fail: it runs each one against known-bad (must exit non-zero) and known-good (must pass) fixture trees under test/fixtures/guards/<guard>/{bad,good}/ via the GBRAIN_GUARD_ROOT env seam, and enforces manifest completeness — a new scripts/check-* script that isn't registered in the manifest fails the build. A guard whose pattern rots into a permanently-green no-op fails CI instead of masquerading as coverage.

A guard may carry extra known-bad trees named bad-<variant>/; each one must fail on its own. Refactor wave 1 uses them to prove that every scanner naming a file the wave splits also scans the new module locations (src/core/engine-sql/, src/core/schema-migrations/, src/commands/sync/, src/commands/doctor/checks/, src/commands/serve-http-*.ts, src/core/minions/handlers/): check-jsonb-pattern.sh, check-engine-dynamic-import.sh, check-source-config-leak.sh, check-no-legacy-getconnection.sh, check-operations-filter-bypass.sh, check-source-id-projection.sh (engine-sql) and check-search-path.sh (the generated PGLite template) each have a bad fixture placed inside the new path. The checklist of every script, workflow, helper and doc that names a split file is docs/designs/refactor-wave-1/path-consumers.md.

Layering guard

scripts/check-layering.ts (bun run check:layering, in bun run verify) parses every file under src/core/engine-sql/ and src/core/schema-migrations/ and fails on any import, type-only included, of an engine façade (pglite-engine.ts, postgres-engine.ts, engine-factory.ts) from engine-sql, or of src/core/migrate.ts from schema-migrations. Those directories are loaded by the engines and by migrate.ts, so an import back up is an ESM cycle that can fail with a temporal-dead-zone error at module load. Take the executor as a parameter and import types from src/core/engine.ts; migration helpers live in schema-migrations/helpers.ts and the Migration type in schema-migrations/types.ts. Fixtures: test/fixtures/guards/check-layering.ts/; forms are driven in test/scripts/layering.test.ts.

Durable-flush guard

scripts/check-durable-flush.ts (bun run check:durable-flush, in bun run verify) fails on an fsyncSync(fd) anywhere in src/ outside src/core/fs-durable.ts whose fd is assigned from a read-only openSync (flags omitted, a flag string without w/a/+, or O_RDONLY without O_WRONLY/O_RDWR), file or directory, and on one whose flags it cannot read. Windows refuses fsync on a read-only handle and has no directory flush (EPERM), which wedges the managed write queue (#5595) and every skill-bundle publication (#5475). Flushes of descriptors opened for writing pass. Each failure prints FAIL [durable_flush_read_handle]: <file>:<line>, the open it traced, a Fix: line and this anchor. Fix: fsync the descriptor you wrote through before closing it (set its final mode with fchmodSync(fd) first), or call flushFile(path) / flushDirectory(path, { bestEffort? }) from src/core/fs-durable.ts. A file that cannot migrate yet goes in the guard's ALLOWLIST with a reason (empty today); an entry whose file no longer needs it fails as durable_flush_stale_allowlist. Fixtures: test/fixtures/guards/check-durable-flush.ts/; forms are driven in test/scripts/durable-flush-guard.test.ts. The helper and the #5595/#5475 regressions run natively on the windows-latest row of the test.yml security-regressions job; test/helpers/win32-flush-semantics.ts makes them discriminate on POSIX hosts too.

Engine-sql ratchet

scripts/check-engine-sql-ratchet.ts (bun run check:engine-sql-ratchet, in bun run verify) keeps each storage domain's SQL in one place, src/core/engine-sql/<domain>.ts, by stopping SQL from growing back into the engines. It parses src/core/pglite-engine.ts, src/core/postgres-engine.ts and every file under src/core/pglite-engine/ and src/core/postgres-engine/, and names each class member (PostgresEngine.getPage), top-level function and top-level variable (insertFact). A unit is SQL-bearing when the literal text of a string, template, tagged template or + chain inside it has SQL structure: SELECT ... FROM <x>, SELECT <fn>(, INSERT INTO <x>, UPDATE <x> [alias] SET, DELETE FROM <x>, WITH <x> AS (, CREATE|ALTER|DROP <object kind>, TRUNCATE <x>, SET LOCAL <x>, ON CONFLICT, WHERE ... ORDER BY|GROUP BY|LIMIT, or a $<n>::type cast. Comments and identifiers never count. Keywords match in upper or lower case but never Title Case, and a lowercase match also needs a second SQL signal (where, returning, $1, ::, ;, * and similar), so "Select a file" or "could not delete from cache" is not SQL.

scripts/engine-sql-baseline.tsv lists migrated<TAB><domain> rows (the domain's module must exist under src/core/engine-sql/) and method<TAB><path><TAB><QualifiedName> rows for the SQL-bearing members that remain. Rows only shrink. The guard fails on:

  • a new SQL-bearing member with no row: move the SQL into src/core/engine-sql/<domain>.ts and delegate, or mark the declaration (on its line or the line above) with // engine-sql-ok: <reason>; an empty reason fails;
  • a stale row, whose member is gone, no longer SQL-bearing or now marked: delete it, or run bun scripts/check-engine-sql-ratchet.ts --prune, which drops stale and duplicate rows and never adds one;
  • a duplicate or malformed row, or a migrated row with no module.

When a domain moves, delete its members' rows and add its migrated row in the same commit. Fixtures: test/fixtures/guards/check-engine-sql-ratchet.ts/; forms are driven in test/scripts/engine-sql-ratchet.test.ts.

Engine-sql dynamic SQL

scripts/check-engine-sql-dynamic.ts (bun run check:engine-sql-dynamic, in bun run verify) parses every file under src/core/engine-sql/ except fragment.ts, the renderer, which writes $n and splices trusted text by design. In engine-sql every value reaches SQL as a bound parameter through sqlFragment, and only constant text is spliced. Trusted text is a string literal; a const in the same file initialized with trusted text or an as const object or array literal (members and element accesses included); a CONSTANT_ALLOWLIST name (ENRICH_ORDER_SQL); a call to a VETTED_BUILDERS entry (pageReadFilter, buildRecencyComponentSql, privatePagesFilterFragment, currentCodeEdgeFilter, buildCJKKeywordSql, currentTextProjectionFilter); a template or + chain whose parts are all trusted or are numbers the same function checked earlier with Number.isFinite(<same expression>); or a conditional whose branches are both trusted. Both registries live in the script with a one-line reason each. The guard fails on:

  • trustedSql(arg) with an arg that is not trusted text: bind the value with ${value} in sqlFragment instead, or register a new builder with its reason after review;
  • an untagged template or + concatenation, passed directly or through a local variable (let appends included) as the SQL of .query(, .unsafe(, .executeRaw( or executeRawJsonb(, with an untrusted part: compose with sqlFragment and run it with executor.run(fragment);
  • a literal $<digit>, or a $ right before a substitution, in a composed string (a template with substitutions, any sqlFragment template, any + operand): let renderFragment number the parameters. A static string passed as-is may carry $1;
  • an expanded list, IN ( right before a substitution or a non-literal + operand: bind the array as one parameter, = ANY(${ids}::text[]), so prepared-statement caches stay bounded.

Fixtures: test/fixtures/guards/check-engine-sql-dynamic.ts/; forms are driven in test/scripts/engine-sql-dynamic.test.ts.

Engine-sql brands

scripts/check-engine-sql-brands.ts (bun run check:engine-sql-brands, in bun run verify) keeps the RLS read brands in src/core/engine-sql/brands.ts unforgeable. ScopedRead records a read that ran inside withScopedReadTransaction on master and LegacyUnscopedRead one that ran unscoped on the pool (EO4), so a forged brand silently changes how a read is scoped. The guard fails on:

  • a brand key (any __obtainVia... name) in a text file under src/, test/ or scripts/ other than brands.ts and the guard's own script, fixtures and test: get a branded executor from scopedRead(tx) inside withScopedReadTransaction, or from unscopedExecutor(executor, '<reason>');
  • in src/, a cast onto ScopedRead or LegacyUnscopedRead outside brands.ts, an as unknown as T where T names SqlExecutor, ScopedRead or LegacyUnscopedRead, or a double cast passed straight to scopedRead( or unscopedExecutor(. Driver-handle casts such as tx as unknown as PgConn in dialect-postgres.ts pass;
  • an import of unscopedExecutor or LegacyUnscopedRead (value, type, alias, re-export or import('...').X type) from outside engine-sql, the two engine façades, doctor (src/commands/doctor.ts, src/commands/doctor/**, src/core/doctor*), maintenance (src/core/maintenance/**), admin (src/commands/admin*.ts, src/core/admin/**), migrations (src/core/migrate.ts, src/core/schema-migrations/**, src/commands/migrations/**) and test/; an import of scopedRead from outside engine-sql, the façades and test/; or a namespace, dynamic or require import of brands.ts from outside that scopedRead list. src/core/ops/**, the MCP-facing surface, is always denied. Take the branded executor from the engine façade instead.

The allowlists live in the script. Fixtures: test/fixtures/guards/check-engine-sql-brands.ts/; forms are driven in test/scripts/engine-sql-brands.test.ts.

Retired-phrase guard

scripts/check-retired-phrases.sh (bun run check:retired-phrases, in bun run verify, under a second) keeps the instructions agents follow literally in step with refactor wave 1. It greps CLAUDE.md, AGENTS.md, CONTRIBUTING.md, docs/ and skills/ for the contributor-workflow phrases the wave retired: the old migrations-array wording and appending to it, the rule that every engine method is written twice, the CLI switch-case step, the migrate.ts region policy, and schema text listed as a hand-synced pair of schema.sql and the PGLite schema module. The patterns and the current instruction for each live in the script's RETIRED table. Historical records may quote them and are exempt: docs/designs/, docs/test-audit/, docs/incidents/, docs/plans/, docs/proposals/, docs/research/, docs/issues/, docs/superpowers/, docs/migrations/, skills/migrations/ and the wave 1 porting kit (docs/architecture/wave-1-*); CHANGELOG.md is not scanned. Each hit prints FAIL: <file:line> retired phrase "<match>", then Why:, Fix: with the current instruction, and See:. The fix is to rewrite the sentence to the current workflow, never to exempt the file. To retire another phrase, add a row to RETIRED. Fixtures: test/fixtures/guards/check-retired-phrases.sh/ (one bad-<location> tree per scanned location); every pattern and the exemptions are driven in test/scripts/check-retired-phrases.test.ts.

Bash 3.2 parse guard

macOS ships GNU bash 3.2.57 as /bin/bash, and its parser rejects shapes bash 5 accepts. The one that broke every Mac (#5810) is a heredoc inside $(...) whose body holds an odd quote. scripts/check-bash32.sh (bun run check:bash32) runs the real 3.2 parser, bash -n, over every tracked *.sh except the guard fixtures under test/fixtures/guards/. It uses the first parser available: GBRAIN_BASH32=<path> (a bash 3.x binary), /bin/bash when it is bash 3.x (stock macOS), or the digest-pinned bash:3.2 Docker image (GBRAIN_BASH32=docker forces the image). With none it prints one skip line and exits 0; GBRAIN_BASH32_REQUIRE=1 makes that exit 2. Each failure prints FAIL: <file:line>, Why: (macOS /bin/bash is 3.2), Fix: (read heredoc text with IFS= read -r -d '' VAR <<'EOF' || true) and See:. To reproduce one file by hand: docker run --rm -v "$PWD":/w -w /w bash:3.2 bash -n <file>.

The guard checks parsing only. It is not in bun run verify, which must not need Docker. The test.yml verify job runs it with GBRAIN_BASH32_REQUIRE=1 together with test/scripts/check-bash32.test.ts, whose real-parser cases feed it the test/fixtures/guards/check-bash32.sh/{bad,good} trees. The macOS 26 job runs it under /bin/bash and then runs bun run verify there, which covers bash-4 runtime features the parser cannot see.

Placeholder assertions

scripts/check-test-placeholders.mjs (bun run check:test-placeholders, in bun run verify) parses every test/**/*.test.ts file outside test/fixtures/ with the TypeScript compiler API and fails on the no-op forms expect(true) with no matcher, expect(true).toBe(true), expect(true).toBeTruthy() and expect(1).toBe(1). Text inside strings and template literals is ignored, and expect(true).toBe(false) fail sentinels are allowed. Remaining sites (type-only contracts enforced by typecheck, skip-arm markers, gates that fail by throwing) sit in a reasoned allowlist in the script, keyed by file, test name and exact count; a site above its count fails as new, and an entry whose file, test or count shrank fails as stale. This is a hygiene check for one pattern, not a detector of low-value tests in general; the authoring gate above owns that.

Function-size ratchet

scripts/check-function-size.ts (bun run check:function-size, in bun run verify, about 1.5 s) measures every function-like node in src/**/*.ts except *.generated.ts and .d.ts with the TypeScript compiler API: function declarations, methods, constructors, accessors, arrow functions and function expressions, including object-literal and class-property forms. A nested function is measured on its own, and its lines also count toward the function that contains it. Code under test/ is out of scope.

scripts/function-size-baseline.tsv holds one row per function over 300 lines: path, name, lines, justification. The name is a path built from declarations, property names and call context, never line numbers, so edits above a function do not touch its row: PGLiteEngine.initSchema, runServeHttp>app.post('/mcp'), MIGRATIONS[v131].handler. > enters a function, . a member, = a call whose result is bound, and a repeated key gets a #2 ordinal. The guard fails when a function over 300 lines has no row, a baselined function grows, a baselined function drops to 300 lines or fewer (remove the row), a row has more than 50 lines of stale slack (lower it), a row names a function that no longer exists, or a row is malformed, duplicated or out of order. A row raised above, or added since, the baseline at the merge-base with origin/master needs an issue or TODO id (#1234, TODOS.md:12, TODO: <slug>) in its justification; the summary prints every raise.

Each failure prints FAIL: <file:line> <what> with the computed key, then one Why: / Fix: / See: block. The fix is extraction: move a cohesive block into a named helper or sibling module (phase, stage or handler-table pattern). After a move-only commit changes a function's key, run bun scripts/check-function-size.ts --transfer. It rewrites a missing row to the one unbaselined over-limit function whose whitespace-normalized text is identical to the old function at HEAD (--from <ref> for another base), apart from an added leading export and module specifiers re-relativized to the new directory (each resolved against its own file, so a retargeted specifier still refuses), keeping lines and justification, and leaves everything else for review. Fixtures: test/fixtures/guards/check-function-size.ts/{bad,good}; every rule is driven in test/scripts/check-function-size.test.ts.

SyncRun state guard

scripts/check-sync-run-state.ts (bun run check:sync-run-state, in bun run verify, well under a second) protects the refactor wave 1 SyncRun rule (A17). SyncRun (src/commands/sync/sync-run.ts) holds the state one incremental sync shares between closures that interleave across awaits: the checkpoint flush and its cadence, the import workers, the stall watchdog and the partial exit. Its mutable fields are the members of interface SyncRun not marked readonly. Over src/commands/sync/**/*.ts the guard fails when a mutable field is destructured from a SyncRun value (const { bankedFiles } = run, or a { checkpointDead }: SyncRun parameter) or copied into a local (const banked = run.bankedFiles), because such a copy goes stale at the next await. A SyncRun value is a binding named run, annotated SyncRun, or initialized from createSyncRun(). Readonly fields (collection references, fixed configuration) may be destructured. Fields tagged @checkpoint in their JSDoc (the flush cadence, banked count, single-flight flag, dead flag, SIGTERM deregistration and yield counter) have one owner: only functions in sync-run.ts may assign them, so the flush, the SIGTERM hook and partial() cannot disagree about checkpoint state. Each failure prints FAIL: <file:line> plus Why: / Fix: / See:; the fix is to use run.<field> at each read and write, and to change checkpoint state through a sync-run.ts function. Fixtures: test/fixtures/guards/check-sync-run-state.ts/{good,bad,bad-alias,bad-param,bad-owner}.

Source reads in tests

test/test-reads-source-smell.test.ts finds test code that reads src/ text: readFileSync, readFile (including fs.promises.readFile) and Bun.file calls whose arguments name a src/ literal, a 'src' path segment, or a constant holding such a path. Each read site needs a tagged marker on its line or within the three lines above:

// test-reads-source-ok[structural]: <why a source read is the right tool>

The category is one of prompt-byte, trust-boundary, generated-artifact, structural or raw-bytes, and every marker must carry one. Files that predate the rule are ratcheted by their exact count of unjustified read sites, so a new untagged read in such a file fails and a count that drops must be lowered. The ratchet counts read sites only: a new assertion over an existing source binding is not detected and remains the authoring gate's job. Rerun with bun test test/test-reads-source-smell.test.ts.

Structural guards over the files that refactor wave 1 decomposes read them through test/helpers/source-surface.ts rather than readFileSync. A surface is one façade plus the modules it is split into (sync, cli, serve-http, jobs, hybrid, autopilot, migrate, pglite-engine, postgres-engine, doctor). surfaceSource(surface) concatenates the surface with file boundary markers and serves containment assertions (toContain, not.toContain, single-line regexes). surfaceFileSource(surface, path) returns one named file and serves positional assertions (indexOf ordering, slice windows, [\s\S] spans, line math); a file outside the surface throws. A lane that moves code adds the destination to the surface in the same commit: a new directory is globbed automatically, while a module in an existing directory or flat file set is listed explicitly so today's assertions are not widened. test/helpers/doctor-source.ts is the doctor instance of the same loaders.

Registry-walking ratchets

Structural suites that walk a registry so the NEXT gap of a known class cannot ship silently. All allowlists below are shrink-only unless noted.

  • test/operations-coverage-ledger.test.ts — every op in src/core/operations.ts maps to a covering test file in a checked-in ledger; the UNCOVERED allowlist only shrinks. Shares one registry-enumeration helper (test/helpers/ops-registry.ts) with the jobs-ops token-redaction sweep so two walkers can't drift.
  • test/operations-source-isolation-matrix.test.ts — every non-localOnly read op runs under a scoped remote ctx and a federated grant; nothing carrying the other source's identity may return. Deliberate brain-wide behavior requires an explicit BRAIN_WIDE_READS entry with a rationale string. Anti-vacuity is mandatory: each op's control call must SEE the cross-source marker before its scoped assertions count; an op that can't be driven is an explicit counted SKIP disposition, never a silent pass.
  • test/scripts/e2e-wiring.test.ts — every test/e2e/*.test.ts must be claimed by a PR-time lane (a scripts/e2e-test-map.ts row, a workflow mention, or the shrink-only test/fixtures/e2e-unmapped-baseline.txt), and every map entry must point at a real file (typo guard).
  • test/engine-surface-coverage.test.ts — two-way census of the BrainEngine interface against the PGLite prototype (new methods force a visible list edit) plus a runtime UNCALLED ratchet scanning the whole test corpus for references, so a never-called engine method can't ship.
  • scripts/check-orphan-modules.mjs (verify battery, guard-manifest registered with bad/good fixtures) — transitive import walk from the cli/mcp/engine entrypoints; see Orphan-module guard.

Orphan-module guard

bun run check:orphan-modules walks static, dynamic and require relative imports from the runtime entrypoints (CLI, MCP server, plugin engines, admin, package exports). Every src/ module it cannot reach needs a disposition:

  • Imported by nothing, not even tests: fails as hard-orphan unless it has a reasoned ALLOWLIST entry (shrink-only).
  • Imported only by tests (or scripts): fails as unpermitted-test-only unless it is named in PERMITTED_TEST_ONLY with a reason. The set may grow only with a reason in a reviewer-visible edit; modules reached from scripts/** use reason script-reachable, which the guard verifies.
  • A permitted entry whose module was deleted, wired into a runtime entrypoint, dropped by every test, or tagged script-reachable without a scripts/** importer fails as stale-permitted-entry. Remove or correct the record; never restore code to satisfy the list.

Each failure prints the rule, the module, the tests that import it, the reason, the remedy, the rerun command and this anchor. Fixture mode (GBRAIN_GUARD_ROOT) reads the permitted set from <root>/permitted-test-only.json; test/scripts/check-orphan-modules.test.ts proves every rule fails on a bad tree.

The takes-bootstrap graduation instrument (evals/takes-bootstrap/: 123-case corpus, scorer, live harness + $0 replay) is CI-guarded keyless by test/eval-takes-bootstrap.test.ts — the guard proves the instrument, not the score; the autopilot tier flips only on a committed GRADUATED live run.

Shell dispatch and Windows

All four of test, verify, ci:local and test:e2e hand off to shell scripts under scripts/, so every check:* entry in package.json invokes its script as bash scripts/<name>.sh instead of relying on the shebang — bun on Windows cannot exec a .sh directly. Add a new shell-script check with that same prefix. The scripts/*.ts entries run under bun and take no prefix.

The scripts must also be on disk with Unix line endings. A strict bash (WSL, Linux CI, macOS) rejects CRLF and dies on the script's first meaningful line; the Cygwin bash that ships with Git for Windows tolerates it, so a green local run is not by itself evidence that a script is CRLF-clean. The root .gitattributes pins *.sh text eol=lf, which overrides the core.autocrlf=true default that Git for Windows installs. It pins *.md the same way, because the frontmatter readers anchor on a --- fence followed by a Unix line ending and a CRLF checkout makes a document parse as having no frontmatter, silently. Working copies cloned before those pins need a one-time git rm --cached -r . -q && git reset --hard to pick them up; see the Windows section of CONTRIBUTING.md.

Wallclock figures in the table above are from a Mac dev box. Windows is substantially slower because each check pays full process-creation cost, and three tree-walking checks (check:privacy, check:test-names, check:test-isolation) plus typecheck can exceed the 120s per-check cap in run-verify-parallel.sh there even though they pass on Linux and macOS.

CI vs local: intentionally divergent file sets

  • CI matrix (.github/workflows/test.yml) runs scripts/test-shard.sh across 8 matrix shards partitioned by weight-aware LPT bin-packing (scripts/sharding.ts; files with no mined weight fall back to the p75 file weight so a new unweighted file can't silently unbalance a shard) and INCLUDES *.slow.test.ts (the dedicated slow files — longmemeval, entity-resolve-perf, entity-card-perf, export-scale, brainbench-e2e — run as dedicated jobs alongside the matrix, and reconcile-crash.slow.test.ts runs only in the persistence-validation reconciliation job) plus evals/**/*.test.ts (keyless-allowlist-gated — test/scripts/evals-collection.test.ts). Each shard's bun process is bounded by --max-concurrency (GBRAIN_TEST_MAX_CONCURRENCY, default 4). Every bun-test job — matrix shards, serial-tests, verify, the slow/eval jobs — activates the PGLite schema snapshot (built in-runner via scripts/lib/test-env.sh; the BrainBench gate uses the separate default-profile snapshot for its in-memory PGLite; the ~42MB tar is also cached across jobs via actions/cache, with the runner's own hash check staying authoritative). CI EXCLUDES *.serial.test.ts from the shards and runs them across four serial-tests workers via bun run test:serial — one bun process per file preserves the mock.module quarantine; the pool runs those processes concurrently. bun run verify gets its own job too, as does the BrainBench memory-conformance gate (brainbench job → scripts/ci-brainbench-gate.sh, hermetic in-memory PGLite, ~15s), which compares HEAD's fresh run against master's committed baseline (evals/brainbench/baselines/main.json) — the test-status aggregate checks its result explicitly. E2E (.github/workflows/e2e.yml) always runs its applicable execution lanes, with the jsonb-parity job in front of tier2 as the token-spend gate, and aggregates through e2e-status. Scheduled runs also require the full-corpus lanes, including each slow suite excluded from the coverage shards (longmemeval, entity-resolve-perf, entity-card-perf, brainbench-e2e, export-scale and reconcile-crash). Both aggregates reject failures, cancellations, and unexpected skips. Dependency caches and validated PGLite snapshots remain; successful test results are never reused. CI is the ground truth for "did everything pass."
  • Local fast loop (scripts/run-unit-shard.sh via the parallel wrapper) uses the same weighted partitioner as CI and EXCLUDES *.slow.test.ts AND *.serial.test.ts. Each shard runs its complete ordered selection with a fresh Bun process per file, without adding workers. Later groups still run after failures; missing summaries or file-completion evidence fail the shard. Local trades coverage for inner-loop speed; CI catches what local skips.

This divergence is intentional. Don't try to make them equal — the two scripts deliberately solve different problems. The regression test at test/scripts/run-unit-shard.test.ts pins what the local fast loop should and shouldn't include, and that no unit-lane file spawning the CLI through test/helpers/cli-spawn.ts hand-pins a per-test timeout below the bunfig default (an explicit test(name, fn, N) ceiling overrides bun's --timeout, so GBRAIN_TEST_TIMEOUT_MULTIPLIER never reaches it — inherit the default instead; cli-spawn's own kill timer still reaps a hung child); test/scripts/run-unit-parallel.test.ts pins the wrapper's memory-adaptive concurrency, and the OOM/external-kill serial rescue pass, and operator-interrupt teardown (a Ctrl-C / SIGTERM to the wrapper while shards are live TERMs then KILLs every shard descendant, so a cancelled run cannot leave gtimeout/bun alive until the shard cap).

Coverage lanes and gates

Line coverage is opt-in via COVERAGE_DIR: when set, the shell lanes (scripts/test-shard.sh, scripts/run-serial-tests.sh, scripts/run-e2e.sh) pass --coverage --coverage-reporter=lcov to bun; when unset, the exec line is byte-identical to a non-coverage run. Every bun process gets its OWN coverage dir ($COVERAGE_DIR/shard, serial-$idx, e2e-$idx) because a reused dir silently overwrites lcov.info — the shard runner also pins xargs to a single batch (-n 100000 -x) so an argv overflow fails loud instead of spawning a second, overwriting bun process. On a green run each lane writes $COVERAGE_DIR/lane-manifest.json ({lane, sha, lcovCount, complete}); a red run writes no manifest, which downstream merging treats as an incomplete lane. E2E shards use distinct e2e-1 through e2e-4 lane names (unsharded runs use e2e), with an executed-files.txt receipt written only after every selected file completes its native Bun report successfully. The runner requires a fresh parent-owned JUnit report for the selected file and matching final console pass/fail/skip totals. Python 3's standard XML parser validates the document and checks its actual testcase counts against each suite and the console; a zero exit without that report cannot borrow a nested child's summary as completion evidence. Each coverage invocation atomically creates its own COVERAGE_DIR; an existing destination, even an empty one, is refused without changing its contents. Use a new path for a rerun. Failed or cancelled runs write no completion receipt, and shorter or skip-only runs cannot inherit old coverage or delete another run's outputs. Skip-only files can emit no LCOV, so execution-file counts and LCOV counts are intentionally different measures. run-e2e.sh specifics: COVERAGE_DIR is normalized to an absolute path against the repo root before HOME moves (the script redirects HOME/GBRAIN_HOME and E2E tests spawn CLI subprocesses with varying cwd — an un-normalized relative dir would scatter output), and E2E_FILE_TIMEOUT_SECS caps each file's wallclock (default 180s; the nightly coverage lane uses 300s for instrumentation overhead). Both env names are deliberately non-GBRAIN_-prefixed so the hermetic env scrub keeps them.

Two corpora.

  • PR corpus (prCorpus) — the 15 coverage-collecting lanes in .github/workflows/test.yml: the 8 matrix shards, four serial-tests partitions, and the three dedicated slow jobs (slow-eval-longmemeval, slow-entity-resolve-perf, slow-brainbench-e2e). Deterministic (runs identically on every PR); this is the corpus the gates run against.
  • fullCorpus — nightly or explicit manual opt-in in .github/workflows/e2e.yml: coverage-full-{unit,serial,slow,e2e} + coverage-full-report. Fully self-contained (every lane re-runs with coverage inside that workflow, including the full default E2E discovery across four isolated Postgres workers) — the honest merged unit+serial+slow+e2e number, kept as the coverage-full-merged trend artifact.

The E2E workflow's manual full_corpus boolean defaults to false. Setting it to true executes the same full unit/serial/slow/E2E profile, report and receipt checks as a schedule; selected E2E receives the same explicit empty sentinel. Ordinary pushes, PRs and default manual runs retain their existing selection. Explicit full manual runs have a separate concurrency group, so they do not cancel ordinary validation on the same branch. To measure a branch before the next scheduled run, dispatch e2e.yml on that branch with full_corpus=true.

The nightly E2E artifacts are coverage-full-e2e-1 through -4, with one manifest per artifact and one coverage directory per Bun process. Lightweight e2e-full-execution-* artifacts also carry the manifest and executed-file list. Full-profile e2e-status requires the matrix job to succeed and validates all four same-commit receipts against the exact expected partitions using scripts/verify-nightly-e2e.ts. Missing artifacts, duplicate identities, wrong commits, omitted or repeated files, failures and cancellations cannot report complete execution. The report independently verifies these receipts before merging; missing execution evidence prevents publishing a full-corpus report. Coverage percentages remain advisory and the report job itself is not an e2e-status dependency. The receipts prove file execution, not execution of every optional assertion within a file.

Merge (scripts/merge-lcov.ts). Walks the input dirs for lcov.info + lane-manifest.json, sums DA hits per file:line, normalizes paths repo-relative, and emits a merged lcov plus a summary JSON: src-only totals/per-dir/per-file percentages, a lineHits map (the diff gate's input), and the never-loaded src file list. --manifest-expect lane,lane,... pins the expected lane set (serial-1 through serial-4 for PRs, serial nightly, and e2e-1 through e2e-4 nightly). The merger checks commit SHA (--sha overrides checkout HEAD for offline artifacts), duplicate identities, and actual per-lane LCOV counts; a missing or complete: false manifest, an unparseable lcov, or a shard lane with lcovCount != 1 marks the summary degraded: true. Degraded is data, not failure: the merge never aborts (exit 0), and both gates print WOULD PASS/WOULD FAIL and exit 0 on a degraded summary instead of enforcing against partial data.

Diff gate (scripts/coverage-diff-gate.ts). Gates the added/changed lines of git diff origin/master...HEAD restricted to gate scope (src/**.ts minus *.test.ts/*.generated.ts/*.d.ts): covered/(covered+uncovered) must be ≥ 80%, AND no gate-scoped changed file may be entirely absent from the coverage data (a never-loaded file is one violation — add a test that imports it). Non-executable lines (no lcov record) don't count against you; empty and doc-only diffs short-circuit to PASS via the select-e2e classifier. Escape hatches: a commit body containing [coverage-exempt: reason] passes with a loud warning, and scripts/coverage-gate-exemptions.txt (exact path or trailing-/ prefix per line; resolved via git show origin/master:scripts/coverage-gate-exemptions.txt, never the working tree, so a PR cannot self-exempt; SHRINK-ONLY — additions need a graduation review in the PR description) excludes paths from the gate while still reporting them ([e2e-exempt], [subprocess-undercount]). Report-only unless COVERAGE_GATE_ENFORCE=1. Exit contract: 0 = pass or report-only, 1 = gate fail while enforcing, 2 = infrastructure error (missing summary, git failure — never conflated with a coverage verdict).

Baseline gate (scripts/coverage-baseline-gate.ts). Anti-erosion floor: reads the baseline via git show origin/master:scripts/coverage-baseline.json — the master copy, never the working tree, so a PR cannot weaken its own bar — and compares like-for-like by corpus (--corpus prCorpus in test.yml, --corpus fullCorpus nightly). A global drop > 0.5pp, a per-directory drop

1.0pp, or a never-loaded-count increase fails (deleting tests shrinks the coverage denominator, which inflates pct for free); a corpus section that is null on master is an ungated first landing. provisional: true in the baseline keeps the gate report-only regardless of enforcement — the committed baseline is currently provisional with both corpus sections unseeded. scripts/update-coverage-baseline.ts --summary <json> --corpus <c> [--promote] writes the working-tree baseline (per-file detail limited to the baseline's watchlist); --promote flips provisional: false at graduation.

CI wiring. The 17 PR lanes upload coverage-* artifacts; the advisory coverage-report job downloads + merges (COVERAGE_CORPUS=prCorpus), renders scripts/render-coverage-summary.ts to the step summary (including the behavioral-vs-structural counts from scripts/structural-suites.tsv), and runs both gates with COVERAGE_GATE_ENFORCE: '0'. It is deliberately NOT in test-status needs — it cannot block a PR until graduation. Test results are never cached: every run executes its required checks, while Bun dependencies and validated PGLite snapshots remain cached. Gitleaks runs independently for all changes, including documentation. e2e-status also requires the four coverage-full-* execution lanes on scheduled runs; only coverage percentage reporting remains advisory.

Bun caveats. Bun/JSC emits line records only, so function coverage is informational (no reliable function names). There is NO subprocess coverage: code exercised only through spawned CLI subprocesses undercounts — src/cli.ts carries a permanent [subprocess-undercount] exemption for this. A src file never imported by any test produces no lcov record at all; the summary reports these as a count + sorted list, deliberately never a percentage (physical lines ≠ executable lines), and the diff gate treats a changed-but-never-loaded file as a violation.

One-command local smoke (one shard of ten, so totals reflect a tenth of the corpus — this checks the plumbing, not the number):

COVERAGE_DIR=$PWD/.coverage bash scripts/test-shard.sh 1 10 \
  && bun scripts/merge-lcov.ts --out-lcov .coverage/merged.lcov --out-json .coverage/summary.json .coverage \
  && bun scripts/render-coverage-summary.ts --summary .coverage/summary.json

Optional flags: coverage-diff-gate.ts --base <ref> overrides the diff base (default origin/master); render-coverage-summary.ts --structural scripts/structural-suites.tsv adds the behavioral-vs-structural split to the rendered summary (both CI lanes pass it); classify-tests.ts --summary prints counts only.

Failure-first logging

When bun run test finds any failure, the wrapper:

  1. Writes failure blocks (each prefixed with --- shard N: <test name> ---) to .context/test-failures.log (workspace-local, gitignored). On systems without a writable .context/, falls back to /tmp/gbrain-test-failures.log.
  2. Prints a loud stderr banner with the absolute log path, plus the last 30 lines of the failure log inlined. Banner survives | head / | tail / agent-side log truncation.
  3. Writes a one-line-per-shard summary to .context/test-summary.txt (shard N/M: pass=X fail=Y skip=Z rc=W).
  4. Exits non-zero. Empty failure log + non-zero exit = infrastructure problem (wedged shard, killed child); the banner says so.

If a shard hits the per-shard GBRAIN_TEST_SHARD_TIMEOUT cap (default 3000s — sized so the heaviest count-balanced shard finishes under 4-way contention; GBRAIN_TEST_SHARD_KILL_AFTER sets the grace after TERM before KILL, default 30s), the wrapper classifies the kill one of two ways:

  • EXIT-HANG → warn-pass. If the shard's log had been silent for ≥300s at kill time AND shows zero (fail) markers, the shard finished all its work, leaked a handle, and never exited (a known PGLite-adjacent handle leak — see TODOS.md "unit-shard exit hang"). The wrapper prints a ⚠️ shard N/M: EXIT-HANG ... Treating as pass-with-warning banner, writes EXIT-HANG (idle Ns, 0 fails) ... warn-pass to the summary, and does NOT fail the run. Its pass counts are undercounted (bun never printed its final summary). Bun's per-test --timeout turns a genuinely hung TEST into a printed (fail) — new output — so this classification cannot mask a hung test; the residual maskable case is a file-level import hang in the very last file, which the banner keeps visible.
  • WEDGED → hard failure. Anything else (failures present, or the log was still growing) writes --- shard N: WEDGED after ${SHARD_TIMEOUT}s --- to the failure log with the last 50 lines of the shard log, marks the run failed, and proceeds with other shards' results.

Triage rule: a warn-pass EXIT-HANG line in .context/test-summary.txt is NOT a test failure — don't burn time bisecting it; a WEDGED line is.

File taxonomy

  • *.test.ts → fast loop (parallel up-to-4-shard fan-out, memory-adaptive).
  • *.slow.test.ts → run via bun run test:slow only (intentional cold-path tests; would dominate the fast loop's wallclock).
  • *.serial.test.ts → run via bun run test:serial after the parallel pass completes; one bun process per file (--max-concurrency=1 within a shared process is not enough — the module registry still leaks mock.module), with those per-file processes POOLED (per-process isolation never required one-at-a-time execution). Files touching machine-global state (launchd/cron) live on the sequential EXCLUSIVE_FILES lane inside scripts/run-serial-tests.sh — growth-guarded to ≤3 entries with justification comments. Quarantine for tests that share file-wide state and race when run alongside other files in the same bun test process. Several dozen files, discovered by the *.serial.test.ts glob — no list to maintain. Typical residents: mock.module(...) users (top-level mocks leak across files in a shard process, e.g. test/embed.serial.test.ts), env-coupled files (e.g. test/brain-registry.serial.test.ts), and process-lifecycle suites that assert on process.exitCode (e.g. test/pglite-engine-disconnect.serial.test.ts). Do not put the parallelism back on a serial file unless you've fixed the contention root cause (it just re-introduces the flake).
  • test/e2e/*.test.ts → real-Postgres E2E. Skipped when DATABASE_URL is unset. One out-of-directory file rides this lane: test/phantom-redirect-engine-parity.test.ts (lives in test/ for its PGLite arm, but its Postgres arm is only reachable through a DATABASE_URL-bearing lane — the unit wrappers strip the URL, so run-e2e.sh's no-args list and CI's parity job carry it). run-e2e.sh wraps each file in a hard outer timeout (default 180s; GBRAIN_E2E_FILE_TIMEOUT=<seconds> overrides) because a synchronously-blocking PGLite WASM call can outlive bun's timer-based --timeout; LLM-bound Tier-2 files (skills.test.ts) automatically get 4× the cap since real provider round-trips legitimately run past 180s.
  • tests/heavy/*.sh → ops-shape shell scripts. Cost minutes per run; NOT in default bun test. Run via bun run test:heavy or scheduled nightly via .github/workflows/heavy-tests.yml. Examples: pg_upgrade matrix (boot legacy brain → walk to head), RSS budget gate (measure peak worker RSS vs committed baseline), read-latency-under-sync (p50/p95/p99 under concurrent writer load), sync lock regression (N concurrent syncs assert 1 winner + N-1 lock-busy + zero leaked gbrain_cycle_locks rows). See tests/heavy/README.md for when to add a script here vs *.slow.test.ts. Files prefixed with _ (e.g. tests/heavy/_build_legacy_fixtures.sh) are helpers/libs invoked by sibling tests — the runner skips them.
  • test/fuzz/*.test.ts → property-based fuzz harness. Pure-validator targets in pure-validators.test.ts are guarded by scripts/check-fuzz-purity.sh (in bun run verify), which bun build --target=bun bundles each target and greps the resulting bundle for banned transitive imports (node:fs, node:child_process, engine modules). Anything that fails the guard moves to mixed-validators.test.ts (still property-tested, but no purity guarantee) or filesystem-validators.test.ts (fs-backed, uses temp dirs). Fuzz tests run in the default bun test loop because they're fast (~3s for ~12 properties × 1000 runs each).

The taxonomy above is LANE-based (where a test runs). A second, orthogonal axis is INTENT:

  • Behavioral tests execute product code and assert on behavior — the default.
  • Structural (source-shape) suites read repo source/doc TEXT and assert on its shape (wiring guards, drift pins, doctorSource() consumers). They are real invariants but execute no product paths, so they inflate the headline test count without adding line coverage. The committed inventory is scripts/structural-suites.tsv, generated by scripts/classify-tests.ts (suite-level, content-based detectors: repo-anchored readFileSync/Bun.file readers, grep-style exec scanners, the doctor-source helpers) and freshness-checked in bun run verify (check:structural-manifest — regenerate with bun scripts/classify-tests.ts when suites change shape). The inventory is approximate by design; fix misclassifications in the classifier's detector list, never by hand-editing the TSV. CI's coverage report renders behavioral vs structural counts side by side.

Guards that pin doctor source text read it through test/helpers/doctor-source.ts (doctorSource() = the façade + every src/commands/doctor/** module, for containment assertions; doctorFileSource(rel) = one named file, for positional/ordering assertions) so peeling doctor.ts into modules can't silently move a pinned string out of a guard's sight.

TTY and interactive-CLI testing

Four escalating tools; reach for the cheapest one that answers the question:

Question Tool Example
Does the TTY/non-TTY branch logic pick right? Inject isTTY into the pure function — no subprocess test/init-provider-picker.test.ts, test/jobs-watch-mode.test.ts
Does the real CLI behave right when stdin is NOT a terminal? Spawn the CLI with piped/ignored stdio test/cli-stdin-hang.test.ts (fast loop); test/init-fresh-pglite.slow.test.ts (slow lane)
Does the real CLI render menus and read typed input under a REAL terminal? launchTty from test/helpers/tty-harness.ts in a *.serial.test.ts file test/init-picker-pty.serial.test.ts
How does the install FEEL (stalls, copy, silence windows)? scripts/dx-explore.ts — instrument, not a test; nothing asserts transcripts under .context/dx-runs/ (see docs/guides/bootstrap.md)

Real-PTY test rules: put the file in the serial lane (*.serial.test.ts — that lane runs in required CI; a new test/e2e/* file does NOT, since unit shards exclude the directory and the e2e workflow runs only explicitly named files, no glob); assert NON-default picker values (bare Enter and each prompt's 60s readLineSafe timeout both resolve to the default, so a defaults-asserting test passes with dead input); always await session.close() in a finally (only close() clears the harness wall timer); and point HOME plus GBRAIN_HOME at a temp root with pass-through auth keys stripped via dropEnv so picker state is machine-independent.

Skills-manifest freshness guard

skills/skills.lock.json is a committed sha256 inventory of every bundled file under skills/ (tamper evidence, not signatures — see src/core/skills-integrity.ts). Any change under skills/ must regenerate it: bun run scripts/generate-skills-manifest.ts. scripts/check-skills-manifest-fresh.sh (bun run check:skills-manifest, wired into bun run verify) regenerates to a tmp file and diffs, failing CI on drift; at runtime gbrain doctor reports the same drift as a warn-only skills_manifest_integrity check.

Docs CLI truth check

test/docs-cli-commands.test.ts checks every gbrain <verb> in code fences and inline code across README, docs and skills against the registered verbs. In docs/guides/, docs/migrations/ and skills/ it also runs each invocation's flags through the CLI's own validator, via test/helpers/cli-command-surface.ts. When a hit is stale, fix the doc. When the example documents an older release, put <!-- gbrain-cli: historical --> on the line above its code fence, or on the line with the inline code. The test's ALLOWLIST is a last resort: it only shrinks, every entry needs a reason, and stale entries fail.

Test-isolation lint and helpers

This section is the canonical home of the test-isolation discipline — CONTRIBUTING.md and other docs link here rather than restating the rules.

The cross-file flake class is enforced statically by scripts/check-test-isolation.sh, wired into bun run verify. Rules (non-serial unit files only; *.serial.test.ts and test/e2e/* are skipped):

Rule What it bans Fix
R1 process.env.X = ..., bracket assignment, delete process.env.X, Object.assign(process.env, ...), Reflect.set(process.env, ...) Use withEnv() from test/helpers/with-env.ts, OR rename file to *.serial.test.ts
R2 mock.module(...) anywhere in the file Rename file to *.serial.test.ts (no DI on production code for testability)
R3 new PGLiteEngine( outside ~50 lines after a beforeAll( line Use the canonical block (below) inside beforeAll(
R4 Files creating new PGLiteEngine( without engine.disconnect( inside an afterAll( block Add afterAll(() => engine.disconnect())

Files that violated these rules at the isolation-lint baseline are listed in scripts/check-test-isolation.allowlist. The allow-list MUST shrink over time — never add new entries.

Canonical PGLite block (R3 + R4 compliant)

Every test file that needs a PGLite engine should use this exact pattern:

import { PGLiteEngine } from '../src/core/pglite-engine.ts';
import { resetPgliteState } from './helpers/reset-pglite.ts';

let engine: PGLiteEngine;

beforeAll(async () => {
  engine = new PGLiteEngine();
  await engine.connect({});
  await engine.initSchema();
});

afterAll(async () => {
  await engine.disconnect();
});

beforeEach(async () => {
  await resetPgliteState(engine);
});

Why this exact shape: beforeAll creates a single engine per file (PGLite WASM cold-start + initSchema is ~20s); beforeEach clears user data via resetPgliteState; afterAll disconnects so the engine doesn't leak across file boundaries within a shard process. Ordinary resets atomically delete rows with cleanup-only trigger suppression and restart owned sequences, retaining table/index storage. The helper restores trigger behavior before reseeding and falls back to TRUNCATE CASCADE for schemas whose triggers, rules, inheritance, external foreign keys or privileges require its original semantics. Schema/generation infrastructure survives, and each reset rotates the logical brain identity.

Every full reset measures aggregate target-table storage, including indexes and TOAST, with pg_total_relation_size. Above 8 MiB it uses the same atomic TRUNCATE path to reclaim storage; no reset counter or stale size estimate is retained. The helper regression suite checks repeated TOAST-heavy resets, cleanup and sequence parity, restored triggers and foreign-key enforcement.

withEnv pattern (R1 fix)

import { withEnv } from './helpers/with-env.ts';

test('reads OPENAI_API_KEY', async () => {
  await withEnv({ OPENAI_API_KEY: 'sk-test' }, async () => {
    expect(loadConfig().openai_key).toBe('sk-test');
  });
});

// Delete a var (override is undefined):
await withEnv({ GBRAIN_HOME: undefined }, fn);

// Multiple keys:
await withEnv({ A: '1', B: '2', C: undefined }, fn);

withEnv saves the prior value of every key it touches and restores via try/finally — including when the callback throws. An absent TZ is restored as the zone that was in effect, because Bun keeps the last explicitly set zone when TZ is deleted; a TZ override therefore never leaks into later files in the same process. It is cross-test safe but NOT intra-file concurrent-safe. process.env is process-global; two test.concurrent() calls in the same file both touching the same key will race. Files using withEnv stay outside the test.concurrent() codemod's eligibility filter.

Speed + environment helpers (test/helpers/)

Reach for these before hand-rolling; the five speed helpers each have their own unit test, and the two environment probes are exercised through their consumer suites:

  • cli-spawn.ts — runCli(argv, opts) (async, hermetic env, timeout-killed), runCliBatch(argvs, {width}) (bounded pool, DEFAULT WIDTH 2 — the cap is per-invocation and 4 shards × width multiplies CLI children machine-wide; each child can boot a ~1.5GB PGLite), runCliMemo (argv-keyed memo for read-only calls like --help; clearCliMemo() drops the memo when a test mutates what a memoized call would observe). Replaces the per-file spawn wrappers; a file of N independent sequential spawns becomes one width-2 batch in beforeAll.
  • wait-for.ts — waitFor(predicate, {timeoutMs, intervalMs}) / waitForValue. Replaces fixed setTimeout sleeps: polls resolve as soon as the condition holds, and generous deadlines make slow-CI runs LESS flaky than a tuned sleep, not more.
  • with-snapshot.ts — withColdPglite(fn): per-TEST scoped GBRAIN_PGLITE_SNAPSHOT opt-out (save/delete/restore); withSnapshotValue(value, fn) is the general form (pin any snapshot path for fn's scope; undefined = deleted). Use instead of a file-level delete process.env.GBRAIN_PGLITE_SNAPSHOT, which forces every engine in the file to cold-boot. Caution: a snapshot-restored engine does not replay migrations on a later initSchema() after a version rewind — rewind-arc tests need the cold path (see test/bootstrap.test.ts).
  • reset-pglite.ts#resetPgliteStateNarrow(engine, tables) — explicit-table truncate for hot loops (the full reset clears the whole catalog). The table list is REQUIRED — a default would silently under-truncate.
  • git-fixture.ts — makeGitFixture(dir): build-once git repo + reset()/commitAll() between tests, replacing per-test git init chains.
  • fs-perms.ts — permsEnforced() / crontabAvailable() probes: some hosts (FUSE/overlay sandboxes, root) don't enforce permission bits or lack a crontab; tests asserting "this write MUST fail" / "cron registered" use test.skipIf(!probe()) so they skip visibly there and still run in CI.
  • git-stderr-probe.ts — gitStderrLeads(): skips raw-git-stderr-slice assertions behind ambient git PATH shims that print their own diagnostics first (e.g. Conductor's auth-broker wrapper).

When to quarantine instead of fix

Rename to *.serial.test.ts when:

  • The file uses mock.module(...) (R2 — there's no clean fix without changing production code).
  • The file is genuinely env-coupled (e.g. gbrain-home-isolation.test.ts, claw-test-cli.test.ts) — module-load env readers + ESM caching defeat dynamic-import-after-env tricks.
  • The file's tests intentionally share state across it() boundaries.

The quarantine has grown to dozens of files — treat it as debt: every addition needs a reason from the list above, and prefer fixing the contention root cause when one exists.

Unit test inventory

bun test runs all tests without a database. E2E tests skip gracefully when DATABASE_URL is not set.

GBRAIN_HOME isolation preload. test/helpers/gbrain-home-preload.ts (bunfig [test] preload) points GBRAIN_HOME at a per-run scratch dir when it isn't already set, so unit tests never read — or clobber — the operator's real ~/.gbrain config/brain. Without it, any config-honoring code path silently changes behavior with whatever the live config.json says. The canonical GBRAIN_HOME convention is config.ts:configDir(): GBRAIN_HOME is a PARENT dir and .gbrain is appended. Subprocess-spawning tests must set BOTH HOME: tmp and GBRAIN_HOME: tmp in the child env (HOME alone loses to the inherited preload value; in-process HOME mutation loses to Bun's cached os.homedir()). The e2e wrapper sets its own GBRAIN_HOME before bun starts, which this preload respects. Because the preload respects a pre-set value, the unit/slow wrappers (run-unit-parallel.sh / run-unit-shard.sh / run-slow-tests.sh) strip an ambient GBRAIN_HOME at their boundary — same discipline as the database-URL vars — so a dev shell configured for a real brain can't ride through. GBRAIN_DEBUG_PRELOAD=1 prints the allocated scratch home for debugging.

Installer fixtures must never delete GBRAIN_HOME to test a fallback against the operator's home. Spawn a disposable child with HOME set before Bun starts, then set GBRAIN_HOME to the specific fixture. real-home-guard-preload.ts compares metadata for the real-home autopilot wrapper, env file, start script, launchd plist and systemd unit around tests. It detects changes rather than intercepting writes and never reads env-file contents. A deliberate one-shot installer test can explicitly set GBRAIN_TEST_ALLOW_REAL_HOME_WRITES=1, which prints a warning; use that only inside an independently isolated child home. test/real-home-guard-preload.test.ts runs the installer suite with fake-live sentinels and verifies they are untouched.

Provider-key strip preload. test/helpers/provider-keys-preload.ts (bunfig [test] preload) strips the ambient provider credentials the canonical fold recognizes, using the explicit test/helpers/provider-env.ts list checked against every recipe credential/endpoint plus compatibility aliases, and defaults GBRAIN_MODEL_DISCOVERY=off (respecting an explicit operator override), so key-aware model routing (resolveTierDefault) resolves identically to keyless CI and latest-model discovery never makes a real network call from a test. Without it, a chat key exported in the dev shell flips default-model assertions AND turns gated paths into live provider calls. Tests that want keys inject them explicitly (configureGateway({env}), withEnv, serial-file process.env) — the preload removes ambient shell state only, before any test file loads. The e2e wrapper (scripts/run-e2e.sh) opts back in at its boundary via GBRAIN_TEST_KEEP_PROVIDER_KEYS=1 — e2e is the lane where real keys are deliberate (live embed/parity tests skip-gate on them). The routing-only qm-provisioning, serve-http-surface-ceiling, serve-stdio-roundtrip and thin-client fixtures still strip provider state in every child, set both HOME and GBRAIN_HOME to their temporary brain, and pass Bun --no-env-file (including provisioned shell commands). They exercise routing without spending provider tokens even when a lane carries keys. Fixture-specific environment overrides apply last; unrelated credentials are preserved rather than removed with a broad key-name pattern.

Database-URL run guard. A bun test invocation REFUSES to start while DATABASE_URL or GBRAIN_DATABASE_URL is ambient in the environment, because some tests run destructive SQL against whatever those URLs point at (a bare bun test with ~/.gbrain/.env sourced would run them against a real brain). The guard is a bunfig [test] preload (test/helpers/database-url-guard-preload.ts); it hard-fails with instructions rather than silently unsetting (a silent unset would turn DATABASE_URL-gated e2e tests into green skips). The e2e wrappers (scripts/run-e2e.sh, the e2e/heavy workflows) opt in at their own boundary via GBRAIN_TEST_ALLOW_DATABASE_URL=1; the unit/slow wrappers instead strip both URL vars at their boundary (unit tests need no database), which keeps bun run test:full working with DATABASE_URL exported. Caveat: bun loads bunfig.toml from the invocation cwd, so the preload layer only applies to runs started at the repo root — the per-file name floor below is the layer that doesn't care about cwd. Two more layers apply after the opt-in: every test that runs destructive SQL on the ambient URL must call assertSafeE2eDatabaseUrl() (test/helpers/db-guard.ts — name floor: the database name must contain "test" as a segment, or be opted in via GBRAIN_E2E_ALLOW_DB) or carry an inline name floor the coverage gate recognizes (test/e2e/schema-drift.test.ts keeps its own looksLikeTestDb, deliberately different because it also accepts *_e2e), and test/db-guard-coverage.test.ts statically scans the suite and fails when a file connects to DATABASE_URL and runs destructive SQL unguarded. Local CI explicitly sets GBRAIN_TEST_DB=1 for its known Docker E2E databases, and run-e2e.sh preserves that opt-in so schema-drift can reset stale fixture schemas on service-name hosts. The hard test-database name floor still applies; this flag only relaxes the localhost requirement. The heavy shell lane gets the same floor outside bun: tests/heavy/_db_floor.sh (sourced by scripts/run-heavy.sh for the whole lane, and by each database-touching heavy script itself, since scripts are documented for direct invocation — the PGLite-based heavy scripts unset the URL instead) checks BOTH URL variables and strips query strings before extracting the database name, so a ?host=/tmp/test-sockets parameter can't smuggle a test-shaped segment past it.

Unit tests and what they cover:

  • test/facts-engine.test.ts / test/consolidate-valid-until.test.ts — facts-list filtering and consolidate correctness: unconsolidatedOnly is applied before the 100-row limit, so newer consolidated facts cannot permanently hide older pending facts; the phase regression seeds 100 consolidated rows plus three older pending rows and requires all three to progress.
  • test/markdown.test.ts — frontmatter parsing; splitBody sentinel precedence, horizontal-rule preservation, inferType wiki subtypes.
  • test/chunkers/recursive.test.ts — chunking.
  • test/parity.test.ts — operations contract parity.
  • test/cli.test.ts — CLI structure.
  • test/cli-finish-teardown.test.ts — the CLI teardown contract: computeTeardownDeadlineMs formula/floor/live-registry scaling + GBRAIN_TEARDOWN_DEADLINE_MS override (garbage/zero/negative values fall back to the formula); finishCliTeardown clean path (drain BEFORE disconnect, no exit, no warn), backstop on hung drain or disconnect (honors an errored op's exit code), throwing drain/disconnect warned + swallowed; the gbrain-owned verdict channel is immune to PGLite WASM process.exitCode writes; flushThenExit unit coverage with mocked streams (exits once after both stream callbacks, non-TTY aliveness grace, blocked-pipe guard, EPIPE-safe, GBRAIN_FLUSH_GRACE_MS override).
  • test/flush-then-exit-harness.test.ts — real spawned-Bun pipe semantics for flushThenExit (fixture: test/fixtures/flush-then-exit-harness.ts): a 4MB piped stdout payload arrives byte-complete with the exit code even with a late reader, small output survives exit with a concurrent reader, and the fence resolves promptly (wall time well under the guard + grace ceiling).
  • test/cli-should-force-exit.test.ts — shouldForceExitAfterMain daemon-survival gate: serve (stdio and --http) never force-exits, including with preceding global flags; op commands / empty / flag-only argv do; space-separated global-flag VALUES can't fake a command (--timeout 30s serve resolves to the serve daemon, not a 30s command).
  • test/cli-exit-verdict-pin.test.ts — structural class pin: greps src/ so the NEXT raw process.exitCode = write fails CI (a raw write bypasses the gbrain-owned verdict channel and gets silently zeroed by the deliberate flush-exit, which would make a FAIL path exit 0). Runtime variants live in test/cli-finish-teardown.test.ts; this is the review-time guard.
  • test/cli-pipe-truncation.test.ts — real-CLI pipe completeness, implementation-agnostic: the actual CLI run the way agents run it (piped stdout) produces complete, parseable, byte-stable --tools-json output and exits deliberately, well under the teardown backstop. Synthetic flush-mechanism coverage stays in test/flush-then-exit-harness.test.ts.
  • test/volunteer-context.test.ts — push-based context core, hermetic in-memory PGLite: parseWindow lenient user:/assistant: parsing, multi-turn window extraction, confidence-gated volunteering (arm confidences, multi-turn/newest-turn boosts, min_confidence gate, max-pages cap), slug-only suppression, privacy (rationales are deterministic templates; synopses pass the takes/facts fence), and the approximate usage-stats join.
  • test/watch-command.test.ts — gbrain watch push transport: streaming loop, rolling window, session dedupe, --json JSONL shape, channel: 'watch' event logging, clean EOF return. Hermetic PGLite + injected line/write deps (no subprocess, no real stdin).
  • test/watch-sigint.serial.test.ts — gbrain watch SIGINT lifecycle against a real spawned CLI subprocess with a tmpdir brain. SERIAL: parallel unit shards flake on concurrent subprocess spawns (same rationale as apply-migrations-pglite-spawn.serial.test.ts).
  • test/init-picker-pty.serial.test.ts — the interactive gbrain init pickers (embedding-provider + search-mode) driven under a REAL pseudo-terminal via launchTty: typed input lands (a NON-default mode choice verified by a follow-up non-TTY config read — bare Enter and the readLineSafe timeout both resolve to defaults, so a defaults-asserting test would pass with dead input), prompt-to-acknowledgement gaps bounded well under the fallback window, plus the Ctrl-D/EOF keyless fallback. On CI, missing PTY support fails loud instead of skipping. Hermetic: HOME + GBRAIN_HOME at a temp root, pass-through auth keys stripped via dropEnv; session.close() in finally. Serial: PTY spawn + full PGLite bootstrap, and the serial lane is what runs in required CI.
  • test/tty-harness.test.ts — the real-PTY harness's pure helpers (stripAnsi, computeStalls, renderStallsReport, parseDriveCommand, buildClaudeTuiSeed) with zero subprocesses; the file's live-PTY smokes are describe.skipIf(!ptySupported())-gated.
  • test/autopilot-launchd-lifecycle.serial.test.ts — autopilot lifecycle behavior, not generated-string assertions: the full install → self-disable → status → reinstall → uninstall arc with launchctl replaced by an argv recorder and the generated wrapper executed by a REAL bash against a genuinely deleted repo (every platform), plus a darwin-only fail-SKIP describe against the real launchd under a per-run unique label (GBRAIN_AUTOPILOT_LABEL) so it can never collide with — or tear down — a real install on the host. Serial: spawns subprocesses and pins HOME/GBRAIN_HOME for the whole file.
  • test/autopilot-fanout.test.ts — Autopilot fan-out and policy pins: targeted idempotency keys reopen per dispatch interval while stable doctor/remediate keys remain unchanged; the 60-minute full-cycle floor wins with a remaining small plan, and an all-fresh restart check advances the process-local clock without masking failed stale-source submissions.
  • test/agent-scheduler-contract.serial.test.ts — the documented external agent-scheduler shell chain (gbrain sync --repo X && gbrain embed --stale, live-sync.md / INSTALL_FOR_AGENTS.md Step 7) driven end-to-end through a real /bin/sh against a keyless PGLite brain: the && short-circuit IS the contract (argv arrays can't exercise it), the keyless bare stale embed exits 0, and the pull-failure case that must break the chain does. Anti-vacuity: the fixture commits a real page and every read-back asserts pages >= 1. Serial: real spawned CLI + tmpdir HOME.
  • test/cli-format-volunteer.test.ts — formatResult's volunteer_context human rendering: pointer lines with confidence/arm/rationale, the empty-result message, the approximate stats summary.
  • test/config.test.ts — config redaction.
  • test/files.test.ts — MIME/hash.
  • test/import-file.test.ts — import pipeline.
  • test/upgrade.serial.test.ts — the gbrain upgrade command via subprocess: --help prints usage and exits 0, install-method detection, resolveBunGlobalRoot, and the self-upgrade marker format (serial: spawns the real CLI).
  • test/file-migration.test.ts — file migration.
  • test/file-resolver.test.ts — file resolution.
  • test/import-resume.test.ts — import checkpoints.
  • test/migrate.test.ts — migration: v8/v9 helper-btree-index SQL structural assertions; 1000-row wall-clock fixtures pinning O(n log n) behavior; v12/v13 SQL shape; sqlFor + transaction:false runner semantics; the max_stalled DEFAULT 1 regression guard; v24 sqlFor.pglite: '' no-op assertion; v117 context_volunteer_events (named + idempotent entry, documented columns + both source-scoped indexes after initSchema, insert + 90-day purgeStaleVolunteerEvents round-trip).
  • test/bootstrap.test.ts — bootstrap contract: no-op on fresh install, idempotent across two initSchema() calls, no-op on modern brain that already has every probed column, full bootstrap path on a simulated legacy brain, fresh-install regression guard, legacy links shape coverage.
  • test/schema-bootstrap-coverage.test.ts — CI guard covering BOTH embedded schema blobs: neither may forward-reference state its engine's bootstrap can't create, and a reference covered on one blob is NOT automatically covered on the other (dream_verdicts exists only in the Postgres blob). PGLite half: REQUIRED_BOOTSTRAP_COVERAGE lists every forward reference in PGLITE_SCHEMA_SQL; the test fails loudly if applyForwardReferenceBootstrap skips one (extend both arrays when adding a column-with-index to the embedded schema blob). Also parses src/core/migrate.ts source text for every ALTER TABLE ... ADD COLUMN (top-level sql:, sqlFor.{postgres,pglite} overrides, AND handler-body engine.runMigration(N, \ALTER TABLE ...`)) and asserts each (table, column) pair is covered by the bootstrap OR by the schema blob's CREATE TABLE bodies — catching the column-only forward-reference class (e.g. sources.archived, oauth_clients.source_id) that a CREATE INDEX parser alone can't see. Postgres half (the class-closure gate): parses every CREATE INDEX column reference in SCHEMA_SQLand requires each to be in the blob's CREATE TABLE body AND not migration-added, or probed + ALTERed bysrc/core/engine-sql/bootstrap.ts(the single bootstrap both engines run; the PGLite half reads it minus itsdialect-only:postgresregions) — a column that is both in the blob's CREATE TABLE and migration-added is still a forward reference for pre-existing brains, whereCREATE TABLE IF NOT EXISTSno-ops and the blob's CREATE INDEX wedgesinitSchemabefore migrations can help. This gate is parser-driven (no registry to extend); intentional non-probes go inPOSTGRES_INDEX_REF_EXEMPTIONSwith a rationale. Honest scope: CREATE INDEX column references only — constraints, views, and trigger bodies are a filed TODOS.md follow-up.parseBaseTableColumns` strips SQL line + block comments before identifying column names so commented-out lines don't hide adjacent columns.
  • test/dream-verdict-cache-ttl.test.ts — dream_verdicts TTL contract on PGLite: put assigns the default TTL, expired rows miss on read and only they are swept, re-judging via upsert refreshes a nearly-expired row, the migration backfill derives expiry from judged_at idempotently, and a NULL-expiry row (the pre-backfill upgrade window) reads as a hit and survives the sweep — the locally-runnable pin for the NULL-tolerant read predicate both engines share.
  • test/helpers/schema-diff.ts + test/helpers/schema-diff.test.ts + test/e2e/schema-drift.test.ts — cross-engine schema parity gate. Helper exports pure snapshotSchema(query) / diffSnapshots(pg, pglite, opts) / formatDiffForFailure(diff) / isCleanDiff(diff) over a four-tuple per column (data_type, udt_name, is_nullable, column_default). E2E test spins up fresh PGLite + Postgres, runs engine.initSchema() on each, snapshots information_schema.columns, then diffs. 2-table allowlist (files, file_migration_ledger) — every other Postgres table must reach PGLite via PGLITE_SCHEMA_SQL or a migration's sqlFor.pglite branch. Sentinels for oauth_clients, mcp_request_log, access_tokens, eval_candidates give tighter blame messages. Skips without DATABASE_URL. Wired into scripts/e2e-test-map.ts so changes to src/schema.sql, the PGLite schema façade or its generated template, or the migration runner trigger it. The failure message names every drift with a paste-ready hint pointing at src/schema.sql (or its TS fragment) plus bun run build:schema, a PGLite rule in scripts/build-schema.ts, or a migration's sqlFor.pglite branch.
  • test/schema-catalog-golden.test.ts + test/e2e/schema-catalog-golden.test.ts — refactor wave 1 E4 catalog goldens (test/fixtures/goldens/catalog/). snapshotCatalog(query) in test/helpers/schema-diff.ts pins columns with ordinal position, defaults, full index definitions, constraints incl. CHECK text, triggers, function signature + body sha256, views, policies, grants (connecting role placeholdered), RLS flags, sequences and extension names; extension-owned objects are excluded. Captured at three configs (CATALOG_CONFIGS in test/helpers/schema-catalog.ts: default OpenAI/1536, 4096 dims which skips the chunk and halfvec HNSW indexes, GBRAIN_FTS_LANGUAGE=portuguese) on PGLite engine init, the PGLite schema blob without migrations, Postgres engine init (fresh throwaway database per capture) and Postgres db.initSchema(), plus an optional PgBouncer arm. Every capture runs twice through the schema-catalog-lines-v1 normalizer. PGLite master-vs-branch compares ordinals (the goldens pin them); PG ↔ PGLite parity stays name-based. Regenerate deliberately with GBRAIN_TEST_UPDATE_GOLDENS=1.
  • test/pglite-upgrade-replay.test.ts — refactor wave 1 EO3: opens the pinned master-built PGLite brain (test/fixtures/goldens/pglite-upgrade-replay/brain.tar.gz + MANIFEST.json, rebuilt only on master with bun scripts/build-pglite-upgrade-fixture.ts) with current code and asserts the catalog before boot equals the E4 golden, the post-boot catalog is pinned, corpus rows and fingerprint survive, and a restart + repeat initSchema() is a no-op (no migrations, identical catalog, unchanged relation/constraint/function OIDs; the triggers the blob recreates on every boot are pinned).
  • test/setup-branching.test.ts — setup flow.
  • test/slug-validation.test.ts — slug validation.
  • test/storage.test.ts — storage backends.
  • test/supabase-admin.test.ts — Supabase admin.
  • test/yaml-lite.test.ts — YAML parsing.
  • test/check-update.test.ts — version check + update CLI.
  • test/pglite-engine.test.ts — PGLite engine, all BrainEngine methods including addLinksBatch / addTimelineEntriesBatch (empty batch, missing optionals, within-batch dedup via ON CONFLICT, missing-slug rows dropped by JOIN, half-existing batch, batch of 100) plus connect() error-wrap assertion (original error nested, #223 link in message, lock released).
  • test/links-timeline-jsonb-poison.test.ts — the PGLite half of the JSONB batch-poison lock (always-on, no DATABASE_URL). Locks the jsonb_to_recordset batch-insert path for links/timeline/takes against free-text "poison" payloads (commas, quotes, backslashes, braces, em-dashes) and asserts NUL is stripped from free-text body fields but rejected in identity fields. Lone-UTF-16-surrogate cases: every free-text field (link context; timeline summary/detail/source; take claim/source) well-forms to U+FFFD across batch + scalar write paths, while a surrogate in an identity field (slug) still fail-closed rejects the batch. The Postgres lane is test/e2e/jsonb-batch-poison-postgres.test.ts.
  • test/engine-factory.test.ts — engine factory + dynamic imports.
  • test/integrations.test.ts — recipe parsing, CLI routing, recipe validation.
  • test/publish.test.ts — content stripping, encryption, password generation, HTML output.
  • test/backlinks.test.ts — entity extraction, back-link detection, timeline entry generation.
  • test/lint.test.ts — LLM artifact detection, code fence stripping, frontmatter validation.
  • test/report.test.ts — report format, directory structure.
  • test/skills-conformance.test.ts — skill frontmatter + required sections validation.
  • test/resolver.test.ts — RESOLVER.md coverage, routing validation; round-trip that every quoted RESOLVER.md trigger matches a frontmatter triggers: entry in the target skill, and every name="<word>" reference in any SKILL.md resolves to a declared op in src/core/operations.ts or a Minions handler in PROTECTED_JOB_NAMES.
  • test/search.test.ts — RRF normalization, compiled truth boost, cosine similarity, dedup key.
  • test/sql-ranking.test.ts — source-boost helpers: longest-prefix-match in SQL CASE, detail=high temporal-bypass, three-meta-char LIKE escape (%, _, \), single-quote SQL-literal doubling, env override parsing for GBRAIN_SOURCE_BOOST + GBRAIN_SEARCH_EXCLUDE, resolveBoostMap / resolveHardExcludes merge semantics.
  • test/dedup.test.ts — source-aware dedup, compiled truth guarantee, layer interactions.
  • test/query-intent-legacy.test.ts — query intent classification: entity/temporal/event/general (the non-concept intents). test/query-intent-concept.test.ts — the concept intent: definitional/landscape cue detection, the proper-noun / quoted-phrase / sub-3-word guards, vector-lean weight routing.
  • test/eval.test.ts — retrieval metrics: precisionAtK, recallAtK, mrr, ndcgAtK, parseQrels.
  • test/brainbench-fixtures.test.ts / test/brainbench-generator.test.ts / test/brainbench-metrics.test.ts / test/brainbench-continuity.test.ts / test/brainbench-writeback.test.ts / test/brainbench-adapters.test.ts / test/brainbench-scoreboard.test.ts — the BrainBench memory-conformance unit suites (src/eval/brainbench/): fixture loader/validator + the sealed-gold seal (a gold key inside a fixture must reject) and committed-corpus integrity; generator determinism (the committed corpus is exactly what gen.ts produces, holdout discipline, category counts); metric formulas over hand-built turn rows (zero should-retrieve turns, empty injections, acceptable-vs-gold asymmetry, micro-averaging); cross-harness continuity (writer's decision persists through the production write-back pipeline, reader recalls on the SAME brain); write-back grading the PRODUCTION conversation→facts pipeline via the injected gold extractor; adapter seam contracts over hermetic PGLite (budget caps, suppression modes); scoreboard + gate governance (baseline determinism, count-aware gating, corpus-bless modes, justification flow, isolation gates-at-zero). test/brainbench-floors.test.ts — the pre-registered quality floors as executable assertions against the committed baseline (a baseline bless can't bank a threshold violation).
  • test/eval-brainbench-e2e.slow.test.ts — BrainBench CLI end-to-end via subprocess against a small tmp corpus: the literal exit codes (0 pass / 1 regression / 2 error-or-inconclusive — the CI product), --out artifact validity incl. _meta.metric_glossary, byte-deterministic --update-baseline, anti-vacuous-pass, and the eval run-all once-per-sweep record. Slow-tiered with its own CI job (slow-brainbench-e2e); independent CLI runs execute once through a width-2 pool in beforeAll. There is no in-process full-corpus completion test: CI's brainbench gate runs the committed corpus fresh on every PR and its baseline compare is the fixtures-hash drift guard.
  • test/check-resolvable.test.ts — resolver reachability, MECE overlap, gap detection, proximity-based DRY detection, extractDelegationTargets coverage.
  • test/dry-fix.test.ts — auto-fix: three shape-aware expander pure-function tests; five guards (working-tree-dirty, no-git-backup, inside-code-fence, already-delegated within 40 lines, ambiguous-multi-match, block-is-callout).
  • test/doctor-fix.test.ts — gbrain doctor --fix CLI integration: dry-run preview, apply path, JSON output shape.
  • test/doctor-registry.test.ts — doctor check registry contract: every entry name and emitted check categorized (FAIL/Why/Fix/See), emits[] equals the AST-walked names of each entry's run, runtime registry equals the static walk, STOP gates at master's early returns.
  • test/doctor-mode-matrix.serial.test.ts — doctor mode matrix through the registry runner: entries run, STOP position, engine calls and --fix mutations for default, --fast, --fix, --fix --dry-run, no engine and connection failure.
  • test/backoff.test.ts — load-aware throttling, concurrency limits, active hours.
  • test/transcription.test.ts — provider detection, format validation, API key errors.
  • test/enrichment-service.test.ts — entity slugification, extraction, tier escalation.
  • test/minions.test.ts — Minions job queue: CRUD, state machine, backoff, stall detection, dependencies, worker lifecycle, lock management, claim mechanics, depth/child-cap, timeouts, cascade kill, idempotency, child_done inbox, attachments, removeOnComplete/Fail, max_stalled clamp/default/plumbing coverage.
  • test/minion-queue-renewlock-signal.test.ts — renewLock forwards its optional AbortSignal to executeRawDirect (stub-engine capture); legacy 3-arg calls unchanged; token-fence miss returns false.
  • test/cycle-drain-renewal.test.ts — runDrainRenewalTick (cycle drain): per-call signal aborted on timeout (slot released), onLost once on a lost fence, throws swallowed, hung renewal resolves at the deadline. Plus two structural source-text pins on inline-drain.ts (the shape guard only covers worker.ts): the renewal must not go back to a raw setInterval(() => queue.renewLock(...)), and the handler invocation must stay wrapped in withChatPhase('job:<name>') so a drained child's gateway spend is attributed to the child rather than absorbed by an enclosing phase: tag.
  • test/queue-probe-cancellation.test.ts — probeQueueState/queryWedgeSignals signal threading: the 1500ms budget CANCELS the losing probe query; fast-path signals never abort; throw still collapses to {probe_failed: true}.
  • test/db-pool-max-lifetime.test.ts — resolveMaxLifetimeSeconds: env forms, 0-disables, 30–60min jitter bounds, warn-once on invalid, per-call jitter variance.
  • test/pool-gauge.test.ts — CheckoutGauge pure semantics + the PostgresEngine seams with fake pools: counted while in flight, released on resolve, on REJECTED queries, and on the SYNCHRONOUS pre-aborted-signal throw (leak guards); getPoolDiagnostics fail-open.
  • test/db-probe.test.ts — runDbProbe verdict matrix (pool_starved / server_unreachable / unknown), honest-disjunction + no-waiter-arithmetic wording pins, hung probes cancelled via their signals, diagnostics absent/throwing fail open.
  • test/postgres-engine-reserved-routing.test.ts — withReservedConnection routing: direct pool when dual-pool active, read pool when kill-switched/in-tx, semaphore cap (directPoolSize−1) with read-pool overflow, permit released on fn throw and reserve failure.
  • test/job-isolation-protocol.test.ts — outcome-file codec round-trip + every decode failure path (missing/malformed/oversize→UnrecoverableError; byte counts, never content), handler-error instanceof reconstruction, child-CLI invocation resolution, and REAL detached-process killProcessGroup tests incl. the grandchild-death guarantee (exercises the Bun negative-pid /bin/kill fallback for real under bun test).
  • test/run-child-entry.test.ts — runChildJobEntry on real in-memory PGLite with a REAL claim-minted token: success (fenced updateProgress lands), handler-failure outcome (exit 0), token-mismatch never runs the handler (exit 14), missing job/handler, parent-death watchdog aborts a live handler.
  • test/child-job-runner.test.ts — runJobInChild against real .mjs children: success + full env contract (incl. GBRAIN_DIRECT_POOL_SIZE=1), error/lease outcome reconstruction, crash, SIGTERM-ignorer → group SIGKILL at the injected grace, pre-aborted signal, spawn ENOENT → ChildSpawnInfraError, worker-shutdown drain (report-during-drain completes; non-reporting kill → ChildWorkerShutdownError).
  • test/worker-job-isolation.test.ts — full parent path on PGLite with the fake-run-child.mjs fixture: claim → child → fenced completeJob (real token over env), error outcome → failJob, crash burns the attempt, spawn failure RELEASES with zero attempts burned, and the serialization-parity pin (unreportable results fail in BOTH modes, never falsely complete).
  • test/jobs-isolation-flag.test.ts — parseJobIsolationFlag: space/= forms, env fallback + flag-wins, empty-env default, other flags untouched.
  • test/extract.test.ts — link extraction, timeline extraction, frontmatter parsing, directory type inference.
  • test/extract-db.test.ts — gbrain extract --source db: typed link inference, idempotency, --type filter, --dry-run JSON output.
  • test/extract-fs.test.ts — gbrain extract --source fs: first-run inserts + second-run reports zero, dry-run dedups candidates across files, second-run perf regression guard for the N+1 dedup bug.
  • test/link-extraction.test.ts — canonical extractEntityRefs both formats, extractPageLinks dedup, inferLinkType heuristics, parseTimelineEntries date variants, isAutoLinkEnabled config.
  • test/graph-query.test.ts — direction in/out/both, type filter, indented tree output.
  • test/features.test.ts — feature scanning, brain_score calculation, CLI routing, persistence.
  • test/file-upload-security.test.ts — symlink traversal, cwd confinement, slug + filename allowlists, remote vs local trust.
  • test/query-sanitization.test.ts — prompt-injection stripping, output sanitization, structural boundary.
  • test/search-limit.test.ts — clampSearchLimit default/cap behavior across list_pages and get_ingest_log.
  • test/repair-jsonb.test.ts — JSONB repair: TARGETS list, idempotency, engine-awareness.
  • test/migrations-v0_12_2.test.ts — JSONB-repair orchestrator phases: schema → repair → verify → record.
  • test/orphans.test.ts — orphans command: detection, pseudo filtering, text/json/count outputs, MCP op.
  • test/postgres-engine.test.ts — statement_timeout scoping: sql.begin + SET LOCAL shape, source-level grep guardrail against a reintroduced bare SET statement_timeout.
  • test/sync.test.ts — sync logic + regression guard asserting top-level engine.transaction is not called.
  • test/sync-pull-failed-anchor.serial.test.ts — a failed internal git pull (local-path origin vs protocol.file.allow=never) with zero imports returns partial/pull_failed (not up_to_date), freezes last_commit + last_sync_at, recovers after a manual pull; fall-through import of local commits preserved. Serial: pins GBRAIN_HOME to a temp dir for the whole file.
  • test/sync-concurrency.test.ts — autoConcurrency() thresholds + PGLite-forces-serial + explicit-override clamping; shouldRunParallel() explicit-bypasses-floor contract; parseWorkers() validation rejecting '0'/'-3'/'foo'/'1.5'/trailing chars.
  • test/sync-parallel.test.ts — PGLite-routed coverage of the bookmark gate under concurrency, head-drift gate, vanished-file failure capture, PGLite-stays-serial, and the gbrain-sync writer-lock contract.
  • test/sync-all-missing-path.test.ts — sync --all --missing-path <fail|skip> pure helpers: parseMissingPathMode (default fail, explicit values, loud rejection of bad/dangling values, never swallows a following flag) and partitionMissingPathSources (classification driven only by the injected pathExists predicate — no fs; null local_path passes through runnable; order preserved).
  • test/sync-failures.test.ts — classifyErrorCode regex coverage for all 12 codes against literal production message strings from markdown.ts and import-file.ts; summarizeFailuresByCode sort + pre-classified-honor; recordSyncFailures code-field persistence; acknowledgeSyncFailures AcknowledgeResult shape + backfill on legacy entries.
  • test/sync-soft-delete.serial.test.ts — removed-file recovery arc: a git rm drained by sync SOFT-deletes the page (deleted_at set; row recoverable, not gone), an already-soft-deleted row isn't re-flipped (purge clock preserved), batch delete failures decompose to per-file batches and the run banks instead of aborting, delete → re-add inside the window revives via upsert (content updated, chunks replaced, no duplicate), soft-deleted pages stay invisible to search/getLinks/getBacklinks, the rename lane converges against an out-of-band soft delete, and full-sync reconcile + the unsyncable lane are SOFT with the purge window honored end-to-end.
  • test/sync-exclude-config.test.ts — persisted sync.exclude reach: honored with no flag on incremental AND first-sync full-walk paths, trailing-slash covers directory contents, a per-call flag narrows without re-opening the persisted scope, mixed comma+newline multi-pattern values, conservative posture (pages imported before the exclusion stay live, incl. full-sync reconcile), and a throwing/unreadable config read degrades to no-persisted-scope instead of breaking the sync.
  • test/sync-include-hidden-config.test.ts — persisted sync.include_hidden reach (the dot-directory waiver's twin to sync.exclude): baseline control (no config, no flag → dot-directory pruned), honored with no flag on incremental AND first-sync full-walk paths, trailing-slash covers nested files (lowercased slug), an unnamed dot-directory stays pruned, a per-call includeHidden unions with the persisted waiver, and a throwing config read degrades to no-waiver instead of breaking the sync.
  • test/doctor.test.ts — doctor command; assertions that jsonb_integrity scans the four JSONB write sites and markdown_body_completeness is present.
  • test/utils.test.ts — shared SQL utilities + tryParseEmbedding null-return and single-warn semantics.
  • test/build-llms.test.ts — llms.txt/llms-full.txt generator: path resolution, idempotence, spec shape, regen-drift guard, content contract, AGENTS.md install-path mirror, size-budget enforcement.
  • test/oauth.test.ts — OAuth 2.1 provider: register, getClient, client_credentials grant exchange, authorization_code flow with PKCE challenge/verifier, refresh token rotation, verifyAccessToken with both OAuth + legacy access_tokens fallback, revokeToken, sweepExpiredTokens; contract test asserting scope + localOnly annotations on all operations; coerceTimestamp unit cases (null/undefined/string/number/throw-on-NaN); NULL-expires_at-as-expired contract for both refresh + access token paths; cascade-delete contract asserting revoke-client purges oauth_tokens + oauth_codes via FK CASCADE; cross-client isolation (wrong-client attempt MUST reject AND rightful owner MUST still succeed atomically afterward); empty-string redirect_uri bypass guard; PKCE DCR public-client gate (token_endpoint_auth_method: "none" returns no client_secret, default client_secret_post clients get the one-time-reveal secret, getClient NULL→undefined normalization, full PKCE /authorize → /token round-trip against a public client).
  • test/mcp-dispatch-summarize.test.ts — summarizeMcpParams invariants: declared-keys allow-list intersection, attacker-key-name leak guard (unknown keys counted not named), 1KB byte bucketing for size-probe defense, missing op falls through to fully-redacted shape, declared-keys sorted for deterministic output.
  • test/trust-boundary-contract.test.ts — fail-closed trust semantics under cast bypass: ctx.remote === undefined treated as remote/untrusted at every flipped call site; as any and Partial<> spreads can't downgrade trust by accident.
  • test/remote-privacy-sweep.test.ts — registry-driven remote privacy sweep: every non-localOnly op dispatched remote-shaped through dispatchToolCall against a corpus seeded with high-entropy private sentinels, in both scalar and federated caller shapes; the full response envelope (structured fields, rendered text, errors, _meta.brain_hot_memory) asserted sentinel-free. Fail-closed maintenance contract: a new op fails the suite until classified in EXPECTED_OUTCOME (+ PARAM_FACTORY if it can return corpus data); localOnly ops asserted denied over non-stdio transports; publish-gated ops must deny naming their gate. Curated static sibling: test/operations-trust-boundary.test.ts.
  • test/check-resolvable-cli.test.ts — CLI wrapper: exit codes, JSON envelope shape, AGENTS.md fallback chain.
  • test/regression-v0_16_4.test.ts — findRepoRoot regression guard, hermetic startDir parameterization.
  • test/repo-root.test.ts — findRepoRoot walk semantics + default-arg parity; the 4-tier autoDetectSkillsDir fallback chain ($OPENCLAW_WORKSPACE → ~/.openclaw/workspace → repo-root → ./skills); RESOLVER.md/AGENTS.md filename precedence; explicit-env-wins-over-repo-root; tier-0 $GBRAIN_SKILLS_DIR valid/invalid/precedence-over-OPENCLAW_WORKSPACE; the install-path walk in autoDetectSkillsDirReadOnly; no-drift on primary success; AUTO_DETECT_HINT + AUTO_DETECT_HINT_READ_ONLY content; regression guard asserting the shared autoDetectSkillsDir MUST NEVER return 'install_path' source (how the read-path/write-path split stays safe).
  • test/resolver-merge.test.ts — multi-file resolver merge: findAllResolverFiles empty / RESOLVER.md-only / AGENTS.md-only / both-present (RESOLVER.md first); checkResolvable merge semantics across skills/RESOLVER.md + ../AGENTS.md for the OpenClaw layout where the skillpack ships a thin RESOLVER.md and the real dispatcher lives at the workspace root; dedup by skillPath (first occurrence wins); AGENTS.md-at-workspace-root works alone.
  • test/filing-audit.test.ts — filing audit: writes_pages / writes_to frontmatter, filing-rules JSON validation.
  • test/skill-brain-first.test.ts — shared frontmatter parser; analyzeSkillBrainFirst compliance ladder across 9 fixtures under test/fixtures/brain-first-skills/ (compliant-callout, compliant-phase, compliant-position, exempt-frontmatter, missing-brain-first, multi-pattern, negation-prose, no-external, typo-frontmatter); offset helpers; external-lookup regex shape; audit snapshot+diff transition logic; FORMERLY_HARDCODED_EXEMPT regression absorption.
  • test/routing-eval.test.ts — fixture parsing, structural routing, ambiguous_with, Haiku tie-break layer.
  • test/skill-manifest.test.ts — skill manifest parser: drift detection, managed-block markers.
  • test/skillify-scaffold.test.ts — gbrain skillify scaffold stubs: SKILL.md, script, tests, routing-eval fixtures.
  • test/skillpack-install.test.ts — skillpack bundle + surviving installer primitives: bundle.ts enumeration (manifest load/validate, dependency closure, --all) and the installer.ts seams that outlived the removed skillpack install command (diffSkill behind gbrain skillpack diff, managed-block build/parse, lockfile concurrency, atomic writes).
  • test/http-transport.test.ts — HTTP transport: bearer auth + missing/no-Bearer/unknown/revoked + /health bypass; dispatch.ts round-trip; invalid_params; application/json response shape (not SSE); CORS default-deny + allowlist; body cap on Content-Length AND chunked; two-bucket rate limit (refill, exhaust+Retry-After, LRU eviction, TTL prune, pre-auth IP fires before DB); mcp_request_log audit on success + auth_failed.
  • test/mcp-expose.test.ts — gbrain mcp expose against a fake Tailscale runner in a tmpdir: dispatch + argument shape (exclusive pairs, invalid --port / --surface), plan + consent (TTY prompt, non-TTY without --yes, declined), every Tailscale step (binary lookup and install plan, login including the refusal to sudo a non-system binary, the HTTPS-certificate / Funnel identity pre-checks, publish with the fail-closed serve status read and the foreign-handler refusal), happy paths on linux-systemd and darwin launchd (app-bundle CLI), service edge cases, --status and --remove including receipt-less recovery and the scoped --set-path=/ off, the occupied-port probe (any answer counts), the PGLite lock-holder warning, receipt shape guard + rollback + path confinement, engine detection + summary variants, and the runMcp dispatch regression; never prints a stack trace.
  • test/serve-service.test.ts — the persistent-service half of mcp expose: paths under gbrainPath('serve') + supervisor target detection, renderServeWrapper (and the rendered wrapper actually running under bash), launchd plist + systemd unit renderers, ensureAdminToken (0600, token shape, exclusive-create race), install / uninstall / state probes with their edges and supervisor hardening, receipt read/write + shape validation.
  • test/tailscale.test.ts — the pure Tailscale helpers: parseTailscaleStatus (tolerant of missing fields), findTailscaleBinary, tailscaleInstallPlan per platform, the argv builders, parseServeStatusStrict + handler lookup (non-JSON or non-object output is null, never an empty config), classifyTailscaleError kinds and defaults, and defaultCommandRunner via real spawns of hermetic commands only.
  • test/restart-sweep.test.ts — recipes/restart-sweep.md inlined script: sentinel-anchored fenced-block extraction with salted tmp filenames to bypass ESM cache; constructor-time env reads (proves no module-load snapshot); idempotency layer load/save/atomic-tmp-rename/corrupt-JSON-recovery/30-day-prune; (sessionKey, lastAlertedAt) cooldown gate with 6h threshold; AGGRESSIVE-gate two-state tests; execFile argv shape proving shell metachars in OPENCLAW_TELEGRAM_GROUP cannot reach /bin/sh; real-\n-not-literal alert formatting; GBRAIN_HOME state path override.
  • test/eval-longmemeval.slow.test.ts + test/eval-longmemeval-e2e.slow.test.ts — LongMemEval harness, hermetic with no DATABASE_URL and no API keys, split in two files so CI's LPT bin-packer can shard them: the pure / harness-shared half (harness lifecycle, PGLite create + resetTables over runtime-enumerated pg_tables with the infrastructure tables preserved, schema-migration robustness of the reset, the warm-create speed gate, haystackToPages, the source-boost regression guard, loadResumeSet, the schema-v2 buildByTypeSummary) and the end-to-end half (every describe that calls runEvalLongMemEval against ONE shared benchmark brain: stubbed-LLM answer-gen and --retrieval-only runs, JSONL format + key contract, per-question failure handling, --resume-from, --by-type + --by-type-floor on a no-op resume, a run where every question errored exits 1, duplicate question_id handling).
  • test/eval-longmemeval-mixedcase.slow.test.ts — the like-for-like harness pinned on the _s-shaped mixed-case fixture (test/fixtures/longmemeval-mixedcase.jsonl, placeholder bodies under scripts/check-fixture-privacy.sh): raw-id join through the per-question slug→raw map, strict recall_all vs any-hit on a two-gold question, abstention exclusion, slug_collision error rows, retrieval_config_hash-gated resume, retrieved[] rows for replay.
  • test/eval-longmemeval-parse-args.test.ts — hermetic table test for gbrain eval longmemeval argument validation: every invalid flag value exits 1 from parseArgs before any work (the dataset path is deliberately non-existent so a case that slipped past the parser fails with a DIFFERENT message), --help exits 0.
  • test/eval-longmemeval-cli-smoke.test.ts — subprocess smoke through the real CLI, pre-dispatch flag validator included: the documented --retrieval-only --by-type --no-trajectory --keyword-only invocation exits 0 and writes a by_type_summary line (the flag registry once attributed these flags to another command's row).
  • test/generate-flag-registry.test.ts — the flag-registry marker-segmentation rule: a --flag literal belongs to the command named in the enclosing command === 'X' head for every dispatch shape (plain, compound && args[0] === 'sub', multi-line compound).
  • test/longmemeval-judge.test.ts + test/eval-longmemeval-judge.slow.test.ts — the judged answer-accuracy lane. Pure half: the evaluate_qa.py::get_anscheck_prompt port (one branch per question type, abstention by _abs suffix, the data-boundary framing + tag neutralisation), the official 'yes'-substring verdict rule vs the runner's malformed class, judge_config_hash sensitivity, cost estimates, the BudgetLedger. End-to-end half: --judge on the mixed-case fixture with a canned reader (ThinkLLMClient) and a canned judge (JudgeChatFn) on in-memory PGLite — rows carry the judge fields + reader pins, the summary headline scores ungradable rows as incorrect, judge-only backfill on --resume-from.
  • test/longmemeval-metrics.test.ts / test/longmemeval-resume.test.ts / test/longmemeval-run-config.test.ts / test/longmemeval-emit.test.ts / test/longmemeval-capture.test.ts / test/longmemeval-reader.test.ts / test/longmemeval-splits-fixture.test.ts — pure pins for the harness modules: the raw-id join + recall_all@k / recall_any@k + schema-v2 buckets; resume re-scoring (recall recomputed, never trusted; gold_missing / collisions counted over the same row set as a live run); loadQuestionIds / loadDataset and the retrieval_config_hash --search-pin fold (an absent fold hashes identically to every existing receipt); the emitter's truncate / append modes, atomic summary rewrite and same-file resume compaction; the --capture-pool receipt fields mirroring hybrid.ts's autocut inputs; the reader receiving WHOLE sessions (the 4000-char sanitizer cap does not apply); integrity of the committed seed-42 splits (ids only, disjoint halves, no _abs ids).
  • test/longmemeval-embed-cache.test.ts — src/eval/shared/embed-cache.ts hermetic pins (counting fake transport, bun:sqlite files under a per-test tmp dir, an explicit openai:text-embedding-3-large @ 4 dims gateway): exact (model@dims, text, side) keying, the hard EmbedCacheIntegrityError on a dims mismatch, bypassed / infra_faults accounting, the canonical hash. Its canonical-hash describe is listed in scripts/structural-suites.tsv.
  • test/longmemeval-diagnostics.test.ts — miss diagnostics (src/eval/longmemeval/diagnostics.ts) pure pins: the class decision table (synthetic arm ranks → class), the frozen clause splitter, the H1 signature and H3a/H3b split, receipt parsing + top-k reading, split membership, the summary and glossary header.
  • test/replay-autocut-floor.test.ts — src/eval/shared/autocut-replay.ts + its CLI on synthetic pools: the live decisions the shipped default (jump 0.2, minKeep 1, floor 0.35) makes on the fixture pools are HARDCODED literals worked by hand (so validateLive is checked against something the code did not produce), the floor sweep, paired deltas, split-half, normalizePoolRow refusing slug-less rows.
  • test/eval-spend-guard.test.ts — scripts/eval-spend-guard.sh subprocess pins with a temp ledger (env passed to spawn, never mutated): a marker file proves the wrapped command ran; two rows per launch (running reservation, done reconciliation); fail-closed on a missing / unparseable ledger, a malformed amount and a cap breach; actual-cost file precedence.
  • test/r1-namedthing-rerank-ab.test.ts — scripts/r1-namedthing-rerank-ab.ts + the NamedThingBench corpus module, hermetic: the embed transport is stubbed to throw so only the OFF arm runs in-process; the ON arm's paid path is covered by the pure verdict / integrity functions and the CLI dry run's "ON arm skipped" contract; the seed-contract engine lives in beforeAll (test-isolation rule R3).
  • test/ai/gateway-chat-temperature.test.ts — ChatOpts.temperature reaches the AI SDK call and the provider-reported snapshot surfaces as ChatResult.responseModel (the judge pins temperature 0; without the field a judge run would have used the provider default), through the __setGenerateTextTransportForTests seam.
  • test/search/fusion-lists.test.ts + test/search/expansion-variant-budget.test.ts — role-tagged fusion arms + budget-normalized weighted RRF: pure pins for composeFusionLists / rrfFusionWeighted (null budget is byte-identical to unweighted fusion; empty arms cast no vote; a missing original makes every text arm a variant), then search.expansion_variant_budget end-to-end through hybridSearch on a discriminating corpus (in-memory PGLite + deterministic basisEmbedding), delta-asserted.
  • test/search/arm-confidence.test.ts + test/search/arm-confidence-hybrid.test.ts — search.keyword_arm_confidence_floor: the pure statistic + decision, its composition through fusion-lists.ts, the knob plane (bundle, config parse, resolution chain, knobs-hash kacf=, registry), and the hermetic end-to-end where a weak keyword arm is down-weighted only with the floor set.
  • test/search/metadata-boost-gate.test.ts + test/search/metadata-boost-gate-hybrid.test.ts — search.metadata_boost_gate: lexicalArmsVoted / decideMetadataBoosts (relaxed rows never count; image modality exempt), the thread through runPostFusionStages, the knob plane (mbg=, dashboard), and the hermetic end-to-end in which a hub page's backlink / recency boosts are skipped when only the vector arm voted.
  • test/search/relational-rerank-pin.test.ts + test/search/relational-rerank-pin-hybrid.serial.test.ts — search.relational_rerank_pin: pure permutation pins for pinRelationalRows (top block in fused order bounded by max; a row the reranker ranked higher keeps that claim; one row per page; text rows keep relative order; every no-op path returns the input) and the ONE range contract, then the end-to-end on the relational corpus (bodies never name the related entity) with a canned reranker on in-memory PGLite — serial lane because it mocks the reranker module.
  • test/search/relational-intent-memo.test.ts — the default relational pattern set is compiled once per process (identity across calls, stateless sharing).
  • test/config-adaptive-return-keys.test.ts — the adaptive-return / autocut / CRAG search knobs AND the four ranker-wave keys (search.expansion_variant_budget, search.relational_rerank_pin, search.keyword_arm_confidence_floor, search.metadata_boost_gate) are registered in KNOWN_CONFIG_KEYS, so gbrain config set on a documented knob is never a silent no-op.
  • test/longmemeval-sanitize.test.ts — sanitization parity pinning that INJECTION_PATTERNS from src/core/think/sanitize.ts is the single source of truth (adding a pattern there must cover both <take> framing and <chat_session> framing, no per-surface regex drift).
  • test/openai-compat-multimodal.test.ts — gateway's openai-compatible multimodal path: happy-path single + multi-input embedding, unauthenticated proxy mode, dimension-mismatch guard (throws AIConfigError with model id + observed + expected pre-storage), default-dim fallback when recipe declares default_dims, HTTP 401 / 400 / malformed-JSON / non-array error paths, and the Voyage /multimodalembeddings recipe still routing through its dedicated path. Hermetic via the __setEmbedTransportForTests seam.
  • test/serve-stdio-lifecycle.test.ts — MCP_STDIO=1 env guard: stdin EOF does NOT trigger shutdown when the env is set, SIGTERM still does (guard scope is correct), unset env preserves the CLI lifecycle. Exercises the ServeOptions.mcpStdio?: boolean test seam directly so tests don't mutate process.env.
  • test/db-lock-fencing.test.ts — fenced lock identity: a DbLockHandle carries its acquisition fence, refresh() returns true while owned and false after a steal (0-row fenced UPDATE), a stolen-from handle's release() is a fenced no-op that leaves the successor's row intact, and startCycleLockRefresher aborts its controller with LockStolenError on a fenced miss while serializing ticks (a slow refresh never overlaps the next).
  • test/cycle-lock-steal.serial.test.ts — runCycle steal-abort arc end-to-end: a mid-run steal produces a structured partial report (reason: 'lock_stolen'), runs no further phases, and never touches the successor's lock row; a steal-free cycle completes and releases normally.
  • test/cycle-any-abort-signal.test.ts — anyAbortSignal combining: pre-aborted inputs, late aborts propagating their reason, duck-typed signal stubs (no addEventListener) observed via poll, and dispose() detaching the caller-signal listener + clearing the poll timer (the daemon leak class).
  • test/cycle-triage-rescue.test.ts — the dream triage gate: passesTriageGate band arithmetic (floor inclusive, at/above threshold never "rescued"), content-type allowlisting, segment verification through normForGrounding (case/curly-quote/dash folding matches, fabricated segments never do), the ≥40-char + dedupe-by-normalized-quote rules, and fail-closed behavior on every malformed verdict shape (null score, missing/short/non-string segments) with minSegments: 0 as the kill switch.
  • test/cycle-synthesize-verify.test.ts — the mechanical claim verification pass: span extraction with code fences / inline code / wikilinks / link targets masked, odd-mark paragraphs counted unbalanced, the quote ladder (exact keep → normalized replace with the verbatim slice → near-match replace trimmed to the matched tokens inside one speaker turn), refusal of any match that crosses a speaker label, wrong-speaker attribution, canonical number and date grounding, claim units (a failing sentence, list item or table row leaves the body whole and lands in unverified_claims), source-span + speaker provenance in grounding.quotes, diff-scoped verification of pre-existing pages against their pre-run page_versions revision, and per-page fail-open.
  • test/cycle-write-path-mini-eval.test.ts — the hermetic $0 write-path mini-eval. A frozen 3-transcript mini-corpus (high / buried / routine bands, placeholder names, deliberately disjoint from the paid Cat 35 corpus so there is no tuning coupling) drives the REAL runPhaseSynthesize on PGLite: real triage parse + gate incl. the rescue, real fan-out + oneshot drain, real quote verify/repair, real provenance stamp + reverse-write + telemetry. The ONLY stub is the gateway chat transport (__setChatTransportForTests), serving a scripted judge and a scripted child. Scope honesty matters here: a scripted child CANNOT measure whether a prompt change improved model output — that stays the paid benchmark's job (receipts in docs/eval/FIX_WAVE_BASELINES.md). This is the no-API-key regression pin for the MECHANICAL write path, and its salient-unit presence score is the canary that catches emission, chunk-slug-rewrite, and repair-over-deletion regressions in the normal unit lane.
  • test/cycle-repeated-consolidation.test.ts + test/helpers/repeated-consolidation.ts — hermetic three-cycle consolidation pin: real runCycle (synthesize → extract → extract_facts) on PGLite with a scripted agentic child (gateway tool loop) that writes supported claims, fabricated quotes, speaker swaps and invented numbers into new and existing pages. Pins zero checkable invented claims in active memory (page text, timeline rows, facts), every supported claim kept, and the two unquoted inventions the mechanical check cannot see. scripts/repeated-consolidation-experiment.ts prints the same measurements as JSON for a before/after comparison.
  • test/cycle-synthesize-triage.test.ts / test/cycle-synthesize-triage-calibration.test.ts — triage gate wiring inside runTriagePass (reports carry rescued/verified_segments, details.triage rescue + token counters, dry-run parity), plus the 25-fixture calibration corpus (10 high / 10 low / 5 buried, all synthetic placeholders) enforcing band-consistent parsing, a ≥80% band-accuracy rubric-drift pin, and that ≥4 of the 5 buried fixtures reach the gate. TRIAGE_VERSION participates in cache validity, so a rubric bump re-judges rather than serving stale verdicts.
  • test/dream-retriage.test.ts — gbrain dream retriage reads THE shared gate: reconcile-queue never cancels a rescued job, --audit-rejects excludes rescued files from the reject sample, alongside the spend-gate / dry-run / liveness arcs.
  • test/facts-extract-idea-kind.test.ts — the idea extractor kind: taxonomy coercion (known kinds survive verbatim, idea stays idea, unknown kinds coerce to fact), prompt shape (the two precomputed system-prompt variants differ in EXACTLY one clause — the low-tier line — both carry the idea definition and the widened enum, and repeat calls return the identical string so prompt caching still hits), and admission wiring (no admission or an admission allowing low → label-honestly; a high-only admission → skip-low).
  • test/migrations-v145.test.ts — the facts.kind CHECK widening: the migration's structure (canonical name, idempotent flag, probe + widened predicate), a fresh PGLite schema admitting an idea INSERT, and an upgrade from a pre-v145 brain swapping the 5-kind constraint for the widened one with a re-run applying nothing.
  • test/queue-stall-parent-unblock.test.ts — the shared killJobs tail: a stall-exhausted child lands child_done(dead) in its parent's inbox and unblocks the parent, a requeued child doesn't touch the parent, all three reapers route through the tail with their own outcome, and the idempotent stranded-parent sweep self-heals parents whose children were already dead (without unblocking parents that still have a live child).
  • test/queue-started-at-retry.test.ts — every automatic re-run path clears started_at (failJob delayed branch, stall requeue, lease release, promoteDelayed, parent re-claim) so a retried job's wall-clock budget measures execution, not backoff wait; end-to-end survival of the wall-clock sweep on a fresh attempt.
  • test/embed-modality-preserved.test.ts — carryChunkMetadata carries modality + all code-metadata fields through re-embed merges (an image chunk stays image), plus the write-side contract that omitting modality resets it to text (why the shared list is load-bearing).
  • test/embed-oversize-heal.test.ts — oversize-chunk healing pure core: healOversizedChunks split/reindex/metadata-carry (fenced_code chunk_source never coerced), healedChunksToStaleRows remap (only rows still needing embeddings survive), and the healOversizedPageChunks orchestrator incl. the freshness guard (a concurrent rewrite between snapshot and write skips the heal — no clobber).
  • test/embed-oversize-heal-drain.serial.test.ts — the heal wired into all three real drains against PGLite: embedStaleForSource, embedStalePages (phase-end closure), and runEmbedCore --stale (embedAllStale) heal oversized stored rows in place while preserving the embedded sibling's vector.
  • test/embed-stall.test.ts — the embed stall watchdog unit: resolveEmbedStallAbortSeconds env resolution (default 900; garbage → default; <= 0 disables), createEmbedStallWatchdog fire/reset/stop semantics, the run-scoped embedding-API liveness clock, and the assertEmbedNotStalled handler contract (clean result no-op, stalled result throws).
  • test/jobs-embed-stall-wiring.serial.test.ts — the stall contract at the minion boundary: an embed job whose core result carries reason: 'stall_timeout' THROWS (job marked failed, banked progress in the message); a clean result resolves with the embed report.
  • test/embed.serial.test.ts — runEmbedCore lifecycle on PGLite with mock.module seams: abort-signal threading, the stall-watchdog arc against the real lock table (stall fires → single-flight locks released + summary flushed + reason: 'stall_timeout' surfaces; non-CLI callers get the error RESULT, no process exit; live progress keeps the watchdog quiet), cleanup aborting an in-flight heartbeat refresh, and the heartbeat tick-timeout arc (a never-settling refresh times out per tick via GBRAIN_EMBED_LOCK_HEARTBEAT_TIMEOUT_MS; 3 consecutive failures → lock_lost + drain abort).
  • test/handlers-embed-backfill.test.ts — the embed-backfill job handler's budget-cap classification matrix: default cap dropped for unpriced models, pricing.overrides restores enforceability, off uncaps, explicit caps fail closed on unpriced models (incl. an explicit $10 equal to the default), the defaulted cap still enforces for priced models, a present-but-garbage cap value keeps the $10 default FAIL-CLOSED (never droppable), and the handler-lane stall watchdog (a wedged drain aborts and fails the job).
  • test/ai/reranker-readiness.test.ts — the pure rerankerReadiness predicate: the default Voyage model with and without the key (the fix names the key AND the disable command; an empty-string key counts as absent), shape failures that never throw (unknown provider, no reranker touchpoint, unlisted model, keyless local recipe, garbage input), and agreement with gateway.isAvailable('reranker', model) on an env × model matrix.
  • test/rerank-no-key.serial.test.ts — the gateway no_key preflight: RerankError('no_key') before any HTTP call, ONE audit row per process per model with no stderr line, the per-model memo and its test seam, no budget reservation for a skipped rerank (and BudgetExhausted before the transport call when the key IS present), HTTP 401 staying auth; plus applyReranker on no_key — results unchanged, no per-query rows, onSkip fires, a throwing hook never breaks search, and genuine failures never fire it.
  • test/hybrid-reranker-skipped.serial.test.ts — balanced search on PGLite without VOYAGE_API_KEY: reranker_skipped (no_key) stamped on the meta with results kept and nothing printed, fresh searches retain the stamp while shared result caching stays disabled with no cache writes even when requested, and with the key present the reranker runs (rerank_score stamped, no skip entry).
  • test/degraded-stages-recall.test.ts — affectsRecall / RANKING_ONLY_DEGRADED_STAGES: reranker_skipped is ranking-only, every other closed-vocabulary stage affects recall, and a mixed list is degraded iff a recall-affecting stage is present.
  • test/cli-explain-degraded-render.test.ts — formatResult --explain threads the captured retrieval meta: a reranker_skipped stamp renders as the degraded header, a clean run prints no header, and the plain (non-explain) renderer is unchanged.
  • test/doctor-reranker-health.test.ts — the readiness-aware reranker_health check: key absent → warn naming the key and the disable command; key present → ok "ready"; disabled by config row or by the conservative bundle → ok; no_key skip rows informational once ready; a DB-plane key the CLI folded into the gateway counts; the no-gateway fallback to env > file > DB plane; unknown model, self-host override and auth rows; a brain with no embedding provider is never blamed; audit rows for another model never warn on the active default.
  • test/modes-report-reranker.test.ts — buildModesReport attributes the five reranker knobs, the readiness block flips on the key, config overrides are reflected; formatModesText prints the runtime Reranker: line plus the per-bundle reranker=… autocut=… line and surfaces a gateway base-URL override as self_hosted; redactReadinessForRemote strips required_key / key_present / fix and keeps the verdict.
  • test/init-reranker-default.test.ts — writeNewInstallRerankerDefault: a Voyage key present → no write regardless of embedding pick; a keyed non-Voyage pick without the key → explicit search.reranker.enabled=false; a keyless install → no write; a Voyage key that lives only in the DB config plane or only in config.json counts; never-clobber on an existing explicit row.
  • test/import-abort-error.test.ts — runImport preflight/argv failures throw typed ImportAbortError instead of exiting the process; the calling process survives the abort.
  • test/lint-fix-single-pass.test.ts — gbrain lint --fix walks the tree once and total_fixed reports the fixes THIS run applied.
  • test/snapshot-shape-guard.test.ts — PGLite snapshot loader refusal matrix: shape-less version files, dims/model mismatches, and stale schema hashes are all refused; matching hash + shape loads; a migration-handler edit changes the hash.
  • test/stats-health-source-scope.test.ts — source-scoped stats/health/identity: engine-level scoping (every counter confined; degrees/denominators/the islanded predicate scope BOTH edge endpoints; mutating an excluded source moves nothing a scoped caller sees) plus the op layer (remote scalar + federated grants confine get_stats/get_health/get_brain_identity; remote unscoped and the __all__ sentinel fail closed to zeros; trusted local keeps the brain-wide view).
  • test/takes-list-subcommand.test.ts — takes list routing (list is a subcommand, not a slug) + the --limit/--offset flags: cap/skip/paging, engine default without --limit, bare list unchanged, and invalid values (0, non-numeric) exit 1 with the positive-integer message.
  • test/stale-takes-bigint.test.ts + test/take-proposals.test.ts — 64-bit row normalization at the engine boundary: listStaleTakes rows and takes propose --json / loadProposal rows come back as NUMBERS (never bigint/string ids or string weights) on both engines, so takes embed/propose survive real Postgres int8 rows.
  • test/llm-json-reasoning-ladder.test.ts — parseLlmJson's reasoning-block recovery ladder: strips a closed or truncated <think> block ONLY after a raw parse fails (valid JSON containing the tag text is untouched), case-insensitive, array payloads, and the facts/atoms extractors routing through it (the ORIGINAL failure reason is preserved when the retry also fails).
  • test/models-per-task-extract-atoms.serial.test.ts — gbrain models reports models.dream.extract_atoms through the phase's own resolver (pins the narrow-resolver divergence: models.tier.utility is deliberately ignored; unconfigured falls back to the same tier default the runtime uses).
  • test/conversation-facts-pricing-wiring.test.ts — pricing.overrides reaches every conversation-facts entry point: the strict config registry accepts the key, and direct extraction, the cycle backfill, and transcripts --facts all price through the operator override.
  • test/budget/no-pricing-registration.test.ts — explicit cost cap + unpriced model: the refusal carries the lookup-and-register guidance as text and structured fields on BudgetTracker, enrich, conversation facts (core and cycle phase) and skillopt; default caps warn (naming gbrain pricing set) and run; gbrain pricing set merges without clobbering other overrides, validates rates, warns on $0, and the retried run is priced and capped; list/unset; and the trust boundary (no operation registers prices, thin clients refuse gbrain pricing). PGLite + gateway test transports.
  • test/extract-atoms-explicit-cap-no-pricing.test.ts — extract_atoms with an explicit cycle.extract_atoms.budget_usd: an unpriced chat model or embed route is refused (status warn, no model call, details.no_pricing with register_command), the rollup records an expected limit not a halt, doctor's extract_health overlay names the command, and after gbrain pricing set the retried run extracts and the record clears; a default cap still warns and runs. PGLite + _chat seam + embed transport stub.
  • test/cycle/extract-atoms-model-config-fail-soft.test.ts — a throwing getConfig during extract_atoms model resolution falls back to the tier default instead of rejecting the phase.

Lane-move pilot (2026-09)

The 20 heaviest PGLite-only files in test/e2e/ (by scripts/e2e-weights.json) moved out of the sequential Postgres runner into the lanes that run on every PR. Each met the move criterion: it constructs PGLite (or spawns a PGLite CLI) directly, imports nothing from test/e2e/helpers.ts, has no DATABASE_URL/hasDatabase gate, and its header confirmed no Postgres use. sync-delegation-under-serve.serial and dream-synthesize-pglite stayed in test/e2e/ because named e2e.yml jobs run them. Assertions are unchanged; executed-test counts match the E2E runs. A file becomes serial when it mutates process-global state and slow when it takes about 30 s or more. The remaining PGLite-only E2E files are a TODOS.md item decided from the pilot measurement.

Source files whose only E2E owner moved (src/commands/claw-test.ts, src/core/claw-test/**, src/core/brain-resolver.ts, src/commands/mounts.ts, src/commands/connect.ts, src/core/connect-probe.ts, src/commands/embed-facts-delegate.ts) no longer have an E2E_TEST_MAP row, so a change to them selects all E2E (fail-closed); their owners run in every PR.

Former path New path Lane Command Lane reason
test/e2e/claw-test.test.ts test/claw-test.slow.test.ts slow bash scripts/run-slow-tests.sh test/claw-test.slow.test.ts over 30 s (harness subprocess runs)
test/e2e/init-fresh-pglite.test.ts test/init-fresh-pglite.slow.test.ts slow bash scripts/run-slow-tests.sh test/init-fresh-pglite.slow.test.ts over 30 s (CLI subprocesses)
test/e2e/mounts-routing-pglite.test.ts test/mounts-routing-pglite.slow.test.ts slow bash scripts/run-slow-tests.sh test/mounts-routing-pglite.slow.test.ts about 30 s (two persistent PGLite brains, CLI spawns)
test/e2e/qm-provisioning.test.ts test/qm-provisioning.test.ts unit bun test test/qm-provisioning.test.ts no process-global state, under 30 s
test/e2e/minions-field-report-repro.test.ts test/minions-field-report-repro.test.ts unit bun test test/minions-field-report-repro.test.ts no process-global state, under 30 s
test/e2e/fresh-install-pglite.test.ts test/fresh-install-pglite.serial.test.ts serial bash scripts/run-serial-tests.sh test/fresh-install-pglite.serial.test.ts mutates process.env and console
test/e2e/remote-privacy-journeys.test.ts test/remote-privacy-journeys.serial.test.ts serial bash scripts/run-serial-tests.sh test/remote-privacy-journeys.serial.test.ts constructs PGLite outside beforeAll (isolation rule R3)
test/e2e/serve-stdio-roundtrip.test.ts test/serve-stdio-roundtrip.test.ts unit bun test test/serve-stdio-roundtrip.test.ts no process-global state, under 30 s
test/e2e/serve-http-surface-ceiling.test.ts test/serve-http-surface-ceiling.test.ts unit bun test test/serve-http-surface-ceiling.test.ts no process-global state, under 30 s
test/e2e/skillpack-flow.test.ts test/skillpack-flow.test.ts unit bun test test/skillpack-flow.test.ts no process-global state, under 30 s
test/e2e/v0_28_5-fix-wave.test.ts test/v0_28_5-fix-wave.serial.test.ts serial bash scripts/run-serial-tests.sh test/v0_28_5-fix-wave.serial.test.ts mutates process.env
test/e2e/backfill-perf-pglite.test.ts test/backfill-perf-pglite.test.ts unit bun test test/backfill-perf-pglite.test.ts no process-global state, under 30 s
test/e2e/connect-bearer.test.ts test/connect-bearer.test.ts unit bun test test/connect-bearer.test.ts no process-global state, under 30 s
test/e2e/bootstrap-hook-under-serve.serial.test.ts test/bootstrap-hook-under-serve.serial.test.ts serial bash scripts/run-serial-tests.sh test/bootstrap-hook-under-serve.serial.test.ts mutates process.env; already a serial file
test/e2e/upgrade-bun-link-arc.serial.test.ts test/upgrade-bun-link-arc.serial.test.ts serial bash scripts/run-serial-tests.sh test/upgrade-bun-link-arc.serial.test.ts mutates process.argv; already a serial file
test/e2e/dream-synthesize-chunking.test.ts test/dream-synthesize-chunking.serial.test.ts serial bash scripts/run-serial-tests.sh test/dream-synthesize-chunking.serial.test.ts mutates process.env
test/e2e/search-readiness-http.test.ts test/search-readiness-http.test.ts unit bun test test/search-readiness-http.test.ts no process-global state, under 30 s
test/e2e/transcripts-ingest-pglite.test.ts test/transcripts-ingest-pglite.test.ts unit bun test test/transcripts-ingest-pglite.test.ts no process-global state, under 30 s
test/e2e/bootstrap-harness-lifecycle.serial.test.ts test/bootstrap-harness-lifecycle.serial.test.ts serial bash scripts/run-serial-tests.sh test/bootstrap-harness-lifecycle.serial.test.ts mutates console; already a serial file
test/e2e/fact-backfill-resident.test.ts test/fact-backfill-resident.test.ts unit bun test test/fact-backfill-resident.test.ts no process-global state, under 30 s

E2E test inventory

E2E tests live in test/e2e/ and run against real Postgres+pgvector (require DATABASE_URL), except where noted as PGLite in-memory (no DATABASE_URL needed). One file outside the directory also rides the e2e lane: test/phantom-redirect-engine-parity.test.ts (Postgres arm; see the file taxonomy above).

  • test/e2e/write-attribution-postgres.test.ts — runs test/write-attribution.test.ts on Postgres, direct and through transaction-mode PgBouncer (scripts/e2e-backend-matrix.txt): the coordinator's transaction-local actor reaches page versions, live revisions, facts, takes and timeline rows.

  • test/e2e/facts-separation-postgres.test.ts — real-Postgres parity for cross-session facts, supersession, and the pre-limit unconsolidatedOnly predicate used by consolidation.

  • bun run test:e2e runs Tier 1 (mechanical, all operations, no API keys). Includes dedicated cases for the postgres-engine addLinksBatch / addTimelineEntriesBatch bind path — postgres-js's JSONB bind (jsonb_to_recordset(($1::jsonb)->'rows')) differs from PGLite's and gets its own coverage.

  • test/e2e/search-quality.test.ts — search quality against PGLite (no API keys, in-memory).

  • test/e2e/graph-quality.test.ts — knowledge graph pipeline (auto-link via put_page, reconciliation, traversePaths) against PGLite in-memory.

  • test/e2e/jsonb-batch-poison-postgres.test.ts — the real-Postgres half of the JSONB batch-poison lock (the engine whose bind path differs). Seeds free-text "poison" context (Zoom URL with ?pwd=, commas, quotes, Windows backslash path, braces, em-dash) and asserts the links/timeline/takes batch writers do not error with "malformed array literal"; also asserts NUL is stripped from free-text bodies (context/summary/detail/claim) and still rejected in identity fields. Lone-surrogate lock: a lone UTF-16 surrogate in free text (the 22P02 class on Supabase) well-forms to U+FFFD across batch + scalar paths (incl. timeline + take source), while a surrogate in an identity field still rejects the batch. DATABASE_URL-gated.

  • test/e2e/postgres-jsonb.test.ts — round-trips all 5 JSONB write sites (pages.frontmatter, raw_data.data, ingest_log.pages_updated, files.metadata, page_versions.frontmatter) against real Postgres and asserts jsonb_typeof='object' plus ->>'key' returns the expected scalar. Guards against the double-encode bug.

  • test/e2e/integrity-batch.test.ts — parity for scanIntegrity's batch-load fast path vs sequential. Cases (dedup, hits, validate, topPages) seed a fixture and assert both paths return identical results. Dedup case uses raw SQL via getConn().unsafe() to seed a (test-source-2, people/alice) row alongside the default-source row, since engine.putPage doesn't take a source_id. Pins multi-source overcounting; the "multi-source duplicate slugs scan once" case expects both batch + sequential paths to report 2.

  • test/e2e/jsonb-roundtrip.test.ts — companion regression against the 4 doctor-scanned JSONB sites. Assertion-level overlap with postgres-jsonb.test.ts is intentional defense-in-depth: if doctor's scan surface drifts from the actual write surface, one of these tests catches it.

  • test/e2e/sync.test.ts — --skip-failed failure-loop test alongside happy-path tests: broken file → performSync returns blocked_by_failures with grouped breakdown → performSync({skipFailed: true}) advances bookmark and returns AcknowledgeResult with code summary → second broken file → second cycle. Saves and restores the user's real ~/.gbrain/sync-failures.jsonl so the test is hermetic. Asserts bookmark gating, JSONL state, dedup across paths, summary aggregation, and the literal doctor-rendering string format.

  • test/e2e/upgrade.test.ts — check-update against real GitHub API (network required).

  • test/e2e/minions-shell-pglite.test.ts — PGLite --follow inline shell-job path (in-memory, no DATABASE_URL required) — the path the minion-orchestrator skill documents for dev use.

  • test/e2e/job-isolation.test.ts — process isolation on real Postgres (DATABASE_URL-gated, wired EXPLICITLY into .github/workflows/e2e.yml tier1 — the workflow runs only named files): a concurrency-3 isolated drain through real child processes (the fake-run-child.mjs fixture — real spawns, no child DB pools), and the REAL jobs run-child CLI entrypoint end-to-end (engine bootstrap incl. the child's own pools, quiet handler registry, token validation, outcome protocol).

  • test/e2e/sync-reconcile-postgres.test.ts — the sync reconcile's real-Postgres array-parameter binding path (DATABASE_URL-gated). Wired EXPLICITLY into .github/workflows/e2e.yml tier1 beside job-isolation, and listed in the selected-e2e EXCLUDE set so a PR touching sync.ts doesn't run it a second time there.

  • test/e2e/pglite-cli-exit.serial.test.ts — real spawned-CLI exit behavior on PGLite (in-memory, no DATABASE_URL): read commands (search/get/query) exit 0 promptly; CLI_ONLY capture exits clean and frees the single-writer lock; the teardown describes pin every disconnect site — a failed op exits 1 with the error on stderr, and the dashboard, read-only-timeout, doctor, and dream --dry-run paths all exit with no force-exit banner.

  • test/e2e/pgbouncer-teardown.test.ts — PgBouncer TRANSACTION-mode teardown. Pins the bug CLASS, not timings: a CLI op against a txn-mode pooled URL exits 0 with intact stdout and does NOT ride the 10s hard-deadline backstop (the engine.disconnect() did not return banner is the smoking gun). Gated by GBRAIN_PGBOUNCER_URL + GBRAIN_PGBOUNCER_DIRECT_URL (NOT DATABASE_URL) — set automatically by bun run ci:local's pgbouncer compose service. Both URLs survive the E2E runner and preload scrub, while CLI children clear ordinary database overrides so the pooled URL in their isolated config wins. Selected CI runs require a nonzero executed-test count (GBRAIN_CI_REQUIRE_PGBOUNCER=1); missing targets or an all-skipped file fail the gate. It skips gracefully elsewhere. Uses a DEDICATED gbrain_pgbouncer_test database so it never races the gbrain_test TRUNCATE fixtures.

  • test/e2e/volunteer-context-postgres.test.ts — volunteer_context on REAL Postgres (engine parity beyond the hermetic PGLite unit suite): resolution arms through the actual op handler, the fire-and-forget volunteer-event sink landing rows, the stats join, and the RLS pin that context_volunteer_events has ROW LEVEL SECURITY enabled (keeps the v35 auto-RLS event trigger honest for migration-created tables). DATABASE_URL-gated.

  • test/e2e/openclaw-reference-compat.test.ts — check-resolvable + skillpack install-model against a minimal AGENTS.md workspace fixture (test/fixtures/openclaw-reference-minimal/), regression guard for the OpenClaw deployment shape.

  • test/e2e/workspace-generic-compat.test.ts — always-on (PGLite, no binary): pins the INSTALL_FOR_AGENTS.md "any repo with a workspace" contract against test/fixtures/generic-agents-workspace/ (Hermes is the motivating consumer): cwd_walk_up detection, the GBRAIN_SKILLS_DIR override, check-resolvable on a root AGENTS.md, and scaffold additivity + refuse-overwrite. The real Hermes-behavior proof is the door suite below.

  • test/e2e/install-real-hermes.serial.test.ts — the hermes "door": real hermes binary + real hermes mcp add handshake (full-catalog tool discovery; the count tracks the op catalog, so the test asserts discovery happened, not a number) + a paid hermes -z recall turn against a seeded brain. Triple-gated: GBRAIN_REAL_HERMES_E2E=1 (explicit opt-in — run-e2e.sh scrubs GBRAIN_*, so it can never fire under bun run test:e2e) + resolvable binary + non-empty ANTHROPIC key (anthropic-pinned on purpose: a second provider key flips hermes provider-auto into a mis-routed 401). Hermetic HOME + HERMES_HOME with a tripwire on the operator's real config; evidence copies to GBRAIN_E2E_EVIDENCE_DIR for CI upload. Venue: heavy-tests.yml (real-agent-e2e + hermes-door jobs).

  • test/e2e/install-real-grok.serial.test.ts — the grok "door" (xAI Grok Build; every asserted shape observed against the pin in docs/mcp/GROK-CLI-PIN.md). SPLIT-GATED, a deliberate divergence from the hermes door: grok's mcp add/list/doctor run keyless, so the compat tier (version-shape pin, documented-shape grok mcp add gbrain -- gbrain serve --surface verbs via a PATH-staged bin dir, saved-TOML asserts via Bun.TOML.parse, mcp doctor handshake proving the seven-verb surface, vendor-fallback provenance guard, direct-TOML surface) needs only GBRAIN_REAL_GROK_E2E=1 + a resolvable binary; the paid SMOKE additionally needs a non-empty XAI_API_KEY and asserts a PER-RUN NONCE fact (grok has fs/shell tools — the committed fact is greppable, so recall of it proves nothing) with web search disabled. mcp add is lazy (exit 0 always) — mcp doctor <name> --json is the honest discriminator (exit 0/1 observed). Hermetic HOME + GROK_HOME + tmp cwd on every spawn (grok reads vendor MCP configs for trusted folders and loads .envrc from cwd); bounded tripwire over the operator's real ~/.grok config/credential files (volatile paths excluded — grok rewrites logs/sessions/bin/docs every run) + a checkout guard that no .grok//.mcp.json appeared in the repo root. Venue: heavy-tests.yml (real-agent-e2e + grok-door jobs); run directly via GBRAIN_REAL_GROK_E2E=1 bun test test/e2e/install-real-grok.serial.test.ts.

  • test/e2e/install-real-opencode.serial.test.ts — the opencode "door" (SST opencode; every asserted shape observed against the pin in docs/mcp/OPENCODE-CLI-PIN.md). SPLIT-GATED a step past the grok door: opencode's anonymous FREE TIER drives MCP tool calls keyless, so even the nonce SMOKE runs in the keyless tier — T1 bare-semver version pin (the SST-vs-claimant discriminator), T2 documented-shape opencode mcp add gbrain --env … -- gbrain serve --surface verbs + the honest opencode mcp list discriminator (it SPAWNS every server; ✓/✗ text is the assertion surface — exit code is 0 even on failure, and mcp debug is OAuth-only), T2b spawn-gate CANARY (a project-config decoy is spawn-attempted with NO trust prompt — if this ever gates, the bootstrap user-global scope default's rationale changed: re-observe), T3 writer parity (gbrain's opencode-json.ts output handshakes through the real binary; cross-tool preservation both ways), T4 keyless SMOKE (per-run nonce + STRUCTURAL gbrain_* tool_use proof via parseOpencodeJsonl, --format json). The paid T5 anthropic leg additionally needs a non-empty ANTHROPIC_API_KEY and self-validates the pinned model id against the authed opencode models list BEFORE any spend. Hermetic HOME + both XDG dirs + tmp cwd on every spawn; --pure on every probe (mcp list autoloads plugins — a code-execution surface); bounded tripwire over the operator's real opencode configs/auth.json + a repo-root checkout guard. Venue: heavy-tests.yml (real-agent-e2e + opencode-door jobs, plus the schedule-only opencode-door-canary latest-version leg — continue-on-error, a pin-refresh signal, never a gate); run directly via GBRAIN_REAL_OPENCODE_E2E=1 bun test test/e2e/install-real-opencode.serial.test.ts.

Door cadence policy: the NEWEST door agent runs at nightly/schedule cadence (currently opencode, whose canary leg also tracks latest); a door drops to label-only (real-agent-e2e) after 2 stable monthly cycles with unchanged pins. Rationale: churn concentrates in the newest integration; steady-state doors pay for themselves on demand, not nightly.

  • test/helpers/tty-harness.ts + test/tty-harness.test.ts — the DX real-PTY harness (Bun.spawn({terminal:})): pure text/timing helpers unit-tested with zero subprocesses, plus three live PTY smokes against sh guarded by describe.skipIf(!ptySupported()). The harness itself is a dev instrument surface — its consumer scripts/dx-explore.ts never runs in CI (transcripts land in gitignored .context/dx-runs/); see docs/guides/bootstrap.md for the scenario runbook.
  • test/e2e/search-swamp.test.ts — reproduces the source-swamp case. Seeds a curated originals/talks/article-outline-fat-code page against two <fork>/chat/ pages stuffed with the same multi-word phrase. Asserts the article wins keyword AND vector ranking, that detail=high lets the chat swamp re-surface, and that source_id passes through the two-stage CTE intact. PGLite in-memory.
  • test/e2e/search-exclude.test.ts — test/ + archive/ pages hidden by default, include_slug_prefixes opts back in, caller-supplied exclude_slug_prefixes adds to defaults. Both keyword and vector search paths.
  • test/e2e/engine-parity.test.ts — Postgres ↔ PGLite top-result and result-set parity for searchKeyword + searchVector (Postgres ranks pages then picks best chunk while PGLite returns chunks directly, so the source-boost behavior needs parity coverage). Skips without DATABASE_URL.
  • test/e2e/postgres-bootstrap.test.ts — exercises PostgresEngine.initSchema() directly against a real Postgres database: bootstrap → SCHEMA_SQL → migrations converge from a legacy brain shape, and a brain already at LATEST is an idempotent no-op. Live wedge-class convergence cases rewind a brain to an old schema shape and assert full initSchema convergence: pre-v121 timeline, pre-v143 dream_verdicts (including that pre-existing rows keep their judged_at-derived TTL instead of gaining a fresh 30 days), and pre-v7/pre-v136 minion_jobs shapes. Also covers the standalone db.initSchema replay path from src/core/db.ts, which shares the same bootstrap. Skips without DATABASE_URL.
  • test/e2e/http-transport.test.ts — gbrain serve --http end-to-end against real Postgres: bearer auth round-trip, last_used_at SQL-level debounce, mcp_request_log row insertion on success and auth_failed paths, /health DB-down → 503 (DB-probing health check), and the dispatch round-trip with a real operation. Skips without DATABASE_URL.
  • test/e2e/serve-http-oauth.test.ts — real-Postgres E2E against gbrain serve --http with full OAuth 2.1. Spawns a subprocess server, registers a client via the CLI, mints client_credentials tokens, exercises the /mcp JSON-RPC pipeline. Real DCR /register HTTP-level response-shape test (asserts typeof body.client_id_issued_at === 'number' over the wire, RFC 7591 §3.2.1); real CLI subprocess test for revoke-client (registers → mints token → revokes via execSync → asserts token rejected at /mcp → asserts re-run exits 1); server fixture flips on --enable-dcr so /register is reachable. bun execSync env-inheritance contract: bun's execSync does NOT inherit env mutations done via process.env.X = ..., only OS-level env from before bun started. helpers.ts loads .env.testing and sets DATABASE_URL via process.env mutation, which is invisible to subprocesses unless env: { ...process.env } is passed explicitly — every subprocess call in this file passes env: { ...process.env }. The same contract applies to the sibling sync/cycle/dream/claw-test E2Es. afterAll cleanup is guarded on clientId (won't throw if beforeAll failed before registration); cleanup errors surface to stderr without throwing so real test failures aren't masked. Also covers the trust boundary: an HTTP MCP submit_job for name: "shell" MUST reject with a permission error (request handler sets remote: true and submit_job's protected-name guard fires), and the same guard rejects subagent submission. Skips without DATABASE_URL.
  • test/e2e/sync-parallel.test.ts — DATABASE_URL-gated. 60-file Postgres sync at concurrency=4 imports all + no connection leak (probes pg_stat_activity before/after to confirm worker engines disconnected). 120-file serial-vs-parallel benchmark prints SYNC_PARALLEL_BENCH N files | serial=Xms | parallel(4)=Yms | speedup=Zx. Asserts parallel ≤ serial × 1.5 (CI-noise tolerant; not a strict speedup gate).
  • test/e2e/multi-source-bug-class.test.ts — PGLite in-memory regression suite pinning every multi-source bug site: listAllPageRefs ordering by (source_id, slug), getPage with sourceId picks the right (source, slug) row, extract-takes processes both overlapping people/alice rows independently, listPages filters correctly with PageFilters.sourceId, addLinksBatch with from/to_source_id targets the right rows, validateSourceId rejects path traversal, reverse-write disk layout uses brainDir/.sources/<id>/<slug>.md for non-default sources, copyMigrationSources lands source metadata before overlapping-slug pages. No DATABASE_URL needed. Wired into scripts/e2e-test-map.ts so changes to extract-takes / patterns / synthesize / embed / extract / migrate-engine auto-trigger it.
  • test/e2e/migrate-engine-sources-postgres.test.ts — DATABASE_URL-gated companion for the legacy engine copier (gbrain migrate --to pglite, and --to postgres with migrate.graduation false): migrates a PGLite brain carrying two non-default sources with overlapping slugs into real Postgres and asserts copyMigrationSources created every sources FK parent (config JSONB intact, not double-encoded) before any page write. Unit-level manifest identity (crash manifest resumes only against the SAME target; legacy engine-only manifests start fresh) is test/migrate-engine-resume.test.ts.
  • test/e2e/facts-fence-reconcile-postgres.test.ts — DATABASE_URL-gated round-trip for the escape-aware fence parser: renders a ## Facts fence whose cells carry literal pipes, backslashes (Windows paths), and empty cells via renderFactsTable, runs the wipe-and-reinsert reconcile (runExtractFacts) on real Postgres, and asserts every cell survives byte-identically with no column shift.
  • test/e2e/source-isolation-pglite.test.ts — PGLite in-memory regression suite pinning the source-isolation seal at two layers. Engine layer: searchKeyword / searchVector / searchKeywordChunks / listPages / getPage / traverseGraph / traversePaths apply sourceId (scalar fast path) and sourceIds (array path) correctly across both engines. Op-handler layer: routes through sourceScopeOpts(ctx) so a read+write-scoped OAuth client bound to --source dept-x cannot see rows from neighboring sources via search, query, list_pages, get_page, or find_experts. Covers both ctx.sourceId (single-source clients) and ctx.auth.allowedSources (federated_read clients) precedence; federated array wins over scalar wins over nothing. No DATABASE_URL needed.
  • test/e2e/think-source-isolation-pglite.test.ts — PGLite in-memory suite pinning the think gather stage's source scope: seeds three sources with cross-source links and embedded takes, then asserts runGather under a federated sourceIds grant (and under a scalar sourceId) keeps every stream — hybrid retrieval, takes keyword + vector (searchTakes/searchTakesVector), and the traversePaths graph walk — inside the grant while still reaching authorized neighboring sources. No DATABASE_URL needed.
  • test/e2e/skill-brain-first.test.ts — doctor reports skill_brain_first check with structured issues; --fix --dry-run previews insertion without writing; --fix applies the canonical Convention callout idempotently; brain_first: exempt frontmatter resolves the warn; brain_first_typo surfaces a paste-ready hint; audit JSONL records detected / resolved / fixed transitions; stable brain emits 0 audit lines/run.
  • Journey suites (each claimed by an scripts/e2e-test-map.ts row; DATABASE_URL-gated unless noted): migrate-engine-pglite-to-postgres.test.ts (whole-brain runMigrateEngine transfer incl. the child-process failure arm — config not flipped), takes-write-ops-postgres.test.ts (takes op layer + withPageLock serialization), propose-takes-jsonb-postgres.test.ts + calibration-profile-write.test.ts (JSONB bind shape on real Postgres), engine-parity-cjk.test.ts (cross-engine CJK keyword parity on an identical corpus — both engines route hasCJK() queries through the shared ILIKE builder in src/core/search/cjk-keyword-sql.ts; top-slug agreement, chunk-grain parity, mixed-query AND semantics, nonexistent-term strictness), code-edges-read-parity.test.ts / ontology-merge-parity.test.ts / chronicle-event-projection-parity.test.ts / health-parity-postgres.test.ts (read-path + getHealth parity), sync-sigkill-resume-postgres.test.ts (real SIGKILL mid-sync; DB-polled checkpoint, stranded-lock reclaim, exactly-once resume), serve-http-source-grant.test.ts (legacy no-grant federated widening vs granted confinement over real /mcp), autopilot-linux-lifecycle.serial.test.ts (PATH-shimmed crontab/systemctl arc, hermetic), and the thin-client daily-driver verb extension inside thin-client.test.ts.
  • Tier 2 (test/e2e/skills.test.ts) requires OpenClaw + API keys, runs nightly in CI.
  • test/claw-test.slow.test.ts (slow lane since the lane-move pilot) also covers live mode token-free via shim agents (OPENCLAW_BIN=<sh script>): the success-oracle break path (a do-nothing agent FAILS), the E0 child-friction merge surviving tempdir cleanup, and the upgrade staging + schema-version probe.
  • If .env.testing doesn't exist in this directory, check sibling worktrees: find ../ -maxdepth 2 -name .env.testing -print -quit and copy it here if found.
  • Run E2E tests without asking permission. When you want to verify behavior, there's a relevant E2E test, or you're shipping anything covered by an E2E suite — spin up the test DB, run the tests, tear down. Don't ask, don't propose it, don't defer. The lifecycle is short (~2-30s startup, sub-minute tests, instant teardown) and the gate value is high. Skipping with "DATABASE_URL unset" is silent regression, not caution.

API keys and running ALL tests

ALWAYS source the user's shell profile before running tests:

source ~/.zshrc 2>/dev/null || true

This loads OPENAI_API_KEY and ANTHROPIC_API_KEY. Without these, Tier 2 tests skip silently. Do NOT skip Tier 2 tests just because they require API keys — load the keys and run them.

When asked to "run all E2E tests" or "run tests", that means ALL tiers:

  • Tier 1: bun run test:e2e (mechanical, sync, upgrade — no API keys needed)
  • Tier 2: test/e2e/skills.test.ts (requires OpenAI + Anthropic + openclaw CLI)
  • Always spin up the test DB, source zshrc, run everything, tear down.

Key-gated live files that no CI job has keys for are left out of the scripts/run-e2e.sh default glob, so the nightly full corpus and the local gates stop counting their skips as discovered coverage. Naming a file on the command line still runs it (the runner keeps provider keys):

File Required key Command
test/e2e/openrouter-anthropic-subagent-replay.live.test.ts OPENROUTER_API_KEY OPENROUTER_API_KEY=... bash scripts/run-e2e.sh test/e2e/openrouter-anthropic-subagent-replay.live.test.ts
test/e2e/openrouter-deepseek-subagent-replay.live.test.ts OPENROUTER_API_KEY OPENROUTER_API_KEY=... bash scripts/run-e2e.sh test/e2e/openrouter-deepseek-subagent-replay.live.test.ts
test/e2e/voyage-rerank-live.test.ts VOYAGE_API_KEY VOYAGE_API_KEY=... bash scripts/run-e2e.sh test/e2e/voyage-rerank-live.test.ts
test/e2e/voyage-multimodal.test.ts VOYAGE_API_KEY VOYAGE_API_KEY=... bash scripts/run-e2e.sh test/e2e/voyage-multimodal.test.ts
test/live/decide-typesafe.live.test.ts TYPESAFE_API_KEY (or JEV_TYPESAFE_API_KEY) + GBRAIN_LIVE_TYPESAFE=1 GBRAIN_TEST_KEEP_PROVIDER_KEYS=1 GBRAIN_LIVE_TYPESAFE=1 TYPESAFE_API_KEY=... bun test test/live/decide-typesafe.live.test.ts

E2E test DB lifecycle (ALWAYS follow this)

The sequential E2E runner requires Python 3 for standard-library XML validation of Bun's native JUnit reports. CI and the local Docker runner provide it; direct host runs must have python3 on PATH before launching tests.

setupDB() clears rows while preserving physical schema. Fixtures that seed fixed legacy-width text vectors use setupLegacyEmbeddingDB() instead: it establishes the canonical test shape after clearing the database, including facts and takes, so a preceding CLI-init test cannot change their assumptions. Custom-dimension and migration tests continue using ordinary setupDB(). For fixtures testing schema/index creation or source-scoped cleanup, preserve that lifecycle and derive incidental text-vector widths from the database.

You are responsible for spinning up and tearing down the test Postgres container. Do not leave containers running after tests. Do not skip E2E tests, do not ask permission to run them — see the "run without asking" rule above.

  1. Check for .env.testing — if missing, copy from sibling worktree. Read it to get the DATABASE_URL (it has the port number).
  2. Check if the port is free: docker ps --filter "publish=PORT" — if another container is on that port, pick a different port (try 5435, 5436, 5437) and start on that one instead.
  3. Start the test DB:
    docker run -d --name gbrain-test-pg \
      -e POSTGRES_USER=postgres -e POSTGRES_PASSWORD=postgres \
      -e POSTGRES_DB=gbrain_test \
      -p PORT:5432 pgvector/pgvector:pg16
    Wait for ready: docker exec gbrain-test-pg pg_isready -U postgres
  4. Bootstrap the schema (required — fresh containers have no oauth_clients, mcp_request_log, pages etc.; tests like serve-http-oauth.test.ts will fail with relation "oauth_clients" does not exist if you skip this):
    DATABASE_URL=postgresql://postgres:postgres@localhost:PORT/gbrain_test \
      bun run src/cli.ts doctor --json > /dev/null 2>&1
    gbrain doctor triggers initSchema() on first connect, which is the canonical way to bring a fresh DB to head. apply-migrations --yes alone does NOT seed the base schema — it runs ALTER-style migrations on top of initSchema. Tests that bypass the engine (raw execSync-spawned auth register-client) hit the schema directly and need this step to have run first.
  5. Run E2E tests: DATABASE_URL=postgresql://postgres:postgres@localhost:PORT/gbrain_test bun run test:e2e
  6. Tear down immediately after tests finish (pass or fail): docker stop gbrain-test-pg && docker rm gbrain-test-pg

Never leave gbrain-test-pg running. If you find a stale one from a previous run, stop and remove it before starting a new one.

Authorization regression gates

test/data-frontmatter.test.ts and test/frontmatter-security.test.ts pin inert frontmatter parsing, opaque serialization, scalar compatibility, and import errors. test/authorization-boundaries.test.ts covers scalar source grants, foreign/private facts, and delegated tool exclusions.

test/oauth-consent-security.test.ts covers pending consent, CSRF, policy changes, duplicate decisions, and uncertain completion. Production HTTP flows live in test/e2e/serve-http-consent.test.ts; client-lock races and grant rollback on Postgres live in test/e2e/oauth-grant-transactions.test.ts.

test/minions-submission-authority.test.ts covers submission schemas, durable policy, lifecycle operations, legacy approval snapshots, file confinement, and both workers. test/e2e/minions-authority-parity.test.ts exercises real Postgres JSONB and authorization behavior. Filesystem write concurrency is pinned by the existing fence and timeline suites.

test/guarded-http.test.ts and test/guarded-http-tls.serial.test.ts cover DNS, TLS identity and ports, redirects, deadlines, body limits, and cleanup. CI runs these boundaries on Bun 1.4.0 and 1.4.2, audits root and admin dependencies, and executes scripts/test-gitleaks-config.sh to prove fixture exceptions still report an unrelated secret in the same file. scripts/scan-worktree-secrets.sh scans tracked files plus new files eligible for commit; tracked ignored files remain included. Full-history scans use gitleaks git . --log-opts=--all from a complete clone and reports must remain private.

The Docker gate sets GBRAIN_CI_DISABLE_TEST_ENV_FILE=1 so a bind-mounted developer .env.testing cannot add credentials or change the isolated test database. Explicit local provider E2E runs can continue using that file.