test/local-install-lifecycle.serial.test.ts exercises isolated Bun-link install,
keyless memory write/read/search, process reopen and migration/post-upgrade with
service-command tripwires. test/e2e/grandfather-projection-postgres.test.ts
checks guarded metadata-only grandfathering, duplicate source slugs, preserved
valid text projections and refusal to seal previously unsealed rows on Postgres.
test/reindex-markdown-persistence.slow.test.ts retains the bounded 3,600-page
real-CLI reindex, SIGKILL and resume workload; the diagnostic benchmark launcher
is scripts/bench-reindex-markdown.ts.
On-demand reference (see CLAUDE.md Reference map). Current behavior + invariants only.
Repository-owned Linux validation jobs use ephemeral Ubicloud runners pinned to Ubuntu 24.04. The Ubicloud Managed Runners GitHub App must have access to this repository and active billing in its connected project; runner labels alone do not grant access. No Ubicloud API token is passed to workflow jobs.
| Workload | Runner | Capacity |
|---|---|---|
| Unit shards, slow and eval jobs, BrainBench, admin browser, shared-skills compatibility, persistence soak, reconciliation crashes and read latency, native Linux cells, OpenClaw startup, JSONB parity, selected E2E and Tier 2 | ubicloud-standard-4-ubuntu-2404 |
4 vCPU, 16 GB RAM |
Serial pool (PR and nightly coverage), verify, PgBouncer/RLS deployment matrix |
ubicloud-standard-8-ubuntu-2404 |
8 vCPU, 32 GB RAM |
E2E Tier 1 (its CLI init spawns exceed their timeouts on 4 vCPUs), label-gated and nightly heavy-tests jobs |
ubicloud-standard-16-ubuntu-2404 |
16 vCPU, 64 GB RAM |
| Label-gated heavy test suite | ubicloud-standard-30-ubuntu-2404 |
30 vCPU, 120 GB RAM |
| Native ARM64 glibc and musl tests | ubicloud-standard-4-arm-ubuntu-2404 |
4 vCPU |
| Coverage reports and Semgrep | ubicloud-standard-4-ubuntu-2404 |
4 vCPU, 16 GB RAM |
| Planning, status aggregation, dependency audit, gitleaks, security regressions and actionlint | ubicloud-standard-2-ubuntu-2404 |
2 vCPU, 8 GB RAM |
macOS and Windows matrices stay on GitHub-hosted runners. Release building and publishing also stay unchanged. The pinned upstream OSV reusable workflow does not expose a runner override, so its runner remains upstream-owned.
Sizes come from measured CPU use, not guesses. Each Ubicloud project shares one
vCPU quota between pull-request CI and agent ci:ubicloud VMs, so an oversized
runner makes every other job wait. On matched VMs (September 2026), a unit
shard is one Bun process that averaged 1.1-1.5 busy cores and took 361s on 4,
8 and 16 vCPUs (488s on 2); the serial pool and verify took the same time on
8 and 16 vCPUs (serial shard 2: 222s and 224s; 277s on 4); the 2,500-write
PGLite soak averaged 1.3-1.8 busy cores and took 523s on 4 vCPUs and 500s on
16. Re-measure with scripts/ubicloud/ubi-runner.sh run -s standard-N before
growing a runner: more CPU does not shorten a single-process job.
test/scripts/ci-runner-routing.test.ts pins capacity and platform routing;
.github/actionlint.yaml declares the exact custom runner labels. Actual
GitHub job records and completed checks establish runner availability; local
workflow tests do not.
Every test file runs on every push to master, on the nightly schedule and on manual dispatch. Pull requests run a narrower matrix of the same files:
| Lane | Pull request | Push to master, nightly, manual |
|---|---|---|
| Security regressions | Linux, macOS and Windows on Bun 1.4.2 | Also Bun 1.4.0 |
| Persistence read latency, deployment matrix, soak, reconciliation crashes | Bun 1.4.2 | Bun 1.4.0 and 1.4.2 |
| Persistence soak size | 2,500 writes | 10,000 writes |
| Native writer locks, native paths changed | Every target on Bun 1.4.2, musl, both Windows probes, OpenClaw | Every target, musl and Windows probe on Bun 1.4.0 and 1.4.2, OpenClaw |
| Native writer locks, other changes | linux-x64-glibc / Bun 1.4.2 smoke cell (full native step list) |
Same as above |
test/export-scale.slow.test.ts |
10,001 pages | 100,001 pages |
The changes job classifies a pull request's changed files with
scripts/ci-native-scope.sh: native lock sources, the native toolchain, IPC,
persistence, publication, backup, export and sync sources, their native tests,
package.json, bun.lock and the workflow files select every target. An
unreadable file list selects every target too. Skipped cells never report a
failure: test-status needs the native-locks and persistence-validation
workflow calls, which succeed when their remaining cells do, so the required
check names are unchanged. test/scripts/ci-pr-scope.test.ts pins every scope.
Shared-skill tests distinguish canonical publication, protocol delivery, installed
files and native harness use. test/shared-skills-transports.test.ts and
test/e2e/shared-skills-transports.test.ts use real HTTP authentication, OAuth
issuance and a new stdio process; they do not prove vendor-native activation.
test/persistence-skill-bundles.serial.test.ts and
test/persistence-skill-crash.slow.test.ts exercise typed file-set CAS and
independent-process publication/restoration kills on PGLite in the unit lane and
PostgreSQL through test/e2e/persistence-skill-bundles-postgres.test.ts.
Shared persistence suites use test/helpers/test-backends.ts: direct invocation
defaults to PGLite, and a safe DATABASE_URL opts into both engines. Their E2E
wrappers select PostgreSQL before registering tests, refusing a missing or unsafe
database instead of silently running only the local backend. Backend selection is
captured at registration so hooks retain it after the import environment restores.
Every backend's assertions remain in the shared suites; engine-specific cases run
in their owning lane.
Ordinary PostgreSQL setupDB() clears fixture data, operator configuration and
source sync identity while retaining config.version and the stored embedding
identity, avoiding historical migration
replay against an already-current schema. Migration-focused fixtures use
setupDB({ replayMigrations: true }); an absent ledger also runs the cold chain.
test/e2e/fixture-reset-postgres.test.ts checks both paths, cleanup and vector-shape
preservation, including the deliberate legacy-width restoration helper. That
helper aligns both the physical columns and stored embedding identity with the
legacy test configuration; ordinary resets preserve that identity.
The required shared-skills-compatibility CI job builds the pinned pre-feature
executable with scripts/build-shared-skills-baseline.sh and supplies
GBRAIN_TEST_OLD_BINARY to test/persistence-skill-old-binary.slow.test.ts.
An absent old executable is an explicit skip, never old-writer evidence.
test/shared-skills-catalog-performance.test.ts runs the reproducible
10/100/1,000-skill read benchmark when GBRAIN_TEST_SHARED_SKILLS_BENCHMARK=1;
its timings and database-call counts are recorded diagnostics, while identity
and catalog-size assertions are deterministic.
The shared-skills cases under evals/harness-instructions/ test interpretation
separately from executed calls and native sessions.
scripts/shared-skills/lifecycle.ts measures real authenticated HTTP enrollment,
revision/asset reads, publication, missed-notification polling, acknowledgments,
queue/recovery bytes, and concurrent read latency at 10/100/1,000 skills. Run it on
a quiet machine with the protocol and fixture boundaries in its
README. Its additional body/asset latency
comparison is experimental; report each engine's measured result without
substituting it for the existing scripts/persistence/performance.ts memory-read
gate. The five pure accounting tests run normally; the small runtime smoke is
explicitly opt-in and never counts as a full performance pass.
test/persistence-reconcile-merge.test.ts pins loss-preserving field choices.
test/persistence-reconcile.test.ts runs the guarded repair and replay contracts
on PGLite and, with an explicit safe DATABASE_URL, isolated PostgreSQL databases.
It covers stale preconditions, current/original grants, private facts, retained
backups, ordinary mutations after repair, and competing publications.
test/reconcile-owner-journey.serial.test.ts drives real CLI requests through
HTTP and stdio PGLite owners before and after activation, restarts the owner, and
independently reads the newly remembered private fact and provenance.
test/reconcile-crash.slow.test.ts and test/e2e/reconcile-crash*.test.ts kill real
processes at all eight publication boundaries with activation off/on. PostgreSQL
uses one file per activation state to stay within the unchanged per-file cap. Optional
GBRAIN_TEST_RECONCILE_CRASH_MANIFEST_DIR retains executed-case evidence.
test/e2e/reconcile-pgbouncer.test.ts requires the transaction-mode pooler when
GBRAIN_CI_REQUIRE_PGBOUNCER=1 and proves repair followed by a new private memory
write. The durable-persistence workflow runs these contracts on both supported
Bun versions and uploads the crash manifests; local CI runs the slow and E2E lanes.
test/docs-navigation.test.ts checks local links and fragments in the primary
install/memory guides and all docs/architecture/key-files/ references, requires
every subsystem to be linked from KEY_FILES.md, and guards against blanket
graph-write and preference-routing claims. The fixture suite
test/scripts/check-key-files-current-state.test.ts covers history markers,
cross-subsystem duplicate entries, and byte caps for the entry docs and references.
Search reliability has real-planner and transport regressions in
test/e2e/vector-candidate-safety-postgres.test.ts,
test/e2e/search-query-contract-postgres.test.ts,
test/e2e/projection-statistics-postgres.test.ts, and
test/search-readiness-http.test.ts. The statistics tests include owner,
restricted-reader and FORCE-RLS roles; the candidate tests distinguish natural
plans from forced-HNSW controls and prove server cancellation of exact fallback.
test/e2e/vector-plan-real-column-postgres.test.ts is the #5824 plan proof on
the real embedding column: in a dedicated 64-dim database (it needs CREATEDB
and CREATE EXTENSION vector, so it runs under bun run ci:local) it first
shows the legacy-guard statement seq-scans on its fixture, then that the
emitted statement uses idx_chunks_embedding across a filter matrix incl. RLS
scope binding, then the stale-heavy escalation, exact-fallback and short-window
cases and every doctor vector_plan outcome. The PGLite side is
test/search/vector-freshness.test.ts; the SQL shape and lockstep are
test/search/vector-statement.test.ts. The opt-in reporter-scale bench is
scripts/bench/vector-plan-5824.ts.
test/e2e/projection-recovery-parity.test.ts runs the shared Markdown/code
recovery, graph-edge preservation and migration-origin contracts against
PostgreSQL; their root suites cover PGLite in the unit lane. PGLite work caps never
count a Promise race as cancellation evidence.
The recovery parity entry also runs symbol-resolver-projection-race.test.ts:
paused resolver/rebuild ordering, atomic rollback, candidate revalidation, and
a real PostgreSQL lock-wait receipt before releasing the competing writer.
test/pglite-in-memory-create-retry.serial.test.ts injects create failures while
using real PGLite instances and a validated schema snapshot. It pins one cold
retry only before an in-memory database has opened, both failure diagnostics,
schema replay after snapshot fallback, post-open cleanup and close poisoning,
concurrent connect/disconnect ordering, exit-code preservation, and exclusion of
the persistent repair path. Run it in its own Bun process because it mocks the
PGLite module. Its recovery cases discriminate against the no-retry base; its
post-open cleanup cases discriminate against a retry that replaces a live database.
test/e2e/serve-http-oauth.test.ts additionally pins confidential POST/Basic revocation, public-client SDK fallthrough, malformed/mixed authentication rejection, cross-client isolation, unknown-token opacity, metadata auth methods, no-store responses, strict post-revoke 401, and retryable backend 503 semantics. SDK-driven discovery and real owner-approved PKCE also pin read-only bootstrap, explicit writer requests, scope clamping, and DCR delegation refusal. test/oauth-scope-hint.test.ts exercises the actual SDK middleware over HTTP without requiring a database.
test/put-page-persistence.test.ts and test/e2e/put-page-persistence-postgres.test.ts
pin durable page acceptance and ordinary-error publication: native contention
returns an accepted pending receipt without changing the page, and replay of its
original UUID commits exactly once after release. Filesystem or required
source-path failure rolls back the database transaction. Embedding failure
preserves the canonical receipt; a delayed result superseded by another revision
cannot install vectors. The PGLite suite also covers scoped physical file paths,
unchanged-content no-ops, legacy hashes, deletion/recreation, and sanitized
diagnostics. Actual process-death boundaries belong to the crash suites below.
test/subagent-required-writes.test.ts and
test/subagent-put-page-rejection.serial.test.ts distinguish a persisted write
from prose-only completion, rejected imports, and historical rejected ledger
envelopes across the Anthropic, gateway, and oneshot lanes. Unchanged saves,
optional-write jobs, and saved pages with failed enrichment are positive controls.
test/cycle/global-freshness-postcondition.serial.test.ts exercises the registered
maintenance handler with failed phases, incomplete children, budget deferrals,
abort/lock loss, and successful warning-only controls.
Assign ownership to an assertion and its execution boundary, not to a test filename or a shared helper. Record the contract, backend, runtime version, OS/architecture/libc, source-versus-compiled artifact, transport/authentication, process/storage/crash boundary, workload size and required cadence. Shared scenario code across two engines is not duplicate engine coverage: PostgreSQL JSONB, locking and pooler behavior are not established by a PGLite pass.
| Responsibility | Execution owner | What it does not establish |
|---|---|---|
| Keyless behavior, structural guards and shared contracts | Unit shards and verify in test.yml; process-isolated serial and dedicated slow lanes where required |
Real PostgreSQL, native activation or compiled behavior |
| PostgreSQL behavior and engine parity | Named and diff-selected jobs in e2e.yml; the complete nightly runner corpus |
Execution of key-gated or native-door assertions merely because their files were discovered |
| Durable publication and recovery under sustained load | persistence-validation.yml and scripts/persistence/README.md |
Power-loss safety, production authentication or equivalence to two smaller databases |
| Native lock ABI and compiled-process exclusion | native-locks.yml, compiled smoke and release validation |
All compiled CLI features or native-harness activation |
| Browser journeys | Required admin-browser job and admin/e2e/*.pw.ts |
Vendor-native agent behavior |
| Native agent doors and heavier operational scenarios | Explicit jobs in heavy-tests.yml |
A passing skipped door or generic protocol test is not native activation |
| Live-provider and optional recipe/eval behavior | Their explicitly configured opt-in commands/jobs | A missing key, early return or skipped assertion is not live-provider evidence |
| Line-coverage accounting | PR prCorpus and nightly fullCorpus reports |
Subprocess coverage, all platforms or proof that every discovered case executed |
Before removing repeated work, identify the surviving owner for the same contract and every relevant boundary, prove that owner actually executes, and retain its cadence, failure gate and coverage artifacts. A shared fixture can reduce maintenance while keeping both engine arms. Making one crash lane authoritative or collecting LCOV in a named owner requires a separate ownership change; nightly sharding alone makes neither change.
The 2026-09-29 test audit's lane reports, inventories and mutation-probe logs are committed under docs/test-audit/2026-09-29/; cite them for the surviving-owner and probe evidence behind a consolidation.
Recorded ownership changes:
-
test/e2e/reconcile-crash.test.tsandtest/e2e/reconcile-crash-unactivated.test.ts: the PR owner ispersistence-validation.yml, called fromtest.ymlon every PR on both supported Bun versions against pg16. Its "Require all reconciliation crash boundaries" step runs both files by name and uploads the crash manifests, unchanged. Both files are inE2E_EXCLUSIONS(PERSISTENCE_VALIDATION_OWNEDinscripts/e2e-matrix.ts), so PRselected-e2edoes not run them a second time;scripts/select-e2e.tsprintsexcluded: <file> (owned by persistence-validation.yml)on stderr when a mapped source changes. The nightly full-corpus E2E run and the local gates (ci:local,ci:ubicloud, their:diffforms) still run them. Run them locally with the same command the workflow uses, withDATABASE_URLexported for the test database from "E2E test DB lifecycle":GBRAIN_TEST_ALLOW_DATABASE_URL=1 \ GBRAIN_TEST_RECONCILE_CRASH_MANIFEST_DIR=.context/reconcile-crashes \ bun --no-env-file test --timeout=180000 \ test/e2e/reconcile-crash.test.ts test/e2e/reconcile-crash-unactivated.test.ts -
Attendance parity (
test/attendance-retrieval.test.ts,test/attendance-repair.test.ts,test/extract-timeline-attendance.test.ts): the unit lane owns the PGLite arm; thetest/e2e/*-postgres.test.tswrappers load the scenarios throughregisterPostgresTests, so E2E runs only the PostgreSQL arm.
Name the profile when reporting “all tests.” The local fast loop, test:full,
ci:local, required PR checks and nightly fullCorpus are not interchangeable
supersets. Native matrices, sustained persistence validation, browser tests and
optional recipe/eval commands have separate responsibilities. A faster nightly
E2E schedule does not shorten a PR critical path dominated by persistence.
Report matched executed timings separately from dry-run partition estimates,
including setup, queueing and retries; never count skip-only output as coverage.
The gate shape and cadence are defined once, by O-CEO-16 (with O-ENG-16 and
O-CEO-9) in the Foundations 1 plan; scripts/scale/gates.ts and
.github/workflows/scale-tier.yml implement it. This section says how to run it.
bun run test:scale -- --pages 10000 [--engine pglite|postgres] [--seed 1] [--corpus-dir <dir>] [--import-mode cli|content] [--enforce] [--out <file.json>]
(scripts/scale/run.ts) generates a deterministic two-source brain from the
seed (scripts/scale/fixture.ts: links, timeline bullets, ## Facts and
## Takes fences, partly overlapping bodies, island pages, a seeded dense vector in 16 dimensions
per page), and imports it into a fresh brain under a temporary GBRAIN_HOME:
PGLite in a fresh data dir, or Postgres in a fresh database created from
DATABASE_URL and dropped afterwards. The default --import-mode cli writes
the Markdown corpus once into --corpus-dir (reused while its manifest
matches) and runs the real gbrain import per source, timing each file from
its progress events; --import-mode content keeps the per-page
importFromContent loop. It then extracts links, timeline, facts and takes,
writes the vectors onto every chunk so the vector arm runs keylessly through
queryEmbedFn, and measures p50 over five runs after a warmup for each op,
each with a known-answer check: get_health, list_pages, local and MCP-path
search, a source-scoped grant search, hybrid query with an injected
vector, traverse_graph, get_backlinks and find_orphans, plus a
cold-process first query and two concurrent receipt-bearing put_pages.
find_orphans is checked with one call at the op's maximum page size: its
rows must hold every fixture island and match total_orphans. It
captures every statement each op sends and replays the reads under
EXPLAIN ANALYZE for the planner check. With the brain closed, it then runs
the large-brain operational ceilings through the real CLI
(scripts/scale/f4d.ts), each an enforced data check with its measurement in
the report's f4d section: a gbrain sync of a fresh source (1,000 files at 10k and up) past a 1 s
progress-aware deadline completes (f4d_sync_deadline); gbrain embed --stale against a local stub embedding endpoint stops at its time budget with
exit 11, the remaining count and the resume command (f4d_embed_budget_stop;
nothing leaves the machine); gbrain serve answers initialize and a search
without hitting its boot deadline (f4d_serve_boot); and at 20,000 pages and
up, gbrain sources add registers a 20,000-file checkout on a fresh managed
brain (f4d_sources_add_20k). Each phase logs its start and end, so a run
stopped by a job timeout shows where it was. The report leads with the headline
metric, MCP search p50 at the run's size as shipped (no manual ANALYZE).
Exit codes: 0 when every enforced gate passes, or always without --enforce;
1 when an enforced gate fails (each failure names the gate and op and prints
its EXPLAIN; the JSON report and a .explain.txt land next to --out);
2 on a usage error; 3 when the harness itself crashed (not a verdict).
Enforced under --enforce: import rate, known answers, no-op re-import,
no duplicates across sources, the import phase timer, and the stats-dependent
gates (PLANNER_HEALTH_ENFORCED in scripts/scale/gates.ts): the Nested Loop
inner-loop gate on the key plans, the budgets phase timer, and planner stats.
Planner stats are probed after the first timed op, because F4b analyzes on the
first planner-sensitive read, and only hot tables above 500 rows must have
pg_stats rows (PGLite only: on Postgres autovacuum owns statistics, so a missing row is report-only).
Interactive ceilings and calibrated budgets (scripts/scale/budgets.json,
written by --calibrate) stay report-only until
bun scripts/scale/trend.ts prints "ceilings stable" over the last five
nightly runs; a reviewer then sets the repo variable
GBRAIN_SCALE_ENFORCE_CEILINGS=1. The same script picks the nightly sizes.
Reproduce any report with the command it prints. The fixture's determinism
is pinned by test/scripts/scale-fixture.test.ts, the gate policy by
test/scripts/scale-gates.test.ts and scale-trend.test.ts, the
find_orphans known answer by test/scripts/scale-orphans-verifier.test.ts,
and a 40-page enforced run by test/scripts/scale-harness.slow.test.ts.
Before adding a test, answer four questions in the PR description or the test header:
- What observable behavior or contract does it protect?
- What credible regression makes it fail?
- Why does existing coverage not already catch it?
- Does it need a production seam that no production caller needs?
A regression test must fail when its fix is reverted; prove it with
scripts/check-test-discriminates.sh (see CONTRIBUTING.md). If question 1 has
no answer, or the answer to question 3 names an existing owner at the same
boundary, do not add the test.
Good: a test that runs gbrain remote ping against a fake MCP server returning
{ status: 'failed' } and asserts exit code 1 with the failure reason in the
JSON output. It protects a user-visible contract, fails if the poll loop reads
the wrong field, and needs no seam.
Bad: a test that reads src/commands/remote.ts and asserts it contains
job.status. It passes when the loop is broken in a way that keeps the token,
fails on a harmless rename, and duplicates the behavioral test above.
Delete or merge a test only with evidence, recorded in the PR body:
- Name the contract the test claims to protect and classify the evidence case below.
- Probe it: make a behavior-breaking edit to the production code (or, for a vacuous assertion, show that such an edit passes), run the test and the surviving owner, then revert. Behavior-preserving edits that fail the test are useful extra evidence of implementation coupling.
- Confirm the surviving owner executes (executed-test counts, not skip output) at the same or a more frequent cadence, with an equal or stronger failure gate, per "Coverage responsibilities before consolidation" above.
- Remove the deleted file's entries from
scripts/ubicloud/weights.json,scripts/test-weights.json,scripts/serial-weights.jsonandscripts/e2e-weights.json, grepscripts/,.github/,scripts/e2e-test-map.tsandtest/fixtures/e2e-unmapped-baseline.txtfor the path, and regeneratescripts/structural-suites.tsv(bun scripts/classify-tests.ts).
Evidence cases:
- Retained contract: the contract still matters. Evidence is a surviving owner at the same boundary plus an executed mutation that fails it.
- Intentionally abandoned contract: the behavior is being removed or was
never shipped. Evidence is the approved disposition plus reachability proof
(no production caller) and a check that no user-facing promise (docs, skills,
--help, CHANGELOG) still describes it. - Vacuous assertion: the test asserts nothing about product behavior (a
constant compared to itself, a copied function, a
typeofprobe that typecheck already enforces). Evidence is a demonstration that a behavior-breaking edit leaves it passing, or that it imports no product code.
Evidence template:
| Deleted test | Probe edit | Result | Surviving owner | Owner result |
|---|---|---|---|---|
test/x.test.ts › "name" |
src/y.ts: what changed |
deleted test passes (blind) | test/z.test.ts › "name" |
fails (N of M) |
The sequential E2E runner gives each test file a fresh HOME and GBRAIN_HOME.
Configuration written by a CLI initialization or schema migration remains
available within that file, but cannot change a later file's selected schema or
harness state. Each file's home is removed after it exits, including failures;
the runner's exit trap also cleans up interrupted runs.
Test command tiers, each with a clear scope:
| Command | What it runs | Wallclock | When to use |
|---|---|---|---|
bun run test |
Parallel unit-test fast loop. Sharded fan-out via scripts/run-unit-parallel.sh (default 4 shards — CPU-detected, clamped to a max of 8; 4 limits local PGLite WASM-init contention; GitHub CI uses 8 unit shards), then a serial pass over *.serial.test.ts. Excludes *.slow.test.ts and test/e2e/*. No pre-checks, no typecheck. Builds/refreshes the PGLite schema snapshot BEFORE the shard fan-out and exports GBRAIN_PGLITE_SNAPSHOT so PGLite-booting files restore a baked schema instead of replaying every migration (~3.5x per booting file; see "PGLite schema snapshot" below). Opt out: GBRAIN_NO_SNAPSHOT=1. Memory-safe by default: total concurrency (shards × intra-shard width) is capped to available memory at GBRAIN_TEST_MEM_PER_FILE_MB (default 1536 — a PGLite WASM instance) per concurrent slot, shedding INTRA-SHARD width first and shards only after it (bun's --max-concurrency bounds only test.concurrent tests — 1 file in the corpus — so intra width is nearly free to shed, while every dropped shard removes a whole bun process of real fan-out; shedding shards first would collapse a 16GB box to a serial 1×4 run, measured 3.25× slower than 4×1 on the same machine). Two phantom-failure classes are automatically re-run serially (the rescue pass): failures carrying the WASM out-of-memory signature, and shards killed externally (SIGTERM/SIGKILL well before the shard timeout — sibling workspaces' process cleanup, memory jetsam). On machines without coreutils timeout, the fallback watchdog drops a .watchdog sentinel before TERMing a shard at the cap so the WEDGED/EXIT-HANG classifier stays reachable there (a bare rc=143 would otherwise read as a plain failure). Phantoms pass serially and the run goes green with an oom_rescued note; real failures fail again serially and stay red. Knobs: GBRAIN_TEST_NO_MEM_ADAPT=1, GBRAIN_TEST_NO_OOM_FALLBACK=1, GBRAIN_TEST_MAX_CONCURRENCY (intra-shard, default 4), GBRAIN_TEST_SHARD_TIMEOUT / GBRAIN_TEST_SHARD_KILL_AFTER, plus --shards N / --max-concurrency N / --dry-run script args. |
a few minutes on a Mac dev box | Inner edit loop. Default. |
bun run verify |
CI's authoritative pre-test gate set, fanned out by scripts/run-verify-parallel.sh through a bounded worker pool (default detect_cpus; override GBRAIN_VERIFY_MAX_PARALLEL) with the heavy checks ordered first (typecheck, the two compile-embed checks, admin build, fuzz bundles, guard self-tests, the PGLite-booting chronicle eval check, whole-tree greps). The battery includes the deterministic check:eval-chronicle eval gate; check:eval-canary is deliberately NOT in the battery (its test-file twin test/eval-canary.test.ts spawns the identical runner in the unit matrix, and CI's verify job and matrix always run together — the package script stays for on-demand runs, so verify-only local callers should know the canary rides the unit lane instead). The CHECKS array in that script is the single source of truth — CI literally calls bun run verify in a dedicated job. |
~50s (pool-bounded; longest check dominates) | Before pushing; before /ship. |
bun run test:full |
verify && bun run test && bun run test:slow && [smart e2e]. Smart e2e runs only when DATABASE_URL is set and propagates its failure; otherwise it prints a skip notice to stderr. Use ci:local to provision the databases and require PgBouncer execution. |
~3-5min depending on slow + e2e | Pre-merge sanity, before opening a PR. |
bun run ci:local |
Independent host gitleaks scans, then frozen dependencies, guards/typecheck, the complete serial and slow lanes, and four unit/E2E shards inside Docker. Each E2E shard has its own pgvector database; selected PgBouncer tests must execute against the transaction-mode pooler. Unit, serial, and slow lanes have database URL overrides unset. Any failed stage fails the command. Complete shard logs survive container teardown under .context/ci-local-shards/. ci:local:diff narrows E2E selection; --no-shard runs unit/E2E sequentially. Doc-only diffs still require successful gitleaks scans. |
Depends on the full corpus | Full local gate before shipping. |
bun run ci:ubicloud |
The ci:local lanes (gitleaks, guards/typecheck, serial, slow, unit, all E2E with required PgBouncer execution) fanned out across ephemeral Ubicloud VMs from one heaviest-first work queue; ci:ubicloud:diff narrows E2E like ci:local:diff. Needs UBICLOUD_API_KEY or UBICLOUD_API_TOKEN, no local Docker. See "Ubicloud fan-out" below. |
~5 min (floor: the longest single file) | Full gate before shipping when a Ubicloud token is available. |
bun run test:slow |
Just the *.slow.test.ts set (intentional cold-path correctness checks). |
seconds-to-minutes | When touching slow-path code. |
bun run test:serial |
Just the *.serial.test.ts set (cross-file-contention quarantine; one bun process per file for true module-registry isolation), run through a POOL of concurrent per-file processes — the isolation is per-process, not per-machine. Dispatch is heaviest-first (LPT) from the advisory scripts/serial-weights.json (seconds; mined from the .context/serial-durations.txt table each run banks; absent/corrupt weights fall back to discovery order, absent keys to the corpus p75 — scheduling only, never correctness; LPT order + the corrupt-weights fail-soft are pinned by test/scripts/run-serial-pool.test.ts). Pool defaults to min(detect_cpus, 4) then memory-adapts (same doctrine as the parallel runner); a small growth-guarded set of files (machine-global state or contention-critical timing — see the justified EXCLUSIVE_FILES list in scripts/run-serial-tests.sh, capped at 3 by test/scripts/serial-files.test.ts) runs on a sequential EXCLUSIVE lane after the pool. Per-test timeout 120s (pooled contention headroom); each pooled file is wall-clock-killed at 300s (timeout -k, exit-hang containment). SHARD=N/M partitions pooled files by duration; the three exclusive files run only on shard 1. Unset runs the complete corpus. Routing variables are cleared before tests start, so nested runners remain independent. Externally-killed files (exit 143/137 or a missing exit sentinel — sibling-workspace cleanup, memory jetsam) get ONE sequential rescue re-run, mirroring the parallel runner's doctrine: phantoms stay green with a rescue note, real failures stay red. Prints per-file PASS lines plus a top-10 slowest-files list. Knobs: GBRAIN_SERIAL_POOL=N (explicit pool width — bypasses the memory clamp; 1 restores fully-sequential), GBRAIN_SERIAL_FILE_TIMEOUT. |
a few minutes for all ~220 files at pool=4 | Debugging quarantined files; CI's serial-tests job. |
bun run test:e2e |
Real Postgres E2E. Requires Docker + DATABASE_URL. Sequential within a shard; SHARD=N/M fans out against separate databases (ci-local runs 4 containers). Activates the PGLite snapshot like every other runner (per-file cold-path opt-outs where the test asserts the path TO post-initSchema state), exporting it as an ABSOLUTE path so CLI children spawned with varying cwd still find it. |
~5-10min | Pre-ship; nightly. |
bun run test:compile-smoke |
Self-update integrity verify under a REAL bun build --compile binary, offline (sets GBRAIN_SELFUPDATE_COMPILE_SMOKE=1). The unit suite mocks the network seams; this proves the dependency-free crypto/base64/JSON verify path survives compilation — the failure mode sigstore-js would have hit. |
~5s (one compile) | When touching src/core/binary-self-update.ts; pre-ship on self-update changes. |
bun run test:admin |
Pinned Playwright Chromium tests for the production embedded admin UI, served with an isolated temporary home/cwd and in-memory PGLite. Exercises owner login, OAuth consent, registration, setup, and lifecycle actions. | seconds-to-minutes | When touching the admin browser flow; required admin-browser CI job. |
For the admin browser lane, install frozen dependencies in the repository and
admin/, run bunx playwright install --with-deps chromium on Linux, then run
bun run build:admin before bun run test:admin. Tests live in
admin/e2e/*.pw.ts so Bun's unit-test discovery does not execute them. The
browser suite proves the GBrain dashboard journey; it does not establish
activation inside a native vendor harness.
There is no check:all script: a second, hand-synced guard registry would
drift from verify, leaving checks that never run anywhere. The CHECKS
array in scripts/run-verify-parallel.sh is the single execution list
(including check:newlines, check:exports-count,
check:no-legacy-getconnection). The guard REGISTRY is scripts/guards-manifest.tsv (see "Guard registry and
self-test" below).
bun run typecheck uses TypeScript's native incremental analysis in
node_modules/.cache/gbrain-typecheck.tsbuildinfo. Every invocation still runs
the compiler; source, root-file, configuration and dependency changes invalidate
the affected analysis, and cached diagnostics remain failures. The cache is local
and ignored by Git; CI does not restore prior typecheck results.
The local Docker runner isolates root and admin node_modules, plus the generated
admin bundle, in named volumes. Admin build dependencies, Vite's generated cache
and build output stay inside container volumes instead of replacing host files
or leaving root-owned directories behind. ci:local --clean removes these volumes
too; build the admin app on the host when updating its committed bundle.
scripts/ci-ubicloud.ts runs the ci:local gate on ephemeral Ubicloud VMs
instead of one Docker host. It packs the working tree once (tracked files,
untracked files that are not ignored, and .git) and streams it to every VM, so
uncommitted edits are tested. scripts/ubicloud/ubi-runner.sh creates and
destroys the VMs; every VM is destroyed on exit, including Ctrl-C, and any
ubirun-* VM older than 12 hours is garbage-collected by the next run.
Each VM runs scripts/ubicloud/setup-ci-vm.sh: the pinned Bun from
docker-compose.ci.yml, the runner container's test prerequisites plus Node,
frozen dependencies, both PGLite snapshot fixtures, and one
pgvector/pgvector:pg16 server fronted by a transaction-mode PgBouncer per
slot. Every slot's schema is bootstrapped with setupLegacyEmbeddingDB(), the
same step nightly full-corpus E2E workers run, so no E2E file depends on which
file reaches a database first. Setup takes 70-90 seconds, including VM boot.
Scheduling is dynamic. Every unit, serial, slow and E2E file is one item in a
global queue ordered by weight, heaviest first. Each idle slot on any VM takes
the next item, so a slow VM or a mis-weighted file delays only the slot holding
it. Light items leave in same-lane batches to amortize SSH round trips, and the
batch target shrinks as the queue drains. Items of 60 seconds or more are the
run's long poles, so they spread one per VM before any VM takes a second one.
The first VM to finish setup runs the
machine-level work first: gitleaks, verify, then the serial lane's
machine-exclusive files one at a time with nothing else on that VM. After that
it joins the pool. Items run through the ci:local wrappers
(scripts/ubicloud/ci-item.sh): run-unit-shard.sh, run-serial-tests.sh and
run-slow-tests.sh accept explicit file arguments for this purpose, and
run-e2e.sh runs each E2E file against its slot's own server and pooler. The
unit, serial and slow lanes run with database URLs unset. Tests run natively as
a non-root user on Ubuntu 24.04, the same OS as the CI runners, instead of as
root in the oven/bun container.
Weights come from, in order: .context/ci-ubicloud/weights.json (merged after
every run), the committed scripts/ubicloud/weights.json (refresh it with
--record-weights on a green full run), then the lane weight files mined from
GitHub CI. Unknown files get their lane's p75. Per-item logs, failure logs and
summary.json land in .context/ci-ubicloud/<run>/. The exit status is non-zero
when any item fails, an item never produces a result, or no VM becomes ready.
An item whose SSH batch dies without a result is retried once on any slot, and a
VM with three such infrastructure errors is retired.
Defaults are four standard-16 VMs (64 vCPUs) in eu-central-h1 with 8 slots
each, one per two vCPUs (--vms, --size, --slots, --location). The
Ubicloud project's vCPU quota (256) is shared with pull-request CI, so the
default leaves room for about two concurrent PR runs; the former default of
ten VMs took 160 vCPUs and queued PR jobs for up to 28 minutes. A VM that the
quota refuses fails to provision and the run continues on the VMs that did
start, so a busy project shrinks the fleet instead of failing. Pass --vms 10
only when the quota is idle. Slow-lane items run test/export-scale.slow.test.ts
at the pull-request scale (GBRAIN_TEST_EXPORT_SCALE_PAGES=10001). --lanes runs a
subset, --keep leaves the VMs up for debugging, and --diff follows
ci:local:diff (a doc-only diff runs gitleaks alone). The corpus is roughly
8,000 seconds of test compute at that density, so 80 slots finish everything
but the longest files about two minutes after setup; more slots per VM add CPU
contention that slows timing-sensitive files without shortening the run. Wall time is bounded
by setup plus the longest single file,
test/reindex-markdown-persistence.slow.test.ts (one test, about 230 seconds),
so adding VMs past the default does not shorten a run.
scripts/e2e-backend-matrix.txt lists the E2E files that must pass on direct
Postgres and through a transaction-mode PgBouncer: the E5 executor binding
matrix (test/e2e/executor-binding-matrix.test.ts, whose PGLite arm is
test/executor-binding-matrix.test.ts) and every test/e2e/*parity* file.
When GBRAIN_PGBOUNCER_E2E_URL is set, scripts/run-e2e.sh runs each listed
file twice: first against DATABASE_URL with
GBRAIN_TEST_BACKEND=postgres-direct, then with DATABASE_URL set to the
pooled URL and GBRAIN_TEST_BACKEND=pgbouncer. The PGLite arm inside each
parity file runs in both passes. Both passes must execute the same, non-zero
number of tests, and the summary prints the per-backend counts. The pooled
URL must carry ?prepare=false, because resolvePrepare only auto-detects
port 6543 and CI poolers listen elsewhere; the runner refuses a pooled URL
without it. With GBRAIN_CI_REQUIRE_PGBOUNCER=1, a listed file fails when no
pooled URL is configured.
Instead of a full URL, a lane may set GBRAIN_PGBOUNCER_E2E_DB=<name>: the
runner then reaches that database through the pooler in
GBRAIN_PGBOUNCER_URL, pins prepare=false itself, and creates the database
on first use through GBRAIN_PGBOUNCER_DIRECT_URL
(scripts/lib/ensure-e2e-database.ts). ci:ubicloud routes each slot's own
pooler at the slot database, ci:local gives each shard a
gbrain_pooled_<N>_test database behind its single pooler, and e2e.yml
tier1 runs the list against a pgbouncer service. An entry may carry
<TAB>pooled-timeout=<seconds> when its pooled pass needs more than the
per-file cap; !path<TAB>reason records a parity file deliberately left out.
test/scripts/e2e-backend-matrix.test.ts pins the list's completeness, the CI
wiring and the runner's count assertion.
The engine-sql executor (src/core/engine-sql/, refactor wave 1 W1) is pinned
by these tests; the *-parity and RLS files run on every backend in the matrix
above, and each E2E file keeps a PGLite arm in the unit lane.
test/executor-binding-matrix.test.ts/test/e2e/executor-binding-matrix.test.ts: the E5 case table runs twice per backend, throughengine.executeRawand through the dialect adapters (engineSqlExecutorfactory), so the adapters bind, count, fail and cancel exactly like master's raw path.test/engine-sql-executor.test.ts:sqlFragmentrenders the same text and values as the postgres.js tagged template (vendored serializer); Postgres driver options (prepare: true, simple: falsefor converted statements, master's options forexecuteRaw/unsafe), gauge bypass, EO1 transaction lane, brand@ts-expect-errorfixtures.test/e2e/engine-sql-prepare-parity.test.ts:pg_prepared_statementsholds a converted statement on direct Postgres and nothing through PgBouncer; a zero-parameter multi-statement string is rejected on every backend.test/engine-sql-transaction.test.ts/test/e2e/engine-sql-transaction-parity.test.ts: per-domain write-then-throw rollback throughengine.transaction()andtransactionDirect()(dual pool on Postgres), with a concurrent pool read. Add a case totest/helpers/engine-sql-rollback-cases.tsfor every migrated domain write. A mutation that caches the executor on the engine fails it (DISCRIMINATE_BASE=<mutation> bash scripts/check-test-discriminates.sh).test/engine-sql-capabilities.test.ts/test/e2e/engine-sql-capabilities-parity.test.ts: each dialect capability with a boundary-size and a concurrent-write case.test/e2e/engine-sql-normalize-parity.test.ts: every declared column kind ofnormalize.tsdecodes to one shape on each backend.test/e2e/engine-sql-rls-scope.test.ts:ScopedReadreads under a non-ownerNOBYPASSRLSrole (cross-source denial, concurrent isolation, nested rollback restoration, connection reuse).- SQL text:
test/engine-sql-sql-text.test.tsgoldens must stay byte-identical after a conversion; onlysql-text/_driver.jsonmoves (tagged ->runUnsafe).
bun test test/native-lock.test.ts test/scripts/native-lock-prebuilds.test.ts
checks real process exclusion, crash handoff, retained files, cancellation,
missing-addon failure and source/binary manifest integrity. Tests use isolated
temporary paths and never open an operator datastore. The required
native-locks.yml lane rebuilds and executes all eight OS/architecture/libc
targets on Bun 1.4.0 and 1.4.2, including native musl Docker userspace.
Every pair also runs bun scripts/native/compiled-smoke.ts to prove compiled
process locking. Release CI verifies the shipped CLI embeds the matching
addon and runs the compiled smoke on its two release platforms. Rebuild
instructions and the precise packaging/runtime distinction are in
native/locks/README.md.
Release compilation uses Bun 1.4.2; strict Darwin codesign verification must
pass before publication. The native macOS 26.2 smoke is not macOS 27
certification, and Linux fault injection is not a full native Windows backup
create/restore test.
The glibc Linux, macOS and Windows matrix also runs
persistence-publication-native.serial.test.ts,
persistence-git-publication.test.ts,
persistence-sync-origin-native.serial.test.ts and
backup-portability-native.serial.test.ts in separate Bun processes. These
exercise real Git publication, historical source origins and backup restoration
with fresh-process reopen. macOS and Windows set
GBRAIN_TEST_REQUIRE_CASE_INSENSITIVE=1; a case-sensitive fixture must fail
rather than silently skip the publication regression. Inspect executed test
counts before claiming native coverage; a configured lane alone is not evidence.
The publication, sync-origin, processing-option and company-sync suites also
run separately with an explicit PostgreSQL URL in the persistence deployment
matrix. test/e2e/persistence-publication-parity.test.ts and
test/e2e/persistence-sync-{origin,options,company}-parity.test.ts include them
in the full local E2E gate; the keyless serial runner alone cannot exercise
their PostgreSQL branches.
The OpenClaw 2026.9.4 / Node 24.18.0 native-host fixture proves plugin startup, restarted-turn saved-page pointer retrieval and same-slug source isolation with a deterministic loopback provider:
GBRAIN_TEST_OPENCLAW_BIN=<absolute-installed-cli> \
GBRAIN_TEST_OPENCLAW_DATABASE_URL=<isolated-postgres-test-db> \
bun test test/openclaw-context-engine-native.serial.test.tsThe database user needs CREATEDB; fixtures create/drop unique databases
rather than truncating shared rows. Real-provider recall and macOS 27 behavior
remain unverified.
Focused safety coverage: test/apply-migrations-safety.serial.test.ts checks
force dry-run previews before DB/ledger access and failed-phase partial exit;
test/real-home-guard-preload.test.ts pins the test-home fingerprint backstop
(detection, not prevention). Managed retry, durable diagnostics, restart and
PGLite/Postgres parity are covered by test/persistence-sync-failures.serial.test.ts
and test/e2e/managed-sync-failures.test.ts. Backup remote readback and fsync
fault cases run in test/backup-verification.serial.test.ts and
test/backup-fsync.serial.test.ts; test/e2e/backup-coverage-parity.test.ts
covers PGLite/Postgres page/fact/config parity. Output redaction uses
test/search/output-redaction.serial.test.ts and
test/search/output-redaction.test.ts, including unchanged internal capture.
Managed writer fixtures use isolated PGLite and guarded disposable Postgres:
test/e2e/fact-vector-repair-parity.test.ts,
test/e2e/fact-embedding-backfill-parity.test.ts, and
test/fact-backfill-resident.test.ts cover preserved vectors, bounded
NULL-only fact backfill, selected-config refusal and owner-held PGLite IPC;
test/ai/google-embed-batch-items.test.ts pins 100-item provider batches.
test/persistence-embedding-effects.test.ts,
test/persistence-effect-retry.test.ts, and
test/embedding-completion-atomic.serial.test.ts cover partial vector
completion, exhausted durable attempts and state-bound explicit retry.
test/managed-extract-atoms.test.ts, test/managed-facts-backstop.test.ts
and their test/e2e/ counterparts exercise admitted atom/fact replay,
including fresh-process facts authority. test/persistence-connectors.test.ts
covers managed bound/unbound Google/GitHub sources, API pagination and
source-scoped deletions. test/persistence-connector-retry.test.ts covers
explicit retry, compaction, checkpoint dependency identity, concurrent approval
and lost acknowledgements. Each suite creates its own home, engines and
lifecycle through test/helpers/connector-fixture.ts; the helper shares no
live engine or mutable suite state. Their separate E2E entry points,
test/e2e/managed-connector-routing.test.ts and
test/e2e/managed-connector-retry.test.ts, retain the runner's default
180-second per-file cap without duplicating the base cases in the retry lane.
Linux root runners execute the complete EACCES case in an isolated setpriv
child and assert UID 65534 before testing permissions. This needs a readable
checkout, not changes to the parent process identity or checkout permissions;
the CI runner image supplies setpriv.
test/managed-maintenance.test.ts and
test/helpers/maintenance-restart.ts cover local synthesize/patterns/
consolidation, restart replay, retired takes and semantic snapshots;
test/managed-unsupported-preflight.serial.test.ts checks that the managed
facts-family bulk lanes refuse an unaccepted writer before spend, and
test/managed-facts-writers.test.ts proves each of them (fence reconcile,
phantom redirect, fence writes, loops extraction, bulk conversation facts)
publishes through the coordinator on PGLite and Postgres. These use synthetic provider/API transports, not
paid model calls or production connectors. PGLite dream/job CLI with an active
owner is not proven delegated by the live fact-backfill IPC test.
test/facts-worker-config.test.ts and its PostgreSQL E2E counterpart dispose
the original consumer before executing a real facts-absorb job. They verify the
worker passes trusted selected configuration, ignores job-supplied configuration
and settles the entity-page effect with zero fact or chunk embedding calls when
disabled. Fact extraction still captures the generated fact with a NULL embedding.
test/managed-facts-embedding.test.ts and its PostgreSQL counterpart bind retained
fact vectors to the selected brain's model and dimensions, including equal-width
host/mount mismatches, keyless capture, policy changes and replay without new spend.
test/managed-atom-regressions.test.ts and its PostgreSQL counterpart preserve
later target edits through explicit retries and honor database-only storage policy
without relaxing source authority. test/managed-synthesis-postprocess.test.ts
and its E2E wrapper verify that completed quote/provenance work never rewrites a
later user edit, while unfinished work resumes against its original revision.
The synthesis suite also preserves the existing same-date summary on replay and
rebuilds a complete index after partial recovery. test/managed-atom-compaction.test.ts
and its PostgreSQL counterpart age and compact real receipts: permanent completion
identity still prevents repeated extraction, while expired retry payloads produce
an explicit refusal without changing terminal outcomes or compaction accounting.
test/managed-facts-compaction.test.ts and its PostgreSQL counterpart cover the
same lifetime boundary for explicit and derived fact-batch identities, including
failed or partially committed batches and successful replay without new spend.
Connector sweep fencing and physical-path normalization have separate parity
coverage in test/persistence-connector-fencing.test.ts. Standalone crash/recovery
cases live in test/persistence-connector-recovery.test.ts and their own E2E
wrapper so they do not share the routing file's wall-clock budget; their original
assertions, child watchdogs and per-file timeout are unchanged.
test/managed-atoms-cli.slow.test.ts exercises real disk-backed PGLite CLI
recovery with a loopback provider: live-owner refusal, graceful owner stop,
malformed extraction, explicit same-input retry, idempotent replay and owner
restart. Fresh-process readback checks the private canonical file, searchable
chunk, retained failure receipt, committed completion and released leases.
test/managed-connector-routing.serial.test.ts pins actual activation and
performSync routing for API sources; the maintenance suite also drives
runCycle with eligible facts in two sources and proves the other source is
unchanged. The E2E wrapper files ensure these optional PostgreSQL arms execute
in the database lane rather than only passing their PGLite controls.
The persistence invariant jobs run the complete scripts/persistence/validate.ts
gate (10,000-write soak) on pushes to master and manual dispatches. Pull requests
run the same schedules and crash boundaries with a 2,500-write soak; the full
PGLite soak alone takes 15-28 minutes and would otherwise set every PR's wall
time. test/scripts/data-safety-native-workflow.test.ts pins the split. The
crash robot (generated sequences of real operations, SIGKILL at every crash
seam they reach, process faults, the reference model) runs as its own job
beside the soak: 150 s on pull requests, 600 s elsewhere, Postgres through a
transaction-mode PgBouncer. That job also replays the shrunk crash-robot
regressions (test/persistence-crash-robot.slow.test.ts) and the history
fixture test on both engines; see scripts/persistence/README.md.
For platform-only feedback, dispatch
gh workflow run test.yml --ref <branch> -f native_only=true. This explicit manual option uses a separate concurrency
group so it does not cancel an ongoing full persistence soak. Its
native-only-validation-scope artifact records the exact commit and
full_ci: false; it never emits the required test-status check for unrun full
CI. Omitting the option preserves every normal PR, push and full manual gate.
test/pglite-lock.test.ts proves process pause/crash handoff, metadata damage,
legacy migration refusal and stable ownership across datastore replacement.
test/pglite-engine-disconnect.serial.test.ts uses actual disk-backed PGLite
for concurrent opens, consumer/statement drains, persisted reopen, delayed
close and failed close. A close deadline retains the kernel lock; it is never
successful shutdown evidence. Watchdog and telemetry regression suites cover
loop starvation and background statement teardown.
test/db-lock-concurrency.test.ts proves unique identities even when two
acquisitions have identical database timestamps, exact successor-safe cleanup,
renewal cancellation/late-completion drain and mandatory loss propagation.
test/e2e/db-lock-acquisition-token.test.ts repeats acquisition/cleanup
invariants against real Postgres; the E2E map selects it for lease and engine
changes. test/engine-control-routing.test.ts pins direct/shared pool routing,
nested transaction confinement and the Postgres resident-stop barrier.
test/persistence-consumer-scheduling.test.ts pins completion wake-ups,
including a wake-up arriving during an active tick, without lowering the idle
poll interval. Per-root deadlines preserve blocked/retryable backoff even while
another root keeps committing; expired deadlines permit retries. Shutdown drains
active preparation without starting another request.
test/persistence-root-refresh.test.ts checks that unchanged root registrations
do not replace their durable files while unbound, moved and original bound paths
all remain fenced.
test/e2e/persistence-http-liveness.test.ts drives real authenticated HTTP MCP
with legacy and source-bound OAuth credentials against an isolated PostgreSQL
database. It covers pending page/fact writes, lost acknowledgments, owner
interruption, exact canonical readback, 44 pages over three clients, and current
receipt authorization. Set GBRAIN_TEST_OLD_BINARY to a retained compatible
executable to additionally exercise old/new owner handoffs in both directions
for queued, actually claimed and recovery-bearing rows. Interrupted owners are
reaped and native exclusion is verified before their successors start; only the
abandoned running lease is advanced by the fixture. An omitted executable skips
these compatibility cases, not proves them.
test/e2e/persistence-phase-liveness.test.ts holds real table/row locks and
ordinary/direct pool slots to verify phase cancellation (including capacity
marking), tracked queued BEGIN, renewal, expired-head FIFO and shutdown fences.
A loopback TCP gate separately delays or rejects cold direct initialization,
checking retained work through shutdown and same-ID completion after retry.
The memory-mutation tests use a
deterministic loopback embedding transport for slow and aborted preparation;
they do not contact a paid provider. Receipt contract, MCP parser, IPC and CLI
tests pin the same nested allowlist and advisory age policy. Diagnostic tests
exercise both fresh/index-upgrade parity and a genuinely interrupted PostgreSQL
concurrent index build when a PostgreSQL fixture is supplied.
test/persistence-chaos.slow.test.ts and test/e2e/persistence-chaos.test.ts
execute real journal/coordinator schedules and eight SIGKILL publication
boundaries, followed by a small multi-process soak. The Postgres test creates
and drops fresh test databases, requiring CREATEDB on the explicit test URL.
It never truncates the shared E2E database. The reusable
persistence-validation.yml gate runs 1,000 schedules and 10,000 writes per
engine under Bun 1.4.0 and 1.4.2 and uploads actual executed-case manifests.
See scripts/persistence/README.md for
workloads, reruns, performance measurements and the process-crash scope.
test/e2e/persistence-runtime-matrix.test.ts additionally requires the real
transaction-mode PgBouncer fixture. Its 24 cells exercise direct/pooler
connections, enforced RLS under a non-bypass role, ordinary pool sizes 1/2/3,
and shared pools or a separate one-connection direct route. It verifies
reserved short control capacity while bulk connections remain held, then
drains and commits the original request. Ownership cases cover mismatched
successor manifests, stale owners, root replacement under a held kernel
lock, and actual source deletion/recreation. The reusable persistence lane
runs this matrix on both supported Bun versions and uploads its manifest.
The required persistence lane also runs scripts/persistence/performance.ts
on both engines and Bun versions. Three independent instances use the
existing 500-page/200-query read-latency corpus, with public put_page
mutations and actual in-flight interval coverage of at least 90%. Any read
or write failure invalidates the sample. CI uses --informational: the 50%
loaded-versus-idle p99 threshold is advisory, not a merge blocker. Loaded
reads compete with additional write work, so this ratio alone is not evidence
of a change regressing the same workload. Manifests retain the original
threshold verdict, each sample, admission/commit latency, queue age, RSS,
recovery bytes and pool activity.
The CLI without --informational still enforces the threshold. The original
heavy shell entry invokes this harness; its optional STRICT_LATENCY=1
flag affects only the latency threshold, never validity requirements.
Graduation (gbrain migrate --to postgres) is tested against two fixtures.
test/fixtures/graduation/legacy-brain.ts builds a small brain by hand-written
SQL on a fresh schema, and its expected outcomes live in expected.json
beside it (test/graduation-legacy-fixture.test.ts proves the build matches
and that the target checker discriminates).
scripts/persistence/graduation-fixture.ts wraps buildHistoryFixture for
the 1k and 10k history brains, builds keylessly in a child process, and
caches each build as a tarball keyed by the fixture sources, schema version,
seed and size (GBRAIN_GRADUATION_FIXTURE_CACHE, default
~/.cache/gbrain-graduation-fixtures). A restore re-homes paths and owner
stamps and marks sources synced, so the source doctor stays green.
The E2E suites drive the real CLI in child processes:
graduation-crash (SIGKILL at every run and rollback boundary),
graduation-faults (ENOSPC on a tmpfs tablespace, which needs Docker;
password rotation; DDL route mismatch), graduation-clients (older
releases, respawned and resident serve, stale CLI and MCP configs) and
graduation-cli (agent flow, zero-mutation --plan/--status, --force,
PgBouncer through GBRAIN_PGBOUNCER_URL, a NOSUPERUSER role, the 1k round
trip). Kill and pause points come from graduationBoundary() hooks that only
test/helpers/graduation-hooks-preload.ts registers. Older release binaries
are built once per tag under GBRAIN_OLDER_RELEASE_DIR. The crash suite takes
about 40 seconds per case, so run it with GBRAIN_E2E_FILE_TIMEOUT=3600.
scripts/persistence/graduation-ttv.ts records the commands and wall time
from the first plan to a green doctor, the run's phase timings and query
p50/p95 on both engines. The 1k-page gate is five minutes;
tests/heavy/graduation_10k.sh reports the 10k run.
scripts/build-pglite-snapshot.ts (bun run build:pglite-snapshot) bakes a
post-initSchema() PGLite data dir into test/fixtures/pglite-snapshot.tar
plus a version file; PGLiteEngine.initSchema() restores the tar instead of
replaying the embedded schema + all migrations when the env var
GBRAIN_PGLITE_SNAPSHOT points at it. Runners activate it through the shared
ensure_pglite_snapshot helper in scripts/lib/test-env.sh (also home of
detect_cpus and detect_available_mem_mb), sourced by
run-unit-parallel.sh, test-shard.sh, run-slow-tests.sh,
run-serial-tests.sh, run-verify-parallel.sh, and run-e2e.sh (which
re-exports the path as ABSOLUTE — its tests spawn CLI children with varying
cwd); scripts/ci-local.sh calls the builder directly. The helper builds/refreshes the snapshot and
exports the env var, no-ops on GBRAIN_NO_SNAPSHOT=1 or an already-inherited
path, and is non-fatal on build failure — tests fall back to cold init, with
a one-line "active" echo so a silent fallback stays visible in CI logs.
Measured effect: ~3.5x per PGLite-booting file (a cold boot replays every
migration, ~3.1s each on a CI shard). Properties:
-
Idempotent. A hash short-circuit exits in ~40ms when the snapshot is fresh, and REBUILDS a stale one. The hash covers the raw file bytes of the static import closure of
pglite-schema.ts, the schema-migration registry,migrate.tsand the forward-reference bootstrap (engine-sql/bootstrap.ts), pluspglite-engine.ts(src/core/snapshot-schema-inputs.tscomputes the list; no hand list). Imported SQL and handler changes invalidate the fixture; coverage instrumentation does not change the hash.test/snapshot-inputs-closure.test.tschecks that list against an independent TS-AST closure, requires every literal dynamic import in the closure to be classified, and discovers all 13pglite-snapshot-*CI cache keys: identicalhashFilesinputs covering every hash input, each profile restoring its own tar. A failure names the missing file and both workflow files to edit. -
Concurrency-safe. Each profile has its own lock with a PID/token owner and host/process-namespace identity. Only a confirmed dead local owner using the current retirement protocol can be reclaimed. Both normal release and crash recovery retain a nonempty owner tombstone so a delayed observer cannot remove the next builder's lock. Keep those records while builders may run. Live, foreign, ownerless, or older-protocol locks time out without building; callers visibly fall back to cold initialization. Temporary tar/version files are atomically renamed, with the version last.
-
Never authoritative. The loader (
tryLoadSnapshotinsrc/core/pglite-engine.ts) verifies the schema hash AND the embedding shape the snapshot was baked with (dims=/model=lines in the version file) against what this process would create; any mismatch — including a version file without shape lines — warns once and falls through to normal cold init. A wrong fixture can never poison the suite. -
Opt out.
GBRAIN_NO_SNAPSHOT=1skips the build + env export for a run; the migration-replay canary tests clear the env themselves regardless.
Pinned by test/snapshot-shape-guard.test.ts (hash + shape refusal matrix,
imported SQL/handler dependency hash sensitivity).
The builder accepts --profile legacy|default (legacy remains the default).
Legacy uses the unit preload's embedding shape. Default uses the CLI's canonical
embedding shape and writes pglite-snapshot-default.tar plus its .version.
The artifacts, locks, and CI caches are separate. ensure_default_pglite_snapshot
exports an absolute GBRAIN_TEST_DEFAULT_SNAPSHOT; BrainBench applies it only
to CLI children, including run-all. The parent unit process retains its legacy
snapshot. The slow runner and direct BrainBench test invocation prepare the
default profile automatically. GBRAIN_NO_SNAPSHOT=1 clears both paths and
survives test preloads.
scripts/pglite-checkpoint-harness/supervisor.ts reproduces the large-store
PGLite freeze. It spawns worker.ts, which imports 40 KB pages through
PGLiteEngine.transaction() into one long-lived store. It then watches that
process from outside: CPU from /proc, committed pages from the worker's
progress file, and WAL and checkpoint activity from the data directory. A
worker that burns CPU with no committed page and no checkpoint progress for
--stall-sec is reported as wedged. A completed run also asserts that WAL
since the last redo point never exceeded the guard threshold plus the largest
single transaction. --shared-buffers and --max-wal-size scale the store
down with ALTER SYSTEM, so a small machine reaches the same trigger in
minutes:
bun scripts/pglite-checkpoint-harness/supervisor.ts --dir /tmp/h --fresh --pages 3000 --stall-sec 600 --timeout-sec 1800
bun scripts/pglite-checkpoint-harness/supervisor.ts --dir /tmp/h --fresh --pages 1500 \
--shared-buffers 16MB --max-wal-size 160MB --stall-sec 120 --expect-wedge # baseline check--min-store-gb <n> also fails a completed run whose data directory is
smaller than n GiB; the result reports the store size with and without WAL.
On macOS the supervisor samples worker CPU with ps instead of /proc.
.github/workflows/macos-validation.yml runs on a GitHub-hosted macos-26
runner nightly, on manual dispatch, and on pull requests that carry the
macos-validation label. It uses no secrets and has a 90-minute cap. It checks
the device-identity re-stamp (#5604) on real APFS (the step first asserts an
APFS volume with a non-zero birth time and inode; the device-number change
itself stays simulated because a runner never reboots), runs the checkpoint
harness above on a store of at least 2 GiB with the WAL-bound assertion
(#5449), and verifies the latest published signed darwin-arm64 release
binary with codesign --verify --strict and --version against its release
tag (#5286). It also runs the bash 3.2 parse guard under the
runner's /bin/bash, then bun run verify (#5810), so a script or guard that
only works under bash 4 or later fails there. Maintainers with triage or write access apply the label; an
outside contributor whose change touches macOS-specific persistence, locking
or release code asks for it in the pull request. Scheduled and dispatched runs
use the default branch's workflow file, so the label is the way to get this
evidence for a change before it merges. test.yml's security matrix,
release.yml's darwin build and native-locks.yml's darwin cells pin
macos-26 / macos-26-intel rather than macos-latest.
Required CI runs eight weighted unit workers, four serial workers with bounded
per-file pools, and up to four selected E2E workers. E2E selection and exclusions
run once before setup; the resulting file lists are frozen and executed against
separate Postgres services. An explicit empty selection launches no tests;
selection errors, failed workers, cancellations, and unexpected skips fail the
existing aggregate checks. Nightly full-corpus E2E uses four independent
Postgres jobs with the same weighted partitioner and one fresh Bun process per
file, sequential within each job. It does not use the selected-E2E exclusion
list: default discovery includes every test/e2e/*.test.ts and
test/phantom-redirect-engine-parity.test.ts.
Each full-profile worker first initializes its own service schema with the
guarded setupLegacyEmbeddingDB() helper, in a temporary home with provider
keys stripped and local environment-file loading disabled. No partition relies
on a preceding file to create shared tables. Bootstrap is a separate timed CI
step and must be included in end-to-end comparisons.
Refresh after a large test wave or when the longest shard repeatedly exceeds the mean shard execution time by 25%:
bun run weights:mine --lane unit --run <successful-test-run>
bun run weights:mine --lane serial --run <successful-test-run>
bun run weights:mine --lane e2e --run <successful-e2e-run>
bun run weights:mine --lane e2e --e2e-profile full --run <successful-full-corpus-run>The miner accepts --from-file or stdin for timestamped GitHub-format logs and
--out for inspection before replacing a checked-in map. Unit timing uses only
unit matrix jobs, includes evals/, and closes the final file at the Bun summary.
Serial timing uses runner durations, never timestamps of buffered output. E2E
uses each file's Bun summary and merges partial selections into known weights.
File/stdin imports also merge unobserved entries; only a complete GitHub unit,
serial or explicit full-profile E2E run replaces that lane's entire map. Captured artifacts need their final
successful completion marker, and GitHub imports verify every expected job.
Incomplete or failed inputs leave the existing map intact. Sidecar metadata
records the source run/commit, units, and counts. Unit/E2E weights are milliseconds;
serial weights remain seconds. New files receive the corpus p75 estimate. Empty
maps and zero-cost ties distribute files deterministically; corrupt serial
weights warn and retain safe fallback scheduling.
The default E2E miner reads selected jobs and merges partial observations.
--e2e-profile full instead requires a successful GitHub run with a source SHA
matching checkout HEAD. It pins the run attempt and full-job IDs, reconstructs
default discovery using the committed runner and tracked test paths, and
requires every source file exactly once across complete successful job logs.
Missing/extra files, ambiguous basenames, duplicate execution or failed evidence
leave weights and metadata untouched. Basename resolution preserves the
outside-directory parity entry. Full mode replaces the complete map and records
source SHA, attempt, jobs, log hash and corpus hash; file/stdin imports cannot
claim full-profile provenance. These weights schedule work; they are not an
observed parallel runtime.
CI retains timestamped unit/E2E logs, frozen E2E selection, and serial attempt records for 14 days. Compare push-to-required-green time including queueing, first failure, rescues/reruns, runner minutes, unique file counts, and coverage completeness. Compare cold and warm caches separately. Snapshot timings and partition estimates are projections until matched workflow runs confirm them; successful test results are never cached.
Outputs captured on master before refactor wave 1 moves any code live under
test/fixtures/goldens/; test/fixtures/goldens/README.md maps each file to
its owning test and named normalizer. test/helpers/golden.ts writes and
compares them (expectGolden) and proves each normalizer by capturing twice
(expectNormalizerStable). Regenerate only deliberately, never in a refactor
commit: GBRAIN_TEST_UPDATE_GOLDENS=1 bun test <file> (the switch carries the
GBRAIN_TEST_ prefix because the unit preload scrubs other GBRAIN_*
overrides). Performance baselines are a bench, not a test:
docs/designs/refactor-wave-1/perf-baseline.md.
gbrain doctor runs DOCTOR_CHECK_REGISTRY (src/commands/doctor/registry.ts)
in order: one { name, emits, run(ctx) } entry per topic block under
src/commands/doctor/checks/, each returning its checks or STOP_DOCTOR.
test/doctor-registry.test.ts fails with a FAIL / Why / Fix / See
block when an entry's name or any emits[] name is missing from
src/core/doctor-categories.ts, when emits[] differs from what the entry's
run can push (AST walk in test/helpers/doctor-registry-ast.ts), or when a
STOP gate moves away from where master's buildChecks returned early.
test/doctor-mode-matrix.serial.test.ts wraps every entry and the engine
with recorders and asserts, per mode (default, --fast, --fix,
--fix --dry-run, no engine, connection failure), which entries ran, where
the run stopped, which engine calls happened and which mutations landed (the
SKILL.md DRY auto-repair, the dead-holder lock reap). The W0 registry,
early-stop and --json goldens pin the output itself.
scripts/verify-move-only.ts proves a commit tagged Move-Only: yes moves code
without editing it: every top-level statement of every touched TS file on the
base side reappears token for token on the head side (tokens from
scripts/lib/normalize-tokens.ts, so whitespace and comments are ignored and
string/SQL text is exact). Imports, export ... from lines and toggling the
export modifier on a moved statement are allowed and counted. Run
bun scripts/verify-move-only.ts <commit> (default HEAD~1..HEAD);
--wrapper migration inlines export const vNNN: Migration = {...} files into
the generated registry array so the W3 split must reproduce the base side's
single MIGRATIONS literal entry for entry; --wrapper doctor-entry inlines each
run<Topic>(ctx: DoctorContext): Promise<Check[]> body (minus its ctx
destructure / connectedEngine / const checks prologue and return checks;)
at its checks.push(...(await runX(ctx))); call in buildChecks, drops the
const ctx: DoctorContext = {...}; glue and resolves relative import() /
require() specifiers to repo paths, so the W4 doctor peel must reproduce the
original buildChecks body; and --rename-map <json> applies identifier rewrites for
Mechanical-Rename: yes commits. Failures print FAIL: <file:line> with the
first differing token. Pinned by test/scripts/verify-move-only.test.ts.
Schema migrations live one per file in src/core/schema-migrations/v<NNN>-<name>.ts
(NNN zero-padded to 3, name = the slug with - → _, one
export const v<NNN>: Migration = {...} per file). bun run new:migration <snake_name>
scaffolds the next version; bun run build:schema-migrations regenerates the committed
static-import registry registry.generated.ts (regenerate, never hand-merge). The
array order is master's historical order (HISTORICAL_ARRAY_ORDER in
scripts/build-schema-migrations.ts), then ascending; the runner sorts by version.
Two guards run in bun run verify:
check:schema-migrations(scripts/check-schema-migrations-fresh.sh) regenerates the registry into a temp file and diffs it; the generator also fails on a filename/version/name mismatch and on a version defined twice, naming both files with thegit mv+version:+ regenerate recipe.check:schema-migration-order(scripts/check-schema-migration-order.ts) fails when a migration origin/master does not have is numbered at or below origin/master's latest version (it would be skipped forever on current brains) or reuses a version with a different name. Base ref:GBRAIN_MIGRATION_BASE_REF(defaultorigin/master); skipped with a notice when the ref is missing, failed underCI=true.
Collision recovery: an unapplied branch migration is renumbered (git mv, edit
version, regenerate); one already applied to a disposable dev DB means rebuilding
that DB and replaying; one applied to retained data needs explicit schema_version
reconciliation, never just a counter edit. Pinned by
test/scripts/build-schema-migrations.test.ts and test/migrations-golden.test.ts.
check:schema-fresh (scripts/check-schema-fresh.sh) runs scripts/build-schema.ts --out-dir <tmp> (fragments -> src/schema.sql regions -> schema-embedded.generated.ts
-> pglite-schema.generated.ts) and diffs every output, naming the source to edit.
Canonical sources and PGLite capability rules: docs/ENGINES.md#canonical-schema-sources.
Pinned by test/scripts/build-schema.test.ts; the end state by the E4 catalog goldens.
The privacy and test-isolation guards use scripts/lib/guard-candidates.sh to
scan fresh file contents in bounded batches before applying their detailed
per-file rules. They do not cache passing results. Candidate scanner failures
fail the guard, and matching files retain the same allowlists and diagnostics.
scripts/guards-manifest.tsv is THE single registry of scripts/check-*
guards (currently 68), each classified scanner (greps/parses repo sources —
must eventually carry fixtures), buildfresh, or repostate (build/freshness
guards are exempt-with-reason, not fixture-tested).
scripts/guard-self-test.sh (bun run check:guard-self-test, wired into
bun run verify) proves every selftest=yes scanner CAN fail: it runs each
one against known-bad (must exit non-zero) and known-good (must pass) fixture
trees under test/fixtures/guards/<guard>/{bad,good}/ via the
GBRAIN_GUARD_ROOT env seam, and enforces manifest completeness — a new
scripts/check-* script that isn't registered in the manifest fails the
build. A guard whose pattern rots into a permanently-green no-op fails CI
instead of masquerading as coverage.
A guard may carry extra known-bad trees named bad-<variant>/; each one must
fail on its own. Refactor wave 1 uses them to prove that every scanner naming
a file the wave splits also scans the new module locations
(src/core/engine-sql/, src/core/schema-migrations/, src/commands/sync/,
src/commands/doctor/checks/, src/commands/serve-http-*.ts,
src/core/minions/handlers/): check-jsonb-pattern.sh,
check-engine-dynamic-import.sh, check-source-config-leak.sh,
check-no-legacy-getconnection.sh, check-operations-filter-bypass.sh,
check-source-id-projection.sh (engine-sql) and check-search-path.sh (the
generated PGLite template) each have a bad fixture placed inside the new path. The checklist
of every script, workflow, helper and doc that names a split file is
docs/designs/refactor-wave-1/path-consumers.md.
scripts/check-layering.ts (bun run check:layering, in bun run verify)
parses every file under src/core/engine-sql/ and src/core/schema-migrations/
and fails on any import, type-only included, of an engine façade
(pglite-engine.ts, postgres-engine.ts, engine-factory.ts) from
engine-sql, or of src/core/migrate.ts from schema-migrations. Those
directories are loaded by the engines and by migrate.ts, so an import back
up is an ESM cycle that can fail with a temporal-dead-zone error at module
load. Take the executor as a parameter and import types from
src/core/engine.ts; migration helpers live in schema-migrations/helpers.ts
and the Migration type in schema-migrations/types.ts. Fixtures:
test/fixtures/guards/check-layering.ts/; forms are driven in
test/scripts/layering.test.ts.
scripts/check-durable-flush.ts (bun run check:durable-flush, in
bun run verify) fails on an fsyncSync(fd) anywhere in src/ outside
src/core/fs-durable.ts whose fd is assigned from a read-only openSync
(flags omitted, a flag string without w/a/+, or O_RDONLY without
O_WRONLY/O_RDWR), file or directory, and on one whose flags it cannot
read. Windows refuses fsync on a read-only handle and has no directory flush
(EPERM), which wedges the managed write queue (#5595) and every skill-bundle
publication (#5475). Flushes of descriptors opened for writing pass. Each
failure prints FAIL [durable_flush_read_handle]: <file>:<line>, the open it
traced, a Fix: line and this anchor. Fix: fsync the descriptor you wrote
through before closing it (set its final mode with fchmodSync(fd) first), or
call flushFile(path) / flushDirectory(path, { bestEffort? }) from
src/core/fs-durable.ts. A file that cannot migrate yet goes in the guard's
ALLOWLIST with a reason (empty today); an entry whose file no longer needs it fails as
durable_flush_stale_allowlist. Fixtures:
test/fixtures/guards/check-durable-flush.ts/; forms are driven in
test/scripts/durable-flush-guard.test.ts. The helper and the #5595/#5475
regressions run natively on the windows-latest row of the test.yml
security-regressions job; test/helpers/win32-flush-semantics.ts makes them
discriminate on POSIX hosts too.
scripts/check-engine-sql-ratchet.ts (bun run check:engine-sql-ratchet, in
bun run verify) keeps each storage domain's SQL in one place,
src/core/engine-sql/<domain>.ts, by stopping SQL from growing back into the
engines. It parses src/core/pglite-engine.ts, src/core/postgres-engine.ts
and every file under src/core/pglite-engine/ and src/core/postgres-engine/,
and names each class member (PostgresEngine.getPage), top-level function and
top-level variable (insertFact). A unit is SQL-bearing when the literal text
of a string, template, tagged template or + chain inside it has SQL
structure: SELECT ... FROM <x>, SELECT <fn>(, INSERT INTO <x>,
UPDATE <x> [alias] SET, DELETE FROM <x>, WITH <x> AS (,
CREATE|ALTER|DROP <object kind>, TRUNCATE <x>, SET LOCAL <x>,
ON CONFLICT, WHERE ... ORDER BY|GROUP BY|LIMIT, or a $<n>::type cast.
Comments and identifiers never count. Keywords match in upper or lower case
but never Title Case, and a lowercase match also needs a second SQL signal
(where, returning, $1, ::, ;, * and similar), so "Select a file"
or "could not delete from cache" is not SQL.
scripts/engine-sql-baseline.tsv lists migrated<TAB><domain> rows (the
domain's module must exist under src/core/engine-sql/) and
method<TAB><path><TAB><QualifiedName> rows for the SQL-bearing members that
remain. Rows only shrink. The guard fails on:
- a new SQL-bearing member with no row: move the SQL into
src/core/engine-sql/<domain>.tsand delegate, or mark the declaration (on its line or the line above) with// engine-sql-ok: <reason>; an empty reason fails; - a stale row, whose member is gone, no longer SQL-bearing or now marked:
delete it, or run
bun scripts/check-engine-sql-ratchet.ts --prune, which drops stale and duplicate rows and never adds one; - a duplicate or malformed row, or a
migratedrow with no module.
When a domain moves, delete its members' rows and add its migrated row in
the same commit. Fixtures: test/fixtures/guards/check-engine-sql-ratchet.ts/;
forms are driven in test/scripts/engine-sql-ratchet.test.ts.
scripts/check-engine-sql-dynamic.ts (bun run check:engine-sql-dynamic, in
bun run verify) parses every file under src/core/engine-sql/ except
fragment.ts, the renderer, which writes $n and splices trusted text by
design. In engine-sql every value reaches SQL as a bound parameter through
sqlFragment, and only constant text is spliced. Trusted text is a string
literal; a const in the same file initialized with trusted text or an
as const object or array literal (members and element accesses included); a
CONSTANT_ALLOWLIST name (ENRICH_ORDER_SQL); a call to a VETTED_BUILDERS
entry (pageReadFilter, buildRecencyComponentSql,
privatePagesFilterFragment, currentCodeEdgeFilter, buildCJKKeywordSql,
currentTextProjectionFilter); a template or + chain whose parts are all
trusted or are numbers the same function checked earlier with
Number.isFinite(<same expression>); or a conditional whose branches are both
trusted. Both registries live in the script with a one-line reason each. The
guard fails on:
trustedSql(arg)with an arg that is not trusted text: bind the value with${value}insqlFragmentinstead, or register a new builder with its reason after review;- an untagged template or
+concatenation, passed directly or through a local variable (letappends included) as the SQL of.query(,.unsafe(,.executeRaw(orexecuteRawJsonb(, with an untrusted part: compose withsqlFragmentand run it withexecutor.run(fragment); - a literal
$<digit>, or a$right before a substitution, in a composed string (a template with substitutions, anysqlFragmenttemplate, any+operand): letrenderFragmentnumber the parameters. A static string passed as-is may carry$1; - an expanded list,
IN (right before a substitution or a non-literal+operand: bind the array as one parameter,= ANY(${ids}::text[]), so prepared-statement caches stay bounded.
Fixtures: test/fixtures/guards/check-engine-sql-dynamic.ts/; forms are
driven in test/scripts/engine-sql-dynamic.test.ts.
scripts/check-engine-sql-brands.ts (bun run check:engine-sql-brands, in
bun run verify) keeps the RLS read brands in
src/core/engine-sql/brands.ts unforgeable. ScopedRead records a read that
ran inside withScopedReadTransaction on master and LegacyUnscopedRead one
that ran unscoped on the pool (EO4), so a forged brand silently changes how a
read is scoped. The guard fails on:
- a brand key (any
__obtainVia...name) in a text file undersrc/,test/orscripts/other thanbrands.tsand the guard's own script, fixtures and test: get a branded executor fromscopedRead(tx)insidewithScopedReadTransaction, or fromunscopedExecutor(executor, '<reason>'); - in
src/, a cast ontoScopedReadorLegacyUnscopedReadoutsidebrands.ts, anas unknown as TwhereTnamesSqlExecutor,ScopedReadorLegacyUnscopedRead, or a double cast passed straight toscopedRead(orunscopedExecutor(. Driver-handle casts such astx as unknown as PgConnindialect-postgres.tspass; - an import of
unscopedExecutororLegacyUnscopedRead(value, type, alias, re-export orimport('...').Xtype) from outside engine-sql, the two engine façades, doctor (src/commands/doctor.ts,src/commands/doctor/**,src/core/doctor*), maintenance (src/core/maintenance/**), admin (src/commands/admin*.ts,src/core/admin/**), migrations (src/core/migrate.ts,src/core/schema-migrations/**,src/commands/migrations/**) andtest/; an import ofscopedReadfrom outside engine-sql, the façades andtest/; or a namespace, dynamic orrequireimport ofbrands.tsfrom outside thatscopedReadlist.src/core/ops/**, the MCP-facing surface, is always denied. Take the branded executor from the engine façade instead.
The allowlists live in the script. Fixtures:
test/fixtures/guards/check-engine-sql-brands.ts/; forms are driven in
test/scripts/engine-sql-brands.test.ts.
scripts/check-retired-phrases.sh (bun run check:retired-phrases, in
bun run verify, under a second) keeps the instructions agents follow
literally in step with refactor wave 1. It greps CLAUDE.md, AGENTS.md,
CONTRIBUTING.md, docs/ and skills/ for the contributor-workflow phrases
the wave retired: the old migrations-array wording and appending to it, the
rule that every engine method is written twice, the CLI switch-case step, the
migrate.ts region policy, and schema text listed as a hand-synced pair of
schema.sql and the PGLite schema module. The patterns and the current
instruction for each live in the script's RETIRED table. Historical records
may quote them and are exempt: docs/designs/, docs/test-audit/,
docs/incidents/, docs/plans/, docs/proposals/, docs/research/,
docs/issues/, docs/superpowers/, docs/migrations/, skills/migrations/
and the wave 1 porting kit (docs/architecture/wave-1-*); CHANGELOG.md is
not scanned. Each hit prints FAIL: <file:line> retired phrase "<match>",
then Why:, Fix: with the current instruction, and See:. The fix is to
rewrite the sentence to the current workflow, never to exempt the file. To
retire another phrase, add a row to RETIRED. Fixtures:
test/fixtures/guards/check-retired-phrases.sh/ (one bad-<location> tree per
scanned location); every pattern and the exemptions are driven in
test/scripts/check-retired-phrases.test.ts.
macOS ships GNU bash 3.2.57 as /bin/bash, and its parser rejects shapes
bash 5 accepts. The one that broke every Mac (#5810) is a heredoc inside
$(...) whose body holds an odd quote. scripts/check-bash32.sh
(bun run check:bash32) runs the real 3.2 parser, bash -n, over every
tracked *.sh except the guard fixtures under test/fixtures/guards/. It
uses the first parser available: GBRAIN_BASH32=<path> (a bash 3.x binary),
/bin/bash when it is bash 3.x (stock macOS), or the digest-pinned bash:3.2
Docker image (GBRAIN_BASH32=docker forces the image). With none it prints
one skip line and exits 0; GBRAIN_BASH32_REQUIRE=1 makes that exit 2. Each
failure prints FAIL: <file:line>, Why: (macOS /bin/bash is 3.2), Fix:
(read heredoc text with IFS= read -r -d '' VAR <<'EOF' || true) and See:.
To reproduce one file by hand:
docker run --rm -v "$PWD":/w -w /w bash:3.2 bash -n <file>.
The guard checks parsing only. It is not in bun run verify, which must not
need Docker. The test.yml verify job runs it with GBRAIN_BASH32_REQUIRE=1
together with test/scripts/check-bash32.test.ts, whose real-parser cases
feed it the test/fixtures/guards/check-bash32.sh/{bad,good} trees. The
macOS 26 job runs it under /bin/bash and then runs bun run verify there,
which covers bash-4 runtime features the parser cannot see.
scripts/check-test-placeholders.mjs (bun run check:test-placeholders, in
bun run verify) parses every test/**/*.test.ts file outside
test/fixtures/ with the TypeScript compiler API and fails on the no-op forms
expect(true) with no matcher, expect(true).toBe(true),
expect(true).toBeTruthy() and expect(1).toBe(1). Text inside strings and
template literals is ignored, and expect(true).toBe(false) fail sentinels
are allowed. Remaining sites (type-only contracts enforced by typecheck,
skip-arm markers, gates that fail by throwing) sit in a reasoned allowlist in
the script, keyed by file, test name and exact count; a site above its count
fails as new, and an entry whose file, test or count shrank fails as stale.
This is a hygiene check for one pattern, not a detector of low-value tests in
general; the authoring gate above owns that.
scripts/check-function-size.ts (bun run check:function-size, in
bun run verify, about 1.5 s) measures every function-like node in
src/**/*.ts except *.generated.ts and .d.ts with the TypeScript compiler
API: function declarations, methods, constructors, accessors, arrow functions
and function expressions, including object-literal and class-property forms.
A nested function is measured on its own, and its lines also count toward the
function that contains it. Code under test/ is out of scope.
scripts/function-size-baseline.tsv holds one row per function over 300
lines: path, name, lines, justification. The name is a path built from
declarations, property names and call context, never line numbers, so edits
above a function do not touch its row: PGLiteEngine.initSchema,
runServeHttp>app.post('/mcp'), MIGRATIONS[v131].handler. > enters a
function, . a member, = a call whose result is bound, and a repeated key
gets a #2 ordinal. The guard fails when a function over 300 lines has no
row, a baselined function grows, a baselined function drops to 300 lines or
fewer (remove the row), a row has more than 50 lines of stale slack (lower
it), a row names a function that no longer exists, or a row is malformed,
duplicated or out of order. A row raised above, or added since, the baseline
at the merge-base with origin/master needs an issue or TODO id (#1234,
TODOS.md:12, TODO: <slug>) in its justification; the summary prints every
raise.
Each failure prints FAIL: <file:line> <what> with the computed key, then one
Why: / Fix: / See: block. The fix is extraction: move a cohesive block
into a named helper or sibling module (phase, stage or handler-table pattern).
After a move-only commit changes a function's key, run
bun scripts/check-function-size.ts --transfer. It rewrites a missing row to
the one unbaselined over-limit function whose whitespace-normalized text is
identical to the old function at HEAD (--from <ref> for another base),
apart from an added leading export and module specifiers re-relativized to
the new directory (each resolved against its own file, so a retargeted
specifier still refuses), keeping lines and justification, and leaves
everything else for review.
Fixtures: test/fixtures/guards/check-function-size.ts/{bad,good}; every rule
is driven in test/scripts/check-function-size.test.ts.
scripts/check-sync-run-state.ts (bun run check:sync-run-state, in
bun run verify, well under a second) protects the refactor wave 1 SyncRun
rule (A17). SyncRun (src/commands/sync/sync-run.ts) holds the state one
incremental sync shares between closures that interleave across awaits: the
checkpoint flush and its cadence, the import workers, the stall watchdog and
the partial exit. Its mutable fields are the members of interface SyncRun
not marked readonly. Over src/commands/sync/**/*.ts the guard fails when a
mutable field is destructured from a SyncRun value (const { bankedFiles } = run, or a { checkpointDead }: SyncRun parameter) or copied into a local
(const banked = run.bankedFiles), because such a copy goes stale at the next
await. A SyncRun value is a binding named run, annotated SyncRun, or
initialized from createSyncRun(). Readonly fields (collection references,
fixed configuration) may be destructured. Fields tagged @checkpoint in their
JSDoc (the flush cadence, banked count, single-flight flag, dead flag, SIGTERM
deregistration and yield counter) have one owner: only functions in
sync-run.ts may assign them, so the flush, the SIGTERM hook and partial()
cannot disagree about checkpoint state. Each failure prints
FAIL: <file:line> plus Why: / Fix: / See:; the fix is to use
run.<field> at each read and write, and to change checkpoint state through a
sync-run.ts function. Fixtures:
test/fixtures/guards/check-sync-run-state.ts/{good,bad,bad-alias,bad-param,bad-owner}.
test/test-reads-source-smell.test.ts finds test code that reads src/ text:
readFileSync, readFile (including fs.promises.readFile) and Bun.file
calls whose arguments name a src/ literal, a 'src' path segment, or a
constant holding such a path. Each read site needs a tagged marker on its line
or within the three lines above:
// test-reads-source-ok[structural]: <why a source read is the right tool>The category is one of prompt-byte, trust-boundary, generated-artifact,
structural or raw-bytes, and every marker must carry one. Files that
predate the rule are ratcheted by their exact count of unjustified read sites,
so a new untagged read in such a file fails and a count that drops must be
lowered. The ratchet counts read sites only: a new assertion over an existing
source binding is not detected and remains the authoring gate's job. Rerun with
bun test test/test-reads-source-smell.test.ts.
Structural guards over the files that refactor wave 1 decomposes read them
through test/helpers/source-surface.ts rather than readFileSync. A surface
is one façade plus the modules it is split into (sync, cli, serve-http,
jobs, hybrid, autopilot, migrate, pglite-engine, postgres-engine,
doctor). surfaceSource(surface) concatenates the surface with file
boundary markers and serves containment assertions (toContain,
not.toContain, single-line regexes). surfaceFileSource(surface, path)
returns one named file and serves positional assertions (indexOf ordering,
slice windows, [\s\S] spans, line math); a file outside the surface
throws. A lane that moves code adds the destination to the surface in the
same commit: a new directory is globbed automatically, while a module in an
existing directory or flat file set is listed explicitly so today's
assertions are not widened. test/helpers/doctor-source.ts is the doctor
instance of the same loaders.
Structural suites that walk a registry so the NEXT gap of a known class cannot ship silently. All allowlists below are shrink-only unless noted.
test/operations-coverage-ledger.test.ts— every op insrc/core/operations.tsmaps to a covering test file in a checked-in ledger; theUNCOVEREDallowlist only shrinks. Shares one registry-enumeration helper (test/helpers/ops-registry.ts) with the jobs-ops token-redaction sweep so two walkers can't drift.test/operations-source-isolation-matrix.test.ts— every non-localOnly read op runs under a scoped remote ctx and a federated grant; nothing carrying the other source's identity may return. Deliberate brain-wide behavior requires an explicitBRAIN_WIDE_READSentry with a rationale string. Anti-vacuity is mandatory: each op's control call must SEE the cross-source marker before its scoped assertions count; an op that can't be driven is an explicit counted SKIP disposition, never a silent pass.test/scripts/e2e-wiring.test.ts— everytest/e2e/*.test.tsmust be claimed by a PR-time lane (ascripts/e2e-test-map.tsrow, a workflow mention, or the shrink-onlytest/fixtures/e2e-unmapped-baseline.txt), and every map entry must point at a real file (typo guard).test/engine-surface-coverage.test.ts— two-way census of theBrainEngineinterface against the PGLite prototype (new methods force a visible list edit) plus a runtimeUNCALLEDratchet scanning the whole test corpus for references, so a never-called engine method can't ship.scripts/check-orphan-modules.mjs(verify battery, guard-manifest registered with bad/good fixtures) — transitive import walk from the cli/mcp/engine entrypoints; see Orphan-module guard.
bun run check:orphan-modules walks static, dynamic and require relative
imports from the runtime entrypoints (CLI, MCP server, plugin engines, admin,
package exports). Every src/ module it cannot reach needs a disposition:
- Imported by nothing, not even tests: fails as
hard-orphanunless it has a reasonedALLOWLISTentry (shrink-only). - Imported only by tests (or scripts): fails as
unpermitted-test-onlyunless it is named inPERMITTED_TEST_ONLYwith areason. The set may grow only with a reason in a reviewer-visible edit; modules reached fromscripts/**use reasonscript-reachable, which the guard verifies. - A permitted entry whose module was deleted, wired into a runtime
entrypoint, dropped by every test, or tagged
script-reachablewithout ascripts/**importer fails asstale-permitted-entry. Remove or correct the record; never restore code to satisfy the list.
Each failure prints the rule, the module, the tests that import it, the
reason, the remedy, the rerun command and this anchor. Fixture mode
(GBRAIN_GUARD_ROOT) reads the permitted set from
<root>/permitted-test-only.json; test/scripts/check-orphan-modules.test.ts
proves every rule fails on a bad tree.
The takes-bootstrap graduation instrument (evals/takes-bootstrap/: 123-case
corpus, scorer, live harness + $0 replay) is CI-guarded keyless by
test/eval-takes-bootstrap.test.ts — the guard proves the instrument, not
the score; the autopilot tier flips only on a committed GRADUATED live run.
All four of test, verify, ci:local and test:e2e hand off to shell scripts
under scripts/, so every check:* entry in package.json invokes its script as
bash scripts/<name>.sh instead of relying on the shebang — bun on Windows cannot
exec a .sh directly. Add a new shell-script check with that same prefix. The
scripts/*.ts entries run under bun and take no prefix.
The scripts must also be on disk with Unix line endings. A strict bash (WSL, Linux
CI, macOS) rejects CRLF and dies on the script's first meaningful line; the Cygwin
bash that ships with Git for Windows tolerates it, so a green local run is not by
itself evidence that a script is CRLF-clean.
The root .gitattributes pins *.sh text eol=lf, which overrides the
core.autocrlf=true default that Git for Windows installs. It pins *.md the
same way, because the frontmatter readers anchor on a --- fence followed by a
Unix line ending and a CRLF checkout makes a document parse as having no
frontmatter, silently. Working copies cloned
before those pins need a one-time git rm --cached -r . -q && git reset --hard to
pick them up; see the Windows section of CONTRIBUTING.md.
Wallclock figures in the table above are from a Mac dev box. Windows is
substantially slower because each check pays full process-creation cost, and three
tree-walking checks (check:privacy, check:test-names, check:test-isolation)
plus typecheck can exceed the 120s per-check cap in run-verify-parallel.sh
there even though they pass on Linux and macOS.
- CI matrix (
.github/workflows/test.yml) runsscripts/test-shard.shacross 8 matrix shards partitioned by weight-aware LPT bin-packing (scripts/sharding.ts; files with no mined weight fall back to the p75 file weight so a new unweighted file can't silently unbalance a shard) and INCLUDES*.slow.test.ts(the dedicated slow files — longmemeval, entity-resolve-perf, entity-card-perf, export-scale, brainbench-e2e — run as dedicated jobs alongside the matrix, andreconcile-crash.slow.test.tsruns only in the persistence-validation reconciliation job) plusevals/**/*.test.ts(keyless-allowlist-gated —test/scripts/evals-collection.test.ts). Each shard's bun process is bounded by--max-concurrency(GBRAIN_TEST_MAX_CONCURRENCY, default 4). Every bun-test job — matrix shards, serial-tests, verify, the slow/eval jobs — activates the PGLite schema snapshot (built in-runner viascripts/lib/test-env.sh; the BrainBench gate uses the separate default-profile snapshot for its in-memory PGLite; the ~42MB tar is also cached across jobs via actions/cache, with the runner's own hash check staying authoritative). CI EXCLUDES*.serial.test.tsfrom the shards and runs them across fourserial-testsworkers viabun run test:serial— one bun process per file preserves themock.modulequarantine; the pool runs those processes concurrently.bun run verifygets its own job too, as does the BrainBench memory-conformance gate (brainbenchjob →scripts/ci-brainbench-gate.sh, hermetic in-memory PGLite, ~15s), which compares HEAD's fresh run against master's committed baseline (evals/brainbench/baselines/main.json) — thetest-statusaggregate checks its result explicitly. E2E (.github/workflows/e2e.yml) always runs its applicable execution lanes, with the jsonb-parity job in front of tier2 as the token-spend gate, and aggregates throughe2e-status. Scheduled runs also require the full-corpus lanes, including each slow suite excluded from the coverage shards (longmemeval, entity-resolve-perf, entity-card-perf, brainbench-e2e, export-scale and reconcile-crash). Both aggregates reject failures, cancellations, and unexpected skips. Dependency caches and validated PGLite snapshots remain; successful test results are never reused. CI is the ground truth for "did everything pass." - Local fast loop (
scripts/run-unit-shard.shvia the parallel wrapper) uses the same weighted partitioner as CI and EXCLUDES*.slow.test.tsAND*.serial.test.ts. Each shard runs its complete ordered selection with a fresh Bun process per file, without adding workers. Later groups still run after failures; missing summaries or file-completion evidence fail the shard. Local trades coverage for inner-loop speed; CI catches what local skips.
This divergence is intentional. Don't try to make them equal — the two scripts deliberately solve different problems. The regression test at test/scripts/run-unit-shard.test.ts pins what the local fast loop should and shouldn't include, and that no unit-lane file spawning the CLI through test/helpers/cli-spawn.ts hand-pins a per-test timeout below the bunfig default (an explicit test(name, fn, N) ceiling overrides bun's --timeout, so GBRAIN_TEST_TIMEOUT_MULTIPLIER never reaches it — inherit the default instead; cli-spawn's own kill timer still reaps a hung child); test/scripts/run-unit-parallel.test.ts pins the wrapper's memory-adaptive concurrency, and the OOM/external-kill serial rescue pass, and operator-interrupt teardown (a Ctrl-C / SIGTERM to the wrapper while shards are live TERMs then KILLs every shard descendant, so a cancelled run cannot leave gtimeout/bun alive until the shard cap).
Line coverage is opt-in via COVERAGE_DIR: when set, the shell lanes
(scripts/test-shard.sh, scripts/run-serial-tests.sh, scripts/run-e2e.sh)
pass --coverage --coverage-reporter=lcov to bun; when unset, the exec line is
byte-identical to a non-coverage run. Every bun process gets its OWN coverage
dir ($COVERAGE_DIR/shard, serial-$idx, e2e-$idx) because a reused dir
silently overwrites lcov.info — the shard runner also pins xargs to a single
batch (-n 100000 -x) so an argv overflow fails loud instead of spawning a
second, overwriting bun process. On a green run each lane writes
$COVERAGE_DIR/lane-manifest.json ({lane, sha, lcovCount, complete}); a red
run writes no manifest, which downstream merging treats as an incomplete lane.
E2E shards use distinct e2e-1 through e2e-4 lane names (unsharded runs use
e2e), with an executed-files.txt receipt written only after every selected
file completes its native Bun report successfully. The runner requires a fresh
parent-owned JUnit report for the selected file and matching final console
pass/fail/skip totals. Python 3's standard XML parser validates the document
and checks its actual testcase counts against each suite and the console;
a zero exit without that report cannot borrow a nested child's summary as
completion evidence. Each coverage invocation atomically creates its own
COVERAGE_DIR; an existing destination, even an empty one, is refused without
changing its contents. Use a new path for a rerun. Failed or cancelled runs
write no completion receipt, and shorter or skip-only runs cannot inherit old
coverage or delete another run's outputs.
Skip-only files can emit no LCOV, so execution-file counts and LCOV counts are
intentionally different measures.
run-e2e.sh specifics: COVERAGE_DIR is normalized to an absolute path
against the repo root before HOME moves (the script redirects
HOME/GBRAIN_HOME and E2E tests spawn CLI subprocesses with varying cwd —
an un-normalized relative dir would scatter output), and E2E_FILE_TIMEOUT_SECS
caps each file's wallclock (default 180s; the nightly coverage lane uses 300s
for instrumentation overhead). Both env names are deliberately
non-GBRAIN_-prefixed so the hermetic env scrub keeps them.
Two corpora.
- PR corpus (
prCorpus) — the 15 coverage-collecting lanes in.github/workflows/test.yml: the 8 matrix shards, fourserial-testspartitions, and the three dedicated slow jobs (slow-eval-longmemeval,slow-entity-resolve-perf,slow-brainbench-e2e). Deterministic (runs identically on every PR); this is the corpus the gates run against. - fullCorpus — nightly or explicit manual opt-in in
.github/workflows/e2e.yml:coverage-full-{unit,serial,slow,e2e}+coverage-full-report. Fully self-contained (every lane re-runs with coverage inside that workflow, including the full default E2E discovery across four isolated Postgres workers) — the honest merged unit+serial+slow+e2e number, kept as thecoverage-full-mergedtrend artifact.
The E2E workflow's manual full_corpus boolean defaults to false. Setting it
to true executes the same full unit/serial/slow/E2E profile, report and receipt
checks as a schedule; selected E2E receives the same explicit empty sentinel.
Ordinary pushes, PRs and default manual runs retain their existing selection.
Explicit full manual runs have a separate concurrency group, so they do not
cancel ordinary validation on the same branch. To measure a branch before the
next scheduled run, dispatch e2e.yml on that branch with full_corpus=true.
The nightly E2E artifacts are coverage-full-e2e-1 through -4, with one
manifest per artifact and one coverage directory per Bun process. Lightweight
e2e-full-execution-* artifacts also carry the manifest and executed-file list.
Full-profile e2e-status requires the matrix job to succeed and validates all four
same-commit receipts against the exact expected partitions using
scripts/verify-nightly-e2e.ts. Missing artifacts, duplicate identities, wrong
commits, omitted or repeated files, failures and cancellations cannot report
complete execution. The report independently verifies these receipts before
merging; missing execution evidence prevents publishing a full-corpus report.
Coverage percentages remain advisory and the report job itself is not an
e2e-status dependency. The receipts prove file execution, not execution of
every optional assertion within a file.
Merge (scripts/merge-lcov.ts). Walks the input dirs for lcov.info +
lane-manifest.json, sums DA hits per file:line, normalizes paths
repo-relative, and emits a merged lcov plus a summary JSON: src-only
totals/per-dir/per-file percentages, a lineHits map (the diff gate's input),
and the never-loaded src file list. --manifest-expect lane,lane,... pins the
expected lane set (serial-1 through serial-4 for PRs, serial nightly,
and e2e-1 through e2e-4 nightly).
The merger checks commit SHA (--sha overrides checkout HEAD for offline
artifacts), duplicate identities, and actual per-lane LCOV counts; a missing or complete: false manifest, an unparseable
lcov, or a shard lane with lcovCount != 1 marks the summary
degraded: true. Degraded is data, not failure: the merge never aborts (exit
0), and both gates print WOULD PASS/WOULD FAIL and exit 0 on a degraded
summary instead of enforcing against partial data.
Diff gate (scripts/coverage-diff-gate.ts). Gates the added/changed lines
of git diff origin/master...HEAD restricted to gate scope (src/**.ts minus
*.test.ts/*.generated.ts/*.d.ts): covered/(covered+uncovered) must be
≥ 80%, AND no gate-scoped changed file may be entirely absent from the
coverage data (a never-loaded file is one violation — add a test that imports
it). Non-executable lines (no lcov record) don't count against you; empty and
doc-only diffs short-circuit to PASS via the select-e2e classifier. Escape
hatches: a commit body containing [coverage-exempt: reason] passes with a
loud warning, and scripts/coverage-gate-exemptions.txt (exact path or
trailing-/ prefix per line; resolved via
git show origin/master:scripts/coverage-gate-exemptions.txt, never the
working tree, so a PR cannot self-exempt; SHRINK-ONLY — additions need a
graduation review in the PR description) excludes paths from the gate while
still reporting them
([e2e-exempt], [subprocess-undercount]). Report-only unless
COVERAGE_GATE_ENFORCE=1. Exit contract: 0 = pass or report-only, 1 = gate
fail while enforcing, 2 = infrastructure error (missing summary, git failure —
never conflated with a coverage verdict).
Baseline gate (scripts/coverage-baseline-gate.ts). Anti-erosion floor:
reads the baseline via git show origin/master:scripts/coverage-baseline.json
— the master copy, never the working tree, so a PR cannot weaken its own bar —
and compares like-for-like by corpus (--corpus prCorpus in test.yml,
--corpus fullCorpus nightly). A global drop > 0.5pp, a per-directory drop
1.0pp, or a never-loaded-count increase fails (deleting tests shrinks the coverage denominator, which inflates pct for free); a corpus section that is
nullon master is an ungated first landing.provisional: truein the baseline keeps the gate report-only regardless of enforcement — the committed baseline is currently provisional with both corpus sections unseeded.scripts/update-coverage-baseline.ts --summary <json> --corpus <c> [--promote]writes the working-tree baseline (per-file detail limited to the baseline'swatchlist);--promoteflipsprovisional: falseat graduation.
CI wiring. The 17 PR lanes upload coverage-* artifacts; the advisory
coverage-report job downloads + merges (COVERAGE_CORPUS=prCorpus), renders
scripts/render-coverage-summary.ts to the step summary (including the
behavioral-vs-structural counts from scripts/structural-suites.tsv), and
runs both gates with COVERAGE_GATE_ENFORCE: '0'. It is deliberately NOT in
test-status needs — it cannot block a PR until graduation. Test results are
never cached: every run executes its required checks, while Bun dependencies
and validated PGLite snapshots remain cached. Gitleaks runs independently for
all changes, including documentation. e2e-status also requires the four
coverage-full-* execution lanes on scheduled runs; only coverage percentage
reporting remains advisory.
Bun caveats. Bun/JSC emits line records only, so function coverage is
informational (no reliable function names). There is NO subprocess coverage:
code exercised only through spawned CLI subprocesses undercounts — src/cli.ts
carries a permanent [subprocess-undercount] exemption for this. A src file
never imported by any test produces no lcov record at all; the summary reports
these as a count + sorted list, deliberately never a percentage (physical
lines ≠ executable lines), and the diff gate treats a changed-but-never-loaded
file as a violation.
One-command local smoke (one shard of ten, so totals reflect a tenth of the corpus — this checks the plumbing, not the number):
COVERAGE_DIR=$PWD/.coverage bash scripts/test-shard.sh 1 10 \
&& bun scripts/merge-lcov.ts --out-lcov .coverage/merged.lcov --out-json .coverage/summary.json .coverage \
&& bun scripts/render-coverage-summary.ts --summary .coverage/summary.jsonOptional flags: coverage-diff-gate.ts --base <ref> overrides the diff base
(default origin/master); render-coverage-summary.ts --structural scripts/structural-suites.tsv adds the behavioral-vs-structural split to the
rendered summary (both CI lanes pass it); classify-tests.ts --summary prints
counts only.
When bun run test finds any failure, the wrapper:
- Writes failure blocks (each prefixed with
--- shard N: <test name> ---) to.context/test-failures.log(workspace-local, gitignored). On systems without a writable.context/, falls back to/tmp/gbrain-test-failures.log. - Prints a loud stderr banner with the absolute log path, plus the last 30 lines of the failure log inlined. Banner survives
| head/| tail/ agent-side log truncation. - Writes a one-line-per-shard summary to
.context/test-summary.txt(shard N/M: pass=X fail=Y skip=Z rc=W). - Exits non-zero. Empty failure log + non-zero exit = infrastructure problem (wedged shard, killed child); the banner says so.
If a shard hits the per-shard GBRAIN_TEST_SHARD_TIMEOUT cap (default 3000s — sized so the heaviest count-balanced shard finishes under 4-way contention; GBRAIN_TEST_SHARD_KILL_AFTER sets the grace after TERM before KILL, default 30s), the wrapper classifies the kill one of two ways:
- EXIT-HANG → warn-pass. If the shard's log had been silent for ≥300s at kill time AND shows zero
(fail)markers, the shard finished all its work, leaked a handle, and never exited (a known PGLite-adjacent handle leak — see TODOS.md "unit-shard exit hang"). The wrapper prints a⚠️ shard N/M: EXIT-HANG ... Treating as pass-with-warningbanner, writesEXIT-HANG (idle Ns, 0 fails) ... warn-passto the summary, and does NOT fail the run. Its pass counts are undercounted (bun never printed its final summary). Bun's per-test--timeoutturns a genuinely hung TEST into a printed(fail)— new output — so this classification cannot mask a hung test; the residual maskable case is a file-level import hang in the very last file, which the banner keeps visible. - WEDGED → hard failure. Anything else (failures present, or the log was still growing) writes
--- shard N: WEDGED after ${SHARD_TIMEOUT}s ---to the failure log with the last 50 lines of the shard log, marks the run failed, and proceeds with other shards' results.
Triage rule: a warn-pass EXIT-HANG line in .context/test-summary.txt is NOT a test failure — don't burn time bisecting it; a WEDGED line is.
*.test.ts→ fast loop (parallel up-to-4-shard fan-out, memory-adaptive).*.slow.test.ts→ run viabun run test:slowonly (intentional cold-path tests; would dominate the fast loop's wallclock).*.serial.test.ts→ run viabun run test:serialafter the parallel pass completes; one bun process per file (--max-concurrency=1within a shared process is not enough — the module registry still leaksmock.module), with those per-file processes POOLED (per-process isolation never required one-at-a-time execution). Files touching machine-global state (launchd/cron) live on the sequentialEXCLUSIVE_FILESlane insidescripts/run-serial-tests.sh— growth-guarded to ≤3 entries with justification comments. Quarantine for tests that share file-wide state and race when run alongside other files in the samebun testprocess. Several dozen files, discovered by the*.serial.test.tsglob — no list to maintain. Typical residents:mock.module(...)users (top-level mocks leak across files in a shard process, e.g.test/embed.serial.test.ts), env-coupled files (e.g.test/brain-registry.serial.test.ts), and process-lifecycle suites that assert onprocess.exitCode(e.g.test/pglite-engine-disconnect.serial.test.ts). Do not put the parallelism back on a serial file unless you've fixed the contention root cause (it just re-introduces the flake).test/e2e/*.test.ts→ real-Postgres E2E. Skipped whenDATABASE_URLis unset. One out-of-directory file rides this lane:test/phantom-redirect-engine-parity.test.ts(lives intest/for its PGLite arm, but its Postgres arm is only reachable through a DATABASE_URL-bearing lane — the unit wrappers strip the URL, sorun-e2e.sh's no-args list and CI's parity job carry it).run-e2e.shwraps each file in a hard outer timeout (default 180s;GBRAIN_E2E_FILE_TIMEOUT=<seconds>overrides) because a synchronously-blocking PGLite WASM call can outlive bun's timer-based--timeout; LLM-bound Tier-2 files (skills.test.ts) automatically get 4× the cap since real provider round-trips legitimately run past 180s.tests/heavy/*.sh→ ops-shape shell scripts. Cost minutes per run; NOT in defaultbun test. Run viabun run test:heavyor scheduled nightly via.github/workflows/heavy-tests.yml. Examples: pg_upgrade matrix (boot legacy brain → walk to head), RSS budget gate (measure peak worker RSS vs committed baseline), read-latency-under-sync (p50/p95/p99 under concurrent writer load), sync lock regression (N concurrent syncs assert 1 winner + N-1 lock-busy + zero leakedgbrain_cycle_locksrows). Seetests/heavy/README.mdfor when to add a script here vs*.slow.test.ts. Files prefixed with_(e.g.tests/heavy/_build_legacy_fixtures.sh) are helpers/libs invoked by sibling tests — the runner skips them.test/fuzz/*.test.ts→ property-based fuzz harness. Pure-validator targets inpure-validators.test.tsare guarded byscripts/check-fuzz-purity.sh(inbun run verify), whichbun build --target=bunbundles each target and greps the resulting bundle for banned transitive imports (node:fs,node:child_process, engine modules). Anything that fails the guard moves tomixed-validators.test.ts(still property-tested, but no purity guarantee) orfilesystem-validators.test.ts(fs-backed, uses temp dirs). Fuzz tests run in the defaultbun testloop because they're fast (~3s for ~12 properties × 1000 runs each).
The taxonomy above is LANE-based (where a test runs). A second, orthogonal axis is INTENT:
- Behavioral tests execute product code and assert on behavior — the default.
- Structural (source-shape) suites read repo source/doc TEXT and assert on its shape (wiring guards, drift pins,
doctorSource()consumers). They are real invariants but execute no product paths, so they inflate the headline test count without adding line coverage. The committed inventory isscripts/structural-suites.tsv, generated byscripts/classify-tests.ts(suite-level, content-based detectors: repo-anchoredreadFileSync/Bun.filereaders, grep-style exec scanners, the doctor-source helpers) and freshness-checked inbun run verify(check:structural-manifest— regenerate withbun scripts/classify-tests.tswhen suites change shape). The inventory is approximate by design; fix misclassifications in the classifier's detector list, never by hand-editing the TSV. CI's coverage report renders behavioral vs structural counts side by side.
Guards that pin doctor source text read it through test/helpers/doctor-source.ts (doctorSource() = the façade + every src/commands/doctor/** module, for containment assertions; doctorFileSource(rel) = one named file, for positional/ordering assertions) so peeling doctor.ts into modules can't silently move a pinned string out of a guard's sight.
Four escalating tools; reach for the cheapest one that answers the question:
| Question | Tool | Example |
|---|---|---|
| Does the TTY/non-TTY branch logic pick right? | Inject isTTY into the pure function — no subprocess |
test/init-provider-picker.test.ts, test/jobs-watch-mode.test.ts |
| Does the real CLI behave right when stdin is NOT a terminal? | Spawn the CLI with piped/ignored stdio | test/cli-stdin-hang.test.ts (fast loop); test/init-fresh-pglite.slow.test.ts (slow lane) |
| Does the real CLI render menus and read typed input under a REAL terminal? | launchTty from test/helpers/tty-harness.ts in a *.serial.test.ts file |
test/init-picker-pty.serial.test.ts |
| How does the install FEEL (stalls, copy, silence windows)? | scripts/dx-explore.ts — instrument, not a test; nothing asserts |
transcripts under .context/dx-runs/ (see docs/guides/bootstrap.md) |
Real-PTY test rules: put the file in the serial lane (*.serial.test.ts — that
lane runs in required CI; a new test/e2e/* file does NOT, since unit shards
exclude the directory and the e2e workflow runs only explicitly named files,
no glob);
assert NON-default picker values (bare Enter and each prompt's 60s
readLineSafe timeout both resolve to the default, so a defaults-asserting
test passes with dead input); always await session.close() in a finally
(only close() clears the harness wall timer); and point HOME plus
GBRAIN_HOME at a temp root with pass-through auth keys stripped via
dropEnv so picker state is machine-independent.
skills/skills.lock.json is a committed sha256 inventory of every bundled file under
skills/ (tamper evidence, not signatures — see src/core/skills-integrity.ts).
Any change under skills/ must regenerate it: bun run scripts/generate-skills-manifest.ts.
scripts/check-skills-manifest-fresh.sh (bun run check:skills-manifest, wired into
bun run verify) regenerates to a tmp file and diffs, failing CI on drift; at runtime
gbrain doctor reports the same drift as a warn-only skills_manifest_integrity check.
test/docs-cli-commands.test.ts checks every gbrain <verb> in code fences and
inline code across README, docs and skills against the registered verbs. In
docs/guides/, docs/migrations/ and skills/ it also runs each invocation's
flags through the CLI's own validator, via test/helpers/cli-command-surface.ts.
When a hit is stale, fix the doc. When the example documents an older release,
put <!-- gbrain-cli: historical --> on the line above its code fence, or on the
line with the inline code. The test's ALLOWLIST is a last resort: it only
shrinks, every entry needs a reason, and stale entries fail.
This section is the canonical home of the test-isolation discipline — CONTRIBUTING.md and other docs link here rather than restating the rules.
The cross-file flake class is enforced statically by scripts/check-test-isolation.sh, wired into bun run verify. Rules (non-serial unit files only; *.serial.test.ts and test/e2e/* are skipped):
| Rule | What it bans | Fix |
|---|---|---|
| R1 | process.env.X = ..., bracket assignment, delete process.env.X, Object.assign(process.env, ...), Reflect.set(process.env, ...) |
Use withEnv() from test/helpers/with-env.ts, OR rename file to *.serial.test.ts |
| R2 | mock.module(...) anywhere in the file |
Rename file to *.serial.test.ts (no DI on production code for testability) |
| R3 | new PGLiteEngine( outside ~50 lines after a beforeAll( line |
Use the canonical block (below) inside beforeAll( |
| R4 | Files creating new PGLiteEngine( without engine.disconnect( inside an afterAll( block |
Add afterAll(() => engine.disconnect()) |
Files that violated these rules at the isolation-lint baseline are listed in scripts/check-test-isolation.allowlist. The allow-list MUST shrink over time — never add new entries.
Every test file that needs a PGLite engine should use this exact pattern:
import { PGLiteEngine } from '../src/core/pglite-engine.ts';
import { resetPgliteState } from './helpers/reset-pglite.ts';
let engine: PGLiteEngine;
beforeAll(async () => {
engine = new PGLiteEngine();
await engine.connect({});
await engine.initSchema();
});
afterAll(async () => {
await engine.disconnect();
});
beforeEach(async () => {
await resetPgliteState(engine);
});Why this exact shape: beforeAll creates a single engine per file (PGLite WASM cold-start + initSchema is ~20s); beforeEach clears user data via resetPgliteState; afterAll disconnects so the engine doesn't leak across file boundaries within a shard process. Ordinary resets atomically delete rows with cleanup-only trigger suppression and restart owned sequences, retaining table/index storage. The helper restores trigger behavior before reseeding and falls back to TRUNCATE CASCADE for schemas whose triggers, rules, inheritance, external foreign keys or privileges require its original semantics. Schema/generation infrastructure survives, and each reset rotates the logical brain identity.
Every full reset measures aggregate target-table storage, including indexes and
TOAST, with pg_total_relation_size. Above 8 MiB it uses the same atomic TRUNCATE
path to reclaim storage; no reset counter or stale size estimate is retained.
The helper regression suite checks repeated TOAST-heavy resets, cleanup and
sequence parity, restored triggers and foreign-key enforcement.
import { withEnv } from './helpers/with-env.ts';
test('reads OPENAI_API_KEY', async () => {
await withEnv({ OPENAI_API_KEY: 'sk-test' }, async () => {
expect(loadConfig().openai_key).toBe('sk-test');
});
});
// Delete a var (override is undefined):
await withEnv({ GBRAIN_HOME: undefined }, fn);
// Multiple keys:
await withEnv({ A: '1', B: '2', C: undefined }, fn);withEnv saves the prior value of every key it touches and restores via try/finally — including when the callback throws. An absent TZ is restored as the zone that was in effect, because Bun keeps the last explicitly set zone when TZ is deleted; a TZ override therefore never leaks into later files in the same process. It is cross-test safe but NOT intra-file concurrent-safe. process.env is process-global; two test.concurrent() calls in the same file both touching the same key will race. Files using withEnv stay outside the test.concurrent() codemod's eligibility filter.
Reach for these before hand-rolling; the five speed helpers each have their own unit test, and the two environment probes are exercised through their consumer suites:
cli-spawn.ts—runCli(argv, opts)(async, hermetic env, timeout-killed),runCliBatch(argvs, {width})(bounded pool, DEFAULT WIDTH 2 — the cap is per-invocation and 4 shards × width multiplies CLI children machine-wide; each child can boot a ~1.5GB PGLite),runCliMemo(argv-keyed memo for read-only calls like--help;clearCliMemo()drops the memo when a test mutates what a memoized call would observe). Replaces the per-file spawn wrappers; a file of N independent sequential spawns becomes one width-2 batch inbeforeAll.wait-for.ts—waitFor(predicate, {timeoutMs, intervalMs})/waitForValue. Replaces fixedsetTimeoutsleeps: polls resolve as soon as the condition holds, and generous deadlines make slow-CI runs LESS flaky than a tuned sleep, not more.with-snapshot.ts—withColdPglite(fn): per-TEST scopedGBRAIN_PGLITE_SNAPSHOTopt-out (save/delete/restore);withSnapshotValue(value, fn)is the general form (pin any snapshot path for fn's scope;undefined= deleted). Use instead of a file-leveldelete process.env.GBRAIN_PGLITE_SNAPSHOT, which forces every engine in the file to cold-boot. Caution: a snapshot-restored engine does not replay migrations on a laterinitSchema()after a version rewind — rewind-arc tests need the cold path (seetest/bootstrap.test.ts).reset-pglite.ts#resetPgliteStateNarrow(engine, tables)— explicit-table truncate for hot loops (the full reset clears the whole catalog). The table list is REQUIRED — a default would silently under-truncate.git-fixture.ts—makeGitFixture(dir): build-once git repo +reset()/commitAll()between tests, replacing per-testgit initchains.fs-perms.ts—permsEnforced()/crontabAvailable()probes: some hosts (FUSE/overlay sandboxes, root) don't enforce permission bits or lack a crontab; tests asserting "this write MUST fail" / "cron registered" usetest.skipIf(!probe())so they skip visibly there and still run in CI.git-stderr-probe.ts—gitStderrLeads(): skips raw-git-stderr-slice assertions behind ambient git PATH shims that print their own diagnostics first (e.g. Conductor's auth-broker wrapper).
Rename to *.serial.test.ts when:
- The file uses
mock.module(...)(R2 — there's no clean fix without changing production code). - The file is genuinely env-coupled (e.g.
gbrain-home-isolation.test.ts,claw-test-cli.test.ts) — module-load env readers + ESM caching defeat dynamic-import-after-env tricks. - The file's tests intentionally share state across
it()boundaries.
The quarantine has grown to dozens of files — treat it as debt: every addition needs a reason from the list above, and prefer fixing the contention root cause when one exists.
bun test runs all tests without a database. E2E tests skip gracefully when DATABASE_URL is not set.
GBRAIN_HOME isolation preload. test/helpers/gbrain-home-preload.ts (bunfig
[test] preload) points GBRAIN_HOME at a per-run scratch dir when it isn't
already set, so unit tests never read — or clobber — the operator's real
~/.gbrain config/brain. Without it, any config-honoring code path silently
changes behavior with whatever the live config.json says. The
canonical GBRAIN_HOME convention is config.ts:configDir(): GBRAIN_HOME is a
PARENT dir and .gbrain is appended. Subprocess-spawning tests must set BOTH
HOME: tmp and GBRAIN_HOME: tmp in the child env (HOME alone loses to the
inherited preload value; in-process HOME mutation loses to Bun's cached
os.homedir()). The e2e wrapper sets its own GBRAIN_HOME before bun starts,
which this preload respects. Because the preload respects a pre-set value, the
unit/slow wrappers (run-unit-parallel.sh / run-unit-shard.sh /
run-slow-tests.sh) strip an ambient GBRAIN_HOME at their boundary — same
discipline as the database-URL vars — so a dev shell configured for a real
brain can't ride through. GBRAIN_DEBUG_PRELOAD=1 prints the allocated
scratch home for debugging.
Installer fixtures must never delete GBRAIN_HOME to test a fallback against
the operator's home. Spawn a disposable child with HOME set before Bun starts,
then set GBRAIN_HOME to the specific fixture. real-home-guard-preload.ts
compares metadata for the real-home autopilot wrapper, env file, start script,
launchd plist and systemd unit around tests. It detects changes rather than
intercepting writes and never reads env-file contents. A deliberate one-shot
installer test can explicitly set GBRAIN_TEST_ALLOW_REAL_HOME_WRITES=1, which
prints a warning; use that only inside an independently isolated child home.
test/real-home-guard-preload.test.ts runs the installer suite with fake-live
sentinels and verifies they are untouched.
Provider-key strip preload. test/helpers/provider-keys-preload.ts (bunfig
[test] preload) strips the ambient provider credentials the canonical fold
recognizes, using the explicit test/helpers/provider-env.ts list checked
against every recipe credential/endpoint plus compatibility aliases, and
defaults GBRAIN_MODEL_DISCOVERY=off (respecting an explicit operator
override), so key-aware model routing (resolveTierDefault) resolves
identically to keyless CI and latest-model discovery never makes a real
network call from a test. Without it, a chat key exported in the dev shell
flips default-model assertions AND turns gated paths into live provider calls.
Tests that want keys inject them explicitly
(configureGateway({env}), withEnv, serial-file process.env) — the
preload removes ambient shell state only, before any test file loads. The e2e
wrapper (scripts/run-e2e.sh) opts back in at its boundary via
GBRAIN_TEST_KEEP_PROVIDER_KEYS=1 — e2e is the lane where real keys are
deliberate (live embed/parity tests skip-gate on them). The routing-only
qm-provisioning, serve-http-surface-ceiling, serve-stdio-roundtrip and
thin-client fixtures still strip provider state in every child, set both
HOME and GBRAIN_HOME to their temporary brain, and pass Bun
--no-env-file (including provisioned shell commands). They exercise routing
without spending provider tokens even when a lane carries keys.
Fixture-specific environment overrides apply last; unrelated credentials are
preserved rather than removed with a broad key-name pattern.
Database-URL run guard. A bun test invocation REFUSES to start while
DATABASE_URL or GBRAIN_DATABASE_URL is ambient in the environment, because some
tests run destructive SQL against whatever those URLs point at (a bare bun test
with ~/.gbrain/.env sourced would run them against a real brain). The guard is a bunfig
[test] preload (test/helpers/database-url-guard-preload.ts); it hard-fails with
instructions rather than silently unsetting (a silent unset would turn
DATABASE_URL-gated e2e tests into green skips). The e2e wrappers
(scripts/run-e2e.sh, the e2e/heavy workflows) opt in at their own boundary via
GBRAIN_TEST_ALLOW_DATABASE_URL=1; the unit/slow wrappers instead strip both
URL vars at their boundary (unit tests need no database), which keeps
bun run test:full working with DATABASE_URL exported. Caveat: bun loads
bunfig.toml from the invocation cwd, so the preload layer only applies to
runs started at the repo root — the per-file name floor below is the layer
that doesn't care about cwd. Two more layers apply after the opt-in: every
test that runs destructive SQL on the ambient URL must call
assertSafeE2eDatabaseUrl() (test/helpers/db-guard.ts — name floor: the database
name must contain "test" as a segment, or be opted in via GBRAIN_E2E_ALLOW_DB)
or carry an inline name floor the coverage gate recognizes
(test/e2e/schema-drift.test.ts keeps its own looksLikeTestDb, deliberately
different because it also accepts *_e2e), and test/db-guard-coverage.test.ts
statically scans the suite and fails when a file connects to DATABASE_URL and
runs destructive SQL unguarded. Local CI explicitly sets GBRAIN_TEST_DB=1
for its known Docker E2E databases, and run-e2e.sh preserves that opt-in so
schema-drift can reset stale fixture schemas on service-name hosts. The hard
test-database name floor still applies; this flag only relaxes the localhost
requirement. The heavy shell lane gets the same floor outside
bun: tests/heavy/_db_floor.sh (sourced by scripts/run-heavy.sh for the whole
lane, and by each database-touching heavy script itself, since scripts are
documented for direct invocation — the PGLite-based heavy scripts unset the URL
instead) checks BOTH URL variables and strips query strings before extracting
the database name, so a ?host=/tmp/test-sockets parameter can't smuggle a
test-shaped segment past it.
Unit tests and what they cover:
test/facts-engine.test.ts/test/consolidate-valid-until.test.ts— facts-list filtering and consolidate correctness:unconsolidatedOnlyis applied before the 100-row limit, so newer consolidated facts cannot permanently hide older pending facts; the phase regression seeds 100 consolidated rows plus three older pending rows and requires all three to progress.test/markdown.test.ts— frontmatter parsing;splitBodysentinel precedence, horizontal-rule preservation,inferTypewiki subtypes.test/chunkers/recursive.test.ts— chunking.test/parity.test.ts— operations contract parity.test/cli.test.ts— CLI structure.test/cli-finish-teardown.test.ts— the CLI teardown contract:computeTeardownDeadlineMsformula/floor/live-registry scaling +GBRAIN_TEARDOWN_DEADLINE_MSoverride (garbage/zero/negative values fall back to the formula);finishCliTeardownclean path (drain BEFORE disconnect, no exit, no warn), backstop on hung drain or disconnect (honors an errored op's exit code), throwing drain/disconnect warned + swallowed; the gbrain-owned verdict channel is immune to PGLite WASMprocess.exitCodewrites;flushThenExitunit coverage with mocked streams (exits once after both stream callbacks, non-TTY aliveness grace, blocked-pipe guard, EPIPE-safe,GBRAIN_FLUSH_GRACE_MSoverride).test/flush-then-exit-harness.test.ts— real spawned-Bun pipe semantics forflushThenExit(fixture:test/fixtures/flush-then-exit-harness.ts): a 4MB piped stdout payload arrives byte-complete with the exit code even with a late reader, small output survives exit with a concurrent reader, and the fence resolves promptly (wall time well under the guard + grace ceiling).test/cli-should-force-exit.test.ts—shouldForceExitAfterMaindaemon-survival gate:serve(stdio and--http) never force-exits, including with preceding global flags; op commands / empty / flag-only argv do; space-separated global-flag VALUES can't fake a command (--timeout 30s serveresolves to theservedaemon, not a30scommand).test/cli-exit-verdict-pin.test.ts— structural class pin: grepssrc/so the NEXT rawprocess.exitCode =write fails CI (a raw write bypasses the gbrain-owned verdict channel and gets silently zeroed by the deliberate flush-exit, which would make a FAIL path exit 0). Runtime variants live intest/cli-finish-teardown.test.ts; this is the review-time guard.test/cli-pipe-truncation.test.ts— real-CLI pipe completeness, implementation-agnostic: the actual CLI run the way agents run it (piped stdout) produces complete, parseable, byte-stable--tools-jsonoutput and exits deliberately, well under the teardown backstop. Synthetic flush-mechanism coverage stays intest/flush-then-exit-harness.test.ts.test/volunteer-context.test.ts— push-based context core, hermetic in-memory PGLite:parseWindowlenientuser:/assistant:parsing, multi-turn window extraction, confidence-gated volunteering (arm confidences, multi-turn/newest-turn boosts,min_confidencegate, max-pages cap), slug-only suppression, privacy (rationales are deterministic templates; synopses pass the takes/facts fence), and the approximate usage-stats join.test/watch-command.test.ts—gbrain watchpush transport: streaming loop, rolling window, session dedupe,--jsonJSONL shape,channel: 'watch'event logging, clean EOF return. Hermetic PGLite + injected line/write deps (no subprocess, no real stdin).test/watch-sigint.serial.test.ts—gbrain watchSIGINT lifecycle against a real spawned CLI subprocess with a tmpdir brain. SERIAL: parallel unit shards flake on concurrent subprocess spawns (same rationale asapply-migrations-pglite-spawn.serial.test.ts).test/init-picker-pty.serial.test.ts— the interactivegbrain initpickers (embedding-provider + search-mode) driven under a REAL pseudo-terminal vialaunchTty: typed input lands (a NON-default mode choice verified by a follow-up non-TTY config read — bare Enter and thereadLineSafetimeout both resolve to defaults, so a defaults-asserting test would pass with dead input), prompt-to-acknowledgement gaps bounded well under the fallback window, plus the Ctrl-D/EOF keyless fallback. On CI, missing PTY support fails loud instead of skipping. Hermetic: HOME + GBRAIN_HOME at a temp root, pass-through auth keys stripped viadropEnv;session.close()infinally. Serial: PTY spawn + full PGLite bootstrap, and the serial lane is what runs in required CI.test/tty-harness.test.ts— the real-PTY harness's pure helpers (stripAnsi,computeStalls,renderStallsReport,parseDriveCommand,buildClaudeTuiSeed) with zero subprocesses; the file's live-PTY smokes aredescribe.skipIf(!ptySupported())-gated.test/autopilot-launchd-lifecycle.serial.test.ts— autopilot lifecycle behavior, not generated-string assertions: the full install → self-disable → status → reinstall → uninstall arc withlaunchctlreplaced by an argv recorder and the generated wrapper executed by a REAL bash against a genuinely deleted repo (every platform), plus a darwin-only fail-SKIP describe against the real launchd under a per-run unique label (GBRAIN_AUTOPILOT_LABEL) so it can never collide with — or tear down — a real install on the host. Serial: spawns subprocesses and pins HOME/GBRAIN_HOME for the whole file.test/autopilot-fanout.test.ts— Autopilot fan-out and policy pins: targeted idempotency keys reopen per dispatch interval while stable doctor/remediate keys remain unchanged; the 60-minute full-cycle floor wins with a remaining small plan, and an all-fresh restart check advances the process-local clock without masking failed stale-source submissions.test/agent-scheduler-contract.serial.test.ts— the documented external agent-scheduler shell chain (gbrain sync --repo X && gbrain embed --stale, live-sync.md / INSTALL_FOR_AGENTS.md Step 7) driven end-to-end through a real/bin/shagainst a keyless PGLite brain: the&&short-circuit IS the contract (argv arrays can't exercise it), the keyless bare stale embed exits 0, and the pull-failure case that must break the chain does. Anti-vacuity: the fixture commits a real page and every read-back asserts pages >= 1. Serial: real spawned CLI + tmpdir HOME.test/cli-format-volunteer.test.ts—formatResult'svolunteer_contexthuman rendering: pointer lines with confidence/arm/rationale, the empty-result message, the approximate stats summary.test/config.test.ts— config redaction.test/files.test.ts— MIME/hash.test/import-file.test.ts— import pipeline.test/upgrade.serial.test.ts— thegbrain upgradecommand via subprocess:--helpprints usage and exits 0, install-method detection,resolveBunGlobalRoot, and the self-upgrade marker format (serial: spawns the real CLI).test/file-migration.test.ts— file migration.test/file-resolver.test.ts— file resolution.test/import-resume.test.ts— import checkpoints.test/migrate.test.ts— migration: v8/v9 helper-btree-index SQL structural assertions; 1000-row wall-clock fixtures pinning O(n log n) behavior; v12/v13 SQL shape;sqlFor+transaction:falserunner semantics; themax_stalled DEFAULT 1regression guard; v24sqlFor.pglite: ''no-op assertion; v117context_volunteer_events(named + idempotent entry, documented columns + both source-scoped indexes afterinitSchema, insert + 90-daypurgeStaleVolunteerEventsround-trip).test/bootstrap.test.ts— bootstrap contract: no-op on fresh install, idempotent across twoinitSchema()calls, no-op on modern brain that already has every probed column, full bootstrap path on a simulated legacy brain, fresh-install regression guard, legacylinksshape coverage.test/schema-bootstrap-coverage.test.ts— CI guard covering BOTH embedded schema blobs: neither may forward-reference state its engine's bootstrap can't create, and a reference covered on one blob is NOT automatically covered on the other (dream_verdictsexists only in the Postgres blob). PGLite half:REQUIRED_BOOTSTRAP_COVERAGElists every forward reference inPGLITE_SCHEMA_SQL; the test fails loudly ifapplyForwardReferenceBootstrapskips one (extend both arrays when adding a column-with-index to the embedded schema blob). Also parsessrc/core/migrate.tssource text for everyALTER TABLE ... ADD COLUMN(top-levelsql:,sqlFor.{postgres,pglite}overrides, AND handler-bodyengine.runMigration(N, \ALTER TABLE ...`)) and asserts each (table, column) pair is covered by the bootstrap OR by the schema blob's CREATE TABLE bodies — catching the column-only forward-reference class (e.g.sources.archived,oauth_clients.source_id) that a CREATE INDEX parser alone can't see. Postgres half (the class-closure gate): parses every CREATE INDEX column reference inSCHEMA_SQLand requires each to be in the blob's CREATE TABLE body AND not migration-added, or probed + ALTERed bysrc/core/engine-sql/bootstrap.ts(the single bootstrap both engines run; the PGLite half reads it minus itsdialect-only:postgresregions) — a column that is both in the blob's CREATE TABLE and migration-added is still a forward reference for pre-existing brains, whereCREATE TABLE IF NOT EXISTSno-ops and the blob's CREATE INDEX wedgesinitSchemabefore migrations can help. This gate is parser-driven (no registry to extend); intentional non-probes go inPOSTGRES_INDEX_REF_EXEMPTIONSwith a rationale. Honest scope: CREATE INDEX column references only — constraints, views, and trigger bodies are a filed TODOS.md follow-up.parseBaseTableColumns` strips SQL line + block comments before identifying column names so commented-out lines don't hide adjacent columns.test/dream-verdict-cache-ttl.test.ts—dream_verdictsTTL contract on PGLite: put assigns the default TTL, expired rows miss on read and only they are swept, re-judging via upsert refreshes a nearly-expired row, the migration backfill derives expiry fromjudged_atidempotently, and a NULL-expiry row (the pre-backfill upgrade window) reads as a hit and survives the sweep — the locally-runnable pin for the NULL-tolerant read predicate both engines share.test/helpers/schema-diff.ts+test/helpers/schema-diff.test.ts+test/e2e/schema-drift.test.ts— cross-engine schema parity gate. Helper exports puresnapshotSchema(query)/diffSnapshots(pg, pglite, opts)/formatDiffForFailure(diff)/isCleanDiff(diff)over a four-tuple per column (data_type,udt_name,is_nullable,column_default). E2E test spins up fresh PGLite + Postgres, runsengine.initSchema()on each, snapshotsinformation_schema.columns, then diffs. 2-table allowlist (files,file_migration_ledger) — every other Postgres table must reach PGLite viaPGLITE_SCHEMA_SQLor a migration'ssqlFor.pglitebranch. Sentinels foroauth_clients,mcp_request_log,access_tokens,eval_candidatesgive tighter blame messages. Skips withoutDATABASE_URL. Wired intoscripts/e2e-test-map.tsso changes tosrc/schema.sql, the PGLite schema façade or its generated template, or the migration runner trigger it. The failure message names every drift with a paste-ready hint pointing atsrc/schema.sql(or its TS fragment) plusbun run build:schema, a PGLite rule inscripts/build-schema.ts, or a migration'ssqlFor.pglitebranch.test/schema-catalog-golden.test.ts+test/e2e/schema-catalog-golden.test.ts— refactor wave 1 E4 catalog goldens (test/fixtures/goldens/catalog/).snapshotCatalog(query)intest/helpers/schema-diff.tspins columns with ordinal position, defaults, full index definitions, constraints incl. CHECK text, triggers, function signature + body sha256, views, policies, grants (connecting role placeholdered), RLS flags, sequences and extension names; extension-owned objects are excluded. Captured at three configs (CATALOG_CONFIGSintest/helpers/schema-catalog.ts: default OpenAI/1536, 4096 dims which skips the chunk and halfvec HNSW indexes,GBRAIN_FTS_LANGUAGE=portuguese) on PGLite engine init, the PGLite schema blob without migrations, Postgres engine init (fresh throwaway database per capture) and Postgresdb.initSchema(), plus an optional PgBouncer arm. Every capture runs twice through theschema-catalog-lines-v1normalizer. PGLite master-vs-branch compares ordinals (the goldens pin them); PG ↔ PGLite parity stays name-based. Regenerate deliberately withGBRAIN_TEST_UPDATE_GOLDENS=1.test/pglite-upgrade-replay.test.ts— refactor wave 1 EO3: opens the pinned master-built PGLite brain (test/fixtures/goldens/pglite-upgrade-replay/brain.tar.gz+MANIFEST.json, rebuilt only on master withbun scripts/build-pglite-upgrade-fixture.ts) with current code and asserts the catalog before boot equals the E4 golden, the post-boot catalog is pinned, corpus rows and fingerprint survive, and a restart + repeatinitSchema()is a no-op (no migrations, identical catalog, unchanged relation/constraint/function OIDs; the triggers the blob recreates on every boot are pinned).test/setup-branching.test.ts— setup flow.test/slug-validation.test.ts— slug validation.test/storage.test.ts— storage backends.test/supabase-admin.test.ts— Supabase admin.test/yaml-lite.test.ts— YAML parsing.test/check-update.test.ts— version check + update CLI.test/pglite-engine.test.ts— PGLite engine, all BrainEngine methods includingaddLinksBatch/addTimelineEntriesBatch(empty batch, missing optionals, within-batch dedup via ON CONFLICT, missing-slug rows dropped by JOIN, half-existing batch, batch of 100) plusconnect()error-wrap assertion (original error nested, #223 link in message, lock released).test/links-timeline-jsonb-poison.test.ts— the PGLite half of the JSONB batch-poison lock (always-on, noDATABASE_URL). Locks thejsonb_to_recordsetbatch-insert path for links/timeline/takes against free-text "poison" payloads (commas, quotes, backslashes, braces, em-dashes) and asserts NUL is stripped from free-text body fields but rejected in identity fields. Lone-UTF-16-surrogate cases: every free-text field (link context; timeline summary/detail/source; take claim/source) well-forms to U+FFFD across batch + scalar write paths, while a surrogate in an identity field (slug) still fail-closed rejects the batch. The Postgres lane istest/e2e/jsonb-batch-poison-postgres.test.ts.test/engine-factory.test.ts— engine factory + dynamic imports.test/integrations.test.ts— recipe parsing, CLI routing, recipe validation.test/publish.test.ts— content stripping, encryption, password generation, HTML output.test/backlinks.test.ts— entity extraction, back-link detection, timeline entry generation.test/lint.test.ts— LLM artifact detection, code fence stripping, frontmatter validation.test/report.test.ts— report format, directory structure.test/skills-conformance.test.ts— skill frontmatter + required sections validation.test/resolver.test.ts— RESOLVER.md coverage, routing validation; round-trip that every quoted RESOLVER.md trigger matches a frontmattertriggers:entry in the target skill, and everyname="<word>"reference in any SKILL.md resolves to a declared op insrc/core/operations.tsor a Minions handler inPROTECTED_JOB_NAMES.test/search.test.ts— RRF normalization, compiled truth boost, cosine similarity, dedup key.test/sql-ranking.test.ts— source-boost helpers: longest-prefix-match in SQL CASE,detail=hightemporal-bypass, three-meta-char LIKE escape (%,_,\), single-quote SQL-literal doubling, env override parsing forGBRAIN_SOURCE_BOOST+GBRAIN_SEARCH_EXCLUDE,resolveBoostMap/resolveHardExcludesmerge semantics.test/dedup.test.ts— source-aware dedup, compiled truth guarantee, layer interactions.test/query-intent-legacy.test.ts— query intent classification: entity/temporal/event/general (the non-concept intents).test/query-intent-concept.test.ts— theconceptintent: definitional/landscape cue detection, the proper-noun / quoted-phrase / sub-3-word guards, vector-lean weight routing.test/eval.test.ts— retrieval metrics:precisionAtK,recallAtK,mrr,ndcgAtK,parseQrels.test/brainbench-fixtures.test.ts/test/brainbench-generator.test.ts/test/brainbench-metrics.test.ts/test/brainbench-continuity.test.ts/test/brainbench-writeback.test.ts/test/brainbench-adapters.test.ts/test/brainbench-scoreboard.test.ts— the BrainBench memory-conformance unit suites (src/eval/brainbench/): fixture loader/validator + the sealed-gold seal (agoldkey inside a fixture must reject) and committed-corpus integrity; generator determinism (the committed corpus is exactly whatgen.tsproduces, holdout discipline, category counts); metric formulas over hand-built turn rows (zero should-retrieve turns, empty injections, acceptable-vs-gold asymmetry, micro-averaging); cross-harness continuity (writer's decision persists through the production write-back pipeline, reader recalls on the SAME brain); write-back grading the PRODUCTION conversation→facts pipeline via the injected gold extractor; adapter seam contracts over hermetic PGLite (budget caps, suppression modes); scoreboard + gate governance (baseline determinism, count-aware gating, corpus-bless modes, justification flow, isolation gates-at-zero).test/brainbench-floors.test.ts— the pre-registered quality floors as executable assertions against the committed baseline (a baseline bless can't bank a threshold violation).test/eval-brainbench-e2e.slow.test.ts— BrainBench CLI end-to-end via subprocess against a small tmp corpus: the literal exit codes (0 pass / 1 regression / 2 error-or-inconclusive — the CI product),--outartifact validity incl._meta.metric_glossary, byte-deterministic--update-baseline, anti-vacuous-pass, and theeval run-allonce-per-sweep record. Slow-tiered with its own CI job (slow-brainbench-e2e); independent CLI runs execute once through a width-2 pool inbeforeAll. There is no in-process full-corpus completion test: CI'sbrainbenchgate runs the committed corpus fresh on every PR and its baseline compare is the fixtures-hash drift guard.test/check-resolvable.test.ts— resolver reachability, MECE overlap, gap detection, proximity-based DRY detection,extractDelegationTargetscoverage.test/dry-fix.test.ts— auto-fix: three shape-aware expander pure-function tests; five guards (working-tree-dirty, no-git-backup, inside-code-fence, already-delegated within 40 lines, ambiguous-multi-match, block-is-callout).test/doctor-fix.test.ts—gbrain doctor --fixCLI integration: dry-run preview, apply path, JSON output shape.test/doctor-registry.test.ts— doctor check registry contract: every entry name and emitted check categorized (FAIL/Why/Fix/See),emits[]equals the AST-walked names of each entry'srun, runtime registry equals the static walk, STOP gates at master's early returns.test/doctor-mode-matrix.serial.test.ts— doctor mode matrix through the registry runner: entries run, STOP position, engine calls and--fixmutations for default,--fast,--fix,--fix --dry-run, no engine and connection failure.test/backoff.test.ts— load-aware throttling, concurrency limits, active hours.test/transcription.test.ts— provider detection, format validation, API key errors.test/enrichment-service.test.ts— entity slugification, extraction, tier escalation.test/minions.test.ts— Minions job queue: CRUD, state machine, backoff, stall detection, dependencies, worker lifecycle, lock management, claim mechanics, depth/child-cap, timeouts, cascade kill, idempotency,child_doneinbox, attachments, removeOnComplete/Fail,max_stalledclamp/default/plumbing coverage.test/minion-queue-renewlock-signal.test.ts—renewLockforwards its optional AbortSignal toexecuteRawDirect(stub-engine capture); legacy 3-arg calls unchanged; token-fence miss returns false.test/cycle-drain-renewal.test.ts—runDrainRenewalTick(cycle drain): per-call signal aborted on timeout (slot released), onLost once on a lost fence, throws swallowed, hung renewal resolves at the deadline. Plus two structural source-text pins oninline-drain.ts(the shape guard only coversworker.ts): the renewal must not go back to a rawsetInterval(() => queue.renewLock(...)), and the handler invocation must stay wrapped inwithChatPhase('job:<name>')so a drained child's gateway spend is attributed to the child rather than absorbed by an enclosingphase:tag.test/queue-probe-cancellation.test.ts—probeQueueState/queryWedgeSignalssignal threading: the 1500ms budget CANCELS the losing probe query; fast-path signals never abort; throw still collapses to{probe_failed: true}.test/db-pool-max-lifetime.test.ts—resolveMaxLifetimeSeconds: env forms, 0-disables, 30–60min jitter bounds, warn-once on invalid, per-call jitter variance.test/pool-gauge.test.ts—CheckoutGaugepure semantics + the PostgresEngine seams with fake pools: counted while in flight, released on resolve, on REJECTED queries, and on the SYNCHRONOUS pre-aborted-signal throw (leak guards);getPoolDiagnosticsfail-open.test/db-probe.test.ts—runDbProbeverdict matrix (pool_starved / server_unreachable / unknown), honest-disjunction + no-waiter-arithmetic wording pins, hung probes cancelled via their signals, diagnostics absent/throwing fail open.test/postgres-engine-reserved-routing.test.ts—withReservedConnectionrouting: direct pool when dual-pool active, read pool when kill-switched/in-tx, semaphore cap (directPoolSize−1) with read-pool overflow, permit released on fn throw and reserve failure.test/job-isolation-protocol.test.ts— outcome-file codec round-trip + every decode failure path (missing/malformed/oversize→UnrecoverableError; byte counts, never content), handler-error instanceof reconstruction, child-CLI invocation resolution, and REAL detached-processkillProcessGrouptests incl. the grandchild-death guarantee (exercises the Bun negative-pid/bin/killfallback for real underbun test).test/run-child-entry.test.ts—runChildJobEntryon real in-memory PGLite with a REAL claim-minted token: success (fenced updateProgress lands), handler-failure outcome (exit 0), token-mismatch never runs the handler (exit 14), missing job/handler, parent-death watchdog aborts a live handler.test/child-job-runner.test.ts—runJobInChildagainst real .mjs children: success + full env contract (incl.GBRAIN_DIRECT_POOL_SIZE=1), error/lease outcome reconstruction, crash, SIGTERM-ignorer → group SIGKILL at the injected grace, pre-aborted signal, spawn ENOENT →ChildSpawnInfraError, worker-shutdown drain (report-during-drain completes; non-reporting kill →ChildWorkerShutdownError).test/worker-job-isolation.test.ts— full parent path on PGLite with thefake-run-child.mjsfixture: claim → child → fenced completeJob (real token over env), error outcome → failJob, crash burns the attempt, spawn failure RELEASES with zero attempts burned, and the serialization-parity pin (unreportable results fail in BOTH modes, never falsely complete).test/jobs-isolation-flag.test.ts—parseJobIsolationFlag: space/= forms, env fallback + flag-wins, empty-env default, other flags untouched.test/extract.test.ts— link extraction, timeline extraction, frontmatter parsing, directory type inference.test/extract-db.test.ts—gbrain extract --source db: typed link inference, idempotency,--typefilter,--dry-runJSON output.test/extract-fs.test.ts—gbrain extract --source fs: first-run inserts + second-run reports zero, dry-run dedups candidates across files, second-run perf regression guard for the N+1 dedup bug.test/link-extraction.test.ts— canonicalextractEntityRefsboth formats,extractPageLinksdedup,inferLinkTypeheuristics,parseTimelineEntriesdate variants,isAutoLinkEnabledconfig.test/graph-query.test.ts— direction in/out/both, type filter, indented tree output.test/features.test.ts— feature scanning, brain_score calculation, CLI routing, persistence.test/file-upload-security.test.ts— symlink traversal, cwd confinement, slug + filename allowlists, remote vs local trust.test/query-sanitization.test.ts— prompt-injection stripping, output sanitization, structural boundary.test/search-limit.test.ts—clampSearchLimitdefault/cap behavior acrosslist_pagesandget_ingest_log.test/repair-jsonb.test.ts— JSONB repair: TARGETS list, idempotency, engine-awareness.test/migrations-v0_12_2.test.ts— JSONB-repair orchestrator phases: schema → repair → verify → record.test/orphans.test.ts— orphans command: detection, pseudo filtering, text/json/count outputs, MCP op.test/postgres-engine.test.ts—statement_timeoutscoping:sql.begin+SET LOCALshape, source-level grep guardrail against a reintroduced bareSET statement_timeout.test/sync.test.ts— sync logic + regression guard asserting top-levelengine.transactionis not called.test/sync-pull-failed-anchor.serial.test.ts— a failed internalgit pull(local-path origin vsprotocol.file.allow=never) with zero imports returnspartial/pull_failed(notup_to_date), freezeslast_commit+last_sync_at, recovers after a manual pull; fall-through import of local commits preserved. Serial: pinsGBRAIN_HOMEto a temp dir for the whole file.test/sync-concurrency.test.ts—autoConcurrency()thresholds + PGLite-forces-serial + explicit-override clamping;shouldRunParallel()explicit-bypasses-floor contract;parseWorkers()validation rejecting'0'/'-3'/'foo'/'1.5'/trailing chars.test/sync-parallel.test.ts— PGLite-routed coverage of the bookmark gate under concurrency, head-drift gate, vanished-file failure capture, PGLite-stays-serial, and thegbrain-syncwriter-lock contract.test/sync-all-missing-path.test.ts—sync --all --missing-path <fail|skip>pure helpers:parseMissingPathMode(default fail, explicit values, loud rejection of bad/dangling values, never swallows a following flag) andpartitionMissingPathSources(classification driven only by the injected pathExists predicate — no fs; nulllocal_pathpasses through runnable; order preserved).test/sync-failures.test.ts—classifyErrorCoderegex coverage for all 12 codes against literal production message strings frommarkdown.tsandimport-file.ts;summarizeFailuresByCodesort + pre-classified-honor;recordSyncFailurescode-field persistence;acknowledgeSyncFailuresAcknowledgeResultshape + backfill on legacy entries.test/sync-soft-delete.serial.test.ts— removed-file recovery arc: agit rmdrained by sync SOFT-deletes the page (deleted_atset; row recoverable, not gone), an already-soft-deleted row isn't re-flipped (purge clock preserved), batch delete failures decompose to per-file batches and the run banks instead of aborting, delete → re-add inside the window revives via upsert (content updated, chunks replaced, no duplicate), soft-deleted pages stay invisible to search/getLinks/getBacklinks, the rename lane converges against an out-of-band soft delete, and full-sync reconcile + the unsyncable lane are SOFT with the purge window honored end-to-end.test/sync-exclude-config.test.ts— persistedsync.excludereach: honored with no flag on incremental AND first-sync full-walk paths, trailing-slash covers directory contents, a per-call flag narrows without re-opening the persisted scope, mixed comma+newline multi-pattern values, conservative posture (pages imported before the exclusion stay live, incl. full-sync reconcile), and a throwing/unreadable config read degrades to no-persisted-scope instead of breaking the sync.test/sync-include-hidden-config.test.ts— persistedsync.include_hiddenreach (the dot-directory waiver's twin tosync.exclude): baseline control (no config, no flag → dot-directory pruned), honored with no flag on incremental AND first-sync full-walk paths, trailing-slash covers nested files (lowercased slug), an unnamed dot-directory stays pruned, a per-callincludeHiddenunions with the persisted waiver, and a throwing config read degrades to no-waiver instead of breaking the sync.test/doctor.test.ts— doctor command; assertions thatjsonb_integrityscans the four JSONB write sites andmarkdown_body_completenessis present.test/utils.test.ts— shared SQL utilities +tryParseEmbeddingnull-return and single-warn semantics.test/build-llms.test.ts—llms.txt/llms-full.txtgenerator: path resolution, idempotence, spec shape, regen-drift guard, content contract, AGENTS.md install-path mirror, size-budget enforcement.test/oauth.test.ts— OAuth 2.1 provider: register, getClient,client_credentialsgrant exchange,authorization_codeflow with PKCE challenge/verifier, refresh token rotation,verifyAccessTokenwith both OAuth + legacyaccess_tokensfallback,revokeToken,sweepExpiredTokens; contract test assertingscope+localOnlyannotations on all operations;coerceTimestampunit cases (null/undefined/string/number/throw-on-NaN); NULL-expires_at-as-expired contract for both refresh + access token paths; cascade-delete contract assertingrevoke-clientpurgesoauth_tokens+oauth_codesvia FK CASCADE; cross-client isolation (wrong-client attempt MUST reject AND rightful owner MUST still succeed atomically afterward); empty-stringredirect_uribypass guard; PKCE DCR public-client gate (token_endpoint_auth_method: "none"returns noclient_secret, defaultclient_secret_postclients get the one-time-reveal secret,getClientNULL→undefined normalization, full PKCE/authorize→/tokenround-trip against a public client).test/mcp-dispatch-summarize.test.ts—summarizeMcpParamsinvariants: declared-keys allow-list intersection, attacker-key-name leak guard (unknown keys counted not named), 1KB byte bucketing for size-probe defense, missing op falls through to fully-redacted shape, declared-keys sorted for deterministic output.test/trust-boundary-contract.test.ts— fail-closed trust semantics under cast bypass:ctx.remote === undefinedtreated as remote/untrusted at every flipped call site;as anyandPartial<>spreads can't downgrade trust by accident.test/remote-privacy-sweep.test.ts— registry-driven remote privacy sweep: every non-localOnly op dispatched remote-shaped throughdispatchToolCallagainst a corpus seeded with high-entropy private sentinels, in both scalar and federated caller shapes; the full response envelope (structured fields, rendered text, errors,_meta.brain_hot_memory) asserted sentinel-free. Fail-closed maintenance contract: a new op fails the suite until classified inEXPECTED_OUTCOME(+PARAM_FACTORYif it can return corpus data); localOnly ops asserted denied over non-stdio transports; publish-gated ops must deny naming their gate. Curated static sibling:test/operations-trust-boundary.test.ts.test/check-resolvable-cli.test.ts— CLI wrapper: exit codes, JSON envelope shape, AGENTS.md fallback chain.test/regression-v0_16_4.test.ts—findRepoRootregression guard, hermetic startDir parameterization.test/repo-root.test.ts—findRepoRootwalk semantics + default-arg parity; the 4-tierautoDetectSkillsDirfallback chain ($OPENCLAW_WORKSPACE→~/.openclaw/workspace→ repo-root →./skills); RESOLVER.md/AGENTS.md filename precedence; explicit-env-wins-over-repo-root; tier-0$GBRAIN_SKILLS_DIRvalid/invalid/precedence-over-OPENCLAW_WORKSPACE; the install-path walk inautoDetectSkillsDirReadOnly; no-drift on primary success;AUTO_DETECT_HINT+AUTO_DETECT_HINT_READ_ONLYcontent; regression guard asserting the sharedautoDetectSkillsDirMUST NEVER return'install_path'source (how the read-path/write-path split stays safe).test/resolver-merge.test.ts— multi-file resolver merge:findAllResolverFilesempty / RESOLVER.md-only / AGENTS.md-only / both-present (RESOLVER.md first);checkResolvablemerge semantics acrossskills/RESOLVER.md+../AGENTS.mdfor the OpenClaw layout where the skillpack ships a thin RESOLVER.md and the real dispatcher lives at the workspace root; dedup byskillPath(first occurrence wins); AGENTS.md-at-workspace-root works alone.test/filing-audit.test.ts— filing audit:writes_pages/writes_tofrontmatter, filing-rules JSON validation.test/skill-brain-first.test.ts— shared frontmatter parser;analyzeSkillBrainFirstcompliance ladder across 9 fixtures undertest/fixtures/brain-first-skills/(compliant-callout, compliant-phase, compliant-position, exempt-frontmatter, missing-brain-first, multi-pattern, negation-prose, no-external, typo-frontmatter); offset helpers; external-lookup regex shape; audit snapshot+diff transition logic;FORMERLY_HARDCODED_EXEMPTregression absorption.test/routing-eval.test.ts— fixture parsing, structural routing,ambiguous_with, Haiku tie-break layer.test/skill-manifest.test.ts— skill manifest parser: drift detection, managed-block markers.test/skillify-scaffold.test.ts—gbrain skillify scaffoldstubs: SKILL.md, script, tests, routing-eval fixtures.test/skillpack-install.test.ts— skillpack bundle + surviving installer primitives:bundle.tsenumeration (manifest load/validate, dependency closure,--all) and theinstaller.tsseams that outlived the removedskillpack installcommand (diffSkillbehindgbrain skillpack diff, managed-block build/parse, lockfile concurrency, atomic writes).test/http-transport.test.ts— HTTP transport: bearer auth + missing/no-Bearer/unknown/revoked +/healthbypass; dispatch.ts round-trip; invalid_params; application/json response shape (not SSE); CORS default-deny + allowlist; body cap on Content-Length AND chunked; two-bucket rate limit (refill, exhaust+Retry-After, LRU eviction, TTL prune, pre-auth IP fires before DB);mcp_request_logaudit on success + auth_failed.test/mcp-expose.test.ts—gbrain mcp exposeagainst a fake Tailscale runner in a tmpdir: dispatch + argument shape (exclusive pairs, invalid--port/--surface), plan + consent (TTY prompt, non-TTY without--yes, declined), every Tailscale step (binary lookup and install plan, login including the refusal tosudoa non-system binary, the HTTPS-certificate / Funnel identity pre-checks, publish with the fail-closedserve statusread and the foreign-handler refusal), happy paths on linux-systemd and darwin launchd (app-bundle CLI), service edge cases,--statusand--removeincluding receipt-less recovery and the scoped--set-path=/ off, the occupied-port probe (any answer counts), the PGLite lock-holder warning, receipt shape guard + rollback + path confinement, engine detection + summary variants, and therunMcpdispatch regression; never prints a stack trace.test/serve-service.test.ts— the persistent-service half ofmcp expose: paths undergbrainPath('serve')+ supervisor target detection,renderServeWrapper(and the rendered wrapper actually running under bash), launchd plist + systemd unit renderers,ensureAdminToken(0600, token shape, exclusive-create race), install / uninstall / state probes with their edges and supervisor hardening, receipt read/write + shape validation.test/tailscale.test.ts— the pure Tailscale helpers:parseTailscaleStatus(tolerant of missing fields),findTailscaleBinary,tailscaleInstallPlanper platform, the argv builders,parseServeStatusStrict+ handler lookup (non-JSON or non-object output isnull, never an empty config),classifyTailscaleErrorkinds and defaults, anddefaultCommandRunnervia real spawns of hermetic commands only.test/restart-sweep.test.ts—recipes/restart-sweep.mdinlined script: sentinel-anchored fenced-block extraction with salted tmp filenames to bypass ESM cache; constructor-time env reads (proves no module-load snapshot); idempotency layer load/save/atomic-tmp-rename/corrupt-JSON-recovery/30-day-prune;(sessionKey, lastAlertedAt)cooldown gate with 6h threshold; AGGRESSIVE-gate two-state tests; execFile argv shape proving shell metachars inOPENCLAW_TELEGRAM_GROUPcannot reach/bin/sh; real-\n-not-literal alert formatting;GBRAIN_HOMEstate path override.test/eval-longmemeval.slow.test.ts+test/eval-longmemeval-e2e.slow.test.ts— LongMemEval harness, hermetic with noDATABASE_URLand no API keys, split in two files so CI's LPT bin-packer can shard them: the pure / harness-shared half (harness lifecycle, PGLite create +resetTablesover runtime-enumeratedpg_tableswith the infrastructure tables preserved, schema-migration robustness of the reset, the warm-create speed gate,haystackToPages, the source-boost regression guard,loadResumeSet, the schema-v2buildByTypeSummary) and the end-to-end half (every describe that callsrunEvalLongMemEvalagainst ONE shared benchmark brain: stubbed-LLM answer-gen and--retrieval-onlyruns, JSONL format + key contract, per-question failure handling,--resume-from,--by-type+--by-type-flooron a no-op resume, a run where every question errored exits 1, duplicatequestion_idhandling).test/eval-longmemeval-mixedcase.slow.test.ts— the like-for-like harness pinned on the_s-shaped mixed-case fixture (test/fixtures/longmemeval-mixedcase.jsonl, placeholder bodies underscripts/check-fixture-privacy.sh): raw-id join through the per-question slug→raw map, strictrecall_allvs any-hit on a two-gold question, abstention exclusion,slug_collisionerror rows,retrieval_config_hash-gated resume,retrieved[]rows for replay.test/eval-longmemeval-parse-args.test.ts— hermetic table test forgbrain eval longmemevalargument validation: every invalid flag value exits 1 fromparseArgsbefore any work (the dataset path is deliberately non-existent so a case that slipped past the parser fails with a DIFFERENT message),--helpexits 0.test/eval-longmemeval-cli-smoke.test.ts— subprocess smoke through the real CLI, pre-dispatch flag validator included: the documented--retrieval-only --by-type --no-trajectory --keyword-onlyinvocation exits 0 and writes aby_type_summaryline (the flag registry once attributed these flags to another command's row).test/generate-flag-registry.test.ts— the flag-registry marker-segmentation rule: a--flagliteral belongs to the command named in the enclosingcommand === 'X'head for every dispatch shape (plain, compound&& args[0] === 'sub', multi-line compound).test/longmemeval-judge.test.ts+test/eval-longmemeval-judge.slow.test.ts— the judged answer-accuracy lane. Pure half: theevaluate_qa.py::get_anscheck_promptport (one branch per question type, abstention by_abssuffix, the data-boundary framing + tag neutralisation), the official'yes'-substring verdict rule vs the runner'smalformedclass,judge_config_hashsensitivity, cost estimates, theBudgetLedger. End-to-end half:--judgeon the mixed-case fixture with a canned reader (ThinkLLMClient) and a canned judge (JudgeChatFn) on in-memory PGLite — rows carry the judge fields + reader pins, the summary headline scores ungradable rows as incorrect, judge-only backfill on--resume-from.test/longmemeval-metrics.test.ts/test/longmemeval-resume.test.ts/test/longmemeval-run-config.test.ts/test/longmemeval-emit.test.ts/test/longmemeval-capture.test.ts/test/longmemeval-reader.test.ts/test/longmemeval-splits-fixture.test.ts— pure pins for the harness modules: the raw-id join +recall_all@k/recall_any@k+ schema-v2 buckets; resume re-scoring (recall recomputed, never trusted;gold_missing/collisionscounted over the same row set as a live run);loadQuestionIds/loadDatasetand theretrieval_config_hash--search-pinfold (an absent fold hashes identically to every existing receipt); the emitter's truncate / append modes, atomic summary rewrite and same-file resume compaction; the--capture-poolreceipt fields mirroring hybrid.ts's autocut inputs; the reader receiving WHOLE sessions (the 4000-char sanitizer cap does not apply); integrity of the committed seed-42 splits (ids only, disjoint halves, no_absids).test/longmemeval-embed-cache.test.ts—src/eval/shared/embed-cache.tshermetic pins (counting fake transport,bun:sqlitefiles under a per-test tmp dir, an explicitopenai:text-embedding-3-large @ 4 dimsgateway): exact(model@dims, text, side)keying, the hardEmbedCacheIntegrityErroron a dims mismatch,bypassed/infra_faultsaccounting, the canonical hash. Its canonical-hash describe is listed inscripts/structural-suites.tsv.test/longmemeval-diagnostics.test.ts— miss diagnostics (src/eval/longmemeval/diagnostics.ts) pure pins: the class decision table (synthetic arm ranks → class), the frozen clause splitter, the H1 signature and H3a/H3b split, receipt parsing + top-k reading, split membership, the summary and glossary header.test/replay-autocut-floor.test.ts—src/eval/shared/autocut-replay.ts+ its CLI on synthetic pools: the live decisions the shipped default (jump 0.2, minKeep 1, floor 0.35) makes on the fixture pools are HARDCODED literals worked by hand (sovalidateLiveis checked against something the code did not produce), the floor sweep, paired deltas, split-half,normalizePoolRowrefusing slug-less rows.test/eval-spend-guard.test.ts—scripts/eval-spend-guard.shsubprocess pins with a temp ledger (env passed to spawn, never mutated): a marker file proves the wrapped command ran; two rows per launch (runningreservation,donereconciliation); fail-closed on a missing / unparseable ledger, a malformed amount and a cap breach; actual-cost file precedence.test/r1-namedthing-rerank-ab.test.ts—scripts/r1-namedthing-rerank-ab.ts+ the NamedThingBench corpus module, hermetic: the embed transport is stubbed to throw so only the OFF arm runs in-process; the ON arm's paid path is covered by the pure verdict / integrity functions and the CLI dry run's "ON arm skipped" contract; the seed-contract engine lives inbeforeAll(test-isolation rule R3).test/ai/gateway-chat-temperature.test.ts—ChatOpts.temperaturereaches the AI SDK call and the provider-reported snapshot surfaces asChatResult.responseModel(the judge pins temperature 0; without the field a judge run would have used the provider default), through the__setGenerateTextTransportForTestsseam.test/search/fusion-lists.test.ts+test/search/expansion-variant-budget.test.ts— role-tagged fusion arms + budget-normalized weighted RRF: pure pins forcomposeFusionLists/rrfFusionWeighted(nullbudget is byte-identical to unweighted fusion; empty arms cast no vote; a missing original makes every text arm a variant), thensearch.expansion_variant_budgetend-to-end throughhybridSearchon a discriminating corpus (in-memory PGLite + deterministicbasisEmbedding), delta-asserted.test/search/arm-confidence.test.ts+test/search/arm-confidence-hybrid.test.ts—search.keyword_arm_confidence_floor: the pure statistic + decision, its composition throughfusion-lists.ts, the knob plane (bundle, config parse, resolution chain, knobs-hashkacf=, registry), and the hermetic end-to-end where a weak keyword arm is down-weighted only with the floor set.test/search/metadata-boost-gate.test.ts+test/search/metadata-boost-gate-hybrid.test.ts—search.metadata_boost_gate:lexicalArmsVoted/decideMetadataBoosts(relaxed rows never count; image modality exempt), the thread throughrunPostFusionStages, the knob plane (mbg=, dashboard), and the hermetic end-to-end in which a hub page's backlink / recency boosts are skipped when only the vector arm voted.test/search/relational-rerank-pin.test.ts+test/search/relational-rerank-pin-hybrid.serial.test.ts—search.relational_rerank_pin: pure permutation pins forpinRelationalRows(top block in fused order bounded bymax; a row the reranker ranked higher keeps that claim; one row per page; text rows keep relative order; every no-op path returns the input) and the ONE range contract, then the end-to-end on the relational corpus (bodies never name the related entity) with a canned reranker on in-memory PGLite — serial lane because it mocks the reranker module.test/search/relational-intent-memo.test.ts— the default relational pattern set is compiled once per process (identity across calls, stateless sharing).test/config-adaptive-return-keys.test.ts— the adaptive-return / autocut / CRAG search knobs AND the four ranker-wave keys (search.expansion_variant_budget,search.relational_rerank_pin,search.keyword_arm_confidence_floor,search.metadata_boost_gate) are registered inKNOWN_CONFIG_KEYS, sogbrain config seton a documented knob is never a silent no-op.test/longmemeval-sanitize.test.ts— sanitization parity pinning thatINJECTION_PATTERNSfromsrc/core/think/sanitize.tsis the single source of truth (adding a pattern there must cover both<take>framing and<chat_session>framing, no per-surface regex drift).test/openai-compat-multimodal.test.ts— gateway's openai-compatible multimodal path: happy-path single + multi-input embedding, unauthenticated proxy mode, dimension-mismatch guard (throwsAIConfigErrorwith model id + observed + expected pre-storage), default-dim fallback when recipe declaresdefault_dims, HTTP 401 / 400 / malformed-JSON / non-array error paths, and the Voyage/multimodalembeddingsrecipe still routing through its dedicated path. Hermetic via the__setEmbedTransportForTestsseam.test/serve-stdio-lifecycle.test.ts—MCP_STDIO=1env guard: stdin EOF does NOT trigger shutdown when the env is set, SIGTERM still does (guard scope is correct), unset env preserves the CLI lifecycle. Exercises theServeOptions.mcpStdio?: booleantest seam directly so tests don't mutateprocess.env.test/db-lock-fencing.test.ts— fenced lock identity: aDbLockHandlecarries its acquisition fence,refresh()returns true while owned and false after a steal (0-row fenced UPDATE), a stolen-from handle'srelease()is a fenced no-op that leaves the successor's row intact, andstartCycleLockRefresheraborts its controller withLockStolenErroron a fenced miss while serializing ticks (a slow refresh never overlaps the next).test/cycle-lock-steal.serial.test.ts— runCycle steal-abort arc end-to-end: a mid-run steal produces a structured partial report (reason: 'lock_stolen'), runs no further phases, and never touches the successor's lock row; a steal-free cycle completes and releases normally.test/cycle-any-abort-signal.test.ts—anyAbortSignalcombining: pre-aborted inputs, late aborts propagating their reason, duck-typed signal stubs (noaddEventListener) observed via poll, anddispose()detaching the caller-signal listener + clearing the poll timer (the daemon leak class).test/cycle-triage-rescue.test.ts— the dream triage gate:passesTriageGateband arithmetic (floor inclusive, at/above threshold never "rescued"), content-type allowlisting, segment verification throughnormForGrounding(case/curly-quote/dash folding matches, fabricated segments never do), the ≥40-char + dedupe-by-normalized-quote rules, and fail-closed behavior on every malformed verdict shape (null score, missing/short/non-string segments) withminSegments: 0as the kill switch.test/cycle-synthesize-verify.test.ts— the mechanical claim verification pass: span extraction with code fences / inline code / wikilinks / link targets masked, odd-mark paragraphs countedunbalanced, the quote ladder (exact keep → normalized replace with the verbatim slice → near-match replace trimmed to the matched tokens inside one speaker turn), refusal of any match that crosses a speaker label, wrong-speaker attribution, canonical number and date grounding, claim units (a failing sentence, list item or table row leaves the body whole and lands inunverified_claims), source-span + speaker provenance ingrounding.quotes, diff-scoped verification of pre-existing pages against their pre-runpage_versionsrevision, and per-page fail-open.test/cycle-write-path-mini-eval.test.ts— the hermetic $0 write-path mini-eval. A frozen 3-transcript mini-corpus (high / buried / routine bands, placeholder names, deliberately disjoint from the paid Cat 35 corpus so there is no tuning coupling) drives the REALrunPhaseSynthesizeon PGLite: real triage parse + gate incl. the rescue, real fan-out + oneshot drain, real quote verify/repair, real provenance stamp + reverse-write + telemetry. The ONLY stub is the gateway chat transport (__setChatTransportForTests), serving a scripted judge and a scripted child. Scope honesty matters here: a scripted child CANNOT measure whether a prompt change improved model output — that stays the paid benchmark's job (receipts indocs/eval/FIX_WAVE_BASELINES.md). This is the no-API-key regression pin for the MECHANICAL write path, and its salient-unit presence score is the canary that catches emission, chunk-slug-rewrite, and repair-over-deletion regressions in the normal unit lane.test/cycle-repeated-consolidation.test.ts+test/helpers/repeated-consolidation.ts— hermetic three-cycle consolidation pin: realrunCycle(synthesize→extract→extract_facts) on PGLite with a scripted agentic child (gateway tool loop) that writes supported claims, fabricated quotes, speaker swaps and invented numbers into new and existing pages. Pins zero checkable invented claims in active memory (page text, timeline rows, facts), every supported claim kept, and the two unquoted inventions the mechanical check cannot see.scripts/repeated-consolidation-experiment.tsprints the same measurements as JSON for a before/after comparison.test/cycle-synthesize-triage.test.ts/test/cycle-synthesize-triage-calibration.test.ts— triage gate wiring insiderunTriagePass(reports carryrescued/verified_segments,details.triagerescue + token counters, dry-run parity), plus the 25-fixture calibration corpus (10 high / 10 low / 5 buried, all synthetic placeholders) enforcing band-consistent parsing, a ≥80% band-accuracy rubric-drift pin, and that ≥4 of the 5 buried fixtures reach the gate.TRIAGE_VERSIONparticipates in cache validity, so a rubric bump re-judges rather than serving stale verdicts.test/dream-retriage.test.ts—gbrain dream retriagereads THE shared gate: reconcile-queue never cancels a rescued job,--audit-rejectsexcludes rescued files from the reject sample, alongside the spend-gate / dry-run / liveness arcs.test/facts-extract-idea-kind.test.ts— theideaextractor kind: taxonomy coercion (known kinds survive verbatim,ideastaysidea, unknown kinds coerce tofact), prompt shape (the two precomputed system-prompt variants differ in EXACTLY one clause — the low-tier line — both carry the idea definition and the widened enum, and repeat calls return the identical string so prompt caching still hits), and admission wiring (no admission or an admission allowinglow→ label-honestly; a high-only admission → skip-low).test/migrations-v145.test.ts— thefacts.kindCHECK widening: the migration's structure (canonical name, idempotent flag, probe + widened predicate), a fresh PGLite schema admitting anideaINSERT, and an upgrade from a pre-v145 brain swapping the 5-kind constraint for the widened one with a re-run applying nothing.test/queue-stall-parent-unblock.test.ts— the sharedkillJobstail: a stall-exhausted child landschild_done(dead)in its parent's inbox and unblocks the parent, a requeued child doesn't touch the parent, all three reapers route through the tail with their own outcome, and the idempotent stranded-parent sweep self-heals parents whose children were already dead (without unblocking parents that still have a live child).test/queue-started-at-retry.test.ts— every automatic re-run path clearsstarted_at(failJob delayed branch, stall requeue, lease release, promoteDelayed, parent re-claim) so a retried job's wall-clock budget measures execution, not backoff wait; end-to-end survival of the wall-clock sweep on a fresh attempt.test/embed-modality-preserved.test.ts—carryChunkMetadatacarries modality + all code-metadata fields through re-embed merges (an image chunk stays image), plus the write-side contract that omitting modality resets it to text (why the shared list is load-bearing).test/embed-oversize-heal.test.ts— oversize-chunk healing pure core:healOversizedChunkssplit/reindex/metadata-carry (fenced_codechunk_sourcenever coerced),healedChunksToStaleRowsremap (only rows still needing embeddings survive), and thehealOversizedPageChunksorchestrator incl. the freshness guard (a concurrent rewrite between snapshot and write skips the heal — no clobber).test/embed-oversize-heal-drain.serial.test.ts— the heal wired into all three real drains against PGLite:embedStaleForSource,embedStalePages(phase-end closure), andrunEmbedCore --stale(embedAllStale) heal oversized stored rows in place while preserving the embedded sibling's vector.test/embed-stall.test.ts— the embed stall watchdog unit:resolveEmbedStallAbortSecondsenv resolution (default 900; garbage → default;<= 0disables),createEmbedStallWatchdogfire/reset/stop semantics, the run-scoped embedding-API liveness clock, and theassertEmbedNotStalledhandler contract (clean result no-op, stalled result throws).test/jobs-embed-stall-wiring.serial.test.ts— the stall contract at the minion boundary: an embed job whose core result carriesreason: 'stall_timeout'THROWS (job marked failed, banked progress in the message); a clean result resolves with the embed report.test/embed.serial.test.ts—runEmbedCorelifecycle on PGLite withmock.moduleseams: abort-signal threading, the stall-watchdog arc against the real lock table (stall fires → single-flight locks released + summary flushed +reason: 'stall_timeout'surfaces; non-CLI callers get the error RESULT, no process exit; live progress keeps the watchdog quiet), cleanup aborting an in-flight heartbeat refresh, and the heartbeat tick-timeout arc (a never-settling refresh times out per tick viaGBRAIN_EMBED_LOCK_HEARTBEAT_TIMEOUT_MS; 3 consecutive failures →lock_lost+ drain abort).test/handlers-embed-backfill.test.ts— theembed-backfilljob handler's budget-cap classification matrix: default cap dropped for unpriced models,pricing.overridesrestores enforceability,offuncaps, explicit caps fail closed on unpriced models (incl. an explicit $10 equal to the default), the defaulted cap still enforces for priced models, a present-but-garbage cap value keeps the $10 default FAIL-CLOSED (never droppable), and the handler-lane stall watchdog (a wedged drain aborts and fails the job).test/ai/reranker-readiness.test.ts— the purererankerReadinesspredicate: the default Voyage model with and without the key (the fix names the key AND the disable command; an empty-string key counts as absent), shape failures that never throw (unknown provider, no reranker touchpoint, unlisted model, keyless local recipe, garbage input), and agreement withgateway.isAvailable('reranker', model)on an env × model matrix.test/rerank-no-key.serial.test.ts— the gatewayno_keypreflight:RerankError('no_key')before any HTTP call, ONE audit row per process per model with no stderr line, the per-model memo and its test seam, no budget reservation for a skipped rerank (andBudgetExhaustedbefore the transport call when the key IS present), HTTP 401 stayingauth; plusapplyRerankeronno_key— results unchanged, no per-query rows,onSkipfires, a throwing hook never breaks search, and genuine failures never fire it.test/hybrid-reranker-skipped.serial.test.ts— balanced search on PGLite withoutVOYAGE_API_KEY:reranker_skipped (no_key)stamped on the meta with results kept and nothing printed, fresh searches retain the stamp while shared result caching stays disabled with no cache writes even when requested, and with the key present the reranker runs (rerank_scorestamped, no skip entry).test/degraded-stages-recall.test.ts—affectsRecall/RANKING_ONLY_DEGRADED_STAGES:reranker_skippedis ranking-only, every other closed-vocabulary stage affects recall, and a mixed list is degraded iff a recall-affecting stage is present.test/cli-explain-degraded-render.test.ts—formatResult --explainthreads the captured retrieval meta: areranker_skippedstamp renders as the degraded header, a clean run prints no header, and the plain (non-explain) renderer is unchanged.test/doctor-reranker-health.test.ts— the readiness-awarereranker_healthcheck: key absent → warn naming the key and the disable command; key present → ok "ready"; disabled by config row or by the conservative bundle → ok;no_keyskip rows informational once ready; a DB-plane key the CLI folded into the gateway counts; the no-gateway fallback to env > file > DB plane; unknown model, self-host override andauthrows; a brain with no embedding provider is never blamed; audit rows for another model never warn on the active default.test/modes-report-reranker.test.ts—buildModesReportattributes the five reranker knobs, the readiness block flips on the key, config overrides are reflected;formatModesTextprints the runtimeReranker:line plus the per-bundlereranker=… autocut=…line and surfaces a gateway base-URL override asself_hosted;redactReadinessForRemotestripsrequired_key/key_present/fixand keeps the verdict.test/init-reranker-default.test.ts—writeNewInstallRerankerDefault: a Voyage key present → no write regardless of embedding pick; a keyed non-Voyage pick without the key → explicitsearch.reranker.enabled=false; a keyless install → no write; a Voyage key that lives only in the DB config plane or only inconfig.jsoncounts; never-clobber on an existing explicit row.test/import-abort-error.test.ts—runImportpreflight/argv failures throw typedImportAbortErrorinstead of exiting the process; the calling process survives the abort.test/lint-fix-single-pass.test.ts—gbrain lint --fixwalks the tree once andtotal_fixedreports the fixes THIS run applied.test/snapshot-shape-guard.test.ts— PGLite snapshot loader refusal matrix: shape-less version files, dims/model mismatches, and stale schema hashes are all refused; matching hash + shape loads; a migration-handler edit changes the hash.test/stats-health-source-scope.test.ts— source-scoped stats/health/identity: engine-level scoping (every counter confined; degrees/denominators/the islanded predicate scope BOTH edge endpoints; mutating an excluded source moves nothing a scoped caller sees) plus the op layer (remote scalar + federated grants confineget_stats/get_health/get_brain_identity; remote unscoped and the__all__sentinel fail closed to zeros; trusted local keeps the brain-wide view).test/takes-list-subcommand.test.ts—takes listrouting (listis a subcommand, not a slug) + the--limit/--offsetflags: cap/skip/paging, engine default without--limit, barelistunchanged, and invalid values (0, non-numeric) exit 1 with the positive-integer message.test/stale-takes-bigint.test.ts+test/take-proposals.test.ts— 64-bit row normalization at the engine boundary:listStaleTakesrows andtakes propose --json/loadProposalrows come back as NUMBERS (never bigint/string ids or string weights) on both engines, so takes embed/propose survive real Postgres int8 rows.test/llm-json-reasoning-ladder.test.ts—parseLlmJson's reasoning-block recovery ladder: strips a closed or truncated<think>block ONLY after a raw parse fails (valid JSON containing the tag text is untouched), case-insensitive, array payloads, and the facts/atoms extractors routing through it (the ORIGINAL failure reason is preserved when the retry also fails).test/models-per-task-extract-atoms.serial.test.ts—gbrain modelsreportsmodels.dream.extract_atomsthrough the phase's own resolver (pins the narrow-resolver divergence:models.tier.utilityis deliberately ignored; unconfigured falls back to the same tier default the runtime uses).test/conversation-facts-pricing-wiring.test.ts—pricing.overridesreaches every conversation-facts entry point: the strict config registry accepts the key, and direct extraction, the cycle backfill, andtranscripts --factsall price through the operator override.test/budget/no-pricing-registration.test.ts— explicit cost cap + unpriced model: the refusal carries the lookup-and-register guidance as text and structured fields on BudgetTracker, enrich, conversation facts (core and cycle phase) and skillopt; default caps warn (naminggbrain pricing set) and run;gbrain pricing setmerges without clobbering other overrides, validates rates, warns on $0, and the retried run is priced and capped; list/unset; and the trust boundary (no operation registers prices, thin clients refusegbrain pricing). PGLite + gateway test transports.test/extract-atoms-explicit-cap-no-pricing.test.ts— extract_atoms with an explicitcycle.extract_atoms.budget_usd: an unpriced chat model or embed route is refused (statuswarn, no model call,details.no_pricingwithregister_command), the rollup records an expected limit not a halt, doctor'sextract_healthoverlay names the command, and aftergbrain pricing setthe retried run extracts and the record clears; a default cap still warns and runs. PGLite +_chatseam + embed transport stub.test/cycle/extract-atoms-model-config-fail-soft.test.ts— a throwinggetConfigduring extract_atoms model resolution falls back to the tier default instead of rejecting the phase.
The 20 heaviest PGLite-only files in test/e2e/ (by scripts/e2e-weights.json)
moved out of the sequential Postgres runner into the lanes that run on every PR.
Each met the move criterion: it constructs PGLite (or spawns a PGLite CLI)
directly, imports nothing from test/e2e/helpers.ts, has no
DATABASE_URL/hasDatabase gate, and its header confirmed no Postgres use.
sync-delegation-under-serve.serial and dream-synthesize-pglite stayed in
test/e2e/ because named e2e.yml jobs run them. Assertions are unchanged;
executed-test counts match the E2E runs. A file becomes serial when it mutates
process-global state and slow when it takes about 30 s or more. The remaining
PGLite-only E2E files are a TODOS.md item decided from the pilot measurement.
Source files whose only E2E owner moved (src/commands/claw-test.ts,
src/core/claw-test/**, src/core/brain-resolver.ts, src/commands/mounts.ts,
src/commands/connect.ts, src/core/connect-probe.ts,
src/commands/embed-facts-delegate.ts) no longer have an E2E_TEST_MAP row,
so a change to them selects all E2E (fail-closed); their owners run in every PR.
| Former path | New path | Lane | Command | Lane reason |
|---|---|---|---|---|
test/e2e/claw-test.test.ts |
test/claw-test.slow.test.ts |
slow | bash scripts/run-slow-tests.sh test/claw-test.slow.test.ts |
over 30 s (harness subprocess runs) |
test/e2e/init-fresh-pglite.test.ts |
test/init-fresh-pglite.slow.test.ts |
slow | bash scripts/run-slow-tests.sh test/init-fresh-pglite.slow.test.ts |
over 30 s (CLI subprocesses) |
test/e2e/mounts-routing-pglite.test.ts |
test/mounts-routing-pglite.slow.test.ts |
slow | bash scripts/run-slow-tests.sh test/mounts-routing-pglite.slow.test.ts |
about 30 s (two persistent PGLite brains, CLI spawns) |
test/e2e/qm-provisioning.test.ts |
test/qm-provisioning.test.ts |
unit | bun test test/qm-provisioning.test.ts |
no process-global state, under 30 s |
test/e2e/minions-field-report-repro.test.ts |
test/minions-field-report-repro.test.ts |
unit | bun test test/minions-field-report-repro.test.ts |
no process-global state, under 30 s |
test/e2e/fresh-install-pglite.test.ts |
test/fresh-install-pglite.serial.test.ts |
serial | bash scripts/run-serial-tests.sh test/fresh-install-pglite.serial.test.ts |
mutates process.env and console |
test/e2e/remote-privacy-journeys.test.ts |
test/remote-privacy-journeys.serial.test.ts |
serial | bash scripts/run-serial-tests.sh test/remote-privacy-journeys.serial.test.ts |
constructs PGLite outside beforeAll (isolation rule R3) |
test/e2e/serve-stdio-roundtrip.test.ts |
test/serve-stdio-roundtrip.test.ts |
unit | bun test test/serve-stdio-roundtrip.test.ts |
no process-global state, under 30 s |
test/e2e/serve-http-surface-ceiling.test.ts |
test/serve-http-surface-ceiling.test.ts |
unit | bun test test/serve-http-surface-ceiling.test.ts |
no process-global state, under 30 s |
test/e2e/skillpack-flow.test.ts |
test/skillpack-flow.test.ts |
unit | bun test test/skillpack-flow.test.ts |
no process-global state, under 30 s |
test/e2e/v0_28_5-fix-wave.test.ts |
test/v0_28_5-fix-wave.serial.test.ts |
serial | bash scripts/run-serial-tests.sh test/v0_28_5-fix-wave.serial.test.ts |
mutates process.env |
test/e2e/backfill-perf-pglite.test.ts |
test/backfill-perf-pglite.test.ts |
unit | bun test test/backfill-perf-pglite.test.ts |
no process-global state, under 30 s |
test/e2e/connect-bearer.test.ts |
test/connect-bearer.test.ts |
unit | bun test test/connect-bearer.test.ts |
no process-global state, under 30 s |
test/e2e/bootstrap-hook-under-serve.serial.test.ts |
test/bootstrap-hook-under-serve.serial.test.ts |
serial | bash scripts/run-serial-tests.sh test/bootstrap-hook-under-serve.serial.test.ts |
mutates process.env; already a serial file |
test/e2e/upgrade-bun-link-arc.serial.test.ts |
test/upgrade-bun-link-arc.serial.test.ts |
serial | bash scripts/run-serial-tests.sh test/upgrade-bun-link-arc.serial.test.ts |
mutates process.argv; already a serial file |
test/e2e/dream-synthesize-chunking.test.ts |
test/dream-synthesize-chunking.serial.test.ts |
serial | bash scripts/run-serial-tests.sh test/dream-synthesize-chunking.serial.test.ts |
mutates process.env |
test/e2e/search-readiness-http.test.ts |
test/search-readiness-http.test.ts |
unit | bun test test/search-readiness-http.test.ts |
no process-global state, under 30 s |
test/e2e/transcripts-ingest-pglite.test.ts |
test/transcripts-ingest-pglite.test.ts |
unit | bun test test/transcripts-ingest-pglite.test.ts |
no process-global state, under 30 s |
test/e2e/bootstrap-harness-lifecycle.serial.test.ts |
test/bootstrap-harness-lifecycle.serial.test.ts |
serial | bash scripts/run-serial-tests.sh test/bootstrap-harness-lifecycle.serial.test.ts |
mutates console; already a serial file |
test/e2e/fact-backfill-resident.test.ts |
test/fact-backfill-resident.test.ts |
unit | bun test test/fact-backfill-resident.test.ts |
no process-global state, under 30 s |
E2E tests live in test/e2e/ and run against real Postgres+pgvector (require DATABASE_URL), except where noted as PGLite in-memory (no DATABASE_URL needed). One file outside the directory also rides the e2e lane: test/phantom-redirect-engine-parity.test.ts (Postgres arm; see the file taxonomy above).
-
test/e2e/write-attribution-postgres.test.ts— runstest/write-attribution.test.tson Postgres, direct and through transaction-mode PgBouncer (scripts/e2e-backend-matrix.txt): the coordinator's transaction-local actor reaches page versions, live revisions, facts, takes and timeline rows. -
test/e2e/facts-separation-postgres.test.ts— real-Postgres parity for cross-session facts, supersession, and the pre-limitunconsolidatedOnlypredicate used by consolidation. -
bun run test:e2eruns Tier 1 (mechanical, all operations, no API keys). Includes dedicated cases for the postgres-engineaddLinksBatch/addTimelineEntriesBatchbind path — postgres-js's JSONB bind (jsonb_to_recordset(($1::jsonb)->'rows')) differs from PGLite's and gets its own coverage. -
test/e2e/search-quality.test.ts— search quality against PGLite (no API keys, in-memory). -
test/e2e/graph-quality.test.ts— knowledge graph pipeline (auto-link via put_page, reconciliation, traversePaths) against PGLite in-memory. -
test/e2e/jsonb-batch-poison-postgres.test.ts— the real-Postgres half of the JSONB batch-poison lock (the engine whose bind path differs). Seeds free-text "poison" context (Zoom URL with?pwd=, commas, quotes, Windows backslash path, braces, em-dash) and asserts the links/timeline/takes batch writers do not error with "malformed array literal"; also asserts NUL is stripped from free-text bodies (context/summary/detail/claim) and still rejected in identity fields. Lone-surrogate lock: a lone UTF-16 surrogate in free text (the22P02class on Supabase) well-forms to U+FFFD across batch + scalar paths (incl. timeline + takesource), while a surrogate in an identity field still rejects the batch.DATABASE_URL-gated. -
test/e2e/postgres-jsonb.test.ts— round-trips all 5 JSONB write sites (pages.frontmatter,raw_data.data,ingest_log.pages_updated,files.metadata,page_versions.frontmatter) against real Postgres and assertsjsonb_typeof='object'plus->>'key'returns the expected scalar. Guards against the double-encode bug. -
test/e2e/integrity-batch.test.ts— parity forscanIntegrity's batch-load fast path vs sequential. Cases (dedup, hits, validate, topPages) seed a fixture and assert both paths return identical results. Dedup case uses raw SQL viagetConn().unsafe()to seed a(test-source-2, people/alice)row alongside the default-source row, sinceengine.putPagedoesn't take asource_id. Pins multi-source overcounting; the "multi-source duplicate slugs scan once" case expects both batch + sequential paths to report 2. -
test/e2e/jsonb-roundtrip.test.ts— companion regression against the 4 doctor-scanned JSONB sites. Assertion-level overlap withpostgres-jsonb.test.tsis intentional defense-in-depth: if doctor's scan surface drifts from the actual write surface, one of these tests catches it. -
test/e2e/sync.test.ts—--skip-failedfailure-loop test alongside happy-path tests: broken file →performSyncreturnsblocked_by_failureswith grouped breakdown →performSync({skipFailed: true})advances bookmark and returnsAcknowledgeResultwith code summary → second broken file → second cycle. Saves and restores the user's real~/.gbrain/sync-failures.jsonlso the test is hermetic. Asserts bookmark gating, JSONL state, dedup across paths, summary aggregation, and the literal doctor-rendering string format. -
test/e2e/upgrade.test.ts— check-update against real GitHub API (network required). -
test/e2e/minions-shell-pglite.test.ts— PGLite--followinline shell-job path (in-memory, noDATABASE_URLrequired) — the path the minion-orchestrator skill documents for dev use. -
test/e2e/job-isolation.test.ts— process isolation on real Postgres (DATABASE_URL-gated, wired EXPLICITLY into.github/workflows/e2e.ymltier1 — the workflow runs only named files): a concurrency-3 isolated drain through real child processes (thefake-run-child.mjsfixture — real spawns, no child DB pools), and the REALjobs run-childCLI entrypoint end-to-end (engine bootstrap incl. the child's own pools, quiet handler registry, token validation, outcome protocol). -
test/e2e/sync-reconcile-postgres.test.ts— the sync reconcile's real-Postgres array-parameter binding path (DATABASE_URL-gated). Wired EXPLICITLY into.github/workflows/e2e.ymltier1 beside job-isolation, and listed in the selected-e2e EXCLUDE set so a PR touching sync.ts doesn't run it a second time there. -
test/e2e/pglite-cli-exit.serial.test.ts— real spawned-CLI exit behavior on PGLite (in-memory, noDATABASE_URL): read commands (search/get/query) exit 0 promptly; CLI_ONLYcaptureexits clean and frees the single-writer lock; the teardown describes pin every disconnect site — a failed op exits 1 with the error on stderr, and the dashboard, read-only-timeout, doctor, anddream --dry-runpaths all exit with no force-exit banner. -
test/e2e/pgbouncer-teardown.test.ts— PgBouncer TRANSACTION-mode teardown. Pins the bug CLASS, not timings: a CLI op against a txn-mode pooled URL exits 0 with intact stdout and does NOT ride the 10s hard-deadline backstop (theengine.disconnect() did not returnbanner is the smoking gun). Gated byGBRAIN_PGBOUNCER_URL+GBRAIN_PGBOUNCER_DIRECT_URL(NOTDATABASE_URL) — set automatically bybun run ci:local'spgbouncercompose service. Both URLs survive the E2E runner and preload scrub, while CLI children clear ordinary database overrides so the pooled URL in their isolated config wins. Selected CI runs require a nonzero executed-test count (GBRAIN_CI_REQUIRE_PGBOUNCER=1); missing targets or an all-skipped file fail the gate. It skips gracefully elsewhere. Uses a DEDICATEDgbrain_pgbouncer_testdatabase so it never races thegbrain_testTRUNCATE fixtures. -
test/e2e/volunteer-context-postgres.test.ts—volunteer_contexton REAL Postgres (engine parity beyond the hermetic PGLite unit suite): resolution arms through the actual op handler, the fire-and-forget volunteer-event sink landing rows, the stats join, and the RLS pin thatcontext_volunteer_eventshas ROW LEVEL SECURITY enabled (keeps the v35 auto-RLS event trigger honest for migration-created tables).DATABASE_URL-gated. -
test/e2e/openclaw-reference-compat.test.ts—check-resolvable+ skillpack install-model against a minimal AGENTS.md workspace fixture (test/fixtures/openclaw-reference-minimal/), regression guard for the OpenClaw deployment shape. -
test/e2e/workspace-generic-compat.test.ts— always-on (PGLite, no binary): pins the INSTALL_FOR_AGENTS.md "any repo with a workspace" contract againsttest/fixtures/generic-agents-workspace/(Hermes is the motivating consumer):cwd_walk_updetection, theGBRAIN_SKILLS_DIRoverride,check-resolvableon a root AGENTS.md, and scaffold additivity + refuse-overwrite. The real Hermes-behavior proof is the door suite below. -
test/e2e/install-real-hermes.serial.test.ts— the hermes "door": realhermesbinary + realhermes mcp addhandshake (full-catalog tool discovery; the count tracks the op catalog, so the test asserts discovery happened, not a number) + a paidhermes -zrecall turn against a seeded brain. Triple-gated:GBRAIN_REAL_HERMES_E2E=1(explicit opt-in — run-e2e.sh scrubs GBRAIN_*, so it can never fire underbun run test:e2e) + resolvable binary + non-empty ANTHROPIC key (anthropic-pinned on purpose: a second provider key flips hermes provider-auto into a mis-routed 401). Hermetic HOME + HERMES_HOME with a tripwire on the operator's real config; evidence copies toGBRAIN_E2E_EVIDENCE_DIRfor CI upload. Venue: heavy-tests.yml (real-agent-e2e+hermes-doorjobs). -
test/e2e/install-real-grok.serial.test.ts— the grok "door" (xAI Grok Build; every asserted shape observed against the pin indocs/mcp/GROK-CLI-PIN.md). SPLIT-GATED, a deliberate divergence from the hermes door: grok'smcp add/list/doctorrun keyless, so the compat tier (version-shape pin, documented-shapegrok mcp add gbrain -- gbrain serve --surface verbsvia a PATH-staged bin dir, saved-TOML asserts viaBun.TOML.parse,mcp doctorhandshake proving the seven-verb surface, vendor-fallback provenance guard, direct-TOML surface) needs onlyGBRAIN_REAL_GROK_E2E=1+ a resolvable binary; the paid SMOKE additionally needs a non-emptyXAI_API_KEYand asserts a PER-RUN NONCE fact (grok has fs/shell tools — the committed fact is greppable, so recall of it proves nothing) with web search disabled.mcp addis lazy (exit 0 always) —mcp doctor <name> --jsonis the honest discriminator (exit 0/1 observed). Hermetic HOME + GROK_HOME + tmp cwd on every spawn (grok reads vendor MCP configs for trusted folders and loads.envrcfrom cwd); bounded tripwire over the operator's real~/.grokconfig/credential files (volatile paths excluded — grok rewrites logs/sessions/bin/docs every run) + a checkout guard that no.grok//.mcp.jsonappeared in the repo root. Venue: heavy-tests.yml (real-agent-e2e+grok-doorjobs); run directly viaGBRAIN_REAL_GROK_E2E=1 bun test test/e2e/install-real-grok.serial.test.ts. -
test/e2e/install-real-opencode.serial.test.ts— the opencode "door" (SST opencode; every asserted shape observed against the pin indocs/mcp/OPENCODE-CLI-PIN.md). SPLIT-GATED a step past the grok door: opencode's anonymous FREE TIER drives MCP tool calls keyless, so even the nonce SMOKE runs in the keyless tier — T1 bare-semver version pin (the SST-vs-claimant discriminator), T2 documented-shapeopencode mcp add gbrain --env … -- gbrain serve --surface verbs+ the honestopencode mcp listdiscriminator (it SPAWNS every server;✓/✗text is the assertion surface — exit code is 0 even on failure, andmcp debugis OAuth-only), T2b spawn-gate CANARY (a project-config decoy is spawn-attempted with NO trust prompt — if this ever gates, the bootstrap user-global scope default's rationale changed: re-observe), T3 writer parity (gbrain'sopencode-json.tsoutput handshakes through the real binary; cross-tool preservation both ways), T4 keyless SMOKE (per-run nonce + STRUCTURALgbrain_*tool_use proof viaparseOpencodeJsonl,--format json). The paid T5 anthropic leg additionally needs a non-emptyANTHROPIC_API_KEYand self-validates the pinned model id against the authedopencode modelslist BEFORE any spend. Hermetic HOME + both XDG dirs + tmp cwd on every spawn;--pureon every probe (mcp listautoloads plugins — a code-execution surface); bounded tripwire over the operator's real opencode configs/auth.json + a repo-root checkout guard. Venue: heavy-tests.yml (real-agent-e2e+opencode-doorjobs, plus the schedule-onlyopencode-door-canarylatest-version leg — continue-on-error, a pin-refresh signal, never a gate); run directly viaGBRAIN_REAL_OPENCODE_E2E=1 bun test test/e2e/install-real-opencode.serial.test.ts.
Door cadence policy: the NEWEST door agent runs at nightly/schedule cadence (currently opencode, whose canary leg also tracks latest); a door drops to label-only (real-agent-e2e) after 2 stable monthly cycles with unchanged pins. Rationale: churn concentrates in the newest integration; steady-state doors pay for themselves on demand, not nightly.
test/helpers/tty-harness.ts+test/tty-harness.test.ts— the DX real-PTY harness (Bun.spawn({terminal:})): pure text/timing helpers unit-tested with zero subprocesses, plus three live PTY smokes againstshguarded bydescribe.skipIf(!ptySupported()). The harness itself is a dev instrument surface — its consumerscripts/dx-explore.tsnever runs in CI (transcripts land in gitignored.context/dx-runs/); seedocs/guides/bootstrap.mdfor the scenario runbook.test/e2e/search-swamp.test.ts— reproduces the source-swamp case. Seeds a curatedoriginals/talks/article-outline-fat-codepage against two<fork>/chat/pages stuffed with the same multi-word phrase. Asserts the article wins keyword AND vector ranking, thatdetail=highlets the chat swamp re-surface, and thatsource_idpasses through the two-stage CTE intact. PGLite in-memory.test/e2e/search-exclude.test.ts—test/+archive/pages hidden by default,include_slug_prefixesopts back in, caller-suppliedexclude_slug_prefixesadds to defaults. Both keyword and vector search paths.test/e2e/engine-parity.test.ts— Postgres ↔ PGLite top-result and result-set parity forsearchKeyword+searchVector(Postgres ranks pages then picks best chunk while PGLite returns chunks directly, so the source-boost behavior needs parity coverage). Skips withoutDATABASE_URL.test/e2e/postgres-bootstrap.test.ts— exercisesPostgresEngine.initSchema()directly against a real Postgres database: bootstrap → SCHEMA_SQL → migrations converge from a legacy brain shape, and a brain already at LATEST is an idempotent no-op. Live wedge-class convergence cases rewind a brain to an old schema shape and assert fullinitSchemaconvergence: pre-v121 timeline, pre-v143dream_verdicts(including that pre-existing rows keep theirjudged_at-derived TTL instead of gaining a fresh 30 days), and pre-v7/pre-v136minion_jobsshapes. Also covers the standalonedb.initSchemareplay path fromsrc/core/db.ts, which shares the same bootstrap. Skips withoutDATABASE_URL.test/e2e/http-transport.test.ts—gbrain serve --httpend-to-end against real Postgres: bearer auth round-trip,last_used_atSQL-level debounce,mcp_request_logrow insertion on success and auth_failed paths,/healthDB-down → 503 (DB-probing health check), and the dispatch round-trip with a real operation. Skips withoutDATABASE_URL.test/e2e/serve-http-oauth.test.ts— real-Postgres E2E againstgbrain serve --httpwith full OAuth 2.1. Spawns a subprocess server, registers a client via the CLI, mintsclient_credentialstokens, exercises the/mcpJSON-RPC pipeline. Real DCR/registerHTTP-level response-shape test (assertstypeof body.client_id_issued_at === 'number'over the wire, RFC 7591 §3.2.1); real CLI subprocess test forrevoke-client(registers → mints token → revokes viaexecSync→ asserts token rejected at/mcp→ asserts re-run exits 1); server fixture flips on--enable-dcrso/registeris reachable. bun execSync env-inheritance contract: bun'sexecSyncdoes NOT inherit env mutations done viaprocess.env.X = ..., only OS-level env from before bun started. helpers.ts loads.env.testingand setsDATABASE_URLviaprocess.envmutation, which is invisible to subprocesses unlessenv: { ...process.env }is passed explicitly — every subprocess call in this file passesenv: { ...process.env }. The same contract applies to the sibling sync/cycle/dream/claw-test E2Es.afterAllcleanup is guarded onclientId(won't throw ifbeforeAllfailed before registration); cleanup errors surface to stderr without throwing so real test failures aren't masked. Also covers the trust boundary: an HTTP MCPsubmit_jobforname: "shell"MUST reject with a permission error (request handler setsremote: trueandsubmit_job's protected-name guard fires), and the same guard rejects subagent submission. Skips withoutDATABASE_URL.test/e2e/sync-parallel.test.ts—DATABASE_URL-gated. 60-file Postgres sync at concurrency=4 imports all + no connection leak (probespg_stat_activitybefore/after to confirm worker engines disconnected). 120-file serial-vs-parallel benchmark printsSYNC_PARALLEL_BENCH N files | serial=Xms | parallel(4)=Yms | speedup=Zx. Asserts parallel ≤ serial × 1.5 (CI-noise tolerant; not a strict speedup gate).test/e2e/multi-source-bug-class.test.ts— PGLite in-memory regression suite pinning every multi-source bug site:listAllPageRefsordering by(source_id, slug),getPagewith sourceId picks the right(source, slug)row,extract-takesprocesses both overlappingpeople/alicerows independently,listPagesfilters correctly withPageFilters.sourceId,addLinksBatchwithfrom/to_source_idtargets the right rows,validateSourceIdrejects path traversal, reverse-write disk layout usesbrainDir/.sources/<id>/<slug>.mdfor non-default sources,copyMigrationSourceslands source metadata before overlapping-slug pages. NoDATABASE_URLneeded. Wired intoscripts/e2e-test-map.tsso changes to extract-takes / patterns / synthesize / embed / extract / migrate-engine auto-trigger it.test/e2e/migrate-engine-sources-postgres.test.ts—DATABASE_URL-gated companion for the legacy engine copier (gbrain migrate --to pglite, and--to postgreswithmigrate.graduation false): migrates a PGLite brain carrying two non-default sources with overlapping slugs into real Postgres and assertscopyMigrationSourcescreated everysourcesFK parent (config JSONB intact, not double-encoded) before any page write. Unit-level manifest identity (crash manifest resumes only against the SAME target; legacy engine-only manifests start fresh) istest/migrate-engine-resume.test.ts.test/e2e/facts-fence-reconcile-postgres.test.ts—DATABASE_URL-gated round-trip for the escape-aware fence parser: renders a## Factsfence whose cells carry literal pipes, backslashes (Windows paths), and empty cells viarenderFactsTable, runs the wipe-and-reinsert reconcile (runExtractFacts) on real Postgres, and asserts every cell survives byte-identically with no column shift.test/e2e/source-isolation-pglite.test.ts— PGLite in-memory regression suite pinning the source-isolation seal at two layers. Engine layer:searchKeyword/searchVector/searchKeywordChunks/listPages/getPage/traverseGraph/traversePathsapplysourceId(scalar fast path) andsourceIds(array path) correctly across both engines. Op-handler layer: routes throughsourceScopeOpts(ctx)so aread+write-scoped OAuth client bound to--source dept-xcannot see rows from neighboring sources viasearch,query,list_pages,get_page, orfind_experts. Covers bothctx.sourceId(single-source clients) andctx.auth.allowedSources(federated_read clients) precedence; federated array wins over scalar wins over nothing. NoDATABASE_URLneeded.test/e2e/think-source-isolation-pglite.test.ts— PGLite in-memory suite pinning thethinkgather stage's source scope: seeds three sources with cross-source links and embedded takes, then assertsrunGatherunder a federatedsourceIdsgrant (and under a scalarsourceId) keeps every stream — hybrid retrieval, takes keyword + vector (searchTakes/searchTakesVector), and thetraversePathsgraph walk — inside the grant while still reaching authorized neighboring sources. NoDATABASE_URLneeded.test/e2e/skill-brain-first.test.ts— doctor reportsskill_brain_firstcheck with structured issues;--fix --dry-runpreviews insertion without writing;--fixapplies the canonical Convention callout idempotently;brain_first: exemptfrontmatter resolves the warn;brain_first_typosurfaces a paste-ready hint; audit JSONL recordsdetected/resolved/fixedtransitions; stable brain emits 0 audit lines/run.- Journey suites (each claimed by an
scripts/e2e-test-map.tsrow; DATABASE_URL-gated unless noted):migrate-engine-pglite-to-postgres.test.ts(whole-brainrunMigrateEnginetransfer incl. the child-process failure arm — config not flipped),takes-write-ops-postgres.test.ts(takes op layer +withPageLockserialization),propose-takes-jsonb-postgres.test.ts+calibration-profile-write.test.ts(JSONB bind shape on real Postgres),engine-parity-cjk.test.ts(cross-engine CJK keyword parity on an identical corpus — both engines routehasCJK()queries through the shared ILIKE builder insrc/core/search/cjk-keyword-sql.ts; top-slug agreement, chunk-grain parity, mixed-query AND semantics, nonexistent-term strictness),code-edges-read-parity.test.ts/ontology-merge-parity.test.ts/chronicle-event-projection-parity.test.ts/health-parity-postgres.test.ts(read-path + getHealth parity),sync-sigkill-resume-postgres.test.ts(real SIGKILL mid-sync; DB-polled checkpoint, stranded-lock reclaim, exactly-once resume),serve-http-source-grant.test.ts(legacy no-grant federated widening vs granted confinement over real/mcp),autopilot-linux-lifecycle.serial.test.ts(PATH-shimmed crontab/systemctl arc, hermetic), and the thin-client daily-driver verb extension insidethin-client.test.ts. - Tier 2 (
test/e2e/skills.test.ts) requires OpenClaw + API keys, runs nightly in CI. test/claw-test.slow.test.ts(slow lane since the lane-move pilot) also covers live mode token-free via shim agents (OPENCLAW_BIN=<sh script>): the success-oracle break path (a do-nothing agent FAILS), the E0 child-friction merge surviving tempdir cleanup, and the upgrade staging + schema-version probe.- If
.env.testingdoesn't exist in this directory, check sibling worktrees:find ../ -maxdepth 2 -name .env.testing -print -quitand copy it here if found. - Run E2E tests without asking permission. When you want to verify behavior, there's a relevant E2E test, or you're shipping anything covered by an E2E suite — spin up the test DB, run the tests, tear down. Don't ask, don't propose it, don't defer. The lifecycle is short (~2-30s startup, sub-minute tests, instant teardown) and the gate value is high. Skipping with "DATABASE_URL unset" is silent regression, not caution.
ALWAYS source the user's shell profile before running tests:
source ~/.zshrc 2>/dev/null || trueThis loads OPENAI_API_KEY and ANTHROPIC_API_KEY. Without these, Tier 2 tests
skip silently. Do NOT skip Tier 2 tests just because they require API keys — load
the keys and run them.
When asked to "run all E2E tests" or "run tests", that means ALL tiers:
- Tier 1:
bun run test:e2e(mechanical, sync, upgrade — no API keys needed) - Tier 2:
test/e2e/skills.test.ts(requires OpenAI + Anthropic + openclaw CLI) - Always spin up the test DB, source zshrc, run everything, tear down.
Key-gated live files that no CI job has keys for are left out of the
scripts/run-e2e.sh default glob, so the nightly full corpus and the local
gates stop counting their skips as discovered coverage. Naming a file on the
command line still runs it (the runner keeps provider keys):
| File | Required key | Command |
|---|---|---|
test/e2e/openrouter-anthropic-subagent-replay.live.test.ts |
OPENROUTER_API_KEY |
OPENROUTER_API_KEY=... bash scripts/run-e2e.sh test/e2e/openrouter-anthropic-subagent-replay.live.test.ts |
test/e2e/openrouter-deepseek-subagent-replay.live.test.ts |
OPENROUTER_API_KEY |
OPENROUTER_API_KEY=... bash scripts/run-e2e.sh test/e2e/openrouter-deepseek-subagent-replay.live.test.ts |
test/e2e/voyage-rerank-live.test.ts |
VOYAGE_API_KEY |
VOYAGE_API_KEY=... bash scripts/run-e2e.sh test/e2e/voyage-rerank-live.test.ts |
test/e2e/voyage-multimodal.test.ts |
VOYAGE_API_KEY |
VOYAGE_API_KEY=... bash scripts/run-e2e.sh test/e2e/voyage-multimodal.test.ts |
test/live/decide-typesafe.live.test.ts |
TYPESAFE_API_KEY (or JEV_TYPESAFE_API_KEY) + GBRAIN_LIVE_TYPESAFE=1 |
GBRAIN_TEST_KEEP_PROVIDER_KEYS=1 GBRAIN_LIVE_TYPESAFE=1 TYPESAFE_API_KEY=... bun test test/live/decide-typesafe.live.test.ts |
The sequential E2E runner requires Python 3 for standard-library XML validation
of Bun's native JUnit reports. CI and the local Docker runner provide it; direct
host runs must have python3 on PATH before launching tests.
setupDB() clears rows while preserving physical schema. Fixtures that seed
fixed legacy-width text vectors use setupLegacyEmbeddingDB() instead: it
establishes the canonical test shape after clearing the database, including
facts and takes, so a preceding CLI-init test cannot change their assumptions.
Custom-dimension and migration tests continue using ordinary setupDB().
For fixtures testing schema/index creation or source-scoped cleanup, preserve
that lifecycle and derive incidental text-vector widths from the database.
You are responsible for spinning up and tearing down the test Postgres container. Do not leave containers running after tests. Do not skip E2E tests, do not ask permission to run them — see the "run without asking" rule above.
- Check for
.env.testing— if missing, copy from sibling worktree. Read it to get the DATABASE_URL (it has the port number). - Check if the port is free:
docker ps --filter "publish=PORT"— if another container is on that port, pick a different port (try 5435, 5436, 5437) and start on that one instead. - Start the test DB:
Wait for ready:
docker run -d --name gbrain-test-pg \ -e POSTGRES_USER=postgres -e POSTGRES_PASSWORD=postgres \ -e POSTGRES_DB=gbrain_test \ -p PORT:5432 pgvector/pgvector:pg16
docker exec gbrain-test-pg pg_isready -U postgres - Bootstrap the schema (required — fresh containers have no
oauth_clients,mcp_request_log,pagesetc.; tests likeserve-http-oauth.test.tswill fail withrelation "oauth_clients" does not existif you skip this):DATABASE_URL=postgresql://postgres:postgres@localhost:PORT/gbrain_test \ bun run src/cli.ts doctor --json > /dev/null 2>&1
gbrain doctortriggersinitSchema()on first connect, which is the canonical way to bring a fresh DB to head.apply-migrations --yesalone does NOT seed the base schema — it runs ALTER-style migrations on top ofinitSchema. Tests that bypass the engine (rawexecSync-spawnedauth register-client) hit the schema directly and need this step to have run first. - Run E2E tests:
DATABASE_URL=postgresql://postgres:postgres@localhost:PORT/gbrain_test bun run test:e2e - Tear down immediately after tests finish (pass or fail):
docker stop gbrain-test-pg && docker rm gbrain-test-pg
Never leave gbrain-test-pg running. If you find a stale one from a previous run,
stop and remove it before starting a new one.
test/data-frontmatter.test.ts and test/frontmatter-security.test.ts pin inert
frontmatter parsing, opaque serialization, scalar compatibility, and import
errors. test/authorization-boundaries.test.ts covers scalar source grants,
foreign/private facts, and delegated tool exclusions.
test/oauth-consent-security.test.ts covers pending consent, CSRF, policy
changes, duplicate decisions, and uncertain completion. Production HTTP flows
live in test/e2e/serve-http-consent.test.ts; client-lock races and grant rollback
on Postgres live in test/e2e/oauth-grant-transactions.test.ts.
test/minions-submission-authority.test.ts covers submission schemas, durable
policy, lifecycle operations, legacy approval snapshots, file confinement, and
both workers. test/e2e/minions-authority-parity.test.ts exercises real Postgres
JSONB and authorization behavior. Filesystem write concurrency is pinned by the
existing fence and timeline suites.
test/guarded-http.test.ts and test/guarded-http-tls.serial.test.ts cover DNS,
TLS identity and ports, redirects, deadlines, body limits, and cleanup. CI runs
these boundaries on Bun 1.4.0 and 1.4.2, audits root and admin dependencies, and
executes scripts/test-gitleaks-config.sh to prove fixture exceptions still
report an unrelated secret in the same file. scripts/scan-worktree-secrets.sh
scans tracked files plus new files eligible for commit; tracked ignored files
remain included. Full-history scans use gitleaks git . --log-opts=--all from a
complete clone and reports must remain private.
The Docker gate sets GBRAIN_CI_DISABLE_TEST_ENV_FILE=1 so a bind-mounted
developer .env.testing cannot add credentials or change the isolated test
database. Explicit local provider E2E runs can continue using that file.