Skip to content

Rehearse recovering from a failed migration, not just restoring from one - #13

Merged
GunsNR merged 3 commits into
mainfrom
claude/phase2-migration-recovery-drill-qxgyn0
Sep 1, 2026
Merged

GunsNR merged 3 commits into
mainfrom
claude/phase2-migration-recovery-drill-qxgyn0

Conversation

@GunsNR

@GunsNR GunsNR commented Aug 31, 2026 •

Copy link
Copy Markdown
Owner

Closes the recovery half of Phase 2 criterion 3 — "Restoration and rollback or forward-fix have been rehearsed."

Restoration has been rehearsed since 2026-08-25 and runs in CI. The recovery half existed only as prose in the runbook, and prose is not a rehearsal: the entire value of a recovery procedure is whether it works when someone follows it under pressure.

npm run db:recovery-drill stages a real incident against disposable PostgreSQL and executes the documented recovery against it, asserting 34 steps in ~25s. CI runs it on every push.

Strategy: forward-fix, with restore as the escape hatch

By necessity rather than preference. prisma migrate resolve --rolled-back is valid only for a migration Prisma recorded as failed — it errors on one that succeeded. Undoing a successful migration would need a hand-written down script (migrate diff --to-migrations --script + db execute) plus a hand-edit of _prisma_migrations: two unsupported operations that manufacture exactly the drift db:drift exists to catch. And no down script can recover data — re-adding a dropped column yields a column of nulls.

Restore is rehearsed alongside it, because it is the only remedy when a migration destroyed or transformed data.

The incident

A product contradiction rather than a contrivance. The staged migration denormalises the UTC day onto MeasurementRun, then asserts one run per project per day. Gate 1 requires two runs on one day to stay two distinct runs, and the shared fixture contains exactly that — so it fails with a real 23505 for a reason a reviewer would recognise:

Error: P3018   Database error code: 23505
ERROR: could not create unique index "MeasurementRun_projectId_runDay_key"
DETAIL: Key ("projectId", "runDay")=(…, 2026-06-15) is duplicated.

It is the migration shape that actually breaks deployments: additive, reviewed, green on an empty database and on a sparse local copy, red only where the data is real.

What measuring the failure corrected

The drill observes what the failure left behind rather than assuming it, and that changed the documentation. For a migration whose statements can all run inside a transaction, PostgreSQL rolls the DDL back entirely, so the schema is untouched and the history is what breaks: _prisma_migrations keeps a failed row and every subsequent migrate deploy refuses with P3009. For that class, the incident is a wedged deployment pipeline, not a corrupt schema.

Scoped deliberately. CREATE INDEX CONCURRENTLY and several ALTER TYPE forms cannot run in a transaction and can leave a half-applied schema. The drill rehearses the transactional case only, and all three documents now say so — the operations runbook included, since it is the one an operator reads under pressure.

Isolation

The rehearsal migrations live in supertool/scripts/rehearsal-migrations/, never in prisma/migrations/. The drill copies the real history into a temp directory and appends them there — the real directory is only ever a cpSync source, so the product's history is read and never written. Adding a failing migration to the real history to satisfy a test would ship an unreviewed schema change to every environment to prove something about a database nobody has. Both tests/migration-recovery-drill.test.ts and a runtime precondition in the drill fail if a fixture ever lands there.

prisma/migrations/ is unchanged: its git tree hash is identical on both sides of this diff (d9a2a39…). No migration was rewritten or deleted.

What is proven

  • A migration failing on real data is recoverable without data loss, using only commands Prisma supports.
  • Recovery returns the database to zero drift against the declared schema; migrate deploy and migrate status are clean afterwards.
  • Tenant data, relationships and integrity are unchanged — four integrity snapshots (pre-incident, post-recovery, post-forward-fix, restored) compared field by field: row counts, orphan checks, observed/failed/unavailable provenance counts, same-day run identity, API-key scopes and quotas, live job-idempotency enforcement. All four identical.
  • The forward-fix applies cleanly and the app can still write through it — verified by writing and reading back org/project/run/observation through the ordinary generated Prisma client.
  • The forward-fix's blast radius is exactly what it claims (migrate diff --script touches runDay and nothing else).
  • The restore path works: pre-migration dump restores, real history deploys onto it with zero drift, integrity matches.

Fail-closed, verified

Deliberately broken three ways, each failing the drill with a message naming what went wrong:

Sabotage Result
Remove UNIQUE from the incident migration "The incident migration SUCCEEDED … proves nothing"
Make the forward-fix column NOT NULL "runDay is NOT NULL; the running application could not insert a run"
Make the forward-fix DELETE duplicate rows "Same-day runs disappeared; the forward-fix destroyed the data it was fixing"

What remains unproven

Written out in docs/evidence/2026-08-31-migration-recovery-drill.md §6: no hosted-provider run; no production volume (so nothing about lock duration or backfill time); synthetic data only; one incident class only; no non-transactional DDL case; no downtime measurement.

Scope

Both rehearsals now share one fixture and one set of integrity checks (scripts/lib/rehearsal-support.mjs) so a claim proven by one means the same thing in the other. db-rehearse.mjs shrinks accordingly — behaviour unchanged, re-run and passing.

No capability status changed. roadmap.ts, capabilities.ts and prisma/migrations/ are untouched. Nothing was deployed, no Railway resource modified, Phase 2 is not marked complete. Merging cannot deploy anything: the repository has one workflow, it has no deploy step, it references no secrets., and its database endpoints are localhost service containers.

Verification

typecheck ✅ · lint ✅ · test ✅ 810 passed / 44 files · build ✅ · PHP lint ✅ · Elementor JSON ✅ · db:rehearse ✅ · db:recovery-drill ✅ 34/34 steps

CI green on b82e6f2 — run 33501252443.

Evidence: docs/evidence/2026-08-31-migration-recovery-drill.{md,json} (local run PostgreSQL 16.13; CI PostgreSQL 16.15; Prisma 6.16.2). Decision recorded as ADR-019.

claude added 3 commits August 31, 2026 21:17
Phase 2 criterion 3 asks that restoration *and* rollback or forward-fix be
rehearsed. Restoration has been rehearsed since 2026-08-25 and runs in CI. The
recovery half existed only as prose in the runbook, and prose is not a
rehearsal: the whole value of a recovery procedure is whether it works when
someone follows it under pressure.

`npm run db:recovery-drill` stages a real incident against disposable
PostgreSQL and executes the documented recovery against it, asserting 34 steps.
CI runs it on every push.

The incident is a product contradiction rather than a contrivance: a migration
that denormalises the UTC day onto MeasurementRun and then asserts one run per
project per day. Gate 1 requires two runs on one day to stay two runs, and the
shared fixture contains exactly that, so the migration fails with a real 23505
for a reason a reviewer would recognise.

Forward-fix is the strategy, by necessity rather than preference.
`migrate resolve --rolled-back` is valid only for a migration Prisma recorded
as failed; it errors on one that succeeded. Undoing a successful migration
would need a hand-written down script plus a hand-edit of _prisma_migrations —
two unsupported operations that manufacture exactly the drift db:drift exists
to catch — and no down script can recover data a destructive migration removed.
Restore is rehearsed alongside it as the escape hatch for that one case.

Measuring the failure rather than assuming it corrected the documentation. On
PostgreSQL, DDL is transactional and Prisma sends a migration as one implicit
transaction, so a failed migration leaves the schema untouched and the history
damaged: every subsequent deploy refuses with P3009. The runbook had implied an
operator should hunt for half-applied DDL. Usually there is none, and the
incident is a wedged deployment pipeline.

The rehearsal migrations live in scripts/rehearsal-migrations/, never in
prisma/migrations/. Adding a failing migration to the real history to satisfy a
test would ship an unreviewed schema change to every environment to prove
something about a database nobody has. A test and a runtime precondition both
fail if a fixture ever lands there.

Both rehearsals now share one fixture and one set of integrity checks, so a
claim proven by one means the same thing in the other.

Fail-closed behaviour verified by deliberately breaking the drill three ways:
removing UNIQUE from the incident migration, making the forward-fix column NOT
NULL, and making the forward-fix DELETE duplicate rows. Each fails the drill
with a message naming what went wrong.

Criterion 3 is satisfied against disposable PostgreSQL and not against a hosted
provider or production-sized data. What remains unproven is written out in
docs/evidence/2026-08-31-migration-recovery-drill.md §6. No capability status
changed and roadmap.ts is untouched: this is durability evidence, not a new
customer-facing capability.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015iMfdYSDKhvMCoZbVQFCBz
The checked-in report JSON was captured locally before the commit that
introduced the drill existed, so its commitSha names the parent. Say so, and
cite CI's run on the real commit instead — where the service log records the
staged 23505 and the job still concluded successfully, which is only possible
if every recovery and integrity assertion passed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015iMfdYSDKhvMCoZbVQFCBz
The runbook said flatly that "a migration that fails partway leaves the schema
untouched", and told the operator not to go looking for half-applied DDL. That
is true only for migrations whose statements can all run inside a transaction,
which is the only class the drill rehearses. CREATE INDEX CONCURRENTLY and
several ALTER TYPE forms cannot, and can genuinely leave a half-applied schema.

The evidence document and the hosted-validation runbook both carried that
caveat already. The operations runbook did not — and it is the one an operator
actually reads under pressure, so it was the worst place for the unqualified
version to live.

Scope the claim to what was observed and name the exception explicitly. A test
now fails the build if the caveat goes missing again.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015iMfdYSDKhvMCoZbVQFCBz
@GunsNR
GunsNR marked this pull request as ready for review September 1, 2026 11:15
@GunsNR
GunsNR merged commit f5ce4f3 into main Sep 1, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants