Repository navigation
Rehearse recovering from a failed migration, not just restoring from one - #13
Merged
Merged
Conversation
Phase 2 criterion 3 asks that restoration *and* rollback or forward-fix be rehearsed. Restoration has been rehearsed since 2026-08-25 and runs in CI. The recovery half existed only as prose in the runbook, and prose is not a rehearsal: the whole value of a recovery procedure is whether it works when someone follows it under pressure. `npm run db:recovery-drill` stages a real incident against disposable PostgreSQL and executes the documented recovery against it, asserting 34 steps. CI runs it on every push. The incident is a product contradiction rather than a contrivance: a migration that denormalises the UTC day onto MeasurementRun and then asserts one run per project per day. Gate 1 requires two runs on one day to stay two runs, and the shared fixture contains exactly that, so the migration fails with a real 23505 for a reason a reviewer would recognise. Forward-fix is the strategy, by necessity rather than preference. `migrate resolve --rolled-back` is valid only for a migration Prisma recorded as failed; it errors on one that succeeded. Undoing a successful migration would need a hand-written down script plus a hand-edit of _prisma_migrations — two unsupported operations that manufacture exactly the drift db:drift exists to catch — and no down script can recover data a destructive migration removed. Restore is rehearsed alongside it as the escape hatch for that one case. Measuring the failure rather than assuming it corrected the documentation. On PostgreSQL, DDL is transactional and Prisma sends a migration as one implicit transaction, so a failed migration leaves the schema untouched and the history damaged: every subsequent deploy refuses with P3009. The runbook had implied an operator should hunt for half-applied DDL. Usually there is none, and the incident is a wedged deployment pipeline. The rehearsal migrations live in scripts/rehearsal-migrations/, never in prisma/migrations/. Adding a failing migration to the real history to satisfy a test would ship an unreviewed schema change to every environment to prove something about a database nobody has. A test and a runtime precondition both fail if a fixture ever lands there. Both rehearsals now share one fixture and one set of integrity checks, so a claim proven by one means the same thing in the other. Fail-closed behaviour verified by deliberately breaking the drill three ways: removing UNIQUE from the incident migration, making the forward-fix column NOT NULL, and making the forward-fix DELETE duplicate rows. Each fails the drill with a message naming what went wrong. Criterion 3 is satisfied against disposable PostgreSQL and not against a hosted provider or production-sized data. What remains unproven is written out in docs/evidence/2026-08-31-migration-recovery-drill.md §6. No capability status changed and roadmap.ts is untouched: this is durability evidence, not a new customer-facing capability. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015iMfdYSDKhvMCoZbVQFCBz
The checked-in report JSON was captured locally before the commit that introduced the drill existed, so its commitSha names the parent. Say so, and cite CI's run on the real commit instead — where the service log records the staged 23505 and the job still concluded successfully, which is only possible if every recovery and integrity assertion passed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015iMfdYSDKhvMCoZbVQFCBz
The runbook said flatly that "a migration that fails partway leaves the schema untouched", and told the operator not to go looking for half-applied DDL. That is true only for migrations whose statements can all run inside a transaction, which is the only class the drill rehearses. CREATE INDEX CONCURRENTLY and several ALTER TYPE forms cannot, and can genuinely leave a half-applied schema. The evidence document and the hosted-validation runbook both carried that caveat already. The operations runbook did not — and it is the one an operator actually reads under pressure, so it was the worst place for the unqualified version to live. Scope the claim to what was observed and name the exception explicitly. A test now fails the build if the caveat goes missing again. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015iMfdYSDKhvMCoZbVQFCBz
GunsNR
marked this pull request as ready for review
September 1, 2026 11:15
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes the recovery half of Phase 2 criterion 3 — "Restoration and rollback or forward-fix have been rehearsed."
Restoration has been rehearsed since 2026-08-25 and runs in CI. The recovery half existed only as prose in the runbook, and prose is not a rehearsal: the entire value of a recovery procedure is whether it works when someone follows it under pressure.
npm run db:recovery-drillstages a real incident against disposable PostgreSQL and executes the documented recovery against it, asserting 34 steps in ~25s. CI runs it on every push.Strategy: forward-fix, with restore as the escape hatch
By necessity rather than preference.
prisma migrate resolve --rolled-backis valid only for a migration Prisma recorded as failed — it errors on one that succeeded. Undoing a successful migration would need a hand-written down script (migrate diff --to-migrations --script+db execute) plus a hand-edit of_prisma_migrations: two unsupported operations that manufacture exactly the driftdb:driftexists to catch. And no down script can recover data — re-adding a dropped column yields a column of nulls.Restore is rehearsed alongside it, because it is the only remedy when a migration destroyed or transformed data.
The incident
A product contradiction rather than a contrivance. The staged migration denormalises the UTC day onto
MeasurementRun, then asserts one run per project per day. Gate 1 requires two runs on one day to stay two distinct runs, and the shared fixture contains exactly that — so it fails with a real23505for a reason a reviewer would recognise:It is the migration shape that actually breaks deployments: additive, reviewed, green on an empty database and on a sparse local copy, red only where the data is real.
What measuring the failure corrected
The drill observes what the failure left behind rather than assuming it, and that changed the documentation. For a migration whose statements can all run inside a transaction, PostgreSQL rolls the DDL back entirely, so the schema is untouched and the history is what breaks:
_prisma_migrationskeeps a failed row and every subsequentmigrate deployrefuses withP3009. For that class, the incident is a wedged deployment pipeline, not a corrupt schema.Scoped deliberately.
CREATE INDEX CONCURRENTLYand severalALTER TYPEforms cannot run in a transaction and can leave a half-applied schema. The drill rehearses the transactional case only, and all three documents now say so — the operations runbook included, since it is the one an operator reads under pressure.Isolation
The rehearsal migrations live in
supertool/scripts/rehearsal-migrations/, never inprisma/migrations/. The drill copies the real history into a temp directory and appends them there — the real directory is only ever acpSyncsource, so the product's history is read and never written. Adding a failing migration to the real history to satisfy a test would ship an unreviewed schema change to every environment to prove something about a database nobody has. Bothtests/migration-recovery-drill.test.tsand a runtime precondition in the drill fail if a fixture ever lands there.prisma/migrations/is unchanged: its git tree hash is identical on both sides of this diff (d9a2a39…). No migration was rewritten or deleted.What is proven
migrate deployandmigrate statusare clean afterwards.migrate diff --scripttouchesrunDayand nothing else).Fail-closed, verified
Deliberately broken three ways, each failing the drill with a message naming what went wrong:
UNIQUEfrom the incident migrationNOT NULLDELETEduplicate rowsWhat remains unproven
Written out in
docs/evidence/2026-08-31-migration-recovery-drill.md§6: no hosted-provider run; no production volume (so nothing about lock duration or backfill time); synthetic data only; one incident class only; no non-transactional DDL case; no downtime measurement.Scope
Both rehearsals now share one fixture and one set of integrity checks (
scripts/lib/rehearsal-support.mjs) so a claim proven by one means the same thing in the other.db-rehearse.mjsshrinks accordingly — behaviour unchanged, re-run and passing.No capability status changed.
roadmap.ts,capabilities.tsandprisma/migrations/are untouched. Nothing was deployed, no Railway resource modified, Phase 2 is not marked complete. Merging cannot deploy anything: the repository has one workflow, it has no deploy step, it references nosecrets., and its database endpoints are localhost service containers.Verification
typecheck✅ ·lint✅ ·test✅ 810 passed / 44 files ·build✅ · PHP lint ✅ · Elementor JSON ✅ ·db:rehearse✅ ·db:recovery-drill✅ 34/34 stepsCI green on
b82e6f2— run 33501252443.Evidence:
docs/evidence/2026-08-31-migration-recovery-drill.{md,json}(local run PostgreSQL 16.13; CI PostgreSQL 16.15; Prisma 6.16.2). Decision recorded as ADR-019.