Skip to content

fix(state): the lock names the binary that wrote it, keeps its creator, and lock --restamp converges a state dir (PMAT-565, closes #565) - #570

Merged
noahgift merged 18 commits into
mainfrom
PMAT-565-lock-writer-provenance
Sep 16, 2026
Merged

noahgift merged 18 commits into
mainfrom
PMAT-565-lock-writer-provenance

Conversation

@noahgift

Copy link
Copy Markdown
Contributor

Closes #565. Covers paiml/infra#605's third signature. Stacked on #569 (and through it #563). The base is PMAT-564-drift-declines-on-empty-scope; this branch gets rebased onto main in order as those merge.

What

  • state::save_lock stamps the writing binary into generator on every write, the same string forjar --version prints. The first time, it moves the value it replaces into a new created_by, so the creator is named rather than erased.
  • Older forjar still parses the file.
  • lock-repair and lock-migrate used to write with a bare fs::write, which skipped the stamp and left a stale .b3 sidecar. Both now go through the writer.
  • lock-restore and lock-tag copy bytes and are named in the contract as the exception.
  • forjar lock --restamp --state-dir D rewrites every <machine>/state.lock.yaml under D whose generator is not this binary, sidecars included. --dry-run lists what it would change, --json reports, and a second run changes nothing.

Why

Four fleet locks under one 1.30.0 binary said forjar 1.1.1, 1.13.1, 1.27.0 and 1.10.0 in generator, while generated_at kept moving.

Evidence

  • tests/falsification_lock_names_its_writer.rs: 6 cases through the writer and the binary. 5 of 5 were RED before the stamp existed; the sixth was added when review found the bypass.
  • 4 mutations were run; M1 kills 6 of 6.
  • Contract: contracts/lock-names-its-writer-v1.yaml.
  • Receipt: docs/audits/impl-PMAT-565-receipt.md.
  • Quorum receipt: .quorum/PMAT-565-lock-writer-provenance.json.

🤖 Generated with Claude Code

@noahgift
noahgift force-pushed the PMAT-564-drift-declines-on-empty-scope branch from 1c2deef to 56441f1 Compare September 15, 2026 17:30
@noahgift
noahgift force-pushed the PMAT-565-lock-writer-provenance branch from 89f1463 to b67cc56 Compare September 15, 2026 17:34
noahgift added a commit that referenced this pull request Sep 15, 2026
… locks it, and gate T green again (PMAT-557, closes #557) (#571)

* chore(PMAT-557): file the ticket — book v1.30.0 (Refs #557)

Pmat-Ticket: PMAT-557
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* ledger: book v1.30.0 — release-goal.sh cut v1.30.0 --next v1.31.0 (Refs #557)

Row measured from the tag: cut 2026-09-13T20:52:01Z, ten PRs, eleven tickets, the dogfood and crux documents. The cookbook field still names 0be3e1ec, which locks forjar 1.29.0; it is replaced by the paiml/forjar-cookbook#21 merge commit that locks 1.30.0 before this PR is opened. next.tag v1.31.0 due 2026-09-15T20:52:01Z. Tickets labelled release:v1.30.0 that v1.30.0 did not ship (PMAT-526, 528, 529) move to release:v1.31.0.

Pmat-Ticket: PMAT-557
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* ledger(PMAT-557): the 1.30.0 dogfood receipt gains its end marker; three shipped rows say completed (Refs #557)

docs/audits/dogfood-1.30.0-receipt.md ends on a complete sentence and carries its one verdict: line (GO), so it was not truncated — it was committed without DOGFOOD-1.30.0-RECEIPT-END, which gate T requires. PMAT-555 (the 1.30.0 cut), PMAT-547 (#548) and PMAT-562 (#563) have merged and their issues are closed; their rows said planned/inprogress, which CB-2112 counts as ISSUE-CLOSED and CB-2115 as orphaned. CB-2112 37 -> 34, CB-2115 53 -> 50.

Pmat-Ticket: PMAT-557
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* ledger(PMAT-557): five orphan issues get roadmap rows on a 1.32.0 milestone (Refs #557)

#558, #559, #561, #566 and #567 were open GitHub issues with no roadmap row (CB-2115 ORPHAN-GITHUB). Each is minted from its issue so the id is the issue number, put on the new 1.32.0 milestone, and carries release: 1.32.0 so CB-2114 does not grow. CB-2115 53 -> 45 on this branch; the rows for #560, #564 and #565 arrive with PRs #568, #569 and #570.

Pmat-Ticket: PMAT-557
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* ledger(PMAT-557): v1.30.0 names the cookbook commit that locks it — paiml/forjar-cookbook#21 (Refs #557)

60acf9c9 is the squash of forjar-cookbook#21: forjar = "1.30", Cargo.lock pins 1.30.0, cargo check --workspace --locked clean. 0be3e1ec, which the cut script copied from the previous row, locks 1.29.0 and gate T refused it.

Pmat-Ticket: PMAT-557
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* ledger(PMAT-557): release-goal.sh sync — tickets merged since v1.30.0 carry release:v1.31.0 (Refs #557)

Pmat-Ticket: PMAT-557
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(PMAT-557): receipt, quorum evidence, estimates row (Refs #557)

Pmat-Ticket: PMAT-557
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* quorum(PMAT-557): write every citation whole (Refs #557)

Pmat-Ticket: PMAT-557
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* quorum(PMAT-557): no citation span wraps a line (Refs #557)

Pmat-Ticket: PMAT-557
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* quorum(PMAT-557): committed receipt, kind triage — 3 lanes, 5 confirmed, 3 refuted (Refs #557)

Pmat-Ticket: PMAT-557
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* quorum(PMAT-557): the pmat lane names analyze_vacuous_tests, which it ran (Refs #557)

Pmat-Ticket: PMAT-557
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* ledger(PMAT-557): the ticket states the repair scope; the digest records the merge rail's two findings (Refs #557)

The merge rail's lane 1 was right twice: the digest said the three shipped rows changed only status (the cut and sync steps also bump updated: and add labels on two of them), and the ticket asked only for the booking while the branch also repairs the records gates T and B refuse. The acceptance criteria now name that repair. Lane 3's claim that PMAT-545 and PMAT-527 were edited instead of PMAT-555 and PMAT-526 was checked by attributing every changed line to its row, and refuted.

Pmat-Ticket: PMAT-557
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* quorum(PMAT-557): re-bind the receipt after the merge rail's findings (Refs #557)

Pmat-Ticket: PMAT-557
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* quorum(PMAT-557): settle which rows changed by id, not by hunk layout (Refs #557)

A merge-rail lane twice read hunk context as the edited row and claimed PMAT-545 and PMAT-527 were changed. Parsing roadmap.yaml at origin/main and at HEAD and comparing rows by id: six rows added, none removed, six changed (PMAT-526/528/529/547/555/562), PMAT-545 and PMAT-527 identical. Recorded as the instrument in the evidence.

Pmat-Ticket: PMAT-557
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* quorum(PMAT-557): re-bind after the per-row evidence (Refs #557)

Pmat-Ticket: PMAT-557
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
@noahgift
noahgift force-pushed the PMAT-564-drift-declines-on-empty-scope branch from 172bde2 to e91639e Compare September 15, 2026 19:55
@noahgift
noahgift force-pushed the PMAT-565-lock-writer-provenance branch from b67cc56 to a0222b1 Compare September 15, 2026 20:00
Base automatically changed from PMAT-564-drift-declines-on-empty-scope to main September 15, 2026 21:23
@noahgift
noahgift force-pushed the PMAT-565-lock-writer-provenance branch from a0222b1 to b286233 Compare September 15, 2026 21:30
@noahgift

Copy link
Copy Markdown
Contributor Author

quorum-review (AD-04): NOT agreed (auto_merge: checked=true was_armed=false disarmed=false)

{
 "ticket": "PMAT-565",
 "head": "b286233e1006c25a48a4b4381ae559b4b9d3a7e6",
 "width": 3,
 "executor": "agy",
 "agreed": false,
 "auto_merge": {
  "checked": true,
  "was_armed": false,
  "disarmed": false,
  "note": "auto-merge not armed"
 },
 "lanes": [
  {
   "lane": 1,
   "verdict": "PASS",
   "findings": 4
  },
  {
   "lane": 2,
   "verdict": "NO-VERDICT",
   "findings": 0
  },
  {
   "lane": 3,
   "verdict": "PASS",
   "findings": 3
  }
 ]
}

noahgift added a commit that referenced this pull request Sep 16, 2026
* release: forjar 1.31.0 — the cut (PMAT-574, refs #574)

Version, lock, the CHANGELOG heading over the three-PR window, and the crux
comparison each behaviour owes.

The window is paiml/infra#605 answered end to end: drift's verdict reaches the
exit code (PMAT-562, #563), a run that inspected none of the resources it was
asked about declines instead of reporting zero drift (PMAT-564, #569), and a
service is converged only while the loaded unit executes the declared program
(PMAT-560, #568).

GATE H PASS 3 of 3 behaviour bullet(s) under [1.31.0] reconciled in
docs/audits/crux-1.31.0.md, each naming >= 3 of the 28 surveyed systems. The
rows carry no backticks inside the key span, because the gate greps the bullet's
first six words LITERALLY and `forjar drift` does not contain "forjar drift".

The bookkeeping gate B measures is part of the cut, not around it. Arm 7 was RED
on this branch: CB-2112 37 against 35, CB-2114 36 against 34, CB-2115 49 against
43. Repaired rather than raised, by the baseline's own convention (a row whose id
tail is the issue number, a milestone, and a bare `release:`):

  PMAT-557 #557, PMAT-560 #560, PMAT-564 #564  ->  status completed; their
                                                   issues are closed, so the
                                                   open rows read as
                                                   ORPHAN-ROADMAP and
                                                   ISSUE-CLOSED both
  PMAT-560, PMAT-574                           ->  release: 1.31.0
  #565 #572 #573                               ->  rows minted under their
                                                   issue numbers, release:
                                                   1.32.0, milestone 1.32.0
  PMAT-559                                     ->  its row's title was
                                                   truncated when it was
                                                   minted, which is what the
                                                   DRIFT finding was

Measured after, on this tree: CB-2112 34 (ceiling 35), CB-2114 34 (34),
CB-2115 43 (43). Nothing is lowered here — a ceiling may only be lowered from a
measurement of the COMMITTED tree.

PMAT-574 stays inprogress: a ticket is open in its own PR and the booking PR
marks it completed, which is the rule the commit-msg hook enforced when this
commit first said completed.

PMAT-565 is NOT in this release: #570 was still open when the cut was made, so
its behaviour owes no crux row and its ticket carries release: 1.32.0.

Pmat-Ticket: PMAT-574
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* release(1.31.0): the dogfood receipt, the cut log, and the README's version lines (Refs PMAT-574)

`make dogfood-release` exited 0 on the committed tree at 9e3a474 with all nine
gates green:

GATE A PASS 5 of 5 merged PR(s) since v1.30.0 carry a harness receipt
GATE B PASS ... CB-2112=34/35 CB-2114=34/34 CB-2115=42/43 held
GATE C PASS 211 CLI name(s), 12 MCP tool(s), 12 HTTP verb(s)
GATE D PASS 18 documented invocation(s); version claims reconcile with Cargo.toml
GATE E PASS 5 of 5 merged PR(s) since v1.30.0 carry a quorum receipt
GATE F PASS line coverage 96.43% >= 95%; nothing to mutate (no .rs differs)
GATE G PASS 43 contract(s) validate
GATE H PASS 3 of 3 behaviour bullet(s) reconciled in docs/audits/crux-1.31.0.md
GATE T PASS ... cut in flight: Cargo.toml is at 1.31.0

The three reds cleared BEFORE that run are in docs/audits/logs/PMAT-574-cut.log
with the measurement either side, because a gate that was never red proves
nothing about itself.

The README's two version lines move 1.30 -> 1.31. Gate D passed either way —
Cargo reads `forjar = "1.30"` as `>=1.30.0, <2.0.0`, which admits what this cut
ships — so this is the cut keeping the documented version equal to the shipped
one, not a gate forcing it. The first draft of the receipt claimed the lines
were already correct; they were not, and the claim is replaced by what the file
says.

Pmat-Ticket: PMAT-574
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(PMAT-574): the receipt's diff list names itself; the round's evidence (Refs PMAT-574)

Lane 1 (gemini-3.1-pro-high, cited at impl-PMAT-574-receipt.md:1) read the
receipt's `## The diff` section and found it did not name the receipt doing the
listing, nor the quorum artifacts. True, and corrected.

The lane's other half — that the receipt is an unrequested file — is refuted by
scripts/dogfood/harness.sh: gate A fails any merged PR whose ticket has no
docs/audits/impl-<ticket>-receipt.md at HEAD ending in
IMPL-<ticket>-RECEIPT-END. Both halves are recorded in
.quorum/evidence/release-1.31.0-lanes.md rather than the convenient one.

The other two lanes returned NO-VERDICT (one wrote no verdict object, one hit
`UNAVAILABLE (code 503): No capacity available for model gpt-oss-120b-medium`),
and the evidence says so. A NO-VERDICT is not a PASS.

Pmat-Ticket: PMAT-574
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* quorum(PMAT-574): committed receipt — 1 round of 3 lanes, 6 confirmed, 4 refuted (Refs PMAT-574)

The round did not agree and the receipt says so: lane 1 FAIL with one cited
finding, lanes 2 and 3 NO-VERDICT (no verdict object; 503 with no capacity for
gpt-oss-120b-medium). lane_errors=2.

Three of the four refutations came from instruments that ruled before any lane
saw the branch — gate H on the crux keys, the commit-msg hook on a completed
ticket, and README.md itself on the receipt's claim that its version lines were
already right. judges_note says so rather than letting judges=3 read as three
independent tiers.

Pmat-Ticket: PMAT-574
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(PMAT-574): the after-line sums, and the 42-vs-43 gap gets its real cause (Refs PMAT-574)

Round 3's lane 1 (gemini-3.1-pro-high) refuted two claims this branch had
written:

  1. the cut log's after-composition read `CB-2115 43 (ORPHAN-ROADMAP 24,
     ORPHAN-GITHUB 8, DRIFT 10)`, and 24 + 8 + 10 = 42. A total that
     contradicts its own breakdown is the black box this format exists to
     refuse.
  2. both the cut log and the dogfood receipt explained the 42-vs-43 difference
     as "the receipts this commit had not yet added". A markdown file cannot
     move ORPHAN-ROADMAP, ORPHAN-GITHUB or DRIFT.

Re-measured at 06:40Z: CB-2115 42, composition 24 + 8 + 10, which sums. `pmat
comply` reads GitHub LIVE and stamps every run with `snapshot: gh paiml/forjar
taken <ts>`, so 43 at 05:00Z and 42 since are two measurements of a moving
source. Both documents say that now.

The same lane's four other findings are off-by-one citation claims and are
refuted in .quorum/evidence/release-1.31.0-judges.md by re-reading each cited
file — CHANGELOG.md:12 IS the first bullet, crux-1.31.0.md:31 IS the first row,
README.md:96 IS `forjar = "1.31"`, and CB-2114's repair is two rows rather than
four because three were minted already carrying `release: 1.32.0`.

Receipt now records three rounds, 6 confirmed and 6 refuted, and names
gemini-3.8-flash-* as the lane that returned a SUCCESS envelope with no verdict
object in all three.

Pmat-Ticket: PMAT-574
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* quorum(PMAT-574): round 4 — gate R's six PRs, the CHANGELOG's citation convention (Refs PMAT-574)

Round 4: lane 1 FAIL, lane 2 PASS, lane 3 FAIL. Two new claims, both refuted by
measurement, both now in the digest:

  gate R reports 6 PR(s) since v1.30.0 while gates A, E and T report 5. Both are
  right about different windows: gh lists five merged after the tag's timestamp
  (#571 #569 #568 #563 #548), and gate R's six are those plus #556 — the release
  PR whose squash commit IS the tag (`git rev-list -n1 v1.30.0` and #556's merge
  commit are both ddd0c44). The receipt says so where it quotes gate R.

  the CHANGELOG should cite #569 (the PR) rather than #564 for PMAT-564. The
  file's convention is the ISSUE, unbroken through the 1.30.0 section:
  (PMAT-549, #549), (PMAT-534, #534), (PMAT-540, #540), (PMAT-542, #542),
  (PMAT-535, #535).

Two lanes now share the same line-number error — the first bullet at
CHANGELOG.md:13 and the version line at README.md:97. Re-measured at HEAD by
numbering each blob out of git: CHANGELOG.md:10 heading, :11 blank, :12 bullet;
README.md:95 comment, :96 version. Two lanes agreeing does not move a
measurement.

6 confirmed, 8 refuted, 4 rounds, none agreed — recorded as it happened.

Pmat-Ticket: PMAT-574
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* quorum(PMAT-574): round 5 — the evidence prose now matches its own tables (Refs PMAT-574)

Round 5's lane 1 was right, three times in one finding: lanes.md still opened
"Three rounds were run" and called 203d8a6 the final head, agy.md still said
"One round", and claims.md still said "The round reviewed head 203d8a6" — while
the table underneath and the receipt said four. Prose contradicting its own
table is exactly the shape this format exists to refuse.

All three now describe every round on the head it ran against (203d8a6,
e65a1fa twice, 1da75f2, b0ccb46) and say why the count moves: a round that
raises a real finding produces a fix, the fix moves the head, and the next round
reviews a head no earlier round saw. The rail's final round reviews the final
head and cannot, by construction, be described in a file it is reviewing.

6 confirmed, 9 refuted, 5 rounds, none agreed.

Pmat-Ticket: PMAT-574
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* quorum(PMAT-574): rounds 6-8, and the receipt's own counts (Refs PMAT-574)

Rounds 6 and 7 raised NO finding: every lane that returned a verdict passed, and
each round collapsed only because one lane hit `UNAVAILABLE (code 503): No
capacity available` for its model. That is the state of the agy backend this
morning, not a property of this branch, and the evidence says so.

Round 8's lane 1 raised four findings and all four were this receipt's own stale
counts: `recorded_at` said three author claims where four had been refuted by
lanes, `judges_note` said three rounds and four refutations where the digest
said five and nine, and two "three rounds" phrases had survived incremental
editing. Every one is corrected. That is the second time a lane has caught the
bookkeeping OF the bookkeeping, which is the argument for running it.

8 rounds, 6 confirmed, 10 refuted, none agreed.

Pmat-Ticket: PMAT-574
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* quorum(PMAT-574): the round count lives in ONE place now (Refs PMAT-574)

Rounds 8 and 9 refuted the same thing, and round 9 did it from two lanes at
once: the round count and the refutation count were restated in four files, and
had gone stale in three of them. agy.md still said five rounds, claims.md still
listed heads up to b0ccb46, and judges.md's opening said nine refutations over
ten items.

Fixed structurally rather than by another pass of hand-editing four numbers:

  - the per-round table in .quorum/evidence/release-1.31.0-lanes.md is the ONLY
    place rounds are counted and heads are named
  - agy.md and claims.md refer to that table and restate no number
  - judges.md states the adjudication count once, where the gate reads it
  - the receipt keeps rounds / claims_confirmed / claims_refuted, which the
    gate validates against the digest's own items

The lanes table also says out loud what it cannot cover: the merge rail runs its
round on the FINAL head, and that round cannot be described in a file it is
reviewing.

6 confirmed, 11 refuted.

Pmat-Ticket: PMAT-574
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* quorum(PMAT-574): the stale strings a silent no-op edit left behind (Refs PMAT-574)

Round 11 refuted, from all three lanes, the claim that the structural fix had
landed:

  - .quorum/evidence/release-1.31.0-lanes.md still opened "Eight rounds were
    run" and named b0ccb46 as the last head. The edit meant to replace that
    paragraph matched nothing and reported success, which is the failure mode
    worth naming: a replacement that silently applies to zero bytes.
  - agy_teamwork.mode still read "one round of three sandboxed agy quorum
    lanes" beside "rounds": 9.
  - docs/audits/impl-PMAT-574-receipt.md still said "the round that produced
    it".

All three fixed, and this time the tree was SWEPT — every count beside the word
"round" in the evidence, the receipt and both audit documents — rather than
assumed. The sweep also separated the live claims from the QUOTATIONS: the
digest quotes stale strings so its findings can be checked, and the lanes table
now says so in its own words.

Round 10's three findings are recorded as false: it claimed the evidence
sha256s and byte counts were stale, every one matches byte for byte at HEAD,
and the total it proposed does not equal the sum it listed.

6 confirmed, 12 refuted.

Pmat-Ticket: PMAT-574
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
@noahgift
noahgift force-pushed the PMAT-565-lock-writer-provenance branch from b286233 to a4e5014 Compare September 16, 2026 08:59
noahgift and others added 16 commits September 16, 2026 13:03
…on every write (Refs #565)

StateLock gains created_by (serde default, absent on every fleet lock today); 106+31 literal constructions patched with created_by: None. Five cases: a write stamps the real writer and keeps the creator; the creator is set once and the writer every time; a legacy lock migrates its stale generator into created_by; apply rewrites the writer; lock --restamp converges every lock in one run (dry-run writes nothing, idempotent, sidecars rewritten). 5 of 5 RED.

Pmat-Ticket: PMAT-565
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…eator, and lock --restamp converges a fleet in one run (Refs #565)

save_lock stamps generator with this binary (writer_stamp, what forjar --version prints) and moves the value it replaces into created_by the first time — the one writer every path goes through, so no caller can forget. forjar lock --restamp walks --state-dir and writes every lock whose generator is not this binary through the same writer, sidecar included; --dry-run lists, --json reports, a second run is a no-op. Three LockArgs literals gain restamp: false.

Pmat-Ticket: PMAT-565
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Pmat-Ticket: PMAT-565
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… restamp's scope is stated (Refs #565)

Two of three review lanes read src/cli/lock_repair.rs and lock_audit.rs: the minimal repaired lock, the normalised form and the migrated schema were written with a bare fs::write — unstamped, and with the .b3 sidecar left stale for the next apply to refuse. All three go through save_lock now, with a falsifier through the binary. lock-restore and lock-tag copy bytes and are named as the exception. The CHANGELOG, book and contract no longer say 'every lock under a state dir': restamp walks <machine>/state.lock.yaml one level, once per workspace dir, and leaves the global lock (restamped by every apply) and .yaml.age alone — stated, not implied.

Pmat-Ticket: PMAT-565
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…rning a verb it does not (Refs #565)

Pmat-Ticket: PMAT-565
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…f 6, measured (Refs #565)

Pmat-Ticket: PMAT-565
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…565)

Pmat-Ticket: PMAT-565
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… (Refs #565)

Pmat-Ticket: PMAT-565
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…(Refs #565)

Pmat-Ticket: PMAT-565
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ed PMAT-564 (Refs #565)

Pmat-Ticket: PMAT-565
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…(Refs #565)

The branch's own diff was replayed with `git rebase --onto origin/main
e91639e`; its patch-id is unchanged (4313f3d4), so base_commit moves to
d324c57 and diff_sha256 to the hash the gate prints for the new base.

Pmat-Ticket: PMAT-565
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…efs #565)

file-health refused this branch:

  ##[error]exceeds the 500-line limit (500 -> 501). Split it:
  pmat split tests/falsification_planner_proof_reversibility.rs

The file held four falsifications (FJ-1385 proof obligations, FJ-1382
reversibility, FJ-1379 --why, FJ-004 hash_desired_state) and sat at EXACTLY the
limit, so PMAT-565's one added struct field — `created_by: None` in the
`make_lock` helper, owed by every literal of a struct that gained a field —
pushed it over.

Split on the boundary the file's own section comments already draw: FJ-1385 and
FJ-1382 stay (297 lines), FJ-1379 and FJ-004 move to
tests/falsification_planner_proof_why_hash.rs (221 lines) with `make_lock`,
which only the moved tests use. Each file's imports are now what it actually
needs. No test is added, removed or changed: 10 pass in the new binary and the
old one keeps the rest.

Pmat-Ticket: PMAT-565
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…Unreleased (Refs #565)

The 1.31.0 cut minted a PMAT-565 roadmap row (the issue was open with none, an
ORPHAN-GITHUB against gate B's CB-2115) while this branch carries its own. After
the rebase the file held BOTH. Kept the branch's row — it has the acceptance
criteria — and moved it to `release: 1.32.0`, which is where this ships now that
it missed the cut; dropped the minted stub.

The union resolver put this branch's CHANGELOG paragraph under `## [1.31.0]`,
where it would have claimed a released behaviour that is not in the release.
Moved under `## [Unreleased]`, which is what gate H will read for 1.32.0.

Pmat-Ticket: PMAT-565
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
#565)

The branch was replayed onto 8285616 (the v1.31.0 cut). base_commit and
diff_sha256 move to the new base; the review this receipt records is unchanged,
and the merge rail runs its own round on the new head before anything merges.

Pmat-Ticket: PMAT-565
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The rebase onto 583a58e resolved five both-edited roadmap rows in main's
favour, which is right for four of them — PMAT-574 completed, and the three
release labels the booking moved to v1.32.0. For PMAT-565 it was not: main's
copy is the STUB the 1.31.0 cut minted to clear an ORPHAN-GITHUB finding, with
`acceptance_criteria: []`, and it replaced this branch's own row.

Restored the branch's row (its acceptance criteria and priority), keeping
`release: 1.32.0` from the cut, because that is the release this ships in. The
status stays `inprogress`: a ticket is open in its own PR.

Pmat-Ticket: PMAT-565
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Replayed onto 583a58e. base_commit and diff_sha256 move; the review this
receipt records is unchanged, and the merge rail runs its own round on the new
head.

Pmat-Ticket: PMAT-565
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@noahgift

Copy link
Copy Markdown
Contributor Author

Two data points from paiml/infra's state dir, measured 2026-09-16 while auditing the 1.31.0 fleet pin. I filed #580 for this before finding #565 and have closed it as a duplicate; posting the evidence here so it lands on the ticket that is actually fixing it.

A fifth lock, and a worse value than the four in your description. All five were last written by a 1.30.0 binary, confirmed against forjar history --json, which stamps a correct forjar_version per apply_started event:

lock generator: says generated_at: binary that ran it (event log) run id
intel forjar 0.1.0 2026-09-16T02:11:15Z 1.30.0 r-ab35e7e5f928
gx10 forjar 1.13.1 2026-09-16T02:12:41Z 1.30.0 r-ab49ff0c7517
yoga forjar 1.1.1 2026-09-15T09:19:01Z 1.30.0 r-73f8cc5cbc7d
lambda-labs forjar 1.4.0 2026-09-14T16:47:03Z 1.30.0 r-3da7029e1237
mini forjar 1.6.1 2026-09-14T17:16:17Z 1.30.0 r-3f0d731dbfb7

intel reading 0.1.0 is the one I would point --restamp's test at: it is not merely an old release, it is a version that predates most of the file format, and it survived every apply since.

The check that should have caught this cannot. src/cli/lock_audit.rs:98 is the only place that looks at the field:

if !lock.generator.starts_with("forjar") {

forjar 0.1.0 passes that. So the audit for generator is structurally incapable of failing on a stale generator — a green with no red behind it. Worth a case in the falsification suite asserting the audit does go red for a lock whose stamp is not the running binary, otherwise the restamp can regress silently later.

Unrelated to the stamp but found in the same pass, in case it is in scope for the same file: forjar history --json already carries the correct per-event forjar_version, so the information needed to restamp retroactively exists in the event log even for locks whose generator was never right.

noahgift added a commit that referenced this pull request Sep 16, 2026
#570

I filed #580 for the lock generator never being refreshed
without checking for an existing ticket. PMAT-565 says the same thing
more precisely and #570 already implements a larger fix: it stamps the
writer on every write, preserves the original creator in a new
created_by, routes lock-repair and lock-migrate through the writer, and
adds lock --restamp. It is booked for 1.32.0.

#580 is closed as a duplicate and the measurement that was in it — a
fifth lock reading forjar 0.1.0, and lock_audit.rs:98 checking only
starts_with("forjar") so it cannot fail on a stale stamp — is now a
comment on #570 where it can be used.

Three rows remain: PMAT-579, PMAT-581, PMAT-582.

Refs PMAT-579, PMAT-581, PMAT-582

Pmat-Ticket: PMAT-579
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@noahgift

Copy link
Copy Markdown
Contributor Author

quorum-review (AD-04): NOT agreed (auto_merge: checked=true was_armed=false disarmed=false)

{
 "ticket": "PMAT-565",
 "head": "bf4c9fcda97897ccd8e397e23bf613d066039f2f",
 "width": 3,
 "executor": "agy",
 "agreed": false,
 "auto_merge": {
  "checked": true,
  "was_armed": false,
  "disarmed": false,
  "note": "auto-merge not armed"
 },
 "lanes": [
  {
   "lane": 1,
   "verdict": "FAIL",
   "findings": 3
  },
  {
   "lane": 2,
   "verdict": "PASS",
   "findings": 0
  },
  {
   "lane": 3,
   "verdict": "PASS",
   "findings": 5
  }
 ]
}

noahgift and others added 2 commits September 16, 2026 14:31
…g, merge (Refs #565)

The merge-rail quorum read this branch's own claim — "Every path that writes a
StateLock goes through save_lock" — and refuted it with three citations:

  src/cli/destroy.rs:93          cleanup_succeeded_entries rewrote a pruned lock
                                 with serde_yaml_ng::to_string + fs::write, then
                                 re-sealed by hand
  src/cli/lock_lifecycle.rs:185  lock-defrag MIRRORED save_lock (a helper that
                                 wrote atomically and refreshed the sidecar)
                                 instead of calling it — and the mirror never
                                 grew the writer stamp
  src/cli/lock_merge.rs:61,70,138 lock-merge wrote its output with a bare
                                 fs::write and NO .b3 sidecar at all, so the
                                 next apply over a merged state dir failed its
                                 integrity check

All three go through `state::save_lock` now. The dead mirror
(`write_lock_and_sidecar`) is gone: a mirror of the writer is a second thing to
keep in step, and this is the second time one drifted.

Three new cases, and the two that drive the binary were RED on the unfixed
source in a scratch clone with the fake first-writer string surviving:

  lock_defrag_writes_through_the_writer
  lock_merge_writes_through_the_writer_sidecar_included   (also asserts the .b3)
  cleanup_succeeded_entries_writes_through_the_writer     (in-crate: the fn is
                                                           pub(crate))

The contract and the receipt no longer assert more than the cases cover: both
now name all five verbs that used a bare fs::write and say which quorum round
found which. GATE G PASS 44 contracts.

Pmat-Ticket: PMAT-565
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…-write claim (Refs #565)

Round 2, on the rail at head bf4c9fc: lane 1 FAIL with three citations, lanes 2
and 3 PASS. The finding was true and is fixed in 0a4192c — destroy, lock-defrag
and lock-merge each wrote a StateLock by hand while the contract said every write
went through save_lock.

claims_refuted 4 -> 5. The fifth is the branch's own clause, which is the most
useful kind of refutation this rail produces: the suite was green, the gates were
green, and the claim was still false.

Pmat-Ticket: PMAT-565
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@noahgift

Copy link
Copy Markdown
Contributor Author

quorum-review (AD-04): three PASS — agreed (auto_merge: checked=true was_armed=false disarmed=false)

{
 "ticket": "PMAT-565",
 "head": "137cb17e7440893c4c39482fbd9344945cb6e395",
 "width": 3,
 "executor": "agy",
 "agreed": true,
 "auto_merge": {
  "checked": true,
  "was_armed": false,
  "disarmed": false,
  "note": "auto-merge not armed"
 },
 "lanes": [
  {
   "lane": 1,
   "verdict": "PASS",
   "findings": 3
  },
  {
   "lane": 2,
   "verdict": "PASS",
   "findings": 0
  },
  {
   "lane": 3,
   "verdict": "PASS",
   "findings": 0
  }
 ]
}

@noahgift
noahgift enabled auto-merge (squash) September 16, 2026 12:41
@noahgift

Copy link
Copy Markdown
Contributor Author

The quorum receipt check is red on this PR and I am letting auto-merge proceed. Here is the reasoning, so it is a judgment on the record rather than a silent bypass.

The check is not in the required gate's needs list, so GitHub does not block on it; that is a fact about the config, not a justification, so here is the measurement.

The receipt binds the correct diff. Verified three independent ways at head 137cb17e:

where hash
.quorum/PMAT-565-lock-writer-provenance.json fe4fd4119e86b66c…
PRINT_HASH=1 scripts/quorum-gate.sh in the worktree fe4fd4119e86b66c…
the same gate in a FRESH clone of origin (earlier head) matched its receipt exactly
CI eec84777c5e4ba42…

CI's hash matches no diff I can construct. Ruled out by measurement, not by argument: 15 candidate merge-bases (every main commit back through cb94fc2), 4 candidate heads including every intermediate commit of this branch and the PR merge commit, the trailing-newline capture difference, four diff algorithms, -U1/-U5/--no-prefix/--ignore-space-change, and four pathspec variants including dropping the .quorum exclusion entirely. None produces CI's value.

This is forjar#573, now at eight occurrences across four PRs. On #563, #569 and #575 a rerun cleared it; on this PR three reruns did not, so "rerun it" is no longer a workaround. The gate prints nothing that would name its own cause — which is clause 1 of #573 and should ship before it costs another release.

What IS measured green on this branch: the quorum gate locally (5 confirmed, 5 refuted, redaction clean), GATE G PASS 44 contracts, the ticket's falsification suite at 8 of 8 with two cases proven RED on the unfixed source in a scratch clone, and a 3/3 merge-rail quorum on this exact head.

clippy (self-hosted, clean-room, X64) is forjar#567 (cargo-clippy is not installed for toolchain 1.89.0) and mutation is forjar#566 (failed to compile cargo-mutants v27.1.0). Both are fleet toolchain provisioning, neither is this branch, neither is required.

@noahgift
noahgift merged commit 0edfe2d into main Sep 16, 2026
27 of 35 checks passed
@noahgift
noahgift deleted the PMAT-565-lock-writer-provenance branch September 16, 2026 13:59
noahgift added a commit that referenced this pull request Sep 16, 2026
#570

I filed #580 for the lock generator never being refreshed
without checking for an existing ticket. PMAT-565 says the same thing
more precisely and #570 already implements a larger fix: it stamps the
writer on every write, preserves the original creator in a new
created_by, routes lock-repair and lock-migrate through the writer, and
adds lock --restamp. It is booked for 1.32.0.

#580 is closed as a duplicate and the measurement that was in it — a
fifth lock reading forjar 0.1.0, and lock_audit.rs:98 checking only
starts_with("forjar") so it cannot fail on a stale stamp — is now a
comment on #570 where it can be used.

Three rows remain: PMAT-579, PMAT-581, PMAT-582.

Refs PMAT-579, PMAT-581, PMAT-582

Pmat-Ticket: PMAT-579
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
noahgift added a commit that referenced this pull request Sep 16, 2026
#583)

* docs(roadmap): four rows from auditing a fleet pin with forjar 1.31.0

Found while converging paiml/infra from forjar 1.30.0 to 1.31.0. Rows
only, no implementation.

PMAT-579 apply and plan do not check that the running binary is the
version the manifest pins. The same pin was applied by 1.30.0 on one host
and 1.31.0 on the next, minutes apart, silently.

PMAT-580 finalize_machine refreshes generated_at but never assigns
lock.generator, so every per-machine receipt names whichever version
first created the file — intel says forjar 0.1.0 for an apply 1.30.0
performed. lock_audit only asserts the string starts with forjar, so it
cannot fail on the thing that is wrong with it.

PMAT-581 is the capability half of forjar#497: a remote resource has no
content key, so an SSH target plans NoOp over changed sources. infra
carries 35 hand-rolled activation checks standing in for the primitive.

PMAT-582 drift declines only at 0 of N, so a missing state dir inspected
40 of 153 and still returned a drift verdict. One host at one commit read
2, 3 and 9 across three sessions, every run exiting 1.

Refs PMAT-579, PMAT-580, PMAT-581, PMAT-582

Pmat-Ticket: PMAT-579
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(roadmap): drop PMAT-580 — it duplicates PMAT-565, already fixed in #570

I filed #580 for the lock generator never being refreshed
without checking for an existing ticket. PMAT-565 says the same thing
more precisely and #570 already implements a larger fix: it stamps the
writer on every write, preserves the original creator in a new
created_by, routes lock-repair and lock-migrate through the writer, and
adds lock --restamp. It is booked for 1.32.0.

#580 is closed as a duplicate and the measurement that was in it — a
fifth lock reading forjar 0.1.0, and lock_audit.rs:98 checking only
starts_with("forjar") so it cannot fail on a stale stamp — is now a
comment on #570 where it can be used.

Three rows remain: PMAT-579, PMAT-581, PMAT-582.

Refs PMAT-579, PMAT-581, PMAT-582

Pmat-Ticket: PMAT-579
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(roadmap): each row's criterion describes THIS diff, not the work it registers

infra#633, the sibling of this PR, was failed 3/3 by its quorum lanes
for exactly this: a registration-only diff carrying a criterion that
described the implementation it does not contain. The lanes were right,
and these three rows had the same shape.

Each criterion now states what this diff does — it registers the issue
and carries the measurement that justifies it — and the implementation
moved to notes as future scope rather than an unmet promise.

No measurement changed; only which of them the criterion claims.

Refs PMAT-579, PMAT-581, PMAT-582

Pmat-Ticket: PMAT-579
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(audits): impl receipts for PMAT-579, PMAT-581, PMAT-582

Written BEFORE the quorum round, not after a failure. The sibling infra
PR was failed by a lane that read the immutable title as the diff's
contract, and quorum-review reads one receipt per ticket for exactly the
multi-ticket case: on #37 a round went 3/3 FAIL with every finding naming
work another ticket asked for.

Each receipt states the measurement its row carries, what would make the
row false, and why the title reads like an order.

Refs PMAT-579, PMAT-581, PMAT-582

Pmat-Ticket: PMAT-579
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(audits): separate the diff's triage rail from each row's kind:code label

Quorum round 1 on this PR: lane 1 FAIL, lanes 2 and 3 PASS. Lane 1 was
right that the receipts read as a contradiction: they said the rows were
"all on the kind: triage rail" while each row carries kind:code.

Those are two different kinds. The DIFF is triage in the quorum gate's
path-rail sense; each ROW's label classifies the work it registers, which
is code. Measured on this tree: 115 rows carry kind:code and 7 kind:triage,
every planned row registering future code work is kind:code, and
kind:triage is reserved for classify-and-link work. So the label stays and
the receipts now name the two axes separately. Lane 1's first suggested
fix, relabelling the rows, would have made the label false.

Refs PMAT-579, PMAT-581, PMAT-582

Pmat-Ticket: PMAT-579
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(quorum): evidence for the PMAT-579/581/582 registration — claims, lanes, judges, agy, pmat, crux

Round 3 of three sandboxed agy lanes AGREED 3/3 on 5a6d607. The receipts'
label count is corrected: 115 kind:code was the merge base, this diff makes
it 118. Refs PMAT-579

Pmat-Ticket: PMAT-579
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* docs(quorum): evidence names the rebase onto 0edfe2d (forjar#570)

Refs PMAT-579

Pmat-Ticket: PMAT-579
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* docs(quorum): receipt for the PMAT-579/581/582 registration (kind: triage), rebased onto 0edfe2d

Refs PMAT-579

Pmat-Ticket: PMAT-579
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
noahgift added a commit that referenced this pull request Sep 20, 2026
…the quorum binding it fixed (PMAT-592, closes #592 #594 #573) (#593)

Cuts v1.32.0, 50 hours past its cadence due date. One behaviour ships: the
per-machine lock names the binary that wrote it on every write, keeps its
creator, and `lock --restamp` converges a state dir (PMAT-565, #570) — the third
signature of paiml/infra#605.

Three blockers were found and cleared on the way, none of them planned:

GATE B WAS RED ON MAIN (#594), not on this branch. Measured in a detached
worktree before touching anything: origin/main CB-2115 55/43 and CB-2114 37/34,
against the cut branch at 52 and 37. `make dogfood-release` could not exit 0 for
anyone. Repaired rather than raised — no ceiling edited — and 12 of the 14 excess
findings turned out to be DRIFT title mismatches, the cheapest thing in the file.
Final: CB-2112 33/35, CB-2114 33/34, CB-2115 38/43.

GATE A REFUSED THE #583 RECEIPTS. All three were missing both halves of the
contract — zero `verdict:` lines and no END marker. The gate names one file at a
time, so fixing only the one it printed would have cost two more red runs.

THE QUORUM BINDING WAS NOT PORTABLE BETWEEN CLONES (#573). CI refused this
branch's receipt as STALE on a head the local gate had just passed. Root cause:
`quorum-gate.sh` hashed `git diff` without `--full-index`, so git abbreviated the
index lines to a length chosen from the OBJECT COUNT of whichever clone ran it —
this workstation (42982 objects) to 8, the runner to 9, and `core.abbrev=9`
locally reproduced the runner's hash exactly. Filed as a flake ("passes on
rerun"); it is per-clone deterministic, and a rerun only ever appeared to fix it
by landing on a runner that happened to agree with the author. Fixed with
`tests/falsification_quorum_hash_is_abbrev_independent.rs`, which carries an
anti-vacuity arm requiring that WITHOUT the flag the two settings disagree.

WHAT THE QUORUM CAUGHT, and it is not what a quorum is usually for. Three lanes
(2 agy + 1 claude, read-only over the recorded gate result, no build and no
token). All eight structural claims about the diff held. Every kill was PROSE:
"the fourth signature" where the CHANGELOG says third; a 50-hour-late cut called
"two days after 1.31.0" against its own dates; 55 and 53 quoted as one
measurement from two unnamed instruments; and one reason given for holding back
both #590 and #591 when #591's cause is named to the line in this very diff.
Nine gates had passed and the release still contained four checkable falsehoods
in its own audit trail.

  make dogfood-release: nine gates, exit 0, zero error lines
  quorum gate passed (blocking): lanes=3 refuters=3 judges=3 confirmed=3 refuted=10

NOT in this tag, and named so the next reader does not have to infer it: #590
(apply --refresh ignores completion_check for a templated resource) has a
reproduction and no root cause, and #591 (nothing compares live against the
declaration) is two halves neither of which is safe to ship alone. Both target
v1.33.0, so 1.32.1 is owed for the arbiter P0 chain.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

1 participant