Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
153 changes: 153 additions & 0 deletions .quorum/PMAT-574-release-1.31.0.json

Large diffs are not rendered by default.

32 changes: 32 additions & 0 deletions .quorum/evidence/release-1.31.0-agy.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,32 @@
# PMAT-574 — the agy rounds

Every round is three sandboxed `agy` quorum lanes, review-only, in their own
clones with push removed, `writes=false`, against `main` at 9884334a. THE
PER-ROUND TABLE IN `release-1.31.0-lanes.md` IS THE ONLY PLACE THIS RECEIPT
COUNTS ROUNDS: this file and `release-1.31.0-claims.md` deliberately restate no
number, because a count repeated in four files is a count that goes stale in
three of them — which two separate rounds caught, and which is the reason for
this rule.

Models used: `gemini-3.1-pro-high`, `gemini-3.8-flash-high`,
`gemini-3.8-flash-medium`, `gemini-3.6-flash-high`, `gemini-3.6-flash-medium`,
`gpt-oss-120b-medium`. Every round ran three DISTINCT model ids, none in the
author's family (the author is `opus`/claude). The distinctness was chosen
deliberately: the PMAT-565 round ran `gemini-3.1-pro-high` twice and the
validator marked it `partial` — "lanes sharing an id are resamples, not
independent reviewers (PMAT-125)". `gemini-3.8-flash-*` returned a SUCCESS
envelope with no verdict object in every round it ran in, and was dropped after
round 3 for that reason and no other.

Each lane reviewed in its own full clone with push removed. No lane may write to
this tree; the `writes=false` column is a property of the harness, not a promise
from the lane.

## What a lane can and cannot rule on

A lane reads a diff. It cannot run the release gates, cannot ask GitHub what
merged, and cannot measure the CB-21xx ratchet — so on a release cut, whose diff
is almost entirely a record of things measured elsewhere, the lanes are the
WEAKER half of the adjudication and this receipt does not pretend otherwise.
The refutations that mattered came from the gates and the hooks, and
`release-1.31.0-judges.md` names them.
45 changes: 45 additions & 0 deletions .quorum/evidence/release-1.31.0-claims.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,45 @@
# PMAT-574 — the claims put to the round

forjar 1.31.0 is a release cut: three behaviour PRs since v1.30.0 (plus two that
change none), no Rust changed, and the whole of the diff is the version, the
record of what shipped, and the roadmap bookkeeping the release gates demanded
before they would go green.

The claims put to the three lanes were:

1. The version is bumped everywhere it is declared — `Cargo.toml`, `Cargo.lock`
and the README's two dependency lines — and nowhere is left behind.
2. The CHANGELOG's `[1.31.0]` section describes the PRs of this window and no
others, and each behaviour it names is one the window actually shipped.
3. Every behaviour bullet has a crux row naming at least three of the 28
surveyed systems, and the audit says truthfully how it was produced.
4. The roadmap changes are bookkeeping the release gates require, they are
confined to rows and fields, and no ceiling in `scripts/ratchets/` is raised.
5. `PMAT-557`, `PMAT-560` and `PMAT-564` are marked completed because their PRs
MERGED in this window, not to make a gate green; `PMAT-565`, `PMAT-572` and
`PMAT-573` are minted under their own issue numbers carrying
`release: 1.32.0` because they are NOT in this release.
6. The dogfood receipt reports the gate lines the run actually printed, and the
cut log's REDs are reds this branch really had.
7. Nothing in the diff changes forjar's behaviour, so no falsification test is
owed for a behaviour change, and the receipt says so rather than pointing at
an unrelated test.
8. PMAT-565 (#570) is excluded from this release, and the exclusion is STATED —
in the crux document, in the impl receipt and in the ticket's own
`release: 1.32.0` — rather than left for a reader to notice.

These claims were put to every round in the table in
`release-1.31.0-lanes.md` — each on the head that existed when that round
started, because a round that raises a real finding produces a fix and the fix
moves the head. Every round ran three lanes, review-only, sandboxed,
`writes=false`, with three distinct model ids, none in the author's family. The
count of rounds is not restated here on purpose; the table is the one place it
lives.

## What the round could NOT adjudicate, and who did

A lane reads the diff. It cannot run `make dogfood-release`, cannot ask GitHub
what merged, and cannot measure the ratchet. Those refutations came from the
gates and the hooks, before the lanes saw the branch, and
`.quorum/evidence/release-1.31.0-judges.md` names them as the refuting
authority.
40 changes: 40 additions & 0 deletions .quorum/evidence/release-1.31.0-crux.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,40 @@
# PMAT-574 — the crux comparison this release owes

`docs/audits/crux-1.31.0.md`, three rows, one per behaviour bullet, each naming
at least three of the 28 systems `scripts/dogfood/crux-reconcile.sh` surveys.
Written by the release orchestrator from documentation memory rather than
surveyed by an `agy` lane as 1.29.0's was; every third-party claim is marked
`[X]` and the audit's Method section says so in its first paragraph.

## The three rows, and what each comparison found

1. **drift declines over an empty scope (PMAT-564)** — Bazel, Ansible,
Terraform, Kubernetes. Bazel treats a target pattern matching nothing as an
error `[X]` and modern Ansible refuses a `--limit` that matches no host
`[X]`; `kubectl get` over an empty selection prints `No resources found` and
exits 0 `[X]`, which is the shape forjar had. forjar's defect was worse than
the Kubernetes shape, because the empty selection was not empty: the run
graded two resources from a manifest it had never been given.

2. **a service is converged only while the loaded unit runs the declared
program (PMAT-560)** — systemd, Ansible, Puppet, Chef, Podman. systemd itself
distinguishes the file on disk from the unit it has loaded `[X]`; the
configuration managers all converge on `state`/`enabled` and manage the unit
FILE separately `[X]`, which is exactly the gap; the container world already
identifies what runs by DIGEST `[X]`, which is what `exec_sha256` is — hashed
at the LIVE path, never the declared one.

3. **drift exits 1 on any DRIFTED line (PMAT-562)** — Terraform, Puppet,
Kubernetes, Ansible. `terraform plan -detailed-exitcode` exits 2 on changes
`[X]`, Puppet's `--detailed-exitcodes` exits 2 `[X]`, `kubectl diff` exits 1
`[X]`; Ansible's `--check` exits 0 and makes you parse the recap `[X]`. The
finding: the two systems with an opt-in flag both default to "changes are not
an error" for a PLAN, whereas forjar's `drift` is a CHECK, whose only reason
to run is the question the exit code now answers.

## What the comparison cannot show

That it is correct in detail. No reference system was invoked; every `[X]` is a
documentation-memory claim. The gate agrees about itself: it asserts the
reconciliation was WRITTEN and is COMPLETE over the bullets, and catches the
cheapest failure — a behaviour shipped with no comparison at all.
Loading
Loading