Conversation
…vs sm_121 GPU) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…the first over-budget one Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… sm_121 GPU GPU sm_89 vs sm_121 pos0 logits agree to 6.5e-7; x86 AVX2 vs aarch64 NEON CPU references differ by 2.6e-2. aarch64 budget = measured worst 7.105e-2 rounded up (8e-2); x86_64 stays 7e-2. QWEN35_E2E_DUMP now fails after printing, never passes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…y gx10) + findings yoga-eph has no models, so the module SKIPped at PR time and the whole-forward budgets were only ever read on sm_89/x86_64. The gx10 nightly holds the model; the step refuses a skip or an empty selection. Verified: 17/17 on GB10. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
noahgift
enabled auto-merge
September 27, 2026 10:31
|
§13.11 rung 1 — quorum shadow verdict Shadow mode: this records a verdict and merges nothing. A refusal |
Contributor
Author
|
Moved into #4607 (PRCAP fold, cop order 14:02Z). Fold = MOVE: head |
auto-merge was automatically disabled
September 28, 2026 14:24
Pull request was closed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #4523
What
The qwen35 e2e CUDA-vs-CPU parity test measured 7.105e-2 relative L-inf on gx10 (sm_121/aarch64), against a 7e-2 budget. This is not an sm_121 defect:
The GPUs agree with each other. The difference comes from the CPU reference (Q8_K activation quantization, SIMD path per arch). So the budget belongs to the host CPU arch. It is derived from measurement (the worst reading rounded up to one significant figure), not widened by hand:
7e-2(unchanged)8e-2The falsifier: the old 7e-2 budget fails 5/5 on gx10. With the fix, gx10 runs the module 17/17 with 0 skips, and lambda passes at 7e-2.
Also
QWEN35_E2E_DUMP=<dir>is a measurement mode that writes per-position GPU/CPU logits. It always fails, so it can never read as a pass..github/workflows/cuda-nightly.ymlgets a new step in the gx10 job that runs the qwen35 CUDA parity module on sm_121. The step is RED on a non-zero rc, on any SKIP/absent line, or on 0 passed. The PR-timecuda-unitlane has no models, so the module SKIPs there. This is a workflow edit, with no runner-host or secret changes.docs/findings/4523-sm121-logits.jsonl(PREREG, M1–M3). Roadmap:PMAT-4523.docs/audits/quorum-PMAT-4523.json. AGREED 3/3 PASS (gemini-3.1-pro-high ×2 and claude-haiku-4-5; the author is claude-opus-5-5).🤖 Generated with Claude Code