Conversation
…n 1) The fixture is a Python project whose test_mean fails; the fix is a one-line edit to stats.py. task.txt asks the agent to make it and run the test. scripts/apr_code_edit_verify.sh runs one cell (model x host) under the fleet GPU lock and choom 1000, and scripts/lib/apr_code_edit_verify.py judges it from artifacts only: the working-copy diff, an independent test re-run, the --emit-trace tool calls, and the serve child's own output. `apr code` has no GPU flag and its driver hides the child's output on success, so APR_BIN points the driver at a wrapper that tees the child's stdout and stderr. The first failing mechanism is named with #3719's vocabulary. The fragment is aprender-c3's (043333a), reassigned to aprender-f8. Refs #3719 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…_tokens) Per aprender-62 (#3712) and aprender-97 (#3715): a code row keys on (file, verb, thinking, context). thinking is read from the serve child's own line, since the driver strips <think> before parsing, and is `unknown` until the child prints one. context is `4k` only if one request's measured prompt reached 4096 tokens; otherwise it is `task`, which no rung owes. session_end.tokens_in is summed over turns (agent/result.rs), so it is kept as tokens_in_total and never used as a prompt size. Refs #3719 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…nt rule and the row-key rulings Refs #3719 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
First real cell (lambda, Qwen3.5-4B, apr 0.69.0 856009cc9): the serve child printed `Model ready: 0 layers` and `gpu-layers: requested=all resolved=0 total=0 (backend=cuda)`, then `CUDA optimized model ready`, and the first completion failed with HTTP 500 "Model architecture not supported for GPU-resident path". The judge took the ready line as backend=cuda and blamed "tool call not parsed". Now: a child reporting 0 layers did not load; backend is cuda only when every layer is resident on CUDA (the gpu-layers line, which is the evidence cited) AND a completion came back; an HTTP error from a fully loaded child is named "serve child refused the forward". Refs #3719 #3571 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…de effects; commit the case table `apr code --emit-trace` writes four records and the assistant turn is one text block whatever the agent did (agent/code.rs emit_ccpa_trace), and the PreToolUse/PostToolUse hooks are not wired into the loop. The judge read "tool call parsed" and "the agent ran the test" from tool_use blocks in that trace, so it could never pass a correct run. Now: - "test not run" reads a log written by python3/python shims put first on the agent's PATH (ShellTool runs `sh -c` with the inherited environment); the harness's own re-run calls the real interpreter by path. - "tool call not parsed" needs stats.py unchanged while the final answer still carries the markup the driver parses (<tool_call>, ```json). - The Qwen35Session route's load line (#3571 step 2, aprender-c7: `Model ready: Qwen3.5 hybrid, N layers resident on the GPU|CPU, ...`) is read for layers and residency alongside the generic route's gpu-layers line. scripts/check_apr_code_edit_verify.sh is the judge's case table: 19 rows, no model, no GPU, dispatched by guard_tree --no-cargo. Two mutants were run against it and each turns it red: trusting `CUDA optimized model ready` for residency (rows zero, partial) and dropping the test-not-run check (no-test). Refs #3719 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…lue other than on/off keys onto no cell aprender-c7: after #3755 the template name is whatever the shared selection picks and the thinking value is one of off | on | the model's choice. Only on/off key a ladder row; anything else is kept as thinking_raw. Refs #3719 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
run_single_prompt replaced the manifest's system prompt, for EVERY model
size, with COMPACT_SYSTEM_PROMPT ("Answer the question. Be direct."), which
names no tool and no <tool_call> format. The tools stayed registered, but
no model was told they existed, so `apr code -p` could not edit a file.
Measured (#3719 cell gx10 / Qwen3.5-4B-Q4_K_M, apr 0.69.0 cc3892acd, with
the #3571 step (2) serve route): the serve child loaded 32 layers on CUDA,
and the model answered the edit-and-verify task with "Without seeing the
actual code, I'll assume…". num_turns 1, prompt 81 tokens, zero tool calls,
stats.py unchanged. A fake `apr serve` capturing the request body showed the
system message was the 31-character COMPACT prompt.
-p now keeps the manifest's prompt, which cmd_code has already scaled
(PMAT-198: COMPACT below 2B, the full tool table otherwise). PMAT-197's
small-model concern is that scaling, which still applies. The two comments
that said COMPACT "keeps tool format" now say it names no tools.
Falsifier: falsify_3719_single_prompt_run_tells_the_model_about_its_tools
drives run_single_prompt through a driver that records request.system.
Green here; RED with the old override restored ("a -p run sent
COMPACT_SYSTEM_PROMPT, which names no tool"). cargo test -p
aprender-orchestrate --lib agent:: 884 passed; clippy --lib --tests
-D warnings clean.
Refs #3719
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…n the #3571 route baseline-856009cc9 (origin/main 52f43da + #3726 c57c260 + harness, --features cuda): all four cells, Qwen3.5-{4B,9B}-Q4_K_M x {lambda, gx10}, FAIL "serve child did not load". The child printed `Model ready: 0 layers` and `gpu-layers: resolved=0 total=0 (backend=cuda)`, then `CUDA optimized model ready`, then answered HTTP 500. Fix: #3571. route-cc3892acd (+ #3571 step (2) 5a4a8e1, pre-receipt): gx10 / 4B loads 32 layers on CUDA and generates, but `apr code -p` sent the system prompt "Answer the question. Be direct.", so the model made no tool call. The fake serve request capture is included. Fix: b2b89d6 in this branch. Refs #3719 #3571 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…v 5) The cop's gpu-q puts a (priority, arrival) queue in front of the same /tmp/apr-gpu.lock and itself runs the job under flock + choom -n 1000, so with --gpu-q the harness hands its apr code step to gpu-q instead of taking the lock a second time (two flocks on one file deadlock). The lock-acquired marker is still written once the lock is held. Checked against a scratch GPUQ_DIR/GPUQ_LOCK: marker written, env passed, oom_score_adj 1000. Refs #3719 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…der-62's condition for the ladder) aprender-62 (#3712) will call the harness from model_ladder.sh OUTSIDE apr_locked with --gpu-q 1, so "every GPU apr call is locked" rests on the harness. Condition set: its case table must prove the apr call runs with the lock held and choom 1000, and that a held lock declines rather than hangs. - --gpu-q had no wait bound (gpu-q has none): the gated step now runs under timeout LOCK_WAIT+TIMEOUT+5, so a run that never got the lock declines. - APR_GPU_LOCK overrides the lock path so the case table never queues on the fleet lock; gpu-q gets it as GPUQ_LOCK. - Four harness rows run the real harness against a fake apr (serve child lines through the APR_BIN wrapper, lock+oom probe, the task done right) and a stub gpu-q with gpu-q's exec contract: free lock -> PASS with `held=yes oom=1000`, held lock -> DECLINE in bound, in both gate modes. Each row has its own timeout, so a hang fails the row. Mutants, each turns the table red: choom dropped from the flock gate (h-flock-free sees oom=0); the gpu-q bound dropped (h-gpuq-held rc 124); a second flock around gpu-q (h-gpuq-free and h-gpuq-held deadlock, rc 124). Refs #3719 #3712 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…able matches without a pipe into grep -q guard_tree --no-cargo on the branch: check_roadmap_fragment_required saw the new fragment without a regenerated aggregate, and check_no_pipe_into_grep_q counted 75 sites against a ceiling of 74 (the case table's expect() piped printf into grep -q, which pipefail can fail on SIGPIPE). The aggregate is regenerated (make roadmap-aggregate), and expect() uses [[ ]] matching. Refs #3719 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… 6); a gpu-q without `wait` falls back to the bounded flock aprender-62: gpu-q v3 (lambda + gx10, md5 e2c25ec8) bounds the whole wait, queue + flock, with GPUQ_WAIT, exits 75 without running the command, and lists `wait` in --caps. The harness now passes GPUQ_WAIT=LOCK_WAIT instead of wrapping gpu-q in one timeout that also bounded the run, and falls back to the bounded flock when --caps lacks `wait`, as model_ladder.sh does (#3771). Two rows are added for that fallback; the stub keeps v3's contract. Mutants, each red: GPUQ_WAIT dropped (h-gpuq-held rc 124); --caps check dropped (h-old-gpuq-held rc 124). Refs #3719 #3712 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ets is executed (#3719) Measured (#3719 cell gx10 / Qwen3.5-4B-Q4_K_M, apr 0.69.0 206ad8421, CUDA, #3571 step (2) route): the model's final turn was a file_edit call with the exact right edit, but its JSON lacked the outer `}`: <tool_call> {"name": "file_edit", "input": {"path": "stats.py", "old": "return sum(values) / (len(values) - 1)", "new": "return sum(values) / len(values)"} </tool_call> The envelope parser failed to parse it and returned it as the answer text, so stats.py was never edited ("tool call not parsed"). repair_unclosed_tool_call: when a <tool_call> is DELIMITED by its </tool_call>, the brackets still open there are closed in order. It is conservative like the CCPA-m296 salvage parser: it returns None unless the scan ends outside a string, nothing closes a bracket it did not open, at least one bracket is open, and the result has a string `name` and an explicit `input`. Undelimited calls, unterminated strings, mismatched closers, trailing commas, and a missing name or input are all still refused. Falsifier falsify_3719_delimited_tool_call_missing_outer_brace_is_executed uses the captured text verbatim. It and repair_closes_nested_brackets_in_order go RED with the repair's call site disabled; repair_refuses_every_other_ malformation pins what stays refused. agent:: 887 passed; clippy --lib --tests -D warnings clean. Refs #3719 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A v4 cell's evidence quoted 'gpu-q: waiting (prio 1, 2/5 in queue)' (the queue talking, not apr). stderr lines starting 'gpu-q: ' are dropped before any excerpt; a case-table row pins it. Refs #3719 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ept (#3719) The coding prompt's tool table and few-shot Example 2 taught `file_edit {"path", "old", "new"}`; FileEditTool requires `["path", "old_string", "new_string"]`. AprServeDriver strips the JSON schemas from the prompt, so the table is all a served model sees. Measured on Qwen3.5-4B-Q4_K_M. On CUDA (lambda, v4/v5 cells): 6 iterations, 5 tool calls, no edit, empty answer. On CPU, a logging proxy between apr code and its serve child shows why: the model read test_stats.py and stats.py, then called file_edit with the exact right edit using "old"/"new". The tool answered `missing required field 'old_string'`, and the model repeated the identical call until the loop guard ended the turn. (Qwen3.5-9B recovered from the error and PASSED.) The falsifier checks every tool example in CODE_SYSTEM_PROMPT (table rows and <tool_call> examples) against the REGISTERED tool's input_schema `required` list. It also found a second drift nobody had hit: `memory` was taught with key/value, but it requires `content`. Both are fixed. falsify_3719_prompt_tool_examples_supply_every_required_field was RED before (5 findings) and is green after. agent:: 888 passed; fmt + clippy --lib --tests -D warnings clean. Refs #3719 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…743); implementation receipt
Qwen3.5-{4B,9B}-Q4_K_M × {lambda, gx10}, CUDA, through gpu-q prio 1: 4/4
PASS. Each cell carries the serve child's residency lines (32 layers
resident on the GPU, gpu-layers resolved=32 total=32 backend=cuda, zero
fallback lines), the exact one-line edit, the agent's own unittest run and
an independent re-run OK. The binary is the baseline + #3571 step (2)
5a4a8e1 + this branch. docs/audits/impl-PMAT-3719-receipt.md maps each
done_when to its evidence and falsifiers.
Refs #3719 #3571
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…xity ratchet) guard_tree --no-cargo at b9d7812: check_complexity_ratchet saw parse_tool_calls_envelope GROWN (cognitive 32 -> 34, cyclomatic 10 -> 11) and repair_unclosed_tool_call NEW over a threshold (cognitive 28). The envelope parser now calls parse_call_json in the same `if let` shape it had before. The repair is unclosed_brackets (the scan) + string_step (one character of string state) + the validation. Behaviour is unchanged: the three repair tests and agent:: 888 pass. Refs #3719 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Refs #3719 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… union The orphaned PMAT-3719-apr-code-qwen35-cuda (@dcf9f6b85, claim dropped 2026-09-22, never landed) carried onto main 761d624 for a fold. Conflicts: - crates/aprender-orchestrate/src/agent/code_tests.rs: both sides only appended a section at the end (#3775 on main, #3719 here). Resolved as main's file plus this branch's appended #3719 block, unedited. - docs/roadmaps/roadmap.yaml: generated; regenerated with `make roadmap-aggregate` (PMAT-3719 present once). Checked on the merged tree, private target dir: aprender-orchestrate agent::code 104/104, agent::driver 154/154, clippy -D warnings clean, cargo fmt --check clean, check_apr_code_edit_verify.sh case table OK. The four-cell receipt (evidence, apr 0.69.0 7b161d743) is NOT re-measured at this merge. Refs #3719 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…79ae2 Re-measured after carrying the orphaned branch onto main: apr 0.69.3 (bb79ae2) --features cuda, RTX 4090 sm_89, scripts/apr_code_edit_verify.sh. - lambda · Qwen3.5-4B-Q4_K_M: PASS. The serve child reports cuda with 32 layers resident and no fallback; the agent ran the test (shim log); the harness re-ran it here with rc 0. - lambda · Qwen3.5-9B-Q4_K_M: PASS, same checks. The gx10 cells are not re-measured; they still stand on v6 (0.69.0 7b161d743). The receipt doc gains a v7 section saying so. The same 14 files per cell as v6 are committed; the project copy and the shims are not. Refs #3719 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Lanes claude-sonnet-5 x2 (the gemini seat moved to Claude by the agy tier-1-only policy) and claude-haiku-4-5; author claude-opus-5-5. The advisory apr lane did not run (brief 199,619 bytes > 24,576). Refs #3719 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…eports a cut as "length" (#3718) #3718 done_when 3: a prompt too long for the context says so, never a silent cut. The CPU (effective_max_tokens), CUDA and Qwen3.5 (Session) paths already refuse it. The wgpu handler was the gap: it capped max_tokens at 4096 and nothing else, so an over-length prompt prefilled anyway, the decode loop grew the KV cache past the model's window, and the stream's final chunk hardcoded finish_reason "stop". - WgpuInferenceState carries the model's context_length (GGUF config). - wgpu_chat_completion refuses prompt_len >= context_length with HTTP 400 and an OpenAI-shaped body (error.code = "context_length_exceeded", both counts), before any prefill, and clamps max_tokens to the room left. - The streaming done chunk reports "length" when the budget was spent (the blocking path already did). - context_token_budget / context_length_exceeded_body are pure and not gated on the wgpu feature, so default-feature CI runs their 4 case rows. Verified: cargo test -p apr-cli --lib context_budget_3718 (4 pass); cargo clippy -p apr-cli --lib --no-default-features --features wgpu -D warnings clean. Note: `--features wgpu` WITH default features does not compile on main (finetune.rs uses entrenar wgpu types without enabling entrenar/wgpu); that is pre-existing and not touched here. Refs #3718 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… reported a cut as "stop" (#3718) Quorum round 1 (gemini-3.1-pro-high, FAIL): handler_gpu_completion.rs hardcoded "finish_reason": "stop" on the CUDA non-streaming reply and on the CUDA->CPU fallback, so a reply cut at max_tokens still read as finished. The SafeTensors /v1/chat/completions path (chat.rs build_chat_response) did the same whenever there were no tool calls. All four serve paths now take finish_reason from ONE rule, finish_reason_for(generated, max_tokens): "length" when the budget ran out, else "stop" (tool_calls still wins on the SafeTensors path). The two inline copies in the wgpu handler are replaced by it. Tests: finish_reason_for's table, and a SafeTensors reply of 16/16 tokens reads "length". cargo test -p apr-cli --lib (filtered) 25 pass; clippy -D warnings clean on default and on --no-default-features --features wgpu; cargo check --features cuda clean. Refs #3718 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… out hashing the test binary (#4550) x86-main's mutation section went red with "cargo test failed in an unmutated tree": cargo-mutants' baseline `cargo test -p apr-cli --lib` hit its 300 s timeout with 7434/7435 tests done. The stragglers were all 17 test_llm_band tests: provenance hashes current_exe() (PP-25), which in a test is the 472 MB apr-cli lib test binary, and sha2 at opt-level 0 runs ~11 MB/s, so each test paid a ~40 s SHA-256 and one never finished (cuda_without_the_server_feature_is_refused). Deterministic, not a flake: not in any flake ledger, so fixed rather than rerun. Same hunk as 49/x86-slow-2 (88449ce) and #4554, byte-identical, so the two PRs merge in either order. Optimizing only sha2 keeps the hash real. Measured: `cargo test -p apr-cli --lib test_llm_band` 65 passed in 1.39 s (CI baseline: 17 tests >60 s, one past the 300 s kill). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…uorum), it lands once via #4554 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…_length_exceeded, finish_reason judged against the clamped budget (#3718) Quorum finding on #4550: the SafeTensors chat/completions paths judged finish_reason against the requested max_tokens while Session clamps the budget to context_length - prompt_len silently, so a context cut read as "stop"; a prompt >= the window surfaced as a 500. Both now go through context_token_budget / context_length_exceeded_body via st_context_budget. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-59 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Agent: aprender-59 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
§13.11 rung 1 — quorum shadow verdict Shadow mode: this records a verdict and merges nothing. A refusal |
Contributor
Author
|
quorum-review (AD-04): three PASS — agreed (auto_merge: checked=true was_armed=false disarmed=false) {
"ticket": "PMAT-3718,PMAT-3719",
"head": "3c07f85b5191ca3088bf439acd36c00cc003a4d2",
"width": 3,
"executor": "agy",
"agreed": true,
"auto_merge": {
"checked": true,
"was_armed": false,
"disarmed": false,
"note": "auto-merge not armed"
},
"lanes": [
{
"lane": 1,
"verdict": "PASS",
"findings": 0
},
{
"lane": 2,
"verdict": "PASS",
"findings": 4
},
{
"lane": 3,
"verdict": "PASS",
"findings": 2
}
]
} |
Contributor
Author
|
quorum-review (AD-04): three PASS — agreed (auto_merge: checked=true was_armed=false disarmed=false) {
"ticket": "PMAT-3718,PMAT-3719",
"head": "3c07f85b5191ca3088bf439acd36c00cc003a4d2",
"width": 3,
"executor": "agy",
"agreed": true,
"auto_merge": {
"checked": true,
"was_armed": false,
"disarmed": false,
"note": "auto-merge not armed"
},
"lanes": [
{
"lane": 1,
"verdict": "PASS",
"findings": 2
},
{
"lane": 2,
"verdict": "PASS",
"findings": 1
},
{
"lane": 3,
"verdict": "PASS",
"findings": 0
}
]
} |
…t-5, haiku-4-5; degraded: same-family, no author seat) Agent: aprender-59 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
#4614) R2 asked whether #4606's over-length refusal duplicates D5 (#4614, PMAT-4616). It does not: D5 maps aprender-serve api errors; these are apr-cli's own SafeTensors handlers. Driven through the real safetensors_app router, a 200-token prompt against a 64-token window: - main + #4614 (bd46408): /generate and /v1/chat/completions -> 500 "Generation failed: ... refused whole rather than truncated" (RED) - #4606 (9a92e2f): both -> 400 error.code context_length_exceeded (GREEN) Run on intel, cargo test -p apr-cli --lib tests_st_overlength_router_3718. Refs PMAT-3718 Agent: aprender-59 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Contributor
Author
|
Closed: adds build-path Python, against the operator ruling of 2026-10-04. The work returns after the Python port ticket. Branch kept. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
PRCAP fold (cop order 14:02Z, 24/10 CRIT). Fold = MOVE: each source PR's head commit is an ancestor of this branch (merged with
--no-ff, so the original commits are kept). Built on origin/main15f1b2a496. Fold head:3c07f85b5191ca3088bf439acd36c00cc003a4d2.a6acb67602f7c18b1f26Clean merge, fmt clean.
Carried quorum/receipt files move with their commits. A fold changes the head, so this PR needs a new quorum receipt signed with
--ticketlisting every source ticket. Not armed (workers don't arm).Agent: aprender-59
🤖 Generated with Claude Code