Adversarial QA for LLM-RL environments. It tries to get paid without doing the task.
An RL environment is a grader, and graders have bugs. The expensive ones are quiet:
a format reward that pays for empty <answer></answer> tags, a substring match that
accepts every option at once, a reward function that throws on an empty completion and
gets silently scored 0.0 by the framework. You find out after the policy has spent a
week learning the bug instead of the task.
envdoctor attacks the grader directly with policies that never look at the question, and grades the results against what a correct answer actually earns.
No model, no API key, no credits. A scan is a few hundred local function calls.
pip install rl-envdoctor
envdoctor scan envdoctor.demo::gameableThe distribution is rl-envdoctor; the command and the import package are envdoctor.
envdoctor gameable-mcq
8 cases x 26 policies
reward ceiling 1.000 via oracle_answer_tag, payout threshold 0.500
policy kind mean max pay
------------------------- -------- ----- ----- ----
oracle_adaptive baseline 1.000 1.000 100%
oracle_corrupted_adaptive probe 1.000 1.000 100%
mc_shotgun exploit 0.700 0.700 100%
markdown_skeleton exploit 0.350 0.700 50%
copy_prompt exploit 0.300 0.300 0%
empty_answer_tag exploit 0.300 0.300 0%
empty exploit 0.000 0.000 0%
...
!! [critical] `mc_shotgun` earns 0.700 mean reward (70% of ceiling)
Emits every multiple-choice letter, so substring matching always hits.
repro: envdoctor scan envdoctor.demo::gameable --policy mc_shotgun --n 5 --show-trials
! [high] `empty_answer_tag` earns 0.300 mean reward (30% of ceiling)
Emits empty <answer></answer> tags.
!! [critical] Grader accepts corrupted answers (100% of full reward)
The grader is matching on shape, not on the answer.
Exit code is 1 when anything at or above --fail-on (default high) is found, so
envdoctor scan works as a CI gate with no wrapper.
We scanned 149 environments sampled from the 1,512 public environments on the Prime Intellect Environments Hub, then had an independent agent re-derive every finding from the environment's source and try to refute it.
| count | of | |
|---|---|---|
| scanned model-free | 87 | 149 sampled |
| confirmed reward defect | 28 | 87 scanned (32%) |
| a task-blind policy earns at least what a correct completion earns | 16 | 87 scanned (18%) |
| refuted as an envdoctor artifact | 13 | 41 flagged (32%) |
Confirmed root causes: format rewards paid for scaffolding (12), substring and fuzzy matching (5), constraints that an empty string satisfies vacuously (4), a reward function that never reads the completion (1), a reward function that raises and is silently scored 0.0 (1).
Three of them, in the authors' own code:
bhogan94/q-programming-languagekeys its test-case lookup on the raw prompt string, which stops matching once prompts are formatted as chat, so a fallback pays a flat 0.5 to any non-empty output. The constant"42"and the reference Q solution both score 0.500 and the Q interpreter never runs.- Six
maziyar/OpenMed_*environments print a worked example ending in\boxed{B}in their own system prompt. Echoing the prompt is parsed as the answerBand collects full credit on every row whose gold answer is B. maziyar/OpenMed_MedKnowledgedelivers its answer column as an int, so both correctness reward functions raiseAttributeErroron every rollout. The framework swallows that as 0.0, leaving only content-blind format rewards: an echoed prompt scores 0.275 while a correct\boxed{C}scores 0.150.
The full write-up, including the 13 cases where this tool was wrong and why, is in audit/AUDIT.md. Per-environment records are in audit/results/.
Detection quality is a claim, so it is measured. envdoctor selftest scores the checks
against a zoo of environments whose defects are known by construction: every defect class
the Hub audit confirmed, plus the healthy shapes most likely to be flagged by mistake.
$ envdoctor selftest
ok pays_for_empty empty_ties_correct detected
ok worked_example_in_prompt prompt_leaks_answer detected
ok dense_but_correct healthy quiet
...
defects caught 15/15 (recall 100%), false positives 0/8
The healthy half is the half that matters. Building it caught a real false positive: on a
partial-credit grader, an empty <answer></answer> tag collects 0.44 of the ceiling from
string similarity alone, which is the metric's floor rather than a defect. Dense graders
now require a policy to take half of what a correct answer earns before it is reported.
| check | what it means |
|---|---|
empty_ties_correct |
Nothing outscores a blank page. The strongest thing this tool says. |
completion_independent_reward |
The reward never reads the completion, so no policy differs from any other. |
exploit_pays_out |
A policy that ignores the task earned reward. |
grader_accepts_wrong |
A wrong answer, built by perturbing the right one, still scores. |
prompt_leaks_answer |
The prompt hands the answer over, so echoing it is a strategy. |
judge_pays_for_tone |
Confident phrasing beats hedged phrasing at the same length and no content. |
reward_saturation |
Every completion gets the same reward, so there is no gradient. |
nondeterministic_reward |
The same completion scored twice gave two answers. |
grader_crash |
The reward function raised. Frameworks swallow this and return 0.0. |
invalid_reward |
NaN, infinity, or a reward outside [0, 1]. |
slow_grader |
Grading alone is slow enough to pace a training run. |
baseline_sanity |
A correct answer scored nothing, so nothing else here is calibrated. |
empty_scan |
No cases were graded, so a clean result would mean nothing was tested. |
envdoctor checks lists them; envdoctor policies lists the attacks.
"The empty string scored 0.3" means nothing on its own. Every scan replays the dataset's own answer to find the reward ceiling, and reports exploits as a share of it.
Environments state their answer format in the prompt: <reversed_text> tags, \boxed{},
#### 42. envdoctor reads that instruction out of the prompt and wraps the ground truth
the same way, which is how it calibrates environments whose format it has never seen.
If a correct answer still scores nothing, the run says so rather than reporting a clean bill of health. That environment probably needs a real rollout, with tools, multiple turns or code execution, and its grader has not been tested by this scan.
pip install 'rl-envdoctor[verifiers]'
envdoctor scan gsm8k --n 20
envdoctor scan my-env --env-arg num_examples=200 --env-arg use_think=trueBoth generations work and the right one is picked automatically. A module exposing
load_environment() is scanned through its rubric; a module exposing a Taskset is
scanned through its tasks' @vf.reward methods, using Task.score(trace, runtime=None).
Either way it is a dataset row in, a synthetic completion graded, a reward out, with no
rollout and no model.
That is why it is free, and it is also the limit. Environments scored during a rollout come back uncalibrated, and on v1 the rewards that need a container are skipped silently by the framework, so envdoctor records which ones it could not exercise:
skipped_runtime_rewards: ["executes"]
scored_model_free: []
An environment in that state has not really been scanned, and the report says so instead of reporting that nothing paid out.
One scan tells you about one environment. A sweep is what turns this into a claim about a collection, which is how the Hub audit was run:
envdoctor sweep env_a env_b env_c --json sweep.json
envdoctor sweep --from-file environments.txt --fail-on criticaltarget worst findings ceiling calibrated top exploit
----------------------------- -------- -------- ------- ---------- ----------------
envdoctor.demo::gameable critical 6 1.0 yes mc_shotgun 0.700
envdoctor.demo::healthy clean 0 1.0 yes -
1 of 2 failed at high or worse
A target that fails to load is recorded and the sweep continues, because one broken environment should not cost you the other forty nine.
pip install 'rl-envdoctor[inspect]'
envdoctor scan mypkg.evals::my_taskInspect scorers have the same failure modes as RL graders, because they are the same
thing under a different name. includes() accepts the target anywhere in the output, so
a completion that lists every option satisfies it without answering.
Model-graded scorers are skipped rather than scored, because grading with a model costs money and is not reproducible. The report names the ones it skipped instead of returning a zero that looks like a strict grader.
Wrap the grader and scan it in your own test suite:
from envdoctor import Case, FunctionTarget, run_checks, scan
target = FunctionTarget(my_reward_fn, [Case(id="0", prompt=..., answer=...), ...])
report = run_checks(scan(target))
assert not [f for f in report.findings if f.severity == "critical"]Or point the CLI at it: envdoctor scan mypkg.envs::target.
- uses: really-notabot/envdoctor@v0
with:
target: my-env
fail-on: high
- uses: github/codeql-action/upload-sarif@v3
if: always()
with:
sarif_file: envdoctor.sarifThe SARIF upload puts each finding in the Security tab and on the pull request diff, with a stable fingerprint so a finding is tracked across runs rather than reopened every time the wording changes.
An environment with findings already in it should not fail on day one, or the gate gets switched off. Record what you know about and fail only on what is new:
envdoctor scan my-env --json envdoctor-baseline.json # once, then commit it
envdoctor scan my-env --baseline envdoctor-baseline.jsonbaseline: 0 new, 6 known, 1 no longer found
fixed, drop from the baseline: exploit_pays_out:copy_prompt
Installing the package registers a pytest plugin, so gameability can be an ordinary assertion next to your other tests:
from envdoctor.pytest_plugin import assert_not_gameable
def test_my_env_is_not_gameable(envdoctor_scan):
assert_not_gameable(envdoctor_scan("mypkg.envs::target"))
def test_my_reward_function(envdoctor_scan_function):
report = envdoctor_scan_function(my_reward_fn, my_cases)
assert_not_gameable(report, fail_on="medium")The failure message carries the repro command for the policy that scored.
Scanning an environment runs its code. Installing the package runs its build, importing
the module runs everything at module scope, and load_environment() or a Taskset
constructor runs whatever the author wrote. envdoctor does not sandbox any of that and
cannot: calling the grader is the entire technique.
So treat envdoctor scan some-hub-env as equivalent to running an untrusted program,
because it is. For anything off a public hub, run it in a container or a throwaway VM,
with no credentials in the environment. The Hub audit in audit/ ran every environment
in a subprocess with a timeout, which contains a hang or a crash; it does not contain a
hostile author, and nothing here should be read as claiming otherwise.
Environment quality is a known problem, mostly written up rather than tooled.
arXiv:2606.16062 found that 28.5% of sampled
SWE-bench-Verified tasks accept incorrect patches; verifiers v1 ships a per-task
validate() hook for checking that a task's own verifier accepts the gold answer.
envdoctor is the adversarial counterpart, and it is meant to be run on every commit
rather than once for a paper.
Built with Claude Code.
MIT