Probity rules that know which tests reach the code the agent just changed.
Probity's enforceTdd knows that a failing test was observed. It cannot know that the file being edited is covered by 29 other test files in four packages. These rules read a graphify code graph, walk it backwards from the changed files, and refuse to let the agent commit or end its turn until those tests have run green.
Deterministic, no LLM calls, no API key. A reverse traversal on a 15k-edge graph takes about 25ms.
TDAD measured three agent setups on 100 SWE-bench Verified tasks:
| setup | test-level regressions | pass-to-pass failures |
|---|---|---|
| vanilla agent | 6.08% | 562 |
| TDD prompting only | 9.94% | 799 |
| graph-selected impacted tests + TDD | 1.82% | 155 |
Prompting an agent to do TDD made regressions worse. The reduction came from selecting impacted tests from a code graph and running them before submitting. That selection step is what this package adds.
Same issue (tRPC #7468), same model, three runs:
| no rules | rules bypassed | all four rules | |
|---|---|---|---|
| TDD order | implementation first | invisible | test first |
| source lines changed | 26 | 26 | 16 |
unrequested isDone guard |
8 lines | 8 lines | none |
| test lines | 102 | 72 | 105 |
| blocks | 1 | 0 | 4 |
The point is the third column. The agent wrote less production code and more test, because every write had to be justified by a failing assertion and every commit had to be verified against the tests that reach it.
What the agent saw in run 1, at the commit it tried to make:
Probity: 29 test file(s) reach your changes and have not run green since your last write:
packages/tests/server/regression/issue-7468-http-links-abort-on-unsubscribe.test.ts (written this session)
packages/react-query/test/abortOnUnmount.test.tsx (imports, depth 1)
packages/tests/server/errors.test.ts (imports, depth 1)
... and 26 more
Run: pnpm vitest run packages/tests/server/regression/issue-7468-...
then retry.
It had hand-picked three test files to run. abortOnUnmount.test.tsx was not one of them, on an abort-on-unsubscribe fix.
Not on npm yet. For now:
git clone https://github.com/Sandbye/probity-plugin-graphify.git
cd probity-plugin-graphify && pnpm install && pnpm buildThen in your own repo:
pnpm add -D @nizos/probity
pnpm add -D file:../probity-plugin-graphify
pip install graphifyy # or: uv tool install graphifyy
graphify extract . --code-only --no-clusterThe last command writes graphify-out/graph.json. Rebuild it with graphify update . after a refactor, or install graphify's git hook to keep it warm.
// probity.config.ts
import { defineConfig, enforceTdd } from '@nizos/probity'
import {
forbidShellFileWrites,
requireImpactedTests,
requireReachingTest,
} from 'probity-plugin-graphify'
export default defineConfig({
rules: [
forbidShellFileWrites(),
requireImpactedTests({ depth: 1, runner: 'pnpm vitest run {files}' }),
{
files: ['src/**', 'test/**'],
rules: [enforceTdd(), requireReachingTest()],
},
],
})For the turn-end gate, add a Stop hook alongside Probity's PreToolUse one:
{
"hooks": {
"Stop": [
{
"hooks": [
{
"type": "command",
"command": "npx probity-graphify stop --depth 1 --runner 'pnpm vitest run {files}'"
}
]
}
]
}
}Gates a command (default git commit) on every test file that reaches this session's changes having run green since the last code change. Changed files come from the session transcript and git working-tree state, so an edit made any way at all is covered.
| option | default | |
|---|---|---|
before |
/git commit/ |
which commands to gate |
depth |
2 |
reverse traversal depth |
runner |
npx vitest run {files} |
suggested command, {files} substituted |
maxListed |
12 |
how many files to itemize |
graph |
nearest graphify-out/graph.json |
explicit graph path |
ignoreGit |
false |
use transcript writes only |
narrow |
false |
seed from the changed symbols instead of whole files |
depth counts file boundaries crossed, not edges. A call chain inside one file is part of the same change surface, so it costs nothing, and depth: 1 keeps meaning "the modules that use this directly" wherever in the file the edit landed. Use depth: 1 on repos with hub modules: in tRPC, one file in packages/server/src/observable/ reaches 22 test files at depth 1 and 79 at depth 2, out of 170.
demo/benchmark-recall.mjs scores selection against real history. When a commit changes source and test files, the tests its author touched are tests a human judged relevant, so the question becomes: from the source diff alone, how much of that judgement does the graph recover, and how many test files does it make the agent run?
51 tRPC commits, 177 test files in the suite:
| selection | recall of author-chosen tests | commits with nothing found | test files run |
|---|---|---|---|
sibling naming (x.ts -> x.test.ts) |
8% | - | 0.5 |
| graph, depth 1 | 58% | 18 of 51 | 9.4 (5% of suite) |
| graph, depth 2 | 82% | 5 of 51 | 49.8 (28% of suite) |
graph, depth 2, narrow |
73% | 8 of 51 | 43.2 |
Three things to take from that, in order of how much they should change your mind:
- A graph is worth having. Naming convention recovers 8% of what a human would pick. Depth 2 recovers 82%.
- This is not safe selection. At depth 2, 5 of 51 changes would have had no relevant test named. Use it to make the agent run the tests it would otherwise skip, not as a claim that a green gate means no regression.
narrowis off by default because the benchmark said so. Seeding from the changed symbols looked like the obvious win, and it cost 9 points of recall to avoid 6.6 test files. Cheaper is the wrong direction for a guard.
Caveat on the ground truth: it credits only tests the author touched. A test they should have touched and did not is invisible, and a test they edited for an unrelated reason counts as a miss. It measures agreement with human judgement, not regression prevention.
depth counts file boundaries crossed, not edges. A call chain inside one file is part of the same change surface, so it costs nothing, and depth: 1 keeps meaning "the modules that use this directly" wherever in the file the edit landed.
Blocks a write to a source file that no test reaches. Paired with enforceTdd, this makes the failing test have to be a related one. Passes for test files, for files the graph does not know (new files, which enforceTdd owns), and when a test naming the file was written earlier in the session.
Blocks shell file mutation (cat > f, heredocs, sed -i, tee, inline scripts calling write_text) so edits go through the Write and Edit tools, where every write-gated rule can see them.
The bypass is real, not hypothetical. With Claude Code's auto mode active the agent is told to make file changes through the shell: in run 2 above, all 42 tool calls were Bash, zero were Write or Edit, and every write-gated rule including Probity's own enforceTdd saw an idle session while three files changed and a commit landed.
This rule is a stopgap, and a weak one. Probity issue #62 already tracks the underlying gap and argues for first-class command-to-write mapping inside Probity, with a better analysis than this regex: tokenizing the command into argv, following bash -c payloads, and stripping wrapper prefixes. Measured against the shapes that issue lists, this rule catches 7 of 15. It misses perl -i, awk into a file, dd of=, cp, git apply, sed --in-place, python -c open(), and anything wrapped in bash -c. Use it as a speed bump against the auto-mode default, not as a security boundary, and watch #62 for the real fix.
Claude Code Stop hook. Refuses the end of a turn while tests reaching the session's changes are unverified. Catches the common case where the agent never commits at all, because you do.
Every path fails open. No graph, an unreadable graph, a graph.json in a format this version does not know, an unknown file, a git command that fails: the rule passes and carries the reason. Probity fail-closes on a throwing rule, so failing open matters, otherwise a graphify upgrade would block every commit in the repo.
- Static. Imports and calls in source. Reflection, string-keyed dispatch and runtime DI are invisible. TDAD covered this with coverage-learned links; not implemented here.
- Symbol-level, but only where the diff allows it. A change touching module header lines widens to the whole file, and cross-package
imports_fromedges that land on a file node rather than a symbol cannot be narrowed at all. - Workspace packages that export
distneed either tsconfigpaths(tRPC has them) or a"source": "./src/index.ts"condition inexports, or cross-package reach is invisible. Adding that condition is inert for node and tsc. - The graph lags the session. A test written this session is handled; a new caller added this session is not seen until
graphify updateruns. enforceTddcosts a turn per write in scope, around 5s. These rules cost nothing.forbidShellFileWritescatches 7 of the 15 shell-write shapes listed in Probity issue #62. A speed bump, not a boundary.
Tested against @nizos/probity 1.10.0 and graphifyy 0.9.53. A weekly CI job runs the suite plus a contract check against the latest release of both, so upstream drift surfaces here rather than in your session.
- Probity by Nizar Selander (MIT), which these rules plug into. The shell-redirect detection in
forbidShellFileWritesfollows the pattern used by Probity's ownforbidCommandPatternconfig. - graphify by Graphify-Labs (Apache-2.0), which builds the code graph.
DEFAULT_RELATIONSmirrors graphify's ownDEFAULT_AFFECTED_RELATIONSsographify affectedand these rules agree. - TDAD by Rafael Alonso (MIT), the source of the approach and the regression numbers.
MIT