A lie detector for AI coding agents. It runs your checks itself and lets the
real exit code overrule the agent's self-report β PROVEN β
or OVERCLAIM β.
receipts over self-report
nocap (slang, "no cap"): no lie, for real. An agent is "capping" when it claims a task is done that it didn't actually finish. nocap catches that.
Your AI agent says "bug fixed, all tests pass β " β but the tests never ran, the build is broken, or the file wasn't actually changed. You trust it, merge, and break prod. Every Cursor / Claude Code / Cline / Aider user has been here.
nocap is the receipt. It doesn't trust what the agent says β it runs the check itself and tells you whether the claim is PROVEN or an OVERCLAIM.
Zero install, zero config β try it right now:
npx @zoeyx/nocap demoAdjudicate a single command (exit 0 β PROVEN, anything else β OVERCLAIM):
npx @zoeyx/nocap verify "npm test" --claim "bug fixed, tests pass"
# exits 0 if the real exit code is 0, exits 1 on OVERCLAIMCreate a .nocap.yml in your repo (or run npx @zoeyx/nocap init):
claim: "login form validates email and shows errors"
verify:
- id: unit-tests
run: "npm test"
- id: build
run: "npm run build"Then:
npx @zoeyx/nocap checknocap runs every verify step itself and adjudicates: all green β PROVEN,
any red β OVERCLAIM. The agent's narrative never counts β only the receipt.
# .github/workflows/nocap.yml
on: [pull_request]
jobs:
nocap:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with: { node-version: 20 }
- run: npm ci
- uses: zhouyuanxinand/nocap@v1
with:
command: checkThe check fails (nonzero exit) on OVERCLAIM, so a lying agent can't get its PR merged.
npx @zoeyx/nocap hook installInstalls a pre-push hook that runs nocap check and blocks the push while
any verify step is red. Remove it with nocap hook uninstall.
nocap --json check{
"verdict": "PROVEN",
"claim": "login form validates email and shows errors",
"steps": [
{ "id": "unit-tests", "command": "npm test", "exitCode": 0, "provenance": "harness" }
]
}nocap distills a single principle β receipt > observation β into a tiny CLI:
- An observation is what the agent says ("done", "tests pass").
- A receipt is a command nocap ran itself, capturing the real exit code and hashing the output.
- Only a receipt can prove a claim. The agent cannot self-approve.
That's it. No prompts, no models, no theory β just a deterministic check whose exit code is the verdict. The agent is free to claim whatever it wants; nocap runs the proof.
For the curious: receipts are persisted under .nocap/runs/<id>/receipts/ with
SHA-256 hashes of stdout/stderr, so a captured result can be re-verified later
(tamper detection).
| Command | What it does |
|---|---|
nocap demo |
self-contained "catch an agent lying" scenario β no setup |
nocap verify "<cmd>" |
run one command, adjudicate it |
nocap check |
read .nocap.yml, run every verify step, adjudicate |
nocap init |
scaffold a .nocap.yml (auto-detects your test command) |
nocap hook install / uninstall |
manage the pre-push git hook |
--json |
machine-readable output (any command) |
Exit code: 0 on PROVEN, 1 on OVERCLAIM, 2 on usage error.
nocap is model- and host-agnostic. It doesn't care which agent wrote the code β Cursor, Claude Code, Cline, Aider, Windsurf, Continue, GitHub Copilot Workspace, or your own. If it can lie about being done, nocap can catch it.
npm install -g @zoeyx/nocap # global
# or just use it ad hoc:
npx @zoeyx/nocap checkRequires Node β₯ 18.
| what it checks | how it decides | |
|---|---|---|
| nocap | an agent's completion claim | re-runs the receipt (real exit code) |
| Guardrails AI / NeMo | LLM output content | validates generated text |
| OpenAI Evals / promptfoo | model quality offline | scores a dataset |
| Langfuse / Opik | agent traces | observes, doesn't adjudicate |
nocap is the only one that exists to answer one question: did the agent actually do what it said it did?
The "receipt > observation" idea is inspired by agent-governance (MIT) β a deeper governance protocol. nocap is the distilled, model-agnostic, 30-seconds-to-set-up version of its core check.
MIT