Defensible answers about your AI system, instead of vibes.
A skill for any AI coding agent that tells you when you need an eval,
where it fits, how to build it, and how to read what comes out.
English · 简体中文 · Español · 한국어 · 日本語
Evals are suddenly everywhere. Every AI talk, every launch post, every hiring thread says you need them. Then you sit down to actually do it, and the questions start.
Does a prompt tweak really need an eval, or is that overkill? At which point in the build does the first one go in? What is an "eval harness", concretely, beyond a folder named evals/? Which metric, how many examples, does a judge model even count? And when a number finally comes out, is 34 out of 40 good? Is a 3-point gain real, or noise? Or did the run just crash quietly and report a pass?
Answer these by feel and you get exactly what most teams have: a benchmark nobody trusts, a gate that has been green for a month because it grades nothing, and a number in the README that would not survive one sharp question.
Eval Genius is that missing judgment, packaged as a skill your AI agent runs with you. It thinks like a measurement engineer: decide what "better" means before you look, push every check you can down to plain code, treat a crash as unmeasured rather than a pass, and trust the number last.
It is not a course you have to read first. You describe where you are, in plain words, and it takes the next step, whether that step is "you don't need one yet" or "here is the gate, and here is why this run cannot be trusted."
It is tool-agnostic and dependency-free: a SKILL.md plus a few standard-library Python scripts. It runs in Claude Code, or any agent that loads skills, or from your terminal on its own.
Real questions, answered from wherever you actually are:
- "Do I need evals for my chatbot, or is that overkill right now?"
- "Where does an eval even go in my build?"
- "I got 34 out of 40, is that good?"
- "Is this 3-point gain real, or noise?"
- "Calibrate my LLM judge against some human labels."
No setup ritual, no vocabulary you have to learn first. Describe the situation, get the next move.
It walks the whole path, and meets you at any point on it, including the start:
- Decides whether you need an eval at all, and which kind belongs at your stage, from first prototype to production.
- Picks the eval: what to measure, which grader (code first, a judge only where no assertion works), which metric, how many examples, adopt a public benchmark or build your own.
- Builds and gates it: fixture, runner, scorer, reporter, a bar written before the run, and a CI gate that ends in PASS, FAIL, or CANNOT-MEASURE and refuses to compare mismatched runs.
- Reads the result with you: against the bar you wrote, with noise bounds, per-item diffs, and a harness-bug check before any surprising number is believed.
- Writes it up honestly, with caveats, tiers, and the comparison rule stated out loud.
- Refuses the shortcuts that produce pretty lies: bars moved after the fact, blended scores, run-until-green, and judges nobody calibrated.
Three standard-library scripts ship with the skill and run standalone:
| Script | What it settles |
|---|---|
check_gate.py |
Compares a change against its baseline per item; exits 0 PASS, 1 FAIL, 2 CANNOT-MEASURE, so a crash can never masquerade as a pass |
paired_bootstrap.py |
Puts a confidence interval on the difference, so "it improved" actually means something |
judge_agreement.py |
Measures how much your LLM judge agrees with human labels, before you let it grade anything |
Skills are just a folder. Put it where your agent looks for them:
# Claude Code
cp -R skill ~/.claude/skills/eval-genius
# Any other agent: point it at skill/SKILL.md, or add the skill/ folder to its skills pathThen talk to it in plain language ("do I need evals for my chatbot?", "is this delta real?", "calibrate my judge"). The scripts also run on their own:
python3 skill/scripts/check_gate.py --baseline base.json --treatment treat.jsonThe full method behind the skill, in one plain-language document, lives in METHODOLOGY.md. You do not need it to use the skill; it is there if you want to see the thinking.





