Skip to content

About

Stop guessing at evals. A skill for any AI agent that tells you when you need one, which to run, how to build it, and whether the number would survive a hard question. Deterministic-first, tool-agnostic.

Resources

Stars

1 star

Watchers

0 watching

Forks

 
 

Repository files navigation

Eval Genius, the star-cloaked measurement sage

Eval Genius

Defensible answers about your AI system, instead of vibes.

A skill for any AI coding agent that tells you when you need an eval,
where it fits, how to build it, and how to read what comes out.

License: Apache 2.0 Works across any agent Stdlib Python, zero dependencies

English · 简体中文 · Español · 한국어 · 日本語


Everyone says "you need evals." Almost nobody says when.

The problem

Eval Genius at a crossroads of floating paths, unsure which way the evals go

Evals are suddenly everywhere. Every AI talk, every launch post, every hiring thread says you need them. Then you sit down to actually do it, and the questions start.

Does a prompt tweak really need an eval, or is that overkill? At which point in the build does the first one go in? What is an "eval harness", concretely, beyond a folder named evals/? Which metric, how many examples, does a judge model even count? And when a number finally comes out, is 34 out of 40 good? Is a 3-point gain real, or noise? Or did the run just crash quietly and report a pass?

Answer these by feel and you get exactly what most teams have: a benchmark nobody trusts, a gate that has been green for a month because it grades nothing, and a number in the README that would not survive one sharp question.

Meet Eval Genius

Eval Genius with a compass and a star-map, weighing results on a set of scales

Eval Genius is that missing judgment, packaged as a skill your AI agent runs with you. It thinks like a measurement engineer: decide what "better" means before you look, push every check you can down to plain code, treat a crash as unmeasured rather than a pass, and trust the number last.

It is not a course you have to read first. You describe where you are, in plain words, and it takes the next step, whether that step is "you don't need one yet" or "here is the gate, and here is why this run cannot be trusted."

It is tool-agnostic and dependency-free: a SKILL.md plus a few standard-library Python scripts. It runs in Claude Code, or any agent that loads skills, or from your terminal on its own.

What you can ask it

Eval Genius working hands-on, pulling a star-map into a laptop

Real questions, answered from wherever you actually are:

  • "Do I need evals for my chatbot, or is that overkill right now?"
  • "Where does an eval even go in my build?"
  • "I got 34 out of 40, is that good?"
  • "Is this 3-point gain real, or noise?"
  • "Calibrate my LLM judge against some human labels."

No setup ritual, no vocabulary you have to learn first. Describe the situation, get the next move.

What it does for you

Eval Genius crossing floating platforms through a pass/fail gate toward the results

It walks the whole path, and meets you at any point on it, including the start:

  • Decides whether you need an eval at all, and which kind belongs at your stage, from first prototype to production.
  • Picks the eval: what to measure, which grader (code first, a judge only where no assertion works), which metric, how many examples, adopt a public benchmark or build your own.
  • Builds and gates it: fixture, runner, scorer, reporter, a bar written before the run, and a CI gate that ends in PASS, FAIL, or CANNOT-MEASURE and refuses to compare mismatched runs.
  • Reads the result with you: against the bar you wrote, with noise bounds, per-item diffs, and a harness-bug check before any surprising number is believed.
  • Writes it up honestly, with caveats, tiers, and the comparison rule stated out loud.
  • Refuses the shortcuts that produce pretty lies: bars moved after the fact, blended scores, run-until-green, and judges nobody calibrated.

The parts that are easiest to get wrong, handled

Eval Genius at a desk with a checklist, a bell curve, and a judge-vs-human scale

Three standard-library scripts ship with the skill and run standalone:

Script What it settles
check_gate.py Compares a change against its baseline per item; exits 0 PASS, 1 FAIL, 2 CANNOT-MEASURE, so a crash can never masquerade as a pass
paired_bootstrap.py Puts a confidence interval on the difference, so "it improved" actually means something
judge_agreement.py Measures how much your LLM judge agrees with human labels, before you let it grade anything

Install

Skills are just a folder. Put it where your agent looks for them:

# Claude Code
cp -R skill ~/.claude/skills/eval-genius

# Any other agent: point it at skill/SKILL.md, or add the skill/ folder to its skills path

Then talk to it in plain language ("do I need evals for my chatbot?", "is this delta real?", "calibrate my judge"). The scripts also run on their own:

python3 skill/scripts/check_gate.py --baseline base.json --treatment treat.json

Curious about the reasoning?

The full method behind the skill, in one plain-language document, lives in METHODOLOGY.md. You do not need it to use the skill; it is there if you want to see the thinking.


Built by Alex Greenshpun. If it helps, a star or a share helps others find it.

GitHub Website X

License: Apache 2.0

About

Stop guessing at evals. A skill for any AI agent that tells you when you need one, which to run, how to build it, and whether the number would survive a hard question. Deterministic-first, tool-agnostic.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages