Framework for evaluating and improving agents
-
Updated
Aug 22, 2026 - Python
Framework for evaluating and improving agents
A curated, non-BS library of the best resources for building and evaluating AI agents — papers, blogs, talks, tools, benchmarks. Maintained by BenchFlow.
A Universal Platform for Training and Evaluation of Mobile Interaction
A graphical interface for reinforcement learning and gym-based environments.
Interoperating between (Deep) Reiforcement Learning libraries
Gymnasium-style API standard for RL environment creation in JAX
Workspace manager for coding agents. Interactively solve and develop Harbor tasks.
Create new gridworld gym environments easily
Adversarial QA for LLM-RL environments: find out what reward an empty answer earns. Model-free, zero API cost.
Turn any real software into a replayable RL environment for training AI agents — deterministic replay, verifiable rewards, TRL & verifiers adapters.
Foundry Lite: a public runnable sample of Veyl’s local environment harness for software-engineering agent evals.
Sound error bounds, symbolic GPU safety checks and targeted falsification for Triton kernels. 88% of planted bugs pass the standard fixed-shape allclose test; Litmus catches 100% with 0% false positives, on CPU.
A lightweight, open-source framework that turns historical GitHub pull requests into reproducible, verifiable software-engineering tasks for training and evaluating coding agents.
Comprehensive AI agent evaluation platform — searchable benchmark catalog, comparison matrices, automated scanner, interactive dashboards, and community-curated best practices for LLM evaluation.
Agent-evaluation environments: planted-truth worlds, ungameable graders, calibrated difficulty
Open-source RL environments and evals for the capabilities we want AI to have. Judge-free scoring, mandatory baselines, and an enforced defensive-asymmetry gate.
Outcome-verified agent trajectories, benchmarks, and RL environments — with a live leaderboard and a CI gate for your agents. Offline-first, MIT.
A Harbor evaluation task for coding agents, published with two working attacks on its own grader and the fixes that close them.
Claude Code Agent Skills for building, red-teaming and tuning agentic RL evaluation environments — a four-skill pattern (guardian, validation-debugger, score-tuner, iteration-loop) plus a 24-point adversarial reviewer.
Add a description, image, and links to the rl-environments topic page so that developers can more easily learn about it.
To associate your repository with the rl-environments topic, visit your repo's landing page and select "manage topics."