A multi-agent framework for automated mechanistic interpretability research
Seesaw automates the hypothesis → experiment → critique loop in mechanistic interpretability. A human researcher provides a research question; three agents handle the rest.
Scout research question → Research Plan → Lens Research Plan → ExperimentBundle → Quill ExperimentBundle → CritiqueReport
Human-in-the-loop checkpoints sit between every stage, so you stay in control of what actually runs. Each agent is also independently usable as a standalone MCP server — see DOCS.md for how to run each one on its own.
uv venv .venv --python 3.12 && source .venv/bin/activate
uv pip install -e .
cp .env.example .env # add ANTHROPIC_API_KEY (required), FIRECRAWL_API_KEY (optional)
python -m orchestrator.src.main --question "What attention heads mediate indirect object identification in GPT-2 Small?"Full setup, per-agent usage, and deployment notes live in DOCS.md.
|
The Research Planner Searches arXiv and the web, then produces a structured Research Plan — hypotheses, target models, and concrete experiments to run. Tools
Stack LangGraph ReAct · Claude · Firecrawl |
The Experiment Runner Executes mechanistic interpretability experiments from the Research Plan via TransformerLens, then generates an LLM interpretation of each result.
Stack LangGraph StateGraph · Claude · TransformerLens |
The Reviewer Reviews the ExperimentBundle like a scientific peer reviewer — flags methodological issues, unsupported conclusions, and missing experiments. Outputs follow-up specs Lens can execute directly. Stack LangGraph StateGraph · Claude Opus |
Every agent shares the same internal structure — app/, config/, models/, tools/, routers/, resources/, prompts/, server.py — so the codebase reads consistently whether you're in Scout, Lens, or Quill. Details in DOCS.md.
The pipeline itself is deliberately simple: a fixed-order prompt-chaining workflow, not autonomous multi-agent routing. Scout always runs first, Lens second, Quill third — no agent decides who runs next. That keeps the system predictable and easy to debug, with room to grow into an Evaluator-Optimizer loop between Lens and Quill later.
Most mechanistic interpretability work is done by hand — identify a circuit, ablate it, write it up. Seesaw is an experiment in automating that research loop, with the goal of accelerating exploratory safety research on small models.
