Multi-agent orchestrated code review with adversarial debate and W&B Weave tracing.
Built for the AGI House Γ W&B Multi-Agent Orchestration Build Day, May 2026.
Paste a code diff β 4 specialist AI agents review it in parallel β they debate each other's findings β a Lead agent consolidates into a final ranked review with a risk score.
βββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β CODE DIFF INPUT β
ββββββββββββββββββββββββ¬βββββββββββββββββββββββββββββββ
β
βββββββββββββββΌββββββββββββββ
β β β
ββββββΌβββββ ββββββΌβββββ ββββββΌβββββ βββββββββββ
βπSecurityβ ββ‘ Perf β βπ§ Logic β βπ¨ Style β
ββββββ¬βββββ ββββββ¬βββββ ββββββ¬βββββ ββββββ¬βββββ
β β β β
βββββββββββββββΌββββββββββββββββββββββββββββ
β
ββββββββββΌβββββββββ
β DEBATE ROUND β β Agents challenge each other
β (adversarial) β
ββββββββββ¬βββββββββ
β
ββββββββββΌβββββββββ
β LEAD AGENT β β Deduplicates, ranks, scores
β (consolidate) β
ββββββββββ¬βββββββββ
β
ββββββββββΌβββββββββ
β FINAL REVIEW β β Risk score 0-10
β + Weave Trace β β Full orchestration visible
βββββββββββββββββββ
| Phase | Agents | What Happens |
|---|---|---|
| 1. Parallel Review | Security, Performance, Logic, Style | Each specialist runs independently on the same diff |
| 2. Adversarial Debate | All specialists + Moderator | Each agent sees others' findings and can challenge false positives |
| 3. Lead Consolidation | Lead Engineer | Deduplicates, ranks by severity Γ confidence, produces final score |
Multi-agent orchestration patterns used:
- Parallel execution (Phase 1)
- Adversarial/critic loops (Phase 2)
- Hierarchical consolidation (Phase 3)
- All traced as nested Weave operations
cd /path/to/this/repo
python -m venv .venv
source .venv/bin/activate
pip install -e .export WANDB_API_KEY="your-wandb-key"
export WEAVE_PROJECT="code-review-swarm"
# Choose your backend:
export REVIEW_SWARM_BACKEND=wandb_inference # W&B Serverless Inference (recommended β uses W&B credits)
# export REVIEW_SWARM_BACKEND=anthropic # Claude (needs ANTHROPIC_API_KEY)
# export REVIEW_SWARM_BACKEND=simulate # Mock mode for offline demosOr copy .env.example to .env and fill in values.
# Default killer demo β reviews real Daytona.io workspace manager code
review-swarm demo
# List all demo scenarios
review-swarm demo --list-scenarios
# Run a specific scenario
review-swarm demo -s daytona-sandbox-api
review-swarm demo -s daytona-sdk-auth
review-swarm demo -s daytona-workspace # (default β has SQL injection + cmd injection + N+1)
review-swarm demo -s basic # synthetic examples# By URL
review-swarm github https://github.com/daytonaio/daytona/pull/4822
# By shorthand
review-swarm github daytonaio/daytona#4858# From a file
review-swarm review -f path/to/diff.patch
# From git diff via stdin
git diff | review-swarm review --stdin
# From git diff of staged changes
git diff --cached | review-swarm review --stdin
# Review a GitHub PR directly
review-swarm review -g daytonaio/daytona#4858# Runs against 8 labeled known-buggy diffs, scores precision/recall
review-swarm evaluate# Get a PR diff and pipe it in
gh pr diff 42 | review-swarm review -s# Explicit project
review-swarm review --demo -p "my-team/hackathon-demo"
# Or set WEAVE_PROJECT env var
export WEAVE_PROJECT="my-team/hackathon-demo"
review-swarm demoAfter running, open the link printed in the terminal to see:
- The full agent orchestration graph
- Each specialist's findings as nested traces
- The debate round with challenges and verdicts
- Token usage and latency per agent
- The final consolidated review
Every operation is decorated with @weave.op(), so you get:
| Trace | What It Shows |
|---|---|
π Code Review Swarm Pipeline |
Top-level orchestration β the full pipeline |
π Specialist: {category} (Γ4) |
Each specialist's input/output, tokens, latency |
π£οΈ Adversarial Debate |
Challenges raised, verdicts, false positives caught |
βοΈ Challenge: {X} vs {Y} (ΓN) |
Individual debate challenges as separate traced nodes |
π Lead Consolidation |
Final deduplication and ranking logic |
π Review Metrics |
Structured metrics: severity breakdown, token usage |
π― Quality Scorer |
Inline quality scores for Weave Monitors |
This project deeply integrates with the W&B platform across 7 features:
Every function in the pipeline is traced with descriptive call_display_name labels. The trace tree shows the full multi-agent orchestration graph with timing, inputs, and outputs.
All 6 agent prompts are published to Weave as versioned StringPrompt objects:
security_specialist,performance_specialist,logic_specialist,style_specialistlead_consolidator,debate_moderator
Edit prompts in the Weave UI β changes take effect on next run without code changes.
review-swarm publish-prompts # Explicitly publish/update promptsCodeReviewSwarmModel is a weave.Model subclass with versioned parameters (model name, specialists list, debate toggle). Published to Weave for comparison across experiments.
- Published
Dataset(known-bugs-eval-set) with 8 labeled buggy diffs EvaluationLoggerwith 6 scored metrics per sample- Results appear in the Evals tab with comparison tables
review-swarm evaluate # Run eval suite β see results in Weave Evals tabEvery review call runs an inline π― Quality Scorer that emits:
has_critical_findings,average_confidence,debate_was_meaningfulfalse_positives_caught,quality_score
Set up a Monitor in the Weave UI β select the review_diff op β auto-score production traffic.
Uses meta-llama/Llama-3.3-70B-Instruct via W&B's OpenAI-compatible inference API:
- No Anthropic key needed β uses W&B credits ($100 included)
- Full prompt/completion traces logged automatically by Weave
- Shows real LLM latency, tokens, and model responses in the trace
Structured metrics logged per call: risk score, severity breakdown, false positive rate, token usage. Enables dashboards and trend analysis in Weave.
The eval suite tests against 8 known-buggy code samples:
| Metric | What It Measures |
|---|---|
| Category Recall | Did the swarm find the right type of bug? |
| Severity Accuracy | Did it rate the severity correctly? |
| Precision | What fraction of findings are actually relevant? |
| Detection Rate | Did it find anything at all? |
- Show the architecture β explain the 3-phase pipeline (30 sec)
- Run
review-swarm demoβ reviews real Daytona.io code with W&B Inference, watch the orchestration tree with findings (60 sec) - Open Weave β show the trace graph with all agents visible as named nested operations (30 sec)
- Show Prompts in Weave β demonstrate editing a prompt in the UI (15 sec)
- Run
review-swarm evaluateβ show precision/recall metrics in the Evals tab (30 sec) - Show Quality Scorer β point to the inline scoring op and how Monitors would pick it up (15 sec)
Key talking points:
- "4 specialist agents run in parallel β security, performance, logic, style"
- "Then they DEBATE each other β challenge false positives adversarially"
- "The lead agent consolidates after debate, only upheld findings survive"
- "Every step is traced in Weave β you can see the full orchestration graph"
- "Prompts are versioned in Weave β edit them in the UI, no code deploy needed"
- "We use W&B Serverless Inference β real LLM calls with Llama 3.3 70B"
- "Evaluation suite with published Dataset and 6 quality metrics"
- "Quality scorer runs inline for continuous monitoring via Weave Monitors"
- W&B Weave β tracing, evaluations, prompt management, model versioning, monitors
- W&B Serverless Inference β hosted LLM calls (Llama 3.3 70B), uses W&B credits
- Anthropic Claude (claude-sonnet-4-20250514) β optional alternative backend
- Python asyncio β parallel agent execution
- Rich β beautiful terminal UI for the demo
- Pydantic β structured data models for findings
- OpenAI SDK β W&B Inference client (OpenAI-compatible API)
review_swarm/
βββ __init__.py # Package init
βββ models.py # Pydantic data models (Finding, Review, etc.)
βββ prompts.py # System prompts for each specialist agent
βββ agents.py # Core orchestration engine (parallel β debate β lead)
βββ weave_prompts.py # Weave prompt publishing and loading
βββ mock_responses.py # Mock responses for simulate mode
βββ demos.py # Curated real-world diffs from Daytona.io
βββ github.py # GitHub PR diff fetcher (any public repo)
βββ evaluation.py # Eval harness with Weave EvaluationLogger + Dataset
βββ ui.py # Rich terminal rendering
βββ cli.py # Click CLI entry point
| Criterion | How We Address It |
|---|---|
| Agent Orchestration | 4 parallel specialists β adversarial debate β lead consolidation. Clear multi-agent handoffs with named trace nodes. |
| Utility | Solves real code review β finds security vulns, perf issues, logic bugs, style problems on real OSS code. |
| Technical Execution | Async parallel execution, multi-backend (W&B Inference + Anthropic + simulate), structured eval with 6 metrics. |
| Creativity | Adversarial debate between agents is novel β they challenge each other's false positives, reducing noise. |
| Sponsor Usage | 7 W&B features: Weave Tracing, Prompt Management, Model Versioning, Evaluations, Quality Scorer, W&B Inference, Metrics Dashboard. |