Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

Β 

History

4 Commits

Folders and files

Repository files navigation

🐝 Code Review Swarm

Multi-agent orchestrated code review with adversarial debate and W&B Weave tracing.

Built for the AGI House Γ— W&B Multi-Agent Orchestration Build Day, May 2026.

What It Does

Paste a code diff β†’ 4 specialist AI agents review it in parallel β†’ they debate each other's findings β†’ a Lead agent consolidates into a final ranked review with a risk score.

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                  CODE DIFF INPUT                      β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                       β”‚
         β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
         β”‚             β”‚             β”‚
    β”Œβ”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”
    β”‚πŸ”’Securityβ”‚  β”‚βš‘ Perf  β”‚  β”‚πŸ§  Logic β”‚  β”‚πŸŽ¨ Style β”‚
    β””β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”˜
         β”‚             β”‚             β”‚             β”‚
         β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                       β”‚
              β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”
              β”‚  DEBATE ROUND   β”‚  ← Agents challenge each other
              β”‚  (adversarial)  β”‚
              β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                       β”‚
              β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”
              β”‚  LEAD AGENT     β”‚  ← Deduplicates, ranks, scores
              β”‚  (consolidate)  β”‚
              β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                       β”‚
              β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”
              β”‚  FINAL REVIEW   β”‚  β†’ Risk score 0-10
              β”‚  + Weave Trace  β”‚  β†’ Full orchestration visible
              β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Architecture

Phase Agents What Happens
1. Parallel Review Security, Performance, Logic, Style Each specialist runs independently on the same diff
2. Adversarial Debate All specialists + Moderator Each agent sees others' findings and can challenge false positives
3. Lead Consolidation Lead Engineer Deduplicates, ranks by severity Γ— confidence, produces final score

Multi-agent orchestration patterns used:

  • Parallel execution (Phase 1)
  • Adversarial/critic loops (Phase 2)
  • Hierarchical consolidation (Phase 3)
  • All traced as nested Weave operations

Quick Start

1. Install

cd /path/to/this/repo
python -m venv .venv
source .venv/bin/activate
pip install -e .

2. Set Environment Variables

export WANDB_API_KEY="your-wandb-key"
export WEAVE_PROJECT="code-review-swarm"

# Choose your backend:
export REVIEW_SWARM_BACKEND=wandb_inference  # W&B Serverless Inference (recommended β€” uses W&B credits)
# export REVIEW_SWARM_BACKEND=anthropic      # Claude (needs ANTHROPIC_API_KEY)
# export REVIEW_SWARM_BACKEND=simulate       # Mock mode for offline demos

Or copy .env.example to .env and fill in values.

3. Run the Demo

# Default killer demo β€” reviews real Daytona.io workspace manager code
review-swarm demo

# List all demo scenarios
review-swarm demo --list-scenarios

# Run a specific scenario
review-swarm demo -s daytona-sandbox-api
review-swarm demo -s daytona-sdk-auth
review-swarm demo -s daytona-workspace   # (default β€” has SQL injection + cmd injection + N+1)
review-swarm demo -s basic               # synthetic examples

4. Review a Real GitHub PR

# By URL
review-swarm github https://github.com/daytonaio/daytona/pull/4822

# By shorthand
review-swarm github daytonaio/daytona#4858

5. Review Your Own Code

# From a file
review-swarm review -f path/to/diff.patch

# From git diff via stdin
git diff | review-swarm review --stdin

# From git diff of staged changes
git diff --cached | review-swarm review --stdin

# Review a GitHub PR directly
review-swarm review -g daytonaio/daytona#4858

6. Run the Evaluation Suite

# Runs against 8 labeled known-buggy diffs, scores precision/recall
review-swarm evaluate

Usage Examples

Review a PR diff

# Get a PR diff and pipe it in
gh pr diff 42 | review-swarm review -s

Review with Weave tracing

# Explicit project
review-swarm review --demo -p "my-team/hackathon-demo"

# Or set WEAVE_PROJECT env var
export WEAVE_PROJECT="my-team/hackathon-demo"
review-swarm demo

View traces in Weave

After running, open the link printed in the terminal to see:

  • The full agent orchestration graph
  • Each specialist's findings as nested traces
  • The debate round with challenges and verdicts
  • Token usage and latency per agent
  • The final consolidated review

What Gets Traced in Weave

Every operation is decorated with @weave.op(), so you get:

Trace What It Shows
🐝 Code Review Swarm Pipeline Top-level orchestration β€” the full pipeline
πŸ” Specialist: {category} (Γ—4) Each specialist's input/output, tokens, latency
πŸ—£οΈ Adversarial Debate Challenges raised, verdicts, false positives caught
βš”οΈ Challenge: {X} vs {Y} (Γ—N) Individual debate challenges as separate traced nodes
πŸ“Š Lead Consolidation Final deduplication and ranking logic
πŸ“ˆ Review Metrics Structured metrics: severity breakdown, token usage
🎯 Quality Scorer Inline quality scores for Weave Monitors

W&B / Weave Integration (Sponsor Usage)

This project deeply integrates with the W&B platform across 7 features:

1. Weave Tracing (@weave.op())

Every function in the pipeline is traced with descriptive call_display_name labels. The trace tree shows the full multi-agent orchestration graph with timing, inputs, and outputs.

2. Weave Prompt Management

All 6 agent prompts are published to Weave as versioned StringPrompt objects:

  • security_specialist, performance_specialist, logic_specialist, style_specialist
  • lead_consolidator, debate_moderator

Edit prompts in the Weave UI β†’ changes take effect on next run without code changes.

review-swarm publish-prompts  # Explicitly publish/update prompts

3. Weave Model Versioning

CodeReviewSwarmModel is a weave.Model subclass with versioned parameters (model name, specialists list, debate toggle). Published to Weave for comparison across experiments.

4. Weave Evaluations

  • Published Dataset (known-bugs-eval-set) with 8 labeled buggy diffs
  • EvaluationLogger with 6 scored metrics per sample
  • Results appear in the Evals tab with comparison tables
review-swarm evaluate  # Run eval suite β†’ see results in Weave Evals tab

5. Quality Scorer (for Weave Monitors)

Every review call runs an inline 🎯 Quality Scorer that emits:

  • has_critical_findings, average_confidence, debate_was_meaningful
  • false_positives_caught, quality_score

Set up a Monitor in the Weave UI β†’ select the review_diff op β†’ auto-score production traffic.

6. W&B Serverless Inference

Uses meta-llama/Llama-3.3-70B-Instruct via W&B's OpenAI-compatible inference API:

  • No Anthropic key needed β€” uses W&B credits ($100 included)
  • Full prompt/completion traces logged automatically by Weave
  • Shows real LLM latency, tokens, and model responses in the trace

7. Review Metrics Dashboard

Structured metrics logged per call: risk score, severity breakdown, false positive rate, token usage. Enables dashboards and trend analysis in Weave.

Evaluation Metrics

The eval suite tests against 8 known-buggy code samples:

Metric What It Measures
Category Recall Did the swarm find the right type of bug?
Severity Accuracy Did it rate the severity correctly?
Precision What fraction of findings are actually relevant?
Detection Rate Did it find anything at all?

Demo Script (for the hackathon presentation)

  1. Show the architecture β€” explain the 3-phase pipeline (30 sec)
  2. Run review-swarm demo β€” reviews real Daytona.io code with W&B Inference, watch the orchestration tree with findings (60 sec)
  3. Open Weave β€” show the trace graph with all agents visible as named nested operations (30 sec)
  4. Show Prompts in Weave β€” demonstrate editing a prompt in the UI (15 sec)
  5. Run review-swarm evaluate β€” show precision/recall metrics in the Evals tab (30 sec)
  6. Show Quality Scorer β€” point to the inline scoring op and how Monitors would pick it up (15 sec)

Key talking points:

  • "4 specialist agents run in parallel β€” security, performance, logic, style"
  • "Then they DEBATE each other β€” challenge false positives adversarially"
  • "The lead agent consolidates after debate, only upheld findings survive"
  • "Every step is traced in Weave β€” you can see the full orchestration graph"
  • "Prompts are versioned in Weave β€” edit them in the UI, no code deploy needed"
  • "We use W&B Serverless Inference β€” real LLM calls with Llama 3.3 70B"
  • "Evaluation suite with published Dataset and 6 quality metrics"
  • "Quality scorer runs inline for continuous monitoring via Weave Monitors"

Tech Stack

  • W&B Weave β€” tracing, evaluations, prompt management, model versioning, monitors
  • W&B Serverless Inference β€” hosted LLM calls (Llama 3.3 70B), uses W&B credits
  • Anthropic Claude (claude-sonnet-4-20250514) β€” optional alternative backend
  • Python asyncio β€” parallel agent execution
  • Rich β€” beautiful terminal UI for the demo
  • Pydantic β€” structured data models for findings
  • OpenAI SDK β€” W&B Inference client (OpenAI-compatible API)

Project Structure

review_swarm/
β”œβ”€β”€ __init__.py         # Package init
β”œβ”€β”€ models.py           # Pydantic data models (Finding, Review, etc.)
β”œβ”€β”€ prompts.py          # System prompts for each specialist agent
β”œβ”€β”€ agents.py           # Core orchestration engine (parallel β†’ debate β†’ lead)
β”œβ”€β”€ weave_prompts.py    # Weave prompt publishing and loading
β”œβ”€β”€ mock_responses.py   # Mock responses for simulate mode
β”œβ”€β”€ demos.py            # Curated real-world diffs from Daytona.io
β”œβ”€β”€ github.py           # GitHub PR diff fetcher (any public repo)
β”œβ”€β”€ evaluation.py       # Eval harness with Weave EvaluationLogger + Dataset
β”œβ”€β”€ ui.py               # Rich terminal rendering
└── cli.py              # Click CLI entry point

Judging Criteria Alignment

Criterion How We Address It
Agent Orchestration 4 parallel specialists β†’ adversarial debate β†’ lead consolidation. Clear multi-agent handoffs with named trace nodes.
Utility Solves real code review β€” finds security vulns, perf issues, logic bugs, style problems on real OSS code.
Technical Execution Async parallel execution, multi-backend (W&B Inference + Anthropic + simulate), structured eval with 6 metrics.
Creativity Adversarial debate between agents is novel β€” they challenge each other's false positives, reducing noise.
Sponsor Usage 7 W&B features: Weave Tracing, Prompt Management, Model Versioning, Evaluations, Quality Scorer, W&B Inference, Metrics Dashboard.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages