Skip to content

Repository files navigation

loom

Agentic engineering for the rest of us.
You design the system. Loom builds it, checks every piece, and learns from its mistakes.

Loop engineering is sold as the whole of software development: state a high-level goal, and an agent loops until finished software comes out. It does get there, by brute force, paying in tokens for every wrong turn and rediscovering what an engineer on the team already knew.

Loom is for engineers who would rather spend their own judgment than someone else's compute. You discuss, theorise, test, explore, and plan the system, and put your experience into that plan. Loom implements it with Claude Code (and optionally Codex) sessions running in parallel across isolated git worktrees. It checks every stage's work itself, then checks all of it together in a final integration-verify stage. Along the way it records its own mistakes and distils them into a knowledge base the next plan starts from. You can watch from a terminal dashboard or a browser, steer any session, or take the keyboard at any time.

A loom run in the web dashboard: the plan as a dependency graph, a stage's detail dialog, its live terminal with control taken over, the ledger, the settings, and a stage completing and freeing the stages that depend on it

A real run in loom status --web, sped up: watch a stage's session, take the keyboard, and see a finished stage merge and start the ones that depend on it.

Why Loom

  • Your expertise where it counts, tokens everywhere else — you do the thinking and put it into the plan. Loom executes it unattended and comes back to you only when a stage needs a person; you can monitor, steer, or take over whenever you like. (Human Expertise Where It Matters)
  • It learns your project — every session records what it got wrong and what it decided. Each plan ends by distilling that into a knowledge base in your repo, so every plan starts from what the earlier ones learned. (It Learns From Every Plan)
  • Parallel by default — stages form a dependency graph, everything independent runs at once in its own git worktree, and each stage merges back as it finishes. (Parallel execution and progressive merge)
  • Done means verified — loom runs each stage's acceptance criteria itself and inspects the tree for stubs, unwired code, and code nothing calls. The agent's own account of its work does not count. Once every stage has merged, an integration-verify stage reviews and tests the combined work as one system. (Verification Model)
  • Rules enforced by hooks — commit discipline, worktree boundaries, and subagent limits fire from shell hooks, so they hold whatever the model intends. (Deterministic guardrails)
  • Contained by default — sessions run in a filesystem and network sandbox and cannot write loom's own state, hooks, or config. (Sandbox Configuration)
  • Expensive models only where judgment is needed — the orchestrator plans and verifies; implementation goes to the cheapest subagent that can do the piece, Claude or Codex. (Model Allocation)
  • Survives crashes and context limits — all state is plain files. The daemon spots dead and hung sessions and retries them, and a handoff is written before a session runs out of context. (Crash recovery and liveness)
  • Watch it live, steer when you want — loom status --live is a dashboard in your terminal. loom status --web is the same run in a browser: the plan as a dependency graph, a ledger of every stage, and terminals you can watch read-only or take over. (Primary Commands, Web Dashboard)

Quick Start

1. Install

curl -fsSL https://raw.githubusercontent.com/cosmix/loom/main/install.sh | bash

You need git, the claude CLI, and jq. Signed binaries cover Linux x86_64 and macOS Apple Silicon; every other platform builds from source — see Installation.

2. Write a plan

Plans are how loom knows what to build. Open Claude Code in your target project and use the /loom-plan-writer skill to create one:

cd /path/to/project
claude  # start Claude Code CLI

Inside the Claude Code session:

  1. Load the plan-writing skill by typing /loom-plan-writer (or just mention that you're writing a loom plan).
  2. Describe what you want to build and discuss with Claude, as you would in plan mode.
  3. When done, Claude will write the plan to doc/plans/PLAN-<name>.md.

Validate the draft before you spend anything on it:

loom plan verify doc/plans/PLAN-<name>.md

Then read it yourself and check that it captured your intent in enough detail. For longer plans, loom pressure doc/plans/PLAN-<name>.md has two model families (Claude and Codex) review it adversarially and harden it. Without the codex CLI, run /pressure doc/plans/PLAN-<name>.md inside a Claude Code session for a Claude-only pressure test.

3. Run it

loom init doc/plans/PLAN-<name>.md
loom run
loom status --live
loom stop

loom init parses the plan, creates stage state, and installs the project hook wiring. loom run starts the daemon and orchestrator. loom status --live follows the run in your terminal, and loom status --web puts it in a browser. Add --terminals (tmux backend only) to watch or take over the Claude Code sessions from the browser. --host <addr> binds the dashboard to an address other than 127.0.0.1; read Web Dashboard Remote Access first, since a non-loopback bind serves plain HTTP.

loom stop stops the daemon, and the orchestrator with it.

Next: the two ideas loom is built on, human expertise where it matters and learning from every plan. Then How It Works and the Feature Tour, a map of everything below.

Human Expertise Where It Matters

Loop engineering presents itself as the end of the story: give an agent a high-level goal, let it try, fail, read the error, and try again until the checks pass, and finished software comes out. It does produce software, at a high price and with little elegance. Every wrong turn is paid for in tokens, and much of what those tokens buy is the rediscovery of something an engineer on the team already knew. A frontier lab can afford that. Most organisations cannot.

Human engineers have a lot to offer: experience, judgment, and a working theory of the system that no loop recovers cheaply. Loom puts that first and spends tokens on the rest:

  • Up front, in the plan. You discuss the design, form theories and test them, explore the code, and decide what gets built, how it splits into stages, and what proves each stage is done. loom plan verify and loom pressure find the weak spots before anything is spent on execution. This is where an hour of your time saves the most tokens.
  • During the run, only when needed. A stage that needs a person says so: WaitingForInput when an agent asks a question, NeedsHumanReview when it is escalated, a criteria dispute when an agent believes a check is wrong. Everything else runs unattended.
  • Whenever you choose to look. The terminal dashboard (loom status --live) and the web dashboard (loom status --web) show every stage, what it is doing, which models it is running, and how much context it has used. You can open any session, from tmux with loom attach or from a terminal in the browser, and watch it or take the keyboard and steer it.

What you hand over comes back checked. Loom verifies every stage's work before it merges, then the plan's integration-verify stage reviews and tests everything together, so stages that each pass on their own still have to work as one system.

Steering is also teaching. Agents record a human correction as a memory entry before they do anything else, so what you tell one session is distilled into the knowledge base and reaches every later one.

Details: Human-in-the-loop where it matters, Web Dashboard, Terminal Backends.

It Learns From Every Plan

Most agent setups start every session from zero. The same architecture is rediscovered, the same wrong turn is taken, and the same correction is given again. Loom keeps what each session learns and gives it to the next one.

  1. Every session keeps a journal. As they work, agents record mistakes with a prevention rule, decisions with their rationale, and whatever surprised them about the codebase (loom memory note, decision, change, question).
  2. Every plan ends by distilling it. The knowledge-distill stage reads every stage's journal and curates it into doc/loom/knowledge/: architecture, entry points, patterns, conventions, stack, concerns, and a mistakes file where each entry says what happened, why, how to prevent it, and how it was fixed.
  3. Every later session starts from it. A stage's signal carries a Knowledge Brief: the sections retrieval judged relevant to that stage, already quoted. Agents pull anything more with loom knowledge context --query, and look code up in a source graph (loom map) before they open files.
  4. Wrong knowledge gets corrected. When the code contradicts the knowledge base, the code wins. The agent records the contradiction and the next distillation applies it.

The knowledge base is plain markdown in your repository, so it is reviewed, versioned, and shared with your team like any other file. It is tiered, so it can keep growing while each session loads only the part it needs.

The effect compounds. The first plan on a project pays for discovery. After a few plans the knowledge base holds the project's structure, its conventions, and the mistakes already made on it, and later plans make far fewer of them: fewer failed verifications, fewer retries, fewer tokens.

Details: Knowledge System.

How It Works

flowchart LR
    A[Plan] --> B[loom init]
    B --> C[Stages run in parallel worktrees]
    C --> D[Loom verifies]
    D --> E[Merge]
    E --> I[Integration verify]
    I --> F[Distill knowledge]
    F -. read first by the next plan .-> A
Loading

A plan is a markdown file holding a list of stages, each with its dependencies and the commands that prove it is done.

  1. Init. loom init <plan-path> parses the plan and creates the state for every stage under .loom/work/.
  2. Schedule. loom run starts a daemon and an orchestrator. Every stage whose dependencies are met gets its own worktree (.worktrees/<stage-id>, branch loom/<stage-id>) and its own Claude Code session, briefed by a signal file: the assignment, the relevant knowledge, and the memory of earlier sessions.
  3. Work. The session's main agent decomposes the stage, delegates implementation to cheaper subagents, and records what it learns with loom memory.
  4. Verify. loom stage complete runs the stage's acceptance criteria and the goal-backward checks. A failure leaves the stage Executing, and the agent has to fix the work and try again.
  5. Merge. A verified stage merges back to the target branch, which frees the stages that depend on it. A real conflict gets a dedicated resolution session.
  6. Integrate. Once the implementation stages have merged, an integration-verify stage reviews the combined work and runs the full suite against it.
  7. Distill. The plan's final knowledge-distill stage curates every stage's memory into doc/loom/knowledge/, which the next plan's sessions read first.

You follow along with loom status --live or loom status --web, and step in only when a stage asks for a person: recover, verify, merge, or retry stages as needed with the stage commands.

Stage Lifecycle

WaitingForDeps → Queued → Executing → Completed

Everything else is an explicit, inspectable outcome rather than a hang:

State Meaning
Blocked A before_stage check or an explicit block stopped the stage
NeedsHandoff Context ceiling reached; a handoff was written
WaitingForInput The agent asked a question (raised automatically by the AskUser hooks)
MergeConflict Auto-merge hit a real conflict; a resolution session is spawned
MergeBlocked Merge cannot proceed (e.g. another merge is in progress)
CompletedWithFailures Work finished but acceptance did not pass
NeedsHumanReview Escalated to a person
NeedsAdjudication A disputed acceptance criterion is awaiting a verdict
Skipped Explicitly skipped

Feature Tour

Everything loom does, one line each, grouped by what you are doing at the time. Each line links to the section that documents it.

Plan

  • Draft a plan interactively with the /loom-plan-writer skill. (Write a plan)
  • Validate a plan with no side effects: schema, sandbox limits, and worker-ownership overlaps. (Plan Commands)
  • Harden a plan through adversarial review rounds from two model families. (Primary Commands)
  • Give each stage a type: knowledge, standard, integration-verify, or knowledge-distill, each with its own verification rules. (Stage Type Behavior)
  • Amend a stage's criteria mid-run through an audited path that snapshots the plan and logs the change. (Disputes, Adjudication, and Amendments)

Run

Verify

Learn

  • Journal notes, decisions, and mistakes to a per-stage memory as work happens. (Knowledge System)
  • Distill every stage's memory into permanent, tiered knowledge at the end of a plan. (Knowledge System)
  • Pull a token-budgeted, quoted context pack for one specific question on demand. (Knowledge System)
  • Query a source graph of the code's own symbols for outlines, impact, and callers. (Knowledge / Memory)

Observe

  • Follow the run in a live terminal dashboard, one row per stage with its state, models, activity, context use, and merge status. (Primary Commands)
  • Watch the same ledger and a dependency graph in the browser, with settings and remote access. (Web Dashboard)
  • Take control of a stage's live session from a browser terminal. (Web Dashboard Terminals)
  • See Claude and Codex quota bars with reset countdowns wherever the ledger renders. (Web Dashboard)
  • Report what a run actually spent per provider, and compare a policy change against paired runs. (Other Commands)
  • List and attach to any live session by stage or session ID. (Other Commands)

Contain

  • Confine each session's filesystem reads/writes and network domains by plan and per-stage rules. (Sandbox Configuration)
  • Rebuild a minimal, allowlisted environment for every command loom runs from your plan. (Command Confinement)
  • Deny a session write access to loom's own state, hooks, and config, relaying the rare legitimate exception through the daemon. (Session State Confinement)
  • Enforce commit discipline, worktree boundaries, and subagent limits from shell hooks. (Deterministic guardrails)
  • Run stages in auto permission mode by default and tighten it per plan or per stage; bypass-permissions is rejected. (Permission Mode)
  • Enable Claude Code remote control on spawned sessions automatically when its prerequisites are met. (Remote Control)

Operate

  • Hold, release, skip, retry, reset, or hand a stuck stage to a human. (Stage Commands)
  • Diagnose and repair a broken install or missing hook wiring. (Other Commands)
  • Clear worktrees, sessions, or state selectively without touching the rest. (Other Commands)
  • Check for and install a newer loom release. (Other Commands)
  • Install shell tab-completion for bash, zsh, or fish. (Shell Completions)

Spend

  • Run each stage's main agent at its stage type's configured model and effort. (Model Allocation)
  • Override the model or effort for one stage explicitly, without touching the rest of the plan. (Stage Fields)
  • Keep each signal's prefix byte-identical across sessions, so the large doctrine block is a cache hit. (Cost control by construction)
  • Inject at most 5 matched skills per stage out of 73 installed; 10 core skills are always loaded and the other 63 load on demand. (Cost control by construction)

Contents

What Loom Solves

Autonomous agent work fails in a small number of predictable ways. Loom answers each with a mechanism, not a paragraph of prompt.

Failure mode What actually happens Loom's answer
False completion Tests were never run, the module was written but never imported, the fix is a TODO Loom runs the acceptance criteria itself, then checks artifacts for stubs, wiring for real integration, and dead-code patterns for orphaned work. The bypass flags need a token the agent cannot read.
Instruction drift Rules decay the moment they scroll out of attention Shell hooks enforce the rules that matter deterministically — commit discipline, staging scope, worktree boundaries, subagent limits — outside the model's control.
Amnesia Every session rediscovers the same architecture and repeats the same mistakes A per-stage memory journal feeds a distillation stage that curates permanent, tiered knowledge; later sessions read it before touching code.
Cost scaling with tokens, not value Expensive models doing cheap work; re-reading everything, every time Judgment stays on an orchestrator; bulk implementation is delegated to cheap subagents. Signals are laid out for KV-cache reuse and knowledge is tiered, so agents load only what they need.
Context exhaustion The session degrades into an expensive compaction loop Context budgets are monitored per stage; a handoff is written before compaction and the resumed session is re-anchored to its assignment.
Lost runs A crashed or hung session takes the work with it All state is files under .loom/work/. The daemon detects dead and hung sessions, classifies the failure, and retries or escalates.
Serialization Multi-stage work runs one-at-a-time, or collides on the same files A dependency DAG schedules independent stages concurrently in separate worktrees, with progressive auto-merge and dedicated conflict-resolution sessions.

Key Capabilities

Deterministic guardrails

The rules that matter are not left to the model. Loom installs Claude Code hooks, a Codex-native hook subset, and a git pre-commit hook that fire regardless of what an agent intends:

  • commit-guard.sh blocks a session from ending with uncommitted work or a stage still Executing
  • git-add-guard.sh blocks git add -A / git add .; git-pre-commit-hook.sh blocks commits containing .loom/work or .worktrees
  • worktree-isolation.sh / worktree-file-guard.sh block cross-worktree writes, reads, and path traversal
  • commit-filter.sh blocks subagent git operations (a subagent commit loses the main agent's work) and blocks AI attribution in commit messages
  • subagent-verify-guard.sh blocks subagents from running project-wide build/test/lint suites, so verification stays with the one agent that can see the whole tree — with integration-verify stages carved out, and no opt-out environment variable
  • pre-compact.sh blocks compaction, writes a handoff, then allows it; session-start.sh re-anchors the resumed agent to its signal file
  • plans-path-guard.sh keeps plans in doc/plans/ where loom and git can see them

Subagent detection is a live process-tree ancestry check, not a PPID comparison. See Verification Is the Main Agent's Job.

Verification that outlives the agent's opinion

loom stage complete is not a self-report. Loom executes the stage's acceptance criteria in-process and refuses completion on failure, leaving the stage Executing so the agent must fix and retry. On top of that, goal-backward verification asks whether the outcome exists:

  • artifacts — files exist and contain real implementation (stub detection rejects TODO, FIXME, unimplemented!, todo!, pass, NotImplementedError)
  • wiring — regex proof that new code is actually referenced: module registered, route mounted, component rendered
  • wiring_tests — runtime commands proving the integration behaves
  • dead_code_check — command output patterns catching code that exists but is never called
  • before_stage / after_stage — pre-spawn and post-acceptance gates; a failed pre-check blocks the stage before a session is even spawned

The escape hatches (--no-verify, --force-unsafe, --assume-merged) require a one-time operator proof bound to the project, stage, action, and exact flag set. The operator supplies the daemon secret only while minting the proof; the target command cannot fetch that credential for its caller or reuse the proof for another action.

Knowledge capture and distillation

Loom treats what agents learn as a durable artifact with a pipeline behind it, rather than a scratch file.

  1. Capture — during execution, agents record to a per-stage journal: loom memory note (gotchas, mistakes-with-prevention), decision (with rationale), change, question. The journal is injected into the recitation section at the end of the next signal, where model attention is highest.
  2. Distill — a knowledge-distill stage runs at the end of a plan, reads every stage memory, and curates it into permanent knowledge — mistakes rewritten as actionable prevention rules, decisions with their rationale, reusable patterns and conventions.
  3. Retrieve — the result is a tiered base under doc/loom/knowledge/: a generated INDEX.md, seven tier-1 summaries, and tier-2 topic files. Inside a stage, the per-stage Knowledge Brief comes first. Otherwise, agents read INDEX.md for orientation, then the tier-1 summary for their area, then only the topics they touch; a specific question is pulled with loom knowledge context --query, which returns the matching sections quoted — so the base can grow without every session paying to load it.

Knowledge lives in doc/loom/knowledge/; agents write it through loom, and loom knowledge sync rebuilds the derived retrieval artifacts after the tree changes, including the one-time flat-to-hierarchical upgrade. Details: Knowledge System.

Cost control by construction

Loom's savings come from delegation, not downgrade:

  • Orchestration's model and effort come from the stage type's default — every stage's main agent plans, decomposes, verifies, and commits, the judgment-heavy work that is worst to economize on — and both are overridable, in [models] in either config file or per stage; see Model Allocation.
  • Delegation is a cost decision, tokens times model tier. A stage's main agent makes a change itself only when it is at most 20 lines in at most 2 files it has already read, needs no further exploration, and one command proves it; a Fable main session delegates even those. Everything larger goes to the cheapest subagent that can do it, spawned by agent type so the choice is explicit rather than inherited: Haiku for mechanical edits such as a rename; Fable only for visual/UI design, a bug that survived a delegated fix attempt, and extremely challenging algorithmic design (no agent type pins it — the model override is stated explicitly at spawn); Opus for mainstream architecture and algorithm implementation; Sonnet or Codex GPT-5.6 Terra for common implementation and integration tests; Codex GPT-6 Luna for boilerplate, scaffolding, and simple unit tests. The codex tiers are licensed only on stages listing codex in implementers, and additionally require the codex CLI and its plugin to be installed — when either is missing, loom run prints an advisory warning at startup (it never aborts) and terra-/luna-tier work falls back to Sonnet.
  • Signals are built for cache reuse. Each signal is a four-section layout with a per-stage-type stable prefix that is byte-identical across sessions, so the large doctrine block is a cache hit rather than a re-read.
  • The orchestrator's rulebook loads when it is needed. Delegation, briefs, file ownership, waiting on subagents and commit timing live in the loom-orchestration core skill, which a session loads before it fans out; the installed CLAUDE.md keeps the hard stops and a pointer, under a 20 KB ceiling. spawn-guard.sh prepends the subagent preamble to every typed spawn, so an orchestrator no longer pastes it.
  • Context budgets prevent compaction, which is the expensive failure: an uncached re-read that costs more and produces worse work.
  • Tiered knowledge and a skill index keep the working set small — at most 5 matched skills are injected per stage, out of 73 installed.
  • Waits and repeat reads are settled by receipts, not by polling. An orchestrator waits on a backgrounded Codex forward by its exact receipt (loom subagents wait --receipt <id>), and repeated loom subagents list polling is counted by the poll guard. A repeated file read is warned or denied only when a transcript receipt proves the earlier result was delivered.
  • Consumption is measured, not assumed. loom usage reports Claude and Codex separately from provider-native telemetry, and loom usage --compare judges a candidate policy offline against paired runs. A token-proxy gain alone never counts as a subscription saving, and any quality or latency regression rejects the candidate; see the evaluation protocol.

Per-stage model, reasoning_effort, and ultracode fields let you override any of this explicitly.

Parallel execution and progressive merge

Stages form a dependency DAG; everything independent runs at once, each in its own worktree (.worktrees/<stage-id>, branch loom/<stage-id>). Completed stages merge back progressively under a file lock, and a real conflict spawns a dedicated resolution session rather than stalling the run.

Crash recovery and liveness

All orchestration state is plain files in .loom/work/, so nothing is lost when a process dies. The daemon polls every 5s, tracks PID liveness and per-session heartbeats, flags hung sessions after 300s, and classifies failures across ten types into retryable (exponential backoff) and needs-diagnosis. Tool-call telemetry drives a stuck-session signal when a session's recent calls are overwhelmingly failures. Orphaned sessions are recovered on daemon restart.

A stage completes only on an authenticated evidence record that survives handoffs, so a session that finished its work but died before the daemon observed it does not loop between retries. Each session records why it exited — completed, crashed, context ceiling, stalled, operator stop, criteria blocked, or replaced — and loom status, loom status --live, and the web dashboard surface a completion-pending or blocked state with that reason ahead of generic activity.

Sandboxing and plan hardening

Plan-level defaults and per-stage overrides control filesystem reads/writes, network domains, and permission mode for the agent session, and commands loom runs from your plan get a rebuilt, allowlisted environment so they cannot read ambient credentials (Sandbox Configuration). Before you spend anything, loom plan verify validates a plan with no side effects — running the same sandbox validation that would otherwise only fail at loom init — and loom pressure hardens it through adversarial review rounds run by two different model families. A sandboxed session's own writes are confined the same way: it cannot touch loom's state, hooks, or config directly, and the handful of legitimate exceptions go through a one-shot request the daemon applies (Session State Confinement).

Human-in-the-loop where it matters

Thirteen stage states make "needs a person" an explicit outcome rather than a hang: WaitingForInput (raised automatically when an agent asks a question, or when a v2 contract freeze is refused), NeedsHumanReview, Blocked, MergeConflict. Operators get loom stage hold/release/skip/retry/human-review, and an agent that believes a criterion is wrong can escalate with loom stage dispute-criteria instead of quietly weakening it.

Installation

Loom is under active development. Signed binaries are published for Linux x86_64 and macOS Apple Silicon; every other platform builds from source with the Rust toolchain installed.

Prerequisites

Tool Needed for Required?
Rust toolchain (cargo) building the loom binary only when building from source
git worktrees, merges, crash reports yes
claude (Claude Code CLI) every orchestrated session yes
jq every loom hook parses the Claude Code hook payload with it yes; install.sh and loom run refuse to proceed without it, loom repair reports it
rg (ripgrep) and fd the installed CLAUDE.md steers agents to these over grep/find recommended; install.sh and loom run warn when missing
tmux the tmux terminal backend optional
sccache shared dependency compiles across stage worktrees optional
codex CLI the codex implementer lane and loom pressure optional

Install script

curl -fsSL https://raw.githubusercontent.com/cosmix/loom/main/install.sh | bash

This downloads the signed release binary for your platform to ~/.local/bin/loom, then installs Loom's agents, skills, commands, hooks, and orchestration rules into the default asset roots, ~/.claude/ and ~/.codex/, from the assets embedded in that binary. Codex asks you to review new or changed non-managed hooks with /hooks before they run.

Custom asset roots

Set either variable to install Loom assets outside the default root:

Variable Default asset root
LOOM_CLAUDECODE_INSTALL_DIR ~/.claude
LOOM_CODEX_INSTALL_DIR ~/.codex

A nonempty variable replaces its matching default; an empty variable behaves as unset. Every install that passes neither --claude-dir nor --codex-dir records the roots it used in ~/.config/loom/install-roots.toml, and later installs and loom update reuse them, so the variables are needed only when choosing or changing a root. For each root, --claude-dir or --codex-dir takes precedence over its variable, then the recorded root, then the default. An install given either flag does not change the record.

Use absolute paths such as $HOME/.local/share/loom/claude:

export LOOM_CLAUDECODE_INSTALL_DIR="$HOME/.local/share/loom/claude"
export LOOM_CODEX_INSTALL_DIR="$HOME/.local/share/loom/codex"
curl -fsSL https://raw.githubusercontent.com/cosmix/loom/main/install.sh | bash

A later loom update in a shell without these variables refreshes the same roots.

This selects Loom's installation destinations only. It does not configure Claude Code or Codex, or relocate Loom's other runtime discovery paths.

Build from source

Required on Linux ARM64, and what you want when working on loom itself:

git clone https://github.com/cosmix/loom.git
cd loom
bash ./dev-install.sh

dev-install.sh builds the release binary (cargo build --release) and runs install.sh, which installs loom-* prefixed agents and skills (non-destructively, preserving user customizations), hooks, and configuration into the default ~/.claude/ root and the CLI binary to ~/.local/bin/loom. Orchestration rules are written directly to the default ~/.claude/CLAUDE.md (existing file is backed up). It inherits the custom asset-root variables above when they are exported.

install.sh takes an optional --skills core|all flag (default core): core installs a small set of always-loaded core skills to ~/.claude/skills/ and catalogs the rest under ~/.claude/loom-skill-catalog/, loaded on demand; all installs every loom skill directly to ~/.claude/skills/.

Project hook wiring

loom init installs and configures project hook wiring automatically. For an existing repo that is missing Claude Code hook setup, run loom repair --fix.

Platform Support

  • Linux x86_64: primary development and full CI test runs; signed release binary
  • macOS (Apple Silicon): supported for build/terminal integration, CI does build-only verification; signed release binary
  • macOS (Intel): builds from source; no release binary is published
  • Windows via WSL2: supported — WSL2 runs the Linux x86_64 binary unmodified. Use the tmux backend, since a stock WSL install has no GUI terminal emulator for the native backend to find — see Running under WSL. Native Windows, outside WSL, is unsupported.
  • Linux ARM64: builds from source; no release binary is published yet
  • Headless (SSH, no terminal emulator): supported via the tmux backend — see Terminal Backends

What Gets Installed

Default location Contents
~/.claude/agents/loom-*.md 5 specialized subagents (per-item, non-destructive)
~/.claude/skills/loom-*/ 10 core domain knowledge modules, always loaded (per-item, non-destructive)
~/.claude/loom-skill-catalog/loom-*/ 63 more domain knowledge modules, loaded on demand (--skills core, the default)
~/.claude/commands/*.md Loom slash commands (/pressure, /address, /distill)
~/.claude/hooks/loom/ Embedded lifecycle and guardrail hooks + shared libraries
~/.claude/CLAUDE.md Orchestration rules
~/.codex/skills/pressure/ Codex pressure-testing skill ($pressure)
~/.codex/hooks/loom/ Loom hook assets used by Codex-native registrations
~/.codex/hooks.json Non-destructively merged Codex hook registrations
~/.codex/AGENTS.md Codex navigation and execution doctrine
~/.local/bin/loom Loom CLI

CLI Reference

Primary Commands

loom init <plan-path> [--clean] [--backend native|tmux]
loom run [--manual] [--max-parallel N] [--foreground] [--watch] [--no-merge] [--backend native|tmux]
loom status [--live] [--compact] [--verbose] [--web [PORT] [--host HOST] [--terminals]]
loom stop
loom resume <stage-id>
loom check <stage-id> [--suggest] [--no-cache]
loom pressure <plan-path> [--rounds N] [--claude-model M] [--claude-effort E] [--codex-model M] [--codex-effort E] [--address-model M] [--address-effort E] [--dry-run]

When loom run stops — every stage settled, or interrupted by loom stop — it prints stages by outcome: Completed; Failed (Blocked, MergeConflict, CompletedWithFailures, MergeBlocked, NeedsHumanReview — terminal, needing intervention); Unfinished (WaitingForDeps, Queued, Executing, WaitingForInput, NeedsAdjudication — still in progress when the run stopped); and Needs Handoff. The run succeeds only when Failed, Unfinished, and Needs Handoff are all empty.

loom pressure hardens a plan before you run it by combining two external agents over --rounds rounds (default 2). Each round runs both pressure-tests in parallel: Claude /pressure edits the plan in place in the foreground (you watch it live), while Codex $pressure writes an independent review next to it (codex-<plan>.md) in the background (its output is captured to a temp log to keep the terminal clean). Once both finish, Claude /address folds the review back in. Claude stays interactive and auto-closes when done; Codex runs from the repo root. Requires both the claude and codex CLIs on PATH. --dry-run prints the exact commands without spawning anything.

Each of the three steps spawns with an independently selectable model and reasoning effort: --claude-model/--claude-effort for /pressure (model accepts haiku, sonnet, opus, or fable; effort accepts low, medium, high, xhigh, or max), --address-model/--address-effort for /address (same value sets), and --codex-model/--codex-effort for $pressure (model accepts gpt-6-astra, gpt-6.1-sol, gpt-6-sol, gpt-5.6-terra, or gpt-6-luna; effort accepts low, medium, high, or xhigh — no max, that value is Claude-only). Absent a flag, each key falls back independently: the project config's .loom/work/config.toml [pressure] section first (if it sets that key), then ~/.loom/config.toml (loom config -k pressure.claude_model <value>), then its built-in default. [pressure] resolves per key, so a project section that sets only claude_model still lets codex_effort fall through to your user config.

When the codex model is unavailable to the account (codex rejects it as not supported or not found), loom pressure retries with the same family's nearest newer version first and then the nearest older one, never another tier; the default gpt-6.1-sol falls back to gpt-6-sol. The fallback notice prints after the Claude session ends. When every model in the family is unavailable the run fails and shows the codex error.

Key Default
pressure.claude_model opus
pressure.claude_effort xhigh
pressure.codex_model gpt-6.1-sol
pressure.codex_effort xhigh
pressure.address_model opus
pressure.address_effort high

loom status --live renders a live ledger dashboard: one row per stage across eight columns (STATE, STAGE, DEPENDS ON, MODELS, ACTIVITY, CONTEXT, TIME, MERGE). MODELS lists the orchestrator's own model first, then the models any subagents it spawned ran on. Columns drop in priority order as the terminal narrows; below a 64x16 (columns x rows) terminal a notice replaces the dashboard entirely. Press ? to toggle a legend overlay explaining every state icon. When Claude or Codex quota data is available, a footer line above the legend shows a percent bar and reset countdown for each.

loom status --web [PORT] [--host HOST] [--terminals] serves the same ledger in the browser, alongside a dependency-graph view of the plan — see Web Dashboard.

Plan Commands

loom plan verify <plan-path> [--strict] [--json] [--no-color]

loom plan verify validates a plan file without touching .loom/work/ or requiring a git repo. It runs the same fatal validation as loom init (schema errors, unknown or retired fields at every nested policy layer, duplicate IDs, unknown dependencies, path safety) plus advisory warnings (structural issues, a stage without a summary, missing knowledge-bootstrap stage, sandbox gaps). A retired top-level truths block is rejected; move behavioral commands to acceptance. It additionally rejects requirements a sandboxed stage cannot satisfy — an allow_write grant under /tmp or missing on the host, a TMPDIR=<absolute path> override, a hardcoded /tmp/ path, and a mkdir, touch or output redirect aimed outside the worktree in acceptance, setup, wiring_tests, before_stage or after_stage commands — because the stage sandbox already provides a writable $TMPDIR. It rejects commands that cannot mean what they say inside a sandbox — an exit status masked by a final || true, || :, || exit 0 or ; true, HOME assigned from a variable or substitution, and a bare mktemp -d — and warns on network tools, reads of doc/plans/, rg -r, and hardcoded /tmp/ paths. A single rg or grep acceptance criterion that already passes at HEAD earns a baseline warning, since it cannot tell a stage that did its work from one that did nothing. A stage's skills: list is checked against the skill index: an empty, duplicate or unknown name is an error. It also warns when a stage description's worker/files-owned table claims overlapping paths between workers, or a claim outside the stage's declared files. Exits 0 on success, non-zero on fatal errors; --strict promotes warnings to errors.

Stage Commands

loom stage complete <stage-id> [--session <id>] [--no-verify] [--force-unsafe --assume-merged] [--no-cache]
loom stage block <stage-id> <reason>
loom stage reset <stage-id> [--hard] [--kill-session]
loom stage waiting <stage-id>
loom stage resume <stage-id>
loom stage hold <stage-id>
loom stage release <stage-id>
loom stage skip <stage-id> [--reason <text>]
loom stage retry <stage-id> [--force] [--context <message>]
loom stage merge [stage-id] [--resolved]
loom stage human-review <stage-id> [--approve|--force-complete|--reject <reason>]
loom stage dispute-criteria <stage-id> --criterion-index N --reason <text> [--evidence-commit <sha>] [--failure-output <path>]
loom stage dispute-findings <stage-id> --finding <id>... --reason <text>       # Plan version 2: challenge open review findings
loom stage dispute-contract <stage-id> --contract <id> --reason <text>        # Plan version 2: challenge a frozen contract
loom stage dispute-integrity <stage-id> --event <id>... --reason <text>       # Plan version 2: challenge a test-integrity event
loom stage adjudicate --stage <stage-id> --dispute <n> --verdict-file <path>
loom stage contracts freeze|show|restore <stage-id> [--contract <id>]         # Plan version 2: freeze (contract session), inspect, or restore frozen contract files
loom stage review status|integrity <stage-id>                                 # Plan version 2: review rounds and open findings; test-integrity events
loom stage admin-proof [stage-id] [--daemon-stop] [--no-verify] [--force-unsafe] [--assume-merged]
loom stage amend <stage-id> --field acceptance|wiring|wiring-tests|contracts --op replace|insert|delete --index N [--value <yaml>] [--reason <text>]

loom stage dispute-criteria is the sanctioned way for an agent to challenge a criterion it believes is wrong or impossible, instead of quietly weakening it. The daemon writes request.md and moves the stage to NeedsAdjudication; the verdict is daemon-written and never authored by the agent. loom stage adjudicate records that verdict from the adjudication session; loom stage admin-proof mints an operator's proof for a trusted broker; loom stage amend is the audited, operator-facing way to edit a stage's acceptance, wiring, or wiring_tests array directly. See Disputes, Adjudication, and Amendments.

Stage Outputs

loom stage output set <stage-id> <key> <value> [--description <text>]
loom stage output get <stage-id> <key>
loom stage output list <stage-id>
loom stage output remove <stage-id> <key>

Knowledge / Memory

loom map [--outline <path>] [--find-all <symbol>] [--impact <symbol|path>] [--callers <symbol>] [--callees <symbol>] [--json]
                                                                # Query the derived source graph: file outlines, symbol lookup, impact/caller/callee analysis
loom knowledge context --query <text> [--stage <id>] [--budget-tokens <n>] [--explain] [--json]  # Token-budgeted context pack for a question
loom knowledge update <file> [content]                        # Append a section to a tier-1 file or tier-2 topic (<category>/<slug>)
loom knowledge replace-section <file> <heading> [content]      # Rewrite one section's body in place, at whatever level it's found
loom knowledge delete-section <file> <heading>                # Remove a section and its nested subsections (a heading rename is delete-section then update)
loom knowledge annotate <target> [--state <s>] [--section <heading>] [--source <path>]... [--verified <rev|HEAD>] [--alias <name>]... [--blurb <text>]
                                                                # Set lifecycle state (whole file, or one `##` section), evidence sources, verified revision, aliases, or a topic's blurb
loom knowledge telemetry [--stage <id>] [--json]                # Summarise recorded context delivery, prompt briefs, abstentions and pulls
loom knowledge sync [--structural-only] [--json]              # Rebuild derived retrieval artifacts after editing knowledge
loom knowledge check [--strict] [--strict-evidence] [--baseline <file>] [--json]  # Report knowledge-base diagnostics (read-only; never opens the context store); --strict-evidence also fails on changed or unassessable declared sources; --baseline makes --strict fail only on structural issues the file does not record
loom knowledge check --write-baseline <file>                  # Record every current structural issue in <file> and exit 0, to adopt the size limits on a tree that cannot clear them yet

loom memory note <text> [--stage <id>]
loom memory decision <text> [--context <why>] [--stage <id>]
loom memory change <text> [--stage <id>]
loom memory question <text> [--stage <id>]
loom memory query <search> [--stage <id>]
loom memory list [--stage <id>] [--entry-type <type>]
loom memory show [--stage <id>] [--all]
loom memory resolve <event-id> --outcome <promoted|merged|discarded|deferred> [--target <file#heading>] [--reason <text>]
                                                                # Record how a captured note/decision/question was processed
loom memory pending [--stage <id>] [--group] [--strict] [--json]  # List notes, decisions, and questions that have no receipt; --group prints corrections (by target file and heading), mistakes, decisions, other

A plan's .loom/work/ state is archived to <main>/.loom/memory/archive/<plan-id>-<timestamp>/ when the plan completes; loom clean/loom init --clean are what actually delete it.

See Knowledge System for how these fit together.

Other Commands

loom review [--ai-summary]                                                   # Generate a code-review doc from stage memories; --ai-summary uses headless `claude -p`
loom usage [--since <duration|date>] [--until <rfc3339>] [--provider claude|codex|all] [--project <path> | --all] [--stage <id>] [--plan <name>] [--windows 5h|week] [--json]
                                                                             # Report what agent sessions actually consumed, per provider (Claude and Codex are never summed), including peak resident context by scope in a `peaks` section (main vs subagent p50/p90/max, share above 250k and 400k tokens)
loom usage [--claude-root <dir>] [--codex-root <dir>] [--receipts-root <dir>] [--forward-receipts-root <dir>]
                                                                             # Read explicit telemetry roots; a supplied root never falls back
loom usage --compare <artifact.json> [--json]                                # Offline paired evaluation: exit 0 supported, 1 rejected, 2 inconclusive (doc/token-optimization-evaluation.md)
loom subagents list | harvest [--id <id>] [--json]                           # One-shot diagnostics only: list shows liveness, harvest prints a terminal report (optionally one agent's); watch/wait below are the blocking waits
loom subagents watch --worker <kind>:<id> [--worker <kind>:<id> ...] [--session <claude-parent-uuid>] --timeout <secs> [--json]
                                                                             # One owned wait, inside a loom stage only (needs LOOM_STAGE_ID/LOOM_SESSION_ID/LOOM_WORK_DIR): exit 0 all bound workers have fresh correlated success; 2 deadline passed; 3 worker failed/cancelled
                                                                             # kind is claude (spawned agent ID) or codex (--unit-id); exit 4 AlreadyWaiting/Busy (no second monitor)
                                                                             # Exit 5 identity/evidence unknown; --dir or no --worker is rejected with migration guidance
loom subagents wait --receipt <id> [--timeout <secs>] [--json]               # Wait on one exact forwarded Codex job; exit 0 succeeded, 1 failed/canceled, 2 still running/unknown
loom attach [stage-id]                                                       # tmux backend only; omit the id for a tiled overview
loom sessions list
loom sessions kill <session-id...> | --stage <stage-id>
loom handoff [--stage <id>] [--session <id>] [--trigger <type>] [--message <text>]
                                                                             # Capture session state to a handoff file; stage and session default to $LOOM_STAGE_ID/$LOOM_SESSION_ID, --trigger defaults to manual
loom worktree list
loom worktree remove <stage-id>
loom graph
loom project detect [PATH] [--json]                                          # Per package: language kinds, the test runner loom would use (or unsupported), and the matching skills
loom context record-edit --stage <id> --path <path> [--path <path>...]       # Keep a stage's context overlay current
loom hook user-prompt                                                        # UserPromptSubmit entry point; invoked by loom's hooks
loom request status <id> [--session <id>]                                    # Plumbing: has the daemon applied a request relayed through the sandbox? <id> is printed after the originating command
loom skill-index                                                             # Plumbing: rebuild the skill keyword index the skill-trigger hook reads
loom repair [--fix]
loom clean [--all|--worktrees|--sessions|--state]
loom update
loom config [-k <key> [<value>] | --list | --print]                          # Read or write ~/.loom/config.toml; bare in a terminal it opens the settings screen (see Configuration)
loom install-assets [--claude-dir <path>] [--codex-dir <path>] [--skills core|all]  # Install loom's agents, skills, commands, hooks and doctrine files; see Install script below for what --skills core|all installs where
loom completions [<shell>] [--install] [--migrate]

loom handoff writes the document a successor session resumes from, under .loom/work/handoffs/. Loom's pre-compact hook calls it automatically before a compaction, and an agent that reaches its context ceiling calls it explicitly with --trigger ceiling.

loom request status and loom skill-index are plumbing for loom's own hooks and its sandbox relay rather than commands a plan author types; they are listed so hook output that names them is traceable.

Configuration

Loom keeps its settings in two TOML files. A project file overrides the user file, and an explicit per-invocation value (a stage's own model field in the plan, a loom pressure flag) overrides both:

Tier File Applies to
user ~/.loom/config.toml Every workspace on this machine
project <repo>/.loom/work/config.toml This workspace only

Neither file needs to exist: every key has a built-in default. LOOM_HOME relocates the user file ($LOOM_HOME/config.toml).

There are three ways to change a setting:

  1. loom config. Run bare in a terminal it opens a settings screen; with flags it is scriptable. loom config -k <key> prints one key, loom config -k <key> <value> writes it (validated against the key's type and value set), loom config --list prints every key with its value and where it came from, and loom config --print prints the resolved user config as TOML. It reads and writes the user file only. The settings screen edits each row by its type: ↑↓/k/j move, ←→/h/l/space cycle an enum's variants or toggle a bool, Enter opens a text editor for a number or free-text key (and steps a bool/enum forward otherwise, so no keystroke opens a text field on a closed vocabulary), s saves, Esc/q quits, * marks a pending change.
  2. Edit the files. Both files use the same [section] / key = value layout as the table below; project sections may be partial. loom init writes a [context] section into the project file, everything else is opt-in.
  3. The dashboard. loom status --web has a settings dialog that edits both files, one key at a time, with the resolution shown per key (Web Dashboard Settings).

The keys, with their built-in defaults:

Key Default Values Project tier
update.check true true, false no
update.check_interval_hours 24 integer no
terminal.backend native native, tmux whole section
context.ceiling_tokens 800000 integer whole section
pressure.claude_model opus haiku, sonnet, opus, fable per key
pressure.claude_effort xhigh low, medium, high, xhigh, max per key
pressure.codex_model gpt-6.1-sol gpt-6-astra, gpt-6.1-sol, gpt-6-sol, gpt-5.6-terra, gpt-6-luna per key
pressure.codex_effort xhigh low, medium, high, xhigh per key
pressure.address_model opus haiku, sonnet, opus, fable per key
pressure.address_effort high low, medium, high, xhigh, max per key
models.standard_model opus haiku, sonnet, opus, fable per key
models.standard_effort high low, medium, high, xhigh, max per key
models.knowledge_model opus haiku, sonnet, opus, fable per key
models.knowledge_effort medium low, medium, high, xhigh, max per key
models.knowledge_distill_model sonnet haiku, sonnet, opus, fable per key
models.knowledge_distill_effort high low, medium, high, xhigh, max per key
models.integration_verify_model opus haiku, sonnet, opus, fable per key
models.integration_verify_effort xhigh low, medium, high, xhigh, max per key

"Whole section" means a project [terminal] or [context] section replaces the user tier's section outright, so a key it omits takes the built-in. "Per key" means a project [pressure] or [models] section overrides only the keys it names and the rest fall through to the user file. The pressure.* keys are explained under Primary Commands, the models.* keys under Model Allocation.

Web Dashboard

loom status --web [PORT] [--host HOST] [--terminals]

The dashboard binds to 127.0.0.1 by default and serves the live ledger over a WebSocket in the browser; see Web Dashboard Remote Access for --host. Two views share the same data: the graph (below) draws the plan as a dependency graph colored by stage state, and /ledger puts one row per stage in a table with the same columns as loom status --live. Both carry a "needs attention" panel naming the stages that need a person, with a suggested command for each.

A daemon poller checks Claude and Codex quota roughly every 180 seconds; whenever either has data, the dashboard's sticky footer shows a percent bar and reset countdown for each, the same meter drawn above the legend in loom status --live.

The loom web dashboard showing a plan as a dependency graph of stages

The loom web dashboard ledger view, one row per stage

The /ledger view: state, dependencies, models, activity and context for every stage.

Without PORT, it starts at port 7373 and automatically tries the next available port when a candidate is occupied. Supplying a nonzero PORT requests that exact port; PORT 0 asks the OS for any free port. It works without the daemon by polling .loom/work/ files directly when the daemon socket is unreachable. Besides the ledger, it exposes a settings dialog for editing loom's configuration and, with --terminals, a way to open a stage's terminal from the browser.

Web Dashboard Remote Access

loom status --web --host 0.0.0.0        # every interface
loom status --web --host 192.168.1.5    # one concrete address
loom status --web --host ::             # every interface, IPv6

--host accepts an IPv4/IPv6 literal or the exact alias localhost; it rejects DNS names, URLs, and anything carrying its own port (--web already owns the port). Left unset it binds 127.0.0.1, which keeps the dashboard's original behavior exactly.

Any bind that is not loopback — a concrete address or a wildcard (0.0.0.0/::) — turns on remote mode: loom mints one process token and prints a URL containing it at startup. Every route then requires a cookie exchanged for that token by opening the printed URL once; a wrong or missing token gets refused. Each request's Host must name the connection's own accepted socket address exactly, and a supplied Origin must match it too, so a DNS name, wrong port, or a proxy that rewrites Host all get refused.

Limits:

  • Plain HTTP, no TLS — anyone who observes the token or cookie on the network gets dashboard and settings access, and terminal control if --terminals is also set. Use it only on trusted networks, or tunnel a loopback bind instead: ssh -L 7373:127.0.0.1:7373 host.
  • No reverse-proxy support: a proxy that rewrites Host to a DNS name is refused.
  • No DNS hostnames for --host, and no config key for the bind host — it is a CLI flag only.
  • The token is per process; restarting the dashboard invalidates every existing cookie.

Web Dashboard Settings

The dashboard header has a settings button; opening it (or navigating to ?settings=1) edits loom's configuration in place, so the browser's back button closes the dialog. Each row shows all three tiers — built-in, user, project — side by side as table columns, or as one card per row below 700px wide. It edits every key in the registry loom config does — eighteen keys across update, terminal, context, pressure, and models: update.check / update.check_interval_hours (loom's self-update check), terminal.backend (see Terminal Backends), context.ceiling_tokens (see Plan-Level Context Fields), the six pressure.* model and effort picks described under Primary Commands, and the eight models.* model and effort picks described under Model Allocation. loom config --list prints every key with its current value and origin. Every control here writes immediately on change, one key at a time; there is no separate Save step, and validation errors from the server surface next to the control that triggered them.

Above that config table sits a dashboard section that never touches those files — two preferences kept entirely in the browser's localStorage: a theme picker (Ledger, Aubergine, Pacific, Graphite, each drawn as a miniature of the stage graph in that theme's own colors) and a desktop-notification switch that raises an OS notification when a stage needs a human, needs a handoff, or the whole run finishes. The section renders even when /api/config is unreachable, and stays visible under any filter that matches it. Desktop notifications need a secure context — https, localhost, or a loopback address — so the switch disables itself and explains why when the dashboard is served over plain HTTP to a non-loopback host.

Two scopes:

Scope File Applies to
user ~/.loom/config.toml Every workspace on this machine
project <repo>/.loom/work/config.toml This workspace only

Sixteen of the eighteen keys have a project tier: terminal.backend, context.ceiling_tokens, and every pressure.* and models.* key; the dialog marks the remaining update.* keys as machine-wide rather than offering a project control that would do nothing. A key's effective value resolves project → user → built-in default, and each row shows which tier is currently in force plus what clearing an override would fall back to.

The fallback is not uniform across sections. For [pressure] and [models], the project tier resolves per key — a project section that sets only one key still lets every other key in that section fall through to your user config. For [terminal] and [context], a project override still replaces the whole .loom/work/config.toml section, not just one key: clearing the override removes the key and, if that empties the section, the section too — but if a sibling key is still in there (as with [context], since loom init writes ceiling_tokens alongside subagent_ceiling_tokens), the section still wins as a whole and the value resolves to the built-in default rather than falling through to your user setting. The dialog reports each case accurately; for [terminal]/[context] that section-level fallback just may not be what you expected.

On the default loopback bind, the dashboard stays an unauthenticated tool for the person running it — the settings endpoint adds no login. On a remote (--host) bind, every route including settings requires the startup-token cookie in addition to the checks below. Writes are gated by the same Host check as the rest of the dashboard, plus a strict Origin check (loopback: must be present and loopback; remote: must equal the connection's own address exactly) and a per-process CSRF token issued on load.

Web Dashboard Terminals

loom status --web --terminals (tmux backend only) adds a "Take control" view to each stage's detail dialog: an xterm.js terminal opens right in the browser, attached to that stage's live tmux session. It starts in a read-only View mode; switching to Control sends every keystroke to the running agent. Enabling --terminals mints a one-time token and prints it in the startup URL — open that exact link once to set an auth cookie for the dashboard; a plain --web link never gets the "Take control" option. With a non-loopback --host, the same startup token already covers the dashboard cookie; the cookie alone still never grants "Take control" — that capability comes only from passing --terminals.

The loom web dashboard's terminal view attached to a live stage session

A stage's terminal, opened from its detail dialog: live output, viewable read-only or with control handed over.

Plan Format

Plans live in doc/plans/ with metadata in fenced YAML between loom markers.

# PLAN-0001: Feature Name

<!-- loom METADATA -->

```yaml
loom:
  version: 1
  sandbox:
    enabled: true
  stages:
    - id: implement-api
      name: Implement API
      summary: Adds the endpoint and its tests.
      description: Add endpoint + tests
      working_dir: "."
      stage_type: standard
      dependencies: []
      acceptance:
        - "cargo test"
        - command: "cargo test api_integration::returns_200"
          stdout_contains: ["test result: ok"]
      files:
        - "loom/src/**/*.rs"
      artifacts:
        - "loom/src/api/*.rs"
      wiring:
        - source: "loom/src/main.rs"
          pattern: "mod api;"
          description: "API module registered"
      execution_mode: team

    - id: integration-verify
      name: Integration Verify
      working_dir: "."
      stage_type: integration-verify
      dependencies: ["implement-api"]
      acceptance:
        - "cargo test --all-targets"
        - command: "cargo test api_integration::returns_200"
          stdout_contains: ["test result: ok"]
```

<!-- END loom METADATA -->

Plan-Level Context Fields

Set in the loom: block alongside version and stages, these supply the default ceiling for every stage that does not declare its own. Both are absolute resident-token counts with a minimum of 60000, and both are persisted into .loom/work/config.toml's [context] section at loom init.

Field Required Notes
context_ceiling_tokens No Default ceiling for a stage's main agent session (default 150000)
subagent_ceiling_tokens No Ceiling for subagents spawned by a stage session (default 120000); never read from a stage
auto_merge No Plan-wide default for automatic merge on completion; a stage's own auto_merge overrides it, and both fall back to the orchestrator's own setting (loom run --no-merge disables it) when unset
change_impact No Nested block comparing before/after state; see Change Impact

Stage Fields

Field Required Notes
id Yes Stage identifier
name Yes Human-readable title
working_dir Yes Relative execution directory (. allowed)
description No Task brief for the agent: the full spec its session works from
summary No 1-3 sentences for a human on what the stage delivers; the web dashboard shows it instead of description. loom plan verify warns on a stage without one
dependencies No Upstream stage IDs
parallel_group No Optional label grouping related stages; the dependency graph alone still decides scheduling order
auto_merge No Per-stage override for automatic merge on completion; takes priority over the plan-level and orchestrator defaults
acceptance Conditionally required Shell criteria (strings or extended objects with stdout_contains etc.)
setup No Setup commands
files No File glob scope
stage_type No standard (default), knowledge, integration-verify, knowledge-distill
artifacts / wiring Conditionally required Required for standard and integration-verify (acceptance OR goal-backward)
wiring_tests / dead_code_check No Extended verification
before_stage No Pre-spawn checks (TruthCheck list); stage → Blocked if any fail
after_stage No Post-acceptance checks (TruthCheck list); completion fails if any fail
code_review No integration-verify only: dimensions (string list) and require_all (bool); rendered as checklist in agent signal
bug_fix No Marks this stage as a bug fix; requires regression_test (Bug-Fix Stages)
regression_test Conditionally required Required when bug_fix is true: file (test path, relative to working_dir) plus optional must_contain patterns
model No Model for this stage's main agent; omit to use the stage type's configured default (Model Allocation), overridable via [models] in either config file, or set here as a deliberate per-stage override
reasoning_effort No low, medium, high, xhigh, max; omit to use the stage type's configured default (Model Allocation), overridable the same way
implementers No Licensed agent lanes as a list, first = preferred for routine work: ["codex", "claude"]. Default ["claude"]. Listing a lane makes it available, not mandatory — a stage mixes lanes per subagent
ultracode No License this stage for large multi-agent fan-out; per-stage opt-in (default false)
skills No Skills this stage's agents need, by catalog name (["loom-rust", "loom-debugging"]). loom plan verify rejects an empty, duplicate or unknown name; the signal lists each as a skill to load before the detected ones, and a subagent's brief names them
subagent_timeout_secs No Seconds of tool silence before the monitor warns appears hung (default 300); the advisory idle budget death is judged against — never the --timeout on the single owned loom subagents watch --worker ... --timeout 3600 call
context_ceiling_tokens No Absolute resident-token ceiling for this stage's session (minimum 60000). Resolved stage value → plan-level context_ceiling_tokens → 150000. The session hook warns at 80% and blocks at 100%; the daemon forces a handoff at 125%
plan_overview No Set false to suppress the embedded plan overview in this stage's signal
sandbox No Per-stage sandbox override
sandbox.permission_mode No auto (default), accept-edits, plan, default — resolves stage > plan > stage-type default; bypass-permissions is rejected at init
execution_mode No single (default) or team hint

Stage Type Behavior

  • knowledge: knowledge/bootstrap work, different verification expectations
  • standard: implementation stage; must define goal-backward checks
  • integration-verify: final quality gate combining code review and functional verification; must define goal-backward checks. Define code_review.dimensions to render a checklist of review dimensions in the agent's signal.
  • knowledge-distill: final stage; curates stage memories into permanent knowledge files

Plan Version 2

version: 2 turns on stricter verification; version: 1 keeps the rules above unchanged, and a v1 plan that uses a v2-only field is rejected (`<field>` requires `version: 2`). The v2 fields:

Field On Notes
contracts stage Tests written and frozen before implementation: id, file, test, optional runner, scenario, and rejects (the plausible wrong implementation the test must fail on). Every standard stage of a v2 plan needs at least one; knowledge, knowledge-distill and integration-verify stages take none
harness stage Globs of extra test-only files the contract writer may edit. Every matched file is frozen with the contracts, so do not glob production files
reachable stage symbol, from, description, optional min_confidence: the new unit must be reachable from an entry point in the source graph
literal: true, glob source wiring entry Match the pattern as literal text; let source be a glob (*, ?, [). A pattern that matches only a definition of a name defined in that file is rejected, so point it at a consumer
ratchet_files plan Exact paths of baseline or ledger files a stage could loosen; any change to one raises a test-integrity event
loom:
  version: 2
  ratchet_files: ["loom/maintainability-baseline.txt"]
  stages:
    - id: implement-api
      stage_type: standard
      contracts:
        - id: rejects-empty-name
          file: "tests/api_contract.rs"
          test: "rejects_empty_name"
          scenario: "POST /users with an empty name"
          rejects: "an implementation that stores the empty name"
      reachable:
        - symbol: "create_user"
          from: "main"
          description: "the handler is reachable from main"

A v2 standard stage starts with a separate contract session that writes the contract tests, runs them red, and ends with loom stage contracts freeze; loom then hands the stage to the implementing session. loom project detect shows which test runner loom will use for each package. Runners without an adapter fall back to running the contract's test string and judging its exit code.

The contract phase is visible wherever loom status renders: a magenta contracts tag next to the stage's [model] tag (added to the legend while such a stage exists), the live TUI's activity cell reading writing contract tests or, once stale, contract writer idle <duration>, and the web dashboard's violet contracts tag with a dashed node outline, the same activity text, a ... · contract phase stage-strip label, and a contract writer row in the stage dialog's session type.

A refused freeze — a file changed outside the contracts and harness globs, a contract not reported red, or (checked by the daemon) a FIFO, socket or device node left in the worktree — parks the stage WaitingForInput with the problems as its review_reason, the reason loom status shows for it. The daemon parks the stage as soon as it refuses a relayed freeze; a refusal from the writer's own pre-check parks it when the writer stops. Typing into the writer's session resumes it, and so do loom stage resume <stage-id> and a later freeze the daemon accepts; loom stage retry still refuses a stage in WaitingForInput. A writer that stops before its first freeze attempt is not parked.

Verification Model

loom check <stage-id> validates outcomes, not just compilation/tests:

  • acceptance: shell criteria (simple strings or extended objects with stdout_contains, exit_code, etc.)
  • artifacts: real implementation files exist
  • wiring: critical integration links exist
  • wiring_tests: runtime integration checks
  • dead_code_check: detect unused code via command output patterns

For standard and integration-verify stages, acceptance criteria or at least one goal-backward check must be defined.

In a version 2 plan, loom stage complete also runs, in order: the contract check (every frozen contract file unchanged, every contract test passing), test integrity (fewer test declarations or assertions, removed or changed assertion lines, or a changed ratchet_files entry each raise an event that must be fixed or accepted through an adjudicated dispute), the tests reached from the changed code in the source graph, and the review gate. A criterion that runs a recognised test runner and selects zero tests fails. The gate needs a loom-code-reviewer round on the final change fingerprint with every finding fixed or ruled on; editing any file after the last round, formatting included, needs another round, while committing does not. The loom daemon that owns the worktree computes this fingerprint, once, for both the round a review records and the completion that checks it; a sandboxed loom stage complete cannot reach the daemon's socket, so it skips test integrity and the review gate locally, and the daemon runs both when it applies the completion. Integration-verify re-checks every stage's reachable and wiring on the merged tree and lists the full test command.

Verification Is Enforced, Not Self-Reported

loom stage complete is the only way a stage finishes, and it runs the acceptance criteria itself before doing anything else. If they fail, the stage stays Executing — the agent must fix the work and re-run, and fix_attempts is incremented so repeated failures surface rather than accumulate silently. after_stage checks then run post-acceptance, and goal-backward verification runs before the progressive merge. A passing criterion is cached against an exact fingerprint of its command and the tree, so an unchanged criterion is not re-run on the next attempt; --no-cache on loom check or loom stage complete forces a real run regardless.

Artifact verification treats a stub as a failure: a file that exists but contains TODO, FIXME, unimplemented!, todo!, a bare pass, or raise NotImplementedError does not count as delivered.

Normal stage completion crosses a narrow control boundary. The stage runs one exact pinned loom stage complete <stage-id> command; a PostToolUse bridge accepts only the matching stage and session, requires Loom's verification marker, and sends a non-extensible CompleteStage request to the daemon. The request cannot carry commands, paths, or bypass flags.

The three bypass flags — --no-verify, --force-unsafe, --assume-merged — are the operator's, and cost the operator nothing: loom stage complete <stage> --no-verify just works from your shell. It authorizes itself against .loom/work/admin.token, which you can already read and a sandboxed agent cannot (the sandbox binds the whole process tree, so a loom an agent spawns is denied the same read). The proof is still bound to the project, stage, action, and exact flag set, and consumed on first use — you simply never handle it.

The same applies to loom stop. Nothing asks a human to mint a credential they already hold; making them carry an HMAC between two commands added ceremony, not security.

loom stage admin-proof remains for the case it was actually built for: a trusted broker minting a narrowly-scoped capability for another process. It takes the secret through LOOM_ADMIN_TOKEN and never reads the token file, so a caller that can invoke loom but cannot read that file gains nothing.

An agent that genuinely believes a criterion is wrong or impossible has a sanctioned path — loom stage dispute-criteria — rather than an incentive to weaken it.

Verification Is the Main Agent's Job

Subagents do not verify. A subagent may run at most one narrowly-scoped check covering the files it just changed; project-wide builds, full test suites, and repo-wide lint or typecheck runs belong to the main agent — the only party that can see the whole tree and act on the result.

This is enforced, not just advised: loom-hooks/subagent-verify-guard.sh (a PreToolUse:Bash hook) blocks project-wide runners — cargo build, cargo test, make test, tsc, go build and friends — when the caller is detected as a subagent. Scoped invocations pass, quoted mentions are ignored, and unrecognised commands are always allowed: a false block would strand a subagent mid-task.

Two things worth knowing:

  • integration-verify stages are carved out. That stage type exists to run the complete suite, so its subagents may. The carve-out is read from the stage file and fails safe — an ambiguous or missing stage file means no relaxation.
  • There is deliberately no opt-out environment variable. The main agent is never affected, so an escape hatch would only serve to defeat the rule.

Bug-Fix Stages

A stage with bug_fix: true requires a regression_test block naming the test that pins the fix: file (path to the test, relative to working_dir) and optional must_contain (patterns the file's content must include). Verification checks the file exists and carries those patterns, so a bug-fix stage cannot complete on a fix with no test guarding the regression.

Change Impact

A plan-level change_impact block separates the failures a stage introduced from the ones it inherited. Loom runs baseline_command when a stage starts executing and saves the output; at loom stage complete it runs the command again and compares the lines matching failure_patterns:

loom:
  version: 1
  change_impact:
    baseline_command: "cargo test --no-fail-fast 2>&1 || true"
    compare_command: "cargo test --no-fail-fast 2>&1 || true" # optional; defaults to baseline_command
    failure_patterns:
      - "FAILED"
      - "error\\[E"
    policy: fail # fail (default) | warn | skip

The comparison reports new failures and fixed failures. policy decides what a new failure does to the stage: fail blocks completion, warn reports it and continues, skip turns the check off. Failures already present in the baseline do not count against the stage.

Disputes, Adjudication, and Amendments

loom stage dispute-criteria is how an agent challenges a criterion instead of quietly weakening it. The daemon moves the stage to NeedsAdjudication and spawns a real session — with the full tool surface, running the disputed criterion itself — to judge it; loom stage adjudicate records that session's verdict, and a stage's own worktree session is refused so it can never judge its own dispute.

A verdict does one of three things: Accept patches acceptance or wiring and re-queues the stage; NeedsMoreEvidence appends the judge's questions to the next signal and re-queues, up to a capped number of rounds; Reject moves the stage to NeedsHumanReview. Every accepted amendment writes a numbered snapshot under .loom/work/plan_versions/ plus an audit row.

Version 2 plans add three more dispute kinds, each filed once per review round with several ids: dispute-findings (a judge rules each finding uphold, dismiss or defer to a dependent stage that is not yet finished; integration-verify never defers), dispute-contract (Accept re-freezes the contract and may amend contracts; Reject sends the agent to loom stage contracts restore) and dispute-integrity (Accept records the event as accepted at its current count or hash). Each kind has its own budget of three disputes per stage.

The plan-level adjudication.max_amendments_per_stage field bounds autonomous amendments per stage (default 10), so a multi-fix plan is not sent to a human part-way through. loom stage amend gives an operator the same audited path directly: it replaces, inserts, or deletes one element of a stage's acceptance, wiring, or wiring_tests array, snapshotting and logging the change the same way.

Knowledge System

Loom's answer to "every session starts from zero" is a three-stage pipeline: capture during execution, distill at the end of a plan, retrieve cheaply forever after.

1. Capture — session memory

While a stage runs, its agent journals to .loom/work/memory/<session>.md:

loom memory note "gotcha: worktree exclude lives at <worktree>/.git/info/exclude, not <dir>/.git/..."
loom memory decision "centralized plan lookup in plan/parser" --context "avoids an orchestrator→commands layering violation"

Entries are typed (note, decision, change, question). The most recent are embedded in the recitation section at the end of the next signal — the position with the highest model attention — so a later stage inherits the detail an earlier stage paid for instead of rediscovering it.

Memory is deliberately cheap and disposable. It is a working journal, not the deliverable.

2. Distill — memories become knowledge

A knowledge-distill stage runs at the end of a plan and performs the reduce step: it reads every stage memory, dedupes, and curates the survivors into permanent knowledge. Mistakes are rewritten as actionable prevention rules rather than anecdotes:

## [Short description]

**What happened:** ...
**Why:** [root cause]
**Prevention:** [how to detect it earlier]
**Fix:** [what to do instead]

Procedural noise ("spawned agents", "ran tests") and anything recoverable from git history is dropped. The stage starts from loom memory pending --group, which prints corrections first, sorted by target file and heading so each one is applied with replace-section, then mistakes, decisions and the rest; a mistake that repeats one already in the tree becomes a proposal for a hook or a loom plan verify check rather than another paragraph. Every entry ends with a loom memory resolve receipt. loom review turns the same memories into a human-readable code-review document.

3. Retrieve — a tiered base agents can afford to read

Knowledge lives in doc/loom/knowledge/ and is tiered: a generated INDEX.md (tier 0) maps the seven curated summary files (tier 1), which link out to per-category topic files (tier 2, e.g. architecture/merge-flow.md). Tier-1 files stay navigable summaries; detail lives in topics. The index is regenerated automatically on every knowledge write.

Reading protocol — agents read INDEX.md first for orientation, then the tier-1 summary for the area they are working in, then only the tier-2 topics they actually touch; a specific question is pulled with loom knowledge context --query, which returns the matching sections quoted. Inside a stage, the per-stage Knowledge Brief comes first. Loading the whole base defeats the point of tiering. Every session that starts inside a repository holding doc/loom/knowledge/INDEX.md also receives a one-paragraph pointer to it from the knowledge-orient.sh SessionStart hook, installed globally by loom init/loom repair --fix.

Writing protocol — when a tier-1 section grows past roughly 40 lines, move its body into a topic with loom knowledge update <category>/<slug> and leave a 2-4 line summary plus a relative link behind. Write the link as [Title](category/slug.md) in a tier-1 file: that is the tree's convention for references.

There is no aggregate line budget across the knowledge base. What matters is per-file size, and loom knowledge check --strict enforces it: a tier-1 summary is at most 250 lines with sections of at most 40, a tier-2 topic at most 400 lines with sections of at most 80, and the generated INDEX.md at most 16 KB. A tree that cannot clear those yet records its current issues once with --write-baseline and gates with --baseline, so only new issues fail. loom memory note refuses a mistake: note without a Prevention: and a stale-knowledge: note that does not name <file>#<heading> and a Correction:, so a correction can be applied mechanically.

Retrieval is deterministic and offline. loom knowledge context --query <text> returns a token-budgeted context pack: the tool chunks the curated prose, scores each chunk, fuses the per-channel rankings, and takes whole chunks in order until the budget is spent, always reporting what it left out. There is no embedding model, no network call and no randomness — a pack is a pure function of the bytes on disk and the query string, so the same query returns the same pack.

loom knowledge context --query "how does merge cleanup order work" --budget-tokens 3000
loom knowledge context --query "source graph coverage" --explain   # per-item scores and why each was selected
loom knowledge context --query "sandbox rules" --json              # machine-readable

Stage sessions do not have to ask. Signal generation embeds a per-stage Knowledge Brief built through the same single entry point, so what a stage receives at spawn and what you get from the CLI are produced identically. Loom records what each recipient was given, so a second retrieval in the same session skips what the first already quoted rather than repeating it.

--scope selects which channels to search: knowledge searches the curated prose, source searches the derived source graph of symbols extracted from the code, and all (the default) fuses both. The two are ranked separately — prose over chunk text, symbols over their scope and signature — and then fused, so one pack can mix curated prose with the exact symbols a query names. A symbol whose file the parser could not fully read is still returned, but without a high-confidence claim.

The source graph maintains itself. loom init and loom run publish a base layer for the current revision before anything else starts, and each stage's working-tree overlay is refreshed immediately before its signal is written — so a stage's brief describes the code as that stage will actually find it. Publication is advisory: when it cannot run, loom prints one line and carries on, because a missing graph must degrade retrieval rather than block a run. A base layer is keyed to a revision and so is never published from a dirty tree; there, loom knowledge sync builds a working-tree overlay instead and tells you which layer it wrote.

Codex's native PostToolUse:apply_patch hook records every patched path into the stage overlay, and its UserPromptSubmit hook uses the same retrieval entry point as Claude. Shell-based edits remain outside file-tool hook visibility and are reconciled by the next explicit graph refresh.

Run loom knowledge sync after editing knowledge outside the CLI; it reports whether both derived layers are current, naming the source-graph layer it produced — base, local-overlay or skipped, with the reason.

Bootstrapping and maintenance

Agents populate knowledge during orchestration: a knowledge-bootstrap stage writes CONTENT, while loom init scaffolds the directory automatically. loom knowledge sync rebuilds the derived retrieval artifacts.

Knowledge directories created before the hierarchy existed stay flat and keep working unchanged — neither reading nor updating migrates them behind your back. loom knowledge sync performs the opt-in upgrade, creating INDEX.md the first time it runs on a flat directory.

Knowledge writes are protected by the sandbox defaults: agents update knowledge through loom knowledge ..., never by editing the files directly.

Bootstrapping a repo not managed by loom

loom knowledge bootstrap [--structural-only] [--refresh] [--dry-run] [--model M] [--effort E] builds doc/loom/knowledge/ from scratch for a repository that has no loom plans — retrieval needs the directory to exist before sync/context can do anything. It runs a deterministic host phase (scaffold the tree, rebuild the catalog and source graph, partition the repo into directory clusters with content digests) followed by an interactive foreground Claude session that writes knowledge only through loom knowledge ..., then a host finalization that refreshes the index and writes a receipt.

  • --structural-only stops after the deterministic phase — no model session is launched.
  • --refresh narrows the work to clusters whose digest changed since the last receipt, any removed cluster, and any tier-1 file still at its template content; with nothing changed it prints knowledge is current and spawns nothing.
  • --dry-run prints the work plan and the exact claude command it would run, without spawning anything.
  • --model/--effort pick the model and effort for the semantic session, same values as elsewhere in loom.

The command refuses to run inside a stage session — stages write knowledge through loom knowledge update instead. The committed receipt doc/loom/knowledge/.bootstrap-receipt.json records one digest per cluster; commit it alongside the generated knowledge files so --refresh reports current across clones, not just locally.

Model Allocation

Every stage's main agent is an orchestrator; the model and effort it runs come from its stage type's default, which the operator can override:

Stage type Default model Default effort
standard opus high
knowledge opus medium
knowledge-distill sonnet high
integration-verify opus xhigh

Configure either default per stage type in the [models] section of ~/.loom/config.toml (user tier) or <repo>/.loom/work/config.toml (project tier, resolved per key — a project section that sets only one of these eight keys still lets the rest fall through to the user config):

# ~/.loom/config.toml or .loom/work/config.toml
[models]
standard_model = "opus"
standard_effort = "high"
knowledge_model = "opus"
knowledge_effort = "medium"
knowledge_distill_model = "sonnet"
knowledge_distill_effort = "high"
integration_verify_model = "opus"
integration_verify_effort = "xhigh"

A plan stage's model / reasoning_effort field overrides both config tiers for that one stage. Merge and base-conflict sessions stay pinned at opus/high and adjudication keeps its own [adjudication] model; neither is configurable through [models].

The orchestrator decomposes the work, hands each subagent full context, then verifies and commits. It makes a change itself only when that is cheaper than a spawn: at most 20 changed lines in at most 2 files it has already read, proven by one command; a fable session delegates even those. Implementation is delegated to as few subagents as the work allows, each spawned by agent type so the model choice is explicit:

Agent Model Use for
loom-software-engineer Sonnet Common implementation and integration tests to detailed instructions
loom-codex-forwarder Codex GPT-5.6 Terra or GPT-6 Luna Codex lane, licensed only on stages listing codex in implementers: Terra for common implementation/integration tests, Luna for boilerplate, scaffolding, and simple unit tests
loom-senior-software-engineer Opus Mainstream architecture and algorithm implementation, complex debugging, security-sensitive or cross-cutting work
loom-code-reviewer Opus Read-only code, security, and architecture review
loom-advisor Fable Diagnosis after a repeated failure — advice returned, nothing written

Fable-tier implementation — major bugs, visual/UI design, extremely challenging algorithmic design — has no dedicated agent type; it is spawned with an explicit model override rather than relying on inheritance.

The loom-codex-forwarder row additionally depends on the codex CLI and its plugin's companion runtime being installed. loom run checks this at startup and prints an advisory warning if either is missing — it never blocks the run — and terra-/luna-tier work falls back to Sonnet for the duration; the stage signal states the fallback explicitly and does not spawn loom-codex-forwarder.

This is why savings come from delegation rather than downgrade: an untyped subagent silently inherits the stage's own (usually Opus) model, making every worker expensive. Two failures on the same task should produce a loom-advisor diagnosis, not a blind retry at a larger model.

A stage omits model and reasoning_effort by default, so the stage type's configured default applies. Set either field only as a deliberate per-stage override (low, medium, high, xhigh, max for effort). ultracode: true licenses a stage for large multi-agent fan-out; it is per-stage opt-in so the cost decision stays explicit.

Sandbox Configuration

Loom supports plan-level defaults plus stage-level overrides.

loom:
  version: 1
  sandbox:
    enabled: true
    auto_allow: true
    filesystem:
      deny_read:
        - "~/.ssh/**"
        - "~/.aws/**"
        - "../../**"
        - "../.worktrees/**"
      deny_write:
        - "../../**"
        - "doc/loom/knowledge/**"
      allow_write:
        - "src/**"
    network:
      allowed_domains: ["github.com", "crates.io"]
      additional_domains: []
      allow_local_binding: false
      allow_unix_sockets: []

Note: knowledge file writes are intentionally protected by sandbox defaults; knowledge updates should be done via loom knowledge ... commands. Plan-configured excluded_commands are rejected because broad executable exemptions bypass the host sandbox. When sandboxing is enabled, generated settings use host denyRead rules for sensitive paths and failIfUnavailable: true; failure to write those settings blocks session spawn. Unit tests pin the generated policy and blocked-spawn behavior. A credentialed Claude host-runtime canary across Bash, interpreters, build scripts, symlinks, and file tools remains a manual release check.

Command Confinement

The sandbox: block above bounds the agent session. A separate control bounds the commands loom itself runs from your plan — every acceptance criterion, setup command, truth check, wiring test, dead-code check and change-impact command:

loom:
  version: 1
  sandbox:
    command_confinement: confined # plan-level default; `inherit` to opt out

  stages:
    - id: build
      sandbox:
        command_confinement: inherit # per-stage override
Level Behavior
confined Default. The child process environment is cleared and rebuilt from a fixed allowlist
inherit The child inherits loom's ambient environment

Plans are trusted artifacts, but trusted is not privileged: under confined, a plan line cannot read GITHUB_TOKEN, AWS_* or ANTHROPIC_API_KEY merely because you started loom from a shell that had them. The allowlist carries what a build toolchain needs to find itself — HOME, PATH, CARGO_HOME, RUSTUP_HOME, locale and terminal variables, TMPDIR, the proxy variables and the CA-bundle locations (SSL_CERT_FILE, SSL_CERT_DIR, NIX_SSL_CERT_FILE). SSH_AUTH_SOCK is deliberately withheld, so an acceptance criterion that needs SSH auth fails by design rather than silently borrowing your agent.

What confinement is not. It is environment scrubbing — least-privilege hygiene, not a security boundary. Loom applies no namespace, seccomp, landlock, cgroup or network isolation to the commands it spawns: a confined command shares your network namespace, can read and write any path your user can, and can reach any Unix socket on the host. The network: settings above are emitted into the agent session's sandbox and do not restrict plan-authored commands. Use confined to keep ambient credentials out of plan commands; do not use it to run code you would not run yourself.

Session State Confinement

Every spawned session — stage, merge, base-conflict, adjudication — launches from a settings capsule that denies writes to .loom/, .claude/, .worktrees/, the hooks directories, and git hooks and config, regardless of the filesystem rules above.

A handful of commands still need to reach that state from inside a sandboxed session: loom memory ..., loom stage block, loom stage dispute-criteria, loom handoff, loom stage merge --resolved, and loom worktree remove. Each writes a one-shot ticket instead, picked up by a PostToolUse relay hook and applied by the daemon at most once through a per-session ledger. loom request status <id> reports whether the daemon has applied a given ticket yet.

The daemon also refuses to merge, or hand to a conflict-resolution session, any branch whose diff touches .claude/, .mcp.json, .loom/, or the in-repo hooks directory — the stage moves to NeedsHumanReview naming the offending paths instead.

Permission Mode

All stages default to auto (agents auto-accept any action their heuristics deem safe, since loom stages run autonomously with no human to answer prompts; the sandbox deny/allow rules are the safety boundary). Override per-plan or per-stage to tighten control:

loom:
  version: 1
  sandbox:
    permission_mode: accept-edits # plan-level override

  stages:
    - id: my-stage
      sandbox:
        permission_mode: plan # stage-level override (takes precedence)

Valid values: auto (default), accept-edits, plan, default. bypass-permissions is rejected at init time.

Remote Control

Claude Code's --remote-control flag lets the loom orchestrator drive spawned Claude sessions programmatically. Loom enables it automatically when prerequisites are met — no configuration required.

Prerequisites (preflight check):

  • Claude version ≥ 2.1.51
  • Auth: claude.ai login — loom accepts either credential store: ~/.claude/.credentials.json, or (on macOS) a Claude Code-credentials entry in the Keychain, which is where Claude Code stores credentials on macOS instead of the file. Additionally, none of these env vars may be set: ANTHROPIC_API_KEY, CLAUDE_CODE_OAUTH_TOKEN, CLAUDE_CODE_USE_BEDROCK, CLAUDE_CODE_USE_VERTEX, CLAUDE_CODE_USE_FOUNDRY

The flag exits non-zero when its prerequisites are not met, so loom never passes it blindly. When preflight fails, loom falls back silently to standard mode and prints a one-line advisory at orchestrator startup (e.g. ⚠ Remote Control disabled: <reason>).

Configuration — the [remote_control] section of .loom/work/config.toml carries a single switch:

# .loom/work/config.toml
[remote_control]
mode = "auto"   # default: enable whenever preflight passes
# mode = "off"  # never enable, regardless of preflight

Toggling mode takes effect on the next session spawn — no daemon restart needed.

Fast crashes — a session that crashes within 15 seconds of spawn while Remote Control is active disables Remote Control for the rest of that daemon run, logs it once, and retries the stage without the flag. Nothing is written to disk: a daemon restart tries Remote Control again from scratch. Set mode = "off" above to stop it from trying at all.

Session naming — every spawned session is named after its stage in the Remote Control UI: the stage name for stage sessions, and Merge: <stage name>, Base conflict: <stage name>, Knowledge: <stage name> for merge, base-conflict, and knowledge sessions respectively. Claude binaries whose --remote-control flag doesn't accept a name argument automatically fall back to the bare flag — detected via a one-time claude --help capability check, no configuration needed.

Terminal Backends

Loom spawns each stage's Claude Code session through a terminal backend. Two are available.

Backend Default Sessions run in Needs a GUI?
native ✅ yes a host terminal emulator window yes
tmux opt-in a detached tmux server (no window) no

The native backend opens a real terminal window per session — you watch stages run in your own terminal emulator. It requires a detectable emulator, so it cannot run headless.

The tmux backend spawns each session into a detached tmux server instead, which makes loom usable over SSH, on a headless Linux box, or anywhere no terminal emulator exists. tmux must be installed and on PATH.

Selecting a backend

# .loom/work/config.toml
[terminal]
backend = "tmux"   # or "native" (default)

Or from the CLI:

loom init <plan> --backend tmux   # skips the interactive backend prompt
loom run --backend tmux           # persists the choice to [terminal]

loom init prompts for a backend when run interactively; with no TTY it defaults to native. Changing the backend while the daemon is running is refused with a hint — loom stop first, then re-run with --backend. Selecting a backend takes effect on the next spawn.

If tmux is selected but not installed, loom init prints an advisory warning; loom run (and loom run --foreground) refuses to start — the configured backend is never silently swapped for the other lane.

Running under WSL

WSL2 runs the published loom-linux-x86_64 binary unmodified — it is an ordinary glibc ELF, and git worktrees, the .loom/work/ Unix socket and PID liveness checks all behave as they do on native Linux.

What does not carry over is the native backend. It opens a real terminal-emulator window per session, and a stock WSL install ships no Linux GUI stack — without WSLg or an X server there is no emulator to detect. Select tmux explicitly:

sudo apt install tmux jq          # both required; ripgrep and fd are recommended
loom init doc/plans/PLAN-<name>.md --backend tmux
loom run --backend tmux

Sessions then run in detached tmux servers, and you watch them from a Windows Terminal tab:

loom status --live       # live ledger dashboard
loom status --web        # browser dashboard: starts at 7373, then the next free port
loom attach              # tiled overview of every live session
loom attach <stage-id>   # attach to one stage

On Windows 11 with WSLg, an emulator installed inside WSL is detectable and the native backend does work, but tmux remains the better fit: a run survives closing the window, and loom attach reconnects to one already in progress.

Two WSL-specific notes:

  • Keep the repository on the Linux filesystem (~/src/…), not under /mnt/c. Loom creates a worktree per parallel stage and the daemon polls stage files every 5s; 9p latency across the Windows drive makes both crawl.
  • ARM64 Windows has no published Linux ARM64 binary. Install the Rust toolchain inside WSL and build from source.

One tmux server per session

Loom does not put every stage in one shared tmux server. Each session gets its own server on its own socket, named loom-<session-id> under $TMUX_TMPDIR (else /tmp).

This is deliberate: a wedged or killed server takes down exactly one stage, instead of every stage running in parallel. Liveness is tracked from PID files rather than by asking tmux, because a tmux server whose agent process has died still reports the session as existing — which would hide the crash from loom's monitor and prevent the retry.

Attaching

loom attach              # tiled overview of every live session
loom attach <stage-id>   # attach directly to one stage's session

With no argument, loom attach builds a per-repo viewer window with one pane per live session, tiled. With a stage id it attaches straight to that stage. Both require a real terminal (a TTY), and both work only for sessions spawned by the tmux backend — with the native backend it tells you so.

Panes in the overview are live, writable terminals, not read-only views: keystrokes go to the agent, and C-b x closes that stage's pane. Detach the normal way with C-b d.

Tmux unavailable

The configured backend is authoritative: nothing on disk can swap it for the other lane. If tmux is configured but not on PATH, loom run refuses to start (see Selecting a backend); if it becomes unavailable mid-run, a spawn fails with an error naming the fix instead of silently falling back to native.

Agent Teams (Experimental)

Loom enables agent teams in spawned sessions (CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1) and injects team-usage guidance into stage signals.

Use teams when work needs coordination/discussion across agents (multi-dimension review, exploratory analysis). Use subagents for independent, concrete file-level tasks.

State Layout

project/
├── .loom/
│   ├── work/                  # orchestration state (mode 0700)
│   │   ├── config.toml
│   │   ├── stages/
│   │   ├── sessions/
│   │   ├── signals/
│   │   ├── handoffs/
│   │   ├── memory/            # per-session journals
│   │   ├── archive/
│   │   ├── crashes/
│   │   ├── logs/              # per-session captured stderr
│   │   ├── pids/
│   │   ├── wrappers/
│   │   ├── orchestrator.sock  # daemon IPC socket
│   │   └── orchestrator.pid
│   ├── memory/archive/        # state of completed plans
│   └── cache/                 # derived caches
├── .worktrees/
│   └── <stage-id>/            # one per stage; its .loom/work links to the main one
├── doc/plans/
└── doc/loom/knowledge/
    ├── INDEX.md            # generated tier-0 map
    ├── architecture.md     # tier-1 summaries
    ├── patterns.md
    ├── ...
    └── architecture/       # tier-2 topics, one directory per category
        └── merge-flow.md

Shell Completions

Loom provides context-aware tab completions for all commands, subcommands, flags, and dynamic values (stage IDs, plan files, session IDs, knowledge files).

Quick Install

loom completions --install

Auto-detects your shell from $SHELL and writes completions to the standard location:

Shell Install Path
Bash ~/.local/share/bash-completion/completions/loom
Zsh ~/.zfunc/_loom
Fish ~/.config/fish/completions/loom.fish

Follow the printed post-install instructions to activate (e.g., for zsh, ensure fpath=(~/.zfunc $fpath) appears before compinit in ~/.zshrc).

Manual Setup

You can also write the completion script to a file yourself:

# bash
loom completions bash > ~/.local/share/bash-completion/completions/loom

# zsh — ensure ~/.zfunc is in fpath (add before compinit in ~/.zshrc):
#   fpath=(~/.zfunc $fpath)
#   autoload -Uz compinit && compinit
mkdir -p ~/.zfunc
loom completions zsh > ~/.zfunc/_loom

# fish
loom completions fish > ~/.config/fish/completions/loom.fish

Migrating from Older Versions

Older versions of loom used clap_complete and required an eval line in your shell RC file that ran a subprocess on every shell startup. The new system writes a static script to disk and only calls loom at actual tab-completion time, which means faster shell startup and completions that work even before loom is in your PATH.

To check whether you need to migrate:

loom completions --migrate

This scans for two things:

  1. eval lines in RC files (.bashrc, .zshrc, etc.) like eval "$(loom completions zsh)" — these should be removed
  2. Stale completion files containing old clap_complete markers — these need to be regenerated

If issues are found, follow the printed instructions. Typically: remove the eval line from your RC file, then run loom completions --install to write the new file-based completion script.

What's Completed

  • Commands and subcommands (loom stage <TAB> shows all stage subcommands)
  • Flags (loom run --<TAB> shows available flags)
  • Stage IDs with smart filtering (loom stage complete <TAB> shows only executing stages)
  • Plan files, session IDs, knowledge files (including aliases like deps, tech)
  • Model names, trigger types, and more

License

MIT

About

A Rust-based agent orchestrator enabling a swarm of Claude Code instances building software.

Topics

Resources

Contributing

Stars

56 stars

Watchers

4 watching

Forks

Releases

Packages

Used by

Contributors

Languages