Agentic engineering for the rest of us.
You design the system. Loom builds it, checks every piece, and learns from its mistakes.
Loop engineering is sold as the whole of software development: state a high-level goal, and an agent loops until finished software comes out. It does get there, by brute force, paying in tokens for every wrong turn and rediscovering what an engineer on the team already knew.
Loom is for engineers who would rather spend their own judgment than someone else's compute. You discuss, theorise, test, explore, and plan the system, and put your experience into that plan. Loom implements it with Claude Code (and optionally Codex) sessions running in parallel across isolated git worktrees. It checks every stage's work itself, then checks all of it together in a final integration-verify stage. Along the way it records its own mistakes and distils them into a knowledge base the next plan starts from. You can watch from a terminal dashboard or a browser, steer any session, or take the keyboard at any time.
A real run in loom status --web, sped up: watch a stage's session, take the keyboard, and see a finished stage merge and start the ones that depend on it.
- Your expertise where it counts, tokens everywhere else — you do the thinking and put it into the plan. Loom executes it unattended and comes back to you only when a stage needs a person; you can monitor, steer, or take over whenever you like. (Human Expertise Where It Matters)
- It learns your project — every session records what it got wrong and what it decided. Each plan ends by distilling that into a knowledge base in your repo, so every plan starts from what the earlier ones learned. (It Learns From Every Plan)
- Parallel by default — stages form a dependency graph, everything independent runs at once in its own git worktree, and each stage merges back as it finishes. (Parallel execution and progressive merge)
- Done means verified — loom runs each stage's acceptance criteria itself and inspects the tree for stubs, unwired code, and code nothing calls. The agent's own account of its work does not count. Once every stage has merged, an
integration-verifystage reviews and tests the combined work as one system. (Verification Model) - Rules enforced by hooks — commit discipline, worktree boundaries, and subagent limits fire from shell hooks, so they hold whatever the model intends. (Deterministic guardrails)
- Contained by default — sessions run in a filesystem and network sandbox and cannot write loom's own state, hooks, or config. (Sandbox Configuration)
- Expensive models only where judgment is needed — the orchestrator plans and verifies; implementation goes to the cheapest subagent that can do the piece, Claude or Codex. (Model Allocation)
- Survives crashes and context limits — all state is plain files. The daemon spots dead and hung sessions and retries them, and a handoff is written before a session runs out of context. (Crash recovery and liveness)
- Watch it live, steer when you want —
loom status --liveis a dashboard in your terminal.loom status --webis the same run in a browser: the plan as a dependency graph, a ledger of every stage, and terminals you can watch read-only or take over. (Primary Commands, Web Dashboard)
curl -fsSL https://raw.githubusercontent.com/cosmix/loom/main/install.sh | bashYou need git, the claude CLI, and jq. Signed binaries cover Linux x86_64 and macOS Apple Silicon; every other platform builds from source — see Installation.
Plans are how loom knows what to build. Open Claude Code in your target project and use the /loom-plan-writer skill to create one:
cd /path/to/project
claude # start Claude Code CLIInside the Claude Code session:
- Load the plan-writing skill by typing
/loom-plan-writer(or just mention that you're writing a loom plan). - Describe what you want to build and discuss with Claude, as you would in plan mode.
- When done, Claude will write the plan to
doc/plans/PLAN-<name>.md.
Validate the draft before you spend anything on it:
loom plan verify doc/plans/PLAN-<name>.mdThen read it yourself and check that it captured your intent in enough detail. For longer plans, loom pressure doc/plans/PLAN-<name>.md has two model families (Claude and Codex) review it adversarially and harden it. Without the codex CLI, run /pressure doc/plans/PLAN-<name>.md inside a Claude Code session for a Claude-only pressure test.
loom init doc/plans/PLAN-<name>.md
loom run
loom status --live
loom stoploom init parses the plan, creates stage state, and installs the project hook wiring. loom run starts the daemon and orchestrator. loom status --live follows the run in your terminal, and loom status --web puts it in a browser. Add --terminals (tmux backend only) to watch or take over the Claude Code sessions from the browser. --host <addr> binds the dashboard to an address other than 127.0.0.1; read Web Dashboard Remote Access first, since a non-loopback bind serves plain HTTP.
loom stop stops the daemon, and the orchestrator with it.
Next: the two ideas loom is built on, human expertise where it matters and learning from every plan. Then How It Works and the Feature Tour, a map of everything below.
Loop engineering presents itself as the end of the story: give an agent a high-level goal, let it try, fail, read the error, and try again until the checks pass, and finished software comes out. It does produce software, at a high price and with little elegance. Every wrong turn is paid for in tokens, and much of what those tokens buy is the rediscovery of something an engineer on the team already knew. A frontier lab can afford that. Most organisations cannot.
Human engineers have a lot to offer: experience, judgment, and a working theory of the system that no loop recovers cheaply. Loom puts that first and spends tokens on the rest:
- Up front, in the plan. You discuss the design, form theories and test them, explore the code, and decide what gets built, how it splits into stages, and what proves each stage is done.
loom plan verifyandloom pressurefind the weak spots before anything is spent on execution. This is where an hour of your time saves the most tokens. - During the run, only when needed. A stage that needs a person says so:
WaitingForInputwhen an agent asks a question,NeedsHumanReviewwhen it is escalated, a criteria dispute when an agent believes a check is wrong. Everything else runs unattended. - Whenever you choose to look. The terminal dashboard (
loom status --live) and the web dashboard (loom status --web) show every stage, what it is doing, which models it is running, and how much context it has used. You can open any session, from tmux withloom attachor from a terminal in the browser, and watch it or take the keyboard and steer it.
What you hand over comes back checked. Loom verifies every stage's work before it merges, then the plan's integration-verify stage reviews and tests everything together, so stages that each pass on their own still have to work as one system.
Steering is also teaching. Agents record a human correction as a memory entry before they do anything else, so what you tell one session is distilled into the knowledge base and reaches every later one.
Details: Human-in-the-loop where it matters, Web Dashboard, Terminal Backends.
Most agent setups start every session from zero. The same architecture is rediscovered, the same wrong turn is taken, and the same correction is given again. Loom keeps what each session learns and gives it to the next one.
- Every session keeps a journal. As they work, agents record mistakes with a prevention rule, decisions with their rationale, and whatever surprised them about the codebase (
loom memory note,decision,change,question). - Every plan ends by distilling it. The
knowledge-distillstage reads every stage's journal and curates it intodoc/loom/knowledge/: architecture, entry points, patterns, conventions, stack, concerns, and a mistakes file where each entry says what happened, why, how to prevent it, and how it was fixed. - Every later session starts from it. A stage's signal carries a Knowledge Brief: the sections retrieval judged relevant to that stage, already quoted. Agents pull anything more with
loom knowledge context --query, and look code up in a source graph (loom map) before they open files. - Wrong knowledge gets corrected. When the code contradicts the knowledge base, the code wins. The agent records the contradiction and the next distillation applies it.
The knowledge base is plain markdown in your repository, so it is reviewed, versioned, and shared with your team like any other file. It is tiered, so it can keep growing while each session loads only the part it needs.
The effect compounds. The first plan on a project pays for discovery. After a few plans the knowledge base holds the project's structure, its conventions, and the mistakes already made on it, and later plans make far fewer of them: fewer failed verifications, fewer retries, fewer tokens.
Details: Knowledge System.
flowchart LR
A[Plan] --> B[loom init]
B --> C[Stages run in parallel worktrees]
C --> D[Loom verifies]
D --> E[Merge]
E --> I[Integration verify]
I --> F[Distill knowledge]
F -. read first by the next plan .-> A
A plan is a markdown file holding a list of stages, each with its dependencies and the commands that prove it is done.
- Init.
loom init <plan-path>parses the plan and creates the state for every stage under.loom/work/. - Schedule.
loom runstarts a daemon and an orchestrator. Every stage whose dependencies are met gets its own worktree (.worktrees/<stage-id>, branchloom/<stage-id>) and its own Claude Code session, briefed by a signal file: the assignment, the relevant knowledge, and the memory of earlier sessions. - Work. The session's main agent decomposes the stage, delegates implementation to cheaper subagents, and records what it learns with
loom memory. - Verify.
loom stage completeruns the stage's acceptance criteria and the goal-backward checks. A failure leaves the stageExecuting, and the agent has to fix the work and try again. - Merge. A verified stage merges back to the target branch, which frees the stages that depend on it. A real conflict gets a dedicated resolution session.
- Integrate. Once the implementation stages have merged, an
integration-verifystage reviews the combined work and runs the full suite against it. - Distill. The plan's final
knowledge-distillstage curates every stage's memory intodoc/loom/knowledge/, which the next plan's sessions read first.
You follow along with loom status --live or loom status --web, and step in only when a stage asks for a person: recover, verify, merge, or retry stages as needed with the stage commands.
WaitingForDeps → Queued → Executing → Completed
Everything else is an explicit, inspectable outcome rather than a hang:
| State | Meaning |
|---|---|
Blocked |
A before_stage check or an explicit block stopped the stage |
NeedsHandoff |
Context ceiling reached; a handoff was written |
WaitingForInput |
The agent asked a question (raised automatically by the AskUser hooks) |
MergeConflict |
Auto-merge hit a real conflict; a resolution session is spawned |
MergeBlocked |
Merge cannot proceed (e.g. another merge is in progress) |
CompletedWithFailures |
Work finished but acceptance did not pass |
NeedsHumanReview |
Escalated to a person |
NeedsAdjudication |
A disputed acceptance criterion is awaiting a verdict |
Skipped |
Explicitly skipped |
Everything loom does, one line each, grouped by what you are doing at the time. Each line links to the section that documents it.
Plan
- Draft a plan interactively with the
/loom-plan-writerskill. (Write a plan) - Validate a plan with no side effects: schema, sandbox limits, and worker-ownership overlaps. (Plan Commands)
- Harden a plan through adversarial review rounds from two model families. (Primary Commands)
- Give each stage a type:
knowledge,standard,integration-verify, orknowledge-distill, each with its own verification rules. (Stage Type Behavior) - Amend a stage's criteria mid-run through an audited path that snapshots the plan and logs the change. (Disputes, Adjudication, and Amendments)
Run
- Schedule independent stages across a dependency DAG, each in its own git worktree. (Parallel execution and progressive merge)
- Merge finished stages back progressively, spawning a dedicated session on a real conflict. (Parallel execution and progressive merge)
- Spawn sessions in a native terminal window or a headless tmux backend for SSH and WSL. (Terminal Backends)
- Delegate implementation to Claude or Codex subagent lanes, chosen per stage. (Model Allocation)
- Use subagents for concrete file-level tasks and an agent team for work that needs discussion across agents. (Agent Teams (Experimental))
- Write a context handoff before a session's token ceiling forces a restart. (Other Commands)
- Detect crashed or hung sessions and retry or escalate them automatically. (Crash recovery and liveness)
- Complete a stage only on an attested evidence record that survives a handoff. (Crash recovery and liveness)
Verify
- Run every acceptance criterion itself before letting a stage finish. (Verification Is Enforced, Not Self-Reported)
- Check that the outcome actually exists: real artifacts, live wiring, passing wiring tests, no dead code. (Verification Model)
- Gate a stage before spawn and after acceptance with
before_stage/after_stagechecks. (Verification Model) - Capture a baseline when a stage starts and charge the stage only for failures it introduced. (Change Impact)
- Require a regression test on any stage marked as a bug fix. (Bug-Fix Stages)
- Let an agent dispute a wrong criterion instead of weakening it, through adjudication to a verdict. (Disputes, Adjudication, and Amendments)
- Reserve the verification bypass flags for a one-time, operator-only proof. (Verification Is Enforced, Not Self-Reported)
- Skip re-running a criterion whose pass is already cached for an identical tree. (Verification Is Enforced, Not Self-Reported)
Learn
- Journal notes, decisions, and mistakes to a per-stage memory as work happens. (Knowledge System)
- Distill every stage's memory into permanent, tiered knowledge at the end of a plan. (Knowledge System)
- Pull a token-budgeted, quoted context pack for one specific question on demand. (Knowledge System)
- Query a source graph of the code's own symbols for outlines, impact, and callers. (Knowledge / Memory)
Observe
- Follow the run in a live terminal dashboard, one row per stage with its state, models, activity, context use, and merge status. (Primary Commands)
- Watch the same ledger and a dependency graph in the browser, with settings and remote access. (Web Dashboard)
- Take control of a stage's live session from a browser terminal. (Web Dashboard Terminals)
- See Claude and Codex quota bars with reset countdowns wherever the ledger renders. (Web Dashboard)
- Report what a run actually spent per provider, and compare a policy change against paired runs. (Other Commands)
- List and attach to any live session by stage or session ID. (Other Commands)
Contain
- Confine each session's filesystem reads/writes and network domains by plan and per-stage rules. (Sandbox Configuration)
- Rebuild a minimal, allowlisted environment for every command loom runs from your plan. (Command Confinement)
- Deny a session write access to loom's own state, hooks, and config, relaying the rare legitimate exception through the daemon. (Session State Confinement)
- Enforce commit discipline, worktree boundaries, and subagent limits from shell hooks. (Deterministic guardrails)
- Run stages in
autopermission mode by default and tighten it per plan or per stage;bypass-permissionsis rejected. (Permission Mode) - Enable Claude Code remote control on spawned sessions automatically when its prerequisites are met. (Remote Control)
Operate
- Hold, release, skip, retry, reset, or hand a stuck stage to a human. (Stage Commands)
- Diagnose and repair a broken install or missing hook wiring. (Other Commands)
- Clear worktrees, sessions, or state selectively without touching the rest. (Other Commands)
- Check for and install a newer loom release. (Other Commands)
- Install shell tab-completion for bash, zsh, or fish. (Shell Completions)
Spend
- Run each stage's main agent at its stage type's configured model and effort. (Model Allocation)
- Override the model or effort for one stage explicitly, without touching the rest of the plan. (Stage Fields)
- Keep each signal's prefix byte-identical across sessions, so the large doctrine block is a cache hit. (Cost control by construction)
- Inject at most 5 matched skills per stage out of 73 installed; 10 core skills are always loaded and the other 63 load on demand. (Cost control by construction)
- Why Loom
- Quick Start
- Human Expertise Where It Matters
- It Learns From Every Plan
- How It Works
- Feature Tour
- What Loom Solves
- Key Capabilities
- Installation
- CLI Reference
- Configuration
- Web Dashboard
- Plan Format
- Verification Model
- Knowledge System
- Model Allocation
- Sandbox Configuration
- Terminal Backends
- Agent Teams (Experimental)
- State Layout
- Shell Completions
- License
Autonomous agent work fails in a small number of predictable ways. Loom answers each with a mechanism, not a paragraph of prompt.
| Failure mode | What actually happens | Loom's answer |
|---|---|---|
| False completion | Tests were never run, the module was written but never imported, the fix is a TODO |
Loom runs the acceptance criteria itself, then checks artifacts for stubs, wiring for real integration, and dead-code patterns for orphaned work. The bypass flags need a token the agent cannot read. |
| Instruction drift | Rules decay the moment they scroll out of attention | Shell hooks enforce the rules that matter deterministically — commit discipline, staging scope, worktree boundaries, subagent limits — outside the model's control. |
| Amnesia | Every session rediscovers the same architecture and repeats the same mistakes | A per-stage memory journal feeds a distillation stage that curates permanent, tiered knowledge; later sessions read it before touching code. |
| Cost scaling with tokens, not value | Expensive models doing cheap work; re-reading everything, every time | Judgment stays on an orchestrator; bulk implementation is delegated to cheap subagents. Signals are laid out for KV-cache reuse and knowledge is tiered, so agents load only what they need. |
| Context exhaustion | The session degrades into an expensive compaction loop | Context budgets are monitored per stage; a handoff is written before compaction and the resumed session is re-anchored to its assignment. |
| Lost runs | A crashed or hung session takes the work with it | All state is files under .loom/work/. The daemon detects dead and hung sessions, classifies the failure, and retries or escalates. |
| Serialization | Multi-stage work runs one-at-a-time, or collides on the same files | A dependency DAG schedules independent stages concurrently in separate worktrees, with progressive auto-merge and dedicated conflict-resolution sessions. |
The rules that matter are not left to the model. Loom installs Claude Code hooks, a Codex-native hook subset, and a git pre-commit hook that fire regardless of what an agent intends:
commit-guard.shblocks a session from ending with uncommitted work or a stage stillExecutinggit-add-guard.shblocksgit add -A/git add .;git-pre-commit-hook.shblocks commits containing.loom/workor.worktreesworktree-isolation.sh/worktree-file-guard.shblock cross-worktree writes, reads, and path traversalcommit-filter.shblocks subagent git operations (a subagent commit loses the main agent's work) and blocks AI attribution in commit messagessubagent-verify-guard.shblocks subagents from running project-wide build/test/lint suites, so verification stays with the one agent that can see the whole tree — withintegration-verifystages carved out, and no opt-out environment variablepre-compact.shblocks compaction, writes a handoff, then allows it;session-start.shre-anchors the resumed agent to its signal fileplans-path-guard.shkeeps plans indoc/plans/where loom and git can see them
Subagent detection is a live process-tree ancestry check, not a PPID comparison. See Verification Is the Main Agent's Job.
loom stage complete is not a self-report. Loom executes the stage's acceptance criteria in-process and refuses completion on failure, leaving the stage Executing so the agent must fix and retry. On top of that, goal-backward verification asks whether the outcome exists:
artifacts— files exist and contain real implementation (stub detection rejectsTODO,FIXME,unimplemented!,todo!,pass,NotImplementedError)wiring— regex proof that new code is actually referenced: module registered, route mounted, component renderedwiring_tests— runtime commands proving the integration behavesdead_code_check— command output patterns catching code that exists but is never calledbefore_stage/after_stage— pre-spawn and post-acceptance gates; a failed pre-check blocks the stage before a session is even spawned
The escape hatches (--no-verify, --force-unsafe, --assume-merged) require a one-time operator proof bound to the project, stage, action, and exact flag set. The operator supplies the daemon secret only while minting the proof; the target command cannot fetch that credential for its caller or reuse the proof for another action.
Loom treats what agents learn as a durable artifact with a pipeline behind it, rather than a scratch file.
- Capture — during execution, agents record to a per-stage journal:
loom memory note(gotchas, mistakes-with-prevention),decision(with rationale),change,question. The journal is injected into the recitation section at the end of the next signal, where model attention is highest. - Distill — a
knowledge-distillstage runs at the end of a plan, reads every stage memory, and curates it into permanent knowledge — mistakes rewritten as actionable prevention rules, decisions with their rationale, reusable patterns and conventions. - Retrieve — the result is a tiered base under
doc/loom/knowledge/: a generatedINDEX.md, seven tier-1 summaries, and tier-2 topic files. Inside a stage, the per-stage Knowledge Brief comes first. Otherwise, agents readINDEX.mdfor orientation, then the tier-1 summary for their area, then only the topics they touch; a specific question is pulled withloom knowledge context --query, which returns the matching sections quoted — so the base can grow without every session paying to load it.
Knowledge lives in doc/loom/knowledge/; agents write it through loom, and loom knowledge sync rebuilds the derived retrieval artifacts after the tree changes, including the one-time flat-to-hierarchical upgrade. Details: Knowledge System.
Loom's savings come from delegation, not downgrade:
- Orchestration's model and effort come from the stage type's default — every stage's main agent plans, decomposes, verifies, and commits, the judgment-heavy work that is worst to economize on — and both are overridable, in
[models]in either config file or per stage; see Model Allocation. - Delegation is a cost decision, tokens times model tier. A stage's main agent makes a change itself only when it is at most 20 lines in at most 2 files it has already read, needs no further exploration, and one command proves it; a Fable main session delegates even those. Everything larger goes to the cheapest subagent that can do it, spawned by agent type so the choice is explicit rather than inherited: Haiku for mechanical edits such as a rename; Fable only for visual/UI design, a bug that survived a delegated fix attempt, and extremely challenging algorithmic design (no agent type pins it — the model override is stated explicitly at spawn); Opus for mainstream architecture and algorithm implementation; Sonnet or Codex GPT-5.6 Terra for common implementation and integration tests; Codex GPT-6 Luna for boilerplate, scaffolding, and simple unit tests. The codex tiers are licensed only on stages listing codex in
implementers, and additionally require thecodexCLI and its plugin to be installed — when either is missing,loom runprints an advisory warning at startup (it never aborts) and terra-/luna-tier work falls back to Sonnet. - Signals are built for cache reuse. Each signal is a four-section layout with a per-stage-type stable prefix that is byte-identical across sessions, so the large doctrine block is a cache hit rather than a re-read.
- The orchestrator's rulebook loads when it is needed. Delegation, briefs, file ownership, waiting on subagents and commit timing live in the
loom-orchestrationcore skill, which a session loads before it fans out; the installedCLAUDE.mdkeeps the hard stops and a pointer, under a 20 KB ceiling.spawn-guard.shprepends the subagent preamble to every typed spawn, so an orchestrator no longer pastes it. - Context budgets prevent compaction, which is the expensive failure: an uncached re-read that costs more and produces worse work.
- Tiered knowledge and a skill index keep the working set small — at most 5 matched skills are injected per stage, out of 73 installed.
- Waits and repeat reads are settled by receipts, not by polling. An orchestrator waits on a backgrounded Codex forward by its exact receipt (
loom subagents wait --receipt <id>), and repeatedloom subagents listpolling is counted by the poll guard. A repeated file read is warned or denied only when a transcript receipt proves the earlier result was delivered. - Consumption is measured, not assumed.
loom usagereports Claude and Codex separately from provider-native telemetry, andloom usage --comparejudges a candidate policy offline against paired runs. A token-proxy gain alone never counts as a subscription saving, and any quality or latency regression rejects the candidate; see the evaluation protocol.
Per-stage model, reasoning_effort, and ultracode fields let you override any of this explicitly.
Stages form a dependency DAG; everything independent runs at once, each in its own worktree (.worktrees/<stage-id>, branch loom/<stage-id>). Completed stages merge back progressively under a file lock, and a real conflict spawns a dedicated resolution session rather than stalling the run.
All orchestration state is plain files in .loom/work/, so nothing is lost when a process dies. The daemon polls every 5s, tracks PID liveness and per-session heartbeats, flags hung sessions after 300s, and classifies failures across ten types into retryable (exponential backoff) and needs-diagnosis. Tool-call telemetry drives a stuck-session signal when a session's recent calls are overwhelmingly failures. Orphaned sessions are recovered on daemon restart.
A stage completes only on an authenticated evidence record that survives handoffs, so a session that finished its work but died before the daemon observed it does not loop between retries. Each session records why it exited — completed, crashed, context ceiling, stalled, operator stop, criteria blocked, or replaced — and loom status, loom status --live, and the web dashboard surface a completion-pending or blocked state with that reason ahead of generic activity.
Plan-level defaults and per-stage overrides control filesystem reads/writes, network domains, and permission mode for the agent session, and commands loom runs from your plan get a rebuilt, allowlisted environment so they cannot read ambient credentials (Sandbox Configuration). Before you spend anything, loom plan verify validates a plan with no side effects — running the same sandbox validation that would otherwise only fail at loom init — and loom pressure hardens it through adversarial review rounds run by two different model families. A sandboxed session's own writes are confined the same way: it cannot touch loom's state, hooks, or config directly, and the handful of legitimate exceptions go through a one-shot request the daemon applies (Session State Confinement).
Thirteen stage states make "needs a person" an explicit outcome rather than a hang: WaitingForInput (raised automatically when an agent asks a question, or when a v2 contract freeze is refused), NeedsHumanReview, Blocked, MergeConflict. Operators get loom stage hold/release/skip/retry/human-review, and an agent that believes a criterion is wrong can escalate with loom stage dispute-criteria instead of quietly weakening it.
Loom is under active development. Signed binaries are published for Linux x86_64 and macOS Apple Silicon; every other platform builds from source with the Rust toolchain installed.
| Tool | Needed for | Required? |
|---|---|---|
Rust toolchain (cargo) |
building the loom binary |
only when building from source |
git |
worktrees, merges, crash reports | yes |
claude (Claude Code CLI) |
every orchestrated session | yes |
jq |
every loom hook parses the Claude Code hook payload with it | yes; install.sh and loom run refuse to proceed without it, loom repair reports it |
rg (ripgrep) and fd |
the installed CLAUDE.md steers agents to these over grep/find |
recommended; install.sh and loom run warn when missing |
tmux |
the tmux terminal backend | optional |
sccache |
shared dependency compiles across stage worktrees | optional |
codex CLI |
the codex implementer lane and loom pressure |
optional |
curl -fsSL https://raw.githubusercontent.com/cosmix/loom/main/install.sh | bashThis downloads the signed release binary for your platform to ~/.local/bin/loom, then installs Loom's agents, skills, commands, hooks, and orchestration rules into the default asset roots, ~/.claude/ and ~/.codex/, from the assets embedded in that binary. Codex asks you to review new or changed non-managed hooks with /hooks before they run.
Set either variable to install Loom assets outside the default root:
| Variable | Default asset root |
|---|---|
LOOM_CLAUDECODE_INSTALL_DIR |
~/.claude |
LOOM_CODEX_INSTALL_DIR |
~/.codex |
A nonempty variable replaces its matching default; an empty variable behaves as unset. Every install that passes neither --claude-dir nor --codex-dir records the roots it used in ~/.config/loom/install-roots.toml, and later installs and loom update reuse them, so the variables are needed only when choosing or changing a root. For each root, --claude-dir or --codex-dir takes precedence over its variable, then the recorded root, then the default. An install given either flag does not change the record.
Use absolute paths such as $HOME/.local/share/loom/claude:
export LOOM_CLAUDECODE_INSTALL_DIR="$HOME/.local/share/loom/claude"
export LOOM_CODEX_INSTALL_DIR="$HOME/.local/share/loom/codex"
curl -fsSL https://raw.githubusercontent.com/cosmix/loom/main/install.sh | bashA later loom update in a shell without these variables refreshes the same roots.
This selects Loom's installation destinations only. It does not configure Claude Code or Codex, or relocate Loom's other runtime discovery paths.
Required on Linux ARM64, and what you want when working on loom itself:
git clone https://github.com/cosmix/loom.git
cd loom
bash ./dev-install.shdev-install.sh builds the release binary (cargo build --release) and runs install.sh, which installs loom-* prefixed agents and skills (non-destructively, preserving user customizations), hooks, and configuration into the default ~/.claude/ root and the CLI binary to ~/.local/bin/loom. Orchestration rules are written directly to the default ~/.claude/CLAUDE.md (existing file is backed up). It inherits the custom asset-root variables above when they are exported.
install.sh takes an optional --skills core|all flag (default core): core installs a small set of always-loaded core skills to ~/.claude/skills/ and catalogs the rest under ~/.claude/loom-skill-catalog/, loaded on demand; all installs every loom skill directly to ~/.claude/skills/.
loom init installs and configures project hook wiring automatically. For an existing repo that is missing Claude Code hook setup, run loom repair --fix.
- Linux x86_64: primary development and full CI test runs; signed release binary
- macOS (Apple Silicon): supported for build/terminal integration, CI does build-only verification; signed release binary
- macOS (Intel): builds from source; no release binary is published
- Windows via WSL2: supported — WSL2 runs the Linux x86_64 binary unmodified. Use the tmux backend, since a stock WSL install has no GUI terminal emulator for the native backend to find — see Running under WSL. Native Windows, outside WSL, is unsupported.
- Linux ARM64: builds from source; no release binary is published yet
- Headless (SSH, no terminal emulator): supported via the tmux backend — see Terminal Backends
| Default location | Contents |
|---|---|
~/.claude/agents/loom-*.md |
5 specialized subagents (per-item, non-destructive) |
~/.claude/skills/loom-*/ |
10 core domain knowledge modules, always loaded (per-item, non-destructive) |
~/.claude/loom-skill-catalog/loom-*/ |
63 more domain knowledge modules, loaded on demand (--skills core, the default) |
~/.claude/commands/*.md |
Loom slash commands (/pressure, /address, /distill) |
~/.claude/hooks/loom/ |
Embedded lifecycle and guardrail hooks + shared libraries |
~/.claude/CLAUDE.md |
Orchestration rules |
~/.codex/skills/pressure/ |
Codex pressure-testing skill ($pressure) |
~/.codex/hooks/loom/ |
Loom hook assets used by Codex-native registrations |
~/.codex/hooks.json |
Non-destructively merged Codex hook registrations |
~/.codex/AGENTS.md |
Codex navigation and execution doctrine |
~/.local/bin/loom |
Loom CLI |
loom init <plan-path> [--clean] [--backend native|tmux]
loom run [--manual] [--max-parallel N] [--foreground] [--watch] [--no-merge] [--backend native|tmux]
loom status [--live] [--compact] [--verbose] [--web [PORT] [--host HOST] [--terminals]]
loom stop
loom resume <stage-id>
loom check <stage-id> [--suggest] [--no-cache]
loom pressure <plan-path> [--rounds N] [--claude-model M] [--claude-effort E] [--codex-model M] [--codex-effort E] [--address-model M] [--address-effort E] [--dry-run]When loom run stops — every stage settled, or interrupted by loom stop — it prints stages by outcome: Completed; Failed (Blocked, MergeConflict, CompletedWithFailures, MergeBlocked, NeedsHumanReview — terminal, needing intervention); Unfinished (WaitingForDeps, Queued, Executing, WaitingForInput, NeedsAdjudication — still in progress when the run stopped); and Needs Handoff. The run succeeds only when Failed, Unfinished, and Needs Handoff are all empty.
loom pressure hardens a plan before you run it by combining two external agents over --rounds rounds (default 2). Each round runs both pressure-tests in parallel: Claude /pressure edits the plan in place in the foreground (you watch it live), while Codex $pressure writes an independent review next to it (codex-<plan>.md) in the background (its output is captured to a temp log to keep the terminal clean). Once both finish, Claude /address folds the review back in. Claude stays interactive and auto-closes when done; Codex runs from the repo root. Requires both the claude and codex CLIs on PATH. --dry-run prints the exact commands without spawning anything.
Each of the three steps spawns with an independently selectable model and reasoning effort: --claude-model/--claude-effort for /pressure (model accepts haiku, sonnet, opus, or fable; effort accepts low, medium, high, xhigh, or max), --address-model/--address-effort for /address (same value sets), and --codex-model/--codex-effort for $pressure (model accepts gpt-6-astra, gpt-6.1-sol, gpt-6-sol, gpt-5.6-terra, or gpt-6-luna; effort accepts low, medium, high, or xhigh — no max, that value is Claude-only). Absent a flag, each key falls back independently: the project config's .loom/work/config.toml [pressure] section first (if it sets that key), then ~/.loom/config.toml (loom config -k pressure.claude_model <value>), then its built-in default. [pressure] resolves per key, so a project section that sets only claude_model still lets codex_effort fall through to your user config.
When the codex model is unavailable to the account (codex rejects it as not supported or not found), loom pressure retries with the same family's nearest newer version first and then the nearest older one, never another tier; the default gpt-6.1-sol falls back to gpt-6-sol. The fallback notice prints after the Claude session ends. When every model in the family is unavailable the run fails and shows the codex error.
| Key | Default |
|---|---|
pressure.claude_model |
opus |
pressure.claude_effort |
xhigh |
pressure.codex_model |
gpt-6.1-sol |
pressure.codex_effort |
xhigh |
pressure.address_model |
opus |
pressure.address_effort |
high |
loom status --live renders a live ledger dashboard: one row per stage across eight columns (STATE, STAGE, DEPENDS ON, MODELS, ACTIVITY, CONTEXT, TIME, MERGE). MODELS lists the orchestrator's own model first, then the models any subagents it spawned ran on. Columns drop in priority order as the terminal narrows; below a 64x16 (columns x rows) terminal a notice replaces the dashboard entirely. Press ? to toggle a legend overlay explaining every state icon. When Claude or Codex quota data is available, a footer line above the legend shows a percent bar and reset countdown for each.
loom status --web [PORT] [--host HOST] [--terminals] serves the same ledger in the browser, alongside a dependency-graph view of the plan — see Web Dashboard.
loom plan verify <plan-path> [--strict] [--json] [--no-color]loom plan verify validates a plan file without touching .loom/work/ or requiring a git repo. It runs the same fatal validation as loom init (schema errors, unknown or retired fields at every nested policy layer, duplicate IDs, unknown dependencies, path safety) plus advisory warnings (structural issues, a stage without a summary, missing knowledge-bootstrap stage, sandbox gaps). A retired top-level truths block is rejected; move behavioral commands to acceptance. It additionally rejects requirements a sandboxed stage cannot satisfy — an allow_write grant under /tmp or missing on the host, a TMPDIR=<absolute path> override, a hardcoded /tmp/ path, and a mkdir, touch or output redirect aimed outside the worktree in acceptance, setup, wiring_tests, before_stage or after_stage commands — because the stage sandbox already provides a writable $TMPDIR. It rejects commands that cannot mean what they say inside a sandbox — an exit status masked by a final || true, || :, || exit 0 or ; true, HOME assigned from a variable or substitution, and a bare mktemp -d — and warns on network tools, reads of doc/plans/, rg -r, and hardcoded /tmp/ paths. A single rg or grep acceptance criterion that already passes at HEAD earns a baseline warning, since it cannot tell a stage that did its work from one that did nothing. A stage's skills: list is checked against the skill index: an empty, duplicate or unknown name is an error. It also warns when a stage description's worker/files-owned table claims overlapping paths between workers, or a claim outside the stage's declared files. Exits 0 on success, non-zero on fatal errors; --strict promotes warnings to errors.
loom stage complete <stage-id> [--session <id>] [--no-verify] [--force-unsafe --assume-merged] [--no-cache]
loom stage block <stage-id> <reason>
loom stage reset <stage-id> [--hard] [--kill-session]
loom stage waiting <stage-id>
loom stage resume <stage-id>
loom stage hold <stage-id>
loom stage release <stage-id>
loom stage skip <stage-id> [--reason <text>]
loom stage retry <stage-id> [--force] [--context <message>]
loom stage merge [stage-id] [--resolved]
loom stage human-review <stage-id> [--approve|--force-complete|--reject <reason>]
loom stage dispute-criteria <stage-id> --criterion-index N --reason <text> [--evidence-commit <sha>] [--failure-output <path>]
loom stage dispute-findings <stage-id> --finding <id>... --reason <text> # Plan version 2: challenge open review findings
loom stage dispute-contract <stage-id> --contract <id> --reason <text> # Plan version 2: challenge a frozen contract
loom stage dispute-integrity <stage-id> --event <id>... --reason <text> # Plan version 2: challenge a test-integrity event
loom stage adjudicate --stage <stage-id> --dispute <n> --verdict-file <path>
loom stage contracts freeze|show|restore <stage-id> [--contract <id>] # Plan version 2: freeze (contract session), inspect, or restore frozen contract files
loom stage review status|integrity <stage-id> # Plan version 2: review rounds and open findings; test-integrity events
loom stage admin-proof [stage-id] [--daemon-stop] [--no-verify] [--force-unsafe] [--assume-merged]
loom stage amend <stage-id> --field acceptance|wiring|wiring-tests|contracts --op replace|insert|delete --index N [--value <yaml>] [--reason <text>]loom stage dispute-criteria is the sanctioned way for an agent to challenge a criterion it believes is wrong or impossible, instead of quietly weakening it. The daemon writes request.md and moves the stage to NeedsAdjudication; the verdict is daemon-written and never authored by the agent. loom stage adjudicate records that verdict from the adjudication session; loom stage admin-proof mints an operator's proof for a trusted broker; loom stage amend is the audited, operator-facing way to edit a stage's acceptance, wiring, or wiring_tests array directly. See Disputes, Adjudication, and Amendments.
loom stage output set <stage-id> <key> <value> [--description <text>]
loom stage output get <stage-id> <key>
loom stage output list <stage-id>
loom stage output remove <stage-id> <key>loom map [--outline <path>] [--find-all <symbol>] [--impact <symbol|path>] [--callers <symbol>] [--callees <symbol>] [--json]
# Query the derived source graph: file outlines, symbol lookup, impact/caller/callee analysis
loom knowledge context --query <text> [--stage <id>] [--budget-tokens <n>] [--explain] [--json] # Token-budgeted context pack for a question
loom knowledge update <file> [content] # Append a section to a tier-1 file or tier-2 topic (<category>/<slug>)
loom knowledge replace-section <file> <heading> [content] # Rewrite one section's body in place, at whatever level it's found
loom knowledge delete-section <file> <heading> # Remove a section and its nested subsections (a heading rename is delete-section then update)
loom knowledge annotate <target> [--state <s>] [--section <heading>] [--source <path>]... [--verified <rev|HEAD>] [--alias <name>]... [--blurb <text>]
# Set lifecycle state (whole file, or one `##` section), evidence sources, verified revision, aliases, or a topic's blurb
loom knowledge telemetry [--stage <id>] [--json] # Summarise recorded context delivery, prompt briefs, abstentions and pulls
loom knowledge sync [--structural-only] [--json] # Rebuild derived retrieval artifacts after editing knowledge
loom knowledge check [--strict] [--strict-evidence] [--baseline <file>] [--json] # Report knowledge-base diagnostics (read-only; never opens the context store); --strict-evidence also fails on changed or unassessable declared sources; --baseline makes --strict fail only on structural issues the file does not record
loom knowledge check --write-baseline <file> # Record every current structural issue in <file> and exit 0, to adopt the size limits on a tree that cannot clear them yet
loom memory note <text> [--stage <id>]
loom memory decision <text> [--context <why>] [--stage <id>]
loom memory change <text> [--stage <id>]
loom memory question <text> [--stage <id>]
loom memory query <search> [--stage <id>]
loom memory list [--stage <id>] [--entry-type <type>]
loom memory show [--stage <id>] [--all]
loom memory resolve <event-id> --outcome <promoted|merged|discarded|deferred> [--target <file#heading>] [--reason <text>]
# Record how a captured note/decision/question was processed
loom memory pending [--stage <id>] [--group] [--strict] [--json] # List notes, decisions, and questions that have no receipt; --group prints corrections (by target file and heading), mistakes, decisions, otherA plan's .loom/work/ state is archived to <main>/.loom/memory/archive/<plan-id>-<timestamp>/ when the plan completes; loom clean/loom init --clean are what actually delete it.
See Knowledge System for how these fit together.
loom review [--ai-summary] # Generate a code-review doc from stage memories; --ai-summary uses headless `claude -p`
loom usage [--since <duration|date>] [--until <rfc3339>] [--provider claude|codex|all] [--project <path> | --all] [--stage <id>] [--plan <name>] [--windows 5h|week] [--json]
# Report what agent sessions actually consumed, per provider (Claude and Codex are never summed), including peak resident context by scope in a `peaks` section (main vs subagent p50/p90/max, share above 250k and 400k tokens)
loom usage [--claude-root <dir>] [--codex-root <dir>] [--receipts-root <dir>] [--forward-receipts-root <dir>]
# Read explicit telemetry roots; a supplied root never falls back
loom usage --compare <artifact.json> [--json] # Offline paired evaluation: exit 0 supported, 1 rejected, 2 inconclusive (doc/token-optimization-evaluation.md)
loom subagents list | harvest [--id <id>] [--json] # One-shot diagnostics only: list shows liveness, harvest prints a terminal report (optionally one agent's); watch/wait below are the blocking waits
loom subagents watch --worker <kind>:<id> [--worker <kind>:<id> ...] [--session <claude-parent-uuid>] --timeout <secs> [--json]
# One owned wait, inside a loom stage only (needs LOOM_STAGE_ID/LOOM_SESSION_ID/LOOM_WORK_DIR): exit 0 all bound workers have fresh correlated success; 2 deadline passed; 3 worker failed/cancelled
# kind is claude (spawned agent ID) or codex (--unit-id); exit 4 AlreadyWaiting/Busy (no second monitor)
# Exit 5 identity/evidence unknown; --dir or no --worker is rejected with migration guidance
loom subagents wait --receipt <id> [--timeout <secs>] [--json] # Wait on one exact forwarded Codex job; exit 0 succeeded, 1 failed/canceled, 2 still running/unknown
loom attach [stage-id] # tmux backend only; omit the id for a tiled overview
loom sessions list
loom sessions kill <session-id...> | --stage <stage-id>
loom handoff [--stage <id>] [--session <id>] [--trigger <type>] [--message <text>]
# Capture session state to a handoff file; stage and session default to $LOOM_STAGE_ID/$LOOM_SESSION_ID, --trigger defaults to manual
loom worktree list
loom worktree remove <stage-id>
loom graph
loom project detect [PATH] [--json] # Per package: language kinds, the test runner loom would use (or unsupported), and the matching skills
loom context record-edit --stage <id> --path <path> [--path <path>...] # Keep a stage's context overlay current
loom hook user-prompt # UserPromptSubmit entry point; invoked by loom's hooks
loom request status <id> [--session <id>] # Plumbing: has the daemon applied a request relayed through the sandbox? <id> is printed after the originating command
loom skill-index # Plumbing: rebuild the skill keyword index the skill-trigger hook reads
loom repair [--fix]
loom clean [--all|--worktrees|--sessions|--state]
loom update
loom config [-k <key> [<value>] | --list | --print] # Read or write ~/.loom/config.toml; bare in a terminal it opens the settings screen (see Configuration)
loom install-assets [--claude-dir <path>] [--codex-dir <path>] [--skills core|all] # Install loom's agents, skills, commands, hooks and doctrine files; see Install script below for what --skills core|all installs where
loom completions [<shell>] [--install] [--migrate]loom handoff writes the document a successor session resumes from, under .loom/work/handoffs/. Loom's pre-compact hook calls it automatically before a compaction, and an agent that reaches its context ceiling calls it explicitly with --trigger ceiling.
loom request status and loom skill-index are plumbing for loom's own hooks and its sandbox relay rather than commands a plan author types; they are listed so hook output that names them is traceable.
Loom keeps its settings in two TOML files. A project file overrides the user file, and an explicit per-invocation value (a stage's own model field in the plan, a loom pressure flag) overrides both:
| Tier | File | Applies to |
|---|---|---|
user |
~/.loom/config.toml |
Every workspace on this machine |
project |
<repo>/.loom/work/config.toml |
This workspace only |
Neither file needs to exist: every key has a built-in default. LOOM_HOME relocates the user file ($LOOM_HOME/config.toml).
There are three ways to change a setting:
loom config. Run bare in a terminal it opens a settings screen; with flags it is scriptable.loom config -k <key>prints one key,loom config -k <key> <value>writes it (validated against the key's type and value set),loom config --listprints every key with its value and where it came from, andloom config --printprints the resolved user config as TOML. It reads and writes the user file only. The settings screen edits each row by its type:↑↓/k/jmove,←→/h/l/space cycle an enum's variants or toggle a bool,Enteropens a text editor for a number or free-text key (and steps a bool/enum forward otherwise, so no keystroke opens a text field on a closed vocabulary),ssaves,Esc/qquits,*marks a pending change.- Edit the files. Both files use the same
[section]/key = valuelayout as the table below; project sections may be partial.loom initwrites a[context]section into the project file, everything else is opt-in. - The dashboard.
loom status --webhas a settings dialog that edits both files, one key at a time, with the resolution shown per key (Web Dashboard Settings).
The keys, with their built-in defaults:
| Key | Default | Values | Project tier |
|---|---|---|---|
update.check |
true |
true, false |
no |
update.check_interval_hours |
24 |
integer | no |
terminal.backend |
native |
native, tmux |
whole section |
context.ceiling_tokens |
800000 |
integer | whole section |
pressure.claude_model |
opus |
haiku, sonnet, opus, fable |
per key |
pressure.claude_effort |
xhigh |
low, medium, high, xhigh, max |
per key |
pressure.codex_model |
gpt-6.1-sol |
gpt-6-astra, gpt-6.1-sol, gpt-6-sol, gpt-5.6-terra, gpt-6-luna |
per key |
pressure.codex_effort |
xhigh |
low, medium, high, xhigh |
per key |
pressure.address_model |
opus |
haiku, sonnet, opus, fable |
per key |
pressure.address_effort |
high |
low, medium, high, xhigh, max |
per key |
models.standard_model |
opus |
haiku, sonnet, opus, fable |
per key |
models.standard_effort |
high |
low, medium, high, xhigh, max |
per key |
models.knowledge_model |
opus |
haiku, sonnet, opus, fable |
per key |
models.knowledge_effort |
medium |
low, medium, high, xhigh, max |
per key |
models.knowledge_distill_model |
sonnet |
haiku, sonnet, opus, fable |
per key |
models.knowledge_distill_effort |
high |
low, medium, high, xhigh, max |
per key |
models.integration_verify_model |
opus |
haiku, sonnet, opus, fable |
per key |
models.integration_verify_effort |
xhigh |
low, medium, high, xhigh, max |
per key |
"Whole section" means a project [terminal] or [context] section replaces the user tier's section outright, so a key it omits takes the built-in. "Per key" means a project [pressure] or [models] section overrides only the keys it names and the rest fall through to the user file. The pressure.* keys are explained under Primary Commands, the models.* keys under Model Allocation.
loom status --web [PORT] [--host HOST] [--terminals]The dashboard binds to 127.0.0.1 by default and serves the live ledger over a WebSocket in the browser; see Web Dashboard Remote Access for --host. Two views share the same data: the graph (below) draws the plan as a dependency graph colored by stage state, and /ledger puts one row per stage in a table with the same columns as loom status --live. Both carry a "needs attention" panel naming the stages that need a person, with a suggested command for each.
A daemon poller checks Claude and Codex quota roughly every 180 seconds; whenever either has data, the dashboard's sticky footer shows a percent bar and reset countdown for each, the same meter drawn above the legend in loom status --live.
The /ledger view: state, dependencies, models, activity and context for every stage.
Without PORT, it starts at port 7373 and automatically tries the next available port when a candidate is occupied. Supplying a nonzero PORT requests that exact port; PORT 0 asks the OS for any free port. It works without the daemon by polling .loom/work/ files directly when the daemon socket is unreachable. Besides the ledger, it exposes a settings dialog for editing loom's configuration and, with --terminals, a way to open a stage's terminal from the browser.
loom status --web --host 0.0.0.0 # every interface
loom status --web --host 192.168.1.5 # one concrete address
loom status --web --host :: # every interface, IPv6--host accepts an IPv4/IPv6 literal or the exact alias localhost; it rejects DNS names, URLs, and anything carrying its own port (--web already owns the port). Left unset it binds 127.0.0.1, which keeps the dashboard's original behavior exactly.
Any bind that is not loopback — a concrete address or a wildcard (0.0.0.0/::) — turns on remote mode: loom mints one process token and prints a URL containing it at startup. Every route then requires a cookie exchanged for that token by opening the printed URL once; a wrong or missing token gets refused. Each request's Host must name the connection's own accepted socket address exactly, and a supplied Origin must match it too, so a DNS name, wrong port, or a proxy that rewrites Host all get refused.
Limits:
- Plain HTTP, no TLS — anyone who observes the token or cookie on the network gets dashboard and settings access, and terminal control if
--terminalsis also set. Use it only on trusted networks, or tunnel a loopback bind instead:ssh -L 7373:127.0.0.1:7373 host. - No reverse-proxy support: a proxy that rewrites
Hostto a DNS name is refused. - No DNS hostnames for
--host, and no config key for the bind host — it is a CLI flag only. - The token is per process; restarting the dashboard invalidates every existing cookie.
The dashboard header has a settings button; opening it (or navigating to ?settings=1) edits loom's configuration in place, so the browser's back button closes the dialog. Each row shows all three tiers — built-in, user, project — side by side as table columns, or as one card per row below 700px wide. It edits every key in the registry loom config does — eighteen keys across update, terminal, context, pressure, and models: update.check / update.check_interval_hours (loom's self-update check), terminal.backend (see Terminal Backends), context.ceiling_tokens (see Plan-Level Context Fields), the six pressure.* model and effort picks described under Primary Commands, and the eight models.* model and effort picks described under Model Allocation. loom config --list prints every key with its current value and origin. Every control here writes immediately on change, one key at a time; there is no separate Save step, and validation errors from the server surface next to the control that triggered them.
Above that config table sits a dashboard section that never touches those files — two preferences kept entirely in the browser's localStorage: a theme picker (Ledger, Aubergine, Pacific, Graphite, each drawn as a miniature of the stage graph in that theme's own colors) and a desktop-notification switch that raises an OS notification when a stage needs a human, needs a handoff, or the whole run finishes. The section renders even when /api/config is unreachable, and stays visible under any filter that matches it. Desktop notifications need a secure context — https, localhost, or a loopback address — so the switch disables itself and explains why when the dashboard is served over plain HTTP to a non-loopback host.
Two scopes:
| Scope | File | Applies to |
|---|---|---|
user |
~/.loom/config.toml |
Every workspace on this machine |
project |
<repo>/.loom/work/config.toml |
This workspace only |
Sixteen of the eighteen keys have a project tier: terminal.backend, context.ceiling_tokens, and every pressure.* and models.* key; the dialog marks the remaining update.* keys as machine-wide rather than offering a project control that would do nothing. A key's effective value resolves project → user → built-in default, and each row shows which tier is currently in force plus what clearing an override would fall back to.
The fallback is not uniform across sections. For [pressure] and [models], the project tier resolves per key — a project section that sets only one key still lets every other key in that section fall through to your user config. For [terminal] and [context], a project override still replaces the whole .loom/work/config.toml section, not just one key: clearing the override removes the key and, if that empties the section, the section too — but if a sibling key is still in there (as with [context], since loom init writes ceiling_tokens alongside subagent_ceiling_tokens), the section still wins as a whole and the value resolves to the built-in default rather than falling through to your user setting. The dialog reports each case accurately; for [terminal]/[context] that section-level fallback just may not be what you expected.
On the default loopback bind, the dashboard stays an unauthenticated tool for the person running it — the settings endpoint adds no login. On a remote (--host) bind, every route including settings requires the startup-token cookie in addition to the checks below. Writes are gated by the same Host check as the rest of the dashboard, plus a strict Origin check (loopback: must be present and loopback; remote: must equal the connection's own address exactly) and a per-process CSRF token issued on load.
loom status --web --terminals (tmux backend only) adds a "Take control" view to each stage's detail dialog: an xterm.js terminal opens right in the browser, attached to that stage's live tmux session. It starts in a read-only View mode; switching to Control sends every keystroke to the running agent. Enabling --terminals mints a one-time token and prints it in the startup URL — open that exact link once to set an auth cookie for the dashboard; a plain --web link never gets the "Take control" option. With a non-loopback --host, the same startup token already covers the dashboard cookie; the cookie alone still never grants "Take control" — that capability comes only from passing --terminals.
A stage's terminal, opened from its detail dialog: live output, viewable read-only or with control handed over.
Plans live in doc/plans/ with metadata in fenced YAML between loom markers.
# PLAN-0001: Feature Name
<!-- loom METADATA -->
```yaml
loom:
version: 1
sandbox:
enabled: true
stages:
- id: implement-api
name: Implement API
summary: Adds the endpoint and its tests.
description: Add endpoint + tests
working_dir: "."
stage_type: standard
dependencies: []
acceptance:
- "cargo test"
- command: "cargo test api_integration::returns_200"
stdout_contains: ["test result: ok"]
files:
- "loom/src/**/*.rs"
artifacts:
- "loom/src/api/*.rs"
wiring:
- source: "loom/src/main.rs"
pattern: "mod api;"
description: "API module registered"
execution_mode: team
- id: integration-verify
name: Integration Verify
working_dir: "."
stage_type: integration-verify
dependencies: ["implement-api"]
acceptance:
- "cargo test --all-targets"
- command: "cargo test api_integration::returns_200"
stdout_contains: ["test result: ok"]
```
<!-- END loom METADATA -->Set in the loom: block alongside version and stages, these supply the
default ceiling for every stage that does not declare its own. Both are
absolute resident-token counts with a minimum of 60000, and both are persisted
into .loom/work/config.toml's [context] section at loom init.
| Field | Required | Notes |
|---|---|---|
context_ceiling_tokens |
No | Default ceiling for a stage's main agent session (default 150000) |
subagent_ceiling_tokens |
No | Ceiling for subagents spawned by a stage session (default 120000); never read from a stage |
auto_merge |
No | Plan-wide default for automatic merge on completion; a stage's own auto_merge overrides it, and both fall back to the orchestrator's own setting (loom run --no-merge disables it) when unset |
change_impact |
No | Nested block comparing before/after state; see Change Impact |
| Field | Required | Notes |
|---|---|---|
id |
Yes | Stage identifier |
name |
Yes | Human-readable title |
working_dir |
Yes | Relative execution directory (. allowed) |
description |
No | Task brief for the agent: the full spec its session works from |
summary |
No | 1-3 sentences for a human on what the stage delivers; the web dashboard shows it instead of description. loom plan verify warns on a stage without one |
dependencies |
No | Upstream stage IDs |
parallel_group |
No | Optional label grouping related stages; the dependency graph alone still decides scheduling order |
auto_merge |
No | Per-stage override for automatic merge on completion; takes priority over the plan-level and orchestrator defaults |
acceptance |
Conditionally required | Shell criteria (strings or extended objects with stdout_contains etc.) |
setup |
No | Setup commands |
files |
No | File glob scope |
stage_type |
No | standard (default), knowledge, integration-verify, knowledge-distill |
artifacts / wiring |
Conditionally required | Required for standard and integration-verify (acceptance OR goal-backward) |
wiring_tests / dead_code_check |
No | Extended verification |
before_stage |
No | Pre-spawn checks (TruthCheck list); stage → Blocked if any fail |
after_stage |
No | Post-acceptance checks (TruthCheck list); completion fails if any fail |
code_review |
No | integration-verify only: dimensions (string list) and require_all (bool); rendered as checklist in agent signal |
bug_fix |
No | Marks this stage as a bug fix; requires regression_test (Bug-Fix Stages) |
regression_test |
Conditionally required | Required when bug_fix is true: file (test path, relative to working_dir) plus optional must_contain patterns |
model |
No | Model for this stage's main agent; omit to use the stage type's configured default (Model Allocation), overridable via [models] in either config file, or set here as a deliberate per-stage override |
reasoning_effort |
No | low, medium, high, xhigh, max; omit to use the stage type's configured default (Model Allocation), overridable the same way |
implementers |
No | Licensed agent lanes as a list, first = preferred for routine work: ["codex", "claude"]. Default ["claude"]. Listing a lane makes it available, not mandatory — a stage mixes lanes per subagent |
ultracode |
No | License this stage for large multi-agent fan-out; per-stage opt-in (default false) |
skills |
No | Skills this stage's agents need, by catalog name (["loom-rust", "loom-debugging"]). loom plan verify rejects an empty, duplicate or unknown name; the signal lists each as a skill to load before the detected ones, and a subagent's brief names them |
subagent_timeout_secs |
No | Seconds of tool silence before the monitor warns appears hung (default 300); the advisory idle budget death is judged against — never the --timeout on the single owned loom subagents watch --worker ... --timeout 3600 call |
context_ceiling_tokens |
No | Absolute resident-token ceiling for this stage's session (minimum 60000). Resolved stage value → plan-level context_ceiling_tokens → 150000. The session hook warns at 80% and blocks at 100%; the daemon forces a handoff at 125% |
plan_overview |
No | Set false to suppress the embedded plan overview in this stage's signal |
sandbox |
No | Per-stage sandbox override |
sandbox.permission_mode |
No | auto (default), accept-edits, plan, default — resolves stage > plan > stage-type default; bypass-permissions is rejected at init |
execution_mode |
No | single (default) or team hint |
knowledge: knowledge/bootstrap work, different verification expectationsstandard: implementation stage; must define goal-backward checksintegration-verify: final quality gate combining code review and functional verification; must define goal-backward checks. Definecode_review.dimensionsto render a checklist of review dimensions in the agent's signal.knowledge-distill: final stage; curates stage memories into permanent knowledge files
version: 2 turns on stricter verification; version: 1 keeps the rules above unchanged, and a v1 plan that uses a v2-only field is rejected (`<field>` requires `version: 2`). The v2 fields:
| Field | On | Notes |
|---|---|---|
contracts |
stage | Tests written and frozen before implementation: id, file, test, optional runner, scenario, and rejects (the plausible wrong implementation the test must fail on). Every standard stage of a v2 plan needs at least one; knowledge, knowledge-distill and integration-verify stages take none |
harness |
stage | Globs of extra test-only files the contract writer may edit. Every matched file is frozen with the contracts, so do not glob production files |
reachable |
stage | symbol, from, description, optional min_confidence: the new unit must be reachable from an entry point in the source graph |
literal: true, glob source |
wiring entry |
Match the pattern as literal text; let source be a glob (*, ?, [). A pattern that matches only a definition of a name defined in that file is rejected, so point it at a consumer |
ratchet_files |
plan | Exact paths of baseline or ledger files a stage could loosen; any change to one raises a test-integrity event |
loom:
version: 2
ratchet_files: ["loom/maintainability-baseline.txt"]
stages:
- id: implement-api
stage_type: standard
contracts:
- id: rejects-empty-name
file: "tests/api_contract.rs"
test: "rejects_empty_name"
scenario: "POST /users with an empty name"
rejects: "an implementation that stores the empty name"
reachable:
- symbol: "create_user"
from: "main"
description: "the handler is reachable from main"A v2 standard stage starts with a separate contract session that writes the contract tests, runs them red, and ends with loom stage contracts freeze; loom then hands the stage to the implementing session. loom project detect shows which test runner loom will use for each package. Runners without an adapter fall back to running the contract's test string and judging its exit code.
The contract phase is visible wherever loom status renders: a magenta contracts tag next to the stage's [model] tag (added to the legend while such a stage exists), the live TUI's activity cell reading writing contract tests or, once stale, contract writer idle <duration>, and the web dashboard's violet contracts tag with a dashed node outline, the same activity text, a ... · contract phase stage-strip label, and a contract writer row in the stage dialog's session type.
A refused freeze — a file changed outside the contracts and harness globs, a contract not reported red, or (checked by the daemon) a FIFO, socket or device node left in the worktree — parks the stage WaitingForInput with the problems as its review_reason, the reason loom status shows for it. The daemon parks the stage as soon as it refuses a relayed freeze; a refusal from the writer's own pre-check parks it when the writer stops. Typing into the writer's session resumes it, and so do loom stage resume <stage-id> and a later freeze the daemon accepts; loom stage retry still refuses a stage in WaitingForInput. A writer that stops before its first freeze attempt is not parked.
loom check <stage-id> validates outcomes, not just compilation/tests:
acceptance: shell criteria (simple strings or extended objects withstdout_contains,exit_code, etc.)artifacts: real implementation files existwiring: critical integration links existwiring_tests: runtime integration checksdead_code_check: detect unused code via command output patterns
For standard and integration-verify stages, acceptance criteria or at least one goal-backward check must be defined.
In a version 2 plan, loom stage complete also runs, in order: the contract check (every frozen contract file unchanged, every contract test passing), test integrity (fewer test declarations or assertions, removed or changed assertion lines, or a changed ratchet_files entry each raise an event that must be fixed or accepted through an adjudicated dispute), the tests reached from the changed code in the source graph, and the review gate. A criterion that runs a recognised test runner and selects zero tests fails. The gate needs a loom-code-reviewer round on the final change fingerprint with every finding fixed or ruled on; editing any file after the last round, formatting included, needs another round, while committing does not. The loom daemon that owns the worktree computes this fingerprint, once, for both the round a review records and the completion that checks it; a sandboxed loom stage complete cannot reach the daemon's socket, so it skips test integrity and the review gate locally, and the daemon runs both when it applies the completion. Integration-verify re-checks every stage's reachable and wiring on the merged tree and lists the full test command.
loom stage complete is the only way a stage finishes, and it runs the acceptance criteria itself before doing anything else. If they fail, the stage stays Executing — the agent must fix the work and re-run, and fix_attempts is incremented so repeated failures surface rather than accumulate silently. after_stage checks then run post-acceptance, and goal-backward verification runs before the progressive merge. A passing criterion is cached against an exact fingerprint of its command and the tree, so an unchanged criterion is not re-run on the next attempt; --no-cache on loom check or loom stage complete forces a real run regardless.
Artifact verification treats a stub as a failure: a file that exists but contains TODO, FIXME, unimplemented!, todo!, a bare pass, or raise NotImplementedError does not count as delivered.
Normal stage completion crosses a narrow control boundary. The stage runs one exact pinned
loom stage complete <stage-id> command; a PostToolUse bridge accepts only the matching stage and
session, requires Loom's verification marker, and sends a non-extensible CompleteStage request to
the daemon. The request cannot carry commands, paths, or bypass flags.
The three bypass flags — --no-verify, --force-unsafe, --assume-merged — are the operator's, and
cost the operator nothing: loom stage complete <stage> --no-verify just works from your shell. It
authorizes itself against .loom/work/admin.token, which you can already read and a sandboxed agent
cannot (the sandbox binds the whole process tree, so a loom an agent spawns is denied the same
read). The proof is still bound to the project, stage, action, and exact flag set, and consumed on
first use — you simply never handle it.
The same applies to loom stop. Nothing asks a human to mint a credential they already hold; making
them carry an HMAC between two commands added ceremony, not security.
loom stage admin-proof remains for the case it was actually built for: a trusted broker minting a
narrowly-scoped capability for another process. It takes the secret through LOOM_ADMIN_TOKEN and
never reads the token file, so a caller that can invoke loom but cannot read that file gains nothing.
An agent that genuinely believes a criterion is wrong or impossible has a sanctioned path — loom stage dispute-criteria — rather than an incentive to weaken it.
Subagents do not verify. A subagent may run at most one narrowly-scoped check covering the files it just changed; project-wide builds, full test suites, and repo-wide lint or typecheck runs belong to the main agent — the only party that can see the whole tree and act on the result.
This is enforced, not just advised: loom-hooks/subagent-verify-guard.sh (a PreToolUse:Bash hook) blocks project-wide runners — cargo build, cargo test, make test, tsc, go build and friends — when the caller is detected as a subagent. Scoped invocations pass, quoted mentions are ignored, and unrecognised commands are always allowed: a false block would strand a subagent mid-task.
Two things worth knowing:
integration-verifystages are carved out. That stage type exists to run the complete suite, so its subagents may. The carve-out is read from the stage file and fails safe — an ambiguous or missing stage file means no relaxation.- There is deliberately no opt-out environment variable. The main agent is never affected, so an escape hatch would only serve to defeat the rule.
A stage with bug_fix: true requires a regression_test block naming the test that pins the fix: file (path to the test, relative to working_dir) and optional must_contain (patterns the file's content must include). Verification checks the file exists and carries those patterns, so a bug-fix stage cannot complete on a fix with no test guarding the regression.
A plan-level change_impact block separates the failures a stage introduced from the ones it inherited. Loom runs baseline_command when a stage starts executing and saves the output; at loom stage complete it runs the command again and compares the lines matching failure_patterns:
loom:
version: 1
change_impact:
baseline_command: "cargo test --no-fail-fast 2>&1 || true"
compare_command: "cargo test --no-fail-fast 2>&1 || true" # optional; defaults to baseline_command
failure_patterns:
- "FAILED"
- "error\\[E"
policy: fail # fail (default) | warn | skipThe comparison reports new failures and fixed failures. policy decides what a new failure does to the stage: fail blocks completion, warn reports it and continues, skip turns the check off. Failures already present in the baseline do not count against the stage.
loom stage dispute-criteria is how an agent challenges a criterion instead of quietly weakening it. The daemon moves the stage to NeedsAdjudication and spawns a real session — with the full tool surface, running the disputed criterion itself — to judge it; loom stage adjudicate records that session's verdict, and a stage's own worktree session is refused so it can never judge its own dispute.
A verdict does one of three things: Accept patches acceptance or wiring and re-queues the stage; NeedsMoreEvidence appends the judge's questions to the next signal and re-queues, up to a capped number of rounds; Reject moves the stage to NeedsHumanReview. Every accepted amendment writes a numbered snapshot under .loom/work/plan_versions/ plus an audit row.
Version 2 plans add three more dispute kinds, each filed once per review round with several ids: dispute-findings (a judge rules each finding uphold, dismiss or defer to a dependent stage that is not yet finished; integration-verify never defers), dispute-contract (Accept re-freezes the contract and may amend contracts; Reject sends the agent to loom stage contracts restore) and dispute-integrity (Accept records the event as accepted at its current count or hash). Each kind has its own budget of three disputes per stage.
The plan-level adjudication.max_amendments_per_stage field bounds autonomous amendments per stage (default 10), so a multi-fix plan is not sent to a human part-way through. loom stage amend gives an operator the same audited path directly: it replaces, inserts, or deletes one element of a stage's acceptance, wiring, or wiring_tests array, snapshotting and logging the change the same way.
Loom's answer to "every session starts from zero" is a three-stage pipeline: capture during execution, distill at the end of a plan, retrieve cheaply forever after.
While a stage runs, its agent journals to .loom/work/memory/<session>.md:
loom memory note "gotcha: worktree exclude lives at <worktree>/.git/info/exclude, not <dir>/.git/..."
loom memory decision "centralized plan lookup in plan/parser" --context "avoids an orchestrator→commands layering violation"Entries are typed (note, decision, change, question). The most recent are embedded in the recitation section at the end of the next signal — the position with the highest model attention — so a later stage inherits the detail an earlier stage paid for instead of rediscovering it.
Memory is deliberately cheap and disposable. It is a working journal, not the deliverable.
A knowledge-distill stage runs at the end of a plan and performs the reduce step: it reads every stage memory, dedupes, and curates the survivors into permanent knowledge. Mistakes are rewritten as actionable prevention rules rather than anecdotes:
## [Short description]
**What happened:** ...
**Why:** [root cause]
**Prevention:** [how to detect it earlier]
**Fix:** [what to do instead]Procedural noise ("spawned agents", "ran tests") and anything recoverable from git history is dropped. The stage starts from loom memory pending --group, which prints corrections first, sorted by target file and heading so each one is applied with replace-section, then mistakes, decisions and the rest; a mistake that repeats one already in the tree becomes a proposal for a hook or a loom plan verify check rather than another paragraph. Every entry ends with a loom memory resolve receipt. loom review turns the same memories into a human-readable code-review document.
Knowledge lives in doc/loom/knowledge/ and is tiered: a generated INDEX.md (tier 0) maps the seven curated summary files (tier 1), which link out to per-category topic files (tier 2, e.g. architecture/merge-flow.md). Tier-1 files stay navigable summaries; detail lives in topics. The index is regenerated automatically on every knowledge write.
Reading protocol — agents read INDEX.md first for orientation, then the tier-1 summary for the area they are working in, then only the tier-2 topics they actually touch; a specific question is pulled with loom knowledge context --query, which returns the matching sections quoted. Inside a stage, the per-stage Knowledge Brief comes first. Loading the whole base defeats the point of tiering. Every session that starts inside a repository holding doc/loom/knowledge/INDEX.md also receives a one-paragraph pointer to it from the knowledge-orient.sh SessionStart hook, installed globally by loom init/loom repair --fix.
Writing protocol — when a tier-1 section grows past roughly 40 lines, move its body into a topic with loom knowledge update <category>/<slug> and leave a 2-4 line summary plus a relative link behind. Write the link as [Title](category/slug.md) in a tier-1 file: that is the tree's convention for references.
There is no aggregate line budget across the knowledge base. What matters is per-file size, and loom knowledge check --strict enforces it: a tier-1 summary is at most 250 lines with sections of at most 40, a tier-2 topic at most 400 lines with sections of at most 80, and the generated INDEX.md at most 16 KB. A tree that cannot clear those yet records its current issues once with --write-baseline and gates with --baseline, so only new issues fail. loom memory note refuses a mistake: note without a Prevention: and a stale-knowledge: note that does not name <file>#<heading> and a Correction:, so a correction can be applied mechanically.
Retrieval is deterministic and offline. loom knowledge context --query <text> returns a token-budgeted context pack: the tool chunks the curated prose, scores each chunk, fuses the per-channel rankings, and takes whole chunks in order until the budget is spent, always reporting what it left out. There is no embedding model, no network call and no randomness — a pack is a pure function of the bytes on disk and the query string, so the same query returns the same pack.
loom knowledge context --query "how does merge cleanup order work" --budget-tokens 3000
loom knowledge context --query "source graph coverage" --explain # per-item scores and why each was selected
loom knowledge context --query "sandbox rules" --json # machine-readableStage sessions do not have to ask. Signal generation embeds a per-stage Knowledge Brief built through the same single entry point, so what a stage receives at spawn and what you get from the CLI are produced identically. Loom records what each recipient was given, so a second retrieval in the same session skips what the first already quoted rather than repeating it.
--scope selects which channels to search: knowledge searches the curated prose, source searches the derived source graph of symbols extracted from the code, and all (the default) fuses both. The two are ranked separately — prose over chunk text, symbols over their scope and signature — and then fused, so one pack can mix curated prose with the exact symbols a query names. A symbol whose file the parser could not fully read is still returned, but without a high-confidence claim.
The source graph maintains itself. loom init and loom run publish a base layer for the current revision before anything else starts, and each stage's working-tree overlay is refreshed immediately before its signal is written — so a stage's brief describes the code as that stage will actually find it. Publication is advisory: when it cannot run, loom prints one line and carries on, because a missing graph must degrade retrieval rather than block a run. A base layer is keyed to a revision and so is never published from a dirty tree; there, loom knowledge sync builds a working-tree overlay instead and tells you which layer it wrote.
Codex's native PostToolUse:apply_patch hook records every patched path into the stage overlay, and its UserPromptSubmit hook uses the same retrieval entry point as Claude. Shell-based edits remain outside file-tool hook visibility and are reconciled by the next explicit graph refresh.
Run loom knowledge sync after editing knowledge outside the CLI; it reports whether both derived layers are current, naming the source-graph layer it produced — base, local-overlay or skipped, with the reason.
Agents populate knowledge during orchestration: a knowledge-bootstrap stage writes CONTENT, while loom init scaffolds the directory automatically. loom knowledge sync rebuilds the derived retrieval artifacts.
Knowledge directories created before the hierarchy existed stay flat and keep working unchanged — neither reading nor updating migrates them behind your back. loom knowledge sync performs the opt-in upgrade, creating INDEX.md the first time it runs on a flat directory.
Knowledge writes are protected by the sandbox defaults: agents update knowledge through loom knowledge ..., never by editing the files directly.
loom knowledge bootstrap [--structural-only] [--refresh] [--dry-run] [--model M] [--effort E] builds doc/loom/knowledge/ from scratch for a repository that has no loom plans — retrieval needs the directory to exist before sync/context can do anything. It runs a deterministic host phase (scaffold the tree, rebuild the catalog and source graph, partition the repo into directory clusters with content digests) followed by an interactive foreground Claude session that writes knowledge only through loom knowledge ..., then a host finalization that refreshes the index and writes a receipt.
--structural-onlystops after the deterministic phase — no model session is launched.--refreshnarrows the work to clusters whose digest changed since the last receipt, any removed cluster, and any tier-1 file still at its template content; with nothing changed it printsknowledge is currentand spawns nothing.--dry-runprints the work plan and the exactclaudecommand it would run, without spawning anything.--model/--effortpick the model and effort for the semantic session, same values as elsewhere in loom.
The command refuses to run inside a stage session — stages write knowledge through loom knowledge update instead. The committed receipt doc/loom/knowledge/.bootstrap-receipt.json records one digest per cluster; commit it alongside the generated knowledge files so --refresh reports current across clones, not just locally.
Every stage's main agent is an orchestrator; the model and effort it runs come from its stage type's default, which the operator can override:
| Stage type | Default model | Default effort |
|---|---|---|
standard |
opus |
high |
knowledge |
opus |
medium |
knowledge-distill |
sonnet |
high |
integration-verify |
opus |
xhigh |
Configure either default per stage type in the [models] section of ~/.loom/config.toml (user tier) or <repo>/.loom/work/config.toml (project tier, resolved per key — a project section that sets only one of these eight keys still lets the rest fall through to the user config):
# ~/.loom/config.toml or .loom/work/config.toml
[models]
standard_model = "opus"
standard_effort = "high"
knowledge_model = "opus"
knowledge_effort = "medium"
knowledge_distill_model = "sonnet"
knowledge_distill_effort = "high"
integration_verify_model = "opus"
integration_verify_effort = "xhigh"A plan stage's model / reasoning_effort field overrides both config tiers for that one stage. Merge and base-conflict sessions stay pinned at opus/high and adjudication keeps its own [adjudication] model; neither is configurable through [models].
The orchestrator decomposes the work, hands each subagent full context, then verifies and commits. It makes a change itself only when that is cheaper than a spawn: at most 20 changed lines in at most 2 files it has already read, proven by one command; a fable session delegates even those. Implementation is delegated to as few subagents as the work allows, each spawned by agent type so the model choice is explicit:
| Agent | Model | Use for |
|---|---|---|
loom-software-engineer |
Sonnet | Common implementation and integration tests to detailed instructions |
loom-codex-forwarder |
Codex GPT-5.6 Terra or GPT-6 Luna | Codex lane, licensed only on stages listing codex in implementers: Terra for common implementation/integration tests, Luna for boilerplate, scaffolding, and simple unit tests |
loom-senior-software-engineer |
Opus | Mainstream architecture and algorithm implementation, complex debugging, security-sensitive or cross-cutting work |
loom-code-reviewer |
Opus | Read-only code, security, and architecture review |
loom-advisor |
Fable | Diagnosis after a repeated failure — advice returned, nothing written |
Fable-tier implementation — major bugs, visual/UI design, extremely challenging algorithmic design — has no dedicated agent type; it is spawned with an explicit model override rather than relying on inheritance.
The loom-codex-forwarder row additionally depends on the codex CLI and its plugin's companion runtime being installed. loom run checks this at startup and prints an advisory warning if either is missing — it never blocks the run — and terra-/luna-tier work falls back to Sonnet for the duration; the stage signal states the fallback explicitly and does not spawn loom-codex-forwarder.
This is why savings come from delegation rather than downgrade: an untyped subagent silently inherits the stage's own (usually Opus) model, making every worker expensive. Two failures on the same task should produce a loom-advisor diagnosis, not a blind retry at a larger model.
A stage omits model and reasoning_effort by default, so the stage type's configured default applies. Set either field only as a deliberate per-stage override (low, medium, high, xhigh, max for effort). ultracode: true licenses a stage for large multi-agent fan-out; it is per-stage opt-in so the cost decision stays explicit.
Loom supports plan-level defaults plus stage-level overrides.
loom:
version: 1
sandbox:
enabled: true
auto_allow: true
filesystem:
deny_read:
- "~/.ssh/**"
- "~/.aws/**"
- "../../**"
- "../.worktrees/**"
deny_write:
- "../../**"
- "doc/loom/knowledge/**"
allow_write:
- "src/**"
network:
allowed_domains: ["github.com", "crates.io"]
additional_domains: []
allow_local_binding: false
allow_unix_sockets: []Note: knowledge file writes are intentionally protected by sandbox defaults; knowledge updates should be done via loom knowledge ... commands. Plan-configured excluded_commands are rejected because broad executable exemptions bypass the host sandbox. When sandboxing is enabled, generated settings use host denyRead rules for sensitive paths and failIfUnavailable: true; failure to write those settings blocks session spawn. Unit tests pin the generated policy and blocked-spawn behavior. A credentialed Claude host-runtime canary across Bash, interpreters, build scripts, symlinks, and file tools remains a manual release check.
The sandbox: block above bounds the agent session. A separate control bounds the commands loom itself runs from your plan — every acceptance criterion, setup command, truth check, wiring test, dead-code check and change-impact command:
loom:
version: 1
sandbox:
command_confinement: confined # plan-level default; `inherit` to opt out
stages:
- id: build
sandbox:
command_confinement: inherit # per-stage override| Level | Behavior |
|---|---|
confined |
Default. The child process environment is cleared and rebuilt from a fixed allowlist |
inherit |
The child inherits loom's ambient environment |
Plans are trusted artifacts, but trusted is not privileged: under confined, a plan line cannot read GITHUB_TOKEN, AWS_* or ANTHROPIC_API_KEY merely because you started loom from a shell that had them. The allowlist carries what a build toolchain needs to find itself — HOME, PATH, CARGO_HOME, RUSTUP_HOME, locale and terminal variables, TMPDIR, the proxy variables and the CA-bundle locations (SSL_CERT_FILE, SSL_CERT_DIR, NIX_SSL_CERT_FILE). SSH_AUTH_SOCK is deliberately withheld, so an acceptance criterion that needs SSH auth fails by design rather than silently borrowing your agent.
What confinement is not. It is environment scrubbing — least-privilege hygiene, not a security boundary. Loom applies no namespace, seccomp, landlock, cgroup or network isolation to the commands it spawns: a confined command shares your network namespace, can read and write any path your user can, and can reach any Unix socket on the host. The
network:settings above are emitted into the agent session's sandbox and do not restrict plan-authored commands. Useconfinedto keep ambient credentials out of plan commands; do not use it to run code you would not run yourself.
Every spawned session — stage, merge, base-conflict, adjudication — launches from a settings capsule that denies writes to .loom/, .claude/, .worktrees/, the hooks directories, and git hooks and config, regardless of the filesystem rules above.
A handful of commands still need to reach that state from inside a sandboxed session: loom memory ..., loom stage block, loom stage dispute-criteria, loom handoff, loom stage merge --resolved, and loom worktree remove. Each writes a one-shot ticket instead, picked up by a PostToolUse relay hook and applied by the daemon at most once through a per-session ledger. loom request status <id> reports whether the daemon has applied a given ticket yet.
The daemon also refuses to merge, or hand to a conflict-resolution session, any branch whose diff touches .claude/, .mcp.json, .loom/, or the in-repo hooks directory — the stage moves to NeedsHumanReview naming the offending paths instead.
All stages default to auto (agents auto-accept any action their heuristics deem safe, since loom stages run autonomously with no human to answer prompts; the sandbox deny/allow rules are the safety boundary). Override per-plan or per-stage to tighten control:
loom:
version: 1
sandbox:
permission_mode: accept-edits # plan-level override
stages:
- id: my-stage
sandbox:
permission_mode: plan # stage-level override (takes precedence)Valid values: auto (default), accept-edits, plan, default. bypass-permissions is rejected at init time.
Claude Code's --remote-control flag lets the loom orchestrator drive spawned Claude sessions programmatically. Loom enables it automatically when prerequisites are met — no configuration required.
Prerequisites (preflight check):
- Claude version ≥ 2.1.51
- Auth: claude.ai login — loom accepts either credential store:
~/.claude/.credentials.json, or (on macOS) aClaude Code-credentialsentry in the Keychain, which is where Claude Code stores credentials on macOS instead of the file. Additionally, none of these env vars may be set:ANTHROPIC_API_KEY,CLAUDE_CODE_OAUTH_TOKEN,CLAUDE_CODE_USE_BEDROCK,CLAUDE_CODE_USE_VERTEX,CLAUDE_CODE_USE_FOUNDRY
The flag exits non-zero when its prerequisites are not met, so loom never passes it blindly. When preflight fails, loom falls back silently to standard mode and prints a one-line advisory at orchestrator startup (e.g. ⚠ Remote Control disabled: <reason>).
Configuration — the [remote_control] section of .loom/work/config.toml carries a single switch:
# .loom/work/config.toml
[remote_control]
mode = "auto" # default: enable whenever preflight passes
# mode = "off" # never enable, regardless of preflightToggling mode takes effect on the next session spawn — no daemon restart needed.
Fast crashes — a session that crashes within 15 seconds of spawn while Remote Control is active disables Remote Control for the rest of that daemon run, logs it once, and retries the stage without the flag. Nothing is written to disk: a daemon restart tries Remote Control again from scratch. Set mode = "off" above to stop it from trying at all.
Session naming — every spawned session is named after its stage in the Remote Control UI: the stage name for stage sessions, and Merge: <stage name>, Base conflict: <stage name>, Knowledge: <stage name> for merge, base-conflict, and knowledge sessions respectively. Claude binaries whose --remote-control flag doesn't accept a name argument automatically fall back to the bare flag — detected via a one-time claude --help capability check, no configuration needed.
Loom spawns each stage's Claude Code session through a terminal backend. Two are available.
| Backend | Default | Sessions run in | Needs a GUI? |
|---|---|---|---|
native |
✅ yes | a host terminal emulator window | yes |
tmux |
opt-in | a detached tmux server (no window) | no |
The native backend opens a real terminal window per session — you watch stages run in your own
terminal emulator. It requires a detectable emulator, so it cannot run headless.
The tmux backend spawns each session into a detached tmux server instead, which makes loom usable
over SSH, on a headless Linux box, or anywhere no terminal emulator exists. tmux must be installed
and on PATH.
# .loom/work/config.toml
[terminal]
backend = "tmux" # or "native" (default)Or from the CLI:
loom init <plan> --backend tmux # skips the interactive backend prompt
loom run --backend tmux # persists the choice to [terminal]loom init prompts for a backend when run interactively; with no TTY it defaults to native.
Changing the backend while the daemon is running is refused with a hint — loom stop first, then
re-run with --backend. Selecting a backend takes effect on the next spawn.
If tmux is selected but not installed, loom init prints an advisory warning; loom run (and
loom run --foreground) refuses to start — the configured backend is never silently swapped for
the other lane.
WSL2 runs the published loom-linux-x86_64 binary unmodified — it is an ordinary glibc ELF, and git
worktrees, the .loom/work/ Unix socket and PID liveness checks all behave as they do on native Linux.
What does not carry over is the native backend. It opens a real terminal-emulator window per
session, and a stock WSL install ships no Linux GUI stack — without WSLg or an X server there is no
emulator to detect. Select tmux explicitly:
sudo apt install tmux jq # both required; ripgrep and fd are recommended
loom init doc/plans/PLAN-<name>.md --backend tmux
loom run --backend tmuxSessions then run in detached tmux servers, and you watch them from a Windows Terminal tab:
loom status --live # live ledger dashboard
loom status --web # browser dashboard: starts at 7373, then the next free port
loom attach # tiled overview of every live session
loom attach <stage-id> # attach to one stageOn Windows 11 with WSLg, an emulator installed inside WSL is detectable and the native backend does
work, but tmux remains the better fit: a run survives closing the window, and loom attach
reconnects to one already in progress.
Two WSL-specific notes:
- Keep the repository on the Linux filesystem (
~/src/…), not under/mnt/c. Loom creates a worktree per parallel stage and the daemon polls stage files every 5s; 9p latency across the Windows drive makes both crawl. - ARM64 Windows has no published Linux ARM64 binary. Install the Rust toolchain inside WSL and build from source.
Loom does not put every stage in one shared tmux server. Each session gets its own server on its
own socket, named loom-<session-id> under $TMUX_TMPDIR (else /tmp).
This is deliberate: a wedged or killed server takes down exactly one stage, instead of every stage running in parallel. Liveness is tracked from PID files rather than by asking tmux, because a tmux server whose agent process has died still reports the session as existing — which would hide the crash from loom's monitor and prevent the retry.
loom attach # tiled overview of every live session
loom attach <stage-id> # attach directly to one stage's sessionWith no argument, loom attach builds a per-repo viewer window with one pane per live session, tiled.
With a stage id it attaches straight to that stage. Both require a real terminal (a TTY), and both
work only for sessions spawned by the tmux backend — with the native backend it tells you so.
Panes in the overview are live, writable terminals, not read-only views: keystrokes go to the
agent, and C-b x closes that stage's pane. Detach the normal way with C-b d.
The configured backend is authoritative: nothing on disk can swap it for the other lane. If tmux is
configured but not on PATH, loom run refuses to start (see Selecting a backend);
if it becomes unavailable mid-run, a spawn fails with an error naming the fix instead of silently
falling back to native.
Loom enables agent teams in spawned sessions (CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1) and injects team-usage guidance into stage signals.
Use teams when work needs coordination/discussion across agents (multi-dimension review, exploratory analysis). Use subagents for independent, concrete file-level tasks.
project/
├── .loom/
│ ├── work/ # orchestration state (mode 0700)
│ │ ├── config.toml
│ │ ├── stages/
│ │ ├── sessions/
│ │ ├── signals/
│ │ ├── handoffs/
│ │ ├── memory/ # per-session journals
│ │ ├── archive/
│ │ ├── crashes/
│ │ ├── logs/ # per-session captured stderr
│ │ ├── pids/
│ │ ├── wrappers/
│ │ ├── orchestrator.sock # daemon IPC socket
│ │ └── orchestrator.pid
│ ├── memory/archive/ # state of completed plans
│ └── cache/ # derived caches
├── .worktrees/
│ └── <stage-id>/ # one per stage; its .loom/work links to the main one
├── doc/plans/
└── doc/loom/knowledge/
├── INDEX.md # generated tier-0 map
├── architecture.md # tier-1 summaries
├── patterns.md
├── ...
└── architecture/ # tier-2 topics, one directory per category
└── merge-flow.md
Loom provides context-aware tab completions for all commands, subcommands, flags, and dynamic values (stage IDs, plan files, session IDs, knowledge files).
loom completions --installAuto-detects your shell from $SHELL and writes completions to the standard location:
| Shell | Install Path |
|---|---|
| Bash | ~/.local/share/bash-completion/completions/loom |
| Zsh | ~/.zfunc/_loom |
| Fish | ~/.config/fish/completions/loom.fish |
Follow the printed post-install instructions to activate (e.g., for zsh, ensure fpath=(~/.zfunc $fpath) appears before compinit in ~/.zshrc).
You can also write the completion script to a file yourself:
# bash
loom completions bash > ~/.local/share/bash-completion/completions/loom
# zsh — ensure ~/.zfunc is in fpath (add before compinit in ~/.zshrc):
# fpath=(~/.zfunc $fpath)
# autoload -Uz compinit && compinit
mkdir -p ~/.zfunc
loom completions zsh > ~/.zfunc/_loom
# fish
loom completions fish > ~/.config/fish/completions/loom.fishOlder versions of loom used clap_complete and required an eval line in your shell RC file that ran a subprocess on every shell startup. The new system writes a static script to disk and only calls loom at actual tab-completion time, which means faster shell startup and completions that work even before loom is in your PATH.
To check whether you need to migrate:
loom completions --migrateThis scans for two things:
evallines in RC files (.bashrc,.zshrc, etc.) likeeval "$(loom completions zsh)"— these should be removed- Stale completion files containing old
clap_completemarkers — these need to be regenerated
If issues are found, follow the printed instructions. Typically: remove the eval line from your RC file, then run loom completions --install to write the new file-based completion script.
- Commands and subcommands (
loom stage <TAB>shows all stage subcommands) - Flags (
loom run --<TAB>shows available flags) - Stage IDs with smart filtering (
loom stage complete <TAB>shows only executing stages) - Plan files, session IDs, knowledge files (including aliases like
deps,tech) - Model names, trigger types, and more
MIT


