Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 3 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -236,8 +236,9 @@ flow; `arbor doctor` diagnoses the install. Both `quickstart` and `setup` write
`~/.arbor/config.yaml`, so day-to-day you can just run `arbor`
with no flags. The first thing Arbor does is an **intake conversation** that turns your
goal, target directory, metric, baseline, budget, dev/test discipline, and artifact
paths into a one-screen **Arbor Research Contract**. Once you confirm it, the live
dashboard takes over.
paths into a one-screen **Arbor Research Contract**. Read/discuss requests stay in a
scoped, read-only mode; experiment plans are staged and launch only after a later explicit
confirmation. Once you confirm the resolved contract, the live dashboard takes over.

```bash
# Point at a benchmark directory and a config
Expand Down
2 changes: 1 addition & 1 deletion README.zh-CN.md
Original file line number Diff line number Diff line change
Expand Up @@ -194,7 +194,7 @@ arbor # 在当前目录启动交互式会话

已有模型提供商?`arbor setup` 提供完整的 提供商 / 模型 / base_url / API 密钥 配置流程;
`arbor doctor` 用于诊断安装。`quickstart` 与 `setup` 都会将配置写入
`~/.arbor/config.yaml`,此后日常使用直接运行 `arbor` 即可,无需额外参数。Arbor 启动后首先进行一次**任务摄入对话**,将你的目标、目标目录、评估指标、基线、预算、开发/测试纪律和产物路径整理成一份简洁的 **Arbor 研究合同**。确认后,实时仪表盘随即接管。
`~/.arbor/config.yaml`,此后日常使用直接运行 `arbor` 即可,无需额外参数。Arbor 启动后首先进行一次**任务摄入对话**,将你的目标、目标目录、评估指标、基线、预算、开发/测试纪律和产物路径整理成一份简洁的 **Arbor 研究合同**。阅读/讨论请求保持在受限的只读模式;实验计划会先暂存,只有你在后续消息中明确确认后才会启动。确认完整合同后,实时仪表盘随即接管。

```bash
# 指定基准目录和配置文件
Expand Down
13 changes: 8 additions & 5 deletions docs/cli.md
Original file line number Diff line number Diff line change
Expand Up @@ -30,10 +30,13 @@ changing eval or data"`). Omit it to start with the intake chat.

### Default flow

1. Open an interactive chat with the intake agent.
2. The agent confirms which project directory to work on (the `--cwd` flag is only a hint).
3. When you agree on a plan, the agent launches the experiment.
4. You confirm the research contract shown in the terminal.
1. Open an interactive chat with the intake agent. Read/discuss requests stay in scoped,
read-only discussion mode; launch planning is enabled only for an optimization/run goal.
2. The agent confirms which project directory to work on. `--cwd` is the default boundary;
external paths must be named explicitly and are re-authorized after chat resume.
3. The agent stages an exact plan and shows it to you. A later, explicit confirmation launches
that same staged plan; visible questions never execute hidden tool calls.
4. You confirm the fully resolved research contract shown in the terminal.
5. A quick preflight runs against the chosen project.
6. The coordinator runs to completion and writes `REPORT.md`.

Expand All @@ -45,7 +48,7 @@ changing eval or data"`). Omit it to start with the intake chat.
| `--config, -c PATH` | Project YAML config. Defaults to `research_config.yaml` / `arbor.yaml` / `autoresearch.yaml` in the target project. |
| `--max-cycles N` | Max completed/skipped/failed idea experiments before finalizing. |
| `--max-turns N` | Hard cap on coordinator ReAct turns — a cost/runaway safety valve. |
| `--intake-max-turns N` | Max planning-chat turns before launch (default `30`). |
| `--intake-max-turns N` | Max internal agent turns for each intake message (default `30`). |
| `--run-name NAME` | Session name under `.arbor/sessions/`. Defaults to a timestamp. |
| `--resume` | Resume an interrupted run from its checkpoint in the existing workspace/session. |
| `--workspace-dir PATH` | Session/artifact directory override. Default `<target>/.arbor/sessions/<run_name>`. |
Expand Down
10 changes: 5 additions & 5 deletions docs/cli.zh.md
Original file line number Diff line number Diff line change
Expand Up @@ -29,10 +29,10 @@ data"`)。省略它则从接入对话开始。

### 默认流程

1. 与接入智能体打开一段交互式对话。
2. 智能体确认要在哪个项目目录上工作(`--cwd` 参数只是一个提示)。
3. 当你与智能体就计划达成一致后,智能体启动实验。
4. 你确认终端里展示的研究契约。
1. 与接入智能体打开一段交互式对话。阅读/讨论请求保持在有范围限制的只读讨论模式;只有优化或启动目标才进入运行规划模式。
2. 智能体确认目标项目目录。`--cwd` 是默认权限边界;外部路径必须由你明确点名,恢复旧对话后也要重新授权。
3. 智能体暂存并展示一份精确计划;你在后续消息中明确确认后,CLI 才启动同一份计划。显示问题时不会在后台继续执行工具。
4. 你确认终端里展示的完整研究契约。
5. 针对所选项目跑一次快速预检。
6. Coordinator 运行至完成并写出 `REPORT.md`。

Expand All @@ -44,7 +44,7 @@ data"`)。省略它则从接入对话开始。
| `--config, -c PATH` | 项目 YAML 配置。默认取目标项目里的 `research_config.yaml` / `arbor.yaml` / `autoresearch.yaml`。 |
| `--max-cycles N` | 定稿前最多完成/跳过/失败多少个想法实验。 |
| `--max-turns N` | Coordinator ReAct 轮次的硬上限——一个成本/失控安全阀。 |
| `--intake-max-turns N` | 启动前规划对话的最多轮次(默认 `30`)。 |
| `--intake-max-turns N` | 每条 intake 消息允许的最大内部 Agent 轮数(默认 `30`)。 |
| `--run-name NAME` | `.arbor/sessions/` 下的会话名。默认是时间戳。 |
| `--resume` | 在现有工作空间/会话里从检查点续跑一次被中断的运行。 |
| `--workspace-dir PATH` | 会话/产物目录覆盖。默认 `<target>/.arbor/sessions/<run_name>`。 |
Expand Down
14 changes: 8 additions & 6 deletions docs/preparing-a-benchmark.md
Original file line number Diff line number Diff line change
Expand Up @@ -33,18 +33,20 @@ score: 0.8123
A simple working baseline already in the repo is ideal — it gives Arbor a number to beat
and confirms the eval actually runs.

!!! tip "Only have code? The intake agent can build the eval for you"
!!! tip "Only have code? Intake can plan the eval setup with you"
You don't need an eval script — or even a dev/test split — *before* you start. If your
repo is just code, launch `arbor` anyway: in the intake chat the agent asks what
"better" means, then offers to **scaffold a minimal eval** (and, if you want, carve a
**dev/held-out split**) for you to confirm. No held-out set and don't want one? You can
"better" means, then stages a plan whose first coordinator task can **scaffold a minimal
eval** (and, if you want, carve a **dev/held-out split**) after you confirm. No held-out
set and don't want one? You can
iterate on a single split — the agent will just note that the final score has no
held-out guard. See [Describe the task](#2-describe-the-task-readme-or-just-tell-the-cli).

!!! tip "Let the intake agent do the plumbing"
!!! tip "Let Arbor do the plumbing after approval"
You don't have to pre-initialize git, `chmod +x` your script, or run the eval yourself.
When you launch `arbor`, the intake agent will quietly do those setup steps for you
(and confirm the eval produces a score) before the study starts.
Intake itself is read-only: it inspects files inside the approved scope and stages the
contract. After you confirm, deterministic preflight and the coordinator handle setup
and evaluation work; intake never runs shell commands behind a displayed question.

## 2. Describe the task — README *or* just tell the CLI

Expand Down
13 changes: 7 additions & 6 deletions docs/preparing-a-benchmark.zh.md
Original file line number Diff line number Diff line change
Expand Up @@ -31,16 +31,17 @@ score: 0.8123

仓库里已经有一个能跑通的简单基线是最理想的——它为 Arbor 提供了一个要超越的基准数字,也确认评测确实能跑。

!!! tip "只有代码?接入智能体能替你搭好评测"
!!! tip "只有代码?接入智能体能和你规划评测准备"
在你开始*之前*,你并不需要评测脚本——甚至不需要 dev/test 划分。如果你的仓库只有代码,照样
启动 `arbor`:在接入对话里,智能体会询问“更好”意味着什么,然后提出**搭建一个最小评测**
(如果你愿意,还会**切分出 dev/留出划分**)供你确认。没有留出集,也不打算准备?你可以在单一划分上
启动 `arbor`:在接入对话里,智能体会询问“更好”意味着什么,并暂存一份计划;经你确认后,
coordinator 的首个任务可以**搭建一个最小评测**(如果你愿意,还会**切分出 dev/留出划分**)。
没有留出集,也不打算准备?你可以在单一划分上
迭代——智能体只会提示最终分数缺少留出护栏。见
[描述任务](#2-describe-the-task-readme-or-just-tell-the-cli)。

!!! tip "让接入智能体处理准备工作"
你不必预先初始化 git、给脚本 `chmod +x`,或自己跑评测。当你启动 `arbor` 时,接入智能体会在
研究开始前自动替你完成这些准备步骤(并确认评测确实产出一个分数)。
!!! tip "确认后让 Arbor 处理准备工作"
接入阶段本身是只读的:只会在你批准的范围内检查文件并暂存研究契约。你确认之后,确定性的预检和
coordinator 才会处理环境与评测工作;接入智能体不会在显示问题后偷偷执行 shell 命令。

## 2. 描述任务——用 README *或者* 直接告诉 CLI { #2-describe-the-task-readme-or-just-tell-the-cli }

Expand Down
22 changes: 17 additions & 5 deletions src/cli/commands/run.py
Original file line number Diff line number Diff line change
Expand Up @@ -44,7 +44,7 @@ def run_command(
max_turns: int | None = typer.Option(None, "--max-turns",
help="Hard cap on coordinator ReAct turns. Use as a cost/runaway safety valve."),
intake_max_turns: int = typer.Option(30, "--intake-max-turns",
help="Max planning-chat turns before launch. Rarely needed."),
help="Max internal agent turns per intake message. Rarely needed."),
run_name: str | None = typer.Option(None, "--run-name", help="Session name under .arbor/sessions/. Defaults to timestamp."),
resume: bool = typer.Option(
False, "--resume",
Expand Down Expand Up @@ -263,6 +263,11 @@ def _probe_config() -> CoordinatorConfig:
plan_cwd = Path(outcome.cwd).resolve()
unloaded_skills = list(outcome.unloaded_skills)
refined_instruction = outcome.instruction
if outcome.notes:
refined_instruction += (
"\n\nAdditional user constraints:\n- "
+ "\n- ".join(outcome.notes)
)
selected_plugin = outcome.plugin
selected_plugin_profile = outcome.plugin_profile
selected_plugin_mode = outcome.plugin_mode
Expand Down Expand Up @@ -302,10 +307,9 @@ def _probe_config() -> CoordinatorConfig:

# ── 2. Quick essentials check against plan_cwd ─────────────────
#
# The intake agent already runs its own sanity pass (git status,
# eval.sh, etc.) inline during the chat, so we don't show a
# ceremonial preflight section here — just run silently and only
# surface fatal issues as a red panel.
# Intake is deliberately read-only. Run the deterministic essentials
# preflight after the staged plan is approved, and surface only failures by
# default so planning stays conversational without hiding execution.
#
# Resolve the plugin's eval_contract so the contamination preflight can
# warn (zero-network) when the benchmark is likely in pretraining data.
Expand Down Expand Up @@ -790,6 +794,14 @@ def _resolve_effective_options(
max_cycles: int | None,
max_turns: int | None,
) -> dict[str, Any]:
# ``load_layered_defaults`` intentionally preserves structured blocks.
# Normalize the documented ``llm: {provider, model, ...}`` form onto the
# flat lookup surface used below, while keeping legacy top-level values as
# higher-precedence compatibility aliases.
nested_llm = project_defaults.get("llm")
if isinstance(nested_llm, dict):
project_defaults = {**nested_llm, **project_defaults}
project_defaults.pop("llm", None)
if project_defaults.get("provider") is not None:
provider_source = "project"
eff_provider = project_defaults.get("provider")
Expand Down
90 changes: 83 additions & 7 deletions src/cli/intake/conversation_store.py
Original file line number Diff line number Diff line change
Expand Up @@ -28,6 +28,7 @@

import json
import os
import re
import tempfile
from dataclasses import dataclass
from datetime import datetime, timezone
Expand All @@ -45,6 +46,7 @@
META_NAME = "meta.json"

_TITLE_MAX = 80
_CONV_ID_RE = re.compile(r"conv_\d{8}_\d{6}(?:_\d+)?")


def _utc_now_iso() -> str:
Expand Down Expand Up @@ -135,8 +137,20 @@ def save_conversation(
Updates ``rec`` in place (``updated_at``, ``turns``, ``title``, ``launched``)
so the caller's handle stays current across repeated saves.
"""
rec.dir.mkdir(parents=True, exist_ok=True)
write_messages(rec.messages_path, messages)
if not _CONV_ID_RE.fullmatch(rec.conv_id):
raise ValueError(f"invalid conversation id: {rec.conv_id!r}")
root = conversations_root(rec.cwd)
root.mkdir(parents=True, exist_ok=True)
if root.is_symlink():
raise OSError(f"refusing symlinked conversation root: {root}")
try:
root.resolve(strict=True).relative_to(rec.cwd.resolve(strict=True))
except (OSError, ValueError) as exc:
raise OSError(f"conversation root escapes project: {root}") from exc
rec.dir.mkdir(parents=False, exist_ok=True)
if rec.dir.is_symlink():
raise OSError(f"refusing symlinked conversation directory: {rec.dir}")
write_messages(rec.messages_path, _messages_for_disk(messages))

rec.updated_at = _utc_now_iso()
rec.turns = _count_user_turns(messages)
Expand All @@ -158,15 +172,33 @@ def find_conversations(cwd: str | os.PathLike[str]) -> list[ConversationRecord]:
Defensive: a dir without a parseable ``meta.json`` is skipped, never fatal.
"""
root = conversations_root(cwd)
if not root.is_dir():
if not root.is_dir() or root.is_symlink():
return []
try:
resolved_root = root.resolve(strict=True)
resolved_root.relative_to(Path(cwd).resolve(strict=True))
except (OSError, ValueError):
return []

records: list[ConversationRecord] = []
for conv_dir in root.iterdir():
if not conv_dir.is_dir():
if (
not conv_dir.is_dir()
or conv_dir.is_symlink()
or not _CONV_ID_RE.fullmatch(conv_dir.name)
):
continue
try:
conv_dir.resolve(strict=True).relative_to(resolved_root)
except (OSError, ValueError):
continue
data = _load_json(conv_dir / META_NAME)
if not isinstance(data, dict) or "conv_id" not in data:
if (
not isinstance(data, dict)
or data.get("conv_id") != conv_dir.name
or (conv_dir / META_NAME).is_symlink()
or (conv_dir / MESSAGES_NAME).is_symlink()
):
continue
try:
records.append(ConversationRecord.from_meta(Path(cwd), data))
Expand All @@ -189,13 +221,17 @@ def latest_unfinished(cwd: str | os.PathLike[str]) -> ConversationRecord | None:


def _count_user_turns(messages: list[dict[str, Any]]) -> int:
return sum(1 for m in messages if m.get("role") == "user")
return sum(
1
for m in messages
if m.get("role") == "user" and not m.get("_internal")
)


def _derive_title(messages: list[dict[str, Any]]) -> str:
"""First user message, flattened to a short single line."""
for m in messages:
if m.get("role") != "user":
if m.get("role") != "user" or m.get("_internal"):
continue
text = _message_text(m.get("content"))
text = " ".join(text.split())
Expand All @@ -217,6 +253,46 @@ def _message_text(content: Any) -> str:
return ""


def _messages_for_disk(messages: list[dict[str, Any]]) -> list[dict[str, Any]]:
"""Copy history while removing file contents returned by intake tools.

The live agent keeps full results in memory for the current conversation.
Persisted chat is resumable context, not a second copy of every file the
user authorized Arbor to inspect.
"""

sanitized: list[dict[str, Any]] = []
for message in messages:
if message.get("_internal") == "context_summary":
sanitized.append({
"role": "user",
"_internal": "context_summary",
"content": (
"[compacted context omitted from persisted intake history; "
"restate the current goal and re-authorize any needed paths]"
),
})
continue
content = message.get("content")
if not isinstance(content, list):
sanitized.append(dict(message))
continue
blocks: list[Any] = []
for block in content:
if isinstance(block, dict) and block.get("type") == "tool_result":
blocks.append({
**block,
"content": (
"[tool result omitted from persisted intake history; "
"ask the user to re-authorize the path before re-reading]"
),
})
else:
blocks.append(block)
sanitized.append({**message, "content": blocks})
return sanitized


def _atomic_write_json(path: Path, data: dict[str, Any]) -> None:
path.parent.mkdir(parents=True, exist_ok=True)
payload = json.dumps(data, indent=2, ensure_ascii=False)
Expand Down
Loading
Loading