Summary
#2954 made a full memory-index rebuild resumable by capping each call at one slice
(MEMORY_INDEX_REBUILD_DEFAULT_SLICE = 500), but the next slice only runs on the next
daemon indexer round, and a round now spans hours. Result: bounded rebuild jobs start and
never finish, and because outbox consumption is skipped while a job is open, agent cursors
stay frozen exactly as before #2954.
Observed on a live daemon running 885367b9 (v0.39.0, includes #2954 + #2958), started
2026-09-13T13:26:36Z.
Evidence
1. Every open rebuild job is stuck at exactly one slice
~/.holon/.holon/indexes/memory.v2.sqlite3, read-only at 17:00Z:
select count(*), max(documents_processed), sum(documents_processed > 500),
sum(phase = 'finalize') from memory_index_rebuild_jobs;
→ 36 | 500 | 0 | 0
min(started_at) = 2026-09-13T13:27:51Z max(started_at) = 2026-09-13T16:58:29Z
All 36 jobs are phase = scan, source_kind_index = 1, documents_processed = 500, with
last_progress_at ≈ started_at + 1s. A new job starts every ~4–5 minutes; not one job has
ever received a second slice in 3.5 hours.
2. A round is hours long, so a slice is scheduled hours apart
~/.holon/agents: 630 agent directories, 573 of them tmp_child_* (E2E leftovers).
runtime_index_outbox: 63 distinct agent ids, 44 are tmp_child_*.
- These ephemeral agents are enumerated into the work set every round
(src/host.rs:1468-1505: identity records ∪ outbox watermark agents ∪ pending sources),
and each visit is expensive: 36 agent refresh exceeded watchdog threshold (>30 s)
warnings in this window.
So one serialized round over the work set spans hours; the per-agent rebuild cadence is
"one 500-document slice per round", which is slower than the producers.
3. An unfinished job freezes the cursor, not just indexing freshness
refresh_memory_index_bounded (src/memory/index.rs:376-383) skips normal outbox
consumption whenever a rebuild job exists:
let handled_rebuild = index.rebuild_job(&agent_id)?.is_some()
|| !index.rebuild_intents_for_agent(&agent_id)?.is_empty();
index.advance_rebuild(storage, batch_limit.max(1))?;
if !handled_rebuild {
refresh_memory_index_for_storage(...)?; // never runs while the job is open
}
Live index_status for holon-web at 16:55Z: cursor still 947440 (unchanged since
2026-09-12T16:55:57Z — 24 h), high_watermark 970644, lag 19654 → 23204 and growing,
pending_count 1167, stale_reasons includes rebuild_in_progress,
results_may_be_incomplete: true.
4. Collateral cost
database is locked 4187 times in 3.5 h (~0.33/s) while the daemon averages 240 % CPU;
in the same window the front-end read path degraded (221 roster.snapshot spans up to
16.1 s, 243 events.backfill spans totalling 1956 s up to 63.9 s, 24 × 503 on
snapshot/projection-snapshot). Note also did_work |= status.rebuild_phase.is_some()
(src/host.rs:1562-1564) keeps rounds back-to-back while any job is open, so the loop burns
CPU without making progress — the same spin shape as #2939.
Candidate directions (needs design decision)
- Give the indexer round a time budget plus a round-robin continuation position, so a
round cannot span hours and every agent gets a slice per round.
- Enumerate open rebuild jobs explicitly and resume existing jobs before starting new
ones (observed: 36 started, 0 resumed).
- Keep leaked ephemeral agents (
tmp_child_*) out of the work set — and/or clean them up at
E2E teardown; 573 directories and 44 outbox agents currently consume round budget forever.
- Reconsider the
handled_rebuild short-circuit so a long rebuild bounds, rather than
freezes, cursor advancement.
Summary
#2954 made a full memory-index rebuild resumable by capping each call at one slice
(
MEMORY_INDEX_REBUILD_DEFAULT_SLICE = 500), but the next slice only runs on the nextdaemon indexer round, and a round now spans hours. Result: bounded rebuild jobs start and
never finish, and because outbox consumption is skipped while a job is open, agent cursors
stay frozen exactly as before #2954.
Observed on a live daemon running
885367b9(v0.39.0, includes #2954 + #2958), started2026-09-13T13:26:36Z.
Evidence
1. Every open rebuild job is stuck at exactly one slice
~/.holon/.holon/indexes/memory.v2.sqlite3, read-only at 17:00Z:All 36 jobs are
phase = scan,source_kind_index = 1,documents_processed = 500, withlast_progress_at ≈ started_at + 1s. A new job starts every ~4–5 minutes; not one job hasever received a second slice in 3.5 hours.
2. A round is hours long, so a slice is scheduled hours apart
~/.holon/agents: 630 agent directories, 573 of themtmp_child_*(E2E leftovers).runtime_index_outbox: 63 distinct agent ids, 44 aretmp_child_*.(
src/host.rs:1468-1505: identity records ∪ outbox watermark agents ∪ pending sources),and each visit is expensive: 36
agent refresh exceeded watchdog threshold(>30 s)warnings in this window.
So one serialized round over the work set spans hours; the per-agent rebuild cadence is
"one 500-document slice per round", which is slower than the producers.
3. An unfinished job freezes the cursor, not just indexing freshness
refresh_memory_index_bounded(src/memory/index.rs:376-383) skips normal outboxconsumption whenever a rebuild job exists:
Live
index_statusforholon-webat 16:55Z:cursorstill947440(unchanged since2026-09-12T16:55:57Z — 24 h),
high_watermark970644,lag19654 → 23204 and growing,pending_count1167,stale_reasonsincludesrebuild_in_progress,results_may_be_incomplete: true.4. Collateral cost
database is locked4187 times in 3.5 h (~0.33/s) while the daemon averages 240 % CPU;in the same window the front-end read path degraded (221
roster.snapshotspans up to16.1 s, 243
events.backfillspans totalling 1956 s up to 63.9 s, 24 × 503 onsnapshot/projection-snapshot). Note also
did_work |= status.rebuild_phase.is_some()(
src/host.rs:1562-1564) keeps rounds back-to-back while any job is open, so the loop burnsCPU without making progress — the same spin shape as #2939.
Candidate directions (needs design decision)
round cannot span hours and every agent gets a slice per round.
ones (observed: 36 started, 0 resumed).
tmp_child_*) out of the work set — and/or clean them up atE2E teardown; 573 directories and 44 outbox agents currently consume round budget forever.
handled_rebuildshort-circuit so a long rebuild bounds, rather thanfreezes, cursor advancement.