Describe the bug
I keep hitting a corrupt .moon/cache/states/workspaceGraph.json in a monorepo where several moon processes overlap, and I had an AI agent dig into it on my behalf. Everything below — the measurements, the scripts, the reading of the code — is its work, run on my machine, and I'm passing it on in case it saves you time.
The symptom, roughly a dozen times in one day:
Failed to parse JSON file .moon/cache/states/workspaceGraph.json
missing field `projects` at line 1 column 267866
The column equals the file's own size. Once it happens, every later invocation fails identically — json::read_file(&cache_path)? in WorkspaceBuilderAsync::new_with_cache propagates, so nothing rewrites the file. rm -rf .moon/cache buys exactly one clean run. Because the message repeats byte for byte, it reads as a deterministic serialization bug, which is what sent us the wrong way for a day.
What the agent found instead was two things about concurrent invocations.
1. Two moon processes can be inside the graph-cache lock at the same time.
CacheEngine::create_lock sets remove_on_unlock, and starbase_utils's FileLock::unlock removes the lock file before releasing the flock. A process that opens the path after that unlink, while the previous holder is still in the critical section, gets a new inode and locks that one instead.
2. The cache files are written in place.
json::write_file ends in std::fs::write — O_TRUNC, then write. A reader can therefore observe a zero-length or partial file, and a moon killed mid-write leaves a partial file on disk.
Steps to reproduce
Setup — a synthetic workspace big enough that graph building takes a moment:
mkdir -p /tmp/moonrepro && cd /tmp/moonrepro && git init -q .
mkdir -p .moon && printf 'projects:\n - "packages/*/moon.yml"\nvcs:\n defaultBranch: "main"\n' > .moon/workspace.yml
python3 - <<'PY'
import os
for i in range(600):
d = f'/tmp/moonrepro/packages/pkg{i:03d}'
os.makedirs(d, exist_ok=True)
tasks = '\n'.join(f" task{j}:\n command: 'echo hello-from-pkg{i:03d}-task{j}-padding'\n options:\n cache: false" for j in range(6))
open(f'{d}/moon.yml', 'w').write(f"id: 'pkg{i:03d}'\nlanguage: 'unknown'\ntasks:\n{tasks}\n")
PY
git add -A && git commit -qm init
Then staggered invocations, while a loop rewrites one config so every run has to rebuild and rewrite the cache:
( for i in $(seq 1 400); do
printf "id: 'pkg000'\nlanguage: 'unknown'\ntasks:\n task0:\n command: 'echo hi-%0*d'\n options:\n cache: false\n" $((i*53)) 0 > packages/pkg000/moon.yml
sleep 0.1
done ) &
for i in $(seq 1 40); do
moon query projects --log debug > /tmp/run_$i.log 2>&1 &
sleep 0.05
done
wait
The overlap is visible on every run. A run's lock-held window is bracketed by Cache hit, reading item ... workspaceGraphStateV1.json (just after the lock is taken) and Writing cache item ... workspaceGraphStateV1.json (just before the graph is written, still holding it), so overlapping windows are overlapping critical sections:
python3 - <<'PY'
import re, glob
def sec(s):
h, m, rest = s.split(':'); return int(h)*3600 + int(m)*60 + float(rest)
ivs = []
for f in glob.glob('/tmp/run_*.log'):
txt = open(f, errors='replace').read()
s = re.search(r'\[\w+ ([\d:.]+)\][^\n]*Cache hit, reading item[^\n]*workspaceGraphStateV1', txt)
e = re.search(r'\[\w+ ([\d:.]+)\][^\n]*Writing cache item[^\n]*workspaceGraphStateV1', txt)
if s and e: ivs.append((sec(s.group(1)), sec(e.group(1)), f))
ivs.sort()
pairs = [(a, b) for i, a in enumerate(ivs) for b in ivs[i+1:] if b[0] < a[1]]
print('runs measured:', len(ivs), 'pairs holding the lock simultaneously:', len(pairs))
for a, b in pairs[:5]:
print(' %s [%.3f-%.3f] || %s [%.3f-%.3f]' % (a[2], a[0], a[1], b[2], b[0], b[1]))
PY
Five consecutive passes on moon 2.5.2, macOS arm64, reported overlapping pairs every time: 28, 163, 199, 161, 168 pairs out of ~40 runs. The count moves with machine load; the overlap does not.
What we could and could not reproduce
The overlap above: every run, five for five.
The corruption itself: intermittent. Running the same loop without --log debug, 400 runs with ~8 in flight, one run failed outright —
× Failed to parse JSON file /private/tmp/moonrepro/.moon/cache/states/workspaceGraphStateV1.json.
╰─▶ EOF while parsing a value at line 1 column 0
— which is a reader seeing the file between the truncate and the write. Two further 400-run passes produced none.
We could not reproduce the exact "missing field" text from the top of this issue on demand. Truncating the file by hand gives EOF while parsing…, and splicing a short write over a longer one gives trailing characters. The corrupt file from the original incident was deleted before anyone thought to keep it, so that specific byte pattern is unexplained — worth saying plainly rather than glossing over.
Also unclear whether an interrupted writer (Ctrl-C, a reaped background run) accounts for some of our incidents, since it would leave the same partial file with no concurrency involved. That half is easy to produce on demand, if it is useful for testing: run moon query projects under ulimit -f 2048 and moon takes SIGXFSZ partway through its own write, leaving 1048576 bytes of valid-looking JSON that stops mid-string.
With such a file in place, every graph-touching command — query, run, ci, project — fails at the same offset with the same message, and the task never runs. moon clean exits 0 and reports "Deleted 0 files", leaving the file where it is; moon clean --lifetime "0 seconds" does clear it.
Expected behavior
Overlapping moon invocations in one workspace don't leave the cache unreadable.
Environment
moon 2.5.2 (latest release as of 2026-08-19; unchanged since 2.5.1, and the same on master)
macOS 26.5.2, Darwin 25.5.0, arm64
Node 24.11.1, pnpm 11.1.2 (nodeLinker: hoisted)
Additional context
The concurrency in our repo is ordinary: a long-lived moon run of a persistent dev-server task, a moon ci from a pre-push hook, and ad hoc moon run invocations, all in one worktree. Nothing exotic, and nothing that hints at the cause from the outside — the parse error surfaces in place of whatever was being run, so on our side it hid a failing formatting check across three git push attempts, because moon ci died in "Building action graph" before any task executed.
Doesn't look like #2524 to us — that was a per-OFD flock deadlock inside a single process. This is across processes, and it ends in an unreadable cache rather than a hang.
Describe the bug
I keep hitting a corrupt
.moon/cache/states/workspaceGraph.jsonin a monorepo where several moon processes overlap, and I had an AI agent dig into it on my behalf. Everything below — the measurements, the scripts, the reading of the code — is its work, run on my machine, and I'm passing it on in case it saves you time.The symptom, roughly a dozen times in one day:
The column equals the file's own size. Once it happens, every later invocation fails identically —
json::read_file(&cache_path)?inWorkspaceBuilderAsync::new_with_cachepropagates, so nothing rewrites the file.rm -rf .moon/cachebuys exactly one clean run. Because the message repeats byte for byte, it reads as a deterministic serialization bug, which is what sent us the wrong way for a day.What the agent found instead was two things about concurrent invocations.
1. Two moon processes can be inside the graph-cache lock at the same time.
CacheEngine::create_locksetsremove_on_unlock, andstarbase_utils'sFileLock::unlockremoves the lock file before releasing theflock. A process that opens the path after that unlink, while the previous holder is still in the critical section, gets a new inode and locks that one instead.2. The cache files are written in place.
json::write_fileends instd::fs::write—O_TRUNC, then write. A reader can therefore observe a zero-length or partial file, and a moon killed mid-write leaves a partial file on disk.Steps to reproduce
Setup — a synthetic workspace big enough that graph building takes a moment:
Then staggered invocations, while a loop rewrites one config so every run has to rebuild and rewrite the cache:
The overlap is visible on every run. A run's lock-held window is bracketed by
Cache hit, reading item ... workspaceGraphStateV1.json(just after the lock is taken) andWriting cache item ... workspaceGraphStateV1.json(just before the graph is written, still holding it), so overlapping windows are overlapping critical sections:Five consecutive passes on moon 2.5.2, macOS arm64, reported overlapping pairs every time: 28, 163, 199, 161, 168 pairs out of ~40 runs. The count moves with machine load; the overlap does not.
What we could and could not reproduce
The overlap above: every run, five for five.
The corruption itself: intermittent. Running the same loop without
--log debug, 400 runs with ~8 in flight, one run failed outright —— which is a reader seeing the file between the truncate and the write. Two further 400-run passes produced none.
We could not reproduce the exact "missing field" text from the top of this issue on demand. Truncating the file by hand gives
EOF while parsing…, and splicing a short write over a longer one givestrailing characters. The corrupt file from the original incident was deleted before anyone thought to keep it, so that specific byte pattern is unexplained — worth saying plainly rather than glossing over.Also unclear whether an interrupted writer (Ctrl-C, a reaped background run) accounts for some of our incidents, since it would leave the same partial file with no concurrency involved. That half is easy to produce on demand, if it is useful for testing: run
moon query projectsunderulimit -f 2048and moon takesSIGXFSZpartway through its own write, leaving 1048576 bytes of valid-looking JSON that stops mid-string.With such a file in place, every graph-touching command —
query,run,ci,project— fails at the same offset with the same message, and the task never runs.moon cleanexits 0 and reports "Deleted 0 files", leaving the file where it is;moon clean --lifetime "0 seconds"does clear it.Expected behavior
Overlapping moon invocations in one workspace don't leave the cache unreadable.
Environment
Additional context
The concurrency in our repo is ordinary: a long-lived
moon runof a persistent dev-server task, amoon cifrom a pre-push hook, and ad hocmoon runinvocations, all in one worktree. Nothing exotic, and nothing that hints at the cause from the outside — the parse error surfaces in place of whatever was being run, so on our side it hid a failing formatting check across threegit pushattempts, becausemoon cidied in "Building action graph" before any task executed.Doesn't look like #2524 to us — that was a per-OFD
flockdeadlock inside a single process. This is across processes, and it ends in an unreadable cache rather than a hang.