Skip to content

[bug] Concurrent moon invocations can leave workspaceGraph.json unreadable, and every later run then fails identically #2684

Description

@jdillon

Describe the bug

I keep hitting a corrupt .moon/cache/states/workspaceGraph.json in a monorepo where several moon processes overlap, and I had an AI agent dig into it on my behalf. Everything below — the measurements, the scripts, the reading of the code — is its work, run on my machine, and I'm passing it on in case it saves you time.

The symptom, roughly a dozen times in one day:

Failed to parse JSON file .moon/cache/states/workspaceGraph.json
  missing field `projects` at line 1 column 267866

The column equals the file's own size. Once it happens, every later invocation fails identically — json::read_file(&cache_path)? in WorkspaceBuilderAsync::new_with_cache propagates, so nothing rewrites the file. rm -rf .moon/cache buys exactly one clean run. Because the message repeats byte for byte, it reads as a deterministic serialization bug, which is what sent us the wrong way for a day.

What the agent found instead was two things about concurrent invocations.

1. Two moon processes can be inside the graph-cache lock at the same time.

CacheEngine::create_lock sets remove_on_unlock, and starbase_utils's FileLock::unlock removes the lock file before releasing the flock. A process that opens the path after that unlink, while the previous holder is still in the critical section, gets a new inode and locks that one instead.

2. The cache files are written in place.

json::write_file ends in std::fs::write — O_TRUNC, then write. A reader can therefore observe a zero-length or partial file, and a moon killed mid-write leaves a partial file on disk.

Steps to reproduce

Setup — a synthetic workspace big enough that graph building takes a moment:

mkdir -p /tmp/moonrepro && cd /tmp/moonrepro && git init -q .
mkdir -p .moon && printf 'projects:\n  - "packages/*/moon.yml"\nvcs:\n  defaultBranch: "main"\n' > .moon/workspace.yml
python3 - <<'PY'
import os
for i in range(600):
    d = f'/tmp/moonrepro/packages/pkg{i:03d}'
    os.makedirs(d, exist_ok=True)
    tasks = '\n'.join(f"  task{j}:\n    command: 'echo hello-from-pkg{i:03d}-task{j}-padding'\n    options:\n      cache: false" for j in range(6))
    open(f'{d}/moon.yml', 'w').write(f"id: 'pkg{i:03d}'\nlanguage: 'unknown'\ntasks:\n{tasks}\n")
PY
git add -A && git commit -qm init

Then staggered invocations, while a loop rewrites one config so every run has to rebuild and rewrite the cache:

( for i in $(seq 1 400); do
    printf "id: 'pkg000'\nlanguage: 'unknown'\ntasks:\n  task0:\n    command: 'echo hi-%0*d'\n    options:\n      cache: false\n" $((i*53)) 0 > packages/pkg000/moon.yml
    sleep 0.1
  done ) &
for i in $(seq 1 40); do
  moon query projects --log debug > /tmp/run_$i.log 2>&1 &
  sleep 0.05
done
wait

The overlap is visible on every run. A run's lock-held window is bracketed by Cache hit, reading item ... workspaceGraphStateV1.json (just after the lock is taken) and Writing cache item ... workspaceGraphStateV1.json (just before the graph is written, still holding it), so overlapping windows are overlapping critical sections:

python3 - <<'PY'
import re, glob
def sec(s):
    h, m, rest = s.split(':'); return int(h)*3600 + int(m)*60 + float(rest)
ivs = []
for f in glob.glob('/tmp/run_*.log'):
    txt = open(f, errors='replace').read()
    s = re.search(r'\[\w+ ([\d:.]+)\][^\n]*Cache hit, reading item[^\n]*workspaceGraphStateV1', txt)
    e = re.search(r'\[\w+ ([\d:.]+)\][^\n]*Writing cache item[^\n]*workspaceGraphStateV1', txt)
    if s and e: ivs.append((sec(s.group(1)), sec(e.group(1)), f))
ivs.sort()
pairs = [(a, b) for i, a in enumerate(ivs) for b in ivs[i+1:] if b[0] < a[1]]
print('runs measured:', len(ivs), 'pairs holding the lock simultaneously:', len(pairs))
for a, b in pairs[:5]:
    print('  %s [%.3f-%.3f] || %s [%.3f-%.3f]' % (a[2], a[0], a[1], b[2], b[0], b[1]))
PY

Five consecutive passes on moon 2.5.2, macOS arm64, reported overlapping pairs every time: 28, 163, 199, 161, 168 pairs out of ~40 runs. The count moves with machine load; the overlap does not.

What we could and could not reproduce

The overlap above: every run, five for five.

The corruption itself: intermittent. Running the same loop without --log debug, 400 runs with ~8 in flight, one run failed outright —

× Failed to parse JSON file /private/tmp/moonrepro/.moon/cache/states/workspaceGraphStateV1.json.
╰─▶ EOF while parsing a value at line 1 column 0

— which is a reader seeing the file between the truncate and the write. Two further 400-run passes produced none.

We could not reproduce the exact "missing field" text from the top of this issue on demand. Truncating the file by hand gives EOF while parsing…, and splicing a short write over a longer one gives trailing characters. The corrupt file from the original incident was deleted before anyone thought to keep it, so that specific byte pattern is unexplained — worth saying plainly rather than glossing over.

Also unclear whether an interrupted writer (Ctrl-C, a reaped background run) accounts for some of our incidents, since it would leave the same partial file with no concurrency involved. That half is easy to produce on demand, if it is useful for testing: run moon query projects under ulimit -f 2048 and moon takes SIGXFSZ partway through its own write, leaving 1048576 bytes of valid-looking JSON that stops mid-string.

With such a file in place, every graph-touching command — query, run, ci, project — fails at the same offset with the same message, and the task never runs. moon clean exits 0 and reports "Deleted 0 files", leaving the file where it is; moon clean --lifetime "0 seconds" does clear it.

Expected behavior

Overlapping moon invocations in one workspace don't leave the cache unreadable.

Environment

moon 2.5.2 (latest release as of 2026-08-19; unchanged since 2.5.1, and the same on master)
macOS 26.5.2, Darwin 25.5.0, arm64
Node 24.11.1, pnpm 11.1.2 (nodeLinker: hoisted)

Additional context

The concurrency in our repo is ordinary: a long-lived moon run of a persistent dev-server task, a moon ci from a pre-push hook, and ad hoc moon run invocations, all in one worktree. Nothing exotic, and nothing that hints at the cause from the outside — the parse error surfaces in place of whatever was being run, so on our side it hid a failing formatting check across three git push attempts, because moon ci died in "Building action graph" before any task executed.

Doesn't look like #2524 to us — that was a per-OFD flock deadlock inside a single process. This is across processes, and it ends in an unreadable cache rather than a hang.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

bugSomething isn't working

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions