What happened
CI run 34555573487 (the "Real Harness (0.1.2-alpha.5)" check on PR #147, head 071e2e5 — a branch that touches none of the harness runtime fixtures) failed the captain-idle-wakeup scenario. Every behavioral assertion was green — member executed, task terminal, captain actually yielded, notified after idle — and then the process exited 1 with a single stderr line from the preserved evidence:
Error: session "01271676-d9f8-4ee6-a438-0565a3d2a40a" is not live in this store
The session id in the error matches the worker member's subagent session.
Likely mechanism
scripts/fixtures/harness-runtime-idle.mjs ends with a flush loop after captain-notified-after-yield:
for (const agent of ctx.agents.list()) {
await ctx.sessions.flush(agent.session);
}
If a member's continuable session has already been retired from the store by the time the loop reaches it, the host session store throws session ... is not live in this store and the scenario exits 1 even though the behavior under test completed successfully. Since the compatibility gate treats any scenario failure as fatal, one hit of this race turns an otherwise-green PR red.
Reproduction status
Not reproducible locally: on the documented alpha.5 flow (pnpm pack + harness-runtime-verify.mjs --host-version 0.1.2-alpha.5), the branch candidate passed all 7 scenarios, and so did a candidate built from plain 426024f, with the same request counts as the CI failure run (12 captain / 6 member). So this reads as a timing race between the teardown flush loop and member session retirement rather than a behavioral regression — but it did fire once on CI, and the window looks structural (the loop flushes whatever agents.list() returns, without tolerating sessions that retired mid-teardown).
Suggestion
Make the teardown flush tolerant of already-retired sessions (skip or catch the "not live" error per agent) so the race window cannot fail an otherwise-green run. If that direction looks right, happy to send the small PR.
What happened
CI run 34555573487 (the "Real Harness (0.1.2-alpha.5)" check on PR #147, head 071e2e5 — a branch that touches none of the harness runtime fixtures) failed the captain-idle-wakeup scenario. Every behavioral assertion was green — member executed, task terminal, captain actually yielded, notified after idle — and then the process exited 1 with a single stderr line from the preserved evidence:
The session id in the error matches the worker member's subagent session.
Likely mechanism
scripts/fixtures/harness-runtime-idle.mjsends with a flush loop aftercaptain-notified-after-yield:If a member's continuable session has already been retired from the store by the time the loop reaches it, the host session store throws
session ... is not live in this storeand the scenario exits 1 even though the behavior under test completed successfully. Since the compatibility gate treats any scenario failure as fatal, one hit of this race turns an otherwise-green PR red.Reproduction status
Not reproducible locally: on the documented alpha.5 flow (
pnpm pack+harness-runtime-verify.mjs --host-version 0.1.2-alpha.5), the branch candidate passed all 7 scenarios, and so did a candidate built from plain426024f, with the same request counts as the CI failure run (12 captain / 6 member). So this reads as a timing race between the teardown flush loop and member session retirement rather than a behavioral regression — but it did fire once on CI, and the window looks structural (the loop flushes whateveragents.list()returns, without tolerating sessions that retired mid-teardown).Suggestion
Make the teardown flush tolerant of already-retired sessions (skip or catch the "not live" error per agent) so the race window cannot fail an otherwise-green run. If that direction looks right, happy to send the small PR.