Skip to content

feat(control): in-place workflow reruns (GitHub-faithful) + attempt history - #374

Draft
Bnjoroge1 wants to merge 6 commits into
mainfrom
gha-reruns-port
Draft

Bnjoroge1 wants to merge 6 commits into
mainfrom
gha-reruns-port

Conversation

@Bnjoroge1

@Bnjoroge1 Bnjoroge1 commented Oct 5, 2026 •

Copy link
Copy Markdown
Collaborator

Re-run workflow attempts in place (GitHub-faithful)

Base: Bnjoroge/lock-refactor · Head: gha-reruns-port

Gate green on this head: fmt-check clean, clippy -D warnings clean, full lib 848 passed/0 failed/3 ignored at --test-threads=1, 11 integration binaries all 0 failed, preloop-cli 206/0, focused migration+rerun lite 15/0 + PG 14/0.

Draft because of merge order, not code state: the schema migration below has been ported to the refinery migration branch (#372, V2026100503). Before merge this branch must rebase onto #372 and drop its in-branch migration plumbing (migrate_3_to_4/migrate_4_to_5, SCHEMA_VERSION bumps). The feature code and tests are final.

Problem

preloop's "re-run" resubmitted the recorded submission as a brand-new run: new run_id, run_attempt always 1, the old run left terminal. GitHub re-runs in place — same run id, run_attempt incremented, the previous attempt archived to history — and "Re-run failed jobs" / per-job re-runs only re-execute the failed cone. Tools (gh run rerun, Checks-page rerequests, the REST shims) expect that.

Design

  • One write transaction resets the mode-selected jobs on the same run_id: attempt snapshot into run_history/job_history keyed by run_attempt, runs.run_attempt += 1, github.run_attempt bumped in the stored context, fresh jobId/timeline.id/requestId per selected job.
  • Modes: All, Failed (failed/cancelled/timed-out seeds + dependent closure), Job(id) (+ dependents); reusable-caller subtrees reset wholesale; expanded matrix parents are never selected (their legs are).
  • Re-admission mirrors submit_run step for step: workflow concurrency, if: (needs-less inline, needs-gated in promotion), platform/pool labels, environment gate (re-armed per attempt), job concurrency, max-parallel.
  • Jobs outside the selection keep status + outputs, so dependents see the carried-forward needs context.
  • Surfaces: native POST /api/v1/runs/:id/rerun {mode, job_id?} (absent body = all), CLI preloop rerun [RUN_ID] [--failed|--job ID], and GitHub-compat POST /repos/{o}/{r}/actions/runs/{id}/rerun, /rerun-failed-jobs, /actions/jobs/{job_id}/rerun, /actions/runs/{id}/cancel (dispatch auth, actions: write; 201/202/403/404/409/422; /repos/... errors carry GitHub's message + documentation_url beside preloop's error).
  • check_run.rerequested (job mode) and check_suite.rerequested (all mode) now re-run in place; check-run reporting is scoped to the reset set so carried-forward jobs keep their existing checks.
  • Archiver hold: PRELOOP_RERUN_WINDOW_DAYS (default 30, 0 disables) keeps a completed run with ≥1 failed/cancelled/timed-out job in the live tables (60 s grace, 5 s ticks; retention still wins). A fully archived run falls back to the old behavior for all (new run from the recorded submission) and 409 for failed/job.
  • Schema migration: run_history/job_history gain run_attempt in their primary keys, backfilling run_attempt = 1 for pre-change rows and preserving identities. The tested SQL lives in the migrations branch as V2026100503 (refinery, forward-only); this branch's own version-stamp plumbing will be dropped at rebase.

Conflict points

Known gaps

  • In-place rehydration of fully archived runs is a follow-up (all starts a new run; partial modes 409).
  • Per-attempt log separation: logs stay per run (all attempts accumulate), not per attempt like github.com's UI.

Summary by cubic

Rerunning a completed workflow now replays it in place on the same run id, like GitHub's "Re-run jobs": run_attempt increments, the previous attempt is snapshotted into run_history/job_history, and selected jobs are reset and re-admitted through the same concurrency, if:, platform, and environment-gate chain a fresh submit uses. Previously a rerun started a brand-new run with a new run_id.

  • Three modes mirror GitHub's buttons: all jobs, failed/cancelled jobs plus dependents, and one job plus dependents.
  • Re-run surfaces: POST /api/v1/runs/:id/rerun (absent body = all), the new preloop rerun CLI, GitHub-compatible /repos/.../actions/runs/{id}/rerun, /rerun-failed-jobs, /actions/jobs/{job_id}/rerun, and /cancel shims, plus in-place handling of check_run/check_suite rerequested. The shims return GitHub-shaped errors (message + documentation_url) alongside preloop's error key and reject rerun bodies with enable_debug_logging/enable_debugger as 422.
  • Jobs outside the selection keep their results and outputs, so dependents see the carried-forward needs context.
  • A fully archived run can only be re-run in full, which starts a new run (old behavior); partial modes return 409. PRELOOP_RERUN_WINDOW_DAYS (default 30, 0 disables) holds completed runs with failed jobs in the live tables so the in-place path stays reachable; retention still wins.
  • Schema: run_history/job_history add run_attempt to their primary keys. SQLite 4→5 and PostgreSQL 5→6 migrate on first open — one rollback-safe transaction — backfilling pre-change rows to attempt 1 while keeping their identities.

Written for commit e65da3f. Summary will update on new commits.

Review in cubic Turn on auto-fix

@coderabbitai

coderabbitai Bot commented Oct 5, 2026

Copy link
Copy Markdown

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Bnjoroge1 and others added 5 commits October 7, 2026 00:24
Re-run a completed run as a new attempt on the same run id: run_attempt
increments, the previous attempt is snapshotted into run_history/job_history,
and the mode-selected jobs (all / failed+dependents / one job+dependents) are
reset and re-admitted through the same workflow-concurrency, if:, platform,
environment-gate, job-concurrency and max-parallel chain a submit uses.
Non-selected jobs keep their results and outputs for dependents' needs.

Native POST /api/v1/runs/:id/rerun takes {mode,job_id?}; check_run.rerequested
re-runs the affected job in place and check_suite.rerequested re-runs the
suite's run in place. GitHub-compatible shims (dispatch auth, actions:write):
actions/runs/{id}/rerun (201), actions/runs/{id}/rerun-failed-jobs (201),
actions/jobs/{job_id}/rerun (201), actions/runs/{id}/cancel (202, 409 when
terminal). New preloop rerun CLI.

Fully archived runs (no live rows) keep lock-refactor's behavior for a full
re-run (new run from the recorded submission) and return 409 for partial
modes; in-place rehydration is a follow-up. PRELOOP_RERUN_WINDOW_DAYS
(default 30, 0 disables) holds rerun-useful completed runs in the live tables
so the in-place path is reachable; retention still wins. Fixes the PG
archiver's job_history insert for the new run_attempt column.
- Close stale NULL-result job_requests before a rerun mints its next attempt
  (SQLite; mirrors the PG path) so a forced-terminal run can rerun.
- Bump control schema versions (SQLite 4, PostgreSQL 5) for the run_attempt
  primary-key change; add the shared-suite backdate hook.
- Validate the GitHub rerun body (enable_debug_logging/enable_debugger) with
  422, and add a GitHub-shaped error envelope (message + documentation_url)
  for /repos/... while keeping the legacy error key.
- Clippy: needless borrow in cancel_run_inner; RunHead struct for the rerun
  head tuple.
SQLite 3->4 and PostgreSQL 4->5 migrate on first open instead of refusing:
run_history/job_history gain run_attempt in their primary keys, pre-change
rows are attempt-1 history and keep their identities, and the whole step is
one transaction so a crash rolls it back and the next open retries. Older
versions are still refused. Tests build populated pre-change databases
(lite: schema downgrade on disk incl. a rolled-back partial step and an
end-to-end rerun+archive; pg: raw-client fixture with a rolled-back ALTER,
plus fresh-creation stamp). CHANGELOG updated from 'recreate' to the
supported migration.
…assertions

- broker: a rerequest reopens the run (non-terminal); assert that instead of
  pinning the projected status.
- dispatch_tests: after archiving, the original run is no longer live, so the
  live working set holds only the newly resubmitted run.
The rebase onto main dropped the per-operation queue count: the sampler
snapshot owns `state.queue_depth` now (main removed the submit-path store
for the same reason). Drop `RerunOutcome.queue_depth` and the pg
`queue_depth()` call it needed (that helper is gone from main), and wake
`sampler_notify` beside the waiter wake so the gauge refreshes early.
@Bnjoroge1
Bnjoroge1 changed the base branch from Bnjoroge/lock-refactor to main October 7, 2026 04:44
main already stamped SQLite v4 / Postgres v5 (check_run_updates gained
version/lease_owner), so the rebased history-PK change could not reuse the
old 3->4 / 4->5 numbers: a database main had just created would have been
accepted without the new columns. Move to SQLite 5 / Postgres 6 with the
previous version bumped accordingly, stamp the pg migration from
SCHEMA_VERSION instead of a literal, and retarget both in-place migration
tests (fixtures are the current schema minus the history-PK change, stamped
at the new previous version).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant