Skip to content

RFC: compact PR gates and scheduled full CI on GitHub Actions #312

Description

@thomasklemm

Summary

I'd like feedback on a GitHub Actions-first CI shape: a compact, real PR gate that includes Rails → compiled Spinel execution; targeted extra checks; regular full validation of the latest main; and publication bundled with that full run.

This follows the queue-pressure discussion in #273 and overlaps with @dai199's Bazel/BuildBuddy work in #309. Thank you for that pilot — there are useful build-graph and artifact-sharing ideas we can adopt independently of the execution platform. This proposal is not a request to close #309 or a conclusion that remote execution has no value.

The direction below is what we're preparing. The workflow mechanics and optimizations still need review and an Actions pilot; this is not a landed CI replacement.

Revision: The original nine-execution proposal did not include a required Spinel AOT path. Since Rails → compiled Spinel is a core use case, the proposed default gate now adds native Spinel tests and a Rails-vs-compiled-app comparison. Ruby execution alone cannot establish that the Spinel output works.

Proposed shape

Event Validation Publication
Draft PR Fixture generation + unit/default suite None
Normal PR Compact baseline described below, including the required Spinel reference path None
Clearly target-specific PR Compact gate + relevant target/WASM/IDE/Spinel lanes None
Deliberately requested full PR run Full validation of that PR's exact tested merge tree, including broader Spinel coverage None
Four-hour schedule Full matrix on a snapshot of the latest canonical main, when relevant inputs changed; retain broad Spinel coverage and explicitly observe current upstream Spinel Site and downloads from that same snapshot
Manual full-main run Fresh full validation of the selected canonical-main revision Subject to the publication policy below

Scheduled runs would not queue one full matrix for every intermediate main commit. Superseded PR runs would still cancel, including draft/ready transitions. Background matrix parallelism should be bounded so it does not recreate the PR queue pressure we're trying to reduce; the appropriate bound needs measurement.

Proposed execution flow

The diagrams show the desired proposal, including the required Spinel reference checks. They are not a claim that every branch is already implemented.

PR path — select coverage, execute it, then aggregate actual results

┌──────────────────────────────┐
│ PR event / tested merge tree │
└──────────────┬───────────────┘
               ▼
┌──────────────────────────────────────────────┐
│ Plan coverage from the whole PR change      │
│ Draft? Explicit full request? Owned targets?│
└──────────────┬───────────────────────────────┘
               │
       ┌───────┴─────────┐
       ▼                 ▼
┌────────────────┐  ┌───────────────────────────────────────┐
│ Draft          │  │ Non-draft: original compact baseline │
│ Fixtures + unit│  │ + required Spinel reference path     │
└───────┬────────┘  └───────────────────┬───────────────────┘
        │                              │
        │          ┌───────────────────┴───────────────────┐
        │          ▼                                       ▼
        │  ┌──────────────────────────┐  ┌─────────────────────────┐
        │  │ Fixed, verified Spinel   │  │ Other selected coverage │
        │  │ toolchain, prepared once│  │ Owned changes → extras  │
        │  └─────────────┬────────────┘  │ Full request → full set │
        │                ▼               └─────────────┬───────────┘
        │  ┌──────────────────────────┐                │
        │  │ Native model/controller │                │
        │  │ and runtime tests       │                │
        │  │ + Rails DOM/JSON compare│                │
        │  └─────────────┬────────────┘                │
        │                │                             │
        └────────────────┴──────────────┬──────────────┘
                                        ▼
                    ┌───────────────────────────────────┐
                    │ Stable PR completion gate        │
                    │ Every selected required check ran│
                    │ and passed; advisory shown apart │
                    └───────────────────────────────────┘

The non-draft baseline always includes reference Spinel; target routing only adds coverage. Drafts remain fixture + unit. Native tests and the server comparison are separately reported validations, not necessarily sequential jobs. A toolchain failure or missing selected consumer must not produce a green required gate. Full requests retain the advisory/required distinction rather than silently treating every lane as blocking.

Full-main path — one source snapshot, shared outputs, explicit publication gate

┌───────────────────────────────────────┐
│ Four-hour schedule / manual full-main│
└──────────────────┬────────────────────┘
                   ▼
┌─────────────────────────────────────────────────┐
│ Select canonical-main source snapshot           │
│ Record reference pin + observed upstream Spinel │
│ and other relevant input identities             │
└──────────────────┬──────────────────────────────┘
                   ▼
┌─────────────────────────────────────────────────┐
│ Schedule: reuse only valid completed evidence   │
│ for matching relevant inputs                   │
│ Manual/rerun/uncertain inputs: fresh execution  │
└────────────┬─────────────────────────┬──────────┘
             │ Work needed             │ Valid match
             ▼                         ▼
┌──────────────────────────────┐  ┌────────────────────────┐
│ Run full selected coverage   │  │ No duplicate full run │
│ Required compact baseline   │  │ Report evidence reused│
│ incl. reference Spinel      │  └────────────────────────┘
│ + broader target coverage   │
│ + advisory upstream Spinel  │
└──────────────┬───────────────┘
               ▼
┌─────────────────────────────────────────────────┐
│ Shared producer artifacts → archive/browser     │
│ validation → reports + assembled site           │
│ Preserve source/toolchain identity and bytes    │
└──────────────────┬──────────────────────────────┘
                   ▼
┌─────────────────────────────────────────────────┐
│ Retain produced reports and repro archives      │
│ even when validation fails; label outcomes      │
└──────────────────┬──────────────────────────────┘
                   ▼
┌─────────────────────────────────────────────────┐
│ Publication eligible?                           │
│ Canonical main + publishing enabled             │
│ + required compact baseline passed              │
│ + assembled site available                      │
│ + validated revision still eligible to publish  │
└────────────┬─────────────────────────┬───────────┘
             │ Yes                     │ No
             ▼                         ▼
┌────────────────────────────┐  ┌─────────────────────────┐
│ Publish the same artifacts │  │ No live publication    │
│ No rebuild from newer main│  │ Keep diagnostic outputs│
└────────────────────────────┘  └─────────────────────────┘

This is a dependency sketch, not a instruction to serialize all full-matrix work. Independent tests and artifact consumers should run in parallel within a measured concurrency budget. Failures in extra/advisory target coverage must remain visible but need not block publication when its required baseline passed. Archives built with observed upstream Spinel must report that toolchain and their own validation outcome; passing reference checks cannot certify those different binaries.

Implementation note: Initial local work covers coverage planning/routing, scheduled/manual orchestration, selected archive production, result aggregation, and publication boundaries. It is still unfinished and unverified, with no published implementation PR or live Actions pilot yet. The required fixed-reference Spinel producer + native tests + Rails comparison are proposed additions, not yet implemented. Resolving upstream master to one SHA for a full run is useful run consistency, but is not the same as establishing a verified reference or a required Spinel merge gate.

Compact PR gate

Retain the original nine-execution baseline:

  • Fixture generation and unit/default suite — including cargo test --all-targets and the existing emitted-Ruby regression coverage.
  • Store check.
  • Rails DOM comparisons: Ruby, Rust, TypeScript.
  • TypeScript SharedWorker browser smoke.
  • Campfire conformance.
  • Campfire comparison / DB differential.

Add a required Spinel reference path:

  • Provide one complete Spinel toolchain built from a fixed, verified upstream revision, shared by its consumers.
  • Native tests (toolchain-spinel): emit the real-blog project and compile/run its model/controller and additional runtime tests as native programs.
  • Rails comparison (compare-spinel): emit and compile the blog server, then compare its five HTML pages and two JSON endpoints against live Rails.

These checks complement each other: native tests cover operations beyond the selected page responses, while the comparison verifies actual server behavior against Rails. This is a small end-to-end path, not a claim of complete Rails support. The ordinary unit suite does not replace these explicitly ignored toolchain tests.

The final number of Actions jobs depends on producer sharing and grouping; we should not promise a job count or speedup before measuring the implementation.

Clearly owned emitter/runtime changes would add their corresponding lanes. WASM/IDE changes would add relevant browser coverage. Broad shared analyzer/lowerer/runtime changes would retain the compact floor including Spinel, rather than automatically expanding every PR to the full matrix; reviewers could deliberately request full validation for risky work.

The trade-off is explicit: some currently required target lanes would leave the default PR gate. Uncommon target-specific regressions could therefore first appear after merge. Full coverage remains available regularly and on request, but this is not equivalent to full pre-merge validation of every change. We'd especially welcome feedback on whether the proposed floor is representative enough.

Spinel reference validation vs upstream observation

Current Spinel validation lanes follow upstream master and are advisory. Merely retaining those advisory jobs would not make Spinel a merge gate. The proposal separates two roles:

Role Toolchain Failure policy
Default PR reference path Fixed Spinel revision, first verified against the current canonical Roundhouse baseline Required: toolchain preparation, native tests, and comparison must actually succeed; no silent fallback to advisory
Upstream compatibility observation Current upstream Spinel revision, resolved once and recorded for the run Advisory, with individual failures and unavailable/skipped consumers reported explicitly

The reference revision still needs to be selected and verified; an old green snapshot is not evidence that it works with today's Roundhouse. Updates to the reference should be deliberate and tested. This must not hide upstream regressions: the scheduled/manual full-validation capability retains moving-upstream observation.

The broader Spinel coverage remains in the full matrix: framework and focused runtime tests (parameters, signed values, crypto, DB connection leases), blog archive/README/browser smoke, Campfire compiled Rails/WebSocket comparison across GC modes, Campfire DB differential, and Campfire archive/browser/Docker smoke. Directly affected PRs should request the relevant additional coverage. Existing documented exclusions and unsupported cases remain visible.

Each result and published/repro artifact must identify the Spinel revision actually used. A successful reference-toolchain check cannot certify an archive built with a different upstream compiler.

Correctness boundaries to settle

1. A selected check must really have run

Use a stable PR completion gate that verifies every selected required lane completed successfully. A conditionally skipped or missing selected job must not count as a passing test. Supplemental required-lane failures must fail the PR gate; explicitly advisory failures remain visible and non-blocking. The required Spinel producer and both reference validations are part of this rule.

Routing must consider the whole PR change, including renames/deletions, not just the latest commit. Changes to CI selection and shared test infrastructure need an explicit conservative policy. Requested full PR validation must test the actual PR merge tree, not accidentally dispatch a test of main or an unrelated branch head.

2. Unchanged main is not necessarily unchanged input

The upstream-observation Spinel path tracks moving code, and fixture/package generation can resolve newer external dependencies. Resolve shared moving inputs once per full run and record their identities so all consumers use the same snapshot within each toolchain role.

The scheduled skip decision must account for relevant external inputs, not only the Roundhouse SHA. Without a reliable input identity, the affected drift checks should execute rather than assume freshness. An incomplete/cancelled prior run is not proof of completed validation. Manual full runs should remain a way to force fresh execution.

Pinned/reproducible inputs would help, but should not silently remove intentional coverage of new Rails or Spinel versions. Exact checkpoint and drift-check mechanics are still open.

3. Publication and diagnostic artifacts serve different purposes

Build archives once, smoke-test those bytes, and publish those same artifacts from the same main snapshot. Record the source revision, actual toolchain identity, and validation outcomes; don't rebuild against a newer main during publication.

The public site/downloads may lag main until the next background cycle. Publication should require the compact baseline including the Spinel reference path, not blanket success of every target/advisory lane. Individual target or upstream-Spinel failures must not suppress useful repro archives: preserve produced diagnostic artifacts on failures even if live publication is blocked, and distinguish available for reproduction from validated successfully.

External bench-data refreshes are another input to site publication; they need not automatically force an unrelated full compiler matrix.

Ideas to carry over from #309

  1. Share current-run builds, not merely dependency caches. We already share fixtures, WASM, and Spinel. Explore compatible native Roundhouse binaries as producer artifacts too, with explicit debug/release, build-flag, provenance, and native-library boundaries. Consumers must actually use them; ordinary cache hits alone don't eliminate repeated compilation. Pilot a small compatible group first.
  2. Make the producer/consumer graph explicit. Keep archive generation, archive validation, and publication connected through actual outputs. Avoid rebuilding the same emit independently for each stage, but also avoid introducing a single central producer that unnecessarily serializes unrelated checks.
  3. Evaluate versioned toolchain images after measuring setup cost. Purpose-built, digest-pinned images could reduce repeated installation on Actions. Start with expensive common environments rather than assuming one all-language image is best; download time and maintenance may offset the benefit.

These are candidates to evaluate, not commitments to implement all three. We should measure queue-to-first-step time separately from setup/build/test time and compare end-to-end PR feedback, not just job count.

Feedback requested

@rubys @dai199 — I'd appreciate your thoughts before we lock in the workflow boundaries:

  • Is the compact floor plus native Spinel tests and Rails comparison the right pre-merge safety net? Which heavier checks, if any, should also be default?
  • Does separating a required verified Spinel reference from advisory upstream-master observation fit the co-development workflow? Which reference revision should we validate first?
  • Which changes should always require additional lanes or a full PR run, especially CI/harness changes?
  • Does the four-hour full-main/publication cycle preserve the freshness and repro-archive availability you need?
  • Which shared-build or environment improvements from CI on Bazel + BuildBuddy, beside ci.yml (#273) #309 should we pilot first, and where would Bazel/BuildBuddy still offer the clearest additional benefit?

We should validate routing and missing/failed/skipped-job behavior with tests, then run an actual Actions pilot — including fork and full-PR paths, verified-reference Spinel failures, and advisory upstream failures — before changing required branch-protection checks or retiring the existing coverage. Let's coordinate the overlapping work rather than duplicate it.

Activity

  1. thomasklemm commented on Oct 2, 2026

    @thomasklemm
    CollaboratorAuthor

    Actually Spinel-related steps should also be in the PR CI Gate I guess

  2. thomasklemm commented on Oct 2, 2026

    @thomasklemm
    CollaboratorAuthor

    @dai199 Plan above just got an update. Putting an agent to work, guess it makes more sense to see it in a PR, and compare. Would also work several things from #309 in

  3. thomasklemm commented on Oct 2, 2026

    @thomasklemm
    CollaboratorAuthor

    Closing since PR is open and this RFC doesn't match the implementation. You can just comment directly in #313

  4. thomasklemm commented on Oct 4, 2026

    @thomasklemm
    CollaboratorAuthor

    Done in #313

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions