Skip to content

Latest commit

 

History

History
209 lines (163 loc) · 10.4 KB

File metadata and controls

209 lines (163 loc) · 10.4 KB

ci-cost — architecture

How the pieces fit, and where the line between deterministic computation and LLM judgment sits. For why the tool exists and what it deliberately doesn't do, see README.md; for the history of decisions and why they changed, CHANGELOG.md.

The load-bearing constraint

No LLM is ever in the cost-computation path. Every dollar figure comes from a Python script reading a versioned rate table. Given a pinned catalog_version and a usage profile, the computed comparison is byte-for-byte reproducible, and any number in a report can be traced to a catalog entry and a measured minute count.

An LLM appears at exactly two points, both outside that path, both behind a human gate:

Where What it decides The gate
add-vendor skill which of a vendor's own machine classes corresponds to which ci-cost runner class a reviewed pull request
ci-cost skill which computed facts matter for this repo, and how to say them the numbers are already fixed; only the prose is generated

Neither can produce a rate. The first proposes one and a person approves it; the second reads one and explains it.

Data flow

                    GitHub REST API
                          │
              pull_github_data.py          measured
                          │                 ────────
                  usage-profile.json
                          │
                          ├────────────────────────────┐
                          │                            │
          compute_current_vendor_cost.py      build_comparison.py
              (single vendor)                  (every vendor)
                          │                            │
                  ci-cost-report.md            comparison.json     facts
                                                       │           ─────
                                            skills/ci-cost/SKILL.md
                                                       │
                                               ci-cost-report.md   prose

Both report paths exist deliberately: the scripts must stay runnable and verifiable without an LLM present, so compute_current_vendor_cost.py is not dead weight — it's the proof that the math works standalone.

Layout

skills/ci-cost/          the installable skill — what `npx skills add` copies
  SKILL.md
  scripts/               stdlib-only; runnable and tested without an LLM
  references/            the versioned pricing catalog, bundled

maintainers/ci-cost/     catalog authoring; never copied into an install
  SKILL.md               the add-vendor skill
  scripts/               scraping, draft validation, catalog PRs
  references/            vendor drafts, scrape staging

tests/                   flat, one test module per feature

Paths inside the skill resolve from __file__ via skill_paths.py, never from the working directory — once installed, cwd is the user's repository, so a relative path points into their project rather than at the bundled catalog.

Maintainer scripts put the skill's scripts/ on sys.path rather than keeping a second copy of the catalog schema, which would drift.

Scripts

Script Responsibility
pull_github_data.py Samples workflow runs and jobs over the GitHub REST API, classifies each job's runner labels, extrapolates sampled minutes per class to a monthly figure. Flags low_confidence when the observed history is too short to trust.
load_pricing_catalog.py Loads and validates skills/ci-cost/references/pricing-catalog.json. Archetype-dependent required fields, finite non-negative rates, duplicate-JSON-key rejection, and the equivalence-map agreement rules.
replay_cross_vendor_cost.py One pure function per pricing archetype over (usage_profile, vendor_entry). Partitions usage into priced / verified-unavailable / unmapped. No network, no LLM.
build_comparison.py Prices a usage profile against every catalog vendor, emits JSON facts. Takes no position on ranking.
compute_current_vendor_cost.py Single-vendor cost plus a rendered markdown report. The standalone, LLM-free path.
scrape_pricing_data.py Maintainer-only. Scrapes vendor pricing pages via Firecrawl into gitignored staging for manual verification. Extracts no rates itself.
validate_vendor_draft.py Gates a vendor draft before it can become a PR. Validates proposed_entry with the real catalog validator, so a passing draft cannot fail once merged.
open_vendor_pr.py Merges a validated draft into the catalog on a branch and opens the PR, with the evidence in the body.

The runner-class vocabulary

Everything keys on the class names classify_runner_labels emits: linux-x64-default, linux-arm64, windows-x64, windows-arm64, macos, self-hosted, plus sized variants (linux-x64-4-core, macos-12-core, …).

Sized names are generated, not enumerated. runner_class_name builds them by interpolating whatever core count appeared in the label, so linux-x64-64-core is valid and no catalog can list every possibility in advance. An equivalence map is therefore always partial by design, and the fail-loudly path is the mechanism that tells the maintainer which entry to add. is_known_runner_class is the single source of truth for what is a legal name.

The catalog

skills/ci-cost/references/pricing-catalog.json is a frozen snapshot, never fetched live. Determinism is a claim about the math, not the rates — which is why every report must print catalog_version, the snapshot date, and a verify-before- acting disclaimer.

Rate fields depend on the archetype, so a credit-priced vendor's arithmetic stays auditable in the file rather than pre-multiplied into a flat rate:

pricing_model Rate fields
per_minute_by_runner_class rates: {class: $/min}
credit_based credits_per_minute: {class: n} + usd_per_credit
self_hosted_compute_estimate rates: {class: $/min}

per_active_user_flat_plus_usage is deliberately not accepted: a per-seat cost isn't a function of the sampled usage profile, so nothing could compute it, and accepting it would let an uncomputable entry pass validation.

Three states, not two

runner_class_equivalents records which of the vendor's own classes each runner class was priced against — and encodes the distinction the whole comparison rests on:

In the map Meaning Effect
"docker/medium" mapped and priced charged
null verified: this vendor has no equivalent excluded from the total, disclosed in the report
absent nobody has mapped it yet hard error

Collapsing the last two would let a forgotten mapping silently drop minutes from a vendor's total and make it look cheaper than it is. That is the single most dangerous failure this tool can have, because nothing downstream can detect it.

Why partial coverage is a first-class concept

A vendor that cannot run part of the workload has a smaller total, so a naive ranking puts it first and it reads as the bargain. Starsling has no macOS runners: for an iOS repo it scores $20.00/mo against GitHub Actions' $49.84 while being unable to build the app.

build_comparison.py therefore emits is_like_for_like, the excluded classes with their minutes, and excluded_cost_on_current_vendor — the excluded minutes priced at the current vendor's rates, because "excludes 320 macOS minutes" only becomes legible as "$19.84/month". Vendors are ordered cheapest-first for stable output, and the payload carries an ordering_note stating that this is not a ranking.

Why fixed fees are a first-class concept

A vendor's per-minute rates are only part of what it charges, and the missing part fails in the same direction as partial coverage: the total comes out smaller than the vendor really is, so it sorts first and reads as the bargain.

base_monthly_cost is therefore a required field on every catalog entry, with the same three-state discipline as runner_class_equivalents:

Value Means Effect on the total
0 verified: no fixed charge before usage total is the usage leg
<n> verified: this much every month added to the usage leg
null a fee exists that cannot be stated as one number total is a floor; has_unmodeled_fees is true
absent — hard error

RunsOn is why this exists. Its per-minute rates are an AWS spot pass-through with no markup, which makes it the cheapest row in most tables, and it separately charges a tiered annual license fee denominated in euros. Neither a snapshotted FX rate nor a guess at the volume tier belongs in a catalog whose claim is reproducibility, so the fee is null with a base_cost_note, and has_unmodeled_fees blocks the vendor from being presented as cheapest.

build_comparison.py emits usage_monthly_cost and base_monthly_cost alongside total_monthly_cost so the arithmetic stays auditable rather than arriving as one figure to be taken on trust.

Why the free tier is priced, and the allowance is not

GitHub does not charge for standard hosted runners on public repositories, so pull_github_data.py samples the repo's visibility and the pricing layer zeroes those classes when the catalog vendor declares standard_runners_free_for_public_repos. Without it, an open-source project gets told it spends four figures a month on runners nobody bills it for.

Two boundaries hold that in place. Larger (sized) runners are billed on public repos at the usual rate, so only non-sized classes are zeroed. And the flag is per-vendor rather than global: a third-party runner host bills an open-source project like anyone else, which is exactly what makes the cross-vendor comparison meaningful for a public repo.

A plan's included minutes are handled the opposite way — disclosed via included_minutes_note, never subtracted. The allowance belongs to the account and is shared across every repository under it, and a one-repo audit cannot see how much of it the other repos consumed. Subtracting it here would mean attributing an account-wide allowance to a single repo, which is a guess dressed as arithmetic.

Tests

Flat under tests/, one module per feature (test_<feature>.py), with a single root conftest.py putting scripts/ on the path.

Run everything from the repo root:

python3 -m pytest tests/ -q