How the pieces fit, and where the line between deterministic computation and
LLM judgment sits. For why the tool exists and what it deliberately doesn't
do, see README.md; for the history of decisions and why they
changed, CHANGELOG.md.
No LLM is ever in the cost-computation path. Every dollar figure comes from
a Python script reading a versioned rate table. Given a pinned
catalog_version and a usage profile, the computed comparison is byte-for-byte
reproducible, and any number in a report can be traced to a catalog entry and a
measured minute count.
An LLM appears at exactly two points, both outside that path, both behind a human gate:
| Where | What it decides | The gate |
|---|---|---|
add-vendor skill |
which of a vendor's own machine classes corresponds to which ci-cost runner class | a reviewed pull request |
ci-cost skill |
which computed facts matter for this repo, and how to say them | the numbers are already fixed; only the prose is generated |
Neither can produce a rate. The first proposes one and a person approves it; the second reads one and explains it.
GitHub REST API
│
pull_github_data.py measured
│ ────────
usage-profile.json
│
├────────────────────────────┐
│ │
compute_current_vendor_cost.py build_comparison.py
(single vendor) (every vendor)
│ │
ci-cost-report.md comparison.json facts
│ ─────
skills/ci-cost/SKILL.md
│
ci-cost-report.md prose
Both report paths exist deliberately: the scripts must stay runnable and
verifiable without an LLM present, so compute_current_vendor_cost.py is not
dead weight — it's the proof that the math works standalone.
skills/ci-cost/ the installable skill — what `npx skills add` copies
SKILL.md
scripts/ stdlib-only; runnable and tested without an LLM
references/ the versioned pricing catalog, bundled
maintainers/ci-cost/ catalog authoring; never copied into an install
SKILL.md the add-vendor skill
scripts/ scraping, draft validation, catalog PRs
references/ vendor drafts, scrape staging
tests/ flat, one test module per feature
Paths inside the skill resolve from __file__ via skill_paths.py, never from
the working directory — once installed, cwd is the user's repository, so a
relative path points into their project rather than at the bundled catalog.
Maintainer scripts put the skill's scripts/ on sys.path rather than keeping
a second copy of the catalog schema, which would drift.
| Script | Responsibility |
|---|---|
pull_github_data.py |
Samples workflow runs and jobs over the GitHub REST API, classifies each job's runner labels, extrapolates sampled minutes per class to a monthly figure. Flags low_confidence when the observed history is too short to trust. |
load_pricing_catalog.py |
Loads and validates skills/ci-cost/references/pricing-catalog.json. Archetype-dependent required fields, finite non-negative rates, duplicate-JSON-key rejection, and the equivalence-map agreement rules. |
replay_cross_vendor_cost.py |
One pure function per pricing archetype over (usage_profile, vendor_entry). Partitions usage into priced / verified-unavailable / unmapped. No network, no LLM. |
build_comparison.py |
Prices a usage profile against every catalog vendor, emits JSON facts. Takes no position on ranking. |
compute_current_vendor_cost.py |
Single-vendor cost plus a rendered markdown report. The standalone, LLM-free path. |
scrape_pricing_data.py |
Maintainer-only. Scrapes vendor pricing pages via Firecrawl into gitignored staging for manual verification. Extracts no rates itself. |
validate_vendor_draft.py |
Gates a vendor draft before it can become a PR. Validates proposed_entry with the real catalog validator, so a passing draft cannot fail once merged. |
open_vendor_pr.py |
Merges a validated draft into the catalog on a branch and opens the PR, with the evidence in the body. |
Everything keys on the class names classify_runner_labels emits:
linux-x64-default, linux-arm64, windows-x64, windows-arm64, macos,
self-hosted, plus sized variants (linux-x64-4-core, macos-12-core, …).
Sized names are generated, not enumerated. runner_class_name builds them
by interpolating whatever core count appeared in the label, so
linux-x64-64-core is valid and no catalog can list every possibility in
advance. An equivalence map is therefore always partial by design, and the
fail-loudly path is the mechanism that tells the maintainer which entry to add.
is_known_runner_class is the single source of truth for what is a legal name.
skills/ci-cost/references/pricing-catalog.json is a frozen snapshot, never fetched live.
Determinism is a claim about the math, not the rates — which is why every
report must print catalog_version, the snapshot date, and a verify-before-
acting disclaimer.
Rate fields depend on the archetype, so a credit-priced vendor's arithmetic stays auditable in the file rather than pre-multiplied into a flat rate:
pricing_model |
Rate fields |
|---|---|
per_minute_by_runner_class |
rates: {class: $/min} |
credit_based |
credits_per_minute: {class: n} + usd_per_credit |
self_hosted_compute_estimate |
rates: {class: $/min} |
per_active_user_flat_plus_usage is deliberately not accepted: a per-seat
cost isn't a function of the sampled usage profile, so nothing could compute it,
and accepting it would let an uncomputable entry pass validation.
runner_class_equivalents records which of the vendor's own classes each runner
class was priced against — and encodes the distinction the whole comparison
rests on:
| In the map | Meaning | Effect |
|---|---|---|
"docker/medium" |
mapped and priced | charged |
null |
verified: this vendor has no equivalent | excluded from the total, disclosed in the report |
| absent | nobody has mapped it yet | hard error |
Collapsing the last two would let a forgotten mapping silently drop minutes from a vendor's total and make it look cheaper than it is. That is the single most dangerous failure this tool can have, because nothing downstream can detect it.
A vendor that cannot run part of the workload has a smaller total, so a naive ranking puts it first and it reads as the bargain. Starsling has no macOS runners: for an iOS repo it scores $20.00/mo against GitHub Actions' $49.84 while being unable to build the app.
build_comparison.py therefore emits is_like_for_like, the excluded classes
with their minutes, and excluded_cost_on_current_vendor — the excluded
minutes priced at the current vendor's rates, because "excludes 320 macOS
minutes" only becomes legible as "$19.84/month". Vendors are ordered
cheapest-first for stable output, and the payload carries an ordering_note
stating that this is not a ranking.
A vendor's per-minute rates are only part of what it charges, and the missing part fails in the same direction as partial coverage: the total comes out smaller than the vendor really is, so it sorts first and reads as the bargain.
base_monthly_cost is therefore a required field on every catalog entry,
with the same three-state discipline as runner_class_equivalents:
| Value | Means | Effect on the total |
|---|---|---|
0 |
verified: no fixed charge before usage | total is the usage leg |
<n> |
verified: this much every month | added to the usage leg |
null |
a fee exists that cannot be stated as one number | total is a floor; has_unmodeled_fees is true |
| absent | — | hard error |
RunsOn is why this exists. Its per-minute rates are an AWS spot pass-through
with no markup, which makes it the cheapest row in most tables, and it
separately charges a tiered annual license fee denominated in euros. Neither a
snapshotted FX rate nor a guess at the volume tier belongs in a catalog whose
claim is reproducibility, so the fee is null with a base_cost_note, and
has_unmodeled_fees blocks the vendor from being presented as cheapest.
build_comparison.py emits usage_monthly_cost and base_monthly_cost
alongside total_monthly_cost so the arithmetic stays auditable rather than
arriving as one figure to be taken on trust.
GitHub does not charge for standard hosted runners on public repositories, so
pull_github_data.py samples the repo's visibility and the pricing layer
zeroes those classes when the catalog vendor declares
standard_runners_free_for_public_repos. Without it, an open-source project
gets told it spends four figures a month on runners nobody bills it for.
Two boundaries hold that in place. Larger (sized) runners are billed on public repos at the usual rate, so only non-sized classes are zeroed. And the flag is per-vendor rather than global: a third-party runner host bills an open-source project like anyone else, which is exactly what makes the cross-vendor comparison meaningful for a public repo.
A plan's included minutes are handled the opposite way — disclosed via
included_minutes_note, never subtracted. The allowance belongs to the
account and is shared across every repository under it, and a one-repo audit
cannot see how much of it the other repos consumed. Subtracting it here would
mean attributing an account-wide allowance to a single repo, which is a guess
dressed as arithmetic.
Flat under tests/, one module per feature (test_<feature>.py), with a
single root conftest.py putting scripts/ on the path.
Run everything from the repo root:
python3 -m pytest tests/ -q