Skip to content

Commit fd62a66

Browse files
os-litantclaude
andauthored
feat(pm): archive the open board before the closed history, and move the snapshot cron off the patrol's hour (#17442)
The first scheduled run of the archiver walked `state=all` oldest-first: 800 requests bought 137 closed issues and 268 closed pull requests from February and zero open cards. The records an account suspension destroys are the OPEN board, so that half was being archived last. A first walk now runs two phases in one order. The open phase asks the `/issues` listing for `state=open` (it carries pull requests too) and walks it to a short page; only when it is complete does the closed-history walk start from its own cursor. The two cursors are kept apart, so a run that runs out of budget inside the open set resumes inside it and the history cursor waits untouched. `walk_phase` and `walk.open_set` in the manifest separate "open set complete, history resuming" from "still inside the open set". An archive written before this order existed re-walks the open board first and keeps its history cursor rather than discarding it. The count check's pending predicate moves with it: its arithmetic needs only the open set, so the run that finishes the open phase has a reading even while the history is still resuming, and a later history run does not — that enumeration is from an earlier run. Every pending verdict now carries the reason it is pending. Also fixed, in the same file and required by the negative control: the nested `board.read_at` stamp escaped `materialManifest`, so a steady-state run whose whole job is to write nothing would move it and commit a manifest-only diff on every scheduled run. The cron moves from `37 1,7,13,19` — the half-state patrol's cron to the minute — to `7 2,8,14,20`. Both are scheduled board readers spending one per-repository `GITHUB_TOKEN` hour of 1,000 requests. The header now names that pool, the other standing spenders, and why the cap is 800. Claude-Session: https://claude.ai/code/session_01YKEjmbYNvYWJvWGSWx26zK Co-authored-by: Claude <noreply@anthropic.com>
1 parent 4280055 commit fd62a66

2 files changed

Lines changed: 421 additions & 104 deletions

File tree

‎.github/workflows/board-snapshot.yml‎

Lines changed: 68 additions & 16 deletions
Original file line numberDiff line numberDiff line change
@@ -28,16 +28,46 @@ name: Board Snapshot
2828
# self-test cases read its own source and hold that structurally — and
2929
# `--restore` PRINTS a recreate payload for a seat to post, or for nobody to.
3030
#
31-
# ## Budget, and why a run may stop before it is finished
31+
# ## The budget this run is capped against, and who else spends it
3232
#
33-
# `GITHUB_TOKEN` is limited to 1,000 requests per hour PER REPOSITORY, and this
34-
# job is not that budget's only caller (the half-state patrol and the closed-card
35-
# sweep share it). A first full snapshot of this board — several thousand
36-
# numbers, each with a comment thread — does not fit one run and must not try, so
37-
# the script stops at its own `--max-requests` ceiling, writes a resume cursor
38-
# into the manifest and exits 0. The next scheduled run continues from that
39-
# cursor. Four runs a day walk the backlog in a few days without ever exceeding
40-
# the budget, and a steady-state incremental run costs a few hundred requests.
33+
# `GITHUB_TOKEN` is limited to 1,000 requests per hour PER REPOSITORY. That is
34+
# ONE pool for every workflow in this repo that calls the API — not a per-workflow
35+
# allowance — and this job is far from its only caller:
36+
#
37+
# this archiver 800 per run (the `--max-requests` cap below), 4/day
38+
# half-state patrol a board sweep, 4/day at :37 on hours 1,7,13,19
39+
# required-set patrol 2/day at :23 on hours 4,16
40+
# release-coverage patrol daily at 04:19; platform-checklist watchdog at 02:51
41+
# merged-branch reaper Mondays 04:37, an API sweep over branches and PRs
42+
# every CI, lint and smoke run
43+
# a handful of requests each, hourly and per push
44+
#
45+
# 800 is chosen against that pool rather than against this job's appetite: it is
46+
# the largest cap that still leaves a fifth of the hour to whatever else lands in
47+
# the same window, and the schedule below keeps the two heavy board readers out
48+
# of each other's hours entirely. ⛔ Raising the cap is not a tuning decision,
49+
# ⛔ there is no retry loop, and ⛔ there is no second token — a run that wants
50+
# more requests waits for the next slot, which is what the resume cursor is for.
51+
#
52+
# ## Why a run may stop before it is finished — and what it reads FIRST
53+
#
54+
# A first full snapshot of this board — several thousand numbers, each with a
55+
# comment thread — does not fit one run and must not try, so the script stops at
56+
# its own `--max-requests` ceiling, writes its phase cursor into the manifest and
57+
# exits 0. The next scheduled run continues from that cursor. Four runs a day
58+
# walk the backlog in a few days without ever exceeding the budget, and a
59+
# steady-state incremental run costs a few hundred requests.
60+
#
61+
# WHICH records it reads first is the part that matters, and it is measured
62+
# rather than assumed. The first scheduled run of this workflow walked
63+
# `state=all` oldest-first and spent all 800 requests on 137 closed issues and
64+
# 268 closed pull requests from February — ZERO open cards, with days of runs
65+
# still to go. The open board is exactly what a suspension destroys (#17374 F3),
66+
# so it was the half being archived last. A first walk now runs the OPEN phase to
67+
# completion — issues and pull requests, one listing — before the closed history
68+
# starts, and a run that runs out of budget inside the open set resumes inside
69+
# it. The manifest's `walk_phase` says which phase a run is in, so "open set
70+
# complete, history resuming" is never read as "still inside the open set".
4171
#
4272
# ⛔ On a real rate-limit refusal the script does NOT retry: it stops, writes the
4373
# cursor and exits non-zero with the reset time. A loop against a spent budget
@@ -85,13 +115,35 @@ name: Board Snapshot
85115

86116
on:
87117
schedule:
88-
# The half-state patrol's cadence, deliberately: four times a day, six hours
89-
# apart, at :37 past the hour — offset from the top of the hour where the
90-
# hourly triage Routine runs, so the two do not contend for the same minute
91-
# of the shared request budget. Six-hourly is the loss window this accepts:
92-
# a card created and destroyed inside one interval was never archived, and
93-
# nothing cheaper than a webhook closes that, which is a different card.
94-
- cron: '37 1,7,13,19 * * *'
118+
# Four times a day, six hours apart, at :07 past the hour. The HOURS are
119+
# chosen against the other standing spenders listed above, not for tidiness.
120+
#
121+
# This job first shipped at `37 1,7,13,19` — the half-state patrol's cron to
122+
# the minute. Both are scheduled board readers, both authenticate as this
123+
# repository's `GITHUB_TOKEN`, and that budget is per repository per hour, so
124+
# a snapshot run spending its full 800 leaves the patrol the remainder of the
125+
# same window. Measured on the day this moved: the 13:37Z patrol run and the
126+
# 13:41Z snapshot run were BOTH green, so nothing had been starved yet — the
127+
# offset is PREVENTION, taken while the archive was still small enough that
128+
# no run had yet spent its whole cap against the patrol's window.
129+
#
130+
# Of the five hour sets that share no hour with the patrol's 1,7,13,19, this
131+
# is the least contended on this repository's cron inventory: 08, 14 and 20
132+
# UTC carry no other scheduled workflow at all, and 02 UTC carries only the
133+
# platform-checklist watchdog at :51 — 44 minutes after this run starts —
134+
# plus CodeQL on Mondays. `4,10,16,22` was the obvious alternative and is
135+
# worse on the same measurement: it drops an 800-request run into the busiest
136+
# scheduled hour on this board (04:00 rerun-safety, 04:19 release-coverage
137+
# patrol, 04:23 required-set patrol, 04:37 branch reaper on Mondays, 04:41
138+
# create-smoke) and shares hour 16 with the required-set patrol.
139+
#
140+
# The minute is off the top of the hour on purpose, as in every patrol here:
141+
# scheduled workflows queue behind everyone else's :00 cron.
142+
#
143+
# Six-hourly is the loss window this accepts: a card created and destroyed
144+
# inside one interval was never archived, and nothing cheaper than a webhook
145+
# closes that, which is a different card.
146+
- cron: '7 2,8,14,20 * * *'
95147
workflow_dispatch: {}
96148
# Changes to the archiver itself get exercised before they merge. The paths
97149
# name every file the run actually loads — the archiver, the module it imports

0 commit comments

Comments
 (0)