Repository navigation
Properly leverage pre-commit cache #332
Description
Activity
- added a parent issue
on Sep 9, 2026 is keyed on the .pre-commit-config.yaml contents and is therefore updated only when we update that file in a PR.
The official
pre-commitGitHub action already does this: https://github.com/pre-commit/actionIt's so small that it fits inline here, and we could probably inline something similar in
shared-workflowsif we wanted to avoid a dependency on it:name: pre-commit description: run pre-commit inputs: extra_args: description: options to pass to pre-commit run required: false default: '--all-files' runs: using: composite steps: - run: python -m pip install pre-commit shell: bash - run: python -m pip freeze --local shell: bash - uses: actions/cache@v4 with: path: ~/.cache/pre-commit key: pre-commit-3|${{ env.pythonLocation }}|${{ hashFiles('.pre-commit-config.yaml') }} - run: pre-commit run --show-diff-on-failure --color=always ${{ inputs.extra_args }} shell: bash
ref: https://github.com/pre-commit/action/blob/1b06ec171f2f6faa71ed760c4042bd969e4f8b43/action.yml
For any repo where
pre-commit run --all-filesis self-contained (doesn't actually need to be in the conda env we set up by convention incheck_style.sh), switching would be a quick win.But that'd mean unwinding some of the conventions around how
check_style.shworks and figuring out where to stick that action (inchecksworkflow, probably?, but idk about new job vs. new step).I'm not exactly sure what you're proposing, but I implemented this "manually" in cudf in NVIDIA/cudf#24093 by just running the checks.yaml workflow in build.yaml. Does that track with what you'd expect?
Here's an example of a recent PR (NVIDIA/cudf#24101) where the first run of the checks.yaml pre-commit job hits the cache and doesn't need to re-setup the pre-commit environments: https://github.com/NVIDIA/cudf/actions/runs/34427112171/job/102714611203?pr=24101
For comparison, here are two random PRs along with their very first runs of the style check job:
- Fix Parquet writer partition size computation for lists and structs NVIDIA/cudf#23907: https://github.com/NVIDIA/cudf/actions/runs/33453269377/job/99687733439
- Avoid serializing I/O requests from Parquet reader NVIDIA/cudf#23823: https://github.com/NVIDIA/cudf/actions/runs/33129053736/job/98714044187
🫠 I didn't actually know that
checks.yamlwas already saving and restoring thepre-commitcache, and apparently has been for 4 years: rapidsai/shared-workflows#20ignore my comments, they all were made assuming that that wasn't already happening.
Got it never mind then! Yes, the issue isn't that we haven't been saving, it's that we've been caching in a way that doesn't share across PRs and also dramatically bloats our cache usage.
- added 7 commits that reference this issue
on Oct 1, 2026
Currently the checks.yaml shared workflow is only run on PRs. That is problematic because the resulting cache is only shared within the PR. What we really want is a cache that is shared across PRs. Such a cache should generally be cheap to maintain, it just requires that we run pre-commit in our build.yaml so that we can have a cache tied to main that is available to all PRs and is keyed on the
.pre-commit-config.yamlcontents and is therefore updated only when we update that file in a PR. As one data point, currently the per-PR pre-commit caches are occupying >7Gb of the 10Gb per-repo GHA cache quota in cudf, so cleaning this up is is critical to unlocking usage of the cache for other purposes.PRs to seed the pre-commit cache from main builds
release/26.10tomain