Repository navigation
Add vram argument to cluv submit to make it easier to use MIG GPUs - #220
Merged
Merged
Conversation
lebrice
added this pull request to stack #218
September 24, 2026 19:53
lebrice
force-pushed
the
add_vram_argument
branch
from
September 24, 2026 19:56
4b8a01d to
a4ce494
Compare
lebrice
force-pushed
the
add_vram_argument
branch
3 times, most recently
from
September 24, 2026 20:13
d9fd92c to
67babdb
Compare
cluv submit to use MIG GPUscluv submit to make it easy to use MIG GPUs
cluv submit to make it easy to use MIG GPUscluv submit to make it easier to use MIG GPUs
lebrice
force-pushed
the
add_vram_argument
branch
from
September 24, 2026 20:32
67babdb to
82e1d02
Compare
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## master #220 +/- ##
===========================================
+ Coverage 72.58% 86.84% +14.25%
===========================================
Files 22 23 +1
Lines 2444 2592 +148
===========================================
+ Hits 1774 2251 +477
+ Misses 670 341 -329 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
lebrice
force-pushed
the
add_vram_argument
branch
from
September 25, 2026 13:51
82e1d02 to
de7aa97
Compare
lebrice
force-pushed
the
add_vram_argument
branch
3 times, most recently
from
September 25, 2026 21:21
abee7bd to
14d22a0
Compare
lebrice
marked this pull request as ready for review
September 25, 2026 21:27
lebrice
force-pushed
the
add_vram_argument
branch
from
September 28, 2026 17:00
14d22a0 to
319c544
Compare
hvdbm
approved these changes
Sep 28, 2026
Signed-off-by: Fabrice Normandin <normandf@mila.quebec>
Co-authored-by: hvdbm <hugo.vandenbroucke-menu@mila.quebec>
lebrice
force-pushed
the
add_vram_argument
branch
from
September 29, 2026 17:29
28e2487 to
0c722a9
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #62.
What this does
cluv submit rorqual scripts/job.sh --gpus=1 --vram=10GBsubmits one job per GPU type of the cluster that has at least 10GB of VRAM and keeps the first one that starts, cancelling the others (reusing the
submit_and_keep_firstmachinery from #161).On Rorqual that's a race between the
1g.10gb,2g.20gband3g.40gbMIG slices and a full H100. MIG slices are never allocated unless you ask for them explicitly, and they are mostly idle, so this makes single-GPU jobs start much sooner.--gpus=h100:1), only that model and its MIG slices are raced, as described in the issue.--gres=gpu:...stays--gres), and when the request comes from the job script header, the flag is passed on the command line, where it overrides the#SBATCHdirective.Where the GPU types come from
Nothing is hard-coded: the GPU types are read from the cluster with
sinfo --noheader --format='%f|%G' | sort -uand cached (per cluster, for a week) so this doesn't add an SSH round-trip to every submission.
The VRAM of each type is resolved from, in order:
nvidia_h100_80gb_hbm3_1g.10gb→ 10GB,a100_4g.20gb→ 20GB),nvidia_h100_80gb_hbm3→ 80GB),x86_64,volta,nvlink,dgx,32gb→ 32GB,h200 … 150gb→ 150GB),h100on Fir/Rorqual/Killarney).Testing
Unit tests —
tests/test_vram.py(62 tests) uses realsinfooutput captured on Rorqual, Narval, Mila and Tamia.Integration tests — two new ones in
tests/test_integration.py, both lightweight (no job is ever submitted):test_get_gpu_types: lists the GPU types of the cluster over SSH and asserts we know how much VRAM each one has. This is the canary for a cluster getting a GPU model our VRAM logic can't resolve. It already earned its keep: it found that Nibi hasa5000andmi300a, which are now in the fallback table.test_vram_sbatch_args_are_valid: runs the flags that--vramgenerates throughsbatch --test-onlyon the cluster (batched into a single SSH command) and asserts the GPU request isn't malformed. This is what caught-G=h100:1being rejected with "Invalid Trackable RESource" — short options need their value as a separate argument.It deliberately doesn't assert the job would schedule: Tamia only allocates whole GPU nodes and Mila's H100s live in their own partition, so a valid request can still be refused for reasons that have nothing to do with the flag.
Both skip on clusters with no GPUs (Trillium). Locally they pass against the 10 clusters I have connections to (Mila, Tamia, Rorqual, Fir, Narval, Nibi, Killarney, Vulcan, …).
On Rorqual,
--gpus=1 --vram=10GBgives:i.e. the MIG slices would start ~5 minutes before the full H100.
Notes / follow-ups
--vramisn't exposed through the Hydra launcher yet; that would be a small follow-up.shard:l40s:16(GPU sharding) alongsidegpu:l40s:4. Shards are ignored for now.🤖 Generated with Claude Code