Skip to content

Add vram argument to cluv submit to make it easier to use MIG GPUs - #220

Merged
lebrice merged 2 commits into
masterfrom
add_vram_argument
Sep 29, 2026
Merged

lebrice merged 2 commits into
masterfrom
add_vram_argument

Conversation

@lebrice

@lebrice lebrice commented Sep 24, 2026

Copy link
Copy Markdown
Contributor

Closes #62.

What this does

cluv submit rorqual scripts/job.sh --gpus=1 --vram=10GB

submits one job per GPU type of the cluster that has at least 10GB of VRAM and keeps the first one that starts, cancelling the others (reusing the submit_and_keep_first machinery from #161).

On Rorqual that's a race between the 1g.10gb, 2g.20gb and 3g.40gb MIG slices and a full H100. MIG slices are never allocated unless you ask for them explicitly, and they are mostly idle, so this makes single-GPU jobs start much sooner.

  • When a GPU model is already requested (--gpus=h100:1), only that model and its MIG slices are raced, as described in the issue.
  • Jobs asking for more than one GPU are left alone (MIG slices can only be used one at a time), as are clusters where no GPU type is big enough. Both cases print a warning and submit the job as it was.
  • The flag form of the existing GPU request is kept (--gres=gpu:... stays --gres), and when the request comes from the job script header, the flag is passed on the command line, where it overrides the #SBATCH directive.

Where the GPU types come from

Nothing is hard-coded: the GPU types are read from the cluster with

sinfo --noheader --format='%f|%G' | sort -u

and cached (per cluster, for a week) so this doesn't add an SSH round-trip to every submission.

The VRAM of each type is resolved from, in order:

  1. the MIG profile in the GRES name (nvidia_h100_80gb_hbm3_1g.10gb → 10GB, a100_4g.20gb → 20GB),
  2. the memory in the GRES name (nvidia_h100_80gb_hbm3 → 80GB),
  3. the node features, which is where Mila and Tamia put it (x86_64,volta,nvlink,dgx,32gb → 32GB, h200 … 150gb → 150GB),
  4. a small fallback table, for the names that say nothing (h100 on Fir/Rorqual/Killarney).

Testing

Unit tests — tests/test_vram.py (62 tests) uses real sinfo output captured on Rorqual, Narval, Mila and Tamia.

Integration tests — two new ones in tests/test_integration.py, both lightweight (no job is ever submitted):

  • test_get_gpu_types: lists the GPU types of the cluster over SSH and asserts we know how much VRAM each one has. This is the canary for a cluster getting a GPU model our VRAM logic can't resolve. It already earned its keep: it found that Nibi has a5000 and mi300a, which are now in the fallback table.

  • test_vram_sbatch_args_are_valid: runs the flags that --vram generates through sbatch --test-only on the cluster (batched into a single SSH command) and asserts the GPU request isn't malformed. This is what caught -G=h100:1 being rejected with "Invalid Trackable RESource" — short options need their value as a separate argument.

    It deliberately doesn't assert the job would schedule: Tamia only allocates whole GPU nodes and Mila's H100s live in their own partition, so a valid request can still be refused for reasons that have nothing to do with the flag.

Both skip on clusters with no GPUs (Trillium). Locally they pass against the 10 clusters I have connections to (Mila, Tamia, Rorqual, Fir, Narval, Nibi, Killarney, Vulcan, …).

On Rorqual, --gpus=1 --vram=10GB gives:

--gpus=nvidia_h100_80gb_hbm3_1g.10gb:1  -> Job to start at 18:49:03 on rg12502
--gpus=nvidia_h100_80gb_hbm3_2g.20gb:1  -> Job to start at 18:49:04 on rg12501
--gpus=nvidia_h100_80gb_hbm3_3g.40gb:1  -> Job to start at 18:49:05 on rg12502
--gpus=h100:1                           -> Job to start at 18:54:19 on rg31701

i.e. the MIG slices would start ~5 minutes before the full H100.

Notes / follow-ups

  • --vram isn't exposed through the Hydra launcher yet; that would be a small follow-up.
  • Vulcan exposes shard:l40s:16 (GPU sharding) alongside gpu:l40s:4. Shards are ignored for now.

🤖 Generated with Claude Code

@lebrice
lebrice added this pull request to stack #218 September 24, 2026 19:53
@lebrice
lebrice force-pushed the add_vram_argument branch 3 times, most recently from d9fd92c to 67babdb Compare September 24, 2026 20:13
@lebrice lebrice changed the title Add vram argument to cluv submit to use MIG GPUs Add vram argument to cluv submit to make it easy to use MIG GPUs Sep 24, 2026
@lebrice lebrice changed the title Add vram argument to cluv submit to make it easy to use MIG GPUs Add vram argument to cluv submit to make it easier to use MIG GPUs Sep 24, 2026
@codecov-commenter

codecov-commenter commented Sep 24, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 95.20958% with 8 lines in your changes missing coverage. Please review.
✅ Project coverage is 86.84%. Comparing base (d5f553c) to head (0c722a9).

Files with missing lines Patch % Lines
cluv/cache.py 80.00% 4 Missing ⚠️
cluv/cli/submit_utils/vram.py 96.80% 4 Missing ⚠️
Additional details and impacted files
@@             Coverage Diff             @@
##           master     #220       +/-   ##
===========================================
+ Coverage   72.58%   86.84%   +14.25%     
===========================================
  Files          22       23        +1     
  Lines        2444     2592      +148     
===========================================
+ Hits         1774     2251      +477     
+ Misses        670      341      -329     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@lebrice
lebrice force-pushed the add_vram_argument branch 3 times, most recently from abee7bd to 14d22a0 Compare September 25, 2026 21:21
@lebrice
lebrice marked this pull request as ready for review September 25, 2026 21:27
Base automatically changed from refactor_submit to master September 28, 2026 17:00
Comment thread cluv/__main__.py Outdated
lebrice and others added 2 commits September 29, 2026 13:29
Signed-off-by: Fabrice Normandin <normandf@mila.quebec>
Co-authored-by: hvdbm <hugo.vandenbroucke-menu@mila.quebec>
@lebrice
lebrice merged commit d96a2dc into master Sep 29, 2026
23 of 24 checks passed
@lebrice
lebrice deleted the add_vram_argument branch September 29, 2026 20:58
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[FEAT]: Make it easier to use MIG GPUs with a --vram argument to cluv submit

3 participants