My setup
I run ninfer-serve (qwen3.8-27b NVFP4, ~29 GB VRAM on an RTX 5090 D) as the
local brain for an agent that also drives ComfyUI for image generation on
the same GPU. Only one of the two can hold the VRAM at a time.
The pain
Today the handoff means killing the server and restarting it: ~40–60 s cold
start each way (24 GB of weights from NVMe + re-init). vLLM's sleep mode
(weights to pinned host RAM, ~1–2 s each way) is exactly what I'd want.
Request
An on-demand unload/wake for the model weights — free the VRAM while the
process stays up, re-load on demand. KV/state can just be dropped, it is
rebuilt.
Happy to help test and provide details on the use case.
My setup
I run ninfer-serve (qwen3.8-27b NVFP4, ~29 GB VRAM on an RTX 5090 D) as the
local brain for an agent that also drives ComfyUI for image generation on
the same GPU. Only one of the two can hold the VRAM at a time.
The pain
Today the handoff means killing the server and restarting it: ~40–60 s cold
start each way (24 GB of weights from NVMe + re-init). vLLM's sleep mode
(weights to pinned host RAM, ~1–2 s each way) is exactly what I'd want.
Request
An on-demand unload/wake for the model weights — free the VRAM while the
process stays up, re-load on demand. KV/state can just be dropped, it is
rebuilt.
Happy to help test and provide details on the use case.