How to load the model successfully through multi-card in vllm? #15626
SmallHappyJerry
announced in
Q&A
Replies: 1 comment
|
The flag you want is Working examples# 2 GPUs
vllm serve meta-llama/Llama-3.1-8B-Instruct --tensor-parallel-size 2from vllm import LLM
llm = LLM("meta-llama/Llama-3.1-8B-Instruct", tensor_parallel_size=2)Rules of thumb that avoid the common failures
How to learn this (recommended path)
One more tip: if load time scales with TP size, convert to a sharded checkpoint ( Start with |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
How to load the model successfully through multi-card in vllm? How can I learn and understand this problem?Is there a recommended study blog?
All reactions