Repository navigation
llamacpp:sycl — first-class Intel GPU backend #3523
TravisJ33304
started this conversation in
Request for Comment (RFC)
Replies: 1 comment 1 reply
|
@TravisJ33304 I converted your issue to an RFC (see new policy landing soon: #3521 ) Your description above could use a note about maintainership: how do we ensure @kenvandine please review this RFC. Reference implementation available here: #3509 |
1 reply
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Summary
Lemonade documents Intel GPUs as Vulkan, but
/system-infonever detectsintel_gpu, so llama.cpp never auto-selects a GPU on Arc. We want a first-classllamacpp:syclvariant (same shape as CUDA/ROCm/Vulkan).Why SYCL, not only Vulkan
On Intel Arc Pro B70, llama.cpp SYCL is the validated high-performance path.
Vulkan on Arc has quality issues (garbled output unless
GGML_VK_DISABLE_F16).Vulkan can remain the no-oneAPI fallback.
Scope
intel_gpu(Linux: PCI0x8086,xe/i915; prefer discrete over iGPU).syclsupport row +llamacpp.sycl_binBYO.ggml-org
llama-*-sycl*asset exists; otherwise requiresycl_bin.--device SYCL0, oneAPI/Level Zero env.get_global_vram_usage_pct()for Intel soauto_evictworks.server_llm.py --backend sycl(no GitHub-hosted SYCL CI).Hardware
We can test on Intel Arc Pro B70 (Linux,
xedriver).Maintenance
All reactions