Summary
Lemonade documents Intel GPUs as Vulkan, but /system-info never detects
intel_gpu, so llama.cpp never auto-selects a GPU on Arc. We want a first-class
llamacpp:sycl variant (same shape as CUDA/ROCm/Vulkan).
Why SYCL, not only Vulkan
On Intel Arc Pro B70, llama.cpp SYCL is the validated high-performance path.
Vulkan on Arc has quality issues (garbled output unless GGML_VK_DISABLE_F16).
Vulkan can remain the no-oneAPI fallback.
Scope
- Detect
intel_gpu (Linux: PCI 0x8086, xe/i915; prefer discrete over iGPU).
sycl support row + llamacpp.sycl_bin BYO.
- Linux: do not download a Vulkan zip as SYCL. Builtin install only if a real
ggml-org llama-*-sycl* asset exists; otherwise require sycl_bin.
- Launch:
--device SYCL0, oneAPI/Level Zero env.
- Extend
get_global_vram_usage_pct() for Intel so auto_evict works.
- Tests: C++ helpers + hardware
server_llm.py --backend sycl (no GitHub-hosted SYCL CI).
Out of scope
ComfyUI, whisper SYCL, Kokoro GPU, hexzai Ansible.
Hardware
We can test on Intel Arc Pro B70 (Linux, xe driver).
Summary
Lemonade documents Intel GPUs as Vulkan, but
/system-infonever detectsintel_gpu, so llama.cpp never auto-selects a GPU on Arc. We want a first-classllamacpp:syclvariant (same shape as CUDA/ROCm/Vulkan).Why SYCL, not only Vulkan
On Intel Arc Pro B70, llama.cpp SYCL is the validated high-performance path.
Vulkan on Arc has quality issues (garbled output unless
GGML_VK_DISABLE_F16).Vulkan can remain the no-oneAPI fallback.
Scope
intel_gpu(Linux: PCI0x8086,xe/i915; prefer discrete over iGPU).syclsupport row +llamacpp.sycl_binBYO.ggml-org
llama-*-sycl*asset exists; otherwise requiresycl_bin.--device SYCL0, oneAPI/Level Zero env.get_global_vram_usage_pct()for Intel soauto_evictworks.server_llm.py --backend sycl(no GitHub-hosted SYCL CI).Out of scope
ComfyUI, whisper SYCL, Kokoro GPU, hexzai Ansible.
Hardware
We can test on Intel Arc Pro B70 (Linux,
xedriver).