Skip to content

llamacpp:sycl — first-class Intel GPU backend #3508

Description

@TravisJ33304

Summary

Lemonade documents Intel GPUs as Vulkan, but /system-info never detects
intel_gpu, so llama.cpp never auto-selects a GPU on Arc. We want a first-class
llamacpp:sycl variant (same shape as CUDA/ROCm/Vulkan).

Why SYCL, not only Vulkan

On Intel Arc Pro B70, llama.cpp SYCL is the validated high-performance path.
Vulkan on Arc has quality issues (garbled output unless GGML_VK_DISABLE_F16).
Vulkan can remain the no-oneAPI fallback.

Scope

  • Detect intel_gpu (Linux: PCI 0x8086, xe/i915; prefer discrete over iGPU).
  • sycl support row + llamacpp.sycl_bin BYO.
  • Linux: do not download a Vulkan zip as SYCL. Builtin install only if a real
    ggml-org llama-*-sycl* asset exists; otherwise require sycl_bin.
  • Launch: --device SYCL0, oneAPI/Level Zero env.
  • Extend get_global_vram_usage_pct() for Intel so auto_evict works.
  • Tests: C++ helpers + hardware server_llm.py --backend sycl (no GitHub-hosted SYCL CI).

Out of scope

ComfyUI, whisper SYCL, Kokoro GPU, hexzai Ansible.

Hardware

We can test on Intel Arc Pro B70 (Linux, xe driver).

Activity

  1. added
    engine::llamacppllama.cpp backend (LlamaCppServer); GPU/CPU LLM inference (Vulkan, ROCm, Metal)
    enhancementNew feature or request
    on Sep 6, 2026
  2. locked and limited conversation to collaborators on Sep 8, 2026
  3. converted this issue into a discussion #3523 on Sep 8, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    engine::llamacppllama.cpp backend (LlamaCppServer); GPU/CPU LLM inference (Vulkan, ROCm, Metal)enhancementNew feature or requestruntime::vulkanVulkan runtime / GPU backend

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions