[hardware wanted] Strix Halo / Ryzen AI MAX+: CPU + iGPU + NPU in Colibri — Windows and Linux notes #1590
Replies: 4 comments 3 replies
|
First numbers from the Linux side. Machine: GMKtec EVO-X2, Ryzen AI Max+ 395, 128 GB, Nobara 44, kernel 7.2 with in-tree All numbers below are from the firmware's Balanced mode (85 W socket limit). A 120 W round follows. How power is measured, and the limits of itOn Linux the NPU has no sensor of its own. Q2: Linux portability of the #1261 laneI ported the loader to Linux: Q1a: the shared-expert gate/up pair (K=6144, N=2048, fmt=4 gs64), CPU vs iGPUThe bench is steady state with warm-up excluded, and it uses the engine's own
At M=1 the CPU wins. From M=16 up the iGPU is 17–33 % cheaper per operation. But the iGPU only reaches ~380 GFLOPS at M≥16, a small fraction of what gfx1151 can do. The reason is structural: Q1b: whole model, Gemma 4 E4B, NPU (FastFlowLM) vs iGPU (llama.cpp Vulkan)512 tokens, three runs each, run-to-run spread under 3 %:
The NPU does not save energy per token. Above idle the two cost the same. The iGPU draws three times the power and finishes 4.7x sooner. Including the rest of the machine, which keeps drawing power for the whole run, the iGPU is a third cheaper. What the NPU does have is low power: 22 W, quiet, and it leaves the iGPU free. Caveat: these are different runtimes and probably different quantisations of the same weights. Memory bandwidth (spec: 256 GB/s, shared)
So on this part a bandwidth-bound decode belongs on the iGPU, and the CPU cannot use the whole memory system even when it is idle otherwise. A CPU+iGPU concurrent figure is still to come; my first attempt reported best-of rather than same-window numbers and is discarded. Next
|
|
A follow-up on the Vulkan numbers above. It is a proposal rather than a PR, because I can't measure the gain on my box right now. What limits the iGPU at M>1. Proposal. Add a tiled path for S>1 alongside the current GEMV, which stays as it is for decode. Each workgroup would dequantise a weight tile (fmt 4 grouped int4, and presumably fmt 1/2/5) into shared memory once, then apply it to a tile of token rows through cooperative-matrix MMA or packed int8 dot products. llama.cpp's What I'd expect it to buy (estimate, not measured).
It also matters for the NPU question: a placement comparison is only fair against a GPU kernel that actually uses the matrix hardware. I'm not planning to write it myself right now; GLM is the only family on this path and doesn't fit this box usefully. I'm happy to benchmark and energy-measure a candidate on gfx1151 if someone picks it up. |
|
Sorry for the slow reply. I held the artifacts back until #1261 landed, so what you verify against is upstream rather than a PR branch. It merged into The four qualified files are attached ( They're whole-array BF16 GEMM kernels built for XDNA2 against the Windows XRT toolchain, so whether they load under XRT 2.26 on Linux is exactly the open question. For comparison, on Windows On the port: now that the lane is in On your measurements: thank you, and for being explicit about what the socket figure can and can't tell us. Three things stood out:
Looking forward to the 120 W series and the concurrent bandwidth run. The tiled Vulkan proposal deserves JustVugg's eyes, and it's buried in a hardware thread here. I checked your reading against the code: the dispatch really is one set of workgroups per token row, and there's no cooperative-matrix path anywhere. A 35–40% shorter time to first token, with decode unchanged, is a strong case, and your offer to benchmark and energy-measure a candidate on gfx1151 is exactly what someone picking it up would need. It might be worth its own Ideas discussion, since it affects every AMD/Intel Vulkan user, not only this box. |
|
Thank you both, this thread is exactly the kind of evidence that moves decisions. On the NPU, to be clear about where I stand: when #959 asked, the answer was no, because what was on the table did not fit how colibri works. That was about that proposal, not about the hardware. If someone can make the NPU do useful work inside colibri, measured and fail-closed the way #1261 is, the answer is yes. #1261 is in @huppiflupp, the Linux port: yes, as its own PR against
The attached zip checks out: I hashed the four files and they match On the numbers: energy per token above idle is the same for the NPU and the iGPU (1.08 and 1.10 J), which settles the question at the top of the thread and is consistent with #1261 claiming neither speed nor energy. The NPU column in the shared-expert table (M = 1, 16, 64, 256) is the head-to-head I'd most like to see next, along with the 120 W series. The bandwidth split, 108 GB/s for the CPU against 226 GB/s for the iGPU, is the number that matters most on this part: a bandwidth-bound decode wants the iGPU. The tiled Vulkan path for S>1: I checked your reading too. |
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
A thread for anyone running Colibri on AMD Strix Halo (Ryzen AI MAX+ 395, Radeon 8060S / gfx1151, XDNA2 NPU), to compare notes on using the three processors together. This started as a side conversation with @huppiflupp on #1511, and moved here so it doesn't clutter that PR and so others with the same hardware can join in.
Prior context. NPU support was asked about before in #959, and the answer then was a clear no. XDNA is a graph-compiler target while Colibri decides at runtime which experts to move, and on this hardware the bottleneck is disk and memory bandwidth, not compute. Both points shaped #1261. Its kernels are compiled ahead of time for fixed shapes, but weights and activations are passed at runtime, and it only targets the shared expert, which is dense and runs for every token, so nothing depends on routing. The bandwidth point still stands: that is why #1261 makes no speed claim, and why power, not throughput, is the open question below.
Where things stand
iGPU (HIP, Windows). Native Windows HIP support for gfx1151 landed in #788.
NPU (XDNA2, Windows). #1261 (under review) adds an optional, explicitly requested lane for one operation family, the GLM shared expert's gate/up projections. It is deliberately small: explicit opt-in, qualified shapes only, and a fail-closed fallback to exactly what the normal path computes. It makes no speed claim. Raw operations on the NPU beat the CPU path on real GLM shapes, but at model level weight conversion and device setup eat most of that, and end to end it's roughly a wash so far.
NPU vs iGPU, whole-model decode. @huppiflupp measured the iGPU at about 4.5x the NPU's decode throughput on the same box (numbers and caveats here): Gemma4 via Lemonade rather than Colibri, and power not measured.
Linux. @huppiflupp has the NPU working under Linux with XRT 2.26.0 and the
amdxdnadriver, and published the setup: https://github.com/huppiflupp/strix-halo-npu-linuxThe direction
The aim isn't NPU-only or GPU-only, but CPU + iGPU + NPU together, with each operation placed where it pays. Locally I've measured that one process can hold live XRT/XDNA2 and HIP/Radeon contexts at the same time, in either init order, and that HIP and XDNA work genuinely overlaps rather than serialising. Neither of those is a result about model speed yet. They only show the combination is mechanically possible.
Open questions
.xclbinartifacts (whole-array BF16 GEMM kernels built for XDNA2 on Windows) load and run under XRT on Linux. Nobody has tried yet.If you have this hardware
Numbers are more useful with context: OS, driver and runtime versions, which backend, model and quantisation, what exactly was timed, and power if you can get it.
All reactions