Arc Pro B60 for Local LLMs

A thread for me to throw some benchmarks of the performance I’m getting from this card. I purchased it to experiment with LLMs and SRIOV. I’ve seen a few questions about this in other places. I’ll put this here so that its separate from the discussions about the BattleMatrix setup.

System:

Using llama.cpp build 8175 SYCL I get the following numbers

Qwen3.5 9B Q4 K XL

llama-cpp-sycl
–bench -m /models/Qwen3.5-9B-UD-Q4_K_XL.gguf -ngl 100 -r 3 -p 2048,16384 -o md
load_backend: loaded SYCL backend from /app/libggml-sycl.so
load_backend: loaded CPU backend from /app/libggml-cpu-alderlake.so

model size params backend ngl test t/s
qwen35 ?B Q4_K - Medium 5.55 GiB 8.95 B SYCL 100 pp2048 328.71 ± 11.00
qwen35 ?B Q4_K - Medium 5.55 GiB 8.95 B SYCL 100 pp16384 317.41 ± 1.76
qwen35 ?B Q4_K - Medium 5.55 GiB 8.95 B SYCL 100 tg128 23.01 ± 0.03

Qwen3.5 35B-A3B Q4 K XL

llama-cpp-sycl
–bench -m /models/Qwen3.5-35B-A3B-UD-Q4_K_XL.gguf -ngl 100 -r 3 -p 2048,16384 -o md
load_backend: loaded SYCL backend from /app/libggml-sycl.so
load_backend: loaded CPU backend from /app/libggml-cpu-alderlake.so

model size params backend ngl test t/s
qwen35moe ?B Q8_0 19.16 GiB 34.66 B SYCL 100 pp2048 112.62 ± 0.93
qwen35moe ?B Q8_0 19.16 GiB 34.66 B SYCL 100 pp16384 111.61 ± 1.11
qwen35moe ?B Q8_0 19.16 GiB 34.66 B SYCL 100 tg128 8.34 ± 0.22

For reference of performance here is Qwen3.5 35B-A3B running on my RTX Pro 4500 (obv more expensive card but both nominally 200W GPUs)

local/llama.cpp:full-cuda
–bench -m /models/Qwen3.5-35B-A3B-UD-Q4_K_XL.gguf -r 3 -p 2048,16384 -o md
ggml_cuda_init: found 1 CUDA devices:
Device 0: NVIDIA RTX PRO 4500 Blackwell, compute capability 12.0, VMM: yes
load_backend: loaded CUDA backend from /app/libggml-cuda.so
load_backend: loaded CPU backend from /app/libggml-cpu-alderlake.so

model size params backend ngl test t/s
qwen35moe ?B Q8_0 19.16 GiB 34.66 B CUDA 99 pp2048 3807.62 ± 143.98
qwen35moe ?B Q8_0 19.16 GiB 34.66 B CUDA 99 pp16384 3185.18 ± 15.34
qwen35moe ?B Q8_0 19.16 GiB 34.66 B CUDA 99 tg128 133.47 ± 0.33

(I’m doing intentionally default config here for easy comparisons, I can tweak the batch/microbatch configs and get more t/s out of the nvidia card, but it wouldn’t be a reasonable comparison)

My experience is that the B60 is ok for a chat length context, on smaller models. A card like this with opencode/claude code ends up with you waiting 5 or 10mins for it to start responding to the initial prompt.

I am happy to run some other llama-bench on the card if folks have something of interest.

have you tried vulkan backend for the intel card? it seems much better for my a750 ive been testing. im a rookie though

Qwen3.5 9B Q4 K XL - Vulkan

llama-cpp-vulkan-bench
–bench -m /models/Qwen3.5-9B-UD-Q4_K_XL.gguf -ngl 100 -r 3 -p 2048,16384 -o md
ggml_vulkan: Found 1 Vulkan devices:
ggml_vulkan: 0 = Intel(R) Arc™ Pro B60 Graphics (BMG G21) (Intel open-source Mesa driver) | uma: 0 | fp16: 1 | bf16: 0 | warp size: 32 | shared memory: 131072 | int dot: 0 | matrix cores: KHR_coopmat
load_backend: loaded Vulkan backend from /app/libggml-vulkan.so
load_backend: loaded CPU backend from /app/libggml-cpu-alderlake.so

model size params backend ngl test t/s
qwen35 ?B Q4_K - Medium 5.55 GiB 8.95 B Vulkan 100 pp2048 559.62 ± 0.35
qwen35 ?B Q4_K - Medium 5.55 GiB 8.95 B Vulkan 100 pp16384 517.64 ± 0.14
qwen35 ?B Q4_K - Medium 5.55 GiB 8.95 B Vulkan 100 tg128 17.82 ± 0.00

build: d903f30e2 (8175)

So it looks like vulkan backend for llama.cpp runs prefill faster but token generation slower relative to SYCL backend for Qwen3.5-9B Q4 K XL. I’d really love to be able to run this model faster than this so that the 2nd GPU has some other purpose than ‘just’ 1 or 2 SRIOV VMs. This vulkan backend is getting closer to what I’d like to see. If token generation was 25+ t/s I’d be a little happier. Worth exploring some more though!

SYCL works now? How complicated was the setup?

I cloned the llama.cpp repo and built docker container with:

docker build -t llama-cpp-sycl --build-arg=“GGML_SYCL_F16=ON” –target full -f .devops/intel.Dockerfile .
1 Like

Intel also merged their OpenVINO backend just a few days ago, which in theory should provide the best experience for Intel integrated graphics as far back as 6th gen and up to the latests Intel dGPUs and iGPUs. I haven’t been able to get it to work on my 9900K (not for lack of trying!)

But I’d love to see some numberes on these. Build instructions can be found here.

1 Like

Would be interesting to see if performance could be improved by processing the context on the iGPU on a model that only barely fits in VRAM. (265k iGPU+B60)


I’ll get back to you on that lmao

Would be interesting to see if performance could be improved by processing the context on the iGPU on a model that only barely fits in VRAM. (265k iGPU+B60)


I’ll get back to you on that lmao

Does the 9900K iGPU support the intel xe driver?

No, but I believe OpenVino is supported on all Intel graphics which are supported by Intels OpenCL icd. I could be wrong though.

Edit: I was wrong. It seems to be restricted to more modern GPUs and their support for older CPUs is AVX2 only.

Interesting to see the openvino developments, looks like a few rough edges still, to be expected given how fresh this is!

I tried to follow the docker build instructions and it built without any errors.

I tried to run with qwen3.5 9B and it threw an error and I found this - Eval bug: OpenVino: Cant load Qwen3.5 · Issue #20562 · ggml-org/llama.cpp · GitHub

I saw a reply in the B70 thread about Qn_0 quants potentially being faster on intel GPU.

Gave it a go just now and it seems to have some merit, mostly on token generation:

model size params backend ngl test t/s
qwen35 ?B Q4_K - Medium 2.54 GiB 4.21 B Vulkan 100 pp2048 821.08 ± 0.28
qwen35 ?B Q4_K - Medium 2.54 GiB 4.21 B Vulkan 100 tg512 30.87 ± 0.03
qwen35 ?B Q4_K - Medium 2.54 GiB 4.21 B Vulkan 100 pp2048+tg512 132.27 ± 0.70
model size params backend ngl test t/s
qwen35 ?B Q4_0 2.40 GiB 4.21 B Vulkan 100 pp2048 812.10 ± 0.80
qwen35 ?B Q4_0 2.40 GiB 4.21 B Vulkan 100 tg512 42.89 ± 0.07
qwen35 ?B Q4_0 2.40 GiB 4.21 B Vulkan 100 pp2048+tg512 174.50 ± 0.43

llama.cpp build: d903f30e2 (8175) with vulkan backend in docker on B60

I now have my opencode setup running with qwen3.5-27B running on RTX Pro 4500 and with the small model that opencode uses for session name generation and other small tasks using the qwen3.5 4B model on the B60 while the B60 also runs SRIOV for a windows VM.

1 Like

They tend to run significantly faster on AMD GPUs, too. The trouble is, those quants are legacy for a reason - the quality is quite a bit lower than _K and _K_M variants etc.

The latest 8685+ llama.cpp has some nice SYCL improvements. Q8_0 support was added by - PR#21527

llama-cpp-sycl
–bench -m /models/qwen3.5_9B/Qwen3.5-9B-UD-Q4_K_XL.gguf -ngl 100 -r 3 -o md -p 2048,16384
load_backend: loaded SYCL backend from /app/libggml-sycl.so
load_backend: loaded CPU backend from /app/libggml-cpu-alderlake.so

model size params backend ngl test t/s
qwen35 9B Q4_K - Medium 5.55 GiB 8.95 B SYCL 100 pp2048 1633.72 ± 0.59
qwen35 9B Q4_K - Medium 5.55 GiB 8.95 B SYCL 100 pp16384 1464.48 ± 0.36
qwen35 9B Q4_K - Medium 5.55 GiB 8.95 B SYCL 100 tg128 33.41 ± 0.00

build: 71a81f6fc (8688)

On my B60 Pro with qwen3.5 9B Q4_K_XL I’m not getting the full 3x TG speedup but definitely prefill speeds are up a long way which is very handy for interactivity with long system prompts like when using claude code or opencode etc

Very cool to see these optimizations still coming in!

1 Like

running Qwen3.5 35B A3B again, with newer build. New build is the real deal.

llama-cpp-sycl
–bench -m /models/Qwen3.5_35B/Qwen3.5-35B-A3B-UD-Q4_K_XL.gguf -ngl 100 -r 3 -o md -p 2048
load_backend: loaded SYCL backend from /app/libggml-sycl.so
load_backend: loaded CPU backend from /app/libggml-cpu-alderlake.so

model size params backend ngl test t/s
qwen35moe 35B.A3B Q4_K - Medium 20.70 GiB 34.66 B SYCL 100 pp2048 400.93 ± 1.63
qwen35moe 35B.A3B Q4_K - Medium 20.70 GiB 34.66 B SYCL 100 tg128 37.24 ± 0.01

build: 71a81f6fc (8688)

1 Like

Stress tested things by running a benchmark in windows via sriov and llama.cpp in Linux at the same time. This card is only on pcie 4.0 4x slot so things aren’t exactly smooth sailing when loading llm and running inference at the same time.

I bumped up the vram for QXL output and things are quite usable for light desktop apps in the VM at 4k. Pretty happy with the setup with the B60 running VMs and smaller LLMs to offload things.

Here’s two PRs to track or try early