A thread for me to throw some benchmarks of the performance I’m getting from this card. I purchased it to experiment with LLMs and SRIOV. I’ve seen a few questions about this in other places. I’ll put this here so that its separate from the discussions about the BattleMatrix setup.
Using llama.cpp build 8175 SYCL I get the following numbers
Qwen3.5 9B Q4 K XL
llama-cpp-sycl
–bench -m /models/Qwen3.5-9B-UD-Q4_K_XL.gguf -ngl 100 -r 3 -p 2048,16384 -o md
load_backend: loaded SYCL backend from /app/libggml-sycl.so
load_backend: loaded CPU backend from /app/libggml-cpu-alderlake.so
model
size
params
backend
ngl
test
t/s
qwen35 ?B Q4_K - Medium
5.55 GiB
8.95 B
SYCL
100
pp2048
328.71 ± 11.00
qwen35 ?B Q4_K - Medium
5.55 GiB
8.95 B
SYCL
100
pp16384
317.41 ± 1.76
qwen35 ?B Q4_K - Medium
5.55 GiB
8.95 B
SYCL
100
tg128
23.01 ± 0.03
Qwen3.5 35B-A3B Q4 K XL
llama-cpp-sycl
–bench -m /models/Qwen3.5-35B-A3B-UD-Q4_K_XL.gguf -ngl 100 -r 3 -p 2048,16384 -o md
load_backend: loaded SYCL backend from /app/libggml-sycl.so
load_backend: loaded CPU backend from /app/libggml-cpu-alderlake.so
model
size
params
backend
ngl
test
t/s
qwen35moe ?B Q8_0
19.16 GiB
34.66 B
SYCL
100
pp2048
112.62 ± 0.93
qwen35moe ?B Q8_0
19.16 GiB
34.66 B
SYCL
100
pp16384
111.61 ± 1.11
qwen35moe ?B Q8_0
19.16 GiB
34.66 B
SYCL
100
tg128
8.34 ± 0.22
For reference of performance here is Qwen3.5 35B-A3B running on my RTX Pro 4500 (obv more expensive card but both nominally 200W GPUs)
local/llama.cpp:full-cuda
–bench -m /models/Qwen3.5-35B-A3B-UD-Q4_K_XL.gguf -r 3 -p 2048,16384 -o md
ggml_cuda_init: found 1 CUDA devices:
Device 0: NVIDIA RTX PRO 4500 Blackwell, compute capability 12.0, VMM: yes
load_backend: loaded CUDA backend from /app/libggml-cuda.so
load_backend: loaded CPU backend from /app/libggml-cpu-alderlake.so
model
size
params
backend
ngl
test
t/s
qwen35moe ?B Q8_0
19.16 GiB
34.66 B
CUDA
99
pp2048
3807.62 ± 143.98
qwen35moe ?B Q8_0
19.16 GiB
34.66 B
CUDA
99
pp16384
3185.18 ± 15.34
qwen35moe ?B Q8_0
19.16 GiB
34.66 B
CUDA
99
tg128
133.47 ± 0.33
(I’m doing intentionally default config here for easy comparisons, I can tweak the batch/microbatch configs and get more t/s out of the nvidia card, but it wouldn’t be a reasonable comparison)
My experience is that the B60 is ok for a chat length context, on smaller models. A card like this with opencode/claude code ends up with you waiting 5 or 10mins for it to start responding to the initial prompt.
I am happy to run some other llama-bench on the card if folks have something of interest.
So it looks like vulkan backend for llama.cpp runs prefill faster but token generation slower relative to SYCL backend for Qwen3.5-9B Q4 K XL. I’d really love to be able to run this model faster than this so that the 2nd GPU has some other purpose than ‘just’ 1 or 2 SRIOV VMs. This vulkan backend is getting closer to what I’d like to see. If token generation was 25+ t/s I’d be a little happier. Worth exploring some more though!
Intel also merged their OpenVINO backend just a few days ago, which in theory should provide the best experience for Intel integrated graphics as far back as 6th gen and up to the latests Intel dGPUs and iGPUs. I haven’t been able to get it to work on my 9900K (not for lack of trying!)
But I’d love to see some numberes on these. Build instructions can be found here.
Would be interesting to see if performance could be improved by processing the context on the iGPU on a model that only barely fits in VRAM. (265k iGPU+B60)
Would be interesting to see if performance could be improved by processing the context on the iGPU on a model that only barely fits in VRAM. (265k iGPU+B60)
I saw a reply in the B70 thread about Qn_0 quants potentially being faster on intel GPU.
Gave it a go just now and it seems to have some merit, mostly on token generation:
model
size
params
backend
ngl
test
t/s
qwen35 ?B Q4_K - Medium
2.54 GiB
4.21 B
Vulkan
100
pp2048
821.08 ± 0.28
qwen35 ?B Q4_K - Medium
2.54 GiB
4.21 B
Vulkan
100
tg512
30.87 ± 0.03
qwen35 ?B Q4_K - Medium
2.54 GiB
4.21 B
Vulkan
100
pp2048+tg512
132.27 ± 0.70
model
size
params
backend
ngl
test
t/s
qwen35 ?B Q4_0
2.40 GiB
4.21 B
Vulkan
100
pp2048
812.10 ± 0.80
qwen35 ?B Q4_0
2.40 GiB
4.21 B
Vulkan
100
tg512
42.89 ± 0.07
qwen35 ?B Q4_0
2.40 GiB
4.21 B
Vulkan
100
pp2048+tg512
174.50 ± 0.43
llama.cpp build: d903f30e2 (8175) with vulkan backend in docker on B60
I now have my opencode setup running with qwen3.5-27B running on RTX Pro 4500 and with the small model that opencode uses for session name generation and other small tasks using the qwen3.5 4B model on the B60 while the B60 also runs SRIOV for a windows VM.
They tend to run significantly faster on AMD GPUs, too. The trouble is, those quants are legacy for a reason - the quality is quite a bit lower than _K and _K_M variants etc.
The latest 8685+ llama.cpp has some nice SYCL improvements. Q8_0 support was added by - PR#21527
llama-cpp-sycl
–bench -m /models/qwen3.5_9B/Qwen3.5-9B-UD-Q4_K_XL.gguf -ngl 100 -r 3 -o md -p 2048,16384
load_backend: loaded SYCL backend from /app/libggml-sycl.so
load_backend: loaded CPU backend from /app/libggml-cpu-alderlake.so
model
size
params
backend
ngl
test
t/s
qwen35 9B Q4_K - Medium
5.55 GiB
8.95 B
SYCL
100
pp2048
1633.72 ± 0.59
qwen35 9B Q4_K - Medium
5.55 GiB
8.95 B
SYCL
100
pp16384
1464.48 ± 0.36
qwen35 9B Q4_K - Medium
5.55 GiB
8.95 B
SYCL
100
tg128
33.41 ± 0.00
build: 71a81f6fc (8688)
On my B60 Pro with qwen3.5 9B Q4_K_XL I’m not getting the full 3x TG speedup but definitely prefill speeds are up a long way which is very handy for interactivity with long system prompts like when using claude code or opencode etc
Very cool to see these optimizations still coming in!
Stress tested things by running a benchmark in windows via sriov and llama.cpp in Linux at the same time. This card is only on pcie 4.0 4x slot so things aren’t exactly smooth sailing when loading llm and running inference at the same time.
I bumped up the vram for QXL output and things are quite usable for light desktop apps in the VM at 4k. Pretty happy with the setup with the B60 running VMs and smaller LLMs to offload things.