Arc Pro B60 for Local LLMs

I saw a reply in the B70 thread about Qn_0 quants potentially being faster on intel GPU.

Gave it a go just now and it seems to have some merit, mostly on token generation:

model size params backend ngl test t/s
qwen35 ?B Q4_K - Medium 2.54 GiB 4.21 B Vulkan 100 pp2048 821.08 ± 0.28
qwen35 ?B Q4_K - Medium 2.54 GiB 4.21 B Vulkan 100 tg512 30.87 ± 0.03
qwen35 ?B Q4_K - Medium 2.54 GiB 4.21 B Vulkan 100 pp2048+tg512 132.27 ± 0.70
model size params backend ngl test t/s
qwen35 ?B Q4_0 2.40 GiB 4.21 B Vulkan 100 pp2048 812.10 ± 0.80
qwen35 ?B Q4_0 2.40 GiB 4.21 B Vulkan 100 tg512 42.89 ± 0.07
qwen35 ?B Q4_0 2.40 GiB 4.21 B Vulkan 100 pp2048+tg512 174.50 ± 0.43

llama.cpp build: d903f30e2 (8175) with vulkan backend in docker on B60

I now have my opencode setup running with qwen3.5-27B running on RTX Pro 4500 and with the small model that opencode uses for session name generation and other small tasks using the qwen3.5 4B model on the B60 while the B60 also runs SRIOV for a windows VM.

1 Like