I saw a reply in the B70 thread about Qn_0 quants potentially being faster on intel GPU.
Gave it a go just now and it seems to have some merit, mostly on token generation:
| model | size | params | backend | ngl | test | t/s |
|---|---|---|---|---|---|---|
| qwen35 ?B Q4_K - Medium | 2.54 GiB | 4.21 B | Vulkan | 100 | pp2048 | 821.08 ± 0.28 |
| qwen35 ?B Q4_K - Medium | 2.54 GiB | 4.21 B | Vulkan | 100 | tg512 | 30.87 ± 0.03 |
| qwen35 ?B Q4_K - Medium | 2.54 GiB | 4.21 B | Vulkan | 100 | pp2048+tg512 | 132.27 ± 0.70 |
| model | size | params | backend | ngl | test | t/s |
|---|---|---|---|---|---|---|
| qwen35 ?B Q4_0 | 2.40 GiB | 4.21 B | Vulkan | 100 | pp2048 | 812.10 ± 0.80 |
| qwen35 ?B Q4_0 | 2.40 GiB | 4.21 B | Vulkan | 100 | tg512 | 42.89 ± 0.07 |
| qwen35 ?B Q4_0 | 2.40 GiB | 4.21 B | Vulkan | 100 | pp2048+tg512 | 174.50 ± 0.43 |
llama.cpp build: d903f30e2 (8175) with vulkan backend in docker on B60
I now have my opencode setup running with qwen3.5-27B running on RTX Pro 4500 and with the small model that opencode uses for session name generation and other small tasks using the qwen3.5 4B model on the B60 while the B60 also runs SRIOV for a windows VM.