The latest 8685+ llama.cpp has some nice SYCL improvements. Q8_0 support was added by - PR#21527
llama-cpp-sycl
–bench -m /models/qwen3.5_9B/Qwen3.5-9B-UD-Q4_K_XL.gguf -ngl 100 -r 3 -o md -p 2048,16384
load_backend: loaded SYCL backend from /app/libggml-sycl.so
load_backend: loaded CPU backend from /app/libggml-cpu-alderlake.so
| model | size | params | backend | ngl | test | t/s |
|---|---|---|---|---|---|---|
| qwen35 9B Q4_K - Medium | 5.55 GiB | 8.95 B | SYCL | 100 | pp2048 | 1633.72 ± 0.59 |
| qwen35 9B Q4_K - Medium | 5.55 GiB | 8.95 B | SYCL | 100 | pp16384 | 1464.48 ± 0.36 |
| qwen35 9B Q4_K - Medium | 5.55 GiB | 8.95 B | SYCL | 100 | tg128 | 33.41 ± 0.00 |
build: 71a81f6fc (8688)
On my B60 Pro with qwen3.5 9B Q4_K_XL I’m not getting the full 3x TG speedup but definitely prefill speeds are up a long way which is very handy for interactivity with long system prompts like when using claude code or opencode etc
Very cool to see these optimizations still coming in!