Thoughts from existing B70 users?

Hey, so funny story, after fiddling a bunch and rebooting a lot, I was getting a similar issue. It looks like the memory init might be a tad funny on these cards. (Not surprising) I was able to work around the issue by disabling and then re-enabling the GPU in device manager to force a fresh re-init after Windows has booted. After doing that I can use the full 32G of RAM in Vulkan. See if that helps any.

my vm has been running for 24h now, under full load and brutal memory pressure and nothing has crashed

i swear i am never rebooting that machine

I originally got llama.cpp setup with some cherrypicked changes from sycl: Battlemage (BMG) optimizations — AOT, Q5_K reorder, PAD stride fix, new ops, oneMKL routing by aicss-genai · Pull Request #22066 · ggml-org/llama.cpp · GitHub. I was getting about 60-70 tok/s on Qwen 3.6 35B A3B Q4K_M GGUF from Unsloth. That felt extremely fast.

However, I have web search enabled, and so parallel tasks were very slow! I had heard the vLLM was faster for concurrent throughput, so I had claude deep dive into getting that working. Here is my V1 approach on my nixos machine: vllm setup with qwen3.6 35B A3B by jasonboukheir · Pull Request #13 · jasonboukheir/dotfiles · GitHub.

You can see in there I have a docs/VLLM.md file that the agent was using to keep track of info. Notably, are these results:

```
Working setup: intel/llm-scaler-vllm:0.14.0-b8.2 + Qwen/Qwen3.6-35B-A3B (BF16 base) + --quantization sym_int4 (online IPEX/GGML Q4_0 inline-pack) + --kv-cache-dtype fp8. Validated end-to-end via raw podman before translation:

  • VRAM: 19.01 GiB model footprint
  • KV cache: 206,720 tokens at --max-model-len 32768 (with fp8 KV; doubles the bf16-KV baseline of 103,104)
  • Throughput (200-token completions, single benchmark, no soak):
    • 1-stream: 20.2 tok/s
    • 4-stream: 79 tok/s aggregate (19.8 / stream)
    • 8-stream: 155 tok/s aggregate (19.4 / stream)
    • 16-stream: 310 tok/s aggregate (19.4 / stream)
    • 32-stream: 551 tok/s aggregate (17.2 / stream)
    • 64-stream: 868 tok/s aggregate (13.6 / stream)

Per-stream stays within 5% of single-stream through 16-way — the card is far from the kernel ceiling on this model.
```

I’m happy with 20 tok/s single stream as long as I get good throughput on parallel streams.

This is all running on a single arc b70 pro. Looking at how I can get better kv cache quantization in the future, and waiting for official support for AutoRound int4 to get better quality.

I know nothing about graphics programming. However, I have a Claude Max subscription.

Today, I

  1. got a working vllm pipeline with palmfuture/Qwen3.6-35B-A3B-GPTQ-Int4
  2. setup turboquant on it
  3. turned on the XPU graph and got that working. Single stream tok/s went up to 60 tok/s, and kept the multi stream perf benefits of vllm.

Still working on upstream-ing things. If you wanna see what I’m doing via nix, you can here: dotfiles/hosts/brutus/services/local-llm.nix at 9826e0c4715d4b8a96316cc084b9a0c27e3d526b · jasonboukheir/dotfiles · GitHub

So TIL that the biggest bottleneck is CPU <> GPU round trip in the perf. The XPU Graph caches a bunch of operations so that the CPU can send bigger batches in one trip – however, it eats up a lot of VRAM for the workspace.

Next is to improve some of the XPU kernels to natively reduce this batching. This should just give some perf wins across the board.

Just wanted to NOTE that running LLAMA.CPP with the Vulkan Back End under Windows is so much faster than either SYCL or VULKAN under Linux (UBUNTU 26.04). The Intel ARC Drivers (for Windows) version 32.0.101.8737 is stable for LLAMA.CPP use as long as you install the drivers with the Clean Installation option otherwise you’ll get crashes.

Here’s the difference between VULKAN under Linux (Ubuntu 26.04) vs WINDOWS

[WINDOWS]

c:\Development\tools\llama.cpp>llama-bench -m ..\models\Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf
load_backend: loaded RPC backend from c:\Development\tools\llama.cpp\ggml-rpc.dll
ggml_vulkan: Found 1 Vulkan devices:
ggml_vulkan: 0 = Intel(R) Arc™ Pro B70 Graphics (Intel Corporation) | uma: 0 | fp16: 1 | bf16: 0 | warp size: 32 | shared memory: 49152 | int dot: 1 | matrix cores: KHR_coopmat
load_backend: loaded Vulkan backend from c:\Development\tools\llama.cpp\ggml-vulkan.dll
load_backend: loaded CPU backend from c:\Development\tools\llama.cpp\ggml-cpu-haswell.dll

model size params backend ngl test t/s
qwen35moe 35B.A3B Q4_K - Medium 20.81 GiB 34.66 B Vulkan 99 pp512 1859.81 ± 253.07
qwen35moe 35B.A3B Q4_K - Medium 20.81 GiB 34.66 B Vulkan 99 tg128 100.75 ± 0.11

build: c3c150539 (8996)

[LINUX]

root@nas:/storage/src/llama.cpp/build/bin# ./llama-bench -m /data/llm/models/Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf
ggml_vulkan: Found 1 Vulkan devices:
ggml_vulkan: 0 = Intel(R) Graphics (BMG G31) (Intel open-source Mesa driver) | uma: 0 | fp16: 1 | bf16: 1 | warp size: 32 | shared memory: 49152 | int dot: 1 | matrix cores: KHR_coopmat

model size params backend ngl test t/s
qwen35moe 35B.A3B Q4_K - Medium 20.81 GiB 34.66 B Vulkan 99 pp512 1355.78 ± 7.63
qwen35moe 35B.A3B Q4_K - Medium 20.81 GiB 34.66 B Vulkan 99 tg128 44.42 ± 0.00

build: 63d93d173 (9007)

It’s pretty easy to run llama-server ( as a WINDOWS service ) with NSSM (the non-sucking service manager) utility. I personally just run it this way and use ‘opencode’ to connect to it as an OPENAI provider.

Quite interesting how much faster the Windows Vulkan driver is compared to the Linux one for this use-case.

One last note, the development branch of the MESA Vulkan Drivers - 3.2-Dev is considerably faster than the one shipped with Ubuntu 26.04 (on a B70). It basically doubles the TG under Linux and is faster than the LLAMACPP SYCL back end. Looks like the MESA team implemented ‘NV_coopmat2’ extension and exposed it in MESA and software that looks for this extension (LLAMA.CPP) gets a considerable performance boost. Additionally, it looks like the B70 cards are “downclocking” when running LLAMA CPP vulkan workloads as the firmware doesn’t consider the type of workload to warrant the clock speeds (specially in TG). Some people are getting better performance when forcing a locked GPU frequency.

Here’s how I compiled an updated libvulkan_intel.so

Created a ubuntu 26.04 VM using ubuntu-server as the base using libvirtd

Logged into the VM via VNC

then did the following in the VM

> apt install meson glslang-tools pkg-config libclc-21-dev python-is-python3 python3-mako libdrm-dev llvm-dev libllvmspirvlib-21-dev spirv-tools-dev clang libclang-dev libwayland-dev libwayland-client0 wayland-client wayland-protocols wayland-scanner++ xcb libxcb1-dev libxcb-randr0-dev libx11-xcb-dev libxcb-dri3-dev libxcb-present-dev libxcb-shm0-dev libxshmfence-dev libxrandr-dev

> mkdir /opt/src

> cd /opt/src

> git clone https://gitlab.freedesktop.org/mesa/mesa.git

> cd mesa

> meson setup builddir/ -Dbuildtype=release -Dgallium-drivers=[] -Dvulkan-drivers=intel -Dopengl=false -Dglx=disabled -Degl=disabled -Dgbm=disabled -Dgles1=disabled -Dgles2=disabled

> meson compile -C builddir/

Finally, copied the binary ‘builddir/src/intel/vulkan/libvulkan_intel.so’ SOMEWHERE and overwrite /lib/x86_64-linux-gnu/libvulkan_intel.so on the OS i’m trying to run llama-server (vulkan) (Ubuntu 26.04).

That’s the only file needed from the MESA 26.2-dev

QWEN 3.6

root@nas:/storage/src/llama.cpp/build/bin# ./llama-bench -m /data/llm/models/Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf
ggml_vulkan: Found 1 Vulkan devices:
ggml_vulkan: 0 = Intel(R) Graphics (BMG G31) (Intel open-source Mesa driver) | uma: 0 | fp16: 1 | bf16: 1 | warp size: 32 | shared memory: 49152 | int dot: 1 | matrix cores: NV_coopmat2

model size params backend ngl test t/s
qwen35moe 35B.A3B Q4_K - Medium 20.81 GiB 34.66 B Vulkan 99 pp512 1370.89 ± 8.39
qwen35moe 35B.A3B Q4_K - Medium 20.81 GiB 34.66 B Vulkan 99 tg128 81.12 ± 0.13

build: 2496f9c (9049)

GEMMA 4

root@nas:/storage/src/llama.cpp/build/bin# ./llama-bench -m /data/llm/models/gemma-4-26B-A4B-it-UD-Q4_K_M.gguf
ggml_vulkan: Found 1 Vulkan devices:
ggml_vulkan: 0 = Intel(R) Graphics (BMG G31) (Intel open-source Mesa driver) | uma: 0 | fp16: 1 | bf16: 1 | warp size: 32 | shared memory: 49152 | int dot: 1 | matrix cores: NV_coopmat2

model size params backend ngl test t/s
gemma4 26B.A4B Q4_K - Medium 15.70 GiB 25.23 B Vulkan 99 pp512 1750.07 ± 8.30
gemma4 26B.A4B Q4_K - Medium 15.70 GiB 25.23 B Vulkan 99 tg128 78.55 ± 0.01

build: 2496f9c (9049)