Thoughts from existing B70 users?

I’ve been messing around with the B70 for the last week. The price point is rather difficult to ignore…that is until you see for 300 more dollars you could get a R9700 which seems to be performing at 2-3x the PP and TPS at single use.

I know Intel is new to the game, but I’m a bit worried at the way it appears they are handling the B70 rollout. It’s like it has no identifiable target market…both in software or consumer type.

I think Wendell nailed it in the Battlemage video that Intel is basically selling this B70 card on a “promise”… and there are 3 major things to me that make it feel like it was a $1000 gamble…

  1. Intel marketed this as an AI workstation card… and targeted existing RTX Linux infrastructure, but really positioned this card’s price point perfectly for local enthusiast, then seemed to drop the ball on backend Windows integration.
  2. It’s not clear there is a definitive software path. Seems like Intel has settled on SCYL, but OpenVino sits on oneDNN and is supposed to leverage XMX, of which integration was pulled back and archived because of security flaws on Windows AFAIK. Leaving Vulkan as the most optimized version for at least single user token and prompt processing.
  3. Intel is focused on gaming with the B70 now? But….it was a workstation card right? Currently it’s seems there is no way to leverage XMX…at least on Windows. I would think that this would be a priority for Intel over most anything else?

I want to experiment with the OpenVeno route because according to all documentation I’m reading this is the most optimized stack for the XMX matrices in theroy. But as of now on Llama CPP the most optimized outputs seem to be coming from GGUF running Llama CPP on a Vulkan backend…and this really shouldn’t be the case.

I’m going in on building a dedicated Ubuntu server hoping the stack is more mature there…the documentation suggests that it is, but I’m not seeing the raw numbers indicate that any Linux stack is outperforming Vulkan either.

I guess I’m reaching out to people that are using the B70 trying to get thoughts on your experiences…this was a bit of a chunk, so Id like to see it succeed, but I’m becoming discouraged about the best path forward based on what I’m seeing…

1 Like

I removed my post because I want to make absolutely sure that I’m correct but I think I did something…. I think I did a thing

i have seen the absolute best numbers on windows with llama.cpp and vulkan backend.

2-3 times faster than on linux.

I tried the backends. all of them. and vllm. and openvino. all sucked donkey’s balls, performance wise, especially compared to windows vulkan.

i sh*t you not.

I have had a similar experience. Performance with Arc seems to be the best on windows generally. I’m willing to use it if it works well, but I would prefer Linux if there’s an easy way to manage it.

How is support for multiple intel pro cards, I was pretty dismayed when I tried pooling together a B580 and a A310 I had into a workstation to test some LLM capabilities with llama.cpp only to figure out intel doesn’t support pooling memory with the vulkan backend ( I think ? ), only for the OpenVINO backend to not really work either.

Really considering getting a B70 or multiple of the lower tier pro cards.

Don’t get me wrong, l love my b70, dealing with one more headless windows install is the tiniest price to pay for the performance and capability it offers even now. I have hopes that it will improve a lot but I am happy even if it doesn’t.

i’ve hit two brick walls:

  1. Wasn’t able to utilize VRAM above ~22GB, apparently that’s because WDDM doesn’t allow to use more than 75% of VRAM? llama.cpp just crashes if I try to use more.

  2. on gemma-4-26B-A4B-it-UD-Q4_K_M I wasn’t able to use image recognition. llama.cpp for me just crashes like this:

4.55.684.463 I slot update_slots: id 0 | task 3031 | n_tokens = 3082, memory_seq_rm [3082, end)
4.55.684.750 I srv process_chun: processing image…
4.55.685.112 I encoding image slice…
D:\a\llama.cpp\llama.cpp\ggml\src\ggml-vulkan\ggml-vulkan.cpp:1974: GGML_ASSERT(!src0 || get_misalign_bytes(ctx, src0) == 0) failed

Another issue, probably related to the “invisible” VRAM wall, is crash on prompt saving:

Never had any of these issues in Linux, but in Linux I only get half the performance, at best.

Since you’re on Windows, try using LM Studio. I can fill the VRAM all the way. Do you have an iGPU you can use maybe for display? My machine is headless, so maybe that is it.

With LLAMA.CPP on LInux (Ubuntu 26.04 which fully supports the card), I’m getting much better performance with LLAMA-CPP SYCL than LLAMA-CPP VULKAN on the Gemma4 and Qwen 3.6 MOE models (for example - Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf). 70 TOK/SEC generation (SYCL) vs 45 TOK/SEC generation (VULKAN).

UBUNTU 26.04

Commands

mkdir ~/src
cd ~/src
git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp  

editted the file ./devops/intel.DockerFile

  • modified LINE 7 from “ARG GGML_SYCL_F16=OFF” → "ARG GGML_SYCL_F16=“ON”

  • modified LINE 64 from “COPY --from=build /app/lib/ /app” → “COPY --from=build /app/lib/ /app/”

  • modified LINE 65 from “COPY --from=build /app/full /app” → "COPY --from=build /app/full /app/

  • modified LINE 92 from “COPY --from=build /app/lib/ /app” → “COPY --from=build /app/lib/ /app/”

  • modified LINE 93 from “COPY --from=build /app/full /llama-cli” → "COPY --from=build /app/full /app/

  • modified LINE 104 from “COPY --from=build /app/lib/ /app” → “COPY --from=build /app/lib/ /app/”

  • modified line 105 from “COPY --from=build /app/full/llama-server /app” → “COPY --from=build app/full/llama-server /app/”

built the container file

docker build -t local/rmfllama.cpp:full-intel --target full -f .devops/intel.Dockerfile .

downloaded a model

mkdir /data/model/llm
cd /data/model/llm
wget https://huggingface.co/unsloth/Qwen3.6-35B-A3B-GGUF/resolve/main/Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf?download=true
mv Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf\?download\=true Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf

then deployed / ran the container

note: change the --group-add 141 command to the right group number for “render” in /etc/groups

docker run -d --name "llama-cpp-server" -v /data/llm/models:/models --restart unless-stopped -p 8080:8080 --device /dev/dri/renderD128:/dev/dri/renderD128 --device /dev/dri/card0:/dev/dri/card0 --group-add 141 local/llama.cpp:full-intel --server -m /models/Qwen3.6-27B-UD-Q4_K_XL.gguf --port 8080 --host 0.0.0.0 -t 1 -c 131072 --n-gpu-layers 99 --temp 0.6 --top-p 0.95 --top-k 20 --min-p 0.00 --repeat-penalty 1.00 --presence-penalty 0.0 --chat-template-kwargs '{"preserve_thinking": true}'

you can then use a browser and open http://%YOUR_IP_ADDRESS%:8080 to get to the chat interface. If you want to enable the model to be used via coding agent, you may want to add “–api-key %SOME_API_KEY%” to the docker command line.

to check docker container processes

> docker ps -a

to stop a running instance

> docker stop %CONTAINER_ID%

to remove a stopped instance

> docker rm %CONTAINER_ID%

Note: it will auto restart during reboots, if you don’t want that then stop it using the “docker stop %CONTAINER_ID%”. You can start it manually by doing **“docker start %CONTAINER_ID%”.

Here’s how I benchmarked it

docker run -it --rm -v /data/llm/models:/models --device /dev/dri/renderD128:/dev/dri/renderD128 --device /dev/dri/card0:/dev/dri/card0 --group-add 141 local/llama.cpp:full-intel --bench -m /models/Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf

and here’s the output I received

docker run -it --rm -v /data/llm/models:/models --device /dev/dri/renderD128:/dev/dri/renderD128 --device /dev/dri/card0:/dev/dri/card0 --group-add 141 local/llama.cpp:full-intel --bench -m /models/Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf
load_backend: loaded SYCL backend from /app/libggml-sycl.so
load_backend: loaded CPU backend from /app/libggml-cpu-haswell.so

model size params backend ngl test t/s
qwen35moe 35B.A3B Q4_K - Medium 20.81 GiB 34.66 B SYCL 99 pp512 606.52 ± 5.92
qwen35moe 35B.A3B Q4_K - Medium 20.81 GiB 34.66 B SYCL 99 tg128 71.92 ± 0.09

build: 983ca8992 (8952)

bare-metal or as a VM? if VM: proxmox?

what hardware (CPU/Motherboard/RAM)?

Appreciated!

AMD 5900XT CPU, MSI X570-UNIFY Motherboard, 96GB of RAM, Intel Arc B70..

Benchmarks

root@nas:/data/llm/models# docker run -it --rm -v /data/llm/models:/models --device /dev/dri/renderD128:/dev/dri/renderD128 --device /dev/dri/card0:/dev/dri/card0 --group-add 141 local/llama.cpp:full-intel --bench -m /models/Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf
load_backend: loaded SYCL backend from /app/libggml-sycl.so
load_backend: loaded CPU backend from /app/libggml-cpu-haswell.so

model size params backend ngl test t/s
qwen35moe 35B.A3B Q4_K - Medium 20.81 GiB 34.66 B SYCL 99 pp512 609.43 ± 6.10
qwen35moe 35B.A3B Q4_K - Medium 20.81 GiB 34.66 B SYCL 99 tg128 71.79 ± 0.07

build: 983ca8992 (8952)

root@nas:/data/llm/models# docker run -it --rm -v /data/llm/models:/models --device /dev/dri/renderD128:/dev/dri/renderD128 --device /dev/dri/card0:/dev/dri/card0 --group-add 141 local/llama.cpp:full-intel --bench -m /models/Qwen3.6-27B-UD-Q4_K_XL.gguf
load_backend: loaded SYCL backend from /app/libggml-sycl.so
load_backend: loaded CPU backend from /app/libggml-cpu-haswell.so

model size params backend ngl test t/s
qwen35 27B Q4_K - Medium 16.39 GiB 26.90 B SYCL 99 pp512 795.74 ± 4.08
qwen35 27B Q4_K - Medium 16.39 GiB 26.90 B SYCL 99 tg128 19.97 ± 0.08

build: 983ca8992 (8952)

nope, uses llama.cpp-vulkan in the background and is unable to fill more than 21.5GB of VRAM before it errors out gracefully (LMStudio doesn’t crash though).

no igpu in that system, just a virtual display from proxmox. i could try setting that to “none”…

I’m able to pool memory of a B70 with a B580 with the SYCL backend and Vulkan backend.

You may have run into some issues with how the memory was fit since you do have to be a little hands-on with some of the settings for cards of different capacities. Alternatively, you could also be having issues since Alchemist has different hardware features than Battlemage (i.e. native double precision isn’t available on Alchemist).

vLLM seems to offer better perf for the Intel cards on Linux, but definitely less friendly to setup vs llama.cpp

This is very odd as I can fill the memory to the brim… You’re giving the VM at least 32g of system RAM too, right?

I’m guessing that’s probably it, I didn’t try the SYCL backend but it might as well be worth a try.

yes, it has exactly 32GB. what windows version and what driver exactly are you using? Pro or non Pro? Which version?
Thx!

Details, and proof that it lets me exercise that VRAM. Maybe this will help?

today i tested windows 10 bare metal on the same hardware, instead of windows server 2022 VM on proxmox with gpu passthrough.

i was able to utilize the full amount of VRAM but weirdly enough the dedicated gpu memory block just disappeared from the B70. the two B50s which were sitting alongside the B70 in the same server did have the dedicated GPU memory metric in Task manager.

I didn’t have enough time to find out what exactly was going on but I was able to load gemma4 26B MoE q4_k_m with 256k context and it performed as expected, so…

maybe was put into some kind of different “compute” state by LMStudio, because it was the only one of the three GPUs that I had ticked in LMStudio for inference usage.

:man_shrugging:

today i tested windows 10 bare metal on the same hardware, instead of windows server 2022 VM on proxmox with gpu passthrough.

i was able to utilize the full amount of VRAM but weirdly enough the dedicated gpu memory block just disappeared from the B70. the two B50s which were sitting alongside the B70 in the same server did have the dedicated GPU memory metric in Task manager.

I didn’t have enough time to find out what exactly was going on but I was able to load gemma4 26B MoE q4_k_m with 256k context and it performed as expected, so…

maybe was put into some kind of different “compute” state by LMStudio, because it was the only one of the three GPUs that I had ticked in LMStudio for inference usage.

:man_shrugging:

edit: what windows version/edition are you running?

On this system, Windows 11 Pro 25H2 with Atlas tweaks.

1 Like