Intel B70 Launch - Unboxed and Tested

same it has what i would call a buzzing/coil whine when in use

Yeah that sounds about right.

Glad to hear it’s normal. Hopefully they get that sorted out in the next one.

Ive spent the better part of the last week having my AI test my dual B70 system for inference. Got llama.cpp working.
Going to throw some benchmark results down below, from my results file in .md . My AI wrote inside << and >>

<<## Phase 1: SYCL vs Vulkan Baselines

### Methodology

Every model was tested on both backends with identical settings (q8_0/q8_0 KV cache, flash attention ON). Three metrics measured per run: prompt processing speed at 512 tokens (pp512), long prompt processing at 4096 tokens (pp4096), and generation throughput for 128 tokens (tg128).

Results

Model Backend tg128 tok/s pp512 tok/s pp4096 tok/s
Qwen3.6-27B-Q4_K_S SYCL 25.3 751 981
Qwen3.6-27B-Q4_K_S Vulkan 18.9 483 341
Qwen3.6-35B-MXFP4_MOE SYCL 52.3 929 1610
Qwen3.6-35B-MXFP4_MOE Vulkan 33.5 1185* 1018
gemma-26B-MoE SYCL 59.6 1502 2139
gemma-26B-MoE Vulkan 47.7 910 275*
gemma-31B-dense SYCL 23.5 570 736
gemma-31B-dense Vulkan 15.7 254 68*

*Anomaly: Vulkan pp512 for Qwen35B and pp4096 collapse for gemma models likely caused by VRAM pressure triggering CPU fallback.

### Breakdown

SYCL dominates across the board. The tg128 gap is consistent at 1.3–1.6× faster on every model. Vulkan has one anomalous win (Qwen35B pp512) but collapses catastrophically on long prompts for gemma models — pp4096 drops to 275 tok/s on gemma-26B and 68 tok/s on gemma-31B, compared to SYCL’s 2139 and 736 respectively.

**Verdict: Always use SYCL.** Vulkan is not viable for production inference on Arc B70.

## Phase 2: KV Cache Quantization (f16/f16 vs q4_0/q4_0)

### Methodology

All tests ran on SYCL with flash attention ON, batch/ubatch 2048/2048. Three cache types tested: f16/f16 (full precision), q4_0/q4_0 (aggressive quantization), and q8_0/f16 (mixed).

Results — tg128 Generation Throughput

Model f16/f16 tok/s q4_0/q4_0 tok/s Δ
Qwen3.6-27B-Q4_K_S 25.7 25.3 +1.4% (f16)
gemma-26B-MoE 65.8 59.3 +11.0% (f16) :white_check_mark:
gemma-31B-dense 25.4 23.4 +8.5% (f16)

Results — pp512 Prompt Processing

Model f16/f16 tok/s q4_0/q4_0 tok/s Δ
Qwen3.6-27B-Q4_K_S 781 752 +3.9% (f16)
gemma-26B-MoE 1706 1495 +14.1% (f16) :white_check_mark:
gemma-31B-dense 707 569 +24.2% (f16) :white_check_mark::white_check_mark:

Results — pp4096 Long Prompt Processing

Model f16/f16 tok/s q4_0/q4_0 tok/s Δ
Qwen3.6-27B-Q4_K_S 987 982 +0.5% (f16)
gemma-26B-MoE 2182 2147 +1.6% (f16)
gemma-31B-dense 755 734 +2.9% (f16)

### Breakdown

f16/f16 KV cache is consistently faster — up to **+24%** on pp512 for gemma-31B-dense and **+11%** on tg128 for gemma-26B-MoE. The benefit is most pronounced during prompt processing (where the cache is actively written) rather than generation (where it’s mostly read).

The tradeoff: f16/f16 uses roughly 4× more VRAM per cached token than q4_0/q4_0. On a 32 GB GPU, that matters for long-context workloads — but the speed gain is real and measurable.

**Verdict: Use f16/f16 KV cache when VRAM allows.** The +1–24% speed boost across all metrics is worth the memory cost on most models.

## Phase 3: Flash Attention ON vs OFF

### Methodology

All tests ran on SYCL with f16/f16 KV cache, batch/ubatch 2048/2048. Only flash attention toggled between runs.

### Results — tg128 Generation Throughput

Model FA ON tok/s FA OFF tok/s Δ
Qwen3.6-27B-Q4_K_S 25.6 25.9 +1.2% (FA Off)
Qwen3.6-35B-MXFP4_MOE 52.5 53.5 +1.9% (FA Off)
gemma-26B-MoE 65.8 63.4 -3.7% (FA On wins)
gemma-31B-dense 25.4 24.8 -2.4% (FA On wins)

### Breakdown

Flash attention has negligible impact on tg128 — within ±4%. Interestingly, Qwen models are slightly faster with FA off (possibly less overhead for short generation), while gemma models prefer FA on. Since flash attention saves VRAM and helps pp4096 performance, keeping it ON as default makes sense despite the marginal tg128 tradeoff.

**Verdict: Keep flash attention ON.** The ±4% impact on generation is noise; the VRAM savings and long-prompt benefits outweigh any micro-optimization.

## Phase 5: Concurrency Saturation Curves

### Methodology

We tested streams 1 through 20 (with fine-grained testing at 2, 3, and 4) on all four models. Each stream fires a concurrent request with `max_tokens=256` and temperature 0.0. Aggregate throughput is the sum of all streams’ tok/s; per-stream tok/s shows individual degradation.

### Full Results

Model Stream 1 Stream 2 Stream 3 Stream 4 Stream 5 Stream 10 Stream 15 Stream 20 Ceiling
gemma-4-26B-A4B-it-QAT-Q4_0 52.5 75.3 89.6 100.0 85.8 94.6 97.8 98.8 ~99 tok/s
gemma-4-12B-it-QAT-Q4_0 52.6 73.1 94.6 116.3 90.0 92.2 97.2 100.7 ~101 tok/s
Qwen3.6-35B-A3B-MXFP4_MOE 39.6 48.3 61.4 72.9 69.2 66.2 69.7 71.6 ~72 tok/s
Qwen3.6-27B-Q8_0 12.7 20.4 27.4 32.4 24.6 28.9 31.8 33.1 ~34 tok/s

>>

From my other chats with google AI, intel cards seems to have 2 compute pathways or something, so it makes sense for 2 or 4 concurrency streams to have the best numbers (2 cards?), and it looks like scheduling bandwidth hits hard at 5 streams

I completely failed, or just got junk data with MTP results. LM studio shows MTP working, I must be passing the wrong flags or something, my AI built the llama.cpp handler…

Ive also been struggling to get testing done on both GPU’s, no matter what I tell my bot to do, those models end up on a single GPU. I want to try to get it across both so I can see what performance degradation/uplift there is between single and dual layer split and dual tensor split.

I think theres also a slowdown because I used Q8 for the 27b model, I read somewhere that q8 is a specific slowdown until something gets an update.

Forgot some testing:

<<

Phase 6: Batch Sizing (ubatch 1024 vs 2048 vs 4096)

Methodology

All tests ran on SYCL with f16/f16 KV cache, flash attention ON. Three ubatch sizes tested per model: 1024, 2048, and 4096 (batch_size matched). tg128 measures generation throughput; pp4096 measures long prompt processing speed.

Results — tg128 Generation Throughput

Model ubatch=1024 ubatch=2048 ubatch=4096 Δ
Qwen3.6-27B-Q4_K_S 25.8 25.7 25.7 ±0% (identical)
Qwen35B-MXFP4_MOE 53.9 53.7 53.9 ±0% (identical)
gemma-26B-MoE 65.8 65.8 65.8 ±0% (identical)
gemma-31B-dense 25.4 25.4 25.4 ±0% (identical)

Generation is completely unaffected by batch sizing. Every combo shows identical tg128 within ±0.2 tok/s across all models. Generation is compute-bound, not memory-bandwidth bound.

Results — pp4096 Long Prompt Processing

Model ubatch=1024 ubatch=2048 ubatch=4096 Winner
gemma-31B-dense 605 tok/s 752 tok/s 755 tok/s +25% at ubatch≥2048 :white_check_mark:
gemma-26B-MoE 1709 tok/s 2187 tok/s 2299 tok/s +34% at ubatch=4096 :white_check_mark::white_check_mark:
Qwen3.6-27B-Q4_K_S 887 tok/s 988 tok/s 1007 tok/s +13% at ubatch=4096 :white_check_mark:
Qwen35B-MXFP4_MOE 1230 tok/s 1616 tok/s 1906 tok/s +55% at ubatch=4096 :white_check_mark::white_check_mark::white_check_mark:

Breakdown

ubatch size matters a LOT for long prompt processing. The pattern is clear: larger ubatch = more tokens processed per forward pass during prompt evaluation. This directly impacts pp4096 but has zero effect on tg128 (generation).

The sweet spot: batch=4096, ubatch=4096 gives the best pp4096 across all models with no downside on generation. Qwen35B-MOE sees a massive +55% boost — from 1230 to 1906 tok/s — just by increasing ubatch from 1024 to 4096.

Verdict: Use batch_size=4096, ubatch_size=4096. It’s a free +13–55% boost on long prompt processing with zero cost to generation throughput.

This part of testing was very helpful to me.
Note that some model quants were different in different phases. This has led to some frustration with my AI, but , at least it did the tests while I was doing other stuff (note Qwen3.6-27b at q4 and q8).

1 Like

The best Intel llm performance is unlocked with the urakozz github gist - I can’t post links but you can bing that.

without mtp

| model       |   test |            t/s |     peak t/s |      ttfr (ms) |   est_ppt (ms) |   e2e_ttft (ms) |
|:------------|-------:|---------------:|-------------:|---------------:|---------------:|----------------:|
| Qwen3.8-27B | pp4096 | 2313.29 ± 8.87 |              | 1860.31 ± 6.81 | 1771.10 ± 6.81 |  1860.31 ± 6.81 |
| Qwen3.8-27B |  tg256 |   30.76 ± 0.01 | 31.67 ± 0.47 |                |                |                 |

and with mtp

| model       |   test |             t/s |     peak t/s |       ttfr (ms) |    est_ppt (ms) |   e2e_ttft (ms) |
|:------------|-------:|----------------:|-------------:|----------------:|----------------:|----------------:|
| Qwen3.8-27B | pp4096 | 2148.60 ± 16.39 |              | 2005.37 ± 14.75 | 1906.78 ± 14.75 | 2005.37 ± 14.75 |
| Qwen3.8-27B |  tg256 |    45.40 ± 1.98 | 55.00 ± 0.82 |                 |                 |                 |

Also 12k pp/s 100tok/s on qwen 3.6 35b a3b - although that model isn’t as capable as I’d like.

I’m more than happy with daily driving 3.8 27b with hermes

It does take a bit of fiddling. I patched my vllm with some open PR’s to fix prefix caching and something else I’ve forgotten. Currently fumbling my way through enabling kv cache offloading - right now the choice is mtp or cache offload, and mtp wins for me.

In fact I’m so happy with the b70 that I have a second one arriving this week, to be run on a bifurcated pcie 4.0 slot - either tensor parallelism works or I run 2 independent models

If there’s any interest I’m happy to go into more detail

What Quant did you use?

Vishva007/Qwen3.8-27B-W4A16-AutoRound-GPTQ

1 Like

Mine draws 45w in idle. Do you have different numbers or just a different pain tolerance?:sweat_smile:

Unless nvtop is wildly inaccurate, I’d say somethings wrong if yours idles at 45 watts. By idling, do you mean keeping the vram loaded with a model?

1 Like

Unfortunately they do idle around 45 watts compared to Nvidia 10w or 15w unloaded.

For anyone trying to virtualize proxmox is stuck in an old kernel 7.0 which can have an impact on performance. I’ve pivoted to IncusOS for virtualization. IncusOS makes it trivial to you stay with the latest kernel (kernel 7.1.9-zabbly+).

I’m curious what kind of hardware everyone’s using Motherboard and CPU choice. Ideally PCI Express 16x gen 4 or 5. For me im on GIGABYTE MC62-G40-00 with AMD Ryzen Threadripper PRO 3945WX 12-Cores.

I’m struggling to get Qwen3.8-27B to run on 2x intel-b70 pros on vLLM. Anyone else have success that outperforms a single card?

Please do!

Yes. I see your numbers, so there is hope. I’ll try something different, maybe it’s the mainboards firmware.

I havent yet deployed this, still doing benchmarking/testing and using my other server for my agent. But, I had my agent create a llm-runner program that recieves the API calls from my agent and spins up the appropriate vllm instance, then when it hasnt recieved calls for a while, unloads vram and idles the GPU’s.

Didnt test the 45W thing, but makes sense. Cards cant sleep if they have to constantly refresh the ram. Will see if dealing with potential bugs I made by having my agent build a middleman is worth the $100/year power savings.

I thought they were only usable in 1,2,4,8, etc , you can have three?

this is correct. except for how llama.cpp scales, not with tensor parallelism. it has more options for scaling than TP