Intel B70 Launch - Unboxed and Tested

B70 Info

A B70?!? In this economy? And yet this may be the “best” path to 128gb vram for around/less than the cost of DGX spark. Assuming you have a machine to plug them into.

B70 32gb AI llm-scaler (vLLM) testing

 vllm serve /llm/models/hub/models--Qwen--Qwen3.5-27B/snapshots/b7ca741b86de18df552fd2cc952861e04621a4bd   --served-model-name Qwen/Qwen3.5-27B   --port 8000 --no-enable-prefix-caching --enable-chunked-prefill --max-num-seqs 128 --block-size 64 --enforce-eager  --dtype bfloat16 --disable-custom-all-reduce --tensor-parallel-size 4

Avg generation throughput: 540.0 tokens/s, Running: 50 reqs,

DANG!

============ Serving Benchmark Result ============
Successful requests:                     50
Failed requests:                         0
Benchmark duration (s):                  69.22
Total input tokens:                      51200
Total generated tokens:                  25600
Request throughput (req/s):              0.72
Output token throughput (tok/s):         369.83
Peak output token throughput (tok/s):    550.00
Peak concurrent requests:                50.00
Total token throughput (tok/s):          1109.48
---------------Time to First Token----------------
Mean TTFT (ms):                          11467.51
Median TTFT (ms):                        11316.84
P99 TTFT (ms):                           21193.65
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms):                          110.70
Median TPOT (ms):                        111.14
P99 TPOT (ms):                           121.26
---------------Inter-token Latency----------------
Mean ITL (ms):                           110.70
Median ITL (ms):                         92.52
P99 ITL (ms):                            567.33
==================================================

Keep in mind this is 50 requests at once, however.

For a single request the floor performance with dynamic FP8 quant of qwen 27b on a single gpu was

Avg prompt throughput: 85.4 tokens/s, , Avg generation throughput: 13.4 tokens/s,

HOWEVER: I would not recommend a single B70 for Qwen 27B dense in fp8 dynamic quant. For vLLM benchmarking I had to lower the context and set the max gpu memory utilization to 0.8 or it was unstable. Two B70s for the Q8 Qwen 3.5 27b was fine. Similarly, there was simply no room to work with Qwen 27b bf16 on two B70s.

The model perplexity/stability was reasonable in dynamic Q8. The native FP8 model crashed on startup, however. So if you have trouble, download the native bf16 and use the dynamic quant? That seems like an odd happenstance to me, but I assume it’s a me problem.

Two B70s and Qwen 3.5 27b in dynamic FP8 mode

 vllm bench serve   --model Qwen/Qwen3.5-27B   --dataset-name random   --random-input-len 1024   --random-output-len 512   --num-prompts 8   --request-rate inf   --port 8000

============ Serving Benchmark Result ============
Successful requests:                     8
Failed requests:                         0
Benchmark duration (s):                  41.86
Total input tokens:                      8192
Total generated tokens:                  4096
Request throughput (req/s):              0.19
Output token throughput (tok/s):         97.84
Peak output token throughput (tok/s):    112.00
Peak concurrent requests:                8.00
Total token throughput (tok/s):          293.53
---------------Time to First Token----------------
Mean TTFT (ms):                          1961.21
Median TTFT (ms):                        2113.37
P99 TTFT (ms):                           2826.61
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms):                          77.84
Median TPOT (ms):                        77.56
P99 TPOT (ms):                           80.33
---------------Inter-token Latency----------------
Mean ITL (ms):                           77.84
Median ITL (ms):                         76.27
P99 ITL (ms):                            78.92
==================================================

Interestingly, this is the type of latency I was looking for (hoping for) in the video. 8 requests at once, though.

Single request is still capped around 14 t/s


============ Serving Benchmark Result ============
Successful requests:                     1
Failed requests:                         0
Benchmark duration (s):                  38.63
Total input tokens:                      1024
Total generated tokens:                  512
Request throughput (req/s):              0.03
Output token throughput (tok/s):         13.25
Peak output token throughput (tok/s):    14.00
Peak concurrent requests:                1.00
Total token throughput (tok/s):          39.76
---------------Time to First Token----------------
Mean TTFT (ms):                          443.08
Median TTFT (ms):                        443.08
P99 TTFT (ms):                           443.08
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms):                          74.73
Median TPOT (ms):                        74.73
P99 TPOT (ms):                           74.73
---------------Inter-token Latency----------------
Mean ITL (ms):                           74.73
Median ITL (ms):                         74.87
P99 ITL (ms):                            76.95
==================================================

all other parameters being equal.

Gaming?

Watch this space. Currently (3/25) we have

UPDATE 3/26: yay driver works now! Here are the results

Cyberpunk 2077: Phantom Liberty - Ultra

Shadow of the Tomb Raider - Highest

Monster Hunter Wilds - Medium

23 Likes

what can it do virtualization wise? how many VFs can it do ?

6 Likes

How does comfy stack up on one of these? Flux1 and Qwen Image 2512 would be great

2 Likes

I thought BMG-G31 support was upstreamed to the Xe Linux driver many months back when the first round of “it’s almost here” rumors began. There was also that leaked HP driver/firmware package that might be fun to try on one of these boards.

2 Likes

what is the estimated price-tag ? more than 2 kidneys and a liver ?

1 Like

inb4 this doesn’t make it to the UK market for over a year. B50s are juuuuust barely turning up over here…

1 Like

Given they both seem to be competing for the same market, it’ll be interesting to see how the B70 compares to the R9700, especially with rocm 7.2. I need to upgrade from my 6900xt and want at least 32gb of vram without having to run multiple cards.

2 Likes

Would love to see a llama.cpp vulkan benchmark on a few models if possible

3 Likes

Msrp is (was?) $949 USD per card

2 Likes

:thinking: with 128 GB VRAM ??? are you sure ?

I thought it was 32x4 for 128gb total

2 Likes

They don’t have 128GB VRAM. They’re just the cheapest way to get to 128GB VRAM.

I wouldn’t think they’d perform too well - Intel cards don’t do all that well with llama.cpp (at least, the mainline version).

That’s kind of the problem…you’re still stuck with infrequent updates to Intel’s forks of vLLM and llama.cpp, which means never being able to run the latest models. If only Intel would actually contribute back to the projects, their cards would be so much more usable.

2 Likes

Less than four 5090s and less than a Pro 6000, which is good enough for me. I’m on the market for a new workstation over the next couple weeks and my plans got flipped by this drop.

2 Likes

Honestly, the same could be said of the R9700. They’re a bit more expensive than the B70, but they also have a proven record of good performance. Don’t buy Intel cards on the promise of future performance, because it probably won’t come to be.

Been there, got burned.

4 Likes

A thousand USD for 32GB vram…I mean who cares if it works at that point, right?

1 Like

Here to second this.

1 Like

I am considering 3 tbh if performance and support continues
Currently using Strix Halo, but pp512 takes way too long for large contexts

2 Likes

I can’t wait to get 3x of these for myself. :grinning_face:

1 Like

I am also considering a trio. I was about to fire on 3x R9700 Pro AI, but if this can contend, I am all for saving $1,000!

2 Likes

Grumble Intel! I wish my last attempt with getting our OpenCL code running in a container with Intel GPU hadn’t left me such a bad taste in my mouth.