LLM Adventures with Intel Arc Pro B60s, P2P, VLLM

I’m not sure how to begin this so here’s my life story and ramblings as I try and justify my intel purchase.

I feel I need to get this information out to a place that can be scraped by agents or found by future adventurers rather than lost in scattered thoughts and sporadic messages.

I’ll hopefully keep updating this as I go


The build began in December 2025, an awful time to start, made do with what I had already laying around.

A Maxun Z890, an intel 225F and 64gb of questionable latency DDR5.

I bought two Intel B60s around Christmas just to see what it’s like, no way you can buy brand new intel cards cheaper than a used 3090 and run AI well, right? turns out you can, but you really have to work for it.

I had four cards by January, and now in July a PEX88096 10-port switch, 11-slot pcie “backplane” with plans to bring my total up to 8 cards by August.

Platform

Naively I began with the official intel endorsed config which was not a good time, Intel’s BKC/platform at that time was Ubuntu Server 25.04 with pinned dependencies, an offline installer and overall it was just a bad experience, I should’ve known better, they themselves would have known better, no one did better.

currently my platform and recommendation is kernel 7.1+ and compute-runtime being 26.22.38646.4 or newer and realistically any OS where you can install the dependencies and a modern kernel.
(GitHub - intel/compute-runtime: Intel® Graphics Compute Runtime for oneAPI Level Zero and OpenCL™ Driver · GitHub)

if you’re lazy or starting out, with Ubuntu 26.04 LTS you can follow:

Which is relatively new documentation and was not what intel were preaching earlier this year or even what is probably still in the llm-scaler documentation, avoid that.

Current Hardware

OS: Ubuntu 26.04 LTS (Resolute Raccoon) x86_64
Host: X870 Pro-A WiFi
Kernel: Linux 7.1.3-070103-generic
CPU: AMD Ryzen 9 9900X (24) @ 5.66 GHz
GPU 1: Intel Arc Pro B60 @ 2.00 GHz [Discrete]
GPU 2: Intel Arc Pro B60 @ 2.00 GHz [Discrete]
GPU 3: Intel Arc Pro B60 @ 2.00 GHz [Discrete]
GPU 4: Intel Arc Pro B60 @ 2.00 GHz [Discrete]
GPU 5: Intel Arc Pro B60 @ 2.00 GHz [Discrete]
GPU 6: Intel Arc Pro B60 @ 2.00 GHz [Discrete]
GPU 7: Intel Arc Pro B60 @ 2.00 GHz [Discrete]
GPU 8: Intel Arc Pro B60 @ 2.00 GHz [Discrete]

I recently bought 4 more cards bringing my total to 8, it required swapping the 265k plus/Z890 for a 9900x/X870.

having multiple cards on a consumer platform is a bad time, before buying the pcie switch it was running four cards, two at 5.0 x8 and two at 4.0 x4 across different root ports, the performance was awful and the latency was even worse.

but it ran and it was cheap.

I eventually purchased a PEX88096 10-port pcie switch and an 11-slot 8654 to PCIE board off aliexpress, with the current price of ECC memory and server platforms it was about a 10th the price of doing things properly.

example listings of the parts I bought, I did not use these stores, this is an example:

https://www.aliexpress.com/item/1005011868399389.html  
https://www.aliexpress.com/item/1005012272021905.html

It’s worth noting that battlemage is pcie 5.0 and I halved my theoretical bandwidth using a 4.0 switch, however the results seem great and I can’t justify the price of 5.0 switches with the cards themselves being so cheap.

Added bonus of the B60 in particular is it’s wired x8, I only had to run a single 8654 8i cable per slot, per card, that’s at least 20 bucks of savings each.

With the particular 10 port switch I bought, it conveniently gives me the possibility of up to 10 x8 cards with dip switches on the pcie card itself for x16, x8/x8, 4*x4

xpu-smi topology -m

              GPU 0/0     GPU 1/0     GPU 2/0     GPU 3/0     GPU 4/0     GPU 5/0     GPU 6/0     GPU 7/0     NIC0     CPU Affinity
  GPU 0/0     S           PIX         PIX         PIX         PIX         PIX         PIX         PIX         PHB      0-23
  GPU 1/0     PIX         S           PIX         PIX         PIX         PIX         PIX         PIX         PHB      0-23
  GPU 2/0     PIX         PIX         S           PIX         PIX         PIX         PIX         PIX         PHB      0-23
  GPU 3/0     PIX         PIX         PIX         S           PIX         PIX         PIX         PIX         PHB      0-23
  GPU 4/0     PIX         PIX         PIX         PIX         S           PIX         PIX         PIX         PHB      0-23
  GPU 5/0     PIX         PIX         PIX         PIX         PIX         S           PIX         PIX         PHB      0-23
  GPU 6/0     PIX         PIX         PIX         PIX         PIX         PIX         S           PIX         PHB      0-23
  GPU 7/0     PIX         PIX         PIX         PIX         PIX         PIX         PIX         S           PHB      0-23
  NIC0        PHB         PHB         PHB         PHB         PHB         PHB         PHB         PHB         S        0-23

lspci -tv

-[0000:00]-+-00.0  Advanced Micro Devices, Inc. [AMD] Raphael/Granite Ridge Root Complex
           +-00.2  Advanced Micro Devices, Inc. [AMD] Raphael/Granite Ridge IOMMU
           +-01.0  Advanced Micro Devices, Inc. [AMD] Raphael/Granite Ridge Dummy Host Bridge
           +-01.1-[01-2f]----00.0-[02-2f]--+-00.0-[03-0c]----00.0-[04-0c]--+-10.0-[05-08]----00.0-[06-08]--+-01.0-[07]----00.0  Intel Corporation Battlemage G21 [Arc Pro B60]
           |                               |                               |                               \-02.0-[08]----00.0  Intel Corporation Device e2f7
           |                               |                               \-18.0-[09-0c]----00.0-[0a-0c]--+-01.0-[0b]----00.0  Intel Corporation Battlemage G21 [Arc Pro B60]
           |                               |                                                               \-02.0-[0c]----00.0  Intel Corporation Device e2f7
           |                               +-04.0-[0d-1e]----00.0-[0e-1e]--+-00.0-[0f-12]----00.0-[10-12]--+-01.0-[11]----00.0  Intel Corporation Battlemage G21 [Arc Pro B60]
           |                               |                               |                               \-02.0-[12]----00.0  Intel Corporation Device e2f7
           |                               |                               +-08.0-[13-16]----00.0-[14-16]--+-01.0-[15]----00.0  Intel Corporation Battlemage G21 [Arc Pro B60]
           |                               |                               |                               \-02.0-[16]----00.0  Intel Corporation Device e2f7
           |                               |                               +-10.0-[17-1a]----00.0-[18-1a]--+-01.0-[19]----00.0  Intel Corporation Battlemage G21 [Arc Pro B60]
           |                               |                               |                               \-02.0-[1a]----00.0  Intel Corporation Device e2f7
           |                               |                               \-18.0-[1b-1e]----00.0-[1c-1e]--+-01.0-[1d]----00.0  Intel Corporation Battlemage G21 [Arc Pro B60]
           |                               |                                                               \-02.0-[1e]----00.0  Intel Corporation Device e2f7
           |                               +-08.0-[1f-2a]----00.0-[20-2a]--+-00.0-[21]--
           |                               |                               +-08.0-[22-25]----00.0-[23-25]--+-01.0-[24]----00.0  Intel Corporation Battlemage G21 [Arc Pro B60]
           |                               |                               |                               \-02.0-[25]----00.0  Intel Corporation Device e2f7
           |                               |                               +-10.0-[26-29]----00.0-[27-29]--+-01.0-[28]----00.0  Intel Corporation Battlemage G21 [Arc Pro B60]
           |                               |                               |                               \-02.0-[29]----00.0  Intel Corporation Device e2f7
           |                               |                               \-18.0-[2a]--
           |                               +-0c.0-[2b-2e]----00.0-[2c-2e]--+-14.0-[2d]--
           |                               |                               \-15.0-[2e]--
           |                               \-1c.0-[2f]----00.0  Broadcom / LSI PEX880xx PCIe Gen 4 Switch

nvtop

 Device 0 [Battlemage G21 (Arc Pro B60)] PCIe GEN 4@ 8x RX: N/A TX: N/A
 GPU 2000MHz MEM N/A MHz  TEMP  33°C FAN    0RPM POW   7 / 400 W
 GPU[                            0%(eff 0%)] MEM[                     0.039Gi/23.906Gi]

 Device 1 [Battlemage G21 (Arc Pro B60)] PCIe GEN 4@ 8x RX: N/A TX: N/A
 GPU 2000MHz MEM N/A MHz  TEMP  31°C FAN 1305RPM POW   9 / 400 W
 GPU[                            0%(eff 0%)] MEM[                     0.039Gi/23.906Gi]

 Device 2 [Battlemage G21 (Arc Pro B60)] PCIe GEN 4@ 8x RX: N/A TX: N/A
 GPU 2000MHz MEM N/A MHz  TEMP  30°C FAN 1305RPM POW   9 / 400 W
 GPU[                            0%(eff 0%)] MEM[                     0.039Gi/23.906Gi]

 Device 3 [Battlemage G21 (Arc Pro B60)] PCIe GEN 4@ 8x RX: N/A TX: N/A
 GPU 2000MHz MEM N/A MHz  TEMP  33°C FAN    0RPM POW   8 / 400 W
 GPU[                            0%(eff 0%)] MEM[                     0.039Gi/23.906Gi]

 Device 4 [Battlemage G21 (Arc Pro B60)] PCIe GEN 4@ 8x RX: N/A TX: N/A
 GPU 400MHz  MEM N/A MHz  TEMP  32°C FAN    0RPM POW   9 / 400 W
 GPU[                            0%(eff 0%)] MEM[                     0.039Gi/23.906Gi]

 Device 5 [Battlemage G21 (Arc Pro B60)] PCIe GEN 4@ 8x RX: N/A TX: N/A
 GPU 400MHz  MEM N/A MHz  TEMP  32°C FAN    0RPM POW   8 / 440 W
 GPU[                            0%(eff 0%)] MEM[                     0.039Gi/23.906Gi]

 Device 6 [Battlemage G21 (Arc Pro B60)] PCIe GEN 4@ 8x RX: N/A TX: N/A
 GPU 400MHz  MEM N/A MHz  TEMP  32°C FAN    0RPM POW   9 / 400 W
 GPU[                            0%(eff 0%)] MEM[                     0.039Gi/23.906Gi]

 Device 7 [Battlemage G21 (Arc Pro B60)] PCIe GEN 4@ 8x RX: N/A TX: N/A
 GPU 400MHz  MEM N/A MHz  TEMP  32°C FAN    0RPM POW   9 / 440 W
 GPU[                            0%(eff 0%)] MEM[                     0.039Gi/23.906Gi]
old intel topology:

I hit the limitation of this intel consumer platform and only 6 cards will post, any more throws a D4 qcode and fails to post, MMIO space can’t be user adjusted on intel consumer platforms, at least not mine. 6/8 was the final total.

xpu-smi topology -m

              GPU 0/0     GPU 1/0     GPU 2/0     GPU 3/0     GPU 4/0     GPU 5/0     NIC0     NIC1     CPU Affinity    
  GPU 0/0     S           PIX         PIX         PIX         PIX         PIX         SYS      SYS      0-19            
  GPU 1/0     PIX         S           PIX         PIX         PIX         PIX         SYS      SYS      0-19            
  GPU 2/0     PIX         PIX         S           PIX         PIX         PIX         SYS      SYS      0-19            
  GPU 3/0     PIX         PIX         PIX         S           PIX         PIX         SYS      SYS      0-19            
  GPU 4/0     PIX         PIX         PIX         PIX         S           PIX         SYS      SYS      0-19            
  GPU 5/0     PIX         PIX         PIX         PIX         PIX         S           SYS      SYS      0-19            
  NIC0        SYS         SYS         SYS         SYS         SYS         SYS         S        PHB      0-19            
  NIC1        SYS         SYS         SYS         SYS         SYS         SYS         PHB      S        0-19       

lspci -tv

-+-[0000:00]-+-00.0  Intel Corporation Device 7d1b
 |           +-04.0  Intel Corporation Device ad03
 |           +-06.0-[01-29]----00.0-[02-29]--+-00.0-[03-0c]----00.0-[04-0c]--+-10.0-[05-08]----00.0-[06-08]--+-01.0-[07]----00.0  Intel Corporation Battlemage G21 [Arc Pro B60]
 |           |                               |                               |                               \-02.0-[08]----00.0  Intel Corporation Device e2f7
 |           |                               |                               \-18.0-[09-0c]----00.0-[0a-0c]--+-01.0-[0b]----00.0  Intel Corporation Battlemage G21 [Arc Pro B60]
 |           |                               |                                                               \-02.0-[0c]----00.0  Intel Corporation Device e2f7
 |           |                               +-04.0-[0d-1b]----00.0-[0e-1b]--+-00.0-[0f-12]----00.0-[10-12]--+-01.0-[11]----00.0  Intel Corporation Battlemage G21 [Arc Pro B60]
 |           |                               |                               |                               \-02.0-[12]----00.0  Intel Corporation Device e2f7
 |           |                               |                               +-08.0-[13-16]----00.0-[14-16]--+-01.0-[15]----00.0  Intel Corporation Battlemage G21 [Arc Pro B60]
 |           |                               |                               |                               \-02.0-[16]----00.0  Intel Corporation Device e2f7
 |           |                               |                               +-10.0-[17-1a]----00.0-[18-1a]--+-01.0-[19]----00.0  Intel Corporation Battlemage G21 [Arc Pro B60]
 |           |                               |                               |                               \-02.0-[1a]----00.0  Intel Corporation Device e2f7
 |           |                               |                               \-18.0-[1b]--
 |           |                               +-08.0-[1c-24]----00.0-[1d-24]--+-00.0-[1e]--
 |           |                               |                               +-08.0-[1f-22]----00.0-[20-22]--+-01.0-[21]----00.0  Intel Corporation Battlemage G21 [Arc Pro B60]
 |           |                               |                               |                               \-02.0-[22]----00.0  Intel Corporation Device e2f7
 |           |                               |                               +-10.0-[23]--
 |           |                               |                               \-18.0-[24]--
 |           |                               +-0c.0-[25-28]----00.0-[26-28]--+-14.0-[27]--
 |           |                               |                               \-15.0-[28]--
 |           |                               \-1c.0-[29]----00.0  Broadcom / LSI PEX880xx PCIe Gen 4 Switch
 |           +-06.1-[2a]----00.0  INNOGRIT Corporation NVMe SSD Controller IG5236 [RainierPC]
 |           +-07.0-[2b-54]--
 |           +-07.1-[55-7e]--

nvtop

 Device 0 [Battlemage G21 (Arc Pro B60)] PCIe GEN 4@ 8x RX: N/A TX: N/A
 GPU 400MHz  MEM N/A MHz  TEMP  37°C FAN    0RPM POW   6 / 400 W
 GPU[                            0%(eff 0%)] MEM[                     0.039Gi/23.906Gi]

 Device 1 [Battlemage G21 (Arc Pro B60)] PCIe GEN 4@ 8x RX: N/A TX: N/A
 GPU 400MHz  MEM N/A MHz  TEMP  36°C FAN    0RPM POW   8 / 400 W
 GPU[                            0%(eff 0%)] MEM[                     0.039Gi/23.906Gi]

 Device 2 [Battlemage G21 (Arc Pro B60)] PCIe GEN 4@ 8x RX: N/A TX: N/A
 GPU 400MHz  MEM N/A MHz  TEMP  34°C FAN    0RPM POW  11 / 400 W
 GPU[                            0%(eff 0%)] MEM[                     0.039Gi/23.906Gi]

 Device 3 [Battlemage G21 (Arc Pro B60)] PCIe GEN 4@ 8x RX: N/A TX: N/A
 GPU 400MHz  MEM N/A MHz  TEMP  36°C FAN    0RPM POW   8 / 400 W
 GPU[                            0%(eff 0%)] MEM[                     0.039Gi/23.906Gi]

 Device 4 [Battlemage G21 (Arc Pro B60)] PCIe GEN 4@ 8x RX: N/A TX: N/A
 GPU 400MHz  MEM N/A MHz  TEMP  37°C FAN    0RPM POW   9 / 400 W
 GPU[                            0%(eff 0%)] MEM[                     0.039Gi/23.906Gi]

 Device 5 [Battlemage G21 (Arc Pro B60)] PCIe GEN 4@ 8x RX: N/A TX: N/A
 GPU 400MHz  MEM N/A MHz  TEMP  38°C FAN    0RPM POW   9 / 400 W
 GPU[                            0%(eff 0%)] MEM[                     0.039Gi/23.906Gi]

P2P isn’t working on this awful intel platform existence

This one was rough, I ended up having to do all the ACS in grub and escape them for some reason, confirm they’re all there after a reboot in /proc/cmdline

GRUB_CMDLINE_LINUX_DEFAULT="pci=disable_acs_redir=0000:07:01.0\\;0000:05:10.0\\;0000:03:00.0\\;0000:15:01.0\\;0000:0f:08.0\\;0000:03:04.0\\;0000:11:01.0\\;0000:0f:00.0\\;0000:0b:01.0\\;0000:05:18.0"

the addresses will be different, easiest way is to fire up VLLM or another backend that’ll attempt P2P and extract the system’s cries from dmesg.

may also not hurt to try intel_iommu=off while you’re at it.

There’s probably a better way to do this than manually adding all the addresses like this, but it worked.

I have a feeling this just works on AMD but I had intel.

tips, tricks, problems and solutions

underclocking and performance

These cards don’t draw a lot of power individually and the 220wish TDP is quite extreme for AI workloads.

Under load/decoding it’s about 100w, 150w doing batches quite aggresssively.

I’ve found reducing maximum clock from 2400mhz to 2000mhz nets around a 15% (80-85w vs 100w) reduction in power for a negligible loss in single request LLM performance, haven’t tested further, should be explored.

idle without working ASPM can be rough, around 35w per card.

With working ASPM and an active process such as nvtop in the backround they’re closer to 10w, without a process they end up closer to 35w again.

cards don’t show up in xpu-smi discovery | dmesg complains about bar space/allocations

GRUB_CMDLINE_LINUX_DEFAULT="pci=realloc=off"

turns out this is an XE quirk that I don’t care to investigate further, likely related to the XE VF passthrough mode, didn’t go away with sr-iov enabled either.

this relates to: a vbios update broke my cards from showing up in xpu-smi/xpumanager.

There’s likely performance implications to this but implications are better than a non-functioning system.

P2P isn’t working AMD edition

This one was rough, I ended up having to do all the ACS in grub and escape them for some reason, confirm they’re all there after a reboot in /proc/cmdline

GRUB_CMDLINE_LINUX_DEFAULT="pci=disable_acs_redir=0000:07:01.0\\;0000:05:10.0\\;0000:03:00.0\\;0000:15:01.0\\;0000:0f:08.0\\;0000:03:04.0\\;0000:11:01.0\\;0000:0f:00.0\\;0000:0b:01.0\\;0000:05:18.0"

the addresses will be different, easiest way is to fire up VLLM or another backend that’ll attempt P2P and extract the system’s cries from dmesg.

There’s probably a better way to do this than manually adding all the addresses like this, but it worked.

I had a feeling this was an intel quirk but it happens on AMD, I likely need to configure the switch itself but it works for now so good enough.

fans don’t update when the cards are idle, temperature doesn’t update while idle

For the last 7 months this has remained a problem for me, I currently run nvtop in a screen as part of my startup jobs, however the most reliable way to update the cards is to simply probe them by running something such as xpu-smi discovery.

your user will require video/render groups for this to work sudoless.

high power idle after a load, cards wedge, errors when waking cards

this one was relatively common sense but I forgot at some point during my platform migration and noticed they were idling at 31w each, fix was enabling

iommu=pt

to stop them getting wedged when P2P is enabled, I also noticed errors reported with this on/full.

another problem I had was any kind of load caused the cards to wedge and never return to a proper idle, “solution” was to disable d3cold

#!/bin/bash
# Find all BDFs for Intel Arc B60 (8086:e211)
GPU_BDFS=$(lspci -d 8086:e211 -D | awk '{print $1}')

if [ -z "$GPU_BDFS" ]; then
  echo "No Intel Arc GPUs (8086:e211) found."
  exit 1
fi

for bdf in $GPU_BDFS; do
  echo "Pinning $bdf out of D3cold..."
  echo 0  | sudo tee /sys/bus/pci/devices/$bdf/d3cold_allowed
  echo on | sudo tee /sys/bus/pci/devices/$bdf/power/control
done
2 Likes

LLMs

splitting this post out to save myself heartache in the future when I update this

VLLM

Currently this is all I use and all I know of that will run multiple cards well, aside from llm-scaler which has its own problems listed below.

Six months ago VLLM was in a sorry state, however with the vllm-xpu-kernels migration that started earlier this year and became fruitful around May-June, the experience has been fantastic and it’s hard to recommend anything else.

base config used throughout this section:

export OMP_NUM_THREADS=4 #specific to my CPU
export VLLM_WORKER_MULTIPROC_METHOD=spawn
export VLLM_XPU_ENABLE_XPU_GRAPH=0 #still broken as of 11th July 2026.
#P2P - fixes oddly aggressive allocations/OOM crashes
export PYTORCH_ALLOC_CONF=expandable_segments:True
vllm serve google/gemma-4-31B-it \
  --gpu-memory-utilization 0.90 \
  --max_num_seqs=16 \
  --max-num-batched-tokens=8192 \
  --port 8000 \
  --host 0.0.0.0  \
  --max-model-len=262144 \
  --block-size 64 \
  -tp=4 \
  --quantization fp8 \
  --enable-auto-tool-choice \
  --tool-call-parser gemma4 \
  --reasoning-parser gemma4 \
  --limit-mm-per-prompt '{"image":3,"video":1,"audio":0}' \
  --enable-prefix-caching

Vague performance results

I only really run Gemma-4-31B with the above config, minus prefix-caching for benchmarks.

I’d guesstimate performance at 1750~2000 PP and 30~35 TG for single requests with P2P.

P2P is around 65~70% faster on prefill and TTFT than without.

Here’s some loose benchmarks

Benchmarks - P2P P2P 512:
============ Serving Benchmark Result ============
Successful requests:                     10
Failed requests:                         0
Maximum request concurrency:             1
Benchmark duration (s):                  78.31
Total input tokens:                      5120
Total generated tokens:                  2560
Request throughput (req/s):              0.13
Output token throughput (tok/s):         32.69
Peak output token throughput (tok/s):    35.00
Peak concurrent requests:                2.00
Total token throughput (tok/s):          98.07
---------------Time to First Token----------------
Mean TTFT (ms):                          262.60
Median TTFT (ms):                        247.62
P99 TTFT (ms):                           469.86
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms):                          29.68
Median TPOT (ms):                        29.68
P99 TPOT (ms):                           29.69
---------------Inter-token Latency----------------
Mean ITL (ms):                           29.68
Median ITL (ms):                         29.69
P99 ITL (ms):                            30.44
==================================================

P2P 8K:

============ Serving Benchmark Result ============
Successful requests:                     10
Failed requests:                         0
Maximum request concurrency:             1
Benchmark duration (s):                  106.76
Total input tokens:                      81920
Total generated tokens:                  2560
Request throughput (req/s):              0.09
Output token throughput (tok/s):         23.98
Peak output token throughput (tok/s):    34.00
Peak concurrent requests:                2.00
Total token throughput (tok/s):          791.28
---------------Time to First Token----------------
Mean TTFT (ms):                          3132.23
Median TTFT (ms):                        3459.50
P99 TTFT (ms):                           3563.09
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms):                          29.58
Median TPOT (ms):                        29.59
P99 TPOT (ms):                           29.59
---------------Inter-token Latency----------------
Mean ITL (ms):                           29.58
Median ITL (ms):                         29.61
P99 ITL (ms):                            30.03
==================================================
Benchmarks - Without P2P NO-P2P 512:
============ Serving Benchmark Result ============
Successful requests:                     10        
Failed requests:                         0         
Maximum request concurrency:             1         
Benchmark duration (s):                  88.95     
Total input tokens:                      5120      
Total generated tokens:                  2560      
Request throughput (req/s):              0.11      
Output token throughput (tok/s):         28.78     
Peak output token throughput (tok/s):    32.00     
Peak concurrent requests:                2.00      
Total token throughput (tok/s):          86.34     
---------------Time to First Token----------------
Mean TTFT (ms):                          653.25    
Median TTFT (ms):                        716.49    
P99 TTFT (ms):                           719.76    
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms):                          32.32     
Median TPOT (ms):                        32.32     
P99 TPOT (ms):                           32.35     
---------------Inter-token Latency----------------
Mean ITL (ms):                           32.32     
Median ITL (ms):                         32.32     
P99 ITL (ms):                            33.31     
==================================================

NO-P2P 8K:

============ Serving Benchmark Result ============
Successful requests:                     10        
Failed requests:                         0         
Maximum request concurrency:             1         
Benchmark duration (s):                  183.28    
Total input tokens:                      81920     
Total generated tokens:                  2560      
Request throughput (req/s):              0.05      
Output token throughput (tok/s):         13.97     
Peak output token throughput (tok/s):    32.00     
Peak concurrent requests:                2.00      
Total token throughput (tok/s):          460.94    
---------------Time to First Token----------------
Mean TTFT (ms):                          10139.18  
Median TTFT (ms):                        11057.92  
P99 TTFT (ms):                           11701.94  
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms):                          32.11     
Median TPOT (ms):                        32.11     
P99 TPOT (ms):                           32.15     
---------------Inter-token Latency----------------
Mean ITL (ms):                           32.11     
Median ITL (ms):                         32.12     
P99 ITL (ms):                            32.72     
==================================================

My biggest enemy has been the Triton attention backend and it’s horrendous performance in general for intel.

after crying on github the developers have graciously allowed us to run --attention-backend FLASH_ATTN overriding triton and other defaults even on MM models.

Triton vs platform default

vllm bench with 8k tokens in, 256 tokens out, 5 iterations, 1 warmup, single concurrency.

Metric TRITON_ATTN FLASH_ATTN P2P (FLASH_ATTN)
Output token throughput 2.21 t/s 8.33 t/s 23.98 t/s
Total token throughput 143.48 t/s 541.20 t/s 791.28 t/s
Peak output throughput 11.00 t/s 30.00 t/s 34.00 t/s
TTFT 46,165.38 ms 11,087.75 ms 3,132.23 ms
TPOT, ITL 93.10 ms/t 33.74 ms/t 29.58 ms/t

Please keep in mind I cherry picked the worst scenario for this comparison and the comparison was run before I upgraded the host to working P2P.

I’ve added P2P results to this table after the fact to demonstrate the general uplift, I reran FLASH_ATTN without P2P on the new hardware topology and results were within 5% of original results so I’ve left them as is.

I didn’t waste my time rerunning Triton with P2P, however, I believe the P2P latency improvements would help mask how awful it is and potentially bring it closer to the FA P2P results.

intel/llm-scaler

It’s the official place to start, especially if you really want to use Docker and it already supports the model you want to run.

main benefit here is the dynamic online quantization specific to llm-scaler such as int4/sym_int4

intel provide a few goodies with their custom ESIMD package and other patches, which can be peeped on via intel/llm-scaler/vllm/{patches,custom-esimd-kernels-vllm}

however post June 2026 this backend is a lot harder to sell with how well VLLM upstream support is going for the intel guys and how behind llm-scaler are by comparison, new models come fast and so does the technology, it’s a third class citizen by comparison.

before July 2026 llm-scaler was running VLLM v0.14.0 with today marking the first major upgrade to v0.21.0

The lag in upstream versioning is rough, for context:

March 31st, 2026 - Gemma4 released
April 3rd, 2026 - VLLM released model support
July 11th, 2026 - llm-scaler released model support

it’s my opinion that llm-scaler still exists, for now, purely as a platform evaluation tool for prospective customers and beginners, and it should be treated as such.

llama.cpp (SYCL & Vulkan & OpenVino)

This should be the easiest to use for a single card however performance dramatically falls off when you try and use multiple cards.

if you can’t fit the model in a single intel card, you should not be using this.

There’s a good chance you’re coming in with experience of llama.cpp from Nvidia/AMD, throw that perception away, it’s not good for intel.

tips, tricks, problems and solutions

P2P OOMs, P2P crashes with device lost, the system OOM kills my VLLM processes

export PYTORCH_ALLOC_CONF=expandable_segments:True

This one little line cut down my allocations for gemma-4-31B-it from 32gb to maybe 9gb at runtime, I have a feeling this is a VLLM quirk with MM and there might be implications to lazy/under allocating like this but it hasn’t crashed yet and I’ve tried to make it crash.

if they weren’t 24gb cards and had more VRAM I probably never would have noticed this.

VLLM v0.27.0 or newer crash during warmup or first request, GuC reset, hangs after large allreduce/allgather:

When I upgraded from v0.26.0 to v0.27.0 P2P errored out and the XE drivers hung in dmesg, turns out the bump to torch 2.13 brought oneccl >=2022.0 which has some quirks/edge cases for weird system topologies, below fixes worked out for me.

export CCL_SYCL_ALLREDUCE_TMP_BUF=1
export CCL_SYCL_ALLGATHERV_TMP_BUF=1
or
export CCL_SYCL_ALLGATHERV_SIMPLE_THRESHOLD=1073741824
export CCL_SYCL_ALLREDUCE_SIMPLE_THRESHOLD=1073741824

dumping VLLM output somewhere I can tail

this is the bare minimum but you get the idea

VLLM_CONFIGURE_LOGGING=1 \
VLLM_LOGGING_CONFIG_PATH=/path/to/logger.json
{
  "formatters": {
    "json": {
      "class": "pythonjsonlogger.jsonlogger.JsonFormatter"
    }
  },
  "handlers": {
    "file": {
      "class": "logging.FileHandler",
      "formatter": "json",
      "level": "INFO",
      "filename": "/tmp/vllm.log"
    }
  },
  "loggers": {
    "vllm": {
      "handlers": ["file"],
      "level": "INFO",
      "propagate": false
    }
  },
  "version": 1
}

bonus script to try and update the fans when they may be idle.

#!/bin/bash

LOG_FILE="/tmp/vllm.log"
SEARCH_STRING="Avg prompt throughput"
COMMAND_TO_RUN="sleep 15 && xpu-smi discovery > /dev/null"

tail -F "$LOG_FILE" | while IFS= read -r line; do
    if [[ "$line" == *"$SEARCH_STRING"* ]]; then
        eval "$COMMAND_TO_RUN"
    fi
done

2 Likes

reserved

(video, image generation coming soon, still benchmarking vllm-scaler-omni against comfyui main)

Out of curiosity, are you still primarily using vllm over vllm-scaler on these cards?

Yeah, in fact between then and now it has gotten even better

Sometime around V26 they started building xpu docker images from main, nightly or releases for users less technical/easier testing adoption

vllm-xpu-kernels has recently had a couple big PRs that’ll speed up gemma4 and another patch is nearly landed (will be in 0.14.0) for GDN/MTP qwen 3_5 (includes 3.8 27b)

2 Likes