Thanks for the great write-up, it’s what finally got me to try this on my own two CMP 170HX.
Same rough setup as yours: 2x CMP 170HX, vLLM W4A16, TP2 + EP, MTP4. The difference is I got BAR1 P2P working between the two cards even though they sit on separate PCIe root complexes, which is normally reported as a dead end (topo shows NODE, topo -p2p shows GNS, cudaDeviceCanAccessPeer returns 0).
Short version: it’s not actually a hardware limit, just a driver-side fix (a registry key, no rebuild). Can post the details if anyone wants them. I did check it properly with a content test: cuMemcpyPeer both directions plus SM remote read/write with different patterns, checksums line up, no Xid or AER.
My results, on 2x CMP 170HX / EPYC / ROMED8-2T, PCIe Gen2 x16, driver 610.57.04, vLLM 0.30.0, W4A16 TP2 + EP, MTP4, with --disable-custom-all-reduce and NCCL_P2P_LEVEL=SYS. Cache-free (prefix caching off, vllm bench serve random dataset), single stream, median of 3 runs:
| Context | vLLM W4A16, no P2P | vLLM W4A16 + P2P |
|---|---|---|
| 8k prefill | 2,476 tok/s | 4,714 |
| 32k prefill | 2,545 | 4,661 |
| 65k prefill | 2,510 | 4,640 |
| 128k prefill | 2,443 | 4,378 |
| decode (256-tok class) | 111.3 | ~120 |
So roughly 1.8 to 1.9x on prefill, about +8% on decode. Lines up with the idea that TP2’s per-token all-reduce is what holds prefill back without P2P (host-staged at ~2 GB/s per card); open the peer path and prompt processing nearly doubles, while decode is bandwidth-bound so it barely moves. The whole cache-free sweep (5 context sizes, 3 reps each, single stream) ran in about 8.5 minutes.
A couple of things worth flagging:
- It’s a cross-setup comparison, not a clean P2P on/off A/B on the same box. Same model, same vLLM config, same PCIe Gen2 x16 though, so P2P is really the only big variable.
- You need --disable-custom-all-reduce. With P2P on, vLLM’s custom IPC all-reduce OOMs the VRAM or throws custom_all_reduce.cuh ‘invalid argument’ across root complexes. PyNCCL still uses P2P fine.
- FP8 KV cache is a no-go here. vLLM takes the --kv-cache-dtype fp8 flag, but Flash-Next’s QWEN4_EXP state-attention backend crashes on load.
Can drop the full raw vllm bench outputs (commands plus complete result blocks, all 3 reps) if anyone wants to reproduce.
BTW, both cards are frequency and power throttled (capped at 150W and 200W) because I don’t have proper cooling on them yet.
