Couldn't resist grabbing a CMP 170HX, and now I'm in a sticky position

Thanks for the great write-up, it’s what finally got me to try this on my own two CMP 170HX.

Same rough setup as yours: 2x CMP 170HX, vLLM W4A16, TP2 + EP, MTP4. The difference is I got BAR1 P2P working between the two cards even though they sit on separate PCIe root complexes, which is normally reported as a dead end (topo shows NODE, topo -p2p shows GNS, cudaDeviceCanAccessPeer returns 0).

Short version: it’s not actually a hardware limit, just a driver-side fix (a registry key, no rebuild). Can post the details if anyone wants them. I did check it properly with a content test: cuMemcpyPeer both directions plus SM remote read/write with different patterns, checksums line up, no Xid or AER.

My results, on 2x CMP 170HX / EPYC / ROMED8-2T, PCIe Gen2 x16, driver 610.57.04, vLLM 0.30.0, W4A16 TP2 + EP, MTP4, with --disable-custom-all-reduce and NCCL_P2P_LEVEL=SYS. Cache-free (prefix caching off, vllm bench serve random dataset), single stream, median of 3 runs:

Context vLLM W4A16, no P2P vLLM W4A16 + P2P
8k prefill 2,476 tok/s 4,714
32k prefill 2,545 4,661
65k prefill 2,510 4,640
128k prefill 2,443 4,378
decode (256-tok class) 111.3 ~120

So roughly 1.8 to 1.9x on prefill, about +8% on decode. Lines up with the idea that TP2’s per-token all-reduce is what holds prefill back without P2P (host-staged at ~2 GB/s per card); open the peer path and prompt processing nearly doubles, while decode is bandwidth-bound so it barely moves. The whole cache-free sweep (5 context sizes, 3 reps each, single stream) ran in about 8.5 minutes.

A couple of things worth flagging:

  • It’s a cross-setup comparison, not a clean P2P on/off A/B on the same box. Same model, same vLLM config, same PCIe Gen2 x16 though, so P2P is really the only big variable.
  • You need --disable-custom-all-reduce. With P2P on, vLLM’s custom IPC all-reduce OOMs the VRAM or throws custom_all_reduce.cuh ‘invalid argument’ across root complexes. PyNCCL still uses P2P fine.
  • FP8 KV cache is a no-go here. vLLM takes the --kv-cache-dtype fp8 flag, but Flash-Next’s QWEN4_EXP state-attention backend crashes on load.

Can drop the full raw vllm bench outputs (commands plus complete result blocks, all 3 reps) if anyone wants to reproduce.

BTW, both cards are frequency and power throttled (capped at 150W and 200W) because I don’t have proper cooling on them yet.

Thanks, this is extremely interesting — and the timing is perfect because my setup has changed quite a bit since my earlier posts.

I can now confirm that I also have real P2P working on the CMP 170HX, not just nvidia-smi topo -p2p reporting support.

My current setup is:

  • 3× CMP 170HX 64 GB (4th card incoming) All gen2 16x behind a PLX Chip

  • NVIDIA 610.43.02

  • Ubuntu 24.04

  • Broadcom/LSI PEX88096 PCIe switch

  • vLLM W4A16

  • TP2 + EP

  • MTP4

  • Qwen3.8-Flash-Next W4A16 INT4PLE

  • 262k context

I spent quite a bit of time chasing BAR1, REBAR, bridge windows and even custom kernels because initially NCCL would deadlock as soon as actual P2P transport was enabled. Host-staged/SHM worked, but real P2P did not.

We eventually got the driver side sorted out and now CUDA P2P is actually functional and survives proper peer-memory tests.

Interestingly, before getting P2P working, my best Qwen numbers were already roughly:

  • ~2,295 tok/s long-context prefill

  • ~109 tok/s median decode

  • ~115 tok/s best decode

  • MTP4 acceptance around 34%

So your ~4.4–4.7k tok/s P2P prefill result is the part that really caught my attention.

It strongly suggests that I haven’t yet extracted the full prefill benefit from the working peer path. Your result gives me a very useful target for a controlled A/B test on my own machine.

I also find your --disable-custom-all-reduce result particularly interesting. We previously had NCCL/P2P deadlocks and IPC-related issues, so forcing PyNCCL while retaining CUDA P2P may actually be the cleaner production configuration for these cards.

Could you post the exact registry key you used?

I’d especially like to compare it with what we’re currently doing in the 610 driver. If the same result can be achieved with a simple persistent RM registry setting instead of maintaining parts of our current CMP unlock/driver modifications, that would simplify things considerably.

It would also be useful if you could post:

nvidia-smi topo -m

nvidia-smi topo -p2p r

nvidia-smi topo -p2p w

and, if possible, the relevant NCCL debug lines showing that NCCL actually selected the P2P path.

I’m currently hardening the setup so P2P and the PCIe configuration remain persistent across driver reloads/reboots/model changes. Once that’s finished I’ll run the same cache-free P2P ON vs OFF prefill sweep on identical hardware.

That should give us the clean A/B comparison that neither of our current datasets has yet.

Also: with the fourth CMP 170HX coming, I want to see how this behaves with TP4 / EP4, particularly for DeepSeek-V4.1-Flash. That’s probably where fixing the peer path will become even more important.

Very nice result — especially the SM remote read/write verification. That’s much stronger evidence than relying on topo -p2p alone.

Four CMPs? Nice. If you found a cheap source for those, I would genuinely love to hear it, asking for a friend (me). :slight_smile:

The key that matters is ForceP2P=0x11, passed through the driver’s NVreg_RegistryDwords:

options nvidia NVreg_RegistryDwords="ForceP2P=0x11"

0x11 is NV_REG_STR_CL_FORCE_P2P with READ bits [1:0]=1 and WRITE bits [5:4]=1. Persist it, update-initramfs -u, then reload the module or reboot. That one key is enough. I also had RMForceStaticBar1, RMPcieP2PType and RMForceP2PType set from earlier attempts, but those turned out to be no-ops for this.

The reason you need it is an init-order bug in the cmpunlocker patch (0011). It sets p2pOverride = bCmp170hx ? 0x11 : NOT_OVERRIDEN in kernel_bif.c _kbifInitRegistryOverrides, but that runs in kbifConstructEngine before pGpu->idInfo.PCIDeviceID is filled in. So devId reads 0, bCmp170hx is false, and p2pOverride stays 0xffffffff, i.e. never set. With no override the driver uses the GSP caps, and those hard-report GPU_NOT_SUPPORTED for the mining CMP, so you land on GNS and cudaDeviceCanAccessPeer=0. The same function also reads the ForceP2P regkey and sets p2pOverride directly, and a module param is available at load time while PCIDeviceID isn’t, so it dodges the ordering problem. No code patch or rebuild.

Topology here is two CMPs on separate EPYC 7443 root ports (01:00.0 and 81:00.0):

$ nvidia-smi topo -m
        GPU0  GPU1  CPU Affinity  NUMA Affinity
GPU0     X    NODE  0-47          0
GPU1    NODE   X    0-47          0
$ nvidia-smi topo -p2p r
        GPU0  GPU1
GPU0     X    OK
GPU1     OK   X
$ nvidia-smi topo -p2p w
        GPU0  GPU1
GPU0     X    OK
GPU1     OK   X

So NODE, and p2p read/write OK both ways once the regkey is in.

I also ran a content test in torch (from the vLLM venv), not just the topo status: can_device_access_peer True both ways, cuMemcpyPeer 0<->1 and 1<->0 with pattern checksums, plus an SM kernel reading peer memory, all correct, no Xid/AER/NVRM. So it is actually moving data across the root complexes, not just reporting OK.

For serving I run vLLM with NCCL_P2P_LEVEL=SYS and --disable-custom-all-reduce. The custom IPC all-reduce OOMs and throws custom_all_reduce.cuh:164 “invalid argument” across root complexes; PYNCCL uses the P2P path without that. FP8 KV cache is a separate dead end on this model (the state-attention backend crashes on load), nothing to do with P2P.

I will follow up with the NCCL_DEBUG=INFO lines that show it picking the P2P path, both cards are busy serving right now so I cannot pull them this minute.

1 Like

Nah I bought 3pc at 1250€ a piece and the fourth at 1900€ https://www.alibaba.com/x/1lBHW4U?ck=pdp

They send fast with FedEx and Vat is Low you can Talk with them :slight_smile: and you can Pay Secure