Dual R9700 llama.cpp Rocm tensor-split, working P2P on Proxmox, Qwen3.8-27B

A little preamble, I have two R9700’s on an AM4 board with a 3900X both of which are hung off the CPU now at x8x8 Gen 4. For layer split it is not as important to be CPU direct but for Tensor parallelism as many people on here have stated its pretty important. The thing I have been struggling with is getting this all to behave inside a VM on Proxmox. below is the some technical details behind it. Now I will admit I am not an expert in this but I am sharing my experience getting Qwen3.8 27B (Unsloth’s UD-Q8_K_XL quant) to run without going insane with CPU usage and limiting host memory footprint.

The resulting setup.
Host: Proxmox VE 9
Guest: Ubuntu 26.04
GPUs: 2x Radeon AI PRO R9700 32 GB
Model: Qwen3.8 27B UD-Q8_K_XL GGUF
Backend: ROCm
Split: tensor 1:1
P2P copy: ~13 GiB/s each direction
Non-MTP: ~31 t/s single stream
MTP n=3: ~49-50 t/s single stream
P3: ~19-25 t/s/request with 3 simultaneous streams

The problems I needed to solve as I found them.

  1. Pass through only the gpu functions in proxmox, if you pass through the entire device it flagged it without atomic operations being supported inside the guest and as this was for AI I didn’t need the audio anyway. If you look in dmesg and get this will fix that.
    amdgpu … PCIE atomic ops is not supported
  2. Next thing I figured out was that the BAR mapping caused it to get mapped too high I believe (again not an expert, ai was involved here) which caused P2P not to be possible as it was mapped at something like ~0x380000000000 and DMA can apparently only reach below 0x100000000000? Again cannot caveat this enough, deeper than I typically wonder into the inner workings. Setting cpu: host,guest-phys-bits=44 in the proxmox config which puts the BAR low enough that P2P started working.
    At this point cat /sys/module/amdgpu/parameters/pcie_p2p returned Y.
  3. This gave me better performance already but still was hard on the CPU and host memory. So the next step was taking advantage of a direct-P2P AllReduce patch by JohnTDI-cpu on github, I was able to replicate the stated results for this after building llama.cpp with this patch and got me finally to where I am at today.
    With the patch:
    PP512: 871.65 t/s
    TG128: 31.42 t/s
    Without the patch:
    PP512: 934.58 t/s
    TG128: 28.58 t/s
    This hurt prompt-processing a bit but decode has a measured improvement which for me is preferable.

Some other things that became rather troublesome along the way. Enabling mmproj in any one of Qwen 35B or 27B 3.6 or 3.8 models seems to give you a random chance of the repeating ////////’s of annoyance as I like to call them. This is something I experienced across atleast 6 different quants of different models from different people including the originals and atleast another 6 different builds of llama.cpp from lemonade, unsloth, custom, release builds, vulkan and rocm backends. The only thing that seemed to stop it so far(fingers crossed) is disabling mmproj.

I am not an expert, there is likely some things in here wrong or incomplete and I am willing to tinker some more and learn. I don’t do this stuff for work, I am purely doing it for the fun. I’ve done a lot more reading here than posting and let me tell you this place has been an invaluable resource.

1 Like

Would like to give shout out to StillDeadcode/vllm-radiance Following usage text would be.

It was hard for me to make it run at first. The problem was untill P2P was not working. But then it does it as easy as download model and run:

sudo docker run --rm -it --device /dev/kfd --device /dev/dri --group-add "$(getent group render | cut -d: -f3)" --group-add "$(getent group video | cut -d: -f3)" --shm-size 4g --cap-add SYS_PTRACE --security-opt seccomp=unconfined -v /mnt/pve/raid0:/models:ro -v "$PWD/vllm-cache:/cache" -p 8080:8080 -e HIP_VISIBLE_DEVICES=0,1,2 -e VLLM_ROCM_USE_AITER=1 -e VLLM_ROCM_USE_AITER_UNIFIED_ATTENTION=1 -e VLLM_ROCM_USE_AITER_MHA=0 -e VLLM_ROCM_USE_AITER_MLA=0 -e VLLM_ROCM_USE_AITER_MOE=0 -e VLLM_ROCM_USE_AITER_LINEAR=0 -e VLLM_ROCM_USE_AITER_FP8BMM=0 -e VLLM_ROCM_USE_AITER_FP4BMM=0 -e VLLM_ROCM_USE_AITER_RMSNORM=0 -e NCCL_PROTO=Simple -e RADIANCE_PRESHUFFLE=1 -e RADIANCE_ATTN_TUNE=1 -e RADIANCE_GDN_WMMA=1 -e RADIANCE_VIT_FLASH=1 -e RADIANCE_FAST_REDUCE=1 -e RADIANCE_AR_MAX_KB=32768 -e RADIANCE_AR_QUANT=1 -e RADIANCE_FUSE_RMS_QUANT=1 -e RADIANCE_DYNAMIC_DRAFT=1 -e VLLM_CACHE_ROOT=/cache/vllm -e TORCHINDUCTOR_CACHE_DIR=/cache/inductor -e TRITON_CACHE_DIR=/cache/triton -e AITER_ROOT_DIR=/cache/aiter -e TRITON_CACHE_AUTOTUNING=1 stilldeadcode/vllm-radiance:0.5.7 --model /models/Qwen3.8-27B-FP8 --served-model-name Qwen/Qwen3.8-27B-FP8 --quantization fp8 --kv-cache-dtype fp8 --tensor-parallel-size 2 --gpu-memory-utilization 0.92 --attention-backend ROCM_AITER_UNIFIED_ATTN --enable-prefix-caching --mamba-cache-mode align --speculative-config ‘{“method”:“mtp”,“num_speculative_tokens”:4,“attention_backend”:“ROCM_AITER_UNIFIED_ATTN”,“disable_padded_drafter_batch”:true}’ --no-async-scheduling --host 0.0.0.0 --port 8080 --max-model-len $((1024*256)) --max-num-seqs 2 --enable-auto-tool-choice --tool-call-parser qwen3_coder --reasoning-parser qwen3

mtp works, vision works. no loops so far.

Summary

This text will be hidden

So it’s too early to tell for sure but I ad forgotten that I had my memory overclocked on the gpu’s which seemed stable but after hours of use eventually it might have been the reason for the devolution into //’s because it started happening again even with vision disabled so it was a red herring of sorts. I now seem to have it roughly stable with rocWMMA re-enabled, for better prompt processing speed on tp2 along with the same tg/s rates and mtp settings and such.