A little preamble, I have two R9700’s on an AM4 board with a 3900X both of which are hung off the CPU now at x8x8 Gen 4. For layer split it is not as important to be CPU direct but for Tensor parallelism as many people on here have stated its pretty important. The thing I have been struggling with is getting this all to behave inside a VM on Proxmox. below is the some technical details behind it. Now I will admit I am not an expert in this but I am sharing my experience getting Qwen3.8 27B (Unsloth’s UD-Q8_K_XL quant) to run without going insane with CPU usage and limiting host memory footprint.
The resulting setup.
Host: Proxmox VE 9
Guest: Ubuntu 26.04
GPUs: 2x Radeon AI PRO R9700 32 GB
Model: Qwen3.8 27B UD-Q8_K_XL GGUF
Backend: ROCm
Split: tensor 1:1
P2P copy: ~13 GiB/s each direction
Non-MTP: ~31 t/s single stream
MTP n=3: ~49-50 t/s single stream
P3: ~19-25 t/s/request with 3 simultaneous streams
The problems I needed to solve as I found them.
- Pass through only the gpu functions in proxmox, if you pass through the entire device it flagged it without atomic operations being supported inside the guest and as this was for AI I didn’t need the audio anyway. If you look in dmesg and get this will fix that.
amdgpu … PCIE atomic ops is not supported - Next thing I figured out was that the BAR mapping caused it to get mapped too high I believe (again not an expert, ai was involved here) which caused P2P not to be possible as it was mapped at something like ~0x380000000000 and DMA can apparently only reach below 0x100000000000? Again cannot caveat this enough, deeper than I typically wonder into the inner workings. Setting
cpu: host,guest-phys-bits=44in the proxmox config which puts the BAR low enough that P2P started working.
At this pointcat /sys/module/amdgpu/parameters/pcie_p2preturned Y. - This gave me better performance already but still was hard on the CPU and host memory. So the next step was taking advantage of a direct-P2P AllReduce patch by JohnTDI-cpu on github, I was able to replicate the stated results for this after building llama.cpp with this patch and got me finally to where I am at today.
With the patch:
PP512: 871.65 t/s
TG128: 31.42 t/s
Without the patch:
PP512: 934.58 t/s
TG128: 28.58 t/s
This hurt prompt-processing a bit but decode has a measured improvement which for me is preferable.
Some other things that became rather troublesome along the way. Enabling mmproj in any one of Qwen 35B or 27B 3.6 or 3.8 models seems to give you a random chance of the repeating ////////’s of annoyance as I like to call them. This is something I experienced across atleast 6 different quants of different models from different people including the originals and atleast another 6 different builds of llama.cpp from lemonade, unsloth, custom, release builds, vulkan and rocm backends. The only thing that seemed to stop it so far(fingers crossed) is disabling mmproj.
I am not an expert, there is likely some things in here wrong or incomplete and I am willing to tinker some more and learn. I don’t do this stuff for work, I am purely doing it for the fun. I’ve done a lot more reading here than posting and let me tell you this place has been an invaluable resource.