Nice build and good compromises. There may be a silver lining in your platform choice, especially in the context of four GPUs hanging off one board. I run a 3× Nvidia Blackwell rig on WRX90, and enumerating all cards at Gen 5 is a circus act at best. I hit it from both ends — boot-time enumeration failures with hung boot loops, and dropped workloads once the system was up and under load. Forcing all PCIe slots down to Gen 4 in BIOS solved both completely.
For my workloads the cost is basically a rounding error — maybe a 2–4% hit versus Gen 5. One caveat: I’m sitting on 190+ GB of VRAM across two cards, so my models fit without sharding and the bus barely gets touched. If you’re splitting models across all four R9700s, PCIe traffic goes up and the penalty may be somewhat larger — though llama.cpp’s layer-split hands off sequentially between GPUs rather than syncing constantly, so it’s still fairly light on the bus.
The honest truth is that even on WRX90 you’d likely have ended up negotiating down to Gen 4 for stability anyway, so the Gen 4 hit isn’t really a compromise you made — it’s just where multi-GPU rigs land in practice. You skipped the debugging tax and arrived at the same place. Well done!