I’m currently on a quest to simulate large-area industrial magnetron sputtering. Recently, I pivoted from COMSOL Multiphysics to WarpX, which has sent me down the rabbit hole of hardware optimization. I’m chasing higher fidelity simulations with more particles and faster simulation steps per second, but I feel like I might have bitten off more than I can chew.
Due to the current insane state of the hardware market, I’ve focused my recent obsession on PCIe Gen 4 hardware. Here is my journey so far and the crossroads I am currently at.
Attempt 1: The FP64 Bottleneck
I started with a Gigabyte G242-Z11 server outfitted with 2x NVIDIA L40S GPUs.
Unfortunately, I learned a couple of hard lessons very quickly:
-
WarpX demands FP64: The L40S has massive FP32 capabilities, but its dismal FP64 performance left me severely bottlenecked.
-
Topology limitations: The G242-Z11 accommodates up to 4 GPUs, but they all connect directly to the CPU. There is no direct GPU-to-GPU communication over the PCIe bus, leading to painfully slow data transfer speeds between the cards.
Attempt 2: The AMD Pivot and Physical Constraints
Evaluating my options for FP64 compute in the PCIe Gen 4 landscape, I found that the AMD MI210 is my best bet, and WarpX natively supports ROCm.
However, I immediately hit a physical wall. The G242-Z11 is a 2U server where the GPUs lay flat rather than plugging into standard slots next to each other. This eliminates my ability to use the Infinity Fabric connectors between the MI210s, which I desperately need for the VRAM and bandwidth benefits.
The Tradeoff: Frankenstein Build vs. High Idle Power
Running my own business means every decision is a tradeoff, and I am currently torn between two paths:
Option A: Rerouting to a custom 4U box
Each PCIe 4.0 x16 slot in the G242-Z11 is served by 2x MCIO connectors. I am considering rerouting these connections into a dedicated, smaller 4U enclosure using a PCIe backplane with standard double-wide GPU spacing. This would let me bridge the MI210s with an Infinity connector and reap the performance benefits without buying an entirely new server.
Option B: The ASUS ESC8000A-E11
I’ve explored just buying a new server, and the ASUS ESC8000A-E11 is the only chassis that comes close to fitting the bill. My primary hesitation here is power and scale. I really don’t want to run dual CPUs and 32 sticks of RAM. I already have the RAM on hand, but I cannot justify sustaining that kind of power load at idle for my current scale.
Has anyone here attempted an MCIO-to-backplane reroute like this, or have any advice on optimizing an MI210 topology for WarpX without getting killed by dual-socket idle power draw?
I am all ears. Thanks in advance!