Must Go Faster... Enter WarpX

​I’m currently on a quest to simulate large-area industrial magnetron sputtering. Recently, I pivoted from COMSOL Multiphysics to WarpX, which has sent me down the rabbit hole of hardware optimization. I’m chasing higher fidelity simulations with more particles and faster simulation steps per second, but I feel like I might have bitten off more than I can chew.

​Due to the current insane state of the hardware market, I’ve focused my recent obsession on PCIe Gen 4 hardware. Here is my journey so far and the crossroads I am currently at.

​Attempt 1: The FP64 Bottleneck

​I started with a Gigabyte G242-Z11 server outfitted with 2x NVIDIA L40S GPUs.

Unfortunately, I learned a couple of hard lessons very quickly:

  • ​WarpX demands FP64: The L40S has massive FP32 capabilities, but its dismal FP64 performance left me severely bottlenecked.

  • ​Topology limitations: The G242-Z11 accommodates up to 4 GPUs, but they all connect directly to the CPU. There is no direct GPU-to-GPU communication over the PCIe bus, leading to painfully slow data transfer speeds between the cards.

​Attempt 2: The AMD Pivot and Physical Constraints

​Evaluating my options for FP64 compute in the PCIe Gen 4 landscape, I found that the AMD MI210 is my best bet, and WarpX natively supports ROCm.

​However, I immediately hit a physical wall. The G242-Z11 is a 2U server where the GPUs lay flat rather than plugging into standard slots next to each other. This eliminates my ability to use the Infinity Fabric connectors between the MI210s, which I desperately need for the VRAM and bandwidth benefits.

​The Tradeoff: Frankenstein Build vs. High Idle Power

​Running my own business means every decision is a tradeoff, and I am currently torn between two paths:

​Option A: Rerouting to a custom 4U box

Each PCIe 4.0 x16 slot in the G242-Z11 is served by 2x MCIO connectors. I am considering rerouting these connections into a dedicated, smaller 4U enclosure using a PCIe backplane with standard double-wide GPU spacing. This would let me bridge the MI210s with an Infinity connector and reap the performance benefits without buying an entirely new server.

​Option B: The ASUS ESC8000A-E11

I’ve explored just buying a new server, and the ASUS ESC8000A-E11 is the only chassis that comes close to fitting the bill. My primary hesitation here is power and scale. I really don’t want to run dual CPUs and 32 sticks of RAM. I already have the RAM on hand, but I cannot justify sustaining that kind of power load at idle for my current scale.

​Has anyone here attempted an MCIO-to-backplane reroute like this, or have any advice on optimizing an MI210 topology for WarpX without getting killed by dual-socket idle power draw?

​I am all ears. Thanks in advance!

2 Likes

i mean… with the asus option, do not let it idle? suspend and wol might be good

I was thinking of going to the Asus and putting together a separate ultra low power storage and data server that I would use to keep the data for it and my other HPC system on. Right now my G242-z11 is also my bulk storage server as it had plenty of room for a large SAS disk array. This would allow me to spool up the Asus beast when needed and keep it sleeping whenever idle. IPMI makes this easy. The Asus also boost my multiple L40S performance as it allows for direct GPU to GPU communication over the PCIe switch. By the way, running both CUDA and ROCm on the same system isn’t too hard. They don’t work together in any fashion, but they can easily run their own programs at the same time on the same system.

yeah or just ethtool -s enpasdasd wol g

how long is a typical sim ? i am only knowledgeable in bioinfo stuff but if you need f64 it is probably hours

In COMSOL it started off taking me about a month to run a simulation a few years ago, now I can do it in about a week in COMSOL. WarpX allows me to run the same simulation in about a day. Tracking 10s of millions of particles in pico second time steps for 10s of microseconds take a lot…

it probably does, as it would be potentially catastrophic if not, does it have every x amount of steps backups?

I save a subset of data every 0.1 microseconds and dump the rest. No reason to save 40-80 GB of data for thousands of steps when all I need is the telemetery data.

if the data is not sensitive you might what to look at preemptive instances in ie gcloud

This might be a dumb question, but why can’t you compile for FP32?
Or is it that solution stability suffers too much from this?

I spent a lot of time trying various solutions for FP32. The short answers is, as the electrons lose energy in ionization events there just aren’t enough decimal places in FP32 to keep the electrons from snapping to 0eV as a rounding error which turns them into essentially zombie particles that build and build until it eventually crashes the simulation. I couldn’t figure out a way to eliminate these zero energy particles while retaining the fidelity of the plasma and preserving a strict conservation of energy.