Multicore CPU, vs Multiple CPU, vs Cluster

"As a long-time COMSOL user and COMSOL Certified Consultant who frequently deals with very large models, I’ve been perpetually testing the fastest way to solve problems that are easily parallelizable.

It seems like the landscape is always shifting:
Years ago, multi-socket systems were the only way to get high core counts.

Now, we have insanely high core counts on a single CPU, and almost unfathomable counts on clusters.

I personally have found a single high-core-count CPU to be faster than multiple sockets with the same total core count due to data transfer latency and overhead.

Given this, I’m curious about the community’s experience with clusters. Specifically, what type of models (e.g., highly coupled physics, simple structural analysis, large mesh size) run well enough on a cluster to overcome the networking communication overhead required to compile data and handle the serial steps between each parallelized phase?

Any concrete examples of scaling performance would be greatly appreciated!"

2 Likes

Hey long time no see!

The only time I’ve found clusters useful is during batch sweeps where the simulations running on the different computers are isolated from each other and don’t have any MPI going on.

I’m actually seeing alot of performance scaling problems in just a dual socket machine in most of my problems, so much so that I can often get better performance by “turning off” one of the CPUs via numactl to get better performance executing on only 1 CPU; this is with highly coupled MEF-Turbulent Flow physics setup in the high tens of millions of DoF on a segregated solver.

​​​ ​ ​

​​​ ​ ​
6.4’s support of GPU compute on all the physics modules will be interesting and hopefully reduce the need for traditional cluster compute… Although I’m skeptical of the actual speedup cuDSS will provide (especially for problems that don’t fit in GPU memory).

7.x is supposed to get multi-GPU solver support which might help with the GPU memory issue.

1 Like

I have been experimenting with 6.4 starting today. I have found that cuDSS on my RTX A4000 is just about as fast as my TR-7995 CPU. The models do eat up the video RAM fast but I did notice that the option for multiple GPUs is already there in the solver settings. I ordered an RTX 6000 Blackwell to see how that goes but it does run in hybrid mode fairly quickly using system RAM.

I will let you know if the RTX 6000 works well.

1 Like

!!!

is GPU memory usage similar to what CPU memory usage would be for the same problem? I suppose this would depend on the solver settings somewhat.

CuDSS should have support for both single and double precision, if you can switch between the two and take advantage of FP32, the RTX 6000 is basically the fastest card money can buy.
IIRC 6.3 was locked to FP32 when GPU solving on the handful of physics 6.3 was limited to and you didn’t even have a choice to enable FP64.

I think this might only be for the pressure acoustics interface, atleast if the release notes are to be believed

I ordered the RTX 6000 Pro Blackwell 4 days ago and paid for overnight shipping but alas no card yet… COMSOL does allow me to choose between single and double precision and the RTX 6000 Pro has great FP32 performance but still very limited FP64 performance, bad enough that I may need to look at how to get a A100 GPU if the double precision FP64 is ultimately required for my simulations. My charged particle simulation are not large sparse systems of linear equations but ODEs which cuDSS is not great for. I have confirmed that I can use a cuDSS solver for the electric field space charge calculation after each charged particle time step which may save time.

The GPU memory using the cuDSS looks to be less than the memory required for PARDISO, which is my go-to solver on my system.

I will keep you updated on what I can find. If anyone out there has a A100 I can borrow please let me know as I would love to test it with COMSOL.

1 Like

That works until the socket core count outstrips the socket memory bandwidth for that application code and case. Cache sizes complicate things further still. For instance I use OpenFOAM, cases are 3.2Miilion cells and 40 cores of Xeon 6138 Gold across two sockets and now on 32 cores effectively single socket AMD Epyc 7532 The Xeon jobs took 55 minutes and the Epyc 26 minutes. Here the L3 cache of the Epyc at 8MB per core vs the Xeon 1.2MB per core is playing its part as well as the 50% extra memory bandwidth.