Very nice rig! Welcome to the party!
We have a big thread over here with some more llama.cpp settings and examples. Some folks have had luck with --numa distribute and dialing in the exact number of CPU cores to use by searching with llama-bench --threads $(seq -s, 32 8 64) etc.
I don’t have access to such a rig, but curious what is the output on it in NPS0 with numactl --hardware ? Does it present 2 nodes, one for each socket, or really just one big single node weaved together with AMD magic?
Finally, if you have even a single CUDA GPU with 16GB VRAM or more installed, check into my ktransformers guide as folks are seeing almost 2x inference speeds with ktransformers on big MoE models over latest llama.cpp.
Keep us posted with your results!