Intel Xeon Max 9480

Hi @monstercameron, were you able to fine-tune your setup any better? If so, could you please please please post some more benchmark numbers? And do you have any interest with trying out Linux with the 9480? Windows is known to possibly not play well with large core counts. Also, have you tried playing around with HBM-only mode?

The 9480 in hbm-only mode could potentially be a good host CPU to offload some layers off-onto for performing inference on larger models where even 2-4 top gaming GPUs fall short on the vram front.

@wendell any interest in this? I’m seeing some listings for this cpu on ebay for around 2500 usd.

must…resist…temptation…

patrick @ sth was going to send me a couple to play with but it only works in pretty specific boards, too. at least the microcode versions out there vary a fair bit and some are better than others. he’s much more read in on that particular thing than I am

1 Like

Also FYI these have Tcase set to just 64C - even beefy cooler without forced chilled air might be not be enough.

how mutch memory for how many threads ?

They were selling for 900USD brand new in trays not long ago, 2500USD is the scalper pricing.

Linux shows better memory bandwidth over Windows, almost 2TB/s for stream triad-ish loads using AVX512 in HBM-only mode:


I’d imagine/hope that number could go up another 400-500GB/s in flat mode; I’m even more curious what happens to the latencies in flat mode.

The ASUS board that can run them has some very weird quirks like no PCIe bifurcation, no SNC4 mode and NVMe VROC is only applicable to the MCIO ports instead of the PCIe slots.

I’m about to swap out my Dynatron S6’s for a moderately goofy water cooling setup. The Dynatrons barely kept temps below 70C (die temp, not case) using 20C air, but this is with the CPUs running above their 350w TDP.

Each CPU has 64GB of HBM, and has 56 cores, hyper threading isn’t going to be used for the kinds of workloads these CPUs are for.

2 Likes

Any idea as to why? Its very unusual to see this kind of hardware let go this cheap way before any refresh cycle.

It top end part less than two years old !

If you have one could you please run some llama.cpp benchmarks if you’re into that stuff (totally cool if not)? HBM2e+AMX should be a winner but on openbechmark the only 9480 score is 2-3 token/s for TWO of them with the llama2-7Bq4 model, which is so comically bad/off that it honestly feels like misinformation…

Ill respond to the rest of this thread later, just wanted to say I built this system speficially to test out amx perf boosts, bitnet but with a high bandwidth cpu and lavapipe on ubuntu (CPU graphics)

I do get like decent TPS but a lowend Nvidia would unfortunately still beat this machine :frowning:

AFAIKS there is no real way to know if the code is taking the AMX path as I dont see any perf increases. Here were my build options so one of the amx flags would even show up

cmake -B build -DGGML_NATIVE=ON -DGGML_AMX_TILE=ON -DGGML_AMX_INT8=ON -DGGML_AMX_BF16=ON

llama-cli -m C:\Users\Cam\.cache\lm-studio\models\NousResearch\Hermes-3-Llama-3.1-8B-GGUF\Hermes-3-Llama-3.1-8B.Q4_K_M.gguf -p "I believe the meaning of life is" -n 128 -t 112

Another note about scaling:
-t 24 → 8.59 tokens per second
-t 56 → 10.35 tokens per second
-t 112 → 11.30 tokens per second

I think I was getting 20Tk/s when the cores would turbo and I limited the run to 24 cores.

These parts are locked I think but if I could get an all core 3.5GHz clock and fully populated DDR5, I reckon I could squeeze out more perf.

1 Like

Apparently HBM-only is better for perf than HBM+DDR5

But thanks for replying and the insight.

1 Like

From someone on reddit, looks pretty good if real

llama-cli -m ~/Llama-3.1-8B-Instruct-Q8_0.gguf -p ā€œI believe the meaning of life isā€ -n 128 --numa distribute --device none -fa -t 110

llama_perf_sampler_print: sampling time = 17.70 ms / 136 runs ( 0.13 ms per token, 7683.62 tokens per second)

llama_perf_context_print: load time = 665.57 ms

llama_perf_context_print: prompt eval time = 63.69 ms / 8 tokens ( 7.96 ms per token, 125.62 tokens per second)

llama_perf_context_print: eval time = 2774.31 ms / 127 runs ( 21.84 ms per token, 45.78 tokens per second)

llama_perf_context_print: total time = 2885.02 ms / 135 tokens


llama-cli -m ~/Meta-Llama-3.1-8B-Instruct-Q4_0_8_8.gguf -p ā€œI believe the meaning of life isā€ -n 128 --numa distribute --device none -fa -t 110

llama_perf_sampler_print: sampling time = 17.53 ms / 136 runs ( 0.13 ms per token, 7758.57 tokens per second)

llama_perf_context_print: load time = 614.69 ms

llama_perf_context_print: prompt eval time = 39.34 ms / 8 tokens ( 4.92 ms per token, 203.36 tokens per second)

llama_perf_context_print: eval time = 1991.15 ms / 127 runs ( 15.68 ms per token, 63.78 tokens per second)

llama_perf_context_print: total time = 2078.71 ms / 135 tokens

Some more info from the gentleman who was able to have the 9480 work decently well with llama.cpp-

Yep. llama.cpp version 4200 (46c69e0e) It used to be pretty slow compared to OpenVINO and Intel converted models, but llama.cpp has improved recently.

It gets ~9.6 tok/s on Llama-3.3-70B-Instruct-Q4_0_8_8.gguf.

I’m not sure why, but I have a suspicion it had something to do with the aurora supercomputer.

Yeah I can do that. Got any good guides or resources for it? The most I’ve done with local LLMs up to this point is run chatRTX.
I imagine it might take some NUMA tweaking to get it running best; I’m getting the best performance out of my system by forcing applications to run in 8 NUMA nodes (a NUMA node for each tile).

I believe the AMX support is there already their ā€œnormalā€ x86 support.
Here’s if you wanna build it-

Here’s if you want to get a precompiled binary-

As for models…
some recommendations-

  1. Qwen 2.5 Coder 32B (top performing open-weight model for coding)
    bartowski/Qwen2.5-Coder-32B-Instruct-GGUF Ā· Hugging Face
    Recommended quantizations- Any of the Q5/Q6 ones (K_M/K_L)

  2. Another interesting model is QwQ 32B, again from Qwen, but this essentially is like OpenAI’s o1 ā€œthinkerā€ model, but obv probably smaller and is a proof of concept.
    bartowski/QwQ-32B-Preview-GGUF Ā· Hugging Face

  3. Not really going to recommend any ā€œsmallā€ models, because too easy tbh.
    So if you have a gpu or two, maybe it makes more sense to see how the 9480 performs when it’s used to support gpus by offloading some layers off of inference for larger models-
    bartowski/Llama-3.1-Nemotron-70B-Instruct-HF-GGUF Ā· Hugging Face

There’s another special thing that would be interesting called speculative decoding, which essentially uses a smaller model (of the same series) in conjunction with the main model to essentially speed up inference without losing quality. Is useful for bandwidth-limited scenarios, the article has more details below and some instructions on how to use it-

Here’s the draft model to use with Qwen2.5 Coder 34B- bartowski/Qwen2.5-Coder-0.5B-GGUF Ā· Hugging Face

And with the LLaMa one- bartowski/Llama-3.2-1B-Instruct-GGUF Ā· Hugging Face

All of this in HBM-only mode only, I think in ddr5+hbm mode it’ll be slow.

2 Likes

Im on vacation until next year, I’ll ping you when I get back and I can que up all these benchmarks for you.

1 Like

New update for y’all. Unrelated to the Intel portion, but there’s a new open-weight LLM from Deepseek called the Deepseek V3, which matches or exceeds GPT-4o or Claude 3.5 Sonnet on seemingly all types of tasks :open_mouth:

What’s the catch? Well, it’s 671B parameters but it’s a mixture of experts model, so while you need gobs of VRAM, inference only activates 37B parameters. Hence the Xeon Max or any other modern server CPU should comfortably give very decent results as long as you have enough RAM. (Well, that’s the prevailing theory)

llama.cpp support hasn’t landed yet, and we’ve yet to see quantized versions of this model, so I wouldn’t be nervous for needing ~768GB-1TB of RAM to run the unquantized version.

I hope you guys realize…this is state of the art level performance in all aspects that can run on your own server right now! Didn’t see this coming, what a great time to be alive.

@wendell any interest in this one? I know you have the Epyc 9005s and enough RAM to make a small datacenter shy :face_with_peeking_eye:

I need to buy more RAM for deepseek. I also still need to performance tune this Xeon.

1 Like

are you running Windows 11 pro on a Xeon Max 9480???

I discovered that hwinfo64 is giving bad numbers on power usage for these CPUs, mine also reports incorrect low power consumption figures when I’m in Windows.

Turns out I’m wrong on this, there’s no option to enable SNC4, but it is on by default. Its possible the ā€œWorkloadā€ setting that you choose in BIOS automatically applies the SNC setting and just doesn’t tell you though:

I’m not sure how much performance tuning you’ll be able to do on Windows, the only way I was able to affect the scheduling of what cores were running code on mine, which then dictates the frequency it runs at, was to go into BIOS and adjust the CPU bitmap.

I’ve been playing with numactl configs (which is a linux thing) and it’s been really powerful at tuning how the processor runs, I can make specific cores access specific HBM stacks.

yeah, I had an install from a gen 2 Epyc chip that just booted and I only booted into an ubuntu livecd to poke around.

1 Like

Im ready to move off windows but I found the network stack to be buggy in ubuntu. Should I be using latest Ubuntu or some other distro in your experience? While I have been using Ubuntu since gnome 2 days, I’ve never done an Arch level distro.

1 Like