Ryzen 5950X
Asus Dark Hero
32G DDR4
RTX 6000 Pro 96G 600W (Power Limited to 300W)
RTX 6000 Pro 96G 600W (Power Limited to 300W)
Corsair HX1200i
Arch Linux
Problem:
After 5-60m, the system locks up, I am unable to see video output from the display port, my ssh connections get disconnected and it stops responding to network activity. If I start the machine, it will load up in a very short time if I load a llm model with vllm or even if I just start the server and do absolutely nothing.
Observations:
This is what bothers me, it is 100% stable while loading a model with sglang and has run for months and continues to run without issue if I immediately load a model with sglang. I have run it for months even under heavy load (95% memory usage and 100% gpu usage) with MiniMax M2.1 and GLM Air, it runs flawless. I only have a problem with sglang is not running, like system sitting idle on fresh reboot or running vllm instead.
Decode above 4G works, the system will not post if resize bar is enabled.
After lots of troubleshooting, I have narrowed it down to dropping the PCIE to GEN3 and the system becomes stable regardless what is running. GEN4 is only stable while a model is loaded w/ sglang, and is unstable doing nothing, running a model with vllm or llama.
I am not using risers. I have confirmed the gpus are running at 8x.
Do you have AER enabled and getting logged for pcie bus errors?
Any ability to capture video output before the hang? Is there a kernel panic or anything?
I do not have AER enabled.
I have JetKVM monitoring the display and have tailed logs over ssh and neither show any symptom before it happens, it just locks up.
I faced similar issues with RTX 5090 on an X670E board.
I would get random mouse lag followed by frozen/black screen after booting into Windows 11 25H2, on a weekly basis. Eventually, the system would reboot.
The issue happened despite forcing the PCIE slot to Gen5/4/3, setting Windows PCIE Link State Power Management to off, clean installing multiple driver versions, and even changing the DisplayPort cable.
With GPU-Z, I noticed the issue occurred when the PCIE slot was downgraded to Gen2, despite forcing it to faster speeds in the BIOS.
Eventually, what resolved the issue (4 weeks so far) was to enable PCIE Spread Spectrum in combination with PCIE Gen4. If it ran at Gen5, the issue would still occur.
Turn AER on, this smells like a PCIE signaling problem and the only way to see it is to report errors. I have had boards/cpus/cards where it would negotiate gen4 or gen5 but when you put enough load on it errors would increment. Most pcie errors are auto-corrected resulting in a small performance hit, but if there is an uncorrectable one that’s what causes the crash.
Will turn that on and see if I can see something. Not sure there is anything I can do if I do see it, as it works fine with GEN3, but GEN4 it crashes pretty rapidly. What really bothers me though I can max out my GPUS load and ram and have zero problems as long as I use sglang, anything else bombs even nothing at all. I thought the gpus were going in some low power state which was causing it as it I am never doing anything when it happened, but system was stable for 3+ months because I always had sglang running.
I’d be curious if you managed to find out the actual PCIE speed your GPU was momentarily running at before it froze, without Spread Spectrum. Not just the PCIE speed specified in the BIOS.
I never had to enable Spread Spectrum for my RTX 3090 on the same motherboard in order for it to run at Gen4 setting in the BIOS.
I have a suspicion Blackwell GPUs have issues running at slower than Gen3 speeds, unlike Ampere, so any EMI which causes it to downgrade to Gen2 would cause a crash.
I am not sure I do know it is always idle when it happens. I’ve tried monitoring logs and other things like nvidia p state but nothing has been helpful in pointing me to the issue.
The fan curve for the rtx pro 6000 workstation cards is not aggressive enough for more then one card to get heat soaked in that time frame. Even in my open air case I will see the top card start throttling due to thermals in the first hour. LACT has been a life saver for me; a little more noise for alot more comfort.