My home machine has been giving me some issues and I need some help figuring out what the problem is. (i7 265K, 96GB Corsair 6400 DDR5, RTX3090 FE)
My day-to-day work is Blender & Houdini and that’s where I’m noticing issues.
Lots of crashes when rendering and the machine tends to become unstable when RAM usage goes up.
I had a feeling it might be a bad stick of memory so I’ve run memtest overnight but after 3 full test runs everything still passes.
When I render the same scenes on a different machine at the office (AMD 7900, RTX5080) I have zero issues, both machines are running up-to-date Arch with the same drivers. Blender has always been pretty stable for me and the problem persists between versions, so I feel like that’s not the root of the cause.
So I’m quite certain it’s a hardware issue on the Intel machine. My question is how do I figure out the next steps of testing. I want to try and test the memory controller on the Intel CPU and the memory on the RTX 3090 FE, although that GPU has been running as a render machine for about 5 years and I haven’t had any issues until the surrounding hardware changed. (It was in a Threadripper machine before)
One thing I did notice is that when I updated my bios recently the memory training for the ram’s XMP profile failed, so I rolled it back to a previous version and it was fine. Could it be passing the training but sill be too close for comfort? I’d think memtest would catch it?
Any other suggestions on how to (and what to) further diagnose would be welcome.
I’m sure their dev’s will be happy to assist getting to the core (see what i did ) of the issue. May cost you an ARM and a leg, given you’re apparently a commercial enterprise.
If you genuinely believe the issue is ram, and memtest isnt showing anything, then you should try to memtest for longer. As a sanity check you could also apply the XMP and pull maybe 400mhz out of the clocks. If theres no improvement then its not ram.
I’ve also run in to problems compiling software or encoding videos with ffmpeg to a ramdisk (long story). With the compiling I can even resume and it just keeps going and then breaks at a different point in the compilation process, which is odd.
Who knows, the Core Ultras might have the same degradation as the previous generations. Should’ve known the deal was too good to be true
You can use mprime to run stress tests of CPU and/or RAM. (No account is needed if you only want to run the stress tests.) This is supposedly much more taxing on the RAM than memtest86.
Then I ran a Prime95 CPU stress test and noticed something funky…
So yeah the CPU apparently just shot up to max temp every time it was loaded. I hadn’t checked temps since I put the machine together. After reseating the air cooler and repasting the problem was still there albeit taking a little longer to reach the max temperature.
I’ve now installed a beefy AIO and under load 2 cores reach slightly over 90C in synthetic stress tests while the rest stay below that and everything seems stable for now.
Looks like the CPU was just happily cooking itself and I had no idea. Fingers crossed this solved the issue. Thanks for all your suggestions on what to test! If the problem returns in the near future I’ll add to this post if necessary.