System started to experience random crashes on Linux on multiple distros forcing hard shutdown

I have a pcie usb card to add some extra ports to the case, I’ll try removing it later to see what happens

1 Like

Better upgrade to Windows and never look back. Hope you can take a little humor since this is what I hear every day the other way around. Hope you get your problems figured out

Yes, I can take it and understand where you are coming from. Even if I do not completely agree with the sentence speaking of compatibility Windows tends to be better. However Linux and Windows both have strengths and weaknesses and none is clearly the superior choice compared to the other, that’s the reason why I have both installed on my machine, for my needs Linux is a better primary OS however sometimes I do need to do stuff that requires Windows

Also for the first time since a very long time the system went into hyprlock, so it might be that the PCI-E USB card was causing the issue and the system has been stable since it was removed. But since the crashes appear even many hours after boot I’ll consider the issue solved only after a 24h stability test

1 Like

It crashed again, same error message plus soon after another one
kernel: mt7921e 0000:08:00.0: Message 00002ced (seq 8) timeout

remove the mediatek hardware.
that really ought to be the first thing in every computer manual ever.

possibly related: Random mt7921e driver crashes leading to complete system freeze / Kernel & Hardware / Arch Linux Forums

Unfortunately it is soldered to the motherboard, I’ll see if there is a way to disable it in the BIOS but if there isn’t I do not think that would be possible

1 Like

blacklist the mt7921e driver then?

I added mt7921e.disable_aspm=1 to GRUB_CMDLINE_LINUX_DEFAULT if that’s what you mean, let’s see if it prevents crashes, thank you for your patience btw, I am pretty technical but on this issue I am really lost

1 Like

mediatek has always been the boogeyman
that aspm trick is worth a try. but what i meant was to blacklist the driver and prevent the module form even loading to begin with.

Ok, I added a file in /etc/modprobe.d blaklisting mt7921e and btmtk, thankfully I do not use wifi or bluetooth on this machine so I can keep using it normally while testing

1 Like

dont forget to regenerate your initial ramdisk after you mess with modprobe configs!

I know, I rebooted right after creating the file

1 Like

Still crashes, now I tried disabling a setting I just found in the bios regarding XHCI, let’s see if it fixes something

1 Like

Nope, still crashes, no new errors in journalctl

also it also crashed in text only mode

1 Like

Tomorrow I’ll create a new partition and run the cachyos installer to make a fresh install with only the basic packages and KDE to check if that’s stable

1 Like

Hello @FC3243D4 and welcome to the Forums.

Sorry that you are encountering these issues. I actually had similar problems to you, even with BIOS crashing. The culprit for me was ASUS memory context restore. However, I do not believe 5000 series compat motherboards have/need this setting. But wouldn’t hurt to see if it exists and try disabling it.

I looked at your above logs and saw that you are using a 3090. I had an IRL friend also have issues with the recent Nvidia driver release on bazzite. You could try downgrading to 590.44.01-1 and seeing if that resolves the error. I’m not 100% how you do that on Cachy/Arch based distros.

1 Like

next reboot I’ll try to see if I find a setting similar in the bios thank you

I tried downgrading to 580 driver (I do not recall the exact release) and it did not fix the issue

Right now the clean install seems pretty stable, it’s been running for 7h but I’ll wait until I see an uptime of at least 24h since the crashes sometimes happened very late (the latest I saw a crash was about 10h). If it’s stable I’ll probably look to coolercontrol since it is the only package I recall installing on all distros (right now it is not installed and I am running the system passively thanks to the very big radiators I have for the CPU and the GPU)

1 Like

It passed the 24h stability “test” at idle

Right after, I installed coolercontrol and started a combinedstress test on the CPU (stress-ng) and GPU (furmark) and it went well for blabout 30/40min

Then I went to add a custom offset sensor for the GPU and a couple of minutes later the whole system crashed and I experienced multiple crashes in BIOS, I’ll try removing the offset sensor to see if that’s the problem