Truenas Scale disks errors and unreliable pool

I have a server with TrueNAS Scale 25.10.6 as a Proxmox VM, and I’m having issues with write errors on the disks in the main pool. These almost always occur on weekends, presumably during backups. The strange thing is that if I turn the server off and on again, the disks may not show any errors for months, but when I perform updates and restart the TrueNAS VM or even the Proxmox server, errors can appear on the disks in the pool over the next few days (it usually targets one disk and gives errors only to that one until it disconnects and degrades the pool).

The disks aren’t faulty, and these errors appear randomly, sometimes on one disk, sometimes on another. Since I’m used to these errors, one of these disks failed in the past few months, and I mistook it for the usual problem, cleaning up the ZFS pool and crashing the TrueNAS VM because the disk had been set to read-only mode.

I’ve tried everything, and at this point I no longer know how to achieve a stable and reliable configuration.

The pool consists of four 10TB ST10000NM013G drives in RAIDZ2 with a log cache of two U.2 drives, with the controller passed to the VM by Proxmox (I used to do disk passthrough with the same result.).

My configuration is as follows:

CPU – A-EPYC9004-2G5

Motherboard – MBD-H13SSL-NT latest BIOS

RAM – 128GB

HBA Controller – AOC-S3816L-16IT latest firmware

Drives – RAIDZ2 4x ST10000NM013G Log 2x U.2-NVME-TLC-4TB-ST-G4

The drives do not have any active power management options, either in the BIOS or TrueNAS.

Can anyone point me in the right direction?

this is the truenas vm conf file in proxmox:

agent: 1
balloon: 0
boot: order=scsi0
cores: 12
cpu: host
hostpci0: 0000:01:00.0,pcie=1,rombar=0
hotplug: network,usb
ide2: none,media=cdrom
machine: q35
memory: 40960
meta: creation-qemu=8.0.2,ctime=1690897446
name: TrueNas
net0: virtio=D6:B9:5B:5E:69:6C,bridge=vmbr0,firewall=1
numa: 0
onboot: 1
ostype: l26
scsi0: local-zfs:vm-101-disk-0,discard=on,iothread=1,size=32G,ssd=1
scsi1: /dev/disk/by-id/nvme-eui.00000000000000000026b728302d15e5,backup=0,discard=on,iothread=1,serial=50026B728xxxxx>
scsi2: /dev/disk/by-id/nvme-eui.00000000000000000026b728302d1575,backup=0,discard=on,iothread=1,serial=50026B728yyyyy>
scsihw: virtio-scsi-single
smbios1: uuid=9d2f59b0-b601-42e3-958c-0b8c0d45414d
sockets: 1
startup: order=2,up=100
tags: debian;linux
vmgenid: df12398e-68d7-43eb-8be9-9ed086f452bb

Is the hba receiving enough air so it does not not overheat? Even a normal (non-server) 40mm fan slapped on it is good enough if other fans cant keep it cool

In the Proxmox virtual machine configuration, you can change scsihw from virtio-scsi-single to virtio-scsi-pci.

I don’t see how it could help since the scsihw parameter only manages the virtual controller where the SSDs of the vdev log are connected, which is not the one causing me problems. I did passthrough of the entire HBA controller so TrueNAS has direct access to the hardware.

1 Like

the machine is in a server grade case with good airflow

Do you have another brand/model of controller that you’re sure works perfectly? Otherwise, I’ll throw this controller away and buy a new one.
I usually prefer to throw money at the problem and fix it quickly :slight_smile:

In the kernel logs i found this error:
[Mon Aug 31 02:45:48 2026] sd 2:0:4:0: [sdg] tag#1906 FAILED Result: hostbyte=DID_SOFT_ERROR driverbyte=DRIVER_OK cmd_age=0s
[Mon Aug 31 02:45:48 2026] I/O error, dev sdg, sector 941747712 op 0x1:(WRITE) flags 0x0 phys_seg 15 prio class 2
[Mon Aug 31 02:45:48 2026] zio pool=Main-pool vdev=/dev/disk/by-partuuid/2b142e8b-3136-11ee-a874-992e5664a07f error=5 type=2 offset=480027279360 size=106496 flags=2148533376

Have you checked temperatures? I’d install lm-sensors and see which cards give you data you can monitor. And review drive temps as well.

I’d also get extra airflow over the hardware as well. Drop a spare case fan in there blowing over the HBA, maybe the NIC - heck, an old school solution is to open the case and point a big box fan at it.

Is AER enabled in the BIOS? Have you tried forcing the HBA’s PCIe slot to 3.0 speed? Have you tried new SAS cables?

And finally, have you tried running TrueNAS on bare metal rather than in a VM?

These kinds of hardware problems can be very annoying. Here’s the Hail Mary list of things I’d generally try:

  1. Update all firmwares and softwares.
  2. Reseat all the hardware.
  3. Reproduce the problem repeatedly to identify a reliable trigger.
  4. Swap hardware to isolate the fault (move HDDs between slots, swap backplane controllers, etc).
  5. If it’s a hiesenbug that is easily verified, with a quick fix, automate it.

In your situation the first thing I’d do is a memtest, then firmware update, then swap HDDs. Make sure you document everything you’ve done.

That said, once your hardware is cursed, maybe it’s time to start planning a replacement.

My understanding is that ZFS needs direct physical access to the disks and the controller in order to work reliably. Have you blocklisted the HBA controller in Proxmox?

truenas is very sensitive to disk firmware versions. check those.