I have a server with TrueNAS Scale 25.10.6 as a Proxmox VM, and I’m having issues with write errors on the disks in the main pool. These almost always occur on weekends, presumably during backups. The strange thing is that if I turn the server off and on again, the disks may not show any errors for months, but when I perform updates and restart the TrueNAS VM or even the Proxmox server, errors can appear on the disks in the pool over the next few days (it usually targets one disk and gives errors only to that one until it disconnects and degrades the pool).
The disks aren’t faulty, and these errors appear randomly, sometimes on one disk, sometimes on another. Since I’m used to these errors, one of these disks failed in the past few months, and I mistook it for the usual problem, cleaning up the ZFS pool and crashing the TrueNAS VM because the disk had been set to read-only mode.
I’ve tried everything, and at this point I no longer know how to achieve a stable and reliable configuration.
The pool consists of four 10TB ST10000NM013G drives in RAIDZ2 with a log cache of two U.2 drives, with the controller passed to the VM by Proxmox (I used to do disk passthrough with the same result.).
My configuration is as follows:
CPU – A-EPYC9004-2G5
Motherboard – MBD-H13SSL-NT latest BIOS
RAM – 128GB
HBA Controller – AOC-S3816L-16IT latest firmware
Drives – RAIDZ2 4x ST10000NM013G Log 2x U.2-NVME-TLC-4TB-ST-G4
The drives do not have any active power management options, either in the BIOS or TrueNAS.
Can anyone point me in the right direction?
this is the truenas vm conf file in proxmox:
agent: 1
balloon: 0
boot: order=scsi0
cores: 12
cpu: host
hostpci0: 0000:01:00.0,pcie=1,rombar=0
hotplug: network,usb
ide2: none,media=cdrom
machine: q35
memory: 40960
meta: creation-qemu=8.0.2,ctime=1690897446
name: TrueNas
net0: virtio=D6:B9:5B:5E:69:6C,bridge=vmbr0,firewall=1
numa: 0
onboot: 1
ostype: l26
scsi0: local-zfs:vm-101-disk-0,discard=on,iothread=1,size=32G,ssd=1
scsi1: /dev/disk/by-id/nvme-eui.00000000000000000026b728302d15e5,backup=0,discard=on,iothread=1,serial=50026B728xxxxx>
scsi2: /dev/disk/by-id/nvme-eui.00000000000000000026b728302d1575,backup=0,discard=on,iothread=1,serial=50026B728yyyyy>
scsihw: virtio-scsi-single
smbios1: uuid=9d2f59b0-b601-42e3-958c-0b8c0d45414d
sockets: 1
startup: order=2,up=100
tags: debian;linux
vmgenid: df12398e-68d7-43eb-8be9-9ed086f452bb