Buddy Backup And More Storage Server: Part 1, Construction

The Build

eBay Affiliate Listing for Disk Shelf:

The Fitness Tests

Smartctl

The first tool in your toolbox is smartctl

It’s going to take 900 minutes for that test to complete! A watched pot never boils:

And a selection of failures:


Error 7 occurred at disk power-on lifetime: 47833 hours (1993 days + 1 hours)
  When the command that caused the error occurred, the device was active or idle.

image

sudo apt-get install smartmontools lsscsi mailutils ksh lvm2
git clone https://github.com/ezonakiusagi/bht.git 
cd bht 
./bht /dev/drive1 /dev/drive2 /dev/drive3 

To check on the status of the test run, execute bht with --status option:

$ cd /path/to/test/data
$ bht --status

So on my system it shook out like:

INFO: changing to /home/w/fitdrive/bht.
no pools available
no pools available
no pools available
no pools available
no pools available
no pools available
no pools available
ATTN: this hard drive testing process will wipe all data
ATTN: on the following hard drives:
ATTN: /dev/sdi /dev/sdj /dev/sdk /dev/sdl /dev/sdm /dev/sdn /dev/sdo
ATTN: ARE YOU SURE YOU WANT TO PROCEED?: yes
INFO: collecting SMART data from each drive.
INFO: WDCWD30EFRX-68EUZN0 / WD-WMC4N0E9FHA7
INFO: WDCWD30EFRX-68EUZN0 / WD-WMC4N0DA443P
INFO: WDCWD30EFRX-68EUZN0 / WD-WMC4N0E1PN2L
INFO: HGSTHDN724030ALE640 / PK1234P8JMU31P
INFO: WDCWD30EFRX-68EUZN0 / WD-WCC4NHSL2LUL
INFO: WDCWD30EFRX-68EUZN0 / WD-WCC4NFLDULC4
INFO: WDCWD30EFRX-68EUZN0 / WD-WMC4N0E6V977
INFO: Running badblocks on /dev/sdi.
INFO: ==> output to /home/w/fitdrive/bht/disks/WDCWD30EFRX-68EUZN0_WD-WMC4N0E9FHA7.
INFO: ==> [WDCWD30EFRX-68EUZN0|WD-WMC4N0E9FHA7]
INFO: Running badblocks on /dev/sdj.
INFO: ==> output to /home/w/fitdrive/bht/disks/WDCWD30EFRX-68EUZN0_WD-WMC4N0DA443P.
INFO: ==> [WDCWD30EFRX-68EUZN0|WD-WMC4N0DA443P]
INFO: Running badblocks on /dev/sdk.
INFO: ==> output to /home/w/fitdrive/bht/disks/WDCWD30EFRX-68EUZN0_WD-WMC4N0E1PN2L.
INFO: ==> [WDCWD30EFRX-68EUZN0|WD-WMC4N0E1PN2L]
INFO: Running badblocks on /dev/sdl.
INFO: ==> output to /home/w/fitdrive/bht/disks/HGSTHDN724030ALE640_PK1234P8JMU31P.
INFO: ==> [HGSTHDN724030ALE640|PK1234P8JMU31P]
INFO: Running badblocks on /dev/sdm.
INFO: ==> output to /home/w/fitdrive/bht/disks/WDCWD30EFRX-68EUZN0_WD-WCC4NHSL2LUL.
INFO: ==> [WDCWD30EFRX-68EUZN0|WD-WCC4NHSL2LUL]
INFO: Running badblocks on /dev/sdn.
INFO: ==> output to /home/w/fitdrive/bht/disks/WDCWD30EFRX-68EUZN0_WD-WCC4NFLDULC4.
INFO: ==> [WDCWD30EFRX-68EUZN0|WD-WCC4NFLDULC4]
INFO: Running badblocks on /dev/sdo.
INFO: ==> output to /home/w/fitdrive/bht/disks/WDCWD30EFRX-68EUZN0_WD-WMC4N0E6V977.
INFO: ==> [WDCWD30EFRX-68EUZN0|WD-WMC4N0E6V977]

Note This Testing IS Destructive.

Bonus Rounds – what about Multipath?

With dual-controller SAS shelves, the same physical drive may be reachable through two SAS paths. Without multipath, Linux can expose the same disk twice. Multipath combines those paths into one block device so you do not accidentally write to the same disk through two independent /dev/sdX devices.

You might be able to enable multipath on your setup!

sudo apt update
sudo apt install -y multipath-tools lsscsi sg3-utils

sudo tee /etc/multipath.conf >/dev/null <<'EOF'
defaults {
    user_friendly_names yes
    find_multipaths no
}
EOF

sudo systemctl enable --now multipathd
sudo systemctl restart multipathd
sudo multipath -r

Check what can be seen by the system:

lsblk -S -o NAME,HCTL,VENDOR,MODEL,SERIAL
lsscsi -t
sudo multipath -ll
ls -l /dev/disk/by-id/dm-*

It can be worth it to setup multipath for redundancy and better performance. Typically SAS drives can do multipath connections natively but SATA drives cannot unless you have a sata interposer card either built-in to your enclosure OR that sits between your drive and the backplane. (Netapp is usually an extra-card-between-drive-and-backplane type of setup.)

Setting Up the ZFS Pool

ZFS AnyRaid is not quite ready yet? For this kind of use case? But It was fun to experiment with it in its current state.

zpool status tank
  pool: tank
 state: ONLINE
status: Some supported and requested features are not enabled on the pool.
        The pool has uneven levels of redundancy across top-level vdevs.
       action: Enable all features using 'zpool upgrade'. 
  scan: scrub repaired 0B in 18:42:11 with 0 errors on Sun May  3 06:42:11 2026
config:

        NAME             STATE     READ WRITE CKSUM
        tank             ONLINE       0     0     0
          raidz3-0       ONLINE       0     0     0
            mpath3t00    ONLINE       0     0     0
            mpath3t01    ONLINE       0     0     0
            mpath3t02    ONLINE       0     0     0
            mpath3t03    ONLINE       0     0     0
            mpath3t04    ONLINE       0     0     0
            mpath3t05    ONLINE       0     0     0
            mpath3t06    ONLINE       0     0     0
            mpath3t07    ONLINE       0     0     0
            mpath3t08    ONLINE       0     0     0
            mpath3t09    ONLINE       0     0     0
            mpath3t10    ONLINE       0     0     0
            mpath3t11    ONLINE       0     0     0
            mpath3t12    ONLINE       0     0     0
            mpath3t13    ONLINE       0     0     0
            mpath3t14    ONLINE       0     0     0
            mpath3t15    ONLINE       0     0     0
            mpath3t16    ONLINE       0     0     0
            mpath3t17    ONLINE       0     0     0
            mpath3t18    ONLINE       0     0     0
            mpath3t19    ONLINE       0     0     0
          raidz3-1       ONLINE       0     0     0
            mpath4t00    ONLINE       0     0     0
            mpath4t01    ONLINE       0     0     0
            mpath4t02    ONLINE       0     0     0
            mpath4t03    ONLINE       0     0     0
            mpath4t04    ONLINE       0     0     0
            mpath4t05    ONLINE       0     0     0
            mpath4t06    ONLINE       0     0     0
            mpath4t07    ONLINE       0     0     0
            mpath4t08    ONLINE       0     0     0
            mpath4t09    ONLINE       0     0     0
            mpath4t10    ONLINE       0     0     0
            mpath4t11    ONLINE       0     0     0
            mpath4t12    ONLINE       0     0     0
            mpath4t13    ONLINE       0     0     0
            mpath4t14    ONLINE       0     0     0
            mpath4t15    ONLINE       0     0     0
          raidz2-2       ONLINE       0     0     0
            mpath10t00   ONLINE       0     0     0
            mpath10t01   ONLINE       0     0     0
            mpath10t02   ONLINE       0     0     0
            mpath10t03   ONLINE       0     0     0
            mpath10t04   ONLINE       0     0     0
            mpath10t05   ONLINE       0     0     0

NOTE: The uneven (mixed raidz2 and 3) is not recommended as explained in the video. If one vdev has a failure, the entire pool is lost.

I created another pool with the other drives:

zpool status suspool
  pool: suspool
 state: ONLINE
status: 
  scan: scrub repaired 0B in 05:16:33 with 0 errors on Sun May  3 11:16:33 2026
config:

        NAME             STATE     READ WRITE CKSUM
        suspool          ONLINE       0     0     0
          raidz1-0       ONLINE       0     0     0
            mpath8t00    ONLINE       0     0     0
            mpath8t01    ONLINE       0     0     0
            mpath8t02    ONLINE       0     0     0
            mpath8t03    ONLINE       0     0     0
          raidz1-1       ONLINE       0     0     0
            mpath6t00    ONLINE       0     0     0
            mpath6t01    ONLINE       0     0     0
            mpath6t02    ONLINE       0     0     0
            mpath6t03    ONLINE       0     0     0


If you really want “one big pool” I could suggest mergerfs but it has a lot of downsides. Better to treat the two pools like datasets in one big pool.

Performance

We are able to saturate 10 gigabit networking! That was not entirely expected.

I mean theoretically we could do 42 × 180 MB/s = 7,560 MB/s

But there are bottlenecks in this netapp diskshelf, especially with sata disks. Real world?

SAS shelf ceiling: ~10 GB/s usable
Drive sequential read: ~7.5 GB/s ideal
Real sequential read: ~2.5–4.0 GB/s

RAIDZ sequential write: ~6.1 GB/s ideal
Real sequential write: ~2.0–3.5 GB/s

Random read IOPS: ~3k–4k best-case
Random write IOPS: ~100–300 realistic

This is actually.. not awful! The performance is very uneven however:

Still.. this isn’t as bad as I was expecting and the I/O almost doesn’t drop below 3 gigabytes/sec.

What about the small SSDs?

The disk shelf was more of a bottleneck than expected. Each Sata SSD drive can do 500 megabytes per second so spreading them around can make a big difference. That is by far the best part of the pool.

Final Notes

Remember to try to locate the sata SSD drives elsewhere than the front row of each shelf of disks for airflow reasons. And try to populate the front four spots of each drawer if you’re using this disk shelf.

15 Likes

Thx for sharing! :+1:

video soon!

3 Likes

Totally understood all that.
Uh huh.
Like, yeah bro.

(well i got the saturation part.. consolation prize, lol)

So, not trying to self promote or anything, just trying to be helpful in general, the code is AI slop that I made to solve a problem for myself, but if it helps others, then cool.

I made a util ( GitHub - gcs8/truenas-jbod-ui · GitHub ) for giving a good visual display of the disks in a system, you can add on history/stat tracking if you want, but it will run more or less stateless. It started as just a way to get the right LED to blink on a given disk as that is not always clear and I hated the jbod display that TrueNAS made and the order trays where displayed in did not match with my supermicro gear.

I hope this helps someone other than just me, other than LED control it’s just setup as least permissions use and the admin side car will tell you (or do it for you) just the limited commands you need to get it going, and bad AI written Wiki on the project.

Sorry Wendell, if you think this is clutter you can purge this post, no hard feelings. I just know keeping track of disks and/or blinking an ID light for someone hours away can be a lifesaver.

1 Like

Also, to share what I am doing for my offsite replica. It’s a Hypervisor with a TrueNAS VM with the HBA passed through to it, a Quantasotr VM with a iSCSI LUN from the TN box for it’s storage, a HeadScale exit node VM to phone back home and get linked up. It’s been working great.

The HW is just one of these 36 bay units and some 14T refurb SAS disks I got a great deal for 26 of them back in 2024.

3 Likes

yas, this is what the community is about, high fives!

3 Likes

I’d do things differently. This kinda trash belongs to the dumpster. If you have free energy (idk, solar?) then it’s still usable, but with a big caveat.

There’s so many drives that can go bad in here, it’s not even funny. Try to build 4 or 5 different vdevs. A raid-z3 pool with 16x 4TB and 20x 3TB drives is terrible. Even assuming 2 stripped raid-z3 (so 13/16 4TB and 17/20 3T usable disks), you still get a fault tolerance of 6 drives out of 36 (that’s 1/6 fault tolerance).

Personally I’d not trust that. Given how many of these drives are expected to fail in the next 1 to 2 years, I’d go waaaay more redundancy. I’d even be willing to go triple mirrors with multiple stripped vdevs, but if you want a bit more capacity, you can do something like 2x raid-z2 vdevs with 10x 3TB drives (so 2 drives can fail in each vdev) and 2x raid-z2 vdevs, 1 with 10x 4TB and another with 6x 4TB (so another 2 drives in each vdev, with the smaller 1 being filled with the potentially less reliable drives).

That’ll raise the fault tolerance from 6 disks to 8 and you’ll have less chances of fully losing the pool. What I mean is that when you lose 4 disks in a raid-z3 of 20 disks (which is very likely), the whole thing goes poof. Buy if you lose 2 disks in one vdev and 2 in another vdev (e.g. in the 10x 3TB vdevs), you’re still fine. Of course, lose 3 disks in a single vdev and good-bye data, but IMO smaller vdevs are less likely to get screwed up.

And whatever you do, don’t try to use the 3 pools separately. The only way I’d agree with Wendell with the big capacity raid-z3 pool is if you do zfs replication from the raid-z3 pool over to the raid-z1 and raid-z2 pools (e.g. you have a dataset on the z3 pool called “vidprojectsbkp” that gets zfs sent to the raid-z1 pool and another dataset “immichbkp” that gets zfs sent to the raid-z2 pool). With that kind of redundancy, having a big pool that could crash at any time isn’t a big deal (since you have more than 1 copy of your data).

That’s why I always used my label printer to print the serial of the disks on the chassis. If a disk is screwed, I’d just note the serial and look in front of the server for the proper code. Replace a disk, replace the label with the serial.

Not bashing on your project. I’m glad something like that exists (even if I’m not likely to use it), but I hope that will help others coming. You should definitely make a bigger thread announcing it (maybe blog about updates in that thread).

This might be a decent use case for draid, would need enough spare HW to play with that though.

I am still working on it, might do that in a couple weeks.

It’s all good, it’s more than a one trick pony, but it’s meant to be sorta whatever you need/want it to be. Eg: See the attached zip, it’s a fully offline snapshot (just a HTML file) that you can attach to a ticket or use as a system state export if you had to fully disassemble the setup and move it you can use this as auto documentation for it, or a before and after, etc…

jbod-snapshot-archive-core-lsi-f-sas3x48front-0c04-lsi-r-sas3x48rear-0c04-20260514t210853z.zip (431.5 KB)

Yes! i knew my hardware hoarding would be useful one day. I have recently started to run out of space on my current setup and have a massive pile of old hdds and had been thinking about something like this using unraid. I have a license as ive used it in the past (several years ago) but the performance was terrible, but for a backup for my backups :thinking: , guess i have a project for this weekend now.

Pictures or it didn’t happen! :smiling_face_with_sunglasses: :+1:

(Always helpful to visually see what others are doing.)

Most disks come pre-labeled from the factory with their serials on the front, so if you have caddieless disk retention (like top loading chassis) you can see the labels clearly which is super convenient:

5 Likes

Does this even make sense financially?

200 watts with out any drives
320 watts idle
800 watts load

Let’s assume 400 watts for average load.
400Wx8760h= 3500 kWh/year

power cost: between 0.2-0.4 euro/kwh
700 - 1400 euro just for the power cost.

I see these low capacity high millage used HDD going for 5-10 euro/TB

If we get 2 years use out of them that would be 2.5-5 euro/TB/year

Lower estimate for 200 TB is 700+500= 1200 euro/year
high estimate for 200 TB is 1400+1000= 2400 euro/year

One can end up paying more for power then the HDD. I don’t think it makes sense to go out and buy a disk shelf and small 3-4 TB drives. If someone owns the drives it might make sense. Otherwise Serverpartsdeals etc sells 28TB HDD for 20 euro/TB, the possible savings in power and replacement rate will probably make it cheaper.

2 Likes

I love ewaste backup servers! Here’s one of mine - it was the best price, free.

Datto/ASRock “Beebox” KBL-NUC with a Celeron 3865U, 8 GiB of RAM, a 16 GB Optane boot disk, and a very old, very suspicious 5 TB 15mm 2.5” disk as a single-drive ZFS pool.

Syncs my important VMs from my main PBS server at home, sits happily on my desk at the office and connects to HQ with NetBird. At the moment, it contains 39658 snapshots and is using just 3.25 TB of 4.86 TB (praise be PBS dedup and small VMs).

These Datto boxes were shipped running Ubuntu with ZFS. This one had a 2 TB 15mm spinner from the factory, and already spent quite a few years as a backup target. It was replaced a year or two ago and I thought it would be fitting to press it back into service with a bigger drive.

Additionally, in the spirit of the thread, here’s my favorite disk fitness test.. extremely sophisticated:

for disk in /dev/sd{c,d,e,f,g,h,i,j,k,l,m,n,o,p,q,r,s,t,u,v,w,x,y,z,aa,ab,ac,ad,ae,af,ag,ah,ai,aj,ak,al,am,an,ao,ap,aq,ar,as,at,au,av,aw,ax,ay,az,ba,bb,bc,bd,be,bf,bg,bh,bi,bj,bk,bl,bm,bn,bo,bp,bq,br,bs,bt,bu,bv,bw,bx,by,bz,ca,cb,cc,cd,ce,cf,cg,ch,ci,cj,ck,cl,cm,cn,co,cp,cq,cr,cs,ct}; do
    sudo badblocks -wsv "$disk" -b 4096 -o ./badblocks_${disk##*/}.log > ./badblocks_status_${disk##*/}.log 2>&1 &
done
1 Like

You know you can simplify this:

by this:

for $disk in `cat /dev/sd*` ; do
    sudo badblocks -wsv "$disk" -b 4096 -o ./badblocks_$disk/}.log > ./badblocks_status_$disk/}.log 2>&1 &
done

Wildcards are such a blessing :wink:

Wouldn’t this wipe out your OS if your OS was /dev/sda?

2 Likes