I use ZFS 2.3.0 on Rocky Linux 9.5 with ELRepo 6.12 kernel.
In my home environment I’ve migrated 10 TB of data from a 2x18TB mirror to 3x18TB raidz1 and added a special vdev for metadata to the new zpool.
The old mirror zpool had compression=on and default settings for recordsize, the new raidz1 zpool was created with recordsize=1M and compression=on.
The special vdev for metadata is two mirrored 1 TB SSD’s which is overkill sizewise but I hade them on the shelf unused. That also meant that there’s plenty of room for small files so I set special_small_blocks=64K before migrating the data to the new zpool, having calculated from “zdb -Lbbbs” statistics from the old zpool that the total size of metadata (Metadata Total: 7.18G ASIZE) and files 64kB and smaller (64K asize Cum. 96.3G) was a little over 100 GB.
After migration (rsync), the special vdev was barely in use, 1.3 GB according to “zpool list -v”. The data I used to calculate must either be incorrect or I’ve misunderstood what they mean. I guess that the change from default recordsize to 1M meant that the metadata was a lot smaller, Metadata Total: 1.04G ASIZE, which is understandable since there’s a lot of large files in the zpool.
I’m having a harder time understanding that setting special_small_blocks=64K made basically no difference and was nowhere near the old number, after migration 64K asize Cum. was 1.43G.
I have tried setting special_small_blocks=512K and doing a zfs send/receive to a new dataset, expecting that all files 512 kB and smaller would end up on the special vdev, but I had to cancel the operation after 100 GB was copied since basically all 100 GB ended up on the special vdev. Tried zfs send/receive again with special_small_blocks=256K, aborted after 350 GB since 340 GB had ended up on the special vdev.
Changed special_small_blocks back to 512K on the new datasets and did a rsync to the new datasets, that didn’t fill the special vdev but I ended up with only about 5 GB usage on the special vdev after all 10 TB had been rsynced to new datasets with special_small_blocks=512K.
Can anyone explain what happened when I did zfs send/receive? Also what I’m doing wrong or misunderstanding? I have full “zdb -Lbbbs” output from the old zpool and the new zpool if there’s any more info needed.
No, unfortunately I still don’t understand why using zfs send/receive meant that almost all data ended up on the special vdev. Perhaps the special vdev is used as a form of cache during transfer or it is simply a bug. Using rsync avoided filling the special vdev.
I can’t make sense of the current output from “zdb -Lbbbs” either:
Metadata total is 1.10 GB (mirrored, meaning the actual data is 551 MB)
The cumulative size of all blocks 512K and smaller is 11.9 GB (raidz1, the actual data is 2/3 = 7.9 GB)
I assume that 11.9 GB is actually 7.9 GB (2/3 of 11.9) since I know that the zpool contains 10.5 TB of data but the cumulative size is listed as 15.7 TB (including parity data).
So 551 MB of metadata + 7.9 GB of special_small_blocks = 8.45 GB … but the last table lists 5.09 GB used on “mirror (special)” and that (I assume) is the total mirrored size so the actual datasize is 2.545 GB which means that just under 2 GB of special_small_blocks are actually stored on the special vdev.
There might be some compression happening that I’m not aware of? I created the special vdev using the same parameters as Wendell did in ZFS Metadata Special Device: Z, i e only “-o ashift=12”, no “-O compression=on” unless it is inherited from the zpool it was added to.
Zfs send receive sends and receives blocks as is, so if you send a 128k block and the new pool recives a 128k block it just writes it as is. ZFS send/recieve has no knowledge of what other blocks it goes with or how to reassemble the file that the block contains a part of, so it can try to re arrange it in a new block size. Rsync on the other hand is working at a file level. So its giving the zfs the whole file, where zfs can then store it as it pleases.
Your on the right track here, rather any fragmented files that made small blocks, now that your having zfs re-decide how to store the file having rewritten the whole file rather than parts allows it to rearrange the fragments into whole blocks, in effect getting rid of the some amount of small blocks.
ok so you got the expected behavior of zfs here, zfs sent a stream of 128k blocks and recived a stream of 128kblocks, all smallrer than 512k so they all went on the special. Again zfs send/recv is not aware of the files or their size realative to the blocks there stored in. its just taking one block after another and plunking it in the new dataset or pool.
Thanks, that explains the behaviour. I wasn’t sure how to redistribute existing small files to the special vdev after changing special_small_blocks and found this example using send/receive, I understand now that send/receive works if you want to move metadata for existing files to a special vdev, not if you want to redistribute small files.
EDIT: Unless you use the --large-block option for zfs send, see below.
you mean small blocks, what size block zfs stores a file in is up to it. a 48k file may be stored in 3 16k blocks and not 2 32k blocks or 1 64k block. less wasted space.
Im not aware of the exact allgorithm zfs uses to determine this so take in merly as an example of one way it could be doing things. The point is zfs thinks in terms of blocks not files. if it puts your file in smaller blocks then yes it would end up on the special but it may not.
Ok, I’m probably mixing up record size, block size and file size.
My understanding was that setting recordsize to 1M meant that files larger than 1 MB are stored in one or more blocks of 1M. Files smaller than 1 MB are stored in blocks the size of the closest (larger) multiple of 2, for example a file that is 50 kB will be stored in a 64K block.
But send/receive sends all the data in 128k blocks? Even when recordsize was set to 1M when the zpool was created?
EDIT: Might have found the answer in the zfs-send man page:
-L, --large-block
Generate a stream which may contain blocks larger than 128 KiB. This flag has no effect if the large_blocks pool feature is disabled, or
if the recordsize property of this filesystem has never been set above 128 KiB. The receiving system must have the large_blocks pool fea‐
ture enabled as well.
So, if I had used the above option with zfs send, it may have redistributed small blocks (only) to the special vdev.
Ok, so the zfs pool has a block size, each dataset has a record size, and then each file has its own record size, and these all can be different.
So i make a pool tank with a block size of 64k, then make a dataset shared with a record size of 1m on tank. when i save a large video file to tank/shared zfs will try to save the file in the datasets record size of 1m not the pools block size. Also though if i save a word document of 2kb to tank/shared and zfs uses a 2k block the recordsize in metadata for the files blocks is that the file is stored in 2k blocks not a 1m size which is the dataset target or 64k the pool block size.
No it streams the blocks as is, so it does not have to decompress or recompress, or rehash the blocks one after the other in order. This is quickest. then on the receiving pool if the pool simply writes it to the pool as is, again for the same reason less work. this means the snapshot feature works sending snapshots because the snapshot only has to send changed or new blocks and the receiving pool since it did not change anything when writing just overwrites changed blocks or writes new blocks as is. This is why zfs send/recv is so blazingly fast compared to rsync cp and other methods of moving data. its shortcutting majority of the work(compression, deduplication, and parity) by copying the origin pools homework, instead of having to undo and redo it on each end.
No, it just is a flag letting zfs send know the origin pool has blocks larger than the zfs default.
I’m having trouble differentiating between small files and small blocks, one example from Klara Systems:
Optionally, the SPECIAL may also be configured to store small data blocks. For example, with special_small_blocks=4K, individual files 4KiB or smaller will be stored entirely on the SPECIAL; with special_small_blocks=64K, files 64KiB or smaller are stored entirely on the SPECIAL.
That’s basically what I expected when setting special_small_blocks=512K, all files 512 kB and smaller would end up on the special vdev.
Just to recap, my pool was created with recordsize=1M and special_small_blocks=64K. After initial data transfer (using rsync) I changed special_small_blocks from 64K to 512K and tried to use zfs send/receive from one dataset to another in the same pool to trigger a redistribution of small blocks to the special vdev, with the unexpected (for me at least) result that pretty much all data ended up on the special vdev. Makes sense if zfs send/receive used 128K blocks, not if it used the actual recordsize of 1M.
If I understand you correctly, zfs send/receive can’t be used to redistribute small blocks after special_small_blocks has been changed, even with --large-block?
Copying the data using rsync redistributed the data. You need free space to do it, since I had (a lot) more free space than data I just rsynced all the data to a new dataset, destroyed the old dataset and renamed the new dataset to the same name as the old one.
If I would do it today I would have tried the new zfs rewrite command, available in ZFS 2.3.4 and later. Not 100% sure it would move small files to the special vdev but it sounds like it might.