Experiments with dm-integrity and T10-PI on Linux

Just want to say that this is a fantastic open minded thread and I really appreciate seeing others explore options outside the norm here.

I’ve been testing with LVM Integrity RAID which somewhat simplifies the manual Device Mapper work.

I included some additional detail in this thread:

In my testing the write penalty has been a bit less favorable in both 7k sata HDD and NVMe integrity mirrors.

  • In 7k SATA HDD testing (R1 mirror of 2 drives) I would see a 90% decrease in write performance but only about 15% decrease in read performance.
  • In NVMe testing (R1 mirror of 2 drives) I would see a 73% decrease in write performance and about 18% decrease in read performance

7K drives are Seagate EXOS (ST10000NM0146)

NVMe are WD SN850

Physical disk and Integrity mirror block sizes aligned at 4k. Details on setup here:

I’ve been testing primarily in Rocky 9 and Alma 9.

Additional findings - I’ve been hard pressed to get to a reasonable workflow for LVM Integrity RAID plus writecache. Helpful thread with awesome LVM devs here:

  • One dev even published a testing rpm for a custom LVM2 build that allows integrity RAID with caching
  • I’m not sure how this will fair with regular system updates but definitely something to test more

I’ve also been testing with Stratis layered above LVM Integrity RAID

  • simple read Caching setup with incredible performance
  • but still no write caching so writes remain in the toilet on HDD based arrays

There is a way to test write cache performance but it doesn’t persist a reboot yet. More here:

When Stratis matures to the point of having write caching and some turnkey snapshot and replication automations that is the direction I plan to go.

  • but at this point it does not yet offer built in raid capability (has to be layered over LVM or MDADM RAID)
  • and has no provisions for integrity checksums (again must be layered)
2 Likes

Hey, thanks for joining in. Hmmm… I’ve been leaning towards doing an integrity/mdraid1 setup on a new fedora machine (since I have some nvme drives that don’t support T10-PI) but your report of a 70-90% write performance hit has me worried and wanting to try this on real hardware fairly soon (and before I actually start moving real files onto the mirror). I’ll do so when I get a chance and will report back.

1 Like

If you stick with SSD for your LVM / DM-integrity raid and then layer Stratis on top I think you will be more than happy with the performance. Whatever secret sauce Stratis has going is phenomenal for SSD performance.

BcacheFS is likely to deliver sooner than Stratis on all their goals lol. But I would definitely add Stratis that to the testing mix. Most of us are living on a lot of read IO and not that much write anyway.

1 Like

On many of the major distributions, there is a default cron job or timer to regularly perform resync. On Ubuntu, for example, it is handled by mdcheck_start.timer; on Debian, it is /etc/cron.d/mdadm; and on Centos, it is /etc/cron.d/raid-check.

2 Likes

Related:

2 Likes

I’ve never been happier to catch a hardware error… :smiling_face_with_tear:

T10 works. The SSD (Intel Optane P4800X) itself takes the checksum generated by the OS and verifies the data.


I only noticed because the virtual machines were behaving strangely. Particular applications would become sluggish or outright freeze, persisting between reboots of the virtual machine and the entire system.

The data in /home directory was fine and all the virtual machine images were fine too. But yeah… time to swap in a fresh drive and put that one in the triage bin. And I think it’s about time to put T10 to good use in my daily driver systems.

3 Likes

@LiKenun interesting, was this silent coruption on disk or corruption in the signal during a write? was there anything in dmesg? I’m wondering how to do a “scrub” on a T10 disk, basically.

I should have updated this sooner, but since I passed my WiFi device into my guest, my host has no internet connection. The modem is sufficiently far away that I don’t yet have a cable to connect it via Ethernet.

In a nutshell: The sector was technically corrupt, but it wasn’t the data. I dug a lot deeper after seeing the problem again. /boot/grub/grubenv was the problem. While Linux is able to ensure that checksums for each sector are generated, it’s not true of the GRUB environment which simply writes the data part of the sector. Thus, the log entries about data corruption would always appear right around reading that file because the checksum would never check out.

That means I still don’t really know if the hardware is checking that the write payloads have been received correctly. Perhaps it does, but only when a capable OS requests it to be so. That means I’ll be waiting a while for a real write corruption to trigger the T10 integrity checks. :sweat_smile:

How I found out:

  • entries=$(nvme smart-log /dev/nvme0n1 | grep --color=never -oP '^num_err_log_entries\s*:\s\K[[:digit:]]+$'); while [ $entries -gt 0 ]; do nvme error-log /dev/nvme0n1 -e $entries | grep --color=never -oP '^lba\s*:\s\K0x[[:xdigit:]]+$'; : $((entries--)); done | sort | uniq gave me the unique LBAs (in hexadecimal format) which encountered the checksum error.
  • nvme read /dev/nvme0n1 -s 0xb7e700 -r 0xb7e700 -c 0 -z 4096 -p 15 | od -t x1 performed a read with verification. (The LBA given to -s and -r must match as the checksum guards the sector number which is written into the metadata. -c is the count of sectors to read minus 1; -c 1 would read a count of 2 sectors. -z is the length of the data portion of the extended sector. -p is the protection information mode as a sum of four bit flags. Removing the -r and -p options would read out the data without verification.)
  • find / \( -path /dev -prune -o -path /proc -prune -o -path /run -prune -o -path /sys -prune \) -o -type f -exec cat "{}" + > /dev/null was my sanity check to see which file tripped the read verification.

This was not entirely reproducible, and definitely another issue unrelated to storage. The sluggish application was fixed by setting a flag (which I don’t remember unsetting before). Not sure if the underlying QCOW2 becoming fragmented has anything to do with the residual degraded guest performance. That’s a rabbit hole I’ll jump into another time…

1 Like

Got it. Thanks for posting your investigation process, it’s clever. An interesting saga, this has been. You probably thought of this but you could perhaps format your nvme device with two namespaces, use one for efi, leaving that formatted without T10-PI, and a second for linux, that uses T10-PI

1 Like

Gonna have to put /boot on a physically separate device. Between the Optane P4800X and P5800X, the former supports OPAL but only 1 namespace, and the latter supports 128 namespaces but no OPAL. Can’t have it all… :face_holding_back_tears::joy:

2 Likes

I’ve been exploring either replacing the md mirror on my main workstation with btrfs or zfs, or adding dm-integrity to it. I do a lot of zfs work so that’s the obvious choice, but I’ve had some bad experiences with zfs root, which is why I usually use a simple md mirror (and lvm) for root.

Anyway, my primary concern with dm-integrity is write amplification, i.e. writing x+8 bytes to flash for every x bytes of data, so I’ve been exploring if I could format an NVMe or SAS SSD to 520 or 4,104 byte sectors and alleviate the write amplification issue. That lead me here, to a much more interesting solution. :smile:

Fortunately I have a PM1725a handy that supports multiple namespaces and T10-PI and x+8 LBA formats, so I’m going to attempt to replicate all y’all’s amazing work to figure out the optimal configuration before pulling the trigger on a 9500-8i or something and some DC-class NVMe or SAS SSDs to replace my 2TB SATA SSDs. Thanks for the head start! I still need to figure out the best way to measure write amplification…

# nvme id-ctrl -H /dev/nvme1n1
<snip>
nn        : 32
<snip>
# nvme id-ns -H /dev/nvme1n1
<snip>
mc      : 0x3
  [1:1] : 0x1   Metadata Pointer Supported
  [0:0] : 0x1   Metadata as Part of Extended Data LBA Supported

dpc     : 0x1f
  [4:4] : 0x1   Protection Information Transferred as Last 8 Bytes of Metadata Supported
  [3:3] : 0x1   Protection Information Transferred as First 8 Bytes of Metadata Supported
  [2:2] : 0x1   Protection Information Type 3 Supported
  [1:1] : 0x1   Protection Information Type 2 Supported
  [0:0] : 0x1   Protection Information Type 1 Supported

dps     : 0
  [3:3] : 0     Protection Information is Transferred as Last 8 Bytes of Metadata
  [2:0] : 0     Protection Information Disabled
<snip>
LBA Format  0 : Metadata Size: 0   bytes - Data Size: 512 bytes - Relative Performance: 0x1 Better 
LBA Format  1 : Metadata Size: 8   bytes - Data Size: 512 bytes - Relative Performance: 0x3 Degraded 
LBA Format  2 : Metadata Size: 0   bytes - Data Size: 4096 bytes - Relative Performance: 0 Best (in use)
LBA Format  3 : Metadata Size: 8   bytes - Data Size: 4096 bytes - Relative Performance: 0x2 Good 
3 Likes

Of the people who’ve spoken up, LiKenun is the one who is experienced with doing this on real hardware. I’ve tried it on a SAS HDD (there’s a link somewhere in this thread to a different thread in which I discuss that) and I think it’s working, but SAS DIF and T10-PI for NVMe are different. Good luck and please share any results or instructions that you figure out!

I posted details in this thread, but I just had Type 1 DIF protection catch an error on a SAS disk, in case anyone is interested.

1 Like

I’m fairly confident that any enterprise drive that lists “end-to-end data path protection” or something similar in its data/spec sheet is good to go, and I’m about to buy a few Micron 9200 Pro drives to test this theory. Fingers crossed!

That isn’t always the case. If it isn’t specific about which “ends” I would inquire further about where the protection starts and what modes of failure the protection covers. For absolute assurance, I always grep the keywords T10, PI, DIF, DIX, and variable-sized sectors (including specific numbers like 520 and 4104). It’s kind of like “AES encryption.” There are countless drives that have it on their sticker, but you can’t actually control access to the data with a key, and it’s only there to make crypto erase possible; unless TCG Opal is supported, the AES doesn’t really guard your data at rest.

I don’t suppose you’re maintaining a list of what works and what doesn’t?

I have a nearly exhaustive list of which Optanes support it. (Not hard since there aren’t many SKUs anyway.) All of them are from the P480*X series or P580*X series. I suspect the H20s might have it on the Optane half of the device, but I have stopped trying to get that working.

For NAND M.2 SSDs, it’s an incredibly short list. It’s basically enterprise territory, and you aren’t getting your hands on it unless you are ready to buy in bulk (10 MOQ for the Netlist ones).

As for the U.2 NAND SSDs, last I looked, those which have any mention of T10 come from Kioxia and—of course—NetList again. Fadu’s website mentions an impressive array of features supported, although nowhere is anyone selling their SSDs. Upon scrutinizing the spec sheet very closely, Haishen’s SSDs probably support it (“Haishen5 Series supports multiple VSS sector formats”). Kioxia is the most recognizable brand and also the most obtainable of the bunch. There aren’t any avenues I know of which can be used to obtain the other ones.

Incredibly ironic is that in my quest to keep tabs on such drives on the market, I’ve turned to Perplexity.AI to do some of the web sleuthing for me. Its answer leads right back to my own post (on Tom’s Hardware). :joy: So it didn’t tell me anything I didn’t already know. :roll_eyes:

Screenshot

FWIW, Intel P4510s claim “end to end” protection and I can detect no evidence of T10-PI support.

Yep and sure enough the 9200 Pros are a bust, although the firmware is quite old so there’s still the faintest glimmer of hope that it will show support after an update. If I can find the bin…

1 Like

Add the Micron 7300 MAX/Pro SSDs to the list of probable T10-supporting M.2 SSDs. Micron’s own materials do not specifically contain “T10” though.

Sources:

Screenshot


Add the Western Digital Ultrastar DC SN730/SN100 (“T10 DIF”).


Add the Intel SSD DC P3500 (“T10 DIF protection” and “Variable Sector Size:512, 520, 528, 4096, 4104, 4160, 4224 Bytes”).


Add the Solidigm D5-P5336 series (“Variable sector size and end-to-end data-path protection”). But I should caution that Solidigm itself never mentions T10 or variable sector size support. This document may be describing a variant customized for Lenovo’s systems.


Kioxia specifically mentions its “FL6, CM7, and PM7” series.

  • The FL6 is a pretty close analog to the Intel Optane P5800X series and has TCG OPAL on top of that. (Though amusingly, it is actually not much cheaper than its Optane equivalent at this time.)
  • The CM7-V series is PCIe Gen 5, making it the fastest T10-DIF-supporting series known.
2 Likes