OpenZFS 2.3, Direct IO, and NVMe

The 2.3.0 release of OpenZFS adds support for direct IO, which is briefly described in the release notes:

More, from the merged pull request:

"By adding Direct IO support to ZFS, the ARC can be bypassed when issuing reads/writes. There are certain cases where caching data in the ARC can decrease overall performance. In particular the performance of ZPool’s composed of NVMe devices displayed poor read/write performance due to the extra overhead of memcpy’s issued to the ARC.

There are also cases where caching in the ARC may not make sense such as when data will not be referenced later. By using the O_DIRECT flag, unnecessary data copies to the ARC can be avoided."

And then described in greater detail in this talk from a few years ago:

It sounds compelling to me. Has anyone had time to experiment with a direct IO-enabled pool? (Is it yet too new?)

  • What are the performance characteristics of direct IO in practice?
  • How does it / does it not improve performance of NVMe-only ZFS topologies?
  • Is it a situational optimization or is it more generally helpful?

Thanks.

2 Likes

I’ve been using direct=standard for awhile now and didn’t really notice any difference, however out of pure morbid curiosity I switched to direct=always today.

dd if=/dev/zero of=/mnt/zero bs=4M
530669+0 records in
530669+0 records out
2225787109376 bytes (2.2 TB, 2.0 TiB) copied, 250.458 s, 8.9 GB/s

Definitely see a difference there. Almost an extra 4 GB/s on sequential writes.

4 Likes

A half year later, I’m curious if anyone has gained any greater experience with direct IO?

Hello,
I did not investigate direct io since it does not support ZVOLs, and in my case I use ZVOLs.
For me, SSD performance is crucial, and despite using ZFS for 10 years on HDDs… it’s shortcomings regarding SSD and ZVOLs are kind of important.
Of course, everyone’s needs differ, so please do not see this as a straight-up recommendation against ZFS.

I am currently setting up a new server, this machine will exclusively host VMs, some of which will require high IO perf when dealing with random reads (DBs and such).

For the high-IOPS subsection of SSDs I decided to go with (bottom-up, devs first):

  • SSDs
  • dm-integrity
  • md-raid10
  • LVM
  • LVM-thin with auto-resize

I need dm-integrity to be absolutely sure that the data is correct (RAID10 on its own doesn’t care for that). I understand that this costs some compute, but so would ZFS checksums.

LVM-thin with auto-resize: This is mostly a work-around to be able to trim free space. When You allocate everything to LVM-thin LV, and a need to trim free space comes up - there is now way to do it (or it is unknown to me). With this setup I can create an “ordinary” LV and then TRIM it.

1 Like