committed 10:03PM - 20 Aug 26 UTC
Linux's FIDEDUPERANGE ioctl lets out-of-band dedupe tools such as
duperemove and… bees ask a filesystem to share byte-identical file
ranges. ZFS returned EOPNOTSUPP for it. Implement it on top of block
cloning, so matching ranges share storage without the memory cost of
the dedup table.
zfs_clone_range() is split into a precheck and a locked helper so the
dedupe path can reuse the clone loop. zfs_dedupe_range() compares the
two ranges and, only if they are byte-for-byte identical, clones the
source blocks over the destination, all under a single RL_READER and
RL_WRITER hold, so the bytes that get cloned are exactly the bytes
that were compared. Source and destination must share a block size,
which the clone step requires anyway.
The comparison settles a block from its block pointers when it can:
equal strong checksums prove it identical, and, for blocks stored
verbatim, unequal checksums prove it different. Two embedded blocks
are compared by their payloads without decompressing anything, and a
matching first DVA at a matching physical birth proves equality by
itself, with no checksum involved, so it works even for encrypted or
weakly checksummed blocks. Only what the pointers cannot settle is
read, and the read compares the dbufs in place rather than copies of
them, a whole block at a time, since a block is read and decompressed
in full whatever part of it is asked for. The compare loops break out
with EINTR on a pending signal, like the clone loop does.
A range that is already fully shared, which is what a previously
deduped pair of files looks like to a tool that cannot see ZFS-level
sharing, skips the clone step entirely: the request reports its full
length as deduped and changes nothing on disk. Redoing the clone
would dirty every block only to install the pointers it already has,
restamp its logical birth, making the next incremental send carry
every such block again, and push a BRT incref plus a deferred decref
through sync for a net refcount change of zero.
Both ranges have to lie within their files. Trimming a range down to
the destination EOF would be invisible to the caller, because the
ioctl reports the length it asked for rather than the one the
filesystem returns, so a shortened dedupe cannot be told apart from a
complete one; generic_remap_checks() rejects this for the filesystems
that use it and we now do the same. The full requested range is
compared before it is aligned down to a block boundary, so a differing
trailing sub-block reports DIFFERS instead of being rounded away and
called identical.
A dirty block makes the clone step wait for the txg regardless of
zfs_bclone_wait_dirty: a dedupe has no copy fallback to absorb a
shortened range, so EAGAIN would surface to the ioctl caller, and the
comparison has already read the very data it would be refusing to
wait for.
A dedupe leaves the destination content alone, so it does not update
mtime or ctime, strip setid bits, or write back any other attribute,
and it is not logged to the ZIL: losing it to a crash costs only the
sharing, which is all the ioctl promises, and avoids restamping
timestamps on replay. It also leaves the page cache alone. Both
files are flushed before the comparison, and the blocks the clone
installs hold the bytes that were already there, so a cached page is
still correct; the only page a refresh could change is one carrying an
mmap store made after the flush, and overwriting that would discard a
write.
On the Linux side both the remap_file_range(REMAP_FILE_DEDUP) path and
the pre-4.20 dedupe_file_range fop are wired up, EBADE maps to
FILE_DEDUPE_RANGE_DIFFERS, and both remap paths now take the two inode
locks in address order. The clone path has had that ordering bug
since it was written and the dedupe path would have inherited it, so
opposite-direction requests could deadlock.
Reviewed-by: Brian Behlendorf <[email protected]>
Reviewed-by: Alexander Motin <[email protected]>
Reviewed-by: Chunwei Chen <[email protected]>
Signed-off-by: MorganaFuture <[email protected]>
Closes #11065
Closes #18745