Ask Level1Techs: Why not qcow2 on ZFS? What's double copy-on-write?

The question came up.. why would you NOT want to use qcow2 on zfs? Why does that make performance suck exactly?

Ah, it’s time to level up your understanding of how that particular sausage is made!

Where CoW Happens

Key point: qcow2 itself is a CoW allocator+mapper, and ZFS also CoWs every modified block. They each generate metadata updates and new allocations.

Guest filesystem (NTFS/ext4/XFS)
  └─ guest writes 4K/8K/16K block
     └─ Virtual disk layer (qcow2)
        ├─ L2 tables (mapping guest offsets → host clusters)
        ├─ Refcount table/blocks (cluster allocation/free bookkeeping)
        ├─ Data clusters (e.g., 64K clusters by default, configurable)
        └─ CoW rule: modifying an allocated cluster may require:
             - allocate a *new* cluster
             - write new data there
             - update mapping (L2) to point to it
             - update refcount metadata

Host filesystem layer (ZFS dataset)
  ├─ ZPL (POSIX layer) / DMU objects
  ├─ CoW rule: *every* changed block is written to new space
  ├─ Metadata CoW: indirect blocks, dnodes, space maps, etc.
  └─ TXG commit groups changes; sync semantics govern when it must hit stable storage

Physical vdevs (mirror / RAIDZ / etc.)
  └─ actual IO to NVMe/SAS/SATA

The “small random write” path

Assume:

  • Guest does a 4K write at some virtual disk offset.
  • qcow2 cluster size is 64K (common default).
  • qcow2 file resides on a ZFS dataset.

Case A: qcow2 cluster already allocated, overwriting part of it

In qcow2, you can’t just overwrite in-place if it would violate snapshot/refcount semantics; the typical path involves allocating new storage and remapping.

Guest 4K write @ VMOFF
   |
   v
qcow2:
  1) Read/locate L2 entry for VMOFF (may require reading L2 table)
  2) Allocate new 64K cluster (if CoW needed)
  3) Write 64K cluster data (often read-modify-write: old cluster → patch 4K → write new)
  4) Update L2 table entry to new cluster location
  5) Update refcount metadata for new cluster (+1) and old cluster (-1 possibly)
   |
   v
ZFS (because qcow2 is a file):
  For EACH qcow2 write above, ZFS CoW triggers:
  - new data blocks for:
      • cluster data write (64K-ish, may fragment into recordsize blocks)
      • L2 table write (metadata blocks)
      • refcount blocks write (metadata blocks)
  - plus ZFS metadata CoW:
      • indirect blocks up the tree
      • dnode updates
      • space maps / metaslabs bookkeeping
      • (optionally) intent log (ZIL) activity for sync writes

So one guest 4K write can become:

  • qcow2: data + multiple metadata writes
  • ZFS: data + multiple metadata writes again, because all those qcow2 updates are file writes that themselves trigger ZFS CoW and space accounting.

That’s the “double CoW” in practice.

So can you break down amplification for me?

Sure:

Guest I/O:                 4K write

qcow2 layer:
  - data cluster write:     up to 64K (cluster granularity / RMW dependent)
  - L2 update(s):           4K–64K total (depends on caching and table locality)
  - refcount update(s):     4K–64K total
  => qcow2 IO:              ~tens of KB to 100KB+ per 4K guest write (worst-ish case)

ZFS layer (CoW on each of the above host writes):
  - rewrites new blocks for each changed qcow2 region
  - plus tree metadata churn per txg
  => ZFS IO:                qcow2 IO + additional metadata amplification

Not every workload hits worst case every time (cache, sequentiality, preallocation, and snapshot state matter), but this is exactly why small random write workloads can get hammered when you put CoW-on-CoW.

So why does raw on zvol work?

Guest FS → (virtio-blk/scsi) → raw disk
  └─ zvol (block device) in ZFS
       ├─ ZFS still CoWs (integrity + snapshots + replication)
       └─ BUT you removed qcow2’s CoW metadata layer:
            no L2 tables
            no refcount blocks
            no qcow2 cluster allocator
  • qcow2 is always CoW-capable, but whether a given write causes a CoW-style allocate+remap depends on:
  • whether internal snapshots exist
  • refcount state (shared clusters)
  • preallocation mode (falloc, metadata, etc.)
  • whether clusters are already allocated and unshared
  • Even when qcow2 can overwrite an unshared allocated cluster, you still typically pay qcow2 metadata overhead and ZFS CoW on the modified file blocks.
  • ZFS’s own CoW + metadata churn is unavoidable; the point is to avoid stacking CoW allocators.

7 Likes