Proxmox Storage Replication?

I’ve run oVirt on GlusterFS as a hyperconverged infrastructure to support international corporate research teams for many years in a variety of settings, some simple some very complex.

The HCI concepts more or less introduced via Nutanix and then copied by vSphere and others were extremely attractive, because there were plenty of old servers with local storage, mostly 10 and some 100Gbit/s networks, but no SAN or IF fabrics available.

The basic setup was always a 3-node cluster, the minimum for automatic fault recovery, but often with the 3rd node only used for a quorum, which GlusterFS supported. Somewhat less write amplification, less storage needed overall, but more complex in the operator brain, never good in a critical environment.

For kicks and because many of the expected data sets were highly compressible ML data, I also worked with VDO (dedup+compression) as well as mixed HDD and SSD storage.

For the home lab, which did all the functional testing QA, I typically just went with full 3-way replication, because the oVirt part was complex enough and I didn’t want the extra mental complexity of the arbiter-only nodes I had to use in production to save disks.

When IBM/Redhat killed that project, I switched to Proxmox/Ceph simply because there was no alternative and I’d come to love the non-stop normal operational mode of 3-node HCI.

Ceph was and is more obscure than GlusterFS in my mind, on the other hand it has just worked and a lot better than GlusterFS ever did. Yes, gaining insight into GlusterFS and fixing things is easier in theory, because there is always a normal xfs file system at the base (using extended attributes for the cluster meta-data), but I also needed to intervene often enough.

Should anything ever break in a Ceph cluster, perusing the manual for fixing was enough to convince me that I’d never even try. For me a broken Ceph cluster is a disaster scenario, not a fault tolerance case. So I’d have to restore things from backup on a newly wiped ceph storage volume.

In the mean-time, Ceph hasn’t failed me, it’s always come back, even on triple node power failures, but also with only a light load on the cluster.

As with any cluster, it puts the responsibility on the fabric between them, it needs to be way more reliable than the nodes, otherwise you get a very expensive random number generator.

Everbody tells you you need to separate East-West (replication and migration) and client networks, what you make of that in a home lab is your fault.

One issue with GlusterFS was that while you could configure every permutation of error correction to data bits for redundancy, scaling from 1 via 1+1+A to any of the more advanced redundancy codes wasn’t supported. I guess it’s easier in Ceph, but it’s only home-lab now, so all tests beyond the 3-node cluster were only done via nested virtualization on VMware: that doesn’t give you a lot of insight into practicality or performance. 3-node full write amplification has a real cost, but pays in simple.

Think as early as possible how your cluster could grow, in member size, in number of clusters, in terms of disaster resilience. I have no idea if that old Nutanix legend, where you could just add bricks and rebalance ever worked as well as they advertised, it’s certainly never been true for the cheap knock-offs, vSphere, RedHat (RHV/oVirt) or Xen-Server/XCP-ng.

Investigate your use case! Test the migration/scaling in a lab, aim as close to the expected capacity as you can afford from the start, growing the machines not their number is far easier.

You don’t need to spread Ceph and storage across all Proxmox compute nodes: those can be “stateles” compute only. But if you also use that for capacity management (e.g. turning off unused nodes), you may need to adjust your Proxmox internal quorum vote allocations (wonder how I learned?), so that alway-on (Ceph members) nodes get more votes vs. standby (no Ceph) machines.

Functionally, I can only recommend that you test these things via nested virtualization first, much less effort, much faster results. And not everybody has access to hardware like Wendel.

What oVirt supported and where Proxmox is far more simplistic in design is real policies on how related VMs should be monitored, and managed automatically, e.g. load based migration that would maintain software clusters on different hardware nodes, starting stand-by systems or rebooting failed VMs on a spare host.

That complexity didn’t come for free on oVirt, so in the home-lab that may be a blessing and saves a lot of RAM for the management engine and its agents.

I’ve always thought it crazy that two companies like Proxmox and Linbit, who seem to be within walking distance in Vienna evidently don’t offer a tightly integrated product when Proxmox absolutely needed a native redundant storage, but that’s Germans and Austrians for you…

With the integration of Ceph in Proxmox, that no longer matters.

So long story short: I’ve had an excellent experience using Ceph with Proxmox as a HCI solution. It works, and it’s exceptionally easy to use.

It manages to deal with single node faults (mostly auto-recoverable) and single node maintenance really well, dual node failure is a disaster and you need to be ready and trained for both: they are not the same thing.

I repeat, they are not the same thing: you need to design and plan your disaster recovery separately and it may require as much of its own redundancy and resilience as you require for your use case, potentially several clusters and spread across the globe.

In a disaster case you restore on blank hardware from a backup and it’s not a fully automated response in all disaster scenarios.

Application vs. system cluster: they are just not the same thing, even if you could layer them onto each other. I’ve done plenty of Oracle MAA in my career and designed and run hyper critical national electronic payment schemes across Europe. Ultimately you can’t escape the CAP theorem proposed and proven by a Google CTO Eric Brewer and only the application can decide which two of the three qualities it prefers: the OS and the hypervisor are further away and thus more prone to follow a static preference.

ZFS replication: To me that’s a backup mechanism, ideally bordering on fault tolerance, if you add some agency on top. But in the case of database files or similar, transactional control on the replicated data is lost. Oracle’s Golden Gate replication is an expensive option for a reason.

It’s a little sad that Ceph doesn’t seem to have async replication yet (GlusterFS at least had it in theory), nor native CIFS/SMB support. Then again, complexitiy kills…

BTW I do like using CephFS (file system layer on top of the native Ceph block abstraction) in some cases, even if it’s not recommended for performance reasons. But it eases capacity management.

I run Univention UCS on top of my Proxmox and NextCloud integrated and managed by it, to get Windows and Linux file services and a unified user management including NextCloud for everything (you may still want to segregate management and user accounts).

That was a bit of a bother for a while, UCS doesn’t support Debian well enough, but if you’re looking for a jack-of-all-trades solution, those two will you get far beyond anything TrueNAS dares to dream about.

1 Like