Proxmox Storage Replication?

hello community!

Ceph is cool and all but what are people actually using for storage replication with proxmox?

im hesitant to recommend ceph for a 3 node cluster. zfs replication and native replication (i.e. MySQL or postgresql replication by having each node have one db instance vm that’s pinned to that hardware) it seems can do what most people would need anyway.

but then there’s the whole linstor path… and what else?

1 Like

I had a 7 node hyperconverged Proxmox Ceph cluster for a while. Each node had a Toshiba XD5 2TB M.2 and a 10TB HDD. Most of my performance issues were due to the erasure code pools with cephfs I had running on the HDDs. I always got decent performance from the SSD pool, with both 3 copy RBD pools and some wider erasure code RBD pools performing not quite as good. I eventually scaled the cluster down (3 nodes), and switched back to a ZFS box for bulk storage. There was a small period at the beginning where I was using Samsung consumer SSDs with Ceph, that was a mistake.

I’ve found this Ceph performance guide to be helpful: https://yourcmc.ru/wiki/Ceph_performance. One of the biggest takeaways is to use SSDs with power loss protection.

The author of the above guide ended up writing his own SDS, Vitastor. I’d like to test it out, but I haven’t had the chance to look into it yet. https://vitastor.io/en/index.html#/

1 Like

So to start off, I am the crazy person using Ceph at home because it’s stable and scalable, if a little (okay, maybe a lot) slow. Two tiers, NVME, and HDD w/ db+wal on SSD. I think Wendell is right here, it’s hard to recommend Ceph if you’re not going big because Ceph really wants to go big. It wasn’t until I got into the 6+ node, 30+ OSD range that it got to truly usable speeds and even this is a bit low performance for what a lot of us would like.

Before going off on the tentacle-y adventures, when I only had a few machines, the go to was ZFS on NVME with snapshot replication every 15m. On regular consumer “bought these at Microcenter” m.2 SSDs, this replication only took on average about 4-6 seconds for each VM. Proxmox does a good job of keeping the snapshot sync schedules on track and live migration works, only marginally slower than with Ceph because you have to wait for the replication every time. It’s hard on the NAND, but so is Ceph.

My non-prod machines actually still do replicated ZFS because ZFS on an m.2 is easily an order of magnitude faster than Ceph, even backed by NVME, even with 6 nodes and 50gbit on every node. And sometimes, you just gotta go fast.

Honorable mention to TrueNAS, having the option to move bulk data onto ZFS outside the cluster entirely has saved my bacon on many occasions. Before Proxmox Datacenter Manager existed, it was the only real way to move a VM across cluster boundaries or onto a standalone, non-clustered node.

It’s funny that Wendell mentions native service/database clustering, with pinned VMs because even with Ceph backing the storage, I have plans to do this for select services in the future because you can have HA in addition to active-active load balancing, which is in my opinion, “The Way”. As soon as Jellyfin adds external database support, this will be my next project.

2 Likes

For me the big problem with a three node Ceph cluster is uptime. One node dies while you’re doing maintenance on a second? Good bye VMs! So I like to have at least 4 nodes, so you can keep 3 nodes on while during maintaince.

ZFS replication (through proxmox) is nice for an emergency, but I never enable auto HA. I’d rather deploy critical apps with redundancy and only recover from replica if original image is unrecoverable. Doesn’t make it viable for big customers or critical systems

Though now with Proxmox 9 having SAN support I’m recommending 3 nodes + SAN. Just makes things simple.

Really wish there was a more flexible data replication system like VSAN for proxmox, so you could do 2 node + witness. Then you could run on one node in a pinch.

Better question, do you even need replication? Replication eats up a ton of resources and isn’t really needed for most people. Proxmox doesn’t need replicated storage to work and you can even do live transfers even though it does take a bit more time.

How you manage data is more interesting. For larger datasets like movies and photos I use a virtualized TrueNAS that has two 2TB SSD’s connected via a separate controller that has PCIe passthough enabled. This allows me to have small (ish) VMs with the larger data sets that are mounted via NFS. I think the important thing to realize is that spreading VMs across nodes is much heavier than using a lighter weight container based solution. With containers you just have a volume with data instead of all the overhead with VMs. In short, think cattle not pets.

From a hosted software perspective, I run everything in Podman Quadlets. I have two VMs with one being for stuff that doesn’t take much storage like Radicale and Forgejo and the other for stuff that has large amounts of data tied to it like Immich. My Proxmox cluster is 3 nodes and in total I have about 9 VMs and containers running various things like my storage and my desktop (vfio).

What I’m working on is getting a full IaC (Infrastructure as code) setup going where I manage everything with Ansible. Right now I’m still creating and managing my containers manually but I would like to have a setup where Ansible deploys everything and manages resources like DNS. To get to that point I’m looking to deploying Garage so that my data can be pulled to the VM via rclone and then backed up once a day while the VM is active. I’m using Forgejo and Woodpecker CI to run my playbooks so in theory I could get to the point where everything is managed via git.

TL;DR
I think you would be much better off making a video about container/application storage solutions instead of VM storage solutions. I’m planning on just using local storage with some automation to pull/push to S3 via rclone. I haven’t set it up yet but I’m looking into garage: https://garagehq.deuxfleurs.fr

Any clustered/replicated storage solution would need to be very lightweight and portable for me to consider using it. It would need to fit within 512mb of ram and not require much in the way of dedicated hardware. Ideally it should be trivial to either mount as or sync with container storage.

I’ve also toyed briefly with the idea of using Kubernetes and longhorn but Kubernetes is way overkill for the time being.

I’m just using rsync over ssh instead of s3 as it is much simpler to maintain

1 Like

I’ve run oVirt on GlusterFS as a hyperconverged infrastructure to support international corporate research teams for many years in a variety of settings, some simple some very complex.

The HCI concepts more or less introduced via Nutanix and then copied by vSphere and others were extremely attractive, because there were plenty of old servers with local storage, mostly 10 and some 100Gbit/s networks, but no SAN or IF fabrics available.

The basic setup was always a 3-node cluster, the minimum for automatic fault recovery, but often with the 3rd node only used for a quorum, which GlusterFS supported. Somewhat less write amplification, less storage needed overall, but more complex in the operator brain, never good in a critical environment.

For kicks and because many of the expected data sets were highly compressible ML data, I also worked with VDO (dedup+compression) as well as mixed HDD and SSD storage.

For the home lab, which did all the functional testing QA, I typically just went with full 3-way replication, because the oVirt part was complex enough and I didn’t want the extra mental complexity of the arbiter-only nodes I had to use in production to save disks.

When IBM/Redhat killed that project, I switched to Proxmox/Ceph simply because there was no alternative and I’d come to love the non-stop normal operational mode of 3-node HCI.

Ceph was and is more obscure than GlusterFS in my mind, on the other hand it has just worked and a lot better than GlusterFS ever did. Yes, gaining insight into GlusterFS and fixing things is easier in theory, because there is always a normal xfs file system at the base (using extended attributes for the cluster meta-data), but I also needed to intervene often enough.

Should anything ever break in a Ceph cluster, perusing the manual for fixing was enough to convince me that I’d never even try. For me a broken Ceph cluster is a disaster scenario, not a fault tolerance case. So I’d have to restore things from backup on a newly wiped ceph storage volume.

In the mean-time, Ceph hasn’t failed me, it’s always come back, even on triple node power failures, but also with only a light load on the cluster.

As with any cluster, it puts the responsibility on the fabric between them, it needs to be way more reliable than the nodes, otherwise you get a very expensive random number generator.

Everbody tells you you need to separate East-West (replication and migration) and client networks, what you make of that in a home lab is your fault.

One issue with GlusterFS was that while you could configure every permutation of error correction to data bits for redundancy, scaling from 1 via 1+1+A to any of the more advanced redundancy codes wasn’t supported. I guess it’s easier in Ceph, but it’s only home-lab now, so all tests beyond the 3-node cluster were only done via nested virtualization on VMware: that doesn’t give you a lot of insight into practicality or performance. 3-node full write amplification has a real cost, but pays in simple.

Think as early as possible how your cluster could grow, in member size, in number of clusters, in terms of disaster resilience. I have no idea if that old Nutanix legend, where you could just add bricks and rebalance ever worked as well as they advertised, it’s certainly never been true for the cheap knock-offs, vSphere, RedHat (RHV/oVirt) or Xen-Server/XCP-ng.

Investigate your use case! Test the migration/scaling in a lab, aim as close to the expected capacity as you can afford from the start, growing the machines not their number is far easier.

You don’t need to spread Ceph and storage across all Proxmox compute nodes: those can be “stateles” compute only. But if you also use that for capacity management (e.g. turning off unused nodes), you may need to adjust your Proxmox internal quorum vote allocations (wonder how I learned?), so that alway-on (Ceph members) nodes get more votes vs. standby (no Ceph) machines.

Functionally, I can only recommend that you test these things via nested virtualization first, much less effort, much faster results. And not everybody has access to hardware like Wendel.

What oVirt supported and where Proxmox is far more simplistic in design is real policies on how related VMs should be monitored, and managed automatically, e.g. load based migration that would maintain software clusters on different hardware nodes, starting stand-by systems or rebooting failed VMs on a spare host.

That complexity didn’t come for free on oVirt, so in the home-lab that may be a blessing and saves a lot of RAM for the management engine and its agents.

I’ve always thought it crazy that two companies like Proxmox and Linbit, who seem to be within walking distance in Vienna evidently don’t offer a tightly integrated product when Proxmox absolutely needed a native redundant storage, but that’s Germans and Austrians for you…

With the integration of Ceph in Proxmox, that no longer matters.

So long story short: I’ve had an excellent experience using Ceph with Proxmox as a HCI solution. It works, and it’s exceptionally easy to use.

It manages to deal with single node faults (mostly auto-recoverable) and single node maintenance really well, dual node failure is a disaster and you need to be ready and trained for both: they are not the same thing.

I repeat, they are not the same thing: you need to design and plan your disaster recovery separately and it may require as much of its own redundancy and resilience as you require for your use case, potentially several clusters and spread across the globe.

In a disaster case you restore on blank hardware from a backup and it’s not a fully automated response in all disaster scenarios.

Application vs. system cluster: they are just not the same thing, even if you could layer them onto each other. I’ve done plenty of Oracle MAA in my career and designed and run hyper critical national electronic payment schemes across Europe. Ultimately you can’t escape the CAP theorem proposed and proven by a Google CTO Eric Brewer and only the application can decide which two of the three qualities it prefers: the OS and the hypervisor are further away and thus more prone to follow a static preference.

ZFS replication: To me that’s a backup mechanism, ideally bordering on fault tolerance, if you add some agency on top. But in the case of database files or similar, transactional control on the replicated data is lost. Oracle’s Golden Gate replication is an expensive option for a reason.

It’s a little sad that Ceph doesn’t seem to have async replication yet (GlusterFS at least had it in theory), nor native CIFS/SMB support. Then again, complexitiy kills…

BTW I do like using CephFS (file system layer on top of the native Ceph block abstraction) in some cases, even if it’s not recommended for performance reasons. But it eases capacity management.

I run Univention UCS on top of my Proxmox and NextCloud integrated and managed by it, to get Windows and Linux file services and a unified user management including NextCloud for everything (you may still want to segregate management and user accounts).

That was a bit of a bother for a while, UCS doesn’t support Debian well enough, but if you’re looking for a jack-of-all-trades solution, those two will you get far beyond anything TrueNAS dares to dream about.

1 Like

Clusters need a quorum with a clear majority. A four node cluster may end up with a split vote… you always need odd numbers (and a greater network to be sure).

You don’t need odd number. You need at least 3. With 4 you still only support 1 node down, like with 3. With 5 and 6 nodes it handles 2 nodes being down, etc. Since lost quorum locks the cluster, nothin else happens except your cluster fails to run.

1 Like

I guess that’s true and I might have been misdirected by deliberations done long ago with regards to HA environments, that included only a single physical path between two “availability zones”: some CIOs were still trying to save cost…

Nobody would to that in these days of cloud, back then a link failure resulted in both HA pairs going down, while using a 3+2 setup made sure the left side would continue to process transactions.

I don’t do any HA stuff so I don’t really need the real time nature of it all, but I do use zfs send / recv to duplicate the data when needed. It is very sleek. I love how it just does the incremental blocks since the last common snapshot.

1 Like

After your video on Proxmox and a Ceph cluster, I was expecting a new top level thread…

But there is nothing here!

Clusters are tricky, no doubt. They rely on every communication path being more reliable than any device, and that includes a) tight response time windows b) proper fencing, otherwise you get an expensive random number generator.

So yes, if any device goes hunting butterflies without responding, you need a proper time-out instance to stop the wait and then you’ll need to cut that device off, reliably.

If it’s a network link, it needs to be severed, if it’s a storage device, it needs to be taken offline.

And then you need enough survivors to continue operations and a well tested prodedure to heal. They aren’t a freeby.

The main issue with clusters is that they are used for scale-out and for resilience. But you don’t just get both for free. You need to clearly define what your primary objective is: scale-out or availability.

And with HCI the other issue is that you’re often running two totally disjunct clusters, here it’s Proxmox and Ceph. On oVirt it was Gluster and, well the original oVirt cluster. They don’t fail the same, if you can avoid mixing them, you actually should (more below).

If it’s redundancy, your availability goes down the drain because error probability multiplies. So cluster redundancy starts with a debt and needs a triplicate as a minimum only to compensate that automatic inital hit on reliability.

If you think you should get more capacity or performance out of a triple than a single, you’re already falling into a trap: all you get ideally is fault tolerance against a single fault in storage (which ZFS replication doesn’t deliver). For the Proxmox or VM side of things, the best you can hope for is automatic restart. (There used to be Marthon systems, but that’s another story. And Non-Stop is still a little pricey.)

If it’s (storage) scale-out, you should ask yourself if more systems are really better. Are you trying to reach aggregate bandwidth ceilings a lesser number of storage nodes are unable to achieve? Then that still comes typically with the cost of failure probability: if you have enough storage nodes, failures become a certainty and your time-outs and fencing need to match.

HCI has this irresistable LEGO charm, just add blocks with CPU, RAM and storage and bigger=better! But Nutanix marketing totally oversold a promise that doesn’t really work. That’s why none of the hyperscalers do that: they scale compute and storage separately. In fact they never put both of the into a single chassis. Instead they put a perfect fabric between them, which understands 100% what those nodes are doing and differentiates all network and storage flows with timeout controls and fencing.

Unfortunately they don’t sell that fabric and the software and it doesn’t come in tiny little packages so you can use it at home. The smallest you can do is typically a triple node HCI.

And that’s why I’ve put my IAS and my NextCloud servers on a Proxmox triple, using very modest machines, production runs on a set of NUCs (gen 8, 10 and 12) with 10Gbit, while QA runs on three N5005 Atoms with 2.5Gbit. I accept that three-way replication has them go 1/3 on writes, while the networks are far too slow for what the SSDs could do on reads.

What I get is a) storage fault tolerance for a single node defect and b) automatic restart for VMs, that might have run on a host. I means I can be on a trip while the home lab keeps runnning and typically fix things later.

Power failures obviously aren’t allowed to happen, and that absolutely includes the network, which, as I said at the beginning, needs to be way more reliable than the nodes.

And that is already the case for Proxmox itself and it’s cluster.

I did this mostly, because I used to run a bigger research lab with dozens of people working on a somewhat bigger infrastructure then still using oVirt, which has since been killed off. There it did rather well, survive critical hardware faults, which would have resulted in loss of data and outages otherwise. But it was also very nearly a full job to keep it running well.

The simplest Proxmox triple is much less complex, but hard to recommend if all you want is some resilience, unfortunately.

I guess you could build some Ceph augmented “hyperscaler light” switches, that have compute nodes with just enough smarts to boot off them and allow you to plug in drives.

Proxmox could perhaps use something making it simpler to run for minimal setups, but having it go into managing storage, may not be ideal for consumers, schools, clubs, non-profits etc.

Mind you, the company sells support, so I’m not sure how easy they dare make it. Not that it’s really ever easy underneath the hood.