Proxmox: Everything You Wish You'd Known Sooner (especially for VMWare Refugees)

With oVirt/RHV gone, there are fewer contenders anyway, but Proxmox with Ceph finally ticks the HCI box, which was critical for me and the reason I went with oVirt originally… Ironic that HCI on oVirt died before the project itself went titsup.

I agree also on the community, the TrueNAS community was pretty narrowminded when it was still all BSD and neither the product nor the community seem to progress as fast as they’d have to in the current environment.

Xcp-ng has a lot of great people, but they didn’t appreciate that I believe it’s a dead horse still on a canter.

The main disadvantage of Proxmox was that it’s so primitive, doesn’t have a management engine running in the background with its own HA and recovery challenges.

But that’s also an advantage, especially if running a compute cluster isn’t your exclusive day job but simply a means to an end to reduce your effort vs not having an HA capable solution at all.

So today it’s mostly the best fit for my needs and I just pray they’ll resist letting them be bought up like SuSE, and continue to provide some urgently needed “sovereignty”.

1 Like

Just use the zabbix-agent.

I haven’t thought of using zabbix to monitor VM status. I just installed the zabbix-agent inside each VM. I don’t have any opinion on which of the available zfs templates to use (as I haven’t used zabbix in a while), but whichever you use, pick one that has the feature to monitor for faulted vdevs.

If the volume group name is not the same, you should see it in vgscan && vgs. If the name is the same, you may not see, but because the PV UUID are not the same, it shouldn’t be an issue.

You’d see a physical volume (or more if you are doing multi-partitions and merging them with lvm), a volume group and a bunch of logical volumes for VMs. And no, these would not be mapped to anything. You’d need to grab your VM templates from the prod server.

Like Wendell mentioned in the video, either use zfs with replication on each host or use ceph (that’s the most supported options, together with lvm-thin). Of course you have the option to go crazier with a chunky SAN (NetApp / Pure / 3PAR), but you’d miss on a bunch of features that you wouldn’t use (because proxmox doesn’t support them and has no reason to). In theory, you could leverage SAN level snapshots, if you use a LUN per VM, but it’d be a bit of a clusterf***.

But you could also use a SAN with just a large LUN dedicated to proxmox and slap LVM on it (zfs wouldn’t make a ton of sense on a SAN LUN, unless maybe you want to use zfs-send, but some of the large vendors, like NetApp, have built-in cloning options, like SnapMirror - if you didn’t want to pay for VMWare, you probably wouldn’t want to pay for NetApp either IMO).

IMO the next best thing for proxmox is to do a separate ceph cluster that acts as a SAN. Connect proxmox to those ceph arrays through a dedicated storage vlan and call it a day. With this, you’d get the most out of proxmox and you’d avoid some pitfalls of hyperconvergence (mentioned in this thread).

I wouldn’t be able to live with myself with that.

I think TrueNAS was better with FreeBSD. I don’t like appliances anyway, so I wouldn’t be using TrueNAS, but rather just plain FreeBSD, but I digress.

I posted on your earlier storage question here.

Here are my comments after watching the “7 important things” video.

I’m ever so glad we got SSDs these days (well actually, at least we did when I bought mine…), somehow stacks of spinning rust have dominated during the most formative decades of my career and that’s where ZFS held quite a few aces up the sleeve, while it was never HA without Lustre or similar. And then came SANs or redundant NFS filers managed by parts of my team at the time, so I never had as much hands-on with ZFS as I wanted.

Post retirement in the family-home-lab I still maintain stacks of big disks for bulk data while VMs or anything hot is SSDs, but still not under any type of serious load.

From those oVirt days I keep two clusters around that were switched to Proxmox a few years ago designed at the time for functional testing on the cheap, so one is actually Gemini Lake Atoms, the other G8-11 NUCs, with the Atoms basically on cold standby. What’s the usable physical minimum for practical function and failure testing, was the challenge.

That cluster runs mostly infrastructure services 24x7 for homelab and family and it’s all about not having to worry and plenty of time after the first warning, like a disk fault on RAID6.

Using Ceph based HCI wasn’t just the natural transition from oVirt, it also means that storage isn’t an extra problem domain with very different failure patterns, that I would have to manage in my head when I wasn’t looking for extra work.

But that doesn’t mean I keep all data on it, 3x write amplification with NVMe storage on a 10Gbit network is painful to think about, quite ok on functional VMs, but not everything else.

My CUDA workstations are mostly dual boot, ran Linux for work and Windows for gaming: today they are dual boot Windows and Proxmox (potentially with GPU pass-through VMs), but those won’t run 24x7.

So I have ZFS on their local storage and I use NFS and CIFS/SMB for the bulk data and simply have everything/everyone just use what’s most appropriate for the use case.

If you’re not big enough to afford really good storage at enterprise SAN level (and the brain capacity to manage that separately), I’d recommend anyone to at least consider Ceph for perhaps a critical functions and management cluster, it doesn’t close any door on using all the other storage technology where that’s more appropriate.

2 Likes

Actually, another one of my last projects was trying to aim even below the Atoms cluster to Raspberry PIs or similar, something like a very economical out-of-the box HCI cluster for use cases from Outback to Coalmine, Camper to Granny’s old farm, or just a black box you give to kids moving out with their gaming rigs, that ties them easily into your home-lab without compromising on security or managability…

Since the RP5 only went up to 8GB RAM at the time, I got an Orange Pi 5+ with 32GB to run the bigger share of the VMs (around €200 at the time). The OP5+ offers an M.2 slot for NVMe and “native” dual 2.5Gbit Ethernet, the RP5 took both a super fast Kingston DataTraveller MAX 1TB drive and 2.5 GB/s Ethernet on USB3, both matched my Atoms cluster well enough.

A great developer somewhere in Asia had already done the brunt of the work and ported Proxmox to ARM, so I could limit myself to the typical functional testing, which included trying to bring up VMs with Windows for ARM (to be used via RDP) as well as the various Linux variants. There I sadly missed Cendio’s great ThinLinc in an ARM server variant as the Linux RDP equivalent, btw. the OP5+ does quite a lot better even on a 4k screen than any RP5, none of them are Fruity Cult number chrunchers, obviously, not at €200 for 32GB of RAM and the storage of your choice.

Things worked really well, managing even mixed x86/ARM clusters was just fine, accessing CEPH or any other type of storage from the ARM nodes was flawless, only live migration was a bit of an issue, since the Raspberry and the Orange didn’t have absolutely identical CPUs and there are as yet no abstraction levels in KVM for ARM, as there are for x86, which have made live migration e.g. between AMD and Intel or different ISA generations much less painful.

And yes, live migration between ARM and x86 failed properly, just as “dead migration” of VMs with non-matching ISAs brought a proper error message when trying to launch them, but no crash: I just had to try :laughing:

Again, I’d love to have the ability to operate a lego brick HCI cluster that’s really cheap and fully passive solid state to boot-strap a resilient management environment, that other more powerful and more specialized pieces of equipment can then attach to or even install from. the very bottom-up seeding counter design to from-the-top raining and ruling clouds.

But Proxmox themselves don’t seem ready to maintain an ARM branch themselves and a single Asian dude can’t well carry that burden alone.

Their focus has helped them survive for a very long time, so we can’t really complain.

I did some more testing with FIO, using iscsi, nfs and cifs. For nfs and cifs I tried both raw and qcow2. It seems it was indeed caching as raw and qcow2 are pretty much the same. iscsi is a little better but I doubt it will matter too much on my 10gbit network.

It’s quite frustrating that cache can lead you down the wrong path. Is there a way to force cache to be off for testing? Would be helpful for benchmarking protocols.

As for creating a dataset that aligns with the filesystem, my Truenas instance warns me that 16k is the lowest I should be going. I’ll test out that on the weekend.

Thanks a lot for pointing me in the right direction

What is the keyboard shown in the video?

Missing in Proxmox: scale-out management, grouping, nesting, trees..

A frequent nuisance with HCI: management clusters and storage clusters are ill defined and often somewhat antagonistic.

Please note: I am recording this, while my memories with regards to vSphere, oVirt and Xcp-ng are quickly fading… I could be wrong on the details!

Problem: IT infrastructure needs more or less constant rebuilding, new pieced of hardware are moved in, old ones get removed, things are being upgraded etc., but you want to maintain continuity for your customer base, be they family or paying the rent.

So you may just want to be able to have something like an ‘old farm’ and a ‘new farm’, or a ‘production farm’ and a ‘DR/QA/Transition farm, distinct, if you can afford it’ and then be able to move live VMs or container between them, to prepare for major infrastructure upgrades, physical or logical. In other words: management domains and failure domains are distinct and failure domains might correlate more with storage.

That’s where I think Proxmox falls seriously short in that it only knows servers and clusters of servers, not collections of servers, nor collections of servers and clusters, nor does it understand storage clustering.

Since the only way to group servers in Proxmox is to cluster them, that’s what you do. But clusters need to be constantly up in order to work (share the Corosync ‘brain’), partitioning a cluster is a no-no, but that’s exactly what you may want to operate, groups of clusters where availability dependencies are constained to the members of a group, while visibility is between all e.g. to support migrations.

Perhaps put more simply, you may want clusters A and B and the ability to shift workloads and then flip roles.

In my case, I maintain a primary HCI Proxmox triple cluster with CEPH, with another such triple in cold standby, that are using very modest hardware (NUCs and Atoms) to run critical admin services like Univention IAS. The big home-lab work VMs run on workstations, that aren’t actually running 24x7, while most of them actually boot serveral distinct operating systems, Proxmox is just one of them.

Proxmox doesn’t like such volatility, but I trick it into acceptance by giving the core cluster higher numbers of votes, so a quorum is maintained, even if all those workstations are doing something else or shut down.

But that fails when it comes to managing two distinct clusters to differentiate failure or operations domains, i.e. to live migrate workloads between them.

BTW. oVirt wasn’t much better there, not sure about vSphere.

So why do I mention this? Because that’s the one area, where Xcp-ng did shine, I could manage distinct cells or availability groups from a single management pane and live migrate VMs beween them.

Given the Proxmox design based on corosync, I don’t see an easy way out of this and I’m sure the Proxmox guys are painfully aware of this: in the mean-time I’d like to know how relevant this sort of thing is for you and how you manage

I currently use a very manual approach, where I shut down and backup all VMs I want to move between my two clusters and then restore them on the other. Since the critical VMs are HA at the applicaton level, it doesn’t even cause service interruptions.

On the primary cluster I also maintain a ‘transition’ node with local storage, that allows me to temporarily hold a VM I want to keep in operation, even while I do work on the CEPH cluster. Here the abiltiy to move VM disks from CEPH to local storage on the transition node helps me de-couple that dependency.

There I wish I could run ZFS replication on top of CEPH or have CEPH support async replication itself… it’s so easy to want everthing without paying or contributing :slight_smile:

1 Like

Use Proxmox Datacenter Manager and just live migrate the VM directly like you want? I used to do the backup and restore dance but that hasn’t been required for at least a year.

3 Likes

The backend KVM supports live migration between any system. It’s just that Proxmox doesn’t provide a GUI for it. Corosync is only used in a single cluster. You could live migrate or clone it using the qemu kvm backend. This has been a feature available in virt-manager (with libvirt) for many years.

Then there’s proxmox datacenter manager, which strawberry mentioned. You have a single pane of glass where you can migrate VMs between separate hosts and separate clusters (the capability to live migrate should be there, but idk how pdm does that - note that sometimes different CPU types, assuming you aren’t using a predefined virtual CPU type, may require the VM be rebooted).

2 Likes

I know that KVM can do more, but Proxmox warns you against mixing their instrumentation with any other: they are completely unware of each other.

DataCenter manager: I haven’t looked at it yet, since somehow was under the impression that it was mostly a re-skin, not a functional expansion: I’ll have a look, thanks!

I’ve now looked at the DCM V1, just before 1.1. came out :face_with_raised_eyebrow:

It’s mostly nagware, cross-cluster migration the only actual function expansion I was able to discover.

Apart from the permanent nagging, it’s full of bugs, e.g. you can’t delete servers that are shut down or have in fact been migrated to a cluster in the mean-time.

I am worried about the direction Proxmox may be taking with DCM, which seems to point to complete refactoring of the management stack, but with the nagging turned up to intolerable or “denial of usability unless you pay”.

Corosync remains an architectural Achiles heel IMHO, especially when not all nodes are permanently up in your environment and you need to balance that with giving weighed numbers of votes. Additional nodes entering and leaving or moving within groups or clusters are normal life cycle events in any non-trivial operation, but not managed as such. E.g. for topolgy/membership changes all nodes better be up; if they are not, it’s non-trival to discover why they are upset when they try to come online again, and how to fix that.

No, the DCM doesn’t deal with that at all, or [hint, hint!] fix that automagically.

On top Proxmox and Ceph cluster menberships are completely ignorant of each other, which means your head needs to manage the gap, potentially after being woken up in the middle of the night: ok for hobbyist use, but Proxmox aims higher, so they must earn it, too.

1 Like

Apart from the HCI cluster running Ceph for the mission critical services like IAS (Univention based for NextCloud integration), I also run some dual-boot workstations with local storage used for VM/ML/CUDA experiments. And more recently I’ve added boxes I use to develop a “universal appliance”, which includes ZFS storage and replication which will spread across “sites”, basically the new homes my kids are building and which is intended to provide some type of family cloud where important data is replicated across while less important data can be accessed/streamd securely across.

The main issue is that I am trying to make this work across Windows and Linux desktops and Linux servers as transparently as possible and that means CIFS/NFS working pretty near identically with identity and access rights management across the file share space.

That turns out to be the real nightmare, mostly because there is way too many players and options on the Unix side while just shoving CIFS/SMB onto/into/under Linux doesn’t seem to work, either. Anyhow, that’s separate story, mostly I wanted to report on something positive: the joy of VirtioFS.

So far I had resorted to passing hard disks from the Proxmox host as raw devices to the Unvention file servers, which best understand user IDs and permit central management of shares und their UI.

Obviously that wasn’t going to work with RAIDZ ZFS volumes the host was also going to use for VMs and whatnots, not give away exclusively as a pass-through to a VM.

That’s where I tried VirtioFS, which oddly is tied to the hypervisor, but I’m not complaining since it works.

Well, most of the time, because obviously I was immediately going to share the ZFS volumes passed through to the Univention VM file server via CIFS and NFS4, right?

“stale file handle” is the result, evidently that’s by design somehow, nothing easily fixed at least on the NFS side. It seems to work well enough for Windows clients, where the performance is also rather impressive or just plain native.

I’m still a bit scared about running ZFS on a system without ECC RAM, the Zen 5825U in theory supports that, the Topton board has options for it in the BIOS, but the SO-DIMM sockets might simply miss the circuit board traces or the socket not support it… in the end it just plain doesn’t work. I mean the RAM works as RAM, but no tool can find ECC turned on, even if enabled in the BIOS and supported by the DIMMs, unlike on my normal slew of X570 and X670E boards, and rather more normally the Xeons.

Which is why I’m still relying on hardware RAID controllers for the primary storage, those univeral appliances will mostly act as local caches so bit-rot ideally should be contained.

That’s where I am also trying to integrate Syncthing, to control a bit better as to what’s going on, I’ve never actually worked with ZFS log shipping yet, and not sure it will offer the type of control and granularity that I want.

But Syncthing also adds an extra UID and access rights layer and translation issue, when it crosses Windows/Unix boundaries..

Sorry for the rambling, but Wendel called out.

1 Like

That means you should only use CIFS. Don’t mix-and-match, that’s never going to end well. You should just samba mount the share on all your linux servers.

I much rather prefer NFS, but if windows gets involved, I’m absolutely switching everything to samba.

1 Like

That seems sage advice, thank you!

I’m not sure if I can make it work e.g. with NextCloud and Syncthing, but it seems a better way forward already and not the outright dread I was facing so far.

Hi All

I have been working on migrating from VMware to Proxmox and have been playing around with it for about 2 months now, but I’m still feeling a bit lost on the storage side of things. In VMware, it was really easy and straightforward to set up and get working, but with all the options that Proxmox gives you, I’m feeling lost and not really sure what the best way is to do it.

We have a 3-node cluster running on Dell PowerEdge R440 servers with 2 x 4**-**port 10G Nic cards and a 40TB Dell PowerVault ME4024 SAN.

In VMware, we had the SAN attached to the 3 nodes via iSCSI, providing 2 x 20 TB datastores for virtual disks. I don’t have any budget to replace the storage hardware at this time, so I have been trying to get it working with Proxmox. Whoever we would like to keep the ability to thin-provision virtual disks for VMs and to take a VM snapshot before installing an update.

In my research, I found you can do Clustered QCOW2 over iSCSI, which I managed to set up and create a file system on the attached LUNs, but I ran into issues with ISOs loaded on one node not being visible on the other 2 nodes. When I tried to fix the cluster config so it would sync correctly, I somehow broke the storage cluster and haven’t been able to restore it. So I really don’t want to proceed with this option, as I’m not happy with its stability in a production environment.

From what I understand, we can connect the SAN to Proxmox using the standard iSCSI multipath method and mount the LUNs as LVMs, but then they work as block-level storage, and we can’t thin-provision virtual disks for VMs or take snapshots of VMs.

So if my only option is to continue using our 40TB Dell PowerVault ME4024 SAN as shared storage for the cluster, what’s the best, most stable way to set that up, and is it possible to still enable thin-provisioned virtual disks and take VM snapshots before installing an update?

Secondly, when we do have a budget to replace the storage, what’s the best, most stable way to implement shared storage for the cluster?
ZFS sounds cool, but looks a bit complicated, so I’m wondering if NFS may be better. We’re a small team and time-poor, so it needs to be as set-and-forget reliable as possible, like VMware was.

Thanks.

I’m not sure how vmware handled 2 iSCSI LUNs shared between 3 hosts (probably their overlay stuff with their datastore). Proxmox only gives you the raw storage and it’s up to you to configure them (which is what Wendell kinda warned in the video IIRC).

You could have the SAN split into 3 LUNs, 1 for each host, format them with LVM and make it a Thin Pool volume group. That way, vdisks will be thin-provisioned.

ZFS or Ceph would be ideal, but because you don’t have direct-attached storage (and you’re using a SAN), then thin-LVM is probably the best option you have.

In my own experience with proxmox in production, I unexpectedly found that qcow2 on NFS was the best option (for our situation): all hosts share the same NFS servers (we had multiple NAS’es), so they all have direct access to the VM vdisks.

That made live-migration incredibly fast (only transfer the RAM contents - the qcow2 disk stays in the same place and doesn’t get copied between hosts). It was incredible to upgrade the infrastructure (live migrate VMs around, upgrade the host, boot it back up, move VMs back along with more VMs, upgrade the other host, reboot it back, move VMs from the last couple of hosts on those two, upgrade, reboot, move VMs back - VMs always stay up).

And mind you, we were using 2x 1Gbps LACP for both the storage and for the corosync network (separate LAggs, so 4 gigabit ports per host, minus IPMI and out-of-band web + ssh management ports, we had at most 6 ports used).

iSCSI (or direct FC connection if you can do that) with 3 LUN (1 per host) formatted with LVM-Thin would probably be faster than NFS (iSCSI would probably only be marginally faster) in day-to-day performance, but then if hosts fail, you’ll have to do fencing or other HA stuff and move the LUN to a working box, then launch the VMs in the new host.

But with NFS, all hypervisors can access the vol where the qcow2 files are. Technically you’re supposed to do fencing with NFS as well and kick the failed host off of the NFS server, maybe by reloading the NFS export to not allow the failed host’s IP to mount the qcow2 volume.

But I’ve never found myself in a situation where hosts stopped communicating with each other, but not the NFS server (part of the reason is that we were just lucky and didn’t have instability, but another thing is that we didn’t have redundant network cards - sure, we had a minimum of 4 eth ports per hypervisor, but they were on a single pci-e card, so if the slot failed, the whole host goes offline, while the LACP per 2 ports accounted for single ports on the host or on the redundant switches going down).

So, Proxmox recently added an option that let’s you do snapshots on LVM shared storage. Roadmap - Proxmox VE. We have used this in production multiple times and it works really well.

The caveat is this doesn’t do thin provisioning (kind of). LVs are allocated thickly, but qcow2 writes data thinly. This option works best with a SAN / Storage system that can allocate LUNs thinly (ideally with Discard / UNMAP support). So a 1TB snapshot would look like its using 1TB in Proxmox, but on the SAN / Storage it would barely register until writes happen.

vmware has a shared disk filesystem called vmfs, so the iscsi luns are formatted with that.

On linux there are a couple cluster filesystems, ocfs2 and gfs2.

Neither is natively integrated in Proxmox, somewhat baffling since this would allow easy migration of existing vmware clusters.

You can set them up manually of course.

No its not. Thin lvm is not usable for a shared block storage.

The only thing ive seen for replicating vmware setup in Proxmox without just going full manual install of ocfs2 or gfs2 setups is to use Starwind storage plugin. It is not vendor specific so it works with whatever block storage. Never used it myself, see I Almost Gave Up on Proxmox iSCSI Storage in My Home Lab Then This Worked - Virtualization Howto

And here the official page https://www.starwindsoftware.com/resource-library/starwind-x-proxmox-san-integration-configuration-guide/

1 Like

I was not talking about shared block storage, but single LUN (I explicitly said 1 LUN per host, not shared LUNs) and I would think you would figure that out, when I was talking about NFS instant live-migration (aside from RAM) vs having to transfer data between LUNs. Yea, with lvm thinpool you’d have to move the vdisks from 1 host to another, for live migration.

In the case of a complete host fault, you would fence the LUN in the SAN (to prevent the faulted host, when coming back online to access the same resource) and attach the LUN to another host (the one that plans take over the hosts, the one with the higher corosync vote).

GFS2 works better at vm-level than hypervisor-level. At the hypervisor (leave aside the fact that proxmox doesn’t support it, even though debian would), you’d still need to do fencing on gfs2, to prevent the faulted host from accessing the shared resource.

Ceph stuff

But Proxmox has some idiosyncrasies you should be aware of, in particular in split-brain scenarios and it’s a thing that’s not been accounted for even in their own ceph implementation.

You get into a situation where 1 host goes offline in corosync, but still runs the VMs and has access to ceph (and that’s applicable to a SAN as well). It’ll keep running those VMs (not try to shut them down or suspend), while on the online corosync network, other hosts will have replaced the (partially) faulted host and start VMs on other hosts. Guess what happens next: boom, 2 instances of the same VM try accessing the same vdisk.

With a SAN (even a ceph SAN, assuming proxmox is only a ceph initiator), you can fence off the faulted (or partially faulted) host. In hyperconverged ceph, you can’t fence off a partially faulted host (because it’s part of the ceph cluster and has direct access to itself). But you’d still have the same problem with GFS2 (need for fencing).

I’ve heard of people “fixing” the problem of partly-fauled hosts by using a smart PDU and killing power to the host, but that just sounds really dirty to me.

Never heard of it, but it sounds like it might work for you. Cool stuff and good find.

Mindless ramble about clustered hypervisors and the problems with shared resources that cause contingency

I gotta be honest here. That’s just the old way of doing things and I don’t find hypervisor HA that great anymore. Applications should be resilient themselves. You do that and you don’t need hypervisor HA or live-migration for that matter and you won’t even need to think about sharing resources between hosts.

Everyone fleeing from vmware should start looking at modernizing their software stack and only try to keep hypervisor-level stuff limited for legacy applications. Why would you need to run a particular VM on a clustered (fault-tolerant) SAN and when a compute node crashes, have to go through the mental workflow gymnastics of segregating the faulted node and starting the VM on a different host?

For the same task have 2 VMs (or more) on 2 hosts (or more) on at least 2 different storage boxes. A VM running postgres went down? Just fail-over to (one of) the standby postgres server(s) on a different VM. Nginx VM went down? Even easier to move to another VM (and that’s assuming you aren’t already using k8s which can do a lot more with a lot less resources).

This kind of setup makes it irrelevant whether the VM itself crashed, the hypervisor crashed, the corosync link crashed (and in fact, you can run this on independent hypervisors w/o clustering, since the VMs talk to each other; the hosts don’t have to worry about shared resources at all), the VM network link crashed on one of the hosts, all the FC or iSCSI paths failed, or one of the storages failed.

Bonus points for this: your storage crashes or experiences degradation, just fail-over to the the other storage (at the VM’s OS HA level) and you may shut down all VMs associated with this box, while you work on rebuilding it. If you have the option, have the storage replicated to another box and re-launch the VMs with the other (replicated) backend storage.

Once it’s fixed or replaced, start the VMs and have the OS layer begin the replication from the active nodes to the new standby, freshly turned on.

Even updates becomes easier. Update the standby(s), reboot, have it continue replicating until it catches up, fail-over and make it the primary, update the now new standby, reboot. You may fail-over based on resource availability (to free up one host and storage and have the other pull a bit more work, instead of just sitting idle and just do writes to the storage that gets replicated from inside the VMs).

I am aware this is very opinionated, but I believe this to be better than shared resources on the hypervisors, with the only exception being NFS, because all hypervisors are equal in that regard. But I also don’t really like assigning storage to the VM at the hypervisor level and much rather prefer that a VM obtains its own disks straight from the source (if it can and if I can help it).