Proxmox: Everything You Wish You'd Known Sooner (especially for VMWare Refugees)

Outline

  1. Storage – it’s far expansive and flexible and that means there can be landmines for the inexperienced.
  2. Linux – ESXi is locked down. Linux under the hood is not nearly so.
  3. Backup & Disaster Recovery – Actually Easier than vmware thanks to PBS.
  4. Proxmox Datacenter Manager: Not vCenter — Yet — But You Should Be Using It
  5. Proxmox SDN: Powerful, Young, and Very Linux
  6. Memory Deduplication Is On — and You Control It
  7. Community – What’s YOUR lesson learned?

Storage – The Big One

Understand that the storage layer in Proxmox are usually Linux primitives: ZFS, LVM/LVM-thin (dm-thin), mdraid/lvm-raid, CEPH or possibly NFS/ISCSI

If you’re coming from VMWare, you might be familiar wtih VMFS, VSAN and how storage works there for hyperconvergence.. or perhaps in a general SAN type context.

The first thing to understand is that ZFS is big. And different. It’s not a file system. It’s a full storage stack that does end-to-end checksums (data+metadata), scrubs w/self-heal (less like LSI raidcard scrub more like what IBM or Pure Storage implement in their enterprise products. ZFS has excellent snapshots and replication built in, and is tried-and-true copy-on-write. ZFS (and many of the technologies here) are old open-source technologies.

There are pitfalls with ZFS: There are lots of scenarios where writes can be amplified and this is often the source of significant I/O pressure/stalls and poor performance after migration from another virtualization platform.

As an administrator, one must come to understand some aspects of their storage geometry to understand the failure modes and diagnose performance issues.

VMFS mindset: “My array is sacred; ESXi is a consumer.”

ZFS mindset: “My hypervisor is now part of the storage trust boundary.”

For example, qcow2 virtual machine files on top of ZFS is a double copy-on-write scenario. Instead, a raw zvol block storage device is (generally) best for most workloads.

See also:

ZFS has vastly different performance characteristics depending on if it is RaidZ vs a mirror (or striped mirror), and how many vdevs exist in the overall zfs storage pool.

Storage Migration

I see an alarming number of guides on the internet for migration from vmdk > qcow2 in a zfs context that doesn’t talk about sector misalignment. What’s sector misalignment in this context? Generally, with ZFS, it works in 128k chunks by default (and some use cases make sense for 1mb and other use cases make sense for ~16kb). A given ZFS dataset can be created inside of a ZFS volume with a different record size.

Generally when mapping virtual machine storage there are alignment considerations for both the partition and filesystem clusters in the VM.

Operational consequences VMware admins need to hear:

  • Scrubs and SMART matter now (schedule them, alert on them)
  • Snapshots/replication are cheap—until they aren’t (snapshot sprawl and retention policy become your job)
  • “Just add disks” is not always a free win (recordsize/volblocksize, special vdevs, slog decisions, and ashift permanence)

If VMFS was “a datastore,” ZFS is “a storage system.” Proxmox makes you own that—but then rewards you for it.

Example Migration & Check

# Install tools (Debian/Proxmox host)
apt-get update
apt-get install -y libguestfs-tools virt-v2v

# Convert an OVA to local KVM output (creates disks + XML metadata)
virt-v2v -i ova /path/to/vm.ova -o local -os /var/lib/v2v-out

Example: 200G disk, 16K volblocksize (common starting point; tune per workload).

zfs create -o mountpoint=none tank/vmdata

zfs create -V 200G -b 16K -o volmode=dev tank/vmdata/vm-100-disk-0

# then to convert vmware vmdk

# Source can be VMDK, qcow2, etc.
qemu-img convert -p -f vmdk -O raw /path/to/source.vmdk /dev/zvol/tank/vmdata/vm-100-disk-0

# run sync twice to flush buffers out to disk 
sync
sync

Then in proxmox you can do something like:
qm set 100 -scsi0 /dev/zvol/tank/vmdata/vm-100-disk-0,discard=on,ssd=1

assuming you’ve already created a VM id # 100 in proxmox

Note: Ideally the storage is also mapped in your Proxmox host configuration because if you’re using ZFS and a Proxmox cluster ZFS makes it easy to manage snapshots and migrations from one cluster host to another. This example short-circuits Proxmox management a bit and TODO comeback and fix this if there are questions.

The Proxmox Wiki is also useful. Read it.

If you want to learn more about ZFS

Not ready to talk about migrations, tools, and it’s a strange new world for you talking about file system cluster sizes? Not wo worry – we have other threads and external resources:

Understanding ZFS from Klara

ZFS Performance Tuning from Klara

ZFS Lessons Learned with Allan Jude and Tom Lawrence

What about Not ZFS for storage on Proxmox?

First, LVM.

On VMWare, it gives you nothing for raid. This sucks because basically every fast nvme raid system out there goes directly through the CPU, not an add-in pcie card, because add-in pcie cards are the bottleneck when your SSDs can manage 15 gigabytes per second transfer.

And ZFS is a lot of overhead. LVM is a volume manager. It can sit on top of almost anything.

mdraid (Linux MD) is the Linux software raid implementation. That can present a raid device, typically /dev/md0, from a number of underlying physical disks.

LVM Raid are built on top of the Linux MD code via device mapper, but it is not really the same as you’d get running the md raid admin tools directly (mdadm). In other words the Raid1/5/6/10 that LVM supports is essentially the same code path as linux MD but.. one gets there a different way. Tried and true code, though).

Intel VROC on Linux is extremely close to mdraid behavior, but one onotable difference is support for plugging the raid5 write-hole via a journal/mailbox style mechanism. GRAID Technology has taken over VROC and this software raid on VMWare natively is the only fast software nvme raid that I know of.

Guest Drivers Storage Performance

Don’t overlook guest driver performance. There is a new storage driver coming for windows that’s better optimized for more threads. I can clear 100k i/ops in a VM easily now.

9gb/sec from NFS over RDMA on Proxmox 9 is entirely possible:

Hyperconvergence?

If you came from vSAN, you’ll be tempted to map 1:1 with CEPH. Don’t. At least until you’re fully read in.

Proxmox supports modern Ceph releases and documents upgrade paths heavily. Understand that, like ZFS, Ceph predates Proxmox.

Success comes from CRUSH/failure domain design, network design, and hardware homogeneity—not from assuming “HCI magic.” Depending on your administrative experience, this could be a blessing or a curse. Or both.

Ceph is quite good in my experience, especially with many (7+) cluster hosts, and can deliver a similar HCI experience

2 - You’re not “in vCenter”; you’re in Linux (and that’s the point)

This storage talk may make it seem like a lot, but this is the power of Linux. It’s still Debian 13 under the hood. I like this because it is less opaque. When vSAN first launched everyone assumed that it would be efficiently locating blocks in the storage pool. Narrator: It was not doing that.

If storage is the first mental model shift for VMware admins, Linux is the second—and bigger—one.

Proxmox is not a sealed appliance like ESXi.
It’s a Debian-based Linux system, and that freedom cuts both ways.

  • ESXi: locked-down hypervisor, tightly controlled surface area
  • Proxmox: full Linux OS with a web UI on top
  • You are an OS administrator again, not just a hypervisor operator

Proxmox deliberately does not re-implement or “own” core technologies:

  • ZFS, LVM, mdraid, Ceph, corosync, Open vSwitch

These are upstream Linux projects with long histories

Proxmox integrates them, but does not abstract away their behavior

This is more flexible than VMware—but it assumes you know what you’re doing.

Customizing the host: sometimes good, sometimes a mistake

Because it’s Linux, you can:

  • install packages
  • tune kernel behavior
  • add monitoring agents
  • script automation

But you should not treat the Proxmox host like a general-purpose server.

Common pitfall: Docker on the Proxmox host

Docker itself works fine… the problem is networking generally. Docker rewrites iptables / nftables

Proxmox assumes it controls bridges, firewalling, and routing

Result: subtle breakage, hard-to-debug networking issues

Best practice:
Run Docker inside a VM (or container), not on the Proxmox host.

This is not a performance issue—it’s an operational sanity issue.

Containers: Proxmox has something VMware never really did

Proxmox supports:

  • Full VMs (KVM/QEMU)
    *LXC system containers (not Docker)

LXC:

  • very low overhead
  • excellent for services, appliances, and hardware access
  • often a better fit for things like:
  • media servers
  • lightweight services
  • direct access to host devices (GPUs, encoders, etc.)

VMware never had a clean equivalent here.
Its Kubernetes integrations were .. uh… lacking. And license-driven.

Filesystems and sharing: improving, but still evolving

Running Docker in a VM often means you want filesystem passthrough

  • Historically awkward
    virtio-fs now makes this practical!
    relatively recent, but works well

This is another example of Proxmox benefiting directly from Linux ecosystem progress

#3 – Backup & Disaster Recovery: Proxmox Gets This Right

This is where Proxmox quietly embarrasses both VMware and Microsoft.

VMware’s historical failure mode

  • Early free ESXi had usable snapshot and backup hooks
  • VMware deliberately removed or crippled them to force licensing
    • This happened well before Broadcom
  • VMware’s philosophy prioritized:
    • HA
    • replication
    • complex clustering features
      over:
    • bare-metal restore
    • simple, reliable backups

Result:

  • Backup became a third-party problem
  • This vacuum is exactly why Veeam exploded in popularity

This wasn’t an accident—it was a product decision.

Windows followed the same pattern:

  • “We ship the OS, backup is your problem”
  • Windows Backup became unreliable sometime around ~2007
  • Enterprises stopped trusting it long ago

Proxmox takes the opposite approach

Proxmox Backup Server (PBS) is not an afterthought.

It is:

  • a first-class product
  • designed alongside Proxmox VE
  • focused on backup and restore as core workflows

PBS provides:

  • incremental-forever backups
  • block-level deduplication
  • verified restores
  • fast file-level recovery
  • native encryption
  • hardened authentication
  • simple, transparent retention policies

No licensing games. No feature sabotage.


Architecture matters

  • PBS works best on bare metal, but can run as a VM
  • Designed to scale independently from compute
  • Supports self-replication
    • makes true 3-2-1 backups trivial
  • Explicitly designed to reduce ransomware blast radius
    • backup server auth model is isolated from VE hosts

This is the kind of design you get when backup is treated as a primary concern.


“But we already use Veeam”

Good news: Veeam now officially supports Proxmox VE.

That means:

  • You don’t lose your existing backup workflows
  • You can mix:
    • PBS for fast local + replication backups
    • Veeam for enterprise workflows, compliance, and cross-platform consistency

This removes one of the last serious blockers for VMware refugees.

#4 - #4 – Proxmox Datacenter Manager: Not vCenter (Yet), But You Should Be Using It

Let’s get this out of the way first:

Proxmox Datacenter Manager is not a vCenter replacement.
And that’s okay.

What it actually is

  • A central management plane for:
    • multiple Proxmox VE instances
    • multiple clusters
    • Proxmox Backup Server
  • A way to get fleet-level visibility without logging into every cluster

Most real management still happens at the cluster or node level.


Why it exists

Datacenter Manager exists because:

  • people run more than one Proxmox cluster
  • VMware refugees expect some central visibility
  • Proxmox needed a way to scale operational awareness without rebuilding vCenter

And development velocity here is high—largely driven by the influx of VMware users.


What it can do today

  • Single UI for multiple Proxmox VE instances
  • View:
    • cluster health
    • node status
    • VM and container inventories
  • Integrates Proxmox Backup Server
  • Supports cross-cluster VM migration
  • Provides delegated access / role-based views

Think of it as:

“Unified visibility and light orchestration,” not “full policy-driven automation.”


What it does not do (yet)

  • No deep lifecycle automation
  • No DRS equivalent
  • No advanced policy engines
  • Limited built-in health analytics

If you’re expecting:

  • automatic load balancing
  • aggressive hands-off optimization
    you won’t find it here.

That’s still on you.


Monitoring: the missing piece

Proxmox has:

  • solid built-in graphs
  • good logs
  • per-node visibility

What it lacks is rich, opinionated health monitoring.

This is where Linux shines:

  • tools like Netdata work extremely well on Proxmox hosts
  • real-time visibility into:
    • CPU contention
    • memory pressure
    • I/O stalls
    • network saturation

Best practice:

  • install monitoring agents on hosts
  • bind dashboards to internal networks only
  • treat monitoring as infrastructure, not a Proxmox feature request

VMware contrast

  • vCenter bundles:
    • orchestration
    • automation
    • policy engines
    • telemetry
      into one heavy control plane
  • Proxmox intentionally avoids this monolith
  • Instead, it exposes clean integration points

This is a design choice, not a missing feature. Datacenter manager might someday bridge the gap more, especially for cluster-wide aspects, but not yet.

#5 - Networking Best Practices, SDN and Beyond

#5 – Networking & Proxmox SDN: Powerful, Linux-Native, and Not NSX

Networking is one of the areas where VMware admins bring the most assumptions with them.

Some of those assumptions will hurt you.


Start with the basics: Proxmox cluster networking

Proxmox relies heavily on corosync for:

  • cluster membership
  • quorum
  • fencing coordination

Important clarifications:

  • Corosync does not need high bandwidth
  • It does need reliability and low jitter
  • 1 GbE is fine; bonded 1 GbE is better
  • Do not share corosync with:
    • live migration
    • VM traffic
    • storage replication

This should feel familiar if you’ve built VMware clusters—but Proxmox does not enforce this separation for you.


Separate your traffic classes

A sane Proxmox design usually includes:

  • Corosync network (reliable, boring)
  • Migration network (high bandwidth)
  • VM data networks (often VLAN-based)

This can be:

  • multiple physical NICs
  • bonded NICs
  • VLANs over high-speed links

Proxmox gives you the primitives; it does not dictate topology.


Linux networking under the hood

Proxmox networking is built from:

  • Linux bridges
  • Linux bonding
  • Open vSwitch (optional)
  • iptables / nftables
  • standard Linux routing

Nothing proprietary. Nothing hidden.

This means:

  • excellent performance
  • predictable behavior
  • deep observability
  • and zero hand-holding

Proxmox SDN: what it actually is

Proxmox SDN is:

  • a management layer over Linux networking
  • based on Open vSwitch
  • supports:
    • VLAN
    • Q-in-Q
    • routed networks
    • NATed private networks

It shines at:

  • creating private VM networks
  • building lab or tenant isolation
  • standardizing network definitions across nodes

What Proxmox SDN is not

This is where VMware admins need to recalibrate expectations.

Proxmox SDN:

  • does not configure top-of-rack switches
  • does not offload policy to hardware fabrics
  • does not have an NSX equivalent
  • does not integrate with SmartNICs for policy enforcement

VMware could do some of this because:

  • Broadcom owns the switching silicon
  • VMware built tight hardware partnerships

Proxmox intentionally stays hardware-agnostic.


Why that’s not a dealbreaker

For most environments:

  • VM-level networking
  • VLAN segmentation
  • NAT isolation
  • software routing

are more than sufficient.

And because it’s Linux:

  • you can inspect every packet path
  • you can debug with tcpdump, ip, ethtool, ovs-vsctl
  • you can integrate external networking tools if needed

Proxmox HA is cluster-manager-driven orchestration. It’s powerful, but:

  • fencing/watchdog strategy is your responsibility
  • shared storage expectations differ depending on ZFS replication vs Ceph vs shared SAN
  • you’ll want to design failure domains intentionally rather than expecting “it’ll just do what vSphere does”

#6 – Subtle Knobs & Tunables: Power Tools, Not First Tools

Once storage and networking are sane, Proxmox exposes a class of second-order tunables that VMware mostly hid—or slowly removed.

This is just something I want you to be aware of that can impact performance, but not always.

These knobs matter, but they are not where your first performance problems will be.


Memory deduplication (KSM)

Proxmox uses KSM (Kernel Samepage Merging):

  • merges identical memory pages under pressure
  • enabled by default
  • tunable and observable

It helps when:

  • you run many similar VMs
  • memory is actually contested

It can hurt:

  • CPU-heavy or latency-sensitive workloads

This is the Linux equivalent of VMware’s old TPS—but honest, explicit, and under your control.


Huge pages

Proxmox lets you use:

  • 2 MB huge pages
  • 1 GB huge pages (hardware permitting)

Benefits:

  • fewer TLB misses
  • lower page table overhead
  • modest but real gains for large, memory-heavy VMs

Costs:

  • reduced flexibility
  • harder memory overcommit
  • planning required

Huge pages are not magic—they’re workload-specific.


CPU topology & scheduling

Unlike VMware, Proxmox makes CPU layout explicit:

  • NUMA awareness
  • CPU pinning
  • host vs passthrough CPU models
  • scheduler behavior is Linux, not a black box

This matters for:

  • multi-socket systems
  • memory-locality-sensitive workloads
  • performance consistency

Misuse can make things worse, not better.


I/O tuning outside storage

Beyond ZFS/LVM geometry, you still have:

  • virtio queue depth
  • multi-queue networking
  • interrupt distribution
  • I/O thread placement

These are real knobs, but again:

If you’re touching these before fixing storage alignment, you’re optimizing the wrong layer.

7 Likes

How are YOU using storage options in Proxmox?

There are some other good proxmox threads here I should link

Also, this:
Proxmox VE Administration Guide

CPU security mitigations (i.e. spectre/meltdown should NOT run in your virtual machine!)

2 Likes

Thanks for the video. I sub’d to your youtube channel 6 or so months ago and only made a forum account today.

This topic is aligns perfectly with my project to find a VMware replacement prior to the middle of next year for our two cluster 12 host environment. Not sure Proxmox/ZFS is the right choice for us but more information is always beneficial.

Thanks again for all the great videos and knowledge.

I work in a Datacenter / ISP and we have multiple proxmox clusters, some internal, some customer facing, some “private clouds” we resell.
Were currently midway through a migration from a legacy hyperv cluster of around 80 servers which ran around 300 webhosting vm’s.

Weve skipped zfs almost entirely and ether ran qcow2 on the legacy dell raid controllers or straight to ceph with ssd’s and nvmes.

we would never go back.

We have a mix of clusers which are high performance (epycs with nvmes) and older xeons with enterprise ssd’s.

The biggest learnings we had were with corosync issues and making sure to have fully seperated network infra from fronend and storage with corosync running on both for safety.

Honnestly if you think your going to grow your cluster more than 5 nodes. Just go with ceph and enjoy it. just dont do less than 25gb storage switches IMHO, 10gb is the absolute minimum and honnestly we don’t even consider that for new clusters.

What has been really interesting recently is the interest in “Private Cloud’s”. We have had some large companys coming to us to get large parts of their workload out of AWS or GCP. The massive discounts customers seemed to be getting for 10+ years are going away and the financials make sense for many of them to buy gear and go colo or managed private clouds

3 Likes

Great overview of the current state of Proxmox.

One thing I didn’t realise was Proxmox SDN being OpenVSwitch-based. I always understood OpenVSwtich as a node-level alternative to Linux-native bonds/bridges/VLANs with SDN capabilities. I thought ‘Proxmox SDN’ at the Datacenter/cluster-level was based on OpenFabric / FRR.

BTW I’ve linked a useful diagram of the Linux Network Stack (with Proxmox flavouring): Entire Linux Network stack diagram

2 Likes

Log time watcher, first time poster. At work we have been moving away from VMware for over a year, and between new products and acquisition of another company, we’ve picked up two other virtualization platforms to offer our customers. We are looking at using Proxmox to host all of our management, apparently back ended by object storage (even though we don’t have a decent object storage solution yet). I have a single spare PC at home that I would not mind converting to run proxmox - I don’t have the resources to build a cluster at home. I have an old QNAP that has about 13 tb total usable capacity, and a much newer NvME solution running Open Media Vault that has about he same across four 4tb NvME drives. I have options for storage, at least, and if I have to blow one array or another away and do it differently to support proxmox, I can.

I need to spend some time with this post, really dig into it to try to re-wrap my brain around the concepts.

“If storage is the first mental model shift for VMware admins, Linux is the second—and bigger—one.”

Maybe - maybe not. At work, we’ve been bitching for YEARS about the lack of ability to basically drop into Linux on vmware appliances and “just do.” I’m old, I’ve been doing this a long time, and honestly, having Linux underpin virtualization directly is what most of my co-workers have wanted - whether they know it or not - for years. “appliances,” whether physical or virtual, can be EXTREMELY limiting. I look forward to digging in on this.

I have lots of questions, but I need to do the reading to see how many questions are already answered above. Thanks for putting this together. Very timely for me at least :slight_smile:

Luckily for me we’re doing a hardware refresh and the EOL hardware will be my test bed for options to migrate off VMware.

As Wendell said the options seem to pretty much be Nutanix, HyperV (Azure on Prem), Proxmox, Xen-ng. The last two have the issue of being European companies and those of us in the states particularly in my line of work have to factor that into the equation (time difference for non-priorty 1 issues, geopolitics, etc). Enterprise 24x7x365 4 hr response is a must for us so support is a big deal. VMware’s support is basically worthless these days so even if we had time zone issues I’m not sure we’d be loosing much.

HyperV feels to me like it stalled and is on the edge of being abandoned by MS for years when it makes sense for them to push those customers to the cloud.

Then there is Nutanix which really isn’t that much cheaper than VMware and just feels like they are just as likely to get bought out by private equity and have the screws turned on us again.

Bot the bottom line is we can’t justify continuing using VMware at their current pricing/offerings so we have to move to something else. So I’ll be taking a hybrid approach by testing solutions over the next year while while simultaneously building out our Azure presence.

1 Like

I’m writing because Wendell saw my post on the corresponding YouTube-Video and seems curious. :slight_smile:

We are running a three node Cluster with Ceph and all the HA stuff.
Storage is all NVMe.
Each node has 1TB of 2933-DDR4 RAM and is powered by a Epyc 7443P.
Everything is overpowered because we wanted a system that we can make use of the whole seven years Dell provides services.
Ceph is running over fully meshed (ovs) 100Gbit/s Mellanox Cards - unfortunately software only. Proxmox doesn’t support Hardware-Offloading, as far as I know.
Corosync is running over dedicated also ovs fully meshed 1Gbit/ Intel NICs.
The public network is a 25Gbit/s broadcom, curently running at 10Gbit/s - as soon as the network is updated they will run at 25.

And we are running a pbs in another room at the building.

We haven’t had a unplanned shutdown of a VM over the last almost five years. That is because all VMs have rules within the HA manager. So I can just reboot or power off one node and everything else is handled by the HA manager. Once you have configured everything it is that simple.
We are using the 100Gbit/s network for live migration.
On a maintenance day, when I have to update firmware on the hardware; the only thing I’m watching out for is between reboots, that ceph is in good health. But “rebuild” is really fast.

During a BIOS-Update the “Predictable Network Interface Names” changed! Predictable my A**. But it wasn’t that bad. Just change the names in the networking interfaces.

Updating ceph and proxmox - which has happened twice since the initial setup - was without any hiccup. Just follow the documentation.

Our professors have access to their machines; access is granted over LDAP.

And we are currently working on a script to securely shutdown everything if the UPS is low on battery. Proxmox GmbH knows that I want that feature build into Proxmox, but it seems it is no priority to them.

Any questions?

3 Likes

I use NUT to handle graceful shutdowns with lots of UPSs and even more nodes. The perks of Proxmox living on top of Debian Linux. You might not require anything custom if your UPS(s?) supports it. Not that I would try to talk you out of a fun project.

2 Likes

Years ago I used to copy vmdk files over to my proxmox server and rename the files to .raw and edit the vm config file to change it to raw device and the vm would just start up and work.

1 Like

The moment you are using ceph on your cluster, there is no graceful shutdown possibility with Proxmox. Due to the fact that you cannot make sure that all nodes take the same amount of time to shut down, there will be a moment in time where only one node is still alive and possibly still running VMs or are shutting down. At that point in time, ceph will go into read-only mode, which will prevent to shut down all the VMs gracefully. I already talked to someone over at Proxmox about this issue and he got me the right insights at which point we should access the API to shut down the VMs. Also, we have to make sure that the VM running NUT is the last one to shut down and of course I want to run NUT on my cluster so that I can be sure that it is always running.
If there was an easier or more straightforward way to accomplish that, I would argue that the support at Proxmox would be aware of that.

2 Likes

One thing I forgot was we actually migrated about 30 VMs from a VMWare machine to Proxmox using an export function of the ESXi, which was terribly slow because it only used one CPU and you could not deactivate compression. But import on Proxmox worked like 10 times faster and all the Linux VMs worked fine. The only issue we had was with one Windows Server VM, which we fortunately could decommission since then.

1 Like

Two node cluster with Dell Powerstore and 32Gb FC network. All storage related needs to be configured in cli. Shared storage works with thick LVM. Only thing I’m unsure is with Powerstore immutable snapshots, in Vmware snapshot can be mount as new volume and import VM:s but I’m not sure about Proxmox if it can import VM:s from thick LVM volume and specially if I already duplicate versions of those VM’s running.

Proxmox has ESXi import tool which is bit clunky but I’ve used it succesfully with production VMs. Biggest problems come with Windows iscsi drivers, I think pvscsi works fine in Proxmox but with other there is mess of injecting VirtIO drivers to VM. Also Vmware tools refure to uninstall if VM is not on Vmware and needs to be uninstalled before migration or with script.

notable*

Back when I managed proxmox (v5, 6 and 7) we used a couple of centos NAS’es with md raid10 and shared them via NFS, then on all the proxmox hosts we had the NFS paths added to them and used qcow2 VM disk format.

When we collocated our servers, the guys at the datacenter told us that they have never seen so many ethernet cables coming from a single server, lmao! The minimum was 5, with a max of 7 on some servers (2x NICs bond pairs, 1 for storage LAN, 1 for VM traffic and corosync and a single port for IPMI, with some servers having a separate corosync bond - all were LACP gigabit, except for the storage LAN, which 10G bond).

I would’ve done thing differently (if the budget was approved - I especially wouldn’t have combined VM and corosync traffic, but it worked alright).

I never understood, even with the previous prices, why everything moved to cloud. The only somewhat “reasonable” thing was lowering the IT department (no need for the hardware janitors to go to the baremetal servers). But hosting your own infrastructure was way more economical (and still is, even with RAM shortages).

You don’t have to build a cluster at home. If it’s a lab, just virtualize everything (run proxmox or virt-manager on a single server and put all the proxmox hosts in VMs, if you want to play with it).

That’s exactly the reason why I personally don’t use stuff like pfsense / opnsense, truenas (core / scale) and others and prefer to stick to general purpose OS. Even proxmox annoys me sometimes, but I still recommend it to people - at least it’s more open and it’s open source as well. Not the best thing out there, but it’s the best we have and TBH the easiest to get into (when someone doesn’t like the software, but still recommends it, you know either something is wrong with the industry or something’s really good about the software).

I’d say that if you go FOSS in your environment, your admins need to know very well what they’re doing. You’re paying for more qualified computer janitors, as opposed to paying prohibitively expensive support contracts (and the fault falls within the company as opposed to outsourcing fault - IMHO this makes companies be more honest instead of playing the blame game).

?!?

Literally free to use, if you don’t use the enterprise repo and access to the enterprise repo shouldn’t be as expensive as VMWare. Support contracts, probably, since they do the same dumb per socket / per CPU cores licensing for support (I still don’t think it’s as expensive as vmware).

I have to agree to this. I feel like a lot of bigger fish are seeing the growth of proxmox and may be eyeing to buy it. But given all the current track-record, when a company gets bought, their products get ensh*ttified.

I actually had that happen to me on Proxmox version upgrade (v5 to 6), when they broke some udev rules (didn’t get fixed until v7, so I had to work with NICs with names like “rename6” and “rename7”).

This kinda leads in to what Wendell was saying.

I was always a fan of keeping the hypervisor do as little as possible and I’m not a fan of the CPU + RAM usage charts in Proxmox, VMWare and other platforms. They’re a useful indicator in rare scenarios, but it’s better that you just install prometheus / zabbix agents on your hypervisors and inside your VMs and you monitor them on a dedicated NMS. That’s where you should keep an eye on SMART alerts, ZFS degradation, CPU spikes, RAM usage, space utilization etc.

In the same vein, NUT server is great for graceful shutdowns. Similarly, I prefer backups to happen at OS level, not VM / hypervisor level. You could do snapshots and zfs-send or-what-have-you, but if you’re running programs inside your VMs that require some data consistency, don’t put your money on snapshots. And that way, you save on storage consumption as well, since you can avoid backing up useless cruft.

Ugh… I never thought about this, but I always found the hyperconverged infrastructure to be undesirable (better to have a storage LAN, dedicated storage servers and have the compute connect to storage through TCP/IP or fibrechannel).

I don’t think you can import / activate 2 LVM volume groups with the same name on the system. You’d maybe need some kind of workflow on a second server to activate the VG and clone / export the LV to another VG, on another LUN, then use that on Proxmox (or you might be able to hack your way around with SSH to send the LV data straight to the VG that’s already on the other server).

Reading a bit deeper into this, with the snapshot of the LUN, you could make a cloned volume and you might be able to pvchange the UUID and vgrename the VG name and then import it on your proxmox server.

But you’ll never want to have the same dev PV UUID and VG present at the same time. You’ll trash both file systems (you may be able to snap revert on powerstore, but… you know… massive downtime for a bunch of VMs for no good reason).

1 Like

Proxmox is awesome.

It won out of a three-contender race between Proxmox, xcp-ng, and TrueNAS when I had my mass consolidation project of January 2023.

virtiofs allows Host ↔ VM communication if you want/need it, and as far as I can tell, bypasses the network stack entirely (but has some RAM/memory footprint ramifications).

Native support for LXCs also means that I can share GPU resources between LXCs, neither of which I can do with xcp-ng nor TrueNAS (as a virtualisation platform).

The one downside, that I have found, so far, with Proxmox is that being that it’s Debian-based, the opensm doesn’t support virtualisation for IB networks. For that, you’ll have to use the RHEL (or RHEL-derivative) opensm, which is a minor PITA. Also unfortunately, that’s neither a Debian nor a Proxmox problem as I reached out to the linux-rdma core group and they have no intention of enabling virtualisation in the Debian versions of opensm.

There are a lot of things that I’ve been able to do with Proxmox as it checked pretty much all of the boxes in terms of what I wanted the system to do.

As a sidebar, the virtio-nic shows up in Windows 7 as a 100 Gbps NIC. In Windows 10+, Linux, and MacOS, it shows up as a 10 Gbps NIC.

Proxmox is and can be insanely powerful.

Proxmox Data Center Manager is new, but I am positive that they will keep improving that.

It is a very competent platform and the people over at their forums are also more helpful than the people in either the TrueNAS or the xcp-ng forums.

2 Likes

Realizing I made a mistake using cqow2 I converted an image to .raw. However after booting it up and benchmarking the iops were halved?

The storage is on a truenas instance with SSD’s over CIFS/SMB. Any ideas why this slowdown might occur?

did you run some successive benchmarks? the first time on qcow2 it’s likely a lot was cached. to really compare you’d have to reboot the host as the cache can persist there outside the vm

the other advantage of raw is creating a dataset that aligns with the filesystem cluster size

1 Like

I’m using Zabbix to monitor Proxmox via API with Zabbix offical template but I noticed with my home Proxmox server that it doesn’t detect disk failure in ZFS array in any way. So it would be better to install linux agent on Proxmox host and use generic ZFS template to monitor ZFS and then monitor VM aspect via API?

Lets say I get separate DR server which doesn’t have any datastores mounted and then I mount snapshot of Thick LVM volume on it. What should I see? Bunch of unmapped VM disks or would they need to be imported after that? Current hardware was bought only to be used with Vmware beforce Broadcom deal and after part of the cluster was split into Proxmox. After current hardware expires need to think what would be most robust storage option for Proxmox.

Ceph is about high availability so shutting it all down isn’t really part of the main design specs, because that’s no availability, neither normal nor high; least of all under load.

Now, of course, in my home lab, e.g. with three small Atoms running what used to be an oVirt cluster now on Proxmox (another uses faster NUCs), even I do shut them down when I go on a trip or re-organize/rebuild my “data center” (which is lots of boxes under my desk).

For truly critcal services like (Univention) IDM, I run an extra instance on a stand-alone system while I do infrastructure rebuilds of the cluster, since that might take some time.

But there it’s just all about shutting down all guests first, so there is zero updates on the Ceph cluster and then you can just shut down those nodes one-by-one.

Likewise, I just turn the cluster members on after, letting them reach synced state (Ceph does a quick scrub) and only then bring up the VMs.

For all maintenance I just try to make sure I never down more than one node and ensure there isn’t much activity going on in the mean-time so resynch won’t take long

Worked just fine every time so far, was quite a bit more nail biting with GlusterFS and oVirt, even if it’s the same principle. And there I sometimes had to wait for days e.g. until the data center team in the country next door where some of my oVirt clusters ran, were able to swap out the defective hardware.

1 Like