Nicholas TongProxmox VE

47 / 56

the hypervisor that wouldn't start without a third vote

6 min read 1,308 words

Screenshot of the Proxmox VE web interface on the datacenter summary page, showing a two-node cluster like the one in this post

The power blipped at 2:40am on a Tuesday in June. Two seconds, maybe three, the kind of flicker that doesn't reset a microwave clock. I know because I came downstairs after and the microwave still had the right time, which felt like an insult given what the shelf two feet away was doing.

Both Proxmox nodes rebooted. One came back. The other didn't, because two years earlier I had set "Restore on AC Power Loss" to "Power Off" in its BIOS, meaning when mains returns, stay dark and wait for a human. I set that deliberately, for reasons I can no longer reconstruct, and then forgot the setting existed. So at 2:41am my "redundant" cluster was one used office PC holding one corosync vote out of two.

The survivor would run. It just wouldn't do anything. The cluster filesystem at /etc/pve had flipped read-only because the cluster lost quorum, and the VM that runs my SIEM, the one box in the house whose entire job is watching, refused to start. Proxmox's docs say it flatly: the cluster "switches to read-only mode if it loses quorum". The docs do not describe how that sentence reads at 2:41am.

I fixed it from my phone, over SSH, thumbs, by the light of the flashlight on my keyring: pvecm expected 1. That command tells the cluster to act as if it had the votes it needed. The SIEM started. The second node needed a physical finger on its power button, and that was a morning problem.

The gap in my SIEM's logs that night is twenty-six minutes long. It is the only weather-shaped hole in a year of data. I kept it. Scars are documentation.

the part everyone gets wrong

Everyone hears that clustering is free and turns it on. Two nodes, corosync, done, and now it feels like redundancy. A two-node cluster with no third vote is less available than two machines that ignore each other. A standalone host doesn't care that its sibling lost power. A clustered host refuses to start your VMs because it cannot prove its sibling is dead. Same hardware, opposite outcomes, and the whole difference is one quorum rule.

I got this wrong for about two years, and the reason is that the button is free. Proxmox doesn't paywall clustering. There's no license dialog warning you off. The wizard accepts two node names as happily as five, and nothing in the interface says "you have just built a system that one power blip can stun". The failure isn't in the software. It's in the assumption people carry to it, which is that free redundancy is redundancy.

To be clear about the product, because the product deserves it: Proxmox VE is good. Version 9.2 shipped in May with Debian 13.5 underneath and a new dynamic load balancer, and in August the project ported it to arm64 with help from Nvidia and Supermicro, which The Register covered the next day. It's free under AGPLv3. The optional subscriptions, which mainly buy the enterprise package repository and support, are priced per CPU socket per year: EUR 120 for community, EUR 370 basic, EUR 550 standard, EUR 1,100 premium. I could be wrong about this, and if I am, someone will tell me, but those numbers crept up around four percent this year without an announcement; I found it by comparing Wayback snapshots. Not a scandal. Just what "free" looks like when you squint.

how the votes actually work

Corosync runs on every node and keeps a membership list. The votequorum layer inside it hands each node one vote by default, then requires the cluster to hold more than half of all votes to be "quorate". Two nodes means two votes, so quorum is two. Lose a node and the survivor holds one of two. Fifty percent. Fifty is not more than half, so the config filesystem goes read-only and the scheduler stops starting guests.

pvecm expected 1 lowers the expected vote count. It exists for exactly my night, a node that will be back and a survivor that should keep working meanwhile. It is also how you build a split brain: run it on both halves of a partitioned network and you have two clusters, each certain it's the cluster. The command is a scalpel. I had been treating it as a painkiller.

The sanctioned fix is a QDevice, a third vote that lives somewhere else. A Raspberry Pi runs corosync's qnetd daemon and casts the deciding ballot whenever the nodes disagree, which in practice means whenever one of them goes quiet. The docs treat this as a first-class option for two-node clusters, and it is. Setup took me twenty minutes.

I didn't do it for three months.

That's the admission I'd rather not type. After the June night I knew the exact fix, knew it down to the command name, and kept reaching for pvecm expected 1 instead, because it was faster and because each individual blip felt handled. Three months of a known twenty-minute fix, skipped, on a problem that had already woken me once. Detection engineering is my day job. The gap between knowing and doing is not a homelab quirk. It's the whole discipline, and I failed it on hardware I own.

where it breaks

The QDevice moves the failure instead of deleting it. Read the cluster docs closely and the arithmetic gets worse, not better: if the QDevice's daemon dies, no node may fail either, or the cluster loses quorum again. On fifteen nodes that's an edge case. On my two, it's the entire design. There's also a documented landmine in votequorum itself: the auto_tie_breaker flag can't be combined with a QDevice, so the shortcut you were about to try at midnight is already tried and ruled out.

Patch discipline matters more here than on a laptop, because a hypervisor console is a view of everything. In April the project fixed CVE-2026-51082, a race condition that could let one VNC session hijack another, CVSS 7.2, patched in pve-manager 9.1.9 and 8.4.19. A shared console between tenants is a boundary, and this was a hole in the boundary. It got a patch within the month, which is the part that earns trust.

And the reason I'm here at all: I migrated off ESXi in a weekend last year, after Broadcom ended the free version in February 2024. They later restored a limited free build, 8.0U3e, in April 2025. I don't need that to be a villain story. Companies reprice products; you re-plan. What I'll say is what it cost: two used 1-liter office PCs, one weekend, about fourteen hours counting the storage re-learning, and Proxmox Datacenter Manager 1.0 shipped that December and would have shaved a couple of those hours off. The migration was fine. The part nobody budgets for is quorum, and quorum is what billed me the one night that mattered.

what to do before the next blip

  1. Set "Restore on AC Power Loss" to restore power on every machine you can't physically walk to. Costs nothing, takes one reboot.
  2. Running two nodes? Add the QDevice tonight. Twenty minutes, a Pi you already own, pvecm qdevice setup and done.
  3. Break it on purpose. Pull power on node two on a Saturday afternoon and watch which VMs refuse to start. A failure you haven't rehearsed is just a future incident with better lighting.
  4. Check pve-manager is past 9.1.9. Then check again next month, because it will not stay put.

Total bill for this post: one weekend to migrate, one night to learn what quorum means, twenty-six minutes of missing logs, and twenty minutes of fix I deferred by a quarter. Hurricane season runs through November, and the Pi now sits on its own little UPS. The BIOS settings on both machines read "Power On". I checked twice.

The microwave clock was fine the whole time.