Engineering note
My Three-Node Proxmox Cluster
I built this cluster so I could patch or lose one host without taking down the smart-home services everyone in the house notices immediately.
It runs on three small Proxmox nodes. Home automation lives on Node 1, DNS on Node 2, and virtual machines on Node 3. A separate Ubuntu machine handles GPU work.
Layout
flowchart TB
subgraph PROXMOX["PROXMOX CLUSTER"]
N1["Node 1<br/>LXC Containers<br/>━━━━━━━━<br/>Homebridge<br/>Scrypted (Ring)<br/>so-co (Sonos)"]
N2["Node 2<br/>DNS Services"]
N3["Node 3<br/>Virtual Machines"]
end
GPU["GPU Server<br/>Ollama API<br/>Local Models"]
N1 -->|API Calls| GPU
N1 -->|Passthrough| AppleHome[Apple Home]
style GPU fill:#1b1c19,stroke:#c6a663
style PROXMOX fill:#1b1c19,stroke:#3e3e37
The cluster nodes are modified Dell OptiPlex Micro machines with i5-10500T processors and 16 GB of RAM. They are small, quiet, and easy to replace. That matters more here than peak CPU performance.
I measured Node 1 at about 4 W idle. Node 3, which carries the VM load, peaked at 24 W during my testing. The separate GPU host has an RTX 3070 and runs Ubuntu with Ollama. Keeping it outside Proxmox avoids PCIe passthrough and lets the three cluster nodes stay low-power.
Why three nodes
I installed Proxmox VE on each machine, assigned static IPs, created the cluster on Node 1, and joined the other two. Corosync handles membership and quorum.
With three votes, the cluster can keep quorum with any two nodes online. An even two-node setup would need another vote or special handling when the nodes disagree. Three physical nodes made the failure behaviour easier to understand and test.
Corosync also made the network part of the cluster. During early testing, a switch briefly dropped packets, quorum disappeared, and the nodes fenced themselves. Since then I have kept cluster traffic wired and the network path deliberately boring.
Network and services
All three nodes connect to a UniFi US-8-150W switch. The management network is still a flat /24; I have not moved it to a separate VLAN yet. Corosync and normal traffic share the primary Ethernet interface, and I disabled the unused secondary interfaces after they caused routing confusion during setup.
Node 1 runs Homebridge, Scrypted, and Sonos services in LXC containers. These services need mDNS and HomeKit broadcasts, so their virtual interfaces bridge directly onto the physical network instead of sitting behind NAT.
Node 2 runs Pi-hole. A second DNS instance carries the same configuration, so DNS can fail over separately from Proxmox HA. Node 3 holds the virtual machines, including a Windows VM that stays tied to that host.
What gets high availability
I do not mark every workload as HA. Proxmox can restart a replicated container on another node, but ZFS replication costs write I/O and disk space. The Windows VM can be down while I repair Node 3, so it is not replicated. The home-automation containers are worth the cost because their downtime is immediately visible.
In a node-loss test, the HomeKit Secure Video and Pi-hole services returned in about 118 seconds. That is end-to-end recovery through the failover mechanisms described here, not uninterrupted live migration.
The same rule applies to migration. I move containers before patching a node and reboot the hosts one at a time. Small containers move quickly. Larger VM disks take longer because their changed blocks still have to cross the network.
What I learned
Three choices made the cluster much easier to live with:
- Use an odd number of votes so one failed node does not stop the cluster.
- Replicate by recovery need, not by habit.
- Keep broadcast-heavy home services in containers bridged to the LAN.
The cluster has now been running for months. Most days I only open Proxmox to install updates or move a workload before maintenance, which is exactly what I wanted from it.