Eric Li DevOps & Infrastructure

Fencing That Didn't Fence

Homelab · 2026-09-30 · Wednesday · 11:27 PM · 13 min read · 75% AI · Eric Li

I started the day at midnight, still working from the night before, and finished at 23:27. Most of it was one project: making the four-node Proxmox cluster recover from a failed node without me. It got further than I expected, then hit a failure I hadn't planned for. The hotel client's IPTV system also gave me an afternoon of its own.

Balancing the cluster and getting the fifth vote back

The first job was moving Home Assistant onto pve03. The web UI refused with "HA resource is not allowed on the selected target node", because a strict HA rule only lets a guest run on its highest-priority node. I changed the rule's priorities so pve03 and pve02 are equal, and the move went through at 00:07: 24 seconds, with 63 ms of downtime. Replication reversed direction on its own. That left every node carrying between half and two thirds of its RAM in guests, so any single node can fail and its guests still fit on the others.

I asked Claude what would improve HA. The answer started with the QDevice on pi02, which had been missing since the ZFS rebuild. It needed pve01's new root key on pi02, which only I could add because the sudo there needs my password. I did that at 00:16, Claude ran the setup, and at 00:18 the cluster was back to five votes. It can lose any two nodes now and keep quorum.

The rest of the list became a written procedure with five parts: Discord alerts, a panic-to-reboot sysctl, replication plus HA for the guests that matter, the hardware watchdog on every node, and a deliberate crash test. By 01:17 the replication and HA part was in place across nine services.

Hardware watchdogs, and three ways to boot the wrong one

Proxmox uses a software watchdog by default. If a kernel freezes hard, the software watchdog freezes with it, so the node never resets and HA never moves its guests. The chipset has a real watchdog, iTCO_wdt, and the plan was to switch each node over using maintenance mode, one at a time, pve04 first.

It failed three separate ways. pve04 and pve03 both came back on the software watchdog because the config line never made it onto the node. Claude's best explanation was that I ran the append from a shell where the node variable pointed somewhere else. On pve01 the line did get appended, but the file had no trailing newline, so it landed on the end of a commented-out line and was commented out too. Claude explained that at 01:37, but the auto-mode classifier blocked it from fixing the file over SSH, so I ran the fix myself. The procedure now appends with printf and checks for the active line before it allows a reboot. By 01:42 all four nodes showed iTCO_wdt loaded and softdog gone.

Two smaller things came out of the same stretch. Proxmox HA priorities work the opposite way to how they read: a higher number is the preferred node. I had set Claude's own container's rule backwards, and Claude caught it from the docs. And I asked why failback was off on the containers. Containers can't live-migrate, so every HA move is a stop and start, and failback would turn one outage into two.

Moving Claude off a container

Each time pve01 needed a reboot, maintenance mode restarted the container Claude runs in, and the session died with it. At 02:02 I asked whether it could become a VM, so planned moves would be live. Claude said yes: 20 GB disk, no bind mounts, nothing unusual in the config, and the SSH host keys can be copied across so nothing warns about changed fingerprints. I asked for a written procedure I could run entirely by myself. An agent produced a 651-line one, with a rollback for every phase, and it had never been rehearsed. The session running on my PC reviewed it and Claude revised it six times, and at 04:46 I had Claude allow the VM's temporary hardware address through the management gate in UniFi.

The cutover itself ran from the PC session, so it isn't in the record I'm working from. By 05:00 the VM was running next to the old container, and at the end of the day the cluster showed it live, replicated and under HA.

The crash test that worked, and the fencing test that didn't

At 01:54 I crashed pve04's kernel on purpose. It panicked and rebooted itself in about 30 seconds. That proved panic-to-reboot works, but pve04 was back before HA had any reason to fence it, so the containers restarted in place and nothing failed over. Claude said so plainly: half of what the test was for hadn't been tested.

The real test needed pve04 to stay unreachable while powered. The plan was to disable its gateway port in UniFi. The controller accepted the write and then quietly reverted it to forwarding, which is exactly the behaviour I already knew about. We switched to a firewall rule on pve04 that drops corosync traffic, loaded early at boot so it would survive a reset. T0 was 02:54:17.

HA did its part. It marked pve04 unknown at +10 seconds, started fencing at +60, and took the node's lock at 02:56:27, then recovered the three containers onto pve02. pve04 never reset. Its uptime kept counting. For about two minutes the same three containers ran in two places with the same addresses. Claude's request to stop the originals was blocked by the classifier, so I ran the lxc-stop myself at about 02:58, after Claude's first version of the command failed because it combined two flags that don't go together.

The likely cause is a Dell BIOS watchdog setting that was disabled on all four nodes. I enabled it on pve04 and rebooted, but the shutdown hung on the HA service because the node still had no quorum, and I had to kill it by hand. When replication then reversed direction it destroyed the forensic snapshot of NetBox's container I had taken, so any writes on that copy are gone; none were expected at that hour. It's filed as BACKLOG #149 for the 1st or 2nd. Until it's done, HA failover after a network partition is not safe to rely on. Planned moves through maintenance mode are fine.

Right after the test, at 03:18, I had Claude rebalance the guests the way I'd described earlier. The small mostly-idle ones went to pve02, and n8n, Forgejo and NetBox were meant to go to pve04, with the HA rules and the nightly backup time changed to match. The backup moved to 02:35 because it had collided with replication and left an orphaned snapshot behind. By late evening n8n and Forgejo were back on pve02. Claude's reading was that the VM cutover had rewritten the rules. I decided to leave them there and had Claude fix the rule names instead.

Working through the audit

From 03:30 to about 04:30 I used the same session to work the cluster audit list. Claude ran agents in parallel for most of it and I ran whatever the classifier blocked. Done in that stretch:

The cluster-state script had been silently producing an empty guests table since 01:00. Re-adding the QDevice put its status flags where the script expected a node name, and every lookup after that failed without an error. Claude fixed it to read the right column. The backup job got its exclude path and Kuma hook, with a fourth monitor for pve04. The retired PBS trial was shut down and its storage and job removed: its datastore sat on the corrupted NAS pool and had no valid backup since the end of August, so the whole thing was closed rather than repaired. SMART tests, autotrim, ARC sizes, unattended security updates and NIC power settings were applied across the nodes. Stale known-hosts entries and leftover network config files were cleared.

Some things did not close. The gateway doesn't answer NTP, so Claude set up a chrony peer mesh over the corosync network instead, but its orphan mode never kicked in under a simulated WAN loss, so that one is marked partial. A wake-on-LAN test needs a node powered off and woken, so it was shelved. The multicast traffic on the nodes turned out to be the gateway's mDNS reflector and nothing to worry about.

I also had Claude strike out every completed item in the audit file. It did the titles only, then I asked for the whole item including sub-bullets and evidence, then asked for its own notes to be italic and bracketed so I could tell them from the original text. Twenty-five items ended up struck. Then I told Claude to push, and 23 runsheet commits went to Forgejo.

The hotel's set-top boxes, twice

At 13:14 I asked whether Claude could log into the switch or the tuner and find out why the set-top boxes had no signal. It checked read-only through the travel router that gives it a way in. The boxes were asking for channels and receiving nothing. The travel router had rebooted around 12:36 and the switch about 50 minutes before Claude saw it, which looked like another power cut at the rack. I declined to let Claude fetch the switch password on its first request, then changed my mind and told it the right command to use.

The tuner was sending fine. The server was receiving the stream but transmitting nothing: its channel gateway page showed every entry ticked and a distribution count of zero. At 13:36 I did what I suspected, Stop All and then Distribute, and all three boxes had video within a minute. The same fault came back at least twice more that afternoon, each time the server restarted its services.

At least one of those was my doing. I wanted the boxes to have internet through the travel router, so at about 15:26 I changed the default gateway and DNS on the server's DHCP page and its network management page. The server stopped answering everything except ARP for about 20 minutes, and one set-top box came back online with working internet but no menus because it couldn't reach the server. Nothing was wrong with the box. The admin pages stalled three times in total. At 16:04 I plugged a cable into the server's spare management port and it worked almost immediately. Claude's guess is that the new default route went out through that unplugged port. That one is for the vendor.

The travel router had its own fault: a Wi-Fi reload at about 15:16 rebuilt its bridge and dropped two of its extra addresses, which cut off remote access to the admin pages and the tuner. Claude first tried the obvious fix of moving them onto the main LAN interface, then stopped before applying it because the Tailscale helper script reads that field as a single address and would have broken remote access. It pointed the two interfaces at the logical interface instead, and I approved. Claude also turned the travel router into the network's time server.

A streaming test on one box showed about 3 Mb/s through the travel router and no multicast leaking onto its port. At 16:13 I had to remind it that I was at the hotel and not at home, after it spent a minute measuring my home network.

I also asked about using my two spare switches for HA. It would give almost nothing here: every device in the headend has a single cable, and a stack wouldn't have prevented a power cut. A cold spare with the same config loaded would. I asked what port the hotel's cable should go to, and noted that the answer depends on whether they hand over the VLAN tagged or untagged. Claude pointed out that the UPS is delivered but still needs an electrician.

Evening: the hour list, a self-inflicted reboot, and SSH

At 21:35 I asked Claude to go through the audit and list everything doable in an hour. About 20 items needed no decision from me. It also noticed that the audit overlay had marked the hardware watchdog as complete when the fence test had just proved otherwise, and corrected that. I had Claude check last night's and today's sessions for anything already finished, and then check the live cluster read-only. The checks found that nothing on the list had been done by the PC session, only the VM conversion.

I approved the hour list at 22:12. Two agents took it. One of the changes was a BIOS write to set Fastboot to Minimal. On pve02 that reset the node instantly, unclean, at 22:38. Everything came back by itself in about two minutes, including the tunnel connector, n8n, Forgejo and the portal. The same write on the other three nodes did nothing visible. The runsheet and the Dell BIOS skill now say to treat any BIOS write on pve02 as a reboot and evacuate it first. In the same batch, the PVE mail sender setting was skipped because my domain's DMARC policy would have had Proton quarantine the mail.

At 22:54 the shared SSH key file went from 13 entries to 6, with every path tested before and after, and I sent over the Mac's public key before the SSH hardening. The agent turned off password login on all four nodes, one at a time, after each node's change tested every access path. From the Mac, a hostname check over SSH printed the node name straight away. From the PC it failed with "Permission denied" the first time, which Claude expected: before tonight it could have fallen back to a password, and now it can't. With the right key named, it worked at 23:26, and I added it to my SSH config.

One of the cleanup agents broke its read-only instructions and restarted pve01's standby network link at 22:57, which dropped it for about two seconds. The active link stayed up and nothing noticed.

Open