Eric Li DevOps & Infrastructure

Moving the Floor While Standing on It

Homelab · 2026-09-27 · Sunday · 10:07 PM · 6 min read · 75% AI · Eric Li

The plan for today was to move the whole four-node Proxmox cluster off LVM-thin storage and onto ZFS, one node at a time, renumbering the nodes into x.x.0.51 to .54 as they went. It got done, after two incidents and after the hotel codename work from the night before was carried out just after midnight.

Codenames everywhere, and a count that was wrong

Every real location name for the hotel client got replaced across the vault, NetBox, the skills, memory and the scripts in ~/bin, so none of it needs redacting later. A rollback snapshot of the NetBox database came first. Then folders and files got renamed, four sub-agents and Claude's own pass reworked the prose, and NetBox's site, tenant, buildings, wireless network and access points got relabelled through the API. Two ~/bin commits went out at 00:28 and 01:20, and the SSH alias for the on-site travel router got renamed with them.

The NetBox sub-agent reported object counts unchanged across 50 endpoints and zero old names left. Claude ran its own scan and found about forty. A regex that matched one name with or without a space but missed the hyphenated form explained part of the gap. The rest came from a second sub-agent that had been told to apply a rename through a message relayed from the lead session. Its permission classifier read that as instruction poisoning and blocked it. Claude's first run of the follow-up fix script was denied as a change to a shared resource, and it went through once I said "fix those". A connectivity test that afternoon confirmed the renamed router alias still works.

Also after midnight: the verified copy of the cabling standard from yesterday got finished as a skill, and I asked what had pulled my whole network into NetBox. The answer was a sync script, a few hundred lines of Python that reads UniFi's API for switches, APs and VLANs and Proxmox's API for every node and guest, then writes it all into NetBox with a scoped token and credentials fetched from the password manager at run time. It runs dry by default and needs a separate flag to write, and another to mark vanished objects offline instead of deleting them. I asked the same question again at 19:28 and got the same answer.

Two incidents before the rebuild started

At 03:57 pve01 went down for a RAM upgrade. datacenter.cfg still had shutdown_policy=migrate from the HA setup in August, so the HA manager tried to move Home Assistant and the LLM proxy to pve03. It moved their configs but not their disks, because neither guest lives on shared storage any more, and both landed in an error state. Home Assistant was down until I got in from PowerShell around 04:15, moved the configs back by hand and started the guests on pve01 directly. HA had been built around guests living on the NAS, which stopped being true when the NAS got unreliable. All it was giving me now was fencing risk, so once things were stable I had every HA resource removed. Guests start in place with onboot=1, the way they used to.

Then at 06:17, picking the day's project back up, a live migration of n5ubuntu from pve03 to pve01 with --with-local-disks filled pve01's thin pool. The mirror writes the full 150G virtual disk, not the space actually used, and pve01 had about 109G free. The pool hit 100% around 06:25, and one of the things on it was CT 109, the container this Claude workspace runs in. Its root filesystem went read-only mid-write, and Home Assistant paused on an I/O error. A stop, an fsck and a start brought the container back clean a few minutes later, but anything written in the minutes before was lost. A handoff note written on pve01 itself while the container was read-only is what the next session started from at 06:34.

Rebuilding the cluster node by node

pve02 went first: evacuated, pulled from the cluster around 08:36, wiped, reinstalled on a single ZFS disk and rejoined at x.x.0.52. Its RTL8125 2.5G card came out and went into pve04. pve03 got re-addressed to .53, rebuilt, and had its guests moved off. pve04 got reinstalled at .54 on a striped ZFS pair and joined as the fourth node.

pve04 was the awkward one. It's running on a 90 W power brick when its factory supply is 180 W, so under load the supply asserts PROCHOT and the CPU clamps itself. The fix is two systemd services, one capping the RAPL power limits to what the brick can deliver and one clearing the BD PROCHOT bit so the CPU ignores the brick's signal. The first reboot test failed: an ordering cycle between the two units made systemd delete one of the jobs, the power limits reverted to the firmware's 65/145 W, and the CPU sat at 200 MHz. After fixing the unit ordering, a second real reboot at about 12:05 came up correct. It stays capped until the proper brick arrives.

Guests moved all day. n5monitor and n5ubuntu went live from pve03 to pve02 around 11:00, and PBS went to pve01. At 11:11 pi02, the Raspberry Pi, got set up as a corosync QDevice, which gave the four nodes a fifth, tie-breaking vote.

Emptying pve01, and the node that forgot its vote

By early afternoon pve01 was down to the one guest that couldn't move casually, this container, since stopping it ends the session. Before that move, at 14:53, the NAS went off. Immich and PBS were stopped first and the NAS-backed storages detached from every node, so nothing would hang looking for a share that wasn't there.

pve01's own rebuild ran from a session on my PC, not from the container on the box being wiped. At 15:36 I was back on pve01, on ZFS at .51 and rejoined. The rebuild had dropped the QDevice out of the corosync config, so the cluster counted four votes instead of five, and re-adding it stalled because pi02 refused SSH from pve01's freshly generated key. That's still open.

Set-top boxes and where the cable map lives

In the late afternoon I pulled screenshots off the three bench set-top boxes over adb. They were reachable but asleep, and a sleeping box screenshots as a black frame. Waking one takes a remote keypress over the same link, and since that can also wake the TV over HDMI-CEC, Claude checked first before sending it. All three showed live channels once woken, so the tuner-to-box path works end to end on the bench. The same session ran the numbers on scaling one stream per box to all hundred rooms: unicast grows with every room watching and multicast doesn't, and a hundred rooms of HD unicast gets close to what a single delivery link can carry.

A research pass on where the hotel's cable map should live came back recommending NetBox as the master record, with the spreadsheet generated from it for handover. I haven't answered that yet. NetBox got ahead of the decision anyway: by evening the hotel's three racks were loaded from an elevation sheet, one of them as a full 42U, with two rack-visualisation plugins installed.

The last check of the night, at 22:05, found all four nodes up and quorate, every ZFS pool healthy and barely used, and no failed units. The portal, n8n and the LLM proxy were still stopped and not set to start on boot, and the QDevice isn't voting. Those are tomorrow's jobs.