I was up from just before 1am and most of the work was done by about 7:15. It began with a question about which bit of network gear to upgrade next and ended with a Kubernetes cluster running on the spare node. There was a short stretch at 15:20 to tidy up before heading to the hotel.
The gateway and its memory
I started by working out which piece of network gear most needs upgrading. My guess was the gateway, since it sometimes drops into safe mode when utilisation gets too high. The 30-day metrics point at memory. CPU averages 33% and doesn't move with traffic, and the WAN averages 3.7 Mbps. Memory sits near 70% after a reboot, climbs about a point a day, and reaches 91 to 95% after two or three weeks. It spent 23% of the month above 85%, which is where community reports say safe mode kicks in. The gateway has 3 GB of RAM that can't be upgraded, and the camera recorder and intrusion prevention both run on it.
So I'm not replacing it yet. The cheaper fixes are trimming intrusion prevention or moving the recorder off the gateway. I haven't done either. The runsheet is updated, including the firmware version, which was 5.1.33 where the old note said 5.1.19.
Network audit and firewall changes
Next I audited how everything is cabled and configured. Captures showed clean separation between the networks. The weak spots were mostly in the firewall: the IoT and Home Wi-Fi networks use the same password, IoT and camera devices can reach the gateway's admin page, and two rules were wider than they needed to be. One server is also sending DNS straight out to a public resolver, which will skip my own DNS once I switch over.
I applied the changes one at a time, each previewed first and read back afterwards, and again about a minute later to catch any silent revert:
- Disabled the rule letting cameras reach the gateway. The doorbell hub kept working without it.
- Narrowed the Home-to-server rule from every port to 13 service ports on three hosts, and dropped a dead entry.
- Turned on logging for four IoT rules so I can narrow them after a week of data.
- Disabled one camera rule that never matched anything, and deleted a rule tied to my phone's MAC address, which changes constantly.
- Moved the TV and the house alarm onto IoT. The TV picked up an address straight away. The alarm still had a reservation from the management network, so I fixed that and it still hasn't taken a lease.
The Hue bridge
At 03:27 I moved the Hue bridge onto the IoT network, assuming that was safe. The bridge has a static address set on the device itself, so it ended up on the wrong network with an address that doesn't work there, and neither the Stream Deck nor the Apple Home app could reach the lights. I only noticed at about 05:53, when Home Assistant looked broken, and my first guess was the firewall changes, which turned out to be wrong.
I moved the port back to Home, the bridge answered again, and it's staying on Home. If I ever try again, the order is: set the bridge to DHCP in the Hue app, move the port, then repoint Home Assistant.
Home Assistant through the second tunnel
ha.n5hq.me had been failing for some requests. Cloudflare spreads traffic across two tunnel connectors, and Home Assistant's trusted proxy list only had one of them, so the other got a 400. I added the second connector, restarted Home Assistant (about a minute offline), and both now return 200. I ruled out any Cloudflare-side change such as extra access policies, because it risks breaking Home Assistant and Apple Home.
Emptying pve04
Around 05:00 I destroyed the old container my workstation VM was converted from. It had been sitting stopped with a stale lock, so I unmounted and removed it by hand. That means the VM now has no container to roll back to, only its replica and the nightly backup.
Then I cleared the guests off pve04, because I want it as a Talos test node. That meant six new replication jobs to pve01 and pve03, each synced cleanly twice, then renaming four HA rules so none of them mention pve04, then deleting the six old jobs. pve04 now has 468 GB free, 12 GiB of RAM, no HA rules and no replicas. I checked it by counting pve04 mentions in the HA config, which came to zero.
Building kuben5
I worked from a long build walkthrough video and the official Talos docs for v1.14, which I mirrored locally. At about 06:10 I created three VMs on pve04: one control plane with 4 cores and 6 GiB, and two workers with 2 cores and 3 GiB each. All three got fixed addresses and were added to the admin allow-list in the firewall and skipped in the nightly backup.
Generating the config, applying it and bootstrapping took until 06:27. I hadn't expected v1.14 to set the hostname through its own config document instead of the old network setting. The first health check complained it couldn't find the control plane node, and a second run passed. The cluster, kuben5, runs Talos v1.14.2 and Kubernetes v1.37.1, and all three nodes are Ready. The cluster secrets are in the password manager and kept off the shared folder.
After that I registered the VMs in NetBox, which also refreshed 11 VM records and 39 client records, and a second dry run found nothing left to change. I wrote up the build procedure so I can repeat it at another site. Then I added Helm v3.22.0, local-path storage, MetalLB v0.16.1, Traefik v3.7.13, and a test Uptime Kuma on its own address. I deleted its pod to check that a file written to the volume survived.
Still open
- The gateway memory problem has a diagnosis but no fix yet.
- The alarm on IoT still has no address, and I need to replug it or change its own network settings.
- The two disabled firewall rules can be deleted after a quiet few days, and the IoT rules get narrowed after a week of logs.
- The Hue bridge still has a stale reservation on the IoT network that does nothing.
- The Proxmox kernel update on pve01 is waiting on me, because it needs a reboot that takes down my workstation VM and others.
- The test Kuma in the cluster has no DNS name and hasn't been set up yet.
- I'm going to the hotel later to work out what the servers there actually run.