Eric Li DevOps & Infrastructure

Green Dashboards, Two Real Faults

Homelab · 2026-09-21 · Monday · 11:33 PM · 3 min read · 80% AI · Eric Li

Late in the afternoon a lot of the Linux ISO pipeline had disconnected again, mainly the download client from the two apps that feed it. I asked Claude to find out what was actually working before anything got touched.

Two faults no monitor saw

Three diagnostic agents went out in parallel: one on the download-client health of both apps, one on the seedbox and its quota, and one sweeping the rest of the lab. Most of what looked broken had just moved hosts. Under that noise were two real faults, and every monitor watching them was green.

The first was one app holding a single dead connection to the download client for close to 29 hours, reusing the same broken socket for every request. The other app hit the same endpoint with the same credentials and was fine, which ruled out a credential problem and made the first one worth chasing. The second fault was the subtitle app. It still had the old API keys for both main apps, left over from a key rotation that had been swept everywhere except there, and twice an hour it threw a JSON parsing error trying to read a 401 response as data.

Once I said fix everything, Claude repaired both and tested each with real work. The first came back after a docker restart, followed by a real search and grab that landed a point release that had been failing all day, and an actual import a couple of minutes later. The second needed the subtitle app stopped first, since it rewrites its own config on shutdown and would have overwritten a live edit. Then its two stored keys got corrected and the app restarted. Its item count settled to match the other app's file count exactly, which showed the sync was live again. A stuck queue of 17 items turned up along the way: torrents whose release name looks like a normal file but whose payload is a Windows executable. Claude left those alone, since deleting files off the seedbox wasn't something I'd asked for.

Why nothing caught it

Later that evening I asked the obvious follow-up: how do we stop these from breaking again. Claude went back through the monitoring this lab already has before suggesting anything new. Of 46 monitors on the primary dashboard, 33 only prove something is listening, and several of the *arr apps are set to count an authentication failure as up, so a dead credential stays green. The subtitle app and the collection-cleanup app have no monitor at all.

The freshness checks were the part that stung. They were built for this kind of failure, after an earlier incident where a database sat corrupt for thirteen days while its container reported healthy. One confirms a scan finished and the other confirms a sync job ran. Both of today's faults broke the step right after the one being checked, so the jobs ran and the checks passed while nothing got done.

Claude wrote up three changes in priority order. First, check each app's own internal health verdict, because its front door answers even when the work behind it is broken. Second, extend an existing credential-check pattern to every place an API key gets copied, where today it covers only the one that caused the last incident. Third, flag queue items stuck behind a bad file type. None of it is built yet. The write-up is filed until I decide which to do first.

The next morning, checking it held

A leftover task notification from one of the original diagnostic agents came through overnight with nothing new in it. Since a night had passed, Claude used it to check that both fixes had held. The first app had grabbed three more point releases on its own since the repair, so automatic acquisition was working again. The subtitle app showed zero parse errors in twelve hours, so the twice-hourly failures were gone. The stuck queue of fake executables and the choice of which prevention work to build first were still waiting on me.