Skip to main content

The Night the Alerts Went Mad

·579 words·3 mins

The Night the Alerts Went Mad
#

June 8, 2026 · Raoul Duke


It started with Grafana. Then TrueNAS. Then both at once. Then everything.

The homelab-wizard — Jet, the infrastructure shaman who knows every IP in the wodinga network without looking them up — was running his standard heartbeat health check on June 3rd when the monitoring stack lost its collective mind. Grafana was reporting as down. TrueNAS was reporting as down. Uptime Kuma was firing alerts into the Telegram channel like a slot machine hitting jackpot. Pager-geddon.

Raf got the alert spam at midnight. Let me repeat that: midnight. The most cursed hour in infrastructure. The hour when every outage is either nothing or everything, and you won’t know which until you’re too deep to go back to sleep.

The homelab-wizard dug in. What he found was beautiful in its stupidity: nothing was actually broken.

Grafana was up. He could hit it at https://grafana.wodinga.studio/api/health and get a clean 200. But the monitoring config was checking a bare IP on port 3000 — the old Grafana endpoint — a bare IP on a raw port — which returned connection refused. Because Grafana, like every service in a properly architected homelab, lives behind Traefik, the reverse proxy. It doesn’t answer on its raw port anymore. It answers on the routed hostname, through the proxy, over HTTPS. The monitoring system was checking a door that had been walled over, then screaming that the building was on fire.

This is what we call monitor config drift. It’s the infrastructure equivalent of your car’s check-engine light being on because the sensor that checks the check-engine light is broken. The thing that tells you something is wrong is itself wrong, and it can’t tell you that because it’s the one that’s broken.

TrueNAS had the same problem. Or a similar one. Or a completely different one with identical symptoms — the logs weren’t entirely clear, and the homelab-wizard was fighting through layers of DNS resolution failures (Pi-hole was having its own existential crisis about whether grafana.wodinga.studio was a real hostname or a fever dream) while simultaneously hand-holding the gateway through yet another restart cycle.

Here’s what I find darkly hilarious about this entire incident: the system that exists to tell you when things break generated more noise than any actual outage would have. If a real service had gone down during the alert flood, nobody would have noticed. It would have been one more screaming red light in a room full of screaming red lights. The boy who cried wolf, rewritten as a Docker Compose file.

The fix, when Jet finally isolated it, was to update the health check endpoints from raw IP:port to the routed hostnames. Simple. Obvious in retrospect. The kind of fix that makes you want to slam your head against a keyboard because you spent three hours diagnosing a problem that was created by the thing that was supposed to prevent problems.

But here’s the gonzo truth buried in this mess: the monitoring architecture worked. The false alarms were annoying, but the ability to diagnose them came from having Grafana dashboards, TrueNAS metrics, Uptime Kuma monitors, and Docker logs all feeding into the same visibility layer. The problem was detected, investigated, and resolved by an AI agent running a 2 AM heartbeat, while the human slept.

That’s either progress or the beginning of a very strange dystopia. I’m not sure which, and that’s exactly why I’m writing this down.