Known blind spots

This lab borrows production patterns, but it is still one person's home lab on a residential connection. Some failure modes are deliberately not covered, because closing them would cost more than an outage does here. This page lists them.

The monitoring runs on the thing it monitors

Prometheus, Alertmanager, and the pipeline that pushes alerts to my phone all run inside the Kubernetes cluster, and the blackbox probe that checks https://khider.fr runs inside the same network as the site. If the cluster or the connection goes down, nothing outside the lab can page me. I find out by noticing the outage myself, and for these workloads, discovering a full outage a few hours late costs nothing.

One site, one ISP, one power feed

Everything lives in one place, behind one residential connection and one public IP, with no failover. Cloudflare hides the origin, but a power cut or an ISP outage takes everything down until it comes back.

Every byte of the cluster's state sits on one NAS

The Kubernetes cluster has three nodes, but all of their persistent volumes are served over iSCSI by a single Synology NAS, already the root cause of one postmortem and a contributing factor in another. RAID1 covers a dead drive, not a dead power supply, a bad DSM update, or a controller bug. Nightly off-site backups to R2 cap the damage at restore time plus the last day of data. The few containers that still run outside the cluster (qBittorrent and the monitoring exporters) run on the NAS itself, so a NAS failure takes those down with the cluster.

One hypervisor for the whole cluster

The three Kubernetes node VMs all live on the same single Proxmox server. Kubernetes can lose one node and keep going, but not the machine all three run on. The reverse proxy, the CI runner, and the DNS resolver the whole LAN uses all run inside the cluster now, so losing the hypervisor also takes name resolution down for every device at home. A second server used to spread this risk, but I retired it because the heat and noise at home outweighed the redundancy.

On-call is one person

There is no rotation and no escalation. If something breaks while I am asleep or away, it stays broken until I get to it. Alerts re-notify every twelve hours, but there is no bound on time to recovery.

The deploy path assumes the lab is healthy

Flux runs inside the cluster it reconciles, and the self-hosted GitHub Actions runner that applies Ansible and Terraform changes runs in that same cluster. If the cluster is down, every deploy path is down with it, and recovery has to start by hand.

Why publish this

Writing these down makes me decide whether each risk is accepted or just ignored. The ones that stop being acceptable end up on the roadmap of the repo, and the ones that cause incidents end up in the postmortems.