Postmortems
- A stale Flux source URL left my cluster open to repojacking After a GitHub username change, Flux kept syncing the cluster from the old repository URL for two months. Anyone who registered the old username could have taken over the cluster.
- A disabled QEMU guest agent puts my Talos nodes in a reboot loop A Talos node ran the QEMU guest agent while Proxmox had it turned off, and reset itself every 70 minutes. Turning the agent off everywhere right after then created the same mismatch on the other two nodes, control plane included.
- A dying PoE port corrupts packets and crashes my storage A decayed PoE port on a Ubiquiti USW Flex Mini corrupted large packets toward the NAS for three days, dropping storage sessions every 12 hours, while every network status light stayed green.
- An NVMe failure corrupts Postgres and app volumes over iSCSI One NVMe SSD died and dropped iSCSI write sessions mid-flight, leaving several XFS volumes unmountable. Every filesystem was repaired with no data loss.
- An SSD failure takes down my homelab production server The single SSD backing the Proxmox LVM thin pool failed and took down all ten VMs. There was no RAID and no backups, so what survived was copied out of the disk images by hand.