An NVMe failure corrupts Postgres and app volumes over iSCSI

2026-07-03 Resolved SEV1 storage hardware iscsi

One of the two NVMe SSDs in my Synology RAID1 pool died. It dropped active iSCSI write sessions mid-flight, which left the XFS filesystems on several LUNs with dirty logs. Kubernetes could no longer mount them, and every workload backed by those volumes went down. I repaired all of them and recovered every byte. No data loss.

Impact

  • Down: authentik (Postgres would not start), sonarr, prowlarr, prometheus.
  • Duration: volumes were unmountable until each XFS log was zeroed and repaired.
  • Data loss: none. The databases replayed their own WAL once the filesystems came back.

Trigger

The drive dropped out of the array. The RAID1 pool went Degraded and the iSCSI sessions serving the LUNs were interrupted mid-write. XFS on those LUNs was left with an unclean log, so the next mount failed with can't read superblock.

Root cause

The failed drive (nvme1n1) did not wear out. SMART showed defective NAND.

  • available_spare: 0% (below threshold)
  • media_errors: 2478
  • percentage_used: 2% (barely any wear)
  • critical_warning: spare-below-threshold bit set

This was a hardware defect, not write wear. I had suspected the monitoring stack of wearing the SSD out with its write volume, but the surviving drive (Samsung 970 EVO Plus 1TB) reports 0% used and 0 errors.

Contributing factors

  • No SMART alerting. The drive was failing silently. I only learned the spare was at 0% by SSHing into the NAS after the outage.
  • Single-session iSCSI targets. A phantom session held the Prometheus target open and blocked the node from logging back in, which added a step before that volume could be repaired.

Detection

authentik-postgresql-0 was stuck in ContainerCreating with mount: can't read superblock on /dev/sdd. Digging into that surfaced the same dirty-log failure on the sonarr, prowlarr, and prometheus LUNs.

Resolution

Talos is immutable, so there is no host shell to run filesystem tools. I ran a privileged rescue pod (Alpine + xfsprogs, hostPath /dev) pinned to the node holding the sessions, then for each LUN:

  1. Confirmed the filesystem with blkid and a read-only xfs_repair -n.
  2. Zeroed the dirty log and repaired with xfs_repair -L (destructive to the log only, not the data).
  3. Verified with a read-only test mount before releasing the volume.

Prometheus needed one extra step. A phantom Synology iSCSI session held the single-session target open and blocked login (error 19). Toggling the target Disable/Enable in SAN Manager cleared it, then the XFS repair proceeded. After repair, all pods returned to Running and the databases recovered their WAL.

What went well

  • No data lost. RAID1 kept one good copy, and xfs_repair -L only sacrifices the journal.
  • Read-only dry runs before every destructive step meant no guesswork.

What went wrong

  • No off-array backups of the Postgres data. Recovery depended entirely on the surviving RAID1 member.

Follow-ups

  • Replace the failed NVMe and rebuild the mirror.
  • Add SMART alerting (available_spare, media_errors, critical_warning) into the monitoring stack.
  • Off-site backups for authentik and the arr apps. Velero takes daily CSI snapshots and moves the volume data to Cloudflare R2, off the NVMe pool and off the NAS.
  • Reclaim thin-provisioned space with fstrim once the pool is healthy, and add discard to the storage class so reclamation stays automatic.