Nobody scheduled it. That is the part that matters.

A total power loss hit the estate. Every physical host, the distributed storage layer, the database cluster, the hypervisor control plane, and the Kubernetes cluster went down hard, at once, with no graceful shutdown. One host then failed to come back at all — its BIOS boot order had been lost and needed repairing at the console.

This is not a drill write-up. A planned drill six days earlier had produced the recovery order, but this was the real thing: the kind of event that finds what a drill cannot.

The recovery, in dependency order

The whole lesson is the order. You recover bottom-up, and you verify each layer before you start the next:

  1. Physical hosts — power on; repair the lost boot entry at the console.
  2. Distributed block storage — nothing to do. A hardening change made days earlier (disabling automatic eviction) meant zero evictions; the satellites reconnected and resynced unaided. Verify only.
  3. Database cluster — self-formed into a primary component, 3/3 synced, no manual bootstrap, thanks to an earlier hardening pass. Manual bootstrap is still the fallback: self-forming is not guaranteed when nodes die at different times.
  4. Hypervisor control plane — came up on its own once the database was there.
  5. Virtual machines — resumed in waves: gateway, then control plane, then storage and workers, then service VMs. Expected, and hit, a retry pass — resumes issued while replicated disks are still reconnecting fail, because the volume has no connected up-to-date peer yet. Six of sixteen hit it; a second pass ten minutes later succeeded for all.
  6. Network gateway pair — the hard part. More below.
  7. Kubernetes — started by hand on every node, by design. Control-plane nodes took 5–15 minutes each for quorum and defragmentation; one was IO-starved by three VMs cold-booting simultaneously against network-remote disks. Patience, not a defect.
  8. In-cluster convergence — largely automatic once the API was up, except the secrets chain and an image-availability trap.
ElapsedPhase
0:00–0:15Assess; verify storage, database and hypervisor self-recovery
0:15–0:35VM resume waves including the retry pass; start Kubernetes
0:35–1:05Wait out an IO storm; diagnose and fix the gateway failure
1:05–1:50In-cluster recovery, secrets unseal, image steering, verification

Result: everything recovered. No data loss. No split-brain. All VMs running, all Kubernetes nodes ready, secrets management HA and unsealed, every GitOps application synced and healthy, storage fully up-to-date, monitoring confirmed fresh, the CI runner re-registered on its own.

This ran under the estate’s operating model — AI-directed, human-owned: the agents do the work under direction, and a human owns design, prioritization, and acceptance. The disaster-recovery plan of record here is an agent with a recovery skill rather than a human runbook, because a runbook is useless if the person holding it is asleep.

What the unplanned outage exposed

The value of an unscheduled outage is that it finds what a drill cannot. This one found four things:

  • The “HA” gateway pair was a cold-start single point of failure. A known reload race killed the load balancer on both gateways during boot, and the watchdog meant to catch exactly that had never started — its log was empty. The active gateway held the virtual IP as a black hole: answering ARP, serving nothing. Fixed the same evening with file-locking around config regeneration and the watchdog rebuilt under a supervisor with heartbeat logging, then validated by hard-rebooting both gateways. Failover measured at ≤13 seconds.
  • Egress lockdown plus an incomplete internal mirror is a reschedule trap. Pods rescheduled onto a node without a cached image cannot pull, because the outside world is deliberately unreachable. Hit the CNI operator, cert-manager, the ingress controller, and more. Worked around live by locating the node holding the cache and steering pods to it. Fixed by populating the mirror and configuring every node to mirror upstream registries.
  • A certificate was missing a DNS subject alternative name, so the cluster had to be reached by address rather than name mid-recovery. Fixed at the template level so new nodes inherit it.
  • A race between disk reconnection and VM resume, self-inflicted by recovering faster than storage could reconnect. Now documented: expect a retry pass, or wait for the replication links to establish.

The same evening, the theory-written runbook was replaced with a verified one, and what the outage exposed was fixed.

The honest comparison

My previous-generation platform, on an immutable Kubernetes distribution, recovered from the same outage completely hands-off. This one needed manual gates. Most are deliberate — I want a human deciding when a control plane rejoins. One was not: an automatic VM-resume hook I had specified and never built would have removed roughly twenty minutes of manual work and the disk race entirely.

I am telling you that because it is true, and because the distinction between a deliberate manual gate and an unbuilt one is exactly the judgment you want in the person running your stack.

What this means for you

This is what “Run” looks like when it is real: a platform that has taken a total, unscheduled power loss and come back with zero data loss — and an operating discipline that converts every incident into a verified runbook and a permanent fix, the same evening.

The full, unabridged incident record — this one and five others, written the way I would want to read them — is in the public kws repository. Every claim in this post links back to it.