← Back to events
ActiveTech

A Server Lost Power at 00:32. We Found Out at 08:18

What happened

One machine stopped, and S3, the container registry and every deployment stopped with it for eight hours. Our monitoring caught it in three minutes and told nobody. Here is the full postmortem: why single-machine redundancy was not what we thought it was, the near-miss we nearly caused ourselves, and the pager and hardware watchdog we shipped the same day.

Summary assembled by rule from the sources below

Why it's spreading

Timeline

  1. First appeared on Hacker NewsHacker News

Sources