A Server Lost Power at 00:32. We Found Out at 08:18
What happened
One machine stopped, and S3, the container registry and every deployment stopped with it for eight hours. Our monitoring caught it in three minutes and told nobody. Here is the full postmortem: why single-machine redundancy was not what we thought it was, the near-miss we nearly caused ourselves, and the pager and hardware watchdog we shipped the same day.
Summary assembled by rule from the sources below
Why it's spreading
Timeline
- First appeared on Hacker NewsHacker News