← 返回事件
持续讨论科技

A Server Lost Power at 00:32. We Found Out at 08:18

发生了什么

One machine stopped, and S3, the container registry and every deployment stopped with it for eight hours. Our monitoring caught it in three minutes and told nobody. Here is the full postmortem: why single-machine redundancy was not what we thought it was, the near-miss we nearly caused ourselves, and the pager and hardware watchdog we shipped the same day.

摘要按规则整理自下方来源原文

为什么在扩散

时间线

  1. Hacker News 最先出现Hacker News

来源