Files
nexus/sreweekly/markdown/200/02-a-conjecture-on-why-reliable-systems-fail.md
2026-09-12 17:23:01 +08:00

37 lines
2.4 KiB
Markdown
Raw Blame History

This file contains invisible Unicode characters
This file contains invisible Unicode characters that are indistinguishable to humans but may be processed differently by a computer. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# A conjecture on why reliable systems fail
- **期号**: SRE Weekly Issue #200(2019-12-29)
- **作者**: Lorin Hochstein
- **链接**: https://lorinhochstein.wordpress.com/2017/06/24/a-conjecture-on-why-reliable-systems-fail/
## 简介
> Once a system reaches a certain level of reliability, most major incidents will involve:A manual intervention that was intended to mitigate a minor incident, orUnexpected behavior of a subsystem whose primary purpose was to improve reliability
## 正文
*(Some of my co-workers call this [Lorin’s Law](https://twitter.com/theckman/status/1185233751028060161))*
Even highly reliable systems go down occasionally. After having read over the details of several incidents, I’ve started to notice a pattern, which has led me to the following conjecture:
Once a system reaches a certain level of reliability, most major incidents will involve:
- A manual intervention that was intended to mitigate a minor incident, *or*
- Unexpected behavior of a subsystem *whose primary purpose was to improve reliability*
Here are three examples from Amazon’s post-mortem write-ups of major AWS outages:
The [S3 outage](https://aws.amazon.com/message/41926/) on February 28, 2017 involved a *manual intervention* to debug an issue that was causing the S3 billing system to progress more slowly than expected.
The [DynamoDB outage](https://aws.amazon.com/message/5467D2/) on September 20, 2015 (which also affected SQS, auto scaling, and CloudWatch) involved healthy storage servers taking themselves out of service by executing a distributed protocol that was (presumably) designed that way *for fault tolerance*.
The [EBS outage](https://aws.amazon.com/message/680342/) on October 22, 2012 (which also affected EC2, RDS, and ELBs) involved a memory leak bug in an agent that *monitors the health of EBS servers*.
In the fun book “The systems bible” (which I recommend) it is observed that (among other antics), “Fail-safe systems fail by failing unsafely”.
It is quite frequent that safety in a system is obtained by pulling in some kind of subsystem designed for that goal. Therefore, when the subsystem fails… I suppose you can design a system such that safety is derived from how the system is structured (rather than added as an extra), but I’m not sure if you can do that on purpose!
I’m curious about the status of the conjecture now, four years later.
I think it has held up pretty well.