SRE weekly 所有文章

This commit is contained in:
2026-09-12 17:23:01 +08:00
parent 409b40ddcb
commit af7633f9dc
8486 changed files with 4489990 additions and 7 deletions

View File

@@ -0,0 +1,147 @@
# Failure—Is It A Matter Of When?
- **期号**: SRE Weekly Issue #179(2019-08-04)
- **作者**: Barry O’Reilly
- **链接**: https://barryoreilly.com/failure-is-it-a-matter-of-when/
## 简介
This is an engrossing write-up of the Chernobyl incident from the perspective of complex systems and failure analysis.
## 正文
I’ve been watching the outstanding HBO series Chernobyl which details the worst nuclear reactor meltdown in human history—an event that was approximately 400 times more potent than the atomic bomb dropped on Hiroshima.
What occurred to me, and what I discovered over the course of the TV series, this was a cataphoric failure destined to happen. A slow drift into failure from the beginning—only expedited by the people executing the experiment their rules-based culture required. But why? Is failure just a matter of when?
![Chernobyl Disaster](https://barryoreilly.com/wp-content/uploads/2019/07/Cherno_Title.jpg)
## The Mystery Is How Anything Ever Works At All
In the pursuit of success in our dynamic, ever-changing and complex business environment with limited resources and many conflicting desired outcomes, a succession of tiny decisions eventually can produce breakdowns—a domino action of latent failures—on a tremendous scale.
From news feeds to newspapers, daily debacles highlight how the systems we design—with positive intent—can create more unintended consequences and negative effects on the society those very systems are designed to support.
From [Facebook hacking](https://en.wikipedia.org/wiki/Facebook%E2%80%93Cambridge_Analytica_data_scandal) to [Boeing 737 Max accidents](https://en.wikipedia.org/wiki/Boeing_737_MAX_groundings), [algorithmic autonomous-bot arguments](https://www.wired.com/2017/03/internet-bots-fight-theyre-human/) to legacy top-down management structures, information flows and decision-making we struggle to cope with much of our context, to the point its amazing that anything ever works as intended at all.
When problems occur [we hunt for a single root cause](https://www.verica.io/inhumanity-of-root-cause-analysis/), that one broken piece or person to hold accountable. Our analyses of complex system breakdowns remains linear, componential and reductive. In short, it’s *inhumane*.
The growth of complexity in society has outpaced our understanding of how complex systems succeed and fail. Or as Sidney Dekker human factors and safety author said, “Our technologies have gotten ahead of our theories.”
## Modeling The Drift Into Failure
Another pioneering safety researcher Jens Rasmussen identified this failure-mode phenomenon which he called “[drift to danger](https://risk-engineering.org/concept/Rasmussen-practical-drift)”, or the “systemic migration of organizational behavior toward accident under the influence of pressure toward cost-effectiveness in an aggressive, competing environment” ![Rasmussen](https://barryoreilly.com/wp-content/uploads/2019/07/Rasmussen.001-e1563854913448.jpeg)
*Rasmussen illustrated the competing priorities and constraints that affect sociotechnical systems, as shown above.*
Any major initiative is subjected to multiple pressures and our responsibility is to operate within the space of possibilities formed by economic, workload and safety constraints to navigate towards the desired outcomes we hope to achieve at a given time.
Yet, our capitalist landscape encourages decision-makers to focus on short-term incentives, financial success and survival over long-term criteria such as safety, security and scalability. Workers must be more productive to stay ahead and become “cheaper, faster, better”. Customer expectations accelerate exponentially with each compounding innovation cycle of progress. These pressures push against and migrate teams towards the limits of acceptable (safe) performance. Accidents occur when the system’s activity crosses the boundary into unacceptable safety conditions.
Rasmussen’s model helps us to map and navigate complexity, toward properties for which we wish to optimize. For example, if we want to optimize for Safety, then we need to understand where our safety boundary is in his model for our work. For instance, optimizing for Safety is the primary, explicit outcome of [Chaos Engineering](https://principlesofchaos.org/).
## Spoiler Alerts (Or The Lack Thereof)
The crew at Chernobyl were performing a low power test to understand if residual turbine spin could generate enough electric power to keep the cooling system running as the reactor was shutting down. The standard for the planning of such an experiment should have been detailed and conservative. It was not. It was a “see what happens”, poorly designed experiment—with no criteria established ahead of time for when to abort the experiment. The test had also failed multiple times previously.
Design engineering assistance was not requested, therefore the crew proceeded without safety precautions and without properly coordinating or communicating the procedure with safety personnel. Chernobyl was also the award-winning, top-performance reactor site in the Soviet Union.
The experiment went out of control.
In order to keep the reactor from shutting down completely, the crew shut down several safety systems. Then, when the remaining alarm signaled ignored it for 20 seconds.
While the behavior of the team was questionable, there was also a deeper, unbeknown latent failure in waiting. The reactor at Chernobyl had a unique engineering flaw that caused the reactor to overheat during the test, one which had beed obfuscated from the scientists by the Governments policy to insure State secrets remainder so—the graphite components they selected for the reactor design also believed to offer similar safety standards at cheaper costs. It did not.
Under standard operating conditions reactor No.4’s max power output was 3,200 MWt (megawatt thermal) during the power surge that followed the reactors output spiked to over 320,000 MWt. This caused the reactor housing to rupture, resulting in a massive steam explosion and fire that demolished the reactor building and released large amounts of radiation into the atmosphere.
The first official explanation of the Chernobyl accident was quickly published in August 1986, 3 months after the accident. It effectively placed the blame on the power plant operators noting that the catastrophe was caused by gross violations of operating rules and regulations. The operator error was due to their lack of knowledge of nuclear reactor physics and engineering, as well as lack of experience and training. The hunt for the single root cause and individuals error complete, case closed.
It wasn’t until later and the International Atomic Energy Agency’s 1993 revised analysis that debate around the reactor’s design was called into question.
One reason there is such contradictory viewpoints and debate about the causes of the Chernobyl accident was that the primary data covering the disaster, as registered by the instruments and sensors, were not completely published in the official sources.
Much of the low level information for how the plant was designed was also kept from operators due to secrecy and censorship by the Soviet government.
The four Chernobyl reactors were pressurized water reactors of the Soviet RBMK design, very different from standard commercial designs and employed a unique combination of a graphite moderator and water coolant. This makes the RBMK design very unstable at low power levels, and prone to suddenly increasing energy production to a dangerous level. This behaviour is counter-intuitive, and was unknown to the operating crew.
Additionally, the Chernobyl plant did not have the fortified containment structure common to most nuclear power plants elsewhere in the world. Without this protection, radioactive material escaped into the environment.
## Contributing Factors To Failure To Consider
**KPI’s drive behavior**
- The crew at Chernobyl were required to complete the test to confirm to standard operation rules to ‘be safe’
- They had a narrow focus on what mattered in terms of safety e.g. completing the test versus operating the plant safety
- They did not define boundaries, success and failure criteria for the experiment in advance of performing it
- They didn’t have all the information to set themselves up for success
- They disregard other indicators flagged by the system as anomalies and pushed ahead to complete the experiment with the timeline they were assigned
**Flow of information**
- The quality of your decisions is based on the quality of your information, supported by a good process to make decisions. The operators were following a bad process with missing information
- Often we find that people setting policy are not the people doing the actual work, and this causes breakdowns between work-as-expected and work-as-done, such as the graphite design of the reactor
- Often policy violations are chosen by workers because they are in a double-bind position, and so they choose to optimize for one value (like timeliness, or efficiency) at the expense of another (like verification)
**Values guide behavior** 
- What the company tells you are the behaviors that lead to success, and to be followed
- What behaviors you believe lead to success?
- What would you do when following the rules goes against your values?
**Limited resource put pressure on behavior** 
- When KPIs and behaviors are set in such a way that pressure put behaviors under stress and create unsafe systems
- The team that were fully prepared to run the test ultimately had to be replaced by a night shift crew when the test was delayed, and that night shift had far less preparation
- The chief engineer was thus put under further pressure to complete the test (and avoid further delay) ultimately clouding his judgment around the inherent risk of using an alternate crew
## How To Drift Into Value (Over Failure)
Performance variability may introduce a drift in your situation. However, we can drift to success over failure by creating experiences and social structures for people to safely learn how to handle uncertainty, and navigate towards the desired outcomes we hope to achieve at a given time.
Here’s a set of principles and practices to consider;
- Try to encourage sharing of high quality information as frequently and liberally as possible
- Look into the layers of your organization—where are decisions made?
- How can you move authority to where the information is richest, the context most current, and the employees closest to customers or the situation at hand?
- Proactively review ‘Things That Went Right’ (i.e- positive investigations such as the Thai Soccer Team Diving Rescue) and examine the ‘near misses’ instead of waiting for situations that ‘went wrong’
- Be aware of conflicting KPIs, compare them with the organization’s values and the behaviors they might drive
- Have explicit and communicated boundaries for economic, workload and safety constraints. For example, in Rasmussen’s mode these exist implicitly, whether you acknowledge them or not. Make sure the people doing the work understand where all three boundaries actually are in your context
- Explore how values are documented, sense-checked and evolve over time—just as your organization values shifts over time
- Proactively seek out what process are being used, sidetracked or short cut. Could you enable the shortcuts processes that work safely? Could you turn the expedite process into the process?
- Rule bound versus rule guided cultures give people flexibility. Can you do it better? Why not share it?
- Have fast feedback mechanism in place to tell you when you’re hitting your pre-defined risk and experiment boundaries
## Conclusion
Complex systems have emergent properties, which means explaining accidents by working backwards from the particular part of the system which has failed will never provide a full explanation of what went wrong.
There will rarely be a single source of failure but many tiny acts and decisions along the way that eventually unearth the latent failures in your system. This is why it’s important to remember, that our work doesn’t only progress through time, it progresses through new information, understanding and knowledge.
The questions is what systems do you have in place to safely control your problem domain (not your people)?
Chernobyl while the world’s worst nuclear accident did lead to major changes in safety culture and in industry cooperation, particularly between East and West before the end of the Soviet Union. Former President Gorbachev said that the Chernobyl accident was a more important factor in the fall of the Soviet Union than Perestroika – his program of liberal reform.
What mental models, theories and methods are you using but not driving the outcomes you’re seeking—[and which must be unlearned](https://barryoreilly.com/unlearn-let-go-of-past-success-to-achieve-extraordinary-results-is-officially-available/)?
### References
Special thanks to [Mark Da Silva](https://www.linkedin.com/in/mark-da-silva-20892010/), [Qiu Yi Khut](https://www.linkedin.com/in/qiuyikhut/), [Erinn Collier](https://www.linkedin.com/in/erinncollier/) and [Casey Rosenthal](https://www.linkedin.com/in/caseyrosenthal/) for their thoughtful reviews
This post was also inspired by The Lund University Learning Lab on Resilience Engineering I attended at Slack, lead by  [David Woods](https://www.linkedin.com/in/davidwoods3/), [John Allspaw,](https://www.linkedin.com/in/jallspaw/) [Laura Maguire,](https://www.linkedin.com/in/lauramaguire/) [Nora Jones](https://www.linkedin.com/in/norajones1) and [Richard I. Cook](https://www.linkedin.com/in/richardcookmd/).
![Resilience Engineering](https://barryoreilly.com/wp-content/uploads/2019/07/Image-from-iOS-e1563853766642.jpg)
Ideas worth exploring (curated by [Lorin Hochstein)](https://twitter.com/lhochstein):
- [The adaptive universe](https://github.com/lorin/resilience-engineering#the-adaptive-universe) (David Woods)
- [Dynamic safety model](https://github.com/lorin/resilience-engineering#dynamic-safety-model) (Jens Rasmussen)
- [Safety-II](https://github.com/lorin/resilience-engineering#safety-i-vs-safety-ii) (Erik Hollnagel)
- [Graceful extensibility](https://github.com/lorin/resilience-engineering#graceful-extensibility) (David Woods)
- [ETTO: Efficiency-tradeoff principle](https://github.com/lorin/resilience-engineering#etto-principle) (Erik Hollnagel)
- [Drift into failure](https://github.com/lorin/resilience-engineering#drift-into-failure) (Sidney Dekker)
- Robust yet fragile (John C. Doyle)
- [STAMP: Systems-Theoretic Accident Model & Process](https://github.com/lorin/resilience-engineering#stamp) (Nancy Leveson)
- Polycentric governance (Elinor Ostrom)

View File

@@ -0,0 +1,81 @@
# Disasterpiece Theater: Slack’s process for approachable Chaos Engineering
- **期号**: SRE Weekly Issue #179(2019-08-04)
- **作者**: Richard Crowley — Slack
- **链接**: https://slack.engineering/disasterpiece-theater-slacks-process-for-approachable-chaos-engineering-3434422afb54
## 简介
Slack’s Disasterpiece Theater isn’t quite chaos engineering, but it’s arguably better in some ways. They carefully craft scenarios to test their system’s resiliency, verifying (or disproving!) their hypothesis that a given disruption will be handled by the system without an incident. They share three riveting stories of lessons learned from past exercises.
> The process each Disasterpiece Theater exercise follows is designed to maximize learning while minimizing risk of a production incident.
## 正文
Slack is a large and complex piece of software that’s been added to and changed many times over the last five years. We added features, grew to 10,000,000 DAUs, and made major architectural changes. We made assumptions and tested them with processes that often resembled science.
Whenever we launch features or make changes, we test the fault tolerance of that new code. Unfortunately, we seldom get to repeat these tests as the environment continues to change around that no-longer-new code. As the sands shift, those initial test results lose value. We remain confident in the resilience and robustness of our most critical systems but that confidence is less well-founded as time progresses. And luck is not an availability strategy, so something must be done.
If we were starting from scratch, we’d probably be practicing Chaos Engineering. After all, [“the best way to test the failure path is never to shut the service down normally.”](https://www.usenix.org/legacy/event/lisa07/tech/full_papers/hamilton/hamilton_html/index.html) But we’re not starting from scratch — we operate a large-scale, business-critical service. So what do we need, right now? We need to make Slack as reliable as possible. We need our development environment to be a more confidence-inspiring place to test for fault tolerance, and we believe that testing the fault tolerance of *all* our systems — not just new systems — will help us meet these needs. We don’t want to cause user-impacting incidents, so whatever we do needs to be safe as well. We also don’t need false confidence, so whatever we do needs to be in production.
In January of 2018, we started a rigorous process of identifying failures that are likely to happen and that we must be able to tolerate, and then purposely causing them to happen in production. This isn’t (yet) Chaos Engineering as practiced and evangelized by Netflix. It’s the first step; we call it Disasterpiece Theater.
## Preparing for an exercise
The process each Disasterpiece Theater exercise follows is designed to maximize learning while minimizing risk of a production incident. Each exercise takes place at a well-publicized time and place with all of the relevant experts in the same room or on the same video conference — we’re not (yet) trying to test our monitoring during these exercises. Before the exercise, one or two hosts write a detailed plan and share it widely. The plan is critical to the safety of the exercise but the plan on its own doesn’t teach us much about our fault tolerance.
The hosts are responsible for doing a “tabletop” exercise in which they think through the entire operation. They document precisely how they’re going to incite the failure, right down to the commands they’re going to run and how they’re going to select which EC2 instances are involved (we’ve taken to calling them “tributes”). We ask the hosts to go on the record for how confident they are that fault tolerance in the dev environment predicts fault tolerance in the prod environment for this exercise. They also document all the logs, metrics, and alerts that should be monitored, as well as runbooks that may be necessary during this exercise. Most importantly, they make a specific hypothesis explaining how the failure will be experienced by upstream and downstream systems and by Slack clients. An example of this might be, “Termination of a MySQL master will result in 20 seconds of increased latency for requests that depend on that database but no increase in latency for other requests and less than 1,000 failed API requests, all of which are retried by clients.”
## Disaster strikes: the exercise in motion
We start each exercise by reviewing the plan and projecting/sharing dashboards in Grafana and searches in Kibana. Now we’re ready to incite failure.
We announce the exercise in our **#ops** channel where more than 700 people hang out. We don’t stop deploys or any other normal activities during the exercise but we do make those folks aware of our plans. We broadcast a few coarse status updates in **#ops** throughout the exercise and keep our play-by-play in **#disasterpiece-theater**.
Exercises always begin by inciting the failure in dev. Then we inspect logs and metrics to confirm the failure is visible in all the ways we expect it to be and not visible in others. It’s a common instinct to want to *go fix something* but we control ourselves and watch the system take care of itself. We look for load balancers and other traffic management to route around the failure or for capacity to be replaced. Occasionally we have to follow runbooks to restore service.
Once the failure has been dealt with in the development environment, we pause to make a go or no-go decision about proceeding to production. The exercise isn’t considered a failure or a waste of time if we don’t proceed to production; in fact, some of our most valuable lessons have come from the development environment. We seriously consider aborting if automated remediations didn’t work, didn’t work *perfectly*, took too long or, most importantly, if the failure would result in more disruption than a short and minor increase in latency for customers. If we’re aborting, we announce the abort in **#ops**.
![](https://d34u8crftukxnk.cloudfront.net/slackpress/prod/sites/7/0_clR6h7WJovYPH8XC.png)
Hopefully, though, we’re encouraged by the results in development and are ready to incite failure in the production environment. We project/share the production dashboards in Grafana and searches in Kibana. We announce in **#ops** that we’re moving on to production.
![](https://d34u8crftukxnk.cloudfront.net/slackpress/prod/sites/7/0_q7q9Sgn77BYhwFz9.png)
Finally, the moment of truth arrives. We incite failure in production. Just like we did in development, we inspect logs and metrics, looking to confirm our hypothesis. We give automated remediation time to do its work. Usually, this moment that’s theoretically terrifying is actually quite calm. When we’re finished, we announce the all-clear in **#ops**.
Then, we debrief: What was the time to detect and time to resolve? Did any users notice? Did any humans have to intervene? What was terrifying? Was any of our documentation wrong? Were any dashboards in Grafana out of date?
## Our results to date
We’ve run dozens of Disasterpiece Theater exercises at Slack. The majority of them have gone roughly according to plan, expanding our confidence in existing systems and proving the correct functioning of new ones. Some, however, have identified serious vulnerabilities to the availability or correctness of Slack and given us the opportunity to fix them before impacting customers. Here are summaries of three particularly successful exercises:
### Avoid cache inconsistency
The first time Disasterpiece Theater turned its attention to memcached it was to demonstrate in production that automatic instance replacement worked properly. The exercise was simple, opting to disconnect a memcached instance from the network to observe a spare take its place. Next, we restored its network connectivity and terminated the replacement instance.
During our review of the plan we recognized a vulnerability in the instance replacement algorithm and soon confirmed its existence in the development environment. As it was originally implemented, if an instance loses its lease on a range of cache keys and then gets that same lease back, it does not flush its cache entries. However, in this case, another instance had served that range of cache keys in the interim, meaning the data in the original instance had become stale and possibly incorrect.
We addressed this in the exercise by manually flushing the cache at the appropriate moment and then, immediately after the exercise, changed the algorithm and tested it again. Without this result, we may have lived unknowingly with a small risk of cache corruption for quite a while.
### Try, try again (for safety)
In early 2019 we planned a series of ten exercises to demonstrate Slack’s tolerance of zonal failures and network partitions in AWS. One of these exercises concerned Channel Server, a system responsible for broadcasting newly sent messages and metadata to all connected Slack client WebSockets. The goal was simply to partition 25% of the Channel Servers from the network to observe that the failures were detected and the instances were replaced by spares.
The first attempt to create this network partition failed to fully account for the overlay network that provides transparent transit encryption. In effect, we isolated each Channel Server *far more* than anticipated, creating a situation closer to disconnecting them from the network than a network partition. We stopped early to regroup and get the network partition just right.
The second attempt showed promise but was also ended before reaching production. This exercise did offer a positive result, though: It showed Consul was quite adept at routing around network partitions. This inspired confidence but doomed this exercise as we ended up doing a lot of work to not even cause any Channel Servers to fail.
The third and final attempt finally brought along a complete arsenal of iptables(8) rules and succeeded in partitioning 25% of the Channel Servers from the network. Consul detected the failures quickly and replacements were thrown into action. Most importantly, the load this massive automated reconfiguration brought on the Slack API was well within that system’s capacity. At the end of a long road, it was positive results all around!
### Impossibility result
There have also been negative results. Incident response often involves making configuration changes using an internally developed system called Confabulator. During one particularly bad incident, Confabulator didn’t operate as expected and we had to make and deploy the configuration change manually. I thought this was worthy of further investigation. The maintainers and I planned an exercise to directly mimic the situation we encountered. Confabulator would be partitioned from the Slack service but otherwise left completely intact. Then we would try to make a no-op configuration change.
We reproduced the error without any trouble and started tracing through our code. It didn’t take long to find the problem. The system’s authors anticipated the situation in which Slack itself was down and thus was unable to validate the proposed configuration change; they offered an emergency mode that skipped that validation. However, both normal and emergency modes attempted to post a notice of the configuration change to a Slack channel. There was no timeout on this action but there was a timeout on the overall configuration API action. As a result, even in emergency mode, the request could never make it as far as making the requested configuration change if Slack itself was down. Since then, we’ve made many improvements to code and configuration deploy and have audited timeout and retry policies in these critical systems.
## Into the future
Disasterpiece Theater has made the regular, safe testing of the fault tolerance of Slack’s most critical systems approachable and non-terrifying. It helps us understand and improve Slack’s basic reliability, one of the most important factors in earning and keeping our customers’ trust, even as we expand and evolve the product.
Exercises like the three highlighted above helped us to improve Slack’s reliability and built (or corrected) our confidence in our systems’ fault tolerance. Our Resilience Engineering team continues to expand and evolve this process all the time and, of course, is planning to run many more Disasterpiece Theater exercises. If you find this interesting and want to be a part of the next exercise, [come join us](https://slack.com/careers/)!

View File

@@ -0,0 +1,13 @@
# Resilience Engineering, Cognitive Systems Engineering, and Human Factors Concepts in Software Contexts
- **期号**: SRE Weekly Issue #179(2019-08-04)
- **作者**: —
- **链接**: https://www.youtube.com/playlist?list=PLb1aZTnPf3-OMChMkrr6WsokRI6LOnuem
## 简介
The above is the title of this YouTube playlist curated by John Allspaw.
## 正文
[About](https://www.youtube.com/about/)[Press](https://www.youtube.com/about/press/)[Copyright](https://www.youtube.com/about/copyright/)[Contact us](https://www.youtube.com/t/contact_us/)[Creators](https://www.youtube.com/creators/)[Advertise](https://www.youtube.com/ads/)[Developers](https://developers.google.com/youtube)[Terms](https://www.youtube.com/t/terms)[Privacy](https://www.youtube.com/t/privacy)[Policy & Safety](https://www.youtube.com/about/policies/)[How YouTube works](https://www.youtube.com/howyoutubeworks?utm_campaign=ytgen&utm_source=ythp&utm_medium=LeftNav&utm_content=txt&u=https%3A%2F%2Fwww.youtube.com%2Fhowyoutubeworks%3Futm_source%3Dythp%26utm_medium%3DLeftNav%26utm_campaign%3Dytgen)[Test new features](https://www.youtube.com/new)[NFL Sunday Ticket](https://tv.youtube.com/learn/nflsundayticket)Resilience Engineering, Cognitive Systems Engineering, and Human Factors Concepts in Software Contexts - YouTube

View File

@@ -0,0 +1,113 @@
# “It’s dead, Jim”: How we write an incident postmortem
- **期号**: SRE Weekly Issue #179(2019-08-04)
- **作者**: Fran Garcia — HostedGraphite
- **链接**: https://www.hostedgraphite.com/blog/its-dead-jim-how-we-write-an-incident-postmortem
## 简介
My favorite sentence:
> If you think an incident is “too common” to get its own postmortem that’s a good indicator that there’s a deeper issue that we need to address, and an excellent opportunity to apply our postmortem process to it.
## 正文
No matter how hard we try to offer an uninterrupted service, outages are inevitable. Fixing the underlying issue and [notifying customers](https://www.hostedgraphite.com/blog/how-to-write-a-status-page-update) is crucial, but it shouldn’t end there. There must be a process in place to learn from what happened and make sure it doesn’t happen again.
In the penultimate part of our SRE process series, we look at how to write an incident postmortem–what it is, why it’s important, who should write it, and considerations to keep in mind before putting pen to paper.
## What is a postmortem?
A postmortem is the written record of an incident, including its impact, actions taken to mitigate it, and lessons learned from it (including followup tasks). While the focus of our incident management is usually in mitigating a currently ongoing incident, the goal of the postmortem is to look forward, and try to make sure we have learned as much as we can from a given incident, so we can be in a better position than we were before to avoid similar issues from reoccurring in the future. If we don't do this we'll be fighting fires every day, and that's no fun.
In other words, a postmortem is the process by which we learn from failure, and a way to document and communicate those lessons. The more we fail, the more learning opportunities we have.
## Why are postmortems important?
There are several reasons postmortems are an incredibly important tool:
- It allows us to document the incident, ensuring that it won't be forgotten. A well-documented incident is invaluable because it includes not only a description of what happened but of what actions we took and the things we believed to be true at the time, which can help inform our actions during future incidents.
- They are the most effective mechanism we can use to drive improvement in our infrastructure. Nothing like seeing our services and processes fail in new and interesting ways to realise what areas need improvement.
- It helps shift the focus from the immediate **now** ("we need to mitigate the impact from this incident now") to the future ("what can we do to improve our systems so this incident doesn't reoccur?").
- When they're posted publicly, it lets our users know that we take every outage seriously, and that we're doing all we can to learn from them and prevent any future disruptions to the service we provide.
## What's the goal of a postmortem?
The number one goal of a postmortem is to learn things from it. It's not a great sign when you sit down to write a postmortem and you already know everything you're going to say. If we're not learning anything, we aren't digging deep enough.
The final postmortem document is just a (small) part of our postmortem process, and its value lies in sharing (both with the rest of the team and the outside world) the important lessons we have learned, so the goal of our postmortem process is not just to produce a document. This document is merely the conduit by which we share what we learned on our journey of discovery. In a way, you could say that the real postmortem was the friends we made along the way.
## Why do we share our postmortems?
We believe in being open with our customers, and we take this very seriously with our [customer communication during incidents](https://www.hostedgraphite.com/blog/how-to-write-a-status-page-update), so publishing our lessons learned after an incident is just an extension of this. Our customers deserve to know why their service wasn't working the way they expect it to work and that when we tell them we'll do better in the future we're not just saying it, and we have actual steps we'll take to ensure that's the case.
## Postmortem process
So we just had an incident. It probably required posting something on our [status page](https://status.hostedgraphite.com/), and now status page is giving you the option to write and publish a postmortem for this incident. This is where the fun begins.
## Remember, a postmortem is not just a document
We already said that the goal of a postmortem process is not just to produce a document, but to learn from failure as much as we can. This means that part of this process is going to involve asking some hard questions to try to extract as much [learning juice](https://www.youtube.com/watch?v=quTn9pL39Cw) as we can from failure. This means that if we feel we don't have much to say on a given postmortem it could very well be because we haven't dug too deeply into this particular incident and everything surrounding it.
For example, we shouldn't be satisfied with identifying what triggered an incident (after all, [there is no root cause](https://www.kitchensoap.com/2012/02/10/each-necessary-but-only-jointly-sufficient/)), but should use the opportunity to investigate all the contributing factors that made it possible, and/or how our automation might have been able to prevent this from ever happening. The lessons we learn from an incident only stop coming when we stop digging, so an incident with no lessons learned only means we didn't look hard enough.
## When should I write a postmortem?
Postmortems are such a good learning opportunity that we should take every chance we get to write one, but the decision on writing one or not usually falls on the incident commander (normally the on-call engineer at the time). If we aren't sure if a given incident "deserves" a postmortem, **it's never a bad choice to write one anyway**, it's best to err on the side of oversharing than to give the impression that we don't care enough about communicating about our incidents, and we should always be happy to have another learning opportunity.
If you think an incident is "too common" to get its own postmortem that's a good indicator that there's a deeper issue that we need to address, and an excellent opportunity to apply our postmortem process to it. Sometimes a single instance of an incident can't give you enough information to get any meaningful lessons out of it, but when looking at a group of seemingly related incidents as an aggregate they might start to paint a clearer picture.
If we know that we'll want to write a postmortem before officially resolving the incident on our status page, it's always a good idea to tell our customers to expect a postmortem. A postmortem doesn't need to go out on the same day the incident happened, and there's certainly no expectation of staying until late or over the weekend writing one. Having a postmortem ready on the next business day after an incident is a good goal, but in some cases (such as particularly complex incidents, or times where we're still very busy dealing with the fallout) this could be delayed a bit more. Ideally, it should never take more than a week after the incident is resolved for the postmortem to be published.
It's also worth noting that not all postmortems need to be published on our status page or be tied to an actual incident. Sometimes we'll want to write a postmortem around near-incidents or incidents that didn't have enough of a visible impact to warrant updating our customers. A postmortem doesn't need to be published externally to be useful.
## Who should be writing this postmortem?
It's usually up to the on-call engineer to write a postmortem for any incidents that happened during [their watch](https://pics.me.me/when-your-shift-is-over-and-your-manager-asks-you-4760512.png), but as with many other things regarding on-call, this too can be delegated. The on-call engineer is still responsible for ensuring that we produced a postmortem and that it's shared both internally and publicly, but they don't necessarily need to write it themselves.
That said, just because one person is leading this process doesn't mean a postmortem is a one-person job. A postmortem is a team effort and you'll want input from everybody that was involved in the incident (and others that weren't). We all have a different perspective and a different mental model of what our systems look like, so only by combining them all, you'll get closer to the full picture of what really happened during an incident.
## Is this going to be a finger-pointing exercise?
[No](https://www.youtube.com/watch?v=GGn25URIss8). We could fill pages talking about blameless postmortems and how important we are, but the main takeaway is that we all make mistakes, and we're not here to point at those and say "our problem is that someone made a mistake, we'll try making zero mistakes next time". What we want is to learn why our processes allowed for that mistake to happen, to understand if the person that made a mistake was operating under wrong assumptions (and how our people can have the necessary information to make better decisions) or even why they were doing what they were doing in the first place (instead of that process being fully automated).
Nobody gets blamed when something goes wrong, but the more we share about these experiences, the more we'll learn about them. Despite not assigning blame, we can (and should!) explicitly identify times where a mistake was made.
It's important to call out mistakes, but our focus should be on the mistake itself and what can be learned from it, as opposed to the person making the mistake. That person becomes our leading expert on that particular mistake, so we'll want to learn from everything they have to teach us.
## So where do I start?
The first (and most important) steps of the postmortem process don't require us to write a single word. Before we can put everything we've learned in a document we have to truly understand the incident.
1. Compile a timeline of the incident. It's really useful to see what actions we took and when we took them, and the things we thought to be true at any point during the incident.
2. Ask yourself (and others) **a lot of questions** . We know[there is no (single) root cause](https://www.kitchensoap.com/2012/02/10/each-necessary-but-only-jointly-sufficient/) , and that the story of an incident is composed of[infinite hows](https://www.oreilly.com/ideas/the-infinite-hows) , which means that a postmortem will only be useful if we continue digging and challenging any assumptions we have about or systems. Some examples of useful questions would include (but are certainly not limited to):
- How did this failure go unnoticed for XX minutes? Do we not have alerts that cover this failure scenario? Did they work as expected?
- Even in cases when we still don't know why something happened and remains a mystery, what kind of instrumentation/diagnostics do we think we'd need to be able to identify it the next time it happens?
- Did we accurately assess the impact originally? If we didn't, how can we make sure we do it better the next time?
- Could the incident have been worse but maybe we got lucky somehow? What could happen if the next time we don't have that kind of luck?
- Did we get unlucky and an incident that shouldn't have been a major issue somehow became one? Then we need to dig into what were the contributing factors to that, since "have better luck next time" is not the best strategy.
- Was the incident caused or made worse by something we did? What led us to believe that was the right course of action? Could our systems/tooling have prevented us from taking that action or mitigate its impact? Remember, our postmortems are blameless so this is not a finger-pointing exercise, but we need to be able to identify these instances so we can look at all the contributing factors.
- Did we, at some point, make the wrong call? Did we have invalid/incomplete information at the time? Maybe our documentation was the issue?
- What kind of information would we need to do better next time?
- What was each of us thinking during the incident? How did we feel? Did we feel we had the right information/context at all times? The people involved in the incident are also a part of the system we're trying to learn about, and as such it's important not to overlook them.
While working through the timeline and asking questions you should be making a note of everything that makes you think "hmmm maybe this could have gone better if only we had X", as those will end up becoming our follow-up actions for this postmortem.
Next up, we’ll take a deeper look at the structure of a postmortem–section by section–with a helpful template, writing tips, and some pointers to keep in mind.
## Related reading
There are too many great resources out there to list, but the following should be considered required reading (or watching!) on the topic:
["The infinite hows" - John Allspaw](https://www.oreilly.com/ideas/the-infinite-hows)
[Incidents as we Imagine Them Versus How They Actually Are - John Allspaw (video)](https://www.youtube.com/watch?v=8DtzmV1jiyQ)
[How complex systems fail - Richard Cook](https://web.mit.edu/2.75/resources/random/How%20Complex%20Systems%20Fail.pdf)
[The Multiple Audiences and Purposes of Post-Incident Reviews](https://www.adaptivecapacitylabs.com/blog/2018/10/08/the-multiple-audiences-and-purposes-of-post-incident-reviews/)
[Some Observations On the Messy Realities of Incident Reviews](https://www.adaptivecapacitylabs.com/blog/2019/06/17/some-observations-on-the-messy-realities-of-incident-reviews/)
[Hindsight and sacrifice decisions](https://www.adaptivecapacitylabs.com/blog/2019/03/03/hindsight-and-sacrifice-decisions/)

View File

@@ -0,0 +1,13 @@
# Building a real-time anomaly detection system for time series at Pinterest
- **期号**: SRE Weekly Issue #179(2019-08-04)
- **作者**: —
- **链接**: https://medium.com/@Pinterest_Engineering/building-a-real-time-anomaly-detection-system-for-time-series-at-pinterest-a833e6856ddd
## 简介
> In this post, we’ll share the algorithms and infrastructure that we developed to build a real-time, scalable anomaly detection system for Pinterest’s key operational timeseries metrics. Read on to hear about our learnings, lessons, and plans for the future.
## 正文
> ⚠️ 抓取失败:HTTP 403

View File

@@ -0,0 +1,13 @@
# Ably Debugging Tales Part 1 — An Elixir Erlang Mystery
- **期号**: SRE Weekly Issue #179(2019-08-04)
- **作者**: Eve Harris — Ably
- **链接**: https://medium.com/ably-realtime/ably-debugging-tales-part-1-an-elixir-erlang-mystery-ably-blog-data-in-motion-d279a18a2eb
## 简介
I sure do love a good debugging story.
## 正文
> ⚠️ 抓取失败:HTTP 403

View File

@@ -0,0 +1,93 @@
# Incident investigation: Learning vs Blaming
- **期号**: SRE Weekly Issue #179(2019-08-04)
- **作者**: Phillip Dowland — Safety Differently
- **链接**: http://www.safetydifferently.com/incident-investigation-learning-vs-blaming/
## 简介
When an incident occurs, your company is faced with a choice: do you seek to learn as much as possible about how it happened, or do you seek to find out who messed up?
## 正文
![](http://www.safetydifferently.com/wp-content/uploads/2019/07/blank-conifers-crossroad-1578750-1024x687.jpg)
I would like to tell a story. The story is not in itself unique and I suppose many persons have seen or experienced similar. There are two ways to tell this story with two very different outcomes for both the organisation and individuals involved.
Let me set the scene, a difficult project is under pressure to produce, machinery is not in one hundred percent working order, operatives have been moved to different sites due to the production difficulties. The crew are continually having to correct a mechanical issue located in an area at height. This has not gone unnoticed, in fact two days previous a senior leader had observed the crew struggling with this issue. However, the mechanical problem had not been reported through the organisational reporting system.
So, we can see a site that is struggling and in hindsight perhaps ripe for suffering some kind of incident. In fact, on the day of the story no one was hurt, however the potential was certainly there.
So, the mechanical issue at height had again occurred and to rectify it the operative, newly brought in to try to pull back the production issues, climbed without any fall protection controls.
## Story 1.
The person stood watching the works waiting for someone to make a health and safety violation, when out of the corner of his eye he saw the operative climb without any fall prevention/protection controls in place. In a blink of an eye the photo was taken and shared through the online reporting as a high risk dangerous condition. Due to the high risk categorisation this was cascaded through the organisation to senior management.
Within hours the senior manager was demanding the answer to why the operative had carried this violation out, why had he thought that working at height like this was a good idea? With this level of emotion the investigation process was put into motion, it began with questions such as “what were you thinking of?”; “why did you break the rules” and by the end of the discussion the operative had been blamed and found guilty. Post investigation was almost like the sentencing at a crown court with the operative finding himself in front of the Operational and HR managers, receiving a written warning and being told that the rules are there to keep him safe and he must follow them.
## Story 2.
The person stood watching the works waiting for someone to make a health and safety violation, when out of the corner of his eye he saw the operative climb without any fall prevention/protection controls in place. In a blink of an eye the photo was taken and shared through the online reporting as a high risk dangerous condition. Due to the high risk categorisation this was cascaded through the organisation to senior management.
Upon receiving the information, the senior manager, realising that the company had, in his words, “been given a get out of jail free card,” set in motion a period of essential learning for the organisation. During this time, a learning group was assembled to look at the facts around the incident. The group included not only health and safety representatives, but also operational staff both management and site based and independent personnel from a different division, the one thing in common was that all persons involved in this team new the work and how to do the work.
During the [learning team](https://safetydifferently.com/contemporary-safety-glossary/#learning-team) a real understanding of how the job was a. set up b. performed and c. where the systems in place had failed, was gained. This allowed the business to put system controls in place to help prevent any reoccurrence, thus mitigating against possible injury of staff.
## The Importance of Learning
What are the differences in the two approaches and how do they affect the outcome and the organisations continual improvement? Lets first look at the blame culture, when an individual or work group makes a mistake, in apportioning blame for this error we are, as Todd Conklin states in his book The 5 Principles of Human Performance, making it a choice of the individual/group. This approach does nothing for the improvement of the company and can in fact just be like a sticky plaster over the issue due to not actually identifying the system contributing factors and in turn can lead to a repeat of the incident; (Sidney Dekker; [Just Culture](https://safetydifferently.com/contemporary-safety-glossary/#just-culture); 2016). 
Let’s take for example the true story of a 24 year old graduate medical student, who in his first month was looking after a 16-year old boy who was undergoing palliative chemotherapy. This boy needed to have two different injections, one intravenously and a second by lumbar puncture into the spine. The intravenous drug was highly toxic, in fact it would be fatal to the patient if administered to the spine, it did however arrive on the ward in a nearly identical syringe to the other injection. Both these syringes were handed to the young doctor for the lumbar puncture procedure and both injected into the patient’s spine. Despite the efforts of the medical staff the young boy died a week later. The graduate was blamed and prosecuted for this tragic incident. Fortunately, this conviction was overturned, but because of the sticky plaster approach the real system failure was not learned from and the same mistake was made in a different hospital and another patient unnecessarily died.
So, in just focusing on the operative working at height the system failures; non reporting of mechanical issues; production issues; new personnel and added pressure of the hope that they could turn the production around, were missed/not considered.
What then does the approach to the initial issue from story 2 tell us? In the first instance a learning team was established. Evidence from [RoSPA research in the UK](https://www.rospa.com/occupational-safety/advice/safety-failure/) show that a team based approach to incident learning is extremely powerful and can:
- provide access to local, ‘expert’ knowledge, particularly about operational issues;
- support the building of trust and the development of ‘just’ (open, fair) cultures;
- promote learning about how to investigate in general (i.e. not just H&S failures).
This helps give ownership and shows a level of trust from management that the workers play a huge part in the success of the operations. I wrote in a recent article how lean theory can tie into the new [safety differently](https://safetydifferently.com/contemporary-safety-glossary/#safety-differently) way. In this article it was stated that the importance of realising that the workers are a hub of knowledge that when tapped into through a Kaizen (continuous improvement) method can smooth ongoing operations. This same process in incident investigations is vital. By respecting and trusting the workforce to take part in these learning sessions encourages openness whereas blaming is more than likely to make non reporting of problems a reality. Also utilising this operational knowledge can really lead to robust practices being put in place by identifying where systems are perhaps weak and breaking down during operations.
Imagine then if the emotional response to the death of the 16 year old patient was changed to one of learning from the failure, looking at where the system had failed and led to the error from the young graduate and sharing these lessons with other medical facilities then perhaps another life would have been saved.
In conclusion, we can see that response to incidents cause a level of emotions that can lead to investigations becoming witch hunts rather than true learning sessions for the *Don’t find blame, learn and fix systems!* 
Nicely written! If you can establish and maintain that tone of learning over blame, the investigation will be much more productive. In my experience, sometimes the person who made the error is the one who has the hardest time letting go because, even if no one else on the team does, they blame themselves. It takes a powerful level of introspection and self-forgiveness – particularly if the outcome was egregious.
Can only agree to learnings in story 2, which is what we should strive to achieve. Another point is the absence of the possibility of stopping the work, rather than just reporting a possible hazardous situation. Both stories could have a fatal outcome regardless of the reporting. This is also a part of a healthy just and safety culture: The ability, trust and authority for anyone to stop a job, in case of a possible hazardous situation. The reporting person in both stories could have taken the picture and then stopped the job, and the doctor could have stopped if unsure of the contents of the syringes. I believe that a company, which has such Stop Work Authority and accepts it to the fullest, has the highest potential for maintaining a just safety culture.
Incident investigations may find that human error is to blame. However human error is often the result of systemic failure and should be the starting point of an investigation and not the end.
We need to ask how or why a person or group behaved as they did by first examining the system, environment or other factors that may have contributed to the behavior.
This once again presumes that Dr’s. Conklin and Dekker’s characterization of an RCA/Investigation is accurate…according to them all RCA is 1) linear, 2) component-based and 3) results in a single cause. Unfortunately, that is a mischaracterization as field-proven, effective RCA approaches are conducted by veteran analysts on a daily basis.
Such analyses do incorporate and REQUIRE this learning perspective. They seek to understand why good people, make poor decisions at the time they did. They seek to understand the flawed organizational systems that influenced the decision-maker. They seek to understand the external socio-technical systems that influence internal management systems.
What they also require is evidence to back up what is concluded as opposed to allowing hearsay to fly as fact.
Hybrids of the two approaches exist and are being applied effectively, however that ‘fact’ seems to be ignored. Perhaps we should amass a Learning Team to determine why that is?
Do Human Performance Teams Make RCA Obsolete?
https://reliability.com/pdf/rca-vs-hpi-2017-rci.pdf
We all have unity in purpose and should work with each other to attain the same ‘ends’.
Maybe there is a Story 3…
The operative was considering engaging in a necessary task without any fall prevention/protection controls in place. Because he trusts his supervisor he has a conversation with her about the falling hazards, pre-task planning, and how to do the job with adequate tools and equipment. They plan the work together and engage in the task safely. The supervisor and the employee share the story with upper management as an example of safe thinking and team work. Upper management recognizes this is a good story to tell throughout the organization.
If you really want cooperation and openness, stop using the work “investigation”. What will get you better results,
“I hear you made a mistake last night and I’m here to investigate the incident.”
or
“I hear you were involved in an event last night and I’m here to learn from your experience and try to prevent a recurrence.”
???
I couldn’t agree more with You Jim. Language is absolutely key. If we use less punitive language we get a lot more richer conversation. Which in turn creates the space to learning rather than finger pointing.