SRE weekly 所有文章

This commit is contained in:
2026-09-12 17:23:01 +08:00
parent 409b40ddcb
commit af7633f9dc
8486 changed files with 4489990 additions and 7 deletions

View File

@@ -0,0 +1,787 @@
<!DOCTYPE html>
<html>
<head>
<meta charset="utf-8" />
<meta http-equiv="X-UA-Compatible" content="IE=edge"><title>Incidents — Trends from the Trenches</title><link rel="icon" type="image/png" href=/favicon-16x16.png /><meta name="viewport" content="width=device-width, initial-scale=1">
<meta name="description" content="Most publicized production incidents are war stories. Each involves drama with dead ends, twists and turns, and a victory at the end…">
<meta itemprop="name" content="Incidents — Trends from the Trenches">
<meta itemprop="description" content="Most publicized production incidents are war stories. Each involves drama with dead ends, twists and turns, and a victory at the end…">
<meta itemprop="datePublished" content="2019-02-26T18:26:22+00:00">
<meta itemprop="dateModified" content="2019-02-26T18:26:22+00:00">
<meta itemprop="wordCount" content="2291">
<meta itemprop="image" content="https://www.subbu.org/subbu.jpg"><meta property="og:url" content="https://www.subbu.org/essays/2019/incidents-trends-from-the-trenches/">
<meta property="og:site_name" content="Writing is clarifying">
<meta property="og:title" content="Incidents — Trends from the Trenches">
<meta property="og:description" content="Most publicized production incidents are war stories. Each involves drama with dead ends, twists and turns, and a victory at the end…">
<meta property="og:locale" content="en">
<meta property="og:type" content="article">
<meta property="og:image" content="https://www.subbu.org/subbu.jpg">
<meta property="article:published_time" content="2019-02-26T18:26:22Z"><meta property="article:modified_time" content="2019-02-26T18:26:22Z"><link href='https://fonts.googleapis.com/css?family=Playfair+Display:700' rel='stylesheet' type='text/css'>
<link rel="stylesheet" type="text/css" media="screen" href="https://www.subbu.org/css/normalize.css" />
<link rel="stylesheet" type="text/css" media="screen" href="https://www.subbu.org/css/main.css" />
<link rel="stylesheet" type="text/css" href="https://www.subbu.org/css/custom.css" />
<link id="dark-scheme" rel="stylesheet" type="text/css" href="https://www.subbu.org/css/dark.css" />
<script src="https://www.subbu.org/js/feather.min.js"></script>
<script src="https://www.subbu.org/js/main.js"></script>
<script>
window.goatcounter = {
path: function(p) { return location.pathname; }
}
</script>
<script data-goatcounter="https://www.subbu.org/gc/count"
async src="https://www.subbu.org/gc/count.js"></script>
<noscript><img src="https://www.subbu.org/gc/count?p=%2fessays%2f2019%2fincidents-trends-from-the-trenches%2f"></noscript>
</head>
<body>
<div class="container wrapper">
<div class="header">
<div class="avatar">
<a href="https://www.subbu.org/">
<img src="/subbu.jpg" alt="Writing is clarifying" />
</a>
</div>
<h1 class="site-title"><a href="https://www.subbu.org/">Writing is clarifying</a></h1>
<div class="site-description"><p>Writing on leadership and technology</p><nav class="nav social">
<ul class="flat"><li><a href="https://www.linkedin.com/in/subbu/" title="LinkedIn"><svg xmlns="http://www.w3.org/2000/svg" width="24" height="24" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M16 8a6 6 0 0 1 6 6v7h-4v-7a2 2 0 0 0-2-2 2 2 0 0 0-2 2v7h-4v-7a6 6 0 0 1 6-6z"></path><rect x="2" y="9" width="4" height="12"></rect><circle cx="4" cy="4" r="2"></circle></svg></a></li><li><a href="https://bsky.app/profile/sallamar.bsky.social" title="Bluesky"><svg xmlns="http://www.w3.org/2000/svg" width="24" height="24" viewBox="0 0 568 501" fill="currentColor"><path d="M123.121 33.664C188.241 82.553 258.281 181.68 284 234.873c25.719-53.192 95.759-152.32 160.879-201.21C491.866-1.611 568-28.906 568 57.947c0 17.346-9.945 145.713-15.778 166.555-20.275 72.453-94.155 90.933-159.875 79.748C507.222 323.8 536.444 388.56 473.333 453.32c-119.86 122.992-172.272-30.859-185.702-70.281-2.462-7.227-3.614-10.608-3.631-7.733-.017-2.875-1.169.506-3.631 7.733-13.43 39.422-65.842 193.273-185.702 70.281-63.111-64.76-33.89-129.52 80.986-149.071-65.72 11.185-139.6-7.295-159.875-79.748C10.945 203.659 1 75.291 1 57.946 1-28.906 76.135-1.612 123.121 33.664z"></path></svg></a></li><li><a href="/index.xml" title="RSS"><i data-feather="rss"></i></a></li><li><a href="#" class="scheme-toggle" id="scheme-toggle"></a></li></ul>
</nav>
</div>
<nav class="nav">
<ul class="flat">
<li>
<a href="/">Home</a>
</li>
<li>
<a href="/coaching/">Coaching</a>
</li>
<li>
<a href="/essays">Writing</a>
</li>
<li>
<a href="/series/mountain">The Hold</a>
</li>
<li>
<a href="/subscribe">Subscribe</a>
</li>
<li>
<a href="/about">About</a>
</li>
</ul>
</nav>
</div>
<div class="post article">
<div class="post-header">
<div class="matter">
<h1 class="title">Incidents — Trends from the Trenches</h1>
<p class="article-date"> Tuesday, February 26, 2019 </p>
</div>
</div>
<div class="series-note home-coaching-note">
<p><strong>New:</strong> <a href="/coaching/">one-on-one coaching</a> for tech professionals working through hard leadership problems—whether you manage a team or the whitespace between them.</p>
</div>
<div class="markdown">
<p>Most publicized production incidents are war stories. Each involves drama with dead ends, twists and turns, and a victory at the end. Something innocuous happens, that then snowballs across several layers to take down some parts of a business. A big chunk of internal or external customers gets impacted. Several teams spend long hours on a conference call or in the war room to mitigate the customer impact.</p>
<p>You may recall well-publicized incidents like the <a href="https://aws.amazon.com/message/41926/">AWS S3 outage</a> in 2017 that impacted several AWS customers, including Apple iCloud, or the <a href="https://en.m.wikipedia.org/wiki/2016_Dyn_cyberattack">cyber attack</a> on <a href="https://dyn.com/dns/">Dyn DNS</a> that affected several American and European sites, or last year’s <a href="https://www.amazon.com/">Amazon.com</a>’s <a href="https://www.cnbc.com/2018/07/19/amazon-internal-documents-what-caused-prime-day-crash-company-scramble.html">Prime Day</a> outage.</p>
<p>Such incidents are rare, and yet they remain in our memories for years. In reality, most production environments encounter incidents almost every day. As you see below, <strong>the cumulative cost and customer impact of such incidents can be much larger than the infrequent dramatic ones.</strong></p>
<p>During the fall of 2018, I set out to develop informed opinions on how to improve the availability of production systems at work. There is no dearth of architecture patterns, tools, techniques and processes available to improve availability. How do you determine which ones to focus on and when, and make continuous improvements? That was the question I was grappling with. More important, I also needed a way to challenge some of my own prior opinions.</p>
<h2 id="incident-analysis">Incident analysis<a class="heading-anchor" href="#incident-analysis" aria-label="Link to this section"><svg xmlns="http://www.w3.org/2000/svg" width="0.7em" height="0.7em" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2.8" stroke-linecap="round" stroke-linejoin="round"><path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/><path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/></svg></a></h2>
<p>In order for this, I could think of no better way than to study incidents to spot patterns. During November and December of 2018, I spent several weeks to study several hundred production incidents. This sample set covered a very large set of customer-facing apps and services running on-prem and cloud, including some that are yet to be modernized, as well as the on-prem infrastructure. I meticulously went through each critical incident, read incident logs, and where available, reviewed postmortem reports and classified incidents based on a few categories of potential <em>triggers</em>.</p>
<p><em>A clarification on the terminology here.</em> I’m using the term “trigger” and not “root cause” to classify incidents. This is to emphasize the fact that most production incidents have several root causes. A trigger may just have surfaced an incident.</p>
<p>This analysis was time-consuming and laborious. Yet, the insights I gathered were well worth the time I spent. In this article, I want to share my findings, offer some hypotheses to explain the findings, and what could be done to improve availability.</p>
<p>The chart below summarizes my findings. It shows the top 5 triggers behind these incidents, ordered by the cumulative customer impact.</p>
<p><img src="/img/1__I6igQWlZj4me8MLStVlbmw.png" alt=""></p>
<p>The size of each slice represents the customer impact as measured by certain metrics, and not the number of incidents.</p>
<p>Contrast this chart to the one below, which shows the incidents by number under each category.</p>
<p><img src="/img/1__OfcU7R9KvXP2jiPHp6whLA.png" alt=""></p>
<p>I omitted some categories in these charts due to those not being relevant for this article. A similar analysis of a different sample set of incidents might produce a different set of triggers, though I suspect that the above shows common trends across most large enterprises undergoing constant change. Few peers in the industry also conferred that they notice similar patterns.</p>
<h3 id="observation-1-change-is-the-most-commontrigger">Observation 1: Change is the most common trigger<a class="heading-anchor" href="#observation-1-change-is-the-most-commontrigger" aria-label="Link to this section"><svg xmlns="http://www.w3.org/2000/svg" width="0.7em" height="0.7em" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2.8" stroke-linecap="round" stroke-linejoin="round"><path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/><path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/></svg></a></h3>
<p>About a third of the impact was triggered by changes. Of this, about 50% was due to software deployments. In my classification, a change could be any of the following:</p>
<ul>
<li>Automated CI/CD releases</li>
<li>Semi-automated deployments legacy apps</li>
<li>Manual changes</li>
<li>Configuration changes, such as traffic routing, or ingress/egress filters</li>
<li>Experiments (A/B tests)</li>
</ul>
<p>As I showed in <a href="https://m.subbu.org/taming-the-rate-of-change-439e3dccbb5d">Taming the Rate of Change</a>, given that the production environment at work undergoes a few thousand changes every working day, the change failure rate is still low. The impact, nonetheless, is significant.</p>
<p>This observation supports the anecdotal evidence for a low number of incidents during long weekends and holidays when production changes are low. Just last week, a colleague of mine quipped that production systems were mostly stable during the recent <a href="https://www.google.com/search?q=seattle+snowmageddon+2019&amp;source=lnms&amp;tbm=isch&amp;sa=X&amp;ved=0ahUKEwjwkcqKucngAhWX_YMKHX2wBBYQ_AUIDigB&amp;biw=958&amp;bih=1089">Seattle Snowmageddon 2019</a> because most people could not get to work. Some areas also lost power and Internet access during that time.</p>
<p>A couple of months before this analysis that produced the above pie charts, I analyzed a smaller sample of just over 100 critical incidents that covered a particular set of business functions. For each incident, I asked a simple question — was there a change that preceded the incident. I grouped all incidents with a “yes” into one bucket, and everything else into another bucket. The result is below.</p>
<p><img src="/img/0__LEX7qUs7lSVxU12e.jpg" alt=""></p>
<p>The result was surprising and extremely alarming. Over two-thirds of the sample of incidents was triggered by one or more changes. This finding led me to the latter analysis of the larger sample of incidents. Change is still at the top, by customer impact.</p>
<p>There is prior research to support this observation. An 2016 ACM paper titled <a href="https://dl.acm.org/citation.cfm?id=2934891">Evolve or Die: High-Availability Design Principles Drawn from Google’s Network Infrastructure</a> makes the following observation based on a detailed analysis of over 100 high-impact network failure events:</p>
<blockquote>
<p>a large number of failures happen when a network management operation is in progress within the network.</p>
</blockquote>
<h3 id="observation-2-config-drift-accumulates-over-time-and-masks-potential-future-incidents">Observation 2: Config drift accumulates over time and masks potential future incidents<a class="heading-anchor" href="#observation-2-config-drift-accumulates-over-time-and-masks-potential-future-incidents" aria-label="Link to this section"><svg xmlns="http://www.w3.org/2000/svg" width="0.7em" height="0.7em" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2.8" stroke-linecap="round" stroke-linejoin="round"><path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/><path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/></svg></a></h3>
<p>The second trigger from the top is config drift, which contributed to about one-fifth of the impact.</p>
<p>For those not familiar with config drift, consider a cluster of nodes each of which is expected to maintain a certain configuration. The configuration may include the OS, OS level or application level dependencies, security groups and such access controls, config files etc.</p>
<p>The cluster could be a SQL database in an active-passive configuration, a Zookeeper cluster, or pair of network switches. In order for the cluster to stay healthy in case of failures of any one node, each is expected to be in a certain configuration. Now, say, due to someone manually making changes, or an automation defect, one of the nodes does not have the expected configuration. This is config drift.</p>
<p>It is fairly common for config drift to stay dormant for weeks or months and surface only when some other event happens. In one particular incident, one of the network switches configured in a pair drifted from its configuration. Months later, the other switch failed for some of the reason, and the drifted switch could not take over. This lead to network disruption. I’ve witnessed similar incidents in the past with other types of clusters, and have stories to tell.</p>
<h3 id="observation-3-we-dont-always-know-why-systemsfail">Observation 3: We don’t always know why systems fail<a class="heading-anchor" href="#observation-3-we-dont-always-know-why-systemsfail" aria-label="Link to this section"><svg xmlns="http://www.w3.org/2000/svg" width="0.7em" height="0.7em" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2.8" stroke-linecap="round" stroke-linejoin="round"><path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/><path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/></svg></a></h3>
<p>The next biggest in my finding was a large number of incidents that recovered on their own after a while. Though this category was the third by customer impact (per the first pie chart in this article) on my list, it accounted for over 40% of the incidents (per the second pie chart) I examined.</p>
<p>To reiterate, for over 40% of incidents, there was an alert of customer impact, an incident was declared, relevant people got on the incident bridge, and while the investigation was ongoing, the impact mitigated by itself.</p>
<p>Unfortunately, such incidents don’t get the attention of postmortem analysis, and hence corrective actions.</p>
<h3 id="observation-4-infrastructure-issues-are-less-frequent-than-commonlybelieved">Observation 4: Infrastructure issues are less frequent than commonly believed<a class="heading-anchor" href="#observation-4-infrastructure-issues-are-less-frequent-than-commonlybelieved" aria-label="Link to this section"><svg xmlns="http://www.w3.org/2000/svg" width="0.7em" height="0.7em" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2.8" stroke-linecap="round" stroke-linejoin="round"><path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/><path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/></svg></a></h3>
<p>Infrastructure related failures like data center power, disk or other hardware, WAN link etc. are less frequent than most people believe. The same is true for public cloud service or region failures. Such issues accounted for a smaller percentage of customer impact in my analysis.</p>
<p>In some of the incidents I reviewed, while initial investigations pointed to misbehaving infrastructure (such as a particular vendor’s appliance failing), further analysis revealed botched changes (see Observation 1) or config drift (see Observation 3).</p>
<h3 id="observation-5-certificate-related-issues-continue-to-be-aheadache"><strong>Observation 5: Certificate related issues continue to be a headache</strong><a class="heading-anchor" href="#observation-5-certificate-related-issues-continue-to-be-aheadache" aria-label="Link to this section"><svg xmlns="http://www.w3.org/2000/svg" width="0.7em" height="0.7em" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2.8" stroke-linecap="round" stroke-linejoin="round"><path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/><path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/></svg></a></h3>
<p>Finally, the fifth in my list is incidents related to certificate handling. There were just a handful of incidents in this category, and yet the impact was not insignificant. The issues related to forgetting to renew certificates in time or not coordinating the renewal across multiple systems. While these are easily fixable through automation or even processes, such errors continue to happen in complex production environments.</p>
<h2 id="what-is-goingon">What is going on<a class="heading-anchor" href="#what-is-goingon" aria-label="Link to this section"><svg xmlns="http://www.w3.org/2000/svg" width="0.7em" height="0.7em" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2.8" stroke-linecap="round" stroke-linejoin="round"><path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/><path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/></svg></a></h2>
<p>Given the large sample size covering a diverse set of apps, services, and technologies, analysis like this provides an opportunity to better understand contemporary production environments at a high level. Below is my hypotheses of what might be contributing to these trends.</p>
<p><strong>First, we trip on ourselves when making changes.</strong> The biggest risk to the availability of production systems is constant change. Due to the adoption of microservices, and investments into containers, CI/CD, and the cloud, our ability to make changes in production environments has been rapidly increasing. There is no turning back from this trend due to productivity gains. However, change safety is not always an inherent feature in the tools used to make changes.</p>
<p>As I argued in <a href="https://m.subbu.org/taming-the-rate-of-change-439e3dccbb5d">Taming the Rate of Change</a>, these technology trends are contributing to the following:</p>
<ol>
<li><strong>Hyperconnectedness:</strong> Enterprises are increasingly deriving value from connecting various services in numerous ways. In a sense, the value of the enterprise is slowly shifting from nodes (systems doing particular things) to edges (interconnectedness). This is increasing possibilities for both success and failure.</li>
<li><strong>Side effects:</strong> Amidst hundreds or thousands of services, anyone making a change to a particular microservice is unlikely to know all the consumers of that service across multiple layers.</li>
<li><strong>Hope driven releases:</strong> Production environments are often the only reliable environments to test a change. As most enterprises are decentralizing once-common release engineering discipline, pre-production environments are becoming stale, unreliable, and lightly monitored. Consequently testing in production is increasingly becoming vogue.</li>
</ol>
<p><strong>Second, the desire for speed may be stealing focus from automation.</strong> This analysis makes it clear that automation is rarely complete, with less frequently used parts of any workflow getting the least amount of attention.</p>
<p>Furthermore, as we move on from one generation of technology and architecture to the next one, we rarely leave the prior generation in the best possible shape.</p>
<p>Consequently, as systems age, less frequently used parts accumulate config drift. Unlike the other form of bugs, drift tends to remain dormant until some other event occurs before leading to a fault.</p>
<p>This trend is not limited to on-prem services. Apps and services deployed on the cloud are also subject to config drift. Teams adopting new technology usually start with automation to get going quickly, but not necessarily automate manageability tasks that come up in the future. This keeps the door open for drift to creep in.</p>
<p><strong>Finally, the large number of incidents in the unknown category shows that our ability to comprehend the physics of hyperconnected systems is limited.</strong> Furthermore, as systems seem to recover on their own, we’re also losing the opportunity to learn from such incidents.</p>
<h2 id="potential-ways-toimprove">Potential ways to improve<a class="heading-anchor" href="#potential-ways-toimprove" aria-label="Link to this section"><svg xmlns="http://www.w3.org/2000/svg" width="0.7em" height="0.7em" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2.8" stroke-linecap="round" stroke-linejoin="round"><path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71"/><path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71"/></svg></a></h2>
<p>This analysis certainly helped me refine my opinions on areas of investments. I want to highlight a few techniques to help deal with the trends I noticed.</p>
<p><img src="/img/1____WXFrLY1zX7ih3eq9IiQAg.png" alt=""></p>
<p>First, the most important take away from this analysis is improving <strong>change safety</strong>. Progressive deployments (i.e., introducing the change bit by bit), feature flags, <a href="https://martinfowler.com/bliki/BlueGreenDeployment.html">blue-green deployments</a>, predictable rollbacks, and shadow testing are some of the ways to improve change safety. Anyone interested in increasing deployment frequency must also invest in such safety strategies.</p>
<p>The second area of investment is <strong>fault containment and redundancy</strong>. Some of the complex incidents take time to restore, and traffic shifting to a redundant copy (active or passive) may provide a faster and reliable alternative to in-place fire-fighting. See my article on <a href="https://m.subbu.org/fault-domains-and-the-vegas-rule-923fc037119">Fault Domains and the Vegas Rule</a> for a description of how redundant fault domains can help reduce time to restore. Another excellent article to read on this topic is <a href="https://www.allthingsdistributed.com/">Werner Vogels</a>’ <a href="https://www.allthingsdistributed.com/2018/03/ten-years-of-aws-compartimentalization.html">Looking back at 10 years of compartmentalization at AWS</a> which describes how AWS uses “compartmentalization” for horizontal scalability as well as to contain faults to smaller domains.</p>
<p>However, maintaining redundancy is non-trivial. Apart from designing for redundancy, periodic traffic-shifting practice drills are essential for maintaining fault domain integrity and readiness to shift traffic.</p>
<p>The third area of investment is to <strong>either commit fully to automate systems or use a cloud-managed service to take care of most of the automation</strong>. I always recommend the latter due to increased time to market and lower operational overhead. Though this does not fully eliminate the possibility of config drift, it can at least help reduce the number of moving parts you’ve to automate yourself.</p>
<p>Next in the list of areas of investment is <strong>observability</strong>, in particular, tracing, to improve steady-state understanding of today’s hyperconnected production environments. Traces and service graphs help improve a team’s understanding of how their services are used and how they are behaving during the steady-state.</p>
<p>The fifth and the second most important area after change safety is investing to <strong>increase the time spent after incidents</strong> through post-incident rituals. Across the industry, most teams treat incidents as distractions and are eager to get back to regularly scheduled work as soon as systems are restored. This trend needs to change as incidents teach us about non-linear behaviors of complex hyperconnected systems. At work, we’re experimenting incorporation of a few post-incident rituals like peer-reviews of postmortem reports, and in some cases, subjecting the system in production to the similar triggers after fixes have been made.</p>
<p>Prior to this analysis, my approach to improving the availability of production systems involved adopting defensive strategies like <a href="https://github.com/Netflix/Hystrix">Hystrix</a>, ensuring redundancy, and adopting chaos testing. These are all essential techniques in a toolbox. This analysis gave a perspective on where to zoom in, and of course, boldly highlighted the need for change safety.</p>
<p>Let me end with a caveat. Any analysis like this will highlight some broad strokes while obscuring specifics. Take such findings as one of several inputs.</p>
</div>
<div class="subscribe-cta">
<p><em>If you enjoyed this essay, consider subscribing for future essays. I write about
technology and leadership. I will never share your email with anyone. You can unsubscribe
at any time.</em></p>
<form
name="newsletter"
method="POST"
action="/api/subscribe"
>
<input class="hidden" name="gotcha" />
<input type="email" name="email" id="bd-email" placeholder="Your email" required />
<div class="cf-turnstile" data-sitekey="0x4AAAAAACqsW1JNiZtH5Pc7"></div>
<input type="submit" value="Subscribe" />
</form>
<script src="https://challenges.cloudflare.com/turnstile/v0/api.js" async defer></script>
</div>
<nav class="post-nav">
<a class="post-nav-prev" href="https://www.subbu.org/essays/2019/the-value-is-in-dealing-with-the-messy-stuff/">&#8592; The Value is in Dealing with the Messy Stuff</a>
<a class="post-nav-next" href="https://www.subbu.org/essays/2019/opinions/">Opinions &#8594;</a>
</nav>
<h2 class="related-title">See Also</h2>
<h3 class="related-subtitle">Related</h3>
<ul class="related-posts">
<li>
<a href="https://www.subbu.org/essays/2019/studying-an-incident/">Studying an Incident</a> <small>(92% match, Dec 30, 2019)</small>
</li>
<li>
<a href="https://www.subbu.org/essays/2019/why-learn-from-incidents/">Why Learn from Incidents</a> <small>(92% match, Dec 17, 2019)</small>
</li>
<li>
<a href="https://www.subbu.org/essays/2019/if-only-production-incidents-could-speak/">If Only Production Incidents Could Speak</a> <small>(92% match, Jul 18, 2019)</small>
</li>
<li>
<a href="https://www.subbu.org/essays/2017/fault-domains-and-the-vegas-rule/">Fault Domains and the Vegas Rule</a> <small>(90% match, Feb 17, 2017)</small>
</li>
<li>
<a href="https://www.subbu.org/essays/2019/forming-failure-hypothesis/">Forming Failure Hypothesis</a> <small>(89% match, Sep 27, 2019)</small>
</li>
</ul>
<h3 class="related-subtitle">More to Read</h3>
<ul class="related-posts">
<li>
<a href="/essays/2026/the-elephant-in-the-brownfield/">The Elephant in the Brownfield</a> <small>(Aug 21, 2026)</small>
</li>
<li>
<a href="/essays/2026/two-bars/">Two Bars, Spreading Apart</a> <small>(Jul 28, 2026)</small>
</li>
<li>
<a href="/essays/2026/a-lesson-from-the-feedlot/">A Lesson From the Feedlot</a> <small>(Jun 3, 2026)</small>
</li>
<li>
<a href="/essays/2026/a-lesson-from-the-cockpit/">A Lesson From the Cockpit</a> <small>(May 8, 2026)</small>
</li>
<li>
<a href="/essays/2026/the-hold/">The Hold</a> <small>(Apr 22, 2026)</small>
</li>
</ul>
</div>
</div>
<div class="footer wrapper">
<nav class="nav">
<div>2026 © Subbu Allamaraju </div>
</nav>
</div><script>feather.replace()</script>
<!-- Cloudflare Pages Analytics --><script defer src='https://static.cloudflareinsights.com/beacon.min.js' data-cf-beacon='{"token": "a90d59c0629844c39a1d6db2cb98e81a"}'></script><!-- Cloudflare Pages Analytics --><script>(function(){function c(){var b=a.contentDocument||(a.contentWindow&&a.contentWindow.document);if(b){var d=b.createElement('script');d.innerHTML="window.__CF$cv$params={r:'a39ab76e1b4ecba0',t:'MTc4OTE3MjExMw=='};var a=document.createElement('script');a.src='/cdn-cgi/challenge-platform/scripts/jsd/main.js';document.getElementsByTagName('head')[0].appendChild(a);";b.getElementsByTagName('head')[0].appendChild(d)}}if(document.body){var a=document.createElement('iframe');a.height=1;a.width=1;a.style.position='absolute';a.style.top=0;a.style.left=0;a.style.border='none';a.style.visibility='hidden';document.body.appendChild(a);if('loading'!==document.readyState)c();else if(window.addEventListener)document.addEventListener('DOMContentLoaded',c);else{var e=document.onreadystatechange||function(){};document.onreadystatechange=function(b){e(b);'loading'!==document.readyState&&(document.onreadystatechange=e,c())}}}})();</script></body>
</html>