SRE weekly 所有文章
This commit is contained in:
@@ -0,0 +1,320 @@
|
||||
<!DOCTYPE html>
|
||||
<html lang="en">
|
||||
<head>
|
||||
<meta charset="utf-8">
|
||||
<title>Failure is Familiar, Safety is Surprising</title>
|
||||
<!-- https://web.dev/articles/responsive-web-design-basics -->
|
||||
<meta name="viewport" content="width=device-width, initial-scale=1">
|
||||
<link rel="stylesheet" href="/css/style.css">
|
||||
<link rel="alternate" type="application/rss+xml" href="/rss.xml">
|
||||
<link rel="preload" href="/fonts/MonaSans.woff2" as="font" type="font/woff2" crossorigin>
|
||||
<style>
|
||||
body:hover {
|
||||
border-image: url("https://assets.ryanfrantz.com/hit/posts/failure-is-familiar-safety-is-surprising.html");
|
||||
}
|
||||
</style>
|
||||
</head>
|
||||
<body>
|
||||
<banner>
|
||||
<a href="/">Ryan Frantz</a>
|
||||
</banner>
|
||||
<header>
|
||||
<nav>
|
||||
<ul>
|
||||
<li><a href="/archive/">Archive</a></li>
|
||||
<li><a href="/contact">Contact</a></li>
|
||||
<li><a href="/cv">CV</a></li>
|
||||
<li><a href="/papers/">Papers</a></li>
|
||||
<li><a href="/reading/">Reading</a></li>
|
||||
<li><a href="/tags/">Tags</a></li>
|
||||
<li><a href="/talks/">Talks</a></li>
|
||||
</ul>
|
||||
</nav>
|
||||
</header>
|
||||
<main>
|
||||
<div class="title">
|
||||
Failure is Familiar, Safety is Surprising
|
||||
</div>
|
||||
<time datetime="2019-05-05">
|
||||
5 May 2019
|
||||
</time>
|
||||
|
||||
<blockquote>
|
||||
Knowledge and error flow from the same mental sources; only success can tell
|
||||
one from the other.
|
||||
<footer>
|
||||
Ernst Mach
|
||||
</footer>
|
||||
</blockquote>
|
||||
|
||||
<blockquote>
|
||||
How did this ever work?!
|
||||
<footer>
|
||||
Every person that has ever attempted to write software.
|
||||
</footer>
|
||||
</blockquote>
|
||||
|
||||
<p>The complexity of our systems means they are not easily decomposable; it is rare
|
||||
that modern systems can be broken down to a single service running on a single
|
||||
host, for example. Further, they are always changing and the conditions in which
|
||||
they operate are dynamic. Crucially, new behavior is constantly emerging. By the
|
||||
time we think we’ve mapped them, the terrain has updated underneath us. It is
|
||||
surprising that our systems work at all. Yet they do and we take it for granted.</p>
|
||||
|
||||
<p>We’ve been studying failure for a very long time, cataloging the ways in which
|
||||
systems break down; it’s what allows us to build more robust systems that can
|
||||
navigate known issues. This is what Hollnagel and others refer to as Safety-I,
|
||||
defined as “the absence of accidents and incidents”. We have learned to identify
|
||||
numerous scenarios and to build in specific defenses to address them. The absence
|
||||
of incidents can only be achieved if we have complete knowledge of all possible
|
||||
failure states. If I have even a superficial grasp of the law of requisite
|
||||
variety, I believe I could say that our systems do not obey it. That is, as
|
||||
long as the set of failures are known and thus finite, we can count on them
|
||||
being addressed (as they present themselves) and we can, therefore, claim safety
|
||||
as the absence of accidents. But there are always new failures yet to be
|
||||
discovered.</p>
|
||||
|
||||
<p>While those failures tend to capture our imagination there is ample opportunity to
|
||||
learn more about our systems by studying how “things go right” (Hollnagel).
|
||||
By taking a different perspective on safety, Safety-II, defined as “the ability
|
||||
to succeed under varying conditions”, we may surface the capacity within
|
||||
organizations that allows them to adapt to uncertain events and situations and
|
||||
keep their systems running. Feebly attempting to apply the law of requisite
|
||||
variety, here, the more we understand about how people interoperate with their
|
||||
systems the more we can apply that knowledge to regulate them. The presence of
|
||||
expertise is the driving force behind the safety of our systems.</p>
|
||||
|
||||
<h2 id="a-failure-by-any-other-name">A Failure by Any Other Name…</h2>
|
||||
|
||||
<p>Our perception of how our systems operate is often binary: they either
|
||||
“just work” or they fail. And when things do fail, we are often surprised.
|
||||
However, if we give it some consideration we will find that failure is very
|
||||
familiar to us.</p>
|
||||
|
||||
<p>Software engineers constantly experience a world filled with failure. So much so,
|
||||
that we’ve erected soft monuments to them such as the <code class="language-plaintext highlighter-rouge">errno.h</code> header in the
|
||||
Linux kernel code. If I run a program on a Linux host, it can fail in one of
|
||||
several hundred known and defined ways. Those definitions are assigned a unique
|
||||
number that may be provided as an exit status. See some examples in the snippet
|
||||
below:</p>
|
||||
|
||||
<pre>
|
||||
#define EPERM 1 /* Operation not permitted */
|
||||
#define ENOENT 2 /* No such file or directory */
|
||||
...
|
||||
#define EMFILE 24 /* Too many open files */
|
||||
#define ENOTTY 25 /* Not a typewriter */
|
||||
...
|
||||
/*
|
||||
* This error code is special: arch syscall entry code will return
|
||||
* -ENOSYS if users try to call a syscall that doesn't exist. To keep
|
||||
* failures of syscalls that really do exist distinguishable from
|
||||
* failures due to attempts to use a nonexistent syscall, syscall
|
||||
* implementations should refrain from returning -ENOSYS.
|
||||
*/
|
||||
#define ENOSYS 38 /* Invalid system call number */
|
||||
</pre>
|
||||
|
||||
<p>Note the fine-grained (possibly esoteric) errors like <code class="language-plaintext highlighter-rouge">Not a typewriter</code> and
|
||||
cases where specifying an error, <code class="language-plaintext highlighter-rouge">ENOSYS</code>, may be invalid (an error within an
|
||||
error!).</p>
|
||||
|
||||
<p>Indeed, failure is familiar to us. We know its various names. The more we
|
||||
encounter it the more robust we can build our systems. Robustness, as a quality
|
||||
of our systems, is a static property. A robust system is expected to withstand
|
||||
known variables and bounded quantities; it is not guaranteed to be safe under
|
||||
dynamic conditions.</p>
|
||||
|
||||
<h2 id="success-is-invisible">Success Is Invisible</h2>
|
||||
|
||||
<p>A single exit status defines success for a Linux program: <code class="language-plaintext highlighter-rouge">0</code>, that is, “zero”.
|
||||
Simply. Zero. There are no status codes indicating “passed with flying colors”
|
||||
or “passed by the seat of its pants”. Success is simply success,
|
||||
regardless of what work the program performed or what obstacles it avoided. The
|
||||
details of how the code works, how it was able to address potential failures,
|
||||
and what contributed to its success are likely invisible to the operator.</p>
|
||||
|
||||
<p>Many of you reading this likely are able to drive a car and understand how it
|
||||
works, on the surface, at least. But I’m willing to bet, unless you’re a
|
||||
mechanic (or in training to be one) you likely don’t know much about what makes
|
||||
it “go” beyond turning the ignition, pushing the accelerator and brake pedals,
|
||||
and programming your favorite station into the radio.
|
||||
If you own a car, it’s likely been more than 3 months or 3000 miles since you’ve
|
||||
had the oil changed. Unfortunately, for many of us, learning how crucial oil
|
||||
is to the operation of an automobile comes through failure, when we have our
|
||||
vehicle towed to the closest garage because the oil pan is dry and all the metal
|
||||
bits that should not have rubbed together, did. But if you know more about what
|
||||
makes a car run successfully, such as the fact that oil helps lubricate the
|
||||
engine, keep it cool, and remove particulates, you’re more likely to keep to a
|
||||
regular maintenance schedule.</p>
|
||||
|
||||
<p>Success is invisible. That is, the work that goes into creating the conditions
|
||||
for success can be difficult to describe or see. It is driven by our expertise
|
||||
and collective tacit knowledge. This seems a paradox, that we could be
|
||||
successful yet not fully understand the factors that contribute to things
|
||||
going “right”. <a href="https://www.researchgate.net/publication/224753269_The_Messy_Details_Insights_From_the_Study_of_Technical_Work_in_Healthcare">Nemeth et al</a>
|
||||
describe the dynamics of complex systems as “messy details” that “operators
|
||||
navigate and negotiate” to “create success.” They go on to inform us why it is
|
||||
difficult to get at the reasons for that success:</p>
|
||||
|
||||
<blockquote>
|
||||
[A] basic difficulty arises and is captured by the law of fluency in
|
||||
cognitive systems: “well adapted cognitive work occurs with a facility that
|
||||
belies the difficulty of the demands resolved and the dilemmas balanced” .
|
||||
</blockquote>
|
||||
|
||||
<!--
|
||||
Lest I forget how clever I am, the below unicode character represents the math
|
||||
symbol for a strict superset. That is, A is a superset of B but B is not equal
|
||||
to A.
|
||||
-->
|
||||
<h2 id="safety-ii--safety-i">Safety-II ⊃ Safety-I</h2>
|
||||
|
||||
<p>Failures are dramatic and eye-catching. They’re also often very local
|
||||
and narrowly focused. When the dust settles and we move on, we should not take
|
||||
for granted that our systems are “back to normal.” We should strive to
|
||||
understand what “normal” is for us. We must study our systems and organizations
|
||||
from the Safety-II perspective: seeking to understand the totality of what it
|
||||
means to operate.</p>
|
||||
|
||||
<blockquote cite="http://www.safetydifferently.com/what-safety-ii-isnt/">
|
||||
Safety-II is about all possible outcomes: involving normal, everyday, routine
|
||||
performance; exceptionally good performance: and near-misses accidents and
|
||||
disasters. Our traditional approach, Safety-I, has largely limited itself to the
|
||||
latter – the accidents (actual or potential) at the tail end of the distribution.
|
||||
Safety-II is about the whole distribution, and its profile.<br />
|
||||
<footer>
|
||||
Steven Shorrock, <cite>http://www.safetydifferently.com/what-safety-ii-isnt/</cite>
|
||||
</footer>
|
||||
</blockquote>
|
||||
|
||||
<p>Safety II and Safety-I are not mutually exclusive. Indeed, as we endeavor to
|
||||
discover the numerous ways that a system works successfully, we are bound to
|
||||
uncover even more failure scenarios:</p>
|
||||
|
||||
<blockquote cite="https://how.complexsystems.fail/#4">
|
||||
Complex systems contain changing mixtures of failures latent within them.<br />
|
||||
<footer>
|
||||
Dr. Richard Cook, <cite>How Complex Systems Fail</cite>
|
||||
</footer>
|
||||
</blockquote>
|
||||
|
||||
<p>How do our people adapt to those scenarios? Cook continues by stating that our
|
||||
systems “run in degraded mode.” How is that possible if not for the capacities
|
||||
that people bring? There are numerous stories we can tell each other about the
|
||||
creative ways we’ve used bubble gum and duct tape. Many engineers know of, or
|
||||
have implemented, a <code class="language-plaintext highlighter-rouge">cron</code> job to restart a process every <code class="language-plaintext highlighter-rouge">N-1</code> days where <code class="language-plaintext highlighter-rouge">N</code>
|
||||
is the count of days when that process tends to fail.</p>
|
||||
|
||||
<p>We should also learn why our systems exist in the first place: what
|
||||
purpose they serve; what benefit they provide. As we do so we will begin to
|
||||
imagine the systems more “in the world” <a href="#fn_1">[1]</a> and begin to understand how
|
||||
they behave in a larger context. What are the intended uses of these systems?
|
||||
How have new uses and expectations about operation accreted over time? How does
|
||||
its behavior rely on or impact that of others? This may be overwhelming at first
|
||||
but in time, we will come to see that this expanded view helps to clarify what
|
||||
normal performance is for our systems and their role within their environment.</p>
|
||||
|
||||
<h2 id="we-have-work-to-do">We Have Work to Do</h2>
|
||||
|
||||
<p>Success need not be surprising. The more we study how our systems behave, the
|
||||
more expertly we will be able to operate them. That expertise will power the
|
||||
resilience within our organizations that we use to keep our systems running. An
|
||||
outcome of this is safer systems.</p>
|
||||
|
||||
<p>These learnings will not come easy. What people have internalized during their
|
||||
careers, what they’ve learned about the organizations they work within, and how
|
||||
they’ve adapted as the systems they operate have changed requires dedicated
|
||||
effort. Luckily, getting started is not difficult.</p>
|
||||
|
||||
<p>Your organization likely has some sort of artifacts laying around like design
|
||||
documents, <a href="/posts/architecture-reviews.html">architecture review</a> meeting minutes,
|
||||
or post-incident reports that capture critical context or decision-making
|
||||
that has occurred in the past. This content will lay the groundwork for what
|
||||
people’s initial expectations were. This is the baseline from which we can
|
||||
understand any deviations from initial intent and models that people use to do
|
||||
their work. But, for many reasons, these artifacts will be incomplete so your
|
||||
next step <strong>must</strong> be to <em>talk to your people</em>.</p>
|
||||
|
||||
<p>Go, talk to your people. Understand what daily operations look like, what
|
||||
obstacles they’ve encountered and how they’ve worked around them. Then,
|
||||
socialize that knowledge so that others can learn from it and use it to adapt
|
||||
their own work. The more we do this, the more likely we come to see success an
|
||||
an everyday part of our work.</p>
|
||||
|
||||
<h2 id="footnotes">Footnotes</h2>
|
||||
|
||||
<p><a name="fn_1"></a>
|
||||
[1]</p>
|
||||
<blockquote cite="http://www.humanfactors.lth.se/fileadmin/lusa/Sidney_Dekker/articles/2007/SafetyScienceMonitor.pdf">
|
||||
[N]ew view stories... tend[] to end up in the world, in the system in which
|
||||
people work[], systems which people made work in the first place.
|
||||
<footer>
|
||||
Siegenthaler and Laursen,
|
||||
<cite>
|
||||
<a href="http://www.humanfactors.lth.se/fileadmin/lusa/Sidney_Dekker/articles/2007/SafetyScienceMonitor.pdf">Six stages to the new view of human error</a>
|
||||
</cite>
|
||||
</footer>
|
||||
</blockquote>
|
||||
|
||||
|
||||
|
||||
|
||||
<hr>
|
||||
<small>
|
||||
Tags
|
||||
<ul>
|
||||
|
||||
<li><a href="/tags#Engineering">Engineering</a>
|
||||
|
||||
<li><a href="/tags#Human Factors">Human Factors</a>
|
||||
|
||||
<li><a href="/tags#Learning">Learning</a>
|
||||
|
||||
<li><a href="/tags#Systems Thinking">Systems Thinking</a>
|
||||
|
||||
</ul>
|
||||
</small>
|
||||
|
||||
|
||||
</main>
|
||||
<footer>
|
||||
<figure>
|
||||
<figcaption>
|
||||
<a rel="me" href="https://hachyderm.io/@frantz">
|
||||
Mastodon
|
||||
</a>
|
||||
</figcaption>
|
||||
</figure>
|
||||
<figure>
|
||||
<figcaption>
|
||||
<a href="https://github.com/RyanFrantz">
|
||||
GitHub
|
||||
</a>
|
||||
</figcaption>
|
||||
</figure>
|
||||
<figure>
|
||||
<figcaption>
|
||||
<a href="https://bsky.app/profile/ryanfrantz.bsky.social">
|
||||
Bluesky
|
||||
</a>
|
||||
</figcaption>
|
||||
</figure>
|
||||
<figure>
|
||||
<figcaption>
|
||||
<a href="https://www.linkedin.com/in/theryanfrantz/">
|
||||
LinkedIn
|
||||
</a>
|
||||
</figcaption>
|
||||
</figure>
|
||||
<figure>
|
||||
<figcaption>
|
||||
<a href="/rss.xml">
|
||||
RSS
|
||||
</a>
|
||||
</figcaption>
|
||||
</figure>
|
||||
</footer>
|
||||
|
||||
</body>
|
||||
</html>
|
||||
Reference in New Issue
Block a user