Files
nexus/sreweekly/articles/172/03-failure-is-familiar-safety-is-surprising.html
2026-09-12 17:23:01 +08:00

321 lines
14 KiB
HTML
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="utf-8">
<title>Failure is Familiar, Safety is Surprising</title>
<!-- https://web.dev/articles/responsive-web-design-basics -->
<meta name="viewport" content="width=device-width, initial-scale=1">
<link rel="stylesheet" href="/css/style.css">
<link rel="alternate" type="application/rss+xml" href="/rss.xml">
<link rel="preload" href="/fonts/MonaSans.woff2" as="font" type="font/woff2" crossorigin>
<style>
body:hover {
border-image: url("https://assets.ryanfrantz.com/hit/posts/failure-is-familiar-safety-is-surprising.html");
}
</style>
</head>
<body>
<banner>
<a href="/">Ryan Frantz</a>
</banner>
<header>
<nav>
<ul>
<li><a href="/archive/">Archive</a></li>
<li><a href="/contact">Contact</a></li>
<li><a href="/cv">CV</a></li>
<li><a href="/papers/">Papers</a></li>
<li><a href="/reading/">Reading</a></li>
<li><a href="/tags/">Tags</a></li>
<li><a href="/talks/">Talks</a></li>
</ul>
</nav>
</header>
<main>
<div class="title">
Failure is Familiar, Safety is Surprising
</div>
<time datetime="2019-05-05">
5 May 2019
</time>
<blockquote>
Knowledge and error flow from the same mental sources; only success can tell
one from the other.
<footer>
Ernst Mach
</footer>
</blockquote>
<blockquote>
How did this ever work?!
<footer>
Every person that has ever attempted to write software.
</footer>
</blockquote>
<p>The complexity of our systems means they are not easily decomposable; it is rare
that modern systems can be broken down to a single service running on a single
host, for example. Further, they are always changing and the conditions in which
they operate are dynamic. Crucially, new behavior is constantly emerging. By the
time we think we’ve mapped them, the terrain has updated underneath us. It is
surprising that our systems work at all. Yet they do and we take it for granted.</p>
<p>We’ve been studying failure for a very long time, cataloging the ways in which
systems break down; it’s what allows us to build more robust systems that can
navigate known issues. This is what Hollnagel and others refer to as Safety-I,
defined as “the absence of accidents and incidents”. We have learned to identify
numerous scenarios and to build in specific defenses to address them. The absence
of incidents can only be achieved if we have complete knowledge of all possible
failure states. If I have even a superficial grasp of the law of requisite
variety, I believe I could say that our systems do not obey it. That is, as
long as the set of failures are known and thus finite, we can count on them
being addressed (as they present themselves) and we can, therefore, claim safety
as the absence of accidents. But there are always new failures yet to be
discovered.</p>
<p>While those failures tend to capture our imagination there is ample opportunity to
learn more about our systems by studying how “things go right” (Hollnagel).
By taking a different perspective on safety, Safety-II, defined as “the ability
to succeed under varying conditions”, we may surface the capacity within
organizations that allows them to adapt to uncertain events and situations and
keep their systems running. Feebly attempting to apply the law of requisite
variety, here, the more we understand about how people interoperate with their
systems the more we can apply that knowledge to regulate them. The presence of
expertise is the driving force behind the safety of our systems.</p>
<h2 id="a-failure-by-any-other-name">A Failure by Any Other Name…</h2>
<p>Our perception of how our systems operate is often binary: they either
“just work” or they fail. And when things do fail, we are often surprised.
However, if we give it some consideration we will find that failure is very
familiar to us.</p>
<p>Software engineers constantly experience a world filled with failure. So much so,
that we’ve erected soft monuments to them such as the <code class="language-plaintext highlighter-rouge">errno.h</code> header in the
Linux kernel code. If I run a program on a Linux host, it can fail in one of
several hundred known and defined ways. Those definitions are assigned a unique
number that may be provided as an exit status. See some examples in the snippet
below:</p>
<pre>
#define EPERM 1 /* Operation not permitted */
#define ENOENT 2 /* No such file or directory */
...
#define EMFILE 24 /* Too many open files */
#define ENOTTY 25 /* Not a typewriter */
...
/*
* This error code is special: arch syscall entry code will return
* -ENOSYS if users try to call a syscall that doesn't exist. To keep
* failures of syscalls that really do exist distinguishable from
* failures due to attempts to use a nonexistent syscall, syscall
* implementations should refrain from returning -ENOSYS.
*/
#define ENOSYS 38 /* Invalid system call number */
</pre>
<p>Note the fine-grained (possibly esoteric) errors like <code class="language-plaintext highlighter-rouge">Not a typewriter</code> and
cases where specifying an error, <code class="language-plaintext highlighter-rouge">ENOSYS</code>, may be invalid (an error within an
error!).</p>
<p>Indeed, failure is familiar to us. We know its various names. The more we
encounter it the more robust we can build our systems. Robustness, as a quality
of our systems, is a static property. A robust system is expected to withstand
known variables and bounded quantities; it is not guaranteed to be safe under
dynamic conditions.</p>
<h2 id="success-is-invisible">Success Is Invisible</h2>
<p>A single exit status defines success for a Linux program: <code class="language-plaintext highlighter-rouge">0</code>, that is, “zero”.
Simply. Zero. There are no status codes indicating “passed with flying colors”
or “passed by the seat of its pants”. Success is simply success,
regardless of what work the program performed or what obstacles it avoided. The
details of how the code works, how it was able to address potential failures,
and what contributed to its success are likely invisible to the operator.</p>
<p>Many of you reading this likely are able to drive a car and understand how it
works, on the surface, at least. But I’m willing to bet, unless you’re a
mechanic (or in training to be one) you likely don’t know much about what makes
it “go” beyond turning the ignition, pushing the accelerator and brake pedals,
and programming your favorite station into the radio.
If you own a car, it’s likely been more than 3 months or 3000 miles since you’ve
had the oil changed. Unfortunately, for many of us, learning how crucial oil
is to the operation of an automobile comes through failure, when we have our
vehicle towed to the closest garage because the oil pan is dry and all the metal
bits that should not have rubbed together, did. But if you know more about what
makes a car run successfully, such as the fact that oil helps lubricate the
engine, keep it cool, and remove particulates, you’re more likely to keep to a
regular maintenance schedule.</p>
<p>Success is invisible. That is, the work that goes into creating the conditions
for success can be difficult to describe or see. It is driven by our expertise
and collective tacit knowledge. This seems a paradox, that we could be
successful yet not fully understand the factors that contribute to things
going “right”. <a href="https://www.researchgate.net/publication/224753269_The_Messy_Details_Insights_From_the_Study_of_Technical_Work_in_Healthcare">Nemeth et al</a>
describe the dynamics of complex systems as “messy details” that “operators
navigate and negotiate” to “create success.” They go on to inform us why it is
difficult to get at the reasons for that success:</p>
<blockquote>
[A] basic difficulty arises and is captured by the law of fluency in
cognitive systems: “well adapted cognitive work occurs with a facility that
belies the difficulty of the demands resolved and the dilemmas balanced” .
</blockquote>
<!--
Lest I forget how clever I am, the below unicode character represents the math
symbol for a strict superset. That is, A is a superset of B but B is not equal
to A.
-->
<h2 id="safety-ii--safety-i">Safety-II ⊃ Safety-I</h2>
<p>Failures are dramatic and eye-catching. They’re also often very local
and narrowly focused. When the dust settles and we move on, we should not take
for granted that our systems are “back to normal.” We should strive to
understand what “normal” is for us. We must study our systems and organizations
from the Safety-II perspective: seeking to understand the totality of what it
means to operate.</p>
<blockquote cite="http://www.safetydifferently.com/what-safety-ii-isnt/">
Safety-II is about all possible outcomes: involving normal, everyday, routine
performance; exceptionally good performance: and near-misses accidents and
disasters. Our traditional approach, Safety-I, has largely limited itself to the
latter – the accidents (actual or potential) at the tail end of the distribution.
Safety-II is about the whole distribution, and its profile.<br />
<footer>
Steven Shorrock, <cite>http://www.safetydifferently.com/what-safety-ii-isnt/</cite>
</footer>
</blockquote>
<p>Safety II and Safety-I are not mutually exclusive. Indeed, as we endeavor to
discover the numerous ways that a system works successfully, we are bound to
uncover even more failure scenarios:</p>
<blockquote cite="https://how.complexsystems.fail/#4">
Complex systems contain changing mixtures of failures latent within them.<br />
<footer>
Dr. Richard Cook, <cite>How Complex Systems Fail</cite>
</footer>
</blockquote>
<p>How do our people adapt to those scenarios? Cook continues by stating that our
systems “run in degraded mode.” How is that possible if not for the capacities
that people bring? There are numerous stories we can tell each other about the
creative ways we’ve used bubble gum and duct tape. Many engineers know of, or
have implemented, a <code class="language-plaintext highlighter-rouge">cron</code> job to restart a process every <code class="language-plaintext highlighter-rouge">N-1</code> days where <code class="language-plaintext highlighter-rouge">N</code>
is the count of days when that process tends to fail.</p>
<p>We should also learn why our systems exist in the first place: what
purpose they serve; what benefit they provide. As we do so we will begin to
imagine the systems more “in the world” <a href="#fn_1">[1]</a> and begin to understand how
they behave in a larger context. What are the intended uses of these systems?
How have new uses and expectations about operation accreted over time? How does
its behavior rely on or impact that of others? This may be overwhelming at first
but in time, we will come to see that this expanded view helps to clarify what
normal performance is for our systems and their role within their environment.</p>
<h2 id="we-have-work-to-do">We Have Work to Do</h2>
<p>Success need not be surprising. The more we study how our systems behave, the
more expertly we will be able to operate them. That expertise will power the
resilience within our organizations that we use to keep our systems running. An
outcome of this is safer systems.</p>
<p>These learnings will not come easy. What people have internalized during their
careers, what they’ve learned about the organizations they work within, and how
they’ve adapted as the systems they operate have changed requires dedicated
effort. Luckily, getting started is not difficult.</p>
<p>Your organization likely has some sort of artifacts laying around like design
documents, <a href="/posts/architecture-reviews.html">architecture review</a> meeting minutes,
or post-incident reports that capture critical context or decision-making
that has occurred in the past. This content will lay the groundwork for what
people’s initial expectations were. This is the baseline from which we can
understand any deviations from initial intent and models that people use to do
their work. But, for many reasons, these artifacts will be incomplete so your
next step <strong>must</strong> be to <em>talk to your people</em>.</p>
<p>Go, talk to your people. Understand what daily operations look like, what
obstacles they’ve encountered and how they’ve worked around them. Then,
socialize that knowledge so that others can learn from it and use it to adapt
their own work. The more we do this, the more likely we come to see success an
an everyday part of our work.</p>
<h2 id="footnotes">Footnotes</h2>
<p><a name="fn_1"></a>
[1]</p>
<blockquote cite="http://www.humanfactors.lth.se/fileadmin/lusa/Sidney_Dekker/articles/2007/SafetyScienceMonitor.pdf">
[N]ew view stories... tend[] to end up in the world, in the system in which
people work[], systems which people made work in the first place.
<footer>
Siegenthaler and Laursen,
<cite>
<a href="http://www.humanfactors.lth.se/fileadmin/lusa/Sidney_Dekker/articles/2007/SafetyScienceMonitor.pdf">Six stages to the new view of human error</a>
</cite>
</footer>
</blockquote>
<hr>
<small>
Tags
<ul>
<li><a href="/tags#Engineering">Engineering</a>
<li><a href="/tags#Human Factors">Human Factors</a>
<li><a href="/tags#Learning">Learning</a>
<li><a href="/tags#Systems Thinking">Systems Thinking</a>
</ul>
</small>
</main>
<footer>
<figure>
<figcaption>
<a rel="me" href="https://hachyderm.io/@frantz">
Mastodon
</a>
</figcaption>
</figure>
<figure>
<figcaption>
<a href="https://github.com/RyanFrantz">
GitHub
</a>
</figcaption>
</figure>
<figure>
<figcaption>
<a href="https://bsky.app/profile/ryanfrantz.bsky.social">
Bluesky
</a>
</figcaption>
</figure>
<figure>
<figcaption>
<a href="https://www.linkedin.com/in/theryanfrantz/">
LinkedIn
</a>
</figcaption>
</figure>
<figure>
<figcaption>
<a href="/rss.xml">
RSS
</a>
</figcaption>
</figure>
</footer>
</body>
</html>