321 lines
14 KiB
HTML
321 lines
14 KiB
HTML
<!DOCTYPE html>
|
||
<html lang="en">
|
||
<head>
|
||
<meta charset="utf-8">
|
||
<title>Failure is Familiar, Safety is Surprising</title>
|
||
<!-- https://web.dev/articles/responsive-web-design-basics -->
|
||
<meta name="viewport" content="width=device-width, initial-scale=1">
|
||
<link rel="stylesheet" href="/css/style.css">
|
||
<link rel="alternate" type="application/rss+xml" href="/rss.xml">
|
||
<link rel="preload" href="/fonts/MonaSans.woff2" as="font" type="font/woff2" crossorigin>
|
||
<style>
|
||
body:hover {
|
||
border-image: url("https://assets.ryanfrantz.com/hit/posts/failure-is-familiar-safety-is-surprising.html");
|
||
}
|
||
</style>
|
||
</head>
|
||
<body>
|
||
<banner>
|
||
<a href="/">Ryan Frantz</a>
|
||
</banner>
|
||
<header>
|
||
<nav>
|
||
<ul>
|
||
<li><a href="/archive/">Archive</a></li>
|
||
<li><a href="/contact">Contact</a></li>
|
||
<li><a href="/cv">CV</a></li>
|
||
<li><a href="/papers/">Papers</a></li>
|
||
<li><a href="/reading/">Reading</a></li>
|
||
<li><a href="/tags/">Tags</a></li>
|
||
<li><a href="/talks/">Talks</a></li>
|
||
</ul>
|
||
</nav>
|
||
</header>
|
||
<main>
|
||
<div class="title">
|
||
Failure is Familiar, Safety is Surprising
|
||
</div>
|
||
<time datetime="2019-05-05">
|
||
5 May 2019
|
||
</time>
|
||
|
||
<blockquote>
|
||
Knowledge and error flow from the same mental sources; only success can tell
|
||
one from the other.
|
||
<footer>
|
||
Ernst Mach
|
||
</footer>
|
||
</blockquote>
|
||
|
||
<blockquote>
|
||
How did this ever work?!
|
||
<footer>
|
||
Every person that has ever attempted to write software.
|
||
</footer>
|
||
</blockquote>
|
||
|
||
<p>The complexity of our systems means they are not easily decomposable; it is rare
|
||
that modern systems can be broken down to a single service running on a single
|
||
host, for example. Further, they are always changing and the conditions in which
|
||
they operate are dynamic. Crucially, new behavior is constantly emerging. By the
|
||
time we think we’ve mapped them, the terrain has updated underneath us. It is
|
||
surprising that our systems work at all. Yet they do and we take it for granted.</p>
|
||
|
||
<p>We’ve been studying failure for a very long time, cataloging the ways in which
|
||
systems break down; it’s what allows us to build more robust systems that can
|
||
navigate known issues. This is what Hollnagel and others refer to as Safety-I,
|
||
defined as “the absence of accidents and incidents”. We have learned to identify
|
||
numerous scenarios and to build in specific defenses to address them. The absence
|
||
of incidents can only be achieved if we have complete knowledge of all possible
|
||
failure states. If I have even a superficial grasp of the law of requisite
|
||
variety, I believe I could say that our systems do not obey it. That is, as
|
||
long as the set of failures are known and thus finite, we can count on them
|
||
being addressed (as they present themselves) and we can, therefore, claim safety
|
||
as the absence of accidents. But there are always new failures yet to be
|
||
discovered.</p>
|
||
|
||
<p>While those failures tend to capture our imagination there is ample opportunity to
|
||
learn more about our systems by studying how “things go right” (Hollnagel).
|
||
By taking a different perspective on safety, Safety-II, defined as “the ability
|
||
to succeed under varying conditions”, we may surface the capacity within
|
||
organizations that allows them to adapt to uncertain events and situations and
|
||
keep their systems running. Feebly attempting to apply the law of requisite
|
||
variety, here, the more we understand about how people interoperate with their
|
||
systems the more we can apply that knowledge to regulate them. The presence of
|
||
expertise is the driving force behind the safety of our systems.</p>
|
||
|
||
<h2 id="a-failure-by-any-other-name">A Failure by Any Other Name…</h2>
|
||
|
||
<p>Our perception of how our systems operate is often binary: they either
|
||
“just work” or they fail. And when things do fail, we are often surprised.
|
||
However, if we give it some consideration we will find that failure is very
|
||
familiar to us.</p>
|
||
|
||
<p>Software engineers constantly experience a world filled with failure. So much so,
|
||
that we’ve erected soft monuments to them such as the <code class="language-plaintext highlighter-rouge">errno.h</code> header in the
|
||
Linux kernel code. If I run a program on a Linux host, it can fail in one of
|
||
several hundred known and defined ways. Those definitions are assigned a unique
|
||
number that may be provided as an exit status. See some examples in the snippet
|
||
below:</p>
|
||
|
||
<pre>
|
||
#define EPERM 1 /* Operation not permitted */
|
||
#define ENOENT 2 /* No such file or directory */
|
||
...
|
||
#define EMFILE 24 /* Too many open files */
|
||
#define ENOTTY 25 /* Not a typewriter */
|
||
...
|
||
/*
|
||
* This error code is special: arch syscall entry code will return
|
||
* -ENOSYS if users try to call a syscall that doesn't exist. To keep
|
||
* failures of syscalls that really do exist distinguishable from
|
||
* failures due to attempts to use a nonexistent syscall, syscall
|
||
* implementations should refrain from returning -ENOSYS.
|
||
*/
|
||
#define ENOSYS 38 /* Invalid system call number */
|
||
</pre>
|
||
|
||
<p>Note the fine-grained (possibly esoteric) errors like <code class="language-plaintext highlighter-rouge">Not a typewriter</code> and
|
||
cases where specifying an error, <code class="language-plaintext highlighter-rouge">ENOSYS</code>, may be invalid (an error within an
|
||
error!).</p>
|
||
|
||
<p>Indeed, failure is familiar to us. We know its various names. The more we
|
||
encounter it the more robust we can build our systems. Robustness, as a quality
|
||
of our systems, is a static property. A robust system is expected to withstand
|
||
known variables and bounded quantities; it is not guaranteed to be safe under
|
||
dynamic conditions.</p>
|
||
|
||
<h2 id="success-is-invisible">Success Is Invisible</h2>
|
||
|
||
<p>A single exit status defines success for a Linux program: <code class="language-plaintext highlighter-rouge">0</code>, that is, “zero”.
|
||
Simply. Zero. There are no status codes indicating “passed with flying colors”
|
||
or “passed by the seat of its pants”. Success is simply success,
|
||
regardless of what work the program performed or what obstacles it avoided. The
|
||
details of how the code works, how it was able to address potential failures,
|
||
and what contributed to its success are likely invisible to the operator.</p>
|
||
|
||
<p>Many of you reading this likely are able to drive a car and understand how it
|
||
works, on the surface, at least. But I’m willing to bet, unless you’re a
|
||
mechanic (or in training to be one) you likely don’t know much about what makes
|
||
it “go” beyond turning the ignition, pushing the accelerator and brake pedals,
|
||
and programming your favorite station into the radio.
|
||
If you own a car, it’s likely been more than 3 months or 3000 miles since you’ve
|
||
had the oil changed. Unfortunately, for many of us, learning how crucial oil
|
||
is to the operation of an automobile comes through failure, when we have our
|
||
vehicle towed to the closest garage because the oil pan is dry and all the metal
|
||
bits that should not have rubbed together, did. But if you know more about what
|
||
makes a car run successfully, such as the fact that oil helps lubricate the
|
||
engine, keep it cool, and remove particulates, you’re more likely to keep to a
|
||
regular maintenance schedule.</p>
|
||
|
||
<p>Success is invisible. That is, the work that goes into creating the conditions
|
||
for success can be difficult to describe or see. It is driven by our expertise
|
||
and collective tacit knowledge. This seems a paradox, that we could be
|
||
successful yet not fully understand the factors that contribute to things
|
||
going “right”. <a href="https://www.researchgate.net/publication/224753269_The_Messy_Details_Insights_From_the_Study_of_Technical_Work_in_Healthcare">Nemeth et al</a>
|
||
describe the dynamics of complex systems as “messy details” that “operators
|
||
navigate and negotiate” to “create success.” They go on to inform us why it is
|
||
difficult to get at the reasons for that success:</p>
|
||
|
||
<blockquote>
|
||
[A] basic difficulty arises and is captured by the law of fluency in
|
||
cognitive systems: “well adapted cognitive work occurs with a facility that
|
||
belies the difficulty of the demands resolved and the dilemmas balanced” .
|
||
</blockquote>
|
||
|
||
<!--
|
||
Lest I forget how clever I am, the below unicode character represents the math
|
||
symbol for a strict superset. That is, A is a superset of B but B is not equal
|
||
to A.
|
||
-->
|
||
<h2 id="safety-ii--safety-i">Safety-II ⊃ Safety-I</h2>
|
||
|
||
<p>Failures are dramatic and eye-catching. They’re also often very local
|
||
and narrowly focused. When the dust settles and we move on, we should not take
|
||
for granted that our systems are “back to normal.” We should strive to
|
||
understand what “normal” is for us. We must study our systems and organizations
|
||
from the Safety-II perspective: seeking to understand the totality of what it
|
||
means to operate.</p>
|
||
|
||
<blockquote cite="http://www.safetydifferently.com/what-safety-ii-isnt/">
|
||
Safety-II is about all possible outcomes: involving normal, everyday, routine
|
||
performance; exceptionally good performance: and near-misses accidents and
|
||
disasters. Our traditional approach, Safety-I, has largely limited itself to the
|
||
latter – the accidents (actual or potential) at the tail end of the distribution.
|
||
Safety-II is about the whole distribution, and its profile.<br />
|
||
<footer>
|
||
Steven Shorrock, <cite>http://www.safetydifferently.com/what-safety-ii-isnt/</cite>
|
||
</footer>
|
||
</blockquote>
|
||
|
||
<p>Safety II and Safety-I are not mutually exclusive. Indeed, as we endeavor to
|
||
discover the numerous ways that a system works successfully, we are bound to
|
||
uncover even more failure scenarios:</p>
|
||
|
||
<blockquote cite="https://how.complexsystems.fail/#4">
|
||
Complex systems contain changing mixtures of failures latent within them.<br />
|
||
<footer>
|
||
Dr. Richard Cook, <cite>How Complex Systems Fail</cite>
|
||
</footer>
|
||
</blockquote>
|
||
|
||
<p>How do our people adapt to those scenarios? Cook continues by stating that our
|
||
systems “run in degraded mode.” How is that possible if not for the capacities
|
||
that people bring? There are numerous stories we can tell each other about the
|
||
creative ways we’ve used bubble gum and duct tape. Many engineers know of, or
|
||
have implemented, a <code class="language-plaintext highlighter-rouge">cron</code> job to restart a process every <code class="language-plaintext highlighter-rouge">N-1</code> days where <code class="language-plaintext highlighter-rouge">N</code>
|
||
is the count of days when that process tends to fail.</p>
|
||
|
||
<p>We should also learn why our systems exist in the first place: what
|
||
purpose they serve; what benefit they provide. As we do so we will begin to
|
||
imagine the systems more “in the world” <a href="#fn_1">[1]</a> and begin to understand how
|
||
they behave in a larger context. What are the intended uses of these systems?
|
||
How have new uses and expectations about operation accreted over time? How does
|
||
its behavior rely on or impact that of others? This may be overwhelming at first
|
||
but in time, we will come to see that this expanded view helps to clarify what
|
||
normal performance is for our systems and their role within their environment.</p>
|
||
|
||
<h2 id="we-have-work-to-do">We Have Work to Do</h2>
|
||
|
||
<p>Success need not be surprising. The more we study how our systems behave, the
|
||
more expertly we will be able to operate them. That expertise will power the
|
||
resilience within our organizations that we use to keep our systems running. An
|
||
outcome of this is safer systems.</p>
|
||
|
||
<p>These learnings will not come easy. What people have internalized during their
|
||
careers, what they’ve learned about the organizations they work within, and how
|
||
they’ve adapted as the systems they operate have changed requires dedicated
|
||
effort. Luckily, getting started is not difficult.</p>
|
||
|
||
<p>Your organization likely has some sort of artifacts laying around like design
|
||
documents, <a href="/posts/architecture-reviews.html">architecture review</a> meeting minutes,
|
||
or post-incident reports that capture critical context or decision-making
|
||
that has occurred in the past. This content will lay the groundwork for what
|
||
people’s initial expectations were. This is the baseline from which we can
|
||
understand any deviations from initial intent and models that people use to do
|
||
their work. But, for many reasons, these artifacts will be incomplete so your
|
||
next step <strong>must</strong> be to <em>talk to your people</em>.</p>
|
||
|
||
<p>Go, talk to your people. Understand what daily operations look like, what
|
||
obstacles they’ve encountered and how they’ve worked around them. Then,
|
||
socialize that knowledge so that others can learn from it and use it to adapt
|
||
their own work. The more we do this, the more likely we come to see success an
|
||
an everyday part of our work.</p>
|
||
|
||
<h2 id="footnotes">Footnotes</h2>
|
||
|
||
<p><a name="fn_1"></a>
|
||
[1]</p>
|
||
<blockquote cite="http://www.humanfactors.lth.se/fileadmin/lusa/Sidney_Dekker/articles/2007/SafetyScienceMonitor.pdf">
|
||
[N]ew view stories... tend[] to end up in the world, in the system in which
|
||
people work[], systems which people made work in the first place.
|
||
<footer>
|
||
Siegenthaler and Laursen,
|
||
<cite>
|
||
<a href="http://www.humanfactors.lth.se/fileadmin/lusa/Sidney_Dekker/articles/2007/SafetyScienceMonitor.pdf">Six stages to the new view of human error</a>
|
||
</cite>
|
||
</footer>
|
||
</blockquote>
|
||
|
||
|
||
|
||
|
||
<hr>
|
||
<small>
|
||
Tags
|
||
<ul>
|
||
|
||
<li><a href="/tags#Engineering">Engineering</a>
|
||
|
||
<li><a href="/tags#Human Factors">Human Factors</a>
|
||
|
||
<li><a href="/tags#Learning">Learning</a>
|
||
|
||
<li><a href="/tags#Systems Thinking">Systems Thinking</a>
|
||
|
||
</ul>
|
||
</small>
|
||
|
||
|
||
</main>
|
||
<footer>
|
||
<figure>
|
||
<figcaption>
|
||
<a rel="me" href="https://hachyderm.io/@frantz">
|
||
Mastodon
|
||
</a>
|
||
</figcaption>
|
||
</figure>
|
||
<figure>
|
||
<figcaption>
|
||
<a href="https://github.com/RyanFrantz">
|
||
GitHub
|
||
</a>
|
||
</figcaption>
|
||
</figure>
|
||
<figure>
|
||
<figcaption>
|
||
<a href="https://bsky.app/profile/ryanfrantz.bsky.social">
|
||
Bluesky
|
||
</a>
|
||
</figcaption>
|
||
</figure>
|
||
<figure>
|
||
<figcaption>
|
||
<a href="https://www.linkedin.com/in/theryanfrantz/">
|
||
LinkedIn
|
||
</a>
|
||
</figcaption>
|
||
</figure>
|
||
<figure>
|
||
<figcaption>
|
||
<a href="/rss.xml">
|
||
RSS
|
||
</a>
|
||
</figcaption>
|
||
</figure>
|
||
</footer>
|
||
|
||
</body>
|
||
</html>
|