Antifragility is a Fragile Concept

Antifragility is a Fragile Concept

People new to Chaos Engineering sometimes ask me if it’s just another word for fault injection, failure testing, or antifragility. It’s quite distinct from all three, but let’s focus on the last one.

First, I feel a responsibility to say something about the book where the concept originated.

The term “antifragility” was coined by the author Nassim Taleb in his book Antifragile: Things that Gain from Disorder. As the journalist Tom Barlett wrote while profiling Taleb, “Antifragile feels like a compendium of people and things Taleb doesn't like.” He is unabashedly condescending, and rails conspiratorially against the “Soviet-Harvard Delusion.” In tone, the book makes an art of sour-grapes and thin-skinned swipes at a people, institutions, and even places. It reminds me of Ayn Rand’s book The New Left: The Anti-Industrial Revolution, in which she invites us to imagine “computers programmed by a bunch of hippies” in a vain attempt to mock people who do things because they care about one another. It is not a good book. I found it unpleasant to read. I do not recommend it.

Putting aside the tone and focusing on the content, what the book lacks in civility it mirrors in foundational research. This is understandable given the author’s open contempt for academia, but it leads to naive conclusions, misinformation, and many large errors. Some well-established theories and processes are presented as original creations of the author. Others, such as the nuances of systemic drift, resource allocation, organizational predispositions, and priority contention are papered over, maligned, or ignored completely.

The English language has many ways of describing things that grow stronger after being exposed to stress. The word “hormesis” is what Taleb was looking for, but while he acknowledges that word later in the book, along with hypertrophy and other ways of describing the phenomena of growing, he is unsatisfied with the treatment of previous descriptions.

He devises three categories: things that are fragile, robust, and antifragile. The former break under stress, the middle withstand stress up to some point, and the latter get stronger under stress. It’s a tidy categorization on the surface. Applying this ontology in any meaningful way immediately exposes problems.

For example: a system can be both fragile and antifragile, or robust and antifragile, because the first two categories describe a property at one point in time while the third describes a reaction after the fact. Think of a child’s bones: they are fragile, because they break easily, but antifragile, because they grow stronger when stressed. Think of a runner’s quadriceps: robust, because they can withstand a marathon, and antifragile, because they grow stronger with exercise. The framework of classification Taleb proposes is not internally consistent.

The quote that I most often see repeated from the book is, “Antifragility is beyond resilience or robustness. The resilient resists shocks and stays the same; the antifragile gets better.” This is just wrong. There is an entire field, High Reliability Organization (HRO), that specifically studies the resilience of high criticality systems like nuclear power plants and aircraft carriers using language around adaptation to previously unknown circumstance. There in another field, Resilience Engineering, which again is entirely focused on how organizations ‘get better’ at optimizing for safety. The list goes on. Nassim’s definition of “resilience” is shamelessly self-serving, so narrow that it wouldn’t be recognized by the industry that describes itself in terms of that very word.

So how does one become antifragile? Taleb proposes that one starts by finding instances of fragility by exposing a system to volatility and then remove those fragilities. I get why people confuse that with Chaos Engineering. There is a conceptual overlap, and the descriptions sound similar. Chaos Engineering programs facilitate experiments to uncover systemic weaknesses by exposing complex systems to “turbulent conditions in production.” Presumably this is done so those exposed weaknesses can be fixed. In that sense, Chaos Engineering could be interpreted as a tool for creating antifragile systems.

This is where the concept of antifragility veers from a truism into bad advice. 

The book suggests first focusing on removing fragility as a tactic for improving a system’s stability. If we embrace that as the foundational step, then we ignore the empirical evidence of most of what is known about organizational behavior and sociology.

As Rochlin says in a review of High Reliability Organization: “To the extent that regulators, systems designers, and analysts continue to focus their attention on the avoidance of error and the control of risk, and to seek objective and positivistic indicators of performance, safety becomes marginalized as a residual property. This not only neglects the importance of expressed and perceived safety as a constitutive property of safe operation, but may actually interfere with the means and processes by which it is created and maintained.” [Rochlin’s paper “Safe operation as a social construct” p1558]

You don’t create alignment within an organization by spending all of your time seeking misalignments and then removing them. Appreciative Inquiry comes to mind as a practice from organizational psychology that stands as an antithesis to Taleb’s antifragility. There are countless others. They share the property that they are additive and encourage the conditions for just-in-time adaptations that allow human operators to carry complex systems safely through unforeseen circumstances.

That’s how complex, distributed systems at scale survive, and it’s nothing at all like the mental model Taleb proposes in the concept ‘antifragility.’

He goes on to encourage other arm-chair strategies, like adding redundancies, etc. -- all of which are equally naive. Actual research suggests that: “Despite designers’ best intentions, redundancy can unwittingly increase the chances of an accident by encouraging operators to push safety limits well beyond where they would have, had such redundancies not been installed.” [Snook’s book “Friendly Fire” p239] The empirical analysis and research is significantly more sophisticated, actionable, and pragmatic than Taleb’s folk theory.

Chaos Engineering can be used to remove points of fragility in a system; however, that’s an unnecessary limitation on the value of Chaos Engineering. The more powerful property of Chaos Engineering is that it provides humans with a signal of the safety margin within which they operate. It empowers the people doing the actual work of building and maintaining the system. It puts them in a better position to navigate the chaos inherent in a complex system, prioritize and optimize for the stated and unstated business goals as they go, and respond to unforeseen events when they do happen.

Antifragile is a confused concept that ignores a wealth of knowledge we have about building and operating resilient systems. Chaos Engineering, by contrast, is a well-defined, pragmatic discipline for navigating complex systems. In my mind, they are quite far apart.

JJ Ruescas I know you would like this article

It would be great to read an expanded, more detailed version of your second to last paragraph; e.g., specifically how does chaos engineering empower the people doing the actual work?

Like
Reply

I think there's a background distinction between Taleb's work and the cases you describe. The important thing to get about Taleb's viewpoint is that it comes from a place (trading) where systems are almost completely opaque - there is nothing but uncertainty. In Engineering, we can see the mechanism and we can formulate strategies to be reactive by working with the mechanism. In trading, you can't bolster the mechanism. You don't get to see it or touch it. Model error can be disastrous, so you either look for a better model or develop an ontology that helps you navigate 'model uncertainty' and domains where we can't design reactive apparatus.

Late to the party, but I brought my own whiskey, so hopefully you'll let me in the door. On reading the comments, I was reminded of a couple of experiences neck deep in the Swamp of Resilience. The first is that a lot of people come to similar concept sets via disparate paths in their own ways. I have a lot to say about this phenomenon with respect to resilience in particular, but in particular that conceptual mirror suggests there's a lot of potential trade between disciplines that can raise everyone up. The second, as a counterpoint, is that some of these concept sets seem to come from a philosophical place and others from a problem solving place. The former are often not very useful, and later mired in the complexity of trying to get stuff done in a given field, so they're less sweepingly applicable. Antifragility is in the philosophical bucket. And, to back Casey up, it's frankly an ontological hot mess, even if there are some kernels of wisdom you can take with you. Stephen Johnson and John Day define a really useful ontology of failure in a pair of papers published by AIAA about 10 years ago, which we also integrated into a NASA Fault Management handbook ... I personally haven't found a concept set that is more useful.

To view or add a comment, sign in

More articles by Casey Rosenthal

  • Beatles Agents Problem

    John Allspaw included the following quote in his chapter, “People in the Loop,” in the Chaos Engineering book: “An…

    1 Comment
  • Decline and Fall of a Billionaire's Twitter

    "I can understand perfectly how the report of my illness got about. I have even heard on good authority that I was dead.

    7 Comments
  • Appreciating a Friend

    David Hussman is one of those people who makes you want to be a little bit better version of yourself. I met him at…

    1 Comment
  • Would a Chaos by any other Name

    A few people have asked me why Netflix changed the name of the Chaos Engineering Team, which I managed up until this…

    4 Comments
  • Personal Mission

    Utah Phillips once told me to keep track of the people that I owe. For that reason, I drop some names in this post.

    3 Comments
  • Engineering Personas

    SUMMARY Mental models can help us make decisions in complex situations. Choosing whom to hire is among the most complex…

    15 Comments
  • My Paternity Leave at Netflix

    Netflix offers one year of paid paternity leave to new fathers. I just had a baby in December.

    29 Comments

Others also viewed

Explore content categories