21 lines
20 KiB
HTML
21 lines
20 KiB
HTML
<!doctype html><html lang=en dir=auto data-theme=auto><head><meta charset=utf-8><meta http-equiv=X-UA-Compatible content="IE=edge"><meta name=viewport content="width=device-width,initial-scale=1,shrink-to-fit=no"><meta name=robots content="index, follow"><title>Designed to Fail | Brave New Geek</title><meta name=keywords content="availability,distributed systems,fault tolerance,messaging,microservices,ops,resilience engineering,soa,software engineering,systems,systems theory"><meta name=description content="When it comes to reliability engineering, people often talk about things like fault injection, monitoring, and operations runbooks. These are all critical pieces for building systems which can withstand failure, but what’s less talked about is the need to design systems which deliberately fail.
|
||
Reliability design has a natural progression which closely follows that of architectural design. With monolithic systems, we care more about preventing failure from occurring. With service-oriented architectures, controlling failure becomes less manageable, so instead we learn to anticipate it. With highly distributed microservice architectures where failure is all but guaranteed, we embrace it."><meta name=author content><link rel=canonical href=https://bravenewgeek.com/designed-to-fail/><link crossorigin=anonymous href=/assets/css/stylesheet.4861a452a4c13a9a1fbf2085400b74a7de96b1beeb94dee57b4273e5dffcf337.css integrity="sha256-SGGkUqTBOpofvyCFQAt0p96Wsb7rlN7le0Jz5d/88zc=" rel="preload stylesheet" as=style><link rel=icon href=https://bravenewgeek.com/favicon.ico><link rel=icon type=image/png sizes=16x16 href=https://bravenewgeek.com/favicon.ico><link rel=icon type=image/png sizes=32x32 href=https://bravenewgeek.com/favicon.ico><link rel=apple-touch-icon href=https://bravenewgeek.com/favicon.ico><link rel=mask-icon href=https://bravenewgeek.com/favicon.ico><meta name=theme-color content="#2e2e33"><meta name=msapplication-TileColor content="#2e2e33"><link rel=alternate hreflang=en href=https://bravenewgeek.com/designed-to-fail/><noscript><style>#theme-toggle,.top-link{display:none}</style><style>@media(prefers-color-scheme:dark){:root{--theme:rgb(29, 30, 32);--entry:rgb(46, 46, 51);--primary:rgb(218, 218, 219);--secondary:rgb(155, 156, 157);--tertiary:rgb(65, 66, 68);--content:rgb(196, 196, 197);--code-block-bg:rgb(46, 46, 51);--code-bg:rgb(55, 56, 62);--border:rgb(51, 51, 51);color-scheme:dark}.list{background:var(--theme)}.toc{background:var(--entry)}}</style></noscript><script>localStorage.getItem("pref-theme")==="dark"?document.querySelector("html").dataset.theme="dark":localStorage.getItem("pref-theme")==="light"?document.querySelector("html").dataset.theme="light":window.matchMedia("(prefers-color-scheme: dark)").matches?document.querySelector("html").dataset.theme="dark":document.querySelector("html").dataset.theme="light"</script><link rel=preconnect href=https://fonts.googleapis.com><link rel=preconnect href=https://fonts.gstatic.com crossorigin><link rel=stylesheet href="https://fonts.googleapis.com/css2?family=JetBrains+Mono:wght@400;500&family=Source+Serif+4:ital,opsz,wght@0,8..60,400;0,8..60,600;1,8..60,400&family=Space+Grotesk:wght@500;600;700&display=swap"><meta property="og:url" content="https://bravenewgeek.com/designed-to-fail/"><meta property="og:site_name" content="Brave New Geek"><meta property="og:title" content="Designed to Fail"><meta property="og:description" content="When it comes to reliability engineering, people often talk about things like fault injection, monitoring, and operations runbooks. These are all critical pieces for building systems which can withstand failure, but what’s less talked about is the need to design systems which deliberately fail.
|
||
Reliability design has a natural progression which closely follows that of architectural design. With monolithic systems, we care more about preventing failure from occurring. With service-oriented architectures, controlling failure becomes less manageable, so instead we learn to anticipate it. With highly distributed microservice architectures where failure is all but guaranteed, we embrace it."><meta property="og:locale" content="en_us"><meta property="og:type" content="article"><meta property="article:section" content="posts"><meta property="article:published_time" content="2015-07-21T20:17:17-05:00"><meta property="article:modified_time" content="2016-12-20T22:36:57-05:00"><meta property="article:tag" content="Availability"><meta property="article:tag" content="Distributed Systems"><meta property="article:tag" content="Fault Tolerance"><meta property="article:tag" content="Messaging"><meta property="article:tag" content="Microservices"><meta property="article:tag" content="Ops"><meta name=twitter:card content="summary"><meta name=twitter:title content="Designed to Fail"><meta name=twitter:description content="When it comes to reliability engineering, people often talk about things like fault injection, monitoring, and operations runbooks. These are all critical pieces for building systems which can withstand failure, but what’s less talked about is the need to design systems which deliberately fail.
|
||
Reliability design has a natural progression which closely follows that of architectural design. With monolithic systems, we care more about preventing failure from occurring. With service-oriented architectures, controlling failure becomes less manageable, so instead we learn to anticipate it. With highly distributed microservice architectures where failure is all but guaranteed, we embrace it."><script type=application/ld+json>{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Posts","item":"https://bravenewgeek.com/posts/"},{"@type":"ListItem","position":2,"name":"Designed to Fail","item":"https://bravenewgeek.com/designed-to-fail/"}]}</script><script type=application/ld+json>{"@context":"https://schema.org","@type":"BlogPosting","headline":"Designed to Fail","name":"Designed to Fail","description":"When it comes to reliability engineering, people often talk about things like fault injection, monitoring, and operations runbooks. These are all critical pieces for building systems which can withstand failure, but what’s less talked about is the need to design systems which deliberately fail.\nReliability design has a natural progression which closely follows that of architectural design. With monolithic systems, we care more about preventing failure from occurring. With service-oriented architectures, controlling failure becomes less manageable, so instead we learn to anticipate it. With highly distributed microservice architectures where failure is all but guaranteed, we embrace it.\n","keywords":["availability","distributed systems","fault tolerance","messaging","microservices","ops","resilience engineering","soa","software engineering","systems","systems theory"],"articleBody":"When it comes to reliability engineering, people often talk about things like fault injection, monitoring, and operations runbooks. These are all critical pieces for building systems which can withstand failure, but what’s less talked about is the need to design systems which deliberately fail.\nReliability design has a natural progression which closely follows that of architectural design. With monolithic systems, we care more about preventing failure from occurring. With service-oriented architectures, controlling failure becomes less manageable, so instead we learn to anticipate it. With highly distributed microservice architectures where failure is all but guaranteed, we embrace it.\nWhat does it mean to embrace failure? Anticipating failure is understanding the behavior when things go wrong, building systems to be resilient to it, and having a game plan for when it happens, either manual or automated. Embracing failure means making a conscious decision to purposely fail, and it’s essential for building highly available large-scale systems.\nA microservice architecture typically means a complex web of service dependencies. One of SOA’s goals is to isolate failure and allow for graceful degradation. The key to being highly available is learning to be partially available. Frequently, one of the requirements for partial availability is telling the client “no_.”_ Outright rejecting service requests is often better than allowing them to back up because, when dealing with distributed services, the latter usually results in cascading failure across dependent systems.\nWhile designing our distributed messaging service at Workiva, we made explicit decisions to drop messages on the floor if we detect the system is becoming overloaded. As queues become backed up, incoming messages are discarded, a statsd counter is incremented, and a backpressure notification is sent to the client. Upon receiving this notification, the client can respond accordingly by failing fast, exponentially backing off, or using some other flow-control strategy. By bounding resource utilization, we maintain predictable performance, predictable (and measurable) lossiness, and impede cascading failure.\nOther techniques include building kill switches into service calls and routers. If an overloaded service is not essential to core business, we fail fast on calls to it to prevent availability or latency problems upstream. For example, a spam-detection service is not essential to an email system, so if it’s unavailable or overwhelmed, we can simply bypass it. Netflix’s Hystrix has a set of really nice patterns for handling this.\nIf we’re not careful, we can often be our own worst enemy. Many times, it’s our own internal services which cause the biggest DoS attacks on ourselves. By isolating and controlling it, we can prevent failure from becoming widespread and unpredictable. By building in backpressure mechanisms and other types of intentional “failure” modes, we can ensure better availability and reliability for our systems through graceful degradation. Sometimes it’s better to fight fire with fire and failure with failure.\n","wordCount":"465","inLanguage":"en","datePublished":"2015-07-21T20:17:17-05:00","dateModified":"2016-12-20T22:36:57-05:00","mainEntityOfPage":{"@type":"WebPage","@id":"https://bravenewgeek.com/designed-to-fail/"},"publisher":{"@type":"Organization","name":"Brave New Geek","logo":{"@type":"ImageObject","url":"https://bravenewgeek.com/favicon.ico"}}}</script></head><body id=top><header class=site-header><div class="wrap site-header-inner"><a class=wordmark href=https://bravenewgeek.com/ accesskey=h title="Brave New Geek (Alt + H)"><span class=wordmark-name>Brave New Geek</span>
|
||
<span class=wordmark-tag>Introspections of a software engineer</span></a><nav class=site-nav aria-label=Primary><a href=https://bravenewgeek.com/archive/>Archive</a>
|
||
<a href=https://bravenewgeek.com/tags/>Tags</a>
|
||
<a href=https://bravenewgeek.com/about-me/>About</a>
|
||
<button id=theme-toggle class=theme-toggle accesskey=t title="Toggle theme (Alt + T)" aria-label="Toggle light/dark theme">
|
||
<svg class="moon" width="18" height="18" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M21 12.79A9 9 0 1111.21 3 7 7 0 0021 12.79z"/></svg>
|
||
<svg class="sun" width="18" height="18" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><circle cx="12" cy="12" r="5"/><line x1="12" y1="1" x2="12" y2="3"/><line x1="12" y1="21" x2="12" y2="23"/><line x1="4.22" y1="4.22" x2="5.64" y2="5.64"/><line x1="18.36" y1="18.36" x2="19.78" y2="19.78"/><line x1="1" y1="12" x2="3" y2="12"/><line x1="21" y1="12" x2="23" y2="12"/><line x1="4.22" y1="19.78" x2="5.64" y2="18.36"/><line x1="18.36" y1="5.64" x2="19.78" y2="4.22"/></svg></button></nav></div></header><main class=main><article class="post wrap"><div class=post-return><a href=https://bravenewgeek.com/><span class=pager-arrow>←</span> the log</a></div><header class=post-header><div class=post-meta><span class=post-offset>#43</span><time datetime=2015-07-21>2015-07-21</time><span>3 min read</span>
|
||
<span class=post-cats><a href=https://bravenewgeek.com/category/distributed-systems-2/>Distributed Systems</a><a href=https://bravenewgeek.com/category/software-engineering/>Software Engineering</a><a href=https://bravenewgeek.com/category/systems-theory/>Systems Theory</a></span></div><h1 class=post-title>Designed to Fail</h1></header><div class="post-content md-content"><p>When it comes to reliability engineering, people often talk about things like fault injection, monitoring, and operations runbooks. These are all critical pieces for building systems which can withstand failure, but what’s less talked about is the need to design systems which <em>deliberately</em> fail.</p><p>Reliability design has a natural progression which closely follows that of architectural design. With monolithic systems, we care more about preventing failure from occurring. With service-oriented architectures, controlling failure becomes less manageable, so instead we learn to anticipate it. With highly distributed microservice architectures where failure is all but guaranteed, we <em>embrace</em> it.</p><p>What does it mean to embrace failure? Anticipating failure is understanding the behavior when things go wrong, building systems to be resilient to it, and having a game plan for when it happens, either manual or automated. <em>Embracing</em> failure means making a conscious decision to purposely fail, and it’s essential for building highly available large-scale systems.</p><p>A microservice architecture typically means a complex web of service dependencies. One of SOA’s goals is to isolate failure and allow for graceful degradation. The key to being highly available is learning to be <em>partially</em> available. Frequently, one of the requirements for partial availability is telling the client <em>“no</em>_.”_ Outright rejecting service requests is often better than allowing them to back up because, when dealing with distributed services, the latter usually results in cascading failure across dependent systems.</p><p>While designing our distributed messaging service at Workiva, we made explicit decisions to <em>drop</em> messages on the floor if we detect the system is becoming overloaded. As queues become backed up, incoming messages are discarded, a statsd counter is incremented, and a backpressure notification is sent to the client. Upon receiving this notification, the client can respond accordingly by failing fast, exponentially backing off, or using some other flow-control strategy. By bounding resource utilization, we maintain predictable performance, predictable (and measurable) lossiness, and impede cascading failure.</p><p>Other techniques include building kill switches into service calls and routers. If an overloaded service is not essential to core business, we fail fast on calls to it to prevent availability or latency problems upstream. For example, a spam-detection service is not essential to an email system, so if it’s unavailable or overwhelmed, we can simply bypass it. Netflix’s <a href=https://github.com/Netflix/Hystrix>Hystrix</a> has a set of really nice patterns for handling this.</p><p>If we’re not careful, we can often be our own worst enemy. Many times, it’s our own internal services which <a href=http://caitiem.com/2015/06/23/clients-are-jerks-aka-how-halo-4-dosed-the-services-at-launch-how-we-survived/>cause the biggest DoS attacks on ourselves</a>. By isolating and controlling it, we can prevent failure from becoming widespread and unpredictable. By building in backpressure mechanisms and other types of intentional “failure” modes, we can ensure better availability and reliability for our systems through graceful degradation. Sometimes it’s better to fight fire with fire and failure with failure.</p></div><footer class=post-footer><ul class=post-tags><li><a href=https://bravenewgeek.com/tag/availability/>Availability</a></li><li><a href=https://bravenewgeek.com/tag/distributed-systems/>Distributed Systems</a></li><li><a href=https://bravenewgeek.com/tag/fault-tolerance/>Fault Tolerance</a></li><li><a href=https://bravenewgeek.com/tag/messaging/>Messaging</a></li><li><a href=https://bravenewgeek.com/tag/microservices/>Microservices</a></li><li><a href=https://bravenewgeek.com/tag/ops/>Ops</a></li><li><a href=https://bravenewgeek.com/tag/resilience-engineering/>Resilience Engineering</a></li><li><a href=https://bravenewgeek.com/tag/soa/>Soa</a></li><li><a href=https://bravenewgeek.com/tag/software-engineering-2/>software engineering</a></li><li><a href=https://bravenewgeek.com/tag/systems/>Systems</a></li><li><a href=https://bravenewgeek.com/tag/systems-theory/>Systems Theory</a></li></ul><nav class=post-nav aria-label="Adjacent posts"><a class=post-nav-link href=https://bravenewgeek.com/what-you-want-is-what-you-dont-understanding-trade-offs-in-distributed-messaging/><span class=post-nav-dir><span class=pager-arrow>←</span> newer</span>
|
||
<span class=post-nav-title>What You Want Is What You Don’t: Understanding Trade-Offs in Distributed Messaging</span>
|
||
</a><a class="post-nav-link post-nav-right" href=https://bravenewgeek.com/service-disoriented-architecture/><span class=post-nav-dir>older <span class=pager-arrow>→</span></span>
|
||
<span class=post-nav-title>Service-Disoriented Architecture</span></a></nav></footer></article></main><footer class=site-footer><div class="wrap site-footer-inner"><div class=footer-meta><span class=footer-copy>© 2026 Tyler Treat</span>
|
||
<span class="footer-sep footer-dot">·</span>
|
||
<span class=footer-links><a href=/feed/>rss</a>
|
||
<span class=footer-sep>·</span>
|
||
<a href=https://github.com/tylertreat target=_blank rel="noopener noreferrer me">github</a>
|
||
<span class=footer-sep>·</span>
|
||
<a href=https://www.linkedin.com/in/ttreat/ target=_blank rel="noopener noreferrer me">linkedin</a></span></div></div></footer><a href=#top id=top-link class="top-link hidden" aria-label="go to top" title="Go to Top (Alt + G)" accesskey=g><svg viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="feather feather-chevrons-up"><polyline points="17 11 12 6 7 11"/><polyline points="17 18 12 13 7 18"/></svg>
|
||
</a><script>let menu=document.getElementById("menu");if(menu){const e=localStorage.getItem("menu-scroll-position");e&&(menu.scrollLeft=parseInt(e,10)),menu.onscroll=function(){localStorage.setItem("menu-scroll-position",menu.scrollLeft)}}document.querySelectorAll('a[href^="#"]').forEach(e=>{e.addEventListener("click",function(e){e.preventDefault();var t=this.getAttribute("href").substr(1);window.matchMedia("(prefers-reduced-motion: reduce)").matches?document.querySelector(`[id='${decodeURIComponent(t)}']`).scrollIntoView():document.querySelector(`[id='${decodeURIComponent(t)}']`).scrollIntoView({behavior:"smooth"}),t==="top"?history.replaceState(null,null," "):history.pushState(null,null,`#${t}`)})})</script><script>var toplink=document.getElementById("top-link");window.onscroll=function(){const e=window.innerHeight;document.body.scrollTop>e||document.documentElement.scrollTop>e?toplink.classList.remove("hidden"):toplink.classList.add("hidden")}</script><script>document.getElementById("theme-toggle").addEventListener("click",()=>{const e=document.querySelector("html");e.dataset.theme==="dark"?(e.dataset.theme="light",localStorage.setItem("pref-theme","light")):(e.dataset.theme="dark",localStorage.setItem("pref-theme","dark"))})</script><script>document.querySelectorAll("pre > code").forEach(e=>{const n=e.parentNode.parentNode,t=document.createElement("button");t.classList.add("copy-code"),t.innerHTML="copy";function s(){t.innerHTML="copied!",setTimeout(()=>{t.innerHTML="copy"},2e3)}t.addEventListener("click",t=>{if("clipboard"in navigator){navigator.clipboard.writeText(e.textContent),s();return}const n=document.createRange();n.selectNodeContents(e);const o=window.getSelection();o.removeAllRanges(),o.addRange(n);try{document.execCommand("copy"),s()}catch{}o.removeRange(n)}),n.classList.contains("highlight")?n.appendChild(t):n.parentNode.firstChild==n||(e.parentNode.parentNode.parentNode.parentNode.parentNode.nodeName=="TABLE"?e.parentNode.parentNode.parentNode.parentNode.parentNode.appendChild(t):e.parentNode.appendChild(t))})</script></body></html> |