Files
nexus/sreweekly/articles/53/01-take-it-to-the-limit-considerations-for-building-reliable-systems.html
2026-09-12 17:23:01 +08:00

26 lines
33 KiB
HTML
Raw Blame History

This file contains invisible Unicode characters
This file contains invisible Unicode characters that are indistinguishable to humans but may be processed differently by a computer. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
<!doctype html><html lang=en dir=auto data-theme=auto><head><meta charset=utf-8><meta http-equiv=X-UA-Compatible content="IE=edge"><meta name=viewport content="width=device-width,initial-scale=1,shrink-to-fit=no"><meta name=robots content="index, follow"><title>Take It to the Limit: Considerations for Building Reliable Systems | Brave New Geek</title><meta name=keywords content="availability,distributed systems,fault tolerance,message queues,messaging,microservices,ops,resilience engineering,soa,software engineering,systems,systems theory"><meta name=description content="Complex systems usually operate in failure mode. This is because a complex system typically consists of many discrete pieces, each of which can fail in isolation (or in concert). In a microservice architecture where a given function potentially comprises several independent service calls, high availability hinges on the ability to be partially available. This is a core tenet behind resilience engineering. If a function depends on three services, each with a reliability of 90%, 95%, and 99%, respectively, partial availability could be the difference between 99.995% reliability and 84% reliability (assuming failures are independent). Resilience engineering means designing with failure as the normal."><meta name=author content><link rel=canonical href=https://bravenewgeek.com/take-it-to-the-limit-considerations-for-building-reliable-systems/><link crossorigin=anonymous href=/assets/css/stylesheet.4861a452a4c13a9a1fbf2085400b74a7de96b1beeb94dee57b4273e5dffcf337.css integrity="sha256-SGGkUqTBOpofvyCFQAt0p96Wsb7rlN7le0Jz5d/88zc=" rel="preload stylesheet" as=style><link rel=icon href=https://bravenewgeek.com/favicon.ico><link rel=icon type=image/png sizes=16x16 href=https://bravenewgeek.com/favicon.ico><link rel=icon type=image/png sizes=32x32 href=https://bravenewgeek.com/favicon.ico><link rel=apple-touch-icon href=https://bravenewgeek.com/favicon.ico><link rel=mask-icon href=https://bravenewgeek.com/favicon.ico><meta name=theme-color content="#2e2e33"><meta name=msapplication-TileColor content="#2e2e33"><link rel=alternate hreflang=en href=https://bravenewgeek.com/take-it-to-the-limit-considerations-for-building-reliable-systems/><noscript><style>#theme-toggle,.top-link{display:none}</style><style>@media(prefers-color-scheme:dark){:root{--theme:rgb(29, 30, 32);--entry:rgb(46, 46, 51);--primary:rgb(218, 218, 219);--secondary:rgb(155, 156, 157);--tertiary:rgb(65, 66, 68);--content:rgb(196, 196, 197);--code-block-bg:rgb(46, 46, 51);--code-bg:rgb(55, 56, 62);--border:rgb(51, 51, 51);color-scheme:dark}.list{background:var(--theme)}.toc{background:var(--entry)}}</style></noscript><script>localStorage.getItem("pref-theme")==="dark"?document.querySelector("html").dataset.theme="dark":localStorage.getItem("pref-theme")==="light"?document.querySelector("html").dataset.theme="light":window.matchMedia("(prefers-color-scheme: dark)").matches?document.querySelector("html").dataset.theme="dark":document.querySelector("html").dataset.theme="light"</script><link rel=preconnect href=https://fonts.googleapis.com><link rel=preconnect href=https://fonts.gstatic.com crossorigin><link rel=stylesheet href="https://fonts.googleapis.com/css2?family=JetBrains+Mono:wght@400;500&family=Source+Serif+4:ital,opsz,wght@0,8..60,400;0,8..60,600;1,8..60,400&family=Space+Grotesk:wght@500;600;700&display=swap"><meta property="og:url" content="https://bravenewgeek.com/take-it-to-the-limit-considerations-for-building-reliable-systems/"><meta property="og:site_name" content="Brave New Geek"><meta property="og:title" content="Take It to the Limit: Considerations for Building Reliable Systems"><meta property="og:description" content="Complex systems usually operate in failure mode. This is because a complex system typically consists of many discrete pieces, each of which can fail in isolation (or in concert). In a microservice architecture where a given function potentially comprises several independent service calls, high availability hinges on the ability to be partially available. This is a core tenet behind resilience engineering. If a function depends on three services, each with a reliability of 90%, 95%, and 99%, respectively, partial availability could be the difference between 99.995% reliability and 84% reliability (assuming failures are independent). Resilience engineering means designing with failure as the normal."><meta property="og:locale" content="en_us"><meta property="og:type" content="article"><meta property="article:section" content="posts"><meta property="article:published_time" content="2016-12-20T19:55:52-06:00"><meta property="article:modified_time" content="2016-12-23T12:05:50-06:00"><meta property="article:tag" content="Availability"><meta property="article:tag" content="Distributed Systems"><meta property="article:tag" content="Fault Tolerance"><meta property="article:tag" content="Message Queues"><meta property="article:tag" content="Messaging"><meta property="article:tag" content="Microservices"><meta name=twitter:card content="summary"><meta name=twitter:title content="Take It to the Limit: Considerations for Building Reliable Systems"><meta name=twitter:description content="Complex systems usually operate in failure mode. This is because a complex system typically consists of many discrete pieces, each of which can fail in isolation (or in concert). In a microservice architecture where a given function potentially comprises several independent service calls, high availability hinges on the ability to be partially available. This is a core tenet behind resilience engineering. If a function depends on three services, each with a reliability of 90%, 95%, and 99%, respectively, partial availability could be the difference between 99.995% reliability and 84% reliability (assuming failures are independent). Resilience engineering means designing with failure as the normal."><script type=application/ld+json>{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Posts","item":"https://bravenewgeek.com/posts/"},{"@type":"ListItem","position":2,"name":"Take It to the Limit: Considerations for Building Reliable Systems","item":"https://bravenewgeek.com/take-it-to-the-limit-considerations-for-building-reliable-systems/"}]}</script><script type=application/ld+json>{"@context":"https://schema.org","@type":"BlogPosting","headline":"Take It to the Limit: Considerations for Building Reliable Systems","name":"Take It to the Limit: Considerations for Building Reliable Systems","description":"Complex systems usually operate in failure mode. This is because a complex system typically consists of many discrete pieces, each of which can fail in isolation (or in concert). In a microservice architecture where a given function potentially comprises several independent service calls, high availability hinges on the ability to be partially available. This is a core tenet behind resilience engineering. If a function depends on three services, each with a reliability of 90%, 95%, and 99%, respectively, partial availability could be the difference between 99.995% reliability and 84% reliability (assuming failures are independent). Resilience engineering means designing with failure as the normal.\n","keywords":["availability","distributed systems","fault tolerance","message queues","messaging","microservices","ops","resilience engineering","soa","software engineering","systems","systems theory"],"articleBody":"Complex systems usually operate in failure mode. This is because a complex system typically consists of many discrete pieces, each of which can fail in isolation (or in concert). In a microservice architecture where a given function potentially comprises several independent service calls, high availability hinges on the ability to be partially available. This is a core tenet behind resilience engineering. If a function depends on three services, each with a reliability of 90%, 95%, and 99%, respectively, partial availability could be the difference between 99.995% reliability and 84% reliability (assuming failures are independent). Resilience engineering means designing with failure as the normal.\nAnticipating failure is the first step to resilience zen, but the second is embracing it. Telling the client “no” and failing on purpose is better than failing in unpredictable or unexpected ways. Backpressure is another critical resilience engineering pattern. Fundamentally, it’s about enforcing limits. This comes in the form of queue lengths, bandwidth throttling, traffic shaping, message rate limits, max payload sizes, etc. Prescribing these restrictions makes the limits explicit when they would otherwise be implicit (eventually your server will exhaust its memory, but since the limit is implicit, it’s unclear exactly when or what the consequences might be). Relying on unbounded queues and other implicit limits is like someone saying they know when to stop drinking because they eventually pass out.\nRate limiting is important not just to prevent bad actors from DoSing your system, but also yourself. Queue limits and message size limits are especially interesting because they seem to confuse and frustrate developers who haven’t fully internalized the motivation behind them. But really, these are just another form of rate limiting or, more generally, backpressure. Let’s look at max message size as a case study.\nImagine we have a system of distributed actors. An actor can send messages to other actors who, in turn, process the messages and may choose to send messages themselves. Now, as any good software engineer knows, the eighth fallacy of distributed computing is “the network is homogenous.” This means not all actors are using the same hardware, software, or network configuration. We have servers with 128GB RAM running Ubuntu, laptops with 16GB RAM running macOS, mobile clients with 2GB RAM running Android, IoT edge devices with 512MB RAM, and everything in between, all running a hodgepodge of software and network interfaces.\nWhen we choose not to put an upper bound on message sizes, we are making an implicit assumption (recall the discussion on implicit/explicit limits from earlier). Put another way, you and everyone you interact with (likely unknowingly) enters an unspoken contract of which neither party can opt out. This is because any actor may send a message of arbitrary size. This means any downstream consumers of this message, either directly or indirectly, must also support arbitrarily large messages.\nHow can we test something that is arbitrary? We can’t. We have two options: either we make the limit explicit or we keep this implicit, arbitrarily binding contract. The former allows us to define our operating boundaries and gives us something to test. The latter requires us to test at some undefined production-level scale. The second option is literally gambling reliability for convenience. The limit is still there, it’s just hidden. When we don’t make it explicit, we make it easy to DoS ourselves in production. Limits become even more important when dealing with cloud infrastructure due to their multitenant nature. They prevent a bad actor (or yourself) from bringing down services or dominating infrastructure and system resources.\nIn our heterogeneous actor system, we have messages bound for mobile devices and web browsers, which are often single-threaded or memory-constrained consumers. Without an explicit limit on message size, a client could easily doom itself by requesting too much data or simply receiving data outside of its control—this is why the contract is unspoken but binding.\nLet’s look at this from a different kind of engineering perspective. Consider another type of system: the US National Highway System. The US Department of Transportation uses the Federal Bridge Gross Weight Formula as a means to prevent heavy vehicles from damaging roads and bridges. It’s really the same engineering problem, just a different discipline and a different type of infrastructure.\nThe August 2007 collapse of the Interstate 35W Mississippi River bridge in Minneapolis brought renewed attention to the issue of truck weights and their relation to bridge stress. In November 2008, the National Transportation Safety Board determined there had been several reasons for the bridge’s collapse, including (but not limited to): faulty gusset plates, inadequate inspections, and the extra weight of heavy construction equipment combined with the weight of rush hour traffic.\nThe DOT relies on weigh stations to ensure trucks comply with federal weight regulations, fining those that exceed restrictions without an overweight permit.\nThe federal maximum weight is set at 80,000 pounds. Trucks exceeding the federal weight limit can still operate on the country’s highways with an overweight permit, but such permits are only issued before the scheduled trip and expire at the end of the trip. Overweight permits are only issued for loads that cannot be broken down to smaller shipments that fall below the federal weight limit, and if there is no other alternative to moving the cargo by truck.\nWeight limits need to be enforced so civil engineers have a defined operating range for the roads, bridges, and other infrastructure they build. Computers are no different. This is the reason many systems enforce these types of limits. For example, Amazon clearly publishes the limits for its Simple Queue Service—the max in-flight messages for standard queues is 120,000 messages and 20,000 messages for FIFO queues. Messages are limited to 256KB in size. Amazon Kinesis, Apache Kafka, NATS, and Google App Engine pull queues all limit messages to 1MB in size. These limits allow the system designers to optimize their infrastructure and ameliorate some of the risks of multitenancy—not to mention it makes capacity planning much easier.\nUnbounded anything—whether its queues, message sizes, queries, or traffic—is a resilience engineering anti-pattern. Without explicit limits, things fail in unexpected and unpredictable ways. Remember, the limits exist, they’re just hidden. By making them explicit, we restrict the failure domain giving us more predictability, longer mean time between failures, and shorter mean time to recovery at the cost of more upfront work or slightly more complexity.\nIt’s better to be explicit and handle these limits upfront than to punt on the problem and allow systems to fail in unexpected ways. The latter might seem like less work at first but will lead to more problems long term. By requiring developers to deal with these limitations directly, they will think through their APIs and business logic more thoroughly and design better interactions with respect to stability, scalability, and performance.\n","wordCount":"1132","inLanguage":"en","datePublished":"2016-12-20T19:55:52-06:00","dateModified":"2016-12-23T12:05:50-06:00","mainEntityOfPage":{"@type":"WebPage","@id":"https://bravenewgeek.com/take-it-to-the-limit-considerations-for-building-reliable-systems/"},"publisher":{"@type":"Organization","name":"Brave New Geek","logo":{"@type":"ImageObject","url":"https://bravenewgeek.com/favicon.ico"}}}</script></head><body id=top><header class=site-header><div class="wrap site-header-inner"><a class=wordmark href=https://bravenewgeek.com/ accesskey=h title="Brave New Geek (Alt + H)"><span class=wordmark-name>Brave New Geek</span>
<span class=wordmark-tag>Introspections of a software engineer</span></a><nav class=site-nav aria-label=Primary><a href=https://bravenewgeek.com/archive/>Archive</a>
<a href=https://bravenewgeek.com/tags/>Tags</a>
<a href=https://bravenewgeek.com/about-me/>About</a>
<button id=theme-toggle class=theme-toggle accesskey=t title="Toggle theme (Alt + T)" aria-label="Toggle light/dark theme">
<svg class="moon" width="18" height="18" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M21 12.79A9 9 0 1111.21 3 7 7 0 0021 12.79z"/></svg>
<svg class="sun" width="18" height="18" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><circle cx="12" cy="12" r="5"/><line x1="12" y1="1" x2="12" y2="3"/><line x1="12" y1="21" x2="12" y2="23"/><line x1="4.22" y1="4.22" x2="5.64" y2="5.64"/><line x1="18.36" y1="18.36" x2="19.78" y2="19.78"/><line x1="1" y1="12" x2="3" y2="12"/><line x1="21" y1="12" x2="23" y2="12"/><line x1="4.22" y1="19.78" x2="5.64" y2="18.36"/><line x1="18.36" y1="5.64" x2="19.78" y2="4.22"/></svg></button></nav></div></header><main class=main><article class="post wrap"><div class=post-return><a href=https://bravenewgeek.com/><span class=pager-arrow>&larr;</span> the log</a></div><header class=post-header><div class=post-meta><span class=post-offset>#57</span><time datetime=2016-12-20>2016-12-20</time><span>6 min read</span>
<span class=post-cats><a href=https://bravenewgeek.com/category/design-patterns/>Design Patterns</a><a href=https://bravenewgeek.com/category/distributed-systems-2/>Distributed Systems</a><a href=https://bravenewgeek.com/category/messaging/>Messaging</a><a href=https://bravenewgeek.com/category/software-engineering/>Software Engineering</a><a href=https://bravenewgeek.com/category/systems-theory/>Systems Theory</a></span></div><h1 class=post-title>Take It to the Limit: Considerations for Building Reliable Systems</h1></header><div class="post-content md-content"><p>Complex systems usually operate in failure mode. This is because a complex system typically consists of many discrete pieces, each of which can fail in isolation (or in concert). In a microservice architecture where a given function potentially comprises several independent service calls, <em>high</em> availability hinges on the ability to be <em>partially</em> available. This is a core tenet behind resilience engineering. If a function depends on three services, each with a reliability of 90%, 95%, and 99%, respectively, partial availability could be the difference between 99.995% reliability and 84% reliability (assuming failures are independent). Resilience engineering means designing with failure as the normal.</p><p>Anticipating failure is the first step to resilience zen, but the second is <a href=https://bravenewgeek.com/designed-to-fail/><em>embracing</em> it</a>. Telling the client “no” and failing on purpose is better than failing in unpredictable or unexpected ways. Backpressure is another critical resilience engineering pattern. Fundamentally, it’s about enforcing limits. This comes in the form of queue lengths, bandwidth throttling, traffic shaping, message rate limits, max payload sizes, etc. Prescribing these restrictions makes the limits explicit when they would otherwise be implicit (<em>eventually</em> your server will exhaust its memory, but since the limit is implicit, it’s unclear exactly <em>when</em> or <em>what</em> the consequences might be). Relying on unbounded queues and other implicit limits is like someone saying they know when to stop drinking because they eventually pass out.</p><p>Rate limiting is important not just to prevent bad actors from DoSing your system, but also yourself. Queue limits and message size limits are especially interesting because they seem to confuse and frustrate developers who haven’t fully internalized the motivation behind them. But really, these are just another form of rate limiting or, more generally, backpressure. Let’s look at max message size as a case study.</p><p>Imagine we have a system of distributed actors. An actor can send messages to other actors who, in turn, process the messages and may choose to send messages themselves. Now, as any good software engineer knows, the <a href=https://en.wikipedia.org/wiki/Fallacies_of_distributed_computing>eighth fallacy of distributed computing</a> is “the network is homogenous.” This means not all actors are using the same hardware, software, or network configuration. We have servers with 128GB RAM running Ubuntu, laptops with 16GB RAM running macOS, mobile clients with 2GB RAM running Android, IoT edge devices with 512MB RAM, and everything in between, all running a hodgepodge of software and network interfaces.</p><p>When we choose <em>not</em> to put an upper bound on message sizes, we are making an implicit assumption (recall the discussion on implicit/explicit limits from earlier). Put another way, you and everyone you interact with (likely unknowingly) enters an unspoken contract of which neither party can opt out. This is because any actor may send a message of arbitrary size. This means any downstream consumers of this message, either directly or indirectly, must also support arbitrarily large messages.</p><p>How can we test something that is arbitrary? We can’t. We have two options: either we make the limit explicit or we keep this implicit, arbitrarily binding contract. The former allows us to define our operating boundaries and gives us something to test. The latter requires us to test at some undefined production-level scale. The second option is literally gambling reliability for convenience. The limit is still there, it’s just hidden. When we don’t make it explicit, we make it easy to DoS ourselves in production. Limits become even more important when dealing with cloud infrastructure due to their multitenant nature. They prevent a bad actor (or yourself) from bringing down services or dominating infrastructure and system resources.</p><p>In our heterogeneous actor system, we have messages bound for mobile devices and web browsers, which are often single-threaded or memory-constrained consumers. Without an explicit limit on message size, a client could easily doom itself by requesting too much data or simply receiving data outside of its control—this is why the contract is unspoken but binding.</p><p>Let’s look at this from a different kind of engineering perspective. Consider another type of system: the US National Highway System. The US Department of Transportation uses the <a href=https://en.wikipedia.org/wiki/Federal_Bridge_Gross_Weight_Formula>Federal Bridge Gross Weight Formula</a> as a means to prevent heavy vehicles from damaging roads and bridges. It’s really the same engineering problem, just a different discipline and a different type of infrastructure.</p><blockquote><p>The August 2007 collapse of the Interstate 35W Mississippi River bridge in Minneapolis brought renewed attention to the issue of truck weights and their relation to bridge stress. In November 2008, the National Transportation Safety Board determined there had been several reasons for the bridge’s collapse, including (but not limited to): faulty gusset plates, inadequate inspections, and the extra weight of heavy construction equipment combined with the weight of rush hour traffic.</p></blockquote><p>The DOT relies on <a href=https://en.wikipedia.org/wiki/Weigh_station>weigh stations</a> to ensure trucks comply with federal weight regulations, fining those that exceed restrictions without an overweight permit.</p><blockquote><p>The federal maximum weight is set at 80,000 pounds. Trucks exceeding the federal weight limit can still operate on the country’s highways with an overweight permit, but such permits are only issued before the scheduled trip and expire at the end of the trip. Overweight permits are only issued for loads that cannot be broken down to smaller shipments that fall below the federal weight limit, and if there is no other alternative to moving the cargo by truck.</p></blockquote><p>Weight limits need to be enforced so civil engineers have a defined operating range for the roads, bridges, and other infrastructure they build. Computers are no different. This is the reason many systems enforce these types of limits. For example, Amazon clearly publishes the <a href=http://docs.aws.amazon.com/AWSSimpleQueueService/latest/SQSDeveloperGuide/sqs-limits.html>limits for its Simple Queue Service</a>—the max in-flight messages for standard queues is 120,000 messages and 20,000 messages for FIFO queues. Messages are limited to 256KB in size. <a href=http://docs.aws.amazon.com/streams/latest/dev/service-sizes-and-limits.html>Amazon Kinesis</a>, <a href=https://kafka.apache.org/documentation/>Apache Kafka</a>, <a href=https://nats.io/documentation/faq/#msgsize>NATS</a>, and <a href=https://cloud.google.com/appengine/docs/quotas#Task_Queue>Google App Engine pull queues</a> all limit messages to 1MB in size. These limits allow the system designers to optimize their infrastructure and ameliorate some of the risks of multitenancy—not to mention it makes capacity planning <em>much</em> easier.</p><p>Unbounded <em>anything</em>—whether its queues, message sizes, queries, or traffic—is a resilience engineering anti-pattern. Without explicit limits, things fail in unexpected and unpredictable ways. Remember, the limits exist, they’re just hidden. By making them explicit, we restrict the failure domain giving us more predictability, longer mean time between failures, and shorter mean time to recovery at the cost of more upfront work or slightly more complexity.</p><p>It’s better to be explicit and handle these limits upfront than to punt on the problem and allow systems to fail in unexpected ways. The latter might seem like less work at first but will lead to more problems long term. By requiring developers to deal with these limitations directly, they will think through their APIs and business logic more thoroughly and design better interactions with respect to stability, scalability, and performance.</p></div><footer class=post-footer><ul class=post-tags><li><a href=https://bravenewgeek.com/tag/availability/>Availability</a></li><li><a href=https://bravenewgeek.com/tag/distributed-systems/>Distributed Systems</a></li><li><a href=https://bravenewgeek.com/tag/fault-tolerance/>Fault Tolerance</a></li><li><a href=https://bravenewgeek.com/tag/message-queues/>Message Queues</a></li><li><a href=https://bravenewgeek.com/tag/messaging/>Messaging</a></li><li><a href=https://bravenewgeek.com/tag/microservices/>Microservices</a></li><li><a href=https://bravenewgeek.com/tag/ops/>Ops</a></li><li><a href=https://bravenewgeek.com/tag/resilience-engineering/>Resilience Engineering</a></li><li><a href=https://bravenewgeek.com/tag/soa/>Soa</a></li><li><a href=https://bravenewgeek.com/tag/software-engineering-2/>software engineering</a></li><li><a href=https://bravenewgeek.com/tag/systems/>Systems</a></li><li><a href=https://bravenewgeek.com/tag/systems-theory/>Systems Theory</a></li></ul><nav class=post-nav aria-label="Adjacent posts"><a class=post-nav-link href=https://bravenewgeek.com/fast-topic-matching/><span class=post-nav-dir><span class=pager-arrow>&larr;</span> newer</span>
<span class=post-nav-title>Fast Topic Matching</span>
</a><a class="post-nav-link post-nav-right" href=https://bravenewgeek.com/benchmarking-commit-logs/><span class=post-nav-dir>older <span class=pager-arrow>&rarr;</span></span>
<span class=post-nav-title>Benchmarking Commit Logs</span></a></nav></footer><section class=wp-comments><h2>Comments</h2><p class=wp-comments-notice>Comments are from this blog's WordPress era and are preserved read-only.</p><article class=wp-comment><header><span class=wp-comment-author>Timothy Pote</span>
<time class=wp-comment-date>December 22, 2016</time></header><div class=wp-comment-body><p>Spot on!</p><p>I have a tiny correction for you: SQS has no limits on queue depth. The numbers you mention are limits on the in-flight messages you&#8217;re permitted, but you can have queues of arbitrary length.</p><p>While that might seem to violate the principles in this post, I would assume that this guarantee is a result of the host of other restrictions they put on the system (message size, in-flight limit, limited retention, etc).</p><p>Source: <a href=https://aws.amazon.com/sqs/faqs/ rel="nofollow ugc">https://aws.amazon.com/sqs/faqs/</a> (search for &#8220;How large can Amazon SQS message queues be?&#8221;)</p></div><div class=wp-comment-replies><article class=wp-comment><header><span class=wp-comment-author>Tyler Treat</span>
<time class=wp-comment-date>December 22, 2016</time></header><div class=wp-comment-body><p>You&#8217;re right, thanks for pointing that out, Timothy.</p></div></article></div></article><article class=wp-comment><header><span class=wp-comment-author>Daniel J</span>
<time class=wp-comment-date>December 22, 2016</time></header><div class=wp-comment-body><p>Thanks for a great post, I think it really drives the point home.</p><p>A minor point, I don&#8217;t know if I misunderstand your meaning when talking about SQS, but the limits they speak of are in-flight messages, messages that have been read but not yet deleted. The queues are theoretically unbounded.</p></div><div class=wp-comment-replies><article class=wp-comment><header><span class=wp-comment-author>Daniel J</span>
<time class=wp-comment-date>December 22, 2016</time></header><div class=wp-comment-body><p>Hmm reading the post on a plane (in-flight!) apparently meant I didn&#8217;t see the comments. The point was already made :D</p></div></article></div></article><article class=wp-comment><header><span class=wp-comment-author>valentin</span>
<time class=wp-comment-date>December 26, 2016</time></header><div class=wp-comment-body><p>Wait, doesn&#8217;t embrace aka LetItCrush principle mean that we do not enforce the graceful shutdown cases and limits? We just let it crush when it goes wrong naturally instead of complicating our code with graceful termination?</p></div></article><article class=wp-comment><header><span class=wp-comment-author>David Leonhartsberger</span>
<time class=wp-comment-date>December 27, 2016</time></header><div class=wp-comment-body><p>Agree very much</p></div></article><article class=wp-comment><header><span class=wp-comment-author>Jeroen Pluimers</span>
<time class=wp-comment-date>February 18, 2017</time></header><div class=wp-comment-body><p>I think a lot of systems don&#8217;t have these limits defined as people have a hard time coming up with them. So it would be very helpful to have some guidance on finding and designing for these limits.</p></div></article><article class=wp-comment><header><span class=wp-comment-author>Mr. Opa Opa</span>
<time class=wp-comment-date>March 8, 2019</time></header><div class=wp-comment-body><p>Please, could you explain the following numbers? How is 99.995% and 84% related to all reliability of three services mentioned?</p><p>&#8220;This is a core tenet behind resilience engineering. If a function depends on three services, each with a reliability of 90%, 95%, and 99%, respectively, partial availability could be the difference between 99.995% reliability and 84% reliability (assuming failures are independent).&#8221;</p></div></article></section></article></main><footer class=site-footer><div class="wrap site-footer-inner"><div class=footer-meta><span class=footer-copy>&copy; 2026 Tyler Treat</span>
<span class="footer-sep footer-dot">·</span>
<span class=footer-links><a href=/feed/>rss</a>
<span class=footer-sep>·</span>
<a href=https://github.com/tylertreat target=_blank rel="noopener noreferrer me">github</a>
<span class=footer-sep>·</span>
<a href=https://www.linkedin.com/in/ttreat/ target=_blank rel="noopener noreferrer me">linkedin</a></span></div></div></footer><a href=#top id=top-link class="top-link hidden" aria-label="go to top" title="Go to Top (Alt + G)" accesskey=g><svg viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="feather feather-chevrons-up"><polyline points="17 11 12 6 7 11"/><polyline points="17 18 12 13 7 18"/></svg>
</a><script>let menu=document.getElementById("menu");if(menu){const e=localStorage.getItem("menu-scroll-position");e&&(menu.scrollLeft=parseInt(e,10)),menu.onscroll=function(){localStorage.setItem("menu-scroll-position",menu.scrollLeft)}}document.querySelectorAll('a[href^="#"]').forEach(e=>{e.addEventListener("click",function(e){e.preventDefault();var t=this.getAttribute("href").substr(1);window.matchMedia("(prefers-reduced-motion: reduce)").matches?document.querySelector(`[id='${decodeURIComponent(t)}']`).scrollIntoView():document.querySelector(`[id='${decodeURIComponent(t)}']`).scrollIntoView({behavior:"smooth"}),t==="top"?history.replaceState(null,null," "):history.pushState(null,null,`#${t}`)})})</script><script>var toplink=document.getElementById("top-link");window.onscroll=function(){const e=window.innerHeight;document.body.scrollTop>e||document.documentElement.scrollTop>e?toplink.classList.remove("hidden"):toplink.classList.add("hidden")}</script><script>document.getElementById("theme-toggle").addEventListener("click",()=>{const e=document.querySelector("html");e.dataset.theme==="dark"?(e.dataset.theme="light",localStorage.setItem("pref-theme","light")):(e.dataset.theme="dark",localStorage.setItem("pref-theme","dark"))})</script><script>document.querySelectorAll("pre > code").forEach(e=>{const n=e.parentNode.parentNode,t=document.createElement("button");t.classList.add("copy-code"),t.innerHTML="copy";function s(){t.innerHTML="copied!",setTimeout(()=>{t.innerHTML="copy"},2e3)}t.addEventListener("click",t=>{if("clipboard"in navigator){navigator.clipboard.writeText(e.textContent),s();return}const n=document.createRange();n.selectNodeContents(e);const o=window.getSelection();o.removeAllRanges(),o.addRange(n);try{document.execCommand("copy"),s()}catch{}o.removeRange(n)}),n.classList.contains("highlight")?n.appendChild(t):n.parentNode.firstChild==n||(e.parentNode.parentNode.parentNode.parentNode.parentNode.nodeName=="TABLE"?e.parentNode.parentNode.parentNode.parentNode.parentNode.appendChild(t):e.parentNode.appendChild(t))})</script></body></html>