Files
nexus/sreweekly/articles/502/06-what-now-handling-errors-in-large-systems.html
2026-09-12 17:23:01 +08:00

1446 lines
27 KiB
HTML
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
<!DOCTYPE html>
<html xmlns="http://www.w3.org/1999/xhtml" xml:lang="en" lang="en-us">
<head>
<meta http-equiv="content-type" content="text/html; charset=utf-8" />
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<title>What Now? Handling Errors in Large Systems - Marc's Blog</title>
<meta name="author" content="Marc Brooker" />
<!-- Homepage CSS -->
<link rel="stylesheet" href="/blog/css/screen.css" type="text/css" media="screen, projection" />
<link rel="stylesheet" href="/blog/css/syntax.css" type="text/css" media="screen, projection" />
</head>
<body>
<div class="site">
<div class="title">
<h1><a href="/blog/">Marc's Blog</a></h1>
</div>
<div class="about">
<h1>About Me</h1>
My name is Marc Brooker. I like to build things that work, and do cool stuff. I like building big things. I also dabble in machining, welding, cooking, and skiing.<br/><br/>
I am an engineer at Amazon Web Services (AWS) in Seattle, where I work on agentic AI, especially safety and policy for agentic AI. Before that, I worked on EC2, EBS, databases, serverless, and serverless databases.<br/>
All opinions are my own.
<h1>Links</h1>
<a href="https://brooker.co.za/blog/publications.html">My Publications and Videos</a><br/>
<a rel="me" href="https://fediscience.org/@marcbrooker">@marcbrooker on Mastodon</a>
<a href="https://twitter.com/MarcJBrooker">@MarcJBrooker on Twitter</a>
<br/><br/><br/>
<a href="https://brooker.co.za/blog/2026/06/18/my-blog-and-ai.html">Is this blog written by AI?</a>
</div>
<div id="post">
<h1 id="what-now-handling-errors-in-large-systems">What Now? Handling Errors in Large Systems</h1>
<script>
MathJax = {
tex: {inlineMath: [['$', '$'], ['\\(', '\\)']]}
};
</script>
<script id="MathJax-script" async="" src="https://cdn.jsdelivr.net/npm/mathjax@3/es5/tex-mml-chtml.js"></script>
<script>
function vote(btn, choice) {
const widget = btn.closest('.vote-widget');
widget.querySelector('.user-vote').textContent = choice === 'tick' ? '✅' : '❌';
widget.querySelector('.result').style.display = 'block';
widget.querySelectorAll('button').forEach(b => b.style.display = 'none');
}
function showAllAnswers() {
document.querySelectorAll('.vote-widget .result').forEach(result => {
result.style.display = 'block';
});
document.querySelectorAll('.vote-widget button').forEach(btn => {
if (btn.onclick && btn.onclick.toString().includes('vote(')) {
btn.style.display = 'none';
}
});
}
</script>
<style>
.justification {
font-style: italic;
color: #444;
}
</style>
<p class="meta">More options means more choices.</p>
<p>Cloudflare’s deep <a href="https://blog.cloudflare.com/18-november-2025-outage/">postmortem for their November 18 outage</a> triggered a ton of online chatter about error handling, caused by a single line in the postmortem:</p>
<figure class="highlight"><pre><code class="language-rust" data-lang="rust"><span class="nf">.unwrap</span><span class="p">()</span></code></pre></figure>
<p>If you’re not familiar with Rust, you need to know about <a href="https://doc.rust-lang.org/std/result/enum.Result.html#method.unwrap">Result</a>, a kind of struct that can contain either a successful result, or an error. <code class="language-plaintext highlighter-rouge">unwrap</code> says basically “return the successful results if there is one, otherwise crash the program”<sup><a href="#foot1">1</a></sup>. You can think of it like an <code class="language-plaintext highlighter-rouge">assert</code>.</p>
<p>There’s a ton of debate about whether <code class="language-plaintext highlighter-rouge">assert</code>s are good in production<sup><a href="#foot2">2</a></sup>, but most are missing the point. Quite simply, this isn’t a question about a single program. It’s not a local property. Whether <code class="language-plaintext highlighter-rouge">assert</code>s are appropriate for a given component is a global property of the system, and the way it handles data.</p>
<p>Let’s play a little error handling game. Click the ✅ if you think crashing the process or server is appropriate, and the ❌ if you don’t. Then you’ll see my vote and justification.</p>
<ul>
<li> One of ten web servers behind a load balancer encounters uncorrectable memory errors, and takes itself out of service. <div class="vote-widget">
<button onclick="vote(this, 'tick')">✅</button>
<button onclick="vote(this, 'cross')">❌</button>
<div class="result" style="display:none;">
<p>Your vote: <span class="user-vote"></span></p>
<p>My vote: <span class="my-vote">✅</span></p>
<p class="justification">Uncorrectable memory errors are independent, and do not depend on user-provided content. In the presence of bad memory, it's impossible for a program to proceed safely. Taking the machine out of service is the safest course of action.</p>
</div>
</div>
</li>
<li> One of ten multi-threaded application servers behind a load balancer encounters a null pointer in business logic while processing a customer request. <div class="vote-widget">
<button onclick="vote(this, 'tick')">✅</button>
<button onclick="vote(this, 'cross')">❌</button>
<div class="result" style="display:none;">
<p>Your vote: <span class="user-vote"></span></p>
<p>My vote: <span class="my-vote">❌</span></p>
<p class="justification">Customer requests triggering bugs in business logic isn't a good reason to bring the whole server down. Instead, fail that particular request (returning an HTTP 5xx error), and continue with other user requests. In approaches like Erlang, or even Lambda, it may be the right approach to crash the whole application in response to a bad request, because this crash is handled at a higher layer in the architecture. This is also why I prefer languages like Rust and Java to languages like C and C++ for services: the ability to continue after a `NullPointerException` or getting a `Option::None` is much better than the ability to continue after (say) a segfault. It is possible to write C that's safe in all the same cases, but the explicit handling of errors in Rust make it much easier.</p>
</div>
</div>
</li>
<li> One database replica receives a logical replication record from the primary that it doesn't know how to process <div class="vote-widget">
<button onclick="vote(this, 'tick')">✅</button>
<button onclick="vote(this, 'cross')">❌</button>
<div class="result" style="display:none;">
<p>Your vote: <span class="user-vote"></span></p>
<p>My vote: <span class="my-vote">✅</span></p>
<p class="justification">In general, replicas in this position can't continue, because applying future updates can cause arbitrary state corruption, and return arbitrarily wrong results to clients. For a successful system, ensuring the primary doesn't send bad records to replicas must be a system invariant.</p>
</div>
</div>
</li>
<li> One web server receives a global configuration file from the control plane that appears malformed. <div class="vote-widget">
<button onclick="vote(this, 'tick')">✅</button>
<button onclick="vote(this, 'cross')">❌</button>
<div class="result" style="display:none;">
<p>Your vote: <span class="user-vote"></span></p>
<p>My vote: <span class="my-vote">❌</span></p>
<p class="justification">The right answer here will vary based on the needs of the system, but in most systems the best design would be for the server to continue with the last known good version of configuration, while alerting an operator that the latest version can't be processed. This is subtly different from the previous case: configuration doesn't tend to have the same consistency requirements as state, and tends to be entirely replaced with each new version, and so treating configuration currency as an invariant of the system reduces resilience unnecessarily.</p>
</div>
</div>
</li>
<li> One web server fails to write its log file because of a full disk. <div class="vote-widget">
<button onclick="vote(this, 'tick')">✅</button>
<button onclick="vote(this, 'cross')">❌</button>
<div class="result" style="display:none;">
<p>Your vote: <span class="user-vote"></span></p>
<p>My vote: <span class="my-vote">❌</span></p>
<p class="justification">It may seem like this is an uncorrelated condition, and it could be. The local log rotation agent could have crashed, for example. But it also could be because of a global condition, like a prior deployment of a bad log rotation configuration, or ongoing load spike. Unless there are specific requirements (e.g. legal requirements) for log retention, it's likely best to continue and inform an operator.</p>
</div>
</div>
</li>
</ul>
<p>If you don’t want to play, and just see my answers, click here: <button onclick="showAllAnswers()">Show All Answers</button>.</p>
<p>There are three unifying principles behind my answers here.</p>
<p><strong>Are failures correlated?</strong> If the decision is a local one that’s highly likely to be uncorrelated between machines, then crashing is the cleanest thing to do. Crashing has the advantage of reducing the complexity of the system, by removing the <em>working in degraded mode</em> state. On the other hand, if failures can be correlated (including by adversarial user behavior), its best to design the system to reject the cause of the errors and continue.</p>
<p><strong>Can they be handled at a higher layer?</strong> This is where you need to understand your architecture. Traditional web service architectures can handle low rates of errors at a higher layer (e.g. by replacing instances or containers as they fail load balancer health checks using <a href="https://aws.amazon.com/autoscaling/">AWS Autoscaling</a>), but can’t handle high rates of crashes (because they are limited in how quickly instances or containers can be replaced). Fine-grained architectures, starting with Lambda-style serverless all the way to Erlang’s approach, are designed to handle higher rates of errors, and crashing rather the continuing is appropriate in more cases.</p>
<p><strong>Is it possible to meaningfully continue?</strong> This is where you need to understand your business logic. In most cases with configuration, and some cases with data, its possible to continue with the last-known good version. This adds complexity, by introducing the behavior mode of running with that version, but that complexity may be worth the additional resilience. On the other hand, in a database that handles updates via operations (e.g. <code class="language-plaintext highlighter-rouge">x = x + 1</code>) or conditional operations (<code class="language-plaintext highlighter-rouge">if x == 1 then y = y + x</code>) then continuing after skipping some records could cause arbitrary state corruption. In the latter case, the system must be designed (including its operational practices) to ensure the invariant that replicas only get records they understand. These kinds of invariants make the system less resilient, but are needed to avoid state divergence.</p>
<p>The bottom line is that error handling in systems isn’t a local property. The right way to handle errors is a global property of the system, and error handling needs to be built into the system from the beginning.</p>
<p>Getting this right is hard, and that’s where blast radius reduction techniques like cell-based architectures, independent regions, and <a href="https://aws.amazon.com/blogs/architecture/shuffle-sharding-massive-and-magical-fault-isolation/">shuffle sharding</a> come in. Blast radius reduction means that if you do the wrong thing you affect less than all your traffic - ideally a small percentage of traffic. Blast radius reduction is humility in the face of complexity.</p>
<p><em>Footnotes</em></p>
<ol>
<li><a name="foot1"></a> Yes, I know a <code class="language-plaintext highlighter-rouge">panic</code> <a href="https://doc.rust-lang.org/book/ch09-01-unrecoverable-errors-with-panic.html">isn’t necessarily a crash</a>, but it’s close enough for our purposes here. If you’d like to explain the difference to me, feel free.</li>
<li><a name="foot2"></a> And a ton of debate about whether Rust helped here. I think Rust does two things very well in this case: it makes the <code class="language-plaintext highlighter-rouge">unwrap</code> case explicit in the code (the programmer can see that this line has “succeed or die behavior”, entirely locally on this one line of code), and prevents action-at-a-distance behavior (which silently continuing with a <code class="language-plaintext highlighter-rouge">NULL</code> pointer could cause). What Rust doesn’t do perfectly here is make this explicit enough. Some suggested that <code class="language-plaintext highlighter-rouge">unwrap</code> should be called <code class="language-plaintext highlighter-rouge">or_panic</code>, which I like. Others suggested lints like <code class="language-plaintext highlighter-rouge">clippy</code> should be more explicit about requiring <code class="language-plaintext highlighter-rouge">unwrap</code> to come with some justification, which may be helpful in some code bases. Overall, I’d rather be writing Rust than C here.</li>
</ol>
</div>
<div id="related">
&laquo; <a href="/blog">Back to the blog index</a><br>
<br>
<!-- Similar Posts -->
<h4>Similar Posts</h4>
<ul class="posts">
<li><span>24 May 2021</span> &raquo; <a href="/blog/2021/05/24/metastable.html">Metastability and Distributed Systems</a></li>
<li><span>16 Feb 2022</span> &raquo; <a href="/blog/2022/02/16/circuit-breakers.html">Will circuit breakers solve my problems?</a></li>
<li><span>03 Jan 2016</span> &raquo; <a href="/blog/2016/01/03/correlation.html">Why Must Systems Be Operated?</a></li>
</ul>
<!-- Dissimilar Posts -->
<h4>Something Completely Different</h4>
<ul class="posts">
<li><span>28 Jul 2020</span> &raquo; <a href="/blog/2020/07/28/fish.html">A Story About a Fish</a></li>
</ul>
</div>
<div class="footer">
<div class="contact">
<p>
Marc Brooker<br />
The opinions on this site are my own. They do not necessarily represent those of my employer.<br />
marcbrooker@gmail.com
</p>
<p>
<a href="https://brooker.co.za/blog/rss.xml"><img src="/blog/images/feed-icon-14x14.png" /> RSS</a>
<a href="https://brooker.co.za/blog/atom.xml"><img src="/blog/images/feed-icon-14x14.png" /> Atom</a>
</p>
</div>
<div class="license">
<!-- <a rel="license" href="http://creativecommons.org/licenses/by/4.0/"><img alt="Creative Commons License" style="border-width:0" src="https://i.creativecommons.org/l/by/4.0/88x31.png" /></a><br /> -->
This work is licensed under a <a rel="license" href="http://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</a>.
</div>
</div>
</div>
</body>
</html>