Files
nexus/sreweekly/articles/260/02-increment-reliability-the-process-implementing-yelp-s-failover-strateg.html
2026-09-12 17:23:01 +08:00

19 lines
37 KiB
HTML
Raw Blame History

This file contains invisible Unicode characters
This file contains invisible Unicode characters that are indistinguishable to humans but may be processed differently by a computer. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
<!doctype html><html><head><meta charset=utf-8><title>Implementing Yelp’s Traffic Failover Strategy - Increment</title><meta name=description content='Learn how Yelp site reliability engineers implemented a traffic failover strategy, finding a balance between reliability, performance, and cost efficiency.'><link rel=canonical href=http://localhost:3000/reliability/yelp-traffic-failover-strategy/ ><link rel=apple-touch-icon-precomposed href=/img/icon-571805a1.png><meta property=og:title content='The process: Implementing Yelp’s failover strategy – Increment: Reliability'><meta property=og:url content=http://localhost:3000/reliability/yelp-traffic-failover-strategy/ ><meta property=og:description content='How Yelp engineers orchestrated their traffic failover process and effected a delicate balance between reliability, performance, and cost efficiency.'><meta property=og:image content='https://images.ctfassets.net/3njn2qm7rrbs/6Q3D0w8nis0u66FQeCTjbg/843b26ce0013258589c76a15a49461d4/cover-issue16.png?w=1000'><meta name=twitter:card content=summary_large_image><meta name=twitter:image content='https://images.ctfassets.net/3njn2qm7rrbs/6Q3D0w8nis0u66FQeCTjbg/843b26ce0013258589c76a15a49461d4/cover-issue16.png?w=1000'><meta name=twitter:site content=@IncrementMag><meta name=twitter:title content='The process: Implementing Yelp’s failover strategy – Increment: Reliability'><meta name=twitter:description content='How Yelp engineers orchestrated their traffic failover process and effected a delicate balance between reliability, performance, and cost efficiency.'><link rel=alternate type=application/rss+xml title=Increment href=/feed.xml><meta name=viewport content='width=device-width,initial-scale=1'><link rel=preload href=/fonts/baton-turbo/400-30a55d66.woff2 as=font type=font/woff2 crossorigin=anonymous><link rel=preload href=/fonts/baton-turbo/500-1603c0e8.woff2 as=font type=font/woff2 crossorigin=anonymous><link rel=preload href=/fonts/tiempos-text/400-c4810745.woff2 as=font type=font/woff2 crossorigin=anonymous><link rel=preload href=/fonts/tiempos-head/700-383ede62.woff2 as=font type=font/woff2 crossorigin=anonymous><link rel=stylesheet type=text/css href=/css/bundle-289d885f.css><link rel=stylesheet type=text/css href=/css/issues/16-c2b7b183.css><script>// Don't fade art if it loads ~instantly
setTimeout(()=>{document.documentElement.classList.add('fadeArt')},250);</script><script defer src=/js/defer-737fbd90.js></script><script>const INCREMENT_META={issueNumber:16,issueSlug:'reliability',articleSlug:'yelp-traffic-failover-strategy'};</script></head><body class='Issue_reliability Article_yelp-traffic-failover-strategy'><nav class=PageNav><div class=u-Container><div class=column><h1 class=logo><a href=/ ><img src=/img/logo-ae2c55d5.svg alt=Increment></a></h1><a class=out-now href=https://store.increment.com/ style=color:#595959><div class='IssueTitle tiny'><div class='t-Caps meta tiny'><span>NEW</span></div><h3 class='t-IssueTitle title'>Buy the print edition</h3></div></a><ul class=nav><li><a href=/issues/ ><span>Issues</span></a></li><li><a href=/topics/ ><span>Topics</span></a></li><li><a href=https://store.increment.com/ ><span>Store</span></a></li><li><a href=/about/ ><span>About</span></a></li></ul></div></div></nav><div class=ArticlePage itemscope itemtype=http://schema.org/Article><script type=application/ld+json>{
"@context": "http://schema.org",
"@type": "Article",
"headline": "The process: Implementing Yelp’s failover strategy",
"image": " https://images.ctfassets.net/3njn2qm7rrbs/6Q3D0w8nis0u66FQeCTjbg/843b26ce0013258589c76a15a49461d4/cover-issue16.png?w&#x3D;1000",
"datePublished": "Thu, 25 Feb 2021 19:00:00 GMT",
"dateModified": "Thu, 25 Feb 2021 19:00:00 GMT",
"publisher": {
"@type": "Organization",
"name": "Increment",
"logo": {
"@type": "ImageObject",
"url": "https://increment.com/img/logo.png"
}
},
"description": "How Yelp engineers orchestrated their traffic failover process and effected a delicate balance between reliability, performance, and cost efficiency.",
"mainEntityOfPage": "http://localhost:3000/reliability/yelp-traffic-failover-strategy/"
}</script><header class='u-Container ArticleHeader'><div class=column><div class=u-Grid><div class=main><h4 class='t-Byline byline large'><a href=#authors class=j-SmoothScroll itemprop=author><span>Mathieu Frappier</span>, <span>Dorothy Jung</span>, and <span>Qui Nguyen</span></a></h4><h1 class='t-TitleSerif large title' itemprop=name>The process: Implementing Yelp’s failover strategy</h1><div class='t-BodySans large intro' itemprop=description>How Yelp engineers orchestrated their traffic failover process and effected a delicate balance between reliability, performance, and cost efficiency.</div></div><a class=issue href=/reliability/ ><span class='t-Caps tiny part-of'>Part of</span><div class='IssueTitle small' style=color:#863051><div class='t-Caps meta'><span>Issue 16</span> <span>February 2021</span></div><h2 class='t-IssueTitle title'>Reliability</h2></div></a></div></div></header><div class='u-Container ArticleContent'><article class='ContentBody column' itemprop=articleBody><div class=ArticleLayout><p>On the surface, the process is straightforward: Site reliability engineers at Yelp sometimes shift traffic to prevent user-facing errors. Under the hood, however, it involves a complex choreography between production systems, infrastructure teams, and hundreds of developers and their services. This is the story of how Yelp’s production engineering and compute infrastructure teams implemented a failover strategy by finding a balance between reliability, performance, and cost efficiency.</p></div><div class='ArticleLayout flipped'><h2>What’s a traffic failover?</h2><p>Yelp serves requests out of two regional AWS data centers located on both U.S. coasts. Read-only requests, which constitute the bulk of user traffic, are sent to the nearest data center, with additional logic to ensure the load is evenly distributed between both regions. Sometimes, one region becomes unhealthy due to a bad infrastructure configuration, an impaired critical data store, or, on rare occasions, an AWS issue. When any of these happen, we could be serving users HTTP 500 errors and need to act quickly.</p><p>To mitigate such outages, one tool at Yelp’s disposal is the failover: the ability to quickly shift traffic from the unhealthy region to the healthy one. A partial traffic shift alleviates pressure on impaired systems and allows them to recover. The shift can also be total: a full failover. All we need to do to shift traffic is update a Git-controlled YAML file. But even during an emergency, merging and pushing the change requires approvals, typically from the secondary on-call production engineer, a manager, or an engineer involved in the ongoing incident.</p><p class=image><picture><source srcset='https://images.ctfassets.net/3njn2qm7rrbs/tc4nFsnfvc8vDC4y9LwD5/93ee959fe49fa1c62710a6a3dfba3970/image1-2000-2739b88d.jpeg?w=2000 2000w, https://images.ctfassets.net/3njn2qm7rrbs/tc4nFsnfvc8vDC4y9LwD5/93ee959fe49fa1c62710a6a3dfba3970/image1-2000-2739b88d.jpeg?w=1000 1000w, https://images.ctfassets.net/3njn2qm7rrbs/tc4nFsnfvc8vDC4y9LwD5/93ee959fe49fa1c62710a6a3dfba3970/image1-2000-2739b88d.jpeg?w=500 500w, https://images.ctfassets.net/3njn2qm7rrbs/tc4nFsnfvc8vDC4y9LwD5/93ee959fe49fa1c62710a6a3dfba3970/image1-2000-2739b88d.jpeg?w=100 100w' sizes=750px><img class=j-SmoothLoad src='https://images.ctfassets.net/3njn2qm7rrbs/tc4nFsnfvc8vDC4y9LwD5/93ee959fe49fa1c62710a6a3dfba3970/image1-2000-2739b88d.jpeg?w=1000' sizes=750px width=750px alt='<i>An extract from the traffic management configuration file</i>'></picture><span class=Caption><i>An extract from the traffic management configuration file</i></span></p><p>On-call engineers at Yelp regularly practice partial and full failovers to ensure our infrastructure can handle the sudden change in load and our teams remain comfortable executing the procedure. While the failover itself is a simple config change, situations that require full failovers are often stressful and unpredictable. The primary on-call engineer needs to be familiar with the process to avoid additional strain.</p><h2>When failovers fail</h2><p>Major shifts in traffic patterns can overwhelm the healthy region that’s now serving global traffic. In Yelp’s early days, we “melted” a healthy region on numerous occasions by sending too much traffic to one region too quickly. Most of our services and clusters of machines can scale up in minutes, assuming all systems are working as they should, but these are crucial minutes we can’t spare. Our response needs to be instant.</p><p>Furthermore, in a constantly evolving infrastructure, adding capacity to production can be complicated—a recent change in low-level configuration could slow or even prevent us from getting new, healthy machines. This can quickly devolve into a worst-case scenario where we’re unable to scale up the healthy region and end up serving HTTP 500 errors to users.</p></div><div class=ArticleLayout><h2>Keeping double capacity</h2><p>A good way to prevent meltdowns is to keep extra compute capacity around at all times.</p><p>One way to do this is to have more machines available. By doubling the number of running machines, we always have the compute capacity we need to handle failovers. This also means we don’t need to add machines in an emergency, which removes one step in the failover process and, more importantly, reduces dependency on the compute infrastructure team if something goes wrong provisioning these instances.</p><p>However, keeping idle machines around just in case of a failover can seem like a waste of resources, so we put them to work by distributing the containers we need among all available machines. This way, each machine has just 50 percent of its resources allocated to services in normal situations, allowing it to absorb load spikes and maintain more consistent performance—and it costs the same.</p><p class=image><picture><source srcset='https://images.ctfassets.net/3njn2qm7rrbs/5HxFhiCsA6FrFhtgRzso5d/6c13e86b6844b43127253f533ca11c76/image5-2000-60e276dd.jpeg?w=2000 2000w, https://images.ctfassets.net/3njn2qm7rrbs/5HxFhiCsA6FrFhtgRzso5d/6c13e86b6844b43127253f533ca11c76/image5-2000-60e276dd.jpeg?w=1000 1000w, https://images.ctfassets.net/3njn2qm7rrbs/5HxFhiCsA6FrFhtgRzso5d/6c13e86b6844b43127253f533ca11c76/image5-2000-60e276dd.jpeg?w=500 500w, https://images.ctfassets.net/3njn2qm7rrbs/5HxFhiCsA6FrFhtgRzso5d/6c13e86b6844b43127253f533ca11c76/image5-2000-60e276dd.jpeg?w=100 100w' sizes=900px><img class=j-SmoothLoad src='https://images.ctfassets.net/3njn2qm7rrbs/5HxFhiCsA6FrFhtgRzso5d/6c13e86b6844b43127253f533ca11c76/image5-2000-60e276dd.jpeg?w=1000' sizes=900px width=900px alt='<i>Spreading containers evenly across multiple machines gives services more headroom.</i>'></picture><span class=Caption><i>Spreading containers evenly across multiple machines gives services more headroom.</i></span></p><p>With enough machines to handle failover conditions, we’ve gained reliability (spreading containers on multiple hosts means a single host failure will impact fewer services) and improved performance consistency. However, we still need to address the critical minutes it takes to schedule more copies of our services during an emergency failover. We need traffic shifts to happen in seconds, not minutes, to minimize the number of 500s we may be serving.</p><h2>Always be ready to shift traffic</h2><p>Precious time can be wasted during failovers as we wait for additional containers to be scheduled to handle the new load, download their Docker image, and warm up their workers before finally being able to serve traffic. To shave off those extra minutes, we decided to keep the extra capacity inside the containers as part of the normal service configuration, ensuring we don’t need to add any more during failover. While this may seem like a benign detail, it’s the key to our reliability strategy, since it allows us to always be ready for failover.</p><p class=image><picture><source srcset='https://images.ctfassets.net/3njn2qm7rrbs/3MU5SxkY29GLVMkUmnsrdq/b1b1af84bf5fc6bf7950fb2b437af979/image6-2000-54036bc6.jpeg?w=2000 2000w, https://images.ctfassets.net/3njn2qm7rrbs/3MU5SxkY29GLVMkUmnsrdq/b1b1af84bf5fc6bf7950fb2b437af979/image6-2000-54036bc6.jpeg?w=1000 1000w, https://images.ctfassets.net/3njn2qm7rrbs/3MU5SxkY29GLVMkUmnsrdq/b1b1af84bf5fc6bf7950fb2b437af979/image6-2000-54036bc6.jpeg?w=500 500w, https://images.ctfassets.net/3njn2qm7rrbs/3MU5SxkY29GLVMkUmnsrdq/b1b1af84bf5fc6bf7950fb2b437af979/image6-2000-54036bc6.jpeg?w=100 100w' sizes=900px><img class=j-SmoothLoad src='https://images.ctfassets.net/3njn2qm7rrbs/3MU5SxkY29GLVMkUmnsrdq/b1b1af84bf5fc6bf7950fb2b437af979/image6-2000-54036bc6.jpeg?w=1000' sizes=900px width=900px></picture></p><p>Having additional containers in use has a few key advantages:</p><ul><li><p>We can guarantee high performance consistency to users by keeping free resources on all machines. These resources are already reserved by the services.</p></li><li><p>We already have enough containers to handle a sudden 2x traffic increase.</p></li><li><p>We’ve cut down our dependency on the compute infrastructure’s ability to schedule containers at a critical time.</p></li></ul><p>In order for this strategy to work, however, containers must be correctly sized and ask for the right amount of resources from the compute platform. In a service-oriented architecture, developers are directly in charge of the configuration of their service. This configuration has to reflect our failover strategy, and every service needs to be configured to use exactly 50 percent of its allocated resources, which is what’s required to handle the doubled load during failover.</p><h2>Setting the right values</h2><div class='Sidebar small'><h3>When less is more</h3><p>This example applies to most services with single-threaded workers. Even though this autotuning behavior is documented internally, it can be unintuitive. Only a handful of people who’ve already spent time trying to understand this behavior are aware of it.</p><p>Take a service with a limited number of workers—let’s say four (e.g., one container can serve four requests at a time)—asking for four CPUs. The autoscaler is configured to maintain this service at 50 percent CPU usage.</p><p>Now, imagine this service is underperforming. A developer might naturally think, “Just give it more CPU to make it faster or less constrained.” So they bump the CPU to eight.</p><p>But such a change would likely further slow down the service for at least two reasons:</p><ul><li><p>Workers are mostly single-threaded: No matter what, one container will never use more than four CPU cores.</p></li><li><p>By setting the CPU to eight and keeping a 50 percent target usage for autoscaling, the service will never reach the target usage (four) and the autoscaler will start to scale down the number of containers for that service, decreasing the total capacity of the service to serve traffic.</p></li></ul><p>Perhaps counterintuitively, a better configuration would be to decrease the number of CPUs you’re asking for to two. This is because:</p><ul><li><p>In a service-oriented architecture, most services spend their time waiting on other services, so it’s unlikely you’ll need a full CPU core per worker.</p></li><li><p>Now you’ve reached the autoscaler target and a worker uses an average of 50 percent of one CPU core. The autoscaler increases the number of containers, upping total worker capacity.</p></li></ul><p>This demonstrates why production autoscaling settings should always be adapted to services’ internal architecture.</p></div><p>The main advantage of a service-oriented architecture is greater developer velocity. Developers can deploy the services they own many times a day without having to worry about batching code changes, merge conflicts, release schedules, and so on. Teams fully own their services and associated configurations, including resource allocation like CPU, memory, and autoscaling settings.</p><p class=image><picture><source srcset='https://images.ctfassets.net/3njn2qm7rrbs/7myIkSiWTPiVBnRn7yZY2f/6e42b29c59115149e4e4687b7d581110/image7-2000-8b213b62.jpeg?w=2000 2000w, https://images.ctfassets.net/3njn2qm7rrbs/7myIkSiWTPiVBnRn7yZY2f/6e42b29c59115149e4e4687b7d581110/image7-2000-8b213b62.jpeg?w=1000 1000w, https://images.ctfassets.net/3njn2qm7rrbs/7myIkSiWTPiVBnRn7yZY2f/6e42b29c59115149e4e4687b7d581110/image7-2000-8b213b62.jpeg?w=500 500w, https://images.ctfassets.net/3njn2qm7rrbs/7myIkSiWTPiVBnRn7yZY2f/6e42b29c59115149e4e4687b7d581110/image7-2000-8b213b62.jpeg?w=100 100w' sizes=750px><img class=j-SmoothLoad src='https://images.ctfassets.net/3njn2qm7rrbs/7myIkSiWTPiVBnRn7yZY2f/6e42b29c59115149e4e4687b7d581110/image7-2000-8b213b62.jpeg?w=1000' sizes=750px width=750px alt='<i>Enabling service autoscaling on Yelp’s compute platform, PaaSTA, can be deceptively easy.</i>'></picture><span class=Caption><i>Enabling service autoscaling on Yelp’s compute platform, PaaSTA, can be deceptively easy.</i></span></p><p>Resource allocation is a challenge. In order to be failover-ready, every service needs precise resource allocation to ensure containers use exactly 50 percent of their resources in normal situations. How can we expect software engineers across dozens of teams to fiddle with these settings in production until they get it exactly right?</p><h2>Finding the optimal settings</h2><p>Some core components of Yelp, like the search infrastructure, have been optimized for performance over time. Their owners know exactly how the application behaves under heavy production load, how many threads they can use, what the typical wait versus actual CPU time is, how frequent and expensive garbage collection operations are, and so on. However, most teams at Yelp don’t have all this knowledge. Most services work with the default values, and the only “tuning” we need to do consists of bumping up a few resources here and there as the service grows in popularity and complexity.</p><h2>Abstracting resource declaration</h2><p>We needed to help our developers find the right settings for their services, so we invested in tooling to monitor, recommend, and even abstract away autoscaling settings and resource requirements for all production services. We now analyze services’ resource usage and generate optimized settings for CPU, memory, disk, and so on. Developers can opt out if they want to: A manually set value will always have priority over templated safe defaults and optimized defaults.</p><p class=image><picture><source srcset='https://images.ctfassets.net/3njn2qm7rrbs/4N9N6g6Jf1tXGUm857N7rE/c3f883484cf300f159ee71215a6f4c2e/image2-2000-a371044c.jpeg?w=2000 2000w, https://images.ctfassets.net/3njn2qm7rrbs/4N9N6g6Jf1tXGUm857N7rE/c3f883484cf300f159ee71215a6f4c2e/image2-2000-a371044c.jpeg?w=1000 1000w, https://images.ctfassets.net/3njn2qm7rrbs/4N9N6g6Jf1tXGUm857N7rE/c3f883484cf300f159ee71215a6f4c2e/image2-2000-a371044c.jpeg?w=500 500w, https://images.ctfassets.net/3njn2qm7rrbs/4N9N6g6Jf1tXGUm857N7rE/c3f883484cf300f159ee71215a6f4c2e/image2-2000-a371044c.jpeg?w=100 100w' sizes=900px><img class=j-SmoothLoad src='https://images.ctfassets.net/3njn2qm7rrbs/4N9N6g6Jf1tXGUm857N7rE/c3f883484cf300f159ee71215a6f4c2e/image2-2000-a371044c.jpeg?w=1000' sizes=900px width=900px alt='<i>New containers based on configuration recommendations have 50 percent usage.</i>'></picture><span class=Caption><i>New containers based on configuration recommendations have 50 percent usage.</i></span></p><p>Generating optimized settings for CPU allows us to ensure that services autoscale correctly and have enough capacity for failover. We also reap reliability benefits. For example, if a service hits its memory limit and is killed, its optimized defaults will automatically be updated within two hours with a higher memory allocation. It now takes us less time to fix this particular issue than it took us to diagnose it in the days before autotuning.</p><h3>Autotuned defaults for all</h3><p>After offering optimized settings as opt-ins for a few weeks and building tooling to monitor the system’s success metrics, we felt confident shifting from a recommendation system to optimized values enabled by default. We call this “autotuned defaults.” Unless it’s specifically opted out, every service now automatically declares and uses an optimal amount of resources. A side effect of this change is a significant reduction in compute cost due to the rightsizing of previously oversized services. Best of all, developers enjoy simpler service configurations.</p></div><div class='ArticleLayout flipped'><h2>Setting ourselves up for success</h2><p>Properly allocating the extra capacity required for an emergency failover means we always have the compute capacity we need. Automating service resource allocation removes a considerable burden for service owners and improves the process as a whole. In combination, these strategies simplify incident response and reduce downtime during the most serious outages.</p><p>One of the most interesting aspects of our failover and autoscaling strategy is organizational. By including the extra margin for failover inside every container, multiple teams have become more efficient. The production engineering team now has control over all service configurations, which is a prerequisite for a successful failover. The compute infrastructure team can focus on enhancing the platform without worrying too much about its ability to handle a failover. And developers don’t need to go through the time-consuming process of tuning the resource allocation or autoscaling configuration for their services. Instead, they can focus on what they do best: building a better product for our users and community.</p></div></article></div><div class='u-Container ArticleFooter'><div class=column><div class='u-Grid ContentBody small content'><div class=authors id=authors><div><div class=photo><figure style='background-image:url(https://images.ctfassets.net/3njn2qm7rrbs/2BhSzoJqoDAjhSCM6RfD6s/e40fcd07a64fd5aa86a576a3b465046e/Mathieu_Frappier.png?w=500)'></figure></div><div class=text><h4>About the author</h4><p><b>Mathieu Frappier</b> joined Yelp in 2015 and has worked on many aspects of its production infrastructure. He thrives on improving the efficiency of complex systems.</p><p><a href=https://twitter.com/mat_fra target=_blank>@mat_fra</a></p></div></div><div><div class=photo><figure style='background-image:url(https://images.ctfassets.net/3njn2qm7rrbs/DIoS0MqnlAQQYMQAY86Xb/61832108fc5dc301335c2e41f045b79d/Dorothy_Jung.jpeg?w=500)'></figure></div><div class=text><h4>About the author</h4><p><b>Dorothy Jung</b> is an engineering manager at Yelp. She’s presented on reliability best practices at LISA and SREcon.</p><p><a href=https://twitter.com/ydorothyjung target=_blank>@ydorothyjung</a></p></div></div><div><div class=photo><figure style='background-image:url(https://images.ctfassets.net/3njn2qm7rrbs/4vvkst8fYcGWGXc7oTmb0w/e719df450cb90a3ceffdde1682427220/Qui_Nguyen.png?w=500)'></figure></div><div class=text><h4>About the author</h4><p><b>Qui Nguyen</b> is a senior engineer and tech lead at Yelp. She enjoys solving problems and getting computers to help solve them.</p></div></div></div><div class=topics><div class=text><h4>Topics</h4><p><a href=/topics/learn/ >Learn Something New</a></p></div></div></div></div></div></div><div class=SubscribeBox><a id=newsletter class=anchor href=#newsletter></a><div class='ContentBody inverted store'><div class=u-Container><div class='u-Grid column'><div class=box style=background:#4c70b1><div class=text><h2>Buy the print edition</h2><p class=j-TextBalance>Visit the Increment Store to purchase print issues.</p><p><a class='t-Caps u-Arrow' href=https://store.increment.com/ >Store</a></p></div><a href=https://store.increment.com/ class=magazine><figure class=j-MaskedImage data-fill=/art/19/19-cutout-1000-7ffb5dba.png data-mask=/art/19/fill-cms-1000-83522f38.png></figure></a></div></div></div></div><div class='ContentBody email'><div class=u-Container><div class='u-Grid column'><div class='j-EmailForm box' style='box-shadow:0 -5px 0 #4c70b1'></div></div></div></div></div><div class=ContinueReading><div class=u-Container><div class=column><h2 class='t-Caps xlarge'>Continue Reading</h2><ul class='u-Grid articles'><li class='ArticleBlock footer' style=color:#863051><a class=j-Preload href=/reliability/observability-distributed-tracing/ ><div class='IssueTitle tiny' style=color:#863051><div class='t-Caps meta'><span>16</span></div><h3 class='t-IssueTitle title'>Reliability</h3></div><div class=text><div class=head><h4 class='t-Byline byline'><span>Mads Hartmann</span></h4><h3 class='t-TitleSans title'>Tracing a path to observability</h3></div><div class='t-BodySerif small intro'>A chronicle of Glitch’s efforts to gain visibility into its production systems—and make them more&nbsp;reliable.</div></div></a></li><li class='ArticleBlock footer' style=color:#f5684d><a class=j-Preload href=/documentation/the-complex-world-of-life-saving-safety-critical-software/ ><div class='IssueTitle tiny' style=color:#f5684d><div class='t-Caps meta'><span>6</span></div><h3 class='t-IssueTitle title'>Documentation</h3></div><div class=text><div class=head><h4 class='t-Byline byline'><span>David J. Lumb</span></h4><h3 class='t-TitleSans title'>Inside the complex world of life-saving software</h3></div><div class='t-BodySerif small intro'>Nuclear power plants. Medical devices. Airplanes. Self-driving cars. Developing software for safety-critical projects takes documentation to the next&nbsp;level.</div></div></a></li><li class='ArticleBlock footer' style=color:#c096ca><a class=j-Preload href=/internationalization/paystack-building-a-better-checkout/ ><div class='IssueTitle tiny' style=color:#c096ca><div class='t-Caps meta'><span>8</span></div><h3 class='t-IssueTitle title'>Internationalization</h3></div><div class=text><div class=head><h4 class='t-Byline byline'><span>Opemipo Aikomo</span></h4><h3 class='t-TitleSans title'>The process: Building a better checkout</h3></div><div class='t-BodySerif small intro'>How Paystack reimagined its checkout experience to better serve Africa-based&nbsp;users.</div></div></a></li><li class='ArticleBlock footer' style=color:#40af9e><a class=j-Preload href=/software-architecture/in-space-no-one-can-hear-you-kernel-panic/ ><div class='IssueTitle tiny' style=color:#40af9e><div class='t-Caps meta'><span>12</span></div><h3 class='t-IssueTitle title'>Software Architecture</h3></div><div class=text><div class=head><h4 class='t-Byline byline'><span>Glenn Fleishman</span></h4><h3 class='t-TitleSans title'>In space, no one can hear you kernel panic</h3></div><div class='t-BodySerif small intro'>For NASA, redundancy is all-important. (Why send a single server beyond the stratosphere when you can send&nbsp;five?)</div></div></a></li><li class='ArticleBlock footer' style=color:#863051><a class=j-Preload href=/reliability/testing-beyond-coverage/ ><div class='IssueTitle tiny' style=color:#863051><div class='t-Caps meta'><span>16</span></div><h3 class='t-IssueTitle title'>Reliability</h3></div><div class=text><div class=head><h4 class='t-Byline byline'><span>Benoit Baudry</span> and <span>Martin Monperrus</span></h4><h3 class='t-TitleSans title'>Testing beyond coverage</h3></div><div class='t-BodySerif small intro'>Pseudo-tested methods can be a reliability risk. Here, the authors explain how they developed a methodology and tool to uncover them in Java applications.</div></div></a></li><li class='ArticleBlock footer' style=color:#863051><a class=j-Preload href=/reliability/chaos-engineering/ ><div class='IssueTitle tiny' style=color:#863051><div class='t-Caps meta'><span>16</span></div><h3 class='t-IssueTitle title'>Reliability</h3></div><div class=text><div class=head><h4 class='t-Byline byline'><span>Ana Margarita Medina</span></h4><h3 class='t-TitleSans title'>Chaotic good</h3></div><div class='t-BodySerif small intro'>As software systems become ever more complex, chaos engineering provides a (not-actually-so-chaotic) tool kit for building more reliable and resilient&nbsp;systems.</div></div></a></li><li class='ArticleBlock footer' style=color:#863051><a class=j-Preload href=/reliability/solar-storm-impact/ ><div class='IssueTitle tiny' style=color:#863051><div class='t-Caps meta'><span>16</span></div><h3 class='t-IssueTitle title'>Reliability</h3></div><div class=text><div class=head><h4 class='t-Byline byline'><span>Ian Steadman</span></h4><h3 class='t-TitleSans title'>Earth, wind, and solar fire</h3></div><div class='t-BodySerif small intro'>If a major solar storm were to sweep across Earth, would today’s electrical and communications infrastructure be resilient enough to endure its&nbsp;impact?</div></div></a></li><li class='ArticleBlock footer' style=color:#863051><a class=j-Preload href=/reliability/home-network-isp/ ><div class='IssueTitle tiny' style=color:#863051><div class='t-Caps meta'><span>16</span></div><h3 class='t-IssueTitle title'>Reliability</h3></div><div class=text><div class=head><h4 class='t-Byline byline'><span>Poornima Apte</span></h4><h3 class='t-TitleSans title'>Home sweet home network</h3></div><div class='t-BodySerif small intro'>Facing dramatic shifts in residential usage, internet service providers are working to keep latency low and connectivity&nbsp;high.</div></div></a></li><li class='ArticleBlock footer' style=color:#ef766e><a class=j-Preload href=/on-call/when-the-pager-goes-off/ ><div class='IssueTitle tiny' style=color:#ef766e><div class='t-Caps meta'><span>1</span></div><h3 class='t-IssueTitle title'>On-Call</h3></div><div class=text><div class=head><h4 class='t-Byline byline'><span>Increment Staff</span></h4><h3 class='t-TitleSans title'>What happens when the pager goes off?</h3></div><div class='t-BodySerif small intro'>To discover the state of incident response across the tech industry, we surveyed over thirty industry leaders (including Amazon, Dropbox, Facebook, Google, and Netflix) about their incident response processes.</div></div></a></li></ul><h3 class='t-Caps large'>Explore Topics</h3><ul class='t-BodySerif large topics'><li><a href=/topics/learn/ >Learn Something New</a></li><li><a href=/topics/scaling/ >Scaling &amp; Growth</a></li><li><a href=/topics/ask-an-expert/ >Ask an Expert</a></li><li><a href=/topics/interviews/ >Interviews &amp; Surveys</a></li><li><a href=/topics/guides/ >Guides &amp; Best Practices</a></li><li><a href=/topics/opinion/ >Essays &amp; Opinion</a></li><li><a href=/topics/culture/ >Workplace &amp; Culture</a></li></ul><h3 class='t-Caps large'>All Issues</h3><ul class=issues><li><a href=/planning/ ><div class='IssueTitle large' style=color:#4c70b1><div class='t-Caps meta'><span>Issue 19</span> <span>November 2021</span></div><h1 class='t-IssueTitle title'>Planning</h1></div></a></li><li><a href=/mobile/ ><div class='IssueTitle large' style=color:#439eab><div class='t-Caps meta'><span>Issue 18</span> <span>August 2021</span></div><h1 class='t-IssueTitle title'>Mobile</h1></div></a></li><li><a href=/containers/ ><div class='IssueTitle large' style=color:#443d79><div class='t-Caps meta'><span>Issue 17</span> <span>May 2021</span></div><h1 class='t-IssueTitle title'>Containers</h1></div></a></li><li><a href=/reliability/ ><div class='IssueTitle large' style=color:#863051><div class='t-Caps meta'><span>Issue 16</span> <span>February 2021</span></div><h1 class='t-IssueTitle title'>Reliability</h1></div></a></li><li><a href=/remote/ ><div class='IssueTitle large' style=color:#29386a><div class='t-Caps meta'><span>Issue 15</span> <span>November 2020</span></div><h1 class='t-IssueTitle title'>Remote</h1></div></a></li><li><a href=/apis/ ><div class='IssueTitle large' style=color:#00afbe><div class='t-Caps meta'><span>Issue 14</span> <span>August 2020</span></div><h1 class='t-IssueTitle title'>APIs</h1></div></a></li><li><a href=/frontend/ ><div class='IssueTitle large' style=color:#5ebe92><div class='t-Caps meta'><span>Issue 13</span> <span>May 2020</span></div><h1 class='t-IssueTitle title'>Frontend</h1></div></a></li><li><a href=/software-architecture/ ><div class='IssueTitle large' style=color:#40af9e><div class='t-Caps meta'><span>Issue 12</span> <span>February 2020</span></div><h1 class='t-IssueTitle title'>Software Architecture</h1></div></a></li><li><a href=/teams/ ><div class='IssueTitle large' style=color:#8e65bf><div class='t-Caps meta'><span>Issue 11</span> <span>November 2019</span></div><h1 class='t-IssueTitle title'>Teams</h1></div></a></li><li><a href=/testing/ ><div class='IssueTitle large' style=color:#e89e00><div class='t-Caps meta'><span>Issue 10</span> <span>August 2019</span></div><h1 class='t-IssueTitle title'>Testing</h1></div></a></li><li><a href=/open-source/ ><div class='IssueTitle large' style=color:#4a5ad3><div class='t-Caps meta'><span>Issue 9</span> <span>May 2019</span></div><h1 class='t-IssueTitle title'>Open Source</h1></div></a></li><li><a href=/internationalization/ ><div class='IssueTitle large' style=color:#c096ca><div class='t-Caps meta'><span>Issue 8</span> <span>February 2019</span></div><h1 class='t-IssueTitle title'>Internationalization</h1></div></a></li><li><a href=/security/ ><div class='IssueTitle large' style=color:#4dbac5><div class='t-Caps meta'><span>Issue 7</span> <span>October 2018</span></div><h1 class='t-IssueTitle title'>Security</h1></div></a></li><li><a href=/documentation/ ><div class='IssueTitle large' style=color:#f5684d><div class='t-Caps meta'><span>Issue 6</span> <span>August 2018</span></div><h1 class='t-IssueTitle title'>Documentation</h1></div></a></li><li><a href=/programming-languages/ ><div class='IssueTitle large' style=color:#d69336><div class='t-Caps meta'><span>Issue 5</span> <span>April 2018</span></div><h1 class='t-IssueTitle title'>Programming Languages</h1></div></a></li><li><a href=/energy-environment/ ><div class='IssueTitle large' style=color:#d6658e><div class='t-Caps meta'><span>Issue 4</span> <span>February 2018</span></div><h1 class='t-IssueTitle title'>Energy & Environment</h1></div></a></li><li><a href=/development/ ><div class='IssueTitle large' style=color:#53a88e><div class='t-Caps meta'><span>Issue 3</span> <span>October 2017</span></div><h1 class='t-IssueTitle title'>Development</h1></div></a></li><li><a href=/cloud/ ><div class='IssueTitle large' style=color:#707aed><div class='t-Caps meta'><span>Issue 2</span> <span>July 2017</span></div><h1 class='t-IssueTitle title'>Cloud</h1></div></a></li><li><a href=/on-call/ ><div class='IssueTitle large' style=color:#ef766e><div class='t-Caps meta'><span>Issue 1</span> <span>April 2017</span></div><h1 class='t-IssueTitle title'>On-Call</h1></div></a></li></ul></div></div></div><footer class=PageFooter><div class='u-Container ContentBody small'><svg style=display:none><symbol id=twitterIcon viewBox='0 0 32 32'><path d='M32.1 6c-1.2.5-2.5.9-3.8 1 1.4-.8 2.4-2.1 2.9-3.6-1.3.8-2.7 1.3-4.2 1.6a6.8 6.8 0 0 0-4.8-2c-3.6 0-6.6 3-6.6 6.6 0 .5.1 1 .2 1.5-5.5-.3-10.4-3-13.6-6.9-.6 1-.9 2.1-.9 3.3 0 2.3 1.2 4.3 2.9 5.5-1.1 0-2.1-.3-3-.8v.1c0 3.2 2.3 5.9 5.3 6.5-.6.2-1.1.2-1.7.2-.4 0-.8 0-1.2-.1.8 2.6 3.3 4.5 6.2 4.6-2.3 1.8-5.1 2.8-8.2 2.8-.5 0-1.1 0-1.6-.1C2.9 28 6.3 29 10 29c12.1 0 18.7-10 18.7-18.7v-.9c1.4-.9 2.5-2 3.4-3.4z' fill=currentColor /></symbol><symbol id=facebookIcon viewBox='0 0 32 32'><path d='M30.2 0H1.8C.8 0 0 .8 0 1.8v28.5c0 1 .8 1.8 1.8 1.8h15.3V19.6h-4.2v-4.8h4.2v-3.6c0-4.1 2.5-6.4 6.2-6.4 1.8 0 3.3.2 3.7.2v4.3h-2.6c-2 0-2.4 1-2.4 2.4v3.1h4.8l-.6 4.8H22V32h8.2c1 0 1.8-.8 1.8-1.8V1.8c0-1-.8-1.8-1.8-1.8z' fill=currentColor /></symbol><symbol id=rssIcon viewBox='0 0 32 32'><path d='M10.7 25.6c0 2.4-2 4.4-4.4 4.4S2 28 2 25.6s2-4.4 4.4-4.4 4.3 2 4.3 4.4zM6.1 2c-.6 0-1.3 0-2 .1-1.3.1-2.2 1.2-2.1 2.4.1 1.2 1.2 2.2 2.4 2.1.6 0 1.1-.1 1.6-.1 10.7 0 19.4 8.7 19.4 19.4 0 .5 0 1-.1 1.6-.1 1.2.8 2.3 2.1 2.4h.2c1.2 0 2.1-.9 2.2-2.1.1-.7.1-1.4.1-2C30 12.7 19.3 2 6.1 2zm-.7 9.6c-.4 0-.8 0-1.3.1-1.3.1-2.2 1.2-2.1 2.5.1 1.3 1.2 2.2 2.5 2.1h.9c5.7 0 10.4 4.7 10.4 10.4v.9c-.1 1.3.8 2.4 2.1 2.5h.2c1.2 0 2.2-.9 2.3-2.1 0-.4.1-.8.1-1.2-.1-8.4-6.9-15.2-15.1-15.2z' fill=currentColor /></symbol><symbol id=linkedInIcon viewBox='0 0 32 32'><path d='M29.6,0H2.4C1.1,0,0,1,0,2.3v27.4C0,31,1.1,32,2.4,32h27.3c1.3,0,2.4-1,2.4-2.3V2.3C32,1,30.9,0,29.6,0z M9.5,27.3H4.7V12 h4.8V27.3z M7.1,9.9c-1.5,0-2.8-1.2-2.8-2.8c0-1.5,1.2-2.8,2.8-2.8c1.5,0,2.8,1.2,2.8,2.8C9.9,8.7,8.6,9.9,7.1,9.9z M27.3,27.3 h-4.7v-7.4c0-1.8,0-4-2.5-4c-2.5,0-2.8,1.9-2.8,3.9v7.6h-4.7V12H17v2.1h0.1c0.6-1.2,2.2-2.5,4.5-2.5c4.8,0,5.7,3.2,5.7,7.3V27.3z' fill=currentColor /></symbol></svg><div class='column main'><section class=social><a href=https://twitter.com/incrementmag class=twitter><svg viewBox='0 0 32 32'><use xlink:href=#twitterIcon x=0 y=0></use></svg> <span>@incrementmag</span> </a><a href=https://facebook.com/incrementmag class=facebook><svg viewBox='0 0 32 32'><use xlink:href=#facebookIcon x=0 y=0></use></svg> <span>incrementmag</span> </a><a href=/feed.xml class=rss><svg viewBox='0 0 32 32'><use xlink:href=#rssIcon x=0 y=0></use></svg> <span>RSS Feed</span></a></section><section><h4>About</h4><p><em>Increment</em> is a print and digital magazine about how teams build and operate software systems at scale. <a href=/about/ >Learn more</a></p></section><section><h4>Work with us</h4><p>Interested in joining the team at Stripe? <a href=https://stripe.com/jobs>View job openings</a></p></section></div><p class='column copyright'><span>&copy; 2022 <em>Increment</em></span> <a href=https://stripe.com>Published by Stripe</a> <a href=https://stripe.com/privacy/media-policy>Privacy policy</a></p></div></footer></body></html>