Files
nexus/sreweekly/articles/263/01-increment-reliability-tracing-a-path-to-observability.html
2026-09-12 17:23:01 +08:00

19 lines
30 KiB
HTML
Raw Permalink Blame History

This file contains invisible Unicode characters
This file contains invisible Unicode characters that are indistinguishable to humans but may be processed differently by a computer. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
<!doctype html><html><head><meta charset=utf-8><title>Tracing a Path to Observability - Increment</title><meta name=description content='Find out how observability and distributed tracing helped software engineers at Glitch gain visibility into their production systems and make them more reliable.'><link rel=canonical href=http://localhost:3000/reliability/observability-distributed-tracing/ ><link rel=apple-touch-icon-precomposed href=/img/icon-571805a1.png><meta property=og:title content='Tracing a path to observability – Increment: Reliability'><meta property=og:url content=http://localhost:3000/reliability/observability-distributed-tracing/ ><meta property=og:description content='A chronicle of Glitch’s efforts to gain visibility into its production systems—and make them more reliable.'><meta property=og:image content='https://images.ctfassets.net/3njn2qm7rrbs/6Q3D0w8nis0u66FQeCTjbg/843b26ce0013258589c76a15a49461d4/cover-issue16.png?w=1000'><meta name=twitter:card content=summary_large_image><meta name=twitter:image content='https://images.ctfassets.net/3njn2qm7rrbs/6Q3D0w8nis0u66FQeCTjbg/843b26ce0013258589c76a15a49461d4/cover-issue16.png?w=1000'><meta name=twitter:site content=@IncrementMag><meta name=twitter:title content='Tracing a path to observability – Increment: Reliability'><meta name=twitter:description content='A chronicle of Glitch’s efforts to gain visibility into its production systems—and make them more reliable.'><link rel=alternate type=application/rss+xml title=Increment href=/feed.xml><meta name=viewport content='width=device-width,initial-scale=1'><link rel=preload href=/fonts/baton-turbo/400-30a55d66.woff2 as=font type=font/woff2 crossorigin=anonymous><link rel=preload href=/fonts/baton-turbo/500-1603c0e8.woff2 as=font type=font/woff2 crossorigin=anonymous><link rel=preload href=/fonts/tiempos-text/400-c4810745.woff2 as=font type=font/woff2 crossorigin=anonymous><link rel=preload href=/fonts/tiempos-head/700-383ede62.woff2 as=font type=font/woff2 crossorigin=anonymous><link rel=stylesheet type=text/css href=/css/bundle-289d885f.css><link rel=stylesheet type=text/css href=/css/issues/16-c2b7b183.css><script>// Don't fade art if it loads ~instantly
setTimeout(()=>{document.documentElement.classList.add('fadeArt')},250);</script><script defer src=/js/defer-737fbd90.js></script><script>const INCREMENT_META={issueNumber:16,issueSlug:'reliability',articleSlug:'observability-distributed-tracing'};</script></head><body class='Issue_reliability Article_observability-distributed-tracing'><nav class=PageNav><div class=u-Container><div class=column><h1 class=logo><a href=/ ><img src=/img/logo-ae2c55d5.svg alt=Increment></a></h1><a class=out-now href=https://store.increment.com/ style=color:#595959><div class='IssueTitle tiny'><div class='t-Caps meta tiny'><span>NEW</span></div><h3 class='t-IssueTitle title'>Buy the print edition</h3></div></a><ul class=nav><li><a href=/issues/ ><span>Issues</span></a></li><li><a href=/topics/ ><span>Topics</span></a></li><li><a href=https://store.increment.com/ ><span>Store</span></a></li><li><a href=/about/ ><span>About</span></a></li></ul></div></div></nav><div class=ArticlePage itemscope itemtype=http://schema.org/Article><script type=application/ld+json>{
"@context": "http://schema.org",
"@type": "Article",
"headline": "Tracing a path to observability",
"image": " https://images.ctfassets.net/3njn2qm7rrbs/6Q3D0w8nis0u66FQeCTjbg/843b26ce0013258589c76a15a49461d4/cover-issue16.png?w&#x3D;1000",
"datePublished": "Thu, 25 Feb 2021 19:00:00 GMT",
"dateModified": "Thu, 25 Feb 2021 19:00:00 GMT",
"publisher": {
"@type": "Organization",
"name": "Increment",
"logo": {
"@type": "ImageObject",
"url": "https://increment.com/img/logo.png"
}
},
"description": "A chronicle of Glitch’s efforts to gain visibility into its production systems—and make them more reliable.",
"mainEntityOfPage": "http://localhost:3000/reliability/observability-distributed-tracing/"
}</script><header class='u-Container ArticleHeader'><div class=column><div class=u-Grid><div class=main><h4 class='t-Byline byline large'><a href=#authors class=j-SmoothScroll itemprop=author><span>Mads Hartmann</span></a></h4><h1 class='t-TitleSerif large title' itemprop=name>Tracing a path to observability</h1><div class='t-BodySans large intro' itemprop=description>A chronicle of Glitch’s efforts to gain visibility into its production systems—and make them more&nbsp;reliable.</div></div><a class=issue href=/reliability/ ><span class='t-Caps tiny part-of'>Part of</span><div class='IssueTitle small' style=color:#863051><div class='t-Caps meta'><span>Issue 16</span> <span>February 2021</span></div><h2 class='t-IssueTitle title'>Reliability</h2></div></a></div></div></header><div class='u-Container ArticleContent'><article class='ContentBody column' itemprop=articleBody><div class=ArticleLayout><p>September 2019. For months, we’d been seeing a steep increase in the number of projects being built on Glitch, and more and more traffic headed to them. We struggled to keep up. What had started as performance edge cases had become commonplace. Many unknown unknowns had become known unknowns.</p><p>This reconstructed diary chronicles our efforts to better understand our production systems and ultimately make them more reliable, a journey that led us to adopt distributed tracing and use an observability vendor. Hopefully reading about our experience can offer some inspiration and insight for your own.</p><h2>Day 1: We have a problem</h2><p>On Glitch, you can build everything from simple static sites to full-stack Node.js projects. When full-stack projects aren’t worked on or visited for a few minutes, we “stop” them to save resources and “start” them again on demand. How long a project takes to start depends on its complexity, but recently we’ve noticed something strange: The same project might start in a couple of seconds one day and a minute the next. To make matters worse, an increasing number of users are reporting these fluctuating start times.</p><p>We can’t find any obvious reason or pattern that would explain this. It’s time to start debugging.</p><h2>Day 2: Digging in</h2><p>We have metrics that might help us debug this issue. I’m quite fond of one in particular—we go way back—that measures how long it takes to start a project. In the past, we’ve used it to identify full-system performance regressions.</p><p>The problem is we only have access to aggregated values such as the mean, median, and 95th percentile. Since the slowdown isn’t affecting every project, these aren’t very useful. The slow-to-start projects haven’t made a dent in the aggregates, and we can’t use the metric to narrow down which projects are affected.</p><p>We might be able to pinpoint the cause if we can slice the metric by a few more dimensions: project ID, release version of the service, instance ID, and so on. We’ll add a few more labels and debug further tomorrow.</p><h2>Day 3: You’re charging us how much?</h2><p>Oh no. We’ve blown through the number of time series our metrics vendor allows per month. Unless we get the situation under control, they’ll charge us an additional $200,000—monthly.</p><p>Project ID, it turns out, is a high-cardinality label—that is, a label for which there’s a very high range of possible values. For each unique combination of labels, a new time series is created. The total possible number of time series can be calculated as <code># of projects * # of hosts * # of service release tags</code>. With millions of projects, thousands of hosts, and several releases each day, we’re producing millions of time series.</p><p>Without support for high-cardinality labels, we’re greatly restricted in the kinds of questions we can ask of our systems. Metrics won’t be the golden ticket to debugging this narrow performance regression. Better remove those labels again…</p><h2>Day 4: Let’s check the logs</h2><p>Luckily, we have structured logs. Our services emit blobs of JSON objects that are processed by our log pipeline. There’s no limit to the number or cardinality of labels we can attach to them.</p><p>We’ve instrumented our services such that a unique request ID is generated for each request and propagated through our systems; this allows us to easily view all logs for a single request. We’ve also instrumented our most important code paths with log lines that trace the timing of important operations, resulting in log lines that read <code>&lt;operation name&gt; started</code> and <code>&lt;operation name&gt; ended &lt;duration&gt;</code>. They allow us to look at individual requests and identify where latencies are introduced. We’ve used them to debug some pretty gnarly performance regressions. It’s not the easiest way to do it, but it’s better than nothing.</p><p>There’s one problem, though: scalability. Currently, 65 percent of our log lines are for timing operations. As traffic increases, this is becoming prohibitively expensive. There must be a better way.</p></div><div class='ArticleLayout flipped'><h2>Day 23: Observability to the rescue</h2><p>The term “observability” has been popping up on my Twitter feed a lot recently—and it turns out there’s a whole community of people who’ve been through this before. Many have invested their talents, time, and resources into figuring out how best to build systems that are easy to inspect when operational hiccups, performance regressions, and other unexpected things happen.</p><p>Observability is all about being able to ask questions of your systems and get answers based on the telemetry they produce. We’re definitely in need of some answers.</p><h2>Day 25: Telemetry talks</h2><p>After countless hours of research, we believe distributed tracing is a promising solution. It addresses both of the issues we experienced with metrics and logs, allowing for high-cardinality labels and supporting sampling, which keeps costs down.</p><p>A distributed trace tracks the progression of a single request as it’s handled by your services. A trace is structured as a tree of spans, which are named, timed operations to which you can attach an arbitrary number of labels. If you squint, you could argue that our current logging infrastructure is a bespoke implementation of distributed traces. However, a proper distributed tracing solution allows for sampling, and the traces contain a richer structure, which allows for more sophisticated visualizations and analysis.</p><p>So, how do we go about getting our services to emit trace data?</p><h2>Day 30: Tuning our instrumentation</h2><p>Instrumenting your services to emit distributed traces is more conceptually advanced than logs or metrics, since you have to propagate contexts between services and think about sampling. In our case, though, this isn’t too heavy a lift. Context propagation isn’t a new concept to us—we already propagate request IDs for our logs—and the ability to perform sampling is one of the reasons we wanted to look into distributed traces. It also helps that we only have a handful of microservices.</p><p>Since instrumenting services is time-consuming, we’ve decided to use the vendor-agnostic <a href=https://opencensus.io/ target=_blank rel='noopener noreferrer'>OpenCensus</a> project (later superseded by <a href=https://opentelemetry.io/ target=_blank rel='noopener noreferrer'>OpenTelemetry</a>). That way, we can always decide to change vendors without having to re-instrument everything. We use the Node SDK to instrument our services and have them send traces to the OpenCensus collector. The SDK takes care of context propagating and sampling, while the OpenCensus collector ingests the trace data and sends it off to a vendor or tool, in our case <a href=https://www.honeycomb.io/ target=_blank rel='noopener noreferrer'>Honeycomb</a>. We decided to use a vendor rather than self-host a tool to save our small team of engineers time.</p><h2>Day 32: Catching the culprit</h2><p>Now that we have access to highly detailed traces and sophisticated tooling to query them, we’ve been able to identify the cause of the slow-to-start projects: caching. We visualized a heat map of the duration of all traces for starting projects, and inspected the trace-view of one of the slowest. This showed us that latencies were introduced in the <code>project.initialization</code> span. We then visualized a heat map of the duration of all <code>project.initialization</code> spans and started grouping by various labels until we found a pattern. Spans labeled <code>project.reuse = false</code> were consistently slower.</p><p>The tool couldn’t tell us why—for that we’d have to rely on our expertise and knowledge of our systems, or pull out the source code and start digging. In this case, we knew that to speed up project start time, we try to place projects on the host they last ran on to reduce the amount of initialization we have to do. This shortens start times and reduces strain on the hosts. However, now that we’re running many more projects, the previous host is far less likely to have available capacity. Projects increasingly have to perform cold starts, which makes it more likely that other projects will start up simultaneously, creating a vicious circle. We’ll have to revisit that scheduling algorithm.</p></div><div class=ArticleLayout><h2>Day 176: Welcome side effects</h2><p>December 2019. We’ve had all of our services instrumented for about six months now—it’s time to take stock of our observability journey.</p><p>We’re still iterating on our sampling configuration. We use head-based sampling and only retain about 3 percent of traces. This has allowed us to debug most issues, but the low sampling rate can be problematic if we’re trying to investigate something less common, as it’s unlikely to get sampled. There are more sophisticated sampling techniques we could adopt to alleviate this, but for now we still rely on logs in these cases.</p><p>There’s always a learning curve when introducing a new tool, and adoption has to be nurtured. We had an initial training session with Honeycomb, performed regular show-and-tell sessions internally to share tips and tricks, and highlighted interesting queries during incident retrospectives. Most of our engineers are now comfortable with Honeycomb and use it regularly.</p><p>So, was it worth the investment in time and resources? I believe so. The observability practices we’ve adopted to help debug slow project starts have had positive side effects. With richer telemetry and more sophisticated tooling, we’re quicker to identify problems and disprove hypotheses. We’ve greatly reduced the duration of incidents and rarely have to make code changes to get answers to our questions. This is where observability really shines: It enhances your ability to ask questions of your systems, including ones you hadn’t originally thought to ask.</p><p>But perhaps the most valuable thing we’ve gotten out of this journey is a shift in perspective. Now that we have deeper visibility into our production systems, we’re much more confident carrying out experiments. We’ve started releasing smaller changes, guarded by feature flags, and observing how our systems behave. This has reduced the blast radius of code changes: We can shut down experiments before they develop into full-fledged incidents. We think these practices, and the mindset shift they’ve prompted, will benefit our systems and our users—and we hope our learnings will benefit you, too.</p></div></article></div><div class='u-Container ArticleFooter'><div class=column><div class='u-Grid ContentBody small content'><div class=authors id=authors><div><div class=photo><figure style='background-image:url(https://images.ctfassets.net/3njn2qm7rrbs/5iPEAntjfUNzh9IsJaOIG7/e2a7f9e7b6c6d083c697bdff43e4e8d2/Mads_Hartmann.png?w=500)'></figure></div><div class=text><h4>About the author</h4><p><b>Mads Hartmann</b> is an engineering manager at Glitch. When he’s not rummaging around in production, he’s making it easier to understand and more reliable.</p><p><a href=https://twitter.com/Mads_Hartmann target=_blank>@Mads_Hartmann</a></p></div></div></div><div class=topics><div class=text><h4>Topics</h4><p><a href=/topics/learn/ >Learn Something New</a></p></div></div></div></div></div></div><div class=SubscribeBox><a id=newsletter class=anchor href=#newsletter></a><div class='ContentBody inverted store'><div class=u-Container><div class='u-Grid column'><div class=box style=background:#4c70b1><div class=text><h2>Buy the print edition</h2><p class=j-TextBalance>Visit the Increment Store to purchase print issues.</p><p><a class='t-Caps u-Arrow' href=https://store.increment.com/ >Store</a></p></div><a href=https://store.increment.com/ class=magazine><figure class=j-MaskedImage data-fill=/art/19/19-cutout-1000-7ffb5dba.png data-mask=/art/19/fill-cms-1000-83522f38.png></figure></a></div></div></div></div><div class='ContentBody email'><div class=u-Container><div class='u-Grid column'><div class='j-EmailForm box' style='box-shadow:0 -5px 0 #4c70b1'></div></div></div></div></div><div class=ContinueReading><div class=u-Container><div class=column><h2 class='t-Caps xlarge'>Continue Reading</h2><ul class='u-Grid articles'><li class='ArticleBlock footer' style=color:#863051><a class=j-Preload href=/reliability/yelp-traffic-failover-strategy/ ><div class='IssueTitle tiny' style=color:#863051><div class='t-Caps meta'><span>16</span></div><h3 class='t-IssueTitle title'>Reliability</h3></div><div class=text><div class=head><h4 class='t-Byline byline'><span>Mathieu Frappier</span>, <span>Dorothy Jung</span>, and <span>Qui Nguyen</span></h4><h3 class='t-TitleSans title'>The process: Implementing Yelp’s failover strategy</h3></div><div class='t-BodySerif small intro'>How Yelp engineers orchestrated their traffic failover process and effected a delicate balance between reliability, performance, and cost efficiency.</div></div></a></li><li class='ArticleBlock footer' style=color:#863051><a class=j-Preload href=/reliability/solar-storm-impact/ ><div class='IssueTitle tiny' style=color:#863051><div class='t-Caps meta'><span>16</span></div><h3 class='t-IssueTitle title'>Reliability</h3></div><div class=text><div class=head><h4 class='t-Byline byline'><span>Ian Steadman</span></h4><h3 class='t-TitleSans title'>Earth, wind, and solar fire</h3></div><div class='t-BodySerif small intro'>If a major solar storm were to sweep across Earth, would today’s electrical and communications infrastructure be resilient enough to endure its&nbsp;impact?</div></div></a></li><li class='ArticleBlock footer' style=color:#863051><a class=j-Preload href=/reliability/home-network-isp/ ><div class='IssueTitle tiny' style=color:#863051><div class='t-Caps meta'><span>16</span></div><h3 class='t-IssueTitle title'>Reliability</h3></div><div class=text><div class=head><h4 class='t-Byline byline'><span>Poornima Apte</span></h4><h3 class='t-TitleSans title'>Home sweet home network</h3></div><div class='t-BodySerif small intro'>Facing dramatic shifts in residential usage, internet service providers are working to keep latency low and connectivity&nbsp;high.</div></div></a></li><li class='ArticleBlock footer' style=color:#707aed><a class=j-Preload href=/cloud/weather-control-as-a-service/ ><div class='IssueTitle tiny' style=color:#707aed><div class='t-Caps meta'><span>2</span></div><h3 class='t-IssueTitle title'>Cloud</h3></div><div class=text><div class=head><h4 class='t-Byline byline'><span>Ingrid Burrington</span></h4><h3 class='t-TitleSans title'>Weather control as a service: The scaling and seeding of cloud infrastructure</h3></div><div class='t-BodySerif small intro'>The fierce race between the cloud computing giants — Amazon, Microsoft, and Google — and their quests to reshape the technological and physical landscape.</div></div></a></li><li class='ArticleBlock footer' style=color:#f5684d><a class=j-Preload href=/documentation/the-complex-world-of-life-saving-safety-critical-software/ ><div class='IssueTitle tiny' style=color:#f5684d><div class='t-Caps meta'><span>6</span></div><h3 class='t-IssueTitle title'>Documentation</h3></div><div class=text><div class=head><h4 class='t-Byline byline'><span>David J. Lumb</span></h4><h3 class='t-TitleSans title'>Inside the complex world of life-saving software</h3></div><div class='t-BodySerif small intro'>Nuclear power plants. Medical devices. Airplanes. Self-driving cars. Developing software for safety-critical projects takes documentation to the next&nbsp;level.</div></div></a></li><li class='ArticleBlock footer' style=color:#40af9e><a class=j-Preload href=/software-architecture/in-space-no-one-can-hear-you-kernel-panic/ ><div class='IssueTitle tiny' style=color:#40af9e><div class='t-Caps meta'><span>12</span></div><h3 class='t-IssueTitle title'>Software Architecture</h3></div><div class=text><div class=head><h4 class='t-Byline byline'><span>Glenn Fleishman</span></h4><h3 class='t-TitleSans title'>In space, no one can hear you kernel panic</h3></div><div class='t-BodySerif small intro'>For NASA, redundancy is all-important. (Why send a single server beyond the stratosphere when you can send&nbsp;five?)</div></div></a></li><li class='ArticleBlock footer' style=color:#40af9e><a class=j-Preload href=/software-architecture/illuminating-the-grid/ ><div class='IssueTitle tiny' style=color:#40af9e><div class='t-Caps meta'><span>12</span></div><h3 class='t-IssueTitle title'>Software Architecture</h3></div><div class=text><div class=head><h4 class='t-Byline byline'><span>Chris Stokel-Walker</span></h4><h3 class='t-TitleSans title'>Illuminating the grid</h3></div><div class='t-BodySerif small intro'>While the way we consume electricity has changed dramatically, utilities have been slow to catch up. Here’s a look at the challenges facing the software powering our electrical grids—and some of the proposed solutions.</div></div></a></li><li class='ArticleBlock footer' style=color:#863051><a class=j-Preload href=/reliability/testing-beyond-coverage/ ><div class='IssueTitle tiny' style=color:#863051><div class='t-Caps meta'><span>16</span></div><h3 class='t-IssueTitle title'>Reliability</h3></div><div class=text><div class=head><h4 class='t-Byline byline'><span>Benoit Baudry</span> and <span>Martin Monperrus</span></h4><h3 class='t-TitleSans title'>Testing beyond coverage</h3></div><div class='t-BodySerif small intro'>Pseudo-tested methods can be a reliability risk. Here, the authors explain how they developed a methodology and tool to uncover them in Java applications.</div></div></a></li><li class='ArticleBlock footer' style=color:#863051><a class=j-Preload href=/reliability/high-performing-team-trust/ ><div class='IssueTitle tiny' style=color:#863051><div class='t-Caps meta'><span>16</span></div><h3 class='t-IssueTitle title'>Reliability</h3></div><div class=text><div class=head><h4 class='t-Byline byline'><span>Tess Donnelly</span> and <span>Tiarnán de Burca</span></h4><h3 class='t-TitleSans title'>Trust is an enabling technology</h3></div><div class='t-BodySerif small intro'>To build a high-performing software delivery system, your stack’s capabilities are just one part of the&nbsp;picture.</div></div></a></li></ul><h3 class='t-Caps large'>Explore Topics</h3><ul class='t-BodySerif large topics'><li><a href=/topics/learn/ >Learn Something New</a></li><li><a href=/topics/scaling/ >Scaling &amp; Growth</a></li><li><a href=/topics/ask-an-expert/ >Ask an Expert</a></li><li><a href=/topics/interviews/ >Interviews &amp; Surveys</a></li><li><a href=/topics/guides/ >Guides &amp; Best Practices</a></li><li><a href=/topics/opinion/ >Essays &amp; Opinion</a></li><li><a href=/topics/culture/ >Workplace &amp; Culture</a></li></ul><h3 class='t-Caps large'>All Issues</h3><ul class=issues><li><a href=/planning/ ><div class='IssueTitle large' style=color:#4c70b1><div class='t-Caps meta'><span>Issue 19</span> <span>November 2021</span></div><h1 class='t-IssueTitle title'>Planning</h1></div></a></li><li><a href=/mobile/ ><div class='IssueTitle large' style=color:#439eab><div class='t-Caps meta'><span>Issue 18</span> <span>August 2021</span></div><h1 class='t-IssueTitle title'>Mobile</h1></div></a></li><li><a href=/containers/ ><div class='IssueTitle large' style=color:#443d79><div class='t-Caps meta'><span>Issue 17</span> <span>May 2021</span></div><h1 class='t-IssueTitle title'>Containers</h1></div></a></li><li><a href=/reliability/ ><div class='IssueTitle large' style=color:#863051><div class='t-Caps meta'><span>Issue 16</span> <span>February 2021</span></div><h1 class='t-IssueTitle title'>Reliability</h1></div></a></li><li><a href=/remote/ ><div class='IssueTitle large' style=color:#29386a><div class='t-Caps meta'><span>Issue 15</span> <span>November 2020</span></div><h1 class='t-IssueTitle title'>Remote</h1></div></a></li><li><a href=/apis/ ><div class='IssueTitle large' style=color:#00afbe><div class='t-Caps meta'><span>Issue 14</span> <span>August 2020</span></div><h1 class='t-IssueTitle title'>APIs</h1></div></a></li><li><a href=/frontend/ ><div class='IssueTitle large' style=color:#5ebe92><div class='t-Caps meta'><span>Issue 13</span> <span>May 2020</span></div><h1 class='t-IssueTitle title'>Frontend</h1></div></a></li><li><a href=/software-architecture/ ><div class='IssueTitle large' style=color:#40af9e><div class='t-Caps meta'><span>Issue 12</span> <span>February 2020</span></div><h1 class='t-IssueTitle title'>Software Architecture</h1></div></a></li><li><a href=/teams/ ><div class='IssueTitle large' style=color:#8e65bf><div class='t-Caps meta'><span>Issue 11</span> <span>November 2019</span></div><h1 class='t-IssueTitle title'>Teams</h1></div></a></li><li><a href=/testing/ ><div class='IssueTitle large' style=color:#e89e00><div class='t-Caps meta'><span>Issue 10</span> <span>August 2019</span></div><h1 class='t-IssueTitle title'>Testing</h1></div></a></li><li><a href=/open-source/ ><div class='IssueTitle large' style=color:#4a5ad3><div class='t-Caps meta'><span>Issue 9</span> <span>May 2019</span></div><h1 class='t-IssueTitle title'>Open Source</h1></div></a></li><li><a href=/internationalization/ ><div class='IssueTitle large' style=color:#c096ca><div class='t-Caps meta'><span>Issue 8</span> <span>February 2019</span></div><h1 class='t-IssueTitle title'>Internationalization</h1></div></a></li><li><a href=/security/ ><div class='IssueTitle large' style=color:#4dbac5><div class='t-Caps meta'><span>Issue 7</span> <span>October 2018</span></div><h1 class='t-IssueTitle title'>Security</h1></div></a></li><li><a href=/documentation/ ><div class='IssueTitle large' style=color:#f5684d><div class='t-Caps meta'><span>Issue 6</span> <span>August 2018</span></div><h1 class='t-IssueTitle title'>Documentation</h1></div></a></li><li><a href=/programming-languages/ ><div class='IssueTitle large' style=color:#d69336><div class='t-Caps meta'><span>Issue 5</span> <span>April 2018</span></div><h1 class='t-IssueTitle title'>Programming Languages</h1></div></a></li><li><a href=/energy-environment/ ><div class='IssueTitle large' style=color:#d6658e><div class='t-Caps meta'><span>Issue 4</span> <span>February 2018</span></div><h1 class='t-IssueTitle title'>Energy & Environment</h1></div></a></li><li><a href=/development/ ><div class='IssueTitle large' style=color:#53a88e><div class='t-Caps meta'><span>Issue 3</span> <span>October 2017</span></div><h1 class='t-IssueTitle title'>Development</h1></div></a></li><li><a href=/cloud/ ><div class='IssueTitle large' style=color:#707aed><div class='t-Caps meta'><span>Issue 2</span> <span>July 2017</span></div><h1 class='t-IssueTitle title'>Cloud</h1></div></a></li><li><a href=/on-call/ ><div class='IssueTitle large' style=color:#ef766e><div class='t-Caps meta'><span>Issue 1</span> <span>April 2017</span></div><h1 class='t-IssueTitle title'>On-Call</h1></div></a></li></ul></div></div></div><footer class=PageFooter><div class='u-Container ContentBody small'><svg style=display:none><symbol id=twitterIcon viewBox='0 0 32 32'><path d='M32.1 6c-1.2.5-2.5.9-3.8 1 1.4-.8 2.4-2.1 2.9-3.6-1.3.8-2.7 1.3-4.2 1.6a6.8 6.8 0 0 0-4.8-2c-3.6 0-6.6 3-6.6 6.6 0 .5.1 1 .2 1.5-5.5-.3-10.4-3-13.6-6.9-.6 1-.9 2.1-.9 3.3 0 2.3 1.2 4.3 2.9 5.5-1.1 0-2.1-.3-3-.8v.1c0 3.2 2.3 5.9 5.3 6.5-.6.2-1.1.2-1.7.2-.4 0-.8 0-1.2-.1.8 2.6 3.3 4.5 6.2 4.6-2.3 1.8-5.1 2.8-8.2 2.8-.5 0-1.1 0-1.6-.1C2.9 28 6.3 29 10 29c12.1 0 18.7-10 18.7-18.7v-.9c1.4-.9 2.5-2 3.4-3.4z' fill=currentColor /></symbol><symbol id=facebookIcon viewBox='0 0 32 32'><path d='M30.2 0H1.8C.8 0 0 .8 0 1.8v28.5c0 1 .8 1.8 1.8 1.8h15.3V19.6h-4.2v-4.8h4.2v-3.6c0-4.1 2.5-6.4 6.2-6.4 1.8 0 3.3.2 3.7.2v4.3h-2.6c-2 0-2.4 1-2.4 2.4v3.1h4.8l-.6 4.8H22V32h8.2c1 0 1.8-.8 1.8-1.8V1.8c0-1-.8-1.8-1.8-1.8z' fill=currentColor /></symbol><symbol id=rssIcon viewBox='0 0 32 32'><path d='M10.7 25.6c0 2.4-2 4.4-4.4 4.4S2 28 2 25.6s2-4.4 4.4-4.4 4.3 2 4.3 4.4zM6.1 2c-.6 0-1.3 0-2 .1-1.3.1-2.2 1.2-2.1 2.4.1 1.2 1.2 2.2 2.4 2.1.6 0 1.1-.1 1.6-.1 10.7 0 19.4 8.7 19.4 19.4 0 .5 0 1-.1 1.6-.1 1.2.8 2.3 2.1 2.4h.2c1.2 0 2.1-.9 2.2-2.1.1-.7.1-1.4.1-2C30 12.7 19.3 2 6.1 2zm-.7 9.6c-.4 0-.8 0-1.3.1-1.3.1-2.2 1.2-2.1 2.5.1 1.3 1.2 2.2 2.5 2.1h.9c5.7 0 10.4 4.7 10.4 10.4v.9c-.1 1.3.8 2.4 2.1 2.5h.2c1.2 0 2.2-.9 2.3-2.1 0-.4.1-.8.1-1.2-.1-8.4-6.9-15.2-15.1-15.2z' fill=currentColor /></symbol><symbol id=linkedInIcon viewBox='0 0 32 32'><path d='M29.6,0H2.4C1.1,0,0,1,0,2.3v27.4C0,31,1.1,32,2.4,32h27.3c1.3,0,2.4-1,2.4-2.3V2.3C32,1,30.9,0,29.6,0z M9.5,27.3H4.7V12 h4.8V27.3z M7.1,9.9c-1.5,0-2.8-1.2-2.8-2.8c0-1.5,1.2-2.8,2.8-2.8c1.5,0,2.8,1.2,2.8,2.8C9.9,8.7,8.6,9.9,7.1,9.9z M27.3,27.3 h-4.7v-7.4c0-1.8,0-4-2.5-4c-2.5,0-2.8,1.9-2.8,3.9v7.6h-4.7V12H17v2.1h0.1c0.6-1.2,2.2-2.5,4.5-2.5c4.8,0,5.7,3.2,5.7,7.3V27.3z' fill=currentColor /></symbol></svg><div class='column main'><section class=social><a href=https://twitter.com/incrementmag class=twitter><svg viewBox='0 0 32 32'><use xlink:href=#twitterIcon x=0 y=0></use></svg> <span>@incrementmag</span> </a><a href=https://facebook.com/incrementmag class=facebook><svg viewBox='0 0 32 32'><use xlink:href=#facebookIcon x=0 y=0></use></svg> <span>incrementmag</span> </a><a href=/feed.xml class=rss><svg viewBox='0 0 32 32'><use xlink:href=#rssIcon x=0 y=0></use></svg> <span>RSS Feed</span></a></section><section><h4>About</h4><p><em>Increment</em> is a print and digital magazine about how teams build and operate software systems at scale. <a href=/about/ >Learn more</a></p></section><section><h4>Work with us</h4><p>Interested in joining the team at Stripe? <a href=https://stripe.com/jobs>View job openings</a></p></section></div><p class='column copyright'><span>&copy; 2022 <em>Increment</em></span> <a href=https://stripe.com>Published by Stripe</a> <a href=https://stripe.com/privacy/media-policy>Privacy policy</a></p></div></footer></body></html>