Files
nexus/sreweekly/articles/68/06-on-call-at-any-size.html
2026-09-12 17:23:01 +08:00

19 lines
42 KiB
HTML
Raw Blame History

This file contains invisible Unicode characters
This file contains invisible Unicode characters that are indistinguishable to humans but may be processed differently by a computer. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
<!doctype html><html><head><meta charset=utf-8><title>On-call at any size – Increment: On-Call</title><meta name=description content='We take a close look at how to make on-call work at any scale, sharing industry best practices that apply to companies at any size, from tiny startups in garages to companies the size of Amazon, Facebook, and Google.'><link rel=canonical href=http://localhost:3000/on-call/on-call-at-any-size/ ><link rel=apple-touch-icon-precomposed href=/img/icon-571805a1.png><meta property=og:title content='On-call at any size – Increment: On-Call'><meta property=og:url content=http://localhost:3000/on-call/on-call-at-any-size/ ><meta property=og:description content='We take a close look at how to make on-call work at any scale, sharing industry best practices that apply to companies at any size, from tiny startups in garages to companies the size of Amazon, Facebook, and Google.'><meta property=og:image content='https://images.ctfassets.net/3njn2qm7rrbs/8REmHSiJuGJOi8ZP09Inu/c7245e42aed6c587b824caa299e76db7/collaboration.png?w=1000'><meta name=twitter:card content=summary_large_image><meta name=twitter:image content='https://images.ctfassets.net/3njn2qm7rrbs/8REmHSiJuGJOi8ZP09Inu/c7245e42aed6c587b824caa299e76db7/collaboration.png?w=1000'><meta name=twitter:site content=@IncrementMag><meta name=twitter:title content='On-call at any size – Increment: On-Call'><meta name=twitter:description content='We take a close look at how to make on-call work at any scale, sharing industry best practices that apply to companies at any size, from tiny startups in garages to companies the size of Amazon, Facebook, and Google.'><link rel=alternate type=application/rss+xml title=Increment href=/feed.xml><meta name=viewport content='width=device-width,initial-scale=1'><link rel=preload href=/fonts/baton-turbo/400-30a55d66.woff2 as=font type=font/woff2 crossorigin=anonymous><link rel=preload href=/fonts/baton-turbo/500-1603c0e8.woff2 as=font type=font/woff2 crossorigin=anonymous><link rel=preload href=/fonts/tiempos-text/400-c4810745.woff2 as=font type=font/woff2 crossorigin=anonymous><link rel=preload href=/fonts/tiempos-head/700-383ede62.woff2 as=font type=font/woff2 crossorigin=anonymous><link rel=stylesheet type=text/css href=/css/bundle-289d885f.css><link rel=stylesheet type=text/css href=/css/issues/1-a2738f50.css><script>// Don't fade art if it loads ~instantly
setTimeout(()=>{document.documentElement.classList.add('fadeArt')},250);</script><script defer src=/js/defer-737fbd90.js></script><script>const INCREMENT_META={issueNumber:1,issueSlug:'on-call',articleSlug:'on-call-at-any-size'};</script></head><body class='Issue_on-call Article_on-call-at-any-size'><nav class=PageNav><div class=u-Container><div class=column><h1 class=logo><a href=/ ><img src=/img/logo-ae2c55d5.svg alt=Increment></a></h1><a class=out-now href=https://store.increment.com/ style=color:#595959><div class='IssueTitle tiny'><div class='t-Caps meta tiny'><span>NEW</span></div><h3 class='t-IssueTitle title'>Buy the print edition</h3></div></a><ul class=nav><li><a href=/issues/ ><span>Issues</span></a></li><li><a href=/topics/ ><span>Topics</span></a></li><li><a href=https://store.increment.com/ ><span>Store</span></a></li><li><a href=/about/ ><span>About</span></a></li></ul></div></div></nav><div class=ArticlePage itemscope itemtype=http://schema.org/Article><script type=application/ld+json>{
"@context": "http://schema.org",
"@type": "Article",
"headline": "On-call at any size",
"image": " https://images.ctfassets.net/3njn2qm7rrbs/8REmHSiJuGJOi8ZP09Inu/c7245e42aed6c587b824caa299e76db7/collaboration.png?w&#x3D;1000",
"datePublished": "Thu, 13 Apr 2017 09:00:00 GMT",
"dateModified": "Thu, 13 Apr 2017 09:00:00 GMT",
"publisher": {
"@type": "Organization",
"name": "Increment",
"logo": {
"@type": "ImageObject",
"url": "https://increment.com/img/logo.png"
}
},
"description": "We take a close look at how to make on-call work at any scale, sharing industry best practices that apply to companies at any size, from tiny startups in garages to companies the size of Amazon, Facebook, and Google.",
"mainEntityOfPage": "http://localhost:3000/on-call/on-call-at-any-size/"
}</script><header class='u-Container ArticleHeader'><div class='u-Container art'><div class=u-Art><div class=placeholder style='background:#f2f2f2;background-position:center 70%'></div><picture><source srcset='https://images.ctfassets.net/3njn2qm7rrbs/8REmHSiJuGJOi8ZP09Inu/c7245e42aed6c587b824caa299e76db7/collaboration.png?w=2000 2000w, https://images.ctfassets.net/3njn2qm7rrbs/8REmHSiJuGJOi8ZP09Inu/c7245e42aed6c587b824caa299e76db7/collaboration.png?w=1000 1000w, https://images.ctfassets.net/3njn2qm7rrbs/8REmHSiJuGJOi8ZP09Inu/c7245e42aed6c587b824caa299e76db7/collaboration.png?w=500 500w, https://images.ctfassets.net/3njn2qm7rrbs/8REmHSiJuGJOi8ZP09Inu/c7245e42aed6c587b824caa299e76db7/collaboration.png?w=100 100w' sizes=100vw><img class=j-SmoothLoad src='https://images.ctfassets.net/3njn2qm7rrbs/8REmHSiJuGJOi8ZP09Inu/c7245e42aed6c587b824caa299e76db7/collaboration.png?w=1000' sizes=100vw alt='On-call at any size' style='object-position:center 70%'></picture></div></div><div class=column><div class=u-Grid><div class=main><h4 class='t-Byline byline large'><a href=#authors class=j-SmoothScroll itemprop=author><span>Increment Staff</span></a></h4><h1 class='t-TitleSerif large title' itemprop=name>On-call at any size</h1><div class='t-BodySans large intro' itemprop=description>We take a close look at how to make on-call work at any scale, sharing industry best practices that apply to companies at any size, from tiny startups in garages to companies the size of Amazon, Facebook, and&nbsp;Google.</div></div><a class=issue href=/on-call/ ><span class='t-Caps tiny part-of'>Part of</span><div class='IssueTitle small' style=color:#ef766e><div class='t-Caps meta'><span>Issue 1</span> <span>April 2017</span></div><h2 class='t-IssueTitle title'>On-Call</h2></div></a></div></div></header><div class='u-Container ArticleContent'><article class='ContentBody column' itemprop=articleBody><div class=ArticleLayout><p>Creating and running an on-call rotation is the very first step in building truly reliable infrastructure at any scale, but it’s far from the last. As a company grows and scales, the systems that it builds and runs grow in complexity and require more sophisticated on-call practices. While there is no universal approach, there are industry best practices for setting up on-call and incident response at any and every size.</p><p>In the sections that follow, we take a close look at how to make on-call work at any scale. We’ll examine how to design, support, and empower on-call and incident response for each size, starting with how tiny garage startups with a handful of engineers can run an on-call rotation and making our way up to best practices for companies the size of Amazon, Facebook, or Google.</p><h2>Garage startup (1-10 employees)</h2><p>Even the smallest of startups needs to ensure that their products are readily available to users, and building an on-call rotation that sets the new company up for future success can help them meet this goal. Having an on-call rotation at this scale is essential for understanding whether or not the product is actually working as expected and promised, and for fixing it when it breaks. It’s all well and good for the servers to go down when the only users are friends and family, but after a customer has paid for the service and when investors want to know about the service’s performance, letting the servers go down for two hours during someone’s date night is not going to cut it.</p><p>Startups bootstrap their on-call rotations out of three primitives. The first is a <strong>notification tool</strong> that will notify them when there is a problem, like <a href=https://www.pagerduty.com/ target=_blank rel='noopener noreferrer'>PagerDuty</a>. The second is a <strong>collaboration tool</strong> engineers can use for communication during outages, such as <a href=https://slack.com/ target=_blank rel='noopener noreferrer'>Slack</a>, Hipchat or IRC. The final primitive is a <strong>monitoring</strong> <strong>system</strong> to know when there is a problem, which can be a bit harder to set up than the first two so we’ll take a closer look at it.</p><p>Good, accurate monitoring is made up of three pieces: <strong>healthchecks</strong>, <strong>metrics</strong>, and <strong>dashboards</strong>. <strong>Heathchecks</strong> should be run against both services and servers, verifying they are reachable and responding. Many startups have unfortunately managed to create high-availability, responsive, and scalable ways to serve error pages, so true confidence that the business is operating normally requires collecting <strong>metrics</strong>. The next step is to create <strong>dashboards</strong> of these metrics so that engineers can visually verify the system is behaving normally (and spot correlated events, like a botched deploy, when it is not). These needs can be met by a variety of hosted and open-source options. On the hosted side, various platforms such as <a href=https://www.datadoghq.com/ target=_blank rel='noopener noreferrer'>Datadog</a> provide healthchecks, metrics and dashboards. If a company decides to go the open-source route, a combination of tools such as <a href=https://www.nagios.org/ target=_blank rel='noopener noreferrer'>Nagios</a> and <a href=https://graphiteapp.org/ target=_blank rel='noopener noreferrer'>Graphite</a> will do the trick.</p><p>There are significant advantages to being small. When something does go wrong, it tends to be quite clear who broke what, and the entire team is familiar enough with the system to be able to fix problems. The business impact of something breaking is also smaller: early adopters have a higher tolerance for outages, and the absolute number of users affected is much, much lower than when operating at scale. The culture and habits adopted when the company is established will grow along with the company it as it scales, making this the best window of opportunity to foster and evolve a culture of reliability.</p><p>However, being scrappy is not without its disadvantages. Hardening alert configurations and expanding them to detect a wide array of potential outages is labor intensive work, and it’s an ongoing challenge to prioritize better monitoring over feature development. The consequences of a small team expand beyond prioritization, and providing 24/7 coverage with only two or three engineers can be a punishing experience.</p><h3>Key practices</h3><ul><li><p><strong>Notification tools</strong> like <a href=https://www.pagerduty.com/ target=_blank rel='noopener noreferrer'>PagerDuty</a> so that problems can be surfaced to the team quickly. <strong>Collaboration tools</strong> to coordinate and communicate during outages, like <a href=https://slack.com/ target=_blank rel='noopener noreferrer'>Slack</a>, <a href=https://www.hipchat.com/ target=_blank rel='noopener noreferrer'>Hipchat</a> or IRC.</p></li><li><p><strong>Healthchecks</strong> to automatically detect degradation of services or servers: <a href=https://www.nagios.org/ target=_blank rel='noopener noreferrer'>Nagios</a>, <a href=https://www.datadoghq.com/ target=_blank rel='noopener noreferrer'>Datadog</a>.</p></li><li><p><strong>Metrics</strong> and <strong>Dashboards</strong> to establish baselines for system behavior and allow operators to spot anomalies that healthchecks miss, such as <a href=https://graphiteapp.org/ target=_blank rel='noopener noreferrer'>Graphite</a>, <a href=https://www.datadoghq.com/ target=_blank rel='noopener noreferrer'>Datadog</a>.</p></li></ul><h2>A real office (11-100 employees)</h2><p>As a company grows, so do the demands of its infrastructure, its operational workload, and its on-call and incident response. For many companies, this is when dozens of servers grow to hundreds, and, in the interests of scalability, a monolithic service may be broken into multiple microservices.</p><p>The engineering team will be large enough at this point to expand the on-call rotation; a <strong>sustainable rotation</strong> might require, for example, eight engineers who run a two layer rotation while restricting time on-call to one week per month. It is critical that incidents are routed to the rotation–and not consistently bypassed to be solved by a subset of experienced engineers–to avoid burning out key engineers, expand understanding of the systems, and train the seed members for expanding to multiple rotations. Moving from a shared on-call rotation to <strong>multiple rotations</strong> is a key scaling challenge, requiring active attention to ensure practices remain healthy and consistent.</p><p>The consequences of not having business continuity become frighteningly real, leading to an investment in <strong>disaster recovery</strong> to be able to recover from the dire and unexpected. More than a one-time implementation, recovery must be validated at least on an annual basis, initially as a best practice, and eventually for compliance purposes.</p><p>At this size, it’s no longer effective to inspect logs on individual servers, and it’s time to adopt a centralized <strong>log aggregator</strong>, such as <a href=https://www.elastic.co/products target=_blank rel='noopener noreferrer'>ELK</a> or <a href=https://www.splunk.com/ target=_blank rel='noopener noreferrer'>Splunk</a>. Similarly, an <strong>error aggregator (</strong>like <a href=https://sentry.io/ target=_blank rel='noopener noreferrer'>Sentry</a>)<a href=https://sentry.io/ target=_blank rel='noopener noreferrer'> </a>will make it easier to understand where errors are coming from. Deployment tooling and strategies should be expanded to support easy <strong>rollbacks</strong>, and broken deployments should be rolled back rather than repaired in the production environment.</p><p>The tooling adopted at this phase is the same tooling deployed at companies that have scaled to thousands of engineers and major revenue. Despite this fact, the infrastructure footprint is still small enough to allow for major shifts, such as moving from a physical datacenter to the cloud, or from “manual orchestration” to programmatic orchestration like <a href=https://kubernetes.io/ target=_blank rel='noopener noreferrer'>Kubernetes</a> or <a href=http://mesos.apache.org/ target=_blank rel='noopener noreferrer'>Mesos</a>. Making large-scale changes to the infrastructure becomes more difficult when the company grows larger, so companies at this scale have a unique opportunity to quickly test out new infrastructure and set themselves up for future scalability and reliability—an opportunity that won’t be there for long.</p><p>The process surrounding on-call will start to evolve as well. <strong>Incident classification</strong> is the practice of classifying incidents according to their user impact, often labeling a minor incident as “Level 4” and a widespread outage as “Level 1.” <strong>Postmortems</strong> to discuss the timeline and impact of large incidents will create a feedback loop to guide learning about system reliability: <a href=https://codeascraft.com/2016/11/17/debriefing-facilitation-guide/ target=_blank rel='noopener noreferrer'>Etsy’s Debriefing Facilitation Guide</a> describes how to run healthy postmortems that create a permanent culture of learning from failures.</p><p>Managing an incident will become complex enough to require two or more individuals collaborating, and clearly defined roles and responsibilities makes collaboration during stressful times easier. In particular, defining the role of <strong>incident commander</strong>, who is not involved in debugging but instead coordinates communication is valuable. One approach to incident commanding is covered in <a href=https://response.pagerduty.com/training/incident_commander/ target=_blank rel='noopener noreferrer'>Pagerduty’s Incident Response Documentation</a>.</p><h3>Key practices</h3><ul><li><p><strong>Sustainable rotation</strong> comprised of at least eight engineers.</p></li><li><p><strong>Multiple rotations</strong> make it likely the person paged can debug the impacted system.</p></li><li><p><strong>Disaster recovery</strong> will make it possible to recover from the dire and unexpected.</p></li><li><p><strong>Log aggregator</strong> to collect logs from all servers, such as <a href=https://www.elastic.co/products target=_blank rel='noopener noreferrer'>ELK</a> or <a href=https://www.splunk.com/ target=_blank rel='noopener noreferrer'>Splunk</a>.</p></li><li><p><strong>Error aggregator</strong> to track the source of errors, like <a href=https://sentry.io/ target=_blank rel='noopener noreferrer'>Sentry</a>.</p></li><li><p><strong>Rollback</strong> broken deployments, instead of repairing in production.</p></li><li><p><strong>Incident classification</strong> will clarify the user impact of outages.</p></li><li><p><strong>Postmortems</strong> after large incidents will create a feedback loop to guide learning about system reliability.</p></li><li><p><strong>Incident commanders</strong> who handles coordination but don’t debug.</p></li></ul><h2>Beyond Dunbar’s Number (101-1,000 employees)</h2><p>As a company grows and its business becomes increasingly valuable, the stress created by incidents increases in tandem. There will be engineers at the company who remember back when two hours of downtime was smoothed over with a company T-shirt to the one affected customer, while now downtime can cost tens of thousands of dollars per minute.</p><p>Fortunately, as a larger company there are more resources and tools to offset the stress of being on-call. The first–obvious, but greatly underutilized–is explicitly <strong>training</strong> engineers about on-call process and tooling. Beyond a presentation, the best training involves running “<strong>game days</strong>“ with manually triggered failures or generated load to create realistic opportunities for debugging.</p><p>At this point, each team should be responsible for the software they write, with a <strong>backstop rotation</strong> of your most experienced engineers for exceptional circumstances. An effective combination of author ownership with a backstop is necessary, as increasingly no individual engineer has a complete mental map of how the entire system works. There are some exceptions—Airbnb, for example, relies on an elite volunteer rotation for their backstop, and Slack continues to rely on their operations team for primary on-call.</p><p>This is also when there are enough resources to selectively move beyond generic open-source solutions, starting to roll out specialized tooling. This includes <strong>request tracing</strong> tooling such as <a href=http://zipkin.io/ target=_blank rel='noopener noreferrer'>Zipkin</a> or <a href=http://lightstep.com/ target=_blank rel='noopener noreferrer'>Lightstep</a>, which will greatly simplify debugging, particularly in environments where a monolithic service has been unwound into a swarm of microservices. It’s also time to write an <strong>ownership metadata service</strong> that makes it possible to have a centralized source of truth for who owns what; this metadata will be the backbone of routing alerts, rolling out cost accounting and frantic lookups during outages. Moving past a wiki documenting the steps to take during an outage, an <strong>incident registration tool</strong> to coordinate the incident process from a single place: sending alerts and notifications in chat, and filing a ticket to create a postmortem. Most companies start making a major investment in <strong>runbooks</strong>, ideally including a link to a runbook in every alert that is sent.</p><p>In addition to tools for managing incidents, the disruption and financial impact will drive increasing focus on avoiding incidents to begin with. This includes more rigorous deployment strategies, such as deploying to a small number of <strong>canary hosts</strong> before deploying widely, and <strong>incremental deployments</strong> that abort automatically if error rates elevate unexpectedly. Software will be automatically tested in a quality assurance or <strong>staging environment</strong>, running integration tests against every deployment before it is promoted to your production environment.</p><p>For outages that can’t be prevented entirely, there are mechanisms to limit the blast radius, such as <strong>rate limiters</strong> which shed excessive traffic and <strong>circuit breakers</strong> to prevent cascading failures from failing downstream services.</p><p>Process should explicitly include steps to <strong>communicating externally</strong>, building trust with customers and users by giving timely visibility into issues. Some companies, like Gitlab, take this to the extent of providing public postmortems. At this phase, it’s important to have clearly defined <strong>reliability metrics</strong> to measure and improve against, particularly for PaaS and SaaS companies.</p><p>Some challenges are exacerbated by size, and alert escalations across teams–particularly for legacy systems–will become increasingly frequent and a source of contention.</p><h3>Key practices</h3><ul><li><p><strong>Training</strong> for engineers on on-call process and tooling.</p></li><li><p><strong>“Game days</strong>“ to create realistic opportunities for debugging.</p></li><li><p><strong>Backstop rotation</strong> of your most experienced engineers for exceptional circumstances.</p></li><li><p><strong>Request tracing</strong> to debug across service boundaries, such as <a href=http://zipkin.io/ target=_blank rel='noopener noreferrer'>Zipkin</a> or <a href=http://lightstep.com/ target=_blank rel='noopener noreferrer'>Lightstep</a>.</p></li><li><p><strong>Ownership metadata service</strong> maintaining a centralized source of truth for who owns what.</p></li><li><p><strong>Incident registration tool</strong> to coordinate the incident process from a single place.</p></li><li><p><strong>Runbooks</strong> for common issues, linked to every delivered alert.</p></li><li><p><strong>Canary hosts</strong> to validate software in production before deploying widely.</p></li><li><p><strong>Incremental deployments</strong> that abort automatically if error rates elevate unexpectedly.</p></li><li><p><strong>Staging environment</strong> that automatically test software before it reaches production.</p></li><li><p><strong>Rate limiters</strong> to shed excessive traffic.</p></li><li><p><strong>Circuit breakers</strong> to prevent cascading failures from failing downstream services.</p></li><li><p><strong>Communicating externally</strong> to build trust with customers and users.</p></li><li><p><strong>Reliability metrics</strong> to measure and improve against.</p></li></ul><h2>Large companies (1,000-10,000 employees)</h2><p>At this scale a company has become truly massive, moving from a handful of on-call rotations to dozens of them. Outages are immensely expensive, and when one does occur, prioritize <strong>recovery over understanding</strong>, as it’s too expensive to debug in production.</p><p>Servers or VMs are sufficiently numerous that even acknowledging alerts issued from individual servers is infeasible, requiring investment into <strong>alert deduplication</strong>, with each root cause triggering a single alert. The frequency of untuned alerts firing will overwhelm on-call engineers, leading to tracking <strong>alert metrics</strong>, which guide tuning and fixing alerts that fire frequently: you aim for every alert to be an <strong>actionable alert</strong>.</p><p>Failure isn’t unusual, it’s a constant. Four servers breaking will go from an exceptional day to a normal Tuesday; switches and entire datacenters will start failing with frequency. Don’t be content with entropy—start to actively <strong>inject failures</strong> into systems to ensure they fail properly using tools like <a href=https://github.com/Netflix/SimianArmy/wiki/Chaos-Monkey target=_blank rel='noopener noreferrer'>Chaos Monkey</a>. Because failures are too frequent to remediate manually, <strong>deprecate runbooks</strong> and move towards <strong>automated remediation</strong> as widely as possible. A <strong>multi-region</strong> or multi-datacenter strategy makes it possible to continue providing service, even under the most dire circumstances. As a consequence of this investment into tooling, humans can increasingly delegate on-call responsibilities to computers, only becoming involved under exceptional circumstances.</p><p>Things change behind the technology as well. 24/7 coverage is replaced with a <strong>follow the sun</strong> rotation to reduce midnight calls, increasing quality of life, and also reducing time to remediation. An <strong>executive sponsor</strong> for reliability talks about why reliability matters at company all-hands meetings. Incidents become frequently enough to hire a <strong>technical program manager</strong> to coordinate the postmortem process.</p><h3>Key practices</h3><ul><li><p><strong>Recovery over understanding</strong>, it’s too expensive to debug in production.</p></li><li><p><strong>Alert deduplication</strong> so each root cause triggers a single alert.</p></li><li><p><strong>Alert metrics</strong> to guide tuning and fixing alerts that fire frequently.</p></li><li><p><strong>Actionable alerts</strong> are the only alerts.</p></li><li><p><strong>Inject failures</strong> into systems to ensure they fail properly, using tools like <a href=https://github.com/Netflix/SimianArmy/wiki/Chaos-Monkey target=_blank rel='noopener noreferrer'>Chaos Monkey</a>.</p></li><li><p><strong>Deprecate runbooks</strong> and move towards <strong>automated remediation</strong>.</p></li><li><p><strong>Multi-region</strong> provides real-time business continuity.</p></li><li><p><strong>Follow the sun</strong> rotation to reduce midnight calls.</p></li><li><p><strong>Executive sponsor</strong> for reliability to champion efforts.</p></li><li><p><strong>Technical program manager</strong> to coordinate the postmortem process.</p></li></ul><h2>Leviathan (10,000+ employees)</h2><p>Companies at this size operate in rarefied air, having enough engineers to build any tool imaginable, and to <strong>automate</strong> away as much of the operational workload as possible. The overall system’s complexity is beyond comprehension of any individual, causing the benefits of having written part of the system to be surpassed by the benefits of experience with on-call tooling. Consequently, it’s time to move away from the previously effective “who wrote it, is on-call for it” model to a <strong>centralized on-call</strong>, operated by specialized engineers.</p><h3>Key practices</h3><ul><li><p>True <strong>automation</strong> of the operational workload for stable parts of the system.</p></li><li><p><strong>Centralized on-call</strong> operated by specialized engineers.</p></li></ul></div></article></div><div class='u-Container ArticleFooter'><div class=column><div class='u-Grid ContentBody small content'><div class=authors id=authors><div><div class=photo><figure style='background-image:url(https://images.ctfassets.net/3njn2qm7rrbs/3buorQfLVwQiD58sXNILJE/d59d2cbaf87ce0d40e499f26f586ad77/pager.png?w=500)'></figure></div><div class=text><h4>Artwork by</h4><p><b>Mark Conlan</b></p><p><a href=https://markconlan.com/ target=_blank>markconlan.com</a></p></div></div></div><div class=topics><div class=text><h4>Topics</h4><p><a href=/topics/scaling/ >Scaling &amp; Growth</a></p><p><a href=/topics/guides/ >Guides &amp; Best Practices</a></p></div></div></div></div></div></div><div class=SubscribeBox><a id=newsletter class=anchor href=#newsletter></a><div class='ContentBody inverted store'><div class=u-Container><div class='u-Grid column'><div class=box style=background:#4c70b1><div class=text><h2>Buy the print edition</h2><p class=j-TextBalance>Visit the Increment Store to purchase print issues.</p><p><a class='t-Caps u-Arrow' href=https://store.increment.com/ >Store</a></p></div><a href=https://store.increment.com/ class=magazine><figure class=j-MaskedImage data-fill=/art/19/19-cutout-1000-7ffb5dba.png data-mask=/art/19/fill-cms-1000-83522f38.png></figure></a></div></div></div></div><div class='ContentBody email'><div class=u-Container><div class='u-Grid column'><div class='j-EmailForm box' style='box-shadow:0 -5px 0 #4c70b1'></div></div></div></div></div><div class=ContinueReading><div class=u-Container><div class=column><h2 class='t-Caps xlarge'>Continue Reading</h2><ul class='u-Grid articles'><li class='ArticleBlock footer' style=color:#ef766e><a class=j-Preload href=/on-call/ask-an-expert/ ><div class='IssueTitle tiny' style=color:#ef766e><div class='t-Caps meta'><span>1</span></div><h3 class='t-IssueTitle title'>On-Call</h3></div><div class=text><div class=head><h4 class='t-Byline byline'><span>Increment Staff</span></h4><h3 class='t-TitleSans title'>Ask an expert: How should startups approach on-call and incident response?</h3></div><div class='t-BodySerif small intro'>We asked several industry experts if they had any advice for small companies who are just starting to set up their on-call and incident response processes, and here’s what they&nbsp;said.</div></div></a></li><li class='ArticleBlock footer' style=color:#ef766e><a class=j-Preload href=/on-call/when-the-pager-goes-off/ ><div class='IssueTitle tiny' style=color:#ef766e><div class='t-Caps meta'><span>1</span></div><h3 class='t-IssueTitle title'>On-Call</h3></div><div class=text><div class=head><h4 class='t-Byline byline'><span>Increment Staff</span></h4><h3 class='t-TitleSans title'>What happens when the pager goes off?</h3></div><div class='t-BodySerif small intro'>To discover the state of incident response across the tech industry, we surveyed over thirty industry leaders (including Amazon, Dropbox, Facebook, Google, and Netflix) about their incident response processes.</div></div></a></li><li class='ArticleBlock footer' style=color:#ef766e><a class=j-Preload href=/on-call/who-owns-on-call/ ><div class='IssueTitle tiny' style=color:#ef766e><div class='t-Caps meta'><span>1</span></div><h3 class='t-IssueTitle title'>On-Call</h3></div><div class=text><div class=head><h4 class='t-Byline byline'><span>Increment Staff</span></h4><h3 class='t-TitleSans title'>Who owns on-call?</h3></div><div class='t-BodySerif small intro'>Industry leaders like Google, Amazon, Dropbox, Spotify, and Netflix are putting developers on-call for their software. We spoke to these companies to find out&nbsp;why.</div></div></a></li><li class='ArticleBlock footer' style=color:#ef766e><a class=j-Preload href=/on-call/the-benefits-of-transparency/ ><div class='IssueTitle tiny' style=color:#ef766e><div class='t-Caps meta'><span>1</span></div><h3 class='t-IssueTitle title'>On-Call</h3></div><div class=text><div class=head><h4 class='t-Byline byline'><span>Increment Staff</span></h4><h3 class='t-TitleSans title'>The benefits of transparency: Interview with Sytse “Sid” Sijbrandij, CEO of GitLab</h3></div><div class='t-BodySerif small intro'>We spoke with Sid about GitLab’s recent outage, its aftermath, and the impressive level of transparency GitLab offered the public during the&nbsp;outage.</div></div></a></li><li class='ArticleBlock footer' style=color:#53a88e><a class=j-Preload href=/development/what-its-like-to-be-a-developer-at/ ><div class='IssueTitle tiny' style=color:#53a88e><div class='t-Caps meta'><span>3</span></div><h3 class='t-IssueTitle title'>Development</h3></div><div class=text><div class=head><h4 class='t-Byline byline'><span>Increment Staff</span></h4><h3 class='t-TitleSans title'>What it’s like to be a developer at …</h3></div><div class='t-BodySerif small intro'>From popular tools to code review, deployment to daily life, here’s a look at the developer experience at Slack, Lyft, DigitalOcean, and&nbsp;more.</div></div></a></li><li class='ArticleBlock footer' style=color:#4a5ad3><a class=j-Preload href=/open-source/open-source-at-scale/ ><div class='IssueTitle tiny' style=color:#4a5ad3><div class='t-Caps meta'><span>9</span></div><h3 class='t-IssueTitle title'>Open Source</h3></div><div class=text><div class=head><h4 class='t-Byline byline'><span>Increment Staff</span></h4><h3 class='t-TitleSans title'>Open source at scale</h3></div><div class='t-BodySerif small intro'>Technical leaders at Microsoft, Kickstarter, DigitalOcean, and Red Hat answer questions about when and why to opt for OSS, how open source has influenced their organizations, and its future role in corporations.</div></div></a></li><li class='ArticleBlock footer' style=color:#e89e00><a class=j-Preload href=/testing/testing-at-scale/ ><div class='IssueTitle tiny' style=color:#e89e00><div class='t-Caps meta'><span>10</span></div><h3 class='t-IssueTitle title'>Testing</h3></div><div class=text><div class=head><h4 class='t-Byline byline'><span>Increment Staff</span></h4><h3 class='t-TitleSans title'>Testing at scale</h3></div><div class='t-BodySerif small intro'>Rob Zuber (CircleCI), Greg Bell and Claudiu Coman (Hootsuite), and Scott Triglia (Yelp) talk test suite times, manual versus automated testing, and how to build testing infrastructure.</div></div></a></li><li class='ArticleBlock footer' style=color:#8e65bf><a class=j-Preload href=/teams/engineering-teams-at-scale/ ><div class='IssueTitle tiny' style=color:#8e65bf><div class='t-Caps meta'><span>11</span></div><h3 class='t-IssueTitle title'>Teams</h3></div><div class=text><div class=head><h4 class='t-Byline byline'><span>Increment Staff</span></h4><h3 class='t-TitleSans title'>Engineering teams at scale</h3></div><div class='t-BodySerif small intro'>Leaders at 1Password, Atlassian, and Intuit share how they measure success, determine focus, and foster mentorship, sustainability, and collaboration in their organizations.</div></div></a></li><li class='ArticleBlock footer' style=color:#40af9e><a class=j-Preload href=/software-architecture/architecture-at-scale/ ><div class='IssueTitle tiny' style=color:#40af9e><div class='t-Caps meta'><span>12</span></div><h3 class='t-IssueTitle title'>Software Architecture</h3></div><div class=text><div class=head><h4 class='t-Byline byline'><span>Increment Staff</span></h4><h3 class='t-TitleSans title'>Software architecture at scale</h3></div><div class='t-BodySerif small intro'>Leaders at Foursquare, Hulu, and Twitter discuss early architecture decisions, downstream effects, and architectural philosophies.</div></div></a></li></ul><h3 class='t-Caps large'>Explore Topics</h3><ul class='t-BodySerif large topics'><li><a href=/topics/learn/ >Learn Something New</a></li><li><a href=/topics/scaling/ >Scaling &amp; Growth</a></li><li><a href=/topics/ask-an-expert/ >Ask an Expert</a></li><li><a href=/topics/interviews/ >Interviews &amp; Surveys</a></li><li><a href=/topics/guides/ >Guides &amp; Best Practices</a></li><li><a href=/topics/opinion/ >Essays &amp; Opinion</a></li><li><a href=/topics/culture/ >Workplace &amp; Culture</a></li></ul><h3 class='t-Caps large'>All Issues</h3><ul class=issues><li><a href=/planning/ ><div class='IssueTitle large' style=color:#4c70b1><div class='t-Caps meta'><span>Issue 19</span> <span>November 2021</span></div><h1 class='t-IssueTitle title'>Planning</h1></div></a></li><li><a href=/mobile/ ><div class='IssueTitle large' style=color:#439eab><div class='t-Caps meta'><span>Issue 18</span> <span>August 2021</span></div><h1 class='t-IssueTitle title'>Mobile</h1></div></a></li><li><a href=/containers/ ><div class='IssueTitle large' style=color:#443d79><div class='t-Caps meta'><span>Issue 17</span> <span>May 2021</span></div><h1 class='t-IssueTitle title'>Containers</h1></div></a></li><li><a href=/reliability/ ><div class='IssueTitle large' style=color:#863051><div class='t-Caps meta'><span>Issue 16</span> <span>February 2021</span></div><h1 class='t-IssueTitle title'>Reliability</h1></div></a></li><li><a href=/remote/ ><div class='IssueTitle large' style=color:#29386a><div class='t-Caps meta'><span>Issue 15</span> <span>November 2020</span></div><h1 class='t-IssueTitle title'>Remote</h1></div></a></li><li><a href=/apis/ ><div class='IssueTitle large' style=color:#00afbe><div class='t-Caps meta'><span>Issue 14</span> <span>August 2020</span></div><h1 class='t-IssueTitle title'>APIs</h1></div></a></li><li><a href=/frontend/ ><div class='IssueTitle large' style=color:#5ebe92><div class='t-Caps meta'><span>Issue 13</span> <span>May 2020</span></div><h1 class='t-IssueTitle title'>Frontend</h1></div></a></li><li><a href=/software-architecture/ ><div class='IssueTitle large' style=color:#40af9e><div class='t-Caps meta'><span>Issue 12</span> <span>February 2020</span></div><h1 class='t-IssueTitle title'>Software Architecture</h1></div></a></li><li><a href=/teams/ ><div class='IssueTitle large' style=color:#8e65bf><div class='t-Caps meta'><span>Issue 11</span> <span>November 2019</span></div><h1 class='t-IssueTitle title'>Teams</h1></div></a></li><li><a href=/testing/ ><div class='IssueTitle large' style=color:#e89e00><div class='t-Caps meta'><span>Issue 10</span> <span>August 2019</span></div><h1 class='t-IssueTitle title'>Testing</h1></div></a></li><li><a href=/open-source/ ><div class='IssueTitle large' style=color:#4a5ad3><div class='t-Caps meta'><span>Issue 9</span> <span>May 2019</span></div><h1 class='t-IssueTitle title'>Open Source</h1></div></a></li><li><a href=/internationalization/ ><div class='IssueTitle large' style=color:#c096ca><div class='t-Caps meta'><span>Issue 8</span> <span>February 2019</span></div><h1 class='t-IssueTitle title'>Internationalization</h1></div></a></li><li><a href=/security/ ><div class='IssueTitle large' style=color:#4dbac5><div class='t-Caps meta'><span>Issue 7</span> <span>October 2018</span></div><h1 class='t-IssueTitle title'>Security</h1></div></a></li><li><a href=/documentation/ ><div class='IssueTitle large' style=color:#f5684d><div class='t-Caps meta'><span>Issue 6</span> <span>August 2018</span></div><h1 class='t-IssueTitle title'>Documentation</h1></div></a></li><li><a href=/programming-languages/ ><div class='IssueTitle large' style=color:#d69336><div class='t-Caps meta'><span>Issue 5</span> <span>April 2018</span></div><h1 class='t-IssueTitle title'>Programming Languages</h1></div></a></li><li><a href=/energy-environment/ ><div class='IssueTitle large' style=color:#d6658e><div class='t-Caps meta'><span>Issue 4</span> <span>February 2018</span></div><h1 class='t-IssueTitle title'>Energy & Environment</h1></div></a></li><li><a href=/development/ ><div class='IssueTitle large' style=color:#53a88e><div class='t-Caps meta'><span>Issue 3</span> <span>October 2017</span></div><h1 class='t-IssueTitle title'>Development</h1></div></a></li><li><a href=/cloud/ ><div class='IssueTitle large' style=color:#707aed><div class='t-Caps meta'><span>Issue 2</span> <span>July 2017</span></div><h1 class='t-IssueTitle title'>Cloud</h1></div></a></li><li><a href=/on-call/ ><div class='IssueTitle large' style=color:#ef766e><div class='t-Caps meta'><span>Issue 1</span> <span>April 2017</span></div><h1 class='t-IssueTitle title'>On-Call</h1></div></a></li></ul></div></div></div><footer class=PageFooter><div class='u-Container ContentBody small'><svg style=display:none><symbol id=twitterIcon viewBox='0 0 32 32'><path d='M32.1 6c-1.2.5-2.5.9-3.8 1 1.4-.8 2.4-2.1 2.9-3.6-1.3.8-2.7 1.3-4.2 1.6a6.8 6.8 0 0 0-4.8-2c-3.6 0-6.6 3-6.6 6.6 0 .5.1 1 .2 1.5-5.5-.3-10.4-3-13.6-6.9-.6 1-.9 2.1-.9 3.3 0 2.3 1.2 4.3 2.9 5.5-1.1 0-2.1-.3-3-.8v.1c0 3.2 2.3 5.9 5.3 6.5-.6.2-1.1.2-1.7.2-.4 0-.8 0-1.2-.1.8 2.6 3.3 4.5 6.2 4.6-2.3 1.8-5.1 2.8-8.2 2.8-.5 0-1.1 0-1.6-.1C2.9 28 6.3 29 10 29c12.1 0 18.7-10 18.7-18.7v-.9c1.4-.9 2.5-2 3.4-3.4z' fill=currentColor /></symbol><symbol id=facebookIcon viewBox='0 0 32 32'><path d='M30.2 0H1.8C.8 0 0 .8 0 1.8v28.5c0 1 .8 1.8 1.8 1.8h15.3V19.6h-4.2v-4.8h4.2v-3.6c0-4.1 2.5-6.4 6.2-6.4 1.8 0 3.3.2 3.7.2v4.3h-2.6c-2 0-2.4 1-2.4 2.4v3.1h4.8l-.6 4.8H22V32h8.2c1 0 1.8-.8 1.8-1.8V1.8c0-1-.8-1.8-1.8-1.8z' fill=currentColor /></symbol><symbol id=rssIcon viewBox='0 0 32 32'><path d='M10.7 25.6c0 2.4-2 4.4-4.4 4.4S2 28 2 25.6s2-4.4 4.4-4.4 4.3 2 4.3 4.4zM6.1 2c-.6 0-1.3 0-2 .1-1.3.1-2.2 1.2-2.1 2.4.1 1.2 1.2 2.2 2.4 2.1.6 0 1.1-.1 1.6-.1 10.7 0 19.4 8.7 19.4 19.4 0 .5 0 1-.1 1.6-.1 1.2.8 2.3 2.1 2.4h.2c1.2 0 2.1-.9 2.2-2.1.1-.7.1-1.4.1-2C30 12.7 19.3 2 6.1 2zm-.7 9.6c-.4 0-.8 0-1.3.1-1.3.1-2.2 1.2-2.1 2.5.1 1.3 1.2 2.2 2.5 2.1h.9c5.7 0 10.4 4.7 10.4 10.4v.9c-.1 1.3.8 2.4 2.1 2.5h.2c1.2 0 2.2-.9 2.3-2.1 0-.4.1-.8.1-1.2-.1-8.4-6.9-15.2-15.1-15.2z' fill=currentColor /></symbol><symbol id=linkedInIcon viewBox='0 0 32 32'><path d='M29.6,0H2.4C1.1,0,0,1,0,2.3v27.4C0,31,1.1,32,2.4,32h27.3c1.3,0,2.4-1,2.4-2.3V2.3C32,1,30.9,0,29.6,0z M9.5,27.3H4.7V12 h4.8V27.3z M7.1,9.9c-1.5,0-2.8-1.2-2.8-2.8c0-1.5,1.2-2.8,2.8-2.8c1.5,0,2.8,1.2,2.8,2.8C9.9,8.7,8.6,9.9,7.1,9.9z M27.3,27.3 h-4.7v-7.4c0-1.8,0-4-2.5-4c-2.5,0-2.8,1.9-2.8,3.9v7.6h-4.7V12H17v2.1h0.1c0.6-1.2,2.2-2.5,4.5-2.5c4.8,0,5.7,3.2,5.7,7.3V27.3z' fill=currentColor /></symbol></svg><div class='column main'><section class=social><a href=https://twitter.com/incrementmag class=twitter><svg viewBox='0 0 32 32'><use xlink:href=#twitterIcon x=0 y=0></use></svg> <span>@incrementmag</span> </a><a href=https://facebook.com/incrementmag class=facebook><svg viewBox='0 0 32 32'><use xlink:href=#facebookIcon x=0 y=0></use></svg> <span>incrementmag</span> </a><a href=/feed.xml class=rss><svg viewBox='0 0 32 32'><use xlink:href=#rssIcon x=0 y=0></use></svg> <span>RSS Feed</span></a></section><section><h4>About</h4><p><em>Increment</em> is a print and digital magazine about how teams build and operate software systems at scale. <a href=/about/ >Learn more</a></p></section><section><h4>Work with us</h4><p>Interested in joining the team at Stripe? <a href=https://stripe.com/jobs>View job openings</a></p></section></div><p class='column copyright'><span>&copy; 2022 <em>Increment</em></span> <a href=https://stripe.com>Published by Stripe</a> <a href=https://stripe.com/privacy/media-policy>Privacy policy</a></p></div></footer></body></html>