Files
nexus/sreweekly/articles/261/05-increment-reliability-reliability-at-scale.html
2026-09-12 17:23:01 +08:00

19 lines
36 KiB
HTML
Raw Blame History

This file contains invisible Unicode characters
This file contains invisible Unicode characters that are indistinguishable to humans but may be processed differently by a computer. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
<!doctype html><html><head><meta charset=utf-8><title>Software Reliability at Scale - Increment</title><meta name=description content='Engineering leaders at Deliveroo, DigitalOcean, Fastly, and Headspace offer insights from their organizations and practical advice for engineering teams embarking on reliability journeys.'><link rel=canonical href=http://localhost:3000/reliability/reliability-at-scale/ ><link rel=apple-touch-icon-precomposed href=/img/icon-571805a1.png><meta property=og:title content='Reliability at scale – Increment: Reliability'><meta property=og:url content=http://localhost:3000/reliability/reliability-at-scale/ ><meta property=og:description content='Leaders at Deliveroo, DigitalOcean, Fastly, and Headspace share how their organizations think about reliability and resiliency and their advice to engineering orgs embarking on reliability journeys.'><meta property=og:image content='https://images.ctfassets.net/3njn2qm7rrbs/6Q3D0w8nis0u66FQeCTjbg/843b26ce0013258589c76a15a49461d4/cover-issue16.png?w=1000'><meta name=twitter:card content=summary_large_image><meta name=twitter:image content='https://images.ctfassets.net/3njn2qm7rrbs/6Q3D0w8nis0u66FQeCTjbg/843b26ce0013258589c76a15a49461d4/cover-issue16.png?w=1000'><meta name=twitter:site content=@IncrementMag><meta name=twitter:title content='Reliability at scale – Increment: Reliability'><meta name=twitter:description content='Leaders at Deliveroo, DigitalOcean, Fastly, and Headspace share how their organizations think about reliability and resiliency and their advice to engineering orgs embarking on reliability journeys.'><link rel=alternate type=application/rss+xml title=Increment href=/feed.xml><meta name=viewport content='width=device-width,initial-scale=1'><link rel=preload href=/fonts/baton-turbo/400-30a55d66.woff2 as=font type=font/woff2 crossorigin=anonymous><link rel=preload href=/fonts/baton-turbo/500-1603c0e8.woff2 as=font type=font/woff2 crossorigin=anonymous><link rel=preload href=/fonts/tiempos-text/400-c4810745.woff2 as=font type=font/woff2 crossorigin=anonymous><link rel=preload href=/fonts/tiempos-head/700-383ede62.woff2 as=font type=font/woff2 crossorigin=anonymous><link rel=stylesheet type=text/css href=/css/bundle-289d885f.css><link rel=stylesheet type=text/css href=/css/issues/16-c2b7b183.css><script>// Don't fade art if it loads ~instantly
setTimeout(()=>{document.documentElement.classList.add('fadeArt')},250);</script><script defer src=/js/defer-737fbd90.js></script><script>const INCREMENT_META={issueNumber:16,issueSlug:'reliability',articleSlug:'reliability-at-scale'};</script></head><body class='Issue_reliability Article_reliability-at-scale'><nav class=PageNav><div class=u-Container><div class=column><h1 class=logo><a href=/ ><img src=/img/logo-ae2c55d5.svg alt=Increment></a></h1><a class=out-now href=https://store.increment.com/ style=color:#595959><div class='IssueTitle tiny'><div class='t-Caps meta tiny'><span>NEW</span></div><h3 class='t-IssueTitle title'>Buy the print edition</h3></div></a><ul class=nav><li><a href=/issues/ ><span>Issues</span></a></li><li><a href=/topics/ ><span>Topics</span></a></li><li><a href=https://store.increment.com/ ><span>Store</span></a></li><li><a href=/about/ ><span>About</span></a></li></ul></div></div></nav><div class=ArticlePage itemscope itemtype=http://schema.org/Article><script type=application/ld+json>{
"@context": "http://schema.org",
"@type": "Article",
"headline": "Reliability at scale",
"image": " https://images.ctfassets.net/3njn2qm7rrbs/6Q3D0w8nis0u66FQeCTjbg/843b26ce0013258589c76a15a49461d4/cover-issue16.png?w&#x3D;1000",
"datePublished": "Thu, 25 Feb 2021 19:00:00 GMT",
"dateModified": "Thu, 25 Feb 2021 19:00:00 GMT",
"publisher": {
"@type": "Organization",
"name": "Increment",
"logo": {
"@type": "ImageObject",
"url": "https://increment.com/img/logo.png"
}
},
"description": "Leaders at Deliveroo, DigitalOcean, Fastly, and Headspace share how their organizations think about reliability and resiliency and their advice to engineering orgs embarking on reliability journeys.",
"mainEntityOfPage": "http://localhost:3000/reliability/reliability-at-scale/"
}</script><header class='u-Container ArticleHeader'><div class=column><div class=u-Grid><div class=main><h4 class='t-Byline byline large'><a href=#authors class=j-SmoothScroll itemprop=author><span>Increment Staff</span></a></h4><h1 class='t-TitleSerif large title' itemprop=name>Reliability at scale</h1><div class='t-BodySans large intro' itemprop=description>Leaders at Deliveroo, DigitalOcean, Fastly, and Headspace share how their organizations think about reliability and resiliency and their advice to engineering orgs embarking on reliability&nbsp;journeys.</div></div><a class=issue href=/reliability/ ><span class='t-Caps tiny part-of'>Part of</span><div class='IssueTitle small' style=color:#863051><div class='t-Caps meta'><span>Issue 16</span> <span>February 2021</span></div><h2 class='t-IssueTitle title'>Reliability</h2></div></a></div></div></header><div class='u-Container ArticleContent'><article class='ContentBody column' itemprop=articleBody><div class='ArticleLayout four-columns grid centered'><div><h4>Victoria Puscas</h4><p class=small><strong>Engineering manager, pricing and logistics algorithms<br></strong><i><strong>Deliveroo</strong></i></p><p class=small>2,500 employees</p></div><div><h4>Laura Thomson</h4><p class=small><strong>VP of engineering<br></strong><i><strong>Fastly</strong></i></p><p class=small>750+ employees</p></div><div><h4>Al Sene</h4><p class=small><strong>VP of engineering<br></strong><i><strong>DigitalOcean</strong></i></p><p class=small>500+ employees</p></div><div><h4>Bhavini Soneji</h4><p class=small><strong>VP of engineering<br></strong><i><strong>Headspace</strong></i></p><p class=small>250+ employees</p></div></div><div class=ArticleLayout><hr><h2>How does your organization think about resiliency and reliability writ large?</h2><p>Our engineering principles reflect our mission to prioritize high-quality code and emphasize our commitment to learning from mistakes. We conduct production incident reviews in a safe, nonjudgmental environment to understand both the root cause of a problem and the necessary steps to prepare the system and organization to successfully cope with similar problems in the future.</p><p>That’s also why we instrument our systems for observability. You don’t know you have a problem if you don’t monitor the health and performance of your systems—and have actionable, well-documented steps to mitigate issues. No service or system goes to production unless it has some basic monitoring in place.</p><p>Finally, we build to last. When building products and solutions, we ensure services will perform and scale given the business doubles in size annually. We run experiments and make data-driven decisions, which also means building services and products that can tolerate frequent change and are easy to extend or build on.</p><p class=small>— Victoria Puscas, Deliveroo</p><div class='Sidebar large'><blockquote><p>A resilient team’s foundation requires uniquely human traits like empathy, vulnerability, and understanding.</p></blockquote></div><p>We fall back on the resilience of our systems and people when reliability fails, so both are critical to our success. While some elements of a resilient team mirror a resilient network—avoiding single points of failure, for example—we know a resilient team’s foundation requires uniquely human traits like empathy, vulnerability, and understanding.</p><p class=small>— Laura Thomson, Fastly</p><p>We encourage a culture of learning to continually get better at preventing incidents, reducing impact, and shortening recovery times. Incidents are an opportunity to grow and improve resiliency—we run blameless postmortems that focus on lessons learned and preventative measures.</p><p>Scaling reliability requires a focus on technology, people, and process. In terms of technology, we architect and design in a way that’s conducive to more resilience and fault tolerance when things go wrong. Time to recover after failure detection is crucial, so the incident management and recovery process must be easy to deploy. Finally, it’s important to ensure people are well-prepared and invested in making the overall system better.</p><p class=small>— Al Sene, DigitalOcean</p><p>Our roadmap has two categories: innovation (product features) and tune-up items, which drive reliability and scale to speed of innovation. Tune-up items comprise the backbone of our innovation and include investing in automation of quality gates in order to release frequently, fast incident detection and response, and refactoring to ensure components are secure, scalable, and available, with low latency and continuous delivery.</p><p>This amounts to a self-perpetuating flywheel. Investing in tune-up items—and therefore reliability and speed of innovation—leads to higher development velocity, which improves product innovation, which leads to scaling the team, which then leads to more investment in the speed of innovation.</p><p class=small>— Bhavini Soneji, Headspace</p></div><div class='ArticleLayout flipped'><h2>Does your organization have dedicated reliability engineers?</h2><p>We don’t have DevOps or support engineers, per se. We have engineers who build solutions for our platforms and infrastructure and are responsible for building systems that enable our product teams to design, build, and ship changes quickly. These solutions include CI/CD, our event bus internal libraries, federated authentication services, and more.</p><p class=small>— Victoria Puscas, Deliveroo</p><p>We have reliability engineers <i>and</i> resilience engineers. Reliability engineers own our platform’s integrity, which represents customers’ trust in our services in every respect, including stability, security, and data integrity. Resilience engineers look for brittle points across all our systems and rebuild accordingly.</p><p class=small>— Laura Thomson, Fastly</p><p>We currently have a small, specialized group of dedicated reliability engineers, with the intention to grow into a larger SRE team. They work with product teams to enhance the customer experience and ensure our services meet their target availability.</p><p class=small>— Al Sene, DigitalOcean</p><p>We don’t have dedicated reliability engineers. Our approach is to give engineers and development teams end-to-end ownership of their scenario or component. We have an on-call rotation for each platform (iOS, Android, API, and web) with weekly rotations and a warm hand-off. The on-call engineer is responsible for initial triage and routing to the team that owns that scenario or component if there’s no standard playbook, as well as for both production and non-production environments and manning mobile client releases.</p><p class=small>— Al Sene, DigitalOcean</p></div><div class=ArticleLayout><h2>What measures or metrics do you use to capture investment in reliability?</h2><p>At a team level, we look at a prioritized list of repair items and reliability-related tickets every sprint. We also track our teams’ service uptime, SEVs by level, time in SEV, and time to restore. At a higher level, when we plan for the quarter or longer time horizons, we create engineering goals around making our platform more stable and scalable. These become part of our roadmap.</p><p>There are also times when we need to take a stance on some of our systems’ older parts. For example, we might have a company-level initiative to improve the tools we give restaurants to manage their commercials or menus. For such initiatives, we typically prioritize quality over profitability as metrics. These changes are intended to significantly improve the experiences of our restaurants, riders, or internal partners by giving them useful, reliable, mature tools and systems to grow and work with.</p><p class=small>— Victoria Puscas, Deliveroo</p><p>We measure common metrics like time to detect and time to recover from issues. We’re also always looking for ways to reframe base reliability metrics around customer impact. It’s important not to default to only measuring easy things, but to really dig in and ask teams, “What are your goals? What do you need to do to get there?” Metric choice is critical to driving the right system optimizations.</p><p class=small>— Laura Thomson, Fastly</p><p>We strive to monitor and measure everything we can to encapsulate all aspects of the process, find errors, determine the health of our systems, and identify opportunities for future growth. We enact our commitment to reliability by making it everyone’s responsibility.</p><p>Typically, we measure overall system availability, subsystems health, and customer experience against internal objectives. We also take into account industry-standard metrics like time to detect, time to recover, change fail rates, and outage durations. As we continue to make investments, we monitor how these metrics trend over time.</p><p class=small>— Al Sene, DigitalOcean</p><p>We evaluate reliability investment by asking:</p><ul><li><p>What’s the customer impact? (e.g., bugs, latency, time to value)</p></li><li><p>What’s the business impact? (e.g., brand trust, revenue impact due to downtime)</p></li><li><p>What’s the impact on the operational efficiency of the business? (e.g., developer and staff productivity)</p></li></ul><p>We use different sets of metrics. The first is a high-level reliability view, which includes count of incidents and customer impact, revenue implications, crash rate, and app store review rating. The second is around production incidents and follow-through, which includes mean time to resolution, fix rate for production bug fixes and incident action items, percentage of incidents reported through customer support versus detected through alerting, and more. The third is proactive software development life cycle pipeline strengthening, including releases canceled or rolled back, quality gates, count of end-to-end test automation and coverage, pre-production environment uptime, and more.</p><p class=small>— Bhavini Soneji, Headspace</p></div><div class='ArticleLayout flipped'><h2>When it’s been a while since your last incident, how do you keep your teams sharp and ensure continued investment in reliability?</h2><p>We celebrate our “longest since last SEV” moments, and if there are no incidents to discuss during our weekly live service health review sessions, we share best practices and celebrate good work.</p><p>We also proactively prepare for Q4, our busiest time of year. In September, our growth and daily order volume typically exceed expectations, which might come with unexpected production problems. During these periods, we prepare our systems to cope with roughly 10 percent more load every week. This involves looking at the system’s weakest points and proactively mitigating any risks.</p><p class=small>— Victoria Puscas, Deliveroo</p><p>Along with regular onboarding and training, we run tabletop exercises, or “pre-mortems,” which help employees hone their skills and get ahead of potential problems. These exercises take an operational team through a plausible scenario—usually complex, worst-case–style scenarios—and the team works through how they’d respond, what they’d investigate, and potential fixes. Then, if the worst really happens, people might have already thought through what to do and will be calmer and better prepared to respond. These exercises are also fun and great for team building.</p><p class=small>— Laura Thomson, Fastly</p><div class='Sidebar large'><blockquote><p>Reliability is a culture, and it has to be embraced by engineers, product managers, and designers.</p></blockquote></div><p>We keep our team sharp by improving existing processes and conducting readiness reviews to anticipate what could go wrong with new code and how to avoid it in the first place. We want to ensure the team is proactive, not just reactive, and we know there’s never a shortage of areas for improvement. We focus on better detection, prevention, and testability of our recovery procedures. It’s crucial to have a surefire testing method that easily understands incidents as they occur so they can be mitigated in the future.</p><p class=small>— Al Sene, DigitalOcean</p><p>Reliability is a culture, and it has to be embraced by engineers, product managers, and designers. We want to instill a culture of data-driven decision-making, and we want teams to proactively inspect the health of their releases and ensure they can analyze and detect issues.</p><p>We also prioritize transparency around incidents and learnings by adopting biweekly live site meetings with leads, fixing calendar slots for postmortems with learnings broadcast to the org, and quarterly chaos simulations in pre-production, which we aim to automate.</p><p class=small>— Bhavini Soneji, Headspace</p></div><div class=ArticleLayout><h2>How do you think about and/or fund projects to address low-probability but high-risk events?</h2><p>Some events can be mitigated with a clear procedure and line of escalation. We have a few documents that describe what we might need to do and who we might need to involve in case of a data breach or security vulnerability. (Fortunately, I can’t remember anything like that happening in the four years I’ve been here.)</p><p>There’s also a risk that your third-party provider goes down or a data center experiences an outage. Then, the question becomes about investing the time to mitigate these issues now versus later. We’ve started planning to build an infrastructure that’s resilient to local failures in the next few years by isolating issues to a specific region, rather than letting them affect customers elsewhere.</p><p class=small>— Victoria Puscas, Deliveroo</p><p>These events are best prevented at the architectural level, whether in the design phase or via the systems thinking and hacking projects our resilience engineering team takes on. We also work through simulations to explore prevention and mitigation strategies, including the tabletop exercises mentioned prior, as well as network simulations using in-house tools.</p><p class=small>— Laura Thomson, Fastly</p><p>Just because an event is low-probability doesn’t mean you shouldn’t prepare for it. You should invest in contingency plans to stay ahead of any incident that might pop up. This ensures you’re taking care of customers regardless of what’s happening behind the scenes.</p><div class='Sidebar large'><blockquote><p>Strengthening cross-functional teams, processes, and foundations isn’t necessarily exciting, but it’s key to innovation.</p></blockquote></div><p>We continue to invest heavily in building reliable, secure and highly available services across our portfolio of products. This is table stakes for the cloud industry.</p><p class=small>— Al Sene, DigitalOcean</p><p>For cloud region and data breaches, the first order is architecting and designing the system correctly, ensuring we have data encryption and data backups, and having stateless services driven through configuration. We ensure our roadmap and strategic planning prioritize strengthening our quality gates to mitigate incidents proactively with fast detection and response, chaos simulation, and cloud region and data security. Strengthening cross-functional teams, processes, and foundations isn’t necessarily exciting, but it’s key to innovation.</p><p class=small>— Bhavini Soneji, Headspace</p></div><div class='ArticleLayout flipped'><h2>What would you share with rapidly growing tech companies to help them on their own reliability journeys?</h2><p>Identify the most critical areas of your systems and proactively look at what needs to be done to support growth by at least one order of magnitude. I also recommend avoiding making huge changes in one go—anything bigger than a few weeks of work is probably too big, unless you have a specific reliability or scalability problem and know exactly what you’re doing. Why? Because this particular work, improvement, tech stack, etc. might not be the solution, and it may be difficult to convince the company to let a team take a six-month journey without a guarantee of success or improvement.</p><p>Things will still go sideways. That’s normal. Try to iterate and experiment quickly. Find a problem, prioritize it accordingly, and test it. There is no other secret sauce!</p><p class=small>— Victoria Puscas, Deliveroo</p><p>As you scale, you have to reinvent what you’re doing. You can’t just scale it up linearly—you have to get creative. Think divergently, and don’t be afraid to try something completely different. At Fastly, we’re scaling our performance test platform so we can squeeze every drop of performance from our systems under production-like traffic. Figuring out how to simulate realistic load has been an interesting challenge.</p><p class=small>— Laura Thomson, Fastly</p><p>My recommendation to other high-growth tech companies is to build a culture of learning within your organizations. Be transparent with your teams about incidents and the lessons learned, and continue to iterate. Honesty and humility about your shortcomings is the first step to ensuring future success.</p><p class=small>— Al Sene, DigitalOcean</p><p>You have to have the same mindset for reliability as product development. Many factors will shape your particular investment, but the key is making it part of the company’s DNA to find alignment between product and engineering teams during company strategy planning.</p><p>As feature teams broaden to include different product lines, staffing horizontal teams becomes critical to laying the technology foundation, driving reusability, and maintaining consistency. Platform teams lay out the common building blocks that application teams can build on or reuse, while infrastructure teams lay out the framework that application teams will integrate to drive continuous releases with quality gates and enable fast detection and response.</p><p>The bottom line: Have transparency and clear communication around decisions and trade-offs, while having the flexibility to align with business priorities and meet hard deadlines.</p><p class=small>— Bhavini Soneji, Headspace</p></div></article></div><div class='u-Container ArticleFooter'><div class=column><div class='u-Grid ContentBody small content'><div class=authors id=authors></div><div class=topics><div class=text><h4>Topics</h4><p><a href=/topics/scaling/ >Scaling &amp; Growth</a></p><p><a href=/topics/culture/ >Workplace &amp; Culture</a></p></div></div></div></div></div></div><div class=SubscribeBox><a id=newsletter class=anchor href=#newsletter></a><div class='ContentBody inverted store'><div class=u-Container><div class='u-Grid column'><div class=box style=background:#4c70b1><div class=text><h2>Buy the print edition</h2><p class=j-TextBalance>Visit the Increment Store to purchase print issues.</p><p><a class='t-Caps u-Arrow' href=https://store.increment.com/ >Store</a></p></div><a href=https://store.increment.com/ class=magazine><figure class=j-MaskedImage data-fill=/art/19/19-cutout-1000-7ffb5dba.png data-mask=/art/19/fill-cms-1000-83522f38.png></figure></a></div></div></div></div><div class='ContentBody email'><div class=u-Container><div class='u-Grid column'><div class='j-EmailForm box' style='box-shadow:0 -5px 0 #4c70b1'></div></div></div></div></div><div class=ContinueReading><div class=u-Container><div class=column><h2 class='t-Caps xlarge'>Continue Reading</h2><ul class='u-Grid articles'><li class='ArticleBlock footer' style=color:#8e65bf><a class=j-Preload href=/teams/engineering-teams-at-scale/ ><div class='IssueTitle tiny' style=color:#8e65bf><div class='t-Caps meta'><span>11</span></div><h3 class='t-IssueTitle title'>Teams</h3></div><div class=text><div class=head><h4 class='t-Byline byline'><span>Increment Staff</span></h4><h3 class='t-TitleSans title'>Engineering teams at scale</h3></div><div class='t-BodySerif small intro'>Leaders at 1Password, Atlassian, and Intuit share how they measure success, determine focus, and foster mentorship, sustainability, and collaboration in their organizations.</div></div></a></li><li class='ArticleBlock footer' style=color:#29386a><a class=j-Preload href=/remote/remote-work-at-scale-google-hashicorp-invision-range/ ><div class='IssueTitle tiny' style=color:#29386a><div class='t-Caps meta'><span>15</span></div><h3 class='t-IssueTitle title'>Remote</h3></div><div class=text><div class=head><h4 class='t-Byline byline'><span>Increment Staff</span></h4><h3 class='t-TitleSans title'>Remote at scale</h3></div><div class='t-BodySerif small intro'>Leaders at Google Cloud, HashiCorp, InVision, and Range discuss their remote ethos, supporting remote employees, and best practices for remote&nbsp;work.</div></div></a></li><li class='ArticleBlock footer' style=color:#4a5ad3><a class=j-Preload href=/open-source/open-source-at-scale/ ><div class='IssueTitle tiny' style=color:#4a5ad3><div class='t-Caps meta'><span>9</span></div><h3 class='t-IssueTitle title'>Open Source</h3></div><div class=text><div class=head><h4 class='t-Byline byline'><span>Increment Staff</span></h4><h3 class='t-TitleSans title'>Open source at scale</h3></div><div class='t-BodySerif small intro'>Technical leaders at Microsoft, Kickstarter, DigitalOcean, and Red Hat answer questions about when and why to opt for OSS, how open source has influenced their organizations, and its future role in corporations.</div></div></a></li><li class='ArticleBlock footer' style=color:#e89e00><a class=j-Preload href=/testing/testing-at-scale/ ><div class='IssueTitle tiny' style=color:#e89e00><div class='t-Caps meta'><span>10</span></div><h3 class='t-IssueTitle title'>Testing</h3></div><div class=text><div class=head><h4 class='t-Byline byline'><span>Increment Staff</span></h4><h3 class='t-TitleSans title'>Testing at scale</h3></div><div class='t-BodySerif small intro'>Rob Zuber (CircleCI), Greg Bell and Claudiu Coman (Hootsuite), and Scott Triglia (Yelp) talk test suite times, manual versus automated testing, and how to build testing infrastructure.</div></div></a></li><li class='ArticleBlock footer' style=color:#40af9e><a class=j-Preload href=/software-architecture/architecture-at-scale/ ><div class='IssueTitle tiny' style=color:#40af9e><div class='t-Caps meta'><span>12</span></div><h3 class='t-IssueTitle title'>Software Architecture</h3></div><div class=text><div class=head><h4 class='t-Byline byline'><span>Increment Staff</span></h4><h3 class='t-TitleSans title'>Software architecture at scale</h3></div><div class='t-BodySerif small intro'>Leaders at Foursquare, Hulu, and Twitter discuss early architecture decisions, downstream effects, and architectural philosophies.</div></div></a></li><li class='ArticleBlock footer' style=color:#5ebe92><a class=j-Preload href=/frontend/frontend-at-scale/ ><div class='IssueTitle tiny' style=color:#5ebe92><div class='t-Caps meta'><span>13</span></div><h3 class='t-IssueTitle title'>Frontend</h3></div><div class=text><div class=head><h4 class='t-Byline byline'><span>Increment Staff</span></h4><h3 class='t-TitleSans title'>Frontend at scale</h3></div><div class='t-BodySerif small intro'>Leaders at Atlassian, Canva, Tinder, and Vimeo discuss frameworks, tooling, and rapidly evolving technologies.</div></div></a></li><li class='ArticleBlock footer' style=color:#00afbe><a class=j-Preload href=/apis/apis-at-scale-adobe-airbnb-kong-pubnub/ ><div class='IssueTitle tiny' style=color:#00afbe><div class='t-Caps meta'><span>14</span></div><h3 class='t-IssueTitle title'>APIs</h3></div><div class=text><div class=head><h4 class='t-Byline byline'><span>Increment Staff</span></h4><h3 class='t-TitleSans title'>APIs at scale</h3></div><div class='t-BodySerif small intro'>Leaders at Adobe, Airbnb, Kong, and PubNub talk API design, documentation, and development.</div></div></a></li><li class='ArticleBlock footer' style=color:#443d79><a class=j-Preload href=/containers/containerization-at-scale/ ><div class='IssueTitle tiny' style=color:#443d79><div class='t-Caps meta'><span>17</span></div><h3 class='t-IssueTitle title'>Containers</h3></div><div class=text><div class=head><h4 class='t-Byline byline'><span>Increment Staff</span></h4><h3 class='t-TitleSans title'>Containers at scale</h3></div><div class='t-BodySerif small intro'>Engineering leaders at Datadog, Braze, and BetterUp discuss container tools, testing, and monitoring, and how they’ve approached container migrations.</div></div></a></li><li class='ArticleBlock footer' style=color:#439eab><a class=j-Preload href=/mobile/mobile-development-at-scale/ ><div class='IssueTitle tiny' style=color:#439eab><div class='t-Caps meta'><span>18</span></div><h3 class='t-IssueTitle title'>Mobile</h3></div><div class=text><div class=head><h4 class='t-Byline byline'><span>Increment Staff</span></h4><h3 class='t-TitleSans title'>Mobile development at scale</h3></div><div class='t-BodySerif small intro'>Engineering leaders at adidas Runtastic, Eventbrite, and Citymapper discuss app performance, how mobile fits into their org structures, and native versus cross-platform development.</div></div></a></li></ul><h3 class='t-Caps large'>Explore Topics</h3><ul class='t-BodySerif large topics'><li><a href=/topics/learn/ >Learn Something New</a></li><li><a href=/topics/scaling/ >Scaling &amp; Growth</a></li><li><a href=/topics/ask-an-expert/ >Ask an Expert</a></li><li><a href=/topics/interviews/ >Interviews &amp; Surveys</a></li><li><a href=/topics/guides/ >Guides &amp; Best Practices</a></li><li><a href=/topics/opinion/ >Essays &amp; Opinion</a></li><li><a href=/topics/culture/ >Workplace &amp; Culture</a></li></ul><h3 class='t-Caps large'>All Issues</h3><ul class=issues><li><a href=/planning/ ><div class='IssueTitle large' style=color:#4c70b1><div class='t-Caps meta'><span>Issue 19</span> <span>November 2021</span></div><h1 class='t-IssueTitle title'>Planning</h1></div></a></li><li><a href=/mobile/ ><div class='IssueTitle large' style=color:#439eab><div class='t-Caps meta'><span>Issue 18</span> <span>August 2021</span></div><h1 class='t-IssueTitle title'>Mobile</h1></div></a></li><li><a href=/containers/ ><div class='IssueTitle large' style=color:#443d79><div class='t-Caps meta'><span>Issue 17</span> <span>May 2021</span></div><h1 class='t-IssueTitle title'>Containers</h1></div></a></li><li><a href=/reliability/ ><div class='IssueTitle large' style=color:#863051><div class='t-Caps meta'><span>Issue 16</span> <span>February 2021</span></div><h1 class='t-IssueTitle title'>Reliability</h1></div></a></li><li><a href=/remote/ ><div class='IssueTitle large' style=color:#29386a><div class='t-Caps meta'><span>Issue 15</span> <span>November 2020</span></div><h1 class='t-IssueTitle title'>Remote</h1></div></a></li><li><a href=/apis/ ><div class='IssueTitle large' style=color:#00afbe><div class='t-Caps meta'><span>Issue 14</span> <span>August 2020</span></div><h1 class='t-IssueTitle title'>APIs</h1></div></a></li><li><a href=/frontend/ ><div class='IssueTitle large' style=color:#5ebe92><div class='t-Caps meta'><span>Issue 13</span> <span>May 2020</span></div><h1 class='t-IssueTitle title'>Frontend</h1></div></a></li><li><a href=/software-architecture/ ><div class='IssueTitle large' style=color:#40af9e><div class='t-Caps meta'><span>Issue 12</span> <span>February 2020</span></div><h1 class='t-IssueTitle title'>Software Architecture</h1></div></a></li><li><a href=/teams/ ><div class='IssueTitle large' style=color:#8e65bf><div class='t-Caps meta'><span>Issue 11</span> <span>November 2019</span></div><h1 class='t-IssueTitle title'>Teams</h1></div></a></li><li><a href=/testing/ ><div class='IssueTitle large' style=color:#e89e00><div class='t-Caps meta'><span>Issue 10</span> <span>August 2019</span></div><h1 class='t-IssueTitle title'>Testing</h1></div></a></li><li><a href=/open-source/ ><div class='IssueTitle large' style=color:#4a5ad3><div class='t-Caps meta'><span>Issue 9</span> <span>May 2019</span></div><h1 class='t-IssueTitle title'>Open Source</h1></div></a></li><li><a href=/internationalization/ ><div class='IssueTitle large' style=color:#c096ca><div class='t-Caps meta'><span>Issue 8</span> <span>February 2019</span></div><h1 class='t-IssueTitle title'>Internationalization</h1></div></a></li><li><a href=/security/ ><div class='IssueTitle large' style=color:#4dbac5><div class='t-Caps meta'><span>Issue 7</span> <span>October 2018</span></div><h1 class='t-IssueTitle title'>Security</h1></div></a></li><li><a href=/documentation/ ><div class='IssueTitle large' style=color:#f5684d><div class='t-Caps meta'><span>Issue 6</span> <span>August 2018</span></div><h1 class='t-IssueTitle title'>Documentation</h1></div></a></li><li><a href=/programming-languages/ ><div class='IssueTitle large' style=color:#d69336><div class='t-Caps meta'><span>Issue 5</span> <span>April 2018</span></div><h1 class='t-IssueTitle title'>Programming Languages</h1></div></a></li><li><a href=/energy-environment/ ><div class='IssueTitle large' style=color:#d6658e><div class='t-Caps meta'><span>Issue 4</span> <span>February 2018</span></div><h1 class='t-IssueTitle title'>Energy & Environment</h1></div></a></li><li><a href=/development/ ><div class='IssueTitle large' style=color:#53a88e><div class='t-Caps meta'><span>Issue 3</span> <span>October 2017</span></div><h1 class='t-IssueTitle title'>Development</h1></div></a></li><li><a href=/cloud/ ><div class='IssueTitle large' style=color:#707aed><div class='t-Caps meta'><span>Issue 2</span> <span>July 2017</span></div><h1 class='t-IssueTitle title'>Cloud</h1></div></a></li><li><a href=/on-call/ ><div class='IssueTitle large' style=color:#ef766e><div class='t-Caps meta'><span>Issue 1</span> <span>April 2017</span></div><h1 class='t-IssueTitle title'>On-Call</h1></div></a></li></ul></div></div></div><footer class=PageFooter><div class='u-Container ContentBody small'><svg style=display:none><symbol id=twitterIcon viewBox='0 0 32 32'><path d='M32.1 6c-1.2.5-2.5.9-3.8 1 1.4-.8 2.4-2.1 2.9-3.6-1.3.8-2.7 1.3-4.2 1.6a6.8 6.8 0 0 0-4.8-2c-3.6 0-6.6 3-6.6 6.6 0 .5.1 1 .2 1.5-5.5-.3-10.4-3-13.6-6.9-.6 1-.9 2.1-.9 3.3 0 2.3 1.2 4.3 2.9 5.5-1.1 0-2.1-.3-3-.8v.1c0 3.2 2.3 5.9 5.3 6.5-.6.2-1.1.2-1.7.2-.4 0-.8 0-1.2-.1.8 2.6 3.3 4.5 6.2 4.6-2.3 1.8-5.1 2.8-8.2 2.8-.5 0-1.1 0-1.6-.1C2.9 28 6.3 29 10 29c12.1 0 18.7-10 18.7-18.7v-.9c1.4-.9 2.5-2 3.4-3.4z' fill=currentColor /></symbol><symbol id=facebookIcon viewBox='0 0 32 32'><path d='M30.2 0H1.8C.8 0 0 .8 0 1.8v28.5c0 1 .8 1.8 1.8 1.8h15.3V19.6h-4.2v-4.8h4.2v-3.6c0-4.1 2.5-6.4 6.2-6.4 1.8 0 3.3.2 3.7.2v4.3h-2.6c-2 0-2.4 1-2.4 2.4v3.1h4.8l-.6 4.8H22V32h8.2c1 0 1.8-.8 1.8-1.8V1.8c0-1-.8-1.8-1.8-1.8z' fill=currentColor /></symbol><symbol id=rssIcon viewBox='0 0 32 32'><path d='M10.7 25.6c0 2.4-2 4.4-4.4 4.4S2 28 2 25.6s2-4.4 4.4-4.4 4.3 2 4.3 4.4zM6.1 2c-.6 0-1.3 0-2 .1-1.3.1-2.2 1.2-2.1 2.4.1 1.2 1.2 2.2 2.4 2.1.6 0 1.1-.1 1.6-.1 10.7 0 19.4 8.7 19.4 19.4 0 .5 0 1-.1 1.6-.1 1.2.8 2.3 2.1 2.4h.2c1.2 0 2.1-.9 2.2-2.1.1-.7.1-1.4.1-2C30 12.7 19.3 2 6.1 2zm-.7 9.6c-.4 0-.8 0-1.3.1-1.3.1-2.2 1.2-2.1 2.5.1 1.3 1.2 2.2 2.5 2.1h.9c5.7 0 10.4 4.7 10.4 10.4v.9c-.1 1.3.8 2.4 2.1 2.5h.2c1.2 0 2.2-.9 2.3-2.1 0-.4.1-.8.1-1.2-.1-8.4-6.9-15.2-15.1-15.2z' fill=currentColor /></symbol><symbol id=linkedInIcon viewBox='0 0 32 32'><path d='M29.6,0H2.4C1.1,0,0,1,0,2.3v27.4C0,31,1.1,32,2.4,32h27.3c1.3,0,2.4-1,2.4-2.3V2.3C32,1,30.9,0,29.6,0z M9.5,27.3H4.7V12 h4.8V27.3z M7.1,9.9c-1.5,0-2.8-1.2-2.8-2.8c0-1.5,1.2-2.8,2.8-2.8c1.5,0,2.8,1.2,2.8,2.8C9.9,8.7,8.6,9.9,7.1,9.9z M27.3,27.3 h-4.7v-7.4c0-1.8,0-4-2.5-4c-2.5,0-2.8,1.9-2.8,3.9v7.6h-4.7V12H17v2.1h0.1c0.6-1.2,2.2-2.5,4.5-2.5c4.8,0,5.7,3.2,5.7,7.3V27.3z' fill=currentColor /></symbol></svg><div class='column main'><section class=social><a href=https://twitter.com/incrementmag class=twitter><svg viewBox='0 0 32 32'><use xlink:href=#twitterIcon x=0 y=0></use></svg> <span>@incrementmag</span> </a><a href=https://facebook.com/incrementmag class=facebook><svg viewBox='0 0 32 32'><use xlink:href=#facebookIcon x=0 y=0></use></svg> <span>incrementmag</span> </a><a href=/feed.xml class=rss><svg viewBox='0 0 32 32'><use xlink:href=#rssIcon x=0 y=0></use></svg> <span>RSS Feed</span></a></section><section><h4>About</h4><p><em>Increment</em> is a print and digital magazine about how teams build and operate software systems at scale. <a href=/about/ >Learn more</a></p></section><section><h4>Work with us</h4><p>Interested in joining the team at Stripe? <a href=https://stripe.com/jobs>View job openings</a></p></section></div><p class='column copyright'><span>&copy; 2022 <em>Increment</em></span> <a href=https://stripe.com>Published by Stripe</a> <a href=https://stripe.com/privacy/media-policy>Privacy policy</a></p></div></footer></body></html>