363 lines
21 KiB
HTML
363 lines
21 KiB
HTML
<!DOCTYPE html>
|
|
|
|
|
|
<html class="no-js" lang="en">
|
|
<head>
|
|
<meta charset="utf-8">
|
|
<title>Some notes on running new software in production</title>
|
|
<meta name="author" content="Julia Evans">
|
|
<meta name="HandheldFriendly" content="True">
|
|
<meta name="MobileOptimized" content="320">
|
|
<meta name="description" content="Some notes on running new software in production">
|
|
<meta name="viewport" content="width=device-width, initial-scale=1">
|
|
|
|
<meta property="og:title" content='Some notes on running new software in production'>
|
|
<meta property="og:type" content="website" />
|
|
<meta property="og:url" content="https://jvns.ca/blog/2018/11/11/understand-the-software-you-use-in-production/" />
|
|
<meta property="og:site_name" content="Julia Evans" />
|
|
|
|
<link rel="canonical" href="https://jvns.ca/blog/2018/11/11/understand-the-software-you-use-in-production/">
|
|
<link href="/favicon.ico" rel="icon">
|
|
|
|
<link href="/stylesheets/screen.css" rel="preload" type="text/css" as="style">
|
|
|
|
<link href="/stylesheets/screen.css" media="screen, projection" rel="stylesheet" type="text/css">
|
|
<link href="/stylesheets/print.css" media="print" rel="stylesheet" type="text/css">
|
|
|
|
|
|
<link href="/atom.xml" rel="alternate" title="Julia Evans" type="application/atom+xml">
|
|
|
|
|
|
<link rel="stylesheet" href="https://cdn.jsdelivr.net/npm/katex@0.16.4/dist/katex.min.css" integrity="sha384-vKruj+a13U8yHIkAyGgK1J3ArTLzrFGBbBc0tDp4ad/EyewESeXE/Iv67Aj8gKZ0" crossorigin="anonymous">
|
|
<script defer data-domain="jvns.ca" src="https://plausible.io/js/script.js"></script>
|
|
<script defer src="https://cdn.jsdelivr.net/npm/katex@0.16.4/dist/katex.min.js" integrity="sha384-PwRUT/YqbnEjkZO0zZxNqcxACrXe+j766U2amXcgMg5457rve2Y7I6ZJSm2A0mS4" crossorigin="anonymous"></script>
|
|
<script defer src="https://cdn.jsdelivr.net/npm/katex@0.16.4/dist/contrib/auto-render.min.js" integrity="sha384-+VBxd3r6XgURycqtZ117nYw44OOcIax56Z4dCRWbxyPt0Koah1uHoK0o4+/RRE05" crossorigin="anonymous" onload="renderMathInElement(document.body);"></script>
|
|
|
|
<script defer type="text/javascript">
|
|
window.heap=window.heap||[],heap.load=function(e,t){window.heap.appid=e,window.heap.config=t=t||{};var r=document.createElement("script");r.type="text/javascript",r.async=!0,r.src="https://cdn.heapanalytics.com/js/heap-"+e+".js";var a=document.getElementsByTagName("script")[0];a.parentNode.insertBefore(r,a);for(var n=function(e){return function(){heap.push([e].concat(Array.prototype.slice.call(arguments,0)))}},p=["addEventProperties","addUserProperties","clearEventProperties","identify","resetIdentity","removeEventProperty","setEventProperties","track","unsetEventProperty"],o=0;o<p.length;o++)heap[p[o]]=n(p[o])};
|
|
heap.load("2242143965");
|
|
</script>
|
|
</head>
|
|
<body>
|
|
<div id="skiptocontent">
|
|
<a href="#main">Skip to main content</a>
|
|
</div>
|
|
<div id="wrap">
|
|
<header role="banner">
|
|
<hgroup>
|
|
<h1><a href="/">Julia Evans</a></h1>
|
|
</hgroup>
|
|
<ul class="header-links">
|
|
<li><a href="/about">About</a></li>
|
|
<li><a href="/talks">Talks</a></li>
|
|
<li><a href="/projects/">Projects</a></li>
|
|
<li><a rel="me" href="https://social.jvns.ca/@b0rk">Mastodon</a></li>
|
|
<li><a href="https://bsky.app/profile/b0rk.jvns.ca">Bluesky</a></li>
|
|
<li><a href="https://github.com/jvns">Github</a></li>
|
|
</ul>
|
|
</header>
|
|
<nav role="navigation" class="header-nav"><ul class="main-navigation">
|
|
<li><a href="/categories/favorite/">Favorites</a></li>
|
|
<li><a href="/til/">TIL</a></li>
|
|
<li><a href="https://wizardzines.com">Zines</a></li>
|
|
<li class="subscription" data-subscription="rss"><a href="/atom.xml" rel="subscribe-rss" title="subscribe via RSS">RSS</a></li>
|
|
</ul>
|
|
</nav>
|
|
<div id="main">
|
|
<div id="content">
|
|
|
|
|
|
<div>
|
|
<article class="hentry" role="article">
|
|
<header>
|
|
<h1 class="entry-title">Some notes on running new software in production</h1>
|
|
|
|
<div class="post-tags">
|
|
|
|
</div>
|
|
<p class="meta sans">
|
|
<time class="date" datetime="2018-11-11T11:00:01" pubdate data-updated="true">
|
|
|
|
November 11, 2018
|
|
|
|
</time>
|
|
</p>
|
|
</header>
|
|
<main>
|
|
<p>I’m working on a talk for kubecon in December! One of the points I want to get across is the amount
|
|
of time/investment it takes to use new software in production without causing really serious
|
|
incidents, and what that’s looked like for us in our use of Kubernetes.</p>
|
|
<p>To start out, this post isn’t blanket advice. There are lots of times when it’s totally fine to just
|
|
use software and not worry about <strong>how</strong> it works exactly. So let’s start by talking about when it’s
|
|
important to invest.</p>
|
|
<h3 id="when-it-matters-99-99" class="post-heading">
|
|
<a href="#when-it-matters-99-99">
|
|
when it matters: 99.99%
|
|
</a>
|
|
</h3>
|
|
<p>If you’re running a service with a low SLO like 99% I don’t think it matters that much to understand
|
|
the software you run in production. You can be down for like 2 hours a month! If something goes
|
|
wrong, just fix it and it’s fine.</p>
|
|
<p>At 99.99%, it’s different. That’s 45 minutes / year of downtime, and if you find out about a serious
|
|
issue for the first time in production it could easily take you 20 minutes or to revert the change.
|
|
That’s half your uptime budget for the year!</p>
|
|
<h3 id="when-it-matters-software-that-you-re-using-heavily" class="post-heading">
|
|
<a href="#when-it-matters-software-that-you-re-using-heavily">
|
|
when it matters: software that you’re using heavily
|
|
</a>
|
|
</h3>
|
|
<p>Also, even if you’re running a service with a 99.99% SLO, it’s impossible to develop a super deep
|
|
understanding of every single piece of software you’re using. For example, a web service might use:</p>
|
|
<ul>
|
|
<li>100 library dependencies</li>
|
|
<li>the filesystem (so there’s linux filesystem code!)</li>
|
|
<li>the network (linux networking code!)</li>
|
|
<li>a database (like postgres)</li>
|
|
<li>a proxy (like nginx/haproxy)</li>
|
|
</ul>
|
|
<p>If you’re only reading like 2 files from disk, you don’t need to do a super deep dive into Linux
|
|
filesystems internals, you can just read the file from disk.</p>
|
|
<p>What I try to do in practice is identify the components which we rely on the (or have the most
|
|
unusual use cases for!), and invest time into understanding those. These are usually pretty easy to
|
|
identify because they’re the ones which will cause the most problems :)</p>
|
|
<h3 id="when-it-matters-new-software" class="post-heading">
|
|
<a href="#when-it-matters-new-software">
|
|
when it matters: new software
|
|
</a>
|
|
</h3>
|
|
<p>Understanding your software especially matters for newer/less mature software projects, because it’s
|
|
more likely to have bugs & or just not have matured enough to be used by most people without
|
|
having to worry. I’ve spent a bunch of time recently with Kubernetes/Envoy which are both relatively
|
|
new projects, and neither of those are remotely in the category of “oh, it’ll just work, don’t worry
|
|
about it”. I’ve spent many hours debugging weird surprising edge cases with both of them and
|
|
learning how to configure them in the right way.</p>
|
|
<h3 id="a-playbook-for-understanding-your-software" class="post-heading">
|
|
<a href="#a-playbook-for-understanding-your-software">
|
|
a playbook for understanding your software
|
|
</a>
|
|
</h3>
|
|
<p>The playbook for understanding the software you run in production is pretty simple. Here it is:</p>
|
|
<ol>
|
|
<li>Start using it in production in a non-critical capacity (by sending a small percentage of traffic
|
|
to it, on a less critical service, etc)</li>
|
|
<li>Let that bake for a few weeks.</li>
|
|
<li>Run into problems.</li>
|
|
<li>Fix the problems. Go to step 3.</li>
|
|
</ol>
|
|
<p>Repeat until you feel like you have a good handle on this software’s failure modes and are
|
|
comfortable running it in a more critical capacity. Let’s talk about that in a little more detail,
|
|
though:</p>
|
|
<h3 id="what-running-into-bugs-looks-like" class="post-heading">
|
|
<a href="#what-running-into-bugs-looks-like">
|
|
what running into bugs looks like
|
|
</a>
|
|
</h3>
|
|
<p>For example, I’ve been spending a lot of time with Envoy in the last year. Some of the issues we’ve
|
|
seen along the way are: (in no particular order)</p>
|
|
<ul>
|
|
<li>One of the default settings resulted in retry & timeout headers not being respected</li>
|
|
<li>Envoy (as a client) doesn’t support TLS session resumption, so servers with a large amount of Envoy clients get DDOSed by TLS handshakes</li>
|
|
<li>Envoy’s active healthchecking means that you services get healthchecked by every client. This is
|
|
mostly okay but (again) services with many clients can get overwhelmed by it.</li>
|
|
<li>Having every client independently healthcheck every server interacts somewhat poorly with services
|
|
which are under heavy load, and can exacerbate performance issues by removing up-but-slow clients
|
|
from the load balancer rotation.</li>
|
|
<li>Envoy doesn’t retry failed connections by default</li>
|
|
<li>it frequently segfaults when given incorrect configuration</li>
|
|
<li>various issues with it segfaulting because of resource leaks / memory safety issues</li>
|
|
<li>hosts running out of disk space between we didn’t rotate Envoy log files often enough</li>
|
|
</ul>
|
|
<p>A lot of these aren’t bugs – they’re just cases where what we expected the default configuration
|
|
to do one thing, and it did another thing. This happens all the time, and it can result in really
|
|
serious incidents. Figuring out how to configure a complicated piece of software appropriately takes
|
|
a lot of time, and you just have to account for that.</p>
|
|
<p>And Envoy is great software! The maintainers are incredibly responsive, they fix bugs quickly and
|
|
its performance is good. It’s overall been quite stable and it’s done well in production. But just
|
|
because something is great software doesn’t mean you won’t also run into 10 or 20 relatively serious
|
|
issues along the way that need to be addressed in one way or another. And it’s helpful to understand
|
|
those issues <strong>before</strong> putting the software in a really critical place.</p>
|
|
<h3 id="try-to-have-each-incident-only-once" class="post-heading">
|
|
<a href="#try-to-have-each-incident-only-once">
|
|
try to have each incident only once
|
|
</a>
|
|
</h3>
|
|
<p>My view is that running new software in production inevitably results in incidents. The trick:</p>
|
|
<ol>
|
|
<li>Make sure the incidents aren’t too serious (by making ‘production’ a less critical system first)</li>
|
|
<li>Whenever there’s an incident (even if it’s not that serious!!!), spend the time necessary to
|
|
understand exactly why it happened and how to make sure it doesn’t happen again</li>
|
|
</ol>
|
|
<p>My experience so far has been that it’s actually relatively possible to pull off “have every
|
|
incident only once”. When we investigate issues and implement remediations, usually that issue
|
|
<strong>never comes back</strong>. The remediation can either be:</p>
|
|
<ul>
|
|
<li>a configuration change</li>
|
|
<li>reporting a bug upstream and either fixing it ourselves or waiting for a fix</li>
|
|
<li>a workaround (“this software doesn’t work with 10,000 clients? ok, we just won’t use it with in
|
|
cases where there are that many clients for now!”, “oh, a memory leak? let’s just restart it every
|
|
hour”)</li>
|
|
</ul>
|
|
<p>Knowledge-sharing is really important here too – it’s always unfortunate when one person finds an
|
|
incident in production, fixes it, but doesn’t explain the issue to the rest of the team so somebody
|
|
else ends up causing the same incident again later because they didn’t hear about the original
|
|
incident.</p>
|
|
<h3 id="understand-what-is-ok-to-break-and-isn-t" class="post-heading">
|
|
<a href="#understand-what-is-ok-to-break-and-isn-t">
|
|
Understand what is ok to break and isn’t
|
|
</a>
|
|
</h3>
|
|
<p>Another huge part of understanding the software I run in production is understanding which parts
|
|
are OK to break (aka “if this breaks, it won’t result in a production incident”) and which aren’t.
|
|
This lets me <strong>focus</strong>: I can put big boxes around some components and decide “ok, if this breaks it
|
|
doesn’t matter, so I won’t pay super close attention to it”.</p>
|
|
<p>For example, with Kubernetes:</p>
|
|
<p>ok to break:</p>
|
|
<ul>
|
|
<li>any stateless control plane component can crash or be cycled out or go down for 5 minutes at any
|
|
time. If we had 95% uptime for the kubernetes control plane that would probably be fine, it just
|
|
needs to be working most of the time.</li>
|
|
<li>kubernetes networking (the system where you give every pod an IP addresses) can break as much as
|
|
it wants because we decided not to use it to start</li>
|
|
</ul>
|
|
<p>not ok:</p>
|
|
<ul>
|
|
<li>for us, if etcd goes down for 10 minutes, that’s ok. If it goes down for 2 hours, it’s not</li>
|
|
<li>containers not starting or crashing on startup (iam issues, docker not starting containers, bugs
|
|
in the scheduler, bugs in other controllers) is serious and needs to be looked at immediately</li>
|
|
<li>containers not having access to the resources they need (because of permissions issues, etc)</li>
|
|
<li>pods being terminated unexpectedly by Kubernetes (if you configure kubernetes wrong it can
|
|
terminate your pods!)</li>
|
|
</ul>
|
|
<p>with Envoy, the breakdown is pretty different:</p>
|
|
<p>ok to break:</p>
|
|
<ul>
|
|
<li>if the envoy control plane goes down for 5 minutes, that’s fine (it’ll keep working with stale
|
|
data)</li>
|
|
<li>segfaults on startup due to configuration errors are sort of okay because they manifest so early
|
|
and they’re unlikely to surprise us (if the segfault doesn’t happen the 1st time, it shouldn’t
|
|
happen the 200th time)</li>
|
|
</ul>
|
|
<p>not ok:</p>
|
|
<ul>
|
|
<li>Envoy crashes / segfaults are not good – if it crashes, network connections don’t happen</li>
|
|
<li>if the control server serves incorrect or incomplete data that’s extremely dangerous and can
|
|
result in serious production incidents. (so downtime is fine, but serving incorrect data is not!)</li>
|
|
</ul>
|
|
<p>Neither of these lists are complete at all, but they’re examples of what I mean by “understand your
|
|
sofware”.</p>
|
|
<h3 id="sharing-ok-to-break-not-ok-lists-is-useful" class="post-heading">
|
|
<a href="#sharing-ok-to-break-not-ok-lists-is-useful">
|
|
sharing ok to break / not ok lists is useful
|
|
</a>
|
|
</h3>
|
|
<p>I think these “ok to break” / “not ok” lists are really useful to share, because even if they’re not
|
|
100% the same for every user, the lessons are pretty hard won. I’d be curious to hear about your
|
|
breakdown of what kinds of failures are ok / not ok for software you’re using!</p>
|
|
<p>Figuring out all the failure modes of a new piece of software and how they apply to your situation
|
|
can take months. (this is why when you ask your database team “hey can we just use NEW DATABASE”
|
|
they look at you in such a pained way). So anything we can do to help other people learn faster is
|
|
amazing</p>
|
|
|
|
</main>
|
|
|
|
<footer>
|
|
|
|
<style type="text/css">
|
|
#mc_embed_signup{background:#fff; clear:left; font:14px Helvetica,Arial,sans-serif; display: inline;}
|
|
#mc_embed_signup {
|
|
display: inline;
|
|
}
|
|
#mc_embed_signup input.button {
|
|
background: #ff5e00;
|
|
display: inline;
|
|
color: white;
|
|
padding: 6px 12px;
|
|
}
|
|
</style>
|
|
<div class="sharing">
|
|
|
|
<style>
|
|
.form-inline {
|
|
display:flex; flex-flow: row wrap; justify-content: center;
|
|
}
|
|
.form-inline input, .form-inline span {
|
|
padding: 10px;
|
|
}
|
|
.form-inline input {
|
|
display:inline;
|
|
max-width:30%;
|
|
margin: 0 10px 0 0;
|
|
background-color: #fff;
|
|
border: 1px solid #ddd;
|
|
border-radius: 5px;
|
|
padding: 10px;
|
|
}
|
|
button {
|
|
background-color: #f50;
|
|
box-shadow: none;
|
|
border: 0;
|
|
border-radius: 5px;
|
|
color: white;
|
|
padding: 5px 10px;
|
|
}
|
|
@media (max-width: 800px) {
|
|
.form-inline input {
|
|
margin: 10px 0;
|
|
max-width:100% !important;
|
|
}
|
|
.form-inline {
|
|
flex-direction: column;
|
|
align-items: stretch;
|
|
}
|
|
}
|
|
</style>
|
|
|
|
<div align="center">
|
|
<form class="form-inline" action="https://app.convertkit.com/forms/1052396/subscriptions" method="post" data-uid="8884355abb" data-format="inline" data-version="5">
|
|
<span> Want a weekly digest of this blog?</span>
|
|
<input name="email_address" type="text" placeholder="Email address" />
|
|
<button type="submit" data-element="submit">Subscribe</button>
|
|
</form>
|
|
</div>
|
|
|
|
|
|
</div>
|
|
|
|
<p class="meta">
|
|
|
|
<a class="basic-alignment left" href="https://jvns.ca/blog/2018/11/01/tailwind--write-css-without-the-css/" title="Previous Post: Tailwind: style your site without writing any CSS!">Tailwind: style your site without writing any CSS!</a>
|
|
|
|
|
|
<a class="basic-alignment right" href="https://jvns.ca/blog/2018/11/18/c---destructors---really-useful/" title="Next Post: An example of how C++ destructors are useful in Envoy">An example of how C++ destructors are useful in Envoy</a>
|
|
|
|
</p>
|
|
</footer>
|
|
|
|
</article>
|
|
</div>
|
|
|
|
</div>
|
|
</div>
|
|
<nav role="navigation" class="footer-nav"> <a href="/">Archives</a>
|
|
</nav>
|
|
<footer role="contentinfo"><span class="credit">© Julia Evans. </span>
|
|
<span>If you like this, you may like <a href="https://web.archive.org/web/20181228051203/http://www.uliaea.ca/">Ulia Ea</a> or, more seriously, this list of <a href="https://jvns.ca/blogroll">blogs I love</a> or some <a href="https://jvns.ca/bookshelf">books I've read</a>. <br>
|
|
<p class="rc-scout__text"><i class="rc-scout__logo"></i>
|
|
You might also like the <a class="rc-scout__link" href="https://www.recurse.com/scout/click?t=546ea46360584b522270b8c3e5d830f8">Recurse Center</a>, my very favorite programming community <a href="/categories/hackerschool/">(my posts about it)</a></p>
|
|
</span>
|
|
<style class="rc-scout__style" type="text/css">.rc-scout{display:block;padding:0;border:0;margin:0;}.rc-scout__text{display:block;padding:0;border:0;margin:0;height:100%;font-size:100%;}.rc-scout__logo{display:inline-block;padding:0;border:0;margin:0;width:0.85em;height:0.85em;background:no-repeat center url('data:image/svg+xml;utf8,%3Csvg%20xmlns%3D%22http%3A%2F%2Fwww.w3.org%2F2000%2Fsvg%22%20viewBox%3D%220%200%2012%2015%22%3E%3Crect%20x%3D%220%22%20y%3D%220%22%20width%3D%2212%22%20height%3D%2210%22%20fill%3D%22%23000%22%3E%3C%2Frect%3E%3Crect%20x%3D%221%22%20y%3D%221%22%20width%3D%2210%22%20height%3D%228%22%20fill%3D%22%23fff%22%3E%3C%2Frect%3E%3Crect%20x%3D%222%22%20y%3D%222%22%20width%3D%228%22%20height%3D%226%22%20fill%3D%22%23000%22%3E%3C%2Frect%3E%3Crect%20x%3D%222%22%20y%3D%223%22%20width%3D%221%22%20height%3D%221%22%20fill%3D%22%2361ae24%22%3E%3C%2Frect%3E%3Crect%20x%3D%224%22%20y%3D%223%22%20width%3D%221%22%20height%3D%221%22%20fill%3D%22%2361ae24%22%3E%3C%2Frect%3E%3Crect%20x%3D%226%22%20y%3D%223%22%20width%3D%221%22%20height%3D%221%22%20fill%3D%22%2361ae24%22%3E%3C%2Frect%3E%3Crect%20x%3D%223%22%20y%3D%225%22%20width%3D%222%22%20height%3D%221%22%20fill%3D%22%2361ae24%22%3E%3C%2Frect%3E%3Crect%20x%3D%226%22%20y%3D%225%22%20width%3D%222%22%20height%3D%221%22%20fill%3D%22%2361ae24%22%3E%3C%2Frect%3E%3Crect%20x%3D%224%22%20y%3D%229%22%20width%3D%224%22%20height%3D%223%22%20fill%3D%22%23000%22%3E%3C%2Frect%3E%3Crect%20x%3D%221%22%20y%3D%2211%22%20width%3D%2210%22%20height%3D%224%22%20fill%3D%22%23000%22%3E%3C%2Frect%3E%3Crect%20x%3D%220%22%20y%3D%2212%22%20width%3D%2212%22%20height%3D%223%22%20fill%3D%22%23000%22%3E%3C%2Frect%3E%3Crect%20x%3D%222%22%20y%3D%2213%22%20width%3D%221%22%20height%3D%221%22%20fill%3D%22%23fff%22%3E%3C%2Frect%3E%3Crect%20x%3D%223%22%20y%3D%2212%22%20width%3D%221%22%20height%3D%221%22%20fill%3D%22%23fff%22%3E%3C%2Frect%3E%3Crect%20x%3D%224%22%20y%3D%2213%22%20width%3D%221%22%20height%3D%221%22%20fill%3D%22%23fff%22%3E%3C%2Frect%3E%3Crect%20x%3D%225%22%20y%3D%2212%22%20width%3D%221%22%20height%3D%221%22%20fill%3D%22%23fff%22%3E%3C%2Frect%3E%3Crect%20x%3D%226%22%20y%3D%2213%22%20width%3D%221%22%20height%3D%221%22%20fill%3D%22%23fff%22%3E%3C%2Frect%3E%3Crect%20x%3D%227%22%20y%3D%2212%22%20width%3D%221%22%20height%3D%221%22%20fill%3D%22%23fff%22%3E%3C%2Frect%3E%3Crect%20x%3D%228%22%20y%3D%2213%22%20width%3D%221%22%20height%3D%221%22%20fill%3D%22%23fff%22%3E%3C%2Frect%3E%3Crect%20x%3D%229%22%20y%3D%2212%22%20width%3D%221%22%20height%3D%221%22%20fill%3D%22%23fff%22%3E%3C%2Frect%3E%3C%2Fsvg%3E');}.rc-scout__link:link,.rc-scout__link:visited{color:#61ae24;text-decoration:underline;}.rc-scout__link:hover,.rc-scout__link:active{color:#4e8b1d;}</style>
|
|
</footer>
|
|
<script type="text/rocketscript">
|
|
(function(){
|
|
var twitterWidgets = document.createElement('script');
|
|
twitterWidgets.type = 'text/javascript';
|
|
twitterWidgets.async = true;
|
|
twitterWidgets.src = 'http://platform.twitter.com/widgets.js';
|
|
document.getElementsByTagName('head')[0].appendChild(twitterWidgets);
|
|
})();
|
|
</script>
|
|
</div>
|
|
</body>
|
|
</html>
|
|
|