195 lines
8.8 KiB
HTML
195 lines
8.8 KiB
HTML
<!DOCTYPE html>
|
||
<html xmlns="http://www.w3.org/1999/xhtml" xml:lang="en-us" lang="en-us">
|
||
<head>
|
||
<link href="http://gmpg.org/xfn/11" rel="profile">
|
||
<meta http-equiv="content-type" content="text/html; charset=utf-8">
|
||
<meta name="generator" content="Hugo 0.60.1" />
|
||
|
||
|
||
<meta name="viewport" content="width=device-width, initial-scale=1.0, maximum-scale=1">
|
||
|
||
|
||
<title>Production readiness · jbd.dev</title>
|
||
|
||
|
||
|
||
<link rel="stylesheet" href="https://jbd.dev/css/custom.css">
|
||
<link rel="stylesheet" href="https://jbd.dev/css/syntax.css">
|
||
<link rel="stylesheet" href="https://jbd.dev/css/style.css">
|
||
|
||
|
||
|
||
<link rel="apple-touch-icon-precomposed" sizes="144x144" href="/apple-touch-icon-144-precomposed.png">
|
||
<link rel="shortcut icon" type="image/png" href="/favicon.png">
|
||
|
||
|
||
<link href="/index.xml" rel="alternate" type="application/rss+xml" title="jbd.dev" />
|
||
|
||
<script src="https://code.jquery.com/jquery-1.12.0.min.js"></script>
|
||
<script src="/js/anchorify.min.js"></script>
|
||
<script>
|
||
jQuery(window).load(function () {
|
||
jQuery('.post-content').anchorify();
|
||
});
|
||
</script>
|
||
|
||
</head>
|
||
|
||
|
||
<script async src="https://www.googletagmanager.com/gtag/js?id=UA-134230136-1"></script>
|
||
<script>
|
||
window.dataLayer = window.dataLayer || [];
|
||
function gtag(){dataLayer.push(arguments);}
|
||
gtag('js', new Date());
|
||
|
||
gtag('config', 'UA-134230136-1');
|
||
</script>
|
||
|
||
<body class=" ">
|
||
<div class="sidebar">
|
||
<div class="container">
|
||
<h1>
|
||
<a href="https://jbd.dev/">jbd.dev</a></h1>
|
||
|
||
<h2>linux</h2>
|
||
<ul class="sidebar-nav">
|
||
|
||
<li><a href="/numa/">NUMA</a></li>
|
||
<li><a href="/core-dumps/">Core dumps</a></li>
|
||
<li><a href="/persistent-disks/">Google: Persistent disks</a></li>
|
||
</ul>
|
||
<h2>practices</h2>
|
||
<ul class="sidebar-nav">
|
||
|
||
<li><a href="/prod-readiness/">Production readiness</a></li>
|
||
<li><a href="/benchmarks-are-hard/">Benchmarks are hard</a></li>
|
||
<li><a href="/microservices-instrumentation/">Microservices instrumentation</a></li>
|
||
<li><a href="/debugging-latency/">Debugging latency</a></li>
|
||
<li><a href="/sre/">Google: SRE</a></li>
|
||
</ul>
|
||
|
||
<p class="small footnote">
|
||
Written by <a href="https://twitter.com/rakyll/">JBD</a> mostly in San Francisco, sometimes in the air.
|
||
See <a href="/about">about</a> for more and contact information.
|
||
</p>
|
||
|
||
</div>
|
||
</div>
|
||
|
||
|
||
|
||
|
||
<div class="content container">
|
||
<div class="post">
|
||
<h1>Production readiness</h1>
|
||
<div class="post-content"><p>Have you ever launched a new service to production?
|
||
Have you ever been maintaining a production service?
|
||
If you answer “yes” to one of these questions, have you
|
||
been guided during the process? What's good or bad to do
|
||
in production? And how do you transfer knowledge when new
|
||
team members want to release production services or
|
||
take the ownership of existing services?</p>
|
||
<p>Most companies end up having organically grown
|
||
approaches when it comes to production practices.
|
||
Each team would figure out their tools and best practices
|
||
themselves with trial-error. This reality often has
|
||
a real tax not only on the success of the projects
|
||
but also
|
||
on engineers.</p>
|
||
<p>Trial-error culture creates an environment where
|
||
finger pointing and blaming is more common.
|
||
Once these behaviors are common,
|
||
it becomes harder to learn from mistakes or
|
||
not to repeat them again.</p>
|
||
<p>Successful organizations:</p>
|
||
<ul>
|
||
<li>acknowledge the need of production guidelines</li>
|
||
<li>spend time on researching practices that apply to them</li>
|
||
<li>start having production readiness discussions when designing new systems or components</li>
|
||
<li>enforce production readiness practices</li>
|
||
</ul>
|
||
<p>Production readiness involve a “review” process.
|
||
Reviews can be a
|
||
checklist or a questionnaire. Reviews can be done manually, automatically or both.
|
||
Organizations can produce checklist
|
||
templates rather than a static list of requirements
|
||
that can be customized based on the needs. By doing
|
||
so, it is possible to give engineers a way to inherit
|
||
knowledge but also enough flexibility when it is required.</p>
|
||
<h2 id="when-to-review-a-service-for-production-readiness">When to review a service for production readiness?</h2>
|
||
<p>Production readiness reviews is not only useful right
|
||
before pushing to production, they can be a protocol when
|
||
handing off operational responsibilities to a different
|
||
team or to a new hire. Use reviews when:</p>
|
||
<ul>
|
||
<li>Launching a new production service.</li>
|
||
<li>Handing off the operations of an existing production
|
||
service to another team such as SRE.</li>
|
||
<li>Handing off the operations of an existing production
|
||
service to new individuals.</li>
|
||
<li>Preparing oncall support.</li>
|
||
</ul>
|
||
<h2 id="production-readiness-checklists">Production readiness checklists</h2>
|
||
<p>A while ago, I <a href="https://medium.com/google-cloud/production-guideline-9d5d10c8f1e">published</a>
|
||
an example checklist for production
|
||
readiness as an example of what they can cover.
|
||
Even though the list came to existence when working with
|
||
Google Cloud customers, it is useful and applicable
|
||
outside of Google Cloud.</p>
|
||
<h3 id="design-and-development">Design and Development</h3>
|
||
<ul>
|
||
<li>Have reproducible builds, your build shouldn’t require access to external services and shouldn’t be affected by an outage of an external system.</li>
|
||
<li>Define and set SLOs for your service at design time.</li>
|
||
<li>Document the availability expectations of external services you depend on.</li>
|
||
<li>Avoid single points of failures by not depending on single global resource. Have the resource replicated or have a proper fallback (e.g. hardcoded value) when resource is not available.</li>
|
||
</ul>
|
||
<h3 id="configuration-management">Configuration Management</h3>
|
||
<ul>
|
||
<li>Static, small and non-secret configuration can be command-line flags. Use a configuration delivery service for everything else.</li>
|
||
<li>Dynamic configuration should have a reasonable fallback in the case of unavailability of the configuration system.</li>
|
||
<li>Development environment configuration shouldn’t inherit from production configuration. This may lead access to production services from development and can cause privacy issues and data leaks.</li>
|
||
<li>Document what can be configured dynamically and explain the fallback behavior if configuration delivery system is not available.</li>
|
||
</ul>
|
||
<h3 id="release-management">Release Management</h3>
|
||
<ul>
|
||
<li>Document all details about your release process. Document how releases affect SLOs (e.g. temporary higher latency due to cache misses).</li>
|
||
<li>Document your canary release process.</li>
|
||
<li>Have a canary analysis plan and setup mechanisms to automatically revert canaries if possible.</li>
|
||
<li>Ensure rollbacks can use the same process that rollouts use.</li>
|
||
</ul>
|
||
<h3 id="observability">Observability</h3>
|
||
<ul>
|
||
<li>Ensure the collection of metrics that are required by your SLOs are collected and exported from your binaries.</li>
|
||
<li>Make sure client- and server-side of the observability data can be differentiated. This is important to debug issues in production.</li>
|
||
<li>Tune alerts to reduce toil, for example remove alerts triggered by the routine events.</li>
|
||
<li>Include underlying platform metrics in your dashboards. Setup alerting for your external service dependencies.</li>
|
||
<li>Always propagate the incoming trace context.
|
||
Even if you are not participating in the trace, this will allow lower-level services to debug debug production issues.</li>
|
||
</ul>
|
||
<h3 id="security-and-protection">Security and Protection</h3>
|
||
<ul>
|
||
<li>Make sure all external requests are encrypted.</li>
|
||
<li>Make sure your production projects have proper IAM configuration.</li>
|
||
<li>Use networks within projects to isolate groups of VM instances.</li>
|
||
<li>Use VPN to securely connect remote networks.</li>
|
||
<li>Document and monitor user data access. Ensure that all user data access is logged and audited.</li>
|
||
<li>Ensure debugging endpoints are limited by ACL.</li>
|
||
<li>Sanitize user input. Have payload size restrictions for user input.</li>
|
||
<li>Ensure your service can block incoming traffic selectively per user. This allows to block the abuse cases without impacting other users.</li>
|
||
<li>Avoid external endpoints that triggers a large number of internal fan-outs.</li>
|
||
</ul>
|
||
<h3 id="capacity-planning">Capacity planning</h3>
|
||
<ul>
|
||
<li>Document how your service scales. Examples: number of users, size of incoming payload, number of incoming messages.</li>
|
||
<li>Document resource requirements for your service. Examples: number of dedicated VM instances, number of Spanner instances, specialized hardware such as GPUs or TPUs.</li>
|
||
<li>Document resource constraints: resource type, region, etc.</li>
|
||
<li>Document quota restrictions to create new resources. For example, document the rate limit of GCE API if you are creating new instances via the API.</li>
|
||
<li>Consider having load tests for performance regressions where possible.</li>
|
||
</ul>
|
||
</div>
|
||
</div>
|
||
</div>
|
||
|
||
</body>
|
||
</html>
|