Files
nexus/sreweekly/articles/456/08-uptime-status-pages-and-transparency-calculus.html
2026-09-12 17:23:01 +08:00

507 lines
17 KiB
HTML
Raw Permalink Blame History

This file contains invisible Unicode characters
This file contains invisible Unicode characters that are indistinguishable to humans but may be processed differently by a computer. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
<!DOCTYPE html>
<html class="no-js" lang="en">
<head>
<meta charset="utf-8">
<title>Uptime, status pages, and transparency calculus | Lawrence Jones</title>
<meta name="description"
content=" From the evergreen AWS status page to hardcoded 100% uptime, no one fully trusts a status page anymore. But why is this? Companies often start with good ...">
<meta name="viewport" content="width=device-width, initial-scale=1">
<!-- If this is an external_url, we want to redirect -->
<!-- Preload Google fonts, which is defined in sass -->
<link rel="preload" as="font" crossorigin href="https://fonts.gstatic.com/s/sourcesanspro/v14/6xK1dSBYKcSV-LCoeQqfX1RYOo3qPZ7nsDc.ttf">
<link rel="preload" as="font" crossorigin href="https://fonts.gstatic.com/s/sourcesanspro/v14/6xKwdSBYKcSV-LCoeQqfX1RYOo3qPZZclSds18E.ttf">
<link rel="preload" as="font" crossorigin href="https://fonts.gstatic.com/s/sourcesanspro/v14/6xK3dSBYKcSV-LCoeQqfX1RYOo3qOK7g.ttf">
<link rel="preload" as="font" crossorigin href="https://fonts.gstatic.com/s/sourcesanspro/v14/6xKydSBYKcSV-LCoeQqfX1RYOo3ig4vwlxdr.ttf">
<!--
Preload any font-awesome assets we might want to use
Identify these URLs by watching the network panel in Chrome when loading pages. Update
them whenever we change font-awesome version.
-->
<link rel="preload" as="font" crossorigin href="https://use.fontawesome.com/releases/v5.8.2/webfonts/fa-brands-400.woff2">
<link rel="preload" as="font" crossorigin href="https://use.fontawesome.com/releases/v5.8.2/webfonts/fa-solid-900.woff2">
<!-- CSS -->
<link rel="stylesheet" href="/assets/css/main.css">
<!-- Favicon -->
<link rel="shortcut icon" href="/assets/favicon.ico" type="image/x-icon">
<!-- RSS -->
<link rel="alternate" type="application/atom+xml" title="Lawrence Jones"
href="/feed.xml" />
<!--
Font Awesome
Configured to lazily load, so it doesn't block the page
-->
<link
rel="preload"
as="style"
onload="this.rel='stylesheet'"
href="https://use.fontawesome.com/releases/v5.8.2/css/all.css"
integrity="sha384-oS3vJWv+0UjzBfQzYUhtDYW+Pj2yciDJxpsK1OYPAYjqT085Qq/1cq5FLXAZQ7Ay"
crossorigin="anonymous">
<!-- KaTeX -->
<!-- Google Analytics, fast loading version -->
<script async src="https://www.googletagmanager.com/gtag/js?id=G-2FV47623W0"></script>
<script>
window.dataLayer = window.dataLayer || [];
function gtag(){dataLayer.push(arguments);}
gtag('js', new Date());
gtag('config', 'G-2FV47623W0');
</script>
<!-- Begin Jekyll SEO tag v2.8.0 -->
<title>Uptime, status pages, and transparency calculus</title>
<meta name="generator" content="Jekyll v4.2.2" />
<meta property="og:title" content="Uptime, status pages, and transparency calculus" />
<meta property="og:locale" content="en_US" />
<meta name="description" content="From the evergreen AWS status page to hardcoded 100% uptime, no one fully trusts a status page anymore. But why is this? Companies often start with good intentions, aiming for full transparency. So why do so many change along the way: what pressures people into an evergreen status page with poorly-reflective uptime numbers?" />
<meta property="og:description" content="From the evergreen AWS status page to hardcoded 100% uptime, no one fully trusts a status page anymore. But why is this? Companies often start with good intentions, aiming for full transparency. So why do so many change along the way: what pressures people into an evergreen status page with poorly-reflective uptime numbers?" />
<link rel="canonical" href="https://blog.lawrencejones.dev/status-pages/" />
<meta property="og:url" content="https://blog.lawrencejones.dev/status-pages/" />
<meta property="og:image" content="https://blog.lawrencejones.dev/assets/images/status-pages.png" />
<meta property="og:type" content="article" />
<meta property="article:published_time" content="2023-01-30T12:00:00+00:00" />
<meta name="twitter:card" content="summary_large_image" />
<meta property="twitter:image" content="https://blog.lawrencejones.dev/assets/images/status-pages.png" />
<meta property="twitter:title" content="Uptime, status pages, and transparency calculus" />
<meta name="twitter:site" content="@lawrjones" />
<script type="application/ld+json">
{"@context":"https://schema.org","@type":"BlogPosting","dateModified":"2023-01-30T12:00:00+00:00","datePublished":"2023-01-30T12:00:00+00:00","description":"From the evergreen AWS status page to hardcoded 100% uptime, no one fully trusts a status page anymore. But why is this? Companies often start with good intentions, aiming for full transparency. So why do so many change along the way: what pressures people into an evergreen status page with poorly-reflective uptime numbers?","headline":"Uptime, status pages, and transparency calculus","image":"https://blog.lawrencejones.dev/assets/images/status-pages.png","mainEntityOfPage":{"@type":"WebPage","@id":"https://blog.lawrencejones.dev/status-pages/"},"url":"https://blog.lawrencejones.dev/status-pages/"}</script>
<!-- End Jekyll SEO tag -->
</head>
<body>
<header class="site-header">
<div class="header-content">
<div class="branding">
<a href="/">
<img class="avatar" src="https://secure.gravatar.com/avatar/a3d694b39e0e33fc479832b00dc128dc?s=105" alt="Gravatar picture of Lawrence">
</a>
<h1 class="site-title">
<a href="/">Lawrence Jones</a>
</h1>
</div>
<nav class="site-nav">
<ul>
<li>
<a class="page-link" href="/about/">
About
</a>
</li>
<!-- Social icons from Font Awesome, if enabled -->
<li>
<a href="/feed.xml" title="Follow RSS feed">
<i class="fas fa-fw fa-rss"></i>
</a>
</li>
<li>
<a href="/cdn-cgi/l/email-protection#4a272f0a262b3d382f24292f2025242f39642e2f3c" title="Email">
<i class="fas fa-fw fa-envelope"></i>
</a>
</li>
<li>
<a href="https://github.com/lawrencejones" title="Follow on GitHub" target="_blank" rel="noopener noreferrer">
<i class="fab fa-fw fa-github"></i>
</a>
</li>
<li>
<a href="https://twitter.com/lawrjones" title="Follow on Twitter" target="_blank" rel="noopener noreferrer">
<i style="color: rgba(29,161,242,1.00);" class="fab fa-fw fa-twitter"></i>
</a>
</li>
<!-- Search bar -->
</ul>
</nav>
</div>
</header>
<div class="content">
<article>
<header style="background-image: url('/')">
<h1 class="title">Uptime, status pages, and transparency calculus</h1>
<p class="meta">
January 30, 2023
</p>
</header>
<section class="post-content">
<p>When you first create a status page, it’s probably because you want to
communicate outages to your customers. The faster you can share details about an
outage, the sooner your customers know what’s going on, and the more effectively
they can handle the outage.</p>
<p>Communicating promptly – in clear language – builds trust. And as a young
company with a customer centric focus, that’s your top priority.</p>
<p>So why is it that as an industry, we no longer fully trust the status page of
large service providers?</p>
<p>Take AWS, for example: an industry joke is that their status page is evergreen.
It got so bad that Corey Quinn created <a href="https://stop.lying.cloud/" target="_blank" rel="noopener noreferrer">stop.lying.cloud</a> as a
simplified, more ‘truthful’ version of the AWS status page. While discontinued
now, Corey’s site used to filter the ‘sea of green’ for the services that were
broken, helping AWS customers to quickly navigate components during stressful
outages.</p>
<p>In a similar theme, just this week Gergely Orosz observed <a href="https://twitter.com/GergelyOrosz/status/1617965847338975232" target="_blank" rel="noopener noreferrer">“Slack’s status page
reports 100% uptime since Feb 2022”</a> despite widespread DNS issues
and many customers seeing Slack blackouts. It seems clearly wrong to read 100%
if you’re one of the customers who were impacted, after you (presumably) fell
back to smoke signals and handwritten notes for communication earlier that week.</p>
<p>But if you follow that thread, you’ll find a Slack engineer hinting at a very
different perspective…</p>
<div style="min-height: 660px">
<blockquote class="twitter-tweet tw-align-center">
<p lang="en" dir="ltr">Still, that's the rules we're playing with. Does 100% uptime seem weird to me as an engineer? Sure. But the primary purpose of this number is to convey to customers whether they'll be receiving refunds or not.</p>— cooper b (@cooperb) <a href="https://twitter.com/cooperb/status/1617978304698646528?ref_src=twsrc%5Etfw" target="_blank" rel="noopener noreferrer">January 24, 2023</a>
</blockquote>
<script data-cfasync="false" src="/cdn-cgi/scripts/5c5dd728/cloudflare-static/email-decode.min.js"></script><script async="" src="https://platform.twitter.com/widgets.js" charset="utf-8"></script>
</div>
<p>What’s going on here, then? What does Cooper mean by “primary purpose … is to
convey to customers whether they’ll be receiving refunds”?</p>
<h2 id="b2b-and-elas">B2B and ELAs</h2>
<p>Slack is a B2B company, meaning their customers are businesses rather than
individuals. When B2B companies sell software, they sign enterprise license
agreements with their customers which include provisions around uptime.</p>
<p>Often termed service level agreements (SLAs), they are a commitment to a level
of service – usually expressed as ‘99.9% uptime’ or similar – that if breached,
entitles the customer to compensation.</p>
<p>This might seem unrelated until you realise that as a company grows, you need a
standardised way to communicate uptime to your customers that everyone –
business and customer – can agree is accurate, and relevant to SLA measurement
and compensation. Now perhaps because it’s already there, or because having a
‘status page’ that reports more downtime than an alternative legal measure isn’t
tenable, the status page often <em>becomes</em> the official measure.</p>
<p>From that moment on, it becomes difficult to use the page in the same way. As
publishing an incident or uptime statistics can expose the company to financial
penalties, you need an ever-increasing amount of buy-in from senior (and
sometimes non-technical) leadership before you can give an update. Even if the
end result is the same – and you end up publishing the update – adding
executives to the incident loop delays comms, often meaning customers have
already had to fend for themselves before you notify them.</p>
<p>Not as simple as it may have seemed. And it’s not just B2B customers who find
themselves in this situation.</p>
<h2 id="it-sucks-for-b2c-too">It sucks for B2C too</h2>
<p>While B2C companies rarely bake SLAs into contracts with individual customers,
they have a different set of challenges when it comes to publicly updating
status pages.</p>
<p>Firstly, and this concern is shared with any company that has large customer
base, publishing an incident is not free. If you have 1M customers and you
notify them all of an incident via your status page, just 0.1% of those
customers need open a support ticket for it to bury your ops team for days.</p>
<p>You might say “so what, who cares? that’s the price of doing business”. But the
truth is most incidents only impact a handful of customers, and if you notify
for every incident, you’re creating a huge amount of unnecessary stress and
worry for those who aren’t impacted. On a human level, if all those customers
spend just 1 minute reading your email, that sums to about 2 years of human life
spent on an issue that may not impact them.</p>
<p>As if that wasn’t enough of an incentive to carefully consider updates, being
‘too’ open can hurt your business beyond SLAs and wasted time.</p>
<h2 id="sales">Sales</h2>
<p>When in an RFP process with prospective customers, you’re in a negotiation where
the buyer will look for reasons to lower the price. It’s not unusual for the
buyer to look through a company’s status page and find incidents to strengthen
their argument: “you’ve had several incidents over the last few months, just
look at your status page!”.</p>
<p>For product people involved in these processes, this can be an uncomfortable
discussion. No doubt you published those updates to do right by your customers,
aiming to be transparent and help resolve the issue as quickly as possible, but
now a prospect – who would benefit from this behaviour if they become a customer
– is using it against you.</p>
<p>You might respond by being even more open about how you measure uptime, how your
SLA works, why the customer shouldn’t worry about this. I’ve been there before,
in one case reviewing our process for calculating uptime and building a data
model that calculated – from request logs – the exact uptime for each customer,
helping identify who qualified for credits and how much.</p>
<p>Sadly, this didn’t go as I’d hoped. The truth is that service quality is
extremely hard to quantify in a single number, and if you’ve got anything even
remotely accurate on a per-customer basis then it’s likely pretty complicated.
Sharing the details of this mechanism with the prospect, who was a non-technical
but legally savvy buyer, did not help. In fact they got so tangled in the
details that it made negotiation even more sticky, in a nasty situation where
being fully transparent had made the prospect trust us even less.</p>
<p>Finally, being open about your incidents can feel a bit like a mug’s game. No
matter your industry, you’ll have competitors who are less open than you about
their issues, keeping a spotless status page despite you knowing they have
frequent, severe outages. That can work against you in sales processes, or if
you’re unlucky enough to be in a regulated industry, can even have your
regulators question why you are so ‘bad’ in comparison.</p>
<h2 id="where-do-we-go-from-here">Where do we go from here?</h2>
<p>Clearly it is in everyone’s interest to have transparent, prompt communication
and useful/accurate uptime reporting for software. But I hope you see that by
applying penalties and building uptime into contracts, we’ve created a number of
incentives that work against this.</p>
<p>It is a real-world example of Goodhart’s Law, in that as soon as we began using
uptime as a target, it stopped being a useful measure.</p>
<p>So what is an ideal alternative? For me, as a naive engineer, I’d love to see
the industry start viewing clear and transparent communication in past incidents
as positive signal about a working relationship. After all, we know incidents
are a fact of life, and it’s much better to be honest about them than hide.</p>
<p>If we could couple that philosophy with break clauses in software contracts in
case of poor service, I think we’re be a good step away from the nickle-and-dime
culture of service credits that can compromise transparency. After all, if the
service is really that bad, surely you’d prefer to find another provider than
fight for a 10% refund?</p>
<p>Sadly, with large companies placing millions on the line, that’s a difficult
change. Lawyers write and sign-off on these contracts and are hired to minimise
company exposure, which makes it hard to view a software service contract as
more of a partnership than a transaction. Additionally, while this might work
well for services like Slack where downtime means you lose an hour of
productivity but soon bounce back, it’s less suited to critical infrastructure
that can seriously harm a business if down for more than a couple of hours.</p>
<p>Perhaps we can’t change the system, but I have hope we might change the game.
Two options that come to mind are increasing the perceived value of great
incident comms, or improving the tooling companies use to communicate to ease
the difficulties this post has outlined.</p>
<p>It’s something we spend a lot of time thinking about at incident.io, and already
have a few ideas in the pipeline. So watch this space!</p>
<p>
<em>
If you liked this post and want to see more, follow me on <a target="_blank" href="https://www.linkedin.com/in/lawrence2jones/" rel="noopener noreferrer">LinkedIn</a>.
</em>
</p>
</section>
</article>
<!-- Disqus -->
<!-- Post navigation -->
<div id="post-nav">
<div id="previous-post" class="post-nav-post">
<p>Previous post</p>
<a href="/ulid/">
Using ULIDs at incident.io
</a>
</div>
<div id="next-post" class="post-nav-post">
<p>Next post</p>
<a href="/catalog/">
Three months building a catalog
</a>
</div>
</div>
</div>
<footer class="site-footer">
<a href="https://www.linkedin.com/in/lawrence2jones/" target="_blank" rel="noopener noreferrer">
<i class="fab fa-linkedin" style="color: #0077b5;"></i> LinkedIn
</a>
</footer>
<script type="module" src="https://static.cloudflareinsights.com/beacon.min.js/v31edd6df95cf4e85bb4c19e7a9bdbcba1788362987495" integrity="sha512-iIg7k2xntmwu6/uSb5tpc/hySgZc4eoL31yB29W6tJFo2akwjPWcEqnCEdJvGexCL0KEQwVYv5BlowfhVz26hg==" data-cf-beacon='{"version":"2024.11.0","token":"292392cdf194479a91b07f11fd03de82","r":1,"spa":2}' crossorigin="anonymous"></script>
</body>
</html>