507 lines
17 KiB
HTML
507 lines
17 KiB
HTML
<!DOCTYPE html>
|
||
<html class="no-js" lang="en">
|
||
<head>
|
||
<meta charset="utf-8">
|
||
<title>Uptime, status pages, and transparency calculus | Lawrence Jones</title>
|
||
<meta name="description"
|
||
content=" From the evergreen AWS status page to hardcoded 100% uptime, no one fully trusts a status page anymore. But why is this? Companies often start with good ...">
|
||
<meta name="viewport" content="width=device-width, initial-scale=1">
|
||
|
||
|
||
<!-- If this is an external_url, we want to redirect -->
|
||
|
||
|
||
<!-- Preload Google fonts, which is defined in sass -->
|
||
|
||
<link rel="preload" as="font" crossorigin href="https://fonts.gstatic.com/s/sourcesanspro/v14/6xK1dSBYKcSV-LCoeQqfX1RYOo3qPZ7nsDc.ttf">
|
||
|
||
<link rel="preload" as="font" crossorigin href="https://fonts.gstatic.com/s/sourcesanspro/v14/6xKwdSBYKcSV-LCoeQqfX1RYOo3qPZZclSds18E.ttf">
|
||
|
||
<link rel="preload" as="font" crossorigin href="https://fonts.gstatic.com/s/sourcesanspro/v14/6xK3dSBYKcSV-LCoeQqfX1RYOo3qOK7g.ttf">
|
||
|
||
<link rel="preload" as="font" crossorigin href="https://fonts.gstatic.com/s/sourcesanspro/v14/6xKydSBYKcSV-LCoeQqfX1RYOo3ig4vwlxdr.ttf">
|
||
|
||
|
||
<!--
|
||
Preload any font-awesome assets we might want to use
|
||
|
||
Identify these URLs by watching the network panel in Chrome when loading pages. Update
|
||
them whenever we change font-awesome version.
|
||
-->
|
||
|
||
<link rel="preload" as="font" crossorigin href="https://use.fontawesome.com/releases/v5.8.2/webfonts/fa-brands-400.woff2">
|
||
|
||
<link rel="preload" as="font" crossorigin href="https://use.fontawesome.com/releases/v5.8.2/webfonts/fa-solid-900.woff2">
|
||
|
||
|
||
<!-- CSS -->
|
||
<link rel="stylesheet" href="/assets/css/main.css">
|
||
|
||
<!-- Favicon -->
|
||
<link rel="shortcut icon" href="/assets/favicon.ico" type="image/x-icon">
|
||
|
||
<!-- RSS -->
|
||
<link rel="alternate" type="application/atom+xml" title="Lawrence Jones"
|
||
href="/feed.xml" />
|
||
|
||
<!--
|
||
Font Awesome
|
||
|
||
Configured to lazily load, so it doesn't block the page
|
||
-->
|
||
<link
|
||
rel="preload"
|
||
as="style"
|
||
onload="this.rel='stylesheet'"
|
||
href="https://use.fontawesome.com/releases/v5.8.2/css/all.css"
|
||
integrity="sha384-oS3vJWv+0UjzBfQzYUhtDYW+Pj2yciDJxpsK1OYPAYjqT085Qq/1cq5FLXAZQ7Ay"
|
||
crossorigin="anonymous">
|
||
|
||
<!-- KaTeX -->
|
||
|
||
|
||
<!-- Google Analytics, fast loading version -->
|
||
|
||
<script async src="https://www.googletagmanager.com/gtag/js?id=G-2FV47623W0"></script>
|
||
<script>
|
||
window.dataLayer = window.dataLayer || [];
|
||
function gtag(){dataLayer.push(arguments);}
|
||
gtag('js', new Date());
|
||
|
||
gtag('config', 'G-2FV47623W0');
|
||
</script>
|
||
|
||
|
||
|
||
<!-- Begin Jekyll SEO tag v2.8.0 -->
|
||
<title>Uptime, status pages, and transparency calculus</title>
|
||
<meta name="generator" content="Jekyll v4.2.2" />
|
||
<meta property="og:title" content="Uptime, status pages, and transparency calculus" />
|
||
<meta property="og:locale" content="en_US" />
|
||
<meta name="description" content="From the evergreen AWS status page to hardcoded 100% uptime, no one fully trusts a status page anymore. But why is this? Companies often start with good intentions, aiming for full transparency. So why do so many change along the way: what pressures people into an evergreen status page with poorly-reflective uptime numbers?" />
|
||
<meta property="og:description" content="From the evergreen AWS status page to hardcoded 100% uptime, no one fully trusts a status page anymore. But why is this? Companies often start with good intentions, aiming for full transparency. So why do so many change along the way: what pressures people into an evergreen status page with poorly-reflective uptime numbers?" />
|
||
<link rel="canonical" href="https://blog.lawrencejones.dev/status-pages/" />
|
||
<meta property="og:url" content="https://blog.lawrencejones.dev/status-pages/" />
|
||
<meta property="og:image" content="https://blog.lawrencejones.dev/assets/images/status-pages.png" />
|
||
<meta property="og:type" content="article" />
|
||
<meta property="article:published_time" content="2023-01-30T12:00:00+00:00" />
|
||
<meta name="twitter:card" content="summary_large_image" />
|
||
<meta property="twitter:image" content="https://blog.lawrencejones.dev/assets/images/status-pages.png" />
|
||
<meta property="twitter:title" content="Uptime, status pages, and transparency calculus" />
|
||
<meta name="twitter:site" content="@lawrjones" />
|
||
<script type="application/ld+json">
|
||
{"@context":"https://schema.org","@type":"BlogPosting","dateModified":"2023-01-30T12:00:00+00:00","datePublished":"2023-01-30T12:00:00+00:00","description":"From the evergreen AWS status page to hardcoded 100% uptime, no one fully trusts a status page anymore. But why is this? Companies often start with good intentions, aiming for full transparency. So why do so many change along the way: what pressures people into an evergreen status page with poorly-reflective uptime numbers?","headline":"Uptime, status pages, and transparency calculus","image":"https://blog.lawrencejones.dev/assets/images/status-pages.png","mainEntityOfPage":{"@type":"WebPage","@id":"https://blog.lawrencejones.dev/status-pages/"},"url":"https://blog.lawrencejones.dev/status-pages/"}</script>
|
||
<!-- End Jekyll SEO tag -->
|
||
|
||
|
||
</head>
|
||
|
||
<body>
|
||
<header class="site-header">
|
||
<div class="header-content">
|
||
<div class="branding">
|
||
|
||
<a href="/">
|
||
<img class="avatar" src="https://secure.gravatar.com/avatar/a3d694b39e0e33fc479832b00dc128dc?s=105" alt="Gravatar picture of Lawrence">
|
||
</a>
|
||
|
||
<h1 class="site-title">
|
||
<a href="/">Lawrence Jones</a>
|
||
</h1>
|
||
</div>
|
||
<nav class="site-nav">
|
||
<ul>
|
||
|
||
|
||
|
||
|
||
|
||
<li>
|
||
<a class="page-link" href="/about/">
|
||
About
|
||
</a>
|
||
</li>
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
<!-- Social icons from Font Awesome, if enabled -->
|
||
|
||
<li>
|
||
<a href="/feed.xml" title="Follow RSS feed">
|
||
<i class="fas fa-fw fa-rss"></i>
|
||
</a>
|
||
</li>
|
||
|
||
|
||
|
||
<li>
|
||
<a href="/cdn-cgi/l/email-protection#4a272f0a262b3d382f24292f2025242f39642e2f3c" title="Email">
|
||
<i class="fas fa-fw fa-envelope"></i>
|
||
</a>
|
||
</li>
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
<li>
|
||
<a href="https://github.com/lawrencejones" title="Follow on GitHub" target="_blank" rel="noopener noreferrer">
|
||
<i class="fab fa-fw fa-github"></i>
|
||
</a>
|
||
</li>
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
<li>
|
||
<a href="https://twitter.com/lawrjones" title="Follow on Twitter" target="_blank" rel="noopener noreferrer">
|
||
<i style="color: rgba(29,161,242,1.00);" class="fab fa-fw fa-twitter"></i>
|
||
</a>
|
||
</li>
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
<!-- Search bar -->
|
||
|
||
</ul>
|
||
</nav>
|
||
</div>
|
||
</header>
|
||
|
||
<div class="content">
|
||
<article>
|
||
<header style="background-image: url('/')">
|
||
<h1 class="title">Uptime, status pages, and transparency calculus</h1>
|
||
|
||
<p class="meta">
|
||
January 30, 2023
|
||
|
||
</p>
|
||
</header>
|
||
<section class="post-content">
|
||
<p>When you first create a status page, it’s probably because you want to
|
||
communicate outages to your customers. The faster you can share details about an
|
||
outage, the sooner your customers know what’s going on, and the more effectively
|
||
they can handle the outage.</p>
|
||
|
||
<p>Communicating promptly – in clear language – builds trust. And as a young
|
||
company with a customer centric focus, that’s your top priority.</p>
|
||
|
||
<p>So why is it that as an industry, we no longer fully trust the status page of
|
||
large service providers?</p>
|
||
|
||
<p>Take AWS, for example: an industry joke is that their status page is evergreen.
|
||
It got so bad that Corey Quinn created <a href="https://stop.lying.cloud/" target="_blank" rel="noopener noreferrer">stop.lying.cloud</a> as a
|
||
simplified, more ‘truthful’ version of the AWS status page. While discontinued
|
||
now, Corey’s site used to filter the ‘sea of green’ for the services that were
|
||
broken, helping AWS customers to quickly navigate components during stressful
|
||
outages.</p>
|
||
|
||
<p>In a similar theme, just this week Gergely Orosz observed <a href="https://twitter.com/GergelyOrosz/status/1617965847338975232" target="_blank" rel="noopener noreferrer">“Slack’s status page
|
||
reports 100% uptime since Feb 2022”</a> despite widespread DNS issues
|
||
and many customers seeing Slack blackouts. It seems clearly wrong to read 100%
|
||
if you’re one of the customers who were impacted, after you (presumably) fell
|
||
back to smoke signals and handwritten notes for communication earlier that week.</p>
|
||
|
||
<p>But if you follow that thread, you’ll find a Slack engineer hinting at a very
|
||
different perspective…</p>
|
||
|
||
<div style="min-height: 660px">
|
||
<blockquote class="twitter-tweet tw-align-center">
|
||
<p lang="en" dir="ltr">Still, that's the rules we're playing with. Does 100% uptime seem weird to me as an engineer? Sure. But the primary purpose of this number is to convey to customers whether they'll be receiving refunds or not.</p>— cooper b (@cooperb) <a href="https://twitter.com/cooperb/status/1617978304698646528?ref_src=twsrc%5Etfw" target="_blank" rel="noopener noreferrer">January 24, 2023</a>
|
||
</blockquote>
|
||
<script data-cfasync="false" src="/cdn-cgi/scripts/5c5dd728/cloudflare-static/email-decode.min.js"></script><script async="" src="https://platform.twitter.com/widgets.js" charset="utf-8"></script>
|
||
</div>
|
||
|
||
<p>What’s going on here, then? What does Cooper mean by “primary purpose … is to
|
||
convey to customers whether they’ll be receiving refunds”?</p>
|
||
|
||
<h2 id="b2b-and-elas">B2B and ELAs</h2>
|
||
|
||
<p>Slack is a B2B company, meaning their customers are businesses rather than
|
||
individuals. When B2B companies sell software, they sign enterprise license
|
||
agreements with their customers which include provisions around uptime.</p>
|
||
|
||
<p>Often termed service level agreements (SLAs), they are a commitment to a level
|
||
of service – usually expressed as ‘99.9% uptime’ or similar – that if breached,
|
||
entitles the customer to compensation.</p>
|
||
|
||
<p>This might seem unrelated until you realise that as a company grows, you need a
|
||
standardised way to communicate uptime to your customers that everyone –
|
||
business and customer – can agree is accurate, and relevant to SLA measurement
|
||
and compensation. Now perhaps because it’s already there, or because having a
|
||
‘status page’ that reports more downtime than an alternative legal measure isn’t
|
||
tenable, the status page often <em>becomes</em> the official measure.</p>
|
||
|
||
<p>From that moment on, it becomes difficult to use the page in the same way. As
|
||
publishing an incident or uptime statistics can expose the company to financial
|
||
penalties, you need an ever-increasing amount of buy-in from senior (and
|
||
sometimes non-technical) leadership before you can give an update. Even if the
|
||
end result is the same – and you end up publishing the update – adding
|
||
executives to the incident loop delays comms, often meaning customers have
|
||
already had to fend for themselves before you notify them.</p>
|
||
|
||
<p>Not as simple as it may have seemed. And it’s not just B2B customers who find
|
||
themselves in this situation.</p>
|
||
|
||
<h2 id="it-sucks-for-b2c-too">It sucks for B2C too</h2>
|
||
|
||
<p>While B2C companies rarely bake SLAs into contracts with individual customers,
|
||
they have a different set of challenges when it comes to publicly updating
|
||
status pages.</p>
|
||
|
||
<p>Firstly, and this concern is shared with any company that has large customer
|
||
base, publishing an incident is not free. If you have 1M customers and you
|
||
notify them all of an incident via your status page, just 0.1% of those
|
||
customers need open a support ticket for it to bury your ops team for days.</p>
|
||
|
||
<p>You might say “so what, who cares? that’s the price of doing business”. But the
|
||
truth is most incidents only impact a handful of customers, and if you notify
|
||
for every incident, you’re creating a huge amount of unnecessary stress and
|
||
worry for those who aren’t impacted. On a human level, if all those customers
|
||
spend just 1 minute reading your email, that sums to about 2 years of human life
|
||
spent on an issue that may not impact them.</p>
|
||
|
||
<p>As if that wasn’t enough of an incentive to carefully consider updates, being
|
||
‘too’ open can hurt your business beyond SLAs and wasted time.</p>
|
||
|
||
<h2 id="sales">Sales</h2>
|
||
|
||
<p>When in an RFP process with prospective customers, you’re in a negotiation where
|
||
the buyer will look for reasons to lower the price. It’s not unusual for the
|
||
buyer to look through a company’s status page and find incidents to strengthen
|
||
their argument: “you’ve had several incidents over the last few months, just
|
||
look at your status page!”.</p>
|
||
|
||
<p>For product people involved in these processes, this can be an uncomfortable
|
||
discussion. No doubt you published those updates to do right by your customers,
|
||
aiming to be transparent and help resolve the issue as quickly as possible, but
|
||
now a prospect – who would benefit from this behaviour if they become a customer
|
||
– is using it against you.</p>
|
||
|
||
<p>You might respond by being even more open about how you measure uptime, how your
|
||
SLA works, why the customer shouldn’t worry about this. I’ve been there before,
|
||
in one case reviewing our process for calculating uptime and building a data
|
||
model that calculated – from request logs – the exact uptime for each customer,
|
||
helping identify who qualified for credits and how much.</p>
|
||
|
||
<p>Sadly, this didn’t go as I’d hoped. The truth is that service quality is
|
||
extremely hard to quantify in a single number, and if you’ve got anything even
|
||
remotely accurate on a per-customer basis then it’s likely pretty complicated.
|
||
Sharing the details of this mechanism with the prospect, who was a non-technical
|
||
but legally savvy buyer, did not help. In fact they got so tangled in the
|
||
details that it made negotiation even more sticky, in a nasty situation where
|
||
being fully transparent had made the prospect trust us even less.</p>
|
||
|
||
<p>Finally, being open about your incidents can feel a bit like a mug’s game. No
|
||
matter your industry, you’ll have competitors who are less open than you about
|
||
their issues, keeping a spotless status page despite you knowing they have
|
||
frequent, severe outages. That can work against you in sales processes, or if
|
||
you’re unlucky enough to be in a regulated industry, can even have your
|
||
regulators question why you are so ‘bad’ in comparison.</p>
|
||
|
||
<h2 id="where-do-we-go-from-here">Where do we go from here?</h2>
|
||
|
||
<p>Clearly it is in everyone’s interest to have transparent, prompt communication
|
||
and useful/accurate uptime reporting for software. But I hope you see that by
|
||
applying penalties and building uptime into contracts, we’ve created a number of
|
||
incentives that work against this.</p>
|
||
|
||
<p>It is a real-world example of Goodhart’s Law, in that as soon as we began using
|
||
uptime as a target, it stopped being a useful measure.</p>
|
||
|
||
<p>So what is an ideal alternative? For me, as a naive engineer, I’d love to see
|
||
the industry start viewing clear and transparent communication in past incidents
|
||
as positive signal about a working relationship. After all, we know incidents
|
||
are a fact of life, and it’s much better to be honest about them than hide.</p>
|
||
|
||
<p>If we could couple that philosophy with break clauses in software contracts in
|
||
case of poor service, I think we’re be a good step away from the nickle-and-dime
|
||
culture of service credits that can compromise transparency. After all, if the
|
||
service is really that bad, surely you’d prefer to find another provider than
|
||
fight for a 10% refund?</p>
|
||
|
||
<p>Sadly, with large companies placing millions on the line, that’s a difficult
|
||
change. Lawyers write and sign-off on these contracts and are hired to minimise
|
||
company exposure, which makes it hard to view a software service contract as
|
||
more of a partnership than a transaction. Additionally, while this might work
|
||
well for services like Slack where downtime means you lose an hour of
|
||
productivity but soon bounce back, it’s less suited to critical infrastructure
|
||
that can seriously harm a business if down for more than a couple of hours.</p>
|
||
|
||
<p>Perhaps we can’t change the system, but I have hope we might change the game.
|
||
Two options that come to mind are increasing the perceived value of great
|
||
incident comms, or improving the tooling companies use to communicate to ease
|
||
the difficulties this post has outlined.</p>
|
||
|
||
<p>It’s something we spend a lot of time thinking about at incident.io, and already
|
||
have a few ideas in the pipeline. So watch this space!</p>
|
||
|
||
|
||
<p>
|
||
<em>
|
||
|
||
If you liked this post and want to see more, follow me on <a target="_blank" href="https://www.linkedin.com/in/lawrence2jones/" rel="noopener noreferrer">LinkedIn</a>.
|
||
</em>
|
||
</p>
|
||
</section>
|
||
</article>
|
||
|
||
<!-- Disqus -->
|
||
|
||
|
||
<!-- Post navigation -->
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
<div id="post-nav">
|
||
|
||
<div id="previous-post" class="post-nav-post">
|
||
<p>Previous post</p>
|
||
<a href="/ulid/">
|
||
Using ULIDs at incident.io
|
||
</a>
|
||
</div>
|
||
|
||
|
||
<div id="next-post" class="post-nav-post">
|
||
<p>Next post</p>
|
||
<a href="/catalog/">
|
||
Three months building a catalog
|
||
</a>
|
||
</div>
|
||
|
||
</div>
|
||
|
||
|
||
|
||
</div>
|
||
|
||
|
||
|
||
<footer class="site-footer">
|
||
<a href="https://www.linkedin.com/in/lawrence2jones/" target="_blank" rel="noopener noreferrer">
|
||
<i class="fab fa-linkedin" style="color: #0077b5;"></i> LinkedIn
|
||
</a>
|
||
</footer>
|
||
|
||
|
||
<script type="module" src="https://static.cloudflareinsights.com/beacon.min.js/v31edd6df95cf4e85bb4c19e7a9bdbcba1788362987495" integrity="sha512-iIg7k2xntmwu6/uSb5tpc/hySgZc4eoL31yB29W6tJFo2akwjPWcEqnCEdJvGexCL0KEQwVYv5BlowfhVz26hg==" data-cf-beacon='{"version":"2024.11.0","token":"292392cdf194479a91b07f11fd03de82","r":1,"spa":2}' crossorigin="anonymous"></script>
|
||
</body>
|
||
</html>
|