192 lines
60 KiB
HTML
192 lines
60 KiB
HTML
<!DOCTYPE html><html lang="en"><head><meta charSet="utf-8"/><meta name="viewport" content="width=device-width, initial-scale=1"/><meta name="theme-color" content="#111111"/><meta name="user-signed-in" content="false"/><title>The Real Failure Rate of EBS — PlanetScale</title><meta name="description" content="Our experience running AWS EBS at scale for critical workloads"/><meta name="robots"/><meta property="og:url" content="https://planetscale.com/blog/the-real-fail-rate-of-ebs"/><meta property="og:type" content="website"/><meta property="og:title" content="The Real Failure Rate of EBS — PlanetScale"/><meta property="og:image" content="https://planetscale.com/assets/the-real-fail-rate-of-ebs-social-Cl1a18Am.jpg"/><meta property="og:description" content="Our experience running AWS EBS at scale for critical workloads"/><meta property="twitter:card" content="summary_large_image"/><meta property="twitter:site" content="@PlanetScale"/><meta property="twitter:creator" content="@PlanetScale"/><meta property="twitter:url" content="https://planetscale.com/blog/the-real-fail-rate-of-ebs"/><meta property="twitter:title" content="The Real Failure Rate of EBS — PlanetScale"/><meta property="twitter:description" content="Our experience running AWS EBS at scale for critical workloads"/><meta property="twitter:image" content="https://planetscale.com/assets/the-real-fail-rate-of-ebs-social-Cl1a18Am.jpg"/><link rel="canonical" href="https://planetscale.com/blog/the-real-fail-rate-of-ebs"/><link rel="preconnect" href="https://planetscale-images.imgix.net"/><link nonce="EZdXajdlI2DRCjbnoKJCwfF5KNgWOgdluJsujYSRAY4=" rel="icon" href="/favicon.ico" type="image/x-icon" sizes="16x16"/><link nonce="EZdXajdlI2DRCjbnoKJCwfF5KNgWOgdluJsujYSRAY4=" rel="icon" href="/icon.png" type="image/png" sizes="32x32"/><link nonce="EZdXajdlI2DRCjbnoKJCwfF5KNgWOgdluJsujYSRAY4=" rel="apple-touch-icon" href="/apple-touch-icon.png" type="image/png" sizes="32x32"/><link nonce="EZdXajdlI2DRCjbnoKJCwfF5KNgWOgdluJsujYSRAY4=" rel="manifest" href="/manifest.webmanifest"/><link rel="modulepreload" href="/assets/entry.client-3vubyXrk.js" nonce="EZdXajdlI2DRCjbnoKJCwfF5KNgWOgdluJsujYSRAY4="/><link rel="modulepreload" href="/assets/jsx-runtime-DwfQwkRq.js" nonce="EZdXajdlI2DRCjbnoKJCwfF5KNgWOgdluJsujYSRAY4="/><link rel="modulepreload" href="/assets/components-_bNmAApg.js" nonce="EZdXajdlI2DRCjbnoKJCwfF5KNgWOgdluJsujYSRAY4="/><link rel="modulepreload" href="/assets/_.well-known_.mcp.server-card_.json_-Sx7XeH3e.js" nonce="EZdXajdlI2DRCjbnoKJCwfF5KNgWOgdluJsujYSRAY4="/><link rel="modulepreload" href="/assets/index-mKTXLmHu.js" nonce="EZdXajdlI2DRCjbnoKJCwfF5KNgWOgdluJsujYSRAY4="/><link rel="modulepreload" href="/assets/errorBoundaries-DhW4jVYt.js" nonce="EZdXajdlI2DRCjbnoKJCwfF5KNgWOgdluJsujYSRAY4="/><link rel="modulepreload" href="/assets/root-DbOv4-98.js" nonce="EZdXajdlI2DRCjbnoKJCwfF5KNgWOgdluJsujYSRAY4="/><link rel="modulepreload" href="/assets/lib-Dg89tQ22.js" nonce="EZdXajdlI2DRCjbnoKJCwfF5KNgWOgdluJsujYSRAY4="/><link rel="modulepreload" href="/assets/analytics.client-DM6E8o1h.js" nonce="EZdXajdlI2DRCjbnoKJCwfF5KNgWOgdluJsujYSRAY4="/><link rel="modulepreload" href="/assets/SiteHeader-C2U5gvDH.js" nonce="EZdXajdlI2DRCjbnoKJCwfF5KNgWOgdluJsujYSRAY4="/><link rel="modulepreload" href="/assets/current-9yDxj94E.js" nonce="EZdXajdlI2DRCjbnoKJCwfF5KNgWOgdluJsujYSRAY4="/><link rel="modulepreload" href="/assets/clsx-eT0YPcGk.js" nonce="EZdXajdlI2DRCjbnoKJCwfF5KNgWOgdluJsujYSRAY4="/><link rel="modulepreload" href="/assets/bugs-38ilEoW0.js" nonce="EZdXajdlI2DRCjbnoKJCwfF5KNgWOgdluJsujYSRAY4="/><link rel="modulepreload" href="/assets/keyboard-D-uXZORL.js" nonce="EZdXajdlI2DRCjbnoKJCwfF5KNgWOgdluJsujYSRAY4="/><link rel="modulepreload" href="/assets/use-tab-direction-dKm-S3Ck.js" nonce="EZdXajdlI2DRCjbnoKJCwfF5KNgWOgdluJsujYSRAY4="/><link rel="modulepreload" href="/assets/blog-pXH7ptHJ.js" nonce="EZdXajdlI2DRCjbnoKJCwfF5KNgWOgdluJsujYSRAY4="/><link rel="modulepreload" href="/assets/blog._slug-Ch_92qsH.js" nonce="EZdXajdlI2DRCjbnoKJCwfF5KNgWOgdluJsujYSRAY4="/><link rel="modulepreload" href="/assets/ContentImage-Dh6VEOUl.js" nonce="EZdXajdlI2DRCjbnoKJCwfF5KNgWOgdluJsujYSRAY4="/><link rel="modulepreload" href="/assets/BlogCategoryLink-DmQyn0gp.js" nonce="EZdXajdlI2DRCjbnoKJCwfF5KNgWOgdluJsujYSRAY4="/><link rel="modulepreload" href="/assets/Details-BSB_b6hI.js" nonce="EZdXajdlI2DRCjbnoKJCwfF5KNgWOgdluJsujYSRAY4="/><link rel="modulepreload" href="/assets/Skittle-CDFOPRjH.js" nonce="EZdXajdlI2DRCjbnoKJCwfF5KNgWOgdluJsujYSRAY4="/><link rel="modulepreload" href="/assets/SiteFooter-B2Gq9u2j.js" nonce="EZdXajdlI2DRCjbnoKJCwfF5KNgWOgdluJsujYSRAY4="/><link rel="modulepreload" href="/assets/Vimeo-00PQJDli.js" nonce="EZdXajdlI2DRCjbnoKJCwfF5KNgWOgdluJsujYSRAY4="/><link rel="modulepreload" href="/assets/YouTube-CMfaljVr.js" nonce="EZdXajdlI2DRCjbnoKJCwfF5KNgWOgdluJsujYSRAY4="/><link rel="modulepreload" href="/assets/date-CJTFH3uT.js" nonce="EZdXajdlI2DRCjbnoKJCwfF5KNgWOgdluJsujYSRAY4="/><link rel="modulepreload" href="/assets/use-inert-others-BMJ6-xOX.js" nonce="EZdXajdlI2DRCjbnoKJCwfF5KNgWOgdluJsujYSRAY4="/><link rel="modulepreload" href="/assets/description-Cf6FZmDe.js" nonce="EZdXajdlI2DRCjbnoKJCwfF5KNgWOgdluJsujYSRAY4="/><link rel="modulepreload" href="/assets/use-is-mounted-uQsUZyP9.js" nonce="EZdXajdlI2DRCjbnoKJCwfF5KNgWOgdluJsujYSRAY4="/><link rel="modulepreload" href="/assets/types-DvonrUFF.js" nonce="EZdXajdlI2DRCjbnoKJCwfF5KNgWOgdluJsujYSRAY4="/><link rel="stylesheet" href="/assets/styles-ns8XBZ1D.css"/><script nonce="EZdXajdlI2DRCjbnoKJCwfF5KNgWOgdluJsujYSRAY4=">window.ENV = {"IMAGE_CDN":"https://planetscale-images.imgix.net","IMAGE_CDN_ENABLED":"true","INTERNAL_API":"https://api.planetscale.com","RELEASE":"117b8aaf-965c-42bc-b013-5f72770de4d9","SENTRY_DSN":"https://bd81903b44804e22a06bdc0c1a91b303@o499952.ingest.us.sentry.io/4504531942572032"}</script></head><body class="flex min-h-screen flex-col"><div class="bg-neki px-3 py-1 text-center font-medium text-gray-900 dark:font-semibold"><span>Neki, sharded Postgres, is now available.</span> <span class="whitespace-nowrap"><a href="https://auth.planetscale.com/sign-up" class="whitespace-nowrap bg-gray-900 px-sm font-semibold text-white">Get started</a></span></div><header class="relative mb-6 mt-4 bg-primary"><div class="flex flex-col gap-y-3 px-3 sm:px-5 container max-w-7xl"><div class="grid w-full grid-cols-[auto_1fr] grid-rows-1 items-center lg:items-start lg:gap-3"><a aria-label="Go to homepage" class="col-start-1 col-end-2 h-4 w-4 rounded-full text-primary lg:hidden" href="/" data-discover="true"><svg xmlns="http://www.w3.org/2000/svg" width="32" height="32" fill="none" viewBox="0 0 40 40"><path fill="currentColor" d="M0 20C0 8.954 8.954 0 20 0c8.121 0 15.112 4.84 18.245 11.794l-26.45 26.45a20 20 0 0 1-3.225-1.83L24.984 20H20L5.858 34.142A19.94 19.94 0 0 1 0 20M39.999 20.007 20.006 40c11.04-.004 19.99-8.953 19.993-19.993"></path></svg></a><div class="group col-start-2 col-end-3 row-start-1 flex shrink-0 items-center justify-end gap-1.5 lg:gap-3"><div class="flex flex-row gap-2 lg:flex-col lg:gap-1 xl:flex-row"><div class="flex items-center justify-end gap-1 lg:h-4"><a href="https://auth.planetscale.com/sign-in" class="font-semibold text-primary hover:text-orange">Sign in</a></div><div class="flex items-center justify-end gap-0.5 lg:h-4"><form class="btn-sm hidden sm:inline-flex" action="/api/demo-sessions" method="post"><button type="submit" class="btn btn-outline btn-sm hidden sm:inline-flex">View sandbox</button></form><a class="btn btn-sm" href="/contact" data-discover="true">Get in touch</a></div></div></div><div class="col-start-1 col-end-2 flex items-center gap-x-3 lg:row-start-1 lg:h-4"><a aria-label="Go to homepage" class="col-start-1 col-end-2 hidden h-4 w-4 rounded-full text-primary lg:block" href="/" data-discover="true"><svg xmlns="http://www.w3.org/2000/svg" width="32" height="32" fill="none" viewBox="0 0 40 40"><path fill="currentColor" d="M0 20C0 8.954 8.954 0 20 0c8.121 0 15.112 4.84 18.245 11.794l-26.45 26.45a20 20 0 0 1-3.225-1.83L24.984 20H20L5.858 34.142A19.94 19.94 0 0 1 0 20M39.999 20.007 20.006 40c11.04-.004 19.99-8.953 19.993-19.993"></path></svg></a><nav aria-label="Main" data-orientation="horizontal" class="hidden items-center lg:flex"><ul class="flex flex-wrap gap-x-1 md:flex-nowrap"><li><div data-headlessui-state=""><button class="font-semibold text-primary hover:text-contrast focus-visible:ring-0 ui-open:text-orange" type="button" aria-expanded="false" data-headlessui-state="">Platform<span class="ml-sm inline-block ui-open:rotate-180">▾</span></button></div><span hidden="" style="position:fixed;top:1px;left:1px;width:1px;height:0;padding:0;margin:-1px;overflow:hidden;clip:rect(0, 0, 0, 0);white-space:nowrap;border-width:0;display:none"></span></li><li class="text-decoration" role="presentation">|</li><li><div data-headlessui-state=""><button class="font-semibold text-primary hover:text-contrast focus-visible:ring-0 ui-open:text-orange" type="button" aria-expanded="false" data-headlessui-state="">Resources<span class="ml-sm inline-block ui-open:rotate-180">▾</span></button></div><span hidden="" style="position:fixed;top:1px;left:1px;width:1px;height:0;padding:0;margin:-1px;overflow:hidden;clip:rect(0, 0, 0, 0);white-space:nowrap;border-width:0;display:none"></span></li><li class="text-decoration" role="presentation">|</li><li><a class="font-semibold text-primary hover:text-contrast" href="/docs">Documentation</a></li><li class="text-decoration" role="presentation">|</li><li><a class="font-semibold text-primary hover:text-contrast" href="/pricing" data-discover="true">Pricing</a></li><li class="text-decoration" role="presentation">|</li><li><a class="font-semibold text-primary hover:text-contrast" href="/migrate" data-discover="true">Migrate</a></li></ul></nav></div></div><details class="lg:hidden"><summary>Navigation</summary><nav class="dashed-box mt-1 p-3"><ul class="flex flex-wrap gap-x-1 md:flex-nowrap"><li><div data-headlessui-state=""><button class="font-semibold text-primary hover:text-contrast focus-visible:ring-0 ui-open:text-orange" type="button" aria-expanded="false" data-headlessui-state="">Platform<span class="ml-sm inline-block ui-open:rotate-180">▾</span></button></div><span hidden="" style="position:fixed;top:1px;left:1px;width:1px;height:0;padding:0;margin:-1px;overflow:hidden;clip:rect(0, 0, 0, 0);white-space:nowrap;border-width:0;display:none"></span></li><li class="text-decoration" role="presentation">|</li><li><div data-headlessui-state=""><button class="font-semibold text-primary hover:text-contrast focus-visible:ring-0 ui-open:text-orange" type="button" aria-expanded="false" data-headlessui-state="">Resources<span class="ml-sm inline-block ui-open:rotate-180">▾</span></button></div><span hidden="" style="position:fixed;top:1px;left:1px;width:1px;height:0;padding:0;margin:-1px;overflow:hidden;clip:rect(0, 0, 0, 0);white-space:nowrap;border-width:0;display:none"></span></li><li class="text-decoration" role="presentation">|</li><li><a class="font-semibold text-primary hover:text-contrast" href="/docs">Documentation</a></li><li class="text-decoration" role="presentation">|</li><li><a class="font-semibold text-primary hover:text-contrast" href="/pricing" data-discover="true">Pricing</a></li><li class="text-decoration" role="presentation">|</li><li><a class="font-semibold text-primary hover:text-contrast" href="/migrate" data-discover="true">Migrate</a></li></ul></nav></details></div></header><main class="container mb-6 flex max-w-7xl flex-1 flex-col px-3 sm:px-5 lg:px-12"><section class=""><p class="block"><a class="pr-sm text-primary hover:text-contrast" href="/blog" data-discover="true">Blog</a><span class="px-sm text-decoration">|</span><a class="px-sm text-blue hover:bg-blue-100 dark:hover:bg-blue-900" href="/blog/category/engineering" data-discover="true">Engineering</a></p><div class="flex lg:flex-row-reverse lg:gap-x-6"><div class="lg:sticky lg:top-2 lg:self-start"><button class="absolute right-0 bg-gray-100 px-sm md:block lg:hidden dark:bg-gray-800 -mt-9 hidden"><span class="inline">Table of contents «</span><span class="hidden">Close »</span></button><aside class="tree-nav w-full shrink-0 space-y-3 lg:w-36 hidden lg:block"><div><h4 class="text-secondary">Table of contents</h4><ul><li><a class="font-semibold text-primary hover:text-blue" href="/blog/the-real-fail-rate-of-ebs#defining-failure" data-discover="true">Defining Failure</a></li><li><a class="font-semibold text-primary hover:text-blue" href="/blog/the-real-fail-rate-of-ebs#the-true-rate-of-failure" data-discover="true">The True Rate of Failure</a></li><li><a class="font-semibold text-primary hover:text-blue" href="/blog/the-real-fail-rate-of-ebs#handling-failure" data-discover="true">Handling Failure</a></li></ul><div class="mb-3 mt-6 border bg-blue-50 p-3 font-semibold text-contrast dark:bg-blue-900"><p>PlanetScale, the fastest cloud Postgres, from $5/month.</p><p><a href="https://app.planetscale.com/new">Start now</a></p></div><p>Get the <a href="/blog/feed.atom">RSS feed</a></p></div></aside></div><article class="min-w-0 flex-grow"><h1>The Real Failure Rate of EBS</h1><p class="text-secondary"><a class="text-contrast no-underline" href="/blog/author/nick" data-discover="true">Nick Van Wiggeren</a> <!-- -->[<a class="no-underline hover:bg-blue-100 dark:hover:bg-blue-900" href="https://x.com/NickVanWig" rel="noopener noreferrer" target="_blank" title="@NickVanWig on X">@<!-- -->NickVanWig</a>]<!-- --> |<!-- --> <time dateTime="2025-03-18">March 18, 2025</time></p><div class="blog-post-body"><p>PlanetScale has deployed millions of Amazon Elastic Block Store (EBS) volumes across the world. We create and destroy tens of thousands of them every day as we stand up databases for customers, take backups, and test our systems end-to-end. Through this experience, we have an unique viewpoint into the failure rate and mechanisms of EBS, and have spent a lot of time working on how to mitigate them.</p><p>In complex systems, failure isn’t a binary outcome. Cloud native systems are built without single paths of failure, but partial failure can still result in degraded performance, loss of user-facing availability, and undefined behavior. Often, minor failure in one part of the stack appears as a full failure in others.</p><p>For example, if a single instance inside of a multi-node distributed caching system runs out of networking resources, the downstream application will interpret error cases as cache misses to avoid failing the request. This will overwhelm the database when the application floods it with queries to fetch data as though it was missing. In this, a partial failure at one level results in a full failure of the database tier, causing downtime.</p><h2 id="defining-failure"><a href="#defining-failure">Defining Failure</a></h2><p>While full failure and data loss is very rare with EBS, “slow” is often as bad as “failed”, and that happens much much more often.</p><p>Here’s what “slow” looks like, from the AWS Console:</p><p><button type="button" aria-haspopup="dialog" aria-expanded="false" aria-label="Enlarge image: EBS Volume Degrading" class="focus-visible-ring group relative block w-fit max-w-[min(100%,800px)] cursor-zoom-in text-left [&_picture]:contents"><picture class="block"><source media="(prefers-color-scheme: light), (prefers-color-scheme: no-preference)" srcSet="https://planetscale-images.imgix.net/assets/ebs-volume-degrading-D7MtBxSa.png?auto=compress%2Cformat"/><img alt="EBS Volume Degrading" src="https://planetscale-images.imgix.net/assets/ebs-volume-degrading-D7MtBxSa.png?auto=compress%2Cformat" width="1600" height="781" loading="lazy" class="w-auto max-w-full"/></picture><span aria-hidden="true" class="pointer-events-none absolute right-1 top-1 z-10 flex h-5 w-5 items-center justify-center border border-white/25 bg-black/70 text-white backdrop-blur-sm transition-colors transition-opacity group-hover:bg-black/90 group-hover:opacity-100 group-focus-visible:opacity-100 motion-reduce:transition-none [@media(hover:hover)_and_(pointer:fine)]:opacity-0"><svg width="14" height="14" viewBox="0 0 16 16" fill="none" stroke="currentColor" stroke-width="1.5" stroke-linecap="round" stroke-linejoin="round"><path d="M10 2h4v4M6 14H2v-4M14 2l-4.5 4.5M2 14l4.5-4.5"></path></svg></span></button><span hidden="" style="position:fixed;top:1px;left:1px;width:1px;height:0;padding:0;margin:-1px;overflow:hidden;clip:rect(0, 0, 0, 0);white-space:nowrap;border-width:0;display:none"></span></p><p>This volume has been operating steadily for at least 10 hours. AWS has reported it at 67% idle, with write latency measuring at single-digit ms/operation. Well within expectations. Suddenly, at around 16:00, write latency spikes to 200ms-500ms/operation, idle time races to zero, and the volume is effectively blocked from reading and writing data.</p><p>To the application running on top of this database: this is a failure. To the user, this is a 500 response on a webpage after a 10 second wait. To you, this is an incident. At PlanetScale, we consider this full failure because our customers do.</p><p>The <a href="https://docs.aws.amazon.com/ebs/latest/userguide/general-purpose.html#gp3-ebs-volume-type">EBS documentation</a> is useful in helping us understand what promises AWS’ gp3 is able to make:</p><blockquote><p>When attached to an EBS–optimized instance, General Purpose SSD (gp2 and gp3) volumes are designed to deliver at least 90 percent of their provisioned IOPS performance 99 percent of the time in a given year</p></blockquote><p>This means a volume is expected to experience <em>under 90%</em> of its provisioned performance 1% of the time. That’s 14 minutes of every day or 86 hours out of the year of <em>potential</em> impact. This rate of degradation <em>far</em> exceeds that of a single disk drive or SSD. This is the cost of separating storage and compute and the sheer complexity of the software and networking components between the client and the backing disks for the volume.</p><p>In our experience, the documentation is accurate: sometimes volumes pass in and out of their provisioned performance in small time windows:</p><p><button type="button" aria-haspopup="dialog" aria-expanded="false" aria-label="Enlarge image: EBS volume having a short blip" class="focus-visible-ring group relative block w-fit max-w-[min(100%,800px)] cursor-zoom-in text-left [&_picture]:contents"><picture class="block"><source media="(prefers-color-scheme: light), (prefers-color-scheme: no-preference)" srcSet="https://planetscale-images.imgix.net/assets/ebs-volume-short-blip-B86N4wvl.png?auto=compress%2Cformat"/><img alt="EBS volume having a short blip" src="https://planetscale-images.imgix.net/assets/ebs-volume-short-blip-B86N4wvl.png?auto=compress%2Cformat" width="1542" height="694" loading="lazy" class="w-auto max-w-full"/></picture><span aria-hidden="true" class="pointer-events-none absolute right-1 top-1 z-10 flex h-5 w-5 items-center justify-center border border-white/25 bg-black/70 text-white backdrop-blur-sm transition-colors transition-opacity group-hover:bg-black/90 group-hover:opacity-100 group-focus-visible:opacity-100 motion-reduce:transition-none [@media(hover:hover)_and_(pointer:fine)]:opacity-0"><svg width="14" height="14" viewBox="0 0 16 16" fill="none" stroke="currentColor" stroke-width="1.5" stroke-linecap="round" stroke-linejoin="round"><path d="M10 2h4v4M6 14H2v-4M14 2l-4.5 4.5M2 14l4.5-4.5"></path></svg></span></button><span hidden="" style="position:fixed;top:1px;left:1px;width:1px;height:0;padding:0;margin:-1px;overflow:hidden;clip:rect(0, 0, 0, 0);white-space:nowrap;border-width:0;display:none"></span></p><p>However, these short windows are enough to have impact on real-time workloads:</p><p><button type="button" aria-haspopup="dialog" aria-expanded="false" aria-label="Enlarge image: Database blip due to EBS" class="focus-visible-ring group relative block w-fit max-w-[min(100%,800px)] cursor-zoom-in text-left [&_picture]:contents"><picture class="block"><source media="(prefers-color-scheme: light), (prefers-color-scheme: no-preference)" srcSet="https://planetscale-images.imgix.net/assets/database-blip-ebs-CFXXjXsh.png?auto=compress%2Cformat"/><img alt="Database blip due to EBS" src="https://planetscale-images.imgix.net/assets/database-blip-ebs-CFXXjXsh.png?auto=compress%2Cformat" width="1547" height="623" loading="lazy" class="w-auto max-w-full"/></picture><span aria-hidden="true" class="pointer-events-none absolute right-1 top-1 z-10 flex h-5 w-5 items-center justify-center border border-white/25 bg-black/70 text-white backdrop-blur-sm transition-colors transition-opacity group-hover:bg-black/90 group-hover:opacity-100 group-focus-visible:opacity-100 motion-reduce:transition-none [@media(hover:hover)_and_(pointer:fine)]:opacity-0"><svg width="14" height="14" viewBox="0 0 16 16" fill="none" stroke="currentColor" stroke-width="1.5" stroke-linecap="round" stroke-linejoin="round"><path d="M10 2h4v4M6 14H2v-4M14 2l-4.5 4.5M2 14l4.5-4.5"></path></svg></span></button><span hidden="" style="position:fixed;top:1px;left:1px;width:1px;height:0;padding:0;margin:-1px;overflow:hidden;clip:rect(0, 0, 0, 0);white-space:nowrap;border-width:0;display:none"></span></p><p>Production systems are not built to handle this level of sudden variance. When there are no guarantees, even overprovisioning doesn’t solve the problem. If this were a once-in-a-million chance, it would be different, but as we’ll discuss below, that is far from the case.</p><h2 id="the-true-rate-of-failure"><a href="#the-true-rate-of-failure">The True Rate of Failure</a></h2><p>At PlanetScale, we see failures like this on a daily basis - the rate of failure is frequent enough that we’ve built systems that monitor EBS volumes directly to minimize impact.</p><p>This is not a secret, it's from the documentation. AWS doesn’t describe how failure is distributed for gp3 volumes, but in our experience it tends to last 1-10 minutes at a time. This is likely the time needed for a failover in a network or compute component.</p><p>Let's assume the following: Each degradation event is random, meaning the level of reduced performance is somewhere between 1% and 89% of provisioned, and your application is designed to withstand losing 50% of its expected throughput before erroring. If each individual failure event lasts 10 minutes, every volume would experience about 43 events per month, with at least 21 of them causing downtime!</p><p>In a large database composed of many shards, this failure compounds. Assume a 256 shard database where each shard has one primary and two replicas: a total of 768 gp3 EBS volumes provisioned. If we take the 50% threshold from above, there is a 99.65% chance you have at least one node experiencing a production-impacting event at any given time.</p><p>Even if you use io2, which AWS sells at 4x to 10x the price, you’d still be expected to be in a failure condition roughly one third of the time in any given year on just that one database!</p><p>To make matters worse, we also see these frequently as correlated failure inside of a single zone, even using <code>io2</code> volumes:</p><p><button type="button" aria-haspopup="dialog" aria-expanded="false" aria-label="Enlarge image: Correlated EBS Failure" class="focus-visible-ring group relative block w-fit max-w-[min(100%,800px)] cursor-zoom-in text-left [&_picture]:contents"><picture class="block"><source media="(prefers-color-scheme: light), (prefers-color-scheme: no-preference)" srcSet="https://planetscale-images.imgix.net/assets/correlated-ebs-failure-BcoLdHtC.png?auto=compress%2Cformat"/><img alt="Correlated EBS Failure" src="https://planetscale-images.imgix.net/assets/correlated-ebs-failure-BcoLdHtC.png?auto=compress%2Cformat" width="1600" height="580" loading="lazy" class="w-auto max-w-full"/></picture><span aria-hidden="true" class="pointer-events-none absolute right-1 top-1 z-10 flex h-5 w-5 items-center justify-center border border-white/25 bg-black/70 text-white backdrop-blur-sm transition-colors transition-opacity group-hover:bg-black/90 group-hover:opacity-100 group-focus-visible:opacity-100 motion-reduce:transition-none [@media(hover:hover)_and_(pointer:fine)]:opacity-0"><svg width="14" height="14" viewBox="0 0 16 16" fill="none" stroke="currentColor" stroke-width="1.5" stroke-linecap="round" stroke-linejoin="round"><path d="M10 2h4v4M6 14H2v-4M14 2l-4.5 4.5M2 14l4.5-4.5"></path></svg></span></button><span hidden="" style="position:fixed;top:1px;left:1px;width:1px;height:0;padding:0;margin:-1px;overflow:hidden;clip:rect(0, 0, 0, 0);white-space:nowrap;border-width:0;display:none"></span></p><p>With enough volumes, the rate of experiencing EBS failure is 100%: our automated mitigations are consistently recycling underperforming EBS volumes to reduce customer-impact, and we expect to see multiple events on a daily basis.</p><p>That’s the true rate of failure of EBS: it’s constant, variable, and all by design. Because there are no performance guarantees when volumes are not operating to their specifications, it is extremely difficult to plan around for workloads that require consistent performance. You can pay for additional nines, but with enough drives over a long enough timeframe, failure is guaranteed.</p><h2 id="handling-failure"><a href="#handling-failure">Handling Failure</a></h2><p>At PlanetScale, our mitigations have clamped down on the expected maximum time for an impact window. We monitor metrics such as read/write latency and idle % closely, and we've even developed basic tests like making sure we can write to a file. This allows us to respond quickly to performance issues, and ensures that an EBS volume isn’t ‘stuck’.</p><p>When we detect that an EBS volume is in a degraded state using these heuristics, we can perform a zero-downtime reparent in seconds to another node in the cluster, and automatically bring up a replacement volume. This doesn’t reduce the impact to zero, as it’s impossible to detect this failure before it happens, but it does ensure the majority of the cases don’t require a human to remediate and are over before users notice.</p><p>This is why we built <a href="/metal">PlanetScale Metal</a>. With a shared-nothing architecture that uses local storage instead of network-attached storage like EBS, the rest of the shards and nodes in a database are able to continue to operate without problem.</p></div></article></div></section></main><footer class="mb-6 mt-10 px-3 sm:px-5 container max-w-7xl"><nav class="grid grid-cols-1 text-left sm:grid-cols-2 lg:grid-cols-5 lg:mx-7"><div class="dashed-box dashed-box-x-t sm:dashed-box-l-t lg:dashed-box-y-l p-3"><h2 class="font-semibold">Company</h2><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/about" data-discover="true">About</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/brand" data-discover="true">Brand</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/blog" data-discover="true">Blog</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/changelog" data-discover="true">Changelog</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/careers" data-discover="true">Careers</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/events" data-discover="true">Events</a></div><div class="dashed-box dashed-box-x-t lg:dashed-box-y-l p-3"><h2 class="font-semibold">Product</h2><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/case-studies" data-discover="true">Case studies</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/enterprise" data-discover="true">Enterprise</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/pricing" data-discover="true">Pricing</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/benchmarks" data-discover="true">Benchmarks</a></div><div class="dashed-box dashed-box-x-t sm:dashed-box-l-t lg:dashed-box-y-l p-3"><h2 class="font-semibold">Resources</h2><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/docs">Documentation</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/migrate" data-discover="true">Migrate</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="https://support.planetscale.com/hc/en-us" rel="nofollow noopener noreferrer" target="_blank">Support</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="https://planetscalestatus.com" rel="nofollow noopener noreferrer" target="_blank">Status</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="https://trust.planetscale.com" rel="nofollow noopener noreferrer" target="_blank">Trust Center</a></div><div class="dashed-box dashed-box-x-t lg:dashed-box-y-l p-3"><h2 class="font-semibold">Courses</h2><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/learn/courses/mysql-for-developers" data-discover="true">MySQL for Developers</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/learn/courses/database-scaling" data-discover="true">Database Scaling</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/learn/courses/vitess" data-discover="true">Learn Vitess</a></div><div class="dashed-box p-3 sm:col-span-2 lg:col-span-1"><h2 class="font-semibold text-primary hover:text-contrast">Open source</h2><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/vitess" data-discover="true">Vitess</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="https://vitess.io/slack" rel="nofollow noopener noreferrer" target="_blank">Vitess community</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="https://github.com/planetscale" rel="me nofollow noopener noreferrer" target="_blank">GitHub</a></div></nav><div class="dashed-box dashed-box-x-b p-3 lg:mx-7"><p class="mb-3 md:mb-0"><a class="text-primary" rel="nofollow" href="/legal/privacy" data-discover="true">Privacy</a><span class="text-decoration" role="presentation"> <!-- -->|<!-- --> </span><a class="text-primary" rel="nofollow" href="/legal/siteterms" data-discover="true">Terms</a><span class="text-decoration" role="presentation"> <!-- -->|<!-- --> </span><a class="text-primary" rel="nofollow" href="/legal/cookies" data-discover="true">Cookies</a><span class="text-decoration" role="presentation"> <!-- -->|<!-- --> </span><a class="text-primary" rel="nofollow" href="/legal/patents" data-discover="true">Patents</a><span class="text-decoration" role="presentation"> <!-- -->|<!-- --> </span><a class="text-primary" rel="nofollow" href="/legal/privacy#privacy-rights-and-choices" data-discover="true">Do Not Share My Personal Information</a></p><p class="text-secondary">© <!-- -->2026<!-- --> PlanetScale, Inc. All rights reserved.</p></div><p class="mb-0 mt-3 break-normal lg:mx-7"><a class="text-primary" href="https://github.com/planetscale" rel="me nofollow noopener noreferrer" target="_blank">GitHub</a><span class="text-decoration" role="presentation"> <!-- -->|<!-- --> </span><a aria-label="X (formerly Twitter)" class="text-primary" href="https://twitter.com/planetscale" rel="me nofollow noopener noreferrer" target="_blank">X</a><span class="text-decoration" role="presentation"> <!-- -->|<!-- --> </span><a aria-label="LinkedIn" class="text-primary" href="https://www.linkedin.com/company/planetscale" target="_blank" rel="noreferrer">LinkedIn</a><span class="text-decoration" role="presentation"> <!-- -->|<!-- --> </span><a class="text-primary" href="https://www.youtube.com/planetscale" rel="me nofollow noopener noreferrer" target="_blank">YouTube</a><span class="text-decoration" role="presentation"> <!-- -->|<!-- --> </span><a aria-label="Discord" class="text-primary" href="https://pscale.link/community" rel="nofollow noopener noreferrer" target="_blank">Discord</a><span class="text-decoration" role="presentation"> <!-- -->|<!-- --> </span><a class="text-primary" href="https://www.facebook.com/planetscaledata" rel="me nofollow noopener noreferrer" target="_blank">Facebook</a></p></footer><script nonce="EZdXajdlI2DRCjbnoKJCwfF5KNgWOgdluJsujYSRAY4=">((storageKey2, restoreKey) => {
|
||
if (!window.history.state || !window.history.state.key) {
|
||
let key2 = Math.random().toString(32).slice(2);
|
||
window.history.replaceState({ key: key2 }, "");
|
||
}
|
||
try {
|
||
let storedY = JSON.parse(sessionStorage.getItem(storageKey2) || "{}")[restoreKey || window.history.state.key];
|
||
if (typeof storedY === "number") window.scrollTo(0, storedY);
|
||
} catch (error2) {
|
||
console.error(error2);
|
||
sessionStorage.removeItem(storageKey2);
|
||
}
|
||
})("react-router-scroll-positions", null)</script><script nonce="EZdXajdlI2DRCjbnoKJCwfF5KNgWOgdluJsujYSRAY4=">window.__reactRouterContext = {"basename":"/","future":{"unstable_enableNodeReadableStream":false,"unstable_optimizeDeps":true},"routeDiscovery":{"mode":"lazy","manifestPath":"/__manifest"},"ssr":true,"isSpaMode":false};window.__reactRouterContext.stream = new ReadableStream({start(controller){window.__reactRouterContext.streamController = controller;}}).pipeThrough(new TextEncoderStream());</script><script nonce="EZdXajdlI2DRCjbnoKJCwfF5KNgWOgdluJsujYSRAY4=" type="module" async="">;
|
||
import * as route0 from "/assets/root-DbOv4-98.js";
|
||
import * as route1 from "/assets/blog-pXH7ptHJ.js";
|
||
import * as route2 from "/assets/blog._slug-Ch_92qsH.js";
|
||
window.__reactRouterManifest = {
|
||
"entry": {
|
||
"module": "/assets/entry.client-3vubyXrk.js",
|
||
"imports": [
|
||
"/assets/jsx-runtime-DwfQwkRq.js",
|
||
"/assets/components-_bNmAApg.js",
|
||
"/assets/_.well-known_.mcp.server-card_.json_-Sx7XeH3e.js",
|
||
"/assets/index-mKTXLmHu.js",
|
||
"/assets/errorBoundaries-DhW4jVYt.js"
|
||
],
|
||
"css": []
|
||
},
|
||
"routes": {
|
||
"root": {
|
||
"id": "root",
|
||
"path": "",
|
||
"hasAction": false,
|
||
"hasLoader": true,
|
||
"hasClientAction": false,
|
||
"hasClientLoader": false,
|
||
"hasClientMiddleware": false,
|
||
"hasDefaultExport": true,
|
||
"hasErrorBoundary": true,
|
||
"module": "/assets/root-DbOv4-98.js",
|
||
"imports": [
|
||
"/assets/jsx-runtime-DwfQwkRq.js",
|
||
"/assets/components-_bNmAApg.js",
|
||
"/assets/_.well-known_.mcp.server-card_.json_-Sx7XeH3e.js",
|
||
"/assets/index-mKTXLmHu.js",
|
||
"/assets/errorBoundaries-DhW4jVYt.js",
|
||
"/assets/lib-Dg89tQ22.js",
|
||
"/assets/analytics.client-DM6E8o1h.js",
|
||
"/assets/SiteHeader-C2U5gvDH.js",
|
||
"/assets/current-9yDxj94E.js",
|
||
"/assets/clsx-eT0YPcGk.js",
|
||
"/assets/bugs-38ilEoW0.js",
|
||
"/assets/keyboard-D-uXZORL.js",
|
||
"/assets/use-tab-direction-dKm-S3Ck.js"
|
||
],
|
||
"css": []
|
||
},
|
||
"routes/blog": {
|
||
"id": "routes/blog",
|
||
"parentId": "root",
|
||
"path": "blog",
|
||
"hasAction": false,
|
||
"hasLoader": false,
|
||
"hasClientAction": false,
|
||
"hasClientLoader": false,
|
||
"hasClientMiddleware": false,
|
||
"hasDefaultExport": false,
|
||
"hasErrorBoundary": false,
|
||
"module": "/assets/blog-pXH7ptHJ.js",
|
||
"imports": [
|
||
"/assets/_.well-known_.mcp.server-card_.json_-Sx7XeH3e.js"
|
||
],
|
||
"css": []
|
||
},
|
||
"routes/blog.$slug": {
|
||
"id": "routes/blog.$slug",
|
||
"parentId": "routes/blog",
|
||
"path": ":slug",
|
||
"hasAction": false,
|
||
"hasLoader": true,
|
||
"hasClientAction": false,
|
||
"hasClientLoader": false,
|
||
"hasClientMiddleware": false,
|
||
"hasDefaultExport": true,
|
||
"hasErrorBoundary": false,
|
||
"module": "/assets/blog._slug-Ch_92qsH.js",
|
||
"imports": [
|
||
"/assets/components-_bNmAApg.js",
|
||
"/assets/lib-Dg89tQ22.js",
|
||
"/assets/jsx-runtime-DwfQwkRq.js",
|
||
"/assets/ContentImage-Dh6VEOUl.js",
|
||
"/assets/clsx-eT0YPcGk.js",
|
||
"/assets/BlogCategoryLink-DmQyn0gp.js",
|
||
"/assets/_.well-known_.mcp.server-card_.json_-Sx7XeH3e.js",
|
||
"/assets/Details-BSB_b6hI.js",
|
||
"/assets/Skittle-CDFOPRjH.js",
|
||
"/assets/SiteFooter-B2Gq9u2j.js",
|
||
"/assets/SiteHeader-C2U5gvDH.js",
|
||
"/assets/Vimeo-00PQJDli.js",
|
||
"/assets/YouTube-CMfaljVr.js",
|
||
"/assets/date-CJTFH3uT.js",
|
||
"/assets/errorBoundaries-DhW4jVYt.js",
|
||
"/assets/keyboard-D-uXZORL.js",
|
||
"/assets/use-tab-direction-dKm-S3Ck.js",
|
||
"/assets/index-mKTXLmHu.js",
|
||
"/assets/use-inert-others-BMJ6-xOX.js",
|
||
"/assets/description-Cf6FZmDe.js",
|
||
"/assets/use-is-mounted-uQsUZyP9.js",
|
||
"/assets/types-DvonrUFF.js",
|
||
"/assets/current-9yDxj94E.js",
|
||
"/assets/analytics.client-DM6E8o1h.js",
|
||
"/assets/bugs-38ilEoW0.js"
|
||
],
|
||
"css": []
|
||
},
|
||
"routes/_index": {
|
||
"id": "routes/_index",
|
||
"parentId": "root",
|
||
"index": true,
|
||
"hasAction": false,
|
||
"hasLoader": true,
|
||
"hasClientAction": false,
|
||
"hasClientLoader": false,
|
||
"hasClientMiddleware": false,
|
||
"hasDefaultExport": true,
|
||
"hasErrorBoundary": false,
|
||
"module": "/assets/_index-BfA6EnlR.js",
|
||
"imports": [
|
||
"/assets/components-_bNmAApg.js",
|
||
"/assets/lib-Dg89tQ22.js",
|
||
"/assets/jsx-runtime-DwfQwkRq.js",
|
||
"/assets/Logo-Gm9TLYAs.js",
|
||
"/assets/SiteFooter-B2Gq9u2j.js",
|
||
"/assets/SiteHeader-C2U5gvDH.js",
|
||
"/assets/_.well-known_.mcp.server-card_.json_-Sx7XeH3e.js",
|
||
"/assets/bugs-38ilEoW0.js",
|
||
"/assets/keyboard-D-uXZORL.js",
|
||
"/assets/use-is-mounted-uQsUZyP9.js",
|
||
"/assets/use-tab-direction-dKm-S3Ck.js",
|
||
"/assets/errorBoundaries-DhW4jVYt.js",
|
||
"/assets/clsx-eT0YPcGk.js",
|
||
"/assets/current-9yDxj94E.js",
|
||
"/assets/analytics.client-DM6E8o1h.js",
|
||
"/assets/index-mKTXLmHu.js"
|
||
],
|
||
"css": []
|
||
},
|
||
"routes/blog._index": {
|
||
"id": "routes/blog._index",
|
||
"parentId": "routes/blog",
|
||
"index": true,
|
||
"hasAction": false,
|
||
"hasLoader": true,
|
||
"hasClientAction": false,
|
||
"hasClientLoader": false,
|
||
"hasClientMiddleware": false,
|
||
"hasDefaultExport": true,
|
||
"hasErrorBoundary": false,
|
||
"module": "/assets/blog._index-DcvTTuDd.js",
|
||
"imports": [
|
||
"/assets/components-_bNmAApg.js",
|
||
"/assets/jsx-runtime-DwfQwkRq.js",
|
||
"/assets/social-Cd2AtOZM.js",
|
||
"/assets/BlogCategoryLink-DmQyn0gp.js",
|
||
"/assets/BlogPostLink-DC1SPKBJ.js",
|
||
"/assets/BlogCategoryNav-CB7TJ3IB.js",
|
||
"/assets/Paginator-xlPA_JNt.js",
|
||
"/assets/SiteFooter-B2Gq9u2j.js",
|
||
"/assets/SiteHeader-C2U5gvDH.js",
|
||
"/assets/date-CJTFH3uT.js",
|
||
"/assets/_.well-known_.mcp.server-card_.json_-Sx7XeH3e.js",
|
||
"/assets/lib-Dg89tQ22.js",
|
||
"/assets/errorBoundaries-DhW4jVYt.js",
|
||
"/assets/clsx-eT0YPcGk.js",
|
||
"/assets/types-DvonrUFF.js",
|
||
"/assets/enumerator-2YLGh-nT.js",
|
||
"/assets/current-9yDxj94E.js",
|
||
"/assets/analytics.client-DM6E8o1h.js",
|
||
"/assets/bugs-38ilEoW0.js",
|
||
"/assets/keyboard-D-uXZORL.js",
|
||
"/assets/use-tab-direction-dKm-S3Ck.js",
|
||
"/assets/index-mKTXLmHu.js"
|
||
],
|
||
"css": []
|
||
}
|
||
},
|
||
"url": "/assets/manifest-e17deb94.js",
|
||
"version": "e17deb94"
|
||
};
|
||
window.__reactRouterRouteModules = {"root":route0,"routes/blog":route1,"routes/blog.$slug":route2};
|
||
|
||
import("/assets/entry.client-3vubyXrk.js");</script><script type="application/ld+json" nonce="EZdXajdlI2DRCjbnoKJCwfF5KNgWOgdluJsujYSRAY4=">{"@context":"https://schema.org","@type":"Organization","name":"PlanetScale, Inc.","url":"https://planetscale.com","sameAs":["https://twitter.com/PlanetScale","https://www.facebook.com/planetscaledata/","https://www.instagram.com/planetscale/"],"address":{"@type":"PostalAddress","streetAddress":"WeWork c/o PlanetScale, 535 Mission Street, 14th Floor","addressLocality":"San Francisco","addressRegion":"CA","postalCode":"94105","addressCountry":"US"}}</script><!--$--><script nonce="EZdXajdlI2DRCjbnoKJCwfF5KNgWOgdluJsujYSRAY4=">window.__reactRouterContext.streamController.enqueue("[{\"_1\":2,\"_3\":-5,\"_4\":-5},\"loaderData\",{\"_5\":6,\"_7\":8},\"actionData\",\"errors\",\"root\",{\"_326\":327},\"routes/blog.$slug\",{\"_9\":10,\"_5\":11},\"blog\",{\"_12\":13,\"_14\":15,\"_16\":-7,\"_17\":18,\"_19\":20,\"_21\":22,\"_23\":24,\"_25\":26,\"_27\":28,\"_29\":30,\"_31\":32},\"https://planetscale.com\",\"body\",[61,62,63,64,65,66,67,68,69,70,71,72,73,74,75,76,77,78,79,80,81,82,83,84,85,86,87,88,89,90,91],\"body_text\",\"PlanetScale has deployed millions of Amazon Elastic Block Store (EBS) volumes across the world. We create and destroy tens of thousands of them every day as we stand up databases for customers, take backups, and test our systems end-to-end. Through this experience, we have an unique viewpoint into the failure rate and mechanisms of EBS, and have spent a lot of time working on how to mitigate them.\\nIn complex systems, failure isn’t a binary outcome. Cloud native systems are built without single paths of failure, but partial failure can still result in degraded performance, loss of user-facing availability, and undefined behavior. Often, minor failure in one part of the stack appears as a full failure in others.\\nFor example, if a single instance inside of a multi-node distributed caching system runs out of networking resources, the downstream application will interpret error cases as cache misses to avoid failing the request. This will overwhelm the database when the application floods it with queries to fetch data as though it was missing. In this, a partial failure at one level results in a full failure of the database tier, causing downtime.\\nDefining Failure\\nWhile full failure and data loss is very rare with EBS, “slow” is often as bad as “failed”, and that happens much much more often.\\nHere’s what “slow” looks like, from the AWS Console:\\n\\nThis volume has been operating steadily for at least 10 hours. AWS has reported it at 67% idle, with write latency measuring at single-digit ms/operation. Well within expectations. Suddenly, at around 16:00, write latency spikes to 200ms-500ms/operation, idle time races to zero, and the volume is effectively blocked from reading and writing data.\\nTo the application running on top of this database: this is a failure. To the user, this is a 500 response on a webpage after a 10 second wait. To you, this is an incident. At PlanetScale, we consider this full failure because our customers do.\\nThe EBS documentation is useful in helping us understand what promises AWS’ gp3 is able to make:\\nWhen attached to an EBS–optimized instance, General Purpose SSD (gp2 and gp3) volumes are designed to deliver at least 90 percent of their provisioned IOPS performance 99 percent of the time in a given year\\nThis means a volume is expected to experience under 90% of its provisioned performance 1% of the time. That’s 14 minutes of every day or 86 hours out of the year of potential impact. This rate of degradation far exceeds that of a single disk drive or SSD. This is the cost of separating storage and compute and the sheer complexity of the software and networking components between the client and the backing disks for the volume.\\nIn our experience, the documentation is accurate: sometimes volumes pass in and out of their provisioned performance in small time windows:\\n\\nHowever, these short windows are enough to have impact on real-time workloads:\\n\\nProduction systems are not built to handle this level of sudden variance. When there are no guarantees, even overprovisioning doesn’t solve the problem. If this were a once-in-a-million chance, it would be different, but as we’ll discuss below, that is far from the case.\\nThe True Rate of Failure\\nAt PlanetScale, we see failures like this on a daily basis - the rate of failure is frequent enough that we’ve built systems that monitor EBS volumes directly to minimize impact.\\nThis is not a secret, it's from the documentation. AWS doesn’t describe how failure is distributed for gp3 volumes, but in our experience it tends to last 1-10 minutes at a time. This is likely the time needed for a failover in a network or compute component.\\nLet's assume the following: Each degradation event is random, meaning the level of reduced performance is somewhere between 1% and 89% of provisioned, and your application is designed to withstand losing 50% of its expected throughput before erroring. If each individual failure event lasts 10 minutes, every volume would experience about 43 events per month, with at least 21 of them causing downtime!\\nIn a large database composed of many shards, this failure compounds. Assume a 256 shard database where each shard has one primary and two replicas: a total of 768 gp3 EBS volumes provisioned. If we take the 50% threshold from above, there is a 99.65% chance you have at least one node experiencing a production-impacting event at any given time.\\nEven if you use io2, which AWS sells at 4x to 10x the price, you’d still be expected to be in a failure condition roughly one third of the time in any given year on just that one database!\\nTo make matters worse, we also see these frequently as correlated failure inside of a single zone, even using io2 volumes:\\n\\nWith enough volumes, the rate of experiencing EBS failure is 100%: our automated mitigations are consistently recycling underperforming EBS volumes to reduce customer-impact, and we expect to see multiple events on a daily basis.\\nThat’s the true rate of failure of EBS: it’s constant, variable, and all by design. Because there are no performance guarantees when volumes are not operating to their specifications, it is extremely difficult to plan around for workloads that require consistent performance. You can pay for additional nines, but with enough drives over a long enough timeframe, failure is guaranteed.\\nHandling Failure\\nAt PlanetScale, our mitigations have clamped down on the expected maximum time for an impact window. We monitor metrics such as read/write latency and idle % closely, and we've even developed basic tests like making sure we can write to a file. This allows us to respond quickly to performance issues, and ensures that an EBS volume isn’t ‘stuck’.\\nWhen we detect that an EBS volume is in a degraded state using these heuristics, we can perform a zero-downtime reparent in seconds to another node in the cluster, and automatically bring up a replacement volume. This doesn’t reduce the impact to zero, as it’s impossible to detect this failure before it happens, but it does ensure the majority of the cases don’t require a human to remediate and are over before users notice.\\nThis is why we built PlanetScale Metal. With a shared-nothing architecture that uses local storage instead of network-attached storage like EBS, the rest of the shards and nodes in a database are able to continue to operate without problem.\",\"aside\",\"toc\",[45,46,47],\"title\",\"The Real Failure Rate of EBS\",\"authors\",[39],\"categories\",[38],\"excerpt\",\"Our experience running AWS EBS at scale for critical workloads\",\"createdAt\",\"2025-03-18\",\"slug\",\"the-real-fail-rate-of-ebs\",\"meta\",{\"_33\":34,\"_35\":26,\"_36\":37,\"_19\":20},\"canonical\",\"https://planetscale.com/blog/the-real-fail-rate-of-ebs\",\"description\",\"image\",\"/assets/the-real-fail-rate-of-ebs-social-Cl1a18Am.jpg\",\"engineering\",{\"_29\":40,\"_41\":42,\"_43\":44},\"nick\",\"name\",\"Nick Van Wiggeren\",\"x\",\"NickVanWig\",{\"_48\":58,\"_50\":59,\"_52\":53,\"_19\":60},{\"_48\":55,\"_50\":56,\"_52\":53,\"_19\":57},{\"_48\":49,\"_50\":51,\"_52\":53,\"_19\":54},\"children\",[],\"id\",\"handling-failure\",\"level\",2,\"Handling Failure\",[],\"the-true-rate-of-failure\",\"The True Rate of Failure\",[],\"defining-failure\",\"Defining Failure\",[\"SingleFetchClassInstance\",322],[\"SingleFetchClassInstance\",318],[\"SingleFetchClassInstance\",314],[\"SingleFetchClassInstance\",306],[\"SingleFetchClassInstance\",302],[\"SingleFetchClassInstance\",298],[\"SingleFetchClassInstance\",286],[\"SingleFetchClassInstance\",282],[\"SingleFetchClassInstance\",278],[\"SingleFetchClassInstance\",267],[\"SingleFetchClassInstance\",258],[\"SingleFetchClassInstance\",235],[\"SingleFetchClassInstance\",231],[\"SingleFetchClassInstance\",218],[\"SingleFetchClassInstance\",214],[\"SingleFetchClassInstance\",201],[\"SingleFetchClassInstance\",197],[\"SingleFetchClassInstance\",189],[\"SingleFetchClassInstance\",185],[\"SingleFetchClassInstance\",181],[\"SingleFetchClassInstance\",177],[\"SingleFetchClassInstance\",173],[\"SingleFetchClassInstance\",169],[\"SingleFetchClassInstance\",158],[\"SingleFetchClassInstance\",134],[\"SingleFetchClassInstance\",130],[\"SingleFetchClassInstance\",126],[\"SingleFetchClassInstance\",117],[\"SingleFetchClassInstance\",113],[\"SingleFetchClassInstance\",109],[\"SingleFetchClassInstance\",92],{\"_93\":94,\"_41\":95,\"_96\":97,\"_48\":98},\"$$mdtype\",\"Tag\",\"p\",\"attributes\",{},[99,100,101],\"This is why we built \",[\"SingleFetchClassInstance\",102],\". With a shared-nothing architecture that uses local storage instead of network-attached storage like EBS, the rest of the shards and nodes in a database are able to continue to operate without problem.\",{\"_93\":94,\"_41\":103,\"_96\":104,\"_48\":105},\"a\",{\"_107\":108},[106],\"PlanetScale Metal\",\"href\",\"/metal\",{\"_93\":94,\"_41\":95,\"_96\":110,\"_48\":111},{},[112],\"When we detect that an EBS volume is in a degraded state using these heuristics, we can perform a zero-downtime reparent in seconds to another node in the cluster, and automatically bring up a replacement volume. This doesn’t reduce the impact to zero, as it’s impossible to detect this failure before it happens, but it does ensure the majority of the cases don’t require a human to remediate and are over before users notice.\",{\"_93\":94,\"_41\":95,\"_96\":114,\"_48\":115},{},[116],\"At PlanetScale, our mitigations have clamped down on the expected maximum time for an impact window. We monitor metrics such as read/write latency and idle % closely, and we've even developed basic tests like making sure we can write to a file. This allows us to respond quickly to performance issues, and ensures that an EBS volume isn’t ‘stuck’.\",{\"_93\":94,\"_41\":118,\"_96\":119,\"_48\":120},\"h2\",{\"_50\":51},[121],[\"SingleFetchClassInstance\",122],{\"_93\":94,\"_41\":103,\"_96\":123,\"_48\":124},{\"_107\":125},[54],\"#handling-failure\",{\"_93\":94,\"_41\":95,\"_96\":127,\"_48\":128},{},[129],\"That’s the true rate of failure of EBS: it’s constant, variable, and all by design. Because there are no performance guarantees when volumes are not operating to their specifications, it is extremely difficult to plan around for workloads that require consistent performance. You can pay for additional nines, but with enough drives over a long enough timeframe, failure is guaranteed.\",{\"_93\":94,\"_41\":95,\"_96\":131,\"_48\":132},{},[133],\"With enough volumes, the rate of experiencing EBS failure is 100%: our automated mitigations are consistently recycling underperforming EBS volumes to reduce customer-impact, and we expect to see multiple events on a daily basis.\",{\"_93\":94,\"_41\":95,\"_96\":135,\"_48\":136},{},[137],[\"SingleFetchClassInstance\",138],{\"_93\":94,\"_41\":139,\"_96\":140,\"_48\":141},\"ContentImage\",{\"_142\":143,\"_144\":145,\"_146\":147,\"_148\":149,\"_150\":151,\"_152\":153},[],\"alt\",\"Correlated EBS Failure\",\"height\",580,\"loading\",\"lazy\",\"src\",\"https://planetscale-images.imgix.net/assets/correlated-ebs-failure-BcoLdHtC.png?auto=compress%2Cformat\",\"srcs\",[154],\"width\",1600,{\"_155\":149,\"_156\":157},\"srcSet\",\"media\",\"(prefers-color-scheme: light), (prefers-color-scheme: no-preference)\",{\"_93\":94,\"_41\":95,\"_96\":159,\"_48\":160},{},[161,162,163],\"To make matters worse, we also see these frequently as correlated failure inside of a single zone, even using \",[\"SingleFetchClassInstance\",164],\" volumes:\",{\"_93\":94,\"_41\":165,\"_96\":166,\"_48\":167},\"code\",{},[168],\"io2\",{\"_93\":94,\"_41\":95,\"_96\":170,\"_48\":171},{},[172],\"Even if you use io2, which AWS sells at 4x to 10x the price, you’d still be expected to be in a failure condition roughly one third of the time in any given year on just that one database!\",{\"_93\":94,\"_41\":95,\"_96\":174,\"_48\":175},{},[176],\"In a large database composed of many shards, this failure compounds. Assume a 256 shard database where each shard has one primary and two replicas: a total of 768 gp3 EBS volumes provisioned. If we take the 50% threshold from above, there is a 99.65% chance you have at least one node experiencing a production-impacting event at any given time.\",{\"_93\":94,\"_41\":95,\"_96\":178,\"_48\":179},{},[180],\"Let's assume the following: Each degradation event is random, meaning the level of reduced performance is somewhere between 1% and 89% of provisioned, and your application is designed to withstand losing 50% of its expected throughput before erroring. If each individual failure event lasts 10 minutes, every volume would experience about 43 events per month, with at least 21 of them causing downtime!\",{\"_93\":94,\"_41\":95,\"_96\":182,\"_48\":183},{},[184],\"This is not a secret, it's from the documentation. AWS doesn’t describe how failure is distributed for gp3 volumes, but in our experience it tends to last 1-10 minutes at a time. This is likely the time needed for a failover in a network or compute component.\",{\"_93\":94,\"_41\":95,\"_96\":186,\"_48\":187},{},[188],\"At PlanetScale, we see failures like this on a daily basis - the rate of failure is frequent enough that we’ve built systems that monitor EBS volumes directly to minimize impact.\",{\"_93\":94,\"_41\":118,\"_96\":190,\"_48\":191},{\"_50\":56},[192],[\"SingleFetchClassInstance\",193],{\"_93\":94,\"_41\":103,\"_96\":194,\"_48\":195},{\"_107\":196},[57],\"#the-true-rate-of-failure\",{\"_93\":94,\"_41\":95,\"_96\":198,\"_48\":199},{},[200],\"Production systems are not built to handle this level of sudden variance. When there are no guarantees, even overprovisioning doesn’t solve the problem. If this were a once-in-a-million chance, it would be different, but as we’ll discuss below, that is far from the case.\",{\"_93\":94,\"_41\":95,\"_96\":202,\"_48\":203},{},[204],[\"SingleFetchClassInstance\",205],{\"_93\":94,\"_41\":139,\"_96\":206,\"_48\":207},{\"_142\":208,\"_144\":209,\"_146\":147,\"_148\":210,\"_150\":211,\"_152\":212},[],\"Database blip due to EBS\",623,\"https://planetscale-images.imgix.net/assets/database-blip-ebs-CFXXjXsh.png?auto=compress%2Cformat\",[213],1547,{\"_155\":210,\"_156\":157},{\"_93\":94,\"_41\":95,\"_96\":215,\"_48\":216},{},[217],\"However, these short windows are enough to have impact on real-time workloads:\",{\"_93\":94,\"_41\":95,\"_96\":219,\"_48\":220},{},[221],[\"SingleFetchClassInstance\",222],{\"_93\":94,\"_41\":139,\"_96\":223,\"_48\":224},{\"_142\":225,\"_144\":226,\"_146\":147,\"_148\":227,\"_150\":228,\"_152\":229},[],\"EBS volume having a short blip\",694,\"https://planetscale-images.imgix.net/assets/ebs-volume-short-blip-B86N4wvl.png?auto=compress%2Cformat\",[230],1542,{\"_155\":227,\"_156\":157},{\"_93\":94,\"_41\":95,\"_96\":232,\"_48\":233},{},[234],\"In our experience, the documentation is accurate: sometimes volumes pass in and out of their provisioned performance in small time windows:\",{\"_93\":94,\"_41\":95,\"_96\":236,\"_48\":237},{},[238,239,240,241,242,243,244],\"This means a volume is expected to experience \",[\"SingleFetchClassInstance\",254],\" of its provisioned performance 1% of the time. That’s 14 minutes of every day or 86 hours out of the year of \",[\"SingleFetchClassInstance\",250],\" impact. This rate of degradation \",[\"SingleFetchClassInstance\",245],\" exceeds that of a single disk drive or SSD. This is the cost of separating storage and compute and the sheer complexity of the software and networking components between the client and the backing disks for the volume.\",{\"_93\":94,\"_41\":246,\"_96\":247,\"_48\":248},\"em\",{},[249],\"far\",{\"_93\":94,\"_41\":246,\"_96\":251,\"_48\":252},{},[253],\"potential\",{\"_93\":94,\"_41\":246,\"_96\":255,\"_48\":256},{},[257],\"under 90%\",{\"_93\":94,\"_41\":259,\"_96\":260,\"_48\":261},\"blockquote\",{},[262],[\"SingleFetchClassInstance\",263],{\"_93\":94,\"_41\":95,\"_96\":264,\"_48\":265},{},[266],\"When attached to an EBS–optimized instance, General Purpose SSD (gp2 and gp3) volumes are designed to deliver at least 90 percent of their provisioned IOPS performance 99 percent of the time in a given year\",{\"_93\":94,\"_41\":95,\"_96\":268,\"_48\":269},{},[270,271,272],\"The \",[\"SingleFetchClassInstance\",273],\" is useful in helping us understand what promises AWS’ gp3 is able to make:\",{\"_93\":94,\"_41\":103,\"_96\":274,\"_48\":275},{\"_107\":277},[276],\"EBS documentation\",\"https://docs.aws.amazon.com/ebs/latest/userguide/general-purpose.html#gp3-ebs-volume-type\",{\"_93\":94,\"_41\":95,\"_96\":279,\"_48\":280},{},[281],\"To the application running on top of this database: this is a failure. To the user, this is a 500 response on a webpage after a 10 second wait. To you, this is an incident. At PlanetScale, we consider this full failure because our customers do.\",{\"_93\":94,\"_41\":95,\"_96\":283,\"_48\":284},{},[285],\"This volume has been operating steadily for at least 10 hours. AWS has reported it at 67% idle, with write latency measuring at single-digit ms/operation. Well within expectations. Suddenly, at around 16:00, write latency spikes to 200ms-500ms/operation, idle time races to zero, and the volume is effectively blocked from reading and writing data.\",{\"_93\":94,\"_41\":95,\"_96\":287,\"_48\":288},{},[289],[\"SingleFetchClassInstance\",290],{\"_93\":94,\"_41\":139,\"_96\":291,\"_48\":292},{\"_142\":293,\"_144\":294,\"_146\":147,\"_148\":295,\"_150\":296,\"_152\":153},[],\"EBS Volume Degrading\",781,\"https://planetscale-images.imgix.net/assets/ebs-volume-degrading-D7MtBxSa.png?auto=compress%2Cformat\",[297],{\"_155\":295,\"_156\":157},{\"_93\":94,\"_41\":95,\"_96\":299,\"_48\":300},{},[301],\"Here’s what “slow” looks like, from the AWS Console:\",{\"_93\":94,\"_41\":95,\"_96\":303,\"_48\":304},{},[305],\"While full failure and data loss is very rare with EBS, “slow” is often as bad as “failed”, and that happens much much more often.\",{\"_93\":94,\"_41\":118,\"_96\":307,\"_48\":308},{\"_50\":59},[309],[\"SingleFetchClassInstance\",310],{\"_93\":94,\"_41\":103,\"_96\":311,\"_48\":312},{\"_107\":313},[60],\"#defining-failure\",{\"_93\":94,\"_41\":95,\"_96\":315,\"_48\":316},{},[317],\"For example, if a single instance inside of a multi-node distributed caching system runs out of networking resources, the downstream application will interpret error cases as cache misses to avoid failing the request. This will overwhelm the database when the application floods it with queries to fetch data as though it was missing. In this, a partial failure at one level results in a full failure of the database tier, causing downtime.\",{\"_93\":94,\"_41\":95,\"_96\":319,\"_48\":320},{},[321],\"In complex systems, failure isn’t a binary outcome. Cloud native systems are built without single paths of failure, but partial failure can still result in degraded performance, loss of user-facing availability, and undefined behavior. Often, minor failure in one part of the stack appears as a full failure in others.\",{\"_93\":94,\"_41\":95,\"_96\":323,\"_48\":324},{},[325],\"PlanetScale has deployed millions of Amazon Elastic Block Store (EBS) volumes across the world. We create and destroy tens of thousands of them every day as we stand up databases for customers, take backups, and test our systems end-to-end. Through this experience, we have an unique viewpoint into the failure rate and mechanisms of EBS, and have spent a lot of time working on how to mitigate them.\",\"current\",{\"_328\":329,\"_330\":331,\"_332\":329},\"development\",false,\"env\",{\"_333\":334,\"_335\":336,\"_337\":338,\"_339\":340,\"_341\":342},\"userSignedIn\",\"IMAGE_CDN\",\"https://planetscale-images.imgix.net\",\"IMAGE_CDN_ENABLED\",\"true\",\"INTERNAL_API\",\"https://api.planetscale.com\",\"RELEASE\",\"117b8aaf-965c-42bc-b013-5f72770de4d9\",\"SENTRY_DSN\",\"https://bd81903b44804e22a06bdc0c1a91b303@o499952.ingest.us.sentry.io/4504531942572032\"]\n");</script><!--$--><script nonce="EZdXajdlI2DRCjbnoKJCwfF5KNgWOgdluJsujYSRAY4=">window.__reactRouterContext.streamController.close();</script><!--/$--><!--/$--></body></html> |