192 lines
90 KiB
HTML
192 lines
90 KiB
HTML
<!DOCTYPE html><html lang="en"><head><meta charSet="utf-8"/><meta name="viewport" content="width=device-width, initial-scale=1"/><meta name="theme-color" content="#111111"/><meta name="user-signed-in" content="false"/><title>Anatomy of a Throttler, part 1 — PlanetScale</title><meta name="description" content="Learn about some design considerations for implementing a database throttler."/><meta name="robots"/><meta property="og:url" content="https://planetscale.com/blog/anatomy-of-a-throttler-part-1"/><meta property="og:type" content="website"/><meta property="og:title" content="Anatomy of a Throttler, part 1 — PlanetScale"/><meta property="og:image" content="https://planetscale.com/assets/anatomy-of-a-throttler-part-1-social-C0hGHHNN.jpg"/><meta property="og:description" content="Learn about some design considerations for implementing a database throttler."/><meta property="twitter:card" content="summary_large_image"/><meta property="twitter:site" content="@PlanetScale"/><meta property="twitter:creator" content="@PlanetScale"/><meta property="twitter:url" content="https://planetscale.com/blog/anatomy-of-a-throttler-part-1"/><meta property="twitter:title" content="Anatomy of a Throttler, part 1 — PlanetScale"/><meta property="twitter:description" content="Learn about some design considerations for implementing a database throttler."/><meta property="twitter:image" content="https://planetscale.com/assets/anatomy-of-a-throttler-part-1-social-C0hGHHNN.jpg"/><link rel="canonical" href="https://planetscale.com/blog/anatomy-of-a-throttler-part-1"/><link rel="preconnect" href="https://planetscale-images.imgix.net"/><link nonce="nyDuN2IITQLGs2NvcHjdoIBrtda9b+1NY1U/+JH3x2E=" rel="icon" href="/favicon.ico" type="image/x-icon" sizes="16x16"/><link nonce="nyDuN2IITQLGs2NvcHjdoIBrtda9b+1NY1U/+JH3x2E=" rel="icon" href="/icon.png" type="image/png" sizes="32x32"/><link nonce="nyDuN2IITQLGs2NvcHjdoIBrtda9b+1NY1U/+JH3x2E=" rel="apple-touch-icon" href="/apple-touch-icon.png" type="image/png" sizes="32x32"/><link nonce="nyDuN2IITQLGs2NvcHjdoIBrtda9b+1NY1U/+JH3x2E=" rel="manifest" href="/manifest.webmanifest"/><link rel="modulepreload" href="/assets/entry.client-3vubyXrk.js" nonce="nyDuN2IITQLGs2NvcHjdoIBrtda9b+1NY1U/+JH3x2E="/><link rel="modulepreload" href="/assets/jsx-runtime-DwfQwkRq.js" nonce="nyDuN2IITQLGs2NvcHjdoIBrtda9b+1NY1U/+JH3x2E="/><link rel="modulepreload" href="/assets/components-_bNmAApg.js" nonce="nyDuN2IITQLGs2NvcHjdoIBrtda9b+1NY1U/+JH3x2E="/><link rel="modulepreload" href="/assets/_.well-known_.mcp.server-card_.json_-Sx7XeH3e.js" nonce="nyDuN2IITQLGs2NvcHjdoIBrtda9b+1NY1U/+JH3x2E="/><link rel="modulepreload" href="/assets/index-mKTXLmHu.js" nonce="nyDuN2IITQLGs2NvcHjdoIBrtda9b+1NY1U/+JH3x2E="/><link rel="modulepreload" href="/assets/errorBoundaries-DhW4jVYt.js" nonce="nyDuN2IITQLGs2NvcHjdoIBrtda9b+1NY1U/+JH3x2E="/><link rel="modulepreload" href="/assets/root-DbOv4-98.js" nonce="nyDuN2IITQLGs2NvcHjdoIBrtda9b+1NY1U/+JH3x2E="/><link rel="modulepreload" href="/assets/lib-Dg89tQ22.js" nonce="nyDuN2IITQLGs2NvcHjdoIBrtda9b+1NY1U/+JH3x2E="/><link rel="modulepreload" href="/assets/analytics.client-DM6E8o1h.js" nonce="nyDuN2IITQLGs2NvcHjdoIBrtda9b+1NY1U/+JH3x2E="/><link rel="modulepreload" href="/assets/SiteHeader-C2U5gvDH.js" nonce="nyDuN2IITQLGs2NvcHjdoIBrtda9b+1NY1U/+JH3x2E="/><link rel="modulepreload" href="/assets/current-9yDxj94E.js" nonce="nyDuN2IITQLGs2NvcHjdoIBrtda9b+1NY1U/+JH3x2E="/><link rel="modulepreload" href="/assets/clsx-eT0YPcGk.js" nonce="nyDuN2IITQLGs2NvcHjdoIBrtda9b+1NY1U/+JH3x2E="/><link rel="modulepreload" href="/assets/bugs-38ilEoW0.js" nonce="nyDuN2IITQLGs2NvcHjdoIBrtda9b+1NY1U/+JH3x2E="/><link rel="modulepreload" href="/assets/keyboard-D-uXZORL.js" nonce="nyDuN2IITQLGs2NvcHjdoIBrtda9b+1NY1U/+JH3x2E="/><link rel="modulepreload" href="/assets/use-tab-direction-dKm-S3Ck.js" nonce="nyDuN2IITQLGs2NvcHjdoIBrtda9b+1NY1U/+JH3x2E="/><link rel="modulepreload" href="/assets/blog-pXH7ptHJ.js" nonce="nyDuN2IITQLGs2NvcHjdoIBrtda9b+1NY1U/+JH3x2E="/><link rel="modulepreload" href="/assets/blog._slug-Ch_92qsH.js" nonce="nyDuN2IITQLGs2NvcHjdoIBrtda9b+1NY1U/+JH3x2E="/><link rel="modulepreload" href="/assets/ContentImage-Dh6VEOUl.js" nonce="nyDuN2IITQLGs2NvcHjdoIBrtda9b+1NY1U/+JH3x2E="/><link rel="modulepreload" href="/assets/BlogCategoryLink-DmQyn0gp.js" nonce="nyDuN2IITQLGs2NvcHjdoIBrtda9b+1NY1U/+JH3x2E="/><link rel="modulepreload" href="/assets/Details-BSB_b6hI.js" nonce="nyDuN2IITQLGs2NvcHjdoIBrtda9b+1NY1U/+JH3x2E="/><link rel="modulepreload" href="/assets/Skittle-CDFOPRjH.js" nonce="nyDuN2IITQLGs2NvcHjdoIBrtda9b+1NY1U/+JH3x2E="/><link rel="modulepreload" href="/assets/SiteFooter-B2Gq9u2j.js" nonce="nyDuN2IITQLGs2NvcHjdoIBrtda9b+1NY1U/+JH3x2E="/><link rel="modulepreload" href="/assets/Vimeo-00PQJDli.js" nonce="nyDuN2IITQLGs2NvcHjdoIBrtda9b+1NY1U/+JH3x2E="/><link rel="modulepreload" href="/assets/YouTube-CMfaljVr.js" nonce="nyDuN2IITQLGs2NvcHjdoIBrtda9b+1NY1U/+JH3x2E="/><link rel="modulepreload" href="/assets/date-CJTFH3uT.js" nonce="nyDuN2IITQLGs2NvcHjdoIBrtda9b+1NY1U/+JH3x2E="/><link rel="modulepreload" href="/assets/use-inert-others-BMJ6-xOX.js" nonce="nyDuN2IITQLGs2NvcHjdoIBrtda9b+1NY1U/+JH3x2E="/><link rel="modulepreload" href="/assets/description-Cf6FZmDe.js" nonce="nyDuN2IITQLGs2NvcHjdoIBrtda9b+1NY1U/+JH3x2E="/><link rel="modulepreload" href="/assets/use-is-mounted-uQsUZyP9.js" nonce="nyDuN2IITQLGs2NvcHjdoIBrtda9b+1NY1U/+JH3x2E="/><link rel="modulepreload" href="/assets/types-DvonrUFF.js" nonce="nyDuN2IITQLGs2NvcHjdoIBrtda9b+1NY1U/+JH3x2E="/><link rel="stylesheet" href="/assets/styles-ns8XBZ1D.css"/><script nonce="nyDuN2IITQLGs2NvcHjdoIBrtda9b+1NY1U/+JH3x2E=">window.ENV = {"IMAGE_CDN":"https://planetscale-images.imgix.net","IMAGE_CDN_ENABLED":"true","INTERNAL_API":"https://api.planetscale.com","RELEASE":"117b8aaf-965c-42bc-b013-5f72770de4d9","SENTRY_DSN":"https://bd81903b44804e22a06bdc0c1a91b303@o499952.ingest.us.sentry.io/4504531942572032"}</script></head><body class="flex min-h-screen flex-col"><div class="bg-neki px-3 py-1 text-center font-medium text-gray-900 dark:font-semibold"><span>Neki, sharded Postgres, is now available.</span> <span class="whitespace-nowrap"><a href="https://auth.planetscale.com/sign-up" class="whitespace-nowrap bg-gray-900 px-sm font-semibold text-white">Get started</a></span></div><header class="relative mb-6 mt-4 bg-primary"><div class="flex flex-col gap-y-3 px-3 sm:px-5 container max-w-7xl"><div class="grid w-full grid-cols-[auto_1fr] grid-rows-1 items-center lg:items-start lg:gap-3"><a aria-label="Go to homepage" class="col-start-1 col-end-2 h-4 w-4 rounded-full text-primary lg:hidden" href="/" data-discover="true"><svg xmlns="http://www.w3.org/2000/svg" width="32" height="32" fill="none" viewBox="0 0 40 40"><path fill="currentColor" d="M0 20C0 8.954 8.954 0 20 0c8.121 0 15.112 4.84 18.245 11.794l-26.45 26.45a20 20 0 0 1-3.225-1.83L24.984 20H20L5.858 34.142A19.94 19.94 0 0 1 0 20M39.999 20.007 20.006 40c11.04-.004 19.99-8.953 19.993-19.993"></path></svg></a><div class="group col-start-2 col-end-3 row-start-1 flex shrink-0 items-center justify-end gap-1.5 lg:gap-3"><div class="flex flex-row gap-2 lg:flex-col lg:gap-1 xl:flex-row"><div class="flex items-center justify-end gap-1 lg:h-4"><a href="https://auth.planetscale.com/sign-in" class="font-semibold text-primary hover:text-orange">Sign in</a></div><div class="flex items-center justify-end gap-0.5 lg:h-4"><form class="btn-sm hidden sm:inline-flex" action="/api/demo-sessions" method="post"><button type="submit" class="btn btn-outline btn-sm hidden sm:inline-flex">View sandbox</button></form><a class="btn btn-sm" href="/contact" data-discover="true">Get in touch</a></div></div></div><div class="col-start-1 col-end-2 flex items-center gap-x-3 lg:row-start-1 lg:h-4"><a aria-label="Go to homepage" class="col-start-1 col-end-2 hidden h-4 w-4 rounded-full text-primary lg:block" href="/" data-discover="true"><svg xmlns="http://www.w3.org/2000/svg" width="32" height="32" fill="none" viewBox="0 0 40 40"><path fill="currentColor" d="M0 20C0 8.954 8.954 0 20 0c8.121 0 15.112 4.84 18.245 11.794l-26.45 26.45a20 20 0 0 1-3.225-1.83L24.984 20H20L5.858 34.142A19.94 19.94 0 0 1 0 20M39.999 20.007 20.006 40c11.04-.004 19.99-8.953 19.993-19.993"></path></svg></a><nav aria-label="Main" data-orientation="horizontal" class="hidden items-center lg:flex"><ul class="flex flex-wrap gap-x-1 md:flex-nowrap"><li><div data-headlessui-state=""><button class="font-semibold text-primary hover:text-contrast focus-visible:ring-0 ui-open:text-orange" type="button" aria-expanded="false" data-headlessui-state="">Platform<span class="ml-sm inline-block ui-open:rotate-180">▾</span></button></div><span hidden="" style="position:fixed;top:1px;left:1px;width:1px;height:0;padding:0;margin:-1px;overflow:hidden;clip:rect(0, 0, 0, 0);white-space:nowrap;border-width:0;display:none"></span></li><li class="text-decoration" role="presentation">|</li><li><div data-headlessui-state=""><button class="font-semibold text-primary hover:text-contrast focus-visible:ring-0 ui-open:text-orange" type="button" aria-expanded="false" data-headlessui-state="">Resources<span class="ml-sm inline-block ui-open:rotate-180">▾</span></button></div><span hidden="" style="position:fixed;top:1px;left:1px;width:1px;height:0;padding:0;margin:-1px;overflow:hidden;clip:rect(0, 0, 0, 0);white-space:nowrap;border-width:0;display:none"></span></li><li class="text-decoration" role="presentation">|</li><li><a class="font-semibold text-primary hover:text-contrast" href="/docs">Documentation</a></li><li class="text-decoration" role="presentation">|</li><li><a class="font-semibold text-primary hover:text-contrast" href="/pricing" data-discover="true">Pricing</a></li><li class="text-decoration" role="presentation">|</li><li><a class="font-semibold text-primary hover:text-contrast" href="/migrate" data-discover="true">Migrate</a></li></ul></nav></div></div><details class="lg:hidden"><summary>Navigation</summary><nav class="dashed-box mt-1 p-3"><ul class="flex flex-wrap gap-x-1 md:flex-nowrap"><li><div data-headlessui-state=""><button class="font-semibold text-primary hover:text-contrast focus-visible:ring-0 ui-open:text-orange" type="button" aria-expanded="false" data-headlessui-state="">Platform<span class="ml-sm inline-block ui-open:rotate-180">▾</span></button></div><span hidden="" style="position:fixed;top:1px;left:1px;width:1px;height:0;padding:0;margin:-1px;overflow:hidden;clip:rect(0, 0, 0, 0);white-space:nowrap;border-width:0;display:none"></span></li><li class="text-decoration" role="presentation">|</li><li><div data-headlessui-state=""><button class="font-semibold text-primary hover:text-contrast focus-visible:ring-0 ui-open:text-orange" type="button" aria-expanded="false" data-headlessui-state="">Resources<span class="ml-sm inline-block ui-open:rotate-180">▾</span></button></div><span hidden="" style="position:fixed;top:1px;left:1px;width:1px;height:0;padding:0;margin:-1px;overflow:hidden;clip:rect(0, 0, 0, 0);white-space:nowrap;border-width:0;display:none"></span></li><li class="text-decoration" role="presentation">|</li><li><a class="font-semibold text-primary hover:text-contrast" href="/docs">Documentation</a></li><li class="text-decoration" role="presentation">|</li><li><a class="font-semibold text-primary hover:text-contrast" href="/pricing" data-discover="true">Pricing</a></li><li class="text-decoration" role="presentation">|</li><li><a class="font-semibold text-primary hover:text-contrast" href="/migrate" data-discover="true">Migrate</a></li></ul></nav></details></div></header><main class="container mb-6 flex max-w-7xl flex-1 flex-col px-3 sm:px-5 lg:px-12"><section class=""><p class="block"><a class="pr-sm text-primary hover:text-contrast" href="/blog" data-discover="true">Blog</a><span class="px-sm text-decoration">|</span><a class="px-sm text-blue hover:bg-blue-100 dark:hover:bg-blue-900" href="/blog/category/engineering" data-discover="true">Engineering</a></p><div class="flex lg:flex-row-reverse lg:gap-x-6"><div class="lg:sticky lg:top-2 lg:self-start"><button class="absolute right-0 bg-gray-100 px-sm md:block lg:hidden dark:bg-gray-800 -mt-9 hidden"><span class="inline">Table of contents «</span><span class="hidden">Close »</span></button><aside class="tree-nav w-full shrink-0 space-y-3 lg:w-36 hidden lg:block"><div><h4 class="text-secondary">Table of contents</h4><ul><li><a class="font-semibold text-primary hover:text-blue" href="/blog/anatomy-of-a-throttler-part-1#which-requests-do-you-throttle" data-discover="true">Which requests do you throttle?</a></li><li><a class="font-semibold text-primary hover:text-blue" href="/blog/anatomy-of-a-throttler-part-1#what-does-the-throttler-throttle-on" data-discover="true">What does the throttler throttle on?</a><ul><li><a class="text-primary hover:text-blue" href="/blog/anatomy-of-a-throttler-part-1#replication-lag" data-discover="true">Replication lag</a></li><li><a class="text-primary hover:text-blue" href="/blog/anatomy-of-a-throttler-part-1#other-database-metrics-whats-in-a-metric" data-discover="true">Other database metrics: what's in a metric?</a></li><li><a class="text-primary hover:text-blue" href="/blog/anatomy-of-a-throttler-part-1#a-closer-examination-queues" data-discover="true">A closer examination: queues</a></li><li><a class="text-primary hover:text-blue" href="/blog/anatomy-of-a-throttler-part-1#pool-usage" data-discover="true">Pool usage</a></li><li><a class="text-primary hover:text-blue" href="/blog/anatomy-of-a-throttler-part-1#the-case-for-multiple-metrics" data-discover="true">The case for multiple metrics</a></li></ul></li><li><a class="font-semibold text-primary hover:text-blue" href="/blog/anatomy-of-a-throttler-part-1#what-does-a-throttled-system-look-like" data-discover="true">What does a throttled system look like?</a></li><li><a class="font-semibold text-primary hover:text-blue" href="/blog/anatomy-of-a-throttler-part-1#check-intervals-and-metric-granularity" data-discover="true">Check intervals and metric granularity</a></li><li><a class="font-semibold text-primary hover:text-blue" href="/blog/anatomy-of-a-throttler-part-1#to-be-continued" data-discover="true">To be continued</a></li></ul><div class="mb-3 mt-6 border bg-blue-50 p-3 font-semibold text-contrast dark:bg-blue-900"><p>PlanetScale, the fastest cloud Postgres, from $5/month.</p><p><a href="https://app.planetscale.com/new">Start now</a></p></div><p>Get the <a href="/blog/feed.atom">RSS feed</a></p></div></aside></div><article class="min-w-0 flex-grow"><h1>Anatomy of a Throttler, part 1</h1><p class="text-secondary"><a class="text-contrast no-underline" href="/blog/author/shlomi" data-discover="true">Shlomi Noach</a> |<!-- --> <time dateTime="2024-08-29">August 29, 2024</time></p><div class="blog-post-body"><p>A throttler is a service or component that pushes back against an <a href="https://en.wikipedia.org/wiki/Flow_control_(data)">incoming flow</a> of requests to ensure the system has the capacity to handle all received requests without being overwhelmed. In this series of posts, we illustrate design considerations for a database system throttler, whose purpose is to keep the database system healthy overall. We discuss choice of metrics, granularity, behavior, impact, prioritization, and other topics.</p><h2 id="which-requests-do-you-throttle"><a href="#which-requests-do-you-throttle">Which requests do you throttle?</a></h2><p>There are different approaches to throttling requests in a database. We focus on throttling asynchronous, batch, and massive operations that are not time critical. Examples could be ETLs, data imports, online DDL operations, mass purges of data, resharding, and so forth. The throttler will push back on those operations that can span minutes, hours, or days of operation. Other forms of throttling may push back on OLTP production traffic. This discussion applies equally to both.</p><p>By way of illustration, consider a job that needs to import 10 million rows into the database. Instead of attempting to apply all 10 million in one go, the job breaks down the task into much smaller subtasks: it will try to import (write) 100 rows at a time. Before any such import, it will request access from the throttler.</p><p>Some throttler implementations are collaborative, meaning they assume clients will respect their instructions. Others act as barriers between the app and the database. Either way, if the throttler indicates that the database is overloaded, the job should hold back for a period of time and then request access again. This process repeats until granted. Each subtask should be small enough so as not to single-handedly tank the database's serving capacity, while large enough to compensate for the added throttler overhead and to enable meaningful progress.</p><h2 id="what-does-the-throttler-throttle-on"><a href="#what-does-the-throttler-throttle-on">What does the throttler throttle on?</a></h2><p>Some generic throttlers only allow a regulated rate of requests, in anticipation that the consuming job will be able to process them at some known, fixed rate. With databases, things are less clear. A database can only handle so many queries at any given point in time, or over some period of time. However, not all queries are created equal. The database capacity of serving queries depends on the scope of queries, any hot spots or cold spots in affected data, the state of the page cache, overlap or lack thereof of data served by queries, to name a few factors.</p><p>We therefore need to be able to determine: how do we consider our database to be "healthy"? How do we determine if it is being overwhelmed?</p><p>To do this, we look for metrics that define or predict service level objectives (SLO) for the database. But things are not always so simple. Let's start with a popular metric which is widely used as a throttling indicator, and see what's so special about that metric.</p><h3 id="replication-lag"><a href="#replication-lag">Replication lag</a></h3><p>Replication is often used in database clusters, especially ones using a <a href="https://dev.mysql.com/doc/refman/8.4/en/group-replication-primary-secondary-replication.html">primary-secondary</a> (aka leader-follower) architecture. Sometimes the replication is set to be asynchronous, and at other times group-communication based. In these scenarios <strong>replication lag</strong> is defined as the time passing between a write on the primary server and the time it is applied or made visible on the replica/secondary server.</p><p>In the MySQL world, replication lag is probably the single most used throttling indicator, as multiple third-party and community tools use it to push back against long-running jobs. This is for good reasons: it is easy to measure and it has clear impact on the product and the business. For example, in the case of a database failover, replication lag impacts the time it takes for a replica/standby server to be promoted and made available to receive write requests. <a href="https://jepsen.io/consistency/models/read-your-writes">Read-after-write</a> can be simplified when replication lag is low, allowing secondary servers to serve some of the read traffic.</p><p>We can thus have a business limitation on the acceptable replication lag, which we can use in the throttler: below this lag, allow requests. Above this lag, push back.</p><h3 id="other-database-metrics-whats-in-a-metric"><a href="#other-database-metrics-whats-in-a-metric">Other database metrics: what's in a metric?</a></h3><p>Another common metric in the MySQL world is the value of <code>threads_running</code>. On any given server, this is the number of concurrent, actively executing queries (not to be confused with concurrent open transactions, some of which could be idle in between queries). This metric is frequently seen on database dashboards, and is an indicator for the database load.</p><p>But, what's an acceptable value? Are <code>50</code> concurrent queries OK? Are <code>100</code> OK? Pick a number, and you'll soon find it doesn't hold water. Some values are acceptable in early morning, while others are just normal during peak traffic hours. As your product evolves and its adoption increases, so do the queries on your database. What was true 3 months ago is not true today. And again, not all queries are created equal.</p><p>What's different about this metric compared with replication lag is that it is much more of a <em>symptom</em> than an actual <em>cause</em>. If all of a sudden we see a sharp spike in active queries, this can indicate some possible causes: perhaps all are held by the commit queue, which for some reason stalls. Or, the queries happen to compete over a specific hotspot and wait on locks. Or, they don't, they all happen to compete on very different pages, none of which is in memory, and they all congest while waiting on the page cache, etc. So what is it exactly that we need to monitor? Is the metric itself useless?</p><p>Not necessarily. An experienced administrator may only need to take one look at this metric on the database dashboard to say "we're having an issue".</p><h3 id="a-closer-examination-queues"><a href="#a-closer-examination-queues">A closer examination: queues</a></h3><p>Software relies heavily on <a href="https://en.wikipedia.org/wiki/Queueing_theory">queues</a>. They are fundamental not only in software design but also in hardware access. Requests queue on network access. They queue on disk access. They queue on CPU access. They queue on locks.</p><p>Circling back to replication lag, much like concurrent queries, it is a <em>symptom</em>. E.g. disk I/O is saturated on the replica, hence the replica cannot keep up replaying the changelog, thereby accumulating lag. Or perhaps the lag is caused by slow network. Or both! Whatever the case is, what's interesting is that the replication mechanism itself is a queue: the changelog event queue. A new write on the primary manifests as a "write event", which is shipped to the replica, and waits to be consumed (processed, replayed) by the replica. Replication lag is the event's time spent in the queue, where our queue is a combination of the network queue, local disk write queue, actual wait time, and finally the event's processing time. Each of these can be the major contributor to the overall replication lag, and yet, we can still look at replication lag as a whole — as a clear indicator for database health.</p><p>Armed with this insight, we take another attempt at understanding other metrics. In the case for concurrent writes, we understand a major contributor to a spike in concurrent queries is their inability to <em>complete</em>. Normally, this means they're held back at <em>commit time</em>, i.e. they wait to be written to the transaction (redo) log. And that means they're in the transaction <em>queue</em>, and we can hence measure the transaction queue latency (aka queue delay).</p><p>But, what's a good threshold? Transaction commit delay is typically caused by disk write/flush time, and that <a href="https://hackmysql.com/commit-latency-aurora-vs-rds-mysql-8.0/">changes dramatically</a> across hardware. It is a matter of knowing your metrics. Yet again, an experienced administrator should know what values to expect. But now the values are more tightly bound to hardware and slightly less affected by the app.</p><p>Queue delay is not the only metric. Another common one is the queue length: the number of entries waiting in the queue. A long queue at the airport isn't in itself a bad thing, some queues move quite fast, and yet it's often a predictor to wait times. Where wait time is impossible or difficult to measure, queue length can be an alternative.</p><p>An operating system's <em>Load Average</em> metric evaluation includes the number of processes waiting for CPU time. This changes by the number of CPUs available. A common rough indicator is a <code>1</code> threshold for <code>(load average)/(num CPUs)</code>. This is again a metric that must agree with your own systems. Some database deployments famously push their servers to their limits with load averages soaring far above <code>1</code> per CPU.</p><h3 id="pool-usage"><a href="#pool-usage">Pool usage</a></h3><p>Another indicator is pool usage. The single most common pool with regard to databases must be the application's database connection pool. To run a query, the app will take a connection from the pool, use it to execute the query, then return the connection to the pool. If the pool has connections to spare, getting that connection comes at no cost. But if the pool is exhausted, then either the app needs to wait for the creation of a new connection, or it gets rejected. Similarly to concurrent queries, a high pool usage indicates a congestion of operations. However, pooled connections can be used across multiple queries in a transaction, as well as across multiple transactions, and the app may run its own logic in between running queries, while still holding on to the connection.</p><p>An exhausted pool is a strong indication of excessive load, while the difference between a <code>60%</code> and an <code>80%</code> used pool is not as clear an indication. Taking a step back, what does it mean that we exhaust some pool? Who decides the size of the pool in the first place? If someone picked a number such as <code>50</code> or <code>100</code>, isn't that number just artificial?</p><p>It may well be, but pool size was likely chosen for some good reason(s). It is perhaps derived from some database configuration, which is itself derived from some hardware limitation. And while the choice of metric could possibly change arbitrarily, it is still sensible, as far as throttling goes, to push back when the pool is exhausted. The throttler thereby relies on the greater system configuration and does not introduce any new artificial thresholds.</p><h3 id="the-case-for-multiple-metrics"><a href="#the-case-for-multiple-metrics">The case for multiple metrics</a></h3><p>A throttler should be able to push back based on a combination of metrics, and not limit itself to just one metric. We've illustrated some metrics above, and every environment may yet have its own load predicting metrics. The administrator should be able to choose an assorted set of metrics the throttler should work with, be able to set specific thresholds for each such metric, and possibly be able to introduce new metrics either programmatically or dynamically.</p><h2 id="what-does-a-throttled-system-look-like"><a href="#what-does-a-throttled-system-look-like">What does a throttled system look like?</a></h2><p>Many software developers will be familiar with the next scenario: you have a multithreaded or otherwise highly concurrent app. There's a bug, likely a race condition or a synchronization issue, and you wish to find it. You choose to print informative debug messages to standard output, and hope to find the bug by examining the log. Alas, when you do so, the bug does not reproduce. Or it may manifest elsewhere.</p><p>In adding writes to standard output, you have introduced new locks. Your debug messages now compete over those locks, which in turn incurs different context switches.</p><p>Introducing a throttler into your infrastructure shows resemblances to this synchronization example. All of a sudden, there is less contention on the database, and certain apps that used to run just fine, exhibit contention/latency behavior. The appearance of a new job suddenly affects the progress of another. But where previously you could clearly analyze database queries to find the root cause, the database now tells you little to nothing. It's now down to the throttler to give you that information. But even the throttler is limited, because all the apps do is to check the throttler for health status. They do not yet actually <em>do</em> anything.</p><p>Let's say we throttle based on replication lag, and let's assume that we want to run an operation so massive that it is <em>bound</em> to drive replication lag high if let loose. With the throttler keeping it under control, though, the operation will only run small batches of subtasks. But an interesting behavior emerges: the operation will push replication lag up to the throttler's threshold, then back down, and push again. As we start the operation, we expect to see the replication lag graph jump up to the threshold value, and then more or less stabilize around that value, slightly higher and slightly lower, for the duration of the operation, which could be hours.</p><p>During that time, the operation will be granted access thousands of times or more, and will likewise also be rejected access thousands of times or more. That is how a healthy system looks with a throttler engaged. No matter how many more concurrent operations we run, we expect to contain replication lag at about the same slight offset above or below the threshold. More on this when we discuss granularity.</p><p>It is not uncommon for a system to run one or two operations for very long periods, which means what we consider as the throttling threshold (say, a <code>5sec</code> replication lag) becomes the actual standard. Thankfully, not all operations and workloads are so aggressive that they necessarily push the metrics as high as their thresholds.</p><h2 id="check-intervals-and-metric-granularity"><a href="#check-intervals-and-metric-granularity">Check intervals and metric granularity</a></h2><p>A throttler collects the metrics asynchronously from check requests, so that it has an immediate answer available upon request. The intervals at which the throttler collects metrics can have a significant effect on how the throttler is being put to use. Let's consider a case where the throttler collects a metric at a large interval, say every 5 seconds. The metric could be anything at all during those 5 seconds, but it is the specific sampling that takes place at the end of that period that counts.</p><p>Similarly, other metrics could have some granularity. Namely, replication lag can be measured in different methods, and the most common one is by deliberate injection of heartbeat events on the primary, and by capturing them on a replica. More on this in another post, but the intervals in which the heartbeat events are generated dictate the granularity or the accuracy of the measured lag. Let's assume we inject heartbeats at one second intervals, and we've just injected a heartbeat at precisely noon. Let's also assume we <em>sample</em> the metrics once per second, and we happen to make that sample at <code>12:00:00.995</code>. The sample still reads <code>12:00:00.000</code> as this was our last injected metric. A client then checks the throttler at <code>12:00:01.990</code>. By now there will have been a new metric value, but one which we have not sampled yet. The throttler responds by using its last sample that is almost, but not quite, one second old, and which in itself represents a metric that is now almost, but not quite, two seconds old.</p><p>Long heartbeat intervals and outdated information have negative impacts on both our system health as well as the throttler's utilization.</p><p>On one hand, it is possible that in the duration of the interval, we miss noticing a significant uptick in system load. We'd only find out about it a few seconds later, at which time the throttler would be engaged. However, by that time, the system performance already degrades. It will take a few seconds before it comes down to acceptable values. But then again, once we do catch that the metrics exceed their thresholds, and for the duration of the next interval, we reject all further requests. If the metrics do turn healthy sooner than that, that's a missed opportunity to make some progress. Thus, we degrade the database's operations capacity.</p><p>When multiple operations attempt to make progress all at once, all will be throttled while metrics are above threshold, and possibly all released at once when metrics return to low values, thus all pushing the metrics up at once.</p><p>Borrowing from the world of networking hardware, it is recommended that metric interval and granularity oversample the range of allowed thresholds. For example, if the acceptable replication lag is at 5 seconds, then it's best to have a heartbeat/sampling interval of 1-2 seconds.</p><p>Lower intervals and more accurate metrics reduce spikes and spread the workload more efficiently. That, too, comes at a cost, which we will discuss in a later post.</p><h2 id="to-be-continued"><a href="#to-be-continued">To be continued</a></h2><p>In the next part of this series we will be looking into singular vs. distributed throttler design, as well as the impact the throttler itself may have on your environment.</p></div></article></div></section></main><footer class="mb-6 mt-10 px-3 sm:px-5 container max-w-7xl"><nav class="grid grid-cols-1 text-left sm:grid-cols-2 lg:grid-cols-5 lg:mx-7"><div class="dashed-box dashed-box-x-t sm:dashed-box-l-t lg:dashed-box-y-l p-3"><h2 class="font-semibold">Company</h2><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/about" data-discover="true">About</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/brand" data-discover="true">Brand</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/blog" data-discover="true">Blog</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/changelog" data-discover="true">Changelog</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/careers" data-discover="true">Careers</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/events" data-discover="true">Events</a></div><div class="dashed-box dashed-box-x-t lg:dashed-box-y-l p-3"><h2 class="font-semibold">Product</h2><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/case-studies" data-discover="true">Case studies</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/enterprise" data-discover="true">Enterprise</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/pricing" data-discover="true">Pricing</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/benchmarks" data-discover="true">Benchmarks</a></div><div class="dashed-box dashed-box-x-t sm:dashed-box-l-t lg:dashed-box-y-l p-3"><h2 class="font-semibold">Resources</h2><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/docs">Documentation</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/migrate" data-discover="true">Migrate</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="https://support.planetscale.com/hc/en-us" rel="nofollow noopener noreferrer" target="_blank">Support</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="https://planetscalestatus.com" rel="nofollow noopener noreferrer" target="_blank">Status</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="https://trust.planetscale.com" rel="nofollow noopener noreferrer" target="_blank">Trust Center</a></div><div class="dashed-box dashed-box-x-t lg:dashed-box-y-l p-3"><h2 class="font-semibold">Courses</h2><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/learn/courses/mysql-for-developers" data-discover="true">MySQL for Developers</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/learn/courses/database-scaling" data-discover="true">Database Scaling</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/learn/courses/vitess" data-discover="true">Learn Vitess</a></div><div class="dashed-box p-3 sm:col-span-2 lg:col-span-1"><h2 class="font-semibold text-primary hover:text-contrast">Open source</h2><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/vitess" data-discover="true">Vitess</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="https://vitess.io/slack" rel="nofollow noopener noreferrer" target="_blank">Vitess community</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="https://github.com/planetscale" rel="me nofollow noopener noreferrer" target="_blank">GitHub</a></div></nav><div class="dashed-box dashed-box-x-b p-3 lg:mx-7"><p class="mb-3 md:mb-0"><a class="text-primary" rel="nofollow" href="/legal/privacy" data-discover="true">Privacy</a><span class="text-decoration" role="presentation"> <!-- -->|<!-- --> </span><a class="text-primary" rel="nofollow" href="/legal/siteterms" data-discover="true">Terms</a><span class="text-decoration" role="presentation"> <!-- -->|<!-- --> </span><a class="text-primary" rel="nofollow" href="/legal/cookies" data-discover="true">Cookies</a><span class="text-decoration" role="presentation"> <!-- -->|<!-- --> </span><a class="text-primary" rel="nofollow" href="/legal/patents" data-discover="true">Patents</a><span class="text-decoration" role="presentation"> <!-- -->|<!-- --> </span><a class="text-primary" rel="nofollow" href="/legal/privacy#privacy-rights-and-choices" data-discover="true">Do Not Share My Personal Information</a></p><p class="text-secondary">© <!-- -->2026<!-- --> PlanetScale, Inc. All rights reserved.</p></div><p class="mb-0 mt-3 break-normal lg:mx-7"><a class="text-primary" href="https://github.com/planetscale" rel="me nofollow noopener noreferrer" target="_blank">GitHub</a><span class="text-decoration" role="presentation"> <!-- -->|<!-- --> </span><a aria-label="X (formerly Twitter)" class="text-primary" href="https://twitter.com/planetscale" rel="me nofollow noopener noreferrer" target="_blank">X</a><span class="text-decoration" role="presentation"> <!-- -->|<!-- --> </span><a aria-label="LinkedIn" class="text-primary" href="https://www.linkedin.com/company/planetscale" target="_blank" rel="noreferrer">LinkedIn</a><span class="text-decoration" role="presentation"> <!-- -->|<!-- --> </span><a class="text-primary" href="https://www.youtube.com/planetscale" rel="me nofollow noopener noreferrer" target="_blank">YouTube</a><span class="text-decoration" role="presentation"> <!-- -->|<!-- --> </span><a aria-label="Discord" class="text-primary" href="https://pscale.link/community" rel="nofollow noopener noreferrer" target="_blank">Discord</a><span class="text-decoration" role="presentation"> <!-- -->|<!-- --> </span><a class="text-primary" href="https://www.facebook.com/planetscaledata" rel="me nofollow noopener noreferrer" target="_blank">Facebook</a></p></footer><script nonce="nyDuN2IITQLGs2NvcHjdoIBrtda9b+1NY1U/+JH3x2E=">((storageKey2, restoreKey) => {
|
|
if (!window.history.state || !window.history.state.key) {
|
|
let key2 = Math.random().toString(32).slice(2);
|
|
window.history.replaceState({ key: key2 }, "");
|
|
}
|
|
try {
|
|
let storedY = JSON.parse(sessionStorage.getItem(storageKey2) || "{}")[restoreKey || window.history.state.key];
|
|
if (typeof storedY === "number") window.scrollTo(0, storedY);
|
|
} catch (error2) {
|
|
console.error(error2);
|
|
sessionStorage.removeItem(storageKey2);
|
|
}
|
|
})("react-router-scroll-positions", null)</script><script nonce="nyDuN2IITQLGs2NvcHjdoIBrtda9b+1NY1U/+JH3x2E=">window.__reactRouterContext = {"basename":"/","future":{"unstable_enableNodeReadableStream":false,"unstable_optimizeDeps":true},"routeDiscovery":{"mode":"lazy","manifestPath":"/__manifest"},"ssr":true,"isSpaMode":false};window.__reactRouterContext.stream = new ReadableStream({start(controller){window.__reactRouterContext.streamController = controller;}}).pipeThrough(new TextEncoderStream());</script><script nonce="nyDuN2IITQLGs2NvcHjdoIBrtda9b+1NY1U/+JH3x2E=" type="module" async="">;
|
|
import * as route0 from "/assets/root-DbOv4-98.js";
|
|
import * as route1 from "/assets/blog-pXH7ptHJ.js";
|
|
import * as route2 from "/assets/blog._slug-Ch_92qsH.js";
|
|
window.__reactRouterManifest = {
|
|
"entry": {
|
|
"module": "/assets/entry.client-3vubyXrk.js",
|
|
"imports": [
|
|
"/assets/jsx-runtime-DwfQwkRq.js",
|
|
"/assets/components-_bNmAApg.js",
|
|
"/assets/_.well-known_.mcp.server-card_.json_-Sx7XeH3e.js",
|
|
"/assets/index-mKTXLmHu.js",
|
|
"/assets/errorBoundaries-DhW4jVYt.js"
|
|
],
|
|
"css": []
|
|
},
|
|
"routes": {
|
|
"root": {
|
|
"id": "root",
|
|
"path": "",
|
|
"hasAction": false,
|
|
"hasLoader": true,
|
|
"hasClientAction": false,
|
|
"hasClientLoader": false,
|
|
"hasClientMiddleware": false,
|
|
"hasDefaultExport": true,
|
|
"hasErrorBoundary": true,
|
|
"module": "/assets/root-DbOv4-98.js",
|
|
"imports": [
|
|
"/assets/jsx-runtime-DwfQwkRq.js",
|
|
"/assets/components-_bNmAApg.js",
|
|
"/assets/_.well-known_.mcp.server-card_.json_-Sx7XeH3e.js",
|
|
"/assets/index-mKTXLmHu.js",
|
|
"/assets/errorBoundaries-DhW4jVYt.js",
|
|
"/assets/lib-Dg89tQ22.js",
|
|
"/assets/analytics.client-DM6E8o1h.js",
|
|
"/assets/SiteHeader-C2U5gvDH.js",
|
|
"/assets/current-9yDxj94E.js",
|
|
"/assets/clsx-eT0YPcGk.js",
|
|
"/assets/bugs-38ilEoW0.js",
|
|
"/assets/keyboard-D-uXZORL.js",
|
|
"/assets/use-tab-direction-dKm-S3Ck.js"
|
|
],
|
|
"css": []
|
|
},
|
|
"routes/blog": {
|
|
"id": "routes/blog",
|
|
"parentId": "root",
|
|
"path": "blog",
|
|
"hasAction": false,
|
|
"hasLoader": false,
|
|
"hasClientAction": false,
|
|
"hasClientLoader": false,
|
|
"hasClientMiddleware": false,
|
|
"hasDefaultExport": false,
|
|
"hasErrorBoundary": false,
|
|
"module": "/assets/blog-pXH7ptHJ.js",
|
|
"imports": [
|
|
"/assets/_.well-known_.mcp.server-card_.json_-Sx7XeH3e.js"
|
|
],
|
|
"css": []
|
|
},
|
|
"routes/blog.$slug": {
|
|
"id": "routes/blog.$slug",
|
|
"parentId": "routes/blog",
|
|
"path": ":slug",
|
|
"hasAction": false,
|
|
"hasLoader": true,
|
|
"hasClientAction": false,
|
|
"hasClientLoader": false,
|
|
"hasClientMiddleware": false,
|
|
"hasDefaultExport": true,
|
|
"hasErrorBoundary": false,
|
|
"module": "/assets/blog._slug-Ch_92qsH.js",
|
|
"imports": [
|
|
"/assets/components-_bNmAApg.js",
|
|
"/assets/lib-Dg89tQ22.js",
|
|
"/assets/jsx-runtime-DwfQwkRq.js",
|
|
"/assets/ContentImage-Dh6VEOUl.js",
|
|
"/assets/clsx-eT0YPcGk.js",
|
|
"/assets/BlogCategoryLink-DmQyn0gp.js",
|
|
"/assets/_.well-known_.mcp.server-card_.json_-Sx7XeH3e.js",
|
|
"/assets/Details-BSB_b6hI.js",
|
|
"/assets/Skittle-CDFOPRjH.js",
|
|
"/assets/SiteFooter-B2Gq9u2j.js",
|
|
"/assets/SiteHeader-C2U5gvDH.js",
|
|
"/assets/Vimeo-00PQJDli.js",
|
|
"/assets/YouTube-CMfaljVr.js",
|
|
"/assets/date-CJTFH3uT.js",
|
|
"/assets/errorBoundaries-DhW4jVYt.js",
|
|
"/assets/keyboard-D-uXZORL.js",
|
|
"/assets/use-tab-direction-dKm-S3Ck.js",
|
|
"/assets/index-mKTXLmHu.js",
|
|
"/assets/use-inert-others-BMJ6-xOX.js",
|
|
"/assets/description-Cf6FZmDe.js",
|
|
"/assets/use-is-mounted-uQsUZyP9.js",
|
|
"/assets/types-DvonrUFF.js",
|
|
"/assets/current-9yDxj94E.js",
|
|
"/assets/analytics.client-DM6E8o1h.js",
|
|
"/assets/bugs-38ilEoW0.js"
|
|
],
|
|
"css": []
|
|
},
|
|
"routes/_index": {
|
|
"id": "routes/_index",
|
|
"parentId": "root",
|
|
"index": true,
|
|
"hasAction": false,
|
|
"hasLoader": true,
|
|
"hasClientAction": false,
|
|
"hasClientLoader": false,
|
|
"hasClientMiddleware": false,
|
|
"hasDefaultExport": true,
|
|
"hasErrorBoundary": false,
|
|
"module": "/assets/_index-BfA6EnlR.js",
|
|
"imports": [
|
|
"/assets/components-_bNmAApg.js",
|
|
"/assets/lib-Dg89tQ22.js",
|
|
"/assets/jsx-runtime-DwfQwkRq.js",
|
|
"/assets/Logo-Gm9TLYAs.js",
|
|
"/assets/SiteFooter-B2Gq9u2j.js",
|
|
"/assets/SiteHeader-C2U5gvDH.js",
|
|
"/assets/_.well-known_.mcp.server-card_.json_-Sx7XeH3e.js",
|
|
"/assets/bugs-38ilEoW0.js",
|
|
"/assets/keyboard-D-uXZORL.js",
|
|
"/assets/use-is-mounted-uQsUZyP9.js",
|
|
"/assets/use-tab-direction-dKm-S3Ck.js",
|
|
"/assets/errorBoundaries-DhW4jVYt.js",
|
|
"/assets/clsx-eT0YPcGk.js",
|
|
"/assets/current-9yDxj94E.js",
|
|
"/assets/analytics.client-DM6E8o1h.js",
|
|
"/assets/index-mKTXLmHu.js"
|
|
],
|
|
"css": []
|
|
},
|
|
"routes/blog._index": {
|
|
"id": "routes/blog._index",
|
|
"parentId": "routes/blog",
|
|
"index": true,
|
|
"hasAction": false,
|
|
"hasLoader": true,
|
|
"hasClientAction": false,
|
|
"hasClientLoader": false,
|
|
"hasClientMiddleware": false,
|
|
"hasDefaultExport": true,
|
|
"hasErrorBoundary": false,
|
|
"module": "/assets/blog._index-DcvTTuDd.js",
|
|
"imports": [
|
|
"/assets/components-_bNmAApg.js",
|
|
"/assets/jsx-runtime-DwfQwkRq.js",
|
|
"/assets/social-Cd2AtOZM.js",
|
|
"/assets/BlogCategoryLink-DmQyn0gp.js",
|
|
"/assets/BlogPostLink-DC1SPKBJ.js",
|
|
"/assets/BlogCategoryNav-CB7TJ3IB.js",
|
|
"/assets/Paginator-xlPA_JNt.js",
|
|
"/assets/SiteFooter-B2Gq9u2j.js",
|
|
"/assets/SiteHeader-C2U5gvDH.js",
|
|
"/assets/date-CJTFH3uT.js",
|
|
"/assets/_.well-known_.mcp.server-card_.json_-Sx7XeH3e.js",
|
|
"/assets/lib-Dg89tQ22.js",
|
|
"/assets/errorBoundaries-DhW4jVYt.js",
|
|
"/assets/clsx-eT0YPcGk.js",
|
|
"/assets/types-DvonrUFF.js",
|
|
"/assets/enumerator-2YLGh-nT.js",
|
|
"/assets/current-9yDxj94E.js",
|
|
"/assets/analytics.client-DM6E8o1h.js",
|
|
"/assets/bugs-38ilEoW0.js",
|
|
"/assets/keyboard-D-uXZORL.js",
|
|
"/assets/use-tab-direction-dKm-S3Ck.js",
|
|
"/assets/index-mKTXLmHu.js"
|
|
],
|
|
"css": []
|
|
}
|
|
},
|
|
"url": "/assets/manifest-e17deb94.js",
|
|
"version": "e17deb94"
|
|
};
|
|
window.__reactRouterRouteModules = {"root":route0,"routes/blog":route1,"routes/blog.$slug":route2};
|
|
|
|
import("/assets/entry.client-3vubyXrk.js");</script><script type="application/ld+json" nonce="nyDuN2IITQLGs2NvcHjdoIBrtda9b+1NY1U/+JH3x2E=">{"@context":"https://schema.org","@type":"Organization","name":"PlanetScale, Inc.","url":"https://planetscale.com","sameAs":["https://twitter.com/PlanetScale","https://www.facebook.com/planetscaledata/","https://www.instagram.com/planetscale/"],"address":{"@type":"PostalAddress","streetAddress":"WeWork c/o PlanetScale, 535 Mission Street, 14th Floor","addressLocality":"San Francisco","addressRegion":"CA","postalCode":"94105","addressCountry":"US"}}</script><!--$--><script nonce="nyDuN2IITQLGs2NvcHjdoIBrtda9b+1NY1U/+JH3x2E=">window.__reactRouterContext.streamController.enqueue("[{\"_1\":2,\"_3\":-5,\"_4\":-5},\"loaderData\",{\"_5\":6,\"_7\":8},\"actionData\",\"errors\",\"root\",{\"_560\":561},\"routes/blog.$slug\",{\"_9\":10,\"_5\":11},\"blog\",{\"_12\":13,\"_14\":15,\"_16\":-7,\"_17\":18,\"_19\":20,\"_21\":22,\"_23\":24,\"_25\":26,\"_27\":28,\"_29\":30,\"_31\":32},\"https://planetscale.com\",\"body\",[88,89,90,91,92,93,94,95,96,97,98,99,100,101,102,103,104,105,106,107,108,109,110,111,112,113,114,115,116,117,118,119,120,121,122,123,124,125,126,127,128,129,130,131,132,133,134,135],\"body_text\",\"A throttler is a service or component that pushes back against an incoming flow of requests to ensure the system has the capacity to handle all received requests without being overwhelmed. In this series of posts, we illustrate design considerations for a database system throttler, whose purpose is to keep the database system healthy overall. We discuss choice of metrics, granularity, behavior, impact, prioritization, and other topics.\\nWhich requests do you throttle?\\nThere are different approaches to throttling requests in a database. We focus on throttling asynchronous, batch, and massive operations that are not time critical. Examples could be ETLs, data imports, online DDL operations, mass purges of data, resharding, and so forth. The throttler will push back on those operations that can span minutes, hours, or days of operation. Other forms of throttling may push back on OLTP production traffic. This discussion applies equally to both.\\nBy way of illustration, consider a job that needs to import 10 million rows into the database. Instead of attempting to apply all 10 million in one go, the job breaks down the task into much smaller subtasks: it will try to import (write) 100 rows at a time. Before any such import, it will request access from the throttler.\\nSome throttler implementations are collaborative, meaning they assume clients will respect their instructions. Others act as barriers between the app and the database. Either way, if the throttler indicates that the database is overloaded, the job should hold back for a period of time and then request access again. This process repeats until granted. Each subtask should be small enough so as not to single-handedly tank the database's serving capacity, while large enough to compensate for the added throttler overhead and to enable meaningful progress.\\nWhat does the throttler throttle on?\\nSome generic throttlers only allow a regulated rate of requests, in anticipation that the consuming job will be able to process them at some known, fixed rate. With databases, things are less clear. A database can only handle so many queries at any given point in time, or over some period of time. However, not all queries are created equal. The database capacity of serving queries depends on the scope of queries, any hot spots or cold spots in affected data, the state of the page cache, overlap or lack thereof of data served by queries, to name a few factors.\\nWe therefore need to be able to determine: how do we consider our database to be \\\"healthy\\\"? How do we determine if it is being overwhelmed?\\nTo do this, we look for metrics that define or predict service level objectives (SLO) for the database. But things are not always so simple. Let's start with a popular metric which is widely used as a throttling indicator, and see what's so special about that metric.\\nReplication lag\\nReplication is often used in database clusters, especially ones using a primary-secondary (aka leader-follower) architecture. Sometimes the replication is set to be asynchronous, and at other times group-communication based. In these scenarios replication lag is defined as the time passing between a write on the primary server and the time it is applied or made visible on the replica/secondary server.\\nIn the MySQL world, replication lag is probably the single most used throttling indicator, as multiple third-party and community tools use it to push back against long-running jobs. This is for good reasons: it is easy to measure and it has clear impact on the product and the business. For example, in the case of a database failover, replication lag impacts the time it takes for a replica/standby server to be promoted and made available to receive write requests. Read-after-write can be simplified when replication lag is low, allowing secondary servers to serve some of the read traffic.\\nWe can thus have a business limitation on the acceptable replication lag, which we can use in the throttler: below this lag, allow requests. Above this lag, push back.\\nOther database metrics: what's in a metric?\\nAnother common metric in the MySQL world is the value of threads_running. On any given server, this is the number of concurrent, actively executing queries (not to be confused with concurrent open transactions, some of which could be idle in between queries). This metric is frequently seen on database dashboards, and is an indicator for the database load.\\nBut, what's an acceptable value? Are 50 concurrent queries OK? Are 100 OK? Pick a number, and you'll soon find it doesn't hold water. Some values are acceptable in early morning, while others are just normal during peak traffic hours. As your product evolves and its adoption increases, so do the queries on your database. What was true 3 months ago is not true today. And again, not all queries are created equal.\\nWhat's different about this metric compared with replication lag is that it is much more of a symptom than an actual cause. If all of a sudden we see a sharp spike in active queries, this can indicate some possible causes: perhaps all are held by the commit queue, which for some reason stalls. Or, the queries happen to compete over a specific hotspot and wait on locks. Or, they don't, they all happen to compete on very different pages, none of which is in memory, and they all congest while waiting on the page cache, etc. So what is it exactly that we need to monitor? Is the metric itself useless?\\nNot necessarily. An experienced administrator may only need to take one look at this metric on the database dashboard to say \\\"we're having an issue\\\".\\nA closer examination: queues\\nSoftware relies heavily on queues. They are fundamental not only in software design but also in hardware access. Requests queue on network access. They queue on disk access. They queue on CPU access. They queue on locks.\\nCircling back to replication lag, much like concurrent queries, it is a symptom. E.g. disk I/O is saturated on the replica, hence the replica cannot keep up replaying the changelog, thereby accumulating lag. Or perhaps the lag is caused by slow network. Or both! Whatever the case is, what's interesting is that the replication mechanism itself is a queue: the changelog event queue. A new write on the primary manifests as a \\\"write event\\\", which is shipped to the replica, and waits to be consumed (processed, replayed) by the replica. Replication lag is the event's time spent in the queue, where our queue is a combination of the network queue, local disk write queue, actual wait time, and finally the event's processing time. Each of these can be the major contributor to the overall replication lag, and yet, we can still look at replication lag as a whole — as a clear indicator for database health.\\nArmed with this insight, we take another attempt at understanding other metrics. In the case for concurrent writes, we understand a major contributor to a spike in concurrent queries is their inability to complete. Normally, this means they're held back at commit time, i.e. they wait to be written to the transaction (redo) log. And that means they're in the transaction queue, and we can hence measure the transaction queue latency (aka queue delay).\\nBut, what's a good threshold? Transaction commit delay is typically caused by disk write/flush time, and that changes dramatically across hardware. It is a matter of knowing your metrics. Yet again, an experienced administrator should know what values to expect. But now the values are more tightly bound to hardware and slightly less affected by the app.\\nQueue delay is not the only metric. Another common one is the queue length: the number of entries waiting in the queue. A long queue at the airport isn't in itself a bad thing, some queues move quite fast, and yet it's often a predictor to wait times. Where wait time is impossible or difficult to measure, queue length can be an alternative.\\nAn operating system's Load Average metric evaluation includes the number of processes waiting for CPU time. This changes by the number of CPUs available. A common rough indicator is a 1 threshold for (load average)/(num CPUs). This is again a metric that must agree with your own systems. Some database deployments famously push their servers to their limits with load averages soaring far above 1 per CPU.\\nPool usage\\nAnother indicator is pool usage. The single most common pool with regard to databases must be the application's database connection pool. To run a query, the app will take a connection from the pool, use it to execute the query, then return the connection to the pool. If the pool has connections to spare, getting that connection comes at no cost. But if the pool is exhausted, then either the app needs to wait for the creation of a new connection, or it gets rejected. Similarly to concurrent queries, a high pool usage indicates a congestion of operations. However, pooled connections can be used across multiple queries in a transaction, as well as across multiple transactions, and the app may run its own logic in between running queries, while still holding on to the connection.\\nAn exhausted pool is a strong indication of excessive load, while the difference between a 60% and an 80% used pool is not as clear an indication. Taking a step back, what does it mean that we exhaust some pool? Who decides the size of the pool in the first place? If someone picked a number such as 50 or 100, isn't that number just artificial?\\nIt may well be, but pool size was likely chosen for some good reason(s). It is perhaps derived from some database configuration, which is itself derived from some hardware limitation. And while the choice of metric could possibly change arbitrarily, it is still sensible, as far as throttling goes, to push back when the pool is exhausted. The throttler thereby relies on the greater system configuration and does not introduce any new artificial thresholds.\\nThe case for multiple metrics\\nA throttler should be able to push back based on a combination of metrics, and not limit itself to just one metric. We've illustrated some metrics above, and every environment may yet have its own load predicting metrics. The administrator should be able to choose an assorted set of metrics the throttler should work with, be able to set specific thresholds for each such metric, and possibly be able to introduce new metrics either programmatically or dynamically.\\nWhat does a throttled system look like?\\nMany software developers will be familiar with the next scenario: you have a multithreaded or otherwise highly concurrent app. There's a bug, likely a race condition or a synchronization issue, and you wish to find it. You choose to print informative debug messages to standard output, and hope to find the bug by examining the log. Alas, when you do so, the bug does not reproduce. Or it may manifest elsewhere.\\nIn adding writes to standard output, you have introduced new locks. Your debug messages now compete over those locks, which in turn incurs different context switches.\\nIntroducing a throttler into your infrastructure shows resemblances to this synchronization example. All of a sudden, there is less contention on the database, and certain apps that used to run just fine, exhibit contention/latency behavior. The appearance of a new job suddenly affects the progress of another. But where previously you could clearly analyze database queries to find the root cause, the database now tells you little to nothing. It's now down to the throttler to give you that information. But even the throttler is limited, because all the apps do is to check the throttler for health status. They do not yet actually do anything.\\nLet's say we throttle based on replication lag, and let's assume that we want to run an operation so massive that it is bound to drive replication lag high if let loose. With the throttler keeping it under control, though, the operation will only run small batches of subtasks. But an interesting behavior emerges: the operation will push replication lag up to the throttler's threshold, then back down, and push again. As we start the operation, we expect to see the replication lag graph jump up to the threshold value, and then more or less stabilize around that value, slightly higher and slightly lower, for the duration of the operation, which could be hours.\\nDuring that time, the operation will be granted access thousands of times or more, and will likewise also be rejected access thousands of times or more. That is how a healthy system looks with a throttler engaged. No matter how many more concurrent operations we run, we expect to contain replication lag at about the same slight offset above or below the threshold. More on this when we discuss granularity.\\nIt is not uncommon for a system to run one or two operations for very long periods, which means what we consider as the throttling threshold (say, a 5sec replication lag) becomes the actual standard. Thankfully, not all operations and workloads are so aggressive that they necessarily push the metrics as high as their thresholds.\\nCheck intervals and metric granularity\\nA throttler collects the metrics asynchronously from check requests, so that it has an immediate answer available upon request. The intervals at which the throttler collects metrics can have a significant effect on how the throttler is being put to use. Let's consider a case where the throttler collects a metric at a large interval, say every 5 seconds. The metric could be anything at all during those 5 seconds, but it is the specific sampling that takes place at the end of that period that counts.\\nSimilarly, other metrics could have some granularity. Namely, replication lag can be measured in different methods, and the most common one is by deliberate injection of heartbeat events on the primary, and by capturing them on a replica. More on this in another post, but the intervals in which the heartbeat events are generated dictate the granularity or the accuracy of the measured lag. Let's assume we inject heartbeats at one second intervals, and we've just injected a heartbeat at precisely noon. Let's also assume we sample the metrics once per second, and we happen to make that sample at 12:00:00.995. The sample still reads 12:00:00.000 as this was our last injected metric. A client then checks the throttler at 12:00:01.990. By now there will have been a new metric value, but one which we have not sampled yet. The throttler responds by using its last sample that is almost, but not quite, one second old, and which in itself represents a metric that is now almost, but not quite, two seconds old.\\nLong heartbeat intervals and outdated information have negative impacts on both our system health as well as the throttler's utilization.\\nOn one hand, it is possible that in the duration of the interval, we miss noticing a significant uptick in system load. We'd only find out about it a few seconds later, at which time the throttler would be engaged. However, by that time, the system performance already degrades. It will take a few seconds before it comes down to acceptable values. But then again, once we do catch that the metrics exceed their thresholds, and for the duration of the next interval, we reject all further requests. If the metrics do turn healthy sooner than that, that's a missed opportunity to make some progress. Thus, we degrade the database's operations capacity.\\nWhen multiple operations attempt to make progress all at once, all will be throttled while metrics are above threshold, and possibly all released at once when metrics return to low values, thus all pushing the metrics up at once.\\nBorrowing from the world of networking hardware, it is recommended that metric interval and granularity oversample the range of allowed thresholds. For example, if the acceptable replication lag is at 5 seconds, then it's best to have a heartbeat/sampling interval of 1-2 seconds.\\nLower intervals and more accurate metrics reduce spikes and spread the workload more efficiently. That, too, comes at a cost, which we will discuss in a later post.\\nTo be continued\\nIn the next part of this series we will be looking into singular vs. distributed throttler design, as well as the impact the throttler itself may have on your environment.\",\"aside\",\"toc\",[43,44,45,46,47],\"title\",\"Anatomy of a Throttler, part 1\",\"authors\",[39],\"categories\",[38],\"excerpt\",\"Learn about some design considerations for implementing a database throttler.\",\"createdAt\",\"2024-08-29\",\"slug\",\"anatomy-of-a-throttler-part-1\",\"meta\",{\"_33\":34,\"_35\":26,\"_36\":37,\"_19\":20},\"canonical\",\"https://planetscale.com/blog/anatomy-of-a-throttler-part-1\",\"description\",\"image\",\"/assets/anatomy-of-a-throttler-part-1-social-C0hGHHNN.jpg\",\"engineering\",{\"_29\":40,\"_41\":42},\"shlomi\",\"name\",\"Shlomi Noach\",{\"_48\":85,\"_50\":86,\"_52\":53,\"_19\":87},{\"_48\":61,\"_50\":62,\"_52\":53,\"_19\":63},{\"_48\":58,\"_50\":59,\"_52\":53,\"_19\":60},{\"_48\":55,\"_50\":56,\"_52\":53,\"_19\":57},{\"_48\":49,\"_50\":51,\"_52\":53,\"_19\":54},\"children\",[],\"id\",\"to-be-continued\",\"level\",2,\"To be continued\",[],\"check-intervals-and-metric-granularity\",\"Check intervals and metric granularity\",[],\"what-does-a-throttled-system-look-like\",\"What does a throttled system look like?\",[64,65,66,67,68],\"what-does-the-throttler-throttle-on\",\"What does the throttler throttle on?\",{\"_48\":82,\"_50\":83,\"_52\":71,\"_19\":84},{\"_48\":79,\"_50\":80,\"_52\":71,\"_19\":81},{\"_48\":76,\"_50\":77,\"_52\":71,\"_19\":78},{\"_48\":73,\"_50\":74,\"_52\":71,\"_19\":75},{\"_48\":69,\"_50\":70,\"_52\":71,\"_19\":72},[],\"the-case-for-multiple-metrics\",3,\"The case for multiple metrics\",[],\"pool-usage\",\"Pool usage\",[],\"a-closer-examination-queues\",\"A closer examination: queues\",[],\"other-database-metrics-whats-in-a-metric\",\"Other database metrics: what's in a metric?\",[],\"replication-lag\",\"Replication lag\",[],\"which-requests-do-you-throttle\",\"Which requests do you throttle?\",[\"SingleFetchClassInstance\",549],[\"SingleFetchClassInstance\",541],[\"SingleFetchClassInstance\",537],[\"SingleFetchClassInstance\",533],[\"SingleFetchClassInstance\",529],[\"SingleFetchClassInstance\",521],[\"SingleFetchClassInstance\",517],[\"SingleFetchClassInstance\",513],[\"SingleFetchClassInstance\",509],[\"SingleFetchClassInstance\",501],[\"SingleFetchClassInstance\",483],[\"SingleFetchClassInstance\",472],[\"SingleFetchClassInstance\",468],[\"SingleFetchClassInstance\",460],[\"SingleFetchClassInstance\",450],[\"SingleFetchClassInstance\",436],[\"SingleFetchClassInstance\",421],[\"SingleFetchClassInstance\",417],[\"SingleFetchClassInstance\",409],[\"SingleFetchClassInstance\",398],[\"SingleFetchClassInstance\",388],[\"SingleFetchClassInstance\",366],[\"SingleFetchClassInstance\",355],[\"SingleFetchClassInstance\",351],[\"SingleFetchClassInstance\",324],[\"SingleFetchClassInstance\",316],[\"SingleFetchClassInstance\",312],[\"SingleFetchClassInstance\",284],[\"SingleFetchClassInstance\",280],[\"SingleFetchClassInstance\",271],[\"SingleFetchClassInstance\",267],[\"SingleFetchClassInstance\",259],[\"SingleFetchClassInstance\",255],[\"SingleFetchClassInstance\",251],[\"SingleFetchClassInstance\",241],[\"SingleFetchClassInstance\",231],[\"SingleFetchClassInstance\",227],[\"SingleFetchClassInstance\",217],[\"SingleFetchClassInstance\",209],[\"SingleFetchClassInstance\",205],[\"SingleFetchClassInstance\",175],[\"SingleFetchClassInstance\",171],[\"SingleFetchClassInstance\",167],[\"SingleFetchClassInstance\",163],[\"SingleFetchClassInstance\",159],[\"SingleFetchClassInstance\",155],[\"SingleFetchClassInstance\",144],[\"SingleFetchClassInstance\",136],{\"_137\":138,\"_41\":139,\"_140\":141,\"_48\":142},\"$$mdtype\",\"Tag\",\"p\",\"attributes\",{},[143],\"In the next part of this series we will be looking into singular vs. distributed throttler design, as well as the impact the throttler itself may have on your environment.\",{\"_137\":138,\"_41\":145,\"_140\":146,\"_48\":147},\"h2\",{\"_50\":51},[148],[\"SingleFetchClassInstance\",149],{\"_137\":138,\"_41\":150,\"_140\":151,\"_48\":152},\"a\",{\"_153\":154},[54],\"href\",\"#to-be-continued\",{\"_137\":138,\"_41\":139,\"_140\":156,\"_48\":157},{},[158],\"Lower intervals and more accurate metrics reduce spikes and spread the workload more efficiently. That, too, comes at a cost, which we will discuss in a later post.\",{\"_137\":138,\"_41\":139,\"_140\":160,\"_48\":161},{},[162],\"Borrowing from the world of networking hardware, it is recommended that metric interval and granularity oversample the range of allowed thresholds. For example, if the acceptable replication lag is at 5 seconds, then it's best to have a heartbeat/sampling interval of 1-2 seconds.\",{\"_137\":138,\"_41\":139,\"_140\":164,\"_48\":165},{},[166],\"When multiple operations attempt to make progress all at once, all will be throttled while metrics are above threshold, and possibly all released at once when metrics return to low values, thus all pushing the metrics up at once.\",{\"_137\":138,\"_41\":139,\"_140\":168,\"_48\":169},{},[170],\"On one hand, it is possible that in the duration of the interval, we miss noticing a significant uptick in system load. We'd only find out about it a few seconds later, at which time the throttler would be engaged. However, by that time, the system performance already degrades. It will take a few seconds before it comes down to acceptable values. But then again, once we do catch that the metrics exceed their thresholds, and for the duration of the next interval, we reject all further requests. If the metrics do turn healthy sooner than that, that's a missed opportunity to make some progress. Thus, we degrade the database's operations capacity.\",{\"_137\":138,\"_41\":139,\"_140\":172,\"_48\":173},{},[174],\"Long heartbeat intervals and outdated information have negative impacts on both our system health as well as the throttler's utilization.\",{\"_137\":138,\"_41\":139,\"_140\":176,\"_48\":177},{},[178,179,180,181,182,183,184,185,186],\"Similarly, other metrics could have some granularity. Namely, replication lag can be measured in different methods, and the most common one is by deliberate injection of heartbeat events on the primary, and by capturing them on a replica. More on this in another post, but the intervals in which the heartbeat events are generated dictate the granularity or the accuracy of the measured lag. Let's assume we inject heartbeats at one second intervals, and we've just injected a heartbeat at precisely noon. Let's also assume we \",[\"SingleFetchClassInstance\",200],\" the metrics once per second, and we happen to make that sample at \",[\"SingleFetchClassInstance\",196],\". The sample still reads \",[\"SingleFetchClassInstance\",192],\" as this was our last injected metric. A client then checks the throttler at \",[\"SingleFetchClassInstance\",187],\". By now there will have been a new metric value, but one which we have not sampled yet. The throttler responds by using its last sample that is almost, but not quite, one second old, and which in itself represents a metric that is now almost, but not quite, two seconds old.\",{\"_137\":138,\"_41\":188,\"_140\":189,\"_48\":190},\"code\",{},[191],\"12:00:01.990\",{\"_137\":138,\"_41\":188,\"_140\":193,\"_48\":194},{},[195],\"12:00:00.000\",{\"_137\":138,\"_41\":188,\"_140\":197,\"_48\":198},{},[199],\"12:00:00.995\",{\"_137\":138,\"_41\":201,\"_140\":202,\"_48\":203},\"em\",{},[204],\"sample\",{\"_137\":138,\"_41\":139,\"_140\":206,\"_48\":207},{},[208],\"A throttler collects the metrics asynchronously from check requests, so that it has an immediate answer available upon request. The intervals at which the throttler collects metrics can have a significant effect on how the throttler is being put to use. Let's consider a case where the throttler collects a metric at a large interval, say every 5 seconds. The metric could be anything at all during those 5 seconds, but it is the specific sampling that takes place at the end of that period that counts.\",{\"_137\":138,\"_41\":145,\"_140\":210,\"_48\":211},{\"_50\":56},[212],[\"SingleFetchClassInstance\",213],{\"_137\":138,\"_41\":150,\"_140\":214,\"_48\":215},{\"_153\":216},[57],\"#check-intervals-and-metric-granularity\",{\"_137\":138,\"_41\":139,\"_140\":218,\"_48\":219},{},[220,221,222],\"It is not uncommon for a system to run one or two operations for very long periods, which means what we consider as the throttling threshold (say, a \",[\"SingleFetchClassInstance\",223],\" replication lag) becomes the actual standard. Thankfully, not all operations and workloads are so aggressive that they necessarily push the metrics as high as their thresholds.\",{\"_137\":138,\"_41\":188,\"_140\":224,\"_48\":225},{},[226],\"5sec\",{\"_137\":138,\"_41\":139,\"_140\":228,\"_48\":229},{},[230],\"During that time, the operation will be granted access thousands of times or more, and will likewise also be rejected access thousands of times or more. That is how a healthy system looks with a throttler engaged. No matter how many more concurrent operations we run, we expect to contain replication lag at about the same slight offset above or below the threshold. More on this when we discuss granularity.\",{\"_137\":138,\"_41\":139,\"_140\":232,\"_48\":233},{},[234,235,236],\"Let's say we throttle based on replication lag, and let's assume that we want to run an operation so massive that it is \",[\"SingleFetchClassInstance\",237],\" to drive replication lag high if let loose. With the throttler keeping it under control, though, the operation will only run small batches of subtasks. But an interesting behavior emerges: the operation will push replication lag up to the throttler's threshold, then back down, and push again. As we start the operation, we expect to see the replication lag graph jump up to the threshold value, and then more or less stabilize around that value, slightly higher and slightly lower, for the duration of the operation, which could be hours.\",{\"_137\":138,\"_41\":201,\"_140\":238,\"_48\":239},{},[240],\"bound\",{\"_137\":138,\"_41\":139,\"_140\":242,\"_48\":243},{},[244,245,246],\"Introducing a throttler into your infrastructure shows resemblances to this synchronization example. All of a sudden, there is less contention on the database, and certain apps that used to run just fine, exhibit contention/latency behavior. The appearance of a new job suddenly affects the progress of another. But where previously you could clearly analyze database queries to find the root cause, the database now tells you little to nothing. It's now down to the throttler to give you that information. But even the throttler is limited, because all the apps do is to check the throttler for health status. They do not yet actually \",[\"SingleFetchClassInstance\",247],\" anything.\",{\"_137\":138,\"_41\":201,\"_140\":248,\"_48\":249},{},[250],\"do\",{\"_137\":138,\"_41\":139,\"_140\":252,\"_48\":253},{},[254],\"In adding writes to standard output, you have introduced new locks. Your debug messages now compete over those locks, which in turn incurs different context switches.\",{\"_137\":138,\"_41\":139,\"_140\":256,\"_48\":257},{},[258],\"Many software developers will be familiar with the next scenario: you have a multithreaded or otherwise highly concurrent app. There's a bug, likely a race condition or a synchronization issue, and you wish to find it. You choose to print informative debug messages to standard output, and hope to find the bug by examining the log. Alas, when you do so, the bug does not reproduce. Or it may manifest elsewhere.\",{\"_137\":138,\"_41\":145,\"_140\":260,\"_48\":261},{\"_50\":59},[262],[\"SingleFetchClassInstance\",263],{\"_137\":138,\"_41\":150,\"_140\":264,\"_48\":265},{\"_153\":266},[60],\"#what-does-a-throttled-system-look-like\",{\"_137\":138,\"_41\":139,\"_140\":268,\"_48\":269},{},[270],\"A throttler should be able to push back based on a combination of metrics, and not limit itself to just one metric. We've illustrated some metrics above, and every environment may yet have its own load predicting metrics. The administrator should be able to choose an assorted set of metrics the throttler should work with, be able to set specific thresholds for each such metric, and possibly be able to introduce new metrics either programmatically or dynamically.\",{\"_137\":138,\"_41\":272,\"_140\":273,\"_48\":274},\"h3\",{\"_50\":70},[275],[\"SingleFetchClassInstance\",276],{\"_137\":138,\"_41\":150,\"_140\":277,\"_48\":278},{\"_153\":279},[72],\"#the-case-for-multiple-metrics\",{\"_137\":138,\"_41\":139,\"_140\":281,\"_48\":282},{},[283],\"It may well be, but pool size was likely chosen for some good reason(s). It is perhaps derived from some database configuration, which is itself derived from some hardware limitation. And while the choice of metric could possibly change arbitrarily, it is still sensible, as far as throttling goes, to push back when the pool is exhausted. The throttler thereby relies on the greater system configuration and does not introduce any new artificial thresholds.\",{\"_137\":138,\"_41\":139,\"_140\":285,\"_48\":286},{},[287,288,289,290,291,292,293,294,295],\"An exhausted pool is a strong indication of excessive load, while the difference between a \",[\"SingleFetchClassInstance\",308],\" and an \",[\"SingleFetchClassInstance\",304],\" used pool is not as clear an indication. Taking a step back, what does it mean that we exhaust some pool? Who decides the size of the pool in the first place? If someone picked a number such as \",[\"SingleFetchClassInstance\",300],\" or \",[\"SingleFetchClassInstance\",296],\", isn't that number just artificial?\",{\"_137\":138,\"_41\":188,\"_140\":297,\"_48\":298},{},[299],\"100\",{\"_137\":138,\"_41\":188,\"_140\":301,\"_48\":302},{},[303],\"50\",{\"_137\":138,\"_41\":188,\"_140\":305,\"_48\":306},{},[307],\"80%\",{\"_137\":138,\"_41\":188,\"_140\":309,\"_48\":310},{},[311],\"60%\",{\"_137\":138,\"_41\":139,\"_140\":313,\"_48\":314},{},[315],\"Another indicator is pool usage. The single most common pool with regard to databases must be the application's database connection pool. To run a query, the app will take a connection from the pool, use it to execute the query, then return the connection to the pool. If the pool has connections to spare, getting that connection comes at no cost. But if the pool is exhausted, then either the app needs to wait for the creation of a new connection, or it gets rejected. Similarly to concurrent queries, a high pool usage indicates a congestion of operations. However, pooled connections can be used across multiple queries in a transaction, as well as across multiple transactions, and the app may run its own logic in between running queries, while still holding on to the connection.\",{\"_137\":138,\"_41\":272,\"_140\":317,\"_48\":318},{\"_50\":74},[319],[\"SingleFetchClassInstance\",320],{\"_137\":138,\"_41\":150,\"_140\":321,\"_48\":322},{\"_153\":323},[75],\"#pool-usage\",{\"_137\":138,\"_41\":139,\"_140\":325,\"_48\":326},{},[327,328,329,330,331,332,333,334,335],\"An operating system's \",[\"SingleFetchClassInstance\",347],\" metric evaluation includes the number of processes waiting for CPU time. This changes by the number of CPUs available. A common rough indicator is a \",[\"SingleFetchClassInstance\",344],\" threshold for \",[\"SingleFetchClassInstance\",340],\". This is again a metric that must agree with your own systems. Some database deployments famously push their servers to their limits with load averages soaring far above \",[\"SingleFetchClassInstance\",336],\" per CPU.\",{\"_137\":138,\"_41\":188,\"_140\":337,\"_48\":338},{},[339],\"1\",{\"_137\":138,\"_41\":188,\"_140\":341,\"_48\":342},{},[343],\"(load average)/(num CPUs)\",{\"_137\":138,\"_41\":188,\"_140\":345,\"_48\":346},{},[339],{\"_137\":138,\"_41\":201,\"_140\":348,\"_48\":349},{},[350],\"Load Average\",{\"_137\":138,\"_41\":139,\"_140\":352,\"_48\":353},{},[354],\"Queue delay is not the only metric. Another common one is the queue length: the number of entries waiting in the queue. A long queue at the airport isn't in itself a bad thing, some queues move quite fast, and yet it's often a predictor to wait times. Where wait time is impossible or difficult to measure, queue length can be an alternative.\",{\"_137\":138,\"_41\":139,\"_140\":356,\"_48\":357},{},[358,359,360],\"But, what's a good threshold? Transaction commit delay is typically caused by disk write/flush time, and that \",[\"SingleFetchClassInstance\",361],\" across hardware. It is a matter of knowing your metrics. Yet again, an experienced administrator should know what values to expect. But now the values are more tightly bound to hardware and slightly less affected by the app.\",{\"_137\":138,\"_41\":150,\"_140\":362,\"_48\":363},{\"_153\":365},[364],\"changes dramatically\",\"https://hackmysql.com/commit-latency-aurora-vs-rds-mysql-8.0/\",{\"_137\":138,\"_41\":139,\"_140\":367,\"_48\":368},{},[369,370,371,372,373,374,375],\"Armed with this insight, we take another attempt at understanding other metrics. In the case for concurrent writes, we understand a major contributor to a spike in concurrent queries is their inability to \",[\"SingleFetchClassInstance\",384],\". Normally, this means they're held back at \",[\"SingleFetchClassInstance\",380],\", i.e. they wait to be written to the transaction (redo) log. And that means they're in the transaction \",[\"SingleFetchClassInstance\",376],\", and we can hence measure the transaction queue latency (aka queue delay).\",{\"_137\":138,\"_41\":201,\"_140\":377,\"_48\":378},{},[379],\"queue\",{\"_137\":138,\"_41\":201,\"_140\":381,\"_48\":382},{},[383],\"commit time\",{\"_137\":138,\"_41\":201,\"_140\":385,\"_48\":386},{},[387],\"complete\",{\"_137\":138,\"_41\":139,\"_140\":389,\"_48\":390},{},[391,392,393],\"Circling back to replication lag, much like concurrent queries, it is a \",[\"SingleFetchClassInstance\",394],\". E.g. disk I/O is saturated on the replica, hence the replica cannot keep up replaying the changelog, thereby accumulating lag. Or perhaps the lag is caused by slow network. Or both! Whatever the case is, what's interesting is that the replication mechanism itself is a queue: the changelog event queue. A new write on the primary manifests as a \\\"write event\\\", which is shipped to the replica, and waits to be consumed (processed, replayed) by the replica. Replication lag is the event's time spent in the queue, where our queue is a combination of the network queue, local disk write queue, actual wait time, and finally the event's processing time. Each of these can be the major contributor to the overall replication lag, and yet, we can still look at replication lag as a whole — as a clear indicator for database health.\",{\"_137\":138,\"_41\":201,\"_140\":395,\"_48\":396},{},[397],\"symptom\",{\"_137\":138,\"_41\":139,\"_140\":399,\"_48\":400},{},[401,402,403],\"Software relies heavily on \",[\"SingleFetchClassInstance\",404],\". They are fundamental not only in software design but also in hardware access. Requests queue on network access. They queue on disk access. They queue on CPU access. They queue on locks.\",{\"_137\":138,\"_41\":150,\"_140\":405,\"_48\":406},{\"_153\":408},[407],\"queues\",\"https://en.wikipedia.org/wiki/Queueing_theory\",{\"_137\":138,\"_41\":272,\"_140\":410,\"_48\":411},{\"_50\":77},[412],[\"SingleFetchClassInstance\",413],{\"_137\":138,\"_41\":150,\"_140\":414,\"_48\":415},{\"_153\":416},[78],\"#a-closer-examination-queues\",{\"_137\":138,\"_41\":139,\"_140\":418,\"_48\":419},{},[420],\"Not necessarily. An experienced administrator may only need to take one look at this metric on the database dashboard to say \\\"we're having an issue\\\".\",{\"_137\":138,\"_41\":139,\"_140\":422,\"_48\":423},{},[424,425,426,427,428],\"What's different about this metric compared with replication lag is that it is much more of a \",[\"SingleFetchClassInstance\",433],\" than an actual \",[\"SingleFetchClassInstance\",429],\". If all of a sudden we see a sharp spike in active queries, this can indicate some possible causes: perhaps all are held by the commit queue, which for some reason stalls. Or, the queries happen to compete over a specific hotspot and wait on locks. Or, they don't, they all happen to compete on very different pages, none of which is in memory, and they all congest while waiting on the page cache, etc. So what is it exactly that we need to monitor? Is the metric itself useless?\",{\"_137\":138,\"_41\":201,\"_140\":430,\"_48\":431},{},[432],\"cause\",{\"_137\":138,\"_41\":201,\"_140\":434,\"_48\":435},{},[397],{\"_137\":138,\"_41\":139,\"_140\":437,\"_48\":438},{},[439,440,441,442,443],\"But, what's an acceptable value? Are \",[\"SingleFetchClassInstance\",447],\" concurrent queries OK? Are \",[\"SingleFetchClassInstance\",444],\" OK? Pick a number, and you'll soon find it doesn't hold water. Some values are acceptable in early morning, while others are just normal during peak traffic hours. As your product evolves and its adoption increases, so do the queries on your database. What was true 3 months ago is not true today. And again, not all queries are created equal.\",{\"_137\":138,\"_41\":188,\"_140\":445,\"_48\":446},{},[299],{\"_137\":138,\"_41\":188,\"_140\":448,\"_48\":449},{},[303],{\"_137\":138,\"_41\":139,\"_140\":451,\"_48\":452},{},[453,454,455],\"Another common metric in the MySQL world is the value of \",[\"SingleFetchClassInstance\",456],\". On any given server, this is the number of concurrent, actively executing queries (not to be confused with concurrent open transactions, some of which could be idle in between queries). This metric is frequently seen on database dashboards, and is an indicator for the database load.\",{\"_137\":138,\"_41\":188,\"_140\":457,\"_48\":458},{},[459],\"threads_running\",{\"_137\":138,\"_41\":272,\"_140\":461,\"_48\":462},{\"_50\":80},[463],[\"SingleFetchClassInstance\",464],{\"_137\":138,\"_41\":150,\"_140\":465,\"_48\":466},{\"_153\":467},[81],\"#other-database-metrics-whats-in-a-metric\",{\"_137\":138,\"_41\":139,\"_140\":469,\"_48\":470},{},[471],\"We can thus have a business limitation on the acceptable replication lag, which we can use in the throttler: below this lag, allow requests. Above this lag, push back.\",{\"_137\":138,\"_41\":139,\"_140\":473,\"_48\":474},{},[475,476,477],\"In the MySQL world, replication lag is probably the single most used throttling indicator, as multiple third-party and community tools use it to push back against long-running jobs. This is for good reasons: it is easy to measure and it has clear impact on the product and the business. For example, in the case of a database failover, replication lag impacts the time it takes for a replica/standby server to be promoted and made available to receive write requests. \",[\"SingleFetchClassInstance\",478],\" can be simplified when replication lag is low, allowing secondary servers to serve some of the read traffic.\",{\"_137\":138,\"_41\":150,\"_140\":479,\"_48\":480},{\"_153\":482},[481],\"Read-after-write\",\"https://jepsen.io/consistency/models/read-your-writes\",{\"_137\":138,\"_41\":139,\"_140\":484,\"_48\":485},{},[486,487,488,489,490],\"Replication is often used in database clusters, especially ones using a \",[\"SingleFetchClassInstance\",496],\" (aka leader-follower) architecture. Sometimes the replication is set to be asynchronous, and at other times group-communication based. In these scenarios \",[\"SingleFetchClassInstance\",491],\" is defined as the time passing between a write on the primary server and the time it is applied or made visible on the replica/secondary server.\",{\"_137\":138,\"_41\":492,\"_140\":493,\"_48\":494},\"strong\",{},[495],\"replication lag\",{\"_137\":138,\"_41\":150,\"_140\":497,\"_48\":498},{\"_153\":500},[499],\"primary-secondary\",\"https://dev.mysql.com/doc/refman/8.4/en/group-replication-primary-secondary-replication.html\",{\"_137\":138,\"_41\":272,\"_140\":502,\"_48\":503},{\"_50\":83},[504],[\"SingleFetchClassInstance\",505],{\"_137\":138,\"_41\":150,\"_140\":506,\"_48\":507},{\"_153\":508},[84],\"#replication-lag\",{\"_137\":138,\"_41\":139,\"_140\":510,\"_48\":511},{},[512],\"To do this, we look for metrics that define or predict service level objectives (SLO) for the database. But things are not always so simple. Let's start with a popular metric which is widely used as a throttling indicator, and see what's so special about that metric.\",{\"_137\":138,\"_41\":139,\"_140\":514,\"_48\":515},{},[516],\"We therefore need to be able to determine: how do we consider our database to be \\\"healthy\\\"? How do we determine if it is being overwhelmed?\",{\"_137\":138,\"_41\":139,\"_140\":518,\"_48\":519},{},[520],\"Some generic throttlers only allow a regulated rate of requests, in anticipation that the consuming job will be able to process them at some known, fixed rate. With databases, things are less clear. A database can only handle so many queries at any given point in time, or over some period of time. However, not all queries are created equal. The database capacity of serving queries depends on the scope of queries, any hot spots or cold spots in affected data, the state of the page cache, overlap or lack thereof of data served by queries, to name a few factors.\",{\"_137\":138,\"_41\":145,\"_140\":522,\"_48\":523},{\"_50\":62},[524],[\"SingleFetchClassInstance\",525],{\"_137\":138,\"_41\":150,\"_140\":526,\"_48\":527},{\"_153\":528},[63],\"#what-does-the-throttler-throttle-on\",{\"_137\":138,\"_41\":139,\"_140\":530,\"_48\":531},{},[532],\"Some throttler implementations are collaborative, meaning they assume clients will respect their instructions. Others act as barriers between the app and the database. Either way, if the throttler indicates that the database is overloaded, the job should hold back for a period of time and then request access again. This process repeats until granted. Each subtask should be small enough so as not to single-handedly tank the database's serving capacity, while large enough to compensate for the added throttler overhead and to enable meaningful progress.\",{\"_137\":138,\"_41\":139,\"_140\":534,\"_48\":535},{},[536],\"By way of illustration, consider a job that needs to import 10 million rows into the database. Instead of attempting to apply all 10 million in one go, the job breaks down the task into much smaller subtasks: it will try to import (write) 100 rows at a time. Before any such import, it will request access from the throttler.\",{\"_137\":138,\"_41\":139,\"_140\":538,\"_48\":539},{},[540],\"There are different approaches to throttling requests in a database. We focus on throttling asynchronous, batch, and massive operations that are not time critical. Examples could be ETLs, data imports, online DDL operations, mass purges of data, resharding, and so forth. The throttler will push back on those operations that can span minutes, hours, or days of operation. Other forms of throttling may push back on OLTP production traffic. This discussion applies equally to both.\",{\"_137\":138,\"_41\":145,\"_140\":542,\"_48\":543},{\"_50\":86},[544],[\"SingleFetchClassInstance\",545],{\"_137\":138,\"_41\":150,\"_140\":546,\"_48\":547},{\"_153\":548},[87],\"#which-requests-do-you-throttle\",{\"_137\":138,\"_41\":139,\"_140\":550,\"_48\":551},{},[552,553,554],\"A throttler is a service or component that pushes back against an \",[\"SingleFetchClassInstance\",555],\" of requests to ensure the system has the capacity to handle all received requests without being overwhelmed. In this series of posts, we illustrate design considerations for a database system throttler, whose purpose is to keep the database system healthy overall. We discuss choice of metrics, granularity, behavior, impact, prioritization, and other topics.\",{\"_137\":138,\"_41\":150,\"_140\":556,\"_48\":557},{\"_153\":559},[558],\"incoming flow\",\"https://en.wikipedia.org/wiki/Flow_control_(data)\",\"current\",{\"_562\":563,\"_564\":565,\"_566\":563},\"development\",false,\"env\",{\"_567\":568,\"_569\":570,\"_571\":572,\"_573\":574,\"_575\":576},\"userSignedIn\",\"IMAGE_CDN\",\"https://planetscale-images.imgix.net\",\"IMAGE_CDN_ENABLED\",\"true\",\"INTERNAL_API\",\"https://api.planetscale.com\",\"RELEASE\",\"117b8aaf-965c-42bc-b013-5f72770de4d9\",\"SENTRY_DSN\",\"https://bd81903b44804e22a06bdc0c1a91b303@o499952.ingest.us.sentry.io/4504531942572032\"]\n");</script><!--$--><script nonce="nyDuN2IITQLGs2NvcHjdoIBrtda9b+1NY1U/+JH3x2E=">window.__reactRouterContext.streamController.close();</script><!--/$--><!--/$--></body></html> |