Files
nexus/sreweekly/articles/519/03-on-benchmarking.html
2026-09-12 17:23:01 +08:00

202 lines
138 KiB
HTML

<!DOCTYPE html><html lang="en"><head><meta charSet="utf-8"/><meta name="viewport" content="width=device-width, initial-scale=1"/><meta name="theme-color" content="#111111"/><meta name="user-signed-in" content="false"/><title>On benchmarking — PlanetScale</title><meta name="description" content="Benchmarking is hard. Done wrong it is very misleading, and unfortunately it is frequently done wrong. Let&#x27;s explore how not to make silly mistakes."/><meta name="robots"/><meta property="og:url" content="https://planetscale.com/blog/on-benchmarking"/><meta property="og:type" content="website"/><meta property="og:title" content="On benchmarking — PlanetScale"/><meta property="og:image" content="https://planetscale.com/assets/on-benchmarking-social-JmCxOY2L.png"/><meta property="og:description" content="Benchmarking is hard. Done wrong it is very misleading, and unfortunately it is frequently done wrong. Let&#x27;s explore how not to make silly mistakes."/><meta property="twitter:card" content="summary_large_image"/><meta property="twitter:site" content="@PlanetScale"/><meta property="twitter:creator" content="@PlanetScale"/><meta property="twitter:url" content="https://planetscale.com/blog/on-benchmarking"/><meta property="twitter:title" content="On benchmarking — PlanetScale"/><meta property="twitter:description" content="Benchmarking is hard. Done wrong it is very misleading, and unfortunately it is frequently done wrong. Let&#x27;s explore how not to make silly mistakes."/><meta property="twitter:image" content="https://planetscale.com/assets/on-benchmarking-social-JmCxOY2L.png"/><link rel="canonical" href="https://planetscale.com/blog/on-benchmarking"/><link rel="preconnect" href="https://planetscale-images.imgix.net"/><link nonce="1NjQdEMwrEHFpdD1zVsfJP6ksuIJ/ecPax+EDqlkDLg=" rel="icon" href="/favicon.ico" type="image/x-icon" sizes="16x16"/><link nonce="1NjQdEMwrEHFpdD1zVsfJP6ksuIJ/ecPax+EDqlkDLg=" rel="icon" href="/icon.png" type="image/png" sizes="32x32"/><link nonce="1NjQdEMwrEHFpdD1zVsfJP6ksuIJ/ecPax+EDqlkDLg=" rel="apple-touch-icon" href="/apple-touch-icon.png" type="image/png" sizes="32x32"/><link nonce="1NjQdEMwrEHFpdD1zVsfJP6ksuIJ/ecPax+EDqlkDLg=" rel="manifest" href="/manifest.webmanifest"/><link rel="modulepreload" href="/assets/entry.client-3vubyXrk.js" nonce="1NjQdEMwrEHFpdD1zVsfJP6ksuIJ/ecPax+EDqlkDLg="/><link rel="modulepreload" href="/assets/jsx-runtime-DwfQwkRq.js" nonce="1NjQdEMwrEHFpdD1zVsfJP6ksuIJ/ecPax+EDqlkDLg="/><link rel="modulepreload" href="/assets/components-_bNmAApg.js" nonce="1NjQdEMwrEHFpdD1zVsfJP6ksuIJ/ecPax+EDqlkDLg="/><link rel="modulepreload" href="/assets/_.well-known_.mcp.server-card_.json_-Sx7XeH3e.js" nonce="1NjQdEMwrEHFpdD1zVsfJP6ksuIJ/ecPax+EDqlkDLg="/><link rel="modulepreload" href="/assets/index-mKTXLmHu.js" nonce="1NjQdEMwrEHFpdD1zVsfJP6ksuIJ/ecPax+EDqlkDLg="/><link rel="modulepreload" href="/assets/errorBoundaries-DhW4jVYt.js" nonce="1NjQdEMwrEHFpdD1zVsfJP6ksuIJ/ecPax+EDqlkDLg="/><link rel="modulepreload" href="/assets/root-DbOv4-98.js" nonce="1NjQdEMwrEHFpdD1zVsfJP6ksuIJ/ecPax+EDqlkDLg="/><link rel="modulepreload" href="/assets/lib-Dg89tQ22.js" nonce="1NjQdEMwrEHFpdD1zVsfJP6ksuIJ/ecPax+EDqlkDLg="/><link rel="modulepreload" href="/assets/analytics.client-DM6E8o1h.js" nonce="1NjQdEMwrEHFpdD1zVsfJP6ksuIJ/ecPax+EDqlkDLg="/><link rel="modulepreload" href="/assets/SiteHeader-C2U5gvDH.js" nonce="1NjQdEMwrEHFpdD1zVsfJP6ksuIJ/ecPax+EDqlkDLg="/><link rel="modulepreload" href="/assets/current-9yDxj94E.js" nonce="1NjQdEMwrEHFpdD1zVsfJP6ksuIJ/ecPax+EDqlkDLg="/><link rel="modulepreload" href="/assets/clsx-eT0YPcGk.js" nonce="1NjQdEMwrEHFpdD1zVsfJP6ksuIJ/ecPax+EDqlkDLg="/><link rel="modulepreload" href="/assets/bugs-38ilEoW0.js" nonce="1NjQdEMwrEHFpdD1zVsfJP6ksuIJ/ecPax+EDqlkDLg="/><link rel="modulepreload" href="/assets/keyboard-D-uXZORL.js" nonce="1NjQdEMwrEHFpdD1zVsfJP6ksuIJ/ecPax+EDqlkDLg="/><link rel="modulepreload" href="/assets/use-tab-direction-dKm-S3Ck.js" nonce="1NjQdEMwrEHFpdD1zVsfJP6ksuIJ/ecPax+EDqlkDLg="/><link rel="modulepreload" href="/assets/blog-pXH7ptHJ.js" nonce="1NjQdEMwrEHFpdD1zVsfJP6ksuIJ/ecPax+EDqlkDLg="/><link rel="modulepreload" href="/assets/blog._slug-Ch_92qsH.js" nonce="1NjQdEMwrEHFpdD1zVsfJP6ksuIJ/ecPax+EDqlkDLg="/><link rel="modulepreload" href="/assets/ContentImage-Dh6VEOUl.js" nonce="1NjQdEMwrEHFpdD1zVsfJP6ksuIJ/ecPax+EDqlkDLg="/><link rel="modulepreload" href="/assets/BlogCategoryLink-DmQyn0gp.js" nonce="1NjQdEMwrEHFpdD1zVsfJP6ksuIJ/ecPax+EDqlkDLg="/><link rel="modulepreload" href="/assets/Details-BSB_b6hI.js" nonce="1NjQdEMwrEHFpdD1zVsfJP6ksuIJ/ecPax+EDqlkDLg="/><link rel="modulepreload" href="/assets/Skittle-CDFOPRjH.js" nonce="1NjQdEMwrEHFpdD1zVsfJP6ksuIJ/ecPax+EDqlkDLg="/><link rel="modulepreload" href="/assets/SiteFooter-B2Gq9u2j.js" nonce="1NjQdEMwrEHFpdD1zVsfJP6ksuIJ/ecPax+EDqlkDLg="/><link rel="modulepreload" href="/assets/Vimeo-00PQJDli.js" nonce="1NjQdEMwrEHFpdD1zVsfJP6ksuIJ/ecPax+EDqlkDLg="/><link rel="modulepreload" href="/assets/YouTube-CMfaljVr.js" nonce="1NjQdEMwrEHFpdD1zVsfJP6ksuIJ/ecPax+EDqlkDLg="/><link rel="modulepreload" href="/assets/date-CJTFH3uT.js" nonce="1NjQdEMwrEHFpdD1zVsfJP6ksuIJ/ecPax+EDqlkDLg="/><link rel="modulepreload" href="/assets/use-inert-others-BMJ6-xOX.js" nonce="1NjQdEMwrEHFpdD1zVsfJP6ksuIJ/ecPax+EDqlkDLg="/><link rel="modulepreload" href="/assets/description-Cf6FZmDe.js" nonce="1NjQdEMwrEHFpdD1zVsfJP6ksuIJ/ecPax+EDqlkDLg="/><link rel="modulepreload" href="/assets/use-is-mounted-uQsUZyP9.js" nonce="1NjQdEMwrEHFpdD1zVsfJP6ksuIJ/ecPax+EDqlkDLg="/><link rel="modulepreload" href="/assets/types-DvonrUFF.js" nonce="1NjQdEMwrEHFpdD1zVsfJP6ksuIJ/ecPax+EDqlkDLg="/><link rel="stylesheet" href="/assets/styles-ns8XBZ1D.css"/><script nonce="1NjQdEMwrEHFpdD1zVsfJP6ksuIJ/ecPax+EDqlkDLg=">window.ENV = {"IMAGE_CDN":"https://planetscale-images.imgix.net","IMAGE_CDN_ENABLED":"true","INTERNAL_API":"https://api.planetscale.com","RELEASE":"117b8aaf-965c-42bc-b013-5f72770de4d9","SENTRY_DSN":"https://bd81903b44804e22a06bdc0c1a91b303@o499952.ingest.us.sentry.io/4504531942572032"}</script></head><body class="flex min-h-screen flex-col"><div class="bg-neki px-3 py-1 text-center font-medium text-gray-900 dark:font-semibold"><span>Neki, sharded Postgres, is now available.</span> <span class="whitespace-nowrap"><a href="https://auth.planetscale.com/sign-up" class="whitespace-nowrap bg-gray-900 px-sm font-semibold text-white">Get started</a></span></div><header class="relative mb-6 mt-4 bg-primary"><div class="flex flex-col gap-y-3 px-3 sm:px-5 container max-w-7xl"><div class="grid w-full grid-cols-[auto_1fr] grid-rows-1 items-center lg:items-start lg:gap-3"><a aria-label="Go to homepage" class="col-start-1 col-end-2 h-4 w-4 rounded-full text-primary lg:hidden" href="/" data-discover="true"><svg xmlns="http://www.w3.org/2000/svg" width="32" height="32" fill="none" viewBox="0 0 40 40"><path fill="currentColor" d="M0 20C0 8.954 8.954 0 20 0c8.121 0 15.112 4.84 18.245 11.794l-26.45 26.45a20 20 0 0 1-3.225-1.83L24.984 20H20L5.858 34.142A19.94 19.94 0 0 1 0 20M39.999 20.007 20.006 40c11.04-.004 19.99-8.953 19.993-19.993"></path></svg></a><div class="group col-start-2 col-end-3 row-start-1 flex shrink-0 items-center justify-end gap-1.5 lg:gap-3"><div class="flex flex-row gap-2 lg:flex-col lg:gap-1 xl:flex-row"><div class="flex items-center justify-end gap-1 lg:h-4"><a href="https://auth.planetscale.com/sign-in" class="font-semibold text-primary hover:text-orange">Sign in</a></div><div class="flex items-center justify-end gap-0.5 lg:h-4"><form class="btn-sm hidden sm:inline-flex" action="/api/demo-sessions" method="post"><button type="submit" class="btn btn-outline btn-sm hidden sm:inline-flex">View sandbox</button></form><a class="btn btn-sm" href="/contact" data-discover="true">Get in touch</a></div></div></div><div class="col-start-1 col-end-2 flex items-center gap-x-3 lg:row-start-1 lg:h-4"><a aria-label="Go to homepage" class="col-start-1 col-end-2 hidden h-4 w-4 rounded-full text-primary lg:block" href="/" data-discover="true"><svg xmlns="http://www.w3.org/2000/svg" width="32" height="32" fill="none" viewBox="0 0 40 40"><path fill="currentColor" d="M0 20C0 8.954 8.954 0 20 0c8.121 0 15.112 4.84 18.245 11.794l-26.45 26.45a20 20 0 0 1-3.225-1.83L24.984 20H20L5.858 34.142A19.94 19.94 0 0 1 0 20M39.999 20.007 20.006 40c11.04-.004 19.99-8.953 19.993-19.993"></path></svg></a><nav aria-label="Main" data-orientation="horizontal" class="hidden items-center lg:flex"><ul class="flex flex-wrap gap-x-1 md:flex-nowrap"><li><div data-headlessui-state=""><button class="font-semibold text-primary hover:text-contrast focus-visible:ring-0 ui-open:text-orange" type="button" aria-expanded="false" data-headlessui-state="">Platform<span class="ml-sm inline-block ui-open:rotate-180">▾</span></button></div><span hidden="" style="position:fixed;top:1px;left:1px;width:1px;height:0;padding:0;margin:-1px;overflow:hidden;clip:rect(0, 0, 0, 0);white-space:nowrap;border-width:0;display:none"></span></li><li class="text-decoration" role="presentation">|</li><li><div data-headlessui-state=""><button class="font-semibold text-primary hover:text-contrast focus-visible:ring-0 ui-open:text-orange" type="button" aria-expanded="false" data-headlessui-state="">Resources<span class="ml-sm inline-block ui-open:rotate-180">▾</span></button></div><span hidden="" style="position:fixed;top:1px;left:1px;width:1px;height:0;padding:0;margin:-1px;overflow:hidden;clip:rect(0, 0, 0, 0);white-space:nowrap;border-width:0;display:none"></span></li><li class="text-decoration" role="presentation">|</li><li><a class="font-semibold text-primary hover:text-contrast" href="/docs">Documentation</a></li><li class="text-decoration" role="presentation">|</li><li><a class="font-semibold text-primary hover:text-contrast" href="/pricing" data-discover="true">Pricing</a></li><li class="text-decoration" role="presentation">|</li><li><a class="font-semibold text-primary hover:text-contrast" href="/migrate" data-discover="true">Migrate</a></li></ul></nav></div></div><details class="lg:hidden"><summary>Navigation</summary><nav class="dashed-box mt-1 p-3"><ul class="flex flex-wrap gap-x-1 md:flex-nowrap"><li><div data-headlessui-state=""><button class="font-semibold text-primary hover:text-contrast focus-visible:ring-0 ui-open:text-orange" type="button" aria-expanded="false" data-headlessui-state="">Platform<span class="ml-sm inline-block ui-open:rotate-180">▾</span></button></div><span hidden="" style="position:fixed;top:1px;left:1px;width:1px;height:0;padding:0;margin:-1px;overflow:hidden;clip:rect(0, 0, 0, 0);white-space:nowrap;border-width:0;display:none"></span></li><li class="text-decoration" role="presentation">|</li><li><div data-headlessui-state=""><button class="font-semibold text-primary hover:text-contrast focus-visible:ring-0 ui-open:text-orange" type="button" aria-expanded="false" data-headlessui-state="">Resources<span class="ml-sm inline-block ui-open:rotate-180">▾</span></button></div><span hidden="" style="position:fixed;top:1px;left:1px;width:1px;height:0;padding:0;margin:-1px;overflow:hidden;clip:rect(0, 0, 0, 0);white-space:nowrap;border-width:0;display:none"></span></li><li class="text-decoration" role="presentation">|</li><li><a class="font-semibold text-primary hover:text-contrast" href="/docs">Documentation</a></li><li class="text-decoration" role="presentation">|</li><li><a class="font-semibold text-primary hover:text-contrast" href="/pricing" data-discover="true">Pricing</a></li><li class="text-decoration" role="presentation">|</li><li><a class="font-semibold text-primary hover:text-contrast" href="/migrate" data-discover="true">Migrate</a></li></ul></nav></details></div></header><main class="container mb-6 flex max-w-7xl flex-1 flex-col px-3 sm:px-5 lg:px-12"><section class=""><p class="block"><a class="pr-sm text-primary hover:text-contrast" href="/blog" data-discover="true">Blog</a><span class="px-sm text-decoration">|</span><a class="px-sm text-blue hover:bg-blue-100 dark:hover:bg-blue-900" href="/blog/category/engineering" data-discover="true">Engineering</a></p><div class="flex lg:flex-row-reverse lg:gap-x-6"><div class="lg:sticky lg:top-2 lg:self-start"><button class="absolute right-0 bg-gray-100 px-sm md:block lg:hidden dark:bg-gray-800 -mt-9 hidden"><span class="inline">Table of contents «</span><span class="hidden">Close »</span></button><aside class="tree-nav w-full shrink-0 space-y-3 lg:w-36 hidden lg:block"><div><h4 class="text-secondary">Table of contents</h4><ul><li><a class="font-semibold text-primary hover:text-blue" href="/blog/on-benchmarking#client-server-architecture" data-discover="true">Client-server architecture</a></li><li><a class="font-semibold text-primary hover:text-blue" href="/blog/on-benchmarking#choosing-resources" data-discover="true">Choosing resources</a></li><li><a class="font-semibold text-primary hover:text-blue" href="/blog/on-benchmarking#workload" data-discover="true">Workload</a></li><li><a class="font-semibold text-primary hover:text-blue" href="/blog/on-benchmarking#closed-and-open-loop" data-discover="true">Closed and open loop</a></li><li><a class="font-semibold text-primary hover:text-blue" href="/blog/on-benchmarking#what-to-measure" data-discover="true">What to measure?</a></li><li><a class="font-semibold text-primary hover:text-blue" href="/blog/on-benchmarking#throughput" data-discover="true">Throughput</a></li><li><a class="font-semibold text-primary hover:text-blue" href="/blog/on-benchmarking#latency" data-discover="true">Latency</a></li><li><a class="font-semibold text-primary hover:text-blue" href="/blog/on-benchmarking#warmup" data-discover="true">Warmup</a></li><li><a class="font-semibold text-primary hover:text-blue" href="/blog/on-benchmarking#configuration" data-discover="true">Configuration</a></li><li><a class="font-semibold text-primary hover:text-blue" href="/blog/on-benchmarking#inconsistency" data-discover="true">(In)consistency</a></li><li><a class="font-semibold text-primary hover:text-blue" href="/blog/on-benchmarking#apples-to-apples-to-oranges" data-discover="true">Apples to apples to oranges</a></li><li><a class="font-semibold text-primary hover:text-blue" href="/blog/on-benchmarking#document-everything" data-discover="true">Document everything</a></li><li><a class="font-semibold text-primary hover:text-blue" href="/blog/on-benchmarking#benchmark-crimes" data-discover="true">Benchmark crimes</a></li><li><a class="font-semibold text-primary hover:text-blue" href="/blog/on-benchmarking#go-forth-and-benchmark" data-discover="true">Go forth and benchmark</a></li></ul><div class="mb-3 mt-6 border bg-blue-50 p-3 font-semibold text-contrast dark:bg-blue-900"><p>PlanetScale, the fastest cloud Postgres, from $5/month.</p><p><a href="https://app.planetscale.com/new">Start now</a></p></div><p>Get the <a href="/blog/feed.atom">RSS feed</a></p></div></aside></div><article class="min-w-0 flex-grow"><h1>On benchmarking</h1><p class="text-secondary"><a class="text-contrast no-underline" href="/blog/author/ben" data-discover="true">Ben Dicken</a> <!-- -->[<a class="no-underline hover:bg-blue-100 dark:hover:bg-blue-900" href="https://x.com/BenjDicken" rel="noopener noreferrer" target="_blank" title="@BenjDicken on X">@<!-- -->BenjDicken</a>]<!-- --> |<!-- --> <time dateTime="2026-05-05">May 5, 2026</time></p><div class="blog-post-body"><p>Benchmarking is hard.<!-- --> <!-- -->There are many ways to do it wrong and few to do it right.</p><p>But zooming out from any single system or harness, there are broad principles that should be applied to all benchmarking.<!-- --> <!-- -->Using these correctly makes it difficult to produce biased results.</p><p>Am I the world&#x27;s best benchmarker?<!-- --> <!-- -->Certainly not.<!-- --> <!-- -->I invented the <a href="https://x.com/BenjDicken/status/1861072804239847914">language balls</a>, after all.<!-- --> <!-- -->But correctness and precision are important parts of PlanetScale&#x27;s culture.<!-- --> <!-- -->We&#x27;ve spent considerable time learning the art of benchmarking, and are here to share best-practices.</p><p>Here, we&#x27;re focusing primarily on benchmarking <em>databases</em>, but these principles apply to many domains.</p><h2 id="client-server-architecture"><a href="#client-server-architecture">Client-server architecture</a></h2><p>Databases typically operate in a client-server model.<!-- --> <!-- -->The database server is started, accepts connections from clients, executes queries, and returns results.</p><p>To benchmark, we need a client that establishes the connections, generates queries, and takes measurements.<!-- --> <!-- -->Since both sides consume resources and we want to give the <em>database</em> its full share of the host server, it&#x27;s common to set up a distinct server for benchmark execution.</p><p>As usual, <em>there&#x27;s a catch</em>.<!-- --> <!-- -->This introduces latency between the two machines.</p><p><button type="button" aria-haspopup="dialog" aria-expanded="false" aria-label="Enlarge image: Client-server" class="focus-visible-ring group relative block w-fit max-w-[min(100%,800px)] cursor-zoom-in text-left [&amp;_picture]:contents"><picture class="block"><source media="(prefers-color-scheme: light), (prefers-color-scheme: no-preference)" srcSet="https://planetscale-images.imgix.net/assets/client-server-DSv7LBqp.png?auto=compress%2Cformat"/><source media="(prefers-color-scheme: dark)" srcSet="https://planetscale-images.imgix.net/assets/client-server-darkmode-l5yZMo25.png?auto=compress%2Cformat"/><img alt="Client-server" src="https://planetscale-images.imgix.net/assets/client-server-DSv7LBqp.png?auto=compress%2Cformat" width="2400" height="1880" loading="lazy" class="w-auto max-w-full"/></picture><span aria-hidden="true" class="pointer-events-none absolute right-1 top-1 z-10 flex h-5 w-5 items-center justify-center border border-white/25 bg-black/70 text-white backdrop-blur-sm transition-colors transition-opacity group-hover:bg-black/90 group-hover:opacity-100 group-focus-visible:opacity-100 motion-reduce:transition-none [@media(hover:hover)_and_(pointer:fine)]:opacity-0"><svg width="14" height="14" viewBox="0 0 16 16" fill="none" stroke="currentColor" stroke-width="1.5" stroke-linecap="round" stroke-linejoin="round"><path d="M10 2h4v4M6 14H2v-4M14 2l-4.5 4.5M2 14l4.5-4.5"></path></svg></span></button><span hidden="" style="position:fixed;top:1px;left:1px;width:1px;height:0;padding:0;margin:-1px;overflow:hidden;clip:rect(0, 0, 0, 0);white-space:nowrap;border-width:0;display:none"></span></p><p>How much this skews the results of the benchmark depends quite a bit on how &quot;far apart&quot; the benchmark server and database server are<!-- --> <!-- -->(network latency)<!-- --> <!-- -->and how long the queries / transactions take on the database<!-- --> <!-- -->(execution latency).</p><p><button type="button" aria-haspopup="dialog" aria-expanded="false" aria-label="Enlarge image: Network latency and execution latency" class="focus-visible-ring group relative block w-fit max-w-[min(100%,800px)] cursor-zoom-in text-left [&amp;_picture]:contents"><picture class="block"><source media="(prefers-color-scheme: light), (prefers-color-scheme: no-preference)" srcSet="https://planetscale-images.imgix.net/assets/network-latency-execution-latency-D7jzvLOo.png?auto=compress%2Cformat"/><source media="(prefers-color-scheme: dark)" srcSet="https://planetscale-images.imgix.net/assets/network-latency-execution-latency-darkmode-DonplkLE.png?auto=compress%2Cformat"/><img alt="Network latency and execution latency" src="https://planetscale-images.imgix.net/assets/network-latency-execution-latency-D7jzvLOo.png?auto=compress%2Cformat" width="3256" height="1444" loading="lazy" class="w-auto max-w-full"/></picture><span aria-hidden="true" class="pointer-events-none absolute right-1 top-1 z-10 flex h-5 w-5 items-center justify-center border border-white/25 bg-black/70 text-white backdrop-blur-sm transition-colors transition-opacity group-hover:bg-black/90 group-hover:opacity-100 group-focus-visible:opacity-100 motion-reduce:transition-none [@media(hover:hover)_and_(pointer:fine)]:opacity-0"><svg width="14" height="14" viewBox="0 0 16 16" fill="none" stroke="currentColor" stroke-width="1.5" stroke-linecap="round" stroke-linejoin="round"><path d="M10 2h4v4M6 14H2v-4M14 2l-4.5 4.5M2 14l4.5-4.5"></path></svg></span></button><span hidden="" style="position:fixed;top:1px;left:1px;width:1px;height:0;padding:0;margin:-1px;overflow:hidden;clip:rect(0, 0, 0, 0);white-space:nowrap;border-width:0;display:none"></span></p><p>Let&#x27;s consider a scenario where each query takes ~10ms to execute on the database.<!-- --> <!-- -->If the network round-trip time is 2.5 milliseconds, then we can execute approximately 80 queries per second over a single connection.<!-- --> <!-- -->On the other hand, what if the round-trip is 15 milliseconds?<!-- --> <!-- -->We&#x27;ve now cut our single-threaded QPS capability in ~half, resulting in 40 QPS.</p><p><button type="button" aria-haspopup="dialog" aria-expanded="false" aria-label="Enlarge image: How network latency impacts throughput" class="focus-visible-ring group relative block w-fit max-w-[min(100%,800px)] cursor-zoom-in text-left [&amp;_picture]:contents"><picture class="block"><source media="(prefers-color-scheme: light), (prefers-color-scheme: no-preference)" srcSet="https://planetscale-images.imgix.net/assets/network-latency-difference-C_4xTSar.png?auto=compress%2Cformat"/><source media="(prefers-color-scheme: dark)" srcSet="https://planetscale-images.imgix.net/assets/network-latency-difference-darkmode-CTYmZVN1.png?auto=compress%2Cformat"/><img alt="How network latency impacts throughput" src="https://planetscale-images.imgix.net/assets/network-latency-difference-C_4xTSar.png?auto=compress%2Cformat" width="3288" height="1404" loading="lazy" class="w-auto max-w-full"/></picture><span aria-hidden="true" class="pointer-events-none absolute right-1 top-1 z-10 flex h-5 w-5 items-center justify-center border border-white/25 bg-black/70 text-white backdrop-blur-sm transition-colors transition-opacity group-hover:bg-black/90 group-hover:opacity-100 group-focus-visible:opacity-100 motion-reduce:transition-none [@media(hover:hover)_and_(pointer:fine)]:opacity-0"><svg width="14" height="14" viewBox="0 0 16 16" fill="none" stroke="currentColor" stroke-width="1.5" stroke-linecap="round" stroke-linejoin="round"><path d="M10 2h4v4M6 14H2v-4M14 2l-4.5 4.5M2 14l4.5-4.5"></path></svg></span></button><span hidden="" style="position:fixed;top:1px;left:1px;width:1px;height:0;padding:0;margin:-1px;overflow:hidden;clip:rect(0, 0, 0, 0);white-space:nowrap;border-width:0;display:none"></span></p><p>Same database.<!-- --> <!-- -->Same benchmark client.<!-- --> <!-- -->The only difference is the speed at which bytes can go over the wire between the two.</p><p>This latency variation will always have an impact on latency measurements.</p><p>It <em>can</em> also impact throughput.<!-- --> <!-- -->We often don&#x27;t run benchmarks on a single connection.<!-- --> <!-- -->We&#x27;ll do 10, 50, or 100 simultaneous connections to best utilize the parallelism of the machine and database.<!-- --> <!-- -->But if we have a fixed connection count, and are not making it dynamic to account for round-trip latency, we can end up allowing the elevated latency to hurt throughput.</p><p>Finally, you should double-check that the client server is not a bottleneck.<!-- --> <!-- -->While benchmarking, ensure that CPU and network utilization are well under their capacity.<!-- --> <!-- -->We want to be straining the database server, not the client.</p><h2 id="choosing-resources"><a href="#choosing-resources">Choosing resources</a></h2><p>It&#x27;s easy to make one database look better than another with an imbalance of resources.<!-- --> <!-- -->Postgres running on a 16-core server will almost always perform better than on an 8-core server.</p><p>An important prerequisite to proper benchmarking is setting up the compute, storage, and networking resources to allow for a fair fight.</p><p>This isn&#x27;t as easy as it sounds, especially when we&#x27;re talking about running things in the hyperscaler clouds like AWS and GCP.<!-- --> <!-- -->For example, the Geekbench results for an AWS <a href="https://browser.geekbench.com/v6/cpu/2119560">r7g.2xlarge</a> are ~15% lower than the results for an <a href="https://browser.geekbench.com/v6/cpu/11335856">r8g.2xlarge</a>.<!-- --> <!-- -->Both have 8 vCPUs and 64 GB RAM.<!-- --> <!-- -->But move one generation newer, and there&#x27;s a ~15% CPU improvement.</p><p><button type="button" aria-haspopup="dialog" aria-expanded="false" aria-label="Enlarge image: Geekbench results" class="focus-visible-ring group relative block w-fit max-w-[min(100%,800px)] cursor-zoom-in text-left [&amp;_picture]:contents"><picture class="block"><source media="(prefers-color-scheme: light), (prefers-color-scheme: no-preference)" srcSet="https://planetscale-images.imgix.net/assets/geekbench-DFb0s8q0.png?auto=compress%2Cformat"/><source media="(prefers-color-scheme: dark)" srcSet="https://planetscale-images.imgix.net/assets/geekbench-darkmode-zAs-8HQc.png?auto=compress%2Cformat"/><img alt="Geekbench results" src="https://planetscale-images.imgix.net/assets/geekbench-DFb0s8q0.png?auto=compress%2Cformat" width="2000" height="1264" loading="lazy" class="w-auto max-w-full"/></picture><span aria-hidden="true" class="pointer-events-none absolute right-1 top-1 z-10 flex h-5 w-5 items-center justify-center border border-white/25 bg-black/70 text-white backdrop-blur-sm transition-colors transition-opacity group-hover:bg-black/90 group-hover:opacity-100 group-focus-visible:opacity-100 motion-reduce:transition-none [@media(hover:hover)_and_(pointer:fine)]:opacity-0"><svg width="14" height="14" viewBox="0 0 16 16" fill="none" stroke="currentColor" stroke-width="1.5" stroke-linecap="round" stroke-linejoin="round"><path d="M10 2h4v4M6 14H2v-4M14 2l-4.5 4.5M2 14l4.5-4.5"></path></svg></span></button><span hidden="" style="position:fixed;top:1px;left:1px;width:1px;height:0;padding:0;margin:-1px;overflow:hidden;clip:rect(0, 0, 0, 0);white-space:nowrap;border-width:0;display:none"></span></p><p>You might then be tempted to just use the same instance for everything, but this breaks down too.<!-- --> <!-- -->The availability of instance types varies over time, region, and database provider.<!-- --> <!-- -->In some cases, it&#x27;s not possible to match.</p><p>In an ideal world, we&#x27;d run everything on the <em>exact</em> same instance.<!-- --> <!-- -->In reality, we sometimes have to settle for matching CPUs and RAM as best we can, and living with the differences.<!-- --> <!-- -->However, you must give this your best effort.<!-- --> <!-- -->Purposefully choosing to benchmark <em>your</em> product on 2025-gen CPU and then comparing to a competitor&#x27;s product on a 2022 CPU, when the alternate was readily available, is intentionally misleading.</p><h2 id="workload"><a href="#workload">Workload</a></h2><p>Even once we know that our infrastructure is set up sanely, there&#x27;s a lot to consider for the workload we run.</p><p>The easiest way to think about this is in terms of traffic ratios.</p><ul><li>How many queries are hitting RAM vs disk?</li><li>What % of the data is hot (frequently queried) vs cold (rarely queried)?</li><li>What&#x27;s the ratio of reads to writes?</li></ul><p>All of these impact performance, especially when combined with the variations of underlying hardware.</p><p>Queries executed on a relational database often require some amount of I/O work.<!-- --> <!-- -->Writing data must always be persisted to disk.<!-- --> <em>Reading</em> data can come from the in-memory cache, or disk on cache misses.</p><p><button type="button" aria-haspopup="dialog" aria-expanded="false" aria-label="Enlarge image: RAM, disk latency" class="focus-visible-ring group relative block w-fit max-w-[min(100%,800px)] cursor-zoom-in text-left [&amp;_picture]:contents"><picture class="block"><source media="(prefers-color-scheme: light), (prefers-color-scheme: no-preference)" srcSet="https://planetscale-images.imgix.net/assets/ram-disk-B91Wyn3l.png?auto=compress%2Cformat"/><source media="(prefers-color-scheme: dark)" srcSet="https://planetscale-images.imgix.net/assets/ram-disk-darkmode-BsqMcUR2.png?auto=compress%2Cformat"/><img alt="RAM, disk latency" src="https://planetscale-images.imgix.net/assets/ram-disk-B91Wyn3l.png?auto=compress%2Cformat" width="2400" height="1352" loading="lazy" class="w-auto max-w-full"/></picture><span aria-hidden="true" class="pointer-events-none absolute right-1 top-1 z-10 flex h-5 w-5 items-center justify-center border border-white/25 bg-black/70 text-white backdrop-blur-sm transition-colors transition-opacity group-hover:bg-black/90 group-hover:opacity-100 group-focus-visible:opacity-100 motion-reduce:transition-none [@media(hover:hover)_and_(pointer:fine)]:opacity-0"><svg width="14" height="14" viewBox="0 0 16 16" fill="none" stroke="currentColor" stroke-width="1.5" stroke-linecap="round" stroke-linejoin="round"><path d="M10 2h4v4M6 14H2v-4M14 2l-4.5 4.5M2 14l4.5-4.5"></path></svg></span></button><span hidden="" style="position:fixed;top:1px;left:1px;width:1px;height:0;padding:0;margin:-1px;overflow:hidden;clip:rect(0, 0, 0, 0);white-space:nowrap;border-width:0;display:none"></span></p><p>Some databases operate on local SSDs, while others use network-attached storage like AWS EBS or Google Persistent Disk.<!-- --> <!-- -->Some even take a hybrid approach.<!-- --> <!-- -->Either way, the percent of read traffic hitting RAM vs disk impacts performance due to I/O wait times.</p><p>Consider a benchmark like <a href="https://github.com/akopytov/sysbench/blob/master/src/lua/oltp_read_only.lua">sysbench OLTP read-only</a>.<!-- --> <!-- -->This is a simple, read-only benchmark that runs a handful of select query patterns repeatedly.<!-- --> <!-- -->As benchmarks often do, the data size is configurable in the preparation phase.<!-- --> <!-- -->If we run this benchmark on a server with 64 GB of RAM and a 32 GB data size, the entire data set will fit in RAM after warming.<!-- --> <!-- -->The same benchmark run with a 320 GB data size will generate significant I/O and inevitably run slower.</p><p>This is related to, but not the same as, data distribution.</p><p>Even for a fixed data size, access patterns can vary widely.<!-- --> <!-- -->The simplest examples are uniform and Zipfian.</p><p><button type="button" aria-haspopup="dialog" aria-expanded="false" aria-label="Enlarge image: Types of data distributions" class="focus-visible-ring group relative block w-fit max-w-[min(100%,800px)] cursor-zoom-in text-left [&amp;_picture]:contents"><picture class="block"><source media="(prefers-color-scheme: light), (prefers-color-scheme: no-preference)" srcSet="https://planetscale-images.imgix.net/assets/distributions-DHT64hC8.png?auto=compress%2Cformat"/><source media="(prefers-color-scheme: dark)" srcSet="https://planetscale-images.imgix.net/assets/distributions-darkmode-Bgx07A5I.png?auto=compress%2Cformat"/><img alt="Types of data distributions" src="https://planetscale-images.imgix.net/assets/distributions-DHT64hC8.png?auto=compress%2Cformat" width="2252" height="1256" loading="lazy" class="w-auto max-w-full"/></picture><span aria-hidden="true" class="pointer-events-none absolute right-1 top-1 z-10 flex h-5 w-5 items-center justify-center border border-white/25 bg-black/70 text-white backdrop-blur-sm transition-colors transition-opacity group-hover:bg-black/90 group-hover:opacity-100 group-focus-visible:opacity-100 motion-reduce:transition-none [@media(hover:hover)_and_(pointer:fine)]:opacity-0"><svg width="14" height="14" viewBox="0 0 16 16" fill="none" stroke="currentColor" stroke-width="1.5" stroke-linecap="round" stroke-linejoin="round"><path d="M10 2h4v4M6 14H2v-4M14 2l-4.5 4.5M2 14l4.5-4.5"></path></svg></span></button><span hidden="" style="position:fixed;top:1px;left:1px;width:1px;height:0;padding:0;margin:-1px;overflow:hidden;clip:rect(0, 0, 0, 0);white-space:nowrap;border-width:0;display:none"></span></p><p>A <em>uniform</em> access pattern gives every row the same chance of being queried on each request.<!-- --> <!-- -->If we have 100 rows, each has a 1% chance of being read for each operation.</p><p>A <em>Zipfian</em> access pattern is skewed: the k-th most popular key is accessed roughly proportional to <code>1/k</code>.<!-- --> <!-- -->A small number of hot rows receive a large share of requests, while most rows are accessed rarely.</p><p>These are only simple models.<!-- --> <!-- -->Real workloads often have messier shapes: recently inserted rows might be hotter than old rows, one tenant might dominate traffic, or a small working set might receive most reads for a period of time.</p><p>Which pattern the benchmark operates with significantly impacts performance, because it in turn impacts how frequently we need to access disk vs RAM and the amount of cache churn.</p><h2 id="closed-and-open-loop"><a href="#closed-and-open-loop">Closed and open loop</a></h2><p>There are two types of benchmark workload shapes: open and closed loops.</p><p>In a closed-loop benchmark, the client sends requests and then waits for a response before sending the next.</p><div class="code-block" data-language="python"><div class="min-w-0 max-w-full"><pre class="shiki shiki-themes planetscale-light planetscale-dark" style="--shiki-light:#2b2b2b;--shiki-dark:#e1e1e1;--shiki-light-bg:#ebebeb;--shiki-dark-bg:#1a1a1a" tabindex="0"><code><span class="line"><span style="--shiki-light:#F35815;--shiki-dark:#F35815">while</span><span style="--shiki-light:#7D5903;--shiki-dark:#FED54A"> True</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">:</span></span>
<span class="line"><span style="--shiki-light:#818181;--shiki-dark:#A1A1A1"> # wait for response</span></span>
<span class="line"><span style="--shiki-light:#2B2B2B;--shiki-dark:#E1E1E1"> response </span><span style="--shiki-light:#F35815;--shiki-dark:#F35815">=</span><span style="--shiki-light:#0B6EC5;--shiki-dark:#73C7F9"> send_bench_request</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">()</span></span>
<span class="line"><span style="--shiki-light:#818181;--shiki-dark:#A1A1A1"> # then send next</span></span>
<span class="line"><span style="--shiki-light:#0B6EC5;--shiki-dark:#73C7F9"> process</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">(</span><span style="--shiki-light:#0B6EC5;--shiki-dark:#73C7F9">response</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">)</span></span>
<span class="line"></span></code></pre></div></div><p>We may do this in parallel across many connections, but each individual connection sends a controlled sequence of queries.<!-- --> <!-- -->A closed loop can also hide a failure mode called coordinated omission: when the database stalls, the client stops issuing new requests too, so the benchmark only records the stalled request and omits the work that would have queued behind it.<!-- --> <!-- -->This is especially misleading for tail latency, where the missing queued requests are exactly the ones that would have made p95/p99 look worse (more on latency and percentiles soon).</p><p>Open loop on the other hand has a fixed pace of sending requests, regardless of how quickly the database responds.</p><div class="code-block" data-language="python"><div class="min-w-0 max-w-full"><pre class="shiki shiki-themes planetscale-light planetscale-dark" style="--shiki-light:#2b2b2b;--shiki-dark:#e1e1e1;--shiki-light-bg:#ebebeb;--shiki-dark-bg:#1a1a1a" tabindex="0"><code><span class="line"><span style="--shiki-light:#F35815;--shiki-dark:#F35815">while</span><span style="--shiki-light:#7D5903;--shiki-dark:#FED54A"> True</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">:</span></span>
<span class="line"><span style="--shiki-light:#818181;--shiki-dark:#A1A1A1"> # fire and forget</span></span>
<span class="line"><span style="--shiki-light:#0B6EC5;--shiki-dark:#73C7F9"> send_bench_request</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">()</span></span>
<span class="line"><span style="--shiki-light:#818181;--shiki-dark:#A1A1A1"> # fixed pace</span></span>
<span class="line"><span style="--shiki-light:#2B2B2B;--shiki-dark:#E1E1E1"> time</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#0B6EC5;--shiki-dark:#73C7F9">sleep</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">(</span><span style="--shiki-light:#D92038;--shiki-dark:#FF7082">0.1</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">)</span></span>
<span class="line"></span></code></pre></div></div><p>This can be fixed throughout the entire benchmark duration, or vary in a controlled way:</p><p><button type="button" aria-haspopup="dialog" aria-expanded="false" aria-label="Enlarge image: Open vs Closed loop benchmark" class="focus-visible-ring group relative block w-fit max-w-[min(100%,800px)] cursor-zoom-in text-left [&amp;_picture]:contents"><picture class="block"><source media="(prefers-color-scheme: light), (prefers-color-scheme: no-preference)" srcSet="https://planetscale-images.imgix.net/assets/open-closed-loop-B0_3fnSO.png?auto=compress%2Cformat"/><source media="(prefers-color-scheme: dark)" srcSet="https://planetscale-images.imgix.net/assets/open-closed-loop-darkmode-CYoXgW5C.png?auto=compress%2Cformat"/><img alt="Open vs Closed loop benchmark" src="https://planetscale-images.imgix.net/assets/open-closed-loop-B0_3fnSO.png?auto=compress%2Cformat" width="3172" height="1416" loading="lazy" class="w-auto max-w-full"/></picture><span aria-hidden="true" class="pointer-events-none absolute right-1 top-1 z-10 flex h-5 w-5 items-center justify-center border border-white/25 bg-black/70 text-white backdrop-blur-sm transition-colors transition-opacity group-hover:bg-black/90 group-hover:opacity-100 group-focus-visible:opacity-100 motion-reduce:transition-none [@media(hover:hover)_and_(pointer:fine)]:opacity-0"><svg width="14" height="14" viewBox="0 0 16 16" fill="none" stroke="currentColor" stroke-width="1.5" stroke-linecap="round" stroke-linejoin="round"><path d="M10 2h4v4M6 14H2v-4M14 2l-4.5 4.5M2 14l4.5-4.5"></path></svg></span></button><span hidden="" style="position:fixed;top:1px;left:1px;width:1px;height:0;padding:0;margin:-1px;overflow:hidden;clip:rect(0, 0, 0, 0);white-space:nowrap;border-width:0;display:none"></span></p><p>Open-loop benchmarks tend to be more realistic.<!-- --> <!-- -->In production systems, database load is applied at the rate that the clients demand, regardless of how well the database is keeping up.</p><p>Closed-loop benchmarks are more commonly seen in academic and performance comparisons, as they offer a more controlled environment for comparing things like QPS across a fixed amount of concurrency.</p><p>Both are beneficial, but they are useful for different things.<!-- --> <!-- -->Important to decide up front what the purpose of a benchmark is, then choose the type accordingly.</p><h2 id="what-to-measure"><a href="#what-to-measure">What to measure?</a></h2><p>Broadly, there are two things we like to measure when benchmarking: <em>throughput</em> and <em>latency</em>.<!-- --> <!-- -->Any good database benchmark will report on both of these things.</p><h2 id="throughput"><a href="#throughput">Throughput</a></h2><p><em>Throughput</em> is the amount of work completed in a slice of time.<!-- --> <!-- -->In databases, the most common measures are Queries Per Second (QPS) or Transactions Per Second (TPS).<!-- --> <!-- -->For many popular benchmarks like <a href="https://www.tpc.org/tpcc/">TPC-C</a> and <a href="https://www.tpc.org/tpch/">TPC-H</a>, <code>TPS &lt; QPS</code> because there are typically multiple queries within single transactions.<!-- --> <!-- -->Either works fine as a measure.</p><p>To measure throughput, choose a workload, a period of time to run it for (say, 5 minutes / 300 seconds), and then execute with TPS / QPS sampling.<!-- --> <!-- -->As a benchmark runs, samples are taken of how many queries or transactions complete each second.<!-- --> <!-- -->We then display this as a graph, showing every collected data point:</p><p><button type="button" aria-haspopup="dialog" aria-expanded="false" aria-label="Enlarge image: Throughput line chart" class="focus-visible-ring group relative block w-fit max-w-[min(100%,800px)] cursor-zoom-in text-left [&amp;_picture]:contents"><picture class="block"><source media="(prefers-color-scheme: light), (prefers-color-scheme: no-preference)" srcSet="https://planetscale-images.imgix.net/assets/throughput-lines-DG3lJ36Y.png?auto=compress%2Cformat"/><source media="(prefers-color-scheme: dark)" srcSet="https://planetscale-images.imgix.net/assets/throughput-lines-darkmode-B-LbVppT.png?auto=compress%2Cformat"/><img alt="Throughput line chart" src="https://planetscale-images.imgix.net/assets/throughput-lines-DG3lJ36Y.png?auto=compress%2Cformat" width="3172" height="1416" loading="lazy" class="w-auto max-w-full"/></picture><span aria-hidden="true" class="pointer-events-none absolute right-1 top-1 z-10 flex h-5 w-5 items-center justify-center border border-white/25 bg-black/70 text-white backdrop-blur-sm transition-colors transition-opacity group-hover:bg-black/90 group-hover:opacity-100 group-focus-visible:opacity-100 motion-reduce:transition-none [@media(hover:hover)_and_(pointer:fine)]:opacity-0"><svg width="14" height="14" viewBox="0 0 16 16" fill="none" stroke="currentColor" stroke-width="1.5" stroke-linecap="round" stroke-linejoin="round"><path d="M10 2h4v4M6 14H2v-4M14 2l-4.5 4.5M2 14l4.5-4.5"></path></svg></span></button><span hidden="" style="position:fixed;top:1px;left:1px;width:1px;height:0;padding:0;margin:-1px;overflow:hidden;clip:rect(0, 0, 0, 0);white-space:nowrap;border-width:0;display:none"></span></p><p>A more compact way of displaying this is via a bar chart with error bars.</p><p><button type="button" aria-haspopup="dialog" aria-expanded="false" aria-label="Enlarge image: Throughput bar chart" class="focus-visible-ring group relative block w-fit max-w-[min(100%,800px)] cursor-zoom-in text-left [&amp;_picture]:contents"><picture class="block"><source media="(prefers-color-scheme: light), (prefers-color-scheme: no-preference)" srcSet="https://planetscale-images.imgix.net/assets/throughput-bars-DAVyhM1x.png?auto=compress%2Cformat"/><source media="(prefers-color-scheme: dark)" srcSet="https://planetscale-images.imgix.net/assets/throughput-bars-darkmode-DRPfZNGr.png?auto=compress%2Cformat"/><img alt="Throughput bar chart" src="https://planetscale-images.imgix.net/assets/throughput-bars-DAVyhM1x.png?auto=compress%2Cformat" width="2000" height="1340" loading="lazy" class="w-auto max-w-full"/></picture><span aria-hidden="true" class="pointer-events-none absolute right-1 top-1 z-10 flex h-5 w-5 items-center justify-center border border-white/25 bg-black/70 text-white backdrop-blur-sm transition-colors transition-opacity group-hover:bg-black/90 group-hover:opacity-100 group-focus-visible:opacity-100 motion-reduce:transition-none [@media(hover:hover)_and_(pointer:fine)]:opacity-0"><svg width="14" height="14" viewBox="0 0 16 16" fill="none" stroke="currentColor" stroke-width="1.5" stroke-linecap="round" stroke-linejoin="round"><path d="M10 2h4v4M6 14H2v-4M14 2l-4.5 4.5M2 14l4.5-4.5"></path></svg></span></button><span hidden="" style="position:fixed;top:1px;left:1px;width:1px;height:0;padding:0;margin:-1px;overflow:hidden;clip:rect(0, 0, 0, 0);white-space:nowrap;border-width:0;display:none"></span></p><p>This communicates similar information in a more compact way, but it&#x27;s ideal to show a full line graph, as that also better visualizes inconsistencies or spikiness of performance throughout a benchmark run.<!-- --> <!-- -->More on this later.</p><p>Error bars are only one way to summarize variance.<!-- --> <a href="https://en.wikipedia.org/wiki/Coefficient_of_variation">Coefficient of variation</a>, <a href="https://en.wikipedia.org/wiki/Interquartile_range">interquartile range</a>, and <a href="https://en.wikipedia.org/wiki/Histogram">histograms</a> are different lenses on the same samples, each helping show whether a benchmark was stable, noisy, or hiding outliers.<!-- --> <!-- -->It&#x27;s helpful to include these or provide the data so readers can compute them themselves.</p><p>Throughput only tells half the story.</p><h2 id="latency"><a href="#latency">Latency</a></h2><p><em>Latency</em> is the amount of time it takes to complete an operation, query, or transaction.<!-- --> <!-- -->We can look at individual latencies (&quot;How long did this particular <code>SELECT * FROM...</code> take?&quot;), but more often we assess latencies in aggregate.</p><p>The standard language for communicating about latencies in distributed systems is with <em>percentiles</em> over some span of time (1 second, 1 minute, etc.).<!-- --> <!-- -->For example:</p><ul><li>p50 - The median latency.<!-- --> <!-- -->During this time period, half of the requests executed faster than this, the other half slower.</li><li>p90 - The 90th percentile.<!-- --> <!-- -->During this time period, 9 out of 10 requests executed faster, 1 out of 10 slower.</li><li>p99 - The 99th percentile.<!-- --> <!-- -->During this time period, 99 out of 100 requests executed faster, 1 out of 100 slower.</li></ul><p>We can measure any latency percentile we want, but these are the most common, along with p95 and p99.9.<!-- --> <!-- -->When benchmarking, we typically measure one or more of these in a series of small windows over the entire benchmark period.<!-- --> <!-- -->Say, sample p50, p90, and p99 once per second over a 5-minute (300-second) execution.<!-- --> <!-- -->Then, we plot the results.</p><p><button type="button" aria-haspopup="dialog" aria-expanded="false" aria-label="Enlarge image: Latency line chart" class="focus-visible-ring group relative block w-fit max-w-[min(100%,800px)] cursor-zoom-in text-left [&amp;_picture]:contents"><picture class="block"><source media="(prefers-color-scheme: light), (prefers-color-scheme: no-preference)" srcSet="https://planetscale-images.imgix.net/assets/latency-lines-BNBoL30i.png?auto=compress%2Cformat"/><source media="(prefers-color-scheme: dark)" srcSet="https://planetscale-images.imgix.net/assets/latency-lines-darkmode-BI3lGrIV.png?auto=compress%2Cformat"/><img alt="Latency line chart" src="https://planetscale-images.imgix.net/assets/latency-lines-BNBoL30i.png?auto=compress%2Cformat" width="3208" height="1416" loading="lazy" class="w-auto max-w-full"/></picture><span aria-hidden="true" class="pointer-events-none absolute right-1 top-1 z-10 flex h-5 w-5 items-center justify-center border border-white/25 bg-black/70 text-white backdrop-blur-sm transition-colors transition-opacity group-hover:bg-black/90 group-hover:opacity-100 group-focus-visible:opacity-100 motion-reduce:transition-none [@media(hover:hover)_and_(pointer:fine)]:opacity-0"><svg width="14" height="14" viewBox="0 0 16 16" fill="none" stroke="currentColor" stroke-width="1.5" stroke-linecap="round" stroke-linejoin="round"><path d="M10 2h4v4M6 14H2v-4M14 2l-4.5 4.5M2 14l4.5-4.5"></path></svg></span></button><span hidden="" style="position:fixed;top:1px;left:1px;width:1px;height:0;padding:0;margin:-1px;overflow:hidden;clip:rect(0, 0, 0, 0);white-space:nowrap;border-width:0;display:none"></span></p><p>In some cases, the line graphs are overkill.<!-- --> <!-- -->As with throughput, the visual can be compressed using a bar chart showing the median (or mean), with error bars.</p><p><button type="button" aria-haspopup="dialog" aria-expanded="false" aria-label="Enlarge image: Latency bar chart" class="focus-visible-ring group relative block w-fit max-w-[min(100%,800px)] cursor-zoom-in text-left [&amp;_picture]:contents"><picture class="block"><source media="(prefers-color-scheme: light), (prefers-color-scheme: no-preference)" srcSet="https://planetscale-images.imgix.net/assets/latency-bars-D6X_kFOI.png?auto=compress%2Cformat"/><source media="(prefers-color-scheme: dark)" srcSet="https://planetscale-images.imgix.net/assets/latency-bars-darkmode-DtAqZb6n.png?auto=compress%2Cformat"/><img alt="Latency bar chart" src="https://planetscale-images.imgix.net/assets/latency-bars-D6X_kFOI.png?auto=compress%2Cformat" width="2000" height="1380" loading="lazy" class="w-auto max-w-full"/></picture><span aria-hidden="true" class="pointer-events-none absolute right-1 top-1 z-10 flex h-5 w-5 items-center justify-center border border-white/25 bg-black/70 text-white backdrop-blur-sm transition-colors transition-opacity group-hover:bg-black/90 group-hover:opacity-100 group-focus-visible:opacity-100 motion-reduce:transition-none [@media(hover:hover)_and_(pointer:fine)]:opacity-0"><svg width="14" height="14" viewBox="0 0 16 16" fill="none" stroke="currentColor" stroke-width="1.5" stroke-linecap="round" stroke-linejoin="round"><path d="M10 2h4v4M6 14H2v-4M14 2l-4.5 4.5M2 14l4.5-4.5"></path></svg></span></button><span hidden="" style="position:fixed;top:1px;left:1px;width:1px;height:0;padding:0;margin:-1px;overflow:hidden;clip:rect(0, 0, 0, 0);white-space:nowrap;border-width:0;display:none"></span></p><p>We now have a way of communicating both <em>how much work</em> we accomplished and <em>how quickly</em> each unit of work was completed.</p><h2 id="warmup"><a href="#warmup">Warmup</a></h2><p>We&#x27;ve now settled the prep work and know <em>what</em> we should be measuring.<!-- --> <!-- -->Now let&#x27;s get tactical.<!-- --> <!-- -->How do we ensure that we are fair when running the benchmark?<!-- --> <!-- -->There&#x27;s a lot to consider for the executions themselves.</p><p>A big one is cache warmup.<!-- --> <!-- -->If we&#x27;ve recently booted up our database, the various caches are not full of pages (<code>buffer_cache</code> in Postgres, <code>buffer_pool</code> in MySQL).<!-- --> <!-- -->These require time and query load to warm, during which time latency and throughput will slowly be brought up to full potential.</p><p><button type="button" aria-haspopup="dialog" aria-expanded="false" aria-label="Enlarge image: Cache warming in databases" class="focus-visible-ring group relative block w-fit max-w-[min(100%,800px)] cursor-zoom-in text-left [&amp;_picture]:contents"><picture class="block"><source media="(prefers-color-scheme: light), (prefers-color-scheme: no-preference)" srcSet="https://planetscale-images.imgix.net/assets/warming-CmzsXHI3.png?auto=compress%2Cformat"/><source media="(prefers-color-scheme: dark)" srcSet="https://planetscale-images.imgix.net/assets/warming-darkmode-BPg8wwcI.png?auto=compress%2Cformat"/><img alt="Cache warming in databases" src="https://planetscale-images.imgix.net/assets/warming-CmzsXHI3.png?auto=compress%2Cformat" width="3172" height="1416" loading="lazy" class="w-auto max-w-full"/></picture><span aria-hidden="true" class="pointer-events-none absolute right-1 top-1 z-10 flex h-5 w-5 items-center justify-center border border-white/25 bg-black/70 text-white backdrop-blur-sm transition-colors transition-opacity group-hover:bg-black/90 group-hover:opacity-100 group-focus-visible:opacity-100 motion-reduce:transition-none [@media(hover:hover)_and_(pointer:fine)]:opacity-0"><svg width="14" height="14" viewBox="0 0 16 16" fill="none" stroke="currentColor" stroke-width="1.5" stroke-linecap="round" stroke-linejoin="round"><path d="M10 2h4v4M6 14H2v-4M14 2l-4.5 4.5M2 14l4.5-4.5"></path></svg></span></button><span hidden="" style="position:fixed;top:1px;left:1px;width:1px;height:0;padding:0;margin:-1px;overflow:hidden;clip:rect(0, 0, 0, 0);white-space:nowrap;border-width:0;display:none"></span></p><p>We typically run databases without measurement for a few minutes to ensure all caches are <em>warmed</em> before starting benchmark measurement.<!-- --> <!-- -->This ensures non-full caches and other startup costs don&#x27;t impact the numbers.</p><h2 id="configuration"><a href="#configuration">Configuration</a></h2><p>Even when warm, there are a number of configuration options that impact performance over long stretches of time.<!-- --> <!-- -->Though there are many, a good example of this is <code>checkpoint_timeout</code> in Postgres.</p><p>This and <code>max_wal_size</code> determine how frequently we need to flush table / index changes to disk (I/O checkpointing).<!-- --> <!-- -->If we set these to low / aggressive values, we may trigger it once every minute, causing regular performance dips.<!-- --> <!-- -->If we set it lax to only trigger once every ten minutes, we may not even notice it in the results of a 5-minute benchmark execution.</p><p><button type="button" aria-haspopup="dialog" aria-expanded="false" aria-label="Enlarge image: Checkpointing in database benchmarks" class="focus-visible-ring group relative block w-fit max-w-[min(100%,800px)] cursor-zoom-in text-left [&amp;_picture]:contents"><picture class="block"><source media="(prefers-color-scheme: light), (prefers-color-scheme: no-preference)" srcSet="https://planetscale-images.imgix.net/assets/checkpoints-DPQC9Z0v.png?auto=compress%2Cformat"/><source media="(prefers-color-scheme: dark)" srcSet="https://planetscale-images.imgix.net/assets/checkpoints-darkmode-BWJJBKrU.png?auto=compress%2Cformat"/><img alt="Checkpointing in database benchmarks" src="https://planetscale-images.imgix.net/assets/checkpoints-DPQC9Z0v.png?auto=compress%2Cformat" width="3172" height="1416" loading="lazy" class="w-auto max-w-full"/></picture><span aria-hidden="true" class="pointer-events-none absolute right-1 top-1 z-10 flex h-5 w-5 items-center justify-center border border-white/25 bg-black/70 text-white backdrop-blur-sm transition-colors transition-opacity group-hover:bg-black/90 group-hover:opacity-100 group-focus-visible:opacity-100 motion-reduce:transition-none [@media(hover:hover)_and_(pointer:fine)]:opacity-0"><svg width="14" height="14" viewBox="0 0 16 16" fill="none" stroke="currentColor" stroke-width="1.5" stroke-linecap="round" stroke-linejoin="round"><path d="M10 2h4v4M6 14H2v-4M14 2l-4.5 4.5M2 14l4.5-4.5"></path></svg></span></button><span hidden="" style="position:fixed;top:1px;left:1px;width:1px;height:0;padding:0;margin:-1px;overflow:hidden;clip:rect(0, 0, 0, 0);white-space:nowrap;border-width:0;display:none"></span></p><p>We can end up with graphs like this in these cases.<!-- --> <!-- -->But run for another 10 minutes, and we&#x27;d likely see a large performance dip on the green line.</p><p>Background jobs, I/O checkpointing, autovacuum, and other work can impact the throughput, skewing the benchmark results.</p><p>It&#x27;s important to consider the impact database configurations have on performance.<!-- --> <!-- -->An identical benchmark on the same hardware can perform very differently with different tunings.<!-- --> <!-- -->DBMSs give us these tunings so we can trade off things like performance, durability, data size, and resource consumption on a case-by-case basis.<!-- --> <!-- -->It&#x27;s generally best to either (a) ensure all configuration options are aligned or (b) for pre-tuned situations (like most database-as-a-service providers) leave things at the pre-tuned defaults.</p><h2 id="inconsistency"><a href="#inconsistency">(In)consistency</a></h2><p>Another important consideration, especially in the cloud, is (in)consistency.<!-- --> <!-- -->Even with the same benchmark instance and same client machine, latency and throughput can vary from run to run.<!-- --> <!-- -->This can be due to contention on the network or noisy neighbors that are co-occupying the same hardware you are running on.</p><p><button type="button" aria-haspopup="dialog" aria-expanded="false" aria-label="Enlarge image: Repeating the same benchmark" class="focus-visible-ring group relative block w-fit max-w-[min(100%,800px)] cursor-zoom-in text-left [&amp;_picture]:contents"><picture class="block"><source media="(prefers-color-scheme: light), (prefers-color-scheme: no-preference)" srcSet="https://planetscale-images.imgix.net/assets/repeating-DBCnml8x.png?auto=compress%2Cformat"/><source media="(prefers-color-scheme: dark)" srcSet="https://planetscale-images.imgix.net/assets/repeating-darkmode-BqI9nytL.png?auto=compress%2Cformat"/><img alt="Repeating the same benchmark" src="https://planetscale-images.imgix.net/assets/repeating-DBCnml8x.png?auto=compress%2Cformat" width="3172" height="1416" loading="lazy" class="w-auto max-w-full"/></picture><span aria-hidden="true" class="pointer-events-none absolute right-1 top-1 z-10 flex h-5 w-5 items-center justify-center border border-white/25 bg-black/70 text-white backdrop-blur-sm transition-colors transition-opacity group-hover:bg-black/90 group-hover:opacity-100 group-focus-visible:opacity-100 motion-reduce:transition-none [@media(hover:hover)_and_(pointer:fine)]:opacity-0"><svg width="14" height="14" viewBox="0 0 16 16" fill="none" stroke="currentColor" stroke-width="1.5" stroke-linecap="round" stroke-linejoin="round"><path d="M10 2h4v4M6 14H2v-4M14 2l-4.5 4.5M2 14l4.5-4.5"></path></svg></span></button><span hidden="" style="position:fixed;top:1px;left:1px;width:1px;height:0;padding:0;margin:-1px;overflow:hidden;clip:rect(0, 0, 0, 0);white-space:nowrap;border-width:0;display:none"></span></p><p>It&#x27;s advisable to do multiple runs to measure consistency.</p><h2 id="apples-to-apples-to-oranges"><a href="#apples-to-apples-to-oranges">Apples to apples to oranges</a></h2><p>The best benchmarks are the ones that compare apples-to-apples.<!-- --> <!-- -->In other words, ones that create data-driven comparisons between products that have the same or very similar characteristics and feature sets.</p><p>Examples of this are:</p><ul><li>Comparing 4 different Postgres configurations to determine workload suitability</li><li>Comparing 3 different cloud MySQL platforms to determine which is most performant</li><li>Comparing MySQL and Postgres on an identical workload (different databases, but same stated purpose)</li></ul><p>People sometimes draw comparisons between vastly different database engines, resulting in wild claims.<!-- --> <!-- -->Things like:</p><ul><li>Analytics queries run 100x faster on Apache Pinot than Postgres</li><li>Achieve 100x higher QPS on a purpose-built realtime database compared to a Postgres relational database</li><li>SQLite latency is 80% lower than MySQL</li></ul><p>These are comparing databases that were distinctly optimized for different purposes.<!-- --> <!-- -->It&#x27;s easy to make one look better than the other, especially when cherry-picking the workload.</p><p>Don&#x27;t do this.<!-- --> <!-- -->Ensure comparisons are between comparable technologies and workloads that fit the DBMS&#x27;s stated purpose.<!-- --> <!-- -->The one exception may be as an internal test to determine which technology, amongst ones with vastly different goals, is best-suited for a system.</p><h2 id="document-everything"><a href="#document-everything">Document everything</a></h2><p>Good benchmarks should be reproducible.<!-- --> <!-- -->Document the client and target setups as exhaustively as possible: hardware (or cloud instance type), OS, software versions, build flags, configurations, benchmark tool, exact command line, etc.<!-- --> <!-- -->After looking at the results of a benchmark, an engineer should be able to reproduce the results.</p><h2 id="benchmark-crimes"><a href="#benchmark-crimes">Benchmark crimes</a></h2><p>As you can see, there&#x27;s a lot to good benchmarking.<!-- --> <!-- -->Missing any one of these steps leads to bias.<!-- --> <!-- -->Some of the most common mistakes:</p><ul><li>Reporting only averages, without percentiles, variance, or the full time-series</li><li>Leaving out hardware, instance type, etc.</li><li>Measuring before the system reaches steady state</li><li>Reporting a percentage difference without the surrounding variance</li><li>Forgetting to check whether the benchmark client is the bottleneck</li></ul><p>That last one is easy to miss!</p><p>If the client machine has maxed out on CPU or network connections, the graph may look like the database has plateaued.<!-- --> <!-- -->But all you&#x27;ve really measured is the limit of the load generator.</p><h2 id="go-forth-and-benchmark"><a href="#go-forth-and-benchmark">Go forth and benchmark</a></h2><p>You now have an elementary understanding of database benchmarking.</p><p>When presenting results, don&#x27;t stop at the numbers.<!-- --> <!-- -->If two runs differ meaningfully, offer a hypothesis for why: hardware, configuration, workload shape, cache behavior, network latency, or something else.<!-- --> <!-- -->The reader should not have to invent the causal story themselves.</p><p>Apply all these to your next round of benchmarks, and you&#x27;re less likely to veer off-course.</p></div></article></div></section></main><footer class="mb-6 mt-10 px-3 sm:px-5 container max-w-7xl"><nav class="grid grid-cols-1 text-left sm:grid-cols-2 lg:grid-cols-5 lg:mx-7"><div class="dashed-box dashed-box-x-t sm:dashed-box-l-t lg:dashed-box-y-l p-3"><h2 class="font-semibold">Company</h2><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/about" data-discover="true">About</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/brand" data-discover="true">Brand</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/blog" data-discover="true">Blog</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/changelog" data-discover="true">Changelog</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/careers" data-discover="true">Careers</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/events" data-discover="true">Events</a></div><div class="dashed-box dashed-box-x-t lg:dashed-box-y-l p-3"><h2 class="font-semibold">Product</h2><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/case-studies" data-discover="true">Case studies</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/enterprise" data-discover="true">Enterprise</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/pricing" data-discover="true">Pricing</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/benchmarks" data-discover="true">Benchmarks</a></div><div class="dashed-box dashed-box-x-t sm:dashed-box-l-t lg:dashed-box-y-l p-3"><h2 class="font-semibold">Resources</h2><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/docs">Documentation</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/migrate" data-discover="true">Migrate</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="https://support.planetscale.com/hc/en-us" rel="nofollow noopener noreferrer" target="_blank">Support</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="https://planetscalestatus.com" rel="nofollow noopener noreferrer" target="_blank">Status</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="https://trust.planetscale.com" rel="nofollow noopener noreferrer" target="_blank">Trust Center</a></div><div class="dashed-box dashed-box-x-t lg:dashed-box-y-l p-3"><h2 class="font-semibold">Courses</h2><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/learn/courses/mysql-for-developers" data-discover="true">MySQL for Developers</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/learn/courses/database-scaling" data-discover="true">Database Scaling</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/learn/courses/vitess" data-discover="true">Learn Vitess</a></div><div class="dashed-box p-3 sm:col-span-2 lg:col-span-1"><h2 class="font-semibold text-primary hover:text-contrast">Open source</h2><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/vitess" data-discover="true">Vitess</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="https://vitess.io/slack" rel="nofollow noopener noreferrer" target="_blank">Vitess community</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="https://github.com/planetscale" rel="me nofollow noopener noreferrer" target="_blank">GitHub</a></div></nav><div class="dashed-box dashed-box-x-b p-3 lg:mx-7"><p class="mb-3 md:mb-0"><a class="text-primary" rel="nofollow" href="/legal/privacy" data-discover="true">Privacy</a><span class="text-decoration" role="presentation"> <!-- -->|<!-- --> </span><a class="text-primary" rel="nofollow" href="/legal/siteterms" data-discover="true">Terms</a><span class="text-decoration" role="presentation"> <!-- -->|<!-- --> </span><a class="text-primary" rel="nofollow" href="/legal/cookies" data-discover="true">Cookies</a><span class="text-decoration" role="presentation"> <!-- -->|<!-- --> </span><a class="text-primary" rel="nofollow" href="/legal/patents" data-discover="true">Patents</a><span class="text-decoration" role="presentation"> <!-- -->|<!-- --> </span><a class="text-primary" rel="nofollow" href="/legal/privacy#privacy-rights-and-choices" data-discover="true">Do Not Share My Personal Information</a></p><p class="text-secondary">© <!-- -->2026<!-- --> PlanetScale, Inc. All rights reserved.</p></div><p class="mb-0 mt-3 break-normal lg:mx-7"><a class="text-primary" href="https://github.com/planetscale" rel="me nofollow noopener noreferrer" target="_blank">GitHub</a><span class="text-decoration" role="presentation"> <!-- -->|<!-- --> </span><a aria-label="X (formerly Twitter)" class="text-primary" href="https://twitter.com/planetscale" rel="me nofollow noopener noreferrer" target="_blank">X</a><span class="text-decoration" role="presentation"> <!-- -->|<!-- --> </span><a aria-label="LinkedIn" class="text-primary" href="https://www.linkedin.com/company/planetscale" target="_blank" rel="noreferrer">LinkedIn</a><span class="text-decoration" role="presentation"> <!-- -->|<!-- --> </span><a class="text-primary" href="https://www.youtube.com/planetscale" rel="me nofollow noopener noreferrer" target="_blank">YouTube</a><span class="text-decoration" role="presentation"> <!-- -->|<!-- --> </span><a aria-label="Discord" class="text-primary" href="https://pscale.link/community" rel="nofollow noopener noreferrer" target="_blank">Discord</a><span class="text-decoration" role="presentation"> <!-- -->|<!-- --> </span><a class="text-primary" href="https://www.facebook.com/planetscaledata" rel="me nofollow noopener noreferrer" target="_blank">Facebook</a></p></footer><script nonce="1NjQdEMwrEHFpdD1zVsfJP6ksuIJ/ecPax+EDqlkDLg=">((storageKey2, restoreKey) => {
if (!window.history.state || !window.history.state.key) {
let key2 = Math.random().toString(32).slice(2);
window.history.replaceState({ key: key2 }, "");
}
try {
let storedY = JSON.parse(sessionStorage.getItem(storageKey2) || "{}")[restoreKey || window.history.state.key];
if (typeof storedY === "number") window.scrollTo(0, storedY);
} catch (error2) {
console.error(error2);
sessionStorage.removeItem(storageKey2);
}
})("react-router-scroll-positions", null)</script><script nonce="1NjQdEMwrEHFpdD1zVsfJP6ksuIJ/ecPax+EDqlkDLg=">window.__reactRouterContext = {"basename":"/","future":{"unstable_enableNodeReadableStream":false,"unstable_optimizeDeps":true},"routeDiscovery":{"mode":"lazy","manifestPath":"/__manifest"},"ssr":true,"isSpaMode":false};window.__reactRouterContext.stream = new ReadableStream({start(controller){window.__reactRouterContext.streamController = controller;}}).pipeThrough(new TextEncoderStream());</script><script nonce="1NjQdEMwrEHFpdD1zVsfJP6ksuIJ/ecPax+EDqlkDLg=" type="module" async="">;
import * as route0 from "/assets/root-DbOv4-98.js";
import * as route1 from "/assets/blog-pXH7ptHJ.js";
import * as route2 from "/assets/blog._slug-Ch_92qsH.js";
window.__reactRouterManifest = {
"entry": {
"module": "/assets/entry.client-3vubyXrk.js",
"imports": [
"/assets/jsx-runtime-DwfQwkRq.js",
"/assets/components-_bNmAApg.js",
"/assets/_.well-known_.mcp.server-card_.json_-Sx7XeH3e.js",
"/assets/index-mKTXLmHu.js",
"/assets/errorBoundaries-DhW4jVYt.js"
],
"css": []
},
"routes": {
"root": {
"id": "root",
"path": "",
"hasAction": false,
"hasLoader": true,
"hasClientAction": false,
"hasClientLoader": false,
"hasClientMiddleware": false,
"hasDefaultExport": true,
"hasErrorBoundary": true,
"module": "/assets/root-DbOv4-98.js",
"imports": [
"/assets/jsx-runtime-DwfQwkRq.js",
"/assets/components-_bNmAApg.js",
"/assets/_.well-known_.mcp.server-card_.json_-Sx7XeH3e.js",
"/assets/index-mKTXLmHu.js",
"/assets/errorBoundaries-DhW4jVYt.js",
"/assets/lib-Dg89tQ22.js",
"/assets/analytics.client-DM6E8o1h.js",
"/assets/SiteHeader-C2U5gvDH.js",
"/assets/current-9yDxj94E.js",
"/assets/clsx-eT0YPcGk.js",
"/assets/bugs-38ilEoW0.js",
"/assets/keyboard-D-uXZORL.js",
"/assets/use-tab-direction-dKm-S3Ck.js"
],
"css": []
},
"routes/blog": {
"id": "routes/blog",
"parentId": "root",
"path": "blog",
"hasAction": false,
"hasLoader": false,
"hasClientAction": false,
"hasClientLoader": false,
"hasClientMiddleware": false,
"hasDefaultExport": false,
"hasErrorBoundary": false,
"module": "/assets/blog-pXH7ptHJ.js",
"imports": [
"/assets/_.well-known_.mcp.server-card_.json_-Sx7XeH3e.js"
],
"css": []
},
"routes/blog.$slug": {
"id": "routes/blog.$slug",
"parentId": "routes/blog",
"path": ":slug",
"hasAction": false,
"hasLoader": true,
"hasClientAction": false,
"hasClientLoader": false,
"hasClientMiddleware": false,
"hasDefaultExport": true,
"hasErrorBoundary": false,
"module": "/assets/blog._slug-Ch_92qsH.js",
"imports": [
"/assets/components-_bNmAApg.js",
"/assets/lib-Dg89tQ22.js",
"/assets/jsx-runtime-DwfQwkRq.js",
"/assets/ContentImage-Dh6VEOUl.js",
"/assets/clsx-eT0YPcGk.js",
"/assets/BlogCategoryLink-DmQyn0gp.js",
"/assets/_.well-known_.mcp.server-card_.json_-Sx7XeH3e.js",
"/assets/Details-BSB_b6hI.js",
"/assets/Skittle-CDFOPRjH.js",
"/assets/SiteFooter-B2Gq9u2j.js",
"/assets/SiteHeader-C2U5gvDH.js",
"/assets/Vimeo-00PQJDli.js",
"/assets/YouTube-CMfaljVr.js",
"/assets/date-CJTFH3uT.js",
"/assets/errorBoundaries-DhW4jVYt.js",
"/assets/keyboard-D-uXZORL.js",
"/assets/use-tab-direction-dKm-S3Ck.js",
"/assets/index-mKTXLmHu.js",
"/assets/use-inert-others-BMJ6-xOX.js",
"/assets/description-Cf6FZmDe.js",
"/assets/use-is-mounted-uQsUZyP9.js",
"/assets/types-DvonrUFF.js",
"/assets/current-9yDxj94E.js",
"/assets/analytics.client-DM6E8o1h.js",
"/assets/bugs-38ilEoW0.js"
],
"css": []
},
"routes/_index": {
"id": "routes/_index",
"parentId": "root",
"index": true,
"hasAction": false,
"hasLoader": true,
"hasClientAction": false,
"hasClientLoader": false,
"hasClientMiddleware": false,
"hasDefaultExport": true,
"hasErrorBoundary": false,
"module": "/assets/_index-BfA6EnlR.js",
"imports": [
"/assets/components-_bNmAApg.js",
"/assets/lib-Dg89tQ22.js",
"/assets/jsx-runtime-DwfQwkRq.js",
"/assets/Logo-Gm9TLYAs.js",
"/assets/SiteFooter-B2Gq9u2j.js",
"/assets/SiteHeader-C2U5gvDH.js",
"/assets/_.well-known_.mcp.server-card_.json_-Sx7XeH3e.js",
"/assets/bugs-38ilEoW0.js",
"/assets/keyboard-D-uXZORL.js",
"/assets/use-is-mounted-uQsUZyP9.js",
"/assets/use-tab-direction-dKm-S3Ck.js",
"/assets/errorBoundaries-DhW4jVYt.js",
"/assets/clsx-eT0YPcGk.js",
"/assets/current-9yDxj94E.js",
"/assets/analytics.client-DM6E8o1h.js",
"/assets/index-mKTXLmHu.js"
],
"css": []
},
"routes/blog._index": {
"id": "routes/blog._index",
"parentId": "routes/blog",
"index": true,
"hasAction": false,
"hasLoader": true,
"hasClientAction": false,
"hasClientLoader": false,
"hasClientMiddleware": false,
"hasDefaultExport": true,
"hasErrorBoundary": false,
"module": "/assets/blog._index-DcvTTuDd.js",
"imports": [
"/assets/components-_bNmAApg.js",
"/assets/jsx-runtime-DwfQwkRq.js",
"/assets/social-Cd2AtOZM.js",
"/assets/BlogCategoryLink-DmQyn0gp.js",
"/assets/BlogPostLink-DC1SPKBJ.js",
"/assets/BlogCategoryNav-CB7TJ3IB.js",
"/assets/Paginator-xlPA_JNt.js",
"/assets/SiteFooter-B2Gq9u2j.js",
"/assets/SiteHeader-C2U5gvDH.js",
"/assets/date-CJTFH3uT.js",
"/assets/_.well-known_.mcp.server-card_.json_-Sx7XeH3e.js",
"/assets/lib-Dg89tQ22.js",
"/assets/errorBoundaries-DhW4jVYt.js",
"/assets/clsx-eT0YPcGk.js",
"/assets/types-DvonrUFF.js",
"/assets/enumerator-2YLGh-nT.js",
"/assets/current-9yDxj94E.js",
"/assets/analytics.client-DM6E8o1h.js",
"/assets/bugs-38ilEoW0.js",
"/assets/keyboard-D-uXZORL.js",
"/assets/use-tab-direction-dKm-S3Ck.js",
"/assets/index-mKTXLmHu.js"
],
"css": []
}
},
"url": "/assets/manifest-e17deb94.js",
"version": "e17deb94"
};
window.__reactRouterRouteModules = {"root":route0,"routes/blog":route1,"routes/blog.$slug":route2};
import("/assets/entry.client-3vubyXrk.js");</script><script type="application/ld+json" nonce="1NjQdEMwrEHFpdD1zVsfJP6ksuIJ/ecPax+EDqlkDLg=">{"@context":"https://schema.org","@type":"Organization","name":"PlanetScale, Inc.","url":"https://planetscale.com","sameAs":["https://twitter.com/PlanetScale","https://www.facebook.com/planetscaledata/","https://www.instagram.com/planetscale/"],"address":{"@type":"PostalAddress","streetAddress":"WeWork c/o PlanetScale, 535 Mission Street, 14th Floor","addressLocality":"San Francisco","addressRegion":"CA","postalCode":"94105","addressCountry":"US"}}</script><!--$--><script nonce="1NjQdEMwrEHFpdD1zVsfJP6ksuIJ/ecPax+EDqlkDLg=">window.__reactRouterContext.streamController.enqueue("[{\"_1\":2,\"_3\":-5,\"_4\":-5},\"loaderData\",{\"_5\":6,\"_7\":8},\"actionData\",\"errors\",\"root\",{\"_1238\":1239},\"routes/blog.$slug\",{\"_9\":10,\"_5\":11},\"blog\",{\"_12\":13,\"_14\":15,\"_16\":-7,\"_17\":18,\"_19\":20,\"_21\":22,\"_23\":24,\"_25\":26,\"_27\":28,\"_29\":30,\"_31\":32},\"https://planetscale.com\",\"body\",[105,106,107,108,109,110,111,112,113,114,115,116,117,118,119,120,121,122,123,124,125,126,127,128,129,130,131,132,133,134,135,136,137,138,139,140,141,142,143,144,145,146,147,148,149,150,151,152,153,154,155,156,157,158,159,160,161,162,163,164,165,166,167,168,169,170,171,172,173,174,175,176,177,178,179,180,181,182,183,184,185,186,187,188,189,190,191,192,193,194,195,196,197,198,199,200,201,202,203,204,205,206,207,208,209,210,211],\"body_text\",\"Benchmarking is hard.There are many ways to do it wrong and few to do it right.\\nBut zooming out from any single system or harness, there are broad principles that should be applied to all benchmarking.Using these correctly makes it difficult to produce biased results.\\nAm I the world's best benchmarker?Certainly not.I invented the language balls, after all.But correctness and precision are important parts of PlanetScale's culture.We've spent considerable time learning the art of benchmarking, and are here to share best-practices.\\nHere, we're focusing primarily on benchmarking databases, but these principles apply to many domains.\\nClient-server architecture\\nDatabases typically operate in a client-server model.The database server is started, accepts connections from clients, executes queries, and returns results.\\nTo benchmark, we need a client that establishes the connections, generates queries, and takes measurements.Since both sides consume resources and we want to give the database its full share of the host server, it's common to set up a distinct server for benchmark execution.\\nAs usual, there's a catch.This introduces latency between the two machines.\\n\\nHow much this skews the results of the benchmark depends quite a bit on how \\\"far apart\\\" the benchmark server and database server are(network latency)and how long the queries / transactions take on the database(execution latency).\\n\\nLet's consider a scenario where each query takes ~10ms to execute on the database.If the network round-trip time is 2.5 milliseconds, then we can execute approximately 80 queries per second over a single connection.On the other hand, what if the round-trip is 15 milliseconds?We've now cut our single-threaded QPS capability in ~half, resulting in 40 QPS.\\n\\nSame database.Same benchmark client.The only difference is the speed at which bytes can go over the wire between the two.\\nThis latency variation will always have an impact on latency measurements.\\nIt can also impact throughput.We often don't run benchmarks on a single connection.We'll do 10, 50, or 100 simultaneous connections to best utilize the parallelism of the machine and database.But if we have a fixed connection count, and are not making it dynamic to account for round-trip latency, we can end up allowing the elevated latency to hurt throughput.\\nFinally, you should double-check that the client server is not a bottleneck.While benchmarking, ensure that CPU and network utilization are well under their capacity.We want to be straining the database server, not the client.\\nChoosing resources\\nIt's easy to make one database look better than another with an imbalance of resources.Postgres running on a 16-core server will almost always perform better than on an 8-core server.\\nAn important prerequisite to proper benchmarking is setting up the compute, storage, and networking resources to allow for a fair fight.\\nThis isn't as easy as it sounds, especially when we're talking about running things in the hyperscaler clouds like AWS and GCP.For example, the Geekbench results for an AWS r7g.2xlarge are ~15% lower than the results for an r8g.2xlarge.Both have 8 vCPUs and 64 GB RAM.But move one generation newer, and there's a ~15% CPU improvement.\\n\\nYou might then be tempted to just use the same instance for everything, but this breaks down too.The availability of instance types varies over time, region, and database provider.In some cases, it's not possible to match.\\nIn an ideal world, we'd run everything on the exact same instance.In reality, we sometimes have to settle for matching CPUs and RAM as best we can, and living with the differences.However, you must give this your best effort.Purposefully choosing to benchmark your product on 2025-gen CPU and then comparing to a competitor's product on a 2022 CPU, when the alternate was readily available, is intentionally misleading.\\nWorkload\\nEven once we know that our infrastructure is set up sanely, there's a lot to consider for the workload we run.\\nThe easiest way to think about this is in terms of traffic ratios.\\nHow many queries are hitting RAM vs disk?\\nWhat % of the data is hot (frequently queried) vs cold (rarely queried)?\\nWhat's the ratio of reads to writes?\\nAll of these impact performance, especially when combined with the variations of underlying hardware.\\nQueries executed on a relational database often require some amount of I/O work.Writing data must always be persisted to disk.Reading data can come from the in-memory cache, or disk on cache misses.\\n\\nSome databases operate on local SSDs, while others use network-attached storage like AWS EBS or Google Persistent Disk.Some even take a hybrid approach.Either way, the percent of read traffic hitting RAM vs disk impacts performance due to I/O wait times.\\nConsider a benchmark like sysbench OLTP read-only.This is a simple, read-only benchmark that runs a handful of select query patterns repeatedly.As benchmarks often do, the data size is configurable in the preparation phase.If we run this benchmark on a server with 64 GB of RAM and a 32 GB data size, the entire data set will fit in RAM after warming.The same benchmark run with a 320 GB data size will generate significant I/O and inevitably run slower.\\nThis is related to, but not the same as, data distribution.\\nEven for a fixed data size, access patterns can vary widely.The simplest examples are uniform and Zipfian.\\n\\nA uniform access pattern gives every row the same chance of being queried on each request.If we have 100 rows, each has a 1% chance of being read for each operation.\\nA Zipfian access pattern is skewed: the k-th most popular key is accessed roughly proportional to 1/k.A small number of hot rows receive a large share of requests, while most rows are accessed rarely.\\nThese are only simple models.Real workloads often have messier shapes: recently inserted rows might be hotter than old rows, one tenant might dominate traffic, or a small working set might receive most reads for a period of time.\\nWhich pattern the benchmark operates with significantly impacts performance, because it in turn impacts how frequently we need to access disk vs RAM and the amount of cache churn.\\nClosed and open loop\\nThere are two types of benchmark workload shapes: open and closed loops.\\nIn a closed-loop benchmark, the client sends requests and then waits for a response before sending the next.while True:\\n # wait for response\\n response = send_bench_request()\\n # then send next\\n process(response)\\n\\nWe may do this in parallel across many connections, but each individual connection sends a controlled sequence of queries.A closed loop can also hide a failure mode called coordinated omission: when the database stalls, the client stops issuing new requests too, so the benchmark only records the stalled request and omits the work that would have queued behind it.This is especially misleading for tail latency, where the missing queued requests are exactly the ones that would have made p95/p99 look worse (more on latency and percentiles soon).\\nOpen loop on the other hand has a fixed pace of sending requests, regardless of how quickly the database responds.while True:\\n # fire and forget\\n send_bench_request()\\n # fixed pace\\n time.sleep(0.1)\\n\\nThis can be fixed throughout the entire benchmark duration, or vary in a controlled way:\\n\\nOpen-loop benchmarks tend to be more realistic.In production systems, database load is applied at the rate that the clients demand, regardless of how well the database is keeping up.\\nClosed-loop benchmarks are more commonly seen in academic and performance comparisons, as they offer a more controlled environment for comparing things like QPS across a fixed amount of concurrency.\\nBoth are beneficial, but they are useful for different things.Important to decide up front what the purpose of a benchmark is, then choose the type accordingly.\\nWhat to measure?\\nBroadly, there are two things we like to measure when benchmarking: throughput and latency.Any good database benchmark will report on both of these things.\\nThroughput\\nThroughput is the amount of work completed in a slice of time.In databases, the most common measures are Queries Per Second (QPS) or Transactions Per Second (TPS).For many popular benchmarks like TPC-C and TPC-H, TPS \u003c QPS because there are typically multiple queries within single transactions.Either works fine as a measure.\\nTo measure throughput, choose a workload, a period of time to run it for (say, 5 minutes / 300 seconds), and then execute with TPS / QPS sampling.As a benchmark runs, samples are taken of how many queries or transactions complete each second.We then display this as a graph, showing every collected data point:\\n\\nA more compact way of displaying this is via a bar chart with error bars.\\n\\nThis communicates similar information in a more compact way, but it's ideal to show a full line graph, as that also better visualizes inconsistencies or spikiness of performance throughout a benchmark run.More on this later.\\nError bars are only one way to summarize variance.Coefficient of variation, interquartile range, and histograms are different lenses on the same samples, each helping show whether a benchmark was stable, noisy, or hiding outliers.It's helpful to include these or provide the data so readers can compute them themselves.\\nThroughput only tells half the story.\\nLatency\\nLatency is the amount of time it takes to complete an operation, query, or transaction.We can look at individual latencies (\\\"How long did this particular SELECT * FROM... take?\\\"), but more often we assess latencies in aggregate.\\nThe standard language for communicating about latencies in distributed systems is with percentiles over some span of time (1 second, 1 minute, etc.).For example:\\np50 - The median latency.During this time period, half of the requests executed faster than this, the other half slower.\\np90 - The 90th percentile.During this time period, 9 out of 10 requests executed faster, 1 out of 10 slower.\\np99 - The 99th percentile.During this time period, 99 out of 100 requests executed faster, 1 out of 100 slower.\\nWe can measure any latency percentile we want, but these are the most common, along with p95 and p99.9.When benchmarking, we typically measure one or more of these in a series of small windows over the entire benchmark period.Say, sample p50, p90, and p99 once per second over a 5-minute (300-second) execution.Then, we plot the results.\\n\\nIn some cases, the line graphs are overkill.As with throughput, the visual can be compressed using a bar chart showing the median (or mean), with error bars.\\n\\nWe now have a way of communicating both how much work we accomplished and how quickly each unit of work was completed.\\nWarmup\\nWe've now settled the prep work and know what we should be measuring.Now let's get tactical.How do we ensure that we are fair when running the benchmark?There's a lot to consider for the executions themselves.\\nA big one is cache warmup.If we've recently booted up our database, the various caches are not full of pages (buffer_cache in Postgres, buffer_pool in MySQL).These require time and query load to warm, during which time latency and throughput will slowly be brought up to full potential.\\n\\nWe typically run databases without measurement for a few minutes to ensure all caches are warmed before starting benchmark measurement.This ensures non-full caches and other startup costs don't impact the numbers.\\nConfiguration\\nEven when warm, there are a number of configuration options that impact performance over long stretches of time.Though there are many, a good example of this is checkpoint_timeout in Postgres.\\nThis and max_wal_size determine how frequently we need to flush table / index changes to disk (I/O checkpointing).If we set these to low / aggressive values, we may trigger it once every minute, causing regular performance dips.If we set it lax to only trigger once every ten minutes, we may not even notice it in the results of a 5-minute benchmark execution.\\n\\nWe can end up with graphs like this in these cases.But run for another 10 minutes, and we'd likely see a large performance dip on the green line.\\nBackground jobs, I/O checkpointing, autovacuum, and other work can impact the throughput, skewing the benchmark results.\\nIt's important to consider the impact database configurations have on performance.An identical benchmark on the same hardware can perform very differently with different tunings.DBMSs give us these tunings so we can trade off things like performance, durability, data size, and resource consumption on a case-by-case basis.It's generally best to either (a) ensure all configuration options are aligned or (b) for pre-tuned situations (like most database-as-a-service providers) leave things at the pre-tuned defaults.\\n(In)consistency\\nAnother important consideration, especially in the cloud, is (in)consistency.Even with the same benchmark instance and same client machine, latency and throughput can vary from run to run.This can be due to contention on the network or noisy neighbors that are co-occupying the same hardware you are running on.\\n\\nIt's advisable to do multiple runs to measure consistency.\\nApples to apples to oranges\\nThe best benchmarks are the ones that compare apples-to-apples.In other words, ones that create data-driven comparisons between products that have the same or very similar characteristics and feature sets.\\nExamples of this are:\\nComparing 4 different Postgres configurations to determine workload suitability\\nComparing 3 different cloud MySQL platforms to determine which is most performant\\nComparing MySQL and Postgres on an identical workload (different databases, but same stated purpose)\\nPeople sometimes draw comparisons between vastly different database engines, resulting in wild claims.Things like:\\nAnalytics queries run 100x faster on Apache Pinot than Postgres\\nAchieve 100x higher QPS on a purpose-built realtime database compared to a Postgres relational database\\nSQLite latency is 80% lower than MySQL\\nThese are comparing databases that were distinctly optimized for different purposes.It's easy to make one look better than the other, especially when cherry-picking the workload.\\nDon't do this.Ensure comparisons are between comparable technologies and workloads that fit the DBMS's stated purpose.The one exception may be as an internal test to determine which technology, amongst ones with vastly different goals, is best-suited for a system.\\nDocument everything\\nGood benchmarks should be reproducible.Document the client and target setups as exhaustively as possible: hardware (or cloud instance type), OS, software versions, build flags, configurations, benchmark tool, exact command line, etc.After looking at the results of a benchmark, an engineer should be able to reproduce the results.\\nBenchmark crimes\\nAs you can see, there's a lot to good benchmarking.Missing any one of these steps leads to bias.Some of the most common mistakes:\\nReporting only averages, without percentiles, variance, or the full time-series\\nLeaving out hardware, instance type, etc.\\nMeasuring before the system reaches steady state\\nReporting a percentage difference without the surrounding variance\\nForgetting to check whether the benchmark client is the bottleneck\\nThat last one is easy to miss!\\nIf the client machine has maxed out on CPU or network connections, the graph may look like the database has plateaued.But all you've really measured is the limit of the load generator.\\nGo forth and benchmark\\nYou now have an elementary understanding of database benchmarking.\\nWhen presenting results, don't stop at the numbers.If two runs differ meaningfully, offer a hypothesis for why: hardware, configuration, workload shape, cache behavior, network latency, or something else.The reader should not have to invent the causal story themselves.\\nApply all these to your next round of benchmarks, and you're less likely to veer off-course.\",\"aside\",\"toc\",[45,46,47,48,49,50,51,52,53,54,55,56,57,58],\"title\",\"On benchmarking\",\"authors\",[39],\"categories\",[38],\"excerpt\",\"Benchmarking is hard. Done wrong it is very misleading, and unfortunately it is frequently done wrong. Let's explore how not to make silly mistakes.\",\"createdAt\",\"2026-05-05\",\"slug\",\"on-benchmarking\",\"meta\",{\"_33\":34,\"_35\":26,\"_36\":37,\"_19\":20},\"canonical\",\"https://planetscale.com/blog/on-benchmarking\",\"description\",\"image\",\"/assets/on-benchmarking-social-JmCxOY2L.png\",\"engineering\",{\"_29\":40,\"_41\":42,\"_43\":44},\"ben\",\"name\",\"Ben Dicken\",\"x\",\"BenjDicken\",{\"_59\":102,\"_61\":103,\"_63\":64,\"_19\":104},{\"_59\":99,\"_61\":100,\"_63\":64,\"_19\":101},{\"_59\":96,\"_61\":97,\"_63\":64,\"_19\":98},{\"_59\":93,\"_61\":94,\"_63\":64,\"_19\":95},{\"_59\":90,\"_61\":91,\"_63\":64,\"_19\":92},{\"_59\":87,\"_61\":88,\"_63\":64,\"_19\":89},{\"_59\":84,\"_61\":85,\"_63\":64,\"_19\":86},{\"_59\":81,\"_61\":82,\"_63\":64,\"_19\":83},{\"_59\":78,\"_61\":79,\"_63\":64,\"_19\":80},{\"_59\":75,\"_61\":76,\"_63\":64,\"_19\":77},{\"_59\":72,\"_61\":73,\"_63\":64,\"_19\":74},{\"_59\":69,\"_61\":70,\"_63\":64,\"_19\":71},{\"_59\":66,\"_61\":67,\"_63\":64,\"_19\":68},{\"_59\":60,\"_61\":62,\"_63\":64,\"_19\":65},\"children\",[],\"id\",\"go-forth-and-benchmark\",\"level\",2,\"Go forth and benchmark\",[],\"benchmark-crimes\",\"Benchmark crimes\",[],\"document-everything\",\"Document everything\",[],\"apples-to-apples-to-oranges\",\"Apples to apples to oranges\",[],\"inconsistency\",\"(In)consistency\",[],\"configuration\",\"Configuration\",[],\"warmup\",\"Warmup\",[],\"latency\",\"Latency\",[],\"throughput\",\"Throughput\",[],\"what-to-measure\",\"What to measure?\",[],\"closed-and-open-loop\",\"Closed and open loop\",[],\"workload\",\"Workload\",[],\"choosing-resources\",\"Choosing resources\",[],\"client-server-architecture\",\"Client-server architecture\",[\"SingleFetchClassInstance\",1233],[\"SingleFetchClassInstance\",1228],[\"SingleFetchClassInstance\",1213],[\"SingleFetchClassInstance\",1203],[\"SingleFetchClassInstance\",1195],[\"SingleFetchClassInstance\",1190],[\"SingleFetchClassInstance\",1179],[\"SingleFetchClassInstance\",1169],[\"SingleFetchClassInstance\",1155],[\"SingleFetchClassInstance\",1148],[\"SingleFetchClassInstance\",1133],[\"SingleFetchClassInstance\",1126],[\"SingleFetchClassInstance\",1111],[\"SingleFetchClassInstance\",1105],[\"SingleFetchClassInstance\",1101],[\"SingleFetchClassInstance\",1088],[\"SingleFetchClassInstance\",1082],[\"SingleFetchClassInstance\",1074],[\"SingleFetchClassInstance\",1069],[\"SingleFetchClassInstance\",1065],[\"SingleFetchClassInstance\",1045],[\"SingleFetchClassInstance\",1031],[\"SingleFetchClassInstance\",1025],[\"SingleFetchClassInstance\",1006],[\"SingleFetchClassInstance\",998],[\"SingleFetchClassInstance\",994],[\"SingleFetchClassInstance\",990],[\"SingleFetchClassInstance\",972],[\"SingleFetchClassInstance\",968],[\"SingleFetchClassInstance\",957],[\"SingleFetchClassInstance\",942],[\"SingleFetchClassInstance\",936],[\"SingleFetchClassInstance\",922],[\"SingleFetchClassInstance\",918],[\"SingleFetchClassInstance\",913],[\"SingleFetchClassInstance\",898],[\"SingleFetchClassInstance\",888],[\"SingleFetchClassInstance\",872],[\"SingleFetchClassInstance\",867],[\"SingleFetchClassInstance\",863],[\"SingleFetchClassInstance\",855],[\"SingleFetchClassInstance\",851],[\"SingleFetchClassInstance\",847],[\"SingleFetchClassInstance\",843],[\"SingleFetchClassInstance\",837],[\"SingleFetchClassInstance\",833],[\"SingleFetchClassInstance\",823],[\"SingleFetchClassInstance\",819],[\"SingleFetchClassInstance\",806],[\"SingleFetchClassInstance\",801],[\"SingleFetchClassInstance\",797],[\"SingleFetchClassInstance\",792],[\"SingleFetchClassInstance\",784],[\"SingleFetchClassInstance\",770],[\"SingleFetchClassInstance\",762],[\"SingleFetchClassInstance\",732],[\"SingleFetchClassInstance\",726],[\"SingleFetchClassInstance\",713],[\"SingleFetchClassInstance\",709],[\"SingleFetchClassInstance\",695],[\"SingleFetchClassInstance\",690],[\"SingleFetchClassInstance\",664],[\"SingleFetchClassInstance\",660],[\"SingleFetchClassInstance\",652],[\"SingleFetchClassInstance\",637],[\"SingleFetchClassInstance\",626],[\"SingleFetchClassInstance\",605],[\"SingleFetchClassInstance\",598],[\"SingleFetchClassInstance\",584],[\"SingleFetchClassInstance\",579],[\"SingleFetchClassInstance\",564],[\"SingleFetchClassInstance\",548],[\"SingleFetchClassInstance\",540],[\"SingleFetchClassInstance\",527],[\"SingleFetchClassInstance\",509],[\"SingleFetchClassInstance\",496],[\"SingleFetchClassInstance\",484],[\"SingleFetchClassInstance\",476],[\"SingleFetchClassInstance\",465],[\"SingleFetchClassInstance\",452],[\"SingleFetchClassInstance\",439],[\"SingleFetchClassInstance\",434],[\"SingleFetchClassInstance\",430],[\"SingleFetchClassInstance\",423],[\"SingleFetchClassInstance\",415],[\"SingleFetchClassInstance\",409],[\"SingleFetchClassInstance\",382],[\"SingleFetchClassInstance\",378],[\"SingleFetchClassInstance\",370],[\"SingleFetchClassInstance\",365],[\"SingleFetchClassInstance\",361],[\"SingleFetchClassInstance\",343],[\"SingleFetchClassInstance\",338],[\"SingleFetchClassInstance\",320],[\"SingleFetchClassInstance\",315],[\"SingleFetchClassInstance\",309],[\"SingleFetchClassInstance\",301],[\"SingleFetchClassInstance\",295],[\"SingleFetchClassInstance\",287],[\"SingleFetchClassInstance\",281],[\"SingleFetchClassInstance\",251],[\"SingleFetchClassInstance\",247],[\"SingleFetchClassInstance\",242],[\"SingleFetchClassInstance\",231],[\"SingleFetchClassInstance\",227],[\"SingleFetchClassInstance\",220],[\"SingleFetchClassInstance\",212],{\"_213\":214,\"_41\":215,\"_216\":217,\"_59\":218},\"$$mdtype\",\"Tag\",\"p\",\"attributes\",{},[219],\"Apply all these to your next round of benchmarks, and you're less likely to veer off-course.\",{\"_213\":214,\"_41\":215,\"_216\":221,\"_59\":222},{},[223,224,225,224,226],\"When presenting results, don't stop at the numbers.\",\" \",\"If two runs differ meaningfully, offer a hypothesis for why: hardware, configuration, workload shape, cache behavior, network latency, or something else.\",\"The reader should not have to invent the causal story themselves.\",{\"_213\":214,\"_41\":215,\"_216\":228,\"_59\":229},{},[230],\"You now have an elementary understanding of database benchmarking.\",{\"_213\":214,\"_41\":232,\"_216\":233,\"_59\":234},\"h2\",{\"_61\":62},[235],[\"SingleFetchClassInstance\",236],{\"_213\":214,\"_41\":237,\"_216\":238,\"_59\":239},\"a\",{\"_240\":241},[65],\"href\",\"#go-forth-and-benchmark\",{\"_213\":214,\"_41\":215,\"_216\":243,\"_59\":244},{},[245,224,246],\"If the client machine has maxed out on CPU or network connections, the graph may look like the database has plateaued.\",\"But all you've really measured is the limit of the load generator.\",{\"_213\":214,\"_41\":215,\"_216\":248,\"_59\":249},{},[250],\"That last one is easy to miss!\",{\"_213\":214,\"_41\":252,\"_216\":253,\"_59\":254},\"ul\",{},[255,256,257,258,259],[\"SingleFetchClassInstance\",277],[\"SingleFetchClassInstance\",273],[\"SingleFetchClassInstance\",269],[\"SingleFetchClassInstance\",265],[\"SingleFetchClassInstance\",260],{\"_213\":214,\"_41\":261,\"_216\":262,\"_59\":263},\"li\",{},[264],\"Forgetting to check whether the benchmark client is the bottleneck\",{\"_213\":214,\"_41\":261,\"_216\":266,\"_59\":267},{},[268],\"Reporting a percentage difference without the surrounding variance\",{\"_213\":214,\"_41\":261,\"_216\":270,\"_59\":271},{},[272],\"Measuring before the system reaches steady state\",{\"_213\":214,\"_41\":261,\"_216\":274,\"_59\":275},{},[276],\"Leaving out hardware, instance type, etc.\",{\"_213\":214,\"_41\":261,\"_216\":278,\"_59\":279},{},[280],\"Reporting only averages, without percentiles, variance, or the full time-series\",{\"_213\":214,\"_41\":215,\"_216\":282,\"_59\":283},{},[284,224,285,224,286],\"As you can see, there's a lot to good benchmarking.\",\"Missing any one of these steps leads to bias.\",\"Some of the most common mistakes:\",{\"_213\":214,\"_41\":232,\"_216\":288,\"_59\":289},{\"_61\":67},[290],[\"SingleFetchClassInstance\",291],{\"_213\":214,\"_41\":237,\"_216\":292,\"_59\":293},{\"_240\":294},[68],\"#benchmark-crimes\",{\"_213\":214,\"_41\":215,\"_216\":296,\"_59\":297},{},[298,224,299,224,300],\"Good benchmarks should be reproducible.\",\"Document the client and target setups as exhaustively as possible: hardware (or cloud instance type), OS, software versions, build flags, configurations, benchmark tool, exact command line, etc.\",\"After looking at the results of a benchmark, an engineer should be able to reproduce the results.\",{\"_213\":214,\"_41\":232,\"_216\":302,\"_59\":303},{\"_61\":70},[304],[\"SingleFetchClassInstance\",305],{\"_213\":214,\"_41\":237,\"_216\":306,\"_59\":307},{\"_240\":308},[71],\"#document-everything\",{\"_213\":214,\"_41\":215,\"_216\":310,\"_59\":311},{},[312,224,313,224,314],\"Don't do this.\",\"Ensure comparisons are between comparable technologies and workloads that fit the DBMS's stated purpose.\",\"The one exception may be as an internal test to determine which technology, amongst ones with vastly different goals, is best-suited for a system.\",{\"_213\":214,\"_41\":215,\"_216\":316,\"_59\":317},{},[318,224,319],\"These are comparing databases that were distinctly optimized for different purposes.\",\"It's easy to make one look better than the other, especially when cherry-picking the workload.\",{\"_213\":214,\"_41\":252,\"_216\":321,\"_59\":322},{},[323,324,325],[\"SingleFetchClassInstance\",334],[\"SingleFetchClassInstance\",330],[\"SingleFetchClassInstance\",326],{\"_213\":214,\"_41\":261,\"_216\":327,\"_59\":328},{},[329],\"SQLite latency is 80% lower than MySQL\",{\"_213\":214,\"_41\":261,\"_216\":331,\"_59\":332},{},[333],\"Achieve 100x higher QPS on a purpose-built realtime database compared to a Postgres relational database\",{\"_213\":214,\"_41\":261,\"_216\":335,\"_59\":336},{},[337],\"Analytics queries run 100x faster on Apache Pinot than Postgres\",{\"_213\":214,\"_41\":215,\"_216\":339,\"_59\":340},{},[341,224,342],\"People sometimes draw comparisons between vastly different database engines, resulting in wild claims.\",\"Things like:\",{\"_213\":214,\"_41\":252,\"_216\":344,\"_59\":345},{},[346,347,348],[\"SingleFetchClassInstance\",357],[\"SingleFetchClassInstance\",353],[\"SingleFetchClassInstance\",349],{\"_213\":214,\"_41\":261,\"_216\":350,\"_59\":351},{},[352],\"Comparing MySQL and Postgres on an identical workload (different databases, but same stated purpose)\",{\"_213\":214,\"_41\":261,\"_216\":354,\"_59\":355},{},[356],\"Comparing 3 different cloud MySQL platforms to determine which is most performant\",{\"_213\":214,\"_41\":261,\"_216\":358,\"_59\":359},{},[360],\"Comparing 4 different Postgres configurations to determine workload suitability\",{\"_213\":214,\"_41\":215,\"_216\":362,\"_59\":363},{},[364],\"Examples of this are:\",{\"_213\":214,\"_41\":215,\"_216\":366,\"_59\":367},{},[368,224,369],\"The best benchmarks are the ones that compare apples-to-apples.\",\"In other words, ones that create data-driven comparisons between products that have the same or very similar characteristics and feature sets.\",{\"_213\":214,\"_41\":232,\"_216\":371,\"_59\":372},{\"_61\":73},[373],[\"SingleFetchClassInstance\",374],{\"_213\":214,\"_41\":237,\"_216\":375,\"_59\":376},{\"_240\":377},[74],\"#apples-to-apples-to-oranges\",{\"_213\":214,\"_41\":215,\"_216\":379,\"_59\":380},{},[381],\"It's advisable to do multiple runs to measure consistency.\",{\"_213\":214,\"_41\":215,\"_216\":383,\"_59\":384},{},[385],[\"SingleFetchClassInstance\",386],{\"_213\":214,\"_41\":387,\"_216\":388,\"_59\":389},\"ContentImage\",{\"_390\":391,\"_392\":393,\"_394\":395,\"_396\":397,\"_398\":399,\"_400\":401},[],\"alt\",\"Repeating the same benchmark\",\"height\",1416,\"loading\",\"lazy\",\"src\",\"https://planetscale-images.imgix.net/assets/repeating-DBCnml8x.png?auto=compress%2Cformat\",\"srcs\",[402,403],\"width\",3172,{\"_404\":397,\"_406\":408},{\"_404\":405,\"_406\":407},\"srcSet\",\"https://planetscale-images.imgix.net/assets/repeating-darkmode-BqI9nytL.png?auto=compress%2Cformat\",\"media\",\"(prefers-color-scheme: dark)\",\"(prefers-color-scheme: light), (prefers-color-scheme: no-preference)\",{\"_213\":214,\"_41\":215,\"_216\":410,\"_59\":411},{},[412,224,413,224,414],\"Another important consideration, especially in the cloud, is (in)consistency.\",\"Even with the same benchmark instance and same client machine, latency and throughput can vary from run to run.\",\"This can be due to contention on the network or noisy neighbors that are co-occupying the same hardware you are running on.\",{\"_213\":214,\"_41\":232,\"_216\":416,\"_59\":417},{\"_61\":76},[418],[\"SingleFetchClassInstance\",419],{\"_213\":214,\"_41\":237,\"_216\":420,\"_59\":421},{\"_240\":422},[77],\"#inconsistency\",{\"_213\":214,\"_41\":215,\"_216\":424,\"_59\":425},{},[426,224,427,224,428,224,429],\"It's important to consider the impact database configurations have on performance.\",\"An identical benchmark on the same hardware can perform very differently with different tunings.\",\"DBMSs give us these tunings so we can trade off things like performance, durability, data size, and resource consumption on a case-by-case basis.\",\"It's generally best to either (a) ensure all configuration options are aligned or (b) for pre-tuned situations (like most database-as-a-service providers) leave things at the pre-tuned defaults.\",{\"_213\":214,\"_41\":215,\"_216\":431,\"_59\":432},{},[433],\"Background jobs, I/O checkpointing, autovacuum, and other work can impact the throughput, skewing the benchmark results.\",{\"_213\":214,\"_41\":215,\"_216\":435,\"_59\":436},{},[437,224,438],\"We can end up with graphs like this in these cases.\",\"But run for another 10 minutes, and we'd likely see a large performance dip on the green line.\",{\"_213\":214,\"_41\":215,\"_216\":440,\"_59\":441},{},[442],[\"SingleFetchClassInstance\",443],{\"_213\":214,\"_41\":387,\"_216\":444,\"_59\":445},{\"_390\":446,\"_392\":393,\"_394\":395,\"_396\":447,\"_398\":448,\"_400\":401},[],\"Checkpointing in database benchmarks\",\"https://planetscale-images.imgix.net/assets/checkpoints-DPQC9Z0v.png?auto=compress%2Cformat\",[449,450],{\"_404\":447,\"_406\":408},{\"_404\":451,\"_406\":407},\"https://planetscale-images.imgix.net/assets/checkpoints-darkmode-BWJJBKrU.png?auto=compress%2Cformat\",{\"_213\":214,\"_41\":215,\"_216\":453,\"_59\":454},{},[455,456,457,224,458,224,459],\"This and \",[\"SingleFetchClassInstance\",460],\" determine how frequently we need to flush table / index changes to disk (I/O checkpointing).\",\"If we set these to low / aggressive values, we may trigger it once every minute, causing regular performance dips.\",\"If we set it lax to only trigger once every ten minutes, we may not even notice it in the results of a 5-minute benchmark execution.\",{\"_213\":214,\"_41\":461,\"_216\":462,\"_59\":463},\"code\",{},[464],\"max_wal_size\",{\"_213\":214,\"_41\":215,\"_216\":466,\"_59\":467},{},[468,224,469,470,471],\"Even when warm, there are a number of configuration options that impact performance over long stretches of time.\",\"Though there are many, a good example of this is \",[\"SingleFetchClassInstance\",472],\" in Postgres.\",{\"_213\":214,\"_41\":461,\"_216\":473,\"_59\":474},{},[475],\"checkpoint_timeout\",{\"_213\":214,\"_41\":232,\"_216\":477,\"_59\":478},{\"_61\":79},[479],[\"SingleFetchClassInstance\",480],{\"_213\":214,\"_41\":237,\"_216\":481,\"_59\":482},{\"_240\":483},[80],\"#configuration\",{\"_213\":214,\"_41\":215,\"_216\":485,\"_59\":486},{},[487,488,489,224,490],\"We typically run databases without measurement for a few minutes to ensure all caches are \",[\"SingleFetchClassInstance\",491],\" before starting benchmark measurement.\",\"This ensures non-full caches and other startup costs don't impact the numbers.\",{\"_213\":214,\"_41\":492,\"_216\":493,\"_59\":494},\"em\",{},[495],\"warmed\",{\"_213\":214,\"_41\":215,\"_216\":497,\"_59\":498},{},[499],[\"SingleFetchClassInstance\",500],{\"_213\":214,\"_41\":387,\"_216\":501,\"_59\":502},{\"_390\":503,\"_392\":393,\"_394\":395,\"_396\":504,\"_398\":505,\"_400\":401},[],\"Cache warming in databases\",\"https://planetscale-images.imgix.net/assets/warming-CmzsXHI3.png?auto=compress%2Cformat\",[506,507],{\"_404\":504,\"_406\":408},{\"_404\":508,\"_406\":407},\"https://planetscale-images.imgix.net/assets/warming-darkmode-BPg8wwcI.png?auto=compress%2Cformat\",{\"_213\":214,\"_41\":215,\"_216\":510,\"_59\":511},{},[512,224,513,514,515,516,517,224,518],\"A big one is cache warmup.\",\"If we've recently booted up our database, the various caches are not full of pages (\",[\"SingleFetchClassInstance\",523],\" in Postgres, \",[\"SingleFetchClassInstance\",519],\" in MySQL).\",\"These require time and query load to warm, during which time latency and throughput will slowly be brought up to full potential.\",{\"_213\":214,\"_41\":461,\"_216\":520,\"_59\":521},{},[522],\"buffer_pool\",{\"_213\":214,\"_41\":461,\"_216\":524,\"_59\":525},{},[526],\"buffer_cache\",{\"_213\":214,\"_41\":215,\"_216\":528,\"_59\":529},{},[530,531,532,224,533,224,534,224,535],\"We've now settled the prep work and know \",[\"SingleFetchClassInstance\",536],\" we should be measuring.\",\"Now let's get tactical.\",\"How do we ensure that we are fair when running the benchmark?\",\"There's a lot to consider for the executions themselves.\",{\"_213\":214,\"_41\":492,\"_216\":537,\"_59\":538},{},[539],\"what\",{\"_213\":214,\"_41\":232,\"_216\":541,\"_59\":542},{\"_61\":82},[543],[\"SingleFetchClassInstance\",544],{\"_213\":214,\"_41\":237,\"_216\":545,\"_59\":546},{\"_240\":547},[83],\"#warmup\",{\"_213\":214,\"_41\":215,\"_216\":549,\"_59\":550},{},[551,552,553,554,555],\"We now have a way of communicating both \",[\"SingleFetchClassInstance\",560],\" we accomplished and \",[\"SingleFetchClassInstance\",556],\" each unit of work was completed.\",{\"_213\":214,\"_41\":492,\"_216\":557,\"_59\":558},{},[559],\"how quickly\",{\"_213\":214,\"_41\":492,\"_216\":561,\"_59\":562},{},[563],\"how much work\",{\"_213\":214,\"_41\":215,\"_216\":565,\"_59\":566},{},[567],[\"SingleFetchClassInstance\",568],{\"_213\":214,\"_41\":387,\"_216\":569,\"_59\":570},{\"_390\":571,\"_392\":572,\"_394\":395,\"_396\":573,\"_398\":574,\"_400\":575},[],\"Latency bar chart\",1380,\"https://planetscale-images.imgix.net/assets/latency-bars-D6X_kFOI.png?auto=compress%2Cformat\",[576,577],2000,{\"_404\":573,\"_406\":408},{\"_404\":578,\"_406\":407},\"https://planetscale-images.imgix.net/assets/latency-bars-darkmode-DtAqZb6n.png?auto=compress%2Cformat\",{\"_213\":214,\"_41\":215,\"_216\":580,\"_59\":581},{},[582,224,583],\"In some cases, the line graphs are overkill.\",\"As with throughput, the visual can be compressed using a bar chart showing the median (or mean), with error bars.\",{\"_213\":214,\"_41\":215,\"_216\":585,\"_59\":586},{},[587],[\"SingleFetchClassInstance\",588],{\"_213\":214,\"_41\":387,\"_216\":589,\"_59\":590},{\"_390\":591,\"_392\":393,\"_394\":395,\"_396\":592,\"_398\":593,\"_400\":594},[],\"Latency line chart\",\"https://planetscale-images.imgix.net/assets/latency-lines-BNBoL30i.png?auto=compress%2Cformat\",[595,596],3208,{\"_404\":592,\"_406\":408},{\"_404\":597,\"_406\":407},\"https://planetscale-images.imgix.net/assets/latency-lines-darkmode-BI3lGrIV.png?auto=compress%2Cformat\",{\"_213\":214,\"_41\":215,\"_216\":599,\"_59\":600},{},[601,224,602,224,603,224,604],\"We can measure any latency percentile we want, but these are the most common, along with p95 and p99.9.\",\"When benchmarking, we typically measure one or more of these in a series of small windows over the entire benchmark period.\",\"Say, sample p50, p90, and p99 once per second over a 5-minute (300-second) execution.\",\"Then, we plot the results.\",{\"_213\":214,\"_41\":252,\"_216\":606,\"_59\":607},{},[608,609,610],[\"SingleFetchClassInstance\",621],[\"SingleFetchClassInstance\",616],[\"SingleFetchClassInstance\",611],{\"_213\":214,\"_41\":261,\"_216\":612,\"_59\":613},{},[614,224,615],\"p99 - The 99th percentile.\",\"During this time period, 99 out of 100 requests executed faster, 1 out of 100 slower.\",{\"_213\":214,\"_41\":261,\"_216\":617,\"_59\":618},{},[619,224,620],\"p90 - The 90th percentile.\",\"During this time period, 9 out of 10 requests executed faster, 1 out of 10 slower.\",{\"_213\":214,\"_41\":261,\"_216\":622,\"_59\":623},{},[624,224,625],\"p50 - The median latency.\",\"During this time period, half of the requests executed faster than this, the other half slower.\",{\"_213\":214,\"_41\":215,\"_216\":627,\"_59\":628},{},[629,630,631,224,632],\"The standard language for communicating about latencies in distributed systems is with \",[\"SingleFetchClassInstance\",633],\" over some span of time (1 second, 1 minute, etc.).\",\"For example:\",{\"_213\":214,\"_41\":492,\"_216\":634,\"_59\":635},{},[636],\"percentiles\",{\"_213\":214,\"_41\":215,\"_216\":638,\"_59\":639},{},[640,641,224,642,643,644],[\"SingleFetchClassInstance\",649],\" is the amount of time it takes to complete an operation, query, or transaction.\",\"We can look at individual latencies (\\\"How long did this particular \",[\"SingleFetchClassInstance\",645],\" take?\\\"), but more often we assess latencies in aggregate.\",{\"_213\":214,\"_41\":461,\"_216\":646,\"_59\":647},{},[648],\"SELECT * FROM...\",{\"_213\":214,\"_41\":492,\"_216\":650,\"_59\":651},{},[86],{\"_213\":214,\"_41\":232,\"_216\":653,\"_59\":654},{\"_61\":85},[655],[\"SingleFetchClassInstance\",656],{\"_213\":214,\"_41\":237,\"_216\":657,\"_59\":658},{\"_240\":659},[86],\"#latency\",{\"_213\":214,\"_41\":215,\"_216\":661,\"_59\":662},{},[663],\"Throughput only tells half the story.\",{\"_213\":214,\"_41\":215,\"_216\":665,\"_59\":666},{},[667,224,668,669,670,671,672,673,224,674],\"Error bars are only one way to summarize variance.\",[\"SingleFetchClassInstance\",685],\", \",[\"SingleFetchClassInstance\",680],\", and \",[\"SingleFetchClassInstance\",675],\" are different lenses on the same samples, each helping show whether a benchmark was stable, noisy, or hiding outliers.\",\"It's helpful to include these or provide the data so readers can compute them themselves.\",{\"_213\":214,\"_41\":237,\"_216\":676,\"_59\":677},{\"_240\":679},[678],\"histograms\",\"https://en.wikipedia.org/wiki/Histogram\",{\"_213\":214,\"_41\":237,\"_216\":681,\"_59\":682},{\"_240\":684},[683],\"interquartile range\",\"https://en.wikipedia.org/wiki/Interquartile_range\",{\"_213\":214,\"_41\":237,\"_216\":686,\"_59\":687},{\"_240\":689},[688],\"Coefficient of variation\",\"https://en.wikipedia.org/wiki/Coefficient_of_variation\",{\"_213\":214,\"_41\":215,\"_216\":691,\"_59\":692},{},[693,224,694],\"This communicates similar information in a more compact way, but it's ideal to show a full line graph, as that also better visualizes inconsistencies or spikiness of performance throughout a benchmark run.\",\"More on this later.\",{\"_213\":214,\"_41\":215,\"_216\":696,\"_59\":697},{},[698],[\"SingleFetchClassInstance\",699],{\"_213\":214,\"_41\":387,\"_216\":700,\"_59\":701},{\"_390\":702,\"_392\":703,\"_394\":395,\"_396\":704,\"_398\":705,\"_400\":575},[],\"Throughput bar chart\",1340,\"https://planetscale-images.imgix.net/assets/throughput-bars-DAVyhM1x.png?auto=compress%2Cformat\",[706,707],{\"_404\":704,\"_406\":408},{\"_404\":708,\"_406\":407},\"https://planetscale-images.imgix.net/assets/throughput-bars-darkmode-DRPfZNGr.png?auto=compress%2Cformat\",{\"_213\":214,\"_41\":215,\"_216\":710,\"_59\":711},{},[712],\"A more compact way of displaying this is via a bar chart with error bars.\",{\"_213\":214,\"_41\":215,\"_216\":714,\"_59\":715},{},[716],[\"SingleFetchClassInstance\",717],{\"_213\":214,\"_41\":387,\"_216\":718,\"_59\":719},{\"_390\":720,\"_392\":393,\"_394\":395,\"_396\":721,\"_398\":722,\"_400\":401},[],\"Throughput line chart\",\"https://planetscale-images.imgix.net/assets/throughput-lines-DG3lJ36Y.png?auto=compress%2Cformat\",[723,724],{\"_404\":721,\"_406\":408},{\"_404\":725,\"_406\":407},\"https://planetscale-images.imgix.net/assets/throughput-lines-darkmode-B-LbVppT.png?auto=compress%2Cformat\",{\"_213\":214,\"_41\":215,\"_216\":727,\"_59\":728},{},[729,224,730,224,731],\"To measure throughput, choose a workload, a period of time to run it for (say, 5 minutes / 300 seconds), and then execute with TPS / QPS sampling.\",\"As a benchmark runs, samples are taken of how many queries or transactions complete each second.\",\"We then display this as a graph, showing every collected data point:\",{\"_213\":214,\"_41\":215,\"_216\":733,\"_59\":734},{},[735,736,224,737,224,738,739,740,741,669,742,743,224,744],[\"SingleFetchClassInstance\",759],\" is the amount of work completed in a slice of time.\",\"In databases, the most common measures are Queries Per Second (QPS) or Transactions Per Second (TPS).\",\"For many popular benchmarks like \",[\"SingleFetchClassInstance\",754],\" and \",[\"SingleFetchClassInstance\",749],[\"SingleFetchClassInstance\",745],\" because there are typically multiple queries within single transactions.\",\"Either works fine as a measure.\",{\"_213\":214,\"_41\":461,\"_216\":746,\"_59\":747},{},[748],\"TPS \u003c QPS\",{\"_213\":214,\"_41\":237,\"_216\":750,\"_59\":751},{\"_240\":753},[752],\"TPC-H\",\"https://www.tpc.org/tpch/\",{\"_213\":214,\"_41\":237,\"_216\":755,\"_59\":756},{\"_240\":758},[757],\"TPC-C\",\"https://www.tpc.org/tpcc/\",{\"_213\":214,\"_41\":492,\"_216\":760,\"_59\":761},{},[89],{\"_213\":214,\"_41\":232,\"_216\":763,\"_59\":764},{\"_61\":88},[765],[\"SingleFetchClassInstance\",766],{\"_213\":214,\"_41\":237,\"_216\":767,\"_59\":768},{\"_240\":769},[89],\"#throughput\",{\"_213\":214,\"_41\":215,\"_216\":771,\"_59\":772},{},[773,774,740,775,776,224,777],\"Broadly, there are two things we like to measure when benchmarking: \",[\"SingleFetchClassInstance\",781],[\"SingleFetchClassInstance\",778],\".\",\"Any good database benchmark will report on both of these things.\",{\"_213\":214,\"_41\":492,\"_216\":779,\"_59\":780},{},[85],{\"_213\":214,\"_41\":492,\"_216\":782,\"_59\":783},{},[88],{\"_213\":214,\"_41\":232,\"_216\":785,\"_59\":786},{\"_61\":91},[787],[\"SingleFetchClassInstance\",788],{\"_213\":214,\"_41\":237,\"_216\":789,\"_59\":790},{\"_240\":791},[92],\"#what-to-measure\",{\"_213\":214,\"_41\":215,\"_216\":793,\"_59\":794},{},[795,224,796],\"Both are beneficial, but they are useful for different things.\",\"Important to decide up front what the purpose of a benchmark is, then choose the type accordingly.\",{\"_213\":214,\"_41\":215,\"_216\":798,\"_59\":799},{},[800],\"Closed-loop benchmarks are more commonly seen in academic and performance comparisons, as they offer a more controlled environment for comparing things like QPS across a fixed amount of concurrency.\",{\"_213\":214,\"_41\":215,\"_216\":802,\"_59\":803},{},[804,224,805],\"Open-loop benchmarks tend to be more realistic.\",\"In production systems, database load is applied at the rate that the clients demand, regardless of how well the database is keeping up.\",{\"_213\":214,\"_41\":215,\"_216\":807,\"_59\":808},{},[809],[\"SingleFetchClassInstance\",810],{\"_213\":214,\"_41\":387,\"_216\":811,\"_59\":812},{\"_390\":813,\"_392\":393,\"_394\":395,\"_396\":814,\"_398\":815,\"_400\":401},[],\"Open vs Closed loop benchmark\",\"https://planetscale-images.imgix.net/assets/open-closed-loop-B0_3fnSO.png?auto=compress%2Cformat\",[816,817],{\"_404\":814,\"_406\":408},{\"_404\":818,\"_406\":407},\"https://planetscale-images.imgix.net/assets/open-closed-loop-darkmode-CYoXgW5C.png?auto=compress%2Cformat\",{\"_213\":214,\"_41\":215,\"_216\":820,\"_59\":821},{},[822],\"This can be fixed throughout the entire benchmark duration, or vary in a controlled way:\",{\"_213\":214,\"_41\":824,\"_216\":825,\"_59\":826},\"CodeBlock\",{\"_827\":828,\"_829\":830,\"_831\":832,\"_461\":-7},[],\"html\",\"\u003cpre class=\\\"shiki shiki-themes planetscale-light planetscale-dark\\\" style=\\\"--shiki-light:#2b2b2b;--shiki-dark:#e1e1e1;--shiki-light-bg:#ebebeb;--shiki-dark-bg:#1a1a1a\\\" tabindex=\\\"0\\\"\u003e\u003ccode\u003e\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003ewhile\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#7D5903;--shiki-dark:#FED54A\\\"\u003e True\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e:\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#818181;--shiki-dark:#A1A1A1\\\"\u003e # fire and forget\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#0B6EC5;--shiki-dark:#73C7F9\\\"\u003e send_bench_request\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e()\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#818181;--shiki-dark:#A1A1A1\\\"\u003e # fixed pace\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#2B2B2B;--shiki-dark:#E1E1E1\\\"\u003e time\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#0B6EC5;--shiki-dark:#73C7F9\\\"\u003esleep\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e(\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#D92038;--shiki-dark:#FF7082\\\"\u003e0.1\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e)\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003c/span\u003e\u003c/code\u003e\u003c/pre\u003e\",\"language\",\"python\",\"clipboard\",false,{\"_213\":214,\"_41\":215,\"_216\":834,\"_59\":835},{},[836],\"Open loop on the other hand has a fixed pace of sending requests, regardless of how quickly the database responds.\",{\"_213\":214,\"_41\":215,\"_216\":838,\"_59\":839},{},[840,224,841,224,842],\"We may do this in parallel across many connections, but each individual connection sends a controlled sequence of queries.\",\"A closed loop can also hide a failure mode called coordinated omission: when the database stalls, the client stops issuing new requests too, so the benchmark only records the stalled request and omits the work that would have queued behind it.\",\"This is especially misleading for tail latency, where the missing queued requests are exactly the ones that would have made p95/p99 look worse (more on latency and percentiles soon).\",{\"_213\":214,\"_41\":824,\"_216\":844,\"_59\":845},{\"_827\":846,\"_829\":830,\"_831\":832,\"_461\":-7},[],\"\u003cpre class=\\\"shiki shiki-themes planetscale-light planetscale-dark\\\" style=\\\"--shiki-light:#2b2b2b;--shiki-dark:#e1e1e1;--shiki-light-bg:#ebebeb;--shiki-dark-bg:#1a1a1a\\\" tabindex=\\\"0\\\"\u003e\u003ccode\u003e\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003ewhile\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#7D5903;--shiki-dark:#FED54A\\\"\u003e True\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e:\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#818181;--shiki-dark:#A1A1A1\\\"\u003e # wait for response\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#2B2B2B;--shiki-dark:#E1E1E1\\\"\u003e response \u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e=\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#0B6EC5;--shiki-dark:#73C7F9\\\"\u003e send_bench_request\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e()\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#818181;--shiki-dark:#A1A1A1\\\"\u003e # then send next\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#0B6EC5;--shiki-dark:#73C7F9\\\"\u003e process\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e(\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#0B6EC5;--shiki-dark:#73C7F9\\\"\u003eresponse\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e)\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003c/span\u003e\u003c/code\u003e\u003c/pre\u003e\",{\"_213\":214,\"_41\":215,\"_216\":848,\"_59\":849},{},[850],\"In a closed-loop benchmark, the client sends requests and then waits for a response before sending the next.\",{\"_213\":214,\"_41\":215,\"_216\":852,\"_59\":853},{},[854],\"There are two types of benchmark workload shapes: open and closed loops.\",{\"_213\":214,\"_41\":232,\"_216\":856,\"_59\":857},{\"_61\":94},[858],[\"SingleFetchClassInstance\",859],{\"_213\":214,\"_41\":237,\"_216\":860,\"_59\":861},{\"_240\":862},[95],\"#closed-and-open-loop\",{\"_213\":214,\"_41\":215,\"_216\":864,\"_59\":865},{},[866],\"Which pattern the benchmark operates with significantly impacts performance, because it in turn impacts how frequently we need to access disk vs RAM and the amount of cache churn.\",{\"_213\":214,\"_41\":215,\"_216\":868,\"_59\":869},{},[870,224,871],\"These are only simple models.\",\"Real workloads often have messier shapes: recently inserted rows might be hotter than old rows, one tenant might dominate traffic, or a small working set might receive most reads for a period of time.\",{\"_213\":214,\"_41\":215,\"_216\":873,\"_59\":874},{},[875,876,877,878,776,224,879],\"A \",[\"SingleFetchClassInstance\",884],\" access pattern is skewed: the k-th most popular key is accessed roughly proportional to \",[\"SingleFetchClassInstance\",880],\"A small number of hot rows receive a large share of requests, while most rows are accessed rarely.\",{\"_213\":214,\"_41\":461,\"_216\":881,\"_59\":882},{},[883],\"1/k\",{\"_213\":214,\"_41\":492,\"_216\":885,\"_59\":886},{},[887],\"Zipfian\",{\"_213\":214,\"_41\":215,\"_216\":889,\"_59\":890},{},[875,891,892,224,893],[\"SingleFetchClassInstance\",894],\" access pattern gives every row the same chance of being queried on each request.\",\"If we have 100 rows, each has a 1% chance of being read for each operation.\",{\"_213\":214,\"_41\":492,\"_216\":895,\"_59\":896},{},[897],\"uniform\",{\"_213\":214,\"_41\":215,\"_216\":899,\"_59\":900},{},[901],[\"SingleFetchClassInstance\",902],{\"_213\":214,\"_41\":387,\"_216\":903,\"_59\":904},{\"_390\":905,\"_392\":906,\"_394\":395,\"_396\":907,\"_398\":908,\"_400\":909},[],\"Types of data distributions\",1256,\"https://planetscale-images.imgix.net/assets/distributions-DHT64hC8.png?auto=compress%2Cformat\",[910,911],2252,{\"_404\":907,\"_406\":408},{\"_404\":912,\"_406\":407},\"https://planetscale-images.imgix.net/assets/distributions-darkmode-Bgx07A5I.png?auto=compress%2Cformat\",{\"_213\":214,\"_41\":215,\"_216\":914,\"_59\":915},{},[916,224,917],\"Even for a fixed data size, access patterns can vary widely.\",\"The simplest examples are uniform and Zipfian.\",{\"_213\":214,\"_41\":215,\"_216\":919,\"_59\":920},{},[921],\"This is related to, but not the same as, data distribution.\",{\"_213\":214,\"_41\":215,\"_216\":923,\"_59\":924},{},[925,926,776,224,927,224,928,224,929,224,930],\"Consider a benchmark like \",[\"SingleFetchClassInstance\",931],\"This is a simple, read-only benchmark that runs a handful of select query patterns repeatedly.\",\"As benchmarks often do, the data size is configurable in the preparation phase.\",\"If we run this benchmark on a server with 64 GB of RAM and a 32 GB data size, the entire data set will fit in RAM after warming.\",\"The same benchmark run with a 320 GB data size will generate significant I/O and inevitably run slower.\",{\"_213\":214,\"_41\":237,\"_216\":932,\"_59\":933},{\"_240\":935},[934],\"sysbench OLTP read-only\",\"https://github.com/akopytov/sysbench/blob/master/src/lua/oltp_read_only.lua\",{\"_213\":214,\"_41\":215,\"_216\":937,\"_59\":938},{},[939,224,940,224,941],\"Some databases operate on local SSDs, while others use network-attached storage like AWS EBS or Google Persistent Disk.\",\"Some even take a hybrid approach.\",\"Either way, the percent of read traffic hitting RAM vs disk impacts performance due to I/O wait times.\",{\"_213\":214,\"_41\":215,\"_216\":943,\"_59\":944},{},[945],[\"SingleFetchClassInstance\",946],{\"_213\":214,\"_41\":387,\"_216\":947,\"_59\":948},{\"_390\":949,\"_392\":950,\"_394\":395,\"_396\":951,\"_398\":952,\"_400\":953},[],\"RAM, disk latency\",1352,\"https://planetscale-images.imgix.net/assets/ram-disk-B91Wyn3l.png?auto=compress%2Cformat\",[954,955],2400,{\"_404\":951,\"_406\":408},{\"_404\":956,\"_406\":407},\"https://planetscale-images.imgix.net/assets/ram-disk-darkmode-BsqMcUR2.png?auto=compress%2Cformat\",{\"_213\":214,\"_41\":215,\"_216\":958,\"_59\":959},{},[960,224,961,224,962,963],\"Queries executed on a relational database often require some amount of I/O work.\",\"Writing data must always be persisted to disk.\",[\"SingleFetchClassInstance\",964],\" data can come from the in-memory cache, or disk on cache misses.\",{\"_213\":214,\"_41\":492,\"_216\":965,\"_59\":966},{},[967],\"Reading\",{\"_213\":214,\"_41\":215,\"_216\":969,\"_59\":970},{},[971],\"All of these impact performance, especially when combined with the variations of underlying hardware.\",{\"_213\":214,\"_41\":252,\"_216\":973,\"_59\":974},{},[975,976,977],[\"SingleFetchClassInstance\",986],[\"SingleFetchClassInstance\",982],[\"SingleFetchClassInstance\",978],{\"_213\":214,\"_41\":261,\"_216\":979,\"_59\":980},{},[981],\"What's the ratio of reads to writes?\",{\"_213\":214,\"_41\":261,\"_216\":983,\"_59\":984},{},[985],\"What % of the data is hot (frequently queried) vs cold (rarely queried)?\",{\"_213\":214,\"_41\":261,\"_216\":987,\"_59\":988},{},[989],\"How many queries are hitting RAM vs disk?\",{\"_213\":214,\"_41\":215,\"_216\":991,\"_59\":992},{},[993],\"The easiest way to think about this is in terms of traffic ratios.\",{\"_213\":214,\"_41\":215,\"_216\":995,\"_59\":996},{},[997],\"Even once we know that our infrastructure is set up sanely, there's a lot to consider for the workload we run.\",{\"_213\":214,\"_41\":232,\"_216\":999,\"_59\":1000},{\"_61\":97},[1001],[\"SingleFetchClassInstance\",1002],{\"_213\":214,\"_41\":237,\"_216\":1003,\"_59\":1004},{\"_240\":1005},[98],\"#workload\",{\"_213\":214,\"_41\":215,\"_216\":1007,\"_59\":1008},{},[1009,1010,1011,224,1012,224,1013,224,1014,1015,1016],\"In an ideal world, we'd run everything on the \",[\"SingleFetchClassInstance\",1021],\" same instance.\",\"In reality, we sometimes have to settle for matching CPUs and RAM as best we can, and living with the differences.\",\"However, you must give this your best effort.\",\"Purposefully choosing to benchmark \",[\"SingleFetchClassInstance\",1017],\" product on 2025-gen CPU and then comparing to a competitor's product on a 2022 CPU, when the alternate was readily available, is intentionally misleading.\",{\"_213\":214,\"_41\":492,\"_216\":1018,\"_59\":1019},{},[1020],\"your\",{\"_213\":214,\"_41\":492,\"_216\":1022,\"_59\":1023},{},[1024],\"exact\",{\"_213\":214,\"_41\":215,\"_216\":1026,\"_59\":1027},{},[1028,224,1029,224,1030],\"You might then be tempted to just use the same instance for everything, but this breaks down too.\",\"The availability of instance types varies over time, region, and database provider.\",\"In some cases, it's not possible to match.\",{\"_213\":214,\"_41\":215,\"_216\":1032,\"_59\":1033},{},[1034],[\"SingleFetchClassInstance\",1035],{\"_213\":214,\"_41\":387,\"_216\":1036,\"_59\":1037},{\"_390\":1038,\"_392\":1039,\"_394\":395,\"_396\":1040,\"_398\":1041,\"_400\":575},[],\"Geekbench results\",1264,\"https://planetscale-images.imgix.net/assets/geekbench-DFb0s8q0.png?auto=compress%2Cformat\",[1042,1043],{\"_404\":1040,\"_406\":408},{\"_404\":1044,\"_406\":407},\"https://planetscale-images.imgix.net/assets/geekbench-darkmode-zAs-8HQc.png?auto=compress%2Cformat\",{\"_213\":214,\"_41\":215,\"_216\":1046,\"_59\":1047},{},[1048,224,1049,1050,1051,1052,776,224,1053,224,1054],\"This isn't as easy as it sounds, especially when we're talking about running things in the hyperscaler clouds like AWS and GCP.\",\"For example, the Geekbench results for an AWS \",[\"SingleFetchClassInstance\",1060],\" are ~15% lower than the results for an \",[\"SingleFetchClassInstance\",1055],\"Both have 8 vCPUs and 64 GB RAM.\",\"But move one generation newer, and there's a ~15% CPU improvement.\",{\"_213\":214,\"_41\":237,\"_216\":1056,\"_59\":1057},{\"_240\":1059},[1058],\"r8g.2xlarge\",\"https://browser.geekbench.com/v6/cpu/11335856\",{\"_213\":214,\"_41\":237,\"_216\":1061,\"_59\":1062},{\"_240\":1064},[1063],\"r7g.2xlarge\",\"https://browser.geekbench.com/v6/cpu/2119560\",{\"_213\":214,\"_41\":215,\"_216\":1066,\"_59\":1067},{},[1068],\"An important prerequisite to proper benchmarking is setting up the compute, storage, and networking resources to allow for a fair fight.\",{\"_213\":214,\"_41\":215,\"_216\":1070,\"_59\":1071},{},[1072,224,1073],\"It's easy to make one database look better than another with an imbalance of resources.\",\"Postgres running on a 16-core server will almost always perform better than on an 8-core server.\",{\"_213\":214,\"_41\":232,\"_216\":1075,\"_59\":1076},{\"_61\":100},[1077],[\"SingleFetchClassInstance\",1078],{\"_213\":214,\"_41\":237,\"_216\":1079,\"_59\":1080},{\"_240\":1081},[101],\"#choosing-resources\",{\"_213\":214,\"_41\":215,\"_216\":1083,\"_59\":1084},{},[1085,224,1086,224,1087],\"Finally, you should double-check that the client server is not a bottleneck.\",\"While benchmarking, ensure that CPU and network utilization are well under their capacity.\",\"We want to be straining the database server, not the client.\",{\"_213\":214,\"_41\":215,\"_216\":1089,\"_59\":1090},{},[1091,1092,1093,224,1094,224,1095,224,1096],\"It \",[\"SingleFetchClassInstance\",1097],\" also impact throughput.\",\"We often don't run benchmarks on a single connection.\",\"We'll do 10, 50, or 100 simultaneous connections to best utilize the parallelism of the machine and database.\",\"But if we have a fixed connection count, and are not making it dynamic to account for round-trip latency, we can end up allowing the elevated latency to hurt throughput.\",{\"_213\":214,\"_41\":492,\"_216\":1098,\"_59\":1099},{},[1100],\"can\",{\"_213\":214,\"_41\":215,\"_216\":1102,\"_59\":1103},{},[1104],\"This latency variation will always have an impact on latency measurements.\",{\"_213\":214,\"_41\":215,\"_216\":1106,\"_59\":1107},{},[1108,224,1109,224,1110],\"Same database.\",\"Same benchmark client.\",\"The only difference is the speed at which bytes can go over the wire between the two.\",{\"_213\":214,\"_41\":215,\"_216\":1112,\"_59\":1113},{},[1114],[\"SingleFetchClassInstance\",1115],{\"_213\":214,\"_41\":387,\"_216\":1116,\"_59\":1117},{\"_390\":1118,\"_392\":1119,\"_394\":395,\"_396\":1120,\"_398\":1121,\"_400\":1122},[],\"How network latency impacts throughput\",1404,\"https://planetscale-images.imgix.net/assets/network-latency-difference-C_4xTSar.png?auto=compress%2Cformat\",[1123,1124],3288,{\"_404\":1120,\"_406\":408},{\"_404\":1125,\"_406\":407},\"https://planetscale-images.imgix.net/assets/network-latency-difference-darkmode-CTYmZVN1.png?auto=compress%2Cformat\",{\"_213\":214,\"_41\":215,\"_216\":1127,\"_59\":1128},{},[1129,224,1130,224,1131,224,1132],\"Let's consider a scenario where each query takes ~10ms to execute on the database.\",\"If the network round-trip time is 2.5 milliseconds, then we can execute approximately 80 queries per second over a single connection.\",\"On the other hand, what if the round-trip is 15 milliseconds?\",\"We've now cut our single-threaded QPS capability in ~half, resulting in 40 QPS.\",{\"_213\":214,\"_41\":215,\"_216\":1134,\"_59\":1135},{},[1136],[\"SingleFetchClassInstance\",1137],{\"_213\":214,\"_41\":387,\"_216\":1138,\"_59\":1139},{\"_390\":1140,\"_392\":1141,\"_394\":395,\"_396\":1142,\"_398\":1143,\"_400\":1144},[],\"Network latency and execution latency\",1444,\"https://planetscale-images.imgix.net/assets/network-latency-execution-latency-D7jzvLOo.png?auto=compress%2Cformat\",[1145,1146],3256,{\"_404\":1142,\"_406\":408},{\"_404\":1147,\"_406\":407},\"https://planetscale-images.imgix.net/assets/network-latency-execution-latency-darkmode-DonplkLE.png?auto=compress%2Cformat\",{\"_213\":214,\"_41\":215,\"_216\":1149,\"_59\":1150},{},[1151,224,1152,224,1153,224,1154],\"How much this skews the results of the benchmark depends quite a bit on how \\\"far apart\\\" the benchmark server and database server are\",\"(network latency)\",\"and how long the queries / transactions take on the database\",\"(execution latency).\",{\"_213\":214,\"_41\":215,\"_216\":1156,\"_59\":1157},{},[1158],[\"SingleFetchClassInstance\",1159],{\"_213\":214,\"_41\":387,\"_216\":1160,\"_59\":1161},{\"_390\":1162,\"_392\":1163,\"_394\":395,\"_396\":1164,\"_398\":1165,\"_400\":953},[],\"Client-server\",1880,\"https://planetscale-images.imgix.net/assets/client-server-DSv7LBqp.png?auto=compress%2Cformat\",[1166,1167],{\"_404\":1164,\"_406\":408},{\"_404\":1168,\"_406\":407},\"https://planetscale-images.imgix.net/assets/client-server-darkmode-l5yZMo25.png?auto=compress%2Cformat\",{\"_213\":214,\"_41\":215,\"_216\":1170,\"_59\":1171},{},[1172,1173,776,224,1174],\"As usual, \",[\"SingleFetchClassInstance\",1175],\"This introduces latency between the two machines.\",{\"_213\":214,\"_41\":492,\"_216\":1176,\"_59\":1177},{},[1178],\"there's a catch\",{\"_213\":214,\"_41\":215,\"_216\":1180,\"_59\":1181},{},[1182,224,1183,1184,1185],\"To benchmark, we need a client that establishes the connections, generates queries, and takes measurements.\",\"Since both sides consume resources and we want to give the \",[\"SingleFetchClassInstance\",1186],\" its full share of the host server, it's common to set up a distinct server for benchmark execution.\",{\"_213\":214,\"_41\":492,\"_216\":1187,\"_59\":1188},{},[1189],\"database\",{\"_213\":214,\"_41\":215,\"_216\":1191,\"_59\":1192},{},[1193,224,1194],\"Databases typically operate in a client-server model.\",\"The database server is started, accepts connections from clients, executes queries, and returns results.\",{\"_213\":214,\"_41\":232,\"_216\":1196,\"_59\":1197},{\"_61\":103},[1198],[\"SingleFetchClassInstance\",1199],{\"_213\":214,\"_41\":237,\"_216\":1200,\"_59\":1201},{\"_240\":1202},[104],\"#client-server-architecture\",{\"_213\":214,\"_41\":215,\"_216\":1204,\"_59\":1205},{},[1206,1207,1208],\"Here, we're focusing primarily on benchmarking \",[\"SingleFetchClassInstance\",1209],\", but these principles apply to many domains.\",{\"_213\":214,\"_41\":492,\"_216\":1210,\"_59\":1211},{},[1212],\"databases\",{\"_213\":214,\"_41\":215,\"_216\":1214,\"_59\":1215},{},[1216,224,1217,224,1218,1219,1220,224,1221,224,1222],\"Am I the world's best benchmarker?\",\"Certainly not.\",\"I invented the \",[\"SingleFetchClassInstance\",1223],\", after all.\",\"But correctness and precision are important parts of PlanetScale's culture.\",\"We've spent considerable time learning the art of benchmarking, and are here to share best-practices.\",{\"_213\":214,\"_41\":237,\"_216\":1224,\"_59\":1225},{\"_240\":1227},[1226],\"language balls\",\"https://x.com/BenjDicken/status/1861072804239847914\",{\"_213\":214,\"_41\":215,\"_216\":1229,\"_59\":1230},{},[1231,224,1232],\"But zooming out from any single system or harness, there are broad principles that should be applied to all benchmarking.\",\"Using these correctly makes it difficult to produce biased results.\",{\"_213\":214,\"_41\":215,\"_216\":1234,\"_59\":1235},{},[1236,224,1237],\"Benchmarking is hard.\",\"There are many ways to do it wrong and few to do it right.\",\"current\",{\"_1240\":832,\"_1241\":1242,\"_1243\":832},\"development\",\"env\",{\"_1244\":1245,\"_1246\":1247,\"_1248\":1249,\"_1250\":1251,\"_1252\":1253},\"userSignedIn\",\"IMAGE_CDN\",\"https://planetscale-images.imgix.net\",\"IMAGE_CDN_ENABLED\",\"true\",\"INTERNAL_API\",\"https://api.planetscale.com\",\"RELEASE\",\"117b8aaf-965c-42bc-b013-5f72770de4d9\",\"SENTRY_DSN\",\"https://bd81903b44804e22a06bdc0c1a91b303@o499952.ingest.us.sentry.io/4504531942572032\"]\n");</script><!--$--><script nonce="1NjQdEMwrEHFpdD1zVsfJP6ksuIJ/ecPax+EDqlkDLg=">window.__reactRouterContext.streamController.close();</script><!--/$--><!--/$--></body></html>