Files
nexus/sreweekly/articles/522/06-the-feedback-loops-behind-kubernetes.html
2026-09-12 17:23:01 +08:00

300 lines
314 KiB
HTML

<!DOCTYPE html><html lang="en"><head><meta charSet="utf-8"/><meta name="viewport" content="width=device-width, initial-scale=1"/><meta name="theme-color" content="#111111"/><meta name="user-signed-in" content="false"/><title>The feedback loops behind Kubernetes — PlanetScale</title><meta name="description" content="Kubernetes is a framework for feedback controllers: write down what you want, observe what exists, make the next change, and repeat."/><meta name="robots"/><meta property="og:url" content="https://planetscale.com/blog/the-feedback-loops-behind-kubernetes"/><meta property="og:type" content="website"/><meta property="og:title" content="The feedback loops behind Kubernetes — PlanetScale"/><meta property="og:image" content="https://planetscale.com/assets/the-feedback-loops-behind-kubernetes-social-DXEE2U-o.png"/><meta property="og:description" content="Kubernetes is a framework for feedback controllers: write down what you want, observe what exists, make the next change, and repeat."/><meta property="twitter:card" content="summary_large_image"/><meta property="twitter:site" content="@PlanetScale"/><meta property="twitter:creator" content="@PlanetScale"/><meta property="twitter:url" content="https://planetscale.com/blog/the-feedback-loops-behind-kubernetes"/><meta property="twitter:title" content="The feedback loops behind Kubernetes — PlanetScale"/><meta property="twitter:description" content="Kubernetes is a framework for feedback controllers: write down what you want, observe what exists, make the next change, and repeat."/><meta property="twitter:image" content="https://planetscale.com/assets/the-feedback-loops-behind-kubernetes-social-DXEE2U-o.png"/><link rel="canonical" href="https://planetscale.com/blog/the-feedback-loops-behind-kubernetes"/><link rel="preconnect" href="https://planetscale-images.imgix.net"/><link nonce="h1m8nAKwwY4XGImd6cj2zjoYAypTpKkpoCcuiMZKkXs=" rel="icon" href="/favicon.ico" type="image/x-icon" sizes="16x16"/><link nonce="h1m8nAKwwY4XGImd6cj2zjoYAypTpKkpoCcuiMZKkXs=" rel="icon" href="/icon.png" type="image/png" sizes="32x32"/><link nonce="h1m8nAKwwY4XGImd6cj2zjoYAypTpKkpoCcuiMZKkXs=" rel="apple-touch-icon" href="/apple-touch-icon.png" type="image/png" sizes="32x32"/><link nonce="h1m8nAKwwY4XGImd6cj2zjoYAypTpKkpoCcuiMZKkXs=" rel="manifest" href="/manifest.webmanifest"/><link rel="modulepreload" href="/assets/entry.client-3vubyXrk.js" nonce="h1m8nAKwwY4XGImd6cj2zjoYAypTpKkpoCcuiMZKkXs="/><link rel="modulepreload" href="/assets/jsx-runtime-DwfQwkRq.js" nonce="h1m8nAKwwY4XGImd6cj2zjoYAypTpKkpoCcuiMZKkXs="/><link rel="modulepreload" href="/assets/components-_bNmAApg.js" nonce="h1m8nAKwwY4XGImd6cj2zjoYAypTpKkpoCcuiMZKkXs="/><link rel="modulepreload" href="/assets/_.well-known_.mcp.server-card_.json_-Sx7XeH3e.js" nonce="h1m8nAKwwY4XGImd6cj2zjoYAypTpKkpoCcuiMZKkXs="/><link rel="modulepreload" href="/assets/index-mKTXLmHu.js" nonce="h1m8nAKwwY4XGImd6cj2zjoYAypTpKkpoCcuiMZKkXs="/><link rel="modulepreload" href="/assets/errorBoundaries-DhW4jVYt.js" nonce="h1m8nAKwwY4XGImd6cj2zjoYAypTpKkpoCcuiMZKkXs="/><link rel="modulepreload" href="/assets/root-DbOv4-98.js" nonce="h1m8nAKwwY4XGImd6cj2zjoYAypTpKkpoCcuiMZKkXs="/><link rel="modulepreload" href="/assets/lib-Dg89tQ22.js" nonce="h1m8nAKwwY4XGImd6cj2zjoYAypTpKkpoCcuiMZKkXs="/><link rel="modulepreload" href="/assets/analytics.client-DM6E8o1h.js" nonce="h1m8nAKwwY4XGImd6cj2zjoYAypTpKkpoCcuiMZKkXs="/><link rel="modulepreload" href="/assets/SiteHeader-C2U5gvDH.js" nonce="h1m8nAKwwY4XGImd6cj2zjoYAypTpKkpoCcuiMZKkXs="/><link rel="modulepreload" href="/assets/current-9yDxj94E.js" nonce="h1m8nAKwwY4XGImd6cj2zjoYAypTpKkpoCcuiMZKkXs="/><link rel="modulepreload" href="/assets/clsx-eT0YPcGk.js" nonce="h1m8nAKwwY4XGImd6cj2zjoYAypTpKkpoCcuiMZKkXs="/><link rel="modulepreload" href="/assets/bugs-38ilEoW0.js" nonce="h1m8nAKwwY4XGImd6cj2zjoYAypTpKkpoCcuiMZKkXs="/><link rel="modulepreload" href="/assets/keyboard-D-uXZORL.js" nonce="h1m8nAKwwY4XGImd6cj2zjoYAypTpKkpoCcuiMZKkXs="/><link rel="modulepreload" href="/assets/use-tab-direction-dKm-S3Ck.js" nonce="h1m8nAKwwY4XGImd6cj2zjoYAypTpKkpoCcuiMZKkXs="/><link rel="modulepreload" href="/assets/blog-pXH7ptHJ.js" nonce="h1m8nAKwwY4XGImd6cj2zjoYAypTpKkpoCcuiMZKkXs="/><link rel="modulepreload" href="/assets/blog._slug-Ch_92qsH.js" nonce="h1m8nAKwwY4XGImd6cj2zjoYAypTpKkpoCcuiMZKkXs="/><link rel="modulepreload" href="/assets/ContentImage-Dh6VEOUl.js" nonce="h1m8nAKwwY4XGImd6cj2zjoYAypTpKkpoCcuiMZKkXs="/><link rel="modulepreload" href="/assets/BlogCategoryLink-DmQyn0gp.js" nonce="h1m8nAKwwY4XGImd6cj2zjoYAypTpKkpoCcuiMZKkXs="/><link rel="modulepreload" href="/assets/Details-BSB_b6hI.js" nonce="h1m8nAKwwY4XGImd6cj2zjoYAypTpKkpoCcuiMZKkXs="/><link rel="modulepreload" href="/assets/Skittle-CDFOPRjH.js" nonce="h1m8nAKwwY4XGImd6cj2zjoYAypTpKkpoCcuiMZKkXs="/><link rel="modulepreload" href="/assets/SiteFooter-B2Gq9u2j.js" nonce="h1m8nAKwwY4XGImd6cj2zjoYAypTpKkpoCcuiMZKkXs="/><link rel="modulepreload" href="/assets/Vimeo-00PQJDli.js" nonce="h1m8nAKwwY4XGImd6cj2zjoYAypTpKkpoCcuiMZKkXs="/><link rel="modulepreload" href="/assets/YouTube-CMfaljVr.js" nonce="h1m8nAKwwY4XGImd6cj2zjoYAypTpKkpoCcuiMZKkXs="/><link rel="modulepreload" href="/assets/date-CJTFH3uT.js" nonce="h1m8nAKwwY4XGImd6cj2zjoYAypTpKkpoCcuiMZKkXs="/><link rel="modulepreload" href="/assets/use-inert-others-BMJ6-xOX.js" nonce="h1m8nAKwwY4XGImd6cj2zjoYAypTpKkpoCcuiMZKkXs="/><link rel="modulepreload" href="/assets/description-Cf6FZmDe.js" nonce="h1m8nAKwwY4XGImd6cj2zjoYAypTpKkpoCcuiMZKkXs="/><link rel="modulepreload" href="/assets/use-is-mounted-uQsUZyP9.js" nonce="h1m8nAKwwY4XGImd6cj2zjoYAypTpKkpoCcuiMZKkXs="/><link rel="modulepreload" href="/assets/types-DvonrUFF.js" nonce="h1m8nAKwwY4XGImd6cj2zjoYAypTpKkpoCcuiMZKkXs="/><link rel="stylesheet" href="/assets/styles-ns8XBZ1D.css"/><script nonce="h1m8nAKwwY4XGImd6cj2zjoYAypTpKkpoCcuiMZKkXs=">window.ENV = {"IMAGE_CDN":"https://planetscale-images.imgix.net","IMAGE_CDN_ENABLED":"true","INTERNAL_API":"https://api.planetscale.com","RELEASE":"117b8aaf-965c-42bc-b013-5f72770de4d9","SENTRY_DSN":"https://bd81903b44804e22a06bdc0c1a91b303@o499952.ingest.us.sentry.io/4504531942572032"}</script></head><body class="flex min-h-screen flex-col"><div class="bg-neki px-3 py-1 text-center font-medium text-gray-900 dark:font-semibold"><span>Neki, sharded Postgres, is now available.</span> <span class="whitespace-nowrap"><a href="https://auth.planetscale.com/sign-up" class="whitespace-nowrap bg-gray-900 px-sm font-semibold text-white">Get started</a></span></div><header class="relative mb-6 mt-4 bg-primary"><div class="flex flex-col gap-y-3 px-3 sm:px-5 container max-w-7xl"><div class="grid w-full grid-cols-[auto_1fr] grid-rows-1 items-center lg:items-start lg:gap-3"><a aria-label="Go to homepage" class="col-start-1 col-end-2 h-4 w-4 rounded-full text-primary lg:hidden" href="/" data-discover="true"><svg xmlns="http://www.w3.org/2000/svg" width="32" height="32" fill="none" viewBox="0 0 40 40"><path fill="currentColor" d="M0 20C0 8.954 8.954 0 20 0c8.121 0 15.112 4.84 18.245 11.794l-26.45 26.45a20 20 0 0 1-3.225-1.83L24.984 20H20L5.858 34.142A19.94 19.94 0 0 1 0 20M39.999 20.007 20.006 40c11.04-.004 19.99-8.953 19.993-19.993"></path></svg></a><div class="group col-start-2 col-end-3 row-start-1 flex shrink-0 items-center justify-end gap-1.5 lg:gap-3"><div class="flex flex-row gap-2 lg:flex-col lg:gap-1 xl:flex-row"><div class="flex items-center justify-end gap-1 lg:h-4"><a href="https://auth.planetscale.com/sign-in" class="font-semibold text-primary hover:text-orange">Sign in</a></div><div class="flex items-center justify-end gap-0.5 lg:h-4"><form class="btn-sm hidden sm:inline-flex" action="/api/demo-sessions" method="post"><button type="submit" class="btn btn-outline btn-sm hidden sm:inline-flex">View sandbox</button></form><a class="btn btn-sm" href="/contact" data-discover="true">Get in touch</a></div></div></div><div class="col-start-1 col-end-2 flex items-center gap-x-3 lg:row-start-1 lg:h-4"><a aria-label="Go to homepage" class="col-start-1 col-end-2 hidden h-4 w-4 rounded-full text-primary lg:block" href="/" data-discover="true"><svg xmlns="http://www.w3.org/2000/svg" width="32" height="32" fill="none" viewBox="0 0 40 40"><path fill="currentColor" d="M0 20C0 8.954 8.954 0 20 0c8.121 0 15.112 4.84 18.245 11.794l-26.45 26.45a20 20 0 0 1-3.225-1.83L24.984 20H20L5.858 34.142A19.94 19.94 0 0 1 0 20M39.999 20.007 20.006 40c11.04-.004 19.99-8.953 19.993-19.993"></path></svg></a><nav aria-label="Main" data-orientation="horizontal" class="hidden items-center lg:flex"><ul class="flex flex-wrap gap-x-1 md:flex-nowrap"><li><div data-headlessui-state=""><button class="font-semibold text-primary hover:text-contrast focus-visible:ring-0 ui-open:text-orange" type="button" aria-expanded="false" data-headlessui-state="">Platform<span class="ml-sm inline-block ui-open:rotate-180">▾</span></button></div><span hidden="" style="position:fixed;top:1px;left:1px;width:1px;height:0;padding:0;margin:-1px;overflow:hidden;clip:rect(0, 0, 0, 0);white-space:nowrap;border-width:0;display:none"></span></li><li class="text-decoration" role="presentation">|</li><li><div data-headlessui-state=""><button class="font-semibold text-primary hover:text-contrast focus-visible:ring-0 ui-open:text-orange" type="button" aria-expanded="false" data-headlessui-state="">Resources<span class="ml-sm inline-block ui-open:rotate-180">▾</span></button></div><span hidden="" style="position:fixed;top:1px;left:1px;width:1px;height:0;padding:0;margin:-1px;overflow:hidden;clip:rect(0, 0, 0, 0);white-space:nowrap;border-width:0;display:none"></span></li><li class="text-decoration" role="presentation">|</li><li><a class="font-semibold text-primary hover:text-contrast" href="/docs">Documentation</a></li><li class="text-decoration" role="presentation">|</li><li><a class="font-semibold text-primary hover:text-contrast" href="/pricing" data-discover="true">Pricing</a></li><li class="text-decoration" role="presentation">|</li><li><a class="font-semibold text-primary hover:text-contrast" href="/migrate" data-discover="true">Migrate</a></li></ul></nav></div></div><details class="lg:hidden"><summary>Navigation</summary><nav class="dashed-box mt-1 p-3"><ul class="flex flex-wrap gap-x-1 md:flex-nowrap"><li><div data-headlessui-state=""><button class="font-semibold text-primary hover:text-contrast focus-visible:ring-0 ui-open:text-orange" type="button" aria-expanded="false" data-headlessui-state="">Platform<span class="ml-sm inline-block ui-open:rotate-180">▾</span></button></div><span hidden="" style="position:fixed;top:1px;left:1px;width:1px;height:0;padding:0;margin:-1px;overflow:hidden;clip:rect(0, 0, 0, 0);white-space:nowrap;border-width:0;display:none"></span></li><li class="text-decoration" role="presentation">|</li><li><div data-headlessui-state=""><button class="font-semibold text-primary hover:text-contrast focus-visible:ring-0 ui-open:text-orange" type="button" aria-expanded="false" data-headlessui-state="">Resources<span class="ml-sm inline-block ui-open:rotate-180">▾</span></button></div><span hidden="" style="position:fixed;top:1px;left:1px;width:1px;height:0;padding:0;margin:-1px;overflow:hidden;clip:rect(0, 0, 0, 0);white-space:nowrap;border-width:0;display:none"></span></li><li class="text-decoration" role="presentation">|</li><li><a class="font-semibold text-primary hover:text-contrast" href="/docs">Documentation</a></li><li class="text-decoration" role="presentation">|</li><li><a class="font-semibold text-primary hover:text-contrast" href="/pricing" data-discover="true">Pricing</a></li><li class="text-decoration" role="presentation">|</li><li><a class="font-semibold text-primary hover:text-contrast" href="/migrate" data-discover="true">Migrate</a></li></ul></nav></details></div></header><main class="container mb-6 flex max-w-7xl flex-1 flex-col px-3 sm:px-5 lg:px-12"><section class=""><p class="block"><a class="pr-sm text-primary hover:text-contrast" href="/blog" data-discover="true">Blog</a><span class="px-sm text-decoration">|</span><a class="px-sm text-blue hover:bg-blue-100 dark:hover:bg-blue-900" href="/blog/category/engineering" data-discover="true">Engineering</a></p><div class="flex lg:flex-row-reverse lg:gap-x-6"><div class="lg:sticky lg:top-2 lg:self-start"><button class="absolute right-0 bg-gray-100 px-sm md:block lg:hidden dark:bg-gray-800 -mt-9 hidden"><span class="inline">Table of contents «</span><span class="hidden">Close »</span></button><aside class="tree-nav w-full shrink-0 space-y-3 lg:w-36 hidden lg:block"><div><h4 class="text-secondary">Table of contents</h4><ul><li><a class="font-semibold text-primary hover:text-blue" href="/blog/the-feedback-loops-behind-kubernetes#part-1-running-postgres-by-hand" data-discover="true">Part 1: running Postgres by hand</a><ul><li><a class="text-primary hover:text-blue" href="/blog/the-feedback-loops-behind-kubernetes#one-container-one-machine" data-discover="true">One container, one machine</a></li><li><a class="text-primary hover:text-blue" href="/blog/the-feedback-loops-behind-kubernetes#pick-a-node-by-hand" data-discover="true">Pick a node, by hand</a></li><li><a class="text-primary hover:text-blue" href="/blog/the-feedback-loops-behind-kubernetes#it-needs-a-real-disk" data-discover="true">It needs a real disk</a></li><li><a class="text-primary hover:text-blue" href="/blog/the-feedback-loops-behind-kubernetes#one-isnt-enough" data-discover="true">One isn&#x27;t enough</a></li><li><a class="text-primary hover:text-blue" href="/blog/the-feedback-loops-behind-kubernetes#they-have-to-find-each-other" data-discover="true">They have to find each other</a></li><li><a class="text-primary hover:text-blue" href="/blog/the-feedback-loops-behind-kubernetes#the-watchdog-script" data-discover="true">The watchdog script</a></li><li><a class="text-primary hover:text-blue" href="/blog/the-feedback-loops-behind-kubernetes#changing-a-parameter" data-discover="true">Changing a parameter</a></li><li><a class="text-primary hover:text-blue" href="/blog/the-feedback-loops-behind-kubernetes#what-we-actually-built" data-discover="true">What we actually built</a></li></ul></li><li><a class="font-semibold text-primary hover:text-blue" href="/blog/the-feedback-loops-behind-kubernetes#part-2-how-we-reinvented-kubernetes" data-discover="true">Part 2: how we reinvented Kubernetes</a><ul><li><a class="text-primary hover:text-blue" href="/blog/the-feedback-loops-behind-kubernetes#the-other-loops" data-discover="true">The other loops</a></li><li><a class="text-primary hover:text-blue" href="/blog/the-feedback-loops-behind-kubernetes#the-for-loop-translated-to-kubernetes" data-discover="true">The for-loop translated to Kubernetes</a></li><li><a class="text-primary hover:text-blue" href="/blog/the-feedback-loops-behind-kubernetes#edge-triggered-notifications-level-triggered-logic" data-discover="true">Edge-triggered notifications, level-triggered logic</a></li><li><a class="text-primary hover:text-blue" href="/blog/the-feedback-loops-behind-kubernetes#informers-the-work-queue-and-a-cache" data-discover="true">Informers, the work queue, and a cache</a></li><li><a class="text-primary hover:text-blue" href="/blog/the-feedback-loops-behind-kubernetes#setpoint-and-measured-variable-spec-and-status" data-discover="true">Setpoint and measured variable: spec and status</a></li><li><a class="text-primary hover:text-blue" href="/blog/the-feedback-loops-behind-kubernetes#self-healing-by-design" data-discover="true">Self-healing by design</a></li><li><a class="text-primary hover:text-blue" href="/blog/the-feedback-loops-behind-kubernetes#what-observe-actually-means-in-a-real-operator" data-discover="true">What &quot;observe&quot; actually means in a real operator</a></li><li><a class="text-primary hover:text-blue" href="/blog/the-feedback-loops-behind-kubernetes#not-every-edge-comes-from-the-api-server" data-discover="true">Not every edge comes from the API server</a></li></ul></li><li><a class="font-semibold text-primary hover:text-blue" href="/blog/the-feedback-loops-behind-kubernetes#conclusion" data-discover="true">Conclusion</a></li></ul><div class="mb-3 mt-6 border bg-blue-50 p-3 font-semibold text-contrast dark:bg-blue-900"><p>PlanetScale, the fastest cloud Postgres, from $5/month.</p><p><a href="https://app.planetscale.com/new">Start now</a></p></div><p>Get the <a href="/blog/feed.atom">RSS feed</a></p></div></aside></div><article class="min-w-0 flex-grow"><h1>The feedback loops behind Kubernetes</h1><p class="text-secondary"><a class="text-contrast no-underline" href="/blog/author/fatih" data-discover="true">Fatih Arslan</a> <!-- -->[<a class="no-underline hover:bg-blue-100 dark:hover:bg-blue-900" href="https://x.com/fatih" rel="noopener noreferrer" target="_blank" title="@fatih on X">@<!-- -->fatih</a>]<!-- --> |<!-- --> <time dateTime="2026-06-16">June 16, 2026</time></p><div class="blog-post-body"><p>For the last decade, Kubernetes has been the backdrop to most of my work: operating clusters, helping build hosted Kubernetes, and writing Kubernetes operators. At PlanetScale, that now means running stateful systems like Postgres and MySQL in production. Kubernetes has many faces, but here I want to talk about one face only: why it is so good at running workloads at scale.</p><p>People ask me what an operator actually does. The canonical answer is: &quot;it reconciles desired state.&quot; This is correct, but it also tells you almost nothing.</p><p>An operator is a feedback controller. It&#x27;s the same closed loop that runs a thermostat or keeps your car at a fixed speed on cruise control. In our case, the thing being controlled is a database. I have been building these loops for years, and the best way I know to make them click is to ignore Kubernetes at the beginning. Kubernetes is full of control theory, even if we don&#x27;t call it that in the day-to-day.</p><p>Before we look at a single line of Kubernetes, we&#x27;re going to run a production database by hand and slowly let the feedback loop appear on its own. Then we&#x27;ll map that loop to Kubernetes, with the pieces production needs: a store, watches, queues, retries, and more. At the end, we&#x27;ll look at what one of these loops looks like in a real operator.</p><div class="mb-3 border p-3 border-blue-600 dark:border-blue-500"><p><span class="bg-blue-600 px-sm text-white dark:bg-blue-500 dark:text-black">Note</span></p><p>A working understanding of containers and <code>kubectl</code> helps, but you don&#x27;t need to be a Kubernetes expert. I&#x27;ll use terms like <em>idempotent</em>, <em>fan-in</em>, and <em>eventual consistency</em>, and introduce the parts that matter as we go.</p><p>We&#x27;re going to start slow and gradually ramp things up. Each part builds on the previous.</p></div><hr/><h2 id="part-1-running-postgres-by-hand"><a href="#part-1-running-postgres-by-hand">Part 1: running Postgres by hand</a></h2><h3 id="one-container-one-machine"><a href="#one-container-one-machine">One container, one machine</a></h3><p>Let&#x27;s start from scratch. I want to run Postgres on a Linux box, and I need it inside a container. To start it, we run:</p><div class="code-block" data-language="bash"><div class="min-w-0 max-w-full"><pre class="shiki shiki-themes planetscale-light planetscale-dark" style="--shiki-light:#2b2b2b;--shiki-dark:#e1e1e1;--shiki-light-bg:#ebebeb;--shiki-dark-bg:#1a1a1a" tabindex="0"><code><span class="line"><span style="--shiki-light:#F35815;--shiki-dark:#F35815">docker</span><span style="--shiki-light:#414141;--shiki-dark:#C1C1C1"> run</span><span style="--shiki-light:#0B6EC5;--shiki-dark:#73C7F9"> -</span><span style="--shiki-light:#7D5903;--shiki-dark:#FED54A">d</span><span style="--shiki-light:#0B6EC5;--shiki-dark:#73C7F9"> -</span><span style="--shiki-light:#7D5903;--shiki-dark:#FED54A">-name</span><span style="--shiki-light:#414141;--shiki-dark:#C1C1C1"> pg</span><span style="--shiki-light:#7D5903;--shiki-dark:#FED54A"> \</span></span>
<span class="line"><span style="--shiki-light:#0B6EC5;--shiki-dark:#73C7F9"> -</span><span style="--shiki-light:#7D5903;--shiki-dark:#FED54A">e</span><span style="--shiki-light:#414141;--shiki-dark:#C1C1C1"> POSTGRES_PASSWORD=secret</span><span style="--shiki-light:#7D5903;--shiki-dark:#FED54A"> \</span></span>
<span class="line"><span style="--shiki-light:#414141;--shiki-dark:#C1C1C1"> postgres:18</span></span>
<span class="line"></span></code></pre></div></div><p>That&#x27;s it. Postgres is running. My app connects to it, writes some rows, and everything works fine. But then the machine goes away: the cloud provider reclaims the instance (hardware fails, or a spot instance gets taken back), or I ship a new version of my setup, which means stopping the old container and starting a fresh one in its place. Either way, the container is replaced, and my data is gone. The container storage was ephemeral, and I did not attach any persistent volume to it.</p><p>There is already a gap between what I <em>want</em> (Postgres, running, with my data) and what I <em>have</em> (a container whose storage disappears when the container or node goes away). The rest of this post is about that gap and the machinery we build to close it.</p><h3 id="pick-a-node-by-hand"><a href="#pick-a-node-by-hand">Pick a node, by hand</a></h3><p>Imagine we have hundreds of nodes (servers) we can use. I already have other workloads running on them. I need to decide <em>which one</em> runs this database. So I <code>ssh</code> into the box that looks the least busy and start the container there.</p><div class="code-block" data-language="bash"><div class="min-w-0 max-w-full"><pre class="shiki shiki-themes planetscale-light planetscale-dark" style="--shiki-light:#2b2b2b;--shiki-dark:#e1e1e1;--shiki-light-bg:#ebebeb;--shiki-dark-bg:#1a1a1a" tabindex="0"><code><span class="line"><span style="--shiki-light:#F35815;--shiki-dark:#F35815">ssh</span><span style="--shiki-light:#414141;--shiki-dark:#C1C1C1"> node-07</span><span style="--shiki-light:#13862E;--shiki-dark:#75DB8C"> 'docker run -d --name pg ... postgres:18'</span></span>
<span class="line"></span></code></pre></div></div><p>I picked <code>node-07</code> because it looked idle enough. I start keeping track of it, save it in some sort of config file, and push it to some repo.</p><h3 id="it-needs-a-real-disk"><a href="#it-needs-a-real-disk">It needs a real disk</a></h3><p>Container storage is ephemeral, so I have to attach a real block device. In the cloud this is an EBS volume (e.g. on AWS); on bare metal it&#x27;s a physical disk. Assuming it&#x27;s a block device, this is what we usually do: provision the volume, attach it to the node, format it, mount it, and point Postgres&#x27; data directory at the mount.</p><div class="code-block" data-language="bash"><div class="min-w-0 max-w-full"><pre class="shiki shiki-themes planetscale-light planetscale-dark" style="--shiki-light:#2b2b2b;--shiki-dark:#e1e1e1;--shiki-light-bg:#ebebeb;--shiki-dark-bg:#1a1a1a" tabindex="0"><code><span class="line"><span style="--shiki-light:#818181;--shiki-dark:#A1A1A1"># provision + attach first with cloud CLI, then on the node:</span></span>
<span class="line"><span style="--shiki-light:#F35815;--shiki-dark:#F35815">mkfs.ext4</span><span style="--shiki-light:#414141;--shiki-dark:#C1C1C1"> /dev/nvme1n1</span></span>
<span class="line"><span style="--shiki-light:#F35815;--shiki-dark:#F35815">mkdir</span><span style="--shiki-light:#0B6EC5;--shiki-dark:#73C7F9"> -</span><span style="--shiki-light:#7D5903;--shiki-dark:#FED54A">p</span><span style="--shiki-light:#414141;--shiki-dark:#C1C1C1"> /var/lib/pg-data</span></span>
<span class="line"><span style="--shiki-light:#F35815;--shiki-dark:#F35815">mount</span><span style="--shiki-light:#414141;--shiki-dark:#C1C1C1"> /dev/nvme1n1</span><span style="--shiki-light:#414141;--shiki-dark:#C1C1C1"> /var/lib/pg-data</span></span>
<span class="line"><span style="--shiki-light:#F35815;--shiki-dark:#F35815">docker</span><span style="--shiki-light:#414141;--shiki-dark:#C1C1C1"> run</span><span style="--shiki-light:#0B6EC5;--shiki-dark:#73C7F9"> -</span><span style="--shiki-light:#7D5903;--shiki-dark:#FED54A">d</span><span style="--shiki-light:#0B6EC5;--shiki-dark:#73C7F9"> -</span><span style="--shiki-light:#7D5903;--shiki-dark:#FED54A">-name</span><span style="--shiki-light:#414141;--shiki-dark:#C1C1C1"> pg</span><span style="--shiki-light:#7D5903;--shiki-dark:#FED54A"> \</span></span>
<span class="line"><span style="--shiki-light:#0B6EC5;--shiki-dark:#73C7F9"> -</span><span style="--shiki-light:#7D5903;--shiki-dark:#FED54A">e</span><span style="--shiki-light:#414141;--shiki-dark:#C1C1C1"> POSTGRES_PASSWORD=secret</span><span style="--shiki-light:#7D5903;--shiki-dark:#FED54A"> \</span></span>
<span class="line"><span style="--shiki-light:#0B6EC5;--shiki-dark:#73C7F9"> -</span><span style="--shiki-light:#7D5903;--shiki-dark:#FED54A">v</span><span style="--shiki-light:#414141;--shiki-dark:#C1C1C1"> /var/lib/pg-data:/var/lib/postgresql</span><span style="--shiki-light:#7D5903;--shiki-dark:#FED54A"> \</span></span>
<span class="line"><span style="--shiki-light:#414141;--shiki-dark:#C1C1C1"> postgres:18</span></span>
<span class="line"></span></code></pre></div></div><p>These are a lot of steps, and each one can fail halfway. And if the disk fills up later, Postgres stops accepting writes and we have to resize the volume by hand: first through the cloud provider, then again inside the filesystem.</p><h3 id="one-isnt-enough"><a href="#one-isnt-enough">One isn&#x27;t enough</a></h3><p>A single Postgres instance is a single point of failure. We want high availability: one primary and two replicas. These need to be on three different machines, with streaming replication between them. So we do the same steps again, three times, on <code>node-07</code>, <code>node-12</code>, and <code>node-19</code>. I also wire up replication by hand: <code>primary_conninfo</code>, replication slots, all of it.</p><p>Now we have three nodes with three Postgres instances. One of the instances is the primary (here it&#x27;s <code>node-07</code>). But this raises new problems, like what to do if the primary&#x27;s node dies?</p><h3 id="they-have-to-find-each-other"><a href="#they-have-to-find-each-other">They have to find each other</a></h3><p>Here is another thing we have to solve. The replicas need to reach the primary, and the primary needs to accept their connections. And every one of these addresses is an IP that changes when a container restarts.</p><p>The first thing I do is hard-code the IPs. I write <code>node-07</code>&#x27;s address into the replicas&#x27; config, I list the replicas&#x27; addresses in the primary&#x27;s <code>pg_hba.conf</code>, and I keep a small <code>/etc/hosts</code> table and save it somewhere.</p><div class="code-block" data-language="conf"><div class="min-w-0 max-w-full"><pre class="shiki shiki-themes planetscale-light planetscale-dark" style="--shiki-light:#2b2b2b;--shiki-dark:#e1e1e1;--shiki-light-bg:#ebebeb;--shiki-dark-bg:#1a1a1a" tabindex="0"><code><span class="line"><span># on each replica's postgresql.auto.conf, until the primary is recreated with a new IP</span></span>
<span class="line"><span>primary_conninfo = 'host=10.4.7.21 port=5432 user=replicator ...'</span></span>
<span class="line"><span></span></span></code></pre></div></div><p>But we still have a problem: the first time the primary is recreated with a different IP, the whole cluster falls apart.</p><p><button type="button" aria-haspopup="dialog" aria-expanded="false" aria-label="Enlarge image: Manual Postgres cluster diagram" class="focus-visible-ring group relative block w-fit max-w-[min(100%,800px)] cursor-zoom-in text-left [&amp;_picture]:contents"><picture class="block"><source media="(prefers-color-scheme: light), (prefers-color-scheme: no-preference)" srcSet="https://planetscale-images.imgix.net/assets/part1-manual-cluster-C6iQCdoL.png?auto=compress%2Cformat"/><source media="(prefers-color-scheme: dark)" srcSet="https://planetscale-images.imgix.net/assets/part1-manual-cluster-darkmode-VBZiV-Rl.png?auto=compress%2Cformat"/><img alt="Manual Postgres cluster diagram" src="https://planetscale-images.imgix.net/assets/part1-manual-cluster-C6iQCdoL.png?auto=compress%2Cformat" width="1504" height="1007" loading="lazy" class="w-auto max-w-full"/></picture><span aria-hidden="true" class="pointer-events-none absolute right-1 top-1 z-10 flex h-5 w-5 items-center justify-center border border-white/25 bg-black/70 text-white backdrop-blur-sm transition-colors transition-opacity group-hover:bg-black/90 group-hover:opacity-100 group-focus-visible:opacity-100 motion-reduce:transition-none [@media(hover:hover)_and_(pointer:fine)]:opacity-0"><svg width="14" height="14" viewBox="0 0 16 16" fill="none" stroke="currentColor" stroke-width="1.5" stroke-linecap="round" stroke-linejoin="round"><path d="M10 2h4v4M6 14H2v-4M14 2l-4.5 4.5M2 14l4.5-4.5"></path></svg></span></button><span hidden="" style="position:fixed;top:1px;left:1px;width:1px;height:0;padding:0;margin:-1px;overflow:hidden;clip:rect(0, 0, 0, 0);white-space:nowrap;border-width:0;display:none"></span></p><h3 id="the-watchdog-script"><a href="#the-watchdog-script">The watchdog script</a></h3><p>Now, this is where we start thinking about how to solve these issues. Everything described so far can break, and will continue to break even if I fix it:</p><ul><li>A replica process dies and doesn&#x27;t come back.</li><li>A disk gets full.</li><li>The primary fails and a replica has to be promoted.</li><li>A config I changed on two nodes but forgot on the third one. They are now out of sync.</li></ul><p>Let&#x27;s assume we&#x27;ve set up a simple uptime monitor and we&#x27;re going to get paged for all these cases. To avoid getting paged at night, we do the sensible thing: write a script. So we decide to write a loop that wakes up every few seconds, looks at each node, and fixes whatever&#x27;s wrong.</p><div class="code-block" data-language="bash"><div class="min-w-0 max-w-full"><pre class="shiki shiki-themes planetscale-light planetscale-dark" style="--shiki-light:#2b2b2b;--shiki-dark:#e1e1e1;--shiki-light-bg:#ebebeb;--shiki-dark-bg:#1a1a1a" tabindex="0"><code><span class="line"><span style="--shiki-light:#F35815;--shiki-dark:#F35815">while</span><span style="--shiki-light:#0B6EC5;--shiki-dark:#73C7F9"> true</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">;</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> do</span></span>
<span class="line"><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> for</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> node</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> in</span><span style="--shiki-light:#414141;--shiki-dark:#C1C1C1"> node-07</span><span style="--shiki-light:#414141;--shiki-dark:#C1C1C1"> node-12</span><span style="--shiki-light:#414141;--shiki-dark:#C1C1C1"> node-19</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">;</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> do</span></span>
<span class="line"><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> if</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> !</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> ssh</span><span style="--shiki-light:#13862E;--shiki-dark:#75DB8C"> "</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">$node</span><span style="--shiki-light:#13862E;--shiki-dark:#75DB8C">"</span><span style="--shiki-light:#13862E;--shiki-dark:#75DB8C"> 'pg_isready -q'</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">;</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> then</span></span>
<span class="line"><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> ssh</span><span style="--shiki-light:#13862E;--shiki-dark:#75DB8C"> "</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">$node</span><span style="--shiki-light:#13862E;--shiki-dark:#75DB8C">"</span><span style="--shiki-light:#13862E;--shiki-dark:#75DB8C"> 'docker start pg'</span><span style="--shiki-light:#818181;--shiki-dark:#A1A1A1"> # it died, bring it back</span></span>
<span class="line"><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> fi</span></span>
<span class="line"></span>
<span class="line"><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> usage</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815">=</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">$(</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815">ssh</span><span style="--shiki-light:#13862E;--shiki-dark:#75DB8C"> "</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">$node</span><span style="--shiki-light:#13862E;--shiki-dark:#75DB8C">"</span><span style="--shiki-light:#13862E;--shiki-dark:#75DB8C"> "df --output=pcent /var/lib/pg-data | tail -1 | tr -dc 0-9"</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">)</span></span>
<span class="line"><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> if</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1"> [</span><span style="--shiki-light:#13862E;--shiki-dark:#75DB8C"> "</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">$usage</span><span style="--shiki-light:#13862E;--shiki-dark:#75DB8C">"</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> -gt</span><span style="--shiki-light:#D92038;--shiki-dark:#FF7082"> 80</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1"> ];</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> then</span></span>
<span class="line"><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> grow_volume</span><span style="--shiki-light:#13862E;--shiki-dark:#75DB8C"> "</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">$node</span><span style="--shiki-light:#13862E;--shiki-dark:#75DB8C">"</span><span style="--shiki-light:#818181;--shiki-dark:#A1A1A1"> # disk filling, make it bigger</span></span>
<span class="line"><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> fi</span></span>
<span class="line"><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> done</span></span>
<span class="line"><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> sleep</span><span style="--shiki-light:#D92038;--shiki-dark:#FF7082"> 5</span></span>
<span class="line"><span style="--shiki-light:#F35815;--shiki-dark:#F35815">done</span></span>
<span class="line"></span></code></pre></div></div><p>It&#x27;s written in Bash, and probably has tons of bugs. You notice something here? The loop doesn&#x27;t care <em>how</em> the database got into a bad state. Every five seconds it looks at the current state of the world and asks this question: does reality match what I want?</p><p>If a process is down, start it. If a disk is filling, grow it. Run the loop once or run it a thousand times and the result is the same, because each action is conditional on the current state. The script is <em>idempotent</em>.</p><h3 id="changing-a-parameter"><a href="#changing-a-parameter">Changing a parameter</a></h3><p>Let&#x27;s make things a little more complex. I need to raise <a href="https://www.postgresql.org/docs/current/runtime-config-connection.html#GUC-MAX-CONNECTIONS"><code>max_connections</code></a> from 100 to 500. This one is not a reload-only change. PostgreSQL says it can only be set at server start, so the manual version is to <code>ssh</code> into each box, edit <code>postgresql.conf</code>, restart Postgres, and check that it took on all three.</p><p>Because I know that ssh&#x27;ing into the nodes manually isn&#x27;t a thing I want anymore, I do the same thing we did previously: I write the desired value down in one place and teach the loop to enforce it.</p><div class="code-block" data-language="bash"><div class="min-w-0 max-w-full"><pre class="shiki shiki-themes planetscale-light planetscale-dark" style="--shiki-light:#2b2b2b;--shiki-dark:#e1e1e1;--shiki-light-bg:#ebebeb;--shiki-dark-bg:#1a1a1a" tabindex="0"><code><span class="line"><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">WANT_MAX_CONNECTIONS</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815">=</span><span style="--shiki-light:#414141;--shiki-dark:#C1C1C1">500</span></span>
<span class="line"></span>
<span class="line"><span style="--shiki-light:#F35815;--shiki-dark:#F35815">for</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> node</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> in</span><span style="--shiki-light:#414141;--shiki-dark:#C1C1C1"> node-07</span><span style="--shiki-light:#414141;--shiki-dark:#C1C1C1"> node-12</span><span style="--shiki-light:#414141;--shiki-dark:#C1C1C1"> node-19</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">;</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> do</span></span>
<span class="line"><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> have</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815">=</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">$(</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815">ssh</span><span style="--shiki-light:#13862E;--shiki-dark:#75DB8C"> "</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">$node</span><span style="--shiki-light:#13862E;--shiki-dark:#75DB8C">"</span><span style="--shiki-light:#13862E;--shiki-dark:#75DB8C"> "psql -tAc 'show max_connections'"</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">)</span></span>
<span class="line"><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> if</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1"> [</span><span style="--shiki-light:#13862E;--shiki-dark:#75DB8C"> "</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">$have</span><span style="--shiki-light:#13862E;--shiki-dark:#75DB8C">"</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> !=</span><span style="--shiki-light:#13862E;--shiki-dark:#75DB8C"> "</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">$WANT_MAX_CONNECTIONS</span><span style="--shiki-light:#13862E;--shiki-dark:#75DB8C">"</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1"> ];</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> then</span></span>
<span class="line"><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> ssh</span><span style="--shiki-light:#13862E;--shiki-dark:#75DB8C"> "</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">$node</span><span style="--shiki-light:#13862E;--shiki-dark:#75DB8C">"</span><span style="--shiki-light:#13862E;--shiki-dark:#75DB8C"> "sed -i 's/^max_connections.*/max_connections = </span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">$WANT_MAX_CONNECTIONS</span><span style="--shiki-light:#13862E;--shiki-dark:#75DB8C">/' /var/lib/pg-data/postgresql.conf"</span></span>
<span class="line"><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> ssh</span><span style="--shiki-light:#13862E;--shiki-dark:#75DB8C"> "</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">$node</span><span style="--shiki-light:#13862E;--shiki-dark:#75DB8C">"</span><span style="--shiki-light:#13862E;--shiki-dark:#75DB8C"> "docker restart pg"</span></span>
<span class="line"><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> fi</span></span>
<span class="line"><span style="--shiki-light:#F35815;--shiki-dark:#F35815">done</span></span>
<span class="line"></span></code></pre></div></div><p>This is the same idea as before. I read what I <em>want</em> (a variable). Observe what I <em>have</em> (a query). If they differ, I take an action to close the difference. Again, I don&#x27;t track whether I changed it last time. All I do is compare and <a href="https://dictionary.cambridge.org/dictionary/english/converge">converge</a>, every loop.</p><h3 id="what-we-actually-built"><a href="#what-we-actually-built">What we actually built</a></h3><p>I started with a desired state that was written down in one place: three instances, this disk size, <code>max_connections = 500</code>. Every few seconds I observe the actual state of the system. I compute the difference. I take whatever action closes that difference. Then I do it again, forever.</p><p>That&#x27;s a <strong>closed feedback loop</strong>. The word &quot;closed&quot; matters. It means the output of the system is fed back into the next decision. I don&#x27;t run <code>docker start</code> and assume the database is fine. I check the database again. If it is still wrong, I act again. If it is already correct, I do nothing.</p><p>The nice part is that the same loop works for different problems. It can restart a dead process, grow a disk, or push <code>max_connections = 500</code>. The action changes, but the shape stays the same: read what I want, observe what I have, compare them, act, repeat. If I draw the same thing as a block diagram, with the control theory names added, it would look like this:</p><p><button type="button" aria-haspopup="dialog" aria-expanded="false" aria-label="Enlarge image: Closed feedback loop diagram" class="focus-visible-ring group relative block w-fit max-w-[min(100%,800px)] cursor-zoom-in text-left [&amp;_picture]:contents"><picture class="block"><source media="(prefers-color-scheme: light), (prefers-color-scheme: no-preference)" srcSet="https://planetscale-images.imgix.net/assets/part1-feedback-loop-Z0q4qAPX.png?auto=compress%2Cformat"/><source media="(prefers-color-scheme: dark)" srcSet="https://planetscale-images.imgix.net/assets/part1-feedback-loop-darkmode-BeMyZLwB.png?auto=compress%2Cformat"/><img alt="Closed feedback loop diagram" src="https://planetscale-images.imgix.net/assets/part1-feedback-loop-Z0q4qAPX.png?auto=compress%2Cformat" width="1504" height="564" loading="lazy" class="w-auto max-w-full"/></picture><span aria-hidden="true" class="pointer-events-none absolute right-1 top-1 z-10 flex h-5 w-5 items-center justify-center border border-white/25 bg-black/70 text-white backdrop-blur-sm transition-colors transition-opacity group-hover:bg-black/90 group-hover:opacity-100 group-focus-visible:opacity-100 motion-reduce:transition-none [@media(hover:hover)_and_(pointer:fine)]:opacity-0"><svg width="14" height="14" viewBox="0 0 16 16" fill="none" stroke="currentColor" stroke-width="1.5" stroke-linecap="round" stroke-linejoin="round"><path d="M10 2h4v4M6 14H2v-4M14 2l-4.5 4.5M2 14l4.5-4.5"></path></svg></span></button><span hidden="" style="position:fixed;top:1px;left:1px;width:1px;height:0;padding:0;margin:-1px;overflow:hidden;clip:rect(0, 0, 0, 0);white-space:nowrap;border-width:0;display:none"></span></p><p>Here is how the vocabulary from <a href="https://en.wikipedia.org/wiki/Control_theory">control theory</a> maps cleanly onto my shell script:</p><ul><li>The <strong>setpoint</strong> is my desired state, the variables at the top of the script (disk size, max_connections and so on).</li><li>The <strong>measured output</strong> is what I observe: <code>pg_isready</code>, <code>df</code>, <code>show max_connections</code>.</li><li>The <strong>error</strong> (e) is the difference between them.</li><li>The <strong>controller</strong> is the body of the loop, the <code>if</code> statements that decide what to do. It is not the whole script.</li><li>The <strong>actuator</strong> is what carries out the action: <code>ssh</code> plus <code>docker start</code>.</li><li>The <strong>plant</strong> is the system being controlled, Postgres and its disk.</li></ul><p>That also gives us a nice way to understand <strong>open-loop</strong> control. My very first attempt, <code>ssh</code> in, run the command, and walk away, was open-loop: fire an action and assume it worked. The Bash script is closed-loop because it keeps feeding the measured state back into the next decision.</p><p>A Bash loop is not a production control plane. Just to name a few issues with it:</p><ul><li>It has no concurrency control, so two copies of the script can race each other. Imagine both deciding to promote a different replica.</li><li>It keeps its only real state, &quot;am I mid-failover?&quot;, in a shell variable that could die with the process.</li><li>It polls every node every five seconds whether anything changed or not, which is fine for three nodes, but too expensive for three thousand nodes.</li><li>It has no idea what to do when the <code>ssh</code> itself times out.</li><li>And the moment I want a second kind of resource, a connection pooler, a backup job, a read replica in another region, I&#x27;m copy-pasting this whole structure.</li></ul><p>What if the script also fails? Who runs it then? We could keep hardening this script, but look at where it goes: we would need a real store for the desired state, watches instead of polling, a work queue, retries, leader election. We would be rebuilding Kubernetes. The real platform already exists, and it&#x27;s Kubernetes.</p><hr/><h2 id="part-2-how-we-reinvented-kubernetes"><a href="#part-2-how-we-reinvented-kubernetes">Part 2: how we reinvented Kubernetes</a></h2><p>Now we can map what we hand-rolled in Part 1 to Kubernetes. Almost all of it already exists there. The operator is the part we care about.</p><h3 id="the-other-loops"><a href="#the-other-loops">The other loops</a></h3><p>Let&#x27;s go through some of the pieces we built by hand before the watchdog loop. You already know these components by name. What you might not have noticed is that they also work like controllers.</p><p><strong>Spinning up the container: the kubelet.</strong> First, a quick definition: a Pod is the smallest thing Kubernetes runs, one or more containers scheduled together on a node and sharing its network. For us it&#x27;s the Postgres container. On every node runs an agent called the kubelet. Its desired state is the set of Pods assigned to its node, which it learns from the API server. Its observed state is the set of containers actually running, which it gets from the container runtime. When they differ, it starts the missing container, kills the extra one, or restarts the crashed one. My <code>if ! pg_isready; then docker start; fi</code> is the kubelet&#x27;s job, just done properly. The kubelet doesn&#x27;t shell into anything; it talks to containerd over a gRPC socket, which talks to runc.</p><p><strong>Picking a node: the scheduler.</strong> Remember me choosing <code>node-07</code>? That&#x27;s the scheduler&#x27;s whole reason to exist. It watches for Pods with no node assigned, filters out the nodes that can&#x27;t work, scores the rest, and writes the decision to one field: <code>pod.Spec.NodeName</code>. The scheduler doesn&#x27;t start the container; it records the placement and lets the kubelet pick it up. You will realize that most things in Kubernetes are decoupled like this.</p><p><strong>Attaching the disk: CSI and the PV/PVC sync.</strong> My multi-step <code>mkfs</code> and <code>mount</code> script becomes a <code>PersistentVolumeClaim</code>, which is a declarative request for storage. The <a href="https://github.com/container-storage-interface/spec/blob/master/spec.md">Container Storage Interface (CSI)</a> driver turns that request into a real volume. CSI itself is a set of controllers and sidecars: one provisions, one attaches, one resizes, and so on, while the kubelet calls the driver&#x27;s node plugin to do the actual mount. It&#x27;s a family of controllers. If there is a PVC but no disk behind it, one controller creates the disk. If the PVC size increases, another controller calls the provider API (e.g., AWS <code>ModifyVolume</code>). Again, I write intent, and a controller does the actual work. (note: I wrote one of the early production CSI drivers, <a href="https://github.com/digitalocean/csi-digitalocean">csi-digitalocean</a>, and a <a href="https://arslan.io/2018/06/21/how-to-write-a-container-storage-interface-csi-plugin/">long post about building one</a>.)</p><p><strong>Making them find each other: the CNI and Services.</strong> The <code>/etc/hosts</code> problem is solved at a layer we no longer have to think about. A CNI plugin gives Pods their network identity; Cilium, for example, does this with eBPF instead of a pile of <code>iptables</code> rules. For stateful workloads, a StatefulSet plus a <a href="https://kubernetes.io/docs/concepts/services-networking/service/#headless-services">headless Service</a> gives each replica its own stable DNS name, which is exactly what a Postgres replica needs. The hard-coded IP that broke our cluster becomes a name that keeps working. DNS is only one tool here; other service discovery systems like etcd, ZooKeeper, and Consul solve similar problems.</p><p>As you see, all the problems we solved with <code>ssh</code> and various scripts are replaced by Kubernetes components and drivers. And these are just a few of them:</p><p><button type="button" aria-haspopup="dialog" aria-expanded="false" aria-label="Enlarge image: Mapping manual operations to Kubernetes controllers" class="focus-visible-ring group relative block w-fit max-w-[min(100%,800px)] cursor-zoom-in text-left [&amp;_picture]:contents"><picture class="block"><source media="(prefers-color-scheme: light), (prefers-color-scheme: no-preference)" srcSet="https://planetscale-images.imgix.net/assets/part2-kubernetes-mapping-BqGHE4E0.png?auto=compress%2Cformat"/><source media="(prefers-color-scheme: dark)" srcSet="https://planetscale-images.imgix.net/assets/part2-kubernetes-mapping-darkmode-CAz9s2jY.png?auto=compress%2Cformat"/><img alt="Mapping manual operations to Kubernetes controllers" src="https://planetscale-images.imgix.net/assets/part2-kubernetes-mapping-BqGHE4E0.png?auto=compress%2Cformat" width="1504" height="883" loading="lazy" class="w-auto max-w-full"/></picture><span aria-hidden="true" class="pointer-events-none absolute right-1 top-1 z-10 flex h-5 w-5 items-center justify-center border border-white/25 bg-black/70 text-white backdrop-blur-sm transition-colors transition-opacity group-hover:bg-black/90 group-hover:opacity-100 group-focus-visible:opacity-100 motion-reduce:transition-none [@media(hover:hover)_and_(pointer:fine)]:opacity-0"><svg width="14" height="14" viewBox="0 0 16 16" fill="none" stroke="currentColor" stroke-width="1.5" stroke-linecap="round" stroke-linejoin="round"><path d="M10 2h4v4M6 14H2v-4M14 2l-4.5 4.5M2 14l4.5-4.5"></path></svg></span></button><span hidden="" style="position:fixed;top:1px;left:1px;width:1px;height:0;padding:0;margin:-1px;overflow:hidden;clip:rect(0, 0, 0, 0);white-space:nowrap;border-width:0;display:none"></span></p><p>All of this so far is useful context, but the part we care about is <em>our watchdog</em> loop, because that&#x27;s the one we get to write ourselves.</p><p>That&#x27;s the operator.</p><h3 id="the-for-loop-translated-to-kubernetes"><a href="#the-for-loop-translated-to-kubernetes">The for-loop translated to Kubernetes</a></h3><p>In Kubernetes, our watchdog script is a <strong>controller</strong>, and the standard way to write one in Go is a library called <a href="https://github.com/kubernetes-sigs/controller-runtime">controller-runtime</a>. At its heart, it&#x27;s a function with a basic signature:</p><div class="code-block" data-language="go"><div class="min-w-0 max-w-full"><pre class="shiki shiki-themes planetscale-light planetscale-dark" style="--shiki-light:#2b2b2b;--shiki-dark:#e1e1e1;--shiki-light-bg:#ebebeb;--shiki-dark-bg:#1a1a1a" tabindex="0"><code><span class="line"><span style="--shiki-light:#F35815;--shiki-dark:#F35815">func</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1"> (</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">r </span><span style="--shiki-light:#F35815;--shiki-dark:#F35815">*Reconciler</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">)</span><span style="--shiki-light:#5E49AF;--shiki-dark:#B7A5FB"> Reconcile</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">(</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">ctx</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> context</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815">Context</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">,</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> req</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> reconcile</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815">Request</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">)</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1"> (</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815">reconcile</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815">Result</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">,</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> error</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">)</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1"> {</span></span>
<span class="line"><span style="--shiki-light:#818181;--shiki-dark:#A1A1A1"> // req contains a namespace/name. That's it. That's the whole input.</span></span>
<span class="line"><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">}</span></span>
<span class="line"></span></code></pre></div></div><p>Notice what is missing here: the function isn&#x27;t told what changed. There is no diff. It isn&#x27;t handed the old object and the new object. It isn&#x27;t given an event type. It gets a key, a namespace and a name, and nothing else. It&#x27;s minimal by design, because it has to work for many different controllers. The function&#x27;s job is to fetch the object with that namespace/name, look at the world, and converge to the desired state.</p><h3 id="edge-triggered-notifications-level-triggered-logic"><a href="#edge-triggered-notifications-level-triggered-logic">Edge-triggered notifications, level-triggered logic</a></h3><p>There are two ways to build any closed feedback loop:</p><ul><li><strong>Edge-triggered</strong>: act on transitions, on events. &quot;The disk crossed 80%.&quot; &quot;The Pod was deleted.&quot; &quot;The number of replicas increased by 2.&quot;</li><li><strong>Level-triggered</strong>: act on the current state, regardless of how you got there. &quot;The disk <em>is</em> at 85%.&quot; &quot;The Pod <em>is</em> missing.&quot; &quot;The number of replicas is 3.&quot;</li></ul><p>My first mental model of controllers, and probably yours at some point, was edge-triggered: listen to a stream of changes, and for each change, try to converge.</p><p>The problem is that this is very fragile. In distributed systems, if one component is fragile, the fragility spreads to the rest of the system. Why is edge-triggering fragile? Say your controller is down for thirty seconds. It misses the events from those thirty seconds, and its view of the world is now permanently wrong. If two events arrive out of order, you process them out of order. If an event is delivered twice, you act twice. You&#x27;re rebuilding your state from a stream of events, and you&#x27;ve inherited all of event sourcing&#x27;s hard problems.</p><p>Here is a very concrete example. Assume you have 1 replica, and you increase it to 3 replicas. Because you have only subscribed to changes, either:</p><ol><li>You miss the event (maybe the queue dropped it, or the consumer, your app, dropped it due to a crash or a full buffer).</li><li>You receive it twice.</li></ol><p>In the first case, you won&#x27;t be able to self-correct. In the second case, if your handler blindly applies the delta again, you&#x27;ll end up with 5 replicas (you overshoot), instead of 3.</p><p>The level-triggered model fixes all of that. Remember, our shell script never asked &quot;what changed?&quot; It asked &quot;what <em>is</em> true right now?&quot;, every five seconds, from scratch. Miss a loop, and the next one catches up. Run the loop twice, and you get the same result. The current state of the world is the only input that matters, and it&#x27;s always available to read. So in the level-triggered case, our example above becomes this: you read <code>replicas=3</code>, you check the current number of replicas, which is 1, and you increase by 2.</p><p>If you miss the event, no one cares. In the next reconcile loop you&#x27;ll catch it. If your app crashes, it comes back, reads again and detects that it did not increase it yet, increases it.</p><p>Kubernetes controllers combine both: <strong>edge-triggered notifications, level-triggered logic</strong>.</p><p>Events (the edges) are only a hint that it&#x27;s worth looking again. They tell you <em>when</em> to reconcile, never <em>what</em> to do. The reconcile itself is level-based: it reads the current state (e.g., <code>replicas=3</code>) and drives toward the desired state (e.g., <code>create 2 replicas</code>), ignoring the triggering event completely. That&#x27;s <em>why</em> <code>Reconcile</code> only gets a key. The framework makes it hard to write edge-triggered logic, on purpose. Edge-triggered logic is how you get a controller that&#x27;s fragile and permanently wrong after its first hiccup.</p><p><button type="button" aria-haspopup="dialog" aria-expanded="false" aria-label="Enlarge image: Edge-triggered versus level-triggered scaling" class="focus-visible-ring group relative block w-fit max-w-[min(100%,800px)] cursor-zoom-in text-left [&amp;_picture]:contents"><picture class="block"><source media="(prefers-color-scheme: light), (prefers-color-scheme: no-preference)" srcSet="https://planetscale-images.imgix.net/assets/part2-edge-vs-level-BUiD8D-7.png?auto=compress%2Cformat"/><source media="(prefers-color-scheme: dark)" srcSet="https://planetscale-images.imgix.net/assets/part2-edge-vs-level-darkmode-3DNb92aw.png?auto=compress%2Cformat"/><img alt="Edge-triggered versus level-triggered scaling" src="https://planetscale-images.imgix.net/assets/part2-edge-vs-level-BUiD8D-7.png?auto=compress%2Cformat" width="1504" height="927" loading="lazy" class="w-auto max-w-full"/></picture><span aria-hidden="true" class="pointer-events-none absolute right-1 top-1 z-10 flex h-5 w-5 items-center justify-center border border-white/25 bg-black/70 text-white backdrop-blur-sm transition-colors transition-opacity group-hover:bg-black/90 group-hover:opacity-100 group-focus-visible:opacity-100 motion-reduce:transition-none [@media(hover:hover)_and_(pointer:fine)]:opacity-0"><svg width="14" height="14" viewBox="0 0 16 16" fill="none" stroke="currentColor" stroke-width="1.5" stroke-linecap="round" stroke-linejoin="round"><path d="M10 2h4v4M6 14H2v-4M14 2l-4.5 4.5M2 14l4.5-4.5"></path></svg></span></button><span hidden="" style="position:fixed;top:1px;left:1px;width:1px;height:0;padding:0;margin:-1px;overflow:hidden;clip:rect(0, 0, 0, 0);white-space:nowrap;border-width:0;display:none"></span></p><p>Our bash script stumbled into this property by accident, at least for the sake of the example. But the <code>controller-runtime</code> framework gives it to you on purpose. It&#x27;s why a Kubernetes controller can crash, get restarted ten minutes later, and converge correctly with no special recovery code. There is no recovery code. There is just the loop. The controller can reconstruct the world from scratch.</p><h3 id="informers-the-work-queue-and-a-cache"><a href="#informers-the-work-queue-and-a-cache">Informers, the work queue, and a cache</a></h3><p>So where do the edges come from? And what stops a controller from DDoSing the API server by listing everything every five seconds like my script did?</p><p>The answer is the <strong>informer</strong>. An informer opens a single watch against the API server for a given resource type, streams every add, update, and delete, and keeps a complete in-memory <strong>cache</strong> of the objects we&#x27;re interested in. Two things matter here:</p><p>First, the informer turns each watch event into a key and puts it on a <strong>work queue</strong>. The queue does a lot of work for you.</p><ul><li>It <em>coalesces</em>: if the same object is updated five times before you get to it, you reconcile it once, against the latest state (level-triggered again).</li><li>It <em>rate-limits</em>: an object that keeps erroring backs off exponentially instead of spinning. This is <a href="https://en.wikipedia.org/wiki/Damping">damping</a>, the same reason a crash-looping container backs off instead of restarting hot.</li><li>It lets you run a pool of workers pulling keys in parallel, which is your fan-out. Events fan in from the watch, collapse in the queue, and fan out to the workers. This is something you need to tune. The higher you set the pool, the more pressure you put on the system: more writes, more API calls, more load on the provider, and more CPU usage in the operator.</li></ul><p><button type="button" aria-haspopup="dialog" aria-expanded="false" aria-label="Enlarge image: Informer, work queue, and controller diagram" class="focus-visible-ring group relative block w-fit max-w-[min(100%,800px)] cursor-zoom-in text-left [&amp;_picture]:contents"><picture class="block"><source media="(prefers-color-scheme: light), (prefers-color-scheme: no-preference)" srcSet="https://planetscale-images.imgix.net/assets/part2-informer-queue-yyOXLWWC.png?auto=compress%2Cformat"/><source media="(prefers-color-scheme: dark)" srcSet="https://planetscale-images.imgix.net/assets/part2-informer-queue-darkmode-ANBXf_4R.png?auto=compress%2Cformat"/><img alt="Informer, work queue, and controller diagram" src="https://planetscale-images.imgix.net/assets/part2-informer-queue-yyOXLWWC.png?auto=compress%2Cformat" width="1504" height="1007" loading="lazy" class="w-auto max-w-full"/></picture><span aria-hidden="true" class="pointer-events-none absolute right-1 top-1 z-10 flex h-5 w-5 items-center justify-center border border-white/25 bg-black/70 text-white backdrop-blur-sm transition-colors transition-opacity group-hover:bg-black/90 group-hover:opacity-100 group-focus-visible:opacity-100 motion-reduce:transition-none [@media(hover:hover)_and_(pointer:fine)]:opacity-0"><svg width="14" height="14" viewBox="0 0 16 16" fill="none" stroke="currentColor" stroke-width="1.5" stroke-linecap="round" stroke-linejoin="round"><path d="M10 2h4v4M6 14H2v-4M14 2l-4.5 4.5M2 14l4.5-4.5"></path></svg></span></button><span hidden="" style="position:fixed;top:1px;left:1px;width:1px;height:0;padding:0;margin:-1px;overflow:hidden;clip:rect(0, 0, 0, 0);white-space:nowrap;border-width:0;display:none"></span></p><p>Second, and this is a detail that bites people a lot: <strong>your reads and your writes in Kubernetes don&#x27;t go to the same place.</strong></p><p>In controller-runtime, the client you&#x27;re handed reads from the informer&#x27;s local cache. Cache reads are cheap, they don&#x27;t touch the API server, and that&#x27;s how a controller reconciles thousands of objects without falling over. But your writes go straight to the API server. The cache only learns about your write when the resulting watch event comes back around, a moment later.</p><p>Because of that, a read can be stale. You need to be prepared for this.</p><p>If you write a field of an object and then read the same object again from the cache, the reconciler might think it&#x27;s not updated yet. You write again, and you get a Conflict error. Retrying with a fresh read can be fine, but blindly retrying against the same stale cached view just spins.</p><p>Most of the time, what you want is to drop the call and <em>requeue</em>. In the next reconcile, the <code>GET</code> will see the updated object, and your write will never happen. That&#x27;s how everything self-converges.</p><p>Here is another edge case. Picture this sequence inside a reconcile:</p><div class="code-block" data-language="go"><div class="min-w-0 max-w-full"><pre class="shiki shiki-themes planetscale-light planetscale-dark" style="--shiki-light:#2b2b2b;--shiki-dark:#e1e1e1;--shiki-light-bg:#ebebeb;--shiki-dark-bg:#1a1a1a" tabindex="0"><code><span class="line"><span style="--shiki-light:#818181;--shiki-dark:#A1A1A1">// I want N replicas. I see fewer, so I create the missing ones.</span></span>
<span class="line"><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">existing</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">,</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> _</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> :=</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> r</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#5E49AF;--shiki-dark:#B7A5FB">listChildPods</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">(</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">ctx</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">)</span><span style="--shiki-light:#818181;--shiki-dark:#A1A1A1"> // reads the CACHE</span></span>
<span class="line"><span style="--shiki-light:#F35815;--shiki-dark:#F35815">for</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> i</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> :=</span><span style="--shiki-light:#5E49AF;--shiki-dark:#B7A5FB"> len</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">(</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">existing</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">);</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> i</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> &#x3C;</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> desired</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">;</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> i</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815">++</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1"> {</span></span>
<span class="line"><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> r</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">client</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#5E49AF;--shiki-dark:#B7A5FB">Create</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">(</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">ctx</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">,</span><span style="--shiki-light:#5E49AF;--shiki-dark:#B7A5FB"> newPod</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">(</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">i</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">))</span><span style="--shiki-light:#818181;--shiki-dark:#A1A1A1"> // writes the API SERVER</span></span>
<span class="line"><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">}</span></span>
<span class="line"></span></code></pre></div></div><p>Now an event fires again a second later, before the cache has caught up with the Pods you just created. You list from the cache, and the new Pods aren&#x27;t there yet. Your code decides it still needs to create them, and you create duplicates. This is the classic stale-cache double-create, and it&#x27;s nasty because it only shows up under timing you can&#x27;t reproduce on your laptop.</p><p><button type="button" aria-haspopup="dialog" aria-expanded="false" aria-label="Enlarge image: Kubernetes cache reads versus API writes diagram" class="focus-visible-ring group relative block w-fit max-w-[min(100%,800px)] cursor-zoom-in text-left [&amp;_picture]:contents"><picture class="block"><source media="(prefers-color-scheme: light), (prefers-color-scheme: no-preference)" srcSet="https://planetscale-images.imgix.net/assets/part2-cache-vs-api-4EYpZbTb.png?auto=compress%2Cformat"/><source media="(prefers-color-scheme: dark)" srcSet="https://planetscale-images.imgix.net/assets/part2-cache-vs-api-darkmode-9mbMAIBr.png?auto=compress%2Cformat"/><img alt="Kubernetes cache reads versus API writes diagram" src="https://planetscale-images.imgix.net/assets/part2-cache-vs-api-4EYpZbTb.png?auto=compress%2Cformat" width="1504" height="1071" loading="lazy" class="w-auto max-w-full"/></picture><span aria-hidden="true" class="pointer-events-none absolute right-1 top-1 z-10 flex h-5 w-5 items-center justify-center border border-white/25 bg-black/70 text-white backdrop-blur-sm transition-colors transition-opacity group-hover:bg-black/90 group-hover:opacity-100 group-focus-visible:opacity-100 motion-reduce:transition-none [@media(hover:hover)_and_(pointer:fine)]:opacity-0"><svg width="14" height="14" viewBox="0 0 16 16" fill="none" stroke="currentColor" stroke-width="1.5" stroke-linecap="round" stroke-linejoin="round"><path d="M10 2h4v4M6 14H2v-4M14 2l-4.5 4.5M2 14l4.5-4.5"></path></svg></span></button><span hidden="" style="position:fixed;top:1px;left:1px;width:1px;height:0;padding:0;margin:-1px;overflow:hidden;clip:rect(0, 0, 0, 0);white-space:nowrap;border-width:0;display:none"></span></p><p>There are two ways out. The correct one is the <strong>expectations pattern</strong>, the same trick the built-in ReplicaSet controller uses: you record that you expect to see N creations in memory, and you don&#x27;t act again until the cache has caught up to your own writes. It works, but it&#x27;s not easy to implement and it&#x27;s a fair amount of machinery. Read more <a href="https://ahmet.im/blog/controller-pitfalls/">on Ahmet&#x27;s blog</a>.</p><p>The pragmatic one, which a lot of people use, is to bypass the cache for the reads where a stale view would cause a double-create or double-delete, and go straight to the API server:</p><div class="code-block" data-language="go"><div class="min-w-0 max-w-full"><pre class="shiki shiki-themes planetscale-light planetscale-dark" style="--shiki-light:#2b2b2b;--shiki-dark:#e1e1e1;--shiki-light-bg:#ebebeb;--shiki-dark-bg:#1a1a1a" tabindex="0"><code><span class="line"><span style="--shiki-light:#818181;--shiki-dark:#A1A1A1">// The cached client can be stale right after our own writes, which</span></span>
<span class="line"><span style="--shiki-light:#818181;--shiki-dark:#A1A1A1">// would make us miscount and create duplicates. For this one read,</span></span>
<span class="line"><span style="--shiki-light:#818181;--shiki-dark:#A1A1A1">// go direct to the API server instead of the cache. Slower,</span></span>
<span class="line"><span style="--shiki-light:#818181;--shiki-dark:#A1A1A1">// but consistent for this decision.</span></span>
<span class="line"><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">err</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> :=</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> r</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">apiReader</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#5E49AF;--shiki-dark:#B7A5FB">List</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">(</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">ctx</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">,</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> &#x26;</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">instances</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">,</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> client</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#5E49AF;--shiki-dark:#B7A5FB">InNamespace</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">(</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">ns</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">),</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> labelSelector</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">)</span></span>
<span class="line"></span></code></pre></div></div><p>This is not only a Kubernetes issue. In any system with a read cache and a write-through path, read-after-write is not consistent unless you make it so. Most of the time the cache is exactly what you want: cheap, local, and eventually consistent. Eventual consistency is fine because the loop runs again. But the moment a decision would be destructive or non-idempotent if you acted on a stale read, you need to know which path you&#x27;re on. Kubernetes solves many hard problems, but it also gives you a few new ones.</p><h3 id="setpoint-and-measured-variable-spec-and-status"><a href="#setpoint-and-measured-variable-spec-and-status">Setpoint and measured variable: spec and status</a></h3><p>Back to the control diagram. My script kept its setpoint in shell variables and its measured state in the output of <code>df</code> and <code>psql</code>. Kubernetes gives both a permanent home, on the object itself.</p><p><code>.spec</code> is the <strong>setpoint</strong>, the desired state. It&#x27;s owned by whoever created the object (a human, or another controller), and the reconciler treats it as read-only intent. It&#x27;s an anti-pattern to write to the <code>.spec</code> from inside the controller. If you do it, stop reading, go and fix your codebase. There are only a handful of exceptions, but a controller should generally never set its own setpoint.</p><p><code>.status</code> is the <strong>measured variable</strong>, the observed state. It&#x27;s owned by the controller, written through a separate status subresource, and it&#x27;s where you record what&#x27;s actually true. The better the status, the better the controller can decide. A good <code>.status</code> field is what makes a controller pleasant to operate. The word <em>observability</em> comes from control theory; <a href="https://en.wikipedia.org/wiki/Observability">Kalman coined it</a> around 1960 to ask whether you can infer a system&#x27;s internal state from its outputs. <code>.status</code> is also your response to any third-party system. If someone wants to learn the outcome of your actions, <code>.status</code> is the place to look at.</p><p>That split is the whole declarative model in two fields. It comes with a piece of bookkeeping that&#x27;s pure control theory: <code>.metadata.generation</code> increments when desired state changes, and by convention the controller writes back <code>.status.observedGeneration</code> to say &quot;the state I&#x27;m reporting reflects this version of your intent.&quot;</p><p>When <code>observedGeneration &lt; generation</code>, the status you&#x27;re looking at does not reflect the latest setpoint yet. That one comparison is how you tell &quot;converged&quot; from &quot;still working on it.&quot;</p><p><button type="button" aria-haspopup="dialog" aria-expanded="false" aria-label="Enlarge image: Spec and status as setpoint and measured variable" class="focus-visible-ring group relative block w-fit max-w-[min(100%,800px)] cursor-zoom-in text-left [&amp;_picture]:contents"><picture class="block"><source media="(prefers-color-scheme: light), (prefers-color-scheme: no-preference)" srcSet="https://planetscale-images.imgix.net/assets/part2-spec-status-CinHPict.png?auto=compress%2Cformat"/><source media="(prefers-color-scheme: dark)" srcSet="https://planetscale-images.imgix.net/assets/part2-spec-status-darkmode-DI9X1oFB.png?auto=compress%2Cformat"/><img alt="Spec and status as setpoint and measured variable" src="https://planetscale-images.imgix.net/assets/part2-spec-status-CinHPict.png?auto=compress%2Cformat" width="1504" height="1007" loading="lazy" class="w-auto max-w-full"/></picture><span aria-hidden="true" class="pointer-events-none absolute right-1 top-1 z-10 flex h-5 w-5 items-center justify-center border border-white/25 bg-black/70 text-white backdrop-blur-sm transition-colors transition-opacity group-hover:bg-black/90 group-hover:opacity-100 group-focus-visible:opacity-100 motion-reduce:transition-none [@media(hover:hover)_and_(pointer:fine)]:opacity-0"><svg width="14" height="14" viewBox="0 0 16 16" fill="none" stroke="currentColor" stroke-width="1.5" stroke-linecap="round" stroke-linejoin="round"><path d="M10 2h4v4M6 14H2v-4M14 2l-4.5 4.5M2 14l4.5-4.5"></path></svg></span></button><span hidden="" style="position:fixed;top:1px;left:1px;width:1px;height:0;padding:0;margin:-1px;overflow:hidden;clip:rect(0, 0, 0, 0);white-space:nowrap;border-width:0;display:none"></span></p><p>This is why the reconcile is <strong>stateless</strong>, and why that matters. Our shell script kept &quot;am I mid-failover?&quot; in a variable that died with the process. A Kubernetes controller keeps nothing important in memory. Every fact it needs is on an API object: the spec it&#x27;s driving toward, the status it last observed, the conditions describing where things stand. Kill the controller, restart it on another node, and it picks up exactly where it left off, not because it saved its progress, but because there was never any in-memory progress to lose. The state lives in the cluster (API server, <code>etcd</code> is what holds the state). The controller is just the loop that reads it.</p><h3 id="self-healing-by-design"><a href="#self-healing-by-design">Self-healing by design</a></h3><p>This is the part I like most.</p><p>When a controller creates a child object (a Pod, a PVC), it stamps an <strong>ownerReference</strong> on the child pointing back at the parent. That reference does two things. It sets up garbage collection: delete the parent, and Kubernetes can cascade the delete to its children. And it gives the controller a way to map child changes back to the parent: &quot;when any object I own changes, enqueue my parent for a reconcile.&quot; <code>ownerReference</code> allows you to link controllers to each other and create chains. If done right, all your controllers and systems fit together.</p><p>Here is an example. Follow the loop:</p><ol><li>A node dies and takes a Pod with it.</li><li>The Pod&#x27;s deletion is a watch event, an edge.</li><li>Through the ownership link, that edge becomes a reconcile request for the parent.</li><li>The parent reconciles, observes its children (level-triggered), sees one is missing and the count is below the setpoint, and creates a replacement.</li><li>The replacement is an unscheduled Pod, an edge for the scheduler.</li><li>The scheduler detects the unscheduled Pod, assigns a node.</li><li>The kubelet gets triggered because that&#x27;s an edge for that node&#x27;s kubelet and it starts the container.</li></ol><p>That&#x27;s multiple feedback loops, each watching the layer below, each reacting to an edge and converging to its own level, chained together through the API server with nobody orchestrating the whole thing.</p><p>Control theory has a name for loops stacked like this: <a href="https://en.wikipedia.org/wiki/Proportional%E2%80%93integral%E2%80%93derivative_controller#Cascade_control"><strong>cascade control</strong></a>. The output of an outer loop becomes the setpoint of an inner loop. A controller never writes its own <code>.spec</code>, but it writes <em>other</em> objects&#x27; <code>.spec</code> all the time. My operator writes the PVC&#x27;s spec, and that spec is the setpoint the CSI controllers converge to. Each loop worries only about its own layer and trusts the loop below.</p><p>So we wrote <code>if ! pg_isready; then docker start; fi</code> and maybe thought we&#x27;re good. Kubernetes turns that one line into several independent controllers that have never heard of each other, but still cooperate because they share the API server and watch each other&#x27;s objects. I like this part a lot. Nobody calls a central orchestrator. Nobody passes a private message. The system heals itself.</p><h3 id="what-observe-actually-means-in-a-real-operator"><a href="#what-observe-actually-means-in-a-real-operator">What &quot;observe&quot; actually means in a real operator</a></h3><p>Up to here I&#x27;ve been a little vague about the &quot;measure&quot; step, because in the examples the measured state is just &quot;list the child Pods.&quot; But in a real database operator it&#x27;s a lot more than that.</p><p>When an operator I work on reconciles a single Postgres instance, the first thing it does, before it decides anything, is build a snapshot of reality from every source that knows something true about that instance. Not just Kubernetes. Kubernetes barely knows anything about whether Postgres is actually healthy.</p><p>The sources gathered at the top of every reconcile:</p><ul><li><strong>The Kubernetes cache</strong>: the Pod, its PVC, the PV behind it, the Node it&#x27;s on, the ConfigMap holding its config. These are the cheap local reads, the stuff we already talked about.</li><li><strong>The database&#x27;s effective configuration.</strong> Not what we last wrote down, but what the server has actually loaded, so we can compare the two and detect drift. Other entities can rewrite or reload the config on disk without us knowing, so the only honest source of truth is the running server itself, never our last write.</li><li><strong>The database&#x27;s own view of its health.</strong> Its role, whether it&#x27;s healthy, how far behind its followers are, whether it&#x27;s currently accepting writes. Some of this comes from the agents that sit next to the database and manage it; some we get by opening a connection and asking the database directly. These calls carry a tight timeout and are allowed to fail, more on that below.</li><li><strong>A background collector.</strong> Some signals are too expensive or too rate-limited to fetch on every reconcile: disk usage, or whether a volume operation we kicked off earlier is still in flight and where it sits in its cooldown window. A separate collector, often a background goroutine, gathers these on a slow cadence and keeps the last value per volume in memory. The reconcile reads that value instantly, without blocking on anything. Think of these as custom workqueues you implement.</li></ul><p><button type="button" aria-haspopup="dialog" aria-expanded="false" aria-label="Enlarge image: Observation fan-in for a reconciler" class="focus-visible-ring group relative block w-fit max-w-[min(100%,800px)] cursor-zoom-in text-left [&amp;_picture]:contents"><picture class="block"><source media="(prefers-color-scheme: light), (prefers-color-scheme: no-preference)" srcSet="https://planetscale-images.imgix.net/assets/part2-observation-fan-in-DvyJkB0Y.png?auto=compress%2Cformat"/><source media="(prefers-color-scheme: dark)" srcSet="https://planetscale-images.imgix.net/assets/part2-observation-fan-in-darkmode-KuH1rwY1.png?auto=compress%2Cformat"/><img alt="Observation fan-in for a reconciler" src="https://planetscale-images.imgix.net/assets/part2-observation-fan-in-DvyJkB0Y.png?auto=compress%2Cformat" width="1504" height="1007" loading="lazy" class="w-auto max-w-full"/></picture><span aria-hidden="true" class="pointer-events-none absolute right-1 top-1 z-10 flex h-5 w-5 items-center justify-center border border-white/25 bg-black/70 text-white backdrop-blur-sm transition-colors transition-opacity group-hover:bg-black/90 group-hover:opacity-100 group-focus-visible:opacity-100 motion-reduce:transition-none [@media(hover:hover)_and_(pointer:fine)]:opacity-0"><svg width="14" height="14" viewBox="0 0 16 16" fill="none" stroke="currentColor" stroke-width="1.5" stroke-linecap="round" stroke-linejoin="round"><path d="M10 2h4v4M6 14H2v-4M14 2l-4.5 4.5M2 14l4.5-4.5"></path></svg></span></button><span hidden="" style="position:fixed;top:1px;left:1px;width:1px;height:0;padding:0;margin:-1px;overflow:hidden;clip:rect(0, 0, 0, 0);white-space:nowrap;border-width:0;display:none"></span></p><p>In code, the snapshot is just a struct, and the reconcile&#x27;s first move is to populate it. This is simplified, but faithful to the real shape:</p><div class="code-block" data-language="go"><div class="min-w-0 max-w-full"><pre class="shiki shiki-themes planetscale-light planetscale-dark" style="--shiki-light:#2b2b2b;--shiki-dark:#e1e1e1;--shiki-light-bg:#ebebeb;--shiki-dark-bg:#1a1a1a" tabindex="0"><code><span class="line"><span style="--shiki-light:#818181;--shiki-dark:#A1A1A1">// The observation snapshot: everything we know about this instance, right now.</span></span>
<span class="line"><span style="--shiki-light:#F35815;--shiki-dark:#F35815">type</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> reconcileHandler</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> struct</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1"> {</span></span>
<span class="line"><span style="--shiki-light:#818181;--shiki-dark:#A1A1A1"> // The object (.spec = setpoint, .status = measured).</span></span>
<span class="line"><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> instance</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> *v1</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815">PostgresInstance</span></span>
<span class="line"></span>
<span class="line"><span style="--shiki-light:#818181;--shiki-dark:#A1A1A1"> // Kubernetes objects.</span></span>
<span class="line"><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> pod</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> *corev1</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815">Pod</span></span>
<span class="line"><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> pvc</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> *corev1</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815">PersistentVolumeClaim</span></span>
<span class="line"><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> node</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> *corev1</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815">Node</span></span>
<span class="line"></span>
<span class="line"><span style="--shiki-light:#818181;--shiki-dark:#A1A1A1"> // Database state.</span></span>
<span class="line"><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> dbState</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> DatabaseState</span></span>
<span class="line"></span>
<span class="line"><span style="--shiki-light:#818181;--shiki-dark:#A1A1A1"> // What Postgres actually loaded, not what we last wrote.</span></span>
<span class="line"><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> effectiveConfig</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> map</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">[</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815">string</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">]</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815">string</span></span>
<span class="line"></span>
<span class="line"><span style="--shiki-light:#818181;--shiki-dark:#A1A1A1"> // Collected out-of-band.</span></span>
<span class="line"><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> diskUsage</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> *resource</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815">Quantity</span></span>
<span class="line"></span>
<span class="line"><span style="--shiki-light:#818181;--shiki-dark:#A1A1A1"> // The volume operation already in flight, if any.</span></span>
<span class="line"><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> storageOp</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> *StorageOperation</span></span>
<span class="line"><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">}</span></span>
<span class="line"></span>
<span class="line"><span style="--shiki-light:#F35815;--shiki-dark:#F35815">func</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1"> (</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">r </span><span style="--shiki-light:#F35815;--shiki-dark:#F35815">*Reconciler</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">)</span><span style="--shiki-light:#5E49AF;--shiki-dark:#B7A5FB"> newReconcileHandler</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">(</span></span>
<span class="line"><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> ctx</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> context</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815">Context</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">,</span></span>
<span class="line"><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> inst</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> *v1</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815">PostgresInstance</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">,</span></span>
<span class="line"><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">)</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1"> (</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815">*reconcileHandler</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">,</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> error</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">)</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1"> {</span></span>
<span class="line"><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> h</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> :=</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> &#x26;reconcileHandler</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">{</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">instance</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">:</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> inst</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">}</span></span>
<span class="line"></span>
<span class="line"><span style="--shiki-light:#818181;--shiki-dark:#A1A1A1"> // Cheap local reads.</span></span>
<span class="line"><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> h</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">pod</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">,</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> h</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">pvc</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">,</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> h</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">node</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> =</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> r</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#5E49AF;--shiki-dark:#B7A5FB">fetchKubeObjects</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">(</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">ctx</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">,</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> inst</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">)</span></span>
<span class="line"></span>
<span class="line"><span style="--shiki-light:#818181;--shiki-dark:#A1A1A1"> // Active database calls.</span></span>
<span class="line"><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> h</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">dbState</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> =</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> r</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#5E49AF;--shiki-dark:#B7A5FB">queryDatabase</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">(</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">ctx</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">,</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> h</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">pod</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">)</span></span>
<span class="line"><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> h</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">effectiveConfig</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> =</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> r</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#5E49AF;--shiki-dark:#B7A5FB">readEffectiveConfig</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">(</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">ctx</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">,</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> h</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">pod</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">)</span></span>
<span class="line"></span>
<span class="line"><span style="--shiki-light:#818181;--shiki-dark:#A1A1A1"> // Values from the collector/metric.</span></span>
<span class="line"><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> h</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">diskUsage</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> =</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> r</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">collector</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#5E49AF;--shiki-dark:#B7A5FB">Usage</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">(</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">h</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">pvc</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">)</span></span>
<span class="line"><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> h</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">storageOp</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> =</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> r</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">collector</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#5E49AF;--shiki-dark:#B7A5FB">InFlightOp</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">(</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">h</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">pvc</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">)</span></span>
<span class="line"></span>
<span class="line"><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> return</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> h</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">,</span><span style="--shiki-light:#7D5903;--shiki-dark:#FED54A"> nil</span></span>
<span class="line"><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">}</span></span>
<span class="line"></span></code></pre></div></div><p>A few things about this are deliberate, and only look obvious after you&#x27;ve been burned once or twice.</p><p><strong>Gather once, at the top.</strong> Every sub-decision in the reconcile reads from this one snapshot. We don&#x27;t re-query the database in the middle of the loop, or read the disk usage again three functions deep. If we did, different parts of the same reconcile could see different versions of reality. This sounds like a small detail, but it changes the whole design.</p><p>For example, the database might be the leader when we check at the top and a replica by the time another helper checks again. Then you get decisions that are individually reasonable, but wrong together. We have a rule in the codebase against stashing state back onto this handler mid-reconcile to pass between steps, because it reintroduces exactly the inconsistency we gathered the snapshot to avoid. Making the <code>reconcileHandler</code> immutable is one way to enforce that rule in the type system instead of relying on code review.</p><p><strong>Partial failures are tolerated.</strong> Reaching the database can fail while the Kubernetes reads succeed. That&#x27;s not always an error that aborts the reconcile. It&#x27;s a measured fact: &quot;Postgres is currently unreachable.&quot; That itself is something to record in status. A control loop that gives up entirely whenever one sensor is unavailable is a control loop that&#x27;s down a lot. We degrade instead. Think of a car. If the rain sensor for the wipers is broken, the whole car doesn&#x27;t stop. You can still drive, but you need to turn on a few things yourself.</p><p>Once the data snapshot exists, the reconcile is a sequence of small, idempotent steps, each comparing one slice of desired against observed and acting to close the gap:</p><div class="code-block" data-language="go"><div class="min-w-0 max-w-full"><pre class="shiki shiki-themes planetscale-light planetscale-dark" style="--shiki-light:#2b2b2b;--shiki-dark:#e1e1e1;--shiki-light-bg:#ebebeb;--shiki-dark-bg:#1a1a1a" tabindex="0"><code><span class="line"><span style="--shiki-light:#F35815;--shiki-dark:#F35815">func</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1"> (</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">r </span><span style="--shiki-light:#F35815;--shiki-dark:#F35815">*reconcileHandler</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">)</span><span style="--shiki-light:#5E49AF;--shiki-dark:#B7A5FB"> reconcile</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">(</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">ctx</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> context</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815">Context</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">)</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1"> (</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815">reconcile</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815">Result</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">,</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> error</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">)</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1"> {</span></span>
<span class="line"><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> var</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> rb</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> results</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815">Builder</span></span>
<span class="line"><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> rb</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#5E49AF;--shiki-dark:#B7A5FB">Merge</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">(</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">r</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#5E49AF;--shiki-dark:#B7A5FB">reconcileConfigMap</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">(</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">ctx</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">))</span><span style="--shiki-light:#818181;--shiki-dark:#A1A1A1"> // push desired config</span></span>
<span class="line"><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> rb</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#5E49AF;--shiki-dark:#B7A5FB">Merge</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">(</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">r</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#5E49AF;--shiki-dark:#B7A5FB">reconcileDatabase</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">(</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">ctx</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">))</span><span style="--shiki-light:#818181;--shiki-dark:#A1A1A1"> // reload/restart if params drifted</span></span>
<span class="line"><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> rb</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#5E49AF;--shiki-dark:#B7A5FB">Merge</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">(</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">r</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#5E49AF;--shiki-dark:#B7A5FB">reconcilePVC</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">(</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">ctx</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">))</span><span style="--shiki-light:#818181;--shiki-dark:#A1A1A1"> // grow the disk if needed</span></span>
<span class="line"><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> rb</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#5E49AF;--shiki-dark:#B7A5FB">Merge</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">(</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">r</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#5E49AF;--shiki-dark:#B7A5FB">reconcilePod</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">(</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">ctx</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">))</span><span style="--shiki-light:#818181;--shiki-dark:#A1A1A1"> // create/replace the Pod</span></span>
<span class="line"><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> rb</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#5E49AF;--shiki-dark:#B7A5FB">Merge</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">(</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">r</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#5E49AF;--shiki-dark:#B7A5FB">reconcileStatus</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">(</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">ctx</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">))</span><span style="--shiki-light:#818181;--shiki-dark:#A1A1A1"> // always last: write what we observed</span></span>
<span class="line"><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> return</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> rb</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#5E49AF;--shiki-dark:#B7A5FB">Result</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">()</span></span>
<span class="line"><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">}</span></span>
<span class="line"></span></code></pre></div></div><p>Status is written last on purpose, because it&#x27;s the measured variable: you record what&#x27;s true after you&#x27;ve taken your actions and observed the result. Again, in our operators, it&#x27;s not possible to write the status mid-reconcile.</p><p>Each step is independently idempotent. Each returns a result, either &quot;I&#x27;m done&quot; or &quot;requeue me in 30 seconds, I&#x27;m waiting on something,&quot; and the results merge. It reads almost exactly like the body of my shell loop. The difference is that &quot;observe the state&quot; grew from <code>df</code> and <code>pg_isready</code> into a fan-in across multiple systems, and &quot;take an action&quot; grew from <code>ssh</code> into typed, conflict-aware API writes.</p><p>This is the operator. The kubelet, the scheduler, CSI, and CNI are infrastructure we get by using Kubernetes. This loop, with its messy real-world observe step, is the part we actually write and deal with. Because we know how the underlying system works, we can design it without treating Kubernetes like a black box.</p><h3 id="not-every-edge-comes-from-the-api-server"><a href="#not-every-edge-comes-from-the-api-server">Not every edge comes from the API server</a></h3><p>There&#x27;s one more piece, and it lets me close a loop from Part 1 that I left deliberately: the disk-usage check.</p><p>My shell script polled <code>df</code> on every node every five seconds. For three nodes, fine. For thousands of databases, you can&#x27;t reconcile every one of them every few seconds just to check a number that rarely changes; you&#x27;d spend all your CPU re-deriving &quot;still at 40%, still at 40%, still at 40%.&quot; This is the level-triggered model&#x27;s one real cost: re-checking everything is correct, but it isn&#x27;t free.</p><p>The fix is to add a sensor that emits its own edges. A background collector polls our metrics pipeline for disk usage on a slow cadence, keeps the last value per volume in memory, and only emits an event when usage crosses a threshold, not while it sits above or below one:</p><div class="code-block" data-language="go"><div class="min-w-0 max-w-full"><pre class="shiki shiki-themes planetscale-light planetscale-dark" style="--shiki-light:#2b2b2b;--shiki-dark:#e1e1e1;--shiki-light-bg:#ebebeb;--shiki-dark-bg:#1a1a1a" tabindex="0"><code><span class="line"><span style="--shiki-light:#818181;--shiki-dark:#A1A1A1">// Edge detection. We fire only on the transition across the threshold,</span></span>
<span class="line"><span style="--shiki-light:#818181;--shiki-dark:#A1A1A1">// not every cycle we happen to be above it. Hovering at 81% is silent;</span></span>
<span class="line"><span style="--shiki-light:#818181;--shiki-dark:#A1A1A1">// crossing 80% upward is an event.</span></span>
<span class="line"><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">crossedUp</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> :=</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> previousUsage</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> &#x3C;</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> pvc</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">GrowThreshold</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> &#x26;&#x26;</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> usage</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815"> >=</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> pvc</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">GrowThreshold</span></span>
<span class="line"><span style="--shiki-light:#F35815;--shiki-dark:#F35815">if</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> crossedUp</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1"> {</span></span>
<span class="line"><span style="--shiki-light:#818181;--shiki-dark:#A1A1A1"> // -> generic event -> work queue -> reconcile</span></span>
<span class="line"><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> relay</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#5E49AF;--shiki-dark:#B7A5FB">Send</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">(</span><span style="--shiki-light:#F35815;--shiki-dark:#F35815">Event</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">{</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">Key</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">:</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600"> pvc</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">.</span><span style="--shiki-light:#A78103;--shiki-dark:#F2B600">Key</span><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">})</span></span>
<span class="line"><span style="--shiki-light:#616161;--shiki-dark:#C1C1C1">}</span></span>
<span class="line"></span></code></pre></div></div><p>That event goes into the same work queue as the API watch events and triggers a normal reconcile of the affected instance. Same rule as before: the event wakes us up, the reconcile decides from the current state.</p><p>For example, say we have a 10GiB disk and it&#x27;s using 8GiB. The collector saw it cross the threshold, so it wakes the reconciler. The reconciler reads the current usage, sees that it crossed the 80% threshold, and sets a new size on the PVC. After that, CSI handles the rest.</p><p>And because edges can be missed (the collector could be down, an event could be dropped from a full channel), there&#x27;s a <strong>resyncer</strong>: a periodic timer that enqueues every object for reconcile every minute or so, regardless of events. It&#x27;s the safety net. It&#x27;s our <code>sleep 5</code> loop. There&#x27;s also <code>RequeueAfter</code>, which a reconcile returns to say &quot;wake me again in 30 seconds,&quot; the controller&#x27;s way of polling a slow external operation without holding a worker.</p><p>There are two more questions: <strong>how often should the loop run, and who is allowed to run it?</strong> Control theory calls the first one the <em>sampling interval</em>. The rule of thumb: act faster than the thing you&#x27;re tracking changes, but not faster than it can respond. Reconciling a disk that fills over hours every few milliseconds just burns CPU to learn the same thing again.</p><p>So the operator puts boundaries around it.</p><ul><li>A <strong>coalescing delay</strong> handles noisy edge events: a burst of events for one object becomes one reconcile (think of it like a fan-in), not a thousand.</li><li>The <strong>resyncer</strong> is the safety net: every object gets looked at once in a while, even when nothing fires.</li><li>And <strong>leader election</strong> answers the <em>who</em>: only one copy of the operator runs the loop at a time. Two controllers writing to the same database object is not &quot;more reliable.&quot; Even with idempotent controllers, they&#x27;ll be requeueing due to conflicts and consuming unnecessary compute. In theory, a perfectly written controller should tolerate this. In practice, software is rarely perfect, and the safer boundary is worth it.</li></ul><p>To close out Part 2, let me redraw the control loop again. The diagram in Part 1 had a few basic boxes. Now, the same loop represents a closed feedback loop more realistically:</p><p><button type="button" aria-haspopup="dialog" aria-expanded="false" aria-label="Enlarge image: Production operator feedback loop diagram" class="focus-visible-ring group relative block w-fit max-w-[min(100%,800px)] cursor-zoom-in text-left [&amp;_picture]:contents"><picture class="block"><source media="(prefers-color-scheme: light), (prefers-color-scheme: no-preference)" srcSet="https://planetscale-images.imgix.net/assets/part2-operator-loop-YZG6qYVG.png?auto=compress%2Cformat"/><source media="(prefers-color-scheme: dark)" srcSet="https://planetscale-images.imgix.net/assets/part2-operator-loop-darkmode-jOHBdP3P.png?auto=compress%2Cformat"/><img alt="Production operator feedback loop diagram" src="https://planetscale-images.imgix.net/assets/part2-operator-loop-YZG6qYVG.png?auto=compress%2Cformat" width="1504" height="943" loading="lazy" class="w-auto max-w-full"/></picture><span aria-hidden="true" class="pointer-events-none absolute right-1 top-1 z-10 flex h-5 w-5 items-center justify-center border border-white/25 bg-black/70 text-white backdrop-blur-sm transition-colors transition-opacity group-hover:bg-black/90 group-hover:opacity-100 group-focus-visible:opacity-100 motion-reduce:transition-none [@media(hover:hover)_and_(pointer:fine)]:opacity-0"><svg width="14" height="14" viewBox="0 0 16 16" fill="none" stroke="currentColor" stroke-width="1.5" stroke-linecap="round" stroke-linejoin="round"><path d="M10 2h4v4M6 14H2v-4M14 2l-4.5 4.5M2 14l4.5-4.5"></path></svg></span></button><span hidden="" style="position:fixed;top:1px;left:1px;width:1px;height:0;padding:0;margin:-1px;overflow:hidden;clip:rect(0, 0, 0, 0);white-space:nowrap;border-width:0;display:none"></span></p><p>There is one new arrow in this diagram: <strong>disturbances</strong>. A controller has two jobs. The first is <a href="https://en.wikipedia.org/wiki/Setpoint_%28control_system%29">setpoint tracking</a>: someone edits the <code>.spec</code>, and the loop chases the new intent. The second is <a href="https://en.wikipedia.org/wiki/Control_theory">disturbance rejection</a>: the world changes on its own. A node dies, a customer starts a bulk import, someone deletes a Pod by hand. The level-triggered reconcile treats both the same way: it only sees the gap.</p><p>Our controller doesn&#x27;t always touch Postgres directly. Sometimes it writes a PVC and lets CSI do the storage work. Sometimes it creates a Pod and lets the scheduler and kubelet do their part. This is what a production operator looks like: one loop we write, surrounded by other loops we don&#x27;t write.</p><p>Notice that every decision in this loop has been binary: start the Pod or don&#x27;t, grow the disk or don&#x27;t, rewrite the config or don&#x27;t. That&#x27;s an <em>on/off controller</em>, and it covers most of what an operator does. But not every question is yes/no; once the answer becomes <em>how much</em> rather than <em>whether</em>, you need a controller with memory and a sense of trend: how long you&#x27;ve been off, and how fast it&#x27;s changing. That&#x27;s a separate post.</p><hr/><h2 id="conclusion"><a href="#conclusion">Conclusion</a></h2><p>All of this works, and most of the time it runs without anyone watching it. But the abstractions still leak, and they usually leak at a bad time.</p><p>Eventual consistency and the split between cache reads and API writes mean that a freshly-created object might not be visible to the thing that just created it. When something goes wrong, we&#x27;re debugging Kubernetes objects, database state, metrics, volume operations, and sometimes the cloud provider at the same time. The bug is usually not in one clean place.</p><p>The declarative model is wonderful until it meets an operation that&#x27;s inherently imperative and stateful, like a failover, a major-version upgrade, or a data migration. Then you have to turn a blocking, non-idempotent action into an idempotent one. That&#x27;s a whole other blog post.</p><p>That complexity is easy to underestimate. If you&#x27;re not dealing with sophisticated systems, if you can sacrifice availability, or if you don&#x27;t care about scalability, maybe all this machinery isn&#x27;t needed at all. Operators do not remove complexity. They move it into code someone has to understand.</p><p>I still think it&#x27;s worth it. For running thousands of databases that have to heal themselves without anyone watching, I don&#x27;t know a better alternative. The hard parts are hard because the problem is hard, not because Kubernetes made it hard.</p><p>Kubernetes is not only a container runtime. It&#x27;s not only a YAML processor, or an orchestrator, or whatever word we use that year. For me, the useful way to read Kubernetes is this: <strong>Kubernetes is a framework for feedback controllers</strong>, plus a consistent store to hold their setpoints and a shared event bus to wake them up.</p><p>Once you see that, the rest fits together. The kubelet, the scheduler, CSI, and your operator all read and write facts onto shared objects, and each one tries to move its own small part of the system toward the desired state. The core idea is still the same one we started with: write down what you want, look at what exists, make the next change, and repeat. Events wake the loop up, but the current state decides what happens.</p><p>Kubernetes didn&#x27;t invent these ideas; a thermostat had them long before us. The mapping to control theory is not perfect, and some boundaries are fuzzy. But the core idea holds. We are writing feedback loops in Go and applying them to databases. Mechanical and electrical engineers figured out how to build stable, long-running systems before us. Software engineering is still catching up, and Kubernetes gives us a practical way to use those ideas in production.</p></div></article></div></section></main><footer class="mb-6 mt-10 px-3 sm:px-5 container max-w-7xl"><nav class="grid grid-cols-1 text-left sm:grid-cols-2 lg:grid-cols-5 lg:mx-7"><div class="dashed-box dashed-box-x-t sm:dashed-box-l-t lg:dashed-box-y-l p-3"><h2 class="font-semibold">Company</h2><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/about" data-discover="true">About</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/brand" data-discover="true">Brand</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/blog" data-discover="true">Blog</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/changelog" data-discover="true">Changelog</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/careers" data-discover="true">Careers</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/events" data-discover="true">Events</a></div><div class="dashed-box dashed-box-x-t lg:dashed-box-y-l p-3"><h2 class="font-semibold">Product</h2><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/case-studies" data-discover="true">Case studies</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/enterprise" data-discover="true">Enterprise</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/pricing" data-discover="true">Pricing</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/benchmarks" data-discover="true">Benchmarks</a></div><div class="dashed-box dashed-box-x-t sm:dashed-box-l-t lg:dashed-box-y-l p-3"><h2 class="font-semibold">Resources</h2><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/docs">Documentation</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/migrate" data-discover="true">Migrate</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="https://support.planetscale.com/hc/en-us" rel="nofollow noopener noreferrer" target="_blank">Support</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="https://planetscalestatus.com" rel="nofollow noopener noreferrer" target="_blank">Status</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="https://trust.planetscale.com" rel="nofollow noopener noreferrer" target="_blank">Trust Center</a></div><div class="dashed-box dashed-box-x-t lg:dashed-box-y-l p-3"><h2 class="font-semibold">Courses</h2><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/learn/courses/mysql-for-developers" data-discover="true">MySQL for Developers</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/learn/courses/database-scaling" data-discover="true">Database Scaling</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/learn/courses/vitess" data-discover="true">Learn Vitess</a></div><div class="dashed-box p-3 sm:col-span-2 lg:col-span-1"><h2 class="font-semibold text-primary hover:text-contrast">Open source</h2><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="/vitess" data-discover="true">Vitess</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="https://vitess.io/slack" rel="nofollow noopener noreferrer" target="_blank">Vitess community</a><a class="block pl-1ch -indent-1ch text-primary hover:text-contrast" href="https://github.com/planetscale" rel="me nofollow noopener noreferrer" target="_blank">GitHub</a></div></nav><div class="dashed-box dashed-box-x-b p-3 lg:mx-7"><p class="mb-3 md:mb-0"><a class="text-primary" rel="nofollow" href="/legal/privacy" data-discover="true">Privacy</a><span class="text-decoration" role="presentation"> <!-- -->|<!-- --> </span><a class="text-primary" rel="nofollow" href="/legal/siteterms" data-discover="true">Terms</a><span class="text-decoration" role="presentation"> <!-- -->|<!-- --> </span><a class="text-primary" rel="nofollow" href="/legal/cookies" data-discover="true">Cookies</a><span class="text-decoration" role="presentation"> <!-- -->|<!-- --> </span><a class="text-primary" rel="nofollow" href="/legal/patents" data-discover="true">Patents</a><span class="text-decoration" role="presentation"> <!-- -->|<!-- --> </span><a class="text-primary" rel="nofollow" href="/legal/privacy#privacy-rights-and-choices" data-discover="true">Do Not Share My Personal Information</a></p><p class="text-secondary">© <!-- -->2026<!-- --> PlanetScale, Inc. All rights reserved.</p></div><p class="mb-0 mt-3 break-normal lg:mx-7"><a class="text-primary" href="https://github.com/planetscale" rel="me nofollow noopener noreferrer" target="_blank">GitHub</a><span class="text-decoration" role="presentation"> <!-- -->|<!-- --> </span><a aria-label="X (formerly Twitter)" class="text-primary" href="https://twitter.com/planetscale" rel="me nofollow noopener noreferrer" target="_blank">X</a><span class="text-decoration" role="presentation"> <!-- -->|<!-- --> </span><a aria-label="LinkedIn" class="text-primary" href="https://www.linkedin.com/company/planetscale" target="_blank" rel="noreferrer">LinkedIn</a><span class="text-decoration" role="presentation"> <!-- -->|<!-- --> </span><a class="text-primary" href="https://www.youtube.com/planetscale" rel="me nofollow noopener noreferrer" target="_blank">YouTube</a><span class="text-decoration" role="presentation"> <!-- -->|<!-- --> </span><a aria-label="Discord" class="text-primary" href="https://pscale.link/community" rel="nofollow noopener noreferrer" target="_blank">Discord</a><span class="text-decoration" role="presentation"> <!-- -->|<!-- --> </span><a class="text-primary" href="https://www.facebook.com/planetscaledata" rel="me nofollow noopener noreferrer" target="_blank">Facebook</a></p></footer><script nonce="h1m8nAKwwY4XGImd6cj2zjoYAypTpKkpoCcuiMZKkXs=">((storageKey2, restoreKey) => {
if (!window.history.state || !window.history.state.key) {
let key2 = Math.random().toString(32).slice(2);
window.history.replaceState({ key: key2 }, "");
}
try {
let storedY = JSON.parse(sessionStorage.getItem(storageKey2) || "{}")[restoreKey || window.history.state.key];
if (typeof storedY === "number") window.scrollTo(0, storedY);
} catch (error2) {
console.error(error2);
sessionStorage.removeItem(storageKey2);
}
})("react-router-scroll-positions", null)</script><script nonce="h1m8nAKwwY4XGImd6cj2zjoYAypTpKkpoCcuiMZKkXs=">window.__reactRouterContext = {"basename":"/","future":{"unstable_enableNodeReadableStream":false,"unstable_optimizeDeps":true},"routeDiscovery":{"mode":"lazy","manifestPath":"/__manifest"},"ssr":true,"isSpaMode":false};window.__reactRouterContext.stream = new ReadableStream({start(controller){window.__reactRouterContext.streamController = controller;}}).pipeThrough(new TextEncoderStream());</script><script nonce="h1m8nAKwwY4XGImd6cj2zjoYAypTpKkpoCcuiMZKkXs=" type="module" async="">;
import * as route0 from "/assets/root-DbOv4-98.js";
import * as route1 from "/assets/blog-pXH7ptHJ.js";
import * as route2 from "/assets/blog._slug-Ch_92qsH.js";
window.__reactRouterManifest = {
"entry": {
"module": "/assets/entry.client-3vubyXrk.js",
"imports": [
"/assets/jsx-runtime-DwfQwkRq.js",
"/assets/components-_bNmAApg.js",
"/assets/_.well-known_.mcp.server-card_.json_-Sx7XeH3e.js",
"/assets/index-mKTXLmHu.js",
"/assets/errorBoundaries-DhW4jVYt.js"
],
"css": []
},
"routes": {
"root": {
"id": "root",
"path": "",
"hasAction": false,
"hasLoader": true,
"hasClientAction": false,
"hasClientLoader": false,
"hasClientMiddleware": false,
"hasDefaultExport": true,
"hasErrorBoundary": true,
"module": "/assets/root-DbOv4-98.js",
"imports": [
"/assets/jsx-runtime-DwfQwkRq.js",
"/assets/components-_bNmAApg.js",
"/assets/_.well-known_.mcp.server-card_.json_-Sx7XeH3e.js",
"/assets/index-mKTXLmHu.js",
"/assets/errorBoundaries-DhW4jVYt.js",
"/assets/lib-Dg89tQ22.js",
"/assets/analytics.client-DM6E8o1h.js",
"/assets/SiteHeader-C2U5gvDH.js",
"/assets/current-9yDxj94E.js",
"/assets/clsx-eT0YPcGk.js",
"/assets/bugs-38ilEoW0.js",
"/assets/keyboard-D-uXZORL.js",
"/assets/use-tab-direction-dKm-S3Ck.js"
],
"css": []
},
"routes/blog": {
"id": "routes/blog",
"parentId": "root",
"path": "blog",
"hasAction": false,
"hasLoader": false,
"hasClientAction": false,
"hasClientLoader": false,
"hasClientMiddleware": false,
"hasDefaultExport": false,
"hasErrorBoundary": false,
"module": "/assets/blog-pXH7ptHJ.js",
"imports": [
"/assets/_.well-known_.mcp.server-card_.json_-Sx7XeH3e.js"
],
"css": []
},
"routes/blog.$slug": {
"id": "routes/blog.$slug",
"parentId": "routes/blog",
"path": ":slug",
"hasAction": false,
"hasLoader": true,
"hasClientAction": false,
"hasClientLoader": false,
"hasClientMiddleware": false,
"hasDefaultExport": true,
"hasErrorBoundary": false,
"module": "/assets/blog._slug-Ch_92qsH.js",
"imports": [
"/assets/components-_bNmAApg.js",
"/assets/lib-Dg89tQ22.js",
"/assets/jsx-runtime-DwfQwkRq.js",
"/assets/ContentImage-Dh6VEOUl.js",
"/assets/clsx-eT0YPcGk.js",
"/assets/BlogCategoryLink-DmQyn0gp.js",
"/assets/_.well-known_.mcp.server-card_.json_-Sx7XeH3e.js",
"/assets/Details-BSB_b6hI.js",
"/assets/Skittle-CDFOPRjH.js",
"/assets/SiteFooter-B2Gq9u2j.js",
"/assets/SiteHeader-C2U5gvDH.js",
"/assets/Vimeo-00PQJDli.js",
"/assets/YouTube-CMfaljVr.js",
"/assets/date-CJTFH3uT.js",
"/assets/errorBoundaries-DhW4jVYt.js",
"/assets/keyboard-D-uXZORL.js",
"/assets/use-tab-direction-dKm-S3Ck.js",
"/assets/index-mKTXLmHu.js",
"/assets/use-inert-others-BMJ6-xOX.js",
"/assets/description-Cf6FZmDe.js",
"/assets/use-is-mounted-uQsUZyP9.js",
"/assets/types-DvonrUFF.js",
"/assets/current-9yDxj94E.js",
"/assets/analytics.client-DM6E8o1h.js",
"/assets/bugs-38ilEoW0.js"
],
"css": []
},
"routes/_index": {
"id": "routes/_index",
"parentId": "root",
"index": true,
"hasAction": false,
"hasLoader": true,
"hasClientAction": false,
"hasClientLoader": false,
"hasClientMiddleware": false,
"hasDefaultExport": true,
"hasErrorBoundary": false,
"module": "/assets/_index-BfA6EnlR.js",
"imports": [
"/assets/components-_bNmAApg.js",
"/assets/lib-Dg89tQ22.js",
"/assets/jsx-runtime-DwfQwkRq.js",
"/assets/Logo-Gm9TLYAs.js",
"/assets/SiteFooter-B2Gq9u2j.js",
"/assets/SiteHeader-C2U5gvDH.js",
"/assets/_.well-known_.mcp.server-card_.json_-Sx7XeH3e.js",
"/assets/bugs-38ilEoW0.js",
"/assets/keyboard-D-uXZORL.js",
"/assets/use-is-mounted-uQsUZyP9.js",
"/assets/use-tab-direction-dKm-S3Ck.js",
"/assets/errorBoundaries-DhW4jVYt.js",
"/assets/clsx-eT0YPcGk.js",
"/assets/current-9yDxj94E.js",
"/assets/analytics.client-DM6E8o1h.js",
"/assets/index-mKTXLmHu.js"
],
"css": []
},
"routes/blog._index": {
"id": "routes/blog._index",
"parentId": "routes/blog",
"index": true,
"hasAction": false,
"hasLoader": true,
"hasClientAction": false,
"hasClientLoader": false,
"hasClientMiddleware": false,
"hasDefaultExport": true,
"hasErrorBoundary": false,
"module": "/assets/blog._index-DcvTTuDd.js",
"imports": [
"/assets/components-_bNmAApg.js",
"/assets/jsx-runtime-DwfQwkRq.js",
"/assets/social-Cd2AtOZM.js",
"/assets/BlogCategoryLink-DmQyn0gp.js",
"/assets/BlogPostLink-DC1SPKBJ.js",
"/assets/BlogCategoryNav-CB7TJ3IB.js",
"/assets/Paginator-xlPA_JNt.js",
"/assets/SiteFooter-B2Gq9u2j.js",
"/assets/SiteHeader-C2U5gvDH.js",
"/assets/date-CJTFH3uT.js",
"/assets/_.well-known_.mcp.server-card_.json_-Sx7XeH3e.js",
"/assets/lib-Dg89tQ22.js",
"/assets/errorBoundaries-DhW4jVYt.js",
"/assets/clsx-eT0YPcGk.js",
"/assets/types-DvonrUFF.js",
"/assets/enumerator-2YLGh-nT.js",
"/assets/current-9yDxj94E.js",
"/assets/analytics.client-DM6E8o1h.js",
"/assets/bugs-38ilEoW0.js",
"/assets/keyboard-D-uXZORL.js",
"/assets/use-tab-direction-dKm-S3Ck.js",
"/assets/index-mKTXLmHu.js"
],
"css": []
}
},
"url": "/assets/manifest-e17deb94.js",
"version": "e17deb94"
};
window.__reactRouterRouteModules = {"root":route0,"routes/blog":route1,"routes/blog.$slug":route2};
import("/assets/entry.client-3vubyXrk.js");</script><script type="application/ld+json" nonce="h1m8nAKwwY4XGImd6cj2zjoYAypTpKkpoCcuiMZKkXs=">{"@context":"https://schema.org","@type":"Organization","name":"PlanetScale, Inc.","url":"https://planetscale.com","sameAs":["https://twitter.com/PlanetScale","https://www.facebook.com/planetscaledata/","https://www.instagram.com/planetscale/"],"address":{"@type":"PostalAddress","streetAddress":"WeWork c/o PlanetScale, 535 Mission Street, 14th Floor","addressLocality":"San Francisco","addressRegion":"CA","postalCode":"94105","addressCountry":"US"}}</script><!--$--><script nonce="h1m8nAKwwY4XGImd6cj2zjoYAypTpKkpoCcuiMZKkXs=">window.__reactRouterContext.streamController.enqueue("[{\"_1\":2,\"_3\":-5,\"_4\":-5},\"loaderData\",{\"_5\":6,\"_7\":8},\"actionData\",\"errors\",\"root\",{\"_2119\":2120},\"routes/blog.$slug\",{\"_9\":10,\"_5\":11},\"blog\",{\"_12\":13,\"_14\":15,\"_16\":-7,\"_17\":18,\"_19\":20,\"_21\":22,\"_23\":24,\"_25\":26,\"_27\":28,\"_29\":30,\"_31\":32},\"https://planetscale.com\",\"body\",[125,126,127,128,129,130,131,132,133,134,135,136,137,138,139,140,141,142,143,144,145,146,147,148,149,150,151,152,153,154,155,156,157,158,159,160,161,162,163,164,165,166,167,168,169,170,171,172,173,174,175,176,177,178,179,180,181,182,183,184,185,186,187,188,189,190,191,192,193,194,195,196,197,198,199,200,201,202,203,204,205,206,207,208,209,210,211,212,213,214,215,216,217,218,219,220,221,222,223,224,225,226,227,228,229,230,231,232,233,234,235,236,237,238,239,240,241,242,243,244,245,246,247,248,249,250,251,252,253,254,255,256,257,258,259,260,261,262,263,264,265,266,267,268,269,270,271,272,273,274,275,276,277,278,279,280,281,282,283,284,285],\"body_text\",\"For the last decade, Kubernetes has been the backdrop to most of my work: operating clusters, helping build hosted Kubernetes, and writing Kubernetes operators. At PlanetScale, that now means running stateful systems like Postgres and MySQL in production. Kubernetes has many faces, but here I want to talk about one face only: why it is so good at running workloads at scale.\\nPeople ask me what an operator actually does. The canonical answer is: \\\"it reconciles desired state.\\\" This is correct, but it also tells you almost nothing.\\nAn operator is a feedback controller. It's the same closed loop that runs a thermostat or keeps your car at a fixed speed on cruise control. In our case, the thing being controlled is a database. I have been building these loops for years, and the best way I know to make them click is to ignore Kubernetes at the beginning. Kubernetes is full of control theory, even if we don't call it that in the day-to-day.\\nBefore we look at a single line of Kubernetes, we're going to run a production database by hand and slowly let the feedback loop appear on its own. Then we'll map that loop to Kubernetes, with the pieces production needs: a store, watches, queues, retries, and more. At the end, we'll look at what one of these loops looks like in a real operator.\\nA working understanding of containers and kubectl helps, but you don't need to be a Kubernetes expert. I'll use terms like idempotent, fan-in, and eventual consistency, and introduce the parts that matter as we go.\\nWe're going to start slow and gradually ramp things up. Each part builds on the previous.\\nPart 1: running Postgres by hand\\nOne container, one machine\\nLet's start from scratch. I want to run Postgres on a Linux box, and I need it inside a container. To start it, we run:docker run -d --name pg \\\\\\n -e POSTGRES_PASSWORD=secret \\\\\\n postgres:18\\n\\nThat's it. Postgres is running. My app connects to it, writes some rows, and everything works fine. But then the machine goes away: the cloud provider reclaims the instance (hardware fails, or a spot instance gets taken back), or I ship a new version of my setup, which means stopping the old container and starting a fresh one in its place. Either way, the container is replaced, and my data is gone. The container storage was ephemeral, and I did not attach any persistent volume to it.\\nThere is already a gap between what I want (Postgres, running, with my data) and what I have (a container whose storage disappears when the container or node goes away). The rest of this post is about that gap and the machinery we build to close it.\\nPick a node, by hand\\nImagine we have hundreds of nodes (servers) we can use. I already have other workloads running on them. I need to decide which one runs this database. So I ssh into the box that looks the least busy and start the container there.ssh node-07 'docker run -d --name pg ... postgres:18'\\n\\nI picked node-07 because it looked idle enough. I start keeping track of it, save it in some sort of config file, and push it to some repo.\\nIt needs a real disk\\nContainer storage is ephemeral, so I have to attach a real block device. In the cloud this is an EBS volume (e.g. on AWS); on bare metal it's a physical disk. Assuming it's a block device, this is what we usually do: provision the volume, attach it to the node, format it, mount it, and point Postgres' data directory at the mount.# provision + attach first with cloud CLI, then on the node:\\nmkfs.ext4 /dev/nvme1n1\\nmkdir -p /var/lib/pg-data\\nmount /dev/nvme1n1 /var/lib/pg-data\\ndocker run -d --name pg \\\\\\n -e POSTGRES_PASSWORD=secret \\\\\\n -v /var/lib/pg-data:/var/lib/postgresql \\\\\\n postgres:18\\n\\nThese are a lot of steps, and each one can fail halfway. And if the disk fills up later, Postgres stops accepting writes and we have to resize the volume by hand: first through the cloud provider, then again inside the filesystem.\\nOne isn't enough\\nA single Postgres instance is a single point of failure. We want high availability: one primary and two replicas. These need to be on three different machines, with streaming replication between them. So we do the same steps again, three times, on node-07, node-12, and node-19. I also wire up replication by hand: primary_conninfo, replication slots, all of it.\\nNow we have three nodes with three Postgres instances. One of the instances is the primary (here it's node-07). But this raises new problems, like what to do if the primary's node dies?\\nThey have to find each other\\nHere is another thing we have to solve. The replicas need to reach the primary, and the primary needs to accept their connections. And every one of these addresses is an IP that changes when a container restarts.\\nThe first thing I do is hard-code the IPs. I write node-07's address into the replicas' config, I list the replicas' addresses in the primary's pg_hba.conf, and I keep a small /etc/hosts table and save it somewhere.# on each replica's postgresql.auto.conf, until the primary is recreated with a new IP\\nprimary_conninfo = 'host=10.4.7.21 port=5432 user=replicator ...'\\n\\nBut we still have a problem: the first time the primary is recreated with a different IP, the whole cluster falls apart.\\n\\nThe watchdog script\\nNow, this is where we start thinking about how to solve these issues. Everything described so far can break, and will continue to break even if I fix it:\\nA replica process dies and doesn't come back.\\nA disk gets full.\\nThe primary fails and a replica has to be promoted.\\nA config I changed on two nodes but forgot on the third one. They are now out of sync.\\nLet's assume we've set up a simple uptime monitor and we're going to get paged for all these cases. To avoid getting paged at night, we do the sensible thing: write a script. So we decide to write a loop that wakes up every few seconds, looks at each node, and fixes whatever's wrong.while true; do\\n for node in node-07 node-12 node-19; do\\n if ! ssh \\\"$node\\\" 'pg_isready -q'; then\\n ssh \\\"$node\\\" 'docker start pg' # it died, bring it back\\n fi\\n\\n usage=$(ssh \\\"$node\\\" \\\"df --output=pcent /var/lib/pg-data | tail -1 | tr -dc 0-9\\\")\\n if [ \\\"$usage\\\" -gt 80 ]; then\\n grow_volume \\\"$node\\\" # disk filling, make it bigger\\n fi\\n done\\n sleep 5\\ndone\\n\\nIt's written in Bash, and probably has tons of bugs. You notice something here? The loop doesn't care how the database got into a bad state. Every five seconds it looks at the current state of the world and asks this question: does reality match what I want?\\nIf a process is down, start it. If a disk is filling, grow it. Run the loop once or run it a thousand times and the result is the same, because each action is conditional on the current state. The script is idempotent.\\nChanging a parameter\\nLet's make things a little more complex. I need to raise max_connections from 100 to 500. This one is not a reload-only change. PostgreSQL says it can only be set at server start, so the manual version is to ssh into each box, edit postgresql.conf, restart Postgres, and check that it took on all three.\\nBecause I know that ssh'ing into the nodes manually isn't a thing I want anymore, I do the same thing we did previously: I write the desired value down in one place and teach the loop to enforce it.WANT_MAX_CONNECTIONS=500\\n\\nfor node in node-07 node-12 node-19; do\\n have=$(ssh \\\"$node\\\" \\\"psql -tAc 'show max_connections'\\\")\\n if [ \\\"$have\\\" != \\\"$WANT_MAX_CONNECTIONS\\\" ]; then\\n ssh \\\"$node\\\" \\\"sed -i 's/^max_connections.*/max_connections = $WANT_MAX_CONNECTIONS/' /var/lib/pg-data/postgresql.conf\\\"\\n ssh \\\"$node\\\" \\\"docker restart pg\\\"\\n fi\\ndone\\n\\nThis is the same idea as before. I read what I want (a variable). Observe what I have (a query). If they differ, I take an action to close the difference. Again, I don't track whether I changed it last time. All I do is compare and converge, every loop.\\nWhat we actually built\\nI started with a desired state that was written down in one place: three instances, this disk size, max_connections = 500. Every few seconds I observe the actual state of the system. I compute the difference. I take whatever action closes that difference. Then I do it again, forever.\\nThat's a closed feedback loop. The word \\\"closed\\\" matters. It means the output of the system is fed back into the next decision. I don't run docker start and assume the database is fine. I check the database again. If it is still wrong, I act again. If it is already correct, I do nothing.\\nThe nice part is that the same loop works for different problems. It can restart a dead process, grow a disk, or push max_connections = 500. The action changes, but the shape stays the same: read what I want, observe what I have, compare them, act, repeat. If I draw the same thing as a block diagram, with the control theory names added, it would look like this:\\n\\nHere is how the vocabulary from control theory maps cleanly onto my shell script:\\nThe setpoint is my desired state, the variables at the top of the script (disk size, max_connections and so on).\\nThe measured output is what I observe: pg_isready, df, show max_connections.\\nThe error (e) is the difference between them.\\nThe controller is the body of the loop, the if statements that decide what to do. It is not the whole script.\\nThe actuator is what carries out the action: ssh plus docker start.\\nThe plant is the system being controlled, Postgres and its disk.\\nThat also gives us a nice way to understand open-loop control. My very first attempt, ssh in, run the command, and walk away, was open-loop: fire an action and assume it worked. The Bash script is closed-loop because it keeps feeding the measured state back into the next decision.\\nA Bash loop is not a production control plane. Just to name a few issues with it:\\nIt has no concurrency control, so two copies of the script can race each other. Imagine both deciding to promote a different replica.\\nIt keeps its only real state, \\\"am I mid-failover?\\\", in a shell variable that could die with the process.\\nIt polls every node every five seconds whether anything changed or not, which is fine for three nodes, but too expensive for three thousand nodes.\\nIt has no idea what to do when the ssh itself times out.\\nAnd the moment I want a second kind of resource, a connection pooler, a backup job, a read replica in another region, I'm copy-pasting this whole structure.\\nWhat if the script also fails? Who runs it then? We could keep hardening this script, but look at where it goes: we would need a real store for the desired state, watches instead of polling, a work queue, retries, leader election. We would be rebuilding Kubernetes. The real platform already exists, and it's Kubernetes.\\nPart 2: how we reinvented Kubernetes\\nNow we can map what we hand-rolled in Part 1 to Kubernetes. Almost all of it already exists there. The operator is the part we care about.\\nThe other loops\\nLet's go through some of the pieces we built by hand before the watchdog loop. You already know these components by name. What you might not have noticed is that they also work like controllers.\\nSpinning up the container: the kubelet. First, a quick definition: a Pod is the smallest thing Kubernetes runs, one or more containers scheduled together on a node and sharing its network. For us it's the Postgres container. On every node runs an agent called the kubelet. Its desired state is the set of Pods assigned to its node, which it learns from the API server. Its observed state is the set of containers actually running, which it gets from the container runtime. When they differ, it starts the missing container, kills the extra one, or restarts the crashed one. My if ! pg_isready; then docker start; fi is the kubelet's job, just done properly. The kubelet doesn't shell into anything; it talks to containerd over a gRPC socket, which talks to runc.\\nPicking a node: the scheduler. Remember me choosing node-07? That's the scheduler's whole reason to exist. It watches for Pods with no node assigned, filters out the nodes that can't work, scores the rest, and writes the decision to one field: pod.Spec.NodeName. The scheduler doesn't start the container; it records the placement and lets the kubelet pick it up. You will realize that most things in Kubernetes are decoupled like this.\\nAttaching the disk: CSI and the PV/PVC sync. My multi-step mkfs and mount script becomes a PersistentVolumeClaim, which is a declarative request for storage. The Container Storage Interface (CSI) driver turns that request into a real volume. CSI itself is a set of controllers and sidecars: one provisions, one attaches, one resizes, and so on, while the kubelet calls the driver's node plugin to do the actual mount. It's a family of controllers. If there is a PVC but no disk behind it, one controller creates the disk. If the PVC size increases, another controller calls the provider API (e.g., AWS ModifyVolume). Again, I write intent, and a controller does the actual work. (note: I wrote one of the early production CSI drivers, csi-digitalocean, and a long post about building one.)\\nMaking them find each other: the CNI and Services. The /etc/hosts problem is solved at a layer we no longer have to think about. A CNI plugin gives Pods their network identity; Cilium, for example, does this with eBPF instead of a pile of iptables rules. For stateful workloads, a StatefulSet plus a headless Service gives each replica its own stable DNS name, which is exactly what a Postgres replica needs. The hard-coded IP that broke our cluster becomes a name that keeps working. DNS is only one tool here; other service discovery systems like etcd, ZooKeeper, and Consul solve similar problems.\\nAs you see, all the problems we solved with ssh and various scripts are replaced by Kubernetes components and drivers. And these are just a few of them:\\n\\nAll of this so far is useful context, but the part we care about is our watchdog loop, because that's the one we get to write ourselves.\\nThat's the operator.\\nThe for-loop translated to Kubernetes\\nIn Kubernetes, our watchdog script is a controller, and the standard way to write one in Go is a library called controller-runtime. At its heart, it's a function with a basic signature:func (r *Reconciler) Reconcile(ctx context.Context, req reconcile.Request) (reconcile.Result, error) {\\n // req contains a namespace/name. That's it. That's the whole input.\\n}\\n\\nNotice what is missing here: the function isn't told what changed. There is no diff. It isn't handed the old object and the new object. It isn't given an event type. It gets a key, a namespace and a name, and nothing else. It's minimal by design, because it has to work for many different controllers. The function's job is to fetch the object with that namespace/name, look at the world, and converge to the desired state.\\nEdge-triggered notifications, level-triggered logic\\nThere are two ways to build any closed feedback loop:\\nEdge-triggered: act on transitions, on events. \\\"The disk crossed 80%.\\\" \\\"The Pod was deleted.\\\" \\\"The number of replicas increased by 2.\\\"\\nLevel-triggered: act on the current state, regardless of how you got there. \\\"The disk is at 85%.\\\" \\\"The Pod is missing.\\\" \\\"The number of replicas is 3.\\\"\\nMy first mental model of controllers, and probably yours at some point, was edge-triggered: listen to a stream of changes, and for each change, try to converge.\\nThe problem is that this is very fragile. In distributed systems, if one component is fragile, the fragility spreads to the rest of the system. Why is edge-triggering fragile? Say your controller is down for thirty seconds. It misses the events from those thirty seconds, and its view of the world is now permanently wrong. If two events arrive out of order, you process them out of order. If an event is delivered twice, you act twice. You're rebuilding your state from a stream of events, and you've inherited all of event sourcing's hard problems.\\nHere is a very concrete example. Assume you have 1 replica, and you increase it to 3 replicas. Because you have only subscribed to changes, either:\\nYou miss the event (maybe the queue dropped it, or the consumer, your app, dropped it due to a crash or a full buffer).\\nYou receive it twice.\\nIn the first case, you won't be able to self-correct. In the second case, if your handler blindly applies the delta again, you'll end up with 5 replicas (you overshoot), instead of 3.\\nThe level-triggered model fixes all of that. Remember, our shell script never asked \\\"what changed?\\\" It asked \\\"what is true right now?\\\", every five seconds, from scratch. Miss a loop, and the next one catches up. Run the loop twice, and you get the same result. The current state of the world is the only input that matters, and it's always available to read. So in the level-triggered case, our example above becomes this: you read replicas=3, you check the current number of replicas, which is 1, and you increase by 2.\\nIf you miss the event, no one cares. In the next reconcile loop you'll catch it. If your app crashes, it comes back, reads again and detects that it did not increase it yet, increases it.\\nKubernetes controllers combine both: edge-triggered notifications, level-triggered logic.\\nEvents (the edges) are only a hint that it's worth looking again. They tell you when to reconcile, never what to do. The reconcile itself is level-based: it reads the current state (e.g., replicas=3) and drives toward the desired state (e.g., create 2 replicas), ignoring the triggering event completely. That's why Reconcile only gets a key. The framework makes it hard to write edge-triggered logic, on purpose. Edge-triggered logic is how you get a controller that's fragile and permanently wrong after its first hiccup.\\n\\nOur bash script stumbled into this property by accident, at least for the sake of the example. But the controller-runtime framework gives it to you on purpose. It's why a Kubernetes controller can crash, get restarted ten minutes later, and converge correctly with no special recovery code. There is no recovery code. There is just the loop. The controller can reconstruct the world from scratch.\\nInformers, the work queue, and a cache\\nSo where do the edges come from? And what stops a controller from DDoSing the API server by listing everything every five seconds like my script did?\\nThe answer is the informer. An informer opens a single watch against the API server for a given resource type, streams every add, update, and delete, and keeps a complete in-memory cache of the objects we're interested in. Two things matter here:\\nFirst, the informer turns each watch event into a key and puts it on a work queue. The queue does a lot of work for you.\\nIt coalesces: if the same object is updated five times before you get to it, you reconcile it once, against the latest state (level-triggered again).\\nIt rate-limits: an object that keeps erroring backs off exponentially instead of spinning. This is damping, the same reason a crash-looping container backs off instead of restarting hot.\\nIt lets you run a pool of workers pulling keys in parallel, which is your fan-out. Events fan in from the watch, collapse in the queue, and fan out to the workers. This is something you need to tune. The higher you set the pool, the more pressure you put on the system: more writes, more API calls, more load on the provider, and more CPU usage in the operator.\\n\\nSecond, and this is a detail that bites people a lot: your reads and your writes in Kubernetes don't go to the same place.\\nIn controller-runtime, the client you're handed reads from the informer's local cache. Cache reads are cheap, they don't touch the API server, and that's how a controller reconciles thousands of objects without falling over. But your writes go straight to the API server. The cache only learns about your write when the resulting watch event comes back around, a moment later.\\nBecause of that, a read can be stale. You need to be prepared for this.\\nIf you write a field of an object and then read the same object again from the cache, the reconciler might think it's not updated yet. You write again, and you get a Conflict error. Retrying with a fresh read can be fine, but blindly retrying against the same stale cached view just spins.\\nMost of the time, what you want is to drop the call and requeue. In the next reconcile, the GET will see the updated object, and your write will never happen. That's how everything self-converges.\\nHere is another edge case. Picture this sequence inside a reconcile:// I want N replicas. I see fewer, so I create the missing ones.\\nexisting, _ := r.listChildPods(ctx) // reads the CACHE\\nfor i := len(existing); i \u003c desired; i++ {\\n r.client.Create(ctx, newPod(i)) // writes the API SERVER\\n}\\n\\nNow an event fires again a second later, before the cache has caught up with the Pods you just created. You list from the cache, and the new Pods aren't there yet. Your code decides it still needs to create them, and you create duplicates. This is the classic stale-cache double-create, and it's nasty because it only shows up under timing you can't reproduce on your laptop.\\n\\nThere are two ways out. The correct one is the expectations pattern, the same trick the built-in ReplicaSet controller uses: you record that you expect to see N creations in memory, and you don't act again until the cache has caught up to your own writes. It works, but it's not easy to implement and it's a fair amount of machinery. Read more on Ahmet's blog.\\nThe pragmatic one, which a lot of people use, is to bypass the cache for the reads where a stale view would cause a double-create or double-delete, and go straight to the API server:// The cached client can be stale right after our own writes, which\\n// would make us miscount and create duplicates. For this one read,\\n// go direct to the API server instead of the cache. Slower,\\n// but consistent for this decision.\\nerr := r.apiReader.List(ctx, \u0026instances, client.InNamespace(ns), labelSelector)\\n\\nThis is not only a Kubernetes issue. In any system with a read cache and a write-through path, read-after-write is not consistent unless you make it so. Most of the time the cache is exactly what you want: cheap, local, and eventually consistent. Eventual consistency is fine because the loop runs again. But the moment a decision would be destructive or non-idempotent if you acted on a stale read, you need to know which path you're on. Kubernetes solves many hard problems, but it also gives you a few new ones.\\nSetpoint and measured variable: spec and status\\nBack to the control diagram. My script kept its setpoint in shell variables and its measured state in the output of df and psql. Kubernetes gives both a permanent home, on the object itself.\\n.spec is the setpoint, the desired state. It's owned by whoever created the object (a human, or another controller), and the reconciler treats it as read-only intent. It's an anti-pattern to write to the .spec from inside the controller. If you do it, stop reading, go and fix your codebase. There are only a handful of exceptions, but a controller should generally never set its own setpoint.\\n.status is the measured variable, the observed state. It's owned by the controller, written through a separate status subresource, and it's where you record what's actually true. The better the status, the better the controller can decide. A good .status field is what makes a controller pleasant to operate. The word observability comes from control theory; Kalman coined it around 1960 to ask whether you can infer a system's internal state from its outputs. .status is also your response to any third-party system. If someone wants to learn the outcome of your actions, .status is the place to look at.\\nThat split is the whole declarative model in two fields. It comes with a piece of bookkeeping that's pure control theory: .metadata.generation increments when desired state changes, and by convention the controller writes back .status.observedGeneration to say \\\"the state I'm reporting reflects this version of your intent.\\\"\\nWhen observedGeneration \u003c generation, the status you're looking at does not reflect the latest setpoint yet. That one comparison is how you tell \\\"converged\\\" from \\\"still working on it.\\\"\\n\\nThis is why the reconcile is stateless, and why that matters. Our shell script kept \\\"am I mid-failover?\\\" in a variable that died with the process. A Kubernetes controller keeps nothing important in memory. Every fact it needs is on an API object: the spec it's driving toward, the status it last observed, the conditions describing where things stand. Kill the controller, restart it on another node, and it picks up exactly where it left off, not because it saved its progress, but because there was never any in-memory progress to lose. The state lives in the cluster (API server, etcd is what holds the state). The controller is just the loop that reads it.\\nSelf-healing by design\\nThis is the part I like most.\\nWhen a controller creates a child object (a Pod, a PVC), it stamps an ownerReference on the child pointing back at the parent. That reference does two things. It sets up garbage collection: delete the parent, and Kubernetes can cascade the delete to its children. And it gives the controller a way to map child changes back to the parent: \\\"when any object I own changes, enqueue my parent for a reconcile.\\\" ownerReference allows you to link controllers to each other and create chains. If done right, all your controllers and systems fit together.\\nHere is an example. Follow the loop:\\nA node dies and takes a Pod with it.\\nThe Pod's deletion is a watch event, an edge.\\nThrough the ownership link, that edge becomes a reconcile request for the parent.\\nThe parent reconciles, observes its children (level-triggered), sees one is missing and the count is below the setpoint, and creates a replacement.\\nThe replacement is an unscheduled Pod, an edge for the scheduler.\\nThe scheduler detects the unscheduled Pod, assigns a node.\\nThe kubelet gets triggered because that's an edge for that node's kubelet and it starts the container.\\nThat's multiple feedback loops, each watching the layer below, each reacting to an edge and converging to its own level, chained together through the API server with nobody orchestrating the whole thing.\\nControl theory has a name for loops stacked like this: cascade control. The output of an outer loop becomes the setpoint of an inner loop. A controller never writes its own .spec, but it writes other objects' .spec all the time. My operator writes the PVC's spec, and that spec is the setpoint the CSI controllers converge to. Each loop worries only about its own layer and trusts the loop below.\\nSo we wrote if ! pg_isready; then docker start; fi and maybe thought we're good. Kubernetes turns that one line into several independent controllers that have never heard of each other, but still cooperate because they share the API server and watch each other's objects. I like this part a lot. Nobody calls a central orchestrator. Nobody passes a private message. The system heals itself.\\nWhat \\\"observe\\\" actually means in a real operator\\nUp to here I've been a little vague about the \\\"measure\\\" step, because in the examples the measured state is just \\\"list the child Pods.\\\" But in a real database operator it's a lot more than that.\\nWhen an operator I work on reconciles a single Postgres instance, the first thing it does, before it decides anything, is build a snapshot of reality from every source that knows something true about that instance. Not just Kubernetes. Kubernetes barely knows anything about whether Postgres is actually healthy.\\nThe sources gathered at the top of every reconcile:\\nThe Kubernetes cache: the Pod, its PVC, the PV behind it, the Node it's on, the ConfigMap holding its config. These are the cheap local reads, the stuff we already talked about.\\nThe database's effective configuration. Not what we last wrote down, but what the server has actually loaded, so we can compare the two and detect drift. Other entities can rewrite or reload the config on disk without us knowing, so the only honest source of truth is the running server itself, never our last write.\\nThe database's own view of its health. Its role, whether it's healthy, how far behind its followers are, whether it's currently accepting writes. Some of this comes from the agents that sit next to the database and manage it; some we get by opening a connection and asking the database directly. These calls carry a tight timeout and are allowed to fail, more on that below.\\nA background collector. Some signals are too expensive or too rate-limited to fetch on every reconcile: disk usage, or whether a volume operation we kicked off earlier is still in flight and where it sits in its cooldown window. A separate collector, often a background goroutine, gathers these on a slow cadence and keeps the last value per volume in memory. The reconcile reads that value instantly, without blocking on anything. Think of these as custom workqueues you implement.\\n\\nIn code, the snapshot is just a struct, and the reconcile's first move is to populate it. This is simplified, but faithful to the real shape:// The observation snapshot: everything we know about this instance, right now.\\ntype reconcileHandler struct {\\n // The object (.spec = setpoint, .status = measured).\\n instance *v1.PostgresInstance\\n\\n // Kubernetes objects.\\n pod *corev1.Pod\\n pvc *corev1.PersistentVolumeClaim\\n node *corev1.Node\\n\\n // Database state.\\n dbState DatabaseState\\n\\n // What Postgres actually loaded, not what we last wrote.\\n effectiveConfig map[string]string\\n\\n // Collected out-of-band.\\n diskUsage *resource.Quantity\\n\\n // The volume operation already in flight, if any.\\n storageOp *StorageOperation\\n}\\n\\nfunc (r *Reconciler) newReconcileHandler(\\n ctx context.Context,\\n inst *v1.PostgresInstance,\\n) (*reconcileHandler, error) {\\n h := \u0026reconcileHandler{instance: inst}\\n\\n // Cheap local reads.\\n h.pod, h.pvc, h.node = r.fetchKubeObjects(ctx, inst)\\n\\n // Active database calls.\\n h.dbState = r.queryDatabase(ctx, h.pod)\\n h.effectiveConfig = r.readEffectiveConfig(ctx, h.pod)\\n\\n // Values from the collector/metric.\\n h.diskUsage = r.collector.Usage(h.pvc)\\n h.storageOp = r.collector.InFlightOp(h.pvc)\\n\\n return h, nil\\n}\\n\\nA few things about this are deliberate, and only look obvious after you've been burned once or twice.\\nGather once, at the top. Every sub-decision in the reconcile reads from this one snapshot. We don't re-query the database in the middle of the loop, or read the disk usage again three functions deep. If we did, different parts of the same reconcile could see different versions of reality. This sounds like a small detail, but it changes the whole design.\\nFor example, the database might be the leader when we check at the top and a replica by the time another helper checks again. Then you get decisions that are individually reasonable, but wrong together. We have a rule in the codebase against stashing state back onto this handler mid-reconcile to pass between steps, because it reintroduces exactly the inconsistency we gathered the snapshot to avoid. Making the reconcileHandler immutable is one way to enforce that rule in the type system instead of relying on code review.\\nPartial failures are tolerated. Reaching the database can fail while the Kubernetes reads succeed. That's not always an error that aborts the reconcile. It's a measured fact: \\\"Postgres is currently unreachable.\\\" That itself is something to record in status. A control loop that gives up entirely whenever one sensor is unavailable is a control loop that's down a lot. We degrade instead. Think of a car. If the rain sensor for the wipers is broken, the whole car doesn't stop. You can still drive, but you need to turn on a few things yourself.\\nOnce the data snapshot exists, the reconcile is a sequence of small, idempotent steps, each comparing one slice of desired against observed and acting to close the gap:func (r *reconcileHandler) reconcile(ctx context.Context) (reconcile.Result, error) {\\n var rb results.Builder\\n rb.Merge(r.reconcileConfigMap(ctx)) // push desired config\\n rb.Merge(r.reconcileDatabase(ctx)) // reload/restart if params drifted\\n rb.Merge(r.reconcilePVC(ctx)) // grow the disk if needed\\n rb.Merge(r.reconcilePod(ctx)) // create/replace the Pod\\n rb.Merge(r.reconcileStatus(ctx)) // always last: write what we observed\\n return rb.Result()\\n}\\n\\nStatus is written last on purpose, because it's the measured variable: you record what's true after you've taken your actions and observed the result. Again, in our operators, it's not possible to write the status mid-reconcile.\\nEach step is independently idempotent. Each returns a result, either \\\"I'm done\\\" or \\\"requeue me in 30 seconds, I'm waiting on something,\\\" and the results merge. It reads almost exactly like the body of my shell loop. The difference is that \\\"observe the state\\\" grew from df and pg_isready into a fan-in across multiple systems, and \\\"take an action\\\" grew from ssh into typed, conflict-aware API writes.\\nThis is the operator. The kubelet, the scheduler, CSI, and CNI are infrastructure we get by using Kubernetes. This loop, with its messy real-world observe step, is the part we actually write and deal with. Because we know how the underlying system works, we can design it without treating Kubernetes like a black box.\\nNot every edge comes from the API server\\nThere's one more piece, and it lets me close a loop from Part 1 that I left deliberately: the disk-usage check.\\nMy shell script polled df on every node every five seconds. For three nodes, fine. For thousands of databases, you can't reconcile every one of them every few seconds just to check a number that rarely changes; you'd spend all your CPU re-deriving \\\"still at 40%, still at 40%, still at 40%.\\\" This is the level-triggered model's one real cost: re-checking everything is correct, but it isn't free.\\nThe fix is to add a sensor that emits its own edges. A background collector polls our metrics pipeline for disk usage on a slow cadence, keeps the last value per volume in memory, and only emits an event when usage crosses a threshold, not while it sits above or below one:// Edge detection. We fire only on the transition across the threshold,\\n// not every cycle we happen to be above it. Hovering at 81% is silent;\\n// crossing 80% upward is an event.\\ncrossedUp := previousUsage \u003c pvc.GrowThreshold \u0026\u0026 usage \u003e= pvc.GrowThreshold\\nif crossedUp {\\n // -\u003e generic event -\u003e work queue -\u003e reconcile\\n relay.Send(Event{Key: pvc.Key})\\n}\\n\\nThat event goes into the same work queue as the API watch events and triggers a normal reconcile of the affected instance. Same rule as before: the event wakes us up, the reconcile decides from the current state.\\nFor example, say we have a 10GiB disk and it's using 8GiB. The collector saw it cross the threshold, so it wakes the reconciler. The reconciler reads the current usage, sees that it crossed the 80% threshold, and sets a new size on the PVC. After that, CSI handles the rest.\\nAnd because edges can be missed (the collector could be down, an event could be dropped from a full channel), there's a resyncer: a periodic timer that enqueues every object for reconcile every minute or so, regardless of events. It's the safety net. It's our sleep 5 loop. There's also RequeueAfter, which a reconcile returns to say \\\"wake me again in 30 seconds,\\\" the controller's way of polling a slow external operation without holding a worker.\\nThere are two more questions: how often should the loop run, and who is allowed to run it? Control theory calls the first one the sampling interval. The rule of thumb: act faster than the thing you're tracking changes, but not faster than it can respond. Reconciling a disk that fills over hours every few milliseconds just burns CPU to learn the same thing again.\\nSo the operator puts boundaries around it.\\nA coalescing delay handles noisy edge events: a burst of events for one object becomes one reconcile (think of it like a fan-in), not a thousand.\\nThe resyncer is the safety net: every object gets looked at once in a while, even when nothing fires.\\nAnd leader election answers the who: only one copy of the operator runs the loop at a time. Two controllers writing to the same database object is not \\\"more reliable.\\\" Even with idempotent controllers, they'll be requeueing due to conflicts and consuming unnecessary compute. In theory, a perfectly written controller should tolerate this. In practice, software is rarely perfect, and the safer boundary is worth it.\\nTo close out Part 2, let me redraw the control loop again. The diagram in Part 1 had a few basic boxes. Now, the same loop represents a closed feedback loop more realistically:\\n\\nThere is one new arrow in this diagram: disturbances. A controller has two jobs. The first is setpoint tracking: someone edits the .spec, and the loop chases the new intent. The second is disturbance rejection: the world changes on its own. A node dies, a customer starts a bulk import, someone deletes a Pod by hand. The level-triggered reconcile treats both the same way: it only sees the gap.\\nOur controller doesn't always touch Postgres directly. Sometimes it writes a PVC and lets CSI do the storage work. Sometimes it creates a Pod and lets the scheduler and kubelet do their part. This is what a production operator looks like: one loop we write, surrounded by other loops we don't write.\\nNotice that every decision in this loop has been binary: start the Pod or don't, grow the disk or don't, rewrite the config or don't. That's an on/off controller, and it covers most of what an operator does. But not every question is yes/no; once the answer becomes how much rather than whether, you need a controller with memory and a sense of trend: how long you've been off, and how fast it's changing. That's a separate post.\\nConclusion\\nAll of this works, and most of the time it runs without anyone watching it. But the abstractions still leak, and they usually leak at a bad time.\\nEventual consistency and the split between cache reads and API writes mean that a freshly-created object might not be visible to the thing that just created it. When something goes wrong, we're debugging Kubernetes objects, database state, metrics, volume operations, and sometimes the cloud provider at the same time. The bug is usually not in one clean place.\\nThe declarative model is wonderful until it meets an operation that's inherently imperative and stateful, like a failover, a major-version upgrade, or a data migration. Then you have to turn a blocking, non-idempotent action into an idempotent one. That's a whole other blog post.\\nThat complexity is easy to underestimate. If you're not dealing with sophisticated systems, if you can sacrifice availability, or if you don't care about scalability, maybe all this machinery isn't needed at all. Operators do not remove complexity. They move it into code someone has to understand.\\nI still think it's worth it. For running thousands of databases that have to heal themselves without anyone watching, I don't know a better alternative. The hard parts are hard because the problem is hard, not because Kubernetes made it hard.\\nKubernetes is not only a container runtime. It's not only a YAML processor, or an orchestrator, or whatever word we use that year. For me, the useful way to read Kubernetes is this: Kubernetes is a framework for feedback controllers, plus a consistent store to hold their setpoints and a shared event bus to wake them up.\\nOnce you see that, the rest fits together. The kubelet, the scheduler, CSI, and your operator all read and write facts onto shared objects, and each one tries to move its own small part of the system toward the desired state. The core idea is still the same one we started with: write down what you want, look at what exists, make the next change, and repeat. Events wake the loop up, but the current state decides what happens.\\nKubernetes didn't invent these ideas; a thermostat had them long before us. The mapping to control theory is not perfect, and some boundaries are fuzzy. But the core idea holds. We are writing feedback loops in Go and applying them to databases. Mechanical and electrical engineers figured out how to build stable, long-running systems before us. Software engineering is still catching up, and Kubernetes gives us a practical way to use those ideas in production.\",\"aside\",\"toc\",[44,45,46],\"title\",\"The feedback loops behind Kubernetes\",\"authors\",[39],\"categories\",[38],\"excerpt\",\"Kubernetes is a framework for feedback controllers: write down what you want, observe what exists, make the next change, and repeat.\",\"createdAt\",\"2026-06-16\",\"slug\",\"the-feedback-loops-behind-kubernetes\",\"meta\",{\"_33\":34,\"_35\":26,\"_36\":37,\"_19\":20},\"canonical\",\"https://planetscale.com/blog/the-feedback-loops-behind-kubernetes\",\"description\",\"image\",\"/assets/the-feedback-loops-behind-kubernetes-social-DXEE2U-o.png\",\"engineering\",{\"_29\":40,\"_41\":42,\"_43\":40},\"fatih\",\"name\",\"Fatih Arslan\",\"x\",{\"_47\":90,\"_49\":91,\"_51\":52,\"_19\":92},{\"_47\":54,\"_49\":55,\"_51\":52,\"_19\":56},{\"_47\":48,\"_49\":50,\"_51\":52,\"_19\":53},\"children\",[],\"id\",\"conclusion\",\"level\",2,\"Conclusion\",[57,58,59,60,61,62,63,64],\"part-2-how-we-reinvented-kubernetes\",\"Part 2: how we reinvented Kubernetes\",{\"_47\":87,\"_49\":88,\"_51\":67,\"_19\":89},{\"_47\":84,\"_49\":85,\"_51\":67,\"_19\":86},{\"_47\":81,\"_49\":82,\"_51\":67,\"_19\":83},{\"_47\":78,\"_49\":79,\"_51\":67,\"_19\":80},{\"_47\":75,\"_49\":76,\"_51\":67,\"_19\":77},{\"_47\":72,\"_49\":73,\"_51\":67,\"_19\":74},{\"_47\":69,\"_49\":70,\"_51\":67,\"_19\":71},{\"_47\":65,\"_49\":66,\"_51\":67,\"_19\":68},[],\"not-every-edge-comes-from-the-api-server\",3,\"Not every edge comes from the API server\",[],\"what-observe-actually-means-in-a-real-operator\",\"What \\\"observe\\\" actually means in a real operator\",[],\"self-healing-by-design\",\"Self-healing by design\",[],\"setpoint-and-measured-variable-spec-and-status\",\"Setpoint and measured variable: spec and status\",[],\"informers-the-work-queue-and-a-cache\",\"Informers, the work queue, and a cache\",[],\"edge-triggered-notifications-level-triggered-logic\",\"Edge-triggered notifications, level-triggered logic\",[],\"the-for-loop-translated-to-kubernetes\",\"The for-loop translated to Kubernetes\",[],\"the-other-loops\",\"The other loops\",[93,94,95,96,97,98,99,100],\"part-1-running-postgres-by-hand\",\"Part 1: running Postgres by hand\",{\"_47\":122,\"_49\":123,\"_51\":67,\"_19\":124},{\"_47\":119,\"_49\":120,\"_51\":67,\"_19\":121},{\"_47\":116,\"_49\":117,\"_51\":67,\"_19\":118},{\"_47\":113,\"_49\":114,\"_51\":67,\"_19\":115},{\"_47\":110,\"_49\":111,\"_51\":67,\"_19\":112},{\"_47\":107,\"_49\":108,\"_51\":67,\"_19\":109},{\"_47\":104,\"_49\":105,\"_51\":67,\"_19\":106},{\"_47\":101,\"_49\":102,\"_51\":67,\"_19\":103},[],\"what-we-actually-built\",\"What we actually built\",[],\"changing-a-parameter\",\"Changing a parameter\",[],\"the-watchdog-script\",\"The watchdog script\",[],\"they-have-to-find-each-other\",\"They have to find each other\",[],\"one-isnt-enough\",\"One isn't enough\",[],\"it-needs-a-real-disk\",\"It needs a real disk\",[],\"pick-a-node-by-hand\",\"Pick a node, by hand\",[],\"one-container-one-machine\",\"One container, one machine\",[\"SingleFetchClassInstance\",2115],[\"SingleFetchClassInstance\",2111],[\"SingleFetchClassInstance\",2107],[\"SingleFetchClassInstance\",2103],[\"SingleFetchClassInstance\",2066],[\"SingleFetchClassInstance\",2063],[\"SingleFetchClassInstance\",2055],[\"SingleFetchClassInstance\",2047],[\"SingleFetchClassInstance\",2043],[\"SingleFetchClassInstance\",2039],[\"SingleFetchClassInstance\",2035],[\"SingleFetchClassInstance\",2021],[\"SingleFetchClassInstance\",2013],[\"SingleFetchClassInstance\",1998],[\"SingleFetchClassInstance\",1994],[\"SingleFetchClassInstance\",1985],[\"SingleFetchClassInstance\",1977],[\"SingleFetchClassInstance\",1973],[\"SingleFetchClassInstance\",1969],[\"SingleFetchClassInstance\",1965],[\"SingleFetchClassInstance\",1957],[\"SingleFetchClassInstance\",1931],[\"SingleFetchClassInstance\",1922],[\"SingleFetchClassInstance\",1914],[\"SingleFetchClassInstance\",1910],[\"SingleFetchClassInstance\",1890],[\"SingleFetchClassInstance\",1885],[\"SingleFetchClassInstance\",1881],[\"SingleFetchClassInstance\",1868],[\"SingleFetchClassInstance\",1860],[\"SingleFetchClassInstance\",1856],[\"SingleFetchClassInstance\",1833],[\"SingleFetchClassInstance\",1829],[\"SingleFetchClassInstance\",1825],[\"SingleFetchClassInstance\",1815],[\"SingleFetchClassInstance\",1806],[\"SingleFetchClassInstance\",1798],[\"SingleFetchClassInstance\",1772],[\"SingleFetchClassInstance\",1768],[\"SingleFetchClassInstance\",1763],[\"SingleFetchClassInstance\",1740],[\"SingleFetchClassInstance\",1732],[\"SingleFetchClassInstance\",1723],[\"SingleFetchClassInstance\",1708],[\"SingleFetchClassInstance\",1698],[\"SingleFetchClassInstance\",1684],[\"SingleFetchClassInstance\",1674],[\"SingleFetchClassInstance\",1583],[\"SingleFetchClassInstance\",1568],[\"SingleFetchClassInstance\",1564],[\"SingleFetchClassInstance\",1531],[\"SingleFetchClassInstance\",1527],[\"SingleFetchClassInstance\",1524],[\"SingleFetchClassInstance\",1516],[\"SingleFetchClassInstance\",1512],[\"SingleFetchClassInstance\",1504],[\"SingleFetchClassInstance\",1500],[\"SingleFetchClassInstance\",1486],[\"SingleFetchClassInstance\",1465],[\"SingleFetchClassInstance\",1412],[\"SingleFetchClassInstance\",1384],[\"SingleFetchClassInstance\",1375],[\"SingleFetchClassInstance\",1361],[\"SingleFetchClassInstance\",1351],[\"SingleFetchClassInstance\",1347],[\"SingleFetchClassInstance\",1339],[\"SingleFetchClassInstance\",1323],[\"SingleFetchClassInstance\",1319],[\"SingleFetchClassInstance\",1315],[\"SingleFetchClassInstance\",1307],[\"SingleFetchClassInstance\",1303],[\"SingleFetchClassInstance\",1270],[\"SingleFetchClassInstance\",1266],[\"SingleFetchClassInstance\",1262],[\"SingleFetchClassInstance\",1258],[\"SingleFetchClassInstance\",1245],[\"SingleFetchClassInstance\",1241],[\"SingleFetchClassInstance\",1226],[\"SingleFetchClassInstance\",1222],[\"SingleFetchClassInstance\",1213],[\"SingleFetchClassInstance\",1173],[\"SingleFetchClassInstance\",1159],[\"SingleFetchClassInstance\",1149],[\"SingleFetchClassInstance\",1141],[\"SingleFetchClassInstance\",1137],[\"SingleFetchClassInstance\",1121],[\"SingleFetchClassInstance\",1111],[\"SingleFetchClassInstance\",1075],[\"SingleFetchClassInstance\",1062],[\"SingleFetchClassInstance\",1053],[\"SingleFetchClassInstance\",1049],[\"SingleFetchClassInstance\",1045],[\"SingleFetchClassInstance\",1041],[\"SingleFetchClassInstance\",1025],[\"SingleFetchClassInstance\",1021],[\"SingleFetchClassInstance\",1017],[\"SingleFetchClassInstance\",1013],[\"SingleFetchClassInstance\",999],[\"SingleFetchClassInstance\",982],[\"SingleFetchClassInstance\",978],[\"SingleFetchClassInstance\",974],[\"SingleFetchClassInstance\",970],[\"SingleFetchClassInstance\",962],[\"SingleFetchClassInstance\",948],[\"SingleFetchClassInstance\",930],[\"SingleFetchClassInstance\",887],[\"SingleFetchClassInstance\",871],[\"SingleFetchClassInstance\",861],[\"SingleFetchClassInstance\",848],[\"SingleFetchClassInstance\",832],[\"SingleFetchClassInstance\",824],[\"SingleFetchClassInstance\",820],[\"SingleFetchClassInstance\",805],[\"SingleFetchClassInstance\",801],[\"SingleFetchClassInstance\",762],[\"SingleFetchClassInstance\",758],[\"SingleFetchClassInstance\",727],[\"SingleFetchClassInstance\",717],[\"SingleFetchClassInstance\",709],[\"SingleFetchClassInstance\",705],[\"SingleFetchClassInstance\",701],[\"SingleFetchClassInstance\",697],[\"SingleFetchClassInstance\",654],[\"SingleFetchClassInstance\",640],[\"SingleFetchClassInstance\",636],[\"SingleFetchClassInstance\",632],[\"SingleFetchClassInstance\",628],[\"SingleFetchClassInstance\",619],[\"SingleFetchClassInstance\",609],[\"SingleFetchClassInstance\",600],[\"SingleFetchClassInstance\",596],[\"SingleFetchClassInstance\",592],[\"SingleFetchClassInstance\",588],[\"SingleFetchClassInstance\",567],[\"SingleFetchClassInstance\",563],[\"SingleFetchClassInstance\",554],[\"SingleFetchClassInstance\",550],[\"SingleFetchClassInstance\",540],[\"SingleFetchClassInstance\",536],[\"SingleFetchClassInstance\",526],[\"SingleFetchClassInstance\",522],[\"SingleFetchClassInstance\",518],[\"SingleFetchClassInstance\",497],[\"SingleFetchClassInstance\",481],[\"SingleFetchClassInstance\",477],[\"SingleFetchClassInstance\",433],[\"SingleFetchClassInstance\",429],[\"SingleFetchClassInstance\",402],[\"SingleFetchClassInstance\",371],[\"SingleFetchClassInstance\",367],[\"SingleFetchClassInstance\",344],[\"SingleFetchClassInstance\",340],[\"SingleFetchClassInstance\",329],[\"SingleFetchClassInstance\",325],[\"SingleFetchClassInstance\",321],[\"SingleFetchClassInstance\",317],[\"SingleFetchClassInstance\",313],[\"SingleFetchClassInstance\",309],[\"SingleFetchClassInstance\",298],[\"SingleFetchClassInstance\",294],[\"SingleFetchClassInstance\",286],{\"_287\":288,\"_41\":289,\"_290\":291,\"_47\":292},\"$$mdtype\",\"Tag\",\"p\",\"attributes\",{},[293],\"Kubernetes didn't invent these ideas; a thermostat had them long before us. The mapping to control theory is not perfect, and some boundaries are fuzzy. But the core idea holds. We are writing feedback loops in Go and applying them to databases. Mechanical and electrical engineers figured out how to build stable, long-running systems before us. Software engineering is still catching up, and Kubernetes gives us a practical way to use those ideas in production.\",{\"_287\":288,\"_41\":289,\"_290\":295,\"_47\":296},{},[297],\"Once you see that, the rest fits together. The kubelet, the scheduler, CSI, and your operator all read and write facts onto shared objects, and each one tries to move its own small part of the system toward the desired state. The core idea is still the same one we started with: write down what you want, look at what exists, make the next change, and repeat. Events wake the loop up, but the current state decides what happens.\",{\"_287\":288,\"_41\":289,\"_290\":299,\"_47\":300},{},[301,302,303],\"Kubernetes is not only a container runtime. It's not only a YAML processor, or an orchestrator, or whatever word we use that year. For me, the useful way to read Kubernetes is this: \",[\"SingleFetchClassInstance\",304],\", plus a consistent store to hold their setpoints and a shared event bus to wake them up.\",{\"_287\":288,\"_41\":305,\"_290\":306,\"_47\":307},\"strong\",{},[308],\"Kubernetes is a framework for feedback controllers\",{\"_287\":288,\"_41\":289,\"_290\":310,\"_47\":311},{},[312],\"I still think it's worth it. For running thousands of databases that have to heal themselves without anyone watching, I don't know a better alternative. The hard parts are hard because the problem is hard, not because Kubernetes made it hard.\",{\"_287\":288,\"_41\":289,\"_290\":314,\"_47\":315},{},[316],\"That complexity is easy to underestimate. If you're not dealing with sophisticated systems, if you can sacrifice availability, or if you don't care about scalability, maybe all this machinery isn't needed at all. Operators do not remove complexity. They move it into code someone has to understand.\",{\"_287\":288,\"_41\":289,\"_290\":318,\"_47\":319},{},[320],\"The declarative model is wonderful until it meets an operation that's inherently imperative and stateful, like a failover, a major-version upgrade, or a data migration. Then you have to turn a blocking, non-idempotent action into an idempotent one. That's a whole other blog post.\",{\"_287\":288,\"_41\":289,\"_290\":322,\"_47\":323},{},[324],\"Eventual consistency and the split between cache reads and API writes mean that a freshly-created object might not be visible to the thing that just created it. When something goes wrong, we're debugging Kubernetes objects, database state, metrics, volume operations, and sometimes the cloud provider at the same time. The bug is usually not in one clean place.\",{\"_287\":288,\"_41\":289,\"_290\":326,\"_47\":327},{},[328],\"All of this works, and most of the time it runs without anyone watching it. But the abstractions still leak, and they usually leak at a bad time.\",{\"_287\":288,\"_41\":330,\"_290\":331,\"_47\":332},\"h2\",{\"_49\":50},[333],[\"SingleFetchClassInstance\",334],{\"_287\":288,\"_41\":335,\"_290\":336,\"_47\":337},\"a\",{\"_338\":339},[53],\"href\",\"#conclusion\",{\"_287\":288,\"_41\":341,\"_290\":342,\"_47\":343},\"hr\",{},[],{\"_287\":288,\"_41\":289,\"_290\":345,\"_47\":346},{},[347,348,349,350,351,352,353],\"Notice that every decision in this loop has been binary: start the Pod or don't, grow the disk or don't, rewrite the config or don't. That's an \",[\"SingleFetchClassInstance\",363],\", and it covers most of what an operator does. But not every question is yes/no; once the answer becomes \",[\"SingleFetchClassInstance\",359],\" rather than \",[\"SingleFetchClassInstance\",354],\", you need a controller with memory and a sense of trend: how long you've been off, and how fast it's changing. That's a separate post.\",{\"_287\":288,\"_41\":355,\"_290\":356,\"_47\":357},\"em\",{},[358],\"whether\",{\"_287\":288,\"_41\":355,\"_290\":360,\"_47\":361},{},[362],\"how much\",{\"_287\":288,\"_41\":355,\"_290\":364,\"_47\":365},{},[366],\"on/off controller\",{\"_287\":288,\"_41\":289,\"_290\":368,\"_47\":369},{},[370],\"Our controller doesn't always touch Postgres directly. Sometimes it writes a PVC and lets CSI do the storage work. Sometimes it creates a Pod and lets the scheduler and kubelet do their part. This is what a production operator looks like: one loop we write, surrounded by other loops we don't write.\",{\"_287\":288,\"_41\":289,\"_290\":372,\"_47\":373},{},[374,375,376,377,378,379,380,381,382],\"There is one new arrow in this diagram: \",[\"SingleFetchClassInstance\",398],\". A controller has two jobs. The first is \",[\"SingleFetchClassInstance\",393],\": someone edits the \",[\"SingleFetchClassInstance\",388],\", and the loop chases the new intent. The second is \",[\"SingleFetchClassInstance\",383],\": the world changes on its own. A node dies, a customer starts a bulk import, someone deletes a Pod by hand. The level-triggered reconcile treats both the same way: it only sees the gap.\",{\"_287\":288,\"_41\":335,\"_290\":384,\"_47\":385},{\"_338\":387},[386],\"disturbance rejection\",\"https://en.wikipedia.org/wiki/Control_theory\",{\"_287\":288,\"_41\":389,\"_290\":390,\"_47\":391},\"code\",{},[392],\".spec\",{\"_287\":288,\"_41\":335,\"_290\":394,\"_47\":395},{\"_338\":397},[396],\"setpoint tracking\",\"https://en.wikipedia.org/wiki/Setpoint_%28control_system%29\",{\"_287\":288,\"_41\":305,\"_290\":399,\"_47\":400},{},[401],\"disturbances\",{\"_287\":288,\"_41\":289,\"_290\":403,\"_47\":404},{},[405],[\"SingleFetchClassInstance\",406],{\"_287\":288,\"_41\":407,\"_290\":408,\"_47\":409},\"ContentImage\",{\"_410\":411,\"_412\":413,\"_414\":415,\"_416\":417,\"_418\":419,\"_420\":421},[],\"alt\",\"Production operator feedback loop diagram\",\"height\",943,\"loading\",\"lazy\",\"src\",\"https://planetscale-images.imgix.net/assets/part2-operator-loop-YZG6qYVG.png?auto=compress%2Cformat\",\"srcs\",[422,423],\"width\",1504,{\"_424\":417,\"_426\":428},{\"_424\":425,\"_426\":427},\"srcSet\",\"https://planetscale-images.imgix.net/assets/part2-operator-loop-darkmode-jOHBdP3P.png?auto=compress%2Cformat\",\"media\",\"(prefers-color-scheme: dark)\",\"(prefers-color-scheme: light), (prefers-color-scheme: no-preference)\",{\"_287\":288,\"_41\":289,\"_290\":430,\"_47\":431},{},[432],\"To close out Part 2, let me redraw the control loop again. The diagram in Part 1 had a few basic boxes. Now, the same loop represents a closed feedback loop more realistically:\",{\"_287\":288,\"_41\":434,\"_290\":435,\"_47\":436},\"ul\",{},[437,438,439],[\"SingleFetchClassInstance\",467],[\"SingleFetchClassInstance\",457],[\"SingleFetchClassInstance\",440],{\"_287\":288,\"_41\":441,\"_290\":442,\"_47\":443},\"li\",{},[444,445,446,447,448],\"And \",[\"SingleFetchClassInstance\",453],\" answers the \",[\"SingleFetchClassInstance\",449],\": only one copy of the operator runs the loop at a time. Two controllers writing to the same database object is not \\\"more reliable.\\\" Even with idempotent controllers, they'll be requeueing due to conflicts and consuming unnecessary compute. In theory, a perfectly written controller should tolerate this. In practice, software is rarely perfect, and the safer boundary is worth it.\",{\"_287\":288,\"_41\":355,\"_290\":450,\"_47\":451},{},[452],\"who\",{\"_287\":288,\"_41\":305,\"_290\":454,\"_47\":455},{},[456],\"leader election\",{\"_287\":288,\"_41\":441,\"_290\":458,\"_47\":459},{},[460,461,462],\"The \",[\"SingleFetchClassInstance\",463],\" is the safety net: every object gets looked at once in a while, even when nothing fires.\",{\"_287\":288,\"_41\":305,\"_290\":464,\"_47\":465},{},[466],\"resyncer\",{\"_287\":288,\"_41\":441,\"_290\":468,\"_47\":469},{},[470,471,472],\"A \",[\"SingleFetchClassInstance\",473],\" handles noisy edge events: a burst of events for one object becomes one reconcile (think of it like a fan-in), not a thousand.\",{\"_287\":288,\"_41\":305,\"_290\":474,\"_47\":475},{},[476],\"coalescing delay\",{\"_287\":288,\"_41\":289,\"_290\":478,\"_47\":479},{},[480],\"So the operator puts boundaries around it.\",{\"_287\":288,\"_41\":289,\"_290\":482,\"_47\":483},{},[484,485,486,487,488],\"There are two more questions: \",[\"SingleFetchClassInstance\",493],\" Control theory calls the first one the \",[\"SingleFetchClassInstance\",489],\". The rule of thumb: act faster than the thing you're tracking changes, but not faster than it can respond. Reconciling a disk that fills over hours every few milliseconds just burns CPU to learn the same thing again.\",{\"_287\":288,\"_41\":355,\"_290\":490,\"_47\":491},{},[492],\"sampling interval\",{\"_287\":288,\"_41\":305,\"_290\":494,\"_47\":495},{},[496],\"how often should the loop run, and who is allowed to run it?\",{\"_287\":288,\"_41\":289,\"_290\":498,\"_47\":499},{},[500,501,502,503,504,505,506],\"And because edges can be missed (the collector could be down, an event could be dropped from a full channel), there's a \",[\"SingleFetchClassInstance\",515],\": a periodic timer that enqueues every object for reconcile every minute or so, regardless of events. It's the safety net. It's our \",[\"SingleFetchClassInstance\",511],\" loop. There's also \",[\"SingleFetchClassInstance\",507],\", which a reconcile returns to say \\\"wake me again in 30 seconds,\\\" the controller's way of polling a slow external operation without holding a worker.\",{\"_287\":288,\"_41\":389,\"_290\":508,\"_47\":509},{},[510],\"RequeueAfter\",{\"_287\":288,\"_41\":389,\"_290\":512,\"_47\":513},{},[514],\"sleep 5\",{\"_287\":288,\"_41\":305,\"_290\":516,\"_47\":517},{},[466],{\"_287\":288,\"_41\":289,\"_290\":519,\"_47\":520},{},[521],\"For example, say we have a 10GiB disk and it's using 8GiB. The collector saw it cross the threshold, so it wakes the reconciler. The reconciler reads the current usage, sees that it crossed the 80% threshold, and sets a new size on the PVC. After that, CSI handles the rest.\",{\"_287\":288,\"_41\":289,\"_290\":523,\"_47\":524},{},[525],\"That event goes into the same work queue as the API watch events and triggers a normal reconcile of the affected instance. Same rule as before: the event wakes us up, the reconcile decides from the current state.\",{\"_287\":288,\"_41\":527,\"_290\":528,\"_47\":529},\"CodeBlock\",{\"_530\":531,\"_532\":533,\"_534\":535,\"_389\":-7},[],\"html\",\"\u003cpre class=\\\"shiki shiki-themes planetscale-light planetscale-dark\\\" style=\\\"--shiki-light:#2b2b2b;--shiki-dark:#e1e1e1;--shiki-light-bg:#ebebeb;--shiki-dark-bg:#1a1a1a\\\" tabindex=\\\"0\\\"\u003e\u003ccode\u003e\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#818181;--shiki-dark:#A1A1A1\\\"\u003e// Edge detection. We fire only on the transition across the threshold,\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#818181;--shiki-dark:#A1A1A1\\\"\u003e// not every cycle we happen to be above it. Hovering at 81% is silent;\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#818181;--shiki-dark:#A1A1A1\\\"\u003e// crossing 80% upward is an event.\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003ecrossedUp\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e :=\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e previousUsage\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e \u0026#x3C;\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e pvc\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003eGrowThreshold\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e \u0026#x26;\u0026#x26;\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e usage\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e \u003e=\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e pvc\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003eGrowThreshold\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003eif\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e crossedUp\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e {\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#818181;--shiki-dark:#A1A1A1\\\"\u003e // -\u003e generic event -\u003e work queue -\u003e reconcile\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e relay\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#5E49AF;--shiki-dark:#B7A5FB\\\"\u003eSend\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e(\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003eEvent\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e{\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003eKey\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e:\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e pvc\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003eKey\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e})\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e}\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003c/span\u003e\u003c/code\u003e\u003c/pre\u003e\",\"language\",\"go\",\"clipboard\",false,{\"_287\":288,\"_41\":289,\"_290\":537,\"_47\":538},{},[539],\"The fix is to add a sensor that emits its own edges. A background collector polls our metrics pipeline for disk usage on a slow cadence, keeps the last value per volume in memory, and only emits an event when usage crosses a threshold, not while it sits above or below one:\",{\"_287\":288,\"_41\":289,\"_290\":541,\"_47\":542},{},[543,544,545],\"My shell script polled \",[\"SingleFetchClassInstance\",546],\" on every node every five seconds. For three nodes, fine. For thousands of databases, you can't reconcile every one of them every few seconds just to check a number that rarely changes; you'd spend all your CPU re-deriving \\\"still at 40%, still at 40%, still at 40%.\\\" This is the level-triggered model's one real cost: re-checking everything is correct, but it isn't free.\",{\"_287\":288,\"_41\":389,\"_290\":547,\"_47\":548},{},[549],\"df\",{\"_287\":288,\"_41\":289,\"_290\":551,\"_47\":552},{},[553],\"There's one more piece, and it lets me close a loop from Part 1 that I left deliberately: the disk-usage check.\",{\"_287\":288,\"_41\":555,\"_290\":556,\"_47\":557},\"h3\",{\"_49\":66},[558],[\"SingleFetchClassInstance\",559],{\"_287\":288,\"_41\":335,\"_290\":560,\"_47\":561},{\"_338\":562},[68],\"#not-every-edge-comes-from-the-api-server\",{\"_287\":288,\"_41\":289,\"_290\":564,\"_47\":565},{},[566],\"This is the operator. The kubelet, the scheduler, CSI, and CNI are infrastructure we get by using Kubernetes. This loop, with its messy real-world observe step, is the part we actually write and deal with. Because we know how the underlying system works, we can design it without treating Kubernetes like a black box.\",{\"_287\":288,\"_41\":289,\"_290\":568,\"_47\":569},{},[570,571,572,573,574,575,576],\"Each step is independently idempotent. Each returns a result, either \\\"I'm done\\\" or \\\"requeue me in 30 seconds, I'm waiting on something,\\\" and the results merge. It reads almost exactly like the body of my shell loop. The difference is that \\\"observe the state\\\" grew from \",[\"SingleFetchClassInstance\",585],\" and \",[\"SingleFetchClassInstance\",581],\" into a fan-in across multiple systems, and \\\"take an action\\\" grew from \",[\"SingleFetchClassInstance\",577],\" into typed, conflict-aware API writes.\",{\"_287\":288,\"_41\":389,\"_290\":578,\"_47\":579},{},[580],\"ssh\",{\"_287\":288,\"_41\":389,\"_290\":582,\"_47\":583},{},[584],\"pg_isready\",{\"_287\":288,\"_41\":389,\"_290\":586,\"_47\":587},{},[549],{\"_287\":288,\"_41\":289,\"_290\":589,\"_47\":590},{},[591],\"Status is written last on purpose, because it's the measured variable: you record what's true after you've taken your actions and observed the result. Again, in our operators, it's not possible to write the status mid-reconcile.\",{\"_287\":288,\"_41\":527,\"_290\":593,\"_47\":594},{\"_530\":595,\"_532\":533,\"_534\":535,\"_389\":-7},[],\"\u003cpre class=\\\"shiki shiki-themes planetscale-light planetscale-dark\\\" style=\\\"--shiki-light:#2b2b2b;--shiki-dark:#e1e1e1;--shiki-light-bg:#ebebeb;--shiki-dark-bg:#1a1a1a\\\" tabindex=\\\"0\\\"\u003e\u003ccode\u003e\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003efunc\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e (\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003er \u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e*reconcileHandler\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e)\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#5E49AF;--shiki-dark:#B7A5FB\\\"\u003e reconcile\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e(\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003ectx\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e context\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003eContext\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e)\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e (\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003ereconcile\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003eResult\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e,\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e error\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e)\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e {\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e var\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e rb\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e results\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003eBuilder\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e rb\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#5E49AF;--shiki-dark:#B7A5FB\\\"\u003eMerge\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e(\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003er\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#5E49AF;--shiki-dark:#B7A5FB\\\"\u003ereconcileConfigMap\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e(\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003ectx\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e))\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#818181;--shiki-dark:#A1A1A1\\\"\u003e // push desired config\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e rb\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#5E49AF;--shiki-dark:#B7A5FB\\\"\u003eMerge\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e(\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003er\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#5E49AF;--shiki-dark:#B7A5FB\\\"\u003ereconcileDatabase\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e(\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003ectx\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e))\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#818181;--shiki-dark:#A1A1A1\\\"\u003e // reload/restart if params drifted\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e rb\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#5E49AF;--shiki-dark:#B7A5FB\\\"\u003eMerge\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e(\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003er\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#5E49AF;--shiki-dark:#B7A5FB\\\"\u003ereconcilePVC\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e(\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003ectx\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e))\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#818181;--shiki-dark:#A1A1A1\\\"\u003e // grow the disk if needed\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e rb\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#5E49AF;--shiki-dark:#B7A5FB\\\"\u003eMerge\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e(\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003er\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#5E49AF;--shiki-dark:#B7A5FB\\\"\u003ereconcilePod\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e(\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003ectx\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e))\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#818181;--shiki-dark:#A1A1A1\\\"\u003e // create/replace the Pod\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e rb\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#5E49AF;--shiki-dark:#B7A5FB\\\"\u003eMerge\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e(\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003er\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#5E49AF;--shiki-dark:#B7A5FB\\\"\u003ereconcileStatus\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e(\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003ectx\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e))\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#818181;--shiki-dark:#A1A1A1\\\"\u003e // always last: write what we observed\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e return\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e rb\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#5E49AF;--shiki-dark:#B7A5FB\\\"\u003eResult\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e()\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e}\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003c/span\u003e\u003c/code\u003e\u003c/pre\u003e\",{\"_287\":288,\"_41\":289,\"_290\":597,\"_47\":598},{},[599],\"Once the data snapshot exists, the reconcile is a sequence of small, idempotent steps, each comparing one slice of desired against observed and acting to close the gap:\",{\"_287\":288,\"_41\":289,\"_290\":601,\"_47\":602},{},[603,604],[\"SingleFetchClassInstance\",605],\" Reaching the database can fail while the Kubernetes reads succeed. That's not always an error that aborts the reconcile. It's a measured fact: \\\"Postgres is currently unreachable.\\\" That itself is something to record in status. A control loop that gives up entirely whenever one sensor is unavailable is a control loop that's down a lot. We degrade instead. Think of a car. If the rain sensor for the wipers is broken, the whole car doesn't stop. You can still drive, but you need to turn on a few things yourself.\",{\"_287\":288,\"_41\":305,\"_290\":606,\"_47\":607},{},[608],\"Partial failures are tolerated.\",{\"_287\":288,\"_41\":289,\"_290\":610,\"_47\":611},{},[612,613,614],\"For example, the database might be the leader when we check at the top and a replica by the time another helper checks again. Then you get decisions that are individually reasonable, but wrong together. We have a rule in the codebase against stashing state back onto this handler mid-reconcile to pass between steps, because it reintroduces exactly the inconsistency we gathered the snapshot to avoid. Making the \",[\"SingleFetchClassInstance\",615],\" immutable is one way to enforce that rule in the type system instead of relying on code review.\",{\"_287\":288,\"_41\":389,\"_290\":616,\"_47\":617},{},[618],\"reconcileHandler\",{\"_287\":288,\"_41\":289,\"_290\":620,\"_47\":621},{},[622,623],[\"SingleFetchClassInstance\",624],\" Every sub-decision in the reconcile reads from this one snapshot. We don't re-query the database in the middle of the loop, or read the disk usage again three functions deep. If we did, different parts of the same reconcile could see different versions of reality. This sounds like a small detail, but it changes the whole design.\",{\"_287\":288,\"_41\":305,\"_290\":625,\"_47\":626},{},[627],\"Gather once, at the top.\",{\"_287\":288,\"_41\":289,\"_290\":629,\"_47\":630},{},[631],\"A few things about this are deliberate, and only look obvious after you've been burned once or twice.\",{\"_287\":288,\"_41\":527,\"_290\":633,\"_47\":634},{\"_530\":635,\"_532\":533,\"_534\":535,\"_389\":-7},[],\"\u003cpre class=\\\"shiki shiki-themes planetscale-light planetscale-dark\\\" style=\\\"--shiki-light:#2b2b2b;--shiki-dark:#e1e1e1;--shiki-light-bg:#ebebeb;--shiki-dark-bg:#1a1a1a\\\" tabindex=\\\"0\\\"\u003e\u003ccode\u003e\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#818181;--shiki-dark:#A1A1A1\\\"\u003e// The observation snapshot: everything we know about this instance, right now.\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003etype\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e reconcileHandler\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e struct\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e {\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#818181;--shiki-dark:#A1A1A1\\\"\u003e // The object (.spec = setpoint, .status = measured).\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e instance\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e *v1\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003ePostgresInstance\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#818181;--shiki-dark:#A1A1A1\\\"\u003e // Kubernetes objects.\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e pod\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e *corev1\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003ePod\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e pvc\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e *corev1\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003ePersistentVolumeClaim\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e node\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e *corev1\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003eNode\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#818181;--shiki-dark:#A1A1A1\\\"\u003e // Database state.\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e dbState\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e DatabaseState\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#818181;--shiki-dark:#A1A1A1\\\"\u003e // What Postgres actually loaded, not what we last wrote.\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e effectiveConfig\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e map\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e[\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003estring\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e]\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003estring\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#818181;--shiki-dark:#A1A1A1\\\"\u003e // Collected out-of-band.\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e diskUsage\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e *resource\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003eQuantity\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#818181;--shiki-dark:#A1A1A1\\\"\u003e // The volume operation already in flight, if any.\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e storageOp\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e *StorageOperation\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e}\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003efunc\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e (\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003er \u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e*Reconciler\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e)\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#5E49AF;--shiki-dark:#B7A5FB\\\"\u003e newReconcileHandler\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e(\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e ctx\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e context\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003eContext\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e,\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e inst\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e *v1\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003ePostgresInstance\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e,\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e)\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e (\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e*reconcileHandler\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e,\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e error\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e)\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e {\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e h\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e :=\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e \u0026#x26;reconcileHandler\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e{\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003einstance\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e:\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e inst\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e}\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#818181;--shiki-dark:#A1A1A1\\\"\u003e // Cheap local reads.\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e h\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003epod\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e,\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e h\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003epvc\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e,\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e h\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003enode\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e =\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e r\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#5E49AF;--shiki-dark:#B7A5FB\\\"\u003efetchKubeObjects\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e(\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003ectx\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e,\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e inst\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e)\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#818181;--shiki-dark:#A1A1A1\\\"\u003e // Active database calls.\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e h\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003edbState\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e =\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e r\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#5E49AF;--shiki-dark:#B7A5FB\\\"\u003equeryDatabase\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e(\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003ectx\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e,\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e h\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003epod\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e)\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e h\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003eeffectiveConfig\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e =\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e r\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#5E49AF;--shiki-dark:#B7A5FB\\\"\u003ereadEffectiveConfig\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e(\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003ectx\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e,\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e h\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003epod\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e)\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#818181;--shiki-dark:#A1A1A1\\\"\u003e // Values from the collector/metric.\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e h\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003ediskUsage\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e =\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e r\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003ecollector\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#5E49AF;--shiki-dark:#B7A5FB\\\"\u003eUsage\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e(\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003eh\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003epvc\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e)\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e h\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003estorageOp\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e =\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e r\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003ecollector\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#5E49AF;--shiki-dark:#B7A5FB\\\"\u003eInFlightOp\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e(\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003eh\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003epvc\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e)\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e return\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e h\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e,\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#7D5903;--shiki-dark:#FED54A\\\"\u003e nil\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e}\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003c/span\u003e\u003c/code\u003e\u003c/pre\u003e\",{\"_287\":288,\"_41\":289,\"_290\":637,\"_47\":638},{},[639],\"In code, the snapshot is just a struct, and the reconcile's first move is to populate it. This is simplified, but faithful to the real shape:\",{\"_287\":288,\"_41\":289,\"_290\":641,\"_47\":642},{},[643],[\"SingleFetchClassInstance\",644],{\"_287\":288,\"_41\":407,\"_290\":645,\"_47\":646},{\"_410\":647,\"_412\":648,\"_414\":415,\"_416\":649,\"_418\":650,\"_420\":421},[],\"Observation fan-in for a reconciler\",1007,\"https://planetscale-images.imgix.net/assets/part2-observation-fan-in-DvyJkB0Y.png?auto=compress%2Cformat\",[651,652],{\"_424\":649,\"_426\":428},{\"_424\":653,\"_426\":427},\"https://planetscale-images.imgix.net/assets/part2-observation-fan-in-darkmode-KuH1rwY1.png?auto=compress%2Cformat\",{\"_287\":288,\"_41\":434,\"_290\":655,\"_47\":656},{},[657,658,659,660],[\"SingleFetchClassInstance\",688],[\"SingleFetchClassInstance\",679],[\"SingleFetchClassInstance\",670],[\"SingleFetchClassInstance\",661],{\"_287\":288,\"_41\":441,\"_290\":662,\"_47\":663},{},[664,665],[\"SingleFetchClassInstance\",666],\" Some signals are too expensive or too rate-limited to fetch on every reconcile: disk usage, or whether a volume operation we kicked off earlier is still in flight and where it sits in its cooldown window. A separate collector, often a background goroutine, gathers these on a slow cadence and keeps the last value per volume in memory. The reconcile reads that value instantly, without blocking on anything. Think of these as custom workqueues you implement.\",{\"_287\":288,\"_41\":305,\"_290\":667,\"_47\":668},{},[669],\"A background collector.\",{\"_287\":288,\"_41\":441,\"_290\":671,\"_47\":672},{},[673,674],[\"SingleFetchClassInstance\",675],\" Its role, whether it's healthy, how far behind its followers are, whether it's currently accepting writes. Some of this comes from the agents that sit next to the database and manage it; some we get by opening a connection and asking the database directly. These calls carry a tight timeout and are allowed to fail, more on that below.\",{\"_287\":288,\"_41\":305,\"_290\":676,\"_47\":677},{},[678],\"The database's own view of its health.\",{\"_287\":288,\"_41\":441,\"_290\":680,\"_47\":681},{},[682,683],[\"SingleFetchClassInstance\",684],\" Not what we last wrote down, but what the server has actually loaded, so we can compare the two and detect drift. Other entities can rewrite or reload the config on disk without us knowing, so the only honest source of truth is the running server itself, never our last write.\",{\"_287\":288,\"_41\":305,\"_290\":685,\"_47\":686},{},[687],\"The database's effective configuration.\",{\"_287\":288,\"_41\":441,\"_290\":689,\"_47\":690},{},[691,692],[\"SingleFetchClassInstance\",693],\": the Pod, its PVC, the PV behind it, the Node it's on, the ConfigMap holding its config. These are the cheap local reads, the stuff we already talked about.\",{\"_287\":288,\"_41\":305,\"_290\":694,\"_47\":695},{},[696],\"The Kubernetes cache\",{\"_287\":288,\"_41\":289,\"_290\":698,\"_47\":699},{},[700],\"The sources gathered at the top of every reconcile:\",{\"_287\":288,\"_41\":289,\"_290\":702,\"_47\":703},{},[704],\"When an operator I work on reconciles a single Postgres instance, the first thing it does, before it decides anything, is build a snapshot of reality from every source that knows something true about that instance. Not just Kubernetes. Kubernetes barely knows anything about whether Postgres is actually healthy.\",{\"_287\":288,\"_41\":289,\"_290\":706,\"_47\":707},{},[708],\"Up to here I've been a little vague about the \\\"measure\\\" step, because in the examples the measured state is just \\\"list the child Pods.\\\" But in a real database operator it's a lot more than that.\",{\"_287\":288,\"_41\":555,\"_290\":710,\"_47\":711},{\"_49\":70},[712],[\"SingleFetchClassInstance\",713],{\"_287\":288,\"_41\":335,\"_290\":714,\"_47\":715},{\"_338\":716},[71],\"#what-observe-actually-means-in-a-real-operator\",{\"_287\":288,\"_41\":289,\"_290\":718,\"_47\":719},{},[720,721,722],\"So we wrote \",[\"SingleFetchClassInstance\",723],\" and maybe thought we're good. Kubernetes turns that one line into several independent controllers that have never heard of each other, but still cooperate because they share the API server and watch each other's objects. I like this part a lot. Nobody calls a central orchestrator. Nobody passes a private message. The system heals itself.\",{\"_287\":288,\"_41\":389,\"_290\":724,\"_47\":725},{},[726],\"if ! pg_isready; then docker start; fi\",{\"_287\":288,\"_41\":289,\"_290\":728,\"_47\":729},{},[730,731,732,733,734,735,736,737,738],\"Control theory has a name for loops stacked like this: \",[\"SingleFetchClassInstance\",749],\". The output of an outer loop becomes the setpoint of an inner loop. A controller never writes its own \",[\"SingleFetchClassInstance\",746],\", but it writes \",[\"SingleFetchClassInstance\",742],\" objects' \",[\"SingleFetchClassInstance\",739],\" all the time. My operator writes the PVC's spec, and that spec is the setpoint the CSI controllers converge to. Each loop worries only about its own layer and trusts the loop below.\",{\"_287\":288,\"_41\":389,\"_290\":740,\"_47\":741},{},[392],{\"_287\":288,\"_41\":355,\"_290\":743,\"_47\":744},{},[745],\"other\",{\"_287\":288,\"_41\":389,\"_290\":747,\"_47\":748},{},[392],{\"_287\":288,\"_41\":335,\"_290\":750,\"_47\":751},{\"_338\":757},[752],[\"SingleFetchClassInstance\",753],{\"_287\":288,\"_41\":305,\"_290\":754,\"_47\":755},{},[756],\"cascade control\",\"https://en.wikipedia.org/wiki/Proportional%E2%80%93integral%E2%80%93derivative_controller#Cascade_control\",{\"_287\":288,\"_41\":289,\"_290\":759,\"_47\":760},{},[761],\"That's multiple feedback loops, each watching the layer below, each reacting to an edge and converging to its own level, chained together through the API server with nobody orchestrating the whole thing.\",{\"_287\":288,\"_41\":763,\"_290\":764,\"_47\":765},\"ol\",{},[766,767,768,769,770,771,772],[\"SingleFetchClassInstance\",797],[\"SingleFetchClassInstance\",793],[\"SingleFetchClassInstance\",789],[\"SingleFetchClassInstance\",785],[\"SingleFetchClassInstance\",781],[\"SingleFetchClassInstance\",777],[\"SingleFetchClassInstance\",773],{\"_287\":288,\"_41\":441,\"_290\":774,\"_47\":775},{},[776],\"The kubelet gets triggered because that's an edge for that node's kubelet and it starts the container.\",{\"_287\":288,\"_41\":441,\"_290\":778,\"_47\":779},{},[780],\"The scheduler detects the unscheduled Pod, assigns a node.\",{\"_287\":288,\"_41\":441,\"_290\":782,\"_47\":783},{},[784],\"The replacement is an unscheduled Pod, an edge for the scheduler.\",{\"_287\":288,\"_41\":441,\"_290\":786,\"_47\":787},{},[788],\"The parent reconciles, observes its children (level-triggered), sees one is missing and the count is below the setpoint, and creates a replacement.\",{\"_287\":288,\"_41\":441,\"_290\":790,\"_47\":791},{},[792],\"Through the ownership link, that edge becomes a reconcile request for the parent.\",{\"_287\":288,\"_41\":441,\"_290\":794,\"_47\":795},{},[796],\"The Pod's deletion is a watch event, an edge.\",{\"_287\":288,\"_41\":441,\"_290\":798,\"_47\":799},{},[800],\"A node dies and takes a Pod with it.\",{\"_287\":288,\"_41\":289,\"_290\":802,\"_47\":803},{},[804],\"Here is an example. Follow the loop:\",{\"_287\":288,\"_41\":289,\"_290\":806,\"_47\":807},{},[808,809,810,811,812],\"When a controller creates a child object (a Pod, a PVC), it stamps an \",[\"SingleFetchClassInstance\",817],\" on the child pointing back at the parent. That reference does two things. It sets up garbage collection: delete the parent, and Kubernetes can cascade the delete to its children. And it gives the controller a way to map child changes back to the parent: \\\"when any object I own changes, enqueue my parent for a reconcile.\\\" \",[\"SingleFetchClassInstance\",813],\" allows you to link controllers to each other and create chains. If done right, all your controllers and systems fit together.\",{\"_287\":288,\"_41\":389,\"_290\":814,\"_47\":815},{},[816],\"ownerReference\",{\"_287\":288,\"_41\":305,\"_290\":818,\"_47\":819},{},[816],{\"_287\":288,\"_41\":289,\"_290\":821,\"_47\":822},{},[823],\"This is the part I like most.\",{\"_287\":288,\"_41\":555,\"_290\":825,\"_47\":826},{\"_49\":73},[827],[\"SingleFetchClassInstance\",828],{\"_287\":288,\"_41\":335,\"_290\":829,\"_47\":830},{\"_338\":831},[74],\"#self-healing-by-design\",{\"_287\":288,\"_41\":289,\"_290\":833,\"_47\":834},{},[835,836,837,838,839],\"This is why the reconcile is \",[\"SingleFetchClassInstance\",844],\", and why that matters. Our shell script kept \\\"am I mid-failover?\\\" in a variable that died with the process. A Kubernetes controller keeps nothing important in memory. Every fact it needs is on an API object: the spec it's driving toward, the status it last observed, the conditions describing where things stand. Kill the controller, restart it on another node, and it picks up exactly where it left off, not because it saved its progress, but because there was never any in-memory progress to lose. The state lives in the cluster (API server, \",[\"SingleFetchClassInstance\",840],\" is what holds the state). The controller is just the loop that reads it.\",{\"_287\":288,\"_41\":389,\"_290\":841,\"_47\":842},{},[843],\"etcd\",{\"_287\":288,\"_41\":305,\"_290\":845,\"_47\":846},{},[847],\"stateless\",{\"_287\":288,\"_41\":289,\"_290\":849,\"_47\":850},{},[851],[\"SingleFetchClassInstance\",852],{\"_287\":288,\"_41\":407,\"_290\":853,\"_47\":854},{\"_410\":855,\"_412\":648,\"_414\":415,\"_416\":856,\"_418\":857,\"_420\":421},[],\"Spec and status as setpoint and measured variable\",\"https://planetscale-images.imgix.net/assets/part2-spec-status-CinHPict.png?auto=compress%2Cformat\",[858,859],{\"_424\":856,\"_426\":428},{\"_424\":860,\"_426\":427},\"https://planetscale-images.imgix.net/assets/part2-spec-status-darkmode-DI9X1oFB.png?auto=compress%2Cformat\",{\"_287\":288,\"_41\":289,\"_290\":862,\"_47\":863},{},[864,865,866],\"When \",[\"SingleFetchClassInstance\",867],\", the status you're looking at does not reflect the latest setpoint yet. That one comparison is how you tell \\\"converged\\\" from \\\"still working on it.\\\"\",{\"_287\":288,\"_41\":389,\"_290\":868,\"_47\":869},{},[870],\"observedGeneration \u003c generation\",{\"_287\":288,\"_41\":289,\"_290\":872,\"_47\":873},{},[874,875,876,877,878],\"That split is the whole declarative model in two fields. It comes with a piece of bookkeeping that's pure control theory: \",[\"SingleFetchClassInstance\",883],\" increments when desired state changes, and by convention the controller writes back \",[\"SingleFetchClassInstance\",879],\" to say \\\"the state I'm reporting reflects this version of your intent.\\\"\",{\"_287\":288,\"_41\":389,\"_290\":880,\"_47\":881},{},[882],\".status.observedGeneration\",{\"_287\":288,\"_41\":389,\"_290\":884,\"_47\":885},{},[886],\".metadata.generation\",{\"_287\":288,\"_41\":289,\"_290\":888,\"_47\":889},{},[890,891,892,893,894,895,896,897,898,899,900,901,902,903],[\"SingleFetchClassInstance\",927],\" is the \",[\"SingleFetchClassInstance\",923],\", the observed state. It's owned by the controller, written through a separate status subresource, and it's where you record what's actually true. The better the status, the better the controller can decide. A good \",[\"SingleFetchClassInstance\",920],\" field is what makes a controller pleasant to operate. The word \",[\"SingleFetchClassInstance\",916],\" comes from control theory; \",[\"SingleFetchClassInstance\",911],\" around 1960 to ask whether you can infer a system's internal state from its outputs. \",[\"SingleFetchClassInstance\",908],\" is also your response to any third-party system. If someone wants to learn the outcome of your actions, \",[\"SingleFetchClassInstance\",904],\" is the place to look at.\",{\"_287\":288,\"_41\":389,\"_290\":905,\"_47\":906},{},[907],\".status\",{\"_287\":288,\"_41\":389,\"_290\":909,\"_47\":910},{},[907],{\"_287\":288,\"_41\":335,\"_290\":912,\"_47\":913},{\"_338\":915},[914],\"Kalman coined it\",\"https://en.wikipedia.org/wiki/Observability\",{\"_287\":288,\"_41\":355,\"_290\":917,\"_47\":918},{},[919],\"observability\",{\"_287\":288,\"_41\":389,\"_290\":921,\"_47\":922},{},[907],{\"_287\":288,\"_41\":305,\"_290\":924,\"_47\":925},{},[926],\"measured variable\",{\"_287\":288,\"_41\":389,\"_290\":928,\"_47\":929},{},[907],{\"_287\":288,\"_41\":289,\"_290\":931,\"_47\":932},{},[933,891,934,935,936,937],[\"SingleFetchClassInstance\",945],[\"SingleFetchClassInstance\",941],\", the desired state. It's owned by whoever created the object (a human, or another controller), and the reconciler treats it as read-only intent. It's an anti-pattern to write to the \",[\"SingleFetchClassInstance\",938],\" from inside the controller. If you do it, stop reading, go and fix your codebase. There are only a handful of exceptions, but a controller should generally never set its own setpoint.\",{\"_287\":288,\"_41\":389,\"_290\":939,\"_47\":940},{},[392],{\"_287\":288,\"_41\":305,\"_290\":942,\"_47\":943},{},[944],\"setpoint\",{\"_287\":288,\"_41\":389,\"_290\":946,\"_47\":947},{},[392],{\"_287\":288,\"_41\":289,\"_290\":949,\"_47\":950},{},[951,952,572,953,954],\"Back to the control diagram. My script kept its setpoint in shell variables and its measured state in the output of \",[\"SingleFetchClassInstance\",959],[\"SingleFetchClassInstance\",955],\". Kubernetes gives both a permanent home, on the object itself.\",{\"_287\":288,\"_41\":389,\"_290\":956,\"_47\":957},{},[958],\"psql\",{\"_287\":288,\"_41\":389,\"_290\":960,\"_47\":961},{},[549],{\"_287\":288,\"_41\":555,\"_290\":963,\"_47\":964},{\"_49\":76},[965],[\"SingleFetchClassInstance\",966],{\"_287\":288,\"_41\":335,\"_290\":967,\"_47\":968},{\"_338\":969},[77],\"#setpoint-and-measured-variable-spec-and-status\",{\"_287\":288,\"_41\":289,\"_290\":971,\"_47\":972},{},[973],\"This is not only a Kubernetes issue. In any system with a read cache and a write-through path, read-after-write is not consistent unless you make it so. Most of the time the cache is exactly what you want: cheap, local, and eventually consistent. Eventual consistency is fine because the loop runs again. But the moment a decision would be destructive or non-idempotent if you acted on a stale read, you need to know which path you're on. Kubernetes solves many hard problems, but it also gives you a few new ones.\",{\"_287\":288,\"_41\":527,\"_290\":975,\"_47\":976},{\"_530\":977,\"_532\":533,\"_534\":535,\"_389\":-7},[],\"\u003cpre class=\\\"shiki shiki-themes planetscale-light planetscale-dark\\\" style=\\\"--shiki-light:#2b2b2b;--shiki-dark:#e1e1e1;--shiki-light-bg:#ebebeb;--shiki-dark-bg:#1a1a1a\\\" tabindex=\\\"0\\\"\u003e\u003ccode\u003e\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#818181;--shiki-dark:#A1A1A1\\\"\u003e// The cached client can be stale right after our own writes, which\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#818181;--shiki-dark:#A1A1A1\\\"\u003e// would make us miscount and create duplicates. For this one read,\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#818181;--shiki-dark:#A1A1A1\\\"\u003e// go direct to the API server instead of the cache. Slower,\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#818181;--shiki-dark:#A1A1A1\\\"\u003e// but consistent for this decision.\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003eerr\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e :=\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e r\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003eapiReader\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#5E49AF;--shiki-dark:#B7A5FB\\\"\u003eList\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e(\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003ectx\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e,\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e \u0026#x26;\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003einstances\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e,\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e client\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#5E49AF;--shiki-dark:#B7A5FB\\\"\u003eInNamespace\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e(\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003ens\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e),\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e labelSelector\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e)\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003c/span\u003e\u003c/code\u003e\u003c/pre\u003e\",{\"_287\":288,\"_41\":289,\"_290\":979,\"_47\":980},{},[981],\"The pragmatic one, which a lot of people use, is to bypass the cache for the reads where a stale view would cause a double-create or double-delete, and go straight to the API server:\",{\"_287\":288,\"_41\":289,\"_290\":983,\"_47\":984},{},[985,986,987,988,989],\"There are two ways out. The correct one is the \",[\"SingleFetchClassInstance\",995],\", the same trick the built-in ReplicaSet controller uses: you record that you expect to see N creations in memory, and you don't act again until the cache has caught up to your own writes. It works, but it's not easy to implement and it's a fair amount of machinery. Read more \",[\"SingleFetchClassInstance\",990],\".\",{\"_287\":288,\"_41\":335,\"_290\":991,\"_47\":992},{\"_338\":994},[993],\"on Ahmet's blog\",\"https://ahmet.im/blog/controller-pitfalls/\",{\"_287\":288,\"_41\":305,\"_290\":996,\"_47\":997},{},[998],\"expectations pattern\",{\"_287\":288,\"_41\":289,\"_290\":1000,\"_47\":1001},{},[1002],[\"SingleFetchClassInstance\",1003],{\"_287\":288,\"_41\":407,\"_290\":1004,\"_47\":1005},{\"_410\":1006,\"_412\":1007,\"_414\":415,\"_416\":1008,\"_418\":1009,\"_420\":421},[],\"Kubernetes cache reads versus API writes diagram\",1071,\"https://planetscale-images.imgix.net/assets/part2-cache-vs-api-4EYpZbTb.png?auto=compress%2Cformat\",[1010,1011],{\"_424\":1008,\"_426\":428},{\"_424\":1012,\"_426\":427},\"https://planetscale-images.imgix.net/assets/part2-cache-vs-api-darkmode-9mbMAIBr.png?auto=compress%2Cformat\",{\"_287\":288,\"_41\":289,\"_290\":1014,\"_47\":1015},{},[1016],\"Now an event fires again a second later, before the cache has caught up with the Pods you just created. You list from the cache, and the new Pods aren't there yet. Your code decides it still needs to create them, and you create duplicates. This is the classic stale-cache double-create, and it's nasty because it only shows up under timing you can't reproduce on your laptop.\",{\"_287\":288,\"_41\":527,\"_290\":1018,\"_47\":1019},{\"_530\":1020,\"_532\":533,\"_534\":535,\"_389\":-7},[],\"\u003cpre class=\\\"shiki shiki-themes planetscale-light planetscale-dark\\\" style=\\\"--shiki-light:#2b2b2b;--shiki-dark:#e1e1e1;--shiki-light-bg:#ebebeb;--shiki-dark-bg:#1a1a1a\\\" tabindex=\\\"0\\\"\u003e\u003ccode\u003e\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#818181;--shiki-dark:#A1A1A1\\\"\u003e// I want N replicas. I see fewer, so I create the missing ones.\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003eexisting\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e,\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e _\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e :=\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e r\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#5E49AF;--shiki-dark:#B7A5FB\\\"\u003elistChildPods\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e(\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003ectx\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e)\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#818181;--shiki-dark:#A1A1A1\\\"\u003e // reads the CACHE\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003efor\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e i\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e :=\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#5E49AF;--shiki-dark:#B7A5FB\\\"\u003e len\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e(\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003eexisting\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e);\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e i\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e \u0026#x3C;\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e desired\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e;\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e i\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e++\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e {\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e r\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003eclient\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#5E49AF;--shiki-dark:#B7A5FB\\\"\u003eCreate\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e(\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003ectx\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e,\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#5E49AF;--shiki-dark:#B7A5FB\\\"\u003e newPod\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e(\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003ei\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e))\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#818181;--shiki-dark:#A1A1A1\\\"\u003e // writes the API SERVER\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e}\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003c/span\u003e\u003c/code\u003e\u003c/pre\u003e\",{\"_287\":288,\"_41\":289,\"_290\":1022,\"_47\":1023},{},[1024],\"Here is another edge case. Picture this sequence inside a reconcile:\",{\"_287\":288,\"_41\":289,\"_290\":1026,\"_47\":1027},{},[1028,1029,1030,1031,1032],\"Most of the time, what you want is to drop the call and \",[\"SingleFetchClassInstance\",1037],\". In the next reconcile, the \",[\"SingleFetchClassInstance\",1033],\" will see the updated object, and your write will never happen. That's how everything self-converges.\",{\"_287\":288,\"_41\":389,\"_290\":1034,\"_47\":1035},{},[1036],\"GET\",{\"_287\":288,\"_41\":355,\"_290\":1038,\"_47\":1039},{},[1040],\"requeue\",{\"_287\":288,\"_41\":289,\"_290\":1042,\"_47\":1043},{},[1044],\"If you write a field of an object and then read the same object again from the cache, the reconciler might think it's not updated yet. You write again, and you get a Conflict error. Retrying with a fresh read can be fine, but blindly retrying against the same stale cached view just spins.\",{\"_287\":288,\"_41\":289,\"_290\":1046,\"_47\":1047},{},[1048],\"Because of that, a read can be stale. You need to be prepared for this.\",{\"_287\":288,\"_41\":289,\"_290\":1050,\"_47\":1051},{},[1052],\"In controller-runtime, the client you're handed reads from the informer's local cache. Cache reads are cheap, they don't touch the API server, and that's how a controller reconciles thousands of objects without falling over. But your writes go straight to the API server. The cache only learns about your write when the resulting watch event comes back around, a moment later.\",{\"_287\":288,\"_41\":289,\"_290\":1054,\"_47\":1055},{},[1056,1057],\"Second, and this is a detail that bites people a lot: \",[\"SingleFetchClassInstance\",1058],{\"_287\":288,\"_41\":305,\"_290\":1059,\"_47\":1060},{},[1061],\"your reads and your writes in Kubernetes don't go to the same place.\",{\"_287\":288,\"_41\":289,\"_290\":1063,\"_47\":1064},{},[1065],[\"SingleFetchClassInstance\",1066],{\"_287\":288,\"_41\":407,\"_290\":1067,\"_47\":1068},{\"_410\":1069,\"_412\":648,\"_414\":415,\"_416\":1070,\"_418\":1071,\"_420\":421},[],\"Informer, work queue, and controller diagram\",\"https://planetscale-images.imgix.net/assets/part2-informer-queue-yyOXLWWC.png?auto=compress%2Cformat\",[1072,1073],{\"_424\":1070,\"_426\":428},{\"_424\":1074,\"_426\":427},\"https://planetscale-images.imgix.net/assets/part2-informer-queue-darkmode-ANBXf_4R.png?auto=compress%2Cformat\",{\"_287\":288,\"_41\":434,\"_290\":1076,\"_47\":1077},{},[1078,1079,1080],[\"SingleFetchClassInstance\",1102],[\"SingleFetchClassInstance\",1085],[\"SingleFetchClassInstance\",1081],{\"_287\":288,\"_41\":441,\"_290\":1082,\"_47\":1083},{},[1084],\"It lets you run a pool of workers pulling keys in parallel, which is your fan-out. Events fan in from the watch, collapse in the queue, and fan out to the workers. This is something you need to tune. The higher you set the pool, the more pressure you put on the system: more writes, more API calls, more load on the provider, and more CPU usage in the operator.\",{\"_287\":288,\"_41\":441,\"_290\":1086,\"_47\":1087},{},[1088,1089,1090,1091,1092],\"It \",[\"SingleFetchClassInstance\",1098],\": an object that keeps erroring backs off exponentially instead of spinning. This is \",[\"SingleFetchClassInstance\",1093],\", the same reason a crash-looping container backs off instead of restarting hot.\",{\"_287\":288,\"_41\":335,\"_290\":1094,\"_47\":1095},{\"_338\":1097},[1096],\"damping\",\"https://en.wikipedia.org/wiki/Damping\",{\"_287\":288,\"_41\":355,\"_290\":1099,\"_47\":1100},{},[1101],\"rate-limits\",{\"_287\":288,\"_41\":441,\"_290\":1103,\"_47\":1104},{},[1088,1105,1106],[\"SingleFetchClassInstance\",1107],\": if the same object is updated five times before you get to it, you reconcile it once, against the latest state (level-triggered again).\",{\"_287\":288,\"_41\":355,\"_290\":1108,\"_47\":1109},{},[1110],\"coalesces\",{\"_287\":288,\"_41\":289,\"_290\":1112,\"_47\":1113},{},[1114,1115,1116],\"First, the informer turns each watch event into a key and puts it on a \",[\"SingleFetchClassInstance\",1117],\". The queue does a lot of work for you.\",{\"_287\":288,\"_41\":305,\"_290\":1118,\"_47\":1119},{},[1120],\"work queue\",{\"_287\":288,\"_41\":289,\"_290\":1122,\"_47\":1123},{},[1124,1125,1126,1127,1128],\"The answer is the \",[\"SingleFetchClassInstance\",1133],\". An informer opens a single watch against the API server for a given resource type, streams every add, update, and delete, and keeps a complete in-memory \",[\"SingleFetchClassInstance\",1129],\" of the objects we're interested in. Two things matter here:\",{\"_287\":288,\"_41\":305,\"_290\":1130,\"_47\":1131},{},[1132],\"cache\",{\"_287\":288,\"_41\":305,\"_290\":1134,\"_47\":1135},{},[1136],\"informer\",{\"_287\":288,\"_41\":289,\"_290\":1138,\"_47\":1139},{},[1140],\"So where do the edges come from? And what stops a controller from DDoSing the API server by listing everything every five seconds like my script did?\",{\"_287\":288,\"_41\":555,\"_290\":1142,\"_47\":1143},{\"_49\":79},[1144],[\"SingleFetchClassInstance\",1145],{\"_287\":288,\"_41\":335,\"_290\":1146,\"_47\":1147},{\"_338\":1148},[80],\"#informers-the-work-queue-and-a-cache\",{\"_287\":288,\"_41\":289,\"_290\":1150,\"_47\":1151},{},[1152,1153,1154],\"Our bash script stumbled into this property by accident, at least for the sake of the example. But the \",[\"SingleFetchClassInstance\",1155],\" framework gives it to you on purpose. It's why a Kubernetes controller can crash, get restarted ten minutes later, and converge correctly with no special recovery code. There is no recovery code. There is just the loop. The controller can reconstruct the world from scratch.\",{\"_287\":288,\"_41\":389,\"_290\":1156,\"_47\":1157},{},[1158],\"controller-runtime\",{\"_287\":288,\"_41\":289,\"_290\":1160,\"_47\":1161},{},[1162],[\"SingleFetchClassInstance\",1163],{\"_287\":288,\"_41\":407,\"_290\":1164,\"_47\":1165},{\"_410\":1166,\"_412\":1167,\"_414\":415,\"_416\":1168,\"_418\":1169,\"_420\":421},[],\"Edge-triggered versus level-triggered scaling\",927,\"https://planetscale-images.imgix.net/assets/part2-edge-vs-level-BUiD8D-7.png?auto=compress%2Cformat\",[1170,1171],{\"_424\":1168,\"_426\":428},{\"_424\":1172,\"_426\":427},\"https://planetscale-images.imgix.net/assets/part2-edge-vs-level-darkmode-3DNb92aw.png?auto=compress%2Cformat\",{\"_287\":288,\"_41\":289,\"_290\":1174,\"_47\":1175},{},[1176,1177,1178,1179,1180,1181,1182,1183,1184,1185,1186,1187,1188],\"Events (the edges) are only a hint that it's worth looking again. They tell you \",[\"SingleFetchClassInstance\",1209],\" to reconcile, never \",[\"SingleFetchClassInstance\",1205],\" to do. The reconcile itself is level-based: it reads the current state (e.g., \",[\"SingleFetchClassInstance\",1201],\") and drives toward the desired state (e.g., \",[\"SingleFetchClassInstance\",1197],\"), ignoring the triggering event completely. That's \",[\"SingleFetchClassInstance\",1193],\" \",[\"SingleFetchClassInstance\",1189],\" only gets a key. The framework makes it hard to write edge-triggered logic, on purpose. Edge-triggered logic is how you get a controller that's fragile and permanently wrong after its first hiccup.\",{\"_287\":288,\"_41\":389,\"_290\":1190,\"_47\":1191},{},[1192],\"Reconcile\",{\"_287\":288,\"_41\":355,\"_290\":1194,\"_47\":1195},{},[1196],\"why\",{\"_287\":288,\"_41\":389,\"_290\":1198,\"_47\":1199},{},[1200],\"create 2 replicas\",{\"_287\":288,\"_41\":389,\"_290\":1202,\"_47\":1203},{},[1204],\"replicas=3\",{\"_287\":288,\"_41\":355,\"_290\":1206,\"_47\":1207},{},[1208],\"what\",{\"_287\":288,\"_41\":355,\"_290\":1210,\"_47\":1211},{},[1212],\"when\",{\"_287\":288,\"_41\":289,\"_290\":1214,\"_47\":1215},{},[1216,1217,989],\"Kubernetes controllers combine both: \",[\"SingleFetchClassInstance\",1218],{\"_287\":288,\"_41\":305,\"_290\":1219,\"_47\":1220},{},[1221],\"edge-triggered notifications, level-triggered logic\",{\"_287\":288,\"_41\":289,\"_290\":1223,\"_47\":1224},{},[1225],\"If you miss the event, no one cares. In the next reconcile loop you'll catch it. If your app crashes, it comes back, reads again and detects that it did not increase it yet, increases it.\",{\"_287\":288,\"_41\":289,\"_290\":1227,\"_47\":1228},{},[1229,1230,1231,1232,1233],\"The level-triggered model fixes all of that. Remember, our shell script never asked \\\"what changed?\\\" It asked \\\"what \",[\"SingleFetchClassInstance\",1237],\" true right now?\\\", every five seconds, from scratch. Miss a loop, and the next one catches up. Run the loop twice, and you get the same result. The current state of the world is the only input that matters, and it's always available to read. So in the level-triggered case, our example above becomes this: you read \",[\"SingleFetchClassInstance\",1234],\", you check the current number of replicas, which is 1, and you increase by 2.\",{\"_287\":288,\"_41\":389,\"_290\":1235,\"_47\":1236},{},[1204],{\"_287\":288,\"_41\":355,\"_290\":1238,\"_47\":1239},{},[1240],\"is\",{\"_287\":288,\"_41\":289,\"_290\":1242,\"_47\":1243},{},[1244],\"In the first case, you won't be able to self-correct. In the second case, if your handler blindly applies the delta again, you'll end up with 5 replicas (you overshoot), instead of 3.\",{\"_287\":288,\"_41\":763,\"_290\":1246,\"_47\":1247},{},[1248,1249],[\"SingleFetchClassInstance\",1254],[\"SingleFetchClassInstance\",1250],{\"_287\":288,\"_41\":441,\"_290\":1251,\"_47\":1252},{},[1253],\"You receive it twice.\",{\"_287\":288,\"_41\":441,\"_290\":1255,\"_47\":1256},{},[1257],\"You miss the event (maybe the queue dropped it, or the consumer, your app, dropped it due to a crash or a full buffer).\",{\"_287\":288,\"_41\":289,\"_290\":1259,\"_47\":1260},{},[1261],\"Here is a very concrete example. Assume you have 1 replica, and you increase it to 3 replicas. Because you have only subscribed to changes, either:\",{\"_287\":288,\"_41\":289,\"_290\":1263,\"_47\":1264},{},[1265],\"The problem is that this is very fragile. In distributed systems, if one component is fragile, the fragility spreads to the rest of the system. Why is edge-triggering fragile? Say your controller is down for thirty seconds. It misses the events from those thirty seconds, and its view of the world is now permanently wrong. If two events arrive out of order, you process them out of order. If an event is delivered twice, you act twice. You're rebuilding your state from a stream of events, and you've inherited all of event sourcing's hard problems.\",{\"_287\":288,\"_41\":289,\"_290\":1267,\"_47\":1268},{},[1269],\"My first mental model of controllers, and probably yours at some point, was edge-triggered: listen to a stream of changes, and for each change, try to converge.\",{\"_287\":288,\"_41\":434,\"_290\":1271,\"_47\":1272},{},[1273,1274],[\"SingleFetchClassInstance\",1294],[\"SingleFetchClassInstance\",1275],{\"_287\":288,\"_41\":441,\"_290\":1276,\"_47\":1277},{},[1278,1279,1280,1281,1282,1283],[\"SingleFetchClassInstance\",1290],\": act on the current state, regardless of how you got there. \\\"The disk \",[\"SingleFetchClassInstance\",1287],\" at 85%.\\\" \\\"The Pod \",[\"SingleFetchClassInstance\",1284],\" missing.\\\" \\\"The number of replicas is 3.\\\"\",{\"_287\":288,\"_41\":355,\"_290\":1285,\"_47\":1286},{},[1240],{\"_287\":288,\"_41\":355,\"_290\":1288,\"_47\":1289},{},[1240],{\"_287\":288,\"_41\":305,\"_290\":1291,\"_47\":1292},{},[1293],\"Level-triggered\",{\"_287\":288,\"_41\":441,\"_290\":1295,\"_47\":1296},{},[1297,1298],[\"SingleFetchClassInstance\",1299],\": act on transitions, on events. \\\"The disk crossed 80%.\\\" \\\"The Pod was deleted.\\\" \\\"The number of replicas increased by 2.\\\"\",{\"_287\":288,\"_41\":305,\"_290\":1300,\"_47\":1301},{},[1302],\"Edge-triggered\",{\"_287\":288,\"_41\":289,\"_290\":1304,\"_47\":1305},{},[1306],\"There are two ways to build any closed feedback loop:\",{\"_287\":288,\"_41\":555,\"_290\":1308,\"_47\":1309},{\"_49\":82},[1310],[\"SingleFetchClassInstance\",1311],{\"_287\":288,\"_41\":335,\"_290\":1312,\"_47\":1313},{\"_338\":1314},[83],\"#edge-triggered-notifications-level-triggered-logic\",{\"_287\":288,\"_41\":289,\"_290\":1316,\"_47\":1317},{},[1318],\"Notice what is missing here: the function isn't told what changed. There is no diff. It isn't handed the old object and the new object. It isn't given an event type. It gets a key, a namespace and a name, and nothing else. It's minimal by design, because it has to work for many different controllers. The function's job is to fetch the object with that namespace/name, look at the world, and converge to the desired state.\",{\"_287\":288,\"_41\":527,\"_290\":1320,\"_47\":1321},{\"_530\":1322,\"_532\":533,\"_534\":535,\"_389\":-7},[],\"\u003cpre class=\\\"shiki shiki-themes planetscale-light planetscale-dark\\\" style=\\\"--shiki-light:#2b2b2b;--shiki-dark:#e1e1e1;--shiki-light-bg:#ebebeb;--shiki-dark-bg:#1a1a1a\\\" tabindex=\\\"0\\\"\u003e\u003ccode\u003e\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003efunc\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e (\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003er \u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e*Reconciler\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e)\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#5E49AF;--shiki-dark:#B7A5FB\\\"\u003e Reconcile\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e(\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003ectx\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e context\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003eContext\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e,\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e req\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e reconcile\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003eRequest\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e)\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e (\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003ereconcile\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e.\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003eResult\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e,\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e error\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e)\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e {\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#818181;--shiki-dark:#A1A1A1\\\"\u003e // req contains a namespace/name. That's it. That's the whole input.\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e}\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003c/span\u003e\u003c/code\u003e\u003c/pre\u003e\",{\"_287\":288,\"_41\":289,\"_290\":1324,\"_47\":1325},{},[1326,1327,1328,1329,1330],\"In Kubernetes, our watchdog script is a \",[\"SingleFetchClassInstance\",1335],\", and the standard way to write one in Go is a library called \",[\"SingleFetchClassInstance\",1331],\". At its heart, it's a function with a basic signature:\",{\"_287\":288,\"_41\":335,\"_290\":1332,\"_47\":1333},{\"_338\":1334},[1158],\"https://github.com/kubernetes-sigs/controller-runtime\",{\"_287\":288,\"_41\":305,\"_290\":1336,\"_47\":1337},{},[1338],\"controller\",{\"_287\":288,\"_41\":555,\"_290\":1340,\"_47\":1341},{\"_49\":85},[1342],[\"SingleFetchClassInstance\",1343],{\"_287\":288,\"_41\":335,\"_290\":1344,\"_47\":1345},{\"_338\":1346},[86],\"#the-for-loop-translated-to-kubernetes\",{\"_287\":288,\"_41\":289,\"_290\":1348,\"_47\":1349},{},[1350],\"That's the operator.\",{\"_287\":288,\"_41\":289,\"_290\":1352,\"_47\":1353},{},[1354,1355,1356],\"All of this so far is useful context, but the part we care about is \",[\"SingleFetchClassInstance\",1357],\" loop, because that's the one we get to write ourselves.\",{\"_287\":288,\"_41\":355,\"_290\":1358,\"_47\":1359},{},[1360],\"our watchdog\",{\"_287\":288,\"_41\":289,\"_290\":1362,\"_47\":1363},{},[1364],[\"SingleFetchClassInstance\",1365],{\"_287\":288,\"_41\":407,\"_290\":1366,\"_47\":1367},{\"_410\":1368,\"_412\":1369,\"_414\":415,\"_416\":1370,\"_418\":1371,\"_420\":421},[],\"Mapping manual operations to Kubernetes controllers\",883,\"https://planetscale-images.imgix.net/assets/part2-kubernetes-mapping-BqGHE4E0.png?auto=compress%2Cformat\",[1372,1373],{\"_424\":1370,\"_426\":428},{\"_424\":1374,\"_426\":427},\"https://planetscale-images.imgix.net/assets/part2-kubernetes-mapping-darkmode-CAz9s2jY.png?auto=compress%2Cformat\",{\"_287\":288,\"_41\":289,\"_290\":1376,\"_47\":1377},{},[1378,1379,1380],\"As you see, all the problems we solved with \",[\"SingleFetchClassInstance\",1381],\" and various scripts are replaced by Kubernetes components and drivers. And these are just a few of them:\",{\"_287\":288,\"_41\":389,\"_290\":1382,\"_47\":1383},{},[580],{\"_287\":288,\"_41\":289,\"_290\":1385,\"_47\":1386},{},[1387,1388,1389,1390,1391,1392,1393,1394],[\"SingleFetchClassInstance\",1408],\" The \",[\"SingleFetchClassInstance\",1404],\" problem is solved at a layer we no longer have to think about. A CNI plugin gives Pods their network identity; Cilium, for example, does this with eBPF instead of a pile of \",[\"SingleFetchClassInstance\",1400],\" rules. For stateful workloads, a StatefulSet plus a \",[\"SingleFetchClassInstance\",1395],\" gives each replica its own stable DNS name, which is exactly what a Postgres replica needs. The hard-coded IP that broke our cluster becomes a name that keeps working. DNS is only one tool here; other service discovery systems like etcd, ZooKeeper, and Consul solve similar problems.\",{\"_287\":288,\"_41\":335,\"_290\":1396,\"_47\":1397},{\"_338\":1399},[1398],\"headless Service\",\"https://kubernetes.io/docs/concepts/services-networking/service/#headless-services\",{\"_287\":288,\"_41\":389,\"_290\":1401,\"_47\":1402},{},[1403],\"iptables\",{\"_287\":288,\"_41\":389,\"_290\":1405,\"_47\":1406},{},[1407],\"/etc/hosts\",{\"_287\":288,\"_41\":305,\"_290\":1409,\"_47\":1410},{},[1411],\"Making them find each other: the CNI and Services.\",{\"_287\":288,\"_41\":289,\"_290\":1413,\"_47\":1414},{},[1415,1416,1417,572,1418,1419,1420,1421,1422,1423,1424,1425,1426,1427,1428,1429],[\"SingleFetchClassInstance\",1461],\" My multi-step \",[\"SingleFetchClassInstance\",1457],[\"SingleFetchClassInstance\",1453],\" script becomes a \",[\"SingleFetchClassInstance\",1449],\", which is a declarative request for storage. The \",[\"SingleFetchClassInstance\",1444],\" driver turns that request into a real volume. CSI itself is a set of controllers and sidecars: one provisions, one attaches, one resizes, and so on, while the kubelet calls the driver's node plugin to do the actual mount. It's a family of controllers. If there is a PVC but no disk behind it, one controller creates the disk. If the PVC size increases, another controller calls the provider API (e.g., AWS \",[\"SingleFetchClassInstance\",1440],\"). Again, I write intent, and a controller does the actual work. (note: I wrote one of the early production CSI drivers, \",[\"SingleFetchClassInstance\",1435],\", and a \",[\"SingleFetchClassInstance\",1430],\".)\",{\"_287\":288,\"_41\":335,\"_290\":1431,\"_47\":1432},{\"_338\":1434},[1433],\"long post about building one\",\"https://arslan.io/2018/06/21/how-to-write-a-container-storage-interface-csi-plugin/\",{\"_287\":288,\"_41\":335,\"_290\":1436,\"_47\":1437},{\"_338\":1439},[1438],\"csi-digitalocean\",\"https://github.com/digitalocean/csi-digitalocean\",{\"_287\":288,\"_41\":389,\"_290\":1441,\"_47\":1442},{},[1443],\"ModifyVolume\",{\"_287\":288,\"_41\":335,\"_290\":1445,\"_47\":1446},{\"_338\":1448},[1447],\"Container Storage Interface (CSI)\",\"https://github.com/container-storage-interface/spec/blob/master/spec.md\",{\"_287\":288,\"_41\":389,\"_290\":1450,\"_47\":1451},{},[1452],\"PersistentVolumeClaim\",{\"_287\":288,\"_41\":389,\"_290\":1454,\"_47\":1455},{},[1456],\"mount\",{\"_287\":288,\"_41\":389,\"_290\":1458,\"_47\":1459},{},[1460],\"mkfs\",{\"_287\":288,\"_41\":305,\"_290\":1462,\"_47\":1463},{},[1464],\"Attaching the disk: CSI and the PV/PVC sync.\",{\"_287\":288,\"_41\":289,\"_290\":1466,\"_47\":1467},{},[1468,1469,1470,1471,1472,1473],[\"SingleFetchClassInstance\",1482],\" Remember me choosing \",[\"SingleFetchClassInstance\",1478],\"? That's the scheduler's whole reason to exist. It watches for Pods with no node assigned, filters out the nodes that can't work, scores the rest, and writes the decision to one field: \",[\"SingleFetchClassInstance\",1474],\". The scheduler doesn't start the container; it records the placement and lets the kubelet pick it up. You will realize that most things in Kubernetes are decoupled like this.\",{\"_287\":288,\"_41\":389,\"_290\":1475,\"_47\":1476},{},[1477],\"pod.Spec.NodeName\",{\"_287\":288,\"_41\":389,\"_290\":1479,\"_47\":1480},{},[1481],\"node-07\",{\"_287\":288,\"_41\":305,\"_290\":1483,\"_47\":1484},{},[1485],\"Picking a node: the scheduler.\",{\"_287\":288,\"_41\":289,\"_290\":1487,\"_47\":1488},{},[1489,1490,1491,1492],[\"SingleFetchClassInstance\",1496],\" First, a quick definition: a Pod is the smallest thing Kubernetes runs, one or more containers scheduled together on a node and sharing its network. For us it's the Postgres container. On every node runs an agent called the kubelet. Its desired state is the set of Pods assigned to its node, which it learns from the API server. Its observed state is the set of containers actually running, which it gets from the container runtime. When they differ, it starts the missing container, kills the extra one, or restarts the crashed one. My \",[\"SingleFetchClassInstance\",1493],\" is the kubelet's job, just done properly. The kubelet doesn't shell into anything; it talks to containerd over a gRPC socket, which talks to runc.\",{\"_287\":288,\"_41\":389,\"_290\":1494,\"_47\":1495},{},[726],{\"_287\":288,\"_41\":305,\"_290\":1497,\"_47\":1498},{},[1499],\"Spinning up the container: the kubelet.\",{\"_287\":288,\"_41\":289,\"_290\":1501,\"_47\":1502},{},[1503],\"Let's go through some of the pieces we built by hand before the watchdog loop. You already know these components by name. What you might not have noticed is that they also work like controllers.\",{\"_287\":288,\"_41\":555,\"_290\":1505,\"_47\":1506},{\"_49\":88},[1507],[\"SingleFetchClassInstance\",1508],{\"_287\":288,\"_41\":335,\"_290\":1509,\"_47\":1510},{\"_338\":1511},[89],\"#the-other-loops\",{\"_287\":288,\"_41\":289,\"_290\":1513,\"_47\":1514},{},[1515],\"Now we can map what we hand-rolled in Part 1 to Kubernetes. Almost all of it already exists there. The operator is the part we care about.\",{\"_287\":288,\"_41\":330,\"_290\":1517,\"_47\":1518},{\"_49\":55},[1519],[\"SingleFetchClassInstance\",1520],{\"_287\":288,\"_41\":335,\"_290\":1521,\"_47\":1522},{\"_338\":1523},[56],\"#part-2-how-we-reinvented-kubernetes\",{\"_287\":288,\"_41\":341,\"_290\":1525,\"_47\":1526},{},[],{\"_287\":288,\"_41\":289,\"_290\":1528,\"_47\":1529},{},[1530],\"What if the script also fails? Who runs it then? We could keep hardening this script, but look at where it goes: we would need a real store for the desired state, watches instead of polling, a work queue, retries, leader election. We would be rebuilding Kubernetes. The real platform already exists, and it's Kubernetes.\",{\"_287\":288,\"_41\":434,\"_290\":1532,\"_47\":1533},{},[1534,1535,1536,1537,1538],[\"SingleFetchClassInstance\",1560],[\"SingleFetchClassInstance\",1556],[\"SingleFetchClassInstance\",1552],[\"SingleFetchClassInstance\",1543],[\"SingleFetchClassInstance\",1539],{\"_287\":288,\"_41\":441,\"_290\":1540,\"_47\":1541},{},[1542],\"And the moment I want a second kind of resource, a connection pooler, a backup job, a read replica in another region, I'm copy-pasting this whole structure.\",{\"_287\":288,\"_41\":441,\"_290\":1544,\"_47\":1545},{},[1546,1547,1548],\"It has no idea what to do when the \",[\"SingleFetchClassInstance\",1549],\" itself times out.\",{\"_287\":288,\"_41\":389,\"_290\":1550,\"_47\":1551},{},[580],{\"_287\":288,\"_41\":441,\"_290\":1553,\"_47\":1554},{},[1555],\"It polls every node every five seconds whether anything changed or not, which is fine for three nodes, but too expensive for three thousand nodes.\",{\"_287\":288,\"_41\":441,\"_290\":1557,\"_47\":1558},{},[1559],\"It keeps its only real state, \\\"am I mid-failover?\\\", in a shell variable that could die with the process.\",{\"_287\":288,\"_41\":441,\"_290\":1561,\"_47\":1562},{},[1563],\"It has no concurrency control, so two copies of the script can race each other. Imagine both deciding to promote a different replica.\",{\"_287\":288,\"_41\":289,\"_290\":1565,\"_47\":1566},{},[1567],\"A Bash loop is not a production control plane. Just to name a few issues with it:\",{\"_287\":288,\"_41\":289,\"_290\":1569,\"_47\":1570},{},[1571,1572,1573,1574,1575],\"That also gives us a nice way to understand \",[\"SingleFetchClassInstance\",1579],\" control. My very first attempt, \",[\"SingleFetchClassInstance\",1576],\" in, run the command, and walk away, was open-loop: fire an action and assume it worked. The Bash script is closed-loop because it keeps feeding the measured state back into the next decision.\",{\"_287\":288,\"_41\":389,\"_290\":1577,\"_47\":1578},{},[580],{\"_287\":288,\"_41\":305,\"_290\":1580,\"_47\":1581},{},[1582],\"open-loop\",{\"_287\":288,\"_41\":434,\"_290\":1584,\"_47\":1585},{},[1586,1587,1588,1589,1590,1591],[\"SingleFetchClassInstance\",1666],[\"SingleFetchClassInstance\",1643],[\"SingleFetchClassInstance\",1634],[\"SingleFetchClassInstance\",1620],[\"SingleFetchClassInstance\",1601],[\"SingleFetchClassInstance\",1592],{\"_287\":288,\"_41\":441,\"_290\":1593,\"_47\":1594},{},[460,1595,1596],[\"SingleFetchClassInstance\",1597],\" is the system being controlled, Postgres and its disk.\",{\"_287\":288,\"_41\":305,\"_290\":1598,\"_47\":1599},{},[1600],\"plant\",{\"_287\":288,\"_41\":441,\"_290\":1602,\"_47\":1603},{},[460,1604,1605,1606,1607,1608,989],[\"SingleFetchClassInstance\",1616],\" is what carries out the action: \",[\"SingleFetchClassInstance\",1613],\" plus \",[\"SingleFetchClassInstance\",1609],{\"_287\":288,\"_41\":389,\"_290\":1610,\"_47\":1611},{},[1612],\"docker start\",{\"_287\":288,\"_41\":389,\"_290\":1614,\"_47\":1615},{},[580],{\"_287\":288,\"_41\":305,\"_290\":1617,\"_47\":1618},{},[1619],\"actuator\",{\"_287\":288,\"_41\":441,\"_290\":1621,\"_47\":1622},{},[460,1623,1624,1625,1626],[\"SingleFetchClassInstance\",1631],\" is the body of the loop, the \",[\"SingleFetchClassInstance\",1627],\" statements that decide what to do. It is not the whole script.\",{\"_287\":288,\"_41\":389,\"_290\":1628,\"_47\":1629},{},[1630],\"if\",{\"_287\":288,\"_41\":305,\"_290\":1632,\"_47\":1633},{},[1338],{\"_287\":288,\"_41\":441,\"_290\":1635,\"_47\":1636},{},[460,1637,1638],[\"SingleFetchClassInstance\",1639],\" (e) is the difference between them.\",{\"_287\":288,\"_41\":305,\"_290\":1640,\"_47\":1641},{},[1642],\"error\",{\"_287\":288,\"_41\":441,\"_290\":1644,\"_47\":1645},{},[460,1646,1647,1648,1649,1650,1649,1651,989],[\"SingleFetchClassInstance\",1662],\" is what I observe: \",[\"SingleFetchClassInstance\",1659],\", \",[\"SingleFetchClassInstance\",1656],[\"SingleFetchClassInstance\",1652],{\"_287\":288,\"_41\":389,\"_290\":1653,\"_47\":1654},{},[1655],\"show max_connections\",{\"_287\":288,\"_41\":389,\"_290\":1657,\"_47\":1658},{},[549],{\"_287\":288,\"_41\":389,\"_290\":1660,\"_47\":1661},{},[584],{\"_287\":288,\"_41\":305,\"_290\":1663,\"_47\":1664},{},[1665],\"measured output\",{\"_287\":288,\"_41\":441,\"_290\":1667,\"_47\":1668},{},[460,1669,1670],[\"SingleFetchClassInstance\",1671],\" is my desired state, the variables at the top of the script (disk size, max_connections and so on).\",{\"_287\":288,\"_41\":305,\"_290\":1672,\"_47\":1673},{},[944],{\"_287\":288,\"_41\":289,\"_290\":1675,\"_47\":1676},{},[1677,1678,1679],\"Here is how the vocabulary from \",[\"SingleFetchClassInstance\",1680],\" maps cleanly onto my shell script:\",{\"_287\":288,\"_41\":335,\"_290\":1681,\"_47\":1682},{\"_338\":387},[1683],\"control theory\",{\"_287\":288,\"_41\":289,\"_290\":1685,\"_47\":1686},{},[1687],[\"SingleFetchClassInstance\",1688],{\"_287\":288,\"_41\":407,\"_290\":1689,\"_47\":1690},{\"_410\":1691,\"_412\":1692,\"_414\":415,\"_416\":1693,\"_418\":1694,\"_420\":421},[],\"Closed feedback loop diagram\",564,\"https://planetscale-images.imgix.net/assets/part1-feedback-loop-Z0q4qAPX.png?auto=compress%2Cformat\",[1695,1696],{\"_424\":1693,\"_426\":428},{\"_424\":1697,\"_426\":427},\"https://planetscale-images.imgix.net/assets/part1-feedback-loop-darkmode-BeMyZLwB.png?auto=compress%2Cformat\",{\"_287\":288,\"_41\":289,\"_290\":1699,\"_47\":1700},{},[1701,1702,1703],\"The nice part is that the same loop works for different problems. It can restart a dead process, grow a disk, or push \",[\"SingleFetchClassInstance\",1704],\". The action changes, but the shape stays the same: read what I want, observe what I have, compare them, act, repeat. If I draw the same thing as a block diagram, with the control theory names added, it would look like this:\",{\"_287\":288,\"_41\":389,\"_290\":1705,\"_47\":1706},{},[1707],\"max_connections = 500\",{\"_287\":288,\"_41\":289,\"_290\":1709,\"_47\":1710},{},[1711,1712,1713,1714,1715],\"That's a \",[\"SingleFetchClassInstance\",1719],\". The word \\\"closed\\\" matters. It means the output of the system is fed back into the next decision. I don't run \",[\"SingleFetchClassInstance\",1716],\" and assume the database is fine. I check the database again. If it is still wrong, I act again. If it is already correct, I do nothing.\",{\"_287\":288,\"_41\":389,\"_290\":1717,\"_47\":1718},{},[1612],{\"_287\":288,\"_41\":305,\"_290\":1720,\"_47\":1721},{},[1722],\"closed feedback loop\",{\"_287\":288,\"_41\":289,\"_290\":1724,\"_47\":1725},{},[1726,1727,1728],\"I started with a desired state that was written down in one place: three instances, this disk size, \",[\"SingleFetchClassInstance\",1729],\". Every few seconds I observe the actual state of the system. I compute the difference. I take whatever action closes that difference. Then I do it again, forever.\",{\"_287\":288,\"_41\":389,\"_290\":1730,\"_47\":1731},{},[1707],{\"_287\":288,\"_41\":555,\"_290\":1733,\"_47\":1734},{\"_49\":102},[1735],[\"SingleFetchClassInstance\",1736],{\"_287\":288,\"_41\":335,\"_290\":1737,\"_47\":1738},{\"_338\":1739},[103],\"#what-we-actually-built\",{\"_287\":288,\"_41\":289,\"_290\":1741,\"_47\":1742},{},[1743,1744,1745,1746,1747,1748,1749],\"This is the same idea as before. I read what I \",[\"SingleFetchClassInstance\",1759],\" (a variable). Observe what I \",[\"SingleFetchClassInstance\",1755],\" (a query). If they differ, I take an action to close the difference. Again, I don't track whether I changed it last time. All I do is compare and \",[\"SingleFetchClassInstance\",1750],\", every loop.\",{\"_287\":288,\"_41\":335,\"_290\":1751,\"_47\":1752},{\"_338\":1754},[1753],\"converge\",\"https://dictionary.cambridge.org/dictionary/english/converge\",{\"_287\":288,\"_41\":355,\"_290\":1756,\"_47\":1757},{},[1758],\"have\",{\"_287\":288,\"_41\":355,\"_290\":1760,\"_47\":1761},{},[1762],\"want\",{\"_287\":288,\"_41\":527,\"_290\":1764,\"_47\":1765},{\"_530\":1766,\"_532\":1767,\"_534\":535,\"_389\":-7},[],\"\u003cpre class=\\\"shiki shiki-themes planetscale-light planetscale-dark\\\" style=\\\"--shiki-light:#2b2b2b;--shiki-dark:#e1e1e1;--shiki-light-bg:#ebebeb;--shiki-dark-bg:#1a1a1a\\\" tabindex=\\\"0\\\"\u003e\u003ccode\u003e\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003eWANT_MAX_CONNECTIONS\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e=\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#414141;--shiki-dark:#C1C1C1\\\"\u003e500\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003efor\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e node\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e in\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#414141;--shiki-dark:#C1C1C1\\\"\u003e node-07\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#414141;--shiki-dark:#C1C1C1\\\"\u003e node-12\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#414141;--shiki-dark:#C1C1C1\\\"\u003e node-19\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e;\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e do\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e have\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e=\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e$(\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003essh\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#13862E;--shiki-dark:#75DB8C\\\"\u003e \\\"\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e$node\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#13862E;--shiki-dark:#75DB8C\\\"\u003e\\\"\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#13862E;--shiki-dark:#75DB8C\\\"\u003e \\\"psql -tAc 'show max_connections'\\\"\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e)\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e if\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e [\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#13862E;--shiki-dark:#75DB8C\\\"\u003e \\\"\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e$have\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#13862E;--shiki-dark:#75DB8C\\\"\u003e\\\"\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e !=\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#13862E;--shiki-dark:#75DB8C\\\"\u003e \\\"\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e$WANT_MAX_CONNECTIONS\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#13862E;--shiki-dark:#75DB8C\\\"\u003e\\\"\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e ];\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e then\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e ssh\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#13862E;--shiki-dark:#75DB8C\\\"\u003e \\\"\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e$node\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#13862E;--shiki-dark:#75DB8C\\\"\u003e\\\"\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#13862E;--shiki-dark:#75DB8C\\\"\u003e \\\"sed -i 's/^max_connections.*/max_connections = \u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e$WANT_MAX_CONNECTIONS\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#13862E;--shiki-dark:#75DB8C\\\"\u003e/' /var/lib/pg-data/postgresql.conf\\\"\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e ssh\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#13862E;--shiki-dark:#75DB8C\\\"\u003e \\\"\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e$node\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#13862E;--shiki-dark:#75DB8C\\\"\u003e\\\"\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#13862E;--shiki-dark:#75DB8C\\\"\u003e \\\"docker restart pg\\\"\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e fi\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003edone\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003c/span\u003e\u003c/code\u003e\u003c/pre\u003e\",\"bash\",{\"_287\":288,\"_41\":289,\"_290\":1769,\"_47\":1770},{},[1771],\"Because I know that ssh'ing into the nodes manually isn't a thing I want anymore, I do the same thing we did previously: I write the desired value down in one place and teach the loop to enforce it.\",{\"_287\":288,\"_41\":289,\"_290\":1773,\"_47\":1774},{},[1775,1776,1777,1778,1779,1780,1781],\"Let's make things a little more complex. I need to raise \",[\"SingleFetchClassInstance\",1789],\" from 100 to 500. This one is not a reload-only change. PostgreSQL says it can only be set at server start, so the manual version is to \",[\"SingleFetchClassInstance\",1786],\" into each box, edit \",[\"SingleFetchClassInstance\",1782],\", restart Postgres, and check that it took on all three.\",{\"_287\":288,\"_41\":389,\"_290\":1783,\"_47\":1784},{},[1785],\"postgresql.conf\",{\"_287\":288,\"_41\":389,\"_290\":1787,\"_47\":1788},{},[580],{\"_287\":288,\"_41\":335,\"_290\":1790,\"_47\":1791},{\"_338\":1797},[1792],[\"SingleFetchClassInstance\",1793],{\"_287\":288,\"_41\":389,\"_290\":1794,\"_47\":1795},{},[1796],\"max_connections\",\"https://www.postgresql.org/docs/current/runtime-config-connection.html#GUC-MAX-CONNECTIONS\",{\"_287\":288,\"_41\":555,\"_290\":1799,\"_47\":1800},{\"_49\":105},[1801],[\"SingleFetchClassInstance\",1802],{\"_287\":288,\"_41\":335,\"_290\":1803,\"_47\":1804},{\"_338\":1805},[106],\"#changing-a-parameter\",{\"_287\":288,\"_41\":289,\"_290\":1807,\"_47\":1808},{},[1809,1810,989],\"If a process is down, start it. If a disk is filling, grow it. Run the loop once or run it a thousand times and the result is the same, because each action is conditional on the current state. The script is \",[\"SingleFetchClassInstance\",1811],{\"_287\":288,\"_41\":355,\"_290\":1812,\"_47\":1813},{},[1814],\"idempotent\",{\"_287\":288,\"_41\":289,\"_290\":1816,\"_47\":1817},{},[1818,1819,1820],\"It's written in Bash, and probably has tons of bugs. You notice something here? The loop doesn't care \",[\"SingleFetchClassInstance\",1821],\" the database got into a bad state. Every five seconds it looks at the current state of the world and asks this question: does reality match what I want?\",{\"_287\":288,\"_41\":355,\"_290\":1822,\"_47\":1823},{},[1824],\"how\",{\"_287\":288,\"_41\":527,\"_290\":1826,\"_47\":1827},{\"_530\":1828,\"_532\":1767,\"_534\":535,\"_389\":-7},[],\"\u003cpre class=\\\"shiki shiki-themes planetscale-light planetscale-dark\\\" style=\\\"--shiki-light:#2b2b2b;--shiki-dark:#e1e1e1;--shiki-light-bg:#ebebeb;--shiki-dark-bg:#1a1a1a\\\" tabindex=\\\"0\\\"\u003e\u003ccode\u003e\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003ewhile\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#0B6EC5;--shiki-dark:#73C7F9\\\"\u003e true\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e;\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e do\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e for\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e node\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e in\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#414141;--shiki-dark:#C1C1C1\\\"\u003e node-07\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#414141;--shiki-dark:#C1C1C1\\\"\u003e node-12\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#414141;--shiki-dark:#C1C1C1\\\"\u003e node-19\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e;\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e do\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e if\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e !\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e ssh\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#13862E;--shiki-dark:#75DB8C\\\"\u003e \\\"\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e$node\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#13862E;--shiki-dark:#75DB8C\\\"\u003e\\\"\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#13862E;--shiki-dark:#75DB8C\\\"\u003e 'pg_isready -q'\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e;\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e then\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e ssh\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#13862E;--shiki-dark:#75DB8C\\\"\u003e \\\"\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e$node\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#13862E;--shiki-dark:#75DB8C\\\"\u003e\\\"\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#13862E;--shiki-dark:#75DB8C\\\"\u003e 'docker start pg'\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#818181;--shiki-dark:#A1A1A1\\\"\u003e # it died, bring it back\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e fi\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e usage\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e=\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e$(\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003essh\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#13862E;--shiki-dark:#75DB8C\\\"\u003e \\\"\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e$node\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#13862E;--shiki-dark:#75DB8C\\\"\u003e\\\"\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#13862E;--shiki-dark:#75DB8C\\\"\u003e \\\"df --output=pcent /var/lib/pg-data | tail -1 | tr -dc 0-9\\\"\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e)\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e if\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e [\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#13862E;--shiki-dark:#75DB8C\\\"\u003e \\\"\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e$usage\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#13862E;--shiki-dark:#75DB8C\\\"\u003e\\\"\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e -gt\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#D92038;--shiki-dark:#FF7082\\\"\u003e 80\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#616161;--shiki-dark:#C1C1C1\\\"\u003e ];\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e then\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e grow_volume\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#13862E;--shiki-dark:#75DB8C\\\"\u003e \\\"\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#A78103;--shiki-dark:#F2B600\\\"\u003e$node\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#13862E;--shiki-dark:#75DB8C\\\"\u003e\\\"\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#818181;--shiki-dark:#A1A1A1\\\"\u003e # disk filling, make it bigger\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e fi\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e done\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003e sleep\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#D92038;--shiki-dark:#FF7082\\\"\u003e 5\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003edone\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003c/span\u003e\u003c/code\u003e\u003c/pre\u003e\",{\"_287\":288,\"_41\":289,\"_290\":1830,\"_47\":1831},{},[1832],\"Let's assume we've set up a simple uptime monitor and we're going to get paged for all these cases. To avoid getting paged at night, we do the sensible thing: write a script. So we decide to write a loop that wakes up every few seconds, looks at each node, and fixes whatever's wrong.\",{\"_287\":288,\"_41\":434,\"_290\":1834,\"_47\":1835},{},[1836,1837,1838,1839],[\"SingleFetchClassInstance\",1852],[\"SingleFetchClassInstance\",1848],[\"SingleFetchClassInstance\",1844],[\"SingleFetchClassInstance\",1840],{\"_287\":288,\"_41\":441,\"_290\":1841,\"_47\":1842},{},[1843],\"A config I changed on two nodes but forgot on the third one. They are now out of sync.\",{\"_287\":288,\"_41\":441,\"_290\":1845,\"_47\":1846},{},[1847],\"The primary fails and a replica has to be promoted.\",{\"_287\":288,\"_41\":441,\"_290\":1849,\"_47\":1850},{},[1851],\"A disk gets full.\",{\"_287\":288,\"_41\":441,\"_290\":1853,\"_47\":1854},{},[1855],\"A replica process dies and doesn't come back.\",{\"_287\":288,\"_41\":289,\"_290\":1857,\"_47\":1858},{},[1859],\"Now, this is where we start thinking about how to solve these issues. Everything described so far can break, and will continue to break even if I fix it:\",{\"_287\":288,\"_41\":555,\"_290\":1861,\"_47\":1862},{\"_49\":108},[1863],[\"SingleFetchClassInstance\",1864],{\"_287\":288,\"_41\":335,\"_290\":1865,\"_47\":1866},{\"_338\":1867},[109],\"#the-watchdog-script\",{\"_287\":288,\"_41\":289,\"_290\":1869,\"_47\":1870},{},[1871],[\"SingleFetchClassInstance\",1872],{\"_287\":288,\"_41\":407,\"_290\":1873,\"_47\":1874},{\"_410\":1875,\"_412\":648,\"_414\":415,\"_416\":1876,\"_418\":1877,\"_420\":421},[],\"Manual Postgres cluster diagram\",\"https://planetscale-images.imgix.net/assets/part1-manual-cluster-C6iQCdoL.png?auto=compress%2Cformat\",[1878,1879],{\"_424\":1876,\"_426\":428},{\"_424\":1880,\"_426\":427},\"https://planetscale-images.imgix.net/assets/part1-manual-cluster-darkmode-VBZiV-Rl.png?auto=compress%2Cformat\",{\"_287\":288,\"_41\":289,\"_290\":1882,\"_47\":1883},{},[1884],\"But we still have a problem: the first time the primary is recreated with a different IP, the whole cluster falls apart.\",{\"_287\":288,\"_41\":527,\"_290\":1886,\"_47\":1887},{\"_530\":1888,\"_532\":1889,\"_534\":535,\"_389\":-7},[],\"\u003cpre class=\\\"shiki shiki-themes planetscale-light planetscale-dark\\\" style=\\\"--shiki-light:#2b2b2b;--shiki-dark:#e1e1e1;--shiki-light-bg:#ebebeb;--shiki-dark-bg:#1a1a1a\\\" tabindex=\\\"0\\\"\u003e\u003ccode\u003e\u003cspan class=\\\"line\\\"\u003e\u003cspan\u003e# on each replica's postgresql.auto.conf, until the primary is recreated with a new IP\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan\u003eprimary_conninfo = 'host=10.4.7.21 port=5432 user=replicator ...'\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan\u003e\u003c/span\u003e\u003c/span\u003e\u003c/code\u003e\u003c/pre\u003e\",\"conf\",{\"_287\":288,\"_41\":289,\"_290\":1891,\"_47\":1892},{},[1893,1894,1895,1896,1897,1898,1899],\"The first thing I do is hard-code the IPs. I write \",[\"SingleFetchClassInstance\",1907],\"'s address into the replicas' config, I list the replicas' addresses in the primary's \",[\"SingleFetchClassInstance\",1903],\", and I keep a small \",[\"SingleFetchClassInstance\",1900],\" table and save it somewhere.\",{\"_287\":288,\"_41\":389,\"_290\":1901,\"_47\":1902},{},[1407],{\"_287\":288,\"_41\":389,\"_290\":1904,\"_47\":1905},{},[1906],\"pg_hba.conf\",{\"_287\":288,\"_41\":389,\"_290\":1908,\"_47\":1909},{},[1481],{\"_287\":288,\"_41\":289,\"_290\":1911,\"_47\":1912},{},[1913],\"Here is another thing we have to solve. The replicas need to reach the primary, and the primary needs to accept their connections. And every one of these addresses is an IP that changes when a container restarts.\",{\"_287\":288,\"_41\":555,\"_290\":1915,\"_47\":1916},{\"_49\":111},[1917],[\"SingleFetchClassInstance\",1918],{\"_287\":288,\"_41\":335,\"_290\":1919,\"_47\":1920},{\"_338\":1921},[112],\"#they-have-to-find-each-other\",{\"_287\":288,\"_41\":289,\"_290\":1923,\"_47\":1924},{},[1925,1926,1927],\"Now we have three nodes with three Postgres instances. One of the instances is the primary (here it's \",[\"SingleFetchClassInstance\",1928],\"). But this raises new problems, like what to do if the primary's node dies?\",{\"_287\":288,\"_41\":389,\"_290\":1929,\"_47\":1930},{},[1481],{\"_287\":288,\"_41\":289,\"_290\":1932,\"_47\":1933},{},[1934,1935,1649,1936,1937,1938,1939,1940,1941],\"A single Postgres instance is a single point of failure. We want high availability: one primary and two replicas. These need to be on three different machines, with streaming replication between them. So we do the same steps again, three times, on \",[\"SingleFetchClassInstance\",1954],[\"SingleFetchClassInstance\",1950],\", and \",[\"SingleFetchClassInstance\",1946],\". I also wire up replication by hand: \",[\"SingleFetchClassInstance\",1942],\", replication slots, all of it.\",{\"_287\":288,\"_41\":389,\"_290\":1943,\"_47\":1944},{},[1945],\"primary_conninfo\",{\"_287\":288,\"_41\":389,\"_290\":1947,\"_47\":1948},{},[1949],\"node-19\",{\"_287\":288,\"_41\":389,\"_290\":1951,\"_47\":1952},{},[1953],\"node-12\",{\"_287\":288,\"_41\":389,\"_290\":1955,\"_47\":1956},{},[1481],{\"_287\":288,\"_41\":555,\"_290\":1958,\"_47\":1959},{\"_49\":114},[1960],[\"SingleFetchClassInstance\",1961],{\"_287\":288,\"_41\":335,\"_290\":1962,\"_47\":1963},{\"_338\":1964},[115],\"#one-isnt-enough\",{\"_287\":288,\"_41\":289,\"_290\":1966,\"_47\":1967},{},[1968],\"These are a lot of steps, and each one can fail halfway. And if the disk fills up later, Postgres stops accepting writes and we have to resize the volume by hand: first through the cloud provider, then again inside the filesystem.\",{\"_287\":288,\"_41\":527,\"_290\":1970,\"_47\":1971},{\"_530\":1972,\"_532\":1767,\"_534\":535,\"_389\":-7},[],\"\u003cpre class=\\\"shiki shiki-themes planetscale-light planetscale-dark\\\" style=\\\"--shiki-light:#2b2b2b;--shiki-dark:#e1e1e1;--shiki-light-bg:#ebebeb;--shiki-dark-bg:#1a1a1a\\\" tabindex=\\\"0\\\"\u003e\u003ccode\u003e\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#818181;--shiki-dark:#A1A1A1\\\"\u003e# provision + attach first with cloud CLI, then on the node:\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003emkfs.ext4\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#414141;--shiki-dark:#C1C1C1\\\"\u003e /dev/nvme1n1\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003emkdir\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#0B6EC5;--shiki-dark:#73C7F9\\\"\u003e -\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#7D5903;--shiki-dark:#FED54A\\\"\u003ep\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#414141;--shiki-dark:#C1C1C1\\\"\u003e /var/lib/pg-data\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003emount\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#414141;--shiki-dark:#C1C1C1\\\"\u003e /dev/nvme1n1\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#414141;--shiki-dark:#C1C1C1\\\"\u003e /var/lib/pg-data\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003edocker\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#414141;--shiki-dark:#C1C1C1\\\"\u003e run\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#0B6EC5;--shiki-dark:#73C7F9\\\"\u003e -\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#7D5903;--shiki-dark:#FED54A\\\"\u003ed\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#0B6EC5;--shiki-dark:#73C7F9\\\"\u003e -\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#7D5903;--shiki-dark:#FED54A\\\"\u003e-name\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#414141;--shiki-dark:#C1C1C1\\\"\u003e pg\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#7D5903;--shiki-dark:#FED54A\\\"\u003e \\\\\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#0B6EC5;--shiki-dark:#73C7F9\\\"\u003e -\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#7D5903;--shiki-dark:#FED54A\\\"\u003ee\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#414141;--shiki-dark:#C1C1C1\\\"\u003e POSTGRES_PASSWORD=secret\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#7D5903;--shiki-dark:#FED54A\\\"\u003e \\\\\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#0B6EC5;--shiki-dark:#73C7F9\\\"\u003e -\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#7D5903;--shiki-dark:#FED54A\\\"\u003ev\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#414141;--shiki-dark:#C1C1C1\\\"\u003e /var/lib/pg-data:/var/lib/postgresql\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#7D5903;--shiki-dark:#FED54A\\\"\u003e \\\\\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#414141;--shiki-dark:#C1C1C1\\\"\u003e postgres:18\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003c/span\u003e\u003c/code\u003e\u003c/pre\u003e\",{\"_287\":288,\"_41\":289,\"_290\":1974,\"_47\":1975},{},[1976],\"Container storage is ephemeral, so I have to attach a real block device. In the cloud this is an EBS volume (e.g. on AWS); on bare metal it's a physical disk. Assuming it's a block device, this is what we usually do: provision the volume, attach it to the node, format it, mount it, and point Postgres' data directory at the mount.\",{\"_287\":288,\"_41\":555,\"_290\":1978,\"_47\":1979},{\"_49\":117},[1980],[\"SingleFetchClassInstance\",1981],{\"_287\":288,\"_41\":335,\"_290\":1982,\"_47\":1983},{\"_338\":1984},[118],\"#it-needs-a-real-disk\",{\"_287\":288,\"_41\":289,\"_290\":1986,\"_47\":1987},{},[1988,1989,1990],\"I picked \",[\"SingleFetchClassInstance\",1991],\" because it looked idle enough. I start keeping track of it, save it in some sort of config file, and push it to some repo.\",{\"_287\":288,\"_41\":389,\"_290\":1992,\"_47\":1993},{},[1481],{\"_287\":288,\"_41\":527,\"_290\":1995,\"_47\":1996},{\"_530\":1997,\"_532\":1767,\"_534\":535,\"_389\":-7},[],\"\u003cpre class=\\\"shiki shiki-themes planetscale-light planetscale-dark\\\" style=\\\"--shiki-light:#2b2b2b;--shiki-dark:#e1e1e1;--shiki-light-bg:#ebebeb;--shiki-dark-bg:#1a1a1a\\\" tabindex=\\\"0\\\"\u003e\u003ccode\u003e\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003essh\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#414141;--shiki-dark:#C1C1C1\\\"\u003e node-07\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#13862E;--shiki-dark:#75DB8C\\\"\u003e 'docker run -d --name pg ... postgres:18'\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003c/span\u003e\u003c/code\u003e\u003c/pre\u003e\",{\"_287\":288,\"_41\":289,\"_290\":1999,\"_47\":2000},{},[2001,2002,2003,2004,2005],\"Imagine we have hundreds of nodes (servers) we can use. I already have other workloads running on them. I need to decide \",[\"SingleFetchClassInstance\",2009],\" runs this database. So I \",[\"SingleFetchClassInstance\",2006],\" into the box that looks the least busy and start the container there.\",{\"_287\":288,\"_41\":389,\"_290\":2007,\"_47\":2008},{},[580],{\"_287\":288,\"_41\":355,\"_290\":2010,\"_47\":2011},{},[2012],\"which one\",{\"_287\":288,\"_41\":555,\"_290\":2014,\"_47\":2015},{\"_49\":120},[2016],[\"SingleFetchClassInstance\",2017],{\"_287\":288,\"_41\":335,\"_290\":2018,\"_47\":2019},{\"_338\":2020},[121],\"#pick-a-node-by-hand\",{\"_287\":288,\"_41\":289,\"_290\":2022,\"_47\":2023},{},[2024,2025,2026,2027,2028],\"There is already a gap between what I \",[\"SingleFetchClassInstance\",2032],\" (Postgres, running, with my data) and what I \",[\"SingleFetchClassInstance\",2029],\" (a container whose storage disappears when the container or node goes away). The rest of this post is about that gap and the machinery we build to close it.\",{\"_287\":288,\"_41\":355,\"_290\":2030,\"_47\":2031},{},[1758],{\"_287\":288,\"_41\":355,\"_290\":2033,\"_47\":2034},{},[1762],{\"_287\":288,\"_41\":289,\"_290\":2036,\"_47\":2037},{},[2038],\"That's it. Postgres is running. My app connects to it, writes some rows, and everything works fine. But then the machine goes away: the cloud provider reclaims the instance (hardware fails, or a spot instance gets taken back), or I ship a new version of my setup, which means stopping the old container and starting a fresh one in its place. Either way, the container is replaced, and my data is gone. The container storage was ephemeral, and I did not attach any persistent volume to it.\",{\"_287\":288,\"_41\":527,\"_290\":2040,\"_47\":2041},{\"_530\":2042,\"_532\":1767,\"_534\":535,\"_389\":-7},[],\"\u003cpre class=\\\"shiki shiki-themes planetscale-light planetscale-dark\\\" style=\\\"--shiki-light:#2b2b2b;--shiki-dark:#e1e1e1;--shiki-light-bg:#ebebeb;--shiki-dark-bg:#1a1a1a\\\" tabindex=\\\"0\\\"\u003e\u003ccode\u003e\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#F35815;--shiki-dark:#F35815\\\"\u003edocker\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#414141;--shiki-dark:#C1C1C1\\\"\u003e run\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#0B6EC5;--shiki-dark:#73C7F9\\\"\u003e -\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#7D5903;--shiki-dark:#FED54A\\\"\u003ed\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#0B6EC5;--shiki-dark:#73C7F9\\\"\u003e -\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#7D5903;--shiki-dark:#FED54A\\\"\u003e-name\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#414141;--shiki-dark:#C1C1C1\\\"\u003e pg\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#7D5903;--shiki-dark:#FED54A\\\"\u003e \\\\\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#0B6EC5;--shiki-dark:#73C7F9\\\"\u003e -\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#7D5903;--shiki-dark:#FED54A\\\"\u003ee\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#414141;--shiki-dark:#C1C1C1\\\"\u003e POSTGRES_PASSWORD=secret\u003c/span\u003e\u003cspan style=\\\"--shiki-light:#7D5903;--shiki-dark:#FED54A\\\"\u003e \\\\\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003cspan style=\\\"--shiki-light:#414141;--shiki-dark:#C1C1C1\\\"\u003e postgres:18\u003c/span\u003e\u003c/span\u003e\\n\u003cspan class=\\\"line\\\"\u003e\u003c/span\u003e\u003c/code\u003e\u003c/pre\u003e\",{\"_287\":288,\"_41\":289,\"_290\":2044,\"_47\":2045},{},[2046],\"Let's start from scratch. I want to run Postgres on a Linux box, and I need it inside a container. To start it, we run:\",{\"_287\":288,\"_41\":555,\"_290\":2048,\"_47\":2049},{\"_49\":123},[2050],[\"SingleFetchClassInstance\",2051],{\"_287\":288,\"_41\":335,\"_290\":2052,\"_47\":2053},{\"_338\":2054},[124],\"#one-container-one-machine\",{\"_287\":288,\"_41\":330,\"_290\":2056,\"_47\":2057},{\"_49\":91},[2058],[\"SingleFetchClassInstance\",2059],{\"_287\":288,\"_41\":335,\"_290\":2060,\"_47\":2061},{\"_338\":2062},[92],\"#part-1-running-postgres-by-hand\",{\"_287\":288,\"_41\":341,\"_290\":2064,\"_47\":2065},{},[],{\"_287\":288,\"_41\":2067,\"_290\":2068,\"_47\":2069},\"Callout\",{\"_2101\":2102},[2070,2071],[\"SingleFetchClassInstance\",2076],[\"SingleFetchClassInstance\",2072],{\"_287\":288,\"_41\":289,\"_290\":2073,\"_47\":2074},{},[2075],\"We're going to start slow and gradually ramp things up. Each part builds on the previous.\",{\"_287\":288,\"_41\":289,\"_290\":2077,\"_47\":2078},{},[2079,2080,2081,2082,1649,2083,1937,2084,2085],\"A working understanding of containers and \",[\"SingleFetchClassInstance\",2097],\" helps, but you don't need to be a Kubernetes expert. I'll use terms like \",[\"SingleFetchClassInstance\",2094],[\"SingleFetchClassInstance\",2090],[\"SingleFetchClassInstance\",2086],\", and introduce the parts that matter as we go.\",{\"_287\":288,\"_41\":355,\"_290\":2087,\"_47\":2088},{},[2089],\"eventual consistency\",{\"_287\":288,\"_41\":355,\"_290\":2091,\"_47\":2092},{},[2093],\"fan-in\",{\"_287\":288,\"_41\":355,\"_290\":2095,\"_47\":2096},{},[1814],{\"_287\":288,\"_41\":389,\"_290\":2098,\"_47\":2099},{},[2100],\"kubectl\",\"type\",\"note\",{\"_287\":288,\"_41\":289,\"_290\":2104,\"_47\":2105},{},[2106],\"Before we look at a single line of Kubernetes, we're going to run a production database by hand and slowly let the feedback loop appear on its own. Then we'll map that loop to Kubernetes, with the pieces production needs: a store, watches, queues, retries, and more. At the end, we'll look at what one of these loops looks like in a real operator.\",{\"_287\":288,\"_41\":289,\"_290\":2108,\"_47\":2109},{},[2110],\"An operator is a feedback controller. It's the same closed loop that runs a thermostat or keeps your car at a fixed speed on cruise control. In our case, the thing being controlled is a database. I have been building these loops for years, and the best way I know to make them click is to ignore Kubernetes at the beginning. Kubernetes is full of control theory, even if we don't call it that in the day-to-day.\",{\"_287\":288,\"_41\":289,\"_290\":2112,\"_47\":2113},{},[2114],\"People ask me what an operator actually does. The canonical answer is: \\\"it reconciles desired state.\\\" This is correct, but it also tells you almost nothing.\",{\"_287\":288,\"_41\":289,\"_290\":2116,\"_47\":2117},{},[2118],\"For the last decade, Kubernetes has been the backdrop to most of my work: operating clusters, helping build hosted Kubernetes, and writing Kubernetes operators. At PlanetScale, that now means running stateful systems like Postgres and MySQL in production. Kubernetes has many faces, but here I want to talk about one face only: why it is so good at running workloads at scale.\",\"current\",{\"_2121\":535,\"_2122\":2123,\"_2124\":535},\"development\",\"env\",{\"_2125\":2126,\"_2127\":2128,\"_2129\":2130,\"_2131\":2132,\"_2133\":2134},\"userSignedIn\",\"IMAGE_CDN\",\"https://planetscale-images.imgix.net\",\"IMAGE_CDN_ENABLED\",\"true\",\"INTERNAL_API\",\"https://api.planetscale.com\",\"RELEASE\",\"117b8aaf-965c-42bc-b013-5f72770de4d9\",\"SENTRY_DSN\",\"https://bd81903b44804e22a06bdc0c1a91b303@o499952.ingest.us.sentry.io/4504531942572032\"]\n");</script><!--$--><script nonce="h1m8nAKwwY4XGImd6cj2zjoYAypTpKkpoCcuiMZKkXs=">window.__reactRouterContext.streamController.close();</script><!--/$--><!--/$--></body></html>