From 6f956f9c083b6477f0db8eeabf59fc5cf28e5509 Mon Sep 17 00:00:00 2001 From: weishen Date: Thu, 10 Sep 2026 20:52:28 +0800 Subject: [PATCH] =?UTF-8?q?sreweekly:=20528=20=E6=9C=9F=E6=95=B0=E6=8D=AE?= =?UTF-8?q?=20+=20=E5=85=A8=E6=96=87=E6=8A=93=E5=8F=96=EF=BC=88articles/pa?= =?UTF-8?q?ges=20+=20markdown=20=E6=AD=A3=E6=96=87=E6=89=A9=E5=85=85?= =?UTF-8?q?=EF=BC=89?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- ...gestion-podcast-video-incident-report.html | 12 + ...ing-spotify-podcasts-over-reliability.html | 561 +++ ...s-nobody-has-the-whole-picture-to-ass.html | 553 +++ .../articles/528/04-the-quiet-quarter.html | 484 ++ ...w-we-managed-to-not-wreck-our-systems.html | 4 + ...you-ve-just-had-an-incident-what-next.html | 198 + ...atadog-s-state-of-ai-engineering-2026.html | 3019 ++++++++++++ ...a-read-replica-until-you-ve-read-this.html | 37 + sreweekly/articles/528/index.json | 50 + ...-incident-management-processes-wither.html | 581 +++ .../02-what-comes-after-observability.html | 4 + ...tic-sre-humans-agents-and-reliability.html | 2896 ++++++++++++ ...t-july-2-2026-us-east-services-outage.html | 269 ++ ...-finding-bugs-in-raft-implementations.html | 367 ++ ...our-infrastructure-monitoring-in-2026.html | 106 + ...a-systemd-service-with-privatetmp-yes.html | 78 + ...l-versus-resilience-engineering-views.html | 967 ++++ sreweekly/articles/529/index.json | 50 + ...ted-why-resilience-still-needs-humans.html | 419 ++ .../02-respecting-fatigue-isn-t-coddling.html | 569 +++ ...3-on-building-scalable-control-planes.html | 9 + ...ario-saved-the-eu-but-broke-my-system.html | 198 + sreweekly/articles/530/05-ai-and-sre.html | 403 ++ ...-is-still-taking-down-major-platforms.html | 302 ++ ...state-machines-to-auto-remediate-a-gl.html | 1 + ...turned-off-pub-sub-and-nobody-noticed.html | 32 + sreweekly/articles/530/index.json | 50 + .../531/01-heroic-saves-are-near-misses.html | 565 +++ ...-complexity-tension-in-systems-design.html | 256 + ...-systems-what-most-teams-get-wrong-an.html | 3041 ++++++++++++ .../531/04-seeing-the-people-in-control.html | 1890 ++++++++ .../articles/531/05-type-conversion.html | 395 ++ ...re-tool-q-a-at-incident-fest-adaptive.html | 198 + ...-helped-find-the-sqlite-wal-reset-bug.html | 1 + ...liability-with-topology-spread-constr.html | 1738 +++++++ sreweekly/articles/531/index.json | 50 + ...ecomes-everyone-s-favorite-workaround.html | 545 +++ ...-setup-timeouts-by-tuning-conntrack-g.html | 179 + ...rage-at-scale-what-i-actually-watched.html | 137 + ...oad-the-journey-from-static-rate-limi.html | 748 +++ ...r-and-the-art-of-graceful-degradation.html | 4122 +++++++++++++++++ sreweekly/articles/532/index.json | 44 + ...idents-start-before-the-response-does.html | 563 +++ ...azure-regional-outage-from-july-23-26.html | 983 ++++ ...bases-fail-at-coordination-boundaries.html | 2912 ++++++++++++ .../articles/533/04-the-record-says.html | 484 ++ ...d-automate-and-never-automate-with-ai.html | 157 + ...-slower-how-we-rebuilt-git-serving-at.html | 656 +++ ...ode-review-than-automatable-detection.html | 910 ++++ sreweekly/articles/533/index.json | 50 + sreweekly/html/528-2026-08-02.html | 442 ++ sreweekly/manifest.json | 35 +- ...ingestion-podcast-video-incident-report.md | 71 + ...tting-spotify-podcasts-over-reliability.md | 169 + ...ans-nobody-has-the-whole-picture-to-ass.md | 53 + .../markdown/528/04-the-quiet-quarter.md | 87 + ...how-we-managed-to-not-wreck-our-systems.md | 201 + ...6-you-ve-just-had-an-incident-what-next.md | 73 + ...-datadog-s-state-of-ai-engineering-2026.md | 174 + ...d-a-read-replica-until-you-ve-read-this.md | 197 + ...em-incident-management-processes-wither.md | 58 + .../529/02-what-comes-after-observability.md | 86 + ...entic-sre-humans-agents-and-reliability.md | 106 + ...ort-july-2-2026-us-east-services-outage.md | 96 + ...05-finding-bugs-in-raft-implementations.md | 286 ++ ...-your-infrastructure-monitoring-in-2026.md | 248 + ...f-a-systemd-service-with-privatetmp-yes.md | 54 + ...nal-versus-resilience-engineering-views.md | 26 + ...mated-why-resilience-still-needs-humans.md | 91 + .../02-respecting-fatigue-isn-t-coddling.md | 56 + .../03-on-building-scalable-control-planes.md | 110 + ...-mario-saved-the-eu-but-broke-my-system.md | 102 + sreweekly/markdown/530/05-ai-and-sre.md | 60 + ...ry-is-still-taking-down-major-platforms.md | 10 + ...d-state-machines-to-auto-remediate-a-gl.md | 8 + ...e-turned-off-pub-sub-and-nobody-noticed.md | 163 + .../531/01-heroic-saves-are-near-misses.md | 52 + ...nd-complexity-tension-in-systems-design.md | 179 + ...ed-systems-what-most-teams-get-wrong-an.md | 134 + .../531/04-seeing-the-people-in-control.md | 67 + sreweekly/markdown/531/05-type-conversion.md | 42 + ...-sre-tool-q-a-at-incident-fest-adaptive.md | 62 + ...le-helped-find-the-sqlite-wal-reset-bug.md | 150 + ...reliability-with-topology-spread-constr.md | 131 + ...-becomes-everyone-s-favorite-workaround.md | 44 + ...unethical-ways-to-manage-technical-debt.md | 4 + ...od-setup-timeouts-by-tuning-conntrack-g.md | 288 ++ ...torage-at-scale-what-i-actually-watched.md | 44 + .../05-the-rise-of-cognitive-observability.md | 4 + ...rload-the-journey-from-static-rate-limi.md | 205 + ...ger-and-the-art-of-graceful-degradation.md | 118 + ...ncidents-start-before-the-response-does.md | 48 + ...n-azure-regional-outage-from-july-23-26.md | 53 + ...tabases-fail-at-coordination-boundaries.md | 139 + sreweekly/markdown/533/04-the-record-says.md | 124 + ...uld-automate-and-never-automate-with-ai.md | 74 + ...ng-slower-how-we-rebuilt-git-serving-at.md | 169 + .../533/07-a-tale-of-two-flink-autoscalers.md | 4 + ...-code-review-than-automatable-detection.md | 86 + sreweekly/pages/528.html | 442 ++ 100 files changed, 38563 insertions(+), 5 deletions(-) create mode 100644 sreweekly/articles/528/01-content-ingestion-podcast-video-incident-report.html create mode 100644 sreweekly/articles/528/02-the-pulse-quitting-spotify-podcasts-over-reliability.html create mode 100644 sreweekly/articles/528/03-modern-software-architecture-means-nobody-has-the-whole-picture-to-ass.html create mode 100644 sreweekly/articles/528/04-the-quiet-quarter.html create mode 100644 sreweekly/articles/528/05-30-to-70-prs-a-day-how-we-managed-to-not-wreck-our-systems.html create mode 100644 sreweekly/articles/528/06-you-ve-just-had-an-incident-what-next.html create mode 100644 sreweekly/articles/528/07-an-sre-response-to-datadog-s-state-of-ai-engineering-2026.html create mode 100644 sreweekly/articles/528/08-don-t-add-a-read-replica-until-you-ve-read-this.html create mode 100644 sreweekly/articles/528/index.json create mode 100644 sreweekly/articles/529/01-without-a-program-to-support-them-incident-management-processes-wither.html create mode 100644 sreweekly/articles/529/02-what-comes-after-observability.html create mode 100644 sreweekly/articles/529/03-the-rise-of-agentic-sre-humans-agents-and-reliability.html create mode 100644 sreweekly/articles/529/04-incident-report-july-2-2026-us-east-services-outage.html create mode 100644 sreweekly/articles/529/05-finding-bugs-in-raft-implementations.html create mode 100644 sreweekly/articles/529/06-how-to-build-your-infrastructure-monitoring-in-2026.html create mode 100644 sreweekly/articles/529/07-getting-access-to-the-tmp-of-a-systemd-service-with-privatetmp-yes.html create mode 100644 sreweekly/articles/529/08-traditional-versus-resilience-engineering-views.html create mode 100644 sreweekly/articles/529/index.json create mode 100644 sreweekly/articles/530/01-expertise-can-t-be-automated-why-resilience-still-needs-humans.html create mode 100644 sreweekly/articles/530/02-respecting-fatigue-isn-t-coddling.html create mode 100644 sreweekly/articles/530/03-on-building-scalable-control-planes.html create mode 100644 sreweekly/articles/530/04-mario-saved-the-eu-but-broke-my-system.html create mode 100644 sreweekly/articles/530/05-ai-and-sre.html create mode 100644 sreweekly/articles/530/06-certificate-expiry-is-still-taking-down-major-platforms.html create mode 100644 sreweekly/articles/530/07-how-stripe-uses-graph-search-and-state-machines-to-auto-remediate-a-gl.html create mode 100644 sreweekly/articles/530/08-we-turned-off-pub-sub-and-nobody-noticed.html create mode 100644 sreweekly/articles/530/index.json create mode 100644 sreweekly/articles/531/01-heroic-saves-are-near-misses.html create mode 100644 sreweekly/articles/531/02-control-and-complexity-tension-in-systems-design.html create mode 100644 sreweekly/articles/531/03-structured-logging-in-distributed-systems-what-most-teams-get-wrong-an.html create mode 100644 sreweekly/articles/531/04-seeing-the-people-in-control.html create mode 100644 sreweekly/articles/531/05-type-conversion.html create mode 100644 sreweekly/articles/531/06-my-boss-wants-me-to-pick-an-ai-sre-tool-q-a-at-incident-fest-adaptive.html create mode 100644 sreweekly/articles/531/07-how-tailscale-helped-find-the-sqlite-wal-reset-bug.html create mode 100644 sreweekly/articles/531/08-optimizing-kubernetes-pods-for-reliability-with-topology-spread-constr.html create mode 100644 sreweekly/articles/531/index.json create mode 100644 sreweekly/articles/532/01-when-declaring-an-incident-becomes-everyone-s-favorite-workaround.html create mode 100644 sreweekly/articles/532/03-solving-mysterious-kubernetes-pod-setup-timeouts-by-tuning-conntrack-g.html create mode 100644 sreweekly/articles/532/04-storage-at-scale-what-i-actually-watched.html create mode 100644 sreweekly/articles/532/06-how-uber-conquered-database-overload-the-journey-from-static-rate-limi.html create mode 100644 sreweekly/articles/532/07-voyager-and-the-art-of-graceful-degradation.html create mode 100644 sreweekly/articles/532/index.json create mode 100644 sreweekly/articles/533/01-incidents-start-before-the-response-does.html create mode 100644 sreweekly/articles/533/02-quick-thoughts-on-azure-regional-outage-from-july-23-26.html create mode 100644 sreweekly/articles/533/03-why-distributed-databases-fail-at-coordination-boundaries.html create mode 100644 sreweekly/articles/533/04-the-record-says.html create mode 100644 sreweekly/articles/533/05-what-sres-should-automate-and-never-automate-with-ai.html create mode 100644 sreweekly/articles/533/06-20-the-ci-traffic-without-getting-slower-how-we-rebuilt-git-serving-at.html create mode 100644 sreweekly/articles/533/08-there-is-more-to-code-review-than-automatable-detection.html create mode 100644 sreweekly/articles/533/index.json create mode 100644 sreweekly/html/528-2026-08-02.html create mode 100644 sreweekly/markdown/528/01-content-ingestion-podcast-video-incident-report.md create mode 100644 sreweekly/markdown/528/02-the-pulse-quitting-spotify-podcasts-over-reliability.md create mode 100644 sreweekly/markdown/528/03-modern-software-architecture-means-nobody-has-the-whole-picture-to-ass.md create mode 100644 sreweekly/markdown/528/04-the-quiet-quarter.md create mode 100644 sreweekly/markdown/528/05-30-to-70-prs-a-day-how-we-managed-to-not-wreck-our-systems.md create mode 100644 sreweekly/markdown/528/06-you-ve-just-had-an-incident-what-next.md create mode 100644 sreweekly/markdown/528/07-an-sre-response-to-datadog-s-state-of-ai-engineering-2026.md create mode 100644 sreweekly/markdown/528/08-don-t-add-a-read-replica-until-you-ve-read-this.md create mode 100644 sreweekly/pages/528.html diff --git a/sreweekly/articles/528/01-content-ingestion-podcast-video-incident-report.html b/sreweekly/articles/528/01-content-ingestion-podcast-video-incident-report.html new file mode 100644 index 00000000..bda92db6 --- /dev/null +++ b/sreweekly/articles/528/01-content-ingestion-podcast-video-incident-report.html @@ -0,0 +1,12 @@ +Content Ingestion & Podcast Video Incident Report | Spotify Engineering

Content Ingestion & Podcast Video Incident Report

Feature Image

Over the past two months, podcast creators have experienced a series of reliability issues on Spotify. This report covers the most significant of these, the June 24 publishing delay, in full detail and describes the broader reliability program now underway across our publishing pipeline.

What happened?

When a podcast creator publishes a new episode, the audio and video content goes through a series of processing steps before it becomes available to Spotify users. These steps include transcoding (converting media into the formats our apps need) and content analysis.

On June 24, our video transcoding infrastructure reached maximum capacity. This created a backlog that delayed the publication of video podcast episodes for several hours. Creators reported that their episodes were not appearing on Spotify as expected. A queue built up and video podcast episodes that would normally be published within minutes were delayed for hours. Understandably, some creators re-uploaded episodes that had not appeared, which added further load. That is on us, not them: the system should have confirmed their upload was received and queued, and it did not.

+
+ BlogPostGraph + +
+

The image above shows the build-up and subsequent emptying of the medium-priority (used for new episodes - shown in blue) and low-priority (used for updates to older episodes - shown in yellow) queues.

Four factors converged to create this situation:

  • Our transcoding infrastructure was running with insufficient headroom to handle large spikes of content delivery. During typical periods of content submission our transcoding systems were able to scale and support both low-priority transcoding as well as the high-priority publication of new content. However, there wasn’t sufficient headroom to scale and support large spikes caused by the bulk delivery of new content.

  • A scheduled batch processing job was running. On occasion, we need to re-process existing episodes to be compatible with changes or additions to Spotify’s playback systems. During the June 24 disruption, a routine batch job processing existing content was consuming additional capacity alongside the regular processing for new episodes. While this job appeared fine earlier in the day, it became problematic when combined with increased content submissions.

  • Recent improvements increased per-item processing cost. We had recently changed our video transcoding to deliver better quality at lower bitrates. That change increased the time and processing power each episode requires, and we did not fully account for that added demand in our capacity planning.

  • A software bug was underutilizing available compute resources. Following a recent infrastructure migration to more powerful hardware, a bug in our resource scheduling caused our systems to underuse available processing capacity, reducing throughput by about 10%.

When we identified the issue, we stopped the batch job, deployed a fix for the resource scheduling bug, and added additional processing capacity overnight. By the following morning, all backlogs had cleared and publishing was operating normally. We subsequently added further capacity to provide the headroom that had been missing.

Timeline (UTC)

  • 13:30 — Early alerts fire in our internal monitoring. Not immediately recognized as a broader capacity issue.

  • 15:00 — Video podcast delivery spike pushes transcoding close to maximum capacity.

  • 16:35 — Batch processing job stopped to free capacity.

  • 17:31 — First Creator report of an issue impacting podcast video publishing received.

  • 17:34 — Automated alerts confirm queue backlog exceeding thresholds. Incident response begins.

  • 19:00 — Creator reports escalated to incident team.

  • 20:49 — Software fix deployed to improve resource utilization.

  • 00:14 (Jun 25) — Additional processing cluster brought online.

  • 01:02 — All queues cleared.

  • 07:30 — Full confirmation: all publishing pipelines operating normally.

One thing this timeline makes plain: roughly four hours passed between the first alerts and formal incident response. We want to do better. Engineers investigating the early alerts stopped the batch job at 16:35, but we did not recognize the full scope of the capacity problem until queues breached thresholds at 17:34. The monitoring improvements described below exist to close exactly that gap.

Where do we go from here?

We have already taken several steps to address this specific incident:

  • Increased our transcoding capacity by approximately 67%, providing significantly more headroom for traffic spikes and batch operations.

  • Fixed the resource scheduling bug that was leaving under-utilized compute capacity.

  • Improved our monitoring to alert earlier when capacity is approaching limits.

Beyond this specific incident, we are investing in broader improvements to the reliability of our podcast publishing pipeline. We've formed a dedicated cross-team effort focused on:

  • Building better capacity planning that accounts for not just steady-state traffic, but also burst capacity and incident recovery.

  • Improving prioritization across our publishing systems so that real-time content from creators is always processed ahead of background operations.

  • Extending rate limiting and backpressure mechanisms throughout the pipeline to handle unexpected load gracefully.

  • During this incident, many creators learned something was wrong from their audiences before they heard anything from us. We are improving our processes and technical capabilities so creators get notified as soon as possible when things aren’t working. 

When a creator hits publish, their audience is waiting, and hours matter. We fell short repeatedly this summer, and we know a report like this only counts if the next incident is handled better than the last. The work above is how we intend to earn that trust back.

\ No newline at end of file diff --git a/sreweekly/articles/528/02-the-pulse-quitting-spotify-podcasts-over-reliability.html b/sreweekly/articles/528/02-the-pulse-quitting-spotify-podcasts-over-reliability.html new file mode 100644 index 00000000..e54ce4fb --- /dev/null +++ b/sreweekly/articles/528/02-the-pulse-quitting-spotify-podcasts-over-reliability.html @@ -0,0 +1,561 @@ + + + + + + + The Pulse: Quitting Spotify Podcasts over reliability - The Pragmatic Engineer + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +
+ + + + + +
+ +
+ +
+
+ +
+

The Pulse: Quitting Spotify Podcasts over reliability

+
+ +
+ + +

Hi, this is Gergely with a bonus, free issue of the Pragmatic Engineer Newsletter. In every issue, I cover Big Tech and startups through the lens of senior engineers and engineering leaders. Today, we cover one out of four topics of last week's The Pulse issue. Full subscribers received the article below seven days ago. If you’ve been forwarded this email, you can subscribe here.

You can no longer watch The Pragmatic Engineer Podcast as video in the Spotify app (only as audio) because I have quit publishing video on that streaming platform. This comes after I decided that reliability takes a back seat within that team – and across much of Spotify. Unlike on other platforms such as YouTube, Apple Podcasts, and Substack, I’ve recently encountered a series of reliability issues around Spotify being unable to process video episodes. Even though I enjoyed a direct link with the Podcasts team there, things haven’t improved.

So from now, I will no longer be publishing video episodes on Spotify. You can find videos of my in-depth chats with guests only on YouTube. Apologies for any inconvenience this change causes! Audio episodes of the podcast can still be found on Spotify via the RSS podcast feed hosted on Substack.

Honestly, the decision to quit the streaming giant wasn’t hard, and I reckon there’s a point here about the risk of deprioritizing reliable operations at major companies in order to push on things like AI adoption, as Spotify seems to be doing.

Some context: for the first two years of The Pragmatic Engineer Podcast, it was published on three podcast platforms:

  1. Substack’s podcast platform (audio): this is where the “master” RSS feed is served to the likes of Apple Podcasts, the web, Overcast, Pocket Casts, etc
  2. YouTube (video): video episodes uploaded individually
  3. Spotify (video + audio): every video episode was uploaded individually and then served as video or audio episodes from the platform.

As someone hosting a podcast, there are good reasons to bother doing three separate uploads:

  • Most podcast platforms don’t support video. There will always be a need for a platform that serves the master RSS feed for audio versions while the video ones are elsewhere.
  • YouTube doesn’t integrate with anything. YouTube is the leader in video podcast distribution, and uploading there directly makes sense.
  • I had a direct line to the Spotify team, which was a big plus. Starting out the podcast, I had the unusual privilege of contact with the podcasts team, thanks to the newsletter gaining a decently-size audience. I was persuaded to take the plunge with them.

For eighteen months, nothing major went wrong. The admin portal for podcast publishers (called ‘Spotify Creators’) was pretty wonky; it gave intermittent errors, and was unable to remember me when I signed in, so, each Wednesday, I’d have to sign in with a code sent to my email to publish an episode.

But overall, things worked, until it all went suddenly downhill…

Unable to publish Spotify podcast episodes 3 out of 5 weeks

From late May, I did not include links to Spotify on new episode announcements because their podcasts product or platform seemingly had outages every time one published on Wednesdays at around 9am PST / 12pm EST / 6pm EU time.

Outage #1 (20 May): podcast publishing broke, my episode would not process on Spotify for 2+ hours. When uploading a video file to Spotify, there’s a processing pipeline that runs to create chunks of the podcast in different video and audio formats. This pipeline appeared to stop running, meaning new episodes were not published.

It was not just the publishing that broke: the Creator portal looked absurd, with NaN% values everywhere, during the outage:

During outage #1

I emailed the Spotify team to alert them about the outage and also complained online. I got a response, confirming the outage and pledging to do better:

“The issue was in one of our podcast publishing metadata pipelines. A small subset of episodes completed normal media processing but then missed a downstream publish update because a newly introduced validation signal was not correctly wired into the logic that wakes up the publishing path. In simpler terms: the episode could become eligible to publish, but the final propagation step was not reliably triggered for that class of episodes.

We identified the root cause, deployed a fix, and reprocessed the affected episodes with all-clear called early this morning. We’re also tightening the system so that fields used for publishing eligibility cannot be added without also triggering the relevant downstream updates.

Separately, we’re reviewing how partial creator-impacting publishing delays are surfaced, because even when this is not a broad platform outage, it is still a bad experience for publishers like yourself.

Apologies again that you hit this. It was a real bug, not a wide outage, but it hit some of our most relevant creators.”

Outage #2 (17 June): Spotify down. Four weeks later, when attempting to publish a video episode, all of Spotify went down for many users, including myself.

Spotify’s web player on 17 June

Spotify does not maintain a status page, so it’s impossible to tell how widespread the outage was. I didn’t include a Spotify link in that week’s announcement either.

Outage #3 (24 June): podcast publishing broke – again. Outage #3 in five weeks; deja vu. This time, it was episode publishing not working, yet again. After waiting two hours for the episode to publish on Spotify, I yet again sent out the announcement with no Spotify link.

I also emailed the Spotify Podcasts team, who confirmed the outage. I said I was considering stopping publishing video episodes, and to switch to audio-only publishing (which means pointing Spotify to my master RSS feed.) I said that an apology was appreciated but it wasn’t enough to make it worth publishing video episodes there.

I also asked for the incident review because I had the feeling that reliability was not all that important on this podcast product. For the first outage I got a vague description of what happened, and promises of improvements that were never done – e.g. during this second outage, there was no improved communications to creators, which I was told would happen, after outage #1.

Internally, Spotify’s team surely conducted an incident review as per usual, so I figured I’d hear back in about two weeks’ time, and assumed a reply would be forthcoming because I’d made clear I was ready to leave Spotify Podcasts if reliability didn’t improve.

No incident review three weeks later, so I quit Spotify

The incident review had never arrived as promised by three weeks later, even though there had been time for it to be completed. It was yet another sign of a platform that has become unreliable. Also, the creator portal occasionally threw up this error:

Spotify’s creator portal on 16 July

I checked my Spotify stats: stream plays had been trending downwards unsurprisingly, given the ongoing outages, while the other podcast platforms didn’t show the decline. It made me decide “enough is enough” and to move off Spotify.

Staying on their platform depended on seeing an incident review, but they didn’t prioritize transparency, still had no status page, and nobody had built a feature for episode-processing status like YouTube has had for years. So, I pulled the plug and left:

Offboarding from Spotify’s (video) podcasts product

After I made the switch away from Spotify, the platform’s creators portal became buggier than ever, as in these examples:

My Creators page after I changed the source of my podcasts to the master RSS feed

Comments disappeared:

My show had no comments, suddenly

… even though other parts of the UI showed dozens of comments:

Zero comments, yet episodes with comments

Episode links directed to 404 pages:

404 pages inside the Creator portal, when clicking links

A day or two later, these issues disappeared: I assume no one had tested the flow of moving away from Spotify Podcasts to an RSS feed, and it’s why the experience was so poor.

Incident review finally published, but with a wrong timeline

A few days after offboarding from Spotify, their team published the incident report for outage #3. Reading through it, something did not add up in the timeline:

The original timeline published for the 24 June incident

My email account confirmed that I mailed the Spotify team at around 17:30 about the outage. So, after weeks of creating this report, why did the incident report downplay the fact that customers alerted the team before their own automated alerts fired?I complained to the Podcasts team, and to their credit, the incident report was updated:

The updated incident timeline

I didn’t like how high-level the report is, and how vague the promised improvements were. Specifically, this one:

“During this incident, many creators learned something was wrong from their audiences before they heard anything from us. We are improving our processes and technical capabilities so creators get notified as soon as possible when things aren’t working.”

Overall, I don’t regret the choice to leave, particularly when the focus of Spotify’s leadership is on AI, not reliability.

Does Spotify have “AI psychosis?”

Previously, I used the term “AI psychosis” differently from the usual way of describing when someone starts believing everything an AI model tells them, however outlandish. I applied it to Meta’s rush to develop its own AI model at the cost of the reliability of its profitable business activities. This was based on Instagram’s most embarrassing-ever account takeover incident, which occurred when the team responsible for Instagram’s Trust & Safety was slashed. Soon after, AI-generated, AI-reviewed code caused the hacking of a former US president’s account.

At Spotify, it should have gone the other way. In March, I had the opportunity to meet its Head of Technology & Platforms, Tyson Singer, who said the company puts reliability far ahead of AI adoption, and doesn’t adopt AI for its own sake. So, it was somewhat surprising to read the summary below of a podcast Spotify did with Anthropic:

“Spotify now ships 4,500 production deploys a day, and 73% of PRs are now AI-assisted.

Niklas Gustavsson (VP of Engineering at Spotify) keeps 5 to 10 Claude sessions running in tmux, one per git worktree, agents working in the background. All of it inside a 20M+ line monorepo. He expected agents to struggle at that size, but it’s worked well.

Spotify’s migration codemods grew into thousands of lines of edge cases. Code has too much API surface for static rewrites. Early LLMs barely did better. Adding a judge took PR success from ~25% to 80%.

All of this leans on verification, the single most important thing when agents are used and the place most companies underinvest

Spotify rebuilt their test automation around it so engineers can confidently guide and supervise agents, rather than manually execute repetitive tasks.”

It seems to me that all the talk is about usage of AI, and none about reliability, all while Spotify’s platform becomes less reliable than ever, at the same time as the streamer is going all-in on AI; with AI judges and devs running 5-10 parallel Claude sessions.

All things considered, it’s worth asking if Spotify has the corporate variant of “AI psychosis”, whereby the reliability of a successful operation gets torched in the chase for the next big thing by executives. I don’t even think Spotify is all that different from Meta and other companies in this!

Things look bad, based on the quality and reliability degradation of products. Annoyingly, in many cases, customers don’t really have the choice of going elsewhere. My podcast is an exception, as video podcasts on Spotify never truly took off, so quitting the platform wasn’t a big deal. Even so, I’m particularly disappointed that Spotify has prioritized AI usage over reliability. I know some executives there pushed against this, but I feel safe in assuming that they lost that battle.

Value of staying reliable & “sucking less”

Max Kanat-Alexander, distinguished engineer at Capital One, has written about how a software project can become wildly successful just by “sucking less” in his reflections upon the success of the Bugzilla project, (2004-2009):

“All you have to do to succeed in software is to consistently suck less with every release.

Nobody would say that Bugzilla 2.18 was awesome, but everybody would say that it sucked less than Bugzilla 2.16 did. Bugzilla 2.20 wasn’t perfect, but without a doubt, it sucked less than Bugzilla 2.18. And then Bugzilla 3.0 fixed a whole lot of sucking in Bugzilla, and it got a whole lot more downloads.

Why is it that this worked?

As long as you consistently suck less with every release, you will retain most of your users. You’re fixing the things that bother them, so there’s no reason for them to switch away. Even if you didn’t fix everything in this release, if you sucked less, your users will have faith that eventually, the things that bother them will be fixed. New users will find your software, and they’ll stick with it too. And in this way, your user count will increase steadily over time.

But what happens if you release frequently, but instead of fixing the things in your software that suck, you just add new features that don’t fix the sucking? Well, eventually the patience of the individual user is going to run out. They’re not going to wait forever for your software to stop sucking.”

Personally, I got tired of Spotify’s Podcasts product continually going in the wrong direction on Max’s scale: the poor reliability, frequent errors on the Creators site, and the sense that they don’t really care about improving existing things.


Read the full issue of last week's The Pulse. The full The Pulse additionally covers:

  1. Will Kimi K3 trigger US push for closed-source AI models? Moonshot AI’s latest open model, Kimi K3, is on par with Anthropic’s Fable 5. Could it lead to the US government regulating or banning Chinese open models to protect US labs?
  2. AWS laughs off “heart attack” billing error. AWS customers were billed trillions more than they should have been, due to what was likely a conversion error. But instead of sharing an incident report, AWS saw the funny side.
  3. Industry pulse. OpenAI’s unreleased model tried to hack HuggingFace to improve its test scores, X took more than a year to develop its new Android app, Google’s new AI model flops, and more.

Read the full The Pulse.

+ + + +

+ Subscribe to my weekly newsletter to get articles like this in your inbox. It's a pretty good read - and the #1 software engineering newsletter on Substack. +

+ +
+ + +
+
+ + + + + + + + + +
+ + + + + + + + + + + + + + + + + + + + + + + diff --git a/sreweekly/articles/528/03-modern-software-architecture-means-nobody-has-the-whole-picture-to-ass.html b/sreweekly/articles/528/03-modern-software-architecture-means-nobody-has-the-whole-picture-to-ass.html new file mode 100644 index 00000000..36e63baa --- /dev/null +++ b/sreweekly/articles/528/03-modern-software-architecture-means-nobody-has-the-whole-picture-to-ass.html @@ -0,0 +1,553 @@ + + + + + + + + + + Modern software architecture means nobody has the whole picture. To assemble one in an emergency, you need an incident tech lead. | Brent Chapman + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +
+ + + +
+ +
+
+ +
+
+
+
+
+ + +
+ +

Every large software development organization has made the same bargain: if nobody has to understand the whole system, we can build a more capable system than we otherwise could, even though it grows bigger and more complex. We break systems into components with well-defined interfaces so that each team needs to understand only the pieces it owns, plus the interfaces of its neighbors. That’s the point of every decomposition strategy, whether it’s microservices, bounded contexts, service ownership, or a well-modularized monolith. The architecture deliberately limits what any one person has to hold in their head.

+ + + +

It’s a sound strategy. It’s also why some incidents are so much harder than others.

+ + + +

Lorin Hochstein named this pattern beautifully in a recent post, The demon of the gaps. Failures that stay inside a single component are the easy ones; you page the owning team, they figure out what’s wrong, and they fix it. The hairy incidents emerge from unexpected interactions across components: several services throwing errors at once, or no services throwing errors while customers see broken behavior anyway. As Lorin puts it, “you’ve built an analysis solution but you’re now faced with a synthesis problem.” In order to scale, the architecture deliberately optimized away the need for whole-system understanding; now the whole system isn’t working, but nobody has that understanding to call on.

+ + + +

I’ve watched this play out in incident channels many times. Subject matter experts from six different teams, each reporting that their own service looks healthy. Six dashboards are green, but meanwhile, checkout is still failing for customers. The knowledge needed to explain what’s happening exists, distributed across six heads, but nobody is assembling the pieces. The responders need to understand how the system as a whole is behaving right now. That understanding has to be built live, under pressure, from multiple partial models. That’s synthesis work, and it doesn’t happen on its own.

+ + + +

Lorin observes that guidance on preparing for this work is almost nonexistent. Here’s the encouraging part: closing the structural gap is fairly straightforward. Most companies already use structured incident roles (incident commander, subject matter expert, customer liaison, etc.); they need to add a synthesis role, activated when needed for complex incidents.

+ + + +

Synthesis is a job. Name it.

+ + + +

The incident commander (IC) coordinates the overall response; every incident will have one. On the most complex incidents, though, where you need this synthesis function most, the trick is to also activate an incident tech lead (TL) to lead the technical investigation. The role is analogous to the tech lead role many teams have in their everyday structure, but its scope is the incident rather than one particular service. Most companies have never established the incident TL role, and for routine incidents they don’t miss it: the IC can handle the technical side along with everything else, but for complex incidents, the TL role can be incredibly valuable.

+ + + +

The incident TL job, properly understood, is the synthesis job: connecting observations across component boundaries, correlating the partial models from different subject matter experts, and maintaining the evolving picture of how the system is failing and what we’re doing about it. The TL doesn’t need to be the deepest expert in any single component. They need to be good at building a working model out of other people’s expertise, and much of that work is cross-checking, holding indications from different components up against each other and noticing the discrepancies: “If we’re seeing this in component A, we should be seeing that in component B, but we aren’t; why not?” “If A is doing this and B is doing that, the problem must be upstream of both.” “Wait, A says one thing but B says another; they can’t both be right, can they?”

+ + + +

The separation between IC and TL exists to protect that work, and it cuts both ways. Synthesis requires sustained, heads-down attention; you can’t reconstruct a system model in the gaps between stakeholder updates and staffing decisions. And the same complexity that makes an incident demand serious synthesis also multiplies the outward-facing work: more stakeholders to update, more escalations, more decisions about the response itself. The two loads peak together, and one person can’t carry both.

+ + + +

The IC takes everything outward-facing precisely so the TL can stay immersed in the technical picture, and the TL handles the heads-down focused work so that the IC has time for everything else. When I’m the incident commander, one of the most valuable things I can do for my tech lead is keep everyone else out of their hair. But the separation is a division of labor, not a wall. I like to think of the IC and the TL standing back to back, facing opposite directions, talking over their shoulders to keep each other informed. Each is watching a different part of the horizon, and together they have the whole picture.

+ + + +

The response team crosses the boundaries on purpose

+ + + +

An incident response is a temporary organization: an ad hoc team assembled across ownership boundaries for exactly as long as the incident lasts. Conway’s law observes that systems end up mirroring the communication structures of the organizations that build them, and the mirror works in both directions: your team boundaries and your component boundaries align, which is exactly what you want for everyday work. The incident structure deliberately cuts across those boundaries, because the gaps between components are where the problem lives. Pulling six SMEs into one channel isn’t enough by itself, though. A group of experts in the same room is a meeting; a group of experts with someone responsible for synthesizing what they know is a response.

+ + + +

The communication mechanisms are synthesis tools

+ + + +

The standard incident communication practices may seem like bureaucratic overhead until you see what they’re for. “Going around the horn” (each responder, in turn, briefly reports what they’re seeing and doing) forces the partial models into the open, where the TL can correlate them. A periodic situation report, or SitRep, forces someone to compress the current understanding into a few sentences; writing it is itself an act of synthesis, and reading it gives every responder the same baseline picture to work from. Narrating before you act keeps each responder’s local view visible to the whole room. None of these mechanisms exists for discipline’s sake. They’re how a group of people, each holding a partial model, builds and maintains a shared one.

+ + + +

Wildfires don’t respect organizational boundaries either

+ + + +

As is often the case in incident management, we can look to fire departments for inspiration and solutions. Consider a major wildfire. Dozens of agencies converge: federal, state, tribal, and local, some from hundreds or even thousands of miles away. No single agency understands the whole incident, with its terrain, weather, fuel, crews, and aircraft. The Incident Command System (ICS), the standard structure for emergency response in the US and beyond, is how all these disparate parts get pulled together into a coherent whole. ICS treats building the shared picture as a staffed function: a planning section tracks the situation and the resources, assembles the common operating picture, and distributes it to every responder through the incident action plan. Nobody simply hopes that shared understanding will emerge; somebody owns producing it.

+ + + +

Software companies can borrow that lesson directly: treat synthesis as a named responsibility rather than an emergent property. If the IC role at your company is defined as “project manager of the outage” and nobody is explicitly responsible for assembling the technical picture, the synthesis function is unowned, and it will show in your cross-boundary incidents. Establish the incident tech lead role. Protect it from outward-facing distraction. And practice it: when you run game days or tabletop exercises, choose scenarios that cross team boundaries, because those are the scenarios that exercise synthesis rather than component expertise.

+ + + +

Decomposition made whole-system understanding nobody’s everyday job, and that’s fine; it’s a good strategy with a known cost. Incident management structure is how you pay that cost only when you must, with machinery built for the moment.

+ + + +
+ + + +

I’m writing a book, “Incident Management for DevOps and SRE.” Sign up at im4ds.com to be notified when it’s available, and to get occasional progress updates and early access to selected content.

+ + + +

If your company needs help with incident management right now, that’s the focus of my consulting practice at GreatCircle.com/im.

+
+ +
+ +
+ + +
+
+
+ + +
+ + + + +
+
+ + +
+ + + + + + + + + + + + + + diff --git a/sreweekly/articles/528/04-the-quiet-quarter.html b/sreweekly/articles/528/04-the-quiet-quarter.html new file mode 100644 index 00000000..cb53b390 --- /dev/null +++ b/sreweekly/articles/528/04-the-quiet-quarter.html @@ -0,0 +1,484 @@ + + + + + + + + + + + + The Quiet Quarter - by Tim Irving - Zero Sev Zero + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +
+

Discussion about this post

User's avatar

No posts

Ready for more?

    +
    + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + diff --git a/sreweekly/articles/528/05-30-to-70-prs-a-day-how-we-managed-to-not-wreck-our-systems.html b/sreweekly/articles/528/05-30-to-70-prs-a-day-how-we-managed-to-not-wreck-our-systems.html new file mode 100644 index 00000000..2df0d402 --- /dev/null +++ b/sreweekly/articles/528/05-30-to-70-prs-a-day-how-we-managed-to-not-wreck-our-systems.html @@ -0,0 +1,4 @@ +30 to 70 PRs a Day: How We Managed to Not Wreck Our Systems
    Observability Engineering second edition out now! 27 net-new chapters written for today's observability challenges.Get your copy

    30 to 70 PRs a Day: How We Managed to Not Wreck Our Systems

    The Honeycomb engineering team set out to double our productivity in a year. This is how we did it, what we did to keep things stable, what it cost us, and what we’re still figuring out.

    30 to 70 PRs a Day: How We Managed to Not Wreck Our Systems

    In this two-part blog series, I give a detailed report-out on how our Honeycomb engineering team 2.5x-ed our throughput using AI without breaking everything or lowering our standards for quality. Part 1 explains how we did it and shows data about how that ramp-up happened. Part 2 shares what we learned.

    TL;DR

    • Peak-weekday merges roughly doubled (~30 to ~74) as AI-attributed lines went from near-zero to a floor of 82.6% of new code by June 2026. Incidents grew too, tracking that change volume about as linearly as you'd expect; the goal now is keeping each failure cheap to contain, not holding the count flat.
    • The gains came in three phases: slow experimentation (2025), a tooling-driven adoption bump (October 2025), then a step-change in delegation intensity after Opus 4.6 shipped in February 2026—same engineers, same tools, but they stopped supervising every step.
    • Our core thesis: AI amplifies your existing practices. It makes a dysfunctional org more dysfunctional and a high-autonomy, high-ownership org faster. The practices—continuous delivery, fast and AI-legible CI, closed-loop observability, CLAUDE.md, feature flags—are the actual story, not the multiplier.

    Why did I put that big number in the title? It’s the number that gets you to click. But we’re going to dig into all of the caveats and details beyond just 2.5x-ing our throughput, including whether we let quality slip. This blog and its companion post are our report-out on what we did to achieve that result, how we kept the systems underneath that throughput from falling over, and what we learned in the process.

    Setting out to double productivity in a year

    In mid-August 2025, our founders sent a letter to the whole company, not just engineering: each of us should aim to double our productivity over the next year. It was addressed to individuals; the framing underneath it was a team sport, not an individual race, and nobody was supposed to read it as a contest with their teammates. The letter named Darragh Curran’s Intercom 2x post as the framing inspiration, explicitly. Eight months later, we started to measure and reflect. This post follows on from Darragh’s retrospective on hitting 2x in nine months and Kesha Mykhailov and Niamh Young’s post on safely scaling AI auto-approval. We’re smaller than Intercom and a couple of years younger, and we’ve historically followed similar paths a few months behind them.

    The number of merges on a peak weekday to Honeycomb’s monorepo more than doubled from about 30 in early 2025 to about 74 in April 2026. This codebase doubled in sixteen and a half months, from approximately 0.97 million lines at the end of 2024 past 1.94 million in the first week of May 2026, and it sits at 2.1 million as of early July. The previous doubling had taken nearly three years. Most of that net-new code was co-written with or reviewed by AI somewhere in its history.

    Bee-ing honest about that number

    1. It’s peak weekday, not calendar average. Wednesday no-meeting days are the most productive, both before and after AI; calendar average is roughly half of peak. That’s the peak. Don’t go looking for 70 PRs on a random Tuesday and conclude I lied to you.
    2. It’s a floor, not a true count. One of our heaviest Claude Code users, by token volume, has zero AI-attributed git commits. They ran 22 sessions and 199 million tokens through Claude Code in 30 days, and git saw none of it, because their Co-Authored-By trailer is disabled at the tool level. If we can’t measure their AI usage, we can’t measure several other people’s either.
    3. It’s entangled with org changes happening at the same time. In early January 2026, our founders sent a follow-up letter naming sharper strategic stakes than the August one had: rebuilding Honeycomb’s product surface and market posture to be AI-first, well beyond the original productivity target. The mid-January realignment toward greenfield, higher-AI-leverage work followed from that letter. We onboarded new engineers. We invested in platform engineering. Anyone who tells you a single factor caused their team’s gain is overstating it, including us.

    There’s a big leap from “AI tooling makes engineering teams genuinely faster” to “any team that adopts it will get the same result.” Most of this post lives in the gap between those two claims. The interesting story isn’t the multiplier; it’s what we did to keep things stable underneath it, what it cost us, and what we’re still figuring out. The short version, and the line I keep coming back to every time I’m invited on stage: AI amplifies your existing practices. It can make a dysfunctional org more dysfunctional, or it can bring out the best in an org that already has high autonomy, ownership, and feedback loops. We weren’t setting out to prove that thesis; we took a challenge, and the thesis became visible after the fact.

    If you caught my “AI is like chocolate” talk, you know the bit: chocolate doesn’t belong on everything, too much of it in one sitting will make you sick, and no amount of chocolate substitutes for knowing how to cook. I gave that talk as a pessimist turned realist, and that’s still who I am. What changed between then and now isn’t the metaphor; it’s that the tooling and the practices around it got good enough that chocolate’s rightful place in the kitchen got bigger. It’s more versatile and forgiving than it used to be. The concessions later in this post, in “Where the skeptics are right,” are concessions I’m still making. Being a reformed skeptic doesn’t mean I stopped being one.

    You should take every number on this page, including ours, with a grain of salt. The methodology matters more than the magnitude. Keep your semantic bullshit defenses up against hype and plausible-sounding data, all the way through the rest of this post.

    Join the masterclass with Liz Fong-Jones

    Six live sessions with Liz Fong-Jones
    turn Observability Engineering into practice.
    Starts August 3rd.

    What the numbers do and don’t say about quality

    We haven’t had a spectacular AI-caused failure. No “AI deleted my database,” no “AI shipped code that corrupted user data.” That’s not luck; it’s not new either. It’s a result of designing defensively, whether the chaos agents be human or robot. Stacking agents on top of existing infrastructure with bulkheads between components, reviews, deploy trains, feature flags, and least-privilege access has meant outages are lower-impact, rather than either non-existent or uniformly critical-severity. If you don’t have that in place yet, that’s the thing to fix before you scale up AI usage, not after. A dropped database is a systems design issue: somebody skipped building the guardrail that would have caught it, regardless of who or what wrote the code.

    Incidents are growing in absolute count. From a 2024 baseline of about 18.5 incidents per quarter, Q1 2026 hit 32, 1.7x baseline, against PR throughput at about 2.5x baseline; for one quarter that looked sub-linear. Q2 didn’t hold: 53 incidents, 2.9x baseline, even as PR throughput itself held roughly flat once you back out the May freeze weeks and the January ramp-up swarm. Two quarters in, incident growth tracks change volume about as linearly as you’d expect. Change is the leading driver of incidents industry-wide (see the VOID report, Google’s DORA research), and we’re shipping a lot more of it. The law of large numbers caught up with us.

    What the data actually supports is two things: the absence of spectacular failures, plus some integrated second-order effects we’re picking up on. Broader AI-causation narratives tend to be self-flattering, whichever direction they point; AI is woven in deeply enough at this point that isolating it as a single cause is rarely a meaningful exercise. Attribution-by-cause is a fairy tale we humans tell ourselves to feel better about whatever stance we already hold.

    More changes shipped means more chances for a defect to land somewhere in the batch, at whatever the org’s baseline defect rate happens to be. We’re shipping far more change, so we get far more incidents, in roughly the proportion you’d expect. It’s too early to say whether AI assistance in debugging reduces incident severity once something breaks; in some cases it’s helped us find the root cause fast, in others it’s sent us chasing a confident, wrong answer instead. That volume showed up as real strain on the teams absorbing it, not just as a line going up on a chart. We’re leaning on better automated preflight checks, among other levers, to try to bend that curve back down. The burnout risk that comes with capturing AI’s speed as pure output rather than sustainable pace is a real one, and worth naming rather than assuming away.

    The metric I’d rather put on the wall isn’t PRs per day. Throughput is an input metric, not the product; it’s just the coarse measure we have today to demonstrate step-change. Throughput going up while user outcomes plateau or decline is, by definition, enshittification.

    What does “AI contribution” actually mean?

    There’s no single “AI percentage.” There are at least four different denominators, and you have to be precise about which question you’re asking before you quote a number at anyone.

    (Vendor dashboards will happily hand you an “AI-influenced PRs” number with a confidence toggle. Set to loose confidence, ours cheerfully reported a majority of PRs as AI before we’d done any of the work below. Don’t trust low confidence. The numbers in this post come from our own git history and telemetry, calibrated by hand.)

    These are calibrated floors, not point estimates. We got there by layering three corrections onto git’s raw signal: local branch trailers (245 PRs whose AI attribution was stripped by squash-merge), GitHub-API branch trailers (117 PRs where branches were deleted post-merge but we recovered the trailers via GraphQL across all 7,952 PRs), and a smell-test telemetry override (112 PRs from engineers with zero git AI attribution but heavy Claude Code session telemetry, gated per month against their actual session activity).

    The 95% engineer-level adoption figure, which is closer to our gut estimates, reconciles naturally with the lower PR-level (63%) and line-level (75%) floors: high adoption, selective per-PR use. We can measure floors from git history. We can’t measure ceilings. Being clear about the difference matters more than the specific numbers do.

    AI-attributed floor on top of the unattributed baseline

    Surviving lines at HEAD since 2024. The red band is a floor; lines from PRs with AI attribution somewhere in their history. The blue band is “no attribution,” not “no AI.”

    What happened

    The adoption curve at Honeycomb has three distinct phases. Each one was driven by something different.

    The first phase, April through September 2025, was sustained low-rate experimentation. About one or two new engineers adopted Claude Code per month, with support but no mandate; engineers were exploring on their own, at least until mid-August, when the founders’ letter landed and the picture blurs. AI’s share of new lines climbed modestly, peaked around 18% in June, then drifted back down to 8% by October as the honeymoon wore off. This is what “leadership opened the door” looks like in practice: not a discrete push so much as a sustained green light. The dip in the second half of 2025 was evaluation, not failure: engineers tried Claude with the model available at the time, decided it didn’t yet justify the friction, and pulled back. If the catnip is rotten, herding cats to eat the catnip is even more difficult.

    The second phase began in October 2025 with a Claude Code harness improvement. Seven new adopters that month, with no model release behind it; tooling quality alone moved the needle. November and December were a pause (although some engineers used the holiday break to try the tools in their personal capacities). Then Claude 4.5 landed in late December, and adoption picked back up in January 2026 with eight new adopters, and AI’s share of new lines climbing to 18% by January.

    AI commits per month vs cumulative engineers using AI

    Engineers with at least one AI-attributed commit, cumulative. Each step change lines up with a capability event, not a flat schedule.

    Opus 4.6 launched on Thursday, February 5, 2026. The session-rate takeoff in our Claude Code telemetry lands on February 11-13: the first full work-week after a Thursday launch with a Friday-and-weekend bake-in. Distinct monthly Claude Code users stayed nearly flat across January, February, and March (59, 63, 64). Sessions per user 3.3x’d over the same window (21, 35, 70). AI’s share of new lines went from 18% in January to 46% in February to 65% in March.

    We’ve been calling this confidence to delegate. The same engineers, with the same tooling, with access to the same model family, simply changed how they used it. They stopped consulting or closely monitoring each step, and started delegating. The lift is per-user-intensity, not headcount enabled.

    Engineer count vs sessions vs PRs before and after Opus 4.6

    The Q1 proof, one magnitude per panel: engineer count barely moved, sessions per engineer exploded, and committed PRs-by-model shows the delegation landing on Opus 4.6 specifically.

    While the median engineer’s PR throughput grew about 45% relative to its pre-February baseline (call that baseline 1.0x), the top of the distribution moved much further. P75 went from about 1.8x baseline to about 2.9x, roughly +64% in relative terms; the weekly maximum across our active engineers went from a 3.2x-6.4x baseline range pre-February to a 7.7x-12.3x range in April. The floor barely moved at all (P25 went from about 0.5x baseline to about 0.8x). The AI uplift is concentrated at the top end of the distribution. It is not a universal rising tide that automatically boosts every engineer, and that’s okay, because not all engineering is in the bucket of things AI accelerates. As Charity says, the closer to touching bytes on disk you are, the more cautious you need to be about reviewing everything with a paranoid lens.

    We’d never had a dozen engineers a week shipping 7+ PRs each in our entire 10-year history, until March 2026. Pre-February, two to four engineers a week hit seven or more PRs; in February that was three to seven, in March seven to twelve, in April eight to sixteen. New, and now routine. And the engineers at the top aren’t a stable cast; the March-selected and April-selected top-twelve cohorts only overlap by about half. It’s a rotating cast riding the new ceiling, not a handful of superusers carrying everyone else.

    Weekly merged PRs

    Weekly merged PRs (total/AI-attributed/no-attribution) and the per-engineer weekly distribution for each cut, September 2025 through the week of June 22. The top tail past 20 PRs per engineer-week appears in March and persists; the May dent is the freeze weeks, not decay. Bots excluded, including the autobot.

    Peak-weekday non-AI merges held roughly steady at 25-30 across the entire window; humans didn’t slow down on their best days to make room for AI. But at the org-wide weekly level, non-AI commit-hash volume declined notably, from about 120 PR commits a week pre-February to about 60-80 a week through February to April, while AI commit-hash volume grew from near-zero to 150-200 a week.

    So engineers didn’t slow down on the days they were shipping; they shipped fewer human-attributed PRs in aggregate while shipping many more AI-attributed ones. Some of the +44 peak-weekday delta is genuinely new capacity. Some of it is substitution, where work that used to be human-attributed is now agent-attributed because the same engineers shifted to driving with AI rather than typing by hand. We can’t cleanly separate lift from substitution without a controlled experiment we don’t have.

    May, June, and a third workflow

    For this blog, we pulled a fresh data cut through June 28, rather than waiting the six months we’d originally planned since the May and June presentations. One methodology note before the numbers: this refresh also excludes mechanical bots (e.g. Dependabot) from every throughput denominator, something the April numbers above didn’t do. So the figures in this section aren’t a perfectly clean continuation of the ones above them. Same discipline as the rest of this post: check the methodology before you trust the magnitude.

    On that basis, weekday-average merges went 38.0 in March, 47.9 in April, then dipped to 36.0 in May before climbing back to 41.9 across the four complete weeks of June. The May dip isn’t engineers slowing down; it’s a supply constraint. May carried an intense marketing push plus merge freezes around Innovation Week and O11yCon SF. June rebounded as soon as the freezes lifted, and peak-day merges hit 70 again on June 18, matching and sustaining April’s peak.

    The more interesting news isn’t the wobble in the average. It’s that a third category of work showed up entirely.

    honeycomb-autobot[bot] landed its first commit on main on April 23. It’s Claude Code on AgentCore, dispatched from RWX (the same CI substrate from earlier in this post) and traced by Honeycomb, triggered from a Linear issue or an @honeycomb-autobot mention on a review, with no human anywhere in the commit-generation loop, only the review loop. That’s categorically different from “AI-assisted coding.” Human-in-loop coding still means a person is driving the session and choosing what to commit. The autobot doesn’t have anyone in that seat at all, but instead back-loads the work onto the review cycle where work is more mechanical and a human feels confident going hands-free during the actual coding.

    Three months of data on it: 3 autonomous merges in April (0.3% of the month), 13 in May (1.7%), 70 across the four weeks of June (8.4%). Over the same window, human-authored merges with zero AI attribution kept shrinking, down to 211 in four weeks of June against a 2025 baseline in the 400s a month, while total throughput held at roughly twice the 2025 baseline the whole time. Same pattern as everywhere else in this post: substitution, not addition.

    Merges and lines by workflow

    The three-way split, January 2025 through the four full weeks of June 2026, mechanical bots excluded. The assisted band is a floor; the autonomous band is exact, because bot authorship is self-evident.

    The adoption shape looks like the rest of the story too, not a power-user phenomenon. Of the 96 autobot squashes on main through July 6, 63 carry an explicit “Triggered by” line naming 23 distinct engineers; one of the eng enablement leads who co-authored autobot is the heaviest user at 18, and everyone else in the tail is 2 to 4 each. And 55 of those 96 squashes carry no Claude co-author trailer at all, which means the same trailer-based blind spot from the “Bee-ing honest” section up top shows up here too. If we counted the autobot’s work by trailer the way we count human-driven work, we’d have missed 57% of it. We count it by author identity instead, which for a bot account is exact rather than a floor. But it’s a reminder that every attribution method has exactly one failure mode it’s blind to, and you only find out what it is by checking, not by assuming it doesn’t have one.

    The autonomous workflow rides the frontier model the same way human delegation does: 30 of 33 model-tagged autobot squashes in the four weeks of June cite Opus 4.8. And it isn’t yet a lines-of-code story. Autonomous work has added roughly 5,500 lines total since April, about 0.3% of the codebase’s current size. (Added, not surviving; it’s too early to measure how many of those lines are still alive at HEAD.) Right now this is a merge-count phenomenon, not a codebase-composition one. It’ll be worth watching whether that changes.

    The codebase kept growing underneath all of this. HEAD was at 1.88 million lines in April; by July 6 it’s 2,096,286, more than double the 972,000 lines at the end of 2024. Net-new code since end-2024 is now majority AI-attributed for the first time: 599,000 of 1,124,000 added lines, at least 53%, up from 41% in April. June’s line-level floor, at least 82.6% of new lines AI-attributed, is the highest month on record, ahead of April’s 75%. The projection in the April data I presented on-stage, that AI would cross a third of HEAD “around end of 2026,” turned out to be conservative; the current slope puts that closer to September, with half of HEAD by roughly mid-2027.

    One number needs its own caveat rather than a triumphant read. Human-attributed lines surviving at HEAD actually ticked down slightly, from 1,509,000 in April to 1,497,000 in July. That’s not a clean “AI replaced human code” story. Some of it is genuine replacement; some of it is recalibration retroactively reclassifying PRs that were originally counted as human, as we keep finding hidden AI attribution in old PRs. Don’t read a precise story into that number. Read it as more evidence that the floor keeps rising as we get better at measuring it, which has been true of every number in this post so far.

    Engineers with at least one AI-attributed commit: 70, up from 64 in April. That’s broadening, not just the same 50-plus people going faster. And the frontier-model succession kept stair-stepping exactly the way it did in February: Opus 4.6 gave way to Opus 4.7, which gave way to Opus 4.8 (366 of June’s model-tagged PRs, against 31 for 4.7 and 29 for 4.6, the same one-month displacement pattern each time). Claude Fable 5, the newest Mythos-tier model, shows up in June’s trailers too, on 34 PRs. Sonnet 5 hasn’t landed a merged PR yet, since it only just launched, but it’s already showing up on PRs out for review, which tends to be the leading indicator before it shows up in this table. And the autonomous workflow doesn’t relax the one constraint that’s held the whole way through this post: a person still has to trigger it, via a Linear ticket or an @-mention. Our human names have stopped showing up on the commit messages. The only adoption ceiling is still how many engineers choose to reach for it.

    AI-attributed PRs per week by model

    The succession, extended through June. Each Opus release displaces its predecessor in committed work within about a month of arriving; the February pattern wasn’t a one-off.

    Possible (overlapping) explanations

    We can name several factors that line up with the inflection. None of them, on its own, explains the curve, and we can’t isolate which mattered most without a controlled test we can’t run in retrospect. What follows is a list of overlapping contributions our own team has flagged, not one single cause dressed up as several.

    Frontier-model capability: Opus 4.6, in February 2026, was the specific release that produced the session-rate jump. Earlier Opus releases (4.1 in August 2025, 4.5 in December) shifted the floor without producing a takeoff. The rest of the model family in the same window, Sonnet 4.6 and Haiku 4.5, didn’t move the committable-delegation needle in our telemetry: Sonnet 4.6 launched February 17 with full Honeycomb adoption (41 distinct users) but produced only 35 commits in March and 27 in April, against Opus 4.6’s 232 and 127. It was frontier-model capability specifically that crossed the threshold from consultation to delegation, not “any new model.”

    Leadership signaling, twice: August 2025’s letter set the explicit “experiment, take time to figure out what AI does for your work” frame. January 2026’s follow-up named sharper strategic stakes: rebuild the product surface and the market posture for AI-first. The mid-January realignment toward greenfield, higher-AI-leverage work followed from that. Without either signal, I doubt the throughput curve looks the same.

    Tooling that improved over months: The October 2025 Claude Code harness improvements produced a noticeable adoption step on their own, with no model release behind them. Each release of the tooling added something some engineer needed before they could delegate. Models alone don’t explain the curve; the wrapper around the model matters just as much as the model does.

    Substrate already in place: Continuous delivery, code-ownership practices, fast CI, blameless incident analysis, observability that links shipped code back to the PR that created it: the practices the rest of this post is about. AI work landed in an org that already had the substrate to absorb it. We can’t run the counterfactual, but the practices section below is our best account of why this didn’t go badly.

    Cumulative engineer-level expertise: From April through September 2025, one or two engineers a month adopted Claude Code. By February 2026 we had a base of fifty-plus engineers who’d been using it for months and could mentor everyone else. That cohort effect is hard to pin to any single date; it’s the gradual accumulation of in-house expertise that the February takeoff drew on.

    These factors overlap, and we can’t say any one of them in isolation would have gotten us here. Leadership signaling probably amplifies tooling improvements. Substrate makes it possible for engineers to share what they’re learning. Frontier-model capability matters more in an org that already has the substrate and the cohort to use it well. The factors compound rather than substitute for one another. Anyone telling you a single factor caused their team’s gain is overstating it. Including us.

    We followed Intercom, with caveats

    Three things are worth crediting in Intercom's playbook.

    They set a more realistic and achievable 2x goal, not the 10x-and-up claims that were common currency during the 2024 hype cycle. Discipline matters when the temptation is to oversell what AI delivers. We wanted some of that discipline for ourselves.

    They were transparent about both the methodology and the org-shape changes that came with the gain. The substrate Intercom built, a Claude Code plugin marketplace with 153 contributors representing 31% of their R&D org and 267 skills, is genuinely platform infrastructure that ships its own product. They spun up a dedicated team, team-2x, to build it. We’re smaller and a couple of years younger, but we’re building toward something in the same shape, at our scale.

    They engaged their auditors, Schellman, early, before scaling auto-approval, to confirm that the evidence trail an AI-approved PR produces is the same evidence trail an auditor expects from a human-approved one. The “who” changes. The “what” doesn’t. That’s a model worth following: build for safety first, and compliance follows from it.

    Where we differ is that we’re at zero auto-approval today, and they’re at 19.2% as of April. That’s deliberate sequencing on our part, not a sign that we're “behind.” Before you can safely scale auto-approval, which is automating the bug-catching half of PR review, you need substantial substrate underneath it. The compliance question is necessary but not sufficient; the technical preconditions sit underneath the procedural ones. You need codified rules in CLAUDE.md and skills that an auto-review agent can actually verify against; MCP-mediated dev-loop access so the agent reviews against the same context a human would have had (design intent, ticket history, production behavior); fast and AI-legible CI so the verification loop closes quickly; closed-loop production observability that links shipped code back to the PR that created it; and a dissemination layer so humans stay aware of what’s shipping even when they’re not gating it.

    A 19% auto-approval number means radically different things at an org that invested in the substrate before turning the switch versus one that just turned the switch. Intercom’s number is downstream of substrate they built first. A company that turns on “auto-approve PRs under 20 lines” without equivalent substrate will report the same number, but it’s measuring rubber-stamping against weak constraints, not safe automation against strong ones. Different point on the same trajectory.

    The headline finding from Intercom’s auto-approval post is the most striking parallel to our own data. They report AI-authored backend code reverting at 0.53% and AI-authored frontend code reverting at 0.22%, against human-authored revert rates of 5.39% and 2.00% respectively. It’s worth naming the selection effect here: if the easier, lower-risk changes increasingly get auto-approved, humans are left reviewing the harder residual cases, which would push human revert rates up for reasons that have nothing to do with humans getting worse at their jobs. Their downtime from breaking code changes dropped 35% even as deployment frequency doubled. Theirs is a per-PR claim about strict, decomposed, sub-agent-driven review against an Intercom-specific guidance flywheel. Ours is a per-quarter claim about severity, not volume: incident count is tracking change volume about as linearly as you’d expect, but we haven’t had a spectacular AI failure, and our guardrails are aimed at containing how bad any one incident gets rather than pretending we can hold the count flat. Different denominators, same direction. Both posts are pushing back on the naive “more code, more failures” intuition, from different evidence.

    We had a private conversation with the Intercom team in early May 2026. They hit the same February 2026 inflection point we did. The Opus 4.6 unlock was the gut-call attribution from multiple Intercom engineers, though they noted it was almost impossible to disentangle from their internal mandates and team-2x activity in the same period. They’re partnering with a Stanford research group to try to isolate the variables, and even that group’s initial pre-January analysis missed the inflection entirely. The world’s most data-rich org on this exact question is paying academic researchers to help them figure it out, and they still don’t know.

    That independent-org corroboration is the cleanest natural experiment either of us has. Two separate orgs, same month, same uncertainty about cause, same direction. That’s worth more than either of our individual analyses on its own. It’s also worth weighing against the broader base rate: DX’s longitudinal study across roughly 400 companies found AI usage up 65% translating to only about 8% more PR throughput on average. We’re the outlier case here, not the median one.

    Here's the truth: Nothing I've written here will help you if your underlying org isn't already healthy and functional. AI just amplifies what you're already doing. Read part 2 of this blog series to see what we learned.

    AI Influence Level disclosure

    This post: AIL-3.0 (substantial AI involvement, human steering on every load-bearing call). The data analysis, custom git-of-theseus extensions, commit-history trawling, calibration overrides, the merge-rate and incident charts, was AI-assisted. The July refresh, the May-June numbers and the autonomous-workflow analysis added above, was pulled with Claude Fable 5. The slide deck this post derives from was composed with AI assistance, and AI assistance was used to reformat the slide bullet points and speaker notes into essay form. The voice and tone polish was done in a separate Claude project tuned to my writing style, followed by a very extensive manual editing process where I further added or changed at least 20% of the words.

    Strategic decisions, data interpretation, and judgment calls about what to keep and what to cut are mine. AI helped me move faster on a deadline; it didn’t supply the substance. The talk and this post are themselves an example of the same 2x story they describe: weeks of human work, AI-assisted, not 10x. The substance wouldn’t exist without me, and the level of polish wouldn’t exist without AI.

    AIL framework: danielmiessler.com/blog/ai-influence-level-ail. Illustrations in the original talk: AIL-0, by bbghost.bsky.social. Art should be made by artists, not machines. Illustrations in the blog by our amazing design team.

    Sources and context: Fin/Intercom 2x post; Fin/Intercom AI PR approval safety post; Honeycomb-Intercom case study; Emily Nakashima on AI-amplified engineering leadership. This post is adapted from the talk of the same name, delivered at Sydney Tech Leaders and LDX3 London in 2026.

    \ No newline at end of file diff --git a/sreweekly/articles/528/06-you-ve-just-had-an-incident-what-next.html b/sreweekly/articles/528/06-you-ve-just-had-an-incident-what-next.html new file mode 100644 index 00000000..0a1a6a46 --- /dev/null +++ b/sreweekly/articles/528/06-you-ve-just-had-an-incident-what-next.html @@ -0,0 +1,198 @@ +You've (Just) Had an Incident. What Next? | Uptime Labs + + + + + +

    You've (Just) Had an Incident. What Next?

    Karan Nagarajowda
    |
    Tags:
    Blog
    Incident Management
    IN THIS ARTICLE

    Ready to make incident response your competitive advantage?

    See how Uptime Labs builds provable, scalable incident response capability across your organisation.

    I've run enough major incidents to know that the first hour rarely goes the way people expect. Here's what I've learned about what actually happens inside a team when something breaks, how teams often reflect, and how we think about accountability without losing a blameless culture.

    The first hour of an incident: what's really happening

    Incident timeline illustration reading detect - assess - declare & categorise - communicate - diagnose - resolve - review

    Logistically, after detection, people try to understand what's actually happening, then work out whether how to move forward. Note that incidents do not progress through stages in a linear way. You may come back to assessing severity or revisit assessment based on new information that emerge as incident progresses.

    Similarly, it's also complex from an emotional perspective. Individual engineers often wonder if they broke something. Some people see a symptom and immediately form a theory i.e. a gut instinct about the cause - and then start hunting for evidence to support that theory rather than staying open-minded. Meanwhile, managers want constant updates, and what actually happens in Slack is information overload.

    But in my view, the real challenge in the first hour isn't technical; it's cognitive overload. If you divide an incident into stages, the first part of the time should be owned by the incident commander, and their job isn't to solve the problem. It's to reduce chaos: create the space for people to pick up tasks and start investigating, make decisions explicit, and keep communication flowing. Those first minutes - or even hours, in a bigger incident - are never about fixing the system. They're about making sense of a surprising situation, recruiting people who can help and creating sufficient psychological safety that people will voice their theories and be explicit about certainty of premises of the theory. That's what separates a good incident commander.

    A word about pressure

    I think back to a story from earlier in my career that says more about pressure than any framework could. I was working through a major incident next to a colleague who was technically very strong. Our manager - who was visibly under stress - came up behind us and demanded we check the logs. The pressure of being watched and barked at was so distracting that my colleague forgot the basic syntax of the view command. Our manager ended up spelling it out: "v–i–e–w, and then the file name."

    It's a small moment, but a telling one. Technical skill isn't the bottleneck when the room is hostile. A junior engineer can have all the right preparation, the right runbooks, the right paired senior and still fold if the environment around them is built on pressure and intimidation rather than support.

    What teams could overlook after

    People naturally want to answer "why did this happen?" Uncertainty is uncomfortable, and immediately after an incident there's a lot of it. So people jump to ‘why’, but I think that's the wrong question, because it sends you straight towards root-cause analysis.

    As we explored in The Technical Foundations of Incident Response, the right sequence is: can we stop the customer impact - stop the bleeding? Then, can we stabilise the system? Then, can we preserve the evidence? Only then do you investigate - that's where the actual learning happens.

    However, when a team implements a change and encounters errors, fixation can quickly set in. They observe memory errors and become convinced it's a memory leak, focusing all investigation efforts on proving that diagnosis. The problem is that this fixation causes them to lose sight of the primary incident response objective: restoring service. During an incident, multiple decision paths are available (rolling back the change, restarting the service, modifying configuratio) each with different risk and learning profiles. By pursuing only one diagnostic avenue in parallel, teams sacrifice their ability to understand what actually happened. Gathering information before committing to a single response path, and considering what each approach might reveal, is as important as speed in bringing systems back online.

    The second accident is making several big decisions in parallel: restarting, scaling up, changing configuration, with different people acting independently. I wouldn't say the process needs to be strictly sequential, but it makes it much easier to asses impact of each change if you try one thing, gather information, and only then move to the next. Doing three or four things in parallel might bring the system back faster, but it destroys your ability to learn afterwards what actually happened. Preserving the evidence is just as important as restoring the service.

    Accountability without losing a blameless culture

    A ‘no-blame’ culture in incident management should not mean removing human accountability or glossing over the decisions people make during crises. Rather, it means resisting the urge to simply label events as ‘human error’ and moving on. Every incident involves decisions made by individuals who bear responsibility for those choices, but understanding why they made them is critical.

    The goal is to reconstruct the context in which decisions were made: what information was available at the time, what pressures and constraints existed, and what risks seemed apparent or hidden. When we skip this deeper analysis to avoid blame, we sacrifice the insights that could prevent future incidents. Accountability and learning are not opposites; they work together when we focus on understanding the decision-making environment rather than punishing the decision-maker.

    A good postmortem embraces the human element; one of the the key skills of Postmortem (incident review ) is to conduct in a way that no one is uncomfortable. Part of this means avoiding reducing a complex failure down to one person's mistake. That's what a blameless culture actually means: not pinning a complex failure on one person or team, as we discuss in our regulatory incident response piece. Accountability, on the other hand, is about improving future outcomes, not assigning guilt.

    So in a postmortem, the questions focus should be on, "What set of circumstances led to the incident? What information was available to the human operator at the point of the decision making? Why that decision made sense to them?" .

    Definitely not "who did it, or why did they do it?" If someone skipped a checklist and that caused a major issue, the question isn't "why did they skip it" - it's "was there pressure that made skipping it possible? What barriers should have stopped that?" It's about whether the system allowed the mistake, not about the person. Accountability, meanwhile, is about owning the future outcome - again, without assigning guilt.

    Who owns the learning?

    It's the incident commander's responsibility to make sure the learning happens, but I don't own all of the learning myself. The engineering team should own the technical timeline. Monitoring and operations should look at the alerts and the response coordination. Customer support should explain the customer experience during the incident. My job as incident commander is to make sure all of that comes together.

    One of the most valuable questions I ask isn't "what failed?" It's "what made the incident harder to resolve than it needed to be?" That's the better question, and it usually surfaces poor documentation, confusing dashboards, unclear ownership, missing alerts and communication gaps. That's what comes out of these discussions.

    Add Uptime Labs to your post-incident learning

    None of this is instinctive. Reducing chaos before fixing the system, resisting the pull towards "why", asking what made an incident more difficult to resolve rather than who's to blame - these are habits people get better at by practising them, not by reading about them and hoping they’ll stay in your head while everything is on fire.

    That's exactly why I helped build Uptime Labs. Our drills put your team through the same ambiguity, incomplete information, and communication pressure a real incident throws at you, minus the real-world stakes. You find out how your team actually behaves in that first hour before it costs you a customer. Each drill comes with a personalised report designed to support ongoing skills development.

    If you want to see what that looks like in practice, try a drill or get in touch to book a demo.

    Karan Nagarajowda
    Share this post
    Clear blue sky with scattered white clouds.

    Ready to make incident response your competitive advantage?

    — Chris Voss

    See how Uptime Labs builds provable, scalable incident response capability across your financial services organisation.

    + + + \ No newline at end of file diff --git a/sreweekly/articles/528/07-an-sre-response-to-datadog-s-state-of-ai-engineering-2026.html b/sreweekly/articles/528/07-an-sre-response-to-datadog-s-state-of-ai-engineering-2026.html new file mode 100644 index 00000000..57dfe6e7 --- /dev/null +++ b/sreweekly/articles/528/07-an-sre-response-to-datadog-s-state-of-ai-engineering-2026.html @@ -0,0 +1,3019 @@ + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + An SRE Response to Datadog's State of AI Engineering 2026 + + + + + + + + + + + + +
    +
    +
    +
    + + + +
    + + +
    + + + + + + + + + + +
    +
    +
    + +
    +
    + +
    +
    +
    +

    Could your team report a vulnerability within 24 hours? Find out on September 23.

    +
    + + + +
    +
    +
    + +
    + +
    +
    +
    +
    +
    +
    +
    + + +
    +
    + + + + + +
    +
    +
    + + + +
    +
    +

    Agent Sprawl Is Your Next Production Incident: An SRE Response to Datadog's State of AI Engineering 2026

    +
    + +
    +

    Datadog published the State of AI Engineering 2026 report. Read it. It's the most comprehensive look at AI in production available now.

    +
    + +
    + By  + + + · + Analysis +
    +
    +
    + +
    + + +
    + + + Comment + + +
    + +
    +
    + Save +
    +
    + + + + + +
    +
    + + 3.8K Views +
    +
    +
    + + +
    + +
    + +
    +

    Datadog published the State of AI Engineering 2026 report— real telemetry from over a thousand production environments. Read it. It is the most comprehensive look at AI in production available right now.

    +

    I want to respond from the reliability engineering perspective, because the data reveals a problem the report names but doesn't fully resolve: agent sprawl is now a production reliability crisis, and the SRE discipline does not yet have governance frameworks for it.

    +

    What the Data Shows

    +

    Three findings stand out from an SRE perspective:

    +

    Framework adoption doubled year over year. LangChain, LangGraph, Pydantic AI, Vercel AI SDK — up from 9% of organizations in early 2025 to nearly 18% by 2026. Services using agentic frameworks: more than doubled.

    +

    70%+ of organizations run three or more models. The share running more than six models nearly doubled. Teams are building model portfolios rather than committing to a single provider.

    +

    Teams add models faster than they retire them. Datadog calls this "LLM tech debt." Each overlapping model introduces its own quality, latency, and cost profile. The report is explicit: this becomes a governance problem.

    +

    These three findings combine to describe an environment growing faster than it can be governed. I call this Agent Sprawl.

    +

    Defining Agent Sprawl

    +
    +

    Agent Sprawl — the condition where AI agent infrastructure complexity (frameworks, models, tool layers, orchestration patterns) grows faster than your ability to measure and govern its reliability.

    +
    +

    It is structurally identical to the microservices sprawl problem SRE teams faced between 2015 and 2020. Teams added services faster than they added SLOs. The result: production incidents nobody could attribute because the dependency graph was too complex to observe.

    +

    Agent Sprawl has three specific manifestations:

    +

    1. Framework-Invisible Call Complexity

    +

    When you add LangChain, LangGraph, or any orchestration framework, it adds steps and paths you did not write — retry logic, fallback handlers, context window management, tool routing. All of this happens between your application code and your observability layer.

    +

    Your SLIs measure at the application boundary. Framework-added calls are invisible.

    +

    This means your Tool Invocation Efficiency (TIE) baseline — tool calls per task completion — is measuring a mix of your agent's behavior and your framework's behavior. When you upgrade the framework, both change simultaneously. You cannot separate them.

    +

    In practice, across regulated production environments I've studied, TIE baselines can drift 30 – 40% after a framework major version upgrade with no corresponding change in the agent's task logic. The baseline shift looks like agent degradation. It's actually framework overhead. Teams spend hours on a false RCA.

    +

    The fix: Instrument at the framework output layer, not the application layer. Capture tool invocations after framework processing. Then freeze your TIE baseline before any upgrade and compare shadow traffic before promoting.

    +

    2. Multi-Model SLO Orphaning

    +

    70% of organizations running 3+ models means 70% have at least two additional SLO ownership gaps they haven't acknowledged.

    +

    SLOs are set once — typically when the first model is deployed. As models 2, 3, 4, 5, 6 are added for specific task classes, latency profiles, or cost tiers, nobody revisits the SLO ownership model. Models run in production with no named owner, no baseline, no error budget.

    +

    When model 3 degrades, there is no owner to page, no baseline to compare against, no runbook to execute. The degradation surfaces as a customer complaint, not an alert.

    +

    The fix: Treat every model in your fleet like a microservice. Each model gets: a named owner (not a team — a person), a task-class-specific SLO, and a 30-day observation baseline before the SLO is enforced.

    +

    3. LLM Tech Debt as a Reliability Liability

    +

    Deprecated models running in agent chains create silent compatibility risks. When a provider announces deprecation, teams with models buried inside multi-step chains often miss the migration window. The model ages. Safety training falls behind. Decision Quality Rate declines slowly — too slowly to trigger a threshold alert — until accumulated drift surfaces as a production incident.

    +

    The fix: Treat model deprecation notices the same way you treat dependency CVEs. Automate alerts at 60, 30, and 7 days before end-of-life. Build the migration ticket at announcement time, not at expiry.

    +

    The Governance Framework Agent Sprawl Needs

    +

    The Agent Fleet Inventory

    +

    Before you can govern sprawl, you need to know what you're governing. Maintain a living inventory with, for each component: framework and version, model(s) used, task classes handled, named SLO owner, current TIE/DQR baselines, and deprecation dates.

    +
    +
    +
    +
    + Python +
      +
    +
    +
    from agentsre.sprawl import AgentFleetInventory, FleetComponent, ComponentType
    +
    +inventory = AgentFleetInventory()
    +inventory.register(FleetComponent(
    +    component_id="anthropic.claude-sonnet-4-6",
    +    component_type=ComponentType.MODEL,
    +    agent_id="payment-processor",
    +    task_classes=["payment-routing", "fraud-detection"],
    +    slo_owner="[email protected]",                    # named human — not a team
    +    baseline_established_at="2026-04-01",
    +    deprecation_date="2027-06-01",
    +    last_slo_review="2026-04-01",
    +    current_tie_baseline=2.4,
    +    current_dqr_baseline=91.2,
    +))
    +
    +report = inventory.quarterly_review_report()
    +print(f"Fleet governance score: {report['fleet_governance_score']}/100")
    +
    +
    +
    +


    +

    Framework Version Governance — Canary Before Promotion

    +
    +
    +
    +
    + Python +
      +
    +
    +
    from agentsre.sprawl import FrameworkVersionGovernance
    +
    +gov = FrameworkVersionGovernance(
    +    tie_drift_threshold=1.15,   # block if TIE drifts >15%
    +    dqr_drift_threshold=0.85,   # block if DQR drops >15%
    +    min_shadow_samples=50,
    +)
    +
    +# Before upgrade: snapshot production baseline
    +gov.snapshot_baseline(
    +    agent_id="payment-processor",
    +    task_class="payment-routing",
    +    framework_version="langchain-0.2.x",
    +    tie_values=production_tie_samples,
    +    dqr_values=production_dqr_samples,
    +)
    +
    +# After 48hrs shadow traffic:
    +result = gov.evaluate_upgrade(
    +    agent_id="payment-processor",
    +    task_class="payment-routing",
    +    production_version="langchain-0.2.x",
    +    shadow_version="langchain-0.3.x",
    +)
    +
    +if result.decision == UpgradeDecision.BLOCK:
    +    rollback()   # framework added hidden overhead — don't promote
    +
    +
    +
    +


    +

    The Quarterly Multi-Model SLO Review

    +

    The review should take 30–60 minutes per quarter. For every model in fleet:

    +
      +
    • Verify named owner exists
    • +
    • Verify baseline is current (< 90 days old)
    • +
    • Check deprecation schedule against provider announcements
    • +
    • Review TIE per-model — models with rising TIE relative to task class baseline are drifting
    • +
    +

    Models scoring below 70 on the governance health score are flagged as governance debt requiring a 30-day remediation window.

    +

    The Datadog Report's Implicit Challenge

    +

    The State of AI Engineering 2026 describes an industry in rapid expansion. What it does not fully resolve is the SRE question: who governs all of this, and what does that look like in practice?

    +

    The SRE community has solved exactly this class of problem before — in distributed systems, in microservices, in cloud infrastructure. The discipline already exists. It needs to be applied to the AI agent layer now, before agent sprawl becomes agent chaos.

    +

    The Datadog data tells us the window is closing. Framework adoption doubles in a year. Multi-model fleets become the norm. Model debt accumulates.

    +

    Build the governance layer before the production incidents start.

    +

    Resources

    + +

    What's your biggest agent sprawl challenge right now?

    +
    + +
    + + +
    +

    Opinions expressed by DZone contributors are their own.

    +
    +
    +
    +
    + + + +
    + +
    +
    +
    + +
    + + + +
    +
    +
    +
    +
    +
    +
    + +
    + × + +
    + + +
    +
    + +
    +
    +
    +
    +
    + + +
    + + + + + +
    + +
    + + + + + + + + + + + + + + + + + + + + + + diff --git a/sreweekly/articles/528/08-don-t-add-a-read-replica-until-you-ve-read-this.html b/sreweekly/articles/528/08-don-t-add-a-read-replica-until-you-ve-read-this.html new file mode 100644 index 00000000..c57881ff --- /dev/null +++ b/sreweekly/articles/528/08-don-t-add-a-read-replica-until-you-ve-read-this.html @@ -0,0 +1,37 @@ +Don't add a read replica until you've read this | Blog | incident.io

    Don't add a read replica until you've read this

    July 21, 2026 — 23 min read

    As the size and complexity of their relational database workload grows, every company eventually goes through the process of off-loading work on a read replica. It comes with lots of benefits, but at a cost of increased complexity. This article is about how we dealt with that, a lot of learnings, and some useful techniques.

    incident.io is the software reliability platform built to investigate, respond and prevent incidents, powered by AI that deeply understands your organization. Thousands of customers rely on it to be the thing that supports them through anything from a minor blip to a full outage. Any disruption to that service has a significant impact on those users, and that’s always top of mind for our engineering team. Everything we build is designed to be performant, reliable, and gracefully degrading.

    When large parts of the internet goes down, as they did for the AWS outage on 20 October 2025, we see a meaningful increase in alert volume as engineers across the globe are getting woken up. During events like this, we simply can’t fall over from the increased load. This means we always have to run with significant spare capacity. And so we’re always looking for opportunities to reduce the load on our DB, either through performance improvements or code redesign. While working on projects to control the resource utilization on our primary database, we clearly identified that we, like a lot of online services, run an overall read heavy workload.

    We already had a read replica set up and some queries were already utilizing it, but it was something that we were doing on a case by case basis. We knew we could do more. We set ourselves the goal to move everything over to the read replica. For every query that could be move moved to the read replica, that’s additional capacity for our primary, and additional protection for those events of massive traffic that we design for.

    The benefits of read replicas

    But before that, let’s do a quick recap of why you’d want to introduce read replicas into your stack. They do bring complexity, having two databases and two database connection pools in your app logic is more to think about than just having the one. You also need to manage two things, where both often quickly become critical to running your service. Double the graphs and metrics and warnings to worry about. Not to mention, you’re now paying for two databases.

    They bring a lot of benefits though. They’re the natural step to take to give people access to run operational queries on the production system, without risking those queries disrupting the production workload. A query run on the primary can cause overload, delays, contention, locks, and much more. On the read replica the blast radius is much smaller, although there are notable exceptions to this like hot_standby_feedback. But maybe the most interesting thing is that it opens the door to horizontal scaling. Relational databases generally don’t scale horizontally, just vertically. Writing to two primary databases is slower than to one, and for most of us spending more to get less performance is not a very interesting prospect. But read replicas are different, since they don’t need to coordinate writes, you can just have more of them and load balance read queries across them. That’s pretty cool.

    Read-after-write consistency

    Even before approaching this larger re-think we had already established some useful primitives. Our backend language is Go, but the basic techniques translate to some degree to any language.

    The number one thing you need to deal with as you’re migrating work over to a read replica is read-after-write consistency. Or in other words, avoiding stale reads. The gist of it is, if you write something to the primary and then immediately after read the thing back but from the replica, there’s no guarantee that you get the same thing back. This gets worse the faster you read after writing. Now if you’re hand-crafting some beautiful artisanal code in your walled garden project, you can probably attempt explicitly picking the primary or read replica for each individual query through a code flow. But for practical reasons we often end up just limiting the read replica to the queries and code paths where we feel that nothing can go wrong.

    That’s not what we’re looking for, we want to move everything over. That means we need automated detection of mutating queries, and to automatically fall over to the primary after a write has been detected to ensure that you are able to read your own writes. The basic mechanism we introduced for this uses the Go context to carry a special flag that controls whether we can use the read replica. It defaults to true.

    type taintedKey struct{}
    +
    +// A write taints the context: reads on a tainted context must
    +// go to the primary, since the replica may not have caught up.
    +func Taint(ctx context.Context) context.Context {
    +      return context.WithValue(ctx, taintedKey{}, true)
    +}
    +
    +func CanUseReplica(ctx context.Context) bool {
    +      tainted, _ := ctx.Value(taintedKey{}).(bool)
    +      return !tainted
    +}
    +

    The second concept we introduced was a simple function for detecting whether a given query was “safe” to go to the read replica. Initially we designed this to be conservative, we were ok with some things being misinterpreted and going to the primary, as long as we don’t get accidental and weird bugs where we send the wrong query to the wrong place. Some heuristics got us a long way, like whenever we start a transaction, we send it to the primary and mark the ctx so that any subsequent queries also go to the primary, with the basic assumption that anything in a transaction is probably something you want on the primary (this turns out to not be quite true, more on this later).

    // Transactions probably mean writes: run on primary,
    +// and taint the ctx so everything after follows it there.
    +func (db *DB) Transaction(ctx context.Context, fn func(context.Context) error) error {
    +      ctx = Taint(ctx)
    +      return db.primary.Transaction(ctx, fn)
    +}
    +

    The second heuristic was to strip comments and trim whitespace, then check if the query starts with SELECT. Not very elegant, but it got the job done. There’s a fancier version further down!

    func IsReadOnly(query string) bool {
    +      q := strings.TrimSpace(stripComments(query))
    +      return strings.HasPrefix(strings.ToUpper(q), "SELECT")
    +}
    +

    On top of this we built a basic layer on top of our database handles that created both primary and replica transaction pools and transparently switched between them, using the flag we had previously set. This means it was now generally safe for us to just opt in any code path to use the read replica, trusting this mechanism to direct the queries to the right place and maintain read-after-write consistency.

    func (db *DB) route(ctx context.Context, query string) *sql.DB {
    +      if !IsReadOnly(query) {
    +              return db.primary // and taint the ctx here
    +      }
    +      if !CanUseReplica(ctx) {
    +              return db.primary // we wrote earlier, stay consistent
    +      }
    +      return db.replica
    +}
    +

    Additionally we put a circuit breaker in here. If we’re not able to open connections to the read replica we give up and send all queries to the primary. Note that although we’re happy to have that fail over behavior for now, with enough total load across your databases, your primary might not be able to handle this.

    Event handlers

    incident.io is designed from the ground up on event publication and subscription, the entire system is built around publishing messages and having a fleet of workers processing them. This gives us all kinds of interesting super powers, including the ability to buffer up messages to process them later during periods of overload.

    The workers is exactly where we had been adding some read replica offloading, more specifically in a set of workers that execute low priority workloads that do read heavy work. This is also where we started our work, primarily by moving more and more slices of our workers over to the read replica. Our initial results were really positive, we were almost being too successful and we quickly had to upgrade our read replica because it was taking over so much work from the primary. After having moved some of our heaviest read query workloads, congratulating ourselves for our great work, we noticed something odd. Not a super clear pattern, but the occasional NotFoundError coming out of our database adapter in the subscribers on the workers that we had just moved.

    And that’s when it struck us. We were ensuring read-after-write consistency within the context of a go ctx, but what about across the boundary of publishing and processing a message?

    Let’s say we have a request come in, to create a post-mortem document. But the post-mortem document itself can take a while to create and involves hitting a separate internal service over the network, we don’t do all of that work in line with the client waiting. So we write the row to the database, respond immediately to the client, and then publish a message to a worker to deal with actually creating the document. Incredibly, what we were seeing was the time between publishing a message to the message queue and the worker pulling that message to process it be shorter than the replication lag to the replica. This wasn’t because the replica was falling behind, it was doing just fine, it was just that occasionally the event was processed incredibly quickly. Overall it was a rare occurrence, something between 0.1 and 0.5% of messages, and they were automatically retried, but obviously not something we wanted to keep getting alerted on. Interestingly we had technically had this issue for a few code paths for a while, but it wasn’t until we wholesale moved large parts of our codebase over to the read replica that this problem showed itself.

    We discussed some options and quickly homed in on a well-documented tool in the PostgreSQL world.

    Log sequence numbers to the rescue

    The log sequence number, or LSN, is a special pointer in PostgreSQL that represents a specific position in the Write-Ahead Log. This means that you can use it to check whether the read replica has caught up with a specific write in the primary database. You get the LSN with a built in PostgreSQL function:

    SELECT pg_current_wal_lsn()::text
    +

    The basic technique we want to apply:

    1. Grab the LSN from the primary after having done your write
    2. Pass it to the code running on the worker
    3. Compare the read replica LSN with the LSN from the primary, if the read replica is at or after that LSN, you know that you are past the point where your write happened

    In order to enforce read-after-write consistency across our workers, we started stamping each published message with the LSN from the primary at the time of publishing. Grabbing the LSN is very cheap, it’s just an in-memory operation on the PostgreSQL side.

    Having rolled that out, every message now has this stamp on it. So then we added a check on the subscriber side that compared the LSN on the message with the LSN of the read replica. The LSN comparison can be done inside of a SQL query using the native data type:

    SELECT pg_last_wal_replay_lsn() &gt;= $1::pg_lsn
    +

    If the replica was caught up, we are fine to process it. If it isn’t, we just kick the message back to the queue.

    That last part actually turns out to be a great back pressure mechanism in overload situations. If the read replica is failing to keep up with the primary, for whatever reason, instead of directing all queries to the primary and risking overloading it, we just keep nacking the messages until the read replica is feeling better again, with exponential backoff to avoid overload. It turns a potential outage situation into a graceful degradation instead, all the work is processed as expected, just with a delay.

    If you’re doing this, you want to couple it with some dashboards and alerts. You don’t want to get caught out on the replica lagging behind. We often see people measure read replica lag in time, but that’s not a useful measurement. A read replica can catch up 5 minutes of delay in seconds because the rate of change has slowed down, or it can struggle to close the distance on 30s of delay because the rate of change is going up. Basically, the replay rate can change. A more useful measurement is the number of bytes behind. It still doesn’t quite convey whether you have a problem or not, but it’s more consistent than measuring it as time.

    No more random NotFoundError, great success!

    Stamping LSN on users

    After having tackled the events and subscribers, we looked around for other places where we would expect to bump into the same consistency problem, where we need to be careful to read our own writes, and we identified the API as another surface area to deal with. This included our public API where we support creating resources, returning an ID, and then reading the resource back. Done quickly you’d risk a 404. Additionally, our dashboard relies on our internal private API and we frequently use the same pattern there: create an empty resource and return the ID, queue the work to actually create the resource, frontend polls the API using the ID, resource is eventually created and displayed in the UI.

    There are a few different classic blog posts that talk about using LSN stamping on API requests, so this is not exactly new territory. We want to apply the same principles as we did for our events, but we can be a bit more elegant about it for the API. Thinking about the problem we’re solving, it’s basically focused on the case of a POST followed by a GET. Or more generally, a mutating request followed by a read. Since we’re pretty consistent about HTTP verbs in our API, we can actually limit LSN stamping to mutating requests: POST, PUT, PATCH, DELETE. Additionally we can limit our scope to a given actor, whether it’s a user, an API key, or something like our mobile app, as the domain of consistency. If I create a document I want to see it immediately, but it’s ok for my coworker to have a few milliseconds delay to see the same document.

    We added a middleware that runs after the request handler in our Goa web layer, checks the method of the request, and optionally stamps the actor **with the LSN. We store this directly in PostgreSQL, using the native data type.

    UPDATE users SET read_after_write_lsn = GREATEST(read_after_write_lsn, pg_current_wal_lsn()) WHERE id = ?
    +

    This has the benefit of us already loading that row anyway during auth at the start of every request, so we could include this column there. This is also why we chose not to move this work to a key value store like Redis. Redis would add a network request, looking up the LSN, to every single incoming request, and we still need to get the latest LSN from replica too.

    We shipped the middleware and with our actor tables quickly filling up with LSN stamps, each representing the last action each one has taken, we tackled the other half of this. We added another middleware that runs after auth and takes the LSN stamp from the actor and compares it with the read replica. Unlike subscribers, we can’t just nack the message and send it back to the queue to be processed later. People tend to not want to have to retry all their requests, and we didn’t want to push this on every client to our system, of either having to retry on a certain status code, or carry LSN stamps on headers. So instead of nacking, we fall back to the primary. We may need to rethink that at some point, when we can’t afford to go to primary anymore, but after rolling this out the actual impact on the primary is minimal. Only about 0.01% of requests to our API fail the LSN check.

    With this we could move the last part of our codebase over to the read replica, reversing the trend of increased CPU utilization on primary week over week and getting it back down to a healthy level with all that extra capacity that we aim for.

    Optimizations

    After having rolled this out we went after two optimizations that we had identified while we were working on this project. Although fairly simple, we didn’t implement them before rolling out because we 1. wanted to avoid premature optimizations, and 2. doing it after means we get a pretty graph and clear validation of the efficacy of the optimization. Everyone loves a pretty graph.

    The first one was to limit the number of nacks on the event subscribers. As mentioned before, we saw between 0.1 and 0.5% of processed events getting sent back to the queue due to replica lag. Each nack causes a delay of at least 10 seconds in processing the message, since that’s our default retry delay. That’s very much acceptable, but we had a theory: that when we had replica lag the lag is almost always very small.

    So whenever we check the replica and it’s behind, we added a 100 ms sleep before trying again. The impact was striking, almost eliminating nacks. Over time we saw the nack rate drop to ~0.03%. For what was basically a couple of lines of code. In real terms we’re not talking about a lot of messages, but it was a worthwhile improvement. PostgreSQL 19 actually comes with this functionality built in, with the new WAIT FOR LSN.

    Secondly, when investigating why certain queries, that were actually perfectly safe to run on the read replica, were getting directed to the primary we realized our naive approach to detecting mutating queries was maybe just a little bit too naive. We were missing out on tons of juicy queries that were absolutely eligible for the read replica. Taking inspiration from the https://github.com/pgplex/pgparser library, we created a small tokenizer to replace our heuristics. Armed with the tokenizer, we could now confidently identify any queries that were safe to run on the read replica. When we rolled it out we saw a massive drop in CPU utilization on the primary, and a corresponding increase on the replica side.

    Database connection pool routing strategies

    So now we’re in our new better world where everything is opted into the read replica by default, we have read-after-write consistency, and thanks to LSN stamping, that applies across events and subscribers, and the API as well.

    But we had two remaining concerns: firstly that our code was still explicitly opting into the read replica everywhere, and secondly that internal concerns were leaking: anyone who needed specific database pool behaviors had to understand exactly how all of this worked. We didn’t want to have to maintain this for the team permanently, or introduce friction for everyone. So we went after the last major part of the project: inverting the default and creating a “DSL” for choosing database connection pool routing strategies. Everything goes to the read replica, whether you’re aware or not, and the mechanisms we’ve implemented ensures your code just works.

    The easy part was tweaking our database handles with internal pool routing, we changed it to be enabled by default. The bigger part was giving the engineering team the tools to tweak the behavior where they needed to. We identified four strategies:

    • ReadAfterWrite - this is our default behavior. Track mutations and route to primary after.
    • StaleRead - where you have reads after writes, but you’re fine with them going to the read replica.
    • Primary - pin the context to the primary and send all queries there. We use this where we can’t tolerate any kind of lag, primarily around our on-call product.
    • Replica - pin the context to the replica and send all queries there. This is kind of a nuclear option, it will force queries to the replica even when the replica can’t execute them. Useful where you’d rather fail loudly and fix your code.

    We accept these four strategies on every different level. The database handle takes strategies, and so does the subscriber definitions and the web layer API endpoint definitions. You can also override strategies per context.

    This gives us the inverted default, everything goes to read replica, while still providing the engineering team with the tools they need to keep shipping.

    Outcome

    Once every part of our codebase was “opted in” to the new read replica behavior, we inverted the default. Everything now goes to the read replica, and the people writing code don’t have to worry about it. It just works. Getting there is not hard, but it did come with a lot of learnings.

    So what about the numbers? After the project wrapped up more than 60% of all read queries go to the replica. Of the ones that don’t, the majority were explicitly pinned to the primary. Only a small number of reads are actually routed to primary to preserve read-after-write consistency.

    We cut CPU utilization on the primary in half. Looking back over the last few months CPU utilization had been growing significantly week over week as our customer base has been expanding. This project reversed the trend and we’ve had several weeks of CPU utilization going down week over week. A healthy primary with lots of spare capacity means that we can keep handling disaster situations where half the internet goes down and everyone gets paged.

    C

    Replica lag causes back pressure instead of overload. In a catastrophe situation where our replica is failing to keep up, or going down, all events are safely buffered in our queue service and only picked up when the replica is ready to get back to work. This means that problems with the replica do not spread to other parts of our stack.

    We’ve set the stage for horizontal scaling. The door is now wide open for us to add additional read replicas, giving us a clear path to keep scaling our product for the future.

    Engineering can keep shipping. Our transparent, opt-in by default, database pool routing strategies and read-after-write consistency mechanisms ensure that the read replica does not get in your way. You could work here for months without even realizing it exists.

    We hope this writeup can help demystify read replicas and how to get the most out of them, applying fairly straightforward techniques to ensure read-after-write consistency. We really enjoyed working on this project and we hope you’ve enjoyed reading about it! If you’re interested in this kind of thing, come join us, we’re hiring!

    Picture of Johanna Larsson
    Johanna Larsson
    Product Engineer
    View more

    See related articles

    View all

    So good, you’ll break things on purpose

    Ready for modern incident management? Book a call with one of our experts today.

    Signup image

    We’d love to talk to you about

    • All-in-one incident management
    • Our unmatched speed of deployment
    • Why we’re loved by users and easily adopted
    • How we work for the whole organization
    \ No newline at end of file diff --git a/sreweekly/articles/528/index.json b/sreweekly/articles/528/index.json new file mode 100644 index 00000000..7e757dc1 --- /dev/null +++ b/sreweekly/articles/528/index.json @@ -0,0 +1,50 @@ +[ + { + "idx": 1, + "url": "https://engineering.atspotify.com/2026/7/content-ingestion-and-podcast-video-incident-report/", + "ok": true, + "error": null + }, + { + "idx": 2, + "url": "https://blog.pragmaticengineer.com/the-pulse-quitting-spotify-podcasts-over-reliability/", + "ok": true, + "error": null + }, + { + "idx": 3, + "url": "https://greatcircle.com/blog/2026/06/16/incident-tech-lead/", + "ok": true, + "error": null + }, + { + "idx": 4, + "url": "https://read.zerosevzero.com/p/the-quiet-quarter", + "ok": true, + "error": null + }, + { + "idx": 5, + "url": "https://www.honeycomb.io/blog/30-70-prs-day-how-we-managed-not-wreck-systems", + "ok": true, + "error": null + }, + { + "idx": 6, + "url": "https://www.uptimelabs.io/articles/incidents-and-recovery", + "ok": true, + "error": null + }, + { + "idx": 7, + "url": "https://dzone.com/articles/agent-sprawl-production", + "ok": true, + "error": null + }, + { + "idx": 8, + "url": "https://incident.io/blog/dont-add-a-read-replica-until-youve-read-this", + "ok": true, + "error": null + } +] \ No newline at end of file diff --git a/sreweekly/articles/529/01-without-a-program-to-support-them-incident-management-processes-wither.html b/sreweekly/articles/529/01-without-a-program-to-support-them-incident-management-processes-wither.html new file mode 100644 index 00000000..8943930e --- /dev/null +++ b/sreweekly/articles/529/01-without-a-program-to-support-them-incident-management-processes-wither.html @@ -0,0 +1,581 @@ + + + + + + + + + + Without a program to support them, incident management processes wither | Brent Chapman + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +
    + + + +
    + +
    +
    + +
    +
    +
    +
    +
    + + +
    + +

    Every fire department has a training program. Big-city departments have entire training divisions; even small volunteer departments that can’t spare anyone full time still name a training officer. Not because training is the department’s mission, but because maintaining the capability to do the mission requires sustained, dedicated attention.

    + + + +

    New recruits need to be brought up to speed. Everyone needs to learn about evolving techniques and new equipment. Procedures need to be updated as building codes and materials change. Hard-won lessons from past incidents would survive only as stories told around the kitchen table; the fire service has a strong storytelling tradition, and its legends and cautionary tales carry real value, but oral history is hard to study, standardize, and train on.

    + + + +

    Maintaining operational capability is itself a job, distinct from the operational work it supports, and fire departments size the role to the department rather than leave it unassigned.

    + + + +

    Many software companies haven’t learned this yet. They invest real effort in building an incident management process. They define severity levels, write runbooks, designate incident commanders (ICs), set up communication channels. The project might take weeks or months of focused work, often driven by someone who cares deeply about doing it right (and often done in their “spare time”). When it’s done, it works, at least for a while. Incidents get declared. ICs run the response. Post-incident reviews happen. Everyone takes it for granted.

    + + + +

    Then the person driving it gets promoted, or moves to another team, or leaves the company. The process, which was never really institutionalized because it didn’t need to be while that person was carrying it, begins to decay. Not catastrophically, but more like a garden nobody is tending any more: it doesn’t collapse overnight, it just slowly fills with weeds until one day you look up and realize the original design is barely recognizable.

    + + + +

    The training materials haven’t been updated since the initial rollout. New engineers join but never go through incident training because nobody is scheduling it anymore. The severity level definitions still describe one product, but the company now has three. The IC rotation is running on the same six people it started with, even though the engineering team has doubled in size. The post-incident review template still references a tool the company stopped using a year ago.

    + + + +

    None of these are crises on their own. Each one is easy to defer. But they compound, and the cumulative effect is that the process on paper bears less and less resemblance to what actually happens during incidents. In my experience, six months is roughly how long institutional momentum carries before the absence of active stewardship becomes visible in the quality of your incident responses. And growth accelerates the decay: the company simply grows away from the process, and nobody’s job is to notice.

    + + + +

    This is what happens when you have a process but not a program.

    + + + +

    A process is not a program

    + + + +

    A process is a set of documented procedures: how incidents get declared, who fills which roles, what communication channels to use, how to run a post-incident review. A process can be written down, trained once, and followed.

    + + + +

    A program is the organizational structure that develops, maintains, evolves, and champions the process over time. It’s the thing that keeps the process alive.

    + + + +

    Many companies build the process and assume they’ve built the program. They haven’t. They’ve written a document, and documents don’t train new hires, don’t recruit for on-call rotations, and don’t update themselves when the company reorganizes around them. People do those things, and it only happens reliably when it’s actually somebody’s job.

    + + + +

    “Everybody owns it” means nobody owns it

    + + + +

    When I ask companies who owns their incident management program, the most common answer is some version of “we all do” or “the engineering organization as a whole.” This sounds collaborative. In practice, it means nobody has the explicit responsibility, the dedicated time, or the institutional authority to keep the process alive.

    + + + +

    This organizational challenge isn’t unique to incident management. Companies that are serious about security don’t say “everybody owns security” and leave it at that. They assign ownership because shared responsibility without explicit ownership means the work doesn’t get done.

    + + + +

    Incident management is the same kind of organizational capability. It needs someone whose actual job, not just their passionate side interest, is keeping it healthy.

    + + + +

    What a program actually does

    + + + +

    When I talk about an incident management program, I mean ownership of the full lifecycle of the capability, not just the procedures themselves. That includes keeping everything current as the company grows and changes: process documentation, severity definitions, escalation paths, tooling, runbooks.

    + + + +

    It includes running a training pipeline so new hires are prepared before their first real incident, not thrown into the deep end during it. It includes maintaining the incident commander corps: recruiting new incident commanders, nurturing their development, supporting healthy on-call rotations across teams, and recognizing the people who do this demanding work. My former Slack colleague Scott Nelson Windels likens this to the farm teams and academies that elite sports clubs run: the point isn’t just fielding today’s roster, it’s making sure capable players are always coming up to fill it next quarter, too.

    + + + +

    And it includes owning the post-incident review process and looking across incidents for patterns that no individual team would spot on their own. It includes tracking whether the process is actually being followed, and investigating when it isn’t, not to punish people, but to understand whether the process needs to change.

    + + + +

    No single component is enough on its own, and no component stays healthy without sustained attention.

    + + + +

    The good news

    + + + +

    Building a program doesn’t require hiring a large team or creating a new department. At many companies, especially smaller ones, it starts with one person who has explicit ownership and dedicated time. What matters is that the responsibility is named, visible, and institutionally supported, not just assumed.

    + + + +

    Here’s a quick test. Ask who owns your incident management program. Not who wrote the process, and not who ran the last big incident, but who is accountable, today, for whether the training is current, the rotations are staffed, and the severity levels still match the product. If the answer is a name, the follow-up question is what happens when that person leaves. If the answer is “everybody,” or someone who left the company last year, the process is quietly withering. And if you have a program but it would collapse without you, you haven’t finished building it yet.

    + + + +

    The fire department didn’t name a training officer because it had extra budget. It named a training officer because it understood that maintaining a capability requires ongoing investment. The alternative, assuming trained firefighters stay trained and procedures stay current without anyone specifically owning those things, is how capabilities quietly erode until they fail when you need them most.

    + + + +
    + + + +

    I’m writing a book on Incident Management for DevOps and SRE. If you’d like to know when it’s available, and get occasional updates along the way, you can sign up at im4ds.com.

    + + + +

    If your company needs help building its incident management program, that’s the focus of my consulting practice at Great Circle.

    +
    + +
    + +
    + + +
    +
    +
    + + +
    + + + + +
    +
    + + +
    + + + + + + + + + + + + + + diff --git a/sreweekly/articles/529/02-what-comes-after-observability.html b/sreweekly/articles/529/02-what-comes-after-observability.html new file mode 100644 index 00000000..b51f8042 --- /dev/null +++ b/sreweekly/articles/529/02-what-comes-after-observability.html @@ -0,0 +1,4 @@ +What Comes After Observability?
    Observability Engineering second edition out now! 27 net-new chapters written for today's observability challenges.Get your copy

    What Comes After Observability?

    A year ago, I predicted ways in which AI was about to fundamentally change observability as we knew it. Here's what we've seen happen since—both at Honeycomb and with our customers—and what we're building for the future.

    What Comes After Observability?

    A year ago, I wrote “It’s the End of Observability (and I Feel Fine).” The upshot of that post was that AI was about to fundamentally change the way we approach systems design and operation in the future. In the grand tradition, I’d like to revisit my claims from then and see how my predictions panned out.

    Claim: Agents can zero-shot investigations for less than a dollar

    I asked the agent the same question we’d ask you in a demo, and the agent figured it out with no additional prompts, training, or guidance…and it did it for sixty cents.

    I pointed out in my original blog that the way I evaluated our MCP server was to feed it a prompt based on the same demo that we give to people—e.g., here’s this weird latency spike, investigate it, tell me why it happened. At the time, that demo took eight tool calls and about $0.60 of inference. A year later, we process about 2 million agent-initiated query runs through Honeycomb every month!

    What’s really interesting, though, is that we’ve seen agents start to accomplish significantly longer-horizon tasks. Frontier models—like Sonnet or Fable 5—are increasingly agentic and capable. In the same agent session, we’ve seen tool calls double on average between February and June of this year. Our original demo of eight queries looks downright quaint.

    Watch Austin, on demand

    Watch Austin Parker and the other authors

    of Observability Engineering

    discuss what's changed in the AI era.

    Claim: Humans stay in the loop

    I’m not gonna sit here and say this destroys the idea of humans being involved in the process, though. I don’t think that’s true. The rise of the cloud didn’t destroy the idea of IT. The existence of Rails doesn’t mean we don’t need server programmers. Productivity increases expand the map. There’ll be more software, of all shapes and sizes. We’re going to need more of everything.

    What’s really interesting is that human-initiated queries haven’t shrunk. Agents aren’t taking investigations away from humans, they’re accentuating them. People still use the UI, people still check their boards, but they’re now able to do more than they could before by leveraging AI. We see this a lot in terms of edits—while people are using agents to build boards and update SLOs, the overwhelming majority of tool calls are to just run queries.

    In my blog, I claimed that productivity increases weren’t going to diminish the impact or effect of people, and that’s what the data shows. The map is not the territory, but the map has grown significantly thanks to agents.

    This doesn’t mean that all of these agents are being driven by a human, though. We’re seeing an increasing amount of headless agents using our MCP; it’s actually the fastest-growing segment.

    Claim: It’s only getting cheaper to do this

    Inference costs are only going down…If your product’s value proposition is nice graphs and easy instrumentation, you are cooked. An LLM commoditizes the analysis piece, OpenTelemetry commoditizes the instrumentation piece.

    In my original post, I focused on the idea of inference costs going down. I think that’s, broadly, still true (even if a lot of people are getting sticker shock as they transition from all-you-can-eat to API pricing for development workloads). That said, I’m writing this post on a laptop that effortlessly runs Google’s Gemma4 model, and it’s perfectly capable of using our MCP. I think, long-run, the “cost of AI” is going to keep going down at a fixed level of capability.

    What I didn’t call, though, is that the robots are actually surprisingly efficient when it comes to using Honeycomb! On average, agent queries cost about half as much to serve as human ones. There’s a pretty easy explanation for this: the agents are able to much more accurately figure out what to search for, especially if they’re running with access to your code. They know exactly what to look at, and they don’t need to spend as much time searching across an entire environment to orient themselves.

    Interestingly enough, we also notice that agents—on average—tend to run queries over a 50% shorter time range than humans. Where this gets really interesting is how different the agent runs tend to be. Custom agents tend to look at smaller time windows, while human-driven ones tend to look across longer periods of time. My hypothesis is that those custom agents are probably more task-oriented (“Here’s an error, look around this time for traces”) vs. more exploratory work being done in the development loop.

    The headline number, though? On average, an agent-initiated query fans out to 3.3x fewer lambdas, scans 1.8x fewer bytes, and burns 2.2x less compute. What does this mean for us? Well, from March to June we’ve grown 4x in query volume while only increasing cost by 23%, leading to a 72% reduction in unit costs for agent-initiated queries. I’ll take that!

    Average per-query run of human-initiated queries compared to agent-initiated queries.

    An interesting coda to this, though, is that agents are significantly worse for caching. While we don’t see significant amounts of caching for interactive queries in general, we see almost none for agent-driven queries. This is mostly because agents don’t need to ask the same question twice. I’ve also personally observed that agents prefer to re-run queries rather than reuse existing ones, which is interesting and probably deserves more attention.

    That said, we’ve seen pretty staggering levels of adoption, especially in the enterprise. One of our larger customers read 62 PiB of data in one month's time exclusively through agents, mostly through off-the-shelf tools like Claude Code.

    There’s a downside to this as well, though: a lot of those queries wind up being kinda slow, because the agents aren’t chewing through nicely structured data, they’re grepping across unstructured body fields. Structured, high-cardinality, high-dimensionality data is no longer a nice-to-have for humans, it is a requirement for making agent investigations affordable. All of that data that the agent has to chew through in order to find the needle in the haystack? That’s tokens coming out of your pocket. Do you want to pay Fable rates to look through gigabytes of crappy logs?

    Claim: Fast feedback loops are all that matter

    I’m gonna put a marker out there: the only thing that really matters is fast, tight feedback loops at every stage of development and operations. AI thrives on speed—it’ll outrun you every time.

    Just go ask Claude to summarize the current state of the “Loops” discourse and you’ll see that the rest of the AI thought leadership industry is talking about something you read here a year ago.

    Anyway, there’s no prize for being first, so I instead want to share a story from one of our customers. Guy writes in to tell us how much he loves MCP and how he’s using it. He’s hooked up everything to agents that talk to Honeycomb. When a customer reports an issue or someone files a bug, the agent goes out to verify it using production telemetry. When an SLO burns, agent goes out, looks into it. Agent takes that investigation, hands it off to another agent, which goes out and writes a fix, makes a PR. Yet another agent reviews the PR and pings a human for final review and merge. PR gets deployed, goes into prod, wakes up that original agent to look at production telemetry to see if the problem was fixed. If it wasn’t, enter the loop again.

    That’s the kind of tool that we’re building at Honeycomb today, yes. But what we’re building for tomorrow is so much more than that. I don’t want us to build just a really fast column store (although a really fast column store is cool). I want to make it easy for everyone to have these fast feedback loops. I want you to be able to understand what your agents are doing with your code, and I want to make it easy for you—and your agents—to learn about production and turn those learnings into knowledge.

    If you’ve had the chance to check out Canvas, and our Canvas Agent, you’ve seen the first cut of this (and if you haven’t, you should check it out; it’s free!). This is just the first chapter of what we’re building, though. If you're interested in learning more about how we see AI changing observability and what we're working on here at Honeycomb, join me and the other authors of the Observability Engineering book in our on-demand AMA.

    P.S. We’re hiring if you wanna come work on agents with us.

    Want to learn more?

    Talk to our team about how we're helping organizations build the operational foundation for AI development success.

    \ No newline at end of file diff --git a/sreweekly/articles/529/03-the-rise-of-agentic-sre-humans-agents-and-reliability.html b/sreweekly/articles/529/03-the-rise-of-agentic-sre-humans-agents-and-reliability.html new file mode 100644 index 00000000..e3613333 --- /dev/null +++ b/sreweekly/articles/529/03-the-rise-of-agentic-sre-humans-agents-and-reliability.html @@ -0,0 +1,2896 @@ + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + The Rise of Agentic SRE: Humans, Agents, and Reliability + + + + + + + + + + + + +
    +
    +
    +
    + + + +
    + + +
    + + + + + + + + + + +
    +
    +
    + +
    +
    + +
    +
    +
    +

    Could your team report a vulnerability within 24 hours? Find out on September 23.

    +
    + + + +
    +
    +
    + +
    + +
    +
    +
    +
    +
    +
    +
    + + +
    +
    + + + + + +
    +
    +
    + + + +
    +
    +

    The Rise of Agentic SRE: Humans, Agents, and Reliability

    +
    + +
    +

    Agentic SRE speeds up incident response, but it also requires clear guardrails, strong observability, and human oversight.

    +
    + +
    + By  + + + · + Opinion +
    +
    +
    + +
    + + +
    + + + Comment + + +
    + +
    +
    + Save +
    +
    + + + + + +
    +
    + + 4.6K Views +
    +
    +
    + + +
    + +
    + +
    +

    Site reliability engineering has always been about reducing toil, improving resilience and helping teams respond to incidents with speed and confidence. Agentic SRE takes this idea further, allowing AI systems to observe, reason, and act within operational workflows inside of bounded constraints. The outcome is not a replacement for SREs, but a new operating model in which humans supervise intelligent agents that can help triage, diagnose, and remediate faster than manual processes alone.

    +

    What Agentic SRE Means

    +

    Agentic SRE is the use of AI agents to carry out reliability tasks with some autonomy. The agents are able to capture telemetry, correlate signals across systems, propose likely causes, take safe actions, and hand over to humans when the problem exceeds their authority. In practice, this means an AI assistant that can summarise an incident, pull up relevant dashboards, check recent deploys, compare symptoms against runbooks and even trigger low-risk remediation steps.

    +

    What changes are not the nature of the assistance but the limits of its application. Traditional automation is usually rule-based: if X happens, do Y. Agentic systems are different in that they can adapt to context, select between several paths, and orchestrate steps across tools. This makes them especially useful in complex environments where the same symptom may come from many different root causes.

    +

    Why SRE Needs Agents

    +

    The systems today are too big and too interconnected to be run totally by hand. Teams are contending with noisy alerts, fragmented observability data, constant deployments, and ever more dynamic infrastructure. During incidents, engineers often burn precious minutes just to gather context before they can start a real diagnosis. Agentic SRE is attractive because it shortens that time.

    +

    In a handful of high-friction places, agents can cut toil. They can filter alert storms, enrich alerts with deployment history, draw out meaningful patterns from logs, and surface relevant runbooks. They can also automate repetitive incident response tasks such as opening tickets, notifying owners, checking service health, or validating if a rollback is safe. That doesn't eliminate the need for engineers, but it does take away some of the low-value work that distracts them from judgment-intensive choices.

    +

    The Human Role Remains Central

    +

    One common fear is that SREs will be replaced with autonomous systems. Indeed, the human role becomes more, not less, important. Agents are good at pattern recognition, summarisation, and bounded execution. Humans are still better at trade-offs, risk assessment, organisational context, and deciding when not to act. Reliability is not merely a technical problem. It is a business and coordination problem.

    +

    Humans should set policies, guardrails, and escalation thresholds for agent behaviour. They need to decide which actions can be safely automated, which require approval, and which should never be delegated. So the SRE is evolving from operator to system designer, to policy author, to reliability supervisor. That shift is profound because the skill set you need for the job changes.

    +

    Where Agents Fit Today

    +

    The best place to start is with low-risk, high-frequency jobs. These are the areas where automation can provide immediate value without unacceptable risk. Think incident summarisation, alert enrichment, log correlation, runbook retrieval, change impact analysis, and post-incident report drafting.

    +

    Incident copilots are another strong use case. An agent can also act as a second brain during an outage: it can aggregate timelines, verify recent code changes, search knowledge bases, and suggest next steps. It can help responders avoid duplication of effort and make the first 10 minutes of an incident much more productive. An effective agent can also lessen the cognitive load on on-call engineers by turning the scattered telemetry into a coherent story.

    +

    A third useful area is remediation assistance. Agents can recommend actions such as scaling a service, restarting a failing job, disabling a faulty feature flag, or rolling back a deployment. In mature setups, these actions can be executed automatically for pre-approved scenarios, while more risky actions still require human confirmation. That combination of automation and oversight is where agentic SRE becomes genuinely powerful.

    +

    A Practical Architecture

    +

    An effective agentic SRE system typically has five layers. First, it needs a telemetry layer that includes metrics, logs, traces, events, and deployment data. Without strong observability, the agent is blind and will make incorrect guesses. Second, it requires a reasoning layer, often powered by an LLM, to interpret context and decide what to do next. 

    +

    Third, there should be a tool layer that gives the agent access to safe operational functions, such as querying dashboards, reading configs, opening tickets, or triggering runbooks. Fourth, it needs policies and guardrails that define permissions, approval workflows, rate limits, and failure boundaries. Finally, it should have an audit layer so every decision, action, and recommendation can be traced later.

    +

    That architecture matters because the danger is not the model itself; the danger is uncontrolled action. A reliable agent is not one that knows everything. It is one that acts only within well-defined limits and remains observable, reversible, and accountable.

    +

    Guardrails That Matter

    +

    Trust is the currency of autonomous operations. If teams do not trust the system, they will ignore it. If they trust it too much, they may hand over dangerous actions without oversight. The right answer is neither blind trust nor permanent skepticism. It is a layered trust model built through guardrails.

    +

    Start with permission scoping. An agent should not have broad access by default. Its permissions should be narrow, explicit, and tied to specific tasks. Next, use action tiers. Low-risk actions can be automatic, medium-risk actions can require confirmation, and high-risk actions should remain human-only. You also need strong rollback paths so any automated action can be quickly reversed.

    +

    Another essential safeguard is the observability of the agent itself. Just as production systems need monitoring, agents need monitoring too. Teams should track what the agent saw, what it inferred, what action it proposed, and whether the result improved the situation. That makes the system auditable and helps teams refine its behavior over time.

    +

    The Operating Model Changes

    +

    Agentic SRE changes incident response from a purely human workflow into a human-agent collaboration loop. In the old model, an engineer gets paged, reads alerts, searches dashboards, checks logs, consults teammates, and then acts. In the new model, the agent can do much of the initial gathering and triage before the human even joins. That shortens the path from detection to understanding.

    +

    This also changes how teams design runbooks. Instead of static documents that people read under pressure, runbooks become machine-readable operational playbooks. Some of the best runbooks will be written with automation in mind, including clear preconditions, decision points, and action boundaries. That makes them useful both for humans and for agents.

    +

    Post-incident work also improves. Agents can draft a timeline, collect evidence, identify suspicious changes, and summarize repeated patterns across incidents. That leaves engineers with more time to focus on systemic fixes rather than manual documentation. Over time, the organization develops a stronger feedback loop between incidents, learning, and platform improvements.

    +

    Risks and Failure Modes

    +

    Agentic SRE is not free of risk. One failure mode is confident hallucination, where an agent sounds plausible but is wrong. In operations, a wrong answer is not just inaccurate; it can cause downtime. Another risk is over-automation, where teams let agents act in situations that are not actually safe to delegate.

    +

    There is also the risk of hidden complexity. If an agent stitches together many systems, it can become difficult to understand why it chose a specific action. That opacity can undermine trust and create governance problems. Security is another major concern because an agent with tool access can become an attractive target if permissions are poorly controlled.

    +

    These risks do not mean agents should be avoided. They mean they must be introduced carefully. The best strategy is to start with narrow, well-understood workflows, measure outcomes, and expand only when confidence is earned. Reliability teams already understand progressive delivery, canary releases, and blast-radius reduction; the same principles should apply to agentic operations.

    +

    How to Start

    +

    The easiest entry point is to pick one painful workflow and automate only the first mile. A good candidate is alert triage. An agent can ingest alerts, group duplicates, summarize likely causes, and point responders toward relevant dashboards and runbooks. That alone can save significant time without requiring the agent to make risky changes.

    +

    Another strong starting point is incident summarization. This is low risk, highly useful, and easy for teams to evaluate. A third option is change impact analysis, where an agent compares recent deploys, feature flag changes, and error spikes to highlight likely correlations. These use cases are valuable because they build trust through usefulness rather than hype.

    +

    Measure success with clear operational metrics. Look at time to acknowledge, time to diagnose, time to mitigate, alert volume reduction, and after-hours toil reduction. Also measure negative outcomes, such as false suggestions, unsafe recommendations, or overreliance on the agent. Good SRE practice is about evidence, not enthusiasm.

    +

    A New Reliability Mindset

    +

    The biggest change Agentic SRE brings is a shift in mindset. It encourages teams to stop viewing automation as just a collection of scripts and to see it instead as a supervised operational partner. This partner can observe faster than a person, summarise quickly, and carry out repetitive tasks more reliably. However, it still requires humans to define the purpose, set limits, and determine acceptable risk.

    +

    This is why agentic SRE is not merely “AI in operations". It represents a larger redesign of how reliability work is accomplished. The focus shifts from manual responses to intelligent coordination. It changes from isolated dashboards to context-aware agents. It evolves from static runbooks to flexible playbooks. It transforms reactive tasks into guided independence.

    +

    Organizations that excel with this model will not be the ones that automate everything. They will be the ones that automate thoughtfully, govern effectively, and keep humans involved where decision-making matters most. In this way, Agentic SRE is more about enhancing the reliability system around the engineer than about replacing the engineer themselves.

    +

    Closing Thoughts

    +

    Agentic SRE marks a real change in how we can manage modern systems. It provides a way to respond faster, reduce repetitive work, and handle incidents more consistently, but only with strong observability, clear permissions, and human oversight. The future of reliability is not completely automatic or fully manual; it is collaborative, constrained, and constantly improving. 

    +

    For SRE teams, there's a chance to become designers of this new model. This involves creating agent workflows, writing safer runbooks, setting policy limits, and measuring impact with the same attention given to any production system. Teams that excel in this will not only respond more quickly. They will create systems that are more resilient, more adaptable, and much simpler to operate at scale.

    +
    + +
    + + +
    +

    Opinions expressed by DZone contributors are their own.

    +
    +
    +
    +
    + + + +
    + +
    +
    +
    + +
    + + + +
    +
    +
    +
    +
    +
    +
    + +
    + × + +
    + + +
    +
    + +
    +
    +
    +
    +
    + + +
    + + + + + +
    + +
    + + + + + + + + + + + + + + + + + + + + + + diff --git a/sreweekly/articles/529/04-incident-report-july-2-2026-us-east-services-outage.html b/sreweekly/articles/529/04-incident-report-july-2-2026-us-east-services-outage.html new file mode 100644 index 00000000..949f33fe --- /dev/null +++ b/sreweekly/articles/529/04-incident-report-july-2-2026-us-east-services-outage.html @@ -0,0 +1,269 @@ +Incident Report: July 2, 2026 — US East Services Outage | Railway Blog
    Avatar of Ray Chen
    Ray Chen

    Incident Report: July 2, 2026 — US East Services Outage

    Railway experienced a Major Outage concentrated in one of our US East availability zones on July 2, 2026.

    +

    A network degradation in one of the ISPs connecting our datacenters to the rest of the internet caused elevated latency and packet loss for traffic between our US regions. While rerouting traffic away from degraded ISP, a change at one of our US East availability zones briefly left the site without a stable route to the internet.

    +

    The effects from above exposed hidden bugs that silently pushed storage traffic onto a slow backup network and impacted some private networking tunnels, degrading disk performance and private networking in US East for roughly two hours.

    +

    Impact

    +

    On July 2, 2026 between roughly 07:44 UTC and 12:01 UTC, users may have experienced increased response times and intermittent connectivity issues on traffic between US regions, including private networking. Some workloads in one of our US East availability zones additionally saw degraded disk performance and disrupted private networking for roughly two hours.

    +

    Incident Timeline

    +

    All times are UTC on July 2, 2026.

    +
      +
    • 07:44 — We started observing packet loss in our US East region affecting user traffic. A public incident was declared on our status page. We traced the packet loss to one of our upstream network carriers
    • +
    • 07:44~08:32 — We disconnected from the degraded network carrier at all US border routers. Traffic was successfully rerouted through other carriers. Conditions improved across most US paths, but latency and packet loss into US East had not fully recovered
    • +
    • 08:39 — Paths through a secondary network carrier at the affected US East zone were still showing packet loss, as it was handing traffic back through the degraded carrier on the return leg. We disconnected from the secondary carrier there as well. Unknown to us at the time, that was the only carrier still supplying that site's default route (the catch-all path a network uses to reach the internet). The primary degraded carrier had already been disconnected, and our only remaining carrier’s connection there does not supply a default route. This left the zone without a stable route to the internet for roughly 20 minutes. During this window the disruption was at its most severe as traffic into and out of US East saw failed connections and heavily degraded private networking
    • +
    • 08:59 — We reconnected our secondary carrier. Routing stabilized, US East connectivity began recovering, and the storage cluster returned to a coordinated, healthy state
    • +
    • 09:00~10:45 — Storage performance in the zone remained degraded despite routing looking healthy. Throughput stayed pinned at roughly a third of capacity
    • +
    • 10:45 — Root cause of degraded storage performance identified. Significant amount of storage connections had been established over a slow internal management network during the routing instability and remained stuck there after recovery
    • +
    • 10:45~11:00 — We terminated the stuck connections across all storage and compute hosts in the zone. They reconnected over the correct network within seconds
    • +
    • 11:04 — I/O wait across the zone returned to baseline (58% → under 5%). Storage throughput surged as the cluster caught up on backlogged writes, then settled at normal levels. At the same time, we identified that private networking tunnels had latched onto an incorrect address during the routing instability and never corrected themselves
    • +
    • 11:49 — A fleet-wide restart of the mesh networking agents in the zone forced all tunnels to re-establish with correct addresses. Private networking fully recovered and full connectivity was restored across US regions; we continued monitoring
    • +
    • 12:01 — With latency, packet loss, private networking, and volume performance stable at normal levels, the incident was marked resolved
    • +
    +

    The full incident is available on our Status Page.

    +

    What Happened?

    +

    A few things went wrong that led to unintended cascading effects across our systems. The commonality across the failure cases were traced to stale connections that ended up capturing a bad path during a brief window of instability, and held onto it after the network recovered, because nothing in the system re-asserted the correct state.

    +

    1) Upstream ISP Degradation (US Regions)

    +

    Datacenters connect to the internet by buying connectivity from transit providers; carriers that operate long-haul fiber and agree to deliver your traffic to any destination in the world. The largest of these are called Tier 1 ISPs. Railway connects every Metal datacenter to at least three of them, so that any single carrier can fail without taking us offline.

    +

    On July 2, a carrier carrying our traffic was impacted by a network degradation somewhere in their US backbone. Traffic they normally carried on other paths spilled onto the route that carried our traffic between US West and US East, causing saturation leading to higher latency and packet loss.

    +

    Our own routers showed no errors and healthy connections to every carrier, which told us the problem was upstream. Internal probes caught the degradation nearly two hours before it became visible to user traffic, giving us time to trace the lossy paths, all of which ran through the degraded network provider, while retries and redundant routing absorbed the loss.

    +

    When the loss began reaching user traffic, we declared a public incident and disconnected from the degraded provider at all US borders. Traffic rerouted through other providers and packet loss returned to baseline on most US paths.

    +

    2) Storage Performance Degradation (US East)

    +

    At 08:39, we also disconnected from a secondary carrier at this zone. Paths through the secondary carrier were still showing packet loss, because it was handing traffic back through the primary carrier’s degraded network on the return leg, and disconnecting a carrier on our side does not control the route traffic takes coming back to us.

    +

    In hindsight, this change should not have been made without first verifying the behavior of default routes on our core switches. The zone is one of our first-generation sites, and unlike our newer datacenters, it gets its default route (the catch-all path to the internet) from its carriers rather than generating one itself.

    +

    The primary network carrier was already disconnected at that site, and the only remaining carrier there does not supply a default route. Therefore, disconnecting the secondary carrier removed the last one. For the roughly 20 minutes until we reconnected, the site had no stable route to the internet. This window was the most severe part of the incident for users: traffic into and out of US East saw failed connections and degraded private networking until routing stabilized at 08:59.

    +

    That instability exposed a hidden bug in how our servers behave when their primary network path disappears. Each server has two networks: a high-bandwidth fabric that carries production traffic, and a slow management network used for administrative access.

    +

    When the fabric's default route vanished, the servers' operating systems fell back to the only route left: the one on the management network. A default Linux behavior then allowed servers to answer for their storage addresses on that network too, so storage traffic began flowing over a path with a small fraction of the fabric's capacity.

    +

    Network connections don't re-check their route once established; they keep using the path they started on until they close. So the storage connections created during that 20-minute window stayed stuck on the slow management network even after the fabric was fully restored. The routing tables looked correct, and the storage cluster reported healthy, but storage throughput was capped at roughly a third of normal, and two thirds of servers in the zone sat waiting on disk while the cluster tried to push its backlog through.

    +

    Once we found the connections coming from management-network addresses, we terminated those connections across every storage and compute server in the zone. They reconnected over the correct fabric within seconds, and I/O wait dropped from 58% to under 5% in about 15 minutes.

    +

    3) Private Networking Degradation (US East)

    +

    Railway's private networking runs over encrypted tunnels between servers, and each tunnel learns its peer's address from the packets it receives.

    +

    During the routing disturbance, tunnel traffic was briefly funneled through a device that rewrites the source address of traffic passing through it. Thousands of tunnels learned that device's address as their peer's address, and kept it after routing was rolled back. The mesh only re-verifies peer addresses when its membership changes, and because these tunnels sit silent when idle, a broken one never sends the packet that would have fixed it.

    +

    At peak, roughly 20,000 host-to-host private network links were blackholed. This included inter-region services communicating with services in US East. Recovery required restarting the mesh networking agents across the fleet, forcing every tunnel to re-establish with the correct addresses.

    +

    Preventative Measures

    +

    We have already rolled out the following:

    +
      +
    • Disconnected the degraded network carrier at all US borders. Our fleet is currently operating on other Tier 1 carriers with full headroom. We will reconnect the degraded carrier once their backbone has recovered and we have verified path health
    • +
    • Cleared all stuck management-network storage connections across every storage and compute host in the affected zone
    • +
    • Restarted the mesh networking agents fleet-wide in the zone, re-establishing all private networking tunnels with correct addresses
    • +
    +

    We are additionally working on:

    +
      +
    • Migrating first-generation sites to self-generated default routes. Our newer datacenters generate their own default route at the border rather than depending on carriers to supply one. At those sites, losing any single carrier is a minor path change, not a site-wide routing event. The affected zone predates this design, which is why it was vulnerable today. Bringing all remaining first-generation sites onto this pattern is the biggest structural fix from this incident
    • +
    • Correcting host fallback behavior so that production traffic never takes the management network as a fallback path. Both networks are internal to our infrastructure and within the same trust boundary — traffic that crossed the management network stayed inside our own equipment, and private networking traffic remained encrypted end-to-end — so this was a path-selection and capacity problem. We will fix this by making a missing fabric route fail cleanly and immediately instead of silently taking a slower path, which will give us the ability to detect and resolve it in minutes
    • +
    • Alerting on management network utilization and blackholed private network links, so that either failure mode pages us immediately instead of surfacing as degraded performance
    • +
    +
    +

    A carrier failure is a normal event on the internet, and our multi-carrier design handled it the way it should. The degradation that followed came from our side: an older site design that depended on carriers for its default route, a disconnection made without checking what routing would remain, and systems that captured a bad path during the instability and held onto it silently after the network recovered.

    +

    We apologize for this outage and are actively working to prevent similar issues from happening again. Each of the fixes above targets one of those links, so that the next carrier failure ends where this one should have (with traffic quietly taking a different road).

    \ No newline at end of file diff --git a/sreweekly/articles/529/05-finding-bugs-in-raft-implementations.html b/sreweekly/articles/529/05-finding-bugs-in-raft-implementations.html new file mode 100644 index 00000000..69422964 --- /dev/null +++ b/sreweekly/articles/529/05-finding-bugs-in-raft-implementations.html @@ -0,0 +1,367 @@ +Finding bugs in Raft implementations | Antithesis
    Bug Bash is coming to Copenhagen this fall!Learn more

    July 27, 2026

    Finding bugs in Raft implementations

    Introduction

    +
    TW Lim headshot
    TW Lim Technical Writer
    +

    The Raft consensus protocol is one of the foundations of the internet. Raft offers a formal specification in TLA+ and a concrete, detailed implementation guide, and hundreds of groups have created open-source implementations of the algorithm by following the instructions in the guide. It is, by far, the most widely used consensus algorithm in production systems.

    +

    Despite the paper’s famously accessible style, we’ve found bugs in every Raft implementation we’ve tested, including HashiCorp Raft, Aeron Cluster, OpenRaft, and MicroRaft — despite the investment in formal methods, careful code review, unit testing, and years of testing in production. The bugs we found manifest as violations of Raft’s main invariant (called state machine safety in the paper, commonly referred to elsewhere as total order delivery). If you’re using a Raft implementation, you might want to check it for bugs.

    +

    We’ve sent bug reports upstream. This isn’t intended as a critique of Raft, its authors, its implementers, or any particular implementation. Raft implementations, even with a formal specification and a detailed implementation guide, are not easy to write.

    +

    Rather, this is a story about correctness in distributed systems — the inevitability of bugs, the inadequacy of any single approach, and the high cost of learned helplessness.

    +

    This is a long post. The first section is a position paper, but the rest is a detailed analysis of the issues we found, using one implementation as an example, a discussion of how we found them, and why we think they’re there.

    +

    Why this matters

    +

    If you’ve worked on distributed systems, you’ve been part of a war room at some point (if you haven’t, don’t worry, you will be), dealing with an outage like this one, or this one — both of which were caused by consensus failures. When a consensus issue results in an incident, it’s inevitably far reaching and extremely painful to root cause, replicate, and fix. Many consensus issues remain “unsolved”, because they evade reproduction.

    +

    It’s hard to know exactly what the real world consequences of these problems are, because consensus protocols sit so low in the stack. Are they corrupting the archive of someone’s personal D&D game? Security risks at a nuclear plant? Delaying trains in Germany? It all depends on where the protocol’s deployed.

    +

    What’s certain is that these bugs are a colossal waste of developer time, and affect millions, if not hundreds of millions, of users.

    +

    Yet we accept consensus issues and other deep distributed systems bugs as a fact of life — but once upon a time we accepted cholera, and waiting for your turn on the mainframe, as facts of life as well. Developers deserve something better. Everyone who depends on the software we write deserves something better.

    +

    Bugs are not a fact of life

    +

    We have long accepted consensus bugs as inevitable because consensus protocols are just really, really hard to test properly. To really, thoroughly test a Raft implementation — or any distributed system — you effectively need to imagine every possible thing that could go wrong in your entire stack, and write a test to see if it will break your consensus protocol. It’s virtually impossible to write enough tests to provide this kind of assurance in critical systems.

    +

    We get around this in two ways. First, we inflict the testing on our users, in the form of outages, downstream bugs, war rooms, and so on. It’s impossible for a team of developers to write enough tests, but with enough deployments and enough users across enough system configurations, you’re eventually going to find all the problems.

    +

    Second, we rely on formal verification, and this is how consensus algorithms are built today. We define a model that we can mathematically prove to be correct, and then we… translate this perfect, platonic thing into code. Most Raft implementations are built this way.

    +

    There are degrees here. The Raft protocol is one of very few consensus protocols that meets the strictest standard of “formal verification,” with a manual proof, a model check, and a mechanized proof-of-correctness. Every public consensus protocol we’re aware of has a manual proof/written correctness argument, but only Raft, Paxos, and MultiPaxos have mechanized proofs.

    +
    +

    Specifications are not code

    +

    Formal verification is a great starting point, in that it helps confirm that a core design is sound. Implementing a distributed consensus protocol without starting from a formal spec would result in many orders of magnitude more bugs than any Raft implementation has.

    +

    Conversely, one might wonder what the FoundationDB team might have done if they’d started with a formal spec.

    +
    +

    But the issues here demonstrate the difficulties inherent in going from a formal specification to actual production code. As long as the implementations are being done by unreliable, imperfect programmers (whether humans or LLMs), the actual code, and the systems using it, will only be as solid as the implementors’ assumptions — not the formally verified model.

    +

    Furthermore, formal specifications are rarely truly complete — in the case of Raft, the TLA spec covers the core protocol, but doesn’t speak to details like installSnapshot, or replica replacement, or handling network packet corruption.

    +

    Testing can actually be simple

    +

    Surfacing these bugs didn’t require particularly deep knowledge of consensus or distributed systems. We found these using a simple approach, which a junior engineer could implement in less than a day, literally the simplest state-machine replication workload we could come up with. The key is that this is happening while the system is running in Antithesis, subject to the kind of faults that happen in a real world environment.

    +

    We want to emphasize this point because all too often, we encounter a learned helplessness around testing complex systems. Our profession thinks (with some justification) that writing tests for distributed systems requires more expertise, time, or tokens than we have. This results in a state of perpetual under-testing, which in turn results in extraordinary amounts of developer time being wasted on firefighting (to say nothing of the mental and emotional exhaustion).

    +

    Developers deserve something better. Everyone who depends on the software we write deserves something better.

    +

    Testing Raft

    +
    Marco Primi headshot
    Marco Primi Distributed Systems Engineer
    + +

    About Antithesis

    +

    Antithesis is an autonomous testing platform that runs distributed systems in a deterministic simulation environment, with randomly generated inputs, under aggressive fault injection. To use Antithesis, you deploy a full distributed system to the simulation environment along with a workload — a client that drives the system under test.

    +

    Antithesis exposes the system under test to the kind of unpredictable turbulence it will experience in production, within the safety of a simulation environment. By randomizing the inputs and faults, and intelligently searching the state space of the system to see if system invariants are ever violated.

    +

    We maintain an internal curriculum of systems and bugs we use to benchmark Antithesis’ bug-finding performance. We’ve been adding consensus benchmarks to the curriculum, so we’ve been testing a lot of Raft implementations.

    +

    Part of the difficulty of building something unique is that you also need to build a way to measure its performance.

    +
    + + +

    About Raft

    +

    A common architectural pattern in distributed systems is state machine replication (SMR): all replicas in the system run copies of the same deterministic state machine, and the same sequence of commands is delivered to all of them to transform the data/state. This causes the replicas to progress in soft-lockstep — the state after applying N commands is consistent across all replicas, and all replicas eventually apply all commands.

    +

    Distributed databases (FoundationDB, Aerospike, etc.), message queues (Kafka, NATS, etc.), key-value stores (etcd, Zookeeper), blockchains (Bitcoin, Ethereum), and many other systems are based on SMR.

    +

    A fundamental requirement for implementing SMR is that all replicas must receive the same set of commands in the same order — total order delivery. Total order delivery is simple to describe, but hard to architect in the face of network turbulence, process crashes, and disk failures.

    +

    Because total order delivery needs to be guaranteed for SMR to work, most systems rely on one of a handful of messaging abstractions, with atomic broadcast, also known as total order broadcast, being the most common. Raft is viewed as the friendliest atomic broadcast protocol, Paxos and viewstamped replication are also in wide use.

    + +

    Our testing approach

    +

    To test a Raft implementation, we run a 3-node Raft cluster in Antithesis with a simple workload we call Chain of Blocks.

    +

    Chain of Blocks tests Raft’s most important invariant: after N commands are applied, the state of all replicas matches. It consists of two trivially simple components:

    +
      +
    • A single-state state machine that just hashes the bytes of incoming commands. After each applied command, the state consists of <number of commands applied, current hash>. This is sufficient to allow us to observe state divergence between replicas.
    • +
    • A stateless client whose only responsibility is submitting new commands that consist of random arrays of bytes.
    • +
    +

    You can read more about this workload, and see a sample implementation, here.

    +

    Components like Raft are developed and reviewed by experts and battle-hardened by years of exposure. So it’s striking to us that this single, simple testing strategy — that a junior engineer could write in an afternoon (or Claude could write for a handful of tokens) — still finds bugs when it’s run in Antithesis.

    +

    Results

    +

    In all cases, network partitions and turbulence were sufficient to surface examples of divergence (i.e. no node kill/restart required, or disk corruption, or other faults)

    +

    Here’s an example snippet from logs that shows a state machine divergence:

    +
    [node1]: Applied block 244 (2ad8d...) state: 01fd753a => faf6ba1e
    [node2]: Applied block 244 (c92fb...) state: 01fd753a => dab0977c
    [node3]: Applied block 244 (c92fb...) state: 01fd753a => dab0977c
    +

    After applying 243 blocks, the hashes of all replicas matched (0x01fd753a). Node1 then applies a different 244th command from Node2 and Node3 (0x2ad8d… vs 0xc92fb…). This results in a different state hash after 244 commands applied (0xfaf6ba1e vs 0xdab0977c), violating the primary safety property of Raft (and the underlying atomic broadcast): that all replicas apply the same sequence of command in the same order.

    +

    As an example, here’s a detailed look at the bugs we’ve found in one notable open-source Raft implementation during this process. We emphasize that we’ve found similar bugs in other implementations as well, but are going deep rather than broad here and only presenting one set of bugs to keep the length of this post somewhat manageable.

    +

    We plan to update this post with details of bugs in the other Raft implementations we’ve worked with.

    +

    HashiCorp Raft

    +
    Rohan Padhye headshot
    Rohan Padhye Research Fellow
    +

    HashiCorp Raft is a mature, popular open-source Raft implementation, and the foundational consensus engine underpinning widely-used production infrastructure tools like Consul, Nomad, and Vault.

    +

    We do not believe HashiCorp Raft is any less reliable than the other implementations we tested, or that the bugs in HashiCorp Raft are more serious than the bugs in other implementations — they all violate the same core property of state machine safety.

    +

    We found three distinct bugs in total: one causes numerous safety violations (i.e., data divergence as described above) and two cause liveness violations (where some nodes or the whole cluster cannot make progress unless an operator intervenes). The Chain of Blocks implementation we used is here.

    +

    When run with Antithesis for just one hour of testing, we see a report that looks like this:

    +
    Antithesis report.
    +

    Technical primer on Raft

    +

    To better understand the nature of these bugs, we should define some terms used frequently in the Raft paper:

    +
      +
    • “Term”, “Leader”, “Follower”, “Election” + “RequestVote” RPC: Raft ensures consensus among a cluster of distributed nodes by dividing the protocol into strictly increasing terms starting with term=1. In each term, the nodes attempt to elect exactly one node as the leader by taking a majority vote; all other nodes are called followers. Leader election is conducted using a remote-procedure call (RPC) called RequestVote.
    • +
    • “Log”, “Commit”, “Replication” + “AppendEntries” RPC: Each node maintains a replicated log of data entries (e.g., the commands issued by clients), some prefix of which is said to be committed; that is, the node’s state machine is updated when an entry commits. When a client sends some command to the cluster, only the current leader can service this request, and the leader replicates the associated data entry to all its followers via an RPC called AppendEntries. When a majority of nodes in the cluster have replicated the same data entry in their logs, the leader commits it to its own state machine and broadcasts this fact to its followers. If some follower nodes fall behind because of network faults, or if their entry logs have uncommitted data from previous terms, then the current leader can always catch them up via more AppendEntries RPCs. The AppendEntries RPC is also used as a regular heartbeat mechanism to broadcast that a leader is active.
    • +
    • “Compaction”, “Snapshot” + “InstallSnapshot” RPC: When the entry log grows too large, a Raft node can choose to compact all the committed entries in the log and only store on disk a snapshot of the state machine until that point. If a leader node needs to replicate data entries to followers via AppendEntries but its log has already been compacted, it can instead issue an InstallSnapshot RPC to transit the entire state machine to the follower.
    • +
    +

    Bug 1 - Broken Consensus due to Async Heartbeats

    +

    This is the most serious bug we have found. It can affect four of the five safety properties from the Raft paper (Image 3):

    +
      +
    • Log Matching (violated through mechanism A - seen in the report above)
    • +
    • Leader Completeness (violated through mechanism A - seen in the report above)
    • +
    • State Machine Safety (violated through mechanism A - seen in the report above)
    • +
    • Election Safety (violated through mechanism B - not in the report shown above)
    • +
    • Leader Append-Only (not violated)
    • +
    +
    Root Cause and Trigger
    +
      +
    • The Raft paper assumes that all the operations (handling RPCs, client requests, elections, etc.) are atomic and don’t interleave with each other. Essentially, every node in the protocol is a finite-state machine.
    • +
    • In HashiCorp raft, most state-changing operations are handled by a big switch-loop in the main thread, with several other floating goroutines doing async work that should not affect protocol state (e.g., peer-peer “replication” routines for dispatching append-entries, and per-connection “transport” goroutines that listen for incoming RPCs and hand them off to the main thread for processing).
    • +
    • EXCEPT there is an optimization that diverges from the protocol: incoming heart-beat messages (i.e., AppendEntries without an entry) are handled on the I/O “transport” thread itself, instead of queuing it up for the main thread’s loop like with other RPCs. Presumably, this is to quickly reset the keep-alive timers and reduce spurious re-elections.
    • +
    • CAVEAT is that an incoming heart-beat message (like any other AppendEntries RPC) can change the state if it carries a new term when someone else was elected leader, and this bumps up the global currentTerm which everything else in the main thread relies on, and also sets the current state to be FOLLOWER. It is definitely not sound to perform these state changes concurrently with the main thread, and it can lead to different types of race condition bugs.
    • +
    +
    Mechanism A - Incoming heartbeat races with dispatchLogs() on main thread
    +

    Here’s how the bug causes divergence in a 3-node setup A/B/C:

    +
      +
    • Assume the network link between node A and node B is down.
    • +
    • Node A becomes candidate for term T, requests votes (eventually gets it from Node C)
    • +
    • Node B also runs for term T but no votes
    • +
    • Node B becomes candidate for term T+1, requests votes (gets it from Node C)
    • +
    • Node B wins election for term T+1
    • +
    • Node A wins election for term T (just received the old vote from Node C)
    • +
    • Node B sends out heart-beat AppendEntries with term=T+1, which Node A doesn’t immediately receive
    • +
    • Node A gets a client request to apply a new data entry X, and it still thinks it is a leader for term T, so on the main thread it prepares to create a log entry (data=X, term=T) and send out AppendEntries to peers. +
        +
      • HOWEVER: Just before it can create the log entry, the heart-beat from Node B sent in step 7 above reaches, setting currentTerm=T+1. Because this is done on the fast-path on the network-transport thread, it races with the main thread.
      • +
      • The main thread ends up preparing a log entry (data=X, term=T+1) and also dispatching AppendEntries with these values before the next main-loop iteration where it realizes it is actually now a follower for term T+1.
      • +
      +
    • +
    • Things get really bad from here. Some nodes apply this bogus data=X, term=T+1 to their logs, while others follow whatever Node B (the true leader for term T+1) says, e.g., they might apply data=Y at the same index with term=T+1. Logs diverge in data but get committed since all say term T+1 for the same index. State machines diverge.
    • +
    +
    Mechanism B - Incoming heartbeat races with requestVote() on main thread
    +

    The same bug can cause another kind of safety violation when the async heartbeat handler races with a concurrent handler for the RequestVote RPC on the main thread. If both the incoming heartbeat and the RequestVote carry higher-numbered but different terms T1 and T2 respectively — this is possible when the receiver just recovers from a fault and is catching up to queued messages from peers — and if T1 is less than T2, then it is possible for the receiver to (i) realize the heartbeat’s term T1 is higher than its current term, then (2) on the main thread processing RequestVote realize that T2 is higher than its current term and so set its current term to T2 and then grant a vote; and then (3) back on the heartbeat handler’s thread set its current term to the lower value T1. Needless to say, it should never be possible for a node that has granted a vote for term T2 to regress its current term back down to a lower value T1! When this node at some later point increments its term to T2 again, it has no memory of the fact that it has already voted in this term.

    +

    In summary, the race can eventually cause a node to grant two votes in the same term (T2 above), which in the worst case can lead to two different nodes being elected as leader in the same term! Here are some sample log messages from HashiCorp Raft when we have observed this race condition and its violation of election safety:

    +
    20:11:59.941133 [DEBUG] node-1: vote granted: from="node-2" tally=2 term=286
    [...]
    20:11:59.941133 [INFO] node-1: election won: tally=2 term=286
    [...]
    20:11:59.944273 [DEBUG] node-0: vote granted: from="node-2" tally=2 term=286
    [...]
    20:11:59.944274 [INFO] node-0: election won: tally=2 term=286
    +

    Bug 2 - Deadlock after Leadership Transfer

    +

    This bug causes a liveness issue — a leader node is unable to service client requests or commit new entries while the cluster is healthy. The root cause is a deadlock during a leadership transfer operation (a Raft protocol extension used by Consul where you can ask a current leader to transfer leadership to someone else, say for upgrades) that does not immediately signal any warnings but bites you way into the future by getting the whole cluster stuck.

    +
      +
    • When you initiate a leadership transfer from A to B the current leader A first tries to get B’s logs up to date via an async replication thread.
    • +
    • The leadership transfer logic is a separate goroutine that waits for this replication before telling B to take over.
    • +
    • While all this is happening, A might have to step-down as leader for unrelated reasons (e.g., the lease timer runs out, or it cannot reach a quorum due to network) and say another node C becomes leader.
    • +
    • When A steps-down as leader, it aborts all replication threads (including the A—>B catchup) but the leadership transfer goroutine is still blocked waiting for it to complete (an implementation bug).
    • +
    +

    There is no problem so far, because C is the new leader and it does its job for a while.

    +
      +
    • Unfortunately, the leadership transfer goroutine in A has set a global flag “leadership transfer in progress” and this flag cannot be unset until it is unblocked… but it never will be!!!
    • +
    • Crucially, this flag does not prevent it from participating in elections and the flag does not get reset if it wins a future election.
    • +
    • So in the future, A can legitimately get re-elected as leader but it will refuse to accept any client requests and stall on future commits with the reason “leadership transfer in progress”.
    • +
    • When the network is healthy and A is a leader, there is no way to get out of this other than rebooting the node A to clear the flag.
    • +
    +

    We caught this using an Antithesis eventually test command that checks for commit progress after fault injection is turned off.

    +

    Bug 3 - Livelock in Snapshot Installation

    +

    This bug causes a liveness issue: a follower node becomes unable to replicate log entries and is effectively not able to participate in the consensus protocol. The bug can also cause resource exhaustion, where the node starts creating a potentially unbounded number of temporary snapshot files on disk.

    +

    The root cause for this bug is painfully simple: in the version of HashiCorp Raft we tested, the implementation of the InstallSnapshot RPC simply does not follow Rule 7 from Image 13 in the Raft paper which handles the case when an incoming snapshot represents a state that diverges from the receiver’s uncommitted log entries. In the paper, Rule 7 says:

    +

    “discard the entire log”

    +

    HashiCorp Raft does not discard the existing log, and so after installing the snapshot its state machine can, in some cases, disagree with the data in its own (stale) log entries.

    +

    At first, this might sound benign because the state machine should be the final source of truth. However, when the node receives a subsequent AppendEntries RPC from the leader, it faithfully implements a part of the protocol (Rule 2) that says: “[reject AppendEntries] if log doesn’t contain an entry at prevLogIndex whose term matches prevLogTerm”. So, because of the stale logs, the follower rejects subsequent AppendEntries, causing the leader to respond with a new InstallSnapshot, and this cycle goes on forever.

    +

    We caught this using an Antithesis eventually test command that checks for state-machine convergence after fault injection is turned off. Sample logs from a test run look like this:

    +
    2026-07-07T17:17:13.073Z [INFO] node-0: snapshot network transfer progress: read-bytes=0 percent-complete="0.00%"
    2026-07-07T17:17:15.954Z [ERROR] node-1: failed to appendEntries to: peer="{Voter node-0 raft-node-0:8300}" error="msgpack decode error [pos 29848]: read tcp 10.89.0.7:48504->10.89.0.4:8300: i/o timeout"
    2026-07-07T17:17:23.073Z [INFO] node-0: snapshot network transfer progress: read-bytes=0 percent-complete="0.00%"
    2026-07-07T17:17:26.038Z [ERROR] node-1: failed to appendEntries to: peer="{Voter node-0 raft-node-0:8300}" error="msgpack decode error [pos 29848]: read tcp 10.89.0.7:35706->10.89.0.4:8300: i/o timeout"
    2026-07-07T17:17:33.073Z [INFO] node-0: snapshot network transfer progress: read-bytes=0 percent-complete="0.00%"
    2026-07-07T17:17:36.135Z [ERROR] node-1: failed to appendEntries to: peer="{Voter node-0 raft-node-0:8300}" error="msgpack decode error [pos 30394]: read tcp 10.89.0.7:37032->10.89.0.4:8300: i/o timeout"`
    2026-07-07T17:17:43.072Z [INFO] node-0: snapshot network transfer progress: read-bytes=0 percent-complete="0.00%"
    2026-07-07T17:17:46.249Z [ERROR] node-1: failed to appendEntries to: peer="{Voter node-0 raft-node-0:8300}" error="msgpack decode error [pos 29575]: read tcp 10.89.0.7:54324->10.89.0.4:8300: i/o timeout"`
    2026-07-07T17:17:53.073Z [INFO] node-0: snapshot network transfer progress: read-bytes=0 percent-complete="0.00%"
    2026-07-07T17:17:56.381Z [ERROR] node-1: failed to appendEntries to: peer="{Voter node-0 raft-node-0:8300}" error="msgpack decode error [pos 30303]: read tcp 10.89.0.7:57842->10.89.0.4:8300: i/o timeout"`
    2026-07-07T17:18:03.072Z [INFO] node-0: snapshot network transfer progress: read-bytes=0 percent-complete="0.00%"
    2026-07-07T17:18:06.626Z [ERROR] node-1: failed to appendEntries to: peer="{Voter node-0 raft-node-0:8300}" error="msgpack decode error [pos 31122]: read tcp 10.89.0.7:46978->10.89.0.4:8300: i/o timeout"
    2026-07-07T17:18:13.073Z [INFO] node-0: snapshot network transfer progress: read-bytes=0 percent-complete="0.00%"
    2026-07-07T17:18:17.032Z [ERROR] node-1: failed to appendEntries to: peer="{Voter node-0 raft-node-0:8300}" error="msgpack decode error [pos 30758]: read tcp 10.89.0.7:54192->10.89.0.4:8300: i/o timeout"
    2026-07-07T17:18:23.073Z [INFO] node-0: snapshot network transfer progress: read-bytes=0 percent-complete="0.00%"
    2026-07-07T17:18:27.745Z [ERROR] node-1: failed to appendEntries to: peer="{Voter node-0 raft-node-0:8300}" error="msgpack decode error [pos 31850]: read tcp 10.89.0.7:36218->10.89.0.4:8300: i/o timeout"
    2026-07-07T17:18:33.073Z [INFO] node-0: snapshot network transfer progress: read-bytes=0 percent-complete="0.00%"
    2026-07-07T17:18:39.083Z [ERROR] node-1: failed to appendEntries to: peer="{Voter node-0 raft-node-0:8300}" error="msgpack decode error [pos 33943]: read tcp 10.89.0.7:53700->10.89.0.4:8300: i/o timeout"
    2026-07-07T17:18:43.073Z [INFO] node-0: snapshot network transfer progress: read-bytes=0 percent-complete="0.00%"
    2026-07-07T17:18:51.725Z [ERROR] node-1: failed to appendEntries to: peer="{Voter node-0 raft-node-0:8300}" error="msgpack decode error [pos 38402]: read tcp 10.89.0.7:52390->10.89.0.4:8300: i/o timeout"
    2026-07-07T17:18:53.073Z [INFO] node-0: snapshot network transfer progress: read-bytes=0 percent-complete="0.00%"
    +
    2026-07-07T17:18:55.497Z [ERROR] eventually-fsm-convergence: convergence-check give-up; cluster healthy but did not converge: attempts_taken=24 per_node="map[raft-node-0:8400:map[applied_index:534 last_index:535 reachable:true state:Follower term:74 value:23641] raft-node-1:8400:map[applied_index:569 last_index:569 reachable:true state:Leader term:74 value:25232] raft-node-2:8400:map[applied_index:569 last_index:569 reachable:true state:Follower term:74 value:25232]]"
    +

    Reflections

    +
    TW Lim headshot
    TW Lim Technical Writer
    +

    If there is a single lesson to be learned here, it’s that formal methods alone cannot ensure that software works, because the formal specification still needs to be implemented, and even if you have mechanized verification, you’re verifying the model and not the implementation itself.

    +

    We believe formal methods are useful and necessary — they can confirm the basic soundness of a design, and provide a map that saves engineers from many of the errors that can arise in the implementation of a complex system.

    +

    But as these bugs show, errors continue to arise when translating the formal specification to production code. In the course of our work with various Raft implementations, we identified a number of assumptions in the Raft paper that remain implicit. An implementer who misses any of these details is likely to run into trouble.

    +

    Assumption 1: Each node is a synchronous process

    +
    Marco Primi headshot
    Marco Primi Distributed Systems Engineer
    +

    The paper implicitly assumes that each node is a synchronous process, performing atomic state transitions when handling RPCs and updating its own internal state. The TLA+ spec is designed as such.

    +

    At the same time, none of the implementations we’ve looked at have actually been a synchronous process, and there’s no explicit guidance in the implementation guide as to whether and how one can deviate from the synchronous design.

    +

    Consider the following situation: +Can a leader with an outstanding RPC request respond to a vote, or does it need to wait for a response||timeout? Or should it respond, shut down the request and disregard a future response in order to stay correct?

    +

    HashiCorp Raft is mostly synchronous in handling RPCs, with a single thread in a select loop, except when it isn’t — it does asynchronous heartbeat handling, bypassing its main loop, and this is what allows Bug #1 to creep in.

    +

    Assumption 2: Each response can be mapped to a request

    +

    In the Raft paper, every call is made using RPC, which means that every response can be mapped to a request.

    +

    If one implements Raft on top of simple TCP or UDP, and doesn’t realize that request/response correlation is a crucial part of correctness, bad stuff can easily happen.

    +

    If you’re not working in a language with a nice RPC library, it’s not obvious how to establish this correspondence. For example, the Raft authors have a simulation on their website, which for obvious reasons is often considered a demo/canonical implementation even though they’ve never described it as such. But even their own implementation deviates from the protocol and includes an extra matchIndex field in the AppendEntries response in order to avoid matching requests to responses.

    +

    Assumption 3: currentTerm and votedFor are consistent

    +
    Rohan Padhye headshot
    Rohan Padhye Research Fellow
    +

    Image 2 of the Raft paper states:

    +
    Raft paper figure 2 snippet.
    +

    One key assumption is that currentTerm and votedFor are consistent, because the latter records the vote in the current term. If these go out of sync bad things can happen, especially if votedFor is null.

    +

    For safety, the writing of these values to persistent storage should be done atomically. But in the protocol, these values actually change at different places, so ensuring that both values are updated together is subtle.

    +

    The online simulator in JavaScript does not actually distinguish persistent state from in-memory state, so there is no “reference implementation” of how to do this.

    +

    HashiCorp Raft, for instance, deviates from the protocol and non-atomically stores three separate values for currentTerm, lastVoteTerm and lastVoteCand, with a worst-case risk to performance but not safety (i.e., it might persist an updated lastVoteTerm and then crash before updating lastVoteCand, but when it restarts it will just have a pre-determined vote for that term instead of choosing a candidate as normal).

    +

    Assumption 4: The entire protocol is formally verified

    +

    Figures 2 and 13 in the Raft paper state, respectively:

    +
    Raft paper figure 2 snippet.
    +
    Raft paper figure 13.
    +

    Rule 2 for AppendEntries and Rule 6 for InstallSnapshot creates a potential pitfall. If you implement this naively when you have discarded your log because of a previous InstallSnapshot, you end up doing the wrong thing (most likely creating infinite loops of RPCs between leader and follower, as I have painfully discovered).

    +

    The Raft TLA+ spec and the authors’ simulation don’t actually include InstallSnapshot (though Diego Ongaro’s PhD thesis does discuss the nuances in more depth), so different implementations use different workarounds. HashiCorp Raft, for instance, just avoids discarding logs altogether — which led to Bug #3 above.

    +

    Bug #2 above, the deadlock, is in a feature called “leadership transfer” which is an extension to the core protocol. Like InstallSnapshot, this feature was never formally verified in the TLA+ spec.

    +

    Conclusion

    +
    TW Lim headshot
    TW Lim Technical Writer
    +

    We have long accepted bugs as inevitable. Distributed systems are almost impossible to test thoroughly because their state spaces are so large, and as our stacks get more complex, the problem only gets worse.

    +

    But technological progress also moves the frontier of what’s possible. We ended cholera in the developed world by building water treatment plants and sewage systems. We no longer wait for mainframes because we’ve made computing power cheap and abundant.

    +

    The abundance of compute is a recent phenomenon. It’s held true for maybe the last 15 years, less than a tenth of the history of computing. It’s spurred, among other things, the recent surge of interest in formal verification — the hope that mechanized proof and model checking can finally become easy enough, and cheap enough, that we’ll be able to apply these approaches wherever we need. But what our experience with Raft has shown is that implementations of even mechanically proven models can have flaws, because there’s no mechanical way to check the implementors’ assumptions.

    +

    Fortunately, cheap and abundant compute has also made it possible to test software in ways we couldn’t before. And the other thing our testing of Raft has shown is that when you run systems under fault, even basic testing methods reveal issues that would once have taken months of manual testing to find.

    +

    Developers deserve something better, and so does everyone who depends on our software.

    get started

    Velocity and verification. Finally, both.

    +Talk with one of our product experts to see how Antithesis can help + you ship confidently, no matter who’s writing your code. +

    \ No newline at end of file diff --git a/sreweekly/articles/529/06-how-to-build-your-infrastructure-monitoring-in-2026.html b/sreweekly/articles/529/06-how-to-build-your-infrastructure-monitoring-in-2026.html new file mode 100644 index 00000000..52e36a6a --- /dev/null +++ b/sreweekly/articles/529/06-how-to-build-your-infrastructure-monitoring-in-2026.html @@ -0,0 +1,106 @@ +How to Build Your Infrastructure Monitoring in 2026 ·
    How to Build Your Infrastructure Monitoring in 2026

    How to Build Your Infrastructure Monitoring in 2026

    August 5, 2026 +· observability, monitoring, sre, opentelemetry, victoriametrics, loki, jaeger, vector

    Every year I get asked the same question by teams starting from scratch: “we have Grafana, we have some dashboards, why do we still get paged for things we didn’t see coming?” Most of the time, the answer isn’t a missing tool. It’s a missing method. Teams jump straight to “let’s install Prometheus” or “let’s buy a SaaS observability platform” before answering a much simpler question: what does “healthy” actually mean for this business?

    I’ve built infrastructure monitoring from the ground up for several companies now, and I keep coming back to the same seven steps. This article is that playbook, the way I actually apply it in 2026.

    Requirements

    Before you touch any tool, you need:

    • A clear list of the critical business flows your system supports (payments, checkouts, logins, API calls…)
    • Buy-in from the team on what “acceptable” looks like for those flows
    • A telemetry stack that can handle metrics, logs, and traces (I’ll give you mine below)

    If you get stuck at any point: reach out, I’m happy to help you think through your specific setup.

    1. Start with the business SLI/SLO, not with the tool

    This is the step almost everyone skips, and it’s the one that matters the most. Before deciding what to monitor, decide what “working” means for your business.

    An SLI (Service Level Indicator) is a metric that reflects user-facing behavior. An SLO (Service Level Objective) is the target you set for that metric over a time window.

    Example, if you work for a banking company:

    • SLI: the ratio of successful payment authorizations over total payment authorization attempts
    • SLO: 99.95% of payment authorizations should succeed over a rolling 30-day window

    That single sentence changes everything downstream. It tells you:

    • Which service is “tier 0” (payment authorization service)
    • What your error budget is (0.05% of failed authorizations per month)
    • What should page someone at 3am, and what can wait for Monday morning

    Do this exercise for every critical business flow before writing a single scrape config. If you skip it, you’ll end up monitoring infrastructure CPU graphs while your actual business metric silently burns through its error budget.

    2. Know what to monitor, then pick your stack

    Once your SLIs/SLOs are defined, list what you actually need visibility into to measure them:

    • Infrastructure: nodes, Kubernetes clusters, network, databases, message queues
    • Applications: HTTP servers, HTTP clients, background jobs, gRPC services
    • Business layer: the actual events tied to your SLI (a payment authorization call, a checkout event…)

    Only now do you pick the tech stack, because now you know what it needs to support. Here’s the generic stack I use on most projects:

    PillarToolRole
    MetricsVictoriaMetricsLong-term, cost-efficient metrics storage (Prometheus-compatible)
    Metrics agentvmagentScraping and remote-writing metrics
    LogsLokiLog aggregation, indexed by labels not full text
    Logs & traces ingestionOpenTelemetry CollectorVendor-neutral receiver/processor/exporter pipeline
    TracesJaegerDistributed trace storage and visualization

    The three pillars, and what each is actually for

    It’s worth being explicit about this, because teams often use the wrong pillar to answer the wrong question:

    • Metrics: aggregated, cheap to store, great for trending and alerting. They answer “what” and “how much” (error rate is 2%, p99 latency is 800ms).
    • Logs: high cardinality, detailed, expensive to store at full fidelity. They answer “why” during an investigation (this specific request failed because of X).
    • Traces: the causal chain across services. They answer “where” in a distributed call the latency or error actually happened.

    None of the three replaces the others. Metrics tell you something is wrong, traces tell you where, logs tell you why.

    3. Implement and scrape, favor auto-instrumentation

    Now you build the pipeline. My rule of thumb: instrument automatically first, add manual instrumentation only where auto-instrumentation doesn’t reach (custom business logic, internal queues, batch jobs).

    Use the OpenTelemetry auto-instrumentation libraries for your language. They hook into common frameworks (HTTP servers, HTTP clients, database drivers, gRPC) and emit metrics, traces, and sometimes logs without you writing a single line of instrumentation code.

    Example, a Java service auto-instrumented and shipped straight to your collector, no code change required:

    java -javaagent:opentelemetry-javaagent.jar \
    +  -Dotel.service.name=payment-authorization-service \
    +  -Dotel.exporter.otlp.endpoint=http://otel-collector:4317 \
    +  -Dotel.metrics.exporter=otlp \
    +  -Dotel.traces.exporter=otlp \
    +  -Dotel.logs.exporter=otlp \
    +  -jar payment-service.jar
    +

    On the infrastructure side, vmagent scrapes your Prometheus-format endpoints (node_exporter, kube-state-metrics, cAdvisor, database exporters…):

    # vmagent scrape config
    +scrape_configs:
    +  - job_name: 'kubernetes-pods'
    +    kubernetes_sd_configs:
    +      - role: pod
    +    relabel_configs:
    +      - source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_scrape]
    +        action: keep
    +        regex: true
    +

    4. Build a telemetry pipeline that correlates, not just collects

    Collecting metrics, logs, and traces separately gets you three silos, not observability. The fix is enrichment: attach the same standard labels to everything, so you can pivot from a metric to a trace to a log for the exact same request.

    I always enforce the OpenTelemetry semantic conventions resource attributes as a baseline on every service:

    • service.name: the logical service, e.g. payment-authorization-service
    • service.namespace: the domain/team owning it, e.g. payments
    • deployment.environment: production, staging, etc.

    Enforce this at the OpenTelemetry Collector level so nothing gets ingested without it:

    # otel-collector-config.yaml
    +processors:
    +  resource:
    +    attributes:
    +      - key: service.namespace
    +        value: payments
    +        action: insert
    +      - key: deployment.environment
    +        value: production
    +        action: insert
    +
    +service:
    +  pipelines:
    +    traces:
    +      receivers: [otlp]
    +      processors: [resource, batch]
    +      exporters: [otlp/jaeger]
    +    metrics:
    +      receivers: [otlp]
    +      processors: [resource, batch]
    +      exporters: [prometheusremotewrite/victoriametrics]
    +    logs:
    +      receivers: [otlp]
    +      processors: [resource, batch]
    +      exporters: [loki]
    +

    With these three labels shared across your metrics, logs, and traces, you can go from “error rate spiked on payment-authorization-service in production” straight to the matching traces and logs, without guessing.

    5. Build a RED dashboard, before anything fancier

    Once telemetry is correlated by service_name and service_namespace, build one dashboard before all others: the RED dashboard (Rate, Errors, Duration).

    Apply it to three things per service:

    • HTTP server requests (inbound traffic)
    • HTTP client requests (outbound calls to dependencies)
    • Span metrics (generated from traces, gives you RED per operation, not just per HTTP route)

    Example PromQL for the “R” and “E” of a service, using OpenTelemetry’s standard http.server.request.duration metric:

    # Request rate
    +sum(rate(http_server_request_duration_seconds_count{service_namespace="payments"}[5m])) by (service_name)
    +
    +# Error rate (%)
    +sum(rate(http_server_request_duration_seconds_count{service_namespace="payments", http_response_status_code=~"5.."}[5m])) by (service_name)
    +/
    +sum(rate(http_server_request_duration_seconds_count{service_namespace="payments"}[5m])) by (service_name)
    +* 100
    +
    +# Duration (p99)
    +histogram_quantile(0.99, sum(rate(http_server_request_duration_seconds_bucket{service_namespace="payments"}[5m])) by (service_name, le))
    +

    This one dashboard, applied consistently across every service, gives you a clear signal of application health over time before you’ve written a single custom panel. Everything else (business dashboards, infra dashboards, deep-dive panels) builds on top of it.

    6. Alert on symptoms, not noise, and use multi-window burn rate

    This is where most on-call setups fail. Two rules I don’t compromise on:

    • Alert on actionable conditions, not informative ones. “CPU is at 80%” is informative. “The payment SLO error budget will be exhausted in 2 hours at this burn rate” is actionable. If an alert doesn’t require a human to do something right now, it shouldn’t page anyone; it belongs on a dashboard.
    • Use multi-window, multi-burn-rate alerting so you can tell a critical incident from something that can wait until tomorrow. This is straight out of the Google SRE book’s alerting chapter, and it’s the single highest-leverage thing you can implement for on-call sanity.

    The idea: page immediately only when the error budget is burning fast enough that waiting would breach the SLO. Use a short window to catch fast burns and a longer window to confirm it isn’t a blip, and use a slower threshold for tickets instead of pages.

    Back to our banking example: SLO is 99.95% success over 30 days, meaning an error budget of 0.05%.

    # VictoriaMetrics / Prometheus alerting rules
    +groups:
    +  - name: payment-authorization-slo
    +    rules:
    +      # Fast burn: page immediately.
    +      # Burning 14.4x the allowed rate would exhaust the 30-day budget in ~2 days.
    +      # Confirmed over both a 5m and 1h window to avoid paging on a blip.
    +      - alert: PaymentAuthSLOFastBurn
    +        expr: |
    +          (
    +            sum(rate(payment_authorization_failed_total{service_namespace="payments"}[5m]))
    +            /
    +            sum(rate(payment_authorization_total{service_namespace="payments"}[5m]))
    +          ) > (14.4 * 0.0005)
    +          and
    +          (
    +            sum(rate(payment_authorization_failed_total{service_namespace="payments"}[1h]))
    +            /
    +            sum(rate(payment_authorization_total{service_namespace="payments"}[1h]))
    +          ) > (14.4 * 0.0005)          
    +        labels:
    +          severity: page
    +        annotations:
    +          summary: "Payment authorization burning error budget fast, will breach SLO in ~2 days if it continues"
    +
    +      # Slow burn: create a ticket, review during business hours.
    +      # Burning 3x the allowed rate would exhaust the budget in ~10 days.
    +      - alert: PaymentAuthSLOSlowBurn
    +        expr: |
    +          (
    +            sum(rate(payment_authorization_failed_total{service_namespace="payments"}[1h]))
    +            /
    +            sum(rate(payment_authorization_total{service_namespace="payments"}[1h]))
    +          ) > (3 * 0.0005)
    +          and
    +          (
    +            sum(rate(payment_authorization_failed_total{service_namespace="payments"}[6h]))
    +            /
    +            sum(rate(payment_authorization_total{service_namespace="payments"}[6h]))
    +          ) > (3 * 0.0005)          
    +        labels:
    +          severity: ticket
    +        annotations:
    +          summary: "Payment authorization error budget burning steadily, investigate this week"
    +

    What this buys you on-call:

    • A fast, confirmed burn pages someone at 3am, because at that rate you’ll breach the monthly SLO within days.
    • A slow burn opens a ticket instead of paging, because at that rate you have days to weeks before the budget is exhausted.

    This is the difference between an on-call rotation that burns people out on noise, and one that pages only when it truly matters.

    7. You now have the method, not just the tools

    At this point you have: business-defined SLIs/SLOs, a stack chosen because it fits what you need to monitor, auto-instrumented metrics/logs/traces, a correlated telemetry pipeline via standard labels, a RED dashboard per service, and burn-rate alerting that tells critical from “can wait.”

    That’s not a finished monitoring setup, it’s a solid foundation you can build on: business dashboards, capacity planning, chaos testing against your SLOs, whatever comes next for your organization.

    Conclusion

    If this was helpful, leave a comment and tell me how your monitoring setup looks today. If you’d like help implementing any of these steps for your specific stack, I’m happy to walk through it with you.

    I wish you calm on-call shifts!

    Building something like this in production?

    I help teams turn setups like this into reliable, monitored infrastructure.

    Get a free consulting call

    Get my monitoring stack checklist

    The exact checklist I use when setting up observability for a new team. No spam, unsubscribe anytime.

    \ No newline at end of file diff --git a/sreweekly/articles/529/07-getting-access-to-the-tmp-of-a-systemd-service-with-privatetmp-yes.html b/sreweekly/articles/529/07-getting-access-to-the-tmp-of-a-systemd-service-with-privatetmp-yes.html new file mode 100644 index 00000000..d35c46d6 --- /dev/null +++ b/sreweekly/articles/529/07-getting-access-to-the-tmp-of-a-systemd-service-with-privatetmp-yes.html @@ -0,0 +1,78 @@ + + +You're using a too-old browser + + + +

    You're using a suspiciously old browser

    +

    You're probably reading this page because you've attempted to access +some part of my blog (Wandering Thoughts) or +CSpace, the wiki thing it's part of. Unfortunately +you're using a browser version that my anti-crawler precautions consider +suspicious, most often because it's too old (most often this applies to +versions of Chrome). Unfortunately, as of early 2025 there's a plague +of high volume crawlers (apparently in part to gather data for LLM +training) that use a variety of old browser user agents, especially +Chrome user agents. To reduce the load on +Wandering Thoughts I'm experimenting with +(attempting to) block all of them, and you've run into this.

    + +

    If this is in error and you're using a current version of your +browser of choice, you can contact me at my current place at the +university (you should be able to work out the email address +from that). If possible, please let me know what browser you're +using and so on, ideally with its exact User-Agent string.

    + +

    A special note to people using Inoreader (the feed reader)

    + +

    +I am not blocking Inoreader's feed fetcher or considering it to be +too old, and it routinely fetches feeds from me. I don't know why +Inoreader is showing you this page. It is possible that they're +periodically trying to fetch feeds or pages with an old browser +HTTP User-Agent (or an actual old browser) and taking the results +of that fetch (this page) as what they should show people instead +of the results of their syndication feed fetcher agents. This is a +bad mistake today; the results +of modern HTTP fetches depend partly on the HTTP User-Agent used. +

    + +

    A special note to people using Feedly (the feed reader)

    + +

    Much like Inoreader, Feedly is periodically fetching my syndication +feeds with a fake, old browser HTTP User-Agent header, which fails, +and is then grimly latching on to the results for their actual feed +fetching with their regular Feedly HTTP User-Agent. There is nothing +I can do about this; you should contact Feedly support, if you can +find them. See this +comment of mine in Wandering Thoughts for more details. +

    + +

    A special note for people using Vivaldi

    + +

    Due to an ongoing attack, you may need to change +the +"User Agent Brand Masking" setting so that your Vivaldi identifies +itself as Vivaldi, instead of Google Chrome. This applies to even +the current version of Vivaldi.

    + +

    A special note for people using archive.*

    + +

    You may be seeing this through archive.today, archive.ph, archive.is, +and so on. Unfortunately, archive.* crawls pages to archive in a way that +is impossible to distinguish from malicious actors. They use old Chrome +User-Agent values, crawl from IP address blocks that are widely distributed +and not clearly identified as theirs, and some of their IP addresses have +falsified reverse DNS entries that claim they are googlebot IP addresses +(which is something that is normally done only by quite bad actors). I +suggest that you use archive.org, which is a better behaved archival +crawler and can crawl my blog (Wandering Thoughts). +

    + +
    Chris Siebenmann, 2025-02-17
    + + diff --git a/sreweekly/articles/529/08-traditional-versus-resilience-engineering-views.html b/sreweekly/articles/529/08-traditional-versus-resilience-engineering-views.html new file mode 100644 index 00000000..d2ddaf2f --- /dev/null +++ b/sreweekly/articles/529/08-traditional-versus-resilience-engineering-views.html @@ -0,0 +1,967 @@ + + + + + + + +Traditional versus resilience engineering views – Surfing Complexity + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +
    + + +
    + +
    + + + + +
    +
    + +
    +
    + + + +
    +
    +

    Traditional versus resilience engineering views

    + +
    + +

    As a fan of resilience engineering, I often differ with people on where we should focus our scarce engineering cycles in order to improve reliability.

    + + + +

    I thought it would be a useful exercise to brainstorm some of the differences in focus between what I’ll call the traditional view of reliability, and the resilience engineering view.

    + + + +
    Traditional view focuses on Resilience engineering view focuses on
    accountabilitycoordination
    prioritizationgoal conflicts
    risk mitigationrisk trade-offs
    better processes and conformance thereofmore expertise
    quantitativequalitative
    root causeinteraction of multiple factors
    action itemsinsight
    preventing future incidents,
    ensuring all incidents are novel
    better handling of novel incidents
    reducing complexitynavigating complexity
    objectivesproduction pressure
    robustnessresilience
    human variability as liabilityhuman variability as asset
    building accurate system modelrepairing inevitable model errors
    rigorimprovisation
    explicit knowledgetacit knowledge
    automation, benefits ofautomation, risks introduced by
    + + + +

    +
    + + + + +
    + + + + +
    + + +

    + One thought on “Traditional versus resilience engineering views”

    + + +
      +
    1. + +
    2. +
    + + + + +
    +

    Leave a comment

    + + +
    + + + +

    +

    + +
    + + +
    +
    + + + +
    + + +
    +
    + + + + + + + + + +
    +
    +
    +
    + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + diff --git a/sreweekly/articles/529/index.json b/sreweekly/articles/529/index.json new file mode 100644 index 00000000..ee8905ca --- /dev/null +++ b/sreweekly/articles/529/index.json @@ -0,0 +1,50 @@ +[ + { + "idx": 1, + "url": "https://greatcircle.com/blog/2026/07/14/incident-management-process-versus-program/", + "ok": true, + "error": null + }, + { + "idx": 2, + "url": "https://www.honeycomb.io/blog/what-comes-after-observability", + "ok": true, + "error": null + }, + { + "idx": 3, + "url": "https://dzone.com/articles/rise-of-agentic-sre", + "ok": true, + "error": null + }, + { + "idx": 4, + "url": "https://blog.railway.com/p/incident-report-july-2-2026-us-east-services-outage", + "ok": true, + "error": null + }, + { + "idx": 5, + "url": "https://antithesis.com/blog/2026/finding-bugs-in-raft-implementations/", + "ok": true, + "error": null + }, + { + "idx": 6, + "url": "https://omarghader.github.io/monitoring-infrastructure-guide-2026/", + "ok": true, + "error": null + }, + { + "idx": 7, + "url": "https://utcc.utoronto.ca/~cks/space/blog/linux/SystemdPrivateTmpWhere", + "ok": true, + "error": null + }, + { + "idx": 8, + "url": "https://surfingcomplexity.blog/2026/08/02/traditional-versus-resilience-engineering-views/", + "ok": true, + "error": null + } +] \ No newline at end of file diff --git a/sreweekly/articles/530/01-expertise-can-t-be-automated-why-resilience-still-needs-humans.html b/sreweekly/articles/530/01-expertise-can-t-be-automated-why-resilience-still-needs-humans.html new file mode 100644 index 00000000..b0ddf13b --- /dev/null +++ b/sreweekly/articles/530/01-expertise-can-t-be-automated-why-resilience-still-needs-humans.html @@ -0,0 +1,419 @@ +Expertise Can't Be Automated: Why Resilience Still Needs Humans | Resilience in Software FoundationSkip to content
    Resilience in Software Foundation logo

    Expertise Can't Be Automated: Why Resilience Still Needs Humans

    Published on August 3, 2026

    Abby Wambach, Neil Peart, Kelsey Hightower. These people are all experts in their relevant fields. Through them we recognize expertise as an exceptionally high level of performance on a particular task or within a given domain. The universal inputs to their expertise are time and exposure (aka practice). Expertise in any domain is acquired through doing, and more specifically by continued exposure to varying conditions and inputs inherent to the given domain. The latter part is important: while someone can absolutely become an expert in a very narrow sense (e.g. solving Rubik’s cubes), and someone can also learn about a subject or domain by reading, talking, listening, watching podcasts, etc., high-performance domains exert unpredictable, variable, and dynamic demands on the expert that are only mastered by continued exposure and deliberate practice within that domain.

    +
    +

    ”Experts are the people the team turns to when faced with difficult tasks.”

    +

    —Gary Klein

    +
    +

    I have recently been on a number of podcasts/webinars/panels where the question of what expert incident response looks like comes up, especially how to train for and formalize it. Lurking behind that question for some people is often the hope that in the answer lies the opportunity to productize, scale, and even automate this essential skill for modern software-based businesses. We are currently awash in AI Ops and AI SRE solutions attempting to do just this, and I’m here to convince you that expertise cannot be bottled and sold. Beyond that, I hope to persuade you that expertise is worth investing in and cultivating within your own organization.

    +

    Common Characteristics of Expertise

    +

    It is worth noting that expertise is not simply the accrual of experience over time. As hinted at in the earlier quote from Gary Klein, expertise is a specific kind of knowledge, and it shares a number of things in common.

    +

    Expertise Is Often Invisible

    +

    Cognitive scientists and folks in the field of Resilience Engineering refer to this as the Law of Fluency. This fluency characterizes an activity that is well-adapted, such that the effort and challenges involved in conducting that work are hidden from view, making it appear smooth and effortless. Experts adapt to complex and surprising situations by filling gaps and managing challenges effectively and efficiently, and notably in ways that may not be apparent to other people. In general, fluency is considered a hallmark characteristic of expertise.

    +

    Expertise Is Not Easily Introspected

    +

    In general, experts do not have direct access to the cognitive, physical, and other sources of their expertise—human performance, given the way it is acquired through experience over time, is inherently difficult to explain. This is why many people refer to this kind of skill as implicit, or subconscious. Answering the question “How did you know to do X?” for an expert often leads to a long exploration of “Well, it reminded me of this one time…” or “It seemed similar to when…” In asking someone to explain their expertise, we are effectively asking them to try to unearth every second of practice and effort they have put into acquiring their skill.

    +

    Expertise Enables Adaptation

    +

    People newer to a skill or domain typically rely on documentation and more rigid and rule-based methods for performing a given task. They lack the context and experience to handle variations from standard expectations and outcomes, and tend not to experiment or improvise as much. Experts, on the other hand, demonstrate a unique ability to adapt to novel information, challenges, and surprise situations. Given the breadth and depth of their experience (and having learned by trying many various approaches and strategies over time), they are better equipped (and generally, more confident) to stretch and adjust when presented with novel situations. Think of this as the difference between being able to read sheet music well and being able to play jazz with a group you just met. 

    +

    Consider the following list of characteristics of expertise that Gary Klein and a number of other cognitive scientists have empirically demonstrated. They found that experts:

    +
      +
    • +

      Employ more effective strategies than others, and do so with less effort;

      +
    • +
    • +

      perceive meaning in patterns that others do not notice;

      +
    • +
    • +

      form rich mental models of situations to support sensemaking and anticipatory thinking;

      +
    • +
    • +

      have extensive and highly organized domain knowledge; and

      +
    • +
    • +

      are intrinsically motivated to work on hard problems that stretch their capabilities.

      +
    • +
    +

    If that sounds like some of the fundamental aspects of resilience, you’re on the right track. 

    +

    Expertise Is a Key Ingredient for Resilience

    +

    Given the above characteristics of expertise and how they develop, expertise can’t simply be “bottled” and subsequently scaled or reproduced en masse. As noted above, most high-performance domains are extremely variable, so one element of maintaining expertise is constantly updating mental models and factoring in new variables, challenges, and developments. Layer the nature of modern distributed software systems on top of this, and you begin to see that one cannot simply capture expert incident response and productize it.

    +

    Someone is inevitably going to say “Well we’ll just train our LLMs and agentic models on all the conditions necessary to generate expertise in incident response and then we can have that, right?” As they currently stand, LLMs and AI agents do not acquire expertise the way humans do, in part because they require human model development, training, and refinement. They also are not capable of sensemaking, reflection, coordination across fuzzy boundaries, reciprocity, vicarious learning from others’ experiences, observational learning (picking up things “in the air” from casual discussions), or “seeing the invisible” (perceiving missing vs. present cues)—these are all but a subset of the kinds of cognitive processes involved in developing and maintaining expertise that have been studied for decades by cognitive scientists. 

    +

    AI SRE agents can currently survey a given system’s environment and find inputs and patterns to generate hypotheses about how a given situation may have arisen. They can do this over and over and over, but they do not possess the knowledge or ability to, for example, call Sarah on the database team and see if she knows why things look weird. They can’t factor in the architectural discussions the team had about the recent migration which might explain why things aren’t behaving as expected. They can't notice that the system is behaving exactly like it did six months ago before a cascading failure, because that pattern lives in someone's memory, not in a log file. They can't pick up on the fact that the on-call engineer sounds unusually uncertain on the incident call, or that the team has gone quiet in a way that usually signals something is badly wrong. These aren’t edge cases, they are routine features of how complex incidents actually unfold.

    +

    The prevailing belief in the software industry appears to be that we can use automation and AI to replace what expertise has given us in the past. However, consider the data from the 2024 VOID report, wherein 75% of incidents involving automation required human intervention to comprehend, troubleshoot, and resolve the incident. That is where experts quite often save the day. Here we contend with the ouroboros of expertise and automation, in which experts don’t have as much experience with the system at hand, and then when asked to step in and resolve a problem with said system, they have found their expertise eroded by having less direct access to how it actually functions. 

    +

    A lack of experience with systems due to increased automation and AI can lead to de-skilling of experts, by depriving them of the continued exposure to the complexities of the systems they are expected to understand as experts. Remember: expertise is accumulated through repeated exposure to a wide variety of situations over time. As we add more automation and AI to these systems, it becomes even more difficult for people to build expertise with those systems, much less to be able to introspect how automation and AI are impacting the functioning of these systems. 

    +

    So what can organizations actually do to cultivate and protect the expertise that their resilience depends on?

    +

    Three Things Organizations Can Do to Foster Expertise

    +
      +
    1. +

      Develop and feed a culture where expertise is given the time and space it needs to develop and thrive. This means resisting the pressure to automate away the messy, variable, difficult work that builds expertise in the first place. It means recognizing that the engineer who has been in the weeds with a system for three years is an organizational asset, not just a headcount.

      +
    2. +
    3. +

      Invest in building the skills required to do effective incident analysis. Post-incident reviews done well are one of the most powerful tools organizations have for surfacing, sharing, and building expertise. Not the checkbox RCA that documents what went wrong and assigns blame, but the rich, narrative, learning-focused review that asks how the system actually behaved and how your experts made sense of it under pressure. This is where tacit knowledge gets made visible.

      +
    4. +
    5. +

      Help your experts identify and share the strategies and patterns they use. Incident analysis will help surface these, but the work doesn't stop there. Your organization needs specific approaches like knowledge elicitation, narrative storytelling, and communities of practice for distilling what your experts know into something the broader team can learn from. 

      +
    6. +
    +

    The vendors selling AI SRE and AI Ops solutions aren't necessarily wrong that these tools can help during incident response—they can survey environments, find patterns, and generate hypotheses faster than any human. But they are focused on a different problem than ensuring whether your systems are resilient. Resilience is continuously created by the people who have spent years inside your systems, who know what normal feels like, who remember what happened last time things looked like this. That expertise took time to build, and  it can't be purchased off the shelf. The good news is that it can be cultivated, incentivized, and shared. Expertise is contagious when organizations create the conditions for it to spread.

    +

    References

    +

    Peak: Secrets From the New Science of Expertise (Ericsson & Pool 2016) 

    +

    Seeing What Others Don't: The Remarkable Ways We Gain Insights (Gary Klein, 2015)

    +

    The Cambridge Handbook of Expertise and Expert Performance (Ericsson et al, 2006)

    +

    Why Expertise Matters: A Response to the Challenges (Klein et al., 2017)

    +

    The Ironies of Automation (Bainbridge, 1983)

    +

     

    +

    +

    Courtney Nash

    +

    Vice President, RISF

    +

     

    \ No newline at end of file diff --git a/sreweekly/articles/530/02-respecting-fatigue-isn-t-coddling.html b/sreweekly/articles/530/02-respecting-fatigue-isn-t-coddling.html new file mode 100644 index 00000000..2a4a4e6f --- /dev/null +++ b/sreweekly/articles/530/02-respecting-fatigue-isn-t-coddling.html @@ -0,0 +1,569 @@ + + + + + + + + + + Respecting fatigue isn’t coddling | Brent Chapman + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +
    + + + +
    + +
    +
    + +
    +
    +
    +
    +
    + + +
    + +

    Is it coddling when an on-call engineer takes the next morning off to recover after handling a production incident at 3 a.m., or is it a smart company managing a reliability risk?

    + + + +

    Here’s what that night actually looks like. The engineer gets paged at 3 a.m., then spends two hours diagnosing the problem, coordinating with fellow responders, and restoring service. By 5 a.m., the incident is resolved and they get back to bed, but it takes them a while to settle down and get back to sleep.

    + + + +

    Four hours later, they’re at standup. That afternoon, they’re in a planning meeting. That night, they’re still primary on the pager.

    + + + +

    This is the default at most companies. Nobody made a deliberate decision that it should work this way; it’s just what happens when there’s no explicit policy for post-incident recovery. And it carries more risk than most leaders realize.

    + + + +

    Incident response is more fatiguing than regular work

    + + + +

    Responding to an incident isn’t like a normal day of developing features and chasing bug reports. The cognitive demands are qualitatively different: rapid context-switching under time pressure, high-stakes decisions with incomplete information, coordinating across multiple people and systems, all while knowing that users are affected and stakeholders are watching. And there’s a physiological dimension that regular engineering work rarely triggers: adrenaline. Incident response activates the body’s stress response in a way that writing code or reviewing a design doesn’t. That heightened state feels productive in the moment, but it depletes reserves fast, and the crash afterward is steeper than the apparent effort would justify.

    + + + +

    This, incidentally, is one of the reasons that training and practice matter so much. Responders who’ve rehearsed the process and trust the framework around them experience a less intense stress response when real incidents hit. Turning incident response into a “routine emergency” doesn’t just improve efficiency; it reduces the physiological toll.

    + + + +

    An engineer who’s been actively responding for a few hours isn’t just tired in the way that a long day makes you tired. They’re measurably less effective at exactly the skills incident response demands: integrating new information, evaluating competing hypotheses, making decisions under ambiguity, and recognizing when a current approach isn’t working.

    + + + +

    The degradation is predictable

    + + + +

    Responder fatigue follows a recognizable pattern. As it sets in, people stop processing new information as effectively. They agonize over decisions they’d normally make quickly, or they stop making decisions altogether. They develop tunnel vision, fixating on the one theory they’re already pursuing instead of stepping back to consider alternatives. They fall into ruts, essentially pursuing “Plan A, again, with more feeling this time” instead of asking whether Plan A is still the right plan. They get less creative, more rigid, and more prone to mistakes.

    + + + +

    Fire departments study this, because it’s exactly the scenario firefighters face: interrupted sleep from overnight emergency calls, then back on duty the next day. The research consistently shows that the kind of fragmented, insufficient sleep they get around overnight calls degrades next-day cognitive performance to levels comparable to having had a couple of drinks. That’s why a growing number of fire departments have reconsidered their traditional 48-hour shifts; the performance degradation on day two is bad enough that departments are restructuring around it.

    + + + +

    Self-reporting isn’t enough

    + + + +

    The insidious part is that fatigue undermines exactly the capacity you need to recognize it. A fatigued responder genuinely believes they’re performing normally. That’s simply how fatigue works. “I’m fine, I can keep going” isn’t evidence of fitness. It’s one of the symptoms.

    + + + +

    This is the critical organizational point. If a company leaves fatigue management to individual judgment (“take it easy if you need to”), it’s built a system that depends on impaired people accurately assessing their own impairment.

    + + + +

    Aviation learned this the hard way. The FAA doesn’t ask pilots whether they feel too tired to fly. It sets hard limits on duty time and required rest periods, because decades of accident investigation proved that self-assessment under fatigue is unreliable. Pilots who’d been awake for 20 hours consistently reported feeling capable. The data said otherwise. If you’ve ever had a flight delayed while the airline sought a new crew because the original crew had “timed out,” you’ve seen these rules in action.

    + + + +

    The tech industry hasn’t had its equivalent reckoning yet, but the same cognitive science applies. An engineer who handled a two-hour incident at 3 a.m. and says they’re fine at 9 a.m. may well believe it. That doesn’t mean they’re right, and building your next day around that assumption is a gamble most companies don’t realize they’re taking.

    + + + +

    What active fatigue management looks like

    + + + +

    Companies that take responder fatigue seriously don’t rely on individual heroism or self-assessment. They build a few specific practices into their incident management capability.

    + + + +

    Explicit rest expectations. Not “take it easy if you need to,” but clear guidelines: an engineer who responds to a significant incident overnight is expected to start late or take the morning off, depending on duration and severity. The default is rest; working the next morning is the exception that requires a conscious choice, not the other way around.

    + + + +

    The incident commander (IC) monitors for fatigue. During extended incidents, it’s the IC’s responsibility to watch for fatigue signals in responders: slowed decision-making, tunnel vision, repeated questions, irritability, loss of situational awareness. This is the same responsibility a fire officer has for monitoring crew fatigue on a fireground. A fatigued responder who stays on the line isn’t being dedicated; they’re becoming a risk to the response, their teammates, and themselves. On incidents that stretch beyond a few hours, this includes planning responder reliefs early rather than waiting for someone to admit they’re spent.

    + + + +

    Promote the backup. After a significant overnight incident, consider moving the backup on-call engineer to primary for the next 12 to 24 hours. The person who spent two hours at 3 a.m. restoring service is not the person you want as your first line of defense if something else breaks that afternoon.

    + + + +

    Rethink shift length. Most teams seem to default to week-long on-call shifts with several weeks between shifts, but there’s a strong case for shorter, more frequent shifts. The same logic driving fire departments away from 48-hour shifts applies: shorter shifts mean less accumulated fatigue per shift, even if each person’s total on-call hours per quarter are similar. Shift design is a fatigue management decision, whether your company treats it as one or not.

    + + + +

    Some incident management platforms are starting to build this awareness into their tooling. incident.io, for example, detects overnight pages and proactively asks the responder the next day whether they’d like someone to cover their next shift. That’s the right instinct: making fatigue management a system-level concern rather than leaving it to the judgment of the person who’s least equipped to assess it.

    + + + +

    It’s a reliability decision

    + + + +

    Most companies aren’t actively choosing to ignore a fatigue problem. Rather, they have a fatigue problem that they haven’t noticed yet, because nobody has framed it as an operational risk. When a leader says “we trust our engineers to manage their own energy,” what they’re actually saying is: we have no organizational mechanism for ensuring that the people responding to our next incident are cognitively fit to do so.

    + + + +

    Respecting fatigue isn’t coddling. It’s protecting the quality of everything your engineers do the next day, including the next incident response.

    +
    + +
    + +
    + + +
    +
    +
    + + +
    + + + + +
    +
    + + +
    + + + + + + + + + + + + + + diff --git a/sreweekly/articles/530/03-on-building-scalable-control-planes.html b/sreweekly/articles/530/03-on-building-scalable-control-planes.html new file mode 100644 index 00000000..a9db6726 --- /dev/null +++ b/sreweekly/articles/530/03-on-building-scalable-control-planes.html @@ -0,0 +1,9 @@ +On building scalable control planes | All Things Distributed

    On building scalable control planes

    • 3756 words

    Header image

    Zak van der Merwe has spent his entire career at AWS building control planes. First for EC2 and now for DSQL. On the surface, the control plane looks quite boring: it records what should exist and reconciles that with what actually does. Nobody leaves school dreaming of building one, but Zak will be the first to tell you that if you like solving hard problems in distributed systems, there are few better places to be. It’s where many of those hard problems converge, and where the decisions you make determine whether a service survives its own growth.

    If you’ve been following Marc Brooker’s and Marc Bowes’s writing on DSQL, this is a great companion piece that pulls back the curtain and shows what it means to build a database that was designed from the start with control plane engineers in mind.

    –W


    On building scalable control planes

    I’ve been working at AWS for nearly fourteen years, and for almost all of that time I’ve been building control planes. It’s not the kind of career anyone maps out for themselves. Nobody leaves university thinking “I want to spend the next decade making sure the bookkeeping layer of a cloud service stays up.” But here I am, and I think the reason I’m still here is that control planes turn out to be where many of the interesting problems live, even if it takes a while to see that clearly.

    Before Amazon, I worked at a telecoms company in Cape Town where we had maybe ten servers, all in a room in the back of the office, and every single one had a name. You’d SSH into them, you’d share them with your colleagues, and if something went wrong you could walk over and deal with it. That was my entire mental model of what it meant to run infrastructure. Servers were things you knew individually, took care of deliberately, and could reason about as a set because there were few enough to fit in your head.

    I mention this not because it’s an unusual background but because it was so common less than two decades ago, and I think that’s what makes it worth saying out loud. Maybe your version is a small Kubernetes cluster or a handful of RDS instances where you can visualize the whole thing, you can name the parts, and when something breaks you know which part broke. That feeling of knowing your infrastructure is comfortable, and it makes the next part of the story genuinely hard to describe, because what happened when I joined EC2 was that that feeling just evaporated.

    Honestly, when I started, I didn’t really understand how EC2 worked. I kept trying to map it back to what I knew. If I launch an instance and the underlying server dies, what happens? Does my VM somehow get teleported onto another host? How does the cloud create this illusion that hardware failures don’t matter? I couldn’t square any of it with what I knew about running software.

    My first job at EC2 was health-checking the fleet, pinging every server and trying to figure out if it was healthy or not, and what I found was the opposite of magic. Things were failing constantly. Hosts were going down, hardware misbehaving, disks dying. I had seen the underbelly of EC2 and it was chaotic. My mental model had gone from “servers are precious things you protect” to “everything is on fire all the time.”

    It took a while to shake that feeling, but what I would eventually come to realize was that these failures were tiny drops in an enormous ocean of things working fine. The system was just operating at a scale where failures were a constant, a statistical certainty rather than an emergency. And the thing that made it possible to run a service at that scale without a human responding to every failure, the thing keeping everything humming, was the control plane.

    One way or another, my years at AWS have been spent working on control planes. Every AWS service has one, and I like to think of them as our unsung heroes. The better they work, the less anyone notices them. They’re the reason you don’t have to name your servers, and the reason that when hardware fails, you as a customer never have to deal with it. I’ve gotten to build control planes for two major AWS services: EC2, and DSQL. They’re nearly a decade apart, yet the hard lessons from building one led directly to the design of the other, and that’s the story I want to tell today.

    What is a control plane anyway?

    At this point, I probably owe you a better explanation of what I mean by control plane and why I think they’re interesting. I’ll use EC2 as an example, because that’s where I learned most of what I know.

    The way I think about it is that every service has a data plane and a control plane. The data plane is the set of core capabilities, the raw computing power, the hardware, the networking. The control plane is the conduit between those capabilities and customers. It’s the thing that takes what exists physically in a data center and presents it to you in a format you can actually consume and get value from. Without the control plane, you’d be back to SSH-ing into named servers in a closet somewhere. With it, you can spin up a thousand machines with an API call and never think about where they live.

    EC2 architecture diagram from Cape Town
    (This is how we visualized EC2’s architecture in the Cape Town office. Lots of pen, paper and post-it notes.)

    EC2 involves thousands of engineers and more features than anyone can keep track of, and yet the control plane, conceptually… is pretty simple. Stripped down, EC2 lets you rent a virtual machine (VM) in the cloud, and the control plane’s job is to set up and tear down these VMs for you.

    I like the analogy of a thermostat, because it’s constantly measuring the temperature, it knows where things need to be, and it’s always nudging the system in the right direction. That’s what our control plane does. It’s a continuous loop, watching the state of the world, comparing it to what should be true, and correcting the difference. When you launch a VM, the control plane records that a VM should exist, finds a physical server in the right data center, sets up the image, configures networking, and launches it. Later, if that server disappears for any reason, the control plane notices and updates its records to reflect reality. It’s always reconciling what is with what should be.

    One thing the team talked about constantly, almost to the point where it became a mantra, was that no matter what happens to the control plane, VMs that are already running need to keep working. We call this static stability, and it sounds obvious because of course running VMs should keep running. But at scale, obvious things are the hardest to protect, because every new feature, every change, every dependency is a chance to accidentally violate that guarantee. Maintaining it is the difference between an outage where customers can’t launch new resources and an outage where everything stops. Both are bad, but the second is catastrophically worse. The fact that EC2 was statically stable gave me some comfort in my early days.

    The EC2 team has done a phenomenal job making bad days rare. But understanding what bad days look like shaped a lot of what I know about building control planes.

    Living inside the control plane

    To understand how bad days start, it helps to know how the control plane stores state. At the heart of EC2’s control plane there is a relational database. When customers call the RunInstances API to launch a VM, the most critical thing that happens is that the control plane writes a row into its database: customer X now has VM Y. That’s when the API can safely return.

    In reality, a single RunInstances request triggers hundreds or thousands of internal API calls between micro and macro-services. Many of these services have their own databases recording their own state. It’s hard to exaggerate how complex this has grown over the years, but at the very bottom of all that complexity, there is a MySQL database, and what’s in that database is supposed to match reality.

    The simplest way things went wrong was also the scariest. Sometimes the primary database server just died. Our solution was a hot standby, a backup server continuously replicating from the primary, ideally only milliseconds behind. When the primary failed, we’d cut over to the standby and it could limit the outage to seconds. The team earned that through years of operational practice, building tooling, writing runbooks, training on-call engineers to execute the switchover under pressure. But seconds of outage still meant pagers getting lit up at 3am and asking humans to make decisions with incomplete information. We kept asking ourselves whether the architecture could take humans out of that loop entirely.

    The slower, more chronic problem was making sure our MySQL database kept up with business growth. This is pretty frustrating when you think about it, because the data plane does all the heavy lifting, like downloading VM images, configuring networking, running workloads, while the database is just keeping track of what exists. Every instance we launched meant more inserts, more updates, and more reads against the database, and eventually the bookkeeper couldn’t keep up with the workers.

    So we introduced more servers replicating from the primary and used these as read replicas. Many of the EC2 APIs don’t make any changes, they just describe the state of your current resources (how many VMs do you have, and so on). We sent traffic for these read-only APIs to our new read replicas and this massively reduced the load on our primary database server. This is standard practice for any team trying to scale up a relational database. Incidentally, this fleet of read replicas is why the EC2 API is eventually consistent, and as Marc Brooker has written, this puts an unfortunate cognitive load on our customers. It’s something we wanted to do better with DSQL, which we’ll get to in a bit.

    Read replicas bought us time, but every write still funneled through a single primary server, and eventually we had to shard the database. The first phase of this was visible to customers as we split each AWS region into multiple availability zones (AZs), each with their own independent control plane and separate MySQL databases. This helped with both scaling and availability, since zones fail independently and the blast radius of any single failure shrinks. It also became a fundamental building block that allows AWS customers to build architectures resilient to the loss of a single AZ. The second phase was internal: we sharded each zone into what we call cells. Both of these projects took years of engineering time because they required changes across many services. Every place in the codebase that talks to the database has to know which shard to route to. Simple lookups by primary key are straightforward, but anything else, such as joins across data that doesn’t align with your sharding boundaries, gets much trickier. Even the simplest decisions have consequences at this level. Do you shard by account or by resource? Different services choose differently depending on their access patterns, and there’s no universally right answer.

    There is also a human cost to all of this that I don’t think we talk about enough. In those early years, we didn’t have the automation to handle a lot of what a modern control plane just takes care of. When a security vulnerability was discovered and the whole fleet needed to be patched, we didn’t have a system that could say “go update every host at a safe rate.” We would literally recruit the whole team, subdivide all the hosts, and assign shifts. Everyone in the Cape Town office would get a chunk. Go update every one of your hosts, report status. That’s what life looks like without a mature control plane, and it’s the kind of thing that doesn’t scale. You can patch a fleet of a few hundred hosts that way. You cannot patch a fleet of millions that way. The control plane is what eventually got humans out of that loop entirely.

    If you’ve lived through this progression, the scaling cliffs, the read replica tradeoffs, the sharding projects that always take longer than you think they will, you know it’s a long and painful road, and it’s one that every team building a successful service backed by a relational database eventually walks.

    Searching for Database Xanadu

    After a decade working on EC2, I formed some strong opinions on what my ideal database looks like. It scales with my business without heroics. It is highly available with no downtime for updates, and no servers to babysit. My ideal database lets me leverage the power of the relational data model to model my domain and write software more productively.

    As it turns out, in the early 2020s, a group of experienced engineers on the databases side of AWS were thinking about exactly how to build this type of database. These engineers were expats from services like EC2 and had felt the pain of operating relational databases firsthand. They were also looking at the lessons learned operating massive scale serverless databases like DynamoDB and dreaming up ways to apply them to relational databases.

    They wanted to do for databases what EC2 and really Lambda did to servers. If you operate a traditional database with a “head node” you are in the world of “servers with names” like I was before joining EC2. The ideal database would free you from thinking about “databases with names”. Instead, it would have a control plane that takes care of all of that for you so that you can just think about your database as a logical endpoint that’s always available while it scales up and down.

    Sometime around 2021, this project really started to pick up steam. We’d figured out an architecture which seemed to deliver on this promise of the ideal database. I got the opportunity to join the team and start building its control plane. This service would launch in GA as Amazon Aurora DSQL in 2025.

    Let’s quickly revisit the major pain points that EC2 went through and see how life is different on DSQL—especially for control plane builders.

    In DSQL, there isn’t one server running your database. DSQL spins up a Firecracker micro-VM per connection, which means every connection is its own small head node. If one fails, only that single connection is affected rather than your whole application. Nobody gets paged, no one has to decide to cut over. I don’t manage standbys anymore, because the architecture has removed humans from that painful loop entirely.

    Scaling reads was another problem we spent years on at EC2, adding replicas by hand and accepting eventual consistency as the cost. DSQL adds read replicas automatically, and in fact this is one of the primary jobs of the control plane that I helped build. If your application suddenly sees a spike in read traffic, DSQL handles it, and the reads are strongly consistent, always. After years of telling customers “try again in a moment,” this property still blows my mind. It fundamentally simplifies the architecture of any control plane built on DSQL, and it removes that cognitive tax from the developers using the APIs those control planes expose.

    And then there’s sharding, which was availability zones and cells at EC2 and took us years. When you build AWS control planes for major new services, you have to anticipate that sharding will become necessary, and experience has shown that it’s cheaper to do it from the start than to retrofit it later. This is an ugly dilemma, because you’re extending your time to market on a speculative future problem, and when delivery timelines get tight, I’ve seen many teams give up on sharding just to ship. DSQL removes that dilemma because it automatically partitions your workload and you don’t have to think about it. You can use all the Postgres goodies you’re used to, complex transactions, multi-table joins, secondary indexes, while knowing your database is going to scale with your needs. Many new AWS control planes over the last decade were built on DynamoDB for this same reason, but DSQL offers a world with fewer compromises. You get the scalability of DynamoDB with the relational programming model that developers actually prefer to work with.

    “Self-hosting”

    When it came time to choose a database for the DSQL control plane, we chose DSQL. A team that runs on its own product feels every rough edge before its customers do, but getting there meant taking on the same circular dependency we’d faced at EC2: a control plane can’t depend on the thing it controls.

    We’ve seen two significant benefits from the decision to “self-host”. As customers adopt DSQL, they are creating thousands of databases, and the control plane is continuously scaling their databases up and down based on usage, often very rapidly. All of this customer activity creates “bookkeeping” work for the DSQL control plane, and the amount of this work grows with DSQL adoption. Since the DSQL control plane runs on DSQL, our bookkeeping database scales up to keep up with this increase in demand with minimal work from the team.

    The other benefit is in how we deal with availability zone outages. DSQL was designed from the ground up to survive single zone failures, but just because a zone is down doesn’t mean that customer workloads stop scaling or that customers stop creating databases. In my EC2 days, zone failures were fire storms as control plane databases died and pagers went off. For the DSQL control plane, these unfortunate bad days are much less painful because the DSQL control plane’s database remains available which allows the control plane to keep doing its critical work that ensures customer databases keep chugging along.

    Taking off the rose-tinted glasses

    If you’re still with me, you’re probably thinking to yourself: “what’s the catch?”

    As a relatively new service, there are features that we just don’t support yet. Some of these are gaps that we’re actively filling. Others are more nuanced, and we want to take our time to make sure we build the right thing. A good example is foreign key constraints. Foreign key constraints are a classic database feature that can be very useful and aren’t fundamentally hard to implement. However, foreign keys can also be dangerous at scale. We want to get this right, and that takes time.

    One of the advantages of running Postgres on a single node is that it maintains the working set in memory, and cached reads are insanely fast. Real architectures are more complicated though. For example, a control plane using Postgres would run across multiple availability zones and put a connection multiplexing proxy in front of the database. These are necessary steps for availability and scale, but they increase latency. When you build on DSQL, you don’t need to manage these things yourself. You get good (though not quite single-node Postgres good) latency that remains consistent as your application scales. This is exactly what I want as a control plane builder. Yes, I want fast, but I care even more about predictable latency as my application scales.

    It’s also worth being honest about where things stand for control plane builders at AWS. Migrating something like EC2’s control plane onto DSQL would take years even if we started today, and that’s okay. The ten-odd years I spent on the EC2 control plane taught me that the work that matters most tends to measure its impact in years, not quarters.

    Looking around corners

    We’ve spent most of this post deep in database scaling and life support. It’s a familiar shape for a lot of engineering stories. The problems we faced at EC2, how to go faster without breaking things, how to spend more of our time on the things that matter to customers, how to coordinate across a team that grew from a handful of people to thousands, and how to keep the system reliable while the ground shifted underneath us, are the same problems every engineering organization runs into as it scales. They are close cousins of the problems that produced Amazon’s original distributed computing manifesto back in 1998, and my own focus narrowed over the years to a single version of them, which was how to let individual teams fully own a piece of EC2 and move fast on their most urgent problems without expensive coordination, all while the product still felt like one coherent thing to a customer.

    When I look at the broader industry today, I see echoes of that same pressure playing out at a scale I did not expect, because the arrival of agentic coding has driven the cost of writing software down to almost nothing, and that pushes the hard part of the work somewhere else. When code is cheap, the bottleneck moves to judgment, to figuring out what to build, how to ship it safely, and how to anticipate what your customers will need before they ask. That is the same shift a good control plane makes for the people who build on it, taking the invisible work of keeping infrastructure alive off their plate so they can spend their attention on their customers, only now it is happening to software development as a whole, and even a single-person team feels the need to scale out.

    I am not going to pretend I know what building software will look like a year from now, because we are in the middle of a remodel and the walls are still open. What I do know is that it is much easier to move fast when you are standing on a foundation that will not crack under you, and that the problems worth spending a career on have always been the ones that need your judgment rather than your ability to keep the bookkeeping layer from falling over. My hope is that DSQL gives the next generation of builders that foundation, and gives them back the time to go look around corners for their customers, which is the part I always wished we had more room for at EC2.

    And as Werner says: “Now, go build.”


    \ No newline at end of file diff --git a/sreweekly/articles/530/04-mario-saved-the-eu-but-broke-my-system.html b/sreweekly/articles/530/04-mario-saved-the-eu-but-broke-my-system.html new file mode 100644 index 00000000..3efee541 --- /dev/null +++ b/sreweekly/articles/530/04-mario-saved-the-eu-but-broke-my-system.html @@ -0,0 +1,198 @@ +Mario Saved the EU but Broke My System | Uptime Labs + + + + + +

    Mario Saved the EU but Broke My System

    Hamed Silatani
    |
    Tags:
    Blog
    Incident Management
    Outages
    Simple thumbnail - conference table & headline reading 'Eurozone crisis live: Mario Draghi vows to save the eurozone'
    IN THIS ARTICLE

    Ready to make incident response your competitive advantage?

    See how Uptime Labs builds provable, scalable incident response capability across your organisation.

    Fourteen years ago (almost to the day), I experienced an incident that I unlocked new insight into when I read  Gary Klein's Sources of Power. The following story illustrates why experience, not training, is what develops incident response expertise.

    July 26, 2012: The Setup

    I was working at a trading firm. Everyone in the office knew the Eurozone was in trouble and that an announcement was coming. We'd prepared meticulously: load-testing all our trading systems to 10x their normal capacity, running through every scenario we could imagine.

    Around 11:30 BST, Mario Draghi made his announcement at a global investment conference in London: the ECB would do ‘whatever it takes’ to save the Eurozone. It was a moment of relief for investors. The market started to move rapidly, and we could feel the pressure building on our systems.

    Then, 30 minutes later, everything broke.

    The First Halt

    Around noon, everyone on our vast open-plan engineering floor suddenly jumped up from their desks. The horror on people's faces was unmistakable. Even the Chief Operating Officer rushed up to the engineering floor on the 6th level, which, in the normal order of things, is never a good sign.

    The system had simply stopped & nothing was working. More specifically, load balancers connections had shot up, web servers had exhausted all their connection and middleware services stopped processing. There were no warning signs; no alerts from our expensive monitoring tools. It was utterly inexplicable. A couple of minutes later, it resolved on its own, and everything came back to normal.

    For me, there was a sigh of relief. But the senior engineers who'd been in the industry long enough to know better looked visibly shaken. They knew something: the worst type of incidents are the ones that resolve on their own. That means you don't know what you fixed, and it will happen again.

    The Pattern Emerges

    Sure enough, at 12:30, the exact same outage happened. This time, people came running back from lunch, panic rising. When it resolved again a few minutes later, nobody felt easy. We all knew: judging by the pattern, the next hit would come around 1pm.

    And that was the critical moment. One hour before the US market open at 2pm: the busiest, most consequential time of the trading day when American traders would come online and react to Draghi's announcement. So if the system failed then, the consequences would be catastrophic.

    The Geeky Joke That Saved Everything

    While everyone was frantically scratching their heads, trying to explain what was happening, one engineer made an offhand comment. He said it looked like a ‘major garbage collection’ - a Java thing related to memory management. A stop-the-world garbage collection event where the entire JVM pauses.

    He laughed. No one else did. But in that moment, something shifted. In the absence of any other lead, everyone in the room latched onto this intuition. It was a universal moment of singularity: we had something to pursue.

    Here's what I didn't understand then, but understood after reading Klein: that engineer didn't randomly guess. He saw a pattern. His experience with Java systems allowed him to recognise something that novices would have completely missed. This is the core of what experts see that the rest of us miss during incidents, in the words of Klein:

    ‘Intuition is when we use our experience, and the patterns we have learned, to rapidly size up situations and know how to respond without going through deliberate analysis.”

    Following the Thread

    We had over 1,000 services running across the platform. All the critical ones had garbage collection monitoring and alerting in place, and nothing was alerting. So we shifted focus to tier 2 and tier 3 services: the non-critical systems that might not have been monitored rigorously or might not even have GC logging enabled. We wouldn't have any way of knowing if something happened to them.

    At 1pm, the third outage hit. Now the war room was fully formed. The COO, the Head of Trading, compliance officers - people who rarely set foot on the engineering floor were all there, huddled together, discussing how to respond to clients, how to manage the incoming calls, how to prepare for what would happen at 2pm when every major client would be online watching their trades.

    Ten minutes later, an engineer buried in the tier 3 logs noticed something: a tier 3 service had been performing major garbage collections on a suspicious schedule. It was such a low-importance service that normally no one would have even looked at it. The only reason it got flagged was the time correlation with our outages. But critically, no one could explain how a garbage collection in that service could possibly affect the entire trading flow. We had no proof of causation. But we had no other leads, and we were out of time.

    Decision Under Extreme Pressure

    We quickly huddled and worked out our options. What could we do?

    • Add more memory?
    • Restart the server (rolling or full)?
    • Shut down the service completely?
    • Cut it from the load balancer?

    What fascinated me, and what I only understood years later reading Klein, was how the senior engineers evaluated these options. In less than a minute, they ran through each one mentally, simulating what would happen if we took that action i.e. "If we add more memory, the next garbage collection will just be longer. If we shut down the service, the messaging broker it consumes from will pile up with messages and get flooded. If we isolate it from the load balancer..."

    They settled on isolation. Cut it from the load balancer. It was reversible, surgical, and - as we later confirmed after the crisis - it was the only option that wouldn't have made things worse.

    The Final Wait

    By the time we organised ourselves to isolate the service, we hit another episode at 1:30pm. The system halted again. A few minutes later it recovered, but we were running out of time. The business was already preparing contingency plans: how to apologise to the market, how to think about compensation, how to inform clients if this happened at 2pm.

    Then, about fifteen minutes before 2pm, we finally managed to cut the service from the load balancer.

    We had 15 agonising minutes to wait and hope for the best. We couldn't do anything else. It felt impossible.

    Then 2pm arrived & nothing happened. The sense of relief across the floor was overwhelming.

    Why This Story Matters 14 Years Later

    I didn't fully understand why this story stuck with me until I read Gary Klein's Sources of Power. Suddenly, everything crystallised with new insight:

    Klein identifies 4 cognitive powers that emerge under the kind of pressure we experienced: extreme uncertainty, no clear clues and time pressure. These are the sources of power that separate experts from novices.

    Intuition: That engineer's joke about garbage collection wasn't a wild guess or a moment of whimsy. It was pattern recognition: the ability to rapidly size up a chaotic situation with minimal information. Experts see patterns that novices don't even know to look for.

    Simulation: In less than a minute, senior engineers mentally ran through each option one at a time, simulating outcomes. ‘If we do X, what else might happen downstream?’ This kind of scenario thinking was critical because we couldn't test anything or wait and see. We only had minutes. Novices would need to deliberate; experts could just simulate.

    Metaphor: Previous garbage collection incidents informed their reasoning. They remembered metaphors and examples where adding memory made things worse, which helped them discount those options quickly. Experience creates a library of patterns.

    Storytelling: This story stayed with me for 14 years. It's the kind of vivid incident that surfaces in memory under future pressure, making knowledge available not just to me but to anyone who hears it. This is how expertise spreads, and this is precisely why incident response training must evolve beyond classroom exercises.

    John Allspaw, Founder and Principal, Adaptive Capacity Labs 10mo ·  There are only two ways people learn from incidents:  Personal, first-hand experience Via the experience of others (i.e., vicarious learning)  We cannot influence how or when #1 happens. We CAN influence how and when #2 happens.  Vicarious learning is the only way effective learning from incidents can scale beyond one person.  Creating conditions where this happens can be difficult. However, it's possible as long as there's broad recognition in the organization that...  a. Effective post-incident analysis means building the richest understanding of the event for the broadest possible audience.  b. The quality of post-incident analysis needed for (a) requires skill and expertise, in the same way experienced software engineers can produce higher-quality code more efficiently compared to when they first started.  c. Most organizations do not have this expertise, but these skills can be learned and improved. (This is what we do.)  Incident are being prevented all the time...in many cases, over 99% of the time! This takes effort, skill, and expertise.  Enabling the broadest audience to learn something they didn't know before also takes effort, skill, and expertise.

    John Allspaw’s LinkedIn post advocates for systematic storytelling for vicarious learning - specifically, in the form of post-incident analysis.

    The Core Principle: Experience → Expertise

    Here's what I can articulate now, with Klein's help: the difference between an expert and a novice isn't innate brilliance. It's accumulated experience, whether real or simulated, that builds these 4 powers.

    Most incident responders can't wait decades for real incidents to teach them. That's where simulation comes in. When you give a team carefully designed simulated incident experiences, you're not just teaching them facts. You're building their intuition, their ability to mentally simulate options, their metaphorical library of past incidents and their capacity to tell stories that will resurface under pressure.

    The things you gather from each incident, such as the visceral details, the decisions made, the outcomes, these stay with you forever. They become the patterns your brain recognises instantly. That's where the power is. That's how we reduce recovery time: not through static learning, but through building expertise one incident at a time.

    ‍

    Hamed Silatani

    Hamed is the co-founder and CEO of Uptime Labs. He has 20 years of experience in engineering leadership, reliability engineering and IT operations. Having spent the majority of his career at the sharp end of incident response in financial services, he's looking to help all companies master the unexpected.

    Share this post
    Clear blue sky with scattered white clouds.

    Ready to make incident response your competitive advantage?

    — Chris Voss

    See how Uptime Labs builds provable, scalable incident response capability across your financial services organisation.

    + + + \ No newline at end of file diff --git a/sreweekly/articles/530/05-ai-and-sre.html b/sreweekly/articles/530/05-ai-and-sre.html new file mode 100644 index 00000000..35c3a710 --- /dev/null +++ b/sreweekly/articles/530/05-ai-and-sre.html @@ -0,0 +1,403 @@ + + + + + + + +AI and SRE | bill duncan's blog + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +
    +
    + + + + +
    + +
    +
    + +
    +
    +

    AI and SRE

    + + + +
    +

    AI and SRE: The Same Force Cuts Both Ways

    +

    I’ve spent most of my career in the “in-between” role — the one that sits between software engineering and operations, making sure the systems that carry the weight of the business stay up, stay fast, and stay honest about what they’re actually doing. Listen, follow the data, don’t do anything you can’t undo quickly. That’s been the job for a couple of decades.

    +

    In the last year, the job changed more than it did in the previous ten.

    +

    The pro: a week of work in an afternoon

    +

    I’m not talking about autocomplete. I’m talking about handing a genuinely open-ended problem — “why did this pipeline start drifting eleven months ago and nobody noticed,” “here are forty repos, tell me which ones are safe to migrate first and why” — to an AI agent and getting back, in an hour or two, the kind of analysis that used to take a week of careful, interruption-prone human attention.

    +

    I’ve used this to cut a major cloud cost center by more than 90 percent, to untangle a production bug that had been quietly rotting for the better part of a year across a dozen services, and to review an entire database-upgrade runbook for gaps before we touched a production system carrying terabytes of customer data. None of that work was “prompt a chatbot and copy the answer.” It was iterative, it was checked against real logs and real output, and it was faster than anything I could have done alone at any point in my career.

    +

    I’ve seen AI correctly identify issues with merge requests, pull requests. I’ve seen it correctly root cause incidents in minutes that would have a much longer time to diagnose collecting all the information manually. (I’ve also seen cases where it made the wrong diagnosis too.)

    +

    That’s not a marginal improvement. That’s a different order of magnitude. And it’s real — I’m not describing a demo, I’m describing what shipped.

    +

    The con: a week of work in an afternoon

    +

    Here’s the part our industry is not saying out loud enough: the thing that makes AI a superpower for an individual engineer is the exact same thing that makes fewer engineers necessary. If one person with the right judgment and the right tools can now do what used to take a small team, the small team is the thing that’s at risk — not the work.

    +

    We’re already living in that answer. Layoffs in this field are accelerating again, and I don’t think that’s a coincidence of macroeconomics alone. Some of it is straightforward: the leverage AI gives a skilled engineer is being converted directly into headcount reduction, not just output growth. That’s not a hypothetical for me. It’s not a hypothetical for a lot of people reading this.

    +

    I don’t think wringing our hands about it changes anything. I do think pretending it isn’t happening is worse than useless — it leaves people unprepared for a transition that’s already well underway.

    +

    So what’s actually different about the job now?

    +

    If the leverage is real in both directions, the question worth asking isn’t “will AI take SRE jobs” — some of that has already happened, and more of it will. The question is what part of the job doesn’t compress the same way.

    +

    A few things I keep coming back to:

    +

    Verification doesn’t get automated away — it gets more important.

    +

    AI output is confident and plausible whether or not it’s correct. The instinct that’s always separated a good SRE from a dangerous one — don’t assert a root cause ahead of the evidence, widen the time window before trusting a correlation, verify the actual numbers instead of estimating — matters more now, not less, because the volume of plausible-sounding output you have to check has gone up by an order of magnitude too.

    +

    The bottleneck moved from “can we generate an answer” to “can we trust this one,” and that second skill is still entirely human.

    +

    Judgment about what to build, and what not to automate, gets scarcer and more valuable.

    +

    Anyone can now generate a script. Knowing which problem is actually worth solving, what the blast radius is if it’s wrong, and when the “clever” fix is a trap — that’s still earned the hard way, through incidents you’ve lived through.

    +

    The work shifts from doing to reviewing, and that’s a harder skill to teach.

    +

    Reading someone else’s (or something else’s) work critically, catching the plausible-but-wrong answer, knowing which five lines of a five-hundred-line diff actually matter — that was always a senior skill. It’s now most of the job, for anyone using these tools seriously.

    +

    Communication and trust go up in value as raw output gets cheap.

    +

    When everyone can produce more, the differentiator becomes who people trust to have checked it, and who can explain a tradeoff clearly enough that a room full of stakeholders can make a fast, confident decision. That’s always been true in aviation — the checklist doesn’t fly the plane, the pilot who knows when to deviate from it does — and it’s becoming just as true here.

    +

    Where I land

    +

    I’m not going to pretend this is a comfortable transition, for me or for anyone else in this field right now. The pace of layoffs is real, and anyone telling you AI is purely additive to the job market isn’t looking at the same data I am.

    +

    But I also can’t unsee what’s now possible. A week of work in an afternoon isn’t a slogan for me, it’s what happened, repeatedly, on real production systems. The honest position is that both things are true at once: this is a genuine force multiplier, and it is a genuine threat to how many of us the industry needs. Pretending otherwise, in either direction, doesn’t serve anyone.

    +

    We can’t put the genie back in the bottle. There is a real cost; a human cost and the environmental costs.

    +

    What I’m doing about it is the same thing I’d tell anyone to do with any new tool that changes the shape of a system: don’t assume, follow the data, and don’t do anything you can’t undo quickly — including your assumptions about your own career!

    +
    + + +
    + + +
    + + +
    + + + + +
    + + +
    +
    + + + + + + + diff --git a/sreweekly/articles/530/06-certificate-expiry-is-still-taking-down-major-platforms.html b/sreweekly/articles/530/06-certificate-expiry-is-still-taking-down-major-platforms.html new file mode 100644 index 00000000..3a2f0e20 --- /dev/null +++ b/sreweekly/articles/530/06-certificate-expiry-is-still-taking-down-major-platforms.html @@ -0,0 +1,302 @@ + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + Certificate Expiry Is Still Taking Down Major Platforms | TokenTimer + + + + +
    + + + + + + + diff --git a/sreweekly/articles/530/07-how-stripe-uses-graph-search-and-state-machines-to-auto-remediate-a-gl.html b/sreweekly/articles/530/07-how-stripe-uses-graph-search-and-state-machines-to-auto-remediate-a-gl.html new file mode 100644 index 00000000..25a7d133 --- /dev/null +++ b/sreweekly/articles/530/07-how-stripe-uses-graph-search-and-state-machines-to-auto-remediate-a-gl.html @@ -0,0 +1 @@ +How Stripe uses graph search and state machines to auto-remediate a global database fleet | Stripe Dot Dev Blog

    How Stripe uses graph search and state machines to auto-remediate a global database fleet

    /Article

    About the authors

    /About the authors

    Pragya Mehta

    Pragya Mehta is an engineer on the Document Database Control Plane Platform team at Stripe

      Sai Samant

      Sai Samant is a technical writer at Stripe working across the engineering organization.

      /Related Articles
      [ Fig. 1 ]
      10x
      [ Fig. 2 ]
      10x
      \ No newline at end of file diff --git a/sreweekly/articles/530/08-we-turned-off-pub-sub-and-nobody-noticed.html b/sreweekly/articles/530/08-we-turned-off-pub-sub-and-nobody-noticed.html new file mode 100644 index 00000000..ed25c7ba --- /dev/null +++ b/sreweekly/articles/530/08-we-turned-off-pub-sub-and-nobody-noticed.html @@ -0,0 +1,32 @@ +We turned off Pub/Sub and nobody noticed | Blog | incident.io

      We turned off Pub/Sub and nobody noticed

      August 11, 2026 — 17 min read

      Like many modern software stacks, the incident.io platform is predominantly event-driven. For example, whenever you send us an alert, post a message to our agent on Slack, or update an entry in your Catalog - these are all events that then get enqueued on a message topic, meaning any of our downstream components that are interested in that event can subscribe and react asynchronously, such as sending a push notification or posting a reply to you in Slack.

      As the platform has grown, so has the number of messages flowing through our system, and thus our dependence on our messaging infrastructure. It’s become mission-critical. At the same time, we've also set stricter availability targets for ourselves, like the 99.99% SLA we now commit to for Enterprise customers of our On-call product.

      Until recently, we only used a single provider - Google Cloud Pub/Sub - as our messaging technology. Meaning any blip in Pub/Sub availability meant a blip in our own availability, which isn’t acceptable. So we recently set out on an adventure to make our messaging system more resilient to failure by introducing a secondary message broker to our stack, adding redundancy, and ultimately increasing the availability of our entire platform.

      The goal was to be able to turn off Pub/Sub with zero customer impact. It turned out to be quite the adventure, but last week, we successfully did exactly that.

      This is the story of that adventure.

      The message broker

      Event-driven systems have many benefits, like allowing us to decouple the rate at which we process messages from the rate of ingestion, or have many different components process the same original customer-initiated event, for example, having a user.created topic, and one system listens for events to send a welcome email and another that sets up their initial database state.

      At the heart of such a system is typically a “message broker”, which is responsible for receiving messages from the publishing components, storing them, and forwarding them to any interested subscribers. At incident.io, we’ve historically used Pub/Sub as our message broker of choice; it’s a well-built managed service with a good feature set, and has allowed us to scale with ease over the years.

      Pub/Sub is solid. Its published SLA is 99.95%, and in practice it has comfortably beaten that for us. The problem was that every event in our platform flowed through one broker, operated by one provider, with no way to route around it. Our message broker had become a single point of failure (SPOF), and this was at tension with our own 99.99% availability targets.

      And a SPOF is ultimately a question about accountability. When an escalation doesn't fire, "Sorry, Pub/Sub was down" is not an answer we ever want to give a customer. It's our SLA, and it's our job to meet it, whatever our dependencies are doing that day. We already have redundancy in the other layers of our infrastructure, so why should the message broker be treated any differently?

      The eventadapter

      When we say we have an event-driven architecture, we’re not exaggerating; we currently have ~800+ individual message topics, 1000+ unique subscriptions to those topics, and are processing ~240 million messages a day *(as of August, 2026).

      This large number of topics also means there are thousands of call sites in our codebase which interact with events, which can be quite daunting when you want to, say, replace the underlying technology you use for messaging 🫠.

      Fortunately, we were standing on the shoulders of giants, and the early engineers at incident were wise enough to build code-level abstractions over message publishing and subscribing, which we call the eventadapter. This is a package that exposes some simple but powerful interfaces like:

      // Publisher is the interface for publishing events.
      +type Publisher interface {
      +	Publish(ctx context.Context, ev Eventer, payload []byte) (string, error)
      +}
      +
      +// Subscriber is implemented by all subscribers.
      +type Subscriber interface {
      +	Subscribe(ctx context.Context, topicName string, handler SubscribeHandler, params SubscribeParams) func() error
      +}
      +
      +// SubscribeHandler is what consumers of the package implement to 
      +// handle a single event.
      +type SubscribeHandler[EV Eventer] func(
      +	ctx context.Context, ev *EV, eventMetadata EventMetadata,
      +) error
      +
      +// Eventer is the interface implemented by all events.
      +type Eventer interface {
      +	// Name is how we identify this type of event.
      +	Name() string
      +
      +	// A description of what the event means.
      +	Description() string
      +
      +	// Validate validates the fields of the event before publishing.
      +	Validate() error
      +
      +	// GetOrganisationID returns the organisation ID associated with 
      +	// the event, which we use to add to event telemetry.
      +	GetOrganisationID() string
      +}
      +

      The package exposes a couple of concrete implementations of these interfaces, like a eventadapter.InMemory for use in local development or tests, or eventadapter.PubSub for talking to Pub/Sub.

      Fun fact: we originally had to build the InMemory adapter as, in the early days of incident, running the app locally and opening so many parallel connections to Pub/Sub would crash the office wifi. 🙈

      We conditionally choose and construct which version of the adapter to use at runtime in func main() based on the environment, and pass that down as a dependency to our application components.

      Having the luxury of such an abstraction meant that, to introduce a new message broker, what we needed to do was build a new implementation of the eventadapter interface, swap it in at runtime, without any of the calling code owned by other engineers being aware. This allowed us to hide all of the complexity that comes from load balancing across two different brokers behind the abstraction.

      Choosing a second broker

      One trade-off that came with using the existing interface meant that we are also constrained to the semantics and behaviors of that interface - which, even though it's an abstraction, already had some leakiness from the underlying technology. I.e. we needed to choose something that at least had the same feature set and behaviors as Pub/Sub, as these were semantics that we had come to rely on and reason about in our application.

      Our requirements and process for choosing a second message broker are out of scope of this post. There is a broad landscape of brokers: open-source vs proprietary, managed vs unmanaged, streaming vs non-streaming, ephemeral vs persisted, etc. There is no silver bullet, so our main advice is to document your own requirements and use a decision matrix.

      We chose NATS mainly because it’s a CNCF-adopted project, which means we can have confidence in its future and openly read the source code, is Kubernetes-native (which is where we run our workloads), is a single binary (we’re looking at you, Kafka!), and it is written in Go (the rest of our stack is Go!). Additionally, we had some prior experience running it.

      Dynamic load balancing

      Another design principle was that, when operating at 99.99% of availability, failover between brokers can’t be manual; with ~4 min 23 secs of downtime budget a month, we don’t have time for someone to wake up at 4 am and switch brokers. This meant that, ideally, we had to use both brokers in an active-active setup continuously: messages get balanced across both, and any persistent error rate from one broker would automatically fail over to the other. So we needed to build a dynamic event load balancer! Let’s dig into how we did that.

      As discussed above, the first thing was to create a new concrete implementation of the eventadapter that we called eventadapter.LoadBalancer.

      On the publish side, we pick a broker by hashing the message.ID and rolling a weighted dice against a configurable split. The ID is a ULID (like all IDs in our system), hashed with Go's hash/fnv std-library (the Fowler–Noll–Vo hash function). Because every message hashes independently, a 50/50 split sends roughly half of all messages to each broker — the even distribution you want from a load balancer.

      💡 You may wonder why we sample by message ID here, instead of, say, our typical grouping key, which is organisation ID? Hashing by org ID would mean all of a customer’s events would get pinned to one broker until failover, whereas hashing per-message keeps load even by volume no matter how lopsided any one org is. The trade-off is that there is no per-org or per-operation transport consistency - which we’re happy to live with.

      In normal operation, we’ve chosen to have a 50/50 split across each broker; why invest all this work in a secondary broker if you only use it in an emergency to then find out it's broken? Importantly, the split is configurable without a deploy, in case we need to turn either broker off manually.

      Once we determine the preferred broker for a message, we attempt to publish the event, and if a publish attempt fails (maybe it timed out due to a short network blip), we fail over to the other. All attempts to a given broker also flow through a circuit breaker, so if many attempts in a short period of time start to fail as the broker is degraded, we short-circuit the publishes to that broker early and instantly fail over. Giving us the automated failover we need to reach our availability targets! No one gets woken up; publishes just gracefully start flowing to the other provider.

      Fairness weighted scheduling

      You could argue that the publish-side is a pretty standard load balancer; where it gets more interesting is the subscriber-side and how we handle processing concurrency.

      The incident.io system is a single mono-repo Go program that is then deployed to Kubernetes as a collection of different workload-type-based deployments, such as worker-oncall or worker-ai, this allows us to do things like horizontally scale the number of replicas that receive inbound HTTP alerts independently of, say, our AI-message processing. We’ve talked about this architecture in more detail before: Keep the monolith, but split the workloads.

      In the eventadapater interface, we also have similar controls over the number of concurrent message “handler” functions (or more accurately, goroutines) we run per-machine to process messages for a given topic, a setting called MaxHandlers. So that we can do things like: configure 10 handlers to process webhooks per-machine but only 3 handlers for a lower-priority background cron job.

      Therefore, we needed to think about how we could map the concept of concurrent handlers to the new dual-broker world. The rudimentary solution would have been to simply double the number of handlers, one group of handlers per broker. However, that presents a couple of issues:

      1. It would double the potential throughput of concurrent work on a machine and thus directly increase our resource usage (CPU/memory/network), and each handler needs DB connections to do its work, so we’d also have to increase the connection pool sizes and thus CPU pressure on our database.
      2. More importantly, we might not always need an even split of handler processes per-broker. I.e. if Pub/Sub were to degrade overnight, and the publish rate flipped to an 80/20 split, the majority of messages would flow through NATS, and we would want the majority of our handlers to be consuming from the NATS queue and not Pub/Sub.

      What we really needed was a dynamic scheduler that pulls messages fairly and prioritizes the broker which has more overall work.

      So, like all good computer scientists, we did some research into prior art in this space and took inspiration from some existing queuing-theory algorithms. The most cited paper in this area is the MaxWeight algorithm (Tassiulas and Ephremides (1992)), which can be summarized as: “select the queue with the largest backlog”.

      However, this didn’t align well with our setup, as we had no way to efficiently query each broker for the current queue depth on every pull. That led us to the delay-based variants of MaxWeight, which swap the weight variable from "how many messages are queued" to "how long has the head message been waiting", such as Oldest Cell First (OCF, McKeown, Mekkittikul, Anantharam and Walrand (1999)), and delay-based back-pressure (Ji, Joo and Shroff (2011)). These keep the same throughput guarantees, but only need one piece of information per-queue: the age of the message at the head of the queue.

      With our newfound queueing-theory knowledge, we set out to build our scheduler!

      Each broker (and its underlying Go client) is wrapped in a single-slot "Inbox” interface with just two methods: Peek() to peek at the metadata of the head message stored in the inbox slot, and Receive(), to pop the head message out of the inbox.

      We then have a scheduler goroutine that is responsible for the pool of message handler goroutines; the handlers are bounded by a weighted semaphore, which is sized using the MaxHandlers subscriber setting we discussed before.

      Whenever a slot frees up in the semaphore, the scheduler calls Peek() on both inboxes and dispatches whichever head message has the oldest publish time (a broker-assigned timestamp, so it's immune to clock skew between publishing pods), by calling Receive(), which returns the message and then refills the inbox slot from the broker over the network, in time for the next peek. (We’ve glossed over some detail here, such as each broker’s native client also buffers some messages in-memory, to be more efficient).

      Putting this all together gives us all the properties we desired from our scheduler:

      • Reusing the existing MaxHandlers concurrency gives us one shared concurrency budget across the subscriber for both brokers, and no change in throughput semantics or resource consumption. I.e. our engineers can continue to set MaxHandlers and don't have to worry about the fact that the consumption is occurring via two brokers.
      • Using “oldest message first” scheduling gives us a weighted fairness policy; if we fail over to one broker and its queue backs up, the scheduler will prioritize its inbox to fill the semaphore slots.

      Magic! ✨

      Chaos

      So, we had a working implementation of our event load balancer and were pretty pleased with ourselves. Over the course of the last month we’ve been carefully rolling out the load balancer, first in our staging environment, then production topic-by-topic, until all messages now flow through it - including our most critical On-call escalation events.

      However, as all good reliability engineers know, an unexercised code path or failover mechanism is as useful as not having one. How do we know this would actually protect us against failure in either broker? The only option was to chaos test in production!

      First, with all of our messages flowing 50/50 across both brokers, we intentionally deleted our NATS cluster in our production Kubernetes environment. We watched the messages flow over to Pub/Sub gracefully, without client-facing errors.

      Testing Pub/Sub was a more interesting challenge: we don’t manage it, and can’t really call up Google and ask them to break it intentionally. So we built an automated fault injection system into the load balancer. The fault injector allows us to simulate partial degradation or total outages on either provider, and we can now dynamically increase or decrease the fault tolerance via configuration.

      Equipped with our shiny new chaos tooling, last week we turned off Pub/Sub in production, and no one noticed. The first few publish attempts time out or error and are then automatically retried against NATS; after 30 seconds of failed publish attempts, the circuit breaker opens, and all subsequent publishes short-circuit - not a single message dropped. All of this happened without any of our internal engineers being paged or a single customer noticing. The system worked 🎉

      Importantly, our chaos tooling now allows us to run these failure scenarios in production, continuously, and with ease, just like any other Tuesday.

      We are already seeing the benefits

      We initially set out on this project to make our On-Call product more reliable and eliminate one of the only remaining single points of failure in our system. However, along the way we’ve improved one of our core system primitives and raised the availability bar for our entire platform.

      After a couple of weeks running the system in production, we are already starting to see the benefits:

      • With improved telemetry and observability into our publish failures, we are regularly observing timeouts due to network blips to Pub/Sub, which now instantly fail over. Messages that may have previously caused customer-facing errors are now simply retries.
      • We can more safely operate and maintain both providers independently, without risk of data loss. Need to upgrade the NATS cluster? Easy! Change the Pub/Sub client? Safe!

      We’d love to hear what you think about the design and, as always, if working on these types of reliability challenges interests you - we’re hiring!

      Picture of Patrick Hamann
      Patrick Hamann
      Product Engineer
      View more from Patrick
      Picture of Mike Fisher
      Mike Fisher
      Product Engineer
      View more from Mike

      See related articles

      View all

      So good, you’ll break things on purpose

      Ready for modern incident management? Book a call with one of our experts today.

      Signup image

      We’d love to talk to you about

      • All-in-one incident management
      • Our unmatched speed of deployment
      • Why we’re loved by users and easily adopted
      • How we work for the whole organization
      \ No newline at end of file diff --git a/sreweekly/articles/530/index.json b/sreweekly/articles/530/index.json new file mode 100644 index 00000000..5a5d1974 --- /dev/null +++ b/sreweekly/articles/530/index.json @@ -0,0 +1,50 @@ +[ + { + "idx": 1, + "url": "https://resilienceinsoftware.org/news/11560646", + "ok": true, + "error": null + }, + { + "idx": 2, + "url": "https://greatcircle.com/blog/2026/07/28/respecting-fatigue-isnt-coddling/", + "ok": true, + "error": null + }, + { + "idx": 3, + "url": "https://www.allthingsdistributed.com/2026/08/on-building-scalable-control-planes.html", + "ok": true, + "error": null + }, + { + "idx": 4, + "url": "https://www.uptimelabs.io/articles/hamed-2012-outage-reflections", + "ok": true, + "error": null + }, + { + "idx": 5, + "url": "https://billduncan.org/ai-and-sre/", + "ok": true, + "error": null + }, + { + "idx": 6, + "url": "https://tokentimer.ch/blog/tls-certificate-expiry-outages", + "ok": true, + "error": null + }, + { + "idx": 7, + "url": "https://stripe.dev/blog/how-stripe-uses-graph-search-and-state-machines-to-auto-remediate-a-global-database-fleet", + "ok": true, + "error": null + }, + { + "idx": 8, + "url": "https://incident.io/blog/we-turned-off-pub-sub-and-nobody-noticed", + "ok": true, + "error": null + } +] \ No newline at end of file diff --git a/sreweekly/articles/531/01-heroic-saves-are-near-misses.html b/sreweekly/articles/531/01-heroic-saves-are-near-misses.html new file mode 100644 index 00000000..6f40556f --- /dev/null +++ b/sreweekly/articles/531/01-heroic-saves-are-near-misses.html @@ -0,0 +1,565 @@ + + + + + + + + + + Heroic saves are near misses | Brent Chapman + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +
      + + + +
      + +
      +
      + +
      +
      +
      +
      +
      + + +
      + +

      It’s 3 a.m. in California, where most of the dev team are still snug in their beds. The auth system has started rejecting valid credentials. Early bird East Coast customers are already trying (and failing) to log in for the day, and thousands of users in Europe have already given up and gone elsewhere. In a couple of hours, the West Coast will be waking up too. A brilliant engineer swoops in and saves the day. She has legendary debugging skills and a deep understanding of the auth system, and she puts together a fix in forty minutes that would have taken anyone else hours to even diagnose. Later that morning, leadership is sending thank-you messages in the all-hands channel. Her VP awards her a small spot bonus, and her manager reminds her to include it in the next performance review cycle.

      + + + +

      What doesn’t usually happen is anyone asking: what if she hadn’t been there? Because that heroic save, for all the heartfelt celebration around it, was actually a near miss from a systemic point of view.

      + + + +

      Near misses look like successes

      + + + +

      In aviation and other safety-critical fields, it’s widely accepted that a near miss is an unparalleled opportunity to learn and deserves the same investigation as an actual failure. The reasoning is straightforward: a near miss reveals the same systemic vulnerabilities that a failure does. The only difference between a near miss and a disaster is that the outcome happened to be good this time, often because of luck, timing, or the presence of one specific person.

      + + + +

      A heroic incident response is a similar opportunity. The system nearly failed, and would have failed if that one engineer hadn’t been available or hadn’t known exactly what to do. Her skill, expertise, and dedication are worth appreciating. But her unavailability would have meant a much worse outcome, and that’s worth examining too. Too many companies celebrate the save and stop there.

      + + + +

      The incentive nobody designed

      + + + +

      When a company celebrates a heroic save without examining why the heroics were necessary, it sends a message. The message isn’t intentional, but it’s clear nonetheless: what gets valued is the dramatic rescue, not the boring preparedness work that would have made the rescue unnecessary.

      + + + +

      Over time, that message shapes behavior. The engineer who writes thorough runbook documentation, trains new team members on the auth system, and invests in monitoring improvements doesn’t get the same recognition as the one who swoops in at 3 a.m. and saves the day. Preparedness work is largely invisible in performance reviews. Heroic saves are memorable.

      + + + +

      The result is a perverse incentive loop. Heroics get rewarded, preparedness doesn’t, and the company remains dependent on heroic saves because nobody is investing in the alternative. This isn’t because anyone explicitly decided that preparedness doesn’t matter. It’s because the reward system is quietly rewarding the wrong thing, and nobody has noticed because the heroes keep delivering results. Until they don’t.

      + + + +

      In my experience, this is one of the most common patterns in companies that are struggling with incident management. They have talented, dedicated people who keep delivering heroic results, and because the results keep coming, nobody realizes there’s a growing structural problem underneath.

      + + + +

      The hero as single point of failure

      + + + +

      The incentive loop creates a second problem. The hero gradually becomes a bottleneck and a single point of failure. When that engineer is on vacation and the next auth system incident hits, the team might spend hours just figuring out what’s going wrong, let alone fixing it. When they eventually leave the company (as they likely will; heroes tend to burn out), the team discovers that critical knowledge walked out the door with them.

      + + + +

      I see this pattern regularly in my consulting work. In a company’s most serious incidents, it keeps turning to the same handful of heroic engineers. Those engineers are talented and committed, and their involvement has genuinely saved the company from significant damage. Everyone involved with incidents knows who they are, and breathes a sigh of relief when they join an incident channel. But the company has never seriously examined what its response capability looks like without them. The term that often comes up to describe these people is “indispensable,” which is really another way of saying that the company’s incident response capability depends on specific individuals’ availability.

      + + + +

      Why the problem stays hidden

      + + + +

      The most insidious aspect of this pattern is that it’s invisible to leadership for as long as the heroes keep delivering. Companies at the earliest stages of incident management maturity often don’t realize they’re at risk. Leadership sees consistently good outcomes and assumes the company has strong incident response, when what they actually have is strong individuals (and a certain amount of good luck).

      + + + +

      By the time the fragility surfaces, the gap between where the company thought it was and where it actually was can be startling.

      + + + +

      Heroic is a growth stage, not a compliment

      + + + +

      When I assess incident management capabilities for my consulting clients, one of the dimensions I evaluate is program maturity: where is this company on the growth path from ad hoc response to reliable organizational capability? The first stage on that path is called “Heroic.” It isn’t meant to be flattering. It means that incident response quality is a property of specific talented individuals rather than a property of the company. When those individuals are available, things go well. When they’re not, things go sideways.

      + + + +

      Every company starts here. The question is whether they invest in growing past it, converting individual capability into organizational capability. That transition is what the rest of the maturity model describes, and it’s the core of what effective incident management programs are designed to do.

      + + + +

      What to recognize instead

      + + + +

      None of this means companies should stop recognizing heroic contributions when they happen. When someone saves the day at 3 a.m., thank them. But also investigate why the heroics were necessary, and invest in the answers. That’s a form of recognition too: it says the save mattered enough to learn from.

      + + + +

      To move from “Heroic” to higher levels of organizational capability, you need to shift what gets sustained recognition. Recognize the work that makes heroic saves unnecessary: the runbooks, the training, the well-coordinated responses where nobody had to be heroic.

      + + + +

      If an engineer spent much of their quarter writing runbooks, training new responders, and coordinating incident responses, recognize that work: in performance reviews, in public acknowledgment from leadership, in awards and bonuses. If you don’t, you’re telling your organization that the only incident management work worth noticing is the dramatic save.

      + + + +

      The goal is to make effective incident response something the company can do reliably, regardless of who happens to be on call. Heroes are still welcome, and still admired. They just shouldn’t be required.

      + + + +

      +
      + +
      + +
      + + +
      +
      +
      + + +
      + + + + +
      +
      + + +
      + + + + + + + + + + + + + + diff --git a/sreweekly/articles/531/02-control-and-complexity-tension-in-systems-design.html b/sreweekly/articles/531/02-control-and-complexity-tension-in-systems-design.html new file mode 100644 index 00000000..2e280ef0 --- /dev/null +++ b/sreweekly/articles/531/02-control-and-complexity-tension-in-systems-design.html @@ -0,0 +1,256 @@ + + + + + + + Control and complexity: tension in systems design + + + + + + + + + + + + +
      +
      +

      “My bad opinions”

      + +
    • blog
    • +
    • notes
    • +
    • about
    • +
      +
      + +
      +
      + + 2026/08/10 + +

      Control and complexity: tension in systems design

      +
      + +

      The adoption of LLMs in software development has led countless organizations to rapidly change their practices and structures. Old methods are questioned, replaced, and repurposed as the economics around creating new code get shaken up. Because humans and LLMs aren’t interchangeable, the dynamics in play are also very different. Systems are systems, and so regardless of what is changing, there are known patterns on which we can draw to provide some guidance and warnings.

      + +

      Without taking a step back and looking at the mindset behind the design of the system in which you operate, you’re likely to get somewhat incoherent (as in “clashing” and “conflicting,” not as in “nonsensical”) measures and policies. And so in this post I want to discuss how we organize systems by contrasting two families of approaches.

      + +

      The first is about analytical decomposition that aims to maintain control over a system, and the other is based on a perspective of complex systems that resist analysis, which tend to focus on figuring out interactions and mechanisms to foster desirable emergent behaviour.

      + +

      Comparing these has always been useful to tease apart assumptions and important elements of system design, and it is still relevant now with new types of changes being proposed.

      + +

      The approaches

      + + +

      Analytical Decomposition and Control

      + + +

      At the core of classic science, engineering, and many forms of management, lies the idea that the whole can be understood from its parts. Decompose a complicated thing enough that you can get a thorough and detailed understanding of every component, and you should be able to know how the ensemble works. This approach, analytical decomposition (also sometimes described as “Cartesian-Newtonian”), has been trustworthy and reliable in countless parts of modern life.

      + +

      This ability to divide, analyze, and understand generally extends to understanding causality over time: each action has a reaction, each event has a material cause, and these can be traced and evaluated or tested objectively. It follows that we can turn this around: if we understand an object well enough, then we can predict what it will do when acted upon.

      + +

      This is foundational to building machines and processes with any sort of predictability and reliability. You can have a high-level goal and a lot of disjoint parts, break down the problem, assemble components that are well tested and within tolerances, and have a working solution. A corollary is that if every part in the machine plays its role well, then the machine itself ought to work well.

      + +

      This requires taming a messy, chaotic world, and controlling parameters such that variability can be bounded. Design with enough tolerances and redundancy, and things should work. If not, we can dive in, take it apart, understand what broke, fix it, and be better for it.

      + +

      This approach is everywhere, from signal processing and telecommunications, where lossy information transmission is detected and corrected through redundancy, up to industrial quality control, where statistical processes can be used to define the acceptable boundaries of production.

      + +

      It also exists at the human level: in human factors engineering, concepts such as working memory (how many things the typical operator can hold in mind) or ideal observers (a theoretical person who monitors instruments at an optimal frequency against which we define “complacency”) have been constructed for the purpose of making sure that systems in which people participate will keep them acting within desirable parameters.

      + +

      It’s also visible at organizational levels. +Bureaucratic processes and hierarchies aim to keep alignment top-down such that the whole ensemble works coherently. Mechanisms of discipline and legibility are in play to keep the organization’s evolution under control. At broader scales, organizations often try to control their environment, their market, or the legislative context in which they operate.

      + +

      Basically, by deciding how much of a mess is accepted on the inside of a process, we can define a clearer interface on the outside of it for others to interact with. This abstraction creates a simplified but effective way to group a complicated ensemble into a manageable unit.

      + +

      Software ends up representing a sort of ideal for this mindset: systems can be written in languages that ensure some level of hard-won determinism. Execution is ideally always the same, there is no wear and tear, what worked yesterday will work tomorrow, everywhere. Policy decisions defined far away from the sharp end can be deterministically enforced at all levels.

      + +

      This means systems can be built from components bottom-up, aligning with top-down intent, limiting variability that comes from either machining or human behaviour. The ideal is a highly predictable, controlled, competitive, and reactive system.

      + +

      Complexity and Emergence

      + + +

      The problem is that by definition, complex systems resist analytical decomposition.

      + +

      There are many competing descriptions of complex systems, some of which are behavioural and some of which are structural. They all boil down to something like “things are so interconnected and have so many states that they become either unrepresentable, unpredictable or uncontrollable.”

      + +

      Other key elements are that these systems are dynamic, heavily influenced by their own history, and are also open—they continually change and interact in ways that don’t respect clean boundaries. This creates a tension where many participants have distinct goals, perspectives, representations, and degrees of freedom. By the time you’re done analyzing the system, it’s already something else. Even observing the system changes it in important ways.

      + +

      Put another way, if you find yourself surprised by the system’s behaviour, by the time you’ve pinned down what happened, it’s already a different system and your policy changes will be lagging or contributing to more counterintuitive surprises. Complex systems are more influenced than controlled.

      + +

      This dynamism leads to strategies that encourage equally dynamic adjustments. Since you can’t make these predictable, interventions will often be small and iterative. Alternatively, if you can’t simplify the elements or interactions you’re trying to control, you can increase the variety of control behaviours in order to make ongoing adjustments better. This tends to mean “put a controller—human or otherwise—that has enough internal complexity to cancel out the complexity of the thing it controls.” This, in cybernetics speak, is an attempt at creating more adaptive and dynamic control mechanisms.

      + +

      Balance is attained not by keeping things static, but by keeping them in motion.

      + +

      The ideal system is self-aware and flexible such that it can endlessly adapt and sustain itself, despite ever-increasing challenges. It's unclear whether the ideal can be reached.

      + +

      How they compose (or fail to do so)

      + + +

      Systems generally evolve from a constrained definition of the problem and its potential solutions, something that is tractable and effective. As the scope and scale of operations grow more comprehensive, further interventions trying to steer the system provide diminishing returns, and they increasingly produce unintended effects. These are the effects of complex systems showing up as things become tangled.

      + +

      The coping mechanism I’ve seen the most often is one of doubling down by doing more analysis, more decomposition, and putting more effort into more flexible automation that covers more cases. This in turn changes the nature of success and failure, by creating sometimes less frequent but bigger incidents instead. This type of composition takes place by substituting what breaks when possible, or sometimes by pure accident. It’s rarely been an orderly process.

      + +

      More rarely seen mechanisms seek to find out how much of the analytical and control-centric approaches we can afford to give up, identifying what can’t change at all, and then expanding complexity-aware mechanisms outwards from there. This is far less comfortable because this sort of stance demands that you give up on the idea that you actually are in control—a very unpleasant state of affairs to broadcast for a business.

      + +

      There are in fact long-standing debates as to whether larger scale accidents can actually be avoided. For example, Jean-Christophe Le Coze offers the following categorization:

      +Diagram showing three theoretical explanations for the unpredictability of accidents: technology out of control (Ellul/Perrow), fallible human constructs (Kuhn, Turner, Weick, Vaughan), and self-organizing emergent systems (Ashby, Rasmussen, Snook, Hollnagel). +
        +
      1. A ‘deterministic’ thread, where the properties of the technological systems themselves (such as tight coupling and complexity) will eventually defeat efforts to prevent accidents.
      2. +
      3. An ‘epistemic’ branch that focuses on the idea that organizations will suffer from 'failures of foresight' where weak signals and indicators that accidents are incubating will not be seen or accepted by the structures of power, and worldviews will fail to match new challenges, leading to accidents.
      4. +
      5. A ‘self-organizing’ thread that considers systems as adaptive and therefore frames success and failures as consequences emerging from systems' self-organization, through an exploration of problem and solution spaces with their available resources.
      6. +
      +

      These differing views are not fully incompatible, and authors from one category will frequently borrow from others. +Each perspective will however come with a focal point, a thing that is seen as important and worthy of consideration: the structure of control, the historicity of the system, the dynamics of power structures, the adaptive and changing nature of systems, the limited perspectives of participants, concepts around culture, and so on.

      + +

      Many contributors to these debates, while stating that accidents are unpredictable or hard to avoid, nevertheless seek explanations that can support making them less likely. They look at the limitations of known approaches, and expand the boundaries of what we should consider, adding new perspectives that can reveal new insights.

      + +

      There’s a lot of existing literature across many disciplines to study and get a better grasp on what doesn’t work (and when), and what is contextually useful. The opposition of analytical decomposition for control and complexity for emergence I’m offering here is crude and lacks nuance, but that’s hopefully what makes it an acceptable tool to think about changing systems.

      + +

      Oversimplification is what we’re doing here, and knowing what kind of wrong we’re going for is useful. As George Box (1976) said: “Since all models are wrong [we] must be alert to what is importantly wrong.”

      + +

      Contrasting Approaches in Practice

      + + +

      In a bit of a caricatural manner, the following examples will show relatively stereotypical perspectives to topics relevant to software through both analytical decomposition (with a focus on control) and complexity (with a focus that deliberately limits itself to influence):

      + + + + + + + + + + + + + +
      TopicAnalytical Decomposition / ControlComplexity / Emergence

      Training and education

      Build a well-defined curriculum, best practices for teachers and trainers, and testing mechanisms to ensure predictable performance and uniformity across students.

      Create environments that foster exploration, experimentation, and information exchange; provide guidance and support.

      Safety

      Prevent undesirable behaviours that lead to failure. Hazards are to be contained or designed out, and deviations from procedures or best practices are seen as a risk.

      Foster positive behaviours that lead to success. Find how people bridge gaps in processes, work around obstacles, and recover from problems.

      Correctness

      The software does what the specification or API says it should. Tests pass, it is feature-complete, and operates within known boundaries.

      Users or customers are able to successfully accomplish their tasks; goals can shift based on their needs.

      Reliability

      Uptime is within acceptable range, and is verifiable through SLAs, SLOs, etc. Load testing and thorough verification can prevent outages.

      Nines don’t matter if customers aren’t happy. You also won’t know for sure if software works until you hit production. Plan for recovery and coping with surprise.

      Approach to incidents

      Runbooks define best practices. Protocols and processes are defined to investigate and triage problems as efficiently as possible. Build for clear information and rapid diagnostics. Investigate what broke so recurrence can be prevented.

      Surprises may require improvisation. Who knows what will happen; build capacity to deal with the unknown. Investigations must look into normal work to understand how the system works in the first place.

      Developing features

      Understanding the needs of users and the strengths and gaps in current offerings lets you identify what to build and how to build it.

      Experiments in the field with potential features that you iterate on is how you best find what features may prove useful.

      Standards and norms

      Written unambiguously based on verifiable processes and outcomes to make enforcement tractable, scalable, and clear.

      Written in a goal-oriented manner as to support and guide the people who execute the work and who need to adapt rules to their reality.

      + +

      For each category, the attitude taken can drive people to pick drastically different approaches and activities, some of which may or may never overlap—the drive to control costs and errors can hinder the effectiveness or desire to experiment, and beliefs about how complex systems work may oppose all sorts of measures that are typically used to demonstrate accountability.

      + +

      I say this table is caricatural because in the real world, lines are often not this clean-cut, nor this superficial. It is possible for a control-centric hierarchy to align managers on goals and delegate authority down to cope with system complexity, and for control to be emphasized based on who people in power trust, for example. Centralized control tends to be most effective on the analytical decomposition side, but there are also approaches that aren’t control-centric that benefit from it.

      + +

      In fact, many activities can be used in both approaches, and serve both for distinct people, or even at the same time for any given person:

      + + + + + + + + + + + +
      ActivityAnalytical Decomposition / ControlComplexity / Emergence

      Code Review

      Find bugs and flaws; track and assign accountability; ensure quality.

      Build awareness and provide a space for feedback within and across teams.

      SLO adoption

      Organizational tool to ensure all teams manage their reliability adequately.

      Prioritization tool whose value comes from having teams discuss and define what is an acceptable level of reliability.

      Refactoring

      Pay down technical debt, reduce complexity, improve maintainability and flexibility, normalize used patterns.

      Countering entropy, adapting a code base to changing contexts based on new information available or shifting requirements.

      Chaos Engineering

      Validating that expected failure cases are properly tolerated or recovered from

      Experimentation-driven exercise in which participants form theories about their system’s behaviour in failure scenarios and try to confirm or disconfirm them.

      Using a platform

      A shared platform can encourage good architectural patterns and prevent undesirable ones, while abstracting away complexity for teams that build on it.

      Platforms provide systems with means of commoditizing shared elements to benefit from economies of scale and specialization, and address organizational bottlenecks through self-serve access.

      + +

      Even if activities in this list can serve both analytical decomposition and complexity-aligned approaches, that doesn’t mean that they will.

      + +

      For example, code review approaches that are control-centric and aim to hammer out any deviation from established norms may be adversarial to the point of causing anxiety or hindering actual feedback. Some implementations may still be able to mix automation and the proper social norms to successfully support both purposes to varying degrees of success.

      + +

      My experience has been that for these activities, the underlying position taken truly matters if you want to understand how they play out, and how they sometimes fail to meet someone’s expectations. This underlying position will also matter when it comes to prioritizing one activity against others. If participants or stakeholders do not agree to the higher-level purpose and desired outcome, then there will be a gap in ways these activities are expected to be carried out and how they take place, and in the relative importance they will be given across the system.

      + +

      When someone wants to change, supplement, or remove some of these activities, it’s useful to wonder what’s the nature of the change and what’s the perspective it favours.

      + +

      Flipping across approaches

      + + +

      As a heuristic, when multiple lenses are available, we can either try to find the best one (for some arbitrary criteria), or use a complementary or intersecting approach that uses as many of them as possible. Picking a single lens can lead to seeking implementations that maximize one type of activity contextually—whether control or emergence—whereas a combined approach can seek to make sure chosen activities are able to serve multiple properties, as a sort of tradeoff.

      + +

      Sometimes, what you get is not what you intend. An organization that sets up activities for control may find itself relying on practitioners invisibly repurposing them for complexity-aligned contributions. Meanwhile, the organization’s decision-makers exercise less control than they believe, or misattribute benefits to their own acts. They can then lose what they had when altering control mechanisms and incidentally hindering the hidden adaptations.

      + +

      Conversely, if activities are set up for emergence but are instead done mechanically as if intended for control, they won’t provide the expected benefits and might look and feel like busy work: the organization then neither controls nor benefits from adaptive effects.

      + +

      For broad topics and categories such as reliability or correctness, there are often no clearly defined choices or principles that are written down and that you can use. Organizations however tend to have some general tools that line up on the control-to-emergence spectrum, usually around process design and enforcement mechanisms.

      + +

      If you’re faced with behaviour you dislike, let’s say people from other teams modifying sensitive code your team owns unannounced, you can take measures such as having discussions with them reasserting ownership, and mentioning the expected process. You could require a preliminary RFC document or ticket before any change request is submitted. You can rely on code ownership files to prevent any unexpected change from going further without your agreement. You can move that key code to repositories which other teams cannot access.

      + +

      All of these are relatively local and play on the direct surrounding structure to modify actions and prevent undesirable acts. These approaches may be tremendously effective with little effort, but can also inadvertently fail to make desirable behaviour likelier.

      + +

      Closer to emergence’s perspective, it may be more typical to figure out what drives other teams to send these changes unannounced. What are the constraints and pressures they see that makes their current behaviour reasonable to them? If everyone agrees the process is a good ideal state but it frequently gets ignored, what is perceived as more important than that? Only once this is understood should you then design an intervention. This type of questioning—often informed by patterns such as those highlighted previously by Le Coze—tends to have you pull on a thread that unravels through the whole organization. It can be time consuming and difficult to do without established trust, but it can reshape expectations, and as easily lead to major change as to minor interventions upstream.

      + +

      A combined approach would be one where a broad understanding of the situation is obtained by leaning on complexity-aware methods, and is then used to design simple but high-leverage checks and barriers such that minimal control yields high rewards. This relies on the complexity stance to look not just at the system’s structure and purposes, but at how its various components and participants interact. Once the interactions make more sense, then the analytical approach is hopefully more effective.

      + +

      A risk here is to find yourself with a system that either feels so intractable, resistant to complexity approaches, or inflexible to cross-cutting interventions that you’re back to purely local defences, except they are late, with more work needed to get to the same place.

      + +

      The question then is not which approach is better, but how do we know when the current approach reaches its limits and what do we do then?

      + +

      Pitfalls of uncritical system design

      + + +

      People change their systems all the time, with or without this knowledge. They’re often successful, but not always, or at least not in the ways they had planned. Knowing what to look for doesn’t mean you’ll get it right, but it increases your odds.

      + +

      This might be true in the current LLM-driven shakeups as well. Because the technology is new and design patterns aren’t crystallized yet, a lot of people experiment a bit haphazardly. Many of their ideas have interesting elements or aspects to them that are worth learning from, but glaring omissions from a systems perspective that will still need to be handled.

      + +

      It’s almost impossible not to find examples of wide sweeping changes proposed when reading tech opinion pieces, which I’ll avoid linking to here. But they include ideas such as:

      + +
        +
      • Replacing code reviews with various types of barriers (tests and automated checks), rarely questioning what emergent roles the practice may have nor how static barriers may qualitatively differ from more adaptive ones.
      • +
      • Splitting software work into high-level specs to be translated to code in a black box with external checks only, without offering explanations around how the specs may cover varying abstraction layers, how the external checks can remain tractable, or how information worth learning should cross these boundaries in each direction.
      • +
      • Asking for everyone to become a sort of manager-of-agents while keeping agents under tight control loops, without asking what you may lose (or at least cause as second-order effects) in this analogy by changing the delegation and control mechanisms wholesale.
      • +
      • Focusing on system-level observable outcomes and letting go of imposing the structure within, trusting that the system will self-organize itself adequately.
      • +
      +

      If you design a system with control in mind—the use of barriers (think of the Swiss cheese model), the presence of extensive testing, of processes and procedures guaranteeing best practices—then you should pay as much attention to the mechanisms that will be needed to figure out if control actually works. This means asking questions like:

      + +
        +
      • How do we know our observations remain relevant, and that we surface the right signals?
      • +
      • How can we know if our understanding of the system loses accuracy?
      • +
      • What important elements is our analysis leaving out or obscuring when trying to make things legible?
      • +
      • How much variability is tolerated, and are we suppressing necessary types of it?
      • +
      • Are the things we optimize for creating brittleness elsewhere?
      • +
      • Is our control real or illusory? How would we know if that changes?
      • +
      +

      Well-regulated systems compensate for disruptions in ways that hide or suppress the signals of accumulating problems, both at technical and cultural levels. These questions aim to figure out whether any thought is given to what hides such behaviours.

      + +

      When you design for emergence—think of self-organization, market-like mechanisms, or delegation of decisions to participants with local context—other questions come up:

      + +
        +
      • Are local parts of the system working at cross purposes?
      • +
      • Is goal alignment effective? What maintains coherence?
      • +
      • What capabilities or efficiencies are we sacrificing when giving up on legibility?
      • +
      • Can we afford to lose the efficiency of a control-centric system? When might we need it?
      • +
      • How do we differentiate adaptation from drift?
      • +
      • What preserves dissent and carries information from the edges of the system?
      • +
      +

      Since complexity-aware approaches tend to resist prescriptive stances, there are often risks of increased inertia or widespread misalignment. Emergent properties will be key to success and failure, but without some careful thinking and influence, things can take on a life of their own.

      + +

      Whenever someone pushes for a system design that focuses on analytical decomposition or control, ask how they know they’re doing what’s needed, and the mechanisms by which they adapt. Whenever someone pushes for a design that seems to promise self-regulation and endless flexibility, ask how they’ll maintain coherence and the conditions they rely on for good outcomes. +Whenever someone pushes to switch from one to the other, ask what depends on current behaviour and consider what the second-order impacts might be there.

      + +

      Tech companies often rush to reinvent themselves around the outsized promises of new technology. Integrating new technology into existing workflows generally demands transforming the workflows. These changes often aim at reducing variability and increasing control, but cross subsystem boundaries in ways that disrupt tangled interactions that were dynamically stable.

      + +

      Automation that makes things predictable necessarily removes elements of unpredictability that can be useful to adaptation and evolution. Likewise, trying to make a part of the system more adaptive may necessarily make it less predictable. Both have knock-on effects on the rest of the system.

      + +

      Where and how does the system migrate from one mode of operation to the other? Where is control necessary and where is it not? What do we choose to analyze and decompose and what do we treat like an ecosystem instead?

      + +

      If we don’t have an answer to these, we also don’t have a good answer to how our systems will avoid failure or meet success. Systems are systems. They will keep acting like systems, and failing like systems.

      + +
      +
      +
      + + +
      + + + diff --git a/sreweekly/articles/531/03-structured-logging-in-distributed-systems-what-most-teams-get-wrong-an.html b/sreweekly/articles/531/03-structured-logging-in-distributed-systems-what-most-teams-get-wrong-an.html new file mode 100644 index 00000000..7bb4956f --- /dev/null +++ b/sreweekly/articles/531/03-structured-logging-in-distributed-systems-what-most-teams-get-wrong-an.html @@ -0,0 +1,3041 @@ + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + Structured Logging in Distributed Systems: What Most Teams Get Wrong and How to Fix It + + + + + + + + + + + + +
      +
      +
      +
      + + + +
      + + +
      + + + + + + + + + + +
      +
      +
      + +
      +
      + +
      +
      +
      +

      Could your team report a vulnerability within 24 hours? Find out on September 23.

      +
      + + + +
      +
      +
      + +
      + +
      +
      +
      +
      +
      +
      +
      + + +
      +
      + + + + + +
      +
      +
      + + + +
      +
      +

      Structured Logging in Distributed Systems: What Most Teams Get Wrong and How to Fix It

      +
      + +
      +

      Most teams log, but log badly: wrong severity levels, no trace IDs, inconsistent fields, and logs siloed from traces. Fix that, and incidents go from hours to minutes.

      +
      + +
      + By  + + + · + Analysis +
      +
      +
      + +
      + + +
      + + + Comment + + +
      + +
      +
      + Save +
      +
      + + + + + +
      +
      + + 3.1K Views +
      +
      +
      + + +
      + +
      + +
      +

      Logging is one of the oldest practices in software engineering, yet in distributed systems it remains one of the most poorly implemented. Most teams log, but very few log well. The gap between having logs and having useful logs becomes painfully visible the moment a production incident occurs at 2 AM across a system running dozens of microservices.

      +

      This article focuses on structured logging: what it is, where teams consistently go wrong with it, and the concrete practices that separate log data you can actually act on from log noise that burns engineering hours during incidents. If you are building or operating distributed systems today, structured logging is not optional. It is the foundation on which every other observability signal- traces, metrics, alerts- depends.

      +

      What Structured Logging Actually Means

      +

      Structured logging means emitting log entries as machine-readable key-value pairs rather than arbitrary free-text strings. Instead of this:

      +
      +
      +
      +
      + Plain Text +
        +
      +
      +
      [ERROR] 2026-07-10 03:14:22 - Failed to process payment for user 84729, reason: timeout
      +
      +
      +
      +


      +

      You emit this:

      +
      +
      +
      +
      + JSON +
        +
      +
      +
      {
      +
      +  "timestamp": "2026-07-10T03:14:22Z",
      +
      +  "level": "error",
      +
      +  "service": "payment-service",
      +
      +  "event": "payment_processing_failed",
      +
      +  "user_id": 84729,
      +
      +  "reason": "timeout",
      +
      +  "duration_ms": 3001,
      +
      +  "trace_id": "4bf92f3577b34da6a3ce929d0e0e4736",
      +
      +  "span_id": "00f067aa0ba902b7"
      +
      +}
      +
      +
      +
      +


      +

      The difference sounds cosmetic. It is not. The first format requires regex parsing and string matching to extract meaning. The second is immediately queryable, aggregatable, and, crucially, correlatable with traces and metrics from other services handling the same request.

      +

      The Five Mistakes Distributed Systems Teams Make With Logs

      +

      1. Logging Without Context Propagation

      +

      In a monolith, a single log line tells you where in the codebase an event occurred. In a distributed system, a log line without a correlation identifier tells you almost nothing. If Service A calls Service B which calls Service C, and Service C fails, you need a shared identifier, typically a trace ID, that threads through all three services' logs so you can reconstruct the full request journey.

      +

      The fix is context propagation: passing a trace ID through every request, injecting it into every log entry, and configuring your logging library to include it automatically. In practice, this means integrating your logging setup with OpenTelemetry or a similar tracing framework from day one, not as an afterthought. When your log entries include trace_id and span_id fields, you can jump from a log entry to its full distributed trace in a single query; that capability compresses incident diagnosis from hours to minutes.

      +

      2. Inconsistent Field Naming Across Services

      +

      In a microservices architecture developed by multiple teams, field-naming inconsistencies compound into a real problem at scale. One service logs user_id, another logs userId, a third logs uid. One service logs errors under error, another uses err, another uses exception. When you need to query across services during an incident, this inconsistency forces per-service query variations, slowing everything down.

      +

      Establish and enforce a logging schema across your organization. Define a canonical set of field names for common concepts, user identifiers, request identifiers, error fields, latency fields, and make that schema part of your service standards. Libraries like structlog in Python or logrus/zap in Go make it straightforward to enforce common fields at the logger initialization level, so teams can't easily deviate from the schema accidentally.

      +

      3. Logging at Wrong Severity Levels

      +

      Severity level misuse is endemic. INFO logs that should be DEBUG. Application errors logged as WARN because the developer did not want to trigger alerts. Business logic exceptions logged as ERROR when they are expected and handled. Over time, this degrades the signal value of severity levels to the point where teams stop filtering by level entirely.

      +

      Adopt and document clear severity semantics for your organization:

      +
        +
      • DEBUG: information useful only during active development; should not run in production
      • +
      • INFO: normal operational events (service started, request received, job completed)
      • +
      • WARN: unexpected conditions that are recoverable and do not require immediate action
      • +
      • ERROR: failures that require investigation; every ERROR should eventually be investigated or suppressed with documented justification
      • +
      • FATAL: unrecoverable failures; service cannot continue
      • +
      +

      Treat severity levels as a contract with your future on-call self.

      +

      4. Over-Logging Hot Paths

      +

      High-throughput services that log every incoming request at INFO level generate enormous log volumes that create three problems: storage costs escalate, log search performance degrades, and genuinely important events get buried in noise. A service processing 10,000 requests per second generates over 860 million log lines per day from request logging alone.

      +

      Use sampling for high-frequency, low-severity log events. Most observability platforms and log monitoring tools support log sampling natively; you configure a sampling rate for specific log patterns, keeping representative data without keeping everything. For example, sample 1% of successful payment processing logs but keep 100% of error logs. This dramatically reduces volume while preserving signal fidelity where it matters.

      +

      5. Treating Logs as a Standalone Signal

      +

      Logs become exponentially more powerful when they are correlated with traces and metrics. A spike in error logs is interesting. An error log spike correlated with a latency metric increase correlated with a trace showing a database connection timeout is actionable in seconds. Teams that treat logs as independent from their other observability signals are leaving significant diagnostic capability on the table.

      +

      If you are not already running OpenTelemetry, start there. It provides a unified SDK for instrumenting logs, traces, and metrics in a way that ensures they carry shared context identifiers. Once your logs carry the same trace IDs as your distributed traces, your observability signals become correlated by default, not by manual investigation.

      +

      A Practical Logging Schema to Start With

      +

      Here is a minimal structured logging schema that covers the majority of production use cases across distributed services:

      +
      +
      +
      +
      + JSON +
        +
      +
      +
      {
      +
      +  "timestamp": "ISO-8601 UTC",
      +
      +  "level": "debug|info|warn|error|fatal",
      +
      +  "service": "service-name",
      +
      +  "version": "1.4.2",
      +
      +  "environment": "production",
      +
      +  "event": "snake_case_event_name",
      +
      +  "message": "Human-readable description",
      +
      +  "trace_id": "OpenTelemetry trace ID",
      +
      +  "span_id": "OpenTelemetry span ID",
      +
      +  "user_id": "optional",
      +
      +  "request_id": "optional",
      +
      +  "duration_ms": "optional, numeric",
      +
      +  "error": {
      +
      +    "type": "TimeoutError",
      +
      +    "message": "Connection timed out after 3000ms",
      +
      +    "stack": "optional, omit in high-volume paths"
      +
      +  }
      +
      +}
      +
      +
      +
      +


      +

      This schema is opinionated but extensible. Services add domain-specific fields as needed while every entry maintains the common fields that make cross-service correlation possible.

      +

      Conclusion

      +

      Structured logging in distributed systems is not about logging more; it is about logging intentionally. The practices that separate teams who resolve incidents in minutes from teams who spend hours in log archaeology come down to four things: consistent field naming, trace context propagation, disciplined severity usage, and treating logs as a correlated signal rather than an isolated one.

      +

      Get these right, and your logs become a first-class observability asset during incidents. Get them wrong, and you have the worst of both worlds: high storage costs and low diagnostic value. The patterns outlined here are not theoretical; they are the difference between incident response that feels like detective work and incident response that feels like reading a timeline.

      +
      + +
      + + +
      +

      Opinions expressed by DZone contributors are their own.

      +
      +
      +
      +
      + + + +
      + +
      +
      +
      + +
      + + + +
      +
      +
      +
      +
      +
      +
      + +
      + × + +
      + + +
      +
      + +
      +
      +
      +
      +
      + + +
      + + + + + +
      + +
      + + + + + + + + + + + + + + + + + + + + + + diff --git a/sreweekly/articles/531/04-seeing-the-people-in-control.html b/sreweekly/articles/531/04-seeing-the-people-in-control.html new file mode 100644 index 00000000..7dffdf22 --- /dev/null +++ b/sreweekly/articles/531/04-seeing-the-people-in-control.html @@ -0,0 +1,1890 @@ + + + + + + + + +Seeing the People In Control – Humanistic Systems + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +
      + +
      +
      +
      + + +
      +
      +
      + + + + + + + +
      + +
      +
      + +
      +
      + + + +
      + +
      +
      +

      Seeing the People In Control

      + + +
      + +
      + +

      This article is a slightly edited reproduction of the Editorial published in HindSight magazine issue 36 (Autumn 2024) (all issues available at SKYbrary)

      + + + +
      Image: Steven Shorrock CC BY-NC-SA 2.0 https://flic.kr/p/2fMsKHp
      + + + +

      I joined the world of aviation in the late 1990s as a Human Factors analyst in UK air traffic management. I had just completed my master’s degree in work design and ergonomics, following my bachelor’s degree in applied psychology. For the first half of my career, my focus was mostly on micro interactions: breaking down tasks, procedures, and interactions at a granular level – seconds and minutes, button presses and radio transmissions. This work involved incident analysis, critical incident interviewing, human-machine interface evaluation, and simulation observation, all aimed at identifying episodes of what we might call ‘loss of control’. Breakdowns and breakages in countless human-human and human-machine loops preceded interactions that sometimes led to losses of separation, level busts and runway incursions.

      + + + +

      Looking back, I was primarily using applied cognitive psychology and cognitive ergonomics to understand control through loops of internal mental processes – perception, memory, attention, and decision-making – along with interactions, and feedback and from the environment. This is often depicted in diagrams with boxes and arrows illustrating the processing of information.

      + + + +

      In the second half of my career, my work shifted toward the macro level, zooming out to interactions within and between organisations, over months, years, and even decades. I listen carefully to people in various roles about their unique experiences. Here, the loops involve communication, cultures, and changes over time. These loops are inseparable and interdependent, creating formidable complexity in terms of people, technology, processes, structures, and organisations.

      + + + +

      No single frame of understanding suffices; I draw upon many disciplines, especially humanistic and social psychology, systems thinking, complexity science, and the humanities, in my attempts to understand the world. From this perspective, people seek to maintain control collectively through loops of communication and influence that evolve document them.

      + + + +

      Looking at the big picture, what is incredible is not that we sometimes lose control, but that we manage to maintain control at all. (Note that there are various meanings of ‘control’, from hard – making something happen – to soft – managing or influencing a process or situation – and it is worth thinking about what it means for you.) This brings me to a question that I often pose to groups, including senior managers: If you had to explain to a neighbour why your organisation is so safe, and generally works well, what would you say? The responses vary, but in the best-connected environments, different groups – controllers, engineers, managers, safety specialists – recognise and acknowledge each other’s contributions, forming large, interconnected loops. It’s a vital question to ponder, because if you don’t, how do you know what to nurture and extend…or defend in the face of cost cuts?

      + + + +
      +

      “Looking at the big picture, what is incredible is not that we sometimes lose control, but that we manage to maintain control at all.”

      +
      + + + +

      I recently posed this question to an audience of CEOs and safety directors at a EUROCONTROL conference in Spain. It was heartening to hear some senior leaders acknowledge in detail how people are their organisations’ greatest assets. They emphasised that people need to be in control and in the loop. I was surprised at the level of resonance with the theme of this issue of HindSight.

      + + + +

      The CEOs’ comments took my mind back to a groundbreaking report by Charles Billings, Human-Centered Aviation Automation: Principles and Guidelines, published in 1996 by NASA. Billings was a former flight surgeon and specialist in aviation medicine, who became an influential and distinguished NASA expert in aviation human factors. The principles in his report remain solid to this day, and the first three are so general that they apply regardless of the presence of automation.

      + + + +
        +
      1. The human operator must be in command.
      2. + + + +
      3. To command effectively, the human operator must be involved.
      4. + + + +
      5. To remain involved, the human operator must be appropriately informed.
      6. +
      + + + +

      The remaining principles focus on the relationship between human operators and automated systems:

      + + + +
        +
      1. The human operator must be informed about automated systems behaviour.
      2. + + + +
      3. Automated systems must be predictable.
      4. + + + +
      5. Automated systems must also monitor the human operators.
      6. + + + +
      7. Each agent in an intelligent human-machine system must have knowledge of the intent of the other agents.
      8. + + + +
      9. Functions should be automated only if there is a good reason for doing so.
      10. + + + +
      11. Automation should be designed to be simple to train, to learn, and to operate.
      12. +
      + + + +

      While these principles remain valid, they primarily address the operator-machine dynamic, or ‘joint cognitive system’. This was the focus of my interest in cognitive psychology and cognitive ergonomics. But the humanistic psychologist and systems thinker in me seeks principles that recognise people as more than operators, with control (or influence) distributed throughout organisations, industries, and societies. To this end, I propose the following nine principles to help ‘see’ the people in control:

      + + + +
        +
      1. People are whole and complex beings. We are greater than the sum of our mental, emotional, physical or behavioural ‘parts’, and cannot be fully understood by focusing on tasks, functions, roles, or occupations.
      2. + + + +
      3. People have unique virtues, values, gifts, and passions. For these to be expressed fully, we need a supportive and nurturing environment that values individuality, diversity, and inclusion.
      4. + + + +
      5. People have goals, and seek meaning, purpose, and creativity. We often seek these things through relationships, work, and personal pursuits.
      6. + + + +
      7. People naturally strive to learn, grow, and develop. We tend to flourish in a supportive and enabling environment.
      8. + + + +
      9. People are inherently social beings. We seek meaningful connections with others to find belonging, identity, support, and shared purpose, and are profoundly influenced by social norms, expectations, and pressures.
      10. + + + +
      11. People’s subjective experience is unique. Our experience shapes how we interpret and respond to the world around us and affects our wellbeing.
      12. + + + +
      13. People live in unique and dynamic contexts. These ever-changing contexts – personal, social, organisational, societal, political, environmental, technological, economic, and legal – strongly influence us.
      14. + + + +
      15. People are part of complex adaptive systems. Our interactions are influenced by a dynamic network of interactions, which are interconnected and interdependent, with outcomes that are often unpredictable.
      16. + + + +
      17. People have some choice, control, and responsibility. But agency is distributed among many and shaped by the opportunities and constraints of the contexts in which we exist, along with our capabilities and motivation.
      18. +
      + + + +

      These principles remind us that people are more than operators and need to be considered in the broader context. Although these principles have remained valid over millennia, the contexts and the complex adaptive systems in which we live and work (Principles 7 and 8) have changed dramatically, impacting our choices, control, and responsibilities (Principle 9). I encourage you to consider the principles in the light of any activity or change, inside or outside of an organisation.

      + + + +
      +

      “Things work because people make things work, bridging the gaps in the loops as they arise in order to stay in control.”

      +
      + + + +

      Over the last quarter of a century, one observation has become increasingly clear: everything is connected. In a complex industry like aviation, we can rarely discuss ‘local problems’ in isolation. Even the loss of a single individual – who may possess unique expertise – can significantly impact an organisation. This is equally true for the loss of critical resources. For instance, in our conversation in this issue of HindSight, Captain James Burnell discussed the effects of losing crew rooms at some airports. I revisited this impact through the lens of the nine principles I have just outlined. When I recently shared this story with another pilot from a different country, he was horrified at the prospect. “Crew rooms are sacred!”, he said, “There would be riots!” Crew rooms are shared resources that help crews to stay in the loop and maintain control and have even broader benefits for people.

      + + + +

      Going back to my “If you had to explain to a neighbour…” question, my answer is that things work because people make things work, bridging the gaps in the loops as they arise in order to stay in control. We do this using our remarkable expertise, creativity and connectivity, and do this sometimes to our personal cost. What is amazing is that things work as well as they do. It’s time that we fully acknowledged the reason for this – us – and respect people as so much more than operators and overseers of machines and processes.

      + + +
      + +
      + + + +

      Discover more from Humanistic Systems

      + + + +

      Subscribe to get the latest posts sent to your email.

      + + + +
      +
      +
      +
      +
      +

      + +

      +

      + + + + + + + + +

      +
      +
      +
      +
      + +
      + +
      +
      + + + +
      + + +
      +
      +
      + Unknown's avatar
      +
      + +
      +

      Author: Steven Shorrock

      + +

      + This blog is written by Dr Steven Shorrock. I work as an transdisciplinary humanistic-systems practitioner in safety critical industries. I blog in a personal capacity. Views expressed here are mine and not those of any affiliated organisation. + +Fellow of the British Psychological Society (FBPsS) | Chartered Psychologist (CPsychol) | Chartered Ergonomist and Human Factors Specialist (CErgHF) | BSc (Hons) MSc (Eng) PhD + +LinkedIn: www.linkedin.com/in/steveshorrock/ | Email: contact[at]humanisticsystems[dot]com +

      +
      +
      + + + + +
      + + +
      + +

      + One thought

      + + +
        +
      1. + +
      2. +
      + + +
      + + + +
      +

      Leave a Reply

      + + + + +
      +
      + + + + + +
      + + +
      +
      + + + +
      +
      + + +
      + + + +
      + +
      +
      + +
      + + +
      + +
      + × +
      + +
      +
      + +
      + + + + + +
      +
      + +
      + + +

      Discover more from Humanistic Systems

      + + + +

      Subscribe now to keep reading and get access to the full archive.

      + + +
      +
      +
      +
      +

      + +

      +

      + + + + + + + + +

      +
      +
      +
      +
      + + + +

      Continue reading

      + +
      +
      +
      +
      +
      +
      +
      +
      +

      + + + + + + + + +

      +
      +
      +
      +
      +
      +
      +
      +
      +
      +
      +
      + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + diff --git a/sreweekly/articles/531/05-type-conversion.html b/sreweekly/articles/531/05-type-conversion.html new file mode 100644 index 00000000..14fec354 --- /dev/null +++ b/sreweekly/articles/531/05-type-conversion.html @@ -0,0 +1,395 @@ + + + + + + + +Type Conversion | bill duncan's blog + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +
      +
      + + + + +
      + +
      +
      + +
      +
      +

      Type Conversion

      + + + +
      +

      Type Conversion: What Changes (and What Doesn’t) When You Start a New SRE Job

      +

      Pilots have a specific term for moving from one airplane to another: a “type conversion”. Not “learning to fly” again — you already know how to fly. It’s the process of taking everything you already know and re-mapping it onto a new machine that does the same job with a different cockpit.

      +

      Starting a new SRE job is the same thing. And thinking about it that way has changed how I plan the first few weeks in a new seat.

      +

      The principles don’t change

      +

      Aerodynamics doesn’t care what airplane you’re in. Lift, drag, thrust, weight — every single-engine piston airplane you’ll ever fly balances the same four forces the same way. A stall is a stall. A stabilized approach is a stabilized approach. The control surfaces that do the work — ailerons, elevator, rudder — are present on every one of them, even when they’re shaped differently or hung in a different place on the fuselage.

      +

      SRE is the same underneath. Error budgets, blast radius, the instinct to widen the time window before trusting a correlation, the discipline of not doing anything you can’t undo quickly — none of that is specific to a company. You bring all of it with you on day one. The job isn’t “relearn reliability engineering.” It’s “figure out where this particular airplane keeps its flap controls.”

      +

      The instruments are always there — the panel layout isn’t

      +

      Every airplane you strap into will show you airspeed, altitude, and engine RPM. That’s not optional; you cannot fly safely without them. What changes is where they are on the panel, and maybe whether you’re reading a needle or a glass display.

      +

      Every production system you’ll ever own is going to show you latency, error rate, and saturation, whatever the org calls its version of the USE or RED method. That information exists somewhere, because you can’t run anything at scale without it. What changes is whether it lives in Datadog or Prometheus/Grafana or a homegrown dashboard nobody’s updated the README for. Whether there is an alert on CPU saturation or queue depth, and whether the number you’re staring at is raw or something three layers of aggregation removed from the truth. The first week on a new system is instrument scan practice: find the gauges, confirm what they actually measure, and figure out which ones lie under load.

      +

      The controls you already know may work differently

      +

      This is the part that trips people up, because it’s where confidence and competence quietly come apart. You know how flaps work. You’ve used them hundreds of times. But if you learned on electric flaps — a switch, a preset, done — and the new airplane hands you a mechanical Johnson bar you have to feel your way through by position, “I know how flaps work” isn’t enough anymore. Retractable gear instead of fixed. A constant-speed prop with a blue lever to manage instead of one knob that just goes faster or slower. Same job, same physics, genuinely different procedure, and the procedure is exactly where accidents might happen.

      +

      New job, same story. You know how deploys work. You’ve shipped code for years. But if you learned on a fully automated ArgoCD pipeline and the new shop is still doing blessed-branch deploys by hand through a Jenkins job somebody’s afraid to touch, “I know how deploys work” is not the same as knowing how deploys work *here*. Kubernetes instead of a fleet of long-lived VMs. A change-management process with an approval board instead of merge-and-go. The underlying skill transfers. The muscle memory for the specific lever in front of you does not, and that’s the gap that gets people in trouble in the first month — not lack of skill, but skill applied on autopilot to the wrong control.

      +

      The numbers you have to memorize are airframe-specific

      +

      Every airplane has its own set of V-speeds — best glide, maneuvering speed, gear and flap extension limits — and you memorize them cold for *that* airplane. The number for best glide speed in a 172 will get you killed in a Bonanza. Knowing that V-speeds exist and knowing what they are for this specific one are two completely different kinds of knowledge, and only the second one is useful in an emergency.

      +

      Same with paging thresholds, SLO targets, and escalation policy. You already understand that error budgets exist and that burn-rate alerts are how you catch them early. That’s the concept, and it travels. The actual numbers — what latency triggers a page here, what counts as a SEV1, who gets called at 3 a.m. and in what order — are specific to this system, this traffic pattern, this org’s risk tolerance, and you have to learn them cold before you’re the one holding the controls during an incident. Nobody hands you a card with the new numbers on it before your first on-call shift, any more than they hand a pilot a V-speed card mid-flight. You go find it, in the runbooks and the postmortems and the people who’ve flown this one before you, before you need it.

      +

      How you actually check out on a new type

      +

      Pilots don’t skip the transition training just because they’re experienced. A thousand hours in one airplane buys you good instincts and bad habits in equal measure when you move to a different one, and the honest pilots know it. The process is always the same shape: study the systems (the POH, the equivalent of the wiki nobody’s updated since the last incident), fly with someone who already knows this airplane before you fly it alone, and treat the first hours as information-gathering.

      +

      I’m about to do this again — new company, new stack, new set of V-speeds to learn — and the plan is the same one I’d give anyone else walking into a new seat: don’t assume the panel is laid out the way your last one was, don’t trust a control until you’ve confirmed what it actually does *here*, and spend the first weeks flying the pattern with an instructor in the right seat before you take it up solo. The principles of flight got you your license. They will not, by themselves, tell you where this airplane keeps its flap controls.

      +

      Where I land

      +

      The comfort in all of this is that the hard part — the part that took years to build — is the part that transfers completely. You are not starting over. You’re doing a type conversion, not primary training. The four forces still balance the same way; the instruments still tell the truth if you know how to read them; and the checklist still doesn’t fly the plane — the pilot who knows when to deviate from it does. That pilot is still you. You’re just learning where the controls and instruments are.

      +

       

      +
      + + +
      + + +
      + + +
      + + + + +
      + + +
      +
      + + + + + + + diff --git a/sreweekly/articles/531/06-my-boss-wants-me-to-pick-an-ai-sre-tool-q-a-at-incident-fest-adaptive.html b/sreweekly/articles/531/06-my-boss-wants-me-to-pick-an-ai-sre-tool-q-a-at-incident-fest-adaptive.html new file mode 100644 index 00000000..d841a98b --- /dev/null +++ b/sreweekly/articles/531/06-my-boss-wants-me-to-pick-an-ai-sre-tool-q-a-at-incident-fest-adaptive.html @@ -0,0 +1,198 @@ +'My Boss Wants Me to Pick an AI SRE Tool': Q&A at Incident Fest (Adaptive Capacity Labs) | Uptime Labs + + + + + +

      'My Boss Wants Me to Pick an AI SRE Tool': Q&A at Incident Fest (Adaptive Capacity Labs)

      Sam Salter
      |
      August 11, 2026
      Tags:
      Blog
      AI & Automation
      Best Practices
      Incident Management
      IN THIS ARTICLE

      Ready to make incident response your competitive advantage?

      See how Uptime Labs builds provable, scalable incident response capability across your organisation.

      During Incident Fest 2026, our friends over at Adaptive Capacity Labs, John Allspaw & Beth Adele Long, answered questions about the relationship between AI & humans in our virtual ‘AMA Marquee’. Here are a selection of the questions & answers.

      Q: What are the safe and helpful use cases for AI in incident response based on where technology is today? Real-life examples would be very helpful.

      Beth Adele Long:

      In terms of safety, I’m a proponent of read-only access during incidents. I was going to say “unless it’s a relatively trivial / low-risk scenario,” but any write access that’s powerful enough to be useful is also likely to be dangerous. And incidents are already confusing enough without having to unwind a bizarre decision that was implemented at AI speed.

      With that safety caveat in mind, in long-running incidents, I can certainly see AI being helpful in the same the way it’s already being used during routine work: as a thought partner to help responders figure out what’s going on and explain current behavior. (J. Paul talked about exactly this in his talk, actually, and his point is really important. When AI predictions are bad, they degrade performance much more drastically than its good predictions improve performance.)

      I would love to see AI tools helping responders better with pattern-matching, but again, how that information is connected and then presented to responders very much matters. I’d also love to see AI supporting incident commanders by helping them make sense of the organization itself — who’s the right SME? Who do we page? Who has been working on the current incident long enough that they’re probably burned out, and I should send them on a humanity break? These are aspects we don’t think about enough but that really matter to effective incident response.

      Q: My boss wants me to pick an AI SRE tool to introduce into our org. Where do I start?

      Beth Adele Long:

      Oh boy, this is a tough one. I would start by getting clear about any contrasts between purported aims and real aims. By which I mean: how much is this a pragmatic request based on specific expectations (“As a leader, I believe AI SRE will help us do X, as measured by Y”) and how much is this actually motivated by something along the lines of “the board / my VP / someone in power says we need to be using AI more, and I need you to make me look good.” The more the latter factor is in play, the less room you’ll have to negotiate based on the actual benefit of the tools. You may just have to pick something and let it play out.

      In either case, I recommend looking at how much an AI SRE tool supports integration into everyday work. Are they getting lost in the leftover principle that Stu talked about, promising they’ll do work with no intervention? Or are they making life easier for your SREs? The latter claim is easy to test: do a pilot and see what your engineers say. Operational types are notoriously blunt and usually overloaded, so you’ll probably get a fast and honest take whether the tools are helpful or just annoying.

      Finally: good luck. This is a tough time to be evaluating tools that are still very much in flux and figuring out how to provide genuine value.

      Q: AI has saved a lot of time in the incident review process: sifting through loads of data, nicely constructing the timeline and extracting patterns. It saves a lot of time doing conversations and interviews. I wonder if other folks have seen such time savings.

      John Allspaw:

      The METR study in 2025 on developer productivity helped shine some light on how the perception of time savings/spent can be different than the actual amount of time savings/spent.

      (Before testing, developers guessed AI would make them 24% faster. After using it, they believed they were 20% faster. Turns out it was actually 19% slower.)

      I’m fascinated by this topic, so I have questions for you as well as others:

      When it sifts through data, what data does it dismiss as unimportant?

      Since all timelines are opinionated (because they’re constructed from raw data in ways that make sense to the author of said timeline), same question: how does the AI choose between events to include and events to dismiss?

      Q: When using AI in incident response, people frequently say that you have to second guess whether the AI is saying something sensible or not. Isn't that the same with humans?

      John Allspaw:

      Evaluating what your colleague has said while you’re both responding to an incident is clearly something happens, yep. Whether or not you’re “second guessing” what they’re asserting depends entirely on your experience with the person in the past, what they’ve said earlier in the response, how they described how they arrived at what they’re saying, etc.

      I’d guess that many people with experience responding to incidents can think of other ways that a human coworker’s contributions might be different than an AI agent’s contributions during an incident…?

      Beth Adele Long:

      To expand on John’s remark about “your experience with the person in the past”: with humans we have a deep intuitive sense of how to judge someone’s trustworthiness. The engineer who’s been at the company for 5 years and deeply knows this system gets weighted differently than a new hire; the person who’s “often wrong, never in doubt” gets more skepticism than the quiet person who only speaks up when they really know what’s going on. AI usually falls into the “often wrong, never in doubt” bucket! And my experience is that it’s a lot more expensive to evaluate AI’s wall of text in an incident than to guess the trustworthiness of a colleague’s terse assertion. Humans tend to be more efficient at building a shared context as a group (thanks, evolution). So yes, the second guessing is always happening at some level, but how that process unfolds is different in important ways.

      To see the full range of Q&As, you can explore Incident Fest here.

      Sam Salter

      Sam is the official editor at Uptime Labs, working closely with a global community of engineers and practitioners to surface real-world insights into incident response. Sam aims to help turn hard-won operational experience into clear, practical perspectives aiming to support how organisations prepare for, respond to and learn from incidents.

      Share this post
      Clear blue sky with scattered white clouds.

      Ready to make incident response your competitive advantage?

      — Chris Voss

      See how Uptime Labs builds provable, scalable incident response capability across your financial services organisation.

      + + + \ No newline at end of file diff --git a/sreweekly/articles/531/07-how-tailscale-helped-find-the-sqlite-wal-reset-bug.html b/sreweekly/articles/531/07-how-tailscale-helped-find-the-sqlite-wal-reset-bug.html new file mode 100644 index 00000000..ac73cad4 --- /dev/null +++ b/sreweekly/articles/531/07-how-tailscale-helped-find-the-sqlite-wal-reset-bug.html @@ -0,0 +1 @@ +How Tailscale helped find the SQLite WAL-Reset bug
      Blog|insightsAugust 12, 2026

      How we tracked down a 16-year-old SQLite bug

      Author

      Alex ChanAlex Chan
      Light orange and dark yellow shapes, like ovals, squares, circles, and quarter-circles, against a pale yellow background.

      At the end of last year, our uptime was pretty shaky. You can see this trend on our status page, and that instability continued into the new year. Many of these outages were caused by a single bug, deep in SQLite. It took months of intense forensics to track it down.

      Now we’re in summer, we’re confident that we’ve found the bug, that we understand it—and more importantly, that we’ve fixed it.

      We know our customers expect Tailscale to be a reliable service, and for several months we didn’t live up to that promise. That’s disruptive, and we’re sorry. We’re publishing this blog post to explain what went wrong, how we responded, and how we ultimately helped to uncover a long-standing bug in the heart of the SQLite database.

      Tailscale’s database architecture

      While our clients interact with our control plane as a single public endpoint (controlplane.tailscale.com), internally, our control plane is split into a series of coordination servers (or “shards”). Each tailnet lives on one internal shard at a time, but can migrate seamlessly from one to another. These shards are an internal implementation detail: you don’t know what shard your tailnet is on, and you never need to.

      Each shard has an SQLite database that holds all the information about the tailnets on that shard. A single Go process exclusively accesses that database, and serves the control plane for those tailnets. This single-writer design is exactly how SQLite is meant to be used.

      Architecture diagram illustrating how the Tailscale control plane is made of isolated shards, each of which has an individual SQLite database.

      We’ve used SQLite as our primary database since 2022, and we chose it because it's well-known, reliable, and widely used. SQLite is “boring technology”—in a good way. Many companies use SQLite in much larger deployments without issue, and we expected the same stress-free usage.

      In our current backup pipeline, we take a complete snapshot of the database every few minutes, then upload the entire SQLite file to an S3 bucket. We’d been running this setup without incident since early 2023.

      Fast forward to August last year, when a data pipeline that reads those S3 backups reported an error in one of our databases. We ran SQLite’s PRAGMA integrity_check command against the backup, and found it was indeed corrupted. SQLite corruption is possible, but it’s highly unusual and not something you should encounter in normal operation. We repaired the affected database, and investigated the cause, but to no avail.

      When operating at scale, even rare events can occur with some frequency, so we should have been unsurprised when it happened again—and again, and again, and again. In total, we faced 19 separate instances of database corruption over six months before we finally resolved the underlying bug.

      When you hear the phrase “database corruption”, it’s natural to worry about data loss. Because our control plane only handles configuration data, these databases contain metadata about your tailnet and devices, but never your private encryption keys or network traffic. In the earliest incidents, the recovery process meant a handful of newly added devices or configuration changes didn’t persist, and a small amount of metadata had to be re-entered.

      Architecture diagram illustrating how the Tailscale control plane copes with a database issue on a single shard. When the SQLite database has database corruption, it affects traffic on that shard, but other shards are unaffected.

      Whenever corruption occurred, we had to stop the control plane process on the shard while we repaired or restored the database. This was painful for tailnets on that shard, because their entire control plane disappeared during that recovery window. In the early incidents, that downtime was over an hour, but we gradually sped up the recovery process over subsequent incidents.

      Each tailnet is a mesh network, where devices make peer-to-peer WireGuard® connections to each other. When a device joins the tailnet, it has to get a list of other devices from the control plane before it can establish new connections—so if a device came online during the SQLite downtime, it couldn’t connect. While the database was being repaired, devices already online remained connected to each other, but they couldn’t learn about changes to the network. Those tailnets also temporarily lost access to the web-based admin console and the Tailscale API.

      There’s also a broader impact on trust. We post a global incident on our status page even when only a small number of tailnets are affected. Many people saw a status page event for an incident that didn’t affect them. Indeed, the majority of shards and tailnets were never involved in a database corruption incident! Nonetheless, repeated downtime erodes trust, whether or not you’re directly affected.

      From the very first instance of corruption, we knew this was a serious threat to our reliability, and we threw a lot of engineering time at the problem—but the fix wasn’t easy.

      Trying to find the fault

      This bug resisted all our initial attempts to find it.

      We looked at recent changes, but there weren’t any that seemed relevant. Nobody had been working on our low-level code that interacts with SQLite, because it had all been written years ago and presented no issues up until that point. We re-reviewed all of that code with a fine-toothed comb to look for previously missed bugs, but we didn’t find anything that would cause the corruption we were seeing.

      We looked for common factors between corruption incidents, but we couldn’t find any. It wasn’t tied to a single shard, or customer, or tailnet feature, or time of day, or load level. We were at a loss for what might be triggering the behaviour.

      This lack of reliable trigger conditions meant we couldn’t reproduce the bug synthetically. Instead, we had to rely on deploying passive, forensic telemetry in our live environment to catch the corruption red-handed. Gathering live diagnostics for a database issue is the last thing we wanted to do, but we had no choice.

      As an additional complication, the corruption didn’t occur on a regular schedule. Sometimes incidents would be hours apart, other times weeks. This made it difficult to predict progress or plan further work, because we were never sure when we’d get our next diagnostic dump. We had a six-week period between October and December when there were no corruption incidents, before they returned as an unwelcome Christmas present.

      Because this wouldn’t be a quick or easy fix, we reached out to the SQLite developers for a professional support contract. This was a great decision. It gave us direct access to their deep expertise and experience, and we had many detailed technical conversations about our architecture and our incidents.

      Between Tailscale engineering and the SQLite core developers, we mapped out several theories for what might be causing the corruption—including broken POSIX locks on close(), mismanaging memory owned by SQLite, or accidentally using SQLite from multiple threads while disabling thread safety. After every incident, we gathered more data, added more diagnostics, and systematically ruled out these theories. We were gradually converging on the true bug.

      The transactions that didn’t bark

      While we were investigating the root cause, we still had a live platform to run. We took aggressive steps to automate recovery and minimize downtime:

      • Configuring our control plane shards to hard-stop immediately upon encountering corruption
      • Deploying an automated backup monitor that continuously ran PRAGMA integrity_check over our backups
      • Improving our runbooks and on-call training

      These efforts cut our response time to under an hour—and then we discovered an unexpected clue.

      We wanted a way to restore service that didn’t involve rolling back to the last known-good backup (which would lose a lot of data) or repairing the known-corrupted database (which was potentially risky).

      To do this, we built a transaction logging pipeline. We streamed every SQL statement that modified the database to a separate log file. Because SQLite is a single-writer database with serialisable transactions, our transaction history was completely linear and deterministic. (This wouldn’t be true in a multi-writer database like Postgres or MySQL.) Replaying those transactions against the latest known-good backup should restore the database to its most recent state, safely bypassing the corruption.

      Diagram illustrating how the changes between two backups can be reconstructed by replaying the transactions that occurred between them.

      This pipeline worked, but then it did something even better: it gave us a clue.

      In two incidents, our transaction logs failed to replay cleanly. Upon closer inspection, we discovered that data written and committed by one transaction was inexplicably invisible to later transactions. A write had vanished into thin air without raising an error. That should be impossible!

      The writing on the WAL

      As these incidents were ongoing, the SQLite developers had been developing a new debugging tool. For a while, we’d suspected that the bug was somewhere in the checkpoint process. They were building a new tool to give better visibility into what was happening during checkpoints.

      To understand what this tool found, we need to briefly explain how SQLite checkpoints work.

      A SQLite database is made of a series of “pages”, tiny blocks of information. When you update the database, some of those pages need to be replaced with new pages with the updated information.

      For better performance and greater concurrency, we run SQLite with Write-Ahead Logging, which means new pages aren't written directly to the database file. Instead, they’re written to the "write-ahead log" or "WAL file".

      Architecture diagram illustrating the difference between the database file and the write-ahead log (also known as the “WAL file”). Both are made of individual “pages”, and new pages are written to the WAL file first.

      New pages can't be written to the WAL file indefinitely; at some point they have to be copied back to the main database file. This process is called “checkpointing”.

      Architecture diagram illustrating the SQLite checkpoint procedure. Pages in the WAL file are copied back into the database file. New pages can replace existing pages anywhere in the database file, or be appended to the end of the file.

      In most deployments, SQLite itself decides when to do a checkpoint, and the process is invisible to the end user and developer. In our control plane, we take manual control of the checkpoint process so we can run fast and consistent backups. This non-standard approach seemed suspicious as we steadily eliminated potential causes.

      One clue was that during corruption incidents, our metrics showed that SQLite would report copying more pages from the WAL file than were actually available. If there are 10 pages in the WAL file and 20 pages get copied to the database, something is clearly wrong.

      To understand what was happening during these faulty checkpoints, the SQLite developers created a new debugging tool for the virtual filesystem layer.

      SQLite is split into several layers. The top layer is the parser and code generator, which converts SQL statements into SQLite’s internal data structures. These data structures get passed to the pager, which splits them into the individual pages to be written to disk. Actually writing them to disk is handled by the OS interface, or “virtual filesystem”. Currently SQLite has two mainstream virtual filesystem implementations—Unix and Windows.

      Architecture diagram illustrating the internals of SQLite. It receives SQL statements as input, which pass through three layers: the parser/code generator, the pager, and the OS interface/virtual filesystem. The filesystem layer writes the changes to disk.

      If you're interested in a deeper dive on these internals, I recommend this lecture by Richard Hipp, the primary author of SQLite.

      This approach allows you to replace different layers with different implementations, or wrap an existing layer to get more information. To help diagnose our problem, the SQLite developers created a wrapper around the virtual filesystem that writes additional tracing information and logs about changes to the database. This wrapper is called the tmstmpvfs shim, and the source code is available in the SQLite public repository.

      Architecture diagram illustrating the internals of SQLite with our new debugging layer. The OS interface/virtual filesystem layer has now been wrapped in a tmstmpvfs shim.

      We deployed the shim into our live environment, and waited for the next corruption to occur. Fortunately, we didn't have to wait long.

      The WAL-Reset bug

      After our next corruption incident, the additional logs from the new tmstmpvfs shim allowed the SQLite developers to find and fix the bug: a rare data race in the SQLite source code between a checkpoint and a write transaction.

      In particular, if a write occurs at a specific time during a checkpoint, the checkpointing process gets confused—it thinks some of the pages have been copied from the WAL into the main database file, but they haven’t. Those pages never get written to the database file, and that data is permanently lost. The database file becomes corrupt, because other pages which reference those pages—such as an index—are written to the database.

      The SQLite developers named this the “WAL-Reset bug”, and they estimate it was present in SQLite for at least 16 years. It could exist that long because it was rare—so rare, the SQLite developers had to add code to deliberately trigger it in their testing environments. Their fix adds an additional check to the checkpointing function which detects when the WAL has been reset by another thread.

      They confirmed that this bug caused all of the baffling behaviour we’d seen. It explained the corruption, the transaction logs that wouldn’t apply cleanly, and the inconsistent checkpoint statistics. They also explained why we were more likely to hit the bug than other SQLite users: we take manual control of the checkpointing process, and we checkpoint very aggressively. Even a bug triggered by a rare condition was bound to hit us eventually.

      This was an exciting moment. After months of confusion and uncertainty, we finally had a plausible theory for why the corruption was occurring, and a fix we could deploy to prevent it.

      The SQLite developers released the fix as SQLite 3.52.0, and we prepared to deploy it as soon as it was available.

      Fixed, with a false alarm

      We rolled out SQLite 3.52.0 carefully—first to a few canary shards, then, when we saw it running smoothly, we deployed it to the rest of the control plane.

      Our backup monitor promptly turned red, and reported corruption in 13 different databases. This was extremely alarming, but we followed our recovery procedures to fix all the supposed corruption, and everything was happy. It turned out these databases had not suffered real corruption, but were subject to a second problem in the version of SQLite.

      We shared our errors with the SQLite developers, which uncovered a bug in SQLite related to stale expression indexes. If you create an index on a computed value, and then the computation changes, the index will contain mismatched values, which gets reported as corruption by PRAGMA integrity_check.

      In our case, we were storing some high-precision timestamps as text, converting them to a floating-point number in a VIRTUAL generated column, and the SQLite 3.52.0 release that fixed our data race also made an optimisation that subtly changed the rounding behaviour for text-to-floating-point conversions. Our canary shards didn’t have any timestamps that triggered the changed rounding behaviour, so we missed this in our phased rollout.

      Because this change caused false corruption warnings, the SQLite developers withdrew the 3.52.0 release and instead published 3.51.3, which only contained a fix for the WAL-Reset bug.

      We fixed the issue on our side by reducing the precision of our timestamps to integer seconds; text-to-integer conversions are unambiguous. Meanwhile, the SQLite developers created an automated, self-healing index feature in 3.53.0, which prevents the stale expression index problem.

      Party time!

      With the fix rolled out to our entire control plane, we were ready to declare victory, but we were still cautious. An absence of corruption incidents doesn’t mean things are fixed—we’d already had one six-week period of deceptive calm.

      We wanted positive proof that this data race was actively occurring in our production environment. Now that we understood the cause of the bug—a collision between a write transaction and a WAL-reset—we patched our SQLite driver to log a warning when these two operations overlap. If the warning fired but the database remained uncorrupted, we’d know the fix had saved us from a potential corruption incident.

      We deployed the warning, and we waited. And we waited. And waited. And waited. As weeks slipped by, we began to wonder why we didn’t see it. Was the warning broken? Was our theory wrong? Was the true bug still lurking in the darkness?

      Then, two months later, the alert we were waiting for finally fired:

      Alert Manager notification showing SQLitePartyMode warning: SQLite attempted corruption on shard2.corp.ts.net:8383 in party mode, but the system prevented it. Details include warning code, host, instance, job, namespace, severity, and shard information. Message advises checking server logs for corruption incident details.

      This alert proved that the precise conditions for the WAL-Reset bug do occur in our production environment, which means it was the likely culprit for our six months of shaky uptime.

      Since that weirdly joyous alert fired, we’ve run for another four months without any database incidents, as of this writing. Finally, we could breathe a sigh of relief.

      Off the well-trodden path

      Nobody wanted us to spend six months looking for bugs in SQLite. This was an immensely frustrating experience for both our customers and staff, and we’re all glad to put this instability behind us.

      This investigation is a useful reminder: running boring technology in a non-standard way is a risk. The common paths and standard configurations are incredibly well-tested and reliable. Most people use SQLite in a standard configuration and never face this sort of issue. Everything we were doing was a public, documented, supported configuration—but by taking manual control of the checkpointing process and running at our own aggressive pace, we stepped off the well-trodden operational path.

      Resolving these incidents was a massive, cross-functional effort involving dozens of people—including Tailscale's engineering and support teams, and the core maintainers of SQLite. It is to all of their credit that the impact of these incidents was not much worse.

      We know that repeated downtime erodes trust, no matter how many people are affected, and we’re grateful to our customers for their patience and support while we chased this down.

      Frustrating as this period was, we’re left in a stronger position than we were before. The long-standing bug in SQLite has been patched, and we fixed dozens of other incidental issues that we spotted while looking for it. We funded the open-source SQLite VFS shim that helped isolate the race condition almost immediately, and will help track down similar bugs in the future. Finally, we’ve refined our database backup and recovery processes, and live-tested them over a dozen times.

      Hopefully there won’t be another database incident like this—but if there is, we’ll be ready.

      Share
      Loading...

      Try Tailscale for free

      Schedule a demo
      Contact sales
      cta phone
      mercury
      instacrt
      Retool
      duolingo
      \ No newline at end of file diff --git a/sreweekly/articles/531/08-optimizing-kubernetes-pods-for-reliability-with-topology-spread-constr.html b/sreweekly/articles/531/08-optimizing-kubernetes-pods-for-reliability-with-topology-spread-constr.html new file mode 100644 index 00000000..63a99161 --- /dev/null +++ b/sreweekly/articles/531/08-optimizing-kubernetes-pods-for-reliability-with-topology-spread-constr.html @@ -0,0 +1,1738 @@ +Optimizing Kubernetes pods for reliability with topology spread constraints + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +

      If you’re like many Kubernetes users, you don’t pay much attention to where or how Kubernetes distributes your pods. As long as they’re running, it doesn’t matter where they get deployed, right? Surely Kubernetes will use some complex algorithm to figure out the most reliable way to distribute your pods across the cluster…right?

      Pod distribution plays a much bigger role in reliability than you might think. Fortunately, it’s easy to control when, where, and how Kubernetes distributes pods. By adding a few lines to your manifest, you can ensure your deployments are zone-redundant and evenly scalable. The feature is called topology spread constraints, and in this blog, we’ll explain how it works in full detail.

      ‍

      Why are topology spread constraints important for reliability?

      Topology spread constraints determine how Kubernetes distributes pods across failure domains, such as regions, zones, and nodes. This helps ensure your workloads are truly distributed not just across the cluster, but across your operating environment. You can set cluster-level constraints as a default, or set constraints for individual workloads.

      +
      Note
      +
      We recommend labelling your nodes with topology.kubernetes.io/region and topology.kubernetes.io/zone at minimum.
      +

      ‍

      How to configure topology spread constraints

      Topology spread constraints are defined using the field spec.topologySpreadConstraints. These can be applied to a pod or to the cluster. Constraints have the following fields:

      • maxSkew: the degree to which pods may be unevenly distributed. Its behavior depends on the value of whenUnsatisfiable:
        • If whenUnsatisfiable: DoNotSchedule, this determines the maximum difference between the minimum number of pods in the domain vs. the number of matching pods in the target topology. In other words, this is how far off the minimum a domain is allowed to get.
        • If whenUnsatisfiable: ScheduleAnyway, Kubernetes gives a higher precedence to topologies that would help reduce the skew.
      • minDomains: the minimum number of eligible domains (e.g. availability zones or regions).
      • topologyKey: the node label used to identify nodes used for this constraint. Any nodes that have this label are grouped into topology domains according to their values. For example, using topology.kubernetes.io/zone as a key creates domains based on the availability zones your hosts span.
      • whenUnsatisfiable: how to handle pods that don’t satisfy the spread constraint. By default, it won’t be scheduled (DoNotSchedule). Setting this to ScheduleAnyway schedules the pod regardless, prioritizing nodes that minimize the maxSkew.
      • labelSelector: the pod label used to find matching pods.
      • matchLabelKeys: a list of pod label keys to use to calculate the spreading skew.
      • nodeAffinityPolicy: determines how to treat each pod’s nodeAffinity and nodeSelector settings. Honor (the default) limits the topology calculation to these nodes, while Ignore uses all nodes. 
      • nodeTaintsPolicy: determines whether to include node taints in the topology calculation.

      Note that you can define only one topologySpreadConstraint for a given topologyKey and whenUnsatisfiable pair.

      ‍

      How to add a topology spread constraint to a Kubernetes manifest

      Imagine we have a Kubernetes cluster distributed across three availability zones: us-east-1a, us-east-1b, and us-west-2a. We also have a pod that we want to deploy and replicate for redundancy. We’ll start with the following manifest:

      +
      +
      YAML
      +
      +
      + +
      +
      +
      +
      +
      +
      +apiVersion: apps/v1
      +kind: Deployment
      +metadata:
      +  name: nginx-deployment
      +spec:
      +  selector:
      +    matchLabels:
      +      app: nginx
      +  replicas: 4
      +  template:
      +    metadata:
      +      labels:
      +        app: nginx
      +    spec:
      +      containers:
      +      - name: nginx
      +        image: nginx:1.31.3
      +        ports:
      +        - containerPort: 80
      +
      +
      +

      ‍

      If we deploy four replicas of the pod using a round-robin algorithm, we end up with one node with two pods and two nodes with one pod:

      ‍

      Three nodes spanning three availability zones: two in us-east-1 and one in us-west-2. Two pods are in us-east-1a, one in us-east-1b, and one in us-west-2a.

      ‍

      However, the Kubernetes scheduler might deploy two pods to two nodes, leaving one empty; or it might deploy three pods to us-east-1a and one to us-east-1b, which puts us at risk if the us-east region ever goes down. Or, in the worst case, it could deploy all four to one node and create a single point of failure.

      ‍

      Different potential distributions of the deployment across all three nodes, from even (2-1-1) to extremely skewed (4-0-0).

      ‍

      1. Let’s first limit the pod imbalance by setting maxSkew to 1. This ensures that no single node has more than one additional replica of the pod than any other node.
      2. Next, we’ll set the topologyKey to topology.kubernetes.io/zone, since we want to limit the spread by zone even if our zones span multiple regions.
      3. We want Kubernetes to run the pod even if it can’t satisfy our topology constraints, so we’ll set whenUnsatisfiable to ScheduleAnyway.
      4. We want to match all Nginx pods (in this deployment, anyway), so let’s add a labelSelector that matches app: nginx.

      Now, our manifest looks like this:

      +
      +
      YAML
      +
      +
      + +
      +
      +
      +
      +
      +
      +apiVersion: apps/v1
      +kind: Deployment
      +metadata:
      +  name: nginx-deployment
      +spec:
      +  selector:
      +    matchLabels:
      +      app: nginx
      +  replicas: 4
      +  template:
      +    metadata:
      +      labels:
      +        app: nginx
      +    spec:
      +      topologySpreadConstraints:
      +        - maxSkew: 1
      +          topologyKey: topology.kubernetes.io/zone
      +          whenUnsatisfiable: ScheduleAnyway
      +          labelSelector:
      +            matchLabels:
      +              app: nginx
      +      containers:
      +      - name: nginx
      +        image: nginx:1.31.3
      +        ports:
      +        - containerPort: 80
      +
      +
      +

      ‍

      How to find pods with missing topology spread constraints

      You can use the kubectl command-line tool to retrieve a list of pods, then use the jq command-line tool to filter pods that don’t have topologySpreadConstraints defined. For example:

      +
      +
      SHELL
      +
      +
      + +
      +
      +
      +
      +
      +
      +kubectl get pods -o json | jq -r '.items[] | select(.spec.topologySpreadConstraints == null) | .metadata.name'
      +
      +
      +

      ‍

      You can also use Gremlin’s built-in Detected Risks feature to automatically scan your Kubernetes pods for missing topology spread constraints.

      Once you’ve added your constraints, re-run this command to ensure your pods don’t appear in the output. If you’re using Gremlin, the “topology spread constraints absent” risk status will automatically change from “at-risk” to “mitigated” and your service’s reliability score will increase.

      ‍

      Combining topology spread constraints, node affinity rules, and taints and tolerations

      As we already saw, topology spread constraints can interact with other Kubernetes features, particularly node affinity rules and taints and tolerations. But there are subtle differences between these.

      Affinity rules define the specific criteria for scheduling a pod on a node. For example, a pod running a large language model (LLM) might have an affinity rule that requires a node with a GPU. This way, you can combine affinity rules and topology spread constraints to limit the domain of nodes available to a pod. Just make sure you set nodeAffinityPolicy: Honor (the default).

      Conversely, taints specify where not to schedule a pod unless it has a matching toleration. If a node’s GPU crashes due to a driver issue, you don’t want Kubernetes scheduling LLMs onto that node. Instead, you can apply a taint that prevents Kubernetes from scheduling pods on that node, while also migrating running pods onto new nodes. Like affinity rules, these work in tandem with topology spread constraints by limiting the size of the domain, as long as you set nodeTaintsPolicy: Honor.

      ‍

      Other Kubernetes risks to watch out for

      Topology spread constraints are just one piece of a resilient Kubernetes deployment. If you want to know how to protect yourself against other risks like missing liveness probes, unset resource requests, and improperly configured high-availability clusters, check out our comprehensive ebook, "Kubernetes Reliability at Scale."

      In the meantime, if you'd like a free report of your reliability risks in just a few minutes, you can sign up for a free 30-day Gremlin trial, or use Gremlin's Detected Risks feature to automatically scan your existing Kubernetes pods for missing topology spread constraints.

      ‍

      No items found.
      Start your free trial

      Gremlin's automated reliability platform empowers you to find and fix availability risks before they impact your users. Start finding hidden risks in your systems with a free 30 day trial.

      sTART YOUR TRIAL
      K8s Reliability at Scale

      To learn more about Kubernetes failure modes and how to prevent them at scale, download a copy of our comprehensive ebook

      Get the Ultimate Guide
      Andre Newman
      Andre Newman
      Sr. Reliability Specialist
      + + + + + + + + + + + + + + + + + + \ No newline at end of file diff --git a/sreweekly/articles/531/index.json b/sreweekly/articles/531/index.json new file mode 100644 index 00000000..a4f86780 --- /dev/null +++ b/sreweekly/articles/531/index.json @@ -0,0 +1,50 @@ +[ + { + "idx": 1, + "url": "https://greatcircle.com/blog/2026/07/21/heroic-saves-are-near-misses/", + "ok": true, + "error": null + }, + { + "idx": 2, + "url": "https://ferd.ca/control-and-complexity-tension-in-systems-design.html", + "ok": true, + "error": null + }, + { + "idx": 3, + "url": "https://dzone.com/articles/structured-logging-in-distributed-systems", + "ok": true, + "error": null + }, + { + "idx": 4, + "url": "https://humanisticsystems.com/2025/10/14/seeing-the-people-in-control/", + "ok": true, + "error": null + }, + { + "idx": 5, + "url": "https://billduncan.org/type-conversion/", + "ok": true, + "error": null + }, + { + "idx": 6, + "url": "https://www.uptimelabs.io/articles/incident-response-ai-qa", + "ok": true, + "error": null + }, + { + "idx": 7, + "url": "https://tailscale.com/blog/sqlite-wal-reset-bug", + "ok": true, + "error": null + }, + { + "idx": 8, + "url": "https://www.gremlin.com/blog/optimizing-kubernetes-pod-deployments-for-reliability-with-topology-spread-constraints", + "ok": true, + "error": null + } +] \ No newline at end of file diff --git a/sreweekly/articles/532/01-when-declaring-an-incident-becomes-everyone-s-favorite-workaround.html b/sreweekly/articles/532/01-when-declaring-an-incident-becomes-everyone-s-favorite-workaround.html new file mode 100644 index 00000000..f2207be1 --- /dev/null +++ b/sreweekly/articles/532/01-when-declaring-an-incident-becomes-everyone-s-favorite-workaround.html @@ -0,0 +1,545 @@ + + + + + + + + + + When declaring an incident becomes everyone’s favorite workaround | Brent Chapman + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +
      + + + +
      + +
      +
      + +
      +
      +
      +
      +
      + + +
      + +

      You see someone declare a Sev-2 and you wonder: wait, why is that even an incident? Nothing is down. Customers aren’t affected. But a manager needed to get their team’s problem to the top of another team’s priority queue, and the incident process was a reliable way to make it happen. That’s not really what the incident process is for, but it worked, so where’s the harm?

      + + + +

      The problem is, once folks see that this works, it starts happening more often. A product manager declares an incident because the incident notification is the fastest way to get leadership attention on a problem that’s been stuck in the backlog for weeks. An account team declares one because they need engineering support for a big demo to a major prospect and the incident process is the easiest way to pull engineers out of their sprint work on short notice. An engineer declares one because it’s easier than navigating the formal exception process for the deployment freeze.

      + + + +

      The harm is cumulative. When a growing fraction of your declared “incidents” aren’t real emergencies, the urgency signal degrades. When a genuine Sev-1 arrives, people respond with less urgency because they’ve been conditioned to expect another workaround. And the incentives compound: folks who game the system get their problems solved faster, which teaches everyone else that gaming is how to get things done. Each individual declaration is an understandable decision by someone who needs to get something done; it’s the aggregate that corrodes the process.

      + + + +

      Every one of these non-emergency declarations still carries the full overhead of a real incident. Responders get pulled off their planned work. Someone drops whatever else they were doing to serve as incident commander. Stakeholders context-switch to follow along. When you’re running enough of these, your teams are spending a meaningful fraction of their time in emergency mode for things that aren’t really emergencies, and all the indirect costs of incidents (disrupted projects, context-switching, recovery time) accumulate just the same.

      + + + +

      There’s an irony here: people are reaching for the incident process because it works; they’ve seen that it reliably delivers coordination, prioritization, and urgency on demand.

      + + + +

      The instinctive response is wrong

      + + + +

      When companies notice this pattern, the instinctive response is often to tighten the declaration criteria. They add gatekeeping: maybe you need manager approval to declare an incident, or there’s a pre-declaration checklist you have to complete first, or someone reviews whether the declaration was “warranted” after the fact. The intent is reasonable. The net effect is corrosive.

      + + + +

      Gatekeeping incident declarations is counterproductive. Every speedbump you build also slows down real incidents. The person who hesitates to declare because they’re not sure the problem is “bad enough” is already a common failure mode in incident response. Adding a formal approval step or a post-hoc review of whether the declaration was justified makes that hesitation worse, not better.

      + + + +

      You also miss what the gaming is telling you: people reaching for the incident process are telling you that your normal processes are falling short. If you only crack down on the gaming, you suppress the symptom without learning anything from it, and the underlying problems persist.

      + + + +

      Fix the escape routes, not the escaping

      + + + +

      Instead, look at what side effects people are trying to trigger when they declare questionable incidents, and make those capabilities available through other means.

      + + + +

      If the easiest way to bypass the deployment freeze is to declare an incident, create a non-incident exception process for urgent changes. This doesn’t have to be complicated; a lightweight approval from a designated release manager, with a clear escalation path, covers most cases.

      + + + +

      If the easiest way to get your problem moved up another team’s priority queue is to declare an incident, create a prioritization escalation path that doesn’t require an incident. A cross-team triage meeting, an explicit expedite-request mechanism, or even a dedicated Slack channel that the right people actually monitor can absorb most of the pressure. The bar doesn’t have to be as high as “declare an emergency”; it just has to be lower than “wait six weeks for the next planning cycle.”

      + + + +

      If the easiest way to assemble a cross-functional team on short notice is through the incident process, create a lightweight coordination mechanism for non-incident situations. Some companies call these “swarms” or “tiger teams” or “coordination requests.” The name doesn’t matter; what matters is that people have a way to get the collaboration they need without borrowing the incident process to do it.

      + + + +

      Repeatedly gaming the incident process to get resource prioritization or cross-functional coordination isn’t a series of one-off workarounds; it’s a symptom of a systemic problem that needs a systemic response. Google’s SRE organization built formal Code Yellow and Code Red mechanisms for exactly this: structured ways to rally resources and elevate priority when a problem is serious enough to demand cross-functional attention, but isn’t an incident.

      + + + +

      The diagnostic question

      + + + +

      Look at your last dozen or so incidents and ask, for each one: was this declared because there was an emergency, or because the incident process was the easier path to something the team needed?

      + + + +

      You don’t need a formal audit. Just ask a few experienced incident commanders and on-call engineers; they already know which ones were real and which ones weren’t. Then talk to the folks who called for the questionable ones (in a blameless, fact-finding way, of course). They’ll tell you exactly what’s missing from the normal processes, if you’re willing to listen.

      + + + +

      People gaming the incident process is just a symptom. The underlying problem is usually that normal processes are too rigid, too slow, or too unresponsive, and the incident process is the path of least resistance. Fix the underlying problem and the gaming stops, because there’s nothing left to game around. Your incident urgency signal recovers, your teams stop burning emergency-mode cycles on non-emergencies, and when a real Sev-1 hits, people respond like it matters.

      + + + +

      And if you’re dealing with this, take a moment to appreciate what it says about your incident process: people are borrowing it because it works. The fix isn’t to make it stop working. It’s to make everything else work that well too.

      +
      + +
      + +
      + + +
      +
      +
      + + +
      + + + + +
      +
      + + +
      + + + + + + + + + + + + + + diff --git a/sreweekly/articles/532/03-solving-mysterious-kubernetes-pod-setup-timeouts-by-tuning-conntrack-g.html b/sreweekly/articles/532/03-solving-mysterious-kubernetes-pod-setup-timeouts-by-tuning-conntrack-g.html new file mode 100644 index 00000000..274c195b --- /dev/null +++ b/sreweekly/articles/532/03-solving-mysterious-kubernetes-pod-setup-timeouts-by-tuning-conntrack-g.html @@ -0,0 +1,179 @@ +Solving mysterious Kubernetes pod setup timeouts by tuning conntrack garbage collection - Adyen

      Article

      Inside Cilium CNI: solving mysterious Kubernetes pod setup timeouts

      A deep dive into how Adyen's Data Platform Engineering team investigated and resolved linear-time scaling bottlenecks in Cilium CNI connection tracking garbage collection to fix mysterious Kubernetes pod setup timeouts on high-resource nodes.

      Jorrick Sleijster  ·  Senior Data Platform Engineer, Adyen
      July 8th, 2026
       ·  15 minutes

      I was fully aware a year ago that a single configuration line could break the Kubernetes networking stack. But if they told me that leftovers from Kubernetes pods which terminated hours prior could block new ones from starting, I would have thought they were joking.

      In high-performance networking, 35 seconds is a lifetime. This was the latency required to iterate through our connection tracking table of 7 million entries at a maximum speed of 200,000 entries per second. At our 16-million-entry peak, this sequential lookup could take up to 80 seconds, leading to Cilium CNI timeouts preventing new pods from starting on affected nodes.

      We uncovered this linear-time behavior at Adyen by tracing syscalls, inspecting codebases, and analyzing eBPF internals. This investigation revealed how our varied workloads turned the connection tracking table's garbage collection algorithm into a critical bottleneck.

      Our setup: why we're different

      At Adyen, we run Cilium CNI across all our 100+ Kubernetes clusters. When we switched from Calico to Cilium, we knew we'd face challenges adapting it to our production workloads. Our production big data Kubernetes clusters have a unique usage pattern compared to the other Kubernetes environments within Adyen:

      Data extraction from HDFS. Our infrastructure relies on more than 500 datanodes. Trino represents one of our most demanding HDFS workloads, processing analytical queries against data stored on HDFS. Due to the distributed nature of HDFS, each file you download requires a new connection to any of these 500 nodes. Therefore, during peak hours, a single pod can produce approximately 50,000 connections every minute. +

      Pod churn. Many pods we spawn on the Kubernetes cluster run batch jobs, such as Spark jobs. They stay around for anywhere from a second to a couple of hours.

      Wide variety of workloads. Some workloads are very CPU-intensive, like Spark pods executing complex joins and transformations with relatively few network connections. Others are extremely network-intensive, like Trino pods querying thousands of small files on HDFS, each requiring a new connection. This creates a large number of short-lived connections that stress the connection tracking table.

      Furthermore, our machines are more powerful than most Kubernetes cluster machines within Adyen:

      • Machines with 64 physical cores and 512GB of RAM

      • Machines with 128 physical cores and 2TB of RAM

      +At the time, our environment consisted of Kubernetes 1.31.7 and Cilium 1.16.5. We provide these specific versions to help interested readers correlate our findings with the relevant codebases. 

      To support our unique workloads, we progressively tuned Cilium by increasing DNS proxy timeouts, expanding the connection tracking table capacity, raising API rate limits, and streamlining "security labels". This configuration allowed us to overcome challenge after challenge, except for one:

      `Failed to create pod sandbox: rpc error: code = Unknown desc = failed to setup network for sandbox "0fecf4844d3f8f2df218f09d91f9698bb424e2166551952f46fd7638f4757cf2": plugin type="cilium-cni" failed (add): unable to create endpoint: Cilium API client timeout exceeded`

      Kubernetes would create a pod, spin up the sandbox, and then attempt to set up the networking with Cilium. This operation would time out. From that point onwards, it also timed out any other pod attempting to spawn on the same node. Furthermore, we noticed that on those affected nodes, the API duration for DELETE /v1/endpoint and PUT /v1/endpoint, the routes responsible for adding and deleting a Cilium endpoint, also increased. You can see this clearly in the image below where the endpoint call latency starts to increase linearly over time from 4:20 pm onwards, indicating that endpoint creation and deletion never finish.

      Graph showing payment processing times on Adyen's platform from 9:50 to 17:00.

      It is important to note that these clusters are only used for analytical processing and big data workloads. The issue was identified, analyzed, and fully resolved before any SLO was breached. Furthermore, as our big data environments are isolated from our core transactional flows, this issue never impacted our real-time payment processing pipeline or merchant transactions.

      The symptom: something's wrong, but what?

      First, we found no related timeouts in the Cilium Agent logs. We enabled debug logging, hoping to find hidden errors. Nothing. The logs gave us the big picture but didn't reveal the bottleneck.

      We did learn some valuable things from debugging the broken nodes:

      • Listing all endpoints on a broken node would time out: cilium endpoint list

      • Listing the logs of an endpoint worked fine and showed the endpoint was stuck in the regenerating phase: cilium endpoint log <ENDPOINT_ID>

      • Requesting the endpoint state with  kubectl get ciliumendpoint showed it was still regenerating, while the Kubernetes manifest of the endpoint showed a ready status.

      • Checking the filesystem at /var/run/cilium/state showed endpoint folders with a _next suffix, indicating they didn't complete successfully.

      This validated our suspicion that the issue was in the Cilium agent, not at the kubelet or CNI plugin level. But we still didn't know why.

      Under the hood: how Cilium sets up networking for a pod

      Before we dive into the debugging journey, it's helpful to understand what happens when a pod spawns and Cilium sets up its networking.

      The process starts when Kubernetes assigns a pod to a specific node. The kubelet initiates the pod setup and calls the configured CNI plugin (Cilium, in our case). The Cilium CNI plugin performs several operations:

      1. Allocates an IP using IPAM (IP Address Management)

      2. Creates a link device (in our setup, a veth pair. One end in the host namespace, one in the pod)

      3. Configures the pod network by setting the IP address, configuring routes and setting sysctl parameters

      4. Creates a Cilium endpoint via the Cilium agent API

      5. Retrieves or allocates a security identity using pod labels

      6. Calculates network policy for this endpoint

      7. Generates, compiles, and injects eBPF code into the kernel

      8. Returns success to the kubelet

      The key thing to understand is that when creating a Pod, and thus an endpoint, the CNI plugin calls the Cilium agent, which runs as a DaemonSet on each node. The agent does the heavy lifting: managing eBPF maps, handling connection tracking, applying network policies, and more.

      Here's a simplified view of the flow:

      Flowchart of network setup process involving cabling, CNC plugin, Citrux agent, and kernel.

      A vital part of this process occurs during endpoint creation: Cilium triggers the connection tracking table's garbage collection to verify that the pod's IP is free of residual connections from its previous run. The scrubIPsInConntrackTable function executes this operation, scanning the entire conntrack table to identify and delete relevant entries.

      The consequences of failing garbage collection are significant. Stale entries accumulate without limit, and while our table can accommodate up to 16 million entries, the real bottleneck is the mandatory scan that every new pod requires. This cleaning step is essential to mitigate IP address reuse conflicts, ensuring that a fresh pod doesn't inherit any open connections from a prior one.

      This is where our story really begins.

      Note: Arthur Chiao's excellent deep dive into Cilium's CNI implementation provided much of our understanding of the CNI flow. While Arthur Chiao wrote it a few years ago, the core concepts remain relevant, and it is an invaluable resource for anyone wanting to understand how Cilium works under the hood.

      Down the rabbit hole: tracing the root cause

      What is the agent actually doing?

      We needed to see what the Cilium agent was doing when it hung. Enter pprof, Go's built-in profiler. We enabled pprof in Cilium's configuration and captured traces from a broken node right after spawning 50 pods.

      The CPU and memory profiles didn't reveal much at first. But when we opened the execution traces, the timeline view showed exactly what each goroutine was doing, and everything became clear.

      A dashboard showcasing transaction timelines and payment activity analysis by Adyen

      We saw long-running goroutines spending their entire time in syscalls. Zooming in closer revealed the pattern:

      Timeline of payment processing activities on an Adyen platform dashboard.

      The agent was making BPF syscalls in a tight loop: nextKey(), lookup(), nextKey(), lookup(), over and over again. The agent uses these syscalls to iterate through an eBPF map:

      1. nextKey(currentKey) - Gets the next key in the BPF map after currentKey

      2. lookup(key) - Retrieves the value associated with key

      Let's do some napkin math. We selected a 35-millisecond fragment and counted 14,916 syscall occurrences. That's 426,171 syscalls per second. Since getting each element from an eBPF map takes two syscalls (next + get), we were iterating through roughly 213,000 entries per second. 

      As Cilium is a user-space process, accessing or manipulating the map forces a context switch on every syscall. A context switch is the process of the CPU temporarily halting the user-space process (Cilium) to execute code in the kernel (to handle the BPF syscall) and then resuming the user-space process. This involves saving and restoring the entire state of the CPU registers and memory space, which adds significant overhead and is a major source of latency.

      Here's the critical insight: this is a sequential operation. You can't parallelize it because you need the current key to get the next key. Even on our powerful server CPUs, this was the maximum speed we could achieve. And there was one more important detail in the traces: this was happening in the scrubIPsInConntrackTable function, the one that cleans the connection tracking table when creating an endpoint.

      Why so many syscalls? The connection tracking table

      That's when we remembered something we'd seen in cilium status --verbose:

      Language: plaintext
              BPF Maps:   dynamic sizing: on (ratio: 0.005000)
      +  Name                          Size
      +  TCP connection tracking       16777216
      +  Non-TCP connection tracking   14232516
      +  ...
      +      

      Our TCP connection tracking table had a maximum size of 16 million entries, two times the default of eight million for a machine with 2TB of RAM. We had deliberately configured `cilium_bpf_map_dynamic_size_ratio: 0.0050` in our Helm Chart months earlier, fully expecting the connection tracking capacity to scale proportionally with memory across our node types. At the time, this was a planned scaling adjustment to prevent connection tracking exhaustion on our 512GB RAM machines under high-throughput workloads. This configuration worked as intended to support our unique workload evolution, but as our traffic grew, it created a new scaling challenge in how the larger table interacted with the garbage collection auto-scaling algorithm on our high-resource 2TB RAM nodes.

      But how many entries were actually in the table? We listed them with `cilium bpf ct list global`. This showed about 7 million entries. But here's what made it interesting: when we checked the timestamp of the expires field against the system uptime, we found that the vast majority were already expired, some by several hours. This pointed us directly toward the garbage collection mechanism.

      Now the napkin math gets interesting:

      • Maximum iteration speed: ~200,000 entries per second

      • Current table size: 7 million entries

      • Time to walk the current table: 35 seconds

      • Maximum table size: 16 million entries (at worst)

      • Time to walk the full table: 80 seconds

      +That means every time we created or deleted an endpoint, we would spend as much as 35 seconds iterating through the connection tracking table, and in the worst-case scenario, a maximum of 80 seconds. Keep in mind our CNI timeout constraint is 90 seconds. +

      But wait, if the entries were expired, why wasn’t the garbage collector cleaning them up?

      Why aren't expired entries cleaned up? Garbage collection gone wrong

      Cilium has garbage collection for the connection tracking table. Looking at the metrics, GC was triggering quite often:

      Garbage collection activity over time showing traffic spikes and moderate loads.

      But when we looked at what the garbage collector was actually deleting, we saw a problem:

      Graph showing payment activity spikes over a 24-hour period with Adyen transaction data

      The garbage collector was running frequently but deleting almost nothing most of the time. Only occasionally would it delete significant numbers of entries.

      Digging into the code revealed why. Cilium has two types of GC operations:

      1. Endpoint-specific cleanup: When creating or deleting an endpoint, clean entries matching that endpoint's IP

      2. Periodic expired entry cleanup: Runs on an interval to remove all expired entries

      Pod creation and deletion triggered almost all of the frequent GC runs in the metrics (the type 1 endpoint-specific cleanups). The periodic cleanup (type 2), which removes expired entries, was barely running at all. +

      Why? Because the GC interval auto-scales based on how much it deletes:

      Language: go
              // Simplified from Cilium source
      +func GetInterval(interval time.Duration, maxDeleteRatio float64) time.Duration {
      +    if maxDeleteRatio > 0.25 {
      +        // Deleted > 25% of entries → GC more frequently
      +        interval = time.Duration(float64(interval) * (1.0 - maxDeleteRatio))
      +    } else if maxDeleteRatio < 0.05 {
      +        // Deleted < 5% of entries → GC less frequently
      +        interval = time.Duration(float64(interval) * 1.5)
      +    }
      +
      +    if interval > ConntrackGCMaxLRUInterval {
      +        interval = ConntrackGCMaxLRUInterval  // 12 hours
      +    }
      +    return interval
      +}
      +      

      Here's the problem with a 16-million-entry table:

      • To delete more than 5% (and avoid slowing down), you need to delete 800,000+ entries

      • To delete more than 25% (and speed up), you need to delete 4+ million entries

      • Starting interval: 5 minutes

      • Maximum interval: 12 hours +

      Imagine this scenario:

      1. The node initially experiences a low connection volume when it is newly onboarded or runs only light workloads.

      2. The garbage collector deletes less than 5% of entries during a run, which fails to trigger more frequent cycles.

      3. The garbage collection interval increases progressively from minutes until it reaches the 12-hour maximum. It would start with 7.5 minutes, increase to 11.25 minutes, then to 16.875 minutes, and so on, until eventually reaching the 12-hour limit.

      4. Heavy data workloads eventually land on the node and generate a high volume of network traffic.

      5. New connections fill the tracking table for up to 12 hours before the garbage collector runs again.

      By the time garbage collection eventually executes, the connection tracking table may have already accumulated up to 16 million stale entries. While the GC run might delete enough records to temporarily restore speed, the excessive delay between cycles is inherently problematic. This lag allows the table to accumulate a large number of stale entries again, forcing every pod created during the long interval after the previous GC to endure the significant performance penalty of a sequential scan.

      Graph of real-time payment transaction data from an Adyen terminal or system.

      Why does it cascade? Mutex locks, timeouts, and retries

      While a single slow pod spawn is highly inconvenient, the issue escalated significantly when multiple pods tried to spawn simultaneously.

      We captured a goroutine dump from the Cilium agent using gops and analyzed it with a script that groups similar stack traces. The results were revealing:

      Language: plaintext
              - 57 occurrences:
      +    createEndpoint() → WaitForFirstRegeneration() → waiting on RWMutex
      +- 57 occurrences:  
      +    regenerateBPF() → runPreCompilationSteps() → invoked
      +- 56 occurrences:
      +    scrubIPsInConntrackTable() → garbageCollectConntrack() → waiting for Lock
      +      

      +There's a global mutex on the connection tracking table. When we spawn 50 pods at once:

      • Pod 1 acquires the lock and starts the 80-second table iteration

      • Pods 2-50 queue up waiting for the lock

      • Pod 1 finishes after 80 seconds

      • Pod 2 acquires the lock, starts another 80-second iteration

      • But Pod 2's timer started 80 seconds ago → timeout at 90 seconds

      • Pod 2 times out

      • Pods 3-50 never stand a chance

      +Here's where it gets worse. When the CNI times out after 90 seconds:

      • The timeout returns an error to the caller

      • But the underlying work doesn't stop, the agent keeps iterating

      • The container runtime (containerd) immediately calls DeleteEndpoint()

      • Delete also needs to walk the conntrack table

      • Now the system queues up both create and delete operations

      And then Kubernetes retries:

      • The kubelet's podWorkerLoop retries after 60-90 seconds (with jitter)

      • Each retry adds another endpoint creation and deletion request to the queue

      • The queue grows faster than it drains

      We could see this in the logs. For one pod (cilium-node-breaker-5), we saw:

      • 15:54:30 - Create endpoint (attempt 1)

      • 15:56:00 - Delete endpoint (timeout)

      • 15:57:12 - Create endpoint (attempt 2)

      • 15:59:57 - Create endpoint (attempt 3)

      • 16:02:39 - Create endpoint (attempt 4)

      The node enters a contention cycle: new work arrives faster than old work completes, and the queue never drains.

      Here's the full picture:

      Flowchart illustrating a payment process with Adyen terminals, OLV plugin, and checkout steps.

      The fix: one line to rule them all

      After all that investigation, the fix was anticlimactic in its simplicity. We couldn't rely on the auto-scaling GC interval because it would inevitably grow too long on quiet nodes. Hence, we prevented the GC interval from auto-scaling by setting a fixed value:

      Language: yaml
              conntrackGCInterval: 60s
      +      

      That's it. One configuration line ensures garbage collection runs at least every minute, regardless of how much it deletes. We applied the change at 9:00 am and completed the DaemonSet rollout by 10:00 am. The results speak for themselves:

      Line chart displaying payment transaction data over time at an Adyen endpoint

      The conntrack table size dropped dramatically and stayed stable. More importantly, the API call durations returned to normal:

      Bar chart illustrating transaction volume and payment data for Adyen services over time.
      Graph showing transaction volume and payment activity data from an Adyen system.

      Since the fix, we haven't seen a single instance of the timeout error. Pod spawn times are reliable again.

      Lessons learned

      Scaling parameters can have long-tail interactions. Our proactive tuning of bpf_map_dynamic_size_ratio to support workload scaling on 512GB RAM machines successfully resolved initial capacity limits. However, as our analytical workloads evolved and traffic increased, the larger table size dynamically allocated on our 2TB RAM machines revealed a subtle interaction with the CNI's GC auto-scaling algorithm. These scaling parameters can take months to show their full impact as traffic patterns grow, particularly in environments with adaptive background loops.

      Observability and full-stack understanding are critical. While logs showed symptoms, we needed profiling and tracing to reveal the root cause across the entire stack. The container runtime (timeouts, delete behavior), the CNI plugin (timeout values), the Cilium agent (mutex locks, GC logic), and the Linux kernel (eBPF maps, syscall performance) were all relevant to understanding how we got to the pod spawn timeouts. On top of that, adding napkin math with real numbers proved very powerful. Once we had the key numbers, 200k syscalls/sec, 12M table entries and a 90-second timeout, we determined the root cause far before we understood the full chain. Always measure your system's actual performance characteristics, not just theoretical limits.

      Auto-scaling algorithms need bounds. Cilium's GC interval auto-scaling makes sense for most deployments: if you're deleting lots of entries, run GC more often; if you're deleting few entries, save CPU by running GC less often. But the algorithm didn't account for varying workloads, where a machine has a low connection volume for a prolonged period, after which, with a single pod introduction, it could get a very high connection volume. Nor did the algorithm account for very large tables where "5% of entries" is an enormous absolute number. The 12-hour maximum interval was too long for our workload. Auto-scaling without careful consideration of edge cases can backfire.

      Timeouts don't stop work. When the CNI timed out, we assumed the work would stop. It didn't. The agent kept processing in the background while new requests queued up. This is a common pattern in distributed systems: timeouts protect the caller but don't necessarily cancel the operation. Be explicit about cancellation when needed.

      Treat conntrack health as a first-class operational metric. The difference between a healthy cluster and a contention cycle showed up clearly in some metrics we weren't watching: 

      • GC duration - cilium_datapath_conntrack_gc_duration_seconds - jumped from 1s to 80s

      • Table size - cilium_datapath_conntrack_gc_entries - 7M entries, mostly expired

      Proactively alerting on these metrics is something we now recommend for any Cilium deployment with dynamic workloads, alongside setting `conntrackGCInterval: 60s`. Don't optimise for CPU savings during quiet periods at the expense of pod spawn timeouts during busy periods.

      Conclusion

      A single configuration line ultimately resolved the mysterious timeout error that impacted our ability to spawn new pods on our big data platform: conntrackGCInterval: 60s. Our investigation revealed that the root cause of our pod timeouts was Cilium's auto-scaling garbage collection algorithm, allowing the cleanup interval to grow to 12 hours, leading to a massive accumulation of expired entries and a linear-time iteration trap.

      This experience provided major takeaways regarding system resilience and the necessity of a full-stack understanding. We learned that scaling parameters and resource allocations can have long-tail interactions that only surface months later as workloads evolve. Furthermore, we discovered that auto-scaling algorithms require strict bounds to prevent unexpected performance degradation in edge cases, such as the varying connection volumes we see on our high-resource machines. The investigation also highlighted that timeouts often only protect the caller, without stopping the underlying work, potentially triggering a contention cycle of retries that we could only diagnose through deep observability into mutex locks, syscalls, and eBPF internals.

      As we move forward, we must ask ourselves: are the adaptive behaviours in our infrastructure truly protecting us, or are they masking inefficiencies that only appear at peak capacity? By treating conntrack health as a first-class operational metric and prioritising reliability over minor CPU savings, we can build more robust systems. And remember, if you ever see mysterious timeouts in your CNI: sometimes the answer hides in 426,000 syscalls per second.

      Fresh insights, straight to your inbox

      + diff --git a/sreweekly/articles/532/04-storage-at-scale-what-i-actually-watched.html b/sreweekly/articles/532/04-storage-at-scale-what-i-actually-watched.html new file mode 100644 index 00000000..1af1b406 --- /dev/null +++ b/sreweekly/articles/532/04-storage-at-scale-what-i-actually-watched.html @@ -0,0 +1,137 @@ + Storage at scale: what I actually watched | Sridhar Rajarao + + +

      + ·  +sre, storage, reliability, metrics


      Storage at scale: what I actually watched

      For eight years I ran SRE for a storage system measured in exabytes. The dashboard I checked every morning shrank to seven numbers. Here they are.

      For eight years I ran the SRE team behind a storage system measured in exabytes. Over time, the dashboard I checked every morning shrank to a handful of numbers. These are the seven that told me whether the service was healthy.

      +
      +

      Availability tells you if the system is up. Durability tells you if your data is still there. The two are not the same.

      +
      +

      Here’s the short version.

      + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +
      KPIWhat it measuresHow we tracked it
      AvailabilityPercent of requests succeeding99.99% per region, per service
      DurabilityProbability your data survives11 nines (10^-11 annual loss)
      TTFBTime to first byte returnedp50, p95, p99 latency per object size
      CanariesSynthetic test trafficContinuous PUT/GET from every region
      HotspotsSkew across storage nodesTop-N node load vs cluster median
      IOPSOperations per secondRead/write IOPS per shard, per disk
      DB ShardsMetadata partition healthShard CPU, lag, hot-key skew
      +

      Availability and durability are the two non-negotiables

      +

      Availability is uptime. Durability is whether the data survives. You can be 100% available and lose data, you can be 100% durable and offline. Customers care about both. We hit 11 nines of durability by writing every object to multiple availability domains with erasure coding, and proved it monthly with a recovery drill.

      +

      TTFB is what users actually feel

      +

      Aggregate availability hides slow tails. A 99.99% available service with a p99 TTFB of 2 seconds feels broken. Always track latency by object size bucket. A 10 MB read should not share an SLO with a 100 byte HEAD.

      +

      Canaries are your truth

      +

      Customers don’t tell you when they’re sad. They leave. Canaries are synthetic PUT/GET/LIST traffic running continuously from every region. If a canary fails for 30 seconds, you find out before your customer’s pager goes off.

      +

      Hotspots and IOPS surface the silent failures

      +

      A storage cluster can be 99.99% available while one node is on fire. Track per-node IOPS and bytes-served, and alert on the top-N nodes diverging from cluster median. Hotspots are the leading indicator of a customer key range overwhelming a shard.

      +

      DB shards are the part nobody talks about

      +

      Object storage looks stateless, but the metadata layer is a sharded database. One hot shard, one rebalance gone wrong, and your control plane stalls. Watch shard CPU, replication lag, and hot-key skew the same way you watch the data plane.

      +
      +

      The data plane scales. The control plane bites.

      +
      +

      Those seven numbers, watched together, told me almost everything I needed to know about whether the service was healthy.

      ← All writing

      + \ No newline at end of file diff --git a/sreweekly/articles/532/06-how-uber-conquered-database-overload-the-journey-from-static-rate-limi.html b/sreweekly/articles/532/06-how-uber-conquered-database-overload-the-journey-from-static-rate-limi.html new file mode 100644 index 00000000..28ebf3e0 --- /dev/null +++ b/sreweekly/articles/532/06-how-uber-conquered-database-overload-the-journey-from-static-rate-limi.html @@ -0,0 +1,748 @@ +How Uber Conquered Database Overload: The Journey from Static Rate-Limiting to Intelligent Load Management + + + + + + + +
      Skip to main content
      + +
      April 20, 2026

      How Uber Conquered Database Overload: The Journey from Static Rate-Limiting to Intelligent Load Management

      Dhyanam Vaidya
      Prathamesh Deshpande
      Mike Ma
      1+
      cover-photo-non-ai-17682774511395.jpg
      Share this article

      Introduction

      Uber’s thousands of microservices handle traffic for over 170 million monthly active users: riders, Uber Eats users, drivers, and couriers. At the heart of this infrastructure are Docstore and Schemaless, Uber’s in-house distributed databases built on top of MySQL®. These databases span thousands of clusters, store tens of petabytes of operational data, and serve tens of millions of requests per second with billions of rows read or updated. They back some of the most latency-sensitive and mission-critical workloads, powering every business vertical at Uber: from rides and deliveries to maps, payments, and beyond. 

      At this scale, even minor overloads aren’t isolated events, they cascade. A brief spike in one part of the system can ripple outward: downstream services time out, retries pile up, and degradation amplifies into broader failure. In a multitenant environment, it’s also critical to ensure fairness and prevent any tenant from hogging all the resources. With workloads varying in traffic shape, latency profiles, and system impact, building effective overload protection is a uniquely challenging problem.

      The cost of getting overload protection wrong is steep. This blog shares how we built an intelligent load manager that detects overload from multiple signals to keep our databases stable and fair under pressure.


      Docstore and Schemaless

      Before diving into the load manager that protects Uber’s databases, let’s walk through their architecture.

      While Docstore supports transactions with full CRUD operations and Schemaless is optimized for append-only workloads, both share a common architectural foundation. It comprises three primary layers: a stateless query engine, a stateful storage engine, and a control plane. For the scope of this blog, we’ll focus on the query and storage engine layers.


      Figure 1: Docstore and Schemaless architecture.

      The stateless query engine is responsible for query planning, request routing, sharding, schema management, authorization, request parsing, and validation. It serves as the routing layer: coordinating and validating client requests before handing them off to the storage layer.

      The stateful storage engine handles transaction management, connection pooling, consensus, and replication. Data is sharded across multiple partitions, with each partition consisting of one leader and two followers, coordinated via Raft to ensure strong consistency. Each partition is backed by MySQL nodes with locally attached NVMe SSDs, built to support high-throughput, low-latency workloads at scale.


      Challenges

      Quota-Based Rate Limiting in the Query Engine Layer

      Figure 2: Quota-based rate-limiting setup.

      Initially, we explored a quota based rate-limiting approach within the stateless query engine layer. The concept was simple: assign each read and write request a capacity unit cost based on bytes processed, grant users fixed quotas, and return a 429 when those quotas were exceeded. Since routing nodes were stateless, we stored quota usage in a central Redis® cache. While conceptually sound, this approach didn’t hold up in production.

      First, it added unnecessary complexity. Every request required a Redis call, introducing a new point of failure and the overhead of an additional network hop.

      Further, for the stateless routing layer to accurately shed requests for an overloaded storage partition, it’d need to maintain realtime health and load information for thousands of partitions across the system. This introduced a lot of tracking overhead, undermining the scalability of the architecture.

      The cost model was also too imprecise. In Docstore and Schemaless, due to the way MySQL handles scanning and filtering, a query that performs a full table scan but returns a single row was assigned the same capacity cost as a query that only reads a single row. This fundamental flaw in our metering meant that lightweight and heavyweight operations were treated the same, making quota enforcement unreliable.

      Finally, quotas were defined statically, resulting in frequent requests from stakeholders to adjust their quotas, making them ineffective in multitenant environments. 

      Despite its initial promise, this approach failed. But it gave us a crucial insight: overload management must live as close to the storage nodes as possible. That realization became a cornerstone of the final design in the stateful storage layer.


      Identifying the Right Signal for Overload

      A core challenge in designing a resilient load manager is choosing a reliable signal for overload. Simple QPS-based rate limiting is too coarse. It fails to account for workload variability, often shedding too late or too early. What can be more effective is concurrency: the number of operations currently in flight. It directly reflects system load, following Little’s Law: Concurrency = Throughput × Latency. In stateful systems, it maps closely to resource usage, making it a more dependable indicator.


      Balancing Resilience and Fairness

      Balancing resilience and fairness is a core challenge in multitenant systems. During ‌system-wide stress, we want to shed traffic by priority, dropping low-priority requests first. But when a single noisy actor hogs resources without triggering global overload, we also need per-tenant rate limiting that works independently of the system load. This dual requirement led us to combine dynamic overload detectors with fairness enforcement mechanisms that operate in parallel.


      Building the Foundation of a Unified Load Manager

      Figure 3: Initial load manager setup with CoDel queue.

      Controlled Delay: Smarter Queuing Under Pressure

      The load-shedding journey began with CoDel (Controlled Delay), a concept borrowed from networking to combat bufferbloat. Instead of shedding based on queue length, CoDel looks at how long requests wait in the queue: favoring responsiveness over volume.

      We implemented separate CoDel queues for each operation type:

      • Read queue: for point lookups and light queries
      • Write queue: for insert, update, and upsert operations
      • Slow queue: for long-running and background operations like scans, deletes, or replication

      Each queue was managed independently, giving us better isolation across workloads.


      Figure 4: CoDel queue behavior.

      FIFO queuing wasn’t enough because a pure FIFO queue processes requests in arrival order, which works well when traffic is stable. But under overload, FIFO creates a trap: old requests accumulate, wait too long, and often get abandoned or retried by the client. This results in wasted work. Meanwhile, fresh requests, still relevant and likely to succeed, sit idle at the end of the line.

      CoDel introduces adaptive LIFO to solve this. Figure 5 shows how it works.


      Figure 5: CoDel algorithm.

      Under normal load, the queue behaves as FIFO. Under pressure, it switches to LIFO, favoring newer requests that still have a chance to succeed. This simple shift improves responsiveness by failing fast, shedding stale work, and giving fresh requests priority. 

      Scorecard Engine

      The Scorecard engine is a rule-based admission control component and a lightweight quota system designed to enforce per-tenant concurrency limits in multitenant environments. While load-shedding protects the system during overload, Scorecard ensures that no single tenant can dominate shared infrastructure, even in normal conditions. 

      The configuration is simple and deterministic.


      Figure 6: Scorecard rules.

      The primary benefit of the Scorecard lies in incident containment. It helps pinpoint the source of disruption during outages or traffic spikes. It isolates and caps misbehaving tenants without disrupting others, balances stability during normal load with strict limits under stress, and reduces blast radius during overload events by enforcing boundaries quickly and deterministically.

      The Scorecard provides predictable fairness and blast radius control, especially when multiple tenants are competing for shared resources.

      Regulators

      While Scorecard protects against concurrency-based overuse, it doesn’t cover all the ways a stateful database system can overload. Some forms of skews are subtle. They don’t show up in concurrency saturation, but they can still degrade system performance if left unchecked.

      For example, a low QPS caller can still overload the system by sending large write payloads. Or, traffic skewed to one partition key can overload a single cluster while others sit idle.

      To guard against these skewed behaviors, we introduced plug-in regulators: node-local overload detectors that enforce invariants the system mustn’t violate. They rarely trigger during healthy operation, and that’s by design. At the same time, when users accidentally create hotspots or large data ingestions, regulators kick in to prevent cascading failures.

      We use these regulators:

      • Write bytes regulator: Limits concurrent write volume to prevent I/O saturation
      • Partition key regulator: Throttles traffic targeting hot partition keys
      • Memory regulator: Tracks free process memory and throttles when we’re low on memory
      • Goroutines regulator: Tracks total number of goroutines and throttles when it exceeds threshold

      What Worked Well

      By shedding excess requests, our CoDel queues prevented runaway resource exhaustion, which led to improved stability and a higher success rate for accepted requests. This approach was particularly effective at ensuring that core system functionality remained available during overloads.


      Figure 7: Improved availability.

      The Scorecard engine successfully isolated misbehaving tenants by enforcing per-tenant concurrency limits. This allowed us to quickly contain disruptions from noisy neighbors without penalizing other users, ensuring that shared resources were used fairly.

      Limitations

      While this initial setup laid the foundation for overload protection and fairness, it came with a few limitations. First, CoDel treated all requests equally, dropping low-priority and user-facing traffic alike, leading to a bad customer experience and increased on-call load.

      CoDel also relied on fixed queue timeouts and static inflight concurrency limits, which can be a low-fidelity solution for a dynamic system, requiring frequent manual tuning and leading to operational toil.

      The fixed, static wait times in CoDel led to a thundering herd problem. When requests were eventually rejected, they’d all retry at once, triggering repeated cycles of overload and rejection. During these periods, the lack of traffic differentiation meant even high-priority requests were dropped, leading to customer-visible errors and amplifying the blast radius.

      Ultimately, it kept things from breaking, but lacked the nuance and dynamism required for a high-quality user experience. This highlighted the need for dynamic and priority-aware queues.


      Evolving the Architecture

      Cinnamon Replaces CoDel

      We observed that many overloads stemmed from low-priority, asynchronous jobs: pipelines, aggregators, and internal garbage collection flows. These shouldn’t have the same survivability as ride requests or real-time pricing queries.

      To address this, we replaced CoDel with Cinnamon, a priority-aware load shedder developed by the Delivery team at Uber. Cinnamon makes smarter shedding decisions by considering request rank, dynamic system state, and the relative importance of workloads. 

      Request rank is derived from the priority attached to the request, and if no explicit priority is present, Cinnamon assigns a default based on the calling service. Priority is defined using a tiering model from tier 0 (t0) for the most critical traffic to tier 5 (t5) for the least. While t0 is reserved for a small subset of critical infrastructure services, t1 represents the most important user facing online traffic, the core workloads we aim to protect during overloads. This system allows Cinnamon to shed lower-priority traffic first during overload.

      With request priority awareness in place, we simplified the queue structure to just read and write queues. Long-running and background operations were marked with lower priority instead of having a separate queue.


      Figure 8: Updated load shedder setup with Cinnamon queue.

      Before Cinnamon, the CoDel queue load shedder was priority-agnostic and shedding during overload was indiscriminate.


      Figure 9: Priority Agnostic Load shedder setup with CoDel queue.

      After Cinnamon, the queue load shedder was priority-aware and shedding during overload happened in order of priority.


      Figure 10: Priority Aware Load shedder setup with Cinnamon queue.

      Performance and Stability Gains

      We saw performance and stability gains from the Cinnamon-based design. Requests are ranked, allowing Cinnamon to shed low-priority traffic first, protecting user facing flows. During overloads, critical user-facing requests are better protected with minimal impact. 


      Figure 11: Prioritized shedding in action.

      Cinnamon also adapts queue timeout thresholds using P90 latency metrics, eliminating the need for manual tuning. Moreover, its Auto Tuner dynamically adjusts inflight limits, represented by the available slots in the blue box in Figure 10, to maximize throughput. It does this by continuously monitoring and reacting to realtime latency and error rate signals, ensuring stable and effective load shedding.

      Unlike CoDel’s static approach, which aggressively rejects all requests after a fixed wait time, like 5 milliseconds, Cinnamon’s PID-based control allows the system to absorb pressure without overreacting. It dynamically adjusts queue timeouts and inflight limits based on realtime latency and error signals, shedding only when necessary. This prevents a large class of premature shedding that would otherwise lead to unnecessary rejections, retries, and thundering herd effects. The result is smoother recovery, fewer 429s, and more consistent availability without compromising system health.


      Figure 12: Reduced premature shedding.

      Areas for Improvement

      Despite the gains from Cinnamon, some key challenges remained, highlighting the need for a unified platform.

      The load manager acted based on the local health of the server, tracking signals like inflight concurrency, write bytes, or memory usage. But in distributed systems, overload isn’t always local. A leader node may need to shed traffic because follower nodes are lagging, even if it’s healthy itself. We call this commit index lag. Traditionally, external components using token-bucket-based rate limiters handled such remote shedding decisions. These were easy to build but proved ineffective at scale, introducing split-brain behaviors and globally suboptimal shedding decisions.

      The initial design was excellent for concurrency-based shedding, but it wasn’t built to be a reusable platform for future overload signals that would inevitably arise from a growing system.

      These insights led us to the final evolution of our system: transforming Cinnamon from a concurrency only shedder into a truly general purpose overload control engine. By consolidating all signals into a single, modular decision-making loop, we achieved holistic and consistent overload management.


      The Unified Load Shedding Engine

      Centralizing Overload Decisions

      We enhanced Cinnamon to support pluggable external signals like follower commit lags, enabling the system to make globally informed, priority-aware shedding decisions within the same admission control path. This shift unified local and remote overload logic into a single control loop, closing the gaps that previously caused instability.


      Figure 13: Unified load-shedding engine in Cinnamon.

      But shedding isn’t always a one-size-fits-all decision and that’s where the load manager architecture shines. Built on a BYOS (Bring Your Own Signal) ethos, it provides a pluggable framework that lets the team embed new overload signals and route them to the right control path. Whether the pressure is systemic or actor-specific, the load manager sheds broadly by priority or precisely by caller, based on the signal. 


      Figure 14: Bring your own signal.

      The Payoff: Unified Control, Simplified Load Management

      The shift to a centralized, pluggable architecture made the system more stable and predictable, with real wins.

      Cinnamon sheds excess requests immediately using a PID controller, avoiding the memory and goroutine buildup caused by token bucket limiters. This led to lower tail latencies and a leaner resource usage profile, even under heavy load. We saw: 

      • 80% increase in throughput under overload (QPS average of 5,400 versus 3,000)
      • ~70% reduction in P99 latency (upsert average of 1.0 seconds versus 3.1 seconds)
      • ~93% fewer goroutines during overload (peak 10,000 versus 150,000)
      • ~60% lower heap usage (1 GB max versus 5-6 GB spikes)

      Figure 15A: (Before) Token bucket latency and resource profile.

      Figure 15B: (After) Cinnamon latency and resource profile.


      We also saw smoother, more predictable shedding behavior. Without PID regulation, shedding acts like a hammer: reactive and abrupt. With it, it’s more like a dimmer switch: smooth and stable. The difference is clear when comparing how commit lag stabilizes under a token bucket limiter versus Cinnamon’s PID-based controller.


      Figure 16A: (Before) Token bucket spiky shedding pattern.

      Figure 16B: (After) Cinnamon’s stable shedding pattern.

      Lessons Learned 

      • Prioritization is paramount. Effective load-shedding starts with deciding what matters most. Protect critical, user-facing traffic first. Everything else is secondary.
      • Fail fast, don’t block. Rejecting early is almost always better than holding requests in memory until they expire. It reduces wasted work, keeps latencies predictable, prevents OOMs, and makes the system more resilient under stress.
      • PID regulation for stable shedding. Simple, reactive shedding based solely on current error rates often causes instability, overcorrecting too late, and too hard. PID based regulation brings balance by incorporating system history and directional trends, making it a critical tool for smooth, sustained, and resilient overload control.
      • Place control close to the source of truth. The best shedding decisions happen where the state lives. Protection in the layer that has full context, typically the storage layer in stateful systems.
      • Embrace dynamism. Avoid static configurations wherever possible. Your system should be intelligent enough to adapt to different scenarios, based on the context.
      • Invest in visibility and monitoring. Good observability is the foundation for tuning and trust. Track what’s being shed, why it’s being shed, and how each component contributes to system pressure.
      • Simplicity over complexity. This is a meta principle that guides all the other decisions. 

      Conclusion

      Our journey to a resilient load manager was defined by the unique complexities of a large-scale, stateful, and distributed environment. By unifying disparate components into a single decision-making brain and adopting a Bring Your Own Signal model, we gained the flexibility to handle systemic overloads and localized noisy neighbor issues with precision. The result is a load management system that sheds smarter in a priority-aware manner, keeps tail latencies low, and drastically reduces operational toil. 

      If you like challenges related to distributed systems, databases, storage, and cache, apply for open positions here.

      Acknowledgments

      A project of this scope is rarely accomplished alone. Our sincere thanks to Rich Porter, Jesper Nielsen, Piyush Patel, and the engineers from the Storage and Delivery teams for their guidance and collaboration throughout this journey. From design reviews to on-call insights, their contributions were instrumental in building a resilient system that now safeguards some of Uber’s most critical infrastructure.

      Cover Photo Attribution: “Heavy Traffic Jam in Urban City Center” by Dapur Melodi

      MySQL is a registered trademark of Oracle and/or its affiliates. Other names may be trademarks of their respective owners.

      Redis is a trademark of Redis Labs Ltd. Any rights therein are reserved to Redis Labs Ltd. Any use herein is for referential purposes only and does not indicate any sponsorship, endorsement or affiliation between Redis and Uber.

      Written by

      Dhyanam Vaidya

      Dhyanam Vaidya is a Software Engineer on Uber’s Storage Platform team. He’s contributed to the design and implementation of many Docstore features. His work focuses on improving the reliability, resilience, and operational efficiency of Uber’s distributed databases at scale.

      Prathamesh Deshpande

      Prathamesh Deshpande is a Staff Engineer on Uber’s Storage Platform team, building database features and distributed storage systems that meet Uber’s global reliability and performance requirements. His work focuses on large-scale data management, distributed database storage systems, and platform reliability.

      Mike Ma

      Mike Ma is a Staff Software Engineer on Uber’s Storage Platform team, where he has contributed to multiple core components of both Schemaless and Docstore. His work focuses on scalability, reliability, performance, and operational excellence across Uber’s large scale distributed databases.

      Chaitanya Yalamanchili

      Chaitanya Yalamanchili is a Sr. Manager and technical lead on Uber’s Storage Platform team. He leads the development of online distributed storage systems with a focus on providing a world-class platform that powers all the critical business functions and lines of business at Uber. The platform serves tens of millions of QPS and stores tens of Petabytes of operational data.

      + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + \ No newline at end of file diff --git a/sreweekly/articles/532/07-voyager-and-the-art-of-graceful-degradation.html b/sreweekly/articles/532/07-voyager-and-the-art-of-graceful-degradation.html new file mode 100644 index 00000000..9b50865a --- /dev/null +++ b/sreweekly/articles/532/07-voyager-and-the-art-of-graceful-degradation.html @@ -0,0 +1,4122 @@ + + + + +Voyager and the Art of Graceful Degradation + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +Skip to main content + +
      +
      +
      +
      +
      +
      +
      +
      + +
      +
      +
      +
      +
      +
      +
      +
      +
      +
      + + +

      +Voyager and the Art of Graceful Degradation +

      +
      + +
      +
      +
      + +
      +
      +
      +

      Like a great celestial swan, Voyager 1 is flying — swiftly, boldly, albeit a little stiffly in places. 

      https://assets.science.nasa.gov/dynamicimage/assets/science/missions/voyager/images/1_Voyager_artist_concept.jpg?w=2000&h=1125&fit=crop&crop=faces%2Cfocalpoint
      NASA/JPL-Caltech

      It moves through interstellar space with enormous momentum, far beyond the planets that once defined its mission, carrying instruments that continue to report from a region no man‑made craft has ever reached. Yet every action it takes is constrained by a finite and steadily diminishing supply of energy, each signal carefully weighed against what it costs to send.

      +

      There is a quiet elegance in that balance.

      +

      Voyager does not insist on doing everything it once did. It does not pursue peak capability when conditions no longer allow it. Instead, it adapts -- releasing some functions so that others can continue, prioritizing what matters most over what is merely possible.

      +

      In engineering, we have a name for systems that behave this way.

      +

      We call it graceful degradation.

      On April 17, 2026, engineers at NASA’s Jet Propulsion Laboratory sent a carefully prepared sequence of commands to Voyager 1, asking it to stop doing something it had done almost continuously since 1977: measuring low‑energy charged particles in deep space.
       
      After nearly 49 years of operation, Voyager 1’s Low‑Energy Charged Particle (LECP) detector — an instrument that measures ions, electrons, and cosmic rays to map the structure and pressure of the interstellar medium, helping to define the boundary between the solar system and interstellar space— was shut down to conserve power and extend the spacecraft’s operational life.

      Voyager is powered by a radioisotope thermoelectric generator whose output declines as radioactive fuel decays. Every year, available power drops by a few watts. Unlike systems here on Earth, there is no possibility of provisioning more capacity, no redundancy waiting in reserve, and no “scale out” option.

      +

      Seen through a Site Reliability Engineering lens, Voyager’s power margin is its error budget. It defines how much can go wrong before the mission begins to suffer.

      +

      Early in the mission, that budget was generous. Minor inefficiencies, unexpected behaviors, and non‑optimal configurations could be tolerated. As the decades passed, the margin narrowed. Today, even a modest, unplanned dip of power by a wayward instrument risks triggering Voyager’s undervoltage fault protection — an automated safeguard that will shut components down abruptly to ensure survival.

      +

      In February, a routine roll maneuver caused such a dip. Engineers understood that allowing the spacecraft to cross that line would mean entering a survival mode where system preservation is prioritized over delivering mission value.

      +

      This moment is familiar to anyone who has operated a production system near its limits:

      • CPU saturation turning latency into user-visible slowness
      • +
      • Memory pressure triggering process and container termination
      • +
      • Queues backing up until messages expire undelivered
      • +
      • Storage exhaustion freezing otherwise healthy transactions

      Graceful degradation is about prioritizing your goals and your capabilities, and as you approach a point where you cannot fulfill all your goals, acting before you reach that point.

      • Reduce CPU consumption (lower frame rates, remove animations, disable optional features)
      • +
      • Defer low‑priority work (batch reports, replace live data with aggregates)
      • +
      • Prioritize critical traffic and drop nonessential messages
      • +
      • Reject new transactions when storage thresholds are reached to protect core paths

       In Reliability Engineering, as in much of life, we'd rather have brownouts than blackouts.

      +

      While we never want to disappoint users, we'd rather reduce features rather than take outages. +We'll degrade experience -in a controlled fashion - rather than lose the service entirely. +We shed load in controlled ways instead of letting cascading failures decide the outcome.

      +

      That is exactly what Voyager’s engineers did.

      +

      Years before this moment, scientists and engineers jointly agreed on a shutdown sequence: which instruments would be sacrificed first as power declined, and which capabilities were most critical to preserve. By April 2026, seven of Voyager 1’s ten original science instruments had already been retired. The LECP was simply next on the list — not because it failed, but because its cost‑to‑value ratio was now unfavorable.

      +

      This is the same decision Site Reliability Engineers (SREs) make when:

      +
      • Disabling expensive recommendation pipelines during peak traffic
      • Serving cached or approximate results instead of fully computed ones
      • Temporarily turning off background jobs to protect user‑facing latency
      +

      Nothing is broken, per se. The system is deliberately choosing to do less so that it can continue to succeed in part rather than fail in total. Graceful degradation is not a weakness; it is a sign of maturity.

      +

      Voyager continues to operate the instruments that provide uniquely valuable data — measuring magnetic fields and plasma waves in interstellar space — while relinquishing others whose contribution, though still useful, no longer justifies their cost.

      +

      Even the LECP shutdown was reversible by design. A small motor that rotates the sensor remains powered, preserving the option of reactivation should future power‑saving measures succeed.

      This is graceful degradation with reversibility in mind. The current state is preserved, while recovery paths maintained and, most importantly, options are left open. Granted, the chances of Voyager suddenly being replenished with fresh plutonium for additional power is exactly 0, but Reliability Engineers here on the ground do plan on overcoming their temporary issues which caused the degradation and using the available options to fully restore services. 

      This is why we gate features behind flags instead of deleting code and why we can temporarily change users' capabilities instead of removing them from the system.

      Balancing Performance, Capacity, and Risk

      Reliability is rarely about maximizing performance. It is about continuously balancing performance, capacity, and risk — especially when capacity is finite and margins are thin.

      +

      Voyager operates permanently at this intersection.

      +

      Performance, in Voyager’s case, is scientific throughput: how many instruments are active, how often measurements are taken, and how much data is returned. 
      Capacity is a steadily shrinking power budget that cannot be replenished. 
      Risk grows as margins shrink: a sudden undervoltage event could trigger autonomous shutdowns that are difficult, slow, and dangerous to recover from across a 23‑hour communication delay.

      +

      Graceful degradation is how the Voyager team manages this triangle.

      +

      By shutting down the LECP before power levels became critical, the team deliberately traded peak scientific performance for reduced operational risk and preserved capacity for the instruments that matter most. 

      This illustration shows the various instruments locations on the Voyager spacecraft.
      The status of Voyager's instruments  (NASA/JPL-Caltech)

        

      Voyager does less than it once did — but it does so more safely, more predictably, and for longer.

      +

      This mirrors everyday SRE work:

      +
      • lowering request concurrency to prevent saturation
      • reducing image quality or refresh rates under load
      • shrinking feature scope during high-risk windows
      • renegotiating SLOs instead of pretending nothing has changed
      +

      In each case, performance is intentionally reduced to keep risk within acceptable bounds.

      +

      Because Voyager’s degradation path was defined years in advance, with a healthy system and with management & engineering having time and clarity to make rational trade-offs, the unexpected power dip didn't result in a frantic rush to heroically solve a problem, it triggered a pre-planned process which resulted in a graceful retirement of the instrument chosen ahead of time. No surprises, just good engineering. 

      Graceful degradation is a social and organizational capability as much as a technical one. It requires shared understanding across teams, explicit agreement on priorities, and acceptance that loss is inevitable. It's not about preventing failure forever. It is about ensuring that when degradation occurs, it happens on your terms.

       While few of us work on systems like Voyager, there are many commonalities - 

      • Our platforms are usually far older than the original business model they were designed to support.
      • Our architectures often outlast our architects.
      • Our "temporary" services that we built with “temporary” design decisions have become permanent.
      +

      Our systems survive not by staying perfect, but by letting go gracefully. At least, these are the ones which cause the least stress to their owners and maintainers. 

      Voyager is still returning data from interstellar space not because nothing has failed, but because failures have been managed thoughtfully, incrementally, and with humility. Twenty-five billion kilometers from Earth, Voyager continues to demonstrate a lesson every experienced SRE eventually learns:

      +

      The systems that last longest are not the ones that cling to every feature, but the ones that decide, and well in advance, which parts they are willing to give up.

      +

      If Failure is Not an Option, then Graceful Degradation is Mandatory.

       

      +
      +
      + + +
      +
      +
      +
      + +

      Comments

      +
      +
      + +
      +
      +
      +
      +
      +
      +

      +Popular Posts +

      +
      +
      +
      +

      Excellence Is a Habit

      +
      +
      + +
      +
      +
      +
      + +Image + +
      + + +
      +
      +
      +

      Artemis and Apollo: The Systems That Took Them to the Moon — and Brought Them Home

      +
      +
      + +
      +
      +
      +
      + +Image + +
      + + +
      +
      +
      +
      +
      +
      +
      +
      +
      + + +
      +
      +
      +
      +
      + + + + + + + \ No newline at end of file diff --git a/sreweekly/articles/532/index.json b/sreweekly/articles/532/index.json new file mode 100644 index 00000000..c0846f97 --- /dev/null +++ b/sreweekly/articles/532/index.json @@ -0,0 +1,44 @@ +[ + { + "idx": 1, + "url": "https://greatcircle.com/blog/2026/08/11/declaring-incidents-for-side-effects/", + "ok": true, + "error": null + }, + { + "idx": 2, + "url": "https://queue.acm.org/detail.cfm?ref=rss&id=3830399", + "ok": false, + "error": "HTTP 403" + }, + { + "idx": 3, + "url": "https://www.adyen.com/knowledge-hub/inside-cilium-cni-solving-kubernetes-pod-setup-timeouts", + "ok": true, + "error": null + }, + { + "idx": 4, + "url": "https://sridharrajarao.com/blog/storage-at-scale/", + "ok": true, + "error": null + }, + { + "idx": 5, + "url": "https://pub.towardsai.net/the-rise-of-cognitive-observability-a77e33250037", + "ok": false, + "error": "HTTP 403" + }, + { + "idx": 6, + "url": "https://www.uber.com/us/en/blog/from-static-rate-limiting-to-intelligent-load-management/", + "ok": true, + "error": null + }, + { + "idx": 7, + "url": "https://www.flyingbarron.com/2026/04/voyager-and-art-of-graceful-degradation.html", + "ok": true, + "error": null + } +] \ No newline at end of file diff --git a/sreweekly/articles/533/01-incidents-start-before-the-response-does.html b/sreweekly/articles/533/01-incidents-start-before-the-response-does.html new file mode 100644 index 00000000..25323e68 --- /dev/null +++ b/sreweekly/articles/533/01-incidents-start-before-the-response-does.html @@ -0,0 +1,563 @@ + + + + + + + + + + Incidents start before the response does | Brent Chapman + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +
      + + + +
      + +
      +
      + +
      +
      +
      +
      +
      + + +
      + +

      Your company has probably invested significantly in what happens after an incident is identified: incident response tooling, trained incident commanders, communication protocols, on-call rotations. That investment matters. But what about the gap between when a problem starts and when anyone on your team knows about it?

      + + + +

      During that gap, customer damage is accumulating. The problem is getting worse, the blast radius is expanding, and nobody on the team is doing anything about it because nobody knows yet.

      + + + +

      You can’t eliminate this gap entirely, but you can shrink it. Four investments make the biggest difference.

      + + + +

      Broaden your detection surface

      + + + +

      Automated monitoring is the first and best line of defense, but it can only catch the failure modes someone thought to check for. Human detection isn’t a gap you can eliminate; it’s a permanent and valuable part of your detection capability.

      + + + +

      This means your customer support team is part of your detection infrastructure, whether or not you’ve told them so. So is any part of your company that interacts with customers regularly: account execs, customer success managers, even your social media team. They talk to your customers every day and often see concerns emerge before engineering does. And don’t overlook your customers themselves, who won’t limit their reports to your “official” support channels. If all these folks don’t have clear, fast escalation paths to flag potential problems for engineering, you have a detection gap that no amount of monitoring investment will close.

      + + + +

      If your company is a heavy user of its own product, the detection surface extends even further. When I led Slack’s incident management program, literally anyone in the company might notice a problem while using Slack internally. Not every company is in that position (it depends entirely on what the product is), but those who are should take advantage of it. Make sure everyone (all the way down to the part-time security guard covering the front desk on weekends) knows how to report problems they see.

      + + + +

      And watch for indirect signals. One of Slack’s best harbingers of “something is broken, even if we don’t know what yet” was the page-view rate on our public status page. If it started surging upward, we knew that something was wrong, even if we weren’t getting any other clear signals yet, and we’d start investigating. It was like smelling a light waft of smoke, well before the smoke detectors and fire alarms go off. If you have a public status page, consider adding its traffic patterns to your monitoring. A sudden spike in visits is a low-cost early warning powered by the collective behavior of your user base.

      + + + +

      Lower barriers to reporting

      + + + +

      Most of these detection channels depend on someone raising a concern, and that only works if the barrier to doing so is low. At many companies, the only mechanism for raising an alarm is to declare an incident, which triggers a full coordinated response: pages go out, a channel is created, an incident commander is assigned, people drop what they’re doing.

      + + + +

      That’s appropriate when you know you have a real problem. But if the only way to raise a concern is to trigger that entire response, people will hesitate, and rightfully so. Nobody wants to be the person who launched a full incident response over a hunch that turns out to be wrong. So they wait for more evidence, and the detection gap grows.

      + + + +

      Think of it like calling emergency services. When you call 911 (or 999, 000, 112, or whatever your country’s emergency number is), you don’t have to know whether you need an ambulance, a fire engine, a hazmat team, or a bomb squad. You describe what you see, and a trained dispatcher determines how serious the situation is, what sort of response is warranted, and who to send.

      + + + +

      Your incident detection should work the same way: make it easy for anyone to say “I think something might be wrong,” and let someone with training, experience, and context determine what response is warranted. At Slack, introducing a lightweight mechanism for exactly this was one of the most impactful things we did.

      + + + +

      Continuously right-size your alerting

      + + + +

      It’s tempting to close the detection gap by making your monitoring more aggressive: lower the thresholds, add more alerts, page on anything that twitches. This can backfire badly. Every alert that wakes someone at 3 AM and turns out to be nothing makes it a little more tempting for your on-call engineers to dismiss the next one. Alert fatigue is one of the most insidious threats to detection, precisely because it accumulates gradually. Your alerting system doesn’t fail all at once; it erodes, one false alarm at a time, until the real alerts get lost in the noise.

      + + + +

      The discipline runs in both directions: yes, add monitoring when you discover gaps, but regularly prune alerts that aren’t earning their keep. If a service-owning team can’t get through a review of every alert they received in the past week in a reasonable portion of a weekly ops review meeting, they’re getting too many alerts.

      + + + +

      Examine the gap

      + + + +

      Another way to shrink the detection gap over time is to examine it after every incident. You’re never going to be able to fully automate detection, but it’s still an ideal worth pursuing. Three questions, asked consistently in every post-incident review, create a steady stream of improvements:

      + + + +
        +
      • How long was the gap between when the problem started and when we detected it?
      • + + + +
      • Could we have detected it sooner?
      • + + + +
      • What monitoring would we need to add, or what threshold would we need to adjust, to catch this kind of problem faster next time?
      • +
      + + + +

      The bottom line

      + + + +

      Investing in detection is investing in the foundation of your entire incident management capability. You can have well-trained incident commanders, practiced responders, and polished communication protocols, but none of it matters until you know there’s a problem.

      +
      + +
      + +
      + + +
      +
      +
      + + +
      + + + + +
      +
      + + +
      + + + + + + + + + + + + + + diff --git a/sreweekly/articles/533/02-quick-thoughts-on-azure-regional-outage-from-july-23-26.html b/sreweekly/articles/533/02-quick-thoughts-on-azure-regional-outage-from-july-23-26.html new file mode 100644 index 00000000..18ddf8dd --- /dev/null +++ b/sreweekly/articles/533/02-quick-thoughts-on-azure-regional-outage-from-july-23-26.html @@ -0,0 +1,983 @@ + + + + + + + +Quick thoughts on Azure Regional Outage from July 23, ’26 – Surfing Complexity + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +
      + + +
      + +
      + + + + +
      +
      + +
      +
      + + + +
      +
      +

      Quick thoughts on Azure Regional Outage from July 23, ’26

      + +
      + +

      The folks at Microsoft Azure recently wrote up a post incident review for a networking issue in their West U.S region. From the included timeline, it looks like the impact was on the order of five hours. It’s a pretty short write-up, but let’s take a look at the contributors.

      + + + +
      +

      On 23 July 2026, a break-fix repair was initiated on an optical device to address a network reliability risk.

      +
      + + + +

      The first contributor mentioned in the write-up was work that was done to repair a device in their networking stack. Here I can’t help but think of the first bullet in my conjecture on why reliable systems fail. They made a change to the system in order to fix an ongoing problem, and due to a set of circumstances, things got worse rather than better.

      + + + +
      +

      A defect in our blast radius analysis system incorrectly expanded the scope of the repair event to include all optical devices egressing a specific datacenter. 

      +
      + + + +

      The second contributor mentioned was a (presumably) latent defect in their system. Note the irony of the failure mode here: I suspect this blast radius analysis system usually contributes to reliability, but in this case it hurt reliability by increasing the blast radius.

      + + + +
      +

      The safety validation step, which is designed to confirm that at least one of the two redundant datacenter paths remains available, ran but incorrectly concluded the operation was safe.

      +
      + + + +

      The third contributor mentioned was a safety check (good!) that passed even though the action was unsafe (bad!).

      + + + +
      +

      The checks validated each device individually rather than evaluating the aggregate effect of isolating all devices at once, a scenario that was not accounted for because the system was never designed to process a full datacenter’s worth of devices in a single request.

      +
      + + + +

      The reason it failed was due to an interaction with the second contributor: the blast radius being all of the optical devices egressing the datacenter. The designers never envisioned that the check would have to handle the sort of scenario that occurred as a result of the blast radius analysis system defect.

      + + + +
      +

      As a result, routes were withdrawn from multiple devices simultaneously, disrupting connectivity between the datacenter and the WAN – therefore impacting traffic entering or leaving the West US region.

      +
      + + + +

      It sounds like this change effectively disconnected the West US datacenter from the internet.

      + + + +
      +

      Once the route withdrawals took effect at 14:44 UTC, physical links and routing adjacencies continued to appear healthy, which initially masked the correlation between the break-fix activity and the connectivity disruption

      +
      + + + +

      Here we have our fourth contributor: the operators were receiving misleading signals from the system. The links and routes looked healthy, even though connectivity was broken.

      + + + +
      +

      The impact presented as a WAN routing anomaly, as third-party networks could not reach Azure in the region, rather than as a datacenter connectivity failure.

      +
      + + + +

      Our fifth contributor is another flavor of misleading signals. The symptoms presented as a routing issue between Azure and third-parties.

      + + + +
      +

      Although all physical work in the region was stopped, our engineers could not correlate to this recent change because the preparation activities in advance of the break-fix did not succeed, so the physical layer and traffic appeared healthy.

      +
      + + + +

      This is the sixth contributor mentioned in the writeup. The writing is a little oblique here, but I think what they are saying is that the repair event did not show up in their event log because the repair event didn’t actually complete. It sounds like the preparation activities were the ones that triggered the incident. But, because the repair event didn’t actually happen, the operators looking for events that correlate in time with the onset of the incident didn’t see the triggering event because it didn’t show up in the log of events. That’s my best guess, anyways.

      + + + +
      +

      Our automated recovery and rollback system detected the device failures, and attempted multiple retries to restore the affected devices. However, because that system depended on the same datacenter connectivity that had been disrupted, its automated rollback attempts were unsuccessful.

      +
      + + + +

      This is the seventh and final contributor mentioned. Azure has an automated recovery and rollback system (good!), but the failure mode in this case prevented automated rollback from succeeding (bad!).

      + + + +

      As always, I’d love to know more about how the operators identified what the failure mode actually was, and how they traced it back to the optical device repair work.

      +
      + + + + +
      + + + + +
      + + +

      + One thought on “Quick thoughts on Azure Regional Outage from July 23, ’26”

      + + +
        +
      1. + +
      2. +
      + + + + +
      +

      Leave a comment

      + + +
      + + + +

      +

      + +
      + + +
      +
      + + + +
      + + +
      +
      + + + + + + + + + +
      +
      +
      +
      + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + diff --git a/sreweekly/articles/533/03-why-distributed-databases-fail-at-coordination-boundaries.html b/sreweekly/articles/533/03-why-distributed-databases-fail-at-coordination-boundaries.html new file mode 100644 index 00000000..7dcb2ba4 --- /dev/null +++ b/sreweekly/articles/533/03-why-distributed-databases-fail-at-coordination-boundaries.html @@ -0,0 +1,2912 @@ + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + Why Distributed Databases Fail at Coordination Boundaries + + + + + + + + + + + + +
      +
      +
      +
      + + + +
      + + +
      + + + + + + + + + + +
      +
      +
      + +
      +
      + +
      +
      +
      +

      Could your team report a vulnerability within 24 hours? Find out on September 23.

      +
      + + + +
      +
      +
      + +
      + +
      +
      +
      +
      +
      +
      +
      + + +
      +
      + + + + + +
      +
      +
      + + + +
      +
      +

      Why Distributed Databases Fail at Coordination Boundaries

      +
      + +
      +

      Failures in distributed systems emerge at interfaces where independent components exchange timing, ownership, and state information.

      +
      + +
      + By  + + + · + Analysis +
      +
      +
      + +
      + + +
      + + + Comment + + +
      + +
      +
      + Save +
      +
      + + + + + +
      +
      + + 1.0K Views +
      +
      +
      + + +
      + +
      + +
      +

      Distributed databases are often evaluated through familiar technical dimensions: replication factor, consistency model, partitioning strategy, throughput, latency, and recovery time. These characteristics matter, but they do not fully explain why systems that appear healthy at the component level still experience severe production failures.

      +

      In many cases, the storage engine is not the weakest part of the architecture. The failure occurs at a coordination boundary.

      +

      A coordination boundary is any point where independently operating components must agree on timing, ownership, ordering, configuration, or state. These boundaries appear between replicas, partitions, control planes, data planes, load balancers, clients, metadata services, and background maintenance processes. Each component may behave correctly according to its local rules while the overall system produces an incorrect or unstable result.

      +

      This is why distributed database incidents can be difficult to predict. The database may not fail because a server crashes or a disk becomes unavailable. It may fail because two healthy components temporarily disagree about who owns a partition, whether a node is available, or which version of configuration should be applied.

      +

      Local Correctness Does Not Guarantee System Correctness

      +

      Engineers naturally reason about software components individually. A node accepts requests, writes data, replicates changes, responds to health checks, and reports metrics. If each of those behaviors appears correct, the system is assumed to be healthy.

      +

      Distributed systems challenge that assumption.

      +

      A replica can be healthy but delayed. A coordinator can be available but operating with stale metadata. A load balancer can route traffic correctly according to its current configuration while that configuration no longer reflects the database topology. A client can retry a failed request according to policy while unintentionally amplifying load during a partial outage.

      +

      Each component is locally correct. Their interaction is not.

      +

      Consider a partition ownership transition. One node is being removed, replaced, or scaled down, and another node is taking responsibility for the affected data range. The outgoing node may believe it still owns the partition because it has not received the latest control-plane update. The incoming node may already begin accepting requests because it has received a newer version of the assignment.

      +

      For a brief period, both nodes may behave correctly according to the information available to them. The system, however, has entered an ambiguous ownership state.

      +

      That ambiguity can lead to duplicate processing, inconsistent writes, rejected requests, or unexpected latency. The problem does not exist entirely inside either node. It exists at the boundary where ownership information is exchanged and interpreted.

      +

      Time Is Often the Hidden Coordination Dependency

      +

      Many distributed database designs avoid relying on perfectly synchronized clocks. Even so, time remains embedded throughout the system.

      +

      Timeouts determine when a request is considered failed. Leases determine how long a node retains authority. Heartbeats influence failure detection. Retry intervals shape traffic behavior. Expiration policies determine when data should disappear. Background processes decide when to compact, replicate, repair, or rebalance information.

      +

      These mechanisms create coordination dependencies even when the architecture does not explicitly describe them that way.

      +

      For example, a client sends a write request and does not receive a response before its timeout. The client cannot immediately know whether the write failed, succeeded, or is still being processed. It retries the request through another route.

      +

      If the database supports idempotent request handling, the retry may be safe. If it does not, the same logical operation may be applied twice. The first server and the client both followed their expected behavior. The uncertainty appeared between them because completion and acknowledgment were separated by a network boundary.

      +

      This is a common distributed systems pattern. A timeout provides information about waiting, not about the final outcome of an operation.

      +

      Cloud architects should therefore treat every timeout as an ambiguity boundary. Timeout behavior must be designed together with idempotency, deduplication, retry limits, load shedding, and observability. Configuring a timeout without defining the system’s response to uncertainty simply moves the failure elsewhere.

      +

      Metadata Can Become More Critical Than Data

      +

      Database reliability discussions frequently focus on protecting stored records. Replication, backups, checksums, and repair mechanisms are designed to preserve data durability.

      +

      However, the metadata that describes how data should be accessed can be just as important.

      +

      Partition maps, routing tables, node membership, schema versions, configuration states, and feature capabilities determine how requests travel through the system. If this metadata becomes stale or inconsistent, the underlying data may remain fully intact while applications lose the ability to access it reliably.

      +

      This is particularly important in systems that separate the control plane from the data plane. The control plane decides how infrastructure should be configured. The data plane processes live requests using that configuration.

      +

      Separating these responsibilities improves scalability and operational isolation, but it introduces another coordination boundary. Configuration changes must move safely from the control plane to every affected data-plane component. During that transition, the system may contain multiple valid configuration versions at once.

      +

      The engineering question is not merely whether a configuration update can be delivered. It is whether old and new versions can coexist without violating system correctness.

      +

      Safe configuration rollout often requires versioning, backward compatibility, staged activation, and explicit rollback behavior. Without those protections, a harmless-looking control-plane update can produce a data-plane outage even when no database node has failed.

      +

      Load Balancing Can Amplify Database Instability

      +

      Load balancing is sometimes treated as an infrastructure layer outside the database itself. In practice, routing behavior directly influences distributed database reliability.

      +

      When a node slows down, a load balancer may reduce traffic to it. That appears beneficial, but the remaining traffic must go somewhere. Healthy nodes receive additional load, their latency increases, and health checks may begin failing. The load balancer then removes more nodes, increasing pressure on the smaller remaining pool.

      +

      This creates a feedback loop.

      +

      The database causes routing changes, and the routing changes make the database less stable. Neither system is necessarily defective. The failure emerges from their interaction.

      +

      Aggressive health checks, short timeout thresholds, synchronized retries, and immediate node removal can turn a minor performance issue into a broad outage. A more resilient design considers the rate of change, not only the current health signal.

      +

      Cloud architects should ask whether routing decisions become less reliable during overload. They should also examine whether the database and load-balancing layers use compatible definitions of health. A node capable of serving read traffic may be temporarily unsuitable for writes. A node completing recovery may be reachable but not ready for production load.

      +

      Binary healthy-or-unhealthy classifications often hide these operational differences.

      +

      Background Work Creates Coordination Pressure

      +

      Distributed databases perform significant work outside the direct request path. Replication, compaction, repair, rebalancing, expiration, backup, and cleanup processes compete for shared resources.

      +

      These operations are often independently scheduled, which creates additional coordination boundaries. A compaction process may increase disk activity while a rebalance consumes network bandwidth. A repair job may begin during a traffic peak. Expired records may accumulate faster than cleanup processes can remove them.

      +

      Each mechanism may operate within its configured limits, yet their combined effect can overwhelm the system.

      +

      Time-to-live functionality provides a useful example. Expiring a record appears to be a simple data operation, but at scale it affects storage layout, indexing, replication, read behavior, and cleanup scheduling. The system must determine when an item is logically expired, when it should stop appearing in reads, and when its physical storage can be reclaimed.

      +

      Those events may not occur simultaneously.

      +

      If expiration processing is poorly coordinated, large groups of records can become eligible for deletion at the same time, creating bursts of background work. The feature itself works correctly, but the interaction between expiration timing and resource consumption can destabilize the database.

      +

      The broader lesson is that operational features should be evaluated as distributed workflows, not isolated functions.

      +

      Designing for Boundary Failures

      +

      The most effective way to improve distributed database reliability is to identify coordination boundaries during architecture design.

      +

      For every boundary, engineers should define what information crosses it, how that information is versioned, how long it remains valid, and what happens when delivery is delayed or duplicated. They should also determine whether the receiving component can safely operate with stale information.

      +

      Observability should follow the same structure. Monitoring individual nodes is necessary, but it is not sufficient. Teams need visibility into ownership transitions, metadata propagation delays, retry amplification, routing changes, replication lag, and background-work queues.

      +

      These signals reveal disagreement between components before that disagreement becomes a complete outage.

      +

      Testing must also include transitional states. Steady-state benchmarks show how a system performs when ownership, routing, and configuration are stable. Production failures frequently occur while those conditions are changing.

      +

      Architects should test node replacement, delayed configuration propagation, partial network loss, rolling upgrades, uneven clock behavior, repeated retries, overloaded background workers, and conflicting health signals. These scenarios expose the boundaries where local assumptions stop matching global reality.

      +

      Reliability Lives Between Components

      +

      Distributed databases rarely fail in the clean, isolated ways described by component diagrams. They fail through timing gaps, stale metadata, ambiguous ownership, retry storms, incompatible health decisions, and overlapping maintenance activity.

      +

      The database node that appears responsible may only be the place where the problem becomes visible.

      +

      For cloud architects and engineers, the practical shift is to stop treating coordination as an implementation detail. Coordination is part of the system’s correctness model.

      +

      Storage engines protect data. Replication protects availability. Load balancing distributes work. Control planes manage change. None of these mechanisms can provide reliability independently.

      +

      Reliability emerges from how they coordinate, especially when information is delayed, incomplete, duplicated, or temporarily inconsistent.

      +

      That is where distributed databases are most likely to fail, and where architects should focus first.

      +
      + +
      + + +
      +

      Opinions expressed by DZone contributors are their own.

      +
      +
      +
      +
      + + + +
      + +
      +
      +
      + +
      + + + +
      +
      +
      +
      +
      +
      +
      + +
      + × + +
      + + +
      +
      + +
      +
      +
      +
      +
      + + +
      + + + + + +
      + +
      + + + + + + + + + + + + + + + + + + + + + + diff --git a/sreweekly/articles/533/04-the-record-says.html b/sreweekly/articles/533/04-the-record-says.html new file mode 100644 index 00000000..835f590e --- /dev/null +++ b/sreweekly/articles/533/04-the-record-says.html @@ -0,0 +1,484 @@ + + + + + + + + + + + + The Record Says - by Tim Irving - Zero Sev Zero + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +
      +

      Discussion about this post

      User's avatar

      No posts

      Ready for more?

        +
        + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + diff --git a/sreweekly/articles/533/05-what-sres-should-automate-and-never-automate-with-ai.html b/sreweekly/articles/533/05-what-sres-should-automate-and-never-automate-with-ai.html new file mode 100644 index 00000000..5c82f5d5 --- /dev/null +++ b/sreweekly/articles/533/05-what-sres-should-automate-and-never-automate-with-ai.html @@ -0,0 +1,157 @@ +
        569 reads

        What SREs Should Automate — and Never Automate — with AI

        by
        featured image - What SREs Should Automate — and Never Automate — with AI

        Five key takeaways:

        +
          +
        1. Automate based on impact and recoverability, not on whether the AI is technically capable of doing the task.
        2. +
        3. Alert triage, anomaly detection, incident summaries, capacity forecasting — these are the easy wins. Low risk, high value.
        4. +
        5. Production changes, incident command, security response, severity calls — keep a human's name on these. Always.
        6. +
        7. Reversibility and blast radius are better questions than "can the AI do this."
        8. +
        9. The goal isn't AI replacing engineers. It's AI clearing enough noise that engineers can actually think.
        10. +
        +
        +

        I've sat through the version of this conversation that sounds like a vendor pitch — AI triages everything, drafts your runbooks, predicts outages before they happen, and nobody gets paged at 2 a.m. anymore. I've also watched the other version happen in real time: an automated remediation script restarts the wrong service, confidently, at 11 p.m., and a 20-minute blip turns into a four-hour outage while everyone tries to figure out why the "fix" made things worse.

        +

        Both of those are real. AI is already inside SRE workflows whether or not anyone signed off on it — the question that actually matters is where it belongs, and where a human still needs to be the one holding the decision.

        +

        None of what follows comes from a whitepaper. It's from watching what breaks when teams move too fast with this stuff, and what quietly gets better when they don't.

        +

        Reversibility and blast radius

        +

        Here's the mental model I keep coming back to before automating anything: can you undo it, and how bad is it if you're wrong?

        +

        Restarting a pod — reversible, low stakes. Deleting a database backup — not reversible, at all. Scaling a service up is easy to walk back. Silencing an alert for six hours is technically reversible too, except the six hours where something real happened and nobody saw it isn't something you get back.

        +

        Blast radius is the other half of it, and it's not the same thing as severity. A misclassified low-priority alert costs a few wasted minutes. A misrouted sev-1 costs an hour of response time during an active outage, while the right team sits there not knowing they should be paged. And blast radius scales with what the action touches — one service versus a shared piece of infrastructure everything depends on, even when both look equally "minor" on paper.

        +

        Anything with low reversibility and a wide blast radius shouldn't be running on autopilot. Anything reversible and contained is fair game. The stuff in between is where you actually need judgment — specifically, judgment from the people who'll be the ones on call when it goes sideways.

        +

        Notice this framing never asks whether the AI can do something. It asks what happens if it's wrong. That's the more useful question, and it's the one most teams skip.

        +

        Where this actually works well

        +

        Alert noise. This is the least controversial win there is. Somewhere between 30 and 60% of production alerts are noise by the time a human sees them — duplicates, transients, things that resolved themselves three minutes ago. AI grouping related alerts, suppressing known-flapping signals, correlating spikes with recent deploys — worst case, something gets mislabeled and a human still catches it. Low blast radius, fully reversible. This is exactly the profile you want.

        +

        One catch: it only works well tuned to your environment, not a generic model. An alert that always fires right before a nightly batch job and clears itself a minute later is trivial to suppress — but only if the model actually knows about your batch schedule. Skip that step and you've just added a second layer of noise on top of the first.

        +

        First drafts of runbooks and postmortems. Runbook rot is one of the oldest problems in this field. The doc that was accurate in 2022 is a landmine now — nobody updates it, an incident hits, someone follows it anyway, and step four references a service that got decommissioned eight months ago. AI is genuinely good at pulling together a first draft from past incidents, change logs, whatever documentation exists. Same for postmortems — a draft that someone who actually lived through the incident reviews before it goes out saves real hours.

        +

        Forecasting and anomaly detection. This is pattern matching, and models are good at pattern matching. A holiday traffic spike that happens once a year gives engineers almost no reps to build intuition about — but a model trained across several years of that same spike has plenty. The important part: keep this as a recommendation a human acts on, not something that auto-provisions infrastructure on its own. The moment it stops informing a decision and starts making one, the blast radius changes.

        +

        Narrow, well-understood auto-remediation. This one comes with real caveats, but it earns its place. A specific service that needs a restart when it hits a known stuck state, a queue that needs draining past a defined threshold — fine, if the failure class is precisely defined, tested, and low-blast-radius by design. And there has to be a circuit breaker. If the fix doesn't work within a set window, it stops and escalates instead of retrying forever on a wrong diagnosis. Automation that keeps trying the same broken fix is worse than doing nothing.

        +

        Where it doesn't belong

        +

        Severity calls. Get this wrong either direction and it costs you. A real sev-1 marked as low pulls in the wrong people at the wrong urgency while an SLA clock runs. A minor issue marked critical drags a response team into something that didn't need them at 3 a.m. AI can surface context and flag patterns worth escalating — but the actual call needs a name attached, someone accountable for it. "The model said it was low severity" doesn't hold up in a postmortem.

        +

        Production changes without sign-off. Config changes, scaling decisions, anything touching a database directly, restarts outside that narrow bounded case above — a human authorizes these. AI can prep the change, check it against known-good patterns, even simulate the blast radius. What it shouldn't do is decide the moment is right and pull the trigger itself.

        +

        Security incidents. Different risk shape entirely. Miss something real and an active compromise sits there while the system waits for more confirmation. False-positive and you've locked out legitimate engineers mid-response. AI correlating logs to surface signal fast — genuinely useful. Containment and escalation decisions — that needs someone who can weigh legal and business context a model was never trained on.

        +

        Root cause, as a stated fact. AI narrowing the search space by correlating deploy timing with metric shifts is useful groundwork. But writing "root cause: X" in a postmortem is a claim that shapes what the org fixes next and what it decides to ignore. Get that wrong because a correlation looked convincing, and the actual bug ships again next quarter.

        +

        Who to escalate to. This is context a model just doesn't have — who's already underwater tonight, what else is on fire across the org, whether the responding engineer's confidence is real or performed. Escalation is a trust call as much as a technical one.

        +

        The thing nobody's measuring

        +

        There's a slower cost that never shows up in a single incident review: engineers stop building intuition when AI absorbs all the routine reps. The edge cases are exactly where judgment matters most — and they're exactly the cases you need practice on the boring stuff to be ready for. A team leaning hard on automation can look great for a long stretch, right up until something shows up that doesn't match anything the model — or the team — has seen before.

        +

        This isn't an argument against automating things. It's an argument for being honest about which reps you're willing to give away.

        +

        A few practices worth adopting

        +
          +
        • Decide, as a team, which categories of action AI can take alone versus which need a sign-off — decide this before an incident forces the question at 2 a.m.
        • +
        • Keep an actual human accountable for anything irreversible. Not nominally "in the loop" — actually reviewing before it executes.
        • +
        • Build in a circuit breaker for anything automated. If it doesn't work within a defined window, it escalates instead of retrying.
        • +
        • Rotate people through the routine cases sometimes, even when AI could handle it, so the skill doesn't quietly disappear.
        • +
        • Revisit the boundary as systems change. A failure class that was well-understood six months ago might not be anymore after an architecture shift.
        • +
        +

        Skip this and you end up with automation debt, eroded skills, and a production system nobody fully understands anymore — which is a worse place to be than where you started.

        +

        Where this leaves things

        +

        It's not really a question of whether to use AI. It's whether you're using it somewhere judgment genuinely isn't needed, or somewhere it is and you've just decided waiting for a human is too slow.

        +

        One of those is a real force multiplier. The other is a liability with a delay timer on it.

        About Author

        Sai Joshitha Kathari HackerNoon profile picture
        Senior Site Reliability Engineer at

        Senior Site Reliability Engineer based in Austin, Texas, specializing in distributed systems, Kubernetes, and production reliability.

        TOPICS

        \ No newline at end of file diff --git a/sreweekly/articles/533/06-20-the-ci-traffic-without-getting-slower-how-we-rebuilt-git-serving-at.html b/sreweekly/articles/533/06-20-the-ci-traffic-without-getting-slower-how-we-rebuilt-git-serving-at.html new file mode 100644 index 00000000..0471efc8 --- /dev/null +++ b/sreweekly/articles/533/06-20-the-ci-traffic-without-getting-slower-how-we-rebuilt-git-serving-at.html @@ -0,0 +1,656 @@ + 20× the CI traffic without getting slower: How we rebuilt Git serving at Datadog | Datadog + + +

        Get Started with Datadog

        Engineering

        20× the CI traffic without getting slower: How we rebuilt Git serving at Datadog

        Published

        Read time

        15m

        20× the CI traffic without getting slower: How we rebuilt Git serving at Datadog
        Mike Thompson

        Mike Thompson

        Senior Staff Engineer

        Daniel Esponda

        Daniel Esponda

        Staff Engineer

        If you have ever watched a CI job sit on “Fetching repository …” while nothing seems to happen, you already know the unglamorous truth about continuous integration: Every job begins by getting the code, and getting the code is not free.

        +

        At Datadog, CI fetches code millions of times a week across thousands of repositories. Our largest repositories are monorepos with years of history and hundreds of thousands of files. At that scale, git clone stops being a footnote and becomes a large contributor to CI run times.

        +

        This is the story of gitretriever, the Git mirror we built to serve code to CI at Datadog scale. In its first 4 months, gitretriever served more than a billion Git requests and hundreds of terabytes of code. Today gitretriever handles more than 100 million requests each week. Despite the 20× traffic growth since launch, median latency has remained around 40 ms, while fetch-serving CPU on our previous Git backend has dropped by three to four times.

        +

        Serving Git to CI, and why it gets hard

        +

        Datadog has a unique CI setup: GitHub serves as the authoritative code repository, while almost all of our internal CI workloads run on a self-hosted GitLab installation. CI fetches from GitLab’s Gitaly, fronted by Praefect (Gitaly Cluster’s routing and replication manager) and kept in sync with GitHub by an internal service (aptly named “codesync”). This hybrid architecture has carried us through more than a decade of growth.

        +

        But CI load does not grow smoothly. The expanding use of AI coding agents has driven an order-of-magnitude increase in Git traffic, with agents hitting Git far harder and more often than even our most active contributors ever could. That traffic comes on top of the continually growing load from internal deployment, auditing, and security services. As that growth accelerated, the pressure hit hardest where our code is densest: our large monorepos. Operational load increased, CI run times grew, and multi-hour long CI outages became more frequent. It was clear we needed a more sustainable solution.

        +

        Why the usual fixes don’t scale 

        +

        We tried adding capacity, we tried increasing instance size, we tried placing different repositories on dedicated backends, and we tried optimizing build pipelines. Things would improve for a week or two, but then our CI infrastructure would inevitably end up degraded or outright down. So why didn’t any of the usual approaches work? 

        +

        A single fetch from a large monorepo can consume several seconds of server CPU. At peak, hundreds of jobs perform fetches at the same moment and land on the same handful of nodes. Adding capacity did little to reduce per-node CPU usage. In some cases, adding more nodes made the problem worse.

        +

        Before committing to a new architecture, we had to figure out why none of our previous attempts at fixing the problem had worked:

        +
        • Scale the backend or add nodes: In our replicated setup, every write had to be copied to every replica. Adding a node increased replication overhead instead of relieving it.

        • Put a content delivery network (CDN) or caching proxy in front: The expensive part of a fetch isn’t a static byte range you can cache at the edge. It’s computation that’s specific to each client’s request.

        • Clone on demand from GitHub: That simply moves the thundering herd upstream, where we run into server-side rate limits.

        +

        The common thread was that we had been scaling the wrong axis. Read traffic scales with the number of CI jobs, but in our replicated architecture, write costs scale with the number of replicas. Every time we added replicas to handle more reads, we also increased replication overhead, and more CPU time went to maintaining the system instead of serving fetches.

        +

        To understand why serving those fetches consumed so much CPU in the first place, it helps to look at what happens during a Git fetch.

        +

        Why a Git fetch is expensive

        +

        To understand why our design works, it helps to understand how Git stores data and where a git fetch spends its time.

        +

        Git data types

        +

        Git’s object database is primarily built around immutable objects. For the purposes of this post, we’ll focus on three: 

        +
        • Blobs, which store file contents 

        • Trees, which describe directory entries (for example, folders and blobs) 

        • Commits, which store metadata, a commit message, a reference to a tree, and references to parent commits

        +

        Each object is identified by a hash of its type, size, and contents. SHA-1 remains the default object format, although Git also supports SHA-256 repositories. 

        +

        Finally, there are references, which are mutable names stored separately from objects. For example, refs/heads/main identifies the commit at the tip of the main branch.

        +

        Objects may be stored on disk individually as loose objects or grouped into packfiles. Within a packfile, an object may be stored in full or as a delta against another object (known as delta compression), which allows Git to efficiently store the complete history of changes to files within a repository. Packfiles are immutable to allow for safe concurrent reads.

        +
        How Git references, commits, trees, blobs, and packfiles relate to one another.
        Figure 1: References point to commits, which reference trees and parent commits. Trees reference blobs and other trees. Packfiles store Git objects independently of references.
        How Git references, commits, trees, blobs, and packfiles relate to one another.
        Figure 1: References point to commits, which reference trees and parent commits. Trees reference blobs and other trees. Packfiles store Git objects independently of references.
        +

        Write operations (for example, git push) may introduce new packfiles. A background maintenance process periodically consolidates loose objects and smaller packfiles into new packfiles. Unreachable objects (for example, deleted files) are eventually removed after a certain threshold by being omitted during packfile consolidation.

        +

        Git protocol v2 

        +

        Now that we understand Git’s data types, we can briefly look at how the current (v2) Git protocol works.

        +

        The Git client uses the ls-refs command to learn the current object IDs of references it cares about (for example, all branches). The client and server then begin a multi-round negotiation to determine which objects the server needs to send to the client. You can read more about this negotiation process in the Git protocol v2 documentation. 

        +

        Once the client and server have determined which objects to send, the server creates a packfile containing those objects and sends it to the client.

        +

        Constructing the response packfile can be CPU and I/O-intensive. The server locates objects within packfiles by using an index that Git maintains for each packfile. Some objects can be copied as is into the response packfile, while others must be decompressed and recompressed using delta compression. Under a sufficiently large number of concurrent fetches, this packfile construction work can saturate server CPU and storage capacity. 

        +

        Client behavior, such as requesting weeks’ worth of changes to a large monorepo, can make this more expensive in both CPU and I/O operations. Git attempts to reduce this cost with reachability bitmaps, sparse traversal, multi-pack indexes, and pack reuse. We tried all of these options, but client behavior and the rate at which our monorepos changed still concentrated CPU load on a small number of servers.

        +

        The final step of a fetch or pull from a Git server is for the client to read the received packfile and update its local index of available objects. This requires only a small amount of client-side CPU.

        +

        Our approach: Many independent mirrors, kept fresh

        +

        If the problem is CPU concentrated on a few contended nodes, the solution is to stop concentrating it.

        +

        Gitretriever runs independent pods, each of which maintains a fresh local copy of the repositories it serves without waiting for every node to reach consistency. Each pod serves its local copy directly, with no consensus and no multi-writer replication between peers. Gitretriever pods have two roles, as shown in the following diagram: 

        +
        • Mirrors stay in sync with GitHub. We deliberately keep this fleet small because its job is to be a good GitHub client: a handful of well-behaved pollers rather than thousands of them. 

        • Relays fan out reads to CI jobs. This fleet is larger and autoscaled based on CPU and network load, allowing us to provision enough read capacity to meet demand without turning that growth into additional load on GitHub.

        +
        Git traffic flowing through mirrors and relays between GitHub and CI workloads.
        Figure 2: Mirrors synchronize repositories from GitHub, while relays distribute those repositories to CI jobs and other Git workloads.
        Git traffic flowing through mirrors and relays between GitHub and CI workloads.
        Figure 2: Mirrors synchronize repositories from GitHub, while relays distribute those repositories to CI jobs and other Git workloads.
        +

        Staying fresh and reducing CPU usage

        +

        The architecture works only if every mirror and relay stays close to the latest changes without recreating the CPU bottlenecks we were trying to eliminate. We designed gitretriever around three principles that keep repositories fresh while minimizing repeated work.

        +

        Distribute Git pulls across branches

        +

        Gitretriever mirrors continually poll the upstream in a tight loop for changes. Gitretriever performs a parallel fetch for each reference it detects as changed since the previous synchronization loop iteration. No single request concentrates an expensive delta compression job on GitHub, and each small pack requires far less indexing CPU than one monolithic monorepo pack. Staying close to the tip of each branch also means that, in any given synchronization loop iteration, only a small number of branches have changed, reducing the number of packfiles we need to fetch.

        +

        Spend the sync work once, then reuse it 

        +

        For the busiest repositories, one mirror cannot serve every client, so changes fan out to a fleet of relays. Relays can connect to mirrors or to other relays. Each relay splits its upstream connection into two channels:

        +
        • A signaling gRPC stream: Announces that a pack is ready, propagates reference updates, and communicates mirror and relay topology changes

        • A plain HTTP endpoint: Serves the pack bytes themselves 

        +

        Because Git objects are content-addressed, a relay installs the packfile it receives from its upstream mirror or relay without regenerating, re-indexing, or re-verifying it. It drops the packfile and its index into place, trusting the objects inside by the hashes that identify them. The work of pulling and indexing from GitHub happens once on the mirror, and every relay reuses that work instead of fetching again. As a result, the relay fleet can grow without adding load on GitHub while remaining within single-digit milliseconds of the tip.

        +

        Never build the same pack twice

        +

        Gitretriever is both a Git client and a Git server. The current implementation uses Git’s default backend storage format: packfiles, reference tables, reachability bitmaps, and multi-pack indexes. That means gitretriever has to make serving other Git clients (such as CI jobs) as efficient as possible.

        +

        A fresh push to a busy branch sets off a thundering herd of identical fetches. Gitretriever implements a pack cache, allowing it to reuse previously assembled packfiles for identical client requests. About half of all pack-building fetches are served directly from the cache, skipping the delta compression calculation on mirrors and relays entirely. Cache misses are still served locally by the mirrors and relays, so even a cache miss never becomes a trip to GitHub.

        +

        Underneath these are smaller refinements, including a readiness check that understands Git state and keeps a pod out of rotation until its pack count is healthy, along with background repacking that keeps the packfile count under control while the pod continues serving. But the theme never changes: Take the CPU that used to pile up in one place and either spread it out or stop repeating it.

        +

        Future iterations of gitretriever will build on the relay replication protocol to keep an always-up-to-date copy of our large repositories directly on CI nodes, allowing jobs to skip the initial git clone altogether.

        +

        The bigger surprise: Many use cases don’t need a clone

        +

        Once every repository had a fresh mirror, something in the traffic caught our eye: Most non-CI workloads don’t need a full repository clone. They wanted a single file at a commit, the SHA a branch pointed to, the list of files that changed, or the merge base of two refs. Cloning an entire repository to answer one of those questions was enormous overkill, yet our internal services, developer tools, and AI agents were doing it constantly.

        +

        So we added a small, read-only HTTP API for exactly those queries. Resolving a ref or reading a file takes single-digit to tens of milliseconds. By comparison, a shallow clone of a large monorepo takes on the order of 75 seconds and keeps a CPU core busy for most of that time. Moving these use cases to the API reduces latency and removes load from the entire system. 

        +

        The non-CI workloads changed how we think about gitretriever. It’s less a faster Git server and more the query layer for Git across our engineering systems.

        +

        This is the direction the platform is heading. As workflows become more automated and more AI agents ask questions about code, the cheapest and fastest answer is often another API rather than handing out a repository clone.

        +

        Rolling out gitretriever safely

        +

        Rolling out gitretriever required careful planning. Our CI infrastructure is used by every engineer at Datadog, so one wrong move could bring engineering to a halt. We used feature flags and built in automatic fallback to the old backend into our CI jobs, so if a mirror became unreachable or a fetch failed, the job fell back to the previous path. The worst-case outcome was no worse than before. We then migrated one repository group at a time, starting with the largest monorepo, while watching the old backend’s CPU graph.

        +

        When that first monorepo cut over, we saw an immediate step decrease in CPU usage. That confirmed our understanding of the problem: Gitretriever was absorbing the heaviest, most CPU-dense fetches first. Those were the same ones that had been degrading developer experience and driving outages.

        +

        The metrics matched our expectations:

        +
        • Synchronization time dropped from several seconds to a few hundred milliseconds, making continuous, coordination-free mirroring possible.

        • To date, gitretriever has served more than a billion Git requests and hundreds of terabytes of data across roughly 5,500 repositories, and now handles more than 100 million requests each week.

        • Traffic grew about 20× in 4 months while median serve latency remained around 40 ms (Figure 3). The system became an order of magnitude busier without getting materially slower.

        • The result we care about most: Moving CI fetch traffic to gitretriever reduced the old backend’s fetch-serving CPU by three to four times, even as overall CI activity kept climbing (Figure 4). Its memory footprint dropped in step, which later let us right-size that backend down. The old backend still handles some use cases that gitretriever doesn’t yet support (e.g., rendering the GitLab UI), so we don’t claim we replaced it (yet). But the fetch-path load it had been drowning under is gone.

        +
        Traffic rising substantially from March to July while median latency remains nearly flat.
        Figure 3: Serve volume and median latency, each indexed to launch. Traffic grew about 20× while median latency remained around 40 ms.
        Traffic rising substantially from March to July while median latency remains nearly flat.
        Figure 3: Serve volume and median latency, each indexed to launch. Traffic grew about 20× while median latency remained around 40 ms.
        +
        Fetch-serving CPU dropping sharply during the rollout and remaining substantially lower.
        Figure 4: The old backend’s fetch-serving CPU during the rollout, stepping down as each repository group migrated to gitretriever.
        Fetch-serving CPU dropping sharply during the rollout and remaining substantially lower.
        Figure 4: The old backend’s fetch-serving CPU during the rollout, stepping down as each repository group migrated to gitretriever.
        +

        How we built it: Two engineers, Claude Code, and design doc in nearly every folder

        +

        We chose to use Claude Code on this project to accelerate development and to explore how far AI could responsibly assist with building production infrastructure. What made an AI collaborator trustworthy on a system this central wasn’t the model; it was the discipline around how we used it.

        +

        We planned before we wrote code, designing each change and iterating on the design through several rounds before committing a line of code. We validated every change with integration tests backed by real metrics and logs, not just unit tests, so the bar for “done” was observed behavior rather than a green checkmark. To keep both the AI and ourselves aligned across a dozen packages, we maintained a living design document in nearly every directory, describing its architecture, data flow, concurrency model, and configuration, and updating it alongside the code.

        +

        Those documents ended up serving two purposes. During development, they kept AI-generated changes aligned with the architecture. When ownership of the service transitioned to the team that now maintains it, the same documents became the handoff. 

        +

        The lesson we would pass on is that the design documents became the interface between the engineers, the AI, and the next team. Ultimately, the quality of your tests and telemetry data sets the ceiling on how far you can trust an AI collaborator.

        +

        What’s next

        +

        Gitretriever is not finished. We’re expanding the query API so more workloads can skip cloning entirely, allowing us to fully decommission our old Git backend. We’re also continuing the rollout across the rest of our repositories and building for a future where automated and agent-driven workflows ask even more of Git.

        +

        A few ideas we’ll carry into whatever comes next:

        +
        • Make it disposable so you do not have to make it durable. Some of the hardest parts became much simpler once we made them rebuildable instead of authoritative.

        • Content addressing lets you trust data by name. That’s what makes coordination-free replication safe.

        • The fastest fetch is the one that transfers nothing, whether that’s a fast-path ref update or an API call that answers the real question without a clone.

        +

        More than any single optimization, gitretriever reflects how we approach engineering at Datadog: Push a good system as far as it will go, then, when the scale curve demands it, design the next generation from a better understanding of the problem, validate it against real telemetry data, and write down what you learned so the next team can build on it.

        +

        If this sounds like your kind of problem, we would love to work with you. Take a look at our open roles.

        Start monitoring your metrics in minutes

        \ No newline at end of file diff --git a/sreweekly/articles/533/08-there-is-more-to-code-review-than-automatable-detection.html b/sreweekly/articles/533/08-there-is-more-to-code-review-than-automatable-detection.html new file mode 100644 index 00000000..9143f8db --- /dev/null +++ b/sreweekly/articles/533/08-there-is-more-to-code-review-than-automatable-detection.html @@ -0,0 +1,910 @@ + + + + + + + There is more to code review than (automatable) detection + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +
        + +
        +
        + + +
        + + +
        + +
        +
        + + + +
        + + + +
        + + +

        There is more to code review than (automatable) detection

        + +
        + + + +
        + + + +

        The abstract for article “The End of Code Review: Coding Agents Supersede Human Inspection” paints this picture for the reader…

        + + + +
        +

        Abstract – Code review has been the primary quality gate in software development since Fagan formalised code inspection in 1976. For five decades, having a human examine and comment on a colleague’s changes before merge has been a cornerstone practice at organisations of every size. Coding agents are large language model (LLM)-based autonomous systems capable of reading, writing, testing, and repairing software. We argue that coding agents have crossed a threshold of capability at which traditional human code review is no longer a necessary component of a software quality pipeline. Our argument rests on two claims: every stated goal of code review can be served by agents at lower cost and higher throughput; the naive integration in which agents write code and humans remain the mandatory reviewers is a dead end because it neither provides meaningful assurance nor scales with AI-assisted throughput.

        +
        + + + +

        The article is structured well and quite straightforward for engineers who aren’t used to reading research articles very often. However, I do think the argument critically depends on a problematic framing: the substitution myth.

        + + + +

        The author decomposes peer code review into four stated functions: defect detection, style enforcement, knowledge transfer, and awareness. It argues an agent can perform each one. The conclusion, of course, is that if an agent can execute each of those functions, then the agent has the capability to replace a human reviewer. 

        + + + +

        I think this overlooks some important aspects of peer code review that cannot be reduced to a function:

        + + + +

        A peer reviewer’s confusion

        + + + +

        When an experienced engineer reads a diff and says “I don’t understand this.”, their confusion is the finding. It means the code is either too complex, the abstraction is wrong, or the intent is not clear. An LLM will always ‘understand’ the code in the sense of being able to process it. It can’t give you the signal of legitimate human incomprehension. The article treats comprehensibility as something that is more about style than anything else. It’s not. It’s an emergent property and it shows up in the interaction between a person attempting to understand the artifact and the artifact itself.

        + + + +

        Qualified skepticism about whether the change is even necessary

        + + + +

        Questioning the existence of a change, like:

        + + + + + + + +
        +

        “Should this actually be two PRs?”

        +
        + + + +

        or

        + + + +
        +

        “This solves the symptom, not the problem”

        +
        + + + + + + + +

        These are questions about intent, scope, and appropriateness of the change. All of that comes before whether the code is “correct.” The article’s framing assumes that a) the code change being reviewed is necessary, and b) the main purpose of the review is verification.

        But anybody who has ever had contact with production understands that code review is often the last (or sometimes only) moment when someone can be expected to challenge whether the change is even necessary.

        + + + +

        The ability to see what is not there

        + + + +

        A human reviewer can notice that an API contract has changed but the error handling didn’t. They can notice what is missing. In other words: being able to recognize what is expected to be present, but isn’t. The article doesn’t acknowledge this at all, which is particularly interesting, given that absence blindness is exactly the class of failure that LLMs tend to be quite poor at.

        The agent reviews what is there; engineers with expertise can easily notice what’s missing.

        + + + +

        Who wrote the code influences the scrutiny of the review

        + + + +

        Peer code reviewers have a sort of calibrated attention that comes from past experience with the code’s author, who is often a colleague. For example: a less-tenured engineer’s first commit to, say, a payments module will likely get different attention than a veteran and ‘grey beard’ engineer’s routine refactoring.

        Reviewers typically match the situation’s who, what, when, and where to their own experience of where risk lies.
        The article seems to treat all diffs as equivalent inputs.

        + + + +

        Code review is bidirectional and constructive

        + + + +

        It seems to me that paper reduces knowledge transfer down to just information delivery; the agent simply ‘generates explanations.’ But discussion in a code review is a joint cognitive activity. The peer reviewer learns about the author’s approach, the author learns via the reviewers’ questions, and the result is a shared understanding that neither party had prior to the discussion.

        + + + +

        This is coactive work, not simply a transmission. An agent’s summary isn’t a substitute for a conversation that changes both participants’ mental models.

        + + + +

        Operational context that lives outside repos

        + + + +
        +

        “We just had an incident in this service last Tuesday.”

        + + + +

        “The team that owns this downstream consumer is about to deprecate that interface.” 

        + + + +

        “Legal told us not to log this field anymore.” 

        +
        + + + +

        Human reviewers possess so much more contextual knowledge than they’re aware of, even though they can recognize connections in the wild. People understand the current state of the organization, recent events, and informal agreements that aren’t captured in tests, docs or version control, and they can recognize how these may influence the code under review. This happens so often that it’s all but invisible.

        The article assumes the codebase is the complete context. It never is.

        + + + +

        Accountability for the code isn’t just a beuraucratic formality

        + + + +

        The article treats human responsibility as a compliance artifact, a “named human” for legal or other rule-related purposes. But being aware that you are personally responsible for approving a change shapes how you review it. It is the “skin in the game.” An agent that “signs off” on a pull request bears no consequences and certainly has no incentive structure that fuels an earnest evaluation. While the paper does include ethics concerns in its discussion section, it ends up redirecting it to “requirements engineering and post-deployment monitoring” which seems to me as hand-waving way of kicking the can down the road.

        + + + +

        The most fundamental issue I have with the article is that it assumes code review is a first and foremost a detection process: you find defects, style violations, security issues, etc., and the assumption is that detecting these faster and cheaper is universally better. 

        + + + +

        But code review is also a coordination process, a sensemaking process, and a governance process. 

        + + + +

        The substitution myth often plays out in this same way:

        + + + +
          +
        1. First, decompose the human contribution of work into measurable functions.
        2. + + + +
        3. Show that the machine can replicate this human contribution into measurable functions of its own.
        4. + + + +
        5. Declare the human redundant.
        6. +
        + + + +

        This approach often falls apart at the same point: the human contribution that mattered most was the integration across functions. People’s ability to adapt to unplanned circumstances and contexts and serve the social accountability expected.

        + + + +

        This ability to adapt in those situations aren’t accounted for in the original decomposition step #1, above.
        I don’t think they were accounted for in the original article, either.

        + + + +

        + + + +
        +
        + + +
        + +
        + + +
        + + +
        +
        +
        + +
        +
        + + +
        + + + Scroll to Top +
        + + + + + + + + + + + + + + diff --git a/sreweekly/articles/533/index.json b/sreweekly/articles/533/index.json new file mode 100644 index 00000000..9ebd023f --- /dev/null +++ b/sreweekly/articles/533/index.json @@ -0,0 +1,50 @@ +[ + { + "idx": 1, + "url": "https://greatcircle.com/blog/2026/08/04/detection-gap/", + "ok": true, + "error": null + }, + { + "idx": 2, + "url": "https://surfingcomplexity.blog/2026/08/16/quick-thoughts-on-azure-regional-outage-from-july-23-26/", + "ok": true, + "error": null + }, + { + "idx": 3, + "url": "https://dzone.com/articles/distributed-databases-coordination", + "ok": true, + "error": null + }, + { + "idx": 4, + "url": "https://read.zerosevzero.com/p/the-record-says", + "ok": true, + "error": null + }, + { + "idx": 5, + "url": "https://hackernoon.com/what-sres-should-automate-and-never-automate-with-ai", + "ok": true, + "error": null + }, + { + "idx": 6, + "url": "https://www.datadoghq.com/blog/engineering/gitretriever/", + "ok": true, + "error": null + }, + { + "idx": 7, + "url": "https://netflixtechblog.com/a-tale-of-two-flink-autoscalers-e9f6a1b1492b", + "ok": false, + "error": "HTTP 403" + }, + { + "idx": 8, + "url": "https://www.adaptivecapacitylabs.com/2026/08/24/there-is-more-to-code-review-than-automatable-detection/", + "ok": true, + "error": null + } +] \ No newline at end of file diff --git a/sreweekly/html/528-2026-08-02.html b/sreweekly/html/528-2026-08-02.html new file mode 100644 index 00000000..67cbae5d --- /dev/null +++ b/sreweekly/html/528-2026-08-02.html @@ -0,0 +1,442 @@ + + + + + +SRE Weekly Issue #528 – SRE WEEKLY + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +
        + + + + + + + + +
        + +
        + + + + + + +
        + + +
        +

        SRE Weekly Issue #528

        +
        + + + +
        + + +

        + +
        +

        A message from our sponsor, Planetscale:

        +

        Most database incidents start with one expensive query, not the database being down. PlanetScale gives SRE teams high-availability Postgres and MySQL with automated failover, query insights, and Database Traffic Control to stop runaway queries before they page you.

        +

        → Explore PlanetScale

        +
        + + +
        +
        + +
        +

        Spotify has had some difficulty around podcast publishing, and they shared this analysis of the worst incident.

        +

          Jim Whitehead, Ulrik Mikaelsson, John Lagomarsino, and Saunak Jai Chakrabarti — Spotify

        +
        +
        + + + +
        + +
        +

        …and here’s where it gets interesting. This post shares the user point of view on the Spotify issues, including fact-checking their published timeline.

        +

          Gergely Orosz — The Pragmatic Engineer

        +
        +
        + + + +
        + +
        +

        New incident role unlocked: the incident tech lead. I enjoyed the description of the interplay between the tech lead and the incident commander.

        +

          Brent Chapman

        +
        +
        + + + +
        + +
        +

        This one goes hard: if you try to reduce your incident count, your system will become less reliable, not more. Aim for more incidents, handled well.

        +

          Tim Irving

        +
        +
        + + + +
        + +
        +

        There’s some brutal honesty in here that I find refreshing, especially around the impact on incidents and incident response.

        +

          Liz Fong-Jones — Honeycomb

        +
        +
        + + + +
        + +
        +
        +

        Here’s what I’ve learned about what actually happens inside a team when something breaks, how teams often reflect, and how we think about accountability without losing a blameless culture.

        +
        +

          Karan Nagarajowda — Uptime Labs

        +
        +
        + + + +
        + +
        +
        +

        agent sprawl is now a production reliability crisis, and the SRE discipline does not yet have governance frameworks for it.

        +
        +

           Ajay Devineni — DZone

        +
        +
        + + + +
        + +
        +

        The premise: read replicas can help you scale read load, but they introduce complexity. The article goes into the problems they ran into and how they dealt with them.

        +

          Johanna Larsson — incident.io

        +
        +
        +
        + + + + +
        + +
        + + + +
        + + +
        + + + + + + + + +
        + +
        + + + + +
        + + + + + + + + + + \ No newline at end of file diff --git a/sreweekly/manifest.json b/sreweekly/manifest.json index e4aa35f1..13e72409 100644 --- a/sreweekly/manifest.json +++ b/sreweekly/manifest.json @@ -12,7 +12,9 @@ "extracted": true, "article_count": 8, "markdown_dir": "markdown/533", - "extracted_at": "2026-09-10T07:40:50" + "extracted_at": "2026-09-10T20:30:06", + "articles_fetched": 7, + "articles_failed": 1 }, "532": { "id": "532", @@ -24,7 +26,9 @@ "extracted": true, "article_count": 7, "markdown_dir": "markdown/532", - "extracted_at": "2026-09-10T07:40:50" + "extracted_at": "2026-09-10T20:41:38", + "articles_fetched": 5, + "articles_failed": 2 }, "531": { "id": "531", @@ -36,7 +40,9 @@ "extracted": true, "article_count": 8, "markdown_dir": "markdown/531", - "extracted_at": "2026-09-10T07:40:50" + "extracted_at": "2026-09-10T20:42:12", + "articles_fetched": 8, + "articles_failed": 0 }, "530": { "id": "530", @@ -48,7 +54,9 @@ "extracted": true, "article_count": 8, "markdown_dir": "markdown/530", - "extracted_at": "2026-09-10T07:40:50" + "extracted_at": "2026-09-10T20:42:41", + "articles_fetched": 8, + "articles_failed": 0 }, "529": { "id": "529", @@ -60,7 +68,24 @@ "extracted": true, "article_count": 8, "markdown_dir": "markdown/529", - "extracted_at": "2026-09-10T07:40:50" + "extracted_at": "2026-09-10T20:43:10", + "articles_fetched": 8, + "articles_failed": 0 + }, + "528": { + "id": "528", + "title": "SRE Weekly Issue #528", + "url": "https://sreweekly.com/sre-weekly-issue-528/", + "pub_date": "2026-08-02", + "html_file": "html/528-2026-08-02.html", + "fetched_at": "2026-09-10T20:51:04", + "source": "backfill", + "extracted": true, + "article_count": 8, + "markdown_dir": "markdown/528", + "extracted_at": "2026-09-10T20:51:38", + "articles_fetched": 8, + "articles_failed": 0 } } } \ No newline at end of file diff --git a/sreweekly/markdown/528/01-content-ingestion-podcast-video-incident-report.md b/sreweekly/markdown/528/01-content-ingestion-podcast-video-incident-report.md new file mode 100644 index 00000000..fc7b3d27 --- /dev/null +++ b/sreweekly/markdown/528/01-content-ingestion-podcast-video-incident-report.md @@ -0,0 +1,71 @@ +# Content Ingestion & Podcast Video Incident Report + +- **期号**: SRE Weekly Issue #528(2026-08-02) +- **作者**: Jim Whitehead, Ulrik Mikaelsson, John Lagomarsino, and Saunak Jai Chakrabarti — Spotify +- **链接**: https://engineering.atspotify.com/2026/7/content-ingestion-and-podcast-video-incident-report/ + +## 简介 + +Spotify has had some difficulty around podcast publishing, and they shared this analysis of the worst incident. + +## 正文 + +# Content Ingestion & Podcast Video Incident Report + +![Feature Image](https://images.ctfassets.net/p762jor363g1/3AiqgaI6uFPcF6zc7lYX5t/bbed1b145db94423441163d9f149a02e/incident_report.png) + +Over the past two months, podcast creators have experienced a series of reliability issues on Spotify. This report covers the most significant of these, the June 24 publishing delay, in full detail and describes the broader reliability program now underway across our publishing pipeline. + +## What happened? + +When a podcast creator publishes a new episode, the audio and video content goes through a series of processing steps before it becomes available to Spotify users. These steps include transcoding (converting media into the formats our apps need) and content analysis. + +On June 24, our video transcoding infrastructure reached maximum capacity. This created a backlog that delayed the publication of video podcast episodes for several hours. Creators reported that their episodes were not appearing on Spotify as expected. A queue built up and video podcast episodes that would normally be published within minutes were delayed for hours. Understandably, some creators re-uploaded episodes that had not appeared, which added further load. That is on us, not them: the system should have confirmed their upload was received and queued, and it did not. + +![BlogPostGraph](https://images.ctfassets.net/p762jor363g1/7MLYmEAuLL5pj8xOtqHT6I/bf82c56512ed295e1ddea673040e9f1a/BlogPostGraph.png) + +The image above shows the build-up and subsequent emptying of the medium-priority (used for new episodes - shown in blue) and low-priority (used for updates to older episodes - shown in yellow) queues. + +Four factors converged to create this situation: + +- **Our transcoding infrastructure was running with insufficient headroom to handle large spikes of content delivery** . During typical periods of content submission our transcoding systems were able to scale and support both low-priority transcoding as well as the high-priority publication of new content. However, there wasn’t sufficient headroom to scale and support large spikes caused by the bulk delivery of new content. + +- **A scheduled batch processing job was running.** On occasion, we need to re-process existing episodes to be compatible with changes or additions to Spotify’s playback systems. During the June 24 disruption, a routine batch job processing existing content was consuming additional capacity alongside the regular processing for new episodes. While this job appeared fine earlier in the day, it became problematic when combined with increased content submissions. + +- **Recent improvements increased per-item processing cost.** We had recently changed our video transcoding to deliver better quality at lower bitrates. That change increased the time and processing power each episode requires, and we did not fully account for that added demand in our capacity planning. + +- **A software bug was underutilizing available compute resources.** Following a recent infrastructure migration to more powerful hardware, a bug in our resource scheduling caused our systems to underuse available processing capacity, reducing throughput by about 10%. + +When we identified the issue, we stopped the batch job, deployed a fix for the resource scheduling bug, and added additional processing capacity overnight. By the following morning, all backlogs had cleared and publishing was operating normally. We subsequently added further capacity to provide the headroom that had been missing. + +## Timeline (UTC) + +- **13:30** — Early alerts fire in our internal monitoring. Not immediately recognized as a broader capacity issue. +- **15:00** — Video podcast delivery spike pushes transcoding close to maximum capacity. +- **16:35** — Batch processing job stopped to free capacity. +- **17:31** — First Creator report of an issue impacting podcast video publishing received. +- **17:34** — Automated alerts confirm queue backlog exceeding thresholds. Incident response begins. +- **19:00** — Creator reports escalated to incident team. +- **20:49** — Software fix deployed to improve resource utilization. +- **00:14 (Jun 25)** — Additional processing cluster brought online. +- **01:02** — All queues cleared. +- **07:30** — Full confirmation: all publishing pipelines operating normally. + +One thing this timeline makes plain: roughly four hours passed between the first alerts and formal incident response. We want to do better. Engineers investigating the early alerts stopped the batch job at 16:35, but we did not recognize the full scope of the capacity problem until queues breached thresholds at 17:34. The monitoring improvements described below exist to close exactly that gap. + +## Where do we go from here? + +We have already taken several steps to address this specific incident: + +- Increased our transcoding capacity by approximately 67%, providing significantly more headroom for traffic spikes and batch operations. +- Fixed the resource scheduling bug that was leaving under-utilized compute capacity. +- Improved our monitoring to alert earlier when capacity is approaching limits. + +Beyond this specific incident, we are investing in broader improvements to the reliability of our podcast publishing pipeline. We've formed a dedicated cross-team effort focused on: + +- Building better capacity planning that accounts for not just steady-state traffic, but also burst capacity and incident recovery. +- Improving prioritization across our publishing systems so that real-time content from creators is always processed ahead of background operations. +- Extending rate limiting and backpressure mechanisms throughout the pipeline to handle unexpected load gracefully. +- During this incident, many creators learned something was wrong from their audiences before they heard anything from us. We are improving our processes and technical capabilities so creators get notified as soon as possible when things aren’t working. + +When a creator hits publish, their audience is waiting, and hours matter. We fell short repeatedly this summer, and we know a report like this only counts if the next incident is handled better than the last. The work above is how we intend to earn that trust back. diff --git a/sreweekly/markdown/528/02-the-pulse-quitting-spotify-podcasts-over-reliability.md b/sreweekly/markdown/528/02-the-pulse-quitting-spotify-podcasts-over-reliability.md new file mode 100644 index 00000000..03905cca --- /dev/null +++ b/sreweekly/markdown/528/02-the-pulse-quitting-spotify-podcasts-over-reliability.md @@ -0,0 +1,169 @@ +# The Pulse: Quitting Spotify Podcasts over reliability + +- **期号**: SRE Weekly Issue #528(2026-08-02) +- **作者**: Gergely Orosz — The Pragmatic Engineer +- **链接**: https://blog.pragmaticengineer.com/the-pulse-quitting-spotify-podcasts-over-reliability/ + +## 简介 + +…and here’s where it gets interesting. This post shares the user point of view on the Spotify issues, including fact-checking their published timeline. + +## 正文 + +*Hi, this is Gergely with a bonus, free issue of the Pragmatic Engineer Newsletter. In every issue, I cover Big Tech and startups through the lens of senior engineers and engineering leaders. Today, we cover one out of four topics of* *last week's The Pulse issue**. Full subscribers received the article below seven days ago. If you’ve been forwarded this email, you can* *subscribe here**.* + +You can no longer watch The Pragmatic Engineer Podcast as *video* in the Spotify app (only as audio) because I have quit publishing video on that streaming platform. This comes after I decided that reliability takes a back seat within that team – and across much of Spotify. Unlike on other platforms such as YouTube, Apple Podcasts, and Substack, I’ve recently encountered a series of reliability issues around Spotify being unable to process video episodes. Even though I enjoyed a direct link with the Podcasts team there, things haven’t improved. + +So from now, I will no longer be publishing video episodes on Spotify. You can find videos of my in-depth chats with guests only [on YouTube](https://www.youtube.com/@pragmaticengineer?ref=blog.pragmaticengineer.com). *Apologies for any inconvenience this change causes!* Audio episodes of the podcast can still be found [on Spotify](https://open.spotify.com/show/2Bho9xCbOQMWMJ7UKmqCzD?ref=blog.pragmaticengineer.com) via the RSS podcast feed hosted [on Substack](https://pragmaticpodcast.com/?ref=blog.pragmaticengineer.com). + +Honestly, the decision to quit the streaming giant wasn’t hard, and I reckon there’s a point here about the risk of deprioritizing reliable operations at major companies in order to push on things like AI adoption, as Spotify seems to be doing. + +Some context: for the first two years of The Pragmatic Engineer Podcast, it was published on three podcast platforms: + +1. **Substack’s podcast platform (audio)** : this is where the[“master” RSS feed](https://api.substack.com/feed/podcast/458709.rss) is served to the likes of Apple Podcasts, the web, Overcast, Pocket Casts, etc +2. **YouTube (video):** video episodes uploaded individually +3. **Spotify (video + audio):** every video episode *was* uploaded individually and then served as video or audio episodes from the platform. + +As someone hosting a podcast, there are good reasons to bother doing three separate uploads: + +- **Most podcast platforms don’t support video.** There will always be a need for a platform that serves the master RSS feed for audio versions while the video ones are elsewhere. +- **YouTube doesn’t integrate with anything.** YouTube is the leader in video podcast distribution, and uploading there directly makes sense. +- **I had a direct line to the Spotify team, which was a big plus.** Starting out the podcast, I had the unusual privilege of contact with the podcasts team, thanks to the newsletter gaining a decently-size audience. I was persuaded to take the plunge with them. + +For eighteen months, nothing *major* went wrong. The admin portal for podcast publishers (called ‘Spotify Creators’) was pretty wonky; it gave intermittent errors, and was unable to remember me when I signed in, so, each Wednesday, I’d have to sign in with a code sent to my email to publish an episode. + +But overall, things worked, until it all went suddenly downhill… + +### **Unable to publish Spotify podcast episodes 3 out of 5 weeks** + +From late May, I did not include links to Spotify on new episode announcements because their podcasts product or platform seemingly had outages every time one published on Wednesdays at around 9am PST / 12pm EST / 6pm EU time. + +**Outage #1 (20 May): podcast publishing broke**, my episode would not process on Spotify for 2+ hours. When uploading a video file to Spotify, there’s a processing pipeline that runs to create chunks of the podcast in different video and audio formats. This pipeline appeared to stop running, meaning new episodes were not published. + +It was not just the publishing that broke: the Creator portal looked absurd, with NaN% values everywhere, during the outage: + +![](https://substackcdn.com/image/fetch/$s_!dFJF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8041085f-fc43-4b61-9e99-47d7e11e5a84_1752x1058.png) + + +*During outage #1* +I emailed the Spotify team to alert them about the outage and also [complained online](https://x.com/GergelyOrosz/status/2057127878517526860?s=20&ref=blog.pragmaticengineer.com). I got a response, confirming the outage and pledging to do better: + +“The issue was in one of our podcast publishing metadata pipelines. A small subset of episodes completed normal media processing but then missed a downstream publish update because a newly introduced validation signal was not correctly wired into the logic that wakes up the publishing path. In simpler terms: the episode could become eligible to publish, but the final propagation step was not reliably triggered for that class of episodes. + +We identified the root cause, deployed a fix, and reprocessed the affected episodes with all-clear called early this morning. We’re also tightening the system so that fields used for publishing eligibility cannot be added without also triggering the relevant downstream updates. + +Separately, we’re reviewing how partial creator-impacting publishing delays are surfaced, because even when this is not a broad platform outage, it is still a bad experience for publishers like yourself. + +Apologies again that you hit this. It was a real bug, not a wide outage, but it hit some of our most relevant creators.” + +**Outage #2 (17 June): Spotify down.** Four weeks later, when attempting to publish a video episode, all of Spotify went down for many users, [including myself.](https://x.com/GergelyOrosz/status/2067285989710582271?s=20&ref=blog.pragmaticengineer.com) + +![](https://substackcdn.com/image/fetch/$s_!nelU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fece389ea-6dd6-45c8-b269-718ee5fc0098_1100x620.png) + +Spotify does not maintain a status page, so it’s impossible to tell how widespread the outage was. I didn’t include a Spotify link in that week’s announcement either. + +**Outage #3 (24 June): podcast publishing broke – again.** Outage #3 in five weeks; *deja vu*. This time, it was episode publishing not working, yet again. After waiting two hours for the episode to publish on Spotify, I yet again sent out the announcement with no Spotify link. + +I also emailed the Spotify Podcasts team, who confirmed the outage. I said I was considering stopping publishing video episodes, and to switch to audio-only publishing (which means pointing Spotify to my master RSS feed.) I said that an apology was appreciated but it wasn’t enough to make it worth publishing video episodes there. + +**I also asked for the incident review because I had the feeling that reliability was not all that important on this podcast product.** For the first outage I got a vague description of what happened, and promises of improvements that were never done – e.g. during this second outage, there was no improved communications to creators, which I was told would happen, after outage #1. + +Internally, Spotify’s team surely conducted an incident review as per usual, so I figured I’d hear back in about two weeks’ time, and assumed a reply would be forthcoming because I’d made clear I was ready to leave Spotify Podcasts if reliability didn’t improve. + +### **No incident review three weeks later, so I quit Spotify** + +The incident review had never arrived as promised by three weeks later, even though there had been time for it to be completed. It was yet another sign of a platform that has become unreliable. Also, the creator portal occasionally threw up this error: + +![](https://substackcdn.com/image/fetch/$s_!Y77E!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e90510f-31f3-4ea4-a0d5-b84827c929a0_868x470.png) + + +*Spotify’s creator portal on 16 July* +I checked my Spotify stats: stream plays had been trending downwards unsurprisingly, given the ongoing outages, while the other podcast platforms didn’t show the decline. It made me decide “enough is enough” and to move off Spotify. + +Staying on their platform depended on seeing an incident review, but they didn’t prioritize transparency, still had no status page, and nobody had built a feature for episode-processing status like YouTube has had for years. So, I pulled the plug and left: + +![](https://substackcdn.com/image/fetch/$s_!fWEQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4477d3a8-1d2b-4434-80d8-f8eb7165562a_1358x930.png) + +After I made the switch away from Spotify, the platform’s creators portal became buggier than ever, as in these examples: + +![](https://substackcdn.com/image/fetch/$s_!cc6F!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd0e0887-73fe-4c4b-9ca0-76460aecca8f_2048x1150.png) + +Comments disappeared: + +![](https://substackcdn.com/image/fetch/$s_!Kclv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F182ad478-188d-4f67-ab5b-962fa3196d3b_2048x1546.png) + + +*My show had no comments, suddenly* +… even though other parts of the UI showed dozens of comments: + +![](https://substackcdn.com/image/fetch/$s_!ks8o!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ef9802b-f2d7-466d-8c4d-b59c1972be91_2048x1190.png) + + +*Zero comments, yet episodes with comments* +Episode links directed to 404 pages: + +![](https://substackcdn.com/image/fetch/$s_!Keep!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8c237ccb-e0cb-4bac-903d-83da831ea9e9_2048x1309.png) + +A day or two later, these issues disappeared: I assume no one had tested the flow of moving away from Spotify Podcasts to an RSS feed, and it’s why the experience was so poor. + +### **Incident review finally published, but with a wrong timeline** + +A few days after offboarding from Spotify, their team [published the incident report](https://engineering.atspotify.com/2026/7/content-ingestion-and-podcast-video-incident-report?ref=blog.pragmaticengineer.com) for outage #3. Reading through it, something did not add up in the timeline: + +![](https://substackcdn.com/image/fetch/$s_!5ROX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0bb70a5-df4c-444c-a9a9-9642c084f7c1_1662x866.png) + +My email account confirmed that I mailed the Spotify team at around 17:30 about the outage. So, after weeks of creating this report, why did the incident report downplay the fact that customers alerted the team before their own automated alerts fired?I complained to the Podcasts team, and to their credit, the incident report was updated: + +![](https://substackcdn.com/image/fetch/$s_!a8Ln!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F36fa0a3a-ba07-453f-b1dd-0269b80a7552_2048x932.png) + + +*The updated incident timeline* +I didn’t like how high-level [the report is](https://engineering.atspotify.com/2026/7/content-ingestion-and-podcast-video-incident-report?ref=blog.pragmaticengineer.com), and how vague the promised improvements were. Specifically, this one: + +“During this incident, many creators learned something was wrong from their audiences before they heard anything from us. We are improving our processes and technical capabilities so creators get notified as soon as possible when things aren’t working.” + +Overall, I don’t regret the choice to leave, particularly when the focus of Spotify’s leadership is on AI, not reliability. + +### **Does Spotify have “AI psychosis?”** + +Previously, I used the term “AI psychosis” differently from the usual way of describing when someone starts believing everything an AI model tells them, however outlandish. I [applied it](https://newsletter.pragmaticengineer.com/i/202307236/7-is-ai-psychosis-just-a-meta-issue?ref=blog.pragmaticengineer.com) to Meta’s rush to develop its own AI model at the cost of the reliability of its profitable business activities. This was based on Instagram’s most embarrassing-ever account takeover incident, which [occurred](https://newsletter.pragmaticengineer.com/i/202307236/7-is-ai-psychosis-just-a-meta-issue?ref=blog.pragmaticengineer.com) when the team responsible for Instagram’s Trust & Safety was slashed. Soon after, AI-generated, AI-reviewed code caused the hacking of a former US president’s account. + +At Spotify, it should have gone the other way. In March, I had the opportunity to meet its Head of Technology & Platforms, Tyson Singer, who said the company puts reliability far ahead of AI adoption, and doesn’t adopt AI for its own sake. So, it was somewhat surprising to read the summary below of a podcast Spotify did [with Anthropic](https://x.com/ClaudeDevs/status/2071671418245492926?s=20&ref=blog.pragmaticengineer.com): + +“Spotify now ships 4,500 production deploys a day, and 73% of PRs are now AI-assisted. + +Niklas Gustavsson (VP of Engineering at Spotify) keeps 5 to 10 Claude sessions running in tmux, one per git worktree, agents working in the background. All of it inside a 20M+ line monorepo. He expected agents to struggle at that size, but it’s worked well. + +Spotify’s migration codemods grew into thousands of lines of edge cases. Code has too much API surface for static rewrites. Early LLMs barely did better. Adding a judge took PR success from ~25% to 80%. + +All of this leans on verification, the single most important thing when agents are used and the place most companies underinvest + +Spotify rebuilt their test automation around it so engineers can confidently guide and supervise agents, rather than manually execute repetitive tasks.” + +It seems to me that all the talk is about *usage* of AI, and none about *reliability*, all while Spotify’s platform becomes less reliable than ever, at the same time as the streamer is going all-in on AI; with AI judges and devs running 5-10 parallel Claude sessions. + +All things considered, it’s worth asking if Spotify has the corporate variant of “AI psychosis”, whereby the reliability of a successful operation gets torched in the chase for the next big thing by executives. I don’t even think Spotify is all that different from Meta and other companies in this! + +Things look bad, based on the quality and reliability degradation of products. Annoyingly, in many cases, customers don’t really have the choice of going elsewhere. My podcast is an exception, as video podcasts on Spotify never truly took off, so quitting the platform wasn’t a big deal. Even so, I’m particularly disappointed that Spotify has prioritized AI usage over reliability. I know some executives there pushed against this, but I feel safe in assuming that they lost that battle. + +### **Value of staying reliable & “sucking less”** + +Max Kanat-Alexander, distinguished engineer at Capital One, has [written about](https://www.codesimplicity.com/post/suck-less/?ref=blog.pragmaticengineer.com) how a software project can become wildly successful just by “sucking less” in his reflections upon the success of the Bugzilla project, (2004-2009): + +“All you have to do to succeed in software is to consistently suck less with every release. + +Nobody would say that Bugzilla 2.18 was awesome, but everybody would say that it sucked less than Bugzilla 2.16 did. Bugzilla 2.20 wasn’t perfect, but without a doubt, it sucked less than Bugzilla 2.18. And then Bugzilla 3.0 fixed a whole lot of sucking in Bugzilla, and it got a whole lot more downloads. + +Why is it that this worked? + +As long as you consistently suck less with every release, you will retain most of your users. You’re fixing the things that bother them, so there’s no reason for them to switch away. Even if you didn’t fix everything in this release, if you sucked less, your users will have faith that eventually, the things that bother them will be fixed. New users will find your software, and they’ll stick with it too. And in this way, your user count will increase steadily over time.**But what happens if you release frequently, but instead of fixing the things in your software that suck, you just add new features that don’t fix the sucking?** Well, eventually the patience of the individual user is going to run out. They’re not going to wait forever for your software to stop sucking.” + +Personally, I got tired of Spotify’s Podcasts product continually going in the wrong direction on Max’s scale: the poor reliability, frequent errors on the Creators site, and the sense that they don’t really care about improving *existing* things. + +Read the full issue of [last week's The Pulse](https://pragmaticengineer.substack.com/p/the-pulse-quitting-spotify-podcasts). The full The Pulse additionally covers: + +1. **Will Kimi K3 trigger US push for closed-source AI models?**  Moonshot AI’s latest open model, Kimi K3, is on par with Anthropic’s Fable 5. Could it lead to the US government regulating or banning Chinese open models to protect US labs? +2. **AWS laughs off “heart attack” billing error.**  AWS customers were billed trillions more than they should have been, due to what was likely a conversion error. But instead of sharing an incident report, AWS saw the funny side. +3. **Industry pulse.**  OpenAI’s unreleased model tried to hack HuggingFace to improve its test scores, X took more than a year to develop its new Android app, Google’s new AI model flops, and more. + + [Subscribe to my weekly newsletter](https://newsletter.pragmaticengineer.com/about) to get articles like this in your inbox. It's a pretty good read - and the [#1 software engineering newsletter](https://substack.com/top/technology) on Substack. diff --git a/sreweekly/markdown/528/03-modern-software-architecture-means-nobody-has-the-whole-picture-to-ass.md b/sreweekly/markdown/528/03-modern-software-architecture-means-nobody-has-the-whole-picture-to-ass.md new file mode 100644 index 00000000..8a9db8ce --- /dev/null +++ b/sreweekly/markdown/528/03-modern-software-architecture-means-nobody-has-the-whole-picture-to-ass.md @@ -0,0 +1,53 @@ +# Modern software architecture means nobody has the whole picture. To assemble one in an emergency, you need an incident tech lead. + +- **期号**: SRE Weekly Issue #528(2026-08-02) +- **作者**: Brent Chapman +- **链接**: https://greatcircle.com/blog/2026/06/16/incident-tech-lead/ + +## 简介 + +New incident role unlocked: the incident tech lead. I enjoyed the description of the interplay between the tech lead and the incident commander. + +## 正文 + +Every large software development organization has made the same bargain: if nobody has to understand the whole system, we can build a more capable system than we otherwise could, even though it grows bigger and more complex. We break systems into components with well-defined interfaces so that each team needs to understand only the pieces it owns, plus the interfaces of its neighbors. That’s the point of every decomposition strategy, whether it’s microservices, bounded contexts, service ownership, or a well-modularized monolith. The architecture deliberately limits what any one person has to hold in their head. + +It’s a sound strategy. It’s also why some incidents are so much harder than others. + +Lorin Hochstein named this pattern beautifully in a recent post, [The demon of the gaps](https://surfingcomplexity.blog/2026/06/06/the-demon-of-the-gaps/). Failures that stay inside a single component are the easy ones; you page the owning team, they figure out what’s wrong, and they fix it. The hairy incidents emerge from unexpected interactions across components: several services throwing errors at once, or no services throwing errors while customers see broken behavior anyway. As Lorin puts it, “you’ve built an analysis solution but you’re now faced with a synthesis problem.” In order to scale, the architecture deliberately optimized away the need for whole-system understanding; now the whole system isn’t working, but nobody has that understanding to call on. + +I’ve watched this play out in incident channels many times. Subject matter experts from six different teams, each reporting that their own service looks healthy. Six dashboards are green, but meanwhile, checkout is still failing for customers. The knowledge needed to explain what’s happening exists, distributed across six heads, but nobody is assembling the pieces. The responders need to understand how the system as a whole is behaving right now. That understanding has to be built live, under pressure, from multiple partial models. That’s synthesis work, and it doesn’t happen on its own. + +Lorin observes that guidance on preparing for this work is almost nonexistent. Here’s the encouraging part: closing the structural gap is fairly straightforward. Most companies already use structured incident roles (incident commander, subject matter expert, customer liaison, etc.); they need to add a synthesis role, activated when needed for complex incidents. + +# Synthesis is a job. Name it. + +The **incident commander (IC)** coordinates the overall response; every incident will have one. On the most complex incidents, though, where you need this synthesis function most, the trick is to also activate an **incident tech lead (TL)** to lead the technical investigation. The role is analogous to the tech lead role many teams have in their everyday structure, but its scope is *the incident* rather than one particular service. Most companies have never established the incident TL role, and for routine incidents they don’t miss it: the IC can handle the technical side along with everything else, but for complex incidents, the TL role can be incredibly valuable. + +The incident TL job, properly understood, is the synthesis job: connecting observations across component boundaries, correlating the partial models from different subject matter experts, and maintaining the evolving picture of how the system is failing and what we’re doing about it. The TL doesn’t need to be the deepest expert in any single component. They need to be good at building a working model out of other people’s expertise, and much of that work is cross-checking, holding indications from different components up against each other and noticing the discrepancies: “If we’re seeing this in component A, we should be seeing that in component B, but we aren’t; why not?” “If A is doing this and B is doing that, the problem must be upstream of both.” “Wait, A says one thing but B says another; they can’t both be right, can they?” + +The separation between IC and TL exists to protect that work, and it cuts both ways. Synthesis requires sustained, heads-down attention; you can’t reconstruct a system model in the gaps between stakeholder updates and staffing decisions. And the same complexity that makes an incident demand serious synthesis also multiplies the outward-facing work: more stakeholders to update, more escalations, more decisions about the response itself. The two loads peak together, and one person can’t carry both. + +The IC takes everything outward-facing precisely so the TL can stay immersed in the technical picture, and the TL handles the heads-down focused work so that the IC has time for everything else. When I’m the incident commander, one of the most valuable things I can do for my tech lead is keep everyone else out of their hair. But the separation is a division of labor, not a wall. I like to think of the IC and the TL standing back to back, facing opposite directions, talking over their shoulders to keep each other informed. Each is watching a different part of the horizon, and together they have the whole picture. + +# The response team crosses the boundaries on purpose + +An incident response is a temporary organization: an ad hoc team assembled across ownership boundaries for exactly as long as the incident lasts. Conway’s law observes that systems end up mirroring the communication structures of the organizations that build them, and the mirror works in both directions: your team boundaries and your component boundaries align, which is exactly what you want for everyday work. The incident structure deliberately cuts across those boundaries, because the gaps between components are where the problem lives. Pulling six SMEs into one channel isn’t enough by itself, though. A group of experts in the same room is a meeting; a group of experts with someone responsible for synthesizing what they know is a response. + +# The communication mechanisms are synthesis tools + +The standard incident communication practices may seem like bureaucratic overhead until you see what they’re for. “Going around the horn” (each responder, in turn, briefly reports what they’re seeing and doing) forces the partial models into the open, where the TL can correlate them. A periodic situation report, or SitRep, forces someone to compress the current understanding into a few sentences; writing it is itself an act of synthesis, and reading it gives every responder the same baseline picture to work from. Narrating before you act keeps each responder’s local view visible to the whole room. None of these mechanisms exists for discipline’s sake. They’re how a group of people, each holding a partial model, builds and maintains a shared one. + +# Wildfires don’t respect organizational boundaries either + +As is often the case in incident management, we can look to fire departments for inspiration and solutions. Consider a major wildfire. Dozens of agencies converge: federal, state, tribal, and local, some from hundreds or even thousands of miles away. No single agency understands the whole incident, with its terrain, weather, fuel, crews, and aircraft. The Incident Command System (ICS), the standard structure for emergency response in the US and beyond, is how all these disparate parts get pulled together into a coherent whole. ICS treats building the shared picture as a staffed function: a planning section tracks the situation and the resources, assembles the common operating picture, and distributes it to every responder through the incident action plan. Nobody simply hopes that shared understanding will emerge; somebody owns producing it. + +Software companies can borrow that lesson directly: treat synthesis as a named responsibility rather than an emergent property. If the IC role at your company is defined as “project manager of the outage” and nobody is explicitly responsible for assembling the technical picture, the synthesis function is unowned, and it will show in your cross-boundary incidents. Establish the incident tech lead role. Protect it from outward-facing distraction. And practice it: when you run game days or tabletop exercises, choose scenarios that cross team boundaries, because those are the scenarios that exercise synthesis rather than component expertise. + +Decomposition made whole-system understanding nobody’s everyday job, and that’s fine; it’s a good strategy with a known cost. Incident management structure is how you pay that cost only when you must, with machinery built for the moment. + +*I’m writing a book, “Incident Management for DevOps and SRE.” Sign up at [im4ds.com](https://im4ds.com) to be notified when it’s available, and to get occasional progress updates and early access to selected content.* + +*If your company needs help with incident management right now, that’s the focus of my consulting practice at [GreatCircle.com/im](https://greatcircle.com/im).* + +## Recent Comments diff --git a/sreweekly/markdown/528/04-the-quiet-quarter.md b/sreweekly/markdown/528/04-the-quiet-quarter.md new file mode 100644 index 00000000..fea21a7b --- /dev/null +++ b/sreweekly/markdown/528/04-the-quiet-quarter.md @@ -0,0 +1,87 @@ +# The Quiet Quarter + +- **期号**: SRE Weekly Issue #528(2026-08-02) +- **作者**: Tim Irving +- **链接**: https://read.zerosevzero.com/p/the-quiet-quarter + +## 简介 + +This one goes hard: if you try to reduce your incident count, your system will become less reliable, not more. Aim for more incidents, handled well. + +## 正文 + +![](https://substackcdn.com/image/fetch/$s_!5E6h!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F57a19704-71eb-447c-9d8b-6d3905c538c1_2544x1904.png) + +He was on a video call from a cafe, a product manager, flicking a begleri back and forth between sips of a flat white. I had a knucklebone going, switching it hand to hand while he talked. Two grown professionals, paid reasonable money to be taken seriously, fiddling with bits of string and weighted metal and picking over the companies that got it spectacularly wrong, as though hindsight were a kind of genius. + +We agreed on nearly everything until we didn’t. He had built a tidy picture of the two of us and laid it out before me. He was the irresistible force, leaning on the engineers to ship, to get product to market before the market lost interest. I was the immovable object on the far side of them, the incident barrier, there to slow the whole thing down before somebody broke something that mattered. He meant it as a compliment. + +I waited for him to finish and told him he had me exactly backwards. + +The begleri stopped. He went quiet, the particular quiet of a man recalculating, and then he wrote something on the pad beside his coffee and asked me to explain myself. So I did. I told him I was not the brake. I had never been the brake. If it were up to me the engineers would be breaking more, not less. + +The argument is not complicated, and I have made it often enough to make it fast. There are two ways to run a company that builds software. You change things quickly and break some of them, or you change things slowly and watch the market leave without you. There is no third lane. There is no measured middle where you ship at a sensible pace and nothing ever falls over. That lane is a fiction, sold to executives who find both of the real choices frightening. The companies that bought it are the ones who did not blow up. They sat very still, with a spotless record, and were buried holding it. The cleanest way to die in this business is to stop moving and call it discipline. + +So if you have accepted that you have to keep moving, you have already accepted the incidents. They are not the price of getting it wrong. They are the price of doing anything at all. The team that ships nothing has none of them. The team that ships has them the way a road has potholes, and you can resurface as often as you like, but you do not get the road without the wear. Once you stop pretending the number can reach zero, the useful question changes. It stops being how do we have fewer, which has only bad answers, and becomes what do we get out of the ones we have. The answer turns out to be quite a lot. + +The first thing you get is a team that tells you the truth early. When an incident is an ordinary event and not a permanent mark against your name, people put a hand up while the thing is still small and still cheap, instead of sitting on it and praying, which is what people do when the cost of admitting a fault is a hard conversation with someone who outranks them. Every catastrophe I have stood in the middle of had an earlier, smaller, survivable version that somebody decided not to mention. + +The second thing you get is competence, which is only practice wearing a better word. A team that runs incidents often runs them well. They know where the runbooks are, or they know the runbooks are useless and route around them, and either way they have a rhythm. The team that has not seen an incident in eighteen months has not banked eighteen months of safety. It has banked eighteen months of rust, and it will move through its first real one like a fire drill in a building where nobody can remember which door is the exit. + +The third thing is the one everyone says they want and almost nobody funds. If you run the incident well and tell the truth about it afterwards, you come out the far side knowing something about your system you did not know going in. John Allspaw calls an incident an unplanned investment, and he means it precisely. You did not choose to make it. You made it the moment the thing broke, and you do not even control its size. The only thing left in your hands is whether you collect the return, and most organisations pay the full cost of the outage and then throw away the receipt. + +None of this works while an incident is a thing to be ashamed of. The shame is the whole problem. It is what keeps the hand down, lets the skill go soft, and turns the review afterwards into a hunt for someone to blame instead of something to learn. So I said it to him plainly. I did not want fewer incidents. I wanted a team that had them often, ran them well, learned from them properly, and felt nothing sharper than mild professional interest the entire time. A team like that is more reliable than a team that has only been lucky. And it is a great deal more reliable than a team that has merely been quiet. + +![](https://substackcdn.com/image/fetch/$s_!bkbN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8e5fd40d-1fed-4475-a0dc-aefdbe5aece4_2544x1904.png) + +Here is the part that people accept in the abstract and resist in their bones. If incidents are the price of motion, their absence is not good news by default. A quiet quarter is not a trophy. It is a question, and it has more than one answer, and from the executive chair the answers are impossible to tell apart. + +A team can go quiet for three reasons. The first is that it is healthy. The work is good, the luck is holding, and nothing has broken because, for the moment, nothing had to. This happens. It happens less than anyone wants to believe, and it never holds still, because a healthy team that stops shipping stops being one within a quarter or two. The other two reasons wear the first one’s clothes. They look identical on a graph, and they are both rotten. + +The second reason is that the team is going soft, and there is no villain in this one, which is exactly what makes it hard to see. Nothing is being hidden. The incidents are not happening, the runbooks are quietly going out of date, and the one engineer who understood how the billing system fails at three in the morning has taken a job somewhere warmer, and nobody has noticed the gap because nothing has fallen into it yet. The capability does not announce its departure. It is simply not there on the day you reach for it. The rare large incident always comes, and when it does it lands on a team that has forgotten how to catch it. The quiet did not protect them. It disarmed them. + +The third reason is the one a head of engineering described to me once, quietly, the way people tell you things they have decided not to fix. The incident culture in his department was poisonous. People who caused an incident, or were merely standing near one when it went off, were marked for it, and the mark travelled. It followed them into performance reviews and into rooms they were not in. So his engineers had made the rational choice. They had stopped raising incidents. They had stopped spending time running the ones they could not avoid. And the post-incident review, the thing that turns an outage into knowledge, did not come up at all. He told me this as a problem he was observing, not one he was causing, which is its own kind of tell. + +It took me a moment to register what he had just listed for me. He had named, in order and without meaning to, the three things that make a team good at trouble, and he had explained that his department had switched off every one of them. No early warning. No practised hand. No learning. And the result of switching off all three, the figure sitting proudly at the top of his dashboard, was a low incident count. On paper, his was one of the calmer departments in the building. In truth it was one of the most dangerous, a place where everything that broke was either hidden or survived by luck, and where nobody was getting better at anything. + +This is the problem with the number. Health, rot, and cover-up all produce the same low count, and from the executive chair they are indistinguishable. So when the figure drops, the room relaxes, which is the most dangerous thing a falling incident count can make a room do. A low number is not information. It is the absence of information, wearing the costume of good news. A clean record is not proof that you are safe. Sometimes it is only proof that you are lucky, and sometimes it is proof that someone is lying to you, and the graph will never tell you which. The head of engineering at least knew which quiet he was standing in. Most people reading the dashboard never find out, right up until the quarter that is not quiet at all. + +Software has the good fortune that its quiet quarters usually end in a refund and an apology. Other industries run the same machine with the same blind spot, and when their quiet ends, it ends with bodies. The useful thing about those industries is that they investigate, at length and in public, so none of this has to be taken on faith. The reports exist. They are very long, and they all say a version of the same thing. + +![](https://substackcdn.com/image/fetch/$s_!4ZSH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90d4fd6f-6266-450e-9897-86c418634833_2544x1904.png) + +On the twentieth of April 2010, a group of BP and Transocean executives flew out to the Deepwater Horizon to hand the crew an award for seven years without a lost-time accident. They were on the bridge when the well blew out. The explosion and fire killed eleven people and put the largest oil spill in American history into the Gulf of Mexico. The record was real, and it was worthless, because it measured the wrong thing. Personal safety on the rig was excellent, the kind where the worst thing anyone pictures is a dropped pipe and a crushed foot. Process safety, the slow abstract business of whether the well itself would hold, was a disaster nobody was counting. The well needed twenty-one centralisers to seal correctly. They ran it with six. And the seven-year record was partly an illusion of its own, because the crews understood that raising a concern that delayed the drilling was a good way to lose your job, which means the number measured seven years without a reported problem, in a place where reporting one was punished. They were celebrating the silence at the exact moment it killed them. + +NASA learned the same lesson twice, seventeen years apart, and wrote the textbook in between. In the years before the Challenger, the rubber O-rings that sealed the booster joints kept eroding in flight, and because the shuttle kept coming home anyway, the erosion stopped being treated as a fault and became a known quirk you could fly with. The night before the launch the engineers who built the boosters warned that the cold would stiffen the seals past the point where they could hold. They were overruled, after being asked to do the one thing engineering cannot do on demand: prove the rocket would fail before it had failed. On the twenty-eighth of January 1986 the seal failed and the Challenger came apart seventy-three seconds after lift-off, live on television, in front of the classrooms full of children who had been gathered to watch a schoolteacher fly into space. The sociologist Diane Vaughan gave the pattern its name afterwards. She called it the normalisation of deviance: the slow process by which a warning sign, repeated often enough without disaster, gets quietly reclassified as normal. + +Then NASA did it again. Foam had been breaking off the external tank and striking the orbiter for years. The same foam, from the same ramp, had come away six times before, and because nothing had yet gone fatally wrong, the agency downgraded it from a flight-safety anomaly to a maintenance nuisance, a paperwork item, in the months before one more piece of it punched a hole in Columbia’s wing on the way up. The frequency of the warning had become the reason to ignore it. The engineers saw the strike on the launch film and asked to point a satellite at the wing to check the damage. The request was refused, on the grounds that it had not come through the proper channels. That hole went unexamined, and a fortnight later the orbiter disintegrated over Texas on re-entry, killing all seven aboard. + +The investigation board concluded that NASA’s culture had killed the crew as surely as the foam, a remarkable thing to have to write about the same organisation a second time. But the parallel runs deeper than a repeated blind spot. Both times, the engineers who wanted to stop were made to prove it would fail, while the managers who wanted to fly were asked to prove nothing at all. That is what it means to learn the same lesson twice. Not a forgotten fact, but a burden of proof that sat, on both occasions, on exactly the wrong shoulder. + +There is one industry that looked at the same raw material and drew the opposite conclusion, and it is the reason you can board a plane without thinking about it. Aviation decided, decades ago, that the near-miss was the most valuable thing it owned, and that the only way to get people to report the near-miss was to make reporting safe. The result is the Aviation Safety Reporting System: confidential, voluntary, explicitly non-punitive, with limited immunity for anyone who files. It was built on a single insight, that fear of punishment was suppressing the exact information that could prevent the next crash, so the system strips the reporter’s identity and shields them from enforcement, and in return it has gathered more than two million reports on things that nearly went wrong. The whole edifice runs on the principle that you want more incidents on the record, not fewer. And the neutral party chosen to run it, chosen precisely because it had no power to punish anyone, was NASA. The agency that twice mistook a quiet record for a safe one also operates the finest argument in the world for never doing that again. + +The pattern is consistent enough to be a law. The organisations with the most spotless records are not the safest ones. They are very often the ones that have stopped looking, or stopped listening, or taught their people that looking and listening are career-limiting moves. A clean record is a fact about your reporting, not a fact about your safety, and the two come apart at the worst possible moment. So when you find yourself watching an incident count fall and feeling the room go warm with relief, it is worth knowing whose company you are in. You are standing on the deck of the Deepwater Horizon, on the evening of the twentieth of April, holding the award. + +![](https://substackcdn.com/image/fetch/$s_!a__w!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F60c2a443-cfcf-4b41-b14e-22b46e43b3e6_2544x1904.png) + +It is tempting to call the executives who chase zero foolish, but that lets them off too easily and gets the history wrong. The goal is not stupid. It is old. It is a piece of received wisdom from an era when it made perfect sense, kept alive by people who never noticed the ground underneath it had moved. To understand why the number lies, you have to go back to when it told the truth. + +Reliability did not begin as a software idea. It began in hardware, in the discipline of working out how long a physical thing would run before it broke. The key number was the mean time between failures, and it was a number that meant something, because the thing it described was a component sitting still and wearing out at a knowable rate. A disk had a failure rate you could measure. So you stacked the disks into redundant arrays, the tall humming cabinets that filled the server rooms, and you did the arithmetic, and you could say with a straight face that the odds of losing everything at once were vanishingly small. The software on top of it moved at the same stately pace. It shipped in versions, on discs, in boxes, and between releases it sat as still as the hardware. Change was a rare and deliberate event, scheduled and rehearsed and dreaded. In a world like that, fewer failures was a coherent goal, because failure was a thing that wore out on a schedule, and you could engineer against a schedule. + +Then the ground moved, and almost nobody changed the number. The discs and the boxes went away. Software stopped shipping in versions and started shipping continuously, a hundred or a thousand small changes a day, and the systems it ran on stopped being cabinets you could point at and became sprawling distributed things that no single person could hold in their head or draw on a whiteboard. The hardware problem, the one the old number was built for, was solved so completely by redundancy and the cloud that a failing disk became a non-event. But the failures did not stop. They changed shape. They stopped being a part wearing out and became something stranger, an emergent property of too many moving pieces interacting in a state nobody designed and nobody foresaw. You cannot calculate a mean time between failures for the sentence the system has never executed before. There is no schedule for a surprise. + +What survived all this was the goal. We are still chasing the number that belonged to the cabinets, still treating a low incident count as the mark of a healthy system, in an environment where the thing that number measured no longer exists. Zero incidents is not an ambition. It is a fossil, perfectly shaped for a world that has been gone for twenty years. Chasing it now is not discipline. It is taxidermy. You are keeping a dead thing in a lifelike pose and asking it to guard the house. + +And the inheritance is not harmless, because the world it came from is the one place the strategy worked. When failure was rare and predictable, you could afford to know nothing about it. You could keep your people innocent of how the system broke, because it broke seldom and it broke in familiar ways. None of that holds now. Chase zero in a system that fails by surprise and you do not get a safe organisation, you get an ignorant one, fluent in nothing, practised at nothing, blind to the shapes its own failures take. And the large incident is still coming, because it always is, only now it arrives in a form no one has seen, at a team that has never run one, inside a system no one can reason about under pressure. The pursuit of zero does not protect you from that day. It is the thing that sends you into it unarmed. + +![](https://substackcdn.com/image/fetch/$s_!0ZvI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29277cae-2b39-4175-ba7b-4d08285e4a0f_2544x1904.png) + +So you stop counting and start training. If incidents are the only honest signal you get about how your system fails, the work is not to silence the signal, it is to get fluent in it. You make it safe to raise one, so the small ones surface while they are still small. You run them often enough that the team moves through one the way a good crew moves through a storm, without drama, because they have done it before. You take the post-incident review seriously, as the place where the expensive lesson gets collected instead of binned. And when the system has been quiet for too long, you do not relax. You go and break it yourself, on a Tuesday afternoon, with everyone watching: a drill, a game day, a failure you injected on purpose so you could meet it on your terms instead of its own. The teams that do this are not reckless. They are the least surprised people in the building. + +This is a different definition of reliability than the one on the dashboard, and it is the true one. Reliability was never the absence of failure. A system that has not failed is not reliable, it is untested, and the two feel identical right up until the day they do not. Reliability is what a system and the people around it do when failure arrives, which it will, on a long enough timeline, no matter how clever anyone was at the start. The reliable team is not the one that nothing happens to. It is the one that has made itself hard to surprise and quick to recover, that treats every incident as a rehearsal for the next, and that has therefore turned the thing everyone else is afraid of into the thing it is quietly best at. + +The product manager wanted me to be the immovable object, the thing planted in front of the engineers so they could not break anything. I understand the appeal. It is a tidy picture, and it turns the incident person into a kind of guardian. But the immovable object is the thing this whole essay has been about burying. It is the company that sat still with a spotless record. It is the rig with the award. It is the agency holding the textbook it had already written about itself. The object that refuses to move does not prevent the disaster. It waits for it. + +There is a name on this masthead, and it has been the joke the whole time. Zero sev zero. No incidents, and none of the worst kind, the clean and total silence that every executive has been trained to pray for. I did not call it that because I want it. I called it that because it is the most dangerous condition a system can be in, and almost no one recognises it while they are standing in it. A zero on that line does not tell you that you are safe. It tells you that you are healthy, or that you are rotting, or that someone has stopped telling you the truth, and the graph will not say which, and the day you find out which is not a day you get to choose. So I will leave it where I left it with him, the begleri turning in his fingers and the knucklebone going hand to hand across mine. I do not want fewer incidents. I want more of them, smaller and louder and sooner, run by people who are not afraid of them. + +That is not the absence of trouble. It is the only kind of safety that was ever real. diff --git a/sreweekly/markdown/528/05-30-to-70-prs-a-day-how-we-managed-to-not-wreck-our-systems.md b/sreweekly/markdown/528/05-30-to-70-prs-a-day-how-we-managed-to-not-wreck-our-systems.md new file mode 100644 index 00000000..3f6493b1 --- /dev/null +++ b/sreweekly/markdown/528/05-30-to-70-prs-a-day-how-we-managed-to-not-wreck-our-systems.md @@ -0,0 +1,201 @@ +# 30 to 70 PRs a Day: How We Managed to Not Wreck Our Systems + +- **期号**: SRE Weekly Issue #528(2026-08-02) +- **作者**: Liz Fong-Jones — Honeycomb +- **链接**: https://www.honeycomb.io/blog/30-70-prs-day-how-we-managed-not-wreck-systems + +## 简介 + +There’s some brutal honesty in here that I find refreshing, especially around the impact on incidents and incident response. + +## 正文 + +# 30 to 70 PRs a Day: How We Managed to Not Wreck Our Systems + +The Honeycomb engineering team set out to double our productivity in a year. This is how we did it, what we did to keep things stable, what it cost us, and what we’re still figuring out. + +![Liz Fong-Jones](https://www.honeycomb.io/_next/image?url=https%3A%2F%2Fcdn.sanity.io%2Fimages%2F927dxq0h%2Fproduction%2Ff6f77e2c3753d50fc961f3d1b48b28b3ba91008b-3146x3146.jpg%3Fw%3D80%26h%3D80&w=256&q=75) + +By: [Liz Fong-Jones](https://www.honeycomb.io/author/lizf) + +![The Second Edition of Observability Engineering Is Here](https://www.honeycomb.io/_next/image?url=https%3A%2F%2Fcdn.sanity.io%2Fimages%2F927dxq0h%2Fproduction%2Fee7172fc6d57df70b4b2dd303f6a58fc87e24e93-3840x2160.png&w=640&q=75) + +#### The Second Edition of Observability Engineering Is Here + +The second edition of Observability Engineering is available for download on our website. + +[Learn More](https://www.honeycomb.io/blog/the-second-edition-is-here) + +![30 to 70 PRs a Day: How We Managed to Not Wreck Our Systems](https://www.honeycomb.io/_next/image?url=https%3A%2F%2Fcdn.sanity.io%2Fimages%2F927dxq0h%2Fproduction%2F5e261bbd627f0eede01281d7766cc21a4245155b-3840x2160.png&w=3840&q=75) + +*In this two-part blog series, I give a detailed report-out on how our Honeycomb engineering team 2.5x-ed our throughput using AI without breaking everything or lowering our standards for quality. Part 1 explains how we did it and shows data about how that ramp-up happened. [Part 2 shares what we learned.](https://www.honeycomb.io/blog/ai-amplifies-existing-practices-lessons-ai-first-strategy)* + +## TL;DR + +- Peak-weekday merges roughly doubled (~30 to ~74) as AI-attributed lines went from near-zero to a floor of 82.6% of new code by June 2026. Incidents grew too, tracking that change volume about as linearly as you'd expect; the goal now is keeping each failure cheap to contain, not holding the count flat. +- The gains came in three phases: slow experimentation (2025), a tooling-driven adoption bump (October 2025), then a step-change in delegation intensity after Opus 4.6 shipped in February 2026—same engineers, same tools, but they stopped supervising every step. +- Our core thesis: AI amplifies your existing practices. It makes a dysfunctional org more dysfunctional and a high-autonomy, high-ownership org faster. The practices—continuous delivery, fast and AI-legible CI, closed-loop observability, CLAUDE.md, feature flags—are the actual story, not the multiplier. + +Why did I put that big number in the title? It’s the number that gets you to click. But we’re going to dig into all of the caveats and details beyond just 2.5x-ing our throughput, including whether we let quality slip. This blog and its companion post are our report-out on what we did to achieve that result, how we kept the systems underneath that throughput from falling over, and what we learned in the process. + +## Setting out to double productivity in a year + +In mid-August 2025, our founders sent a letter to the whole company, not just engineering: each of us should aim to double our productivity over the next year. It was addressed to individuals; the framing underneath it was a team sport, not an individual race, and nobody was supposed to read it as a contest with their teammates. The letter named Darragh Curran’s [Intercom 2x post](https://ideas.fin.ai/p/2) as the framing inspiration, explicitly. Eight months later, we started to measure and reflect. This post follows on from Darragh’s [retrospective on hitting 2x in nine months](https://ideas.fin.ai/p/2x-nine-months-later) and Kesha Mykhailov and Niamh Young’s [post on safely scaling AI auto-approval](https://www.intercom.com/blog/ai-is-approving-our-pull-requests-heres-how-we-made-it-safe/). We’re smaller than Intercom and a couple of years younger, and we’ve historically followed similar paths a few months behind them. + +The number of merges on a peak weekday to Honeycomb’s monorepo more than doubled from about 30 in early 2025 to about 74 in April 2026. This codebase doubled in sixteen and a half months, from approximately 0.97 million lines at the end of 2024 past 1.94 million in the first week of May 2026, and it sits at 2.1 million as of early July. The previous doubling had taken nearly three years. Most of that net-new code was co-written with or reviewed by AI somewhere in its history. + +## Bee-ing honest about that number + +1. It’s peak weekday, not calendar average. Wednesday no-meeting days are the most productive, both before and after AI; calendar average is roughly half of peak. That’s the peak. Don’t go looking for 70 PRs on a random Tuesday and conclude I lied to you. +2. It’s a floor, not a true count. One of our heaviest Claude Code users, by token volume, has zero AI-attributed git commits. They ran 22 sessions and 199 million tokens through Claude Code in 30 days, and git saw none of it, because their Co-Authored-By trailer is disabled at the tool level. If we can’t measure their AI usage, we can’t measure several other people’s either. +3. It’s entangled with org changes happening at the same time. In early January 2026, our founders sent a follow-up letter naming sharper strategic stakes than the August one had: rebuilding Honeycomb’s product surface and market posture to be AI-first, well beyond the original productivity target. The mid-January realignment toward greenfield, higher-AI-leverage work followed from that letter. We onboarded new engineers. We invested in platform engineering. Anyone who tells you a single factor caused their team’s gain is overstating it, including us. + +There’s a big leap from “AI tooling makes engineering teams genuinely faster” to “any team that adopts it will get the same result.” Most of this post lives in the gap between those two claims. The interesting story isn’t the multiplier; it’s what we did to keep things stable underneath it, what it cost us, and what we’re still figuring out. The short version, and the line I keep coming back to every time I’m invited on stage: [AI amplifies your existing practices](https://www.honeycomb.io/blog/shipping-is-your-companys-heartbeat-letter-from-cto). It can make a dysfunctional org more dysfunctional, or it can bring out the best in an org that already has high autonomy, ownership, and feedback loops. We weren’t setting out to prove that thesis; we took a challenge, and the thesis became visible after the fact. + +If you caught [my “AI is like chocolate” talk](https://www.honeycomb.io/blog/observability-day-san-francisco-future-ai-observability-is-bright), you know the bit: chocolate doesn’t belong on everything, too much of it in one sitting will make you sick, and no amount of chocolate substitutes for knowing how to cook. I gave that talk as a pessimist turned realist, and that’s still who I am. What changed between then and now isn’t the metaphor; it’s that the tooling and the practices around it got good enough that chocolate’s rightful place in the kitchen got bigger. It’s more versatile and forgiving than it used to be. The concessions later in this post, in “Where the skeptics are right,” are concessions I’m still making. Being a reformed skeptic doesn’t mean I stopped being one. + +You should take every number on this page, including ours, with a grain of salt. The methodology matters more than the magnitude. Keep your semantic bullshit defenses up against hype and plausible-sounding data, all the way through the rest of this post. + +# Join the masterclass with Liz Fong-Jones + +Six live sessions with Liz Fong-Jones + +turn Observability Engineering into practice. + +Starts August 3rd. + +## What the numbers do and don’t say about quality + +We haven’t had a spectacular AI-caused failure. No “AI deleted my database,” no “AI shipped code that corrupted user data.” That’s not luck; it’s not new either. It’s a result of designing defensively, whether the chaos agents be human or robot. Stacking agents on top of existing infrastructure with bulkheads between components, reviews, deploy trains, feature flags, and least-privilege access has meant outages are lower-impact, rather than either non-existent or uniformly critical-severity. If you don’t have that in place yet, that’s the thing to fix before you scale up AI usage, not after. A dropped database is a systems design issue: somebody skipped building the guardrail that would have caught it, regardless of who or what wrote the code. + +Incidents are growing in absolute count. From a 2024 baseline of about 18.5 incidents per quarter, Q1 2026 hit 32, 1.7x baseline, against PR throughput at about 2.5x baseline; for one quarter that looked sub-linear. Q2 didn’t hold: 53 incidents, 2.9x baseline, even as PR throughput itself held roughly flat once you back out the May freeze weeks and the January ramp-up swarm. Two quarters in, incident growth tracks change volume about as linearly as you’d expect. Change is the leading driver of incidents industry-wide (see the VOID report, Google’s DORA research), and we’re shipping a lot more of it. The law of large numbers caught up with us. + +What the data actually supports is two things: the absence of spectacular failures, plus some integrated second-order effects we’re picking up on. Broader AI-causation narratives tend to be self-flattering, whichever direction they point; AI is woven in deeply enough at this point that isolating it as a single cause is rarely a meaningful exercise. Attribution-by-cause is a fairy tale we humans tell ourselves to feel better about whatever stance we already hold. + +More changes shipped means more chances for a defect to land somewhere in the batch, at whatever the org’s baseline defect rate happens to be. We’re shipping far more change, so we get far more incidents, in roughly the proportion you’d expect. It’s too early to say whether AI assistance in debugging reduces incident severity once something breaks; in some cases it’s helped us find the root cause fast, in others it’s sent us chasing a confident, wrong answer instead. That volume showed up as real strain on the teams absorbing it, not just as a line going up on a chart. We’re leaning on better automated preflight checks, among other levers, to try to bend that curve back down. [The burnout risk that comes with capturing AI’s speed as pure output rather than sustainable pace](https://steve-yegge.medium.com/the-ai-vampire-eda6e4f07163) is a real one, and worth naming rather than assuming away. + +The metric I’d rather put on the wall isn’t PRs per day. Throughput is an input metric, not the product; it’s just the coarse measure we have today to demonstrate step-change. Throughput going up while user outcomes plateau or decline is, by definition, enshittification. + +## What does “AI contribution” actually mean? + +There’s no single “AI percentage.” There are at least four different denominators, and you have to be precise about which question you’re asking before you quote a number at anyone. + +(Vendor dashboards will happily hand you an “AI-influenced PRs” number with a confidence toggle. Set to loose confidence, ours cheerfully reported a majority of PRs as AI before we’d done any of the work below. Don’t trust low confidence. The numbers in this post come from our own git history and telemetry, calibrated by hand.) + +These are calibrated floors, not point estimates. We got there by layering three corrections onto git’s raw signal: local branch trailers (245 PRs whose AI attribution was stripped by squash-merge), GitHub-API branch trailers (117 PRs where branches were deleted post-merge but we recovered the trailers via GraphQL across all 7,952 PRs), and a smell-test telemetry override (112 PRs from engineers with zero git AI attribution but heavy Claude Code session telemetry, gated per month against their actual session activity). + +The 95% engineer-level adoption figure, which is closer to our gut estimates, reconciles naturally with the lower PR-level (63%) and line-level (75%) floors: high adoption, selective per-PR use. We can measure floors from git history. We can’t measure ceilings. Being clear about the difference matters more than the specific numbers do. + +![AI-attributed floor on top of the unattributed baseline](https://www.honeycomb.io/_next/image?url=https%3A%2F%2Fcdn.sanity.io%2Fimages%2F927dxq0h%2Fproduction%2Fd82176f5a5aa92a5062fa7d586d20b8aab99f5da-1650x926.png&w=3840&q=75) + +*Surviving lines at HEAD since 2024. The red band is a floor; lines from PRs with AI attribution somewhere in their history. The blue band is “no attribution,” not “no AI.”* + +## What happened + +The adoption curve at Honeycomb has three distinct phases. Each one was driven by something different. + +The first phase, April through September 2025, was sustained low-rate experimentation. About one or two new engineers adopted Claude Code per month, with support but no mandate; engineers were exploring on their own, at least until mid-August, when the founders’ letter landed and the picture blurs. AI’s share of new lines climbed modestly, peaked around 18% in June, then drifted back down to 8% by October as the honeymoon wore off. This is what “leadership opened the door” looks like in practice: not a discrete push so much as a sustained green light. The dip in the second half of 2025 was evaluation, not failure: engineers tried Claude with the model available at the time, decided it didn’t yet justify the friction, and pulled back. If the catnip is rotten, herding cats to eat the catnip is even more difficult. + +The second phase began in October 2025 with a Claude Code harness improvement. Seven new adopters that month, with no model release behind it; tooling quality alone moved the needle. November and December were a pause (although some engineers used the holiday break to try the tools in their personal capacities). Then Claude 4.5 landed in late December, and adoption picked back up in January 2026 with eight new adopters, and AI’s share of new lines climbing to 18% by January. + +![AI commits per month vs cumulative engineers using AI](https://www.honeycomb.io/_next/image?url=https%3A%2F%2Fcdn.sanity.io%2Fimages%2F927dxq0h%2Fproduction%2Fe43cdf5ebadabc6749b9bd8ff0cbdb729ef8aa34-1757x812.png&w=3840&q=75) + +*Engineers with at least one AI-attributed commit, cumulative. Each step change lines up with a capability event, not a flat schedule.* + +Opus 4.6 launched on Thursday, February 5, 2026. The session-rate takeoff in our Claude Code telemetry lands on February 11-13: the first full work-week after a Thursday launch with a Friday-and-weekend bake-in. Distinct monthly Claude Code users stayed nearly flat across January, February, and March (59, 63, 64). Sessions per user 3.3x’d over the same window (21, 35, 70). AI’s share of new lines went from 18% in January to 46% in February to 65% in March. + +We’ve been calling this confidence to delegate. The same engineers, with the same tooling, with access to the same model family, simply changed how they used it. They stopped consulting or closely monitoring each step, and started delegating. The lift is per-user-intensity, not headcount enabled. + +![Engineer count vs sessions vs PRs before and after Opus 4.6](https://www.honeycomb.io/_next/image?url=https%3A%2F%2Fcdn.sanity.io%2Fimages%2F927dxq0h%2Fproduction%2Fe41a18604c09c1d045e90852a85677fd901c6d67-1157x847.png&w=3840&q=75) + +*The Q1 proof, one magnitude per panel: engineer count barely moved, sessions per engineer exploded, and committed PRs-by-model shows the delegation landing on Opus 4.6 specifically.* + +While the median engineer’s PR throughput grew about 45% relative to its pre-February baseline (call that baseline 1.0x), the top of the distribution moved much further. P75 went from about 1.8x baseline to about 2.9x, roughly +64% in relative terms; the weekly maximum across our active engineers went from a 3.2x-6.4x baseline range pre-February to a 7.7x-12.3x range in April. The floor barely moved at all (P25 went from about 0.5x baseline to about 0.8x). The AI uplift is concentrated at the top end of the distribution. It is not a universal rising tide that automatically boosts every engineer, and that’s okay, because not all engineering is in the bucket of things AI accelerates. As Charity says, the closer to touching bytes on disk you are, the more cautious you need to be about reviewing *everything* with a paranoid lens. + +We’d never had a dozen engineers a week shipping 7+ PRs each in our entire 10-year history, until March 2026. Pre-February, two to four engineers a week hit seven or more PRs; in February that was three to seven, in March seven to twelve, in April eight to sixteen. New, and now routine. And the engineers at the top aren’t a stable cast; the March-selected and April-selected top-twelve cohorts only overlap by about half. It’s a rotating cast riding the new ceiling, not a handful of superusers carrying everyone else. + +![Weekly merged PRs](https://www.honeycomb.io/_next/image?url=https%3A%2F%2Fcdn.sanity.io%2Fimages%2F927dxq0h%2Fproduction%2F90d3ef45b4639529a7c2ae95025f66be5cf1f2ef-996x926.png&w=3840&q=75) + +*Weekly merged PRs (total/AI-attributed/no-attribution) and the per-engineer weekly distribution for each cut, September 2025 through the week of June 22. The top tail past 20 PRs per engineer-week appears in March and persists; the May dent is the freeze weeks, not decay. Bots excluded, including the autobot.* + +Peak-weekday non-AI merges held roughly steady at 25-30 across the entire window; humans didn’t slow down on their best days to make room for AI. But at the org-wide weekly level, non-AI commit-hash volume declined notably, from about 120 PR commits a week pre-February to about 60-80 a week through February to April, while AI commit-hash volume grew from near-zero to 150-200 a week. + +So engineers didn’t slow down on the days they were shipping; they shipped fewer human-attributed PRs in aggregate while shipping many more AI-attributed ones. Some of the +44 peak-weekday delta is genuinely new capacity. Some of it is substitution, where work that used to be human-attributed is now agent-attributed because the same engineers shifted to driving with AI rather than typing by hand. We can’t cleanly separate lift from substitution without a controlled experiment we don’t have. + +## May, June, and a third workflow + +For this blog, we pulled a fresh data cut through June 28, rather than waiting the six months we’d originally planned since the May and June presentations. One methodology note before the numbers: this refresh also excludes mechanical bots (e.g. Dependabot) from every throughput denominator, something the April numbers above didn’t do. So the figures in this section aren’t a perfectly clean continuation of the ones above them. Same discipline as the rest of this post: check the methodology before you trust the magnitude. + +On that basis, weekday-average merges went 38.0 in March, 47.9 in April, then dipped to 36.0 in May before climbing back to 41.9 across the four complete weeks of June. The May dip isn’t engineers slowing down; it’s a supply constraint. May carried an intense marketing push plus merge freezes around [Innovation Week and O11yCon SF](https://www.honeycomb.io/resources/topic/innovation-week). June rebounded as soon as the freezes lifted, and peak-day merges hit 70 again on June 18, matching and sustaining April’s peak. + +The more interesting news isn’t the wobble in the average. It’s that a third category of work showed up entirely. + +`honeycomb-autobot[bot]` landed its first commit on main on April 23. It’s Claude Code on AgentCore, dispatched from RWX (the same CI substrate from earlier in this post) and traced by Honeycomb, triggered from a Linear issue or an @honeycomb-autobot mention on a review, with no human anywhere in the commit-generation loop, only the review loop. That’s categorically different from “AI-assisted coding.” Human-in-loop coding still means a person is driving the session and choosing what to commit. The autobot doesn’t have anyone in that seat at all, but instead back-loads the work onto the review cycle where work is more mechanical and a human feels confident going hands-free during the actual coding. + +Three months of data on it: 3 autonomous merges in April (0.3% of the month), 13 in May (1.7%), 70 across the four weeks of June (8.4%). Over the same window, human-authored merges with zero AI attribution kept shrinking, down to 211 in four weeks of June against a 2025 baseline in the 400s a month, while total throughput held at roughly twice the 2025 baseline the whole time. Same pattern as everywhere else in this post: substitution, not addition. + +![Merges and lines by workflow](https://www.honeycomb.io/_next/image?url=https%3A%2F%2Fcdn.sanity.io%2Fimages%2F927dxq0h%2Fproduction%2Ffa5d522b839bcdfd30715b1fddd8bc0399cfa399-1243x928.png&w=3840&q=75) + +*The three-way split, January 2025 through the four full weeks of June 2026, mechanical bots excluded. The assisted band is a floor; the autonomous band is exact, because bot authorship is self-evident.* + +The adoption shape looks like the rest of the story too, not a power-user phenomenon. Of the 96 autobot squashes on main through July 6, 63 carry an explicit “Triggered by” line naming 23 distinct engineers; one of the eng enablement leads who co-authored autobot is the heaviest user at 18, and everyone else in the tail is 2 to 4 each. And 55 of those 96 squashes carry no Claude co-author trailer at all, which means the same trailer-based blind spot from the “Bee-ing honest” section up top shows up here too. If we counted the autobot’s work by trailer the way we count human-driven work, we’d have missed 57% of it. We count it by author identity instead, which for a bot account is exact rather than a floor. But it’s a reminder that every attribution method has exactly one failure mode it’s blind to, and you only find out what it is by checking, not by assuming it doesn’t have one. + +The autonomous workflow rides the frontier model the same way human delegation does: 30 of 33 model-tagged autobot squashes in the four weeks of June cite Opus 4.8. And it isn’t yet a lines-of-code story. Autonomous work has added roughly 5,500 lines total since April, about 0.3% of the codebase’s current size. (Added, not surviving; it’s too early to measure how many of those lines are still alive at HEAD.) Right now this is a merge-count phenomenon, not a codebase-composition one. It’ll be worth watching whether that changes. + +The codebase kept growing underneath all of this. HEAD was at 1.88 million lines in April; by July 6 it’s 2,096,286, more than double the 972,000 lines at the end of 2024. Net-new code since end-2024 is now majority AI-attributed for the first time: 599,000 of 1,124,000 added lines, at least 53%, up from 41% in April. June’s line-level floor, at least 82.6% of new lines AI-attributed, is the highest month on record, ahead of April’s 75%. The projection in the April data I presented on-stage, that AI would cross a third of HEAD “around end of 2026,” turned out to be conservative; the current slope puts that closer to September, with half of HEAD by roughly mid-2027. + +One number needs its own caveat rather than a triumphant read. Human-attributed lines surviving at HEAD actually ticked down slightly, from 1,509,000 in April to 1,497,000 in July. That’s not a clean “AI replaced human code” story. Some of it is genuine replacement; some of it is recalibration retroactively reclassifying PRs that were originally counted as human, as we keep finding hidden AI attribution in old PRs. Don’t read a precise story into that number. Read it as more evidence that the floor keeps rising as we get better at measuring it, which has been true of every number in this post so far. + +Engineers with at least one AI-attributed commit: 70, up from 64 in April. That’s broadening, not just the same 50-plus people going faster. And the frontier-model succession kept stair-stepping exactly the way it did in February: Opus 4.6 gave way to Opus 4.7, which gave way to Opus 4.8 (366 of June’s model-tagged PRs, against 31 for 4.7 and 29 for 4.6, the same one-month displacement pattern each time). Claude Fable 5, the newest Mythos-tier model, shows up in June’s trailers too, on 34 PRs. Sonnet 5 hasn’t landed a merged PR yet, since it only just launched, but it’s already showing up on PRs out for review, which tends to be the leading indicator before it shows up in this table. And the autonomous workflow doesn’t relax the one constraint that’s held the whole way through this post: a person still has to trigger it, via a Linear ticket or an @-mention. Our human names have stopped showing up on the commit messages. The only adoption ceiling is still how many engineers choose to reach for it. + +![AI-attributed PRs per week by model](https://www.honeycomb.io/_next/image?url=https%3A%2F%2Fcdn.sanity.io%2Fimages%2F927dxq0h%2Fproduction%2Ffb5f67d514514013fc8bf8a4d131a7edaf157253-1757x760.png&w=3840&q=75) + +*The succession, extended through June. Each Opus release displaces its predecessor in committed work within about a month of arriving; the February pattern wasn’t a one-off.* + +## Possible (overlapping) explanations + +We can name several factors that line up with the inflection. None of them, on its own, explains the curve, and we can’t isolate which mattered most without a controlled test we can’t run in retrospect. What follows is a list of overlapping contributions our own team has flagged, not one single cause dressed up as several. + +**Frontier-model capability:** Opus 4.6, in February 2026, was the specific release that produced the session-rate jump. Earlier Opus releases (4.1 in August 2025, 4.5 in December) shifted the floor without producing a takeoff. The rest of the model family in the same window, Sonnet 4.6 and Haiku 4.5, didn’t move the committable-delegation needle in our telemetry: Sonnet 4.6 launched February 17 with full Honeycomb adoption (41 distinct users) but produced only 35 commits in March and 27 in April, against Opus 4.6’s 232 and 127. It was frontier-model capability specifically that crossed the threshold from consultation to delegation, not “any new model.” + +**Leadership signaling, twice:** August 2025’s letter set the explicit “experiment, take time to figure out what AI does for your work” frame. January 2026’s follow-up named sharper strategic stakes: rebuild the product surface and the market posture for AI-first. The mid-January realignment toward greenfield, higher-AI-leverage work followed from that. Without either signal, I doubt the throughput curve looks the same. + +**Tooling that improved over months:** The October 2025 Claude Code harness improvements produced a noticeable adoption step on their own, with no model release behind them. Each release of the tooling added something some engineer needed before they could delegate. Models alone don’t explain the curve; the wrapper around the model matters just as much as the model does. + +**Substrate already in place:** Continuous delivery, code-ownership practices, fast CI, blameless incident analysis, observability that links shipped code back to the PR that created it: the practices the rest of this post is about. AI work landed in an org that already had the substrate to absorb it. We can’t run the counterfactual, but the practices section below is our best account of why this didn’t go badly. + +**Cumulative engineer-level expertise:** From April through September 2025, one or two engineers a month adopted Claude Code. By February 2026 we had a base of fifty-plus engineers who’d been using it for months and could mentor everyone else. That cohort effect is hard to pin to any single date; it’s the gradual accumulation of in-house expertise that the February takeoff drew on. + +These factors overlap, and we can’t say any one of them in isolation would have gotten us here. Leadership signaling probably amplifies tooling improvements. Substrate makes it possible for engineers to share what they’re learning. Frontier-model capability matters more in an org that already has the substrate and the cohort to use it well. The factors compound rather than substitute for one another. Anyone telling you a single factor caused their team’s gain is overstating it. Including us. + +## We followed Intercom, with caveats + +Three things are worth crediting in Intercom's playbook. + +They set a more realistic and achievable 2x goal, not the 10x-and-up claims that were common currency during the 2024 hype cycle. Discipline matters when the temptation is to oversell what AI delivers. We wanted some of that discipline for ourselves. + +They were transparent about both the methodology and the org-shape changes that came with the gain. The substrate Intercom built, a Claude Code plugin marketplace with 153 contributors representing 31% of their R&D org and 267 skills, is genuinely platform infrastructure that ships its own product. They spun up a dedicated team, team-2x, to build it. We’re smaller and a couple of years younger, but we’re building toward something in the same shape, at our scale. + +They engaged their auditors, Schellman, early, before scaling auto-approval, to confirm that the evidence trail an AI-approved PR produces is the same evidence trail an auditor expects from a human-approved one. The “who” changes. The “what” doesn’t. That’s a model worth following: build for safety first, and compliance follows from it. + +Where we differ is that we’re at zero auto-approval today, and they’re at 19.2% as of April. That’s deliberate sequencing on our part, not a sign that we're “behind.” Before you can safely scale auto-approval, which is automating the bug-catching half of PR review, you need substantial substrate underneath it. The compliance question is necessary but not sufficient; the technical preconditions sit underneath the procedural ones. You need codified rules in CLAUDE.md and skills that an auto-review agent can actually verify against; MCP-mediated dev-loop access so the agent reviews against the same context a human would have had (design intent, ticket history, production behavior); fast and AI-legible CI so the verification loop closes quickly; closed-loop production observability that links shipped code back to the PR that created it; and a dissemination layer so humans stay aware of what’s shipping even when they’re not gating it. + +A 19% auto-approval number means radically different things at an org that invested in the substrate before turning the switch versus one that just turned the switch. Intercom’s number is downstream of substrate they built first. A company that turns on “auto-approve PRs under 20 lines” without equivalent substrate will report the same number, but it’s measuring rubber-stamping against weak constraints, not safe automation against strong ones. Different point on the same trajectory. + +The headline finding from Intercom’s auto-approval post is the most striking parallel to our own data. They report AI-authored backend code reverting at 0.53% and AI-authored frontend code reverting at 0.22%, against human-authored revert rates of 5.39% and 2.00% respectively. It’s worth naming the selection effect here: if the easier, lower-risk changes increasingly get auto-approved, humans are left reviewing the harder residual cases, which would push human revert rates up for reasons that have nothing to do with humans getting worse at their jobs. Their downtime from breaking code changes dropped 35% even as deployment frequency doubled. Theirs is a per-PR claim about strict, decomposed, sub-agent-driven review against an Intercom-specific guidance flywheel. Ours is a per-quarter claim about severity, not volume: incident count is tracking change volume about as linearly as you’d expect, but we haven’t had a spectacular AI failure, and our guardrails are aimed at containing how bad any one incident gets rather than pretending we can hold the count flat. Different denominators, same direction. Both posts are pushing back on the naive “more code, more failures” intuition, from different evidence. + +We had a private conversation with the Intercom team in early May 2026. They hit the same February 2026 inflection point we did. The Opus 4.6 unlock was the gut-call attribution from multiple Intercom engineers, though they noted it was almost impossible to disentangle from their internal mandates and team-2x activity in the same period. They’re partnering with a Stanford research group to try to isolate the variables, and even that group’s initial pre-January analysis missed the inflection entirely. The world’s most data-rich org on this exact question is paying academic researchers to help them figure it out, and they still don’t know. + +That independent-org corroboration is the cleanest natural experiment either of us has. Two separate orgs, same month, same uncertainty about cause, same direction. That’s worth more than either of our individual analyses on its own. It’s also worth weighing against the broader base rate: [DX’s longitudinal study across roughly 400 companies](https://newsletter.getdx.com/p/ai-productivity-gains-are-10-not) found AI usage up 65% translating to only about 8% more PR throughput on average. We’re the outlier case here, not the median one. + +Here's the truth: Nothing I've written here will help you if your underlying org isn't already healthy and functional. AI just amplifies what you're already doing. [Read part 2 of this blog series to see what we learned.](https://www.honeycomb.io/blog/ai-amplifies-existing-practices-lessons-ai-first-strategy) + +## AI Influence Level disclosure + +This post: AIL-3.0 (substantial AI involvement, human steering on every load-bearing call). The data analysis, custom git-of-theseus extensions, commit-history trawling, calibration overrides, the merge-rate and incident charts, was AI-assisted. The July refresh, the May-June numbers and the autonomous-workflow analysis added above, was pulled with Claude Fable 5. The slide deck this post derives from was composed with AI assistance, and AI assistance was used to reformat the slide bullet points and speaker notes into essay form. The voice and tone polish was done in a separate Claude project tuned to my writing style, followed by a very extensive manual editing process where I further added or changed at least 20% of the words. + +Strategic decisions, data interpretation, and judgment calls about what to keep and what to cut are mine. AI helped me move faster on a deadline; it didn’t supply the substance. The talk and this post are themselves an example of the same 2x story they describe: weeks of human work, AI-assisted, not 10x. The substance wouldn’t exist without me, and the level of polish wouldn’t exist without AI. + +AIL framework: [danielmiessler.com/blog/ai-influence-level-ail](https://danielmiessler.com/blog/ai-influence-level-ail). Illustrations in the original talk: AIL-0, by [bbghost.bsky.social](https://bsky.app/profile/bbghost.bsky.social). Art should be made by artists, not machines. Illustrations in the blog by our amazing design team. + +*Sources and context:* [Fin/Intercom 2x post](https://ideas.fin.ai/p/2x-nine-months-later);[Fin/Intercom AI PR approval safety post](https://www.intercom.com/blog/ai-is-approving-our-pull-requests-heres-how-we-made-it-safe/);[Honeycomb-Intercom case study](https://www.honeycomb.io/resources/case-studies/how-honeycomb-helped-intercom-observe-and-operate-fin-ai);[Emily Nakashima on AI-amplified engineering leadership](https://www.aviator.co/podcast/enineering-leadership-ai-emily-nakashima). This post is adapted from the[talk of the same name](https://leaddev.com/software-quality/30-to-70-prs-a-day-how-we-managed-to-not-wreck-our-systems), delivered at Sydney Tech Leaders and LDX3 London in 2026. diff --git a/sreweekly/markdown/528/06-you-ve-just-had-an-incident-what-next.md b/sreweekly/markdown/528/06-you-ve-just-had-an-incident-what-next.md new file mode 100644 index 00000000..107e674a --- /dev/null +++ b/sreweekly/markdown/528/06-you-ve-just-had-an-incident-what-next.md @@ -0,0 +1,73 @@ +# You’ve (Just) Had an Incident. What Next? + +- **期号**: SRE Weekly Issue #528(2026-08-02) +- **作者**: Karan Nagarajowda — Uptime Labs +- **链接**: https://www.uptimelabs.io/articles/incidents-and-recovery + +## 简介 + +> Here’s what I’ve learned about what actually happens inside a team when something breaks, how teams often reflect, and how we think about accountability without losing a blameless culture. + +## 正文 + +![]() + +### Ready to make incident response your competitive advantage? + +See how Uptime Labs builds provable, scalable incident response capability across your organisation. + +I've run enough major incidents to know that the first hour rarely goes the way people expect. Here's what I've learned about what actually happens inside a team when something breaks, how teams often reflect, and how we think about accountability without losing a blameless culture. + +## The first hour of an incident: what's really happening + +![Incident timeline illustration reading detect - assess - declare & categorise - communicate - diagnose - resolve - review](https://cdn.prod.website-files.com/69eb654f1a323d235a57f701/6a5dd7f8f17f9bc100c5a6bd_Screenshot%202026-07-17%20at%2018.14.33.png) + +Logistically, after detection, people try to understand what's actually happening, then work out whether how to move forward. Note that incidents do not progress through stages in a linear way. You may come back to assessing severity or revisit assessment based on new information that emerge as incident progresses. + +Similarly, it's also complex from an emotional perspective. Individual engineers often wonder if they broke something. Some people see a symptom and immediately form a theory i.e. a gut instinct about the cause - and then start hunting for evidence to support that theory rather than staying open-minded. Meanwhile, managers want constant updates, and what actually happens in Slack is information overload. + +But in my view, the real challenge in the first hour isn't technical; it's cognitive overload. If you [divide an incident into stages](https://www.uptimelabs.io/template/incident-responder-workflow), the first part of the time should be owned by the incident commander, and their job *isn't* to solve the problem. It's to reduce chaos: create the space for people to pick up tasks and start investigating, make decisions explicit, and keep communication flowing. Those first minutes - or even hours, in a bigger incident - are never about fixing the system. They're about making sense of a surprising situation, recruiting people who can help and creating sufficient psychological safety that people will voice their theories and be explicit about certainty of premises of the theory. That's what separates a [good incident commander](https://www.uptimelabs.io/template/training-incident-responders-scaling-teams). + +## A word about pressure + +I think back to a story from earlier in my career that says more about pressure than any framework could. I was working through a major incident next to a colleague who was technically very strong. Our manager - who was visibly under stress - came up behind us and demanded we check the logs. The pressure of being watched and barked at was so distracting that my colleague forgot the basic syntax of the `view` command. Our manager ended up spelling it out: "v–i–e–w, and then the file name." + +It's a small moment, but a telling one. Technical skill isn't the bottleneck when the room is hostile. A junior engineer can have all the right preparation, the right runbooks, the right paired senior and still fold if the environment around them is built on pressure and intimidation rather than support. + +## What teams could overlook after + +People naturally want to answer "*why did this happen?*" Uncertainty is uncomfortable, and immediately after an incident there's a lot of it. So people jump to ‘why’, but I think that's the wrong question, because it sends you straight towards root-cause analysis. + +As we explored in [The Technical Foundations of Incident Response](https://uptimelabs.io/articles/technical-resilience-incident-response/), the right sequence is: can we stop the customer impact - stop the bleeding? Then, can we stabilise the system? Then, can we preserve the evidence? Only then do you investigate - that's where the actual learning happens. + +However, when a team implements a change and encounters errors, fixation can quickly set in. They observe memory errors and become convinced it's a memory leak, focusing all investigation efforts on proving that diagnosis. The problem is that this fixation causes them to lose sight of the primary incident response objective: restoring service. During an incident, multiple decision paths are available (rolling back the change, restarting the service, modifying configuratio) each with different risk and learning profiles. By pursuing only one diagnostic avenue in parallel, teams sacrifice their ability to understand what actually happened. Gathering information before committing to a single response path, and considering what each approach might reveal, is as important as speed in bringing systems back online. + +The second accident is making several big decisions in parallel: restarting, scaling up, changing configuration, with different people acting independently. I wouldn't say the process needs to be strictly sequential, but it makes it much easier to asses impact of each change if you try one thing, gather information, and only then move to the next. Doing three or four things in parallel might bring the system back faster, but it destroys your ability to learn afterwards what actually happened. Preserving the evidence is just as important as restoring the service. + +## Accountability without losing a blameless culture + +A ‘no-blame’ culture in incident management should not mean removing human accountability or glossing over the decisions people make during crises. Rather, it means resisting the urge to simply label events as ‘human error’ and moving on. Every incident involves decisions made by individuals who bear responsibility for those choices, but understanding *why* they made them is critical. + +The goal is to reconstruct the context in which decisions were made: what information was available at the time, what pressures and constraints existed, and what risks seemed apparent or hidden. When we skip this deeper analysis to avoid blame, we sacrifice the insights that could prevent future incidents. Accountability and learning are not opposites; they work together when we focus on understanding the decision-making environment rather than punishing the decision-maker. + +A good postmortem embraces the human element; one of the the key skills of Postmortem (incident review ) is to conduct in a way that no one is uncomfortable. Part of this means avoiding reducing a complex failure down to one person's mistake. That's what a blameless culture actually means: not pinning a complex failure on one person or team, as we discuss in [our regulatory incident response piece](https://uptimelabs.io/articles/regulatory-incident-response/). Accountability, on the other hand, is about improving future outcomes, not assigning guilt. + +So in a postmortem, the questions focus should be on, *"What set of circumstances led to the incident? What information was available to the human operator at the point of the decision making? Why that decision made sense to them?" .* + +Definitely not *"who did it, or why did they do it?"* If someone skipped a checklist and that caused a major issue, the question isn't *"why did they skip it"* - it's "*was there pressure that made skipping it possible? What barriers should have stopped that?"* It's about whether the system allowed the mistake, not about the person. Accountability, meanwhile, is about owning the future outcome - again, without assigning guilt. + +## Who owns the learning? + +It's the incident commander's responsibility to make sure the learning happens, but I don't own all of the learning myself. The engineering team should [own the technical timeline](https://uptimelabs.io/articles/incidents-will-happen-are-you-actually-prepared/). Monitoring and operations should look at the alerts and the response coordination. Customer support should explain the customer experience during the incident. My job as incident commander is to make sure all of that comes together. + +One of the most valuable questions I ask isn't *"what failed?"* It's *"what made the incident harder to resolve than it needed to be?"* That's the better question, and it usually surfaces poor documentation, confusing dashboards, unclear ownership, missing alerts and communication gaps. That's what comes out of these discussions. + +## Add Uptime Labs to your post-incident learning + +None of this is instinctive. Reducing chaos before fixing the system, resisting the pull towards *"why",* asking what made an incident more difficult to resolve rather than who's to blame - these are habits people get better at by practising them, *not* by reading about them and hoping they’ll stay in your head while everything is on fire. + +That's exactly why I helped build Uptime Labs. Our drills put your team through the same ambiguity, incomplete information, and communication pressure a real incident throws at you, minus the real-world stakes. You find out how your team actually behaves in that first hour before it costs you a customer. Each drill comes with a personalised report designed to support ongoing skills development. + +If you want to see what that looks like in practice, [try a drill](https://uptimelabs.io/try-uptimelabs) or get in touch to [book a demo](https://uptimelabs.io/). + +![](https://cdn.prod.website-files.com/69e0a463268ba34093f8b1cb/69f225d89c7c20b5a5ec3c6e_Rectangle%2039651.png) diff --git a/sreweekly/markdown/528/07-an-sre-response-to-datadog-s-state-of-ai-engineering-2026.md b/sreweekly/markdown/528/07-an-sre-response-to-datadog-s-state-of-ai-engineering-2026.md new file mode 100644 index 00000000..b27cdc14 --- /dev/null +++ b/sreweekly/markdown/528/07-an-sre-response-to-datadog-s-state-of-ai-engineering-2026.md @@ -0,0 +1,174 @@ +# An SRE Response to Datadog’s State of AI Engineering 2026 + +- **期号**: SRE Weekly Issue #528(2026-08-02) +- **作者**: Ajay Devineni — DZone +- **链接**: https://dzone.com/articles/agent-sprawl-production + +## 简介 + +> agent sprawl is now a production reliability crisis, and the SRE discipline does not yet have governance frameworks for it. + +## 正文 + +- + ![](https://dz2cdn1.dzone.com/themes/dz20/images/dz-postarticle.svg) [Post an Article](https://dzone.com/content/article/post.html) +- + [Manage My Drafts](https://dzone.com) + +# Agent Sprawl Is Your Next Production Incident: An SRE Response to Datadog's State of AI Engineering 2026 + +Datadog published the State of AI Engineering 2026 report. Read it. It's the most comprehensive look at AI in production available now. + +Join the DZone community and get the full member experience. + +[Join For Free](https://dzone.com/static/registration.html) + +Datadog published the [State of AI Engineering 2026 report](https://www.datadoghq.com/state-of-ai-engineering/)— real telemetry from over a thousand production environments. Read it. It is the most comprehensive look at AI in production available right now. + +I want to respond from the reliability engineering perspective, because the data reveals a problem the report names but doesn't fully resolve: agent sprawl is now a production reliability crisis, and the SRE discipline does not yet have governance frameworks for it. + +## + +Three findings stand out from an SRE perspective: + +**Framework adoption doubled year over year**. LangChain, LangGraph, Pydantic AI, Vercel AI SDK — up from 9% of organizations in early 2025 to nearly 18% by 2026. Services using agentic frameworks: more than doubled. + +**70%+ of organizations run three or more models**. The share running more than six models nearly doubled. Teams are building model portfolios rather than committing to a single provider. + +**Teams add models faster than they retire them**. Datadog calls this "LLM tech debt." Each overlapping model introduces its own quality, latency, and cost profile. The report is explicit: this becomes a governance problem. + +These three findings combine to describe an environment growing faster than it can be governed. I call this **Agent Sprawl**. + +## Defining Agent Sprawl + + **Agent Sprawl** — the condition where AI agent infrastructure complexity (frameworks, models, tool layers, orchestration patterns) grows faster than your ability to measure and govern its reliability. + + +It is structurally identical to the microservices sprawl problem SRE teams faced between 2015 and 2020. Teams added services faster than they added SLOs. The result: production incidents nobody could attribute because the dependency graph was too complex to observe. + +Agent Sprawl has three specific manifestations: + +### + +When you add LangChain, LangGraph, or any orchestration framework, it adds steps and paths you did not write — retry logic, fallback handlers, context window management, tool routing. All of this happens between your application code and your observability layer. + +Your SLIs measure at the application boundary. Framework-added calls are invisible. + +This means your Tool Invocation Efficiency (TIE) baseline — tool calls per task completion — is measuring a mix of your agent's behavior and your framework's behavior. When you upgrade the framework, both change simultaneously. You cannot separate them. + +In practice, across regulated production environments I've studied, TIE baselines can drift 30 – 40% after a framework major version upgrade with no corresponding change in the agent's task logic. The baseline shift looks like agent degradation. It's actually framework overhead. Teams spend hours on a false RCA. + +**The fix**: Instrument at the framework output layer, not the application layer. Capture tool invocations after framework processing. Then freeze your TIE baseline before any upgrade and compare shadow traffic before promoting. + +### + +70% of organizations running 3+ models means 70% have at least two additional SLO ownership gaps they haven't acknowledged. + +[SLOs](https://dzone.com/articles/what-are-slos-slis-and-slas) are set once — typically when the first model is deployed. As models 2, 3, 4, 5, 6 are added for specific task classes, latency profiles, or cost tiers, nobody revisits the SLO ownership model. Models run in production with no named owner, no baseline, no error budget. + +When model 3 degrades, there is no owner to page, no baseline to compare against, no runbook to execute. The degradation surfaces as a customer complaint, not an alert. + +**The fix**: Treat every model in your fleet like a microservice. Each model gets: a named owner (not a team — a person), a task-class-specific SLO, and a 30-day observation baseline before the SLO is enforced. + +### + +Deprecated models running in agent chains create silent compatibility risks. When a provider announces deprecation, teams with models buried inside multi-step chains often miss the migration window. The model ages. Safety training falls behind. Decision Quality Rate declines slowly — too slowly to trigger a threshold alert — until accumulated drift surfaces as a production incident. + +**The fix**: Treat model deprecation notices the same way you treat dependency CVEs. Automate alerts at 60, 30, and 7 days before end-of-life. Build the migration ticket at announcement time, not at expiry. + +## The Governance Framework Agent Sprawl Needs + +### + +Before you can govern sprawl, you need to know what you're governing. Maintain a living inventory with, for each component: framework and version, model(s) used, task classes handled, named SLO owner, current TIE/DQR baselines, and deprecation dates. + + Python + + + +``` +from agentsre.sprawl import AgentFleetInventory, FleetComponent, ComponentType +inventory = AgentFleetInventory() +inventory.register(FleetComponent( + component_id="anthropic.claude-sonnet-4-6", + component_type=ComponentType.MODEL, + agent_id="payment-processor", + task_classes=["payment-routing", "fraud-detection"], + slo_owner=" +``` +[\[email protected\]](https://dzone.com/cdn-cgi/l/email-protection)", # named human — not a team + baseline_established_at="2026-04-01", + deprecation_date="2027-06-01", + last_slo_review="2026-04-01", + current_tie_baseline=2.4, + current_dqr_baseline=91.2, +)) +report = inventory.quarterly_review_report() +print(f"Fleet governance score: {report['fleet_governance_score']}/100") +### + + Python + + + +``` +from agentsre.sprawl import FrameworkVersionGovernance +gov = FrameworkVersionGovernance( + tie_drift_threshold=1.15, # block if TIE drifts >15% + dqr_drift_threshold=0.85, # block if DQR drops >15% + min_shadow_samples=50, +) +# Before upgrade: snapshot production baseline +gov.snapshot_baseline( + agent_id="payment-processor", + task_class="payment-routing", + framework_version="langchain-0.2.x", + tie_values=production_tie_samples, + dqr_values=production_dqr_samples, +) +# After 48hrs shadow traffic: +result = gov.evaluate_upgrade( + agent_id="payment-processor", + task_class="payment-routing", + production_version="langchain-0.2.x", + shadow_version="langchain-0.3.x", +) +if result.decision == UpgradeDecision.BLOCK: +    rollback()   # framework added hidden overhead — don't promote +``` +### + +The review should take 30–60 minutes per quarter. For every model in fleet: + +- Verify named owner exists +- Verify baseline is current (< 90 days old) +- Check deprecation schedule against provider announcements +- Review TIE per-model — models with rising TIE relative to task class baseline are drifting + +Models scoring below 70 on the governance health score are flagged as governance debt requiring a 30-day remediation window. + +## The Datadog Report's Implicit Challenge + +The State of AI Engineering 2026 describes an industry in rapid expansion. What it does not fully resolve is the SRE question: who governs all of this, and what does that look like in practice? + +The SRE community has solved exactly this class of problem before — in distributed systems, in microservices, in cloud infrastructure. The discipline already exists. It needs to be applied to the AI agent layer now, before agent sprawl becomes agent chaos. + +The Datadog data tells us the window is closing. Framework adoption doubles in a year. Multi-model fleets become the norm. Model debt accumulates. + +Build the governance layer before the production incidents start. + +## Resources + +- Open-source implementation: [[https://github.com/Ajay150313/agentsre](https://github.com/Ajay150313/agentsre) ] +- LinkedIn discussion: [[https://www.linkedin.com/posts/ajay-devineni_agenticai-sre-reliability-ugcPost-7455786901673902080-BCRM?utm_source=share&utm_medium=member_desktop&rcm=ACoAACIp55QBRGVmAcEbf0D-1PaR5vEbm2yMcJU](https://www.linkedin.com/posts/ajay-devineni_agenticai-sre-reliability-ugcPost-7455786901673902080-BCRM?utm_source=share&utm_medium=member_desktop&rcm=ACoAACIp55QBRGVmAcEbf0D-1PaR5vEbm2yMcJU) ] + +What's your biggest agent sprawl challenge right now? + + AI + Engineering + Site reliability engineering + + + Opinions expressed by DZone contributors are their own. + +Comments diff --git a/sreweekly/markdown/528/08-don-t-add-a-read-replica-until-you-ve-read-this.md b/sreweekly/markdown/528/08-don-t-add-a-read-replica-until-you-ve-read-this.md new file mode 100644 index 00000000..46a419c1 --- /dev/null +++ b/sreweekly/markdown/528/08-don-t-add-a-read-replica-until-you-ve-read-this.md @@ -0,0 +1,197 @@ +# Don’t add a read replica until you’ve read this + +- **期号**: SRE Weekly Issue #528(2026-08-02) +- **作者**: Johanna Larsson — incident.io +- **链接**: https://incident.io/blog/dont-add-a-read-replica-until-youve-read-this + +## 简介 + +The premise: read replicas can help you scale read load, but they introduce complexity. The article goes into the problems they ran into and how they dealt with them. + +## 正文 + +July 21, 2026 — 23 min read + +As the size and complexity of their relational database workload grows, every company eventually goes through the process of off-loading work on a read replica. It comes with lots of benefits, but at a cost of increased complexity. This article is about how we dealt with that, a lot of learnings, and some useful techniques. + +incident.io is the software reliability platform built to investigate, respond and prevent incidents, powered by AI that deeply understands your organization. Thousands of customers rely on it to be the thing that supports them through anything from a minor blip to a full outage. Any disruption to that service has a significant impact on those users, and that’s always top of mind for our engineering team. Everything we build is designed to be performant, reliable, and gracefully degrading. + +When large parts of the internet goes down, as they did for the [AWS outage on 20 October 2025](https://health.aws.amazon.com/health/status?eventID=arn:aws:health:us-east-1::event/MULTIPLE_SERVICES/AWS_MULTIPLE_SERVICES_OPERATIONAL_ISSUE/AWS_MULTIPLE_SERVICES_OPERATIONAL_ISSUE_BA540_514A652BE1A), we see a meaningful increase in alert volume as engineers across the globe are getting woken up. During events like this, we simply can’t fall over from the increased load. This means we always have to run with significant spare capacity. And so we’re always looking for opportunities to reduce the load on our DB, either through performance improvements or code redesign. While working on projects to control the resource utilization on our primary database, we clearly identified that we, like a lot of online services, run an overall read heavy workload. + +We already had a read replica set up and some queries were already utilizing it, but it was something that we were doing on a case by case basis. We knew we could do more. We set ourselves the goal to move *everything* over to the read replica. For every query that could be move moved to the read replica, that’s additional capacity for our primary, and additional protection for those events of massive traffic that we design for. + +But before that, let’s do a quick recap of why you’d want to introduce read replicas into your stack. They do bring complexity, having two databases and two database connection pools in your app logic is more to think about than just having the one. You also need to manage two things, where both often quickly become critical to running your service. Double the graphs and metrics and warnings to worry about. Not to mention, you’re now paying for two databases. + +They bring a lot of benefits though. They’re the natural step to take to give people access to run operational queries on the production system, without risking those queries disrupting the production workload. A query run on the primary can cause overload, delays, contention, locks, and much more. On the read replica the blast radius is much smaller, although there are notable exceptions to this like `hot_standby_feedback`. But maybe the most interesting thing is that it opens the door to horizontal scaling. Relational databases generally don’t scale horizontally, just vertically. Writing to two primary databases is slower than to one, and for most of us spending more to get less performance is not a very interesting prospect. But read replicas are different, since they don’t need to coordinate writes, you can just have more of them and load balance read queries across them. That’s pretty cool. + +Even before approaching this larger re-think we had already established some useful primitives. Our backend language is Go, but the basic techniques translate to some degree to any language. + +The number one thing you need to deal with as you’re migrating work over to a read replica is [read-after-write consistency](https://jepsen.io/consistency/models/read-your-writes). Or in other words, avoiding stale reads. The gist of it is, if you write something to the primary and then immediately after read the thing back but from the replica, there’s no guarantee that you get the same thing back. This gets worse the faster you read after writing. Now if you’re hand-crafting some beautiful artisanal code in your walled garden project, you can probably attempt explicitly picking the primary or read replica for each individual query through a code flow. But for practical reasons we often end up just limiting the read replica to the queries and code paths where we feel that nothing can go wrong. + +That’s not what we’re looking for, we want to move everything over. That means we need automated detection of mutating queries, and to automatically fall over to the primary after a write has been detected to ensure that you are able to read your own writes. The basic mechanism we introduced for this uses the [Go context](https://pkg.go.dev/context) to carry a special flag that controls whether we can use the read replica. It defaults to true. + +``` +type taintedKey struct{} +// A write taints the context: reads on a tainted context must +// go to the primary, since the replica may not have caught up. +func Taint(ctx context.Context) context.Context { + return context.WithValue(ctx, taintedKey{}, true) +} +func CanUseReplica(ctx context.Context) bool { + tainted, _ := ctx.Value(taintedKey{}).(bool) + return !tainted +} +``` +The second concept we introduced was a simple function for detecting whether a given query was “safe” to go to the read replica. Initially we designed this to be conservative, we were ok with some things being misinterpreted and going to the primary, as long as we don’t get accidental and weird bugs where we send the wrong query to the wrong place. Some heuristics got us a long way, like whenever we start a transaction, we send it to the primary and mark the `ctx` so that any subsequent queries also go to the primary, with the basic assumption that anything in a transaction is probably something you want on the primary (this turns out to not be quite true, more on this later). + +``` +// Transactions probably mean writes: run on primary, +// and taint the ctx so everything after follows it there. +func (db *DB) Transaction(ctx context.Context, fn func(context.Context) error) error { + ctx = Taint(ctx) + return db.primary.Transaction(ctx, fn) +} +``` +The second heuristic was to strip comments and trim whitespace, then check if the query starts with `SELECT`. Not very elegant, but it got the job done. There’s a fancier version further down! + +``` +func IsReadOnly(query string) bool { + q := strings.TrimSpace(stripComments(query)) + return strings.HasPrefix(strings.ToUpper(q), "SELECT") +} +``` +On top of this we built a basic layer on top of our database handles that created both primary and replica transaction pools and transparently switched between them, using the flag we had previously set. This means it was now generally safe for us to just opt in any code path to use the read replica, trusting this mechanism to direct the queries to the right place and maintain read-after-write consistency. + +``` +func (db *DB) route(ctx context.Context, query string) *sql.DB { + if !IsReadOnly(query) { + return db.primary // and taint the ctx here + } + if !CanUseReplica(ctx) { + return db.primary // we wrote earlier, stay consistent + } + return db.replica +} +``` +Additionally we put a [circuit breaker](https://martinfowler.com/bliki/CircuitBreaker.html) in here. If we’re not able to open connections to the read replica we give up and send all queries to the primary. Note that although we’re happy to have that fail over behavior for now, with enough total load across your databases, your primary might not be able to handle this. + +incident.io is designed from the ground up on event publication and subscription, the entire system is built around publishing messages and having a fleet of workers processing them. This gives us all kinds of interesting super powers, including the ability to buffer up messages to process them later during periods of overload. + +The workers is exactly where we had been adding some read replica offloading, more specifically in a set of workers that execute low priority workloads that do read heavy work. This is also where we started our work, primarily by moving more and more slices of our workers over to the read replica. Our initial results were really positive, we were almost being too successful and we quickly had to upgrade our read replica because it was taking over so much work from the primary. After having moved some of our heaviest read query workloads, congratulating ourselves for our great work, we noticed something odd. Not a super clear pattern, but the occasional `NotFoundError` coming out of our database adapter in the subscribers on the workers that we had just moved. + +And that’s when it struck us. We were ensuring read-after-write consistency within the context of a go `ctx`, but what about across the boundary of publishing and processing a message? + +Let’s say we have a request come in, to create a post-mortem document. But the post-mortem document itself can take a while to create and involves hitting a separate internal service over the network, we don’t do all of that work in line with the client waiting. So we write the row to the database, respond immediately to the client, and then publish a message to a worker to deal with actually creating the document. Incredibly, what we were seeing was the time between publishing a message to the message queue and the worker pulling that message to process it be shorter than the replication lag to the replica. This wasn’t because the replica was falling behind, it was doing just fine, it was just that *occasionally* the event was processed incredibly quickly. Overall it was a rare occurrence, something between 0.1 and 0.5% of messages, and they were automatically retried, but obviously not something we wanted to keep getting alerted on. Interestingly we had technically had this issue for a few code paths for a while, but it wasn’t until we wholesale moved large parts of our codebase over to the read replica that this problem showed itself. + +We discussed some options and quickly homed in on a well-documented tool in the PostgreSQL world. + +The log sequence number, or LSN, is a special pointer in PostgreSQL that represents a specific position in the [Write-Ahead Log](https://www.postgresql.org/docs/current/wal-intro.html). This means that you can use it to check whether the read replica has caught up with a specific write in the primary database. You get the LSN with a built in PostgreSQL function: + +``` +SELECT pg_current_wal_lsn()::text +``` +The basic technique we want to apply: + +1. Grab the LSN from the primary after having done your write +2. Pass it to the code running on the worker +3. Compare the read replica LSN with the LSN from the primary, if the read replica is at or after that LSN, you know that you are past the point where your write happened + +In order to enforce read-after-write consistency across our workers, we started stamping each published message with the LSN from the primary at the time of publishing. Grabbing the LSN is very cheap, it’s just an in-memory operation on the PostgreSQL side. + +Having rolled that out, every message now has this stamp on it. So then we added a check on the subscriber side that compared the LSN on the message with the LSN of the read replica. The LSN comparison can be done inside of a SQL query using the [native data type](https://www.postgresql.org/docs/current/datatype-pg-lsn.html): + +``` +SELECT pg_last_wal_replay_lsn() >= $1::pg_lsn +``` +If the replica was caught up, we are fine to process it. If it isn’t, we just kick the message back to the queue. + +That last part actually turns out to be a great back pressure mechanism in overload situations. If the read replica is failing to keep up with the primary, for whatever reason, instead of directing all queries to the primary and risking overloading it, we just keep nacking the messages until the read replica is feeling better again, with exponential backoff to avoid overload. It turns a potential outage situation into a graceful degradation instead, all the work is processed as expected, just with a delay. + +If you’re doing this, you want to couple it with some dashboards and alerts. You don’t want to get caught out on the replica lagging behind. We often see people measure read replica lag in time, but that’s not a useful measurement. A read replica can catch up 5 minutes of delay in seconds because the rate of change has slowed down, or it can struggle to close the distance on 30s of delay because the rate of change is going up. Basically, the replay rate can change. A more useful measurement is the number of bytes behind. It still doesn’t quite convey whether you have a problem or not, but it’s more consistent than measuring it as time. + +No more random `NotFoundError`, great success! + +After having tackled the events and subscribers, we looked around for other places where we would expect to bump into the same consistency problem, where we need to be careful to read our own writes, and we identified the API as another surface area to deal with. This included our public API where we support creating resources, returning an ID, and then reading the resource back. Done quickly you’d risk a 404. Additionally, our dashboard relies on our internal private API and we frequently use the same pattern there: create an empty resource and return the ID, queue the work to actually create the resource, frontend polls the API using the ID, resource is eventually created and displayed in the UI. + +There are a few different [classic blog posts](< https://brandur.org/postgres-reads>) that talk about using LSN stamping on API requests, so this is not exactly new territory. We want to apply the same principles as we did for our events, but we can be a bit more elegant about it for the API. Thinking about the problem we’re solving, it’s basically focused on the case of a `POST` followed by a `GET`. Or more generally, a mutating request followed by a read. Since we’re pretty consistent about HTTP verbs in our API, we can actually limit LSN stamping to mutating requests: `POST`, `PUT`, `PATCH`, `DELETE`. Additionally we can limit our scope to a given actor, whether it’s a user, an API key, or something like our mobile app, as the domain of consistency. If I create a document I want to see it immediately, but it’s ok for my coworker to have a few milliseconds delay to see the same document. + +We added a middleware that runs after the request handler in our [Goa web layer](https://goa.design/), checks the method of the request, and optionally stamps the actor **with the LSN. We store this directly in PostgreSQL, using the native data type. + +``` +UPDATE users SET read_after_write_lsn = GREATEST(read_after_write_lsn, pg_current_wal_lsn()) WHERE id = ? +``` +This has the benefit of us already loading that row anyway during auth at the start of every request, so we could include this column there. This is also why we chose not to move this work to a key value store like Redis. Redis would add a network request, looking up the LSN, to every single incoming request, and we still need to get the latest LSN from replica too. + +We shipped the middleware and with our actor tables quickly filling up with LSN stamps, each representing the last action each one has taken, we tackled the other half of this. We added another middleware that runs after auth and takes the LSN stamp from the actor and compares it with the read replica. Unlike subscribers, we can’t just nack the message and send it back to the queue to be processed later. People tend to not want to have to retry all their requests, and we didn’t want to push this on every client to our system, of either having to retry on a certain status code, or carry LSN stamps on headers. So instead of nacking, we fall back to the primary. We may need to rethink that at some point, when we can’t afford to go to primary anymore, but after rolling this out the actual impact on the primary is minimal. Only about 0.01% of requests to our API fail the LSN check. + +With this we could move the last part of our codebase over to the read replica, reversing the trend of increased CPU utilization on primary week over week and getting it back down to a healthy level with all that extra capacity that we aim for. + +After having rolled this out we went after two optimizations that we had identified while we were working on this project. Although fairly simple, we didn’t implement them before rolling out because we 1. wanted to avoid premature optimizations, and 2. doing it after means we get a pretty graph and clear validation of the efficacy of the optimization. Everyone loves a pretty graph. + +The first one was to limit the number of nacks on the event subscribers. As mentioned before, we saw between 0.1 and 0.5% of processed events getting sent back to the queue due to replica lag. Each nack causes a delay of at least 10 seconds in processing the message, since that’s our default retry delay. That’s very much acceptable, but we had a theory: that when we had replica lag the lag is almost always very small. + +So whenever we check the replica and it’s behind, we added a 100 ms sleep before trying again. The impact was striking, almost eliminating nacks. Over time we saw the nack rate drop to ~0.03%. For what was basically a couple of lines of code. In real terms we’re not talking about a lot of messages, but it was a worthwhile improvement. PostgreSQL 19 actually comes with this [functionality built in](https://rednafi.com/system/wait-for-lsn/), with the new `WAIT FOR LSN`. + +Secondly, when investigating why certain queries, that were actually perfectly safe to run on the read replica, were getting directed to the primary we realized our naive approach to detecting mutating queries was maybe just a little bit too naive. We were missing out on tons of juicy queries that were absolutely eligible for the read replica. Taking inspiration from the [https://github.com/pgplex/pgparser](https://github.com/pgplex/pgparser) library, we created a small tokenizer to replace our heuristics. Armed with the tokenizer, we could now confidently identify any queries that were safe to run on the read replica. When we rolled it out we saw a massive drop in CPU utilization on the primary, and a corresponding increase on the replica side. + +So now we’re in our new better world where everything is opted into the read replica by default, we have read-after-write consistency, and thanks to LSN stamping, that applies across events and subscribers, and the API as well. + +But we had two remaining concerns: firstly that our code was still explicitly opting into the read replica everywhere, and secondly that internal concerns were leaking: anyone who needed specific database pool behaviors had to understand exactly how all of this worked. We didn’t want to have to maintain this for the team permanently, or introduce friction for everyone. So we went after the last major part of the project: inverting the default and creating a “DSL” for choosing database connection pool routing strategies. Everything goes to the read replica, whether you’re aware or not, and the mechanisms we’ve implemented ensures your code just works. + +The easy part was tweaking our database handles with internal pool routing, we changed it to be enabled by default. The bigger part was giving the engineering team the tools to tweak the behavior where they needed to. We identified four strategies: + +- `ReadAfterWrite` - this is our default behavior. Track mutations and route to primary after. +- `StaleRead` - where you have reads after writes, but you’re fine with them going to the read replica. +- `Primary` - pin the context to the primary and send all queries there. We use this where we can’t tolerate any kind of lag, primarily around our on-call product. +- `Replica` - pin the context to the replica and send all queries there. This is kind of a nuclear option, it will force queries to the replica even when the replica can’t execute them. Useful where you’d rather fail loudly and fix your code. + +We accept these four strategies on every different level. The database handle takes strategies, and so does the subscriber definitions and the web layer API endpoint definitions. You can also override strategies per context. + +This gives us the inverted default, everything goes to read replica, while still providing the engineering team with the tools they need to keep shipping. + +Once every part of our codebase was “opted in” to the new read replica behavior, we inverted the default. Everything now goes to the read replica, and the people writing code don’t have to worry about it. It just works. Getting there is not hard, but it did come with a lot of learnings. + +So what about the numbers? After the project wrapped up more than 60% of all read queries go to the replica. Of the ones that don’t, the majority were explicitly pinned to the primary. Only a small number of reads are actually routed to primary to preserve read-after-write consistency. + +**We cut CPU utilization on the primary in half.** Looking back over the last few months CPU utilization had been growing significantly week over week as our customer base has been expanding. This project reversed the trend and we’ve had several weeks of CPU utilization going down week over week. A healthy primary with lots of spare capacity means that we can keep handling disaster situations where half the internet goes down and everyone gets paged. + +**Replica lag causes back pressure instead of overload.** In a catastrophe situation where our replica is failing to keep up, or going down, all events are safely buffered in our queue service and only picked up when the replica is ready to get back to work. This means that problems with the replica do not spread to other parts of our stack. + +**We’ve set the stage for horizontal scaling.** The door is now wide open for us to add additional read replicas, giving us a clear path to keep scaling our product for the future. + +**Engineering can keep shipping.** Our transparent, opt-in by default, database pool routing strategies and read-after-write consistency mechanisms ensure that the read replica does not get in your way. You could work here for months without even realizing it exists. + +We hope this writeup can help demystify read replicas and how to get the most out of them, applying fairly straightforward techniques to ensure read-after-write consistency. We really enjoyed working on this project and we hope you’ve enjoyed reading about it! If you’re interested in this kind of thing, come join us, we’re hiring! + +Johanna Larsson + +Product Engineer + +Our rate limiter depends on Valkey. If Valkey goes down we fail open and stop limiting which isn't good enough for our platform. As an intern, I built per-pod in-memory top-k buffers so we keep rate limiting even with the backing store gone. + +Anthony Oparaocha + +September 2, 2026 + +Our entire event-driven platform ran through a single message broker, which made it a single point of failure. So we added a second one. This is the story of building an event load balancer, the queuing theory behind it, and the final chaos test where we turned off Pub/Sub in production and nobody noticed. + +Patrick Hamann + ++ + +Mike Fisher + +August 11, 2026 + +Today we're launching our new post-mortems experience, and I want to walk you through what we've done and why. + +Pete Hamilton + +March 17, 2026 + +Ready for modern incident management? Book a call with one of our experts today. + +- All-in-one incident management +- Our unmatched speed of deployment +- Why we’re loved by users and easily adopted +- How we work for the whole organization diff --git a/sreweekly/markdown/529/01-without-a-program-to-support-them-incident-management-processes-wither.md b/sreweekly/markdown/529/01-without-a-program-to-support-them-incident-management-processes-wither.md index 758b1db7..8ab41488 100644 --- a/sreweekly/markdown/529/01-without-a-program-to-support-them-incident-management-processes-wither.md +++ b/sreweekly/markdown/529/01-without-a-program-to-support-them-incident-management-processes-wither.md @@ -7,3 +7,61 @@ ## 简介 It’s not enough to define an incident process. You have to spin up and maintain an entire incident management program. + +## 正文 + +Every fire department has a training program. Big-city departments have entire training divisions; even small volunteer departments that can’t spare anyone full time still name a training officer. Not because training is the department’s mission, but because maintaining the *capability* to do the mission requires sustained, dedicated attention. + +New recruits need to be brought up to speed. Everyone needs to learn about evolving techniques and new equipment. Procedures need to be updated as building codes and materials change. Hard-won lessons from past incidents would survive only as stories told around the kitchen table; the fire service has a strong storytelling tradition, and its legends and cautionary tales carry real value, but oral history is hard to study, standardize, and train on. + +Maintaining operational capability is itself a job, distinct from the operational work it supports, and fire departments size the role to the department rather than leave it unassigned. + +Many software companies haven’t learned this yet. They invest real effort in building an incident management process. They define severity levels, write runbooks, designate incident commanders (ICs), set up communication channels. The project might take weeks or months of focused work, often driven by someone who cares deeply about doing it right (and often done in their “spare time”). When it’s done, it works, at least for a while. Incidents get declared. ICs run the response. Post-incident reviews happen. Everyone takes it for granted. + +Then the person driving it gets promoted, or moves to another team, or leaves the company. The process, which was never really institutionalized because it didn’t need to be while that person was carrying it, begins to decay. Not catastrophically, but more like a garden nobody is tending any more: it doesn’t collapse overnight, it just slowly fills with weeds until one day you look up and realize the original design is barely recognizable. + +The training materials haven’t been updated since the initial rollout. New engineers join but never go through incident training because nobody is scheduling it anymore. The severity level definitions still describe one product, but the company now has three. The IC rotation is running on the same six people it started with, even though the engineering team has doubled in size. The post-incident review template still references a tool the company stopped using a year ago. + +None of these are crises on their own. Each one is easy to defer. But they compound, and the cumulative effect is that the process on paper bears less and less resemblance to what actually happens during incidents. In my experience, six months is roughly how long institutional momentum carries before the absence of active stewardship becomes visible in the quality of your incident responses. And growth accelerates the decay: the company simply grows away from the process, and nobody’s job is to notice. + +This is what happens when you have a process but not a program. + +## A process is not a program + +A process is a set of documented procedures: how incidents get declared, who fills which roles, what communication channels to use, how to run a post-incident review. A process can be written down, trained once, and followed. + +A program is the organizational structure that develops, maintains, evolves, and champions the process over time. It’s the thing that keeps the process alive. + +Many companies build the process and assume they’ve built the program. They haven’t. They’ve written a document, and documents don’t train new hires, don’t recruit for on-call rotations, and don’t update themselves when the company reorganizes around them. People do those things, and it only happens reliably when it’s actually somebody’s job. + +## “Everybody owns it” means nobody owns it + +When I ask companies who owns their incident management program, the most common answer is some version of “we all do” or “the engineering organization as a whole.” This *sounds* collaborative. In practice, it means nobody has the explicit responsibility, the dedicated time, or the institutional authority to keep the process alive. + +This organizational challenge isn’t unique to incident management. Companies that are serious about security don’t say “everybody owns security” and leave it at that. They assign ownership because shared responsibility without explicit ownership means the work doesn’t get done. + +Incident management is the same kind of organizational capability. It needs someone whose actual job, not just their passionate side interest, is keeping it healthy. + +## What a program actually does + +When I talk about an incident management program, I mean ownership of the full lifecycle of the capability, not just the procedures themselves. That includes keeping everything current as the company grows and changes: process documentation, severity definitions, escalation paths, tooling, runbooks. + +It includes running a training pipeline so new hires are prepared *before* their first real incident, not thrown into the deep end *during* it. It includes maintaining the incident commander corps: recruiting new incident commanders, nurturing their development, supporting healthy on-call rotations across teams, and recognizing the people who do this demanding work. My former Slack colleague Scott Nelson Windels likens this to the farm teams and academies that elite sports clubs run: the point isn’t just fielding today’s roster, it’s making sure capable players are always coming up to fill it next quarter, too. + +And it includes owning the post-incident review process and looking across incidents for patterns that no individual team would spot on their own. It includes tracking whether the process is actually being followed, and investigating when it isn’t, not to punish people, but to understand whether the process needs to change. + +No single component is enough on its own, and no component stays healthy without sustained attention. + +## The good news + +Building a program doesn’t require hiring a large team or creating a new department. At many companies, especially smaller ones, it starts with one person who has explicit ownership and dedicated time. What matters is that the responsibility is named, visible, and institutionally supported, not just assumed. + +Here’s a quick test. Ask who owns your incident management program. Not who wrote the process, and not who ran the last big incident, but who is accountable, today, for whether the training is current, the rotations are staffed, and the severity levels still match the product. If the answer is a name, the follow-up question is what happens when that person leaves. If the answer is “everybody,” or someone who left the company last year, the process is quietly withering. And if you have a program but it would collapse without you, you haven’t finished building it yet. + +The fire department didn’t name a training officer because it had extra budget. It named a training officer because it understood that maintaining a capability requires ongoing investment. The alternative, assuming trained firefighters stay trained and procedures stay current without anyone specifically owning those things, is how capabilities quietly erode until they fail when you need them most. + +*I’m writing a book on [Incident Management for DevOps and SRE](https://im4ds.com). If you’d like to know when it’s available, and get occasional updates along the way, you can sign up at [im4ds.com](https://im4ds.com).* + +*If your company needs help building its incident management program, that’s the focus of my consulting practice at [Great Circle](https://greatcircle.com/im).* + +## Recent Comments diff --git a/sreweekly/markdown/529/02-what-comes-after-observability.md b/sreweekly/markdown/529/02-what-comes-after-observability.md index 4b93e26b..0d337731 100644 --- a/sreweekly/markdown/529/02-what-comes-after-observability.md +++ b/sreweekly/markdown/529/02-what-comes-after-observability.md @@ -7,3 +7,89 @@ ## 简介 Honeycomb pulls back the curtain a bit to delve into how LLM agents change the way their product is used, and how their query patterns differ from humans’. It’s especially interesting that increasing agent usage has not correlated with decreasing human usage. + +## 正文 + +# What Comes After Observability? + +A year ago, I predicted ways in which AI was about to fundamentally change observability as we knew it. Here's what we've seen happen since—both at Honeycomb and with our customers—and what we're building for the future. + +![Austin Parker](https://www.honeycomb.io/_next/image?url=https%3A%2F%2Fcdn.sanity.io%2Fimages%2F927dxq0h%2Fproduction%2F2a1050e83bac34eea79bd9f2d5bf7c3fc106550f-600x600.jpg%3Fw%3D80%26h%3D80&w=256&q=75) + +By: [Austin Parker](https://www.honeycomb.io/author/austin) + +![The Second Edition of Observability Engineering Is Here](https://www.honeycomb.io/_next/image?url=https%3A%2F%2Fcdn.sanity.io%2Fimages%2F927dxq0h%2Fproduction%2Fee7172fc6d57df70b4b2dd303f6a58fc87e24e93-3840x2160.png&w=640&q=75) + +#### The Second Edition of Observability Engineering Is Here + +The second edition of Observability Engineering is available for download on our website. + +[Download](https://www.honeycomb.io/blog/the-second-edition-is-here) + +![What Comes After Observability?](https://www.honeycomb.io/_next/image?url=https%3A%2F%2Fcdn.sanity.io%2Fimages%2F927dxq0h%2Fproduction%2Fe074790a74ac795de3c41ddaceb84ff80f8393e3-3840x2160.png&w=3840&q=75) + +A year ago, I wrote [“It’s the End of Observability (and I Feel Fine).”](https://www.honeycomb.io/blog/its-the-end-of-observability-as-we-know-it-and-i-feel-fine) The upshot of that post was that AI was about to fundamentally change the way we approach systems design and operation in the future. In the grand tradition, I’d like to revisit my claims from then and see how my predictions panned out. + +## Claim: Agents can zero-shot investigations for less than a dollar + +*I asked the agent the same question we’d ask you in a demo, and the agent figured it out with no additional prompts, training, or guidance…and it did it for sixty cents.* + +I pointed out in my original blog that the way I evaluated our MCP server was to feed it a prompt based on the same demo that we give to people—e.g., here’s this weird latency spike, investigate it, tell me why it happened. At the time, that demo took eight tool calls and about $0.60 of inference. A year later, we process about 2 million agent-initiated query runs through Honeycomb every month! + +What’s really interesting, though, is that we’ve seen agents start to accomplish significantly longer-horizon tasks. Frontier models—like Sonnet or Fable 5—are increasingly agentic and capable. In the same agent session, we’ve seen tool calls double on average between February and June of this year. Our original demo of eight queries looks downright quaint. + +# Watch Austin, on demand + +Watch Austin Parker and the other authors + +of Observability Engineering + +discuss what's changed in the AI era. + +## Claim: Humans stay in the loop + +*I’m not gonna sit here and say this destroys the idea of humans being involved in the process, though. I don’t think that’s true. The rise of the cloud didn’t destroy the idea of IT. The existence of Rails doesn’t mean we don’t need server programmers. Productivity increases expand the map. There’ll be more software, of all shapes and sizes. We’re going to need more of everything.* + +What’s really interesting is that *human-initiated queries haven’t shrunk*. Agents aren’t taking investigations away from humans, they’re *accentuating* them. People still use the UI, people still check their boards, but they’re now able to do more than they could before by leveraging AI. We see this a lot in terms of edits—while people *are* using agents to build boards and update SLOs, the overwhelming majority of tool calls are to just run queries. + +In my blog, I claimed that productivity increases weren’t going to diminish the impact or effect of people, and that’s what the data shows. The map is not the territory, but the map has grown significantly thanks to agents. + +This doesn’t mean that all of these agents are being driven by a human, though. We’re seeing an increasing amount of headless agents using our MCP; it’s actually the fastest-growing segment. + +## Claim: It’s only getting cheaper to do this + +*Inference costs are only going down…If your product’s value proposition is nice graphs and easy instrumentation, you are cooked. An LLM commoditizes the analysis piece, OpenTelemetry commoditizes the instrumentation piece.* + +In my original post, I focused on the idea of inference costs going down. I think that’s, broadly, still true (even if a lot of people are getting sticker shock as they transition from all-you-can-eat to API pricing for development workloads). That said, I’m writing this post on a laptop that effortlessly runs Google’s Gemma4 model, and it’s perfectly capable of using our MCP. I think, long-run, the “cost of AI” is going to keep going down at a fixed level of capability. + +What I didn’t call, though, is that the robots are actually surprisingly efficient when it comes to *using* Honeycomb! On average, agent queries cost about half as much to serve as human ones. There’s a pretty easy explanation for this: the agents are able to much more accurately figure out what to search for, especially if they’re running with access to your code. They know exactly what to look at, and they don’t need to spend as much time searching across an entire environment to orient themselves. + +Interestingly enough, we also notice that agents—on average—tend to run queries over a 50% shorter time range than humans. Where this gets really interesting is how *different* the agent runs tend to be. Custom agents tend to look at smaller time windows, while human-driven ones tend to look across longer periods of time. My hypothesis is that those custom agents are probably more task-oriented (“Here’s an error, look around this time for traces”) vs. more exploratory work being done in the development loop. + +The headline number, though? On average, an agent-initiated query fans out to 3.3x fewer lambdas, scans 1.8x fewer bytes, and burns 2.2x less compute. What does this mean for us? Well, from March to June we’ve grown 4x in query volume while only increasing cost by 23%, leading to a 72% reduction in unit costs for agent-initiated queries. I’ll take that! + +![Average per-query run of human-initiated queries compared to agent-initiated queries.](https://www.honeycomb.io/_next/image?url=https%3A%2F%2Fcdn.sanity.io%2Fimages%2F927dxq0h%2Fproduction%2Ff4a4555365d1b22f17ee5bf3c66bc36905701798-3840x2160.png&w=3840&q=75) + +An interesting coda to this, though, is that agents are *significantly worse* for caching. While we don’t see significant amounts of caching for interactive queries in general, we see almost *none* for agent-driven queries. This is mostly because agents don’t need to ask the same question twice. I’ve also personally observed that agents prefer to re-run queries rather than reuse existing ones, which is interesting and probably deserves more attention. + +That said, we’ve seen pretty [staggering levels of adoption](https://www.honeycomb.io/blog/30-70-prs-day-how-we-managed-not-wreck-systems), especially in the enterprise. One of our larger customers read 62 PiB of data in one month's time *exclusively* through agents, mostly through off-the-shelf tools like Claude Code. + +There’s a downside to this as well, though: a lot of those queries wind up being kinda slow, because the agents aren’t chewing through nicely structured data, they’re grepping across unstructured body fields. Structured, high-cardinality, high-dimensionality data is no longer a nice-to-have for humans, **it is a requirement for making agent investigations affordable**. All of that data that the agent has to chew through in order to find the needle in the haystack? That’s tokens coming out of your pocket. Do you want to pay Fable rates to look through gigabytes of crappy logs? + +## Claim: Fast feedback loops are all that matter + +*I’m gonna put a marker out there: the only thing that really matters is fast, tight feedback loops at every stage of development and operations. AI thrives on speed—it’ll outrun you every time.* + +Just go ask Claude to summarize the current state of the “Loops” discourse and you’ll see that the rest of the AI thought leadership industry is talking about something you read here a year ago. + +Anyway, there’s no prize for being first, so I instead want to share a story from one of our customers. Guy writes in to tell us how much he loves MCP and how he’s using it. He’s hooked up everything to agents that talk to Honeycomb. When a customer reports an issue or someone files a bug, the agent goes out to verify it using production telemetry. When an SLO burns, agent goes out, looks into it. Agent takes that investigation, hands it off to another agent, which goes out and writes a fix, makes a PR. Yet another agent reviews the PR and pings a human for final review and merge. PR gets deployed, goes into prod, wakes up that original agent to look at production telemetry to see if the problem was fixed. If it wasn’t, enter the loop again. + +That’s the kind of tool that we’re building at Honeycomb today, yes. But what we’re building for tomorrow is so much more than that. I don’t want us to build *just* a really fast column store (although a really fast column store is cool). I want to make it easy for everyone to have these fast feedback loops. I want you to be able to understand what your agents are doing with your code, and I want to make it easy for you—and your agents—to learn about production and turn those learnings into knowledge. + +If you’ve had the chance to check out [Canvas](https://www.honeycomb.io/platform/canvas), and our Canvas Agent, you’ve seen the first cut of this (and if you haven’t, you should check it out; it’s free!). This is just the first chapter of what we’re building, though. If you're interested in learning more about how we see AI changing observability and what we're working on here at Honeycomb, join me and the other authors of the Observability Engineering book [in our on-demand AMA](https://www.honeycomb.io/resources/webinars/ama-new-engineering-realities-observability-engineering-authors). + +P.S. [We’re hiring](https://www.honeycomb.io/careers) if you wanna come work on agents with us. + +# Want to learn more? + +Talk to our team about how we're helping organizations build the operational foundation for AI development success. diff --git a/sreweekly/markdown/529/03-the-rise-of-agentic-sre-humans-agents-and-reliability.md b/sreweekly/markdown/529/03-the-rise-of-agentic-sre-humans-agents-and-reliability.md index 3a1a2452..7745cb01 100644 --- a/sreweekly/markdown/529/03-the-rise-of-agentic-sre-humans-agents-and-reliability.md +++ b/sreweekly/markdown/529/03-the-rise-of-agentic-sre-humans-agents-and-reliability.md @@ -9,3 +9,109 @@ I like the approach here, especially measuring both the positive and negative outcomes. > Good SRE practice is about evidence, not enthusiasm. + +## 正文 + +- + ![](https://dz2cdn1.dzone.com/themes/dz20/images/dz-postarticle.svg) [Post an Article](https://dzone.com/content/article/post.html) +- + [Manage My Drafts](https://dzone.com) + +# The Rise of Agentic SRE: Humans, Agents, and Reliability + +Agentic SRE speeds up incident response, but it also requires clear guardrails, strong observability, and human oversight. + +Join the DZone community and get the full member experience. + +[Join For Free](https://dzone.com/static/registration.html) + +Site reliability engineering has always been about reducing toil, improving resilience and helping teams respond to incidents with speed and confidence. [Agentic SRE](https://dzone.com/articles/agentic-ai-sre-copilot-incident-response) takes this idea further, allowing AI systems to observe, reason, and act within operational workflows inside of bounded constraints. The outcome is not a replacement for SREs, but a new operating model in which humans supervise intelligent agents that can help triage, diagnose, and remediate faster than manual processes alone. + +## What Agentic SRE Means + +Agentic SRE is the use of AI agents to carry out reliability tasks with some autonomy. The agents are able to capture telemetry, correlate signals across systems, propose likely causes, take safe actions, and hand over to humans when the problem exceeds their authority. In practice, this means an AI assistant that can summarise an incident, pull up relevant dashboards, check recent deploys, compare symptoms against runbooks and even trigger low-risk remediation steps. + +What changes are not the nature of the assistance but the limits of its application. Traditional automation is usually rule-based: if X happens, do Y. Agentic systems are different in that they can adapt to context, select between several paths, and orchestrate steps across tools. This makes them especially useful in complex environments where the same symptom may come from many different root causes. + +## Why SRE Needs Agents + +The systems today are too big and too interconnected to be run totally by hand. Teams are contending with noisy alerts, fragmented observability data, constant deployments, and ever more dynamic infrastructure. During incidents, engineers often burn precious minutes just to gather context before they can start a real diagnosis. Agentic SRE is attractive because it shortens that time. + +In a handful of high-friction places, agents can cut toil. They can filter alert storms, enrich alerts with deployment history, draw out meaningful patterns from logs, and surface relevant runbooks. They can also automate repetitive incident response tasks such as opening tickets, notifying owners, checking service health, or validating if a rollback is safe. That doesn't eliminate the need for engineers, but it does take away some of the low-value work that distracts them from judgment-intensive choices. + +## The Human Role Remains Central + +One common fear is that SREs will be replaced with autonomous systems. Indeed, the human role becomes more, not less, important. Agents are good at pattern recognition, summarisation, and bounded execution. Humans are still better at trade-offs, risk assessment, organisational context, and deciding when not to act. Reliability is not merely a technical problem. It is a business and coordination problem. + +Humans should set policies, guardrails, and escalation thresholds for agent behaviour. They need to decide which actions can be safely automated, which require approval, and which should never be delegated. So the SRE is evolving from operator to system designer, to policy author, to reliability supervisor. That shift is profound because the skill set you need for the job changes. + +## Where Agents Fit Today + +The best place to start is with low-risk, high-frequency jobs. These are the areas where automation can provide immediate value without unacceptable risk. Think incident summarisation, alert enrichment, log correlation, runbook retrieval, change impact analysis, and post-incident report drafting. + +[Incident copilots](https://dzone.com/articles/rd-infrastructure-and-science?fromrel=true) are another strong use case. An agent can also act as a second brain during an outage: it can aggregate timelines, verify recent code changes, search knowledge bases, and suggest next steps. It can help responders avoid duplication of effort and make the first 10 minutes of an incident much more productive. An effective agent can also lessen the cognitive load on on-call engineers by turning the scattered telemetry into a coherent story. + +A third useful area is remediation assistance. Agents can recommend actions such as scaling a service, restarting a failing job, disabling a faulty feature flag, or rolling back a deployment. In mature setups, these actions can be executed automatically for pre-approved scenarios, while more risky actions still require human confirmation. That combination of automation and oversight is where agentic SRE becomes genuinely powerful. + +## A Practical Architecture + +An effective agentic SRE system typically has five layers. First, it needs a telemetry layer that includes metrics, logs, traces, events, and deployment data. Without strong observability, the agent is blind and will make incorrect guesses. Second, it requires a reasoning layer, often powered by an LLM, to interpret context and decide what to do next. + +Third, there should be a tool layer that gives the agent access to safe operational functions, such as querying dashboards, reading configs, opening tickets, or triggering runbooks. Fourth, it needs policies and guardrails that define permissions, approval workflows, rate limits, and failure boundaries. Finally, it should have an audit layer so every decision, action, and recommendation can be traced later. + +That architecture matters because the danger is not the model itself; the danger is uncontrolled action. A reliable agent is not one that knows everything. It is one that acts only within well-defined limits and remains observable, reversible, and accountable. + +## Guardrails That Matter + +Trust is the currency of autonomous operations. If teams do not trust the system, they will ignore it. If they trust it too much, they may hand over dangerous actions without oversight. The right answer is neither blind trust nor permanent skepticism. It is a layered trust model built through guardrails. + +Start with permission scoping. An agent should not have broad access by default. Its permissions should be narrow, explicit, and tied to specific tasks. Next, use action tiers. Low-risk actions can be automatic, medium-risk actions can require confirmation, and high-risk actions should remain human-only. You also need strong rollback paths so any automated action can be quickly reversed. + +Another essential safeguard is the observability of the agent itself. Just as production systems need monitoring, agents need monitoring too. Teams should track what the agent saw, what it inferred, what action it proposed, and whether the result improved the situation. That makes the system auditable and helps teams refine its behavior over time. + +## The Operating Model Changes + +Agentic SRE changes incident response from a purely human workflow into a human-agent collaboration loop. In the old model, an engineer gets paged, reads alerts, searches dashboards, checks logs, consults teammates, and then acts. In the new model, the agent can do much of the initial gathering and triage before the human even joins. That shortens the path from detection to understanding. + +This also changes how teams design runbooks. Instead of static documents that people read under pressure, runbooks become machine-readable operational playbooks. Some of the best runbooks will be written with automation in mind, including clear preconditions, decision points, and action boundaries. That makes them useful both for humans and for agents. + +Post-incident work also improves. Agents can draft a timeline, collect evidence, identify suspicious changes, and summarize repeated patterns across incidents. That leaves engineers with more time to focus on systemic fixes rather than manual documentation. Over time, the organization develops a stronger feedback loop between incidents, learning, and platform improvements. + +## Risks and Failure Modes + +Agentic SRE is not free of risk. One failure mode is confident hallucination, where an agent sounds plausible but is wrong. In operations, a wrong answer is not just inaccurate; it can cause downtime. Another risk is over-automation, where teams let agents act in situations that are not actually safe to delegate. + +There is also the risk of hidden complexity. If an agent stitches together many systems, it can become difficult to understand why it chose a specific action. That opacity can undermine trust and create governance problems. Security is another major concern because an agent with tool access can become an attractive target if permissions are poorly controlled. + +These risks do not mean agents should be avoided. They mean they must be introduced carefully. The best strategy is to start with narrow, well-understood workflows, measure outcomes, and expand only when confidence is earned. Reliability teams already understand progressive delivery, canary releases, and blast-radius reduction; the same principles should apply to agentic operations. + +## How to Start + +The easiest entry point is to pick one painful workflow and automate only the first mile. A good candidate is alert triage. An agent can ingest alerts, group duplicates, summarize likely causes, and point responders toward relevant dashboards and runbooks. That alone can save significant time without requiring the agent to make risky changes. + +Another strong starting point is incident summarization. This is low risk, highly useful, and easy for teams to evaluate. A third option is change impact analysis, where an agent compares recent deploys, feature flag changes, and error spikes to highlight likely correlations. These use cases are valuable because they build trust through usefulness rather than hype. + +Measure success with clear operational metrics. Look at time to acknowledge, time to diagnose, time to mitigate, alert volume reduction, and after-hours toil reduction. Also measure negative outcomes, such as false suggestions, unsafe recommendations, or overreliance on the agent. Good SRE practice is about evidence, not enthusiasm. + +## A New Reliability Mindset + +The biggest change Agentic SRE brings is a shift in mindset. It encourages teams to stop viewing automation as just a collection of scripts and to see it instead as a supervised operational partner. This partner can observe faster than a person, summarise quickly, and carry out repetitive tasks more reliably. However, it still requires humans to define the purpose, set limits, and determine acceptable risk. + +This is why agentic SRE is not merely “AI in operations". It represents a larger redesign of how reliability work is accomplished. The focus shifts from manual responses to intelligent coordination. It changes from isolated dashboards to context-aware agents. It evolves from static runbooks to flexible playbooks. It transforms reactive tasks into guided independence. + +Organizations that excel with this model will not be the ones that automate everything. They will be the ones that automate thoughtfully, govern effectively, and keep humans involved where decision-making matters most. In this way, Agentic SRE is more about enhancing the reliability system around the engineer than about replacing the engineer themselves. + +## Closing Thoughts + +Agentic SRE marks a real change in how we can manage modern systems. It provides a way to respond faster, reduce repetitive work, and handle incidents more consistently, but only with strong observability, clear permissions, and human oversight. The future of reliability is not completely automatic or fully manual; it is collaborative, constrained, and constantly improving. + +For SRE teams, there's a chance to become designers of this new model. This involves creating agent workflows, writing safer runbooks, setting policy limits, and measuring impact with the same attention given to any production system. Teams that excel in this will not only respond more quickly. They will create systems that are more resilient, more adaptable, and much simpler to operate at scale. + + Site reliability engineering + systems + Trust (business) + + + Opinions expressed by DZone contributors are their own. + +Comments diff --git a/sreweekly/markdown/529/04-incident-report-july-2-2026-us-east-services-outage.md b/sreweekly/markdown/529/04-incident-report-july-2-2026-us-east-services-outage.md index f6074f33..6043d8bb 100644 --- a/sreweekly/markdown/529/04-incident-report-july-2-2026-us-east-services-outage.md +++ b/sreweekly/markdown/529/04-incident-report-july-2-2026-us-east-services-outage.md @@ -7,3 +7,99 @@ ## 简介 The kernel’s route cache: a hidden reliability killer. This is a really intriguing case of self-sustaining impact. + +## 正文 + +![Avatar of Ray Chen](https://cms.railway.com/media/person-ray-chen-dd0fa1297e94.jpg) + +# Incident Report: July 2, 2026 — US East Services Outage + +Railway experienced a Major Outage concentrated in one of our US East availability zones on July 2, 2026. + +A network degradation in one of the ISPs connecting our datacenters to the rest of the internet caused elevated latency and packet loss for traffic between our US regions. While rerouting traffic away from degraded ISP, a change at one of our US East availability zones briefly left the site without a stable route to the internet. + +The effects from above exposed hidden bugs that silently pushed storage traffic onto a slow backup network and impacted some private networking tunnels, degrading disk performance and private networking in US East for roughly two hours. + + +## Impact + +On July 2, 2026 between roughly 07:44 UTC and 12:01 UTC, users may have experienced increased response times and intermittent connectivity issues on traffic between US regions, including private networking. Some workloads in one of our US East availability zones additionally saw degraded disk performance and disrupted private networking for roughly two hours. + + +## Incident Timeline + +*All times are UTC on July 2, 2026.* + +- **07:44** — We started observing packet loss in our US East region affecting user traffic. A public incident was declared on our status page. We traced the packet loss to one of our upstream network carriers +- **07:44~08:32** — We disconnected from the degraded network carrier at all US border routers. Traffic was successfully rerouted through other carriers. Conditions improved across most US paths, but latency and packet loss into US East had not fully recovered +- **08:39** — Paths through a secondary network carrier at the affected US East zone were still showing packet loss, as it was handing traffic back through the degraded carrier on the return leg. We disconnected from the secondary carrier there as well. Unknown to us at the time, that was the only carrier still supplying that site's default route (the catch-all path a network uses to reach the internet). The primary degraded carrier had already been disconnected, and our only remaining carrier’s connection there does not supply a default route. This left the zone without a stable route to the internet for roughly 20 minutes. During this window the disruption was at its most severe as traffic into and out of US East saw failed connections and heavily degraded private networking +- **08:59** — We reconnected our secondary carrier. Routing stabilized, US East connectivity began recovering, and the storage cluster returned to a coordinated, healthy state +- **09:00~10:45** — Storage performance in the zone remained degraded despite routing looking healthy. Throughput stayed pinned at roughly a third of capacity +- **10:45** — Root cause of degraded storage performance identified. Significant amount of storage connections had been established over a slow internal management network during the routing instability and remained stuck there after recovery +- **10:45~11:00** — We terminated the stuck connections across all storage and compute hosts in the zone. They reconnected over the correct network within seconds +- **11:04** — I/O wait across the zone returned to baseline (58% → under 5%). Storage throughput surged as the cluster caught up on backlogged writes, then settled at normal levels. At the same time, we identified that private networking tunnels had latched onto an incorrect address during the routing instability and never corrected themselves +- **11:49** — A fleet-wide restart of the mesh networking agents in the zone forced all tunnels to re-establish with correct addresses. Private networking fully recovered and full connectivity was restored across US regions; we continued monitoring +- **12:01** — With latency, packet loss, private networking, and volume performance stable at normal levels, the incident was marked resolved + +The full incident is available on our [Status Page](https://status.railway.com/incident/TAW34N30). + + +## What Happened? + +A few things went wrong that led to unintended cascading effects across our systems. The commonality across the failure cases were traced to stale connections that ended up capturing a bad path during a brief window of instability, and held onto it after the network recovered, because nothing in the system re-asserted the correct state. + + +### 1) Upstream ISP Degradation (US Regions) + +Datacenters connect to the internet by buying connectivity from transit providers; carriers that operate long-haul fiber and agree to deliver your traffic to any destination in the world. The largest of these are called Tier 1 ISPs. Railway connects every Metal datacenter to at least three of them, so that any single carrier can fail without taking us offline. + +On July 2, a carrier carrying our traffic was impacted by a network degradation somewhere in their US backbone. Traffic they normally carried on other paths spilled onto the route that carried our traffic between US West and US East, causing saturation leading to higher latency and packet loss. + +Our own routers showed no errors and healthy connections to every carrier, which told us the problem was upstream. Internal probes caught the degradation nearly two hours before it became visible to user traffic, giving us time to trace the lossy paths, all of which ran through the degraded network provider, while retries and redundant routing absorbed the loss. + +When the loss began reaching user traffic, we declared a public incident and disconnected from the degraded provider at all US borders. Traffic rerouted through other providers and packet loss returned to baseline on most US paths. + + +### 2) Storage Performance Degradation (US East) + +At 08:39, we also disconnected from a secondary carrier at this zone. Paths through the secondary carrier were still showing packet loss, because it was handing traffic back through the primary carrier’s degraded network on the return leg, and disconnecting a carrier on our side does not control the route traffic takes coming back to us. + +In hindsight, this change should not have been made without first verifying the behavior of default routes on our core switches. The zone is one of our first-generation sites, and unlike our newer datacenters, it gets its default route (the catch-all path to the internet) from its carriers rather than generating one itself. + +The primary network carrier was already disconnected at that site, and the only remaining carrier there does not supply a default route. Therefore, disconnecting the secondary carrier removed the last one. For the roughly 20 minutes until we reconnected, the site had no stable route to the internet. This window was the most severe part of the incident for users: traffic into and out of US East saw failed connections and degraded private networking until routing stabilized at 08:59. + +That instability exposed a hidden bug in how our servers behave when their primary network path disappears. Each server has two networks: a high-bandwidth fabric that carries production traffic, and a slow management network used for administrative access. + +When the fabric's default route vanished, the servers' operating systems fell back to the only route left: the one on the management network. A default Linux behavior then allowed servers to answer for their storage addresses on that network too, so storage traffic began flowing over a path with a small fraction of the fabric's capacity. + +Network connections don't re-check their route once established; they keep using the path they started on until they close. So the storage connections created during that 20-minute window stayed stuck on the slow management network even after the fabric was fully restored. The routing tables looked correct, and the storage cluster reported healthy, but storage throughput was capped at roughly a third of normal, and two thirds of servers in the zone sat waiting on disk while the cluster tried to push its backlog through. + +Once we found the connections coming from management-network addresses, we terminated those connections across every storage and compute server in the zone. They reconnected over the correct fabric within seconds, and I/O wait dropped from 58% to under 5% in about 15 minutes. + + +### 3) Private Networking Degradation (US East) + +Railway's private networking runs over encrypted tunnels between servers, and each tunnel learns its peer's address from the packets it receives. + +During the routing disturbance, tunnel traffic was briefly funneled through a device that rewrites the source address of traffic passing through it. Thousands of tunnels learned that device's address as their peer's address, and kept it after routing was rolled back. The mesh only re-verifies peer addresses when its membership changes, and because these tunnels sit silent when idle, a broken one never sends the packet that would have fixed it. + +At peak, roughly 20,000 host-to-host private network links were blackholed. This included inter-region services communicating with services in US East. Recovery required restarting the mesh networking agents across the fleet, forcing every tunnel to re-establish with the correct addresses. + + +## Preventative Measures + +We have already rolled out the following: + +- **Disconnected the degraded network carrier at all US borders.** Our fleet is currently operating on other Tier 1 carriers with full headroom. We will reconnect the degraded carrier once their backbone has recovered and we have verified path health +- **Cleared all stuck management-network storage connections** across every storage and compute host in the affected zone +- **Restarted the mesh networking agents fleet-wide** in the zone, re-establishing all private networking tunnels with correct addresses + +We are additionally working on: + +- **Migrating first-generation sites to self-generated default routes.** Our newer datacenters generate their own default route at the border rather than depending on carriers to supply one. At those sites, losing any single carrier is a minor path change, not a site-wide routing event. The affected zone predates this design, which is why it was vulnerable today. Bringing all remaining first-generation sites onto this pattern is the biggest structural fix from this incident +- **Correcting host fallback behavior so that production traffic never takes the management network as a fallback path** . Both networks are internal to our infrastructure and within the same trust boundary — traffic that crossed the management network stayed inside our own equipment, and private networking traffic remained encrypted end-to-end — so this was a path-selection and capacity problem. We will fix this by making a missing fabric route fail cleanly and immediately instead of silently taking a slower path, which will give us the ability to detect and resolve it in minutes +- **Alerting on management network utilization and blackholed private network links** , so that either failure mode pages us immediately instead of surfacing as degraded performance + +A carrier failure is a normal event on the internet, and our multi-carrier design handled it the way it should. The degradation that followed came from our side: an older site design that depended on carriers for its default route, a disconnection made without checking what routing would remain, and systems that captured a bad path during the instability and held onto it silently after the network recovered. + +We apologize for this outage and are actively working to prevent similar issues from happening again. Each of the fixes above targets one of those links, so that the next carrier failure ends where this one should have (with traffic quietly taking a different road). diff --git a/sreweekly/markdown/529/05-finding-bugs-in-raft-implementations.md b/sreweekly/markdown/529/05-finding-bugs-in-raft-implementations.md index 14b964b3..e34b8348 100644 --- a/sreweekly/markdown/529/05-finding-bugs-in-raft-implementations.md +++ b/sreweekly/markdown/529/05-finding-bugs-in-raft-implementations.md @@ -7,3 +7,289 @@ ## 简介 > …we rely on formal verification, and this is how consensus algorithms are built today. We define a model that we can mathematically prove to be correct, and then we… translate this perfect, platonic thing into code. + +## 正文 + +## [Introduction](https://antithesis.com#introduction) + +![TW Lim headshot](https://antithesis.com/_astro/069_tw.DHr19spq_1BL2As.jpg) + +The Raft consensus protocol is one of the foundations of the internet. Raft offers a [formal specification in TLA+](https://github.com/ongardie/raft.tla) and [a concrete, detailed implementation guide](https://raft.github.io/raft.pdf), and hundreds of groups have created open-source implementations of the algorithm by following the instructions in the guide. It is, by far, the most widely used consensus algorithm in production systems. + +Despite the paper’s famously accessible style, **we’ve found bugs in *every* Raft implementation we’ve tested**, including HashiCorp Raft, Aeron Cluster, OpenRaft, and MicroRaft — despite the investment in formal methods, careful code review, unit testing, and years of testing in production. The bugs we found manifest as violations of Raft’s main invariant (called *state machine safety* in the paper, commonly referred to elsewhere as *total order delivery*). If you’re using a Raft implementation, you might want to check it for bugs. + +We’ve sent bug reports upstream. This isn’t intended as a critique of Raft, its authors, its implementers, or any particular implementation. Raft implementations, even with a formal specification and a detailed implementation guide, are not easy to write. + +Rather, this is a story about correctness in distributed systems — the inevitability of bugs, the inadequacy of any single approach, and the high cost of learned helplessness. + +This is a long post. The first section is a position paper, but the rest is a detailed analysis of the issues we found, using one implementation as an example, a discussion of how we found them, and *why* we think they’re there. + +## [Why this matters](https://antithesis.com#why-this-matters) + +If you’ve worked on distributed systems, you’ve been part of a war room at some point (if you haven’t, don’t worry, you will be), dealing with an outage like [this one](https://blog.cloudflare.com/a-byzantine-failure-in-the-real-world/), or [this one](https://www.cockroachlabs.com/docs/advisories/a162085) — both of which were caused by consensus failures. When a consensus issue results in an incident, it’s inevitably far reaching and extremely painful to root cause, replicate, and fix. Many consensus issues remain “unsolved”, because they evade reproduction. + +It’s hard to know exactly what the real world consequences of these problems are, because consensus protocols sit so low in the stack. Are they corrupting the archive of someone’s personal D&D game? Security risks at a nuclear plant? Delaying trains in Germany? It all depends on where the protocol’s deployed. + +What’s certain is that these bugs are a colossal waste of developer time, and affect millions, if not hundreds of millions, of users. + +Yet we accept consensus issues and other deep distributed systems bugs as a fact of life — but once upon a time we accepted cholera, and waiting for your turn on the mainframe, as facts of life as well. Developers deserve something better. Everyone who depends on the software we write deserves something better. + +### [Bugs are not a fact of life](https://antithesis.com#bugs-are-not-a-fact-of-life) + +We have long accepted consensus bugs as inevitable because consensus protocols are just really, really hard to test properly. To really, thoroughly test a Raft implementation — or any distributed system — you effectively need to imagine every possible thing that could go wrong in your entire stack, and write a test to see if it will break your consensus protocol. It’s virtually impossible to write enough tests to provide this kind of assurance in critical systems. + +We get around this in two ways. First, we inflict the testing on our users, in the form of outages, downstream bugs, war rooms, and so on. It’s impossible for a team of developers to write enough tests, but with enough deployments and enough users across enough system configurations, you’re eventually going to find all the problems. + +Second, we rely on formal verification, and this is how consensus algorithms are built today. We define a model that we can mathematically *prove* to be correct, and then we… translate this perfect, platonic thing into code. Most Raft implementations are built this way. + +There are degrees here. The Raft protocol is one of very few consensus protocols that meets the strictest standard of “formal verification,” with a manual proof, a model check, and a mechanized proof-of-correctness. Every public consensus protocol we’re aware of has a manual proof/written correctness argument, but only Raft, Paxos, and MultiPaxos have mechanized proofs. + +### [Specifications are not code](https://antithesis.com#specifications-are-not-code) + +Formal verification is a great starting point, in that it helps confirm that a core design is sound. Implementing a distributed consensus protocol without starting from a formal spec would result in many orders of magnitude more bugs than any Raft implementation has. + +Conversely, one might wonder what the FoundationDB team might have done if they’d started with a formal spec. + +But the issues here demonstrate the difficulties inherent in going from a formal specification to actual production code. As long as the implementations are being done by unreliable, imperfect programmers (whether humans or LLMs), the actual code, and the systems using it, will only be as solid as the implementors’ assumptions — not the formally verified model. + +Furthermore, formal specifications are rarely truly complete — in the case of Raft, the TLA spec covers the core protocol, but doesn’t speak to details like `installSnapshot`, or replica replacement, or handling network packet corruption. + +### [Testing can actually be simple](https://antithesis.com#testing-can-actually-be-simple) + +Surfacing these bugs didn’t require particularly deep knowledge of consensus or distributed systems. We found these using a simple approach, which a junior engineer could implement in less than a day, literally the simplest state-machine replication workload we could come up with. The key is that this is happening while the system is running in Antithesis, subject to the kind of faults that happen in a real world environment. + +We want to emphasize this point because all too often, we encounter a learned helplessness around testing complex systems. Our profession thinks (with some justification) that writing tests for distributed systems requires more expertise, time, or tokens than we have. This results in a state of perpetual under-testing, which in turn results in extraordinary amounts of developer time being wasted on firefighting (to say nothing of the mental and emotional exhaustion). + +Developers deserve something better. Everyone who depends on the software we write deserves something better. + +## [Testing Raft](https://antithesis.com#testing-raft) + +![Marco Primi headshot](https://antithesis.com/_astro/001_marco_primi.C8sTreDD_1W9FFT.jpg) + +### [About Antithesis](https://antithesis.com#about-antithesis) + +Antithesis is an autonomous testing platform that runs distributed systems in a deterministic simulation environment, with randomly generated inputs, under aggressive fault injection. To use Antithesis, you deploy a full distributed system to the simulation environment along with a workload — a client that drives the system under test. + +Antithesis exposes the system under test to the kind of unpredictable turbulence it will experience in production, within the safety of a simulation environment. By randomizing the inputs and faults, and intelligently searching the state space of the system to see if system invariants are ever violated. + +We maintain an internal curriculum of systems and bugs we use to benchmark Antithesis’ bug-finding performance. [We’ve been adding consensus benchmarks to the curriculum, so we’ve been testing a lot of Raft implementations.](https://antithesis.com#fn-body-3) + +Part of the difficulty of building something unique is that you also need to build a way to measure its performance. + +### [About Raft](https://antithesis.com#about-raft) + +A common architectural pattern in distributed systems is state machine replication (SMR): all replicas in the system run copies of the same deterministic state machine, and the same sequence of commands is delivered to all of them to transform the data/state. This causes the replicas to progress in soft-lockstep — the state after applying N commands is *consistent* across all replicas, and all replicas *eventually* apply all commands. + +Distributed databases (FoundationDB, Aerospike, etc.), message queues (Kafka, NATS, etc.), key-value stores (etcd, Zookeeper), blockchains (Bitcoin, Ethereum), and many other systems are based on SMR. + +A fundamental requirement for implementing SMR is that all replicas *must* receive the same set of commands in the same order — total order delivery. Total order delivery is simple to describe, but hard to architect in the face of network turbulence, process crashes, and disk failures. + +Because total order delivery needs to be *guaranteed* for SMR to work, most systems rely on one of a handful of messaging abstractions, with atomic broadcast, also known as total order broadcast, being the most common. Raft is viewed as the friendliest atomic broadcast protocol, Paxos and viewstamped replication are also in wide use. + +### [Our testing approach](https://antithesis.com#our-testing-approach) + +To test a Raft implementation, we run a 3-node Raft cluster in Antithesis with a simple workload we call [Chain of Blocks](https://antithesis.com/docs/resources/chain-of-blocks/). + +Chain of Blocks tests Raft’s most important invariant: after N commands are applied, the state of all replicas matches. It consists of two trivially simple components: + +- A single-state state machine that just hashes the bytes of incoming commands. After each applied command, the state consists of . This is sufficient to allow us to observe state divergence between replicas. +- A stateless client whose only responsibility is submitting new commands that consist of random arrays of bytes. + +You can read more about this workload, and see a sample implementation, [here](https://github.com/antithesishq/hashicorp-raft-poc). + +Components like Raft are developed and reviewed by experts and battle-hardened by years of exposure. So it’s striking to us that this single, simple testing strategy — that a junior engineer could write in an afternoon (or Claude could write for a handful of tokens) — still finds bugs when it’s run in Antithesis. + +## [Results](https://antithesis.com#results) + +In all cases, network partitions and turbulence were sufficient to surface examples of divergence (i.e. no node kill/restart required, or disk corruption, or other faults) + +Here’s an example snippet from logs that shows a state machine divergence: + +`[node1]: Applied block 244 (2ad8d...) state: 01fd753a => faf6ba1e[node2]: Applied block 244 (c92fb...) state: 01fd753a => dab0977c[node3]: Applied block 244 (c92fb...) state: 01fd753a => dab0977c` +After applying 243 blocks, the hashes of all replicas matched (0x01fd753a). Node1 then applies a different 244th command from Node2 and Node3 (0x2ad8d… vs 0xc92fb…). This results in a different state hash after 244 commands applied (0xfaf6ba1e vs 0xdab0977c), violating the primary safety property of Raft (and the underlying atomic broadcast): that all replicas apply the same sequence of command in the same order. + +As an example, here’s a detailed look at the bugs we’ve found in one notable open-source Raft implementation during this process. We emphasize that we’ve found similar bugs in other implementations as well, but are going deep rather than broad here and only presenting one set of bugs to keep the length of this post somewhat manageable. + +We plan to update this post with details of bugs in the other Raft implementations we’ve worked with. + +### [HashiCorp Raft](https://antithesis.com#hashicorp-raft) + +![Rohan Padhye headshot](https://antithesis.com/_astro/rohan-padhye.C9feRA4F_Z1va7ov.jpg) + +HashiCorp Raft is a mature, popular open-source Raft implementation, and the foundational consensus engine underpinning widely-used production infrastructure tools like Consul, Nomad, and Vault. + +*We do not believe HashiCorp Raft is any less reliable than the other implementations we tested, or that the bugs in HashiCorp Raft are more serious than the bugs in other implementations* — they all violate the same core property of state machine safety. + +We found three distinct bugs in total: one causes numerous safety violations (i.e., data divergence as described above) and two cause liveness violations (where some nodes or the whole cluster cannot make progress unless an operator intervenes). The Chain of Blocks implementation we used is [here](https://github.com/antithesishq/hashicorp-raft-poc). + +When run with Antithesis for *just one hour* of testing, we see a report that looks like this: + +![Antithesis report.](https://antithesis.com/_astro/findings.CKPiWc3M_ZcPAsG.png) + +#### [Technical primer on Raft](https://antithesis.com#technical-primer-on-raft) + +To better understand the nature of these bugs, we should define some terms used frequently in the [Raft paper](https://raft.github.io/raft.pdf): + +- “Term”, “Leader”, “Follower”, “Election” + “RequestVote” RPC: Raft ensures consensus among a cluster of distributed nodes by dividing the protocol into strictly increasing *terms* starting with term=1. In each term, the nodes attempt to elect exactly one node as the*leader* by taking a majority vote; all other nodes are called*followers* . Leader election is conducted using a remote-procedure call (RPC) called*RequestVote* . +- “Log”, “Commit”, “Replication” + “AppendEntries” RPC: Each node maintains a replicated *log* of data entries (e.g., the commands issued by clients), some prefix of which is said to be*committed* ; that is, the node’s state machine is updated when an entry*commits* . When a client sends some command to the cluster, only the current leader can service this request, and the leader*replicates* the associated data entry to all its followers via an RPC called*AppendEntries* . When a majority of nodes in the cluster have replicated the same data entry in their logs, the leader*commits* it to its own state machine and broadcasts this fact to its followers. If some follower nodes fall behind because of network faults, or if their entry logs have uncommitted data from previous terms, then the current leader can always catch them up via more AppendEntries RPCs. The AppendEntries RPC is also used as a regular heartbeat mechanism to broadcast that a leader is active. +- “Compaction”, “Snapshot” + “InstallSnapshot” RPC: When the entry log grows too large, a Raft node can choose to *compact* all the committed entries in the log and only store on disk a*snapshot* of the state machine until that point. If a leader node needs to replicate data entries to followers via AppendEntries but its log has already been compacted, it can instead issue an*InstallSnapshot* RPC to transit the entire state machine to the follower. + +#### [**Bug 1 - Broken Consensus due to Async Heartbeats**](https://github.com/hashicorp/raft/issues/695) + +**Bug 1 - Broken Consensus due to Async Heartbeats** + +This is the most serious bug we have found. It can affect four of the five safety properties from the [Raft paper (Image 3)](https://raft.github.io/raft.pdf#page=5): + +- **Log Matching** (violated through mechanism A - seen in the report above) +- **Leader Completeness** (violated through mechanism A - seen in the report above) +- **State Machine Safety** (violated through mechanism A - seen in the report above) +- **Election Safety** (violated through mechanism B - not in the report shown above) +- **Leader Append-Only** (not violated) + +##### [Root Cause and Trigger](https://antithesis.com#root-cause-and-trigger) + +- The Raft paper assumes that all the operations (handling RPCs, client requests, elections, etc.) are atomic and don’t interleave with each other. Essentially, every node in the protocol is a finite-state machine. +- In HashiCorp raft, most state-changing operations are handled by a big switch-loop in the main thread, with several other floating goroutines doing async work that should not affect protocol state (e.g., peer-peer “replication” routines for dispatching append-entries, and per-connection “transport” goroutines that listen for incoming RPCs and hand them off to the main thread for processing). +- EXCEPT there is an [optimization that diverges from the protocol](https://github.com/hashicorp/raft/blob/4c8f61ac9255bb95fb3b8319dfcf0ae53ab325b6/docs/divergence.md#asynchronous-heartbeats) : incoming heart-beat messages (i.e., AppendEntries without an entry) are handled on the I/O “transport” thread itself, instead of queuing it up for the main thread’s loop like with other RPCs. Presumably, this is to quickly reset the keep-alive timers and reduce spurious re-elections. +- CAVEAT is that an incoming heart-beat message (like any other AppendEntries RPC) can change the state if it carries a new term when someone else was elected leader, and this bumps up the global currentTerm which everything else in the main thread relies on, and also sets the current state to be FOLLOWER. It is definitely not sound to perform these state changes concurrently with the main thread, and it can lead to different types of race condition bugs. + +##### [Mechanism A - Incoming heartbeat races with dispatchLogs() on main thread](https://antithesis.com#mechanism-a-incoming-heartbeat-races-with-dispatchlogs-on-main-thread) + +Here’s how the bug causes divergence in a 3-node setup A/B/C: + +- Assume the network link between node A and node B is down. +- Node A becomes candidate for term T, requests votes (eventually gets it from Node C) +- Node B also runs for term T but no votes +- Node B becomes candidate for term T+1, requests votes (gets it from Node C) +- Node B wins election for term T+1 +- Node A wins election for term T (just received the old vote from Node C) +- Node B sends out heart-beat AppendEntries with term=T+1, which Node A doesn’t immediately receive +- Node A gets a client request to apply a new data entry X, and it still thinks it is a leader for term T, so on the main thread it prepares to create a log entry (data=X, term=T) and send out AppendEntries to peers. + - HOWEVER: Just before it can create the log entry, the heart-beat from Node B sent in step 7 above reaches, setting currentTerm=T+1. Because this is done on the fast-path on the network-transport thread, it races with the main thread. + - The main thread ends up preparing a log entry (data=X, term=T+1) and also dispatching AppendEntries with these values before the next main-loop iteration where it realizes it is actually now a follower for term T+1. +- Things get really bad from here. Some nodes apply this bogus data=X, term=T+1 to their logs, while others follow whatever Node B (the true leader for term T+1) says, e.g., they might apply data=Y at the same index with term=T+1. Logs diverge in data but get committed since all say term T+1 for the same index. State machines diverge. + +##### [Mechanism B - Incoming heartbeat races with requestVote() on main thread](https://antithesis.com#mechanism-b-incoming-heartbeat-races-with-requestvote-on-main-thread) + +The same bug can cause another kind of safety violation when the async heartbeat handler races with a concurrent handler for the RequestVote RPC on the main thread. If both the incoming heartbeat and the RequestVote carry higher-numbered but different terms T1 and T2 respectively — this is possible when the receiver just recovers from a fault and is catching up to queued messages from peers — and if T1 is less than T2, then it is possible for the receiver to (i) realize the heartbeat’s term T1 is higher than its current term, then (2) on the main thread processing RequestVote realize that T2 is higher than its current term and so set its current term to T2 and then grant a vote; and then (3) back on the heartbeat handler’s thread set its current term to the lower value T1. Needless to say, it should never be possible for a node that has granted a vote for term T2 to regress its current term back down to a lower value T1! When this node at some later point increments its term to T2 again, it has no memory of the fact that it has already voted in this term. + +In summary, the race can eventually cause a node to grant two votes in the same term (T2 above), which in the worst case can lead to two different nodes being elected as leader in the same term! Here are some sample log messages from HashiCorp Raft when we have observed this race condition and its violation of election safety: + + `20:11:59.941133 [DEBUG] node-1: vote granted: from="node-2" tally=2 term=286 [...] 20:11:59.941133 [INFO] node-1: election won: tally=2 term=286 [...] 20:11:59.944273 [DEBUG] node-0: vote granted: from="node-2" tally=2 term=286 [...] 20:11:59.944274 [INFO] node-0: election won: tally=2 term=286` +#### [**Bug 2 - Deadlock after Leadership Transfer**](https://github.com/hashicorp/raft/issues/696) + +**Bug 2 - Deadlock after Leadership Transfer** + +This bug causes a liveness issue — a leader node is unable to service client requests or commit new entries while the cluster is healthy. The root cause is a deadlock during a leadership transfer operation ([a Raft protocol extension used by Consul](https://developer.hashicorp.com/consul/commands/operator/raft#transfer-leader) where you can ask a current leader to transfer leadership to someone else, say for upgrades) that does not immediately signal any warnings but bites you way into the future by getting the whole cluster stuck. + +- When you initiate a leadership transfer from A to B the current leader A first tries to get B’s logs up to date via an async replication thread. +- The leadership transfer logic is a separate goroutine that waits for this replication before telling B to take over. +- While all this is happening, A might have to step-down as leader for unrelated reasons (e.g., the lease timer runs out, or it cannot reach a quorum due to network) and say another node C becomes leader. +- When A steps-down as leader, it aborts all replication threads (including the A—>B catchup) but the leadership transfer goroutine is still blocked waiting for it to complete (an implementation bug). + +There is no problem so far, because C is the new leader and it does its job for a while. + +- Unfortunately, the leadership transfer goroutine in A has set a global flag “*leadership transfer in progress* ” and this flag cannot be unset until it is unblocked… but it never will be!!! +- Crucially, this flag does not prevent it from participating in elections and the flag does not get reset if it wins a future election. +- So in the future, A can legitimately get re-elected as leader but it will refuse to accept any client requests and stall on future commits with the reason “*leadership transfer in progress* ”. +- When the network is healthy and A is a leader, there is no way to get out of this other than rebooting the node A to clear the flag. + +We caught this using an Antithesis [eventually test command](https://antithesis.com/docs/product/writing_tests/test_templates/test_composer_reference/#eventually-command) that checks for commit progress after fault injection is turned off. + +#### [**Bug 3 - Livelock in Snapshot Installation**](https://github.com/hashicorp/raft/issues/697) + +**Bug 3 - Livelock in Snapshot Installation** + +This bug causes a liveness issue: a follower node becomes unable to replicate log entries and is effectively not able to participate in the consensus protocol. The bug can also cause resource exhaustion, where the node starts creating a potentially unbounded number of temporary snapshot files on disk. + +The root cause for this bug is painfully simple: in the version of HashiCorp Raft we tested, the implementation of the InstallSnapshot RPC simply does not follow Rule 7 from [Image 13 in the Raft paper](https://raft.github.io/raft.pdf#page=12) which handles the case when an incoming snapshot represents a state that diverges from the receiver’s uncommitted log entries. In the paper, Rule 7 says: + +“*discard the entire log*” + +HashiCorp Raft does not discard the existing log, and so after installing the snapshot its state machine can, in some cases, disagree with the data in its own (stale) log entries. + +At first, this might sound benign because the state machine should be the final source of truth. However, when the node receives a subsequent `AppendEntries` RPC from the leader, it faithfully implements a part of the protocol (Rule 2) that says: “[reject `AppendEntries`] if log doesn’t contain an entry at `prevLogIndex` whose term matches `prevLogTerm`”. So, because of the stale logs, the follower rejects subsequent `AppendEntries`, causing the leader to respond with a new `InstallSnapshot`, and this cycle goes on forever. + +We caught this using an Antithesis [eventually test command](https://antithesis.com/docs/product/writing_tests/test_templates/test_composer_reference/#eventually-command) that checks for state-machine convergence after fault injection is turned off. Sample logs from a test run look like this: + +``` +2026-07-07T17:17:13.073Z [INFO] node-0: snapshot network transfer progress: read-bytes=0 percent-complete="0.00%"2026-07-07T17:17:15.954Z [ERROR] node-1: failed to appendEntries to: peer="{Voter node-0 raft-node-0:8300}" error="msgpack decode error [pos 29848]: read tcp 10.89.0.7:48504->10.89.0.4:8300: i/o timeout"2026-07-07T17:17:23.073Z [INFO] node-0: snapshot network transfer progress: read-bytes=0 percent-complete="0.00%"2026-07-07T17:17:26.038Z [ERROR] node-1: failed to appendEntries to: peer="{Voter node-0 raft-node-0:8300}" error="msgpack decode error [pos 29848]: read tcp 10.89.0.7:35706->10.89.0.4:8300: i/o timeout"2026-07-07T17:17:33.073Z [INFO] node-0: snapshot network transfer progress: read-bytes=0 percent-complete="0.00%"2026-07-07T17:17:36.135Z [ERROR] node-1: failed to appendEntries to: peer="{Voter node-0 raft-node-0:8300}" error="msgpack decode error [pos 30394]: read tcp 10.89.0.7:37032->10.89.0.4:8300: i/o timeout"`2026-07-07T17:17:43.072Z [INFO] node-0: snapshot network transfer progress: read-bytes=0 percent-complete="0.00%"2026-07-07T17:17:46.249Z [ERROR] node-1: failed to appendEntries to: peer="{Voter node-0 raft-node-0:8300}" error="msgpack decode error [pos 29575]: read tcp 10.89.0.7:54324->10.89.0.4:8300: i/o timeout"`2026-07-07T17:17:53.073Z [INFO] node-0: snapshot network transfer progress: read-bytes=0 percent-complete="0.00%"2026-07-07T17:17:56.381Z [ERROR] node-1: failed to appendEntries to: peer="{Voter node-0 raft-node-0:8300}" error="msgpack decode error [pos 30303]: read tcp 10.89.0.7:57842->10.89.0.4:8300: i/o timeout"`2026-07-07T17:18:03.072Z [INFO] node-0: snapshot network transfer progress: read-bytes=0 percent-complete="0.00%"2026-07-07T17:18:06.626Z [ERROR] node-1: failed to appendEntries to: peer="{Voter node-0 raft-node-0:8300}" error="msgpack decode error [pos 31122]: read tcp 10.89.0.7:46978->10.89.0.4:8300: i/o timeout"2026-07-07T17:18:13.073Z [INFO] node-0: snapshot network transfer progress: read-bytes=0 percent-complete="0.00%"2026-07-07T17:18:17.032Z [ERROR] node-1: failed to appendEntries to: peer="{Voter node-0 raft-node-0:8300}" error="msgpack decode error [pos 30758]: read tcp 10.89.0.7:54192->10.89.0.4:8300: i/o timeout"2026-07-07T17:18:23.073Z [INFO] node-0: snapshot network transfer progress: read-bytes=0 percent-complete="0.00%"2026-07-07T17:18:27.745Z [ERROR] node-1: failed to appendEntries to: peer="{Voter node-0 raft-node-0:8300}" error="msgpack decode error [pos 31850]: read tcp 10.89.0.7:36218->10.89.0.4:8300: i/o timeout"2026-07-07T17:18:33.073Z [INFO] node-0: snapshot network transfer progress: read-bytes=0 percent-complete="0.00%"2026-07-07T17:18:39.083Z [ERROR] node-1: failed to appendEntries to: peer="{Voter node-0 raft-node-0:8300}" error="msgpack decode error [pos 33943]: read tcp 10.89.0.7:53700->10.89.0.4:8300: i/o timeout"2026-07-07T17:18:43.073Z [INFO] node-0: snapshot network transfer progress: read-bytes=0 percent-complete="0.00%"2026-07-07T17:18:51.725Z [ERROR] node-1: failed to appendEntries to: peer="{Voter node-0 raft-node-0:8300}" error="msgpack decode error [pos 38402]: read tcp 10.89.0.7:52390->10.89.0.4:8300: i/o timeout"2026-07-07T17:18:53.073Z [INFO] node-0: snapshot network transfer progress: read-bytes=0 percent-complete="0.00%" +2026-07-07T17:18:55.497Z [ERROR] eventually-fsm-convergence: convergence-check give-up; cluster healthy but did not converge: attempts_taken=24 per_node="map[raft-node-0:8400:map[applied_index:534 last_index:535 reachable:true state:Follower term:74 value:23641] raft-node-1:8400:map[applied_index:569 last_index:569 reachable:true state:Leader term:74 value:25232] raft-node-2:8400:map[applied_index:569 last_index:569 reachable:true state:Follower term:74 value:25232]]" +``` +## [Reflections](https://antithesis.com#reflections) + +![TW Lim headshot](https://antithesis.com/_astro/069_tw.DHr19spq_1BL2As.jpg) + +If there is a single lesson to be learned here, it’s that formal methods alone cannot ensure that software works, because the formal specification still needs to be implemented, and even if you have mechanized verification, you’re verifying the model and not the implementation itself. + +We believe formal methods *are* useful and necessary — they can confirm the basic soundness of a design, and provide a map that saves engineers from many of the errors that can arise in the implementation of a complex system. + +But as these bugs show, errors continue to arise when translating the formal specification to production code. In the course of our work with various Raft implementations, we identified a number of assumptions in the Raft paper that remain implicit. An implementer who misses any of these details is likely to run into trouble. + +### [Assumption 1: Each node is a synchronous process](https://antithesis.com#assumption-1-each-node-is-a-synchronous-process) + +![Marco Primi headshot](https://antithesis.com/_astro/001_marco_primi.C8sTreDD_1W9FFT.jpg) + +The paper implicitly assumes that each node is a synchronous process, performing atomic state transitions when handling RPCs and updating its own internal state. The [TLA+ spec](https://github.com/ongardie/raft.tla) is designed as such. + +At the same time, none of the implementations we’ve looked at have actually been a synchronous process, and there’s no explicit guidance in the implementation guide as to whether and how one can deviate from the synchronous design. + +Consider the following situation: +Can a leader with an outstanding RPC request respond to a vote, or does it need to wait for a `response||timeout`? Or should it respond, shut down the request and disregard a future response in order to stay correct? + +HashiCorp Raft is *mostly* synchronous in handling RPCs, with a single thread in a select loop, except when it isn’t — it does asynchronous heartbeat handling, bypassing its main loop, and this is what allows Bug #1 to creep in. + +### [Assumption 2: Each response can be mapped to a request](https://antithesis.com#assumption-2-each-response-can-be-mapped-to-a-request) + +In the Raft paper, every call is made using RPC, which means that every response can be mapped to a request. + +If one implements Raft on top of simple TCP or UDP, and doesn’t realize that request/response correlation is a crucial part of correctness, bad stuff can easily happen. + +If you’re not working in a language with a nice RPC library, it’s not obvious how to establish this correspondence. For example, the Raft authors have a simulation on their [website](https://raft.github.io/), which for obvious reasons is often considered a demo/canonical implementation even though they’ve never described it as such. But even their own implementation deviates from the protocol and [includes an extra `matchIndex` field in the AppendEntries response](https://github.com/ongardie/raftscope/blob/5b0c10ab51f873721895e7470b49e04c94bf826f/raft.js#L215C5-L215C15) in order to avoid matching requests to responses. + +### [Assumption 3: `currentTerm` and `votedFor` are consistent](https://antithesis.com#assumption-3-currentterm-and-votedfor-are-consistent) + +`currentTerm` and `votedFor` are consistent +![Rohan Padhye headshot](https://antithesis.com/_astro/rohan-padhye.C9feRA4F_Z1va7ov.jpg) + +Image 2 of the Raft paper states: + +![Raft paper figure 2 snippet.](https://antithesis.com/_astro/raft_paper_state.CAYonghL_Z1reAW3.png) + +One key assumption is that `currentTerm` and `votedFor` are consistent, because the latter records the vote in the current term. If these go out of sync bad things can happen, especially if `votedFor` is null. + +For safety, the writing of these values to persistent storage should be done atomically. But in the protocol, these values actually change at different places, so ensuring that both values are updated together is subtle. + +The online simulator in JavaScript does not actually distinguish persistent state from in-memory state, so there is no “reference implementation” of how to do this. + +HashiCorp Raft, for instance, deviates from the protocol and non-atomically stores three separate values for `currentTerm`, `lastVoteTerm` and `lastVoteCand`, with a worst-case risk to performance but not safety (i.e., it might persist an updated `lastVoteTerm` and then crash before updating `lastVoteCand`, but when it restarts it will just have a pre-determined vote for that term instead of choosing a candidate as normal). + +### [Assumption 4: The *entire* protocol is formally verified](https://antithesis.com#assumption-4-the-entire-protocol-is-formally-verified) + +*entire*protocol is formally verified + +Figures 2 and 13 in the Raft paper state, respectively: + +![Raft paper figure 2 snippet.](https://antithesis.com/_astro/raft_paper_append_entries.BaTxItye_Z1X9jN1.png) + +![Raft paper figure 13.](https://antithesis.com/_astro/raft_paper_install_snapshot.Bv6wVgAZ_Z2aSijv.png) + +Rule 2 for `AppendEntries` and Rule 6 for `InstallSnapshot` creates a potential pitfall. If you implement this naively when you have discarded your log because of a previous `InstallSnapshot`, you end up doing the wrong thing (most likely creating infinite loops of RPCs between leader and follower, as I have painfully discovered). + +The Raft TLA+ spec and the authors’ simulation don’t actually include `InstallSnapshot` (though [Diego Ongaro’s PhD thesis](https://web.stanford.edu/~ouster/cgi-bin/papers/OngaroPhD.pdf#page=66) does discuss the nuances in more depth), so different implementations use different workarounds. HashiCorp Raft, for instance, just avoids discarding logs altogether — which led to Bug #3 above. + +Bug #2 above, the deadlock, is in a feature called “leadership transfer” which is an extension to the core protocol. Like `InstallSnapshot`, this feature was never formally verified in the TLA+ spec. + +## [Conclusion](https://antithesis.com#conclusion) + +![TW Lim headshot](https://antithesis.com/_astro/069_tw.DHr19spq_1BL2As.jpg) + +We have long accepted bugs as inevitable. Distributed systems are almost impossible to test thoroughly because their state spaces are so large, and as our stacks get more complex, the problem only gets worse. + +But technological progress also moves the frontier of what’s possible. We ended cholera in the developed world by building water treatment plants and sewage systems. We no longer wait for mainframes because we’ve made computing power cheap and abundant. + +The abundance of compute is a recent phenomenon. It’s held true for maybe the last 15 years, less than a tenth of the history of computing. It’s spurred, among other things, the recent surge of interest in formal verification — the hope that mechanized proof and model checking can finally become easy enough, and cheap enough, that we’ll be able to apply these approaches wherever we need. But what our experience with Raft has shown is that implementations of even mechanically proven models can have flaws, because there’s no mechanical way to check the implementors’ assumptions. + +Fortunately, cheap and abundant compute has also made it possible to test software in ways we couldn’t before. And the other thing our testing of Raft has shown is that when you run systems under fault, even basic testing methods reveal issues that would once have taken months of manual testing to find. + +Developers deserve something better, and so does everyone who depends on our software. diff --git a/sreweekly/markdown/529/06-how-to-build-your-infrastructure-monitoring-in-2026.md b/sreweekly/markdown/529/06-how-to-build-your-infrastructure-monitoring-in-2026.md index 297bcf8c..9d12198a 100644 --- a/sreweekly/markdown/529/06-how-to-build-your-infrastructure-monitoring-in-2026.md +++ b/sreweekly/markdown/529/06-how-to-build-your-infrastructure-monitoring-in-2026.md @@ -7,3 +7,251 @@ ## 简介 I love that this starts with the user. Monitor what matters to your users, and alert on what you can action. + +## 正文 + +![How to Build Your Infrastructure Monitoring in 2026](https://omarghader.github.io/img/it_monitoring_guide_2026.png) + +# How to Build Your Infrastructure Monitoring in 2026 + +Every year I get asked the same question by teams starting from scratch: “we have Grafana, we have some dashboards, why do we still get paged for things we didn’t see coming?” Most of the time, the answer isn’t a missing tool. It’s a missing method. Teams jump straight to “let’s install Prometheus” or “let’s buy a SaaS observability platform” before answering a much simpler question: what does “healthy” actually mean for this business? + +I’ve built infrastructure monitoring from the ground up for several companies now, and I keep coming back to the same seven steps. This article is that playbook, the way I actually apply it in 2026. + +## Requirements + +Before you touch any tool, you need: + +- A clear list of the critical business flows your system supports (payments, checkouts, logins, API calls…) +- Buy-in from the team on what “acceptable” looks like for those flows +- A telemetry stack that can handle metrics, logs, and traces (I’ll give you mine below) + +**If you get stuck at any point**: reach out, I’m happy to help you think through your specific setup. + + +## 1. Start with the business SLI/SLO, not with the tool + +This is the step almost everyone skips, and it’s the one that matters the most. Before deciding what to monitor, decide what “working” means for your business. + +An **SLI (Service Level Indicator)** is a metric that reflects user-facing behavior. An **SLO (Service Level Objective)** is the target you set for that metric over a time window. + +Example, if you work for a banking company: + +- **SLI** : the ratio of successful payment authorizations over total payment authorization attempts +- **SLO** : 99.95% of payment authorizations should succeed over a rolling 30-day window + +That single sentence changes everything downstream. It tells you: + +- Which service is “tier 0” (payment authorization service) +- What your error budget is (0.05% of failed authorizations per month) +- What should page someone at 3am, and what can wait for Monday morning + +Do this exercise for every critical business flow before writing a single scrape config. If you skip it, you’ll end up monitoring infrastructure CPU graphs while your actual business metric silently burns through its error budget. + +## 2. Know what to monitor, then pick your stack + +Once your SLIs/SLOs are defined, list what you actually need visibility into to measure them: + +- **Infrastructure** : nodes, Kubernetes clusters, network, databases, message queues +- **Applications** : HTTP servers, HTTP clients, background jobs, gRPC services +- **Business layer** : the actual events tied to your SLI (a payment authorization call, a checkout event…) + +Only now do you pick the tech stack, because now you know what it needs to support. Here’s the generic stack I use on most projects: + +| Pillar | Tool | Role | +|---|---|---| +| Metrics | VictoriaMetrics | Long-term, cost-efficient metrics storage (Prometheus-compatible) | +| Metrics agent | vmagent | Scraping and remote-writing metrics | +| Logs | Loki | Log aggregation, indexed by labels not full text | +| Logs & traces ingestion | OpenTelemetry Collector | Vendor-neutral receiver/processor/exporter pipeline | +| Traces | Jaeger | Distributed trace storage and visualization | + +### The three pillars, and what each is actually for + +It’s worth being explicit about this, because teams often use the wrong pillar to answer the wrong question: + +- **Metrics** : aggregated, cheap to store, great for trending and alerting. They answer “what” and “how much” (error rate is 2%, p99 latency is 800ms). +- **Logs** : high cardinality, detailed, expensive to store at full fidelity. They answer “why” during an investigation (this specific request failed because of X). +- **Traces** : the causal chain across services. They answer “where” in a distributed call the latency or error actually happened. + +None of the three replaces the others. Metrics tell you something is wrong, traces tell you where, logs tell you why. + +## 3. Implement and scrape, favor auto-instrumentation + +Now you build the pipeline. My rule of thumb: instrument automatically first, add manual instrumentation only where auto-instrumentation doesn’t reach (custom business logic, internal queues, batch jobs). + +Use the [OpenTelemetry auto-instrumentation libraries](https://opentelemetry.io/docs/zero-code/) for your language. They hook into common frameworks (HTTP servers, HTTP clients, database drivers, gRPC) and emit metrics, traces, and sometimes logs without you writing a single line of instrumentation code. + +Example, a Java service auto-instrumented and shipped straight to your collector, no code change required: + +``` +java -javaagent:opentelemetry-javaagent.jar \ + -Dotel.service.name=payment-authorization-service \ + -Dotel.exporter.otlp.endpoint=http://otel-collector:4317 \ + -Dotel.metrics.exporter=otlp \ + -Dotel.traces.exporter=otlp \ + -Dotel.logs.exporter=otlp \ + -jar payment-service.jar +``` +On the infrastructure side, vmagent scrapes your Prometheus-format endpoints (node_exporter, kube-state-metrics, cAdvisor, database exporters…): + +``` +# vmagent scrape config +scrape_configs: + - job_name: 'kubernetes-pods' + kubernetes_sd_configs: + - role: pod + relabel_configs: + - source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_scrape] + action: keep + regex: true +``` +## 4. Build a telemetry pipeline that correlates, not just collects + +Collecting metrics, logs, and traces separately gets you three silos, not observability. The fix is enrichment: attach the same standard labels to everything, so you can pivot from a metric to a trace to a log for the exact same request. + +I always enforce the [OpenTelemetry semantic conventions](https://opentelemetry.io/docs/specs/semconv/) resource attributes as a baseline on every service: + +- `service.name` : the logical service, e.g.`payment-authorization-service` +- `service.namespace` : the domain/team owning it, e.g.`payments` +- `deployment.environment` :`production` ,`staging` , etc. + +Enforce this at the OpenTelemetry Collector level so nothing gets ingested without it: + +``` +# otel-collector-config.yaml +processors: + resource: + attributes: + - key: service.namespace + value: payments + action: insert + - key: deployment.environment + value: production + action: insert +service: + pipelines: + traces: + receivers: [otlp] + processors: [resource, batch] + exporters: [otlp/jaeger] + metrics: + receivers: [otlp] + processors: [resource, batch] + exporters: [prometheusremotewrite/victoriametrics] + logs: + receivers: [otlp] + processors: [resource, batch] + exporters: [loki] +``` +With these three labels shared across your metrics, logs, and traces, you can go from “error rate spiked on `payment-authorization-service` in `production`” straight to the matching traces and logs, without guessing. + +## 5. Build a RED dashboard, before anything fancier + +Once telemetry is correlated by `service_name` and `service_namespace`, build one dashboard before all others: the **RED dashboard** (Rate, Errors, Duration). + +Apply it to three things per service: + +- HTTP server requests (inbound traffic) +- HTTP client requests (outbound calls to dependencies) +- Span metrics (generated from traces, gives you RED per operation, not just per HTTP route) + +Example PromQL for the “R” and “E” of a service, using OpenTelemetry’s standard `http.server.request.duration` metric: + +``` +# Request rate +sum(rate(http_server_request_duration_seconds_count{service_namespace="payments"}[5m])) by (service_name) +# Error rate (%) +sum(rate(http_server_request_duration_seconds_count{service_namespace="payments", http_response_status_code=~"5.."}[5m])) by (service_name) +/ +sum(rate(http_server_request_duration_seconds_count{service_namespace="payments"}[5m])) by (service_name) +* 100 +# Duration (p99) +histogram_quantile(0.99, sum(rate(http_server_request_duration_seconds_bucket{service_namespace="payments"}[5m])) by (service_name, le)) +``` +This one dashboard, applied consistently across every service, gives you a clear signal of application health over time before you’ve written a single custom panel. Everything else (business dashboards, infra dashboards, deep-dive panels) builds on top of it. + +## 6. Alert on symptoms, not noise, and use multi-window burn rate + +This is where most on-call setups fail. Two rules I don’t compromise on: + +- **Alert on actionable conditions, not informative ones.** “CPU is at 80%” is informative. “The payment SLO error budget will be exhausted in 2 hours at this burn rate” is actionable. If an alert doesn’t require a human to do something right now, it shouldn’t page anyone; it belongs on a dashboard. +- **Use multi-window, multi-burn-rate alerting** so you can tell a critical incident from something that can wait until tomorrow. This is straight out of the[Google SRE book’s alerting chapter](https://sre.google/workbook/alerting-on-slos/) , and it’s the single highest-leverage thing you can implement for on-call sanity. + +The idea: page immediately only when the error budget is burning fast enough that waiting would breach the SLO. Use a short window to catch fast burns and a longer window to confirm it isn’t a blip, and use a slower threshold for tickets instead of pages. + +Back to our banking example: SLO is 99.95% success over 30 days, meaning an error budget of 0.05%. + +``` +# VictoriaMetrics / Prometheus alerting rules +groups: + - name: payment-authorization-slo + rules: + # Fast burn: page immediately. + # Burning 14.4x the allowed rate would exhaust the 30-day budget in ~2 days. + # Confirmed over both a 5m and 1h window to avoid paging on a blip. + - alert: PaymentAuthSLOFastBurn + expr: | + ( + sum(rate(payment_authorization_failed_total{service_namespace="payments"}[5m])) + / + sum(rate(payment_authorization_total{service_namespace="payments"}[5m])) + ) > (14.4 * 0.0005) + and + ( + sum(rate(payment_authorization_failed_total{service_namespace="payments"}[1h])) + / + sum(rate(payment_authorization_total{service_namespace="payments"}[1h])) + ) > (14.4 * 0.0005) + labels: + severity: page + annotations: + summary: "Payment authorization burning error budget fast, will breach SLO in ~2 days if it continues" + # Slow burn: create a ticket, review during business hours. + # Burning 3x the allowed rate would exhaust the budget in ~10 days. + - alert: PaymentAuthSLOSlowBurn + expr: | + ( + sum(rate(payment_authorization_failed_total{service_namespace="payments"}[1h])) + / + sum(rate(payment_authorization_total{service_namespace="payments"}[1h])) + ) > (3 * 0.0005) + and + ( + sum(rate(payment_authorization_failed_total{service_namespace="payments"}[6h])) + / + sum(rate(payment_authorization_total{service_namespace="payments"}[6h])) + ) > (3 * 0.0005) + labels: + severity: ticket + annotations: + summary: "Payment authorization error budget burning steadily, investigate this week" +``` +What this buys you on-call: + +- A fast, confirmed burn pages someone at 3am, because at that rate you’ll breach the monthly SLO within days. +- A slow burn opens a ticket instead of paging, because at that rate you have days to weeks before the budget is exhausted. + +This is the difference between an on-call rotation that burns people out on noise, and one that pages only when it truly matters. + +## 7. You now have the method, not just the tools + +At this point you have: business-defined SLIs/SLOs, a stack chosen because it fits what you need to monitor, auto-instrumented metrics/logs/traces, a correlated telemetry pipeline via standard labels, a RED dashboard per service, and burn-rate alerting that tells critical from “can wait.” + +That’s not a finished monitoring setup, it’s a solid foundation you can build on: business dashboards, capacity planning, chaos testing against your SLOs, whatever comes next for your organization. + +## Conclusion + +If this was helpful, leave a comment and tell me how your monitoring setup looks today. If you’d like help implementing any of these steps for your specific stack, I’m happy to walk through it with you. + +I wish you calm on-call shifts! + +## Building something like this in production? + +I help teams turn setups like this into reliable, monitored infrastructure. + +[Get a free consulting call](https://omarghader.github.io/contact/) + +Get my monitoring stack checklist + +The exact checklist I use when setting up observability for a new team. No spam, unsubscribe anytime. diff --git a/sreweekly/markdown/529/07-getting-access-to-the-tmp-of-a-systemd-service-with-privatetmp-yes.md b/sreweekly/markdown/529/07-getting-access-to-the-tmp-of-a-systemd-service-with-privatetmp-yes.md index 3830069f..0fbee5ec 100644 --- a/sreweekly/markdown/529/07-getting-access-to-the-tmp-of-a-systemd-service-with-privatetmp-yes.md +++ b/sreweekly/markdown/529/07-getting-access-to-the-tmp-of-a-systemd-service-with-privatetmp-yes.md @@ -7,3 +7,57 @@ ## 简介 First time I’ve heard of systemd’s PrivateTmp feature. Neat! + +## 正文 + +You're probably reading this page because you've attempted to access +some part of [my blog (Wandering Thoughts)](https://utcc.utoronto.ca/space/blog/) or +[CSpace](https://utcc.utoronto.ca/space/), the wiki thing it's part of. Unfortunately +you're using a browser version that my anti-crawler precautions consider +suspicious, most often because it's too old (most often this applies to +versions of Chrome). Unfortunately, as of early 2025 there's a plague +of high volume crawlers (apparently in part to gather data for LLM +training) that use a variety of old browser user agents, especially +Chrome user agents. To reduce the load on +[Wandering Thoughts](https://utcc.utoronto.ca/space/blog/) I'm experimenting with +(attempting to) block all of them, and you've run into this. + + If this is in error and you're using a current version of your +browser of choice, you can contact me at [my current place at the +university](https://www.cs.toronto.edu/~cks/) (you should be able to work out the email address +from that). If possible, please let me know what browser you're +using and so on, ideally with its exact User-Agent string. + +I am not blocking Inoreader's feed fetcher or considering it to be +too old, and it routinely fetches feeds from me. I don't know why +Inoreader is showing you this page. It is possible that they're +periodically trying to fetch feeds or pages with an old browser +HTTP User-Agent (or an actual old browser) and taking the results +of that fetch (this page) as what they should show people instead +of the results of their syndication feed fetcher agents. This is a +bad mistake today; [the results +of modern HTTP fetches depend partly on the HTTP User-Agent used](https://utcc.utoronto.ca/~cks/space/blog/web/HTTPResultsAndUserAgents). + + Much like Inoreader, Feedly is periodically fetching my syndication +feeds with a fake, old browser HTTP User-Agent header, which fails, +and is then grimly latching on to the results for their actual feed +fetching with their regular Feedly HTTP User-Agent. There is nothing +I can do about this; you should contact Feedly support, if you can +find them. See [this +comment of mine in Wandering Thoughts](https://utcc.utoronto.ca/~cks/space/blog/web/FeedReaderErrorsProblem?showcomments#cks-20260215115901) for more details. + + Due to an ongoing attack, you may need to change +[the +"User Agent Brand Masking" setting](https://help.vivaldi.com/desktop/miscellaneous/user-agent-brand-masking/) so that your Vivaldi identifies +itself as Vivaldi, instead of Google Chrome. This applies to even +the current version of Vivaldi. + + You may be seeing this through archive.today, archive.ph, archive.is, +and so on. Unfortunately, archive.* crawls pages to archive in a way that +is impossible to distinguish from malicious actors. They use old Chrome +User-Agent values, crawl from IP address blocks that are widely distributed +and not clearly identified as theirs, and some of their IP addresses have +falsified reverse DNS entries that claim they are googlebot IP addresses +(which is something that is normally done only by quite bad actors). I +suggest that you use archive.org, which is a better behaved archival +crawler and can crawl [my blog (Wandering Thoughts)](https://utcc.utoronto.ca/space/blog/). diff --git a/sreweekly/markdown/529/08-traditional-versus-resilience-engineering-views.md b/sreweekly/markdown/529/08-traditional-versus-resilience-engineering-views.md index 499d0729..7ded4f56 100644 --- a/sreweekly/markdown/529/08-traditional-versus-resilience-engineering-views.md +++ b/sreweekly/markdown/529/08-traditional-versus-resilience-engineering-views.md @@ -9,3 +9,29 @@ > I thought it would be a useful exercise to brainstorm some of the differences in focus between what I’ll call the traditional view of reliability, and the resilience engineering view. It’s short (just a table), but it definitely made me think. + +## 正文 + +As a fan of resilience engineering, I often differ with people on where we should focus our scarce engineering cycles in order to improve reliability. + +I thought it would be a useful exercise to brainstorm some of the differences in focus between what I’ll call the *traditional* view of reliability, and the resilience engineering view. + +| **Traditional view focuses on** | **Resilience engineering view focuses on** | +| accountability | coordination | +| prioritization | goal conflicts | +| risk mitigation | risk trade-offs | +| better processes and conformance thereof | more expertise | +| quantitative | qualitative | +| root cause | interaction of multiple factors | +| action items | insight | +| preventing future incidents, ensuring all incidents are novel | better handling of novel incidents | +| reducing complexity | navigating complexity | +| objectives | production pressure | +| robustness | resilience | +| human variability as liability | human variability as asset | +| building accurate system model | repairing inevitable model errors | +| rigor | improvisation | +| explicit knowledge | tacit knowledge | +| automation, benefits of | automation, risks introduced by | + +## One thought on “Traditional versus resilience engineering views” diff --git a/sreweekly/markdown/530/01-expertise-can-t-be-automated-why-resilience-still-needs-humans.md b/sreweekly/markdown/530/01-expertise-can-t-be-automated-why-resilience-still-needs-humans.md index fff2d4c8..cbde0c14 100644 --- a/sreweekly/markdown/530/01-expertise-can-t-be-automated-why-resilience-still-needs-humans.md +++ b/sreweekly/markdown/530/01-expertise-can-t-be-automated-why-resilience-still-needs-humans.md @@ -7,3 +7,94 @@ ## 简介 We may improve velocity by handing off tasks to LLM agents, but can that impact resilience? + +## 正文 + +![](https://d21hwc2yj2s6ok.cloudfront.net/shrine_store/uploads/networks/3057/communication_news/11560646/ultra_wide-2f370f4ed36790c72299e57b2ae83da5.webp) + +# Expertise Can't Be Automated: Why Resilience Still Needs Humans + +Abby Wambach, Neil Peart, Kelsey Hightower. These people are all experts in their relevant fields. Through them we recognize expertise as an exceptionally high level of performance on a particular task or within a given domain. The universal inputs to their expertise are time and exposure (aka practice). Expertise in any domain is acquired through doing, and more specifically by continued exposure to varying conditions and inputs inherent to the given domain. The latter part is important: while someone can absolutely become an expert in a very narrow sense (e.g. solving Rubik’s cubes), and someone can also learn about a subject or domain by reading, talking, listening, watching podcasts, etc., high-performance domains exert unpredictable, variable, and dynamic demands on the expert that are only mastered by continued exposure and deliberate practice within that domain. + +”Experts are the people the team turns to when faced with difficult tasks.” + +—Gary Klein + + +I have recently been on a number of podcasts/webinars/panels where the question of what expert incident response looks like comes up, especially how to train for and formalize it. Lurking behind that question for some people is often the hope that in the answer lies the opportunity to productize, scale, and even automate this essential skill for modern software-based businesses. We are currently awash in AI Ops and AI SRE solutions attempting to do just this, and I’m here to convince you that expertise cannot be bottled and sold. Beyond that, I hope to persuade you that expertise is worth investing in and cultivating within your own organization. + +# Common Characteristics of Expertise + +It is worth noting that expertise is not simply the accrual of experience over time. As hinted at in the earlier quote from Gary Klein, expertise is a specific kind of knowledge, and it shares a number of things in common. + +## Expertise Is Often Invisible + +Cognitive scientists and folks in the field of Resilience Engineering refer to this as the [Law of Fluency](https://www.youtube.com/watch?v=2Nc8hgEE7ig). This fluency characterizes an activity that is well-adapted, such that the effort and challenges involved in conducting that work are hidden from view, making it appear smooth and effortless. Experts adapt to complex and surprising situations by filling gaps and managing challenges effectively and efficiently, and notably in ways that may not be apparent to other people. In general, fluency is considered a hallmark characteristic of expertise. + +## Expertise Is Not Easily Introspected + +In general, experts do not have direct access to the cognitive, physical, and other sources of their expertise—human performance, given the way it is acquired through experience over time, is inherently difficult to explain. This is why many people refer to this kind of skill as implicit, or subconscious. Answering the question “How did you know to do X?” for an expert often leads to a long exploration of “Well, it reminded me of this one time…” or “It seemed similar to when…” In asking someone to explain their expertise, we are effectively asking them to try to unearth every second of practice and effort they have put into acquiring their skill. + +## Expertise Enables Adaptation + +People newer to a skill or domain typically rely on documentation and more rigid and rule-based methods for performing a given task. They lack the context and experience to handle variations from standard expectations and outcomes, and tend not to experiment or improvise as much. Experts, on the other hand, demonstrate a unique ability to adapt to novel information, challenges, and surprise situations. Given the breadth and depth of their experience (and having learned by trying many various approaches and strategies over time), they are better equipped (and generally, more confident) to stretch and adjust when presented with novel situations. Think of this as the difference between being able to read sheet music well and being able to play jazz with a group you just met. + +Consider the following [list of characteristics of expertise](https://www.researchgate.net/publication/322693359_Why_Expertise_Matters_A_Response_to_the_Challenges) that Gary Klein and a number of other cognitive scientists have empirically demonstrated. They found that experts: + +- +Employ more effective strategies than others, and do so with less effort; +- +perceive meaning in patterns that others do not notice; +- +form rich mental models of situations to support sensemaking and anticipatory thinking; +- +have extensive and highly organized domain knowledge; and +- +are intrinsically motivated to work on hard problems that stretch their capabilities. + +If that sounds like some of the fundamental aspects of resilience, you’re on the right track. + +# Expertise Is a Key Ingredient for Resilience + +Given the above characteristics of expertise and how they develop, expertise can’t simply be “bottled” and subsequently scaled or reproduced en masse. As noted above, most high-performance domains are extremely variable, so one element of maintaining expertise is constantly updating mental models and factoring in new variables, challenges, and developments. Layer the nature of modern distributed software systems on top of this, and you begin to see that one cannot simply capture expert incident response and productize it. + +Someone is inevitably going to say “Well we’ll just train our LLMs and agentic models on all the conditions necessary to generate expertise in incident response and then we can have that, right?” As they currently stand, LLMs and AI agents do not acquire expertise the way humans do, in part because they require human model development, training, and refinement. They also are not capable of sensemaking, reflection, coordination across fuzzy boundaries, reciprocity, vicarious learning from others’ experiences, observational learning (picking up things “in the air” from casual discussions), or “seeing the invisible” (perceiving missing vs. present cues)—these are all but a subset of the kinds of cognitive processes involved in developing and maintaining expertise that have been studied for decades by cognitive scientists. + +AI SRE agents can currently survey a given system’s environment and find inputs and patterns to generate hypotheses about how a given situation may have arisen. They can do this over and over and over, but they do not possess the knowledge or ability to, for example, call Sarah on the database team and see if she knows why things look weird. They can’t factor in the architectural discussions the team had about the recent migration which might explain why things aren’t behaving as expected. They can't notice that the system is behaving exactly like it did six months ago before a cascading failure, because that pattern lives in someone's memory, not in a log file. They can't pick up on the fact that the on-call engineer sounds unusually uncertain on the incident call, or that the team has gone quiet in a way that usually signals something is badly wrong. These aren’t edge cases, they are routine features of how complex incidents actually unfold. + +The prevailing belief in the software industry appears to be that we can use automation and AI to replace what expertise has given us in the past. However, consider the [data from the 2024 VOID report](https://www.thevoid.community/report-2024), wherein 75% of incidents involving automation required human intervention to comprehend, troubleshoot, and resolve the incident. That is where experts quite often save the day. Here we contend with the ouroboros of expertise and automation, in which experts don’t have as much experience with the system at hand, and then when asked to step in and resolve a problem with said system, they have found their expertise eroded by having less direct access to how it actually functions.  + +A lack of experience with systems due to increased automation and AI can lead to [de-skilling of experts](https://www.complexcognition.co.uk/2021/06/ironies-of-automation.html), by depriving them of the continued exposure to the complexities of the systems they are expected to understand as experts. Remember: expertise is accumulated through repeated exposure to a wide variety of situations over time. As we add more automation and AI to these systems, it becomes even more difficult for people to build expertise with those systems, much less to be able to introspect how automation and AI are impacting the functioning of these systems.  + +So what can organizations actually do to cultivate and protect the expertise that their resilience depends on? + +# Three Things Organizations Can Do to Foster Expertise + +1. +Develop and feed a culture where expertise is given the time and space it needs to develop and thrive. This means resisting the pressure to automate away the messy, variable, difficult work that builds expertise in the first place. It means recognizing that the engineer who has been in the weeds with a system for three years is an organizational asset, not just a headcount. +2. +Invest in building the skills required to do effective incident analysis. Post-incident reviews done well are one of the most powerful tools organizations have for surfacing, sharing, and building expertise. Not the checkbox RCA that documents what went wrong and assigns blame, but the rich, narrative, learning-focused review that asks how the system actually behaved and how your experts made sense of it under pressure. This is where tacit knowledge gets made visible. +3. +Help your experts identify and share the strategies and patterns they use. Incident analysis will help surface these, but the work doesn't stop there. Your organization needs specific approaches like knowledge elicitation, narrative storytelling, and communities of practice for distilling what your experts know into something the broader team can learn from. + +The vendors selling AI SRE and AI Ops solutions aren't necessarily wrong that these tools can help during incident response—they can survey environments, find patterns, and generate hypotheses faster than any human. But they are focused on a different problem than ensuring whether your systems are resilient. Resilience is continuously created by the people who have spent years inside your systems, who know what normal feels like, who remember what happened last time things looked like *this*. That expertise took time to build, and  it can't be purchased off the shelf. The good news is that it can be cultivated, incentivized, and shared. Expertise is contagious when organizations create the conditions for it to spread. + +### References + +[Peak: Secrets From the New Science of Expertise](https://www.amazon.com/Peak-Secrets-New-Science-Expertise/dp/0544456238) (Ericsson & Pool 2016)  + +[Seeing What Others Don't: The Remarkable Ways We Gain Insights](https://www.amazon.com/Seeing-What-Others-Dont-Remarkable/dp/1610393821/ref=asap_bc?ie=UTF8) (Gary Klein, 2015) + +[The Cambridge Handbook of Expertise and Expert Performance](https://www.amazon.com/Cambridge-Expertise-Performance-Handbooks-Psychology/dp/0521600812) (Ericsson et al, 2006) + +[Why Expertise Matters: A Response to the Challenges](https://www.researchgate.net/publication/322693359_Why_Expertise_Matters_A_Response_to_the_Challenges) (Klein et al., 2017) + +[The Ironies of Automation](https://www.complexcognition.co.uk/2021/06/ironies-of-automation.html) (Bainbridge, 1983) + + +![](https://d21hwc2yj2s6ok.cloudfront.net/shrine_store/uploads/networks/3057/communication_news/11560646/compressed-abd9e68f5fc007be1c81f15f1f28e000.webp) + + +**Courtney Nash** + +Vice President, RISF diff --git a/sreweekly/markdown/530/02-respecting-fatigue-isn-t-coddling.md b/sreweekly/markdown/530/02-respecting-fatigue-isn-t-coddling.md index 5994d3e8..b9884359 100644 --- a/sreweekly/markdown/530/02-respecting-fatigue-isn-t-coddling.md +++ b/sreweekly/markdown/530/02-respecting-fatigue-isn-t-coddling.md @@ -7,3 +7,59 @@ ## 简介 Fatigue and burn-out are reliability risks. Fatigue and burn-out are reliability risks. I champion this idea in my SRE practice constantly, and I hope you do too. + +## 正文 + +Is it coddling when an on-call engineer takes the next morning off to recover after handling a production incident at 3 a.m., or is it a smart company managing a reliability risk? + +Here’s what that night actually looks like. The engineer gets paged at 3 a.m., then spends two hours diagnosing the problem, coordinating with fellow responders, and restoring service. By 5 a.m., the incident is resolved and they get back to bed, but it takes them a while to settle down and get back to sleep. + +Four hours later, they’re at standup. That afternoon, they’re in a planning meeting. That night, they’re still primary on the pager. + +This is the default at most companies. Nobody made a deliberate decision that it should work this way; it’s just what happens when there’s no explicit policy for post-incident recovery. And it carries more risk than most leaders realize. + +## Incident response is more fatiguing than regular work + +Responding to an incident isn’t like a normal day of developing features and chasing bug reports. The cognitive demands are qualitatively different: rapid context-switching under time pressure, high-stakes decisions with incomplete information, coordinating across multiple people and systems, all while knowing that users are affected and stakeholders are watching. And there’s a physiological dimension that regular engineering work rarely triggers: adrenaline. Incident response activates the body’s stress response in a way that writing code or reviewing a design doesn’t. That heightened state feels productive in the moment, but it depletes reserves fast, and the crash afterward is steeper than the apparent effort would justify. + +This, incidentally, is one of the reasons that training and practice matter so much. Responders who’ve rehearsed the process and trust the framework around them experience a less intense stress response when real incidents hit. Turning incident response into a “[routine emergency](https://greatcircle.com/blog/2017/12/05/routine-emergencies/)” doesn’t just improve efficiency; it reduces the physiological toll. + +An engineer who’s been actively responding for a few hours isn’t just tired in the way that a long day makes you tired. They’re measurably less effective at exactly the skills incident response demands: integrating new information, evaluating competing hypotheses, making decisions under ambiguity, and recognizing when a current approach isn’t working. + +## The degradation is predictable + +Responder fatigue follows a recognizable pattern. As it sets in, people stop processing new information as effectively. They agonize over decisions they’d normally make quickly, or they stop making decisions altogether. They develop tunnel vision, fixating on the one theory they’re already pursuing instead of stepping back to consider alternatives. They fall into ruts, essentially pursuing “Plan A, again, with more feeling this time” instead of asking whether Plan A is still the right plan. They get less creative, more rigid, and more prone to mistakes. + +Fire departments study this, because it’s exactly the scenario firefighters face: interrupted sleep from overnight emergency calls, then back on duty the next day. The research consistently shows that the kind of fragmented, insufficient sleep they get around overnight calls degrades next-day cognitive performance to levels comparable to having had a couple of drinks. That’s why a growing number of fire departments have reconsidered their traditional 48-hour shifts; the performance degradation on day two is bad enough that departments are restructuring around it. + +## Self-reporting isn’t enough + +The insidious part is that fatigue undermines exactly the capacity you need to recognize it. A fatigued responder genuinely believes they’re performing normally. That’s simply how fatigue works. “I’m fine, I can keep going” isn’t evidence of fitness. It’s one of the symptoms. + +This is the critical organizational point. If a company leaves fatigue management to individual judgment (“take it easy if you need to”), it’s built a system that depends on impaired people accurately assessing their own impairment. + +Aviation learned this the hard way. The FAA doesn’t ask pilots whether they feel too tired to fly. It sets hard limits on duty time and required rest periods, because decades of accident investigation proved that self-assessment under fatigue is unreliable. Pilots who’d been awake for 20 hours consistently reported feeling capable. The data said otherwise. If you’ve ever had a flight delayed while the airline sought a new crew because the original crew had “timed out,” you’ve seen these rules in action. + +The tech industry hasn’t had its equivalent reckoning yet, but the same cognitive science applies. An engineer who handled a two-hour incident at 3 a.m. and says they’re fine at 9 a.m. may well believe it. That doesn’t mean they’re right, and building your next day around that assumption is a gamble most companies don’t realize they’re taking. + +## What active fatigue management looks like + +Companies that take responder fatigue seriously don’t rely on individual heroism or self-assessment. They build a few specific practices into their incident management capability. + +*Explicit rest expectations.* Not “take it easy if you need to,” but clear guidelines: an engineer who responds to a significant incident overnight is expected to start late or take the morning off, depending on duration and severity. The default is rest; working the next morning is the exception that requires a conscious choice, not the other way around. + +*The incident commander (IC) monitors for fatigue.* During extended incidents, it’s the IC’s responsibility to watch for fatigue signals in responders: slowed decision-making, tunnel vision, repeated questions, irritability, loss of situational awareness. This is the same responsibility a fire officer has for monitoring crew fatigue on a fireground. A fatigued responder who stays on the line isn’t being dedicated; they’re becoming a risk to the response, their teammates, and themselves. On incidents that stretch beyond a few hours, this includes planning responder reliefs early rather than waiting for someone to admit they’re spent. + +*Promote the backup.* After a significant overnight incident, consider moving the backup on-call engineer to primary for the next 12 to 24 hours. The person who spent two hours at 3 a.m. restoring service is not the person you want as your first line of defense if something else breaks that afternoon. + +*Rethink shift length.* Most teams seem to default to week-long on-call shifts with several weeks between shifts, but there’s a strong case for shorter, more frequent shifts. The same logic driving fire departments away from 48-hour shifts applies: shorter shifts mean less accumulated fatigue per shift, even if each person’s total on-call hours per quarter are similar. Shift design is a fatigue management decision, whether your company treats it as one or not. + +Some incident management platforms are starting to build this awareness into their tooling. [incident.io](https://incident.io/changelog/escalate-to-the-next-level#prompt-for-cover-after-a-bad-night-on-the-pager), for example, detects overnight pages and proactively asks the responder the next day whether they’d like someone to cover their next shift. That’s the right instinct: making fatigue management a system-level concern rather than leaving it to the judgment of the person who’s least equipped to assess it. + +## It’s a reliability decision + +Most companies aren’t actively choosing to ignore a fatigue problem. Rather, they have a fatigue problem that they haven’t noticed yet, because nobody has framed it as an operational risk. When a leader says “we trust our engineers to manage their own energy,” what they’re actually saying is: we have no organizational mechanism for ensuring that the people responding to our next incident are cognitively fit to do so. + +Respecting fatigue isn’t coddling. It’s protecting the quality of everything your engineers do the next day, including the next incident response. + +## Recent Comments diff --git a/sreweekly/markdown/530/03-on-building-scalable-control-planes.md b/sreweekly/markdown/530/03-on-building-scalable-control-planes.md index 05820206..c6f7f52a 100644 --- a/sreweekly/markdown/530/03-on-building-scalable-control-planes.md +++ b/sreweekly/markdown/530/03-on-building-scalable-control-planes.md @@ -7,3 +7,113 @@ ## 简介 A fun read on how to build control planes for large-scale systems, with some great tidbits on the inner workings of EC2 and Aurora DSQL. + +## 正文 + +## On building scalable control planes + +![Header image](https://www.allthingsdistributed.com/images/on-building-control-planes-that-scale-header.jpg) + + +[Zak van der Merwe](https://www.linkedin.com/in/zak-van-der-merwe-36a7b33a8/) has spent his entire career at AWS building control planes. First for EC2 and now for DSQL. On the surface, the control plane looks quite boring: it records what should exist and reconciles that with what actually does. Nobody leaves school dreaming of building one, but Zak will be the first to tell you that if you like solving hard problems in distributed systems, there are few better places to be. It’s where many of those hard problems converge, and where the decisions you make determine whether a service survives its own growth. + +*If you’ve been following [Marc Brooker’s](https://brooker.co.za/blog/2026/07/19/dsql-paper.html) and [Marc Bowes’s](https://marc-bowes.com/) writing on DSQL, this is a great companion piece that pulls back the curtain and shows what it means to build a database that was designed from the start with control plane engineers in mind.* + +*–W* + +## On building scalable control planes + +I’ve been working at AWS for nearly fourteen years, and for almost all of that time I’ve been building control planes. It’s not the kind of career anyone maps out for themselves. Nobody leaves university thinking “I want to spend the next decade making sure the bookkeeping layer of a cloud service stays up.” But here I am, and I think the reason I’m still here is that control planes turn out to be where many of the interesting problems live, even if it takes a while to see that clearly. + +Before Amazon, I worked at a telecoms company in Cape Town where we had maybe ten servers, all in a room in the back of the office, and every single one had a name. You’d SSH into them, you’d share them with your colleagues, and if something went wrong you could walk over and deal with it. That was my entire mental model of what it meant to run infrastructure. Servers were things you knew individually, took care of deliberately, and could reason about as a set because there were few enough to fit in your head. + +I mention this not because it’s an unusual background but because it was so common less than two decades ago, and I think that’s what makes it worth saying out loud. Maybe your version is a small Kubernetes cluster or a handful of RDS instances where you can visualize the whole thing, you can name the parts, and when something breaks you know which part broke. That feeling of knowing your infrastructure is comfortable, and it makes the next part of the story genuinely hard to describe, because what happened when I joined EC2 was that that feeling just evaporated. + +Honestly, when I started, I didn’t really understand how EC2 worked. I kept trying to map it back to what I knew. If I launch an instance and the underlying server dies, what happens? Does my VM somehow get teleported onto another host? How does the cloud create this illusion that hardware failures don’t matter? I couldn’t square any of it with what I knew about running software. + +My first job at EC2 was health-checking the fleet, pinging every server and trying to figure out if it was healthy or not, and what I found was the opposite of magic. Things were failing constantly. Hosts were going down, hardware misbehaving, disks dying. I had seen the underbelly of EC2 and it was chaotic. My mental model had gone from “servers are precious things you protect” to “everything is on fire all the time.” + +It took a while to shake that feeling, but what I would eventually come to realize was that these failures were tiny drops in an enormous ocean of things working fine. The system was just operating at a scale where failures were a constant, a statistical certainty rather than an emergency. And the thing that made it possible to run a service at that scale without a human responding to every failure, the thing keeping everything humming, was the ***control plane***. + +One way or another, my years at AWS have been spent working on control planes. Every AWS service has one, and I like to think of them as our unsung heroes. The better they work, the less anyone notices them. They’re the reason you don’t have to name your servers, and the reason that when hardware fails, you as a customer never have to deal with it. I’ve gotten to build control planes for two major AWS services: EC2, and DSQL. They’re nearly a decade apart, yet the hard lessons from building one led directly to the design of the other, and that’s the story I want to tell today. + +## What is a control plane anyway? + +At this point, I probably owe you a better explanation of what I mean by control plane and why I think they’re interesting. I’ll use EC2 as an example, because that’s where I learned most of what I know. + +The way I think about it is that every service has a data plane and a control plane. The data plane is the set of core capabilities, the raw computing power, the hardware, the networking. The control plane is the conduit between those capabilities and customers. It’s the thing that takes what exists physically in a data center and presents it to you in a format you can actually consume and get value from. Without the control plane, you’d be back to SSH-ing into named servers in a closet somewhere. With it, you can spin up a thousand machines with an API call and never think about where they live. + +![EC2 architecture diagram from Cape Town](https://www.allthingsdistributed.com/images/on-building-control-planes-that-scale-ec2-diagram-cape-town.jpeg) + + +EC2 involves thousands of engineers and more features than anyone can keep track of, and yet the control plane, conceptually… is pretty simple. Stripped down, EC2 lets you rent a virtual machine (VM) in the cloud, and the control plane’s job is to set up and tear down these VMs for you. + +I like the analogy of a thermostat, because it’s constantly measuring the temperature, it knows where things need to be, and it’s always nudging the system in the right direction. That’s what our control plane does. It’s a continuous loop, watching the state of the world, comparing it to what should be true, and correcting the difference. When you launch a VM, the control plane records that a VM should exist, finds a physical server in the right data center, sets up the image, configures networking, and launches it. Later, if that server disappears for any reason, the control plane notices and updates its records to reflect reality. It’s always reconciling what is with what should be. + +One thing the team talked about constantly, almost to the point where it became a mantra, was that no matter what happens to the control plane, VMs that are already running need to keep working. We call this static stability, and it sounds obvious because of course running VMs should keep running. But at scale, obvious things are the hardest to protect, because every new feature, every change, every dependency is a chance to accidentally violate that guarantee. Maintaining it is the difference between an outage where customers can’t launch new resources and an outage where everything stops. Both are bad, but the second is catastrophically worse. The fact that EC2 was statically stable gave me some comfort in my early days. + +The EC2 team has done a phenomenal job making bad days rare. But understanding what bad days look like shaped a lot of what I know about building control planes. + +## Living inside the control plane + +To understand how bad days start, it helps to know how the control plane stores state. At the heart of EC2’s control plane there is a relational database. When customers call the RunInstances API to launch a VM, the most critical thing that happens is that the control plane writes a row into its database: customer X now has VM Y. That’s when the API can safely return. + +In reality, a single `RunInstances` request triggers hundreds or thousands of internal API calls between micro and macro-services. Many of these services have their own databases recording their own state. It’s hard to exaggerate how complex this has grown over the years, but at the very bottom of all that complexity, there is a MySQL database, and what’s in that database is supposed to match reality. + +The simplest way things went wrong was also the scariest. Sometimes the primary database server just died. Our solution was a hot standby, a backup server continuously replicating from the primary, ideally only milliseconds behind. When the primary failed, we’d cut over to the standby and it could limit the outage to seconds. The team earned that through years of operational practice, building tooling, writing runbooks, training on-call engineers to execute the switchover under pressure. But seconds of outage still meant pagers getting lit up at 3am and asking humans to make decisions with incomplete information. We kept asking ourselves whether the architecture could take humans out of that loop entirely. + +The slower, more chronic problem was making sure our MySQL database kept up with business growth. This is pretty frustrating when you think about it, because the data plane does all the heavy lifting, like downloading VM images, configuring networking, running workloads, while the database is just keeping track of what exists. Every instance we launched meant more inserts, more updates, and more reads against the database, and eventually the bookkeeper couldn’t keep up with the workers. + +So we introduced more servers replicating from the primary and used these as read replicas. Many of the EC2 APIs don’t make any changes, they just describe the state of your current resources (how many VMs do you have, and so on). We sent traffic for these read-only APIs to our new read replicas and this massively reduced the load on our primary database server. This is standard practice for any team trying to scale up a relational database. Incidentally, this fleet of read replicas is why the EC2 API is eventually consistent, and as [Marc Brooker has written](https://brooker.co.za/blog/2025/11/18/consistency.html), this puts an unfortunate cognitive load on our customers. It’s something we wanted to do better with DSQL, which we’ll get to in a bit. + +Read replicas bought us time, but every write still funneled through a single primary server, and eventually we had to shard the database. The first phase of this was visible to customers as we split each AWS region into multiple availability zones (AZs), each with their own independent control plane and separate MySQL databases. This helped with both scaling and availability, since zones fail independently and the blast radius of any single failure shrinks. It also became a fundamental building block that allows AWS customers to build architectures resilient to the loss of a single AZ. The second phase was internal: we sharded each zone into what we call cells. Both of these projects took years of engineering time because they required changes across many services. Every place in the codebase that talks to the database has to know which shard to route to. Simple lookups by primary key are straightforward, but anything else, such as joins across data that doesn’t align with your sharding boundaries, gets much trickier. Even the simplest decisions have consequences at this level. Do you shard by account or by resource? Different services choose differently depending on their access patterns, and there’s no universally right answer. + +There is also a human cost to all of this that I don’t think we talk about enough. In those early years, we didn’t have the automation to handle a lot of what a modern control plane just takes care of. When a security vulnerability was discovered and the whole fleet needed to be patched, we didn’t have a system that could say “go update every host at a safe rate.” We would literally recruit the whole team, subdivide all the hosts, and assign shifts. Everyone in the Cape Town office would get a chunk. Go update every one of your hosts, report status. That’s what life looks like without a mature control plane, and it’s the kind of thing that doesn’t scale. You can patch a fleet of a few hundred hosts that way. You cannot patch a fleet of millions that way. The control plane is what eventually got humans out of that loop entirely. + +If you’ve lived through this progression, the scaling cliffs, the read replica tradeoffs, the sharding projects that always take longer than you think they will, you know it’s a long and painful road, and it’s one that every team building a successful service backed by a relational database eventually walks. + +## Searching for Database Xanadu + +After a decade working on EC2, I formed some strong opinions on what my ideal database looks like. It scales with my business without heroics. It is highly available with no downtime for updates, and no servers to babysit. My ideal database lets me leverage the power of the relational data model to model my domain and write software more productively. + +As it turns out, in the early 2020s, a group of experienced engineers on the databases side of AWS were thinking about exactly how to build this type of database. These engineers were expats from services like EC2 and had felt the pain of operating relational databases firsthand. They were also looking at the lessons learned operating massive scale serverless databases like DynamoDB and dreaming up ways to apply them to relational databases. + +They wanted to do for databases what EC2 and really Lambda did to servers. If you operate a traditional database with a “head node” you are in the world of “servers with names” like I was before joining EC2. The ideal database would free you from thinking about “databases with names”. Instead, it would have a control plane that takes care of all of that *for* you so that you can just think about your database as a logical endpoint that’s always available while it scales up and down. + +Sometime around 2021, this project really started to pick up steam. We’d figured out an architecture which seemed to deliver on this promise of the ideal database. I got the opportunity to join the team and start building its control plane. This service would launch in GA as [Amazon Aurora DSQL](https://aws.amazon.com/rds/aurora/dsql/) in 2025. + +Let’s quickly revisit the major pain points that EC2 went through and see how life is different on DSQL—especially for control plane builders. + +In DSQL, there isn’t one server running your database. DSQL spins up a Firecracker micro-VM per connection, which means every connection is its own small head node. If one fails, only that single connection is affected rather than your whole application. Nobody gets paged, no one has to decide to cut over. I don’t manage standbys anymore, because the architecture has removed humans from that painful loop entirely. + +Scaling reads was another problem we spent years on at EC2, adding replicas by hand and accepting eventual consistency as the cost. DSQL adds read replicas automatically, and in fact this is one of the primary jobs of the control plane that I helped build. If your application suddenly sees a spike in read traffic, DSQL handles it, and the reads are strongly consistent, always. After years of telling customers “try again in a moment,” this property still blows my mind. It fundamentally simplifies the architecture of any control plane built on DSQL, and it removes that cognitive tax from the developers using the APIs those control planes expose. + +And then there’s sharding, which was availability zones and cells at EC2 and took us years. When you build AWS control planes for major new services, you have to anticipate that sharding will become necessary, and experience has shown that it’s cheaper to do it from the start than to retrofit it later. This is an ugly dilemma, because you’re extending your time to market on a speculative future problem, and when delivery timelines get tight, I’ve seen many teams give up on sharding just to ship. DSQL removes that dilemma because it automatically partitions your workload and you don’t have to think about it. You can use all the Postgres goodies you’re used to, complex transactions, multi-table joins, secondary indexes, while knowing your database is going to scale with your needs. Many new AWS control planes over the last decade were built on DynamoDB for this same reason, but DSQL offers a world with fewer compromises. You get the scalability of DynamoDB with the relational programming model that developers actually prefer to work with. + +## “Self-hosting” + +When it came time to choose a database for the DSQL control plane, we chose DSQL. A team that runs on its own product feels every rough edge before its customers do, but getting there meant taking on the same circular dependency we’d faced at EC2: a control plane can’t depend on the thing it controls. + +We’ve seen two significant benefits from the decision to “self-host”. As customers adopt DSQL, they are creating thousands of databases, and the control plane is continuously scaling their databases up and down based on usage, often very rapidly. All of this customer activity creates “bookkeeping” work for the DSQL control plane, and the amount of this work grows with DSQL adoption. Since the DSQL control plane runs on DSQL, our bookkeeping database scales up to keep up with this increase in demand with minimal work from the team. + +The other benefit is in how we deal with availability zone outages. DSQL was designed from the ground up to survive single zone failures, but just because a zone is down doesn’t mean that customer workloads stop scaling or that customers stop creating databases. In my EC2 days, zone failures were fire storms as control plane databases died and pagers went off. For the DSQL control plane, these unfortunate bad days are much less painful because the DSQL control plane’s database remains available which allows the control plane to keep doing its critical work that ensures customer databases keep chugging along. + +## Taking off the rose-tinted glasses + +If you’re still with me, you’re probably thinking to yourself: “what’s the catch?” + +As a relatively new service, there are features that we just don’t support yet. Some of these are gaps that we’re actively filling. Others are more nuanced, and we want to take our time to make sure we build the right thing. A good example is foreign key constraints. Foreign key constraints are a classic database feature that can be very useful and aren’t fundamentally hard to implement. However, foreign keys can also be dangerous at scale. We want to get this right, and that takes time. + +One of the advantages of running Postgres on a single node is that it maintains the working set in memory, and cached reads are insanely fast. Real architectures are more complicated though. For example, a control plane using Postgres would run across multiple availability zones and put a connection multiplexing proxy in front of the database. These are necessary steps for availability and scale, but they increase latency. When you build on DSQL, you don’t need to manage these things yourself. You get good (though not quite single-node Postgres good) latency that remains consistent as your application scales. This is exactly what I want as a control plane builder. Yes, I want fast, but I care even more about predictable latency as my application scales. + +It’s also worth being honest about where things stand for control plane builders at AWS. Migrating something like EC2’s control plane onto DSQL would take years even if we started today, and that’s okay. The ten-odd years I spent on the EC2 control plane taught me that the work that matters most tends to measure its impact in years, not quarters. + +## Looking around corners + +We’ve spent most of this post deep in database scaling and life support. It’s a familiar shape for a lot of engineering stories. The problems we faced at EC2, how to go faster without breaking things, how to spend more of our time on the things that matter to customers, how to coordinate across a team that grew from a handful of people to thousands, and how to keep the system reliable while the ground shifted underneath us, are the same problems every engineering organization runs into as it scales. They are close cousins of the problems that produced Amazon’s original [distributed computing manifesto](https://www.allthingsdistributed.com/2022/11/amazon-1998-distributed-computing-manifesto.html) back in 1998, and my own focus narrowed over the years to a single version of them, which was how to let individual teams fully own a piece of EC2 and move fast on their most urgent problems without expensive coordination, all while the product still felt like one coherent thing to a customer. + +When I look at the broader industry today, I see echoes of that same pressure playing out at a scale I did not expect, because the arrival of agentic coding has driven the cost of writing software down to almost nothing, and that pushes the hard part of the work somewhere else. When code is cheap, the bottleneck moves to judgment, to figuring out what to build, how to ship it safely, and how to anticipate what your customers will need before they ask. That is the same shift a good control plane makes for the people who build on it, taking the invisible work of keeping infrastructure alive off their plate so they can spend their attention on their customers, only now it is happening to software development as a whole, and even a single-person team feels the need to scale out. + +I am not going to pretend I know what building software will look like a year from now, because we are in the middle of a remodel and the walls are still open. What I do know is that it is much easier to move fast when you are standing on a foundation that will not crack under you, and that the problems worth spending a career on have always been the ones that need your judgment rather than your ability to keep the bookkeeping layer from falling over. My hope is that DSQL gives the next generation of builders that foundation, and gives them back the time to go look around corners for their customers, which is the part I always wished we had more room for at EC2. + +And as Werner says: “Now, go build.” diff --git a/sreweekly/markdown/530/04-mario-saved-the-eu-but-broke-my-system.md b/sreweekly/markdown/530/04-mario-saved-the-eu-but-broke-my-system.md index 1ac597fc..57864d73 100644 --- a/sreweekly/markdown/530/04-mario-saved-the-eu-but-broke-my-system.md +++ b/sreweekly/markdown/530/04-mario-saved-the-eu-but-broke-my-system.md @@ -7,3 +7,105 @@ ## 简介 A harrowing incident story underlining the importance of expertise and experience. + +## 正文 + +![Simple thumbnail - conference table & headline reading 'Eurozone crisis live: Mario Draghi vows to save the eurozone']() + +### Ready to make incident response your competitive advantage? + +See how Uptime Labs builds provable, scalable incident response capability across your organisation. + +Fourteen years ago (almost to the day), I experienced an incident that I unlocked new insight into when I read  [Gary Klein's *Sources of Power*](https://mitpress.mit.edu/9780262611466/sources-of-power/). The following story illustrates why experience, not training, is what develops incident response expertise. + +### July 26, 2012: The Setup + +I was working at a trading firm. Everyone in the office knew the Eurozone was in trouble and that an announcement was coming. We'd prepared meticulously: load-testing all our trading systems to 10x their normal capacity, running through every scenario we could imagine. + +Around 11:30 BST, [Mario Draghi made his announcement](https://www.theguardian.com/business/2012/jul/26/eurozone-crisis-greece-bailout-barroso) at a global investment conference in London: the ECB would do ‘whatever it takes’ to save the Eurozone. It was a moment of relief for investors. The market started to move rapidly, and we could feel the pressure building on our systems. + +Then, 30 minutes later, everything broke. + +### The First Halt + +Around noon, everyone on our vast open-plan engineering floor suddenly jumped up from their desks. The horror on people's faces was unmistakable. Even the Chief Operating Officer rushed up to the engineering floor on the 6th level, which, in the normal order of things, is never a good sign. + +The system had simply stopped & nothing was working. More specifically, load balancers connections had shot up, web servers had exhausted all their connection and middleware services stopped processing. There were no warning signs; no alerts from our expensive monitoring tools. It was utterly inexplicable. A couple of minutes later, it resolved on its own, and everything came back to normal. + +For me, there was a sigh of relief. But the senior engineers who'd been in the industry long enough to know better looked visibly shaken. They knew something: *the worst type of incidents are the ones that resolve on their own.* That means you don't know what you fixed, and it will happen again. + +### The Pattern Emerges + +Sure enough, at 12:30, the exact same outage happened. This time, people came running back from lunch, panic rising. When it resolved again a few minutes later, nobody felt easy. We all knew: judging by the pattern, the next hit would come around 1pm. + +And that was the critical moment. One hour before the US market open at 2pm: the busiest, most consequential time of the trading day when American traders would come online and react to Draghi's announcement. So if the system failed then, the consequences would be catastrophic. + +### The Geeky Joke That Saved Everything + +While everyone was frantically scratching their heads, trying to explain what was happening, one engineer made an offhand comment. He said it looked like a ‘major garbage collection’ - a [Java thing](https://www.geeksforgeeks.org/java/garbage-collection-in-java/) related to memory management. A stop-the-world garbage collection event where the entire JVM pauses. + +He laughed. No one else did. But in that moment, something shifted. In the absence of any other lead, everyone in the room latched onto this intuition. It was a universal moment of singularity: we had something to pursue. + +Here's what I didn't understand then, but understood after reading Klein: *that engineer didn't randomly guess.* He saw a pattern. His experience with Java systems allowed him to recognise something that novices would have completely missed. This is the core of what experts see that the rest of us miss during incidents, in the words of Klein: + +*‘Intuition is when we use our experience, and the patterns we have learned, to rapidly size up situations and know how to respond without going through deliberate analysis.”* + +### Following the Thread + +We had over 1,000 services running across the platform. All the critical ones had garbage collection monitoring and alerting in place, and nothing was alerting. So we shifted focus to tier 2 and tier 3 services: the non-critical systems that might not have been monitored rigorously or might not even have GC logging enabled. We wouldn't have any way of knowing if something happened to them. + +At 1pm, the third outage hit. Now the war room was fully formed. The COO, the Head of Trading, compliance officers - people who rarely set foot on the engineering floor were all there, huddled together, discussing how to respond to clients, how to manage the incoming calls, how to prepare for what would happen at 2pm when every major client would be online watching their trades. + +Ten minutes later, an engineer buried in the tier 3 logs noticed something: a tier 3 service had been performing major garbage collections on a suspicious schedule. It was such a low-importance service that normally no one would have even looked at it. The only reason it got flagged was the time correlation with our outages. But critically, no one could explain how a garbage collection in that service could possibly affect the entire trading flow. We had no proof of causation. But we had no other leads, and we were out of time. + +### Decision Under Extreme Pressure + +We quickly huddled and worked out our options. What could we do? + +- Add more memory? +- Restart the server (rolling or full)? +- Shut down the service completely? +- Cut it from the load balancer? + +What fascinated me, and what I only understood years later reading Klein, was how the senior engineers evaluated these options. In less than a minute, they ran through each one mentally, simulating what would happen if we took that action i.e. *"If we add more memory, the next garbage collection will just be longer. If we shut down the service, the messaging broker it consumes from will pile up with messages and get flooded. If we isolate it from the load balancer..."* + +They settled on isolation. Cut it from the load balancer. It was reversible, surgical, and - as we later confirmed after the crisis - it was the only option that wouldn't have made things worse. + +### The Final Wait + +By the time we organised ourselves to isolate the service, we hit another episode at 1:30pm. The system halted again. A few minutes later it recovered, but we were running out of time. The business was already preparing contingency plans: how to apologise to the market, how to think about compensation, how to inform clients if this happened at 2pm. + +Then, about fifteen minutes before 2pm, we finally managed to cut the service from the load balancer. + +We had 15 agonising minutes to wait and hope for the best. We couldn't do anything else. It felt impossible. + +Then 2pm arrived & nothing happened. The sense of relief across the floor was overwhelming. + +### Why This Story Matters 14 Years Later + +I didn't fully understand why this story stuck with me until I read [Gary Klein's *Sources of Power*](https://mitpress.mit.edu/9780262611466/sources-of-power/). Suddenly, everything crystallised with new insight: + +Klein identifies 4 cognitive powers that emerge under the kind of pressure we experienced: extreme uncertainty, no clear clues and time pressure. These are the sources of power that separate experts from novices. + +**Intuition:** That engineer's joke about garbage collection wasn't a wild guess or a moment of whimsy. It was pattern recognition: the ability to rapidly size up a chaotic situation with minimal information. Experts see patterns that novices don't even know to look for. + +**Simulation:** In less than a minute, senior engineers mentally ran through each option one at a time, simulating outcomes. *‘If we do X, what else might happen downstream?’* This kind of scenario thinking was critical because we couldn't test anything or wait and see. We only had minutes. Novices would need to deliberate; experts could just simulate. + +**Metaphor:** Previous garbage collection incidents informed their reasoning. They remembered metaphors and examples where adding memory made things worse, which helped them discount those options quickly. Experience creates a library of patterns. + +**Storytelling:** This story stayed with me for 14 years. It's the kind of vivid incident that surfaces in memory under future pressure, making knowledge available not just to me but to anyone who hears it. This is how expertise spreads, and this is precisely why incident response training must evolve beyond classroom exercises. + +![John Allspaw, Founder and Principal, Adaptive Capacity Labs 10mo · There are only two ways people learn from incidents: Personal, first-hand experience Via the experience of others (i.e., vicarious learning) We cannot influence how or when #1 happens. We CAN influence how and when #2 happens. Vicarious learning is the only way effective learning from incidents can scale beyond one person. Creating conditions where this happens can be difficult. However, it's possible as long as there's broad recognition in the organization that... a. Effective post-incident analysis means building the richest understanding of the event for the broadest possible audience. b. The quality of post-incident analysis needed for (a) requires skill and expertise, in the same way experienced software engineers can produce higher-quality code more efficiently compared to when they first started. c. Most organizations do not have this expertise, but these skills can be learned and improved. (This is what we do.) Incident are being prevented all the time...in many cases, over 99% of the time! This takes effort, skill, and expertise. Enabling the broadest audience to learn something they didn't know before also takes effort, skill, and expertise.](https://cdn.prod.website-files.com/69eb654f1a323d235a57f701/6a74f8f61f33f0f1fe883b76_Screenshot%202026-08-04%20at%2007.27.02.png) + +*John Allspaw’s* *LinkedIn post* *advocates for systematic storytelling for vicarious learning - specifically, in the form of post-incident analysis.* + +### The Core Principle: Experience → Expertise + +Here's what I can articulate now, with Klein's help: the difference between an expert and a novice isn't innate brilliance. It's accumulated experience, whether real or simulated, that builds these 4 powers. + +Most incident responders can't wait decades for real incidents to teach them. That's where [simulation](https://www.uptimelabs.io/solutions/crisis-simulation) comes in. When you give a team carefully designed simulated incident experiences, you're not just teaching them facts. You're building their intuition, their ability to mentally simulate options, their metaphorical library of past incidents and their capacity to tell stories that will resurface under pressure. + +The things you gather from each incident, such as the visceral details, the decisions made, the outcomes, these stay with you forever. They become the patterns your brain recognises instantly. That's where the power is. That's how we reduce recovery time: not through static learning, but through building expertise one incident at a time. + + +![](https://cdn.prod.website-files.com/69e0a463268ba34093f8b1cb/69f225d89c7c20b5a5ec3c6e_Rectangle%2039651.png) diff --git a/sreweekly/markdown/530/05-ai-and-sre.md b/sreweekly/markdown/530/05-ai-and-sre.md index f9f7bb1f..9d2b1310 100644 --- a/sreweekly/markdown/530/05-ai-and-sre.md +++ b/sreweekly/markdown/530/05-ai-and-sre.md @@ -7,3 +7,63 @@ ## 简介 An SRE comes to terms with the way LLM agents are changing our field: what works well, what still requires human involvement, and what the future may look like. + +## 正文 + +# AI and SRE: The Same Force Cuts Both Ways + +I’ve spent most of my career in the “in-between” role — the one that sits between software engineering and operations, making sure the systems that carry the weight of the business stay up, stay fast, and stay honest about what they’re actually doing. Listen, follow the data, don’t do anything you can’t undo quickly. That’s been the job for a couple of decades. + +In the last year, the job changed more than it did in the previous ten. + +## The pro: a week of work in an afternoon + +I’m not talking about autocomplete. I’m talking about handing a genuinely open-ended problem — “why did this pipeline start drifting eleven months ago and nobody noticed,” “here are forty repos, tell me which ones are safe to migrate first and why” — to an AI agent and getting back, in an hour or two, the kind of analysis that used to take a week of careful, interruption-prone human attention. + +I’ve used this to cut a major cloud cost center by more than 90 percent, to untangle a production bug that had been quietly rotting for the better part of a year across a dozen services, and to review an entire database-upgrade runbook for gaps before we touched a production system carrying terabytes of customer data. None of that work was “prompt a chatbot and copy the answer.” It was iterative, it was checked against real logs and real output, and it was faster than anything I could have done alone at any point in my career. + +I’ve seen AI correctly identify issues with merge requests, pull requests. I’ve seen it correctly root cause incidents in minutes that would have a much longer time to diagnose collecting all the information manually. (I’ve also seen cases where it made the wrong diagnosis too.) + +That’s not a marginal improvement. That’s a different order of magnitude. And it’s real — I’m not describing a demo, I’m describing what shipped. + +## The con: a week of work in an afternoon + +Here’s the part our industry is not saying out loud enough: the thing that makes AI a superpower for an individual engineer is the exact same thing that makes fewer engineers necessary. If one person with the right judgment and the right tools can now do what used to take a small team, the small team is the thing that’s at risk — not the work. + +We’re already living in that answer. Layoffs in this field are accelerating again, and I don’t think that’s a coincidence of macroeconomics alone. Some of it is straightforward: the leverage AI gives a skilled engineer is being converted directly into headcount reduction, not just output growth. That’s not a hypothetical for me. It’s not a hypothetical for a lot of people reading this. + +I don’t think wringing our hands about it changes anything. I do think pretending it isn’t happening is worse than useless — it leaves people unprepared for a transition that’s already well underway. + +## So what’s actually different about the job now? + +If the leverage is real in both directions, the question worth asking isn’t “will AI take SRE jobs” — some of that has already happened, and more of it will. The question is what part of the job doesn’t compress the same way. + +A few things I keep coming back to: + +### Verification doesn’t get automated away — it gets more important. + +AI output is confident and plausible whether or not it’s correct. The instinct that’s always separated a good SRE from a dangerous one — don’t assert a root cause ahead of the evidence, widen the time window before trusting a correlation, verify the actual numbers instead of estimating — matters more now, not less, because the volume of plausible-sounding output you have to check has gone up by an order of magnitude too. + +The bottleneck moved from “can we generate an answer” to “can we trust this one,” and that second skill is still entirely human. + +### Judgment about what to build, and what *not* to automate, gets scarcer and more valuable. + +Anyone can now generate a script. Knowing which problem is actually worth solving, what the blast radius is if it’s wrong, and when the “clever” fix is a trap — that’s still earned the hard way, through incidents you’ve lived through. + +### The work shifts from doing to reviewing, and that’s a harder skill to teach. + +Reading someone else’s (or something else’s) work critically, catching the plausible-but-wrong answer, knowing which five lines of a five-hundred-line diff actually matter — that was always a senior skill. It’s now most of the job, for anyone using these tools seriously. + +### Communication and trust go up in value as raw output gets cheap. + +When everyone can produce more, the differentiator becomes who people trust to have checked it, and who can explain a tradeoff clearly enough that a room full of stakeholders can make a fast, confident decision. That’s always been true in aviation — the checklist doesn’t fly the plane, the pilot who knows when to deviate from it does — and it’s becoming just as true here. + +## Where I land + +I’m not going to pretend this is a comfortable transition, for me or for anyone else in this field right now. The pace of layoffs is real, and anyone telling you AI is purely additive to the job market isn’t looking at the same data I am. + +But I also can’t unsee what’s now possible. A week of work in an afternoon isn’t a slogan for me, it’s what happened, repeatedly, on real production systems. The honest position is that both things are true at once: this is a genuine force multiplier, and it is a genuine threat to how many of us the industry needs. Pretending otherwise, in either direction, doesn’t serve anyone. + +We can’t put the genie back in the bottle. There is a real cost; a human cost and the environmental costs. + +What I’m doing about it is the same thing I’d tell anyone to do with any new tool that changes the shape of a system: *don’t assume, follow the data, and don’t do anything you can’t undo quickly* — including your assumptions about your own career! diff --git a/sreweekly/markdown/530/06-certificate-expiry-is-still-taking-down-major-platforms.md b/sreweekly/markdown/530/06-certificate-expiry-is-still-taking-down-major-platforms.md index 57104091..c4970b6c 100644 --- a/sreweekly/markdown/530/06-certificate-expiry-is-still-taking-down-major-platforms.md +++ b/sreweekly/markdown/530/06-certificate-expiry-is-still-taking-down-major-platforms.md @@ -9,3 +9,13 @@ > Recent outages at Tailscale, jsDelivr, ServiceNow, and IPinfo show the same failure pattern: certificate automation broke quietly, while the expiry date kept moving closer. Bonus: they include links to several write-ups of related incidents. + +## 正文 + +Certificate Expiry Is Still Taking Down Major Platforms | TokenTimer + +TokenTimer - Expiration monitoring for tokens, secrets, and credentials + +TokenTimer helps teams track expiring API keys, secrets, service +accounts, certificates, licenses, and subscriptions with ownership, +proactive alerts, and audit history. diff --git a/sreweekly/markdown/530/07-how-stripe-uses-graph-search-and-state-machines-to-auto-remediate-a-gl.md b/sreweekly/markdown/530/07-how-stripe-uses-graph-search-and-state-machines-to-auto-remediate-a-gl.md index ad05ac87..2aa5b538 100644 --- a/sreweekly/markdown/530/07-how-stripe-uses-graph-search-and-state-machines-to-auto-remediate-a-gl.md +++ b/sreweekly/markdown/530/07-how-stripe-uses-graph-search-and-state-machines-to-auto-remediate-a-gl.md @@ -9,3 +9,11 @@ > Our solution treats infrastructure state as a traversable graph and lets a pathfinding algorithm discover recovery sequences at runtime. Whoa, cool trick! + +## 正文 + +Pragya Mehta is an engineer on the Document Database Control Plane Platform team at Stripe + +Sai Samant is a technical writer at Stripe working across the engineering organization. + +Explore our guides and examples to integrate Stripe. diff --git a/sreweekly/markdown/530/08-we-turned-off-pub-sub-and-nobody-noticed.md b/sreweekly/markdown/530/08-we-turned-off-pub-sub-and-nobody-noticed.md index 0798996a..50af41cf 100644 --- a/sreweekly/markdown/530/08-we-turned-off-pub-sub-and-nobody-noticed.md +++ b/sreweekly/markdown/530/08-we-turned-off-pub-sub-and-nobody-noticed.md @@ -9,3 +9,166 @@ Their event-oriented system was based on Google Pub/Sub with its 99.95% SLA, but their own SLA was 99.99%. To resolve that, they moved toward an active-active architecture, load-balancing across 2 message brokers. There’s an interactive simulation of their algorithm midway through that’s fun to play with! + +## 正文 + +August 11, 2026 — 17 min read + +Like many modern software stacks, the incident.io platform is predominantly [event-driven](https://en.wikipedia.org/wiki/Event-driven_architecture). For example, whenever you send us an alert, post a message to our agent on Slack, or update an entry in your Catalog - these are all events that then get enqueued on a message topic, meaning any of our downstream components that are interested in that event can subscribe and react asynchronously, such as sending a push notification or posting a reply to you in Slack. + +As the platform has grown, so has the number of messages flowing through our system, and thus our dependence on our messaging infrastructure. It’s become mission-critical. At the same time, we've also set stricter availability targets for ourselves, like the 99.99% SLA we now commit to for Enterprise customers of our [On-call](https://incident.io/on-call) product. + +Until recently, we only used a single provider - [Google Cloud Pub/Sub](https://cloud.google.com/pubsub) - as our messaging technology. Meaning any blip in Pub/Sub availability meant a blip in our own availability, which isn’t acceptable. So we recently set out on an adventure to make our messaging system more resilient to failure by introducing a secondary message broker to our stack, adding redundancy, and ultimately increasing the availability of our entire platform. + +**The goal was to be able to turn off Pub/Sub with zero customer impact. It turned out to be quite the adventure, but last week, we successfully did exactly that.** + +This is the story of that adventure. + +Event-driven systems have many benefits, like allowing us to decouple the rate at which we process messages from the rate of ingestion, or have many different components process the same original customer-initiated event, for example, having a `user.created` topic, and one system listens for events to send a welcome email and another that sets up their initial database state. + +At the heart of such a system is typically a “message broker”, which is responsible for receiving messages from the publishing components, storing them, and forwarding them to any interested subscribers. At incident.io, we’ve historically used Pub/Sub as our message broker of choice; it’s a [well-built](https://docs.cloud.google.com/pubsub/architecture) managed service with a good feature set, and has allowed us to scale with ease over the years. + +Pub/Sub is solid. Its published SLA is 99.95%, and in practice it has comfortably beaten that for us. The problem was that every event in our platform flowed through one broker, operated by one provider, with no way to route around it. **Our message broker had become a single point of failure (SPOF), and this was at tension with our own 99.99% availability targets.** + +And a SPOF is ultimately a question about accountability. When an escalation doesn't fire, "*Sorry, Pub/Sub was down*" is not an answer we ever want to give a customer. It's our SLA, and it's our job to meet it, whatever our dependencies are doing that day. We already have redundancy in the other layers of our infrastructure, so why should the message broker be treated any differently? + +When we say we have an event-driven architecture, we’re not exaggerating; we currently have ~800+ individual message topics, 1000+ unique subscriptions to those topics, and are processing ~240 million messages a day *(*as of August, 2026).* + +This large number of topics also means there are thousands of call sites in our codebase which interact with events, which can be quite daunting when you want to, say, replace the underlying technology you use for messaging 🫠. + +Fortunately, we were standing on the shoulders of giants, and the early engineers at incident were wise enough to build code-level abstractions over message publishing and subscribing, which we call the `eventadapter`. This is a package that exposes some simple but powerful interfaces like: + +``` +// Publisher is the interface for publishing events. +type Publisher interface { + Publish(ctx context.Context, ev Eventer, payload []byte) (string, error) +} +// Subscriber is implemented by all subscribers. +type Subscriber interface { + Subscribe(ctx context.Context, topicName string, handler SubscribeHandler, params SubscribeParams) func() error +} +// SubscribeHandler is what consumers of the package implement to +// handle a single event. +type SubscribeHandler[EV Eventer] func( + ctx context.Context, ev *EV, eventMetadata EventMetadata, +) error +// Eventer is the interface implemented by all events. +type Eventer interface { + // Name is how we identify this type of event. + Name() string + // A description of what the event means. + Description() string + // Validate validates the fields of the event before publishing. + Validate() error + // GetOrganisationID returns the organisation ID associated with + // the event, which we use to add to event telemetry. + GetOrganisationID() string +} +``` +The package exposes a couple of concrete implementations of these interfaces, like a `eventadapter.InMemory` for use in local development or tests, or `eventadapter.PubSub` for talking to Pub/Sub. + +**Fun fact:** we originally had to build the `InMemory` adapter as, in the early days of incident, running the app locally and opening so many parallel connections to Pub/Sub would crash the office wifi. 🙈 + +We conditionally choose and construct which version of the adapter to use at runtime in `func main()` based on the environment, and pass that down as a dependency to our application components. + +Having the luxury of such an abstraction meant that, to introduce a new message broker, what we needed to do was build a new implementation of the `eventadapter` interface, swap it in at runtime, without any of the calling code owned by other engineers being aware. This allowed us to hide all of the complexity that comes from load balancing across two different brokers behind the abstraction. + +One trade-off that came with using the existing interface meant that we are also constrained to the semantics and behaviors of that interface - which, even though it's an abstraction, already had some leakiness from the underlying technology. I.e. we needed to choose something that at least had the same feature set and behaviors as Pub/Sub, as these were semantics that we had come to rely on and reason about in our application. + +Our requirements and process for choosing a second message broker are out of scope of this post. There is a broad landscape of brokers: open-source vs proprietary, managed vs unmanaged, streaming vs non-streaming, ephemeral vs persisted, etc. There is no silver bullet, so our main advice is to document your own requirements and use a decision matrix. + +We chose [NATS](https://nats.io/) mainly because it’s a CNCF-adopted project, which means we can have confidence in its future and openly read the source code, is Kubernetes-native (which is where we run our workloads), is a single binary (we’re looking at you, Kafka!), and it is written in Go (the rest of our stack is Go!). Additionally, we had some prior experience running it. + +Another design principle was that, when [operating at 99.99% of availability](https://incident.io/blog/humans-arent-fast-enough-for-4-nines), failover between brokers can’t be manual; with ~4 min 23 secs of downtime budget a month, we don’t have time for someone to wake up at 4 am and switch brokers. This meant that, ideally, we had to use both brokers in an active-active setup continuously: messages get balanced across both, and any persistent error rate from one broker would automatically fail over to the other. **So we needed to build a dynamic event load balancer!** Let’s dig into how we did that. + +As discussed above, the first thing was to create a new concrete implementation of the `eventadapter` that we called `eventadapter.LoadBalancer`. + +On the publish side, we pick a broker by hashing the `message.ID` and rolling a weighted dice against a **configurable split**. The ID is a ULID (like all IDs in our system), hashed with Go's `hash/fnv` std-library (the [Fowler–Noll–Vo hash function](https://en.wikipedia.org/wiki/Fowler%E2%80%93Noll%E2%80%93Vo_hash_function)). Because every message hashes independently, a 50/50 split sends roughly half of *all messages* to each broker — the even distribution you want from a load balancer. + +💡 You may wonder why we sample by message ID here, instead of, say, our typical grouping key, which is organisation ID? Hashing by org ID would mean all of a customer’s events would get pinned to one broker until failover, whereas hashing per-message keeps load even by volume no matter how lopsided any one org is. The trade-off is that there is no per-org or per-operation transport consistency - which we’re happy to live with. + +In normal operation, we’ve chosen to have a 50/50 split across each broker; **why invest all this work in a secondary broker if you only use it in an emergency to then find out it's broken**? Importantly, the split is configurable without a deploy, in case we need to turn either broker off manually. + +Once we determine the preferred broker for a message, we attempt to publish the event, and if a publish attempt fails (maybe it timed out due to a short network blip), we fail over to the other. All attempts to a given broker also flow through a circuit breaker, so if many attempts in a short period of time start to fail as the broker is degraded, we short-circuit the publishes to that broker early and instantly fail over. Giving us the automated failover we need to reach our availability targets! No one gets woken up; publishes just gracefully start flowing to the other provider. + +You could argue that the publish-side is a pretty standard load balancer; where it gets more interesting is the subscriber-side and how we handle processing concurrency. + +The incident.io system is a single mono-repo Go program that is then deployed to Kubernetes as a collection of different workload-type-based deployments, such as `worker-oncall` or `worker-ai`, this allows us to do things like horizontally scale the number of replicas that receive inbound HTTP alerts independently of, say, our AI-message processing. We’ve talked about this architecture in more detail before: [Keep the monolith, but split the workloads](https://incident.io/blog/monolith). + +In the `eventadapater` interface, we also have similar controls over the number of concurrent message “handler” functions (or more accurately, goroutines) we run per-machine to process messages for a given topic, a setting called `MaxHandlers`. So that we can do things like: configure 10 handlers to process webhooks per-machine but only 3 handlers for a lower-priority background cron job. + +Therefore, we needed to think about how we could map the concept of concurrent handlers to the new dual-broker world. The rudimentary solution would have been to simply double the number of handlers, one group of handlers per broker. However, that presents a couple of issues: + +1. It would double the potential throughput of concurrent work on a machine and thus directly increase our resource usage (CPU/memory/network), and each handler needs DB connections to do its work, so we’d also have to increase the connection pool sizes and thus CPU pressure on our database. +2. More importantly, we might not always need an even split of handler processes per-broker. I.e. if Pub/Sub were to degrade overnight, and the publish rate flipped to an 80/20 split, the majority of messages would flow through NATS, and we would want the majority of our handlers to be consuming from the NATS queue and not Pub/Sub. + +**What we really needed was a dynamic scheduler that pulls messages fairly and prioritizes the broker which has more overall work.** + +So, like all good computer scientists, we did some research into prior art in this space and took inspiration from some existing queuing-theory algorithms. The most cited paper in this area is the MaxWeight algorithm ([Tassiulas and Ephremides (1992)](https://drum.lib.umd.edu/items/571fda52-aefb-4497-9a2d-69d8c7c907b9)), which can be summarized as: “*select the queue with the largest backlog”*. + +However, this didn’t align well with our setup, as we had no way to ***efficiently*** query each broker for the current queue depth on every pull. That led us to the delay-based variants of MaxWeight, which swap the weight variable from "*how many messages are queued*" to "*how long has the head message been waiting*", such as Oldest Cell First (OCF, [McKeown, Mekkittikul, Anantharam and Walrand (1999)](http://web.mit.edu/modiano/www/6.263/McKeown_copy.pdf)), and delay-based back-pressure ([Ji, Joo and Shroff (2011)](https://arxiv.org/abs/1011.5674)). These keep the same throughput guarantees, but only need one piece of information per-queue: the age of the message at the head of the queue. + +With our newfound queueing-theory knowledge, we set out to build our scheduler! + +Each broker (and its underlying Go client) is wrapped in a single-slot "Inbox” interface with just two methods: `Peek()` to peek at the metadata of the head message stored in the inbox slot, and `Receive()`, to pop the head message out of the inbox. + +We then have a scheduler goroutine that is responsible for the pool of message handler goroutines; the handlers are bounded by a [weighted semaphore](https://pkg.go.dev/golang.org/x/sync/semaphore#Weighted), which is sized using the `MaxHandlers` subscriber setting we discussed before. + +Whenever a slot frees up in the semaphore, the scheduler calls `Peek()` on both inboxes and dispatches whichever head message has the oldest publish time (a broker-assigned timestamp, so it's immune to clock skew between publishing pods), by calling `Receive()`, which returns the message and then refills the inbox slot from the broker over the network, in time for the next peek. (We’ve glossed over some detail here, such as each broker’s native client also buffers some messages in-memory, to be more efficient). + +Putting this all together gives us all the properties we desired from our scheduler: + +- Reusing the existing `MaxHandlers` concurrency gives us one shared concurrency budget across the subscriber for both brokers, and no change in throughput semantics or resource consumption. I.e. our engineers can continue to set`MaxHandlers` and don't have to worry about the fact that the consumption is occurring via two brokers. +- Using “oldest message first” scheduling gives us a weighted fairness policy; if we fail over to one broker and its queue backs up, the scheduler will prioritize its inbox to fill the semaphore slots. + +Magic! ✨ + +So, we had a working implementation of our event load balancer and were pretty pleased with ourselves. Over the course of the last month we’ve been carefully rolling out the load balancer, first in our staging environment, then production topic-by-topic, until all messages now flow through it - including our most critical On-call escalation events. + +However, as all good reliability engineers know, an unexercised code path or failover mechanism is as useful as not having one. How do we know this would actually protect us against failure in either broker? **The only option was to chaos test in production!** + +First, with all of our messages flowing 50/50 across both brokers, **we intentionally deleted our NATS cluster** in our production Kubernetes environment. We watched the messages flow over to Pub/Sub gracefully, without client-facing errors. + +Testing Pub/Sub was a more interesting challenge: we don’t manage it, and can’t really call up Google and ask them to break it intentionally. So we built an automated fault injection system into the load balancer. The fault injector allows us to simulate partial degradation or total outages on either provider, and we can now dynamically increase or decrease the fault tolerance via configuration. + +Equipped with our shiny new chaos tooling, **last week we turned off Pub/Sub in production, and no one noticed.** The first few publish attempts time out or error and are then automatically retried against NATS; after 30 seconds of failed publish attempts, the circuit breaker opens, and all subsequent publishes short-circuit - not a single message dropped. All of this happened without any of our internal engineers being paged or a single customer noticing. **The system worked 🎉** + +Importantly, our chaos tooling now allows us to run these failure scenarios in production, continuously, and with ease, just like any other Tuesday. + +We initially set out on this project to make our On-Call product more reliable and eliminate one of the only remaining single points of failure in our system. However, along the way we’ve improved one of our core system primitives and raised the availability bar for our entire platform. + +After a couple of weeks running the system in production, we are already starting to see the benefits: + +- With improved telemetry and observability into our publish failures, we are regularly observing timeouts due to network blips to Pub/Sub, which now instantly fail over. Messages that may have previously caused customer-facing errors are now simply retries. +- We can more safely operate and maintain both providers independently, without risk of data loss. Need to upgrade the NATS cluster? Easy! Change the Pub/Sub client? Safe! + +We’d love to hear what you think about the design and, as always, if working on these types of reliability challenges interests you - [we’re hiring](https://incident.io/careers)! + +Patrick Hamann + +Product Engineer + +Mike Fisher + +Our learnings from implementing a product-wide read replica migrations, including some useful patterns for routing queries to replica and primary + +Johanna Larsson + +July 21, 2026 + +Today we're launching our new post-mortems experience, and I want to walk you through what we've done and why. + +Pete Hamilton + +March 17, 2026 + +This post is a deep dive into how we improved the P95 latency of an API endpoint from 5s to 0.3s using a niche little computer science trick called a bloom filter. + +November 14, 2025 + +Ready for modern incident management? Book a call with one of our experts today. + +- All-in-one incident management +- Our unmatched speed of deployment +- Why we’re loved by users and easily adopted +- How we work for the whole organization diff --git a/sreweekly/markdown/531/01-heroic-saves-are-near-misses.md b/sreweekly/markdown/531/01-heroic-saves-are-near-misses.md index 495cc4b5..d94824f9 100644 --- a/sreweekly/markdown/531/01-heroic-saves-are-near-misses.md +++ b/sreweekly/markdown/531/01-heroic-saves-are-near-misses.md @@ -7,3 +7,55 @@ ## 简介 Celebrating heroes in incident response can incentivize further heroics. That can prevent the kind of growth that will improve incident response overall. + +## 正文 + +It’s 3 a.m. in California, where most of the dev team are still snug in their beds. The auth system has started rejecting valid credentials. Early bird East Coast customers are already trying (and failing) to log in for the day, and thousands of users in Europe have already given up and gone elsewhere. In a couple of hours, the West Coast will be waking up too. A brilliant engineer swoops in and saves the day. She has legendary debugging skills and a deep understanding of the auth system, and she puts together a fix in forty minutes that would have taken anyone else hours to even diagnose. Later that morning, leadership is sending thank-you messages in the all-hands channel. Her VP awards her a small spot bonus, and her manager reminds her to include it in the next performance review cycle. + +What doesn’t usually happen is anyone asking: what if she hadn’t been there? Because that heroic save, for all the heartfelt celebration around it, was actually a near miss from a systemic point of view. + +## Near misses look like successes + +In aviation and other safety-critical fields, it’s widely accepted that a near miss is an unparalleled opportunity to learn and deserves the same investigation as an actual failure. The reasoning is straightforward: a near miss reveals the same systemic vulnerabilities that a failure does. The only difference between a near miss and a disaster is that the outcome happened to be good this time, often because of luck, timing, or the presence of one specific person. + +A heroic incident response is a similar opportunity. The system nearly failed, and would have failed if that one engineer hadn’t been available or hadn’t known exactly what to do. Her skill, expertise, and dedication are worth appreciating. But her unavailability would have meant a much worse outcome, and that’s worth examining too. Too many companies celebrate the save and stop there. + +## The incentive nobody designed + +When a company celebrates a heroic save without examining why the heroics were necessary, it sends a message. The message isn’t intentional, but it’s clear nonetheless: what gets valued is the dramatic rescue, not the boring preparedness work that would have made the rescue unnecessary. + +Over time, that message shapes behavior. The engineer who writes thorough runbook documentation, trains new team members on the auth system, and invests in monitoring improvements doesn’t get the same recognition as the one who swoops in at 3 a.m. and saves the day. Preparedness work is largely invisible in performance reviews. Heroic saves are memorable. + +The result is a perverse incentive loop. Heroics get rewarded, preparedness doesn’t, and the company remains dependent on heroic saves because nobody is investing in the alternative. This isn’t because anyone explicitly decided that preparedness doesn’t matter. It’s because the reward system is quietly rewarding the wrong thing, and nobody has noticed because the heroes keep delivering results. Until they don’t. + +In my experience, this is one of the most common patterns in companies that are struggling with incident management. They have talented, dedicated people who keep delivering heroic results, and because the results keep coming, nobody realizes there’s a growing structural problem underneath. + +## The hero as single point of failure + +The incentive loop creates a second problem. The hero gradually becomes a bottleneck and a single point of failure. When that engineer is on vacation and the next auth system incident hits, the team might spend hours just figuring out what’s going wrong, let alone fixing it. When they eventually leave the company (as they likely will; heroes tend to burn out), the team discovers that critical knowledge walked out the door with them. + +I see this pattern regularly in my consulting work. In a company’s most serious incidents, it keeps turning to the same handful of heroic engineers. Those engineers are talented and committed, and their involvement has genuinely saved the company from significant damage. Everyone involved with incidents knows who they are, and breathes a sigh of relief when they join an incident channel. But the company has never seriously examined what its response capability looks like without them. The term that often comes up to describe these people is “indispensable,” which is really another way of saying that the company’s incident response capability depends on specific individuals’ availability. + +## Why the problem stays hidden + +The most insidious aspect of this pattern is that it’s invisible to leadership for as long as the heroes keep delivering. Companies at the earliest stages of incident management maturity often don’t realize they’re at risk. Leadership sees consistently good outcomes and assumes the company has strong incident response, when what they actually have is strong individuals (and a certain amount of good luck). + +By the time the fragility surfaces, the gap between where the company *thought* it was and where it *actually* was can be startling. + +## Heroic is a growth stage, not a compliment + +When I assess incident management capabilities for my consulting clients, one of the dimensions I evaluate is program maturity: where is this company on the growth path from ad hoc response to reliable organizational capability? The first stage on that path is called “Heroic.” It isn’t meant to be flattering. It means that incident response quality is a property of specific talented individuals rather than a property of the company. When those individuals are available, things go well. When they’re not, things go sideways. + +Every company starts here. The question is whether they invest in growing past it, converting individual capability into organizational capability. That transition is what the rest of the maturity model describes, and it’s the core of what effective incident management programs are designed to do. + +## What to recognize instead + +None of this means companies should stop recognizing heroic contributions when they happen. When someone saves the day at 3 a.m., thank them. But also investigate why the heroics were necessary, and invest in the answers. That’s a form of recognition too: it says the save mattered enough to learn from. + +To move from “Heroic” to higher levels of organizational capability, you need to shift what gets sustained recognition. Recognize the work that makes heroic saves unnecessary: the runbooks, the training, the well-coordinated responses where nobody had to be heroic. + +If an engineer spent much of their quarter writing runbooks, training new responders, and coordinating incident responses, recognize that work: in performance reviews, in public acknowledgment from leadership, in awards and bonuses. If you don’t, you’re telling your organization that the only incident management work worth noticing is the dramatic save. + +The goal is to make effective incident response something the company can do reliably, regardless of who happens to be on call. Heroes are still welcome, and still admired. They just shouldn’t be required. + +## Recent Comments diff --git a/sreweekly/markdown/531/02-control-and-complexity-tension-in-systems-design.md b/sreweekly/markdown/531/02-control-and-complexity-tension-in-systems-design.md index e105f893..97d0d17d 100644 --- a/sreweekly/markdown/531/02-control-and-complexity-tension-in-systems-design.md +++ b/sreweekly/markdown/531/02-control-and-complexity-tension-in-systems-design.md @@ -7,3 +7,182 @@ ## 简介 What might happen when we quickly adopt LLMs and make sweeping changes in our complex systems? + +## 正文 + +## Control and complexity: tension in systems design + +The adoption of LLMs in software development has led countless organizations to rapidly change their practices and structures. Old methods are questioned, replaced, and repurposed as the economics around creating new code get shaken up. Because humans and LLMs aren’t interchangeable, the dynamics in play are also very different. Systems are systems, and so regardless of what is changing, there are known patterns on which we can draw to provide some guidance and warnings. + +Without taking a step back and looking at the mindset behind the design of the system in which you operate, you’re likely to get somewhat incoherent (as in “clashing” and “conflicting,” not as in “nonsensical”) measures and policies. And so in this post I want to discuss how we organize systems by contrasting two families of approaches. + +The first is about analytical decomposition that aims to maintain control over a system, and the other is based on a perspective of complex systems that resist analysis, which tend to focus on figuring out interactions and mechanisms to foster desirable emergent behaviour. + +Comparing these has always been useful to tease apart assumptions and important elements of system design, and it is still relevant now with new types of changes being proposed. + +### The approaches + +#### Analytical Decomposition and Control + +At the core of classic science, engineering, and many forms of management, lies the idea that the whole can be understood from its parts. Decompose a complicated thing enough that you can get a thorough and detailed understanding of every component, and you should be able to know how the ensemble works. This approach, analytical decomposition (also sometimes described as “Cartesian-Newtonian”), has been trustworthy and reliable in countless parts of modern life. + +This ability to divide, analyze, and understand generally extends to understanding causality over time: each action has a reaction, each event has a material cause, and these can be traced and evaluated or tested objectively. It follows that we can turn this around: if we understand an object well enough, then we can predict what it will do when acted upon. + +This is foundational to building machines and processes with any sort of predictability and reliability. You can have a high-level goal and a lot of disjoint parts, break down the problem, assemble components that are well tested and within tolerances, and have a working solution. A corollary is that if every part in the machine plays its role well, then the machine itself ought to work well. + +This requires taming a messy, chaotic world, and controlling parameters such that variability can be bounded. Design with enough tolerances and redundancy, and things should work. If not, we can dive in, take it apart, understand what broke, fix it, and be better for it. + +This approach is everywhere, from signal processing and telecommunications, where lossy information transmission is detected and corrected through redundancy, up to industrial quality control, where [statistical processes can be used to define the acceptable boundaries](https://surfingcomplexity.blog/2024/11/23/ttr-the-out-of-control-metric/) of production. + +It also exists at the human level: in human factors engineering, concepts such as working memory (how many things the typical operator can hold in mind) or ideal observers (a theoretical person who monitors instruments at an optimal frequency against which we define “complacency”) have been constructed for the purpose of making sure that systems in which people participate will keep them acting within desirable parameters. + +It’s also visible at organizational levels. Bureaucratic processes and hierarchies aim to keep alignment top-down such that the whole ensemble works coherently. Mechanisms of discipline and legibility are in play to keep the organization’s evolution under control. At broader scales, organizations often try to control their environment, their market, or the legislative context in which they operate. + +Basically, by deciding how much of a mess is accepted on the inside of a process, we can define a clearer interface on the outside of it for others to interact with. This abstraction creates a simplified but effective way to group a complicated ensemble into a manageable unit. + +Software ends up representing a sort of ideal for this mindset: systems can be written in languages that ensure some level of hard-won determinism. Execution is ideally always the same, there is no wear and tear, what worked yesterday will work tomorrow, everywhere. Policy decisions defined far away from the sharp end can be deterministically enforced at all levels. + +This means systems can be built from components bottom-up, aligning with top-down intent, limiting variability that comes from either machining or human behaviour. The ideal is a highly predictable, controlled, competitive, and reactive system. + +#### Complexity and Emergence + +The problem is that by definition, complex systems resist analytical decomposition. + +There are many competing descriptions of complex systems, some of which are behavioural and some of which are structural. They all boil down to something like “things are so interconnected and have so many states that they become either unrepresentable, unpredictable or uncontrollable.” + +Other key elements are that these systems are dynamic, heavily influenced by their own history, and are also open—they continually change and interact in ways that don’t respect clean boundaries. This creates a tension where many participants have distinct goals, perspectives, representations, and degrees of freedom. By the time you’re done analyzing the system, it’s already something else. Even observing the system changes it in important ways. + +Put another way, if you find yourself surprised by the system’s behaviour, by the time you’ve pinned down what happened, it’s already a different system and your policy changes will be lagging or contributing to more counterintuitive surprises. Complex systems are more influenced than controlled. + +This dynamism leads to strategies that encourage equally dynamic adjustments. Since you can’t make these predictable, interventions will often be small and iterative. Alternatively, if you can’t simplify the elements or interactions you’re trying to control, you can increase the variety of control behaviours in order to make ongoing adjustments better. This tends to mean “put a controller—human or otherwise—that has [enough internal complexity to cancel out the complexity](https://pespmc1.vub.ac.be/REQVAR.html) of the thing it controls.” This, in cybernetics speak, is an attempt at creating more adaptive and dynamic control mechanisms. + +Balance is attained not by keeping things static, but by keeping them in motion. + +The ideal system is self-aware and flexible such that it can endlessly adapt and sustain itself, despite ever-increasing challenges. It's unclear whether the ideal can be reached. + +### How they compose (or fail to do so) + +Systems generally evolve from a constrained definition of the problem and its potential solutions, something that is tractable and effective. As the scope and scale of operations grow more comprehensive, further interventions trying to steer the system provide diminishing returns, and they increasingly produce unintended effects. These are the effects of complex systems showing up as things become tangled. + +The coping mechanism I’ve seen the most often is one of doubling down by doing more analysis, more decomposition, and putting more effort into more flexible automation that covers more cases. This in turn changes the nature of success and failure, by creating sometimes less frequent but bigger incidents instead. This type of composition takes place by substituting what breaks when possible, or sometimes by pure accident. It’s rarely been an orderly process. + +More rarely seen mechanisms seek to find out how much of the analytical and control-centric approaches we can afford to give up, identifying what can’t change at all, and then expanding complexity-aware mechanisms outwards from there. This is far less comfortable because this sort of stance demands that you give up on the idea that you actually are in control—a very unpleasant state of affairs to broadcast for a business. + +There are in fact long-standing debates as to whether larger scale accidents can actually be avoided. For example, Jean-Christophe Le Coze offers [the following categorization](http://dx.doi.org/10.1016/j.ssci.2014.03.015): + +![Diagram showing three theoretical explanations for the unpredictability of accidents: technology out of control (Ellul/Perrow), fallible human constructs (Kuhn, Turner, Weick, Vaughan), and self-organizing emergent systems (Ashby, Rasmussen, Snook, Hollnagel).](https://ferd.ca/static/img/lecoze-unpredictability.png) + +1. A ‘deterministic’ thread, where the properties of the technological systems themselves (such as tight coupling and complexity) will eventually defeat efforts to prevent accidents. +2. An ‘epistemic’ branch that focuses on the idea that organizations will suffer from 'failures of foresight' where weak signals and indicators that accidents are incubating will not be seen or accepted by the structures of power, and worldviews will fail to match new challenges, leading to accidents. +3. A ‘self-organizing’ thread that considers systems as adaptive and therefore frames success and failures as consequences emerging from systems' self-organization, through an exploration of problem and solution spaces with their available resources. + +These differing views are not fully incompatible, and authors from one category will frequently borrow from others. Each perspective will however come with a focal point, a thing that is seen as important and worthy of consideration: the structure of control, the historicity of the system, the dynamics of power structures, the adaptive and changing nature of systems, the limited perspectives of participants, concepts around culture, and so on. + +Many contributors to these debates, while stating that accidents are unpredictable or hard to avoid, nevertheless seek explanations that can support making them less likely. They look at the limitations of known approaches, and expand the boundaries of what we should consider, adding new perspectives that can reveal new insights. + +There’s a lot of existing literature across many disciplines to study and get a better grasp on what doesn’t work (and when), and what is contextually useful. The opposition of analytical decomposition for control and complexity for emergence I’m offering here is crude and lacks nuance, but that’s hopefully what makes it an acceptable tool to think about changing systems. + +Oversimplification is what we’re doing here, and knowing what kind of wrong we’re going for is useful. As George Box (1976) said: “Since all models are wrong [we] must be alert to what is importantly wrong.” + +### Contrasting Approaches in Practice + +In a bit of a caricatural manner, the following examples will show relatively stereotypical perspectives to topics relevant to software through both analytical decomposition (with a focus on control) and complexity (with a focus that deliberately limits itself to influence): + +| Topic | Analytical Decomposition / Control | Complexity / Emergence | +|---|---|---| +| Training and education | Build a well-defined curriculum, best practices for teachers and trainers, and testing mechanisms to ensure predictable performance and uniformity across students. | Create environments that foster exploration, experimentation, and information exchange; provide guidance and support. | +| Safety | Prevent undesirable behaviours that lead to failure. Hazards are to be contained or designed out, and deviations from procedures or best practices are seen as a risk. | Foster positive behaviours that lead to success. Find how people bridge gaps in processes, work around obstacles, and recover from problems. | +| Correctness | The software does what the specification or API says it should. Tests pass, it is feature-complete, and operates within known boundaries. | Users or customers are able to successfully accomplish their tasks; goals can shift based on their needs. | +| Reliability | Uptime is within acceptable range, and is verifiable through SLAs, SLOs, etc. Load testing and thorough verification can prevent outages. | Nines don’t matter if customers aren’t happy. You also won’t know for sure if software works until you hit production. Plan for recovery and coping with surprise. | +| Approach to incidents | Runbooks define best practices. Protocols and processes are defined to investigate and triage problems as efficiently as possible. Build for clear information and rapid diagnostics. Investigate what broke so recurrence can be prevented. | Surprises may require improvisation. Who knows what will happen; build capacity to deal with the unknown. Investigations must look into normal work to understand how the system works in the first place. | +| Developing features | Understanding the needs of users and the strengths and gaps in current offerings lets you identify what to build and how to build it. | Experiments in the field with potential features that you iterate on is how you best find what features may prove useful. | +| Standards and norms | Written unambiguously based on verifiable processes and outcomes to make enforcement tractable, scalable, and clear. | Written in a goal-oriented manner as to support and guide the people who execute the work and who need to adapt rules to their reality. | + +For each category, the attitude taken can drive people to pick drastically different approaches and activities, some of which may or may never overlap—the drive to control costs and errors can hinder the effectiveness or desire to experiment, and beliefs about how complex systems work may oppose all sorts of measures that are typically used to demonstrate accountability. + +I say this table is caricatural because in the real world, lines are often not this clean-cut, nor this superficial. It is possible for a control-centric hierarchy to align managers on goals and delegate authority down to cope with system complexity, and for control to be emphasized based on who people in power trust, for example. Centralized control tends to be most effective on the analytical decomposition side, but there are also approaches that aren’t control-centric that benefit from it. + +In fact, many activities can be used in both approaches, and serve both for distinct people, or even at the same time for any given person: + +| Activity | Analytical Decomposition / Control | Complexity / Emergence | +|---|---|---| +| Code Review | Find bugs and flaws; track and assign accountability; ensure quality. | Build awareness and provide a space for feedback within and across teams. | +| SLO adoption | Organizational tool to ensure all teams manage their reliability adequately. | Prioritization tool whose value comes from having teams discuss and define what is an acceptable level of reliability. | +| Refactoring | Pay down technical debt, reduce complexity, improve maintainability and flexibility, normalize used patterns. | Countering entropy, adapting a code base to changing contexts based on new information available or shifting requirements. | +| Chaos Engineering | Validating that expected failure cases are properly tolerated or recovered from | Experimentation-driven exercise in which participants form theories about their system’s behaviour in failure scenarios and try to confirm or disconfirm them. | +| Using a platform | A shared platform can encourage good architectural patterns and prevent undesirable ones, while abstracting away complexity for teams that build on it. | Platforms provide systems with means of commoditizing shared elements to benefit from economies of scale and specialization, and address organizational bottlenecks through self-serve access. | + +Even if activities in this list *can* serve both analytical decomposition and complexity-aligned approaches, that doesn’t mean that they *will*. + +For example, code review approaches that are control-centric and aim to hammer out any deviation from established norms may be adversarial to the point of causing anxiety or hindering actual feedback. Some implementations may still be able to mix automation and the proper social norms to successfully support both purposes to varying degrees of success. + +My experience has been that for these activities, the underlying position taken truly matters if you want to understand how they play out, and how they sometimes fail to meet someone’s expectations. This underlying position will also matter when it comes to prioritizing one activity against others. If participants or stakeholders do not agree to the higher-level purpose and desired outcome, then there will be a gap in ways these activities are expected to be carried out and how they take place, and in the relative importance they will be given across the system. + +When someone wants to change, supplement, or remove some of these activities, it’s useful to wonder what’s the nature of the change and what’s the perspective it favours. + +### Flipping across approaches + +As a heuristic, when multiple lenses are available, we can either try to find the best one (for some arbitrary criteria), or use a complementary or intersecting approach that uses as many of them as possible. Picking a single lens can lead to seeking implementations that maximize one type of activity contextually—whether control or emergence—whereas a combined approach can seek to make sure chosen activities are able to serve multiple properties, as a sort of tradeoff. + +Sometimes, what you get is not what you intend. An organization that sets up activities for control may find itself relying on practitioners invisibly repurposing them for complexity-aligned contributions. Meanwhile, the organization’s decision-makers exercise less control than they believe, or misattribute benefits to their own acts. They can then lose what they had when altering control mechanisms and incidentally hindering the hidden adaptations. + +Conversely, if activities are set up for emergence but are instead done mechanically as if intended for control, they won’t provide the expected benefits and might look and feel like busy work: the organization then neither controls nor benefits from adaptive effects. + +For broad topics and categories such as reliability or correctness, there are often no clearly defined choices or principles that are written down and that you can use. Organizations however tend to have some general tools that line up on the control-to-emergence spectrum, usually around process design and enforcement mechanisms. + +If you’re faced with behaviour you dislike, let’s say people from other teams modifying sensitive code your team owns unannounced, you can take measures such as having discussions with them reasserting ownership, and mentioning the expected process. You could require a preliminary RFC document or ticket before any change request is submitted. You can rely on code ownership files to prevent any unexpected change from going further without your agreement. You can move that key code to repositories which other teams cannot access. + +All of these are relatively local and play on the direct surrounding structure to modify actions and prevent undesirable acts. These approaches may be tremendously effective with little effort, but can also inadvertently fail to make desirable behaviour likelier. + +Closer to emergence’s perspective, it may be more typical to figure out what drives other teams to send these changes unannounced. What are the constraints and pressures they see that makes their current behaviour reasonable to them? If everyone agrees the process is a good ideal state but it frequently gets ignored, what is perceived as more important than that? Only once this is understood should you then design an intervention. This type of questioning—often informed by patterns such as those highlighted previously by Le Coze—tends to have you pull on a thread that unravels through the whole organization. It can be time consuming and difficult to do without established trust, but it can reshape expectations, and as easily lead to major change as to minor interventions upstream. + +A combined approach would be one where a broad understanding of the situation is obtained by leaning on complexity-aware methods, and is then used to design simple but high-leverage checks and barriers such that minimal control yields high rewards. This relies on the complexity stance to look not just at the system’s structure and purposes, but at how its various components and participants interact. Once the interactions make more sense, then the analytical approach is hopefully more effective. + +A risk here is to find yourself with a system that either feels so intractable, resistant to complexity approaches, or inflexible to cross-cutting interventions that you’re back to purely local defences, except they are late, with more work needed to get to the same place. + +The question then is not which approach is better, but how do we know when the current approach reaches its limits and what do we do then? + +### Pitfalls of uncritical system design + +People change their systems all the time, with or without this knowledge. They’re often successful, but not always, or at least not in the ways they had planned. Knowing what to look for doesn’t mean you’ll get it right, but it increases your odds. + +This might be true in the current LLM-driven shakeups as well. Because the technology is new and design patterns aren’t crystallized yet, a lot of people experiment a bit haphazardly. Many of their ideas have interesting elements or aspects to them that are worth learning from, but glaring omissions from a systems perspective that will still need to be handled. + +It’s almost impossible not to find examples of wide sweeping changes proposed when reading tech opinion pieces, which I’ll avoid linking to here. But they include ideas such as: + +- Replacing code reviews with various types of barriers (tests and automated checks), rarely questioning what emergent roles the practice may have nor how static barriers may qualitatively differ from more adaptive ones. +- Splitting software work into high-level specs to be translated to code in a black box with external checks only, without offering explanations around how the specs may cover varying abstraction layers, how the external checks can remain tractable, or how information worth learning should cross these boundaries in each direction. +- Asking for everyone to become a sort of manager-of-agents while keeping agents under tight control loops, without asking what you may lose (or at least cause as second-order effects) in this analogy by changing the delegation and control mechanisms wholesale. +- Focusing on system-level observable outcomes and letting go of imposing the structure within, trusting that the system will self-organize itself adequately. + +If you design a system with control in mind—the use of barriers (think of the Swiss cheese model), the presence of extensive testing, of processes and procedures guaranteeing best practices—then you should pay as much attention to the mechanisms that will be needed to figure out if control actually works. This means asking questions like: + +- How do we know our observations remain relevant, and that we surface the right signals? +- How can we know if our understanding of the system loses accuracy? +- What important elements is our analysis leaving out or obscuring when trying to make things legible? +- How much variability is tolerated, and are we suppressing necessary types of it? +- Are the things we optimize for creating brittleness elsewhere? +- Is our control real or illusory? How would we know if that changes? + +Well-regulated systems compensate for disruptions in ways that hide or suppress the signals of accumulating problems, both at technical and cultural levels. These questions aim to figure out whether any thought is given to what hides such behaviours. + +When you design for emergence—think of self-organization, market-like mechanisms, or delegation of decisions to participants with local context—other questions come up: + +- Are local parts of the system working at cross purposes? +- Is goal alignment effective? What maintains coherence? +- What capabilities or efficiencies are we sacrificing when giving up on legibility? +- Can we afford to lose the efficiency of a control-centric system? When might we need it? +- How do we differentiate adaptation from drift? +- What preserves dissent and carries information from the edges of the system? + +Since complexity-aware approaches tend to resist prescriptive stances, there are often risks of increased inertia or widespread misalignment. Emergent properties will be key to success and failure, but without some careful thinking and influence, things can take on a life of their own. + +Whenever someone pushes for a system design that focuses on analytical decomposition or control, ask how they know they’re doing what’s needed, and the mechanisms by which they adapt. Whenever someone pushes for a design that seems to promise self-regulation and endless flexibility, ask how they’ll maintain coherence and the conditions they rely on for good outcomes. Whenever someone pushes to switch from one to the other, ask what depends on current behaviour and consider what the second-order impacts might be there. + +Tech companies often rush to reinvent themselves around the outsized promises of new technology. Integrating new technology into existing workflows generally demands transforming the workflows. These changes often aim at reducing variability and increasing control, but cross subsystem boundaries in ways that disrupt tangled interactions that were dynamically stable. + +Automation that makes things predictable necessarily removes elements of unpredictability that can be useful to adaptation and evolution. Likewise, trying to make a part of the system more adaptive may necessarily make it less predictable. Both have knock-on effects on the rest of the system. + +Where and how does the system migrate from one mode of operation to the other? Where is control necessary and where is it not? What do we choose to analyze and decompose and what do we treat like an ecosystem instead? + +If we don’t have an answer to these, we also don’t have a good answer to how our systems will avoid failure or meet success. Systems are systems. They will keep acting like systems, and failing like systems. diff --git a/sreweekly/markdown/531/03-structured-logging-in-distributed-systems-what-most-teams-get-wrong-an.md b/sreweekly/markdown/531/03-structured-logging-in-distributed-systems-what-most-teams-get-wrong-an.md index d556a76a..c5009edc 100644 --- a/sreweekly/markdown/531/03-structured-logging-in-distributed-systems-what-most-teams-get-wrong-an.md +++ b/sreweekly/markdown/531/03-structured-logging-in-distributed-systems-what-most-teams-get-wrong-an.md @@ -7,3 +7,137 @@ ## 简介 > Most teams log, but log badly: wrong severity levels, no trace IDs, inconsistent fields, and logs siloed from traces. + +## 正文 + +- + ![](https://dz2cdn1.dzone.com/themes/dz20/images/dz-postarticle.svg) [Post an Article](https://dzone.com/content/article/post.html) +- + [Manage My Drafts](https://dzone.com) + +# Structured Logging in Distributed Systems: What Most Teams Get Wrong and How to Fix It + +Most teams log, but log badly: wrong severity levels, no trace IDs, inconsistent fields, and logs siloed from traces. Fix that, and incidents go from hours to minutes. + +Join the DZone community and get the full member experience. + +[Join For Free](https://dzone.com/static/registration.html) + +Logging is one of the oldest practices in software engineering, yet in distributed systems it remains one of the most poorly implemented. Most teams log, but very few log well. The gap between having logs and having useful logs becomes painfully visible the moment a production incident occurs at 2 AM across a system running dozens of microservices. + +This article focuses on structured logging: what it is, where teams consistently go wrong with it, and the concrete practices that separate log data you can actually act on from log noise that burns engineering hours during incidents. If you are building or operating distributed systems today, structured logging is not optional. It is the foundation on which every other observability signal- traces, metrics, alerts- depends. + +## What Structured Logging Actually Means + +[Structured logging](https://dzone.com/articles/structured-logging-spring-boot-improved-logs) means emitting log entries as machine-readable key-value pairs rather than arbitrary free-text strings. Instead of this: + + Plain Text + + + +`[ERROR] 2026-07-10 03:14:22 - Failed to process payment for user 84729, reason: timeout` +You emit this: + + JSON + + + +``` +{ +  "timestamp": "2026-07-10T03:14:22Z", +  "level": "error", +  "service": "payment-service", +  "event": "payment_processing_failed", +  "user_id": 84729, +  "reason": "timeout", +  "duration_ms": 3001, +  "trace_id": "4bf92f3577b34da6a3ce929d0e0e4736", +  "span_id": "00f067aa0ba902b7" +} +``` +The difference sounds cosmetic. It is not. The first format requires regex parsing and string matching to extract meaning. The second is immediately queryable, aggregatable, and, crucially, correlatable with traces and metrics from other services handling the same request. + +## **The Five Mistakes Distributed Systems Teams Make With Logs** + +### **1. Logging Without Context Propagation** + +In a monolith, a single log line tells you where in the codebase an event occurred. In a distributed system, a log line without a correlation identifier tells you almost nothing. If Service A calls Service B which calls Service C, and Service C fails, you need a shared identifier, typically a trace ID, that threads through all three services' logs so you can reconstruct the full request journey. + +The fix is context propagation: passing a trace ID through every request, injecting it into every log entry, and configuring your logging library to include it automatically. In practice, this means integrating your logging setup with OpenTelemetry or a similar tracing framework from day one, not as an afterthought. When your log entries include trace_id and span_id fields, you can jump from a log entry to its full distributed trace in a single query; that capability compresses incident diagnosis from hours to minutes. + +### **2. Inconsistent Field Naming Across Services** + +In a [microservices architecture](https://dzone.com/articles/what-are-microservices-architecture-and-how-do-the) developed by multiple teams, field-naming inconsistencies compound into a real problem at scale. One service logs user_id, another logs userId, a third logs uid. One service logs errors under error, another uses err, another uses exception. When you need to query across services during an incident, this inconsistency forces per-service query variations, slowing everything down. + +Establish and enforce a logging schema across your organization. Define a canonical set of field names for common concepts, user identifiers, request identifiers, error fields, latency fields, and make that schema part of your service standards. Libraries like structlog in Python or logrus/zap in Go make it straightforward to enforce common fields at the logger initialization level, so teams can't easily deviate from the schema accidentally. + +### **3. Logging at Wrong Severity Levels** + +Severity level misuse is endemic. INFO logs that should be DEBUG. Application errors logged as WARN because the developer did not want to trigger alerts. Business logic exceptions logged as ERROR when they are expected and handled. Over time, this degrades the signal value of severity levels to the point where teams stop filtering by level entirely. + +Adopt and document clear severity semantics for your organization: + +- **DEBUG** : information useful only during active development; should not run in production +- **INFO** : normal operational events (service started, request received, job completed) +- **WARN** : unexpected conditions that are recoverable and do not require immediate action +- **ERROR** : failures that require investigation; every ERROR should eventually be investigated or suppressed with documented justification +- **FATAL** : unrecoverable failures; service cannot continue + +Treat severity levels as a contract with your future on-call self. + +### **4. Over-Logging Hot Paths** + +High-throughput services that log every incoming request at INFO level generate enormous log volumes that create three problems: storage costs escalate, log search performance degrades, and genuinely important events get buried in noise. A service processing 10,000 requests per second generates over 860 million log lines per day from request logging alone. + +Use sampling for high-frequency, low-severity log events. Most observability platforms and [log monitoring tools](https://middleware.io/blog/log-monitoring-tools/) support log sampling natively; you configure a sampling rate for specific log patterns, keeping representative data without keeping everything. For example, sample 1% of successful payment processing logs but keep 100% of error logs. This dramatically reduces volume while preserving signal fidelity where it matters. + +### **5. Treating Logs as a Standalone Signal** + +Logs become exponentially more powerful when they are correlated with traces and metrics. A spike in error logs is interesting. An error log spike correlated with a latency metric increase correlated with a trace showing a database connection timeout is actionable in seconds. Teams that treat logs as independent from their other observability signals are leaving significant diagnostic capability on the table. + +If you are not already running [OpenTelemetry](https://dzone.com/refcardz/getting-started-with-opentelemetry), start there. It provides a unified SDK for instrumenting logs, traces, and metrics in a way that ensures they carry shared context identifiers. Once your logs carry the same trace IDs as your distributed traces, your observability signals become correlated by default, not by manual investigation. + +#### **A Practical Logging Schema to Start With** + +Here is a minimal structured logging schema that covers the majority of production use cases across distributed services: + + JSON + + + +``` +{ +  "timestamp": "ISO-8601 UTC", +  "level": "debug|info|warn|error|fatal", +  "service": "service-name", +  "version": "1.4.2", +  "environment": "production", +  "event": "snake_case_event_name", +  "message": "Human-readable description", +  "trace_id": "OpenTelemetry trace ID", +  "span_id": "OpenTelemetry span ID", +  "user_id": "optional", +  "request_id": "optional", +  "duration_ms": "optional, numeric", +  "error": { +    "type": "TimeoutError", +    "message": "Connection timed out after 3000ms", +    "stack": "optional, omit in high-volume paths" +  } +} +``` +This schema is opinionated but extensible. Services add domain-specific fields as needed while every entry maintains the common fields that make cross-service correlation possible. + +## **Conclusion** + +Structured logging in distributed systems is not about logging more; it is about logging intentionally. The practices that separate teams who resolve incidents in minutes from teams who spend hours in log archaeology come down to four things: consistent field naming, trace context propagation, disciplined severity usage, and treating logs as a correlated signal rather than an isolated one. + +Get these right, and your logs become a first-class observability asset during incidents. Get them wrong, and you have the worst of both worlds: high storage costs and low diagnostic value. The patterns outlined here are not theoretical; they are the difference between incident response that feels like detective work and incident response that feels like reading a timeline. + + systems + Observability + + + Opinions expressed by DZone contributors are their own. + +Comments diff --git a/sreweekly/markdown/531/04-seeing-the-people-in-control.md b/sreweekly/markdown/531/04-seeing-the-people-in-control.md index 219a5cf4..d7222857 100644 --- a/sreweekly/markdown/531/04-seeing-the-people-in-control.md +++ b/sreweekly/markdown/531/04-seeing-the-people-in-control.md @@ -9,3 +9,70 @@ > If you had to explain to a neighbour why your organisation is so safe, and generally works well, what would you say? It’s all about people. I really enjoyed the quote from Charles Billings on principles for automation. + +## 正文 + +*This article is a slightly edited reproduction of the [Editorial](https://skybrary.aero/sites/default/files/bookshelf/hs36/HS36-Shorrock-Editorial.pdf) published in HindSight magazine* [issue 36](https://skybrary.aero/articles/hindsight-31) *(Autumn 2024) (all issues available at [SKYbrary](https://www.skybrary.aero/index.php/HindSight_-_EUROCONTROL))* + +![](https://i0.wp.com/humanisticsystems.com/wp-content/uploads/2025/10/47772149581_37635a39ed_k.jpg?resize=730%2C183&ssl=1) + +[https://flic.kr/p/2fMsKHp](https://flic.kr/p/2fMsKHp) + +I joined the world of aviation in the late 1990s as a Human Factors analyst in UK air traffic management. I had just completed my master’s degree in work design and ergonomics, following my bachelor’s degree in applied psychology. For the first half of my career, my focus was mostly on micro interactions: breaking down tasks, procedures, and interactions at a granular level – seconds and minutes, button presses and radio transmissions. This work involved incident analysis, critical incident interviewing, human-machine interface evaluation, and simulation observation, all aimed at identifying episodes of what we might call ‘loss of control’. Breakdowns and breakages in countless human-human and human-machine loops preceded interactions that sometimes led to losses of separation, level busts and runway incursions. + +Looking back, I was primarily using applied cognitive psychology and cognitive ergonomics to understand control through loops of internal mental processes – perception, memory, attention, and decision-making – along with interactions, and feedback and from the environment. This is often depicted in diagrams with boxes and arrows illustrating the processing of information. + +In the second half of my career, my work shifted toward the macro level, zooming out to interactions within and between organisations, over months, years, and even decades. I listen carefully to people in various roles about their unique experiences. Here, the loops involve communication, cultures, and changes over time. These loops are inseparable and interdependent, creating formidable complexity in terms of people, technology, processes, structures, and organisations. + +No single frame of understanding suffices; I draw upon many disciplines, especially humanistic and social psychology, systems thinking, complexity science, and the humanities, in my attempts to understand the world. From this perspective, people seek to maintain control collectively through loops of communication and influence that evolve document them. + +Looking at the big picture, what is incredible is not that we sometimes lose control, but that we manage to maintain control at all. (Note that there are various meanings of ‘control’, from hard – making something happen – to soft – managing or influencing a process or situation – and it is worth thinking about what it means for you.) This brings me to a question that I often pose to groups, including senior managers: If you had to explain to a neighbour why your organisation is so safe, and generally works well, what would you say? The responses vary, but in the best-connected environments, different groups – controllers, engineers, managers, safety specialists – recognise and acknowledge each other’s contributions, forming large, interconnected loops. It’s a vital question to ponder, because if you don’t, how do you know what to nurture and extend…or defend in the face of cost cuts? + +“Looking at the big picture, what is incredible is not that we sometimes lose control, but that we manage to maintain control at all.” + + +I recently posed this question to an audience of CEOs and safety directors at a EUROCONTROL conference in Spain. It was heartening to hear some senior leaders acknowledge in detail how people are their organisations’ greatest assets. They emphasised that people need to be in control and in the loop. I was surprised at the level of resonance with the theme of this issue of HindSight. + +The CEOs’ comments took my mind back to a groundbreaking report by Charles Billings, [Human-Centered Aviation Automation: Principles and Guidelines](https://ntrs.nasa.gov/api/citations/19960016374/downloads/19960016374.pdf), published in 1996 by NASA. Billings was a former flight surgeon and specialist in aviation medicine, who became an influential and distinguished NASA expert in aviation human factors. The principles in his report remain solid to this day, and the first three are so general that they apply regardless of the presence of automation. + +1. The human operator must be in command. +2. To command effectively, the human operator must be involved. +3. To remain involved, the human operator must be appropriately informed. + +The remaining principles focus on the relationship between human operators and automated systems: + +1. The human operator must be informed about automated systems behaviour. +2. Automated systems must be predictable. +3. Automated systems must also monitor the human operators. +4. Each agent in an intelligent human-machine system must have knowledge of the intent of the other agents. +5. Functions should be automated only if there is a good reason for doing so. +6. Automation should be designed to be simple to train, to learn, and to operate. + +While these principles remain valid, they primarily address the operator-machine dynamic, or ‘joint cognitive system’. This was the focus of my interest in cognitive psychology and cognitive ergonomics. But the humanistic psychologist and systems thinker in me seeks principles that recognise people as more than operators, with control (or influence) distributed throughout organisations, industries, and societies. To this end, I propose the following nine principles to help ‘see’ the people in control: + +1. **People are whole and complex beings.** We are greater than the sum of our mental, emotional, physical or behavioural ‘parts’, and cannot be fully understood by focusing on tasks, functions, roles, or occupations. +2. **People have unique virtues, values, gifts, and passions.** For these to be expressed fully, we need a supportive and nurturing environment that values individuality, diversity, and inclusion. +3. **People have goals, and seek meaning, purpose, and creativity.**  We often seek these things through relationships, work, and personal pursuits. +4. **People naturally strive to learn, grow, and develop.**  We tend to flourish in a supportive and enabling environment. +5. **People are inherently social beings.** We seek meaningful connections with others to find belonging, identity, support, and shared purpose, and are profoundly influenced by social norms, expectations, and pressures. +6. **People’s subjective experience is unique.**  Our experience shapes how we interpret and respond to the world around us and affects our wellbeing. +7. **People live in unique and dynamic contexts.**  These ever-changing contexts – personal, social, organisational, societal, political, environmental, technological, economic, and legal – strongly influence us. +8. **People are part of complex adaptive systems.** Our interactions are influenced by a dynamic network of interactions, which are interconnected and interdependent, with outcomes that are often unpredictable. +9. **People have some choice, control, and responsibility.** But agency is distributed among many and shaped by the opportunities and constraints of the contexts in which we exist, along with our capabilities and motivation. + +These principles remind us that people are more than operators and need to be considered in the broader context. Although these principles have remained valid over millennia, the contexts and the complex adaptive systems in which we live and work (Principles 7 and 8) have changed dramatically, impacting our choices, control, and responsibilities (Principle 9). I encourage you to consider the principles in the light of any activity or change, inside or outside of an organisation. + +“Things work because people make things work, bridging the gaps in the loops as they arise in order to stay in control.” + + +Over the last quarter of a century, one observation has become increasingly clear: everything is connected. In a complex industry like aviation, we can rarely discuss ‘local problems’ in isolation. Even the loss of a single individual – who may possess unique expertise – can significantly impact an organisation. This is equally true for the loss of critical resources. For instance, in our conversation in this issue of *HindSight*, Captain James Burnell discussed the effects of losing crew rooms at some airports. I revisited this impact through the lens of the nine principles I have just outlined. When I recently shared this story with another pilot from a different country, he was horrified at the prospect. *“Crew rooms are sacred!”*, he said, *“There would be riots!”* Crew rooms are shared resources that help crews to stay in the loop and maintain control and have even broader benefits for people. + +Going back to my *“If you had to explain to a neighbour…”* question, my answer is that things work because people make things work, bridging the gaps in the loops as they arise in order to stay in control. We do this using our remarkable expertise, creativity and connectivity, and do this sometimes to our personal cost. What is amazing is that things work as well as they do. It’s time that we fully acknowledged the reason for this – us – and respect people as so much more than operators and overseers of machines and processes. + +### Discover more from Humanistic Systems + +Subscribe to get the latest posts sent to your email. + +This one too !!- well done + +T diff --git a/sreweekly/markdown/531/05-type-conversion.md b/sreweekly/markdown/531/05-type-conversion.md index b47a6a9b..d9cd78ec 100644 --- a/sreweekly/markdown/531/05-type-conversion.md +++ b/sreweekly/markdown/531/05-type-conversion.md @@ -7,3 +7,45 @@ ## 简介 Type conversion in aviation involves an experienced pilot training on a new kind of aircraft. This article draws a parallel to transitioning to a new job as an SRE. + +## 正文 + +# Type Conversion: What Changes (and What Doesn’t) When You Start a New SRE Job + +Pilots have a specific term for moving from one airplane to another: a *“type conversion”*. Not *“learning to fly”* again — you already know how to fly. It’s the process of taking everything you already know and re-mapping it onto a new machine that does the same job with a different cockpit. + +Starting a new SRE job is the same thing. And thinking about it that way has changed how I plan the first few weeks in a new seat. + +## The principles don’t change + +Aerodynamics doesn’t care what airplane you’re in. Lift, drag, thrust, weight — every single-engine piston airplane you’ll ever fly balances the same four forces the same way. A stall is a stall. A stabilized approach is a stabilized approach. The control surfaces that do the work — ailerons, elevator, rudder — are present on every one of them, even when they’re shaped differently or hung in a different place on the fuselage. + +SRE is the same underneath. Error budgets, blast radius, the instinct to widen the time window before trusting a correlation, the discipline of not doing anything you can’t undo quickly — none of that is specific to a company. You bring all of it with you on day one. The job isn’t *“relearn reliability engineering.”* It’s *“figure out where this particular airplane keeps its flap controls.”* + +## The instruments are always there — the panel layout isn’t + +Every airplane you strap into will show you airspeed, altitude, and engine RPM. That’s not optional; you cannot fly safely without them. What changes is where they are on the panel, and maybe whether you’re reading a needle or a glass display. + +Every production system you’ll ever own is going to show you latency, error rate, and saturation, whatever the org calls its version of the USE or RED method. That information exists somewhere, because you can’t run anything at scale without it. What changes is whether it lives in Datadog or Prometheus/Grafana or a homegrown dashboard nobody’s updated the README for. Whether there is an alert on CPU saturation or queue depth, and whether the number you’re staring at is raw or something three layers of aggregation removed from the truth. The first week on a new system is instrument scan practice: find the gauges, confirm what they actually measure, and figure out which ones lie under load. + +## The controls you already know may work differently + +This is the part that trips people up, because it’s where confidence and competence quietly come apart. You know how flaps work. You’ve used them hundreds of times. But if you learned on electric flaps — a switch, a preset, done — and the new airplane hands you a mechanical Johnson bar you have to feel your way through by position, “I know how flaps work” isn’t enough anymore. Retractable gear instead of fixed. A constant-speed prop with a blue lever to manage instead of one knob that just goes faster or slower. Same job, same physics, genuinely different procedure, and the procedure is exactly where accidents might happen. + +New job, same story. You know how deploys work. You’ve shipped code for years. But if you learned on a fully automated ArgoCD pipeline and the new shop is still doing blessed-branch deploys by hand through a Jenkins job somebody’s afraid to touch, “I know how deploys work” is not the same as knowing how deploys work *here*. Kubernetes instead of a fleet of long-lived VMs. A change-management process with an approval board instead of merge-and-go. The underlying skill transfers. The muscle memory for the specific lever in front of you does not, and that’s the gap that gets people in trouble in the first month — not lack of skill, but skill applied on autopilot to the wrong control. + +## The numbers you have to memorize are airframe-specific + +Every airplane has its own set of V-speeds — best glide, maneuvering speed, gear and flap extension limits — and you memorize them cold for *that* airplane. The number for best glide speed in a 172 will get you killed in a Bonanza. Knowing that V-speeds exist and knowing what they are for this specific one are two completely different kinds of knowledge, and only the second one is useful in an emergency. + +Same with paging thresholds, SLO targets, and escalation policy. You already understand that error budgets exist and that burn-rate alerts are how you catch them early. That’s the concept, and it travels. The actual numbers — what latency triggers a page here, what counts as a SEV1, who gets called at 3 a.m. and in what order — are specific to this system, this traffic pattern, this org’s risk tolerance, and you have to learn them cold before you’re the one holding the controls during an incident. Nobody hands you a card with the new numbers on it before your first on-call shift, any more than they hand a pilot a V-speed card mid-flight. You go find it, in the runbooks and the postmortems and the people who’ve flown this one before you, before you need it. + +## How you actually check out on a new type + +Pilots don’t skip the transition training just because they’re experienced. A thousand hours in one airplane buys you good instincts and bad habits in equal measure when you move to a different one, and the honest pilots know it. The process is always the same shape: study the systems (the POH, the equivalent of the wiki nobody’s updated since the last incident), fly with someone who already knows this airplane before you fly it alone, and treat the first hours as information-gathering. + +I’m about to do this again — new company, new stack, new set of V-speeds to learn — and the plan is the same one I’d give anyone else walking into a new seat: don’t assume the panel is laid out the way your last one was, don’t trust a control until you’ve confirmed what it actually does *here*, and spend the first weeks flying the pattern with an instructor in the right seat before you take it up solo. The principles of flight got you your license. They will not, by themselves, tell you where this airplane keeps its flap controls. + +## Where I land + +The comfort in all of this is that the hard part — the part that took years to build — is the part that transfers completely. You are not starting over. You’re doing a type conversion, not primary training. The four forces still balance the same way; the instruments still tell the truth if you know how to read them; and the checklist still doesn’t fly the plane — the pilot who knows when to deviate from it does. That pilot is still you. You’re just learning where the controls and instruments are. diff --git a/sreweekly/markdown/531/06-my-boss-wants-me-to-pick-an-ai-sre-tool-q-a-at-incident-fest-adaptive.md b/sreweekly/markdown/531/06-my-boss-wants-me-to-pick-an-ai-sre-tool-q-a-at-incident-fest-adaptive.md index 2129d250..689e8467 100644 --- a/sreweekly/markdown/531/06-my-boss-wants-me-to-pick-an-ai-sre-tool-q-a-at-incident-fest-adaptive.md +++ b/sreweekly/markdown/531/06-my-boss-wants-me-to-pick-an-ai-sre-tool-q-a-at-incident-fest-adaptive.md @@ -7,3 +7,65 @@ ## 简介 Some big names in this Q&A, and they share a couple of delicious morsels. + +## 正文 + +# 'My Boss Wants Me to Pick an AI SRE Tool': Q&A at Incident Fest (Adaptive Capacity Labs) + +![]() + +### Ready to make incident response your competitive advantage? + +See how Uptime Labs builds provable, scalable incident response capability across your organisation. + +*During* *Incident Fest 2026**, our friends over at Adaptive Capacity Labs, John Allspaw & Beth Adele Long, answered questions about the relationship between AI & humans in our virtual ‘AMA Marquee’. Here are a selection of the questions & answers.* + +## Q: What are the safe and helpful use cases for AI in incident response based on where technology is today? Real-life examples would be very helpful. + +**Beth Adele Long:** + +In terms of safety, I’m a proponent of read-only access during incidents. I was going to say “unless it’s a relatively trivial / low-risk scenario,” but any write access that’s powerful enough to be useful is also likely to be dangerous. And incidents are already confusing enough without having to unwind a bizarre decision that was implemented at AI speed. + +With that safety caveat in mind, in long-running incidents, I can certainly see AI being helpful in the same the way it’s already being used during routine work: as a thought partner to help responders figure out what’s going on and explain current behavior. (J. Paul talked about exactly this [in his talk](https://www.uptimelabs.io/incident-fest-26/main-stage), actually, and his point is ***really*** important. When AI predictions are bad, they degrade performance much more drastically than its good predictions *improve* performance.) + +I would love to see AI tools helping responders better with pattern-matching, but again, how that information is connected and then presented to responders very much matters. I’d also love to see AI supporting incident commanders by helping them make sense of the organization itself — who’s the right SME? Who do we page? Who has been working on the current incident long enough that they’re probably burned out, and I should send them on a humanity break? These are aspects we don’t think about enough but that really matter to effective incident response. + +## Q: My boss wants me to pick an AI SRE tool to introduce into our org. Where do I start? + +**Beth Adele Long:** + +Oh boy, this is a tough one. I would start by getting clear about any contrasts between purported aims and real aims. By which I mean: how much is this a pragmatic request based on specific expectations (“As a leader, I believe AI SRE will help us do X, as measured by Y”) and how much is this actually motivated by something along the lines of “the board / my VP / someone in power says we need to be using AI more, and I need you to make me look good.” The more the latter factor is in play, the less room you’ll have to negotiate based on the actual benefit of the tools. You may just have to pick something and let it play out. + +In either case, I recommend looking at how much an AI SRE tool supports *integration into everyday work*. Are they getting lost in the leftover principle that Stu talked about, promising they’ll do work with no intervention? Or are they *making life easier for your SREs*? The latter claim is easy to test: do a pilot and see what your engineers say. Operational types are notoriously blunt and usually overloaded, so you’ll probably get a fast and honest take whether the tools are helpful or just annoying. + +Finally: good luck. This is a tough time to be evaluating tools that are still very much in flux and figuring out how to provide genuine value. + +## Q: AI has saved a lot of time in the incident review process: sifting through loads of data, nicely constructing the timeline and extracting patterns. It saves a lot of time doing conversations and interviews. I wonder if other folks have seen such time savings. + +**John Allspaw:** + +[The METR study in 2025 on developer productivity](https://arxiv.org/pdf/2507.09089) helped shine some light on how the perception of time savings/spent can be different than the actual amount of time savings/spent. + +(Before testing, developers guessed AI would make them 24% faster. After using it, they *believed* they were 20% faster. Turns out it was actually 19% slower.) + +I’m fascinated by this topic, so I have questions for you as well as others: + +When it sifts through data, what data does it dismiss as unimportant? + +Since all timelines are opinionated (because they’re constructed from raw data in ways that make sense to the author of said timeline), same question: how does the AI choose between events to include and events to dismiss? + +## Q: When using AI in incident response, people frequently say that you have to second guess whether the AI is saying something sensible or not. Isn't that the same with humans? + +**John Allspaw:** + +Evaluating what your colleague has said while you’re both responding to an incident is clearly something happens, yep. Whether or not you’re “second guessing” what they’re asserting depends entirely on your experience with the person in the past, what they’ve said earlier in the response, how they described how they arrived at what they’re saying, etc. + +I’d guess that many people with experience responding to incidents can think of other ways that a human coworker’s contributions might be different than an AI agent’s contributions during an incident…? + +**Beth Adele Long:** + +To expand on John’s remark about “your experience with the person in the past”: with humans we have a deep intuitive sense of *how* to judge someone’s trustworthiness. The engineer who’s been at the company for 5 years and deeply knows this system gets weighted differently than a new hire; the person who’s “often wrong, never in doubt” gets more skepticism than the quiet person who only speaks up when they really know what’s going on. AI usually falls into the “often wrong, never in doubt” bucket! And my experience is that it’s a lot more expensive to evaluate AI’s wall of text in an incident than to guess the trustworthiness of a colleague’s terse assertion. Humans tend to be more efficient at building a shared context as a group (thanks, evolution). So yes, the second guessing is always happening at some level, but *how* that process unfolds is different in important ways. + +*To see the full range of Q&As, you can explore* *Incident Fest here**.* + +![](https://cdn.prod.website-files.com/69e0a463268ba34093f8b1cb/69f225d89c7c20b5a5ec3c6e_Rectangle%2039651.png) diff --git a/sreweekly/markdown/531/07-how-tailscale-helped-find-the-sqlite-wal-reset-bug.md b/sreweekly/markdown/531/07-how-tailscale-helped-find-the-sqlite-wal-reset-bug.md index 958233bb..9c749679 100644 --- a/sreweekly/markdown/531/07-how-tailscale-helped-find-the-sqlite-wal-reset-bug.md +++ b/sreweekly/markdown/531/07-how-tailscale-helped-find-the-sqlite-wal-reset-bug.md @@ -9,3 +9,153 @@ A super-engaging deep-dive. > This investigation is a useful reminder: running boring technology in a non-standard way is a risk. + +## 正文 + +At the end of last year, our uptime was [pretty shaky](https://tailscale.com/blog/hypergrowth-isnt-always-easy). You can see this trend on [our status page](https://status.tailscale.com/history), and that instability continued into the new year. Many of these outages were caused by a single bug, deep in [SQLite](https://sqlite.org/index.html). It took months of intense forensics to track it down. + +Now we’re in summer, we’re confident that we’ve found the bug, that we understand it—and more importantly, that we’ve fixed it. + +We know our customers expect Tailscale to be a reliable service, and for several months we didn’t live up to that promise. That’s disruptive, and we’re sorry. We’re publishing this blog post to explain what went wrong, how we responded, and how we ultimately helped to uncover a long-standing bug in the heart of the SQLite database. + +## [Tailscale’s database architecture](https://tailscale.com#tailscales-database-architecture) + +While our clients interact with our [control plane](https://tailscale.com/docs/concepts/control-data-planes) as a single public endpoint ([controlplane.tailscale.com](http://controlplane.tailscale.com)), internally, our control plane is split into a series of coordination servers (or “shards”). Each tailnet lives on one internal shard at a time, but can migrate seamlessly from one to another. These shards are an internal implementation detail: you don’t know what shard your tailnet is on, and you never need to. + +Each shard has an SQLite database that holds all the information about the tailnets on that shard. A single Go process exclusively accesses that database, and serves the control plane for those tailnets. This single-writer design is exactly how SQLite is meant to be used. + +We’ve used SQLite as our primary database [since 2022](https://tailscale.com/blog/database-for-2022), and we chose it because it's well-known, reliable, and widely used. SQLite is [“boring technology”](https://mcfunley.com/choose-boring-technology)—in a good way. Many companies use SQLite in much larger deployments without issue, and we expected the same stress-free usage. + +In our current backup pipeline, we take a complete snapshot of the database every few minutes, then upload the entire SQLite file to an S3 bucket. We’d been running this setup without incident since early 2023. + +Fast forward to August last year, when a data pipeline that reads those S3 backups reported an error in one of our databases. We ran SQLite’s `PRAGMA integrity_check` command against the backup, and found it was indeed corrupted. SQLite corruption [is possible](https://www.sqlite.org/howtocorrupt.html), but it’s highly unusual and not something you should encounter in normal operation. We repaired the affected database, and investigated the cause, but to no avail. + +When operating at scale, even rare events can occur with some frequency, so we should have been unsurprised when it happened again—and again, and again, and again. In total, we faced 19 separate instances of database corruption over six months before we finally resolved the underlying bug. + +When you hear the phrase “database corruption”, it’s natural to worry about data loss. Because our control plane only handles configuration data, these databases contain metadata about your tailnet and devices, but never your private encryption keys or network traffic. In the earliest incidents, the recovery process meant a handful of newly added devices or configuration changes didn’t persist, and a small amount of metadata had to be re-entered. + +Whenever corruption occurred, we had to stop the control plane process on the shard while we repaired or restored the database. This was painful for tailnets on that shard, because their entire control plane disappeared during that recovery window. In the early incidents, that downtime was over an hour, but we gradually sped up the recovery process over subsequent incidents. + +Each tailnet is a mesh network, where devices make peer-to-peer WireGuard® connections to each other. When a device joins the tailnet, it has to get a list of other devices from the control plane before it can establish new connections—so if a device came online during the SQLite downtime, it couldn’t connect. While the database was being repaired, devices already online remained connected to each other, but they couldn’t learn about changes to the network. Those tailnets also temporarily lost access to the web-based admin console and the Tailscale API. + +There’s also a broader impact on trust. We post a global incident on [our status page](https://status.tailscale.com/) even when only a small number of tailnets are affected. Many people saw a status page event for an incident that didn’t affect them. Indeed, the majority of shards and tailnets were never involved in a database corruption incident! Nonetheless, repeated downtime erodes trust, whether or not you’re directly affected. + +From the very first instance of corruption, we knew this was a serious threat to our reliability, and we threw a lot of engineering time at the problem—but the fix wasn’t easy. + +## [Trying to find the fault](https://tailscale.com#trying-to-find-the-fault) + +This bug resisted all our initial attempts to find it. + +We looked at recent changes, but there weren’t any that seemed relevant. Nobody had been working on our low-level code that interacts with SQLite, because it had all been written years ago and presented no issues up until that point. We re-reviewed all of that code with a fine-toothed comb to look for previously missed bugs, but we didn’t find anything that would cause the corruption we were seeing. + +We looked for common factors between corruption incidents, but we couldn’t find any. It wasn’t tied to a single shard, or customer, or tailnet feature, or time of day, or load level. We were at a loss for what might be triggering the behaviour. + +This lack of reliable trigger conditions meant we couldn’t reproduce the bug synthetically. Instead, we had to rely on deploying passive, forensic telemetry in our live environment to catch the corruption red-handed. Gathering live diagnostics for a database issue is the last thing we wanted to do, but we had no choice. + +As an additional complication, the corruption didn’t occur on a regular schedule. Sometimes incidents would be hours apart, other times weeks. This made it difficult to predict progress or plan further work, because we were never sure when we’d get our next diagnostic dump. We had a six-week period between October and December when there were no corruption incidents, before they returned as an unwelcome Christmas present. + +Because this wouldn’t be a quick or easy fix, we reached out to the SQLite developers for a [professional support contract](https://sqlite.org/prosupport.html). This was a great decision. It gave us direct access to their deep expertise and experience, and we had many detailed technical conversations about our architecture and our incidents. + +Between Tailscale engineering and the SQLite core developers, we mapped out several theories for what might be causing the corruption—including [broken POSIX locks on close()](https://www.sqlite.org/howtocorrupt.html#_posix_advisory_locks_canceled_by_a_separate_thread_doing_close_), mismanaging memory owned by SQLite, or accidentally using SQLite from multiple threads [while disabling thread safety](https://sqlite.org/compile.html#threadsafe). After every incident, we gathered more data, added more diagnostics, and systematically ruled out these theories. We were gradually converging on the true bug. + +## [The transactions that didn’t bark](https://tailscale.com#the-transactions-that-didnt-bark) + +While we were investigating the root cause, we still had a live platform to run. We took aggressive steps to automate recovery and minimize downtime: + +- Configuring our control plane shards to hard-stop immediately upon encountering corruption +- Deploying an automated backup monitor that continuously ran `PRAGMA integrity_check` over our backups +- Improving our runbooks and on-call training + +These efforts cut our response time to under an hour—and then we discovered an unexpected clue. + +We wanted a way to restore service that didn’t involve rolling back to the last known-good backup (which would lose a lot of data) or repairing the known-corrupted database (which was potentially risky). + +To do this, we built a transaction logging pipeline. We streamed every SQL statement that modified the database to a separate log file. Because SQLite is [a single-writer database](https://sqlite.org/faq.html#:~:text=only%20one%20process%20can%20be%20making%20changes%20to%20the%20database%20at%20any%20moment%20in%20time) with [serialisable transactions](https://www.sqlite.org/isolation.html), our transaction history was completely linear and deterministic. (This wouldn’t be true in a multi-writer database like Postgres or MySQL.) Replaying those transactions against the latest known-good backup should restore the database to its most recent state, safely bypassing the corruption. + +This pipeline worked, but then it did something even better: it gave us a clue. + +In two incidents, our transaction logs failed to replay cleanly. Upon closer inspection, we discovered that data written and committed by one transaction was inexplicably invisible to later transactions. A write had vanished into thin air without raising an error. That should be impossible! + +## [The writing on the WAL](https://tailscale.com#the-writing-on-the-wal) + +As these incidents were ongoing, the SQLite developers had been developing a new debugging tool. For a while, we’d suspected that the bug was somewhere in the checkpoint process. They were building a new tool to give better visibility into what was happening during checkpoints. + +To understand what this tool found, we need to briefly explain how SQLite checkpoints work. + +A SQLite database is made of a series of “pages”, tiny blocks of information. When you update the database, some of those pages need to be replaced with new pages with the updated information. + +For better performance and greater concurrency, we run SQLite with Write-Ahead Logging, which means new pages aren't written directly to the database file. Instead, they’re written to the "write-ahead log" or "WAL file". + +New pages can't be written to the WAL file indefinitely; at some point they have to be copied back to the main database file. This process is called “checkpointing”. + +In most deployments, SQLite itself decides when to do a checkpoint, and the process is invisible to the end user and developer. In our control plane, we take manual control of the checkpoint process so we can run fast and consistent backups. This non-standard approach seemed suspicious as we steadily eliminated potential causes. + +One clue was that during corruption incidents, our metrics showed that SQLite would report copying more pages from the WAL file than were actually available. If there are 10 pages in the WAL file and 20 pages get copied to the database, something is clearly wrong. + +To understand what was happening during these faulty checkpoints, the SQLite developers created a new debugging tool for the virtual filesystem layer. + +SQLite is split into several layers. The top layer is the parser and code generator, which converts SQL statements into SQLite’s internal data structures. These data structures get passed to the pager, which splits them into the individual pages to be written to disk. Actually writing them to disk is handled by the OS interface, or “virtual filesystem”. Currently SQLite has two mainstream virtual filesystem implementations—Unix and Windows. + +If you're interested in a deeper dive on these internals, I recommend [this lecture by Richard Hipp](https://www.youtube.com/watch?v=ZSKLA81tBis), the primary author of SQLite. + +This approach allows you to replace different layers with different implementations, or wrap an existing layer to get more information. To help diagnose our problem, the SQLite developers created a wrapper around the virtual filesystem that writes additional tracing information and logs about changes to the database. This wrapper is called the `tmstmpvfs` shim, and the source code is available in the [SQLite public repository](https://sqlite.org/src/file/ext/misc/tmstmpvfs.c?proof=394548354). + +We deployed the shim into our live environment, and waited for the next corruption to occur. Fortunately, we didn't have to wait long. + +## [The WAL-Reset bug](https://tailscale.com#the-wal-reset-bug) + +After our next corruption incident, the additional logs from the new `tmstmpvfs` shim allowed the SQLite developers to find and fix the bug: a rare data race in the SQLite source code between a checkpoint and a write transaction. + +In particular, if a write occurs at a specific time during a checkpoint, the checkpointing process gets confused—it thinks some of the pages have been copied from the WAL into the main database file, but they haven’t. Those pages never get written to the database file, and that data is permanently lost. The database file becomes corrupt, because other pages which reference those pages—such as an index—are written to the database. + +The SQLite developers named this the [“WAL-Reset bug”](https://sqlite.org/wal.html#the_wal_reset_bug), and they estimate it was present in SQLite for at least 16 years. It could exist that long because it was rare—so rare, the SQLite developers had to add code to deliberately trigger it in their testing environments. Their fix adds [an additional check to the checkpointing function](https://sqlite.org/src/info/7168988acbec2d8d) which detects when the WAL has been reset by another thread. + +They confirmed that this bug caused all of the baffling behaviour we’d seen. It explained the corruption, the transaction logs that wouldn’t apply cleanly, and the inconsistent checkpoint statistics. They also explained why we were more likely to hit the bug than other SQLite users: we take manual control of the checkpointing process, and we checkpoint very aggressively. Even a bug triggered by a rare condition was bound to hit us eventually. + +This was an exciting moment. After months of confusion and uncertainty, we finally had a plausible theory for why the corruption was occurring, and a fix we could deploy to prevent it. + +The SQLite developers released the fix as [SQLite 3.52.0](https://sqlite.org/changes.html#version_3_52_0), and we prepared to deploy it as soon as it was available. + +## [Fixed, with a false alarm](https://tailscale.com#fixed-with-a-false-alarm) + +We rolled out SQLite 3.52.0 carefully—first to a few canary shards, then, when we saw it running smoothly, we deployed it to the rest of the control plane. + +Our backup monitor promptly turned red, and reported corruption in **13 different databases**. This was extremely alarming, but we followed our recovery procedures to fix all the supposed corruption, and everything was happy. It turned out these databases had not suffered real corruption, but were subject to a second problem in the version of SQLite. + +We shared our errors with the SQLite developers, which uncovered a bug in SQLite related to [stale expression indexes](https://sqlite.org/staleexpridx.html). If you create an index on a computed value, and then the computation changes, the index will contain mismatched values, which gets reported as corruption by `PRAGMA integrity_check`. + +In our case, we were storing some high-precision timestamps as text, converting them to a floating-point number in a VIRTUAL generated column, and the SQLite 3.52.0 release that fixed our data race also made [an optimisation](https://research.swtch.com/fp) that subtly changed the rounding behaviour for text-to-floating-point conversions. Our canary shards didn’t have any timestamps that triggered the changed rounding behaviour, so we missed this in our phased rollout. + +Because this change caused false corruption warnings, the SQLite developers withdrew the 3.52.0 release and instead published 3.51.3, which only contained a fix for the WAL-Reset bug. + +We fixed the issue on our side by reducing the precision of our timestamps to integer seconds; text-to-integer conversions are unambiguous. Meanwhile, the SQLite developers created an automated, [self-healing index feature](https://sqlite.org/staleexpridx.html#selfheal) in 3.53.0, which prevents the stale expression index problem. + +## [Party time!](https://tailscale.com#party-time) + +With the fix rolled out to our entire control plane, we were ready to declare victory, but we were still cautious. An absence of corruption incidents doesn’t mean things are fixed—we’d already had one six-week period of deceptive calm. + +We wanted positive proof that this data race was actively occurring in our production environment. Now that we understood the cause of the bug—a collision between a write transaction and a WAL-reset—we patched our SQLite driver to [log a warning](https://github.com/tailscale/sqlite/commit/070c921dea03f793cbbbd9d39551d0ad26136113#diff-fcff16fa73210758d3a31f4cfa4609ad5b3ca692b66fc7fb7587c08a1ab11048) when these two operations overlap. If the warning fired but the database remained uncorrupted, we’d know the fix had saved us from a potential corruption incident. + +We deployed the warning, and we waited. And we waited. And waited. And waited. As weeks slipped by, we began to wonder why we didn’t see it. Was the warning broken? Was our theory wrong? Was the true bug still lurking in the darkness? + +Then, two months later, the alert we were waiting for finally fired: + +![Alert Manager notification showing SQLitePartyMode warning: SQLite attempted corruption on shard2.corp.ts.net:8383 in party mode, but the system prevented it. Details include warning code, host, instance, job, namespace, severity, and shard information. Message advises checking server logs for corruption incident details.](https://cdn.sanity.io/images/w77i7m8x/production/19ae089ae209ad7e395e188b7f6e4a81be29c431-1458x684.png?w=3840&q=75&fit=clip&auto=format) + +This alert proved that the precise conditions for the WAL-Reset bug *do* occur in our production environment, which means it was the likely culprit for our six months of shaky uptime. + +Since that weirdly joyous alert fired, we’ve run for another four months without any database incidents, as of this writing. Finally, we could breathe a sigh of relief. + +## [Off the well-trodden path](https://tailscale.com#off-the-well-trodden-path) + +Nobody wanted us to spend six months looking for bugs in SQLite. This was an immensely frustrating experience for both our customers and staff, and we’re all glad to put this instability behind us. + +This investigation is a useful reminder: **running boring technology in a non-standard way is a risk.** The common paths and standard configurations are incredibly well-tested and reliable. Most people use SQLite in a standard configuration and never face this sort of issue. Everything we were doing was a public, documented, supported configuration—but by taking manual control of the checkpointing process and running at our own aggressive pace, we stepped off the well-trodden operational path. + +Resolving these incidents was a massive, cross-functional effort involving dozens of people—including Tailscale's engineering and support teams, and the core maintainers of SQLite. It is to all of their credit that the impact of these incidents was not much worse. + +We know that repeated downtime erodes trust, no matter how many people are affected, and we’re grateful to our customers for their patience and support while we chased this down. + +Frustrating as this period was, we’re left in a stronger position than we were before. The long-standing bug in SQLite has been patched, and we fixed dozens of other incidental issues that we spotted while looking for it. We funded the open-source SQLite VFS shim that helped isolate the race condition almost immediately, and will help track down similar bugs in the future. Finally, we’ve refined our database backup and recovery processes, and live-tested them over a dozen times. + +Hopefully there won’t be another database incident like this—but if there is, we’ll be ready. diff --git a/sreweekly/markdown/531/08-optimizing-kubernetes-pods-for-reliability-with-topology-spread-constr.md b/sreweekly/markdown/531/08-optimizing-kubernetes-pods-for-reliability-with-topology-spread-constr.md index b49cafe3..37f61a2b 100644 --- a/sreweekly/markdown/531/08-optimizing-kubernetes-pods-for-reliability-with-topology-spread-constr.md +++ b/sreweekly/markdown/531/08-optimizing-kubernetes-pods-for-reliability-with-topology-spread-constr.md @@ -7,3 +7,134 @@ ## 简介 A handy guide on topology constraints in Kubernetes, with a worked example. + +## 正文 + +![](https://cdn.prod.website-files.com/64a52fe92f1fc7debc3007b0/6a7dd87b9fbd815435905ea9_Blog%20Headers.png) + +# Optimizing Kubernetes pod deployments for reliability with topology spread constraints + +If you’re like many Kubernetes users, you don’t pay much attention to where or how Kubernetes distributes your pods. As long as they’re running, it doesn’t matter where they get deployed, right? Surely Kubernetes will use some complex algorithm to figure out the most reliable way to distribute your pods across the cluster…right? + +Pod distribution plays a much bigger role in reliability than you might think. Fortunately, it’s easy to control when, where, and how Kubernetes distributes pods. By adding a few lines to your manifest, you can ensure your deployments are zone-redundant and evenly scalable. The feature is called *topology spread constraints*, and in this blog, we’ll explain how it works in full detail. + + +## Why are topology spread constraints important for reliability? + +[Topology spread constraints](https://kubernetes.io/docs/concepts/scheduling-eviction/topology-spread-constraints/) determine how Kubernetes distributes pods across failure domains, such as regions, zones, and nodes. This helps ensure your workloads are truly distributed not just across the cluster, but across your operating environment. You can set cluster-level constraints as a default, or set constraints for individual workloads. + + +### How to configure topology spread constraints + +Topology spread constraints are defined using the field spec.topologySpreadConstraints. These can be applied to a pod or to the cluster. Constraints have the following fields: + +- `maxSkew` : the degree to which pods may be unevenly distributed. Its behavior depends on the value of`whenUnsatisfiable` : + - If `whenUnsatisfiable: DoNotSchedule` , this determines the maximum difference between the minimum number of pods in the domain vs. the number of matching pods in the target topology. In other words, this is how far off the minimum a domain is allowed to get. + - If `whenUnsatisfiable: ScheduleAnyway` , Kubernetes gives a higher precedence to topologies that would help reduce the skew. +- If +- `minDomains` : the minimum number of eligible domains (e.g. availability zones or regions). +- `topologyKey` : the node label used to identify nodes used for this constraint. Any nodes that have this label are grouped into topology domains according to their values. For example, using`topology.kubernetes.io/zone` as a key creates domains based on the availability zones your hosts span. +- `whenUnsatisfiable` : how to handle pods that don’t satisfy the spread constraint. By default, it won’t be scheduled (`DoNotSchedule` ). Setting this to`ScheduleAnyway` schedules the pod regardless, prioritizing nodes that minimize the`maxSkew` . +- `labelSelector` : the pod label used to find matching pods. +- `matchLabelKeys` : a list of pod label keys to use to calculate the spreading skew. +- `nodeAffinityPolicy` : determines how to treat each pod’s`nodeAffinity` and`nodeSelector` settings.`Honor` (the default) limits the topology calculation to these nodes, while`Ignore` uses all nodes. +- `nodeTaintsPolicy` : determines whether to include node taints in the topology calculation. + +Note that you can define only one `topologySpreadConstraint` for a given `topologyKey` and `whenUnsatisfiable` pair. + + +## How to add a topology spread constraint to a Kubernetes manifest + +Imagine we have a Kubernetes cluster distributed across three availability zones: us-east-1a, us-east-1b, and us-west-2a. We also have a pod that we want to deploy and replicate for redundancy. We’ll start with the following manifest: + + +If we deploy four replicas of the pod using a round-robin algorithm, we end up with one node with two pods and two nodes with one pod: + + +![](https://cdn.prod.website-files.com/64a52fe92f1fc7debc3007b0/6a7b72f84de133f41d6b4e35_37930b7b.png) + + +However, the Kubernetes scheduler might deploy two pods to two nodes, leaving one empty; or it might deploy three pods to us-east-1a and one to us-east-1b, which puts us at risk if the us-east region ever goes down. Or, in the worst case, it could deploy all four to one node and create a single point of failure. + + +![](https://cdn.prod.website-files.com/64a52fe92f1fc7debc3007b0/6a7b72f84de133f41d6b4e38_2d98338a.png) + + +1. Let’s first limit the pod imbalance by setting `maxSkew` to 1. This ensures that no single node has more than one additional replica of the pod than any other node. +2. Next, we’ll set the `topologyKey` to`topology.kubernetes.io/zone` , since we want to limit the spread by zone even if our zones span multiple regions. +3. We want Kubernetes to run the pod even if it can’t satisfy our topology constraints, so we’ll set `whenUnsatisfiable` to`ScheduleAnyway` . +4. We want to match all Nginx pods (in this deployment, anyway), so let’s add a `labelSelector` that matches`app: nginx` . + +Now, our manifest looks like this: + + +## How to find pods with missing topology spread constraints + +You can use the `kubectl` command-line tool to retrieve a list of pods, then use the `jq` command-line tool to filter pods that don’t have `topologySpreadConstraints` defined. For example: + + +You can also use Gremlin’s built-in [Detected Risks](https://www.gremlin.com/docs/reliability-management-detected-risks#toc-topology-spread-constraints-absent) feature to automatically scan your Kubernetes pods for missing topology spread constraints. + +Once you’ve added your constraints, re-run this command to ensure your pods don’t appear in the output. If you’re using Gremlin, the “[topology spread constraints absent](https://www.gremlin.com/docs/reliability-management-detected-risks#toc-topology-spread-constraints-absent)” risk status will automatically change from “at-risk” to “mitigated” and your service’s reliability score will increase. + + +## Combining topology spread constraints, node affinity rules, and taints and tolerations + +As we already saw, topology spread constraints can interact with other Kubernetes features, particularly [node affinity rules](https://kubernetes.io/docs/concepts/scheduling-eviction/assign-pod-node/#affinity-and-anti-affinity) and [taints and tolerations](https://kubernetes.io/docs/concepts/scheduling-eviction/taint-and-toleration/). But there are subtle differences between these. + +Affinity rules define the specific criteria for scheduling a pod on a node. For example, a pod running a large language model (LLM) might have an affinity rule that requires a node with a GPU. This way, you can combine affinity rules and topology spread constraints to limit the domain of nodes available to a pod. Just make sure you set `nodeAffinityPolicy: Honor` (the default). + +Conversely, taints specify where *not* to schedule a pod unless it has a matching toleration. If a node’s GPU crashes due to a driver issue, you don’t want Kubernetes scheduling LLMs onto that node. Instead, you can apply a taint that prevents Kubernetes from scheduling pods on that node, while also migrating running pods onto new nodes. Like affinity rules, these work in tandem with topology spread constraints by limiting the size of the domain, as long as you set `nodeTaintsPolicy: Honor`. + + +## Other Kubernetes risks to watch out for + +Topology spread constraints are just one piece of a resilient Kubernetes deployment. If you want to know how to protect yourself against other risks like missing liveness probes, unset resource requests, and improperly configured high-availability clusters, check out our comprehensive ebook, "Kubernetes Reliability at Scale." + +In the meantime, if you'd like a free report of your reliability risks in just a few minutes, you can sign up for a free 30-day Gremlin trial, or use Gremlin's [Detected Risks](https://www.gremlin.com/docs/reliability-management-detected-risks#toc-topology-spread-constraints-absent) feature to automatically scan your existing Kubernetes pods for missing topology spread constraints. + + +**Start your free trial** + +Gremlin's automated reliability platform empowers you to find and fix availability risks before they impact your users. Start finding hidden risks in your systems with a free 30 day trial. + +[sTART YOUR TRIAL](https://www.gremlin.com/trial) + +To learn more about Kubernetes failure modes and how to prevent them at scale, download a copy of our comprehensive ebook + +[Get the Ultimate Guide](https://www.gremlin.com/whitepapers/kubernetes-reliability-at-scale-how-to-improve-uptime-with-resiliency-management?utm_source=cta&utm_medium=webpage&utm_campaign=on_page_cta_kub_rel_at_scale+) + +[Back to top](https://www.gremlin.com#single-article) + + +## How to troubleshoot unschedulable Pods in Kubernetes + +Kubernetes is built to scale, and with managed Kubernetes services, you can deploy a Pod without having to worry... + +![How to troubleshoot unschedulable Pods in Kubernetes]() + +Kubernetes is built to scale, and with managed Kubernetes services, you can deploy a Pod without having to worry... + +[Read more](https://www.gremlin.com/blog/how-to-fix-kubernetes-unschedulable-pods) + + +## How to ensure consistent Kubernetes container versions + +One of Kubernetes' killer features is its ability to seamlessly update applications no matter how large your deployment is. Did a developer make a code change, and now you need to update a thousand running containers? Just run kubectl apply -f manifest.yaml and watch as Kubernetes replaces each outdated pod with the new version. + +![How to ensure consistent Kubernetes container versions](https://cdn.prod.website-files.com/64a52fe92f1fc7debc3007b0/6576e67dd94958bc633ef1fb_default-blog-hero.webp) + +One of Kubernetes' killer features is its ability to seamlessly update applications no matter how large your deployment is. Did a developer make a code change, and now you need to update a thousand running containers? Just run kubectl apply -f manifest.yaml and watch as Kubernetes replaces each outdated pod with the new version. + +[Read more](https://www.gremlin.com/blog/kubernetes-container-image-version-uniformity) + + +## Managing slow container starts with Kubernetes readiness probes + +Pods without readiness probes are like engineers without coffee. Learn how readiness probes work, why they’re important, and how to configure them correctly. + +![Managing slow container starts with Kubernetes readiness probes](https://cdn.prod.website-files.com/64a52fe92f1fc7debc3007b0/6a70debb40f15c68d7190eb0_Blog%20Headers.png) + +Pods without readiness probes are like engineers without coffee. Learn how readiness probes work, why they’re important, and how to configure them correctly. + +[Read more](https://www.gremlin.com/blog/managing-slow-container-starts-kubernetes-readiness-probes) diff --git a/sreweekly/markdown/532/01-when-declaring-an-incident-becomes-everyone-s-favorite-workaround.md b/sreweekly/markdown/532/01-when-declaring-an-incident-becomes-everyone-s-favorite-workaround.md index 57d2c959..df0c2c3d 100644 --- a/sreweekly/markdown/532/01-when-declaring-an-incident-becomes-everyone-s-favorite-workaround.md +++ b/sreweekly/markdown/532/01-when-declaring-an-incident-becomes-everyone-s-favorite-workaround.md @@ -7,3 +7,47 @@ ## 简介 Need another team to do something fast? Just use this one weird trick: declare an incident! This article explains why the obvious solution (gating incident declaration) isn’t a good idea. + +## 正文 + +You see someone declare a Sev-2 and you wonder: wait, why is that even an incident? Nothing is down. Customers aren’t affected. But a manager needed to get their team’s problem to the top of another team’s priority queue, and the incident process was a reliable way to make it happen. That’s not really what the incident process is for, but it worked, so where’s the harm? + +The problem is, once folks see that this works, it starts happening more often. A product manager declares an incident because the incident notification is the fastest way to get leadership attention on a problem that’s been stuck in the backlog for weeks. An account team declares one because they need engineering support for a big demo to a major prospect and the incident process is the easiest way to pull engineers out of their sprint work on short notice. An engineer declares one because it’s easier than navigating the formal exception process for the deployment freeze. + +The harm is cumulative. When a growing fraction of your declared “incidents” aren’t real emergencies, the urgency signal degrades. When a genuine Sev-1 arrives, people respond with less urgency because they’ve been conditioned to expect another workaround. And the incentives compound: folks who game the system get their problems solved faster, which teaches everyone else that gaming is how to get things done. Each individual declaration is an understandable decision by someone who needs to get something done; it’s the aggregate that corrodes the process. + +Every one of these non-emergency declarations still carries the full overhead of a real incident. Responders get pulled off their planned work. Someone drops whatever else they were doing to serve as incident commander. Stakeholders context-switch to follow along. When you’re running enough of these, your teams are spending a meaningful fraction of their time in emergency mode for things that aren’t really emergencies, and all the indirect costs of incidents (disrupted projects, context-switching, recovery time) accumulate just the same. + +There’s an irony here: people are reaching for the incident process *because* it works; they’ve seen that it reliably delivers coordination, prioritization, and urgency on demand. + +## The instinctive response is wrong + +When companies notice this pattern, the instinctive response is often to tighten the declaration criteria. They add gatekeeping: maybe you need manager approval to declare an incident, or there’s a pre-declaration checklist you have to complete first, or someone reviews whether the declaration was “warranted” after the fact. The intent is reasonable. The net effect is corrosive. + +Gatekeeping incident declarations is counterproductive. Every speedbump you build also slows down real incidents. The person who hesitates to declare because they’re not sure the problem is “bad enough” is already a common failure mode in incident response. Adding a formal approval step or a post-hoc review of whether the declaration was justified makes that hesitation worse, not better. + +You also miss what the gaming is telling you: people reaching for the incident process are telling you that your normal processes are falling short. If you only crack down on the gaming, you suppress the symptom without learning anything from it, and the underlying problems persist. + +## Fix the escape routes, not the escaping + +Instead, look at what side effects people are trying to trigger when they declare questionable incidents, and make those capabilities available through other means. + +If the easiest way to bypass the deployment freeze is to declare an incident, create a non-incident exception process for urgent changes. This doesn’t have to be complicated; a lightweight approval from a designated release manager, with a clear escalation path, covers most cases. + +If the easiest way to get your problem moved up another team’s priority queue is to declare an incident, create a prioritization escalation path that doesn’t require an incident. A cross-team triage meeting, an explicit expedite-request mechanism, or even a dedicated Slack channel that the right people actually monitor can absorb most of the pressure. The bar doesn’t have to be as high as “declare an emergency”; it just has to be lower than “wait six weeks for the next planning cycle.” + +If the easiest way to assemble a cross-functional team on short notice is through the incident process, create a lightweight coordination mechanism for non-incident situations. Some companies call these “swarms” or “tiger teams” or “coordination requests.” The name doesn’t matter; what matters is that people have a way to get the collaboration they need without borrowing the incident process to do it. + +Repeatedly gaming the incident process to get resource prioritization or cross-functional coordination isn’t a series of one-off workarounds; it’s a symptom of a systemic problem that needs a systemic response. Google’s SRE organization built formal [Code Yellow and Code Red](https://www.theengineeringmanager.com/growth/code-yellow-code-red/) mechanisms for exactly this: structured ways to rally resources and elevate priority when a problem is serious enough to demand cross-functional attention, but isn’t an incident. + +## The diagnostic question + +Look at your last dozen or so incidents and ask, for each one: was this declared because there was an emergency, or because the incident process was the easier path to something the team needed? + +You don’t need a formal audit. Just ask a few experienced incident commanders and on-call engineers; they already know which ones were real and which ones weren’t. Then talk to the folks who called for the questionable ones (in a blameless, fact-finding way, of course). They’ll tell you exactly what’s missing from the normal processes, if you’re willing to listen. + +People gaming the incident process is just a symptom. The underlying problem is usually that normal processes are too rigid, too slow, or too unresponsive, and the incident process is the path of least resistance. Fix the underlying problem and the gaming stops, because there’s nothing left to game around. Your incident urgency signal recovers, your teams stop burning emergency-mode cycles on non-emergencies, and when a real Sev-1 hits, people respond like it matters. + +And if you’re dealing with this, take a moment to appreciate what it says about your incident process: people are borrowing it because it *works*. The fix isn’t to make it stop working. It’s to make everything else work that well too. + +## Recent Comments diff --git a/sreweekly/markdown/532/02-unethical-ways-to-manage-technical-debt.md b/sreweekly/markdown/532/02-unethical-ways-to-manage-technical-debt.md index 2a5d3a23..a55be61a 100644 --- a/sreweekly/markdown/532/02-unethical-ways-to-manage-technical-debt.md +++ b/sreweekly/markdown/532/02-unethical-ways-to-manage-technical-debt.md @@ -7,3 +7,7 @@ ## 简介 Ethics are relative, right? This article is full of genuinely useful tips and framings. + +## 正文 + +> ⚠️ 抓取失败:HTTP 403 diff --git a/sreweekly/markdown/532/03-solving-mysterious-kubernetes-pod-setup-timeouts-by-tuning-conntrack-g.md b/sreweekly/markdown/532/03-solving-mysterious-kubernetes-pod-setup-timeouts-by-tuning-conntrack-g.md index 15c331d5..316220f0 100644 --- a/sreweekly/markdown/532/03-solving-mysterious-kubernetes-pod-setup-timeouts-by-tuning-conntrack-g.md +++ b/sreweekly/markdown/532/03-solving-mysterious-kubernetes-pod-setup-timeouts-by-tuning-conntrack-g.md @@ -7,3 +7,291 @@ ## 简介 Whoa. It’s been quite a few years since my last run-in with an overfull conntrack table, and this is a fun new twist. + +## 正文 + +[Knowledge Hub](https://www.adyen.com/knowledge-hub) + +Article + +# Inside Cilium CNI: solving mysterious Kubernetes pod setup timeouts + +A deep dive into how Adyen's Data Platform Engineering team investigated and resolved linear-time scaling bottlenecks in Cilium CNI connection tracking garbage collection to fix mysterious Kubernetes pod setup timeouts on high-resource nodes. + +I was fully aware a year ago that a single configuration line could break the Kubernetes networking stack. But if they told me that leftovers from Kubernetes pods which terminated hours prior could block new ones from starting, I would have thought they were joking. + +In high-performance networking, 35 seconds is a lifetime. This was the latency required to iterate through our connection tracking table of 7 million entries at a maximum speed of 200,000 entries per second. At our 16-million-entry peak, this sequential lookup could take up to 80 seconds, leading to Cilium CNI timeouts preventing new pods from starting on affected nodes. + +We uncovered this linear-time behavior at Adyen by tracing syscalls, inspecting codebases, and analyzing eBPF internals. This investigation revealed how our varied workloads turned the connection tracking table's garbage collection algorithm into a critical bottleneck. + +## Our setup: why we're different + +At Adyen, we run Cilium CNI across all our 100+ Kubernetes clusters. When we switched from Calico to Cilium, we knew we'd face challenges adapting it to our production workloads. Our production big data Kubernetes clusters have a unique usage pattern compared to the other Kubernetes environments within Adyen: + +**Data extraction from HDFS**. Our infrastructure relies on more than 500 datanodes. Trino represents one of our most demanding HDFS workloads, processing analytical queries against data stored on HDFS. Due to the distributed nature of HDFS, each file you download requires a new connection to any of these 500 nodes. Therefore, during peak hours, a single pod can produce approximately 50,000 connections every minute. + +**Pod churn.** Many pods we spawn on the Kubernetes cluster run batch jobs, such as Spark jobs. They stay around for anywhere from a second to a couple of hours. + +**Wide variety of workloads.** Some workloads are very CPU-intensive, like Spark pods executing complex joins and transformations with relatively few network connections. Others are extremely network-intensive, like Trino pods querying thousands of small files on HDFS, each requiring a new connection. This creates a large number of short-lived connections that stress the connection tracking table. + +Furthermore, our machines are more powerful than most Kubernetes cluster machines within Adyen: + +- Machines with 64 physical cores and 512GB of RAM +- Machines with 128 physical cores and 2TB of RAM + +At the time, our environment consisted of Kubernetes 1.31.7 and Cilium 1.16.5. We provide these specific versions to help interested readers correlate our findings with the relevant codebases. + +To support our unique workloads, we progressively tuned Cilium by increasing DNS proxy timeouts, expanding the connection tracking table capacity, raising API rate limits, and streamlining "security labels". This configuration allowed us to overcome challenge after challenge, except for one: + +**`Failed to create pod sandbox: rpc error: code = Unknown desc = failed to setup network for sandbox "0fecf4844d3f8f2df218f09d91f9698bb424e2166551952f46fd7638f4757cf2": plugin type="cilium-cni" failed (add): unable to create endpoint: Cilium API client timeout exceeded`** + +Kubernetes would create a pod, spin up the sandbox, and then attempt to set up the networking with Cilium. This operation would time out. From that point onwards, it also timed out any other pod attempting to spawn on the same node. Furthermore, we noticed that on those affected nodes, the API duration for DELETE /v1/endpoint and PUT /v1/endpoint, the routes responsible for adding and deleting a Cilium endpoint, also increased. You can see this clearly in the image below where the endpoint call latency starts to increase linearly over time from 4:20 pm onwards, indicating that endpoint creation and deletion never finish. + +![Graph showing payment processing times on Adyen's platform from 9:50 to 17:00.](https://media.ffycdn.net/eu/adyen/gtUYQ1kDX8Tbne4WqAJj.png?format=webp&width=654) + +It is important to note that these clusters are only used for analytical processing and big data workloads. The issue was identified, analyzed, and fully resolved before any SLO was breached. Furthermore, as our big data environments are isolated from our core transactional flows, this issue never impacted our real-time payment processing pipeline or merchant transactions. + +## The symptom: something's wrong, but what? + +First, we found no related timeouts in the Cilium Agent logs. We enabled debug logging, hoping to find hidden errors. Nothing. The logs gave us the big picture but didn't reveal the bottleneck. + +We did learn some valuable things from debugging the broken nodes: + +- Listing all endpoints on a broken node would time out: cilium endpoint list +- Listing the logs of an endpoint worked fine and showed the endpoint was stuck in the regenerating phase: cilium endpoint log +- Requesting the endpoint state with kubectl get ciliumendpoint showed it was still regenerating, while the Kubernetes manifest of the endpoint showed a ready status. +- Checking the filesystem at /var/run/cilium/state showed endpoint folders with a _next suffix, indicating they didn't complete successfully. + +This validated our suspicion that the issue was in the Cilium agent, not at the kubelet or CNI plugin level. But we still didn't know why. + +## Under the hood: how Cilium sets up networking for a pod + +Before we dive into the debugging journey, it's helpful to understand what happens when a pod spawns and Cilium sets up its networking. + +The process starts when Kubernetes assigns a pod to a specific node. The kubelet initiates the pod setup and calls the configured CNI plugin (Cilium, in our case). The Cilium CNI plugin performs several operations: + +1. **Allocates an IP** using IPAM (IP Address Management) +2. **Creates a link device** (in our setup, a veth pair. One end in the host namespace, one in the pod) +3. **Configures the pod network** by setting the IP address, configuring routes and setting sysctl parameters +4. **Creates a Cilium endpoint** via the Cilium agent API +5. **Retrieves or allocates a security identity** using pod labels +6. **Calculates network policy** for this endpoint +7. **Generates, compiles, and injects eBPF****code** into the kernel +8. **Returns success** to the kubelet + +The key thing to understand is that when creating a Pod, and thus an endpoint, the CNI plugin calls the Cilium agent, which runs as a DaemonSet on each node. The agent does the heavy lifting: managing eBPF maps, handling connection tracking, applying network policies, and more. + +Here's a simplified view of the flow: + +![Flowchart of network setup process involving cabling, CNC plugin, Citrux agent, and kernel.](https://media.ffycdn.net/eu/adyen/M187QTMVPa8r7VWxzjje.png?format=webp&width=654) + +A vital part of this process occurs during endpoint creation: Cilium triggers the connection tracking table's garbage collection to verify that the pod's IP is free of residual connections from its previous run. The scrubIPsInConntrackTable function executes this operation, scanning the entire conntrack table to identify and delete relevant entries. + +The consequences of failing garbage collection are significant. Stale entries accumulate without limit, and while our table can accommodate up to 16 million entries, the real bottleneck is the mandatory scan that every new pod requires. This cleaning step is essential to mitigate IP address reuse conflicts, ensuring that a fresh pod doesn't inherit any open connections from a prior one. + +This is where our story really begins. + +*Note:* [Arthur Chiao's excellent deep dive into Cilium's CNI implementation](https://arthurchiao.art/blog/cilium-code-cni-create-network/) *provided much of our understanding of the CNI flow. While Arthur Chiao wrote it a few years ago, the core concepts remain relevant, and it is an invaluable resource for anyone wanting to understand how Cilium works under the hood.* + +## Down the rabbit hole: tracing the root cause + +### What is the agent actually doing? + +We needed to see what the Cilium agent was doing when it hung. Enter pprof, Go's built-in profiler. We enabled pprof in Cilium's configuration and captured traces from a broken node right after spawning 50 pods. + +The CPU and memory profiles didn't reveal much at first. But when we opened the execution traces, the timeline view showed exactly what each goroutine was doing, and everything became clear. + +![A dashboard showcasing transaction timelines and payment activity analysis by Adyen](https://media.ffycdn.net/eu/adyen/MHhSao2fpiwuNajVcGnA.png?format=webp&width=654) + +We saw long-running goroutines spending their entire time in syscalls. Zooming in closer revealed the pattern: + +![Timeline of payment processing activities on an Adyen platform dashboard.](https://media.ffycdn.net/eu/adyen/PkyaDK9owX4GhhFdVJyn.png?format=webp&width=654) + +The agent was making BPF syscalls in a tight loop: nextKey(), lookup(), nextKey(), lookup(), over and over again. The agent uses these syscalls to iterate through an eBPF map: + +1. **nextKey(currentKey)** - Gets the next key in the BPF map after currentKey +2. **lookup(key)** - Retrieves the value associated with key + +Let's do some napkin math. We selected a 35-millisecond fragment and counted 14,916 syscall occurrences. That's 426,171 syscalls per second. Since getting each element from an eBPF map takes two syscalls (next + get), we were iterating through roughly 213,000 entries per second. + +As Cilium is a user-space process, accessing or manipulating the map forces a context switch on every syscall. A context switch is the process of the CPU temporarily halting the user-space process (Cilium) to execute code in the kernel (to handle the BPF syscall) and then resuming the user-space process. This involves saving and restoring the entire state of the CPU registers and memory space, which adds significant overhead and is a major source of latency. + +Here's the critical insight: this is a sequential operation. You can't parallelize it because you need the current key to get the next key. Even on our powerful server CPUs, this was the maximum speed we could achieve. And there was one more important detail in the traces: this was happening in the scrubIPsInConntrackTable function, the one that cleans the connection tracking table when creating an endpoint. + +## Why so many syscalls? The connection tracking table + +That's when we remembered something we'd seen in cilium status --verbose: + + ``` +BPF Maps: dynamic sizing: on (ratio: 0.005000) + Name Size + TCP connection tracking 16777216 + Non-TCP connection tracking 14232516 + ... +``` + Our TCP connection tracking table had a maximum size of 16 million entries, two times the default of eight million for a machine with 2TB of RAM. We had deliberately configured `**cilium_bpf_map_dynamic_size_ratio: 0.0050`** in our Helm Chart months earlier, fully expecting the connection tracking capacity to scale proportionally with memory across our node types. At the time, this was a planned scaling adjustment to prevent connection tracking exhaustion on our 512GB RAM machines under high-throughput workloads. This configuration worked as intended to support our unique workload evolution, but as our traffic grew, it created a new scaling challenge in how the larger table interacted with the garbage collection auto-scaling algorithm on our high-resource 2TB RAM nodes. + +But how many entries were actually in the table? We listed them with `**cilium bpf ct list global`**. This showed about 7 million entries. But here's what made it interesting: when we checked the timestamp of the *expires* field against the system uptime, we found that the vast majority were already expired, some by several hours. This pointed us directly toward the garbage collection mechanism. + +Now the napkin math gets interesting: + +- Maximum iteration speed: ~200,000 entries per second +- Current table size: 7 million entries +- Time to walk the current table: 35 seconds +- Maximum table size: 16 million entries (at worst) +- Time to walk the full table: 80 seconds + +That means every time we created or deleted an endpoint, we would spend as much as 35 seconds iterating through the connection tracking table, and in the worst-case scenario, a maximum of 80 seconds. Keep in mind our CNI timeout constraint is 90 seconds. + +But wait, if the entries were expired, why wasn’t the garbage collector cleaning them up? + +## Why aren't expired entries cleaned up? Garbage collection gone wrong + +Cilium has garbage collection for the connection tracking table. Looking at the metrics, GC was triggering quite often: + +![Garbage collection activity over time showing traffic spikes and moderate loads.](https://media.ffycdn.net/eu/adyen/7Sd3z8m2Z8vrh9AZ2qLS.png?format=webp&width=654) + +But when we looked at what the garbage collector was actually deleting, we saw a problem: + +![Graph showing payment activity spikes over a 24-hour period with Adyen transaction data](https://media.ffycdn.net/eu/adyen/qpPdTL4e1p1RSDrp33o3.png?format=webp&width=654) + +The garbage collector was running frequently but deleting almost nothing most of the time. Only occasionally would it delete significant numbers of entries. + +Digging into the code revealed why. Cilium has two types of GC operations: + +1. **Endpoint-specific cleanup** : When creating or deleting an endpoint, clean entries matching that endpoint's IP +2. **Periodic expired entry cleanup** : Runs on an interval to remove all expired entries + +Pod creation and deletion triggered almost all of the frequent GC runs in the metrics (the type 1 endpoint-specific cleanups). The periodic cleanup (type 2), which removes expired entries, was barely running at all. + +Why? Because the GC interval auto-scales based on how much it deletes: + + ``` +// Simplified from Cilium source +func GetInterval(interval time.Duration, maxDeleteRatio float64) time.Duration { + if maxDeleteRatio > 0.25 { + // Deleted > 25% of entries → GC more frequently + interval = time.Duration(float64(interval) * (1.0 - maxDeleteRatio)) + } else if maxDeleteRatio < 0.05 { + // Deleted < 5% of entries → GC less frequently + interval = time.Duration(float64(interval) * 1.5) + } + if interval > ConntrackGCMaxLRUInterval { + interval = ConntrackGCMaxLRUInterval // 12 hours + } + return interval +} +``` + Here's the problem with a 16-million-entry table: + +- To delete more than 5% (and avoid slowing down), you need to delete 800,000+ entries +- To delete more than 25% (and speed up), you need to delete 4+ million entries +- Starting interval: 5 minutes +- Maximum interval: 12 hours + +Imagine this scenario: + +1. The node initially experiences a low connection volume when it is newly onboarded or runs only light workloads. +2. The garbage collector deletes less than 5% of entries during a run, which fails to trigger more frequent cycles. +3. The garbage collection interval increases progressively from minutes until it reaches the 12-hour maximum. It would start with 7.5 minutes, increase to 11.25 minutes, then to 16.875 minutes, and so on, until eventually reaching the 12-hour limit. +4. Heavy data workloads eventually land on the node and generate a high volume of network traffic. +5. New connections fill the tracking table for up to 12 hours before the garbage collector runs again. + +By the time garbage collection eventually executes, the connection tracking table may have already accumulated up to 16 million stale entries. While the GC run might delete enough records to temporarily restore speed, the excessive delay between cycles is inherently problematic. This lag allows the table to accumulate a large number of stale entries again, forcing every pod created during the long interval after the previous GC to endure the significant performance penalty of a sequential scan. + +![Graph of real-time payment transaction data from an Adyen terminal or system.](https://media.ffycdn.net/eu/adyen/hg44iZQu7MJZUQoadH1L.png?format=webp&width=654) + +## Why does it cascade? Mutex locks, timeouts, and retries + +While a single slow pod spawn is highly inconvenient, the issue escalated significantly when multiple pods tried to spawn simultaneously. + +We captured a goroutine dump from the Cilium agent using gops and analyzed it with a script that groups similar stack traces. The results were revealing: + + ``` +- 57 occurrences: + createEndpoint() → WaitForFirstRegeneration() → waiting on RWMutex +- 57 occurrences: + regenerateBPF() → runPreCompilationSteps() → invoked +- 56 occurrences: + scrubIPsInConntrackTable() → garbageCollectConntrack() → waiting for Lock +``` + +There's a **global mutex on the connection tracking table**. When we spawn 50 pods at once: + +- Pod 1 acquires the lock and starts the 80-second table iteration +- Pods 2-50 queue up waiting for the lock +- Pod 1 finishes after 80 seconds +- Pod 2 acquires the lock, starts another 80-second iteration +- But Pod 2's timer started 80 seconds ago → timeout at 90 seconds +- Pod 2 times out +- Pods 3-50 never stand a chance + +Here's where it gets worse. When the CNI times out after 90 seconds: + +- The timeout returns an error to the caller +- **But the underlying work doesn't stop** , the agent keeps iterating +- The container runtime (containerd) immediately calls DeleteEndpoint() +- Delete also needs to walk the conntrack table +- Now the system queues up both create and delete operations + +And then Kubernetes retries: + +- The kubelet's *podWorkerLoop* retries after 60-90 seconds (with jitter) +- Each retry adds another endpoint creation and deletion request to the queue +- The queue grows faster than it drains + +We could see this in the logs. For one pod (cilium-node-breaker-5), we saw: + +- 15:54:30 - Create endpoint (attempt 1) +- 15:56:00 - Delete endpoint (timeout) +- 15:57:12 - Create endpoint (attempt 2) +- 15:59:57 - Create endpoint (attempt 3) +- 16:02:39 - Create endpoint (attempt 4) + +The node enters a contention cycle: new work arrives faster than old work completes, and the queue never drains. + +Here's the full picture: + +![Flowchart illustrating a payment process with Adyen terminals, OLV plugin, and checkout steps.](https://media.ffycdn.net/eu/adyen/HyPy5BaaAacU1fDny7AG.png?format=webp&width=654) + +## The fix: one line to rule them all + +After all that investigation, the fix was anticlimactic in its simplicity. We couldn't rely on the auto-scaling GC interval because it would inevitably grow too long on quiet nodes. Hence, we prevented the GC interval from auto-scaling by setting a fixed value: + + `conntrackGCInterval: 60s` + That's it. One configuration line ensures garbage collection runs at least every minute, regardless of how much it deletes. We applied the change at 9:00 am and completed the DaemonSet rollout by 10:00 am. The results speak for themselves: + +![Line chart displaying payment transaction data over time at an Adyen endpoint](https://media.ffycdn.net/eu/adyen/A59UxKvaWSPShkLhkfnh.png?format=webp&width=654) + +The conntrack table size dropped dramatically and stayed stable. More importantly, the API call durations returned to normal: + +![Bar chart illustrating transaction volume and payment data for Adyen services over time.](https://media.ffycdn.net/eu/adyen/tMDUBu8rbzUUt74PNTok.png?format=webp&width=654) + +![Graph showing transaction volume and payment activity data from an Adyen system.](https://media.ffycdn.net/eu/adyen/dvCQHAP2N1bS5Yj2WsFE.png?format=webp&width=654) + +Since the fix, we haven't seen a single instance of the timeout error. Pod spawn times are reliable again. + +## Lessons learned + +**Scaling parameters can have long-tail interactions.** Our proactive tuning of bpf_map_dynamic_size_ratio to support workload scaling on 512GB RAM machines successfully resolved initial capacity limits. However, as our analytical workloads evolved and traffic increased, the larger table size dynamically allocated on our 2TB RAM machines revealed a subtle interaction with the CNI's GC auto-scaling algorithm. These scaling parameters can take months to show their full impact as traffic patterns grow, particularly in environments with adaptive background loops. + +**Observability and full-stack understanding are critical.** While logs showed symptoms, we needed profiling and tracing to reveal the root cause across the entire stack. The container runtime (timeouts, delete behavior), the CNI plugin (timeout values), the Cilium agent (mutex locks, GC logic), and the Linux kernel (eBPF maps, syscall performance) were all relevant to understanding how we got to the pod spawn timeouts. On top of that, adding napkin math with real numbers proved very powerful. Once we had the key numbers, 200k syscalls/sec, 12M table entries and a 90-second timeout, we determined the root cause far before we understood the full chain. Always measure your system's actual performance characteristics, not just theoretical limits. + +**Auto-scaling algorithms need bounds.** Cilium's GC interval auto-scaling makes sense for most deployments: if you're deleting lots of entries, run GC more often; if you're deleting few entries, save CPU by running GC less often. But the algorithm didn't account for varying workloads, where a machine has a low connection volume for a prolonged period, after which, with a single pod introduction, it could get a very high connection volume. Nor did the algorithm account for very large tables where "5% of entries" is an enormous absolute number. The 12-hour maximum interval was too long for our workload. Auto-scaling without careful consideration of edge cases can backfire. + +**Timeouts don't stop work**. When the CNI timed out, we assumed the work would stop. It didn't. The agent kept processing in the background while new requests queued up. This is a common pattern in distributed systems: timeouts protect the caller but don't necessarily cancel the operation. Be explicit about cancellation when needed. + +**Treat conntrack health as a first-class operational metric**. The difference between a healthy cluster and a contention cycle showed up clearly in some metrics we weren't watching:  + +- GC duration - cilium_datapath_conntrack_gc_duration_seconds - jumped from 1s to 80s +- Table size - cilium_datapath_conntrack_gc_entries - 7M entries, mostly expired + +Proactively alerting on these metrics is something we now recommend for any Cilium deployment with dynamic workloads, alongside setting `**conntrackGCInterval: 60s`**. Don't optimise for CPU savings during quiet periods at the expense of pod spawn timeouts during busy periods. + +## Conclusion + +A single configuration line ultimately resolved the mysterious timeout error that impacted our ability to spawn new pods on our big data platform: conntrackGCInterval: 60s. Our investigation revealed that the root cause of our pod timeouts was Cilium's auto-scaling garbage collection algorithm, allowing the cleanup interval to grow to 12 hours, leading to a massive accumulation of expired entries and a linear-time iteration trap. + +This experience provided major takeaways regarding system resilience and the necessity of a full-stack understanding. We learned that scaling parameters and resource allocations can have long-tail interactions that only surface months later as workloads evolve. Furthermore, we discovered that auto-scaling algorithms require strict bounds to prevent unexpected performance degradation in edge cases, such as the varying connection volumes we see on our high-resource machines. The investigation also highlighted that timeouts often only protect the caller, without stopping the underlying work, potentially triggering a contention cycle of retries that we could only diagnose through deep observability into mutex locks, syscalls, and eBPF internals. + +As we move forward, we must ask ourselves: are the adaptive behaviours in our infrastructure truly protecting us, or are they masking inefficiencies that only appear at peak capacity? By treating conntrack health as a first-class operational metric and prioritising reliability over minor CPU savings, we can build more robust systems. And remember, if you ever see mysterious timeouts in your CNI: sometimes the answer hides in 426,000 syscalls per second. diff --git a/sreweekly/markdown/532/04-storage-at-scale-what-i-actually-watched.md b/sreweekly/markdown/532/04-storage-at-scale-what-i-actually-watched.md index 0a7d39a4..8a44d81e 100644 --- a/sreweekly/markdown/532/04-storage-at-scale-what-i-actually-watched.md +++ b/sreweekly/markdown/532/04-storage-at-scale-what-i-actually-watched.md @@ -7,3 +7,47 @@ ## 简介 > For eight years I ran SRE for a storage system measured in exabytes. The dashboard I checked every morning shrank to seven numbers. Here they are. + +## 正文 + +For eight years I ran the SRE team behind a storage system measured in exabytes. Over time, the dashboard I checked every morning shrank to a handful of numbers. These are the seven that told me whether the service was healthy. + +Availability tells you if the system is up. Durability tells you if your data is still there. The two are not the same. + + +Here’s the short version. + +| KPI | What it measures | How we tracked it | +|---|---|---| +| Availability | Percent of requests succeeding | 99.99% per region, per service | +| Durability | Probability your data survives | 11 nines (10^-11 annual loss) | +| TTFB | Time to first byte returned | p50, p95, p99 latency per object size | +| Canaries | Synthetic test traffic | Continuous PUT/GET from every region | +| Hotspots | Skew across storage nodes | Top-N node load vs cluster median | +| IOPS | Operations per second | Read/write IOPS per shard, per disk | +| DB Shards | Metadata partition health | Shard CPU, lag, hot-key skew | + +## Availability and durability are the two non-negotiables + +Availability is uptime. Durability is whether the data survives. You can be 100% available and lose data, you can be 100% durable and offline. Customers care about both. We hit 11 nines of durability by writing every object to multiple availability domains with erasure coding, and proved it monthly with a recovery drill. + +## TTFB is what users actually feel + +Aggregate availability hides slow tails. A 99.99% available service with a p99 TTFB of 2 seconds feels broken. Always track latency by object size bucket. A 10 MB read should not share an SLO with a 100 byte HEAD. + +## Canaries are your truth + +Customers don’t tell you when they’re sad. They leave. Canaries are synthetic PUT/GET/LIST traffic running continuously from every region. If a canary fails for 30 seconds, you find out before your customer’s pager goes off. + +## Hotspots and IOPS surface the silent failures + +A storage cluster can be 99.99% available while one node is on fire. Track per-node IOPS and bytes-served, and alert on the top-N nodes diverging from cluster median. Hotspots are the leading indicator of a customer key range overwhelming a shard. + +## DB shards are the part nobody talks about + +Object storage looks stateless, but the metadata layer is a sharded database. One hot shard, one rebalance gone wrong, and your control plane stalls. Watch shard CPU, replication lag, and hot-key skew the same way you watch the data plane. + +The data plane scales. The control plane bites. + + +Those seven numbers, watched together, told me almost everything I needed to know about whether the service was healthy. diff --git a/sreweekly/markdown/532/05-the-rise-of-cognitive-observability.md b/sreweekly/markdown/532/05-the-rise-of-cognitive-observability.md index 75f70c8b..1873f275 100644 --- a/sreweekly/markdown/532/05-the-rise-of-cognitive-observability.md +++ b/sreweekly/markdown/532/05-the-rise-of-cognitive-observability.md @@ -7,3 +7,7 @@ ## 简介 > Traditional observability monitors execution. LLM observability must monitor behavior. + +## 正文 + +> ⚠️ 抓取失败:HTTP 403 diff --git a/sreweekly/markdown/532/06-how-uber-conquered-database-overload-the-journey-from-static-rate-limi.md b/sreweekly/markdown/532/06-how-uber-conquered-database-overload-the-journey-from-static-rate-limi.md index c96ed01d..4c660ebd 100644 --- a/sreweekly/markdown/532/06-how-uber-conquered-database-overload-the-journey-from-static-rate-limi.md +++ b/sreweekly/markdown/532/06-how-uber-conquered-database-overload-the-journey-from-static-rate-limi.md @@ -7,3 +7,208 @@ ## 简介 This one has a lot of great detail on how their approaches to quota management failed and how they iterated. + +## 正文 + +# How Uber Conquered Database Overload: The Journey from Static Rate-Limiting to Intelligent Load Management + +# Introduction + +Uber’s thousands of microservices handle traffic for over 170 million monthly active users: riders, Uber Eats users, drivers, and couriers. At the heart of this infrastructure are [Docstore](https://www.uber.com/us/en/blog/schemaless-sql-database/?uclick_id=0ef55d57-e714-4820-8ab3-c3f30875480f) and [Schemaless](https://www.uber.com/us/en/blog/schemaless-part-one-mysql-datastore/?uclick_id=0ef55d57-e714-4820-8ab3-c3f30875480f), Uber’s in-house distributed databases built on top of MySQL®. These databases span thousands of clusters, store tens of petabytes of operational data, and serve tens of millions of requests per second with billions of rows read or updated. They back some of the most latency-sensitive and mission-critical workloads, powering every business vertical at Uber: from rides and deliveries to maps, payments, and beyond.  + +At this scale, even minor overloads aren’t isolated events, they cascade. A brief spike in one part of the system can ripple outward: downstream services time out, retries pile up, and degradation amplifies into broader failure. In a multitenant environment, it’s also critical to ensure fairness and prevent any tenant from hogging all the resources. With workloads varying in traffic shape, latency profiles, and system impact, building effective overload protection is a uniquely challenging problem. + +The cost of getting overload protection wrong is steep. This blog shares how we built an intelligent load manager that detects overload from multiple signals to keep our databases stable and fair under pressure. + +## Docstore and Schemaless + +Before diving into the load manager that protects Uber’s databases, let’s walk through their architecture. + +While [Docstore](https://www.uber.com/us/en/blog/schemaless-sql-database/?uclick_id=0ef55d57-e714-4820-8ab3-c3f30875480f) supports transactions with full CRUD operations and [Schemaless](https://www.uber.com/us/en/blog/schemaless-part-one-mysql-datastore/?uclick_id=0ef55d57-e714-4820-8ab3-c3f30875480f) is optimized for append-only workloads, both share a common architectural foundation. It comprises three primary layers: a stateless query engine, a stateful storage engine, and a control plane. For the scope of this blog, we’ll focus on the query and storage engine layers. + +The stateless query engine is responsible for query planning, request routing, sharding, schema management, authorization, request parsing, and validation. It serves as the routing layer: coordinating and validating client requests before handing them off to the storage layer. + +The stateful storage engine handles transaction management, connection pooling, consensus, and replication. Data is sharded across multiple partitions, with each partition consisting of one leader and two followers, coordinated via [Raft](https://www.scs.stanford.edu/~zyedidia/docs/papers/raft.pdf) to ensure strong consistency. Each partition is backed by MySQL nodes with locally attached NVMe SSDs, built to support high-throughput, low-latency workloads at scale. + +## Challenges + +### Quota-Based Rate Limiting in the Query Engine Layer + +Initially, we explored a quota based rate-limiting approach within the stateless query engine layer. The concept was simple: assign each read and write request a capacity unit cost based on bytes processed, grant users fixed quotas, and return a 429 when those quotas were exceeded. Since routing nodes were stateless, we stored quota usage in a central Redis® cache. While conceptually sound, this approach didn’t hold up in production. + +First, it added unnecessary complexity. Every request required a Redis call, introducing a new point of failure and the overhead of an additional network hop. + +Further, for the stateless routing layer to accurately shed requests for an overloaded storage partition, it’d need to maintain realtime health and load information for thousands of partitions across the system. This introduced a lot of tracking overhead, undermining the scalability of the architecture. + +The cost model was also too imprecise. In Docstore and Schemaless, due to the way MySQL handles scanning and filtering, a query that performs a full table scan but returns a single row was assigned the same capacity cost as a query that only reads a single row. This fundamental flaw in our metering meant that lightweight and heavyweight operations were treated the same, making quota enforcement unreliable. + +Finally, quotas were defined statically, resulting in frequent requests from stakeholders to adjust their quotas, making them ineffective in multitenant environments. + +Despite its initial promise, this approach failed. But it gave us a crucial insight: overload management must live as close to the storage nodes as possible. That realization became a cornerstone of the final design in the stateful storage layer. + +### Identifying the Right Signal for Overload + +A core challenge in designing a resilient load manager is choosing a reliable signal for overload. Simple QPS-based rate limiting is too coarse. It fails to account for workload variability, often shedding too late or too early. What can be more effective is concurrency: the number of operations currently in flight. It directly reflects system load, following Little’s Law: *Concurrency = Throughput × Latency*. In stateful systems, it maps closely to resource usage, making it a more dependable indicator. + +### Balancing Resilience and Fairness + +Balancing resilience and fairness is a core challenge in multitenant systems. During system-wide stress, we want to shed traffic by priority, dropping low-priority requests first. But when a single noisy actor hogs resources without triggering global overload, we also need per-tenant rate limiting that works independently of the system load. This dual requirement led us to combine dynamic overload detectors with fairness enforcement mechanisms that operate in parallel. + +## Building the Foundation of a Unified Load Manager + +### Controlled Delay: Smarter Queuing Under Pressure + +The load-shedding journey began with [CoDel](https://queue.acm.org/detail.cfm?id=2209336) (Controlled Delay), a concept borrowed from networking to combat bufferbloat. Instead of shedding based on queue length, CoDel looks at how long requests wait in the queue: favoring responsiveness over volume. + +We implemented separate CoDel queues for each operation type: + +- **Read queue** : for point lookups and light queries +- **Write queue** : for insert, update, and upsert operations +- **Slow queue** : for long-running and background operations like scans, deletes, or replication + +Each queue was managed independently, giving us better isolation across workloads. + +FIFO queuing wasn’t enough because a pure FIFO queue processes requests in arrival order, which works well when traffic is stable. But under overload, FIFO creates a trap: old requests accumulate, wait too long, and often get abandoned or retried by the client. This results in wasted work. Meanwhile, fresh requests, still relevant and likely to succeed, sit idle at the end of the line. + +CoDel introduces adaptive LIFO to solve this. Figure 5 shows how it works. + +Under normal load, the queue behaves as FIFO. Under pressure, it switches to LIFO, favoring newer requests that still have a chance to succeed. This simple shift improves responsiveness by failing fast, shedding stale work, and giving fresh requests priority. + +### Scorecard Engine + +The Scorecard engine is a rule-based admission control component and a lightweight quota system designed to enforce per-tenant concurrency limits in multitenant environments. While load-shedding protects the system during overload, Scorecard ensures that no single tenant can dominate shared infrastructure, even in normal conditions. + +The configuration is simple and deterministic. + +The primary benefit of the Scorecard lies in incident containment. It helps pinpoint the source of disruption during outages or traffic spikes. It isolates and caps misbehaving tenants without disrupting others, balances stability during normal load with strict limits under stress, and reduces blast radius during overload events by enforcing boundaries quickly and deterministically. + +The Scorecard provides predictable fairness and blast radius control, especially when multiple tenants are competing for shared resources. + +### Regulators + +While Scorecard protects against concurrency-based overuse, it doesn’t cover all the ways a stateful database system can overload. Some forms of skews are subtle. They don’t show up in concurrency saturation, but they can still degrade system performance if left unchecked. + +For example, a low QPS caller can still overload the system by sending large write payloads. Or, traffic skewed to one partition key can overload a single cluster while others sit idle. + +To guard against these skewed behaviors, we introduced plug-in regulators: node-local overload detectors that enforce invariants the system mustn’t violate. They rarely trigger during healthy operation, and that’s by design. At the same time, when users accidentally create hotspots or large data ingestions, regulators kick in to prevent cascading failures. + +We use these regulators: + +- **Write bytes regulator:** Limits concurrent write volume to prevent I/O saturation +- **Partition key regulator:** Throttles traffic targeting hot partition keys +- **Memory regulator:** Tracks free process memory and throttles when we’re low on memory +- **Goroutines regulator:** Tracks total number of goroutines and throttles when it exceeds threshold + +### What Worked Well + +By shedding excess requests, our CoDel queues prevented runaway resource exhaustion, which led to improved stability and a higher success rate for accepted requests. This approach was particularly effective at ensuring that core system functionality remained available during overloads. + +The Scorecard engine successfully isolated misbehaving tenants by enforcing per-tenant concurrency limits. This allowed us to quickly contain disruptions from noisy neighbors without penalizing other users, ensuring that shared resources were used fairly. + +### Limitations + +While this initial setup laid the foundation for overload protection and fairness, it came with a few limitations. First, CoDel treated all requests equally, dropping low-priority and user-facing traffic alike, leading to a bad customer experience and increased on-call load. + +CoDel also relied on fixed queue timeouts and static inflight concurrency limits, which can be a low-fidelity solution for a dynamic system, requiring frequent manual tuning and leading to operational toil. + +The fixed, static wait times in CoDel led to a thundering herd problem. When requests were eventually rejected, they’d all retry at once, triggering repeated cycles of overload and rejection. During these periods, the lack of traffic differentiation meant even high-priority requests were dropped, leading to customer-visible errors and amplifying the blast radius. + +Ultimately, it kept things from breaking, but lacked the nuance and dynamism required for a high-quality user experience. This highlighted the need for dynamic and priority-aware queues. + +## Evolving the Architecture + +### Cinnamon Replaces CoDel + +We observed that many overloads stemmed from low-priority, asynchronous jobs: pipelines, aggregators, and internal garbage collection flows. These shouldn’t have the same survivability as ride requests or real-time pricing queries. + +To address this, we replaced CoDel with [Cinnamon](https://www.uber.com/us/en/blog/cinnamon-using-century-old-tech-to-build-a-mean-load-shedder/), a priority-aware load shedder developed by the Delivery team at Uber. Cinnamon makes smarter shedding decisions by considering request rank, dynamic system state, and the relative importance of workloads.  + +Request rank is derived from the priority attached to the request, and if no explicit priority is present, Cinnamon assigns a default based on the calling service. Priority is defined using a tiering model from tier 0 (t0) for the most critical traffic to tier 5 (t5) for the least. While t0 is reserved for a small subset of critical infrastructure services, t1 represents the most important user facing online traffic, the core workloads we aim to protect during overloads. This system allows Cinnamon to shed lower-priority traffic first during overload. + +With request priority awareness in place, we simplified the queue structure to just read and write queues. Long-running and background operations were marked with lower priority instead of having a separate queue. + +Before Cinnamon, the CoDel queue load shedder was priority-agnostic and shedding during overload was indiscriminate. + +After Cinnamon, the queue load shedder was priority-aware and shedding during overload happened in order of priority. + +### Performance and Stability Gains + +We saw performance and stability gains from the Cinnamon-based design. Requests are ranked, allowing Cinnamon to shed low-priority traffic first, protecting user facing flows. During overloads, critical user-facing requests are better protected with minimal impact. + +Cinnamon also adapts queue timeout thresholds using P90 latency metrics, eliminating the need for manual tuning. Moreover, its [Auto Tuner](https://www.uber.com/blog/cinnamon-auto-tuner-adaptive-concurrency-in-the-wild/) dynamically adjusts inflight limits, represented by the available slots in the blue box in Figure 10, to maximize throughput. It does this by continuously monitoring and reacting to realtime latency and error rate signals, ensuring stable and effective load shedding. + +Unlike CoDel’s static approach, which aggressively rejects all requests after a fixed wait time, like 5 milliseconds, Cinnamon’s [PID-based control](https://www.uber.com/us/en/blog/pid-controller-for-cinnamon/) allows the system to absorb pressure without overreacting. It dynamically adjusts queue timeouts and inflight limits based on realtime latency and error signals, shedding only when necessary. This prevents a large class of premature shedding that would otherwise lead to unnecessary rejections, retries, and thundering herd effects. The result is smoother recovery, fewer 429s, and more consistent availability without compromising system health. + +### Areas for Improvement + +Despite the gains from Cinnamon, some key challenges remained, highlighting the need for a unified platform. + +The load manager acted based on the local health of the server, tracking signals like inflight concurrency, write bytes, or memory usage. But in distributed systems, overload isn’t always local. A leader node may need to shed traffic because follower nodes are lagging, even if it’s healthy itself. We call this commit index lag. Traditionally, external components using token-bucket-based rate limiters handled such remote shedding decisions. These were easy to build but proved ineffective at scale, introducing split-brain behaviors and globally suboptimal shedding decisions. + +The initial design was excellent for concurrency-based shedding, but it wasn’t built to be a reusable platform for future overload signals that would inevitably arise from a growing system. + +These insights led us to the final evolution of our system: transforming Cinnamon from a concurrency only shedder into a truly general purpose overload control engine. By consolidating all signals into a single, modular decision-making loop, we achieved holistic and consistent overload management. + +## The Unified Load Shedding Engine + +### Centralizing Overload Decisions + +We enhanced Cinnamon to support pluggable external signals like follower commit lags, enabling the system to make globally informed, priority-aware shedding decisions within the same admission control path. This shift unified local and remote overload logic into a single control loop, closing the gaps that previously caused instability. + +But shedding isn’t always a one-size-fits-all decision and that’s where the load manager architecture shines. Built on a BYOS (Bring Your Own Signal) ethos, it provides a pluggable framework that lets the team embed new overload signals and route them to the right control path. Whether the pressure is systemic or actor-specific, the load manager sheds broadly by priority or precisely by caller, based on the signal. + +### The Payoff: Unified Control, Simplified Load Management + +The shift to a centralized, pluggable architecture made the system more stable and predictable, with real wins. + +Cinnamon sheds excess requests immediately using a PID controller, avoiding the memory and goroutine buildup caused by token bucket limiters. This led to lower tail latencies and a leaner resource usage profile, even under heavy load. We saw: + +- 80% increase in throughput under overload (QPS average of 5,400 versus 3,000) +- ~70% reduction in P99 latency (upsert average of 1.0 seconds versus 3.1 seconds) +- ~93% fewer goroutines during overload (peak 10,000 versus 150,000) +- ~60% lower heap usage (1 GB max versus 5-6 GB spikes) + + +We also saw smoother, more predictable shedding behavior. Without PID regulation, shedding acts like a hammer: reactive and abrupt. With it, it’s more like a dimmer switch: smooth and stable. The difference is clear when comparing how commit lag stabilizes under a token bucket limiter versus Cinnamon’s PID-based controller. + +## Lessons Learned + +- **Prioritization is paramount.** Effective load-shedding starts with deciding what matters most. Protect critical, user-facing traffic first. Everything else is secondary. +- **Fail fast, don’t block.** Rejecting early is almost always better than holding requests in memory until they expire. It reduces wasted work, keeps latencies predictable, prevents OOMs, and makes the system more resilient under stress. +- **PID regulation for stable shedding** . Simple, reactive shedding based solely on current error rates often causes instability, overcorrecting too late, and too hard. PID based regulation brings balance by incorporating system history and directional trends, making it a critical tool for smooth, sustained, and resilient overload control. +- **Place control close to the source of truth.** The best shedding decisions happen where the state lives. Protection in the layer that has full context, typically the storage layer in stateful systems. +- **Embrace dynamism.** Avoid static configurations wherever possible. Your system should be intelligent enough to adapt to different scenarios, based on the context. +- **Invest in visibility and monitoring.** Good observability is the foundation for tuning and trust. Track what’s being shed, why it’s being shed, and how each component contributes to system pressure. +- **Simplicity over complexity.** This is a meta principle that guides all the other decisions. + +# Conclusion + +Our journey to a resilient load manager was defined by the unique complexities of a large-scale, stateful, and distributed environment. By unifying disparate components into a single decision-making brain and adopting a Bring Your Own Signal model, we gained the flexibility to handle systemic overloads and localized noisy neighbor issues with precision. The result is a load management system that sheds smarter in a priority-aware manner, keeps tail latencies low, and drastically reduces operational toil. + +If you like challenges related to distributed systems, databases, storage, and cache, apply for open positions [here](https://www.uber.com/us/en/careers/list/?query=storage&department=Engineering). + +## Acknowledgments + +A project of this scope is rarely accomplished alone. Our sincere thanks to Rich Porter, Jesper Nielsen, Piyush Patel, and the engineers from the Storage and Delivery teams for their guidance and collaboration throughout this journey. From design reviews to on-call insights, their contributions were instrumental in building a resilient system that now safeguards some of Uber’s most critical infrastructure. + +*Cover Photo Attribution: “[Heavy Traffic Jam in Urban City Center](https://www.pexels.com/photo/heavy-traffic-jam-in-urban-city-center-32487428/)” by [Dapur Melodi](https://www.pexels.com/@dapur-melodi-192125/)* + +*MySQL is a registered trademark of Oracle and/or its affiliates. Other names may be trademarks of their respective owners.* + +*Redis is a trademark of Redis Labs Ltd. Any rights therein are reserved to Redis Labs Ltd. Any use herein is for referential purposes only and does not indicate any sponsorship, endorsement or affiliation between Redis and Uber.* + +Dhyanam Vaidya + +Dhyanam Vaidya is a Software Engineer on Uber’s Storage Platform team. He’s contributed to the design and implementation of many Docstore features. His work focuses on improving the reliability, resilience, and operational efficiency of Uber’s distributed databases at scale. + +Prathamesh Deshpande + +Prathamesh Deshpande is a Staff Engineer on Uber’s Storage Platform team, building database features and distributed storage systems that meet Uber’s global reliability and performance requirements. His work focuses on large-scale data management, distributed database storage systems, and platform reliability. + +Mike Ma + +Mike Ma is a Staff Software Engineer on Uber’s Storage Platform team, where he has contributed to multiple core components of both Schemaless and Docstore. His work focuses on scalability, reliability, performance, and operational excellence across Uber’s large scale distributed databases. + +Chaitanya Yalamanchili + +Chaitanya Yalamanchili is a Sr. Manager and technical lead on Uber’s Storage Platform team. He leads the development of online distributed storage systems with a focus on providing a world-class platform that powers all the critical business functions and lines of business at Uber. The platform serves tens of millions of QPS and stores tens of Petabytes of operational data. diff --git a/sreweekly/markdown/532/07-voyager-and-the-art-of-graceful-degradation.md b/sreweekly/markdown/532/07-voyager-and-the-art-of-graceful-degradation.md index 45d756a4..31558194 100644 --- a/sreweekly/markdown/532/07-voyager-and-the-art-of-graceful-degradation.md +++ b/sreweekly/markdown/532/07-voyager-and-the-art-of-graceful-degradation.md @@ -7,3 +7,121 @@ ## 简介 This article uses Voyager 1, whose engineers just shut down another instrument to conserve its steadily-decaying power, as an extended analogy for graceful degradation. + +## 正文 + +### Voyager and the Art of Graceful Degradation + +Like a great celestial swan, Voyager 1 is flying `—` swiftly, boldly, albeit a little stiffly in places.  + +| ![https://assets.science.nasa.gov/dynamicimage/assets/science/missions/voyager/images/1_Voyager_artist_concept.jpg?w=2000&h=1125&fit=crop&crop=faces%2Cfocalpoint](https://assets.science.nasa.gov/dynamicimage/assets/science/missions/voyager/images/1_Voyager_artist_concept.jpg?w=2000&h=1125&fit=crop&crop=faces%2Cfocalpoint) | +| [NASA/JPL-Caltech](https://science.nasa.gov/blogs/voyager/2026/04/17/nasa-shuts-off-instrument-on-voyager-1-to-keep-spacecraft-operating/) | + +It moves through interstellar space with enormous momentum, far beyond the planets that once defined its mission, carrying instruments that continue to report from a region no man‑made craft has ever reached. Yet every action it takes is constrained by a finite and steadily diminishing supply of energy, each signal carefully weighed against what it costs to send. + +There is a quiet elegance in that balance. + +Voyager does not insist on doing everything it once did. It does not pursue peak capability when conditions no longer allow it. Instead, it adapts -- releasing some functions so that others can continue, prioritizing what matters most over what is merely possible. + +In engineering, we have a name for systems that behave this way. + +#### We call it **graceful degradation**. + +[Low‑Energy Charged Particle (LECP) detector](https://pds-atmospheres.nmsu.edu/data_and_services/atmospheres_data/Voyager/lecp.html)— an instrument that measures ions, electrons, and cosmic rays to map the structure and pressure of the interstellar medium, helping to define the boundary between the solar system and interstellar space— was shut down to conserve power and extend the spacecraft’s operational life. + +Voyager is powered by a radioisotope thermoelectric generator whose output declines as radioactive fuel decays. Every year, available power drops by a few watts. Unlike systems here on Earth, there is no possibility of provisioning more capacity, no redundancy waiting in reserve, and no “scale out” option. + +Seen through a Site Reliability Engineering lens, Voyager’s power margin is its **error budget.** It defines *how much can go wrong* before the mission begins to suffer. + +Early in the mission, that budget was generous. Minor inefficiencies, unexpected behaviors, and non‑optimal configurations could be tolerated. As the decades passed, the margin narrowed. Today, even a modest, unplanned dip of power by a wayward instrument risks triggering Voyager’s undervoltage fault protection — an automated safeguard that will shut components down abruptly to ensure survival. + +In February, a routine roll maneuver caused such a dip. Engineers understood that allowing the spacecraft to cross that line would mean entering a survival mode where system preservation is prioritized over delivering mission value. + +This moment is familiar to anyone who has operated a production system near its limits: + +- CPU saturation turning latency into user-visible slowness +- Memory pressure triggering process and container termination +- Queues backing up until messages expire undelivered +- Storage exhaustion freezing otherwise healthy transactions + +Graceful degradation is about prioritizing your goals and your capabilities, and as you approach a point where you cannot fulfill all your goals, acting *before* you reach that point. + +- Reduce CPU consumption (lower frame rates, remove animations, disable optional features) +- Defer low‑priority work (batch reports, replace live data with aggregates) +- Prioritize critical traffic and drop nonessential messages +- Reject new transactions when storage thresholds are reached to protect core paths + +In Reliability Engineering, as in much of life, we'd rather have brownouts than blackouts. + +While we never want to disappoint users, we'd rather reduce features rather than take outages. We'll degrade experience -in a controlled fashion - rather than lose the service entirely. We shed load in controlled ways instead of letting cascading failures decide the outcome. + +That is exactly what Voyager’s engineers did. + +Years before this moment, scientists and engineers jointly agreed on a shutdown sequence: which instruments would be sacrificed first as power declined, and which capabilities were most critical to preserve. By April 2026, seven of Voyager 1’s ten original science instruments had already been retired. The LECP was simply next on the list — not because it failed, but because its cost‑to‑value ratio was now unfavorable. + +This is the same decision Site Reliability Engineers (SREs) make when: + +- Disabling expensive recommendation pipelines during peak traffic +- Serving cached or approximate results instead of fully computed ones +- Temporarily turning off background jobs to protect user‑facing latency + +Nothing is broken, *per se*. The system is deliberately choosing to do less so that it can continue to succeed in part rather than fail in total. Graceful degradation is not a weakness; it is a sign of maturity. + +Voyager continues to operate the instruments that provide uniquely valuable data — measuring magnetic fields and plasma waves in interstellar space — while relinquishing others whose contribution, though still useful, no longer justifies their cost. + +Even the LECP shutdown was reversible by design. A small motor that rotates the sensor remains powered, preserving the option of reactivation should future power‑saving measures succeed. + +This is graceful degradation with reversibility in mind. The current state is preserved, while recovery paths maintained and, most importantly, options are left open. Granted, the chances of Voyager suddenly being replenished with fresh plutonium for additional power is exactly 0, but Reliability Engineers here on the ground do plan on overcoming their temporary issues which caused the degradation and using the available options to fully restore services. + +This is why we gate features behind flags instead of deleting code and why we can temporarily change users' capabilities instead of removing them from the system. + +### Balancing Performance, Capacity, and Risk + +Reliability is rarely about maximizing performance. It is about continuously balancing **performance**, **capacity**, and **risk** — especially when capacity is finite and margins are thin. + +Voyager operates permanently at this intersection. + +Performance, in Voyager’s case, is scientific throughput: how many instruments are active, how often measurements are taken, and how much data is returned.  + +Capacity is a steadily shrinking power budget that cannot be replenished.  + +Risk grows as margins shrink: a sudden undervoltage event could trigger autonomous shutdowns that are difficult, slow, and dangerous to recover from across a 23‑hour communication delay. + +Graceful degradation is how the Voyager team manages this triangle. + +By shutting down the LECP before power levels became critical, the team deliberately traded peak scientific performance for reduced operational risk and preserved capacity for the instruments that matter most. + +| ![This illustration shows the various instruments locations on the Voyager spacecraft.](https://science.nasa.gov/wp-content/uploads/2024/03/instruments-3.jpg?w=640) | +| [The status of Voyager's instruments  (NASA/JPL-Caltech)](https://science.nasa.gov/mission/voyager/where-are-voyager-1-and-voyager-2-now/#instrument-status) | + + +**Voyager does less than it once did — but it does so more safely, more predictably, and for longer.** + +This mirrors everyday SRE work: + +- lowering request concurrency to prevent saturation +- reducing image quality or refresh rates under load +- shrinking feature scope during high-risk windows +- renegotiating SLOs instead of pretending nothing has changed + +In each case, performance is intentionally reduced to keep risk within acceptable bounds. + +Because Voyager’s degradation path was defined years in advance, with a healthy system and with management & engineering having time and clarity to make rational trade-offs, the unexpected power dip didn't result in a frantic rush to heroically solve a problem, it triggered a pre-planned process which resulted in a graceful retirement of the instrument chosen ahead of time. No surprises, just good engineering. + +Graceful degradation is a social and organizational capability as much as a technical one. It requires shared understanding across teams, explicit agreement on priorities, and acceptance that loss is inevitable. It's not about preventing failure forever. It is about ensuring that when degradation occurs, it happens on your terms. + +While few of us work on systems like Voyager, there are many commonalities - + +- Our platforms are usually far older than the original business model they were designed to support. +- Our architectures often outlast our architects. +- Our "temporary" services that we built with “temporary” design decisions have become permanent. + +Our systems survive not by staying perfect, but by letting go gracefully. At least, these are the ones which cause the least stress to their owners and maintainers. + +Voyager is still returning data from interstellar space not because nothing has failed, but because failures have been managed thoughtfully, incrementally, and with humility. Twenty-five billion kilometers from Earth, Voyager continues to demonstrate a lesson every experienced SRE eventually learns: + +The systems that last longest are not the ones that cling to every feature, but the ones that decide, and well in advance, which parts they are willing to give up. + +## Comments + +## Post a Comment diff --git a/sreweekly/markdown/533/01-incidents-start-before-the-response-does.md b/sreweekly/markdown/533/01-incidents-start-before-the-response-does.md index a868268a..48dc814b 100644 --- a/sreweekly/markdown/533/01-incidents-start-before-the-response-does.md +++ b/sreweekly/markdown/533/01-incidents-start-before-the-response-does.md @@ -7,3 +7,51 @@ ## 简介 What can you do to shorten the time to detect an incident? Some great ideas in here, especially monitoring your company’s main web page for a sudden uptick in traffic. + +## 正文 + +Your company has probably invested significantly in what happens after an incident is identified: incident response tooling, trained incident commanders, communication protocols, on-call rotations. That investment matters. But what about the gap between when a problem starts and when anyone on your team knows about it? + +During that gap, customer damage is accumulating. The problem is getting worse, the blast radius is expanding, and nobody on the team is doing anything about it because nobody knows yet. + +You can’t eliminate this gap entirely, but you can shrink it. Four investments make the biggest difference. + +## Broaden your detection surface + +Automated monitoring is the first and best line of defense, but it can only catch the failure modes someone thought to check for. Human detection isn’t a gap you can eliminate; it’s a permanent and valuable part of your detection capability. + +This means your customer support team is part of your detection infrastructure, whether or not you’ve told them so. So is any part of your company that interacts with customers regularly: account execs, customer success managers, even your social media team. They talk to your customers every day and often see concerns emerge before engineering does. And don’t overlook your customers themselves, who won’t limit their reports to your “official” support channels. If all these folks don’t have clear, fast escalation paths to flag potential problems for engineering, you have a detection gap that no amount of monitoring investment will close. + +If your company is a heavy user of its own product, the detection surface extends even further. When I led Slack’s incident management program, literally anyone in the company might notice a problem while using Slack internally. Not every company is in that position (it depends entirely on what the product is), but those who are should take advantage of it. Make sure everyone (all the way down to the part-time security guard covering the front desk on weekends) knows how to report problems they see. + +And watch for indirect signals. One of Slack’s best harbingers of “something is broken, even if we don’t know what yet” was the page-view rate on our public status page. If it started surging upward, we knew that *something* was wrong, even if we weren’t getting any other clear signals yet, and we’d start investigating. It was like smelling a light waft of smoke, well before the smoke detectors and fire alarms go off. If you have a public status page, consider adding its traffic patterns to your monitoring. A sudden spike in visits is a low-cost early warning powered by the collective behavior of your user base. + +## Lower barriers to reporting + +Most of these detection channels depend on someone raising a concern, and that only works if the barrier to doing so is low. At many companies, the only mechanism for raising an alarm is to declare an incident, which triggers a full coordinated response: pages go out, a channel is created, an incident commander is assigned, people drop what they’re doing. + +That’s appropriate when you know you have a real problem. But if the only way to raise a concern is to trigger that entire response, people will hesitate, and rightfully so. Nobody wants to be the person who launched a full incident response over a hunch that turns out to be wrong. So they wait for more evidence, and the detection gap grows. + +Think of it like calling emergency services. When you call 911 (or 999, 000, 112, or whatever your country’s emergency number is), you don’t have to know whether you need an ambulance, a fire engine, a hazmat team, or a bomb squad. You describe what you see, and a trained dispatcher determines how serious the situation is, what sort of response is warranted, and who to send. + +Your incident detection should work the same way: make it easy for anyone to say “I think something might be wrong,” and let someone with training, experience, and context determine what response is warranted. At Slack, introducing a lightweight mechanism for exactly this was one of the most impactful things we did. + +## Continuously right-size your alerting + +It’s tempting to close the detection gap by making your monitoring more aggressive: lower the thresholds, add more alerts, page on anything that twitches. This can backfire badly. Every alert that wakes someone at 3 AM and turns out to be nothing makes it a little more tempting for your on-call engineers to dismiss the next one. Alert fatigue is one of the most insidious threats to detection, precisely because it accumulates gradually. Your alerting system doesn’t fail all at once; it erodes, one false alarm at a time, until the real alerts get lost in the noise. + +The discipline runs in both directions: yes, add monitoring when you discover gaps, but regularly prune alerts that aren’t earning their keep. If a service-owning team can’t get through a review of every alert they received in the past week in a reasonable portion of a weekly ops review meeting, they’re getting too many alerts. + +## Examine the gap + +Another way to shrink the detection gap over time is to examine it after every incident. You’re never going to be able to fully automate detection, but it’s still an ideal worth pursuing. Three questions, asked consistently in every post-incident review, create a steady stream of improvements: + +- How long was the gap between when the problem started and when we detected it? +- Could we have detected it sooner? +- What monitoring would we need to add, or what threshold would we need to adjust, to catch this kind of problem faster next time? + +## The bottom line + +Investing in detection is investing in the foundation of your entire incident management capability. You can have well-trained incident commanders, practiced responders, and polished communication protocols, but none of it matters until you know there’s a problem. + +## Recent Comments diff --git a/sreweekly/markdown/533/02-quick-thoughts-on-azure-regional-outage-from-july-23-26.md b/sreweekly/markdown/533/02-quick-thoughts-on-azure-regional-outage-from-july-23-26.md index af487152..6f686bda 100644 --- a/sreweekly/markdown/533/02-quick-thoughts-on-azure-regional-outage-from-july-23-26.md +++ b/sreweekly/markdown/533/02-quick-thoughts-on-azure-regional-outage-from-july-23-26.md @@ -7,3 +7,56 @@ ## 简介 What an interesting incident! I recommend reading Azure’s write-up before reading Lorin’s excellent analysis. + +## 正文 + +The folks at Microsoft Azure recently wrote up a [post incident review for a networking issue in their West U.S region.](https://azure.status.microsoft/en-us/status/history/?trackingId=ZJV6-SGG) From the included timeline, it looks like the impact was on the order of five hours. It’s a pretty short write-up, but let’s take a look at the contributors. + +On 23 July 2026, a break-fix repair was initiated on an optical device to address a network reliability risk. + + +The first contributor mentioned in the write-up was work that was done to repair a device in their networking stack. Here I can’t help but think of the first bullet in my [conjecture on why reliable systems fail](https://surfingcomplexity.blog/2017/06/24/a-conjecture-on-why-reliable-systems-fail/). They made a change to the system in order to fix an ongoing problem, and due to a set of circumstances, things got worse rather than better. + +A defect in our blast radius analysis system incorrectly expanded the scope of the repair event to include all optical devices egressing a specific datacenter. + + +The second contributor mentioned was a (presumably) latent defect in their system. Note the irony of the failure mode here: I suspect this blast radius analysis system usually contributes to reliability, but in this case it hurt reliability by increasing the blast radius. + +The safety validation step, which is designed to confirm that at least one of the two redundant datacenter paths remains available, ran but incorrectly concluded the operation was safe. + + +The third contributor mentioned was a safety check (good!) that passed even though the action was unsafe (bad!). + +The checks validated each device individually rather than evaluating the aggregate effect of isolating all devices at once, a scenario that was not accounted for because the system was never designed to process a full datacenter’s worth of devices in a single request. + + +The reason it failed was due to an interaction with the second contributor: the blast radius being all of the optical devices egressing the datacenter. The designers never envisioned that the check would have to handle the sort of scenario that occurred as a result of the blast radius analysis system defect. + +As a result, routes were withdrawn from multiple devices simultaneously, disrupting connectivity between the datacenter and the WAN – therefore impacting traffic entering or leaving the West US region. + + +It sounds like this change effectively disconnected the West US datacenter from the internet. + +Once the route withdrawals took effect at 14:44 UTC, physical links and routing adjacencies continued to appear healthy, which initially masked the correlation between the break-fix activity and the connectivity disruption + + +Here we have our fourth contributor: the operators were receiving misleading signals from the system. The links and routes looked healthy, even though connectivity was broken. + +The impact presented as a WAN routing anomaly, as third-party networks could not reach Azure in the region, rather than as a datacenter connectivity failure. + + +Our fifth contributor is another flavor of misleading signals. The symptoms presented as a routing issue between Azure and third-parties. + +Although all physical work in the region was stopped, our engineers could not correlate to this recent change because the preparation activities in advance of the break-fix did not succeed, so the physical layer and traffic appeared healthy. + + +This is the sixth contributor mentioned in the writeup. The writing is a little oblique here, but I think what they are saying is that the repair event did not show up in their event log because the repair event didn’t actually complete. It sounds like the preparation activities were the ones that triggered the incident. But, because the repair event didn’t actually happen, the operators looking for events that correlate in time with the onset of the incident didn’t see the triggering event because it didn’t show up in the log of events. That’s my best guess, anyways. + +Our automated recovery and rollback system detected the device failures, and attempted multiple retries to restore the affected devices. However, because that system depended on the same datacenter connectivity that had been disrupted, its automated rollback attempts were unsuccessful. + + +This is the seventh and final contributor mentioned. Azure has an automated recovery and rollback system (good!), but the failure mode in this case prevented automated rollback from succeeding (bad!). + +As always, I’d love to know more about how the operators identified what the failure mode actually was, and how they traced it back to the optical device repair work. + +## One thought on “Quick thoughts on Azure Regional Outage from July 23, ’26” diff --git a/sreweekly/markdown/533/03-why-distributed-databases-fail-at-coordination-boundaries.md b/sreweekly/markdown/533/03-why-distributed-databases-fail-at-coordination-boundaries.md index f9c50d98..ca47c81d 100644 --- a/sreweekly/markdown/533/03-why-distributed-databases-fail-at-coordination-boundaries.md +++ b/sreweekly/markdown/533/03-why-distributed-databases-fail-at-coordination-boundaries.md @@ -7,3 +7,142 @@ ## 简介 > Distributed databases rarely fail in the clean, isolated ways described by component diagrams. They fail through timing gaps, stale metadata, ambiguous ownership, retry storms, incompatible health decisions, and overlapping maintenance activity. + +## 正文 + +- + ![](https://dz2cdn1.dzone.com/themes/dz20/images/dz-postarticle.svg) [Post an Article](https://dzone.com/content/article/post.html) +- + [Manage My Drafts](https://dzone.com) + +# Why Distributed Databases Fail at Coordination Boundaries + +Failures in distributed systems emerge at interfaces where independent components exchange timing, ownership, and state information. + +Join the DZone community and get the full member experience. + +[Join For Free](https://dzone.com/static/registration.html) + +Distributed databases are often evaluated through familiar technical dimensions: replication factor, consistency model, partitioning strategy, throughput, latency, and recovery time. These characteristics matter, but they do not fully explain why systems that appear healthy at the component level still experience severe production failures. + +In many cases, the storage engine is not the weakest part of the architecture. The failure occurs at a coordination boundary. + +A coordination boundary is any point where independently operating components must agree on timing, ownership, ordering, configuration, or state. These boundaries appear between replicas, partitions, control planes, data planes, load balancers, clients, metadata services, and background maintenance processes. Each component may behave correctly according to its local rules while the overall system produces an incorrect or unstable result. + +This is why [distributed database](https://dzone.com/articles/what-is-a-distributed-database) incidents can be difficult to predict. The database may not fail because a server crashes or a disk becomes unavailable. It may fail because two healthy components temporarily disagree about who owns a partition, whether a node is available, or which version of configuration should be applied. + +## Local Correctness Does Not Guarantee System Correctness + +Engineers naturally reason about software components individually. A node accepts requests, writes data, replicates changes, responds to health checks, and reports metrics. If each of those behaviors appears correct, the system is assumed to be healthy. + +Distributed systems challenge that assumption. + +A replica can be healthy but delayed. A coordinator can be available but operating with stale metadata. A load balancer can route traffic correctly according to its current configuration while that configuration no longer reflects the database topology. A client can retry a failed request according to policy while unintentionally amplifying load during a partial outage. + +Each component is locally correct. Their interaction is not. + +Consider a partition ownership transition. One node is being removed, replaced, or scaled down, and another node is taking responsibility for the affected data range. The outgoing node may believe it still owns the partition because it has not received the latest control-plane update. The incoming node may already begin accepting requests because it has received a newer version of the assignment. + +For a brief period, both nodes may behave correctly according to the information available to them. The system, however, has entered an ambiguous ownership state. + +That ambiguity can lead to duplicate processing, inconsistent writes, rejected requests, or unexpected latency. The problem does not exist entirely inside either node. It exists at the boundary where ownership information is exchanged and interpreted. + +## Time Is Often the Hidden Coordination Dependency + +Many distributed database designs avoid relying on perfectly synchronized clocks. Even so, time remains embedded throughout the system. + +Timeouts determine when a request is considered failed. Leases determine how long a node retains authority. Heartbeats influence failure detection. Retry intervals shape traffic behavior. Expiration policies determine when data should disappear. Background processes decide when to compact, replicate, repair, or rebalance information. + +These mechanisms create coordination dependencies even when the architecture does not explicitly describe them that way. + +For example, a client sends a write request and does not receive a response before its timeout. The client cannot immediately know whether the write failed, succeeded, or is still being processed. It retries the request through another route. + +If the database supports idempotent request handling, the retry may be safe. If it does not, the same logical operation may be applied twice. The first server and the client both followed their expected behavior. The uncertainty appeared between them because completion and acknowledgment were separated by a network boundary. + +This is a common distributed systems pattern. A timeout provides information about waiting, not about the final outcome of an operation. + +Cloud architects should therefore treat every timeout as an ambiguity boundary. Timeout behavior must be designed together with idempotency, deduplication, retry limits, load shedding, and observability. Configuring a timeout without defining the system’s response to uncertainty simply moves the failure elsewhere. + +## Metadata Can Become More Critical Than Data + +[Database reliability](https://dzone.com/articles/stop-being-afraid-of-databases) discussions frequently focus on protecting stored records. Replication, backups, checksums, and repair mechanisms are designed to preserve data durability. + +However, the metadata that describes how data should be accessed can be just as important. + +Partition maps, routing tables, node membership, schema versions, configuration states, and feature capabilities determine how requests travel through the system. If this metadata becomes stale or inconsistent, the underlying data may remain fully intact while applications lose the ability to access it reliably. + +This is particularly important in systems that separate the control plane from the data plane. The control plane decides how infrastructure should be configured. The data plane processes live requests using that configuration. + +Separating these responsibilities improves scalability and operational isolation, but it introduces another coordination boundary. Configuration changes must move safely from the control plane to every affected data-plane component. During that transition, the system may contain multiple valid configuration versions at once. + +The engineering question is not merely whether a configuration update can be delivered. It is whether old and new versions can coexist without violating system correctness. + +Safe configuration rollout often requires versioning, backward compatibility, staged activation, and explicit rollback behavior. Without those protections, a harmless-looking control-plane update can produce a data-plane outage even when no database node has failed. + +## Load Balancing Can Amplify Database Instability + +[Load balancing](https://dzone.com/articles/mastering-load-balancers-optimizing-traffic-for-hi) is sometimes treated as an infrastructure layer outside the database itself. In practice, routing behavior directly influences distributed database reliability. + +When a node slows down, a load balancer may reduce traffic to it. That appears beneficial, but the remaining traffic must go somewhere. Healthy nodes receive additional load, their latency increases, and health checks may begin failing. The load balancer then removes more nodes, increasing pressure on the smaller remaining pool. + +This creates a feedback loop. + +The database causes routing changes, and the routing changes make the database less stable. Neither system is necessarily defective. The failure emerges from their interaction. + +Aggressive health checks, short timeout thresholds, synchronized retries, and immediate node removal can turn a minor performance issue into a broad outage. A more resilient design considers the rate of change, not only the current health signal. + +Cloud architects should ask whether routing decisions become less reliable during overload. They should also examine whether the database and load-balancing layers use compatible definitions of health. A node capable of serving read traffic may be temporarily unsuitable for writes. A node completing recovery may be reachable but not ready for production load. + +Binary healthy-or-unhealthy classifications often hide these operational differences. + +## Background Work Creates Coordination Pressure + +Distributed databases perform significant work outside the direct request path. Replication, compaction, repair, rebalancing, expiration, backup, and cleanup processes compete for shared resources. + +These operations are often independently scheduled, which creates additional coordination boundaries. A compaction process may increase disk activity while a rebalance consumes network bandwidth. A repair job may begin during a traffic peak. Expired records may accumulate faster than cleanup processes can remove them. + +Each mechanism may operate within its configured limits, yet their combined effect can overwhelm the system. + +Time-to-live functionality provides a useful example. Expiring a record appears to be a simple data operation, but at scale it affects storage layout, indexing, replication, read behavior, and cleanup scheduling. The system must determine when an item is logically expired, when it should stop appearing in reads, and when its physical storage can be reclaimed. + +Those events may not occur simultaneously. + +If expiration processing is poorly coordinated, large groups of records can become eligible for deletion at the same time, creating bursts of background work. The feature itself works correctly, but the interaction between expiration timing and resource consumption can destabilize the database. + +The broader lesson is that operational features should be evaluated as distributed workflows, not isolated functions. + +## Designing for Boundary Failures + +The most effective way to improve distributed database reliability is to identify coordination boundaries during architecture design. + +For every boundary, engineers should define what information crosses it, how that information is versioned, how long it remains valid, and what happens when delivery is delayed or duplicated. They should also determine whether the receiving component can safely operate with stale information. + +Observability should follow the same structure. Monitoring individual nodes is necessary, but it is not sufficient. Teams need visibility into ownership transitions, metadata propagation delays, retry amplification, routing changes, replication lag, and background-work queues. + +These signals reveal disagreement between components before that disagreement becomes a complete outage. + +Testing must also include transitional states. Steady-state benchmarks show how a system performs when ownership, routing, and configuration are stable. Production failures frequently occur while those conditions are changing. + +Architects should test node replacement, delayed configuration propagation, partial network loss, rolling upgrades, uneven clock behavior, repeated retries, overloaded background workers, and conflicting health signals. These scenarios expose the boundaries where local assumptions stop matching global reality. + +## Reliability Lives Between Components + +Distributed databases rarely fail in the clean, isolated ways described by component diagrams. They fail through timing gaps, stale metadata, ambiguous ownership, retry storms, incompatible health decisions, and overlapping maintenance activity. + +The database node that appears responsible may only be the place where the problem becomes visible. + +For cloud architects and engineers, the practical shift is to stop treating coordination as an implementation detail. Coordination is part of the system’s correctness model. + +Storage engines protect data. Replication protects availability. Load balancing distributes work. Control planes manage change. None of these mechanisms can provide reliability independently. + +Reliability emerges from how they coordinate, especially when information is delayed, incomplete, duplicated, or temporarily inconsistent. + +That is where distributed databases are most likely to fail, and where architects should focus first. + + Database + Load balancing (computing) + + + Opinions expressed by DZone contributors are their own. + +Comments diff --git a/sreweekly/markdown/533/04-the-record-says.md b/sreweekly/markdown/533/04-the-record-says.md index 5be3b263..6c781a99 100644 --- a/sreweekly/markdown/533/04-the-record-says.md +++ b/sreweekly/markdown/533/04-the-record-says.md @@ -13,3 +13,127 @@ I love this concept of a “political incident”: And ouch, I felt this bit: > You have spent forty minutes of the incident on the severity field. + +## 正文 + +![](https://substackcdn.com/image/fetch/$s_!bDmW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe9282a63-57a1-4f70-8294-d2aa867ac79d_1232x928.png) + +The hands went up before I finished. + +Not straight away. There was a gap of about four seconds between the sentence and the first one, which is roughly how long it takes to decide you disagree with someone with enough standing that disagreeing carries a cost. Then a second. Then three more, stacked in the participant list, waiting. + +I was on a call with a room full of incident commanders. The subject was political incidents, by which I mean the ones where the severity arrives before the impact assessment does. What I said was this: in a political incident it is easier not to fight it. Take the escalation. Run the room like you didn’t. Pay lip service to the executive right up until the point they leave the bridge. + +I have said less popular things in that forum. I have not said one that produced a quieter four seconds. + +They waited. That is the part worth recording. Nobody interrupted, nobody typed in the chat, and the hands stayed up while I finished a spiel that had another two minutes in it. This is a room of people trained to hold position under pressure and let the commander finish. They applied the training to me. + +Then it came back, and it was not tactics. It was principle. Push back. Hold the line. Defend the process. The severity matrix exists for a reason, and the reason is that somebody has to be willing to say no to a senior person while the phone is ringing. + +What I had not expected was who it came from. Not the new ones. + +They said it the way you say something about a thing being taken from you. + +I thought they had it precisely wrong. I still do. + +![](https://substackcdn.com/image/fetch/$s_!8jy2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc26ccff2-9a6d-4f8d-91b0-4d6714f9a6bd_1232x928.png) + +Here is what I had done before I stopped doing it. + +I fought them. Every time. I had the matrix, and the matrix was clear, and I could show you the row and the column and the impact definition that put the thing at a three. This is not difficult work. Anybody can read a table. What I did not understand for a long time is that nobody I was arguing with was reading it, or had ever read it, or had at any point agreed to be bound by it. + +So it goes to the manager. The manager takes it up. It comes back down a different pipe entirely and it arrives mid-incident, from a skip two levels removed, and it is four words long. Just do it, please. + +You are on a bridge with fourteen people and an unhealthy service. You have spent forty minutes of the incident on the severity field. And you are going to lose the forty-first as well, because the answer was always going to be yes, and the only thing your forty minutes bought was a record of you being difficult about it. + +Then the hallway. Days later, no meeting invite, no thread. Someone stops you near the kitchen and says you should just do what is asked. Friendly. Actually friendly, which is the part that stays with you. They are not delivering a reprimand, they are doing you a favour, passing on something said about you in a room you were not in. + +It arrives in myriad channels, that irritation. Never the one you fought in. + +I used to read this as a failure of the framework. It is not that. + +They were never working to our severity matrix. They had a pressing need and an escalation of their own and a matter to deal with, and they had learned that the smaller the number the faster their matter moved. That is not a misunderstanding of the process. That is a correct reading of it. + +It is the cost of running a good incident command practice. When urgency is high, people look for the fastest lever in the building. You built it. + +![](https://substackcdn.com/image/fetch/$s_!KMdU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F01fb8d0f-b92c-4187-94a1-bfad5fa5404a_1232x928.png) + +I get paged into a bridge already in progress. This happens. What happens less often is that when I arrive there is someone three levels above me on the call, issuing instructions. + +So something has occurred. I do not yet know what. + +I run the intros anyway, because the ritual is load-bearing and because it buys me ninety seconds of looking at people. Commander on the call. Where are we. What is the sitrep. Who has the deployment. + +The engineers answer. They answer correctly, in order, with the right level of detail. And they are sullen in a way I recognise before I can account for it. Not tired. Not stuck. There is a particular flatness that people carry after they have been told off for a decision that was right, and they know it was right, and they have worked out that saying so a second time will cost them more than the first time did. + +I do not ask what happened before I joined. I have the answer. + +So I smile. I take the plan from the team, and I read it back, and I confirm the deployment window and the rollback trigger. Then I turn to the person three levels up and I ask whether there is anything else. + +There is not. Command is visible, the number is correct, the machine is running. They disembark. + +I count four seconds. + +Then I say the actual thing, which is this. The deployment takes three hours. Plan B is staged. The alerts are rigged and they will fire, and until they fire there is nothing here for any of you to do. + +I ask whether we need to be on this call while it runs. + +They look at me like I have taken the floor out. Twenty minutes ago they were reprimanded for running a three as a three, and now the incident commander is standing them down from three hours of watching a number increment on a dashboard nobody is going to act on. + +We push what matters into the channels. We adjourn. Three and a half hours later it resolves. + +The executive is satisfied. The team is intact. Nobody has had a conversation about the severity matrix, because there was nothing to have a conversation about. The severity is a two. It says so in the record. + +I did not run a two. I ran a three with a two written on it, and I sent six people home. + +![](https://substackcdn.com/image/fetch/$s_!4GAq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7522fd11-8c80-4221-b3c6-f621087bbd90_1232x928.png) + +Here is what I told the room of incident commanders, and here is where it stops being true. + +I told them the loop closes. Absorb the escalation, run the incident properly, and the correction arrives later, in the review, where everybody has hindsight and nobody has adrenaline. The severity gets named as inflated. The record gets amended. It costs nothing because by then it costs nothing. + +I have sat in a great many of those reviews. I have watched that correction arrive perhaps a handful of times. + +What happens instead is that I read a post-incident review for a severity two and the review is not about a severity two. The impact section does not describe an impact. The timeline has a forty minute gap at the start that nobody explains. The whole document has the shape of a thing written around an object rather than about it. + +So I ask why it was escalated. + +Nobody answers. I ask again, differently. Somebody offers a technical reason that is not a reason. I ask a third time, and this is the part I want on the record, because three times is not a rhetorical device, it is the actual number, and by the third one everybody in the room understands that I am not going to stop. + +Then someone says it. Meekly, or with a flash of irritation at having to be the one. A head of engineering escalated this. An executive did not have a report on time. + +And the room changes. Not to relief. To something closer to embarrassment, the collective adjustment of a group of adults who have all been carefully not saying the same thing and have just found out that everyone else was doing it too. + +Unless you were on the bridge, it is not information. It is folklore. It is passed in hallways and hushed the way you hush a name you do not want to summon, and it never once travels far enough to reach a document. + +That is not a review process failing at its job. A review process that cannot write down who escalated something is telling you what it is for. It is for reviewing engineers. + +The absorption works. I have never doubted the absorption. It was the other half I promised them, the half that closes the loop, that I had less evidence for than I sounded like I did. + +Absorption without correction is not neutral. It is a subsidy. I gave that executive a resolved incident, a satisfied bridge, no friction and no consequence, and a review that could not write down what they had done. The next one arrives sooner. Absorb and correct, or do not absorb. + +![](https://substackcdn.com/image/fetch/$s_!NlIf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F719e2f62-be96-4fbb-bcaf-d5d80b512157_1232x928.png) + +So I ask the room whether we can put it in the review. + +They say yes. They mean it at the time, or they mean it enough. Sometimes it appears. More often the document goes out unchanged, and the cycle starts again some weeks later with different people and the identical shape. + +Smile and nod. Agree in the room, decline in the artefact. + +I know that move. I taught it in a forum once and five hands went up. + +What goes in instead is the actions. There are always actions, because a review without actions is a review that failed, and everybody in the room understands the assignment. So they write them. Deliver the next phase of this project faster. Improve the reliability of the report generation. Owners against each one. Due dates. Engineering time apportioned in a planning session six weeks later by people who were not on the bridge and will never be told why the work exists. + +This is administrative violence. The three hour call was not. A three hour call costs an afternoon. This costs a quarter, and it is the only account of the event that will exist in five years. + +The whole thing is measurable, and the measurement is simple. Ask whether the review can record who escalated it and why. Do not argue for it. Just ask, and then watch what the answer is. + +A review process with an aversion to authority will never work. The answer to that one question tells you exactly how safe the people in that building actually feel. Not the survey. Not the values on the wall. Whether a document can contain the sentence “this was escalated by a head of engineering because a report was late,” and survive. + +I stood in front of a room of commanders and they told me I was giving something away. They said it the way you say a thing is being taken from you. + +They were right that something was being taken. Nobody has come for the severity matrix. + +What gets taken is the record. Every absorbed escalation produces a document that is accurate about the outage and silent about the cause. + +And the incident is a two. The record says so. It will say so forever. diff --git a/sreweekly/markdown/533/05-what-sres-should-automate-and-never-automate-with-ai.md b/sreweekly/markdown/533/05-what-sres-should-automate-and-never-automate-with-ai.md index cfe35bf6..4b9fade2 100644 --- a/sreweekly/markdown/533/05-what-sres-should-automate-and-never-automate-with-ai.md +++ b/sreweekly/markdown/533/05-what-sres-should-automate-and-never-automate-with-ai.md @@ -7,3 +7,77 @@ ## 简介 Where can you safely use LLM agents, versus when you should keep things in human hands? This one has some good criteria to consider. + +## 正文 + +**Five key takeaways:** + +1. Automate based on impact and recoverability, not on whether the AI is technically capable of doing the task. +2. Alert triage, anomaly detection, incident summaries, capacity forecasting — these are the easy wins. Low risk, high value. +3. Production changes, incident command, security response, severity calls — keep a human's name on these. Always. +4. Reversibility and blast radius are better questions than "can the AI do this." +5. The goal isn't AI replacing engineers. It's AI clearing enough noise that engineers can actually think. + +I've sat through the version of this conversation that sounds like a vendor pitch — AI triages everything, drafts your runbooks, predicts outages before they happen, and nobody gets paged at 2 a.m. anymore. I've also watched the other version happen in real time: an automated remediation script restarts the wrong service, confidently, at 11 p.m., and a 20-minute blip turns into a four-hour outage while everyone tries to figure out why the "fix" made things worse. + +Both of those are real. AI is already inside SRE workflows whether or not anyone signed off on it — the question that actually matters is where it belongs, and where a human still needs to be the one holding the decision. + +None of what follows comes from a whitepaper. It's from watching what breaks when teams move too fast with this stuff, and what quietly gets better when they don't. + +## Reversibility and blast radius + +Here's the mental model I keep coming back to before automating anything: can you undo it, and how bad is it if you're wrong? + +Restarting a pod — reversible, low stakes. Deleting a database backup — not reversible, at all. Scaling a service up is easy to walk back. Silencing an alert for six hours is technically reversible too, except the six hours where something real happened and nobody saw it isn't something you get back. + +Blast radius is the other half of it, and it's not the same thing as severity. A misclassified low-priority alert costs a few wasted minutes. A misrouted sev-1 costs an hour of response time during an active outage, while the right team sits there not knowing they should be paged. And blast radius scales with what the action touches — one service versus a shared piece of infrastructure everything depends on, even when both look equally "minor" on paper. + +Anything with low reversibility and a wide blast radius shouldn't be running on autopilot. Anything reversible and contained is fair game. The stuff in between is where you actually need judgment — specifically, judgment from the people who'll be the ones on call when it goes sideways. + +Notice this framing never asks whether the AI *can* do something. It asks what happens if it's wrong. That's the more useful question, and it's the one most teams skip. + +## Where this actually works well + +**Alert noise.** This is the least controversial win there is. Somewhere between 30 and 60% of production alerts are noise by the time a human sees them — duplicates, transients, things that resolved themselves three minutes ago. AI grouping related alerts, suppressing known-flapping signals, correlating spikes with recent deploys — worst case, something gets mislabeled and a human still catches it. Low blast radius, fully reversible. This is exactly the profile you want. + +One catch: it only works well tuned to your environment, not a generic model. An alert that always fires right before a nightly batch job and clears itself a minute later is trivial to suppress — but only if the model actually knows about your batch schedule. Skip that step and you've just added a second layer of noise on top of the first. + +**First drafts of runbooks and postmortems.** Runbook rot is one of the oldest problems in this field. The doc that was accurate in 2022 is a landmine now — nobody updates it, an incident hits, someone follows it anyway, and step four references a service that got decommissioned eight months ago. AI is genuinely good at pulling together a first draft from past incidents, change logs, whatever documentation exists. Same for postmortems — a draft that someone who actually lived through the incident reviews before it goes out saves real hours. + +**Forecasting and anomaly detection.** This is pattern matching, and models are good at pattern matching. A holiday traffic spike that happens once a year gives engineers almost no reps to build intuition about — but a model trained across several years of that same spike has plenty. The important part: keep this as a recommendation a human acts on, not something that auto-provisions infrastructure on its own. The moment it stops informing a decision and starts making one, the blast radius changes. + +**Narrow, well-understood auto-remediation.** This one comes with real caveats, but it earns its place. A specific service that needs a restart when it hits a known stuck state, a queue that needs draining past a defined threshold — fine, if the failure class is precisely defined, tested, and low-blast-radius by design. And there has to be a circuit breaker. If the fix doesn't work within a set window, it stops and escalates instead of retrying forever on a wrong diagnosis. Automation that keeps trying the same broken fix is worse than doing nothing. + +## Where it doesn't belong + +**Severity calls.** Get this wrong either direction and it costs you. A real sev-1 marked as low pulls in the wrong people at the wrong urgency while an SLA clock runs. A minor issue marked critical drags a response team into something that didn't need them at 3 a.m. AI can surface context and flag patterns worth escalating — but the actual call needs a name attached, someone accountable for it. "The model said it was low severity" doesn't hold up in a postmortem. + +**Production changes without sign-off.** Config changes, scaling decisions, anything touching a database directly, restarts outside that narrow bounded case above — a human authorizes these. AI can prep the change, check it against known-good patterns, even simulate the blast radius. What it shouldn't do is decide the moment is right and pull the trigger itself. + +**Security incidents.** Different risk shape entirely. Miss something real and an active compromise sits there while the system waits for more confirmation. False-positive and you've locked out legitimate engineers mid-response. AI correlating logs to surface signal fast — genuinely useful. Containment and escalation decisions — that needs someone who can weigh legal and business context a model was never trained on. + +**Root cause, as a stated fact.** AI narrowing the search space by correlating deploy timing with metric shifts is useful groundwork. But writing "root cause: X" in a postmortem is a claim that shapes what the org fixes next and what it decides to ignore. Get that wrong because a correlation looked convincing, and the actual bug ships again next quarter. + +**Who to escalate to.** This is context a model just doesn't have — who's already underwater tonight, what else is on fire across the org, whether the responding engineer's confidence is real or performed. Escalation is a trust call as much as a technical one. + +## The thing nobody's measuring + +There's a slower cost that never shows up in a single incident review: engineers stop building intuition when AI absorbs all the routine reps. The edge cases are exactly where judgment matters most — and they're exactly the cases you need practice on the boring stuff to be ready for. A team leaning hard on automation can look great for a long stretch, right up until something shows up that doesn't match anything the model — or the team — has seen before. + +This isn't an argument against automating things. It's an argument for being honest about which reps you're willing to give away. + +## A few practices worth adopting + +- Decide, as a team, which categories of action AI can take alone versus which need a sign-off — decide this before an incident forces the question at 2 a.m. +- Keep an actual human accountable for anything irreversible. Not nominally "in the loop" — actually reviewing before it executes. +- Build in a circuit breaker for anything automated. If it doesn't work within a defined window, it escalates instead of retrying. +- Rotate people through the routine cases sometimes, even when AI could handle it, so the skill doesn't quietly disappear. +- Revisit the boundary as systems change. A failure class that was well-understood six months ago might not be anymore after an architecture shift. + +Skip this and you end up with automation debt, eroded skills, and a production system nobody fully understands anymore — which is a worse place to be than where you started. + +## Where this leaves things + +It's not really a question of whether to use AI. It's whether you're using it somewhere judgment genuinely isn't needed, or somewhere it is and you've just decided waiting for a human is too slow. + +One of those is a real force multiplier. The other is a liability with a delay timer on it. diff --git a/sreweekly/markdown/533/06-20-the-ci-traffic-without-getting-slower-how-we-rebuilt-git-serving-at.md b/sreweekly/markdown/533/06-20-the-ci-traffic-without-getting-slower-how-we-rebuilt-git-serving-at.md index b1b6ff58..8f1e844b 100644 --- a/sreweekly/markdown/533/06-20-the-ci-traffic-without-getting-slower-how-we-rebuilt-git-serving-at.md +++ b/sreweekly/markdown/533/06-20-the-ci-traffic-without-getting-slower-how-we-rebuilt-git-serving-at.md @@ -7,3 +7,172 @@ ## 简介 I learned a lot about Git while reading this one. Speeding up Git clones in CI may not seem important, but it will when you’re trying to roll out a fix during an incident. + +## 正文 + +![Mike Thompson Mike Thompson](https://web-assets.dd-static.net/42588/1786393483-mike-thompson.jpeg?format=auto&fit=bounds&quality=75&disable=upscale&width=48&dpr=1) + +Mike Thompson + +Senior Staff Engineer + +![Daniel Esponda Daniel Esponda](https://web-assets.dd-static.net/42588/1786393538-daniel-esponda.png?format=auto&fit=bounds&quality=75&disable=upscale&width=48&dpr=1) + +Daniel Esponda + +Staff Engineer + +If you have ever watched a CI job sit on “Fetching repository …” while nothing seems to happen, you already know the unglamorous truth about continuous integration: Every job begins by getting the code, and getting the code is not free. + +At Datadog, CI fetches code millions of times a week across thousands of repositories. Our largest repositories are monorepos with years of history and hundreds of thousands of files. At that scale, `git clone` stops being a footnote and becomes a large contributor to CI run times. + +This is the story of **gitretriever**, the Git mirror we built to serve code to CI at Datadog scale. In its first 4 months, gitretriever served more than a billion Git requests and hundreds of terabytes of code. Today gitretriever handles more than 100 million requests each week. Despite the 20× traffic growth since launch, median latency has remained around 40 ms, while fetch-serving CPU on our previous Git backend has dropped by three to four times. + +## [Serving Git to CI, and why it gets hard](https://www.datadoghq.com#serving-git-to-ci-and-why-it-gets-hard) + +Datadog has a unique CI setup: GitHub serves as the authoritative code repository, while almost all of our internal CI workloads run on a self-hosted GitLab installation. CI fetches from GitLab’s Gitaly, fronted by Praefect (Gitaly Cluster’s routing and replication manager) and kept in sync with GitHub by an internal service (aptly named “codesync”). This hybrid architecture has carried us through more than a decade of growth. + +But CI load does not grow smoothly. The expanding use of AI coding agents has driven an order-of-magnitude increase in Git traffic, with agents hitting Git far harder and more often than even our most active contributors ever could. That traffic comes on top of the continually growing load from internal deployment, auditing, and security services. As that growth accelerated, the pressure hit hardest where our code is densest: our large monorepos. Operational load increased, CI run times grew, and multi-hour long CI outages became more frequent. It was clear we needed a more sustainable solution. + +## [Why the usual fixes don’t scale](https://www.datadoghq.com#why-the-usual-fixes-dont-scale)  + +We tried adding capacity, we tried increasing instance size, we tried placing different repositories on dedicated backends, and we tried optimizing build pipelines. Things would improve for a week or two, but then our CI infrastructure would inevitably end up degraded or outright down. So why didn’t any of the usual approaches work? + +A single fetch from a large monorepo can consume several seconds of server CPU. At peak, hundreds of jobs perform fetches at the same moment and land on the same handful of nodes. Adding capacity did little to reduce per-node CPU usage. In some cases, adding more nodes made the problem worse. + +Before committing to a new architecture, we had to figure out why none of our previous attempts at fixing the problem had worked: + +- **Scale the backend or add nodes** : In our replicated setup, every write had to be copied to every replica. Adding a node increased replication overhead instead of relieving it. +- **Put a content delivery network (CDN) or caching proxy in front** : The expensive part of a fetch isn’t a static byte range you can cache at the edge. It’s computation that’s specific to each client’s request. +- **Clone on demand from GitHub** : That simply moves the thundering herd upstream, where we run into server-side rate limits. + +The common thread was that we had been scaling the wrong axis. Read traffic scales with the number of CI jobs, but in our replicated architecture, write costs scale with the number of replicas. Every time we added replicas to handle more reads, we also increased replication overhead, and more CPU time went to maintaining the system instead of serving fetches. + +To understand why serving those fetches consumed so much CPU in the first place, it helps to look at what happens during a Git fetch. + +## [Why a Git fetch is expensive](https://www.datadoghq.com#why-a-git-fetch-is-expensive) + +To understand why our design works, it helps to understand how Git stores data and where a `git fetch` spends its time. + +### [Git data types](https://www.datadoghq.com#git-data-types) + +Git’s object database is primarily built around immutable objects. For the purposes of this post, we’ll focus on three: + +- **Blobs** , which store file contents +- **Trees** , which describe directory entries (for example, folders and blobs) +- **Commits** , which store metadata, a commit message, a reference to a tree, and references to parent commits + +Each object is identified by a hash of its type, size, and contents. SHA-1 remains the default object format, although Git also supports SHA-256 repositories. + +Finally, there are **references**, which are mutable names stored separately from objects. For example, `refs/heads/main` identifies the commit at the tip of the `main` branch. + +Objects may be stored on disk individually as **loose objects** or grouped into **packfiles**. Within a packfile, an object may be stored in full or as a delta against another object (known as **delta** **compression**), which allows Git to efficiently store the complete history of changes to files within a repository. Packfiles are immutable to allow for safe concurrent reads. + +![How Git references, commits, trees, blobs, and packfiles relate to one another. How Git references, commits, trees, blobs, and packfiles relate to one another.](https://web-assets.dd-static.net/42588/1786393825-f1-gitretriever.png?format=auto&fit=bounds&quality=75&disable=upscale&width=1400&dpr=1) + +Write operations (for example, `git push`) may introduce new packfiles. A background maintenance process periodically consolidates loose objects and smaller packfiles into new packfiles. Unreachable objects (for example, deleted files) are eventually removed after a certain threshold by being omitted during packfile consolidation. + +### [Git protocol v2](https://www.datadoghq.com#git-protocol-v2)  + +Now that we understand Git’s data types, we can briefly look at how the current (v2) Git protocol works. + +The Git client uses the `ls-refs` command to learn the current object IDs of references it cares about (for example, all branches). The client and server then begin a multi-round negotiation to determine which objects the server needs to send to the client. You can read more about this negotiation process in the [Git protocol v2 documentation](https://git-scm.com/docs/gitprotocol-v2).  + +Once the client and server have determined which objects to send, the server creates a packfile containing those objects and sends it to the client. + +Constructing the response packfile can be CPU and I/O-intensive. The server locates objects within packfiles by using an index that Git maintains for each packfile. Some objects can be copied as is into the response packfile, while others must be decompressed and recompressed using delta compression. Under a sufficiently large number of concurrent fetches, this packfile construction work can saturate server CPU and storage capacity. + +Client behavior, such as requesting weeks’ worth of changes to a large monorepo, can make this more expensive in both CPU and I/O operations. Git attempts to reduce this cost with reachability bitmaps, sparse traversal, multi-pack indexes, and pack reuse. We tried all of these options, but client behavior and the rate at which our monorepos changed still concentrated CPU load on a small number of servers. + +The final step of a fetch or pull from a Git server is for the client to read the received packfile and update its local index of available objects. This requires only a small amount of client-side CPU. + +## [Our approach: Many independent mirrors, kept fresh](https://www.datadoghq.com#our-approach-many-independent-mirrors-kept-fresh) + +If the problem is CPU concentrated on a few contended nodes, the solution is to stop concentrating it. + +Gitretriever runs independent pods, each of which maintains a fresh local copy of the repositories it serves without waiting for every node to reach consistency. Each pod serves its local copy directly, with no consensus and no multi-writer replication between peers. Gitretriever pods have two roles, as shown in the following diagram: + +- **Mirrors** stay in sync with GitHub. We deliberately keep this fleet small because its job is to be a good GitHub client: a handful of well-behaved pollers rather than thousands of them. +- **Relays** fan out reads to CI jobs. This fleet is larger and autoscaled based on CPU and network load, allowing us to provision enough read capacity to meet demand without turning that growth into additional load on GitHub. + +![Git traffic flowing through mirrors and relays between GitHub and CI workloads. Git traffic flowing through mirrors and relays between GitHub and CI workloads.](https://web-assets.dd-static.net/42588/1786393883-f2-gitretriever.png?format=auto&fit=bounds&quality=75&disable=upscale&width=1400&dpr=1) + +## [**Staying fresh and reducing CPU usage**](https://www.datadoghq.com#staying-fresh-and-reducing-cpu-usage) + +**Staying fresh and reducing CPU usage** + +The architecture works only if every mirror and relay stays close to the latest changes without recreating the CPU bottlenecks we were trying to eliminate. We designed gitretriever around three principles that keep repositories fresh while minimizing repeated work. + +### [Distribute Git pulls across branches](https://www.datadoghq.com#distribute-git-pulls-across-branches) + +Gitretriever mirrors continually poll the upstream in a tight loop for changes. Gitretriever performs a parallel fetch for each reference it detects as changed since the previous synchronization loop iteration. No single request concentrates an expensive delta compression job on GitHub, and each small pack requires far less indexing CPU than one monolithic monorepo pack. Staying close to the tip of each branch also means that, in any given synchronization loop iteration, only a small number of branches have changed, reducing the number of packfiles we need to fetch. + +### [Spend the sync work once, then reuse it](https://www.datadoghq.com#spend-the-sync-work-once-then-reuse-it)  + +For the busiest repositories, one mirror cannot serve every client, so changes fan out to a fleet of relays. Relays can connect to mirrors or to other relays. Each relay splits its upstream connection into two channels: + +- **A signaling gRPC stream** : Announces that a pack is ready, propagates reference updates, and communicates mirror and relay topology changes +- **A plain HTTP endpoint** : Serves the pack bytes themselves + +Because Git objects are content-addressed, a relay installs the packfile it receives from its upstream mirror or relay without regenerating, re-indexing, or re-verifying it. It drops the packfile and its index into place, trusting the objects inside by the hashes that identify them. The work of pulling and indexing from GitHub happens once on the mirror, and every relay reuses that work instead of fetching again. As a result, the relay fleet can grow without adding load on GitHub while remaining within single-digit milliseconds of the tip. + +### [Never build the same pack twice](https://www.datadoghq.com#never-build-the-same-pack-twice) + +Gitretriever is both a Git client and a Git server. The current implementation uses Git’s default backend storage format: packfiles, reference tables, reachability bitmaps, and multi-pack indexes. That means gitretriever has to make serving other Git clients (such as CI jobs) as efficient as possible. + +A fresh push to a busy branch sets off a thundering herd of identical fetches. Gitretriever implements a **pack cache**, allowing it to reuse previously assembled packfiles for identical client requests. About half of all pack-building fetches are served directly from the cache, skipping the delta compression calculation on mirrors and relays entirely. Cache misses are still served locally by the mirrors and relays, so even a cache miss never becomes a trip to GitHub. + +Underneath these are smaller refinements, including a readiness check that understands Git state and keeps a pod out of rotation until its pack count is healthy, along with background repacking that keeps the packfile count under control while the pod continues serving. But the theme never changes: Take the CPU that used to pile up in one place and either spread it out or stop repeating it. + +Future iterations of gitretriever will build on the relay replication protocol to keep an always-up-to-date copy of our large repositories directly on CI nodes, allowing jobs to skip the initial `git clone` altogether. + +## [The bigger surprise: Many use cases don’t need a clone](https://www.datadoghq.com#the-bigger-surprise-many-use-cases-dont-need-a-clone) + +Once every repository had a fresh mirror, something in the traffic caught our eye: Most non-CI workloads don’t need a full repository clone. They wanted a single file at a commit, the SHA a branch pointed to, the list of files that changed, or the merge base of two refs. Cloning an entire repository to answer one of those questions was enormous overkill, yet our internal services, developer tools, and AI agents were doing it constantly. + +So we added a small, read-only HTTP API for exactly those queries. Resolving a ref or reading a file takes single-digit to tens of milliseconds. By comparison, a shallow clone of a large monorepo takes on the order of 75 seconds and keeps a CPU core busy for most of that time. Moving these use cases to the API reduces latency and removes load from the entire system. + +The non-CI workloads changed how we think about gitretriever. It’s less a faster Git server and more the query layer for Git across our engineering systems. + +This is the direction the platform is heading. As workflows become more automated and more AI agents ask questions about code, the cheapest and fastest answer is often another API rather than handing out a repository clone. + +## [Rolling out gitretriever safely](https://www.datadoghq.com#rolling-out-gitretriever-safely) + +Rolling out gitretriever required careful planning. Our CI infrastructure is used by every engineer at Datadog, so one wrong move could bring engineering to a halt. We used feature flags and built in automatic fallback to the old backend into our CI jobs, so if a mirror became unreachable or a fetch failed, the job fell back to the previous path. The worst-case outcome was no worse than before. We then migrated one repository group at a time, starting with the largest monorepo, while watching the old backend’s CPU graph. + +When that first monorepo cut over, we saw an immediate step decrease in CPU usage. That confirmed our understanding of the problem: Gitretriever was absorbing the heaviest, most CPU-dense fetches first. Those were the same ones that had been degrading developer experience and driving outages. + +The metrics matched our expectations: + +- **Synchronization time dropped from several seconds to a few hundred milliseconds** , making continuous, coordination-free mirroring possible. +- To date, gitretriever has served **more than a billion Git requests and hundreds of terabytes of data** across roughly**5,500 repositories** , and now handles**more than 100 million requests each week** . +- **Traffic grew about 20× in 4 months while median serve latency remained around 40 ms** (Figure 3). The system became an order of magnitude busier without getting materially slower. +- The result we care about most: Moving CI fetch traffic to gitretriever **reduced the old backend’s fetch-serving CPU by three to four times, even as overall CI activity kept climbing** (Figure 4). Its memory footprint dropped in step, which later let us right-size that backend down. The old backend still handles some use cases that gitretriever**doesn’t** yet support (e.g., rendering the GitLab UI), so we**don’t** claim we replaced it (yet). But the fetch-path load it had been drowning under is gone. + +![Traffic rising substantially from March to July while median latency remains nearly flat. Traffic rising substantially from March to July while median latency remains nearly flat.](https://web-assets.dd-static.net/42588/1786543743-f3-gitretriever.png?format=auto&fit=bounds&quality=75&disable=upscale&width=1400&dpr=1) + +![Fetch-serving CPU dropping sharply during the rollout and remaining substantially lower. Fetch-serving CPU dropping sharply during the rollout and remaining substantially lower.](https://web-assets.dd-static.net/42588/1786543827-f4-gitretriever.png?format=auto&fit=bounds&quality=75&disable=upscale&width=1400&dpr=1) + +## [How we built it: Two engineers, Claude Code, and design doc in nearly every folder](https://www.datadoghq.com#how-we-built-it-two-engineers-claude-code-and-design-doc-in-nearly-every-folder) + +We chose to use Claude Code on this project to accelerate development and to explore how far AI could responsibly assist with building production infrastructure. What made an AI collaborator trustworthy on a system this central wasn’t the model; it was the discipline around how we used it. + +We planned before we wrote code, designing each change and iterating on the design through several rounds before committing a line of code. We validated every change with integration tests backed by real metrics and logs, not just unit tests, so the bar for “done” was observed behavior rather than a green checkmark. To keep both the AI and ourselves aligned across a dozen packages, we maintained a living design document in nearly every directory, describing its architecture, data flow, concurrency model, and configuration, and updating it alongside the code. + +Those documents ended up serving two purposes. During development, they kept AI-generated changes aligned with the architecture. When ownership of the service transitioned to the team that now maintains it, the same documents became the handoff. + +The lesson we would pass on is that the design documents became the interface between the engineers, the AI, and the next team. Ultimately, the quality of your tests and telemetry data sets the ceiling on how far you can trust an AI collaborator. + +## [What’s next](https://www.datadoghq.com#whats-next) + +Gitretriever is not finished. We’re expanding the query API so more workloads can skip cloning entirely, allowing us to fully decommission our old Git backend. We’re also continuing the rollout across the rest of our repositories and building for a future where automated and agent-driven workflows ask even more of Git. + +A few ideas we’ll carry into whatever comes next: + +- **Make it disposable so you do not have to make it durable.** Some of the hardest parts became much simpler once we made them rebuildable instead of authoritative. +- **Content addressing lets you trust data by name.** That’s what makes coordination-free replication safe. +- **The fastest fetch is the one that transfers nothing** , whether that’s a fast-path ref update or an API call that answers the real question without a clone. + +More than any single optimization, gitretriever reflects how we approach engineering at Datadog: Push a good system as far as it will go, then, when the scale curve demands it, design the next generation from a better understanding of the problem, validate it against real telemetry data, and write down what you learned so the next team can build on it. + +If this sounds like your kind of problem, we would love to work with you. Take a look at our [open roles](https://careers.datadoghq.com/all-jobs/?s=Infrastructure&child_department_Engineering%5B0%5D=Backend). diff --git a/sreweekly/markdown/533/07-a-tale-of-two-flink-autoscalers.md b/sreweekly/markdown/533/07-a-tale-of-two-flink-autoscalers.md index 155f17f4..139aa7f0 100644 --- a/sreweekly/markdown/533/07-a-tale-of-two-flink-autoscalers.md +++ b/sreweekly/markdown/533/07-a-tale-of-two-flink-autoscalers.md @@ -7,3 +7,7 @@ ## 简介 Switching from their custom-written autoscaler to the new off-the-shelf option made sense, but it wasn’t a simple drop-in replacement. + +## 正文 + +> ⚠️ 抓取失败:HTTP 403 diff --git a/sreweekly/markdown/533/08-there-is-more-to-code-review-than-automatable-detection.md b/sreweekly/markdown/533/08-there-is-more-to-code-review-than-automatable-detection.md index 9dc3a76d..8f8d6f6e 100644 --- a/sreweekly/markdown/533/08-there-is-more-to-code-review-than-automatable-detection.md +++ b/sreweekly/markdown/533/08-there-is-more-to-code-review-than-automatable-detection.md @@ -7,3 +7,89 @@ ## 简介 Can we replace human code review with LLM-based reviews? This article lays out what an LLM can’t replicate, and I’d argue that these are the pieces that matter most for reliability. + +## 正文 + +The abstract for article “[The End of Code Review: Coding Agents Supersede Human Inspection](https://arxiv.org/abs/2606.13175)” paints this picture for the reader… + +**Abstract** – Code review has been the primary quality gate in software development since Fagan formalised code inspection in 1976. For five decades, having a human examine and comment on a colleague’s changes before merge has been a cornerstone practice at organisations of every size. Coding agents are large language model (LLM)-based autonomous systems capable of reading, writing, testing, and repairing software. We argue that coding agents have crossed a threshold of capability at which traditional human code review is no longer a necessary component of a software quality pipeline. Our argument rests on two claims: every stated goal of code review can be served by agents at lower cost and higher throughput; the naive integration in which agents write code and humans remain the mandatory reviewers is a dead end because it neither provides meaningful assurance nor scales with AI-assisted throughput. + + +The article is structured well and quite straightforward for engineers who aren’t used to reading research articles very often. However, I do think the argument critically depends on a problematic framing: ***the substitution myth***. + +The author decomposes peer code review into four stated functions: defect detection, style enforcement, knowledge transfer, and awareness. It argues an agent can perform each one. The conclusion, of course, is that if an agent can execute each of those functions, then the agent has the capability to replace a human reviewer. + +I think this overlooks some important aspects of peer code review that *cannot* be reduced to a function: + +**A peer reviewer’s confusion** + +When an experienced engineer reads a diff and says “I don’t understand this.”, their confusion *is* the finding. It means the code is either too complex, the abstraction is wrong, or the intent is not clear. An LLM will always ‘understand’ the code in the sense of being able to *process* it. It can’t give you the signal of legitimate human incomprehension. The article treats comprehensibility as something that is more about style than anything else. It’s not. It’s an emergent property and it shows up in the interaction between a person attempting to understand the artifact and the artifact itself. + +**Qualified skepticism about whether the change is even necessary** + +Questioning the existence of a change, like: + +“Should this actually be *two* PRs?” + + +or + +“This solves the symptom, not the problem” + + +These are questions about intent, scope, and appropriateness of the change. All of that comes *before* whether the code is “correct.” The article’s framing assumes that a) the code change being reviewed is necessary, and b) the main purpose of the review is verification. + +But anybody who has ever had contact with production understands that code review is *often the last* (or sometimes only) moment when someone can be expected to challenge whether the change is even necessary. + +**The ability to see what is *not* there** + +A human reviewer can notice that an API contract has changed but the error handling didn’t. They can notice what is *missing*. In other words: being able to recognize what is *expected* to be present, but isn’t. The article doesn’t acknowledge this at all, which is particularly interesting, given that [absence blindness](https://absencebench.github.io/) is exactly the class of failure that LLMs tend to be quite poor at. + +The agent reviews what is there; engineers with expertise can easily notice what’s missing. + +*Who* wrote the code influences the scrutiny of the review + +Peer code reviewers have a sort of *calibrated attention* that comes from past experience with the code’s author, who is often a colleague. For example: a less-tenured engineer’s first commit to, say, a payments module will likely get different attention than a veteran and ‘grey beard’ engineer’s routine refactoring. + +Reviewers typically match the situation’s who, what, when, and where to their own experience of where risk lies. + +The article seems to treat all diffs as equivalent inputs. + +**Code review is bidirectional and constructive** + +It seems to me that paper reduces knowledge transfer down to just information delivery; the agent simply ‘generates explanations.’ But discussion in a code review is a *joint* *cognitive activity*. The peer reviewer learns about the author’s approach, the author learns via the reviewers’ questions, and the result is a shared understanding that neither party had prior to the discussion. + +This is **coactive** work, not simply a transmission. An agent’s summary isn’t a substitute for a conversation that changes both participants’ mental models. + +**Operational context that lives outside repos** + +“We just had an incident in this service last Tuesday.” + +“The team that owns this downstream consumer is about to deprecate that interface.” + +“Legal told us not to log this field anymore.” + + +Human reviewers possess so much more contextual knowledge than they’re aware of, even though they can recognize connections in the wild. People understand the current state of the organization, recent events, and informal agreements that aren’t captured in tests, docs or version control, and they can recognize how these may influence the code under review. This happens so often that it’s all but invisible. + +The article assumes the codebase *is* the complete context. It never is. + +**Accountability for the code isn’t just a beuraucratic formality** + +The article treats human responsibility as a compliance artifact, a “named human” for legal or other rule-related purposes. But being aware that you are personally responsible for approving a change shapes how you review it. It is the “skin in the game.” An agent that “signs off” on a pull request bears no consequences and certainly has no incentive structure that fuels an earnest evaluation. While the paper does include ethics concerns in its discussion section, it ends up redirecting it to “requirements engineering and post-deployment monitoring” which seems to me as hand-waving way of kicking the can down the road. + +The most fundamental issue I have with the article is that it assumes code review is a first and foremost a **detection** process: you find defects, style violations, security issues, etc., and the assumption is that detecting these faster and cheaper is universally better.  + +But code review is also a *coordination* process, a *sensemaking* process, and a *governance* process.  + +The ***substitution myth*** often plays out in this same way: + +1. First, decompose the human contribution of work into measurable functions. +2. Show that the machine can replicate this human contribution into measurable functions of its own. +3. Declare the human redundant. + +This approach often falls apart at the same point: the human contribution that mattered most was the *integration* across functions. People’s ability to adapt to unplanned circumstances and contexts and serve the social accountability expected. + +This ability to adapt in those situations aren’t accounted for in the original decomposition step #1, above. + +I don’t think they were accounted for in the original article, either. diff --git a/sreweekly/pages/528.html b/sreweekly/pages/528.html new file mode 100644 index 00000000..67cbae5d --- /dev/null +++ b/sreweekly/pages/528.html @@ -0,0 +1,442 @@ + + + + + +SRE Weekly Issue #528 – SRE WEEKLY + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +
        + + + + + + + + +
        + +
        + + + + + + +
        + + +
        +

        SRE Weekly Issue #528

        +
        + + + +
        + + +

        + +
        +

        A message from our sponsor, Planetscale:

        +

        Most database incidents start with one expensive query, not the database being down. PlanetScale gives SRE teams high-availability Postgres and MySQL with automated failover, query insights, and Database Traffic Control to stop runaway queries before they page you.

        +

        → Explore PlanetScale

        +
        + + +
        +
        + +
        +

        Spotify has had some difficulty around podcast publishing, and they shared this analysis of the worst incident.

        +

          Jim Whitehead, Ulrik Mikaelsson, John Lagomarsino, and Saunak Jai Chakrabarti — Spotify

        +
        +
        + + + +
        + +
        +

        …and here’s where it gets interesting. This post shares the user point of view on the Spotify issues, including fact-checking their published timeline.

        +

          Gergely Orosz — The Pragmatic Engineer

        +
        +
        + + + +
        + +
        +

        New incident role unlocked: the incident tech lead. I enjoyed the description of the interplay between the tech lead and the incident commander.

        +

          Brent Chapman

        +
        +
        + + + +
        + +
        +

        This one goes hard: if you try to reduce your incident count, your system will become less reliable, not more. Aim for more incidents, handled well.

        +

          Tim Irving

        +
        +
        + + + +
        + +
        +

        There’s some brutal honesty in here that I find refreshing, especially around the impact on incidents and incident response.

        +

          Liz Fong-Jones — Honeycomb

        +
        +
        + + + +
        + +
        +
        +

        Here’s what I’ve learned about what actually happens inside a team when something breaks, how teams often reflect, and how we think about accountability without losing a blameless culture.

        +
        +

          Karan Nagarajowda — Uptime Labs

        +
        +
        + + + +
        + +
        +
        +

        agent sprawl is now a production reliability crisis, and the SRE discipline does not yet have governance frameworks for it.

        +
        +

           Ajay Devineni — DZone

        +
        +
        + + + +
        + +
        +

        The premise: read replicas can help you scale read load, but they introduce complexity. The article goes into the problems they ran into and how they dealt with them.

        +

          Johanna Larsson — incident.io

        +
        +
        +
        + + + + +
        + +
        + + + +
        + + +
        + + + + + + + + +
        + +
        + + + + +
        + + + + + + + + + + \ No newline at end of file