This commit is contained in:
2026-09-11 06:20:11 +08:00
parent 8065f42563
commit e80aa747e5
12 changed files with 396 additions and 1259 deletions

View File

@@ -1,96 +0,0 @@
<p><a class="email_only" href="https://sreweekly.com/sre-weekly-issue-529/">View on sreweekly.com</a></p>
<div class="sreweekly-sponsor-message" style="border: 1px solid #b0b0b0; width: 80%;">
<h2 style="text-align: center; font-size: 80%; color: #909090;">A message from our sponsor, <a href="https://sreweekly.com/link/529">Planetscale</a>:</h2>
<p>Your on-call rotation shouldn&#8217;t double as your database&#8217;s HA strategy. PlanetScale databases ship with a primary and two replicas across three AZs, automated failover, and a 99.999% multi-region SLA. Postgres and Vitess available in AWS and GCP.</p>
<p><a href="https://sreweekly.com/link/529">→ Get started with PlanetScale for just $5/mo</a></p>
</div>
<div class="wp-block-group"><div class="wp-block-group__inner-container is-layout-flow wp-block-group-is-layout-flow">
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://greatcircle.com/blog/2026/07/14/incident-management-process-versus-program/" rel="noopener" target="_blank">Without a program to support them, incident management processes wither</a></div>
<div class="sreweekly-description">
<p>It&#8217;s not enough to define an incident process. You have to spin up <em>and maintain</em> an entire incident management program.</p>
<p>&nbsp;&nbsp;<small>Brent Chapman</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://www.honeycomb.io/blog/what-comes-after-observability" rel="noopener" target="_blank">What Comes After Observability?</a></div>
<div class="sreweekly-description">
<p>Honeycomb pulls back the curtain a bit to delve into how LLM agents change the way their product is used, and how their query patterns differ from humans&#8217;. It&#8217;s especially interesting that increasing agent usage has not correlated with decreasing human usage.</p>
<p>&nbsp;&nbsp;<small>Austin Parker &mdash; Honeycomb</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://dzone.com/articles/rise-of-agentic-sre" rel="noopener" target="_blank">The Rise of Agentic SRE: Humans, Agents, and Reliability</a></div>
<div class="sreweekly-description">
<p>I like the approach here, especially measuring both the positive and negative outcomes.</p>
<blockquote>
<p>Good SRE practice is about evidence, not enthusiasm.</p>
</blockquote>
<p>&nbsp;&nbsp;<small> Neel Shah &mdash; DZone</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://blog.railway.com/p/incident-report-july-2-2026-us-east-services-outage" rel="noopener" target="_blank">Incident Report: July 2, 2026 — US East Services Outage</a></div>
<div class="sreweekly-description">
<p>The kernel&#8217;s route cache: a hidden reliability killer. This is a really intriguing case of self-sustaining impact.</p>
<p>&nbsp;&nbsp;<small>Ray Chen &mdash; Railway</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://antithesis.com/blog/2026/finding-bugs-in-raft-implementations/" rel="noopener" target="_blank">Finding bugs in Raft implementations</a></div>
<div class="sreweekly-description">
<blockquote>
<p>&#8230;we rely on formal verification, and this is how consensus algorithms are built today. We define a model that we can mathematically prove to be correct, and then we… translate this perfect, platonic thing into code.</p>
</blockquote>
<p>&nbsp;&nbsp;<small>TW Lim &mdash; Antithesis</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://omarghader.github.io/monitoring-infrastructure-guide-2026/" rel="noopener" target="_blank">How to Build Your Infrastructure Monitoring in 2026 ·</a></div>
<div class="sreweekly-description">
<p>I love that this starts with the user. Monitor what matters to your users, and alert on what you can action.</p>
<p>&nbsp;&nbsp;<small>Omar Ghader</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://utcc.utoronto.ca/~cks/space/blog/linux/SystemdPrivateTmpWhere" rel="noopener" target="_blank">Getting access to the /tmp of a systemd service with PrivateTmp=yes</a></div>
<div class="sreweekly-description">
<p>First time I&#8217;ve heard of systemd&#8217;s PrivateTmp feature. Neat!</p>
<p>&nbsp;&nbsp;<small>Chris Siebenmann</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://surfingcomplexity.blog/2026/08/02/traditional-versus-resilience-engineering-views/" rel="noopener" target="_blank">Traditional versus resilience engineering views</a></div>
<div class="sreweekly-description">
<blockquote>
<p>I thought it would be a useful exercise to brainstorm some of the differences in focus between what I’ll call the traditional view of reliability, and the resilience engineering view.</p>
</blockquote>
<p>It&#8217;s short (just a table), but it definitely made me think.</p>
<p>&nbsp;&nbsp;<small>Lorin Hochstein</small></p>
</div>
</div>
</div></div>

View File

@@ -1,94 +0,0 @@
<p><a class="email_only" href="https://sreweekly.com/sre-weekly-issue-530/">View on sreweekly.com</a></p>
<div class="sreweekly-sponsor-message" style="border: 1px solid #b0b0b0; width: 80%;">
<h2 style="text-align: center; font-size: 80%; color: #909090;">A message from our sponsor, <a href="https://sreweekly.com/link/530">Planetscale</a>:</h2>
<p>Your on-call rotation shouldn&#8217;t double as your database&#8217;s HA strategy. PlanetScale databases ship with a primary and two replicas across three AZs, automated failover, and a 99.999% multi-region SLA. Postgres and Vitess available in AWS and GCP.</p>
<p><a href="https://sreweekly.com/link/530">→ Get started with PlanetScale for just $5/mo</a></p>
</div>
<div class="wp-block-group"><div class="wp-block-group__inner-container is-layout-flow wp-block-group-is-layout-flow">
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://resilienceinsoftware.org/news/11560646" rel="noopener" target="_blank">Expertise Can&#8217;t Be Automated: Why Resilience Still Needs Humans</a></div>
<div class="sreweekly-description">
<p>We may improve velocity by handing off tasks to LLM agents, but can that impact resilience? </p>
<p>&nbsp;&nbsp;<small>Courtney Nash &mdash; Resilience in Software Foundation</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://greatcircle.com/blog/2026/07/28/respecting-fatigue-isnt-coddling/" rel="noopener" target="_blank">Respecting fatigue isn’t coddling</a></div>
<div class="sreweekly-description">
<p>Fatigue and burn-out are reliability risks. <strong>Fatigue and burn-out are reliability risks.</strong> I champion this idea in my SRE practice constantly, and I hope you do too.</p>
<p>  <small>Brent Chapman</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://www.allthingsdistributed.com/2026/08/on-building-scalable-control-planes.html" rel="noopener" target="_blank">On building scalable control planes</a></div>
<div class="sreweekly-description">
<p>A fun read on how to build control planes for large-scale systems, with some great tidbits on the inner workings of EC2 and Aurora DSQL.</p>
<p>  <small>Zak van der Merw</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://www.uptimelabs.io/articles/hamed-2012-outage-reflections" rel="noopener" target="_blank">Mario Saved the EU but Broke My System</a></div>
<div class="sreweekly-description">
<p>A harrowing incident story underlining the importance of expertise and experience.</p>
<p>&nbsp;&nbsp;<small>Hamed Silatani &mdash; Uptime Labs</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://billduncan.org/ai-and-sre/" rel="noopener" target="_blank">AI and SRE</a></div>
<div class="sreweekly-description">
<p>An SRE comes to terms with the way LLM agents are changing our field: what works well, what still requires human involvement, and what the future may look like.</p>
<p>&nbsp;&nbsp;<small>Bill Duncan</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://tokentimer.ch/blog/tls-certificate-expiry-outages" rel="noopener" target="_blank">Certificate Expiry Is Still Taking Down Major Platforms</a></div>
<div class="sreweekly-description">
<blockquote>
<p>Recent outages at Tailscale, jsDelivr, ServiceNow, and IPinfo show the same failure pattern: certificate automation broke quietly, while the expiry date kept moving closer.</p>
</blockquote>
<p>Bonus: they include links to several write-ups of related incidents.</p>
<p>  <small>TokenTimer</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://stripe.dev/blog/how-stripe-uses-graph-search-and-state-machines-to-auto-remediate-a-global-database-fleet" rel="noopener" target="_blank">How Stripe uses graph search and state machines to auto-remediate a global database fleet</a></div>
<div class="sreweekly-description">
<blockquote>
<p>Our solution treats infrastructure state as a traversable graph and lets a pathfinding algorithm discover recovery sequences at runtime.</p>
</blockquote>
<p>Whoa, cool trick!</p>
<p>  <small>Pragya Mehta and Sai Samant — Stripe</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://incident.io/blog/we-turned-off-pub-sub-and-nobody-noticed" rel="noopener" target="_blank">We turned off Pub/Sub and nobody noticed</a></div>
<div class="sreweekly-description">
<p>Their event-oriented system was based on Google Pub/Sub with its 99.95% SLA, but their own SLA was 99.99%. To resolve that, they moved toward an active-active architecture, load-balancing across 2 message brokers.</p>
<p>There&#8217;s an interactive simulation of their algorithm midway through that&#8217;s fun to play with!</p>
<p>  <small>Patrick Hamann and Mike Fisher — incident.io</small></p>
</div>
</div>
</div></div>

View File

@@ -1,96 +0,0 @@
<p><a class="email_only" href="https://sreweekly.com/sre-weekly-issue-531/">View on sreweekly.com</a></p>
<div class="sreweekly-sponsor-message" style="border: 1px solid #b0b0b0; width: 80%;">
<h2 style="text-align: center; font-size: 80%; color: #909090;">A message from our sponsor, <a href="https://sreweekly.com/link/531">Planetscale</a>:</h2>
<p>PlanetScale Metal runs Postgres and Vitess on dedicated NVMe inside AWS and GCP. Get data center speed next to your app, with unlimited IOPS and no throttling. Teams routinely see a 70% drop in p99 and p95 latency after migrating.</p>
<p><a href="https://sreweekly.com/link/531">→ See the benchmarks</a></p>
</div>
<div class="wp-block-group"><div class="wp-block-group__inner-container is-layout-flow wp-block-group-is-layout-flow">
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://greatcircle.com/blog/2026/07/21/heroic-saves-are-near-misses/" target="_blank">Heroic saves are near misses</a></div>
<div class="sreweekly-description">
<p>Celebrating heroes in incident response can incentivize further heroics. That can prevent the kind of growth that will improve incident response overall.</p>
<p>&nbsp;&nbsp;<small>Brent Chapman</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://ferd.ca/control-and-complexity-tension-in-systems-design.html" target="_blank">Control and complexity: tension in systems design</a></div>
<div class="sreweekly-description">
<p>What might happen when we quickly adopt LLMs and make sweeping changes in our complex systems?</p>
<p>&nbsp;&nbsp;<small>Fred Hebert</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://dzone.com/articles/structured-logging-in-distributed-systems" target="_blank">Structured Logging in Distributed Systems: What Most Teams Get Wrong and How to Fix It</a></div>
<div class="sreweekly-description">
<blockquote>
<p>Most teams log, but log badly: wrong severity levels, no trace IDs, inconsistent fields, and logs siloed from traces.</p>
</blockquote>
<p>&nbsp;&nbsp;<small>Ashwini Dave &mdash; DZone</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://humanisticsystems.com/2025/10/14/seeing-the-people-in-control/" target="_blank">Seeing the People In Control</a></div>
<div class="sreweekly-description">
<blockquote>
<p>If you had to explain to a neighbour why your organisation is so safe, and generally works well, what would you say?</p>
</blockquote>
<p>It&#8217;s all about people. I really enjoyed the quote from Charles Billings on principles for automation.</p>
<p>&nbsp;&nbsp;<small>Steven Shorrock</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://billduncan.org/type-conversion/" target="_blank">Type Conversion</a></div>
<div class="sreweekly-description">
<p>Type conversion in aviation involves an experienced pilot training on a new kind of aircraft. This article draws a parallel to transitioning to a new job as an SRE.</p>
<p>&nbsp;&nbsp;<small>Bill Duncan</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://www.uptimelabs.io/articles/incident-response-ai-qa" target="_blank">&#8216;My Boss Wants Me to Pick an AI SRE Tool&#8217;: Q&#038;A at Incident Fest (Adaptive Capacity Labs)</a></div>
<div class="sreweekly-description">
<p>Some big names in this Q&amp;A, and they share a couple of delicious morsels.</p>
<p>&nbsp;&nbsp;<small>Sam Salter &mdash; Uptime Labs, with John Allspaw and Beth Adele Long</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://tailscale.com/blog/sqlite-wal-reset-bug" target="_blank">How Tailscale helped find the SQLite WAL-Reset bug</a></div>
<div class="sreweekly-description">
<p>A super-engaging deep-dive.</p>
<blockquote>
<p>This investigation is a useful reminder: <strong>running boring technology in a non-standard way is a risk.</strong></p>
</blockquote>
<p>&nbsp;&nbsp;<small>Alex Chan &mdash; Tailscale</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://www.gremlin.com/blog/optimizing-kubernetes-pod-deployments-for-reliability-with-topology-spread-constraints" target="_blank">Optimizing Kubernetes pods for reliability with topology spread constraints</a></div>
<div class="sreweekly-description">
<p>A handy guide on topology constraints in Kubernetes, with a worked example.</p>
<p>&nbsp;&nbsp;<small>Andre Newman &mdash; Gremlin</small></p>
</div>
</div>
</div></div>

View File

@@ -1,82 +0,0 @@
<p><a class="email_only" href="https://sreweekly.com/sre-weekly-issue-532/">View on sreweekly.com</a></p>
<div class="sreweekly-sponsor-message" style="border: 1px solid #b0b0b0; width: 80%;">
<h2 style="text-align: center; font-size: 80%; color: #909090;">A message from our sponsor, <a href="https://sreweekly.com/link/532">Planetscale</a>:</h2>
<p>PlanetScale Metal runs Postgres and Vitess on dedicated NVMe inside AWS and GCP. Get data center speed next to your app, with unlimited IOPS and no throttling. Teams routinely see a 70% drop in p99 and p95 latency after migrating.</p>
<p><a href="https://sreweekly.com/link/532">→ See the benchmarks</a></p>
</div>
<div class="wp-block-group"><div class="wp-block-group__inner-container is-layout-flow wp-block-group-is-layout-flow">
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://greatcircle.com/blog/2026/08/11/declaring-incidents-for-side-effects/" target="_blank">When declaring an incident becomes everyone’s favorite workaround</a></div>
<div class="sreweekly-description">
<p>Need another team to do something fast? Just use this one weird trick: declare an incident! This article explains why the obvious solution (gating incident declaration) isn&#8217;t a good idea.</p>
<p>&nbsp;&nbsp;<small>Brent Chapman</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://queue.acm.org/detail.cfm?ref=rss&#038;id=3830399" target="_blank">Unethical Ways to Manage Technical Debt</a></div>
<div class="sreweekly-description">
<p>Ethics are relative, right? This article is full of genuinely useful tips and framings.</p>
<p>&nbsp;&nbsp;<small>Thomas A. Limoncelli &mdash; ACM Queue</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://www.adyen.com/knowledge-hub/inside-cilium-cni-solving-kubernetes-pod-setup-timeouts" target="_blank">Solving mysterious Kubernetes pod setup timeouts by tuning conntrack garbage collection</a></div>
<div class="sreweekly-description">
<p>Whoa. It&#8217;s been quite a few years since my last run-in with an overfull conntrack table, and this is a fun new twist.</p>
<p>&nbsp;&nbsp;<small>Jorrick Sleijster &mdash; Adyen</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://sridharrajarao.com/blog/storage-at-scale/" target="_blank">Storage at scale: what I actually watched</a></div>
<div class="sreweekly-description">
<blockquote>
<p>For eight years I ran SRE for a storage system measured in exabytes. The dashboard I checked every morning shrank to seven numbers. Here they are.</p>
</blockquote>
<p>&nbsp;&nbsp;<small>Sridhar Rajarao</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://pub.towardsai.net/the-rise-of-cognitive-observability-a77e33250037" target="_blank">The Rise of Cognitive Observability</a></div>
<div class="sreweekly-description">
<blockquote>
<p><i>Traditional observability monitors execution. LLM observability must monitor behavior.</i></p>
</blockquote>
<p>&nbsp;&nbsp;<small>Barnadeep Bhowmik</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://www.uber.com/us/en/blog/from-static-rate-limiting-to-intelligent-load-management/" rel="noopener" target="_blank">How Uber Conquered Database Overload: The Journey from Static Rate-Limiting to Intelligent Load Management</a></div>
<div class="sreweekly-description">
<p>This one has a lot of great detail on how their approaches to quota management failed and how they iterated.</p>
<p>  <small>Dhyanam Vaidya, Prathamesh Deshpande, and Mike Ma — Uber</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://www.flyingbarron.com/2026/04/voyager-and-art-of-graceful-degradation.html" target="_blank">Voyager and the Art of Graceful Degradation</a></div>
<div class="sreweekly-description">
<p>This article uses Voyager 1, whose engineers just shut down another instrument to conserve its steadily-decaying power, as an extended analogy for graceful degradation.</p>
<p>&nbsp;&nbsp;<small>Robert Barron</small></p>
</div>
</div>
</div></div>

View File

@@ -1,97 +0,0 @@
<p><a class="email_only" href="https://sreweekly.com/sre-weekly-issue-533/">View on sreweekly.com</a></p>
<div class="sreweekly-sponsor-message" style="border: 1px solid #b0b0b0; width: 80%;">
<h2 style="text-align: center; font-size: 80%; color: #909090;">A message from our sponsor, <a href="https://sreweekly.com/link/533">Planetscale</a>:</h2>
<p>Most database incidents start with one expensive query, not the database being down. PlanetScale gives SRE teams high-availability Postgres and MySQL with automated failover, query insights, and Database Traffic Control to stop runaway queries before they page you.</p>
<p><a href="https://sreweekly.com/link/533">→ Explore PlanetScale</a></p>
</div>
<div class="wp-block-group"><div class="wp-block-group__inner-container is-layout-flow wp-block-group-is-layout-flow">
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://greatcircle.com/blog/2026/08/04/detection-gap/" target="_blank">Incidents start before the response does</a></div>
<div class="sreweekly-description">
<p>What can you do to shorten the time to detect an incident? Some great ideas in here, especially monitoring your company&#8217;s main web page for a sudden uptick in traffic.</p>
<p>&nbsp;&nbsp;<small>Brent Chapman</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://surfingcomplexity.blog/2026/08/16/quick-thoughts-on-azure-regional-outage-from-july-23-26/" target="_blank">Quick thoughts on Azure Regional Outage from July 23, ’26</a></div>
<div class="sreweekly-description">
<p>What an interesting incident! I recommend reading Azure&#8217;s write-up before reading Lorin&#8217;s excellent analysis.</p>
<p>&nbsp;&nbsp;<small>Lorin Hochstein</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://dzone.com/articles/distributed-databases-coordination" target="_blank">Why Distributed Databases Fail at Coordination Boundaries</a></div>
<div class="sreweekly-description">
<blockquote>
<p>Distributed databases rarely fail in the clean, isolated ways described by component diagrams. They fail through timing gaps, stale metadata, ambiguous ownership, retry storms, incompatible health decisions, and overlapping maintenance activity.</p>
</blockquote>
<p>&nbsp;&nbsp;<small> Varsha Ganesh &mdash; DZone</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://read.zerosevzero.com/p/the-record-says" target="_blank">The Record Says</a></div>
<div class="sreweekly-description">
<p>I love this concept of a &#8220;political incident&#8221;:</p>
<blockquote>
<p>The subject was political incidents, by which I mean the ones where the severity arrives before the impact assessment does.</p>
</blockquote>
<p>And ouch, I felt this bit:</p>
<blockquote>
<p>You have spent forty minutes of the incident on the severity field.</p>
</blockquote>
<p>&nbsp;&nbsp;<small>Tim Irving</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://hackernoon.com/what-sres-should-automate-and-never-automate-with-ai" target="_blank">What SREs Should Automate — and Never Automate — with AI</a></div>
<div class="sreweekly-description">
<p>Where can you safely use LLM agents, versus when you should keep things in human hands? This one has some good criteria to consider.</p>
<p>&nbsp;&nbsp;<small>Sai Joshitha Kathari &mdash; HackerNoon</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://www.datadoghq.com/blog/engineering/gitretriever/" target="_blank">20× the CI traffic without getting slower: How we rebuilt Git serving at Datadog</a></div>
<div class="sreweekly-description">
<p>I learned a lot about Git while reading this one. Speeding up Git clones in CI may not seem important, but it will when you&#8217;re trying to roll out a fix during an incident.</p>
<p>&nbsp;&nbsp;<small>Mike Thompson and Daniel Esponda &mdash; Datadog</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://netflixtechblog.com/a-tale-of-two-flink-autoscalers-e9f6a1b1492b" target="_blank">A Tale of Two Flink Autoscalers</a></div>
<div class="sreweekly-description">
<p>Switching from their custom-written autoscaler to the new off-the-shelf option made sense, but it wasn&#8217;t a simple drop-in replacement.</p>
<p>&nbsp;&nbsp;<small>Samuel Yeboah, Francesco Di Chiara and Mingliang Liu &mdash; Netflix</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://www.adaptivecapacitylabs.com/2026/08/24/there-is-more-to-code-review-than-automatable-detection/" target="_blank">There is more to code review than (automatable) detection</a></div>
<div class="sreweekly-description">
<p>Can we replace human code review with LLM-based reviews? This article lays out what an LLM can&#8217;t replicate, and I&#8217;d argue that these are the pieces that matter most for reliability.</p>
<p>&nbsp;&nbsp;<small>John Allspaw &mdash; Adaptive Capacity Labs</small></p>
</div>
</div>
</div></div>