94 lines
4.9 KiB
HTML
94 lines
4.9 KiB
HTML
<p><a class="email_only" href="https://sreweekly.com/sre-weekly-issue-530/">View on sreweekly.com</a></p>
|
||
<div class="sreweekly-sponsor-message" style="border: 1px solid #b0b0b0; width: 80%;">
|
||
<h2 style="text-align: center; font-size: 80%; color: #909090;">A message from our sponsor, <a href="https://sreweekly.com/link/530">Planetscale</a>:</h2>
|
||
<p>Your on-call rotation shouldn’t double as your database’s HA strategy. PlanetScale databases ship with a primary and two replicas across three AZs, automated failover, and a 99.999% multi-region SLA. Postgres and Vitess available in AWS and GCP.</p>
|
||
<p><a href="https://sreweekly.com/link/530">→ Get started with PlanetScale for just $5/mo</a></p>
|
||
</div>
|
||
|
||
|
||
<div class="wp-block-group"><div class="wp-block-group__inner-container is-layout-flow wp-block-group-is-layout-flow">
|
||
<div class="sreweekly-entry">
|
||
<div class="sreweekly-title"><a href="https://resilienceinsoftware.org/news/11560646" rel="noopener" target="_blank">Expertise Can’t Be Automated: Why Resilience Still Needs Humans</a></div>
|
||
<div class="sreweekly-description">
|
||
<p>We may improve velocity by handing off tasks to LLM agents, but can that impact resilience? </p>
|
||
<p> <small>Courtney Nash — Resilience in Software Foundation</small></p>
|
||
</div>
|
||
</div>
|
||
|
||
|
||
|
||
<div class="sreweekly-entry">
|
||
<div class="sreweekly-title"><a href="https://greatcircle.com/blog/2026/07/28/respecting-fatigue-isnt-coddling/" rel="noopener" target="_blank">Respecting fatigue isn’t coddling</a></div>
|
||
<div class="sreweekly-description">
|
||
<p>Fatigue and burn-out are reliability risks. <strong>Fatigue and burn-out are reliability risks.</strong> I champion this idea in my SRE practice constantly, and I hope you do too.</p>
|
||
<p> <small>Brent Chapman</small></p>
|
||
</div>
|
||
</div>
|
||
|
||
|
||
|
||
<div class="sreweekly-entry">
|
||
<div class="sreweekly-title"><a href="https://www.allthingsdistributed.com/2026/08/on-building-scalable-control-planes.html" rel="noopener" target="_blank">On building scalable control planes</a></div>
|
||
<div class="sreweekly-description">
|
||
<p>A fun read on how to build control planes for large-scale systems, with some great tidbits on the inner workings of EC2 and Aurora DSQL.</p>
|
||
<p> <small>Zak van der Merw</small></p>
|
||
</div>
|
||
</div>
|
||
|
||
|
||
|
||
<div class="sreweekly-entry">
|
||
<div class="sreweekly-title"><a href="https://www.uptimelabs.io/articles/hamed-2012-outage-reflections" rel="noopener" target="_blank">Mario Saved the EU but Broke My System</a></div>
|
||
<div class="sreweekly-description">
|
||
<p>A harrowing incident story underlining the importance of expertise and experience.</p>
|
||
<p> <small>Hamed Silatani — Uptime Labs</small></p>
|
||
</div>
|
||
</div>
|
||
|
||
|
||
|
||
<div class="sreweekly-entry">
|
||
<div class="sreweekly-title"><a href="https://billduncan.org/ai-and-sre/" rel="noopener" target="_blank">AI and SRE</a></div>
|
||
<div class="sreweekly-description">
|
||
<p>An SRE comes to terms with the way LLM agents are changing our field: what works well, what still requires human involvement, and what the future may look like.</p>
|
||
<p> <small>Bill Duncan</small></p>
|
||
</div>
|
||
</div>
|
||
|
||
|
||
|
||
<div class="sreweekly-entry">
|
||
<div class="sreweekly-title"><a href="https://tokentimer.ch/blog/tls-certificate-expiry-outages" rel="noopener" target="_blank">Certificate Expiry Is Still Taking Down Major Platforms</a></div>
|
||
<div class="sreweekly-description">
|
||
<blockquote>
|
||
<p>Recent outages at Tailscale, jsDelivr, ServiceNow, and IPinfo show the same failure pattern: certificate automation broke quietly, while the expiry date kept moving closer.</p>
|
||
</blockquote>
|
||
<p>Bonus: they include links to several write-ups of related incidents.</p>
|
||
<p> <small>TokenTimer</small></p>
|
||
</div>
|
||
</div>
|
||
|
||
|
||
|
||
<div class="sreweekly-entry">
|
||
<div class="sreweekly-title"><a href="https://stripe.dev/blog/how-stripe-uses-graph-search-and-state-machines-to-auto-remediate-a-global-database-fleet" rel="noopener" target="_blank">How Stripe uses graph search and state machines to auto-remediate a global database fleet</a></div>
|
||
<div class="sreweekly-description">
|
||
<blockquote>
|
||
<p>Our solution treats infrastructure state as a traversable graph and lets a pathfinding algorithm discover recovery sequences at runtime.</p>
|
||
</blockquote>
|
||
<p>Whoa, cool trick!</p>
|
||
<p> <small>Pragya Mehta and Sai Samant — Stripe</small></p>
|
||
</div>
|
||
</div>
|
||
|
||
|
||
|
||
<div class="sreweekly-entry">
|
||
<div class="sreweekly-title"><a href="https://incident.io/blog/we-turned-off-pub-sub-and-nobody-noticed" rel="noopener" target="_blank">We turned off Pub/Sub and nobody noticed</a></div>
|
||
<div class="sreweekly-description">
|
||
<p>Their event-oriented system was based on Google Pub/Sub with its 99.95% SLA, but their own SLA was 99.99%. To resolve that, they moved toward an active-active architecture, load-balancing across 2 message brokers.</p>
|
||
<p>There’s an interactive simulation of their algorithm midway through that’s fun to play with!</p>
|
||
<p> <small>Patrick Hamann and Mike Fisher — incident.io</small></p>
|
||
</div>
|
||
</div>
|
||
</div></div> |