96 lines
5.0 KiB
HTML
96 lines
5.0 KiB
HTML
<p><a class="email_only" href="https://sreweekly.com/sre-weekly-issue-529/">View on sreweekly.com</a></p>
|
||
|
||
<div class="sreweekly-sponsor-message" style="border: 1px solid #b0b0b0; width: 80%;">
|
||
<h2 style="text-align: center; font-size: 80%; color: #909090;">A message from our sponsor, <a href="https://sreweekly.com/link/529">Planetscale</a>:</h2>
|
||
<p>Your on-call rotation shouldn’t double as your database’s HA strategy. PlanetScale databases ship with a primary and two replicas across three AZs, automated failover, and a 99.999% multi-region SLA. Postgres and Vitess available in AWS and GCP.</p>
|
||
<p><a href="https://sreweekly.com/link/529">→ Get started with PlanetScale for just $5/mo</a></p>
|
||
</div>
|
||
|
||
|
||
<div class="wp-block-group"><div class="wp-block-group__inner-container is-layout-flow wp-block-group-is-layout-flow">
|
||
<div class="sreweekly-entry">
|
||
<div class="sreweekly-title"><a href="https://greatcircle.com/blog/2026/07/14/incident-management-process-versus-program/" rel="noopener" target="_blank">Without a program to support them, incident management processes wither</a></div>
|
||
<div class="sreweekly-description">
|
||
<p>It’s not enough to define an incident process. You have to spin up <em>and maintain</em> an entire incident management program.</p>
|
||
<p> <small>Brent Chapman</small></p>
|
||
</div>
|
||
</div>
|
||
|
||
|
||
|
||
<div class="sreweekly-entry">
|
||
<div class="sreweekly-title"><a href="https://www.honeycomb.io/blog/what-comes-after-observability" rel="noopener" target="_blank">What Comes After Observability?</a></div>
|
||
<div class="sreweekly-description">
|
||
<p>Honeycomb pulls back the curtain a bit to delve into how LLM agents change the way their product is used, and how their query patterns differ from humans’. It’s especially interesting that increasing agent usage has not correlated with decreasing human usage.</p>
|
||
<p> <small>Austin Parker — Honeycomb</small></p>
|
||
</div>
|
||
</div>
|
||
|
||
|
||
|
||
<div class="sreweekly-entry">
|
||
<div class="sreweekly-title"><a href="https://dzone.com/articles/rise-of-agentic-sre" rel="noopener" target="_blank">The Rise of Agentic SRE: Humans, Agents, and Reliability</a></div>
|
||
<div class="sreweekly-description">
|
||
<p>I like the approach here, especially measuring both the positive and negative outcomes.</p>
|
||
<blockquote>
|
||
<p>Good SRE practice is about evidence, not enthusiasm.</p>
|
||
</blockquote>
|
||
<p> <small> Neel Shah — DZone</small></p>
|
||
</div>
|
||
</div>
|
||
|
||
|
||
|
||
<div class="sreweekly-entry">
|
||
<div class="sreweekly-title"><a href="https://blog.railway.com/p/incident-report-july-2-2026-us-east-services-outage" rel="noopener" target="_blank">Incident Report: July 2, 2026 — US East Services Outage</a></div>
|
||
<div class="sreweekly-description">
|
||
<p>The kernel’s route cache: a hidden reliability killer. This is a really intriguing case of self-sustaining impact.</p>
|
||
<p> <small>Ray Chen — Railway</small></p>
|
||
</div>
|
||
</div>
|
||
|
||
|
||
|
||
<div class="sreweekly-entry">
|
||
<div class="sreweekly-title"><a href="https://antithesis.com/blog/2026/finding-bugs-in-raft-implementations/" rel="noopener" target="_blank">Finding bugs in Raft implementations</a></div>
|
||
<div class="sreweekly-description">
|
||
<blockquote>
|
||
<p>…we rely on formal verification, and this is how consensus algorithms are built today. We define a model that we can mathematically prove to be correct, and then we… translate this perfect, platonic thing into code.</p>
|
||
</blockquote>
|
||
<p> <small>TW Lim — Antithesis</small></p>
|
||
</div>
|
||
</div>
|
||
|
||
|
||
|
||
<div class="sreweekly-entry">
|
||
<div class="sreweekly-title"><a href="https://omarghader.github.io/monitoring-infrastructure-guide-2026/" rel="noopener" target="_blank">How to Build Your Infrastructure Monitoring in 2026 ·</a></div>
|
||
<div class="sreweekly-description">
|
||
<p>I love that this starts with the user. Monitor what matters to your users, and alert on what you can action.</p>
|
||
<p> <small>Omar Ghader</small></p>
|
||
</div>
|
||
</div>
|
||
|
||
|
||
|
||
<div class="sreweekly-entry">
|
||
<div class="sreweekly-title"><a href="https://utcc.utoronto.ca/~cks/space/blog/linux/SystemdPrivateTmpWhere" rel="noopener" target="_blank">Getting access to the /tmp of a systemd service with PrivateTmp=yes</a></div>
|
||
<div class="sreweekly-description">
|
||
<p>First time I’ve heard of systemd’s PrivateTmp feature. Neat!</p>
|
||
<p> <small>Chris Siebenmann</small></p>
|
||
</div>
|
||
</div>
|
||
|
||
|
||
|
||
<div class="sreweekly-entry">
|
||
<div class="sreweekly-title"><a href="https://surfingcomplexity.blog/2026/08/02/traditional-versus-resilience-engineering-views/" rel="noopener" target="_blank">Traditional versus resilience engineering views</a></div>
|
||
<div class="sreweekly-description">
|
||
<blockquote>
|
||
<p>I thought it would be a useful exercise to brainstorm some of the differences in focus between what I’ll call the traditional view of reliability, and the resilience engineering view.</p>
|
||
</blockquote>
|
||
<p>It’s short (just a table), but it definitely made me think.</p>
|
||
<p> <small>Lorin Hochstein</small></p>
|
||
</div>
|
||
</div>
|
||
</div></div> |