add sreweekly data dir + openclaw 每日复盘 2026-09-04
This commit is contained in:
@@ -0,0 +1,9 @@
|
||||
# Without a program to support them, incident management processes wither
|
||||
|
||||
- **期号**: SRE Weekly Issue #529(2026-08-10)
|
||||
- **作者**: Brent Chapman
|
||||
- **链接**: https://greatcircle.com/blog/2026/07/14/incident-management-process-versus-program/
|
||||
|
||||
## 简介
|
||||
|
||||
It’s not enough to define an incident process. You have to spin up and maintain an entire incident management program.
|
||||
@@ -0,0 +1,9 @@
|
||||
# What Comes After Observability?
|
||||
|
||||
- **期号**: SRE Weekly Issue #529(2026-08-10)
|
||||
- **作者**: Austin Parker — Honeycomb
|
||||
- **链接**: https://www.honeycomb.io/blog/what-comes-after-observability
|
||||
|
||||
## 简介
|
||||
|
||||
Honeycomb pulls back the curtain a bit to delve into how LLM agents change the way their product is used, and how their query patterns differ from humans’. It’s especially interesting that increasing agent usage has not correlated with decreasing human usage.
|
||||
@@ -0,0 +1,11 @@
|
||||
# The Rise of Agentic SRE: Humans, Agents, and Reliability
|
||||
|
||||
- **期号**: SRE Weekly Issue #529(2026-08-10)
|
||||
- **作者**: Neel Shah — DZone
|
||||
- **链接**: https://dzone.com/articles/rise-of-agentic-sre
|
||||
|
||||
## 简介
|
||||
|
||||
I like the approach here, especially measuring both the positive and negative outcomes.
|
||||
|
||||
> Good SRE practice is about evidence, not enthusiasm.
|
||||
@@ -0,0 +1,9 @@
|
||||
# Incident Report: July 2, 2026 — US East Services Outage
|
||||
|
||||
- **期号**: SRE Weekly Issue #529(2026-08-10)
|
||||
- **作者**: Ray Chen — Railway
|
||||
- **链接**: https://blog.railway.com/p/incident-report-july-2-2026-us-east-services-outage
|
||||
|
||||
## 简介
|
||||
|
||||
The kernel’s route cache: a hidden reliability killer. This is a really intriguing case of self-sustaining impact.
|
||||
@@ -0,0 +1,9 @@
|
||||
# Finding bugs in Raft implementations
|
||||
|
||||
- **期号**: SRE Weekly Issue #529(2026-08-10)
|
||||
- **作者**: TW Lim — Antithesis
|
||||
- **链接**: https://antithesis.com/blog/2026/finding-bugs-in-raft-implementations/
|
||||
|
||||
## 简介
|
||||
|
||||
> …we rely on formal verification, and this is how consensus algorithms are built today. We define a model that we can mathematically prove to be correct, and then we… translate this perfect, platonic thing into code.
|
||||
@@ -0,0 +1,9 @@
|
||||
# How to Build Your Infrastructure Monitoring in 2026 ·
|
||||
|
||||
- **期号**: SRE Weekly Issue #529(2026-08-10)
|
||||
- **作者**: Omar Ghader
|
||||
- **链接**: https://omarghader.github.io/monitoring-infrastructure-guide-2026/
|
||||
|
||||
## 简介
|
||||
|
||||
I love that this starts with the user. Monitor what matters to your users, and alert on what you can action.
|
||||
@@ -0,0 +1,9 @@
|
||||
# Getting access to the /tmp of a systemd service with PrivateTmp=yes
|
||||
|
||||
- **期号**: SRE Weekly Issue #529(2026-08-10)
|
||||
- **作者**: Chris Siebenmann
|
||||
- **链接**: https://utcc.utoronto.ca/~cks/space/blog/linux/SystemdPrivateTmpWhere
|
||||
|
||||
## 简介
|
||||
|
||||
First time I’ve heard of systemd’s PrivateTmp feature. Neat!
|
||||
@@ -0,0 +1,11 @@
|
||||
# Traditional versus resilience engineering views
|
||||
|
||||
- **期号**: SRE Weekly Issue #529(2026-08-10)
|
||||
- **作者**: Lorin Hochstein
|
||||
- **链接**: https://surfingcomplexity.blog/2026/08/02/traditional-versus-resilience-engineering-views/
|
||||
|
||||
## 简介
|
||||
|
||||
> I thought it would be a useful exercise to brainstorm some of the differences in focus between what I’ll call the traditional view of reliability, and the resilience engineering view.
|
||||
|
||||
It’s short (just a table), but it definitely made me think.
|
||||
@@ -0,0 +1,9 @@
|
||||
# Expertise Can’t Be Automated: Why Resilience Still Needs Humans
|
||||
|
||||
- **期号**: SRE Weekly Issue #530(2026-08-17)
|
||||
- **作者**: Courtney Nash — Resilience in Software Foundation
|
||||
- **链接**: https://resilienceinsoftware.org/news/11560646
|
||||
|
||||
## 简介
|
||||
|
||||
We may improve velocity by handing off tasks to LLM agents, but can that impact resilience?
|
||||
@@ -0,0 +1,9 @@
|
||||
# Respecting fatigue isn’t coddling
|
||||
|
||||
- **期号**: SRE Weekly Issue #530(2026-08-17)
|
||||
- **作者**: Brent Chapman
|
||||
- **链接**: https://greatcircle.com/blog/2026/07/28/respecting-fatigue-isnt-coddling/
|
||||
|
||||
## 简介
|
||||
|
||||
Fatigue and burn-out are reliability risks. Fatigue and burn-out are reliability risks. I champion this idea in my SRE practice constantly, and I hope you do too.
|
||||
@@ -0,0 +1,9 @@
|
||||
# On building scalable control planes
|
||||
|
||||
- **期号**: SRE Weekly Issue #530(2026-08-17)
|
||||
- **作者**: Zak van der Merw
|
||||
- **链接**: https://www.allthingsdistributed.com/2026/08/on-building-scalable-control-planes.html
|
||||
|
||||
## 简介
|
||||
|
||||
A fun read on how to build control planes for large-scale systems, with some great tidbits on the inner workings of EC2 and Aurora DSQL.
|
||||
@@ -0,0 +1,9 @@
|
||||
# Mario Saved the EU but Broke My System
|
||||
|
||||
- **期号**: SRE Weekly Issue #530(2026-08-17)
|
||||
- **作者**: Hamed Silatani — Uptime Labs
|
||||
- **链接**: https://www.uptimelabs.io/articles/hamed-2012-outage-reflections
|
||||
|
||||
## 简介
|
||||
|
||||
A harrowing incident story underlining the importance of expertise and experience.
|
||||
9
sreweekly/markdown/530/05-ai-and-sre.md
Normal file
9
sreweekly/markdown/530/05-ai-and-sre.md
Normal file
@@ -0,0 +1,9 @@
|
||||
# AI and SRE
|
||||
|
||||
- **期号**: SRE Weekly Issue #530(2026-08-17)
|
||||
- **作者**: Bill Duncan
|
||||
- **链接**: https://billduncan.org/ai-and-sre/
|
||||
|
||||
## 简介
|
||||
|
||||
An SRE comes to terms with the way LLM agents are changing our field: what works well, what still requires human involvement, and what the future may look like.
|
||||
@@ -0,0 +1,11 @@
|
||||
# Certificate Expiry Is Still Taking Down Major Platforms
|
||||
|
||||
- **期号**: SRE Weekly Issue #530(2026-08-17)
|
||||
- **作者**: TokenTimer
|
||||
- **链接**: https://tokentimer.ch/blog/tls-certificate-expiry-outages
|
||||
|
||||
## 简介
|
||||
|
||||
> Recent outages at Tailscale, jsDelivr, ServiceNow, and IPinfo show the same failure pattern: certificate automation broke quietly, while the expiry date kept moving closer.
|
||||
|
||||
Bonus: they include links to several write-ups of related incidents.
|
||||
@@ -0,0 +1,11 @@
|
||||
# How Stripe uses graph search and state machines to auto-remediate a global database fleet
|
||||
|
||||
- **期号**: SRE Weekly Issue #530(2026-08-17)
|
||||
- **作者**: Pragya Mehta and Sai Samant — Stripe
|
||||
- **链接**: https://stripe.dev/blog/how-stripe-uses-graph-search-and-state-machines-to-auto-remediate-a-global-database-fleet
|
||||
|
||||
## 简介
|
||||
|
||||
> Our solution treats infrastructure state as a traversable graph and lets a pathfinding algorithm discover recovery sequences at runtime.
|
||||
|
||||
Whoa, cool trick!
|
||||
@@ -0,0 +1,11 @@
|
||||
# We turned off Pub/Sub and nobody noticed
|
||||
|
||||
- **期号**: SRE Weekly Issue #530(2026-08-17)
|
||||
- **作者**: Patrick Hamann and Mike Fisher — incident.io
|
||||
- **链接**: https://incident.io/blog/we-turned-off-pub-sub-and-nobody-noticed
|
||||
|
||||
## 简介
|
||||
|
||||
Their event-oriented system was based on Google Pub/Sub with its 99.95% SLA, but their own SLA was 99.99%. To resolve that, they moved toward an active-active architecture, load-balancing across 2 message brokers.
|
||||
|
||||
There’s an interactive simulation of their algorithm midway through that’s fun to play with!
|
||||
@@ -0,0 +1,9 @@
|
||||
# Heroic saves are near misses
|
||||
|
||||
- **期号**: SRE Weekly Issue #531(2026-08-24)
|
||||
- **作者**: Brent Chapman
|
||||
- **链接**: https://greatcircle.com/blog/2026/07/21/heroic-saves-are-near-misses/
|
||||
|
||||
## 简介
|
||||
|
||||
Celebrating heroes in incident response can incentivize further heroics. That can prevent the kind of growth that will improve incident response overall.
|
||||
@@ -0,0 +1,9 @@
|
||||
# Control and complexity: tension in systems design
|
||||
|
||||
- **期号**: SRE Weekly Issue #531(2026-08-24)
|
||||
- **作者**: Fred Hebert
|
||||
- **链接**: https://ferd.ca/control-and-complexity-tension-in-systems-design.html
|
||||
|
||||
## 简介
|
||||
|
||||
What might happen when we quickly adopt LLMs and make sweeping changes in our complex systems?
|
||||
@@ -0,0 +1,9 @@
|
||||
# Structured Logging in Distributed Systems: What Most Teams Get Wrong and How to Fix It
|
||||
|
||||
- **期号**: SRE Weekly Issue #531(2026-08-24)
|
||||
- **作者**: Ashwini Dave — DZone
|
||||
- **链接**: https://dzone.com/articles/structured-logging-in-distributed-systems
|
||||
|
||||
## 简介
|
||||
|
||||
> Most teams log, but log badly: wrong severity levels, no trace IDs, inconsistent fields, and logs siloed from traces.
|
||||
11
sreweekly/markdown/531/04-seeing-the-people-in-control.md
Normal file
11
sreweekly/markdown/531/04-seeing-the-people-in-control.md
Normal file
@@ -0,0 +1,11 @@
|
||||
# Seeing the People In Control
|
||||
|
||||
- **期号**: SRE Weekly Issue #531(2026-08-24)
|
||||
- **作者**: Steven Shorrock
|
||||
- **链接**: https://humanisticsystems.com/2025/10/14/seeing-the-people-in-control/
|
||||
|
||||
## 简介
|
||||
|
||||
> If you had to explain to a neighbour why your organisation is so safe, and generally works well, what would you say?
|
||||
|
||||
It’s all about people. I really enjoyed the quote from Charles Billings on principles for automation.
|
||||
9
sreweekly/markdown/531/05-type-conversion.md
Normal file
9
sreweekly/markdown/531/05-type-conversion.md
Normal file
@@ -0,0 +1,9 @@
|
||||
# Type Conversion
|
||||
|
||||
- **期号**: SRE Weekly Issue #531(2026-08-24)
|
||||
- **作者**: Bill Duncan
|
||||
- **链接**: https://billduncan.org/type-conversion/
|
||||
|
||||
## 简介
|
||||
|
||||
Type conversion in aviation involves an experienced pilot training on a new kind of aircraft. This article draws a parallel to transitioning to a new job as an SRE.
|
||||
@@ -0,0 +1,9 @@
|
||||
# ‘My Boss Wants Me to Pick an AI SRE Tool’: Q&A at Incident Fest (Adaptive Capacity Labs)
|
||||
|
||||
- **期号**: SRE Weekly Issue #531(2026-08-24)
|
||||
- **作者**: Sam Salter — Uptime Labs, with John Allspaw and Beth Adele Long
|
||||
- **链接**: https://www.uptimelabs.io/articles/incident-response-ai-qa
|
||||
|
||||
## 简介
|
||||
|
||||
Some big names in this Q&A, and they share a couple of delicious morsels.
|
||||
@@ -0,0 +1,11 @@
|
||||
# How Tailscale helped find the SQLite WAL-Reset bug
|
||||
|
||||
- **期号**: SRE Weekly Issue #531(2026-08-24)
|
||||
- **作者**: Alex Chan — Tailscale
|
||||
- **链接**: https://tailscale.com/blog/sqlite-wal-reset-bug
|
||||
|
||||
## 简介
|
||||
|
||||
A super-engaging deep-dive.
|
||||
|
||||
> This investigation is a useful reminder: running boring technology in a non-standard way is a risk.
|
||||
@@ -0,0 +1,9 @@
|
||||
# Optimizing Kubernetes pods for reliability with topology spread constraints
|
||||
|
||||
- **期号**: SRE Weekly Issue #531(2026-08-24)
|
||||
- **作者**: Andre Newman — Gremlin
|
||||
- **链接**: https://www.gremlin.com/blog/optimizing-kubernetes-pod-deployments-for-reliability-with-topology-spread-constraints
|
||||
|
||||
## 简介
|
||||
|
||||
A handy guide on topology constraints in Kubernetes, with a worked example.
|
||||
@@ -0,0 +1,9 @@
|
||||
# When declaring an incident becomes everyone’s favorite workaround
|
||||
|
||||
- **期号**: SRE Weekly Issue #532(2026-08-31)
|
||||
- **作者**: Brent Chapman
|
||||
- **链接**: https://greatcircle.com/blog/2026/08/11/declaring-incidents-for-side-effects/
|
||||
|
||||
## 简介
|
||||
|
||||
Need another team to do something fast? Just use this one weird trick: declare an incident! This article explains why the obvious solution (gating incident declaration) isn’t a good idea.
|
||||
@@ -0,0 +1,9 @@
|
||||
# Unethical Ways to Manage Technical Debt
|
||||
|
||||
- **期号**: SRE Weekly Issue #532(2026-08-31)
|
||||
- **作者**: Thomas A. Limoncelli — ACM Queue
|
||||
- **链接**: https://queue.acm.org/detail.cfm?ref=rss&id=3830399
|
||||
|
||||
## 简介
|
||||
|
||||
Ethics are relative, right? This article is full of genuinely useful tips and framings.
|
||||
@@ -0,0 +1,9 @@
|
||||
# Solving mysterious Kubernetes pod setup timeouts by tuning conntrack garbage collection
|
||||
|
||||
- **期号**: SRE Weekly Issue #532(2026-08-31)
|
||||
- **作者**: Jorrick Sleijster — Adyen
|
||||
- **链接**: https://www.adyen.com/knowledge-hub/inside-cilium-cni-solving-kubernetes-pod-setup-timeouts
|
||||
|
||||
## 简介
|
||||
|
||||
Whoa. It’s been quite a few years since my last run-in with an overfull conntrack table, and this is a fun new twist.
|
||||
@@ -0,0 +1,9 @@
|
||||
# Storage at scale: what I actually watched
|
||||
|
||||
- **期号**: SRE Weekly Issue #532(2026-08-31)
|
||||
- **作者**: Sridhar Rajarao
|
||||
- **链接**: https://sridharrajarao.com/blog/storage-at-scale/
|
||||
|
||||
## 简介
|
||||
|
||||
> For eight years I ran SRE for a storage system measured in exabytes. The dashboard I checked every morning shrank to seven numbers. Here they are.
|
||||
@@ -0,0 +1,9 @@
|
||||
# The Rise of Cognitive Observability
|
||||
|
||||
- **期号**: SRE Weekly Issue #532(2026-08-31)
|
||||
- **作者**: Barnadeep Bhowmik
|
||||
- **链接**: https://pub.towardsai.net/the-rise-of-cognitive-observability-a77e33250037
|
||||
|
||||
## 简介
|
||||
|
||||
> Traditional observability monitors execution. LLM observability must monitor behavior.
|
||||
@@ -0,0 +1,9 @@
|
||||
# How Uber Conquered Database Overload: The Journey from Static Rate-Limiting to Intelligent Load Management
|
||||
|
||||
- **期号**: SRE Weekly Issue #532(2026-08-31)
|
||||
- **作者**: Dhyanam Vaidya, Prathamesh Deshpande, and Mike Ma — Uber
|
||||
- **链接**: https://www.uber.com/us/en/blog/from-static-rate-limiting-to-intelligent-load-management/
|
||||
|
||||
## 简介
|
||||
|
||||
This one has a lot of great detail on how their approaches to quota management failed and how they iterated.
|
||||
@@ -0,0 +1,9 @@
|
||||
# Voyager and the Art of Graceful Degradation
|
||||
|
||||
- **期号**: SRE Weekly Issue #532(2026-08-31)
|
||||
- **作者**: Robert Barron
|
||||
- **链接**: https://www.flyingbarron.com/2026/04/voyager-and-art-of-graceful-degradation.html
|
||||
|
||||
## 简介
|
||||
|
||||
This article uses Voyager 1, whose engineers just shut down another instrument to conserve its steadily-decaying power, as an extended analogy for graceful degradation.
|
||||
@@ -0,0 +1,9 @@
|
||||
# Incidents start before the response does
|
||||
|
||||
- **期号**: SRE Weekly Issue #533(2026-09-07)
|
||||
- **作者**: Brent Chapman
|
||||
- **链接**: https://greatcircle.com/blog/2026/08/04/detection-gap/
|
||||
|
||||
## 简介
|
||||
|
||||
What can you do to shorten the time to detect an incident? Some great ideas in here, especially monitoring your company’s main web page for a sudden uptick in traffic.
|
||||
@@ -0,0 +1,9 @@
|
||||
# Quick thoughts on Azure Regional Outage from July 23, ’26
|
||||
|
||||
- **期号**: SRE Weekly Issue #533(2026-09-07)
|
||||
- **作者**: Lorin Hochstein
|
||||
- **链接**: https://surfingcomplexity.blog/2026/08/16/quick-thoughts-on-azure-regional-outage-from-july-23-26/
|
||||
|
||||
## 简介
|
||||
|
||||
What an interesting incident! I recommend reading Azure’s write-up before reading Lorin’s excellent analysis.
|
||||
@@ -0,0 +1,9 @@
|
||||
# Why Distributed Databases Fail at Coordination Boundaries
|
||||
|
||||
- **期号**: SRE Weekly Issue #533(2026-09-07)
|
||||
- **作者**: Varsha Ganesh — DZone
|
||||
- **链接**: https://dzone.com/articles/distributed-databases-coordination
|
||||
|
||||
## 简介
|
||||
|
||||
> Distributed databases rarely fail in the clean, isolated ways described by component diagrams. They fail through timing gaps, stale metadata, ambiguous ownership, retry storms, incompatible health decisions, and overlapping maintenance activity.
|
||||
15
sreweekly/markdown/533/04-the-record-says.md
Normal file
15
sreweekly/markdown/533/04-the-record-says.md
Normal file
@@ -0,0 +1,15 @@
|
||||
# The Record Says
|
||||
|
||||
- **期号**: SRE Weekly Issue #533(2026-09-07)
|
||||
- **作者**: Tim Irving
|
||||
- **链接**: https://read.zerosevzero.com/p/the-record-says
|
||||
|
||||
## 简介
|
||||
|
||||
I love this concept of a “political incident”:
|
||||
|
||||
> The subject was political incidents, by which I mean the ones where the severity arrives before the impact assessment does.
|
||||
|
||||
And ouch, I felt this bit:
|
||||
|
||||
> You have spent forty minutes of the incident on the severity field.
|
||||
@@ -0,0 +1,9 @@
|
||||
# What SREs Should Automate — and Never Automate — with AI
|
||||
|
||||
- **期号**: SRE Weekly Issue #533(2026-09-07)
|
||||
- **作者**: Sai Joshitha Kathari — HackerNoon
|
||||
- **链接**: https://hackernoon.com/what-sres-should-automate-and-never-automate-with-ai
|
||||
|
||||
## 简介
|
||||
|
||||
Where can you safely use LLM agents, versus when you should keep things in human hands? This one has some good criteria to consider.
|
||||
@@ -0,0 +1,9 @@
|
||||
# 20× the CI traffic without getting slower: How we rebuilt Git serving at Datadog
|
||||
|
||||
- **期号**: SRE Weekly Issue #533(2026-09-07)
|
||||
- **作者**: Mike Thompson and Daniel Esponda — Datadog
|
||||
- **链接**: https://www.datadoghq.com/blog/engineering/gitretriever/
|
||||
|
||||
## 简介
|
||||
|
||||
I learned a lot about Git while reading this one. Speeding up Git clones in CI may not seem important, but it will when you’re trying to roll out a fix during an incident.
|
||||
@@ -0,0 +1,9 @@
|
||||
# A Tale of Two Flink Autoscalers
|
||||
|
||||
- **期号**: SRE Weekly Issue #533(2026-09-07)
|
||||
- **作者**: Samuel Yeboah, Francesco Di Chiara and Mingliang Liu — Netflix
|
||||
- **链接**: https://netflixtechblog.com/a-tale-of-two-flink-autoscalers-e9f6a1b1492b
|
||||
|
||||
## 简介
|
||||
|
||||
Switching from their custom-written autoscaler to the new off-the-shelf option made sense, but it wasn’t a simple drop-in replacement.
|
||||
@@ -0,0 +1,9 @@
|
||||
# There is more to code review than (automatable) detection
|
||||
|
||||
- **期号**: SRE Weekly Issue #533(2026-09-07)
|
||||
- **作者**: John Allspaw — Adaptive Capacity Labs
|
||||
- **链接**: https://www.adaptivecapacitylabs.com/2026/08/24/there-is-more-to-code-review-than-automatable-detection/
|
||||
|
||||
## 简介
|
||||
|
||||
Can we replace human code review with LLM-based reviews? This article lays out what an LLM can’t replicate, and I’d argue that these are the pieces that matter most for reliability.
|
||||
Reference in New Issue
Block a user