add sreweekly data dir + openclaw 每日复盘 2026-09-04
This commit is contained in:
@@ -0,0 +1,9 @@
|
||||
# Expertise Can’t Be Automated: Why Resilience Still Needs Humans
|
||||
|
||||
- **期号**: SRE Weekly Issue #530(2026-08-17)
|
||||
- **作者**: Courtney Nash — Resilience in Software Foundation
|
||||
- **链接**: https://resilienceinsoftware.org/news/11560646
|
||||
|
||||
## 简介
|
||||
|
||||
We may improve velocity by handing off tasks to LLM agents, but can that impact resilience?
|
||||
@@ -0,0 +1,9 @@
|
||||
# Respecting fatigue isn’t coddling
|
||||
|
||||
- **期号**: SRE Weekly Issue #530(2026-08-17)
|
||||
- **作者**: Brent Chapman
|
||||
- **链接**: https://greatcircle.com/blog/2026/07/28/respecting-fatigue-isnt-coddling/
|
||||
|
||||
## 简介
|
||||
|
||||
Fatigue and burn-out are reliability risks. Fatigue and burn-out are reliability risks. I champion this idea in my SRE practice constantly, and I hope you do too.
|
||||
@@ -0,0 +1,9 @@
|
||||
# On building scalable control planes
|
||||
|
||||
- **期号**: SRE Weekly Issue #530(2026-08-17)
|
||||
- **作者**: Zak van der Merw
|
||||
- **链接**: https://www.allthingsdistributed.com/2026/08/on-building-scalable-control-planes.html
|
||||
|
||||
## 简介
|
||||
|
||||
A fun read on how to build control planes for large-scale systems, with some great tidbits on the inner workings of EC2 and Aurora DSQL.
|
||||
@@ -0,0 +1,9 @@
|
||||
# Mario Saved the EU but Broke My System
|
||||
|
||||
- **期号**: SRE Weekly Issue #530(2026-08-17)
|
||||
- **作者**: Hamed Silatani — Uptime Labs
|
||||
- **链接**: https://www.uptimelabs.io/articles/hamed-2012-outage-reflections
|
||||
|
||||
## 简介
|
||||
|
||||
A harrowing incident story underlining the importance of expertise and experience.
|
||||
9
sreweekly/markdown/530/05-ai-and-sre.md
Normal file
9
sreweekly/markdown/530/05-ai-and-sre.md
Normal file
@@ -0,0 +1,9 @@
|
||||
# AI and SRE
|
||||
|
||||
- **期号**: SRE Weekly Issue #530(2026-08-17)
|
||||
- **作者**: Bill Duncan
|
||||
- **链接**: https://billduncan.org/ai-and-sre/
|
||||
|
||||
## 简介
|
||||
|
||||
An SRE comes to terms with the way LLM agents are changing our field: what works well, what still requires human involvement, and what the future may look like.
|
||||
@@ -0,0 +1,11 @@
|
||||
# Certificate Expiry Is Still Taking Down Major Platforms
|
||||
|
||||
- **期号**: SRE Weekly Issue #530(2026-08-17)
|
||||
- **作者**: TokenTimer
|
||||
- **链接**: https://tokentimer.ch/blog/tls-certificate-expiry-outages
|
||||
|
||||
## 简介
|
||||
|
||||
> Recent outages at Tailscale, jsDelivr, ServiceNow, and IPinfo show the same failure pattern: certificate automation broke quietly, while the expiry date kept moving closer.
|
||||
|
||||
Bonus: they include links to several write-ups of related incidents.
|
||||
@@ -0,0 +1,11 @@
|
||||
# How Stripe uses graph search and state machines to auto-remediate a global database fleet
|
||||
|
||||
- **期号**: SRE Weekly Issue #530(2026-08-17)
|
||||
- **作者**: Pragya Mehta and Sai Samant — Stripe
|
||||
- **链接**: https://stripe.dev/blog/how-stripe-uses-graph-search-and-state-machines-to-auto-remediate-a-global-database-fleet
|
||||
|
||||
## 简介
|
||||
|
||||
> Our solution treats infrastructure state as a traversable graph and lets a pathfinding algorithm discover recovery sequences at runtime.
|
||||
|
||||
Whoa, cool trick!
|
||||
@@ -0,0 +1,11 @@
|
||||
# We turned off Pub/Sub and nobody noticed
|
||||
|
||||
- **期号**: SRE Weekly Issue #530(2026-08-17)
|
||||
- **作者**: Patrick Hamann and Mike Fisher — incident.io
|
||||
- **链接**: https://incident.io/blog/we-turned-off-pub-sub-and-nobody-noticed
|
||||
|
||||
## 简介
|
||||
|
||||
Their event-oriented system was based on Google Pub/Sub with its 99.95% SLA, but their own SLA was 99.99%. To resolve that, they moved toward an active-active architecture, load-balancing across 2 message brokers.
|
||||
|
||||
There’s an interactive simulation of their algorithm midway through that’s fun to play with!
|
||||
Reference in New Issue
Block a user