add sreweekly data dir + openclaw 每日复盘 2026-09-04
This commit is contained in:
@@ -0,0 +1,9 @@
|
||||
# Incidents start before the response does
|
||||
|
||||
- **期号**: SRE Weekly Issue #533(2026-09-07)
|
||||
- **作者**: Brent Chapman
|
||||
- **链接**: https://greatcircle.com/blog/2026/08/04/detection-gap/
|
||||
|
||||
## 简介
|
||||
|
||||
What can you do to shorten the time to detect an incident? Some great ideas in here, especially monitoring your company’s main web page for a sudden uptick in traffic.
|
||||
@@ -0,0 +1,9 @@
|
||||
# Quick thoughts on Azure Regional Outage from July 23, ’26
|
||||
|
||||
- **期号**: SRE Weekly Issue #533(2026-09-07)
|
||||
- **作者**: Lorin Hochstein
|
||||
- **链接**: https://surfingcomplexity.blog/2026/08/16/quick-thoughts-on-azure-regional-outage-from-july-23-26/
|
||||
|
||||
## 简介
|
||||
|
||||
What an interesting incident! I recommend reading Azure’s write-up before reading Lorin’s excellent analysis.
|
||||
@@ -0,0 +1,9 @@
|
||||
# Why Distributed Databases Fail at Coordination Boundaries
|
||||
|
||||
- **期号**: SRE Weekly Issue #533(2026-09-07)
|
||||
- **作者**: Varsha Ganesh — DZone
|
||||
- **链接**: https://dzone.com/articles/distributed-databases-coordination
|
||||
|
||||
## 简介
|
||||
|
||||
> Distributed databases rarely fail in the clean, isolated ways described by component diagrams. They fail through timing gaps, stale metadata, ambiguous ownership, retry storms, incompatible health decisions, and overlapping maintenance activity.
|
||||
15
sreweekly/markdown/533/04-the-record-says.md
Normal file
15
sreweekly/markdown/533/04-the-record-says.md
Normal file
@@ -0,0 +1,15 @@
|
||||
# The Record Says
|
||||
|
||||
- **期号**: SRE Weekly Issue #533(2026-09-07)
|
||||
- **作者**: Tim Irving
|
||||
- **链接**: https://read.zerosevzero.com/p/the-record-says
|
||||
|
||||
## 简介
|
||||
|
||||
I love this concept of a “political incident”:
|
||||
|
||||
> The subject was political incidents, by which I mean the ones where the severity arrives before the impact assessment does.
|
||||
|
||||
And ouch, I felt this bit:
|
||||
|
||||
> You have spent forty minutes of the incident on the severity field.
|
||||
@@ -0,0 +1,9 @@
|
||||
# What SREs Should Automate — and Never Automate — with AI
|
||||
|
||||
- **期号**: SRE Weekly Issue #533(2026-09-07)
|
||||
- **作者**: Sai Joshitha Kathari — HackerNoon
|
||||
- **链接**: https://hackernoon.com/what-sres-should-automate-and-never-automate-with-ai
|
||||
|
||||
## 简介
|
||||
|
||||
Where can you safely use LLM agents, versus when you should keep things in human hands? This one has some good criteria to consider.
|
||||
@@ -0,0 +1,9 @@
|
||||
# 20× the CI traffic without getting slower: How we rebuilt Git serving at Datadog
|
||||
|
||||
- **期号**: SRE Weekly Issue #533(2026-09-07)
|
||||
- **作者**: Mike Thompson and Daniel Esponda — Datadog
|
||||
- **链接**: https://www.datadoghq.com/blog/engineering/gitretriever/
|
||||
|
||||
## 简介
|
||||
|
||||
I learned a lot about Git while reading this one. Speeding up Git clones in CI may not seem important, but it will when you’re trying to roll out a fix during an incident.
|
||||
@@ -0,0 +1,9 @@
|
||||
# A Tale of Two Flink Autoscalers
|
||||
|
||||
- **期号**: SRE Weekly Issue #533(2026-09-07)
|
||||
- **作者**: Samuel Yeboah, Francesco Di Chiara and Mingliang Liu — Netflix
|
||||
- **链接**: https://netflixtechblog.com/a-tale-of-two-flink-autoscalers-e9f6a1b1492b
|
||||
|
||||
## 简介
|
||||
|
||||
Switching from their custom-written autoscaler to the new off-the-shelf option made sense, but it wasn’t a simple drop-in replacement.
|
||||
@@ -0,0 +1,9 @@
|
||||
# There is more to code review than (automatable) detection
|
||||
|
||||
- **期号**: SRE Weekly Issue #533(2026-09-07)
|
||||
- **作者**: John Allspaw — Adaptive Capacity Labs
|
||||
- **链接**: https://www.adaptivecapacitylabs.com/2026/08/24/there-is-more-to-code-review-than-automatable-detection/
|
||||
|
||||
## 简介
|
||||
|
||||
Can we replace human code review with LLM-based reviews? This article lays out what an LLM can’t replicate, and I’d argue that these are the pieces that matter most for reliability.
|
||||
Reference in New Issue
Block a user