This commit is contained in:
2026-09-11 06:20:11 +08:00
parent 8065f42563
commit e80aa747e5
12 changed files with 396 additions and 1259 deletions

View File

@@ -1,91 +0,0 @@
# SRE Weekly Digest
抓取 [SRE Weekly](https://sreweekly.com) 每期内容,把其中分享的文章逐篇导出为 Markdown,
供后续阅读/筛选/邮件分发使用。脚本由 **n8n 定时调用**,通过不同参数驱动;脚本自身**不发送邮件**。
- 脚本: `scripts/sreweekly.py`(单文件,仅依赖 `feedparser`)
- 依赖: `pip3 install feedparser`
## 目录结构
```
sreweekly/
├── manifest.json # 每期状态:html 是否保存、文章是否已提取、文章数
├── html/<期号>-<日期>.html # 每期原始 HTML(feed 的 content:encoded 原样保存)
└── markdown/<期号>/ # 每篇文章一个 .md 文件
├── 01-<标题>.md
└── ...
```
## 命令一览
```bash
python3 scripts/sreweekly.py # 扫描 feed,新增期保存 HTML + 登记 manifest
python3 scripts/sreweekly.py --check # 检查有无新期 → 输出 NEW:533,532 或 NONE
python3 scripts/sreweekly.py --extract 533 # 提取指定期文章 → markdown/533/
python3 scripts/sreweekly.py --extract latest # 提取最新一期
python3 scripts/sreweekly.py --extract pending # 提取所有尚未提取的期(每日流程推荐)
python3 scripts/sreweekly.py --extract all # 提取全部(已提取过需加 --force 覆盖)
python3 scripts/sreweekly.py --status # 打印状态表
python3 scripts/sreweekly.py --issues # 列出所有已登记期
# 可选参数
--feed <URL> # RSS 地址(默认 https://sreweekly.com/feed/)
--dir <PATH> # 工作目录(默认本目录)
--force # 强制重新提取
```
所有命令**幂等**:重复扫描不重复入库,重复提取自动跳过已提取的期。
## n8n 每日流程(推荐 3 步)
该刊每周一更新,建议每天定时跑一次,有新期才继续:
1. **Execute Command(扫描)**
```
python3 /Users/weishen/Workspace/nexus/Hermes/xingzhi/sreweekly/scripts/sreweekly.py --dir /Users/weishen/Workspace/nexus/Hermes/xingzhi/sreweekly
```
2. **Execute Command(提取未处理期)**
```
python3 .../sreweekly.py --dir ... --extract pending
```
(或先用 `--check` 判断 `NEW:`/`NONE` 再分支,`--extract pending` 本身也天然幂等,空跑无害)
3. **后续节点(邮件)**:读取 `markdown/<新期号>/` 下的 .md 文件,
用 n8n 的 Email / 自定义 AgentMail 节点发送(示例发件收件:star-agent@agentmail.to → billyshen@163.com)。
参考输出(`--check`,可直接作为 n8n IF 条件):
```
NEW:533,532 ← 有新期
NONE ← 无更新
```
## manifest.json 示例
```json
{
"feed": "https://sreweekly.com/feed/",
"last_scanned": "2026-09-10T07:40:50",
"issues": {
"533": {
"id": "533",
"title": "SRE Weekly Issue #533",
"url": "https://sreweekly.com/sre-weekly-issue-533/",
"pub_date": "2026-09-07",
"html_file": "html/533-2026-09-07.html",
"fetched_at": "2026-09-10T07:40:50",
"extracted": true,
"article_count": 8,
"markdown_dir": "markdown/533",
"extracted_at": "2026-09-10T07:41:10"
}
}
}
```
## 说明
- SRE Weekly 的 feed 已在 `content:encoded` 字段携带每期完整 HTML,无需再抓每期页面。
- 页面结构:每篇分享文章是 `<div class="sreweekly-entry">`(标题链接 + 简介 + 作者),
赞助商段落(`sreweekly-sponsor-message`)自动跳过。
- 提取后每篇 md 内容:标题、期号/日期、作者、原文链接、简介(含 blockquote 引用)。

View File

@@ -1,96 +0,0 @@
<p><a class="email_only" href="https://sreweekly.com/sre-weekly-issue-529/">View on sreweekly.com</a></p>
<div class="sreweekly-sponsor-message" style="border: 1px solid #b0b0b0; width: 80%;">
<h2 style="text-align: center; font-size: 80%; color: #909090;">A message from our sponsor, <a href="https://sreweekly.com/link/529">Planetscale</a>:</h2>
<p>Your on-call rotation shouldn&#8217;t double as your database&#8217;s HA strategy. PlanetScale databases ship with a primary and two replicas across three AZs, automated failover, and a 99.999% multi-region SLA. Postgres and Vitess available in AWS and GCP.</p>
<p><a href="https://sreweekly.com/link/529">→ Get started with PlanetScale for just $5/mo</a></p>
</div>
<div class="wp-block-group"><div class="wp-block-group__inner-container is-layout-flow wp-block-group-is-layout-flow">
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://greatcircle.com/blog/2026/07/14/incident-management-process-versus-program/" rel="noopener" target="_blank">Without a program to support them, incident management processes wither</a></div>
<div class="sreweekly-description">
<p>It&#8217;s not enough to define an incident process. You have to spin up <em>and maintain</em> an entire incident management program.</p>
<p>&nbsp;&nbsp;<small>Brent Chapman</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://www.honeycomb.io/blog/what-comes-after-observability" rel="noopener" target="_blank">What Comes After Observability?</a></div>
<div class="sreweekly-description">
<p>Honeycomb pulls back the curtain a bit to delve into how LLM agents change the way their product is used, and how their query patterns differ from humans&#8217;. It&#8217;s especially interesting that increasing agent usage has not correlated with decreasing human usage.</p>
<p>&nbsp;&nbsp;<small>Austin Parker &mdash; Honeycomb</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://dzone.com/articles/rise-of-agentic-sre" rel="noopener" target="_blank">The Rise of Agentic SRE: Humans, Agents, and Reliability</a></div>
<div class="sreweekly-description">
<p>I like the approach here, especially measuring both the positive and negative outcomes.</p>
<blockquote>
<p>Good SRE practice is about evidence, not enthusiasm.</p>
</blockquote>
<p>&nbsp;&nbsp;<small> Neel Shah &mdash; DZone</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://blog.railway.com/p/incident-report-july-2-2026-us-east-services-outage" rel="noopener" target="_blank">Incident Report: July 2, 2026 — US East Services Outage</a></div>
<div class="sreweekly-description">
<p>The kernel&#8217;s route cache: a hidden reliability killer. This is a really intriguing case of self-sustaining impact.</p>
<p>&nbsp;&nbsp;<small>Ray Chen &mdash; Railway</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://antithesis.com/blog/2026/finding-bugs-in-raft-implementations/" rel="noopener" target="_blank">Finding bugs in Raft implementations</a></div>
<div class="sreweekly-description">
<blockquote>
<p>&#8230;we rely on formal verification, and this is how consensus algorithms are built today. We define a model that we can mathematically prove to be correct, and then we… translate this perfect, platonic thing into code.</p>
</blockquote>
<p>&nbsp;&nbsp;<small>TW Lim &mdash; Antithesis</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://omarghader.github.io/monitoring-infrastructure-guide-2026/" rel="noopener" target="_blank">How to Build Your Infrastructure Monitoring in 2026 ·</a></div>
<div class="sreweekly-description">
<p>I love that this starts with the user. Monitor what matters to your users, and alert on what you can action.</p>
<p>&nbsp;&nbsp;<small>Omar Ghader</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://utcc.utoronto.ca/~cks/space/blog/linux/SystemdPrivateTmpWhere" rel="noopener" target="_blank">Getting access to the /tmp of a systemd service with PrivateTmp=yes</a></div>
<div class="sreweekly-description">
<p>First time I&#8217;ve heard of systemd&#8217;s PrivateTmp feature. Neat!</p>
<p>&nbsp;&nbsp;<small>Chris Siebenmann</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://surfingcomplexity.blog/2026/08/02/traditional-versus-resilience-engineering-views/" rel="noopener" target="_blank">Traditional versus resilience engineering views</a></div>
<div class="sreweekly-description">
<blockquote>
<p>I thought it would be a useful exercise to brainstorm some of the differences in focus between what I’ll call the traditional view of reliability, and the resilience engineering view.</p>
</blockquote>
<p>It&#8217;s short (just a table), but it definitely made me think.</p>
<p>&nbsp;&nbsp;<small>Lorin Hochstein</small></p>
</div>
</div>
</div></div>

View File

@@ -1,94 +0,0 @@
<p><a class="email_only" href="https://sreweekly.com/sre-weekly-issue-530/">View on sreweekly.com</a></p>
<div class="sreweekly-sponsor-message" style="border: 1px solid #b0b0b0; width: 80%;">
<h2 style="text-align: center; font-size: 80%; color: #909090;">A message from our sponsor, <a href="https://sreweekly.com/link/530">Planetscale</a>:</h2>
<p>Your on-call rotation shouldn&#8217;t double as your database&#8217;s HA strategy. PlanetScale databases ship with a primary and two replicas across three AZs, automated failover, and a 99.999% multi-region SLA. Postgres and Vitess available in AWS and GCP.</p>
<p><a href="https://sreweekly.com/link/530">→ Get started with PlanetScale for just $5/mo</a></p>
</div>
<div class="wp-block-group"><div class="wp-block-group__inner-container is-layout-flow wp-block-group-is-layout-flow">
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://resilienceinsoftware.org/news/11560646" rel="noopener" target="_blank">Expertise Can&#8217;t Be Automated: Why Resilience Still Needs Humans</a></div>
<div class="sreweekly-description">
<p>We may improve velocity by handing off tasks to LLM agents, but can that impact resilience? </p>
<p>&nbsp;&nbsp;<small>Courtney Nash &mdash; Resilience in Software Foundation</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://greatcircle.com/blog/2026/07/28/respecting-fatigue-isnt-coddling/" rel="noopener" target="_blank">Respecting fatigue isn’t coddling</a></div>
<div class="sreweekly-description">
<p>Fatigue and burn-out are reliability risks. <strong>Fatigue and burn-out are reliability risks.</strong> I champion this idea in my SRE practice constantly, and I hope you do too.</p>
<p>  <small>Brent Chapman</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://www.allthingsdistributed.com/2026/08/on-building-scalable-control-planes.html" rel="noopener" target="_blank">On building scalable control planes</a></div>
<div class="sreweekly-description">
<p>A fun read on how to build control planes for large-scale systems, with some great tidbits on the inner workings of EC2 and Aurora DSQL.</p>
<p>  <small>Zak van der Merw</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://www.uptimelabs.io/articles/hamed-2012-outage-reflections" rel="noopener" target="_blank">Mario Saved the EU but Broke My System</a></div>
<div class="sreweekly-description">
<p>A harrowing incident story underlining the importance of expertise and experience.</p>
<p>&nbsp;&nbsp;<small>Hamed Silatani &mdash; Uptime Labs</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://billduncan.org/ai-and-sre/" rel="noopener" target="_blank">AI and SRE</a></div>
<div class="sreweekly-description">
<p>An SRE comes to terms with the way LLM agents are changing our field: what works well, what still requires human involvement, and what the future may look like.</p>
<p>&nbsp;&nbsp;<small>Bill Duncan</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://tokentimer.ch/blog/tls-certificate-expiry-outages" rel="noopener" target="_blank">Certificate Expiry Is Still Taking Down Major Platforms</a></div>
<div class="sreweekly-description">
<blockquote>
<p>Recent outages at Tailscale, jsDelivr, ServiceNow, and IPinfo show the same failure pattern: certificate automation broke quietly, while the expiry date kept moving closer.</p>
</blockquote>
<p>Bonus: they include links to several write-ups of related incidents.</p>
<p>  <small>TokenTimer</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://stripe.dev/blog/how-stripe-uses-graph-search-and-state-machines-to-auto-remediate-a-global-database-fleet" rel="noopener" target="_blank">How Stripe uses graph search and state machines to auto-remediate a global database fleet</a></div>
<div class="sreweekly-description">
<blockquote>
<p>Our solution treats infrastructure state as a traversable graph and lets a pathfinding algorithm discover recovery sequences at runtime.</p>
</blockquote>
<p>Whoa, cool trick!</p>
<p>  <small>Pragya Mehta and Sai Samant — Stripe</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://incident.io/blog/we-turned-off-pub-sub-and-nobody-noticed" rel="noopener" target="_blank">We turned off Pub/Sub and nobody noticed</a></div>
<div class="sreweekly-description">
<p>Their event-oriented system was based on Google Pub/Sub with its 99.95% SLA, but their own SLA was 99.99%. To resolve that, they moved toward an active-active architecture, load-balancing across 2 message brokers.</p>
<p>There&#8217;s an interactive simulation of their algorithm midway through that&#8217;s fun to play with!</p>
<p>  <small>Patrick Hamann and Mike Fisher — incident.io</small></p>
</div>
</div>
</div></div>

View File

@@ -1,96 +0,0 @@
<p><a class="email_only" href="https://sreweekly.com/sre-weekly-issue-531/">View on sreweekly.com</a></p>
<div class="sreweekly-sponsor-message" style="border: 1px solid #b0b0b0; width: 80%;">
<h2 style="text-align: center; font-size: 80%; color: #909090;">A message from our sponsor, <a href="https://sreweekly.com/link/531">Planetscale</a>:</h2>
<p>PlanetScale Metal runs Postgres and Vitess on dedicated NVMe inside AWS and GCP. Get data center speed next to your app, with unlimited IOPS and no throttling. Teams routinely see a 70% drop in p99 and p95 latency after migrating.</p>
<p><a href="https://sreweekly.com/link/531">→ See the benchmarks</a></p>
</div>
<div class="wp-block-group"><div class="wp-block-group__inner-container is-layout-flow wp-block-group-is-layout-flow">
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://greatcircle.com/blog/2026/07/21/heroic-saves-are-near-misses/" target="_blank">Heroic saves are near misses</a></div>
<div class="sreweekly-description">
<p>Celebrating heroes in incident response can incentivize further heroics. That can prevent the kind of growth that will improve incident response overall.</p>
<p>&nbsp;&nbsp;<small>Brent Chapman</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://ferd.ca/control-and-complexity-tension-in-systems-design.html" target="_blank">Control and complexity: tension in systems design</a></div>
<div class="sreweekly-description">
<p>What might happen when we quickly adopt LLMs and make sweeping changes in our complex systems?</p>
<p>&nbsp;&nbsp;<small>Fred Hebert</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://dzone.com/articles/structured-logging-in-distributed-systems" target="_blank">Structured Logging in Distributed Systems: What Most Teams Get Wrong and How to Fix It</a></div>
<div class="sreweekly-description">
<blockquote>
<p>Most teams log, but log badly: wrong severity levels, no trace IDs, inconsistent fields, and logs siloed from traces.</p>
</blockquote>
<p>&nbsp;&nbsp;<small>Ashwini Dave &mdash; DZone</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://humanisticsystems.com/2025/10/14/seeing-the-people-in-control/" target="_blank">Seeing the People In Control</a></div>
<div class="sreweekly-description">
<blockquote>
<p>If you had to explain to a neighbour why your organisation is so safe, and generally works well, what would you say?</p>
</blockquote>
<p>It&#8217;s all about people. I really enjoyed the quote from Charles Billings on principles for automation.</p>
<p>&nbsp;&nbsp;<small>Steven Shorrock</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://billduncan.org/type-conversion/" target="_blank">Type Conversion</a></div>
<div class="sreweekly-description">
<p>Type conversion in aviation involves an experienced pilot training on a new kind of aircraft. This article draws a parallel to transitioning to a new job as an SRE.</p>
<p>&nbsp;&nbsp;<small>Bill Duncan</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://www.uptimelabs.io/articles/incident-response-ai-qa" target="_blank">&#8216;My Boss Wants Me to Pick an AI SRE Tool&#8217;: Q&#038;A at Incident Fest (Adaptive Capacity Labs)</a></div>
<div class="sreweekly-description">
<p>Some big names in this Q&amp;A, and they share a couple of delicious morsels.</p>
<p>&nbsp;&nbsp;<small>Sam Salter &mdash; Uptime Labs, with John Allspaw and Beth Adele Long</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://tailscale.com/blog/sqlite-wal-reset-bug" target="_blank">How Tailscale helped find the SQLite WAL-Reset bug</a></div>
<div class="sreweekly-description">
<p>A super-engaging deep-dive.</p>
<blockquote>
<p>This investigation is a useful reminder: <strong>running boring technology in a non-standard way is a risk.</strong></p>
</blockquote>
<p>&nbsp;&nbsp;<small>Alex Chan &mdash; Tailscale</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://www.gremlin.com/blog/optimizing-kubernetes-pod-deployments-for-reliability-with-topology-spread-constraints" target="_blank">Optimizing Kubernetes pods for reliability with topology spread constraints</a></div>
<div class="sreweekly-description">
<p>A handy guide on topology constraints in Kubernetes, with a worked example.</p>
<p>&nbsp;&nbsp;<small>Andre Newman &mdash; Gremlin</small></p>
</div>
</div>
</div></div>

View File

@@ -1,82 +0,0 @@
<p><a class="email_only" href="https://sreweekly.com/sre-weekly-issue-532/">View on sreweekly.com</a></p>
<div class="sreweekly-sponsor-message" style="border: 1px solid #b0b0b0; width: 80%;">
<h2 style="text-align: center; font-size: 80%; color: #909090;">A message from our sponsor, <a href="https://sreweekly.com/link/532">Planetscale</a>:</h2>
<p>PlanetScale Metal runs Postgres and Vitess on dedicated NVMe inside AWS and GCP. Get data center speed next to your app, with unlimited IOPS and no throttling. Teams routinely see a 70% drop in p99 and p95 latency after migrating.</p>
<p><a href="https://sreweekly.com/link/532">→ See the benchmarks</a></p>
</div>
<div class="wp-block-group"><div class="wp-block-group__inner-container is-layout-flow wp-block-group-is-layout-flow">
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://greatcircle.com/blog/2026/08/11/declaring-incidents-for-side-effects/" target="_blank">When declaring an incident becomes everyone’s favorite workaround</a></div>
<div class="sreweekly-description">
<p>Need another team to do something fast? Just use this one weird trick: declare an incident! This article explains why the obvious solution (gating incident declaration) isn&#8217;t a good idea.</p>
<p>&nbsp;&nbsp;<small>Brent Chapman</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://queue.acm.org/detail.cfm?ref=rss&#038;id=3830399" target="_blank">Unethical Ways to Manage Technical Debt</a></div>
<div class="sreweekly-description">
<p>Ethics are relative, right? This article is full of genuinely useful tips and framings.</p>
<p>&nbsp;&nbsp;<small>Thomas A. Limoncelli &mdash; ACM Queue</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://www.adyen.com/knowledge-hub/inside-cilium-cni-solving-kubernetes-pod-setup-timeouts" target="_blank">Solving mysterious Kubernetes pod setup timeouts by tuning conntrack garbage collection</a></div>
<div class="sreweekly-description">
<p>Whoa. It&#8217;s been quite a few years since my last run-in with an overfull conntrack table, and this is a fun new twist.</p>
<p>&nbsp;&nbsp;<small>Jorrick Sleijster &mdash; Adyen</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://sridharrajarao.com/blog/storage-at-scale/" target="_blank">Storage at scale: what I actually watched</a></div>
<div class="sreweekly-description">
<blockquote>
<p>For eight years I ran SRE for a storage system measured in exabytes. The dashboard I checked every morning shrank to seven numbers. Here they are.</p>
</blockquote>
<p>&nbsp;&nbsp;<small>Sridhar Rajarao</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://pub.towardsai.net/the-rise-of-cognitive-observability-a77e33250037" target="_blank">The Rise of Cognitive Observability</a></div>
<div class="sreweekly-description">
<blockquote>
<p><i>Traditional observability monitors execution. LLM observability must monitor behavior.</i></p>
</blockquote>
<p>&nbsp;&nbsp;<small>Barnadeep Bhowmik</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://www.uber.com/us/en/blog/from-static-rate-limiting-to-intelligent-load-management/" rel="noopener" target="_blank">How Uber Conquered Database Overload: The Journey from Static Rate-Limiting to Intelligent Load Management</a></div>
<div class="sreweekly-description">
<p>This one has a lot of great detail on how their approaches to quota management failed and how they iterated.</p>
<p>  <small>Dhyanam Vaidya, Prathamesh Deshpande, and Mike Ma — Uber</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://www.flyingbarron.com/2026/04/voyager-and-art-of-graceful-degradation.html" target="_blank">Voyager and the Art of Graceful Degradation</a></div>
<div class="sreweekly-description">
<p>This article uses Voyager 1, whose engineers just shut down another instrument to conserve its steadily-decaying power, as an extended analogy for graceful degradation.</p>
<p>&nbsp;&nbsp;<small>Robert Barron</small></p>
</div>
</div>
</div></div>

View File

@@ -1,97 +0,0 @@
<p><a class="email_only" href="https://sreweekly.com/sre-weekly-issue-533/">View on sreweekly.com</a></p>
<div class="sreweekly-sponsor-message" style="border: 1px solid #b0b0b0; width: 80%;">
<h2 style="text-align: center; font-size: 80%; color: #909090;">A message from our sponsor, <a href="https://sreweekly.com/link/533">Planetscale</a>:</h2>
<p>Most database incidents start with one expensive query, not the database being down. PlanetScale gives SRE teams high-availability Postgres and MySQL with automated failover, query insights, and Database Traffic Control to stop runaway queries before they page you.</p>
<p><a href="https://sreweekly.com/link/533">→ Explore PlanetScale</a></p>
</div>
<div class="wp-block-group"><div class="wp-block-group__inner-container is-layout-flow wp-block-group-is-layout-flow">
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://greatcircle.com/blog/2026/08/04/detection-gap/" target="_blank">Incidents start before the response does</a></div>
<div class="sreweekly-description">
<p>What can you do to shorten the time to detect an incident? Some great ideas in here, especially monitoring your company&#8217;s main web page for a sudden uptick in traffic.</p>
<p>&nbsp;&nbsp;<small>Brent Chapman</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://surfingcomplexity.blog/2026/08/16/quick-thoughts-on-azure-regional-outage-from-july-23-26/" target="_blank">Quick thoughts on Azure Regional Outage from July 23, ’26</a></div>
<div class="sreweekly-description">
<p>What an interesting incident! I recommend reading Azure&#8217;s write-up before reading Lorin&#8217;s excellent analysis.</p>
<p>&nbsp;&nbsp;<small>Lorin Hochstein</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://dzone.com/articles/distributed-databases-coordination" target="_blank">Why Distributed Databases Fail at Coordination Boundaries</a></div>
<div class="sreweekly-description">
<blockquote>
<p>Distributed databases rarely fail in the clean, isolated ways described by component diagrams. They fail through timing gaps, stale metadata, ambiguous ownership, retry storms, incompatible health decisions, and overlapping maintenance activity.</p>
</blockquote>
<p>&nbsp;&nbsp;<small> Varsha Ganesh &mdash; DZone</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://read.zerosevzero.com/p/the-record-says" target="_blank">The Record Says</a></div>
<div class="sreweekly-description">
<p>I love this concept of a &#8220;political incident&#8221;:</p>
<blockquote>
<p>The subject was political incidents, by which I mean the ones where the severity arrives before the impact assessment does.</p>
</blockquote>
<p>And ouch, I felt this bit:</p>
<blockquote>
<p>You have spent forty minutes of the incident on the severity field.</p>
</blockquote>
<p>&nbsp;&nbsp;<small>Tim Irving</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://hackernoon.com/what-sres-should-automate-and-never-automate-with-ai" target="_blank">What SREs Should Automate — and Never Automate — with AI</a></div>
<div class="sreweekly-description">
<p>Where can you safely use LLM agents, versus when you should keep things in human hands? This one has some good criteria to consider.</p>
<p>&nbsp;&nbsp;<small>Sai Joshitha Kathari &mdash; HackerNoon</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://www.datadoghq.com/blog/engineering/gitretriever/" target="_blank">20× the CI traffic without getting slower: How we rebuilt Git serving at Datadog</a></div>
<div class="sreweekly-description">
<p>I learned a lot about Git while reading this one. Speeding up Git clones in CI may not seem important, but it will when you&#8217;re trying to roll out a fix during an incident.</p>
<p>&nbsp;&nbsp;<small>Mike Thompson and Daniel Esponda &mdash; Datadog</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://netflixtechblog.com/a-tale-of-two-flink-autoscalers-e9f6a1b1492b" target="_blank">A Tale of Two Flink Autoscalers</a></div>
<div class="sreweekly-description">
<p>Switching from their custom-written autoscaler to the new off-the-shelf option made sense, but it wasn&#8217;t a simple drop-in replacement.</p>
<p>&nbsp;&nbsp;<small>Samuel Yeboah, Francesco Di Chiara and Mingliang Liu &mdash; Netflix</small></p>
</div>
</div>
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="https://www.adaptivecapacitylabs.com/2026/08/24/there-is-more-to-code-review-than-automatable-detection/" target="_blank">There is more to code review than (automatable) detection</a></div>
<div class="sreweekly-description">
<p>Can we replace human code review with LLM-based reviews? This article lays out what an LLM can&#8217;t replicate, and I&#8217;d argue that these are the pieces that matter most for reliability.</p>
<p>&nbsp;&nbsp;<small>John Allspaw &mdash; Adaptive Capacity Labs</small></p>
</div>
</div>
</div></div>

View File

@@ -1,391 +0,0 @@
#!/usr/bin/env python3
# -*- coding: utf-8 -*-
"""
SRE Weekly 抓取与文章提取脚本
模仿 blogwatcher-daily 的单文件风格:feedparser 抓取 → HTML 落盘 → 逐篇文章导出 Markdown,
并用 manifest.json 记录每期的状态(是否已存 HTML / 是否已提取文章)。
设计给 n8n 定时任务调用,通过不同参数驱动,脚本自身不发送邮件。
目录结构(默认工作目录 ~/Workspace/nexus/Hermes/xingzhi/sreweekly/):
sreweekly/
├── manifest.json # 状态记录(每期: html 是否保存、文章是否提取、文章数)
├── html/<id>-<日期>.html # 每期原始 HTML(feed 的 content:encoded 原样保存)
└── markdown/<id>/ # 每期文章的 markdown 目录(每篇一个文件)
├── 01-<标题slug>.md
└── 02-<标题slug>.md
用法:
python3 sreweekly.py # 扫描 feed,新增期保存为 HTML 并登记 manifest
python3 sreweekly.py --check # 只检查有无新期,输出 NEW:533,532 或 NONE(供 n8n 分支)
python3 sreweekly.py --extract ID # 提取指定期(如 533)的文章 → markdown
python3 sreweekly.py --extract latest # 提取最新一期
python3 sreweekly.py --extract pending # 提取所有尚未提取的期(n8n 每日流程推荐)
python3 sreweekly.py --extract all # 提取全部(含已提取的,需配合 --force 覆盖)
python3 sreweekly.py --status # 打印 manifest 状态表
python3 sreweekly.py --issues # 列出已登记的所有期
选项:
--feed URL RSS 地址(默认 https://sreweekly.com/feed/)
--dir PATH 工作目录(默认 ~/Workspace/nexus/Hermes/xingzhi/sreweekly/)
--force 重新提取已提取过的期(覆盖旧文件)
"""
from html.parser import HTMLParser
import argparse, html, json, os, re, sys, time, urllib.request
try:
import feedparser
except ImportError:
sys.exit("缺少依赖: 请先执行 pip3 install feedparser")
DEFAULT_FEED = "https://sreweekly.com/feed/"
DEFAULT_DIR = os.path.expanduser("~/Workspace/nexus/Hermes/xingzhi/sreweekly")
# ---------------------------------------------------------------- manifest
def load_manifest(base_dir):
path = os.path.join(base_dir, "manifest.json")
if os.path.exists(path):
try:
with open(path, "r", encoding="utf-8") as f:
return json.load(f)
except (json.JSONDecodeError, OSError) as e:
print(f"⚠️ manifest.json 读取失败({e}),按空库处理", file=sys.stderr)
return {"feed": DEFAULT_FEED, "last_scanned": None, "issues": {}}
def save_manifest(base_dir, manifest):
path = os.path.join(base_dir, "manifest.json")
with open(path, "w", encoding="utf-8") as f:
json.dump(manifest, f, ensure_ascii=False, indent=2)
return path
def issue_id_from_link(link):
"""从文章链接提取期号: sre-weekly-issue-533 → 533;失败则用最后一段路径。"""
m = re.search(r"sre-weekly-issue-(\d+)", link or "")
if m:
return int(m.group(1))
m = re.search(r"/([^/]+?)/?$", (link or "").rstrip("/"))
return m.group(1) if m else "unknown"
def format_pubdate(entry):
"""优先用 feed 自带发布日期(YYYY-MM-DD),没有则用当天。"""
if getattr(entry, "published_parsed", None):
return time.strftime("%Y-%m-%d", entry.published_parsed)
return time.strftime("%Y-%m-%d")
def now_iso():
return time.strftime("%Y-%m-%dT%H:%M:%S")
# ---------------------------------------------------------------- feed 扫描
def fetch_feed(feed_url):
d = feedparser.parse(feed_url)
if not d.entries:
sys.exit(f"❌ feed 抓取失败或为空: {feed_url}(bozo={d.bozo})")
return d.entries
def cmd_scan(args, manifest, base_dir):
html_dir = os.path.join(base_dir, "html")
os.makedirs(html_dir, exist_ok=True)
entries = fetch_feed(args.feed)
new_ids = []
for e in entries:
iid = str(issue_id_from_link(e.link))
if iid in manifest["issues"]:
continue # 已登记,跳过(去重)
content_html = e.content[0].value if getattr(e, "content", None) else ""
pubdate = format_pubdate(e)
fname = f"{iid}-{pubdate}.html"
fpath = os.path.join(html_dir, fname)
with open(fpath, "w", encoding="utf-8") as f:
f.write(content_html)
manifest["issues"][iid] = {
"id": iid,
"title": e.title,
"url": e.link,
"pub_date": pubdate,
"html_file": f"html/{fname}",
"fetched_at": now_iso(),
"extracted": False,
"article_count": 0,
"markdown_dir": None,
}
new_ids.append(iid)
print(f"🆕 新期 #{iid} {e.title} → {fname}({len(content_html)} bytes)", file=sys.stderr)
manifest["last_scanned"] = now_iso()
save_manifest(base_dir, manifest)
print(f"📊 扫描完成:新增 {len(new_ids)} 期,共 {len(manifest['issues'])} 期", file=sys.stderr)
return new_ids
def cmd_check(args, manifest):
entries = fetch_feed(args.feed)
known = set(manifest["issues"].keys())
new_ids = []
for e in entries:
iid = str(issue_id_from_link(e.link))
if iid not in known:
new_ids.append(iid)
if new_ids:
print("NEW:" + ",".join(new_ids))
else:
print("NONE")
return 0
# ---------------------------------------------------------------- HTML → 文章
class SREEntryParser(HTMLParser):
"""解析每期的 content:encoded HTML,提取所有 sreweekly-entry 文章条目。
页面结构(见实测):
<div class="sreweekly-sponsor-message">…赞助商…</div> ← 跳过
<div class="sreweekly-entry">
<div class="sreweekly-title"><a href="文章URL">标题</a></div>
<div class="sreweekly-description">
<p>简介…</p>
<blockquote><p>引用…</p></blockquote>
<p>&nbsp;&nbsp;<small>作者</small></p>
</div>
</div>
"""
def __init__(self):
super().__init__(convert_charrefs=True)
self.div_stack = [] # 当前打开中的 div class 栈
self.tag_stack = [] # 当前打开中的标签栈(用于识别 a/blockquote/small)
self.entries = [] # 提取结果: [{title, url, body, author}, ...]
self.cur = None
self.in_title = False
self.in_desc = False
self.in_small = False
self.quote_depth = 0
def _cur_div_class(self):
return self.div_stack[-1] if self.div_stack else None
def handle_starttag(self, tag, attrs):
attrs = dict(attrs)
self.tag_stack.append(tag)
cls = attrs.get("class", "")
if tag == "div":
self.div_stack.append(cls)
if "sreweekly-entry" in cls:
self.cur = {"title": "", "url": "", "body": [], "author": None}
elif "sreweekly-title" in cls and self.cur is not None:
self.in_title = True
elif "sreweekly-description" in cls and self.cur is not None:
self.in_desc = True
return
if self.cur is None:
return
if tag == "a" and self.in_title:
self.cur["url"] = attrs.get("href", "")
elif tag == "blockquote" and self.in_desc:
self.quote_depth += 1
if self.cur["body"] and self.cur["body"][-1] not in ("\n\n", "\n> ", "\n"):
self.cur["body"].append("\n\n")
self.cur["body"].append("\n> ")
elif tag == "small" and self.in_desc:
self.in_small = True
elif tag == "p" and self.in_desc and self.cur["body"] and self.quote_depth == 0:
# 段落间空行(避免在开头及 blockquote 内部加多余空行)
self.cur["body"].append("\n\n")
def handle_endtag(self, tag):
if self.tag_stack:
pop_idx = None
for i in range(len(self.tag_stack) - 1, -1, -1):
if self.tag_stack[i] == tag:
pop_idx = i
break
if pop_idx is not None:
del self.tag_stack[pop_idx:]
if tag == "small" and self.in_small:
self.in_small = False
if tag == "blockquote" and self.quote_depth > 0:
self.quote_depth -= 1
if self.cur is not None and self.in_desc:
self.cur["body"].append("\n")
if tag == "div":
if self.div_stack:
cls = self.div_stack.pop()
if "sreweekly-title" in cls:
self.in_title = False
elif "sreweekly-description" in cls:
self.in_desc = False
elif "sreweekly-entry" in cls:
self._finalize()
def handle_data(self, data):
if self.cur is None:
return
if self.in_title:
self.cur["title"] += data
elif self.in_desc:
if self.in_small:
self.cur["author"] = (self.cur["author"] or "") + data
else:
if not data.strip():
return # 标签间的纯空白(换行/缩进)跳过,段落结构由 p/blockquote 标记生成
if self.cur["body"] and self.cur["body"][-1] in ("\n> ", "\n\n", "\n"):
data = data.lstrip("\n\t \xa0")
self.cur["body"].append(data)
def _finalize(self):
if self.cur is None:
return
title = html.unescape(self.cur["title"]).strip()
body = "".join(self.cur["body"])
body = body.replace("\xa0", " ")
# 压缩多余的连续空行
body = re.sub(r"\n{3,}", "\n\n", body)
# 每行去掉首尾空白,但保留引用行前缀 ">"
lines = []
for ln in body.split("\n"):
stripped = ln.strip()
if stripped.startswith(">"):
lines.append("> " + stripped[1:].strip())
elif stripped:
lines.append(stripped)
else:
lines.append("")
body = "\n".join(lines).strip("\n")
author = html.unescape((self.cur["author"] or "").strip()) or None
self.entries.append({"title": title, "url": self.cur["url"], "body": body, "author": author})
self.cur = None
def parse_issue_html(html_path):
with open(html_path, "r", encoding="utf-8") as f:
p = SREEntryParser()
p.feed(f.read())
return p.entries
def slugify(title, max_len=70):
"""标题 → 文件名 slug:保留字母数字与 CJK,其余转连字符。"""
s = re.sub(r"[^\w\u4e00-\u9fff]+", "-", title.lower(), flags=re.UNICODE).strip("-")
s = re.sub(r"-{2,}", "-", s)
return s[:max_len].strip("-") or "untitled"
def extract_one(iid, entry, base_dir, force=False):
"""把某一期 HTML 解析为每篇文章一个 markdown 文件。返回 (文章数, 输出目录)。"""
md_root = os.path.join(base_dir, "markdown")
os.makedirs(md_root, exist_ok=True)
html_rel = entry.get("html_file")
if not html_rel:
sys.exit(f"❌ 期 #{iid} 没有 html_file 记录,请先扫描")
html_path = os.path.join(base_dir, html_rel)
if not os.path.exists(html_path):
sys.exit(f"❌ HTML 文件不存在: {html_path}")
if entry.get("extracted") and entry.get("article_count", 0) > 0 and not force:
print(f"⏭️ 期 #{iid} 已提取过({entry['article_count']} 篇),跳过(--force 可覆盖)", file=sys.stderr)
return entry.get("article_count", 0), entry.get("markdown_dir"), False
articles = parse_issue_html(html_path)
out_dir = os.path.join(md_root, str(iid))
os.makedirs(out_dir, exist_ok=True)
for idx, a in enumerate(articles, 1):
fname = f"{idx:02d}-{slugify(a['title'])}.md"
fpath = os.path.join(out_dir, fname)
content = render_markdown(a, entry)
with open(fpath, "w", encoding="utf-8") as f:
f.write(content)
entry["extracted"] = True
entry["article_count"] = len(articles)
entry["markdown_dir"] = f"markdown/{iid}"
entry["extracted_at"] = now_iso()
print(f"📄 期 #{iid} → {len(articles)} 篇文章 → {out_dir}", file=sys.stderr)
return len(articles), out_dir, True
def render_markdown(article, entry):
"""文章 → markdown 文件内容。"""
pub = entry.get("pub_date", "")
issue_label = entry.get("title", f"SRE Weekly Issue #{entry.get('id')}")
out = [f"# {article['title']}", ""]
out.append(f"- **期号**: {issue_label}({pub})")
out.append(f"- **作者**: {article['author'] or '—'}")
out.append(f"- **链接**: {article['url']}")
if article["body"]:
out += ["", "## 简介", "", article["body"]]
return "\n".join(out) + "\n"
def resolve_targets(target, manifest):
"""解析 --extract 的目标: ID / latest / pending / all → 期号列表。"""
issues = manifest["issues"]
if target == "latest":
ids = sorted(issues.keys(), key=lambda x: (isinstance(x, str), x))
# 数字优先按数值排,latest 取最大
nums = [i for i in issues if str(i).isdigit()]
return [str(max(int(n) for n in nums))] if nums else [ids[-1]]
if target == "pending":
return [str(i) for i, v in issues.items() if not v.get("extracted") or v.get("article_count", 0) == 0]
if target == "all":
return list(issues.keys())
t = str(target)
if t not in issues:
sys.exit(f"❌ 期 #{t} 不在 manifest 中,请先扫描")
return [t]
def cmd_extract(args, manifest, base_dir):
ids = resolve_targets(args.extract, manifest)
if not ids:
print("✅ 无待提取的期", file=sys.stderr)
return
for iid in ids:
extract_one(iid, manifest["issues"][str(iid)], base_dir, force=args.force)
save_manifest(base_dir, manifest)
# ---------------------------------------------------------------- 展示
def cmd_status(manifest):
issues = manifest["issues"]
print(f"📡 SRE Weekly 共 {len(issues)} 期(最后扫描: {manifest['last_scanned']})")
print()
header = f"{'期号':>6} {'日期':<12} {'HTML':<6} {'提取':<5} {'文章数':<6} 标题"
print(header)
print("-" * len(header))
for iid in sorted(issues.keys(), key=lambda x: int(x) if str(x).isdigit() else 0):
v = issues[iid]
html_ok = "✅" if v.get("html_file") else "—"
ext_ok = "✅" if v.get("extracted") and v.get("article_count", 0) > 0 else "—"
print(f"{iid:>6} {v.get('pub_date',''):<12} {html_ok:<6} {ext_ok:<5} {v.get('article_count',0):<6} {v.get('title','')}")
def cmd_issues(manifest):
for iid in sorted(manifest["issues"].keys(), key=lambda x: int(x) if str(x).isdigit() else 0):
v = manifest["issues"][iid]
print(f"{iid}\t{v.get('pub_date','')}\t{v.get('title','')}\t{v.get('url','')}")
# ---------------------------------------------------------------- main
def main():
ap = argparse.ArgumentParser(description="SRE Weekly 抓取 + 文章提取(供 n8n 定时调用)")
ap.add_argument("--feed", default=DEFAULT_FEED, help=f"RSS 地址(默认 {DEFAULT_FEED})")
ap.add_argument("--dir", default=DEFAULT_DIR, help=f"工作目录(默认 {DEFAULT_DIR})")
ap.add_argument("--check", action="store_true", help="只检查有无新期,输出 NEW:id,id 或 NONE")
ap.add_argument("--extract", metavar="ID|latest|pending|all", help="提取文章 → markdown")
ap.add_argument("--status", action="store_true", help="打印 manifest 状态表")
ap.add_argument("--issues", action="store_true", help="列出已登记的所有期")
ap.add_argument("--force", action="store_true", help="强制重新提取已提取过的期")
args = ap.parse_args()
base_dir = os.path.abspath(os.path.expanduser(args.dir))
os.makedirs(base_dir, exist_ok=True)
if args.extract or args.status or args.issues:
manifest = load_manifest(base_dir)
if args.extract:
cmd_extract(args, manifest, base_dir)
if args.status:
cmd_status(manifest)
if args.issues:
cmd_issues(manifest)
elif args.check:
manifest = load_manifest(base_dir)
cmd_check(args, manifest)
else:
manifest = load_manifest(base_dir)
cmd_scan(args, manifest, base_dir)
return 0
if __name__ == "__main__":
sys.exit(main())