From e80aa747e52e34e705960b9cc33511593b1a563f Mon Sep 17 00:00:00 2001 From: admin Date: Fri, 11 Sep 2026 06:20:11 +0800 Subject: [PATCH] update --- ...Automate — and Never Automate — with AI.md | 87 ++++ Cloud DevOps/Index.md | 3 +- Cloud DevOps/Uptime 99.9% vs 99.99%.md | 280 +++++++++++++ .../Hermes Agent接入阿里云百炼Token Plan.md | 27 ++ sreweekly/README.md | 91 ---- sreweekly/html/529-2026-08-10.html | 96 ----- sreweekly/html/530-2026-08-17.html | 94 ----- sreweekly/html/531-2026-08-24.html | 96 ----- sreweekly/html/532-2026-08-31.html | 82 ---- sreweekly/html/533-2026-09-07.html | 97 ----- sreweekly/scripts/sreweekly.py | 391 ------------------ today-articles-2026-09-01.html | 311 -------------- 12 files changed, 396 insertions(+), 1259 deletions(-) create mode 100644 Clippings/What SREs Should Automate — and Never Automate — with AI.md create mode 100644 Cloud DevOps/Uptime 99.9% vs 99.99%.md delete mode 100644 sreweekly/README.md delete mode 100644 sreweekly/html/529-2026-08-10.html delete mode 100644 sreweekly/html/530-2026-08-17.html delete mode 100644 sreweekly/html/531-2026-08-24.html delete mode 100644 sreweekly/html/532-2026-08-31.html delete mode 100644 sreweekly/html/533-2026-09-07.html delete mode 100644 sreweekly/scripts/sreweekly.py delete mode 100644 today-articles-2026-09-01.html diff --git a/Clippings/What SREs Should Automate — and Never Automate — with AI.md b/Clippings/What SREs Should Automate — and Never Automate — with AI.md new file mode 100644 index 00000000..be8d7a8a --- /dev/null +++ b/Clippings/What SREs Should Automate — and Never Automate — with AI.md @@ -0,0 +1,87 @@ +--- +title: "What SREs Should Automate — and Never Automate — with AI" +source: "https://hackernoon.com/what-sres-should-automate-and-never-automate-with-ai" +author: + - "[[Sai Joshitha Kathari]]" +published: 2026-08-27 +created: 2026-09-10 +description: "Automate based on impact and recoverability, not on whether the AI is technically capable of doing the task." +tags: + - "clippings" +--- +**Five key takeaways:** + +1. Automate based on impact and recoverability, not on whether the AI is technically capable of doing the task. +2. Alert triage, anomaly detection, incident summaries, capacity forecasting — these are the easy wins. Low risk, high value. +3. Production changes, incident command, security response, severity calls — keep a human's name on these. Always. +4. Reversibility and blast radius are better questions than "can the AI do this." +5. The goal isn't AI replacing engineers. It's AI clearing enough noise that engineers can actually think. + +--- + +I've sat through the version of this conversation that sounds like a vendor pitch — AI triages everything, drafts your runbooks, predicts outages before they happen, and nobody gets paged at 2 a.m. anymore. I've also watched the other version happen in real time: an automated remediation script restarts the wrong service, confidently, at 11 p.m., and a 20-minute blip turns into a four-hour outage while everyone tries to figure out why the "fix" made things worse. + +Both of those are real. AI is already inside SRE workflows whether or not anyone signed off on it — the question that actually matters is where it belongs, and where a human still needs to be the one holding the decision. + +None of what follows comes from a whitepaper. It's from watching what breaks when teams move too fast with this stuff, and what quietly gets better when they don't. + +## Reversibility and blast radius + +Here's the mental model I keep coming back to before automating anything: can you undo it, and how bad is it if you're wrong? + +Restarting a pod — reversible, low stakes. Deleting a database backup — not reversible, at all. Scaling a service up is easy to walk back. Silencing an alert for six hours is technically reversible too, except the six hours where something real happened and nobody saw it isn't something you get back. + +Blast radius is the other half of it, and it's not the same thing as severity. A misclassified low-priority alert costs a few wasted minutes. A misrouted sev-1 costs an hour of response time during an active outage, while the right team sits there not knowing they should be paged. And blast radius scales with what the action touches — one service versus a shared piece of infrastructure everything depends on, even when both look equally "minor" on paper. + +Anything with low reversibility and a wide blast radius shouldn't be running on autopilot. Anything reversible and contained is fair game. The stuff in between is where you actually need judgment — specifically, judgment from the people who'll be the ones on call when it goes sideways. + +Notice this framing never asks whether the AI *can* do something. It asks what happens if it's wrong. That's the more useful question, and it's the one most teams skip. + +## Where this actually works well + +**Alert noise.** This is the least controversial win there is. Somewhere between 30 and 60% of production alerts are noise by the time a human sees them — duplicates, transients, things that resolved themselves three minutes ago. AI grouping related alerts, suppressing known-flapping signals, correlating spikes with recent deploys — worst case, something gets mislabeled and a human still catches it. Low blast radius, fully reversible. This is exactly the profile you want. + +One catch: it only works well tuned to your environment, not a generic model. An alert that always fires right before a nightly batch job and clears itself a minute later is trivial to suppress — but only if the model actually knows about your batch schedule. Skip that step and you've just added a second layer of noise on top of the first. + +**First drafts of runbooks and postmortems.** Runbook rot is one of the oldest problems in this field. The doc that was accurate in 2022 is a landmine now — nobody updates it, an incident hits, someone follows it anyway, and step four references a service that got decommissioned eight months ago. AI is genuinely good at pulling together a first draft from past incidents, change logs, whatever documentation exists. Same for postmortems — a draft that someone who actually lived through the incident reviews before it goes out saves real hours. + +**Forecasting and anomaly detection.** This is pattern matching, and models are good at pattern matching. A holiday traffic spike that happens once a year gives engineers almost no reps to build intuition about — but a model trained across several years of that same spike has plenty. The important part: keep this as a recommendation a human acts on, not something that auto-provisions infrastructure on its own. The moment it stops informing a decision and starts making one, the blast radius changes. + +**Narrow, well-understood auto-remediation.** This one comes with real caveats, but it earns its place. A specific service that needs a restart when it hits a known stuck state, a queue that needs draining past a defined threshold — fine, if the failure class is precisely defined, tested, and low-blast-radius by design. And there has to be a circuit breaker. If the fix doesn't work within a set window, it stops and escalates instead of retrying forever on a wrong diagnosis. Automation that keeps trying the same broken fix is worse than doing nothing. + +## Where it doesn't belong + +**Severity calls.** Get this wrong either direction and it costs you. A real sev-1 marked as low pulls in the wrong people at the wrong urgency while an SLA clock runs. A minor issue marked critical drags a response team into something that didn't need them at 3 a.m. AI can surface context and flag patterns worth escalating — but the actual call needs a name attached, someone accountable for it. "The model said it was low severity" doesn't hold up in a postmortem. + +**Production changes without sign-off.** Config changes, scaling decisions, anything touching a database directly, restarts outside that narrow bounded case above — a human authorizes these. AI can prep the change, check it against known-good patterns, even simulate the blast radius. What it shouldn't do is decide the moment is right and pull the trigger itself. + +**Security incidents.** Different risk shape entirely. Miss something real and an active compromise sits there while the system waits for more confirmation. False-positive and you've locked out legitimate engineers mid-response. AI correlating logs to surface signal fast — genuinely useful. Containment and escalation decisions — that needs someone who can weigh legal and business context a model was never trained on. + +**Root cause, as a stated fact.** AI narrowing the search space by correlating deploy timing with metric shifts is useful groundwork. But writing "root cause: X" in a postmortem is a claim that shapes what the org fixes next and what it decides to ignore. Get that wrong because a correlation looked convincing, and the actual bug ships again next quarter. + +**Who to escalate to.** This is context a model just doesn't have — who's already underwater tonight, what else is on fire across the org, whether the responding engineer's confidence is real or performed. Escalation is a trust call as much as a technical one. + +## The thing nobody's measuring + +There's a slower cost that never shows up in a single incident review: engineers stop building intuition when AI absorbs all the routine reps. The edge cases are exactly where judgment matters most — and they're exactly the cases you need practice on the boring stuff to be ready for. A team leaning hard on automation can look great for a long stretch, right up until something shows up that doesn't match anything the model — or the team — has seen before. + +This isn't an argument against automating things. It's an argument for being honest about which reps you're willing to give away. + +## A few practices worth adopting + +- Decide, as a team, which categories of action AI can take alone versus which need a sign-off — decide this before an incident forces the question at 2 a.m. +- Keep an actual human accountable for anything irreversible. Not nominally "in the loop" — actually reviewing before it executes. +- Build in a circuit breaker for anything automated. If it doesn't work within a defined window, it escalates instead of retrying. +- Rotate people through the routine cases sometimes, even when AI could handle it, so the skill doesn't quietly disappear. +- Revisit the boundary as systems change. A failure class that was well-understood six months ago might not be anymore after an architecture shift. + +Skip this and you end up with automation debt, eroded skills, and a production system nobody fully understands anymore — which is a worse place to be than where you started. + +## Where this leaves things + +It's not really a question of whether to use AI. It's whether you're using it somewhere judgment genuinely isn't needed, or somewhere it is and you've just decided waiting for a human is too slow. + +One of those is a real force multiplier. The other is a liability with a delay timer on it. + +--- + diff --git a/Cloud DevOps/Index.md b/Cloud DevOps/Index.md index aceffc10..11cf6ea9 100644 --- a/Cloud DevOps/Index.md +++ b/Cloud DevOps/Index.md @@ -27,7 +27,7 @@ Product Trial/PoC Procedure ## **Deployment & Configuration** [[ITOM Cloud AWS Account Overview]] [[ITOM ESM Cloud Farm Information]] -[[Cloud DevOps/ITOM APM AppPluse Cloud Farm Information]] +[[ITOM APM AppPluse Cloud Farm Information]] Cloud Deployment Automation Product Provision Automation @@ -57,6 +57,7 @@ Audit Compliance Incident Management Change Management Service Health Page +Cloud DevOps ## **Customer Support** diff --git a/Cloud DevOps/Uptime 99.9% vs 99.99%.md b/Cloud DevOps/Uptime 99.9% vs 99.99%.md new file mode 100644 index 00000000..cd4bb60b --- /dev/null +++ b/Cloud DevOps/Uptime 99.9% vs 99.99%.md @@ -0,0 +1,280 @@ +#sla #uptime #downtime +# SLA Downtime 计算详解 + +## 99.9% 与 99.99% 的区别和实际应用 + +--- + +## 一、SLA 基本概念 + +**SLA(Service Level Agreement,服务级别协议)** 是指服务提供商与客户之间约定的服务可用性标准。 + +### SLA 的计算公式 + +``` +SLA % = (总时间 - 宕机时间) / 总时间 × 100% + +反推宕机时间: +宕机时间 = 总时间 × (1 - SLA%) +``` + +**关键点:** SLA 99.9% 意味着允许 0.1% 的宕机时间,99.99% 意味着允许 0.01% 的宕机时间。 + +--- + +## 二、按月计算的 Downtime 时间 + +### 1. 月度总时长(以30天为例) + +|单位|数值| +|---|---| +|总天数|30天| +|总小时数|30 × 24 = 720小时| +|总分钟数|30 × 24 × 60 = 43,200分钟| +|总秒数|30 × 24 × 60 × 60 = 2,592,000秒| + +> 注:实际应用中,有些公司按月份天数计算(28-31天),此处以30天为标准示例。 + +--- + +## 三、SLA 99.9% 详解 + +### 计算公式 + +``` +月度允许宕机时间 = 总时间 × (1 - 99.9%) + = 总时间 × 0.1% + = 总时间 × 0.001 +``` + +### 具体数值 + +| | | +|---|---| +|Daily|1m 26s| +|Weekly|10m 4.8s| +|Monthly|43m 50s| +|Quarterly|2h 11m 29s| +|Yearly|8h 45m 57s| + +### 月度详细计算示例 + +``` +99.9% SLA 月度允许宕机时间: + +方法一(按分钟): + 宕机时间 = 43,200分钟 × 0.001 = 43.2 分钟 + +方法二(按秒钟): + 宕机时间 = 2,592,000秒 × 0.001 = 2,592 秒 = 43分12秒 + +结论:一个月内最多可允许约 43分钟的宕机时间 +``` + +### 99.9% 的通俗理解 + +- **"三个九"** 是常见的说法 +- 一年内最长可宕机约 **8.76小时**(相当于一个工作日) +- 一月内最长可宕机约 **43分钟** +- 一周内最长可宕机约 **1分钟** + +--- + +## 四、SLA 99.99% 详解 + +### 计算公式 + +``` +月度允许宕机时间 = 总时间 × (1 - 99.99%) + = 总时间 × 0.01% + = 总时间 × 0.0001 +``` + +### 具体数值 + +| | | +|---|---| +|Daily|8.6s| +|Weekly|1m 0.48s| +|Monthly|4m 23s| +|Quarterly|13m 8.9s| +|Yearly|52m 36s| + +### 月度详细计算示例 + +``` +99.99% SLA 月度允许宕机时间: + +方法一(按分钟): + 宕机时间 = 43,200分钟 × 0.0001 = 4.32 分钟 = 4分19秒 + +方法二(按秒钟): + 宕机时间 = 2,592,000秒 × 0.0001 = 259.2 秒 ≈ 4分19秒 + +结论:一个月内最多可允许约 26秒的宕机时间 +``` + +### 99.99% 的通俗理解 + +- **"四个九"** 是高端服务的标准 +- 一年内最长可宕机约 **52.6分钟**(接近1小时) +- 一月内最长可宕机约 **4分23秒**(几乎不允许停机) +- 一周内最长可宕机约 **1分钟** +- 对于银行、支付系统等核心业务系统通常要求此级别 + +--- + + + + + +--- + +## 六、Downtime 的统计方法 + +### 1. 宕机时间的计算 + +#### **定义** + +宕机时间 = 服务不可用的时间段 + +#### **统计方式** + +``` +开始宕机时间:用户首次无法访问服务的时刻(UTC时间戳) +结束恢复时间:服务完全恢复可用的时刻(UTC时间戳) + +宕机时长 = 结束恢复时间 - 开始宕机时间 +``` + +#### **实际例子** + +``` +2024年9月5日 14:30:00 - 服务开始故障 +2024年9月5日 14:45:30 - 服务完全恢复 + +宕机时长 = 15分30秒 +``` + +### 2. 月度 SLA 计算 + +``` +月度SLA% = (总时间 - 月度累计宕机时长) / 总时间 × 100% + +例: + 总时间 = 43,200分钟(30天) + 月度累计宕机时长 = 65分钟(多次故障累加) + + 月度SLA% = (43,200 - 65) / 43,200 × 100% + = 43,135 / 43,200 × 100% + = 99.85% + + 结论:该月SLA未达到99.9%的要求 +``` + +### 3. 需要注意的统计细节 + +#### ✅ **应该计入宕机时间** + +- 完全不可用的时间段 +- 部分功能不可用且影响主要业务的时间段 +- 包括故障发现到完全恢复的全部时间 + +#### ❌ **通常不计入宕机时间** + +- 计划内维护时间(如提前通知的升级) +- 客户端问题导致的无法访问 +- 客户网络问题导致的连接失败 +- 用户操作错误导致的故障 + +### 4. 宕机统计的数据来源 + +``` +主要来源: +├─ 服务器日志(访问日志、错误日志) +├─ 监控告警(Prometheus、Datadog等) +├─ 负载均衡器日志 +├─ CDN日志 +├─ 用户反馈/投诉 +└─ 自动化探测(Ping、健康检查) +``` + +--- + +## 七、实际应用案例 + +### 案例1:电商平台(需要99.9%) + +``` +月度时间:43,200分钟 +允许宕机:43分钟 + +场景: +- 8月发生2次故障 + 第一次:15分钟 + 第二次:28分钟 + 累计:43分钟 + +评估:刚好达到99.9% SLA,但已是极限 +``` + +### 案例2:支付系统(需要99.99%) + +``` +月度时间:2,592,000秒 +允许宕机:259秒(约4分钟) + +场景: +- 一整月运行无任何宕机事件 +- 部分时间段响应缓慢但服务可用 +- 计划维护避开业务高峰 + +评估:完全满足99.99% SLA要求 +``` + +### 案例3:内部系统(需要99.5%) + +``` +月度时间:43,200分钟 +允许宕机:216分钟(3.6小时) + +场景: +- 一个月内宕机3小时 +- 仍在容许范围内 + +评估:满足99.5% SLA要求 +``` + +--- + + + +## 九、最佳实践 + +### 💡 监控和告警 + +- 使用专业监控工具(Prometheus、Grafana等) +- 设置实时告警机制 +- 记录所有故障时间和原因 + +### 💡 故障恢复 + +- 制定明确的恢复流程(RTO/RPO) +- 设置自动故障转移(如主备切换) +- 定期进行故障演习 + +### 💡 数据统计 + +- 使用UTC时间戳,避免时区混淆 +- 保存完整的历史日志(建议至少保留1年) +- 定期生成SLA报告,对外透明化 + +### 💡 承诺管理 + +- 清楚标注哪些维护不计入SLA +- 对客户设定合理预期 +- 提供赔偿机制(宕机补偿) + +--- + +**更新日期:** 2024年 **适用范围:** SLA管理、性能评估、架构设计 \ No newline at end of file diff --git a/knowledgebase/Hermes Agent接入阿里云百炼Token Plan.md b/knowledgebase/Hermes Agent接入阿里云百炼Token Plan.md index cc63778f..86d9d5a8 100644 --- a/knowledgebase/Hermes Agent接入阿里云百炼Token Plan.md +++ b/knowledgebase/Hermes Agent接入阿里云百炼Token Plan.md @@ -36,6 +36,33 @@ source ~/.bashrc # 如果使用 zsh,改为 source ~/.zshrc hermes --version ``` +## 配置环境变量 + +将 `OPENAI_API_KEY` 环境变量设置为 Token Plan 个人版专属 API Key。 + +1. 在终端中执行以下命令,查看默认 Shell 类型。 + +```bash +echo $SHELL +``` + +2. 根据 Shell 类型设置环境变量: + ```bash + # 将 YOUR_API_KEY 替换为 Token Plan 个人版 API Key + echo 'export OPENAI_API_KEY="YOUR_API_KEY"' >> ~/.zshrc + ``` + ```bash + # 将 YOUR_API_KEY 替换为 Token Plan 个人版 API Key + echo 'export OPENAI_API_KEY="YOUR_API_KEY"' >> ~/.bash_profile + ``` +3. 执行以下命令使环境变量生效。 + ```bash + source ~/.zshrc + ``` + ```bash + source ~/.bash_profile + ``` + ## 配置接入凭证 通过 `hermes config set` 命令配置接入参数,根据所选方案填入对应的 Base URL 和 API Key: diff --git a/sreweekly/README.md b/sreweekly/README.md deleted file mode 100644 index 611b620e..00000000 --- a/sreweekly/README.md +++ /dev/null @@ -1,91 +0,0 @@ -# SRE Weekly Digest - -抓取 [SRE Weekly](https://sreweekly.com) 每期内容,把其中分享的文章逐篇导出为 Markdown, -供后续阅读/筛选/邮件分发使用。脚本由 **n8n 定时调用**,通过不同参数驱动;脚本自身**不发送邮件**。 - -- 脚本: `scripts/sreweekly.py`(单文件,仅依赖 `feedparser`) -- 依赖: `pip3 install feedparser` - -## 目录结构 - -``` -sreweekly/ -├── manifest.json # 每期状态:html 是否保存、文章是否已提取、文章数 -├── html/<期号>-<日期>.html # 每期原始 HTML(feed 的 content:encoded 原样保存) -└── markdown/<期号>/ # 每篇文章一个 .md 文件 - ├── 01-<标题>.md - └── ... -``` - -## 命令一览 - -```bash -python3 scripts/sreweekly.py # 扫描 feed,新增期保存 HTML + 登记 manifest -python3 scripts/sreweekly.py --check # 检查有无新期 → 输出 NEW:533,532 或 NONE -python3 scripts/sreweekly.py --extract 533 # 提取指定期文章 → markdown/533/ -python3 scripts/sreweekly.py --extract latest # 提取最新一期 -python3 scripts/sreweekly.py --extract pending # 提取所有尚未提取的期(每日流程推荐) -python3 scripts/sreweekly.py --extract all # 提取全部(已提取过需加 --force 覆盖) -python3 scripts/sreweekly.py --status # 打印状态表 -python3 scripts/sreweekly.py --issues # 列出所有已登记期 - -# 可选参数 ---feed # RSS 地址(默认 https://sreweekly.com/feed/) ---dir # 工作目录(默认本目录) ---force # 强制重新提取 -``` - -所有命令**幂等**:重复扫描不重复入库,重复提取自动跳过已提取的期。 - -## n8n 每日流程(推荐 3 步) - -该刊每周一更新,建议每天定时跑一次,有新期才继续: - -1. **Execute Command(扫描)** - ``` - python3 /Users/weishen/Workspace/nexus/Hermes/xingzhi/sreweekly/scripts/sreweekly.py --dir /Users/weishen/Workspace/nexus/Hermes/xingzhi/sreweekly - ``` -2. **Execute Command(提取未处理期)** - ``` - python3 .../sreweekly.py --dir ... --extract pending - ``` - (或先用 `--check` 判断 `NEW:`/`NONE` 再分支,`--extract pending` 本身也天然幂等,空跑无害) - -3. **后续节点(邮件)**:读取 `markdown/<新期号>/` 下的 .md 文件, - 用 n8n 的 Email / 自定义 AgentMail 节点发送(示例发件收件:star-agent@agentmail.to → billyshen@163.com)。 - -参考输出(`--check`,可直接作为 n8n IF 条件): -``` -NEW:533,532 ← 有新期 -NONE ← 无更新 -``` - -## manifest.json 示例 - -```json -{ - "feed": "https://sreweekly.com/feed/", - "last_scanned": "2026-09-10T07:40:50", - "issues": { - "533": { - "id": "533", - "title": "SRE Weekly Issue #533", - "url": "https://sreweekly.com/sre-weekly-issue-533/", - "pub_date": "2026-09-07", - "html_file": "html/533-2026-09-07.html", - "fetched_at": "2026-09-10T07:40:50", - "extracted": true, - "article_count": 8, - "markdown_dir": "markdown/533", - "extracted_at": "2026-09-10T07:41:10" - } - } -} -``` - -## 说明 - -- SRE Weekly 的 feed 已在 `content:encoded` 字段携带每期完整 HTML,无需再抓每期页面。 -- 页面结构:每篇分享文章是 `
`(标题链接 + 简介 + 作者), - 赞助商段落(`sreweekly-sponsor-message`)自动跳过。 -- 提取后每篇 md 内容:标题、期号/日期、作者、原文链接、简介(含 blockquote 引用)。 \ No newline at end of file diff --git a/sreweekly/html/529-2026-08-10.html b/sreweekly/html/529-2026-08-10.html deleted file mode 100644 index a1538320..00000000 --- a/sreweekly/html/529-2026-08-10.html +++ /dev/null @@ -1,96 +0,0 @@ -

- -
-

A message from our sponsor, Planetscale:

-

Your on-call rotation shouldn’t double as your database’s HA strategy. PlanetScale databases ship with a primary and two replicas across three AZs, automated failover, and a 99.999% multi-region SLA. Postgres and Vitess available in AWS and GCP.

-

→ Get started with PlanetScale for just $5/mo

-
- - -
-
- -
-

It’s not enough to define an incident process. You have to spin up and maintain an entire incident management program.

-

  Brent Chapman

-
-
- - - -
- -
-

Honeycomb pulls back the curtain a bit to delve into how LLM agents change the way their product is used, and how their query patterns differ from humans’. It’s especially interesting that increasing agent usage has not correlated with decreasing human usage.

-

  Austin Parker — Honeycomb

-
-
- - - -
- -
-

I like the approach here, especially measuring both the positive and negative outcomes.

-
-

Good SRE practice is about evidence, not enthusiasm.

-
-

   Neel Shah — DZone

-
-
- - - -
- -
-

The kernel’s route cache: a hidden reliability killer. This is a really intriguing case of self-sustaining impact.

-

  Ray Chen — Railway

-
-
- - - -
- -
-
-

…we rely on formal verification, and this is how consensus algorithms are built today. We define a model that we can mathematically prove to be correct, and then we… translate this perfect, platonic thing into code.

-
-

  TW Lim — Antithesis

-
-
- - - -
- -
-

I love that this starts with the user. Monitor what matters to your users, and alert on what you can action.

-

  Omar Ghader

-
-
- - - -
- -
-

First time I’ve heard of systemd’s PrivateTmp feature. Neat!

-

  Chris Siebenmann

-
-
- - - -
- -
-
-

I thought it would be a useful exercise to brainstorm some of the differences in focus between what I’ll call the traditional view of reliability, and the resilience engineering view.

-
-

It’s short (just a table), but it definitely made me think.

-

  Lorin Hochstein

-
-
-
\ No newline at end of file diff --git a/sreweekly/html/530-2026-08-17.html b/sreweekly/html/530-2026-08-17.html deleted file mode 100644 index 09cae6e7..00000000 --- a/sreweekly/html/530-2026-08-17.html +++ /dev/null @@ -1,94 +0,0 @@ -

-
-

A message from our sponsor, Planetscale:

-

Your on-call rotation shouldn’t double as your database’s HA strategy. PlanetScale databases ship with a primary and two replicas across three AZs, automated failover, and a 99.999% multi-region SLA. Postgres and Vitess available in AWS and GCP.

-

→ Get started with PlanetScale for just $5/mo

-
- - -
-
- -
-

We may improve velocity by handing off tasks to LLM agents, but can that impact resilience?

-

  Courtney Nash — Resilience in Software Foundation

-
-
- - - -
- -
-

Fatigue and burn-out are reliability risks. Fatigue and burn-out are reliability risks. I champion this idea in my SRE practice constantly, and I hope you do too.

-

  Brent Chapman

-
-
- - - -
- -
-

A fun read on how to build control planes for large-scale systems, with some great tidbits on the inner workings of EC2 and Aurora DSQL.

-

  Zak van der Merw

-
-
- - - -
- -
-

A harrowing incident story underlining the importance of expertise and experience.

-

  Hamed Silatani — Uptime Labs

-
-
- - - -
- -
-

An SRE comes to terms with the way LLM agents are changing our field: what works well, what still requires human involvement, and what the future may look like.

-

  Bill Duncan

-
-
- - - -
- -
-
-

Recent outages at Tailscale, jsDelivr, ServiceNow, and IPinfo show the same failure pattern: certificate automation broke quietly, while the expiry date kept moving closer.

-
-

Bonus: they include links to several write-ups of related incidents.

-

  TokenTimer

-
-
- - - -
- -
-
-

Our solution treats infrastructure state as a traversable graph and lets a pathfinding algorithm discover recovery sequences at runtime.

-
-

Whoa, cool trick!

-

  Pragya Mehta and Sai Samant — Stripe

-
-
- - - -
- -
-

Their event-oriented system was based on Google Pub/Sub with its 99.95% SLA, but their own SLA was 99.99%. To resolve that, they moved toward an active-active architecture, load-balancing across 2 message brokers.

-

There’s an interactive simulation of their algorithm midway through that’s fun to play with!

-

  Patrick Hamann and Mike Fisher — incident.io

-
-
-
\ No newline at end of file diff --git a/sreweekly/html/531-2026-08-24.html b/sreweekly/html/531-2026-08-24.html deleted file mode 100644 index c9efbeb3..00000000 --- a/sreweekly/html/531-2026-08-24.html +++ /dev/null @@ -1,96 +0,0 @@ -

- -
-

A message from our sponsor, Planetscale:

-

PlanetScale Metal runs Postgres and Vitess on dedicated NVMe inside AWS and GCP. Get data center speed next to your app, with unlimited IOPS and no throttling. Teams routinely see a 70% drop in p99 and p95 latency after migrating.

-

→ See the benchmarks

-
- - -
-
- -
-

Celebrating heroes in incident response can incentivize further heroics. That can prevent the kind of growth that will improve incident response overall.

-

  Brent Chapman

-
-
- - - -
- -
-

What might happen when we quickly adopt LLMs and make sweeping changes in our complex systems?

-

  Fred Hebert

-
-
- - - -
- -
-
-

Most teams log, but log badly: wrong severity levels, no trace IDs, inconsistent fields, and logs siloed from traces.

-
-

  Ashwini Dave — DZone

-
-
- - - -
- -
-
-

If you had to explain to a neighbour why your organisation is so safe, and generally works well, what would you say?

-
-

It’s all about people. I really enjoyed the quote from Charles Billings on principles for automation.

-

  Steven Shorrock

-
-
- - - -
- -
-

Type conversion in aviation involves an experienced pilot training on a new kind of aircraft. This article draws a parallel to transitioning to a new job as an SRE.

-

  Bill Duncan

-
-
- - - -
- -
-

Some big names in this Q&A, and they share a couple of delicious morsels.

-

  Sam Salter — Uptime Labs, with John Allspaw and Beth Adele Long

-
-
- - - -
- -
-

A super-engaging deep-dive.

-
-

This investigation is a useful reminder: running boring technology in a non-standard way is a risk.

-
-

  Alex Chan — Tailscale

-
-
- - - -
- -
-

A handy guide on topology constraints in Kubernetes, with a worked example.

-

  Andre Newman — Gremlin

-
-
-
\ No newline at end of file diff --git a/sreweekly/html/532-2026-08-31.html b/sreweekly/html/532-2026-08-31.html deleted file mode 100644 index c9bbc11e..00000000 --- a/sreweekly/html/532-2026-08-31.html +++ /dev/null @@ -1,82 +0,0 @@ -

- -
-

A message from our sponsor, Planetscale:

-

PlanetScale Metal runs Postgres and Vitess on dedicated NVMe inside AWS and GCP. Get data center speed next to your app, with unlimited IOPS and no throttling. Teams routinely see a 70% drop in p99 and p95 latency after migrating.

-

→ See the benchmarks

-
- - -
-
- -
-

Need another team to do something fast? Just use this one weird trick: declare an incident! This article explains why the obvious solution (gating incident declaration) isn’t a good idea.

-

  Brent Chapman

-
-
- - - -
- -
-

Ethics are relative, right? This article is full of genuinely useful tips and framings.

-

  Thomas A. Limoncelli — ACM Queue

-
-
- - - -
- -
-

Whoa. It’s been quite a few years since my last run-in with an overfull conntrack table, and this is a fun new twist.

-

  Jorrick Sleijster — Adyen

-
-
- - - -
- -
-
-

For eight years I ran SRE for a storage system measured in exabytes. The dashboard I checked every morning shrank to seven numbers. Here they are.

-
-

  Sridhar Rajarao

-
-
- - - -
- -
-
-

Traditional observability monitors execution. LLM observability must monitor behavior.

-
-

  Barnadeep Bhowmik

-
-
- - - -
- -
-

This one has a lot of great detail on how their approaches to quota management failed and how they iterated.

-

  Dhyanam Vaidya, Prathamesh Deshpande, and Mike Ma — Uber

-
-
- - - -
- -
-

This article uses Voyager 1, whose engineers just shut down another instrument to conserve its steadily-decaying power, as an extended analogy for graceful degradation.

-

  Robert Barron

-
-
-
\ No newline at end of file diff --git a/sreweekly/html/533-2026-09-07.html b/sreweekly/html/533-2026-09-07.html deleted file mode 100644 index 44ad1772..00000000 --- a/sreweekly/html/533-2026-09-07.html +++ /dev/null @@ -1,97 +0,0 @@ -

- -
-

A message from our sponsor, Planetscale:

-

Most database incidents start with one expensive query, not the database being down. PlanetScale gives SRE teams high-availability Postgres and MySQL with automated failover, query insights, and Database Traffic Control to stop runaway queries before they page you.

-

→ Explore PlanetScale

-
- - -
-
- -
-

What can you do to shorten the time to detect an incident? Some great ideas in here, especially monitoring your company’s main web page for a sudden uptick in traffic.

-

  Brent Chapman

-
-
- - - -
- -
-

What an interesting incident! I recommend reading Azure’s write-up before reading Lorin’s excellent analysis.

-

  Lorin Hochstein

-
-
- - - -
- -
-
-

Distributed databases rarely fail in the clean, isolated ways described by component diagrams. They fail through timing gaps, stale metadata, ambiguous ownership, retry storms, incompatible health decisions, and overlapping maintenance activity.

-
-

   Varsha Ganesh — DZone

-
-
- - - -
- -
-

I love this concept of a “political incident”:

-
-

The subject was political incidents, by which I mean the ones where the severity arrives before the impact assessment does.

-
-

And ouch, I felt this bit:

-
-

You have spent forty minutes of the incident on the severity field.

-
-

  Tim Irving

-
-
- - - -
- -
-

Where can you safely use LLM agents, versus when you should keep things in human hands? This one has some good criteria to consider.

-

  Sai Joshitha Kathari — HackerNoon

-
-
- - - -
- -
-

I learned a lot about Git while reading this one. Speeding up Git clones in CI may not seem important, but it will when you’re trying to roll out a fix during an incident.

-

  Mike Thompson and Daniel Esponda — Datadog

-
-
- - - -
- -
-

Switching from their custom-written autoscaler to the new off-the-shelf option made sense, but it wasn’t a simple drop-in replacement.

-

  Samuel Yeboah, Francesco Di Chiara and Mingliang Liu — Netflix

-
-
- - - -
- -
-

Can we replace human code review with LLM-based reviews? This article lays out what an LLM can’t replicate, and I’d argue that these are the pieces that matter most for reliability.

-

  John Allspaw — Adaptive Capacity Labs

-
-
-
\ No newline at end of file diff --git a/sreweekly/scripts/sreweekly.py b/sreweekly/scripts/sreweekly.py deleted file mode 100644 index 4762ce36..00000000 --- a/sreweekly/scripts/sreweekly.py +++ /dev/null @@ -1,391 +0,0 @@ -#!/usr/bin/env python3 -# -*- coding: utf-8 -*- -""" -SRE Weekly 抓取与文章提取脚本 - -模仿 blogwatcher-daily 的单文件风格:feedparser 抓取 → HTML 落盘 → 逐篇文章导出 Markdown, -并用 manifest.json 记录每期的状态(是否已存 HTML / 是否已提取文章)。 - -设计给 n8n 定时任务调用,通过不同参数驱动,脚本自身不发送邮件。 - -目录结构(默认工作目录 ~/Workspace/nexus/Hermes/xingzhi/sreweekly/): - sreweekly/ - ├── manifest.json # 状态记录(每期: html 是否保存、文章是否提取、文章数) - ├── html/-<日期>.html # 每期原始 HTML(feed 的 content:encoded 原样保存) - └── markdown// # 每期文章的 markdown 目录(每篇一个文件) - ├── 01-<标题slug>.md - └── 02-<标题slug>.md - -用法: - python3 sreweekly.py # 扫描 feed,新增期保存为 HTML 并登记 manifest - python3 sreweekly.py --check # 只检查有无新期,输出 NEW:533,532 或 NONE(供 n8n 分支) - python3 sreweekly.py --extract ID # 提取指定期(如 533)的文章 → markdown - python3 sreweekly.py --extract latest # 提取最新一期 - python3 sreweekly.py --extract pending # 提取所有尚未提取的期(n8n 每日流程推荐) - python3 sreweekly.py --extract all # 提取全部(含已提取的,需配合 --force 覆盖) - python3 sreweekly.py --status # 打印 manifest 状态表 - python3 sreweekly.py --issues # 列出已登记的所有期 - -选项: - --feed URL RSS 地址(默认 https://sreweekly.com/feed/) - --dir PATH 工作目录(默认 ~/Workspace/nexus/Hermes/xingzhi/sreweekly/) - --force 重新提取已提取过的期(覆盖旧文件) -""" -from html.parser import HTMLParser -import argparse, html, json, os, re, sys, time, urllib.request - -try: - import feedparser -except ImportError: - sys.exit("缺少依赖: 请先执行 pip3 install feedparser") - -DEFAULT_FEED = "https://sreweekly.com/feed/" -DEFAULT_DIR = os.path.expanduser("~/Workspace/nexus/Hermes/xingzhi/sreweekly") - -# ---------------------------------------------------------------- manifest - -def load_manifest(base_dir): - path = os.path.join(base_dir, "manifest.json") - if os.path.exists(path): - try: - with open(path, "r", encoding="utf-8") as f: - return json.load(f) - except (json.JSONDecodeError, OSError) as e: - print(f"⚠️ manifest.json 读取失败({e}),按空库处理", file=sys.stderr) - return {"feed": DEFAULT_FEED, "last_scanned": None, "issues": {}} - -def save_manifest(base_dir, manifest): - path = os.path.join(base_dir, "manifest.json") - with open(path, "w", encoding="utf-8") as f: - json.dump(manifest, f, ensure_ascii=False, indent=2) - return path - -def issue_id_from_link(link): - """从文章链接提取期号: sre-weekly-issue-533 → 533;失败则用最后一段路径。""" - m = re.search(r"sre-weekly-issue-(\d+)", link or "") - if m: - return int(m.group(1)) - m = re.search(r"/([^/]+?)/?$", (link or "").rstrip("/")) - return m.group(1) if m else "unknown" - -def format_pubdate(entry): - """优先用 feed 自带发布日期(YYYY-MM-DD),没有则用当天。""" - if getattr(entry, "published_parsed", None): - return time.strftime("%Y-%m-%d", entry.published_parsed) - return time.strftime("%Y-%m-%d") - -def now_iso(): - return time.strftime("%Y-%m-%dT%H:%M:%S") - -# ---------------------------------------------------------------- feed 扫描 - -def fetch_feed(feed_url): - d = feedparser.parse(feed_url) - if not d.entries: - sys.exit(f"❌ feed 抓取失败或为空: {feed_url}(bozo={d.bozo})") - return d.entries - -def cmd_scan(args, manifest, base_dir): - html_dir = os.path.join(base_dir, "html") - os.makedirs(html_dir, exist_ok=True) - entries = fetch_feed(args.feed) - new_ids = [] - for e in entries: - iid = str(issue_id_from_link(e.link)) - if iid in manifest["issues"]: - continue # 已登记,跳过(去重) - content_html = e.content[0].value if getattr(e, "content", None) else "" - pubdate = format_pubdate(e) - fname = f"{iid}-{pubdate}.html" - fpath = os.path.join(html_dir, fname) - with open(fpath, "w", encoding="utf-8") as f: - f.write(content_html) - manifest["issues"][iid] = { - "id": iid, - "title": e.title, - "url": e.link, - "pub_date": pubdate, - "html_file": f"html/{fname}", - "fetched_at": now_iso(), - "extracted": False, - "article_count": 0, - "markdown_dir": None, - } - new_ids.append(iid) - print(f"🆕 新期 #{iid} {e.title} → {fname}({len(content_html)} bytes)", file=sys.stderr) - manifest["last_scanned"] = now_iso() - save_manifest(base_dir, manifest) - print(f"📊 扫描完成:新增 {len(new_ids)} 期,共 {len(manifest['issues'])} 期", file=sys.stderr) - return new_ids - -def cmd_check(args, manifest): - entries = fetch_feed(args.feed) - known = set(manifest["issues"].keys()) - new_ids = [] - for e in entries: - iid = str(issue_id_from_link(e.link)) - if iid not in known: - new_ids.append(iid) - if new_ids: - print("NEW:" + ",".join(new_ids)) - else: - print("NONE") - return 0 - -# ---------------------------------------------------------------- HTML → 文章 - -class SREEntryParser(HTMLParser): - """解析每期的 content:encoded HTML,提取所有 sreweekly-entry 文章条目。 - - 页面结构(见实测): -
…赞助商…
← 跳过 -
- -
-

简介…

-

引用…

-

  作者

-
-
- """ - - def __init__(self): - super().__init__(convert_charrefs=True) - self.div_stack = [] # 当前打开中的 div class 栈 - self.tag_stack = [] # 当前打开中的标签栈(用于识别 a/blockquote/small) - self.entries = [] # 提取结果: [{title, url, body, author}, ...] - self.cur = None - self.in_title = False - self.in_desc = False - self.in_small = False - self.quote_depth = 0 - - def _cur_div_class(self): - return self.div_stack[-1] if self.div_stack else None - - def handle_starttag(self, tag, attrs): - attrs = dict(attrs) - self.tag_stack.append(tag) - cls = attrs.get("class", "") - if tag == "div": - self.div_stack.append(cls) - if "sreweekly-entry" in cls: - self.cur = {"title": "", "url": "", "body": [], "author": None} - elif "sreweekly-title" in cls and self.cur is not None: - self.in_title = True - elif "sreweekly-description" in cls and self.cur is not None: - self.in_desc = True - return - if self.cur is None: - return - if tag == "a" and self.in_title: - self.cur["url"] = attrs.get("href", "") - elif tag == "blockquote" and self.in_desc: - self.quote_depth += 1 - if self.cur["body"] and self.cur["body"][-1] not in ("\n\n", "\n> ", "\n"): - self.cur["body"].append("\n\n") - self.cur["body"].append("\n> ") - elif tag == "small" and self.in_desc: - self.in_small = True - elif tag == "p" and self.in_desc and self.cur["body"] and self.quote_depth == 0: - # 段落间空行(避免在开头及 blockquote 内部加多余空行) - self.cur["body"].append("\n\n") - - def handle_endtag(self, tag): - if self.tag_stack: - pop_idx = None - for i in range(len(self.tag_stack) - 1, -1, -1): - if self.tag_stack[i] == tag: - pop_idx = i - break - if pop_idx is not None: - del self.tag_stack[pop_idx:] - if tag == "small" and self.in_small: - self.in_small = False - if tag == "blockquote" and self.quote_depth > 0: - self.quote_depth -= 1 - if self.cur is not None and self.in_desc: - self.cur["body"].append("\n") - if tag == "div": - if self.div_stack: - cls = self.div_stack.pop() - if "sreweekly-title" in cls: - self.in_title = False - elif "sreweekly-description" in cls: - self.in_desc = False - elif "sreweekly-entry" in cls: - self._finalize() - - def handle_data(self, data): - if self.cur is None: - return - if self.in_title: - self.cur["title"] += data - elif self.in_desc: - if self.in_small: - self.cur["author"] = (self.cur["author"] or "") + data - else: - if not data.strip(): - return # 标签间的纯空白(换行/缩进)跳过,段落结构由 p/blockquote 标记生成 - if self.cur["body"] and self.cur["body"][-1] in ("\n> ", "\n\n", "\n"): - data = data.lstrip("\n\t \xa0") - self.cur["body"].append(data) - - def _finalize(self): - if self.cur is None: - return - title = html.unescape(self.cur["title"]).strip() - body = "".join(self.cur["body"]) - body = body.replace("\xa0", " ") - # 压缩多余的连续空行 - body = re.sub(r"\n{3,}", "\n\n", body) - # 每行去掉首尾空白,但保留引用行前缀 ">" - lines = [] - for ln in body.split("\n"): - stripped = ln.strip() - if stripped.startswith(">"): - lines.append("> " + stripped[1:].strip()) - elif stripped: - lines.append(stripped) - else: - lines.append("") - body = "\n".join(lines).strip("\n") - author = html.unescape((self.cur["author"] or "").strip()) or None - self.entries.append({"title": title, "url": self.cur["url"], "body": body, "author": author}) - self.cur = None - -def parse_issue_html(html_path): - with open(html_path, "r", encoding="utf-8") as f: - p = SREEntryParser() - p.feed(f.read()) - return p.entries - -def slugify(title, max_len=70): - """标题 → 文件名 slug:保留字母数字与 CJK,其余转连字符。""" - s = re.sub(r"[^\w\u4e00-\u9fff]+", "-", title.lower(), flags=re.UNICODE).strip("-") - s = re.sub(r"-{2,}", "-", s) - return s[:max_len].strip("-") or "untitled" - -def extract_one(iid, entry, base_dir, force=False): - """把某一期 HTML 解析为每篇文章一个 markdown 文件。返回 (文章数, 输出目录)。""" - md_root = os.path.join(base_dir, "markdown") - os.makedirs(md_root, exist_ok=True) - html_rel = entry.get("html_file") - if not html_rel: - sys.exit(f"❌ 期 #{iid} 没有 html_file 记录,请先扫描") - html_path = os.path.join(base_dir, html_rel) - if not os.path.exists(html_path): - sys.exit(f"❌ HTML 文件不存在: {html_path}") - if entry.get("extracted") and entry.get("article_count", 0) > 0 and not force: - print(f"⏭️ 期 #{iid} 已提取过({entry['article_count']} 篇),跳过(--force 可覆盖)", file=sys.stderr) - return entry.get("article_count", 0), entry.get("markdown_dir"), False - - articles = parse_issue_html(html_path) - out_dir = os.path.join(md_root, str(iid)) - os.makedirs(out_dir, exist_ok=True) - for idx, a in enumerate(articles, 1): - fname = f"{idx:02d}-{slugify(a['title'])}.md" - fpath = os.path.join(out_dir, fname) - content = render_markdown(a, entry) - with open(fpath, "w", encoding="utf-8") as f: - f.write(content) - entry["extracted"] = True - entry["article_count"] = len(articles) - entry["markdown_dir"] = f"markdown/{iid}" - entry["extracted_at"] = now_iso() - print(f"📄 期 #{iid} → {len(articles)} 篇文章 → {out_dir}", file=sys.stderr) - return len(articles), out_dir, True - -def render_markdown(article, entry): - """文章 → markdown 文件内容。""" - pub = entry.get("pub_date", "") - issue_label = entry.get("title", f"SRE Weekly Issue #{entry.get('id')}") - out = [f"# {article['title']}", ""] - out.append(f"- **期号**: {issue_label}({pub})") - out.append(f"- **作者**: {article['author'] or '—'}") - out.append(f"- **链接**: {article['url']}") - if article["body"]: - out += ["", "## 简介", "", article["body"]] - return "\n".join(out) + "\n" - -def resolve_targets(target, manifest): - """解析 --extract 的目标: ID / latest / pending / all → 期号列表。""" - issues = manifest["issues"] - if target == "latest": - ids = sorted(issues.keys(), key=lambda x: (isinstance(x, str), x)) - # 数字优先按数值排,latest 取最大 - nums = [i for i in issues if str(i).isdigit()] - return [str(max(int(n) for n in nums))] if nums else [ids[-1]] - if target == "pending": - return [str(i) for i, v in issues.items() if not v.get("extracted") or v.get("article_count", 0) == 0] - if target == "all": - return list(issues.keys()) - t = str(target) - if t not in issues: - sys.exit(f"❌ 期 #{t} 不在 manifest 中,请先扫描") - return [t] - -def cmd_extract(args, manifest, base_dir): - ids = resolve_targets(args.extract, manifest) - if not ids: - print("✅ 无待提取的期", file=sys.stderr) - return - for iid in ids: - extract_one(iid, manifest["issues"][str(iid)], base_dir, force=args.force) - save_manifest(base_dir, manifest) - -# ---------------------------------------------------------------- 展示 - -def cmd_status(manifest): - issues = manifest["issues"] - print(f"📡 SRE Weekly 共 {len(issues)} 期(最后扫描: {manifest['last_scanned']})") - print() - header = f"{'期号':>6} {'日期':<12} {'HTML':<6} {'提取':<5} {'文章数':<6} 标题" - print(header) - print("-" * len(header)) - for iid in sorted(issues.keys(), key=lambda x: int(x) if str(x).isdigit() else 0): - v = issues[iid] - html_ok = "✅" if v.get("html_file") else "—" - ext_ok = "✅" if v.get("extracted") and v.get("article_count", 0) > 0 else "—" - print(f"{iid:>6} {v.get('pub_date',''):<12} {html_ok:<6} {ext_ok:<5} {v.get('article_count',0):<6} {v.get('title','')}") - -def cmd_issues(manifest): - for iid in sorted(manifest["issues"].keys(), key=lambda x: int(x) if str(x).isdigit() else 0): - v = manifest["issues"][iid] - print(f"{iid}\t{v.get('pub_date','')}\t{v.get('title','')}\t{v.get('url','')}") - -# ---------------------------------------------------------------- main - -def main(): - ap = argparse.ArgumentParser(description="SRE Weekly 抓取 + 文章提取(供 n8n 定时调用)") - ap.add_argument("--feed", default=DEFAULT_FEED, help=f"RSS 地址(默认 {DEFAULT_FEED})") - ap.add_argument("--dir", default=DEFAULT_DIR, help=f"工作目录(默认 {DEFAULT_DIR})") - ap.add_argument("--check", action="store_true", help="只检查有无新期,输出 NEW:id,id 或 NONE") - ap.add_argument("--extract", metavar="ID|latest|pending|all", help="提取文章 → markdown") - ap.add_argument("--status", action="store_true", help="打印 manifest 状态表") - ap.add_argument("--issues", action="store_true", help="列出已登记的所有期") - ap.add_argument("--force", action="store_true", help="强制重新提取已提取过的期") - args = ap.parse_args() - - base_dir = os.path.abspath(os.path.expanduser(args.dir)) - os.makedirs(base_dir, exist_ok=True) - - if args.extract or args.status or args.issues: - manifest = load_manifest(base_dir) - if args.extract: - cmd_extract(args, manifest, base_dir) - if args.status: - cmd_status(manifest) - if args.issues: - cmd_issues(manifest) - elif args.check: - manifest = load_manifest(base_dir) - cmd_check(args, manifest) - else: - manifest = load_manifest(base_dir) - cmd_scan(args, manifest, base_dir) - - return 0 - -if __name__ == "__main__": - sys.exit(main()) \ No newline at end of file diff --git a/today-articles-2026-09-01.html b/today-articles-2026-09-01.html deleted file mode 100644 index f936c698..00000000 --- a/today-articles-2026-09-01.html +++ /dev/null @@ -1,311 +0,0 @@ - - -Blogwatcher Daily — 2026-09-01 - - -

Blogwatcher Daily — 2026-09-01

-

共 83 篇文章。


-

【Engadget - Technology News & Expert Reviews】

- -

【Reuters - YouTube】

- -

【AI懶人報 - YouTube】

-
    -
  • EP398 - Zeabur資安大事件,使用者該如何保護自己,防範未然?AI時代的自保指南! -
    🚀 特別感謝贊助本集節目由 VoAI 絕好聲創 提供技術支援。🎤 VoAI 提供最有「台灣味」的 AI 聲音,支援情感語音、台式口音,甚至能一鍵生成虛擬人!🎁 AI懶人報聽眾專屬優惠:👉 輸入優惠碼 AILRB26 立享 95 折!立刻體驗:https://www.voai.ai/━━━━━━━━━━━━━━━━━━━━⏱ 章節00:00 有人打開信箱,看到一筆不是他花的帳單01:04 第一段|這
    -
  • -
  • EP397 - 依賴 AI 記憶是在埋地雷?Claude Memory 2.0 深度解析,別讓錯誤的自動化毀了你的開發專案! -
    🚀 特別感謝贊助本集節目由 VoAI 絕好聲創 提供技術支援。🎤 VoAI 提供最有「台灣味」的 AI 聲音,支援情感語音、台式口音,甚至能一鍵生成虛擬人!🎁 AI懶人報聽眾專屬優惠:👉 輸入優惠碼 AILRB26 立享 95 折!立刻體驗:https://www.voai.ai/Claude 的記憶功能升級了!這次我們不只聊它怎麼變聰明,還要提醒大家,在寫程式時過度依賴 AI 記憶,反而可能害你
    -
  • -
-

【TuTu生活志 - YouTube】

-
    -
  • 尼泊尔灾情信息公示差点就满分?!我用AI优化官方通报图,顺手做了个工具... -
    8月26日,喜马拉雅山区一场严重的跨境冰川灾害席卷尼泊尔与中国边境,造成惨重伤亡。本期我聚焦尼泊尔受灾一侧,深入官方灾情通报,毕竟...我试着用AI优化这些数据呈现,竟意外做出了一套完整工具,衍生出一长串意想不到的内容。真实与美化,在AI时代究竟哪个更重要?看完你会重新思考这个问题。TuTu的网站:https://tutulifestyle.com (聚合了各类信息,包括linktree、哪里买的
    -
  • -
-

【WSJ.com: World News】

- -

【SRE WEEKLY】

-
    -
  • SRE Weekly Issue #532 -
    View on sreweekly.com A message from our sponsor, Planetscale: PlanetScale Metal runs Postgres and Vitess on dedicated NVMe inside AWS and GCP. Get data center speed next to your app, with unlimited I
    -
  • -
-

【AWS DevOps & Developer Productivity Blog】

- -

【Slashdot】

- -

【小众软件】

-
    -
  • OpenClaw 2.0 正式发布:史上规模最大的一次更新:933 位贡献者,16000 个拉取请求 -
    在停更7周以后,OpenClaw 2.0 终于发布,这次更新被开发者称为「迄今为止 OpenClaw 历史上规模最大的一次更新。此次更新由 933 位贡献者共同完成,其中包括 569 位首次贡献者,并包含了超过 16,000 个 pull request。」 OpenClaw 2.0 就是 Open
    -
  • -
  • Steam 12TB 游戏数据泄漏:横跨 10 年,大量未公开内容曝光 -
    Steam 最近爆出了一次规模惊人的游戏数据泄漏:超过 12TB、横跨 2003~2013 年的大量游戏数据被公开。@Appinn 泄漏的内容是 Steam2,里面有《Portal 2》《Left 4 Dead》《CS》等游戏的未公开早期版本,包括废弃地图、模型、对白和游戏机制,甚至还有被取消的 F
    -
  • -
  • 内网穿透,本可以不用那么麻烦 -
    很多人都想随时随地地远程控制在家或办公室的电脑,或者出门在外也能访问 NAS 里的文件、公司内部的 OA 系统,甚至小主机、路由器、打印机等设备。但由于没有公网 IP,这些设备都处于内网,无法直接访问。 内网穿透就可以解决这个问题,它能让你不受局域网的限制,在任何地点,用手机或电脑,也能访问家里或公
    -
  • -
  • 今天,Chrome 应用商店将彻底移除所有 Manifest V2 扩展,趁现在如何备份? -
    这一天终于到了。 根据 Manifest V2 支持时间表,2026 年 8 月 31 日,谷歌将从 Chrome 应用商店中移除所有剩余的 Manifest V2 扩展程序。 注意:安装在 Chrome 138 或更早版本上的 Manifest V2 扩展程序将保留安装状态,但无法接收任何更新,并
    -
  • -
  • 小米17 Pro自带语音输入法会屏蔽脏话 -
    根据酷安网友截图显示,其使用了小米17 Pro自带输入法的语音输入功能,在碰到脏话的时候会提示:已屏蔽敏感内容 该网友还表示:“即使把相关 AI 的模型功能全部关掉,也是无法关闭这个屏蔽功能的。我不知道这个屏蔽功能在这个输入法到底起什么作用。日常口语带一些脏话就会瞬间秒屏蔽。还有一些敏感类的词都会屏
    -
  • -
  • 发现频道:10款大家发现的好评软件[2026年第35期] -
    最近10日,来自小众软件论坛的发现频道的热门排行榜,由系统自动生成,直接列出来: 序号 主题 1️⃣ 一个双击即用的 Windows 软件管理工具,不仅列出所有已安装软件,还能彻底卸载 + 清空注册表 + 删除残留文件夹/文件,不依赖任何运行时。 2️⃣ 【开发者自荐】AI Story – 一句话一
    -
  • -
  • NVIDIA 即将告别 Windows 10:最后一款 Game Ready 驱动 10 月发布 -
    还在用 Windows 10 + NVIDIA 显卡打游戏的同学注意了:按照 NVIDIA 此前公布的支持计划,最后一款支持 Windows 10 的 Game Ready 驱动将在今年 10 月发布。 NVIDIA 早在 2025 年 7 月就公布了这份计划: 10 月之后会怎样? NVIDIA
    -
  • -
-

【异次元软件世界】

-
    -
  • UPDF:AI 加持的全能 PDF 编辑器 (用过才知为啥大家都超级推荐!功能超全) -
    现在的工作与学习几乎离不开 PDF 文件:论文、发票、合同、课件、方案……但现实是,很多人仍在为“PDF 无法编辑”、“格式转换麻烦”、“批注不方便”而无比头疼。 如果你需要寻找一款 功能强、速度快、支持 AI 的 PDF 编辑器,那「UPDF」很可能就是那个最佳答案!它不仅支持电脑和手机 APP,还提供网页版,功能覆盖学习办公中与 PDF 相关的所有需求。重点是,UPDF 界面非常精美,操作现代
    -
  • -
  • 白嫖 68 元 Token 额度!OpenSquilla 开源 Agent 送羊毛,多款模型随便调 -
    玩过 AI Agent 的朋友应该都有同一个感受:这玩意儿好玩好用,但一个稍微复杂点的任务跑下来,多步规划、工具调用、长记忆加载,每一步都在疯狂吃 token,最要命的是很多人图省事,不管任务简单还是复杂,一股脑全都选旗舰贵模型,结果就是账单爆炸。明明一个摘要、格式整理这种轻量活儿,交…… 「 前往查看原文.... 」异次元首页 | 微信公众号 | 关注微博 | 软件精选 | 软件激活码折扣
    -
  • -
  • 微软 Windows Server 2025 LTSC 最新正式版官方 ISO 镜像下载 - 服务器系统 MSDN 原版 -
    随着 Windows 11 新版本发布之后,微软终于也正式推出了 Windows Server 2025 服务器操作系统了!这是一款专门面向企业和服务提供商的先进可靠的服务器系统。 Windows Server 2025 LTSC 是微软迄今为止成功的企业级服务器系统,它继承了 2022 的优点并基于 Win11 作为内核,进一步提升了安全性、性能、稳定性和灵活性,引入了许多服务器相关创新。无论是
    -
  • -
-

【阿榮福利味 - 免費軟體下載】

-
    -
  • [正版購買] Wondershare UniConverter 17.6.0 中文版 - 萬用影片轉檔下載編輯軟體 -
    萬用影片轉檔下載編輯軟體 - Wondershare UniConverter,全能的影片工具組,包含了影片轉檔、影片下載、影片壓縮、影片編輯、影片播放、光碟燒錄、合併轉檔、螢幕錄影等多項基本功能,不只這樣,還有自動剪裁、智慧剪裁、大量刪除片頭片尾、移除或新增浮水印、字幕編輯器、AI 畫像、背景移除器、影片防抖動、圖片轉檔器、GIF 檔案製作等加值功能。(阿榮福利味) ★軟體報價:請「填寫資料」或
    -
  • -
  • Supremo 5.0.3.3883 免安裝中文版 - 遠端遙控軟體 類似 Teamviewer -
    類似「Teamviewer」的遠端遙控軟體 - Supremo,開啟時按「Start」取得一組帳號(ID)跟密碼(Password),如果是要讓人遙控你的電腦,就直接把這組帳號密碼給對方即可,對方連線進來時還會徵求你的同意(Ask for confirmation before allow access),如果是要遙控他人電腦,請切換到「Connect to a remote computer」頁
    -
  • -
  • SMPlayer 26.8.29 免安裝中文版 - 影片播放自由軟體 -
    影片播放軟體 - SMPlayer,是採用MPlayer核心的自由軟體,支援大多數格式的影片及DVD播放,其特色就是會自動記住每個影片的播放位置、字幕、音量...等設定,下次開啟時會依照該影片的記憶設定來播放。(阿榮) 下載連結→ https://www.azofreeware.com/p/smplayer.html 官方網站:SMPlayer 軟體性質:自由軟體(免費) 介面語言:繁體中文(含多
    -
  • -
  • [正版購買] SEO PowerSuite R103.6 - SEO 輔助工具 -
    SEO 輔助工具 - SEO PowerSuite,內建多種網路搜尋引擎最佳化軟體,足以應付你的 SEO 工作需求,包含:SEO 分析工具、關鍵字研究工具、反向連結檢查器、內容編輯器、PPC 廣告最佳化等功能,憑藉其直觀的使用者界面和豐富的專業級功能,是適合新手和專家的 SEO 工具。(阿榮福利味) ※購買請先填寫報價資料:https://www.azofreeware.com/p/quote.h
    -
  • -
  • RisohEditor 6.1.8 免安裝中文版 - 免費的 Windows 資源編輯器 -
    免費的 Win32 資源編輯器 - RisohEditor,類似「ResEdit」、「Resource Hacker」的軟體中文化工具,主要用途是直接編輯 Windows 程式中的資源,例如 RS、RES、EXE、DLL 等檔案內的介面與資源,支援 Unicode 萬國碼。(阿榮福利味) 下載連結→ https://www.azofreeware.com/p/risoheditor.html 官方
    -
  • -
  • NordVPN 2026.09.01 中文版 - 支援多平台的 VPN 軟體 -
    支援多平台的 VPN 軟體 - NordVPN,提供安全與速度兼具的服務,針對流量建立加密通道,涵蓋全球將近兩百個位置,提供八千多台伺服器,能夠防範網路監控、網路威脅、危險的無線網路,可以安裝於家中路由器上,做全方位的家庭網路防護。(阿榮福利味) 下載連結→ https://www.azofreeware.com/p/nordvpn.html 官方網站:Nord Security 軟體性質:共享軟
    -
  • -
-

【BBC News 中文 - YouTube】

-
    -
  • 前線報導:BBC走訪中尼邊境被洪災抹去的小鎮- BBC News 中文 #尼泊爾 #中國 #西藏 -
    尼泊爾上週遭遇毀滅性洪災,死亡人數已達939人,另有3925人失蹤。在重災區拉蘇瓦縣,許多繁榮的城鎮如今面目全非。BBC記者阿扎德·莫希里來到夏布盧貝西,這裡緊鄰中國西藏吉隆,曾經生活在這裡的社區已幾乎不留任何痕跡,取而代之的是大量乾硬的泥土和瓦礫。救援行動仍在持續,隧道內仍有工人受困;而倖存者回憶起洪水來襲的那一天,仍難掩恐懼與悲痛。BBC News 中文: https://www.bbc.co
    -
  • -
  • 尼泊爾與西藏洪災超4200人失蹤 家屬稱「幾天沒有合眼」- BBC News 中文 -
    8月26日,一座災難性泥石流重創尼泊爾北部與中國西藏地區。惡劣天氣和持續上升的水位增加救援難度,兩國救援人員正爭分奪秒搜尋失蹤者。截止週一,尼泊爾的死亡人數已增至903人,另有4200人失蹤。中國方面公佈的西藏死亡人數為16人,另有546人失蹤。BBC News 中文: https://www.bbc.com/zhongwen訂閱BBC News 中文 YouTube:http://bit.ly/
    -
  • -
  • 前線報導:尼泊爾洪災逾900人罹難 民眾絕望尋找失蹤親人- BBC News 中文 -
    尼泊爾上周遭遇毀滅性洪災,已有超過900人死亡,仍有4200人失蹤。有許多人在街頭一張張查看尋找到的遺體照片,絕望地尋找失聯的親人;也有數百人被困隧道,至今生死未卜。BBC記者阿扎德·莫希里從加德滿都出發前往受災嚴重的拉蘇瓦地區。她報導說,儘管救援人員持續在瓦礫與泥濘中搜尋生還者,持續降雨和山區道路受阻,讓救援與物資運送面臨巨大挑戰。BBC News 中文: https://www.bbc.com
    -
  • -
-

【理想生活实验室】

-
    -
  • 今日消费资讯:《牛来》将延长上映到 10 月 4 日、张凌赫成为巴黎欧莱雅美发代言人 -
    《牛来》将延长上映到 10 月 4 日8 月 31 日消息,从 8 月 5 日开始公映的《牛来》密钥确认延期,影片将延长上映到 10 月 4 日结束。截止到 8 月 31 日,影片累计票房已达 5979.8 万元。% ARABICA 科威特贾贝尔购物中心“得来速”店开业8 月 29 日,% ARABICA 在科威特贾贝尔·艾哈迈德城(Jaber Al Ahmad City)的贾贝尔购物中心(Jab
    -
  • -
  • 我们开箱了 Bose QC 消噪耳机 II,时隔三年的产品更新,它改变了什么? -
    8 月 14 日,Bose 推出了旗下新款消噪耳机 QuietComfort 消噪耳机 II(以下简称“QC 消噪耳机 II”),这是 2023 年 9 月 21 日在上海发布的 QuietComfort 消噪耳机的续作,这样时隔三年,这款定位主力市场的产品也更新了。延伸阅读:和 QuietComfort 消噪耳机一起发布的另外两款耳机都在 2025 年完成了升级,它们是 QuietComfort
    -
  • -
-

【零度解说 - YouTube】

- -

【Coursera - YouTube】

-
    -
  • Build a Personal Brand That Gets You Hired -
    Before an interview even begins, employers may already have an impression based on what's online. Learn how to build a professional digital footprint, strengthen a personal brand, and create an online
    -
  • -
-

【灵姐说AI | Ling Talk AI - YouTube】

-
    -
  • 5000万美元爱情大瓜该不该给?我追问4个AI,答案越问越不一样 -
    假如你有88亿美元,毕生挚爱问你要5000万美元,你给不给?我把这道爱情考题交给Claude、ChatGPT、Gemini和豆包。首问,四家都说给;继续加条件,回答开始分岔。从“孙哥和甜姐”的话题出发,这期把争议抽离成假设题,聊的是AI如何理解爱情、金钱与关系边界,不验证或裁定真人私生活。这一期一起看:• 刚交往三个月、不给就分手,AI会怎样回应?• 同样一笔钱,被理解成交易还是家庭保障,会有什么
    -
  • -
-

【TEDx Talks - YouTube】

- -

【Tech With Tim - YouTube】

- -