SRE weekly 所有文章

This commit is contained in:
2026-09-12 17:23:01 +08:00
parent 409b40ddcb
commit af7633f9dc
8486 changed files with 4489990 additions and 7 deletions

View File

@@ -0,0 +1,24 @@
# 2025 VOID Incident Management Survey
- **期号**: SRE Weekly Issue #494(2025-09-14)
- **作者**: Courtney Nash
- **链接**: https://www.linkedin.com/posts/void-incidents_void-survey-activity-7371542977523064832-xNOP
## 简介
Courtney Nash over at The VOID has launched an in-depth survey of incident management practices in tech. Please consider taking the time to fill out this survey. We all stand to benefit hugely from the information it will gather.
## 正文
Announcing the 2025 VOID Incident Management Survey
I’ve been running The VOID for over 4 years now, and during that time I’ve read thousands of incident reports and talked to hundreds of people about their experiences dealing with incidents. And the one thing I strongly suspect (but can’t yet definitively assert) is that most people think that the way their company handles incidents, frankly, isn’t ideal.
But we don’t really know what most companies are doing because no one has ever surveyed the industry to find out.
So that’s what I’m doing, and I need your help for this to work. If I get enough responses—it should take only 15-20 minutes out of your day—I’ll have the data to help make managing incidents better.
Incident responders deserve to be supported in their work. They should have the time, resources, and autonomy to bring their depth of experience to bear without fear of being blamed or ending up a scapegoat. Their critical work suffers when they feel burned out and undervalued, and I aim to show the industry how current beliefs and practices can change for the better.
Take the survey here: [https://lnkd.in/gaxS92QU](https://www.linkedin.com/redir/redirect?url=https%3A%2F%2Flnkd%2Ein%2FgaxS92QU&urlhash=asgV&trk=public_post-text)
Athetai Rajendram Ludmilla Sivanathan Senthooran Sri Amanda A. Michael B. Megan Leach Chasity Parker 👀

View File

@@ -0,0 +1,19 @@
# The VOID Newsletter, September 2025
- **期号**: SRE Weekly Issue #494(2025-09-14)
- **作者**: Courtney Nash
- **链接**: https://www.thevoid.community/newsletter
## 简介
Speaking of The VOID, the first bit of the September issue of the VOID Newsletter stood out to me:
> Back in June, Salesforce had what appeared to be a pretty painful Heroku outage. About a month later, tech blogger Gergely Orosz posted about the incident on BlueSky. I’m bringing this up now because I’ve had over a month to chew on his commentary and I’m still mad about it. As someone who deals in reading public incident reports as a primary feature of my work, I find nothing more infuriating than people arm chair quarterbacking other organizations’ incidents and presuming they actually have any idea_ what really happened_.
As it happens, I also commented on the similarity of Salesforce’s incident to a Datadog incident from the past in issue 482.
I’m with Courtney Nash: we really have to be careful how we opine on public incident write-ups. Not only is it important to avoid blame and hindsight bias, but we also need to be careful not to disincentivize companies from posting public incident write-ups. I highly recommend clicking through to read Courtney’s full analysis.
## 正文
> ⚠️ 抓取失败:trafilatura returned empty

View File

@@ -0,0 +1,231 @@
# What are Error Budgets? A Guide to Managing Reliability
- **期号**: SRE Weekly Issue #494(2025-09-14)
- **作者**: Nawaz Dhandala — OneUpime
- **链接**: https://oneuptime.com/blog/post/2025-09-03-what-are-error-budgets/view
## 简介
> This guide explains what error budgets are, how to manage them effectively, what to look out for, and how they differ from SLOs.
Includes sections on potential pitfalls, real-world examples, and impact on company culture.
## 正文
In the world of software engineering, reliability isn't just about keeping systems running, it's about making smart trade-offs between stability and innovation. Enter error budgets: the secret weapon that helps teams like ours at OneUptime maintain high availability while shipping features at breakneck speed.
But what exactly are error budgets? How do you manage them? And how do they differ from those familiar SLOs? Let's dive in.
## What Are Error Budgets?
An error budget is the acceptable amount of downtime or errors your service can experience before it violates your Service Level Objectives (SLOs). It's essentially permission to fail- within limits.
Think of it this way: If your SLO is 99.9% uptime (allowing for about 8.77 hours of downtime per year), your error budget is that 8.77 hours. As long as you stay within that budget, you're meeting your reliability targets.
"Error budgets give teams the freedom to innovate while maintaining accountability for reliability."
The concept originated at Google as part of their Site Reliability Engineering (SRE) practices.
Instead of aiming for 100% uptime (which is often unrealistic and expensive), SRE teams define acceptable failure rates and use error budgets to track them.
## How Error Budgets Differ from SLOs
While SLOs and error budgets are closely related, they're not the same thing:
- **SLOs** define what "good" looks like (e.g., "99.9% of requests should succeed")
- **Error budgets** quantify how much "bad" is acceptable (e.g., "You can have 0.1% of requests fail")
SLOs are your target, error budgets are your tolerance for missing that target. If you exhaust your error budget, you're violating your SLOs.
The key difference is in mindset: SLOs are about aspiration, error budgets are about reality. They acknowledge that perfect reliability is impossible and expensive, so they give teams breathing room to experiment and learn.
## How to Manage Error Budgets
Managing error budgets effectively requires a combination of measurement, monitoring, and decision-making. Here's how to do it right:
### 1. Define Clear SLOs First
Before you can have an error budget, you need SLOs. Start by identifying what matters most to your users:
- What metrics indicate success for your service?
- What level of performance would make users abandon your product?
- What's the minimum viable reliability for your business?
For example, at OneUptime, our SLOs focus on incident detection time and resolution time, because those directly impact our customers' ability to respond to outages.
### 2. Calculate Your Error Budget
Once you have SLOs, calculating the error budget is straightforward. The formula below shows how to derive your acceptable failure threshold from your reliability target. This calculation is fundamental because it transforms your aspirational SLO into a concrete, measurable budget that your team can track and manage.
```
# Error Budget Formula
# ====================
# This calculates the maximum allowable downtime or failure rate
# that your service can experience while still meeting your SLO.
# Formula:
Error_Budget = 100% - SLO_Target
# Example calculation for a 99.9% availability SLO:
# -------------------------------------------------
SLO_Target = 99.9 # Your service level objective (percentage)
Error_Budget = 100 - SLO_Target # = 0.1%
# Convert to time (assuming 30-day month):
minutes_per_month = 30 * 24 * 60 # = 43,200 minutes
allowed_downtime_minutes = minutes_per_month * (Error_Budget / 100)
# Result: 43.2 minutes of allowed downtime per month
# Key insight: The higher your SLO, the smaller your error budget
# 99% SLO -> 1% budget -> ~7.2 hours/month
# 99.9% SLO -> 0.1% budget -> ~43 minutes/month
# 99.99% SLO -> 0.01% budget -> ~4.3 minutes/month
```
For a 99.9% SLO:
- Error Budget = 0.1% (or 8.77 hours per year)
- This means you can have 0.1% of requests fail or 8.77 hours of downtime
### 3. Track Burn Rate
Burn rate is how quickly you're consuming your error budget. It's calculated as shown below. Understanding burn rate is critical because it tells you not just whether you're failing, but how fast you're approaching your reliability limits. A high burn rate is an early warning signal that demands immediate attention before your error budget is fully exhausted.
```
# Burn Rate Formula
# =================
# This measures how fast you're consuming your error budget.
# It's the ratio of actual failure rate to your allowed failure rate.
# Formula:
Burn_Rate = Actual_Error_Rate / Allowed_Error_Rate
# Expanded formula:
Burn_Rate = (Actual_Errors / Total_Requests) / (Error_Budget / 100)
# Example calculation:
# --------------------
# Given: 99.9% SLO (0.1% error budget), 1,000,000 requests, 1,500 errors
total_requests = 1_000_000
actual_errors = 1_500
error_budget_percent = 0.1 # From our 99.9% SLO
actual_error_rate = actual_errors / total_requests # = 0.0015 = 0.15%
allowed_error_rate = error_budget_percent / 100 # = 0.001 = 0.1%
burn_rate = actual_error_rate / allowed_error_rate # = 1.5
# Interpretation Guide:
# ---------------------
# Burn Rate = 1.0 -> Sustainable: consuming budget at expected rate
# Burn Rate > 1.0 -> Warning: consuming budget faster than planned
# Burn Rate < 1.0 -> Healthy: room for innovation and experimentation
#
# In our example (burn_rate = 1.5):
# We're consuming budget 50% faster than sustainable.
# At this rate, a 30-day budget would be exhausted in 20 days.
# Alert thresholds (common practice):
# - Burn Rate > 2.0 for 1 hour -> Critical alert (fast burn)
# - Burn Rate > 1.0 for 6 hours -> Warning alert (slow burn)
```
A burn rate of 1 means you're consuming your error budget at the expected rate. Anything above 1 means you're on track to exhaust it early.
### 4. Set Up Alerts
Create alerts for different burn rate thresholds:
- 50% of budget consumed (warning)
- 80% of budget consumed (critical)
- 100% of budget consumed (emergency)
These alerts should trigger discussions about whether to slow down feature releases or invest in reliability improvements.
### 5. Make Data-Driven Decisions
Use your error budget to inform development decisions:
- **Green budget** : Full speed ahead on new features
- **Yellow budget** : Proceed with caution, consider reliability impact
- **Red budget** : Focus on stability, pause non-critical features
This creates a natural feedback loop where reliability becomes everyone's responsibility, not just the ops team's.
## What to Look Out For
While error budgets are powerful, they're not without pitfalls. Here are the common traps to avoid:
### 1. Setting Unrealistic SLOs
If your SLOs are too aggressive (like 99.999% uptime), your error budget becomes tiny. This leads to constant alerts and stifles innovation. Start conservative and adjust based on real user needs.
### 2. Ignoring Seasonal Variations
Error budgets should account for different usage patterns. If your service sees 10x traffic during peak hours, your error budget might burn faster then. Consider time-based budgeting or adjusting SLOs seasonally.
### 3. Focusing Only on Availability
Error budgets aren't just about uptime. Consider other dimensions like latency, error rates, and data freshness. A service might be "up" but still violating user expectations.
### 4. Not Communicating Across Teams
Error budgets work best when everyone understands them. Developers need to know how their code affects reliability, and product managers need to understand the trade-offs between features and stability.
### 5. Using Error Budgets as Excuses
"Don't worry about that bug, we have error budget!" is not the right attitude. Error budgets are for planned risk-taking, not sloppiness. Always strive to improve reliability even when you have budget left.
### 6. Forgetting About Recovery
When you do exhaust your error budget, have a plan to recover. This might involve rolling back recent changes, implementing circuit breakers, or temporarily reducing functionality.
## Real-World Examples
Let's look at how error budgets work in practice:
### E-commerce Platform
- SLO: 99.95% availability during business hours
- Error Budget: 0.05% (about 22 minutes per month)
- When budget is low: Team prioritizes stability over new checkout features
### API Service
- SLO: 99.9% success rate for all endpoints
- Error Budget: 0.1% (43.2 minutes of errors per month)
- Burn rate monitoring helps identify problematic endpoints early
### Mobile App
- SLO: 99% of users can complete key flows without errors
- Error Budget: 1% (adjusted for user count)
- Used to balance A/B testing with stability
## The Cultural Impact
Beyond the mechanics, error budgets change how teams think about reliability. They shift the conversation from "How do we prevent all failures?" to "How much failure can we tolerate, and how do we learn from it?"
This mindset encourages:
- Blameless postmortems
- Automated testing and deployment
- Proactive monitoring and alerting
- Cross-functional collaboration
## Getting Started
Ready to implement error budgets? Start small:
1. Pick one service or endpoint
2. Define realistic SLOs based on user needs
3. Calculate your error budget
4. Set up basic monitoring and alerting
5. Use the data to inform your next sprint planning
Remember, error budgets are a tool, not a goal. The real objective is delivering reliable software that delights users while enabling your business to grow.
"Error budgets don't eliminate failures- they make failures productive."
**About OneUptime:** We're building the next generation of observability tools to make SRE practices like error budgets accessible to every engineering team. Learn more about how we can help you implement error budgets and SLOs at [OneUptime.com](https://oneuptime.com).
**Related Reading:**
### Nawaz Dhandala
Author
@nawazdhandala • Sep 03, 2025 •
### Help improve this post
Every OneUptime blog post is open source. Found a typo, an inaccuracy, or have a clearer way to explain something? Anyone can contribute — your edits make this post better for everyone who reads it next.

View File

@@ -0,0 +1,180 @@
# Observability for the Invisible: Tracing Message Drops in Kafka Pipelines
- **期号**: SRE Weekly Issue #494(2025-09-14)
- **作者**: Prakash Wagle — DZone
- **链接**: https://dzone.com/articles/observability-tracing-message-drops-kafka-pipelines
## 简介
> This article explores how backend engineers and DevOps teams can detect, debug, and prevent message loss in Kafka-based streaming pipelines using tools like OpenTelemetry, Fluent Bit, Jaeger, and dead-letter queues.
## 正文
-
![](https://dz2cdn1.dzone.com/themes/dz20/images/dz-postarticle.svg) [Post an Article](https://dzone.com/content/article/post.html)
-
[Manage My Drafts](https://dzone.com)
# Observability for the Invisible: Tracing Message Drops in Kafka Pipelines
Kafka lag lies. Use Fluent Bit, OpenTelemetry, DLQs, and trace IDs to expose missing messages and harden observability in event-driven pipelines.
Join the DZone community and get the full member experience.
[Join For Free](https://dzone.com/static/registration.html)
When an event drops silently in a distributed system, it is not a bug, it is an architectural blind spot. In high-scale messaging platforms, particularly those serving real-time APIs like WhatsApp Business or IoT command chains, telemetry failures are often mistaken for application errors. But the root cause lies deeper: observability gaps in event streams.
This article explores how backend engineers and DevOps teams can detect, debug, and prevent message loss in [Kafka-based streaming pipelines](https://dzone.com/articles/building-robust-real-time-data-pipelines-with-pyth) using tools like [OpenTelemetry](https://dzone.com/articles/guide-to-opentelemetry-intro-to-observability), [Fluent Bit](https://dzone.com/articles/install-fluent-bit-from-source-part-two), [Jaeger](https://dzone.com/articles/tracing-with-opentelemetry-and-jaeger), and dead-letter queues. If your distributed messaging system handles millions of events, this guide outlines exactly how to make those events accountable.
It also briefly surfaces key production concerns around multi-tenant access control, communications security, and the potential for ML-powered anomaly detection in observability pipelines, areas that are foundational to modern, large-scale infrastructure.
## When Kafka Metrics Lie: Lag ≠ Delivery
Say your Kafka consumer group reports zero lag. The pipeline appears stable. But a downstream service dashboard shows stale or missing data.
```
bash
$ kafka-consumer-groups.sh --bootstrap-server kafka:9092 --describe --group user-analytics
TOPIC         PARTITION  CURRENT-OFFSET  LOG-END-OFFSET  LAG
user-events   0          24567           24567           0
```
No visible lag. No alert. Still, payloads are missing.
This is where many distributed pipelines fail, not with crashes, but with silent non-events. The system appears healthy but is semantically broken. Without deep traceability, your metrics are just performance theater.
One senior infra engineer described this gap as the "perfect Kafka illusion", systems that deliver bytes, but not outcomes. This is where architecture, not tooling, must evolve.
## Add Trace Context to Kafka Consumers for Debug Visibility
OpenTelemetry spans must be embedded at the event-handling layer of your consumer logic. This is the only way to establish causal visibility between an incoming payload and its downstream effect (or lack thereof).
Java
```
ConsumerRecords<String, String> records = consumer.poll(Duration.ofMillis(100));
for (ConsumerRecord<String, String> record : records) {
    Span span = tracer.spanBuilder("consume_user_event")
        .setAttribute("topic", record.topic())
        .setAttribute("partition", record.partition())
        .setAttribute("offset", record.offset())
        .startSpan();
    try {
        processRecord(record);
    } catch (Exception e) {
        span.recordException(e);
    } finally {
        span.end();
    }
}
```
By emitting trace data directly within the consumer loop, engineers can detect missing, slow, or corrupt message patterns that Kafka lag metrics alone can never reveal.
## Stop Schema Drift Before It Silently Fails Your Pipeline
One of the most overlooked failure modes in stream processing is schema drift, when producers silently introduce changes that consumers cannot handle.
This issue rarely crashes the system. Instead, it causes semantic degradation: fields go missing, types mismatch, events are misclassified. These failures are quiet, cumulative, and corrosive.
Python
```
from fastavro.validation import validate
from my_schemas import user_event_schema
if not validate(event_data, user_event_schema):
    logger.warning("Schema violation: %s", event_data)
    send_to_dead_letter_queue(event_data)
```
Inline schema validation like this becomes your first line of defense. It converts structural drift into observable exceptions.
Use DLQs as a Debugging Feed, Not a Graveyard, DLQs are too often a compliance checkbox.
But in resilient systems, DLQs are an active observability layer.
JavaScript
```
function sendToDLQ(event, error) {
  const payload = {
    originalEvent: event,
    reason: error.message,
    timestamp: new Date().toISOString()
  };
  dlqProducer.send({
    topic: 'user-events-dlq',
    messages: [{ value: JSON.stringify(payload) }],
  });
}
```
Surface these payloads into dashboards. Create alerts based on DLQ velocity. Feed them into analytics to track error provenance. DLQ monitoring is the canary for Kafka integrity.
## Fluent Bit + Trace ID = Cross-System Log Correlation
Disconnected logs create blind spots. The moment you embed trace IDs into logs, metrics, and spans, they become a causal graph.
`ini` `[PARSER]``Name trace-json``Format json``Time_Key time``Time_Format``%Y-%m-%dT%H:%M:%S``Trace_Key``trace_id`
Now your observability tools, Jaeger, Tempo, or Grafana, can visualize the full lifecycle of an event, across microservices and topics. This is not logging. This is distributed forensics.
## How Message Loss Surfaces in Real Systems
Message loss rarely appears as a 500 error. Instead, it manifests as:
- Stale dashboards: data pipelines silently stall
- Missing audit trails: events never persisted
- Downstream inconsistencies: analytics skewed by phantom gaps
- SLO violations: alerts never trigger because triggers never arrived
These symptoms confuse even experienced SREs. But they almost always trace back to the same root cause: the message was sent, but it was never seen.
## Design Your Messaging Infrastructure for Message Loss, Not Just Load
Most streaming architectures optimize for throughput. Few account for invisible failure conditions like:
- Missing timestamps
- Unparsable payloads
- Clock skew or out-of-order arrival
- Regional lag or partition mismatch
Key observability controls:
- Trace every Kafka consumer event
- Validate schemas at ingest
- Route and monitor DLQ flows
- Correlate logs via trace ID across all layers
Security and data integrity also matter: in multi-tenant architectures, each trace must be attributable to a tenant ID or scoped identity token. This is where IAM enforcement and access-aware observability become critical. Trace IDs should be tied to access contexts.
Looking ahead, some teams are already experimenting with ML anomaly detection over trace flows and DLQ growth rates, flagging novel failure patterns before they affect downstream SLOs.
You cannot eliminate message loss. But you can eliminate its invisibility.
For engineers maintaining distributed systems with Kafka or similar queues, observability must be designed into the system, not added after incident response.
Instrument deeply. Trace cross-service flow. Monitor tenant-level behavior. Learn from what breaks.
Observability
Drops (app)
kafka
Pipeline (software)
Opinions expressed by DZone contributors are their own.
Comments

View File

@@ -0,0 +1,15 @@
# Scaling Prometheus: Managing 80M Metrics Smoothly
- **期号**: SRE Weekly Issue #494(2025-09-14)
- **作者**: Kapil
- **链接**: https://kapillamba4.medium.com/hierarchical-federation-in-prometheus-managing-millions-of-metrics-cleanly-8d8bac940ff3
## 简介
Faced with 80 million time series, these folks found that Statsd + InfluxDB weren’t cutting it, so they switched to Prometheus.
Accessibility note: this article contains a table of text in an image with no alt text.
## 正文
> ⚠️ 抓取失败:HTTP 403

View File

@@ -0,0 +1,77 @@
# A deep dive into Cloudflare’s September 12, 2025 dashboard and API outage
- **期号**: SRE Weekly Issue #494(2025-09-14)
- **作者**: Tom Lianza and Joaquin Madruga — Cloudflare
- **链接**: https://blog.cloudflare.com/deep-dive-into-cloudflares-sept-12-dashboard-and-api-outage/
## 简介
How do these folks keep producing such detailed write-ups the day after an incident?
## 正文
# A deep dive into Cloudflare’s September 12, 2025 dashboard and API outage
![BLOG-2938 1](https://blog.cloudflare.com/_image?href=https%3A%2F%2Fblog.cloudflare.com%2F_emdash%2Fapi%2Fmedia%2Ffile%2F01KW48Q60AN0P6ZP2F3Q70SQJC.png&w=1999&h=1125&f=webp&fit=cover&position=center)
## What Happened
We had an outage in our Tenant Service API which led to a broad outage of many of our APIs and the Cloudflare Dashboard.
The incident’s impact stemmed from several issues, but the immediate trigger was a bug in the dashboard. This bug caused repeated, unnecessary calls to the Tenant Service API. The API calls were managed by a React useEffect hook, but we mistakenly included a problematic object in its dependency array. Because this object was recreated on every state or prop change, React treated it as “always new,” causing the useEffect to re-run each time. As a result, the API call executed many times during a single dashboard render instead of just once. This behavior coincided with a service update to the Tenant Service API, compounding instability and ultimately overwhelming the service, which then failed to recover.
When the Tenant Service became overloaded, it had an impact on other APIs and the dashboard because Tenant Service is part of our API request authorization logic. Without Tenant Service, API request authorization can not be evaluated. When authorization evaluation fails, API requests return 5xx status codes.
We’re very sorry about the disruption. The rest of this blog goes into depth on what happened, and what steps we are taking to prevent it from happening again.
## Timeline
| Time (UTC) | Description |
|---|---|
| 2025-09-12 16:32 | A new version of the Cloudflare Dashboard is released which contains a bug that will trigger many more calls to the /organizations endpoint, including retries in the event of failure. |
| 2025-09-12 17:50 | A new version of the Tenant API Service is deployed. |
| 2025-09-12 17:57 | The Tenant API Service becomes overwhelmed as new versions are deploying. Dashboard Availability begins to drop IMPACT START |
| 2025-09-12 18:17 | After providing more resources to the Tenant API Service, the Cloudflare API climbs to 98% availability, but the dashboard does not recover. IMPACT DECREASE |
| 2025-09-12 18:58 | In an attempt to restore dashboard availability, some erroring codepaths were removed and a new version of the Tenant Service is released. This was ultimately a bad change and causes API Impact again. IMPACT INCREASE |
| 2025-09-12 19:01 | In an effort to relieve traffic against the Tenant API Service, a temporary ratelimiting rule is published. |
| 2025-09-12 19:12 | The problematic changes to the Tenant API Service are reverted, and Dashboard Availability returns to 100%. IMPACT END |
### Dashboard availability
The Cloudflare dashboard was severely impacted throughout the full duration of the incident.
![BLOG-3011 1](https://blog.cloudflare.com/_image?href=https%3A%2F%2Fblog.cloudflare.com%2F_emdash%2Fapi%2Fmedia%2Ffile%2F01KW458QAZTTNR923YP9JBNREG.png&w=715&h=138&f=webp&fit=cover&position=center)
### API availability
The Cloudflare API was severely impacted for two periods during the incident when the Tenant API Service was down.
![BLOG-3011 2](https://blog.cloudflare.com/_image?href=https%3A%2F%2Fblog.cloudflare.com%2F_emdash%2Fapi%2Fmedia%2Ffile%2F01KW44B8NHWE08HXXS6YK6RB9Y.png&w=715&h=135&f=webp&fit=cover&position=center)
## How we responded
Our first goal in an incident is to restore service. Often that involves fixing the underlying issue directly, but not always. In this case we noticed increased usage across our Tenant Service, so we focused on reducing the load and increasing the available resources. We installed a global rate limit on the Tenant Service to help regulate the load. The Tenant Service is a GoLang process that runs on Kubernetes in a subset of our datacenters. We increased the number of pods available as well to help improve throughput. While we did this, we had others on the team continue to investigate why we were seeing the unusually high usage. Ultimately, increasing the resources available to the tenant service helped with availability but was insufficient to restore normal service.
![BLOG-3011 3](https://blog.cloudflare.com/_image?href=https%3A%2F%2Fblog.cloudflare.com%2F_emdash%2Fapi%2Fmedia%2Ffile%2F01KW49B7PHGRB9HFMEFVYH2Z72.png&w=715&h=211&f=webp&fit=cover&position=center)
After the Tenant Service began reporting healthy again and the API largely recovered, we still observed a considerable number of errors being reported from the service. We theorized that these were responsible for the ongoing Dashboard availability issues and made a patch to the service with the expectation that it would improve the API health and restore the dashboard to a healthy state. Ultimately this change degraded service further and was quickly reverted. The second outage can be seen in the graph above.
It’s painful to have an outage like this. That said, there were a few things that helped lessen the impact. Our automatic alerting service quickly identified the correct people to join the call and start working on remediation. Additionally, this was a failure in the control plane which has strict separation of concerns from the data plane. Thus, the outage did not affect services on Cloudflare’s network. The majority of users at Cloudflare were unaffected unless they were making configuration changes or using our dashboard.
## Going forward
We believe it’s important to learn from our mistakes and this incident is an opportunity to make some improvements. Those improvements can be categorized as either ways to reduce / eliminate the impact of a similar change or as improvements to our observability tooling to better inform the team during future events.
### Reducing impact
We use Argo Rollouts for releasing, which monitors deployments for errors and automatically rolls back that service on a detected error. We’ve been migrating our services over to Argo Rollouts but have not yet updated the Tenant Service to use it. Had it been in place, we would have automatically rolled back the second Tenant Service update limiting the second outage. This work had already been scheduled by the team and we’ve increased the priority of the migration.
When we restarted the Tenant Service, everyone’s dashboard began to re-authenticate with the API. This caused the API to become unstable again causing issues with everyone’s dashboard. This pattern is a common one often referred to as a Thundering Herd. Once a resource or service is made available, everyone tries to use it all at once. This is common, but was amplified by the bug in our dashboard logic. The fix for this behavior has already been released via a hotfix shortly after the impact was over. We’ll be introducing changes to the dashboard that include random delays to spread out retries and reduce contention as well.
Finally, the Tenant Service was not allocated sufficient capacity to handle spikes in load like this. We’ve allocated substantially more resources to this service, and are improving the monitoring so that we will be proactively alerted before this service hits capacity limits.
### Improving visibility
We immediately saw an increase in our API usage but found it difficult to identify which requests were retries vs new requests. Had we known that we were seeing a sustained large volume of new requests, it would have made it easier to identify the issue as a loop in the dashboard. We are adding changes to how we call our APIs from our dashboard to include additional information, including if the request is a retry or new request.
We’re very sorry about the disruption. We will continue to investigate this issue and make improvements to our systems and processes.

View File

@@ -0,0 +1,37 @@
# Nothing fails like a history of success
- **期号**: SRE Weekly Issue #494(2025-09-14)
- **作者**: Lorin Hochstein
- **链接**: https://surfingcomplexity.blog/2025/09/07/nothing-fails-like-a-history-of-success/
## 简介
The author ties a recent outage in San Francisco’s BART transit service to a couple of previous incidents by a common thread: confidence placed in a procedure that had been performed successfully previously.
This article also links to BART’s memo which is surprisingly detailed and a great read.
## 正文
The Axiom of Experience: the future will be like the past, because, in the past, the future was like the past. – Gerald M. Weinberg, [An Introduction to General Systems Thinking](https://geraldmweinberg.com/Site/General_Systems.html)
Last Friday, the San Francisco Bay Area Rapid Transit system (known as *BART*) experienced a [multiple hour outage](https://www.kqed.org/news/12054754/bart-outage-shuts-down-entire-system-for-2nd-time-in-months). Later that day, the BART Deputy General Manager released a [memo](https://www.bart.gov/sites/default/files/2025-09/Memo%20%28DGM%20to%20BOD%29%20September%205%2C%202025%20Service%20Disruption%20Update_09.05.2025.pdf) about the outage with some technical details. The memo is brief, but I was honestly surprised to see this amount of detail in a public document that was released so quickly after an incident, especially from a public agency. What I want to focus on in this post is this line (emphasis mine):
Specifically, network engineers were performing a cutover to a new network switch at
Montgomery St. Station… **The team had already successfully performed eight similar cutovers earlier this year.**
This reminded me of something I read in the [Buildkite writeup](https://www.buildkitestatus.com/incidents/txxkzf4r262c) from an incident that happened back in January of this year (emphasis mine):
Given **the confidence gained by initial load testing and the migrations already performed over the past year**, we wanted to allow customers to take advantage of their seasonal low periods to perform shard migrations, as a win-win. **This caused us to discount the risk of performing migrations** during a seasonal low period and what impacts might emerge when regular peak traffic returned.
It also reminded me about the [2022 Rogers Telecommunications outage in Canada](https://crtc.gc.ca/eng/publications/reports/xonarp2023.htm#fn5-rf) (emphasis mine, *[redacted]* comments in the original):
Rogers had assessed the risk for the initial change of this seven-phased process as “High”. Subsequent changes in the series were listed as “Medium.” [redacted] was “Low” risk based on the Rogers algorithm that weighs prior success into the risk assessment value. **Thus, the risk value for [redacted] was reduced to “Low” based on successful completion of prior changes.**
Whenever we make any sort of operational change, we have a mental model of the risk associated with the change. We view novel changes (*I’ve never done something like this before!*) as riskier than changes we’ve performed successfully multiple times in the past (*I’ve done this plenty of times*). I don’t think this sort of thinking is a fallacy: rather, it’s a heuristic, and it’s generally a pretty effective one! But, like all heuristics, it isn’t perfect. As shown in the examples above, the application of this heuristic can result in a miscalibrated mental model of the risk associated with a change.
So, what’s the broader lesson? In practice, our risk models (implicit or otherwise) are always miscalibrated: a history of past successes is just one of multiple avenues that can lead us astray. Trying to achieve a perfect risk model is like trying to deploy software that is guaranteed to have zero bugs: it’s never going to happen. Instead, we need to accept the reality that, like our code, our models of risk will always have defects that are hidden from us until it’s too late. So we’d better get damned good at recovery.

View File

@@ -0,0 +1,144 @@
# How we sped up code search for Graphite Chat
- **期号**: SRE Weekly Issue #494(2025-09-14)
- **作者**: Brandon Willett — Graphite
- **链接**: https://graphite.dev/blog/how-we-sped-up-code-search-graphite-chat
## 简介
The folks at Graphite take us through their discovery of why code search is difficult and the strategies they employed to solve it.
## 正文
Searching a codebase is *usually* a pretty easy problem. You want to know what files use your `MAX_EMAILS_TO_SEND` constant? One command gets you there in milliseconds:
Terminal
grep -r MAX_EMAILS_TO_SEND .
Modern tools like [`ripgrep`](https://github.com/BurntSushi/ripgrep) make it even faster. But this simplicity depends on a couple of things being true:
- Your files all reside on a (hopefully) fast disk.
- There aren't too many of them.
When building out a code search tool for the agentic [Graphite Chat](https://graphite.dev/docs/graphite-chat), we immediately ran into both of these limitations. We needed to support searches across hundreds of thousands (or millions) of files, at any commit, without maintaining a whole traditional VM+disk for each one. Could `grep` even support that kind of use case?
Let’s find out!
## Default-branch-only search isn’t enough
You might read this and think: What's supposed to be novel here? We've had fast code search via API for big repositories for years now. [Sourcegraph](https://sourcegraph.com/) made a whole company out of the idea, and GitHub even gives you something similar for free.
True, but only on the *default* branch. GitHub’s API docs are explicit:
Only the *default branch* is considered. In most cases, this will be the `main` branch.
And Sourcegraph, while it’s closed-source now, still has an archive of their documentation and code as it stood in 2023, which details [the same thing](https://github.com/sourcegraph/sourcegraph-public-snapshot/blob/main/doc/dev/background-information/architecture/index.md):
Sourcegraph also has a fast search path for code that isn't indexed yet, or for code that will never be indexed (for example: code that is not on a default branch). Indexing every branch of every repository isn't a pragmatic use of resources for most customers, so this decision balances optimizing the common case (searching all default branches) with space savings (not indexing everything).
…
Provides on-demand unindexed search for repositories. It scans through a git archive fetched from gitserver to find results, similar in nature to `git grep`.
The searches our model makes almost never target `main`; they’re against arbitrary commits. For us, fast search *at any commit* was the requirement, and that turned out to be a much harder problem.
## What didn’t work
In the spirit of [KISS](https://en.wikipedia.org/wiki/KISS_principle) (or “[Choose Boring Technology](https://boringtechnology.club)”, or “[The Grug Brained Developer](https://grugbrain.dev)”, feel free to pick your favorite decade’s spin on the idea), we figured our first attempt at a solution should also be the most straightforward: a plain old `git grep` on the repository. If that worked, then we could save ourselves a lot of time.
We spent a week running experiments on several block-based storage solutions on AWS:
- On-demand AWS Lambdas to execute searches via one shared EFS volume.
- An ECS cluster where tasks dynamically mount per-repo EFS volumes at query time.
- Persistent EC2 instances which use (faster) EBS with [multi-attach](https://docs.aws.amazon.com/ebs/latest/userguide/ebs-volumes-multi.html) for horizontal scaling.
We saw some interesting results here, for example, we measured the EBS-mounted volumes to consistently perform I/O intensive operations (like searches) about 3x as fast as EFS-mounted ones. From a test we ran searching for a string in the `react-native` repository:
Terminal
# EBSPerformance counter stats for 'git grep interactionManager 8f1ae53':17,541 minor-faults # 0.062 M/sec878 major-faults # 0.003 M/sec282.34 msec task-clock # 0.184 CPUs utilized1.535785602 seconds time elapsed0.208365000 seconds user0.075011000 seconds sys
Terminal
# EFSPerformance counter stats for 'git grep interactionManager 8f1ae53':17,620 minor-faults # 0.050 M/sec53 major-faults # 0.149 K/sec354.59 msec task-clock # 0.078 CPUs utilized4.560184652 seconds time elapsed0.212312000 seconds user0.129052000 seconds sys
However, in spite of the individual advantages each had over the others, they all buckled in the same place: large repositories.
While `react-native` is a decently-sized repo by most measures, it only has a few thousand files:
Terminal
$ git ls-files | wc -l7019
We need to support repositories hundreds of times this size, and in cases like that, we actually end up at the mercy of the Linux page cache. The `major-faults` rows in the snippets above represent times the `grep` process needed to actually reach out to the disk for blocks, and for sample large codebases we tested, searches only returned “fast enough” (under 10 seconds) when `major-faults` was 0 (for instance, when the same search is run twice in a row). In other words, searches are fast only when the *entire repository* is already in the page cache.
At this point, the `grep` operation becomes mostly CPU bound, and perhaps unsurprisingly, this is also where EFS and EBS see their performance disparity vanish:
Terminal
Performance counter stats for 'git grep interactionManager 8f1ae53':17,667 minor-faults # 0.071 M/sec0 major-faults # 0.000 K/sec248.55 msec task-clock # 0.957 CPUs utilized0.259805935 seconds time elapsed0.208592000 seconds user0.040113000 seconds sys
Of course, the only way of ensuring that every block needed for a search is already in the page cache when the search comes in, is to have a crystal ball, which we can’t afford.
So, we have to look into some more earthly solutions!
## What somewhat worked
If we can't brute-force our way through a big repository with `grep`, then it looks like we'll need to do some indexing.
We tested loading the repository’s files into [Elasticsearch](https://www.elastic.co/), which tokenizes documents and executes keyword queries quickly. The results were promising:
Terminal
$ cat search.json{"query": {"wildcard": {"content": {"value": "*interactionManager*"}}}}$ https GET <elastic-endpoint>/blobs/_search?size=100 < search.json{"hits": {"hits": [...],"max_score": 1.0,"total": {"relation": "eq","value": 83}},"timed_out": false,"took": 79}
Even using a tiny node type ([`i4g.large.search`](https://instances.vantage.sh/aws/opensearch/i4g.large.search)), we saw searches over even huge repositories taking well under 500ms. And that's with a cluster of just one node; we could easily split the dataset into shards (ES manages this on its own) such that 10 different nodes handle 1/10 of the files each, and our searches would be even faster.
Problem solved, right? Not quite. What we just tried with Elasticsearch worked great for *one commit*, but our requirement was fast search at *any commit*. The direct approach would be to index the repo’s state at each commit, for easy lookups, but — multiply thousands of repos by thousands of files by thousands of commits, and suddenly you’re managing tens of billions of documents. Running a cluster of that size would be prohibitively complex and expensive.
So, is the “document database” idea dead-on-arrival?
## Back to basics
When looking at the problem through the above lens (the challenge of efficiently managing thousands of files, each with potentially thousands of versions), there *is* one database that's relevant to the discussion: namely, Git itself!
A Git repository seems to store an entire copy of the repository at each commit of a [potentially-decades-long history](https://github.com/torvalds/linux), but does so without terabytes of disk space. This led us to wonder, "how does it do that, and can we somehow adapt that approach to work with a document search database instead of a filesystem?"
For the "how" question, it's honestly very clever if you haven't [read about it](https://git-scm.com/book/en/v2/Git-Internals-Git-Objects) before. The short version is, Git splits the data it stores into several different kinds of objects:
1. **Blobs** → raw file contents,[content-addressed](https://en.wikipedia.org/wiki/Content-addressable_storage) .
2. **Trees** → lists of blob IDs at a given commit.
And as for the "can we" question, it turns out, yes absolutely!
## The life of a search query
The key change was that, like Git, we needed to start storing two kinds of objects instead of just one: the blobs and trees. Then, when it comes time to perform a search, we execute 2 queries:
1. Get the tree object for the commit SHA specified in the query.
2. Get *all* blobs that match the search term specified in the query.
And then, in memory, we filter the results from (2) based on the blob IDs in the result of (1), such that we only return files which are actually in-scope for the request. Easy!
Both of these queries are fast. The first simply retrieves a single document by its ID. While the second does need to consider every version of every file in the repo (which might sound like a lot), our tests show it’s actually not nearly so bad. Some files in the repository have dozens of unique versions, but the vast majority are rarely—if ever—updated. So the total number of documents is generally a small multiple (think ~3x) of the number of files you see in the working tree.
Plus, we’ve made a few optimizations to help things move even faster. If you notice, the above 2 queries are totally independent—you don’t need the result from (1) in order to perform (2), so we do them in parallel. And because a commit’s tree (which just specifies a set of blob IDs) is often much smaller than the result set of blobs, we’ve implemented the in-memory filter as a stream transformer, so that we can start responding to the client the moment we receive our first in-scope match.
## The future
We have a few other ideas for optimizations, some a bit more experimental. For example, if we only ever use the tree for a set-contains operation, do we even need to fetch the entire tree at all? Maybe we can pre-construct the set of blob IDs, and only fetch that at query time, or even pre-construct something that [acts just like that set, most of the time, but is just 10% the size](https://en.wikipedia.org/wiki/Bloom_filter)?
And since we’re using [Turbopuffer](https://turbopuffer.com/) as the document database backing the Graphite code index, adding embeddings is on the table too, which opens the door to semantic search. Maybe we will soon be able to ask PR Chat to check if a pattern you tried is consistent with examples from similar files.
## The now
Today, the system is live, and we’ve already indexed tens of millions of source files across thousands of repositories. For several weeks, Graphite Chat has been using this index to power LLM tool calls, where it’s been a big upgrade from our previous strategy of calling the GitHub API.
We can now fetch as many files as we need (no arbitrary rate limit), our searches accurately target the PR’s branch (instead of only being able to search `main`), and best of all — the index has now served thousands of queries with a median latency consistently under 100 milliseconds.
What started as a “just grep it” idea turned into a deep dive through storage tradeoffs, indexing engines, and Git’s internals. By rethinking search in terms of blobs and trees, we’ve unlocked fast search at any commit—something we couldn’t get off the shelf. And we’ve hardly touched on the ingestion part! It turns out there are ways to have Git shrink a 5GB repo clone down to 20MB, and you'd be surprised how useful an anemic clone like that can still be.
Which is to say, check back soon for a part two!