SRE weekly 所有文章

This commit is contained in:
2026-09-12 17:23:01 +08:00
parent 409b40ddcb
commit af7633f9dc
8486 changed files with 4489990 additions and 7 deletions

View File

@@ -0,0 +1,381 @@
# Scaling Software Systems: 10 Key Factors
- **期号**: SRE Weekly Issue #387(2023-08-27)
- **作者**: Code Reliant
- **链接**: https://www.codereliant.io/scaling-software-systems-10-key-factors/
## 简介
> In this post, we’ll explore 10 areas that are key to designing highly scalable architectures.
The 10 areas they cover in-depth are:
> Horizontal vs. Vertical ScalingLoad BalancingDatabase ScalingAsynchronous ProcessingStateless SystemsCachingNetwork Bandwidth Optimization8, Progressive EnhancementGraceful DegradationCode Scalability
## 正文
![](https://substackcdn.com/image/fetch/$s_!CL35!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5dea17d1-7034-4001-8c58-92321ffb16fa_2000x1439.jpeg)
[Stephen Dawson](https://unsplash.com/@dawson2406?utm_source=ghost&utm_medium=referral&utm_campaign=api-credit)/
[Unsplash](https://unsplash.com/?utm_source=ghost&utm_medium=referral&utm_campaign=api-credit)
As part of my [12-part series](https://www.codereliant.io/p/principles-of-reliable-software-design-part-1) on the Principles of Reliable Software Design, this post will focus on scalability - one of the most critical elements in building robust, future-proof applications.
In today's world of ever-increasing data and users, software needs to be ready to adapt to higher loads. Neglecting scalability is like constructing a beautiful house on weak foundations - it may look great initially but will eventually crumble under strain.
Whether you're building an enterprise system, mobile app or even something for personal use, how do you ensure your software can smoothly handle growth? A scalable system provides a great user experience even during traffic spikes and high usage. An unscalable app is frustrating at best and at worst, becomes unusable or crashes altogether under increased load.
In this post, we'll explore 10 areas that are key to designing highly scalable architectures. By mastering these concepts, you can develop software capable of being deployed on a large scale without expensive rework. Your users will thank you for building an app that delights them today as much as it will tomorrow when your user base has grown 10x.
## Horizontal vs. Vertical Scaling
![Horizontal vs Vertical Scaling Horizontal vs Vertical Scaling](https://substackcdn.com/image/fetch/$s_!5tra!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3000c78b-237a-4d83-83ab-9272a89285a5_1024x768.png)
Vertical scaling involves increasing the power of existing nodes, such as upgrading to servers with faster CPUs, more RAM, or increased storage capacity.
In general, horizontal scaling is preferred because it provides greater reliability through [redundancy](https://www.codereliant.io/p/reliability-foundations-redundancy). If one node fails, other nodes can take over the workload. Horizontal scaling also offers more flexibility to scale out gradually as needed. With vertical scaling, you need to upgrade your hardware altogether to handle increased loads.
However, vertical scaling may be useful when increased computing power is needed for specific tasks like CPU-intensive data processing. Overall, a scalable architecture employs a combination of vertical and horizontal scaling approaches to tune the system resource requirements over time.
## Load Balancing
![](https://substackcdn.com/image/fetch/$s_!lCaF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F318f8492-c889-422d-97a7-87b507e8f037_705x249.png)
This prevents any single server from becoming overwhelmed. The load balancer can implement different algorithms like round-robin, least connections, or IP-hash to determine how to distribute load. More advanced load balancers can detect server health and adaptively shift traffic away from failing nodes.
Load balancing maximizes resource utilization and increases performance. It also provides high availability and reliability. If a server goes down, the load balancer redirects traffic to the remaining online servers. This redundancy makes your system resilient to individual server failures.
[Implementing load balancing](https://www.codereliant.io/p/lets-build-loadbalancer-go) alongside auto-scaling allows your system to scale out smoothly and painlessly. Your application can comfortably handle large traffic variations without running into capacity issues.
## Database Scaling
As your application usage grows, the database backing your system can become a bottleneck. There are several techniques to scale databases to meet high read/write loads. However, databases are one of the hardest components to scale in most systems.
### Database Selection:
The selection of an appropriate database plays a critical role in effectively scaling a database system. It depends on various factors, including the type of data to be stored and the expected query patterns. Different types of data, such as metrics data, logs, enterprise data, graph data, and key/value stores, have distinct characteristics and requirements that demand tailored database solutions.
For metrics data, where high write-throughput is essential to record time-series data, a time-series database like InfluxDB or Prometheus may be more suitable due to their optimized storage and querying mechanisms. On the other hand, for handling large volumes of unstructured data, such as logs, a NoSQL database like Elasticsearch or could provide efficient indexing and searching capabilities.
For enterprise data that requires strict ACID (Atomicity, Consistency, Isolation, Durability) transactions and complex relational querying, a traditional SQL database like PostgreSQL or MySQL might be the right choice. In contrast, for scenarios demanding simple read and write operations, a key/value store such as Redis or Cassandra could offer low-latency data access.
It's essential to thoroughly evaluate the specific requirements of the application and its data characteristics before making a database choice. Sometimes, a combination of databases (polyglot persistence) might be the most effective strategy, utilizing different databases for different parts of the application based on their strengths. Ultimately, the right database selection can significantly impact the scalability, performance, and overall success of the system.
### Vertical Scaling:
Simply throwing more resources at a single database server like CPU, memory and storage can provide temporary relief for increased loads. And it should always be tried out before looking into advanced concepts of scaling the database. In addition, vertical scaling keeps your database stack simple.
However, there is a physical ceiling to how large a single server can scale. Also, a monolithic database remains a single point of failure - if that beefed up server goes down, so does access to the data.
That's why alongside vertical scaling of the database server hardware, it's critical to employ horizontal scaling techniques.
### Replication:
Replication provides redundancy and improves performance by copying data across multiple database instances. A write to the leader node is replicated to read replicas. Reads can be served from the replicas, reducing load on the master. Also, replication copies data across redundant servers, eliminating the single point of failure risk.
![Database Replication Database Replication](https://substackcdn.com/image/fetch/$s_!rlUj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6cbcd4e4-9669-4d86-a400-84f28e1ba74d_1035x500.png)
Sharding partitions your database across multiple smaller servers, allowing you to add more nodes fluidly as needed.
Sharding or partitioning involves splitting your database into multiple smaller databases by a certain criteria like customer ID or geographic region. This allows you to scale horizontally by adding more database servers.
In addition, there other areas that should also be put under light that can help scale database:
- **Schema Denormalization** involves the duplication of data in a database to diminish the need for complex joins in queries, resulting in improved query performance.
- **Caching** frequently accessed data in a fast in-memory cache reduces database queries. A cache hit avoids having to fetch the data from the slower database.
## Asynchronous Processing
Synchronous request-response cycles can create bottlenecks that impede scalability, especially for long running or IO-intensive tasks. Asynchronous processing queues up work to be handled in the background, freeing up resources immediately for other requests.
For example, submitting a video transcoding job could directly block a web request, negatively impacting user experience. Instead, the transcoding task can be published to a queue and handled asynchronously. The user gets an immediate response, while the task processes separately.
![Asynchronous Video Uploading & Transcoding Asynchronous Video Uploading & Transcoding](https://substackcdn.com/image/fetch/$s_!qKhK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1019c9a5-372f-4793-a112-356073938ad9_1142x609.png)
Asynchronous tasks can be executed concurrently by background workers scaled horizontally across many servers. Queue sizes can be monitored to add more workers dynamically. Load is distributed evenly, preventing any single worker from becoming overwhelmed.
Shifting workloads from synchronous to asynchronous allows the application to handle spikes in traffic smoothly without getting bogged down. Systems remain responsive under load using robust queue-based asynchronous processing.
## **Stateless Systems**
Stateless systems are easier to horizontally scale out compared to stateful designs. When application state is persisted in external storage like databases or distributed caches rather than locally on servers, new instances can be spun up as needed.
In contrast, stateful systems require sticky sessions or data replication across instances. A stateless application places no dependency on specific servers. Requests can be routed to any available resource.
Saving state externally also provides better fault tolerance. The loss of any stateless application server is not impactful since it holds no un-persisted critical data. Other servers can seamlessly take over processing.
A stateless architecture improves reliability and scalability. Resources can scale elastically while remaining decoupled from individual instances. However, external state storage adds overhead of cache or database queries. The tradeoffs require careful evaluation when designing web-scale applications.
## **Caching**
Caching frequently accessed data in fast in-memory stores is a powerful technique to optimize scalability. By serving read requests from low latency caches, you can dramatically reduce load on backend databases and improve performance.
For example, product catalog information that rarely changes is ideal for caching. Subsequent product page requests can fetch data from Redis or Memcached rather than overloading your MySQL store. Cache invalidation strategies help keep data consistent.
Caching also benefits compute-heavy processes like template rendering. You can cache the rendered output and bypass redundant rendering for each request. CDNs like Cloudflare cache and serve static assets like images, CSS, and JS globally.
### **Redis Golang Example:**
```
package main
import (
"database/sql"
"encoding/json"
"fmt"
"log"
"net/http"
"time"
"github.com/go-redis/redis"
_ "github.com/go-sql-driver/mysql"
)
const (
dbUser = "your_mysql_username"
dbPassword = "your_mysql_password"
dbName = "your_mysql_dbname"
redisAddr = "localhost:6379"
)
type Product struct {
ID int `json:"id"`
Name string `json:"name"`
Price int `json:"price"`
}
var db *sql.DB
var redisClient *redis.Client
func init() {
// Initialize MySQL connection
dbSource := fmt.Sprintf("%s:%s@/%s", dbUser, dbPassword, dbName)
var err error
db, err = sql.Open("mysql", dbSource)
if err != nil {
log.Fatalf("Error opening database: %s", err)
}
// Initialize Redis client
redisClient = redis.NewClient(&redis.Options{
Addr: redisAddr,
Password: "", // No password set
DB: 0, // Use default DB
})
// Test the Redis connection
_, err = redisClient.Ping().Result()
if err != nil {
log.Fatalf("Error connecting to Redis: %s", err)
}
log.Println("Connected to MySQL and Redis")
}
func getProductFromMySQL(id int) (*Product, error) {
query := "SELECT id, name, price FROM products WHERE id = ?"
row := db.QueryRow(query, id)
var product Product
err := row.Scan(&product.ID, &product.Name, &product.Price)
if err != nil {
return nil, err
}
return &product, nil
}
func getProductFromCache(id int) (*Product, error) {
productJSON, err := redisClient.Get(fmt.Sprintf("product:%d", id)).Result()
if err == redis.Nil {
// Cache miss
return nil, nil
} else if err != nil {
return nil, err
}
var product Product
err = json.Unmarshal([]byte(productJSON), &product)
if err != nil {
return nil, err
}
return &product, nil
}
func cacheProduct(product *Product) error {
productJSON, err := json.Marshal(product)
if err != nil {
return err
}
key := fmt.Sprintf("product:%d", product.ID)
return redisClient.Set(key, productJSON, 10*time.Minute).Err()
}
func getProductHandler(w http.ResponseWriter, r *http.Request) {
productID := 1 // For simplicity, we are assuming product ID 1 here. You can pass it as a query parameter.
// Try getting the product from the cache first
cachedProduct, err := getProductFromCache(productID)
if err != nil {
http.Error(w, "Failed to retrieve product from cache", http.StatusInternalServerError)
return
}
if cachedProduct == nil {
// Cache miss, get the product from MySQL
product, err := getProductFromMySQL(productID)
if err != nil {
http.Error(w, "Failed to retrieve product from database", http.StatusInternalServerError)
return
}
if product == nil {
http.Error(w, "Product not found", http.StatusNotFound)
return
}
// Cache the product for future requests
err = cacheProduct(product)
if err != nil {
log.Printf("Failed to cache product: %s", err)
}
// Respond with the product details
json.NewEncoder(w).Encode(product)
} else {
// Cache hit, respond with the cached product details
json.NewEncoder(w).Encode(cachedProduct)
}
}
func main() {
http.HandleFunc("/product", getProductHandler)
log.Fatal(http.ListenAndServe(":8080", nil))
}
```
Strategically leveraging caching reduces strain on infrastructure and scales horizontally as you add more cache servers. Caching works best for read-heavy workloads with repetitive access patterns. It provides scalability gains alongside database sharding and asynchronous processing.
## Network Bandwidth Optimization
For distributed architectures spread across multiple servers and regions, optimizing network bandwidth utilization is key to scalability. Network calls can become a bottleneck, imposing limits on throughput and latency.
Bandwidth optimization techniques like compression and caching reduce the number of network hops and amount of data transferred. Compressing API and database responses minimizes bandwidth needs.
Persistent connections via HTTP/2 allow multiple requests over one open channel. This reduces round trip overheads, improves resource utilization, and avoid HTTP head of line blocking. However, HTTP/2 still suffers from TCP head of line blocking. So, we can now even use HTTP/3 which is being done over QUIC instead of TCP and TLS, and it avoid TCP head of line blocking.
CDN distribution brings data closer to users by caching assets at edge locations. By serving content from nearby, less data traverses costly long-haul routes.
Gzip Golang Example:
```
package main
import (
"github.com/labstack/echo/v4"
"github.com/labstack/echo/v4/middleware"
)
func main() {
e := echo.New()
// Middleware
e.Use(middleware.Logger())
e.Use(middleware.Recover())
e.Use(middleware.Gzip()) // Add gzip compression middleware
// Routes
e.GET("/", helloHandler)
// Start server
e.Logger.Fatal(e.Start(":8080"))
}
func helloHandler(c echo.Context) error {
return c.String(200, "Hello, Echo!")
}
```
Overall, scaling requires a holistic view encompassing not just compute and storage, but also network connectivity. Optimizing bandwidth usage by minimizing hops, compression, caching and more is invaluable for building high-throughput and low-latency large-scale systems.
## Progressive Enhancement
Progressive enhancement is a strategy that helps improve scalability for web applications. The idea is to build the core functionality first and then progressively enhance the experience for capable browsers and devices.
For example, you can develop the basic HTML/CSS site to ensure accessibility on any browser. Then you can add advanced CSS and JavaScript to incrementally improve interactions for modern browsers with JS support.
Serving basic HTML first provides a fast “time-to-interactive” and works on all platforms. Enhancements load afterwards to optimize the experience without blocking. This balanced approach extends reach while utilizing capabilities. For example, [Qwik](https://qwik.builder.io/docs/concepts/progressive/) bake this concept into the foundation of the framework.
Progressively enhancing in phases also aids scalability. Simple pages require fewer resources and scale better. You can add more advanced features when needed rather than prematurely over-engineering for every possible use case upfront.
Overall, progressive enhancement allows web apps to scale efficiently right from basic to advanced functionality based on device capabilities and user needs.
## Graceful Degradation
In contrast to progressive enhancement, graceful degradation involves starting from an advanced experience and scaling back features when constraints are detected. This allows applications to scale down fluidly when facing resource limitations.
For instance, a graphically-rich app may detect a low-powered mobile device and adapt to downgrade advanced visuals into a more basic presentation. Or a backend system may throttle non-essential operations during peak load to maintain core functionality.
Gracefully degrading preserves critical user workflows even under suboptimal conditions. Errors due to constraints like bandwidth, device capabilities or traffic spikes are minimized. The experience remains operational rather than failing catastrophically.
**Feature degradation** is a valuable tool that should be incorporated and planned for during the initial development of product features. The ability to deactivate features automatically or manually can prove essential in keeping the system functional under various circumstances, such as system overload, migrations, or unexpected performance issues.
When a system experiences high load or is overwhelmed by excessive traffic, dynamically deactivating non-critical features can alleviate strain and prevent complete service failures. This smart use of feature degradation ensures that the core functionalities remain operational and prevents cascading failures across the application.
During database migrations or updates, feature degradation can help maintain system stability. By temporarily disabling certain features, the complexity of the migration process can be reduced, minimizing the risk of data inconsistencies or corruption. Once the migration is complete and verified, the features can be reactivated seamlessly.
Moreover, feature degradation can be a useful mechanism in situations where a critical bug or security vulnerability is discovered in a specific feature. Turning off the affected feature promptly can prevent any further damage while the issue is being addressed, ensuring the overall system's integrity.
Overall, incorporating feature degradation as part of the product's design and development strategy empowers the system to gracefully handle challenging situations, enhance resilience, and maintain an uninterrupted user experience during adverse conditions.
Building in graceful degradation mechanisms like device detection, performance monitoring and throttling improves an application's resilience when scaling up or down. Resources can be dynamically tuned to optimal levels based on real-time constraints and priorities.
## Code Scalability
Scalability best practices focus heavily on infrastructure and architecture. But well-written and optimized code is key for scaling too. Suboptimal code hinders performance and resource utilization even on robust infrastructure.
Tight loops, inefficient algorithms and poorly structured data access can bog down servers. Architectures like microservices increase parallelism, but can multiply these inefficiencies.
Code profilers help identify hot spots and bottlenecks. Refactoring code to scale better optimizes CPU, memory and I/O resource usage. Distributing processing across threads also improves utilization of multi-core servers.
### Example of Unscalable Code (Thread per request):
Inefficient code can hinder scalability even on robust infrastructure. For instance, allocating one thread per request does not scale well - the server will run out of threads under high load.
Better approaches like asynchronous/event-driven programming and non-blocking I/O provide higher scalability. Node.js handles many concurrent requests efficiently on a single thread using this model.
Virtual threads or goroutines are also more scalable than thread pools. Virtual threads are lightweight and managed by the runtime. Examples are goroutines in Go and green threads in Python.
Hundreds of thousands of goroutines can run concurrently vs limited OS threads. The runtime multiplexes goroutines onto real threads automatically. This removes thread lifecycle overhead and resource constraints of thread pools.
Carefully structured code that maximizes asynchronous processing, virtual threads, and minimized overhead is vital for large-scale applications, despite infrastructure.
### Java Example of Virtual Thread Per Task:
```
import java.io.*;
import java.net.ServerSocket;
import java.net.Socket;
import java.util.concurrent.ExecutorService;
import java.util.concurrent.Executors;
public class VirtualThreadServer {
public static void main(String[] args) {
final int portNumber = 8080;
try {
ServerSocket serverSocket = new ServerSocket(portNumber);
System.out.println("Server started on port " + portNumber);
ExecutorService executor = Executors.newVirtualThreadPerTaskExecutor();
while (true) {
// Wait for a client connection
Socket clientSocket = serverSocket.accept();
System.out.println("Client connected: " + clientSocket.getInetAddress());
// Submit the request handling task to the virtual thread executor
executor.submit(() -> handleRequest(clientSocket));
}
} catch (IOException e) {
e.printStackTrace();
}
}
static void handleRequest(Socket clientSocket) {
try (
BufferedReader in = new BufferedReader(new InputStreamReader(clientSocket.getInputStream()));
PrintWriter out = new PrintWriter(clientSocket.getOutputStream(), true)
) {
// Read the request from the client
String request = in.readLine();
// Process the request (you can add your custom logic here)
String response = "HTTP/1.1 200 OK\r\nContent-Type: text/html\r\n\r\nHello, this is a virtual thread server!";
// Send the response back to the client
out.println(response);
// Close the connection
clientSocket.close();
} catch (IOException e) {
e.printStackTrace();
}
}
}
```
Notes if you want to run the code above, make sure you have java 20 installed, copy the code into `VirtualThreadServer.java` and run it using `java --source 20 --enable-preview VirtualThreadServer.java`.
Just as infrastructure needs to scale, so does code. Efficient code ensures servers operate optimally under load. Overloaded servers cripple scalability, irrespective of the surrounding architecture. Optimize code alongside scaling infrastructure for best results.
## Conclusion
Scaling a software system to handle growth is crucial for long-term success. We've explored key techniques like horizontal scaling, load balancing, database sharding, asynchronous processing, caching, and optimized code to design highly scalable architectures.
While scaling requires continual effort, investing early in scalability will prevent painful bottlenecks down the road. Consider your capacity needs well in advance rather than as an afterthought. Build redundancies, monitor usage, expand incrementally, and distribute load across many nodes.
With a robust and adaptive design, your software can continue delighting customers even as usage explodes 10x or 100x. Planning for scale will distinguish your application from the multitudes that crash under growth. Your users will stick around when your platform remains just as fast, available and reliable despite increasing demand.
If you enjoyed this, you will also enjoy all the content we have in the making!

View File

@@ -0,0 +1,136 @@
# Time based vs Event based SLIs
- **期号**: SRE Weekly Issue #387(2023-08-27)
- **作者**: Alex Ewerlöf
- **链接**: https://blog.alexewerlof.com/p/time-based-vs-event-based
## 简介
Are you looking at the number of requests that were served successfully out of the total number of requests? Or the percentage of time the system was up and working properly?
## 正文
[Service level indicator (SLI)](https://blog.alexewerlof.com/p/sli) is defined as the percentage of good divided by valid:
There are two types of Service Level Indicators based on the definition of what good looks like:
- **Time based:** measures good time (uptime)
- **Event based:** measure good events (failures)
This choice has huge implication on how you do error budgeting.
This article talks about these two types of SLI with some examples and when to use which.
# Time Based
Time-based service level indicators measure the percentage of time that the system was behaving good.
The focus is on the time period where the system was behaving correctly.
Although the definition of SLI allows being more specific, usually the valid time is the entire measurement window. In other words:
### Time slice
Although, it is possible to calculate the good time down to any precision, Time-Based SLIs often count the number of good *time slices* (e.g. number of good minutes in a month). That way the formula looks like this:
The length of the time slice defines the granularity of the measurement.
If there are multiple data points in a given time slice, they need to be aggregated to calculate the good/bad status of the time slice. There are many aggregation methods:
- **Sampling:** just take the data point in the selected time slice
- **First value:** take the first data point
- **Last value:** take the last data point
- **Random value:** take a random data point as a representative of the whole time slice
- **Average (mean):** calculate the average of all data points in a given time slice
- **Min:** take the data point with the smallest value in a time slice
- **Max:** take the data point with the largest value in a time slice
- **Range:** calculate Max - Min
- **Mode:** find the most frequent value
- **[Percentile](https://blog.alexewerlof.com/p/percentile):** sort the data points in the time slice based on their values values, then pick the data point closer to the P% of the length of the data set
- **P99:** select the data point at the 99% index (e.g. if there are 612 data points, P99 is the data at position 605 in the sorted data set)
- **P50 (Median):** select the data point at the 50% index. This means 50% of the data points are larger or equal than this data point and 50% are lower or equal to this data point.
- ...
- **Count:** count the number of data points in a given time slice
- **Sum:** add the value of all data points in a given time slice
- **Rate:** calculate the ratio of the number of data points that match a criteria from all the data points in a given time slice
- **[Stddev](https://www.youtube.com/watch?v=esskJJF8pCc):** The standard deviation is a measure of how the data is scattered around the*mean (average)* . It has the same unit as the data point value.
- **[Stdvar](https://www.youtube.com/watch?v=5wJUUgnMGWA):** The standard variance is the squared average diff from the*mean (average)* . The unit is the square of the data point value unit.
- **Topk or Bottomk:** select the top (or bottom) data points in a data set that is sorted in ascending order.
Note: some platforms call this functionality *rollup* (e.g. [Datadog](https://docs.datadoghq.com/dashboards/functions/rollup/)) while others call it aggregation (e.g. [Grafana](https://grafana.com/docs/grafana-cloud/cost-management-and-billing/reduce-costs/metrics-costs/control-metrics-usage-via-adaptive-metrics/define-metrics-aggregation-rules/#supported-aggregation-types) or [Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/current/sql-functions-aggs.html)).
### Example
Suppose we have a database and decided to use the *query response latency* as SLI with a time slice of 1 minute. If the P75 percentile of the data response latencies in a minute is above 1000ms (threshold of good), we consider that entire minute a failed time slice.
This is an example of our data points:
![](https://substackcdn.com/image/fetch/$s_!_h6V!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3be84d4-5a4b-4e52-8e7e-77744df9eaa7_1093x705.png)
### Pros
- Time-based SLI is easier to measure
- Uptime is a bit more intuitive to understand
- At any given point in time, it’s easy to calculate the remaining error budget (e.g. How much downtime is remaining for this month?)
- Time-based SLIs are not sensitive to the amount of load (number of data points in a time slice is aggregated anyway). This means a time-based SLI is a good choice for systems that don’t get too much traffic. On the flip side, time-based SLIs are more forgiving to failures due to a spike in load (as far as the SLI is concerned a minute failed during the night is as significant as a minute failed during peak hours). This property may not be desired.
- Some SLIs are time-based by nature. For example, [percentiles](https://blog.alexewerlof.com/p/percentile) are used to focus the optimization on outliers. Percentile is calculated for a range of values over time (e.g. 5 minutes), which fits perfectly into the concept of time slices. You can also calculate the percentile over the entire SLO window (e.g. P99 of all response latencies in the past 30 days) but this is not the time-based SLI we are talking about. It’s event-based as we’ll get to.
- It is easy to tell the current status based on the recent data. That is because the status of each time slice is calculated separately. You don’t need to aggregate the data through the entire window in order to calculate the current [SLS](https://blog.alexewerlof.com/p/sls) (service level status).
### Cons
- Unless the system receives a uniform load (i.e. ~200 req/sec for the entire month), *time* and*impact* may not be correlated. For example, if the system was down for 30 minutes due to heavy demand right after launching a new product, this disruption is considered as serious as 30 minutes in the middle of the night when most users don’t use the system.
- A few bad events can easily hide in an aggregation period and go under the radar. For example, if the average latency in a minute is in the “good” range, there can be requests in the same aggregation period which their latency in not in the “good” range. For these cases it is better to use [percentile](https://blog.alexewerlof.com/p/percentile) .
- The time-based SLIs usually miss the notion of working hours and assume a global 24/7 service. This doesn't make sense for some products. For example, if a product is supposed to be used only during working hours in a certain time zone, it doesn’t make too much sense to have on-call during night and weekends. As we’ll see in an upcoming article, this is not a show-stopper because the alert notification can be decoupled from the failure detection. In other words, the system can be down outside office hours and the alert can only trigger during the office hours.
- Excluding planned maintenance windows from time-based SLIs can be complicated.
- In 🔴alerting on SLOs we’ll discuss how the speed at which we consume our error budget (burn rate) will be used to trigger an alert. Time-based alerts will have to look back at the predefined time window (based on the burn rate) to detect an incident. This makes the alerts too slow compared to event-based SLIs which trigger as soon as a good/valid ratio drops below the SLO.
# Event Based SLI
The focus is on the number of good events divided by the total number of events.
In the same example as before:
### Example
Let’s reuse the previous example. Suppose that for a database we have decided to use the query response time as SLI. We define a threshold (for example one second) and any individual query that takes longer than one second to response is considered bad.
The diagram below shows the total number of events over time. It changes as the demand for our system fluctuates.
![](https://substackcdn.com/image/fetch/$s_!mOQ5!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F711ee3e1-ef14-4237-913a-2f84a7ddfd55_1094x876.png)
Event based SLIs are interested in the percentage of good events from valid events.
### Pros
- They automatically adjust to the amount of load
- Better map to impact. If 10x more requests got error in the same amount of time, this affects the SLI and error budget 10x.
- In 🔴alerting on SLOs we’ll discuss how the speed at which we consume our error budget (burn rate) will be used to trigger an alert. Event-based SLIs will adjust automatically if the error budget is consumed faster than predicted. For example, if the error budget is being consumed at the rate of 1000x, the event-based alert will quickly pick it up.
### Cons
- Unlike time-based SLI, it is hard to tell the status based on recent data (there’s no time slice). We need the data for the entire evaluation window.
- It is harder to reason about the error budgets for event based SLIs.
- It is more punishing the team when a high number of bad events happen in a short time and can practically burn the entire error budget. One could argue that that these spikes in load are exactly why we’re measuring service levels in the first place.
# Conclusion
Which one you pick boils down to a few question:
- Is the reliability perceived as good time or good events?
- Does the service consumption pattern vary a lot during time? For example does the load drop significantly during the night or weekends? Do you get massive spikes in high season? This means different time slices have different impact and an event-based SLI suits better.
- Time-based is less punishing for the team when incidents happen during high traffic.
- How should the error budget be formulated?
- Time-based SLIs consume the error budget based on the duration of bad time
- Event based SLIs consume the error budget based on the proportion of bad events1.
- Do you want easy math? Time-based has easier in calculations.
- Do you want to know exactly how much error budget is remaining at any given time? Event-based SLIs don’t give an accurate estimate of the remaining error budget if you can’t predict the future load.
- Do you want the SLI to map to the impact? Event-based is more accurate and maps better to the impact. It considers the impact of low/high load and spikes in the load.
# References
- [Request-based and Window-based SLIs](https://cloud.google.com/stackdriver/docs/solutions/slo-monitoring#defn-sli) (Google Cloud SLO Monitoring)
- [SLI Types](https://www.ibm.com/docs/en/instana-observability/current?topic=instana-service-level-objectives-slo#sli-types) (IBM Instana Observability)
- [SRE fundamentals: SLIs, SLAs and SLOs](https://cloud.google.com/blog/products/devops-sre/sre-fundamentals-slis-slas-and-slos) (google cloud blog)
- [Uptime.is](https://uptime.is/) : a simple uptime calculator
If there only 2 types of SLI metrics, namely event and time based SLIs, from which we can derive availability SLO. Does it imply that there are also only ever 2 types of Availability metrics ?

View File

@@ -0,0 +1,13 @@
# I Don’t Alert on Apdex. It Confuses Me
- **期号**: SRE Weekly Issue #387(2023-08-27)
- **作者**: Boris Cherkasky
- **链接**: https://cherkaskyb.medium.com/i-dont-alert-on-apdex-it-confuses-me-5e639242e5db
## 简介
> This is my personal take on something that is considered standard that I just don’t understand. So here we go — the Apdex, what it is, and why I don’t use it!
## 正文
> ⚠️ 抓取失败:HTTP 403

View File

@@ -0,0 +1,13 @@
# Reader: How the video “Three analytical traps in accident investigation” Helps me be a Better Incident Analyst
- **期号**: SRE Weekly Issue #387(2023-08-27)
- **作者**: Randy Horwitz — Learning From Incidents
- **链接**: https://www.learningfromincidents.io/posts/how-the-video-three-analytical-traps-in-accident-investigation-helps-me-be-a-better-incident-analyst
## 简介
Here’s a great explanation of three common cognitive biases we should try to avoid while analyzing incidents.
## 正文
> ⚠️ 抓取失败:URLError: [SSL: SSLV3_ALERT_HANDSHAKE_FAILURE] ssl/tls alert handshake failure (_ssl.c:1032)

View File

@@ -0,0 +1,191 @@
# Lily Cohen (@lily) Re: firefish.lgbt, musician.social, and outdoors.lgbt
- **期号**: SRE Weekly Issue #387(2023-08-27)
- **作者**: Lily Cohen
- **链接**: https://firefish.social/notes/9iqefgi8rzfksnqc
## 简介
A horrifying tale of gitops gone wrong and backups that didn’t back up, leading to catastrophic data loss. This, this is what hugops is for. I’m so sorry, Lily!
## 正文
Premium domain · For sale
# firefish.social
A short, memorable, established domain ready to power your brand. Backed by 1,536 referring domains and 3 years of online authority.
[Buy firefish.social](https://firefish.social/domain/firefish.social/backlink?hash=4b24b8)
[Not interested in buying this domain?Browse relevant content instead](https://firefish.social)
8-character brandable .social
8 characters · 3 years old
Buy-it-now
$100
USD
- Afternic
- GoDaddy checkout
What happens after you buy
Pay
Secure checkout on GoDaddy
Verify
Ownership confirmed
Push
Delivered within 24h
GoDaddy-protected checkout
Every claim below is backed by verified third-party data.
Currently listed below our AI fair-value estimate — room for upside on resale or development.
Asking
$100
AI fair value
$413
TLD
.social
Demand signals indicate strong ranking potential out of the box.
CPC
$0.00
Short, easy to say, easy to type — the foundation of any premium brand.
Length
8
Radio test
Passes
Appeal
3.0
Why this name
**firefish.social** is a compact 8-character name — long enough to be descriptive, short enough to stay memorable. The .social extension is a working SOCIAL name with everything pre-indexed. 1,536 referring domains link back to it — equity you can keep by simply redirecting. For founders launching their next product looking to launch something distinctive, this is the kind of pickup that pays for itself the first time someone reads it out loud.
Verified from public sources at the time of listing. Some advanced metrics require a free account.
🛡 Authority & SEO
Moz DA
20
Moz PA
38
Trust Flow
18
Citation Flow
43
Moz spam score
9
✦ Brandability
Dashes
0
Numbers
Appeal score
🕰 History
Age
3 years
Wayback snapshots
57
First seen
2023
Want to see every metric?
Free account unlocks advanced SEO data, side-by-side comparisons, and price-drop alerts on domains you're watching.
As the world's largest domain registrar with over **20 million customers**, GoDaddy manages more than **84 million domains**, offering a robust platform with **24/7 customer support** and a wide range of TLDs. Its user-friendly interface and comprehensive services, including hosting and website builders, make it ideal for businesses seeking an all-in-one solution.
Professional Trust
Used by SEOs, marketers, and investors all over the world.

View File

@@ -0,0 +1,149 @@
# Authentication slowness or failure to load Duo Prompt on DUO1
- **期号**: SRE Weekly Issue #387(2023-08-27)
- **作者**: —
- **链接**: https://status.duo.com/incidents/rw7g0q7ztj8f
## 简介
Here’s a followup analysis from Duo for an incident they had last week.
## 正文
Incident
Report for [Duo](https://status.duo.com/)
**Authentication Failures on DUO1**
Incident Report - 2023-08-21
Duo Engineering recreated the outage in a test environment, and then successfully validated a fix under load conditions, to the underlying root cause of the outage. The root cause was a specific database query pattern that had a higher probability of deadlocks, getting exacerbated under heavy load. A combination of increased activity, deadlocks, and retries overwhelmed certain shards, causing DUO1 to be unresponsive. The performance test that reproduced this scenario will now be part of our non-functional testing to better gate future releases.
From 9:03 EDT to 14:16 EDT on Aug 21, 2023, the DUO1 deployment experienced increased authentication latency that caused authentication failures for some customer applications protected by the Duo Service. At 9:03 AM Duo’s automated monitoring systems detected and alerted engineering team members to a database latency that was elevated but that did not impact service. Within minutes, while engineering team members investigated and triaged related issues, latency increased to service-impacting levels. A detailed investigation by Duo Engineering identified the application layer as the bottleneck. The team began horizontally scaling the application layer to absorb the increased load. The Duo team began to see recovery around 12:50 EDT, fully resolving by 16:01 EDT. Duo Engineering continues to monitor.
- DUO1
**09:03** Duo Engineering received an alert indicating Authentication Latency on DUO1.
**09:11** Duo Engineering received a Database Backlog Depth alert, indicating replication latency.
**09:14** Duo Engineering started investigating, confirming latency reports via our monitoring tools.
**09:22** Duo Engineering initiated a full incident response process bringing in Customer Support (CS) and other stakeholders to assist with the investigation and communication.**09:34** Status Updated to `Investigating.` 
**09:28** Determined only DUO1 deployment was affected, investigative efforts continue.
**09:36** Back pressure, an automated mechanism that allows our systems to recover safely, came into play, rejecting requests.
**09:40** Duo Engineering verified alerts that pointed to latency in the database layer, validating performance metrics and investigating any slow queries.
**10:01** Duo Engineering also investigated the load balancer and application layer for performance insights.
**10:26** Back pressure began to subside.
**10:44** Multiple impacted customers identified, investigative efforts continued.
**10:45** Options for mitigation were being considered for easing load on service and the potential impact of each option.
**10:45** Additional alerts received, around our AzureAuth service. An investigation was spun up for this alert, to ascertain if they were related.
**11:16** Duo Engineering determined that the AzureAuth incident was caused by the current ongoing incident impacting DUO1.
**11:45** All AzureAuth instances had self-healed.
**12:11** Duo Engineering considered options for scaling DB layer, until investigation pointed to application layer being the bottleneck (at **12:24** EDT).
**12:42** Duo Engineering began increasing application capacity by scaling horizontally.
**12:54** Some customers began recovering as a result of capacity increases.
**13:47** Duo Engineering continued to monitor after the rollout of additional capacity.
**14:16** Additional capacity was in place and taking traffic <- Issue resolved at this point and we entered the observation phase before reporting back to customers.
**14:33** Duo Engineering requested that customers be notified of the update, and asked to re-enable MFA if turned off.
**14:37** Status changed to `Monitoring.` 
**14:58** Duo Engineering identified additional preventative measures to balance load, to prevent recurrence.
**16:01** Status updated to `Resolved.`
Increased load on DUO1 due to significantly increased adoption and simultaneous peak usage across multiple larger customers led to authentication failures. Our monitoring system picked up on latency metrics that alerted Duo Engineering. During this time our automated back-pressure system became active. This system tracks request volume and latency, and if latency is increasing, returns a HTTP Status Code 429 to clients, requesting a retry. This back-pressure mechanism is enabled on all deployments to allow our systems to recover safely and prevent them from being overwhelmed. The team’s investigation considered, evaluated, and ruled out multiple potential causes. This outage was not due to an attack, but the root cause being a significant spike in traffic with a surge of overlapping peak usage, causing capacity issues at the application server tier. Duo Engineering responded by adding more application servers to the load balancer. As a result, the back-pressure began to release and our systems began processing authentications within normal response time ranges.
After service was restored, Duo Engineering continued to investigate and identified overlapping peak usage across multiple customers, with similar usage patterns, leading to contention for shared resources.
Based on peak usage patterns observed during the outage, the team made additional capacity increases to avoid a recurrence. This was completed on Aug 21 at 04:09 EDT. Duo Engineering monitored DUO1 closely the following day, Aug 22nd, between 7:30 EDT and 12:00 EDT, and confirmed that these capacity levels handled the elevating volume experienced, without triggering additional alerts.
Data collected during and after the incident will be used to influence near-term and future capacity-related decisions to ensure that appropriate headroom is available and to improve load distribution with no downtime for customers. Careful capacity planning has been one of pillars of Duo’s service management and that will continue with lessons learned from this incident.
Duo Engineering is in the midst of a multi-quarter re-architecture of our services, which upon completion will enable our services to scale automatically in response to load and be more resilient. The team will also evaluate and enhance current load testing methodologies to better understand system performance and recovery under duress.
Duo Engineering, continuing post incident analysis, identified an excessive write to the database on August 22nd at 01:56 AM, that when eliminated will ease table locking and improve shard performance.
Duo Engineering has identified a number of measures to be integrated to our response runbooks and is implementing additional automation to ensure easy access to these, during an outage, to decrease our response time. Additional training is being added to our training program for current and new engineers, to decrease our response time.
At Duo Engineering, we are committed to learning and improving our service every day, as we go about the business of securing our customers. We will continue to study other observed symptoms and update this document with additional details as they become available.
The issue causing authentication failures on DUO1 has been fully resolved. All authentications are working as expected.
We will be posting a root-cause analysis (RCA) here once our engineering team has finished its thorough investigation of the issue.
Please check back or subscribe to be notified when the RCA is posted.
A fix has been implemented for the authentication failures on DUO1 and we are monitoring the results.
Please check back here or subscribe to updates for any changes.
We are continuing to increase capacity to resolve the authentication failures on DUO1. Systems have started to recover.
Please check back here or subscribe to updates for any changes.
We are continuing to work on a fix to resolve the authentication failures on DUO1.
Please check back here or subscribe to updates for any changes.
We are continuing to work towards a resolution for these errors.
Please check back here or subscribe to updates for any changes.
We have identified the issue causing authentication slowness and failures to load the Duo Prompt and are working toward resolution.
Please check back here or subscribe to updates for any changes.
We are continuing to investigate authentication errors on DUO1 and are working to correct the issue as soon as possible.
Please check back here or subscribe to updates for any changes.
We are currently investigating authentication errors on DUO1 and are working to correct the issue as soon as possible.
Please check back here or subscribe to updates for any changes.
This incident affected: DUO1 (Core Authentication Service, Push Delivery, SSO).

View File

@@ -0,0 +1,104 @@
# Practical Guidance for First-Time Site Reliability Engineers
- **期号**: SRE Weekly Issue #387(2023-08-27)
- **作者**: Ben Wheatley — The New Stack
- **链接**: https://thenewstack.io/practical-guidance-for-first-time-site-reliability-engineers/
## 简介
The first SRE hire at incident.io shares what they learned as they became familiar with the infrastructure and figured out what to do with it.
## 正文
# Practical Guidance for First-Time Site Reliability Engineers
![Featued image for: Practical Guidance for First-Time Site Reliability Engineers](https://cdn.thenewstack.io/media/2023/08/1c0a6c62-reliability-1024x734.jpg)
[Incident.io](https://incident.io/?utm_content=sponsor+disclosure)sponsored this post. Insight Partners is an investor in Incident.io and TNS.
At the beginning of May, I joined [incident.io](http://incident.io) as the first site reliability engineer (SRE), a very exciting but slightly daunting move.
With only some high-level knowledge of what the company and its systems looked like prior to this point, it’s fair to say that I didn’t have much certainty in what exactly I’d be working on or how I’d deliver it.
After joining and having settled in after a day or two, my mission became clear: The primary objective was to build a roadmap for our infrastructure, and then set out to deliver it.
At this company, reliability is something that’s valued even more than at most others; as providers of the tooling you depend on to pick up when your own systems are broken, we become a critical dependency and need to have our product available whenever you need us.
This helps set some initial context for what might be going into the roadmap: an emphasis on [availability and reliability](https://thenewstack.io/experts-weigh-in-on-the-state-of-site-reliability-engineering/).
But if you find yourself in this kind of position, how do you start here and produce a roadmap, starting at zero context?
Here are my tips and advice for breaking down this problem.
## **Getting Started**
Starting with a blank space where you might usually expect to have a roadmap, coming from working in organizations where many years were spent considering these kinds of topics, is a different challenge but not an insurmountable one.
Here are a few strategies that might help you build up context, find the problems that really matter and turn these into a plan of action.
## **Get Well Acquainted with Infrastructure and the Code, Too**
Before diving in and making any changes, it’s obviously pretty vital to get a good feeling for what the current setup is within an organization.
Part of an SRE’s job, especially within a smaller team, is to enable and accelerate engineers, so you’ll need to both build empathy around the daily processes that engineers go through and have a sufficient technical understanding to make small changes to the product when required.
Definitely spend some time in “product land.” You should have a fully functioning [development environment](https://incident.io/blog/developer-environments-should-be-cattle-not-pets), a good understanding of the structure of the primary codebase and be able to put this into practice by picking up some smaller changes and delivering them from start to finish.
If your organization has a model like our [Product Responders](https://incident.io/blog/how-we-leverage-our-product-responder-role), pairing with those will give you a feel for some of the gnarlier issues.
Going through this process, having stepped into the shoes of a product engineer should present some great opportunities to learn not just about the core product, but also the details around deployments, [observability](https://thenewstack.io/observability/) tooling and data stores.
## **Talk with as Many People as You Can**
Now this may sound obvious, but learning as much context as possible from those who’ve been living and breathing the systems you’re taking on responsibility for is going to be crucial.
Whenever you’re not working on onboarding or building up technical knowledge, try to fill the gaps with chats over coffee, going outside for a walk or grabbing lunch together. Beyond building relationships, which is important in itself, this is a great opportunity to find out about current pain points, tools they wish they had and any projects that had been deferred “until we have an SRE.”
Especially important is remembering not just to limit yourself to talking to those grizzled veteran engineers who’ve seen it all, but also the new joiners who may have useful viewpoints, the product and engineering managers who get a good aggregate view from the engineers that they work with, and leadership who will have useful input on longer-term vision and vendor relationships.
## **Keep a Finger on the Pulse of Your Customers**
Keep an eye out for whenever your organization’s customers are getting in touch with any issues that relate to infrastructure or shared concerns, whether that’s through asking the customer support team to keep you in the loop or monitoring your internal incidents channel. If the opportunity is there, you should join discussion and talk with the customer directly, which will allow you to dig into the details further.
These types of interactions may be less frequent than others, but they are very valuable, as they give you an insight into what customers value (such as latency vs. availability) and how they interact with the product, and allow you to start understanding what kinds of trade-offs you can make further down the line.
## **Don’t Limit Yourself to Just Your Peers**
As an SRE at an early-stage company, there’s a good chance that you’ll need to bring in new platforms, tools and processes. There’s also a reasonable chance that these will look different from where you’ve worked previously. Perhaps the business needs are different, or it’s just that industry trends have evolved beyond the systems you’ve worked with previously.
As you start to build up ideas for what kinds of changes you’d like to make, you might find that being the sole SRE makes it tricky to know if you’re on the right track. It’s really useful to validate ideas like this with people outside of your organization too.
Is the hot, new container technology you’re looking at not as good as it’s cracked up to be?
Perhaps contacts at similar-sized companies will have some insight. Similarly, if your company has existing relationships with platform vendors, then lean on them by sending your thoughts and proposals over to the account manager. They may be able to tell you whether you’re following best practices and what their recommendations are based on similar-sized orgs.
## **Pulling It all Together**
If you’ve followed some of the suggestions above, then you should now have a good feeling about the issues and missing building blocks within the organization and be able to make some informed decisions about the next steps.
You’ll also have a lot of diverse inputs from different stakeholders, so you’ll need to distill all of this into a sensible roadmap. The key to this will be picking out common themes in the information by applying some grouping, but after that, you’ll need to make some calls about how to tackle the problems.
The insights you’ve gained into the business, customers, engineer workflows and the product should help you out here. Be careful not to plan too far into the future — you should focus on the problems that are causing pain right now, and then revisit them in a few months’ time, rather than trying to set out a multiyear strategy from the beginning.
To make this all more concrete, here’s the summarized roadmap that I created.
1. **Theme: Compute** — How we deploy and run our core codebase
- Better control of deployments and how we cut over to new versions.
- Improving reliability and performance of routing HTTP requests to our application code.
- Improving observability around the system-level metrics of our containers (such as CPU, memory, open file handles).
2. **Theme: Database** — How we store the data that powers our product
- Gaining the ability to do PostgreSQL major version upgrades with minimal downtime.
- Improving observability about what’s happening within the database.
3. **Theme: Observability** — How we monitor the health of our product and systems
- Gaining the ability to capture application-level metrics. We work with logs and [traces](https://incident.io/blog/tracing) , but we’d like to instrument the application further.
- Improving how we store logs and how they can be used effectively.
4. Gaining the ability to capture application-level metrics. We work with logs and
Distilling all of this into a [document](https://thenewstack.io/why-docs-as-code-should-be-part-of-your-dev-cycle/) and sharing it with key stakeholders should give you the buy-in to go after the problems you need to tackle.
[YOUTUBE.COM/THENEWSTACK
Tech moves fast, don't miss an episode. Subscribe to our YouTube
channel to stream all our podcasts, interviews, demos, and more.](https://youtube.com/thenewstack?sub_confirmation=1)

View File

@@ -0,0 +1,103 @@
# Keeping the Lights On: The On-Call Process that Works
- **期号**: SRE Weekly Issue #387(2023-08-27)
- **作者**: Felix Lopez — The New Stack
- **链接**: https://thenewstack.io/keeping-the-lights-on-the-on-call-process-that-works/
## 简介
This is a story of building a new on-call rotation in a company that didn’t have one. They started out with a pretty awesome list of principles that we could all aspire to.
## 正文
# Keeping the Lights On: The On-Call Process that Works
![Featued image for: Keeping the Lights On: The On-Call Process that Works](https://cdn.thenewstack.io/media/2023/08/c29649dd-on-call-1024x682.jpg)
[Tinybird](https://www.tinybird.co?utm_content=sponsor+disclosure)sponsored this post.
The On-call process is a touchy subject for a SaaS company. On the one hand, you must have it, because your prod server always seems to go down at 2 a.m. on a Saturday. On the other hand, it places a heavy burden on those who must be on call, especially at a small company like [Tinybird](https://www.tinybird.co), where I currently head the engineering team.
I have actively participated in creating the on-call process in three different companies. Two of them worked very well, while the other didn’t. Here, I’m sharing what I’ve learned about making on call successful.
## Before a On-Call Process: Stress and Chaos
When I joined Tinybird, we didn’t have an [on-call system](https://thenewstack.io/how-can-you-tell-if-your-on-call-system-is-broken/). We had automated alerts and a good [monitoring](https://thenewstack.io/observability/) system, but nobody was responsible for an on-call process or a rotation schedule between employees.
Many young companies like ours don’t want to create a formal on-call process. Many employees justifiably shy away from the pressure and individualized responsibility of being on call. It seems better to just handle issues as a hive.
But in reality, this just creates more stress. If nobody is responsible, everybody is responsible.
In the absence of a formal process, Tinybird relied on proactive employees and mobile notifications for some of our alert channels. In other words, it was disorganized, unstructured and stressful. We had multiple alert channels, constant noise and many alerts that weren’t actionable. If that sounds familiar, it’s because this is typical in most companies.
Obviously, this approach to handling production outages doesn’t scale, and it’s a recipe for poor service and disgruntled customers. We knew we needed a formal on-call structure and rotation, but we wanted to avoid overwhelming our relatively small team (at the time, we had less than 10 engineers).
## How It Started: Implementing an On-Call Process
People don’t want to an on-call process. They’re afraid that this on-call experience will look like their last on-call experience, which inevitably sucked. Underneath that fear is insecurity about trying to solve a problem you know little about when nobody is around (or awake) to help. And that burden of responsibility weighs heavily. Sometimes you have to make a decision that can have a big impact. Downtime can be caused by the difference between a 0 and a 1.
Our goal at Tinybird was to assuage those fears and insecurities so that people felt empowered and respected as on-call engineers.
Before we even discussed a process, we outlined some core principles for the on-call system that would provide boundaries and guidance for our implementation.
### Core Principles for an On-Call Process
- **On call is not mandatory.** Some people, for various reasons, do not want or are not able to be on call. We respect that choice.
- **On call is financially compensated.** If you are on call, you get paid for your time and energy.
- **On call is 24/7.** We provide a 24/7 service, and we must have somebody actively on call to maintain our SLAs.
- **Minimize noise.** Noise makes stress. If alerts aren’t actionable, this stress will cause burnout. Alerts must always be actionable.
- **On call isn’t just for SREs (site reliability engineers).** Every engineer should be able to participate. This promotes ownership among all team members and increases everyone’s awareness and understanding of our systems.
- **Every alert should have a runbook.** Since anybody from any function could be on call, we wanted to make sure everyone knew what to do even if the issue wasn’t with their code or system.
- **Minimize the amount of time spent on call.** Our goal was to only have people be on call once every six weeks. Of course, depending on how many people participate, this may not be achievable, but we set it as a target regardless.
- **Have a backup.** Our service-level agreements (SLAs) matter, so we always wanted to have a backup in case our primary on-call personnel were unreachable for whatever reason.
- **Paging someone should be the last resort.** Don’t disrupt somebody outside of working hours unless it is absolutely necessary to maintain our SLAs. Additionally, every time an incident occurs, measures should be taken to prevent its recurrence as much as possible.
### How We Implemented an On-Call Process
So, here’s how we approached our on-call implementation.
First, we made a list of all our existing alerts. We asked two questions:
1. **Are they understandable?** Any of our engineers should be able to see the alert and understand the nature and severity of it very quickly.
2. **Are they actionable?** Alerts that aren’t actionable are just noise. Every alert should demand action. That way, there’s no doubt about whether action should be taken when an on-call alert pops into your inbox.
Second, as much as possible, we made alerts measurable, and each one pointed to the corresponding graph in Grafana that described the anomaly.
In addition, we migrated all of our on-call alerts to a single channel. No more hunting down alerts in different places. We used PagerDuty for raising alerts.
Critically, we created a runbook for each alert that describes the steps to follow to assess and (hopefully) fix the underlying issue. With the runbook, engineers feel empowered to solve the problem without having to dig for more context.
For about two months, every Monday, Tinybird’s CTO and I would meet with the platform team to review each alert with the following objectives:
1. If the alert was not actionable or was a false positive, correct or eliminate it.
2. If the alert was genuine, analyze it to find a long-term solution and give it the necessary priority.
We also started reviewing each incident report collaboratively with the entire engineering team. Before we implemented this process, we would create incident [reports](https://thenewstack.io/why-docs-as-code-should-be-part-of-your-dev-cycle/) (IRs) and share them internally, but we decided to take it a step further.
Now, each IR is presented to the entire engineering team in an open meeting. We want everyone to understand what happened, how it was resolved and what was affected. We use the meeting to identify action points that can prevent future occurrences, such as improving alerts, changing systems, architecture changes, removing single points of failure, etc. This process not only helped us mitigate future issues but also helped increase ownership and overall knowledge of our code and systems across the entire team. The more people know about the codebase, the more they feel empowered to fix something when they are on call.
Initially, we had just three people on call (two engineers and the CTO). We knew this would be challenging for these three, but it was also a temporary way to assess our new process before we rolled it out to the entire team.
Note that we still made on call mandatory during working hours. Each engineer is expected to take an on-call rotation during a normal shift. This has several benefits:
1. **Increased ownership:** Being on call makes you realize the importance of shipping code that is monitored and easily operable. If you know you’re going to be on call to fix something you shipped, you’ll spend more time making sure you know how to operate your code, how to monitor it and how to parse the alerts that get generated.
2. **Knowledge sharing and reduced friction:** Being on call can feel scary when you’re alone. But if you’re on call during working hours, you are not alone. For newcomers, this helps them ease into the on-call process without anxiety. They learn how to respond to common alerts, and they also learn that being on call isn’t as noisy or scary as they think.
Every week, when the on-call shift changes, we review the last shift. We use this time to share knowledge and tricks, identify cross-team initiatives necessary to improve the system as a whole and so on.
Finally, anytime a person is the primary person on call overnight, we give them the next day off.
## How It’s Going: Where Are We Now?
After about a year of implementing this new on-call process, we have nine people rotating on the primary (24/7) on-call system and six people simultaneously on call during working hours.
It has worked exceptionally well. While I won’t go so far as to say that our engineers enjoy being on call, I think it is fair to say that they feel empowered to handle issues that do arise, and they know they have a forum where they can share difficulties about the on-call system and suggestions for how to improve them.
If you’re interested in hearing more about [Tinybird](https://www.tinybird.co) and our on-call system, I’d love to hear from you. Also, if you’re inspired by Tinybird’s on-call culture and think you’d like to work with us, check out our open roles [here](https://www.tinybird.co/about#join-us).
[YOUTUBE.COM/THENEWSTACK
Tech moves fast, don't miss an episode. Subscribe to our YouTube
channel to stream all our podcasts, interviews, demos, and more.](https://youtube.com/thenewstack?sub_confirmation=1)

View File

@@ -0,0 +1,13 @@
# Test In Production
- **期号**: SRE Weekly Issue #387(2023-08-27)
- **作者**: Sven Hans Knecht
- **链接**: https://medium.com/@hans.knechtions/test-in-production-85224e7a82f3?source=rss-f2ac9bbfd2bb------2
## 简介
Why should we test in production? This article gives a really spot-on argument and goes on to explain how to do it.
## 正文
> ⚠️ 抓取失败:HTTP 403