39 lines
5.8 KiB
Markdown
39 lines
5.8 KiB
Markdown
# Increment: Reliability
|
||
|
||
- **期号**: SRE Weekly Issue #259(2021-02-28)
|
||
- **作者**: Stripe
|
||
- **链接**: https://increment.com/reliability/
|
||
|
||
## 简介
|
||
|
||
This quarter’s Increment issue is about Reliability, and I haven’t had this much fun since their first issue about on-call. I’ll include a few of the articles here and more in later issues as I have a chance to review them.
|
||
|
||
## 正文
|
||
|
||
#### Heidi Waterhouse
|
||
|
||
### Everything is broken, and it’s okay
|
||
|
||
Accepting that imperfect things still work is fundamental to preventing failures from becoming catastrophes.
|
||
|
||
This issue shares approaches to reliability and resiliency in our software, technologies, and teams, and offers perspectives on the realities of failure in the systems we build.
|
||
|
||
#### Heidi Waterhouse### Everything is broken, and it’s okay Accepting that imperfect things still work is fundamental to preventing failures from becoming catastrophes.
|
||
#### Ryn Daniels### How to build organizational resilience By encoding resilience into an organization’s culture, engineering teams can be better equipped to tackle the unknown and unexpected.
|
||
#### Tanya Reilly### Embrace your inner incident commander The way we fight fires affects how quickly we can resolve outages. Appointing an incident commander can help—and you (yes, you) can be one.
|
||
#### Benoit Baudry and Martin Monperrus### Testing beyond coverage Pseudo-tested methods can be a reliability risk. Here, the authors explain how they developed a methodology and tool to uncover them in Java applications.
|
||
#### Tess Donnelly and Tiarnán de Burca### Trust is an enabling technology To build a high-performing software delivery system, your stack’s capabilities are just one part of the picture.
|
||
#### Mads Hartmann### Tracing a path to observability A chronicle of Glitch’s efforts to gain visibility into its production systems—and make them more reliable.
|
||
#### Ana Margarita Medina### Chaotic good As software systems become ever more complex, chaos engineering provides a (not-actually-so-chaotic) tool kit for building more reliable and resilient systems.
|
||
#### Safia Abdalla### Open-source excursions: Optimizing for operational resiliency Documentation, automation, and a little sharing-is-caring can help OSS projects maintain their uptime.
|
||
#### Georgios Bakirtzis### Responsible development for the physical world A case for designing consumer software with safety-critical principles and formal methods in mind.
|
||
#### Ipsita Agarwal### Interview: Dr. David D. Woods A discussion of the distinctions (and dependencies) between reliability and resilience, and how to build complex systems that perform under strain and surprise.
|
||
#### Mathieu Frappier, Dorothy Jung, and Qui Nguyen### The process: Implementing Yelp’s failover strategy How Yelp engineers orchestrated their traffic failover process and effected a delicate balance between reliability, performance, and cost efficiency.
|
||
#### John Allspaw, Beth Adele Long, and Dr. Richard Cook### On adaptive capacity in incident response Learnings for tech orgs looking to adopt a resilience engineering perspective.
|
||
#### Lara Hogan### Your brain on progress Strategies for nurturing that feel-good sense of accomplishment when doing largely invisible work.
|
||
#### Ian Steadman### Earth, wind, and solar fire If a major solar storm were to sweep across Earth, would today’s electrical and communications infrastructure be resilient enough to endure its impact?
|
||
#### Poornima Apte### Home sweet home network Facing dramatic shifts in residential usage, internet service providers are working to keep latency low and connectivity high.
|
||
#### Increment Staff### Reliability at scale Leaders at Deliveroo, DigitalOcean, Fastly, and Headspace share how their organizations think about reliability and resiliency and their advice to engineering orgs embarking on reliability journeys.
|
||
#### Ipsita Agarwal### Case study: Resilience as adaptability at Freshworks The company’s disaster preparedness plan, developed in the aftermath of a devastating cyclone, enabled it to adapt and endure during a global pandemic.
|
||
#### Chris Stokel-Walker### Case study: How Akamai weathered a surge in capacity growth Seeing a year’s worth of capacity growth in a matter of weeks, the CDN services provider hustled to build and reinforce the infrastructure it needed to serve its users (and European soccer fans).
|