Files
nexus/sreweekly/markdown/297/02-5-ways-incidents-made-me-a-better-engineer.md
2026-09-12 17:23:01 +08:00

78 lines
6.3 KiB
Markdown
Raw Permalink Blame History

This file contains invisible Unicode characters
This file contains invisible Unicode characters that are indistinguishable to humans but may be processed differently by a computer. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# 5 ways incidents made me a better engineer
- **期号**: SRE Weekly Issue #297(2021-11-21)
- **作者**: Lisa Karlin Curtis — incident.io
- **链接**: https://incident.io/blog/incidents-made-me-a-better-engineer
## 简介
This article really gets to the heart of why I love a good incident. I mean, obviously, I want to minimize, incidents. I swear.
## 正文
November 16, 2021 — 5 min read
Incidents are a great opportunity to gather both context and skill. Understanding the [incident response lifecycle](https://incident.io/blog/what-is-the-incident-response-process) helps teams solve unexpected and challenging problems more effectively.
In my career, I've found incidents can be a great accelerator - for both myself and others around me. It was after leading my first incident at GoCardless that I started to feel really comfortable in the codebase and the team - this is what [building a culture of incident response](https://incident.io/blog/building-a-culture-of-incident-response) does for everyone. I had the same experience joining [incident.io](https://incident.io) (yes we do have incidents, and yes it is quite 🤯).
Incidents often occur at the edges of teams. That makes them a great chance to learn about stuff that isn't in your day-to-day remit.
The obvious example for me is infrastructure: at GoCardless we had an infrastructure group who provided a platform for us to deploy our services. I didn't interact with the infrastructure directly much in my first few months, so didn't have a strong mental model of how any of it fit together. That was a huge limitation on the kinds of problems I could solve. I couldn't make good decisions about how to best use our database, or how to manage asynchronous work, as I didn't understand the trade-offs.
Watching people solve incidents was the entrypoint I needed to start investigating and understanding our infrastructure, and how it connected to my day-to-day trade-offs.
Incidents are usually caused by (or manifest in) the most difficult parts of the systems we interact with. Seeing multiple incidents impact the same component is great way to learn about that component, while simultaneously signalling that understanding the component will be valuable.
I've been introduced to a number of domain areas via incidents including database replication (often the culprit), quorum (terrifying) and DNS (a classic). After getting some initial context during an incident, I could then spend some time reading about these concepts with confidence that it would prove useful.
We're not perfect: our job is hard and our code is very likely to go wrong at some point. Instead of trying to write perfect code, incidents have shown me that it's more important to make code that fails in safe ways. This includes:
- Making code alert loudly and clearly if it sees something that 'can't happen' (famous last words). Ideally, the alert should be easy to trace to a code comment, doc or commit message explaining why (when you wrote the code) you didn't think this would happen.
- Keep the blast radius for failures as small as possible: think carefully about what should be considered 'critical' for a given request, and get everything else out of the way. Being unable to log a user tracking event should never degrade the customer experience.
While it's possible to read this stuff in textbooks, seeing the impact of these choices in real incidents is what taught me how to put this advice into practise.
I've been in many incidents where a graph or set of log lines has been the key bit of information to help diagnose the problem. It's also usually what tells us that the incident is over. Finding components with poor observability can be really stressful: it's like someone blindfolding you and asking you to find the front door. Possible, but not fun or efficient.
Watching more experienced colleagues use [incident response tools](https://incident.io/incident-response-slack) and observability dashboards, and then using them myself, taught me how to get the information I needed quickly. Once you understand how to use the information that's already there, it's easier to understand what other information would be useful when working on other projects.
Incidents help you map your [incident response team](https://incident.io/blog/how-to-structure-incident-response-teams) and meet people outside your day-to-day circle. Many of the colleagues I respected, valued and relied on most were not people I worked with day-to-day. It's a great change to find people who have different skill sets from your usual team mates. Maybe there's someone who knows lots about a particular technology, or someone who is a really great teacher. Having a network of talented people I could ask for advice has been the single most impactful accelerator for my growth.
- Get involved in incidents from day one! Even if you’re only at the very start of your career.
- Be respectful of other people's time and situation. Observe quietly at first, note down questions to ask later.
- Be honest with yourself and others about what you can and can't do alone. [Psychological safety in incident management](https://incident.io/blog/psychological-safety-in-incident-management) means it's safe to say "I want a pair" or "I need help." That's a great way to learn, but depending on the situation might not be appropriate.
Lisa Karlin Curtis
Technical Lead
Our rate limiter depends on Valkey. If Valkey goes down we fail open and stop limiting which isn't good enough for our platform. As an intern, I built per-pod in-memory top-k buffers so we keep rate limiting even with the backing store gone.
Anthony Oparaocha
September 2, 2026
Our entire event-driven platform ran through a single message broker, which made it a single point of failure. So we added a second one. This is the story of building an event load balancer, the queuing theory behind it, and the final chaos test where we turned off Pub/Sub in production and nobody noticed.
Patrick Hamann
+
Mike Fisher
August 11, 2026
Our learnings from implementing a product-wide read replica migrations, including some useful patterns for routing queries to replica and primary
Johanna Larsson
July 21, 2026
Ready for modern incident management? Book a call with one of our experts today.
- All-in-one incident management
- Our unmatched speed of deployment
- Why we’re loved by users and easily adopted
- How we work for the whole organization