Files
nexus/sreweekly/markdown/11/03-inside-atlassian-how-our-site-reliability-engineers-do-incident-manage.md
2026-09-12 17:23:01 +08:00

981 B
Raw Permalink Blame History

Inside Atlassian: how our site reliability engineers do incident management

简介

Atlassian dissects their response to a recent outage and in the process shares a lot of excellent detail on their incident response and SRE process. I love that they’re using the Incident Commander system (though under a different name). This could have (and probably has) come out of my mouth:

The primary goal of the incident team is to restore service. This is not the same as fixing the problem completely – remember that this is a race against the clock and we want to focus first and foremost on restoring customer experience and mitigating the problem. A quick and dirty workaround is often good enough for now – the emphasis is on “now”!

正文

⚠️ 抓取失败:HTTP 403