Files
nexus/sreweekly/markdown/522/01-incident-status-updates-are-a-translation-problem-and-the-right-transl.md
2026-09-12 17:23:01 +08:00

96 lines
13 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Incident status updates are a translation problem, and the right translator probably isn’t in Engineering
- **期号**: SRE Weekly Issue #522(2026-06-21)
- **作者**: Brent Chapman
- **链接**: https://greatcircle.com/blog/2026/03/31/status-update-translation/
## 简介
> […] the fix isn’t “train your engineers to write better status updates.” The fix is to stop asking your engineers to write them, and start asking the right people instead.
## 正文
It’s mid-morning in Europe, and your customers are complaining about stale data. It’s 2 AM for your on-call engineers. Your European customer care team noticed the surge in complaints and paged the on-call incident commander (IC). The IC pulled up the dashboards, saw a spike in processing errors on the data ingest pipeline, and paged the data engineering, platform, and networking on-call engineers. Now the IC is coordinating with three responders, all trying to figure out what’s gone wrong, and whether the problem is actually even worse than it appears. And on top of all that, the IC is expected to write a customer-facing status page update.
The result is predictable. The update either doesn’t happen at all, or it reads something like: “Elevated error rates on the ingest pipeline due to Kafka consumer lag.” Which is perfectly accurate, and perfectly useless to 99% of the people reading it.
This is one of the most common patterns I see in my consulting work, and it’s worth understanding why it happens, because the fix isn’t “train your engineers to write better status updates.” The fix is to stop asking your engineers to write them, and start asking the right people instead.
# The translation problem
Here’s the core issue: incidents generate information in a language that most of your company doesn’t speak, and most of your customers don’t either.
When responders are working an incident, they talk about services, endpoints, error rates, deployment pipelines, and database replicas. They refer to systems by internal codenames. They describe impact in terms of error budgets and percentile latencies. This is exactly the right level of detail for the people doing the troubleshooting, and exactly the wrong level of detail for almost everyone else.
Your customer support team needs to know which customer-visible features are affected, described in the language customers use. Your sales team needs to know whether the demo environment is impacted and what to say to the prospect they’re meeting in an hour. Your executives need to know the business impact: revenue at risk, customers affected, estimated time to resolution. Your customers need to know that you’re aware of the problem, you’re working on it, and roughly when it will be fixed.
None of these audiences need to know that Kafka consumer lag on the ingest pipeline is causing records to back up in the processing queue. They need to know that some recently submitted data may not be appearing yet, and that your team is on it.
Most companies treat this as a writing-quality problem, but it’s a translation problem. Each of these audiences speaks a different dialect: customer-speak, business-impact-speak, relationship-speak, risk-speak. Your engineers are fluent in engineering-speak, which is precisely what you want from them during an incident. Expecting them to also be fluent in four other dialects, under pressure, at 2 AM, is setting everyone up for disappointment.
Consider how much translation is involved in turning “Kafka consumer lag on the ingest pipeline. Stuck consumer group identified; restarting. Backlog clearing, ETA 30min” into “Some users may notice that recently submitted data isn’t appearing yet. Our team has identified the cause and is deploying a fix. We expect this to be resolved within the next 30 minutes.” Internal service names became customer-visible symptoms. A technical explanation became “identified the cause.” An engineering action became an estimated timeline. That’s the work, and it’s work that customer-facing people do better than engineers.
# The right people for the job
The fix is structural. Staff your incident communication functions with people who already speak the audience’s language, and teach them enough incident-speak to do the translation.
A customer support lead who has spent years talking to customers can translate “Kafka consumer lag on the ingest pipeline” into “some recently submitted data may not be appearing yet” far more effectively than an engineer who has never staffed a support queue. An executive liaison who understands how the C-suite thinks about risk can distill a technical sitrep into the three sentences the CEO actually needs, without the incident commander having to figure out what those three sentences are.
Public safety professionals figured this out decades ago. In the Incident Command System used by fire departments and other emergency services, the Public Information Officer (PIO) is a defined incident role with specific training. The principle is straightforward: the people fighting the fire are busy with that, and you want someone else, with special training and experience, speaking to the reporters and the public. The skills are different, the priorities are different, and mixed messaging is dangerous.
# What this looks like in practice
The software world’s equivalent to the PIO is an incident role of Communications Lead, or a set of liaison roles, each serving a specific audience.
The simplest version: when a significant incident is declared, someone from your customer-facing team joins the response as a communications liaison. They follow the technical discussion, work from the incident commander’s sitreps, and translate updates into customer-appropriate language for the status page and support team. The IC reviews external communications for accuracy, but the drafting and posting are done by someone who knows how to talk to customers.
This has a second benefit that’s just as important: it frees the incident commander to focus on the response itself. The IC’s attention is one of the scarcest resources during an incident. Every minute they spend crafting a status page update is a minute they’re not spending on coordination, decision-making, or thinking about what they might be missing. Delegating communication beyond the response to someone better suited for the work is better for the IC, better for the response, and better for the customers.
The same principle extends to every other audience. A support update channel where a customer care liaison posts translated, agent-ready talking points. A senior management update path where an executive liaison posts business-impact summaries. A sales notification that flags which key accounts are affected, so reps know before their next customer call. A broad notification to legal, finance, and HR when a significant incident is declared, so each function can self-select whether to engage based on their own expertise, rather than waiting for the IC to figure out whether this incident has regulatory implications.
# Three things customers actually want to know
Once you have the right people doing the translating, what they actually need to communicate is simpler than you might expect.
In my experience, what customers truly want to know during an incident comes down to three things: that you’re aware of the problem, that you’re working on it, and, if possible, when it will be fixed. If you cover those three, most customers will be satisfied to let you work the problem in peace.
That’s it. They don’t need the technical details. They don’t need the internal coordination. They don’t need the investigation narrative. All of that is for your team, not for them.
# The worst failure isn’t saying the wrong thing
The most damaging communication failure during an incident isn’t saying the wrong thing, it’s saying nothing at all.
A status page that reads “All Systems Operational” while your customers are seeing errors. A support team that has no information to share. Executives finding out about the outage from social media. Each of these erodes trust in a way that even the most jargon-filled status update doesn’t, because silence communicates something very clearly: either you don’t know there’s a problem, or you know and don’t care enough to say anything.
Even an imperfect update is better than no update. Share what you know, be honest about what you don’t, and commit to a cadence for further updates. But calibrate the cadence to the pace of the incident; repetitive content-free “we’re still working on it” updates become noise rather than reassurance. (This is another reason to have skilled communicators handling the job. When I was leading incident management at Slack, the Customer Experience team seemed to have at least a dozen ways of saying “still working on it” without repeating themselves.)
Customers and stakeholders can forgive a rough update. They have a much harder time forgiving radio silence.
# This isn’t a problem Engineering can solve alone
Here’s where this gets real: incidents don’t wait for business hours. If your company has 24/7 on-call coverage for engineering, but your customer support and communications teams work 9-to-5, then at 2 AM on a Saturday your engineer is right back to writing status page updates, because there’s nobody else available to do it.
To solve this, the conversation needs to shift from Engineering’s process to the company’s commitment. Staffing incident communication properly means that Support, Customer Success, or whoever owns customer-facing communication needs their own 24/7 on-call rotation, or at least an escalation path that works at 2 AM, even if that means the Director of Customer Care is effectively always on call for high-severity incidents.
It also means those teams need to be included in incident response training (though not necessarily the full responder training that engineering gets; a focused one-hour session on their role during incidents may be all they need). And it means there needs to be a staffing conversation, and probably a budget conversation, with a VP outside of Engineering who hasn’t thought of incident response as their problem before.
That’s uncomfortable, but it’s also clarifying. If your company truly believes that customer communication during incidents matters, then the teams who are best at customer communication need to be available when incidents happen. If those teams aren’t willing to staff for it, that tells you something about how seriously the company takes it, which is useful information in itself.
For the engineering leader reading this: you shouldn’t try to fix this alone (and you probably can’t anyway). What you can do is make the case. The argument is straightforward: Engineering has invested in 24/7 incident response capability because incidents don’t respect business hours. Customer communication during those incidents is at least as important as the technical response, and it requires a different skill set. The same logic that justifies an engineering on-call rotation justifies a communication on-call rotation.
If the company isn’t willing to make that investment, then it’s making a conscious choice to accept poor customer communication during off-hours incidents, and everyone should be honest about that tradeoff.
# The real question is organizational design
The question to ask isn’t “how do we train our engineers to write better status page updates?” The question is “who in our company already has the skills to communicate with each of these audiences, and how do we bring them into our incident processes?”
That’s an organizational design question, not a training question. And it’s the kind of question that separates companies that handle incidents well from companies that just handle the technical parts well.
Remember that European customer care team from the opening? They were the ones who noticed the problem. They were the ones fielding the complaints. They already knew how to talk to the affected customers. All they needed was a seat at the table, so they could translate what the engineering team was seeing. That’s not a big ask. But it’s one that most companies have never thought to make.
*I’m writing a book about incident management for software organizations, drawing on my experience building and leading incident management programs at companies like Google, Slack, and others. If you’d like to hear about it when it’s available, visit [im4ds.com](https://im4ds.com).*
*If your organization is wrestling with incident communication or other incident management challenges, I do consulting and training on these topics. You can find out more at [greatcircle.com](https://greatcircle.com).*
## Recent Comments