Files
nexus/sreweekly/markdown/519/01-the-problem-with-ai-generated-post-incident-reviews.md
2026-09-12 17:23:01 +08:00

9.4 KiB
Raw Blame History

The Problem with AI-Generated Post-Incident Reviews

简介

They give solid examples to argue that much of the learning happens during the process of writing a post-incident review.

[…] you could throw the post-incident review document away after writing it and still get the vast majority of the value out of the process.

正文

Modern AI tools can produce a competent-looking post-incident review document from a Slack channel transcript and a few prompts. The output will be pleasingly formatted, with a timeline, a list of contributing factors, and a set of action items. It will read coherently, and it will arrive faster than a human-written review would have, with less engineer time spent producing it. For a manager looking at the post-incident review process and seeing engineers grumble about the time it takes, this is tempting.

The catch is that the document was never the point of the review. The real learning comes from analyzing the incident while writing the document, not reading it; the document at the end is the residue of the learning. It’s like studying; you learn a lot more from working the problem sets than you do from just reading a classmate’s summary.

Three layers of learning

The learning happens at three layers: readers of the published document, the writers individually, and the writers as a group.

Readership is the most visible layer, and radiates outward when the document is published. Colleagues throughout the company read the review, see how it says the system behaved and how the team supposedly handled the incident, and perhaps update their own understanding of what’s possible. Most of them weren’t in the incident, so for them, the document is the incident. Their learning is downstream of the writers’ analysis, and weak analysis produces shallow lessons; readers get less than they could have, and some of what they get may be flat-out wrong.

Each writer ends up with a more thorough understanding of the incident than they started with. They start to write “the deploy caused the outage” and realize, as they trace the sequence, that the deploy only surfaced a problem that was already lying in wait. They find themselves describing the dashboard as “down” and stop, because the dashboard wasn’t actually down; it was up, but showing data from the wrong cluster, which is why nothing made sense to the responders for the first eighteen minutes. They write “the team decided to…” and stop, because the team didn’t decide; one person made a call and the others went along, and the gap between those two things turns out to matter.

The writers also learn from each other. Responders fill in their slices of the timeline; the architect annotates the contributing factors; customer success writes the impact section. As they go, someone catches a gap, someone corrects a misremembered moment, a disagreement surfaces in the comments and leads to an enlightening discussion. By the time the writing is done, the group has a grasp of what happened that no individual writer had alone. They built it from reconciling what each of them separately knew.

What AI does to each layer

When AI writes the document, each of these layers fares differently, and none of them fares well.

Readers are still readers. They open the document, take it in, and maybe update their understanding based on what they read. But what they’re absorbing now is the AI’s synthesis, with no human pressure-testing behind it. It may or may not be right, or get to the deeper issues that a group of writers might have uncovered. The document looks like a review, and readers absorb it like a real review. Whether they’re learning anything true or useful is now a function of how well the AI happened to do, with no way to tell from the outside.

There are no writers when the AI does the writing, so there’s no individual learning from the writing process. No one stops mid-sentence to discover that the deploy only surfaced a problem already lying in wait; no one finds that the dashboard was up but showing wrong-cluster data; no one writes “the team decided to” and stops to reconsider. The kind of learning that comes from working through the evidence sentence by sentence doesn’t happen when no one is doing it.

And without writers, you clearly can’t have “writers learning from each other.” The AI conjures up a plausible-sounding narrative from whatever it was given. No one catches a gap; no one corrects a misremembered moment; no disagreement surfaces in the comments to lead to an enlightening discussion. Any tension between what different people actually thought is gone before anyone could surface it.

So you end up with a polished artifact. Your contributing factors are the AI’s guess at what’s plausible, not your team’s hard-won understanding. Your timeline is a transcript reorganization, not a reconstruction. Your “lessons learned” come from the AI pattern-matching against incidents in its training data, not against the ones your company actually had. The document looks like a review, but it’s a fantasy.

You can automate production of the review document. You can’t automate the understanding that the process was supposed to produce. Automating the writing away automates the learning away.

Where AI legitimately helps

This isn’t a case for keeping AI out of the post-incident review process. Rather, it’s a case for being clear about where AI helps the engineers do the work and where it does the work for them. Mechanical support is genuinely useful. Substitution for human thinking is not.

Specific places AI earns its keep:

  • Collating raw material. Pulling content from multiple Slack channels and threads into a unified view, transcribing voice channels (either real-time or recorded), gathering the screenshots and graphs that got posted in-channel during the incident. The grunt work that gives writers a clean starting point.
  • Format and copy editing on a draft a human has written. Tightening prose, suggesting clearer phrasings, catching the inconsistencies that survive a human edit pass. AI is good at this, and using it here saves time without taking anything away from the writer’s engagement with the material.
  • Surfacing gaps. “Your timeline jumps from 14:46 to 15:12 with no entries. Was anything happening then?” That’s a useful prompt for human writers, not a substitute for them.
  • Cross-referencing past incidents. “Three other incidents in the last six months touched the same service; here are the links.” Pattern-matching across a library of past reviews is exactly the kind of mechanical work AI is well-suited to. (This one requires that your library of past reviews actually exists and is structured in a way the tool can search, which is a separate problem worth solving on its own merits.)

AI handles mechanical work that supports the writers. The writers do the thinking. The moment you let the tool do the thinking (generating the contributing factors, drafting the lessons learned, writing the narrative itself), you’ve automated away most of the learning, and arguably the most valuable parts.

The right question for the manager

If you’re an engineering leader evaluating an AI tool that promises to help with post-incident reviews, the question isn’t “does this tool produce a good-looking document faster?” All of them do that, but appearance alone shouldn’t be your criterion. A much better question is “does this tool support my engineers in doing the writing, or does it replace them in doing it?”

Tools in the first category save real time on the parts of the work that don’t produce learning. Tools in the second category save your engineers from the activity that the post-incident review exists for. They give you a polished artifact and an empty experience. The same incidents keep happening, but hey, now you have a faster pipeline for marking off the “Write Post-Incident Review” checkbox.

Here’s a clarifying way to think about it: you could throw the post-incident review document away after writing it and still get the vast majority of the value out of the process. The document is like the scribblings on a whiteboard after a productive working session: interesting, yes, and maybe worth snapping a photo of, but the real value is what leaves the room in the heads of the people who were there. You don’t actually want to throw it away (the document does real work, both immediately and over time, as part of the library of past reviews), but knowing that you could is the right reference point for thinking about which tools genuinely help the process and which ones quietly hollow it out.

AI can produce a good-looking incident review document, but only your engineers can produce the understanding behind a truly good one. Adopt tools that support that distinction; the ones that don’t will leave you with a stack of polished artifacts, but without much actual learning.

Post-incident reviews are one of several topics covered in my forthcoming book, “Incident Management for DevOps and SRE.” Learn more and sign up for updates at im4ds.com.

Need help preventing, preparing for, responding to, and learning from incidents? That’s the focus of my consulting practice at Great Circle.

Recent Comments