SRE Weekly Issue #528

A message from our sponsor, Planetscale:

Most database incidents start with one expensive query, not the database being down. PlanetScale gives SRE teams high-availability Postgres and MySQL with automated failover, query insights, and Database Traffic Control to stop runaway queries before they page you.

→ Explore PlanetScale

Spotify has had some difficulty around podcast publishing, and they shared this analysis of the worst incident.

  Jim Whitehead, Ulrik Mikaelsson, John Lagomarsino, and Saunak Jai Chakrabarti — Spotify

…and here’s where it gets interesting. This post shares the user point of view on the Spotify issues, including fact-checking their published timeline.

  Gergely Orosz — The Pragmatic Engineer

New incident role unlocked: the incident tech lead. I enjoyed the description of the interplay between the tech lead and the incident commander.

  Brent Chapman

This one goes hard: if you try to reduce your incident count, your system will become less reliable, not more. Aim for more incidents, handled well.

  Tim Irving

There’s some brutal honesty in here that I find refreshing, especially around the impact on incidents and incident response.

  Liz Fong-Jones — Honeycomb

Here’s what I’ve learned about what actually happens inside a team when something breaks, how teams often reflect, and how we think about accountability without losing a blameless culture.

  Karan Nagarajowda — Uptime Labs

agent sprawl is now a production reliability crisis, and the SRE discipline does not yet have governance frameworks for it.

   Ajay Devineni — DZone

The premise: read replicas can help you scale read load, but they introduce complexity. The article goes into the problems they ran into and how they dealt with them.

  Johanna Larsson — incident.io