Files
nexus/sreweekly/markdown/536/07-rollback-does-not-erase-distributed-memory.md

96 lines
7.5 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Rollback Does Not Erase Distributed Memory
- **期号**: SRE Weekly Issue #536(2026-09-28)
- **作者**: Balu Kambala
- **链接**: https://storiesfromtheedge.substack.com/p/rollback-does-not-erase-distributed
## 简介
Rollback sounds great in theory, but it doesn’t always work. The CircleCI example really hits hard.
This reminds me of a classic article from the now-defunct company Skyliner, You Can’t Have a Rollback Button.
## 正文
The rollback finished successfully. The deployment dashboard showed 100% of nodes running the previous build. The service confirmed the old configuration was active.
Yet customers were still seeing failures.
In a distributed system, rolling back code can create a false sense of safety. Reverting the change does not revert every component that has already observed the change. Recursive DNS resolvers and other caches may retain the new state until their TTLs expire. A database may contain records written in a format the previous application cannot read. Queues may still hold work created under the failed version.
Your servers are running version A, but they are running inside an environment mutated by version B.*A rollback changes what the source of truth says now. However, it does not erase what the rest of the system has already learned.*
This is why distributed recovery is rarely a single event. Some customers recover immediately. Others recover when caches expire, connections restart, queues drain, or incompatible data is repaired.
So after a rollback, the important question is not only "Is the old version running?" but also "Which parts of the system still retain state from the new version?"
**The Takeaway:** Deploying Version 40 over Version 41 changes the code, but not the environment Version 41 mutated. Recovery begins only when the distributed memory such as DNS caches, client pools, and databases is cleared or reconciled.
Slack’s DNSSEC rollout demonstrates what happens when that memory lives in infrastructure you do not own.
## The Slack DNSSEC rollback
During its [DNSSEC rollout](https://slack.engineering/what-happened-during-slacks-dnssec-rollout) for [slack.com](http://slack.com) Slack observed some DNS resolvers returning unexpected NODATA responses. The team decided to roll back.
DNSSEC depends on records stored in two places. The ‘.com’ parent zone publishes a DS record telling resolvers that `slack.com` must provide cryptographic proof. Slack’s authoritative zone publishes the matching DNSKEY and RRSIG records.
Slack removed the DS record from ‘.com’ through its registrar and later disabled signing on its authoritative servers.
The authoritative configuration looked old again. The resolver caches did not.
Some recursive resolvers had cached the DS record for up to 24 hours. They still expected signed answers from Slack. When the signing material disappeared, validation failed and those resolvers returned SERVFAIL.
**The rollback made the system worse.** Slack had changed the authoritative state, but it could not force resolver caches to forget the DS record.
Slack asked major public DNS providers to flush their caches. Large providers recovered quickly. Other resolvers continued failing until their cached DS records expired.
As shown below, the rollback completed quickly. Recovery followed a long tail as resolver caches expired on different schedules.
![](https://substackcdn.com/image/fetch/$s_!gHJA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F123af348-cb0e-435b-8896-8403d4e1dc1a_2016x864.png)
Slack’s retained state would eventually disappear as resolver caches expired. Database state is harder: records written by the new version may remain indefinitely.
## **The database does not roll back with the code**
CircleCI encountered the [incident](https://discuss.circleci.com/t/incident-report-november-8-2021-jobs-stuck-in-a-not-running-state/41890) in November 2021. A deployment changed the type of a field in its job-distribution database. The new version could not process older jobs, so CircleCI rolled it back. The old version then could not process jobs written during the deployment.
The database now contained two representations, but neither application version understood both.
Engineers recovered by deploying a compatibility change that ignored the field, then draining the accumulated queues.
The sequence below shows why neither application version was a safe recovery point. Version B rejected older records, version A rejected records written by version B, and recovery required code that could handle both formats.
![](https://substackcdn.com/image/fetch/$s_!8N05!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21c2703a-5490-4d00-897e-feda88a0fdd7_2304x896.png)
CircleCI had rolled back the code, but recovery required a compatibility build. Once a new version has written a new representation, the old binary may no longer be a safe recovery target.
## **Rollback is not recovery**
A dashboard can truthfully report that version 41 was replaced by version 40. That does not mean the system is behaving as version 40 did before the deployment.
During an incident, I separate three stages:
1. **The source was rolled back.** The central configuration returned to the old version.
2. **Components applied the rollback.** Systems using that configuration fetched and applied the old version.
3. **The system recovered.** Data and actions created by the new version were made compatible, reversed, or repaired.
Teams often prove the first stage and assume the third. Slack and CircleCI both completed the first stage before reaching the third. Sometimes the safest way to close that gap is to move forward.
## **Why roll-forward can be safer**
These incidents show that rollback is not always enough for recovery. Once the new version writes incompatible data or publishes state that other systems cache, returning to the old version can create another failure.
CircleCI recovered by deploying a compatibility build that ignored the changed field. This allowed it to process jobs written by both versions. For Slack, a safer recovery would have kept the signing material available until cached DS records expired.
The recovery decision should depend on the state already present in the system, not on how easy it is to run the rollback command. The delay between rollback and actual recovery often appears as a long tail in the error graph.
## **Reading the recovery curve**
Slack and CircleCI produced different recovery curves. Slack showed a sharp initial recovery followed by a long tail. CircleCI showed little change after rollback; recovery began only after a compatibility fix and continued as the queue drained.
The curve cannot prove the cause, but it narrows the search. This is what I mean by incident geometry. A sharp recovery may point to a centralized fix. A long tail may indicate components recovering on independent clocks. Little or no improvement after rollback may indicate durable state that requires repair or reconciliation.
I’ll go deeper into incident geometry—how user-impact shapes can lead us to architectural findings—in my talk, [The Geometry of Incidents: What User-Impact Shapes Reveal About Platform Architecture](https://www.wearedevelopers.com/events/world-congress-2026-north-america/sessions/1722-the-geometry-of), at WeAreDevelopers World Congress North America. If you’re attending, please stop by. I’d be glad to continue the discussion in person.