Files
nexus/sreweekly/markdown/108/03-production-test-run-the-self-flagellating-server.md
2026-09-12 17:23:01 +08:00

52 lines
3.5 KiB
Markdown
Raw Blame History

This file contains invisible Unicode characters
This file contains invisible Unicode characters that are indistinguishable to humans but may be processed differently by a computer. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Production Test Run The self flagellating server
- **期号**: SRE Weekly Issue #108(2018-02-04)
- **作者**: —
- **链接**: https://ayende.com/blog/181441-C/production-test-run-the-self-flagellating-server
## 简介
Here’s a fun little distributed system debugging story from the founder of RavenDB.
## 正文
## Production Test RunThe self flagellating server
**2 min**|
**354 words**
Sometimes you see the impossible. In one of our scenarios, we saw a cluster that had such a bad case of split brain that it came near to fracturing the very boundaries of space & time. ![image image](https://ayende.com/blog/Images/Open-Live-Writer/Production-Test-Run-The-self-flagellatin_11019/image_thumb.png)
In a three node cluster, we have one node that looked to be fine. It connected to all the other nodes and was the cluster leader. The other two nodes, however, were *not* in the cluster and in fact, they were showing signs that they never *were* in the cluster.
What was *really* strange was that we took the other two machines down and the first node was still showing a successful cluster. We looked deeper and realized that it wasn’t actually a healthy situation, in fact, this node was very rapidly switching between leader and follower mode.
It took a bit of time to figure out what was going on, but the root cause was DNS. We had the three nodes on separate DNS (a.oren.development.run, b.oren.development.run, c.oren.development.run) and they were setup to point to the three machines. However, we have previously used the same domain names to run a cluster on the first machine only. Because of the way DNS updates, whenever the machine at a.oren.development.run would try to connect to b.oren.development.run it would actually connect to *itself*.
At this point, A would tell B that it is the leader. But A *is* B, so A would respond by becoming a follower (because it was told it should, by itself). Because it became a follower, it disconnected from itself. After a timeout, it would become leader again, and the cycle would continue.
Every time that the server would get up, it would whip itself down again. “I’m a leader”, “No, I’m a leader”, etc.
This is a fun thing to discover. We had to trace pretty deep to figure out that the problem was in the DNS cache (since the DNS itself was properly updated).
We fixed things so we now recognize if we are talking to ourselves and error properly.
####
More posts in
["Production Test Run"](https://ayende.com/blog/posts/series/181345-c/production-test-run)
series:
1. *(26 Jan 2018)*[Overburdened and under provisioned](https://ayende.com/blog/181537-b/production-test-run-overburdened-and-under-provisioned)
2. ***(24 Jan 2018)* [The self flagellating server](https://ayende.com/blog/181441-c/production-test-run-the-self-flagellating-server)**
3. *(22 Jan 2018)*[Rowhammer in Voron](https://ayende.com/blog/181409-a/production-test-run-rowhammer-in-voron)
4. *(18 Jan 2018)*[When your software is configured by a monkey](https://ayende.com/blog/181347-c/production-test-run-when-your-software-is-configured-by-a-monkey)
5. *(17 Jan 2018)*[Too much of a good thing isn’t so good for you](https://ayende.com/blog/181346-c/production-test-run-too-much-of-a-good-thing-isn-t-so-good-for-you)
6. *(16 Jan 2018)*[The worst is yet to come](https://ayende.com/blog/181345-c/production-test-run-the-worst-is-yet-to-come)
## Comments
## Comment preview