Files
nexus/sreweekly/articles/104/02-building-a-distributed-log-from-scratch-part-2-data-replication.html
2026-09-12 17:23:01 +08:00

28 lines
61 KiB
HTML
Raw Blame History

This file contains invisible Unicode characters
This file contains invisible Unicode characters that are indistinguishable to humans but may be processed differently by a computer. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
<!doctype html><html lang=en dir=auto data-theme=auto><head><meta charset=utf-8><meta http-equiv=X-UA-Compatible content="IE=edge"><meta name=viewport content="width=device-width,initial-scale=1,shrink-to-fit=no"><meta name=robots content="index, follow"><title>Building a Distributed Log from Scratch, Part 2: Data Replication | Brave New Geek</title><meta name=keywords content="algorithms,architecture,building a distributed log from scratch,cap theorem,consensus,consistency,data replication,distributed log,distributed systems,kafka,leader election,message queues,message-oriented middleware,messaging,nats,nats streaming,performance,raft,stream processing,write-ahead log,zookeeper"><meta name=description content="In part one of this series we introduced the idea of a message log, touched on why it’s useful, and discussed the storage mechanics behind it. In part two, we discuss data replication.
We have our log. We know how to write data to it and read it back as well as how data is persisted. The caveat to this is, although we have a durable log, it’s a single point of failure (SPOF). If the machine where the log data is stored dies, we’re SOL. Recall that one of our three priorities with this system is high availability, so the question is how do we achieve high availability and fault tolerance?"><meta name=author content><link rel=canonical href=https://bravenewgeek.com/building-a-distributed-log-from-scratch-part-2-data-replication/><link crossorigin=anonymous href=/assets/css/stylesheet.4861a452a4c13a9a1fbf2085400b74a7de96b1beeb94dee57b4273e5dffcf337.css integrity="sha256-SGGkUqTBOpofvyCFQAt0p96Wsb7rlN7le0Jz5d/88zc=" rel="preload stylesheet" as=style><link rel=icon href=https://bravenewgeek.com/favicon.ico><link rel=icon type=image/png sizes=16x16 href=https://bravenewgeek.com/favicon.ico><link rel=icon type=image/png sizes=32x32 href=https://bravenewgeek.com/favicon.ico><link rel=apple-touch-icon href=https://bravenewgeek.com/favicon.ico><link rel=mask-icon href=https://bravenewgeek.com/favicon.ico><meta name=theme-color content="#2e2e33"><meta name=msapplication-TileColor content="#2e2e33"><link rel=alternate hreflang=en href=https://bravenewgeek.com/building-a-distributed-log-from-scratch-part-2-data-replication/><noscript><style>#theme-toggle,.top-link{display:none}</style><style>@media(prefers-color-scheme:dark){:root{--theme:rgb(29, 30, 32);--entry:rgb(46, 46, 51);--primary:rgb(218, 218, 219);--secondary:rgb(155, 156, 157);--tertiary:rgb(65, 66, 68);--content:rgb(196, 196, 197);--code-block-bg:rgb(46, 46, 51);--code-bg:rgb(55, 56, 62);--border:rgb(51, 51, 51);color-scheme:dark}.list{background:var(--theme)}.toc{background:var(--entry)}}</style></noscript><script>localStorage.getItem("pref-theme")==="dark"?document.querySelector("html").dataset.theme="dark":localStorage.getItem("pref-theme")==="light"?document.querySelector("html").dataset.theme="light":window.matchMedia("(prefers-color-scheme: dark)").matches?document.querySelector("html").dataset.theme="dark":document.querySelector("html").dataset.theme="light"</script><link rel=preconnect href=https://fonts.googleapis.com><link rel=preconnect href=https://fonts.gstatic.com crossorigin><link rel=stylesheet href="https://fonts.googleapis.com/css2?family=JetBrains+Mono:wght@400;500&family=Source+Serif+4:ital,opsz,wght@0,8..60,400;0,8..60,600;1,8..60,400&family=Space+Grotesk:wght@500;600;700&display=swap"><meta property="og:url" content="https://bravenewgeek.com/building-a-distributed-log-from-scratch-part-2-data-replication/"><meta property="og:site_name" content="Brave New Geek"><meta property="og:title" content="Building a Distributed Log from Scratch, Part 2: Data Replication"><meta property="og:description" content="In part one of this series we introduced the idea of a message log, touched on why it’s useful, and discussed the storage mechanics behind it. In part two, we discuss data replication.
We have our log. We know how to write data to it and read it back as well as how data is persisted. The caveat to this is, although we have a durable log, it’s a single point of failure (SPOF). If the machine where the log data is stored dies, we’re SOL. Recall that one of our three priorities with this system is high availability, so the question is how do we achieve high availability and fault tolerance?"><meta property="og:locale" content="en_us"><meta property="og:type" content="article"><meta property="article:section" content="posts"><meta property="article:published_time" content="2017-12-27T12:26:55-06:00"><meta property="article:modified_time" content="2018-02-23T16:09:46-06:00"><meta property="article:tag" content="algorithms"><meta property="article:tag" content="Architecture"><meta property="article:tag" content="Building a Distributed Log From Scratch"><meta property="article:tag" content="Cap Theorem"><meta property="article:tag" content="Consensus"><meta property="article:tag" content="Consistency"><meta name=twitter:card content="summary"><meta name=twitter:title content="Building a Distributed Log from Scratch, Part 2: Data Replication"><meta name=twitter:description content="In part one of this series we introduced the idea of a message log, touched on why it’s useful, and discussed the storage mechanics behind it. In part two, we discuss data replication.
We have our log. We know how to write data to it and read it back as well as how data is persisted. The caveat to this is, although we have a durable log, it’s a single point of failure (SPOF). If the machine where the log data is stored dies, we’re SOL. Recall that one of our three priorities with this system is high availability, so the question is how do we achieve high availability and fault tolerance?"><script type=application/ld+json>{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Posts","item":"https://bravenewgeek.com/posts/"},{"@type":"ListItem","position":2,"name":"Building a Distributed Log from Scratch, Part 2: Data Replication","item":"https://bravenewgeek.com/building-a-distributed-log-from-scratch-part-2-data-replication/"}]}</script><script type=application/ld+json>{"@context":"https://schema.org","@type":"BlogPosting","headline":"Building a Distributed Log from Scratch, Part 2: Data Replication","name":"Building a Distributed Log from Scratch, Part 2: Data Replication","description":"In part one of this series we introduced the idea of a message log, touched on why it’s useful, and discussed the storage mechanics behind it. In part two, we discuss data replication.\nWe have our log. We know how to write data to it and read it back as well as how data is persisted. The caveat to this is, although we have a durable log, it’s a single point of failure (SPOF). If the machine where the log data is stored dies, we’re SOL. Recall that one of our three priorities with this system is high availability, so the question is how do we achieve high availability and fault tolerance?\n","keywords":["algorithms","architecture","building a distributed log from scratch","cap theorem","consensus","consistency","data replication","distributed log","distributed systems","kafka","leader election","message queues","message-oriented middleware","messaging","nats","nats streaming","performance","raft","stream processing","write-ahead log","zookeeper"],"articleBody":"In part one of this series we introduced the idea of a message log, touched on why it’s useful, and discussed the storage mechanics behind it. In part two, we discuss data replication.\nWe have our log. We know how to write data to it and read it back as well as how data is persisted. The caveat to this is, although we have a durable log, it’s a single point of failure (SPOF). If the machine where the log data is stored dies, we’re SOL. Recall that one of our three priorities with this system is high availability, so the question is how do we achieve high availability and fault tolerance?\nWith high availability, we’re specifically talking about ensuring continuity of reads and writes. A server failing shouldn’t preclude either of these, or at least unavailability should be kept to an absolute minimum and without the need for operator intervention. Ensuring this continuity should be fairly obvious: we eliminate the SPOF. To do that, we replicate the data. Replication can also be a means for increasing scalability, but for now we’re only looking at this through the lens of high availability.\nThere are a number of ways we can go about replicating the log data. Broadly speaking, we can group the techniques into two different categories: gossip/multicast protocols and consensus protocols. The former includes things like epidemic broadcast trees, bimodal multicast, SWIM, HyParView, and NeEM. These tend to be eventually consistent and/or stochastic. The latter, which I’ve described in more detail here, includes 2PC/3PC, Paxos, Raft, Zab, and chain replication. These tend to favor strong consistency over availability.\nSo there are a lot of ways we can replicate data, but some of these solutions are better suited than others to this particular problem. Since ordering is an important property of a log, consistency becomes important for a replicated log. If we read from one replica and then read from another, it’s important those views of the log don’t conflict with each other. This more or less rules out the stochastic and eventually consistent options, leaving us with consensus-based replication.\nThere are essentially two components to consensus-based replication schemes: 1) designate a leader who is responsible for sequencing writes and 2) replicate the writes to the rest of the cluster.\nDesignating a leader can be as simple as a configuration setting, but the purpose of replication is fault tolerance. If our configured leader crashes, we’re no longer able to accept writes. This means we need the leader to be dynamic. It turns out leader election is a well-understood problem, so we’ll get to this in a bit.\nOnce a leader is established, it needs to replicate the data to followers. In general, this can be done by either waiting for all replicas or waiting for only a quorum (majority) of replicas. There are pros and cons to both approaches.\nPros\nCons\nAll Replicas\nTolerates f failures with f+1 replicas\nLatency pegged to slowest replica\nQuorum\nHides delay from a slow replica\nTolerates f failures with 2f+1 replicas\nWaiting on all replicas means we can make progress as long as at least one replica is available. With quorum, tolerating the same amount of failures requires more replicas because we need a majority to make progress. The trade-off is that the quorum hides any delays from a slow replica. Kafka is an example of a system which uses all replicas (with some conditions on this which we will see later), and NATS Streaming is one that uses a quorum. Let’s take a look at both in more detail.\nReplication in Kafka In Kafka, a leader is selected (we’ll touch on this in a moment). This leader maintains an in-sync replica set (ISR) consisting of all the replicas which are fully caught up with the leader. This is every replica, by definition, at the beginning. All reads and writes go through the leader. The leader writes messages to a write-ahead log (WAL). Messages written to the WAL are considered uncommitted or “dirty” initially. The leader only commits a message once all replicas in the ISR have written it to their own WAL. The leader also maintains a high-water mark (HW) which is the last committed message in the WAL. This gets piggybacked on the replica fetch responses from which replicas periodically checkpoint to disk for recovery purposes. The piggybacked HW then allows replicas to know when to commit.\nOnly committed messages are exposed to consumers. However, producers can configure how they want to receive acknowledgements on writes. It can wait until the message is committed on the leader (and thus replicated to the ISR), wait for the message to only be written (but not committed) to the leader’s WAL, or not wait at all. This all depends on what trade-offs the producer wants to make between latency and durability.\nThe graphic below shows how this replication process works for a cluster of three brokers: b1, b2, and b3. Followers are effectively special consumers of the leader’s log.\nNow let’s look at a few failure modes and how Kafka handles them.\nLeader Fails Kafka relies on Apache ZooKeeper for certain cluster coordination tasks, such as leader election, though this is not actually how the log leader is elected. A Kafka cluster has a single controller broker whose election is handled by ZooKeeper. This controller is responsible for performing administrative tasks on the cluster. One of these tasks is selecting a new log leader (actually partition leader, but this will be described later in the series) from the ISR when the current leader dies. ZooKeeper is also used to detect these broker failures and signal them to the controller.\nThus, when the leader crashes, the cluster controller is notified by ZooKeeper and it selects a new leader from the ISR and announces this to the followers. This gives us automatic failover of the leader. All committed messages up to the HW are preserved and uncommitted messages may be lost during the failover. In this case, b1 fails and b2 steps up as leader.\nFollower Fails The leader tracks information on how “caught up” each replica is. Before Kafka 0.9, this included both how many messages a replica was behind, replica.lag.max.messages, and the amount of time since the replica last fetched messages from the leader, replica.lag.time.max.ms. Since 0.9, replica.lag.max.messages was removed and replica.lag.time.max.ms now refers to both the time since the last fetch request and the amount of time since the replica last caught up.\nThus, when a follower fails (or stops fetching messages for whatever reason), the leader will detect this based on replica.lag.time.max.ms. After that time expires, the leader will consider the replica out of sync and remove it from the ISR. In this scenario, the cluster enters an “under-replicated” state since the ISR has shrunk. Specifically, b2 fails and is removed from the ISR.\nFollower Temporarily Partitioned The case of a follower being temporarily partitioned, e.g. due to a transient network failure, is handled in a similar fashion to the follower itself failing. These two failure modes can really be combined since the latter is just the former with an arbitrarily long partition, i.e. it’s the difference between crash-stop and crash-recovery models.\nIn this case, b3 is partitioned from the leader. As before, replica.lag.time.max.ms acts as our failure detector and causes b3 to be removed from the ISR. We enter an under-replicated state and the remaining two brokers continue committing messages 4 and 5. Accordingly, the HW is updated to 5 on these brokers.\nWhen the partition heals, b3 continues reading from the leader and catching up. Once it is fully caught up with the leader, it’s added back into the ISR and the cluster resumes its fully replicated state.\nWe can generalize this to the crash-recovery model. For example, instead of a network partition, the follower could crash and be restarted later. When the failed replica is restarted, it recovers the HW from disk and truncates its log up to the HW. This preserves the invariant that messages after the HW are not guaranteed to be committed. At this point, it can begin catching up from the leader and will end up with a log consistent with the leader’s once fully caught up.\nReplication in NATS Streaming NATS Streaming relies on the Raft consensus algorithm for leader election and data replication. This sometimes comes as a surprise to some as Raft is largely seen as a protocol for replicated state machines. We’ll try to understand why Raft was chosen for this particular problem in the following sections. We won’t dive deep into Raft itself beyond what is needed for the purposes of this discussion.\nWhile a log is a state machine, it’s a very simple one: a series of appends. Raft is frequently used as the replication mechanism for key-value stores which have a clearer notion of “state machine.” For example, with a key-value store, we have set and delete operations. If we set foo = bar and then later set foo = baz, the state gets rolled up. That is, we don’t necessarily care about the provenance of the key, only its current state.\nHowever, NATS Streaming differs from Kafka in a number of key ways. One of these differences is that NATS Streaming attempts to provide a sort of unified API for streaming and queueing semantics not too dissimilar from Apache Pulsar. This means, while it has a notion of a log, it also has subscriptions on that log. Unlike Kafka, NATS Streaming tracks these subscriptions and metadata associated with them, such as where a client is in the log. These have definite “state machines” affiliated with them, like creating and deleting subscriptions, positions in the log, clients joining or leaving queue groups, and message-redelivery information.\nCurrently, NATS Streaming uses multiple Raft groups for replication. There is a single metadata Raft group used for replicating client state and there is a separate Raft group per topic which replicates messages and subscriptions.\nRaft solves both the problems of leader election and data replication in a single protocol. The Secret Lives of Data provides an excellent interactive illustration of how this works. As you step through that illustration, you’ll notice that the algorithm is actually quite similar to the Kafka replication protocol we walked through earlier. This is because although Raft is used to implement replicated state machines, it actually is a replicated WAL, which is exactly what Kafka is. One benefit of using Raft is we no longer have the need for ZooKeeper or some other coordination service.\nRaft handles electing a leader. Heartbeats are used to maintain leadership. Writes flow through the leader to the followers. The leader appends writes to its WAL and they are subsequently piggybacked onto the heartbeats which get sent to the followers using AppendEntries messages. At this point, the followers append the write to their own WALs, assuming they don’t detect a gap, and send a response back to the leader. The leader commits the write once it receives a successful response from a quorum of followers.\nSimilar to Kafka, each replica in Raft maintains a high-water mark of sorts called the commit index, which is the index of the highest log entry known to be committed. This is piggybacked on the AppendEntries messages which the followers use to know when to commit entries in their WALs. If a follower detects that it missed an entry (i.e. there was a gap in the log), it rejects the AppendEntries and informs the leader to rewind the replication. The Raft paper details how it ensures correctness, even in the face of many failure modes such as the ones described earlier.\nConceptually, there are two logs: the Raft log and the NATS Streaming message log. The Raft log handles replicating messages and, once committed, they are appended to the NATS Streaming log. If it seems like there’s some redundancy here, that’s because there is, which we’ll get to soon. However, keep in mind we’re not just replicating the message log, but also the state machines associated with the log and any clients.\nThere are a few challenges with this replication technique, two of which we will talk about. The first is scaling Raft. With a single topic, there is one Raft group, which means one node is elected leader and it heartbeats messages to followers.\nAs the number of topics increases, so do the number of Raft groups, each with their own leaders and heartbeats. Unless we constrain the Raft group participants or the number of topics, this creates an explosion of network traffic between nodes.\nThere are a couple ways we can go about addressing this. One option is to run a fixed number of Raft groups and use a consistent hash to map a topic to a group. This can work well if we know roughly the number of topics beforehand since we can size the number of Raft groups accordingly. If you expect only 10 topics, running 10 Raft groups is probably reasonable. But if you expect 10,000 topics, you probably don’t want 10,000 Raft groups. If hashing is consistent, it would be feasible to dynamically add or remove Raft groups at runtime, but it would still require repartitioning a portion of topics which can be complicated.\nAnother option is to run an entire node’s worth of topics as a single group using a layer on top of Raft. This is what CockroachDB does to scale Raft in proportion to the number of key ranges using a layer on top of Raft they call MultiRaft. This requires some cooperation from the Raft implementation, so it’s a bit more involved than the partitioning technique but eschews the repartitioning problem and redundant heartbeating.\nThe second challenge with using Raft for this problem is the issue of “dual writes.” As mentioned before, there are really two logs: the Raft log and the NATS Streaming message log, which we’ll call the “store.” When a message is published, the leader writes it to its Raft log and it goes through the Raft replication process.\nOnce the message is committed in Raft, it’s written to the NATS Streaming log and the message is now visible to consumers.\nNote, however, that not only messages are written to the Raft log. We also have subscriptions and cluster topology changes, for instance. These other items are not written to the NATS Streaming log but handled in other ways on commit. That said, messages tend to occur in much greater volume than these other entries.\nMessages end up getting stored redundantly, once in the Raft log and once in the NATS Streaming log. We can address this problem if we think about our logs a bit differently. If you recall from part one, our log storage consists of two parts: the log segment and the log index. The segment stores the actual log data, and the index stores a mapping from log offset to position in the segment.\nAlong these lines, we can think of the Raft log index as a “physical offset” and the NATS Streaming log index as a “logical offset.” Instead of maintaining two logs, we treat the Raft log as our message write-ahead log and treat the NATS Streaming log as an index into that WAL. Particularly, messages are written to the Raft log as usual. Once committed, we write an index entry for the message offset that points back into the log. As before, we use the index to do lookups into the log and can then read sequentially from the log itself.\nRemaining Questions We’ve answered the questions of how to ensure continuity of reads and writes, how to replicate data, and how to ensure replicas are consistent. The remaining two questions pertaining to replication are how do we keep things fast and how do we ensure data is durable?\nThere are several things we can do with respect to performance. The first is we can configure publisher acks depending on our application’s requirements. Specifically, we have three options. The first is the broker acks on commit. This is slow but safe as it guarantees the data is replicated. The second is the broker acks on appending to its local log. This is fast but unsafe since it doesn’t wait on any replica roundtrips but, by that very fact, means that the data is not replicated. If the leader crashes, the message could be lost. Lastly, the publisher can just not wait for an ack at all. This is the fastest but least safe option for obvious reasons. Tuning this all depends on what requirements and trade-offs make sense for your application.\nThe second thing we do is don’t explicitly fsync writes on the broker and instead rely on replication for durability. Both Kafka and NATS Streaming (when clustered) do this. With fsync enabled (in Kafka, this is configured with flush.messages and/or flush.ms and in NATS Streaming, with file_sync), every message that gets published results in a sync to disk. This ends up being very expensive. The thought here is if we are replicating to enough nodes, the replication itself is sufficient for HA of data since the likelihood of more than a quorum of nodes failing is low, especially if we are using rack-aware clustering. Note that data is still periodically flushed in the background by the kernel.\nBatching aggressively is also a key part of ensuring good performance. Kafka supports end-to-end batching from the producer all the way to the consumer. NATS Streaming does not currently support batching at the API level, but it uses aggressive batching when replicating and persisting messages. In my experience, this makes about an order-of-magnitude improvement in throughput.\nFinally, as already discussed earlier in the series, keeping disk access sequential and maximizing zero-copy reads makes a big difference as well.\nThere are a few things worth noting with respect to durability. Quorum is what guarantees durability of data. This comes “for free” with Raft due to the nature of that protocol. In Kafka, we need to do a bit of configuring to ensure this. Namely, we need to configure min.insync.replicas on the broker and acks on the producer. The former controls the minimum number of replicas that must acknowledge a write for it to be considered successful when a producer sets acks to “all.” The latter controls the number of acknowledgments the producer requires the leader to have received before considering a request complete. For example, with a topic that has a replication factor of three, min.insync.replicas needs to be set to two and acks set to “all.” This will, in effect, require a quorum of two replicas to process writes.\nAnother caveat with Kafka is unclean leader elections. That is, if all replicas become unavailable, there are two options: choose the first replica to come back to life (not necessarily in the ISR) and elect this replica as leader (which could result in data loss) or wait for a replica in the ISR to come back to life and elect it as leader (which could result in prolonged unavailability). Initially, Kafka favored availability by default by choosing the first strategy. If you preferred consistency, you needed to set unclean.leader.election.enable to false. However, as of 0.11, unclean.leader.election.enable now defaults to this.\nFundamentally, durability and consistency are at odds with availability. If there is no quorum, then no reads or writes can be accepted and the cluster is unavailable. This is the crux of the CAP theorem.\nIn part three of this series, we will discuss scaling message delivery in the distributed log.\n","wordCount":"3260","inLanguage":"en","datePublished":"2017-12-27T12:26:55-06:00","dateModified":"2018-02-23T16:09:46-06:00","mainEntityOfPage":{"@type":"WebPage","@id":"https://bravenewgeek.com/building-a-distributed-log-from-scratch-part-2-data-replication/"},"publisher":{"@type":"Organization","name":"Brave New Geek","logo":{"@type":"ImageObject","url":"https://bravenewgeek.com/favicon.ico"}}}</script></head><body id=top><header class=site-header><div class="wrap site-header-inner"><a class=wordmark href=https://bravenewgeek.com/ accesskey=h title="Brave New Geek (Alt + H)"><span class=wordmark-name>Brave New Geek</span>
<span class=wordmark-tag>Introspections of a software engineer</span></a><nav class=site-nav aria-label=Primary><a href=https://bravenewgeek.com/archive/>Archive</a>
<a href=https://bravenewgeek.com/tags/>Tags</a>
<a href=https://bravenewgeek.com/about-me/>About</a>
<button id=theme-toggle class=theme-toggle accesskey=t title="Toggle theme (Alt + T)" aria-label="Toggle light/dark theme">
<svg class="moon" width="18" height="18" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M21 12.79A9 9 0 1111.21 3 7 7 0 0021 12.79z"/></svg>
<svg class="sun" width="18" height="18" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><circle cx="12" cy="12" r="5"/><line x1="12" y1="1" x2="12" y2="3"/><line x1="12" y1="21" x2="12" y2="23"/><line x1="4.22" y1="4.22" x2="5.64" y2="5.64"/><line x1="18.36" y1="18.36" x2="19.78" y2="19.78"/><line x1="1" y1="12" x2="3" y2="12"/><line x1="21" y1="12" x2="23" y2="12"/><line x1="4.22" y1="19.78" x2="5.64" y2="18.36"/><line x1="18.36" y1="5.64" x2="19.78" y2="4.22"/></svg></button></nav></div></header><main class=main><article class="post wrap"><div class=post-return><a href=https://bravenewgeek.com/><span class=pager-arrow>&larr;</span> the log</a></div><header class=post-header><div class=post-meta><span class=post-offset>#70</span><time datetime=2017-12-27>2017-12-27</time><span>16 min read</span>
<span class=post-cats><a href=https://bravenewgeek.com/category/distributed-systems-2/>Distributed Systems</a><a href=https://bravenewgeek.com/category/messaging/>Messaging</a><a href=https://bravenewgeek.com/category/software-architecture/>Software Architecture</a><a href=https://bravenewgeek.com/category/software-engineering/>Software Engineering</a></span></div><h1 class=post-title>Building a Distributed Log from Scratch, Part 2: Data Replication</h1></header><div class="post-content md-content"><p>In <a href=https://bravenewgeek.com/building-a-distributed-log-from-scratch-part-1-storage-mechanics/>part one</a> of this series we introduced the idea of a message log, touched on why it’s useful, and discussed the storage mechanics behind it. In part two, we discuss data replication.</p><p>We have our log. We know how to write data to it and read it back as well as how data is persisted. The caveat to this is, although we have a durable log, it’s a single point of failure (SPOF). If the machine where the log data is stored dies, we’re SOL. Recall that one of our three priorities with this system is high availability, so the question is how do we achieve high availability and fault tolerance?</p><p><a href=/wp-content/uploads/2017/12/spof.png><img loading=lazy src=/wp-content/uploads/2017/12/spof.png></a></p><p>With high availability, we’re specifically talking about ensuring continuity of reads and writes. A server failing shouldn’t preclude either of these, or at least unavailability should be kept to an absolute minimum and without the need for operator intervention. Ensuring this continuity should be fairly obvious: we eliminate the SPOF. To do that, we replicate the data. Replication can also be a means for increasing scalability, but for now we’re only looking at this through the lens of high availability.</p><p><a href=/wp-content/uploads/2017/12/replicated_log.png><img loading=lazy src=/wp-content/uploads/2017/12/replicated_log.png></a></p><p>There are a number of ways we can go about replicating the log data. Broadly speaking, we can group the techniques into two different categories: gossip/multicast protocols and consensus protocols. The former includes things like epidemic broadcast trees, bimodal multicast, SWIM, HyParView, and NeEM. These tend to be eventually consistent and/or stochastic. The latter, which I’ve described in more detail <a href=https://bravenewgeek.com/understanding-consensus/>here</a>, includes 2PC/3PC, Paxos, Raft, Zab, and chain replication. These tend to favor strong consistency over availability.</p><p>So there are a lot of ways we can replicate data, but some of these solutions are better suited than others to this particular problem. Since ordering is an important property of a log, consistency becomes important for a <em>replicated</em> log. If we read from one replica and then read from another, it’s important those views of the log don’t conflict with each other. This more or less rules out the stochastic and eventually consistent options, leaving us with consensus-based replication.</p><p>There are essentially two components to consensus-based replication schemes: 1) designate a leader who is responsible for sequencing writes and 2) replicate the writes to the rest of the cluster.</p><p>Designating a leader can be as simple as a configuration setting, but the purpose of replication is fault tolerance. If our configured leader crashes, we’re no longer able to accept writes. This means we need the leader to be dynamic. It turns out leader election is a well-understood problem, so we’ll get to this in a bit.</p><p>Once a leader is established, it needs to replicate the data to followers. In general, this can be done by either waiting for all replicas or waiting for only a quorum (majority) of replicas. There are pros and cons to both approaches.</p><p><strong>Pros</strong></p><p><strong>Cons</strong></p><p><strong>All Replicas</strong></p><p>Tolerates <em>f</em> failures with <em>f+1</em> replicas</p><p>Latency pegged to slowest replica</p><p><strong>Quorum</strong></p><p>Hides delay from a slow replica</p><p>Tolerates <em>f</em> failures with <em>2f+1</em> replicas</p><p>Waiting on all replicas means we can make progress as long as at least one replica is available. With quorum, tolerating the same amount of failures requires more replicas because we need a majority to make progress. The trade-off is that the quorum hides any delays from a slow replica. Kafka is an example of a system which uses all replicas (with some conditions on this which we will see later), and NATS Streaming is one that uses a quorum. Let’s take a look at both in more detail.</p><h3 id=replication-in-kafka>Replication in Kafka<a hidden class=anchor aria-hidden=true href=#replication-in-kafka>#</a></h3><p>In Kafka, a leader is selected (we’ll touch on this in a moment). This leader maintains an in-sync replica set (ISR) consisting of all the replicas which are fully caught up with the leader. This is every replica, by definition, at the beginning. All reads and writes go through the leader. The leader writes messages to a write-ahead log (WAL). Messages written to the WAL are considered uncommitted or “dirty” initially. The leader only commits a message once all replicas in the ISR have written it to their own WAL. The leader also maintains a high-water mark (HW) which is the last committed message in the WAL. This gets piggybacked on the replica fetch responses from which replicas periodically checkpoint to disk for recovery purposes. The piggybacked HW then allows replicas to know when to commit.</p><p>Only committed messages are exposed to consumers. However, producers can configure how they want to receive acknowledgements on writes. It can wait until the message is committed on the leader (and thus replicated to the ISR), wait for the message to only be written (but not committed) to the leader’s WAL, or not wait at all. This all depends on what trade-offs the producer wants to make between latency and durability.</p><p>The graphic below shows how this replication process works for a cluster of three brokers: <em>b1</em>, <em>b2</em>, and <em>b3</em>. Followers are effectively special consumers of the leader’s log.</p><p><a href=/wp-content/uploads/2017/12/kafka_replication.png><img loading=lazy src=/wp-content/uploads/2017/12/kafka_replication.png></a></p><p>Now let’s look at a few failure modes and how Kafka handles them.</p><h4 id=leader-fails>Leader Fails<a hidden class=anchor aria-hidden=true href=#leader-fails>#</a></h4><p>Kafka relies on <a href=https://zookeeper.apache.org/>Apache ZooKeeper</a> for certain cluster coordination tasks, such as leader election, though this is not actually how the log leader is elected. A Kafka cluster has a single controller broker whose election is handled by ZooKeeper. This controller is responsible for performing administrative tasks on the cluster. One of these tasks is selecting a new log leader (actually <em>partition</em> leader, but this will be described later in the series) from the ISR when the current leader dies. ZooKeeper is also used to detect these broker failures and signal them to the controller.</p><p><a href=/wp-content/uploads/2017/12/kafka_leader_failure.png><img loading=lazy src=/wp-content/uploads/2017/12/kafka_leader_failure.png></a></p><p>Thus, when the leader crashes, the cluster controller is notified by ZooKeeper and it selects a new leader from the ISR and announces this to the followers. This gives us automatic failover of the leader. All committed messages up to the HW are preserved and uncommitted messages may be lost during the failover. In this case, <em>b1</em> fails and <em>b2</em> steps up as leader.</p><p><a href=/wp-content/uploads/2017/12/kafka_leader_failover.png><img loading=lazy src=/wp-content/uploads/2017/12/kafka_leader_failover.png></a></p><h4 id=follower-fails>Follower Fails<a hidden class=anchor aria-hidden=true href=#follower-fails>#</a></h4><p>The leader tracks information on how “caught up” each replica is. Before Kafka 0.9, this included both how many messages a replica was behind, <em>replica.lag.max.messages</em>, and the amount of time since the replica last fetched messages from the leader, <em>replica.lag.time.max.ms</em>. Since 0.9, <em>replica.lag.max.messages</em> was removed and <em>replica.lag.time.max.ms</em> now refers to both the time since the last fetch request <em>and</em> the amount of time since the replica last caught up.</p><p><a href=/wp-content/uploads/2017/12/kafka_follower_failure.png><img loading=lazy src=/wp-content/uploads/2017/12/kafka_follower_failure.png></a></p><p>Thus, when a follower fails (or stops fetching messages for whatever reason), the leader will detect this based on <em>replica.lag.time.max.ms</em>. After that time expires, the leader will consider the replica out of sync and remove it from the ISR. In this scenario, the cluster enters an “under-replicated” state since the ISR has shrunk. Specifically, <em>b2</em> fails and is removed from the ISR.</p><p><a href=/wp-content/uploads/2017/12/kafka_follower_failure_removed.png><img loading=lazy src=/wp-content/uploads/2017/12/kafka_follower_failure_removed.png></a></p><h4 id=follower-temporarily-partitioned>Follower Temporarily Partitioned<a hidden class=anchor aria-hidden=true href=#follower-temporarily-partitioned>#</a></h4><p>The case of a follower being temporarily partitioned, e.g. due to a transient network failure, is handled in a similar fashion to the follower itself failing. These two failure modes can really be combined since the latter is just the former with an arbitrarily long partition, i.e. it’s the difference between crash-stop and crash-recovery models.</p><p><a href=/wp-content/uploads/2017/12/kafka_follower_partition.png><img loading=lazy src=/wp-content/uploads/2017/12/kafka_follower_partition.png></a></p><p>In this case, <em>b3</em> is partitioned from the leader. As before, <em>replica.lag.time.max.ms</em> acts as our failure detector and causes <em>b3</em> to be removed from the ISR. We enter an under-replicated state and the remaining two brokers continue committing messages 4 and 5. Accordingly, the HW is updated to 5 on these brokers.</p><p><a href=/wp-content/uploads/2017/12/kafka_follower_partition_removed.png><img loading=lazy src=/wp-content/uploads/2017/12/kafka_follower_partition_removed.png></a></p><p>When the partition heals, <em>b3</em> continues reading from the leader and catching up. Once it is fully caught up with the leader, it’s added back into the ISR and the cluster resumes its fully replicated state.</p><p><a href=/wp-content/uploads/2017/12/kafka_follower_partition_healed.png><img loading=lazy src=/wp-content/uploads/2017/12/kafka_follower_partition_healed.png></a></p><p>We can generalize this to the crash-recovery model. For example, instead of a network partition, the follower could crash and be restarted later. When the failed replica is restarted, it recovers the HW from disk and truncates its log up to the HW. This preserves the invariant that messages after the HW are not guaranteed to be committed. At this point, it can begin catching up from the leader and will end up with a log consistent with the leader’s once fully caught up.</p><h3 id=replication-in-nats-streaming>Replication in NATS Streaming<a hidden class=anchor aria-hidden=true href=#replication-in-nats-streaming>#</a></h3><p>NATS Streaming relies on the <a href=https://raft.github.io/>Raft consensus algorithm</a> for leader election and data replication. This sometimes comes as a surprise to some as Raft is largely seen as a protocol for replicated state machines. We’ll try to understand why Raft was chosen for this particular problem in the following sections. We won’t dive deep into Raft itself beyond what is needed for the purposes of this discussion.</p><p>While a log is a state machine, it’s a very simple one: a series of appends. Raft is frequently used as the replication mechanism for key-value stores which have a clearer notion of “state machine.” For example, with a key-value store, we have <em>set</em> and <em>delete</em> operations. If we set <em>foo = bar</em> and then later set <em>foo = baz</em>, the state gets rolled up. That is, we don’t necessarily care about the provenance of the key, only its current state.</p><p>However, NATS Streaming differs from Kafka in a number of key ways. One of these differences is that NATS Streaming attempts to provide a sort of unified API for streaming and queueing semantics not too dissimilar from <a href=https://pulsar.apache.org/>Apache Pulsar</a>. This means, while it has a notion of a log, it also has subscriptions on that log. Unlike Kafka, NATS Streaming tracks these subscriptions and metadata associated with them, such as where a client is in the log. These have definite “state machines” affiliated with them, like creating and deleting subscriptions, positions in the log, clients joining or leaving queue groups, and message-redelivery information.</p><p>Currently, NATS Streaming uses multiple Raft groups for replication. There is a single metadata Raft group used for replicating client state and there is a separate Raft group per topic which replicates messages and subscriptions.</p><p>Raft solves both the problems of leader election and data replication in a single protocol. The <a href=http://thesecretlivesofdata.com/raft/>Secret Lives of Data</a> provides an excellent interactive illustration of how this works. As you step through that illustration, you’ll notice that the algorithm is actually quite similar to the Kafka replication protocol we walked through earlier. This is because although Raft is used to implement replicated state machines, it actually is a replicated WAL, which is exactly what Kafka is. One benefit of using Raft is we no longer have the need for ZooKeeper or some other coordination service.</p><p>Raft handles electing a leader. Heartbeats are used to maintain leadership. Writes flow through the leader to the followers. The leader appends writes to its WAL and they are subsequently piggybacked onto the heartbeats which get sent to the followers using <em>AppendEntries</em> messages. At this point, the followers append the write to their own WALs, assuming they don’t detect a gap, and send a response back to the leader. The leader commits the write once it receives a successful response from a quorum of followers.</p><p>Similar to Kafka, each replica in Raft maintains a high-water mark of sorts called the <em>commit index</em>, which is the index of the highest log entry known to be committed. This is piggybacked on the <em>AppendEntries</em> messages which the followers use to know when to commit entries in their WALs. If a follower detects that it missed an entry (i.e. there was a gap in the log), it rejects the <em>AppendEntries</em> and informs the leader to rewind the replication. The <a href=https://raft.github.io/raft.pdf>Raft paper</a> details how it ensures correctness, even in the face of many failure modes such as the ones described earlier.</p><p>Conceptually, there are two logs: the Raft log and the NATS Streaming message log. The Raft log handles replicating messages and, once committed, they are appended to the NATS Streaming log. If it seems like there’s some redundancy here, that’s because there is, which we’ll get to soon. However, keep in mind we’re not just replicating the message log, but also the state machines associated with the log and any clients.</p><p>There are a few challenges with this replication technique, two of which we will talk about. The first is scaling Raft. With a single topic, there is one Raft group, which means one node is elected leader and it heartbeats messages to followers.</p><p><a href=/wp-content/uploads/2017/12/raft_single_topic.png><img loading=lazy src=/wp-content/uploads/2017/12/raft_single_topic.png></a></p><p>As the number of topics increases, so do the number of Raft groups, each with their own leaders and heartbeats. Unless we constrain the Raft group participants or the number of topics, this creates an explosion of network traffic between nodes.</p><p><a href=/wp-content/uploads/2017/12/raft_many_topics.png><img loading=lazy src=/wp-content/uploads/2017/12/raft_many_topics.png></a></p><p>There are a couple ways we can go about addressing this. One option is to run a fixed number of Raft groups and use a consistent hash to map a topic to a group. This can work well if we know roughly the number of topics beforehand since we can size the number of Raft groups accordingly. If you expect only 10 topics, running 10 Raft groups is probably reasonable. But if you expect 10,000 topics, you probably don’t want 10,000 Raft groups. If hashing is consistent, it would be feasible to dynamically add or remove Raft groups at runtime, but it would still require repartitioning a portion of topics which can be complicated.</p><p><a href=/wp-content/uploads/2017/12/raft_fixed_groups.png><img loading=lazy src=/wp-content/uploads/2017/12/raft_fixed_groups.png></a></p><p>Another option is to run an entire node’s worth of topics as a single group using a layer on top of Raft. This is what CockroachDB does to scale Raft in proportion to the number of key ranges using a layer on top of Raft they call <a href=https://www.cockroachlabs.com/blog/scaling-raft/>MultiRaft</a>. This requires some cooperation from the Raft implementation, so it’s a bit more involved than the partitioning technique but eschews the repartitioning problem and redundant heartbeating.</p><p><a href=/wp-content/uploads/2017/12/multiraft.png><img loading=lazy src=/wp-content/uploads/2017/12/multiraft.png></a></p><p>The second challenge with using Raft for this problem is the issue of “dual writes.” As mentioned before, there are really two logs: the Raft log and the NATS Streaming message log, which we’ll call the “store.” When a message is published, the leader writes it to its Raft log and it goes through the Raft replication process.</p><p><a href=/wp-content/uploads/2017/12/wal.png><img loading=lazy src=/wp-content/uploads/2017/12/wal.png></a></p><p>Once the message is committed in Raft, it’s written to the NATS Streaming log and the message is now visible to consumers.</p><p><a href=/wp-content/uploads/2017/12/wal_committed.png><img loading=lazy src=/wp-content/uploads/2017/12/wal_committed.png></a></p><p>Note, however, that not only messages are written to the Raft log. We also have subscriptions and cluster topology changes, for instance. These other items are not written to the NATS Streaming log but handled in other ways on commit. That said, messages tend to occur in much greater volume than these other entries.</p><p><a href=/wp-content/uploads/2017/12/dual_writes.png><img loading=lazy src=/wp-content/uploads/2017/12/dual_writes.png></a></p><p>Messages end up getting stored redundantly, once in the Raft log and once in the NATS Streaming log. We can address this problem if we think about our logs a bit differently. If you recall from <a href=https://bravenewgeek.com/building-a-distributed-log-from-scratch-part-1-storage-mechanics/>part one</a>, our log storage consists of two parts: the log segment and the log index. The segment stores the actual log data, and the index stores a mapping from log offset to position in the segment.</p><p>Along these lines, we can think of the Raft log index as a “physical offset” and the NATS Streaming log index as a “logical offset.” Instead of maintaining two logs, we treat the Raft log as our message write-ahead log and treat the NATS Streaming log as an index into that WAL. Particularly, messages are written to the Raft log as usual. Once committed, we write an index entry for the message offset that points back into the log. As before, we use the index to do lookups into the log and can then read sequentially from the log itself.</p><p><a href=/wp-content/uploads/2017/12/raft_index.png><img loading=lazy src=/wp-content/uploads/2017/12/raft_index.png></a></p><h3 id=remaining-questions>Remaining Questions<a hidden class=anchor aria-hidden=true href=#remaining-questions>#</a></h3><p>We’ve answered the questions of how to ensure continuity of reads and writes, how to replicate data, and how to ensure replicas are consistent. The remaining two questions pertaining to replication are how do we keep things fast and how do we ensure data is durable?</p><p>There are several things we can do with respect to performance. The first is we can configure publisher acks depending on our application’s requirements. Specifically, we have three options. The first is the broker acks on commit. This is slow but safe as it guarantees the data is replicated. The second is the broker acks on appending to its local log. This is fast but unsafe since it doesn’t wait on any replica roundtrips but, by that very fact, means that the data is not replicated. If the leader crashes, the message could be lost. Lastly, the publisher can just not wait for an ack at all. This is the fastest but least safe option for obvious reasons. Tuning this all depends on what requirements and trade-offs make sense for your application.</p><p>The second thing we do is don’t explicitly <em>fsync</em> writes on the broker and instead rely on replication for durability. Both Kafka and NATS Streaming (when clustered) do this. With <em>fsync</em> enabled (in Kafka, this is configured with <em>flush.messages</em> and/or <em>flush.ms</em> and in NATS Streaming, with <em>file_sync</em>), every message that gets published results in a sync to disk. This ends up being very expensive. The thought here is if we are replicating to enough nodes, the replication itself is sufficient for HA of data since the likelihood of more than a quorum of nodes failing is low, especially if we are using rack-aware clustering. Note that data is still periodically flushed in the background by the kernel.</p><p>Batching aggressively is also a key part of ensuring good performance. Kafka supports end-to-end batching from the producer all the way to the consumer. NATS Streaming does not currently support batching at the API level, but it uses aggressive batching when replicating and persisting messages. In my experience, this makes about an order-of-magnitude improvement in throughput.</p><p>Finally, as already discussed earlier in the series, keeping disk access sequential and maximizing zero-copy reads makes a big difference as well.</p><p>There are a few things worth noting with respect to durability. Quorum is what guarantees durability of data. This comes “for free” with Raft due to the nature of that protocol. In Kafka, we need to do a bit of configuring to ensure this. Namely, we need to configure <em>min.insync.replicas</em> on the broker and <em>acks</em> on the producer. The former controls the minimum number of replicas that must acknowledge a write for it to be considered successful when a producer sets <em>acks</em> to “all.” The latter controls the number of acknowledgments the producer requires the leader to have received before considering a request complete. For example, with a topic that has a replication factor of three, <em>min.insync.replicas</em> needs to be set to two and <em>acks</em> set to “all.” This will, in effect, require a quorum of two replicas to process writes.</p><p>Another caveat with Kafka is unclean leader elections. That is, if all replicas become unavailable, there are two options: choose the first replica to come back to life (not necessarily in the ISR) and elect this replica as leader (which could result in data loss) or wait for a replica in the ISR to come back to life and elect it as leader (which could result in prolonged unavailability). Initially, Kafka favored availability by default by choosing the first strategy. If you preferred consistency, you needed to set <em>unclean.leader.election.enable</em> to <em>false</em>. However, as of 0.11, <em>unclean.leader.election.enable</em> now defaults to this.</p><p>Fundamentally, durability and consistency are at odds with availability. If there is no quorum, then no reads or writes can be accepted and the cluster is unavailable. This is the crux of the <a href=https://bravenewgeek.com/cap-and-the-illusion-of-choice/>CAP theorem</a>.</p><p>In <a href=https://bravenewgeek.com/building-a-distributed-log-from-scratch-part-3-scaling-message-delivery/>part three</a> of this series, we will discuss scaling message delivery in the distributed log.</p></div><footer class=post-footer><ul class=post-tags><li><a href=https://bravenewgeek.com/tag/algorithms-2/>algorithms</a></li><li><a href=https://bravenewgeek.com/tag/architecture/>Architecture</a></li><li><a href=https://bravenewgeek.com/tag/building-a-distributed-log-from-scratch/>Building a Distributed Log From Scratch</a></li><li><a href=https://bravenewgeek.com/tag/cap-theorem/>Cap Theorem</a></li><li><a href=https://bravenewgeek.com/tag/consensus/>Consensus</a></li><li><a href=https://bravenewgeek.com/tag/consistency/>Consistency</a></li><li><a href=https://bravenewgeek.com/tag/data-replication/>Data Replication</a></li><li><a href=https://bravenewgeek.com/tag/distributed-log/>Distributed Log</a></li><li><a href=https://bravenewgeek.com/tag/distributed-systems/>Distributed Systems</a></li><li><a href=https://bravenewgeek.com/tag/kafka/>Kafka</a></li><li><a href=https://bravenewgeek.com/tag/leader-election/>Leader Election</a></li><li><a href=https://bravenewgeek.com/tag/message-queues/>Message Queues</a></li><li><a href=https://bravenewgeek.com/tag/message-oriented-middleware/>Message-Oriented Middleware</a></li><li><a href=https://bravenewgeek.com/tag/messaging/>Messaging</a></li><li><a href=https://bravenewgeek.com/tag/nats/>Nats</a></li><li><a href=https://bravenewgeek.com/tag/nats-streaming/>Nats Streaming</a></li><li><a href=https://bravenewgeek.com/tag/performance/>Performance</a></li><li><a href=https://bravenewgeek.com/tag/raft/>Raft</a></li><li><a href=https://bravenewgeek.com/tag/stream-processing/>Stream Processing</a></li><li><a href=https://bravenewgeek.com/tag/write-ahead-log/>Write-Ahead Log</a></li><li><a href=https://bravenewgeek.com/tag/zookeeper/>Zookeeper</a></li></ul><nav class=post-nav aria-label="Adjacent posts"><a class=post-nav-link href=https://bravenewgeek.com/building-a-distributed-log-from-scratch-part-3-scaling-message-delivery/><span class=post-nav-dir><span class=pager-arrow>&larr;</span> newer</span>
<span class=post-nav-title>Building a Distributed Log from Scratch, Part 3: Scaling Message Delivery</span>
</a><a class="post-nav-link post-nav-right" href=https://bravenewgeek.com/building-a-distributed-log-from-scratch-part-1-storage-mechanics/><span class=post-nav-dir>older <span class=pager-arrow>&rarr;</span></span>
<span class=post-nav-title>Building a Distributed Log from Scratch, Part 1: Storage Mechanics</span></a></nav></footer><section class=wp-comments><h2>Comments</h2><p class=wp-comments-notice>Comments are from this blog's WordPress era and are preserved read-only.</p><article class=wp-comment><header><span class=wp-comment-author>Shiju Varghese</span>
<time class=wp-comment-date>December 28, 2017</time></header><div class=wp-comment-body><p>Brilliant article.</p></div></article><article class=wp-comment><header><span class=wp-comment-author>Byron Ruth</span>
<time class=wp-comment-date>December 28, 2017</time></header><div class=wp-comment-body><p>Great article Tyler. Two clarifications I have.</p><p>First, you first stated that NATS Streaming currently uses one Raft Group per topic, but then mentioned the problem with this (explosion of network traffic). Is the plan to adopt the MultiRaft approach the Cockroach Labs folks developed?</p><p>Second, in the diagram of the message log being an index of the logical offset which points to the Raft log physical offset, I am assuming this implies the physical message (bytes) are stored only in the Raft log? What is the cost of dereferencing the message while reading the message log? I am making the assumption messages are currently read sequentially from the log files and this approach introduces a new indirection.</p></div><div class=wp-comment-replies><article class=wp-comment><header><span class=wp-comment-author>Tyler Treat</span>
<time class=wp-comment-date>December 28, 2017</time></header><div class=wp-comment-body><p>Yes, the MVP currently does not address the Raft scalability problem. These are some potential solutions but not yet implemented. This goes for the dual writes in the Raft log. However, with this approach, reads would still be sequential from the log (as they are now). These will be two big areas for improvement following an initial release.</p></div></article></div></article><article class=wp-comment><header><span class=wp-comment-author>Ming</span>
<time class=wp-comment-date>January 3, 2018</time></header><div class=wp-comment-body><p>Hi, are you trying to actually build a distributed log in this serials of blog post?</p></div></article><article class=wp-comment><header><span class=wp-comment-author>Chris Young</span>
<time class=wp-comment-date>January 23, 2018</time></header><div class=wp-comment-body><p>&#8220;By default, Kafka favors availability by choosing the second strategy.&#8221; &#8212; you meant the *first* strategy here &#8212; picking a replica not necessarily in the ISR set? But actually, since 0.11, unclean.leader.election.enable defaults to false, so please update that paragraph :)</p></div><div class=wp-comment-replies><article class=wp-comment><header><span class=wp-comment-author>Tyler Treat</span>
<time class=wp-comment-date>January 23, 2018</time></header><div class=wp-comment-body><p>Nice catch, thanks.</p></div></article></div></article><article class=wp-comment><header><span class=wp-comment-author>Aravind</span>
<time class=wp-comment-date>December 23, 2021</time></header><div class=wp-comment-body><p>Regarding these 2 sections of text</p><p>> The thought here is if we are replicating to enough nodes, the replication itself is sufficient for HA of data since the likelihood of more than a quorum of nodes failing is low, especially if we are using rack-aware clustering. &#8230;</p><p>AND</p><p>> Another caveat with Kafka is unclean leader elections. That is, if all replicas become unavailable, there are two options: &#8230;</p><p>Why is it that we&#8217;re worried about all replicas becoming unavailable in Kafka for leader elections, but not as much from the replication/durability perspective?</p></div></article></section></article></main><footer class=site-footer><div class="wrap site-footer-inner"><div class=footer-meta><span class=footer-copy>&copy; 2026 Tyler Treat</span>
<span class="footer-sep footer-dot">·</span>
<span class=footer-links><a href=/feed/>rss</a>
<span class=footer-sep>·</span>
<a href=https://github.com/tylertreat target=_blank rel="noopener noreferrer me">github</a>
<span class=footer-sep>·</span>
<a href=https://www.linkedin.com/in/ttreat/ target=_blank rel="noopener noreferrer me">linkedin</a></span></div></div></footer><a href=#top id=top-link class="top-link hidden" aria-label="go to top" title="Go to Top (Alt + G)" accesskey=g><svg viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="feather feather-chevrons-up"><polyline points="17 11 12 6 7 11"/><polyline points="17 18 12 13 7 18"/></svg>
</a><script>let menu=document.getElementById("menu");if(menu){const e=localStorage.getItem("menu-scroll-position");e&&(menu.scrollLeft=parseInt(e,10)),menu.onscroll=function(){localStorage.setItem("menu-scroll-position",menu.scrollLeft)}}document.querySelectorAll('a[href^="#"]').forEach(e=>{e.addEventListener("click",function(e){e.preventDefault();var t=this.getAttribute("href").substr(1);window.matchMedia("(prefers-reduced-motion: reduce)").matches?document.querySelector(`[id='${decodeURIComponent(t)}']`).scrollIntoView():document.querySelector(`[id='${decodeURIComponent(t)}']`).scrollIntoView({behavior:"smooth"}),t==="top"?history.replaceState(null,null," "):history.pushState(null,null,`#${t}`)})})</script><script>var toplink=document.getElementById("top-link");window.onscroll=function(){const e=window.innerHeight;document.body.scrollTop>e||document.documentElement.scrollTop>e?toplink.classList.remove("hidden"):toplink.classList.add("hidden")}</script><script>document.getElementById("theme-toggle").addEventListener("click",()=>{const e=document.querySelector("html");e.dataset.theme==="dark"?(e.dataset.theme="light",localStorage.setItem("pref-theme","light")):(e.dataset.theme="dark",localStorage.setItem("pref-theme","dark"))})</script><script>document.querySelectorAll("pre > code").forEach(e=>{const n=e.parentNode.parentNode,t=document.createElement("button");t.classList.add("copy-code"),t.innerHTML="copy";function s(){t.innerHTML="copied!",setTimeout(()=>{t.innerHTML="copy"},2e3)}t.addEventListener("click",t=>{if("clipboard"in navigator){navigator.clipboard.writeText(e.textContent),s();return}const n=document.createRange();n.selectNodeContents(e);const o=window.getSelection();o.removeAllRanges(),o.addRange(n);try{document.execCommand("copy"),s()}catch{}o.removeRange(n)}),n.classList.contains("highlight")?n.appendChild(t):n.parentNode.firstChild==n||(e.parentNode.parentNode.parentNode.parentNode.parentNode.nodeName=="TABLE"?e.parentNode.parentNode.parentNode.parentNode.parentNode.appendChild(t):e.parentNode.appendChild(t))})</script></body></html>