24 lines
32 KiB
HTML
24 lines
32 KiB
HTML
<!doctype html><html lang=en dir=auto data-theme=auto><head><meta charset=utf-8><meta http-equiv=X-UA-Compatible content="IE=edge"><meta name=viewport content="width=device-width,initial-scale=1,shrink-to-fit=no"><meta name=robots content="index, follow"><title>Introducing Liftbridge: Lightweight, Fault-Tolerant Message Streams | Brave New Geek</title><meta name=keywords content="cloud-native,distributed log,distributed systems,gnatsd,kafka,liftbridge,message queues,message-oriented middleware,messaging,nats,nats streaming,open source,pulsar,raft,scalability,stream processing,write-ahead log"><meta name=description content="
|
||
Last week I open sourced Liftbridge, my latest project and contribution to the Cloud Native Computing Foundation ecosystem. Liftbridge is a system for lightweight, fault-tolerant (LIFT) message streams built on NATS and gRPC. Fundamentally, it extends NATS with a Kafka-like publish-subscribe log API that is highly available and horizontally scalable.
|
||
I’ve been working on Liftbridge for the past couple of months, but it’s something I’ve been thinking about for over a year. I sketched out the design for it last year and wrote about it in January. It was largely inspired while I was working on NATS Streaming, which I’m currently still the second top contributor to. My primary involvement with NATS Streaming was building out the early data replication and clustering solution for high availability, which has continued to evolve since I left the project. In many ways, Liftbridge is about applying a lot of the things I learned while working on NATS Streaming as well as my observations from being closely involved with the NATS community for some time. It’s also the product of scratching an itch I’ve had since these are the kinds of problems I enjoy working on, and I needed something to code."><meta name=author content><link rel=canonical href=https://bravenewgeek.com/introducing-liftbridge-lightweight-fault-tolerant-message-streams/><link crossorigin=anonymous href=/assets/css/stylesheet.4861a452a4c13a9a1fbf2085400b74a7de96b1beeb94dee57b4273e5dffcf337.css integrity="sha256-SGGkUqTBOpofvyCFQAt0p96Wsb7rlN7le0Jz5d/88zc=" rel="preload stylesheet" as=style><link rel=icon href=https://bravenewgeek.com/favicon.ico><link rel=icon type=image/png sizes=16x16 href=https://bravenewgeek.com/favicon.ico><link rel=icon type=image/png sizes=32x32 href=https://bravenewgeek.com/favicon.ico><link rel=apple-touch-icon href=https://bravenewgeek.com/favicon.ico><link rel=mask-icon href=https://bravenewgeek.com/favicon.ico><meta name=theme-color content="#2e2e33"><meta name=msapplication-TileColor content="#2e2e33"><link rel=alternate hreflang=en href=https://bravenewgeek.com/introducing-liftbridge-lightweight-fault-tolerant-message-streams/><noscript><style>#theme-toggle,.top-link{display:none}</style><style>@media(prefers-color-scheme:dark){:root{--theme:rgb(29, 30, 32);--entry:rgb(46, 46, 51);--primary:rgb(218, 218, 219);--secondary:rgb(155, 156, 157);--tertiary:rgb(65, 66, 68);--content:rgb(196, 196, 197);--code-block-bg:rgb(46, 46, 51);--code-bg:rgb(55, 56, 62);--border:rgb(51, 51, 51);color-scheme:dark}.list{background:var(--theme)}.toc{background:var(--entry)}}</style></noscript><script>localStorage.getItem("pref-theme")==="dark"?document.querySelector("html").dataset.theme="dark":localStorage.getItem("pref-theme")==="light"?document.querySelector("html").dataset.theme="light":window.matchMedia("(prefers-color-scheme: dark)").matches?document.querySelector("html").dataset.theme="dark":document.querySelector("html").dataset.theme="light"</script><link rel=preconnect href=https://fonts.googleapis.com><link rel=preconnect href=https://fonts.gstatic.com crossorigin><link rel=stylesheet href="https://fonts.googleapis.com/css2?family=JetBrains+Mono:wght@400;500&family=Source+Serif+4:ital,opsz,wght@0,8..60,400;0,8..60,600;1,8..60,400&family=Space+Grotesk:wght@500;600;700&display=swap"><meta property="og:url" content="https://bravenewgeek.com/introducing-liftbridge-lightweight-fault-tolerant-message-streams/"><meta property="og:site_name" content="Brave New Geek"><meta property="og:title" content="Introducing Liftbridge: Lightweight, Fault-Tolerant Message Streams"><meta property="og:description" content="Last week I open sourced Liftbridge, my latest project and contribution to the Cloud Native Computing Foundation ecosystem. Liftbridge is a system for lightweight, fault-tolerant (LIFT) message streams built on NATS and gRPC. Fundamentally, it extends NATS with a Kafka-like publish-subscribe log API that is highly available and horizontally scalable.
|
||
I’ve been working on Liftbridge for the past couple of months, but it’s something I’ve been thinking about for over a year. I sketched out the design for it last year and wrote about it in January. It was largely inspired while I was working on NATS Streaming, which I’m currently still the second top contributor to. My primary involvement with NATS Streaming was building out the early data replication and clustering solution for high availability, which has continued to evolve since I left the project. In many ways, Liftbridge is about applying a lot of the things I learned while working on NATS Streaming as well as my observations from being closely involved with the NATS community for some time. It’s also the product of scratching an itch I’ve had since these are the kinds of problems I enjoy working on, and I needed something to code."><meta property="og:locale" content="en_us"><meta property="og:type" content="article"><meta property="article:section" content="posts"><meta property="article:published_time" content="2018-07-27T17:42:49-05:00"><meta property="article:modified_time" content="2018-09-13T23:39:30-05:00"><meta property="article:tag" content="Cloud-Native"><meta property="article:tag" content="Distributed Log"><meta property="article:tag" content="Distributed Systems"><meta property="article:tag" content="Gnatsd"><meta property="article:tag" content="Kafka"><meta property="article:tag" content="Liftbridge"><meta name=twitter:card content="summary"><meta name=twitter:title content="Introducing Liftbridge: Lightweight, Fault-Tolerant Message Streams"><meta name=twitter:description content="Last week I open sourced Liftbridge, my latest project and contribution to the Cloud Native Computing Foundation ecosystem. Liftbridge is a system for lightweight, fault-tolerant (LIFT) message streams built on NATS and gRPC. Fundamentally, it extends NATS with a Kafka-like publish-subscribe log API that is highly available and horizontally scalable.
|
||
I’ve been working on Liftbridge for the past couple of months, but it’s something I’ve been thinking about for over a year. I sketched out the design for it last year and wrote about it in January. It was largely inspired while I was working on NATS Streaming, which I’m currently still the second top contributor to. My primary involvement with NATS Streaming was building out the early data replication and clustering solution for high availability, which has continued to evolve since I left the project. In many ways, Liftbridge is about applying a lot of the things I learned while working on NATS Streaming as well as my observations from being closely involved with the NATS community for some time. It’s also the product of scratching an itch I’ve had since these are the kinds of problems I enjoy working on, and I needed something to code."><script type=application/ld+json>{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Posts","item":"https://bravenewgeek.com/posts/"},{"@type":"ListItem","position":2,"name":"Introducing Liftbridge: Lightweight, Fault-Tolerant Message Streams","item":"https://bravenewgeek.com/introducing-liftbridge-lightweight-fault-tolerant-message-streams/"}]}</script><script type=application/ld+json>{"@context":"https://schema.org","@type":"BlogPosting","headline":"Introducing Liftbridge: Lightweight, Fault-Tolerant Message Streams","name":"Introducing Liftbridge: Lightweight, Fault-Tolerant Message Streams","description":"\nLast week I open sourced Liftbridge, my latest project and contribution to the Cloud Native Computing Foundation ecosystem. Liftbridge is a system for lightweight, fault-tolerant (LIFT) message streams built on NATS and gRPC. Fundamentally, it extends NATS with a Kafka-like publish-subscribe log API that is highly available and horizontally scalable.\nI’ve been working on Liftbridge for the past couple of months, but it’s something I’ve been thinking about for over a year. I sketched out the design for it last year and wrote about it in January. It was largely inspired while I was working on NATS Streaming, which I’m currently still the second top contributor to. My primary involvement with NATS Streaming was building out the early data replication and clustering solution for high availability, which has continued to evolve since I left the project. In many ways, Liftbridge is about applying a lot of the things I learned while working on NATS Streaming as well as my observations from being closely involved with the NATS community for some time. It’s also the product of scratching an itch I’ve had since these are the kinds of problems I enjoy working on, and I needed something to code.\n","keywords":["cloud-native","distributed log","distributed systems","gnatsd","kafka","liftbridge","message queues","message-oriented middleware","messaging","nats","nats streaming","open source","pulsar","raft","scalability","stream processing","write-ahead log"],"articleBody":"\nLast week I open sourced Liftbridge, my latest project and contribution to the Cloud Native Computing Foundation ecosystem. Liftbridge is a system for lightweight, fault-tolerant (LIFT) message streams built on NATS and gRPC. Fundamentally, it extends NATS with a Kafka-like publish-subscribe log API that is highly available and horizontally scalable.\nI’ve been working on Liftbridge for the past couple of months, but it’s something I’ve been thinking about for over a year. I sketched out the design for it last year and wrote about it in January. It was largely inspired while I was working on NATS Streaming, which I’m currently still the second top contributor to. My primary involvement with NATS Streaming was building out the early data replication and clustering solution for high availability, which has continued to evolve since I left the project. In many ways, Liftbridge is about applying a lot of the things I learned while working on NATS Streaming as well as my observations from being closely involved with the NATS community for some time. It’s also the product of scratching an itch I’ve had since these are the kinds of problems I enjoy working on, and I needed something to code.\nAt its core, Liftbridge is a server that implements a durable, replicated message log for the NATS messaging system. Clients create a named stream which is attached to a NATS subject. The stream then records messages on that subject to a replicated write-ahead log. Multiple consumers can read back from the same stream, and multiple streams can be attached to the same subject.\nThe goal is to bridge the gap between sophisticated log-based messaging systems like Apache Kafka and Apache Pulsar and simpler, cloud-native systems. This meant not relying on external coordination services like ZooKeeper, not using the JVM, keeping the API as simple and small as possible, and keeping client libraries thin. The system is written in Go, making it a single static binary with a small footprint (~16MB). It relies on the Raft consensus algorithm to do coordination. It has a very minimal API (just three endpoints at the moment). And the API uses gRPC, so client libraries can be generated for most popular programming languages (there is a Go client which provides some additional wrapper logic, but it’s pretty thin). The goal is to keep Liftbridge very _lightweight—_in terms of runtime, operations, and complexity.\nHowever, the bigger goal of Liftbridge is to extend NATS with a durable, at-least-once delivery mechanism that upholds the NATS tenets of simplicity, performance, and scalability. Unlike NATS Streaming, it uses the core NATS protocol with optional extensions. This means it can be added to an existing NATS deployment to provide message durability with no code changes.\nNATS Streaming provides a similar log-based messaging solution. However, it is an entirely separate protocol built on top of NATS. NATS is an implementation detail—the transport—for NATS Streaming. This means the two systems have separate messaging namespaces—messages published to NATS are not accessible from NATS Streaming and vice versa. Of course, it’s a bit more nuanced than this because, in reality, NATS Streaming is using NATS subjects underneath; technically messages can be accessed, but they are serialized protobufs. These nuances often get confounded by first–time users as it’s not always clear that NATS and NATS Streaming are completely separate systems. NATS Streaming also does not support wildcard subscriptions, which sometimes surprises users since it’s a major feature of NATS.\nAs a result, Liftbridge was built to augment NATS with durability rather than providing a completely separate system. To be clear, it’s still a separate server, but it merely acts as a write-ahead log for NATS subjects. NATS Streaming provides a broader set of features such as durable subscriptions, queue groups, pluggable storage backends, and multiple fault-tolerance modes. Liftbridge aims to have a relatively small API surface area.\nThe key features that differentiate Liftbridge are the shared message namespace, wildcards, log compaction, and horizontal scalability. NATS Streaming replicates channels to the entire cluster through a single Raft group, so adding servers does not help with scalability and actually creates a head-of-line bottleneck since everything is replicated through a single consensus group (n.b. NATS Streaming does have a partitioning mechanism, but it cannot be used in conjunction with clustering). Liftbridge allows replicating to a subset of the cluster, and each stream is replicated independently in parallel. This allows the cluster to scale horizontally and partition workloads more easily within a single, multi-tenant cluster.\nSome of the key features of Liftbridge include:\nLog-based API for NATS Replicated for fault-tolerance Horizontally scalable Wildcard subscription support At-least-once delivery support Message key-value support Log compaction by key (WIP) Single static binary (~16MB) Designed to be high-throughput (more on this to come) Supremely simple Initially, Liftbridge is designed to point to an existing NATS deployment. In the future, there will be support for a “standalone” mode where it can run with an embedded NATS server, allowing for a single deployable process. And in support of the “cloud-native” model, there is work to be done to make Liftbridge play nice with Kubernetes and generally productionalize the system, such as implementing an Operator and providing better instrumentation—perhaps with Prometheus support.\nOver the coming weeks and months, I will be going into more detail on Liftbridge, including the internals of it—such as its replication protocol—and providing benchmarks for the system. Of course, there’s also a lot of work yet to be done on it, so I’ll be continuing to work on that. There are many interesting problems that still need solved, so consider this my appeal to contributors. :)\n","wordCount":"931","inLanguage":"en","datePublished":"2018-07-27T17:42:49-05:00","dateModified":"2018-09-13T23:39:30-05:00","mainEntityOfPage":{"@type":"WebPage","@id":"https://bravenewgeek.com/introducing-liftbridge-lightweight-fault-tolerant-message-streams/"},"publisher":{"@type":"Organization","name":"Brave New Geek","logo":{"@type":"ImageObject","url":"https://bravenewgeek.com/favicon.ico"}}}</script></head><body id=top><header class=site-header><div class="wrap site-header-inner"><a class=wordmark href=https://bravenewgeek.com/ accesskey=h title="Brave New Geek (Alt + H)"><span class=wordmark-name>Brave New Geek</span>
|
||
<span class=wordmark-tag>Introspections of a software engineer</span></a><nav class=site-nav aria-label=Primary><a href=https://bravenewgeek.com/archive/>Archive</a>
|
||
<a href=https://bravenewgeek.com/tags/>Tags</a>
|
||
<a href=https://bravenewgeek.com/about-me/>About</a>
|
||
<button id=theme-toggle class=theme-toggle accesskey=t title="Toggle theme (Alt + T)" aria-label="Toggle light/dark theme">
|
||
<svg class="moon" width="18" height="18" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M21 12.79A9 9 0 1111.21 3 7 7 0 0021 12.79z"/></svg>
|
||
<svg class="sun" width="18" height="18" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><circle cx="12" cy="12" r="5"/><line x1="12" y1="1" x2="12" y2="3"/><line x1="12" y1="21" x2="12" y2="23"/><line x1="4.22" y1="4.22" x2="5.64" y2="5.64"/><line x1="18.36" y1="18.36" x2="19.78" y2="19.78"/><line x1="1" y1="12" x2="3" y2="12"/><line x1="21" y1="12" x2="23" y2="12"/><line x1="4.22" y1="19.78" x2="5.64" y2="18.36"/><line x1="18.36" y1="5.64" x2="19.78" y2="4.22"/></svg></button></nav></div></header><main class=main><article class="post wrap"><div class=post-return><a href=https://bravenewgeek.com/><span class=pager-arrow>←</span> the log</a></div><header class=post-header><div class=post-meta><span class=post-offset>#79</span><time datetime=2018-07-27>2018-07-27</time><span>5 min read</span>
|
||
<span class=post-cats><a href=https://bravenewgeek.com/category/distributed-systems-2/>Distributed Systems</a><a href=https://bravenewgeek.com/category/liftbridge/>Liftbridge</a><a href=https://bravenewgeek.com/category/messaging/>Messaging</a></span></div><h1 class=post-title>Introducing Liftbridge: Lightweight, Fault-Tolerant Message Streams</h1></header><div class="post-content md-content"><p><a href=https://github.com/liftbridge-io/liftbridge><img loading=lazy src=/wp-content/uploads/2018/07/liftbridge.png></a></p><p><a href=https://twitter.com/tyler_treat/status/1019281381493526529>Last week</a> I open sourced <a href=https://github.com/liftbridge-io/liftbridge>Liftbridge</a>, my latest project and contribution to the <a href=https://www.cncf.io/>Cloud Native Computing Foundation</a> ecosystem. Liftbridge is a system for lightweight, fault-tolerant (LIFT) message streams built on <a href=https://nats.io/>NATS</a> and <a href=https://grpc.io/>gRPC</a>. Fundamentally, it extends NATS with a <a href=https://kafka.apache.org/>Kafka</a>-like publish-subscribe log API that is highly available and horizontally scalable.</p><p>I’ve been working on Liftbridge for the past couple of months, but it’s something I’ve been thinking about for over a year. I sketched out the design for it last year and <a href=https://bravenewgeek.com/building-a-distributed-log-from-scratch-part-5-sketching-a-new-system/>wrote about it</a> in January. It was largely inspired while I was working on <a href=https://github.com/nats-io/nats-streaming-server>NATS Streaming</a>, which I’m currently still the second top contributor to. My primary involvement with NATS Streaming was building out the early data replication and clustering solution for high availability, which has continued to evolve since I left the project. In many ways, Liftbridge is about applying a lot of the things I learned while working on NATS Streaming as well as my observations from being closely involved with the NATS community for some time. It’s also the product of scratching an itch I’ve had since these are the kinds of problems I enjoy working on, and I needed something to code.</p><p>At its core, Liftbridge is a server that implements a durable, replicated message log for the NATS messaging system. Clients create a named <em>stream</em> which is attached to a NATS subject. The stream then records messages on that subject to a replicated write-ahead log. Multiple consumers can read back from the same stream, and multiple streams can be attached to the same subject.</p><p><a href=/wp-content/uploads/2018/07/liftbridge-high-level.png><img loading=lazy src=/wp-content/uploads/2018/07/liftbridge-high-level.png></a></p><p>The goal is to bridge the gap between sophisticated log-based messaging systems like Apache Kafka and <a href=https://pulsar.incubator.apache.org/>Apache Pulsar</a> and simpler, cloud-native systems. This meant not relying on external coordination services like ZooKeeper, not using the JVM, keeping the API as simple and small as possible, and keeping client libraries thin. The system is written in Go, making it a single static binary with a small footprint (~16MB). It relies on the <a href=https://raft.github.io/>Raft</a> consensus algorithm to do coordination. It has a <em>very</em> <a href=https://github.com/liftbridge-io/liftbridge-grpc/blob/d658c291552f32ce810995c5e9dca9862ecc44da/api.proto#L104-L120>minimal API</a> (just three endpoints at the moment). And the API uses gRPC, so client libraries can be generated for most popular programming languages (there is a <a href=https://github.com/liftbridge-io/go-liftbridge>Go client</a> which provides some additional wrapper logic, but it’s pretty thin). The goal is to keep Liftbridge very _lightweight—_in terms of runtime, operations, and complexity.</p><p>However, the bigger goal of Liftbridge is to <em>extend</em> NATS with a durable, at-least-once delivery mechanism that upholds the NATS tenets of simplicity, performance, and scalability. Unlike NATS Streaming, it uses the core NATS protocol with optional extensions. This means it can be added to an existing NATS deployment to provide message durability with no code changes.</p><p>NATS Streaming provides a similar log-based messaging solution. However, it is an entirely separate protocol built on top of NATS. NATS is an implementation detail—the <em>transport</em>—for NATS Streaming. This means the two systems have separate messaging namespaces—messages published to NATS are not accessible from NATS Streaming and vice versa. Of course, it’s a bit more nuanced than this because, in reality, NATS Streaming is using NATS subjects underneath; technically messages can be <em>accessed</em>, but they are serialized protobufs. These <a href=https://github.com/nats-io/nats-streaming-server/issues/609>nuances</a> <a href=https://github.com/nats-io/gnatsd/issues/715>often</a> <a href=https://github.com/nats-io/gnatsd/issues/714>get</a> <a href=https://github.com/nats-io/go-nats-streaming/issues/152>confounded</a> <a href=https://github.com/nats-io/gnatsd/issues/713>by</a> <a href=https://github.com/nats-io/gnatsd/issues/514>first</a>–<a href=https://github.com/nats-io/go-nats/issues/251>time</a> <a href=https://github.com/nats-io/go-nats/issues/328>users</a> as it’s not always clear that NATS and NATS Streaming are completely separate systems. NATS Streaming also <a href=https://github.com/nats-io/nats-streaming-server/issues/290>does not support wildcard subscriptions</a>, which sometimes surprises users since it’s a major feature of NATS.</p><p>As a result, Liftbridge was built to <em>augment</em> NATS with durability rather than providing a completely separate system. To be clear, it’s still a separate <em>server</em>, but it merely acts as a write-ahead log for NATS subjects. NATS Streaming provides a broader set of features such as durable subscriptions, queue groups, pluggable storage backends, and multiple fault-tolerance modes. Liftbridge aims to have a relatively small API surface area.</p><p>The key features that differentiate Liftbridge are the shared message namespace, wildcards, log compaction, and horizontal scalability. NATS Streaming replicates channels to the entire cluster through a single Raft group, so adding servers does not help with scalability and actually creates a head-of-line bottleneck since everything is replicated through a single consensus group (n.b. NATS Streaming does have a partitioning mechanism, but it cannot be used in conjunction with clustering). Liftbridge allows replicating to a subset of the cluster, and each stream is replicated independently in parallel. This allows the cluster to scale horizontally and partition workloads more easily within a single, multi-tenant cluster.</p><p><a href=/wp-content/uploads/2018/07/liftbridge-streams.png><img loading=lazy src=/wp-content/uploads/2018/07/liftbridge-streams.png></a></p><p>Some of the key features of Liftbridge include:</p><ul><li>Log-based API for NATS</li><li>Replicated for fault-tolerance</li><li>Horizontally scalable</li><li>Wildcard subscription support</li><li>At-least-once delivery support</li><li>Message key-value support</li><li>Log compaction by key (WIP)</li><li>Single static binary (~16MB)</li><li>Designed to be high-throughput (more on this to come)</li><li>Supremely simple</li></ul><p>Initially, Liftbridge is designed to point to an existing NATS deployment. In the future, there will be support for a “standalone” mode where it can run with an embedded NATS server, allowing for a single deployable process. And in support of the “cloud-native” model, there is work to be done to make Liftbridge play nice with Kubernetes and generally productionalize the system, such as implementing an <a href=https://coreos.com/operators/>Operator</a> and providing better instrumentation—perhaps with <a href=https://prometheus.io/>Prometheus</a> support.</p><p>Over the coming weeks and months, I will be going into more detail on Liftbridge, including the internals of it—such as its replication protocol—and providing benchmarks for the system. Of course, there’s also a lot of work yet to be done on it, so I’ll be continuing to work on that. There are many interesting problems that still need solved, so consider this my appeal to contributors. :)</p></div><footer class=post-footer><ul class=post-tags><li><a href=https://bravenewgeek.com/tag/cloud-native/>Cloud-Native</a></li><li><a href=https://bravenewgeek.com/tag/distributed-log/>Distributed Log</a></li><li><a href=https://bravenewgeek.com/tag/distributed-systems/>Distributed Systems</a></li><li><a href=https://bravenewgeek.com/tag/gnatsd/>Gnatsd</a></li><li><a href=https://bravenewgeek.com/tag/kafka/>Kafka</a></li><li><a href=https://bravenewgeek.com/tag/liftbridge/>Liftbridge</a></li><li><a href=https://bravenewgeek.com/tag/message-queues/>Message Queues</a></li><li><a href=https://bravenewgeek.com/tag/message-oriented-middleware/>Message-Oriented Middleware</a></li><li><a href=https://bravenewgeek.com/tag/messaging/>Messaging</a></li><li><a href=https://bravenewgeek.com/tag/nats/>Nats</a></li><li><a href=https://bravenewgeek.com/tag/nats-streaming/>Nats Streaming</a></li><li><a href=https://bravenewgeek.com/tag/open-source/>Open Source</a></li><li><a href=https://bravenewgeek.com/tag/pulsar/>Pulsar</a></li><li><a href=https://bravenewgeek.com/tag/raft/>Raft</a></li><li><a href=https://bravenewgeek.com/tag/scalability/>Scalability</a></li><li><a href=https://bravenewgeek.com/tag/stream-processing/>Stream Processing</a></li><li><a href=https://bravenewgeek.com/tag/write-ahead-log/>Write-Ahead Log</a></li></ul><nav class=post-nav aria-label="Adjacent posts"><a class=post-nav-link href=https://bravenewgeek.com/the-observability-pipeline/><span class=post-nav-dir><span class=pager-arrow>←</span> newer</span>
|
||
<span class=post-nav-title>The Observability Pipeline</span>
|
||
</a><a class="post-nav-link post-nav-right" href=https://bravenewgeek.com/gcp-and-aws-whats-the-difference/><span class=post-nav-dir>older <span class=pager-arrow>→</span></span>
|
||
<span class=post-nav-title>GCP and AWS: What’s the Difference?</span></a></nav></footer><section class=wp-comments><h2>Comments</h2><p class=wp-comments-notice>Comments are from this blog's WordPress era and are preserved read-only.</p><article class=wp-comment><header><span class=wp-comment-author>Can Gencer</span>
|
||
<time class=wp-comment-date>April 26, 2019</time></header><div class=wp-comment-body><p>After reading through Liftbridge docs and the replication protocol, I’m curious what was the reason you used Kafka’s in-sync replica protocol rather than the “raft log as message log” approach from NATS Streams?</p></div><div class=wp-comment-replies><article class=wp-comment><header><span class=wp-comment-author>Tyler Treat</span>
|
||
<time class=wp-comment-date>April 26, 2019</time></header><div class=wp-comment-body><p>NATS Streaming doesn’t actually use the Raft log as the message log, it uses it as a secondary replication/recovery log meaning it stores messages redundantly. The Kafka-based protocol means we can use the message log itself for replication so no redundancy and it’s faster. It also means we can “tune” the replication based on durability needs (replicate to all or subset of replicas, etc.). Lastly, the ISR approach allows balancing availability with durability. If brokers fail, the ISR can shrink and the cluster can maintain availability. With Raft, you cannot make this trade-off and must retain a quorum.</p></div></article></div></article></section></article></main><footer class=site-footer><div class="wrap site-footer-inner"><div class=footer-meta><span class=footer-copy>© 2026 Tyler Treat</span>
|
||
<span class="footer-sep footer-dot">·</span>
|
||
<span class=footer-links><a href=/feed/>rss</a>
|
||
<span class=footer-sep>·</span>
|
||
<a href=https://github.com/tylertreat target=_blank rel="noopener noreferrer me">github</a>
|
||
<span class=footer-sep>·</span>
|
||
<a href=https://www.linkedin.com/in/ttreat/ target=_blank rel="noopener noreferrer me">linkedin</a></span></div></div></footer><a href=#top id=top-link class="top-link hidden" aria-label="go to top" title="Go to Top (Alt + G)" accesskey=g><svg viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="feather feather-chevrons-up"><polyline points="17 11 12 6 7 11"/><polyline points="17 18 12 13 7 18"/></svg>
|
||
</a><script>let menu=document.getElementById("menu");if(menu){const e=localStorage.getItem("menu-scroll-position");e&&(menu.scrollLeft=parseInt(e,10)),menu.onscroll=function(){localStorage.setItem("menu-scroll-position",menu.scrollLeft)}}document.querySelectorAll('a[href^="#"]').forEach(e=>{e.addEventListener("click",function(e){e.preventDefault();var t=this.getAttribute("href").substr(1);window.matchMedia("(prefers-reduced-motion: reduce)").matches?document.querySelector(`[id='${decodeURIComponent(t)}']`).scrollIntoView():document.querySelector(`[id='${decodeURIComponent(t)}']`).scrollIntoView({behavior:"smooth"}),t==="top"?history.replaceState(null,null," "):history.pushState(null,null,`#${t}`)})})</script><script>var toplink=document.getElementById("top-link");window.onscroll=function(){const e=window.innerHeight;document.body.scrollTop>e||document.documentElement.scrollTop>e?toplink.classList.remove("hidden"):toplink.classList.add("hidden")}</script><script>document.getElementById("theme-toggle").addEventListener("click",()=>{const e=document.querySelector("html");e.dataset.theme==="dark"?(e.dataset.theme="light",localStorage.setItem("pref-theme","light")):(e.dataset.theme="dark",localStorage.setItem("pref-theme","dark"))})</script><script>document.querySelectorAll("pre > code").forEach(e=>{const n=e.parentNode.parentNode,t=document.createElement("button");t.classList.add("copy-code"),t.innerHTML="copy";function s(){t.innerHTML="copied!",setTimeout(()=>{t.innerHTML="copy"},2e3)}t.addEventListener("click",t=>{if("clipboard"in navigator){navigator.clipboard.writeText(e.textContent),s();return}const n=document.createRange();n.selectNodeContents(e);const o=window.getSelection();o.removeAllRanges(),o.addRange(n);try{document.execCommand("copy"),s()}catch{}o.removeRange(n)}),n.classList.contains("highlight")?n.appendChild(t):n.parentNode.firstChild==n||(e.parentNode.parentNode.parentNode.parentNode.parentNode.nodeName=="TABLE"?e.parentNode.parentNode.parentNode.parentNode.parentNode.appendChild(t):e.parentNode.appendChild(t))})</script></body></html> |