29 lines
37 KiB
HTML
29 lines
37 KiB
HTML
<!doctype html><html lang=en dir=auto data-theme=auto><head><meta charset=utf-8><meta http-equiv=X-UA-Compatible content="IE=edge"><meta name=viewport content="width=device-width,initial-scale=1,shrink-to-fit=no"><meta name=robots content="index, follow"><title>Building a Distributed Log from Scratch, Part 1: Storage Mechanics | Brave New Geek</title><meta name=keywords content="architecture,building a distributed log from scratch,data storage,distributed log,distributed systems,kafka,message queues,message-oriented middleware,messaging,nats,nats streaming,performance,stream processing,zero-copy"><meta name=description content="The log is a totally-ordered, append-only data structure. It’s a powerful yet simple abstraction—a sequence of immutable events. It’s something that programmers have been using for a very long time, perhaps without even realizing it because it’s so simple. Whether it’s application logs, system logs, or access logs, logging is something every developer uses on a daily basis. Essentially, it’s a timestamp and an event, a when and a what, and typically appended to the end of a file. But when we generalize that pattern, we end up with something much more useful for a broad range of problems. It becomes more interesting when we look at the log not just as a system of record but a central piece in managing data and distributing it across the enterprise efficiently."><meta name=author content><link rel=canonical href=https://bravenewgeek.com/building-a-distributed-log-from-scratch-part-1-storage-mechanics/><link crossorigin=anonymous href=/assets/css/stylesheet.4861a452a4c13a9a1fbf2085400b74a7de96b1beeb94dee57b4273e5dffcf337.css integrity="sha256-SGGkUqTBOpofvyCFQAt0p96Wsb7rlN7le0Jz5d/88zc=" rel="preload stylesheet" as=style><link rel=icon href=https://bravenewgeek.com/favicon.ico><link rel=icon type=image/png sizes=16x16 href=https://bravenewgeek.com/favicon.ico><link rel=icon type=image/png sizes=32x32 href=https://bravenewgeek.com/favicon.ico><link rel=apple-touch-icon href=https://bravenewgeek.com/favicon.ico><link rel=mask-icon href=https://bravenewgeek.com/favicon.ico><meta name=theme-color content="#2e2e33"><meta name=msapplication-TileColor content="#2e2e33"><link rel=alternate hreflang=en href=https://bravenewgeek.com/building-a-distributed-log-from-scratch-part-1-storage-mechanics/><noscript><style>#theme-toggle,.top-link{display:none}</style><style>@media(prefers-color-scheme:dark){:root{--theme:rgb(29, 30, 32);--entry:rgb(46, 46, 51);--primary:rgb(218, 218, 219);--secondary:rgb(155, 156, 157);--tertiary:rgb(65, 66, 68);--content:rgb(196, 196, 197);--code-block-bg:rgb(46, 46, 51);--code-bg:rgb(55, 56, 62);--border:rgb(51, 51, 51);color-scheme:dark}.list{background:var(--theme)}.toc{background:var(--entry)}}</style></noscript><script>localStorage.getItem("pref-theme")==="dark"?document.querySelector("html").dataset.theme="dark":localStorage.getItem("pref-theme")==="light"?document.querySelector("html").dataset.theme="light":window.matchMedia("(prefers-color-scheme: dark)").matches?document.querySelector("html").dataset.theme="dark":document.querySelector("html").dataset.theme="light"</script><link rel=preconnect href=https://fonts.googleapis.com><link rel=preconnect href=https://fonts.gstatic.com crossorigin><link rel=stylesheet href="https://fonts.googleapis.com/css2?family=JetBrains+Mono:wght@400;500&family=Source+Serif+4:ital,opsz,wght@0,8..60,400;0,8..60,600;1,8..60,400&family=Space+Grotesk:wght@500;600;700&display=swap"><meta property="og:url" content="https://bravenewgeek.com/building-a-distributed-log-from-scratch-part-1-storage-mechanics/"><meta property="og:site_name" content="Brave New Geek"><meta property="og:title" content="Building a Distributed Log from Scratch, Part 1: Storage Mechanics"><meta property="og:description" content="The log is a totally-ordered, append-only data structure. It’s a powerful yet simple abstraction—a sequence of immutable events. It’s something that programmers have been using for a very long time, perhaps without even realizing it because it’s so simple. Whether it’s application logs, system logs, or access logs, logging is something every developer uses on a daily basis. Essentially, it’s a timestamp and an event, a when and a what, and typically appended to the end of a file. But when we generalize that pattern, we end up with something much more useful for a broad range of problems. It becomes more interesting when we look at the log not just as a system of record but a central piece in managing data and distributing it across the enterprise efficiently."><meta property="og:locale" content="en_us"><meta property="og:type" content="article"><meta property="article:section" content="posts"><meta property="article:published_time" content="2017-12-21T15:54:17-06:00"><meta property="article:modified_time" content="2018-02-23T16:09:17-06:00"><meta property="article:tag" content="Architecture"><meta property="article:tag" content="Building a Distributed Log From Scratch"><meta property="article:tag" content="Data Storage"><meta property="article:tag" content="Distributed Log"><meta property="article:tag" content="Distributed Systems"><meta property="article:tag" content="Kafka"><meta name=twitter:card content="summary"><meta name=twitter:title content="Building a Distributed Log from Scratch, Part 1: Storage Mechanics"><meta name=twitter:description content="The log is a totally-ordered, append-only data structure. It’s a powerful yet simple abstraction—a sequence of immutable events. It’s something that programmers have been using for a very long time, perhaps without even realizing it because it’s so simple. Whether it’s application logs, system logs, or access logs, logging is something every developer uses on a daily basis. Essentially, it’s a timestamp and an event, a when and a what, and typically appended to the end of a file. But when we generalize that pattern, we end up with something much more useful for a broad range of problems. It becomes more interesting when we look at the log not just as a system of record but a central piece in managing data and distributing it across the enterprise efficiently."><script type=application/ld+json>{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Posts","item":"https://bravenewgeek.com/posts/"},{"@type":"ListItem","position":2,"name":"Building a Distributed Log from Scratch, Part 1: Storage Mechanics","item":"https://bravenewgeek.com/building-a-distributed-log-from-scratch-part-1-storage-mechanics/"}]}</script><script type=application/ld+json>{"@context":"https://schema.org","@type":"BlogPosting","headline":"Building a Distributed Log from Scratch, Part 1: Storage Mechanics","name":"Building a Distributed Log from Scratch, Part 1: Storage Mechanics","description":"The log is a totally-ordered, append-only data structure. It’s a powerful yet simple abstraction—a sequence of immutable events. It’s something that programmers have been using for a very long time, perhaps without even realizing it because it’s so simple. Whether it’s application logs, system logs, or access logs, logging is something every developer uses on a daily basis. Essentially, it’s a timestamp and an event, a when and a what, and typically appended to the end of a file. But when we generalize that pattern, we end up with something much more useful for a broad range of problems. It becomes more interesting when we look at the log not just as a system of record but a central piece in managing data and distributing it across the enterprise efficiently.\n","keywords":["architecture","building a distributed log from scratch","data storage","distributed log","distributed systems","kafka","message queues","message-oriented middleware","messaging","nats","nats streaming","performance","stream processing","zero-copy"],"articleBody":"The log is a totally-ordered, append-only data structure. It’s a powerful yet simple abstraction—a sequence of immutable events. It’s something that programmers have been using for a very long time, perhaps without even realizing it because it’s so simple. Whether it’s application logs, system logs, or access logs, logging is something every developer uses on a daily basis. Essentially, it’s a timestamp and an event, a when and a what, and typically appended to the end of a file. But when we generalize that pattern, we end up with something much more useful for a broad range of problems. It becomes more interesting when we look at the log not just as a system of record but a central piece in managing data and distributing it across the enterprise efficiently.\nThere are a number of implementations of this idea: Apache Kafka, Amazon Kinesis, NATS Streaming, Tank, and Apache Pulsar to name a few. We can probably credit Kafka with popularizing the idea.\nI think there are at least three key priorities for the effectiveness of one of these types of systems: performance, high availability, and scalability. If it’s not fast enough, the data becomes decreasingly useful. If it’s not highly available, it means we can’t reliably get our data in or out. And if it’s not scalable, it won’t be able to meet the needs of many enterprises.\nWhen we apply the traditional pub/sub semantics to this idea of a log, it becomes a very useful abstraction that applies to a lot of different problems.\nIn this series, we’re not going to spend much time discussing why the log is useful. Jay Kreps has already done the legwork on that with The Log: What every software engineer should know about real-time data’s unifying abstraction. There’s even a book on it. Instead, we will focus on what it takes to build something like this using Kafka and NATS Streaming as case studies of sorts—Kafka because of its ubiquity, NATS Streaming because it’s something with which I have personal experience. We’ll look at a few core components like leader election, data replication, log persistence, and message delivery. Part one of this series starts with the storage mechanics. Along the way, we will also discuss some lessons learned while building NATS Streaming, which is a streaming data layer on top of the NATS messaging system. The intended outcome of this series is threefold: to learn a bit about the internals of a log abstraction, to learn how it can achieve the three goals described above, and to learn some applied distributed systems theory.\nWith that in mind, you will probably never need to build something like this yourself (nor should you), but it helps to know how it works. I also find that software engineering is all about pattern matching. Many types of problems look radically different but are surprisingly similar. Some of these ideas may apply to other things you come across. If nothing else, it’s just interesting.\nLet’s start by looking at data storage since this is a critical part of the log and dictates some other aspects of it. Before we dive into that, though, let’s highlight some first principles we’ll use as a starting point for driving our design.\nAs we know, the log is an ordered, immutable sequence of messages. Messages are atomic, meaning they can’t be broken up. A message is either in the log or not, all or nothing. Although we only ever add messages to the log and never remove them (as with a message queue), the log has a notion of message retention based on some policies, which allows us to control how the log is truncated. This is a practical requirement since otherwise the log will grow endlessly. These policies might be based on time, number of messages, number of bytes, etc.\nThe log can be played back from any arbitrary position. With position, we normally refer to a logical message timestamp rather than a physical wall-clock time, such as an offset into the log. The log is stored on disk, and sequential disk access is actually relatively fast. The graphic below taken from the ACM Queue article The Pathologies of Big Data helps bear this out (this is helpfully pointed out by Kafka’s documentation).\nThat said, modern OS page caches mean that sequential access often avoids going to disk altogether. This is because the kernel keeps cached pages in otherwise unused portions of RAM. This means both reads and writes go to the in-memory page cache instead of disk. With Kafka, for example, we can verify this quite easily by running a simple test that writes some data and reads it back and looking at disk IO using iostat. After running such a test, you will likely see something resembling the following, which shows the number of blocks read and written is exactly zero.\navg-cpu: %user %nice %system %iowait %steal %idle 13.53 0.00 11.28 0.00 0.00 75.19 Device: tps Blk_read/s Blk_wrtn/s Blk_read Blk_wrtn xvda 0.00 0.00 0.00 0 0 With the above in mind, our log starts to look an awful lot like an actual logging file, but instead of timestamps and log messages, we have offsets and opaque data messages. We simply add new messages to the end of the file with a monotonically increasing offset.\nHowever, there are some problems with this approach. Namely, the file is going to get very, very large. Recall that we need to support a few different access patterns: looking up messages by offset and also truncating the log using a variety of different retention policies. Since the log is ordered, a lookup is simply a binary search for the offset, but this is expensive with a large log file. Similarly, aging out data by retention policy is harder.\nTo account for this, we break up the log file into chunks. In Kafka, these are called segments. In NATS Streaming, they are called slices. Each segment is a new file. At a given time, there is a single active segment, which is the segment messages are written to. Once the segment is full (based on some configuration), a new one is created and becomes active.\nSegments are defined by their base offset, i.e. the offset of the first message stored in the segment. In Kafka, the files are also named with this offset. This allows us to quickly locate the segment in which a given message is contained by doing a binary search.\nAlongside each segment file is an index file that maps message offsets to their respective positions in the log segment. In Kafka, the index uses 4 bytes for storing an offset relative to the base offset and 4 bytes for storing the log position. Using a relative offset is more efficient because it means we can avoid storing the actual offset as an int64. In NATS Streaming, the timestamp is also stored to do time-based lookups.\nIdeally, the data written to the log segment is written in protocol format. That is, what gets written to disk is exactly what gets sent over the wire. This allows for zero-copy reads. Let’s take a look at how this otherwise works.\nWhen you read messages from the log, the kernel will attempt to pull the data from the page cache. If it’s not there, it will be read from disk. The data is copied from disk to page cache, which all happens in kernel space. Next, the data is copied into the application (i.e. user space). This all happens with the read system call. Now the application writes the data out to a socket using send, which is going to copy it back into kernel space to a socket buffer before it’s copied one last time to the NIC. All in all, we have four copies (including one from page cache) and two system calls.\nHowever, if the data is already in wire format, we can bypass user space entirely using the sendfile system call, which will copy the data directly from the page cache to the NIC buffer—two copies (including one from page cache) and one system call. This turns out to be an important optimization, especially in garbage-collected languages since we’re bringing less data into application memory. Zero-copy also reduces CPU cycles and memory bandwidth.\nNATS Streaming does not currently make use of zero-copy for a number of reasons, some of which we will get into later in the series. In fact, the NATS Streaming storage layer is actually pluggable in that it can be backed by any number of mediums which implement the storage interface. Out of the box it includes the file-backed storage described above, in-memory, and SQL-backed.\nThere are a few other optimizations to make here such as message batching and compression, but we’ll leave those as an exercise for the reader.\nIn part two of this series, we will discuss how to make this log fault tolerant by diving into data-replication techniques.\n","wordCount":"1493","inLanguage":"en","datePublished":"2017-12-21T15:54:17-06:00","dateModified":"2018-02-23T16:09:17-06:00","mainEntityOfPage":{"@type":"WebPage","@id":"https://bravenewgeek.com/building-a-distributed-log-from-scratch-part-1-storage-mechanics/"},"publisher":{"@type":"Organization","name":"Brave New Geek","logo":{"@type":"ImageObject","url":"https://bravenewgeek.com/favicon.ico"}}}</script></head><body id=top><header class=site-header><div class="wrap site-header-inner"><a class=wordmark href=https://bravenewgeek.com/ accesskey=h title="Brave New Geek (Alt + H)"><span class=wordmark-name>Brave New Geek</span>
|
||
<span class=wordmark-tag>Introspections of a software engineer</span></a><nav class=site-nav aria-label=Primary><a href=https://bravenewgeek.com/archive/>Archive</a>
|
||
<a href=https://bravenewgeek.com/tags/>Tags</a>
|
||
<a href=https://bravenewgeek.com/about-me/>About</a>
|
||
<button id=theme-toggle class=theme-toggle accesskey=t title="Toggle theme (Alt + T)" aria-label="Toggle light/dark theme">
|
||
<svg class="moon" width="18" height="18" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M21 12.79A9 9 0 1111.21 3 7 7 0 0021 12.79z"/></svg>
|
||
<svg class="sun" width="18" height="18" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><circle cx="12" cy="12" r="5"/><line x1="12" y1="1" x2="12" y2="3"/><line x1="12" y1="21" x2="12" y2="23"/><line x1="4.22" y1="4.22" x2="5.64" y2="5.64"/><line x1="18.36" y1="18.36" x2="19.78" y2="19.78"/><line x1="1" y1="12" x2="3" y2="12"/><line x1="21" y1="12" x2="23" y2="12"/><line x1="4.22" y1="19.78" x2="5.64" y2="18.36"/><line x1="18.36" y1="5.64" x2="19.78" y2="4.22"/></svg></button></nav></div></header><main class=main><article class="post wrap"><div class=post-return><a href=https://bravenewgeek.com/><span class=pager-arrow>←</span> the log</a></div><header class=post-header><div class=post-meta><span class=post-offset>#69</span><time datetime=2017-12-21>2017-12-21</time><span>8 min read</span>
|
||
<span class=post-cats><a href=https://bravenewgeek.com/category/distributed-systems-2/>Distributed Systems</a><a href=https://bravenewgeek.com/category/messaging/>Messaging</a><a href=https://bravenewgeek.com/category/software-architecture/>Software Architecture</a><a href=https://bravenewgeek.com/category/software-engineering/>Software Engineering</a></span></div><h1 class=post-title>Building a Distributed Log from Scratch, Part 1: Storage Mechanics</h1></header><div class="post-content md-content"><p>The log is a totally-ordered, append-only data structure. It’s a powerful yet simple abstraction—a sequence of immutable events. It’s something that programmers have been using for a very long time, perhaps without even realizing it because it’s so simple. Whether it’s application logs, system logs, or access logs, logging is something every developer uses on a daily basis. Essentially, it’s a timestamp and an event, a <em>when</em> and a <em>what</em>, and typically appended to the end of a file. But when we generalize that pattern, we end up with something much more useful for a broad range of problems. It becomes more interesting when we look at the log not just as a system of record but a central piece in managing data and distributing it across the enterprise efficiently.</p><p><a href=/wp-content/uploads/2017/12/log.png><img loading=lazy src=/wp-content/uploads/2017/12/log.png></a></p><p>There are a number of implementations of this idea: <a href=https://kafka.apache.org/>Apache Kafka</a>, <a href=https://aws.amazon.com/kinesis/data-streams/>Amazon Kinesis</a>, <a href=https://github.com/nats-io/nats-streaming-server>NATS Streaming</a>, <a href=https://github.com/phaistos-networks/TANK>Tank</a>, and <a href=https://pulsar.apache.org/>Apache Pulsar</a> to name a few. We can probably credit Kafka with popularizing the idea.</p><p>I think there are at least three key priorities for the effectiveness of one of these types of systems: performance, high availability, and scalability. If it’s not fast enough, the data becomes decreasingly useful. If it’s not highly available, it means we can’t reliably get our data in or out. And if it’s not scalable, it won’t be able to meet the needs of many enterprises.</p><p>When we apply the traditional pub/sub semantics to this idea of a log, it becomes a very useful abstraction that applies to a lot of different problems.</p><p><a href=/wp-content/uploads/2017/12/log_use_cases.png><img loading=lazy src=/wp-content/uploads/2017/12/log_use_cases.png></a></p><p>In this series, we’re not going to spend much time discussing <em>why</em> the log is useful. Jay Kreps has already done the legwork on that with <a href=https://engineering.linkedin.com/distributed-systems/log-what-every-software-engineer-should-know-about-real-time-datas-unifying><em>The Log: What every software engineer should know about real-time data’s unifying abstraction</em></a>. There’s even a <a href=https://www.amazon.com/Heart-Logs-Stream-Processing-Integration/dp/1491909382>book</a> on it. Instead, we will focus on what it takes to <em>build</em> something like this using Kafka and NATS Streaming as case studies of sorts—Kafka because of its ubiquity, NATS Streaming because it’s something with which I have personal experience. We’ll look at a few core components like leader election, data replication, log persistence, and message delivery. Part one of this series starts with the storage mechanics. Along the way, we will also discuss some lessons learned while building NATS Streaming, which is a streaming data layer on top of the <a href=https://nats.io/>NATS</a> messaging system. The intended outcome of this series is threefold: to learn a bit about the internals of a log abstraction, to learn how it can achieve the three goals described above, and to learn some applied distributed systems theory.</p><p>With that in mind, you will probably never need to build something like this yourself (nor should you), but it helps to know how it works. I also find that software engineering is all about pattern matching. Many types of problems look radically different but are surprisingly similar. Some of these ideas may apply to other things you come across. If nothing else, it’s just <em>interesting</em>.</p><p>Let’s start by looking at data storage since this is a critical part of the log and dictates some other aspects of it. Before we dive into that, though, let’s highlight some first principles we’ll use as a starting point for driving our design.</p><p>As we know, the log is an ordered, immutable sequence of messages. Messages are <em>atomic</em>, meaning they can’t be broken up. A message is either in the log or not, all or nothing. Although we only ever add messages to the log and never remove them (as with a message queue), the log has a notion of <em>message retention</em> based on some policies, which allows us to control how the log is truncated. This is a practical requirement since otherwise the log will grow endlessly. These policies might be based on time, number of messages, number of bytes, etc.</p><p>The log can be played back from any arbitrary position. With position, we normally refer to a logical message timestamp rather than a physical wall-clock time, such as an offset into the log. The log is stored on disk, and sequential disk access is actually relatively <em>fast</em>. The graphic below taken from the ACM Queue article <a href="http://queue.acm.org/detail.cfm?id=1563874"><em>The Pathologies of Big Data</em></a> helps bear this out (this is helpfully pointed out by Kafka’s <a href=https://kafka.apache.org/documentation/#design_filesystem>documentation</a>).</p><p><a href=/wp-content/uploads/2017/12/disk_access.png><img loading=lazy src=/wp-content/uploads/2017/12/disk_access.png></a></p><p>That said, modern OS page caches mean that sequential access often avoids going to disk altogether. This is because the kernel keeps cached pages in otherwise unused portions of RAM. This means both reads and writes go to the in-memory page cache instead of disk. With Kafka, for example, we can verify this quite easily by running a simple test that writes some data and reads it back and looking at disk IO using <em>iostat</em>. After running such a test, you will likely see something resembling the following, which shows the number of blocks read and written is exactly zero.</p><div class=highlight><pre tabindex=0 class=chroma><code class=language-fallback data-lang=fallback><span class=line><span class=cl>avg-cpu: %user %nice %system %iowait %steal %idle
|
||
</span></span><span class=line><span class=cl> 13.53 0.00 11.28 0.00 0.00 75.19
|
||
</span></span><span class=line><span class=cl>
|
||
</span></span><span class=line><span class=cl>Device: tps Blk_read/s Blk_wrtn/s Blk_read Blk_wrtn
|
||
</span></span><span class=line><span class=cl>xvda 0.00 0.00 0.00 0 0
|
||
</span></span></code></pre></div><p>With the above in mind, our log starts to look an awful lot like an actual logging file, but instead of timestamps and log messages, we have offsets and opaque data messages. We simply add new messages to the end of the file with a monotonically increasing offset.</p><p><a href=/wp-content/uploads/2017/12/log_file.png><img loading=lazy src=/wp-content/uploads/2017/12/log_file.png></a></p><p>However, there are some problems with this approach. Namely, the file is going to get very, very large. Recall that we need to support a few different access patterns: looking up messages by offset and also truncating the log using a variety of different retention policies. Since the log is ordered, a lookup is simply a binary search for the offset, but this is expensive with a large log file. Similarly, aging out data by retention policy is harder.</p><p>To account for this, we break up the log file into chunks. In Kafka, these are called segments. In NATS Streaming, they are called slices. Each segment is a new file. At a given time, there is a single active segment, which is the segment messages are written to. Once the segment is full (based on some configuration), a new one is created and becomes active.</p><p>Segments are defined by their base offset, i.e. the offset of the first message stored in the segment. In Kafka, the files are also named with this offset. This allows us to quickly locate the segment in which a given message is contained by doing a binary search.</p><p><a href=/wp-content/uploads/2017/12/log_segments.png><img loading=lazy src=/wp-content/uploads/2017/12/log_segments.png></a></p><p>Alongside each segment file is an index file that maps message offsets to their respective positions in the log segment. In Kafka, the index uses 4 bytes for storing an offset relative to the base offset and 4 bytes for storing the log position. Using a relative offset is more efficient because it means we can avoid storing the actual offset as an int64. In NATS Streaming, the timestamp is also stored to do time-based lookups.</p><p><a href=/wp-content/uploads/2017/12/log_index.png><img loading=lazy src=/wp-content/uploads/2017/12/log_index.png></a></p><p>Ideally, the data written to the log segment is written in protocol format. That is, what gets written to disk is exactly what gets sent over the wire. This allows for zero-copy reads. Let’s take a look at how this otherwise works.</p><p>When you read messages from the log, the kernel will attempt to pull the data from the page cache. If it’s not there, it will be read from disk. The data is copied from disk to page cache, which all happens in kernel space. Next, the data is copied into the application (i.e. user space). This all happens with the <em>read</em> system call. Now the application writes the data out to a socket using <em>send</em>, which is going to copy it back into kernel space to a socket buffer before it’s copied <em>one last time</em> to the NIC. All in all, we have <em>four</em> copies (including one from page cache) and <em>two</em> system calls.</p><p><a href=/wp-content/uploads/2017/12/read.png><img loading=lazy src=/wp-content/uploads/2017/12/read.png></a></p><p>However, if the data is already in wire format, we can bypass user space entirely using the <em>sendfile</em> system call, which will copy the data directly from the page cache to the NIC buffer—<em>two</em> copies (including one from page cache) and <em>one</em> system call. This turns out to be an important optimization, especially in garbage-collected languages since we’re bringing less data into application memory. Zero-copy also reduces CPU cycles and memory bandwidth.</p><p><a href=/wp-content/uploads/2017/12/sendfile.png><img loading=lazy src=/wp-content/uploads/2017/12/sendfile.png></a></p><p>NATS Streaming does not currently make use of zero-copy for a number of reasons, some of which we will get into later in the series. In fact, the NATS Streaming storage layer is actually <em>pluggable</em> in that it can be backed by any number of mediums which implement the storage interface. Out of the box it includes the file-backed storage described above, in-memory, and SQL-backed.</p><p>There are a few other optimizations to make here such as message batching and compression, but we’ll leave those as an exercise for the reader.</p><p>In <a href=https://bravenewgeek.com/building-a-distributed-log-from-scratch-part-2-data-replication/>part two</a> of this series, we will discuss how to make this log fault tolerant by diving into data-replication techniques.</p></div><footer class=post-footer><ul class=post-tags><li><a href=https://bravenewgeek.com/tag/architecture/>Architecture</a></li><li><a href=https://bravenewgeek.com/tag/building-a-distributed-log-from-scratch/>Building a Distributed Log From Scratch</a></li><li><a href=https://bravenewgeek.com/tag/data-storage/>Data Storage</a></li><li><a href=https://bravenewgeek.com/tag/distributed-log/>Distributed Log</a></li><li><a href=https://bravenewgeek.com/tag/distributed-systems/>Distributed Systems</a></li><li><a href=https://bravenewgeek.com/tag/kafka/>Kafka</a></li><li><a href=https://bravenewgeek.com/tag/message-queues/>Message Queues</a></li><li><a href=https://bravenewgeek.com/tag/message-oriented-middleware/>Message-Oriented Middleware</a></li><li><a href=https://bravenewgeek.com/tag/messaging/>Messaging</a></li><li><a href=https://bravenewgeek.com/tag/nats/>Nats</a></li><li><a href=https://bravenewgeek.com/tag/nats-streaming/>Nats Streaming</a></li><li><a href=https://bravenewgeek.com/tag/performance/>Performance</a></li><li><a href=https://bravenewgeek.com/tag/stream-processing/>Stream Processing</a></li><li><a href=https://bravenewgeek.com/tag/zero-copy/>Zero-Copy</a></li></ul><nav class=post-nav aria-label="Adjacent posts"><a class=post-nav-link href=https://bravenewgeek.com/building-a-distributed-log-from-scratch-part-2-data-replication/><span class=post-nav-dir><span class=pager-arrow>←</span> newer</span>
|
||
<span class=post-nav-title>Building a Distributed Log from Scratch, Part 2: Data Replication</span>
|
||
</a><a class="post-nav-link post-nav-right" href=https://bravenewgeek.com/thrift-on-steroids-a-tale-of-scale-and-abstraction/><span class=post-nav-dir>older <span class=pager-arrow>→</span></span>
|
||
<span class=post-nav-title>Thrift on Steroids: A Tale of Scale and Abstraction</span></a></nav></footer><section class=wp-comments><h2>Comments</h2><p class=wp-comments-notice>Comments are from this blog's WordPress era and are preserved read-only.</p><article class=wp-comment><header><span class=wp-comment-author>Philip Enrico</span>
|
||
<time class=wp-comment-date>December 22, 2017</time></header><div class=wp-comment-body><p>Any free online course so i can learn more deeply</p></div></article><article class=wp-comment><header><span class=wp-comment-author>Ahmed</span>
|
||
<time class=wp-comment-date>December 25, 2017</time></header><div class=wp-comment-body><p>Great</p></div></article><article class=wp-comment><header><span class=wp-comment-author>Rohan</span>
|
||
<time class=wp-comment-date>August 5, 2018</time></header><div class=wp-comment-body><p>Hi Tyler,<br>Thanks for the wonderful series.</p><p>I have a question regarding the idea of splitting the log into segments:<br>Even if we had one single long log file and an index file accompanying it, couldn’t we refer the index file for an offset, that’d give us the byte position and seek to that position?</p></div><div class=wp-comment-replies><article class=wp-comment><header><span class=wp-comment-author>Rohan</span>
|
||
<time class=wp-comment-date>August 20, 2018</time></header><div class=wp-comment-body><p>Understood it myself on thinking more…:)<br>Doing a binary search that refers an index for it’s jumps would still involve logN seeks in the file which’d be expensive.<br>Hence we can’t use that to just “find the last message older than”</p><p>Thanks Tyler!</p></div></article></div></article><article class=wp-comment><header><span class=wp-comment-author>Dio</span>
|
||
<time class=wp-comment-date>April 6, 2022</time></header><div class=wp-comment-body><p>Thanks for this wonderful series of posts! I have some question on the performance of WAL. Apparently Kafka’s strength lies in the fact that the log is append only, therefore there is no disk seek hence being super fast.</p><p>However, my question is:<br>1. AFAIK, two files are written for a single , one is for the actual messages, the other is an index.<br>2. Besides, each broker server (one machine) can hosts many topic + partition, and IIRC Kafka keeps the logs for different separated.<br>#1 combined with #2, I don’t see how it optimizes for disk access, because on a single machine we are still writing to multiple files which now needs disk seek.</p><p>Any insight? Thanks!</p></div></article><article class=wp-comment><header><span class=wp-comment-author>iwa2no</span>
|
||
<time class=wp-comment-date>August 29, 2024</time></header><div class=wp-comment-body><p>Hi, I want to ask why access ordered log in large file is not succifient. AFAIK, if we want to find a log in 10^6 logs, we just need log2(10^6) which approximates 20 times, which I believe is a small number. Have I oversimplified anything ?</p></div></article></section></article></main><footer class=site-footer><div class="wrap site-footer-inner"><div class=footer-meta><span class=footer-copy>© 2026 Tyler Treat</span>
|
||
<span class="footer-sep footer-dot">·</span>
|
||
<span class=footer-links><a href=/feed/>rss</a>
|
||
<span class=footer-sep>·</span>
|
||
<a href=https://github.com/tylertreat target=_blank rel="noopener noreferrer me">github</a>
|
||
<span class=footer-sep>·</span>
|
||
<a href=https://www.linkedin.com/in/ttreat/ target=_blank rel="noopener noreferrer me">linkedin</a></span></div></div></footer><a href=#top id=top-link class="top-link hidden" aria-label="go to top" title="Go to Top (Alt + G)" accesskey=g><svg viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="feather feather-chevrons-up"><polyline points="17 11 12 6 7 11"/><polyline points="17 18 12 13 7 18"/></svg>
|
||
</a><script>let menu=document.getElementById("menu");if(menu){const e=localStorage.getItem("menu-scroll-position");e&&(menu.scrollLeft=parseInt(e,10)),menu.onscroll=function(){localStorage.setItem("menu-scroll-position",menu.scrollLeft)}}document.querySelectorAll('a[href^="#"]').forEach(e=>{e.addEventListener("click",function(e){e.preventDefault();var t=this.getAttribute("href").substr(1);window.matchMedia("(prefers-reduced-motion: reduce)").matches?document.querySelector(`[id='${decodeURIComponent(t)}']`).scrollIntoView():document.querySelector(`[id='${decodeURIComponent(t)}']`).scrollIntoView({behavior:"smooth"}),t==="top"?history.replaceState(null,null," "):history.pushState(null,null,`#${t}`)})})</script><script>var toplink=document.getElementById("top-link");window.onscroll=function(){const e=window.innerHeight;document.body.scrollTop>e||document.documentElement.scrollTop>e?toplink.classList.remove("hidden"):toplink.classList.add("hidden")}</script><script>document.getElementById("theme-toggle").addEventListener("click",()=>{const e=document.querySelector("html");e.dataset.theme==="dark"?(e.dataset.theme="light",localStorage.setItem("pref-theme","light")):(e.dataset.theme="dark",localStorage.setItem("pref-theme","dark"))})</script><script>document.querySelectorAll("pre > code").forEach(e=>{const n=e.parentNode.parentNode,t=document.createElement("button");t.classList.add("copy-code"),t.innerHTML="copy";function s(){t.innerHTML="copied!",setTimeout(()=>{t.innerHTML="copy"},2e3)}t.addEventListener("click",t=>{if("clipboard"in navigator){navigator.clipboard.writeText(e.textContent),s();return}const n=document.createRange();n.selectNodeContents(e);const o=window.getSelection();o.removeAllRanges(),o.addRange(n);try{document.execCommand("copy"),s()}catch{}o.removeRange(n)}),n.classList.contains("highlight")?n.appendChild(t):n.parentNode.firstChild==n||(e.parentNode.parentNode.parentNode.parentNode.parentNode.nodeName=="TABLE"?e.parentNode.parentNode.parentNode.parentNode.parentNode.appendChild(t):e.parentNode.appendChild(t))})</script></body></html> |