988 lines
22 KiB
HTML
988 lines
22 KiB
HTML
<!DOCTYPE html>
|
||
<html xmlns="http://www.w3.org/1999/xhtml" xml:lang="en" lang="en-us">
|
||
<head>
|
||
<meta http-equiv="content-type" content="text/html; charset=utf-8" />
|
||
<meta name="viewport" content="width=device-width, initial-scale=1.0">
|
||
<title>MemoryDB: Speed, Durability, and Composition. - Marc's Blog</title>
|
||
<meta name="author" content="Marc Brooker" />
|
||
|
||
<!-- Homepage CSS -->
|
||
<link rel="stylesheet" href="/blog/css/screen.css" type="text/css" media="screen, projection" />
|
||
<link rel="stylesheet" href="/blog/css/syntax.css" type="text/css" media="screen, projection" />
|
||
|
||
</head>
|
||
<body>
|
||
|
||
<div class="site">
|
||
<div class="title">
|
||
<h1><a href="/blog/">Marc's Blog</a></h1>
|
||
</div>
|
||
|
||
<div class="about">
|
||
<h1>About Me</h1>
|
||
My name is Marc Brooker. I like to build things that work, and do cool stuff. I like building big things. I also dabble in machining, welding, cooking, and skiing.<br/><br/>
|
||
|
||
I am an engineer at Amazon Web Services (AWS) in Seattle, where I work on agentic AI, especially safety and policy for agentic AI. Before that, I worked on EC2, EBS, databases, serverless, and serverless databases.<br/>
|
||
|
||
All opinions are my own.
|
||
<h1>Links</h1>
|
||
<a href="https://brooker.co.za/blog/publications.html">My Publications and Videos</a><br/>
|
||
<a rel="me" href="https://fediscience.org/@marcbrooker">@marcbrooker on Mastodon</a>
|
||
<a href="https://twitter.com/MarcJBrooker">@MarcJBrooker on Twitter</a>
|
||
|
||
<br/><br/><br/>
|
||
<a href="https://brooker.co.za/blog/2026/06/18/my-blog-and-ai.html">Is this blog written by AI?</a>
|
||
</div>
|
||
|
||
|
||
<div id="post">
|
||
<h1 id="memorydb-speed-durability-and-composition">MemoryDB: Speed, Durability, and Composition.</h1>
|
||
|
||
<p class="meta">Blocks are fun.</p>
|
||
|
||
<p>Earlier this week, my colleagues Yacine Taleb, Kevin McGehee, Nan Yan, Shawn Wang, Stefan Mueller, and Allen Samuels published <a href="https://www.amazon.science/publications/amazon-memorydb-a-fast-and-durable-memory-first-cloud-database">Amazon MemoryDB: A fast and durable memory-first cloud database</a><sup><a href="#foot1">1</a></sup>. I’m excited about this paper, both because its a very cool system, and because it gives us an opportunity to talk about the power of composition in distributed systems, and about the power of distributed systems in general.</p>
|
||
|
||
<p>But first, what is <a href="https://aws.amazon.com/memorydb/">MemoryDB</a>?</p>
|
||
|
||
<blockquote>
|
||
<p>Amazon MemoryDB for Redis is a durable database with microsecond reads, low single-digit millisecond writes, scalability, and enterprise security. MemoryDB delivers 99.99% availability and near instantaneous recovery without any data loss.</p>
|
||
</blockquote>
|
||
|
||
<p>or, from the paper:</p>
|
||
|
||
<blockquote>
|
||
<p>We describe how, using this architecture, we are able to remain fully compatible with Redis, while providing single-digit millisecond write and microsecond-scale read latencies, strong consistency, and high availability.</p>
|
||
</blockquote>
|
||
|
||
<p>This is remarkable: MemoryDB keeps compatibility with an existing in-memory data store, adds multi-AZ (multi-datacenter) durability, adds high availability, and adds strong consistency on failover, while still improving read performance and with fairly little cost to write performance.</p>
|
||
|
||
<p>How does that work? As usual, there’s a lot of important details, but the basic idea is composing the in-memory store (Redis) with our existing fast, multi-AZ transaction journal<sup><a href="#foot2">2</a></sup> service (a system we use in many places inside AWS).</p>
|
||
|
||
<p><img src="/blog/images/memorydb_arch.png" alt="" /></p>
|
||
|
||
<p><strong>Composition</strong></p>
|
||
|
||
<p>What’s particularly interesting about this architecture is that the journal service doesn’t only provide durability. Instead, it provides multiple different benefits:</p>
|
||
|
||
<ul>
|
||
<li>durability (by synchronously replicating writes onto storage in multiple AZs),</li>
|
||
<li>fan-out (by being the replication stream replicas can consume),</li>
|
||
<li>leader election (by having strongly-consistent <em>fencing</em> APIs that make it easy to ensure there’s a single leader per shard),</li>
|
||
<li>safety during reconfiguration and resharding (using those same <em>fencing</em> APIs), and</li>
|
||
<li>the ability to move bulk data tasks like snapshotting off the latency-sensitive leader boxes.</li>
|
||
</ul>
|
||
|
||
<p>Moving these concerns into the Journal greatly simplifies the job of the leader, and minimized the amount that the team needed to modify Redis. In turn, this makes keeping up with new Redis (or <a href="https://github.com/valkey-io/valkey">Valkey</a>) developments much easier. From an organizational perspective, it also allows the team that owns Journal to really focus on performance, safety, and cost of the journal without having to worry about the complexities of offering a rich API to customers. Each investment in performance means better performance for a number of AWS services, and similarly for cost, and investments in formal methods, and so on. As an engineer, and engineering leader, I’m always on the look out for these leverage opportunities.</p>
|
||
|
||
<p>Of course, the idea of breaking systems down into pieces separated by interfaces isn’t new. It’s one of the most venerable ideas in computing. Still, this is a great reminder of how composition can reduce overall system complexity. The journal service is a relatively (conceptually) simple system, presenting a simple API. But, by carefully designing that API with affordances like fencing (more on that later), it can remove the need to have complex things like consensus implementations inside its clients (see Section 2.2 of the paper for a great discussion of some of this complexity).</p>
|
||
|
||
<p>As <a href="https://www.aboutamazon.com/news/company-news/amazon-ceo-andy-jassy-2023-letter-to-shareholders">Andy Jassy says</a>:</p>
|
||
|
||
<blockquote>
|
||
<p>Primitives, done well, rapidly accelerate builders’ ability to innovate.</p>
|
||
</blockquote>
|
||
|
||
<p><strong>Distribution</strong></p>
|
||
|
||
<p>It’s well known that distributed systems can improve durability (by making multiple copies of data on multiple machines), availability (by allowing another machine to take over if one fails), integrity (by allowing machines with potentially corrupted data to drop out), and scalability (by allowing multiple machines to do work). However, it’s often incorrectly assumed that this value comes at the cost of complexity and performance. This paper is a great reminder that assumption is not true.</p>
|
||
|
||
<p>Let’s zoom in on one aspect of performance: consistent latency while taking snapshots. MemoryDB moves snapshotting off the database nodes themselves, and into a separate service dedicated to maintaining snapshots.</p>
|
||
|
||
<p><img src="/blog/images/memorydb_snapshotting.png" alt="" /></p>
|
||
|
||
<p>This snapshotting service doesn’t really care about latency (at least not the sub-millisecond read latencies that the database nodes worry about). It’s a throughput-optimized operation, where we want to stream tons of data in the most throughput-efficient way possible. By moving it into a different service, we get to avoid having throughput-optimized and latency-optimized processes running at the same time (with all the cache and scheduling issues that come with that). The system also gets to avoid some implementation complexities of snapshotting in-place. From the paper, talking about the on-box <em>BGSave</em> snapshotting mechanism:</p>
|
||
|
||
<blockquote>
|
||
<p>However, there is a spike on P100 latency reaching up to 67 milliseconds for request response times. This is due to the fork system call which clones the entire memory page table. Based on our internal measurement, this process takes about 12ms per GB of memory.</p>
|
||
</blockquote>
|
||
|
||
<p>and things get worse if there’s not enough memory for the copy-on-write (CoW) copy of the data:</p>
|
||
|
||
<blockquote>
|
||
<p>Once the instance exhausts all the DRAM capacity and starts to use swap to page out memory pages, the latency increases and the throughput drops significantly. […] The tail latency increases over a second
|
||
and throughput drops close to 0…</p>
|
||
</blockquote>
|
||
|
||
<p>the conclusion being that to avoid this effect database nodes need to keep extra RAM around (up to double) just to support snapshotting. An expensive proposition in an in-memory database! Moving snapshotting off-box avoids this cost: memory can be shared between snapshotting tasks, which <a href="https://brooker.co.za/blog/2023/03/23/economics.html">significantly improves utilization of that memory</a>.</p>
|
||
|
||
<p><img src="/blog/images/memorydb_fig7.png" alt="" /></p>
|
||
|
||
<p>The upshot is that, in MemoryDB with off-box snapshotting, performance impact is entirely avoided. Distributed systems can optimize components for the kind of work they do, and can use multi-tenancy to reduce costs.</p>
|
||
|
||
<p><strong>Conclusion</strong></p>
|
||
|
||
<p>Go check out the <a href="https://www.amazon.science/publications/amazon-memorydb-a-fast-and-durable-memory-first-cloud-database">MemoryDB team’s paper</a>. There’s a lot of great content in there, including a smart way to ensure consistency between the leader and the log, a description of the formal methods the team used, and operational concerns around version upgrades. This is what real system building looks like.</p>
|
||
|
||
<p><strong>Bonus: Fencing</strong></p>
|
||
|
||
<p>Above, I mentioned how <em>fencing</em> in the journal service API is something that makes the service much more powerful, and a better building block for real-world distributed systems. To understand what I mean, let’s consider a journal service (a simple ordered stream service) with the following API:</p>
|
||
|
||
<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>write(payload) -> seq
|
||
read() -> (payload, seq) or none
|
||
</code></pre></div></div>
|
||
|
||
<p>You call <em>write</em>, and when the <em>payload</em> has been durably replicated it returns a totally-ordered sequence number for your write. That’s powerful enough, but in most systems would require an additional leader election to ensure that the writes being sent make some logical sense.</p>
|
||
|
||
<p>We can extend the API to avoid this case:</p>
|
||
|
||
<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>write(payload, last_seq) -> seq
|
||
read() -> (payload, seq) or none
|
||
</code></pre></div></div>
|
||
|
||
<p>In this version, writers can ensure they are up-to-date with all reads before doing a write, and make sure they’re not racing with another writer. That’s sufficient to ensure consistency, but isn’t particularly efficient (multiple leaders could always be racing), and doesn’t allow a leader to offer consistent operations that don’t call <em>write</em> (like the in-memory reads the MemoryDB offers). It also makes pipelining difficult (unless the leader can make an assumption about the density of the sequences). An alternative design is to offer a <a href="https://dl.acm.org/doi/10.1145/74851.74870">lease</a> service:</p>
|
||
|
||
<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>try_take_lease() -> (uuid, deadline)
|
||
renew_lease(uuid) -> deadline
|
||
write(payload) -> seq
|
||
read() -> (payload, seq) or none
|
||
</code></pre></div></div>
|
||
|
||
<p>A leader who believes they hold the lease (i.e. their current time is comfortably before the <em>deadline</em>) can assume they’re the only leader, and can go back to using the original write API. If they end up taking the lease, they poll <em>read</em> until the stream is empty, and then can take over as the single leader. This approach offers strong consistency, but only if leaders absolutely obey their contract that they don’t call <em>write</em> unless they hold the lease.</p>
|
||
|
||
<p>That’s easily said, but harder to do. For example, consider the following code:</p>
|
||
|
||
<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>if current_time < deadline:
|
||
<gc or scheduler pause>
|
||
write(payload)
|
||
</code></pre></div></div>
|
||
|
||
<p>Those kinds of pauses are really hard to avoid. They come from GC, from page faults, from swapping, from memory pressure, from scheduling, from background tasks, and many many other things. And that’s not even to mention the possible causes of error on <em>local_time</em>. We can avoid this issue with a small adaptation to our API:</p>
|
||
|
||
<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>try_take_lease() -> (uuid, deadline)
|
||
renew_lease(uuid) -> deadline
|
||
write(payload, lease_holder_uuid) -> seq
|
||
read() -> (payload, seq) or none
|
||
</code></pre></div></div>
|
||
|
||
<p>If <em>write</em> can enforce that the writer is the current lease holder, we can avoid all of these races while still allowing writers to pipeline things as deeply as they like. This still-simple API provides an extremely powerful building block for building systems like MemoryDB.</p>
|
||
|
||
<p>Finally, we may not need to compose our lease service with the journal service, because we may want to use other leader election mechanisms. We can avoid that by offering a relatively simple compare-and-set in the journal API:</p>
|
||
|
||
<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>set_leader_uuid(new_uuid, old_uuid) -> old_uuid
|
||
write(payload, leader_uuid) -> seq
|
||
read() -> (payload, seq) or none
|
||
</code></pre></div></div>
|
||
|
||
<p>Now we have a super powerful composable primitive that can offer both safety to writers, and liveness if the leader election system is reasonably well behaved.</p>
|
||
|
||
<p><em>Footnotes</em></p>
|
||
|
||
<ol>
|
||
<li><a name="foot1"></a> To appear at SIGMOD’24.</li>
|
||
<li><a name="foot2"></a> The paper calls it a <em>log</em> service, which is technically correct, but a term I tend to avoid because its easily confused with logging in the observability sense.</li>
|
||
</ol>
|
||
|
||
</div>
|
||
|
||
<div id="related">
|
||
« <a href="/blog">Back to the blog index</a><br>
|
||
<br>
|
||
<!-- Similar Posts -->
|
||
<h4>Similar Posts</h4>
|
||
<ul class="posts">
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
<li><span>12 Jul 2022</span> » <a href="/blog/2022/07/12/dynamodb.html">The DynamoDB paper</a></li>
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
<li><span>03 Dec 2024</span> » <a href="/blog/2024/12/03/aurora-dsql.html">DSQL Vignette: Aurora DSQL, and A Personal Story</a></li>
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
<li><span>15 Aug 2025</span> » <a href="/blog/2025/08/15/dynamo-dynamodb-dsql.html">Dynamo, DynamoDB, and Aurora DSQL</a></li>
|
||
|
||
|
||
|
||
|
||
|
||
</ul>
|
||
|
||
<!-- Dissimilar Posts -->
|
||
|
||
<h4>Something Completely Different</h4>
|
||
<ul class="posts">
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
|
||
<li><span>28 Jul 2020</span> » <a href="/blog/2020/07/28/fish.html">A Story About a Fish</a></li>
|
||
|
||
|
||
|
||
|
||
</ul>
|
||
|
||
</div>
|
||
|
||
|
||
<div class="footer">
|
||
<div class="contact">
|
||
<p>
|
||
Marc Brooker<br />
|
||
The opinions on this site are my own. They do not necessarily represent those of my employer.<br />
|
||
marcbrooker@gmail.com
|
||
</p>
|
||
<p>
|
||
<a href="https://brooker.co.za/blog/rss.xml"><img src="/blog/images/feed-icon-14x14.png" /> RSS</a>
|
||
<a href="https://brooker.co.za/blog/atom.xml"><img src="/blog/images/feed-icon-14x14.png" /> Atom</a>
|
||
</p>
|
||
</div>
|
||
<div class="license">
|
||
<!-- <a rel="license" href="http://creativecommons.org/licenses/by/4.0/"><img alt="Creative Commons License" style="border-width:0" src="https://i.creativecommons.org/l/by/4.0/88x31.png" /></a><br /> -->
|
||
This work is licensed under a <a rel="license" href="http://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</a>.
|
||
</div>
|
||
</div>
|
||
</div>
|
||
|
||
</body>
|
||
</html>
|