Files
nexus/sreweekly/articles/458/08-snapshot-isolation-vs-serializability.html
2026-09-12 17:23:01 +08:00

628 lines
20 KiB
HTML
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
<!DOCTYPE html>
<html xmlns="http://www.w3.org/1999/xhtml" xml:lang="en" lang="en-us">
<head>
<meta http-equiv="content-type" content="text/html; charset=utf-8" />
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<title>Snapshot Isolation vs Serializability - Marc's Blog</title>
<meta name="author" content="Marc Brooker" />
<!-- Homepage CSS -->
<link rel="stylesheet" href="/blog/css/screen.css" type="text/css" media="screen, projection" />
<link rel="stylesheet" href="/blog/css/syntax.css" type="text/css" media="screen, projection" />
</head>
<body>
<div class="site">
<div class="title">
<h1><a href="/blog/">Marc's Blog</a></h1>
</div>
<div class="about">
<h1>About Me</h1>
My name is Marc Brooker. I like to build things that work, and do cool stuff. I like building big things. I also dabble in machining, welding, cooking, and skiing.<br/><br/>
I am an engineer at Amazon Web Services (AWS) in Seattle, where I work on agentic AI, especially safety and policy for agentic AI. Before that, I worked on EC2, EBS, databases, serverless, and serverless databases.<br/>
All opinions are my own.
<h1>Links</h1>
<a href="https://brooker.co.za/blog/publications.html">My Publications and Videos</a><br/>
<a rel="me" href="https://fediscience.org/@marcbrooker">@marcbrooker on Mastodon</a>
<a href="https://twitter.com/MarcJBrooker">@MarcJBrooker on Twitter</a>
<br/><br/><br/>
<a href="https://brooker.co.za/blog/2026/06/18/my-blog-and-ai.html">Is this blog written by AI?</a>
</div>
<div id="post">
<h1 id="snapshot-isolation-vs-serializability">Snapshot Isolation vs Serializability</h1>
<script>
MathJax = {
tex: {inlineMath: [['$', '$'], ['\\(', '\\)']]}
};
</script>
<script id="MathJax-script" async="" src="https://cdn.jsdelivr.net/npm/mathjax@3/es5/tex-mml-chtml.js"></script>
<p class="meta">Getting into some fundamentals.</p>
<p>In my <a href="https://www.youtube.com/watch?v=huGmR_mi5dQ">re:Invent talk on the internals of Aurora DSQL</a> I mentioned that I think snapshot isolation is a sweet spot in the database isolation spectrum for most kinds of applications. Today, I want to dive in a little deeper into why I think that, and some of the trade-offs of going stronger and weaker.</p>
<p>This post is going to be a little deeper than the last few. If you’re not deeply familiar with SQL’s isolation levels, I recommend checking out Crooks et al’s <a href="https://dl.acm.org/doi/10.1145/3087801.3087802">Seeing is Believing: A Client-Centric Specification of Database Isolation</a>, Berenson et al’s <a href="https://www.microsoft.com/en-us/research/wp-content/uploads/2016/02/tr-95-51.pdf">A Critique of ANSI SQL Isolation Levels</a>, or Adya et al’s <a href="https://pmg.csail.mit.edu/papers/icde00.pdf">Generalized Isolation Level Definitions</a>.</p>
<p>Specifically, I’m going to talk about one very specific mental model of transaction isolation: read-write conflicts, and write-write conflicts.</p>
<p>Let’s start our journey with a transaction, <code class="language-plaintext highlighter-rouge">T1</code>:</p>
<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>BEGIN;
SELECT amnt FROM food WHERE id = 1 OR id = 2;
UPDATE food SET amnt = amnt - 1 WHERE id = 1;
COMMIT;
</code></pre></div></div>
<p>There are a few things to notice about this transaction, when run in a database system that offers interactive (i.e. back-and-forth with the client) transactions:</p>
<ul>
<li>It directly reads two rows from the database, the rows <code class="language-plaintext highlighter-rouge">1</code> and <code class="language-plaintext highlighter-rouge">2</code> from the table <code class="language-plaintext highlighter-rouge">food</code>.</li>
<li>It directly writes one row in the database, the rows <code class="language-plaintext highlighter-rouge">1</code> from the table food.</li>
<li>It starts and ends at different times: starting during <code class="language-plaintext highlighter-rouge">BEGIN</code> and ending during <code class="language-plaintext highlighter-rouge">COMMIT</code>.</li>
<li>Depending on the way that database is implemented, it also likely reads other data (e.g. the system catalog which tells it which tables exist), and may write other data (e.g. a secondary index on the <code class="language-plaintext highlighter-rouge">amnt</code> column of <code class="language-plaintext highlighter-rouge">food</code>).</li>
</ul>
<p>The forth point here is critical in real systems, but let’s ignore it for now and focus only on the first three<sup><a href="#foot2">2</a></sup>. Before we do that, let’s introduce a second transaction, the first one’s parallel universe clone, <code class="language-plaintext highlighter-rouge">T2</code>. We’ll also assume these are the only transactions running at this time.</p>
<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>BEGIN;
SELECT amnt FROM food WHERE id = 1 OR id = 2;
UPDATE food SET amnt = amnt - 1 WHERE id = 2;
COMMIT;
</code></pre></div></div>
<p>We’ll go a step further and explicitly interleave these transactions (as though they were happening at the same time on two different connections to the same database), taking a page from <a href="https://github.com/ept/hermitage/blob/master/postgres.md">Hermitage</a><sup><a href="#foot4">4</a></sup>.</p>
<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>BEGIN; -- T1, t1_start
BEGIN; -- T2, t2_start
SELECT amnt FROM food WHERE id = 1 OR id = 2; -- T1
SELECT amnt FROM food WHERE id = 1 OR id = 2; -- T2
UPDATE food SET amnt = amnt - 1 WHERE id = 1; -- T1
UPDATE food SET amnt = amnt - 1 WHERE id = 2; -- T2
COMMIT; -- T1, t1_commit
COMMIT; -- T2, t2_commit, WHAT HAPPENS HERE?
</code></pre></div></div>
<p>What should happen to the second commit?</p>
<p>Accepting that second commit is an example of <em>write skew</em>, and is generally a good minimal example of the difference between snapshot isolation (SI) and serializability.</p>
<p>But that’s well documented, and not what I’m focussing on. What I’m focussing on is how we might prevent write skew.</p>
<p><em>Read Sets and Write Sets</em></p>
<p>Let’s abstract <code class="language-plaintext highlighter-rouge">T1</code> and <code class="language-plaintext highlighter-rouge">T2</code> a step further: into read sets and write sets. We’ll simplify everything more here by pretending our database has a single table (<code class="language-plaintext highlighter-rouge">food</code>).</p>
<p><code class="language-plaintext highlighter-rouge">T1</code> then has the read set $R_1 = {1,2}$ and the write set $W_1 = {1}$</p>
<p><code class="language-plaintext highlighter-rouge">T2</code> then has the read set $R_2 = {1,2}$ and the write set $W_2 = {2}$</p>
<p>Under serializability, once <code class="language-plaintext highlighter-rouge">T1</code> is committed, <code class="language-plaintext highlighter-rouge">T2</code> can only commit if $R_2 \cap W_1 = \emptyset$ (or, more generally, the set of all writes accepted between <code class="language-plaintext highlighter-rouge">t2_start</code> and <code class="language-plaintext highlighter-rouge">t2_commit</code> does not intersect with <code class="language-plaintext highlighter-rouge">R_2</code><sup><a href="#foot1">1</a></sup>).</p>
<p>Under snapshot isolation, once <code class="language-plaintext highlighter-rouge">T1</code> is committed, <code class="language-plaintext highlighter-rouge">T2</code> can only commit if $W_2 \cap W_1 = \emptyset$ (or, more generally, the set of all writes accepted between <code class="language-plaintext highlighter-rouge">t2_start</code> and <code class="language-plaintext highlighter-rouge">t2_commit</code> does not intersect with <code class="language-plaintext highlighter-rouge">W_2</code>).</p>
<p>Ok?</p>
<p>Notice how similar those two statements are: one about $R_2 \cap W_1$ and one about $W_2 \cap W_1$. That’s the only real difference in the rules.</p>
<p>But it’s a <em>crucial</em> difference.</p>
<p>It’s a crucial difference because of one of the cool and powerful things that SQL databases make easy: <code class="language-plaintext highlighter-rouge">SELECT</code>s. You can grow a transaction’s write set with <code class="language-plaintext highlighter-rouge">UPDATE</code> and <code class="language-plaintext highlighter-rouge">INSERT</code> and friends, but most OLTP applications don’t tend to. You can grow a transaction’s read set with any <code class="language-plaintext highlighter-rouge">SELECT</code>, and many applications do that. If you don’t believe me, go look at the ratio between predicate (i.e. not exact PK equality) <code class="language-plaintext highlighter-rouge">SELECT</code>s in your code base versus predicate <code class="language-plaintext highlighter-rouge">UPDATE</code>s and <code class="language-plaintext highlighter-rouge">INSERT</code>s. If the ratios are even close, you’re a little unusual.</p>
<p><em>Concurrency versus Isolation</em></p>
<p>This is where we enter a world of trade-offs<sup><a href="#foot3">3</a></sup>: avoiding SI’s write skew requires the database to abort (or, sometimes, just block) transactions based on what they <em>read</em>. That’s true for OCC, for PCC, or for nearly any scheme you can devise. To get good performance (and scalability) in the face of concurrency, applications using serializable isolation need to be extremely careful about <em>reads</em>.</p>
<p>The trouble is that isolation primarily exists to simplify the lives of application programmers, and make it so they don’t have to deal with concurrency. SQL-like isolation models do that quite effectively, and are (in my opinion) one of the best ideas in the history of computing. But as we move up the isolation levels to serializability, we start pushing more complexity onto the application programmer by forcing them to worry more about concurrency from a performance and throughput perspective.</p>
<p>This is the cause of my belief that snapshot isolation, combined with strong consistency, is the right default for most applications and most teams of application programmers: it provides a useful minimum in the sum of worries about anomalies and performance.</p>
<p>Fundamentally, it does that by observing that write sets are smaller than read sets, for the majority of OLTP applications (often MUCH smaller).</p>
<p><em>How Low Can We Go? How Low Should We Go?</em></p>
<p>That raises the question of whether its worth going even lower in the isolation spectrum. I don’t have as crisp a general answer, but in Aurora DSQL’s case the answer is, mostly <em>no</em>. Reducing the isolation level from SI to <code class="language-plaintext highlighter-rouge">read committed</code> does not save any distributed coordination, because of the use of physical clocks and MVCC to provide a consistent read snapshot to transactions without coordination. It does save some local coordination on each storage replica, and the implementation of multiversioning in storage, but our measurements indicate those are minimal.</p>
<p>If you’d like to learn more about that architecture, here’s me talking about it at re:Invent 2024:</p>
<iframe width="560" height="315" src="https://www.youtube-nocookie.com/embed/huGmR_mi5dQ?si=qXMxiImNZ5MGf9Co" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen=""></iframe>
<p><em>Optimistic versus Pessimistic Concurrency Control</em></p>
<p>What I said above about SI is, from my perspective, equally true in both OCC and PCC designs, when combined with MVCC to form read snapshots and timestamps to choose a read snapshot. In both cases, the only coordination strictly required for SI is to look for write-write conflicts ($W_2 \cap W_1$) that occured between <code class="language-plaintext highlighter-rouge">t_start</code> and <code class="language-plaintext highlighter-rouge">t_commit</code>, and to choose a commit timestamp based on that detection.</p>
<p>The big advantage OCC has in avoiding coordination has to do with how those write-write conflicts are detected. In backward validation OCC, the only state that is needed to decide whether to commit transaction $T$ is from <em>already committed transactions</em>. This means that all state in the commit protocol is transient, and can be reconstructed from the log of committed transactions, without causing any false aborts. This is a significant benefit!</p>
<p>The other benefit of OCC, <a href="https://brooker.co.za/blog/2024/12/03/aurora-dsql.html">as I covered in my first post on DSQL</a> is that we can avoid all coordination during the process of executing a transaction, and only coordinate at <code class="language-plaintext highlighter-rouge">COMMIT</code> time. This is a huge advantage when coordination latency is considerable, like in the multi-region setting.</p>
<p>As with SI versus serializability, the story of trade-offs between optimistic and pessimistic approaches is a long one. But, for similar reasons, I think OCC (when combined with MVCC) is the right choice for most transactional workloads.</p>
<p><em>Footnotes</em></p>
<ol>
<li><a name="foot1"></a> This general framing is my OCC backward validation bias showing, but it extends to most pessimistic approaches too.</li>
<li><a name="foot2"></a> These special cases, like the catalog and <code class="language-plaintext highlighter-rouge">FOR UPDATE</code>, don’t fit nicely into the academic framework of SI, and most databases will need to handle them differently from regular key accesses.</li>
<li><a name="foot3"></a> You will find, typically online, folks who deny that these trade-offs exist. Those people are wrong, both theoretically and practically.</li>
<li><a name="foot4"></a> You can find Aurora DSQL’s <a href="https://github.com/marcbrooker/hermitage/blob/master/dsql.md">hermitage SQL test set here</a>. You may notice that it’s very similar to PostgreSQL’s <code class="language-plaintext highlighter-rouge">repeatable read</code> set, but non-blocking instead of blocking.</li>
</ol>
</div>
<div id="related">
&laquo; <a href="/blog">Back to the blog index</a><br>
<br>
<!-- Similar Posts -->
<h4>Similar Posts</h4>
<ul class="posts">
<li><span>05 Dec 2024</span> &raquo; <a href="/blog/2024/12/05/inside-dsql-writes.html">DSQL Vignette: Transactions and Durability</a></li>
<li><span>23 Jan 2024</span> &raquo; <a href="/blog/2024/01/23/big-deal.html">Pat's Big Deal, and Transaction Coordination</a></li>
<li><span>05 Feb 2025</span> &raquo; <a href="/blog/2025/02/05/feketes.html">What Fekete's Anomaly Can Teach Us About Isolation</a></li>
</ul>
<!-- Dissimilar Posts -->
<h4>Something Completely Different</h4>
<ul class="posts">
<li><span>21 Jan 2026</span> &raquo; <a href="/blog/2026/01/21/pass-k.html">Pass@k is Mostly Bunk</a></li>
</ul>
</div>
<div class="footer">
<div class="contact">
<p>
Marc Brooker<br />
The opinions on this site are my own. They do not necessarily represent those of my employer.<br />
marcbrooker@gmail.com
</p>
<p>
<a href="https://brooker.co.za/blog/rss.xml"><img src="/blog/images/feed-icon-14x14.png" /> RSS</a>
<a href="https://brooker.co.za/blog/atom.xml"><img src="/blog/images/feed-icon-14x14.png" /> Atom</a>
</p>
</div>
<div class="license">
<!-- <a rel="license" href="http://creativecommons.org/licenses/by/4.0/"><img alt="Creative Commons License" style="border-width:0" src="https://i.creativecommons.org/l/by/4.0/88x31.png" /></a><br /> -->
This work is licensed under a <a rel="license" href="http://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</a>.
</div>
</div>
</div>
</body>
</html>