Files
nexus/sreweekly/articles/419/08-slo-formulas-implementation-in-promql-step-by-step.html
2026-09-12 17:23:01 +08:00

34 lines
20 KiB
HTML
Raw Permalink Blame History

This file contains invisible Unicode characters
This file contains invisible Unicode characters that are indistinguishable to humans but may be processed differently by a computer. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
<!DOCTYPE html> <html lang="en"> <head> <meta http-equiv="Content-Type" content="text/html; charset=UTF-8"> <meta charset="utf-8"> <meta name="viewport" content="width=device-width, initial-scale=1, shrink-to-fit=no"> <meta http-equiv="X-UA-Compatible" content="IE=edge"> <title>SLO formulas implementation in PromQL step by step | Michal Kazmierczak</title> <meta name="author" content="Michal Kazmierczak"> <meta name="description" content="The service level terminology provides a framework for quantifying the quality of a service's reliability. There are plenty of resources available on SLI, SLO, and SLA, most theorizing. This article proposes PromQL implementations for availability and latency SLOs."> <meta name="keywords" content="Prometheus, PromQL, SLO, SLI, availability, latency"> <link href="https://cdn.jsdelivr.net/npm/bootstrap@4.6.1/dist/css/bootstrap.min.css" rel="stylesheet" integrity="sha256-DF7Zhf293AJxJNTmh5zhoYYIMs2oXitRfBjY+9L//AY=" crossorigin="anonymous"> <link rel="stylesheet" href="https://cdn.jsdelivr.net/npm/mdbootstrap@4.20.0/css/mdb.min.css" integrity="sha256-jpjYvU3G3N6nrrBwXJoVEYI/0zw8htfFnhT9ljN3JJw=" crossorigin="anonymous"> <link rel="stylesheet" href="https://cdn.jsdelivr.net/npm/@fortawesome/fontawesome-free@5.15.4/css/all.min.css" integrity="sha256-mUZM63G8m73Mcidfrv5E+Y61y7a12O5mW4ezU3bxqW4=" crossorigin="anonymous"> <link rel="stylesheet" href="https://cdn.jsdelivr.net/npm/academicons@1.9.1/css/academicons.min.css" integrity="sha256-i1+4qU2G2860dGGIOJscdC30s9beBXjFfzjWLjBRsBg=" crossorigin="anonymous"> <link rel="stylesheet" type="text/css" href="https://fonts.googleapis.com/css?family=Roboto:300,400,500,700|Roboto+Slab:100,300,400,500,700|Material+Icons"> <link rel="stylesheet" href="https://cdn.jsdelivr.net/gh/jwarby/jekyll-pygments-themes@master/github.css" media="" id="highlight_theme_light"> <link rel="shortcut icon" href="data:image/svg+xml,&lt;svg%20xmlns=%22http://www.w3.org/2000/svg%22%20viewBox=%220%200%20100%20100%22&gt;&lt;text%20y=%22.9em%22%20font-size=%2290%22&gt;%F0%9F%91%8B&lt;/text&gt;&lt;/svg&gt;"> <link rel="stylesheet" href="/assets/css/main.css"> <link rel="canonical" href="https://mkaz.me/blog/2024/slo-formulas-implementation-in-promql-step-by-step/"> <link rel="stylesheet" href="https://cdn.jsdelivr.net/gh/jwarby/jekyll-pygments-themes@master/native.css" media="none" id="highlight_theme_dark"> <script src="/assets/js/theme.js"></script> <script src="/assets/js/dark_mode.js"></script> </head> <body class="fixed-top-nav "> <header> <nav id="navbar" class="navbar navbar-light navbar-expand-sm fixed-top"> <div class="container"> <a class="navbar-brand title font-weight-lighter" href="/"><span class="font-weight-bold">Michal </span>Kazmierczak</a> <button class="navbar-toggler collapsed ml-auto" type="button" data-toggle="collapse" data-target="#navbarNav" aria-controls="navbarNav" aria-expanded="false" aria-label="Toggle navigation"> <span class="sr-only">Toggle navigation</span> <span class="icon-bar top-bar"></span> <span class="icon-bar middle-bar"></span> <span class="icon-bar bottom-bar"></span> </button> <div class="collapse navbar-collapse text-right" id="navbarNav"> <ul class="navbar-nav ml-auto flex-nowrap"> <li class="nav-item "> <a class="nav-link" href="/">about</a> </li> <li class="nav-item active"> <a class="nav-link" href="/blog/">blog<span class="sr-only">(current)</span></a> </li> <li class="toggle-container"> <button id="light-toggle" title="Change theme"> <i class="fas fa-moon"></i> <i class="fas fa-sun"></i> </button> </li> </ul> </div> </div> </nav> <progress id="progress" value="0"> <div class="progress-container"> <span class="progress-bar"></span> </div> </progress> </header> <div class="container mt-5"> <div class="post"> <header class="post-header"> <h1 class="post-title">SLO formulas implementation in PromQL step by step</h1> <p class="post-meta"> March 25, 2024</p> <p class="post-tags"> <a href="/blog/2024"> <i class="fas fa-calendar fa-sm"></i> 2024 </a>   ·   <a href="/blog/tag/prometheus"> <i class="fas fa-hashtag fa-sm"></i> Prometheus</a>   <a href="/blog/tag/promql"> <i class="fas fa-hashtag fa-sm"></i> PromQL</a>   <a href="/blog/tag/slo"> <i class="fas fa-hashtag fa-sm"></i> SLO</a>   <a href="/blog/tag/sli"> <i class="fas fa-hashtag fa-sm"></i> SLI</a>   <a href="/blog/tag/availability"> <i class="fas fa-hashtag fa-sm"></i> availability</a>   <a href="/blog/tag/latency"> <i class="fas fa-hashtag fa-sm"></i> latency</a>     ·   <a href="/blog/category/monitoring"> <i class="fas fa-tag fa-sm"></i> monitoring</a>   </p> </header> <article class="post-content"> <h2 id="theory">Theory</h2> <p>In engineering, perfection isn’t optimal. Even though a service that is continuously up and always responds with the expected latency<sup><a href="#footnotes">[1]</a></sup> might sound desirable, it would require too expensive resources and overly conservative practices. To maintain the expectations and prioritize engineering resources, it’s common to define and publish <strong>Service Level Objectives (SLO)</strong>.</p> <p>An <strong>SLO defines the tolerable ratio of measurements meeting the target value to all measurements recorded within a specified time interval.</strong></p> <p>To break it down, I believe that a properly formulated <strong>SLO</strong> consists of four segments:</p> <ol> <li>Definition of the metric (what’s being measured).</li> <li>Target value.</li> <li>Anticipated ratio of good (meeting the target) to all measurements, typically expressed in percentages.</li> <li>Time interval over which the SLO is considered.</li> </ol> <p>A simple <strong>availability SLO</strong> could be:</p> <ul> <li> <code class="language-plaintext highlighter-rouge">a service is up 99.99% of the time over a week period</code> <ul> <li>the metric is <em>the service availability</em> (implicit)</li> <li>the target value is that it’s <em>up</em> </li> <li>the anticipated good to all ratio is <em>99.99%</em> </li> <li>the considered time interval is <em>week</em> </li> </ul> </li> </ul> <p>An example of <strong>latency SLO</strong> could be:</p> <ul> <li> <code class="language-plaintext highlighter-rouge">the latency of 99% of requests is less than or equal 250ms over a week period</code> <ul> <li>the metric is the <em>requests latency</em> </li> <li>the target value is <em>less than or equal 250ms</em> </li> <li>the anticipated good to all ratio is <em>99%</em> </li> <li>the considered time interval is <em>week</em> </li> </ul> </li> </ul> <p>Having an SLO defined, a service operator needs to deliver a calculation that tells the reality. The remainder of this article is a step by step guide on how to implement availability and latency SLO formulas with the Prometheus monitoring system.</p> <hr> <p><br></p> <h2 id="practice---slo-formulas-with-promql-step-by-step">Practice - SLO formulas with PromQL step by step</h2> <p>There are plenty of resources out there about SLO calculation. Moreover, there are off-the-shelf solutions for getting SLOs from metrics<sup><a href="#footnotes">[2]</a></sup>. <br> OK then, why add one more resource on the topic? I want to present real-life examples of both availability and latency SLOs, as they are more nuanced than they may initially appear. Also, I find it worthwhile sharing a detailed guide as it showcases uncommon uses of PromQL and demonstrates the language’s versatility.</p> <p><br></p> <h3 id="formula-for-availability-slo">Formula for availability SLO</h3> <p>Let’s look again at our SLO:</p> <div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>A service is up 99.99% of the time over a week period.
</code></pre></div></div> <p>To reason about a service’s availability we need a metric that indicates whether a service is up or down. Additionally, we need to sample this metric at regular time intervals. This allows us to divide the count of samples when the service reported as up by the count of all time intervals in the observed period. This gives us our SLO.</p> <p>Prometheus delivers the <code class="language-plaintext highlighter-rouge">up</code> metric (quite an accurate naming) for every scrape target. It’s a gauge which reports <code class="language-plaintext highlighter-rouge">1</code> when the scrape succeeded and <code class="language-plaintext highlighter-rouge">0</code> if the scrape failed. Given its simplicity, it’s a very good proxy of the service’s availability (however, it’s still prone to false negatives and false positives).</p> <p>Let’s consider the following plot of an <code class="language-plaintext highlighter-rouge">up</code> metric:</p> <p><a href="/assets/img/2024-03-25-a-promql-query-for-slo-calculation-explained/up_my_service.png"> <img src="/assets/img/2024-03-25-a-promql-query-for-slo-calculation-explained/up_my_service.png" alt="plot for the up metric" style="max-width: 100%"> </a></p> <p>This service was sampled every 15s. The time range is from 10:30:00 to 12:00:00 (1h30m inclusive - 361 samples). We can observe two periods when the service was down: <code class="language-plaintext highlighter-rouge">from 11:03:45 to 11:04:00</code> and <code class="language-plaintext highlighter-rouge">from 11:09:45 to 11:21:45</code>. This gives us <code class="language-plaintext highlighter-rouge">2 + 49 = 51</code> samples missing the target and <code class="language-plaintext highlighter-rouge">361 - 51 = 310</code> samples meeting the target. <br> Therefore the SLO is <code class="language-plaintext highlighter-rouge">(310 / 361) * 100 = 85.87%</code>. How to get this number with PromQL? Intuitively, we could count good samples and all samples (with <code class="language-plaintext highlighter-rouge">count_over_time</code>), and then calculate the ratio. However, given that the <code class="language-plaintext highlighter-rouge">up</code> metric is a gauge reporting either <code class="language-plaintext highlighter-rouge">0</code> or <code class="language-plaintext highlighter-rouge">1</code> we can do a trick and use the <code class="language-plaintext highlighter-rouge">avg_over_time</code> function:</p> <div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>avg_over_time(
up{job="my_service"}[90m:]
) * 100
=&gt; 85.87257617728531
</code></pre></div></div> <p>For production readiness, I recommend using <a target="_blank" href="https://prometheus.io/docs/prometheus/latest/querying/api/#instant-queries" rel="external nofollow noopener">instant query</a> and adjust the range. Also, see the footnotes if you suffer from periodic metric absence<sup><a href="#footnotes">[3]</a></sup>.</p> <p><br></p> <h3 id="formula-for-latency-slo">Formula for latency SLO</h3> <p>Let’s recity our latency SLO:</p> <div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>The latency of 99% of requests is less than or equal 250ms over a week period.
</code></pre></div></div> <p>For an accurate calculation, we need the count of requests that took less than or 250ms and the count of all requests. It turns out that it is a rare luxury, though. It actually means that you either have:</p> <ul> <li>the service instrumented upfront with a counter that increments when a request takes less than or 250ms, along with another counter that increments for every request;</li> <li>a histogram metric with one of the buckets set to <code class="language-plaintext highlighter-rouge">0.25</code>, which is the same as the first point as a histogram is essentially a set of counters.</li> </ul> <p>Assuming you have the histogram metric with the <code class="language-plaintext highlighter-rouge">0.25</code> bucket, the query is:</p> <div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>sum(
increase(http_request_latency_seconds_bucket{job="my_service",le="0.25"}[7d])
) * 100 /
sum(
increase(http_request_latency_seconds_bucket{job="my_service",le="+Inf"}[7d])
)
=&gt; 99.958703283089
</code></pre></div></div> <p>Ok, but what to do when such counters aren’t available?</p> <p><br></p> <h3 id="formula-for-latency-slo-with-percentiles">Formula for latency SLO with percentiles</h3> <p>Another approach to latency SLOs is setting expectations regarding percentiles, which is a common real-life example. It gives the flexibility in testing different target values without the need of re-instrumentation. However, I suggest to refrain from relying solely on percentiles in SLOs for several reasons:</p> <ul> <li>it’s an aggregate view, so it may hide certain characteristics of the service; a ratio of two counters offers way more straightforward understanding;</li> <li>combining percentiles from different services can be tricky; for example, averaging percentiles is possible only for services having the same latency distribution (almost impossible) or by calculating a weighted average (also almost impossible);</li> <li>the SLO gets more complicated, making it difficult to reason about — especially for non-SRE people who are also users of the SLO.</li> </ul> <p>Nonetheless, an example of such an SLO is:</p> <div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>99th percentile of requests latency is lower than 400ms 99% of the time over a week period.
</code></pre></div></div> <p>It can be implemented in PromQL in three steps:</p> <ol> <li>Calculate the quantile: <div class="language-plaintext highlighter-rouge"> <div class="highlight"><pre class="highlight"><code>histogram_quantile(
0.99,
sum by (le) (rate(http_server_request_duration_seconds_bucket[1m]))
)
</code></pre></div> </div> <p><a href="/assets/img/2024-03-25-a-promql-query-for-slo-calculation-explained/99th_quantile.png"> <img src="/assets/img/2024-03-25-a-promql-query-for-slo-calculation-explained/99th_quantile.png" alt="plot for the up metric" style="max-width: 100%"> </a></p> </li> <li>Perform a binary quantization to get a vector of <code class="language-plaintext highlighter-rouge">1</code> and <code class="language-plaintext highlighter-rouge">0</code> values corresponding to when the target value is met and not met. It’s a perfect input for our <code class="language-plaintext highlighter-rouge">avg_over_time</code> function. This is achieved with the <code class="language-plaintext highlighter-rouge">bool</code> modifier (<code class="language-plaintext highlighter-rouge">&lt;bool 0.4</code>). <div class="language-plaintext highlighter-rouge"> <div class="highlight"><pre class="highlight"><code>histogram_quantile(
0.99,
sum by (le) (rate(http_server_request_duration_seconds_bucket[1m]))
) &lt;bool 0.4
</code></pre></div> </div> <p><a href="/assets/img/2024-03-25-a-promql-query-for-slo-calculation-explained/99th_quantile_with_binary_quantization.png"> <img src="/assets/img/2024-03-25-a-promql-query-for-slo-calculation-explained/99th_quantile_with_binary_quantization.png" alt="plot for the up metric" style="max-width: 100%"> </a></p> </li> <li>Calculate the percentage: <div class="language-plaintext highlighter-rouge"> <div class="highlight"><pre class="highlight"><code>avg_over_time(
(histogram_quantile(
0.99,
sum by (le) (rate(http_server_request_duration_seconds_bucket[1m]))
) &lt;bool 0.4)[7d:]
) * 100
=&gt; 96.66666666666661
</code></pre></div> </div> </li> </ol> <hr> <p>When you finish prototyping with any of the above formulas, it’s a good practice to define a recording rule for the main metric for an efficient evaluation.</p> <hr> <h2 id="footnotes">Footnotes</h2> <ol> <li>In this article, ‘latency’ refers to the time it takes for the service to generate a response. I want to clarify it because in some resources it’s used as the time that a request is waiting to be handled.</li> <li>On this note, I highly recommend getting familiar with <a target="_blank" href="https://github.com/pyrra-dev/pyrra/" rel="external nofollow noopener">Pyrra</a>.</li> <li>It’s a common issue that the DB misses certain samples when the monitoring is hosted on the same server as the service having problems. While rethinking the architecture would be ideal, a quick workaround is to reduce all labels from the <code class="language-plaintext highlighter-rouge">up</code> metric and <em>fill the gap</em> with the scalar <code class="language-plaintext highlighter-rouge">vector(0)</code>. Beware, the consequence is that the absence of the metric is seen as the service being down. <div class="language-plaintext highlighter-rouge"> <div class="highlight"><pre class="highlight"><code>avg_over_time(
(sum(up{job="my_service"}) or vector(0))[90m:]
) * 100
</code></pre></div> </div> </li> </ol> </article> <hr class="mt-5"> <h4 class="text-3xl font-semibold mb-4 mt-12">Get in touch</h4> <p class="mb-2">If for whatever reason you consider it worthwhile to start a conversation with me, please check out the <a href="/">about</a> page to find the available means of communication.</p> </div> </div> <script src="https://cdn.jsdelivr.net/npm/jquery@3.6.0/dist/jquery.min.js" integrity="sha256-/xUj+3OJU5yExlq6GSYGSHk7tPXikynS7ogEvDej/m4=" crossorigin="anonymous"></script> <script src="https://cdn.jsdelivr.net/npm/bootstrap@4.6.1/dist/js/bootstrap.bundle.min.js" integrity="sha256-fgLAgv7fyCGopR/gBNq2iW3ZKIdqIcyshnUULC4vex8=" crossorigin="anonymous"></script> <script src="https://cdn.jsdelivr.net/npm/mdbootstrap@4.20.0/js/mdb.min.js" integrity="sha256-NdbiivsvWt7VYCt6hYNT3h/th9vSTL4EDWeGs5SN3DA=" crossorigin="anonymous"></script> <script defer src="https://cdn.jsdelivr.net/npm/masonry-layout@4.2.2/dist/masonry.pkgd.min.js" integrity="sha256-Nn1q/fx0H7SNLZMQ5Hw5JLaTRZp0yILA/FRexe19VdI=" crossorigin="anonymous"></script> <script defer src="https://cdn.jsdelivr.net/npm/imagesloaded@4/imagesloaded.pkgd.min.js"></script> <script defer src="/assets/js/masonry.js" type="text/javascript"></script> <script defer src="https://cdn.jsdelivr.net/npm/medium-zoom@1.0.8/dist/medium-zoom.min.js" integrity="sha256-7PhEpEWEW0XXQ0k6kQrPKwuoIomz8R8IYyuU1Qew4P8=" crossorigin="anonymous"></script> <script defer src="/assets/js/zoom.js"></script> <script defer src="/assets/js/common.js"></script> <script defer src="/assets/js/copy_code.js" type="text/javascript"></script> <script async src="https://d1bxh8uas1mnw7.cloudfront.net/assets/embed.js"></script> <script async src="https://badge.dimensions.ai/badge.js"></script> <script type="text/javascript">window.MathJax={tex:{tags:"ams"}};</script> <script defer type="text/javascript" id="MathJax-script" src="https://cdn.jsdelivr.net/npm/mathjax@3.2.0/es5/tex-mml-chtml.js"></script> <script defer src="https://cdnjs.cloudflare.com/polyfill/v3/polyfill.min.js?features=es6"></script> <script async src="https://www.googletagmanager.com/gtag/js?id=G-YV1BKNQ1FR"></script> <script>function gtag(){window.dataLayer.push(arguments)}window.dataLayer=window.dataLayer||[],gtag("js",new Date),gtag("config","G-YV1BKNQ1FR");</script> <script type="text/javascript">function progressBarSetup(){"max"in document.createElement("progress")?(initializeProgressElement(),$(document).on("scroll",function(){progressBar.attr({value:getCurrentScrollPosition()})}),$(window).on("resize",initializeProgressElement)):(resizeProgressBar(),$(document).on("scroll",resizeProgressBar),$(window).on("resize",resizeProgressBar))}function getCurrentScrollPosition(){return $(window).scrollTop()}function initializeProgressElement(){let e=$("#navbar").outerHeight(!0);$("body").css({"padding-top":e}),$("progress-container").css({"padding-top":e}),progressBar.css({top:e}),progressBar.attr({max:getDistanceToScroll(),value:getCurrentScrollPosition()})}function getDistanceToScroll(){return $(document).height()-$(window).height()}function resizeProgressBar(){progressBar.css({width:getWidthPercentage()+"%"})}function getWidthPercentage(){return getCurrentScrollPosition()/getDistanceToScroll()*100}const progressBar=$("#progress");window.onload=function(){setTimeout(progressBarSetup,50)};</script> <script type="module" src="https://static.cloudflareinsights.com/beacon.min.js/v31edd6df95cf4e85bb4c19e7a9bdbcba1788362987495" integrity="sha512-iIg7k2xntmwu6/uSb5tpc/hySgZc4eoL31yB29W6tJFo2akwjPWcEqnCEdJvGexCL0KEQwVYv5BlowfhVz26hg==" data-cf-beacon='{"version":"2024.11.0","token":"98c8107b014f47fbbb9501c7490525fd","r":1,"spa":2}' crossorigin="anonymous"></script>
</body> </html>