1064 lines
65 KiB
HTML
1064 lines
65 KiB
HTML
<!DOCTYPE html><html lang="en">
|
||
<style>
|
||
@import url('https://fonts.googleapis.com/css?family=Nunito');
|
||
@import url('https://fonts.googleapis.com/css?family=Space+Mono');
|
||
</style>
|
||
|
||
<head><!-- Global site tag (gtag.js) - Google Analytics -->
|
||
<script async src="https://www.googletagmanager.com/gtag/js?id=UA-64733189-1"></script>
|
||
<script>
|
||
window.dataLayer = window.dataLayer || [];
|
||
function gtag(){dataLayer.push(arguments);}
|
||
gtag('js', new Date());
|
||
|
||
gtag('config', 'UA-64733189-1');
|
||
|
||
</script><meta charset="utf-8">
|
||
<meta http-equiv="X-UA-Compatible" content="IE=edge">
|
||
<meta name="viewport" content="width=device-width, initial-scale=1, user-scalable=no"><title>ervin's blog</title>
|
||
|
||
<meta name="description" content="This is an attempt to break down the concept of SLO alerting as much as possible.Step-by-step, each concept will be illustrated and occasionally animated.My ...">
|
||
<link rel="canonical" href="https://ervinbarta.com/2021/10/19/slo-alerting-for-mortals/"><link rel="alternate" type="application/rss+xml" title="ervin's blog" href="/feed.xml">
|
||
<!-- begin favicon --><link rel="apple-touch-icon" sizes="180x180" href="/assets/apple-touch-icon.png?v=2"><link rel="icon" type="image/png" sizes="32x32" href="/assets/favicon-32x32.png?v=2"><link rel="icon" type="image/png" sizes="16x16" href="/assets/favicon-16x16.png?v=2"><link rel="manifest" href="/assets/site.webmanifest"><link rel="mask-icon" href="/assets/safari-pinned-tab.svg" color="#fc4d50"><link rel="shortcut icon" href="/assets/favicon.ico">
|
||
|
||
<meta name="msapplication-TileColor" content="#ffc40d"><meta name="msapplication-config" content="/assets/browserconfig.xml">
|
||
|
||
<meta name="theme-color" content="#ffffff">
|
||
<!-- end favicon --><link rel="stylesheet" href="/assets/css/main.css"><link rel="stylesheet" href="https://use.fontawesome.com/releases/v5.0.13/css/all.css" >
|
||
<!-- Begin Jekyll SEO tag v2.7.1 -->
|
||
<title>SLO Alerting for Mortals | ervin’s blog</title>
|
||
<meta name="generator" content="Jekyll v3.9.0" />
|
||
<meta property="og:title" content="SLO Alerting for Mortals" />
|
||
<meta name="author" content="Ervin Barta" />
|
||
<meta property="og:locale" content="en_US" />
|
||
<meta name="description" content="This is an attempt to break down the concept of SLO alerting as much as possible. Step-by-step, each concept will be illustrated and occasionally animated. My hope is that the short material presented here is intuitive enough for it to stick, and to serve as a base for further research. I will try to demistify concepts such as burn rate, error budget, and multi-window alerts. SLOs and SLIs Service level objective (SLO). It represents how reliably the service is delivering “value” to its users. Service level indicator (SLI). A measurement of a specific service metric. We’re using SLIs and math to define an SLO. A 99.9% SLO per month means if 0.1% of requests fail, that’s acceptable, and it won’t raise any fuss. We’re ok with the fact that 1 in every 1K requests will fail. An SLI can be the error rate of the incoming requests. The 0.1% wiggle room is the limit above which we don’t want to go - the error budget. Converting time Most of the math involved here is about converting various time units (i.e 1 month to 720 hours) and checking the ratio between them (i.e. 1h is 0.14% of a month). Then using those ratios with existing SLIs to come to an SLO condition. Time wise, a 0.1% error rate for a month means a 43 minute complete downtime. Burn rate When our budget of 43 minutes a month starts burning, we want to know about it fairly quickly. The green line represents the border between good and evil: if the error rate is exactly 0.1% throughout the month, we’re still fine, but barely. The burn rate is 1 in this case. As soon as the line starts to tilt left (into the danger zone), the error rate is higher than the allowed 0.1% and in turn, and the burn rate also increases - the system eats the error budget faster than it should. Defining the first alert As a start, we define the following alert condition, which is identical to the SLO. Keeping the graph above in mind, this translates to “if the green line starts tilting left: alert!” To reiterate on our timeline, the error budget spans out across the whole month. The alert we defined operates in an hour long time window. It follows that the 1 hour time window is of course not the whole month, but only 0.14% percent of it. At the error rate of 0.1% (our SLO), 0.14% of the budget is consumed during 1 hour. Not really worth to wake up someone over it. Another issue is that the alert will be active for almost an hour. (we will revisit this a bit later) To improve on this, instead of the 0.14% budget burn in an hour, the alert should have a higher threshold and aim at a 2% burn. We need to multiply our threshold with some number to reach this new target. Dividing the target value with the current one gives us the multiplier: 2%/0.14% = 14.3 This magic multiplier is actually the burn rate (or “tilting of the green line to the left” as we defined it a couple of lines above). Simulating an alert Here’s the scenario: 10 requests per 5 minutes is our traffic (constant) 10% error rate for 10 minutes (1 request out of 10 will fail) A snapshot of the state is taken every 5 minutes (scrape_interval in Prometheus terms). In the real world, you will probably have snapshots every 30 seconds or every minute. We’re using 5 minutes here for easier calculation, and to be able to draw square error rate lines, instead of sloping ones. An error rate of 10% in 2 subsequent snapshots (0m-5m, 5m-10m) is enough to trigger the alert. What immediately stands out here is the long running alert. It will be active for 55 minutes, even tough we’re not constantly in an erronous state. The http_error_rate[1h] metric considers the samples in the last 1 hour, and this time window is shifted with each snapshot. The snapshots are happening at the markers on the animation, every 5 minutes (remember, this is how the scenario was defined). As long as both erronous snapshots are inside the window, the alert will be active. When one of them leaves, the error rate drops to 0.8% (as per the calculation above) and the alert stops. This is why the alert is firing for 55 minutes instead of 1 hour. In a more realistic scenario, with a 30s second snapshot interval, it will be active for 1 hour. Multi-window alerts One way to combat the long running alert is to introduce another, shorter time window. It will make sure to end the alert, not long after the error rate goes back to normal. A 5 minute time window with the same error rate as before does just that. A single bad request out of 10 is plenty to trigger the first, 5 minute condition. The condition with the 1 hour window sets off after the second subsequent snapshot with an elevated error rate, as it needs 2 erronous requests out of 120 to surpass the threshold, as we saw in the previous section. Both conditions have the same threshold of 1.4% (0.1% * 14.4); the difference is that the 5 minute one takes 10 samples into consideration, and the 1 hour one takes 120 samples. A bad request has naturally a bigger impact on the smaller sample size than on the bigger one - 1 in 10 vs. 1 in 120. The smaller window is more jittery, where the longer one is slugish, but as they meet at the middle, the result is almost the best of both worlds: we have reasonable sensitivity and decent reset time (i.e the alert stops when the coast is clear). The alert is active only when the snapshots with a high error rate are in both time windows - in our case this is true for 10 minutes. With the addition of the shorter time window, we made sure that alert is matching reality more closely, i.e. it reacts only when there’s an ongoing issue. These were the basics, the next level would be multi-level, multi-burn rate alert which are a bit out of the scope of this post, so refer to the Google SLO alerting documentation for more details." />
|
||
<meta property="og:description" content="This is an attempt to break down the concept of SLO alerting as much as possible. Step-by-step, each concept will be illustrated and occasionally animated. My hope is that the short material presented here is intuitive enough for it to stick, and to serve as a base for further research. I will try to demistify concepts such as burn rate, error budget, and multi-window alerts. SLOs and SLIs Service level objective (SLO). It represents how reliably the service is delivering “value” to its users. Service level indicator (SLI). A measurement of a specific service metric. We’re using SLIs and math to define an SLO. A 99.9% SLO per month means if 0.1% of requests fail, that’s acceptable, and it won’t raise any fuss. We’re ok with the fact that 1 in every 1K requests will fail. An SLI can be the error rate of the incoming requests. The 0.1% wiggle room is the limit above which we don’t want to go - the error budget. Converting time Most of the math involved here is about converting various time units (i.e 1 month to 720 hours) and checking the ratio between them (i.e. 1h is 0.14% of a month). Then using those ratios with existing SLIs to come to an SLO condition. Time wise, a 0.1% error rate for a month means a 43 minute complete downtime. Burn rate When our budget of 43 minutes a month starts burning, we want to know about it fairly quickly. The green line represents the border between good and evil: if the error rate is exactly 0.1% throughout the month, we’re still fine, but barely. The burn rate is 1 in this case. As soon as the line starts to tilt left (into the danger zone), the error rate is higher than the allowed 0.1% and in turn, and the burn rate also increases - the system eats the error budget faster than it should. Defining the first alert As a start, we define the following alert condition, which is identical to the SLO. Keeping the graph above in mind, this translates to “if the green line starts tilting left: alert!” To reiterate on our timeline, the error budget spans out across the whole month. The alert we defined operates in an hour long time window. It follows that the 1 hour time window is of course not the whole month, but only 0.14% percent of it. At the error rate of 0.1% (our SLO), 0.14% of the budget is consumed during 1 hour. Not really worth to wake up someone over it. Another issue is that the alert will be active for almost an hour. (we will revisit this a bit later) To improve on this, instead of the 0.14% budget burn in an hour, the alert should have a higher threshold and aim at a 2% burn. We need to multiply our threshold with some number to reach this new target. Dividing the target value with the current one gives us the multiplier: 2%/0.14% = 14.3 This magic multiplier is actually the burn rate (or “tilting of the green line to the left” as we defined it a couple of lines above). Simulating an alert Here’s the scenario: 10 requests per 5 minutes is our traffic (constant) 10% error rate for 10 minutes (1 request out of 10 will fail) A snapshot of the state is taken every 5 minutes (scrape_interval in Prometheus terms). In the real world, you will probably have snapshots every 30 seconds or every minute. We’re using 5 minutes here for easier calculation, and to be able to draw square error rate lines, instead of sloping ones. An error rate of 10% in 2 subsequent snapshots (0m-5m, 5m-10m) is enough to trigger the alert. What immediately stands out here is the long running alert. It will be active for 55 minutes, even tough we’re not constantly in an erronous state. The http_error_rate[1h] metric considers the samples in the last 1 hour, and this time window is shifted with each snapshot. The snapshots are happening at the markers on the animation, every 5 minutes (remember, this is how the scenario was defined). As long as both erronous snapshots are inside the window, the alert will be active. When one of them leaves, the error rate drops to 0.8% (as per the calculation above) and the alert stops. This is why the alert is firing for 55 minutes instead of 1 hour. In a more realistic scenario, with a 30s second snapshot interval, it will be active for 1 hour. Multi-window alerts One way to combat the long running alert is to introduce another, shorter time window. It will make sure to end the alert, not long after the error rate goes back to normal. A 5 minute time window with the same error rate as before does just that. A single bad request out of 10 is plenty to trigger the first, 5 minute condition. The condition with the 1 hour window sets off after the second subsequent snapshot with an elevated error rate, as it needs 2 erronous requests out of 120 to surpass the threshold, as we saw in the previous section. Both conditions have the same threshold of 1.4% (0.1% * 14.4); the difference is that the 5 minute one takes 10 samples into consideration, and the 1 hour one takes 120 samples. A bad request has naturally a bigger impact on the smaller sample size than on the bigger one - 1 in 10 vs. 1 in 120. The smaller window is more jittery, where the longer one is slugish, but as they meet at the middle, the result is almost the best of both worlds: we have reasonable sensitivity and decent reset time (i.e the alert stops when the coast is clear). The alert is active only when the snapshots with a high error rate are in both time windows - in our case this is true for 10 minutes. With the addition of the shorter time window, we made sure that alert is matching reality more closely, i.e. it reacts only when there’s an ongoing issue. These were the basics, the next level would be multi-level, multi-burn rate alert which are a bit out of the scope of this post, so refer to the Google SLO alerting documentation for more details." />
|
||
<link rel="canonical" href="https://ervinbarta.com/2021/10/19/slo-alerting-for-mortals/" />
|
||
<meta property="og:url" content="https://ervinbarta.com/2021/10/19/slo-alerting-for-mortals/" />
|
||
<meta property="og:site_name" content="ervin’s blog" />
|
||
<meta property="og:type" content="article" />
|
||
<meta property="article:published_time" content="2021-10-19T00:00:00+00:00" />
|
||
<meta name="twitter:card" content="summary" />
|
||
<meta property="twitter:title" content="SLO Alerting for Mortals" />
|
||
<script type="application/ld+json">
|
||
{"author":{"@type":"Person","name":"Ervin Barta"},"datePublished":"2021-10-19T00:00:00+00:00","description":"This is an attempt to break down the concept of SLO alerting as much as possible. Step-by-step, each concept will be illustrated and occasionally animated. My hope is that the short material presented here is intuitive enough for it to stick, and to serve as a base for further research. I will try to demistify concepts such as burn rate, error budget, and multi-window alerts. SLOs and SLIs Service level objective (SLO). It represents how reliably the service is delivering “value” to its users. Service level indicator (SLI). A measurement of a specific service metric. We’re using SLIs and math to define an SLO. A 99.9% SLO per month means if 0.1% of requests fail, that’s acceptable, and it won’t raise any fuss. We’re ok with the fact that 1 in every 1K requests will fail. An SLI can be the error rate of the incoming requests. The 0.1% wiggle room is the limit above which we don’t want to go - the error budget. Converting time Most of the math involved here is about converting various time units (i.e 1 month to 720 hours) and checking the ratio between them (i.e. 1h is 0.14% of a month). Then using those ratios with existing SLIs to come to an SLO condition. Time wise, a 0.1% error rate for a month means a 43 minute complete downtime. Burn rate When our budget of 43 minutes a month starts burning, we want to know about it fairly quickly. The green line represents the border between good and evil: if the error rate is exactly 0.1% throughout the month, we’re still fine, but barely. The burn rate is 1 in this case. As soon as the line starts to tilt left (into the danger zone), the error rate is higher than the allowed 0.1% and in turn, and the burn rate also increases - the system eats the error budget faster than it should. Defining the first alert As a start, we define the following alert condition, which is identical to the SLO. Keeping the graph above in mind, this translates to “if the green line starts tilting left: alert!” To reiterate on our timeline, the error budget spans out across the whole month. The alert we defined operates in an hour long time window. It follows that the 1 hour time window is of course not the whole month, but only 0.14% percent of it. At the error rate of 0.1% (our SLO), 0.14% of the budget is consumed during 1 hour. Not really worth to wake up someone over it. Another issue is that the alert will be active for almost an hour. (we will revisit this a bit later) To improve on this, instead of the 0.14% budget burn in an hour, the alert should have a higher threshold and aim at a 2% burn. We need to multiply our threshold with some number to reach this new target. Dividing the target value with the current one gives us the multiplier: 2%/0.14% = 14.3 This magic multiplier is actually the burn rate (or “tilting of the green line to the left” as we defined it a couple of lines above). Simulating an alert Here’s the scenario: 10 requests per 5 minutes is our traffic (constant) 10% error rate for 10 minutes (1 request out of 10 will fail) A snapshot of the state is taken every 5 minutes (scrape_interval in Prometheus terms). In the real world, you will probably have snapshots every 30 seconds or every minute. We’re using 5 minutes here for easier calculation, and to be able to draw square error rate lines, instead of sloping ones. An error rate of 10% in 2 subsequent snapshots (0m-5m, 5m-10m) is enough to trigger the alert. What immediately stands out here is the long running alert. It will be active for 55 minutes, even tough we’re not constantly in an erronous state. The http_error_rate[1h] metric considers the samples in the last 1 hour, and this time window is shifted with each snapshot. The snapshots are happening at the markers on the animation, every 5 minutes (remember, this is how the scenario was defined). As long as both erronous snapshots are inside the window, the alert will be active. When one of them leaves, the error rate drops to 0.8% (as per the calculation above) and the alert stops. This is why the alert is firing for 55 minutes instead of 1 hour. In a more realistic scenario, with a 30s second snapshot interval, it will be active for 1 hour. Multi-window alerts One way to combat the long running alert is to introduce another, shorter time window. It will make sure to end the alert, not long after the error rate goes back to normal. A 5 minute time window with the same error rate as before does just that. A single bad request out of 10 is plenty to trigger the first, 5 minute condition. The condition with the 1 hour window sets off after the second subsequent snapshot with an elevated error rate, as it needs 2 erronous requests out of 120 to surpass the threshold, as we saw in the previous section. Both conditions have the same threshold of 1.4% (0.1% * 14.4); the difference is that the 5 minute one takes 10 samples into consideration, and the 1 hour one takes 120 samples. A bad request has naturally a bigger impact on the smaller sample size than on the bigger one - 1 in 10 vs. 1 in 120. The smaller window is more jittery, where the longer one is slugish, but as they meet at the middle, the result is almost the best of both worlds: we have reasonable sensitivity and decent reset time (i.e the alert stops when the coast is clear). The alert is active only when the snapshots with a high error rate are in both time windows - in our case this is true for 10 minutes. With the addition of the shorter time window, we made sure that alert is matching reality more closely, i.e. it reacts only when there’s an ongoing issue. These were the basics, the next level would be multi-level, multi-burn rate alert which are a bit out of the scope of this post, so refer to the Google SLO alerting documentation for more details.","mainEntityOfPage":{"@type":"WebPage","@id":"https://ervinbarta.com/2021/10/19/slo-alerting-for-mortals/"},"url":"https://ervinbarta.com/2021/10/19/slo-alerting-for-mortals/","@type":"BlogPosting","headline":"SLO Alerting for Mortals","dateModified":"2021-10-19T00:00:00+00:00","@context":"https://schema.org"}</script>
|
||
<!-- End Jekyll SEO tag -->
|
||
|
||
<script>(function() {
|
||
window.isArray = function(val) {
|
||
return Object.prototype.toString.call(val) === '[object Array]';
|
||
};
|
||
window.isString = function(val) {
|
||
return typeof val === 'string';
|
||
};
|
||
|
||
window.decodeUrl = function(str) {
|
||
return str ? decodeURIComponent(str.replace(/\+/g, '%20')) : '';
|
||
};
|
||
|
||
|
||
window.hasEvent = function(event) {
|
||
return 'on'.concat(event) in window.document;
|
||
};
|
||
|
||
window.isOverallScroller = function(node) {
|
||
return node === document.documentElement || node === document.body || node === window;
|
||
};
|
||
|
||
window.pageLoad = (function () {
|
||
var loaded = false, cbs = [];
|
||
window.addEventListener('load', function () {
|
||
var i, cb; loaded = true;
|
||
if (cbs.length > 0) {
|
||
for (i = 0; i < cbs.length; i++) {
|
||
cb = cbs[i]; cb();
|
||
}
|
||
}
|
||
});
|
||
return {
|
||
then: function(cb) {
|
||
cb && (loaded ? cb() : (cbs.push(cb)));
|
||
}
|
||
};
|
||
})();
|
||
})();(function() {
|
||
window.throttle = function(func, wait) {
|
||
var args, result, thisArg, timeoutId, lastCalled = 0;
|
||
|
||
function trailingCall() {
|
||
lastCalled = new Date;
|
||
timeoutId = null;
|
||
result = func.apply(thisArg, args);
|
||
}
|
||
return function() {
|
||
var now = new Date,
|
||
remaining = wait - (now - lastCalled);
|
||
|
||
args = arguments;
|
||
thisArg = this;
|
||
|
||
if (remaining <= 0) {
|
||
clearTimeout(timeoutId);
|
||
timeoutId = null;
|
||
lastCalled = now;
|
||
result = func.apply(thisArg, args);
|
||
} else if (!timeoutId) {
|
||
timeoutId = setTimeout(trailingCall, remaining);
|
||
}
|
||
return result;
|
||
};
|
||
};
|
||
})();(function() {
|
||
var Set = (function() {
|
||
var add = function(item) {
|
||
var i, data = this._data;
|
||
for (i = 0; i < data.length; i++) {
|
||
if (data[i] === item) {
|
||
return;
|
||
}
|
||
}
|
||
this.size ++;
|
||
data.push(item);
|
||
return data;
|
||
};
|
||
|
||
var Set = function(data) {
|
||
this.size = 0;
|
||
this._data = [];
|
||
var i;
|
||
if (data.length > 0) {
|
||
for (i = 0; i < data.length; i++) {
|
||
add.call(this, data[i]);
|
||
}
|
||
}
|
||
};
|
||
Set.prototype.add = add;
|
||
Set.prototype.get = function(index) { return this._data[index]; };
|
||
Set.prototype.has = function(item) {
|
||
var i, data = this._data;
|
||
for (i = 0; i < data.length; i++) {
|
||
if (this.get(i) === item) {
|
||
return true;
|
||
}
|
||
}
|
||
return false;
|
||
};
|
||
Set.prototype.is = function(map) {
|
||
if (map._data.length !== this._data.length) { return false; }
|
||
var i, j, flag, tData = this._data, mData = map._data;
|
||
for (i = 0; i < tData.length; i++) {
|
||
for (flag = false, j = 0; j < mData.length; j++) {
|
||
if (tData[i] === mData[j]) {
|
||
flag = true;
|
||
break;
|
||
}
|
||
}
|
||
if (!flag) { return false; }
|
||
}
|
||
return true;
|
||
};
|
||
Set.prototype.values = function() {
|
||
return this._data;
|
||
};
|
||
return Set;
|
||
})();
|
||
|
||
window.Lazyload = (function(doc) {
|
||
var queue = {js: [], css: []}, sources = {js: {}, css: {}}, context = this;
|
||
var createNode = function(name, attrs) {
|
||
var node = doc.createElement(name), attr;
|
||
for (attr in attrs) {
|
||
if (attrs.hasOwnProperty(attr)) {
|
||
node.setAttribute(attr, attrs[attr]);
|
||
}
|
||
}
|
||
return node;
|
||
};
|
||
var end = function(type, url) {
|
||
var s, q, qi, cbs, i, j, cur, val, flag;
|
||
if (type === 'js' || type ==='css') {
|
||
s = sources[type], q = queue[type];
|
||
s[url] = true;
|
||
for (i = 0; i < q.length; i++) {
|
||
cur = q[i];
|
||
if (cur.urls.has(url)) {
|
||
qi = cur, val = qi.urls.values();
|
||
qi && (cbs = qi.callbacks);
|
||
for (flag = true, j = 0; j < val.length; j++) {
|
||
cur = val[j];
|
||
if (!s[cur]) {
|
||
flag = false;
|
||
}
|
||
}
|
||
if (flag && cbs && cbs.length > 0) {
|
||
for (j = 0; j < cbs.length; j++) {
|
||
cbs[j].call(context);
|
||
}
|
||
qi.load = true;
|
||
}
|
||
}
|
||
}
|
||
}
|
||
};
|
||
var load = function(type, urls, callback) {
|
||
var s, q, qi, node, i, cur,
|
||
_urls = typeof urls === 'string' ? new Set([urls]) : new Set(urls), val, url;
|
||
if (type === 'js' || type ==='css') {
|
||
s = sources[type], q = queue[type];
|
||
for (i = 0; i < q.length; i++) {
|
||
cur = q[i];
|
||
if (_urls.is(cur.urls)) {
|
||
qi = cur;
|
||
break;
|
||
}
|
||
}
|
||
val = _urls.values();
|
||
if (qi) {
|
||
callback && (qi.load || qi.callbacks.push(callback));
|
||
callback && (qi.load && callback());
|
||
} else {
|
||
q.push({
|
||
urls: _urls,
|
||
callbacks: callback ? [callback] : [],
|
||
load: false
|
||
});
|
||
for (i = 0; i < val.length; i++) {
|
||
node = null, url = val[i];
|
||
if (s[url] === undefined) {
|
||
(type === 'js' ) && (node = createNode('script', { src: url }));
|
||
(type === 'css') && (node = createNode('link', { rel: 'stylesheet', href: url }));
|
||
if (node) {
|
||
node.onload = (function(type, url) {
|
||
return function() {
|
||
end(type, url);
|
||
};
|
||
})(type, url);
|
||
(doc.head || doc.body).appendChild(node);
|
||
s[url] = false;
|
||
}
|
||
}
|
||
}
|
||
}
|
||
}
|
||
};
|
||
return {
|
||
js: function(url, callback) {
|
||
load('js', url, callback);
|
||
},
|
||
css: function(url, callback) {
|
||
load('css', url, callback);
|
||
}
|
||
};
|
||
})(this.document);
|
||
})();</script><script>
|
||
(function() {
|
||
var TEXT_VARIABLES = {
|
||
version: '2.2.1',
|
||
sources: {
|
||
font_awesome: 'https://use.fontawesome.com/releases/v5.0.13/css/all.css',
|
||
jquery: 'https://cdn.bootcss.com/jquery/3.1.1/jquery.min.js',
|
||
leancloud_js_sdk: '//cdn1.lncld.net/static/js/3.4.1/av-min.js',
|
||
chart: 'https://cdn.bootcss.com/Chart.js/2.7.2/Chart.bundle.min.js',
|
||
gitalk: {
|
||
js: 'https://cdn.bootcss.com/gitalk/1.2.2/gitalk.min.js',
|
||
css: 'https://cdn.bootcss.com/gitalk/1.2.2/gitalk.min.css'
|
||
},
|
||
mathjax: 'https://cdn.bootcss.com/mathjax/2.7.4/MathJax.js?config=TeX-MML-AM_CHTML',
|
||
mermaid: 'https://cdn.bootcss.com/mermaid/8.0.0-rc.8/mermaid.min.js'
|
||
},
|
||
site: {
|
||
toc: {
|
||
selectors: 'h1,h2,h3'
|
||
}
|
||
},
|
||
paths: {
|
||
search_js: '/assets/search.js'
|
||
}
|
||
};
|
||
window.TEXT_VARIABLES = TEXT_VARIABLES;
|
||
})();
|
||
</script></head>
|
||
<body>
|
||
<div class="root" data-is-touch="false">
|
||
<div class="layout--page js-page-root"><div class="page__main js-page-main page__viewport has-aside cell cell--auto">
|
||
|
||
<div class="page__main-inner">
|
||
<div></div><div class="page__header"><header class="header"><div class="main">
|
||
<div class="header__title">
|
||
<div class="header__brand"><a title="Tips and tricks to try before you strangle yourself with a wireless mouse
|
||
" href="/">ervin's blog</a></div>
|
||
<button class="button button--secondary button--circle search-button js-search-toggle"><i class="fas fa-search"></i></button>
|
||
</div><nav class="navigation">
|
||
<ul><li class="navigation__link"><a href="/archive/">Archive</a></li><li class="navigation__link"><a href="/about/">About</a></li><li><button class="button button--secondary button--circle search-button js-search-toggle"><i class="fas fa-search"></i></button></li>
|
||
</ul>
|
||
</nav></div>
|
||
</header>
|
||
</div><div class="page__content"><div class ="main"><div class="grid grid--reverse">
|
||
|
||
<div class="col-aside js-col-aside"><aside class="page__aside js-page-aside"><div class="toc-aside js-toc-root"></div></aside></div>
|
||
|
||
<div class="col-main cell--auto"><article itemscope itemtype="http://schema.org/Article"><div class="article__header"><header><h1>SLO Alerting for Mortals</h1></header></div><meta itemprop="headline" content="SLO Alerting for Mortals"><div class="article__info clearfix"><ul class="left-col menu"><li>
|
||
<a class="button button--secondary button--pill button--sm"
|
||
href="/archive.html?tag=sre">sre</a>
|
||
</li></ul><ul class="right-col menu"><li><i class="far fa-calendar-alt"></i> <span>Oct 19, 2021</span>
|
||
</li></ul></div><meta itemprop="author" content="Ervin Barta"/><meta itemprop="datePublished" content="2021-10-19T00:00:00+00:00">
|
||
<meta itemprop="keywords" content="sre"><div class="js-article-content"><div class="layout--article">
|
||
|
||
<div class="article__content" itemprop="articleBody"><p>This is an attempt to break down the concept of SLO alerting as much as possible.
|
||
Step-by-step, each concept will be illustrated and occasionally animated.</p>
|
||
|
||
<p>My hope is that the short material presented here is intuitive enough for it to stick, and to serve as a
|
||
base for further research.</p>
|
||
|
||
<p>I will try to demistify concepts such as burn rate, error budget, and multi-window alerts.</p>
|
||
|
||
<h2 id="slos-and-slis">SLOs and SLIs</h2>
|
||
|
||
<p>Service level objective (SLO). It represents how reliably the service is delivering “value” to its users.</p>
|
||
|
||
<p>Service level indicator (SLI). A measurement of a specific service metric. We’re using SLIs and math to define an SLO.</p>
|
||
|
||
<p>A 99.9% SLO per month means if 0.1% of requests fail, that’s <em>acceptable</em>, and it won’t raise any fuss.</p>
|
||
|
||
<p><img src="/assets/images/slo-alerting/slo-count-999.png" alt="slo-req-count" /></p>
|
||
|
||
<p>We’re ok with the fact that 1 in every 1K requests will fail.</p>
|
||
|
||
<p>An SLI can be the error rate of the incoming requests.
|
||
<img src="/assets/images/slo-alerting/sli-slo-init.png" alt="slo-vs-sli" /></p>
|
||
|
||
<p>The 0.1% wiggle room is the limit above which we don’t want to go - the error budget.</p>
|
||
|
||
<p><img src="/assets/images/slo-alerting/error-budget-timeline.png" alt="error-budget-timeline" /></p>
|
||
|
||
<h2 id="converting-time">Converting time</h2>
|
||
|
||
<p>Most of the math involved here is about converting various time units (i.e 1 month to 720 hours)
|
||
and checking the ratio between them (i.e. 1h is 0.14% of a month). Then using those ratios
|
||
with existing SLIs to come to an SLO condition.</p>
|
||
|
||
<p><img src="/assets/images/slo-alerting/time-conv.png" alt="time-conversion" /></p>
|
||
|
||
<p>Time wise, a 0.1% error rate for a month means a 43 minute complete downtime.</p>
|
||
|
||
<p><img src="/assets/images/slo-alerting/metric-slo-timeline.png" alt="slo-time" /></p>
|
||
|
||
<h2 id="burn-rate">Burn rate</h2>
|
||
|
||
<p>When our budget of 43 minutes a month starts burning, we want to know about it fairly quickly.</p>
|
||
|
||
<p><img src="/assets/images/slo-alerting/error-budget-monthly-graph.png" alt="error-budget-graph" /></p>
|
||
|
||
<p>The green line represents the border between good and evil: if the error rate is <em>exactly</em> 0.1% throughout the month, we’re still fine, but barely.
|
||
The burn rate is 1 in this case. As soon as the line starts to tilt left (into the danger zone), the error rate is higher than the allowed 0.1% and in
|
||
turn, and the burn rate also increases - the system eats the error budget faster than it should.</p>
|
||
|
||
<h2 id="defining-the-first-alert">Defining the first alert</h2>
|
||
|
||
<p>As a start, we define the following alert condition, which is identical to the SLO.</p>
|
||
|
||
<p><img src="/assets/images/slo-alerting/slo-query-first.png" alt="initial-alert-condition" /></p>
|
||
|
||
<p>Keeping the graph above in mind, this translates to “if the green line starts tilting left: alert!”</p>
|
||
|
||
<p>To reiterate on our timeline, the error budget spans out across the whole month. The alert we defined operates in an hour long time window.
|
||
It follows that the 1 hour time window is of course not the whole month, but only 0.14% percent of it.</p>
|
||
|
||
<p>At the error rate of 0.1% (our SLO), 0.14% of the budget is consumed during 1 hour.</p>
|
||
|
||
<p><img src="/assets/images/slo-alerting/metric-slo-sli-timeline.png" alt="slo-vs-sli" /></p>
|
||
|
||
<p>Not really worth to wake up someone over it.</p>
|
||
|
||
<p>Another issue is that the alert will be active for almost an hour. (we will revisit this a bit later)</p>
|
||
|
||
<p>To improve on this, instead of the 0.14% budget burn in an hour, the alert should have a higher threshold and aim at a 2% burn.
|
||
We need to multiply our threshold with some number to reach this new target.</p>
|
||
|
||
<p>Dividing the target value with the current one gives us the multiplier: <code class="language-plaintext highlighter-rouge">2%/0.14% = 14.3</code></p>
|
||
|
||
<p><img src="/assets/images/slo-alerting/slo-query-first-144.png" alt="alert-condition-multiplied" /></p>
|
||
|
||
<p>This magic multiplier is actually the burn rate (or “tilting of the green line to the left” as we defined it a couple of lines above).</p>
|
||
|
||
<h1 id="simulating-an-alert">Simulating an alert</h1>
|
||
|
||
<p>Here’s the scenario:</p>
|
||
<ul>
|
||
<li>10 requests per 5 minutes is our traffic (constant)</li>
|
||
<li>10% error rate for 10 minutes (1 request out of 10 will fail)</li>
|
||
<li>A snapshot of the state is taken every 5 minutes (<code class="language-plaintext highlighter-rouge">scrape_interval</code> in Prometheus terms). In the real world, you will probably have snapshots every 30 seconds or every minute. We’re using 5 minutes here for easier calculation, and to be able to draw square error rate lines, instead of sloping ones.</li>
|
||
</ul>
|
||
|
||
<p><img src="/assets/images/slo-alerting/1h-graph-err-rate.png" alt="1h-graph" /></p>
|
||
|
||
<p>An error rate of 10% in 2 subsequent snapshots (0m-5m, 5m-10m) is enough to trigger the alert.</p>
|
||
|
||
<p><img src="/assets/images/slo-alerting/1h-graph-err-rate-calculation.png" alt="1h-graph-calc" /></p>
|
||
|
||
<p>What immediately stands out here is the long running alert. It will be active for 55 minutes, even tough we’re not constantly in an erronous state.
|
||
The <code class="language-plaintext highlighter-rouge">http_error_rate[1h]</code> metric considers the samples in the last 1 hour, and this time window is shifted with each snapshot. The snapshots are happening at the markers on the animation, every 5 minutes (remember, this is how the scenario was defined).</p>
|
||
|
||
<p><img src="https://imgur.com/q5CUnRk.gif" alt="1h-moving-window-smooth" /></p>
|
||
|
||
<p>As long as both erronous snapshots are inside the window, the alert will be active. When one of them leaves, the error rate drops
|
||
to 0.8% (as per the calculation above) and the alert stops. This is why the alert is firing for 55 minutes instead of 1 hour.
|
||
In a more realistic scenario, with a 30s second snapshot interval, it will be active for 1 hour.</p>
|
||
|
||
<h2 id="multi-window-alerts">Multi-window alerts</h2>
|
||
|
||
<p>One way to combat the long running alert is to introduce another, shorter time window. It will make sure to end the alert, not long after the error rate goes back to normal. A 5 minute time window with the same error rate as before does just that.</p>
|
||
|
||
<p><img src="/assets/images/slo-alerting/1h5m-query.png" alt="1h5m-query" /></p>
|
||
|
||
<p><img src="/assets/images/slo-alerting/1h5m-graph-err-rate.png" alt="1h5m-graph" /></p>
|
||
|
||
<p>A single bad request out of 10 is plenty to trigger the first, 5 minute condition. The condition with the 1 hour window sets off
|
||
after the second subsequent snapshot with an elevated error rate, as it needs 2 erronous requests out of 120 to surpass the threshold, as we saw in the previous section.</p>
|
||
|
||
<p>Both conditions have the same threshold of 1.4% (0.1% * 14.4); the difference is that the 5 minute one takes 10 samples into consideration, and the 1 hour one takes
|
||
120 samples. A bad request has naturally a bigger impact on the smaller sample size than on the bigger one - 1 in 10 vs. 1 in 120. The smaller window
|
||
is more jittery, where the longer one is slugish, but as they meet at the middle, the result is almost the best of both worlds: we have reasonable
|
||
sensitivity and decent reset time (i.e the alert stops when the coast is clear).</p>
|
||
|
||
<p>The alert is active only when the snapshots with a high error rate are in both time windows - in our case this is true for 10 minutes.</p>
|
||
|
||
<p><img src="https://i.imgur.com/2NFzF39.gif" alt="1h5m-moving-window-smooth" /></p>
|
||
|
||
<p>With the addition of the shorter time window, we made sure that alert is matching reality more closely, i.e. it reacts only when there’s an ongoing issue.</p>
|
||
|
||
<p>These were the basics, the next level would be multi-level, multi-burn rate alert which are a bit out of the scope of this post, so refer to
|
||
the <a href="https://sre.google/workbook/alerting-on-slos">Google SLO alerting documentation for more details</a>.</p>
|
||
</div>
|
||
|
||
<footer class="article__footer"><meta itemprop="dateModified" content="2021-10-19T00:00:00+00:00"><!-- this will show at every article content's bottom --><div itemscope itemtype="http://schema.org/Person" class="author-profile card card--flat item"><a href="http://ervinb.github.io" class="item__image"><img itemprop="image" src="https://avatars0.githubusercontent.com/u/2212814?s=400&v=4" /></a><div class="item__content"><meta itemprop="name" content="Ervin Barta">
|
||
<p class="author-profile__name"><meta itemprop="url" content="http://ervinb.github.io">
|
||
<a href="http://ervinb.github.io">Ervin Barta</a></p><div class="author-profile__links"><div class="author-links">
|
||
<ul class="menu menu--nowrap menu--inline"><link itemprop="url" href="https://ervinb.github.io"><li title="Send me Email.">
|
||
<a class="button button--circle mail-button" itemprop="email" href="/cdn-cgi/l/email-protection#5b33321b3e292d3235393a292f3a75383436" target="_blank">
|
||
<i class="fas fa-envelope"></i>
|
||
</a><li title="Follow me on Twitter.">
|
||
<a class="button button--circle twitter-button" itemprop="sameAs" href="https://twitter.com/baer" target="_blank">
|
||
<div class="icon"><svg fill="#000000" width="24px" height="24px" viewBox="0 0 1024 1024" version="1.1" xmlns="http://www.w3.org/2000/svg">
|
||
<path d="M1024.032 194.432c-37.664 16.704-78.176 28-120.672 33.088 43.36-26.016 76.672-67.168 92.384-116.224-40.608 24.064-85.568 41.568-133.408 50.976-38.336-40.832-92.928-66.336-153.344-66.336-116.032 0-210.08 94.048-210.08 210.08 0 16.48 1.856 32.512 5.44 47.872-174.592-8.768-329.408-92.416-433.024-219.52-18.08 31.04-28.448 67.104-28.448 105.632 0 72.896 37.088 137.184 93.472 174.88-34.432-1.088-66.816-10.528-95.168-26.272-0.032 0.864-0.032 1.76-0.032 2.656 0 101.792 72.416 186.688 168.512 205.984-17.632 4.8-36.192 7.36-55.36 7.36-13.536 0-26.688-1.312-39.52-3.776 26.72 83.456 104.32 144.192 196.256 145.888-71.904 56.352-162.496 89.92-260.928 89.92-16.96 0-33.664-0.992-50.112-2.944 92.96 59.616 203.392 94.4 322.048 94.4 386.432 0 597.728-320.128 597.728-597.76 0-9.12-0.192-18.176-0.608-27.168 41.056-29.632 76.672-66.624 104.832-108.736z" />
|
||
</svg>
|
||
</div>
|
||
</a>
|
||
</li><li title="Follow me on Github.">
|
||
<a class="button button--circle github-button" itemprop="sameAs" href="https://github.com/ervinb" target="_blank">
|
||
<div class="icon"><svg fill="#000000" width="24px" height="24px" viewBox="0 0 1024 1024" version="1.1" xmlns="http://www.w3.org/2000/svg">
|
||
<path class="svgpath" data-index="path_0" fill="#272636" d="M0 525.2c0 223.6 143.3 413.7 343 483.5 26.9 6.8 22.8-12.4 22.8-25.4l0-88.7c-155.3 18.2-161.5-84.6-172-101.7-21.1-36-70.8-45.2-56-62.3 35.4-18.2 71.4 4.6 113.1 66.3 30.2 44.7 89.1 37.2 119 29.7 6.5-26.9 20.5-50.9 39.7-69.6C248.8 728.2 181.7 630 181.7 513.2c0-56.6 18.7-108.7 55.3-150.7-23.3-69.3 2.2-128.5 5.6-137.3 66.5-6 135.5 47.6 140.9 51.8 37.8-10.2 80.9-15.6 129.1-15.6 48.5 0 91.8 5.6 129.8 15.9 12.9-9.8 77-55.8 138.8-50.2 3.3 8.8 28.2 66.7 6.3 135 37.1 42.1 56 94.6 56 151.4 0 117-67.5 215.3-228.8 243.7 26.9 26.6 43.6 63.4 43.6 104.2l0 128.8c0.9 10.3 0 20.5 17.2 20.5C878.1 942.4 1024 750.9 1024 525.3c0-282.9-229.3-512-512-512C229.1 13.2 0 242.3 0 525.2L0 525.2z" />
|
||
</svg>
|
||
</div>
|
||
</a>
|
||
</li><li title="Follow me on Linkedin.">
|
||
<a class="button button--circle linkedin-button" itemprop="sameAs" href="https://www.linkedin.com/in/ervinbarta" target="_blank">
|
||
<div class="icon"><svg fill="#000000" width="24px" height="24px" viewBox="0 0 1024 1024" version="1.1" xmlns="http://www.w3.org/2000/svg">
|
||
<path d="M260.096 155.648c0 27.307008-9.899008 50.516992-29.696 69.632-19.796992 19.115008-45.396992 28.672-76.8 28.672-30.036992 0-54.612992-9.556992-73.728-28.672-19.115008-19.115008-28.672-42.324992-28.672-69.632 0-28.672 9.556992-52.224 28.672-70.656 19.115008-18.432 44.372992-27.648 75.776-27.648 31.403008 0 56.32 9.216 74.752 27.648 18.432 18.432 28.331008 41.984 29.696 70.656 0 0 0 0 0 0m-202.752 808.96c0 0 0-632.832 0-632.832 0 0 196.608 0 196.608 0 0 0 0 632.832 0 632.832 0 0-196.608 0-196.608 0 0 0 0 0 0 0m313.344-430.08c0-58.708992-1.364992-126.292992-4.096-202.752 0 0 169.984 0 169.984 0 0 0 10.24 88.064 10.24 88.064 0 0 4.096 0 4.096 0 40.96-68.267008 105.812992-102.4 194.56-102.4 68.267008 0 123.220992 22.868992 164.864 68.608 41.643008 45.739008 62.464 113.664 62.464 203.776 0 0 0 374.784 0 374.784 0 0-196.608 0-196.608 0 0 0 0-350.208 0-350.208 0-91.476992-33.451008-137.216-100.352-137.216-47.787008 0-81.236992 24.576-100.352 73.728-4.096 8.192-6.144 24.576-6.144 49.152 0 0 0 364.544 0 364.544 0 0-198.656 0-198.656 0 0 0 0-430.08 0-430.08 0 0 0 0 0 0" />
|
||
</svg></div>
|
||
</a>
|
||
</li></ul>
|
||
</div>
|
||
</div>
|
||
|
||
</div>
|
||
</div><div class="article__license"><div class="license">
|
||
<p>This work is licensed under a <a itemprop="license" rel="license" href="https://creativecommons.org/licenses/by-nc/4.0/">Attribution-NonCommercial 4.0 International</a> license.
|
||
<a rel="license" href="https://creativecommons.org/licenses/by-nc/4.0/">
|
||
<img alt="Attribution-NonCommercial 4.0 International" src="https://i.creativecommons.org/l/by-nc/4.0/88x31.png" />
|
||
</a>
|
||
</p>
|
||
</div></div></footer><div class="article__section-navigator clearfix"><div class="previous"><span>PREVIOUS</span><a href="/2020/12/05/from-zero-to-encrypted-secrets-in-2-minutes/">From Zero to Encyrpted Secrets in 2 Minutes with SOPS and GPG</a></div></div></div>
|
||
|
||
<script data-cfasync="false" src="/cdn-cgi/scripts/5c5dd728/cloudflare-static/email-decode.min.js"></script><script>(function() {
|
||
var SOURCES = window.TEXT_VARIABLES.sources;
|
||
window.Lazyload.js(SOURCES.jquery, function() {
|
||
$(function() {
|
||
var $this ,$scroll;
|
||
var $articleContent = $('.js-article-content');
|
||
var hasSidebar = $('.js-page-root').hasClass('layout--page--sidebar');
|
||
var scroll = hasSidebar ? '.js-page-main' : 'html, body';
|
||
$scroll = $(scroll);
|
||
|
||
$articleContent.find('.highlight').each(function() {
|
||
$this = $(this);
|
||
$this.attr('data-lang', $this.find('code').attr('data-lang'));
|
||
});
|
||
$articleContent.find('h1[id], h2[id], h3[id], h4[id], h5[id], h6[id]').each(function() {
|
||
$this = $(this);
|
||
$this.append($('<a class="anchor" aria-hidden="true"></a>').html('<i class="fas fa-anchor"></i>'));
|
||
});
|
||
$articleContent.on('click', '.anchor', function() {
|
||
$scroll.scrollToAnchor('#' + $(this).parent().attr('id'), 400);
|
||
});
|
||
});
|
||
});
|
||
})();</script></div><section class="page__comments"><div id="disqus_thread"></div>
|
||
<script>
|
||
/**
|
||
* RECOMMENDED CONFIGURATION VARIABLES: EDIT AND UNCOMMENT THE SECTION BELOW TO INSERT DYNAMIC VALUES FROM YOUR PLATFORM OR CMS.
|
||
* LEARN WHY DEFINING THESE VARIABLES IS IMPORTANT: https://disqus.com/admin/universalcode/#configuration-variables*/
|
||
var disqus_config = function () {
|
||
this.page.url = 'https://ervinbarta.com/2021/10/19/slo-alerting-for-mortals/';
|
||
this.page.identifier = 'slo-alerting';
|
||
};
|
||
(function() { // DON'T EDIT BELOW THIS LINE
|
||
var d = document, s = d.createElement('script');
|
||
s.src = 'https://ervinb.disqus.com/embed.js';
|
||
s.setAttribute('data-timestamp', +new Date());
|
||
(d.head || d.body).appendChild(s);
|
||
})();
|
||
</script>
|
||
<noscript>Please enable JavaScript to view the <a href="https://disqus.com/?ref_noscript">comments powered by Disqus.</a></noscript></section></article>
|
||
</div>
|
||
</div></div></div>
|
||
|
||
<div class="page__footer">
|
||
<div class="footer js-page-footer">
|
||
<div class="main"><aside itemscope itemtype="http://schema.org/Person">
|
||
<meta itemprop="name" content="Ervin Barta"><meta itemprop="url" content="http://ervinb.github.io"><div class="footer__author-links"><div class="author-links">
|
||
<ul class="menu menu--nowrap menu--inline"><link itemprop="url" href="https://ervinb.github.io"><li title="Send me Email.">
|
||
<a class="button button--circle mail-button" itemprop="email" href="/cdn-cgi/l/email-protection#482021082d3a3e21262a293a3c29662b2725" target="_blank">
|
||
<i class="fas fa-envelope"></i>
|
||
</a><li title="Follow me on Twitter.">
|
||
<a class="button button--circle twitter-button" itemprop="sameAs" href="https://twitter.com/baer" target="_blank">
|
||
<div class="icon"><svg fill="#000000" width="24px" height="24px" viewBox="0 0 1024 1024" version="1.1" xmlns="http://www.w3.org/2000/svg">
|
||
<path d="M1024.032 194.432c-37.664 16.704-78.176 28-120.672 33.088 43.36-26.016 76.672-67.168 92.384-116.224-40.608 24.064-85.568 41.568-133.408 50.976-38.336-40.832-92.928-66.336-153.344-66.336-116.032 0-210.08 94.048-210.08 210.08 0 16.48 1.856 32.512 5.44 47.872-174.592-8.768-329.408-92.416-433.024-219.52-18.08 31.04-28.448 67.104-28.448 105.632 0 72.896 37.088 137.184 93.472 174.88-34.432-1.088-66.816-10.528-95.168-26.272-0.032 0.864-0.032 1.76-0.032 2.656 0 101.792 72.416 186.688 168.512 205.984-17.632 4.8-36.192 7.36-55.36 7.36-13.536 0-26.688-1.312-39.52-3.776 26.72 83.456 104.32 144.192 196.256 145.888-71.904 56.352-162.496 89.92-260.928 89.92-16.96 0-33.664-0.992-50.112-2.944 92.96 59.616 203.392 94.4 322.048 94.4 386.432 0 597.728-320.128 597.728-597.76 0-9.12-0.192-18.176-0.608-27.168 41.056-29.632 76.672-66.624 104.832-108.736z" />
|
||
</svg>
|
||
</div>
|
||
</a>
|
||
</li><li title="Follow me on Github.">
|
||
<a class="button button--circle github-button" itemprop="sameAs" href="https://github.com/ervinb" target="_blank">
|
||
<div class="icon"><svg fill="#000000" width="24px" height="24px" viewBox="0 0 1024 1024" version="1.1" xmlns="http://www.w3.org/2000/svg">
|
||
<path class="svgpath" data-index="path_0" fill="#272636" d="M0 525.2c0 223.6 143.3 413.7 343 483.5 26.9 6.8 22.8-12.4 22.8-25.4l0-88.7c-155.3 18.2-161.5-84.6-172-101.7-21.1-36-70.8-45.2-56-62.3 35.4-18.2 71.4 4.6 113.1 66.3 30.2 44.7 89.1 37.2 119 29.7 6.5-26.9 20.5-50.9 39.7-69.6C248.8 728.2 181.7 630 181.7 513.2c0-56.6 18.7-108.7 55.3-150.7-23.3-69.3 2.2-128.5 5.6-137.3 66.5-6 135.5 47.6 140.9 51.8 37.8-10.2 80.9-15.6 129.1-15.6 48.5 0 91.8 5.6 129.8 15.9 12.9-9.8 77-55.8 138.8-50.2 3.3 8.8 28.2 66.7 6.3 135 37.1 42.1 56 94.6 56 151.4 0 117-67.5 215.3-228.8 243.7 26.9 26.6 43.6 63.4 43.6 104.2l0 128.8c0.9 10.3 0 20.5 17.2 20.5C878.1 942.4 1024 750.9 1024 525.3c0-282.9-229.3-512-512-512C229.1 13.2 0 242.3 0 525.2L0 525.2z" />
|
||
</svg>
|
||
</div>
|
||
</a>
|
||
</li><li title="Follow me on Linkedin.">
|
||
<a class="button button--circle linkedin-button" itemprop="sameAs" href="https://www.linkedin.com/in/ervinbarta" target="_blank">
|
||
<div class="icon"><svg fill="#000000" width="24px" height="24px" viewBox="0 0 1024 1024" version="1.1" xmlns="http://www.w3.org/2000/svg">
|
||
<path d="M260.096 155.648c0 27.307008-9.899008 50.516992-29.696 69.632-19.796992 19.115008-45.396992 28.672-76.8 28.672-30.036992 0-54.612992-9.556992-73.728-28.672-19.115008-19.115008-28.672-42.324992-28.672-69.632 0-28.672 9.556992-52.224 28.672-70.656 19.115008-18.432 44.372992-27.648 75.776-27.648 31.403008 0 56.32 9.216 74.752 27.648 18.432 18.432 28.331008 41.984 29.696 70.656 0 0 0 0 0 0m-202.752 808.96c0 0 0-632.832 0-632.832 0 0 196.608 0 196.608 0 0 0 0 632.832 0 632.832 0 0-196.608 0-196.608 0 0 0 0 0 0 0m313.344-430.08c0-58.708992-1.364992-126.292992-4.096-202.752 0 0 169.984 0 169.984 0 0 0 10.24 88.064 10.24 88.064 0 0 4.096 0 4.096 0 40.96-68.267008 105.812992-102.4 194.56-102.4 68.267008 0 123.220992 22.868992 164.864 68.608 41.643008 45.739008 62.464 113.664 62.464 203.776 0 0 0 374.784 0 374.784 0 0-196.608 0-196.608 0 0 0 0-350.208 0-350.208 0-91.476992-33.451008-137.216-100.352-137.216-47.787008 0-81.236992 24.576-100.352 73.728-4.096 8.192-6.144 24.576-6.144 49.152 0 0 0 364.544 0 364.544 0 0-198.656 0-198.656 0 0 0 0-430.08 0-430.08 0 0 0 0 0 0" />
|
||
</svg></div>
|
||
</a>
|
||
</li></ul>
|
||
</div>
|
||
</div>
|
||
</aside>
|
||
<footer class="site-info"><p class="menu menu--center">
|
||
<span>© ervin's blog 2020</span>
|
||
<a type="application/rss+xml" href="/feed.xml">RSS</a>
|
||
</p>
|
||
<p>Powered by <a title="Jekyll is a simple, blog-aware, static site generator." href="http://jekyllrb.com/">Jekyll</a> & <a
|
||
title="TeXt is a succinct theme for blogging." href="https://github.com/kitian616/jekyll-TeXt-theme">TeXt Theme</a>.
|
||
</p>
|
||
</footer>
|
||
</div>
|
||
</div></div>
|
||
</div>
|
||
</div><div class="page__search-panel"><div class="search search--dark">
|
||
<div class="main">
|
||
<div class="search__header">Search</div>
|
||
<div class="search-bar">
|
||
<div class="search-box js-search-box">
|
||
<div class="search-box__icon-search"><i class="fas fa-search"></i></div>
|
||
<input type="text" />
|
||
<div class="search-box__icon-clear js-icon-clear">
|
||
<a><i class="fas fa-times"></i></a>
|
||
</div>
|
||
</div>
|
||
<button class="button button--secondary button--pill search__cancel js-search-toggle">
|
||
Cancel</button>
|
||
</div>
|
||
<div class="search-result js-search-result"></div>
|
||
</div>
|
||
</div>
|
||
|
||
<script data-cfasync="false" src="/cdn-cgi/scripts/5c5dd728/cloudflare-static/email-decode.min.js"></script><script>var SOURCES = window.TEXT_VARIABLES.sources;
|
||
var PAHTS = window.TEXT_VARIABLES.paths;
|
||
window.Lazyload.js([SOURCES.jquery, PAHTS.search_js], function() {
|
||
var searchData = window.TEXT_SEARCH_DATA ? initData(window.TEXT_SEARCH_DATA) : {};
|
||
|
||
function memorize(f) {
|
||
var cache = {};
|
||
return function () {
|
||
var key = Array.prototype.join.call(arguments, ',');
|
||
if (key in cache) return cache[key];
|
||
else return cache[key] = f.apply(this, arguments);
|
||
};
|
||
}
|
||
|
||
function initData(data) {
|
||
var _data = [], i, j, key, keys, cur;
|
||
keys = Object.keys(data);
|
||
for (i = 0; i < keys.length; i++) {
|
||
key = keys[i], _data[key] = [];
|
||
for (j = 0; j < data[key].length; j++) {
|
||
cur = data[key][j];
|
||
cur.title = window.decodeUrl(cur.title);
|
||
cur.url = window.decodeUrl(cur.url);
|
||
_data[key].push(cur);
|
||
}
|
||
}
|
||
return _data;
|
||
}
|
||
|
||
/// search
|
||
function searchByQuery(query) {
|
||
var i, j, key, keys, cur, _title, result = {};
|
||
keys = Object.keys(searchData);
|
||
for (i = 0; i < keys.length; i++) {
|
||
key = keys[i];
|
||
for (j = 0; j < searchData[key].length; j++) {
|
||
cur = searchData[key][j], _title = cur.title;
|
||
if ((result[key] === undefined || result[key] && result[key].length < 4 )
|
||
&& _title.toLowerCase().indexOf(query.toLowerCase()) >= 0) {
|
||
if (result[key] === undefined) {
|
||
result[key] = [];
|
||
}
|
||
result[key].push(cur);
|
||
}
|
||
}
|
||
}
|
||
return result;
|
||
}
|
||
|
||
var renderHeader = memorize(function(header) {
|
||
return $('<p class="search-result__header">' + header + '</p>');
|
||
});
|
||
|
||
var renderItem = function(index, title, url) {
|
||
return $('<li class="search-result__item" data-index="' + index + '"><a class="button" href="' + url + '">' + title + '</a></li>');
|
||
};
|
||
|
||
function render(data) {
|
||
if (!data) {
|
||
return null;
|
||
}
|
||
var $root = $('<ul></ul>'), i, j, key, keys, cur, itemIndex = 0;
|
||
keys = Object.keys(data);
|
||
for (i = 0; i < keys.length; i++) {
|
||
key = keys[i];
|
||
$root.append(renderHeader(key));
|
||
for (j = 0; j < data[key].length; j++) {
|
||
cur = data[key][j];
|
||
$root.append(renderItem(itemIndex++, cur.title, cur.url));
|
||
}
|
||
}
|
||
return $root;
|
||
}
|
||
|
||
// search box
|
||
var $searchBox = $('.js-search-box');
|
||
var $searchInput = $searchBox.children('input');
|
||
var $searchClear = $searchBox.children('.js-icon-clear');
|
||
var $result = $('.js-search-result'), $resultItems;
|
||
var lastActiveIndex, activeIndex;
|
||
|
||
function searchBoxEmpty() {
|
||
$searchBox.removeClass('not-empty'); $result.html(null);
|
||
$resultItems = $('.search-result__item'); activeIndex = 0;
|
||
}
|
||
|
||
$searchInput.on('input', window.throttle(function() {
|
||
var val = $(this).val();
|
||
if (val === '' || typeof val !== 'string') {
|
||
searchBoxEmpty();
|
||
} else {
|
||
$searchBox.addClass('not-empty'); $result.html(render(searchByQuery(val)));
|
||
$resultItems = $('.search-result__item'); activeIndex = 0;
|
||
$resultItems.eq(0).addClass('active');
|
||
}
|
||
}, 400));
|
||
$searchInput.on('focus', function() {
|
||
$(this).addClass('focus');
|
||
});
|
||
$searchInput.on('blur', function() {
|
||
$(this).removeClass('focus');
|
||
});
|
||
$searchClear.on('click', function() {
|
||
$searchInput.val(''); searchBoxEmpty();
|
||
});
|
||
|
||
// search panel
|
||
var $pageRoot = $('.js-page-root');
|
||
var $pageMain = $('.js-page-main');
|
||
var $searchToggle = $('.js-search-toggle');
|
||
var showSearch = false;
|
||
var scrollTop;
|
||
|
||
function openSearchPanel() {
|
||
scrollTop = $(window).scrollTop() || $pageMain.scrollTop();
|
||
$pageRoot.addClass('show-search-panel');
|
||
$pageMain.scrollTop(scrollTop);
|
||
$searchInput[0].focus();
|
||
}
|
||
|
||
function closeSearchPanel() {
|
||
$pageRoot.removeClass('show-search-panel');
|
||
$(window).scrollTop(scrollTop);
|
||
$searchInput[0].blur();
|
||
setTimeout(function() {
|
||
$searchInput.val(''); searchBoxEmpty();
|
||
window.pageAsideAffix && window.pageAsideAffix.refresh();
|
||
}, 400);
|
||
}
|
||
|
||
// Char Code: 13 Enter, 27 ESC, 37 ⬅, 38 ⬆, 39 ➡, 40 ⬇, 83 S, 191 /
|
||
function isFormElement(e) {
|
||
var tagName = e.target.tagName || e.srcElement.tagName;
|
||
return tagName === 'INPUT' || tagName === 'SELECT' || tagName === 'TEXTAREA';
|
||
}
|
||
function charCodeFilter(e) {
|
||
return e.target === $searchInput[0] && (e.which === 13 || e.which === 27 || e.which === 38 || e.which === 40);
|
||
}
|
||
|
||
function updateResultItems() {
|
||
lastActiveIndex >= 0 && $resultItems.eq(lastActiveIndex).removeClass('active');
|
||
activeIndex >= 0 && $resultItems.eq(activeIndex).addClass('active');
|
||
}
|
||
|
||
function moveActiveIndex(direction) {
|
||
var itemsCount = $resultItems ? $resultItems.length : 0;
|
||
if (itemsCount > 1) {
|
||
lastActiveIndex = activeIndex;
|
||
if (direction === 'up') {
|
||
activeIndex = (activeIndex - 1 + itemsCount) % itemsCount;
|
||
} else if (direction === 'down') {
|
||
activeIndex = (activeIndex + 1 + itemsCount) % itemsCount;
|
||
}
|
||
updateResultItems();
|
||
}
|
||
}
|
||
|
||
$(document).on('keyup', function(e) {
|
||
if (!isFormElement(e) || charCodeFilter(e)) {
|
||
if (e.which === 83 || e.which === 191) {
|
||
showSearch || (showSearch = true, openSearchPanel());
|
||
} else if (e.which === 27) {
|
||
showSearch && (showSearch = false, closeSearchPanel());
|
||
} else if (e.which === 38) {
|
||
showSearch && moveActiveIndex('up');
|
||
} else if (e.which === 40) {
|
||
showSearch && moveActiveIndex('down');
|
||
} else if (e.which === 13) {
|
||
showSearch && $resultItems && activeIndex >= 0 && $resultItems.eq(activeIndex).children('a')[0].click();
|
||
}
|
||
}
|
||
});
|
||
|
||
$result.on('mouseover', '.search-result__item > a', function() {
|
||
var itemIndex = $(this).parent().data('index');
|
||
itemIndex >= 0 && (lastActiveIndex = activeIndex, activeIndex = itemIndex, updateResultItems());
|
||
});
|
||
|
||
$searchToggle.on('click', function() {
|
||
showSearch = !showSearch;
|
||
showSearch ? openSearchPanel() : closeSearchPanel();
|
||
});
|
||
});</script></div></div>
|
||
|
||
|
||
<script>(function() {
|
||
var SOURCES = window.TEXT_VARIABLES.sources;
|
||
window.Lazyload.js(SOURCES.jquery, function() {
|
||
function scrollToAnchor(anchor, duration, callback) {
|
||
var $root = this;
|
||
$root.animate({ scrollTop: $(anchor).position().top }, duration, function() {
|
||
window.history.replaceState(null, '', window.location.href.split('#')[0] + anchor);
|
||
callback && callback();
|
||
});
|
||
}
|
||
$.fn.scrollToAnchor = scrollToAnchor;
|
||
});
|
||
})();(function() {
|
||
var SOURCES = window.TEXT_VARIABLES.sources;
|
||
window.Lazyload.js(SOURCES.jquery, function() {
|
||
var $window = $(window), $root, $scrollTarget, $scroller, $scroll;
|
||
var rootTop, rootLeft, rootHeight, scrollBottom, rootBottomTop;
|
||
var offsetBottom = 0, disabled = false, scrollTarget = window, scroller = 'html, body', scroll = window.document;
|
||
var hasInit = false, isOverallScroller = true, curState;
|
||
|
||
function setOptions(options) {
|
||
var _options = options || {};
|
||
_options.offsetBottom && (offsetBottom = _options.offsetBottom);
|
||
_options.scrollTarget && (scrollTarget = _options.scrollTarget);
|
||
_options.scroller && (scroller = _options.scroller);
|
||
_options.scroll && (scroll = _options.scroll);
|
||
_options.disabled !== undefined && (disabled = _options.disabled);
|
||
$scrollTarget = $(scrollTarget);
|
||
$scroller = $(scroller);
|
||
isOverallScroller = window.isOverallScroller($scrollTarget[0]);
|
||
$scroll = $(scroll);
|
||
}
|
||
function initData() {
|
||
top();
|
||
rootHeight = $root.outerHeight();
|
||
rootTop = $root.offset().top + (isOverallScroller ? 0 : $scrollTarget.scrollTop());
|
||
rootLeft = $root.offset().left;
|
||
}
|
||
function calc(needInitData) {
|
||
needInitData && initData();
|
||
scrollBottom = $scroll.outerHeight() - offsetBottom - rootHeight;
|
||
rootBottomTop = scrollBottom - rootTop;
|
||
}
|
||
function top() {
|
||
if (curState !== 'top') {
|
||
$root.removeClass('fixed').css({
|
||
left: 0,
|
||
top: 0
|
||
});
|
||
curState = 'top';
|
||
}
|
||
}
|
||
function fixed() {
|
||
if (curState !== 'fixed') {
|
||
$root.addClass('fixed').css({
|
||
left: rootLeft + 'px',
|
||
top: 0
|
||
});
|
||
curState = 'fixed';
|
||
}
|
||
}
|
||
function bottom() {
|
||
if (curState !== 'bottom') {
|
||
$root.removeClass('fixed').css({
|
||
left: 0,
|
||
top: rootBottomTop + 'px'
|
||
});
|
||
curState = 'bottom';
|
||
}
|
||
}
|
||
function setState() {
|
||
var scrollTop = $scrollTarget.scrollTop();
|
||
if (scrollTop >= rootTop && scrollTop <= scrollBottom) {
|
||
fixed();
|
||
} else if (scrollTop < rootTop) {
|
||
top();
|
||
} else {
|
||
bottom();
|
||
}
|
||
}
|
||
function init() {
|
||
if(!hasInit) {
|
||
var interval, timeout;
|
||
calc(true); setState();
|
||
// run calc every 100 millisecond
|
||
interval = setInterval(function() {
|
||
calc();
|
||
}, 100);
|
||
timeout = setTimeout(function() {
|
||
clearInterval(interval);
|
||
}, 45000);
|
||
window.pageLoad.then(function() {
|
||
setTimeout(function() {
|
||
clearInterval(interval);
|
||
clearTimeout(timeout);
|
||
}, 3000);
|
||
});
|
||
$scrollTarget.on('scroll', function() {
|
||
disabled || setState();
|
||
});
|
||
$window.on('resize', function() {
|
||
disabled || (calc(true), setState());
|
||
});
|
||
hasInit = true;
|
||
}
|
||
}
|
||
|
||
function affix(options) {
|
||
$root = this;
|
||
setOptions(options);
|
||
if (!disabled) {
|
||
init();
|
||
}
|
||
$window.on('resize', window.throttle(function() {
|
||
init();
|
||
}, 200));
|
||
return {
|
||
setOptions: setOptions,
|
||
refresh: function() {
|
||
calc(true); setState();
|
||
}
|
||
};
|
||
}
|
||
$.fn.affix = affix;
|
||
});
|
||
})();(function() {
|
||
var SOURCES = window.TEXT_VARIABLES.sources;
|
||
window.Lazyload.js(SOURCES.jquery, function() {
|
||
var $window = $(window), $root, $scrollTarget, $scroller, $tocUl = $('<ul class="toc"></ul>'), $tocLi, $headings, $activeLast, $activeCur;
|
||
var selectors = 'h1,h2,h3', container = 'body', scrollTarget = window, scroller = 'html, body', disabled = false;
|
||
var headingsPos, scrolling = false, rendered = false, hasInit = false;
|
||
function setOptions(options) {
|
||
var _options = options || {};
|
||
_options.selectors && (selectors = _options.selectors);
|
||
_options.container && (container = _options.container);
|
||
_options.scrollTarget && (scrollTarget = _options.scrollTarget);
|
||
_options.scroller && (scroller = _options.scroller);
|
||
_options.disabled !== undefined && (disabled = _options.disabled);
|
||
$headings = $(container).find(selectors).filter('[id]');
|
||
$scrollTarget = $(scrollTarget);
|
||
$scroller = $(scroller);
|
||
}
|
||
function calc() {
|
||
headingsPos = [];
|
||
$headings.each(function() {
|
||
headingsPos.push(Math.floor($(this).position().top));
|
||
});
|
||
}
|
||
function setState(element, disabled) {
|
||
var scrollTop = $scrollTarget.scrollTop(), i;
|
||
if (disabled || !headingsPos || headingsPos.length < 1) { return; }
|
||
if (element) {
|
||
$activeCur = element;
|
||
} else {
|
||
for (i = 0; i < headingsPos.length; i++) {
|
||
if (scrollTop >= headingsPos[i]) {
|
||
$activeCur = $tocLi.eq(i);
|
||
} else {
|
||
$activeCur || ($activeCur = $tocLi.eq(i));
|
||
break;
|
||
}
|
||
}
|
||
}
|
||
$activeLast && $activeLast.removeClass('active');
|
||
($activeLast = $activeCur).addClass('active');
|
||
}
|
||
function render() {
|
||
if(!rendered) {
|
||
$root.append($tocUl);
|
||
$headings.each(function() {
|
||
var $this = $(this);
|
||
$tocUl.append($('<li></li>').addClass('toc-' + $this.prop('tagName').toLowerCase())
|
||
.append($('<a></a>').text($this.text()).attr('href', '#' + $this.prop('id'))));
|
||
});
|
||
$tocLi = $tocUl.children('li');
|
||
$tocUl.on('click', 'a', function(e) {
|
||
e.preventDefault();
|
||
var $this = $(this);
|
||
scrolling = true;
|
||
setState($this.parent());
|
||
$scroller.scrollToAnchor($this.attr('href'), 400, function() {
|
||
scrolling = false;
|
||
});
|
||
});
|
||
}
|
||
rendered = true;
|
||
}
|
||
function init() {
|
||
var interval, timeout;
|
||
if(!hasInit) {
|
||
render(); calc(); setState(null, scrolling);
|
||
// run calc every 100 millisecond
|
||
interval = setInterval(function() {
|
||
calc();
|
||
}, 100);
|
||
timeout = setTimeout(function() {
|
||
clearInterval(interval);
|
||
}, 45000);
|
||
window.pageLoad.then(function() {
|
||
setTimeout(function() {
|
||
clearInterval(interval);
|
||
clearTimeout(timeout);
|
||
}, 3000);
|
||
});
|
||
$scrollTarget.on('scroll', function() {
|
||
disabled || setState(null, scrolling);
|
||
});
|
||
$window.on('resize', window.throttle(function() {
|
||
if (!disabled) {
|
||
render(); calc(); setState(null, scrolling);
|
||
}
|
||
}, 100));
|
||
}
|
||
hasInit = true;
|
||
}
|
||
function toc(options) {
|
||
$root = this;
|
||
setOptions(options);
|
||
if (!disabled) {
|
||
init();
|
||
}
|
||
$window.on('resize', window.throttle(function() {
|
||
init();
|
||
}, 200));
|
||
return {
|
||
setOptions: setOptions
|
||
};
|
||
}
|
||
$.fn.toc = toc;
|
||
});
|
||
})();/*(function () {
|
||
|
||
})();*/</script><script>
|
||
/* toc must before affix, since affix need to konw toc' height. */(function() {
|
||
var SOURCES = window.TEXT_VARIABLES.sources;
|
||
var TOC_SELECTOR = window.TEXT_VARIABLES.site.toc.selectors;
|
||
window.Lazyload.js(SOURCES.jquery, function() {
|
||
var $window = $(window);
|
||
var $articleContent = $('.js-article-content');
|
||
var $tocRoot = $('.js-toc-root'), $col2 = $('.js-col-aside');
|
||
var toc;
|
||
var tocDisabled = false;
|
||
var hasSidebar = $('.js-page-root').hasClass('layout--page--sidebar');
|
||
var hasToc = $articleContent.find(TOC_SELECTOR).length > 0;
|
||
|
||
function disabled() {
|
||
return $col2.css('display') === 'none' || !hasToc;
|
||
}
|
||
|
||
tocDisabled = disabled();
|
||
|
||
toc = $tocRoot.toc({
|
||
selectors: TOC_SELECTOR,
|
||
container: $articleContent,
|
||
scrollTarget: hasSidebar ? '.js-page-main' : null,
|
||
scroller: hasSidebar ? '.js-page-main' : null,
|
||
disabled: tocDisabled
|
||
});
|
||
|
||
$window.on('resize', window.throttle(function() {
|
||
tocDisabled = disabled();
|
||
toc && toc.setOptions({
|
||
disabled: tocDisabled
|
||
});
|
||
}, 100));
|
||
|
||
});
|
||
})();(function() {
|
||
var SOURCES = window.TEXT_VARIABLES.sources;
|
||
window.Lazyload.js(SOURCES.jquery, function() {
|
||
var $window = $(window), $pageFooter = $('.js-page-footer');
|
||
var $pageAside = $('.js-page-aside');
|
||
var affix;
|
||
var tocDisabled = false;
|
||
var hasSidebar = $('.js-page-root').hasClass('layout--page--sidebar');
|
||
|
||
affix = $pageAside.affix({
|
||
offsetBottom: $pageFooter.outerHeight(),
|
||
scrollTarget: hasSidebar ? '.js-page-main' : null,
|
||
scroller: hasSidebar ? '.js-page-main' : null,
|
||
scroll: hasSidebar ? $('.js-page-main').children() : null,
|
||
disabled: tocDisabled
|
||
});
|
||
|
||
$window.on('resize', window.throttle(function() {
|
||
affix && affix.setOptions({
|
||
disabled: tocDisabled
|
||
});
|
||
}, 100));
|
||
|
||
window.pageAsideAffix = affix;
|
||
});
|
||
})();</script>
|
||
</div>
|
||
<script>(function () {
|
||
var $root = document.getElementsByClassName('root')[0];
|
||
if (window.hasEvent('touchstart')) {
|
||
$root.dataset.isTouch = true;
|
||
document.addEventListener('touchstart', function(){}, false);
|
||
}
|
||
})();</script>
|
||
<script type="module" src="https://static.cloudflareinsights.com/beacon.min.js/v31edd6df95cf4e85bb4c19e7a9bdbcba1788362987495" integrity="sha512-iIg7k2xntmwu6/uSb5tpc/hySgZc4eoL31yB29W6tJFo2akwjPWcEqnCEdJvGexCL0KEQwVYv5BlowfhVz26hg==" data-cf-beacon='{"version":"2024.11.0","token":"be32b0c56dc64a72b1b10e4732e170f4","r":1,"spa":2}' crossorigin="anonymous"></script>
|
||
</body>
|
||
</html> |