23 lines
40 KiB
HTML
23 lines
40 KiB
HTML
<!doctype html><html lang=en dir=auto data-theme=auto><head><meta charset=utf-8><meta http-equiv=X-UA-Compatible content="IE=edge"><meta name=viewport content="width=device-width,initial-scale=1,shrink-to-fit=no"><meta name=robots content="index, follow"><title>Microservice Observability, Part 1: Disambiguating Observability and Monitoring | Brave New Geek</title><meta name=keywords content="cloud,cloud-native,debugging,devops,microservices,monitoring,observability,observability pipeline,ops"><meta name=description content="“Pets versus cattle” has become something of a standard vernacular for describing the shift in how we build systems. It alludes to the elastic and dynamic nature of these (typically, but not necessarily) container-based systems with on-demand scaling and more transparent fault-tolerance. I’ve talked before about this transition before and specifically how it relates to monitoring. In particular, with these more dynamic, microservice-based systems, the conversation starts to shift away from traditional monitoring toward observability. In this series, I’ll describe that distinction, explain why it matters, and share some concrete tactical items for implementing observability in a microservice environment."><meta name=author content><link rel=canonical href=https://bravenewgeek.com/microservice-observability-part-1-disambiguating-observability-and-monitoring/><link crossorigin=anonymous href=/assets/css/stylesheet.4861a452a4c13a9a1fbf2085400b74a7de96b1beeb94dee57b4273e5dffcf337.css integrity="sha256-SGGkUqTBOpofvyCFQAt0p96Wsb7rlN7le0Jz5d/88zc=" rel="preload stylesheet" as=style><link rel=icon href=https://bravenewgeek.com/favicon.ico><link rel=icon type=image/png sizes=16x16 href=https://bravenewgeek.com/favicon.ico><link rel=icon type=image/png sizes=32x32 href=https://bravenewgeek.com/favicon.ico><link rel=apple-touch-icon href=https://bravenewgeek.com/favicon.ico><link rel=mask-icon href=https://bravenewgeek.com/favicon.ico><meta name=theme-color content="#2e2e33"><meta name=msapplication-TileColor content="#2e2e33"><link rel=alternate hreflang=en href=https://bravenewgeek.com/microservice-observability-part-1-disambiguating-observability-and-monitoring/><noscript><style>#theme-toggle,.top-link{display:none}</style><style>@media(prefers-color-scheme:dark){:root{--theme:rgb(29, 30, 32);--entry:rgb(46, 46, 51);--primary:rgb(218, 218, 219);--secondary:rgb(155, 156, 157);--tertiary:rgb(65, 66, 68);--content:rgb(196, 196, 197);--code-block-bg:rgb(46, 46, 51);--code-bg:rgb(55, 56, 62);--border:rgb(51, 51, 51);color-scheme:dark}.list{background:var(--theme)}.toc{background:var(--entry)}}</style></noscript><script>localStorage.getItem("pref-theme")==="dark"?document.querySelector("html").dataset.theme="dark":localStorage.getItem("pref-theme")==="light"?document.querySelector("html").dataset.theme="light":window.matchMedia("(prefers-color-scheme: dark)").matches?document.querySelector("html").dataset.theme="dark":document.querySelector("html").dataset.theme="light"</script><link rel=preconnect href=https://fonts.googleapis.com><link rel=preconnect href=https://fonts.gstatic.com crossorigin><link rel=stylesheet href="https://fonts.googleapis.com/css2?family=JetBrains+Mono:wght@400;500&family=Source+Serif+4:ital,opsz,wght@0,8..60,400;0,8..60,600;1,8..60,400&family=Space+Grotesk:wght@500;600;700&display=swap"><meta property="og:url" content="https://bravenewgeek.com/microservice-observability-part-1-disambiguating-observability-and-monitoring/"><meta property="og:site_name" content="Brave New Geek"><meta property="og:title" content="Microservice Observability, Part 1: Disambiguating Observability and Monitoring"><meta property="og:description" content="“Pets versus cattle” has become something of a standard vernacular for describing the shift in how we build systems. It alludes to the elastic and dynamic nature of these (typically, but not necessarily) container-based systems with on-demand scaling and more transparent fault-tolerance. I’ve talked before about this transition before and specifically how it relates to monitoring. In particular, with these more dynamic, microservice-based systems, the conversation starts to shift away from traditional monitoring toward observability. In this series, I’ll describe that distinction, explain why it matters, and share some concrete tactical items for implementing observability in a microservice environment."><meta property="og:locale" content="en_us"><meta property="og:type" content="article"><meta property="article:section" content="posts"><meta property="article:published_time" content="2019-10-03T10:55:23-05:00"><meta property="article:modified_time" content="2020-01-03T14:21:25-05:00"><meta property="article:tag" content="Cloud"><meta property="article:tag" content="Cloud-Native"><meta property="article:tag" content="Debugging"><meta property="article:tag" content="Devops"><meta property="article:tag" content="Microservices"><meta property="article:tag" content="Monitoring"><meta name=twitter:card content="summary"><meta name=twitter:title content="Microservice Observability, Part 1: Disambiguating Observability and Monitoring"><meta name=twitter:description content="“Pets versus cattle” has become something of a standard vernacular for describing the shift in how we build systems. It alludes to the elastic and dynamic nature of these (typically, but not necessarily) container-based systems with on-demand scaling and more transparent fault-tolerance. I’ve talked before about this transition before and specifically how it relates to monitoring. In particular, with these more dynamic, microservice-based systems, the conversation starts to shift away from traditional monitoring toward observability. In this series, I’ll describe that distinction, explain why it matters, and share some concrete tactical items for implementing observability in a microservice environment."><script type=application/ld+json>{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Posts","item":"https://bravenewgeek.com/posts/"},{"@type":"ListItem","position":2,"name":"Microservice Observability, Part 1: Disambiguating Observability and Monitoring","item":"https://bravenewgeek.com/microservice-observability-part-1-disambiguating-observability-and-monitoring/"}]}</script><script type=application/ld+json>{"@context":"https://schema.org","@type":"BlogPosting","headline":"Microservice Observability, Part 1: Disambiguating Observability and Monitoring","name":"Microservice Observability, Part 1: Disambiguating Observability and Monitoring","description":"“Pets versus cattle” has become something of a standard vernacular for describing the shift in how we build systems. It alludes to the elastic and dynamic nature of these (typically, but not necessarily) container-based systems with on-demand scaling and more transparent fault-tolerance. I’ve talked before about this transition before and specifically how it relates to monitoring. In particular, with these more dynamic, microservice-based systems, the conversation starts to shift away from traditional monitoring toward observability. In this series, I’ll describe that distinction, explain why it matters, and share some concrete tactical items for implementing observability in a microservice environment.\n","keywords":["cloud","cloud-native","debugging","devops","microservices","monitoring","observability","observability pipeline","ops"],"articleBody":"“Pets versus cattle” has become something of a standard vernacular for describing the shift in how we build systems. It alludes to the elastic and dynamic nature of these (typically, but not necessarily) container-based systems with on-demand scaling and more transparent fault-tolerance. I’ve talked before about this transition before and specifically how it relates to monitoring. In particular, with these more dynamic, microservice-based systems, the conversation starts to shift away from traditional monitoring toward observability. In this series, I’ll describe that distinction, explain why it matters, and share some concrete tactical items for implementing observability in a microservice environment.\nIn the past, I’ve used the term “cloud-native” to describe these types of systems, but this buzzword has conflated so many different concepts that it’s been relegated to the likes of “DevOps”—entirely arbitrary and context-dependent. Depending on who you ask, cloud-native means containers, microservices, Kubernetes, elasticity, serverless, automation, or any number of other ideas. The truth, however, is that you can do many of these things on-prem just as much as in the cloud, the difference being largely CapEx versus OpEx. I think the spirit of “cloud-native” really just means architecting systems to take advantage of cloud capabilities, namely higher-level managed services (which may not even have on-prem equivalents), improved elasticity and fault-tolerance (which may or may not mean containers), and reduced operations investment (in part by leveraging managed services).\nBecause there are so many confounding and interrelated-yet-different ideas, I’m going to focus this discussion on elastic microservice architectures. Elastic meaning services that automatically scale up and down as needed (in contrast to static infrastructures), and microservice simply meaning applications comprised of many different—usually smaller—services (in contrast to monoliths or systems comprising just a few coarse-grained services).\nStatic Monolithic Architectures With static monolithic architectures, monitoring is a reasonably well-understood problem. With a monolith, the system is typically in one of two states, up or down, and we can conceivably correlate this to customer impact. Bugs aside, when the monolith is down, we likely have a good idea of how this behavior manifests itself to the user. We can set up Nagios checks and get some meaningful signals out of it. Uptime is mostly a single data point.\nWith a monolith, it’s not unreasonable for ops teams to manage the day-to-day operations of the system and do so effectively. These teams tend to quickly develop a good intuition and “muscle memory” for the application when it’s the only thing they are responsible for, especially when it’s a single deployable unit. Logs can be grepped from a single log file, and if something is wrong with the application, operators might simply SSH into the box to poke at it. Runbooks and standard operating procedures are also common here.\nWith a monolith, we likely have a single runtime such as the JVM, which makes it easier to collect rich telemetry in a centralized way, all the way down to the code level. Tools like Dynatrace and AppDynamics can instrument the JVM itself to collect information on busy and idle threads, garbage collection stats, and request metrics. And because we have just a single deployed artifact running on a handful of static servers, this data can actually be useful and correlated back to customer impact and business metrics.\nElastic Microservice Architectures With elastic microservice architectures, things start to change dramatically. Applications consist of dozens of different microservices. The system is no longer in one of two states but more like one of n-factorial states. In reality, it’s much more because in production you might have different versions of the same service running at the same time as you introduce more sophisticated deployment strategies and rollbacks. Integration testing can’t possibly account for all of these combinations. We can no longer easily correlate system behavior to actual customer impact because system behavior is much more emergent. It can be difficult to pinpoint how the behavior of a given service affects the user’s experience as the system operates in varying states of partial failure and services interact in unique ways. If it’s slow, which part is slow? The frontend service? An upstream service? The database? Some combination of these? Uptime is no longer a single data point but rather a composite of many different data points, but more importantly, what does “up” even mean in the context of a complex microservice architecture?\nWith microservices, it becomes intractable for a single ops team to manage dozens of heterogeneous services beyond anything but in a first-responder, incident-router capacity. There is too much context and specific knowledge needed since microservices are literally the embodiment of the specialization of teams.\nWith microservices, it’s no longer practical or even feasible to grep log files or SSH into the box to debug a problem. There might not even be a box to SSH into if it’s a container that has since been descheduled or a managed serverless runtime. With heterogeneous services, we might have half a dozen languages and runtimes to support, each with differing types of runtime instrumentation. Moreover, because we now have dozens or even hundreds of nodes running many different instances of our services, the value of this low-level, summarized data starts to diminish. It makes for pretty dashboards and can help in answering very specific, predefined questions, but that’s about it. It’s no use for proactive monitoring because it’s too much noise, and it’s no use for reactive debugging because it’s pre-aggregated. There’s not much you can do when all you have are rolled-up time-series metrics, and it’s just as difficult to correlate this data back to customer impact.\nMonitoring and Observability With a complex system, relying on this type of data along with logs can often lead to a deadend when tracking down a particularly insidious bug. And this is where observability comes into play. It picks up where monitoring leaves off.\nWhile monitoring and observability have been getting conflated a lot lately, there’s actually an important distinction to make. Monitoring tends to focus on the overall health of system and business metrics—questions we know in advance. Observability is about providing more granular insights into the behavior of systems and richer context. It’s the difference between “post hoc” versus “ad hoc.”\nIn the top-right corner, we have known knowns. These are things of which we have a high degree of understanding and a large amount of data on, i.e. the things we are aware of and understand. For example, “the system has a 1GB memory limit.” As the designers of this system, this is something that we’re acutely aware of and understand. We know that we know how much memory the system can use before it moves outside of its operating boundaries and bad things happen.\nIn the bottom-right corner, we have known unknowns. These are things we are generally aware of but don’t necessarily understand. For example, “the system exceeded its memory limit and crashed, causing an outage.” As system designers, memory usage is something we know is important and affects system behavior. We can monitor it in production in order to gather lots of data on it, but just having that data often doesn’t help us to understand why memory is being consumed or even how that data manifests itself as system behavior.\nIn the top-left corner, we have unknown knowns, which are things we understand but are not completely aware of. This sounds like a strange, almost oxymoron-like categorization, but it’s basically the things that are gut instinct or intuition. It’s often things we know or think we know without even consciously realizing it. For example, “we implemented an orchestrator to ensure the system is always running.” Intuition tells us that if the process isn’t running, the system isn’t available, so we make sure that it gets restarted when something goes wrong. We might, however, be unaware of the unintended side effects of this decision, and it might be based more on theory and conjecture than data.\nWhich leads us to the bottom-left corner: unknown unknowns. These are the things we are neither aware of nor understand. The events we can’t even predict or foresee happening because if we could foresee them, they wouldn’t be unknown unknowns, they’d be known unknowns. For example, “instances churn because the orchestrator restarts the process when it approaches its memory limit, causing sporadic failures and slowdowns.” This was an unforeseen consequence of our orchestrator implementation. As a result, we could not have tested for it or looked for it with our monitoring tools. Instead, it’s something that happens, we learn from it, and quickly classify it as a known unknown—something we know to look for going forward.\nIn a sense, the known knowns are facts, the known unknowns are hypotheses, the unknown knowns are assumptions, and the unknown unknowns are discoveries. Through this lens, the distinction between observability and monitoring becomes clear. Monitoring is about testing hypotheses and observability is about exploring new discoveries. We monitor known unknowns because these are the things we know to look for, but unknown unknowns are, by definition, unpredictable. We cannot monitor them because we do not know to even look for them in the first place! Instead, we ask questions of our systems in order to understand and categorize these unknown unknowns. Observability is the ability to interrogate our systems after the fact in a data-rich, high-fidelity way. Monitoring, on the other hand, is before the fact and much lower fidelity. These are the dashboards and alerts we set up which usually consist of pre-aggregated metrics. This is what I mean by post hoc versus ad hoc. Observability allows us to ask arbitrary questions of our systems, not questions predefined in advance.\nWith this definition, monitoring is a subset of observability, and observability encompasses many different types of data. For example, things like distributed traces, application logs, system logs, audit logs, and application metrics are all important observability signals. But when we boil it all down, it turns out everything is really just events, of which we want different lenses to view. Some of this data provides context for the event itself, such as logs and metrics, and some of it describes relationships between events, such as traces. It’s important we have a way to collect all this context and store it such that we can query and analyze it using these different lenses. Aggregated metrics alone aren’t enough—they don’t have the granularity nor the context needed. Dashboards are simply answers to specific questions known in advance. Observability needs to go much deeper than this.\nIn part two of this series, we’ll revisit the concept of an observability pipeline as a tactical approach to implementing observability in a microservice environment. As part of this, we’ll discuss some steps that can be taken to incrementally improve observability while iterating toward this pattern.\n","wordCount":"1798","inLanguage":"en","datePublished":"2019-10-03T10:55:23-05:00","dateModified":"2020-01-03T14:21:25-05:00","mainEntityOfPage":{"@type":"WebPage","@id":"https://bravenewgeek.com/microservice-observability-part-1-disambiguating-observability-and-monitoring/"},"publisher":{"@type":"Organization","name":"Brave New Geek","logo":{"@type":"ImageObject","url":"https://bravenewgeek.com/favicon.ico"}}}</script></head><body id=top><header class=site-header><div class="wrap site-header-inner"><a class=wordmark href=https://bravenewgeek.com/ accesskey=h title="Brave New Geek (Alt + H)"><span class=wordmark-name>Brave New Geek</span>
|
||
<span class=wordmark-tag>Introspections of a software engineer</span></a><nav class=site-nav aria-label=Primary><a href=https://bravenewgeek.com/archive/>Archive</a>
|
||
<a href=https://bravenewgeek.com/tags/>Tags</a>
|
||
<a href=https://bravenewgeek.com/about-me/>About</a>
|
||
<button id=theme-toggle class=theme-toggle accesskey=t title="Toggle theme (Alt + T)" aria-label="Toggle light/dark theme">
|
||
<svg class="moon" width="18" height="18" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M21 12.79A9 9 0 1111.21 3 7 7 0 0021 12.79z"/></svg>
|
||
<svg class="sun" width="18" height="18" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><circle cx="12" cy="12" r="5"/><line x1="12" y1="1" x2="12" y2="3"/><line x1="12" y1="21" x2="12" y2="23"/><line x1="4.22" y1="4.22" x2="5.64" y2="5.64"/><line x1="18.36" y1="18.36" x2="19.78" y2="19.78"/><line x1="1" y1="12" x2="3" y2="12"/><line x1="21" y1="12" x2="23" y2="12"/><line x1="4.22" y1="19.78" x2="5.64" y2="18.36"/><line x1="18.36" y1="5.64" x2="19.78" y2="4.22"/></svg></button></nav></div></header><main class=main><article class="post wrap"><div class=post-return><a href=https://bravenewgeek.com/><span class=pager-arrow>←</span> the log</a></div><header class=post-header><div class=post-meta><span class=post-offset>#90</span><time datetime=2019-10-03>2019-10-03</time><span>9 min read</span>
|
||
<span class=post-cats><a href=https://bravenewgeek.com/category/cloud/>Cloud</a><a href=https://bravenewgeek.com/category/devops/>DevOps</a><a href=https://bravenewgeek.com/category/operations/>Operations</a><a href=https://bravenewgeek.com/category/software-engineering/>Software Engineering</a></span></div><h1 class=post-title>Microservice Observability, Part 1: Disambiguating Observability and Monitoring</h1></header><div class="post-content md-content"><p>“Pets versus cattle” has become something of a standard vernacular for describing the shift in how we build systems. It alludes to the elastic and dynamic nature of these (typically, but not necessarily) container-based systems with on-demand scaling and more transparent fault-tolerance. I’ve <a href=https://bravenewgeek.com/the-observability-pipeline/>talked before about this transition</a> before and specifically how it relates to monitoring. In particular, with these more dynamic, microservice-based systems, the conversation starts to shift away from traditional <em>monitoring</em> toward <em>observability</em>. In this series, I’ll describe that distinction, explain why it matters, and share some concrete tactical items for implementing observability in a microservice environment.</p><p>In the past, I’ve used the term “cloud-native” to describe these types of systems, but this buzzword has conflated so many different concepts that it’s been relegated to the likes of “DevOps”—entirely arbitrary and context-dependent. Depending on who you ask, cloud-native means containers, microservices, Kubernetes, elasticity, serverless, automation, or any number of other ideas. The truth, however, is that you can do many of these things on-prem just as much as in the cloud, the difference being largely CapEx versus OpEx. I think the spirit of “cloud-native” really just means architecting systems to take advantage of cloud capabilities, namely higher-level managed services (which may not even have on-prem equivalents), improved elasticity and fault-tolerance (which may or may not mean containers), and reduced operations investment (in part by leveraging managed services).</p><p>Because there are so many confounding and interrelated-yet-different ideas, I’m going to focus this discussion on <em>elastic microservice architectures</em>. <em>Elastic</em> meaning services that automatically scale up and down as needed (in contrast to static infrastructures), and <em>microservice</em> simply meaning applications comprised of many different—usually smaller—services (in contrast to monoliths or systems comprising just a few coarse-grained services).</p><h2 id=static-monolithic-architectures>Static Monolithic Architectures<a hidden class=anchor aria-hidden=true href=#static-monolithic-architectures>#</a></h2><p>With <em>static monolithic architectures</em>, monitoring is a reasonably well-understood problem. With a monolith, the system is typically in one of two states, up or down, and we can conceivably correlate this to customer impact. Bugs aside, when the monolith is down, we likely have a good idea of how this behavior manifests itself to the user. We can set up Nagios checks and get some meaningful signals out of it. Uptime is mostly a single data point.</p><p>With a monolith, it’s not unreasonable for ops teams to manage the day-to-day operations of the system and do so effectively. These teams tend to quickly develop a good intuition and “muscle memory” for the application when it’s the only thing they are responsible for, especially when it’s a single deployable unit. Logs can be grepped from a single log file, and if something is wrong with the application, operators might simply SSH into the box to poke at it. Runbooks and standard operating procedures are also common here.</p><p>With a monolith, we likely have a single runtime such as the JVM, which makes it easier to collect rich telemetry in a centralized way, all the way down to the code level. Tools like Dynatrace and AppDynamics can instrument the JVM itself to collect information on busy and idle threads, garbage collection stats, and request metrics. And because we have just a single deployed artifact running on a handful of static servers, this data can actually be useful and correlated back to customer impact and business metrics.</p><h2 id=elastic-microservice-architectures>Elastic Microservice Architectures<a hidden class=anchor aria-hidden=true href=#elastic-microservice-architectures>#</a></h2><p>With <em>elastic microservice architectures</em>, things start to change dramatically. Applications consist of dozens of different microservices. The system is no longer in one of two states but more like one of <em>n-factorial</em> states. In reality, it’s much more because in production you might have different <em>versions</em> of the same service running at the same time as you introduce more sophisticated deployment strategies and rollbacks. Integration testing can’t possibly account for all of these combinations. We can no longer easily correlate system behavior to actual customer impact because system behavior is much more emergent. It can be difficult to pinpoint how the behavior of a given service affects the user’s experience as the system operates in varying states of partial failure and services interact in unique ways. If it’s slow, which part is slow? The frontend service? An upstream service? The database? Some combination of these? Uptime is no longer a single data point but rather a composite of many different data points, but more importantly, what does “up” even mean in the context of a complex microservice architecture?</p><p>With microservices, it becomes intractable for a single ops team to manage dozens of heterogeneous services beyond anything but in a first-responder, incident-router capacity. There is too much context and specific knowledge needed since microservices are literally the embodiment of the specialization of teams.</p><p>With microservices, it’s no longer practical or even feasible to grep log files or SSH into the box to debug a problem. There might not even <em>be</em> a box to SSH into if it’s a container that has since been descheduled or a managed serverless runtime. With heterogeneous services, we might have half a dozen languages and runtimes to support, each with differing types of runtime instrumentation. Moreover, because we now have dozens or even <em>hundreds</em> of nodes running many different instances of our services, the value of this low-level, summarized data starts to diminish. It makes for pretty dashboards and can help in answering very specific, predefined questions, but that’s about it. It’s no use for <em>proactive</em> monitoring because it’s too much noise, and it’s no use for <em>reactive</em> debugging because it’s pre-aggregated. There’s not much you can do when all you have are rolled-up time-series metrics, and it’s just as difficult to correlate this data back to customer impact.</p><h2 id=monitoring-and-observability>Monitoring and Observability<a hidden class=anchor aria-hidden=true href=#monitoring-and-observability>#</a></h2><p>With a complex system, relying on this type of data along with logs can often lead to a deadend when tracking down a particularly insidious bug. And this is where <em>observability</em> comes into play. It picks up where monitoring leaves off.</p><p>While monitoring and observability have been getting conflated a lot lately, there’s actually an important distinction to make. Monitoring tends to focus on the overall health of system and business metrics—questions we <em>know</em> in advance. Observability is about providing more granular insights into the behavior of systems and richer context. It’s the difference between “post hoc” versus “ad hoc.”</p><p><img loading=lazy src=/wp-content/uploads/2019/10/monitoring_vs_observability-1024x528.png></p><p>In the top-right corner, we have <em>known knowns</em>. These are things of which we have a high degree of understanding and a large amount of data on, i.e. the things we are aware of <em>and</em> understand. For example, “the system has a 1GB memory limit.” As the designers of this system, this is something that we’re acutely aware of and understand. We <em>know</em> that we know how much memory the system can use before it moves outside of its operating boundaries and bad things happen.</p><p>In the bottom-right corner, we have <em>known unknowns</em>. These are things we are generally aware of but don’t necessarily understand. For example, “the system exceeded its memory limit and crashed, causing an outage.” As system designers, memory usage is something we <em>know</em> is important and affects system behavior. We can monitor it in production in order to gather lots of data on it, but just having that data often doesn’t help us to understand <em>why</em> memory is being consumed or even how that data manifests itself as system behavior.</p><p>In the top-left corner, we have <em>unknown knowns</em>, which are things we understand but are not completely aware of. This sounds like a strange, almost oxymoron-like categorization, but it’s basically the things that are gut instinct or intuition. It’s often things we know or <em>think</em> we know without even consciously realizing it. For example, “we implemented an orchestrator to ensure the system is always running.” Intuition tells us that if the process isn’t running, the system isn’t available, so we make sure that it gets restarted when something goes wrong. We might, however, be unaware of the unintended side effects of this decision, and it might be based more on theory and conjecture than data.</p><p>Which leads us to the bottom-left corner: <em>unknown unknowns</em>. These are the things we are neither aware of <em>nor</em> understand. The events we can’t even predict or foresee happening because if we could foresee them, they wouldn’t be unknown unknowns, they’d be <em>known</em> unknowns. For example, “instances churn because the orchestrator restarts the process when it approaches its memory limit, causing sporadic failures and slowdowns.” This was an unforeseen consequence of our orchestrator implementation. As a result, we could not have tested for it or looked for it with our monitoring tools. Instead, it’s something that happens, we learn from it, and quickly classify it as a known unknown—something we know to look for going forward.</p><p><img loading=lazy src=/wp-content/uploads/2019/10/monitoring_vs_observability_overlay-1024x539.png></p><p>In a sense, the known knowns are <em>facts</em>, the known unknowns are <em>hypotheses</em>, the unknown knowns are <em>assumptions</em>, and the unknown unknowns are <em>discoveries</em>. Through this lens, the distinction between observability and monitoring becomes clear. Monitoring is about <em>testing hypotheses</em> and observability is about <em>exploring new discoveries</em>. We monitor known unknowns because these are the things we know to look for, but unknown unknowns are, by definition, <em>unpredictable</em>. We cannot monitor them because we do not know to even look for them in the first place! Instead, we ask questions of our systems in order to understand and categorize these unknown unknowns. Observability is the ability to interrogate our systems <em>after the fact</em> in a data-rich, high-fidelity way. Monitoring, on the other hand, is <em>before the fact</em> and much lower fidelity. These are the dashboards and alerts we set up which usually consist of pre-aggregated metrics. This is what I mean by <em>post hoc</em> versus <em>ad hoc</em>. Observability allows us to ask arbitrary questions of our systems, not questions predefined in advance.</p><p><img loading=lazy src=/wp-content/uploads/2019/10/monitoring_and_observability-1024x537.png></p><p>With this definition, monitoring is a subset of observability, and observability encompasses many different types of data. For example, things like distributed traces, application logs, system logs, audit logs, and application metrics are all important observability signals. But when we boil it all down, it turns out <em>everything is really just events</em>, of which we want different lenses to view. Some of this data provides context for the event itself, such as logs and metrics, and some of it describes relationships <em>between</em> events, such as traces. It’s important we have a way to collect all this context and store it such that we can query and analyze it using these different lenses. Aggregated metrics alone aren’t enough—they don’t have the granularity nor the context needed. Dashboards are simply answers to specific questions known in advance. Observability needs to go much deeper than this.</p><p>In <a href=https://bravenewgeek.com/microservice-observability-part-2-evolutionary-patterns-for-solving-observability-problems/>part two</a> of this series, we’ll revisit the concept of an <a href=https://bravenewgeek.com/the-observability-pipeline/>observability pipeline</a> as a tactical approach to implementing observability in a microservice environment. As part of this, we’ll discuss some steps that can be taken to incrementally improve observability while iterating toward this pattern.</p></div><footer class=post-footer><ul class=post-tags><li><a href=https://bravenewgeek.com/tag/cloud/>Cloud</a></li><li><a href=https://bravenewgeek.com/tag/cloud-native/>Cloud-Native</a></li><li><a href=https://bravenewgeek.com/tag/debugging/>Debugging</a></li><li><a href=https://bravenewgeek.com/tag/devops/>Devops</a></li><li><a href=https://bravenewgeek.com/tag/microservices/>Microservices</a></li><li><a href=https://bravenewgeek.com/tag/monitoring/>Monitoring</a></li><li><a href=https://bravenewgeek.com/tag/observability/>Observability</a></li><li><a href=https://bravenewgeek.com/tag/observability-pipeline/>Observability Pipeline</a></li><li><a href=https://bravenewgeek.com/tag/ops/>Ops</a></li></ul><nav class=post-nav aria-label="Adjacent posts"><a class=post-nav-link href=https://bravenewgeek.com/microservice-observability-part-2-evolutionary-patterns-for-solving-observability-problems/><span class=post-nav-dir><span class=pager-arrow>←</span> newer</span>
|
||
<span class=post-nav-title>Microservice Observability, Part 2: Evolutionary Patterns for Solving Observability Problems</span>
|
||
</a><a class="post-nav-link post-nav-right" href=https://bravenewgeek.com/whats-going-on-with-gke-and-anthos/><span class=post-nav-dir>older <span class=pager-arrow>→</span></span>
|
||
<span class=post-nav-title>What’s Going on with GKE and Anthos?</span></a></nav></footer><section class=wp-comments><h2>Comments</h2><p class=wp-comments-notice>Comments are from this blog's WordPress era and are preserved read-only.</p><article class=wp-comment><header><span class=wp-comment-author>Ruaan erasmus</span>
|
||
<time class=wp-comment-date>October 8, 2019</time></header><div class=wp-comment-body><p>When is Part 2 coming?</p></div><div class=wp-comment-replies><article class=wp-comment><header><span class=wp-comment-author>Tyler Treat</span>
|
||
<time class=wp-comment-date>October 11, 2019</time></header><div class=wp-comment-body><p>Hopefully in the next month or so, stay tuned.</p></div></article></div></article><article class=wp-comment><header><span class=wp-comment-author>Sub</span>
|
||
<time class=wp-comment-date>October 10, 2019</time></header><div class=wp-comment-body><p>Good</p></div></article><article class=wp-comment><header><span class=wp-comment-author>Brandt Tullis</span>
|
||
<time class=wp-comment-date>October 11, 2019</time></header><div class=wp-comment-body><p>Fantastic read, thanks! I appreciated the breakdown of “up/down” and the distinction between observability and monitoring.</p></div></article><article class=wp-comment><header><span class=wp-comment-author>Jean-Mark Wright</span>
|
||
<time class=wp-comment-date>January 6, 2022</time></header><div class=wp-comment-body><p>This was an exceptional article on the topic and really helped me visualize the difference between observability and monitoring. I’m anxious to digest part two! Thanks!</p></div></article></section></article></main><footer class=site-footer><div class="wrap site-footer-inner"><div class=footer-meta><span class=footer-copy>© 2026 Tyler Treat</span>
|
||
<span class="footer-sep footer-dot">·</span>
|
||
<span class=footer-links><a href=/feed/>rss</a>
|
||
<span class=footer-sep>·</span>
|
||
<a href=https://github.com/tylertreat target=_blank rel="noopener noreferrer me">github</a>
|
||
<span class=footer-sep>·</span>
|
||
<a href=https://www.linkedin.com/in/ttreat/ target=_blank rel="noopener noreferrer me">linkedin</a></span></div></div></footer><a href=#top id=top-link class="top-link hidden" aria-label="go to top" title="Go to Top (Alt + G)" accesskey=g><svg viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="feather feather-chevrons-up"><polyline points="17 11 12 6 7 11"/><polyline points="17 18 12 13 7 18"/></svg>
|
||
</a><script>let menu=document.getElementById("menu");if(menu){const e=localStorage.getItem("menu-scroll-position");e&&(menu.scrollLeft=parseInt(e,10)),menu.onscroll=function(){localStorage.setItem("menu-scroll-position",menu.scrollLeft)}}document.querySelectorAll('a[href^="#"]').forEach(e=>{e.addEventListener("click",function(e){e.preventDefault();var t=this.getAttribute("href").substr(1);window.matchMedia("(prefers-reduced-motion: reduce)").matches?document.querySelector(`[id='${decodeURIComponent(t)}']`).scrollIntoView():document.querySelector(`[id='${decodeURIComponent(t)}']`).scrollIntoView({behavior:"smooth"}),t==="top"?history.replaceState(null,null," "):history.pushState(null,null,`#${t}`)})})</script><script>var toplink=document.getElementById("top-link");window.onscroll=function(){const e=window.innerHeight;document.body.scrollTop>e||document.documentElement.scrollTop>e?toplink.classList.remove("hidden"):toplink.classList.add("hidden")}</script><script>document.getElementById("theme-toggle").addEventListener("click",()=>{const e=document.querySelector("html");e.dataset.theme==="dark"?(e.dataset.theme="light",localStorage.setItem("pref-theme","light")):(e.dataset.theme="dark",localStorage.setItem("pref-theme","dark"))})</script><script>document.querySelectorAll("pre > code").forEach(e=>{const n=e.parentNode.parentNode,t=document.createElement("button");t.classList.add("copy-code"),t.innerHTML="copy";function s(){t.innerHTML="copied!",setTimeout(()=>{t.innerHTML="copy"},2e3)}t.addEventListener("click",t=>{if("clipboard"in navigator){navigator.clipboard.writeText(e.textContent),s();return}const n=document.createRange();n.selectNodeContents(e);const o=window.getSelection();o.removeAllRanges(),o.addRange(n);try{document.execCommand("copy"),s()}catch{}o.removeRange(n)}),n.classList.contains("highlight")?n.appendChild(t):n.parentNode.firstChild==n||(e.parentNode.parentNode.parentNode.parentNode.parentNode.nodeName=="TABLE"?e.parentNode.parentNode.parentNode.parentNode.parentNode.appendChild(t):e.parentNode.appendChild(t))})</script></body></html> |