Observablility: Tabs vs. Spaces for Ops
Perhaps there is no greater debate among software engineers than the seemingly endless debate of Tabs vs. Spaces. Even the StackExchange Software Engineering community weighed in with the most votes going to tabs, but the correct answer being marked as spaces!
Divisive debates are not new in IT.
Gif vs. Jif
JASON vs Jay-Sawn vs J-S-O-N (don't spell it out, please!)
A-M-I vs AIM-ee (there is even a t-shirt and an epithet from Corey Quinn aka @QuinnyPig)
VIM vs VI vs Nano vs EMACS
After feeling left out of the debate (VIM vs VI vs Nano vs EMACS aside), TechOps, SREs, and other non-dev operations folks are now embroiled in own controversy.
What observability is not
Charity Majors, CTO at Honeycomb, and I agree that observability is not monitoring. Charity's manifesto defines the monitoring landscape -- "Monitoring is about known-unknowns and actionable alerts." My good friend and SolarWinds HeadGeek, Leon Adato, wrote about actionable alerts in Systems Monitoring for Dummies and again in Monitoring 201, an ebook we collaborated on. (Hint: If you are alerting and neither a system or person is responding, stop alerting!)
Perhaps monitoring is like obscenity as defined by US Supreme Court Justice Potter Stewart who, in a 1964 ruling, stated:
"I shall not today attempt further to define the kinds of material I understand to be embraced within that shorthand description; and perhaps I could never succeed in intelligibly doing so. But I know it when I see it, and the motion picture involved in this case is not that."
While the definition of observability is more ephemeral than many engineers would like, it is certainly not a destination nor is it a product you can purchase. According to the Framework for an Observability Maturity Model,
"The quality of one's observability practice depends upon both technical and social factors. Observability is not a property of the computer system alone or the people alone... If teams feel uncomfortable or unsafe applying their tooling to solve problems, then they won't be able to achieve results."
What exactly is Observability?
To again quote Charity,
"...observability is about unknown-unknowns and empowering you to ask arbitrary new questions and explore where the cookie crumbs take you. Observability means you can understand how your systems are working on the inside just by asking questions from outside." (Observability: A Manifesto)
Charity is indomitable on the topic of observability and often assumes the role of operations futurist, pointing to where we are likely headed even though few are prepared to arrive there today. Observability, like DevOps, SRE, or any other practice or framework, is a journey. Each one of us starts at a different place, uses different modes of transportation, and will experience this journey in a very unique way. However, there are some common elements on our observability journey.
Elements of Observability
Observability is about data, the tooling to ask the questions that need to be asked (whether you know the questions or not), and a culture that supports the continuous improvement loop needed to increase customer satisfaction, whatever that means to your organization.
But what kinds of data do you need? New Relic offers up their M.E.L.T 101 - An Introduction to the Four Essential Telemetry Data Types. If you are trying to complete your journey without at least the first three data types, you probably should have stayed in The Shire, Mr. Baggins!
Metrics
The bedrock of monitoring, metrics are the signals that measure performance. They have a timestamp and a value. Metrics may be raw measures (total CPU % utilization) or may be an aggregate (average CPU % utilization over the last 5 minutes.) Always understand the state of your metrics.
Events
Events are a change in your system from the previous state. This can be a positive change (a user logged in) or a negative change (a node disconnected from the cluster). Unlike metrics, events are never summarized. Discrete events are part of the high-fidelity data mantra of observability.
Logs
Logs are a fire hose of data. (Ironically, PCF refers to their logging mechanism as a firehose and nozzle. A coincidence? Not in my experience!) Logs are as sparse or verbose as defined by the system generating them but, like events, they are also discrete. Logs are never aggregated or summarized in their raw form. (You can aggregated metrics based on the contents of your logs, but those are then log metrics and metrics can be aggregated or summarized.)
Traces
In the debate on observability, traces (aka distributed traces) are the single most controversial item on the list. Traces, well, trace the execution of code and multiple traces form a span that allows us to see exactly how our code performed the tasks associated with the event. When a user logs in, we likely have many discrete actions (AD authentication, handshakes for encryption, maybe a 2FA system), or traces, that would form a span that shows what happened when user joshua.biggley tried to authenticate.
But where do I start?!?
Start by understanding what you want to accomplish and why. If you are using commercial applications then you can collect traces to help you understand the components you control, but you aren't likely to have a vendor volunteer their source code so that you can tweak their app. However, you can collect metrics, events, and logs from that system to give your operations team insights into the environment.
Give teams autonomy to build their own views of the data, to ask their own questions. When New Relic introduced Programmability at FutureStack 2019 in New York City they opened their platform, allowing teams to interact with data on their own terms. If you haven't made friends with a front-end developer (or seven!), now is the time to do it or at least brush up on the basics of ReactJS.
Don't worry about being perfect, the 80/20 rule applies to observability too! The brilliant poet, Maya Angelou, said "I did then what I knew how to do. Now that I know better, I do better." Do what you know how to do. Build on that knowledge. When you know better, do better.
J.
i object to their redefinition of events. cynical and shitty.