--- title: "A Beginner's Guide to SRE (Episode 119)" channel: Slight Reliability date: 2026-03-24 url: "https://www.youtube.com/watch?v=czA9lK-MP5Q" cover: imgs/cover.jpg description: "This week I repurpose a talk I just did at the JuniorDev meetup in Auckland. If you're new to SRE or observability then this is the talk for you. For the more seasoned listeners, it's a chance to see how my perspective and understanding has changed over the years." language: en --- # A Beginner's Guide to SRE (Episode 119) This week I repurpose a talk I just did at the JuniorDev meetup in Auckland. If you're new to SRE or observability then this is the talk for you. For the more seasoned listeners, it's a chance to see how my perspective and understanding has changed over the years. Welcome to SRE Reliability, the show about human beings and their experiences in SRE, observability, and technology leadership. [music] I'm your host, Stephen Townsend. Welcome back to SRE Reliability. I'm Stephen Townsend, and this is the show where we learn about SRE 2 weeks at a time. A couple of weeks ago, maybe three, four weeks ago, I prepared a 20-minute introduction to SRE and observability for a meetup event called Junior Dev here in Auckland, New Zealand. [00:00:01 → 00:00:34] Uh I spent a bit of time preparing for it, and although it's a very introductory topic, I thought it was worth putting it together into a podcast episode as well. So, I think it's interesting to reflect back over the many years that I've been producing this podcast and working in SRE to varying degrees, to reflect on how my perception and understanding of what SRE is changes over those times. >> [snorts] >> So, this is an introduction which is not particularly detailed. I deliberately made it quite high-level. I described it as if SRE is a a loaf of bread, throwing out a few crumbs to give a flavor of what it's about, rather than trying to be completely scientifically accurate in everything that I said. [00:00:32 → 00:01:24] So, I'm going to share this presentation with you. If you were listening to the audio-only version, you're not really missing out on anything. There are some slides, but it's basically just MS Paint pictures that you've probably seen before, just a couple of new ones. There's no visual aids apart from one point which really support what I'm saying. So, you're not missing out if you just listen. [00:01:22 → 00:01:47] But if you want to, you can go and watch this on YouTube where you can see my screen as I present this as well. So, let's see how this goes. I've got my cue cards with me here which I haven't used for a few weeks, but we'll we'll see how this goes. guys. And I'm going to start with an introduction. [00:01:47 → 00:02:04] So, I'm assuming that if you are listening to this podcast, you probably know who I am. But if you don't, this will be an opportunity just to find out a little bit more about me. So, hi there. My name is Stephen Townsend. I live in Auckland, New Zealand with my family. [00:02:04 → 00:02:18] You can see the picture there if you are looking at the slides. I've been working in technology for about 18 years, and my current role is the SRE and DevOps team lead at a company called Blue Card. And if you hadn't heard of Blue Card before, we provide smart metering services for electricity, for water, and for gas across New Zealand and Australia. I studied computer science at the University of Canterbury from 2002 to 2004, and then I remember thinking, I can't do this the rest of my life. I need to have some adventures. [00:02:18 → 00:02:56] I auditioned for Toi Whakaari, the New Zealand Drama School, based in Wellington, New Zealand. Uh and I got in. I got accepted, and I went there and trained for 3 years to be a professional actor. Clearly, I'm not a professional actor. You wouldn't have seen me in anything other than maybe an extra on James Cameron's Avatar, uh the first movie. [00:02:54 → 00:03:19] But drama school was a really formative time in my life. I learned a lot about myself. I grew a lot as a person, and I met my wife, Natasha. She was in my class, so lots was gained from that experience. However, after being a miserable unemployed actor for about 6 to 8 months, I started looking for work in tech. [00:03:17 → 00:03:40] I wanted to take control of my life, and I wanted to get back into that world. I applied for lots of different jobs, had didn't have much luck, but I did eventually land a job as a junior performance test engineer. I'd never heard of performance testing before, or anything about it, but it turned out to be a wonderful mix of the technicals, writing code, uh analyzing huge amounts of data and finding patterns in that, understanding architecture and systems, understanding business context, even a bit of mathematics when it came to things like workload modeling or understanding data. And I did that for 13 years, and I became a deep specialist in that that space. But I got to a certain point where I was really good at what I was doing, right? [00:03:40 → 00:04:27] I was going to different parts of the world to speak at conferences. I was producing a podcast called Performance Time, and creating articles and content online. But I felt like I'd hit this glass ceiling, where as good as I was at the thing that I was doing, I could see other things in the organization happening that I thought I could have a bigger impact in. And that's when I got first into SRE. I was part of an incubator team at IAG at the time. [00:04:25 → 00:04:53] There were just four of us to start with, and we were tasked with the goal of finding opportunities to bring SRE principles and practices into IAG to improve how we operate software. And I was going from a thing that I was really, really good at and specialist at, into something I knew absolutely nothing about. And honestly, I found it so liberating just to let go of all the expectation, all the ego, just throw it aside, and say, ah, okay. Let's figure this out. And as you probably can tell by now, one of the ways which I learn is public learning. [00:04:53 → 00:05:30] So, I started posting questions online, opinions online, observations, and then that's when I started this podcast, SRE Reliability. The idea was, I know nothing about SRE, come and learn with me. And that seemed to resonate with some people. Over the years, obviously, the podcast has changed. It now is predominantly about guests from around the world coming on to talk about their experiences or ideas or challenges, but it's also expanded into other areas, observability, of course, uh technology leadership, mental health, things which are important to me and important to many people working in this world. [00:05:28 → 00:06:09] So, by the time this is released, it'll be something like 120 episodes. Around about 80 different guests from around the world have been on the podcast, so it's been reasonably successful, and it's something that I enjoy doing, and a way to keep connected to that creative part of myself. So, SRE stands for site reliability engineering. It was started by Google back in around 2003, 2004. Uh as Ben Treynor, who was and maybe still is VP of engineering at Google, said, "SRE is what you get when you treat operations as a software problem staffed by software engineers. [00:06:08 → 00:06:47] " Now, I'm probably teaching you to suck eggs here, but since computers have existed, there's been this divide, this chasm between the people who write the code and develop the software, and the people who operate and support it in production. Dev and ops. That's the term DevOps to try and bring them together. And Google had the made this observation that in the development world, the innovation that was happening was many years ahead of what was happening in operations. So, SRE was created as a way to take all those innovative ideas around automation, simplification, abstraction, and apply them to the world of operations. [00:06:47 → 00:07:24] And thus, SRE was born. SRE is now been adopted by organizations around the world. It means different things to different people and different organizations at different times. Up until recently, I didn't say the word SRE much anymore. I talked about reliability engineering because what I was doing and what many people do is so vastly different to what Google did and does that uh I don't know, it didn't feel right to say SRE. [00:07:24 → 00:07:52] But I am literally titled SRE team lead now, so that's something that I I do say now. My own definition was, and I think I'm going to change it pretty soon, is that SRE is about designing, building, and operating reliable services at scale, and making the operation of those services as simple, effective, and dare I say enjoyable as possible. The other important part of SRE is it decouples the scale of an organization from the number of people required to operate its services. So, you think about the Google context as it was exploding in the early 2000s. You might have a service which has 10 times or or more, 100 times more activity on it 1 year to the next. [00:07:52 → 00:08:37] Are you going to hire 10 or 100 times more engineers to operate it? No. That's not cost-effective, and it just doesn't work because as you add more people, you add more coordination overhead, and it just gets out of control. So, by decoupling those two things, being able to operate more complex, busier services with the same amount of people, that unlocks your ability to scale. Now, I'm going to introduce some of the principles and practices, not all of them. [00:08:35 → 00:09:03] Selfishly, I'm picking the things which are most relevant to my role right now. So, I'm pretty new to Blue Card at the time of speaking about this. I've been brought on board to bring SRE to life and in the most amazing way. This is the first time in my career I've really got to get to just truly do the SRE thing and bring it to life in an organization who's ready for it. So, the areas I'm going to talk about today uh this idea of our culture around incidents, uh service level objectives, observability, which isn't, I guess, technically part of SRE, but it's so intertwined with it, I had to speak about it. [00:09:03 → 00:09:45] Toil, and blameless postmortems. And there's so many more aspects to it around architecture and design and all these different principles and ideas, but I'll just start with these to give you those crumbs, to give you the flavor of what SRE is all about. So, let's start with incident culture. An incident, what is an incident? It's some kind of outage or loss of service. [00:09:42 → 00:10:08] It generally negatively impacts your customers in some way or disrupts your business in some way. And it generally requires some human intervention to do something about it. Traditionally, when we think about incidents, we think of them as a bad thing. We must avoid the incidents. They are bad. [00:10:06 → 00:10:24] Let's just stop them from happening, right? And we've also traditionally had this mindset that if we just engineer our systems perfectly, like just amazingly, they'll be 100% reliable and we'll never have an incident. Now, it turns out that that's just simply not true and becoming increasingly less true as the systems that we work with become more complex and more distributed over time. It turns out that incidents are a normal part of running complex systems. SRE embraces this and says incidents will occur, but when they do, what can we learn from them? [00:10:24 → 00:11:03] What can we learn and implement so that we are more prepared for the future and we handle incidents better or our systems are more robust? The thing about incidents is they're often the time we learn the most about our organizations because we're forced to look at things we may have never thought about looking at before. Uh being exposed to different pathways through your systems or behaviors you didn't expect. We learn not just about our technology, but about how uh our different teams actually interact with each other and our processes. So, that kind of attitude around incidents, I think is an important part of SRE. [00:11:01 → 00:11:38] Speaking about incidents is the concept of a blameless postmortem. What are they? They are a formal way of learning from incidents. Now, essentially, they are a written record of an incident which covers a timeline of different events that occurred, uh what happened, what was the impact, but most importantly, what did we learn from the incident? And what actions can we take to be better tomorrow than we were today around how we respond to incidents or how robust our systems are. [00:11:35 → 00:12:10] Now, a blameless postmortem, you might think sounds just like a post-incident review or PIR from ITIL. Now, in a sense, they are both formal ways to learn from incidents. I would say a blameless postmortem is more collaborative. It has a a stronger focus on psychological safety, a strong important part of SRE. And it also has a stronger focus on learning in the broader sense from incidents rather than trying to find the root cause, which is another principle or idea which is dubious at best. [00:12:08 → 00:12:46] Blameless postmortems are one of the things I've actually managed to get off the ground in the first month or two of Blue Current. It's been scary but rewarding to organize, first of all, to get a written account and to encourage people to contribute to it. Started off with just me writing about these things as I was sitting on the incident calls and chats and and pulling things together. Then other people started contributing to it. And now, for the the big important or tricky incidents where there's the stuff to talk about and learn from, pulling in business and technology stakeholders to come together and talk about it. [00:12:45 → 00:13:21] It's been really rewarding to hear, say, for example, people from our business side who were on these calls saying, "I didn't know that that's how our technology worked before. " You know, that that is the point. That is amazing and that's the kind of thing that I'm trying to bring to life, to build a a better understanding and connection between everyone. SLOs stand for service level objectives and it is just an enormous topic and I genuinely can't do it even remote justice in this short talk that I'm doing today. But I'm going to try and give the essence of what they're about. [00:13:21 → 00:13:53] Now, traditionally, when we set up our monitoring and alerting, we set that up around technical signals. Things like CPU usage or garbage collection, memory, network, disk, uh queue lengths, connection or thread pools, those kinds of things, right? And we set up our alerting based on thresholds. For example, if CPU exceeds 80% for more than 5 minutes, fire an alert to someone to have a look at it. That kind of thing. [00:13:52 → 00:14:22] Now, that kind of that sounds sensible and if you've got a a simple system, it kind of works, right? But there's a couple of major challenges with this. The first one is alert noise, where you just get overwhelmed with alerts smashing you every day which aren't necessarily things you need to act on. So, you think about any non-trivial organization, how many different software systems does it have? And then how many different resources does that hardware and software resources do those technology systems have? [00:14:19 → 00:14:51] You're talking about thousands, hundreds of thousands, millions, more than millions of different resources all with these alerts set up potentially just flooding your engineers with alerts. That's the first problem. The second problem is just because a technical signal looks bad, doesn't necessarily mean that your business is disrupted in any way or the customers are even having a problem. It is entirely possible to have a situation where, say, you've got a simple system, that CPU on a particular web or application server is running at 100% all day and yet your customers don't notice anything. They can do whatever they need to do. [00:14:51 → 00:15:29] Um their experience is still great and your business is not suffering in any way. It is equally likely that you have, say, your beautiful dashboard where you're tracking hundreds of different technical signals and they are all green and yet your customers can't consume your services. So, technical signals are not a great way of actually understanding if there's a thing that you need to respond to, if there's a genuine incident or issue that you need to address. I think SLOs are the answer to that solution. So, we still collect technical signals all the time. [00:15:29 → 00:16:03] We collect all the data because we may need it and it will help us diagnose and resolve issues potentially, right? But in terms of alerting, you forget about all those technical signals and instead set up your alerting based on answering one question continuously. Can my customer effectively consume my service? And if the answer is yes, then you don't need to immediately jump on and do anything. And if the answer is no, then you need to do something because you are very confident that your customer is being impacted by something that needs to be addressed. [00:16:03 → 00:16:38] Now, that's about all I can do in this tiny little uh two or three minute talk about SLOs. I haven't even described what they are really, but just to finish off, SLOs are like an internal objective that you set around how reliable you want your services to be for your customer. You might hear of the term SLI or service level indicator. Those are the signals that we use to track whether we're meeting those service level objectives. And you might hear the term error budget. [00:16:37 → 00:17:05] So, let's say that you want you say, "We want the service to be 99. 9% available. " That leaves you just under 9 hours a year in your error budget, which is time that you have technically said it's okay to be down that much each year. So, if you get towards the end of the year or quarter or however you you measure it and you have had no downtime or very little downtime, you've got this error budget that you can use to run riskier experiments which might have a really big payoff. So, you can use your SLOs and your error budgets to drive innovation by being able to to run these experiments which you might otherwise not have felt safe to do. [00:17:05 → 00:17:47] All right, observability, I'm not sure if it's even technically part of SRE, but it's just so intertwined as I see it that I needed to talk about it. The technical definition is, and I'm going to say this wrong, it's something about being able to understand the inner workings of a system by observing its inputs and outputs. Something like that. That's I've got it wrong. I think the key thing to remember is that observability is a quality of a system. [00:17:45 → 00:18:08] It's not a thing that you do. It's a thing that you can achieve. When we have observability, we have insight into what is happening within our systems so that we can understand what's going on and not only understand, but be able to more readily resolve issues or incidents or to improve our system because we understand it better. When we have observability, we have that insight. Service level objectives, SLOs, answer the first most important question, which is is there a problem that I need to respond to? [00:18:08 → 00:18:39] Are my customers able to effectively consume my services? Observability answers the next question, which is if there is a problem, what's going on? What do I need to do? What do I need to look at? When we have that look, that view under the cover, we have a better understanding of how our increasingly complex and distributed systems actually work. [00:18:39 → 00:19:00] And sometimes they work in ways we never would have expected before. It empowers more people to be able to, for example, handle incidents because it no longer depends on an engineer who's been in an org for 20 years and has all the context about everything going on. Because when you have observability, everyone can see quite clearly exactly what's going on and where the problem is located. It allows you to respond to incidents, investigate them, and remediate them much, much faster because you've got clear data which shows you exactly what's going on. There's no more guesswork involved. [00:19:00 → 00:19:36] And um this is a terrible diagram if you're watching the YouTube video. It's the only sort of supporting visual here, but it's just a a diagram trying to show distributed tracing, which I think is the quintessential technique for achieving observability. The idea being that you've got a customer somewhere and they click a button to do something, you're able to trace that flow of requests throughout your system and see exactly what's triggered and when, how long it took, what errors occurred, and any other useful metadata that might help us understand exactly what's happening in a really visual way. Toil is manual, repetitive, tedious work, and operations is rife with toil. Think about you copying and pasting from one file to another or running commands and doing a bunch of forms to make a thing happen. [00:19:35 → 00:20:30] Toil is bad because it costs a lot of money. You often got these very senior engineers who are clicking through and doing this manual tedious stuff. That's expensive, but even more so is the opportunity cost. You could be getting those senior engineers to be working on engineering work, which is going to bring you forward, make your services more reliable, help you scale your organization, or even improve your products. When you have a high amount of toil, your organization isn't able to scale because as you increase the complexity of your services or and add new services, and as you scale up the volume of activity in the business happening, if you're dependent on toil to support that, you're going to have to hire more and more and more people and then have more and more complexity to try and and coordinate all those people, and that's not effective. [00:20:28 → 00:21:18] And toil also leads to potentially unhappy engineers. It's not necessarily satisfying work. It's not helping engineers learn and improve and grow. So, automating and simplifying our processes and ways of working in our technology, uh removing unnecessary or waste wasteful processes and things happening in our systems is a core part of SRE, and it helps not only to uh operate our systems simpler and more effectively and help us scale, it helps with staff retention as well because at least they're happier engineers and more satisfying jobs. So, that's my super fast introduction to SRE and observability. [00:21:16 → 00:21:59] If you want to learn more, there are a ton of books by Google uh that I know the first three here are completely free online if you want to read them. There's nothing stopping you going and learning from them. I'm going to be honest, I've only read the first book cover to cover. Uh I should probably read them some more, especially now that I've got this new role. I think you need to purchase Enterprise Road Map to um to SRE, but that's something I am looking through this year. [00:21:59 → 00:22:27] There are podcasts. There's this podcast here, uh but this I have a particular style. It's not particularly technical. It's more about the human side of SRE and the ideas, but there's other podcasts including Google's own Google SRE podcast. Nice play on words there. [00:22:24 → 00:22:43] But one of the ways that I've learned the most about SRE is just by connecting and following with experts around the world. So, Niall Murphy was one of the co-authors of the original SRE book when he was at Google. And then you've got practitioners like Sebastian and Collette and Amin as well. I I like the way that they think about the work and and the sort of very practical way that they approach it, and I've interacted with them. They've helped me with my thinking. [00:22:42 → 00:23:09] And uh Collette is a co-host of another podcast called This is Fine about resilience engineering, which is a ton of useful thinking that can help with SRE work. So, I recommend checking that out. So, that was my super fast introduction to SRE and observability. I do want to say in my new role, I am incredibly excited about genuinely implementing SRE practices from scratch. I'm actually getting to do it. [00:23:09 → 00:23:39] It's very exciting and I'm facing some interesting challenges. So, I'm planning on doing quite a few more solo episodes this year to explore these different questions and how we might bring SRE and some technology leadership ideas to life as well. Another thing I wanted to mention is that for the first time in a while, I don't have a large backlog of guests for the show. I've got one interview recorded and one guest I'm thinking of approaching, but if you are working in the realm of SRE, particularly if you are implementing it on the front line, I'd love to hear from you if you'd love to come on the show. So, just reach out to me on LinkedIn or Instagram or wherever you follow me. [00:23:36 → 00:24:21] Keeping in mind that this is a show about human beings and how they grapple with this world of SRE, rather than just technical details or particular tools or products. That's not what it's about. So, if that sounds like something you'd love to do, reach out to me. But otherwise, I hope you enjoyed this episode and I hope you have a wonderful fortnight. I'll see you again next time. [00:24:20 → 00:24:44] >> [music] [music] [music] [music] [00:24:42 → 00:25:03]