Insights into a Product SRE team at LinkedIn
The SRE role at LinkedIn lies at the intersection of software development and systems engineering. As SREs, our primary goal is to keep the site up and running at all times. Here at LinkedIn, we have two main types of SREs: embedded and non-embedded. In this blog post, we will focus on the former, which is more integrated and embedded with development teams.
Through this blog post, we will build on a previous SRE culture article, Building the SRE Culture at LinkedIn, while drawing from our experience as SREs in Product, specifically the Ads space. We also want to give an insider look into the day-to-day of Product SREs and a glimpse of the types of problems that we try to solve for any prospective candidates interested in joining SRE at LinkedIn.
SRE’s role is mission-critical
LinkedIn has thousands of deployments and feature ramps every week for thousands of various microservices, all of which must maintain the three nines of availability. In addition to our feed, resume, and job search functionalities for members, LinkedIn also provides products for recruiters and sales teams through Recruiter and Sales Navigator and tools for advertisers through Campaign Manager. This means that if LinkedIn were to have an outage, it would not only impact the 700+ million members who use LinkedIn to find jobs and build their networks but also impact individuals whose jobs are dependent on our enterprise tools. Outages can have both short-term and long-term impacts on the member experience, company brand, and more. Our Ads SRE team deals directly with revenue-impacting incidents, which adds a layer of complexity to our outages.
As a result, there's a lot of responsibilities for SREs in operating the world’s largest professional network. We tackle these challenges by working hand in hand with the other engineering teams to design, build and run large-scale systems that are reliable, efficient, and scalable. In fact, the mission for Product SRE is to
“engineer and drive product reliability by influencing architecture, providing tools, and enhancing observability.”
Each quarter, we align our project planning to this mission statement. We’ll discuss below what some of these projects look like.
SREs are problem solvers
One thing you will notice in engineers who are drawn to SRE is the drive to tackle extremely challenging problems with large-scale systems: for example, the majority of the products that we support in LinkedIn Ads rely on dozens of downstream micro-services to fetch and process data. The overall reliability of these pipelines from beginning to end falls on the hands of Product SREs to ensure the entire workflow is observable and monitored end to end.
Being embedded with our development counterparts means that we have a front-row seat to all of the new features in development and that we get to be an active participant in launching new services and products from the design phase up to the initial deployments, scaling, and ramping process. All of this gives us exposure and intimate familiarity with the various backend and frontend services, datastores, and offline and real-time pipelines involved.
Since our Ads SRE team has a primary on-call rotation for these products, we are able to use the knowledge acquired over time in building and launching the products to more efficiently respond to incidents, and reduce TTR (time to resolve) and TTD (time to detect) for various production issues. We strongly believe that this would not have been possible without the embedded SRE model.
SREs and software engineers foster strong relationships
In order to make stable and robust systems a culture, we work hard to establish good relationships between SREs and application developers. My team does this by having embedded SREs on developer teams in order for SREs to have a good general picture of what is being released, and to know when to step in and provide guidance.
It is everyone’s responsibility to ensure good engineering craftsmanship, but it’s our role as SREs to provide additional perspectives on scaling methods, different system design trade-offs, alternative solutions, and potential architectural pain points. If SREs only step in once the service is ready to launch, it’s either too late in the process to make useful contributions or you may face the additional headaches of being in firefighting mode.
SREs love automation
One of the key assumptions we accept as part of our role is that system failures are inevitable. Therefore, we try to get ahead of these failures by leveraging tooling and automation whenever possible.
Recommended by LinkedIn
As SREs, we sometimes work on problems that aren’t always well-defined. We are usually focused on the whole production stack, ensuring the reliability of systems used by a wide range of internal and external customers. As a result, SREs find that they have a lot of freedom to choose what to work on. This combination of freedom and broad responsibility means that it’s on us to prioritize the projects that will have the biggest impact since we are but a small team of engineers.
As new features are rolled out and the systems that our service relies on change, things can break in new and unexpected ways. When this happens, the number one goal is to provide a working service for all members, whether that means redirecting the traffic to a working datacenter, rolling back to a working version, or implementing a hotfix.
In the short-term, it is ok to manually solve problems, but we also keep in mind that a manual approach isn’t scalable. If we’ve dealt with a problem before, it’s our responsibility to take that pain and convert it into fuel to solve the problem so it doesn’t come back. In fact, a key part of the SRE role is building tooling and processes that automatically fix the most common issues.
SREs embrace reflection and postmortems
Not catching an issue is a loss. Not doing anything to prevent an issue from occurring again is even worse. That’s why whenever there’s a service outage, we always hold a postmortem, during which the on-call SREs and key stakeholders discuss what went wrong and how we can prevent it from happening again. In these meetings, the on-call SRE comes prepared with a timeline of the incident and a root cause analysis, which outlines the systems that impacted each other and contextualize the relevant metrics as evidence for why the outage occurred. This allows us to ask larger questions about the process, capacity, tooling, and/or cultural failures.
However, we can’t always automatically account for and fix a problem just because we know how it manifests. Some complex solutions can take an unjustifiable amount of engineering work to implement, or the problem can arise because of a drawback with the tech stack itself. In these scenarios, we switch focus to how we can detect the problem sooner and reduce the time it takes to fix it.
We have an on-call retrospective tool that gives us a framework for providing feedback on the usefulness of any pages we receive during our rotations so that we can reduce the number of inefficient pages. It also allows us to keep track of how many times we’re getting paged, so we can look at the overall trends to see whether we need to be doing more to keep our services reliable.
SREs combine a diverse skillset
It's no coincidence that SREs at LinkedIn are holistic in both their skills and interests; SREs are known to be knowledgeable in a healthy combination of coding, Linux systems, and networking. SREs at LinkedIn are seen as a subset of software engineers with strong coding skills to write maintainable and efficient code. They are also expected to be knowledgeable in architecture design and have an in-depth understanding of Linux systems and running large-scale distributed systems.
LinkedIn SREs have recently published a School of SRE curriculum, that is focused on onboarding our non-traditional hires and new college grads into the SRE role. The topics covered range from Linux fundamentals and networking to Python and big data concepts as well as system design and security.
As embedded SREs within product teams, one of the key prerequisites aspects of our success is having an in-depth view of our supported products and understanding how they’re used by members. This helps us ensure that we’re monitoring the most critical aspects of the member experience and ensuring that appropriate alerts and monitoring are set up for all critical services and endpoints.
We’re also involved in reviewing and signing-off on RFCs and design documents for any new features or upcoming projects. Building strong relationships with our development partners is key as part of our responsibility to provide guidance and advise our partners on the best practices for production readiness and capacity planning, etc. All of this allows us to be more effective in our roles.
As a result of this dynamic, you will find SREs with a wide range of expertise, ranging from in-depth database performance to video transcoding to frontend development. Some of us are proficient in Python, others in Java or Go, etc. Some of our team projects have involved learning Ember.js and making changes to front-end applications to capture additional metrics for user behavior. We’ve found that being an SRE at LinkedIn means that you’re not boxed into a set of tasks; on the contrary, you have space to find your niche or specialty.
Conclusion
Keeping the site working is a shared responsibility across all engineering; however, SREs are expected to have specialized knowledge as to how to accomplish that. As SRE organizations, such as LinkedIn’s, mature, the focus changes from firefighting to ensuring that we are building robust and reliable systems on a day-to-day basis. It is an exciting time to be an SRE as the reliability of large-scale systems becomes paramount to a lot of organizations. If you want to learn more about what SRE teams are working on at LinkedIn, check out our blog: https://engineering.linkedin.com/blog/topic/sre
Awesome Z!!
Well said Zaina. I will share this with my team.
This is a great article Zaina Afoulki & Lakshmi Namboori. I am particularly very interested about the School of SRE as that is a great pathway for career progression for LinkedIn employee interested in SRE roles