SRE weekly 所有文章
This commit is contained in:
@@ -0,0 +1,26 @@
|
||||
# That Sinking Feeling (The #HugOps Song)
|
||||
|
||||
- **期号**: SRE Weekly Issue #278(2021-07-11)
|
||||
- **作者**: Forrest Brazeal
|
||||
- **链接**: https://www.youtube.com/watch?v=rK_7ozvm53o
|
||||
|
||||
## 简介
|
||||
|
||||
Whoa. This is the best thing ever. I feel like I want to make this the official theme song of SRE Weekly.
|
||||
|
||||
## 正文
|
||||
|
||||
About
|
||||
Press
|
||||
Copyright
|
||||
Contact us
|
||||
Creators
|
||||
Advertise
|
||||
Developers
|
||||
Terms
|
||||
Privacy
|
||||
Policy & Safety
|
||||
How YouTube works
|
||||
Test new features
|
||||
NFL Sunday Ticket
|
||||
© 2026 Google LLC
|
||||
@@ -0,0 +1,13 @@
|
||||
# r/WallStreetBets Incident Anthology (What Worked Edition): Autoscaler
|
||||
|
||||
- **期号**: SRE Weekly Issue #278(2021-07-11)
|
||||
- **作者**: Fran Garcia — Reddit
|
||||
- **链接**: https://www.reddit.com/r/RedditEng/comments/o4ygp0/rwallstreetbets_incident_anthology_what_worked/
|
||||
|
||||
## 简介
|
||||
|
||||
Their auto-scaling algorithm needed a tweak. Before: scale up by N instances. After: scale up by an amount proportional to the current number of instances.
|
||||
|
||||
## 正文
|
||||
|
||||
> ⚠️ 抓取失败:trafilatura returned empty
|
||||
@@ -0,0 +1,79 @@
|
||||
# The Incident Review: 4 Incidents in Outer Space
|
||||
|
||||
- **期号**: SRE Weekly Issue #278(2021-07-11)
|
||||
- **作者**: JJ Tang — Rootly
|
||||
- **链接**: https://rootly.io/blog/the-incident-review-4-incidents-in-outer-space
|
||||
|
||||
## 简介
|
||||
|
||||
> here’s a look at incidents and reliability challenges that have occurred in outer space, and what SREs stand to learn from them.
|
||||
|
||||
## 正文
|
||||
|
||||
Managing incidents like network outages and server failures is tough enough when they happen in a data center.
|
||||
|
||||
But when those failures occur in outer space, the challenges can become stratospheric. When you can’t rely on conventional monitoring and management tools, and even physical access is impossible, you need to think more creatively about incident management.
|
||||
|
||||
With that reality in mind, here’s a look at incidents and reliability challenges that have occurred in outer space, and what SREs stand to learn from them.
|
||||
|
||||
## Computer failure on the Hubble Telescope
|
||||
|
||||
Deployed in 1990, the Hubble Space Telescope has been snapping pictures and collecting other data from outer space for decades.
|
||||
|
||||
But that all stopped in mid-June, when the telescope’s “payload computer,” which manages data collection devices, [stopped responding](https://www.cnn.com/2021/06/24/world/hubble-space-telescope-problem-nasa-scn/index.html).
|
||||
|
||||
At first, NASA engineers attempted the most generic of mitigations: They simply rebooted the computer. That resolved the issue briefly, but it recurred shortly after.
|
||||
|
||||
As of press time, engineers are still investigating the incident, although NASA reports that it thinks hardware failure is likely at fault. If the failed computer can’t be brought back online, the telescope can default to a backup computer that was installed in 2009, but hasn’t actually been used in production.
|
||||
|
||||
From a reliability management perspective, this incident is interesting in that it reflects both foresight and lack of planning on NASA’s part. The fact that NASA installed a backup computer is great, and it may well turn out to be critical for keeping the $4.6 billion telescope operational. Having a backup system already in place is especially important in this context because planning a space mission to deploy a new replacement computer could take years.
|
||||
|
||||
On the other hand, the assessment process that engineers have followed to troubleshoot problems with the primary computer seems a bit slow and disorganized. It doesn’t appear that NASA had a playbook in place for working through an incident like this.
|
||||
|
||||
Of course, when you’re dealing with a totally unique device like the Hubble Telescope, you can’t expect a playbook for everything.
|
||||
|
||||
## When the power goes out in space
|
||||
|
||||

|
||||
|
||||
When the power fails in your data center, you can fall back to generators.
|
||||
|
||||
The fix is not so simple when electricity fails on the International Space Station, which has lost power at least partially on a number of occasions in recent years: [2015](https://www.spaceflightinsider.com/missions/iss/iss-encounters-power-failure-cant-be-fixed-until-2016/), [2019](https://www.theverge.com/2019/4/30/18523864/nasa-international-space-station-spacex-power-channel) and [2020](https://room.eu.com/news/systems-shutdown-as-iss-suffers-from-power-and-oxygen-failures).
|
||||
|
||||
None of these incidents turned out to be fatal to Space Station crew members or cause permanent damage. But they did disrupt many of the Station’s operations. For example, the 2019 incident prevented a SpaceX cargo device from launching as scheduled.
|
||||
|
||||
Recurring power issues have also contributed to calls for the Space Station to be decommissioned, although its future seems to be [safe for now](https://www.washingtonpost.com/technology/2020/12/23/space-station-replace-biden/).
|
||||
|
||||
While SREs no doubt wish that the Space Station had a more reliable power supply, you can’t really fault engineers on this point. It’s not as if power generation in space is easy (the Space Station uses solar panels, but they are managed by complex infrastructure that is prone to failure), and when something does go wrong, obtaining the supplies to fix it is no simple task. But these are the types of reliability issues that will need to be solved if humans are ever to conquer the final frontier of space definitively.
|
||||
|
||||
## Latency in space
|
||||
|
||||
Minimizing network latency is hard enough when your users are just hundreds of miles from your data center.
|
||||
|
||||
But what if they’re 22,000 miles away, as is the case for astronauts in orbit? You end up with some [pretty big latency issues](https://www.theatlantic.com/technology/archive/2015/06/the-internet-in-space-slow-dial-up-lasers-satellites/395618/), it turns out, due to the sheer physical distance that packets must travel to power communications between Earth and devices like the Space Station.
|
||||
|
||||
The good news is that Internet bandwidth, apparently, is relatively good in space -- so good that astronauts can [video chat with their families](https://ny.pbslearningmedia.org/resource/learn-how-astronauts-communicate-in-space/smithsonian-science-starter/). It’s just ping times that are bad.
|
||||
|
||||
This is a challenge that NASA is already working to solve by switching to a laser-based connectivity system, which will deliver much lower latency rates than the satellite-based system astronauts currently use. But until then, space remains a prime reminder that even when bandwidth is excellent, latency problems can lead to a pretty poor end-user experience.
|
||||
|
||||
## The lights go out on the Mars Rover
|
||||
|
||||

|
||||
|
||||
*Opportunity*, a Mars rover launched in 2004, used a pretty obvious strategy for generating power: Solar panels. For years, those panels kept *Opportunity* happily roving across Mars’s surface, collecting all manner of scientific data.
|
||||
|
||||
But in 2018, something happened that engineers hadn’t fully counted on: A major dust storm caused the device’s solar panels to fail. Although system architects had planned for this event by programming the rover to go into hibernation mode in the event of power failure, *Opportunity* never woke back up. In 2019, NASA officially [declared the mission over](https://laist.com/news/jpl-mars-rover-opportunity-battery-is-low-and-its-getting-dark).
|
||||
|
||||
To be sure, *Opportunity* had a good run, and it’s hard to plan for every type of incident that may threaten operations when you’re dealing with a device roaming across the Red Planet. Still, one might wish that the rover had been designed to be just a little more robust in the event of a major dust storm.
|
||||
|
||||
## Conclusion
|
||||
|
||||
Incidents in outer space may seem like the stuff of science fiction -- and they are, for most SREs. But for engineers tasked with managing those systems that humans have deployed into orbit or other planets, incident response is just as important as it would be in any earthly setting.
|
||||
|
||||
## The Incident Review - Previous Posts
|
||||
|
||||
- [The Incident Review: 4 Times When Typos Brought Down Critical Systems](https://rootly.io/blog/the-incident-review-4-times-when-typos-brought-down-critical-systems)
|
||||
- [The Incident Review: 4 Odd Incidents Caused by Animals](https://rootly.io/blog/the-incident-review-4-odd-incidents-caused-by-animals)
|
||||
|
||||
|
||||
{{subscribe-form}}
|
||||
@@ -0,0 +1,66 @@
|
||||
# Prepare for overnight success — with the right load testing approach
|
||||
|
||||
- **期号**: SRE Weekly Issue #278(2021-07-11)
|
||||
- **作者**: Cortex
|
||||
- **链接**: https://www.getcortexapp.com/post/prepare-for-overnight-success-with-the-right-load-testing-approach
|
||||
|
||||
## 简介
|
||||
|
||||
This one includes 3 key things to remember while load testing. My favorite: test the whole system, not just parts.
|
||||
|
||||
## 正文
|
||||
|
||||
What startup hasn’t dreamed about going viral and seeing their usage skyrocket overnight? However, when companies are small, teams tend to focus all their energy on improving their product — preparing for a sudden surge of usage is mostly an afterthought. But beware: when success comes knocking, sometimes it brings a sledgehammer and takes down your entire site.
|
||||
|
||||
When your site goes viral, you probably want to stay up late celebrating, not frantically trying to revive your crashed app. That’s where **load testing** comes in. It’s a great thing to be thinking about while your organization is still in the early stages. Investing in your strategy now will only continue to pay off as you grow and scale.
|
||||
|
||||
### **Three things most startups get wrong about load testing**
|
||||
|
||||
Every engineer knows that load testing is important. Yet, here’s what load testing often looks like at a startup: It’s one engineer’s personal project, where they run scripts on their local machine to test out some theories on what might be slowing the site down. No surprise, then, that teams are often caught unprepared when they finally get that coveted retweet.
|
||||
|
||||
If you’re interested in improving your load testing, there are three things you can focus on that most startups simply don’t do:
|
||||
|
||||
1. Testing your entire production setup, rather than isolated services.
|
||||
2. Defining clear ownership so that someone is responsible for load testing.
|
||||
3. Achieving cultural buy-in within the team so the right actions are taken to improve performance.
|
||||
|
||||
Let’s dig deeper into each of these points.
|
||||
|
||||
#### **1. Testing your production setup**
|
||||
|
||||
There are two kinds of load testing. The first type of load testing is about identifying basic bottlenecks. For example, is your application I/O-bound? Is it CPU-bound? Which microservices fail when you hammer them with concurrent requests? This type of load testing can be good for testing out a hunch based on where you’ve seen latency on your site. But with everything being tested in isolation, you don’t get the full picture of their performance in production.
|
||||
|
||||
That’s why there’s a need for a second type of load testing that’s production-based. This means seeing testing your entire setup in an environment that resembles production as closely as possible. Here, you’re interested in how the whole system takes on load: What are the dependencies and limitations between upstream and downstream services? How do your [SLIs](https://www.cortex.io/post/the-top-3-mistakes-companies-make-with-slos-slas-and-slis) change as you scale up the simulated traffic?
|
||||
|
||||
[The right SRE tools](https://www.cortex.io/post/a-guide-to-the-best-sre-tools) can help you perform production-based load testing. But there’s another challenge: overcoming the cultural and organizational obstacles that prevent you from taking action on your test results.
|
||||
|
||||
#### **2. Defining ownership**
|
||||
|
||||
While many software engineering functions have a dedicated owner, load testing is one of those gray areas, especially at a startup where there’s no dedicated SRE. Should the QA team handle load testing? Or the cloud infrastructure engineering team? How about DevOps?
|
||||
|
||||
If you don’t make a decision about ownership, load testing probably won’t be prioritized — and your on-call engineers will be left cleaning up the mess when services fail in production. Engineering managers can help protect their team from this outcome by making a proactive decision about who will oversee the setup, automation, and reporting on load testing.
|
||||
|
||||
[Read more about how to drive ownership, especially for your microservices](https://www.cortex.io/post/how-to-drive-ownership-in-microservices-608f4ed42be94de59553581e99032537).
|
||||
|
||||
#### **3. Achieving buy-in**
|
||||
|
||||
You can be doing all kinds of automated load testing on your production traffic, but it won’t help if there’s just one engineer on the team who cares about it. (And unfortunately, this is often the case.) If you own load testing at a startup, part of your job is going to be making sure the results have an impact and add value to your team. Here are some suggestions:
|
||||
|
||||
- Avoid tribal knowledge by creating a single source of truth that the whole team has access to. Make the results objective using clear SLIs that are tied to business goals. ( [See our tips here.](https://www.cortex.io/post/the-top-3-mistakes-companies-make-with-slos-slas-and-slis) )
|
||||
- Make sure your results are visual and actionable. Whether it’s with alerts or a dashboard, try to avoid the default of outputting your statistics to a lonely database somewhere.
|
||||
- Automate wherever possible. Create a job that runs, say, every 24 hours. Even better if it automatically notifies engineers when their services fail to meet certain criteria. This makes the results objective and takes you out of the process (no one likes to be a nag).
|
||||
- Be careful about gating releases. For example, adding a load testing regression check to your CI/CD pipeline could be too heavy-handed and alienate your team. Perhaps what you really need at this stage is a way to track how your services are generally performing over time (and whether you are regressing).
|
||||
|
||||
You’ll find more tips about affecting cultural change in [our guide to improving your influence as an SRE](https://www.cortex.io/post/4-ways-to-improve-your-influence-as-an-sre).
|
||||
|
||||
### **Using Cortex for load testing**
|
||||
|
||||
We’ve seen first-hand how some of our customers are using Cortex to scale up their load testing efforts. For example, one customer runs load tests every day and pushes the previous seven days of rolling results to Cortex at all times. As the reporting mechanism, they use a Cortex **Scorecard** they’ve designated specifically for load testing.
|
||||
|
||||
Scorecards are ways to report on service health at a glance. Customers have used Scorecard notifications to set up reports on their load testing results that go straight to the service owners. The SREs can even set new performance goals and track progress using Cortex’s **Initiatives** feature.
|
||||
|
||||
With Scorecards, the whole team has a dashboard for seeing how the system is performing overall. The new **Graph View** lets you contextually overlay your Scorecard data on the graph of your microservices. This feature is especially useful for a load testing Scorecard: You can easily see all the dependencies and identify bottlenecks at a glance.
|
||||
|
||||
If you want to scale up your load testing efforts as an organization (maybe because your userbase is scaling up) and make this a fully integrated part of your QA and 
|
||||
|
||||
[production readiness process](https://www.cortex.io/post/how-to-create-a-great-production-readiness-checklist) — we might be able to help! You can sign up for a free demo of Cortex [here](https://www.cortex.io/demo).
|
||||
@@ -0,0 +1,52 @@
|
||||
# 4 ways to improve your influence as an SRE
|
||||
|
||||
- **期号**: SRE Weekly Issue #278(2021-07-11)
|
||||
- **作者**: Cortex
|
||||
- **链接**: https://www.getcortexapp.com/post/4-ways-to-improve-your-influence-as-an-sre
|
||||
|
||||
## 简介
|
||||
|
||||
SRE is as much about building consensus and earning buy-in as it is about actual engineering.
|
||||
|
||||
## 正文
|
||||
|
||||
On paper, being an SRE may look nothing like being in sales. But in reality, and especially if you work for a startup, there’s more overlap than you might think. We’d bet that every SRE can remember a time when they had to pull out all the stops to convince the team to adopt new tools and processes. In fact, the best SREs are constantly cultivating their influence in different ways.
|
||||
|
||||
Why is influence so important for SREs? Well, if you ask the typical engineer what they care about, “reliability” probably isn’t at the top of that list. Every single team would agree that reliability is essential, *but —*… and there’s always a “but” — there’s an endless number of bugs in the product to fix, there are P0 customer features to work on, and the site *seems* to be up and running. Left to their own devices, engineers will always find excuses to worry about reliability “another day” and so it never gets prioritized.
|
||||
|
||||
Nobody likes to be the bad cop. OK, maybe some people do, but they’re not usually the kind of people you’d want on your team. These tips will help you improve your influence as an SRE so you can do less nagging, add more value, and accomplish your goals with confidence.
|
||||
|
||||
### **1. Tie your work directly to the business bottom line**
|
||||
|
||||
[Why does your organization have SREs in the first place](https://www.cortex.io/post/what-is-sre-site-reliability-engineering)? You might have even asked yourself this question in a moment of frustration when you couldn’t get anyone to listen to your recommendations. And actually, it’s the exact question you *should* be asking.
|
||||
|
||||
Your organization has SREs because reliability is essential. Arguably more than most work that happens at the company, reliability impacts the business bottom line. (As an oversimplified example, when the site is up, you’re making revenue — and when the site is down, you’re not.) The more that you can show that business value, the more you will improve your influence.
|
||||
|
||||
To highlight the value you bring to the organization, make sure that your [SLOs](https://www.cortex.io/post/the-top-3-mistakes-companies-make-with-slos-slas-and-slis) are tied to business processes and requirements. If you can attach a dollar amount to each [SLO](https://www.cortex.io/post/building-reliable-services-a-guide-to-setting-slos), you’re more likely to have an impact. Even something as seemingly abstract as developer morale can be measured. Think about how often your engineers are paged about issues. Does a 1% difference in uptime mean that you’re paying for an extra 10 hours of engineering time per month to keep the site alive?
|
||||
|
||||
### **2. Build bridges quickly and lower gates slowly**
|
||||
|
||||
Building a relationship with engineering is essential to being influential. When you join an organization that hasn’t traditionally had an established SRE team, if you start trying to kick off a top-down process, you’re probably going to meet a lot of resistance. You don’t want to spend all your time chasing people down and telling them what they have to fix.
|
||||
|
||||
Whether you just joined the organization or are simply looking to have more impact, it’s a great idea to schedule time with each team to sit down and actively listen to what their pain points are. Are they having reliability issues? What do they want to improve, and how can you support them in that process? Approach engineers as a partner and show you care about solving their problems and making their lives easier.
|
||||
|
||||
Finally, be careful about how you implement gates. For example, imagine you want to start enforcing some requirements for each build. It might be tempting to add a check to your pipeline that looks at the Dockerfile and blocks the build if these requirements aren’t met. But that’s just going to create resentment and create an “us vs. them” mentality with engineering. Instead, make an announcement saying that in two months, you’re going to add this gate. And over-communicate why the decision was made, explaining how the gate will help the engineering team build better software and have fewer on-call emergencies.
|
||||
|
||||
### **3. Drive conversations on standardization and automation**
|
||||
|
||||
If you’re joining a newly established SRE team at a startup, the engineering organization is most likely growing fast, and the number of services is skyrocketing. Before you know it, you might have 50 different services playing by their own rules — and you could even be trying to report across multiple APM tools and deployment systems.
|
||||
|
||||
As an SRE, you should seize the opportunity to influence decisions on best practices and tooling while the organization’s systems are still immature. Drive conversations with engineering about [what tools the team should use](https://www.cortex.io/post/a-guide-to-the-best-sre-tools), [what on-call rotation best practices should be implemented](https://www.cortex.io/post/best-practices-for-on-call-rotations), and what you want to incorporate into your [production readiness checklist](https://www.cortex.io/post/how-to-create-a-great-production-readiness-checklist).
|
||||
|
||||
Try to reduce friction at every stage of your process with standards and automation. For example, instead of everyone having their documentation in GitHub or markdown files, you might pull this information into a wiki. And if you want all your services to be reporting on certain SLIs, work with your [DevOps team](https://www.cortex.io/post/best-practices-for-devops-teams) to set up dashboards. This will show how much you value the engineering team’s time, improving your relationships and helping your influence grow.
|
||||
|
||||
### **4. Make your process as self-serve as possible**
|
||||
|
||||
As an SRE, instead of yelling “*don’t shoot the messenger*,” it can be a real life-saver to remove yourself from the messenger role altogether. Instead of having to tell people that their services aren’t meeting requirements, you can become more influential if you give them (or their managers/leadership) the tools to see exactly how they’re doing.
|
||||
|
||||
We built Cortex’s [**Scorecards**](https://www.cortex.io/products/scorecard) to address this need. At a glance, Scorecards automatically grade against reliability initiatives you define. Everyone can quickly see whether requirements are being met and know where improvements need to be made. Scorecards enforce very objective requirements, which helps you save time on evaluating services and get buy-in from the team.
|
||||
|
||||
Cortex also has a new feature, called **Initiatives**, to help your team make tactical progress towards your objectives. With Initiatives, you can set a reliability objective and a deadline, and everyone can see precisely what needs to be done for each service and how each scorecard is tracking towards the finish line. Scorecards and Initiatives have been very empowering for engineering teams.
|
||||
|
||||
|
||||
We hope you can put some of these influence-enhancing tips into practice with your team. And if you’re looking for a solution to help you implement these best practices, manage your microservices, and more, you can sign up for a free demo of Cortex [here](https://www.cortex.io/demo).
|
||||
@@ -0,0 +1,53 @@
|
||||
# NoOps: What Does the Future Hold for DevOps Engineers?
|
||||
|
||||
- **期号**: SRE Weekly Issue #278(2021-07-11)
|
||||
- **作者**: Kentaro Wakayama
|
||||
- **链接**: https://codersociety.com/blog/articles/noops
|
||||
|
||||
## 简介
|
||||
|
||||
The definition of NoOps in this article is more clear than others I’ve seen. It’s not about firing your operations team — their skill set is still necessary.
|
||||
|
||||
## 正文
|
||||
|
||||

|
||||
|
||||
Does NoOps mean the end of the DevOps era? Or is it simply the next step in the progression of DevOps? In this article, we’ll explore this question in detail.
|
||||
|
||||
Kentaro Wakayama
|
||||
|
||||
11 July 2021
|
||||
|
||||
With cloud adoption on the rise, the level of abstraction in application architecture has increased— from traditional on-premises servers to containers and serverless deployments. The focus on automation has also increased to the point where manual intervention is no longer preferred, even for infrastructure-related activities like backups, security management, and patch updates. This desired state equates to a NoOps environment, which involves smaller teams that can manage your application lifecycle. Ideally, in such an environment, the efforts required by your operations team will be eliminated.
|
||||
|
||||
It is beyond debate that DevOps is now deeply integrated into the DNA of all cloud-first organizations and is today more of a norm than a rarity. Cloud applications demand agility, and DevOps delivers it. However, does NoOps mean the end of the DevOps era? Or is it simply the next step in the progression of DevOps?
|
||||
|
||||
The success of DevOps greatly depends on the synergy between your development and operations teams, as it brings together system administrators and developers who would otherwise work in silos. Meanwhile, the process of continuous integration and continuous deployment is crucial in DevOps, which helps in identifying issues early on and avoiding them. This in turn results in the faster delivery of solutions. However, note that the operations team is still very much in the picture. There is still a dependency on the operations team to take care of nitty-gritties like: infra config management, security settings, backups, patch management, etc.
|
||||
|
||||
With the cloud, abstraction and automation are scaling to new heights every day. You name it, and you can have it “as-a-service” in the cloud, be it compute, storage, network, or security, to list just a few. Cloud service providers are also investing heavily in their automation ecosystem. You can easily provision your application components using automation templates or just a few API calls. The ongoing management of these components can be automated as well, meaning less overhead to maintain environments and minimal to no human intervention. This leads us to NoOps—the increased abstraction of infrastructure, tightly integrated with development workflows, that requires no operations team to oversee the process.
|
||||
|
||||
NoOps, a term originally coined by [Forrester](https://go.forrester.com/blogs/11-02-07-i_dont_want_devops_i_want_noops/), aims to improve productivity and deliver results much faster than DevOps. In the ideal scenario, developers never have to collaborate with a member of the operations team. Instead, they can use a set of tools and services to responsibly deploy the required cloud components in a secure manner, including both the code and infrastructure. Managed cloud services, like PaaS or serverless, serve as the backbone of NoOps and leverage CI/CD as their core engine for deployment. Hence, note that not all scenarios fit the bill for NoOps.
|
||||
|
||||
NoOps and DevOps essentially try to achieve the same thing: improve the software deployment process and reduce time to market. But while the collaboration between developers and operations team was emphasized in DevOps, the focus has shifted to complete automation in NoOps. This may sound like a silver bullet, but this new approach comes with both advantages and challenges.
|
||||
|
||||
NoOps shifts the focus to services that are deployable by design without manual intervention. From infrastructure to management activities, the aim is to control everything using code, meaning every component should be deployed as part of the code and maintainable in the long run. NoOps essentially seeks to eliminate the manpower required to support the ecosystem for your code.
|
||||
|
||||
NoOps is best suited for born-in-the-cloud environments that leverage PaaS and serverless solutions. Microservices and API-based application architectures fit the bill perfectly, as they offer fine-grained modularity along with automation. Leading cloud service providers like AWS, Azure, and GCP have a laser focus on providing more services and capabilities in PaaS and serverless, which would help accelerate the adoption of NoOps. The current increase in database-as-a-service, container-as-a-service, and [function-as-a-service](https://azure.microsoft.com/en-in/services/functions/) options in the cloud favor this trend as well, and all of these technologies support extreme automation.
|
||||
|
||||
NoOps also shifts the focus from operations to business outcomes. Unlike DevOps, where the dev team and ops team work together to deliver value propositions to the customer, NoOps ideally eliminates any dependency on the operations team, which further reduces time to market. Again, the focus is shifted to priority tasks that deliver value to customers—in other words, “fast beats slow.”
|
||||
|
||||
In theory, not requiring an operations team to take care of your infrastructure might sound lucrative. However, depending on the level of automation that’s achievable, you might still need them around to take care of exceptions or to monitor outcomes. Expecting developers to take care of this would nullify the benefits of NoOps and take away their focus from delivering business outcomes. It is also not a practical approach, considering that developers don’t necessarily have the required skill sets to address operational issues. For example, consider a Disaster Recovery (DR) scenario. You would still need support from an operations team to invoke the DR plan and switch traffic to the failover site.
|
||||
|
||||
Also, not all environments can transition to NoOps. Hybrid deployments and legacy infrastructures would pose a bottleneck—automation is still possible, but human intervention cannot be entirely eliminated in these cases. While aiming for NoOps, PaaS and serverless could become a limiting factor as well, especially during digital transformation. Additional efforts required to refactor legacy monolithic applications to fit the paradigms of PaaS would be counterintuitive. You would need to carefully evaluate the pros and cons, on a case by case basis, before embarking on the NoOps approach.
|
||||
|
||||
Last but not least is the question of security and compliance. Automated deployments aligned with security best practices will not completely eliminate the need for you to take care of security. Traditionally, there is a segregation of duties between operations and development teams. The operations team works together with the security team to enforce controls that protect applications from threats and vulnerabilities. The operations team, meanwhile, is also responsible for handling Identity and Access Management (IAM) solutions.
|
||||
|
||||
Threat vectors and attack methods in the cloud evolve by the day, and so should your cloud security controls. The same goes for compliance. Not all organizations can delegate that responsibility to a set of automated processes. Reducing, or eliminating, the operations team could result in you needing to increase your investment in a security team to ensure the security and compliance of your environments.
|
||||
|
||||
DevOps is considered more of a journey, not a destination, where the focus is on continuous improvement. It would be safer to say that NoOps is the evolution of DevOps, targeting a perfect end-state of extreme automation. It allows organizations to redirect time, effort, and resources from operations to business outcomes. However, this change cannot happen overnight.
|
||||
|
||||
For NoOps to become a reality, a lot of groundwork is involved. You need to identify the right application stack and managed services in the cloud, i.e., PaaS and serverless, for the transition to happen. You need to bake in the component management, configuration, and security controls to get started. Even then, there would be some loose ends like legacy systems that would take more time and effort to transition or that cannot be transitioned at all. And if there’s even a single legacy system left behind, you would still need someone to take care of its operational aspects.
|
||||
|
||||
In a NoOps world, the role of DevOps engineers also changes, as they get the opportunity to learn new skills and processes required for NoOps. Like DevOps, NoOps is more about the shift in culture and process, rather than technology. Organizations need to be intentional about this shift while staying grounded as to the practicalities of the transition.
|
||||
|
||||
For our latest insights and updates, follow us on [LinkedIn](https://www.linkedin.com/company/codersociety)
|
||||
15
sreweekly/markdown/278/07-systems-observability.md
Normal file
15
sreweekly/markdown/278/07-systems-observability.md
Normal file
@@ -0,0 +1,15 @@
|
||||
# Systems Observability
|
||||
|
||||
- **期号**: SRE Weekly Issue #278(2021-07-11)
|
||||
- **作者**: Nishant Modak and Piyush Verma — Last9
|
||||
- **链接**: https://blog.last9.io/need-for-systems-observability/
|
||||
|
||||
## 简介
|
||||
|
||||
Even though I know what observability is, I got a lot out of this article. It has some excellent examples of questions that are hard to answer with traditional dashboards, and includes my new favorite term:
|
||||
|
||||
> The industrial term for this problem is Watermelon Metrics; A situation where individual dashboards look green, but the overall performance is broken and red inside.
|
||||
|
||||
## 正文
|
||||
|
||||
> ⚠️ 抓取失败:URLError: [SSL: UNEXPECTED_EOF_WHILE_READING] EOF occurred in violation of protocol (_ssl.c:1032)
|
||||
@@ -0,0 +1,21 @@
|
||||
# Controlling a process we don’t understand
|
||||
|
||||
- **期号**: SRE Weekly Issue #278(2021-07-11)
|
||||
- **作者**: Lorin Hochstein
|
||||
- **链接**: https://surfingcomplexity.blog/2021/07/10/controlling-a-process-we-dont-understand/
|
||||
|
||||
## 简介
|
||||
|
||||
> Instead, we should consider the fields there where practitioners are responsible for controlling a dynamic process that’s too complex for humans to fully understand.
|
||||
|
||||
## 正文
|
||||
|
||||
I was attending the [Resilience Engineering Association – Naturalistic Decision Making Symposium](https://easychair.org/smart-program/REA-NDM-FONCSI2021/) last month, and one of the talks was by a medical doctor (an anesthesiologist) who was talking about analyzing incidents in anesthesiology. I immediately thought of Dr. Richard Cook, who is also an anesthesiologist, who has been very active in the field of resilience engineering, and I wondered, “what is it with anesthesiology and resilience engineering?” And then it hit me: it’s about *process control*.
|
||||
|
||||
As software engineers in the field we call “tech”, we often discuss whether we are really *engineers* in the same sense that a civil engineer is. But, upon reflection I actually think that’s the wrong question to ask. Instead, we should consider the fields there where **practitioners are responsible for controlling a dynamic process that’s too complex for humans to fully understand.** This type of work involves fields such as spaceflight, aviation, maritime, chemical engineering, power generation (nuclear power in particular), anesthesiology, and, yes, operating software services in the cloud.
|
||||
|
||||
We all have displays to look at to tell us the current state of things, alerts that tell us something is going wrong, and knobs that we can fiddle with when we need to intervene in order to bring the process back into a healthy state. We all feel production pressure, are faced with ambiguity (is that blip really a problem?), are faced with high-pressure situations, and have to make consequential decisions under very high degrees of uncertainty.
|
||||
|
||||
Whether we are engineers or not doesn’t matter. We’re all operators doing our best to bring complex systems under our control. We face similar challenges, and we should recognize that. That is why I’m so fascinated by fields like cognitive systems engineering and resilience engineering. Because it’s so damned relevant to the kind of work that we do in the world of building and operating cloud services.
|
||||
|
||||
## 2 thoughts on “Controlling a process we don’t understand”
|
||||
@@ -0,0 +1,13 @@
|
||||
# Troubleshooting: A journey into the unknown
|
||||
|
||||
- **期号**: SRE Weekly Issue #278(2021-07-11)
|
||||
- **作者**: Pavlos Parissis — Booking.com
|
||||
- **链接**: https://medium.com/booking-com-infrastructure/troubleshooting-a-journey-into-the-unknown-e31b524fa86
|
||||
|
||||
## 简介
|
||||
|
||||
In this epic troubleshooting story, a weird curl bug coupled with Linux memory tuning parameters led to unexpected CPU consumption in an unrelated process.
|
||||
|
||||
## 正文
|
||||
|
||||
> ⚠️ 抓取失败:HTTP 403
|
||||
@@ -0,0 +1,13 @@
|
||||
# How Back Market SREs prepared for Black Friday
|
||||
|
||||
- **期号**: SRE Weekly Issue #278(2021-07-11)
|
||||
- **作者**: Mathieu Garstecki — Back Market
|
||||
- **链接**: https://medium.com/back-market-engineering/how-back-market-sres-prepared-for-black-friday-5f017f343408
|
||||
|
||||
## 简介
|
||||
|
||||
Learning a lesson from a rough Black Friday in 2019, these folks used load testing to gather hard data on how they would likely fare in 2020.
|
||||
|
||||
## 正文
|
||||
|
||||
> ⚠️ 抓取失败:HTTP 403
|
||||
Reference in New Issue
Block a user