SRE weekly 所有文章

This commit is contained in:
2026-09-12 17:23:01 +08:00
parent 409b40ddcb
commit af7633f9dc
8486 changed files with 4489990 additions and 7 deletions

View File

@@ -0,0 +1,17 @@
# SRE@Xero: Managing Incidents Part I
- **期号**: SRE Weekly Issue #97(2017-11-12)
- **作者**: —
- **链接**: https://devblog.xero.com/sre-xero-managing-incidents-part-i-7d02d650a71c
## 简介
Last month, I linked to an article on Xero’s incident response process, and I said:
> I find it interesting that incident response starts off with someone filling out a form.
This article goes into detail on how the form works, why they have it, and the actual questions on the form! Then they go on to explain their “on-call configuration as code” setup, which is really nifty. I can’t wait to see part II and beyond.
## 正文
> ⚠️ 抓取失败:HTTP 403

View File

@@ -0,0 +1,14 @@
# Stretching Spokes
- **期号**: SRE Weekly Issue #97(2017-11-12)
- **作者**: —
- **链接**: https://githubengineering.com/stretching-spokes/
## 简介
Spokes is GitHub’s system for storing distributed replicas of git repositories. This article explains how they can do this over long distances in a reasonable amount of time (and why that’s hard). I especially love the “Spokes checksum” concept.
## 正文
Redirecting…
Click here if you are not redirected.

View File

@@ -0,0 +1,13 @@
# Fly the airplane: Three practices for effective incident response
- **期号**: SRE Weekly Issue #97(2017-11-12)
- **作者**: —
- **链接**: https://www.itproportal.com/features/fly-the-airplane-three-practices-for-effective-incident-response/
## 简介
From the CEO of NS1, a piece on the value of checklists in incident response.
## 正文
> ⚠️ 抓取失败:HTTP 404

View File

@@ -0,0 +1,13 @@
# The Ultimate Guide to Secondary DNS
- **期号**: SRE Weekly Issue #97(2017-11-12)
- **作者**: —
- **链接**: https://dzone.com/articles/the-ultimate-guide-to-secondary-dns
## 简介
Here’s another great guide on the hows and whys of secondary DNS, including options on dealing with nonstandard record types that aren’t compatible with AXFR.
## 正文
> ⚠️ 抓取失败:HTTP 410

View File

@@ -0,0 +1,13 @@
# Availability has a new meaning. And it doesn’t include planned downtime.
- **期号**: SRE Weekly Issue #97(2017-11-12)
- **作者**: —
- **链接**: https://www.itproportal.com/features/availability-has-a-new-meaning-and-it-doesnt-include-planned-downtime/
## 简介
From a customer’s perspective, “planned downtime” and “outage” often mean the same thing.
## 正文
> ⚠️ 抓取失败:HTTP 404

View File

@@ -0,0 +1,70 @@
# Risks of a “serverless” future: dissolving valuable infrastructure
- **期号**: SRE Weekly Issue #97(2017-11-12)
- **作者**: —
- **链接**: https://siliconangle.com/blog/2017/11/08/risks-serverless-future-dissolving-valuable-infrastructure-serverlessconf/
## 简介
“serverless” != “NoOps”
> Willis urges the importance of integration with existing operations processes over replacement. “Serverless is just another form of compute. … All the core principles that we’ve really learned about high-performance organizations apply differently … but the principles stay the same,” he said.
## 正文
![](https://images.siliconangle.com/blogs.dir/1/files/2017/11/John-Willis.jpg) CLOUD
![](https://images.siliconangle.com/blogs.dir/1/files/2017/11/John-Willis.jpg) CLOUD
![](https://images.siliconangle.com/blogs.dir/1/files/2017/11/John-Willis.jpg) CLOUD
### Risks of a ‘serverless’ future: dissolving valuable infrastructure
As technology continues to expand and enable new levels of productivity and growth, the outdated tools and processes that once reigned supreme are naturally cast aside to make room for progress. [John Willis](https://www.linkedin.com/in/johnwillisatlanta/), vice president of DevOps and digital practices at SJ Technologies Inc., is wary of this trend in light of developments towards serverless architecture.
“We make a big mistake to think serverless means we don’t need operations now,” Willis said, adding that he has seen many businesses dissolve valuable infrastructure in their efforts to incorporate this new frontier of tech, a decision he fears will ultimately do a disservice to their functionality.
Willis spoke with Stu Miniman ([@stu](https://twitter.com/stu)), co-host of theCUBE, SiliconANGLE’s mobile livestreaming studio, during the Serverless Conf event in Hell’s Kitchen, New York City. They discussed the issues many businesses have in adapting to serverless architecture, as well as what they can do to ensure a smooth transition and long-term stability. *(* Disclosure below.)*
## ‘Don’t throw the baby out with the bathwater’
Though the draw of total transformation can be attractive to companies looking for an upgrade, Willis urges the importance of integration with existing operations processes over replacement. “Serverless is just another form of compute. … All the core principles that we’ve really learned about high-performance organizations apply differently … but the principles stay the same,” he said.
Willis stresses that this new tech does little to change the fundamentals of ingrained processes or results businesses need to see at every step for full comprehension. Observability, telemetry, repeatable patterns of delivery to ensure there are no code vulnerabilities are all impossible to achieve without the support of ops, according to Willis.
“It’s about supply chain and building repeatable, structured delivery with all the gates and the checks and the units. None of that goes away with serverless, just like it didn’t go away with the cloud … [or] virtualization,” he said.
Without an ops team reinforcing its efforts, Willis fears serverless will soon reveal its foundational cracks. “Serverless is easy to create a function, get it set up, cost effective, but we’re starting to learn all of the complex operational issues,” he said.
Willis sees the human capital side of this work as having a greater utility than just its value to the performance of tech tools. “I will tell you what my definition of ops is: It has really very little to do with technology. It has to do with human capital, how you create high-performing organizations, and the principles and practices that lead to that,” he said.
Though he sees the risks ahead, Willis appears assured that businesses will work wisely in incorporating new technologies. “I think the message is loud and clear that operations still exists; it just has to be thought about,” he concluded.
Watch the complete video interview below, and be sure to check out more of SiliconANGLE’s and theCUBE’s coverage of the [ServerlessConf event](http://siliconangle.tv/serverlessconf-2017/).
##### Photo: SiliconANGLE
# A message from John Furrier, co-founder of SiliconANGLE:
Support our mission to keep content open and free by engaging with theCUBE community. **Join theCUBE’s Alumni Trust Network**, where technology leaders connect, share intelligence and create opportunities.
- **15M+ viewers of theCUBE videos** , powering conversations across AI, cloud, cybersecurity and more
- **11.4k+ theCUBE alumni** — Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network
### Are you an AWS customer?  Support SiliconANGLE financially by buying your AWS services from our Marketplace portal page and links: [https://siliconangle.com/aws-marketplace/](https://siliconangle.com/aws-marketplace/)
##### **About SiliconANGLE Media**
[SiliconANGLE](https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fsiliconangle.com%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=SiliconANGLE&index=9&md5=646b1b564e2259100a2b8638aab0a552),
[theCUBE Network](https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fwww.thecube.net%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=theCUBE+Network&index=10&md5=7de2a85f95ab4a4a495cede20b8cb1da),
[theCUBE Research](https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fthecuberesearch.com%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=theCUBE+Research&index=11&md5=7bb33676722925eb57d588ec343e4f6f),
[CUBE365](https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fwww.cube365.net%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=CUBE365&index=12&md5=d310fb35919714e66ad8d42c9c0c1bc6),
[theCUBE AI](https://cts.businesswire.com/ct/CT?id=smartlink&url=https%3A%2F%2Fwww.thecubeai.com%2F&esheet=54119777&newsitemid=20240910506833&lan=en-US&anchor=theCUBE+AI&index=13&md5=b8b98472f8071b23ebb10ab9a8dd0683)and theCUBE SuperStudios — with flagship locations in Silicon Valley and the New York Stock Exchange — SiliconANGLE Media operates at the intersection of media, technology and AI.
Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.

View File

@@ -0,0 +1,47 @@
# Root Cause Analysis as Storytelling – Wide Awake Developers
- **期号**: SRE Weekly Issue #97(2017-11-12)
- **作者**: —
- **链接**: http://www.michaelnygard.com/blog/2017/11/root-cause-analysis-as-storytelling/
## 简介
When we use root cause analysis, says Michael Nygard, we narrow our focus into counter-factuals that get in the way of finding out what really happened.
CW: hypothetical violent imagery
## 正文
Humans are great storytellers and even better story-listeners. We love to hear stories so much that when there aren't any available, we make them up on our own.
From an early age, children grasp the idea of narrative. Even if they don't understand the forms of storytelling so much, you can hear a four-year-old weave a linked list of events from her day.
We look for stories behind everything. At a deep level, we want the world's events to mean something. Effect follows cause, and causes have an actor to set them in motion.
Our sense of balance also demands that large effects should have large causes, with correspondingly large intent.
A drunk driver speeds through a red light, oblivious. A crossing car stops short. The shaken driver creeps home with a pounding pulse, full of queasy adrenaline. She unbuckles her daughter and hugs her tightly.
A drunk driver speeds through a red light, oblivious. A crossing car is in the intersection. The drunk smashes into it, right at the drivers' side door. The woman's bloody face is hidden behind airbags. Her daughter sits in her new wheelchair for her mother's funeral.
The difference between those stories is a matter of a split second in timing. There is absolutely no change in the motives or desires of anyone in the two vignettes. The first drunk, if caught, would get a jail term and large fine. He would probably lose his driver's license.
But most people would judge the motives of the second driver far more harshly. They would condemn him to a lengthy prison term and a lifetime ban on driving.
When we see a large effect, we expect a large cause, with a large intent.
The idea that some vast, horrible events strike randomly fills us with dread. People can't bear the thought that a single unbalanced nobody can change the course of a nation's history with one rifle shot, so they spend more than 50 years searching for "the truth."
"Root Cause Analysis" expresses a desire for narrative. With the power of hindsight, we want to find out what went wrong, who did it, and how we can make sure it never happens again. But because we have the posterior event, we judge the prior probabilities differently. Any anomaly or blip suddenly becomes suspect.
People don't look as hard at anomalies when nothing bad happens.
They don't notice all the times the same weird log message pops up before … everything continues as normal.
When we look for "root cause," what we are really trying to discern is not "what made this happen." We are looking for something that would have stopped it from happening. We are building a counterfactual narrative—an alternate history—where that drunk driver dropped his keys in the parking lot and was thereby delayed a few crucial seconds.
Peel back the surface on a root cause analysis and you almost always see a formula that goes like this: "factor X" could have prevented this. "Factor X" was not present, therefore the bad event happened.
The catch is that there is usually an endless variety of possible counterfactuals. Often, more than one counterfactual narrative would have prevented the bad outcome equally well. Which one was the root cause? Non-existence of "factor X" or non-existence of "factor Y?"
Next time you have a bad incident, why not try to focus your efforts in a different way? Work on learning from the times that things don't go wrong. And be explicit about looking for many possible interventions that would have prevented the problem. Then select ones with broad ability to prevent or impede many different problems.