SRE weekly 所有文章
This commit is contained in:
99
sreweekly/markdown/32/01-choose-boring-technology.md
Normal file
99
sreweekly/markdown/32/01-choose-boring-technology.md
Normal file
@@ -0,0 +1,99 @@
|
||||
# Choose Boring Technology
|
||||
|
||||
- **期号**: SRE Weekly Issue #32(2016-07-25)
|
||||
- **作者**: —
|
||||
- **链接**: http://mcfunley.com/choose-boring-technology
|
||||
|
||||
## 简介
|
||||
|
||||
It’s tempting to use the newest shiny stack when building a new system. Dan McKinley argues that you should limit yourself to only a few shiny technologies to avoid excessive operational burden.
|
||||
|
||||
> […] the long-term costs of keeping a system working reliably vastly exceed any inconveniences you encounter while building it.
|
||||
|
||||
## 正文
|
||||
|
||||
Probably the single best thing to happen to me in my career was having had [Kellan](http://laughingmeme.org/) placed in charge of me. I stuck around long enough to see Kellan’s technical decisionmaking start to bear fruit. I learned a great deal *from* this, but I also learned a great deal as a *result* of this. I would not have been free to become the engineer that wrote [Data Driven Products Now!](http://mcfunley.com/data-driven-products-lean-startup-2014) if Kellan had not been there to so thoroughly stick the landing on technology choices.
|
||||
|
||||

|
||||
|
||||
In the year since leaving Etsy, I’ve resurrected my ability to care about technology. And my thoughts have crystallized to the point where I can write them down coherently. What follows is a distillation of the Kellan gestalt, which will hopefully serve to horrify him only slightly.
|
||||
|
||||
##### Embrace Boredom.
|
||||
|
||||
Let’s say every company gets about three innovation tokens. You can spend these however you want, but the supply is fixed for a long while. You might get a few more *after* you achieve a [certain level of stability and maturity](http://rc3.org/2015/03/24/the-pleasure-of-building-big-things/), but the general tendency is to overestimate the contents of your wallet. Clearly this model is approximate, but I think it helps.
|
||||
|
||||
If you choose to write your website in NodeJS, you just spent one of your innovation tokens. If you choose to use [MongoDB](http://mcfunley.com/why-mongodb-never-worked-out-at-etsy), you just spent one of your innovation tokens. If you choose to use [service discovery tech that’s existed for a year or less](https://consul.io/), you just spent one of your innovation tokens. If you choose to write your own database, oh god, you’re in trouble.
|
||||
|
||||
Any of those choices might be sensible if you’re a javascript consultancy, or a database company. But you’re probably not. You’re probably working for a company that is at least ostensibly [rethinking global commerce](https://www.etsy.com) or [reinventing payments on the web](https://stripe.com) or pursuing some other suitably epic mission. In that context, devoting any of your limited attention to innovating ssh is an excellent way to fail. Or at best, delay success [\[1\]](http://mcfunley.com#f1).
|
||||
|
||||
What counts as boring? That’s a little tricky. “Boring” should not be conflated with “bad.” There is technology out there that is both boring and bad [\[2\]](http://mcfunley.com#f2). You should not use any of that. But there are many choices of technology that are boring and good, or at least good enough. MySQL is boring. Postgres is boring. PHP is boring. Python is boring. Memcached is boring. Squid is boring. Cron is boring.
|
||||
|
||||
The nice thing about boringness (so constrained) is that the capabilities of these things are well understood. But more importantly, their failure modes are well understood. Anyone who knows me well will understand that it’s only with a overwhelming sense of malaise that I now invoke the spectre of Don Rumsfeld, but I must.
|
||||
|
||||

|
||||
|
||||
When choosing technology, you have both known unknowns and unknown unknowns [\[3\]](http://mcfunley.com#f3).
|
||||
|
||||
- A known unknown is something like: *we don’t know what happens when this database hits 100% CPU.*
|
||||
- An unknown unknown is something like: *geez it didn’t even occur to us that [writing stats would cause GC pauses](http://www.evanjones.ca/jvm-mmap-pause.html).*
|
||||
|
||||
Both sets are typically non-empty, even for tech that’s existed for decades. But for shiny new technology the magnitude of unknown unknowns is significantly larger, and this is important.
|
||||
|
||||
##### Optimize Globally.
|
||||
|
||||
I unapologetically think a bias in favor of boring technology is a good thing, but it’s not the only factor that needs to be considered. Technology choices don’t happen in isolation. They have a scope that touches your entire team, organization, and the system that emerges from the sum total of your choices.
|
||||
|
||||
Adding technology to your company comes with a cost. As an abstract statement this is obvious: if we’re already using Ruby, adding Python to the mix doesn’t feel sensible because the resulting complexity would outweigh Python’s marginal utility. But somehow when we’re talking about Python and Scala or MySQL and Redis people [lose their minds](http://martinfowler.com/bliki/PolyglotPersistence.html), discard all constraints, and start raving about using the best tool for the job.
|
||||
|
||||
[Your function in a nutshell](https://twitter.com/coda/status/580531932393504768) is to map business problems onto a solution space that involves choices of software. If the choices of software were truly without baggage, you could indeed pick a whole mess of locally-the-best tools for your assortment of problems.
|
||||
|
||||
But of course, the baggage exists. We call the baggage “operations” and to a lesser extent “cognitive overhead.” You have to monitor the thing. You have to figure out unit tests. You need to know the first thing about it to hack on it. You need an init script. I could go on for days here, and all of this adds up fast.
|
||||
|
||||
The problem with “best tool for the job” thinking is that it takes a myopic view of the words “best” and “job.” Your job is keeping the company in business, god damn it. And the “best” tool is the one that occupies the “least worst” position for as many of your problems as possible.
|
||||
|
||||
It is basically always the case that the long-term costs of keeping a system working reliably vastly exceed any inconveniences you encounter while building it. Mature and productive developers understand this.
|
||||
|
||||
##### Choose New Technology, Sometimes.
|
||||
|
||||
Taking this reasoning to its *reductio ad absurdum* would mean picking Java, and then trying to implement a website without using anything else at all. And that would be crazy. You need some means to add things to your toolbox.
|
||||
|
||||
An important first step is to acknowledge that this is a process, and a conversation. New tech eventually has company-wide effects, so adding tech is a decision that requires company-wide visibility. Your organizational specifics may force the conversation, or [they may facilitate developers adding new databases and queues without talking to anyone](https://twitter.com/mcfunley/status/578603932949164032). One way or another you have to set cultural expectations that **this is something we all talk about**.
|
||||
|
||||
One of the most worthwhile exercises I recommend here is to **consider how you would solve your immediate problem without adding anything new**. First, posing this question should detect the situation where the “problem” is that someone really wants to use the technology. If that is the case, you should immediately abort.
|
||||
|
||||

|
||||
|
||||
It can be amazing how far a small set of technology choices can go. The answer to this question in practice is almost never “we can’t do it,” it’s usually just somewhere on the spectrum of “well, we could do it, but it would be too hard” [\[4\]](http://mcfunley.com#f4). If you think you can’t accomplish your goals with what you’ve got now, you are probably just not thinking creatively enough.
|
||||
|
||||
It’s helpful to **write down exactly what it is about the current stack that makes solving the problem prohibitively expensive and difficult.** This is related to the previous exercise, but it’s subtly different.
|
||||
|
||||
New technology choices might be purely additive (for example: “we don’t have caching yet, so let’s add memcached”). But they might also overlap or replace things you are already using. If that’s the case, you should **set clear expectations about migrating old functionality to the new system.** The policy should typically be “we’re committed to migrating,” with a proposed timeline. The intention of this step is to keep wreckage at manageable levels, and to avoid proliferating locally-optimal solutions.
|
||||
|
||||
This process is not daunting, and it’s not much of a hassle. It’s a handful of questions to fill out as homework, followed by a meeting to talk about it. I think that if a new technology (or a new service to be created on your infrastructure) can pass through this gauntlet unscathed, adding it is fine.
|
||||
|
||||
##### Just Ship.
|
||||
|
||||
Polyglot programming is sold with the promise that letting developers choose their own tools with complete freedom will make them more effective at solving problems. This is a naive definition of the problems at best, and motivated reasoning at worst. The weight of day-to-day operational [toil](https://twitter.com/handler) this creates crushes you to death.
|
||||
|
||||
Mindful choice of technology gives engineering minds real freedom: the freedom to [contemplate bigger questions](http://mcfunley.com/effective-web-experimentation-as-a-homo-narrans). Technology for its own sake is snake oil.
|
||||
|
||||
*Update, July 27th 2015: I wrote a talk based on this article. You can see it [here](http://boringtechnology.club).*
|
||||
|
||||
1.
|
||||
[required years of effort to amputate](https://www.youtube.com/watch?v=eenrfm50mXw) .
|
||||
Meanwhile, the 90th percentile search latency was about two minutes.[Etsy didn't fail](http://www.sec.gov/Archives/edgar/data/1370637/000119312515077045/d806992ds1.htm) ,
|
||||
but it went several years without shipping anything at all. So it took longer to succeed than it needed to.
|
||||
2.
|
||||
3.
|
||||
[the Socratic Paradox](http://en.wikipedia.org/wiki/I_know_that_I_know_nothing) .
|
||||
Socrates was by all accounts a thoughtful individual in a number of ways that Rumsfeld is not.
|
||||
4.
|
||||
A good example of this from my experience is [Etsy’s activity
|
||||
feeds](https://speakerdeck.com/mcfunley/etsy-activity-feed-architecture) . When we built this feature, we were working pretty hard to consolidate
|
||||
most of Etsy onto PHP, MySQL, Memcached, and Gearman (a PHP job server).
|
||||
It was much more complicated to implement the feature on that stack than it
|
||||
might have been with something like Redis (or[maybe not](https://aphyr.com/posts/283-call-me-maybe-redis) ).
|
||||
But it is absolutely possible to build activity feeds on that stack.An amazing thing happened with that project: our attention turned elsewhere for several years. During that time, activity feeds scaled up 20x while *nobody was watching it at all.* We made no changes
|
||||
whatsoever specifically targeted at activity feeds, but everything worked
|
||||
out fine as usage exploded because we were using a shared platform.
|
||||
This is the long-term benefit of restraint in technology choices in a nutshell.This isn’t an absolutist position--while activity feeds stored in memcached was judged to be practical, implementing full text search with faceting in raw PHP wasn't. So Etsy used Solr.
|
||||
@@ -0,0 +1,53 @@
|
||||
# Postmortem-Report-Reviews/2016-07-20-pshima-stack-exchange-2016-07-20.md
|
||||
|
||||
- **期号**: SRE Weekly Issue #32(2016-07-25)
|
||||
- **作者**: —
|
||||
- **链接**: https://github.com/Operations-Incident-Board/Postmortem-Report-Reviews/blob/master/2016-07-20-pshima-stack-exchange-2016-07-20.md
|
||||
|
||||
## 简介
|
||||
|
||||
Quick on the draw, Pete Shima gives us a review of Stack Exchange’s outage postmortem (linked below) as part of the Operations Incident Board’s Postmortem Report Reviews project. Thanks, Pete!
|
||||
|
||||
## 正文
|
||||
|
||||
Postmortem link: [http://stackstatus.net/post/147710624694/outage-postmortem-july-20-2016](http://stackstatus.net/post/147710624694/outage-postmortem-july-20-2016)
|
||||
|
||||
The customer base of Stack Exchange is largely of technical users looking for solutions to problems and/or to participate in the large community and knowledgebase present in this network.
|
||||
|
||||
From the Stack Exchange about page:
|
||||
|
||||
Stack Exchange is a network of 150+ Q&A communities including Stack Overflow, the preeminent site for programmers to find, ask, and answer questions about software development. Founded in 2008 by Joel Spolsky and Jeff Atwood, the company was built on the premise that serving the developer community at large would lead to a better, smarter Internet. Since then, the Stack Exchange network has grown into a top-50 online destination, with Stack Overflow alone serving more than 40 million professional and novice programmers every month. The broader Stack Exchange Network has expanded to cover topics as diverse as Mathematics, Home Improvement, Statistics, and English Language and Usage.
|
||||
|
||||
I was not personally impacted from the event but the post mortem suggests that this was a major or full outage. It does also not describe what sites were impacted. As Stack Exchange is a network of sites, this would suggest that every site on the network was impaired to the point of being completely unavailable. The target audience of this is likely the people impacted from it so it may be omitted intentionally and the tweets regarding it do clarify this information. This was later clarified that it was only Stack Overflow that was offline, other sites were only impacted by high CPU usage on the webservers.
|
||||
|
||||
The overview also provides some interesting details that give us some insight to where the 34 minutes time was accrued. 10 minutes of identification, 14 minutes to write the fix and 10 minutes to deploy. This is an interesting point of detail that is often not included in post mortems and I feel it gives the reader more understanding of how the event unfolded and where time was spent.
|
||||
|
||||
The next paragraph is where things start to get interesting as it touches on several of the root causes of the issue. The initial one was high CPU usage on the webservers which was caused by the regexp, and this had the knock on effect of slowing down all web requests, which had the knock on effect of causing what I am guessing is time outs on the load balancer health checks, which caused them to deregister all the webservers, which then caused the load balancers to have no available web servers to send traffic, which sent 503 responses to users. Some specific details are left to the user but are not material in understanding the cascading impact of the events that unfolded.
|
||||
|
||||
The lower Technical Details section goes into specifics on why the regular expression caused so much havoc and gives the reader confidence that the Stack Exchange folks have a deep understanding of the problem.
|
||||
|
||||
The incident was first posted to twitter via [https://twitter.com/StackStatus/status/755778941600882688](https://twitter.com/StackStatus/status/755778941600882688) at 14:58 UTC, just 14 minutes after the time noted in the post mortem, retweets were then on the account from [https://twitter.com/Nick_Craver](https://twitter.com/Nick_Craver), with updates noting the fix was being deployed and that it was back online. There is actually considerably more detail from Nick's twitter than from the official Stack Exchange twitter including a graph of CPU usage: [https://twitter.com/Nick_Craver/status/755793398544601088](https://twitter.com/Nick_Craver/status/755793398544601088) and a tweet describing that they are using HAProxy as a load balancer and that health checks can be enabled/disabled in realtime: [https://twitter.com/Nick_Craver/status/755795805798330368](https://twitter.com/Nick_Craver/status/755795805798330368).
|
||||
|
||||
Overall the response was pretty timely and the retweets gave users updates as to what was occuring and multiple follow ups were added even though the outage window was only 34 minutes.
|
||||
|
||||
- Audit our regular expressions and post validation workflow for any similar issues
|
||||
|
||||
This seems like a solid follow up task.
|
||||
|
||||
- Add controls to our load balancer to disable the healthcheck – as we believe everything but the home page would have been accessible if it wasn’t for the the health check
|
||||
|
||||
This was clarified as a control that an operator could use to disable the health checks if the same issue reoccurred. It does not solve the larger fleet utilization issue but it would have prevented this issue from reoccuring in which all the hosts were marked as unhealthy on the load balancer. Despite other sites possibly being slow from the utilization, they still would have been "up".
|
||||
|
||||
- Create a “what to do during an outage” checklist since our StackStatus Twitter notification was later than we would have liked (and a few other outage workflow items we would like to be more consistent on).
|
||||
|
||||
This is an interesting action item and I love seeing action items around improving process. This suggests such events rarely happen at Stack Exchange and when the issue occurred there wasn't a 'what to do during an outage' checklist. If that was the case then this is an impressive recovery time from the events without a procedure. It also shows that while it was 14 minutes to post, they believe they can improve on that time window and response to customers.
|
||||
|
||||
It is great to see this level of visibility from Stack Exchange! The details in the post mortem give the reader a decent understanding of the issue and what caused the full outage. The tweets would suggest Nick played a big part in writing the post mortem and also in recovery of the issue and I would like to thank Nick and the Stack Exchange team for providing this visibility and for releasing the post mortem so quickly after a major incident.
|
||||
|
||||
It is unknown to me if there was a specific site with the post in the Stack Exchange network that caused the issue but without further detail it would suggest that there is a very large blast radius for Stack Exchange sites and a bug on any of the sites could cause a full network outage. This is not a negative, only an observation from this post mortem. It does leave the reader wondering more.
|
||||
|
||||
I would have loved to see some additional information on how the event was discovered and more around improving resiliency of the web server fleet. The root cuased was identified fairly quickly, and more knowledge here could give the reader additional confidence that while this specific issue was resolved, future bugs or unknowns could also be resolved in a similar manner. This is an nit pick of this post mortem and I would consider this post mortem as a good example as a follow up to a major issue. It was later clarified by Nick that details were intentionally left out because the tools they use legally cannot be shared. The clrmd tool ([https://github.com/Microsoft/clrmd](https://github.com/Microsoft/clrmd)) was referenced as something to look at.
|
||||
|
||||
The speed at which the issue was resolved, combined with the communication and immediate follow up to the issue gives the reader confidence that future issues will be resolved in a similar manner and this is a major plus for any good post mortem. I would celebrate this post mortem as a success!
|
||||
|
||||
My name is Pete Shima. [me@peteshima.com](mailto:me@peteshima.com) - petey5k@twitter
|
||||
39
sreweekly/markdown/32/03-chaos-community-day.md
Normal file
39
sreweekly/markdown/32/03-chaos-community-day.md
Normal file
@@ -0,0 +1,39 @@
|
||||
# Chaos Community Day
|
||||
|
||||
- **期号**: SRE Weekly Issue #32(2016-07-25)
|
||||
- **作者**: —
|
||||
- **链接**: http://chaos.community/
|
||||
|
||||
## 简介
|
||||
|
||||
Next month in Seattle will be the second annual Chaos Community Day, an event full of presentations on chaos engineering. I wish I could attend!
|
||||
|
||||
## 正文
|
||||
|
||||
## Notes:
|
||||
|
||||
In this episode we are joined with Cat Swetel, who is a fully vaccinated technology leader. Cat talks with us about devops, feminism, and how epistemic injustice correlates to Continuous Verification.
|
||||
|
||||
[\[Read More\]](https://chaos.community/broadcast/episode-15/)
|
||||
|
||||
Navigating the chaos together
|
||||
|
||||
In this episode we are joined with Jason Cahoon, Site Reliability Engineer at Google. Jason talks with us about all things Google and the DiRT framework (aka Chaos Engineering), from the massive amount of coding happening at Google to experiments reinforcing the value of DiRT itself (plus get the dirt on the DiRT curse!).
|
||||
|
||||
In this episode we are joined with Amy Tobey, Principal SRE at Equinix. Amy discusses with us everything from what an SRE is to whether Chaos Engineering is an advanced practice to the rise of computers.
|
||||
|
||||
In this episode we are joined with Corey Quinn, Chief Cloud Economist at The Duckbill Group. Corey talks with us about many things including Chaos Engineering as a cost optimization strategy for the cloud and his thoughts on AWS Chaos Engineering platforms.
|
||||
|
||||
In this episode, a few of our speakers from GOTOpia 2021 joined us live for a conversation about Chaos Engineering and the benefits, the oddities, and how to undermine the practice.
|
||||
|
||||
In this impromptu episode we take a tour through highlights from each of our 12 previous guests. Season 2, here we come!
|
||||
|
||||
In this episode, we are joined by Christina Yakomin, a vanguard of reliability at Vanguard. Christina dispels the notion that Chaos Engineering is just for startups and can’t be done in high-stakes, regulated environments. This is an episode that hits just the right note—literally.
|
||||
|
||||
In this episode, we are joined by Andy Fleener, Platform Operations Manager at SportsEngine, and contributing author to Chaos Engineering: System Resiliency in Practice.
|
||||
|
||||
We discuss his chapter on humanistic chaos and how we can apply Chaos Engineering to human systems. This leads to other subjects such as Andy’s affinity for humans and how humans approach problems and his preference for salty nut rolls versus peanut brittle.
|
||||
|
||||
In this episode we are joined with Liz Fong-Jones, developer advocate, labor and ethics organizer, and Site Reliability Engineer. She joins in the conversation to discuss observability versus monitoring and how this is critical for Chaos Engineering. Along the way we discuss unionizing software workers, how internet companies handle traffic, and more.
|
||||
|
||||
In this episode, we are joined with Mikolaj Pawlikowski, creator of Powerful Seal and author of the Chaos Engineering book at Manning. This conversation runs the gamut from open source software to seatbelts to the much beloved Toyota Yaris. There is some practical advice on Chaos Engineering and how to help Kubernetes be more reliable through Chaos Engineering practices.
|
||||
@@ -0,0 +1,13 @@
|
||||
# Lives at risk during nationwide weather service meltdown
|
||||
|
||||
- **期号**: SRE Weekly Issue #32(2016-07-25)
|
||||
- **作者**: —
|
||||
- **链接**: http://www.mypalmbeachpost.com/news/weather/lives-at-stake-during-nationwide-weather-service-s/nr3zw/
|
||||
|
||||
## 简介
|
||||
|
||||
As the world becomes more and more dependent on the services we administer, outages become more and more likely to put real people in danger. Here’s a rundown of how dangerous last week’s four-hour outage in US’s national weather service was.
|
||||
|
||||
## 正文
|
||||
|
||||
> ⚠️ 抓取失败:HTTP 403
|
||||
@@ -0,0 +1,13 @@
|
||||
# What We Don’t Get About Microsoft Azure
|
||||
|
||||
- **期号**: SRE Weekly Issue #32(2016-07-25)
|
||||
- **作者**: —
|
||||
- **链接**: http://www.itbusinessedge.com/blogs/unfiltered-opinion/what-we-dont-get-about-microsoft-azure.html
|
||||
|
||||
## 简介
|
||||
|
||||
An interesting opinion piece that argues that Microsoft Azure is more robust than Google and Amazon’s offerings.
|
||||
|
||||
## 正文
|
||||
|
||||
> ⚠️ 抓取失败:HTTP 403
|
||||
@@ -0,0 +1,13 @@
|
||||
# 4 Software Quality Lessons From Pokemon Go’s Wild First Week – DZone Performance
|
||||
|
||||
- **期号**: SRE Weekly Issue #32(2016-07-25)
|
||||
- **作者**: —
|
||||
- **链接**: https://dzone.com/articles/4-software-quality-lessons-from-pokemon-gos-wild-f
|
||||
|
||||
## 简介
|
||||
|
||||
This week, I’m trying to catch all the articles being written about Pokémon GO. Here’s one that supposes the problem might be a lack of sufficient testing.
|
||||
|
||||
## 正文
|
||||
|
||||
> ⚠️ 抓取失败:HTTP 410
|
||||
@@ -0,0 +1,13 @@
|
||||
# Niantic And Nintendo’s Lack Of Communication About ‘Pokémon GO’ Issues Is Inexcusable
|
||||
|
||||
- **期号**: SRE Weekly Issue #32(2016-07-25)
|
||||
- **作者**: —
|
||||
- **链接**: http://www.forbes.com/sites/insertcoin/2016/07/21/niantic-and-nintendos-lack-of-communication-about-pokemon-go-issues-is-inexcusable/#791c57b52e83
|
||||
|
||||
## 简介
|
||||
|
||||
Pokémon GO is blowing up like crazy, and I don’t just mean in popularity. Forbes has a lot to say about the complete lack of communication during and after outages, and we’d do well to listen. This article reads a lot like a recipe for how to communicate well to your userbase about outages.
|
||||
|
||||
## 正文
|
||||
|
||||
> ⚠️ 抓取失败:HTTP 403
|
||||
@@ -0,0 +1,13 @@
|
||||
# Netflix Billing Migration to AWS – Part II
|
||||
|
||||
- **期号**: SRE Weekly Issue #32(2016-07-25)
|
||||
- **作者**: —
|
||||
- **链接**: http://techblog.netflix.com/2016/07/netflix-billing-migration-to-aws-part-ii.html
|
||||
|
||||
## 简介
|
||||
|
||||
Here’s the continuation of last month’s article on Netflix’s billing migration.
|
||||
|
||||
## 正文
|
||||
|
||||
> ⚠️ 抓取失败:HTTP 403
|
||||
Reference in New Issue
Block a user