SRE weekly 所有文章

This commit is contained in:
2026-09-12 17:23:01 +08:00
parent 409b40ddcb
commit af7633f9dc
8486 changed files with 4489990 additions and 7 deletions

View File

@@ -0,0 +1,15 @@
# When downtime is not an option
- **期号**: SRE Weekly Issue #48(2016-11-13)
- **作者**: —
- **链接**: http://saudigazette.com.sa/business/downtime-not-option/
## 简介
A detailed description of Disaster Recovery as a Service (DRaaS), including a discussion of the cost versus creating a DR site oneself. This is the part I always wonder about:
> However, for larger enterprises with complex infrastructures and larger data volumes spread across disparate systems, DRaaS has often been too complicated and expensive to implement.
## 正文
> ⚠️ 抓取失败:HTTP 404

View File

@@ -0,0 +1,15 @@
# The Prime Directive
- **期号**: SRE Weekly Issue #48(2016-11-13)
- **作者**: —
- **链接**: http://retrospectives.com/pages/retroPrimeDirective.html
## 简介
This one’s so short I can almost quote the whole thing here. I love its succinctness:
> Regardless of what we discover, we understand and truly believe that everyone did the best job they could, given what they knew at the time, their skills and abilities, the resources available, and the situation at hand.
## 正文
> ⚠️ 抓取失败:HTTP 503

View File

@@ -0,0 +1,13 @@
# The Netflix Tech Blog: Post-mortem of October 22,2012 AWS degradation
- **期号**: SRE Weekly Issue #48(2016-11-13)
- **作者**: —
- **链接**: http://techblog.netflix.com/2012/10/post-mortem-of-october-222012-aws.html?m=1
## 简介
Just over four years ago, Amazon had a major outage in Elastic Block Store (EBS). Did you see impact? I sure did. Here’s Netflix’s account of how they survived the outage mostly unscathed.
## 正文
> ⚠️ 抓取失败:HTTP 403

View File

@@ -0,0 +1,70 @@
# Serverless promises and the persistent need for critical alerting
- **期号**: SRE Weekly Issue #48(2016-11-13)
- **作者**: —
- **链接**: http://onpage.com/serverless-meets-security-and-critical-alerting/
## 简介
I’m glad to see more people writing that Serverless != #NoOps. This article is well-argued even though it turns into an OnPage ad 3 paragraphs from the end.
## 正文
# Serverless promises and the persistent need for critical alerting
## Why serverless computing doesn’t end the need for security or alerts
Serverless computing provides the advantage of taking away the problem of managing servers. For many small start-ups, this is a huge advantage as the cost of purchasing, maintaining and scaling servers is a real pain point. Serverless also holds forth the prospect of ending the need for Ops as we know it, ending the need for security worries and ending the need for being on-call. But, while this modern-day DevOps marvel known as serverless might seem like a panacea, serverless computing needs to come with a healthy dose of reality.
### The reality of serverless
In an article I recently posted to DZone entitled *How Smart Is Serverless*, I question how smart it is to outsource your security concerns to a third party like AWS. As I note in the article, you cannot abstract security without facing some pretty scary consequences. [Amichai Shulman](http://www.pcworld.com/article/2365602/hacker-puts-full-redundancy-codehosting-firm-out-of-business.html), CTO of Imperva, says this best when he notes:
*“[B]usinesses act on the misconception that when they put data into the cloud they somehow transfer responsibility and liability to the cloud provider, and this is simply not true.”*
While AWS guarantees the physical framework of the information on its property, it doesn’t provide security for anything beyond it. That means all the information in transit is not protected. And there is no provisioning to make sure that you don’t store security or other sensitive information into your code that could be seen in transit.
Furthermore, the servers used by AWS can and have been hacked. CodeSpaces and Ashley-Madison are real examples. The larger a target becomes the more enticing they become to hackers, whether the target is on AWS or on their own servers.
So by embracing serverless and thinking you are entering a NoOps fairy tale world, you *still* need to have Ops or a Dev team member trained in Ops to run scripts against the software, test for security and (importantly) monitor the logs. [Charity Majors](https://charity.wtf/2016/05/31/wtf-is-operations-serverless/) writes it best when she notes:
“I’ve seen what happens when application developers think they don’t have to care about the skills associated with operations engineering. When they forget that **no matter how pretty the abstractions are, you’re still dealing with dusty old concepts** like “persistent state” and “queries” and “unavailability” and so forth, or when they literally just think they can throw money at a service to make it go faster because that’s totally how services work.”
Clearly, doing away with Ops and jumping headfirst into serverless because you think you can avoid all those components of operations you don’t enjoy, is lunacy. Worse, it leads to *really* bad outcomes.
### Critical alerting won’t go away
Dev in a serverless world still requires concern for deployment, security, networking, debugging, monitoring and system scaling. According to [Mike Roberts](http://martinfowler.com/articles/serverless.html) who writes on Martin Fowler’s website, “These problems all still exist with Serverless apps and you’re still going to need a strategy to deal with them.“ As such, you still need critical alerting platforms to sit on your logs coming out of AWS or other serverless providers to notify you how things are going and when they are heading south.
OnPage’s incident alerting platform is ideal for this purpose as it alerts based on emails so anything that can send an email can integrate with the OnPage platform. OnPage will notify for events such as failed deployment, security issues or unusual traffic patterns. By integrating OnPage with log tools such as Logz.io which will monitor your logs for incidents you identify, you can:
- Create persistent alerts that last for up to 8 hours, ensuring you never miss a critical alert
- Tailor alert settings for high-priority messaging (think DNS attacks) as well as for low priority, casual messaging. If an alert is a 401 error, you can set it as a low priority alert so you don’t need to wake up for it in the middle of the night
- Include information on the alert specific to the log error so that information arrives with context
- Send messages between OnPage IDs. That is, one person with the app can send a message to another person with the app
- Enable global coverage so your teams down the block or across the globe can be notified
### Conclusion
So while serverless has many advantages that will help start-ups such as providing significant cost savings or reduce Ops as well as the use of servers, companies should be careful to fully understand the hazards they face by moving to serverless. Critical alerting has been and will be a feature that Devs will need whether they have zero servers in house or a 100.
[Read our blog](https://www.onpage.com/serverless-needs-critical-alerting/) on NoOps and critical alerting.
### About The Author
[Facebook](<https://www.facebook.com/sharer.php?t=Serverless promises and the persistent need for critical alerting&u=https://www.onpage.com/serverless-meets-security-and-critical-alerting/>)
![facebook Share on facebook](https://www.onpage.com/wp-content/plugins/simple-share-buttons-adder/buttons/simple/facebook.png)
[Google]
![google Share on google](https://www.onpage.com/wp-content/plugins/simple-share-buttons-adder/buttons/simple/google.png)
[Twitter](<https://twitter.com/intent/tweet?text=Serverless promises and the persistent need for critical alerting&url=https://www.onpage.com/serverless-meets-security-and-critical-alerting/&via=>)
![twitter Share on twitter](https://www.onpage.com/wp-content/plugins/simple-share-buttons-adder/buttons/simple/twitter.png)
[Linkedin](<https://www.linkedin.com/shareArticle?title=Serverless promises and the persistent need for critical alerting&url=https://www.onpage.com/serverless-meets-security-and-critical-alerting/>)
![linkedin Share on linkedin](https://www.onpage.com/wp-content/plugins/simple-share-buttons-adder/buttons/simple/linkedin.png)

View File

@@ -0,0 +1,13 @@
# Episode 004: Charity Majors – Greater Than Code
- **期号**: SRE Weekly Issue #48(2016-11-13)
- **作者**: —
- **链接**: https://www.greaterthancode.com/2016/10/21/episode-004-charity-majors/
## 简介
What else can we expect from Greater Than Code + Charity Majors? This podcast is 50 minutes of awesome, and there’s a transcription, too! Listen/read for awesome phrases like “stamping out chaos”, find out why Charity says, “I personally hate [the term ‘SRE’] (but I hate a lot of things)”, and hear Conway’s law applied to microservices, #NoOps debunking, and a poignant ending about misogyny and equality.
## 正文
> ⚠️ 抓取失败:URLError: [SSL: TLSV1_ALERT_INTERNAL_ERROR] tlsv1 alert internal error (_ssl.c:1032)

View File

@@ -0,0 +1,50 @@
# Microsoft Announces Azure DNS General Availability
- **期号**: SRE Weekly Issue #48(2016-11-13)
- **作者**: —
- **链接**: https://www.infoq.com/news/2016/10/azure-dns-ga
## 简介
Microsoft released its Route 53 competitor in late September. They say:
> Azure DNS has the scale and redundancy built-in to ensure high availability for your domains. As a global service, Azure DNS is resilient to multiple Azure region failures and network partitioning for both its control plane and DNS serving plane.
## 正文
On September 26<sup>th</sup>, Microsoft announced its Azure DNS service has reached General Availability (GA) in all public Azure regions. The service was initially launched, in preview, at the May 2015 Ignite conference in Chicago.
Sean Wheeler, from the Azure Networking team at Microsoft, [describes](https://azure.microsoft.com/en-us/documentation/articles/dns-overview/) the service in the following way:
“Azure DNS allows customers to host their DNS domain in Azure, so they can manage their DNS records using the same credentials, billing and support contract as their other Azure services.”
Jonathan Tuliani, Azure Networking program manager at Microsoft has [outlined](https://azure.microsoft.com/en-us/blog/azure-dns-general-availability/) some of the benefits of hosting and managing your DNS in Azure:
**Reliability** – Azure DNS has the scale and redundancy built-in to ensure high availability for your domains. As a global service, Azure DNS is resilient to multiple Azure region failures and network partitioning for both its control plane and DNS serving plane.
**Performance** – Our global network of name servers uses ‘anycast’ networking to ensure your DNS queries are always routed to the closest server for the fastest possible response.
**Ease of use** – Your DNS zones and records in Azure DNS can easily be managed via the Azure Portal, Azure PowerShell, or cross-platform Azure CLI. Application integration is supported via our SDK or REST API.
**Security** – Azure DNS benefits from the same authentication and authorization features as other Azure services, including the ability to configure multi-factor authentication and role-based access controls.
**Convenience** – Hosting your DNS in Azure enables you to manage your Azure applications and their DNS records in one place, using a single set of credentials, with a single bill and with end-to-end support.
For customers who purchase their domain name from a third party, they can delegate their DNS to Azure DNS. Customers may want to do this in order to take advantage of Microsoft’s global reach and [99.99% SLA](https://azure.microsoft.com/en-us/support/legal/sla/dns/v1_0/) that they are providing with this service.
Customers can also reduce dependencies on local DNS servers by leveraging this service. Joe Stocker, from the Patriot Consulting Technology Group, [explains](http://www.thecloudtechnologist.com/top-five-reasons-to-consider-azure-dns/):
“If all of your external DNS servers are in the same physical location, then Azure DNS provides an opportunity to migrate to a more resilient solution since Azure DNS is automatically load balanced across multiple regions.”
Microsoft has been working with Third-Party companies including Men & Mice to integrate Azure DNS with management software solutions. Using a solution like Men & Mice allows customers to manage their DNS, DHCP and IPAM configurations in a single tool that now includes support for Azure DNS.
In a recent Men & Mice [blogpost](https://www.menandmice.com/news/microsoft-azure-announces-azure-dns-general-availability-with-men-mice-third-party-support/), the company outlined why it was important to include Azure DNS support:
“Men & Mice has been working closely with the Azure DNS team during the Azure DNS Preview to implement full support for Azure DNS in the Men & Mice Suite. Men & Mice Suite customers can now enjoy the high availability, performance, low cost and convenience of hosting their domains in the cloud with Azure DNS, while maintaining full control of their DNS domains and IP address blocks via the powerful DNS, DHCP and IP Address Management (DDI) tools provided by the Men & Mice Suite.”
![](<https://www.infoq.com/news/2016/10/news/2016/10/azure-dns-ga/en/resources/Azure DNS Mice and Men.png>)
*Image Source: [https://azure.microsoft.com/en-us/blog/azure-dns-general-availability/](<http://Image Source: https://azure.microsoft.com/en-us/blog/azure-dns-general-availability/>)*
Much like many of Microsoft’s other Azure services, Azure DNS will be billed in a usage-based model with no up-front or termination fees. Microsoft has based their billing upon the number of DNS zones that you want to host in Azure and the number of DNS queries that the service receives. For more details on billing, please refer to the following [page](https://azure.microsoft.com/en-us/pricing/details/dns/).

View File

@@ -0,0 +1,13 @@
# Communication Breakdown Leads to Patient Burn
- **期号**: SRE Weekly Issue #48(2016-11-13)
- **作者**: —
- **链接**: https://bwhsafetymatters.org/2016/08/12/communication-breakdown-leads-to-patient-burn/
## 简介
This issue of BWH Safety Matters details an incident in which a communication issue between teams that don’t normally work together resulted in a patient injury. This is exactly the kind of pitfall that becomes more prevalent with the move toward microservices, as siloed teams sometimes come into contact only during an incident.
## 正文
> ⚠️ 抓取失败:HTTP 404

View File

@@ -0,0 +1,19 @@
# New systems will fail: A site outage case study from Envato Market
- **期号**: SRE Weekly Issue #48(2016-11-13)
- **作者**: —
- **链接**: https://envato.com/blog/new-systems-will-fail-unexpected-ways-case-study-envato-market/
## 简介
A detailed postmortem from an outage last month. Lots of takeaways, including one that kept coming up: test your emergency tooling before you need to use it.
## 正文
![AI sound generator guide: How to create sound effects with AI](https://elements.envato.com/learn/wp-content/uploads/2026/06/16x9_JOB-2094_SoundGen_BlogHeader-1536x864-1-768x432.png)
###
[AI sound generator guide: How to create sound effects with AI](https://elements.envato.com/learn/ai-sound-generator-guide)
Discover how Envato's AI sound generator creates custom sound effects from text prompts so creators can skip library searches and edit faster.