SRE weekly 所有文章
This commit is contained in:
@@ -0,0 +1,89 @@
|
||||
# Designing robust and predictable APIs with idempotency
|
||||
|
||||
- **期号**: SRE Weekly Issue #72(2017-05-14)
|
||||
- **作者**: —
|
||||
- **链接**: https://stripe.com/blog/idempotency
|
||||
|
||||
## 简介
|
||||
|
||||
Idempotence is a critically important tool in building a reliable system. Stripe explains the concept and shows how they wrap theoretically non-idempotent actions like charging a credit card into safely idempotent API calls.
|
||||
|
||||
## 正文
|
||||
|
||||
#
|
||||
[Designing robust and predictable APIs with idempotency](https://stripe.com/blog/idempotency)
|
||||
|
||||
|
||||
|
||||
|
||||
Networks are [unreliable](https://en.wikipedia.org/wiki/Fallacies_of_distributed_computing). We’ve all experienced trouble connecting to Wi-Fi, or had a phone call drop on us abruptly.
|
||||
|
||||
The networks connecting our servers are, on average, more reliable than consumer-level last miles like cellular or home ISPs, but given enough information moving across the wire, they’re still going to fail in exotic ways. Outages, routing problems, and other intermittent failures may be statistically unusual on the whole, but still bound to be happening all the time at some ambient background rate.
|
||||
|
||||
To overcome this sort of inherently unreliable environment, it’s important to design APIs and clients that will be robust in the event of failure, and will predictably bring a complex integration to a consistent state despite them. Let’s take a look at a few ways to do that.
|
||||
|
||||
## Planning for failure
|
||||
|
||||
Consider a call between any two nodes. There are a variety of failures that can occur:
|
||||
|
||||
- The initial connection could fail as the client tries to connect to a server.
|
||||
- The call could fail midway while the server is fulfilling the operation, leaving the work in limbo.
|
||||
- The call could succeed, but the connection break before the server can tell its client about it.
|
||||
|
||||
Any one of these leaves the client that made the request in an uncertain situation. In some cases, the failure is definitive enough that the client knows with good certainty that it’s safe to simply retry it. For example, a total failure to even establish a connection to the server. In many others though, the success of the operation is ambiguous from the perspective of the client, and it doesn’t know whether retrying the operation is safe. A connection terminating midway through message exchange is an example of this case.
|
||||
|
||||
This problem is a classic staple of distributed systems, and the definition is broad when talking about a “distributed system” in this sense: as few as two computers connecting via a network that are passing each other messages. Even the Stripe API and just one other server that’s making requests to it comprise a distributed system.
|
||||
|
||||
## Making liberal use of idempotency
|
||||
|
||||
The easiest way to address inconsistencies in distributed state caused by failures is to implement server endpoints so that they’re *idempotent*, which means that they can be called any number of times while guaranteeing that side effects only occur once.
|
||||
|
||||
When a client sees any kind of error, it can ensure the convergence of its own state with the server’s by retrying, and can continue to retry until it verifiably succeeds. This fully addresses the problem of an ambiguous failure because the client knows that it can safely handle any failure using one simple technique.
|
||||
|
||||
As an example, consider the API call for a hypothetical DNS provider that enables us to add subdomains via an HTTP request:
|
||||
|
||||
All the information needed to create a record is included in the call, and it’s perfectly safe for a client to invoke it any number of times. If the server receives a call that it realizes is a duplicate because the domain already exists, it simply ignores the request and responds with a successful status code.
|
||||
|
||||
According to HTTP semantics, [the `PUT` and `DELETE` verbs are idempotent](https://tools.ietf.org/html/rfc7231#section-4.2.2), and [the `PUT` verb](https://tools.ietf.org/html/rfc7231#section-4.3.4) in particular signifies that a target resource should be created or replaced entirely with the contents of a request’s payload (in modern RESTful parlance, [a modification would be represented by a `PATCH`](https://tools.ietf.org/html/rfc5789)).
|
||||
|
||||
## Guaranteeing “exactly once” semantics
|
||||
|
||||
While the inherently idempotent HTTP semantics around `PUT` and `DELETE` are a good fit for many API calls, what if we have an operation that needs to be invoked exactly once and no more? An example might be if we were designing an API endpoint to charge a customer money; accidentally calling it twice would lead to the customer being double-charged, which is very bad.
|
||||
|
||||
This is where *idempotency keys* come into play. When performing a request, a client generates a unique ID to identify just that operation and sends it up to the server along with the normal payload. The server receives the ID and correlates it with the state of the request on its end. If the client notices a failure, it retries the request with the same ID, and from there it’s up to the server to figure out what to do with it.
|
||||
|
||||
If we consider our sample network failure cases from above:
|
||||
|
||||
- On retrying a connection failure, on the second request the server will see the ID for the first time, and process it normally.
|
||||
- On a failure midway through an operation, the server picks up the work and carries it through. The exact behavior is heavily dependent on implementation, but if the previous operation was successfully rolled back by way of an [ACID](https://en.wikipedia.org/wiki/ACID) database, it’ll be safe to retry it wholesale. Otherwise, state is recovered and the call is continued.
|
||||
- On a response failure (i.e. the operation executed successfully, but the client couldn’t get the result), the server simply replies with a cached result of the successful operation.
|
||||
|
||||
The Stripe API implements [idempotency keys](https://stripe.com/docs/api#idempotent_requests) on mutating endpoints (i.e. anything under `POST` in our case) by allowing clients to pass a unique value in with the special `Idempotency-Key` header, which allows a client to guarantee the safety of distributed operations:
|
||||
|
||||
If the above Stripe request fails due to a network connection error, you can safely retry it with the same idempotency key, and the customer is charged only once.
|
||||
|
||||
## Being a good distributed citizen
|
||||
|
||||
Safely handling failure is hugely important, but beyond that, it’s also recommended that it be handled in a considerate way. When a client sees that a network operation has failed, there’s a good chance that it’s due to an intermittent failure that will be gone by the next retry. However, there’s also a chance that it’s a more serious problem that’s going to be more tenacious; for example, if the server is in the middle of an incident that’s causing hard downtime. Not only will retries of the operation not go through, but they may contribute to further degradation.
|
||||
|
||||
It’s usually recommended that clients follow something akin to an [exponential backoff](https://en.wikipedia.org/wiki/Exponential_backoff) algorithm as they see errors. The client blocks for a brief initial wait time on the first failure, but as the operation continues to fail, it waits proportionally to `2^n`, where *n* is the number of failures that have occurred. By backing off exponentially, we can ensure that clients aren’t hammering on a downed server and contributing to the problem.
|
||||
|
||||
Exponential backoff has a long and interesting [history](http://www.cs.utexas.edu/users/lam/NRL/backoff.html) in computer networking.
|
||||
|
||||
Furthermore, it’s also a good idea to mix in an element of randomness. If a problem with a server causes a large number of clients to fail at close to the same time, then even with back off, their retry schedules could be aligned closely enough that the retries will hammer the troubled server. This is known as [the thundering herd problem](https://en.wikipedia.org/wiki/Thundering_herd_problem).
|
||||
|
||||
We can address thundering herd by adding some amount of random “jitter” to each client’s wait time. This will space out requests across all clients, and give the server some breathing room to recover.
|
||||
|
||||
Thundering herd problem when a server faces simultaneous retries from all clients.
|
||||
|
||||
The [Stripe Ruby library](https://github.com/stripe/stripe-ruby) retries on failure automatically with an idempotency key using increasing backoff times and jitter. The implementation for that is pretty simple, and [you can refer to it on GitHub](https://github.com/stripe/stripe-ruby/blob/1bb9ac48b916b1c60591795cdb7ba6d18495e82d/lib/stripe/stripe_client.rb#L78-L92) to see exactly how it works.
|
||||
|
||||
## Codifying the design of robust APIs
|
||||
|
||||
Considering the possibility of failure in a distributed system and how to handle it is of paramount importance in building APIs that are both robust and predictable. Retry logic on clients and idempotency on servers are techniques that are useful in achieving this goal and relatively simple to implement in any technology stack.
|
||||
|
||||
Here are a few core principles to follow while designing your clients and APIs:
|
||||
|
||||
- **Make sure that failures are handled consistently** . Have clients retry operations against remote services. Not doing so could leave data in an inconsistent state that will lead to problems down the road.
|
||||
- **Make sure that failures are handled safely** . Use idempotency and idempotency keys to allow clients to pass a unique value and retry requests as needed.
|
||||
- **Make sure that failures are handled responsibly** . Use techniques like exponential backoff and random jitter. Be considerate of servers that may be stuck in a degraded state.
|
||||
@@ -0,0 +1,13 @@
|
||||
# What’s not Actionable & Business Critical Shouldn’t Ring: Building the Right Alerting System
|
||||
|
||||
- **期号**: SRE Weekly Issue #72(2017-05-14)
|
||||
- **作者**: —
|
||||
- **链接**: https://thoughts.t37.net/whats-not-actionable-business-critical-shouldn-t-ring-building-the-right-alerting-system-e8f4b085a2cb?__s=bwykwk1kcceogszq8abt
|
||||
|
||||
## 简介
|
||||
|
||||
Here’s an account of an effort to move from server-based paging (this server is down) to functional-based alerting (this user action isn’t working), with a resulting impressive reduction in out-of-hours paging.
|
||||
|
||||
## 正文
|
||||
|
||||
> ⚠️ 抓取失败:URLError: [SSL: UNEXPECTED_EOF_WHILE_READING] EOF occurred in violation of protocol (_ssl.c:1032)
|
||||
110
sreweekly/markdown/72/03-cpu-utilization-is-wrong.md
Normal file
110
sreweekly/markdown/72/03-cpu-utilization-is-wrong.md
Normal file
@@ -0,0 +1,110 @@
|
||||
# CPU Utilization is Wrong
|
||||
|
||||
- **期号**: SRE Weekly Issue #72(2017-05-14)
|
||||
- **作者**: —
|
||||
- **链接**: http://www.brendangregg.com/blog/2017-05-09/cpu-utilization-is-wrong.html
|
||||
|
||||
## 简介
|
||||
|
||||
It pays to study up and deeply understand what a simple metric like “cpu utilization” really means.
|
||||
|
||||
## 正文
|
||||
|
||||
The metric we all use for CPU utilization is deeply misleading, and getting worse every year. What is CPU utilization? How busy your processors are? No, that's not what it measures. Yes, I'm talking about the "%CPU" metric used *everywhere*, by *everyone*. In every performance monitoring product. In top(1).
|
||||
|
||||
What you may think 90% CPU utilization means:
|
||||
|
||||

|
||||
|
||||
What it might really mean:
|
||||
|
||||

|
||||
|
||||
Stalled means the processor was not making forward progress with instructions, and usually happens because it is waiting on memory I/O. The ratio I drew above (between busy and stalled) is what I typically see in production. Chances are, you're mostly stalled, but don't know it.
|
||||
|
||||
What does this mean for you? Understanding how much your CPUs are stalled can direct performance tuning efforts between reducing code or reducing memory I/O. Anyone looking at CPU performance, especially on clouds that auto scale based on CPU, would benefit from knowing the stalled component of their %CPU.
|
||||
|
||||
## What really is CPU Utilization?
|
||||
|
||||
The metric we call CPU utilization is really "non-idle time": the time the CPU was not running the idle thread. Your operating system kernel (whatever it is) usually tracks this during context switch. If a non-idle thread begins running, then stops 100 milliseconds later, the kernel considers that CPU utilized that entire time.
|
||||
|
||||
This metric is as old as time sharing systems. The Apollo Lunar Module guidance computer (a pioneering time sharing system) called its idle thread the "DUMMY JOB", and engineers tracked cycles running it vs real tasks as a important computer utilization metric. (I wrote about this [before](http://www.brendangregg.com/usemethod.html#Apollo).)
|
||||
|
||||
So what's wrong with this?
|
||||
|
||||
Nowadays, CPUs have become much faster than main memory, and waiting on memory dominates what is still called "CPU utilization". When you see high %CPU in top(1), you might think of the processor as being the bottleneck – the CPU package under the heat sink and fan – when it's really those banks of DRAM.
|
||||
|
||||
This has been getting worse. For a long time processor manufacturers were scaling their clockspeed quicker than DRAM was scaling its access latency (the "CPU DRAM gap"). That levelled out around 2005 with 3 GHz processors, and since then processors have scaled using more cores and hyperthreads, plus multi-socket configurations, all putting more demand on the memory subsystem. Processor manufacturers have tried to reduce this memory bottleneck with larger and smarter CPU caches, and faster memory busses and interconnects. But we're still usually stalled.
|
||||
|
||||
## How to tell what the CPUs are really doing
|
||||
|
||||
By using Performance Monitoring Counters (PMCs): hardware counters that can be read using [Linux perf](http://www.brendangregg.com/perf.html), and other tools. For example, measuring the entire system for 10 seconds:
|
||||
|
||||
# **perf stat -a -- sleep 10**
|
||||
Performance counter stats for 'system wide':
|
||||
641398.723351 task-clock (msec) # 64.116 CPUs utilized (100.00%)
|
||||
379,651 context-switches # 0.592 K/sec (100.00%)
|
||||
51,546 cpu-migrations # 0.080 K/sec (100.00%)
|
||||
13,423,039 page-faults # 0.021 M/sec
|
||||
1,433,972,173,374 cycles # 2.236 GHz (75.02%)
|
||||
<not supported> stalled-cycles-frontend
|
||||
<not supported> stalled-cycles-backend
|
||||
1,118,336,816,068 instructions # **0.78 insns per cycle** (75.01%)
|
||||
249,644,142,804 branches # 389.218 M/sec (75.01%)
|
||||
7,791,449,769 branch-misses # 3.12% of all branches (75.01%)
|
||||
10.003794539 seconds time elapsed
|
||||
|
||||
The key metric here is **instructions per cycle** (`insns per cycle`: IPC), which shows on average how many instructions we were completed for each CPU clock cycle. The higher, the better (a simplification). The above example of 0.78 sounds not bad (78% busy?) until you realize that this processor's top speed is an IPC of 4.0. This is also known as *4-wide*, referring to the instruction fetch/decode path. Which means, the CPU can retire (complete) four instructions with every clock cycle. So an IPC of 0.78 on a 4-wide system, means the CPUs are running at 19.5% their top speed. Newer Intel processors may move to 5-wide.
|
||||
|
||||
There are hundreds more PMCs you can use to dig further: measuring stalled cycles directly by different types.
|
||||
|
||||
### In the cloud
|
||||
|
||||
If you are in a virtual environment, you might not have access to PMCs, depending on whether the hypervisor supports them for guests. I recently posted about [The PMCs of EC2: Measuring IPC](http://www.brendangregg.com/blog/2017-05-04/the-pmcs-of-ec2.html), showing how PMCs are now available for dedicated host types on the AWS EC2 Xen-based cloud.
|
||||
|
||||
## Interpretation and actionable items
|
||||
|
||||
If your **IPC is < 1.0**, you are likely memory stalled, and software tuning strategies include reducing memory I/O, and improving CPU caching and memory locality, especially on NUMA systems. Hardware tuning includes using processors with larger CPU caches, and faster memory, busses, and interconnects.
|
||||
|
||||
If your **IPC is > 1.0**, you are likely instruction bound. Look for ways to reduce code execution: eliminate unnecessary work, cache operations, etc. [CPU flame graphs](http://www.brendangregg.com/FlameGraphs/cpuflamegraphs.html) are a great tool for this investigation. For hardware tuning, try a faster clock rate, and more cores/hyperthreads.
|
||||
|
||||
For my above rules, I split on an IPC of 1.0. Where did I get that from? I made it up, based on my prior work with PMCs. Here's how you can get a value that's custom for your system and runtime: write two dummy workloads, one that is CPU bound, and one memory bound. Measure their IPC, then calculate their mid point.
|
||||
|
||||
## What performance monitoring products should tell you
|
||||
|
||||
Every performance tool should show IPC along with %CPU. Or break down %CPU into instruction-retired cycles vs stalled cycles, eg, %INS and %STL.
|
||||
|
||||
As for top(1), there is tiptop(1) for Linux, which shows IPC by process:
|
||||
|
||||
tiptop - [root]
|
||||
Tasks: 96 total, 3 displayed screen 0: default
|
||||
**PID** [ %CPU] %SYS P Mcycle Minstr **IPC** %MISS %BMIS %BUS **COMMAND**
|
||||
**3897** 35.3 28.5 4 274.06 178.23 **0.65** 0.06 0.00 0.0 **java**
|
||||
**1319+** 5.5 2.6 6 87.32 125.55 **1.44** 0.34 0.26 0.0 **nm-applet**
|
||||
**900** 0.9 0.0 6 25.91 55.55 **2.14** 0.12 0.21 0.0 **dbus-daemo**
|
||||
|
||||
## Other reasons CPU utilization is misleading
|
||||
|
||||
It's not just memory stall cycles that makes CPU utilization misleading. Other factors include:
|
||||
|
||||
- Temperature trips stalling the processor.
|
||||
- Turboboost varying the clockrate.
|
||||
- The kernel varying the clock rate with speed step.
|
||||
- The problem with averages: 80% utilized over 1 minute, hiding bursts of 100%.
|
||||
- Spin locks: the CPU is utilized, and has high IPC, but the app is not making logical forward progress.
|
||||
|
||||
## Update: is CPU utilization actually wrong?
|
||||
|
||||
There have been hundreds of comments on this post, here (below) and elsewhere ([1](https://news.ycombinator.com/item?id=14301739), [2](https://www.reddit.com/r/programming/comments/6a6v8g/cpu_utilization_is_wrong/)). Thanks to everyone for taking the time and the interest in this topic. To summarize my responses: I'm not talking about iowait at all (that's disk I/O), and there are actionable items if you know you are memory bound (see above).
|
||||
|
||||
But is CPU utilization actually wrong, or just deeply misleading? I think many people interpret high %CPU to mean that the processing unit is the bottleneck, which is wrong (as I said earlier). At that point you don't yet know, and it is often something external. Is the metric technically correct? If the CPU stall cycles can't be used by anything else, aren't they are therefore "utilized waiting" (which sounds like an oxymoron)? In some cases, yes, you could say that %CPU as an OS-level metric is technically correct, but deeply misleading. With hyperthreads, however, those stalled cycles can now be used by another thread, so %CPU may count cycles as utilized that are in fact available. That's wrong. In this post I wanted to focus on the interpretation problem and suggested solutions, but yes, there are technical problems with this metric as well.
|
||||
|
||||
You might just say that utilization as a metric was already broken, as Adrian Cockcroft discussed [previously](http://www.hpts.ws/papers/2007/Cockcroft_HPTS-Useless.pdf).
|
||||
|
||||
## Conclusion
|
||||
|
||||
CPU utilization has become a deeply misleading metric: it includes cycles waiting on main memory, which can dominate modern workloads. Perhaps %CPU should be renamed to %CYC, short for cycles. You can figure out what %CPU really means by using additional metrics, including instructions per cycle (IPC). An IPC < 1.0 likely means memory bound, and an IPC > 1.0 likely means instruction bound. I covered IPC in my [previous post](http://www.brendangregg.com/blog/2017-05-04/the-pmcs-of-ec2.html), including an introduction to the Performance Monitoring Counters (PMCs) needed to measure it.
|
||||
|
||||
Performance monitoring products that show %CPU – which is all of them – should also show PMC metrics to explain what that means, and not mislead the end user. For example, they can show %CPU with IPC, and/or instruction-retired cycles vs stalled cycles. Armed with these metrics, developers and operators can choose how to better tune their applications and systems.
|
||||
|
||||
Click here for Disqus comments (ad supported).
|
||||
13
sreweekly/markdown/72/04-aws-service-health-dashboard.md
Normal file
13
sreweekly/markdown/72/04-aws-service-health-dashboard.md
Normal file
@@ -0,0 +1,13 @@
|
||||
# AWS Service Health Dashboard
|
||||
|
||||
- **期号**: SRE Weekly Issue #72(2017-05-14)
|
||||
- **作者**: —
|
||||
- **链接**: http://status.aws.amazon.com/
|
||||
|
||||
## 简介
|
||||
|
||||
Why am I linking to AWS’s status site? Look closely, and you’ll see that the “green checkmark i” symbol has been replaced with a far more noticeable blue circle with a white diamond. Check out the old icon here for comparison. End of an era, or just another way of presenting the same information?
|
||||
|
||||
## 正文
|
||||
|
||||
> ⚠️ 抓取失败:HTTP 503
|
||||
@@ -0,0 +1,13 @@
|
||||
# Circuit breaker and monitoring of a gRPC service in Ruby (Part 1)
|
||||
|
||||
- **期号**: SRE Weekly Issue #72(2017-05-14)
|
||||
- **作者**: —
|
||||
- **链接**: https://medium.com/@shiladitya16/circuit-breaker-and-monitoring-of-a-grpc-service-in-ruby-part-1-7509c7e1356a
|
||||
|
||||
## 简介
|
||||
|
||||
The author introduces a new Ruby gem, grpc-commons that makes it easy to add circuit breaker and statsd support to a grpc client.
|
||||
|
||||
## 正文
|
||||
|
||||
> ⚠️ 抓取失败:HTTP 403
|
||||
@@ -0,0 +1,13 @@
|
||||
# Introducing distributed tracing in your Python application via Zipkin
|
||||
|
||||
- **期号**: SRE Weekly Issue #72(2017-05-14)
|
||||
- **作者**: —
|
||||
- **链接**: http://t.dripemail2.com/c/eyJhY2NvdW50X2lkIjoiMjY1Njc0MyIsImRlbGl2ZXJ5X2lkIjoiODE3NjA4NzUzIiwidXJsIjoiaHR0cDovL2VjaG9yYW5kLm1lL2ludHJvZHVjaW5nLWRpc3RyaWJ1dGVkLXRyYWNpbmctaW4teW91ci1weXRob24tYXBwbGljYXRpb24tdmlhLXppcGtpbi5odG1sP19fcz1id3lrd2sxa2NjZW9nc3pxOGFidCJ9
|
||||
|
||||
## 简介
|
||||
|
||||
Along with being a tutorial on setting up Zipkin with Python, this article also explains some basic Zipkin concepts.
|
||||
|
||||
## 正文
|
||||
|
||||
> ⚠️ 抓取失败:HTTP 404
|
||||
@@ -0,0 +1,13 @@
|
||||
# Announcing the Modern Incident Resolution Lifecycle
|
||||
|
||||
- **期号**: SRE Weekly Issue #72(2017-05-14)
|
||||
- **作者**: —
|
||||
- **链接**: http://t.dripemail2.com/c/eyJhY2NvdW50X2lkIjoiMjY1Njc0MyIsImRlbGl2ZXJ5X2lkIjoiODE3NjA4NzUzIiwidXJsIjoiaHR0cHM6Ly93d3cucGFnZXJkdXR5LmNvbS9ibG9nL21vZGVybi1pbmNpZGVudC1yZXNvbHV0aW9uLWxpZmVjeWNsZS8_X19zPWJ3eWt3azFrY2Nlb2dzenE4YWJ0In0
|
||||
|
||||
## 简介
|
||||
|
||||
PagerDuty is apparently trying to position itself as more than just a paging service, with a few new features around the entire incident lifecycle. I’m especially interested in checking out the new postmortem tooling.
|
||||
|
||||
## 正文
|
||||
|
||||
> ⚠️ 抓取失败:HTTP 404
|
||||
@@ -0,0 +1,13 @@
|
||||
# How we Upgraded a 22TB MySQL Cluster from 5.6 to 5.7 (in 9 months)
|
||||
|
||||
- **期号**: SRE Weekly Issue #72(2017-05-14)
|
||||
- **作者**: —
|
||||
- **链接**: https://thoughts.t37.net/how-we-upgraded-a-22tb-mysql-cluster-from-5-6-to-5-7-in-9-months-cc41b391895d
|
||||
|
||||
## 简介
|
||||
|
||||
I included this article last week, but my link was outdated and returned a 404. Here’s the corrected link — sorry about that!
|
||||
|
||||
## 正文
|
||||
|
||||
> ⚠️ 抓取失败:URLError: [SSL: UNEXPECTED_EOF_WHILE_READING] EOF occurred in violation of protocol (_ssl.c:1032)
|
||||
@@ -0,0 +1,119 @@
|
||||
# A first look at Elastic’s new Machine Learning Technology
|
||||
|
||||
- **期号**: SRE Weekly Issue #72(2017-05-14)
|
||||
- **作者**: —
|
||||
- **链接**: https://www.linkedin.com/pulse/first-look-elastics-new-machine-learning-technology-robert-cowart
|
||||
|
||||
## 简介
|
||||
|
||||
I put a call out for a review of Elastic’s new beta anomaly detection feature last week, and here one is! Thanks to an Elastic employee for forwarding this link to me.
|
||||
|
||||
## 正文
|
||||
|
||||
# A first look at Elastic's new Machine Learning Technology
|
||||
|
||||
Over the past year the hottest topics in tech have without doubt been machine learning and artificial intelligence. In September of last year [Elastic](https://www.elastic.co) entered the game with its acquisition of Prelert and their machine learning-based anomaly detection technology.
|
||||
|
||||
As I have previously [discussed](https://www.linkedin.com/pulse/wtflow-you-really-still-paying-commercial-solutions-collect-cowart), Elastic Stack is right at home in the world of Digital Infrastructure Management. For a while now they have provided basic alerting functionality with the [Watcher](https://www.elastic.co/products/x-pack/alerting) capabilities of their X-Pack offering, but have lacked more sophisticated methods of extracting deep insights from the massive quantities of data their solutions ingest. Over the last seven months they have been hard at work integrating Prelert into Elastic Stack to fill this gap. Have they succeeded? Let's find out!
|
||||
|
||||
## Now in beta...
|
||||
|
||||
The Prelert technology has been released as an additional [X-Pack](https://www.elastic.co/products/x-pack) component, appropriately named [*Machine Learning*](https://www.elastic.co/products/x-pack/machine-learning). It is considered a "beta" release, which I understand to mean that it shouldn't yet be implemented within production deployments. During my testing I had zero issues with stability, errors or any obvious bugs. However machine learning is by nature a resource intensive process and it is wise to follow these guidelines, until more experience is gained on the additional loads that can be expected.
|
||||
|
||||
## Things that I liked...
|
||||
|
||||
### 1. Wizards
|
||||
|
||||
A few months back I spent a little time playing with the Prelert's offering, and while the potential was obvious, configuring analytics jobs was difficult, especially for those new to the technology, as I was. This has been addressed with very useful wizards for both single and multi metric jobs. For those who "get it" the option to create an advanced job is always available. I started with the wizards and by investigating the resulting jobs was able to much more quickly understand how to leverage the various advanced options.
|
||||
|
||||
### 2. Ability to run jobs against historical data
|
||||
|
||||
The anomaly detection functionality provided by many other solutions works only on new data that occurs after they are activated. Elastic's Machine Learning jobs can be scheduled against all historical data available within your Elasticsearch cluster. Not only does this allow the system to quickly and **accurately** learn what is "normal", it also greatly assists the process of designing and testing jobs without having to figure out how you will recreate the test scenario consistently.
|
||||
|
||||
### 3. Seamless transition to realtime analysis
|
||||
|
||||
Once you are happy that your job is configured and working as intended against historical data, you can schedule it with an open-ended stop time for continuing realtime analysis. The job will periodically analyse recently collected data, identifying any anomalies.
|
||||
|
||||
### 4. Detected anomalies are accessible in an index
|
||||
|
||||
This might seem like a weird thing to mention. After all the provided Anomaly Explorer is a great tool (as you will see later). However for those of us with additional use-cases we can access the analysis results directly from the Elasticsearch index where they are stored. Since this is an index like any other, you can leverage any other tool or method to work with this data. For example, you might configure Watcher to generate anomaly-based alerts, or create Kibana visualizations for use in dashboards.
|
||||
|
||||
The above visualization uses the legacy Prelert swimlane plugin for Kibana, displaying data from the Machine Learning anomalies index.
|
||||
|
||||
### 5. The ability to launch URLs from the Anomaly Explorer
|
||||
|
||||
If you are like me you take a lot of pride in the value provided by the user dashboards that are part of the Elastic Stack based solutions you have developed. Machine Learning jobs can be configured with URLs which can be launched from the Anomaly Explorer. Values from the jobs are passed into variables in the URL to facilitate launch in context. With this capability you can allow users to navigate from an anomaly to dashboards that contain the raw data which caused it.
|
||||
|
||||
## Machine Learning in Action
|
||||
|
||||
By now I imagine you are thinking "Enough banter already... let's see this thing in action!" OK OK! But FIRST... a word about our data source.
|
||||
|
||||
### About BlackRidge
|
||||
|
||||
The data I will feed through Machine Learning is from [BlackRidge](https://www.blackridge.us) Transport Access Control Gateways. BlackRidge's [TAC products](https://www.blackridge.us/products) provide network security and cyber defense that stops cyber-attacks and protects against insider threats. BlackRidge is the only network security technology of its kind that can stop a would-be cyber attack in its tracks, while at the same time permitting all authenticated and authorized traffic flows.
|
||||
|
||||
BlackRidge TAC gateways provide a comprehensive logging capability, which includes logging of all permitted connections as well discarded attempts. The BlackRidge Message Format (BMF) is also IBM LEEF compatible and formatted for easy integration with 3rd-party tools.
|
||||
|
||||
### The Environment
|
||||
|
||||
BlackRidge provides an Elastic Stack integration which consists of the necessary Logstash pipeline and Kibana dashboards. My lab environment includes this integration and the X-Pack Machine Learning beta. Data is arriving live from both a BlackRidge TAC hardware appliance, protecting datacenter resources, and a virtual appliance running in AWS and protecting cloud-based applications.
|
||||
|
||||
### A first Machine Learning job
|
||||
|
||||
As the BlackRidge gateway provides us with logs for all connection attempts, a security related machine learning task is the obvious choice for our first job. So let's look at a job to identify port scanning activity. Note that we will be jumping in with both feet and reviewing the advanced options.
|
||||
|
||||
As mentioned the Job Details allow you to define one or more URLs which are available in the Anomaly Explorer to launch from the anomaly into other dashboards or 3rd-party tools.
|
||||
|
||||
At the heart of the machine learning job is *Detectors*. A Detector defines the data that will be analyzed and the context of that analysis. Detectors are configured in the detector pop-up window.
|
||||
|
||||
The *function* field specifies the [analytical function](https://www.elastic.co/guide/en/x-pack/current/ml-concepts.html#ml-functions) used to process the data. The *field_name* is the field that is processed by the function. To detect a port scan we want to look for excessive counts of unique destination ports that are accessed, or... *high_distinct_count(destPort)*
|
||||
|
||||
By default the entire dataset is handled as a single time-series. This isn't really helpful for detecting port scans. What we need is to break our data out into separate time-series for each destination. We can achieve this by setting the *partition_field_name* to the field holding destination addresses. Each of the resulting time-series will be analyzed independent of the others.
|
||||
|
||||
The *by_field_name* and *over_field_name* fields are used to further define the context within which the data is analyzed for anomalies.
|
||||
|
||||
Using *by_field_name* will cause the data to be further grouped by the unique values (entities) of the specified field. An anomaly in this context is data that deviates significantly from **past data for this entity**.
|
||||
|
||||
The *over_field_name* will also cause data to be further grouped by the unique values (entities) of the specified field. However in this context an anomaly is data that deviates from **the data of the other entities**.
|
||||
|
||||
For the port scan detection job we want to further group our data by the *src* field. For a relevant analysis we want to compare the behavior (the data) of each source against the normal behavior of the other sources. So we want to use the *over_field_name* option.
|
||||
|
||||
Finally, each of the fields that we used to create our Detector would be considered Influencers and should be defined as such in the Analysis Configuration.
|
||||
|
||||
After saving the job you have the opportunity to run it. I had about 10 days of data in my lab environment so I ran it from the beginning of time. After the job completes you can view the results in the Anomaly Explorer or Single Metric Viewer.
|
||||
|
||||
### The Anomaly Explorer
|
||||
|
||||
The Anomaly Explorer contains swim lanes showing the maximum anomaly score over time. There is an overall swim lane that shows the overall score for the job, and also swim lanes for each influencer.
|
||||
|
||||
The Anomaly Explorer allows us to easily identify occurrences of port scans detected by our job. By selecting a block in a swim lane, the anomaly details are displayed alongside the original source data (where applicable).
|
||||
|
||||
To dig deeper into the details of the anomaly, use the URLs that we configured in the job to navigate in-context to the dashboards provided by the BlackRidge integration.
|
||||
|
||||
By leveraging the time-span and field data from the anomaly we land on a dashboard that is focused on the related raw data. We can clearly see the port scanning activity in the raw data. Following our other URL similarly brings us directly to the geo location of the threat.
|
||||
|
||||
The BlackRidge integration for Elastic Stack includes a number of additional dashboards that provide insights into network access attempts in your TAC-protected environment. For example the Connections Analyzer dashboards provide details about frequent conversations. The anomaly detection capabilities of Elastic's Machine Learning allow us to easily focus-in on the port scan within our collected data.
|
||||
|
||||
The great news is that we can confirm that BlackRidge's TAC appliance discarded the connection attempts related to this threat and the environment remained 100% secure!
|
||||
|
||||
CONGRATS! Our Machine Learning job works, is integrated with our user dashboards and is allowing us to quickly identify and focus on anomalies in our environment!!! We can now schedule the job to run continuously and detect port scanning activity as it occurs in the future.
|
||||
|
||||
## Conclusions
|
||||
|
||||
I must admit that I am really impressed with Elastic's Machine Learning beta. It is already well integrated with the rest of the Elastic Stack and delivered valuable insights quickly once the brief learning curve was overcome.
|
||||
|
||||
There are a few minor issues, most of which are related to confusing terminology or missing documentation. But this is a "beta" and I am confident that the Elastic team will get these things corrected in the near future.
|
||||
|
||||
I am definitely looking forward to apply the technology to additional data sources and for a broader set of use-cases!
|
||||
|
||||
I would love to hear your thoughts!
|
||||
|
||||
From what you have written, it seems like xpack is working more along the lines of statistical analysis rather than that of Machine Learning. Of course, I may be wrong but nothing in your article suggests that ML is actually being used. Great article though.
|
||||
|
||||
Great article Robert Cowart, one important thing is that Machine Learning is easy to configure in Elastic Stack, this isn't often the case in other software.
|
||||
|
||||
The beginning of something great.
|
||||
|
||||
Thanks, awesome!
|
||||
|
||||
Thanks Robert, excellent write up
|
||||
33
sreweekly/markdown/72/10-it-outages-who-s-really-at-fault.md
Normal file
33
sreweekly/markdown/72/10-it-outages-who-s-really-at-fault.md
Normal file
@@ -0,0 +1,33 @@
|
||||
# IT Outages, Who’s Really at Fault?
|
||||
|
||||
- **期号**: SRE Weekly Issue #72(2017-05-14)
|
||||
- **作者**: —
|
||||
- **链接**: http://www.informationweek.com/strategic-cio/executive-insights-and-innovation/it-outages-whos-really-at-fault/a/d-id/1328869?_mc=RSS_IWK_EDT
|
||||
|
||||
## 简介
|
||||
|
||||
This article cautions one to be careful to look past an obvious root cause, because a deeper systemic or policy problem may be lurking behind it.
|
||||
|
||||
## 正文
|
||||
|
||||
# IT Outages, Who's Really at Fault?
|
||||
|
||||
Systems do go down, and sometimes the cause seems obvious, but it may be too obvious. Employ root cause analysis methods to find the real cause of failure.
|
||||
|
||||

|
||||
|
||||
As our IT infrastructures grow increasingly complex thanks to advanced technologies such as [virtualization, cloud computing](http://www.informationweek.com/it-infrastructure/achieve-application-and-data-visibility-in-the-cloud) and software defined networking (SDN), understanding the root cause of an IT outage becomes more difficult to achieve. But even more importantly, current troubleshooting techniques to find fault into an outage focuses solely on the technical side of the IT department. In fact, the truth is that many root causes go beyond technology and stem from poor policy and management decisions.
|
||||
|
||||
For anyone who has been involved in enterprise IT support, root cause analysis (RCA) is one of the first troubleshooting methodologies one needs to learn. Thanks to the ever-increasing complexities of network infrastructures and distributed computing platforms, the visible symptom that end users experience is usually not the true cause of the problem. Instead, RCA teaches us to continue drilling down into the cause-and-effect chain to ultimately find the core issue at hand.
|
||||
|
||||
There are multiple [RCA approaches](http://asq.org/learn-about-quality/root-cause-analysis/overview/root-cause-approaches.html) that one can use. The problem that I often see is that when training to use one of the many RCA tools and techniques, the focus of a root cause is typically fixated on one of two areas. First, the root cause reported is often found to be a hardware or software related error on the production infrastructure. Second, the root cause was due to a human error caused by a misconfiguration or poor communication between team members.
|
||||
|
||||
In many cases, one of these two areas are indeed the true root cause of the problem. Once discovered, the cause can be documented and fixed, and the continuous improvement cycle starts over. But in some situations, finding the core of the problem requires a different perspective. Because RCA methods ask us to constantly drill down into a problem, we never take a step back and look at it from a big picture perspective. That’s precisely what needs to be done.
|
||||
|
||||
{image 1}
|
||||
|
||||
For example, if an outage was caused by a hardware failure somewhere on the network, was the true root cause due to faulty gear, or was the hardware past its life expectancy? If the latter is the case, one must then consider why hardware that outlived its [mean time between failure (MTBF)](http://whatis.techtarget.com/definition/MTBF-mean-time-between-failures) is still being relied upon in a production environment. If one continues digging, they may discover that it was previously recommended by IT support staff that this hardware be replaced long ago – but that budget dollars never materialized.
|
||||
|
||||
Another popular root cause that often goes overlooked deals with staffing within the IT department. IT administrators have tremendous responsibilities when it relates to the uptime of an enterprise network. With just a few keystrokes or clicks of a mouse, an admin can [inadvertently bring an infrastructure to its knees](http://www.informationweek.com/it-infrastructure/9-spectacular-cloud-computing-fails). While it’s often easy to simply lay blame on the administrator who made the mistake, it’s important to look more deeply at why the misstep was made in the first place. Do they have to proper training to competently perform their administration duties? Did the admin just complete a marathon work shift and was simply not thinking straight? In situations such as these, policy and proper IT management could have avoided the outage.
|
||||
|
||||
So, the next time you are reviewing an RCA report for an outage, make sure that the root cause indicated truly takes the troubleshooting process as far as the cause-and-effect chain can go. Despite the potentially uncomfortable situation of pointing out faults in management as the root cause, you owe it to your organization to find and fix these types of problems to keep them from occurring repeatedly. Only then does the RCA process perform the way it was intended.
|
||||
@@ -0,0 +1,13 @@
|
||||
# Watch out for serverless computing’s blind spot
|
||||
|
||||
- **期号**: SRE Weekly Issue #72(2017-05-14)
|
||||
- **作者**: —
|
||||
- **链接**: http://www.infoworld.com/article/3196133/cloud-computing/watch-out-for-serverless-computings-blind-spot.html
|
||||
|
||||
## 简介
|
||||
|
||||
Serverless / FaaS abstract away traditional provisioning, and they make it really easy to ignore planning for resource usage.
|
||||
|
||||
## 正文
|
||||
|
||||
> ⚠️ 抓取失败:HTTP 404
|
||||
@@ -0,0 +1,35 @@
|
||||
# Safety Moment – Are Accidents a Failure of Imagination? | PreAccident Investigation Podcast
|
||||
|
||||
- **期号**: SRE Weekly Issue #72(2017-05-14)
|
||||
- **作者**: —
|
||||
- **链接**: http://preaccidentpodcast.podbean.com/e/safety-moment-are-accidents-a-failure-of-imagination/
|
||||
|
||||
## 简介
|
||||
|
||||
Wow, what a concept:
|
||||
|
||||
> you can think of […] reliable systems […] as successfully imagining all of the potential things that could go wrong
|
||||
|
||||
This 2.5-minute podcast from Todd Conklin has a really great question: to achieve reliability, do we have to try to imagine in advance all of the possible ways our systems could fail?
|
||||
|
||||
## 正文
|
||||
|
||||

|
||||
|
||||
May 10, 2017
|
||||
|
||||
# Safety Moment - Are Accidents a Failure of Imagination?
|
||||
|
||||
Best Safety Podcast, Safety Program, Safety Storytelling, Investigations, Human Performance, Safety Differently, Operational Excellence, Resilience Engineering, Safety and Resilience Incentives
|
||||
|
||||
We are good a prevention.
|
||||
|
||||
We are not so good at prediction.
|
||||
|
||||
Can we ever possibly know all the things that could happen wrong in a system that mostly happens right? It is a good question, one that needs some time and thinking by YOU.
|
||||
|
||||
This podcast is the beginning of a long-term discussion about prevention and prediction. Get ready. Fasten your psycholicial seat belts.... we are going for a ride.
|
||||
|
||||
Thanks for listening. Still looking for a sponsor. Any comments? Ping [Toddconklin@gmail.com](mailto:Toddconklin@gmail.com). Thanks for being you - don't stop being you - you are the best you there is... Gracias.
|
||||
|
||||
No comments yet. Be the first to say something!
|
||||
@@ -0,0 +1,13 @@
|
||||
# “The Scariest Moment of My Life” – BWH Safety Matters
|
||||
|
||||
- **期号**: SRE Weekly Issue #72(2017-05-14)
|
||||
- **作者**: —
|
||||
- **链接**: https://bwhsafetymatters.org/2017/05/09/the-scariest-moment-of-my-life/
|
||||
|
||||
## 简介
|
||||
|
||||
A patient was given an incorrect syringe resulting in a 5x insulin overdose. Brigham and Women’s Hospital reports on the accident and what they’re doing to prevent mistakes of this sort in the future.
|
||||
|
||||
## 正文
|
||||
|
||||
> ⚠️ 抓取失败:HTTP 404
|
||||
@@ -0,0 +1,27 @@
|
||||
# PagerDuty’s 2017 State of Digital Operations Report
|
||||
|
||||
- **期号**: SRE Weekly Issue #72(2017-05-14)
|
||||
- **作者**: —
|
||||
- **链接**: https://www.pagerduty.com/resources/reports/digital-operations/
|
||||
|
||||
## 简介
|
||||
|
||||
> Consumers today have increasingly high expectations for digital applications and service performance, but do IT personnel feel equipped to rise to the occasion? In this survey, we uncover the extent of the digital services expectation gap between consumers and IT teams as well as top strategies teams are using to solve digital disruption challenges.
|
||||
|
||||
## 正文
|
||||
|
||||
Stressful. Demanding. Tiring. No matter what you call it, building and operating digital services in 2020 has been a hard job. From unprecedented surges in traffic and incidents to longer hours of unplanned work, teams everywhere are grappling with a harsh new digital reality.
|
||||
|
||||
Check out our recent global survey of 700 developers and IT operations professionals to get a sense of the feelings and realities we’re all facing today.
|
||||
|
||||
Key findings include:
|
||||
|
||||
- 4 out of 5 have seen increased pressure on digital services, and over half say it has reached never-before-seen levels
|
||||
- 3 out of 5 are working an additional 10+ hours per week, and 2 out of 5 think burnout will be an issue in the near future
|
||||
- 4 out of 5 feel automation can help with on-call and incident management—but opinions are more divided about the value in implementing AIOps
|
||||
|
||||
Read the report now!
|
||||
|
||||
"The PagerDuty Operations Cloud is critical for TUI. This is what is actually going to help us grow as a business when it comes to making sure that we provide quality services for our customers."
|
||||
|
||||
- Yasin Quareshy, Head of Technology at TUI
|
||||
Reference in New Issue
Block a user