Files
nexus/sreweekly/markdown/72/09-a-first-look-at-elastic-s-new-machine-learning-technology.md
2026-09-12 17:23:01 +08:00

120 lines
12 KiB
Markdown
Raw Permalink Blame History

This file contains invisible Unicode characters
This file contains invisible Unicode characters that are indistinguishable to humans but may be processed differently by a computer. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# A first look at Elastic’s new Machine Learning Technology
- **期号**: SRE Weekly Issue #72(2017-05-14)
- **作者**: —
- **链接**: https://www.linkedin.com/pulse/first-look-elastics-new-machine-learning-technology-robert-cowart
## 简介
I put a call out for a review of Elastic’s new beta anomaly detection feature last week, and here one is! Thanks to an Elastic employee for forwarding this link to me.
## 正文
# A first look at Elastic's new Machine Learning Technology
Over the past year the hottest topics in tech have without doubt been machine learning and artificial intelligence. In September of last year [Elastic](https://www.elastic.co) entered the game with its acquisition of Prelert and their machine learning-based anomaly detection technology.
As I have previously [discussed](https://www.linkedin.com/pulse/wtflow-you-really-still-paying-commercial-solutions-collect-cowart), Elastic Stack is right at home in the world of Digital Infrastructure Management. For a while now they have provided basic alerting functionality with the [Watcher](https://www.elastic.co/products/x-pack/alerting) capabilities of their X-Pack offering, but have lacked more sophisticated methods of extracting deep insights from the massive quantities of data their solutions ingest. Over the last seven months they have been hard at work integrating Prelert into Elastic Stack to fill this gap. Have they succeeded? Let's find out!
## Now in beta...
The Prelert technology has been released as an additional [X-Pack](https://www.elastic.co/products/x-pack) component, appropriately named [*Machine Learning*](https://www.elastic.co/products/x-pack/machine-learning). It is considered a "beta" release, which I understand to mean that it shouldn't yet be implemented within production deployments. During my testing I had zero issues with stability, errors or any obvious bugs. However machine learning is by nature a resource intensive process and it is wise to follow these guidelines, until more experience is gained on the additional loads that can be expected.
## Things that I liked...
### 1. Wizards
A few months back I spent a little time playing with the Prelert's offering, and while the potential was obvious, configuring analytics jobs was difficult, especially for those new to the technology, as I was. This has been addressed with very useful wizards for both single and multi metric jobs. For those who "get it" the option to create an advanced job is always available. I started with the wizards and by investigating the resulting jobs was able to much more quickly understand how to leverage the various advanced options.
### 2. Ability to run jobs against historical data
The anomaly detection functionality provided by many other solutions works only on new data that occurs after they are activated. Elastic's Machine Learning jobs can be scheduled against all historical data available within your Elasticsearch cluster. Not only does this allow the system to quickly and **accurately** learn what is "normal", it also greatly assists the process of designing and testing jobs without having to figure out how you will recreate the test scenario consistently.
### 3. Seamless transition to realtime analysis
Once you are happy that your job is configured and working as intended against historical data, you can schedule it with an open-ended stop time for continuing realtime analysis. The job will periodically analyse recently collected data, identifying any anomalies.
### 4. Detected anomalies are accessible in an index
This might seem like a weird thing to mention. After all the provided Anomaly Explorer is a great tool (as you will see later). However for those of us with additional use-cases we can access the analysis results directly from the Elasticsearch index where they are stored. Since this is an index like any other, you can leverage any other tool or method to work with this data. For example, you might configure Watcher to generate anomaly-based alerts, or create Kibana visualizations for use in dashboards.
The above visualization uses the legacy Prelert swimlane plugin for Kibana, displaying data from the Machine Learning anomalies index.
### 5. The ability to launch URLs from the Anomaly Explorer
If you are like me you take a lot of pride in the value provided by the user dashboards that are part of the Elastic Stack based solutions you have developed. Machine Learning jobs can be configured with URLs which can be launched from the Anomaly Explorer. Values from the jobs are passed into variables in the URL to facilitate launch in context. With this capability you can allow users to navigate from an anomaly to dashboards that contain the raw data which caused it.
## Machine Learning in Action
By now I imagine you are thinking "Enough banter already... let's see this thing in action!" OK OK! But FIRST... a word about our data source.
### About BlackRidge
The data I will feed through Machine Learning is from [BlackRidge](https://www.blackridge.us) Transport Access Control Gateways. BlackRidge's [TAC products](https://www.blackridge.us/products) provide network security and cyber defense that stops cyber-attacks and protects against insider threats. BlackRidge is the only network security technology of its kind that can stop a would-be cyber attack in its tracks, while at the same time permitting all authenticated and authorized traffic flows.
BlackRidge TAC gateways provide a comprehensive logging capability, which includes logging of all permitted connections as well discarded attempts. The BlackRidge Message Format (BMF) is also IBM LEEF compatible and formatted for easy integration with 3rd-party tools.
### The Environment
BlackRidge provides an Elastic Stack integration which consists of the necessary Logstash pipeline and Kibana dashboards. My lab environment includes this integration and the X-Pack Machine Learning beta. Data is arriving live from both a BlackRidge TAC hardware appliance, protecting datacenter resources, and a virtual appliance running in AWS and protecting cloud-based applications.
### A first Machine Learning job
As the BlackRidge gateway provides us with logs for all connection attempts, a security related machine learning task is the obvious choice for our first job. So let's look at a job to identify port scanning activity. Note that we will be jumping in with both feet and reviewing the advanced options.
As mentioned the Job Details allow you to define one or more URLs which are available in the Anomaly Explorer to launch from the anomaly into other dashboards or 3rd-party tools.
At the heart of the machine learning job is *Detectors*. A Detector defines the data that will be analyzed and the context of that analysis. Detectors are configured in the detector pop-up window.
The *function* field specifies the [analytical function](https://www.elastic.co/guide/en/x-pack/current/ml-concepts.html#ml-functions) used to process the data. The *field_name* is the field that is processed by the function. To detect a port scan we want to look for excessive counts of unique destination ports that are accessed, or... *high_distinct_count(destPort)*
By default the entire dataset is handled as a single time-series. This isn't really helpful for detecting port scans. What we need is to break our data out into separate time-series for each destination. We can achieve this by setting the *partition_field_name* to the field holding destination addresses. Each of the resulting time-series will be analyzed independent of the others.
The *by_field_name* and *over_field_name* fields are used to further define the context within which the data is analyzed for anomalies.
Using *by_field_name* will cause the data to be further grouped by the unique values (entities) of the specified field. An anomaly in this context is data that deviates significantly from **past data for this entity**.
The *over_field_name* will also cause data to be further grouped by the unique values (entities) of the specified field. However in this context an anomaly is data that deviates from **the data of the other entities**.
For the port scan detection job we want to further group our data by the *src* field. For a relevant analysis we want to compare the behavior (the data) of each source against the normal behavior of the other sources. So we want to use the *over_field_name* option.
Finally, each of the fields that we used to create our Detector would be considered Influencers and should be defined as such in the Analysis Configuration.
After saving the job you have the opportunity to run it. I had about 10 days of data in my lab environment so I ran it from the beginning of time. After the job completes you can view the results in the Anomaly Explorer or Single Metric Viewer.
### The Anomaly Explorer
The Anomaly Explorer contains swim lanes showing the maximum anomaly score over time. There is an overall swim lane that shows the overall score for the job, and also swim lanes for each influencer.
The Anomaly Explorer allows us to easily identify occurrences of port scans detected by our job. By selecting a block in a swim lane, the anomaly details are displayed alongside the original source data (where applicable).
To dig deeper into the details of the anomaly, use the URLs that we configured in the job to navigate in-context to the dashboards provided by the BlackRidge integration.
By leveraging the time-span and field data from the anomaly we land on a dashboard that is focused on the related raw data. We can clearly see the port scanning activity in the raw data. Following our other URL similarly brings us directly to the geo location of the threat.
The BlackRidge integration for Elastic Stack includes a number of additional dashboards that provide insights into network access attempts in your TAC-protected environment. For example the Connections Analyzer dashboards provide details about frequent conversations. The anomaly detection capabilities of Elastic's Machine Learning allow us to easily focus-in on the port scan within our collected data.
The great news is that we can confirm that BlackRidge's TAC appliance discarded the connection attempts related to this threat and the environment remained 100% secure!
CONGRATS! Our Machine Learning job works, is integrated with our user dashboards and is allowing us to quickly identify and focus on anomalies in our environment!!! We can now schedule the job to run continuously and detect port scanning activity as it occurs in the future.
## Conclusions
I must admit that I am really impressed with Elastic's Machine Learning beta. It is already well integrated with the rest of the Elastic Stack and delivered valuable insights quickly once the brief learning curve was overcome.
There are a few minor issues, most of which are related to confusing terminology or missing documentation. But this is a "beta" and I am confident that the Elastic team will get these things corrected in the near future.
I am definitely looking forward to apply the technology to additional data sources and for a broader set of use-cases!
I would love to hear your thoughts!
From what you have written, it seems like xpack is working more along the lines of statistical analysis rather than that of Machine Learning. Of course, I may be wrong but nothing in your article suggests that ML is actually being used. Great article though.
Great article Robert Cowart, one important thing is that Machine Learning is easy to configure in Elastic Stack, this isn't often the case in other software.
The beginning of something great.
Thanks, awesome!
Thanks Robert, excellent write up