Detecting Reconnaissance and Discovery Tactics from the MITRE ATT&CK Framework in Zeek Conn Logs Using Spark's Machine Learning in the Big Data Framework

Sikha Bagui; Dustin Mink; Subhash Bagui; Tirthankar Ghosh; Tom McElroy; Esteban Paredes; Nithisha Khasnavis; Russell Plenkers

doi:10.3390/s22207999

Detecting Reconnaissance and Discovery Tactics from the MITRE ATT&CK Framework in Zeek Conn Logs Using Spark's Machine Learning in the Big Data Framework

Sensors (Basel). 2022 Oct 20;22(20):7999. doi: 10.3390/s22207999.

Authors

Sikha Bagui¹, Dustin Mink¹, Subhash Bagui², Tirthankar Ghosh¹, Tom McElroy¹, Esteban Paredes¹, Nithisha Khasnavis¹, Russell Plenkers¹

Affiliations

¹ Department of Computer Science, University of West Florida, Pensacola, FL 32514, USA.
² Department of Mathematics and Statistics, University of West Florida, Pensacola, FL 32514, USA.

Abstract

While computer networks and the massive amount of communication taking place on these networks grow, the amount of damage that can be done by network intrusions grows in tandem. The need is for an effective and scalable intrusion detection system (IDS) to address these potential damages that come with the growth of these networks. A great deal of contemporary research on near real-time IDS focuses on applying machine learning classifiers to labeled network intrusion datasets, but these datasets need be relevant pertaining to the currency of the network intrusions. This paper focuses on a newly created dataset, UWF-ZeekData22, that analyzes data from Zeek's Connection Logs collected using Security Onion 2 network security monitor and labelled using the MITRE ATT&CK framework TTPs. Due to the volume of data, Spark, in the big data framework, was used to run many of the well-known classifiers (naïve Bayes, random forest, decision tree, support vector classifier, gradient boosted trees, and logistic regression) to classify the reconnaissance and discovery tactics from this dataset. In addition to looking at the performance of these classifiers using Spark, scalability and response time were also analyzed.

Keywords: Apache Spark; MITRE ATT&CK® framework; Zeek Connection Logs; big data; intrusion detection systems; machine learning; network traffic analysis.

MeSH terms

Bayes Theorem
Big Data*
Logistic Models
Machine Learning*

Grants and funding

H98230-21-1-0170/National Security Agency