Measuring the Impact of Accurate Feature Selection on the Performance of RBM in Comparison to State of the Art Machine Learning Algorithms

The amassed growth in the size of data, caused by the advancement of technologies and the use of internet of things to collect and transmit data, resulted in the creation of large volumes of data and an increasing variety of data types that need to be processed at very high speeds so that we can ext...

Full description

Bibliographic Details
Main Authors: Tamer Aldwairi, Dilina Perera, Mark A. Novotny
Format: Article
Language:English
Published: MDPI AG 2020-07-01
Series:Electronics
Subjects:
Online Access:https://www.mdpi.com/2079-9292/9/7/1167
_version_ 1797562089068298240
author Tamer Aldwairi
Dilina Perera
Mark A. Novotny
author_facet Tamer Aldwairi
Dilina Perera
Mark A. Novotny
author_sort Tamer Aldwairi
collection DOAJ
description The amassed growth in the size of data, caused by the advancement of technologies and the use of internet of things to collect and transmit data, resulted in the creation of large volumes of data and an increasing variety of data types that need to be processed at very high speeds so that we can extract meaningful information from these massive volumes of unstructured data. The process of mining this data is very challenging since a lot of the data suffers from the problem of high dimensionality. The quandary of high dimensionality represents a great challenge that can be controlled through the process of feature selection. Feature selection is a complex task with multiple layers of difficulty. To be able to grasp and realize the impediments associated with high dimensional data a more and in-depth understanding of feature selection is required. In this study, we examine the effect of appropriate feature selection during the classification process of anomaly network intrusion detection systems. We test its effect on the performance of Restricted Boltzmann Machines and compare its performance to conventional machine learning algorithms. We establish that when certain features that are representative of the model are to be selected the change in the accuracy was always less than 3% across all algorithms. This verifies that the accurate selection of the important features when building a model can have a significant impact on the accuracy level of the classifiers. We also confirmed in this study that the performance of the Restricted Boltzmann Machines can outperform or at least is comparable to other well-known machine learning algorithms. Extracting those important features can be very useful when trying to build a model with datasets with a lot of features.
first_indexed 2024-03-10T18:23:39Z
format Article
id doaj.art-5385a0e51eb448a49756fa75e8f9f2f0
institution Directory Open Access Journal
issn 2079-9292
language English
last_indexed 2024-03-10T18:23:39Z
publishDate 2020-07-01
publisher MDPI AG
record_format Article
series Electronics
spelling doaj.art-5385a0e51eb448a49756fa75e8f9f2f02023-11-20T07:11:30ZengMDPI AGElectronics2079-92922020-07-0197116710.3390/electronics9071167Measuring the Impact of Accurate Feature Selection on the Performance of RBM in Comparison to State of the Art Machine Learning AlgorithmsTamer Aldwairi0Dilina Perera1Mark A. Novotny2Distributed Analytics and Security Institute, High-Performance Computing Collaboratory, Mississippi State University, Mississippi State, MS 39762, USADistributed Analytics and Security Institute, High-Performance Computing Collaboratory, Mississippi State University, Mississippi State, MS 39762, USADepartment of Physics and Astronomy, Mississippi State University, Mississippi State, MS 39762, USAThe amassed growth in the size of data, caused by the advancement of technologies and the use of internet of things to collect and transmit data, resulted in the creation of large volumes of data and an increasing variety of data types that need to be processed at very high speeds so that we can extract meaningful information from these massive volumes of unstructured data. The process of mining this data is very challenging since a lot of the data suffers from the problem of high dimensionality. The quandary of high dimensionality represents a great challenge that can be controlled through the process of feature selection. Feature selection is a complex task with multiple layers of difficulty. To be able to grasp and realize the impediments associated with high dimensional data a more and in-depth understanding of feature selection is required. In this study, we examine the effect of appropriate feature selection during the classification process of anomaly network intrusion detection systems. We test its effect on the performance of Restricted Boltzmann Machines and compare its performance to conventional machine learning algorithms. We establish that when certain features that are representative of the model are to be selected the change in the accuracy was always less than 3% across all algorithms. This verifies that the accurate selection of the important features when building a model can have a significant impact on the accuracy level of the classifiers. We also confirmed in this study that the performance of the Restricted Boltzmann Machines can outperform or at least is comparable to other well-known machine learning algorithms. Extracting those important features can be very useful when trying to build a model with datasets with a lot of features.https://www.mdpi.com/2079-9292/9/7/1167anomaly network intrusion detection systemsmachine learningrestricted boltzmann machineISCX datasetNetFlow trafficcybersecurity
spellingShingle Tamer Aldwairi
Dilina Perera
Mark A. Novotny
Measuring the Impact of Accurate Feature Selection on the Performance of RBM in Comparison to State of the Art Machine Learning Algorithms
Electronics
anomaly network intrusion detection systems
machine learning
restricted boltzmann machine
ISCX dataset
NetFlow traffic
cybersecurity
title Measuring the Impact of Accurate Feature Selection on the Performance of RBM in Comparison to State of the Art Machine Learning Algorithms
title_full Measuring the Impact of Accurate Feature Selection on the Performance of RBM in Comparison to State of the Art Machine Learning Algorithms
title_fullStr Measuring the Impact of Accurate Feature Selection on the Performance of RBM in Comparison to State of the Art Machine Learning Algorithms
title_full_unstemmed Measuring the Impact of Accurate Feature Selection on the Performance of RBM in Comparison to State of the Art Machine Learning Algorithms
title_short Measuring the Impact of Accurate Feature Selection on the Performance of RBM in Comparison to State of the Art Machine Learning Algorithms
title_sort measuring the impact of accurate feature selection on the performance of rbm in comparison to state of the art machine learning algorithms
topic anomaly network intrusion detection systems
machine learning
restricted boltzmann machine
ISCX dataset
NetFlow traffic
cybersecurity
url https://www.mdpi.com/2079-9292/9/7/1167
work_keys_str_mv AT tameraldwairi measuringtheimpactofaccuratefeatureselectionontheperformanceofrbmincomparisontostateoftheartmachinelearningalgorithms
AT dilinaperera measuringtheimpactofaccuratefeatureselectionontheperformanceofrbmincomparisontostateoftheartmachinelearningalgorithms
AT markanovotny measuringtheimpactofaccuratefeatureselectionontheperformanceofrbmincomparisontostateoftheartmachinelearningalgorithms