Design of big data anomaly detection model based on random forest algorithm

Authors

  • Song Shijun
  • Fan Min

DOI:

https://doi.org/10.59782/sidr.v1i1.40

Keywords:

big data clustering, feature extraction, principal component analysis, random forest classifier, decision tree, update weight

Abstract

Aiming at the problem that the anomaly detection process of big data is easily disturbed by edge data, resulting in poor accuracy of big data anomaly detection, a big data anomaly detection model based on random forest algorithm is proposed. Firstly, the improved -means algorithm is used to cluster the big data, and the principal component analysis method is used to extract the big data features; then, a big data anomaly detection model based on random forest classifier is constructed, the extracted features are input into the model, a decision tree is constructed, and the classification accuracy of the classifier is improved by dynamically updating the weight value of the decision tree; finally, the classification result is output to complete the anomaly detection of big data. The experimental results show that the detection time of the proposed model is about 25 s, the average accuracy of big data anomaly detection is, and the false alarm rate is .

How to Cite

Shijun, S., & Min, F. (2024). Design of big data anomaly detection model based on random forest algorithm. Scientific Insights and Discoveries Review, 1, 166–172. https://doi.org/10.59782/sidr.v1i1.40