Synthetic and privacy-preserving traffic trace generation using generative AI models for training Network Intrusion Detection Systems
Синтетическая и конфиденциальная генерация сетевых трасс с помощью генеративных моделей ИИ для обучения систем обнаружения сетевых вторжений
2024-06-20
SCID: 54.1/h2ycyfkx
Discuss with AI
Conditional Variational Autoencoder (CVAE)PCAP and tabular synthetic datasetsnetwork intrusion detection systems (NIDS)privacy-preserving network trafficsynthetic traffic trace generation
Figures from the paper
Abstract (AI)
Network Intrusion Detection Systems (NIDS) are crucial tools for protecting networked devices from cyberattacks. Recent development in the field of Artificial Intelligence (AI) has provided tremendous advantages in implementing NIDSs able to monitor network traffic and block cyberattacks in real-time. In the literature, it is widely recognized that the effective training of a NIDS requires a large quantity of labeled traffic, representative of attacks. Nonetheless, the availability of public and abundant datasets remains remarkably restricted due to the cost of gathering and labeling real traffic traces and privacy concerns for sharing them. To tackle these challenges, in this paper we present a generative AI model capable of synthesizing anonymized traffic traces from real ones, thus dealing with privacy, abundance, and representativeness. The proposal is based on a Conditional Variational Autoencoder (CVAE) and a preprocessing procedure specifically designed for the generation of new traffic traces. To validate our solution, we conduct an extensive empirical study leveraging three recent and publicly-available datasets, containing benign and malicious traffic. The validation is carried out from both the perspectives of classification performance of a robust NIDS and the quality of synthetic data, in comparison to the utilization of real data. We compare our CVAE with two state-of-the-art AI-based traffic data generators and prove that, trained with traces emitted by our generative model, a NIDS has a limited F1-score loss compared to training on real data; competing models instead struggle or fail to generate traces that are as effective for NIDS training and as statistically similar to the original. We make the synthetic datasets available in both PCAP and tabular formats, to facilitate the reproducibility of our findings and encourage further exploration in the field of generative AI for networking.
Key Findings
1
A Conditional Variational Autoencoder (CVAE) plus a specialized preprocessing pipeline can synthesize anonymized traffic traces from real ones, addressing privacy and data scarcity for NIDS training.
2
Extensive empirical validation on three recent public datasets shows the CVAE-generated traces enable NIDS training with only a limited F1-score loss compared to training on real data.
3
The authors release the synthetic datasets in both PCAP and tabular formats to support reproducibility and further research into generative AI for networking.
4
Two state-of-the-art AI-based traffic generators were compared and found to struggle or fail to produce traces as effective for NIDS training or as statistically similar to original data as the proposed CVAE.
Research Object
Synthetic anonymized network traffic traces generated by a Conditional Variational Autoencoder (CVAE)
Research Subject
Effectiveness and privacy-preserving quality of the generated traffic for training Network Intrusion Detection Systems, evaluated via NIDS classification performance (F1-score loss) and statistical similarity to real traffic
Publication Details
Publication Date
2024-06-20
Journal
Publisher
ISSN
Cited by
32
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest