Synthetic and privacy-preserving traffic trace generation using generative AI models for training Network Intrusion Detection Systems

Синтетическая и конфиденциальная генерация сетевых трасс с помощью генеративных моделей ИИ для обучения систем обнаружения сетевых вторжений
Giuseppe Aceto, Antonio Pescapè, Fabio Giampaolo, Francesco Piccialli, Ciro Guida, Stefano Izzo, Edoardo Prezioso
2024-06-20

Conditional Variational Autoencoder (CVAE)PCAP and tabular synthetic datasetsnetwork intrusion detection systems (NIDS)privacy-preserving network trafficsynthetic traffic trace generation
Network Intrusion Detection Systems (NIDS) are crucial tools for protecting networked devices from cyberattacks. Recent development in the field of Artificial Intelligence (AI) has provided tremendous advantages in implementing NIDSs able to monitor network traffic and block cyberattacks in real-time. In the literature, it is widely recognized that the effective training of a NIDS requires a large quantity of labeled traffic, representative of attacks. Nonetheless, the availability of public and abundant datasets remains remarkably restricted due to the cost of gathering and labeling real traffic traces and privacy concerns for sharing them. To tackle these challenges, in this paper we present a generative AI model capable of synthesizing anonymized traffic traces from real ones, thus dealing with privacy, abundance, and representativeness. The proposal is based on a Conditional Variational Autoencoder (CVAE) and a preprocessing procedure specifically designed for the generation of new traffic traces. To validate our solution, we conduct an extensive empirical study leveraging three recent and publicly-available datasets, containing benign and malicious traffic. The validation is carried out from both the perspectives of classification performance of a robust NIDS and the quality of synthetic data, in comparison to the utilization of real data. We compare our CVAE with two state-of-the-art AI-based traffic data generators and prove that, trained with traces emitted by our generative model, a NIDS has a limited F1-score loss compared to training on real data; competing models instead struggle or fail to generate traces that are as effective for NIDS training and as statistically similar to the original. We make the synthetic datasets available in both PCAP and tabular formats, to facilitate the reproducibility of our findings and encourage further exploration in the field of generative AI for networking.
1
A Conditional Variational Autoencoder (CVAE) plus a specialized preprocessing pipeline can synthesize anonymized traffic traces from real ones, addressing privacy and data scarcity for NIDS training.
2
Extensive empirical validation on three recent public datasets shows the CVAE-generated traces enable NIDS training with only a limited F1-score loss compared to training on real data.
3
The authors release the synthetic datasets in both PCAP and tabular formats to support reproducibility and further research into generative AI for networking.
4
Two state-of-the-art AI-based traffic generators were compared and found to struggle or fail to produce traces as effective for NIDS training or as statistically similar to original data as the proposed CVAE.

Synthetic anonymized network traffic traces generated by a Conditional Variational Autoencoder (CVAE)

Effectiveness and privacy-preserving quality of the generated traffic for training Network Intrusion Detection Systems, evaluated via NIDS classification performance (F1-score loss) and statistical similarity to real traffic

Publication Details
Publication Date
2024-06-20
Journal
Publisher
ISSN
Cited by
32
Access Type
Author Information
Authors
Giuseppe Aceto
Antonio Pescapè
Fabio Giampaolo
Francesco Piccialli
Ciro Guida
Stefano Izzo
Edoardo Prezioso
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat →
Make a presentation
100%