Privacy-Preserving Techniques in Generative AI and Large Language Models: A Narrative Review

Методы обеспечения конфиденциальности в генеративном искусственном интеллекте и больших языковых моделях: нарративный обзор
Vassilios S. Verykios, Aris Gkoulalas-Divanis, Georgios Feretzakis, Konstantinos Papaspyridis
2024-11-04

differential privacyfederated learninghomomorphic encryptionmembership inference attackssecure multi-party computation
Generative AI, including large language models (LLMs), has transformed the paradigm of data generation and creative content, but this progress raises critical privacy concerns, especially when models are trained on sensitive data. This review provides a comprehensive overview of privacy-preserving techniques aimed at safeguarding data privacy in generative AI, such as differential privacy (DP), federated learning (FL), homomorphic encryption (HE), and secure multi-party computation (SMPC). These techniques mitigate risks like model inversion, data leakage, and membership inference attacks, which are particularly relevant to LLMs. Additionally, the review explores emerging solutions, including privacy-enhancing technologies and post-quantum cryptography, as future directions for enhancing privacy in generative AI systems. Recognizing that achieving absolute privacy is mathematically impossible, the review emphasizes the necessity of aligning technical safeguards with legal and regulatory frameworks to ensure compliance with data protection laws. By discussing the ethical and legal implications of privacy risks in generative AI, the review underscores the need for a balanced approach that considers performance, scalability, and privacy preservation. The findings highlight the need for ongoing research and innovation to develop privacy-preserving techniques that keep pace with the scaling of generative AI, especially in large language models, while adhering to regulatory and ethical standards.
1
Because absolute privacy is mathematically impossible, technical safeguards must be aligned with legal and regulatory data-protection requirements.
2
Effective privacy preservation requires balancing privacy, model performance, scalability, and ethical considerations as generative AI systems expand.
3
Emerging privacy-enhancing technologies and post-quantum cryptography are identified as promising future directions for generative AI privacy.
4
The review surveys differential privacy, federated learning, homomorphic encryption, and secure multi-party computation for protecting sensitive data in generative AI.
5
These techniques target privacy threats particularly relevant to large language models, including model inversion, data leakage, and membership inference attacks.

Generative AI systems, particularly large language models (LLMs), trained on sensitive data

Privacy preservation, privacy risks, and the trade-offs among privacy, performance, scalability, and regulatory compliance

Publication Details
Publication Date
2024-11-04
Journal
Publisher
ISSN
Cited by
139
Access Type
Author Information
Authors
Vassilios S. Verykios
Aris Gkoulalas-Divanis
Georgios Feretzakis
Konstantinos Papaspyridis
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%