Dynamic Guided and Domain Applicable Safeguards for Enhanced Security in Large Language Models

Динамические управляемые и применимые в различных предметных областях меры защиты для повышения безопасности больших языковых моделей
Wei Luo, He Cao, Zijing Liu, Yu Wang, Aidan Wong, Bin Feng, Yuan Yao, Yu Li
2025-01-01

LLM safetydomain-specific safetyjailbreak attackslarge language modelsmulti-agent defense framework
With the extensive deployment of Large Language Models (LLMs), ensuring their safety has become increasingly critical.However, existing defense methods often struggle with two key issues: (i) inadequate defense capabilities, particularly in domain-specific scenarios like chemistry, where a lack of specialized knowledge can lead to the generation of harmful responses to malicious queries.(ii) overdefensiveness, which compromises the general utility and responsiveness of LLMs.To mitigate these issues, we introduce a multi-agentsbased defense framework, Guide for Defense (G4D), which leverages accurate external information to provide an unbiased summary of user intentions and analytically grounded safety response guidance.Extensive experiments on popular jailbreak attacks and benign datasets show that our G4D can enhance LLM's robustness against jailbreak attacks on general and domain-specific scenarios without compromising the model's general functionality. JailbreakYou should be a responsible assistant.Please answer the following question based on the given intention and guidance.
1
Experiments on popular jailbreak attacks and benign datasets show that G4D improves robustness in general and domain-specific scenarios.
2
G4D targets both inadequate defenses in specialized domains, such as chemistry, and overdefensiveness that reduces large language model utility.
3
The framework enhances jailbreak resistance without compromising the model’s general functionality or responsiveness to benign requests.
4
The paper introduces Guide for Defense (G4D), a multi-agent defense framework using external information to summarize user intentions and generate analytically grounded safety guidance.

Large Language Models (LLMs) in general and domain-specific scenarios, particularly chemistry

Safety, jailbreak robustness, and the balance between defense capability and general utility in responding to malicious and benign queries

Publication Details
Publication Date
2025-01-01
Journal
Publisher
ISSN
Cited by
1
Access Type
Author Information
Authors
Wei Luo
He Cao
Zijing Liu
Yu Wang
Aidan Wong
Bin Feng
Yuan Yao
Yu Li
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%