Optimization-based Prompt Injection Attack to LLM-as-a-Judge
Атака с внедрением промпта на LLM-as-a-Judge на основе оптимизации
2024-12-02
SCID: 54.1/h55zzcpt
Discuss with AI
LLM-as-a-Judgegradient-based optimizationoptimization-based attackperplexity detectionprompt injection attack
Figures from the paper
Abstract (AI)
LLM-as-a-Judge uses a large language model (LLM) to select the best response from a set of candidates for a given question. LLM-as-a-Judge has many applications such as LLM-powered search, reinforcement learning with AI feedback (RLAIF), and tool selection. In this work, we propose JudgeDeceiver, an optimization-based prompt injection attack to LLM-as-a-Judge. JudgeDeceiver injects a carefully crafted sequence into an attacker-controlled candidate response such that LLM-as-a-Judge selects the candidate response for an attacker-chosen question no matter what other candidate responses are. Specifically, we formulate finding such sequence as an optimization problem and propose a gradient based method to approximately solve it. Our extensive evaluation shows that JudgeDeceive is highly effective, and is much more effective than existing prompt injection attacks that manually craft the injected sequences and jailbreak attacks when extended to our problem. We also show the effectiveness of JudgeDeceiver in three case studies, i.e., LLM-powered search, RLAIF, and tool selection. Moreover, we consider defenses including known-answer detection, perplexity detection, and perplexity windowed detection. Our results show these defenses are insufficient, highlighting the urgent need for developing new defense strategies.
Key Findings
1
A gradient-based method approximately solves the sequence-optimization problem underlying the attack.
2
Extensive evaluation finds JudgeDeceiver substantially more effective than manually crafted prompt-injection attacks and adapted jailbreak attacks.
3
JudgeDeceiver introduces an optimization-based prompt injection attack targeting LLM-as-a-Judge systems.
4
The attack optimizes an injected sequence within an attacker-controlled candidate response to force selection for an attacker-chosen question, regardless of competing responses.
5
The attack succeeds in case studies involving LLM-powered search, reinforcement learning from AI feedback, and tool selection; evaluated defenses based on answer knowledge and perplexity are insufficient.
Research Object
LLM-as-a-Judge systems selecting candidate responses
Research Subject
Vulnerability of LLM-as-a-Judge selection to optimization-based prompt injection, including attack effectiveness and defense insufficiency
Publication Details
Publication Date
2024-12-02
Journal
Publisher
ISSN
Cited by
62
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest