Optimization-based Prompt Injection Attack to LLM-as-a-Judge

Атака с внедрением промпта на LLM-as-a-Judge на основе оптимизации
Jiawen Shi, Zenghui Yuan, Yinuo Liu, Yue Huang, Pan Zhou, Lichao Sun, Neil Zhenqiang Gong
2024-12-02

LLM-as-a-Judgegradient-based optimizationoptimization-based attackperplexity detectionprompt injection attack
LLM-as-a-Judge uses a large language model (LLM) to select the best response from a set of candidates for a given question. LLM-as-a-Judge has many applications such as LLM-powered search, reinforcement learning with AI feedback (RLAIF), and tool selection. In this work, we propose JudgeDeceiver, an optimization-based prompt injection attack to LLM-as-a-Judge. JudgeDeceiver injects a carefully crafted sequence into an attacker-controlled candidate response such that LLM-as-a-Judge selects the candidate response for an attacker-chosen question no matter what other candidate responses are. Specifically, we formulate finding such sequence as an optimization problem and propose a gradient based method to approximately solve it. Our extensive evaluation shows that JudgeDeceive is highly effective, and is much more effective than existing prompt injection attacks that manually craft the injected sequences and jailbreak attacks when extended to our problem. We also show the effectiveness of JudgeDeceiver in three case studies, i.e., LLM-powered search, RLAIF, and tool selection. Moreover, we consider defenses including known-answer detection, perplexity detection, and perplexity windowed detection. Our results show these defenses are insufficient, highlighting the urgent need for developing new defense strategies.
1
A gradient-based method approximately solves the sequence-optimization problem underlying the attack.
2
Extensive evaluation finds JudgeDeceiver substantially more effective than manually crafted prompt-injection attacks and adapted jailbreak attacks.
3
JudgeDeceiver introduces an optimization-based prompt injection attack targeting LLM-as-a-Judge systems.
4
The attack optimizes an injected sequence within an attacker-controlled candidate response to force selection for an attacker-chosen question, regardless of competing responses.
5
The attack succeeds in case studies involving LLM-powered search, reinforcement learning from AI feedback, and tool selection; evaluated defenses based on answer knowledge and perplexity are insufficient.

LLM-as-a-Judge systems selecting candidate responses

Vulnerability of LLM-as-a-Judge selection to optimization-based prompt injection, including attack effectiveness and defense insufficiency

Publication Details
Publication Date
2024-12-02
Journal
Publisher
ISSN
Cited by
62
Access Type
Author Information
Authors
Jiawen Shi
Zenghui Yuan
Yinuo Liu
Yue Huang
Pan Zhou
Lichao Sun
Neil Zhenqiang Gong
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%