Not What You've Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection

Не то, на что вы подписывались: компрометация реальных приложений, интегрированных с большими языковыми моделями, с помощью косвенного внедрения промптов
Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz, Mario Fritz
2023-11-21

LLM-integrated applicationsarbitrary code executiondata theftindirect prompt injectioninformation ecosystem contamination
Large Language Models (LLMs) are increasingly being integrated into applications, with versatile functionalities that can be easily modulated via natural language prompts. So far, it was assumed that the user is directly prompting the LLM. But, what if it is not the user prompting? We show that LLM-Integrated Applications blur the line between data and instructions and reveal several new attack vectors, using Indirect Prompt Injection, that enable adversaries to remotely (i.e., without a direct interface) exploit LLM-integrated applications by strategically injecting prompts into data likely to be retrieved at inference time. We derive a comprehensive taxonomy from a computer security perspective to broadly investigate impacts and vulnerabilities, including data theft, worming, information ecosystem contamination, and other novel security risks. We then demonstrate the practical viability of our attacks against both real-world systems, such as Bing Chat and code-completion engines, and GPT-4 synthetic applications. We show how processing retrieved prompts can act as arbitrary code execution, manipulate the application's functionality, and control how and if other APIs are called. Despite the increasing reliance on LLMs, effective mitigations of these emerging threats are lacking. By raising awareness of these vulnerabilities, we aim to promote the safe and responsible deployment of these powerful models and the development of robust defenses that protect users from potential attacks.
1
Existing mitigations are described as ineffective or insufficient despite growing reliance on LLM-integrated applications.
2
Experiments demonstrate practical attacks against Bing Chat, code-completion engines, and GPT-4-based synthetic applications.
3
Indirect prompt injection enables remote attacks on LLM-integrated applications by embedding malicious instructions in data retrieved during inference.
4
Retrieved prompts can function as arbitrary code execution, alter application behavior, and control whether and how external APIs are invoked.
5
The study develops a security-oriented taxonomy covering impacts including data theft, self-propagating attacks, information ecosystem contamination, and other risks.

real-world LLM-integrated applications

vulnerabilities and security impacts caused by indirect prompt injection, including remote exploitation, data theft, functionality manipulation, arbitrary code execution, and API-call control

Publication Details
Publication Date
2023-11-21
Journal
Publisher
ISSN
Cited by
519
Access Type
Author Information
Authors
Kai Greshake
Sahar Abdelnabi
Shailesh Mishra
Christoph Endres
Thorsten Holz
Mario Fritz
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%