Glossary Entry

Prompt Injection

An attack in which text that an AI system reads, such as a web page, email, or file, contains instructions that the model then follows as if they came from its user.

LLMs Deployment

Also called: prompt injections, prompt-injection

Seed source: Simon Willison

Language models cannot reliably tell instructions they should follow from text they are only meant to process, because everything arrives as one sequence of tokens. An agent asked to summarize a page can be told by that page to fetch private data and send it somewhere else.

No filter fully prevents this, so practical defences limit what a successful injection can do: run agents in sandboxes, and avoid giving one agent access to private data, exposure to untrusted content, and the ability to communicate externally all at once.