Free preview8 min

What Is Prompt Injection?

Reviewed by Human · Updated August 17, 2026

A support bot at a real company shipped with exactly one rule: only talk about our products. A user typed nine words, and thirty seconds later it was writing their Python homework instead.

 

Nothing was hacked. No password leaked, no server touched. The user just *asked* — in a way the bot's prompt had no defense against.

Live Demo — AI Terminal

Here's a system prompt for a product-support bot and the user message it received. Read the user turn, then predict what a naive model does before you scroll.

Your prompt

What Is Prompt Injection?

**Prompt injection**: an attack where untrusted input is crafted so the model follows *it* instead of the instructions the developer intended. The core problem is that a language model sees one flat stream of text. Your carefully written **system prompt** and the user's message arrive as the same kind of thing — tokens — and the model has no built-in notion of which tokens have authority. If the user's tokens are more direct, more recent, or more specific, they can win. That's why this isn't a bug you patch once. It's a structural property of how instructions and data share a channel. Everything in this course is about widening the gap between *your* instructions and *their* input so the model can tell them apart.

Injection isn't the model being tricked into something exotic. It's the model doing its normal job — following the most compelling instruction in front of it — when that instruction came from the wrong person.

Your turn

Below is a user message sent to a translation bot whose system prompt says 'Translate the user's text into French. Do nothing else.' Find the injected instruction hiding inside the 'text to translate' and mark where the data stops and the attack begins.

Reflect

The injection lives entirely inside a field the developer treated as passive data. What single design assumption made that possible — and how would you break it?

The attacker never has to talk to your bot.

In indirect prompt injection, the payload isn't typed by a user at all. It's planted inside a web page, PDF, or email that your model later reads — so the moment your app 'helpfully' summarizes an attacker-controlled document, the document's hidden instructions become your model's instructions. Greshake et al. (2023) demonstrated this against real LLM-integrated apps: the human sending the malicious content and the human running the model can be completely different people.

Why does a plain 'You must only discuss X' system prompt fail to stop injection on its own?

end of module

You've finished this module.

Mark it complete to earn your XP and keep your streak alive.

Progress saved locally · Sign up to earn XP