The 4 Failure Categories
Reviewed by Human · Updated August 11, 2026
Everyone believes the same myth: when AI gives you a bad answer, the model failed. A year of debugging prompts teaches you something less comfortable — the model almost always did exactly what your prompt permitted, including the parts you never wrote down.
That gap between what you meant and what you wrote has a shape. Four shapes, actually. Once you can name them, bad output stops feeling random and starts feeling diagnosable.
So if the model isn't broken, what actually happened when you got that terrible answer?
Real world
You get in a taxi and say "take me to Springfield." There are over thirty Springfields in the United States. The driver picks the nearest one, drives competently, and drops you off — 400 miles from the Springfield you meant. The driving wasn't the failure. The destination was underspecified, and the driver resolved it with a reasonable default that happened to be wrong.
AI parallel
A vague prompt works the same way. The model doesn't stop and interrogate you — it picks a reasonable default for every open question in your prompt and executes competently on the wrong interpretation. The output looks confident because the execution was fine. Only the resolved-in-your-absence decisions were off.
Why Does AI Give Bad Answers?
Nearly every bad output falls into one of four categories, and each one maps to a missing element in the anatomy of a prompt. **Ambiguity failure**: The prompt contained a word or instruction with multiple valid readings, and the model picked a reading that didn't match your intent. **Hallucination failure**: The prompt demanded specifics the model couldn't verify, so it generated plausible-sounding content to fill the gap — invented facts, sources, or numbers. **Scope creep failure**: The task was under-constrained, so the model expanded it — you asked for a tweak and got a rewrite, asked for an overview and got an essay. **Style/tone failure**: The content was right but the register was wrong — too formal, too casual, wrong audience — because the prompt never specified who the output was for. The categories matter because each one has a different fix. Adding more detail cures ambiguity but can worsen scope creep. Grounding fixes hallucination but does nothing for tone. Diagnosis comes first; edits come second.
Your turn
Time to run your first diagnosis. Read the scenario below and classify the failure using the four categories. Careful — there may be more than one.
Reflect
You likely found two categories: scope creep ('quick overview' set no length or depth boundary) and hallucination (the prompt gave no competitor list or source, so the model invented entries to complete the pattern). Real failures often stack like this — which is exactly why 'just retry it' so rarely works.
Diagnosis before edits
The instinct when output is bad is to immediately rewrite the prompt — add words, add rules, add a persona. Resist it. An edit made before diagnosis is a guess, and guesses that happen to work teach you nothing. Name the category first. The right edit is usually one sentence once you know which sentence is missing.
You've finished this module.
Mark it complete to earn your XP and keep your streak alive.
Progress saved locally · Sign up to earn XP