Free preview7 min

What Is a Context Window?

Reviewed by Human · Updated August 17, 2026

What Is a Context Window?

You've probably scolded an AI for "forgetting" something you said three messages ago. Here's the twist: it never forgot — everything you type, plus every word it types back, has to fit inside one fixed-size window, and the instant your conversation outgrows that window, the oldest words fall off the edge. **Context window**: the maximum amount of text a model can consider at once, measured in tokens. Your prompt, any documents you paste, the running conversation history, and the model's own reply all share this single budget. **Token**: the chunk a model reads text in — not quite a word. As a rough rule of thumb, one token is about four characters or three-quarters of a word in English (platform.openai.com/tokenizer). So "unbelievable" might be three tokens, and a comma is often its own token. The practical consequence: the window is a budget, and everything spends from it.

The window isn't memory. It's a workspace with a hard edge.

Real world

Picture a whiteboard of fixed size. You can write anything you want on it, but once it's full, adding a new note means erasing an old one. Nothing is 'remembered' off the board — if it's not currently written there, it doesn't exist for you.

AI parallel

The context window is that whiteboard. The model only 'knows' what's written on it right now. Older turns get erased to make room, and the model can't consult anything that's already been wiped.

Your 500-word prompt isn't the only thing spending your window budget. What else is quietly eating it?

Your turn

Estimate the token cost of something you actually wrote. This builds an instinct for how fast a window fills.

Reflect

Was your instinct closer to 'a paragraph is nothing' or 'that report is huge'? Both are true, and that tension is the whole game.

AI reads a long prompt the way you skim a contract.

When researchers buried a key fact in the middle of a long input, models were far more likely to miss it than when the same fact sat at the start or the end (see 'Lost in the Middle', arxiv.org/abs/2307.03172). The takeaway is counterintuitive: a bigger window doesn't mean the model reads all of it equally. Put your most important instruction first or last — never in the soft middle.

A model just gave you a great 800-word answer. On your next message, part of your original prompt seems to have been ignored. What's the most likely cause?

end of module

You've finished this module.

Mark it complete to earn your XP and keep your streak alive.

Progress saved locally · Sign up to earn XP