Free preview6 min

Why 'Think Step by Step' Works

Reviewed by Human · Updated July 22, 2026

In 2022, researchers at the University of Tokyo and Google added one sentence to their prompts and watched accuracy on a math benchmark jump from 17.7% to 78.7% (arxiv.org/abs/2205.11916). The sentence wasn't clever. It was: "Let's think step by step." No examples, no fine-tuning, no new model — just five words that changed *how* the model produced its answer. This module is about why that works, because once you understand the mechanism, you can aim it at problems the researchers never tested.

Live Demo — AI Terminal

Same question, two endings. First try asking it directly in your head — most people (and models) rush to '8'. Then watch what the step-by-step version does.

Your prompt

Why Does Chain-of-Thought Prompting Work?

**Chain-of-thought (CoT) prompting**: instructing an AI model to write out intermediate reasoning steps before giving its final answer. **Reasoning chain**: the visible sequence of steps the model produces on the way to that answer. Here's the mechanical part. A language model generates its response one token at a time, and everything it has already written becomes context for what it writes next. A direct question asks the model to make one giant leap — question straight to answer, all the intermediate computation compressed into a single prediction. CoT breaks that leap into small hops. Each hop is an easy prediction, and each written hop anchors the next one. The multi-step structure the problem *needed* now exists on the page instead of nowhere. That's also why CoT shines on math, logic, planning, and multi-constraint analysis, and does little for simple lookups. If there's no intermediate work to externalize, scratch paper is just paper. Zero-shot CoT is the lazy (complimentary) version: one trigger phrase, no demonstrations. It pairs beautifully with few-shot examples, where you show the model a worked reasoning chain to imitate — that pairing is exactly where the next module goes.

CoT doesn't make the model smarter. It changes the shape of the output so the model's existing ability can actually be applied step by step.

0

Kojima et al. (2022) measured zero-shot accuracy on the MultiArith benchmark. Direct prompting: 17.7%. Add 'Let's think step by step': 78.7%. On the harder GSM8K benchmark, the same phrase lifted accuracy from 10.4% to 40.7% (arxiv.org/abs/2205.11916). Five words, roughly four times the accuracy.

Your turn

Take a question you'd normally fire at an AI directly — a budgeting calculation, a scheduling puzzle, a 'which option is better' decision. Write it below twice: once as a direct question, once with a reasoning trigger. Then run both in your favorite AI and compare where the answers diverge.

Reflect

Look at the CoT output's middle steps. Could you have spotted an error there before trusting the final answer? That checkability is half the value.

end of module

You've finished this module.

Mark it complete to earn your XP and keep your streak alive.

Progress saved locally · Sign up to earn XP