Chain-of-Thought Prompting
Language models generate one token at a time, and each token is computed from what came before. When a model jumps straight to an answer for a multi-step problem (a math word problem, a logic puzzle, a policy decision with several conditions), it has no room to work things out. Chain-of-thought (CoT) prompting asks the model to reason step by step before answering, and the intermediate reasoning measurably improves accuracy on arithmetic, logic, planning, and complex analysis.
The technique evolved from a prompt trick ("Let's think step by step") into a core model capability. Reasoning models and extended thinking modes are trained to produce long internal reasoning before answering, and they can be controlled with thinking budgets or effort settings. Knowing when reasoning helps, and when it just adds latency and cost, is part of modern prompt engineering.
TL;DR
- CoT = having the model write out intermediate reasoning before the final answer, which improves multi-step accuracy.
- Zero-shot CoT: "Think through this step by step". Few-shot CoT: examples that show worked reasoning.
- Use tags (
<thinking>,<answer>) to separate reasoning from the answer your app parses. - Self-consistency: sample several reasoning paths and take the majority answer, for harder problems.
- Reasoning models / extended thinking reason natively. Control them with thinking budgets or effort, and give high-level guidance rather than rigid step scripts.
- Reasoning costs tokens and latency; skip it for simple lookups, formatting, and classification that's already accurate.
Quick Example
Prompted chain-of-thought with a separated answer:
Using extended thinking via the API instead of prompt-level CoT:
Core Concepts
Why Reasoning Helps
Each generated token gives the model more computation to spend on the problem. Writing out intermediate results (sub-totals, which rule applies, what's known versus unknown) lets later tokens condition on them, much like a person using scratch paper. Gains are largest for multi-step math, logical deduction, applying several rules, planning, code debugging, and analysis requiring synthesis. They're smallest for simple factual recall or classification.
Prompting Techniques
Separating Reasoning From Output
Applications usually need only the final answer. Put reasoning and answer in distinct tags, parse the answer tag, and either discard the reasoning or keep it for logging and debugging. With structured outputs, you can include a reasoning field before the answer fields, so the model reasons first. Field order matters, because generation proceeds sequentially.
Reasoning Models and Extended Thinking
Reasoning models (and extended thinking modes in general-purpose models) are trained with reinforcement learning to produce long internal chains of thought, exploring, checking, and backtracking before answering. Practical differences:
- Control via budgets or effort (thinking token budgets,
reasoning_effortlevels) rather than prompt phrases. - Prefer high-level instructions: "Think carefully about edge cases" works better than prescribing every step, since these models often find better reasoning paths than a rigid script.
- Thinking tokens cost money and time. Bigger budgets help hard problems, and waste resources on easy ones.
- Interleaved thinking lets agents reason between tool calls.
See reasoning models.
When Not to Use CoT
- Simple lookups, formatting, translation, and extraction that are already accurate: reasoning adds latency and cost with no gain.
- Latency-critical interactions (voice, autocomplete): prefer direct answers or smaller models.
- When outputs must be minimal: keep reasoning internal (thinking modes), or strip it before display.
Evaluate: compare accuracy, latency, and cost with and without reasoning on your task. See LLM evaluation.
Faithfulness Caveat
Written reasoning isn't guaranteed to reflect how the model actually reached its answer: models can produce plausible-looking rationales that don't match their underlying computation. Treat CoT as a performance aid and a debugging signal, not as a verified explanation. Verify important conclusions independently.
Best Practices
Guide What to Think About
Tell the model which aspects matter (the rules to check, the constraints to verify, the edge cases to consider). Guided reasoning is more reliable than a generic "think step by step".
Ask the Model to Check Its Work
Adding "Before answering, verify the result against the requirements" or "check your calculation" catches many errors, especially for arithmetic and code. Tests and executable checks are even better.
Keep Final Answers Machine-Parseable
Put final answers in tags or structured fields, and validate them. Never parse answers out of free-form reasoning text.
Budget Reasoning by Difficulty
Route easy requests to fast, direct answers, and hard ones to reasoning models or larger thinking budgets. Classification-based routing keeps costs in check. See prompt chaining.
Common Mistakes
Answer Before Reasoning
Showing Raw Reasoning to End Users
Internal reasoning can be long, tentative, or include considerations not meant for users. Display polished answers, and keep reasoning for logs, or offer it as optional detail when appropriate.
Over-Prescribing Steps to Reasoning Models
Forcing a thinking model through a rigid 12-step template can reduce quality compared with giving the goal and key considerations. Start with high-level guidance, and add structure only where evaluation shows it helps.
FAQ
What is chain-of-thought prompting?
A technique that has a language model write out intermediate reasoning steps before its final answer. The extra reasoning tokens let the model break down multi-step problems, which improves accuracy on math, logic, planning, and complex analysis.
Do I still need chain-of-thought with reasoning models?
Not in the same form. Reasoning models think internally before responding, so prompt-level "think step by step" instructions are mostly unnecessary. Guide them with high-level considerations, and control depth with thinking budgets or effort settings.
Does chain-of-thought always improve results?
No. It helps most on problems requiring multiple steps. For simple tasks it adds latency and cost with little benefit, and occasionally it can lead models to overthink. Measure on your task.
What is self-consistency?
Generating several independent reasoning paths (with sampling) for the same question, then selecting the most common final answer. It improves accuracy on problems with a single correct answer, at the cost of multiple generations.
Related Topics
- Prompt Engineering — Techniques overview
- Reasoning Models — Models trained to think before answering
- Few-Shot Prompting — Worked examples with reasoning
- Prompt Chaining — Splitting reasoning across calls
- Structured Outputs — Reasoning fields before answers
- LLM Sampling — Sampling multiple reasoning paths