LLM Context Windows
A model's context window is the maximum number of tokens it can consider at once: system prompt, conversation history, documents, tool definitions and results, and the response it generates. Anything outside the window doesn't exist for the model. It has no memory beyond what you send in the current request.
Windows have grown from about 4,000 tokens in early chat models to hundreds of thousands and even millions today. That changes what's possible (whole codebases, long contracts, hours of transcripts in one prompt), but a bigger window isn't free and isn't automatically better. Cost, latency, and model attention all degrade as you fill it. Deciding what goes into the window is the core of context engineering.
TL;DR
- The context window counts input + output tokens together; there's usually a separate, smaller max output limit.
- LLMs are stateless: chat "memory" is the application re-sending history on every request.
- Longer contexts cost more (per-token pricing) and increase latency, especially time-to-first-token.
- Models may use long contexts unevenly: information buried mid-context can be missed, so test with your data.
- Prompt caching makes repeated long prefixes (documents, tool definitions, instructions) much cheaper and faster.
- For large or changing corpora, retrieval (RAG) usually beats stuffing everything into the window; for one document or a bounded set, long context is often simpler.
Quick Example
Keeping a chat within budget by caching a large stable prefix and trimming old turns:
In production, summarizing old turns usually beats simply dropping them. See compaction below.
Core Concepts
What Counts Against the Window
Everything the model processes in a request:
If input plus requested max_tokens exceeds the window, the API rejects the request, or the output gets truncated.
Statelessness and "Memory"
The model retains nothing between API calls. Chat apps create the illusion of memory by re-sending the conversation each time. Long-term memory features (user preferences, past sessions) work by storing information externally and injecting the relevant parts into the context. See AI agents.
Cost and Latency
- Cost scales with input tokens on every request. A 100k-token prompt re-sent across a 20-turn conversation is 2 million input tokens.
- Latency: prefill time grows with prompt length (attention is quadratic in length; see transformer architecture), so long prompts delay the first token.
- Prompt caching stores the processed state of a stable prefix. Cache reads cost a fraction of normal input tokens and cut latency substantially. Put stable content first (system prompt, tools, documents) and variable content last.
How Well Models Use Long Context
Benchmarks like "needle in a haystack" show modern models can find a single fact anywhere in huge contexts. Real tasks are harder: synthesizing information spread across a long document, following instructions buried deep in context, or reasoning over many similar passages. Quality often degrades as contexts grow, a phenomenon sometimes called context rot, and information in the middle can be underused ("lost in the middle"). Mitigations:
- Put the most important instructions near the start (system prompt) and restate the task at the end, after long documents.
- Structure long inputs with clear delimiters, such as XML tags or headings with document IDs.
- Ask the model to quote relevant passages before answering.
- Remove irrelevant material. Less context is often better context.
Managing Conversation History
Agentic systems accumulate tool results fast. Aggressively trimming stale tool outputs and compacting history keeps long-running agents inside their window and focused.
Long Context vs Retrieval
Many systems combine them: retrieve generously (dozens of chunks or whole documents) into a large window, instead of squeezing into a tiny one. See agentic RAG for retrieval driven by the model itself.
Best Practices
Budget the Window Explicitly
Allocate tokens per component (instructions, tools, retrieved context, history, output) and enforce the budgets with the target model's token counter. Surprises in production usually come from one component, like a giant tool result, crowding out the rest.
Order for Caching and Attention
Stable content first (system prompt, tool definitions, reference documents) with cache breakpoints, then dynamic content (retrieved chunks, history), then the user's question last.
Prefer Relevant Over Everything
Filter, deduplicate, and rank before inserting. Sending 20 highly relevant chunks usually beats 200 loosely related ones, in quality and in cost.
Test With Realistic Long Inputs
Evaluate your use case at the context sizes you'll actually send, with the answer placed in different positions. Headline window sizes don't tell you how well a model reasons at that length on your task. See LLM evaluation.
Common Mistakes
Unbounded Chat History
Enforce a budget and compact or trim older turns.
Forgetting Output and Reasoning Tokens
A prompt that uses 195k of a 200k window leaves only 5k tokens for the response, and reasoning models may spend many tokens thinking first. Reserve output headroom.
Dumping Raw Tool Output Into Context
A tool returning a 50,000-token JSON blob when the agent needed three fields wastes budget and distracts the model. Make tools return concise, relevant results, with pagination or filters.
FAQ
What happens when I exceed the context window?
Most APIs return an error when input plus max_tokens exceeds the limit. Some chat products silently drop or summarize older messages. If the output hits max_tokens, the response is truncated, which the stop reason shows.
Is a bigger context window always better?
No. It enables tasks that need lots of material at once, but each request costs more, runs slower, and can be less accurate if filled with marginally relevant text. The best systems put the right tokens in the window, not the most.
Does the model remember previous conversations?
Not by itself. Each request is independent. Continuity comes from your application re-sending history or injecting stored memories. Provider features like "projects" or "memory" are implemented the same way, outside the model.
How does prompt caching work with the context window?
Caching doesn't enlarge the window. Cached tokens still count toward it. It makes processing a repeated prefix cheaper and faster by reusing computed state for a time window (minutes by default, longer with extended options).
Related Topics
- Large Language Models — How LLMs work end to end
- Context Engineering — Deciding what goes into the window
- RAG — Retrieval as an alternative to stuffing context
- Tokenization — How window sizes are measured
- LLM Inference — KV caches and prompt caching mechanics
- AI Agents — Long-running context management