Prompt Chaining
When a single prompt tries to do everything (read a long document, extract facts, analyze them, draft a report, and format it), quality suffers: the model juggles too many instructions, errors are hard to locate, and one weak step drags down the whole output. Prompt chaining breaks the task into a sequence of focused LLM calls, where each step's output becomes the next step's input, often with code in between to validate, transform, or route.
Chaining is one of the most practical patterns in LLM application design. It sits between single prompts and fully autonomous agents: predictable like software, and flexible like LLMs. It's the "prompt chaining" workflow in agent design patterns, and the backbone of many production pipelines for extraction, analysis, and content generation.
TL;DR
- Split complex tasks into steps with a single clear goal each: extract → analyze → draft → review → format.
- Pass state between steps as structured data (structured outputs), not free-form prose.
- Add gates between steps: validation, schema checks, and business rules, to catch errors early.
- Routing: classify the input first, then send it to a specialized prompt or model.
- Map-reduce / parallelization: process chunks or aspects in parallel, then combine.
- Chains improve quality, debuggability, and control at the cost of more calls and latency. Use them when a single prompt measurably falls short.
Quick Example
A contract review pipeline:
Each step has one job, is testable on its own, and can use a different model sized to its difficulty.
Core Concepts
Why Chain?
- Focus: each prompt has fewer instructions, so the model follows them more reliably.
- Debuggability: when output is wrong, you can see which step failed.
- Control: code between steps enforces schemas, business rules, and routing deterministically.
- Mixed models: cheap, fast models for extraction and classification, strong models for reasoning and writing.
- Parallelism: independent sub-tasks run concurrently.
- Iteration: improve one step without re-tuning a monolithic prompt.
Common Chain Shapes
Passing State Between Steps
Use structured outputs (JSON matching schemas) for intermediate data. They're parseable, validatable, and unambiguous. Include only what the next step needs, which reduces tokens and distraction. Preserve provenance (source IDs and quotes) so later steps and users can trace claims back to inputs. See LLM hallucinations.
Gates and Validation
Between steps, code can check:
- Schema validity (types, required fields, enums), with a retry or fallback on failure.
- Business rules (totals add up, dates are in range, quoted text exists in the source).
- Confidence or classification thresholds, routing low-confidence cases to humans.
- Stop conditions: halt the chain early if a prerequisite fails, rather than generating confident nonsense downstream.
Chains vs Single Prompts vs Agents
Start simple: if one well-written prompt (perhaps with chain-of-thought) meets quality targets, don't chain. Chain when evaluation shows specific steps failing, or when the process is naturally staged.
Best Practices
Design Steps Around Clear Interfaces
Define each step's input and output schema like function signatures. Clear contracts make steps reusable, testable, and replaceable.
Evaluate Each Step and the Whole Chain
Build test sets per step (extraction accuracy, classification accuracy) and end to end. A chain can have great individual steps and still fail at the seams. See LLM evaluation.
Right-Size Models per Step
Use small, fast models for classification, extraction, and routing, and larger or reasoning models for synthesis and judgment. It often cuts cost substantially without hurting quality.
Trace Everything
Log each step's inputs, outputs, model, prompt version, latency, and tokens in a single trace. Observability tools for LLM pipelines make chains debuggable. See LLMOps.
Common Mistakes
Passing Prose Between Steps
Feeding one step's long free-text output into the next compounds verbosity and ambiguity. Use structured, minimal intermediate data.
No Validation Between Steps
Errors from early steps (a hallucinated field, a missed clause) propagate and get amplified by later steps. Validate at each boundary.
Chaining When One Prompt Would Do
Every added step costs latency and tokens, and adds failure points. Split only when it improves measured quality or control.
FAQ
What is prompt chaining?
Decomposing a complex task into a sequence of smaller LLM calls, where each call's output feeds the next, often with code in between for validation, transformation, or routing. It improves reliability and makes each part easier to test and improve.
How is prompt chaining different from an agent?
In a chain, your code defines the sequence of steps in advance. An agent lets the model decide which actions to take, and in what order, in a loop. Chains are more predictable and cheaper for known processes; agents handle open-ended tasks where the steps can't be predetermined.
Does prompt chaining increase cost?
Usually it increases the number of calls, but each call can be smaller and use cheaper models, and parallel steps limit latency growth. Many chains cost about the same as, or less than, one huge prompt using a top-tier model, while producing better results. Measure both quality and cost.
What frameworks support prompt chaining?
You can implement chains with plain code and an LLM SDK. Frameworks like LangChain and LangGraph, LlamaIndex, DSPy, Haystack, and workflow engines (durable execution systems) add abstractions, state management, retries, and tracing for complex pipelines.
Related Topics
- Prompt Engineering — Techniques overview
- Agent Design Patterns — Chaining among other workflows
- Structured Outputs — Passing state between steps
- Chain-of-Thought Prompting — Reasoning within a single step
- Prompt Optimization — Improving each step systematically
- Durable Execution — Reliable long-running pipelines