Prompt Chaining

When a single prompt tries to do everything (read a long document, extract facts, analyze them, draft a report, and format it), quality suffers: the model juggles too many instructions, errors are hard to locate, and one weak step drags down the whole output. Prompt chaining breaks the task into a sequence of focused LLM calls, where each step's output becomes the next step's input, often with code in between to validate, transform, or route.

Chaining is one of the most practical patterns in LLM application design. It sits between single prompts and fully autonomous agents: predictable like software, and flexible like LLMs. It's the "prompt chaining" workflow in agent design patterns, and the backbone of many production pipelines for extraction, analysis, and content generation.

TL;DR

Quick Example

A contract review pipeline:

Each step has one job, is testable on its own, and can use a different model sized to its difficulty.

Core Concepts

Why Chain?

Common Chain Shapes

Passing State Between Steps

Use structured outputs (JSON matching schemas) for intermediate data. They're parseable, validatable, and unambiguous. Include only what the next step needs, which reduces tokens and distraction. Preserve provenance (source IDs and quotes) so later steps and users can trace claims back to inputs. See LLM hallucinations.

Gates and Validation

Between steps, code can check:

Chains vs Single Prompts vs Agents

Start simple: if one well-written prompt (perhaps with chain-of-thought) meets quality targets, don't chain. Chain when evaluation shows specific steps failing, or when the process is naturally staged.

Best Practices

Design Steps Around Clear Interfaces

Define each step's input and output schema like function signatures. Clear contracts make steps reusable, testable, and replaceable.

Evaluate Each Step and the Whole Chain

Build test sets per step (extraction accuracy, classification accuracy) and end to end. A chain can have great individual steps and still fail at the seams. See LLM evaluation.

Right-Size Models per Step

Use small, fast models for classification, extraction, and routing, and larger or reasoning models for synthesis and judgment. It often cuts cost substantially without hurting quality.

Trace Everything

Log each step's inputs, outputs, model, prompt version, latency, and tokens in a single trace. Observability tools for LLM pipelines make chains debuggable. See LLMOps.

Common Mistakes

Passing Prose Between Steps

Feeding one step's long free-text output into the next compounds verbosity and ambiguity. Use structured, minimal intermediate data.

No Validation Between Steps

Errors from early steps (a hallucinated field, a missed clause) propagate and get amplified by later steps. Validate at each boundary.

Chaining When One Prompt Would Do

Every added step costs latency and tokens, and adds failure points. Split only when it improves measured quality or control.

FAQ

What is prompt chaining?

Decomposing a complex task into a sequence of smaller LLM calls, where each call's output feeds the next, often with code in between for validation, transformation, or routing. It improves reliability and makes each part easier to test and improve.

How is prompt chaining different from an agent?

In a chain, your code defines the sequence of steps in advance. An agent lets the model decide which actions to take, and in what order, in a loop. Chains are more predictable and cheaper for known processes; agents handle open-ended tasks where the steps can't be predetermined.

Does prompt chaining increase cost?

Usually it increases the number of calls, but each call can be smaller and use cheaper models, and parallel steps limit latency growth. Many chains cost about the same as, or less than, one huge prompt using a top-tier model, while producing better results. Measure both quality and cost.

What frameworks support prompt chaining?

You can implement chains with plain code and an LLM SDK. Frameworks like LangChain and LangGraph, LlamaIndex, DSPy, Haystack, and workflow engines (durable execution systems) add abstractions, state management, retries, and tracing for complex pipelines.

Related Topics

References