Python Generators & Iterators

Every for loop in Python runs on the iterator protocol, and generators are the easiest way to write iterators yourself. A generator function uses yield to produce values one at a time, pausing between them and keeping its local state. Nothing is computed until someone asks for the next value.

That laziness is the point. You can process a 50 GB log file, an unbounded stream of events, or a paginated API in constant memory, building pipelines of small transformations that each handle one item at a time. The same machinery underlies context managers and, historically, async/await.

TL;DR

Quick Example

Stream a huge CSV file, filter it, and aggregate, holding one row in memory at a time:

Each stage pulls one item from the stage before it. islice stops the whole pipeline after five items, so the rest of the file is never read.

Core Concepts

The Iterator Protocol

A for loop does exactly this and catches StopIteration to end. Lists, dicts, files, ranges, and strings are all iterables. Files and generators are also iterators, meaning they're consumed as you go.

Generator Functions

Each yield suspends the function with its local variables intact. The next next() resumes right after the yield. When the function returns, the generator raises StopIteration.

Generator Expressions

Use a list comprehension when you need the list (indexing, multiple passes, len). Use a generator expression when you'll consume items once.

yield from and Delegation

yield from iterable yields every item from another iterable, and properly forwards send() and throw() to sub-generators. It's how you compose recursive generators cleanly:

Sending Values and Cleanup

Generators are also coroutines in the classic sense: gen.send(value) resumes the generator and makes the current yield expression evaluate to value, and gen.close() raises GeneratorExit inside it so finally blocks run. Modern async code uses async def instead (see asyncio), but send still shows up in some pipeline and state-machine code.

Context Managers From Generators

@contextlib.contextmanager turns a generator with one yield into a with-statement context manager. Code before the yield is setup, and code after it (in a finally) is teardown:

itertools Essentials

Best Practices

Stream Large Data

Read files line by line, page through APIs, and batch database writes with generators rather than loading everything into lists. Memory use then stays flat no matter how large the input grows, which matters for data engineering jobs and workers in memory-limited containers.

Keep Stages Small and Composable

Write one generator per transformation (parse → filter → enrich → batch), then chain them. Each stage is easy to test with a short list as input.

Make Resource Cleanup Deterministic

A generator that opens a file inside with keeps it open until the generator is exhausted or garbage collected. If consumers may stop early, close the generator explicitly (gen.close()) or use contextlib.closing, so cleanup doesn't depend on garbage collection timing.

Return Iterables When Callers Need Multiple Passes

A function returning a generator can only be consumed once. If callers will iterate twice, return a list, or a class whose __iter__ creates a fresh generator each time.

Common Mistakes

Iterating a Generator Twice

Using groupby on Unsorted Data

itertools.groupby groups adjacent equal keys. On unsorted data you get many small groups for the same key. Sort by the same key first, or use a defaultdict(list) when you need true grouping.

Raising StopIteration Inside a Generator

Since Python 3.7, a StopIteration raised inside a generator body becomes a RuntimeError (PEP 479). To end a generator early, just return.

FAQ

What's the difference between an iterable and an iterator?

An iterable can produce an iterator (it has __iter__), and can usually be iterated many times, like a list. An iterator produces values one at a time with __next__ and is consumed as it goes. Every iterator is iterable (its __iter__ returns itself), but not every iterable is an iterator.

When should I use a generator instead of a list?

When the data is large or unbounded, when you only need a single pass, or when you might stop early. Lists are better when you need random access, the length, or multiple passes, or when the data is small enough that laziness buys nothing.

Are generators slower than lists?

Per item, a generator has slightly more overhead than iterating a prebuilt list. Overall they're often faster in practice because they avoid allocating huge lists and can stop early. The memory savings are usually the deciding factor.

How do async generators differ?

An async def function containing yield is an async generator, consumed with async for. It can await between yields, which makes it ideal for streaming paginated API results or database cursors in async code.

Related Topics

References