Overview
Curated: · Written: · Reviewed:
Comprehensions and generators in Python 3.14
Comprehensions and generators both describe iteration, but they make different promises about evaluation time, storage, replay, failure timing, and resource lifetime. The mental model that holds it all together: a comprehension is an eager loop folded into an expression; a generator is a paused stack frame you resume one value at a time. Every interview question in this area is really probing which of those two contracts you're holding.
The four comprehension forms as one grammar
All four comprehensions share one grammar:
{expr for target in iterable [if condition] [for target2 in iterable2 ...]}
Only the brackets and the shape of expr change:
squares = [x*x for x in range(5)] # list: expr is the element
counts = {w: len(w) for w in words} # dict: expr is key: value
unique = {x.lower() for x in tokens} # set: expr is the element
lazy = (x*x for x in range(5)) # genexp: same grammar, no materialization
The first three build a container immediately. The fourth builds a generator object and defers everything except the leftmost iterable. That single difference drives most of the trade-offs below.
Clause order matters the same way in all four forms. Multiple for clauses nest left-to-right, exactly like the equivalent nested loops:
pairs = [(a, b) for a in 'ab' for b in range(2) if b]
# equivalent loops:
# out = []
# for a in 'ab':
# for b in range(2):
# if b:
# out.append((a, b))
print(pairs) # [('a', 1), ('b', 1)]
Read the comprehension left to right and you have the loop nesting. Read it right to left and you get a wrong Cartesian product — a classic interview slip.
One grammar rule trips people constantly: a trailing if filters (it can shrink the output), while an if/else before the for maps (one output per input). [x for x in xs if x] selects; [x if x else 0 for x in xs] transforms. [x for x in xs if x else 0] is a SyntaxError because a filter clause is not a conditional expression.
Scope and evaluation semantics inside a comprehension
A comprehension runs in its own implicit function scope. The loop target never leaks:
x = 'outer'
result = [x for x in range(3)]
print(x) # 'outer' — the comprehension's x is local to it
(Python 2 leaked the target; that's version folklore, not modern semantics.) Names the comprehension only reads resolve through normal enclosing-scope rules, which is how closures over outer variables work. The walrus operator := inside a comprehension binds in the containing scope — the one exception — which is worth stating precisely because it surprises people:
data = [1, 2, 3, 4]
evens = [y for x in data if (y := x) % 2 == 0]
print(evens) # [2, 4]
print(y) # 4 — leaked into the enclosing scope
Evaluation is eager for the container forms: every clause and every element expression runs at construction. A dict comprehension evaluates key and value per surviving item, and later duplicate keys overwrite earlier values while keeping the first key's insertion position. A set comprehension evaluates the element expression for every input before hashing decides membership — deduplication does not save you from an expensive f(x):
calls = 0
def expensive(x):
global calls; calls += 1
return x % 3
s = {expensive(x) for x in range(6)}
print(s, calls) # {0, 1, 2} 6 — all six evaluated, three kept
Lazy vs eager: the core trade-off
The generator expression is the same grammar with different timing:
lst = [n for n in range(3)] # list built now
gen = (n for n in range(3)) # nothing computed yet
The leftmost iterable is evaluated at creation — so gen = (x for x in undefined_name) raises NameError immediately, not lazily. Everything after it defers until the first next().
The decision criteria, stated as a candidate should give them:
- Materialize (list/dict/set) when you need replay, random access,
len(), a stable snapshot, or the size is known-bounded and modest. - Stream (generator) when the consumer may stop early, the input is large or unbounded, or the pipeline should stay memory-flat.
The failure mode of claiming "lazy" without checking the whole chain: a generator feeding sorted() or list() is not streaming — the materialization barrier just moved downstream. Trace the entire consumer chain before claiming bounded memory.
The iterator protocol underneath
Generators sit on the iterator protocol, and interviewers probe whether you can state it precisely:
- An iterable has
__iter__returning an iterator. - An iterator has
__next__producing values and raisesStopIterationwhen exhausted. iter(iterator)returns the iterator itself — iterators are one-shot; containers are replayable because eachiter()call makes a fresh cursor.- A
forloop is sugar: it callsiter(), thennext()repeatedly, catchingStopIterationinternally.
g = (n for n in range(2))
print(iter(g) is g) # True — a generator is its own iterator
print(next(g), next(g)) # 0 1
print(list(g)) # [] — exhausted, not replayed
That one-shot property is the most common production bug shape: passing one generator to two consumers leaves the second silently empty. Fix it with a factory function that builds a fresh generator per consumer, or itertools.tee — but tee buffers everything one branch has consumed that the other hasn't, so widely separated consumers cost memory comparable to materializing.
yield as a resumable state machine
Calling a function containing yield runs none of its body. It returns a generator object; the frame — locals plus instruction pointer — is created frozen:
def g():
print('start')
x = yield 1
print('got', x)
yield 2
gen = g() # nothing printed
print(next(gen)) # prints 'start', returns 1
print(gen.send('a')) # prints 'got a', returns 2
Trace it: g() builds the frozen frame. First next runs to the first yield, emits 1, suspends. send('a') resumes at the yield expression itself, which now evaluates to 'a', assigns it to x, and runs to the next yield. next is exactly send(None).
The rest of the generator protocol, in the order interviewers ask:
return exprin a generator raisesStopIterationwithexpras.value. Aforloop discards it;yield fromcaptures it as the value of the delegation expression.throw(exc)injects the exception at the suspended yield. If the generator handles it and yields again, that value becomesthrow's return value — throw does not guarantee termination.close()raisesGeneratorExitat the suspension point. Cleanup belongs infinally; yielding while handlingGeneratorExitraisesRuntimeError. Catching bareBaseExceptionand continuing to yield accidentally swallows closure.- PEP 479: a
StopIterationescaping a generator body becomesRuntimeError. Never raiseStopIterationmanually inside generator logic — usereturn. yield from subdelegates the whole protocol — values,send,throw,close, and the subgenerator's return value.yield subemits the generator object itself, a different thing entirely.
A practical consequence of deferred bodies: validation placed before the first yield doesn't run at call time. If callers need fail-fast behavior, wrap it:
def read_batches(src):
if not src: # runs immediately
raise ValueError('empty source')
def gen():
yield from src
return gen()
Streaming, early termination, and scale
Lazy pipelines skip work the consumer never requests. any(), all(), break, islice, and takewhile all stop pulling:
import itertools
found = any(x > 90 for x in expensive_scan()) # stops at first hit
first3 = itertools.islice(itertools.count(10), 3) # [10, 11, 12] from an infinite source
With infinite sources (count, cycle, repeat, open-ended generators), the bound must come from the consumer — list(itertools.count()) never returns. Put bounds close to the source and never "debug" an unbounded pipeline with list().
itertools composes without intermediate lists, but read each tool's storage contract: cycle caches the whole input iterable; tee buffers between branches; groupby groups consecutive equal keys only, and each group shares the underlying iterator, so advancing the outer loop invalidates the unconsumed rest of the previous group. Global grouping requires sorting by the same key first — which is itself a materialization barrier.
At scale, the honest answer about performance: generators are not automatically faster. Per-item generator overhead is real (frame resumption costs more than a tight list-comp loop), and a workload that immediately needs every value several times is better served by one materialization. Generators win on memory (O(pipeline state) vs O(n)) and on workloads that stop early. Profile with representative data and consumption depth rather than asserting either direction.
What interviewers probe, and what a weak answer sounds like
- "Is a list comprehension lazy?" Weak: "no, it's eager." Strong: eager, with the leftmost-iterable nuance for genexps and the memory consequence.
- "Why did my second loop over the generator print nothing?" Weak: "generators are one-time." Strong: iterators return themselves from
iter(), exhaustion is permanent, and here are the three fixes with their costs. - "What does
senddo?" Weak: "it passes input." Strong: it sets the value of the suspendedyieldexpression; the first call must besend(None)ornextbecause no yield is suspended yet. - "Is this pipeline streaming?" Weak: "yes, it uses generators." Strong: trace the chain to the final consumer and name any materialization barrier (
sorted,list,groupby-with-sort). - "When not to use a generator?" When you need replay,
len(), indexing, a stable snapshot against a mutating source, or when the consumer will touch every value repeatedly — materialize once.
Worked example: a generator is not a list
rows = (n for n in range(3))
first = list(rows) # [0, 1, 2]
second = list(rows) # []
| consumer | [n for n in range(3)] | (n for n in range(3)) |
|---|---|---|
first list(...) | [0, 1, 2] | [0, 1, 2] |
second list(...) | [0, 1, 2] | [] |
| memory before first consume | all three ints | one generator object (~200 bytes) |
when 0 is computed | construction | first next() |
That table is the interview: lazy is a consumption contract, not a synonym for cheap.
