Skip to content
Tech Interview Prep home
Technical interview guide

Comprehensions & Generators

Concise, often faster ways to build sequences — and the lazy-evaluation alternative that avoids materializing them at all.

Read
47 min
Practice MCQs
25
Interview QA
25
Edition
v2
Editorial status
Reviewed

Scope: Python 3.14, including current generator, StopIteration, and asynchronous-generator semantics.

Overview

Curated: · Written: · Reviewed:

Comprehensions and generators in Python 3.14

Comprehensions and generators both describe iteration, but they make different promises about evaluation time, storage, replay, failure timing, and resource lifetime. The mental model that holds it all together: a comprehension is an eager loop folded into an expression; a generator is a paused stack frame you resume one value at a time. Every interview question in this area is really probing which of those two contracts you're holding.

The four comprehension forms as one grammar

All four comprehensions share one grammar:

{expr for target in iterable [if condition] [for target2 in iterable2 ...]}

Only the brackets and the shape of expr change:

squares  = [x*x for x in range(5)]        # list:   expr is the element
counts   = {w: len(w) for w in words}      # dict:   expr is key: value
unique   = {x.lower() for x in tokens}     # set:    expr is the element
lazy     = (x*x for x in range(5))         # genexp: same grammar, no materialization

The first three build a container immediately. The fourth builds a generator object and defers everything except the leftmost iterable. That single difference drives most of the trade-offs below.

Clause order matters the same way in all four forms. Multiple for clauses nest left-to-right, exactly like the equivalent nested loops:

pairs = [(a, b) for a in 'ab' for b in range(2) if b]
# equivalent loops:
# out = []
# for a in 'ab':
#     for b in range(2):
#         if b:
#             out.append((a, b))
print(pairs)  # [('a', 1), ('b', 1)]

Read the comprehension left to right and you have the loop nesting. Read it right to left and you get a wrong Cartesian product — a classic interview slip.

One grammar rule trips people constantly: a trailing if filters (it can shrink the output), while an if/else before the for maps (one output per input). [x for x in xs if x] selects; [x if x else 0 for x in xs] transforms. [x for x in xs if x else 0] is a SyntaxError because a filter clause is not a conditional expression.

Scope and evaluation semantics inside a comprehension

A comprehension runs in its own implicit function scope. The loop target never leaks:

x = 'outer'
result = [x for x in range(3)]
print(x)  # 'outer' — the comprehension's x is local to it

(Python 2 leaked the target; that's version folklore, not modern semantics.) Names the comprehension only reads resolve through normal enclosing-scope rules, which is how closures over outer variables work. The walrus operator := inside a comprehension binds in the containing scope — the one exception — which is worth stating precisely because it surprises people:

data = [1, 2, 3, 4]
evens = [y for x in data if (y := x) % 2 == 0]
print(evens)  # [2, 4]
print(y)      # 4 — leaked into the enclosing scope

Evaluation is eager for the container forms: every clause and every element expression runs at construction. A dict comprehension evaluates key and value per surviving item, and later duplicate keys overwrite earlier values while keeping the first key's insertion position. A set comprehension evaluates the element expression for every input before hashing decides membership — deduplication does not save you from an expensive f(x):

calls = 0
def expensive(x):
    global calls; calls += 1
    return x % 3
s = {expensive(x) for x in range(6)}
print(s, calls)  # {0, 1, 2} 6 — all six evaluated, three kept

Lazy vs eager: the core trade-off

The generator expression is the same grammar with different timing:

lst = [n for n in range(3)]   # list built now
gen = (n for n in range(3))   # nothing computed yet

The leftmost iterable is evaluated at creation — so gen = (x for x in undefined_name) raises NameError immediately, not lazily. Everything after it defers until the first next().

The decision criteria, stated as a candidate should give them:

  • Materialize (list/dict/set) when you need replay, random access, len(), a stable snapshot, or the size is known-bounded and modest.
  • Stream (generator) when the consumer may stop early, the input is large or unbounded, or the pipeline should stay memory-flat.

The failure mode of claiming "lazy" without checking the whole chain: a generator feeding sorted() or list() is not streaming — the materialization barrier just moved downstream. Trace the entire consumer chain before claiming bounded memory.

The iterator protocol underneath

Generators sit on the iterator protocol, and interviewers probe whether you can state it precisely:

  • An iterable has __iter__ returning an iterator.
  • An iterator has __next__ producing values and raises StopIteration when exhausted.
  • iter(iterator) returns the iterator itself — iterators are one-shot; containers are replayable because each iter() call makes a fresh cursor.
  • A for loop is sugar: it calls iter(), then next() repeatedly, catching StopIteration internally.
g = (n for n in range(2))
print(iter(g) is g)  # True — a generator is its own iterator
print(next(g), next(g))  # 0 1
print(list(g))       # [] — exhausted, not replayed

That one-shot property is the most common production bug shape: passing one generator to two consumers leaves the second silently empty. Fix it with a factory function that builds a fresh generator per consumer, or itertools.tee — but tee buffers everything one branch has consumed that the other hasn't, so widely separated consumers cost memory comparable to materializing.

yield as a resumable state machine

Calling a function containing yield runs none of its body. It returns a generator object; the frame — locals plus instruction pointer — is created frozen:

def g():
    print('start')
    x = yield 1
    print('got', x)
    yield 2

gen = g()            # nothing printed
print(next(gen))     # prints 'start', returns 1
print(gen.send('a')) # prints 'got a', returns 2

Trace it: g() builds the frozen frame. First next runs to the first yield, emits 1, suspends. send('a') resumes at the yield expression itself, which now evaluates to 'a', assigns it to x, and runs to the next yield. next is exactly send(None).

The rest of the generator protocol, in the order interviewers ask:

  • return expr in a generator raises StopIteration with expr as .value. A for loop discards it; yield from captures it as the value of the delegation expression.
  • throw(exc) injects the exception at the suspended yield. If the generator handles it and yields again, that value becomes throw's return value — throw does not guarantee termination.
  • close() raises GeneratorExit at the suspension point. Cleanup belongs in finally; yielding while handling GeneratorExit raises RuntimeError. Catching bare BaseException and continuing to yield accidentally swallows closure.
  • PEP 479: a StopIteration escaping a generator body becomes RuntimeError. Never raise StopIteration manually inside generator logic — use return.
  • yield from sub delegates the whole protocol — values, send, throw, close, and the subgenerator's return value. yield sub emits the generator object itself, a different thing entirely.

A practical consequence of deferred bodies: validation placed before the first yield doesn't run at call time. If callers need fail-fast behavior, wrap it:

def read_batches(src):
    if not src:                 # runs immediately
        raise ValueError('empty source')
    def gen():
        yield from src
    return gen()

Streaming, early termination, and scale

Lazy pipelines skip work the consumer never requests. any(), all(), break, islice, and takewhile all stop pulling:

import itertools
found = any(x > 90 for x in expensive_scan())  # stops at first hit
first3 = itertools.islice(itertools.count(10), 3)  # [10, 11, 12] from an infinite source

With infinite sources (count, cycle, repeat, open-ended generators), the bound must come from the consumer — list(itertools.count()) never returns. Put bounds close to the source and never "debug" an unbounded pipeline with list().

itertools composes without intermediate lists, but read each tool's storage contract: cycle caches the whole input iterable; tee buffers between branches; groupby groups consecutive equal keys only, and each group shares the underlying iterator, so advancing the outer loop invalidates the unconsumed rest of the previous group. Global grouping requires sorting by the same key first — which is itself a materialization barrier.

At scale, the honest answer about performance: generators are not automatically faster. Per-item generator overhead is real (frame resumption costs more than a tight list-comp loop), and a workload that immediately needs every value several times is better served by one materialization. Generators win on memory (O(pipeline state) vs O(n)) and on workloads that stop early. Profile with representative data and consumption depth rather than asserting either direction.

What interviewers probe, and what a weak answer sounds like

  • "Is a list comprehension lazy?" Weak: "no, it's eager." Strong: eager, with the leftmost-iterable nuance for genexps and the memory consequence.
  • "Why did my second loop over the generator print nothing?" Weak: "generators are one-time." Strong: iterators return themselves from iter(), exhaustion is permanent, and here are the three fixes with their costs.
  • "What does send do?" Weak: "it passes input." Strong: it sets the value of the suspended yield expression; the first call must be send(None) or next because no yield is suspended yet.
  • "Is this pipeline streaming?" Weak: "yes, it uses generators." Strong: trace the chain to the final consumer and name any materialization barrier (sorted, list, groupby-with-sort).
  • "When not to use a generator?" When you need replay, len(), indexing, a stable snapshot against a mutating source, or when the consumer will touch every value repeatedly — materialize once.

Worked example: a generator is not a list

rows = (n for n in range(3))
first = list(rows)   # [0, 1, 2]
second = list(rows)  # []
consumer[n for n in range(3)](n for n in range(3))
first list(...)[0, 1, 2][0, 1, 2]
second list(...)[0, 1, 2][]
memory before first consumeall three intsone generator object (~200 bytes)
when 0 is computedconstructionfirst next()

That table is the interview: lazy is a consumption contract, not a synonym for cheap.