Skip to content
Tech Interview Prep home
Technical interview guide

Concurrency (GIL, Threading, Asyncio)

Why Python threads don't parallelize CPU work, and the two real ways around it: multiprocessing and asyncio.

Read
48 min
Practice MCQs
25
Interview QA
25
Edition
v2
Editorial status
Reviewed

Scope: Python 3.14, covering default GIL-enabled CPython and explicitly optional free-threaded builds.

Overview

Curated: · Written: · Reviewed:

Python concurrency in Python 3.14

Concurrency is about overlapping work; parallelism is simultaneous execution. Python offers threads, processes, interpreters, and cooperative asyncio tasks with different memory, scheduling, cancellation, and failure boundaries. A strong interview answer first classifies CPU versus waiting time and shared-state needs, then discusses the GIL—including optional free-threaded builds—without turning an implementation detail into a slogan.

CPython's GIL

In the standard GIL-enabled CPython build, one thread at a time executes Python bytecode within an interpreter.

The lock simplifies parts of the object model; it can be released around blocking I/O and by extension code performing long native operations.

Interview trap. Saying the GIL makes all Python code thread-safe confuses serialized bytecode execution with atomic multi-step invariants.

Engineering practice. Protect shared invariants with locks and choose concurrency from workload behavior rather than treating the GIL as a universal answer.

free-threaded Python 3.14

Python 3.14 supports optional free-threaded CPython builds where the GIL can be disabled, but this is not the default build assumption.

Free threading enables parallel thread execution while introducing extension compatibility checks, different overheads, and genuine simultaneous access to Python objects.

Interview trap. Claiming Python 3.14 removed the GIL everywhere misstates an optional build mode and ignores packages that may re-enable it.

Engineering practice. State build assumptions, test dependencies and synchronization under free threading, and benchmark before changing architecture.

threads for I/O

Threads help when work spends substantial time waiting for I/O because another thread can run while a blocking operation releases the GIL.

The operating system schedules threads, and many socket, file, and native-library waits release interpreter execution to peers.

Interview trap. Adding unbounded threads to slow I/O can exhaust memory, connections, rate limits, or downstream capacity.

Engineering practice. Use bounded pools, deadlines, cancellation signals, and measured concurrency aligned with dependency budgets.

threads for CPU work

CPU-bound pure-Python threads usually do not gain core-level parallelism on a GIL-enabled build.

Threads contend for interpreter execution and add scheduling overhead, although native code that releases the GIL can still run in parallel.

Interview trap. Benchmarking one extension-heavy workload and generalizing its threaded speedup to arbitrary Python loops is unsound.

Engineering practice. Use processes, interpreters, native vectorization, or a verified free-threaded build when CPU parallelism is required.

thread safety and atomicity

Thread-safe design protects complete application invariants, not individual operations that happen to look atomic in one interpreter.

A check-then-act sequence spans several operations and can interleave; free-threaded builds further invalidate accidental GIL-based assumptions.

Interview trap. if key not in d followed by d[key] = value is not an atomic uniqueness transaction.

Engineering practice. Use a lock, queue, immutable message, or storage primitive whose atomic contract matches the invariant.

thread locks

Lock serializes a critical section, while RLock permits the owning thread to acquire recursively and must be released equally.

Context-manager use pairs acquisition and release, but neither primitive defines a fair scheduling order.

Interview trap. Replacing Lock with RLock to hide accidental recursion can conceal a confused ownership design and increase coupling.

Engineering practice. Keep critical sections small, define lock ordering, avoid blocking external calls under lock, and test contention paths.

thread coordination

Events signal state, conditions combine a lock with predicate waiting, and semaphores bound concurrent access.

Condition wait releases and later reacquires its lock and should be used in a predicate loop because notifications do not guarantee the condition remains true.

Interview trap. Treating notify as transfer of a resource lets another thread consume the state before the woken waiter reacquires the lock.

Engineering practice. Encode the predicate explicitly, use wait_for where suitable, and select queues for ownership transfer.

thread pools

ThreadPoolExecutor reuses a bounded set of worker threads and represents submitted work with Future objects.

submit returns immediately; result blocks and propagates the worker exception, while executor shutdown controls acceptance and process exit behavior.

Interview trap. Calling result inside a worker on another future in the same saturated pool can deadlock.

Engineering practice. Avoid pool-internal dependency cycles, bound submission pressure, and collect every future's failure.

multiprocessing isolation

Processes have separate memory spaces, so communication requires serialization, shared-memory constructs, pipes, queues, or managers.

Isolation enables CPU parallelism and failure containment but increases startup, transfer, duplication, and coordination costs.

Interview trap. Passing a normal mutable object to a child and expecting parent-visible mutation confuses copied or serialized state with shared state.

Engineering practice. Send coarse immutable work units, minimize IPC, and establish one clear owner for mutable state.

process start methods

spawn, fork, and forkserver create different inherited state and safety properties, and Python 3.14 platform defaults must not be guessed.

Spawn imports the main module in a fresh interpreter; fork copies process state with thread-related hazards; forkserver delegates safe forks to a server.

Interview trap. Code that works only through inherited globals under fork fails under spawn and is fragile in libraries.

Engineering practice. Guard entry code with if name == 'main', pass explicit arguments, and test the chosen production start method.

process serialization

Process-pool callables, arguments, and results generally must be pickleable and importable in worker interpreters.

Lambdas, nested functions, open handles, and live locks often cannot cross the boundary; large payloads can dominate compute time.

Interview trap. Measuring only worker execution excludes serialization and transfer that can erase parallel speedup.

Engineering practice. Place worker functions at module scope, batch small work, and measure end-to-end throughput including startup and IPC.

process pools

ProcessPoolExecutor offers Future-based CPU parallelism but requires disciplined shutdown and worker-safe task design.

Workers are separate processes; abrupt termination, initializer failures, and a broken pool surface through future exceptions.

Interview trap. Submitting executor methods from within a process-pool task can deadlock because the worker participates in the pool being controlled.

Engineering practice. Keep tasks self-contained, propagate failures, use explicit max_workers and lifecycle bounds, and recover from a broken pool deliberately.

asyncio cooperative scheduling

asyncio tasks share an event-loop thread and switch cooperatively when execution reaches operations that suspend.

Creating a coroutine object does not schedule it; await drives an awaitable, and create_task schedules a coroutine as a Task.

Interview trap. Marking a function async does not make CPU-bound or blocking code nonblocking.

Engineering practice. Keep event-loop work short, use nonblocking libraries, and offload unavoidable blocking calls with controlled executors.

await points

An await is a potential interleaving point only when the awaited operation suspends, so reasoning should still treat shared state across awaits as vulnerable.

Other ready tasks can run during suspension and change state before the original coroutine resumes.

Interview trap. Single-threaded execution does not make read-await-write sequences atomic.

Engineering practice. Avoid holding inconsistent shared state across await; use asyncio locks, ownership queues, or local snapshots.

asyncio task references

create_task returns a Task that should be retained or managed structurally so lifetime and exceptions remain observable.

A Task runs independently of the immediate call stack, and fire-and-forget work can outlive request scope or fail without an awaiting owner.

Interview trap. Discarding the task reference is a reliable background-service pattern because the loop will always report and finish it.

Engineering practice. Use TaskGroup for scoped child tasks or maintain a supervised task collection with completion callbacks.

TaskGroup structured concurrency

TaskGroup binds child-task lifetime to an async-with scope and coordinates sibling cancellation when a child fails.

On exit it waits for children and raises grouped non-cancellation failures according to documented exception-group rules.

Interview trap. TaskGroup is merely gather with different spelling; failure propagation and task creation lifetime differ.

Engineering practice. Use it when sibling operations form one unit of work, and handle exception groups at the boundary that can make a recovery decision.

asyncio cancellation

Task cancellation is cooperative: cancel requests a CancelledError at the next suitable suspension point.

Cleanup runs through finally and context managers; code may temporarily catch cancellation but should normally propagate it after cleanup.

Interview trap. Catching BaseException or CancelledError and returning success can break TaskGroup, timeout, and shutdown protocols.

Engineering practice. Write cancellation-safe cleanup, avoid unbounded non-awaiting loops, and use uncancel only for rare fully understood suppression.

timeouts

asyncio timeout scopes cancel overdue work and translate the cancellation into TimeoutError at the documented boundary.

wait_for targets one awaitable, while the timeout context manager bounds a block; actual completion can include cancellation cleanup time.

Interview trap. A timeout proves the underlying external operation stopped and produced no side effect.

Engineering practice. Combine deadlines with idempotency, dependency-side timeouts, and reconciliation for operations whose outcome can become unknown.

asyncio shield

shield protects an inner awaitable from cancellation of its immediate waiter but does not make it immortal or hide cancellation from the caller.

The outer await still raises CancelledError while the inner task may continue; direct cancellation and shutdown can still affect it.

Interview trap. Using shield without retaining and supervising the task creates detached work with unclear ownership.

Engineering practice. Reserve shielding for critical bounded cleanup or commits, keep a strong reference, and await final outcome elsewhere.

gather behavior

gather runs awaitables concurrently and returns results in input order, which is separate from completion order.

Its exception and cancellation behavior depends on return_exceptions and whether gather itself or a child is cancelled.

Interview trap. Assuming the first raised exception automatically cancels and fully joins every sibling can leak work beyond the caller's expectation.

Engineering practice. Prefer TaskGroup for all-or-fail scopes and use gather when ordered aggregate results and its exact failure contract are intended.

asyncio queues and backpressure

An asyncio Queue coordinates task ownership, and maxsize makes producers await when capacity is exhausted.

put/get transfer items; task_done/join track completion separately from removal, and queue operations need external timeout scopes.

Interview trap. Using an unbounded queue moves overload into memory and latency while allowing producers to outrun consumers indefinitely.

Engineering practice. Choose capacity from service budgets, propagate cancellation, handle shutdown, and measure queue depth and oldest-item age.

asyncio synchronization

asyncio locks and semaphores coordinate tasks in one event loop and are not thread-synchronization primitives.

They suspend tasks rather than block the loop, but they do not protect data accessed concurrently by external threads or processes.

Interview trap. Sharing an asyncio.Lock with code running in to_thread protects cross-thread state.

Engineering practice. Keep ownership in one concurrency domain or add the corresponding thread/process-safe boundary.

offloading blocking work

asyncio.to_thread runs a blocking callable in a thread so the event loop can continue serving other tasks.

Context variables propagate, but cancellation of the awaiting task cannot forcibly stop arbitrary blocking thread code.

Interview trap. Using to_thread for long pure-Python CPU loops provides reliable parallel speedup on every CPython build.

Engineering practice. Use it for bounded blocking I/O, bound executor pressure, and design cooperative stop signals when abandoned work is costly.

async subprocesses

asyncio subprocess APIs integrate process streams with the event loop and require careful pipe consumption.

Waiting on a child with full stdout or stderr pipes can deadlock; communicate coordinates input, output, and process completion.

Interview trap. A cancelled waiter automatically terminates and reaps the operating-system process.

Engineering practice. Define terminate/kill escalation, drain bounded output, await process exit, and avoid shell interpolation of untrusted values.

choosing a concurrency model

Choose threads for compatible blocking I/O, processes or interpreters for isolated CPU parallelism, and asyncio for many cooperative I/O operations.

The decision also includes cancellation, state sharing, startup, serialization, library compatibility, observability, backpressure, and deployment limits.

Interview trap. Selecting the model from task count alone ignores whether tasks wait, compute, share mutable state, or need forceful isolation.

Engineering practice. Characterize the workload, state and failure boundary, prototype representative load, and use the simplest model that meets latency and throughput goals.

Worked example: eight 200 ms jobs on GIL-enabled CPython 3.14

Same machine, default (GIL) build. HTTP is a mock that sleeps in C and releases the GIL; fib is a pure-Python loop.

workloadThreadPoolExecutor(8)ProcessPoolExecutor(8)asyncio + to_thread
8 × 200 ms blocking HTTP~0.22 s~0.9 s (spawn)~0.21 s
8 × 200 ms pure-Python fib~1.6 s~0.28 s~1.6 s

Threads overlap waits. They do not multiply cores for bytecode. A timeout on await does not prove the HTTP POST never landed.