Overview
Curated: · Written: · Reviewed:
Orchestration coordinates outcomes; it does not make tasks correct
A workflow orchestrator records when and where a unit of work should run, which dependencies must be satisfied, what attempt is active, and how operators observe or recover it. The task code still owns business correctness, durable writes, identity, validation and external side effects. A green DAG or flow means the configured terminal-state rule succeeded; it does not prove source completeness, output correctness, or consumer freshness.
This distinction is the single most common interview probe in orchestration questions. Interviewers ask "what does Airflow (or Prefect, or Dagster) guarantee you?" and the weak answer lists features — retries, scheduling, UI, alerting. The strong answer separates what the orchestrator owns — workflow definition, dependency resolution, state tracking, triggering — from what it never owns: the compute the tasks run on, the storage they write to, and the correctness of the data itself. Expect the follow-up: "so what does a green DAG actually prove?" If you can answer that in one sentence — the configured terminal-state rule held, nothing about the data — you've set up the rest of the conversation.
A related probe is vocabulary: scheduler vs. workflow engine vs. orchestrator. The scheduler decides when work is released (cron, events, sensors). The workflow engine resolves the graph and tracks run state. The orchestrator is the product wrapping both with observability, recovery, and operator tooling. In Airflow these are all one system; in a Kafka-and-Jobs design they may be three. Interviewers use this to check whether you've thought about the design or just used the tool.
Model the workflow around data products and explicit contracts. A directed acyclic graph expresses ordering and parallelism, but an edge should represent a real dependency such as required data, schema, quality approval, or publication—not merely a desire to draw boxes left-to-right. Keep tasks cohesive and restartable with stable input/output boundaries. Large payloads belong in durable object or database storage; orchestration metadata should carry references, versions and small control values. Dynamic mapping can fan out partitions, but cardinality, quotas, partial success and aggregation must be bounded.
Task granularity is a classic design question: one task per table, per partition, or per pipeline stage? The weak answer picks one universally. The strong answer ties granularity to retry cost and blast radius — a task that does extract, transform, and load in one body can't retry just the failed load, and a task per row turns the orchestrator into a slow database. Also know how your tool expresses one-time-pass vs. per-partition execution: Airflow's dynamic task mapping and Prefect's .map fan out per-partition attempts with per-partition success, while a plain task is all-or-nothing.
Time has several meanings. A scheduled run has a logical date or data interval that may differ from wall-clock start time. Tasks must parameterize reads and writes by the interval or partition being processed, not by current time, or retry and backfill will corrupt results. Event-triggered workflows need idempotent event identity and deduplication. Schedules, sensors and dataset/asset events are different ways to release work; none proves upstream data is complete without a producer contract. Polling sensors require timeouts, backoff and resource-aware waiting.
Interviewers love the question "a daily job runs at 02:00 — what data does it process?" The weak answer is "yesterday's data." The strong answer names the data interval explicitly (in Airflow, a run with logical date 2026-09-10 covers the interval ending 2026-09-10, and starts after it closes), then handles late-arriving data: does the interval close on a watermark, a producer contract, or hope? Schedule time and event time diverging is the root cause of most "the DAG was green but the table was short" incidents, and saying so is a strong signal.
Retries are for bounded transient failures, with exponential backoff, jitter and a classification that excludes deterministic invalid input or unsafe side effects. Task idempotency means the same logical run and input version converges on one outcome. Use stable operation keys, staging plus atomic publication, transactional or merge-capable sinks, and commit input progress only after the effect. Compensation is not rollback: it is a new, observable business action whose own failure and retry semantics must be designed.
The idempotency probe is nearly guaranteed: "your task ran twice — what happens?" Weak answers say "Airflow retries are safe" (they aren't automatically) or "we use exactly-once" (almost never end-to-end). The strong answer names the mechanism: partitioned overwrite (write the whole 2026-09-10 partition, atomically replace it), merge on a stable key, or staging-plus-publish. Determinism is the sibling property — same input version, same output — and it's what makes backfills trustworthy.
Backfill creates or reruns historical intervals. Define the exact range, code/schema version, source snapshot, reprocessing behavior, run order, concurrency, resource pools, interaction with live schedules and downstream restatement. Dry-run the interval set, isolate or coordinate overlapping writers, validate staged outputs and publish coherently. Catchup and manual task clearing can create many attempts; attempt history, code bundle version and result identity must remain traceable.
Concurrency controls protect dependencies, quotas and fairness. Limit active workflow runs, task parallelism, shared database/API pools and tenant queues separately. A global worker limit does not protect one fragile service, and unlimited catchup can starve current freshness. Priorities need aging or reserved capacity to prevent starvation. Execution infrastructure—local process, containers, Kubernetes Jobs or serverless workers—must use versioned artifacts, bounded resources, short-lived identity, network and secret policies, and compatible cancellation/termination behavior.
Operational health spans scheduler/dispatcher lag, queue time, worker availability, task duration and attempts, deadline misses, sensor age, pool saturation, event-to-run latency, data-interval freshness, output quality and downstream delivery. Alert on consumer impact with runbooks for retry, skip under approved semantics, quarantine, rollback, backfill or capacity isolation — and route those alerts to on-call with an escalation path when the first responder doesn't acknowledge, because a freshness miss at 06:00 that pages nobody until a stakeholder complains at 10:00 is a second failure stacked on the first. Logs, parameters, XCom-like metadata, lineage and error payloads can contain sensitive data and need retention and access controls. When an interviewer asks "how do you know the pipeline is healthy," the weak answer is task success rate; the strong answer leads with freshness and correctness as seen by the consumer, then works backward to the signals.
Deploy workflows as versioned software and data contracts. Parse/lint graph definitions, test task logic, dependency and trigger rules, simulate failures, validate old-run behavior under new code, canary schedules and preserve rollback. In distributed systems, scheduler, workers and definitions may briefly disagree unless runs pin a bundle or image version. Disaster recovery must cover orchestration metadata, schedules, run state, code artifacts, secrets/connections and the actual data systems; rebuilding a scheduler database alone does not restore external effects.
Scheduler, run, and infrastructure mechanics
Airflow separates the scheduler, DAG processor, and workers: the DAG file is parsed into a graph, a DAG run is an instance over a data interval, and a task instance has try_number, state, and timeout. Clearing a task creates another attempt; the attempt number is operational, not a business key. Pools cap slots for a shared resource; a pool of 8 on a 32-worker cluster still overloads a 4-connection OLTP if every task uses that pool. Backfill creates historical DAG runs with explicit ranges, dry-run support, and concurrency that can starve the live schedule unless isolated. Prefect deployments version flow code, parameters, schedules, and concurrency; work pools bind those deployments to infrastructure templates so a flow is not "whatever laptop ran it." Kubernetes Jobs are run-to-completion: backoffLimit, parallel completions, and ttlSecondsAfterFinished define retry and cleanup; a Job that exits 0 after writing a partial object-store prefix is still a successful Job.
OpenLineage's Airflow integration emits job, run, and dataset events from operators that are instrumented. An uninstrumented PythonOperator that copies files is invisible; treating the Airflow graph as lineage will miss it. Pin the DAG bundle or container image on the run so a worker that pulls "latest" mid-backfill does not mix code versions into one interval. Sensors that poke an external API every 30 seconds without timeout become a denial-of-service against that API and a false wait against completeness.
Compatibility, cancellation, and recovery figures
Workflow definition changes must declare whether in-flight runs keep the old graph. Adding a required upstream mid-day can leave existing runs with no way to satisfy the new edge; removing a task can orphan downstream data that still exists. Cancellation and SIGTERM handling differ across local processes, Celery workers, and Kubernetes: a killed task that already committed a sink needs the same idempotent key on retry. Track scheduler heartbeat lag, time-in-queue, sensor duration, pool utilization, DAG parsing time, failed-attempt rate, data-interval age versus SLA, and mean time to recover a stuck backfill. A 99% task-success rate with a 14-hour freshness miss is an orchestration failure. Disaster recovery that restores the metadata database without the object-store outputs, connection secrets, or code artifacts reconstructs a scheduler that cannot re-create the world.
Choose orchestration complexity from outcomes. Simple independent jobs may need only a managed scheduler and run-to-completion platform. Cross-system dependencies, partitioned backfills, human approvals, lineage and recovery may justify a richer orchestrator. Measure trustworthy data delivered within the useful interval, recovery and operator burden—not DAG count or task success percentage. The "do you even need an orchestrator" question is a senior-level probe: the weak answer reaches for the familiar tool; the strong answer sizes the problem first.
Worked example: try_number 2 is not a new business key
Airflow task load_orders POSTs to a warehouse merge API. Attempt 1 times out after the 202.
| DAG run | try_number | idempotency key | warehouse rows for 2026-09-10 | DAG leaf |
|---|---|---|---|---|
orders@2026-09-10 | 1 | new UUID | 1 (maybe) | failed / up_for_retry |
| same interval | 2 | another UUID | 2 | success |
| same interval | 2 | orders:2026-09-10 | 1 | success |
A green DAG with two keys is a duplicate load. The attempt number is operational; the interval is the identity. This is the table to reproduce on a whiteboard when asked "what breaks when a task retries after a partial commit" — the failure is invisible to the orchestrator because both attempts exit 0.
