Overview
Curated: · Written: · Reviewed:
Choose integration from required behavior, not transport fashion
"How would you integrate system X with system Y?" is the opening question for most senior integration design interviews, and the weak answer is a transport: "I'd use REST" or "I'd put a Kafka in between." The strong answer starts with the interaction, not the mechanism. Who owns the fact or operation? Does the caller need a response now? What latency, staleness, volume, burst, ordering, delivery, consistency, security, replay and lifecycle constraints apply? Only then does REST, a queue, a stream, a webhook or a file earn its place.
Interviewers probe this topic by pushing on exactly the seams where a fashionable answer falls apart: what happens when the broker is down, what happens when the same message arrives twice, what happens when a consumer sees events out of order. If you cannot trace those three failures end to end, the rest of the design does not hold. This guide covers the decision axis, then the failure semantics interviewers use to test it.
The choice axis: what each mechanism does to coupling, latency, failure and backpressure
Synchronous request-response (REST, gRPC, RPC). The caller needs an immediate result and both parties share an availability and latency budget. Coupling is tight: the caller is down when the callee is down, and p99 latency is at least the callee's p99. Define timeouts, cancellation, bounded retries, idempotency for retryable mutations, rate limits, response contracts and degraded behavior. Avoid long chains of synchronous calls — tail latency and failure probability compound, and a caller can exhaust threads or connections while waiting. A successful HTTP status is transport evidence, not proof that a downstream business process completed.
Asynchronous messaging (queues, pub-sub). Producers hand off work without waiting. A durable queue buffers bursts and isolates producer availability from consumer availability — this is temporal decoupling, and it is the answer when the caller cannot absorb the callee's failure or latency. Competing consumers increase throughput, but completion order changes and messages get redelivered after crashes or visibility timeouts. Backpressure becomes queue depth and consumer lag, which you must monitor rather than let grow silently.
Streaming (Kafka, Kinesis, Pulsar). An ordered, retained log within partitions, replayable for new consumers and for recovery. Streaming fits when consumers need history, reprocessing, or many independent readers of the same facts. It does not create one global order and does not guarantee exactly-once application at every consumer.
Webhooks. Externally delivered events over HTTP — the producer calls you. Covered in detail below, but on the axis: low setup cost, but delivery reliability is yours to handle, not the producer's.
Batch or file exchange. When volume is high, boundaries are legacy, or latency tolerance is minutes-to-hours, record-at-a-time APIs are inefficient. Coupling is loosest — schema and contract only — but staleness is highest and reconciliation is mandatory.
The interview follow-up is almost always "what breaks if you pick the wrong one?" Picking sync where the callee is slow couples your availability to theirs; picking async where the caller needs a real answer adds a polling or correlation problem you may not want. Say the trade-off out loud before naming a technology.
Delivery semantics: what exactly-once actually costs
Brokers commonly provide at-least-once delivery, which means duplicates are normal. At-most-once means a crash can silently drop work. Exactly-once is the one interviewers push on, and the correct move is to name the boundary: exactly-once broker transfer (achievable with transactional producers and consumer offsets committed atomically with processing, as Kafka's read_committed + send transaction does within the broker) is not the same as exactly-once external side effects. A consumer that writes a database and then calls a payment API can still split those two effects — the broker cannot roll back the payment.
So the real answer is idempotent receivers, not a promise from the broker:
- Idempotency keys on mutating operations, with a uniqueness constraint or a processed-key record so a retry lands on the same result instead of a second effect.
- Naturally idempotent state transitions where possible: "set status to SHIPPED" beats "increment shipped_count".
- Compare-and-set on entity version: reject updates older than the version already applied.
- Transactional inbox/outbox for the dual-write gap (below).
- Reconciliation as the backstop, because deduplication state has a retention horizon and a replay can exceed it.
A weak answer says "we use exactly-once." A strong answer says "at-least-once from the broker, deduplication at the consumer keyed on message ID, dedup state retained past the replay horizon, reconciliation nightly." Interviewers follow up with: how long do you keep dedup records, and what happens on a replay older than that? Have the answer.
Ordering: per-key, not global
Parallel consumers trade global order for throughput — this is the trade, and demanding global order usually kills the throughput that made async attractive in the first place. If an invariant requires order (account balance updates, status transitions for one entity), choose a stable partition or session key and process that key sequentially while other keys run in parallel. That gives you per-entity ordering, which is what business invariants almost always actually need.
Then handle out-of-order arrival anyway, because it happens even within a partition during rebalances or redelivery: carry a sequence number or entity version, reject stale updates, detect gaps. Do not rely on timestamp order across unsynchronized producers. Hot keys, stalled messages and repartitioning need explicit behavior; a single poison record must not block unrelated work indefinitely.
The follow-up probe: "two updates to the same account arrive on different partitions — why?" The answer is a bad or missing partition key, and the fix is keying by entity, not by round-robin or message timestamp.
Failure handling as a pattern set
Interviewers expect you to name the pattern for each failure mode, and to distinguish transient from permanent failure — retrying a 400 Bad Request or a schema violation is a retry storm in the making.
- Timeouts on every hop, budgeted end to end: if the caller allows 2 s, the callee's timeout must be smaller.
- Retries with exponential backoff and jitter for transient failures; jitter prevents synchronized retry storms when a dependency recovers and every client hammers it at the same instant.
- Circuit breakers to stop retrying a dependency that is down and give it room to recover, with a half-open probe before restoring traffic.
- Bulkheads so one slow dependency's connection pool does not starve the rest of the process.
- Dead-letter queues and poison-message quarantine for permanent failures, with alerting and a redrive runbook — a DLQ nobody reads is a silent data loss channel.
- Backpressure and lag monitoring on async paths: queue depth age is an SLO, not a dashboard decoration.
Acknowledgment must follow durable business effect, not merely receipt — acking before the database write is how messages get lost.
The dual-write gap and the outbox
Updating local state and publishing an event in separate operations can leave state without an event or an event without state. A transactional outbox stores the state change and publication intent atomically in the local database; a relay (poller or change-data-capture) publishes and marks progress idempotently. Consumers must still tolerate duplicate publication — the outbox fixes atomicity, not delivery semantics. CDC-based relays make the CDC pipeline's ordering, schema, retention and recovery your production dependency.
Webhooks versus polling versus streaming
A concrete sub-decision interviewers like because it has real security and reliability teeth:
- Polling: simple, but you pay for empty responses and your staleness is your poll interval. Fine for low-frequency, coarse-grained checks.
- Webhooks: push, near-real-time, but delivery is over the public internet to an endpoint you own — so you handle missed, duplicated and reordered deliveries. Verify the sender with a signature over the raw request body (verify before parsing; a re-serialized body will not match the signature), enforce a timestamp/replay window, deduplicate on stable event IDs, return 2xx quickly after durable acceptance and process asynchronously, rotate secrets without downtime, and provide replay/backfill for missed events — a good producer offers a "fetch events since cursor" API to reconcile against. Do not trust source IP alone; do not log sensitive payloads. Rate limits exist on both ends: the producer throttles your endpoint, and your handler must absorb bursts with its own buffer or queue.
- Streaming: when you need history and continuous flow, not just notifications — the retained log is the backfill mechanism.
The weak answer treats a webhook as a reliable function call. The strong answer treats it as an at-least-once, possibly-reordered event feed that you must verify, deduplicate and reconcile.
Contracts, security and operations
Contracts need semantic versioning, not only syntactic compatibility: field meaning, identity, units, nullability, enum evolution, timestamps, privacy classification and producer authority. Prefer additive evolution; consumers ignore unknown fields; producers retain old required behavior through an announced window. Semantic repurposing is breaking even when deserialization succeeds.
Secure each trust boundary: authenticate workloads and users separately, authorize the operation and resource, use least-privilege topics and queues, protect data in transit and at rest, and never take tenant identity from an untrusted payload when it can be bound from authenticated context. Do not forward broad user credentials through every hop.
Operate integrations as data flows: acceptance, completion, end-to-end latency, error reason, retry, duplicate, dead-letter, lag, backlog age, schema/version, reconciliation discrepancy, business outcome. Correlate without exposing payloads. Test dependency loss, timeout, duplicate, reorder, poison, partial batch, expired credential, schema change and regional failure — these are also the follow-up questions.
For multi-step workflows, choose orchestration or choreography deliberately: an orchestrator makes state, timeouts, retries and compensation visible but becomes a dependency; choreography preserves autonomy but hides coupling as participants grow. A saga's compensation is a business operation, not a universal rollback, and can require human resolution. Persist workflow state and make every transition idempotent.
The simplest pattern that meets the required semantics is usually best. Integration platforms standardize identity, policy, observability and connectors; they do not remove your ownership of contracts, data correctness or failure behavior.
Worked example: place order, then email, queue dies
Checkout must persist order 4419 and send a confirmation. The queue broker goes down for 90 seconds right after the HTTP handler returns.
| pattern | order 4419 in DB | confirmation email | after the outage |
|---|---|---|---|
dual write (SQL commit, then Publish) | yes | no | order without email; no retry key, so the email is lost unless a reconciliation job finds it |
| transactional outbox | yes, plus an outbox row in the same commit | no yet | poller sees the outbox row once the broker returns, publishes, email arrives once |
| fire-and-forget HTTP 200 from the broker API | maybe | maybe | 200 means the broker accepted the bytes, not that a consumer processed anything |
Trace the outbox row's life: the handler commits orders and outbox in one transaction, so both exist or neither does. The poller reads unpublished outbox rows, publishes order-4419-confirmation, marks the row published. If the poller publishes but crashes before marking, the row is republished on restart — the email consumer must deduplicate on the order ID, which is why the outbox fixes atomicity, not duplicates.
A 200 from the broker is not the interview. Whether order and email share one write authority is.
