Overview
Curated: · Written: · Reviewed:
A webhook is you making an HTTP request to a stranger's server
A webhook inverts the usual direction: instead of a client polling for changes, the provider sends an HTTP request to a URL the consumer registered. That removes polling latency and wasted requests, which is the appeal.
The consequence, and the source of everything difficult, is that you are now an HTTP client calling an endpoint you do not control, do not monitor, and cannot fix. It may be slow, down, misconfigured, returning 200 while doing nothing, or — since the URL is supplied by a user — pointing somewhere it should not. Every design decision below follows from that.
What interviewers actually probe
The webhook question is a systems-design proxy. Interviewers use it to check whether you reason about delivery guarantees, idempotency, and adversarial input without being asked. Expect these follow-ups in roughly this order:
- "What delivery guarantee do you provide?" — they want at-least-once plus consumer idempotency, said plainly, with the reason why exactly-once is not on the table.
- "Your consumer's endpoint is down for six hours. What happens?" — retry policy, circuit breaking, dead-lettering, and replay.
- "How does the consumer know the request is really from you?" — HMAC over the raw body, timestamp tolerance, constant-time comparison.
- "The consumer registered
http://169.254.169.254/latest/meta-data/. Now what?" — SSRF. If you have not thought about this, this is where the interview goes sideways. - "Why not just use Kafka?" — organisational boundary; you cannot hand a broker to a stranger's account.
A weak answer describes the happy path: "the provider POSTs JSON to the consumer's URL." A strong answer spends most of its time on the unhappy path, because that is the entire engineering problem.
Delivery is at-least-once, and receivers must be idempotent
A delivery can fail after the receiver processed it: the response is lost, the connection drops, a timeout fires while the work has already completed. The sender cannot distinguish that from a genuine failure, so it retries, and the receiver sees the event twice.
That makes at-least-once the only honest guarantee, and receiver idempotency the fundamental requirement. Every event needs a stable identifier the receiver can record, so a repeat is recognisable. A sender that does not provide one has pushed an unsolvable problem onto its consumers.
Ordering is not guaranteed either, and cannot be without serialising delivery per receiver — which one slow endpoint then stalls entirely. So events should carry a timestamp or sequence number, and receivers should ignore anything older than what they have applied. Retries reorder events even when the original sends were ordered, so this holds regardless of what the sender does.
The outbound delivery pipeline
On the sending side, a webhook system is a queue with hostile consumers. The shape that survives production:
- Enqueue on event. The application writes a delivery record (event id, endpoint, payload) to a durable queue and returns. Never deliver inline with the request that generated the event — a slow consumer then stalls your API.
- Partition per endpoint. Deliveries to the same endpoint go to the same partition/queue, so a single worker handles them in order. This gives per-endpoint ordering without a global lock; one dead endpoint stalls only its own partition.
- Delivery workers pull, make the HTTP call with a hard timeout, and record the outcome.
- Retry or dead-letter based on the result, with per-endpoint state (failure counts, circuit state) feeding back into scheduling.
The failure mode to name in an interview: a shared worker pool with no per-endpoint isolation. One large consumer whose endpoint hangs holds connections open until timeout; with a 10-second timeout and 50 workers, 5 hung endpoints consume the entire fleet and delivery for every other consumer stops. Per-endpoint concurrency caps and aggressive timeouts are not optimisations, they are the difference between "degraded for one customer" and "down for all".
Retries, and the cost of doing them badly
Retry with exponential backoff and jitter, over a bounded window — hours or a day, not forever. Immediate retries against a struggling endpoint add load precisely when it is least able to take it, and synchronised retries across many events produce a burst that prevents recovery. Jitter exists specifically to break that synchronisation.
Distinguish retryable from permanent. A timeout, a 5xx or a connection failure should be retried. A 400 or a 410 will fail identically forever, and retrying wastes capacity. A 410 Gone in particular is a receiver saying the endpoint is permanently retired, and a sender that ignores it is generating pure waste.
Circuit-break per endpoint. If a receiver has been failing for an hour, stop attempting immediately and back off hard. Without this, one dead endpoint consumes a share of the delivery capacity indefinitely, and a large receiver going down can degrade delivery for everyone.
After the retry window, dead-letter and tell someone. Silently dropping events is the worst outcome, because the receiver's data is now wrong and neither party knows. Notify the consumer, expose the failures in a dashboard, and provide a replay mechanism.
A concrete policy you can defend: 6 attempts, backoff 30s · 1m · 5m · 30m · 1h · 6h with ±20% jitter, 5-second HTTP timeout, then dead-letter with an email to the consumer and a replay button in the dashboard. The delays sum to 30 + 60 + 300 + 1,800 + 3,600 + 21,600 = 27,390 seconds, about 7.6 hours — so the last attempt lands roughly 7.6 hours after the first, and anything still failing after that goes to the dead-letter queue. The exact numbers matter less than having bounded ones and being able to say where each came from.
Security, in both directions
Signing is not optional. A webhook endpoint is a public URL accepting POSTs, so anyone can send it anything. The receiver must verify the request came from you: an HMAC signature over the raw body with a shared secret, or an asymmetric signature. Include a timestamp in the signed payload and reject old requests, or a captured request can be replayed indefinitely.
Two implementation details matter. The signature must be computed over the raw body before any parsing, since re-serialising changes bytes and breaks verification. And comparison must be constant-time, or the comparison itself leaks the expected value.
Receiver-side verification, in Python 3.12, using the scheme Stripe documents (timestamp + "." + body as the signed payload, 5-minute tolerance):
import hmac, hashlib, time
WEBHOOK_SECRET = b"whsec_..." # shared secret, from config not code
TOLERANCE = 300 # seconds
def verify(raw_body: bytes, stripe_sig_header: str) -> bool:
parts = dict(p.split("=", 1) for p in stripe_sig_header.split(","))
signed_payload = parts["t"].encode() + b"." + raw_body
expected = hmac.new(WEBHOOK_SECRET, signed_payload, hashlib.sha256).hexdigest()
if not hmac.compare_digest(expected, parts["v1"]):
return False
return abs(time.time() - int(parts["t"])) <= TOLERANCE
Note what the code does and does not do: it reads raw_body straight off the socket (parse the JSON after this returns), it uses hmac.compare_digest rather than ==, and the timestamp check happens after the signature check so an attacker cannot probe the tolerance window on unsigned requests. What it does not do: rotate secrets — see below.
Secret rotation. Consumers must be able to rotate the shared secret without downtime. The standard pattern is a grace period where the sender signs with the new secret and the receiver accepts either the old or new signature; Stripe does this by including multiple v1 signatures in the header during rotation. If your answer to "how do you rotate the secret?" is "re-register every consumer at once," you have designed an outage.
The sender has a threat too. Consumers supply the destination URL, so the sender is making requests to arbitrary addresses — which is server-side request forgery. Validate that URLs are public: reject private address ranges, loopback, link-local and cloud metadata endpoints, and re-check after DNS resolution and after every redirect, since a hostname can resolve to a private address and a redirect can point at one. DNS rebinding — a hostname that resolves public at registration and private at delivery time — is why the check belongs at request time, not at registration time only.
Additionally: require HTTPS, set aggressive timeouts, cap response size, limit redirects, and never expose response bodies back to the consumer, since that turns the sender into a proxy for reading internal services.
The receiver's side
Acknowledge fast, process later. Validate the signature, persist the event, return 2xx, and process asynchronously. Processing inline risks exceeding the sender's timeout, which causes a retry, which causes duplicate processing while the first is still running. A receiver that returns quickly is also a receiver the sender does not back off from.
The handler shape, Python 3.12, framework-agnostic:
def webhook_handler(raw_body: bytes, sig: str) -> tuple[int, str]:
if not verify(raw_body, sig):
return 401, "invalid signature"
event = json.loads(raw_body)
if event["type"] not in HANDLED_TYPES:
return 200, "" # unknown type: ack, ignore
if db.insert_event(event["id"]): # unique constraint on event id
enqueue(event) # first sighting: schedule processing
return 200, "" # duplicate: already recorded, ack
The db.insert_event line is the idempotency mechanism made concrete: a unique constraint on the event id means the second delivery takes the same path as the first and returns 2xx either way. If the insert fails on a constraint violation, the event was already seen — ack and move on. If instead you returned 500 on the duplicate, the sender would retry a delivery that has already succeeded, forever.
Return the right status. 2xx means received; anything else invites a retry. Returning 200 while failing internally means the event is lost with no signal to anyone — the receiver-side equivalent of swallowing an exception.
Handle unknown event types gracefully by ignoring them, so a sender adding a new type does not break existing receivers.
Worked example: 200 after charge, timeout, retry
Event payment.succeeded / evt_4419. The receiver charges the card, then the sender's HTTP client times out at 5s. The sender's retry schedule fires again at +30s. Trace both deliveries against the handler above:
| delivery | receiver behaviour | charges | response sender sees |
|---|---|---|---|
| 1 (t=0) | verify ok, insert evt_4419 succeeds, enqueue charge, return 200 — but response lost to the timeout | 1 | timeout → retryable |
| 2 (t=30s) | verify ok, insert evt_4419 hits unique constraint, skip enqueue, return 200 | 1 | 200 → done |
Without the event-id insert, delivery 2 charges the card a second time. With it, the duplicate costs one database write. That is the whole argument for idempotency in one table: the event id is the business key, not the HTTP attempt.
Alternatives, and when each fits
Polling is simple, needs no public endpoint, and the consumer controls the rate. It costs latency and wasted requests, and it does not scale well across many consumers — but for low-volume, latency-tolerant integrations it is genuinely fine, and it is often dismissed too quickly.
Long polling reduces latency by holding a request open until there is something to return. Better than tight polling, still one connection per waiting consumer.
Server-sent events stream one-way from server to client over ordinary HTTP, with automatic reconnection and event identifiers for resumption. Well suited to notification feeds where only the server has something to say.
WebSockets give full-duplex communication, which is what you want for genuinely interactive applications. The cost is a stateful long-lived connection: connection-aware load balancing, capacity management, and reconnection handling. Application-level sticky sessions are needed only when connection state cannot move or be shared.
A message queue between systems you both control is usually better than webhooks, because a durable broker provides buffering and backpressure; a retained log also provides replay without custom delivery machinery. Webhooks are for crossing an organisational boundary where a shared queue is not available.
Providing both webhooks and a polling API is a good default. Webhooks for latency, polling for reconciliation — so a consumer who missed events while their endpoint was down can catch up without asking you to replay, which is the situation that otherwise generates support load.
