Skip to content
Tech Interview Prep home

Top 100 Cloud Architect Interview Questions and Answers

The questions most likely to actually come up in your Cloud Architect interview, ranked by likelihood — with detailed, senior-level answers covering what an interviewer is really listening for.

Curated: · Written: · Reviewed:

Reviewed 99Review pending 1
QA-1Design a multi-region deployment for a service with a 15-minute RTO and a 5-minute RPO. Which topology do you choose, and what does that choice cost you?(show answer)

A 5-minute RPO rules out backup-and-restore based only on nightly or hourly snapshots, because those can lose more data than the objective allows. Synchronous cross-region replication would meet the RPO, potentially with zero committed-data loss, but it may impose unacceptable write latency and reduce write availability during an inter-region partition. If the service can tolerate up to five minutes of data loss, asynchronous replication with a warm standby is usually the better cost, latency, and availability trade.

Concretely: the primary region serves all writes; a read replica in the secondary region applies the log continuously, running minutes behind at most. Infrastructure in the secondary is provisioned and running but scaled down. Failover promotes the replica and repoints DNS or global load balancing.

What this costs: replication lag is your data loss. If the primary dies with 90 seconds of unapplied log, those transactions are gone. You must monitor replica lag as a first-class SLI and alert when it approaches the RPO, because lag silently exceeding 5 minutes means you no longer meet the objective you think you meet.

The 15-minute RTO is what forces "warm" rather than "cold". Cold standby means provisioning capacity during an outage, and capacity acquisition in a regional event is exactly when it is slowest and least reliable. The honest cost is running duplicate infrastructure that earns nothing until the day it does.

The number that matters most is not in the design, it is in the drill: an untested failover has an unknown RTO, not a 15-minute one.

Curated: · Written: · Reviewed:

QA-2A team wants to give an application permanent access keys so it can read from an object store. Why do you push back, and what do you propose instead?(show answer)

Long-lived access keys fail on three counts: they do not expire, so a leak is permanent until someone notices; they are copied into config files, CI variables, and laptops, so the blast radius is unknowable; and they carry no context about which workload used them, so an audit log tells you the key acted, not who did.

The alternative is an identity attached to the compute itself — an instance role, a workload identity, or a service account federated through OIDC. The platform issues short-lived credentials, typically valid for an hour, and rotates them automatically. Nothing is ever written to a file.

For CI systems outside the cloud provider, OIDC federation does the same job: the CI platform presents a signed token asserting which repository and branch is running, and a trust policy exchanges it for temporary credentials scoped to that claim. A stolen token is useless after minutes and only works from the asserted context.

The permission itself should also be narrower than the request implies. "Read from an object store" should be a policy naming the specific bucket and prefix, with read-only actions, not a managed full-access policy. The common failure is granting broad access on day one because it is faster, then never narrowing it — so write the narrow policy first and widen it when something breaks, which surfaces the real requirement.

If a long-lived key is genuinely unavoidable — a third-party SaaS with no federation support — then it needs a documented owner, a rotation schedule that is actually executed, and monitoring for use from unexpected locations.

Curated: · Written: · Reviewed:

QA-3Walk me through what happens when a request enters your VPC and reaches a private database. Which controls does it pass, and what breaks if one is misconfigured?(show answer)

The request arrives at an internet-facing load balancer in a public subnet — public meaning its route table has a default route to an internet gateway. The load balancer's security group allows 443 from the internet.

The load balancer forwards to an application instance in a private subnet, whose route table has no internet gateway route; its outbound traffic goes through a NAT gateway instead. The application's security group allows the application port from the load balancer's security group, referencing the group rather than a CIDR range, so the rule stays correct as instances come and go.

The application connects to the database in a data subnet. The database security group allows the database port from the application's security group only. Network ACLs may sit at the subnet boundary as a coarse second layer; unlike security groups they are stateless, so return traffic needs an explicit rule.

The instructive misconfigurations: a private subnet with an internet gateway route is not private, and nothing in the console labels it wrong. A security group opening the database port to 0.0.0.0/0 is a direct path from anywhere the subnet is reachable. A stateless NACL that allows inbound but not ephemeral outbound ports produces connections that establish and then hang, which reads like an application bug for hours.

The invariant worth stating: a subnet is public or private because of its route table, and a security group is safe because of what it references, not because of what it is named.

Curated: · Written: · Reviewed:

QA-4What actually breaks first when traffic to a service grows ten times, and how do you find out before your users do?(show answer)

Rarely the application code. In practice the first thing to break is whatever has a fixed limit that nobody is watching: the database connection pool, a single-writer primary, a downstream third-party rate limit, or a queue whose consumers cannot keep up.

The database connection pool is the most common. Application instances autoscale, each opens a pool, and the database's max connections is a fixed number. Doubling instances doubles connections until the database refuses them — and the failure looks like a total outage, not gradual degradation, because new connections are rejected while existing ones work.

Finding it before users do means load testing against representative data volumes, not empty tables, and watching utilization against limits rather than against absolute values. A metric of 400 connections tells you nothing; 400 of 500 tells you when you have weeks left.

The specific signals worth alerting on are saturation signals: connection pool utilization, queue depth and its first derivative, replica lag, CPU credit balance on burstable instances, and remaining capacity in any fixed quota. Latency and error rate tell you a problem has already reached users. Saturation tells you one is coming.

The other half of the answer is overload behavior. At ten times traffic something will exceed capacity eventually, and the design question is whether it degrades or collapses. Load shedding, admission control, and bounded queues make the overload survivable; unbounded queues turn a busy period into a cascading failure with a growing backlog nobody can drain.

Curated: · Written: · Reviewed:

QA-5Explain the difference between sharding and replication, and describe a situation where adding replicas makes things worse.(show answer)

They solve different problems. Replication copies the same data to more machines: it buys read capacity, availability, and durability. Sharding splits different data across machines: it buys write capacity and storage beyond one machine.

If your bottleneck is write throughput or dataset size, replicas do not help — every replica applies every write, so the write load per machine is unchanged. If your bottleneck is read throughput, sharding is a much heavier tool than you need.

Replicas make things worse in at least three situations. First, when the application reads its own writes: with asynchronous replication a user updates a profile, is routed to a replica that has not applied the change, and sees stale data. The fix is routing reads that follow a write to the primary, or a session-consistency token — but it must be designed, not discovered.

Second, when write amplification is the real problem: adding replicas increases the total write work in the cluster and the replication traffic, and a primary already saturated by writes gets no relief and some additional load.

Third, when replicas mask a query problem. A missing index turns into "add another replica" because it distributes the pain, and the team pays monthly for what an index would have fixed once.

The diagnostic question is which resource is saturated. Read IOPS and read CPU point to replication; write IOPS, replication lag under normal load, or storage growth point to sharding — which is a much larger commitment, since cross-shard joins and transactions largely stop being available.

Curated: · Written: · Reviewed:

QA-6Your cache hit rate is 95% and the origin still falls over every time you deploy. What is happening?(show answer)

A 95% hit rate means the origin is sized for 5% of traffic. A deploy that empties or invalidates the cache does not increase load by 5% — it multiplies origin load by twenty, instantly. That is a thundering herd, and the origin is being asked to serve a volume it was never provisioned for.

The related failure is a stampede on a single hot key: when a popular entry expires, every concurrent request misses simultaneously and all of them recompute the same value.

Three mitigations, in order of how much they help here. First, do not empty the cache on deploy: version cache keys so old and new entries coexist and the new namespace fills gradually, or keep the cache external to the deployment unit so a rollout does not touch it. Second, request coalescing — the first miss for a key computes the value while the rest wait on that single in-flight computation, turning a thousand origin requests into one. Third, serve-stale-while-revalidate: return the expired entry immediately and refresh in the background, so an expiry never becomes a miss on the request path.

Randomized TTL jitter matters too. Entries populated together expire together, so a uniform TTL reproduces the herd on a schedule.

The underlying point is that a high hit rate is not resilience, it is a dependency. The number to know is what the origin can serve at 0% hit rate, because that is your actual floor.

Curated: · Written: · Reviewed:

QA-7When would you choose eventual consistency, and how do you explain that choice to a product owner who hears "sometimes wrong"?(show answer)

Choose it when the cost of unavailability during a partition exceeds the cost of a brief disagreement between replicas — and when the domain has a sensible merge or last-writer rule.

Good fits: view counts, feed ranking, product recommendations, notification delivery, most analytics. Bad fits: account balances, inventory decrement at the point of sale, uniqueness constraints, and anything where two "correct" answers combine into a real-world loss.

The explanation for a product owner should be in their units, not CAP's. Not "we relax consistency" but "if the network between our regions breaks, would you rather the checkout page show a slightly stale item count, or return an error?" That is the actual trade, and it is a product decision, not an engineering one.

It also helps to name the window. Eventual consistency does not mean unbounded: replication typically converges in tens or hundreds of milliseconds, and the interesting question is what happens during the rare partition. Saying "readers may be up to two seconds behind, and during a regional partition up to a few minutes" is concrete enough to reason about.

Finally, be precise that CAP applies during a partition. When the network is healthy a system can be both consistent and available, and the real day-to-day trade is PACELC's second half: even without partitions, stronger consistency costs latency. That is usually the trade the product owner actually feels.

Curated: · Written: · Reviewed:

QA-8How do you make a payment API safe to retry?(show answer)

With an idempotency key supplied by the client and enforced by the server.

The client generates a unique key per logical operation — not per HTTP attempt — and sends it as a header. The server records the key and a request fingerprint under a unique constraint, then returns the stored result for a completed duplicate and rejects the same key with different parameters.

If the "charge" is only a mutation in the service's own database, record the key and the mutation in one transaction. If it calls an external payment processor, no local database transaction can make both systems atomic. Forward the same idempotency key to a processor that supports idempotent requests and persist an in-progress/succeeded/failed state locally; an outbox or recovery worker must reconcile the uncertain case where the remote call succeeds but the local result is not recorded.

Concurrency needs care: two simultaneous retries can both reach the server. The unique constraint makes one win; the loser must wait for the winner's result rather than returning "in progress" and inviting another retry. Storing a state column — in-progress, succeeded, failed — and having the loser poll or block briefly handles this.

The key must also be scoped to the request body. If a client reuses a key with different parameters, that is a client bug and the server should reject it with a 4xx rather than silently returning the wrong prior response.

Keys need a retention window — 24 hours is typical — long enough to cover any realistic retry and short enough to bound storage. And the response you replay should include the original status code, because a retry of a failed charge must not look like a success.

Curated: · Written: · Reviewed:

QA-9A query that ran in 20ms last month now takes 8 seconds. The code has not changed. Walk me through your diagnosis.(show answer)

Start with the execution plan, not with theories. Run EXPLAIN ANALYZE and compare the plan to what you expect. The question is whether the planner changed its mind or the work genuinely grew.

The common causes, roughly in order of frequency:

Data growth crossed a threshold. A sequential scan on 10,000 rows is fine and on 10 million is not; the plan may be unchanged and simply no longer viable. Look at row counts, not just timing.

Statistics went stale. The planner estimates rows from sampled statistics; after a bulk load, estimates can be badly wrong, and a plan chosen for an estimated 50 rows is catastrophic for an actual 5 million. Compare estimated versus actual rows in the plan — a large divergence is the tell. Running ANALYZE is the fix.

An index stopped being used. A change in a WHERE clause's type, a function wrapped around a column, or an implicit cast makes an index unusable. The plan shows a scan where an index scan used to be.

Bloat or fragmentation from heavy update/delete traffic leaves a table much larger on disk than its live row count implies.

Parameter-sensitive plans: a cached plan chosen for one parameter value is wrong for another with very different selectivity.

Only after the plan is understood does anything else matter — lock contention, a saturated buffer cache, or a noisy neighbor. Those produce variable slowness; a consistently slow query with an 8-second floor is almost always the plan.

Curated: · Written: · Reviewed:

QA-10Your Terraform plan shows it will destroy and recreate a production database. What do you do?(show answer)

Stop, and do not apply. A destroy-and-recreate on a stateful resource is data loss, and the plan is telling you before it happens — which is the entire value of a plan step.

First, find which attribute forced replacement. The plan output names it, marked as forces replacement. Common causes: a change to an immutable field like engine version, availability zone, or subnet group; a resource identifier that changed because a module was refactored; or drift where someone modified the resource in the console and the configuration now disagrees.

The remedy depends on the cause. If a module was renamed or moved, the resource is fine and the state is wrong — terraform state mv (or a moved block, which is reviewable and belongs in the code) fixes the address without touching infrastructure. If a genuinely immutable attribute must change, that is a migration with a data path, not a Terraform apply: create the new resource alongside, replicate, cut over, retire the old one.

Two guardrails should already be in place. prevent_destroy on stateful resources turns this from a scary plan into a hard error. And a CI pipeline that runs plan on every pull request, posts the output, and requires human approval before apply means nobody encounters this for the first time at the keyboard at 2am.

The wider lesson is that console changes create exactly this class of drift, which is why break-glass console access should be time-boxed and reconciled afterward.

Curated: · Written: · Reviewed:

QA-11Design the queue between an order service and a slow fulfilment system. What guarantees do you need, and what do you do with messages that fail repeatedly?(show answer)

The queue exists so a slow downstream cannot hold the order request open. The order service writes the order, publishes a message, and returns; fulfilment consumes at its own rate.

Guarantees: at-least-once delivery is what real brokers give you, so the consumer must be idempotent. Fulfilment keyed by order ID with a uniqueness check gives that for free. Exactly-once is a property you build at the consumer, not one you buy from a broker.

Ordering usually matters less than people assume. Global ordering costs throughput, because it forces a single consumer per partition. Per-order ordering is nearly always what the domain needs, so partition by order ID and let unrelated orders proceed in parallel.

Failure handling has three tiers. Transient failures — a timeout, a 503 — get bounded retries with exponential backoff and jitter. Retries must be bounded, or one poison message consumes the consumer forever. After the retry budget, the message goes to a dead letter queue.

The dead letter queue is only useful if someone looks at it. Depth should be alerted on, not merely graphed, and its messages should carry enough context to replay: the original payload, the failure reason, and the attempt count.

The subtle failure is the dual write: the order is committed to the database but the publish fails, or the reverse. The transactional outbox pattern fixes it — write the message to an outbox table in the same transaction as the order, and a relay publishes from there.

Curated: · Written: · Reviewed:

QA-12How do you choose between a relational database and a document store for a new service?(show answer)

Start from the access patterns and the integrity requirements, not from the data's shape. "Our data is nested" is a weak argument; relational databases store JSON well.

Reach for relational when the data has relationships you will query across, when you need multi-row transactional integrity, when uniqueness and referential constraints protect something that matters, or when the query patterns will change in ways you cannot predict. The relational answer's strength is that an unanticipated query is a new SELECT, not a migration.

Reach for a document store when the aggregate is the unit of access — you read and write a whole document by key, with few cross-document queries — when the schema genuinely varies per record, or when horizontal write scaling beyond a single primary is a near-term certainty rather than a hypothetical.

The honest default for a new service with uncertain requirements is relational. It is much easier to denormalize a normalized schema later than to reconstruct relationships from denormalized documents, and the operational tooling is more mature.

Two traps worth naming. Choosing a document store to avoid schema design does not avoid schema design; it moves it into application code where nothing enforces it. And "we might need to scale" is not a reason on its own — a well-indexed relational primary handles far more than most teams estimate, and premature sharding costs you joins and transactions you will miss immediately.

Curated: · Written: · Reviewed:

QA-13What is the difference between a load balancer health check and a readiness check, and what goes wrong when a team conflates them?(show answer)

A liveness check answers "is this process broken and in need of a restart?" A readiness check answers "should this instance receive traffic right now?" They have different consequences: failing liveness kills the process, failing readiness only removes it from rotation.

Conflating them produces two characteristic outages.

The first is the cascading restart. A readiness check that verifies the database connection is used as a liveness check. The database has a brief hiccup, every instance reports unhealthy, every instance is restarted, and the restart storm — cold caches, reconnection thundering herd — turns a five-second database blip into a twenty-minute outage. Liveness should test only the process itself.

The second is traffic to instances that cannot serve. A health check that returns 200 as soon as the HTTP server binds, before dependencies are connected and caches are warmed, means the load balancer sends traffic during startup and users get errors during every deploy. Readiness must reflect actual ability to serve.

The nuance worth mentioning is that readiness should usually not fail on a downstream dependency either. If every instance marks itself unready because a shared dependency is degraded, you have converted partial degradation into total unavailability, and removed the capacity that could have served the requests not touching that dependency. Better to serve degraded responses and let the dependency's own signals drive the alert.

Curated: · Written: · Reviewed:

QA-14You need to add a NOT NULL column to a 500-million-row table in production. How?(show answer)

Not in one statement. Adding a NOT NULL column with a default historically rewrites the table and holds an exclusive lock for the duration, which on 500 million rows means an outage. Modern PostgreSQL avoids the rewrite for a constant default, but the approach below is what survives across engines and versions.

Break it into steps that each hold a short lock:

Add the column as nullable with no default. This is a metadata-only change and is effectively instant.

Deploy application code that writes the new column on every insert and update, while still tolerating nulls on read. Now the set of null rows is closed and can only shrink.

Backfill in batches — tens of thousands of rows per batch, committing between them, with a pause to let replication catch up. A single UPDATE over 500 million rows produces an enormous transaction, bloats the table, and can push replicas hours behind.

Verify no nulls remain, then add the NOT NULL constraint. In PostgreSQL, add it as a NOT VALID check constraint first and validate it separately, which takes a weaker lock than a direct table alteration.

Finally, remove the null-tolerance from the application.

The general principle is expand-migrate-contract: make the schema accept both shapes, move the data, then narrow it. It is more steps and it is the only version that is safe to abandon halfway.

Curated: · Written: · Reviewed:

QA-15How do you size and structure an autoscaling policy so it actually helps during a traffic spike?(show answer)

Autoscaling helps with gradual load and fails at sharp spikes, so the first honest statement is that it is not a spike-absorption mechanism. Instance provisioning takes minutes; a spike arrives in seconds.

Structure it around three things.

Scale on a signal that leads the failure, not one that follows it. CPU is a poor proxy for an IO-bound service. Requests per instance, queue depth, or concurrent connections track the actual constraint. For a queue consumer, backlog age is the best signal: it directly expresses how far behind you are.

Make scale-up aggressive and scale-down conservative. Adding capacity too early costs money; removing it too early costs an outage. Asymmetric cooldowns — quick to add, slow to remove — reflect that asymmetry, and prevent the oscillation where a scale-down raises per-instance load and immediately triggers a scale-up.

Keep headroom for the provisioning window. If instances take three minutes to become ready, you need enough spare capacity to absorb three minutes of growth. Running at 85% utilization with a three-minute lag means the spike wins.

For known events — a sale, a campaign — schedule the capacity rather than reacting to it. And test the scaled state: an autoscaling group that can reach forty instances but shares a database connection limit sized for ten has simply moved the failure.

Curated: · Written: · Reviewed:

QA-16Explain isolation levels by describing a bug that read committed allows and repeatable read does not.(show answer)

A non-repeatable read. Under read committed, each statement sees a fresh snapshot, so two reads of the same row inside one transaction can return different values if another transaction commits between them.

The classic bug: a transfer routine reads an account balance, checks it against the withdrawal amount, and then re-reads the row to build a receipt. Between the two reads another transfer commits, and the receipt reports a balance inconsistent with the decision that was just made.

Repeatable read pins one snapshot for the whole transaction, so both reads return the same value. It does not prevent everything: write skew survives repeatable read. Two transactions each read the same set of rows, each concludes a constraint still holds — say, at least one doctor remains on call — and each writes a different row. Neither sees the other's write, both commit, and the invariant is violated even though every individual read was consistent.

Serializable is what prevents write skew, at the cost of aborted transactions your application must be prepared to retry.

The practical point for an architect is that isolation level is a contract about what the database prevents, and everything it does not prevent becomes application responsibility. Choosing read committed and then writing check-then-act logic without SELECT FOR UPDATE is the single most common source of correctness bugs that only appear under load.

Curated: · Written: · Reviewed:

QA-17Your monthly cloud bill doubled and nobody knows why. How do you investigate, and how do you prevent a recurrence?(show answer)

Start with the bill's own dimensions before touching infrastructure. Group cost by service, then by region, then by tag or account, and compare against the prior period. A doubling is almost never diffuse; it is one line item.

The usual culprits, in rough order: a forgotten non-production environment left running; data transfer, particularly cross-AZ or egress to the internet, which is invisible in most dashboards because no instance is associated with it; storage that only accumulates — snapshots, log retention, orphaned volumes from terminated instances; an autoscaling group whose maximum was raised during an incident and never lowered; and a runaway process making per-request calls to a metered API.

Cross-AZ transfer deserves specific attention because it is architectural, not operational. A chatty service placed for availability across three zones pays for every hop, and the fix is topology-aware routing, not a smaller instance.

Prevention has three parts. Mandatory tagging enforced at provision time — an untagged resource is an unattributable cost, and tags applied later never get applied. Budget alerts on forecast rather than actual, so you hear about it in week one rather than after the invoice. And cost visibility per team, because a bill nobody owns is a bill nobody reduces.

The structural fix is making cost a review criterion in design, where choosing a topology is cheap, rather than an audit afterward, where changing it is not.

Curated: · Written: · Reviewed:

QA-18When is a window function the right tool instead of GROUP BY, and what does that change about the result set?(show answer)

Use a window function when you need a per-row value computed over a related set of rows while keeping the rows. GROUP BY collapses; a window function does not.

The distinction is easiest to see in a concrete task: show each order with the customer's running total. GROUP BY gives one row per customer and loses the orders. A window function — SUM(amount) OVER (PARTITION BY customer_id ORDER BY ordered_at) — returns every order row with the running total attached.

The same shape covers ranking within a group (ROW_NUMBER, RANK, DENSE_RANK), row-over-row comparison (LAG, LEAD for month-over-month change), and percentiles within a partition. The alternative without windows is a self-join or a correlated subquery, which is both slower and considerably harder to read.

Two details that come up in interviews. First, RANK and DENSE_RANK differ on ties: RANK leaves gaps, DENSE_RANK does not, and ROW_NUMBER breaks ties arbitrarily unless the ORDER BY is deterministic. Second, the frame clause matters: with an ORDER BY, the default frame is the rows from the partition start to the current row, which is what makes a running total work. Without ORDER BY, the frame is the whole partition, which gives a per-group total repeated on every row.

Window functions cannot appear in a WHERE clause, because they are evaluated after filtering — that is why filtering on a rank requires a subquery or CTE.

Curated: · Written: · Reviewed:

QA-19A service you designed must keep serving if an entire cloud region goes down. How do you design for that, and what do RTO and RPO change about your answer?(show answer)

RTO picks the topology, RPO picks the replication. Everything else is downstream of those two numbers. Assume an order-processing API: 5,000 req/s peak, 200 writes/s, RTO 5 minutes, RPO 60 seconds, budget ceiling around 2x baseline run cost.

TargetTopologyReplicationSteady-state cost
RTO 5 min / RPO 1 minactive-active serving, single-writer data per cellasync cross-region~2x
RTO 30-60 min / RPO 5 minwarm standby at 25% sizeasync~1.3x
RTO 2-4 h / RPO 15 minpilot lightsnapshot + log shipping~1.1x
RPO 0multi-region consensussynchronous quorum commit~2.5x plus 60-120 ms per write

For these numbers I go active-active for the stateless tier and single-writer-per-cell for data: clients hit an anycast address (Global Accelerator), which routes to region A or B at roughly 3,250 req/s each; a cell router hashes tenant to cell, and each cell is an Aurora Global Database with the writer in A and a secondary in B.

Failover arithmetic: health checks at 30-second intervals with three consecutive failures is about 90 seconds to detect — I would verify that interval in a test rather than trust the default. Aurora Global Database cross-region promotion is documented at around a minute and is not automatic, so it is a runbook step I automate to overlap detection. Add 10-20 seconds for clients to re-establish connections and I land near 2.5 minutes, with the rest of the RTO for cache warm-up and autoscaling. DNS alone cannot make 5 minutes: a 60-second TTL is honoured inconsistently, and I have watched subsets of clients sit on a dead IP for 10-15 minutes. Anycast removes DNS from the critical path; Route 53 stays as a second layer.

On conflicts: single-writer per cell means there is no conflict resolution to get wrong. Multi-writer stores (DynamoDB global tables, Cassandra) resolve concurrent writes to one key last-writer-wins, which silently discards one of them during a partition — acceptable for a cart, not for order state. The price is that reads outside the writer region either cross the backbone or read slightly stale.

RPO honesty: with async replication, RPO is the replication lag at the moment of failure, so I alert per cell at 10 seconds of lag — the 60-second RPO is a budget being spent. RPO 0 forces synchronous cross-region commit onto every write and roughly doubles the data-tier bill. That is a negotiation with the business, not a config flag.

Capacity: surviving a lost region means each region carries peak, so two regions at 65% cost about 1.3x and the survivor jumps to 1.55x its normal load — it either autoscales inside the RTO or sheds non-critical traffic. Cells bound blast radius: losing one cell is 5% of tenants, and I can fail one cell over without touching the region.

Shared responsibility: the provider compensates failures inside its service boundaries, not your decision to run in one region. The multi-region wiring is mine — replication configured and lag monitored, KMS keys (regional, so an encrypted snapshot does not decrypt in the other region; you re-encrypt with a target-region key), ACM certificates, WAF rules, service quotas and vCPU limits in region B, secrets, and the promotion runbook with a named authorized operator.

Proof: a quarterly game day in business hours that blackholes region A at the network edge rather than stopping the app, so health checks and failover are exercised as one path. Two numbers measured from outside: time from declared failure to p99 error rate back under SLO, and acknowledged writes lost, counted from a synthetic client writing a numbered record every second. Include failback — most DR incidents happen there — and run a single cell first.

Failure modes I design for: split brain if promotion happens before the old writer is fenced; health checks green while writes fail, which is why the check must be a synthetic write; and the single-region dependency nobody inventoried, such as one Kinesis stream or one Lambda layer. The game day is what surfaces those.

Curated: · Written: · Reviewed:

QA-20What does a well-designed blue-green deployment give you that a rolling deployment does not, and what does it cost?(show answer)

It gives you an instant, complete rollback and a fully-formed environment to test before any user reaches it. With a rolling deploy, rollback means another rolling operation over the same minutes, and during the roll two versions serve traffic simultaneously whether you wanted that or not.

Blue-green stands up the new version as a separate fleet, verifies it, then shifts traffic at the router. If the new version misbehaves, the shift is reversed in seconds because the old fleet is still running and warm.

The costs are real. You pay for double capacity during the transition, which for a large fleet is significant. More importantly, the shared database does not get a blue and a green: both versions read and write the same schema, so every schema change must be backward compatible with the version you might roll back to. That constraint — expand, migrate, contract — is the actual discipline; the traffic switch is the easy part.

Stateful connections complicate it too. Long-lived WebSockets or in-flight background jobs do not move at the router, so you need drain behavior and a defined maximum drain time.

A canary sits between the two: shift 1% of traffic, watch error rate and latency against the control, then proceed. It costs less capacity than blue-green and gives real production evidence that a staging test cannot, at the price of a slower rollout and the need for automated analysis to make the go/no-go call.

Curated: · Written: · Reviewed:

QA-21How do you decide what belongs in a shared platform team versus in each product team?(show answer)

Put something in the platform when it is undifferentiated, when getting it wrong is expensive, and when the cost of divergence exceeds the cost of coordination. Leave it with product teams when it is close to the domain or when teams genuinely need different answers.

Clear platform candidates: identity and access patterns, network topology, secret management, the CI/CD path to production, observability plumbing, and the base images and infrastructure modules. Nobody wants six answers to how a service gets credentials, and six answers means five of them are unreviewed.

Clear product-team territory: the service's data model, its API contract, its scaling behavior, its runbook. A platform team that owns these becomes a bottleneck and, worse, becomes accountable for decisions it lacks the context to make.

The design rule that keeps this healthy is the paved road: the platform's offering must be the easiest path, not the mandatory one. If the golden path is genuinely easier than doing it yourself, adoption is voluntary and the platform gets honest feedback. If it is mandatory but bad, teams route around it and you get shadow infrastructure that nobody reviews.

The failure mode to watch for is a platform team measured by tickets closed. That measures the platform as a service desk. The useful measures are adoption of the paved road, lead time from commit to production, and how many teams can deploy without asking anyone.

Curated: · Written: · Reviewed:

QA-22Explain why normalizing to third normal form can hurt, and how you decide to denormalize.(show answer)

Normalization optimizes for write correctness: each fact stored once, so no update can leave two copies disagreeing. It costs read work, because reassembling an entity means joins.

It hurts when a hot read path needs many joins to produce one screen, when a query aggregates across a deep join tree, or when the join keys are large and the tables live on different shards — at which point the join stops being a query-planner problem and becomes an application one.

The decision to denormalize should be driven by a measured query, not by anticipation. Concretely: identify the specific query, confirm it is actually hot, confirm indexing and query structure are exhausted first, and only then duplicate data.

When you do, be explicit about how the copy stays correct. That is the whole cost of denormalization and it is where the bugs live. The options are a trigger, an application-level write path that updates both, a materialized view refreshed on a schedule, or an asynchronous projection built from a change stream. Each has a different staleness window, and you should be able to state it.

A useful middle ground is a covering index or a materialized view, which gives read performance without inventing a second source of truth the application must maintain.

The rule I would state in an interview: normalize until a measurement tells you not to, denormalize deliberately with a named consistency mechanism, and never denormalize by accident because a column was convenient to copy.

Curated: · Written: · Reviewed:

QA-23Your service depends on a third-party API that is occasionally slow. How do you keep that from taking down your service?(show answer)

Three mechanisms, layered.

A timeout on every call, always. A call without a timeout inherits the default, which is often minutes or infinite, and a slow dependency then consumes your threads or connections until you have none. The timeout should be derived from the dependency's real latency distribution — a little above p99 — not chosen as a round number.

A circuit breaker. After a threshold of failures or timeouts, stop calling and fail immediately for a cooling-off period, then let a single probe test recovery. This does two things: it stops you from queuing work against something that cannot serve it, and it stops you from adding load to a dependency that is struggling. Without it, retries make an overloaded dependency worse.

A bulkhead. Give the dependency its own bounded connection pool or concurrency limit, so exhausting it cannot starve unrelated request paths. This is what converts "the recommendation API is slow" into a degraded recommendations panel rather than a site outage.

Then decide the fallback deliberately: cached previous result, a default, a degraded response, or a clear error. The fallback is a product decision and it should be made before the incident.

Retries need care in this context — retry only idempotent operations, bound the attempts, use exponential backoff with jitter, and budget retries as a fraction of traffic so they cannot multiply into a self-inflicted denial of service.

Curated: · Written: · Reviewed:

QA-24What is the difference between authentication and authorization in an API, and where do teams usually get authorization wrong?(show answer)

Authentication establishes who is calling. Authorization decides what that caller may do. They fail differently: an authentication bug lets a stranger in, an authorization bug lets a legitimate user reach someone else's data — which is more common and usually more damaging.

The mistake teams make most often is object-level authorization. The endpoint checks that the token is valid and that the user has the invoices scope, then loads the invoice by the ID in the path — without checking that this invoice belongs to this user. Any authenticated customer can read any invoice by changing a number. This is the top item in the OWASP API risk list for a reason: it is invisible in tests written by people who only use their own data.

The structural fix is to make ownership part of the query rather than a separate check. Selecting the invoice with both its ID and the caller's tenant ID means a mismatch returns nothing, and no developer can forget the check because there is no path around it.

Two more recurring errors. Trusting claims in a JWT without verifying the signature and the issuer — a token is a bearer of assertions, not evidence, until verified. And enforcing authorization only in the UI, so hidden buttons are the only thing preventing an action that the API happily performs.

Scopes limit what an application may attempt; they do not answer whether this user may touch this record. Both checks are required.

Curated: · Written: · Reviewed:

QA-25How would you approach migrating a monolith's database to a separate service without downtime?(show answer)

As a sequence of reversible steps, each of which is safe to stop at.

First, stop new coupling: the tables in question get a single owner in code, and all access routes through that module. Nothing else queries them directly. This is the hard part, and it is code archaeology, not infrastructure.

Second, put the module behind an interface that the rest of the monolith calls. Still one process, one database, but the call graph now matches the intended boundary.

Third, dual write. The new service writes to its own store while the monolith continues writing the old tables, and reads still come from the old path. Compare the two continuously and alert on divergence; this is where you learn what the boundary actually is.

Fourth, shift reads gradually — a percentage of traffic, or by tenant — with the old path still available.

Fifth, stop writing the old tables, and only then drop them, after a retention window long enough to recover from a discovery you have not made yet.

The parts people underestimate: foreign keys crossing the new boundary have to become application-level references, and transactions that spanned the boundary have to become sagas with compensating actions or be redesigned to avoid the need. If neither is acceptable for a given pair of tables, that is evidence the boundary is in the wrong place — better learned at step three than after the split.

Curated: · Written: · Reviewed:

QA-26What index would you create for a query filtering on tenant_id and status and sorting by created_at, and why that column order?(show answer)

A composite index on (tenant_id, status, created_at), in that order.

The rule for a B-tree composite index is equality columns first, then the range or sort column. Both tenant_id and status are equality predicates here, so they lead; created_at trails so the index also satisfies the ORDER BY, letting the planner skip a sort entirely.

Column order within the equality group matters for reuse rather than for this query: an index on (tenant_id, status, ...) can serve a query filtering on tenant_id alone, because a B-tree can seek on any leading prefix. The reverse is not true — filtering on status alone cannot use this index efficiently. So the more broadly useful column leads.

Selectivity is the tiebreaker people usually cite, and it is secondary to the prefix rule. Putting a highly selective column first narrows the scan fastest, but only if queries actually filter on it.

If the query also returns a small fixed set of columns, extending the index to cover them turns it into an index-only scan and avoids the heap lookup entirely — worth it for a genuinely hot query, at the cost of a larger index and slower writes.

Every index is paid for on every insert, update, and delete, so the honest answer includes checking whether an existing index already has this prefix. Three overlapping indexes on the same leading column is a common and expensive smell.

Curated: · Written: · Reviewed:

QA-27Describe how you would design the network for a workload that must not have any route to the internet.(show answer)

Start by removing the paths rather than blocking them. No internet gateway attached to the VPC, no NAT gateway, no egress-only gateway. If the route does not exist, no security group misconfiguration can create one.

That immediately raises the real problem: the workload still needs the cloud provider's own services — object storage, secrets, logging, the metadata endpoint. Those are reached over private endpoints (interface endpoints backed by private IPs in your subnets, or gateway endpoints installed as routes), so traffic to the provider's API stays on the provider's network and never traverses the internet.

Endpoint policies are the second control: an interface endpoint can be restricted so it only reaches specific buckets or accounts, which prevents the exfiltration path where an isolated workload writes to an attacker-controlled bucket in another account over a legitimate-looking endpoint.

Software distribution needs a plan, since the usual answer is a package repository on the internet. A mirrored internal registry, or a pipeline that bakes images outside the isolated environment and promotes artifacts inward, keeps the isolation intact.

Access for operators should be through a session manager or bastion service that connects via the provider's control plane rather than an inbound SSH port.

Finally, verify the property rather than assuming it: flow logs plus a periodic check that no route table in the account carries a default route to a gateway. Isolation that is not continuously verified drifts back within a quarter.

Curated: · Written: · Reviewed:

QA-28Two services need the same data with different access patterns. Do you share a database, duplicate the data, or add an API between them?(show answer)

Shared database is the option to argue against first, because it is the one that looks cheapest. Two services writing the same tables have no boundary: a schema change requires coordinating deploys, and neither team can reason about invariants alone. It is acceptable when the two services are really one system with an artificial split, and rarely otherwise.

An API between them is the default. One service owns the data and the invariants; the other asks. The cost is a runtime dependency — a synchronous call in the request path means the caller's availability is now the product of both — and latency that compounds if the call is in a loop.

Duplication via an event stream or change feed is right when the reader's access pattern is genuinely different: a search index, an analytics store, a denormalized read model. The reader keeps its own copy shaped for its queries, and accepts staleness measured in seconds. It removes the runtime dependency and adds an eventual-consistency contract you must state and monitor.

The deciding question is who owns the invariant. If both services must enforce the same rule, they should not both hold the data — one owns it and the other asks. If the second service only reads, and can tolerate a defined staleness window, duplication is usually the better architecture because it decouples availability.

Read replicas of another team's database are a shared database wearing a disguise; the schema coupling is identical.

Curated: · Written: · Reviewed:

QA-29How do you decide the right granularity for a service boundary?(show answer)

Follow the data and the rate of change, not the org chart or an aesthetic sense of size.

The strongest signal is transactional integrity: things that must change atomically belong together. If splitting two entities forces a distributed transaction or a saga with compensating actions, the split is probably wrong — you have taken a database guarantee and rebuilt it in application code, badly.

The second signal is the deploy coupling. If two services are always released together, they are one service with extra network hops, latency, and failure modes. Count how often a change touches both.

The third is different scaling or availability requirements. A component that needs ten times the capacity of its neighbours, or one that must stay up while the rest is down, has a genuine reason to be separate.

Team ownership matters, but as a constraint rather than the criterion: a boundary no single team can own will erode.

The practical advice is to start coarser than the target. Splitting a module that has clean internal boundaries is straightforward; merging two services whose data has diverged is not. Modules within one deployable, with enforced import rules, give most of the design benefit and none of the distributed-systems cost, and they make the eventual split cheap because the seams are already there.

The failure I would call out is splitting by technical layer — a service for the API, one for business logic, one for data access. Every feature then touches all three.

Curated: · Written: · Reviewed:

QA-30What signals would you put on a dashboard for a service you have never operated before, and what would you alert on?(show answer)

Dashboard and alerts are different questions, and conflating them is why teams have three hundred alerts nobody reads.

For the dashboard, start with the four golden signals: latency (as a distribution, p50/p95/p99 — never a mean), traffic, errors split by cause, and saturation of whatever resource is scarcest. Add the service's dependencies with the same four, because most incidents arrive from below. Add a deploy and config-change annotation overlay, since a large fraction of incidents correlate with a change and this makes that visible in one glance.

For alerts, the rule is: alert on symptoms users feel, not on causes. High CPU is not an alert; it is a cause that may or may not matter. Elevated error rate on a user-facing endpoint is an alert. Latency exceeding the SLO threshold is an alert. Error budget burn rate is the most useful formulation, because it scales urgency to actual damage — a fast burn pages, a slow burn creates a ticket.

The exceptions worth paging on that are not yet symptoms are saturation signals with a lead time: disk filling, certificates expiring, a queue backlog growing faster than it drains, replica lag approaching the RPO. Each of those is an outage with a countdown.

Every alert needs an owner and a runbook. If nobody can say what to do when it fires, it is a graph, not an alert.

Curated: · Written: · Reviewed:

QA-31A stakeholder asks for 99.99% availability. What questions do you ask before agreeing?(show answer)

Four questions, in order.

Availability of what, measured how? 99.99% is 4.4 minutes per month, and it means nothing until you define the indicator: which endpoints, measured from where, counting which responses as failures. Measured at the load balancer, a client-side network failure is not your outage. Measured from synthetic probes in three regions, it might be. Teams routinely agree to a number and then discover they disagree about the denominator.

What is the cost of the gap between 99.9% and 99.99%? The difference is 43 minutes versus 4 minutes per month, and it is usually the difference between a single-region deployment with good practices and a multi-region active-active architecture with all the consistency complexity that implies. Often several times the infrastructure spend, plus a permanent tax on delivery speed. Stated that way, many stakeholders discover 99.9% was what they meant.

What are the dependencies' SLAs? Your availability cannot exceed the product of everything in the critical path. Three dependencies at 99.9% each put a ceiling near 99.7% before your own code fails at all. If the number is not achievable, that is a fact to establish now.

What happens at the boundary? An SLO without a consequence is a wish. What does the team stop doing when the error budget is exhausted?

Then write it down as an SLO with an error budget, and treat the budget as the shared decision-making currency between reliability and feature work.

Curated: · Written: · Reviewed:

QA-32Explain what a saga is and when you would prefer it over a distributed transaction.(show answer)

A saga is a sequence of local transactions across services, where each step has a compensating action that semantically undoes it. There is no global lock and no global commit — the system passes through intermediate states that are visible, and consistency is restored by compensation rather than by rollback.

Prefer it over a two-phase commit essentially whenever the participants are separate services. Two-phase commit requires every participant to hold locks until the coordinator decides, which means one slow or partitioned participant blocks the others, and a coordinator failure can leave resources locked indefinitely. Across service and network boundaries that is an availability liability, and most cloud data stores do not offer a distributed transaction manager anyway.

The cost is that compensation is not rollback. Refunding a payment is not the same as never charging it: the customer saw the charge, and an email may already have been sent. So sagas require the domain to tolerate visible intermediate states, and the compensations have to be designed by someone who understands the business consequences.

Practical requirements: every step and every compensation must be idempotent, because retries are certain. State must be persisted so an orchestrator crash can resume. And you need a plan for a compensation that itself fails, which usually means an alert and a human, not infinite retries.

Choreography — services reacting to each other's events — avoids a central coordinator but makes the overall flow implicit and hard to debug. Orchestration keeps the flow in one readable place, which for anything beyond three steps is worth the coupling.

Curated: · Written: · Reviewed:

QA-33How do you handle secrets for an application running across several environments?(show answer)

The property to preserve is that a secret never exists in a form a human copies. That rules out configuration files in the repository, environment variables set by hand, and CI variables pasted by an engineer.

Store secrets in a managed secret store with per-environment paths, and let the workload fetch them at start-up using its own machine identity — the instance role or workload identity discussed earlier. The application never holds a credential that authorizes reading the secret; it holds an identity the platform vouches for.

Access to each path is an authorization decision: the staging workload's identity can read staging paths only. This is what makes a compromised non-production workload survivable.

Rotation must be automatic and must not require a deploy. For databases, dynamic credentials — the secret store issues a short-lived database user per workload — are the strongest version, because a leaked credential expires on its own. Where that is not available, scheduled rotation with a dual-secret window lets the new value propagate before the old is revoked.

Two operational details that matter more than they sound. Audit every read, so you can answer which workload accessed which secret and when. And ensure secrets do not leak into logs, crash dumps, or exception trackers — that path accounts for a large share of real exposures, and it is prevented by redaction at the logging layer rather than by developer discipline.

Encryption in the store is table stakes; access control and rotation are what actually protect you.

Curated: · Written: · Reviewed:

QA-34Your object storage bucket was found to be publicly readable. Walk me through your response and the follow-up.(show answer)

Immediate containment first, investigation second, but do not destroy evidence in the process.

Remove public access — at the account level if the tooling supports blocking public access globally, which also protects buckets you have not audited yet. Preserve the current policy and access logs before changing anything, because they are what tells you the exposure window.

Then determine scope. Access logs answer what was read and by whom; the object inventory answers what could have been read. Those are different questions and both matter. The exposure window runs from whenever the policy changed — check the configuration history — not from when it was discovered.

If personal data was exposed, notification obligations may start at discovery and are measured in hours in some jurisdictions. That determination belongs with legal and privacy, and the engineering job is to give them accurate scope quickly.

Follow-up has to be structural, because "we fixed the bucket" prevents nothing. Enable account-level public access blocking as a hard control. Add a policy-as-code check that fails a pull request proposing a public bucket, so the control lives where the change is made. Add continuous configuration monitoring so drift from a console change is detected in minutes. And default new buckets to encrypted, private, and versioned through a shared module rather than through documentation.

The blameless postmortem question is not who made the bucket public but why the system permitted it silently — that is the finding worth acting on.

Curated: · Written: · Reviewed:

QA-35When would you use a graph database rather than expressing relationships in SQL?(show answer)

When the queries are about paths of unbounded or unknown depth, and traversal is the dominant access pattern.

SQL handles relationships well at fixed depth — a join is a one-hop traversal, and two or three joins are fine. It degrades when the depth is a variable: find everyone connected to this account through any chain of shared devices, or find the shortest route through a dependency graph. Recursive CTEs can express this, but each level of depth is another join, and the planner's estimates become unreliable quickly. Performance falls off a cliff rather than degrading gracefully.

A graph database stores adjacency with the graph, so a traversal can expand from the current vertices without repeatedly expressing each hop as another relational join. Work is driven mainly by the vertices and edges actually explored, although properties, indexes, storage layout, and distributed partitions still affect the cost. That advantage can be decisive for fraud rings, access-path analysis, recommendation over social graphs, and impact analysis in a dependency network.

The honest counterweight: it is a separate system to operate, the query languages are less familiar, and aggregate reporting over a graph store is generally worse than over a relational one. Many teams that adopted one for a single traversal query would have been better served by a recursive CTE and an index.

The decision test I would apply: are the important queries variable-depth traversals, and are there several of them? One fixed three-hop query does not justify a second database.

Curated: · Written: · Reviewed:

QA-36What is the practical difference between a 502, a 503, and a 504 from a load balancer, and what does each tell you to check?(show answer)

They are useful clues, but their exact causes are load-balancer-specific, so start with that product's reason metrics and documentation.

A 502 means the load balancer reached a backend and got a response it could not parse, or the connection was closed unexpectedly. Check for application crashes, a process running out of memory mid-response, a protocol mismatch such as the balancer speaking HTTPS to a plaintext port, or a keep-alive timeout on the backend that is shorter than the balancer's, so the backend closes a connection the balancer is about to reuse. That last one produces intermittent 502s that correlate with nothing obvious.

A 503 commonly means the balancer cannot currently serve the request—for example, an AWS Application Load Balancer target group with no registered usable targets. Check target registration and state, health checks, routing, capacity, and any product-specific generated-response metrics.

A 504 means the gateway timed out connecting to or waiting for an upstream. It does not prove the application accepted the request: depending on the product, the timeout may be during target connection, TLS handshake, or response. Check the balancer's reason telemetry before moving to application latency, downstream dependencies, pool exhaustion, or locks.

The diagnostic value is narrowing the search, not treating a status code as a universal root cause. Correlate the load balancer's reason code and target metrics with application logs before deciding where the failure occurred.

Curated: · Written: · Reviewed:

QA-37How do you keep an infrastructure-as-code repository maintainable once it covers dozens of services?(show answer)

Three things do most of the work: state boundaries, module design, and a review path.

Split state. One state file for everything means every change locks every service, a mistake can destroy unrelated infrastructure, and plan times grow until nobody reads the output. Split by blast radius and change frequency — networking and identity change rarely and are shared, application infrastructure changes daily and is per-service. Cross-boundary references go through published outputs or data sources, not through a shared state file.

Design modules around a decision, not around a resource. A module that wraps one resource with fifty pass-through variables adds indirection and no value. A module that encodes "how this organisation runs a service" — with the logging, tagging, encryption, and network placement already correct — is what makes the paved road real. Version modules and let consumers upgrade deliberately.

Make the pipeline the only path to apply. Plan on every pull request with the output posted for review, apply only from the main branch, and no human credentials that can apply directly. Otherwise state drifts from reality and the repository becomes documentation rather than truth.

Beyond that: enforce policy as code so a non-compliant plan fails review automatically, and run drift detection on a schedule so console changes surface as findings rather than as surprises during an unrelated apply.

Curated: · Written: · Reviewed:

QA-38A customer reports intermittent errors that you cannot reproduce and that do not appear in your dashboards. How do you proceed?(show answer)

Assume the dashboards are aggregating the signal away, because that is usually what has happened. An error rate of 0.2% is invisible next to a 99.8% success line, and a p99 graph hides a p99.9 that one customer experiences constantly.

So first, break down rather than drill in: split error rate and latency by customer or tenant, by region, by client version, by instance, and by availability zone. Intermittent problems that are invisible in aggregate are usually concentrated in one dimension — a single bad host, one zone with a network issue, one client version with a shorter timeout.

Second, get identifiers from the customer. A request ID or a precise timestamp with a time zone turns an unreproducible report into a trace. If requests are not carrying a correlation ID end to end, that is the finding: the gap is in observability, and fixing it is more valuable than this one incident.

Third, look at what aggregation cannot show: distributed traces for slow requests, and structured logs filtered to that tenant. Tail-based sampling that keeps every erroring trace is what makes this possible after the fact.

Common causes for this shape of report: retries hiding a real failure rate, one unhealthy instance passing a shallow health check, connection pool exhaustion at a specific hour, DNS or TLS handshake failures that never reach application logs, and clients timing out earlier than your latency budget assumes.

Curated: · Written: · Reviewed:

QA-39Explain the trade-off between synchronous and asynchronous replication for a database, using numbers.(show answer)

Synchronous replication commits only after at least one replica has acknowledged the required log position. Your RPO for acknowledged commits can be zero, subject to the system's exact acknowledgement policy, and the cost is that each commit waits for remote coordination.

Within a region, providers engineer availability zones for low-latency connectivity, so synchronous cross-zone replication is often affordable and is a common managed multi-AZ configuration. The actual latency must be measured for the chosen regions, zones, and database rather than assumed from a universal number.

Across regions the numbers change decisively: round trips are commonly tens of milliseconds on the same continent and can exceed 100ms intercontinentally. The replication wait is normally paid when the transaction commits, not once for every SQL statement inside that transaction. Workloads that commit frequently still add an inter-region round trip to each transaction, so latency rises and commit throughput per connection falls.

Asynchronous replication commits locally and ships the log afterward, so write latency is unaffected. The cost is a non-zero RPO equal to the replication lag at the moment of failure, and lag is not constant — it grows under write bursts, during index maintenance, and during long-running transactions on the replica.

So the usual production shape is synchronous within a region for durability and automatic failover, asynchronous across regions for disaster recovery. And whichever you choose, replica lag must be monitored, because with synchronous replication a slow replica becomes a write-latency problem, and with asynchronous replication it silently becomes a data-loss problem.

Curated: · Written: · Reviewed:

QA-40How do you design an API so it can evolve without breaking existing clients?(show answer)

Decide first that additive changes are free and everything else is a version. Adding an optional field or a new endpoint breaks nobody, provided clients ignore unknown fields — which is a contract you should state explicitly in the documentation, because clients that validate strictly will break on your first addition.

Breaking changes are: removing or renaming a field, changing a type, tightening validation, changing the meaning of a value, and adding a required request field. Each of these needs a new version.

For versioning, URL path versioning is the most operationally honest choice — it is visible in logs, cacheable, and trivially routable — even though header-based versioning is more theoretically pure. What matters more than the mechanism is that you can measure usage per version, because you cannot retire a version you cannot see.

Deprecation needs a process, not an announcement: a stated support window, a deprecation header on responses, usage metrics per client, and direct contact with the remaining callers before removal. Most failed API migrations fail because nobody knew who was still calling.

Two design choices that buy the most future room: return objects rather than bare values, so fields can be added; and avoid exposing internal enumerations directly, since a new internal state should not force a client change.

Idempotency, consistent error shapes, and pagination from day one also count as evolvability — retrofitting pagination onto an endpoint that returned an array is a breaking change.

Curated: · Written: · Reviewed:

QA-41What does least privilege actually look like for a CI pipeline that deploys infrastructure?(show answer)

It looks like several narrow identities rather than one broad one, separated along the axes where mistakes are expensive.

Separate by stage: the plan step needs read access and nothing else, and it is the step that runs on untrusted pull-request code. The apply step needs write access and runs only from the protected main branch after approval. Giving the plan step write permissions means an attacker who can open a pull request can modify infrastructure.

Separate by environment: the pipeline identity for staging cannot touch production. Federate via OIDC with a trust policy that pins the repository and the branch or environment, so a workflow on a fork or a feature branch cannot assume the production role even with a valid token.

Scope the permissions themselves to the resource types the pipeline manages, with conditions on tags or naming prefixes where the provider supports it. Administrator access assigned because scoping was tedious is the normal state of affairs and the normal cause of incidents.

Add guardrails that do not depend on the pipeline being correct: service control policies or organisation-level rules that deny destructive actions on protected resources regardless of the caller, and deletion protection on stateful resources.

Finally, treat the pipeline as production infrastructure. It has credentials to your entire estate, so its own configuration deserves code review, its logs should be immutable, and third-party actions should be pinned to a commit hash rather than a moving tag.

Curated: · Written: · Reviewed:

QA-42Your team wants to add a read-through cache in front of a database. What do you ask them to specify before agreeing?(show answer)

Five things, and if they cannot answer them the cache will cause more incidents than it prevents.

What is the staleness budget? Every cache is a decision to serve old data. The answer is a duration, per data type, agreed with whoever owns the product consequence — not a TTL chosen because it sounded reasonable.

What is the invalidation mechanism? TTL-only expiry is fine for tolerant data. Anything a user edits and immediately re-reads needs explicit invalidation on write, and the write path has to be identified. The hardest cases are writes that happen outside the application — a batch job, an admin tool, a database migration.

What happens when the cache is empty or unavailable? If the origin cannot serve the full load, the cache is a single point of failure wearing a performance disguise. This should be a measured number, not an opinion.

What is the key design, and can it collide? Multi-tenant caches with a key that omits the tenant leak data across customers, and this is a real and recurring incident class.

What will you measure? Hit rate alone is not enough — you want hit rate, origin load, and staleness-related error reports, so you can tell a healthy cache from one that is quietly serving wrong data.

If the underlying problem is a missing index or an N+1 query, the cache is hiding a fixable cause behind a permanent operational commitment.

Curated: · Written: · Reviewed:

QA-43How would you approach capacity planning for a launch you have never seen traffic for?(show answer)

Convert unknown demand into a tested capability, then buy margin.

Start with what you can bound: the marketing reach, the size of the user base being notified, historical conversion for similar launches, and the shape of arrival. Peak matters more than total — a campaign email sent at once produces a spike far above the daily average, often 10-20x, in the first ten minutes.

Then load test to find where the system actually breaks, not to confirm it survives an estimate. Ramp until something fails, note which resource saturated first, fix or provision it, repeat. A test that passes at the expected load teaches you nothing about the margin you have. Use representative data volumes; a load test against an empty database measures the wrong system.

Pre-provision rather than relying on autoscaling for the spike, because provisioning lag is longer than the spike's rise time. Warm caches and connection pools before opening the gate.

Then design the overload behavior, because capacity planning that assumes the estimate is right is a wish. Rate limiting at the edge, a queue with admission control, and a graceful degradation path — read-only mode, a static waiting page — turn an underestimate into a slow experience rather than an outage.

Finally, agree the rollback: a feature flag or a staged regional rollout means you can reduce demand, which is the only lever that works when capacity cannot be added fast enough.

Curated: · Written: · Reviewed:

QA-44What is the risk of putting business logic in a database trigger, and when is it still the right choice?(show answer)

The risks are invisibility and testability. A trigger executes on a write that a developer reading application code cannot see, so behavior appears without a call site. It runs inside the writing transaction, so it extends lock duration and can deadlock in ways that are hard to attribute. It does not participate in code review the way application code does unless migrations are reviewed with equal care, and it is awkward to unit test.

Triggers also interact badly with bulk operations: a trigger written for single-row inserts can turn a million-row load into a million procedure calls, and a trigger that writes to another table can cause a bulk operation to deadlock against normal traffic.

They are still the right choice when the guarantee must hold regardless of the writer. Audit trails are the clearest example: if the requirement is that no row can change without an audit record, application code cannot provide it, because a migration, an admin console, or a manual fix bypasses the application. The same reasoning applies to enforcing a temporal invariant, or maintaining a denormalized counter that must never disagree with its source.

The deciding question is whether the rule is a property of the data or a policy of the application. Data properties — this must always be true of any row, however it got there — belong in the database, alongside constraints and foreign keys. Application policy belongs in application code where it is visible, testable, and versioned with the feature it serves.

Curated: · Written: · Reviewed:

QA-45Describe how you would run a design review when two senior engineers disagree about the architecture.(show answer)

Move the discussion from preference to evidence, and make the decision explicit rather than letting it dissolve.

First, establish what is actually being decided. Disagreements at this level are usually about different questions — one person optimizing for delivery speed, the other for operational cost in year two — and once that is visible, much of the heat goes out of it.

Second, write down the constraints that are not negotiable: the latency budget, the compliance requirement, the team's operational capacity, the deadline. Options that violate a constraint leave the discussion; this is often most of the disagreement.

Third, compare the remaining options on the same criteria, and require each side to state what evidence would change their mind. If neither can answer that, the disagreement is about values, not facts, and needs a decision rather than more discussion.

Fourth, where the disagreement is factual and cheap to resolve, resolve it: a prototype, a load test, a spike with a time box. Two days of measurement beats two weeks of argument.

Then decide, name the decider, and record it — the options considered, the criteria, the choice, and the conditions that would cause a revisit. An architecture decision record does most of this work, and the reversal condition matters most, because it converts a contested decision into a testable one.

The failure mode to avoid is deciding by seniority or by attrition. Both leave the losing side uncommitted, and the design gets undermined in implementation.

Curated: · Written: · Reviewed:

QA-46How do you prevent a retry storm from turning a partial outage into a total one?(show answer)

Retries multiply load exactly when a system is least able to absorb it, and every layer that retries multiplies the layers below it. If the client, gateway, and service library each make three total attempts, one logical request can become twenty-seven dependency attempts; three simultaneous client requests can become eighty-one.

The controls, in order of importance.

Retry at one layer only. Decide where retries live and make every other layer fail fast. This single decision removes the multiplication.

Bound the attempts and use exponential backoff with full jitter. Backoff without jitter synchronises clients into waves that arrive together; jitter spreads them. The number of attempts should be small — two or three — because a request that has failed three times is unlikely to succeed on the fourth and is certainly adding load.

Budget retries as a fraction of total traffic, typically 10%. When the budget is exhausted, stop retrying entirely. This is what prevents a widespread failure from producing a self-inflicted denial of service, because a global failure means every request retries at once.

Circuit-break so that once a dependency is clearly down, calls fail immediately rather than waiting for a timeout and then retrying.

On the serving side, shed load explicitly: reject early with a 429 and a Retry-After header rather than queuing work you cannot complete. A fast rejection costs almost nothing; a slow timeout holds a connection for seconds.

Curated: · Written: · Reviewed:

QA-47What is the practical difference between an SLA, an SLO, and an SLI, and why does the distinction matter to an architect?(show answer)

An SLI is a measurement — the proportion of requests served successfully under 300ms, for example. An SLO is a target for that measurement over a window, such as 99.9% over 28 days. An SLA is a contract with a customer, with financial consequences, and it should be looser than the SLO so that missing internal targets does not immediately mean paying out.

The distinction matters architecturally for three reasons.

The SLI definition determines what you build. Measuring at the load balancer and measuring from the client produce different numbers and justify different work — one leads to backend investment, the other to CDN and network investment. Arguments about reliability are often really arguments about the indicator.

The SLO sets the architecture. 99.9% and 99.99% are different systems, as discussed, and the SLO is the input that makes multi-region a requirement rather than an aspiration. Agreeing an SLO before designing prevents building the wrong thing in both directions.

The error budget makes the trade-off explicit. The budget — the allowed 0.1% — is a resource that both feature work and reliability work spend. When it is healthy, ship faster. When it is exhausted, reliability work takes priority. That converts an endless argument between product and engineering into a rule that both agreed to in advance, which is the actual value of the whole framework.

Curated: · Written: · Reviewed:

QA-48You inherit a system with no documentation and an on-call rotation that pages nightly. What do you do in the first month?(show answer)

Stabilise the humans first, then the system, then the design.

Week one: instrument the pages themselves. Every alert gets categorised — did it require action, was it a real user-facing problem, what did the responder do. Most nightly paging is a small number of alerts firing repeatedly, and you cannot fix what you have not counted.

Week two: eliminate the noise. Alerts that never required action get deleted or converted to tickets, not tuned. Alerts on causes rather than symptoms get replaced. Duplicate alerts for one condition get merged. This is usually a large reduction and it buys the team the capacity to do everything else.

Weeks two and three, in parallel: write the runbook as you go. The responder documents what they did while they are doing it. This is the only version of documentation that reliably gets written, and it is the most useful kind.

Weeks three and four: address the top recurring cause. There is almost always one — a disk filling, a leak requiring restarts, a dependency timing out — and fixing it removes a large share of remaining pages.

Only then look at architecture. Draw the actual dependency graph from traces and network flow, not from a diagram someone drew two years ago, and identify the single points of failure.

Resist the urge to rewrite. The system is loud, not necessarily wrong, and you do not yet know which parts encode requirements you have not discovered.

Curated: · Written: · Reviewed:

QA-49How do you decide whether to use a managed service or run something yourself?(show answer)

Compare total cost of ownership honestly, and be specific about what the managed service is actually taking off your plate.

The managed service usually wins on: patching, backup and restore that is tested, failover automation, and the availability engineering that takes a team a year to get right. Those are the parts self-hosting teams consistently underestimate, because the work is invisible until the night it is needed.

Self-hosting wins when you need a version, extension, or configuration the managed service does not expose; when the price at your scale genuinely exceeds the loaded cost of the engineers who would run it; when data residency or regulatory constraints rule it out; or when you already have deep operational expertise in exactly that system.

The costs of managed services that people miss: version upgrades happen on the provider's schedule, some configuration is simply unavailable, debugging is limited to what the provider exposes, and the exit cost grows with adoption. That last one is the strategic consideration — the more provider-specific the service, the more expensive a future move, and that should be a deliberate trade rather than a default.

My general position: use managed services for stateful infrastructure, where the operational burden is highest and the failure modes are worst, and be more willing to self-host stateless components where the operational load is light and portability is cheap.

The question that settles most cases: if this breaks at 3am, who is qualified to fix it, and do we want that to be us?

Curated: · Written: · Reviewed:

QA-50What is connection pooling, and why does it become a problem specifically in serverless architectures?(show answer)

A connection pool amortises the cost of establishing database connections — TCP handshake, TLS negotiation, authentication, session setup — which is expensive relative to a query. A long-lived application process opens a handful of connections and reuses them for thousands of requests.

Serverless breaks the assumption the pattern depends on. Each function instance is its own process with its own pool, instances scale with concurrency rather than with your provisioning decisions, and they are created and destroyed constantly. A thousand concurrent invocations means up to a thousand pools, and a database sized for two hundred connections refuses them.

The failure is abrupt: the database hits max connections and rejects everything, including the healthy long-running services sharing it. It also arrives without warning, because concurrency scaled with traffic and nothing in the function's configuration mentions the database's limit.

The fixes, in order of preference. Put an external connection pooler between the functions and the database — a proxy that maintains a small pool to the database and multiplexes many client connections onto it. This is the direct answer and the reason such proxies exist.

Failing that, use a data API that speaks HTTP rather than the database wire protocol, so there is no persistent connection at all.

Configure reserved concurrency as a hard ceiling so the function cannot scale past what the database can serve, and keep the per-instance pool at one connection, since a single invocation does not run parallel queries.

Curated: · Written: · Reviewed:

QA-51Explain what an availability zone actually is, and what failure it does and does not protect against.(show answer)

An availability zone is one or more physically separate data centres within a region, with independent power, cooling, and physical security, connected to the other zones by provider-engineered low-latency links. That low latency makes synchronous cross-zone replication practical, but the actual bound depends on the provider, region, and service and should be measured.

It protects against a data centre-level failure: a power event, a cooling failure, a fire, a flood, a network device failure confined to one facility. Spreading instances across three zones means one facility's loss removes a third of your capacity rather than all of it.

It does not protect against every regional failure. Regional control-plane or shared-service outages—an API that stops accepting calls, an identity degradation, or a DNS problem—can affect multiple zones together, so zone redundancy is not the same as regional disaster recovery.

It also does not protect against your own mistakes. A bad deploy, a schema migration that corrupts data, a deleted bucket, or an over-broad IAM change propagates to every zone immediately. Nor does it help if the workload is multi-zone but its database has a single writer in one zone.

So: zones are for infrastructure failure, regions are for regional events, and backups and change control are for the failure class that replication faithfully copies.

Curated: · Written: · Reviewed:

QA-52How do you evaluate whether a system is ready for a compliance audit?(show answer)

Audits ask for evidence, not for assurances, so readiness means the evidence exists as a byproduct of how you operate rather than being assembled the week before.

Four categories cover most frameworks. Access: who can reach what, how access was granted, how it is reviewed, and how it is removed when someone leaves. The hard part is usually the removal path and the periodic review, not the granting.

Change: every production change traceable to an approved, reviewed request. A pipeline that requires a pull request and records the approver produces this automatically; a process where engineers can apply from a laptop does not, at any level of discipline.

Data: what personal or regulated data exists, where it lives, how it is encrypted at rest and in transit, how long it is retained, and how deletion is performed and proven. Data mapping is the most common gap, because nobody maintains it unless it is enforced at the point where a new store is created.

Monitoring: log retention that meets the required window, logs that are tamper-evident, and evidence that alerts were acted on.

The readiness test I would apply is a dry run: pick a change from three months ago and produce the approval, the diff, who deployed it, and the access list for the system at that time. If that takes days, the controls exist on paper only.

Curated: · Written: · Reviewed:

QA-53A batch job holds locks that block user traffic every night. How do you fix it?(show answer)

Reduce the lock's scope and duration rather than moving the job an hour later, which only relocates the collision.

First, understand which locks. A batch UPDATE over millions of rows holds row locks for the whole transaction and may escalate; a schema change takes an exclusive table lock; a long-running read in some engines blocks vacuum or holds a snapshot that bloats the table. These need different fixes.

The general fix is batching: process in chunks of a few thousand rows, commit between chunks, and pause briefly. This turns one long transaction into many short ones, so user traffic interleaves instead of queuing. Chunk by primary key range rather than with OFFSET, which re-scans.

Order matters too. If the batch updates rows in a different order from the application, you get deadlocks rather than waits. Making both paths acquire in the same order — usually primary key order — removes them.

For read-heavy analytical work, move it off the primary entirely: a read replica or an analytics store removes the contention rather than managing it.

Where the work must touch the same rows as user traffic, use a lock timeout so the batch yields rather than blocks — the job takes longer and users do not notice. And add a monitor on lock wait time by query, because the reason this was discovered by user complaints rather than by a graph is that nobody was watching the right signal.

Curated: · Written: · Reviewed:

QA-54What would make you choose event-driven architecture over synchronous request/response between services?(show answer)

Choose events when the producer does not need to know who consumes, when the consumer's availability should not affect the producer's, and when the work is genuinely asynchronous from the user's point of view.

Order placed is the canonical example: the order service publishes, and inventory, fulfilment, analytics, and notifications each react. Adding a fifth consumer requires no change to the producer, which is the real architectural benefit — the coupling removed is coupling to the list of consumers.

It also decouples availability. A synchronous call chain has availability equal to the product of every link; an event consumer that is down accumulates a backlog and catches up.

The costs are substantial and worth stating plainly. Debugging becomes harder: there is no stack trace across an event boundary, so distributed tracing with correlation IDs stops being optional. Ordering and duplicate delivery become the consumer's problem. The overall behavior of the system is not written down anywhere — you cannot read one file and know what happens when an order is placed. And eventual consistency becomes user-visible, so the UI has to represent in-progress states honestly.

I would keep synchronous calls where the caller needs the result to proceed and the user is waiting, and use events where the reaction is a side effect. Mixing them deliberately along that line gives most of the benefit; making everything an event because events are decoupled produces a system nobody can reason about.

Curated: · Written: · Reviewed:

QA-55How do you handle schema changes for a service deployed continuously with no maintenance window?(show answer)

With the expand-migrate-contract discipline, and by never deploying a schema change and the code that requires it at the same moment.

Expand: make the schema accept both the old and new shapes. Add the new column as nullable, add the new table, add the new index concurrently. Old code continues to work untouched.

Deploy code that writes both shapes and reads the old one. This is the reversible point — rolling back the application is safe because the old shape is still being maintained.

Migrate the existing data in batches, then switch reads to the new shape, still writing both.

Contract: stop writing the old shape, then remove it, after enough time has passed that a rollback would not need it. Dropping a column the previous release still writes to is how a routine rollback becomes an incident.

Two operational rules make this workable. Build indexes concurrently, so index creation does not hold a write lock. And set a short lock timeout on migrations: a migration waiting behind a long-running query will queue every subsequent query behind itself, converting a fast schema change into an outage. Failing fast and retrying is strictly better.

The version-skew constraint underneath all of it: during any rolling deploy, two versions of the code run against one schema. Every change must be compatible with the version before and after it, which is exactly what expand-migrate-contract guarantees.

Curated: · Written: · Reviewed:

QA-56What are the main ways a cloud bill grows without anyone deploying anything new?(show answer)

Accumulation and traffic shape, mostly.

Storage only grows. Log retention set to never expire, snapshots taken daily and never pruned, orphaned volumes left behind by terminated instances, incomplete multipart uploads that are invisible in the console but billed, and old object versions retained by versioning without a lifecycle rule. Each is small monthly and unbounded over years.

Data transfer scales with usage rather than with deployments. Cross-zone traffic between chatty services, egress to the internet as your user base grows, NAT gateway processing charges on traffic that could have used a private endpoint, and cross-region replication all rise silently.

Autoscaling responds to traffic, so organic growth raises the bill without a deploy. That is expected; the problem is a maximum raised during an incident and never lowered, which turns a temporary ceiling into a permanent floor.

Idle and forgotten resources: non-production environments running at night and at weekends, load balancers with no targets, provisioned capacity for a launch that already happened, and elastic IPs that are billed when unattached.

Committed-use discounts expiring, or reserved capacity lapsing, produces a step change with no infrastructure cause at all.

The systemic fix is lifecycle rules and expiry as defaults rather than as cleanup: retention configured at creation, non-production shut down on a schedule, and a monthly review of the largest line items by tag. Cleanup projects work once; defaults work continuously.

Curated: · Written: · Reviewed:

QA-57Explain what happens under the hood when you add an index, and why it is not free.(show answer)

A B-tree index is a separate on-disk structure holding the indexed column values in sorted order with pointers to the rows. Building it requires reading the whole table and sorting, which on a large table is expensive and, unless built concurrently, holds a lock that blocks writes for the duration.

Once built, it is maintained on every write. An insert adds an entry to every index on the table. An update to an indexed column removes and re-adds an entry, and in some engines an update to any column writes a new row version that every index must point to. A delete marks entries dead in every index. So five indexes mean roughly five times the index maintenance work per write, plus the write-ahead log volume for each.

It also costs memory, because index pages compete with table pages for the buffer cache. An index that does not fit in memory and is not used often can evict pages that are, which slows unrelated queries.

And it costs storage — often more than people expect, since an index on a wide text column can approach the size of the table.

The consequences an architect should draw: index for the queries you actually run, prefer one composite index over three single-column ones where the leading prefix serves both, and periodically drop unused indexes — every database exposes index usage statistics, and unused indexes are pure cost.

Curated: · Written: · Reviewed:

QA-58How do you design a rate limiter for a public API, and where do you put it?(show answer)

Put it at the edge, before requests consume application resources, and make it distributed so a client cannot get n times the limit by hitting n instances.

The algorithm choice matters. A fixed window is simple and allows a client to send two full windows' worth of traffic across a window boundary. A sliding window log is exact but stores a timestamp per request. A token bucket is usually the right answer: it enforces an average rate while allowing a bounded burst, which matches how legitimate clients behave, and it is cheap to implement as a counter plus a timestamp in a shared store.

Decide what you limit by. Per API key is the norm for authenticated traffic; per IP is necessary for unauthenticated endpoints but is crude behind NAT and proxies. Login and password-reset endpoints deserve stricter, separate limits than read endpoints, because the risk there is credential stuffing rather than capacity.

The response contract matters as much as the algorithm: 429 with a Retry-After header, plus headers exposing the limit, remaining, and reset time, so a well-behaved client can pace itself rather than guess. Without those, clients poll and make the problem worse.

Two further considerations: a shared counter store is now in your request path, so decide its failure behavior deliberately — fail open for availability or fail closed for protection. And rate limiting is not abuse prevention on its own; it bounds volume, not intent.

Curated: · Written: · Reviewed:

QA-59Your architecture diagram and your production environment have diverged. Why does that matter, and what do you do about it?(show answer)

It matters because every decision made from the diagram is made from a fiction. Incident response is slower when the dependency you are chasing is not on the map; risk assessments miss the single point of failure that was added last quarter; and new engineers form a mental model that is wrong in exactly the places where being wrong is expensive.

The reason it happens is that a diagram is a document, and documents decay unless something maintains them.

The fix is to derive as much as possible rather than draw it. Service dependencies come from distributed traces, which reflect the calls that actually happen. Network topology comes from the infrastructure code, which is the source of truth if applies only happen through the pipeline. Data stores and their owners come from a service catalogue that is populated at deploy time rather than by hand.

What is left after that is the part worth drawing by hand: the intent. Why this boundary exists, what the trust zones are, which failure the redundancy is protecting against. That changes slowly, and it is the part a generated diagram cannot express.

Then close the loop: drift detection between the infrastructure code and the live environment, so console changes surface within a day, and a rule that the architecture record is updated in the same pull request as the change. A diagram reviewed alongside the code stays current; a diagram in a wiki does not.

Curated: · Written: · Reviewed:

QA-60What is the difference between horizontal and vertical scaling in practice, and why do people reach for vertical first?(show answer)

Vertical scaling makes one machine bigger. Horizontal scaling adds machines. People reach for vertical first because it requires no application change: doubling the instance size is a configuration edit and, for a database, often the only option that does not involve redesign.

Vertical scaling is genuinely the right first move for stateful systems. A relational primary can usually be scaled up substantially before sharding is warranted, and a bigger machine is far cheaper than the engineering cost of splitting data. The limits are that machines have a maximum size, the price curve is superlinear at the top end, and a single machine remains a single point of failure regardless of how large it is. Resizing also usually requires a restart, so it is not a response to a live incident.

Horizontal scaling extends capacity beyond one machine and can improve availability, but it still has ceilings in coordination, partitioning, network, and downstream systems. It also requires the workload to be parallelisable. For stateless application servers that is comparatively easy. For anything holding state — sessions, in-memory caches, a database — it requires designing where the state lives, which is the actual work.

The practical sequence is: make the application layer stateless and scale it horizontally, scale the data layer vertically for as long as that works, and only shard when write throughput or dataset size genuinely exceeds one machine. Sharding early costs you joins, transactions, and operational simplicity in exchange for capacity you are not using.

Curated: · Written: · Reviewed:

QA-61How would you detect and respond to a compromised set of credentials in a cloud account?(show answer)

Detection first, because most compromises are found by their behavior rather than by an alert on the credential itself.

The signals that matter: API calls from an unfamiliar location or ASN, calls to services this identity has never used, reconnaissance patterns such as enumerating identities, buckets, or snapshots, disabling logging or deleting log trails, creating new access keys or identities, and unusual data egress volume. A threat detection service will flag several of these; the ones worth building yourself are the identity-behavior anomalies specific to your estate.

Response, in order. Revoke the session, not just the key — deleting an access key does not invalidate an already-issued temporary session token, and attaching an explicit deny policy or revoking sessions issued before a timestamp is what actually stops activity in progress.

Then preserve evidence before remediating: snapshot the audit log range and any affected instances. Remediation destroys the timeline you will need.

Then scope it: every action that identity took, every resource it touched, and every credential it could have obtained. Assume anything it could read, it read.

Then rotate everything within reach, including credentials stored on resources the identity could access.

The follow-up that prevents recurrence is structural: eliminate the long-lived credential class entirely in favour of federated short-lived identities, require MFA for human access, and alert on credential creation — because the attacker's first move after entry is almost always to establish a second way in.

Curated: · Written: · Reviewed:

QA-62When is it worth introducing a CDN, and what does it not solve?(show answer)

It is worth it as soon as you have geographically distributed users and any cacheable content, which is nearly every public product. The gain is not only bandwidth: terminating TLS close to the user removes handshake round trips, which is often a larger share of perceived latency than the response itself for small requests.

It solves: static asset delivery, image and video distribution, absorbing traffic spikes for cacheable content, and a meaningful share of volumetric attack traffic at the edge.

It does not solve dynamic, personalised responses, which is where the interesting latency usually lives. A per-user dashboard cannot be cached at the edge without careful design — cache keys that include the user, or edge computation, both of which are real work rather than a configuration toggle.

It does not fix a slow origin. Cache misses go to the origin, and a cold cache after an invalidation sends the full load there, so origin capacity still has to be sized for the miss case.

It does not fix write latency at all. Writes go to the origin regardless of edge presence.

And it introduces its own failure modes worth planning for: a bad cache key that mixes users' responses, an invalidation that propagates slowly across hundreds of locations, and a cached error page that outlives the error. Caching a 500 for an hour turns a one-minute incident into a long one, which is why negative caching needs a much shorter TTL.

Curated: · Written: · Reviewed:

QA-63How do you approach right-sizing a fleet without risking performance?(show answer)

Measure over a period long enough to include the peak, change one dimension at a time, and keep the rollback cheap.

Start from utilization percentiles rather than averages. A fleet averaging 15% CPU with a p99 of 85% is not oversized; it is correctly sized for its peak. The interesting cases are those where both the average and the peak are low, which means the instance was chosen for a requirement that is not CPU.

Identify the real constraint before resizing. Many workloads are memory-bound, IO-bound, or limited by connection counts, and shrinking CPU on a memory-bound service saves nothing and risks a lot. Burstable instance families add a further trap: they look comfortable until the credit balance depletes, at which point throughput collapses, so credit balance is the metric that matters there rather than CPU.

Then change in one direction at a time, in a canary or a fraction of the fleet, and hold it through a full traffic cycle including the weekly peak. A week of comfortable behavior on Tuesday says nothing about Monday morning.

Keep headroom deliberately: enough to absorb the loss of one availability zone, plus the provisioning lag of your autoscaling. Right-sizing to the peak leaves nothing for failure.

The cheapest wins usually are not instance sizes at all — they are non-production environments running out of hours, and old-generation instance families that cost more than their newer, faster equivalents.

Curated: · Written: · Reviewed:

QA-64What does idempotency mean for infrastructure code, and where does it commonly break?(show answer)

It means applying the same configuration twice produces the same result as applying it once. Declarative tools give this by construction: the tool compares desired state to actual state and acts only on the difference.

It breaks in several recognisable places.

Provisioners and local scripts embedded in the configuration. A shell script that appends a line to a file runs again on the next apply and appends it twice. Anything imperative inside a declarative tool is where idempotency dies.

Resources with generated names. A resource named with a timestamp or a random value that is regenerated each run produces a new resource on every apply instead of matching the existing one.

State drift from out-of-band changes. Someone edits a resource in the console; the tool now sees a difference and reverts it — which is idempotent with respect to the configuration but surprising to whoever made the change, and destructive if their change was an emergency fix.

Resources the tool cannot fully read back. Where a provider cannot inspect an attribute, it either always shows a diff or never shows one, and both are wrong in different ways.

And ordering assumptions: implicit dependencies that happened to work in the order the tool chose last time and do not this time.

The practical defences are reading every plan rather than skimming it, treating a plan showing changes on an unchanged configuration as a bug to investigate, and keeping imperative steps in a pipeline stage where they can be made idempotent explicitly.

Curated: · Written: · Reviewed:

QA-65How do you decide what to log, given that logging everything is expensive and logging nothing is worse?(show answer)

Decide by asking what question each log line will answer during an incident. A line that answers nothing is cost.

The tiers I would set. Always log, at low volume: every request with a correlation ID, method, path, status, duration, and the authenticated identity. That single structured line answers most operational questions and its volume is predictable. Log every error with enough context to reproduce — inputs, identifiers, the failing dependency — not just a stack trace.

Log every security-relevant event unconditionally and retain it longest: authentication outcomes, authorization denials, permission changes, and access to sensitive data. These are the ones an audit or an investigation needs, and they cannot be reconstructed later.

Sample the high-volume, low-information categories: successful health checks, per-item debug lines in a batch loop, verbose third-party client output. Head sampling for traces with tail-based retention of everything that errored or was slow gives you the interesting tail without the volume.

Make level changeable at runtime, so debug logging can be enabled for one service during an incident without a deploy.

Two hard rules regardless of volume. No secrets, tokens, or personal data in logs — enforced by redaction in the logging layer, since developer discipline fails eventually. And structured output, because grep across terabytes is not an incident response strategy; fields you can filter and aggregate are.

Then set retention by category rather than globally: 7-30 days hot for operational logs, a year or more in cheap storage for audit.

Curated: · Written: · Reviewed:

QA-66Explain the transactional outbox pattern and the problem it solves.(show answer)

It solves the dual-write problem: a handler that must both commit a database change and publish a message has no way to make those two operations atomic, because they are different systems. Commit then publish loses the message if the process dies in between. Publish then commit emits an event for something that never happened.

The outbox removes the second system from the critical path. In the same transaction that writes the order, insert a row into an outbox table containing the message. Either both are committed or neither is — one transaction, one database.

A separate relay then reads unpublished outbox rows and publishes them, marking them sent afterward. If the relay crashes between publishing and marking, it republishes on restart, which is why consumers must be idempotent — but at-least-once is exactly the guarantee message brokers give anyway, so this adds no new burden.

The relay can be a polling loop or, better, a change data capture stream reading the database's replication log, which avoids polling load and reduces latency to milliseconds.

The costs are an extra table, a component to operate, and a small publication delay. What you get is that the database is the single source of truth for both state and the events derived from it, which means no event describes a state that does not exist and no state change silently fails to notify anyone.

The mirror-image pattern on the consumer side is the inbox: record processed message IDs transactionally to make consumption idempotent.

Curated: · Written: · Reviewed:

QA-67What questions do you ask before agreeing to a multi-cloud architecture?(show answer)

What problem is it solving, specifically? The stated reasons are usually availability, negotiating leverage, or avoiding lock-in, and each deserves a different response.

If the reason is availability: is a provider-wide outage actually your risk? Most large outages are regional or service-specific, and multi-region within one provider addresses those at a fraction of the complexity. Multi-cloud active-active means your data layer must be consistent across providers, which is the hardest problem in the design and the one that most such projects fail on.

If the reason is negotiating leverage: is the discount larger than the engineering cost? Running two providers means two sets of expertise, two security models, two IAM systems, two monitoring stacks, and a permanent constraint to the intersection of both feature sets — which is roughly the capability of neither.

If the reason is avoiding lock-in: what is the actual exit scenario, and what would it cost to migrate if you did not prepare? Often the honest answer is that a focused six-month migration is cheaper than a permanent abstraction tax, and that the strongest lock-in is not compute anyway, it is data gravity and managed data services.

What I would usually recommend instead: portability where it is cheap — containers, standard protocols, infrastructure code with a clean module boundary — and provider-specific services where they earn their keep. Then, if a specific regulatory or contractual requirement demands a second provider, scope it to the workload that requires it rather than to everything.

Curated: · Written: · Reviewed:

QA-68How do you keep a shared Kubernetes cluster safe for multiple teams?(show answer)

Establish boundaries at four layers, because a namespace on its own is only a naming convention.

Identity and access: each team gets namespaces with role bindings scoped to them, and nobody gets cluster-admin as a matter of course. Workload identity maps pods to cloud identities per service account, so a pod's cloud permissions are as narrow as its Kubernetes ones.

Resource isolation: resource quotas per namespace so one team cannot consume the cluster, and limit ranges so a pod without explicit requests does not default to unbounded. Requests and limits set properly are also what keeps the scheduler's decisions meaningful.

Network: default-deny network policy, with explicit allows. Without a policy, every pod can reach every other pod and every service in the cluster, which makes one compromised container a lateral movement platform.

Admission control: policy that rejects privileged containers, host mounts, host networking, and images from unapproved registries, plus a requirement for non-root and a read-only root filesystem. This is where the actual container escape risks are addressed, and it must be enforced at admission rather than documented.

Beyond those, node isolation matters for genuinely untrusted workloads — separate node pools, or a sandboxed runtime — because kernel-level isolation between containers is weaker than between virtual machines.

And keep the control plane's audit log, because in a shared cluster the question "who changed this" comes up weekly.

Curated: · Written: · Reviewed:

QA-69Describe how you would test a disaster recovery plan.(show answer)

By executing it, on a schedule, with the results measured against the stated objectives. A plan that has never been run has an unknown RTO, and the gap between the documented procedure and reality is discovered only under load.

Build up in stages. Start with a tabletop walkthrough: the team talks through the steps, which finds missing documentation and unclear ownership cheaply. Then a component test: fail over the database in a non-production environment and time it. Then a full regional failover in production, or in an environment that is a faithful copy, with the clock running.

Measure the actual RTO and RPO rather than the intended ones, and record where the time went. The delays are rarely in the technology — they are in deciding to declare a disaster, in finding the person with the necessary access, in a runbook step that references a system that was renamed, and in DNS TTLs that are longer than anyone remembered.

Test the restore, not the backup. A backup that has never been restored is an assumption. Restore to a fresh environment and verify the data, including that the application actually starts against it.

Test the failure of your recovery mechanism too: what happens if the standby region is where the incident started, or if the credentials needed to fail over are stored in the system that is down.

Then fix the findings and schedule the next one. Quarterly is a reasonable cadence for a serious objective, and the trend in measured RTO is the real indicator of readiness.

Curated: · Written: · Reviewed:

QA-70What is a materialized view, and when does it beat a cache?(show answer)

A materialized view stores the computed result of a query as a real table that the database maintains or refreshes. A cache stores a computed result in a separate system that the application maintains.

The materialized view wins when the computation is expressible in SQL and the consumers are already talking to the database. It keeps one source of truth, it can be indexed, and it participates in the database's own backup story. The application avoids cache-key invalidation logic, but the system still needs a refresh policy and monitoring for refresh failures. For a reporting aggregate refreshed every fifteen minutes, it is often the simpler tool.

The cache wins when the value is not a query result — a rendered fragment, a third-party response, a computed object — when the consumers are distributed and should not touch the database at all, or when you need a sub-millisecond read path and very high throughput that you do not want landing on the database.

The important detail with materialized views is refresh semantics. A plain refresh takes an exclusive lock and blocks readers for its duration; a concurrent refresh avoids that but requires a unique index and does more work. Neither is incremental in most engines, so refreshing a view over a very large table repeatedly can cost more than the queries it replaces.

Staleness is explicit in both cases and should be stated the same way: the view is as old as its last refresh, the cache as old as its TTL. The difference is who is responsible for correctness — the database, or your code.

Curated: · Written: · Reviewed:

QA-71How do you approach a legacy system that must keep running but cannot be modified safely?(show answer)

Contain it rather than confront it. The strangler pattern is the standard approach and it works because it never requires a moment where everything changes.

Put a facade in front — a proxy, a gateway, or an API layer — so callers address the facade rather than the legacy system directly. Nothing changes functionally, but you now have a seam where routing decisions can be made.

Then move functionality outward one capability at a time. Choose the first candidate for low risk and high learning: something with clear boundaries, moderate traffic, and observable behavior. Route that capability to a new implementation, keep the old path available, and compare outputs in production before switching.

Data is the constraint that decides the sequence. Capabilities that read shared state are much easier to extract than ones that write it, so start with reads, and where writes must move, use dual-write with reconciliation until you trust the new path.

Characterisation tests are the safety net for the parts you cannot avoid touching: capture the current behavior, including behavior that is arguably wrong, as tests. You are protecting against unintended change, not asserting correctness.

Two rules that keep this honest. Add no new functionality to the legacy system, or the strangler never closes. And keep the extraction funded as a sequence of small, individually valuable steps, because a multi-year rewrite with no intermediate value is the failure mode this pattern exists to avoid.

Curated: · Written: · Reviewed:

QA-72What is the role of an architecture decision record, and what makes a good one?(show answer)

Its role is to preserve the reasoning, not the decision. Anyone can read the current system and see what was chosen; almost nobody can reconstruct why, and without the why, teams either cargo-cult a constraint that no longer applies or overturn a decision that is still load-bearing.

A good one is short — a page — and contains five things. The context: what forces were in play, what the constraints were, what was actually being decided. The options considered, including the ones rejected, because the rejected options are the part that is impossible to recover later. The decision itself, stated plainly. The consequences, honestly, including what became harder. And the status, with a link to any record that supersedes it.

The two qualities that separate useful records from ceremony: writing the consequences you dislike, and stating what would change the decision. A record that lists only benefits is marketing. One that says "this makes cross-region reads slower, and we would revisit if read latency outside our primary region became a product requirement" is a decision that can be re-evaluated on evidence years later.

They should live in the repository next to the code, be written at the time of the decision rather than reconstructed, and be immutable — superseded by a new record rather than edited, so the history of the reasoning survives.

The test of the practice is whether a new engineer's question "why is it like this?" has an answer that is not a person.

Curated: · Written: · Reviewed:

QA-73How do you prevent one tenant from affecting another in a multi-tenant system?(show answer)

Isolate along the dimensions where one tenant's behavior can consume a shared resource: request capacity, data, storage, and the database's working set.

Request capacity: per-tenant rate limits and concurrency limits, enforced at the edge. Without them, one tenant's runaway integration consumes the whole fleet's capacity, and the incident looks like a general outage.

Data: tenant identity must be part of every query, ideally enforced structurally rather than by convention. Row-level security in the database, or a data access layer where the tenant filter cannot be omitted, is much stronger than reviewing every query. Cross-tenant data leakage is the failure with the worst consequences and it usually comes from one forgotten WHERE clause or one cache key missing the tenant.

Work: background jobs should be scheduled fairly rather than first-in-first-out, or one tenant's bulk import starves everyone else's jobs for hours. Per-tenant queues with round-robin consumption is the usual answer.

Database resources: a single tenant with a hundred times the data skews query plans, fills the buffer cache, and makes shared indexes less effective. Beyond a threshold, very large tenants are better served on dedicated infrastructure, which is why most mature multi-tenant products end up with a tiered model — pooled for the many, silo for the few.

And instrument per tenant. Aggregate metrics conceal exactly the situation you are trying to detect, so error rate and latency should be sliceable by tenant before an incident, not after.

Curated: · Written: · Reviewed:

QA-74When does adding a message queue make a system less reliable rather than more?(show answer)

When it converts a fast, visible failure into a slow, invisible one, and when the team treats it as durable storage rather than as a buffer.

The clearest case is an unbounded queue in front of an under-provisioned consumer. Without the queue, overload produces immediate errors and the problem is obvious within a minute. With it, producers keep succeeding, the backlog grows for hours, and by the time anyone notices, the queue holds work that is no longer relevant — notifications about events users have long since seen — and draining it takes longer than the incident that caused it. The queue turned a capacity problem into a much longer degradation.

It also hurts when the work was genuinely synchronous. Moving a step out of the request path only helps if the user does not need the result. If the UI must poll for completion, you have added a component, an eventual-consistency problem, and a new failure mode in exchange for a faster response that the user cannot act on.

Other ways it degrades reliability: retries without a dead letter queue, so a poison message blocks a partition indefinitely; ordering assumptions that the broker does not guarantee; and a dual write between the database and the broker, which silently loses messages or invents them.

The mitigations are bounded queues with explicit rejection when full, alerting on backlog age rather than depth, and a documented answer to what happens when the queue is unavailable — because for many systems the correct answer is to fail the request rather than to lose the work.

Curated: · Written: · Reviewed:

QA-75How do you evaluate whether to adopt a new managed service that your team has no experience with?(show answer)

Evaluate it as an operational commitment, not as a feature list, and time-box a spike that tests the parts that are expensive to discover late.

The questions I would want answered before committing. What are its failure modes, and what does the provider's status history actually show? What are the hard limits — throughput, size, connections, quotas — and where do they sit relative to our projected use in two years? What is the upgrade and maintenance model, and can we control when it happens? What does it cost at ten times our current scale, since the pricing curve is often the thing that changes the decision? What is the exit path, and how much data would have to move?

Then a spike that exercises the awkward paths rather than the tutorial: a failure injection, a restore from backup, a schema or configuration change under load, and the observability story — can we see what it is doing when it misbehaves, or is it opaque?

Adoption should be staged. First a non-critical workload, in production, long enough to experience a real incident. That is where you learn whether support is responsive and whether the failure modes are survivable.

The decision criterion I would state: adopt when it removes work that is undifferentiated for us and is a core competency for them, and when the failure modes are ones we can detect and survive. Decline when the appeal is novelty, or when the exit cost grows faster than the benefit.

Curated: · Written: · Reviewed:

QA-76What does defense in depth mean concretely for a public-facing web application in the cloud?(show answer)

It means no single control failing results in compromise, and it is best described as the layers a request passes and the layers an attacker would have to defeat after that.

At the edge: DDoS protection and a web application firewall, TLS terminated with modern ciphers, and rate limiting. This layer stops volume and known patterns, not a determined attacker.

At the network: the application in private subnets with no inbound route from the internet, security groups referencing each other rather than CIDR ranges, and egress restricted so a compromised instance cannot freely reach the internet. Egress filtering is the layer most often skipped and it is what limits exfiltration and command-and-control.

At the identity layer: short-lived workload identities, narrowly scoped policies, MFA for humans, and no long-lived keys. This is what limits what a compromised process can do.

At the application: input validation, parameterised queries, output encoding, object-level authorization checks, and dependencies scanned and patched.

At the data: encryption at rest with keys the application cannot export, encryption in transit internally as well as externally, and separate credentials per service so one leak does not reach everything.

And crucially, detection across all of them: audit logs, flow logs, configuration change alerts, and anomaly detection — because the assumption behind defense in depth is that some layer will fail, and the value of the remaining layers depends on someone noticing.

The test is to name a layer and ask what still protects you if it fails entirely.

Curated: · Written: · Reviewed:

QA-77How do you design for graceful degradation rather than total failure?(show answer)

Decide, feature by feature, what the system does when a dependency is unavailable — and make that a design output rather than a runtime accident.

Start by classifying dependencies as critical or optional for each user journey. On a product page, the catalogue is critical; recommendations, reviews, and inventory counts are not. That classification is a product conversation and it is the whole basis of the design.

Then implement the optional ones so they fail independently: a timeout tighter than the page's budget, a circuit breaker, a bulkhead so their exhaustion does not consume shared resources, and a defined fallback — cached content, a default, or simply omitting the section.

Feature flags give you the manual version of the same thing: the ability to turn off an expensive or failing feature during an incident without a deploy. A kill switch on every non-critical, resource-intensive feature is one of the highest-value things to build before you need it.

Read-only mode deserves specific mention. Many systems can serve most of their value with the write path disabled, and having that as a tested mode converts a database primary failure from an outage into a degraded period.

Then test it. Failure injection in a controlled environment — kill the recommendation service and confirm the page still renders — is the only way to know the fallback works. Fallback code paths are exactly the code that is never exercised, which means they are the code most likely to be broken when they are finally needed.

Curated: · Written: · Reviewed:

QA-78Explain the difference between at-least-once, at-most-once, and exactly-once delivery, and which one you can actually buy.(show answer)

At-most-once means a message is delivered zero or one times: the sender does not retry, so nothing is duplicated and messages can be lost. At-least-once means the delivery protocol retries until acknowledgement, so duplicates can occur; it reduces loss from ambiguous acknowledgements but does not defeat retention expiry, unrecoverable broker failure, or application bugs. Exactly-once means precisely one delivery.

You can buy at-least-once. It is what essentially every broker provides, and it follows from the fact that a sender cannot distinguish a lost message from a lost acknowledgement.

Exactly-once delivery, in the strict sense of the message crossing the network exactly once, is not achievable in a distributed system with failures. What you can achieve is exactly-once processing, and it is built at the consumer: deduplicate on a message identifier, or make the effect idempotent so that repeated application is harmless.

Some systems advertise exactly-once semantics, and what they generally provide is transactional processing within their own boundary — a consume, process, and produce cycle committed atomically inside the same platform. That is genuinely useful and it stops at the boundary: the moment the effect is an external API call or a write to a different system, the guarantee no longer covers it.

So the architectural rule is to assume at-least-once everywhere and design consumers to be idempotent. That assumption costs little when duplicates never arrive and saves you entirely when they do — and the failure it prevents, double-charging or double-shipping, is the kind that is visible to customers.

Curated: · Written: · Reviewed:

QA-79What is your approach to choosing partition keys for a sharded data store?(show answer)

Choose the key that appears in the queries you must serve fast, and that distributes volume evenly. Those two requirements conflict often enough that the choice is genuinely a design decision rather than a lookup.

Even distribution first. A key with skew produces a hot partition, and a hot partition means one node is saturated while the rest idle — you have paid the full cost of sharding for a fraction of the benefit. Monotonically increasing keys such as timestamps or sequential IDs are the classic mistake, because all current writes land on the newest partition. Tenant ID is the classic mistake in multi-tenant systems, where the largest customer is often orders of magnitude bigger than the median.

Query alignment second. Any query that does not include the partition key becomes a scatter-gather across every shard, which is slow and scales badly. So the key should be present in the dominant access pattern — user ID for a user-centric workload, order ID where orders are the aggregate.

Where a single attribute cannot do both, composite keys help: tenant ID combined with a hash suffix spreads a large tenant across partitions while keeping their data addressable, at the cost of needing a fan-out for whole-tenant queries.

Two further considerations: cross-partition transactions largely stop being available, so entities that must change together should share a partition; and the key is very expensive to change afterward, so it deserves load-test evidence rather than an assumption about distribution.

Curated: · Written: · Reviewed:

QA-80How would you migrate 50TB of data to the cloud with minimal downtime?(show answer)

Separate the bulk from the delta. Fifty terabytes over a 1Gbps link is roughly five days at full utilisation, which is neither achievable nor acceptable as downtime, so the bulk moves ahead of time and only the delta moves during the cutover.

For the bulk, compare network transfer against a physical transfer appliance. The crossover point depends on available bandwidth: at 10Gbps, 50TB is around twelve hours and network transfer is fine; at 100Mbps it is weeks and a physical device is faster and cheaper. Whichever you choose, the bulk copy happens while the source stays live.

For the delta, use change data capture or incremental replication so the target continuously catches up with changes made since the bulk copy. The cutover window then shrinks to the time needed to drain the last few seconds of changes.

The cutover: stop writes at the source, wait for replication to reach zero lag, verify, repoint the application, resume. That is minutes, not days.

The parts that decide success. Verification must be by checksum or row-count reconciliation per table, not by eyeballing — and it must run before cutover, not after. A rollback path must exist: keep the source intact and writable until you have run on the target long enough to be confident. And bandwidth for the delta must be reserved, because a migration that saturates the link degrades the production system you are migrating.

Test the whole sequence on a subset first; the surprises are in the schema, character encodings, and data that violates constraints nobody knew existed.

Curated: · Written: · Reviewed:

QA-81What makes an on-call rotation sustainable?(show answer)

A page rate low enough that being on call is not a week off from other work, and a set of guarantees that make the burden fair.

Set an explicit page-volume target and track sleep interruption; the appropriate number depends on shift length, severity, and local working-hours coverage. A healthy service should produce very few non-actionable or overnight pages. When the burden regularly disrupts sleep or planned work, fix the paging and its causes rather than merely adding people.

Getting there means every alert is actionable, tied to a user-visible symptom, and has a runbook. Alerts that fire and require no action get deleted; that alone usually removes most of the volume. Alerts on causes get replaced by alerts on symptoms.

Fairness matters as much as volume. A rotation needs enough trained people that shifts and recovery time are sustainable; the number depends on shift length, regional coverage, and escalation design. It also needs explicit compensation or time off in lieu for out-of-hours work, a documented escalation path so nobody is alone with a problem beyond them, and an expectation that a responder who was up at 3am does not work the next morning.

The feedback loop is what keeps it healthy over time: every page reviewed in a weekly operational meeting, recurring causes turned into work items with real priority, and the on-call engineer given time during the shift to fix what paged them rather than only to acknowledge it.

And the responder needs the access to actually resolve incidents. A rotation where the on-call engineer must wake someone else for permissions is a rotation that pages two people every time.

Curated: · Written: · Reviewed:

QA-82How do you think about the trade-off between developer velocity and operational safety?(show answer)

As a false dichotomy in the long run and a real one in any given week, which is why it needs a mechanism rather than a philosophy.

The long-run case: the practices that make deployment safe — small changes, automated tests, fast rollback, good observability — are the same practices that make it fast. Teams that deploy many times a day generally have fewer failures and recover faster than teams that deploy monthly, because small changes are easier to verify and easier to reverse. So most apparent trade-offs are actually a missing capability.

The short-run case is real: adding a required review, a soak period, or a manual gate does slow a specific change and does reduce a specific risk. The mistake is deciding that case by argument.

The mechanism I would use is the error budget. Agree an SLO, and let the remaining budget decide. Budget healthy means ship — the evidence says the current pace is not hurting users. Budget exhausted means reliability work takes priority until it recovers. Both sides agreed the number in advance, so the weekly negotiation disappears.

Alongside it, put safety into the path rather than into approval steps. Automated checks in the pipeline, progressive delivery with automated rollback on error-rate regression, and feature flags decoupling deploy from release make a change safe without making it slow. Manual gates are the tool of last resort — they are slow, and they degrade to rubber-stamping under time pressure, giving you the cost without the protection.

Curated: · Written: · Reviewed:

QA-83Walk me through migrating an on-premises workload to the cloud.(show answer)

The four decisions come in order: what the workload actually depends on, what to do with each piece, how the data moves, and when we cut over and how we come back. Jumping straight to the third is how migrations end with a Sunday-night rollback.

Assessment. The unit of migration is a dependency-consistent group, not a server. Discovery means flow data and config inventory cross-checked against what the owning team believes runs there, and I treat that output as a hypothesis until a dry run proves it. What discovery reliably misses is what breaks cutover night: the 3am batch job, a vendor VPN terminating on a firewall appliance, the licence server, the reporting database someone built in 2019 with a direct link to the primary. Every dependency gets a named owner and an explicit treatment before any date is agreed.

Data gravity decides the sequence and sometimes the strategy. 50 TB over a 100 Mb/s link is 46 days at line rate — 400 Tb ÷ 0.1 Gb/s — before retransmits, change freezes, or the time it takes to notice a bad checksum. So the realistic options are a dedicated link or a shipped device, moving the consumers along with the data, or leaving the data behind until last and streaming it behind continuous replication. If the consumers stay on-prem, every call now crosses a WAN, and I would move the application before the data rather than the reverse.

Strategy per group:

MoveChoose it whenWhat you accept
Retire / retainNothing consumes it, or it cannot legally leaveThe cheapest migration is the one you don't do
RehostCommodity tier, deadline-driven, no scaling or delivery problemKeeps old failure modes, often worse latency and worse licensing economics
ReplatformManaged equivalent exists — OS, database, queueBehavioural differences in failover, backups, extensions, collation, connection limits
RefactorThe code is the constraint — release cadence, scaling, cost per requestMost money and time; the only option that fixes the actual problem

Most estates end up mostly rehost/replatform with two or three refactors. That is a reasonable outcome; the failure is refactoring everything, or rehosting something whose architecture is the reason it needed to move.

Why a rehost fails quietly. Latency arithmetic: 200 sequential database round trips at 0.3 ms LAN latency is 60 ms of network. At 1.5 ms across an availability zone boundary it is 300 ms. Nothing in a low-load lift-and-shift test shows this; it appears under concurrency, as a queueing collapse in the app pool. That is why I measure the call pattern — round trips per request, not just bandwidth — before choosing rehost.

Cutover. Something like:

  • T-2w: target built as code, bulk copy started, smoke and restore tests green against copied data
  • T-2d: continuous replication running, lag under 60 s, reconciliation job proving it
  • T-0: freeze writes or force read-only, final delta, then reconcile — row counts plus a business invariant (total ledger balance, open orders), not just "it's up"
  • T+30m: internal read-only traffic, then 10% of production
  • T+24h: rollback deadline, once reverse replication is running and verified

Rollback is a design decision, not a fallback. The moment the target accepts a write the source doesn't have, "just go back" becomes a data-reconciliation project. So the plan names the exact point after which the recovery path is forward-only, and before that point rollback is a DNS or traffic-weight revert plus reverse sync, tested before cutover, not improvised on the night.

De-risking. Pilot with a low-risk stateless workload first and pay for the runbook that gets. Test backup restore in the target — SAN snapshot DR has no automatic equivalent, and the first place anyone finds out is the first restore. Re-baseline cost with real figures: committed capacity for steady state, cross-AZ traffic, and licensing, where per-socket or per-core entitlements can cost more on cloud vCPUs than the server they replace. Re-check the security baseline, because network paths, key custody and log destinations all change. Finally, set abort thresholds as numbers the go/no-go owner can act on: error rate above 2× baseline for 10 minutes, p99 above 1.5×, replication lag above 60 s — and stop the wave rather than push through it.

Curated: · Written: · Reviewed:

QA-84How would you architect a system that must prove which user performed every action, for a regulator?(show answer)

Make the audit record a byproduct of the action rather than something the application chooses to write, and make it tamper-evident.

The identity must be genuine end to end. That means no shared accounts, no service acting on a user's behalf without carrying the original identity, and no database access path where a human uses a shared credential. Where an operator acts on a user's behalf, the record needs both identities — who acted, and for whom.

The record must be written in the same transaction as the change. An audit log written after commit can be lost; one written before can describe something that never happened. Database triggers or a change data capture stream are the two mechanisms that catch changes regardless of the writer, which is what matters when the regulator asks about a manual fix applied at 2am.

Tamper-evidence: append-only storage, write-once retention policies, and separate credentials such that no identity can both modify data and delete the corresponding audit record. Hash chaining, where each record includes a hash of the previous one, makes silent deletion detectable.

Retention must meet the regulatory window, which is often years, so tiering to cheap storage with the guarantees intact is part of the design.

And it must be queryable within the timeframe the regulator expects. An audit trail that exists but takes a week to search fails the requirement in practice. That usually means indexing by subject and by time, and testing the retrieval path as seriously as the write path.

Curated: · Written: · Reviewed:

QA-85What is the practical difference between a NAT gateway and an internet gateway, and why does it appear on a cost review?(show answer)

An internet gateway is a route attachment that lets resources with public IP addresses communicate with the internet in both directions. It is horizontally scaled by the provider and it is free — you pay for data transfer, not for the gateway.

A NAT gateway lets resources in private subnets, which have no public addresses, make outbound connections while remaining unreachable from outside. It is a managed appliance and it is billed two ways: an hourly charge per gateway, and a per-gigabyte processing charge on everything that passes through it.

That processing charge is why it appears on cost reviews. A workload pulling container images, downloading packages, shipping logs to a third party, or reading from object storage through the NAT gateway pays per gigabyte for traffic that has no business leaving the provider's network at all. On a busy cluster this is regularly one of the largest surprise line items.

On AWS, the fix is often private endpoints. S3 and DynamoDB gateway endpoints have no hourly or data-processing charge and remove that traffic from the NAT path. Interface endpoints for other services have hourly and per-gigabyte charges; whether they cost less than NAT depends on traffic volume, availability-zone placement, and the number of endpoints, so compare the actual pricing rather than assuming.

The other common cause is deploying one NAT gateway per availability zone for resilience — which is correct, since a single gateway is both a single point of failure and a source of cross-zone transfer charges — and then forgetting that the hourly cost is now tripled across several non-production environments.

Curated: · Written: · Reviewed:

QA-86How do you decide whether an incident warrants a postmortem, and what makes one worth reading?(show answer)

Write one whenever the incident consumed a meaningful share of the error budget, affected customers visibly, required an unplanned rollback, or surprised the team — that last criterion matters most, because surprise means the mental model was wrong, which is worth understanding regardless of impact.

What makes one worth reading is that it explains the mechanism, not the sequence. A timeline of who did what at which minute is necessary and is not the value. The value is the answer to why the system permitted this, and why it took as long as it did to detect and to resolve.

Concretely, a good one contains: the customer impact in the customer's terms, the mechanism of the failure explained well enough that a reader could reproduce it, why existing safeguards did not catch it, what made detection slow, what made mitigation slow, and the action items with owners and dates.

Two things separate the useful ones from the ritual ones. It is blameless in a real sense: the question is what about the system made this action reasonable at the time, not who made the mistake. An engineer who fears the document will describe events carefully and reason vaguely, which destroys the value.

And the action items are prioritised honestly. Ten items nobody will do is worse than two that will, because it converts the process into theatre. The measure of the practice is whether the same failure recurs, not whether documents were produced.

Curated: · Written: · Reviewed:

QA-87When would you accept a single point of failure in a design?(show answer)

When the cost of removing it exceeds the expected cost of the outage it causes, and when that trade has been stated explicitly rather than overlooked.

The calculation needs three numbers: how often the component is likely to fail, how long recovery would take, and what an hour of that outage costs the business. A component with a mean time between failures measured in years, a fifteen-minute recovery, and an internal-only impact is frequently not worth eliminating.

Legitimate acceptances I would defend: a single control-plane component whose failure does not affect the data path, so running systems continue while it is down. A single instance of an internal tool with a documented, rehearsed rebuild. A single-region deployment for a product whose customers are in one region and whose contractual availability is 99.9%.

What must accompany the acceptance: monitoring that detects the failure quickly, a tested recovery procedure with a known duration, and a documented decision record so that a future engineer does not assume it was an oversight — or, worse, assume it was deliberate when it was not.

What I would not accept: a single point of failure in the data path that nobody has identified, which is the usual situation. Undocumented single points of failure are found during incidents, and the recovery time is then unbounded because nobody has practised it.

The useful exercise is to walk the request path and name what happens if each component disappears. The surprises in that exercise are the real risk, not the ones already on the diagram.

Curated: · Written: · Reviewed:

QA-88How do you handle personal data in logs and analytics without breaking the product?(show answer)

Decide what identifiers the downstream consumer actually needs, and supply the weakest one that works.

For most operational logging, a stable pseudonymous ID is enough. Debugging needs to correlate a user's requests, not to know their email address, so log a user ID and resolve it to a person only through a separate, access-controlled lookup. That single change removes personal data from the highest-volume, widest-access data store you have.

For fields that must not appear at all — tokens, card numbers, health data, message contents — redact at the logging layer rather than at the call sites. A middleware that strips known-sensitive keys and pattern-matches obvious secrets protects against the case that actually causes incidents, which is a new field added by someone who did not know the rule.

For analytics, aggregate early and retain the raw event for a short window only. Most analytical questions are answered by counts and distributions, and keeping raw per-user events indefinitely creates a deletion obligation and a breach exposure for value that has usually already been extracted.

The deletion path is the part teams underestimate. A user's deletion request must reach every copy: primary store, replicas, backups, search index, cache, analytics warehouse, and the third-party tools. Designing this at the start — a documented inventory of where identifiers propagate, and a deletion job that covers them — is far cheaper than reconstructing it under a regulatory deadline.

Backups are the honest exception, and the standard answer is documented retention windows with deletion applied on restore.

Curated: · Written: · Reviewed:

QA-89What is your process for evaluating whether an architecture will meet a latency requirement before it is built?(show answer)

Build a budget, then check it against measured components rather than guesses.

Start from the user-facing target and decompose it. If the page must render in 500ms at p95, allocate: network round trip to the edge, TLS if the connection is new, the application's own processing, each downstream call, and the database queries beneath those. Write the numbers down. Two things usually become obvious immediately — that sequential calls sum, and that the budget is already spent before the interesting work begins.

Then replace assumptions with measurements. Round-trip times between regions and zones are knowable. Database query latency for a representative dataset is measurable with a prototype in an afternoon. Third-party API latency is published or observable. Anything you cannot measure yet is a risk to be spiked, not a number to be assumed.

Pay particular attention to serial dependencies and to fan-out. A chain of five sequential calls at 20ms each has a floor of 100ms. A fan-out to twenty parallel calls has a latency equal to the slowest of the twenty, which at p95 per call is much worse than the per-call p95 — this is the tail-amplification effect that makes fan-out architectures unexpectedly slow.

Then design against the budget: parallelise what can be parallel, cache what is stable, and move work off the request path where the user does not need the result.

And validate under load, since latency at one request per second tells you nothing about latency at capacity.

Curated: · Written: · Reviewed:

QA-90How do you approach securing service-to-service communication inside a private network?(show answer)

Do not treat the network boundary as the authorization boundary. A private network limits who can reach a service; it says nothing about which caller should be allowed to do what, and a single compromised workload turns the flat internal network into an open one.

Authenticate both ends. Mutual TLS with short-lived certificates issued per workload gives each service a verifiable identity, and the certificate rotation is what makes a stolen credential useless quickly. A service mesh provides this without per-service implementation, which is most of its value.

Authorize per call. The caller's identity should map to a policy stating which services and which operations it may invoke. Default-deny, with explicit allows, so a new service cannot call everything by accident.

Propagate the end-user context separately from the service identity. A downstream service usually needs to know both which service called it and on whose behalf, and conflating those is how internal endpoints end up performing privileged actions for any caller.

Segment the network as a second layer: default-deny network policy so that even an unauthenticated path does not exist between unrelated services. Identity-based authorization and network segmentation catch different failures, which is why both are worth having.

And encrypt internal traffic. The argument against it — the network is trusted — is the assumption being removed. It also matters for compliance, and for the realistic case where the network spans zones or providers.

Curated: · Written: · Reviewed:

QA-91What is the value of chaos engineering, and how would you introduce it without causing an incident?(show answer)

Its value is that it tests the failure paths, which are the only paths that are never exercised by normal traffic and are therefore the most likely to be broken. Every system has a documented behavior for a dependency timing out; a large share of them behave differently in reality, and the difference is discovered during an incident unless you go looking.

Introducing it safely is a matter of sequence and blast radius.

Start with a hypothesis, not an outage: state what you expect to happen when a specific dependency fails, and what signal would show it. An experiment with no prediction is just breaking things.

Start in a non-production environment to validate the tooling and the observation, then move to production, because non-production rarely reproduces the interesting behavior.

In production, start with the smallest blast radius that can teach you something: one instance, one availability zone, a small percentage of traffic, during working hours, with the team watching and an abort mechanism that is tested first. Working hours matter — the point is to learn with everyone present, not to simulate the 3am experience.

Have a stop condition defined in advance, tied to a customer-facing metric, and honour it.

Then broaden gradually as confidence builds. The organisational prerequisite is that the team already has good observability and a functioning incident process — chaos experiments on a system nobody can observe produce outages rather than knowledge, and that is how the practice gets a bad name.

Curated: · Written: · Reviewed:

QA-92How do you evaluate a vendor's claim that their service is highly available?(show answer)

Read the SLA rather than the marketing page, and then read the exclusions.

The number itself is less informative than its definition. What counts as downtime — full unavailability, or degraded performance? Measured over what window: a month, so a four-hour outage may be under the threshold, or a shorter window that reflects real impact? Measured by whom, using what probes? Many SLAs exclude anything not detected by the vendor's own monitoring.

Then the exclusions, which are where the real terms live: scheduled maintenance, which in some contracts is unlimited; issues attributed to customer configuration; regional events; and dependencies the vendor itself relies on.

Then the remedy. Almost every SLA remedy is a service credit proportional to the fee, which is negligible compared to the business impact of an outage. So the SLA is not risk transfer — it is a statement of intent with a small penalty, and you should architect as though the remedy did not exist.

Evidence beats terms. A public status history with real incident detail, and postmortems that are specific rather than apologetic, tell you more than any number. So does asking directly about their largest incident in the last two years and how it was handled.

And architect for their failure regardless: what does your system do when this vendor is down for an hour? If the answer is that you are down too, then their availability is now your ceiling, and that fact belongs in your own SLO calculation.

Curated: · Written: · Reviewed:

QA-93How would you design an API that returns a large result set?(show answer)

With cursor-based pagination, a bounded default page size, and no ability for a caller to request everything.

Cursor pagination — an opaque token encoding the position, typically the last seen sort key — is preferable to offset pagination for two reasons. It is stable under concurrent inserts and deletes, so a caller paging through results does not see duplicates or skip items when the underlying data changes. And it performs consistently: OFFSET 100000 requires the database to scan and discard 100,000 rows, so deep pages get progressively slower, whereas a cursor is an indexed seek.

The sort must be deterministic and must include a tiebreaker, usually the primary key, or the cursor is ambiguous for rows sharing a sort value.

Set a default page size and a maximum, and enforce the maximum server-side. A caller requesting a million items is either a mistake or an attack, and either way the server should not attempt it.

Return the next cursor in the response, and prefer not to return a total count. Counting the full result set is often as expensive as the query itself, and callers rarely need it — where they do, offer it as an explicit, separately-priced option.

For genuinely large exports, pagination is the wrong shape entirely. Offer an asynchronous job: the caller requests an export, receives a job ID, polls or receives a webhook, and downloads a file from object storage. That keeps a multi-gigabyte result out of the request path where it would tie up connections and time out.

Curated: · Written: · Reviewed:

QA-94What would you look at first if a system is slow only for users in one geographic region?(show answer)

Establish whether it is the network path or the serving path, because they lead to entirely different fixes.

Network path first, since regional-only slowness is most often distance or routing. Measure round-trip time from that region to your edge and to your origin. If the region is served from a distant point of presence, or if traffic is being routed to a far region because of a DNS or anycast configuration, latency is structural and no amount of application optimisation will help. Traceroutes from the affected region, and real user monitoring broken down by geography, answer this quickly.

Then check whether the edge is actually serving that region. A CDN with a low hit rate in one region — because content is not being cached there, or because a cache key includes something region-varying — means those users are paying the full origin round trip while everyone else is not.

If the network path is normal, look at the serving path for anything region-correlated: a regional replica with high lag, a dependency deployed in fewer regions, or a database read that crosses regions for those users specifically.

Then the non-technical causes, which are common: peak hours in that region coinciding with a batch job, a regional ISP issue, or regulatory inspection adding latency in some jurisdictions.

The measurement that settles it is a synthetic probe from inside the affected region against both the edge and the origin. Aggregate metrics from your own infrastructure cannot see a problem that occurs before the request arrives.

Curated: · Written: · Reviewed:

QA-95How do you keep infrastructure costs visible to the teams that create them?(show answer)

Attribute cost to the team automatically, report it in units the team recognises, and give them the ability to act on it.

Attribution starts with tagging enforced at creation. A policy that rejects an untagged resource is the only version that works, because tagging campaigns after the fact never complete. Where tags are insufficient — shared clusters, shared databases — allocate by a proxy such as namespace resource usage or request volume, and be transparent that it is an allocation rather than a measurement.

Report in the team's units. A monthly figure of forty thousand dollars means little; cost per thousand requests, or cost per active customer, is a number a team can reason about and can watch move when they change something. Trend matters more than absolute value.

Deliver it where they already look. A monthly finance report is read by finance. A weekly figure in the team's own channel, and a cost estimate on a pull request that changes infrastructure, reaches the person making the decision at the moment they make it.

Give them the lever. Visibility without authority produces resentment; a team that can see its bill but cannot change instance types or retention policies will ignore it. The paved-road modules should make the efficient choice the default.

And be careful what you incentivise. Cost as a target on its own produces under-provisioning and incidents. Cost per unit of value, reviewed alongside reliability, is the pairing that leads to good decisions.

Curated: · Written: · Reviewed:

QA-96Describe a time-bounded plan for reducing risk in a system with known but unquantified problems.(show answer)

Quantify first, because a list of known problems without impact estimates gets prioritised by whoever is loudest.

Week one: inventory and rank. For each known problem, estimate two things — the probability it causes an incident in the next quarter, and the impact if it does. Rough categories are enough; the goal is to separate the three things that matter from the twenty that do not. Include the ones nobody wants to write down, particularly single points of failure and untested recovery paths.

Weeks two to four: buy information cheaply on the top items. A restore test tells you whether your backups work, and takes a day. A failover drill tells you your real RTO. A load test tells you where capacity actually ends. Each converts a guess into a number, and frequently one of them turns out to be far worse than assumed, which reorders the list.

Weeks five to ten: fix in order of expected loss, and prefer detection and recovery over prevention early on. Making an incident detectable in one minute rather than thirty reduces impact across every failure mode at once, whereas preventing one specific failure helps once.

Throughout: put each finding on the backlog with its estimate visible, so the trade against feature work is explicit and made by the people accountable for both.

At the end, re-rank. The value of the exercise is a list that is now evidence-based, and a team that knows which risks it has consciously accepted rather than merely tolerated.

Curated: · Written: · Reviewed:

QA-97What is the difference between a read replica and a standby, and why does it matter operationally?(show answer)

A standby exists for failover; a read replica exists to serve queries. The same underlying replication may power both, but the operational contracts differ in ways that matter under pressure.

A standby is provisioned and operated primarily for failover. Depending on the database product it may be synchronous or asynchronous, may or may not serve reads, and may be promoted automatically or manually. Its recovery point and readiness come from the configured acknowledgement policy, monitored lag, and tested promotion—not from the label alone.

A read replica is typically asynchronous and does serve queries, which means its lag varies with the load you place on it. A heavy analytical query on a replica can delay log application and push lag from milliseconds to minutes. If you have quietly been treating that replica as your disaster-recovery plan, your RPO just moved without anyone noticing.

The operational consequences are concrete. Do not assume a read replica is a valid failover target unless you monitor its lag against your RPO and have tested promotion. Do not run heavy reporting on the replica you intend to promote. And be aware that promoting a replica usually breaks the replication topology — other replicas of the old primary do not automatically follow the new one, which is a detail that turns a five-minute failover into an hour.

The design conclusion is to separate the contracts even if one replica can fulfil both: identify the failover target and protect its recovery objective, then size query-serving replicas for their read load. Monitor lag on each for the reason it exists—read correctness on one, recovery objectives on the other.

Curated: · Written: · Reviewed:

QA-98How do you introduce a significant architectural change to a team that is resistant to it?(show answer)

Find out what the resistance is actually about, because it is usually not the technical merits.

The common reasons, each needing a different response: the team has been through a migration that consumed a year and delivered nothing, and expects a repeat; the change makes someone's expertise less valuable; the operational burden lands on people who were not in the decision; or the team can see a problem with the proposal that the proposer has not addressed. That last one is the most important, and it is the reason to listen first rather than to persuade.

Then reduce the size of the commitment. A proposal to re-architect is a bet nobody can evaluate. A proposal to move one low-risk service, with a defined success measure and a stated point at which you would abandon the approach, is a decision the team can agree to on Tuesday. Evidence from a real migration changes minds in a way that a document does not.

Give the burden to the proposer first. Whoever proposes the change should operate the first instance of it, including being on call for it. This is both fair and clarifying: proposals that will not survive contact with operations tend to be revised at this point.

Make the reversal condition explicit and honour it. A team that has seen an initiative abandoned when the evidence went against it will engage seriously with the next one.

And be honest that some resistance is correct. The right outcome of a design review is sometimes not proceeding.

Curated: · Written: · Reviewed:

QA-99If you could enforce only three technical controls across an entire cloud estate, which would you choose and why?(show answer)

No long-lived credentials. Every workload and every pipeline authenticates with a short-lived, federated identity, and human access requires MFA. Long-lived credentials create a durable, replayable path into an estate and are difficult to inventory and rotate; federation sharply reduces that exposure window. It also forces the identity model to be coherent, because federation does not work without one.

All changes through a reviewed pipeline, with no standing human write access to production. This gives you three things from one control: every change is peer-reviewed before it happens, every change is recorded with an author and an approver, and the environment matches its code, which makes drift detectable. It is also the control that most audits are really asking about. Break-glass access exists, is time-boxed, and alerts loudly.

Comprehensive, tamper-evident audit logging, retained and centralised outside the accounts it describes. This is the control that makes everything else verifiable. Without it you cannot scope an incident, cannot prove a control was effective, and cannot detect the slow drift that precedes most failures. Placing it outside the account being logged is what makes it survive a compromise of that account.

The reasoning behind the selection: the first prevents the most common entry, the second prevents the most common self-inflicted damage and produces evidence as a byproduct, and the third is what lets you detect and understand whatever the first two did not stop. Encryption and network controls matter, but they protect narrower classes and are less often the deciding factor.

Curated: · Written: · Reviewed:

QA-100What is the difference between an API gateway and a load balancer, and when do you need both?(show answer)

A load balancer distributes connections across backends. It is a transport-level concern: health checking, connection draining, TLS termination, and spreading load. It does not know or care what the request means.

An API gateway operates on the API itself. It authenticates callers, enforces authorization and rate limits per consumer, validates requests against a schema, routes by path or version to different services, transforms payloads, and produces per-consumer usage metrics. It is a policy enforcement point.

You need both when API policy and backend traffic distribution are distinct responsibilities: a gateway enforces the public contract and routes to services whose own load balancers distribute traffic. Many managed API gateways include their own highly available front end, and smaller systems may need only a layer-7 load balancer with the required policy features, so two separately operated products are not universal.

You can skip the gateway when there is a single internal consumer, no per-consumer policy, and no public exposure — adding one there is a hop and a component for no benefit.

You should not skip it when you have external consumers. The functions it centralises — authentication, rate limiting, request validation, usage metering, and a stable public contract independent of internal service boundaries — are otherwise reimplemented in every service, inconsistently, and the inconsistencies are where the security gaps appear.

The design caution is that a gateway is on every request path, so it is both a single point of failure and a tempting place to accumulate business logic. Keep it to cross-cutting policy; routing rules that encode domain decisions belong in services, where they can be tested.

Curated: · Written: · Reviewed: