Top 100 Solutions Architect Interview Questions and Answers
The questions most likely to actually come up in your Solutions Architect interview, ranked by likelihood — with detailed, senior-level answers covering what an interviewer is really listening for.
Curated: · Written: · Reviewed:
QA-1A customer says they need a system that is 'fast, secure, and scalable'. How do you turn that into an architecture?(show answer)
Those three words are not requirements, they are the absence of requirements, and accepting them is how a project ends up rebuilt in year two. The job is to convert each into something with a number and a consequence.
Fast becomes a latency budget attached to a named user action: the search results page renders in under 400ms at p95 for users in their primary market. That is testable, and it immediately raises the questions that matter — measured from where, at what concurrency, with how much data.
Secure becomes a threat model and a compliance boundary: who are we defending against, what data classes exist, which regulations apply, what does an auditor need to see. A customer in payments and a customer running an internal tool both say secure and mean entirely different budgets.
Scalable becomes a growth curve with a horizon: 2,000 orders a day now, 20,000 expected within eighteen months, with a seasonal peak of four times the daily average in one week of the year. That last figure usually changes the design more than the average does.
The technique that gets these out of people is to ask what would make this a failure rather than what they want. Customers describe aspirations vaguely and disasters precisely.
Then write the numbers down and get them signed off, because every one of them is a constraint you will be held to, and the ones you inferred silently are the ones that will be disputed.
Curated: · Written: · Reviewed:
QA-2How do you scope a proof of concept so it actually reduces risk rather than becoming a small version of the whole system?(show answer)
Start by naming the single assumption that, if wrong, kills the project. A POC exists to test that assumption and nothing else.
That framing does most of the work. If the risk is whether the legacy mainframe can sustain 500 lookups a second, the POC is a load harness against that interface — no user interface, no authentication, no data model. If the risk is whether the customer's users will accept an asynchronous workflow, the POC is a clickable prototype with no backend at all. These are entirely different artifacts, and confusing them is why POCs run long.
Write the success criterion before building: a number and a threshold, agreed with the customer, that determines pass or fail. Without it the POC ends in a discussion about whether it worked, which the most senior person in the room wins.
Time-box it, and state explicitly what is excluded — error handling, security, scale, operability. Then say clearly, in writing and more than once, that the POC is not the foundation of the production system. The most common failure in this role is a POC that succeeds so visibly that the customer asks to ship it, and the team spends a year paying for shortcuts taken to answer a question that has already been answered.
If the answer turns out to be no, that is a successful POC. Cheap disconfirmation is the entire point.
Curated: · Written: · Reviewed:
QA-3A customer wants to build a component in-house that a vendor already sells. How do you evaluate that decision?(show answer)
Compare on total cost and strategic value, not on the vendor's licence fee against an engineer's optimism about how long it would take.
Build genuinely wins when the component is the customer's differentiator — the thing their business is actually good at — when their requirements are unusual enough that every vendor is a poor fit, or when the vendor relationship creates unacceptable risk in a domain that is core to them.
Buy wins in most other cases, and the reason is the part that never appears in the build estimate: the ongoing cost. A built component needs maintenance, security patching, on-call coverage, documentation, and a bus factor above one, forever. That is typically a meaningful fraction of an engineer's time every year, indefinitely, and it starts the day the project is declared finished.
The build estimate is also reliably wrong in a predictable direction, because it prices the happy path. The vendor's product includes the edge cases their thousand customers found, which yours will find too, one at a time, in production.
The questions I would put to the customer: is this component something a competitor would notice you having? What happens if the person who built it leaves? What is the five-year cost of each option including staff time? And what does exiting the vendor cost if it goes wrong?
That last one is where I would push back on a pure buy answer: buy, but keep the integration behind an interface you own.
Curated: · Written: · Reviewed:
QA-4Walk me through how you build a total cost of ownership estimate for a proposed architecture.(show answer)
Infrastructure is the part everyone estimates and it is rarely the largest number. A credible TCO has four categories.
Infrastructure: compute, storage, network egress, and managed service fees, projected across the growth curve rather than at today's volume. Egress and cross-zone transfer deserve their own lines because they scale with usage and are invisible in per-instance pricing.
Licensing and vendor fees, including how the price changes at the next tier. Many vendor contracts are cheap at pilot volume and step sharply at production volume, and finding that out after signing is a common and avoidable failure.
People: the operational load. How many engineers, at what fraction of their time, to run this? A self-hosted database might save a meaningful monthly figure and consume a quarter of an engineer's year, which is usually a net loss. This is the category that decides most managed-versus-self-hosted arguments honestly.
Change: the cost of the migration itself, the parallel-run period where both systems are paid for, training, and the productivity dip while a team learns something new.
Then present it as a range with the assumptions visible, over a three to five year horizon, and identify which assumption the total is most sensitive to — usually the growth rate. A single number invites false confidence and gets quoted back to you eighteen months later.
And include the do-nothing option's cost, since that is what the spend is actually being compared against.
Curated: · Written: · Reviewed:
QA-5Two systems need to exchange data. How do you choose between a synchronous API, asynchronous messaging, batch, and webhooks?(show answer)
Decide from three properties: whether the caller needs the result to continue, how fresh the data must be, and how much coupling in availability is acceptable.
A synchronous API fits when the caller cannot proceed without the answer — a credit check during checkout, an address validation. The cost is that the caller's availability becomes the product of both systems, and its latency includes the callee's. Anything in a user-facing path needs a timeout and a defined behavior when the call fails.
Asynchronous messaging fits when the second system needs to know something happened but the first does not need its response. It decouples availability: the consumer can be down and catch up. The cost is eventual consistency the user may see, and the operational burden of a broker.
Batch fits when the freshness requirement is measured in hours, when volumes are large, or when the other side is a legacy system that offers a nightly file and nothing else. It is unfashionable and frequently correct — a nightly reconciliation is simpler and more debuggable than a stream nobody needed.
Webhooks fit when the other party owns the event and you want push rather than poll. Assume at-least-once delivery, verify signatures, persist the event, then acknowledge quickly so the sender is not left waiting on your downstream work.
For integration with a customer's existing estate, the deciding constraint is usually what the other system can actually do, not what you would prefer. Design to their real capability, and be explicit when it forces a compromise.
Curated: · Written: · Reviewed:
QA-6You are presenting an architecture to a room containing a CTO, a finance director, and two engineers. How do you structure it?(show answer)
Lead with the decision you want, not with the design. Everyone in that room is there to approve, fund, or object, and burying the ask behind twenty slides of topology means the finance director disengages before you reach it.
Structure it as: here is the problem in business terms, here is what we propose, here is what it costs, here is what it buys, here are the two risks and what we would do about them. Then the technical detail, for the people who want it.
Address each audience's actual question. The CTO wants to know whether this is a sensible bet and what it commits them to. The finance director wants the number, its confidence, and when it is spent. The engineers want to know whether it will work and whether you have thought about the parts that are hard.
The failure that loses credibility fastest is answering a business question with technical detail. If someone asks whether it will be ready by March, the answer is a date and a confidence, not a description of the pipeline.
Diagrams should be layered: one that fits on a slide and shows five boxes, and detailed ones in the appendix for the people who will read them.
And be honest about uncertainty. Stating that the cost estimate is plus or minus thirty percent until the integration is scoped builds more trust than a precise number that turns out wrong — and it is the engineers in the room who will notice which one you did.
Curated: · Written: · Reviewed:
QA-7A customer's stated requirement conflicts with what you believe they actually need. How do you handle it?(show answer)
Assume first that they know something you do not. Stated requirements that look wrong are often carrying a constraint nobody mentioned — a regulator's expectation, a commitment made to a client, a political reality inside the organisation. Ask why the requirement exists before arguing with it.
If, having asked, the requirement still looks like it will not serve them, the useful move is to separate the requirement from the outcome. They asked for a real-time dashboard; what they need is to know within the working day whether a batch failed. Those have very different costs, and once the outcome is on the table the conversation is about how to achieve it rather than about who is right.
Where the conflict is genuine, put both options in front of them with the consequences priced. Not an argument, a comparison: this is what you asked for, here is what it costs and what it constrains later; here is the alternative, cheaper by this much, with this limitation. Then let them decide, because it is their system and their budget.
Record the decision either way, including the case you made. If they choose the option you advised against, that record protects the relationship later — it is a documented decision rather than a surprise.
The one place I would not simply defer is a requirement that creates a security, privacy, or data-integrity problem. Those get raised explicitly, in writing, because the consequence lands on their customers rather than only on them.
Curated: · Written: · Reviewed:
QA-8How do you estimate a migration that involves a legacy system nobody in the room fully understands?(show answer)
Refuse to give a single number, and say why in a way that is useful rather than evasive.
The honest position is that the estimate's uncertainty is dominated by one unknown — what is actually in the legacy system — and the fastest way to reduce it is a short discovery phase rather than a longer estimation meeting. Propose that as the first deliverable: two to four weeks to inventory the interfaces, the data volumes, the undocumented integrations, and the batch jobs nobody remembers scheduling.
While waiting, estimate in ranges tied to explicit assumptions. If there are fewer than twenty integration points, this range; if more, this other one. That gives the customer something to plan with and makes clear what would change it.
Some things can be measured immediately and are worth measuring: row counts and data volumes, the number of distinct client applications hitting the system, peak transaction rates from existing monitoring, and the actual schema. Facts collected in a day narrow a range considerably.
Expect the discoveries to be of a particular kind: undocumented consumers, business logic in stored procedures or a trigger, data that violates constraints the schema does not enforce, and a scheduled job on someone's workstation. Budget for that class of finding rather than being surprised by it.
And structure the commercial arrangement to match the uncertainty. Fixed price on a system nobody understands transfers risk to whoever is most optimistic, which serves neither side.
Curated: · Written: · Reviewed:
QA-9The customer wants 99.99% availability. What do you do before agreeing to it in a proposal?(show answer)
Find out what they mean, what it is worth, and what it will cost them — in that order, before any number goes in a document.
What they mean: 99.99% is 4.4 minutes a month, and it is meaningless until you define which service, which endpoints, measured from where, and what counts as down. Degraded performance? A single failing feature? Planned maintenance? Customers frequently mean "we do not want the outage we had last year", which is a different and more tractable requirement.
What it is worth: ask what an hour of downtime costs them. If the answer is a large number, the architecture is justified. If they cannot answer, that is itself informative — the requirement is probably inherited from a template rather than derived from the business.
What it costs: the gap between 99.9% and 99.99% is usually the gap between a well-run single-region deployment and a multi-region active-active one, with the consistency complexity that implies. Often several times the infrastructure spend and a permanent tax on delivery speed. Presented as a priced option, many customers choose 99.9%.
Also check it is achievable: your availability is capped by the product of every dependency in the path, including their systems and their network. Promising four nines on top of a third-party integration with three nines is a commitment you cannot keep.
Then write it as an SLO with a defined indicator, and make sure the SLA in the contract is looser than the SLO you engineer to.
Curated: · Written: · Reviewed:
QA-10How do you decide what belongs in a statement of work versus what stays flexible?(show answer)
Fix what the customer is buying and what determines done; leave flexible what is genuinely unknown, and be explicit about which is which.
The parts that must be specific: the outcome in the customer's terms, the acceptance criteria, the interfaces and systems in scope, the environments delivered, the assumptions the estimate depends on, and what the customer is responsible for providing — access, data, decisions, and people. That last category causes more overruns than technical difficulty does. A project waiting six weeks for firewall access is a schedule problem nobody wrote down.
The parts that should stay flexible: internal design decisions, technology choices below the level the customer cares about, and sequencing within a phase. Fixing those in a contract means a change request every time you learn something, which is expensive for both sides and discourages the team from improving the design.
The mechanism that makes flexibility safe is a defined change process — what constitutes a change, who approves it, and how it is priced — agreed before it is needed rather than negotiated during a disagreement.
For work with high uncertainty, phase it: a fixed-price discovery producing a firm estimate for the next phase. That gives the customer a real decision point and prices the risk honestly instead of hiding it in a padded fixed bid.
And write the exclusions down. Most disputes are about something both sides assumed the other was doing.
Curated: · Written: · Reviewed:
QA-11A customer asks why they should not just use the vendor's reference architecture unchanged.(show answer)
Because a reference architecture is a demonstration of the vendor's products under assumptions that are unlikely to be theirs, and the assumptions are the part that is not written down.
It encodes a workload shape — read/write ratio, data volume, traffic pattern, latency requirement — and a set of trust boundaries and failure modes. Change any of those and the design's justification changes with it, even though the diagram still looks correct.
It is also, reasonably, optimised for the vendor's revenue. It uses their managed services where a cheaper option exists, and it will not mention that one component could be replaced by something the customer already owns and operates well.
That said, the answer is not to ignore it. A reference architecture is a good starting point and a good check — it encodes real experience about how the products fit together, and diverging from it should be a deliberate decision rather than an oversight.
What I would do is walk it against their specific constraints: their data volumes, their compliance boundary, their existing estate, their team's operational capability, and their cost envelope. Each divergence gets a recorded reason. Often most of it survives, and the three parts that change are the three that would otherwise have caused the project to fail.
The team's operational capability is the constraint most often skipped, and the most expensive to get wrong — a design nobody on the customer's side can run is a design that degrades the moment you leave.
Curated: · Written: · Reviewed:
QA-12How do you handle a customer integration where the other side's API is poorly documented and unreliable?(show answer)
Design so their unreliability is contained, and get the facts you need empirically since the documentation will not give them to you.
Empirically first: call it, record real behavior, and build a small conformance suite. What are the actual response times and their distribution, what error shapes appear, is it idempotent, what are the undocumented rate limits, does it return 200 with an error in the body. That last pattern is common in older enterprise APIs and breaks naive clients silently.
Then contain it. An anti-corruption layer — one module that owns every call to them and translates into your own domain model — means their oddities do not leak into the rest of the system, and a change on their side has one blast radius. Timeouts derived from measured latency, a circuit breaker, bounded retries with backoff, and idempotency on anything that mutates.
Then decide the degraded behavior with the customer, because it is a business decision: queue and retry, fall back to a cached or default value, or fail the operation visibly. For a partner integration, queueing with a visible backlog is often right — the work is not lost and someone can see it is stuck.
Finally, instrument it separately and share the data. When their API is the cause of a problem, having a graph of their error rate and latency turns a dispute into a conversation, and it is frequently the thing that gets the other side to fix something.
And put the observed behavior in writing as an assumption in the design, since it is not in their documentation.
Curated: · Written: · Reviewed:
QA-13What do you do when the customer's team will not be able to operate what you are proposing?(show answer)
Treat it as a design constraint of the same weight as latency or cost, because a system that cannot be operated will degrade to whatever the team can actually run — usually within a year, and usually unsafely.
Assess it honestly and early: what do they run today, what is their on-call maturity, how many people, what is their appetite for learning something new, and what happens when the one person who understands it leaves. Ask what broke last year and how it was fixed; the answer tells you more than a capability matrix.
Then adapt the design rather than the training plan. Managed services over self-hosted, even at higher unit cost, because the operational burden is the constraint. Fewer moving parts, even if a more sophisticated design would be cheaper to run in expert hands. Boring, well-documented technology over the better tool their team has never seen.
Where sophistication is genuinely necessary, price the capability as part of the project: runbooks, a handover period with their engineers doing the work while you watch, and a defined support arrangement with an end date and a plan for what happens after.
And say it plainly to the sponsor. "This design assumes two engineers who can operate Kubernetes; you have none, so the options are to hire, to buy managed support, or to choose the simpler design" is a conversation that is uncomfortable once and cheap. Discovering it after go-live is expensive and lands on the customer.
Curated: · Written: · Reviewed:
QA-14How do you decide whether a proof of concept has succeeded?(show answer)
Against the criterion written before it started, and only against that.
The criterion should be a measurement with a threshold: the legacy interface sustains 500 lookups a second at under 200ms p95 for a sustained hour, or eight of ten pilot users complete the task without assistance. Written down, agreed with the customer, and not revised after seeing the result.
That last constraint matters more than it sounds. The strong pull at the end of a POC is to reinterpret the result favourably, because a POC that fails feels like wasted money. It is not — a cheap disconfirmation has just saved the cost of building the wrong thing, and framing it that way with the customer beforehand is what makes an honest result possible.
Three outcomes are worth distinguishing. The assumption holds: proceed, and be explicit that the POC code is not the production foundation. The assumption fails: the design needs to change, and you now know that for the price of a few weeks. The result is ambiguous: usually the criterion was not sharp enough, or the test was not representative — a load test against empty tables, a pilot with unusually motivated users.
The POC should also produce a written record of what was learned beyond the headline result, because the incidental discoveries — an undocumented rate limit, an authentication quirk, real latency numbers — are often worth as much as the answer itself and are otherwise lost when the branch is deleted.
Curated: · Written: · Reviewed:
QA-15A customer has an existing vendor contract that constrains your design. How do you work with that?(show answer)
Find out the shape of the constraint before designing around it, because contracts are usually less absolute than the person quoting them believes.
The questions: when does it end, what does it actually oblige them to use, is there a minimum spend, what are the exit terms, and who owns the relationship internally. A contract with fourteen months left and no minimum volume is a different constraint from a three-year commitment with a penalty. Frequently the obligation is narrower than the assumption — a committed spend rather than a mandate to use a specific product for everything.
Then design within it deliberately. If they are committed to a platform, use it well rather than routing around it, and put the parts that are genuinely a poor fit behind an interface so they can move later without a rewrite.
Where the constraint imposes a real cost, quantify it and put it in front of the sponsor rather than absorbing it silently. "Using the contracted product here adds this much to the annual run cost and this delay; the alternative would require renegotiating clause X" is information they can act on, and someone in the organisation may already be planning that renegotiation.
The commercial timeline is also a design input. If the contract ends in a year, designing a component to be replaceable at that point is worth real effort, and the cutover can be planned rather than forced.
What I would avoid is treating the constraint as an excuse for a design you cannot defend on its own terms.
Curated: · Written: · Reviewed:
QA-16How do you present a cost estimate you know is uncertain without losing credibility?(show answer)
State the range, the assumptions, and what would move it — and be specific about which parts are firm and which are not.
A single number implies a precision you do not have, and when reality lands outside it your credibility goes with it. A range with visible reasoning survives being wrong, because the customer can see which assumption failed and that you had flagged it.
Structure it in confidence tiers. The infrastructure cost for a known workload is firm within perhaps ten percent. The licensing is firm if the volume assumption holds, and the volume assumption is the customer's. The integration effort with an unexamined legacy system might be plus or minus fifty percent, and that is the number to draw attention to rather than bury.
Name the sensitivity explicitly: this total is dominated by the growth assumption, so if volume is double the projection the figure rises by roughly this much, and here is the point at which the design itself would need to change.
Then propose how to narrow it, which converts the uncertainty from a weakness into a plan: a discovery phase, a load test against their real data, or a vendor quote at the actual tier.
And separate one-off from recurring. Customers routinely conflate them, approve a project on the build number, and then discover a run cost nobody highlighted. Presenting the five-year picture with both lines prevents a specific and common failure of trust twelve months in.
Curated: · Written: · Reviewed:
QA-17Question (locked): A customer wants to move an on-premises monolithic application to the cloud with near-zero downtime and no big-bang risk. How do you structure the migration?(show answer)
Assumptions: one monolith, one primary database, a leased link to the data centre, and a business that can take minutes of read-only but not an outage. Those two constraints rule out moving the whole thing in one window. The shape is discovery, per-component decisions, then waves — each with its own cutover and its own tested rollback.
Start with evidence, not interviews. Flow logs, firewall rules, DNS zones, database session lists and the batch scheduler show what actually talks to what; interviews tell you who owns it. Then establish data gravity: which datastore is the system of record, its size, its change rate, and who reads its tables directly. In a monolith the database is the real integration hub, and the reports reading its tables never appear on the diagram. Those hidden readers are what turn a two-week move into a two-quarter one.
Then a strategy per component, with the reason recorded:
| Strategy | When | Here |
|---|---|---|
| Rehost | nothing needs to change | file/print server, an FTP endpoint |
| Replatform | runtime change, code intact | the bulk: app servers to containers, database to a managed service |
| Refactor | something blocks the move | a Windows-only COM dependency, 2 components of 12 |
| Replace | commodity functionality, liability code | CRM, monitoring, three internal tools |
| Retire | a year of access logs showing no use | usually 10–15% of an estate |
| Retain | regulatory or gravity-bound | adjacent batch system, deferred |
Wave zero is the landing zone — network, identity, observability, keys — not a workload, but it lands first. The pilot wave should be small, self-contained and visible, carrying real users on real data: an internal admin tool or the read-only reporting path. Two to three weeks of parallel running there proves the run model and the reconciliation tooling before anything that matters moves. Then order waves by the dependency graph: leaves first, the gravity centre last, so the datastore with the most consumers moves when everything around it is proven.
Bulk load plus continuous replication — DMS or Debezium. Do the arithmetic first: 1.2 TB over a 500 Mbps VPN is about 5.3 hours at line rate, and a source changing 1 GB an hour needs only 280 KB/s of delta to hold the gap. Bandwidth is rarely the failure mode; a nightly job that rewrites 200 GB is. Keep the source as system of record through the wave and replicate forward; at cutover flip the source of truth and replicate back for a 72-hour window, so rollback is a routing change rather than a restore. Two-way active-active I avoid: split-brain writes and schema drift are how these projects lose data.
Each wave gets a runbook with timings in minutes, a go/no-go at a named point, and rollback triggers with numbers attached: error rate above 1% for five minutes, replication lag over 60 seconds at the freeze, any reconciliation mismatch. Freeze writes for ten minutes, apply the final delta, have the business owner — not the project team — run the smoke checks, then release. Rehearse the rollback too: a rollback nobody has executed is a document, not a rollback.
During transition, run both paths and compare: nightly reconciliation of row counts and financial totals to the cent, synthetic transactions against both sides, and a routing layer moving traffic 5%, 25%, 100% with the old path still hot. Observability must span both environments in one pane, or a regression after cutover cannot be attributed to either side of the cut.
With a hard deadline, trade scope and blast radius, never the rollback. Widen the waves, rehost more and refactor less, cut dual-running to days, shrink the rollback window from a week to 24 hours. Put the two consequences in writing — more rework later, more exposure per wave — and let the sponsor choose. If the deadline is a lease ending, the fallback is rehosting the stragglers and modernising in the cloud afterwards.
Compliance changes sequence more than design. Where the target must be assessed before regulated data lands, the landing zone and its evidence become a gate, and that data moves last or stays. Keep the audit trail unbroken across the cut; a gap an auditor finds nine months later is not a saving. Customer-managed keys, access logging, retention, and a rollback path that preserves the chain of evidence. If an assessor signs off each wave, build the waves around their calendar and start that conversation in week one.
Curated: · Written: · Reviewed:
QA-18How do you evaluate two vendors whose feature lists both cover your customer's requirements?(show answer)
Stop comparing feature lists, because both will claim everything, and evaluate against the customer's actual constraints and the things that only surface under use.
Score against weighted requirements derived from their situation, with the weights agreed in advance so the exercise cannot be rationalised toward a preferred answer afterward. Include the non-functional criteria that feature lists never cover: what the price is at their projected volume and at the next tier, what the exit path costs, what the support model actually is in their time zone, and how the vendor behaves during an incident.
Then get evidence rather than claims. A time-boxed trial against their real data and their real integration, exercising the awkward paths — a failure, a restore, a schema change under load, and the observability story when something misbehaves. Reference calls with customers of similar size and industry, asked specifically about their worst experience rather than their general satisfaction.
Look at the vendor as a company: financial stability, release cadence, the quality of their public incident history, and whether their roadmap depends on the customer's use case remaining a priority.
Present a recommendation with the scoring visible, and name what would change it. Customers make better decisions with a transparent comparison than with a conclusion, and the transparency also protects the decision when a stakeholder with a vendor preference challenges it later.
If the two genuinely tie on merit, choose the one that is cheaper to leave.
Curated: · Written: · Reviewed:
QA-19The customer's security team rejects your design late in the process. How do you recover?(show answer)
First, treat the rejection as information rather than an obstacle, and find out what the actual objection is. Late security objections are usually one of three things: a specific control that is missing, a policy the design violates that nobody surfaced earlier, or a lack of evidence that the design does what you say it does. Those need different responses, and assuming the wrong one wastes the remaining time.
If it is a missing control, price the change and its schedule impact and take it to the sponsor as a decision. If it is a policy violation, find out whether the policy admits an exception process and what evidence it requires; many do, and the path is documented rather than adversarial. If it is an evidence gap, the design may be fine and the artifacts are missing — a data flow diagram, a threat model, a statement of where each data class lives.
Recovering the schedule usually means separating the objection from the whole design. Rarely is everything rejected; establish precisely what is blocked so the rest can proceed.
Then fix the process cause, because a late rejection means the security team was engaged too late. Bringing them into the design review at the point where the trust boundaries are drawn costs an hour and prevents this entirely — and their constraints are frequently easier to satisfy when the design is still fluid.
Being visibly non-defensive matters here. The security team's objection is usually correct, and the relationship is worth more than the point.
Curated: · Written: · Reviewed:
QA-20How do you decide the boundary between what you build and what the customer's team builds?(show answer)
By where the knowledge needs to end up, not by what is convenient to hand over.
The components the customer will own and change most often should be built by their team, with you alongside. If they build it, they understand it; if you build it and hand it over, the knowledge transfer is a document that nobody reads until something breaks. Domain logic in particular belongs with them, because it changes with their business and they are the ones who know what it should do.
What is sensible for you to build: the parts that need expertise they do not have and are not trying to acquire, the parts that are genuinely one-off, and the scaffolding that establishes patterns — the first service, the pipeline, the infrastructure modules — which then serve as the worked example their team extends.
The arrangement that works best is pairing on the boundary rather than splitting it cleanly. Their engineers building with yours produces both a system and a team that can run it, which is the actual deliverable for most engagements.
Two constraints shape it in practice. Their capacity is usually the binding one — a team already fully committed cannot absorb this regardless of the plan, and pretending otherwise produces a schedule that fails. And a hard deadline pushes work toward whoever is fastest, which is a real trade against knowledge transfer that should be made explicitly rather than by default.
Whatever the split, name the owner of every component before go-live. Unowned components are how systems decay.
Curated: · Written: · Reviewed:
QA-21A stakeholder challenges your design in front of the room with an objection you had not considered. What do you do?(show answer)
Say that you had not considered it, and think about it out loud. The instinct is to defend, and defending an objection you have not evaluated is how a design review turns into a contest that the design loses either way.
Then establish what kind of objection it is. If it is factual — a constraint you did not know about, a system you were unaware of — the design may need to change and the useful next step is to capture it and say when you will come back with an answer. If it is a difference in judgement about a trade-off, put both sides on the table and identify what evidence would settle it.
What not to do: accept it immediately to avoid conflict, which is as damaging as defending badly. A design that changes with whoever spoke last is not a design, and the room notices.
Take it offline when the answer needs work, but be specific: what you will look at, and when you will report back. "Let me think about it" without a commitment reads as dismissal.
Being visibly comfortable with not knowing is what builds credibility in this role. The customer is deciding whether to trust your judgement on the hundred decisions they will not see, and someone who defends everything is someone whose confidence carries no information.
Afterward, ask why it was not raised earlier. A significant objection surfacing at the review usually means the wrong people were in the earlier conversations.
Curated: · Written: · Reviewed:
QA-22How do you sequence a multi-phase delivery so the customer gets value before the end?(show answer)
Order phases by value delivered and risk retired, not by architectural layers. The tempting sequence — infrastructure, then platform, then services, then features — delivers nothing usable until the end and gives the customer no evidence that any of it works.
The useful shape is a thin vertical slice first: one real user journey, end to end, in production, however narrow. It proves the architecture, exercises the deployment path, surfaces the integration problems, and gives the sponsor something to show. It also converts a large number of assumptions into facts while there is still time to act on them.
Then extend along whichever axis carries the most risk or value: more journeys, more volume, more of the estate migrated.
Two constraints shape the order. Dependencies that cannot be faked — an identity integration or a network path that everything else needs — have to come early even without visible value, and that should be stated rather than hidden. And the customer's own calendar matters: a retailer cannot cut over in November, and a finance system has quarter-end.
Each phase needs a definition of done that stands alone, and a state that is stable if the next phase is delayed or cancelled. Funding gets cut, priorities change, and a phase boundary that leaves the customer with a half-migrated system is a liability you created.
I would also front-load anything that reduces the estimate's uncertainty, because a firmer number for phase three is worth a great deal at the end of phase one.
Curated: · Written: · Reviewed:
QA-23What questions do you ask on the first call with a new customer before proposing anything?(show answer)
Enough to understand the problem, the constraints, and who actually decides — and none that could be answered by reading their website first.
About the problem: what is happening today that should not be, or what cannot they do that they need to? What have they already tried? What made them start looking now, which usually reveals the real driver — an incident, a contract expiry, an audit, a competitor.
About the constraints: what is the timeline and what is driving it? What is the budget range? What systems must this work with, and are any of them off limits? What compliance regime applies? Who will operate it afterward and what do they run today?
About the decision: who signs off, who else must agree, and has anything like this been proposed before and not proceeded? That last question is the most informative one on the call — a previous attempt that failed tells you about the organisation, and the reason it failed is usually still present.
About success: how will they judge whether this worked in a year? If the answer is vague, that is something to fix before proposing anything, because an unstated success criterion becomes a dispute at the end.
I would also ask what they are worried about. People describe risks more candidly than requirements, and the answer often reshapes the proposal more than anything in the formal brief.
Curated: · Written: · Reviewed:
QA-24How do you handle a request for a fixed-price quote on work with significant unknowns?(show answer)
Do not price uncertainty as if it were known, and do not refuse the customer's need for cost certainty either — they usually have a real budgeting constraint behind the request.
The approach that serves both: split the engagement. A fixed-price discovery phase, small and bounded, whose deliverable is a firm estimate and a design for the next phase. The customer gets a real number before committing the large spend, and you price the part you understand. If discovery reveals the project is unwise, they have found that out cheaply, which is worth paying for.
Where a fixed price on the whole thing is unavoidable, be explicit about what it is contingent on. Every assumption becomes a stated condition, and the change process is agreed up front. A fixed price with twenty documented assumptions is honest; a fixed price with the same twenty assumptions unstated is a dispute scheduled for month four.
Price the risk visibly rather than hiding it in padding. A contingency line the customer can see is defensible, and if the risk does not materialise it can be returned — which is a much better position than a padded number that looks expensive against a competitor who has not priced the risk at all.
And be direct about the incentive problem. Fixed price on unknown scope pushes both sides toward arguing about what was included rather than toward the outcome, and saying that out loud usually improves the conversation about how to structure the work.
Curated: · Written: · Reviewed:
QA-25The customer wants to reuse a system they already paid for, and it is a poor fit. How do you make that case?(show answer)
Not as a technical argument, and not by dismissing the investment. Sunk cost is a real emotional and political fact inside their organisation, and treating it as a fallacy to be pointed out is how a good recommendation gets rejected.
Make the case forward-looking: the question is not what was spent, but which option costs less from today. Price both — adapting the existing system, including the ongoing cost of the compromises it forces, against the alternative including migration. Often the existing system loses on the run cost rather than the build cost, and the five-year view makes that visible where the first-year view does not.
Be specific about the misfit rather than general. "It is not designed for this" is an opinion. "It handles 200 transactions a minute and you need 2,000, and the constraint is a single-threaded component we cannot change" is a fact, and it is much harder to argue with.
Look genuinely for the middle option, because there usually is one: keep the system for the workload it does well, and put the new requirement alongside it behind a common interface. That preserves the investment, limits the change, and is frequently the right answer on the merits rather than merely the diplomatic one.
And find out whether someone's credibility is attached to that system. If so, the conversation needs to happen with them privately before it happens in a room, or the technical merits will not decide it.
Curated: · Written: · Reviewed:
QA-26How do you design an integration that must survive the other system being unavailable for hours?(show answer)
Make the boundary asynchronous and durable, so their downtime becomes your backlog rather than your outage.
Writes to them go through a durable queue that you own. Your side accepts the work, persists the intent, and returns; a worker delivers when they are reachable. Hours of unavailability then cost you a growing backlog and nothing else, provided the queue is bounded and monitored on age rather than only depth.
Reads are harder because a caller may be waiting. The options are a local cache of their data with an explicit staleness window, a degraded response that omits what they provide, or failing the specific operation with a clear message. Which one is right is a business decision, and it should be made with the customer before it is needed rather than by whoever writes the catch block.
Everything delivered must be idempotent, because retry after an outage means duplicates are certain. An operation key that they honour, or one you agree with them, is what makes the recovery safe.
Recovery behavior deserves as much design as the outage. A backlog of four hours released at full rate will knock them over again the moment they return, so the drain should be rate-limited to something they can absorb, and messages that are no longer relevant — a status update superseded by a later one — should be collapsed rather than replayed.
And agree the expectations with their team in advance, including how you will tell each other that something is wrong.
Curated: · Written: · Reviewed:
QA-27What makes an architecture diagram useful to a customer rather than decorative?(show answer)
That it answers a question the reader has, at the level they can act on, and that it is honest about what is not yet decided.
The most common failure is one diagram trying to serve everyone: boxes for every component, every arrow, every AWS icon, in a font nobody can read on a projector. It communicates effort rather than meaning.
Layer instead. A context view showing the system, the people who use it, and the external systems it talks to — five to eight boxes, understandable by a finance director. A container view showing the deployable pieces and the main data flows, for the technical audience. Detail views only where a specific decision needs them.
Label the arrows with what flows and how — synchronous, asynchronous, nightly batch — because the arrow direction alone hides the property that matters most for reliability and cost.
Show the trust boundaries. For any customer with a security or compliance interest, where the data crosses a boundary is the thing they most need to see, and it is usually missing.
Mark what is undecided rather than drawing a confident box. A dashed component labelled as an open question invites the right conversation; a solid one implies a commitment that does not exist.
And keep the version and date on it. Diagrams outlive the discussion that produced them, and an undated diagram will be quoted back as a commitment long after the design moved on.
Curated: · Written: · Reviewed:
QA-28How do you approach a customer whose real constraint is political rather than technical?(show answer)
Recognise it explicitly rather than designing around it silently, and work with the person who owns the constraint rather than over them.
The signals are familiar: a requirement with no technical justification that will not move, a stakeholder who blocks without engaging on the merits, or two groups whose designs are both defensible and mutually exclusive. Continuing to make technical arguments in that situation is the most common way this role wastes months.
What usually helps is finding the underlying interest. A team insisting on owning a component may be protecting headcount, or may have been burned by a dependency before. A director rejecting a platform may have committed publicly to a different one. Those are addressable — often by changing how the work is framed, who presents it, or how ownership is distributed — while the stated objection is not.
Give people a way to change position without losing face. New information is the usual mechanism: a load test, a cost model, a pilot result. It lets someone move because the facts changed rather than because they were wrong.
Escalate carefully and rarely. Going over someone converts a disagreement into an enemy, and this role depends on relationships that outlast the project.
And be honest with your own sponsor about what is happening. A schedule slipping because a decision is blocked internally is information they can act on; a schedule slipping for reasons nobody names is one they cannot.
Curated: · Written: · Reviewed:
QA-29What is the most common way a technically sound architecture fails in delivery?(show answer)
Nobody owns it after the people who designed it leave.
The mechanism is consistent. The design is right, the initial build follows it, and then the pressure of delivery meets a team that understands what the components do but not why the boundaries are where they are. Each individual shortcut is reasonable in isolation — a direct database read across a service boundary because the API does not have the field yet, a shared table because the migration is due Friday. Two years later the architecture on the diagram and the system in production are different things, and no single decision was wrong.
The preventions are unglamorous. Record the decisions with their reasoning and their reversal conditions, so a future engineer meets an argument rather than a rule. Encode boundaries where they can be enforced automatically — module import rules, schema ownership, pipeline checks — because a boundary that depends on discipline erodes at the first deadline. Name an owner for each component before handover, and make sure that owner was in the room during the design rather than receiving it.
The second most common failure is closely related: the design assumed an operational capability the customer's team does not have, so it degrades toward whatever they can actually run.
Both are failures of the handover rather than of the design, which is why in this role the transition out is worth as much attention as the architecture itself, and usually gets far less.
Curated: · Written: · Reviewed:
QA-30How do you validate that a proposed design will actually meet the customer's latency requirement?(show answer)
Build the budget from measured components, not from assumptions, and measure the ones you cannot look up.
Decompose the target. If the requirement is 400ms at p95 for a page, allocate: the user's network round trip to the nearest edge, TLS on a new connection, your processing, each downstream call, and the queries beneath those. Writing the allocation down usually reveals immediately that the budget is spent before the interesting work starts.
Then replace assumptions with facts. Latency between the customer's users and the proposed regions is measurable today with a probe. Their legacy system's response time is measurable by calling it. Third-party latency is observable. Query time against representative data volumes takes an afternoon with a prototype and a generated dataset — and against representative volumes matters, because a query on ten thousand rows tells you nothing about ten million.
Pay attention to two shapes. Sequential calls sum, so a chain of five 30ms hops has a 150ms floor before anything else. And fan-out takes the slowest of the parallel calls, so a modest per-call p95 produces a much worse aggregate p95 as the width grows.
Then test under representative load rather than in isolation, because latency at one request a second says nothing about latency at capacity.
If the budget does not close, that is a finding to raise early. Requirements are frequently negotiable when the cost of meeting them is made concrete; discovering the gap during acceptance testing is not.
Curated: · Written: · Reviewed:
QA-31A customer asks you to compare cloud providers for their workload. How do you approach it without a religious answer?(show answer)
Score against their constraints, and be honest that for most mainstream workloads the providers are close enough that the differentiators are not technical.
The things that usually decide it: existing commitments and discounts, which frequently dominate any technical difference; the skills their team already has, since the operational cost of an unfamiliar platform is real and recurring; region availability where their users and their data are legally required to be; specific managed services they depend on; and procurement or contractual relationships already in place.
The genuinely technical differentiators are narrower than vendor marketing suggests, and worth checking against their actual workload: specific service capabilities, certain performance characteristics, the maturity of a particular managed offering, and the pricing model for their dominant cost line — which for data-heavy workloads is often egress rather than compute, and that can differ materially.
The approach: weight the criteria with the customer before scoring, so the exercise is not reverse-engineered from a preference; price their real workload on each rather than comparing list prices for equivalent instances; and check the compliance and data residency requirements early, since those can eliminate an option outright.
Present the scoring, not just the conclusion. And say plainly where the difference is small, because a recommendation that claims a large technical gap where none exists is one an informed stakeholder will discount entirely — along with the rest of your analysis.
Curated: · Written: · Reviewed:
QA-32How do you handle scope creep on a customer engagement without damaging the relationship?(show answer)
Make it visible and priced, immediately and without drama, every time. The relationship damage comes from surprise, not from saying no.
The mechanics: when a new request appears, acknowledge it as reasonable, then state its impact in the currency the customer cares about — cost, schedule, or something else being dropped. "We can do that; it adds about two weeks and this much, or it fits in the current phase if we defer the reporting module. Which would you prefer?" That is a decision for them rather than a refusal from you, and it usually gets a sensible answer.
The failure mode is absorbing small requests silently to be helpful. Each one is genuinely small, none is worth a conversation, and eventually the schedule slips for reasons nobody documented — at which point the customer is surprised and the team is blamed. Absorbing scope quietly feels generous and produces exactly the dispute it was meant to avoid.
Keep a visible list of everything requested and its disposition — in, deferred, or declined with a reason. Reviewing that list at the regular checkpoint means the conversation happens in small pieces instead of once at the end.
Distinguish creep from discovery. A requirement nobody knew about is not the customer being unreasonable; it is the estimate meeting reality, and it should be handled through the same visible process but framed accurately, because calling it creep implies fault that is not theirs.
And be willing to absorb something genuinely trivial. Insisting on a change request for an hour of work is its own kind of relationship damage.
Curated: · Written: · Reviewed:
QA-33What do you include in a handover so the customer's team can actually run the system?(show answer)
More than documentation, because documentation alone reliably fails.
The artifacts: an architecture record with the decisions and their reasoning, not just the current shape. Runbooks for the operations that will actually be needed — deploy, roll back, restore, rotate credentials, respond to each alert — each one tested by someone other than its author. An inventory of every component with its owner. The infrastructure code, with the pipeline that applies it and no path to production that bypasses it. Access and credentials transferred properly, which in practice is the step most often left half-done.
But the part that determines success is the shadowing period. Their engineers doing the work while yours watch, not the reverse. That means their team performing a deploy, a restore, and ideally a real incident before you leave. A restore drill in particular is worth more than any document, because it exercises the whole chain — access, procedure, and the assumption that the backups work.
Also hand over the failure knowledge: what has broken during the build, what the known weak points are, and what you would fix next with more time. That list is uncomfortable to write and is the most useful thing in the package.
And define what happens after. A support arrangement with an end date, a named contact, and an explicit statement of what is no longer covered — so the boundary is clear on both sides rather than discovered during an incident three months later.
Curated: · Written: · Reviewed:
QA-34The customer asks for a feature that would compromise the security of their own users. How do you respond?(show answer)
Raise it clearly and in writing, propose the alternative that meets the underlying need, and do not implement it quietly if they insist.
First understand what they are actually after, because the request is usually a means rather than an end. A request for support staff to see customer passwords is really a request to help users who are locked out, and that is solvable with impersonation tokens, audited access, and a proper reset flow. Most of these requests dissolve once the underlying need is separated from the proposed mechanism.
Where the conflict is real, be specific about the consequence and who bears it. Not "that is insecure" but "this stores card data in a system that is not in scope for your PCI assessment, which puts your compliance status at risk and exposes your customers' data if that system is breached." The consequence landing on their customers and their regulator, rather than on an abstract principle, is what makes the point land.
Put it in writing. This is partly about the customer's record and partly about yours, and a documented recommendation is what protects both when it is reviewed later.
If they still insist and it is lawful, it is their system and their risk — but the decision should be recorded as theirs, made with the consequence stated, and approved by someone with the authority to accept it. If it is not lawful, or it would harm their users in a way I would not want attached to my name, I would decline that piece of the work and say so plainly.
Curated: · Written: · Reviewed:
QA-35How do you keep a design coherent when several teams are building parts of it in parallel?(show answer)
Fix the contracts early and let the internals vary. The parts that must be agreed before parallel work starts are the interfaces between teams, the data ownership, and the cross-cutting conventions — identity, error shapes, logging, correlation IDs. Everything behind a team's own interface is theirs.
Interfaces should be written down and testable rather than described in a meeting. A schema, an example payload, and a contract test each side runs in their own pipeline turns an agreement into something that fails loudly when it drifts, which is far better than discovering the mismatch during integration.
Data ownership needs to be unambiguous before anyone writes code: exactly one team owns each entity and everyone else reads through their interface. Ambiguity here is what produces two services writing the same table and a coupling that is expensive to unpick.
Integrate continuously rather than at the end. A thin end-to-end path working in week two, even with stubs, converts integration from a phase into a habit — and integration phases at the end of parallel workstreams are where schedules die.
Someone has to hold the whole picture and have the standing to arbitrate. That is usually this role, and it works through regular design reviews across teams rather than through a document nobody reads.
And expect divergence anyway. Periodic checks that the built system still matches the intended boundaries — import rules, schema access, actual call graphs from traces — catch drift while it is still cheap.
Curated: · Written: · Reviewed:
QA-36A customer's board asks why the migration will take nine months when a competitor claims six weeks. How do you answer?(show answer)
By making the comparison specific rather than defending the number. The six-week claim is almost certainly for a different scope, and the useful response is to establish which.
The questions that resolve it: does that six weeks include the data migration, or only standing up the new platform? Does it include the integrations with their existing estate, and how many are there? Does it include the parallel-run period where both systems operate? Does it include decommissioning, or does the old system stay funded? Does it assume the customer's team is available, and to what extent?
In most cases the competitor's number is a platform stand-up and yours is an end-to-end migration, which is a real difference rather than a difference in efficiency.
Then be honest about where the time genuinely goes, because it is rarely where boards assume. It is not writing code; it is data reconciliation, discovering undocumented consumers, waiting for access and decisions, and the parallel-run period that exists to prove the new system before the old one is switched off.
If a shorter timeline is genuinely achievable by reducing scope, put that on the table as an option with its consequences: a phased migration that delivers a subset in ten weeks and the rest later, with the cost of running both for longer.
What I would avoid is competing on the number by compressing the estimate. A schedule agreed under that pressure fails visibly at the end, which is worse for everyone than an uncomfortable conversation now.
Curated: · Written: · Reviewed:
QA-37How do you decide whether a customer needs multi-region, given the cost?(show answer)
From their tolerance for a regional outage, expressed in money and in regulatory obligation, rather than from a preference for robustness.
The question to put to them: if this system were unavailable for four hours, what happens? For some customers that is a bad afternoon. For others it is contractual penalties, a regulatory notification, or trading that cannot happen. Those are different architectures and the customer is the only one who knows which they are.
Then be precise that multi-region addresses one specific failure — a whole region becoming unavailable — which is rarer than the failures most outages come from. Their own deployments, a bad configuration, a dependency failure, and single-zone events cause far more downtime in practice, and multi-zone plus good change control addresses those at a fraction of the cost. Recommending multi-region to a customer whose last three outages were self-inflicted is spending their money on the wrong risk.
If it is warranted, distinguish the tiers honestly, because the cost steps up sharply: restore from backups, a running standby whose data loss is however far replication has lagged, and active-active. Active-active still leaves a product decision about conflicts — two sites selling the last item — and a global database offering moves the replication, not that choice.
And whatever tier is chosen, price the drills. An untested failover has an unknown recovery time, so the ongoing cost of rehearsing it is part of the decision, not an optional extra — and it is the line customers most often cut, which is how they end up paying for a capability that does not work.
Curated: · Written: · Reviewed:
QA-38A customer wants their data to never leave their country. What does that actually constrain?(show answer)
More than the region selector, and the gaps are usually in the supporting services rather than the primary datastore.
The obvious part: compute and storage in an in-country region, which may not exist for their provider, and which may lack some managed services or have only one availability zone — a material availability constraint they need to know about before committing.
The parts that are missed. Backups and their replication targets. Log aggregation, monitoring, and error tracking, which frequently default to a vendor's home region and carry payload data including personal information. Support tooling and any screen-sharing or debugging path where an engineer abroad views production data — that is often a transfer in the regulatory sense. CDN and edge caching, where content is copied by design. Email, SMS, and notification providers. Any third-party SaaS in the data path, including analytics.
Then the question of who can access it, which is separate from where it sits: some regimes care about foreign access to in-country data, so an engineering team abroad with production access may be a problem even though the bytes never move.
The practical approach is a data flow map naming every system that touches each data class, with its residency and its access. That is the artifact a regulator or auditor asks for, it is what makes the gaps visible, and it is almost never in place before someone asks for it.
And confirm the requirement's source — law, contract, or preference — because the three admit very different solutions.
Curated: · Written: · Reviewed:
QA-39How do you explain the consistency trade-off to a business stakeholder who is deciding between two designs?(show answer)
In their scenario, with their consequence, and without the theory.
Not "eventual consistency" but: "if the link between our two sites breaks, we can either keep taking orders and risk two customers buying the last item, or stop taking orders until it is fixed. Which is worse for you?" That is the actual decision, it is theirs to make, and stated that way most stakeholders answer immediately and correctly for their business.
The answer usually varies by operation within the same system, which is the more useful insight to give them. Showing an approximate stock level on a product page can be stale; decrementing inventory at checkout cannot. Framing it per-operation avoids a single system-wide choice that is wrong for half the cases.
Quantify the window so it is not abstract. "Under normal conditions the two sites agree within a second; the case we are discussing is the rare failure where that could stretch to minutes" turns an alarming-sounding property into a bounded risk.
Then name the cost of the stronger option in terms they feel: not milliseconds of coordination latency, but that the checkout takes noticeably longer for every customer, every day, to protect against a failure that happens twice a year.
And write the decision down with the reasoning. This is precisely the choice that gets revisited in an incident, and having a record that the trade was made deliberately — and by whom — is what keeps that conversation productive.
Curated: · Written: · Reviewed:
QA-40The customer's existing database is the bottleneck, but replacing it is out of scope. What options do you present?(show answer)
Options ordered by cost and disruption, with the constraint respected rather than argued with, and the point at which it stops working stated plainly.
Cheapest first: find out what is actually slow. In a large share of cases a small number of queries dominate, and indexing or rewriting them buys a substantial factor for days of work rather than a project. That should always be established before anything architectural, because the alternatives are expensive and may be unnecessary.
Then: move read load off the primary. A read replica for reporting and analytics, which frequently accounts for most of the pressure and none of the business-critical latency. Then caching for the genuinely hot, tolerably stale reads, with the staleness window agreed rather than assumed.
Then: reduce the work reaching it. Batching, moving write-heavy paths behind a queue so peaks are absorbed, and archiving cold data out of hot tables, which often improves things more than expected because the working set shrinks.
Then: scale it vertically, which is a configuration change and usually buys a real multiple before the ceiling.
Then, at the boundary of scope: move one high-volume workload to a store that suits it — a search index, a time-series store — while the primary keeps the transactional core.
And state where this ends. "These get you to roughly three times current volume; beyond that the database itself has to change" gives the customer a planning horizon rather than a surprise, and it is the sentence that makes the constraint their informed choice.
Curated: · Written: · Reviewed:
QA-41How do you scope the observability a customer needs, given they are cost-sensitive?(show answer)
Start from the questions they will need answered during an incident, and buy the cheapest thing that answers each. Observability spend gets out of hand when it is specified as a product tier rather than as a set of questions.
The irreducible set: is the system up and are users affected, which is a small number of synthetic checks and a request-level metric; what changed, which is deploy and configuration change annotations and costs almost nothing; and what is failing, which is structured request logging with a correlation ID and errors captured with context.
That covers most incidents at modest cost. The expensive tiers — full distributed tracing at high sampling, long metric retention at high cardinality, every log line kept hot for months — should be bought where they answer a question the cheap tier cannot, not by default.
The specific economies that work: sample traces heavily but keep every trace that errored or was slow, which preserves the interesting tail at a fraction of the volume; keep operational logs hot for a week or two and tier the rest to cheap storage; and be careful with metric cardinality, since a label with unbounded values is the usual cause of a surprising bill.
Retention should be set by category rather than globally — security and audit records kept long because they cannot be reconstructed, verbose debug output kept briefly.
And point out that the cost of not having it is measured in incident duration. A customer with a four-hour mean recovery time is paying for the gap already, just not on that invoice.
Curated: · Written: · Reviewed:
QA-42A customer asks for a single sign-on integration with their corporate identity provider. What do you need to establish first?(show answer)
Which protocol, who owns which part of the user lifecycle, and what happens to people who are not in their directory.
Protocol: SAML and OIDC are both common in enterprise estates, and the answer is usually dictated by what their identity team supports rather than by preference. Establish it early because it affects the library choices and the effort meaningfully.
The lifecycle question is the one that gets missed. Authentication tells you who someone is; it does not create their account, assign their permissions, or remove them when they leave. Provisioning may be just-in-time on first login, or via SCIM, or manual — and deprovisioning is the part with the compliance consequence. A leaver who retains access because the integration only covered login is a finding waiting to happen.
Then authorization: does their directory carry group membership that maps to your roles, and who maintains that mapping? Groups drifting from intended access is a common source of over-permission, so someone has to own the review.
Then the edge cases, which need answers before build: contractors and partners who are not in the corporate directory, break-glass access when the identity provider is down — which makes their availability your availability — and any service accounts that cannot use interactive login.
Finally, get their identity team into the conversation directly and early. They are frequently a separate group with their own change process and lead times, and that queue is a schedule dependency that surprises projects far more often than the implementation does.
Curated: · Written: · Reviewed:
QA-43How do you decide how much of a customer's estate to modernise versus leave alone?(show answer)
By rate of change and by pain, not by age. A twenty-year-old system that works, nobody modifies, and costs little to run is not a problem to be solved; modernising it spends money and introduces risk to produce the same outcome.
The candidates worth changing are where the business needs to move and the system will not let it: components that block a capability the customer needs, that change frequently and are expensive to change, that carry unacceptable risk — unsupported platforms, unpatchable dependencies, a single person who understands them — or whose run cost is genuinely large.
Map their estate on two axes: how often it changes, and how much pain it causes. High on both is where the return is. Low on both is where the correct action is to leave it, document it, and revisit if the situation changes — which is a recommendation customers rarely receive and often need.
Where change is warranted, prefer containment over replacement. A facade in front of the legacy system, with capabilities extracted one at a time, delivers value incrementally and can be stopped at any point. A full rewrite delivers nothing until the end and is the pattern with the worst track record in this industry.
The honest framing for the customer is that modernisation is an investment with a return, not hygiene. For each component, what does changing it enable, and what does it cost? Components that fail that test should be named as deliberate exclusions in the proposal, so leaving them is a documented decision rather than an omission.
Curated: · Written: · Reviewed:
QA-44What is your approach when the customer's requirement is genuinely unachievable within their budget?(show answer)
Say so early, with the arithmetic, and bring options rather than only the problem.
Early matters most. A gap identified in week two is a scoping conversation; the same gap identified in month five is a failed project. The pressure in this role is to defer the awkward conversation in the hope that something changes, and it does not.
Show the arithmetic rather than asserting the conclusion, because the customer needs to be able to check it. This is what the requirement costs, this is why, this is the largest component of it. Frequently one requirement accounts for most of the gap — a very high availability target, a real-time constraint, a compliance regime — and isolating it changes the conversation from "we cannot afford this" to a decision about one thing.
Then bring options with prices. Reduced scope delivering the highest-value subset within budget. A phased approach spreading the spend across budget periods, if their constraint is annual rather than absolute. A relaxed non-functional requirement, which is usually the cheapest lever and the one nobody thinks to question. Or a different approach entirely — buying rather than building, or accepting a manual process where automation is not worth its cost.
And be willing to say the project should not proceed. A customer who cannot fund a viable version is better served by hearing that than by starting something that fails in month eight, and saying it is the thing that makes the rest of your advice credible.
Curated: · Written: · Reviewed:
QA-45How do you handle a customer who wants to see progress weekly on a project with a long lead time?(show answer)
Give them something real every week, and be honest that early weeks produce evidence rather than features.
The underlying need is usually not features — it is confidence that the money is producing something and that problems will surface early rather than at the end. Meeting that need is straightforward once it is named.
What to show in the early phase: discovery findings, which are genuinely interesting to a customer because they are learning about their own estate; decisions made and the options rejected; risks found and retired; and the environment and pipeline standing up, which is visible progress even without a feature.
Then get a thin vertical slice into a demonstrable state as early as the design allows, even if it is narrow and ugly. One real journey working end to end changes the relationship, because from that point every week has something to show and the customer can see the shape of what they are buying.
Be equally consistent about reporting bad news weekly. A status that is green for four months and then red is worse than useless. Small, early disclosure of a slipping dependency is absorbed easily; the same information at the end is a crisis.
Keep the format light. A weekly demonstration of whatever exists, a short written note of decisions, risks, and what is needed from them — that last item matters, because customer-side dependencies are a common cause of delay and a weekly reminder is the cheapest way to manage them.
Curated: · Written: · Reviewed:
QA-46A customer wants an API for their partners. What do you specify beyond the endpoints?(show answer)
The parts that determine whether partners can actually build on it, most of which are not the endpoints.
Authentication and onboarding: how a partner gets credentials, how they are rotated, and how access is revoked. A manual process that takes two weeks per partner becomes the bottleneck on their partner programme, which is a business constraint rather than a technical one.
Rate limits, stated and enforced, with headers exposing the limit, the remainder, and the reset — so a partner can pace themselves rather than discovering the limit by being blocked.
Versioning and a deprecation policy, agreed before launch. How long is a version supported, how are partners notified, and how do you measure who is still using it? A version whose usage you cannot see cannot be retired, and partner APIs accumulate versions faster than internal ones.
Error contracts: consistent shapes, stable codes, and messages a partner developer can act on without contacting support. This has a direct effect on the customer's support cost.
Idempotency on anything that mutates, since partners will retry.
A sandbox with realistic test data, which is the single biggest factor in how quickly partners integrate, and consistently underestimated.
Then the operational commitments: an SLA, a status page, and a support path with response expectations.
And a documentation and change-communication plan with an owner. A partner API is a product with customers, and treating it as an engineering deliverable that ships once is the most common way these programmes stall.
Curated: · Written: · Reviewed:
QA-47How do you compare two designs when one is cheaper to build and the other cheaper to run?(show answer)
Over a horizon the customer agrees to, with the crossover point identified, and with the reversibility of each stated.
Model both across three to five years including the operational cost — infrastructure, licences, and the staff time to run it, which is the line most often omitted and frequently the largest. Then find the crossover: the month at which the cheaper-to-run option overtakes. If that is eighteen months, the decision is usually easy; if it is seven years, the cheaper build is probably right and the run cost is being over-weighted.
Then adjust for the customer's actual situation rather than the arithmetic alone. A startup with eighteen months of runway should optimise for build cost and speed, because the run cost in year four is a problem they will be fortunate to have. An established organisation with a stable, long-lived workload should weight the run cost heavily. Their cost of capital and whether the spend is capital or operating expenditure also matter more than engineers usually expect, and their finance director will have a view.
Then weigh reversibility, which often decides it. The cheaper build that can be replaced later is a low-risk bet; the cheaper-to-run option that locks in a platform or a data model is a larger commitment and should be held to a higher standard of evidence.
Present it as a model with visible assumptions and a sensitivity on the growth rate, so the customer can test their own beliefs against it rather than accepting a conclusion.
Curated: · Written: · Reviewed:
QA-48What do you do when a customer's team disagrees with your recommendation and they will be operating the system?(show answer)
Take it seriously, because they have information you do not and they are the ones who live with the outcome.
Their objection usually comes from one of three places: they know something about their environment that has not been shared, they have been burned by this approach before, or they have a preference formed by familiarity. The first two are substantive and should change your recommendation. The third deserves examination — familiarity is a legitimate factor, since a team operating something they understand will run it better than something they do not, and that is a real engineering consideration rather than a soft one.
The way to find out which it is: ask what specifically concerns them and what would have to be true for the recommendation to be right. If they can name a concrete failure mode, you have learned something. If the answer is general discomfort, the conversation is about familiarity and can be addressed with evidence or with training — or by changing the recommendation, if the gap is large enough.
Where the disagreement persists and it is their system to run, I would weight their preference heavily. A design they believe in and can operate beats a marginally better design they resent and will erode. That is not capitulation; the operability of a system is part of its quality.
Where I would hold the line is on a decision that creates a security, data-integrity, or availability risk they are not seeing. Then it goes to the sponsor as an explicit trade-off with both positions recorded, rather than being settled by whoever is more persistent.
Curated: · Written: · Reviewed:
QA-49How do you approach the first architecture review with a customer whose previous project failed?(show answer)
Find out why it failed before proposing anything, and expect the stated reason to be incomplete.
Ask directly and without judgement: what happened, what did people conclude at the time, and what would they do differently. The stated cause is usually technical — the wrong platform, an underestimate — and the actual cause is frequently organisational: a decision nobody would make, requirements that kept moving, a sponsor who left, a team that was never available. Those causes are still present, and a technically excellent proposal that ignores them fails the same way.
Then design the engagement around whatever the real failure mode was. If requirements moved constantly, the proposal needs a change process and a phased structure with decision points. If the sponsor left, it needs a broader base of support and value delivered early enough to survive a change of leadership. If the team was never available, the plan must reflect their real capacity rather than the stated one, or that dependency must be removed.
Expect scepticism and treat it as reasonable rather than as an obstacle. Credibility here is earned by early, small, visible delivery rather than by a better document — the previous project also had a good document.
Two things to be careful about: do not criticise the previous work, because the people who did it are usually in the room; and find out whether anything from it is salvageable, since writing off an investment they defended is a poor way to begin and occasionally wrong on the merits.
Curated: · Written: · Reviewed:
QA-50The customer wants real-time reporting on operational data. What do you propose?(show answer)
First, establish what real-time means to them, because the range of answers spans three orders of magnitude in cost. In my experience the requirement usually resolves to within a few minutes, and often to within the working day, once the underlying decision it supports is named.
Ask what action the report triggers and how quickly. A dashboard someone looks at each morning does not need streaming. An alert that dispatches an engineer does, but that is an alert rather than a report. A trading decision genuinely needs seconds. Same words, entirely different architectures.
For most cases the honest answer is a read replica plus scheduled aggregation, refreshed every few minutes. It is cheap, it is operable by an ordinary team, it does not touch the transactional path, and it covers a large share of stated real-time requirements.
Where the freshness requirement is genuinely seconds, propose change data capture into a store shaped for the queries — a stream into a columnar or search store — and be explicit about what that adds: a pipeline to operate, a consistency window to explain, and a materially higher run cost and skill requirement.
The mistake to avoid either way is reporting queries against the transactional primary. It is the cheapest thing to build and the most common cause of an outage six months later, when a report someone wrote grows past what the primary can absorb during business hours.
Whatever is chosen, agree the freshness as a number and show it on the dashboard, so users can see how current the data is rather than assuming.
Curated: · Written: · Reviewed:
QA-51How do you make sure a design review with the customer produces decisions rather than discussion?(show answer)
Send the material in advance, state the decisions being asked for, and make sure the people who can make them are in the room.
The most common failure is a review where the design is presented for the first time. The audience spends the session absorbing it, asks clarifying questions, and leaves with an action to think about it — which means another session and a two-week delay. Circulating a document beforehand, with the specific decisions listed at the top, converts the session into a decision meeting.
Name the decisions explicitly: we are asking you to confirm the availability target, approve the integration approach with the billing system, and choose between these two options for reporting. Three named decisions with their consequences priced is a productive hour. "Here is the architecture, thoughts?" is not.
Check attendance against the decisions. If the security team must approve the trust boundaries and they are not there, that decision will not be made regardless of how well it is presented, and it is better to know that beforehand.
Time-box the discussion of each and be willing to park what cannot be resolved, with an owner and a date. A single unresolvable point can consume a session otherwise.
Record decisions in the room and circulate them the same day, including what was decided against and why. Undocumented decisions get relitigated, usually by whoever was absent, and the reasoning is what stops that from being a fresh argument each time.
Curated: · Written: · Reviewed:
QA-52A customer wants to know what happens if you get hit by a bus. How do you answer?(show answer)
By showing them that the answer does not depend on me, which means the engagement has to be structured that way from the start rather than described that way at the end.
Concretely: the decisions and their reasoning are written down in the repository, not in my head. The infrastructure is code, applied through a pipeline, so the environment can be rebuilt without anyone's private knowledge. The runbooks exist and have been executed by someone other than their author. Their engineers have been building alongside mine rather than receiving the result. There is a named counterpart on their side for each area, and a second person from my side who has been involved enough to continue.
That is a better answer than a contractual continuity clause, because it is verifiable — they can look at the repository.
The question is also worth taking at face value as a concern about key-person risk generally, including on their side. It is a reasonable moment to point out that the same risk exists for whoever on their team is the only person who understands the legacy system, and that the handover plan should address that too. Customers usually recognise the pattern immediately once it is named.
And if the honest answer is that the project does depend on me, that is a finding to act on rather than to reassure away — it means the knowledge transfer has not been happening, and it is much cheaper to correct in month three than in month nine.
Curated: · Written: · Reviewed:
QA-53How do you evaluate whether a customer's proposed timeline is realistic before agreeing to it?(show answer)
Work backward from the deadline through the dependencies that are not under your control, because those are what actually break timelines.
Estimate the work honestly first, then add the things people leave out: the customer-side decisions and how long they historically take, access and environment provisioning, security review and any penetration test with its remediation window, procurement and contract signature, data migration and reconciliation, the parallel-run period, user training, and any change-freeze windows their organisation observes.
That list frequently accounts for more elapsed time than the build, and it is where realistic and optimistic plans diverge. A customer whose security review takes six weeks and who has a December change freeze has a much shorter available window than the calendar suggests.
Then check the critical path for the dependencies that cannot be compressed by adding people, and test the assumptions behind each with the customer directly. "This assumes firewall changes take a week — is that right?" often produces a very different number from someone who knows.
If the timeline does not fit, present the arithmetic rather than an opinion, and bring options: reduced scope for the fixed date, a phased delivery hitting the date with a subset, or a later date. Deadlines usually exist for a reason — an event, a contract, a regulatory date — and knowing which determines whether it can move at all.
What I would not do is agree to a date I do not believe. The credibility cost of a missed commitment is larger than the discomfort of the conversation now.
Curated: · Written: · Reviewed:
QA-54What is the role of a reference architecture inside a consulting practice, and how do you keep it from becoming dogma?(show answer)
Its value is that it encodes what the practice has learned, so each engagement does not re-derive the same decisions and each customer gets the benefit of the last ten projects. That is real, and teams without one produce inconsistent quality that tracks whoever happened to be staffed.
It becomes dogma when it is applied without the reasoning. The failure is recognisable: a customer with a modest workload receiving a design sized for a much larger one, because that is what the reference says, and nobody can explain which requirement justifies each component.
The things that keep it honest. Record the assumptions alongside each decision — this pattern assumes high write volume, this one assumes a regulated environment — so applying it requires checking whether the assumption holds. Make divergence normal and cheap: an engagement that departs from the reference should record why, and those records are the most valuable input to the next revision.
Review it on a cadence against what engagements actually did. If most projects diverge at the same point, the reference is wrong there and should change.
Keep it modular rather than monolithic, so a team can adopt the identity pattern without inheriting the whole stack.
And treat a junior engineer asking why as a signal to check rather than a gap in their knowledge. A reference architecture that cannot be explained from first principles by the person applying it has stopped being knowledge and become ritual, and it will eventually be applied to a customer it does not fit.
Curated: · Written: · Reviewed:
QA-55How do you handle a situation where the customer's data quality is worse than anyone assumed?(show answer)
Quantify it before deciding anything, then treat it as a scope conversation rather than a technical problem to absorb quietly.
Quantify concretely: what proportion of records violate each constraint the new system needs — missing required fields, duplicates, broken references, values outside their domain, inconsistent formats and encodings. A profiling pass over the real data takes days and turns a vague concern into a table the customer can act on.
Then separate the categories, because they have different owners. Some are technical and fixable in the migration: normalising formats, deduplicating on a reliable key. Some require business decisions nobody has made: which of two conflicting addresses is correct, what to do with orders whose customer no longer exists. Those cannot be resolved by the delivery team and will stall the migration if they are not escalated early with a named decision-maker.
Present the options with costs: cleanse before migrating, which is the most work and gives the best outcome; migrate as-is into a system that tolerates it, which defers the problem and often makes it worse; or migrate the clean subset and handle the remainder through a separate process, which is frequently the pragmatic answer.
Be explicit that this changes the estimate. Data quality is one of the most common causes of migration overrun, and absorbing it silently produces a late project with no explanation the customer can accept.
And put quality checks into the pipeline so the state is visible continuously rather than discovered again at cutover.
Curated: · Written: · Reviewed:
QA-56A customer asks whether they should put a cache in front of their database. What do you establish before answering?(show answer)
What problem it is meant to solve, because a cache is a permanent operational commitment being proposed as a performance fix, and half the time the underlying problem is cheaper to fix directly.
Establish first whether the slowness is a small number of queries. If it is, indexing or a rewrite removes the need entirely, for days of work rather than a new component with a staleness contract and an invalidation problem.
If a cache is warranted, the questions the customer has to answer are business questions rather than technical ones. How stale may this data be, per data type, and who accepts that? Which writes must invalidate it, including the ones that happen outside the application — a batch job, an admin tool, a support person editing a record directly? What should users see if the cache is unavailable?
That last one is the question that most often changes the design. If the database cannot serve full traffic without the cache, the cache is a single point of failure and its own availability now sets theirs. The number to establish is what the origin can serve at a zero hit rate, and it should be measured rather than assumed.
Then the multi-tenant trap, which is worth raising explicitly for any customer with distinct client organisations: a cache key that omits the tenant serves one client's data to another, silently, for as long as the entry lives.
If they cannot answer the staleness and invalidation questions, the cache is premature and I would say so.
Curated: · Written: · Reviewed:
QA-57How do you explain to a customer why their reporting queries should not run against the production database?(show answer)
In terms of the risk to the thing they care about, with a concrete mechanism rather than a principle.
The mechanism: reporting queries are large, unpredictable, and written by people optimising for the answer rather than for the cost. One of them scanning a large table consumes the same memory and IO the transactional workload needs, and the effect is not that the report is slow — it is that checkout is slow, during business hours, for reasons that are hard to attribute because the report looks unrelated.
It also creates a coupling nobody intends. A query someone wrote for a monthly board pack becomes load-bearing, and then the schema cannot change without breaking a report that no engineer knows exists. That is a real constraint on their ability to evolve the system, and it accrues quietly.
The remedy is proportionate to their situation. For most customers a read replica is enough: cheap, straightforward, and it removes the contention entirely. Where the reporting is heavy or the shapes are analytical, a separate store fed on a schedule suits the queries better and costs less to run than making the transactional database do both jobs.
One caution worth stating if they already use a replica for recovery: heavy reporting raises its replication lag, which quietly degrades the recovery objective they think they have. If the same replica is both the reporting target and the failover target, those two purposes are in conflict and they should know it.
Curated: · Written: · Reviewed:
QA-58A customer's checkout occasionally double-charges. They want to know how the architecture prevents it. What do you tell them?(show answer)
That preventing it is a property of the payment path's design rather than of care taken by developers, and then show them where the property comes from.
The mechanism is an idempotency key. The client generates one identifier for the logical purchase — not per attempt — and sends it with the request. The server keeps that identifier under a uniqueness constraint so a retry is recognised as the same purchase rather than a new one.
What I would not tell them is that the key and the charge simply go in one database transaction. That is true only when "charge" means a row they own. In checkout the money is taken by a processor they do not own, and a local commit cannot stand in for that processor's own record. Leaving a transaction open while waiting on that call is the wrong shape as well: they pay the processor's latency on a held connection, and a crash after the processor has accepted the payment still has no local row that prevents a second attempt.
So the path I would show them is: persist the key as soon as the request arrives, pass that same key through to the processor, then record what the processor returned. If the process dies in the gap, the retry presents the same key and the processor returns the original payment instead of taking the card again. If two clicks arrive together, one of them must own the key and the other must wait for its result — a uniqueness error sent back as a failure is how the shopper tries again under a new key and is charged twice.
Then the honest framing of why retries happen at all: the client cannot tell a lost request from a lost response, so it must retry, and the payment provider does the same to you. At-least-once behavior is not a bug to be eliminated but a property to be designed for.
I would also ask what their current symptom actually is, because double-charging sometimes has a simpler surface — a button with no client-side guard, or a reconciliation job re-submitting settled transactions. Client-side debounce helps the shopper; it does not replace the server-side key. Reconciliation needs the same identifier, or a uniqueness check on the settlement, or it will create the duplicates the architecture was meant to prevent.
Curated: · Written: · Reviewed:
QA-59The customer wants an audit trail for a compliance requirement. What do you design?(show answer)
Something that records changes regardless of who made them, cannot be quietly altered, and can be searched within the timeframe an auditor expects.
Regardless of who made them is the requirement that shapes the design. An audit trail written by the application misses the migration, the admin tool, and the manual fix applied during an incident — which are precisely the changes an auditor asks about. Capturing at the database level, through triggers or a change data capture stream, catches all of them. That is one of the few cases where I would put logic in a trigger, and the reason is exactly this.
It must record who, not just what. That means no shared accounts and no service acting on a user's behalf without carrying the original identity through. Where a support agent acts for a customer, both identities belong in the record.
Tamper-evidence: append-only storage with a write-once retention policy, and permissions arranged so that no identity can both change data and delete the corresponding record. Hash chaining makes silent deletion detectable if the regime warrants it.
Retention set to the regulatory window, which is often years, so tiering to cheap storage while preserving the guarantees is part of the design rather than an afterthought.
And it must be queryable by subject and by time. An audit trail that exists but takes a week to search fails the requirement in practice, so the retrieval path deserves testing as seriously as the write path — ideally by running a realistic request before the auditor does.
Curated: · Written: · Reviewed:
QA-60How do you decide whether a customer's workload should use a relational database or something else?(show answer)
From the access patterns and the integrity requirements, and with a strong default toward relational unless something specific rules it out.
The default exists for a practical reason in this role: the customer's requirements will change, and a relational schema absorbs an unanticipated query as a new SELECT rather than as a data migration. For a system whose second year is not fully specified — which is most of them — that flexibility is worth more than the performance characteristics being traded away.
What justifies something else. A document store when the aggregate is the unit of access, records genuinely vary in shape, and cross-document queries are rare. A key-value store for high-volume, simple lookups such as sessions. A search index when the requirement is relevance-ranked text search, which relational engines do poorly. A time-series store for high-volume metrics with time-window queries. A graph store when variable-depth traversal is a dominant pattern.
Note that most of these are additions rather than replacements. The usual right answer for a customer is a relational system of record with a specialised store alongside it for the one workload that needs it, fed asynchronously — not choosing one database for everything.
Two things I would push back on. Choosing a document store to avoid designing a schema, which moves the schema into application code where nothing enforces it. And choosing for anticipated scale, since a well-indexed relational primary handles far more than most customers estimate, and the operational cost of a second technology starts immediately.
Curated: · Written: · Reviewed:
QA-61A customer is worried about vendor lock-in. How do you address it proportionately?(show answer)
By identifying where the lock-in actually is, since the concern is usually attached to the wrong thing.
Compute is rarely the problem. Containers and virtual machines move without much difficulty, and the effort is measured in weeks. What genuinely locks a customer in is data gravity — the cost and risk of moving large volumes, and everything integrated with it — along with proprietary managed services with no equivalent elsewhere, and the skills their team has built.
So the proportionate response is to be deliberate where the cost is high and relaxed where it is low. Keep data in portable formats and standard engines where that is reasonable. Put genuinely proprietary services behind an interface you own, so a replacement is a module rather than a rewrite. And use provider-specific services freely where they earn their keep and the exit cost is modest.
What I would argue against is a full abstraction layer designed for portability. It is a permanent tax — more code, the intersection of both providers' capabilities, and a layer that has to be maintained — paid every day against a migration that usually never happens. Teams that build it typically end up with the worst of both: the abstraction's cost, and provider-specific behavior leaking through it anyway.
The reframing I would offer the customer: the useful question is not how to avoid lock-in but what an exit would cost and whether that is acceptable. Estimate it honestly for the two or three largest commitments. Usually the answer is that a focused migration is cheaper than the insurance policy.
Curated: · Written: · Reviewed:
QA-62How do you scope disaster recovery for a customer who has never tested a restore?(show answer)
Start by testing one, before designing anything, because the current state is unknown and frequently worse than assumed.
A restore drill on their existing backups answers questions no design conversation can. Do the backups complete? Can they be restored into a usable system? How long does it take? Does the application start against the restored data? Who has the access required, and is that person available? In my experience this exercise finds a problem more often than not — a backup that has been failing silently, a restore that takes far longer than anyone believed, or credentials nobody can locate.
That result becomes the baseline for the conversation, and it is far more persuasive than a diagram. A customer who has just watched a restore take eleven hours engages with recovery objectives quite differently.
Then set the objectives from the business consequence: how much data can they afford to lose, and how long can they be down, expressed with the cost of each. Those two numbers select the tier — restore from backups, a running standby, or active-active — and the cost steps up sharply, so the numbers deserve scrutiny rather than aspiration. Nightly backups cannot meet a minutes-level loss budget; a standby's loss is the lag at the moment of failure.
Then design to the objectives, and price the drills as an ongoing commitment rather than a one-off. An untested plan has an unknown recovery time, so the rehearsal is not optional extra work; it is the thing that makes the objective real.
And include the recovery of the recovery path: what if the credentials or the runbook live in the system that is down?
Curated: · Written: · Reviewed:
QA-63The customer asks why they need infrastructure as code when their environment rarely changes. How do you make the case?(show answer)
Not on change frequency, because their premise is reasonable. Make it on recovery, review, and the fact that the environment is currently undocumented.
Recovery is the strongest argument. If the environment were lost — a region event, a mistaken deletion, an account compromise — can they rebuild it? Without infrastructure as code the answer is that someone reconstructs it from memory and a diagram, under pressure, and the result differs from the original in ways nobody notices until later. With it, the rebuild is a pipeline run.
Review is the second. A console change is made by one person with no record of what it was, who approved it, or why. The same change as code is reviewed before it happens and recorded afterward, which is also what most compliance regimes are actually asking for when they ask about change management.
The third is that infrastructure as code is the only documentation of an environment that cannot go stale, because it is the thing that produced it.
And there is a rejoinder to the premise: environments that rarely change are the ones where nobody remembers how they were built, so the knowledge is at its most fragile precisely where the change rate is lowest.
The honest cost is the initial import of existing resources, which is tedious, and a discipline that console changes stop. I would propose starting with the next new component rather than a big-bang import, so the practice establishes itself without a project.
Curated: · Written: · Reviewed:
QA-64How do you handle a customer requirement for data deletion on request?(show answer)
By mapping where the identifiers actually propagate, because deletion fails at the copies nobody listed rather than at the primary record.
The inventory is the deliverable: primary datastore, read replicas, backups, search indexes, caches, the analytics warehouse, log aggregation, error tracking, email and messaging providers, the CRM, and any third-party tool that received the data. Most customers have not written this down, and building it usually reveals systems nobody in the room knew were receiving personal data.
Then design deletion as a job that covers each of them, with a record of what was deleted and when — because demonstrating compliance requires evidence, not intent.
The hard cases need explicit decisions rather than defaults. Backups cannot be selectively edited, so the standard position is a documented retention window with deletion applied on restore, and that should be written down and agreed rather than assumed. Data that must be retained for another legal obligation — financial records, fraud investigation — is an exception with a legal basis, and the design needs to keep it while deleting the rest. Analytics data can often be anonymised rather than deleted, provided the anonymisation is genuine.
Then decide between soft and hard deletion deliberately, since a soft delete that leaves the data readable is not deletion in the regulatory sense.
And build the request path itself: how a request arrives, who verifies the requester's identity, and what the timeline is. That process is usually the part that is missing.
Curated: · Written: · Reviewed:
QA-65A customer's peak traffic is 40 times their average, one day a year. How do you design for that?(show answer)
By separating what must scale from what can be shed, and by not sizing the whole system for one day.
First establish the shape precisely: is it 40 times for an hour or for a minute, and which operations does it affect? Frequently the spike hits a narrow path — a product page, a queue join, a payment — while most of the system sees normal load. Designing the whole estate for the peak when a fifth of it needs to scale is how these projects become unaffordable.
Then pre-provision rather than relying on autoscaling for the event itself. Autoscaling responds in minutes and the spike arrives in seconds; on a known date, the capacity should be in place beforehand and the caches warm. Autoscaling remains useful for the recovery and for the unknown days.
Then design the overload behavior, because a 40 times estimate will be wrong in one direction or the other. A queue with admission control, a waiting room for the entry path, rate limiting at the edge, and a defined degraded mode — read-only, or a static page for the parts that cannot serve — turn an underestimate into a slow experience rather than an outage. Customers usually accept queuing far better than errors, and that is a product conversation worth having in advance.
Then load test to failure against representative data, not to confirm the estimate, and rehearse the day: who is watching, what the thresholds are, and who can pull a feature flag.
And check the dependencies. Their payment provider and their own suppliers need to know the date too.
Curated: · Written: · Reviewed:
QA-66How do you decide what to monitor for a customer who currently monitors nothing?(show answer)
Start with the smallest set that answers whether users are affected, and add only what a real incident shows you needed.
The starting set is short. A synthetic check on the main user journey, from outside their network, which is the single most valuable signal because it detects the outage the internal metrics miss. Request rate, error rate, and latency as a distribution for the user-facing endpoints. Saturation on whatever is scarcest — usually the database connection pool and disk. And deploy and configuration change annotations, which cost nothing and correlate with a large share of incidents.
That is enough to know something is wrong and roughly where, which is a transformative improvement over nothing.
Then alerting, kept deliberately small: page on user-visible symptoms only. A customer new to this will be tempted to alert on everything, and the result is an ignored channel within a month. Two or three alerts that are always real is a working system; thirty is noise.
Then add the leading signals with a deadline attached — certificate expiry, disk filling, backup failure — because each is an outage with a countdown and none of them shows up as a symptom until it is too late.
After that, let incidents drive it. Every incident should end with the question of what signal would have shown this sooner, and the answer becomes the next thing instrumented. That produces a monitoring setup shaped by their actual failures rather than by a vendor's checklist, and it stays small enough to be read.
Curated: · Written: · Reviewed:
QA-67A customer wants to expose their internal API to a mobile app. What changes?(show answer)
Almost everything about its assumptions, because an internal API is protected by not being reachable and a mobile client removes that.
Trust: an internal API often assumes callers are well-behaved. A mobile app's traffic is fully controllable by anyone who owns the device — the client can be decompiled, requests can be replayed and modified, and any secret shipped in the binary is public. So every authorization decision must be enforced server-side, and object-level ownership checks in particular, since the app can request any identifier it likes.
Versioning: internal clients are deployed with the server. Mobile clients are not, and old versions persist for months because users do not update. That makes every breaking change expensive and means a deprecation policy and per-version usage measurement are needed from the start rather than later.
Efficiency: an internal chatty API becomes a battery and latency problem over mobile networks. Endpoints usually need to be reshaped so one screen is one request, and payloads trimmed.
Failure behavior: mobile networks fail constantly and partially. The API must be safe to retry — idempotency on mutations — and tolerant of requests arriving late or out of order after a period offline.
Then the operational additions: rate limiting per user and per device, abuse controls on unauthenticated endpoints, and a way to force-upgrade clients whose version has a serious problem.
I would treat it as designing a public API with a single known consumer, not as exposing an existing one.
Curated: · Written: · Reviewed:
QA-68How do you present a risk register to a customer without it becoming a formality nobody reads?(show answer)
Keep it short, make each entry concrete, and bring it to a decision each time rather than reporting it.
Short means the risks that could actually change the outcome — usually five to eight. A register of forty entries is a document that has been optimised for coverage rather than for use, and the effect is that nobody reads any of them, including the three that matter.
Concrete means each risk names a specific event, its consequence in the customer's terms, and what would be done. Not "integration complexity" but "the billing system's API has undocumented rate limits; if they are lower than our peak, the nightly run will not complete in the window, and we would need to negotiate a limit increase or move to a batch file, which adds about three weeks."
Each entry needs an owner and a next action, and several will be owned by the customer rather than by you — decisions pending, access not granted, a dependency on another of their programmes. Making those visible is often the register's main value, because they are the risks the delivery team cannot address alone.
Review it at the regular checkpoint by exception: what changed, what closed, what is newly urgent. Reading the whole list aloud is what trains people to stop listening.
And close risks explicitly when they are retired. A register that only grows reads as a list of complaints; one that visibly closes items demonstrates that raising a risk leads somewhere, which is what keeps people raising them.
Curated: · Written: · Reviewed:
QA-69The customer asks whether microservices are right for them. How do you answer?(show answer)
By asking what problem they expect it to solve, because the honest answer for most customers is not yet, and the reasons they usually give point to cheaper solutions.
The problems microservices genuinely address: independent deployment when multiple teams are blocked on a shared release train, independent scaling when one component needs far more capacity than the rest, and fault isolation when one component's failure must not take the others down.
If they have one team, none of those apply. The coordination cost that microservices remove does not exist, and they will pay the distributed-systems cost — network failures between components, no cross-component transactions, distributed debugging, and an operational surface several times larger — in exchange for nothing.
The reasons customers usually cite are worth examining individually. Deploying is slow: usually a pipeline and test-suite problem, cheaper to fix directly. The codebase is hard to work in: usually a modularity problem, and modules with enforced boundaries inside one deployable give most of the benefit at none of the cost. Scaling: usually one component, which can often be extracted alone.
So my typical recommendation is a well-modularised single deployable, with boundaries enforced by tooling, and extraction of the specific components that have a real independent scaling or deployment need. That keeps the option open, because clean internal boundaries make a later split straightforward, while an unnecessary split is expensive to reverse.
If they have several teams blocking each other, that is the case where I would agree.
Curated: · Written: · Reviewed:
QA-70How do you handle a customer who wants to lift and shift to the cloud with no changes?(show answer)
Support it as a first step if their driver justifies it, and be explicit that it will cost more to run than their data centre unless something follows.
Lift and shift is a legitimate strategy when the driver is a deadline — a data centre lease ending, hardware end of life, an acquisition — because it is the fastest way out and it decouples the move from the modernisation. Trying to redesign under that deadline is how both fail.
What must be said plainly is the economics. Virtual machines sized like physical servers, running continuously, with no autoscaling and no managed services, typically cost more per month than the equivalent owned hardware. The cloud's savings come from elasticity and from operational work you stop doing, and lift and shift captures neither. A customer who expects an immediate saving will be unhappy in month three, and that expectation should be corrected before the contract rather than after the first invoice.
So the recommendation is lift and shift with a named next phase and a business case for it: right-sizing based on measured utilisation, which usually recovers a substantial share; scheduling non-production environments off outside working hours; moving the databases to managed services, which is where most of the operational saving is; and replacing the components with the worst run cost.
Two things to do during the move even under time pressure, because they are far more expensive to retrofit: get the network topology and identity model right, and put the environment into infrastructure as code as it lands.
Curated: · Written: · Reviewed:
QA-71A customer's system works but nobody can change it safely. What do you recommend?(show answer)
Build the ability to change it before changing it, in that order, because every improvement attempted without that ability adds risk to a system already producing fear.
The first thing is a safety net that does not require understanding the code: characterisation tests. Capture the current behavior of the important paths — inputs and outputs, including behavior that is arguably wrong — as automated tests. They assert nothing about correctness; they detect unintended change, which is the actual constraint.
The second is a deployment path that can be reversed quickly. A team that cannot roll back will not change anything, and no amount of encouragement alters that. A pipeline with a tested rollback converts a change from an irreversible commitment into an experiment.
The third is enough observability to know whether a change worked, which for many such systems means starting from almost nothing: error rate, latency, and the key business metric visible within minutes of a deploy.
Only then start changing, beginning somewhere low-risk to prove the loop works.
The organisational half matters as much. A team that has been punished for outages will not take risks regardless of tooling, so the response to the first failure under the new process determines whether the process survives.
I would also expect to find that some fear is well-founded — a component with no tests and real consequences. Naming those explicitly and treating them differently is better than a blanket claim that change is now safe.
Curated: · Written: · Reviewed:
QA-72How do you decide the right level of automation for a customer's deployment process?(show answer)
From their deploy frequency and their tolerance for a bad release, and by automating the steps that are error-prone before the steps that are merely tedious.
For a customer deploying monthly with a change window, a fully automated continuous deployment pipeline is more machinery than the situation warrants, and it will not be trusted or maintained. What they need is the build, the tests, and the deployment itself automated and repeatable, with a human deciding when it runs.
The steps worth automating first are the ones where a human error is likely and consequential: database migrations, configuration differences between environments, and the deployment sequence itself. Manual steps in those places are where releases go wrong, and each one is also an undocumented dependency on the person who knows it.
Rollback should be automated before anything else, because it is the control that makes everything else safe, and it is the step most often left manual — which means it is slowest exactly when it is needed.
Then let frequency drive further investment. A team moving to weekly deploys benefits from automated verification and progressive rollout; a team deploying quarterly does not yet.
The trap to avoid is copying a mature practice wholesale. A customer with no tests who builds a continuous deployment pipeline has automated the delivery of untested changes, which is worse than what they had. The sequence is tests, then repeatable deployment, then rollback, then frequency, then progressive delivery — in that order, because each depends on the one before it.
Curated: · Written: · Reviewed:
QA-73A customer asks you to review an architecture their previous supplier delivered. How do you approach it?(show answer)
As a fitness assessment against their current requirements, not as a critique, and with the incentive problem acknowledged openly.
The incentive problem is real: I am a supplier being asked to evaluate a competitor's work, and a customer who has been through that before will discount everything I say if I only find fault. Saying so at the start, and being specific about what is good, is what makes the criticism credible.
The method: establish the requirements the design was built against, which may not be the ones they have now, because much of what looks wrong in an inherited architecture was correct for a constraint that has since changed. Then assess against today's requirements — does it meet the current load, the current availability need, the current compliance boundary — and separately assess its condition: is it operable by the team that has it, is it patched, is it documented, can it be changed.
Then categorise findings by consequence rather than by severity language: what is a genuine risk to them now, what will become a problem within a year, and what is a difference of style. Style opinions belong at the bottom or not at all; they are what makes a review read as posturing.
For each real finding, give an option and a rough cost, so it is actionable rather than an assessment.
And check whether the customer's real question is technical. Sometimes it is whether to continue with that supplier, and that is worth surfacing directly.
Curated: · Written: · Reviewed:
QA-74How do you specify non-functional requirements so they can actually be tested at acceptance?(show answer)
Each one as a measurement, a threshold, and the conditions under which it is measured. Anything missing one of those three cannot be tested, which means it cannot be accepted or disputed on evidence.
Performance: not "the system will be responsive" but "the order search returns within 800ms at p95, measured at the load balancer, with 200 concurrent users and the production data volume." The load and the data volume are the parts most often omitted, and they are what make the test meaningful — a latency figure against an empty database proves nothing.
Availability: the indicator, the target, the measurement window, and what counts as unavailable, including whether planned maintenance is excluded.
Capacity: the peak transaction rate to sustain, for how long, and the growth horizon the design must accommodate without redesign.
Recovery: the recovery time and data-loss objectives, plus the requirement that they are demonstrated by a drill rather than asserted, which turns a claim into an acceptance test.
Security: specific rather than general — the authentication mechanism, encryption in transit and at rest, the access review process, and the penetration test with an agreed remediation standard for findings.
Operability: what the customer's team must be able to do unaided, which is testable by having them do it.
Then agree who runs the test, in which environment, and with what data. A non-functional requirement tested in an environment a tenth the size of production has not been tested, and that argument at acceptance is entirely avoidable by settling it in the statement of work.
Curated: · Written: · Reviewed:
QA-75The customer wants to know the security posture of the design in one page. What goes on it?(show answer)
The data, the boundaries, the controls at each boundary, and the honest gaps — organised so a security reviewer can find what they need rather than as a list of features.
Data classification first, because it determines everything else: what classes of data exist, where each is stored, and where each flows. A single diagram showing personal, financial, and operational data with their locations answers most of a reviewer's opening questions.
Trust boundaries next: where does data cross from one zone of trust to another — internet to edge, edge to application, application to data, and every third party. Each crossing is where a control belongs.
Then the controls at each boundary, stated concretely: how callers are authenticated, how authorization is enforced, what is encrypted in transit and at rest, how keys are managed, and how credentials are issued and rotated. Naming the mechanism rather than the property matters — short-lived federated workload identities is a fact, "we use least privilege" is a claim.
Then detection and response: what is logged, where the logs live, how long they are kept, and what generates an alert.
Then the compliance mapping if a regime applies, showing which control satisfies which requirement, since that is the artifact the auditor actually wants.
And the gaps, deliberately included: what is not yet in place, what is accepted risk, and what depends on the customer. A one-page posture with no gaps is not credible to anyone experienced, and including them is what makes the rest believable.
Curated: · Written: · Reviewed:
QA-76A customer's workload has to survive the loss of an entire cloud region. How do you choose between backup-restore, pilot light, warm standby and active-active, and how do you tie the choice to RTO/RPO and cost?(show answer)
Pin the failure model and the numbers first, because the strategy is an output of them, not a menu item to pick from. Two figures come before any design: how long the business can be down (RTO) and how much data it can afford to lose (RPO) — and I ask for them per workload tier, not as one number for the estate. Checkout and the reporting cluster almost never deserve the same treatment, and the cheapest way to cut DR spend is to stop paying warm-standby money for components that can wait.
Then the four options, with the planning ranges I'd put in front of a customer. These describe what the design delivers, not what a provider's SLA promises, and they are only real once the customer has measured them:
| Strategy | RTO | RPO | Extra steady-state cost | What has to be true |
|---|---|---|---|---|
| Backup and restore | 4–24 h | Backup interval, often 12–24 h | ~5–15% of primary run cost | Backups copied out of region, a restore target that exists (account, network, IAM), and a measured restore time |
| Pilot light | 30 min – 3 h | Near zero if the data tier replicates continuously | ~20–40% | Data services running in the second region; everything else is code, images and no capacity |
| Warm standby | 5–30 min | Seconds | ~50–100% | A scaled-down but complete environment that can take load, plus tested scale-up |
| Active-active | Effectively zero | Effectively zero | ~100%, plus cross-region data transfer | Write conflicts resolved, state localised, and both regions genuinely serving traffic before the event |
The tie to money is two different conversations. RPO is a data decision: asynchronous replication is cheap but has lag, synchronous or consensus replication buys near-zero RPO at the cost of write latency and engineering effort. RTO is mostly operational — detection, the decision to fail over, the traffic shift, and scale-up. A ten-minute technical failover behind a two-hour escalation path is a 2h10 RTO, and the proposal should commit to the number the business will actually experience.
Worked example, hypothetical figures. Primary runs $25k/month, so $300k/year. Pilot light adds roughly 25%, warm standby roughly 75% — a $150k/year gap. Warm standby moves RTO from about 3 hours to about 30 minutes. Assume a region-loss event every 18 months; that assumption is the softest number in the whole model, and I say so. Warm standby then buys back about 1.7 hours of downtime per year, and it pays for itself against pilot light once an hour of downtime costs more than roughly $90k. Above that, warm standby. Below it, pilot light plus a rehearsed runbook is the honest recommendation, and the money saved should fund the drills that make pilot light's RTO real. Where the loss is regulatory or reputational rather than lost revenue, name that separately — expected-value arithmetic understates it and the board knows it.
Matching strategy to funding usually means tiering the estate, showing the breakeven so the customer chooses their own number, and phasing: backup and restore built and tested this quarter beats warm standby that exists only in a diagram.
The failure modes I raise in the same conversation:
- The failover has never been run. The paper RTO and the measured RTO diverge badly, and the usual gap is restore throughput and scale-up, not the switchover itself. Detected only by scheduled drills with a stopwatch and by timing a real restore from real backups.
- Failback is a second migration. After failing over, the old region has stale or diverged data; without reverse replication and a conflict plan, coming home loses writes. Design it before the event, not during it.
- The dependencies nobody modelled: identity provider, DNS registrar, certificate authority, the deployment pipeline, third-party allowlists, and service quotas in the DR region. If the pipeline runs in the region you just lost, you cannot deploy anywhere.
- Active-active specific: write conflicts across regions, cache and session locality, cross-region data transfer becoming a visible line item, and traffic splitting that misbehaves when one region returns bad health signals.
And the region choice itself — second region of the same provider versus a second provider — is a blast-radius question against duplicated operations and expertise. It is a bigger, separate decision than the strategy tier, and I would not let it hide inside the DR discussion.
Curated: · Written: · Reviewed:
QA-77A customer's nightly batch job has grown until it no longer finishes before the business day. What are the options?(show answer)
Establish where the time goes before changing anything, because the remedies are very different and the assumption is usually wrong. Instrument the phases: extraction, transformation, loading, and any waiting on another system.
If a small number of queries dominate, indexing or rewriting them is the cheapest fix by a wide margin and often buys enough headroom for a couple of years.
If the volume itself has outgrown the window, the options in increasing order of change: process incrementally rather than reprocessing everything, which is usually possible once a reliable change marker exists and is frequently the single biggest win; parallelise across independent partitions, which works when the work is separable by key; or move the heavy transformation to a store designed for it rather than the transactional database.
If it waits on another system, the constraint is not theirs to fix alone and that changes the conversation.
There is also a scoping question worth putting to the customer: does all of it need to run nightly? Batch jobs accumulate steps, and some produce outputs nobody uses any more or that could run weekly. Auditing consumption sometimes removes more work than optimisation does.
The structural answer, if the window keeps shrinking, is incremental or streaming processing — but that is a significant change in operational complexity, and I would want the cheaper options exhausted and the growth curve understood before recommending it, because it commits the customer's team to a different class of system to run.
Curated: · Written: · Reviewed:
QA-78How do you explain database indexing to a non-technical stakeholder who is being asked to fund performance work?(show answer)
With the analogy, then the trade, then the number — in that order, and briefly.
The analogy: without an index, answering a question means reading every record, the way you would find a name by reading every page of a book. An index is the book's index — a sorted list that takes you straight to the right page. The book gets slightly thicker and takes a little longer to print, and every lookup becomes almost instant.
The trade, because it is what justifies the work being scoped rather than applied everywhere: each index makes writing slower and takes storage, so they are added for the questions the system actually asks, which is why this is engineering effort and not a switch.
Then the number, which is what they are funding: this specific report currently takes eleven seconds and will take under a second, and the same change removes the load that has been slowing checkout during business hours. Stakeholders fund outcomes, not mechanisms.
If the question behind the question is why this was not done at the start — which it often is — the honest answer is that indexes are chosen for the query patterns and data volumes that exist, and both have changed since. That is normal system evolution rather than an earlier mistake, and saying so protects the team without being defensive.
I would also set the expectation that this is periodic maintenance rather than a one-off, so the next request is not a surprise.
Curated: · Written: · Reviewed:
QA-79A customer wants to keep seven years of transaction history queryable. How do you design for that?(show answer)
Separate the recent data that the operational system needs from the historical data that is queried occasionally, because keeping all of it in the transactional tables degrades everything.
The problem with the naive approach is not storage cost, which is modest. It is that the working set grows, indexes get larger, query plans degrade, backups take longer, and restore times grow — so the recovery objective quietly worsens every year for data nobody touches.
The design: keep a hot window in the transactional store, sized to what the application genuinely needs — often 90 days or a year — and move older data to storage suited to occasional analytical access, typically columnar files in object storage queried directly, which is dramatically cheaper per terabyte and performs well for the scan-heavy queries history attracts.
Partitioning by time makes the movement cheap: dropping or detaching a partition is a metadata operation rather than a large delete, and large deletes on big tables are their own operational problem.
Then establish what the historical access pattern actually is, because it determines how much effort the query path deserves. Regulatory retrieval of a single record occasionally is a very different requirement from monthly analytical reporting across all seven years.
Two things to confirm with the customer: whether the seven years is a legal obligation, which affects whether it can be deleted and whether it must be immutable, and whether the history must remain identical to what was recorded — because if so, the archive needs write-once storage rather than a copy someone could edit.
Curated: · Written: · Reviewed:
QA-80How do you handle a customer whose developers want a technology the operations team cannot support?(show answer)
Surface it as an organisational decision with a cost attached, rather than letting it be settled by whoever is more insistent — because both sides are usually right about their own constraint.
The developers are typically right that the technology suits the problem. The operations team is typically right that they cannot run it at the standard the business expects. Those are not contradictory, and treating it as a technical argument means one side loses on a point that was never technical.
Put the actual options in front of the sponsor with prices. Adopt it and fund the capability: training, hiring, or a managed service that removes the operational burden — often the cheapest of the three and the one that dissolves the disagreement entirely, because a managed offering moves the work to a vendor. Adopt it with the development team owning operations, which is a legitimate model and requires an explicit on-call commitment rather than a vague promise. Or choose the alternative the operations team can run, and accept whatever that costs in development effort or capability.
The decision belongs to whoever owns both budgets, and framing it that way usually improves the conversation, because each team stops arguing for its own constraint and starts comparing costs.
The failure mode to name explicitly: adopting it without funding the capability. That produces a system nobody is confident operating, which degrades quietly and fails at the worst moment — and it is by far the most common outcome when this disagreement is not resolved deliberately.
Curated: · Written: · Reviewed:
QA-81What do you look for when a customer says their system is slow but has no metrics?(show answer)
Get a measurement before forming any theory, because slow is a report about experience and the cause is rarely where people assume.
The fastest useful measurement is from the outside: time the actual user journey from where users are. That distinguishes a slow server from a network or asset-delivery problem, and it is often the whole answer — a page that is slow because of unoptimised images has no server-side cause to find.
Then narrow by dimension, which is possible even without instrumentation: is it slow for everyone or some users, everywhere or one region, all day or at specific times, all operations or particular ones? Each answer eliminates large categories. Slow only at 9am points to a batch job or a capacity limit; slow only for one customer points to data volume for that tenant.
Then look at what exists even in an uninstrumented system: web server access logs usually carry response times, the database can report its slowest statements, and the infrastructure has basic resource metrics. That is usually enough to find the dominant cause.
The recurring causes in this situation, in rough order: a missing index on a table that has grown, an N+1 query pattern generated by an ORM, an unoptimised front end, a synchronous call to a slow third party in the request path, and connection pool exhaustion at peak.
And instrument as you go, because the same question will be asked again in six months and the answer should not require repeating this exercise.
Curated: · Written: · Reviewed:
QA-82A customer asks for a system that never loses data. How do you respond?(show answer)
Take the requirement seriously, then make it precise, because as stated it is unachievable and the useful conversation is about which loss modes matter.
The precise version has three parts. Durability: once we confirm receipt, the record survives hardware failure — that is achievable with synchronous replication and tested backups. Recovery point: after a disaster, how much recently-written data may be lost, which is the RPO and is a cost decision. And integrity: the record cannot be silently altered or lost through a software fault, which is a different problem entirely and is where most real data loss comes from.
That third one is worth dwelling on with the customer, because replication faithfully copies corruption. The losses I have seen were not hardware: a migration that dropped a column, a bug that overwrote records, a deletion that cascaded further than intended, an integration that silently discarded messages. Against those, the protections are point-in-time recovery, immutable audit records, validation at boundaries, and backups retained long enough to predate a slow-moving corruption.
Then the acknowledgement contract, which is the part systems get wrong: never confirm to a user or a partner until the data is durably committed. Most data loss visible to customers is a system that said yes before it had persisted anything.
I would propose specific numbers — an RPO, a retention period, a tested restore — rather than an absolute, and demonstrate them with a drill. A tested restore is more reassuring than any promise, and the customer can see it.
Curated: · Written: · Reviewed:
QA-83How do you decide whether a customer needs an API gateway?(show answer)
By whether there are cross-cutting concerns that would otherwise be implemented in every service, and by whether the consumers are external.
For a single internal service with one known consumer, a gateway is a hop and a component for no benefit, and I would not recommend one.
For external or partner consumers, it earns its place quickly. Authentication, per-consumer rate limiting, request validation, usage metering, and a stable public contract that does not expose internal service boundaries are all things that otherwise get reimplemented per service — inconsistently, and the inconsistencies are where the gaps appear. Centralising them also means a policy change happens once.
The other case is a customer with several services and a public surface, where the gateway provides routing and versioning independent of how the services are arranged internally, so the internal structure can change without breaking consumers.
The cautions I would state. It is on every request path, so its availability becomes theirs, and it needs the same care as any production dependency. And it accumulates logic: teams put transformation and business rules there because it is convenient, and it becomes an untested tier that nobody owns. The rule I would set is cross-cutting policy only — anything domain-specific belongs in a service where it can be tested.
Also worth checking what they already have. A load balancer with authentication and rate limiting, or a CDN with edge rules, may cover the requirement without another component to operate.
Curated: · Written: · Reviewed:
QA-84The customer's team asks why you are proposing message queues when direct calls are simpler. How do you justify it?(show answer)
By naming the specific failure that direct calls produce in their case, and by conceding the simplicity point, because they are right that it is simpler.
The justification has to be concrete. If their order service calls the fulfilment system synchronously and fulfilment is down for an hour, orders fail — the customer cannot buy, and the business loses revenue for a failure in a system that is not on the critical path of taking money. A queue changes that outcome to a backlog that drains when fulfilment returns. That is the whole argument, and it is a business argument rather than an architectural preference.
The second justification, where it applies, is load smoothing: a spike that would overwhelm a downstream system becomes a queue that drains at whatever rate that system can sustain.
Then be honest about what it costs, because a team that only hears the benefits will resent the consequences. Eventual consistency the user may notice, so the interface has to represent in-progress states. A broker to operate. Duplicate delivery, so consumers must be idempotent. Harder debugging, since there is no stack trace across the boundary and correlation IDs become necessary rather than optional. And a dead letter queue that someone has to watch, or failures accumulate invisibly.
Given that, I would use queues only where the decoupling earns those costs — typically at boundaries where a downstream failure should not fail the user's action — and keep direct calls everywhere the caller genuinely needs the result.
Curated: · Written: · Reviewed:
QA-85How do you approach pricing an ongoing managed service for a customer after the build is complete?(show answer)
From the operational load the design actually creates, measured where possible, rather than from a percentage of the build cost.
The inputs: how many alerts the system is expected to generate and at what hours, how much routine work there is — patching, certificate rotation, capacity review, backup verification — how much change is expected, and what response time the customer needs. A system with a 15-minute response commitment overnight requires a staffed rotation and costs an order of magnitude more than business-hours support, and customers frequently ask for the former while budgeting for the latter.
Establish that expectation explicitly, in the same conversation as the price. Response time, coverage hours, what is included, and what is chargeable change rather than support. Ambiguity there is the most common source of dispute in these arrangements, because both sides read a vague scope in their own favour.
Then price the improvement work separately and visibly. A managed service where the provider profits from the system staying noisy has an incentive problem the customer will eventually notice. Structuring it so that reducing alert volume is rewarded — or at least not penalised — is both more honest and produces a better outcome.
Base the initial figure on the first months' measured load if possible, with a review point, rather than a fixed guess for three years. And be explicit about what the customer must retain on their side, because a support arrangement that assumes they can grant access or make decisions promptly will fail if nobody there is accountable for it.
Curated: · Written: · Reviewed:
QA-86A customer asks what happens to their system when a cloud region has an outage. How do you answer honestly?(show answer)
By describing what their current design actually does, without softening it, and then separating the failure classes so the answer is useful rather than alarming.
For a typical single-region deployment across multiple availability zones, the honest answer is that a genuine regional outage means they are down for its duration, and that duration is not under their control. Multi-zone protects against a data centre failure, not a regional one.
Then give the useful context. Full regional outages are rare; partial ones — a single service degraded, the control plane unavailable while running workloads continue — are more common and often survivable, particularly if their system does not need to scale or deploy during the event. That distinction matters, because a design that keeps serving without the control plane behaves very differently from one that cannot.
Then be specific about what they would actually do: is there a documented procedure, are backups in another region, has a restore been tested, and what is the realistic recovery time. If backups are in the same region, that is the finding to raise immediately, because it is cheap to fix and it is the difference between a long outage and a permanent loss.
Then present the options with costs, and let them decide: cross-region backups as a minimum, warm standby, or active-active.
What I would avoid is implying more protection than exists. This question is often asked because someone at board level has asked it, and an answer that turns out to be optimistic will surface at the worst moment.
Curated: · Written: · Reviewed:
QA-87How do you make sure a customer understands the difference between a pilot and production readiness?(show answer)
By writing down what is deliberately absent from the pilot, and repeating it whenever the pilot succeeds — which is when the distinction is most at risk.
The mechanism is a short, explicit list attached to the pilot from the start: this environment has no high availability, no disaster recovery, no security review, no performance testing at production volume, reduced monitoring, and no support commitment. It exists to answer one question. Each item is something a stakeholder might otherwise reasonably assume is present.
The pressure to skip the gap between pilot and production arrives precisely when the pilot goes well: users like it, a sponsor asks why it cannot simply be opened up, and the cost of saying no rises with the pilot's success. Having named the gaps in advance turns that conversation into a plan with a price rather than a judgement about caution.
Quantify the gap when it arrives, because "it needs hardening" is easy to discount. This much work and this long to add resilience, run a security review, test at real volume, and establish support — with the specific risks of not doing it stated in terms of consequence rather than principle.
Where the pressure is genuinely irresistible, a middle path is sometimes right: extend to a limited, informed audience with explicit constraints and a stated end date, while the production work proceeds. That is better than either refusing or pretending the pilot is ready.
And keep the pilot's own scope honest: a pilot that quietly grows features is no longer testing the question it was built for.
Curated: · Written: · Reviewed:
QA-88What is your approach to designing for a customer whose requirements you expect to change substantially?(show answer)
Optimise for the cost of change rather than for the fit to today's specification, and be explicit that this is the trade being made.
Concretely, that means keeping the data model normalised and general rather than shaped tightly to the current workflow, since a schema fitted precisely to a process that changes is the most expensive thing to unpick. It means putting the volatile parts — pricing rules, workflow steps, business policy — behind interfaces or in configuration rather than distributed through the code. And it means keeping decisions reversible: choosing the option that can be replaced over the one that is marginally better but embeds itself.
It also means resisting premature commitment to an integration surface. Where a boundary is uncertain, an internal module boundary is cheaper to move than a network one.
What it does not mean is building configurable generality everywhere. Speculative flexibility is expensive, usually guesses wrong about which axis will vary, and produces a system that is hard to understand in exchange for options nobody exercises. The evidence for where to invest in flexibility comes from asking the customer which parts of their business are actually changing, and building for that axis alone.
The other half is process. Short phases with decision points, so a change in requirements meets a plan rather than a contract, and a change process agreed in advance so it is routine rather than adversarial.
And say plainly that this design will cost slightly more to build and considerably less to change, so the choice is theirs and made knowingly.
Curated: · Written: · Reviewed:
QA-89A customer wants to know why the estimate for the second phase is higher than the first. How do you explain it?(show answer)
By identifying which of the three usual reasons applies, because they need very different responses and conflating them damages trust.
The first is that the second phase is genuinely harder: the first delivered the straightforward, well-understood journeys, and what remains is the complicated integrations, the edge cases, and the data migration. That is normal and should have been signalled at the start, but frequently was not. Explaining which specific work makes it harder is what makes this credible.
The second is that the first phase revealed something — data quality worse than assumed, an undocumented consumer, a legacy interface that cannot sustain the required rate. This is the estimate improving rather than growing, and framing it that way is honest: the first number was made with less information, and here is what was learned. Customers accept this reasonably well when the finding is specific.
The third is that the first phase was underestimated, and the second is being priced realistically. If that is the truth, say it. It is uncomfortable and it is far better than a fabricated technical justification, which an informed customer will see through and which poisons everything else you tell them.
In every case, show the composition rather than the total, so they can see where it goes and challenge specific lines.
And if the increase changes whether the project is worth doing, say that too. The customer's ability to stop is part of what a phased engagement is for.
Curated: · Written: · Reviewed:
QA-90How do you evaluate whether a customer's proposed SaaS product will meet their compliance obligations?(show answer)
By asking for evidence rather than assurances, and by being clear that the customer remains accountable regardless of what the vendor holds.
The evidence: their current audit report or certification, read rather than noted — the scope section is where the useful information is, because a certification covering a different product or a different region is common and easy to miss. A data processing agreement covering the relevant regimes. Their sub-processor list, since data frequently flows to a fourth party nobody considered. And documented locations for storage, processing, backups, and support access.
Then the questions the paperwork does not answer. Where does support access production data from, and is that a transfer under the applicable regime? What is the deletion process and how is it evidenced? What are the breach notification terms and timelines, and do they meet the customer's own obligations to their regulator? What happens to the data on exit, in what format, and over what period?
Then the customer's own responsibilities under the shared model, which is where most compliance failures actually occur: configuration, access management, and what they put into the product. A compliant product configured carelessly is not a compliant system.
I would also involve their compliance or legal function early rather than presenting a conclusion. This is their decision to make, my role is to gather the evidence and identify the gaps, and a technical recommendation that pre-empts a legal judgement is one they cannot rely on.
Curated: · Written: · Reviewed:
QA-91The customer wants a single system of record but has three departments with conflicting data definitions. What do you do?(show answer)
Treat it as a governance problem with a technical component, and resolve the definitions before building anything, because the conflict does not disappear when it is encoded.
The pattern is familiar: sales counts a customer at the point of signature, finance at first invoice, and support per contract, so three departments report three different customer counts and each is correct within its own definition. A single system with one definition makes two departments wrong.
The first step is to make the conflict explicit and specific: write the three definitions down side by side with worked examples. That alone frequently resolves it, because people have been using one word for three things without noticing.
Then the options. Agree a single canonical definition, which requires a decision-maker senior to all three and is the cleanest outcome when achievable. Or model the underlying reality more precisely — a party with several relationships — and derive each department's view from it, which is usually the better technical answer because each department keeps its correct number and they reconcile to the same base. Or accept separate systems with a documented mapping, which is the honest choice when the definitions are genuinely different business concepts rather than a disagreement.
What I would avoid is choosing a definition by default because a schema had to be written. That is how a system becomes untrusted: two departments see numbers that do not match their own, stop believing it, and keep their spreadsheets — which is the outcome the project was meant to eliminate.
Curated: · Written: · Reviewed:
QA-92How do you decide what belongs in the first release versus later, when the customer wants everything?(show answer)
By separating what is needed for the system to be usable at all from what is needed for it to be complete, and by making the trade visible rather than deciding it privately.
The test I apply to each item: can a real user complete a real task end to end without it? If yes, it is a candidate for later. That is stricter than it sounds and it eliminates a great deal, because much of what feels essential is refinement of a path that already works.
The exceptions that must be in the first release regardless of that test: anything that is expensive to retrofit, and anything a user would be harmed by its absence. Authentication and the permissions model, the audit trail if one is required, and the data model itself belong in the first category — retrofitting a tenancy boundary or an audit requirement is far more expensive than building it in. Data protection obligations belong in the second.
Then put the ranked list in front of the customer rather than presenting a scope decision as made. Ranking their own priorities is a conversation they can have; being told what was cut is one they resist.
Two techniques that help. Ask what they would ship if the date were two months earlier, which produces a sharper ranking than asking what is important. And identify what can be handled by a manual process initially — a report someone runs, an approval by email — because that often defers real engineering effort at very low cost for a low-volume workflow.
Curated: · Written: · Reviewed:
QA-93A customer's cloud bill has doubled in twelve months and they want to cut it. How do you find where the money is going, and what do you change first?(show answer)
Start with the cost-and-usage data and propose nothing until you can name where the growth came from. A bill that doubles in twelve months almost always has two or three dominant causes, and a generic optimisation checklist is worthless until you have found them.
I read the data in three cuts. By service — what grew. By account or project — who to talk to. By tag — which workload, and whether it is production or a dev environment nobody has stopped since March. AWS gives me the Cost and Usage Report into Athena, or Cost Explorer grouped by linked account and cost-allocation tag; GCP is a billing export to BigQuery; Azure is a Cost Management export. The first thing I check is whether the tags exist at all. If a third of the spend is unallocated, that is finding number one, because every conversation after it is guesswork. Then I ask what changed: a doubling is usually a new environment, an unbounded log stream, a storage class nobody revisited, or a data-transfer path that arrived with an architecture change — not prices drifting.
Then the fixes, in rough order of yield per hour of effort, which falls off sharply. Idle and oversized compute first: dev/test running 168 hours a week and used for fifty, production instances at 8% p95 CPU. Storage lifecycle second — everything in hot tiers regardless of age, orphaned snapshots, incomplete multipart uploads. Data transfer third: cross-AZ chatter between services calling each other thousands of times a second, logs shipped uncompressed across boundaries, bytes pushed through a NAT gateway that a gateway endpoint would move at no charge. Autoscaling and scheduling fourth. Commitments last.
Worked example, a B2B SaaS platform of roughly 200 workloads, bill $85k/month growing to $171k (figures rounded and illustrative):
| Line | Month 12 | Cause | Change | After |
|---|---|---|---|---|
| Non-prod compute | $46k | ran 24/7, 30% oversized | stop/start 12h weekdays, rightsize | $19k |
| Object storage | $38k | all hot tier, 40% untouched >180 days | lifecycle to infrequent/archive, purge snapshots | $24k |
| Transfer + NAT | $27k | cross-AZ chatter, uncompressed logs | batch and filter logs, gateway endpoints | $12k |
| Databases | $31k | two idle read replicas, oversized | retire one, rightsize | $26k |
| Logs and monitoring | $18k | debug verbosity in prod | 30 days hot, archive after | $11k |
| Everything else | $11k | — | — | $11k |
$171k down to $103k, about $816k a year. Only then commitments: with the right-sized steady-state compute floor near $55k/month, a one-year no-upfront savings plan is roughly 25% on the committed portion — another $14k/month, and the equivalent of a GCP committed use discount or Azure reservation. The order matters. Buying that commitment against the original $95k of compute would have locked in $40k/month of waste for a year. I cover only 60-70% of the observed floor first and re-measure after a month.
Making it stick is governance, not engineering. Baseline one full month before changing anything, or you cannot prove anything afterwards. Enforce the tag policy through policy-as-code in the pipeline, not a wiki page. Budgets and anomaly alerts per team — a 30% week-on-week anomaly alert finds the rogue training job in days instead of at month end. Publish monthly showback by team; I start with showback rather than chargeback because chargeback needs finance to agree a cost-allocation method, and that stalls the work for a quarter.
On the people: finance gets the number, the date and the named owners in their own worksheet. Engineering gets the change list as work items with an owner on each, and I say explicitly that missing tags are a platform problem, not a team failing. Savings nobody owns come back within two quarters.
Where I weigh cost against reliability and speed, I will not take a saving that quietly moves risk onto the customer. The largest line I refused in that engagement was retiring the second read replica of the orders database: $3k/month, and half their recovery capability for a tier-1 service. I wrote the alternative down instead — keeping that availability target costs about $36k a year, and losing it in an AZ event costs several times that in an afternoon. I have also refused to shrink test environments until engineers stop testing against realistic data, and refused to cut log retention below the compliance window. The rule I give the customer: if the saving is real but it buys back an outage or a slowdown in delivery, name it as a risk with an owner and let them decide with the number visible, rather than deciding quietly on their behalf.
Curated: · Written: · Reviewed:
QA-94How do you handle discovering, mid-project, that the architecture you proposed will not meet a key requirement?(show answer)
Raise it immediately, with the evidence, an assessment of the options, and a clear statement of what it costs — in that order, and to the sponsor rather than only within the team.
Immediately is the whole thing. The instinct is to work the problem quietly and hope it resolves, and every week of delay narrows the options and increases the cost of the eventual conversation. A problem raised in month three with three viable responses is a manageable event; the same problem raised in month six with one is a crisis.
Bring evidence rather than a concern, so the discussion is about what to do rather than whether it is real. A load test result, a measured latency, a limit in a vendor's documentation.
Bring options with costs and consequences: change the design and what that adds; change the requirement, which is sometimes possible once the price of meeting it is visible; accept a limitation, if it is tolerable and documented; or stop, if the project no longer makes sense.
Be straightforward about responsibility without either deflecting or over-apologising. If the assumption was mine and it was wrong, say so once, then spend the time on the remedy — the customer's interest is in what happens next.
Then look at why it was found now rather than earlier. Usually the answer is that the riskiest assumption was not tested first, and the correction for the remainder of the project is to sequence the untested assumptions ahead of the comfortable work.
Curated: · Written: · Reviewed:
QA-95What makes a good architecture decision record for a customer engagement specifically?(show answer)
The same structure as any decision record, plus two things that matter more when the reader will be someone else's team in two years.
The standard content: the context and the forces in play, the options considered including the rejected ones, the decision, and its consequences stated honestly including what became harder. The rejected options are the part that is impossible to reconstruct later and the part most often omitted.
The first addition for a customer engagement is the constraint's source. "We chose this because the customer's security policy prohibits X" or "because their contract with vendor Y runs until 2028" is essential, because those constraints expire. A future engineer finding a decision with no stated source will either preserve a constraint that no longer exists or overturn one that still does.
The second is the reversal condition, written specifically: what would have to change for this to be worth revisiting. That converts the record from a historical note into something actionable, and it is what makes a decision defensible when a new team questions it.
Practically: keep them in the customer's repository rather than in your own systems, because they must outlive the engagement. Write them at the time rather than reconstructing them at handover, when the reasoning has already faded. Keep them to a page. And make them immutable — superseded by a new record rather than edited — so the history of the reasoning survives.
The test is whether a new engineer's "why is it like this?" has an answer that is not a person who has left.
Curated: · Written: · Reviewed:
QA-96How do you decide whether to recommend a customer build on a platform their team already knows but that is a worse technical fit?(show answer)
By pricing the mismatch and the learning curve honestly against each other, over the life of the system rather than the length of the project.
Familiarity is a genuine engineering advantage and it is routinely dismissed as a soft factor. A team operating something they understand recovers faster from incidents, makes better trade-offs under pressure, and does not need a supplier on retainer. That is worth a real amount, and it compounds.
The question is how bad the mismatch is, and specifically whether it is a matter of elegance or of capability. A platform that is merely less suited but can do the job is usually the right choice for a team that knows it. A platform that cannot meet a hard requirement — a throughput ceiling, a compliance boundary, a data model it fundamentally does not support — is not a trade-off, it is a failure, and familiarity does not rescue it.
So the test I apply is whether the mismatch shows up as ongoing cost or as a wall. Ongoing cost can be weighed against the learning curve and the operational risk of the unfamiliar option. A wall cannot.
Where the better-fitting technology is genuinely necessary, the recommendation has to include funding the capability — training, hiring, a managed offering that removes the operational burden, or a support arrangement with an end date — because recommending it without that is recommending a system the customer cannot run.
And I would check how much of the fit argument is mine rather than theirs. It is easy to over-value technical elegance when someone else operates the result.
Curated: · Written: · Reviewed:
QA-97A customer wants to move from quarterly releases to weekly. What has to change first?(show answer)
The things that make a release risky, in the order that each one unblocks the next. Increasing frequency without changing them just delivers the same risk more often.
Automated tests come first, because the quarterly cadence is almost always sustained by a manual regression pass that cannot run weekly. Without a suite that gives real confidence, everything downstream is unsafe.
Then a repeatable, automated deployment, including the database migration path — manual steps are where releases go wrong and they do not scale with frequency.
Then rollback, tested. This is the control that makes the rest safe, and it is the one most often left manual, which means it is slowest exactly when it is needed. A team that can reverse a release in minutes tolerates far more change than one that cannot.
Then enough observability to know within minutes whether a release worked: error rate, latency, and the key business metric, visible per deployment.
Then the release must be decoupled from the deploy, using feature flags, so that shipping code and exposing behavior are separate decisions. This is what allows small, frequent, low-risk deploys of work that is not yet finished.
Only then does frequency increase safely, and it should be gradual — monthly, then fortnightly — with the failure rate watched at each step.
The organisational change matters as much: a change advisory board meeting monthly cannot approve weekly releases, so the approval model has to move from per-release scrutiny to confidence in the pipeline.
Curated: · Written: · Reviewed:
QA-98How do you assess whether a customer's proposed integration will scale to their projected volume?(show answer)
By finding the constraint at the far end, because the bottleneck in an integration is almost always the other system rather than yours.
The measurements to get: what rate does their system actually sustain, what are its documented and undocumented limits, and how does its latency behave as concurrency rises. Call it and find out rather than reading the documentation, because published limits and observed behavior diverge often, particularly with older enterprise systems.
Then do the arithmetic against the projection, and specifically against the peak rather than the average. A daily volume divided by 86,400 seconds is a meaningless number if the traffic arrives in a two-hour window; the peak rate is what has to fit.
Then check for the shapes that break at volume even when the rate fits: an integration that makes one call per record where a batch endpoint exists, a synchronous call in a loop, or pagination that gets slower with depth. Each of these works fine at pilot volume and fails at production volume, which is why pilots are poor evidence for this question.
If the constraint binds, the options are to negotiate a higher limit, which is often possible and worth asking; to reshape the integration — batch instead of per-record, or a bulk file for the initial load; or to decouple with a queue so the peak is smoothed to a rate the far end can absorb.
And state the ceiling explicitly in the design, so the customer knows at what volume this needs revisiting rather than discovering it during a busy period.
Curated: · Written: · Reviewed:
QA-99A customer's finance lead says your proposed architecture is too expensive and asks you to cut the bill by 40% without accepting downtime. How do you respond?(show answer)
Start with the invoice, not the architecture. A 40% target is arithmetic against a breakdown, and in most estates the top three line items carry 70% or more of the spend — so the cut lands in one or two places, not spread evenly across everything. Two questions come first: is the 40% on the total run-rate or on the infrastructure line only, and is the baseline a full business cycle or one peak month? A bill that includes 40 TB of one-off migration egress is not the bill we are optimising.
Then decompose it. Hypothetical but typical, a $96,000/month estate looks like this:
Line | Monthly | What usually sits in it Compute, all on-demand | $38,000 | Oversized instances, non-prod running 24/7, zero commitment coverage Managed database, Multi-AZ plus replica | $21,000 | Provisioned for peak, 20% average utilisation Object and block storage | $9,000 | Everything in hot tier, old snapshots, unattached volumes Internet egress | $12,000 | BI extracts and API responses leaving over the public path Logging and observability | $8,000 | Debug-level logs retained 30 days hot Load balancing, NAT, licences, misc | $8,000 | NAT data processing is the classic hidden item
Now the levers, in order of what they cost the customer in behaviour. Waste first: unattached volumes, orphaned snapshots, idle NAT gateways, load balancers with no targets — normally 5–8%, so about $7,000 here, and it costs nothing but a day of someone's time. Second, rightsizing against measured p95 utilisation over a full business cycle: $8,000, but do not rightsize against a quiet fortnight or the peak day fails. Third, commitment on the steady floor — say $24,000 of the remaining compute is genuinely steady, and a one-year no-upfront Savings Plan takes 15–20% off it, so roughly $4,300. Note that the marketing figure of up to 72% is three-year all-upfront; I quote the number out of the pricing calculator at their actual rates. Fourth, scheduling non-prod: if it is a third of compute and you run it 12 hours on weekdays, that is 60 of 168 hours, so you keep 36% and save $6,400. Fifth, egress shape — 80 of 120 TB behind a CDN at a blended $0.03–0.05/GB instead of $0.09, and the BI tool moved in-region so extracts stop crossing the internet: about $5,000. Sixth, lifecycle policies and gp2 to gp3: 60 TB to Glacier Instant Retrieval at $0.004/GB-month instead of $0.023, plus block storage at $0.08 instead of $0.10: $1,600.
That is $32,300 of resilience-neutral savings, roughly 34%. The remaining $6,000 is where the trade-offs actually live, and this is the part you present rather than perform: drop the read replica and reporting p95 rises from 40 ms to perhaps 300 ms under load; move the reporting store to single-AZ and its RPO becomes the last backup — acceptable, because it is derived data you can rebuild; run the overnight batch on spot capacity and save 60% on it, but only if it is checkpointable and start time is allowed to drift. Deliberately untouched: backup retention, security controls, and the observability floor required to meet the incident response commitments in the contract.
For the two audiences: the engineer gets the mechanism — instance families, commitment coverage percentage, lifecycle rules, the peak-day headroom we are protecting and the number we are protecting it to. The finance lead gets a one-page ladder with three columns: what we cut, what it saves per year, and what it costs in RTO, RPO, latency or staff time. Put the money in the risk column: Multi-AZ at $7,000/month buys failover in around a minute against a snapshot restore measured in hours, and at their transaction rate that outage is worth more than the saving. And be explicit that a Savings Plan is a commitment they owe even if the workload is retired in month seven.
Two failure modes to name out loud. Savings erode without an owner — in my experience 5–10% drifts back per quarter unless someone owns a monthly review — so the deliverable includes an owner and a tag policy, not just a slide. And the over-correction: rightsizing to the point where autoscale cannot react in time, or cutting logs below what the support team needs, turns a cost problem into an availability problem. If the honest ladder tops out at 35%, say so, show where it stops, and let the finance lead choose the last $5,000 with the downtime cost printed next to each option.
Curated: · Written: · Reviewed:
QA-100How do you communicate an incident during a customer engagement when your own work caused it?(show answer)
Early, factually, and with the remedy — and take responsibility without spending the customer's time on contrition.
The sequence in the moment: tell them it is happening and what the impact is before they discover it, even when you do not yet know the cause. A customer who learns about an outage from their own users while their supplier is silently investigating loses confidence in a way that is hard to recover.
Then give them what they need to run their own business: what is affected, what is not, what the workaround is, and when the next update will come — and then actually send the update at that time, even if it only says the investigation continues. Predictable communication is what reduces the anxiety that produces escalation.
When the cause is your work, say so plainly and once. "Our change caused this" is worth more than any explanation, and it should be followed immediately by what is being done, not by a defence. Detailed causal explanation belongs in the postmortem, not in the incident channel.
Afterward, a written postmortem with the mechanism, why it was not caught, and specific actions with owners and dates — and then evidence that the actions happened, because the follow-through is what actually restores trust rather than the document.
The thing to avoid is minimising. Customers usually know roughly how bad it was, and a description that understates it converts a technical failure into a credibility problem that outlasts the incident by a long way.
Curated: · Written: · Reviewed:
