Top 100 Forward Deployed Engineer (FDE) Interview Questions and Answers
The questions most likely to actually come up in your Forward Deployed Engineer (FDE) interview, ranked by likelihood — with detailed, senior-level answers covering what an interviewer is really listening for.
Curated: · Written: · Reviewed:
QA-1What does a forward deployed engineer do that a software engineer or solutions engineer does not?(show answer)
The dividing line is not customer contact — all three roles have that. It is where the code runs and who is on the hook when the customer's workflow does not change.
An FDE writes production code that runs inside one customer's environment and owns that customer's outcome through adoption. A solutions engineer is done when the contract is signed; a product engineer is done when the feature ships to everyone. The FDE is done when the named users at one account actually changed how they work, and is expected to carry whatever repeated across accounts back into the product.
| Solutions engineer | Product / software engineer | Forward deployed engineer | |
|---|---|---|---|
| Where the code runs | demo or POC, usually discarded | your SaaS, all customers | customer's VPC or on-prem, on the hook for it |
| Definition of done | signed deal | shipped feature, adoption on your dashboard | the customer's workflow metric moved |
| Scope of ownership | one technical evaluation | one roadmap area | one customer's whole workflow, end to end |
| Feedback loop | win/loss notes | product telemetry | named users, weekly, plus evidence for the roadmap |
The work itself runs as three loops. Find the real problem by sitting with the people who do the job — not the sponsor who bought. Build and deploy the smallest system that moves their metric, yourself, including the boring parts: auth into their IdP, the nightly batch against their warehouse, the on-call rotation. Then feed the repeats upstream: "three of my last four accounts hand-rolled the same SFTP ingestion, here are the tickets" is a product request with evidence, not a complaint.
Take a worked example, figures illustrative. A claims insurer wants triage automated. Week 1 with three adjusters shows the actual bottleneck is not classification — it is 40 minutes per claim pulling documents from two systems. I ship a read-only connector and a merged view in week 3; handle time drops from 52 to 31 minutes for those three users. Week 5 I widen to the 25-person team, and p95 page load goes from 1.2 s to 9 s because the connector fans out per document. That is my bug to fix, not a ticket I file and wait on. By week 8, 22 of 25 adjusters use it daily and the second account I touch needs the same connector — that is the moment it becomes a product request.
Failure modes, with how you catch them:
- FDE as shadow SE. Lots of decks, no merged code. Detectable in week two: zero deployments, zero commits in the customer's repo.
- FDE as bespoke consultancy. Every engagement forks the product and nothing goes upstream. Detectable: after four accounts, no shared component and no roadmap items traced to field work.
- Blocked on the platform team. The pilot ends with 0 production users because the FDE waited for features instead of shipping around them. Detectable: build queue age on FDE-owned work exceeds a week.
- Unbounded scope. The FDE becomes the customer's de facto on-call for systems they did not build. Detectable: more than roughly a day a week of support that no deployment of ours sits inside.
Do not hire an FDE when the gap is a missing feature — that is a product engineer with a roadmap. Do not hire one when the gap is pre-sale confidence — that is an SE. FDEs pay off when the product needs real integration work at a design partner or lighthouse account and the learning has to come back.
In an interview, ask for one deployment where the candidate wrote the code that went live and the workflow metric it moved. That single example separates all three roles.
Curated: · Written: · Reviewed:
QA-2A customer asks you to build a dashboard. How do you find out what they actually need?(show answer)
Assume I'm deployed into the account: shared channel with their ops team, read access to some of their data, and an engagement measured in weeks rather than a product roadmap I control. Under those conditions a dashboard request is a hypothesis about a decision, and my job in the first call is to name three things — the decision, the person who makes it, and its cadence. If I can't get all three, I don't have a requirement yet.
I ask in this order:
- Who opens it, and what do they do in the ten minutes after? "Watch the numbers" is not an action. "Call the 3PL and re-prioritize the pick wave" is. The verb is the requirement.
- When did this hurt last — walk me through the last three times. Timestamp, what they looked at, what they decided, what it cost. Recounted incidents beat opinions about frequency, and three is enough to see whether it's one decision or three.
- Show me what you use today. Every ops team has a spreadsheet of record — a nightly export with hand-added columns. The columns someone typed in by hand are the requirements. The columns nobody fills in anymore are the ones I can drop.
Then I go read evidence they didn't think to offer: the Teams channel where the question gets asked every morning, the report attachment that gets forwarded with "any update?", the ticket queue. If someone posts "where are the late shipments?" at 08:15 daily, that is a requirement with a timestamp on it.
| Asked for | Decision it serves | Smallest build that changes it |
|---|---|---|
| dashboard of shipments | which 40 late orders to chase at 08:00 | ranked list, refreshed 07:30, with owner and last-touch |
| monthly stock report | whether to reorder this SKU | alert when days-of-cover drops below 5 |
| "real-time" exec dashboard | which of three regions to visit this quarter | weekly four-column CSV |
The gate before I write code: one sentence naming the decision, confirmed in writing by the person who makes it — not the sponsor who requested the build. That distinction is where most of the waste lives.
Failure modes I watch for. Sponsor is not the user: detected when the SOW signer can't say what happens after the number gets read. The 22-chart ops dashboard I shipped early on died this way — usage settled at 2 logins a week because the supervisors needed a list of 40 late shipments each morning, not charts. Illustrative arithmetic: 22 charts × 2 users × 2 visits is 88 renders a week against maybe 40 decisions; the ranked list gets used 40 times before 08:30 every weekday. Second, the dashboard is really a data-integration project: detected when "where does that number come from?" answers "a CSV a 3PL emails us on Tuesdays." Then the pipeline is the deliverable and the dashboard is a debug view. Third, a metric their data can't support: detected by asking for the schema before promising the chart — the field is missing, or it carries PII the customer's security review won't clear for our tenant. Fourth, "everyone needs it": past roughly four named users with genuinely different decisions, it's two products wearing one name.
Sometimes I build the dashboard anyway — a renewal conversation needs an exec-visible artifact, or procurement scoped it that way. Fine. I still ship the ranked list in week one and hang the charts off it in week three, so the thing that changes a decision is in use while the picture is still being styled.
Curated: · Written: · Reviewed:
QA-3Why would you sit with users for a day before designing anything, and what do you look for?(show answer)
Because the gap between the documented process and the one people actually run is where deployments die, and interviews cannot see it. Asked to describe their job, users give the official version from memory and quietly normalize the rest. The time and the errors live in the workarounds they no longer notice: the spreadsheet that reconciles two systems, the phone call that unblocks a queue, the copy-paste bridge into the tool nobody wants to touch. A day of watching is the cheapest way to see the work rather than the story about the work — and for an FDE it doubles as reconnaissance, because every workaround is either an integration requirement or a data-model constraint I have to build against.
How I spend it: follow 10 to 20 real cases end to end, from arrival to resolution, and time each step. Log every system switch and every handoff. Crucially, split working time from waiting time, because the two look identical on a wall clock and mean completely different designs. One case from an intake engagement, medians across the referrals we watched:
| Step | Working | Waiting |
|---|---|---|
| receive referral | 4 min | 0 |
| find patient record | 6 min | 0 |
| wait for clinician review | 3 min | 26 h |
| schedule appointment | 5 min | 4 h |
That is 18 minutes of work in a 30 h 18 min cycle — 99% of the elapsed time is queueing. Halving the typing saves 9 minutes; shaving 4 hours off the clinician review queue saves 26 times that. So the design question is not "how do I make the form faster", it is "who is holding the case, and why". A faster form on that workflow would have been a rounding error.
What I look for, concretely:
- The second screen. The shadow spreadsheet or shared inbox is usually the real system of record for something the primary system can't express. Whatever it holds is a schema requirement.
- System switches and re-keying. Each one is a future integration point and an error source. A single case often traces like:
fax → shared inbox → re-key into EHR → manual duplicate check in spreadsheet → page clinician queue. - The exception path. Ask for the last three cases that went wrong. Happy paths are what people demo; the exceptions are what breaks the tool at week two.
- Vocabulary. Users say "the blue ones" where the spec says "urgent referral type B". Naming the product in their words is most of adoption.
- Who owns the queue. The person who can change the process is the deployment's real owner, and they are rarely the person who booked my visit.
Failure I've seen justify the day: an intake tool designed from interviews alone, which on launch could not accept 35% of referrals — they arrived by fax, a channel nobody mentioned because nobody thought of it as a channel.
The method has its own failure modes. People perform the documented process for an observer, so I ask to watch the last case rather than a demonstration. I shadow whoever is busiest, not whoever is most articulate. A single day misses seasonality — month-end, night shift, the Monday surge — so I pull ticket or audit logs for the surrounding week and check the day was representative. Where I can't be in the room, screen recordings plus ticket and audit logs are a weaker substitute and I label the evidence as such.
What I bring back is a timed map of the observed cases showing where waiting accumulates, the systems touched, and the exceptions seen — reviewed with the people I watched before a single design decision is made.
Curated: · Written: · Reviewed:
QA-4How do you agree with a customer on what success means for a deployment?(show answer)
Assume the customer says "we want to speed up claims triage." My job in the kickoff is to turn that into one number somebody inside their org will be judged on, plus the rules for computing it. Success is a change in a workflow outcome they already track — not a feature shipped, not usage, and not model accuracy. Usage and accuracy are inputs; if they move and the outcome doesn't, the deployment isn't working.
Assumptions I state up front: there is a workflow owner with authority to commit to a target, and the outcome is measurable from data I can reach. If either is missing, the first deliverable is instrumentation, not the model.
The definition I ask the sponsor to sign fits on one page:
| Field | Example |
|---|---|
| Primary metric | median claim triage time, submission to correct queue |
| Baseline | 9 days (measured over the 30 days to 1 Mar) |
| Target | 1 day by 30 Jun |
| Measurement | SQL below, run weekly on the claims warehouse |
| Guardrails | wrong-queue rate < 3%, reopen rate < 5%, p90 triage < 3 days |
| Review date | 15 Jul, decision: expand / fix / stop |
The measurement clause is the part people skip and the part that prevents arguments later. I put the actual query in the doc:
-- PostgreSQL dialect, checked against PG 15
SELECT percentile_cont(0.5) WITHIN GROUP (ORDER BY queued_at - submitted_at) AS median_triage
FROM claims
WHERE submitted_at >= now() - interval '30 days'
AND status <> 'reopened';
Agreeing the exclusions (reopened) and the window (30 days) at the start matters as much as the target. Otherwise at the review someone discovers the baseline included cases the new system rejects, and the comparison is dead.
Arithmetic on the example: 9 days to 1 day is an 89% reduction, at roughly 4,000 claims a month, so a 30-day window gives about 4,000 samples — plenty to distinguish 1 day from 3. If the customer's volume were 40 a month, I would widen the window to a quarter or switch to a per-case trace rather than pretend a median means anything.
Failure modes I watch for:
- Proxy declared the win. I lost a renewal this way: a deployment was called successful on 91% classification accuracy, the claims backlog was unchanged after three months, and the customer reasonably felt nothing had improved. Detection is mechanical — at every review, plot the primary metric next to the proxy. Proxy up, outcome flat, means the model isn't on the critical path of the workflow.
- Baseline nobody can reproduce. The sponsor says "we're at 9 days." I ask them to show me the query before signing. If they can't, we instrument and measure two to four weeks pre-launch and the target date moves. A target against a guessed baseline is a commitment to a coin flip.
- Gaming. Median triage time falls because easy claims auto-triage and hard ones sit untouched. That is why the guardrails carry p90 and wrong-queue rate, not just a second average.
- Signed by someone who can't commit. If the sponsor won't put a date and a number on paper, that is information: the project has no internal owner, and I escalate before writing code.
I don't always demand a signature. For a two-week pilot with one team, the same six fields in a shared doc and a named reviewer is enough — the ceremony scales with the contract value and the number of teams touched. What I won't do is start building against "we'll know it when we see it."
The one thing I want the sponsor to say out loud at the review is what happens if the number misses: extend, change approach, or stop. Agreeing that in advance is what makes the metric binding rather than decorative.
Curated: · Written: · Reviewed:
QA-5Why measure the current process before you ship anything?(show answer)
The baseline is part of the deliverable, not prep work. In an FDE engagement I am dropping software into a customer's live process — their ops team, their data, their SLAs — and the number I will be defending at the pilot review, at the renewal, and in every later stakeholder meeting is a delta. Without a measured "before", that delta is someone's memory of how things used to go, and the ops lead's memory will beat mine.
The assumption worth stating up front: the process produces an outcome I can timestamp before my system touches it — case handling time, throughput, rework rate, escalation volume. If it does, I measure it. If nothing is instrumented, I time live cases: shadow an operator through 20-30 real cases, record start and end, and keep the spread, not just the average.
It shapes the build, not just the verdict. Measuring first tells you where the time actually goes. A client once assumed slow handling was the problem; the timestamps said median handling was 48 minutes but p95 was just over 6 hours, almost all of it queue wait before anyone touched the case. That is a queueing and routing build, not an efficiency build. Shipping the efficiency tool first would have produced a clean-looking 20% win on a number nobody cared about.
What gets recorded before launch is small and specific:
metric: median handling time, ticket opened -> status=resolved
source: reporting.tickets, query baseline-v1 (frozen, version-controlled)
window: 2025-01-06 .. 2025-02-16 (6 weeks, includes month-end close)
n: 412 cases; 11 merged duplicates excluded
signed: customer ops lead, 2025-02-18
The same query, same exclusions and same window length run after launch. Illustrative result:
| Measure | Before (6 wk) | After (6 wk) | Method |
|---|---|---|---|
| median handling time | 48 min | 29 min | ticket timestamps |
| rework rate | 12% | 7% | reopened within 7 days |
Failure modes I watch for, with how I catch them:
- Instrument drift. Before was ticket timestamps, after is app logs, and they disagree by 8 minutes on the same week. Detection: run both instruments over one shared week before launch and reconcile them.
- Seasonality and regression to the mean. Baseline taken during a quiet stretch, or the week after an incident spike. Detection: two comparable cycles, or a matched process in the same org that doesn't get the tool as a control.
- Concurrent changes. The customer re-staffed the queue mid-pilot and the win was partly theirs. Detection: confounders go in the same log as the metrics, dated.
- Selection. Only cases that flowed through the new tool were timed, so abandoned and escalated ones vanished. Detection: measure the whole population, including the ones that never reach it.
The trade-off is time. For a two-day prototype I do not run a six-week study; I take a coarse estimate, say it is coarse, and don't let baseline work gate a ship. Baseline effort scales with the size of the claim I intend to make — a 40% efficiency claim to an executive sponsor earns the full measurement contract; a demo does not.
One failure I've seen: a team presented a 2x speedup, the customer's ops lead produced older figures showing the process was already faster than the team had assumed, and every subsequent result from that engagement was met with a spreadsheet. The claim wasn't wrong so much as unanchored. Record the baseline, its window and its query in the pilot success document, get the ops lead to sign it, and the argument afterwards is about the work instead of about the measurement.
Curated: · Written: · Reviewed:
QA-6The customer wants a platform. How do you decide what to build first?(show answer)
I would anchor scoping the smallest valuable slice in evidence the customer already trusts.
The first release should be one end-to-end path for one user group that changes one real outcome, even if it is narrow. A thin working slice produces evidence and trust, while a broad half-built platform produces neither.
Concretely, list the candidate workflows, score each on value, feasibility with available data, and time to first user, pick the one that can reach real users in weeks, and write explicit non-goals for everything else.
The reason for that specificity is a failure I have seen: A team spent 5 months building ingestion for 14 source systems before any user saw a screen, and the sponsor who funded the project left before the first release.
Choosing the first slice.
| Candidate workflow | Value | Data ready | Weeks to first user |
|---|---|---|---|
| late-order triage | high | yes | 3 |
| demand forecasting | high | partial | 12 |
| supplier scorecards | medium | no | 20 |
I would not consider it settled without evidence: Show a written scope with one workflow, one user group, a date for first real use, and a non-goals list signed by the sponsor.
Narrow and live beats broad and pending.
Curated: · Written: · Reviewed:
QA-7A hospital asks you to reduce emergency department wait times. How do you break the problem down?(show answer)
On decomposing an ambiguous problem, I would rather ship a narrow slice that users rely on than a broad one they ignore.
Decompose the outcome into the stages a patient passes through and measure time in each, because the fix depends on which stage dominates. Arrival, triage, waiting for a bed, treatment, and discharge each have different owners and levers.
Concretely, draw the patient flow, measure median time per stage from existing timestamps, find the stage with the most delay, and propose a targeted intervention for that stage with a way to test it.
The reason for that specificity is a failure I have seen: A team built a triage prediction model for an emergency department where triage took 8 minutes, while patients waited a median of 4 hours for inpatient beds, so the model changed nothing measurable.
Median time per stage.
| Stage | Median time | Owner |
|---|---|---|
| arrival to triage | 8 min | nursing |
| triage to doctor | 55 min | ED staffing |
| decision to admit to bed | 4 h 10 min | bed management |
| discharge | 35 min | ward |
I would not consider it settled without evidence: Quantify time per stage from real timestamps and show that the proposed intervention targets the largest one.
Find the biggest delay before choosing the tool.
Curated: · Written: · Reviewed:
QA-8Operations wants speed, compliance wants more checks, and IT wants fewer integrations. How do you decide?(show answer)
What an interviewer listens for on conflicting stakeholder requests is evidence from real users and real data.
Conflicting requests usually share one underlying goal, and making the trade-off explicit with numbers lets the decision owner choose. The FDE's job is to frame the options, not to pick a winner quietly.
Concretely, write each stakeholder's requirement as a measurable constraint, show two or three designs with their effect on each constraint, and ask the executive sponsor to choose in a single meeting with the trade-offs on one page.
The reason for that specificity is a failure I have seen: An FDE tried to satisfy every group with configurable options, delivered a system with 60 settings that nobody could explain, and the compliance team blocked launch for 7 weeks.
Three designs against three constraints.
| Option | Handling time | Checks performed | New integrations |
|---|---|---|---|
| A, fast path | 12 min | 3 | 1 |
| B, full review | 31 min | 9 | 1 |
| C, risk-based | 15 min | 3 to 9 by risk | 2 |
I would not consider it settled without evidence: Produce a one-page option comparison with each constraint quantified and record which option the sponsor chose.
Make the trade-off visible and let the owner decide.
Curated: · Written: · Reviewed:
QA-9Who are the different people you need on your side at a customer, and why does it matter?(show answer)
I would handle users, buyers, and decision owners so that it still works after I have left the customer.
The buyer funds the work, the decision owner defines success and approves changes, the users decide whether it is adopted, and IT and security decide whether it can run. A deployment fails if any one of them is missing from the plan.
Concretely, map each group by name early, meet the users as often as the sponsor, bring security in during week one, and keep a short weekly update that each group can read in two minutes.
The reason for that specificity is a failure I have seen: A project had strong executive support but never met the night-shift users, who kept using a paper log, and after launch 70 percent of cases still bypassed the new system.
A stakeholder map.
| Group | Cares about | Contact frequency |
|---|---|---|
| executive sponsor | business outcome | every 2 weeks |
| users | their daily effort | weekly, on site |
| IT and security | risk and support load | weekly in the first month |
I would not consider it settled without evidence: Keep a named stakeholder map with each group's success measure and the date of the last conversation with them.
Adoption is decided by the people who do the work.
Curated: · Written: · Reviewed:
QA-10How would you estimate whether a deployment is worth doing before you start?(show answer)
The honest answer on estimating the value of a deployment includes what could go wrong at the customer and how I would know.
Value is the change in an outcome multiplied by its unit value and volume, minus the cost of building and running the system. A rough estimate with stated assumptions is enough to decide which workflow to start with.
Concretely, multiply volume by the expected improvement and the value per unit, compare several candidate workflows the same way, and state the assumption that the answer is most sensitive to.
The reason for that specificity is a failure I have seen: A team chose a technically interesting forecasting project worth about 200,000 dollars a year over a claims triage project worth an estimated 3.3 million, because nobody wrote either estimate down.
Claims triage value estimate.
| Input | Value |
|---|---|
| claims per year | 120,000 |
| handling hours saved per claim | 0.5 |
| cost per hour | $55 |
| annual value | $3.3m |
I would not consider it settled without evidence: Write a one-page estimate for each candidate with volume, improvement, unit value, and cost, and share the assumptions with the sponsor.
Estimate value before choosing where to start.
Curated: · Written: · Reviewed:
QA-11How do you structure a four-week proof of value so it ends with a clear decision?(show answer)
With a time-boxed proof of value, the customer's constraints come first and the architecture follows them.
A proof of value tests one risky assumption on real data with real users and ends on a go or stop decision agreed in advance. Without a fixed end date and criteria it drifts into an unpaid pilot that never ends.
Concretely, name the assumption to test, agree the success threshold and the data before starting, run on a real slice of work, and hold a decision meeting on a fixed date with the sponsor.
The reason for that specificity is a failure I have seen: A proof of concept ran for 7 months with no agreed criteria, and when budget season arrived the customer could not say whether it had worked.
A four-week plan.
| Week | Activity | Output |
|---|---|---|
| 1 | data access and baseline | baseline on 500 cases |
| 2 to 3 | build and run on a live slice | results on 1,200 cases |
| 4 | decision meeting | go, change, or stop |
I would not consider it settled without evidence: Put the assumption, threshold, data slice, and decision date in writing at the start and present results against them at the end.
A proof of value needs an end date and a finish line.
Curated: · Written: · Reviewed:
QA-12When should an FDE refuse or push back on a customer request?(show answer)
I would make saying no to a customer request measurable against the baseline the customer already trusts.
Refuse when the request would break a safety, security, or legal boundary, when it serves one person at the expense of the agreed outcome, or when it would create a one-off system nobody can support. Saying no works better when it comes with an alternative.
Concretely, explain the risk in the customer's terms, offer an option that meets the underlying need, and escalate to the sponsor when the requester insists, with the trade-off written down.
The reason for that specificity is a failure I have seen: An FDE agreed to let a regional manager bypass an approval step for speed, and within 6 weeks 140 payments had been released without the second check the auditors required.
Refusals with alternatives.
| Request | Risk | Alternative |
|---|---|---|
| skip approval for speed | audit failure | pre-approve low-risk payments under $500 |
| export all patient data to Excel | privacy breach | aggregated view with access controls |
I would not consider it settled without evidence: Record the refused request, the stated risk, and the alternative offered, and confirm the sponsor agrees.
A no that comes with an alternative protects both the customer and the deployment.
Curated: · Written: · Reviewed:
QA-13A customer asks for a feature that the platform already supports in another way. What do you do?(show answer)
I would start a request the product already covers differently in the customer's workflow rather than in our product catalogue.
Building a duplicate creates a fork that must be maintained forever, so first check whether the existing capability meets the underlying need with configuration or training. If it does not, the gap is product feedback rather than a custom build.
Concretely, show the user the existing path on their real case, measure whether it meets the need, and if not, file the gap with the product team with examples and frequency instead of coding a private version.
The reason for that specificity is a failure I have seen: Three FDE teams each built their own approval workflow for different customers because nobody checked the platform's workflow engine, and the company later spent 9 months reconciling the three versions.
Three ways to meet the request.
| Option | Build effort | Long-term cost |
|---|---|---|
| configure existing workflow | 2 days | none |
| custom build | 3 weeks | maintained per customer |
| product feature request | 0 for the FDE | shared by all customers |
I would not consider it settled without evidence: Demonstrate the existing capability on the customer's real case and record whether it met the need.
Check the product before forking it.
Curated: · Written: · Reviewed:
QA-14The customer promises their data is clean. What do you check before committing to an outcome?(show answer)
The first thing I want to know about checking the customer's data before promising results is which decision it changes for which user.
Customer data is rarely as complete or consistent as described, and the outcome depends on it. A short profile of the real data before signing a target avoids promising results the data cannot support.
Concretely, get a sample extract early, measure completeness of key fields, duplicates, date coverage, and label quality, and adjust the target or the scope based on what the data shows.
The reason for that specificity is a failure I have seen: A team committed to predicting equipment failures, then found that only 18 percent of past failures had a recorded cause, which made the promised model impossible to train.
Expected against actual completeness.
| Field | Expected complete | Actually complete |
|---|---|---|
| failure date | 100% | 97% |
| failure cause | 90% | 18% |
| machine id | 100% | 88% |
I would not consider it settled without evidence: Profile a real extract and share completeness and quality figures with the sponsor before the target is signed.
Promise what the data can deliver.
Curated: · Written: · Reviewed:
QA-15Why start the customer's security review in the first week?(show answer)
For security review and data access up front, I would name the outcome, the owner, and the measurement before any code.
Security approval and data access are usually the longest items on the critical path, often taking weeks, and they cannot be rushed at the end. Starting them first lets the build run in parallel.
Concretely, send the architecture, data flows, and access requirements in week one, ask for the security questionnaire early, and build with synthetic or sample data while approvals are pending.
The reason for that specificity is a failure I have seen: A deployment was ready to launch in week 6 but waited 11 more weeks for a security review that had not been requested until the code was finished.
Approvals in parallel with the build.
| Item | Typical duration | Start |
|---|---|---|
| security questionnaire | 3 to 6 weeks | week 1 |
| data access approval | 2 to 8 weeks | week 1 |
| build on sample data | 4 weeks | week 1, in parallel |
I would not consider it settled without evidence: Track security and data-access approvals on the project plan from week one with named owners and dates.
Start the slowest approval first.
Curated: · Written: · Reviewed:
QA-16The customer requires everything to run inside their private network with no internet access. How does that change your plan?(show answer)
My approach to deploying inside a customer's network constraints separates what the customer asked for from what they need.
A private or air-gapped deployment removes managed services, public package registries, and remote debugging, so installation, updates, and monitoring must all be planned for the customer's environment. Assumptions from a cloud demo often do not hold.
Concretely, list every external dependency, mirror packages and images into the customer's registry, plan how logs and metrics leave the network if at all, and rehearse the full install on a clean environment that matches theirs.
The reason for that specificity is a failure I have seen: A deployment failed on day one because the application downloaded a model file from the internet at start-up, and the customer's firewall blocked it for 3 days while approvals were obtained.
Dependencies moved inside the network.
| Dependency | Cloud demo | Private network plan |
|---|---|---|
| container images | public registry | customer registry mirror |
| model weights | downloaded at start | shipped in the image |
| monitoring | SaaS tool | customer's log platform |
| outbound calls in isolated test | 7 | 0 |
I would not consider it settled without evidence: Run a full install on a network-isolated test environment and list every outbound call it attempts.
Design for the network you will actually run in.
Curated: · Written: · Reviewed:
QA-17How do you turn a customer's requirements into a spec your team can build from?(show answer)
I treat turning requirements into a technical spec as something I own end to end, from discovery to adoption.
A good spec states the user workflow, inputs and outputs, data sources, rules that must hold, failure behaviour, and how success will be measured. It is short enough that the customer can confirm it is what they meant.
Concretely, write a two-page spec with one worked example that traces a real case through the system, list open questions with owners, and have the customer's workflow owner review it before building.
The reason for that specificity is a failure I have seen: A spec described the approval rules in prose, and 3 weeks into the build the team learned that two regions applied the rules in different orders, which required rewriting the decision logic.
The core of a two-page spec.
| Section | Example content |
|---|---|
| workflow | intake to approval for invoices over $10k |
| rules | two approvers above $50k |
| failure behaviour | queue for manual review, never auto-approve |
| success | median approval time from 2 days to 4 hours |
I would not consider it settled without evidence: Trace at least one real case from each region through the spec and have the workflow owner confirm each step.
A spec is confirmed when the customer recognizes their own case in it.
Curated: · Written: · Reviewed:
QA-18Halfway through a deployment the customer keeps adding requests. How do you handle it?(show answer)
I would anchor managing scope creep in evidence the customer already trusts.
Each added request delays the outcome that was agreed, so it should be traded against something rather than absorbed. A visible backlog with estimates lets the customer choose what matters most.
Concretely, log every new request with an estimate, show its effect on the launch date, and ask the sponsor to either swap it for something in scope or schedule it after launch.
The reason for that specificity is a failure I have seen: A team accepted 23 small requests during a 10-week deployment, and the launch slipped by 9 weeks while the original outcome was still not delivered.
A change log with decisions.
| Request | Estimate | Decision |
|---|---|---|
| extra export format | 2 days | deferred to phase 2 |
| new region | 2 weeks | swapped for reporting module |
| colour changes | 1 day | rejected |
I would not consider it settled without evidence: Keep a change log with each request, its estimate, and the sponsor's decision to swap, defer, or reject it.
Every yes to scope is a no to the date unless something else moves.
Curated: · Written: · Reviewed:
QA-19How do you explain a technical trade-off to a customer executive?(show answer)
On explaining trade-offs to executives, I would rather ship a narrow slice that users rely on than a broad one they ignore.
Executives decide on outcomes, cost, risk, and time, so a trade-off should be stated in those terms with a recommendation. Technical detail belongs in a backup slide.
Concretely, present two or three options in one table with time, cost, risk, and outcome, give a recommendation with the reason, and state what you need from them and by when.
The reason for that specificity is a failure I have seen: An FDE explained a database migration choice in terms of replication modes, the executive postponed the decision to a later meeting, and the project lost 4 weeks.
A trade-off in business terms.
| Option | Time to launch | Extra cost | Risk |
|---|---|---|---|
| launch on current system | 3 weeks | none | slower reports |
| migrate first | 9 weeks | $120k | delay |
| recommended, launch now and migrate in Q3 | 3 weeks | $120k later | low |
I would not consider it settled without evidence: Check that the one-page summary states the options in business terms and ends with a specific decision request.
Ask for a decision in the executive's own terms.
Curated: · Written: · Reviewed:
QA-20You arrive at a skeptical customer. What do you do in the first two weeks?(show answer)
What an interviewer listens for on building trust in the first two weeks is evidence from real users and real data.
Trust comes from understanding the users' problem and delivering something small that works, not from presentations. An early visible win makes users willing to spend time on the larger build.
Concretely, spend the first days with users, pick one annoying problem you can fix within two weeks, ship it, and use the credibility to get access and time for the main work.
The reason for that specificity is a failure I have seen: An FDE spent the first 3 weeks on architecture documents, the operations team saw nothing useful, and they stopped attending the weekly sessions.
The first three weeks.
| Week | Action | Result |
|---|---|---|
| 1 | shadow 6 users | list of 9 pain points |
| 2 | fix duplicate alert emails | 300 fewer emails a day |
| 3 | start the main build | users attend reviews |
I would not consider it settled without evidence: Ship one small fix used by real users within the first two weeks and record who uses it.
Earn the right to build big by fixing something small.
Curated: · Written: · Reviewed:
QA-21When should an FDE escalate an issue to the core product team rather than fix it in the field?(show answer)
I would handle escalating to the core product team so that it still works after I have left the customer.
Escalate when the fix belongs in the shared product, when the issue affects other customers, or when a workaround would create a risk or a permanent fork. Fix locally when it is customer-specific configuration or integration.
Concretely, escalate with a clear write-up of the impact, frequency, affected customers, a reproducible example, and a proposed fix, and agree a temporary mitigation with the customer meanwhile.
The reason for that specificity is a failure I have seen: An FDE patched a data-loss bug locally for one customer without telling the product team, and the same bug corrupted records for 4 other customers over the next month.
Local fix or escalation.
| Issue | Fix locally | Escalate |
|---|---|---|
| customer-specific field mapping | yes | no |
| bug in shared platform code | temporary only | yes |
| feature requested by 3 customers | no | yes, with evidence |
I would not consider it settled without evidence: File each escalation with impact, frequency, reproduction steps, and affected customers, and track it to resolution.
Fix it where it belongs, and tell the people who own it.
Curated: · Written: · Reviewed:
QA-22You take over a deployment that is behind schedule and the customer is unhappy. What do you do?(show answer)
The honest answer on rescuing a failing deployment includes what could go wrong at the customer and how I would know.
First find out what the customer now believes success is and what has failed, then reset to a smaller commitment that can be delivered quickly. Credibility returns through a kept promise, not through a new plan.
Concretely, interview the sponsor and users, list the broken commitments honestly, agree one near-term deliverable with a date, deliver it, and then rebuild the larger plan with them.
The reason for that specificity is a failure I have seen: A new lead presented a revised 6-month roadmap to an unhappy customer without delivering anything first, and the customer cancelled the contract 3 weeks later.
A two-week reset.
| Step | Timeframe |
|---|---|
| listen to sponsor and users | days 1 to 3 |
| agree one deliverable | day 5 |
| deliver it | day 14 |
| reset the roadmap together | day 15 |
I would not consider it settled without evidence: Agree one concrete deliverable with a date within the first two weeks and report whether it was met.
Recover trust with a kept promise, then plan.
Curated: · Written: · Reviewed:
QA-23You need to integrate your product with a customer's legacy internal system with poor or missing API documentation. How do you approach it?(show answer)
With integrating with an undocumented legacy system, the customer's constraints come first and the architecture follows them.
The integration surface is what the running system actually does, not what its documentation claims. Establish it from evidence — code, schemas, live traces, and whoever keeps it alive — and prove it with one thin end-to-end slice before anyone estimates the build.
Concretely, trace one real transaction end to end: capture the HTTP traffic or database writes, read the producing code, and have the maintainer walk the unhappy paths. Write the observed fields, formats, and error behaviour into a contract both teams sign, push one record through, and get the customer to confirm it before deciding where the translation lives.
The reason for that specificity is a failure I have seen: A team built to a documented SOAP contract for six weeks; the WSDL behind the load balancer was three years stale, the live service returned HTTP 200 with an empty result set when an account did not exist, and 8% of orders silently disappeared during the first production week before anyone noticed the totals drifting.
Where to put the translation, with illustrative figures.
| Option | Effort (hypothetical) | Who ships changes | Choose it when |
|---|---|---|---|
| adapter service you own | 2 weeks to build, deployable same day | you | their team runs a 6-week change board |
| change the legacy system | 5 lines of code, 6 weeks to release | their team, needs UAT sign-off | the fix belongs in their domain and they have a fast path |
| extend your product | 2 days plus core-team review, ships in 3 weeks | your product team | 4+ customers need the same behaviour |
I would not consider it settled without evidence: Show one record traced end to end — one log line or packet capture per hop with the same order ID and amount — plus the customer's written sign-off that the final values are correct.
Prove one record all the way through, then estimate the rest.
Curated: · Written: · Reviewed:
QA-24Mid-project the customer asks for a feature that is outside the statement of work, and it is the most interesting engineering problem you have heard all month. How do you respond?(show answer)
I would make an out-of-scope request you want to build measurable against the baseline the customer already trusts.
Your own interest in the work is not a scope decision; the boundary of a deployment is commercial and belongs to the account team and the customer sponsor. Surface the need behind the ask, size it in days, and present a signed trade instead of absorbing it quietly.
Concretely, start by asking what decision or outcome the feature would serve and whether something already in the product reaches it, then estimate it in engineering days, name the milestone or buffer it consumes, and take two or three options to the account team and the sponsor to sign.
The reason for that specificity is a failure I have seen: A request estimated at 3 days took an FDE 9 days once permissions and backfill were included; the migration-testing buffer vanished, the milestone of 1,000 cases processed slipped from week 6 to week 8, and no sponsor had ever agreed to trade it away.
Sizing the ask against the 12 engineering days left before week 6.
| Option | Build days | What it displaces before week 6 |
|---|---|---|
| build as asked | 6 | 4 days of migration testing and the 2-day buffer |
| existing filters plus a saved view | 0.5 | nothing |
| defer to phase two | 0 | nothing |
I would not consider it settled without evidence: Show the one-page change request with the ask, the underlying need, the day estimate, the milestone it displaces, and the signature that accepted the trade.
Curiosity is not scope; a signed trade is.
Curated: · Written: · Reviewed:
QA-25After go-live you keep ownership of the customer's deployment. What does your on-call model look like?(show answer)
I would start on-call model for a customer deployment in the customer's workflow rather than in our product catalogue.
An FDE on call is the incident channel for one customer and the product's early warning system at the same time, and that only works if severity levels, coverage hours and response targets were agreed with the customer before launch rather than improvised in an outage. Every incident should end in two artefacts: a fix the customer can verify and a product ticket with reproduction detail.
Concretely, agree severity levels, coverage hours and response targets with the customer before launch, and route platform alerts and customer reports into one rotation you page on. Post status on a fixed cadence until resolution, then file a product ticket for anything that traces back to the platform.
The reason for that specificity is a failure I have seen: Alerts went to an internal Slack channel and the customer reported a stale dispatch board by email on a Saturday; nobody looked until Monday, by which time 38 hours of orders were missing from the board and the operations manager had gone back to phoning drivers by hand.
Severity and response targets agreed with the customer before launch.
| Severity | Definition on the customer's side | First response | Updates | Escalation |
|---|---|---|---|---|
| Sev1 | dispatch or billing stopped for 50+ users | 30 min, 24/7 | every 30 min | product on-call at 2 h |
| Sev2 | wrong data in a live report, workaround exists | 2 business hours | every 4 h | product ticket same day |
| Sev3 | cosmetic or single-user issue | 1 business day | at resolution | weekly backlog review |
I would not consider it settled without evidence: Show the severity and response table the customer signed before launch, then one drill or real incident where the first status update met its target.
On call for one customer means you own the first hour, not just the fix.
Curated: · Written: · Reviewed:
QA-26The customer's core system has no API, only nightly CSV exports. How do you integrate with it?(show answer)
The first thing I want to know about integrating with a system that has no API is which decision it changes for which user.
File-based integration is common and workable if it is treated as a real interface, with a schema contract, validation, and monitoring. The risks are silent format changes, partial files, and late deliveries.
Concretely, agree the file schema and delivery time, validate each file for row count, columns, and types before loading, load into a staging table, and alert when a file is missing, late, or malformed.
The reason for that specificity is a failure I have seen: A nightly export silently switched its date format from day-first to month-first, and 11 days of records were loaded with the wrong dates before a user noticed.
Checks on each nightly file.
| Check | Rule | Action on failure |
|---|---|---|
| arrival | by 02:00 | alert at 03:00 |
| columns | exactly 18 named | reject file |
| row count | within 20% of 7-day average | hold and alert |
| dates | ISO format only | reject file |
I would not consider it settled without evidence: Reject any file that fails schema, type, or row-count checks and alert the owner the same night.
A file drop is an API with no error messages, so add your own.
Curated: · Written: · Reviewed:
QA-27How do you make a data load safe to rerun?(show answer)
For idempotent ingestion, I would name the outcome, the owner, and the measurement before any code.
An idempotent load produces the same result whether it runs once or five times, which makes retries and backfills safe. The usual method is to upsert on a natural key or replace a whole partition instead of appending.
Concretely, identify a stable key for each record, write loads as upserts or partition overwrites, record which source file or batch each row came from, and test by running the same load twice.
The reason for that specificity is a failure I have seen: A retried pipeline appended the same 60,000 invoices a second time, and the customer's finance team saw revenue doubled for a day before the duplicates were removed by hand.
Three load styles run twice.
| Load style | After first run | After second run |
|---|---|---|
| append | 60,000 rows | 120,000 rows |
| upsert on invoice_id | 60,000 rows | 60,000 rows |
| partition overwrite by date | 60,000 rows | 60,000 rows |
I would not consider it settled without evidence: Run the same load twice on a test dataset and confirm row counts and totals are unchanged.
Every load should be safe to run again.
Curated: · Written: · Reviewed:
QA-28The same customer appears differently in the CRM, billing, and support systems. How do you link them?(show answer)
My approach to entity resolution across customer systems separates what the customer asked for from what they need.
Linking records needs an explicit matching rule, starting with shared identifiers and falling back to normalized fields such as email, phone, and name plus address. Every rule produces some false matches and some missed matches, and both need to be measured.
Concretely, normalize fields, match on exact identifiers first, then on combinations of normalized fields, review a sample of matches by hand, and keep a mapping table recording the rule that produced each link.
The reason for that specificity is a failure I have seen: A matching rule on company name alone merged 1,300 unrelated branches of franchise businesses into single accounts, and support agents saw the wrong customer's history.
Matching rules and their precision.
| Rule | Links made | Precision in review |
|---|---|---|
| tax id exact | 41,000 | 99.8% |
| email domain plus postcode | 12,500 | 96% |
| company name only | 9,800 | 71% |
I would not consider it settled without evidence: Review 200 sampled matches and 200 sampled non-matches by hand and report precision and recall for each rule.
Measure the matching rule before trusting the links.
Curated: · Written: · Reviewed:
QA-29A customer's source system adds and renames fields without telling you. How do you protect the pipeline?(show answer)
I treat schema drift from customer sources as something I own end to end, from discovery to adoption.
Upstream schema changes are inevitable, so the pipeline should detect them at the boundary and fail or quarantine rather than load misaligned data. Silent acceptance is the dangerous case.
Concretely, compare each incoming schema with the expected contract, fail on removed or retyped fields, quarantine records with unexpected fields, and notify the source owner with the difference.
The reason for that specificity is a failure I have seen: A renamed status column arrived empty under its old name for 3 weeks, and every case in the pipeline defaulted to open, inflating the backlog report by 40 percent.
Responses to schema changes.
| Change | Action | Seen last quarter |
|---|---|---|
| new optional column | load, log, notify | 6 |
| column renamed | fail load, notify owner | 2 |
| type changed | fail load, notify owner | 1 |
| column removed | fail load, notify owner | 1 |
I would not consider it settled without evidence: Compare the incoming schema with the contract on every load and alert on any added, removed, or retyped column.
Catch schema changes at the door.
Curated: · Written: · Reviewed:
QA-30Which data quality checks would you put on a customer's data feed?(show answer)
I would anchor data quality checks at ingestion in evidence the customer already trusts.
Checks should cover volume, completeness, uniqueness, validity, and freshness, each with a threshold based on normal behaviour. The goal is to stop bad data before users make decisions on it.
Concretely, set expected ranges from recent history, run checks on every load, block loads that break critical rules, warn on softer ones, and show data freshness to users next to the numbers.
The reason for that specificity is a failure I have seen: A feed delivered half its usual volume for a week because one region's export failed, and a regional manager cut staffing based on the apparent drop in demand.
Feed checks with thresholds.
| Check | Threshold | Severity |
|---|---|---|
| row count vs 7-day median | within 25% | block |
| null customer id | under 1% | block |
| duplicate order id | 0 | block |
| data age | under 6 hours | warn |
I would not consider it settled without evidence: Run volume, null-rate, uniqueness, and freshness checks on every load and record their results.
Check the data before the users see it.
Curated: · Written: · Reviewed:
QA-31The customer's table has 400 million rows. How do you keep your copy up to date without reloading it daily?(show answer)
On incremental loads and change data capture, I would rather ship a narrow slice that users rely on than a broad one they ignore.
Incremental loading moves only new or changed rows, using a reliable updated timestamp or a change-data-capture log from the database. Timestamps miss deletes and back-dated edits, so the method must match how the source changes.
Concretely, use change data capture if the source supports it, otherwise load rows with an updated time after the last watermark minus a safety overlap, and run periodic full reconciliations of counts and totals.
The reason for that specificity is a failure I have seen: An incremental load used created time instead of updated time, so 2.3 million edited records were never refreshed, and reports showed cancelled orders as active.
What each method catches.
| Method | Catches inserts | Catches updates | Catches deletes |
|---|---|---|---|
| created timestamp | yes | no | no |
| updated timestamp | yes | yes | no |
| change data capture | yes | yes | yes |
| rows moved per day, CDC | 1.2m of 400m | none | none |
I would not consider it settled without evidence: Reconcile row counts and key totals against the source weekly and investigate any gap.
Load increments, reconcile totals.
Curated: · Written: · Reviewed:
QA-32You need to reprocess two years of history after a logic fix. How do you do it safely?(show answer)
What an interviewer listens for on running a backfill safely is evidence from real users and real data.
A backfill touches data users rely on, so it should run in bounded batches, be idempotent, avoid overloading the source, and be validated before it replaces what users see.
Concretely, write the backfill to a separate table or partition, run it in date batches with rate limits, compare totals with the old results and explain the differences, then swap it in with a way to revert.
The reason for that specificity is a failure I have seen: A backfill ran against the customer's production database during business hours, slowed their order system for 2 hours, and the customer suspended the team's access.
Backfill totals by month.
| Month | Old total | Backfilled total | Difference |
|---|---|---|---|
| 2025-01 | 1,204,000 | 1,201,900 | -0.2% |
| 2025-02 | 1,188,000 | 1,187,400 | -0.1% |
| 2025-03 | 1,250,000 | 1,061,000 | -15%, investigate |
I would not consider it settled without evidence: Compare backfilled totals with the previous results by month and explain every difference above a set tolerance before switching.
Backfill beside the live data, then swap.
Curated: · Written: · Reviewed:
QA-33You combine event data from three countries' systems and the daily totals look wrong. What would you check?(show answer)
I would handle joining data across time zones and keys so that it still works after I have left the customer.
Systems often store local time without a zone, use different key formats, and close their business day at different hours. Joining without normalizing these produces shifted and duplicated records.
Concretely, convert all timestamps to UTC with the correct source zone, normalize key formats, define the business day explicitly per country, and test the join with known cases.
The reason for that specificity is a failure I have seen: Records from one country stored local time without a zone, were treated as UTC, and appeared 9 hours late, which placed 30 percent of their evening orders on the wrong day.
Normalizing three sources.
| Country | Stored as | Normalized to | Offset in September |
|---|---|---|---|
| UK | local time, no zone | UTC via Europe/London | 1 h |
| Japan | local time, no zone | UTC via Asia/Tokyo | 9 h |
| US | UTC | UTC | 0 h |
I would not consider it settled without evidence: Trace a set of known orders from each country through the join and confirm their dates and keys match the source.
Normalize time and keys before joining.
Curated: · Written: · Reviewed:
QA-34You wrote a Python script to load the customer's data in a day. What must change before it runs in production?(show answer)
The honest answer on hardening a one-day prototype script includes what could go wrong at the customer and how I would know.
A prototype script usually lacks retries, logging, configuration, tests, and a clear failure mode. Production code must be rerunnable, observable, and safe when inputs are bad.
Concretely, move secrets and paths into configuration, add structured logging and exit codes, make it idempotent, add tests on a small real fixture, and schedule it with alerting on failure.
The reason for that specificity is a failure I have seen: A one-day prototype script became the production loader, failed silently on a malformed file, and the customer's reports were 5 days stale before anyone noticed.
Prototype versus production.
| Concern | Prototype | Production |
|---|---|---|
| secrets | in the script | secret store |
| failure | silent exit | alert and non-zero exit |
| rerun | duplicates | idempotent |
| tests | none | fixture of 200 real rows |
I would not consider it settled without evidence: Run the script against a malformed file and confirm it fails loudly, writes nothing partial, and alerts.
A prototype becomes production when it fails loudly.
Curated: · Written: · Reviewed:
QA-35The customer's data includes health records. How does that change how you build?(show answer)
With handling sensitive data in pipelines, the customer's constraints come first and the architecture follows them.
Sensitive data must be minimized, protected in transit and at rest, access-controlled, and logged, and it often must stay inside the customer's environment. Copies made for convenience, such as in logs or on laptops, are the usual leak.
Concretely, collect only needed fields, pseudonymize identifiers where possible, keep processing inside the approved environment, exclude sensitive values from logs, and review access regularly.
The reason for that specificity is a failure I have seen: Debug logging printed full patient records into an application log that was shipped to a third-party monitoring service, exposing 12,000 records.
Minimizing health data.
| Data | Needed | Handling |
|---|---|---|
| patient name | no | not collected |
| date of birth | age only | converted to 10-year age band |
| diagnosis code | yes | encrypted, access logged |
I would not consider it settled without evidence: Scan logs and outputs for sensitive fields before release and after each change.
Keep sensitive data where it is allowed and out of your logs.
Curated: · Written: · Reviewed:
QA-36Several business units share one deployment but must not see each other's data. How do you enforce that?(show answer)
I would make tenant isolation and row-level access measurable against the baseline the customer already trusts.
Access rules belong in the data layer, where every query is filtered by the user's permissions, rather than in the user interface, where a new screen or export can bypass them.
Concretely, attach tenant and role to every request, enforce row-level security in the database or a single query layer, test that each role sees only its data, and cover exports and APIs with the same rules.
The reason for that specificity is a failure I have seen: Access checks were implemented only in the web screens, and a CSV export endpoint returned all 4 business units' records to anyone who could reach it.
Isolation tests by account.
| Test account | Expected rows | Rows returned |
|---|---|---|
| unit A analyst, screen | 12,000 | 12,000 |
| unit B analyst, screen | 8,500 | 8,500 |
| unit A analyst, export | 12,000 | 12,000 |
I would not consider it settled without evidence: Test every endpoint and export with accounts from each tenant and confirm none returns another tenant's rows.
Enforce access where the data is read.
Curated: · Written: · Reviewed:
QA-37The customer's API allows 10 requests a second and fails intermittently. How do you call it reliably?(show answer)
I would start rate limits and retries against customer APIs in the customer's workflow rather than in our product catalogue.
Retries with exponential backoff and jitter recover from transient failures without overwhelming the service, and a client-side rate limiter keeps you under the quota. Retrying calls that change data requires an idempotency key.
Concretely, limit outgoing requests to the quota, retry only on retryable errors with backoff and jitter up to a cap, send an idempotency key on writes, and move repeatedly failing items to a dead-letter queue.
The reason for that specificity is a failure I have seen: A client retried failed calls immediately in a tight loop, hit the customer's API with 400 requests a second during an outage, and was blocked for a day.
Backoff schedule.
| Attempt | Delay before retry |
|---|---|
| 1 | 0.5 s plus jitter |
| 2 | 1 s plus jitter |
| 3 | 2 s plus jitter |
| 4 | dead-letter queue |
I would not consider it settled without evidence: Simulate API errors and throttling in a test and confirm the client stays within quota and does not duplicate writes.
Retry politely and never write twice.
Curated: · Written: · Reviewed:
QA-38The customer's system sends webhooks to your service. What can go wrong and how do you handle it?(show answer)
The first thing I want to know about receiving webhooks reliably is which decision it changes for which user.
Webhooks can arrive late, out of order, more than once, or not at all, so the receiver must verify the signature, acknowledge quickly, process asynchronously, and deduplicate on event id.
Concretely, verify the signature, store the raw event and return success quickly, process from a queue, deduplicate on event id, and run a periodic reconciliation against the source to catch missed events.
The reason for that specificity is a failure I have seen: A webhook handler did slow processing before responding, the sender timed out and resent events, and 2,100 orders were created twice.
Webhook problems and defenses.
| Problem | Defense | Seen per 10,000 events |
|---|---|---|
| duplicate delivery | deduplicate on event id | 42 |
| out-of-order delivery | compare event version | 17 |
| missed delivery | nightly reconciliation | 3 |
I would not consider it settled without evidence: Replay the same event several times and out of order in a test and confirm one correct final state.
Acknowledge fast, process once, reconcile often.
Curated: · Written: · Reviewed:
QA-39In a meeting, the customer asks how many orders were delayed last month. How do you answer quickly and correctly?(show answer)
For answering a customer question with SQL live, I would name the outcome, the owner, and the measurement before any code.
A quick query is only useful if its definition matches the customer's, so confirm what delayed means and which date defines last month before running it. Stating the definition with the number prevents a later dispute.
Concretely, confirm the definition out loud, write the query with the filters visible, sanity check the result against a known total, and send the query and number afterwards so they can check it.
The reason for that specificity is a failure I have seen: An FDE answered 4,200 delayed orders in a meeting using ship date, while the customer's operations team measured by promised delivery date and counted 6,900, and the sponsor questioned every later figure.
Two definitions, two answers.
| Definition of delayed | Delayed orders |
|---|---|
| shipped after promised ship date | 4,200 |
| delivered after promised delivery date | 6,900 |
I would not consider it settled without evidence: State the definition with the number and share the query after the meeting.
Say what you counted when you say the count.
Curated: · Written: · Reviewed:
QA-40Your numbers differ from the customer's ERP. How do you resolve it?(show answer)
My approach to reconciling with the customer's system of record separates what the customer asked for from what they need.
The customer's system of record is the reference by default, so the task is to explain every difference between it and your figures through timing, filters, definitions, or data errors until the gap is zero or understood.
Concretely, pick one period, compare totals, then drill down by category and record until each difference has a named cause, and fix the cause in the pipeline rather than adjusting the number.
The reason for that specificity is a failure I have seen: A deployment reported inventory 3 percent higher than the ERP for months, and the customer stopped trusting any figure until the gap was traced to units counted in cases in one system and in single items in the other.
Bridge to the ERP.
| Step | Units |
|---|---|
| deployment total | 1,030,000 |
| minus in-transit stock counted twice | -18,000 |
| minus case versus single-item mismatch | -12,000 |
| ERP total | 1,000,000 |
I would not consider it settled without evidence: Produce a written bridge from your total to the system of record with each difference quantified.
Explain the gap to zero.
Curated: · Written: · Reviewed:
QA-41How would you model the data for an operations team that manages orders, shipments, and exceptions?(show answer)
I treat designing an operational data model as something I own end to end, from discovery to adoption.
An operational model represents the real things the business acts on, such as orders, shipments, and exceptions, with their relationships and states, so users and applications can work with them directly instead of joining raw tables.
Concretely, identify the core objects with users, define their keys, states, and relationships, map each source table into them, and keep the mapping versioned so source changes do not break applications.
The reason for that specificity is a failure I have seen: Applications were built directly on 40 raw source tables, and when the customer upgraded its ERP, 17 applications broke at once because every one embedded the old table structure.
Core objects for operations.
| Object | Key | States |
|---|---|---|
| order | order_id | open, allocated, shipped, closed |
| shipment | shipment_id | planned, in transit, delivered |
| exception | exception_id | new, assigned, resolved |
| source tables mapped | 40 | into 3 objects |
I would not consider it settled without evidence: Walk a real exception through the model with users and confirm every step can be answered from the objects.
Model what the business acts on, not what the source stores.
Curated: · Written: · Reviewed:
QA-42The customer asks for real-time data. How do you decide whether they need streaming?(show answer)
I would anchor choosing batch or streaming for a customer in evidence the customer already trusts.
Streaming is worth its cost only when a decision must be made within minutes of an event. Most reporting and planning decisions are well served by hourly or daily batches, which are simpler to build and operate.
Concretely, ask how quickly someone acts on the data and what happens if it is an hour old, then pick the slowest refresh that still supports the decision, and use streaming only where it is required.
The reason for that specificity is a failure I have seen: A team built a streaming pipeline for a daily planning meeting, spent 3 months operating it, and the planners still only looked at the data once each morning.
Freshness by decision.
| Use case | Decision latency | Refresh |
|---|---|---|
| fraud block | seconds | streaming |
| dispatch reassignment | 5 minutes | micro-batch |
| daily planning | 1 day | nightly batch |
I would not consider it settled without evidence: Record the decision latency each use case needs and confirm the chosen refresh meets it.
Match freshness to the decision.
Curated: · Written: · Reviewed:
QA-43Field staff work where connectivity drops for hours. How do you design the system?(show answer)
On offline and edge deployments, I would rather ship a narrow slice that users rely on than a broad one they ignore.
Offline work requires a local copy of the needed data, local capture of changes, and a sync process that resolves conflicts when the device reconnects. The conflict rules are business rules and must be agreed with users.
Concretely, decide what data each device needs, store changes locally with timestamps and ids, sync when connected, detect conflicting edits, and resolve them with agreed rules or a manual review queue.
The reason for that specificity is a failure I have seen: A field app saved edits only when online, and crews in rural areas lost an average of 40 minutes of notes per day until they switched back to paper.
Agreed conflict rules.
| Scenario | Rule | Seen per 1,000 syncs |
|---|---|---|
| same field edited twice | latest timestamp wins, logged | 12 |
| job closed on server, edited offline | manual review queue | 3 |
| new record created offline | kept, synced with a new id | 85 |
I would not consider it settled without evidence: Test the app through a 4-hour disconnection with conflicting edits and confirm no data loss and correct conflict handling.
Design for the disconnection, not the demo network.
Curated: · Written: · Reviewed:
QA-44The customer sends 200 GB of CSV a day. Would you change the format, and why?(show answer)
What an interviewer listens for on file formats for large exports is evidence from real users and real data.
Columnar formats such as Parquet store types, compress well, and let queries read only the needed columns, which usually cuts storage and query time sharply compared with CSV. CSV remains useful as a simple exchange format at the boundary.
Concretely, accept CSV at the boundary if the customer cannot change it, convert to Parquet on ingestion with an explicit schema, partition by date, and query the converted data.
The reason for that specificity is a failure I have seen: Analysts queried 200 GB of raw CSV each morning, each query scanned every column, and daily reports took 3 hours to run.
Same data, two formats.
| Format | Daily size | Typical report time |
|---|---|---|
| CSV | 200 GB | 3 h |
| Parquet, compressed | 38 GB | 6 min |
I would not consider it settled without evidence: Compare storage size and query time for a typical report before and after conversion.
Convert once at the door, then query fast.
Curated: · Written: · Reviewed:
QA-45A key report at the customer takes 10 minutes to load. How do you find the cause?(show answer)
I would handle debugging a slow query at a customer site so that it still works after I have left the customer.
Slow queries usually come from scanning too much data, missing indexes or partitions, or joins that multiply rows. The query plan shows which step is expensive, so diagnosis starts there rather than with guesses.
Concretely, capture the actual query, read its execution plan, compare rows scanned with rows returned, add filters, indexes, or pre-aggregation for the expensive step, and measure again.
The reason for that specificity is a failure I have seen: A team doubled the database size to fix a slow report, cost rose by 9,000 dollars a month, and the report stayed slow because it joined two large tables without an index on the join key.
Each fix measured.
| Change | Rows scanned | Runtime |
|---|---|---|
| original | 380 million | 10 min |
| index on join key | 2.1 million | 14 s |
| pre-aggregated daily table | 90,000 | 1.2 s |
I would not consider it settled without evidence: Compare the execution plan and runtime before and after each change.
Read the plan before buying hardware.
Curated: · Written: · Reviewed:
QA-46Your system needs to write decisions back into the customer's ERP. What makes this risky and how do you control it?(show answer)
The honest answer on writing results back to a customer system includes what could go wrong at the customer and how I would know.
Writing back changes the customer's system of record, so errors spread into their operations. Writes must be validated, authorized, idempotent, auditable, and reversible where possible.
Concretely, validate each write against the target's rules, use idempotency keys, write in small batches, log every change with the user and reason, and start with a human-approved mode before automating.
The reason for that specificity is a failure I have seen: An automated write-back sent 3,400 price changes with a decimal-place error to the customer's ERP, and the error reached the online store for 2 hours before it was reversed.
Stages of write-back automation.
| Stage | Writes per day | Approval |
|---|---|---|
| staging test | 500 | all reviewed |
| production, approved mode | 200 | every write |
| production, automated | 3,000 | 2% sampled |
I would not consider it settled without evidence: Run write-backs in a staging copy of the target first and require a sample of changes to be approved by a user before full automation.
Write back slowly, with a trail and an undo.
Curated: · Written: · Reviewed:
QA-47Why does a customer need an audit log of what the system and users did, and what should it contain?(show answer)
With audit logging for decisions, the customer's constraints come first and the architecture follows them.
Regulated and high-stakes decisions must be reconstructable after the fact, which means recording who or what made each decision, when, on which data and version, and why. Application error logs are not an audit trail.
Concretely, record each decision with actor, timestamp, inputs or references to them, rule or model version, output, and any override, in append-only storage with retention matching the customer's policy.
The reason for that specificity is a failure I have seen: A regulator asked why 60 loan applications were declined, and the customer could not answer because the system logged only errors, not decisions.
One audit record.
| Field | Example |
|---|---|
| actor | model v3.2 plus reviewer j.smith |
| inputs | application 88123, bureau file 2026-09-01 |
| output | declined, reason code R14 |
| override | none |
I would not consider it settled without evidence: Reconstruct one past decision entirely from the audit log in front of the customer's compliance team.
If you cannot reconstruct a decision, you cannot defend it.
Curated: · Written: · Reviewed:
QA-48The customer asks why last month's report changed. How do you make that answerable?(show answer)
I would make versioning data and transformations measurable against the baseline the customer already trusts.
Reports change when source data is corrected or transformation logic changes, and both need to be versioned so any past figure can be reproduced and every change explained.
Concretely, keep transformation code in version control, stamp outputs with the code version and data snapshot, retain snapshots for the agreed period, and log logic changes with their expected effect.
The reason for that specificity is a failure I have seen: A monthly revenue figure changed by 4 percent after a logic fix, nobody could reproduce the original, and the customer's finance team spent a week reconciling it by hand.
Two runs of the same month.
| Report run | Code version | Data snapshot | Revenue |
|---|---|---|---|
| Sep 1 | v41 | 2026-08-31 | $4.10m |
| Sep 8 | v42, refund fix | 2026-08-31 | $3.94m |
I would not consider it settled without evidence: Reproduce last month's report exactly from the stored code version and data snapshot.
Every number should be reproducible from a version.
Curated: · Written: · Reviewed:
QA-49How do you test a data pipeline built for a customer's messy data?(show answer)
I would start testing pipelines with real fixtures in the customer's workflow rather than in our product catalogue.
Tests built on idealized data miss the cases that break production, so fixtures should include real examples of nulls, duplicates, bad formats, and edge dates, anonymized where needed.
Concretely, collect examples of every data problem seen, turn them into a small fixture with expected outputs, run the tests on each change, and add a new case every time production data surprises you.
The reason for that specificity is a failure I have seen: A pipeline passed 120 unit tests on synthetic data and failed on its first day because real addresses contained line breaks inside quoted fields.
Fixture cases from real data.
| Fixture case | Expected behaviour |
|---|---|
| line break inside a quoted field | parsed as one row |
| duplicate order id | second copy rejected |
| date 1900-01-01 | treated as missing |
| total cases in fixture | 64 |
I would not consider it settled without evidence: Keep a fixture of real problem records with expected outputs and add each new production surprise to it.
Test with the data that will break you.
Curated: · Written: · Reviewed:
QA-50Sensor readings arrive minutes to hours late and out of order. How do you compute correct hourly totals?(show answer)
The first thing I want to know about late and out-of-order events is which decision it changes for which user.
Grouping by arrival time gives wrong totals when events are late, so events must be grouped by the time they occurred, with a rule for how long to wait before an hour is considered complete and how to correct it afterwards.
Concretely, aggregate by event time, keep hours open for a set lateness window, recompute and restate hours when late data arrives, and mark recent hours as provisional in reports.
The reason for that specificity is a failure I have seen: A dashboard grouped sensor readings by arrival time, a network outage delivered 3 hours of readings at once, and the plant saw a false spike that triggered an unnecessary shutdown.
Three hours delivered at once.
| Hour | By arrival time | By event time |
|---|---|---|
| 10:00 | 0 | 1,200 |
| 11:00 | 0 | 1,180 |
| 12:00 | 3,610 | 1,230 |
I would not consider it settled without evidence: Replay a day of delayed and reordered events and confirm the final hourly totals match the true event-time totals.
Count events when they happened, not when they arrived.
Curated: · Written: · Reviewed:
QA-51A customer wants an AI assistant that answers questions about their internal procedures. Would you use prompting, retrieval, or fine-tuning?(show answer)
For choosing prompting, retrieval, or fine-tuning, I would name the outcome, the owner, and the measurement before any code.
Knowledge that changes and must be cited is best supplied through retrieval at query time, while fine-tuning mainly changes style or format and does not reliably add facts. Start with prompting plus retrieval and move to fine-tuning only for a measured gap.
Concretely, build an evaluation set from real questions, try a strong model with retrieval first, measure accuracy and citation quality, and consider fine-tuning only if a specific, repeated failure remains.
The reason for that specificity is a failure I have seen: A team fine-tuned a model on the customer's procedure manuals, the procedures were updated 6 weeks later, and the assistant kept giving the old instructions with confidence.
Approaches on 300 real questions.
| Approach | Accuracy | Cost of a procedure update |
|---|---|---|
| prompt only | 52% | none |
| prompt plus retrieval | 86% | re-index documents |
| fine-tuned, no retrieval | 71% | retrain |
I would not consider it settled without evidence: Compare approaches on the same evaluation set of real questions and record accuracy, citation correctness, and cost.
Retrieve changing facts, and fine-tune only for behaviour.
Curated: · Written: · Reviewed:
QA-52How do you build an evaluation set for a customer's AI workflow?(show answer)
My approach to building an evaluation set from customer examples separates what the customer asked for from what they need.
An evaluation set should be drawn from real cases, labelled with the correct outcome by the customer's experts, and cover common cases, known hard cases, and cases where the right answer is to refuse or escalate.
Concretely, sample real inputs, have domain experts label the expected output and acceptable variations, include edge and refusal cases, freeze a version for comparisons, and add every production failure to it.
The reason for that specificity is a failure I have seen: A team evaluated on 50 questions they wrote themselves, scored 94 percent, and saw 61 percent correctness when real users asked questions phrased very differently.
An evaluation set of 400 cases.
| Slice | Cases | Purpose |
|---|---|---|
| common questions | 250 | typical accuracy |
| known hard cases | 80 | failure modes |
| should refuse or escalate | 40 | safety |
| past production failures | 30 | regressions |
I would not consider it settled without evidence: Show the evaluation set's source, size, label agreement between experts, and the share of real production failures it contains.
Evaluate on the customer's real cases, labelled by the customer's experts.
Curated: · Written: · Reviewed:
QA-53Your deployment is live and working. How do you hand it off so it keeps working after you leave?(show answer)
I treat handing a live deployment to customer owners who keep it running as something I own end to end, from discovery to adoption.
A handoff is complete when the customer's own named owners can run, diagnose, and change the system without you. Anything they cannot maintain must either move into the core product with a date, or sit inside an explicit and bounded support agreement.
Concretely, name a customer owner for every operational duty, write runbooks keyed to specific alerts rather than to components, rehearse a real incident with the owners before you leave, decide per component whether it stays field-maintained or moves into core, and agree support limits, response times and end dates in writing.
The reason for that specificity is a failure I have seen: The handoff was a 48-page Confluence space and an introduction to the IT manager. Ten days after the FDE left, a vendor schema change broke the nightly load on a Thursday; the alert mailed a shared mailbox nobody owned, nine loads failed before anyone opened a ticket, and the operations team went back to spreadsheets while the renewal was still in review.
Handoff matrix at roll-off.
| Duty | Customer owner | Artefact | Support after handoff |
|---|---|---|---|
| daily ingest and reruns | two named data ops | alert-keyed runbook | business hours, 1 h P1 response, ends 30 Sep |
| prompt and evaluation changes | ML lead | eval runbook plus frozen eval set | field-maintained, review in 90 days |
| custom ERP connector | none — they cannot own it | — | productized into core, ticket PROD-4412, ships Q3 |
| access and offboarding | IT service desk | one-page runbook | customer only |
I would not consider it settled without evidence: Ask the customer's named owner to work a real alert end to end while you watch and say nothing, and record the time to diagnosis, what they handled, and what they had to escalate.
You have handed off when their owner diagnoses the next failure without calling you.
Curated: · Written: · Reviewed:
QA-54The customer's lawyers worry the assistant will make things up. How do you reduce and expose unsupported answers?(show answer)
I would anchor reducing unsupported answers with citations in evidence the customer already trusts.
Requiring answers to cite retrieved passages, checking that each claim is supported by a cited passage, and allowing the system to say it does not know reduce unsupported answers and let users verify the rest.
Concretely, retrieve relevant passages, instruct the model to answer only from them with citations, run an automated check that cited passages support each claim, and return an explicit no-answer when support is missing.
The reason for that specificity is a failure I have seen: An assistant gave confident answers without sources, and a lawyer relied on a clause that did not exist in the contract, which was discovered only in negotiation.
Effect of citation checks.
| Setting | Supported answers | No-answer rate |
|---|---|---|
| no citations required | 78% | 0% |
| citations plus support check | 96% | 9% |
I would not consider it settled without evidence: Measure the share of answers whose claims are supported by their cited passages on a labelled evaluation set.
An answer the user cannot check should not be given.
Curated: · Written: · Reviewed:
QA-55Your assistant often answers from the wrong document. How do you improve retrieval?(show answer)
On improving retrieval quality, I would rather ship a narrow slice that users rely on than a broad one they ignore.
Wrong answers often come from retrieving the wrong passages, caused by poor chunking, missing keyword matching, or no re-ranking. Measuring retrieval separately from the answer shows where to fix.
Concretely, label which passages should be retrieved for a set of questions, measure how often they appear in the top results, and then try chunking by document structure, hybrid keyword and vector search, and a re-ranker, measuring each change.
The reason for that specificity is a failure I have seen: A team spent weeks changing prompts to fix wrong answers, when retrieval returned the correct passage in the top 5 only 48 percent of the time.
Correct passage in the top 5.
| Change | Hit rate |
|---|---|
| fixed 500-token chunks, vector only | 48% |
| chunk by section headings | 63% |
| hybrid keyword plus vector | 79% |
| plus re-ranker | 88% |
I would not consider it settled without evidence: Measure the share of questions whose correct passage appears in the top 5 retrieved results before and after each change.
Fix retrieval before blaming the model.
Curated: · Written: · Reviewed:
QA-56Documents in the customer's knowledge base have different access rights. How do you keep the assistant from leaking them?(show answer)
What an interviewer listens for on permission-aware retrieval is evidence from real users and real data.
Retrieval must apply the same permissions as the source system at query time, so each user can only retrieve passages they are allowed to read. Filtering after the model has seen restricted text is too late.
Concretely, store access metadata with each indexed chunk, filter retrieval by the requesting user's groups, sync permission changes quickly, and test with users of different access levels.
The reason for that specificity is a failure I have seen: An assistant indexed all HR documents without access metadata, and an employee received a summary of a colleague's salary review in answer to a general question.
Retrieval tests by access level.
| User | Can read | Salary review retrieved |
|---|---|---|
| HR manager | all HR documents | yes |
| employee | policies only | no |
| contractor | public documents | no |
| restricted passages leaked in 500 tests | none | 0 |
I would not consider it settled without evidence: Test retrieval with accounts of different access levels and confirm no restricted passage is ever returned to an unauthorized user.
Filter by permission before the model sees anything.
Curated: · Written: · Reviewed:
QA-57The model extracts fields from invoices into JSON. How do you make the output safe for downstream systems?(show answer)
I would handle validating structured AI output so that it still works after I have left the customer.
Model output must be validated like any external input, against a schema and business rules, because models occasionally produce wrong types, missing fields, or plausible but wrong values.
Concretely, use structured output with a schema, validate types and required fields, check business rules such as line items summing to the total, route failures to review, and track the failure rate.
The reason for that specificity is a failure I have seen: Extracted invoice totals were written straight into the payment system, and 38 invoices with a misread decimal were paid at ten times their value before a supplier reported it.
Invoices failing each check.
| Check | Invoices failing |
|---|---|
| schema and types | 0.4% |
| line items sum to total | 1.1% |
| amount above the supplier's usual range | 0.6% |
| routed to human review | 2.1% |
I would not consider it settled without evidence: Check that every extracted invoice passes schema and arithmetic rules before it reaches the payment system, and count the failures.
Treat model output as untrusted input.
Curated: · Written: · Reviewed:
QA-58When should a person approve an AI system's output, and how do you design that step?(show answer)
The honest answer on designing human-in-the-loop review includes what could go wrong at the customer and how I would know.
Human approval belongs where errors are costly or irreversible and where the model's confidence is low. The review step must show the evidence and make disagreement easy, or reviewers will approve by default.
Concretely, route high-impact and low-confidence cases to review, show the source evidence next to the proposed output, record approvals and edits, and measure reviewer agreement and time per case.
The reason for that specificity is a failure I have seen: Reviewers approved 99.6 percent of AI-proposed claim decisions in an average of 4 seconds each, and an audit found 7 percent of approved decisions were wrong.
Routing by case type.
| Case type | Route | Review time |
|---|---|---|
| low value, high confidence | automatic | none |
| high value | human approval | 3 min |
| low confidence | human approval | 5 min |
I would not consider it settled without evidence: Measure how often reviewers change the proposed output and audit a sample of approvals independently.
A review step is only real if reviewers can and do disagree.
Curated: · Written: · Reviewed:
QA-59An AI agent can call tools that change customer records. How do you keep it safe?(show answer)
With guarding AI agents with deterministic controls, the customer's constraints come first and the architecture follows them.
The model can propose actions, but code must enforce which actions are allowed, for whom, with what limits, and with what approvals. Permissions and limits should never depend on the model following instructions.
Concretely, give the agent narrowly scoped tools, check every call against an allowlist and the user's permissions, set limits on amounts and counts, require approval for irreversible actions, and log every call.
The reason for that specificity is a failure I have seen: An agent with a general database tool deleted 1,200 duplicate-looking customer records during a cleanup task because nothing limited which records it could remove.
Tool permissions enforced in code.
| Tool | Allowed | Limit | Approval |
|---|---|---|---|
| read customer record | yes | own region | no |
| update address | yes | 50 per hour | no |
| issue refund | yes | up to $200 | above $200 |
| delete record | no | none | not available |
I would not consider it settled without evidence: Test the agent with instructions to exceed its permissions and confirm the code blocks each attempt.
Let the model suggest, and let code decide what is allowed.
Curated: · Written: · Reviewed:
QA-60The assistant reads emails and documents from outside the company. What is the risk and how do you reduce it?(show answer)
I would make prompt injection from customer documents measurable against the baseline the customer already trusts.
Text inside documents can contain instructions that the model may follow, such as asking it to reveal data or take an action. Content from documents must be treated as data, and the system's powers must be limited so a successful injection does little harm.
Concretely, separate instructions from retrieved content, restrict the tools and data the assistant can access, require approval for sensitive actions, and test with injected instructions in realistic documents.
The reason for that specificity is a failure I have seen: A supplier email contained a hidden line telling the assistant to forward the latest invoice list, and the assistant included it in a draft reply that a busy user almost sent.
Injection tests before and after controls.
| Injected instruction | Without controls | With controls |
|---|---|---|
| reveal other customers' data | leaked | blocked by access filter |
| send an email | drafted and sent | draft requires approval |
| ignore previous rules | followed | no effect on tool limits |
| attacks succeeding out of 120 | 31 | 0 |
I would not consider it settled without evidence: Run a test set of documents containing injected instructions and confirm no sensitive action or disclosure occurs.
Assume documents will try to give orders.
Curated: · Written: · Reviewed:
QA-61Users say the assistant is too slow. How do you set and meet a latency target?(show answer)
I would start latency budgets for AI workflows in the customer's workflow rather than in our product catalogue.
A latency target should come from the workflow, such as a call-centre agent needing an answer while the customer waits. The total time must be split across retrieval, model calls, and post-processing to see where it goes.
Concretely, measure time per stage, set a budget for each, stream partial output, reduce context size, cache repeated work, use a smaller model for simple steps, and track p95 rather than the average.
The reason for that specificity is a failure I have seen: A support assistant averaged 6 seconds but had a p95 of 28 seconds because some questions triggered 5 sequential model calls, and agents stopped using it during busy hours.
p95 latency by stage.
| Stage | Before | After |
|---|---|---|
| retrieval | 1.8 s | 0.6 s |
| model calls | 24.0 s | 5.1 s |
| post-processing | 2.2 s | 0.4 s |
| total | 28.0 s | 6.1 s |
I would not consider it settled without evidence: Measure p50 and p95 latency per stage on real traffic and compare with the budget.
Budget latency per step and watch the slow tail.
Curated: · Written: · Reviewed:
QA-62How do you estimate what an AI workflow will cost to run for a customer?(show answer)
The first thing I want to know about estimating the cost of an AI deployment is which decision it changes for which user.
Running cost is roughly requests times tokens per request times price per token, plus infrastructure and human review. Estimating before launch avoids a system that works but costs more than the value it creates.
Concretely, measure tokens per request on real examples, multiply by expected volume and current prices, add retrieval and hosting costs, compare with the value per case, and look for savings such as smaller models or caching.
The reason for that specificity is a failure I have seen: A document assistant sent entire 200-page contracts with every question, cost 3.40 dollars per question, and the customer paused it after the first monthly bill.
Cost per question by design, at $0.02 per 1,000 tokens.
| Design | Tokens per question | Cost per question |
|---|---|---|
| whole contract in context | 170,000 | $3.40 |
| retrieve top 8 passages | 6,000 | $0.12 |
| plus small model for routing | 4,000 | $0.08 |
I would not consider it settled without evidence: Compute cost per request from measured token counts and compare it with the agreed value per case before scaling.
Price the request before you scale it.
Curated: · Written: · Reviewed:
QA-63What happens to your customer's workflow if the model provider has an outage?(show answer)
For provider and model fallback, I would name the outcome, the owner, and the measurement before any code.
A critical workflow needs a defined degraded mode, such as a secondary model, a cached answer, or a manual path, and the switch must be tested. Without it, a provider outage becomes a customer outage.
Concretely, set timeouts, detect provider errors, fail over to a tested secondary model or a manual queue, tell users the system is in degraded mode, and rehearse the failover regularly.
The reason for that specificity is a failure I have seen: A provider outage lasted 3 hours, the claims workflow had no fallback, and 2,600 claims queued with no way for staff to work them.
Degraded modes and rehearsals.
| Failure | Degraded mode | Rehearsed |
|---|---|---|
| primary model down | secondary model, same prompts | monthly |
| both models down | manual queue with checklist | quarterly |
| retrieval index down | route to manual queue | monthly |
I would not consider it settled without evidence: Simulate a provider outage in a test and measure whether the workflow continues in the defined degraded mode.
Plan the outage before the provider has one.
Curated: · Written: · Reviewed:
QA-64The model provider releases a new version. How do you decide whether to switch?(show answer)
My approach to regressions when the model version changes separates what the customer asked for from what they need.
A new model can improve average results while breaking specific cases the customer depends on, so a switch should be tested on the frozen evaluation set and a sample of recent traffic before rollout.
Concretely, run the evaluation set on both versions, compare overall and per-slice results, review changed outputs, roll out to a small share of traffic first, and keep the old version available for rollback.
The reason for that specificity is a failure I have seen: A team switched to a newer model version for its better benchmark scores, and the invoice extractor started returning dates in a different format, breaking 13 percent of payments.
Old and new version by slice.
| Slice | Old version | New version |
|---|---|---|
| common invoices | 97% | 98% |
| foreign formats | 91% | 77% |
| date fields | 99% | 86% |
I would not consider it settled without evidence: Compare both versions on the evaluation set by slice and review every case whose output changed.
Test the new model on your cases, not on its benchmarks.
Curated: · Written: · Reviewed:
QA-65Should an AI system make decisions automatically or assist a person? How do you decide?(show answer)
I treat deciding between automation and assistance as something I own end to end, from discovery to adoption.
Automate when errors are cheap, reversible, and rare enough to accept, and assist when errors are costly or the system's accuracy on a case type is not yet proven. The choice can differ by case type within one workflow.
Concretely, measure accuracy by case type, estimate the cost of an error for each, automate only the types with low error cost and proven accuracy, and keep a human in the loop for the rest.
The reason for that specificity is a failure I have seen: A team fully automated refund decisions to hit a cost target, the model was wrong on 6 percent of high-value refunds, and losses exceeded the savings within two months.
Mode chosen by case type.
| Case type | Accuracy | Cost of an error | Mode |
|---|---|---|---|
| refund under $50 | 98% | up to $50 | automatic |
| refund over $500 | 94% | $500 or more | assist |
| suspected fraud | 88% | high | assist |
I would not consider it settled without evidence: Show accuracy and error cost by case type and the rule deciding which types are automated.
Automate by case type, not all at once.
Curated: · Written: · Reviewed:
QA-66Design the smallest production architecture for an AI workflow that reads incoming documents and proposes actions.(show answer)
I would anchor the smallest safe production architecture in evidence the customer already trusts.
The smallest safe architecture separates ingestion, interpretation by the model, validation by code, human approval where needed, and a write step with an audit trail. Each part can then fail and recover on its own.
Concretely, receive documents into a queue, extract and propose with the model, validate against schemas and rules, route to approval or automatic handling, write idempotently to the target system, and log each step.
The reason for that specificity is a failure I have seen: A prototype read documents, called the model, and wrote to the customer's system in one request, and a timeout retry created 417 duplicate actions before the integration was disabled.
Failure handling per stage.
| Stage | Failure handling |
|---|---|
| ingest queue | redelivery safe by document id |
| model proposal | timeout returns item to the queue |
| validation | failed items go to review |
| write | one idempotency key per action |
| duplicate effects in 1,000 fault tests | 0 |
I would not consider it settled without evidence: Test duplicate delivery, model timeout, validation failure, and write failure, and confirm exactly one correct effect each time.
Separate the steps so each can fail safely.
Curated: · Written: · Reviewed:
QA-67The customer wants the system deployed in their own cloud account. What do you need to plan?(show answer)
On deploying into a customer's cloud account, I would rather ship a narrow slice that users rely on than a broad one they ignore.
Running in the customer's account means their identity system, network rules, secrets management, and change process apply, and you may not have direct access in production. Infrastructure as code and clear operating responsibilities make this workable.
Concretely, deliver the deployment as code, request least-privilege roles, store secrets in the customer's secret manager, agree who operates what, and rehearse a deployment and rollback in their staging account.
The reason for that specificity is a failure I have seen: A deployment relied on an FDE's personal cloud credentials, and when the FDE left the project the customer could not update or restart the system for 2 weeks.
Ownership in the customer's account.
| Item | Owner |
|---|---|
| infrastructure code | vendor, reviewed by the customer |
| cloud roles | customer IT |
| secrets | customer secret manager |
| on-call | shared, per runbook, 2 named people each side |
I would not consider it settled without evidence: Deploy and roll back the system in the customer's staging account using only the documented roles and code.
Deploy it so the customer can run it without you.
Curated: · Written: · Reviewed:
QA-68How do you ship changes to a customer deployment safely and often?(show answer)
What an interviewer listens for on continuous delivery for a customer deployment is evidence from real users and real data.
Small, frequent, automated releases with tests and a quick rollback are safer than large manual ones. The customer's change process can be satisfied by showing the pipeline's checks and approvals.
Concretely, build a pipeline that runs tests, deploys to staging, runs smoke tests on real workflows, requires the agreed approval, deploys to production, and can roll back in minutes.
The reason for that specificity is a failure I have seen: A team shipped monthly manual releases with 40 changes each, and a failed release took 9 hours to back out because nobody knew which change caused the problem.
Release measures before and after.
| Measure | Before | After |
|---|---|---|
| releases per month | 1 | 18 |
| changes per release | 40 | 2 |
| rollback time | 9 h | 6 min |
I would not consider it settled without evidence: Measure deployment frequency, change failure rate, and time to roll back for each release.
Ship small and roll back fast.
Curated: · Written: · Reviewed:
QA-69What would you monitor for a deployed customer workflow?(show answer)
I would handle workflow service level indicators so that it still works after I have left the customer.
The most useful indicators measure what users experience, such as cases completed on time, errors users see, and latency, rather than only server health. A server can be healthy while the workflow is failing.
Concretely, define indicators for completion, correctness, and latency of the workflow, set targets with the customer, alert on sustained breaches and stuck work, and review them weekly.
The reason for that specificity is a failure I have seen: A dashboard showed 100 percent server uptime while a stuck queue stopped 1,800 cases from progressing for a whole day.
Workflow indicators and targets.
| Indicator | Target |
|---|---|
| cases completed within 4 hours | 95% |
| oldest item in queue | under 30 min |
| user-visible errors | under 0.5% |
| p95 page latency | under 2 s |
I would not consider it settled without evidence: Alert on workflow completion and queue age as well as server health, and review their trends weekly.
Monitor the workflow, not just the servers.
Curated: · Written: · Reviewed:
QA-70How do you roll out a new system to 2,000 users without a big-bang launch?(show answer)
The honest answer on rolling out in cohorts includes what could go wrong at the customer and how I would know.
A cohort rollout starts with a small trained group, compares their results with the current process, and expands only when agreed thresholds are met. Problems stay small and are found early.
Concretely, run in shadow mode first, launch to a pilot group with a support owner, define expansion and rollback thresholds, expand in steps such as 5, 25, and 100 percent, and hold each step for a stable period.
The reason for that specificity is a failure I have seen: A system launched to all 2,000 users on one Monday, missing permissions blocked 30 percent of them, and the support team received 900 tickets in the first day.
A staged rollout.
| Step | Users | Hold period | Expansion condition |
|---|---|---|---|
| shadow | 0 | 2 weeks | agreement over 90% |
| pilot | 100 | 2 weeks | no severe issues |
| wave 2 | 500 | 1 week | error rate under 1% |
| full | 2,000 | ongoing | none |
I would not consider it settled without evidence: Record results at each cohort against the thresholds and the decision to expand, hold, or roll back.
Expand when the evidence says so, not when the calendar does.
Curated: · Written: · Reviewed:
QA-71What makes a rollback plan real rather than a line in a document?(show answer)
With planning a rollback, the customer's constraints come first and the architecture follows them.
A real rollback plan states the trigger, the steps, the owner, the time it takes, and what happens to data created by the new system, and it has been rehearsed. Data changes are usually what make rollback hard.
Concretely, define rollback triggers in numbers, keep the old path available during rollout, plan how new records return to the old system, rehearse the rollback in staging, and time it.
The reason for that specificity is a failure I have seen: A rollback was triggered after launch problems, but 6,000 records created in the new system had no way back to the old one, and staff re-entered them by hand over a weekend.
A rehearsed rollback plan.
| Element | Example |
|---|---|
| trigger | error rate over 2% for 30 min |
| owner | on-call FDE and customer ops lead |
| duration | 20 minutes, rehearsed |
| data plan | new records exported to the old system nightly |
I would not consider it settled without evidence: Rehearse the full rollback, including data created by the new system, and record how long it took.
A rollback you have not rehearsed is only a hope.
Curated: · Written: · Reviewed:
QA-72The system goes down during the customer's busiest hour. What do you do?(show answer)
I would make responding to an incident at a customer measurable against the baseline the customer already trusts.
Restore the customer's ability to work first, using the fallback or rollback, communicate status early and regularly, and investigate the cause afterwards. Diagnosing while users wait extends the damage.
Concretely, declare the incident, switch to the degraded mode or roll back, send updates at fixed intervals, record a timeline, and hold a blameless review with actions that have owners and dates.
The reason for that specificity is a failure I have seen: An FDE spent 90 minutes debugging live while 300 agents could not work, when switching to the manual fallback would have restored service in 5 minutes.
An incident timeline.
| Time | Action |
|---|---|
| 10:02 | alert fires on queue age |
| 10:06 | incident declared, fallback enabled |
| 10:10 | first customer update |
| next day | review with 4 owned actions |
I would not consider it settled without evidence: Record time to detect, time to restore, and the follow-up actions with owners for each incident.
Restore first, diagnose second.
Curated: · Written: · Reviewed:
QA-73The customer expects five times normal volume during storm season. How do you prepare?(show answer)
I would start planning capacity for a peak event in the customer's workflow rather than in our product catalogue.
Peak events expose limits in every dependency, including the customer's systems, the model provider's rate limits, and staff capacity for review. Load testing at the expected peak shows which one breaks first.
Concretely, estimate peak volume from past events, load test end to end at that level with a margin, raise rate limits and quotas in advance, prepare queue-based smoothing, and agree which work can be delayed.
The reason for that specificity is a failure I have seen: A claims system handled normal volume but hit the model provider's rate limit at twice normal load during a storm, and 11,000 claims waited up to 3 days.
Capacity against the expected peak.
| Component | Capacity | Expected peak |
|---|---|---|
| intake API | 200 per second | 90 per second |
| model provider quota | 60 per second | 75 per second, raise quota |
| human review staff | 1,500 per day | 4,000 per day, triage rule |
I would not consider it settled without evidence: Run an end-to-end load test at the expected peak plus a margin and record which component limits throughput.
Test the peak before the peak tests you.
Curated: · Written: · Reviewed:
QA-74You have built the same integration for three customers. What do you do?(show answer)
The first thing I want to know about deciding when to productize field work is which decision it changes for which user.
Work repeated across customers is a signal for the product, and turning it into a shared component reduces cost and risk. The case should be made with evidence of frequency, effort, and customer impact.
Concretely, document each instance, the effort spent, and the differences between them, propose a generalized component with the product team, and help build it using the field examples as test cases.
The reason for that specificity is a failure I have seen: Seven FDEs each built a slightly different connector for the same accounting system, and a single upstream change broke all seven in the same week.
Repeated builds of one connector.
| Customer | Hours spent | Differences |
|---|---|---|
| A | 120 | custom field mapping |
| B | 95 | different auth method |
| C | 140 | extra invoice types |
| shared component estimate | 200, once | configuration |
I would not consider it settled without evidence: Present the product team with the list of repeated builds, the hours spent, and the proposed shared design.
Build it three times, then make it a product.
Curated: · Written: · Reviewed:
QA-75How do you make sure what you learn at customers actually changes the product?(show answer)
For handing field learning back to the product, I would name the outcome, the owner, and the measurement before any code.
Field learning only changes the product when it reaches the people who set priorities, in a form they can act on, with frequency, impact, and examples. Anecdotes from one meeting rarely move a roadmap.
Concretely, keep a running log of field problems with counts and affected customers, summarize the top items monthly for product managers, attach real examples, and follow up on decisions.
The reason for that specificity is a failure I have seen: Users at 5 customers built the same spreadsheet workaround for a missing bulk-edit feature, but the product team never heard about it because feedback went only into individual project notes.
A monthly field summary.
| Field problem | Customers affected | Hours lost per month |
|---|---|---|
| no bulk edit | 5 | 320 |
| slow export | 3 | 90 |
| missing audit filter | 2 | 40 |
I would not consider it settled without evidence: Show the monthly field summary sent to product, with each item's frequency, impact, and the decision taken.
Turn field stories into counted evidence.
Curated: · Written: · Reviewed:
QA-76How do you measure whether users have actually adopted the new system?(show answer)
My approach to measuring adoption properly separates what the customer asked for from what they need.
Logins and page views show access, not adoption, so the useful measure is the share of real work completed through the new system and whether users return to old tools. Adoption should be measured per team, because averages hide groups that never switched.
Concretely, count cases completed in the new system as a share of all cases, track by team and shift, look for work still done in spreadsheets or email, and interview teams with low shares.
The reason for that specificity is a failure I have seen: A deployment reported 95 percent of users logged in weekly, while only 38 percent of cases were actually processed in the new system and the rest still went through email.
Logins against completed work.
| Team | Weekly logins | Cases completed in system |
|---|---|---|
| day shift | 98% | 71% |
| night shift | 92% | 12% |
| overall | 95% | 38% |
I would not consider it settled without evidence: Report the share of cases completed in the new system by team and compare it with the share still handled in old tools.
Adoption is work done in the system, not logins.
Curated: · Written: · Reviewed:
QA-77How do you get busy operations staff to change how they work?(show answer)
I treat training users and managing change as something I own end to end, from discovery to adoption.
People adopt a new tool when it saves them effort on their real work and when their managers expect them to use it. Training on real cases and fixing early friction quickly matter more than long sessions.
Concretely, train in short sessions on the users' own cases, name local champions, sit with users in the first days to fix friction, and agree with managers when the old path will be switched off.
The reason for that specificity is a failure I have seen: Users received a 3-hour classroom training on generic examples, and a month later most had forgotten it and returned to the old system.
A change plan.
| Activity | Timing |
|---|---|
| 30-minute session on own cases | launch day |
| champion per team | from week 1 |
| on-site support | first 5 days |
| old path switched off | week 6, if adoption is over 80% |
I would not consider it settled without evidence: Track adoption by team for four weeks after training and record each friction issue found and fixed.
Train on real work and fix friction fast.
Curated: · Written: · Reviewed:
QA-78After launch, many users go back to their spreadsheets. What do you do?(show answer)
I would anchor users reverting to spreadsheets in evidence the customer already trusts.
Users return to old tools because something the spreadsheet does is missing or slower in the new system. Finding that specific gap and fixing it is more effective than mandating use.
Concretely, ask users to show you the spreadsheet, list what it does that the system does not, rank those gaps by how many users they affect, fix the top ones, and measure whether spreadsheet use falls.
The reason for that specificity is a failure I have seen: Management mandated use of the new system without asking why people reverted, and users kept a shadow spreadsheet for bulk updates, so the official data was 2 days behind reality.
What the spreadsheet still does.
| Spreadsheet feature | Users relying on it | Fix |
|---|---|---|
| bulk update | 42 | bulk edit screen |
| custom filter | 30 | saved views |
| offline copy | 8 | export |
I would not consider it settled without evidence: Document the spreadsheet features users rely on and measure spreadsheet use after the top gaps are fixed.
The spreadsheet tells you what is missing.
Curated: · Written: · Reviewed:
QA-79How do you run a feedback loop with users during a deployment?(show answer)
On running a user feedback loop, I would rather ship a narrow slice that users rely on than a broad one they ignore.
Feedback is useful when it is collected close to the work, grouped by problem, and visibly acted on. Users stop reporting problems when nothing seems to happen.
Concretely, provide a one-click feedback option in the tool, review feedback weekly, group it into issues with counts, fix or explain the top items, and tell users what changed because of their input.
The reason for that specificity is a failure I have seen: A feedback form collected 400 comments in a month that nobody triaged, and users stopped reporting problems, including a data error that affected billing.
Feedback handled week by week.
| Week | Feedback items | Issues fixed | Reported back |
|---|---|---|---|
| 1 | 64 | 5 | yes |
| 2 | 41 | 7 | yes |
| 3 | 22 | 4 | yes |
I would not consider it settled without evidence: Publish a weekly summary to users of the feedback received, the top issues, and what was changed.
Close the loop, or the feedback stops.
Curated: · Written: · Reviewed:
QA-80The sales team wants a demo that shows everything working. How do you keep the demo honest?(show answer)
What an interviewer listens for on keeping a sales demo honest is evidence from real users and real data.
A demo that uses curated data and hides limits creates expectations the deployment cannot meet. Demos should use realistic data, show what is real versus planned, and avoid promising outcomes before discovery.
Concretely, build demos on anonymized realistic data, label planned features clearly, show one real end-to-end path rather than many staged ones, and brief the sales team on what not to promise.
The reason for that specificity is a failure I have seen: A demo showed instant answers on 20 clean documents, the customer signed expecting the same on 2 million messy files, and the first month was spent resetting expectations.
Each demo element labelled.
| Demo element | Status |
|---|---|
| document search | live |
| custom approval flow | configurable in 2 weeks |
| automatic coding of invoices | planned |
I would not consider it settled without evidence: Review the demo script with the FDE lead and mark every capability shown as live, configurable, or planned.
Demo what you can deliver.
Curated: · Written: · Reviewed:
QA-81It is three months in. How do you present results to the customer's executive sponsor?(show answer)
I would handle presenting results to the executive sponsor so that it still works after I have left the customer.
The sponsor needs to know whether the agreed outcome moved, what it is worth, what is at risk, and what decision is needed next. Results should be measured against the baseline agreed at the start.
Concretely, lead with the outcome against the baseline and target, translate it into money or time, name one or two risks with mitigations, and end with a specific ask such as expansion or a decision.
The reason for that specificity is a failure I have seen: An FDE presented a list of features delivered, the sponsor asked what had improved, and without a clear answer the expansion was postponed by two quarters.
The first slide.
| Metric | Baseline | Target | Now |
|---|---|---|---|
| claim triage time | 9 days | 1 day | 1.4 days |
| annual value | none | $3.3m | $2.9m run rate |
| next ask | none | none | expand to 2 regions |
I would not consider it settled without evidence: Check that the first slide shows the outcome metric against the agreed baseline and target.
Report the outcome first, then the work.
Curated: · Written: · Reviewed:
QA-82Why write a case study after a deployment, and what goes in it?(show answer)
The honest answer on writing a deployment case study includes what could go wrong at the customer and how I would know.
A case study turns one deployment into reusable knowledge for future FDEs, sales, and product, covering the problem, what was built, what worked, what failed, and the measured result.
Concretely, write it within two weeks of completion, include the baseline and result, the architecture, the hardest problems and how they were solved, and what you would do differently, and get the customer's approval before sharing it outside the company.
The reason for that specificity is a failure I have seen: A team repeated a 6-week data-access delay at a second customer in the same industry because nobody had written down how the first delay was resolved.
A case study outline.
| Section | Example |
|---|---|
| problem | 9-day claims backlog |
| result | 1.4-day triage |
| hardest problem | data access took 6 weeks |
| lesson | start the security review in week 1 |
I would not consider it settled without evidence: Check that the case study states the baseline, result, key problems, and lessons, and that the next similar project used it.
Write down what you learned while you still remember it.
Curated: · Written: · Reviewed:
QA-83A customer manager says the system is wrong all the time. How do you respond?(show answer)
With a customer complaint about accuracy, the customer's constraints come first and the architecture follows them.
Take the complaint seriously but turn it into specific examples and a measured error rate, because the fix depends on whether errors are frequent, concentrated in one case type, or a perception from one bad case.
Concretely, ask for concrete examples, check them, measure the error rate on a random sample, group errors by cause, fix the largest cause, and report back with the numbers.
The reason for that specificity is a failure I have seen: An FDE argued the system was 95 percent accurate on average, and the manager escalated, because errors on her team's case type were actually 30 percent.
Errors by case type.
| Case type | Sample | Errors | Rate |
|---|---|---|---|
| standard orders | 300 | 9 | 3% |
| returns | 100 | 30 | 30% |
| overall | 400 | 39 | 9.8% |
I would not consider it settled without evidence: Measure the error rate on a random sample by case type and share the results with the manager.
Answer a complaint with examples and a measured rate.
Curated: · Written: · Reviewed:
QA-84The customer insists on a launch date you believe is impossible. What do you do?(show answer)
I would make negotiating a deadline you cannot meet measurable against the baseline the customer already trusts.
Agreeing to a date you cannot meet damages trust more than a hard conversation now. The useful move is to show what can be delivered by that date and what the full scope actually needs.
Concretely, break the scope into parts with estimates, show what fits by the date, identify the risks, offer options such as a smaller first release, and let the sponsor choose with the facts.
The reason for that specificity is a failure I have seen: An FDE accepted a launch date to avoid conflict, the team worked 7 weekends, and the launch still slipped by 5 weeks with most features untested.
Scope against date.
| Option | Scope | Date |
|---|---|---|
| full scope | 3 workflows | 14 weeks |
| first release | 1 workflow | 6 weeks, the requested date |
| rushed full scope | 3 workflows, untested | 6 weeks, high risk |
I would not consider it settled without evidence: Present a written scope-versus-date comparison and record the option the sponsor chose.
Trade scope for date in the open.
Curated: · Written: · Reviewed:
QA-85Requirements are unclear and the customer cannot answer your questions yet. How do you make progress?(show answer)
I would start acting without complete requirements in the customer's workflow rather than in our product catalogue.
Waiting for complete requirements usually means waiting forever, so an FDE builds something concrete from reasonable assumptions and uses it to get feedback. A working draft gets better answers than a questionnaire.
Concretely, write down your assumptions, build a small working version on real data, show it to users within days, and change it based on their reactions.
The reason for that specificity is a failure I have seen: A team sent a 60-question requirements document to the customer, waited 5 weeks for answers, and the answers still did not match what users needed when they saw the first version.
Two ways to learn requirements.
| Approach | Time to first feedback | Assumptions corrected |
|---|---|---|
| requirements questionnaire | 5 weeks | 4 |
| working draft on real data | 4 days | 11 |
I would not consider it settled without evidence: Show users a working draft within the first week and record which assumptions they confirmed or corrected.
Build a draft to discover the requirements.
Curated: · Written: · Reviewed:
QA-86An FDE coding interview gives you a messy data file and 60 minutes. How do you approach it?(show answer)
The first thing I want to know about approaching a practical coding interview is which decision it changes for which user.
These exercises test whether you can explore data, write working code quickly, handle edge cases, and explain trade-offs, not whether you know a clever algorithm. Working, readable code with checks beats an unfinished elegant solution.
Concretely, spend the first minutes inspecting the data and restating the task, get a simple version working end to end, then handle edge cases, add a few checks, and explain what you would improve with more time.
The reason for that specificity is a failure I have seen: A candidate spent 40 minutes designing an elegant class hierarchy and had no working output when time ran out, while the data contained malformed rows they never looked at.
A 60-minute plan.
| Minutes | Activity |
|---|---|
| 0 to 5 | inspect data and restate the task |
| 5 to 30 | simple working version |
| 30 to 50 | edge cases and checks |
| 50 to 60 | explain trade-offs |
I would not consider it settled without evidence: Check that you have a working end-to-end result within the first half of the time.
Make it work, then make it right.
Curated: · Written: · Reviewed:
QA-87Something breaks while you are on site with the customer watching. How do you handle it?(show answer)
For debugging with the customer watching, I would name the outcome, the owner, and the measurement before any code.
Customers judge how you handle problems as much as whether problems happen. Staying calm, explaining what you are checking, restoring their work first, and following up with a clear cause builds more trust than a flawless demo.
Concretely, acknowledge the problem, restore the user's work with a workaround if possible, investigate methodically while narrating briefly, and send a written explanation and fix afterwards.
The reason for that specificity is a failure I have seen: An FDE hid a failing job during a customer visit and promised it was working, the customer found the error in their data the next day, and the relationship took months to recover.
Handling a live failure.
| Step | Time |
|---|---|
| acknowledge and give a workaround | 5 min |
| find the cause | 45 min |
| written explanation | within 24 h |
| permanent fix | within 1 week |
I would not consider it settled without evidence: Send the customer a short written explanation of the cause and the fix within a day of the incident.
Handle the failure in the open.
Curated: · Written: · Reviewed:
QA-88Tell me about a time you persuaded a customer to change their approach.(show answer)
My approach to changing a customer's mind with evidence separates what the customer asked for from what they need.
Interviewers want to hear how you used evidence and an understanding of the customer's goals, rather than authority, to change a decision. A good answer shows the situation, the evidence you gathered, how you presented it, and the result.
Concretely, use the STAR structure, make the evidence concrete, such as a pilot result or measured data, show that you understood the customer's concern, and end with the measured outcome.
The reason for that specificity is a failure I have seen: A candidate described winning an argument by escalating to the customer's boss, and the interviewer marked the answer down because it showed no evidence or understanding.
A STAR answer.
| Part | Example |
|---|---|
| situation | customer wanted to automate all approvals |
| evidence | pilot showed 11% errors on high-value cases |
| action | proposed automation under $1,000 only |
| result | 70% automated, error cost down 90% |
I would not consider it settled without evidence: Rehearse the story in under 2 minutes and confirm it contains the evidence and a measured result.
Persuade with evidence the customer can check.
Curated: · Written: · Reviewed:
QA-89Tell me about a deployment that failed and what you learned.(show answer)
I treat a deployment that failed as something I own end to end, from discovery to adoption.
Interviewers are testing honesty, ownership, and learning, so a strong answer names a real failure you contributed to, explains the cause without blaming others, and shows what you changed afterwards.
Concretely, choose a real failure with consequences, state your part in it, explain the root cause, describe what you did to recover, and give the specific practice you changed afterwards.
The reason for that specificity is a failure I have seen: A candidate described a failure caused entirely by a difficult customer, took no responsibility, and the panel concluded they would repeat the same mistake.
A failure story that shows learning.
| Part | Example |
|---|---|
| failure | launched without a rollback path |
| my part | skipped the rollback rehearsal to meet the date |
| impact | 2 days of manual work for 80 users |
| change | rollback rehearsal is now a launch gate |
I would not consider it settled without evidence: Check that your story states your own mistake and a specific practice you changed.
Own the failure and show the change.
Curated: · Written: · Reviewed:
QA-90How do you make the most of limited time on site with a customer?(show answer)
I would anchor working effectively on site in evidence the customer already trusts.
Time on site is most valuable for observing work, building relationships, and making decisions that are slow remotely. It should be planned around the questions that can only be answered in person.
Concretely, set goals for the visit in advance, schedule time with users as well as managers, observe real work, make decisions in person, and write up findings and actions before leaving.
The reason for that specificity is a failure I have seen: An FDE spent a 3-day visit coding at a desk, met only the project manager, and left without having seen the night shift that caused most of the delays.
A three-day visit plan.
| Day | Focus |
|---|---|
| 1 | shadow day and night shifts |
| 2 | decisions with the sponsor and IT |
| 3 | fix the top friction issue, write the summary |
I would not consider it settled without evidence: Write a visit summary listing the goals, who you met, what you observed, and the actions agreed.
Use the visit for what only a visit can do.
Curated: · Written: · Reviewed:
QA-91Your deployment is live and you are moving on. How do you hand it off?(show answer)
On handing off to support and customer success, I would rather ship a narrow slice that users rely on than a broad one they ignore.
A good handoff leaves the customer and support team able to run, fix, and change the system without you, which requires documentation, runbooks, monitoring, and a period of shadowed support.
Concretely, write runbooks for common problems, confirm monitoring and alerts go to the new owners, run joint on-call for a period, transfer access and credentials, and schedule a check-in after the handoff.
The reason for that specificity is a failure I have seen: An FDE left for a new project with no runbook, and the first failed nightly load took the support team 3 days to fix because nobody knew where the logs were.
A handoff checklist.
| Item | Done |
|---|---|
| runbook for the top 10 failures | yes |
| alerts routed to support | yes |
| 2 weeks of joint on-call | yes |
| access transferred | yes |
I would not consider it settled without evidence: Have the support team resolve a simulated failure using only the runbook before the handoff is complete.
Hand off when someone else can fix it without you.
Curated: · Written: · Reviewed:
QA-92While working at a customer, you discover their system exposes sensitive data. What do you do?(show answer)
What an interviewer listens for on finding a security problem at a customer is evidence from real users and real data.
Report it promptly to the customer's security contact and your own leadership through agreed channels, avoid accessing more data than needed to confirm it, and do not share details widely. Fixing it quietly yourself can breach the customer's processes.
Concretely, stop exploring once the issue is confirmed, document what you saw, report it to the named security contacts, follow the customer's incident process, and help with remediation if asked.
The reason for that specificity is a failure I have seen: An FDE found an open storage bucket and mentioned it in a shared chat channel with 200 members before telling the customer's security team, which widened the exposure.
Reporting timeline.
| Step | Target time |
|---|---|
| confirm the issue | minutes |
| report to customer security | within 1 hour |
| notify own leadership | within 1 hour |
| written record | same day |
I would not consider it settled without evidence: Record the time of discovery, the time of the report to the security contact, and the actions taken.
Report security issues quickly, narrowly, and through the right people.
Curated: · Written: · Reviewed:
QA-93A customer asks you to build something you believe is harmful or against policy. How do you handle it?(show answer)
I would handle a customer request that is unsafe or unethical so that it still works after I have left the customer.
An FDE should not build something that is unsafe, illegal, or against the company's policies, even for an important customer. The right response is to raise it clearly, involve your leadership, and look for a legitimate way to meet the underlying need.
Concretely, state the concern specifically, check the relevant policy or legal guidance, escalate to your leadership, and propose an alternative that meets the legitimate goal.
The reason for that specificity is a failure I have seen: A team built a monitoring feature that tracked individual employees' screens at a customer without review, and the resulting legal complaint cost both companies more than the contract was worth.
Declined requests with alternatives.
| Request | Concern | Alternative |
|---|---|---|
| track individual screen activity | privacy law | team-level workload metrics |
| auto-deny claims by postcode | discrimination risk | rules reviewed by compliance |
| requests declined this year | 2 of 140 | both with alternatives |
I would not consider it settled without evidence: Record the concern, the policy consulted, the escalation, and the alternative proposed.
Some requests must be declined, and declining well is part of the job.
Curated: · Written: · Reviewed:
QA-94The signed statement of work promises more than the data or systems can support. What do you do?(show answer)
The honest answer on contract scope versus engineering reality includes what could go wrong at the customer and how I would know.
Discovering that the contract cannot be delivered as written is common, and raising it early with evidence gives both sides time to adjust. Silently trying to deliver the impossible leads to a failed project and a dispute.
Concretely, document the gap with evidence, estimate what can be delivered, involve the account team and your leadership, and propose a revised scope or phase plan to the customer.
The reason for that specificity is a failure I have seen: A team discovered in week 3 that a promised integration was impossible, kept quiet hoping to find a way, and the gap came out at acceptance in month 5, leading to a contract dispute.
A gap analysis.
| Contract item | Reality | Proposed change |
|---|---|---|
| real-time sync with the ERP | ERP exports nightly only | nightly sync with alerting |
| 99% extraction accuracy | 18% of scans unreadable | 99% on readable scans |
I would not consider it settled without evidence: Share a written gap analysis with the account team within a week of discovering the problem.
Raise contract gaps early, with evidence.
Curated: · Written: · Reviewed:
QA-95You are responsible for three customer deployments at the same time. How do you manage them?(show answer)
With managing several deployments at once, the customer's constraints come first and the architecture follows them.
Running several deployments works when each has a clear current priority, visible status, and someone at the customer who can unblock it. Context switching is expensive, so work should be batched and blockers removed early.
Concretely, keep a single status view with each deployment's next milestone, blockers, and risks, block time by customer, clear blockers first, and tell stakeholders early when priorities force a delay.
The reason for that specificity is a failure I have seen: An FDE split every day across three customers, each deployment slipped a month, and all three sponsors felt neglected.
One week across three customers.
| Customer | Next milestone | Blocker | Days this week |
|---|---|---|---|
| A | first user, week 5 | none | 2 |
| B | data access | IT approval | 0.5 |
| C | cohort 2 rollout | none | 2.5 |
I would not consider it settled without evidence: Review a weekly status view of all deployments with milestones, blockers, and time spent on each.
Batch the work and clear blockers first.
Curated: · Written: · Reviewed:
QA-96When should you recommend stopping a customer deployment?(show answer)
I would make deciding to end a deployment measurable against the baseline the customer already trusts.
Recommend stopping when the agreed outcome is not achievable with the available data, access, or commitment, and further effort would waste both sides' resources. Stopping honestly protects the relationship better than continuing a failing project.
Concretely, compare progress with the agreed criteria, identify the blocking cause, test whether any realistic change would fix it, and present the evidence and options, including stopping, to the sponsor.
The reason for that specificity is a failure I have seen: A deployment continued for 9 months without reliable data access, consumed two FDEs full time, and ended with the customer refusing to renew and telling peers it had failed.
Progress after three months.
| Criterion | Target | Status |
|---|---|---|
| data access | week 2 | still pending |
| first real user | week 5 | not started |
| sponsor time | weekly | twice in 3 months |
I would not consider it settled without evidence: Present the sponsor with progress against the agreed criteria and the specific blocker, with options including stopping.
Stopping on evidence is a result, not a failure.
Curated: · Written: · Reviewed:
QA-97How should an FDE work with the company's product managers?(show answer)
I would start working with product managers in the customer's workflow rather than in our product catalogue.
Product managers decide what goes into the shared product, and FDEs see how it performs at customers. The relationship works when FDEs bring evidence about repeated problems and product managers explain the roadmap and trade-offs.
Concretely, meet product managers regularly, share counted field problems with examples, invite them to customer sessions, explain which workarounds are temporary, and respect the roadmap decisions they make.
The reason for that specificity is a failure I have seen: An FDE team bypassed product managers and asked engineers directly for custom features, and the product accumulated 40 one-off flags that nobody could maintain.
Field requests with product decisions.
| Field request | Customers | Product decision |
|---|---|---|
| bulk edit | 5 | next quarter |
| custom export | 1 | use the API instead |
| offline mode | 3 | under review |
I would not consider it settled without evidence: Keep a shared list of field requests with product decisions and review it monthly with product managers.
Bring product evidence, not demands.
Curated: · Written: · Reviewed:
QA-98Why do you want to be a forward deployed engineer rather than a product engineer?(show answer)
The first thing I want to know about why choose the FDE role is which decision it changes for which user.
A convincing answer connects your experience to what is distinctive about the job, such as working directly with users, owning ambiguous problems end to end, and seeing the impact of your code, while showing you understand the costs, such as travel and context switching.
Concretely, give one or two specific examples where you enjoyed working with users on a messy problem and shipped something they used, and mention the trade-offs you have considered.
The reason for that specificity is a failure I have seen: A candidate said they wanted the role because it involved less coding, and the interviewer ended the interview early because FDEs at that company wrote production code daily.
Motivation backed by evidence.
| Motivation | Example evidence |
|---|---|
| working with users | ran 12 user sessions for an internal tool |
| end-to-end ownership | built and supported a pipeline for 18 months |
| impact | cut a team's weekly reporting from 6 hours to 20 minutes |
I would not consider it settled without evidence: Check that your answer includes a specific past example and names the trade-offs of the role.
Show you want the hard parts of the job, not just the title.
Curated: · Written: · Reviewed:
QA-99You start at a customer in an industry you know nothing about. How do you get up to speed?(show answer)
For learning a new domain quickly, I would name the outcome, the owner, and the measurement before any code.
Learning a domain fast means learning its vocabulary, its key decisions, its regulations, and its data, mostly from the people who do the work. An FDE who understands the domain asks better questions and builds more useful systems.
Concretely, read the customer's own training materials, keep a glossary, shadow experts, ask them to explain their hardest recent decision, and check your understanding by explaining it back to them.
The reason for that specificity is a failure I have seen: An FDE at an insurer confused a claim reserve with a payment for 3 weeks, and a report built on that misunderstanding had to be withdrawn.
A glossary confirmed by experts, first 2 weeks.
| Term | Meaning | Confirmed by |
|---|---|---|
| reserve | estimated future cost of a claim | claims lead |
| subrogation | recovering costs from a third party | legal |
| FNOL | first notice of loss | claims lead |
I would not consider it settled without evidence: Keep a glossary of domain terms confirmed by a customer expert and review it in the first two weeks.
Learn the language before building for it.
Curated: · Written: · Reviewed:
QA-100In 45 minutes, design a deployment that helps a logistics company reduce late deliveries.(show answer)
My approach to an end-to-end deployment design case separates what the customer asked for from what they need.
A strong answer moves from the goal and baseline, to the causes of lateness, to a minimal system, to rollout and measurement. Interviewers look for clarity on data, trade-offs, and how you would know it worked.
Concretely, clarify the lateness metric and baseline, list data such as routes, GPS, and delivery windows, decompose lateness by cause, propose a minimal system such as at-risk alerts for dispatchers, and describe the rollout, monitoring, and success measure.
The reason for that specificity is a failure I have seen: A candidate designed a full route-optimization platform with streaming data and machine learning, and could not say how many deliveries were late or how the first version would be tested.
A 45-minute design on one page.
| Part | Example |
|---|---|
| baseline | 7.8% of stops late |
| largest cause | overloaded routes, 45% of late stops |
| minimal system | at-risk stop alerts for dispatchers |
| rollout | 2 depots, then all 30 |
| success | late stops under 5% |
I would not consider it settled without evidence: Check that the design names the baseline, the largest cause, the minimal system, the rollout plan, and the success metric.
Design the smallest system that moves the number.
Curated: · Written: · Reviewed:
