Top 100 Platform Engineer Interview Questions and Answers
The questions most likely to actually come up in your Platform Engineer interview, ranked by likelihood — with detailed, senior-level answers covering what an interviewer is really listening for.
Curated: · Written: · Reviewed:
QA-1A director asks why the platform team exists when every product team can already deploy. What is your answer?(show answer)
The first thing I would establish about what a platform is for is which developer decision it is meant to remove.
A platform exists to remove work that would otherwise be repeated by every team, and to make the safe path the convenient one. If each team is solving deployment once and cheaply, the platform's justification has to come from somewhere else — repeated cost, inconsistent risk controls, or knowledge that does not scale.
Concretely, name the specific repeated work with its measured cost: how many teams solve it, how long each takes, and how often. Where the number is small, the honest answer is that the capability does not need a platform yet.
The reason for that specificity is a failure I have seen: A platform team built a deployment abstraction for 6 services that each spent about 2 days a year on deployment work; the abstraction cost 9 engineer-months and the estate never grew past 11 services.
Break-even on one capability.
| Estate | Team-days/yr saved | Build cost | Payback |
|---|---|---|---|
| 6 services | 12 | 180 days | 15 yr |
| 40 services | 80 | 180 days | 2.3 yr |
| 140 services | 280 | 180 days | 8 months |
I would not consider it settled without evidence: Count the teams doing the work and the hours it costs them before committing, and state the break-even.
A platform is justified by work it removes, not by work it performs.
Curated: · Written: · Reviewed:
QA-2How does funding a platform as a project rather than a product change the technical decisions you make?(show answer)
I would start platform as product versus platform as project from the journey a developer actually walks, not from the architecture diagram.
An interface creates ongoing obligations — support, migration, incident response — that outlive the work that produced it. Project funding delivers a launch and no capacity for those, so the right technical choice under it is different from the one under product funding.
Concretely, under project funding prefer thin, replaceable wrappers you can hand back; under product funding you may own a deeper abstraction because you will maintain it. State which model is in force before choosing, and revisit the choice when funding changes.
The reason for that specificity is a failure I have seen: A project-funded team shipped a deep deployment abstraction and disbanded; two years later 34 services depended on an interface nobody owned, and the migration off it took longer than the original build.
Abstraction depth against funding model.
| Funding | Right depth | Obligation | Services stranded |
|---|---|---|---|
| project | thin wrapper | low | 0 |
| project, deep | deep | high | 34 |
| product, deep | deep | funded | 0 |
I would not consider it settled without evidence: Write the support and deprecation policy at design time and confirm someone is funded to honour it for the interface's expected life.
Depth of abstraction is a bet on how long you will be there to maintain it.
Curated: · Written: · Reviewed:
QA-3Which numbers would convince you the platform is working, and which would you refuse to report?(show answer)
This is an area where a high adoption number and a good handling of measuring platform success are not the same thing.
Adoption counts who onboarded, not whether the platform helped. The useful measures are outcomes for the teams using it — time to first deploy, change lead time, incident recovery — segmented so an improvement in the mix is not read as an improvement in the work.
Concretely, pair each outcome with a counter-metric that would move if the platform were trading one thing for another, report distributions rather than averages, and refuse to publish anything that ranks individual developers, which changes behaviour without improving it.
The reason for that specificity is a failure I have seen: A platform reported 94 percent adoption while median time to first deploy for a new service had risen from 3 to 6 days, because onboarding was mandatory and the number counted onboarding.
Two views of the same quarter.
| Measure | Q1 | Q2 | Reads as |
|---|---|---|---|
| adoption | 61% | 94% | success |
| median first deploy | 3 d | 6 d | regression |
| p90 first deploy | 9 d | 21 d | regression |
I would not consider it settled without evidence: Report time to first deploy and change lead time by cohort, with the tail, alongside any adoption figure.
Adoption is an input, and the platform is judged on outputs.
Curated: · Written: · Reviewed:
QA-4Two teams report very different lead times on the same pipeline. Before investigating the pipeline, what would you check?(show answer)
My answer to defining lead time for changes begins with who carries the work afterwards, because an abstraction is an obligation.
Lead time can be measured from first commit, from pull-request open, or from merge, and the three differ by hours or days on identical infrastructure. Most disagreement about delivery metrics is disagreement about definitions rather than about performance.
Concretely, write the definition, the event source, and the population next to every figure, and keep them stable. A team that changes where the clock starts has not regressed, and a chart that does not say which start was used cannot distinguish the two.
The reason for that specificity is a failure I have seen: A quarterly review compared one team measuring from merge at 38 minutes against another measuring from first commit at 31 hours, and concluded the second team had a pipeline problem; the pipelines were identical and the difference was review latency.
Same pipeline, three definitions.
| Clock starts at | Team A | Team B |
|---|---|---|
| first commit | 29 h | 31 h |
| PR open | 26 h | 27 h |
| merge | 38 min | 40 min |
I would not consider it settled without evidence: Publish the metric definition with its event source, and recompute both teams on the same definition before comparing.
A delivery metric is a definition before it is a measurement.
Curated: · Written: · Reviewed:
QA-5A team's change failure rate has fallen from 14 percent to 3 percent. Is that good news?(show answer)
I would treat change failure rate and its counter-metric as a product decision with users who can route around me.
Change failure rate and deployment frequency move against each other under pressure. A rate that falls while deployment frequency also falls usually means the team is batching changes and deploying less, which is the opposite of the intended improvement.
Concretely, read the two together and never separately, define failure explicitly so that planned rollbacks are not silently counted, and check batch size, because a falling rate with growing changes per deployment is a larger blast radius rather than a safer one.
The reason for that specificity is a failure I have seen: A team improved change failure rate from 14 to 3 percent by moving from 40 deploys a month to 6; changes per deployment rose from 2 to 14 and the mean incident duration doubled because each failure was harder to attribute.
Where the improvement came from.
| Quarter | Deploys/mo | Changes/deploy | Failure rate | Mean incident |
|---|---|---|---|---|
| Q1 | 40 | 2 | 14% | 41 min |
| Q2 | 6 | 14 | 3% | 88 min |
I would not consider it settled without evidence: Report deployment frequency, changes per deployment, and failure rate on the same chart, so an improvement in one that came from the others is visible.
Any single delivery metric can be improved by making another one worse.
Curated: · Written: · Reviewed:
QA-6The company-wide deployment frequency improved while no team got faster. How?(show answer)
The useful question for organisation-level metrics and Simpson's paradox is what happens on the day it breaks and the developer has to look underneath.
An organisation-wide figure is a weighted mix. When a fast-deploying service grows its share of total deployments, the aggregate improves even if every individual service slowed down, because the change is in the mix rather than in the work.
Concretely, aggregate to the service or team level and report the distribution across them, use the aggregate only for capacity questions, and check the mix whenever an organisation-level number moves without a matching movement in any segment.
The reason for that specificity is a failure I have seen: A company figure rose from 276 to 528 deploys a week while the per-service rate in both cohorts fell, because the high-frequency cohort grew from 2 services to 6 while the slower cohort shrank from 8 to 4; the platform took credit for an improvement that had not happened anywhere.
Aggregate up, every cohort down.
| Cohort | Q1 rate/service | Q2 rate/service | Q1 services | Q2 services |
|---|---|---|---|---|
| high-frequency | 90/wk | 82/wk | 2 | 6 |
| rest | 12/wk | 9/wk | 8 | 4 |
| total deploys | 276/wk | 528/wk | 10 | 10 |
I would not consider it settled without evidence: Show per-service deployment frequency alongside the aggregate and require a segment to have moved before attributing a change.
An aggregate can improve while every part of it gets worse.
Curated: · Written: · Reviewed:
QA-7Your quarterly developer satisfaction score fell 8 points. What do you check before acting?(show answer)
I would settle developer surveys as instruments by watching a team use it rather than by asking whether they liked it.
A survey score is an instrument reading, and it moves when the instrument changes as readily as when the experience does. Question wording, response rate, and which population answered all shift the number independently of anything the platform did.
Concretely, keep item wording stable across waves, report response rate and non-response alongside the score, ask about specific recent tasks rather than general feeling, and protect identifiability in small teams, where a segment of four people is not anonymous whatever the label says.
The reason for that specificity is a failure I have seen: An 8-point fall came from a response rate dropping from 71 to 34 percent after the survey moved out of the onboarding flow; the remaining respondents were disproportionately the people with an active complaint.
What moved between waves.
| Wave | Response rate | Score | Wording changed |
|---|---|---|---|
| Q1 | 71% | 74 | no |
| Q2 | 34% | 66 | no |
| Q2 reweighted | — | 73 | no |
I would not consider it settled without evidence: Publish response rate and respondent composition with every wave, and pair each finding with a behavioural measure where one exists.
A score that moved may be telling you about the survey.
Curated: · Written: · Reviewed:
QA-8Adoption of the golden path is at 40 percent and leadership wants to mandate it. What would you say?(show answer)
The judgement in golden path adoption without mandate is what stays visible and overridable, not how much gets hidden.
Voluntary adoption is the feedback loop that tells you whether the path is good. A mandate replaces that signal with a compliance number and hides the friction, which resurfaces later as shadow tooling and abandonment at the first incident.
Concretely, mandate only the risk controls a security or regulatory review genuinely requires, let the ergonomics compete on merit, and treat a departure as a specification for the next version. Where adoption is low and controls are not the reason, the finding is about the path.
The reason for that specificity is a failure I have seen: A mandate took recorded adoption from 40 to 100 percent in a quarter; a later audit found 23 of 61 services had copied the template once and diverged, so they received no updates and were counted as adopters.
Adoption before and after the mandate.
| Measure | Before | After |
|---|---|---|
| recorded adoption | 40% | 100% |
| tracking current version | 38% | 62% |
| diverged copies | 2 | 23 |
I would not consider it settled without evidence: Measure the fraction of adopters still tracking the current template version, not the fraction that ever onboarded.
A mandate converts a measurement into a number.
Curated: · Written: · Reviewed:
QA-9A team needs something the golden path does not support. What do you offer them?(show answer)
Where platform teams lose the interview on escape hatches is mandating what has not yet earned adoption.
A path with no supported way off it forces any team with a genuinely different requirement to abandon it entirely, taking the security and operational defaults with them. A supported exit from one component keeps the rest.
Concretely, make the hatch explicit and visible, record who takes it and why, and treat a cluster of teams leaving at the same step as the specification for the next version rather than as a governance failure.
The reason for that specificity is a failure I have seen: A path with no hatch lost a team that needed a different runtime; they rebuilt the whole pipeline themselves and in doing so dropped the artifact signing and the audit logging that the path had provided for free.
What a team keeps when it leaves.
| Route | Controls retained | Teams |
|---|---|---|
| on path | 9 of 9 | 44 |
| component hatch | 8 of 9 | 6 |
| full rebuild | 3 of 9 | 1 |
I would not consider it settled without evidence: Track exits by step, and confirm that teams taking a hatch retain the mandatory controls rather than leaving the path entirely.
Without a supported exit, the only exit is the whole platform.
Curated: · Written: · Reviewed:
QA-10How do you ship a breaking change to a template that 60 services already use?(show answer)
I would answer versioning a golden path by separating the risk control that is genuinely required from the preference that is not.
A golden path is versioned software with users, so the ordinary migration obligations apply. A path whose adopters are spread across five versions is five paths under one name, each with its own behaviour and its own security posture.
Concretely, publish a support window, provide an automated upgrade for the mechanical part of the change, measure the share of adopters on each version, and never deprecate without an upgrade route.
The reason for that specificity is a failure I have seen: A breaking template change was announced with a 30-day window and no codemod; after 90 days 41 of 60 services were still on the old version, and the platform team was maintaining both while telling teams only one was supported.
Migration with and without a codemod.
| Day | Manual upgrade | With codemod |
|---|---|---|
| 30 | 12 of 60 | 47 of 60 |
| 60 | 17 of 60 | 58 of 60 |
| 90 | 19 of 60 | 60 of 60 |
I would not consider it settled without evidence: Report version distribution across the estate weekly, and treat a long tail as unfinished work rather than as a team compliance problem.
A version nobody can leave is a version you still support.
Curated: · Written: · Reviewed:
QA-11When should the platform hide a system entirely, and when should it just set a good default?(show answer)
The engineering content of abstraction versus default is the migration path and the support policy, not the templating language.
A default is visible, overridable, and teachable, and it degrades gracefully during an incident. A full abstraction removes the concept, which is worth the cost only when the platform is also prepared to own every failure mode expressed in the underlying system's terms.
Concretely, hide what you will diagnose on the developer's behalf, default what they may need to reason about, and where you do abstract, invest in error messages in the developer's vocabulary and a documented path from the platform symptom to the underlying cause.
The reason for that specificity is a failure I have seen: A deployment abstraction hid the orchestrator until a scheduling failure surfaced as "deployment did not become ready"; the on-call developer needed both the orchestrator knowledge the platform had promised to remove and the platform's mapping onto it.
Diagnosis time by surface design.
| Surface | Median diagnosis | Escalations to platform |
|---|---|---|
| raw orchestrator | 22 min | 4% |
| abstraction, generic errors | 74 min | 61% |
| abstraction, mapped errors | 19 min | 9% |
I would not consider it settled without evidence: Take an incident in the underlying system and check that a developer can diagnose it from the platform's surface alone, without reading the platform's source.
Hiding a concept means owning its failures.
Curated: · Written: · Reviewed:
QA-12How would you check that the platform actually reduced cognitive load rather than moving it?(show answer)
Before rolling cognitive load as the platform hypothesis out I would write down what would tell me it had made things worse.
An abstraction that removes ten decisions from every product team and adds a support queue to the platform team has relocated work, which can still be right because the platform pays once. It is wrong only when nobody counted.
Concretely, observe developers completing specific tasks — bootstrap, deploy, debug, recover — recording tools touched, context switches, and help requested, and track platform support volume and its causes on the other side of the ledger.
The reason for that specificity is a failure I have seen: A platform cut product-team setup steps from 18 to 4 and grew the platform support queue from 30 to 210 tickets a month with unchanged headcount; response time went to 3 days and teams began routing around it.
Both sides of the ledger.
| Measure | Before | After |
|---|---|---|
| product setup steps | 18 | 4 |
| platform tickets/mo | 30 | 210 |
| platform headcount | 4 | 4 |
| median ticket response | 4 h | 3 d |
I would not consider it settled without evidence: Report platform support volume and its top causes alongside every developer-facing improvement.
Load that moved is not load that went away.
Curated: · Written: · Reviewed:
QA-13You are about to let any developer provision a database. What has to be true first?(show answer)
The first thing I would establish about self-service and blast radius is which developer decision it is meant to remove.
Self-service moves an action from a small group who understand its consequences to a large group who reasonably do not. The safety has to come from the interface rather than from the caller's knowledge, because that knowledge is exactly what self-service removes.
Concretely, bound what the interface can do — size, region, retention, cost, and deletion protection — make destructive operations require a distinct and reversible path, and default to the safe option rather than the flexible one.
The reason for that specificity is a failure I have seen: A self-service database interface allowed a size parameter with no ceiling; a copy-pasted value provisioned 8 instances at 64 vCPU for a staging workload, at about 14,000 dollars before anyone noticed the following month.
Worst plausible call, before and after bounds.
| Parameter | Unbounded | Bounded |
|---|---|---|
| instance size | 64 vCPU | 8 vCPU |
| count | unlimited | 2 |
| monthly worst case | $14,000 | $430 |
I would not consider it settled without evidence: Enumerate what the worst plausible call to the interface can cost or destroy, and cap it before opening access.
Self-service moves the risk into the interface, so bound it there.
Curated: · Written: · Reviewed:
QA-14Should a self-service platform let a developer delete a production database?(show answer)
I would start deletion and reversibility in self-service from the journey a developer actually walks, not from the architecture diagram.
Reversibility, not permission, is what makes a destructive operation safe to expose. An action that cannot be undone needs friction proportional to the loss, and that friction belongs in the interface rather than in a policy document.
Concretely, soft-delete with a retention window, require a distinct confirmation naming the resource, protect anything tagged production behind a second approver, and test the restore path on a schedule rather than assuming it works.
The reason for that specificity is a failure I have seen: A delete call removed a production database and its automated backups together, because both lived under the same lifecycle; the restore path had never been exercised and recovery took 11 hours from an offsite copy.
Recovery by deletion design.
| Design | Recovery path | Tested | Elapsed |
|---|---|---|---|
| hard delete | offsite copy | no | 11 h |
| soft delete, 7 d | in place | no | 40 min |
| soft delete, tested | in place | monthly | 6 min |
I would not consider it settled without evidence: Restore from the retained copy on a schedule and record the elapsed time, rather than confirming that backups exist.
The safety of a destructive action is the restore you have actually run.
Curated: · Written: · Reviewed:
QA-15What is your policy when a platform API needs a breaking change?(show answer)
This is an area where a high adoption number and a good handling of platform APIs and backward compatibility are not the same thing.
A platform API's consumers are colleagues who cannot be forced to upgrade on your schedule and who will route around you if the cost is high enough. Compatibility is therefore a product decision rather than a courtesy.
Concretely, version the interface, run old and new in parallel for a published window, instrument per-version usage so the tail is visible, and provide the migration rather than describing it. Remove a version when usage reaches zero, not when the window expires.
The reason for that specificity is a failure I have seen: A field was made required in place with a two-week notice; 19 automated pipelines that nobody had inventoried broke on the same morning, and 6 of them belonged to teams that had already left the company's Slack channel for platform announcements.
Callers remaining on the old version.
| Week | Announced only | With instrumented usage |
|---|---|---|
| 2 | unknown | 31 |
| 6 | unknown | 8 |
| 10 | unknown | 0 |
I would not consider it settled without evidence: Instrument per-version, per-caller usage and drive removal from observed usage rather than from the announcement date.
You may publish a deadline; only usage tells you when it is safe.
Curated: · Written: · Reviewed:
QA-16Why is an operator's reconcile loop written to be idempotent, and what breaks when it is not?(show answer)
My answer to Kubernetes operators and the reconciliation model begins with who carries the work afterwards, because an abstraction is an obligation.
Reconciliation is level-triggered: the loop is handed desired state and observes actual state, and it may run at any time, repeatedly, on the same input. It is not a sequence of edge-triggered events, so any step that assumes it runs once will eventually run twice.
Concretely, compute the action from the difference between desired and observed state rather than from the event that woke you, make every write safe to repeat, and never carry progress in memory between invocations since the process can restart mid-loop.
The reason for that specificity is a failure I have seen: An operator created a cloud resource on each reconcile without checking whether one existed; a controller restart during a rollout produced 7 duplicate load balancers over 40 minutes, none of which were tracked in status.
Resources created over 40 minutes.
| Design | Reconciles | Resources created |
|---|---|---|
| create on event | 7 | 7 |
| create if absent | 7 | 1 |
| create if absent + status | 7 | 1 |
I would not consider it settled without evidence: Run the reconcile loop repeatedly against unchanged state and assert that no write occurs after the first convergence.
The loop will run again, so write it as though it always has.
Curated: · Written: · Reviewed:
QA-17A team wants a CRD with 40 fields mirroring the underlying API. What would you push back on?(show answer)
I would treat custom resource design as a product decision with users who can route around me.
A custom resource is a user interface, and mirroring the underlying API gives the developer every decision the platform was supposed to remove while adding a translation layer they now also have to learn.
Concretely, express the resource in the vocabulary of the developer's intent, default what the platform should decide, and keep an explicit escape for the rare case rather than exposing every knob. Treat schema additions as API changes with the same compatibility rules.
The reason for that specificity is a failure I have seen: A 40-field CRD produced a median 96-line manifest per service; teams copied one from a neighbour without understanding it, and 14 of 22 services ended up in the same misconfigured resource-limit state.
Manifest size and copied misconfiguration.
| Design | Median lines | Fields set | Services misconfigured |
|---|---|---|---|
| 40-field mirror | 96 | 31 | 14 of 22 |
| intent-shaped | 11 | 5 | 1 of 22 |
I would not consider it settled without evidence: Measure the median manifest a real team writes and how many fields they change from the default; a high count means the abstraction has not been made.
A resource that mirrors the API has not abstracted anything.
Curated: · Written: · Reviewed:
QA-18A developer asks why their resource is not ready and the operator logs say nothing useful. What was missing?(show answer)
The useful question for operator status and observability is what happens on the day it breaks and the developer has to look underneath.
A controller's status is the interface through which its users understand it. Progress and failure that live only in operator logs are invisible to the developer, who does not have access to them and should not need it.
Concretely, write conditions that name what is being waited on and why, surface the underlying error in the developer's terms rather than passing through a raw API message, and include the observed generation so a stale status is distinguishable from a current one.
The reason for that specificity is a failure I have seen: A resource sat not-ready for 3 hours with an empty status while the operator retried a quota-denied cloud call; the developer opened a ticket, and the platform on-call found the cause in 4 minutes once they read the operator's own logs.
Time to cause by status design.
| Status content | Developer finds cause | Tickets/mo |
|---|---|---|
| empty | no | 34 |
| raw API error | sometimes | 12 |
| mapped condition | yes | 2 |
I would not consider it settled without evidence: Take each failure path in the controller and confirm a developer can name the cause from status alone.
If the answer is only in your logs, you have not shipped it.
Curated: · Written: · Reviewed:
QA-19A standard is documented and half the estate violates it. What do you do?(show answer)
I would settle admission control versus documentation by watching a team use it rather than by asking whether they liked it.
A rule that is enforced only by documentation is a rule that holds wherever someone read it recently. Where the consequence justifies it, the enforcement belongs in the path that creates the resource.
Concretely, move the rule into admission control or the template, ship it in warn mode first to size the violation set, fix or exempt what it finds, and only then enforce. An exemption should be explicit and expiring rather than a silent gap.
The reason for that specificity is a failure I have seen: A documented resource-limit standard was violated by 61 of 140 workloads; enforcing it without a warn phase would have blocked deployments for 24 teams simultaneously on a Monday morning.
Rollout by phase.
| Phase | Violations | Deploys blocked |
|---|---|---|
| documented only | 61 | 0 |
| warn mode | 61 | 0 |
| after remediation | 3 | 0 |
| enforce | 3 | 3, expected |
I would not consider it settled without evidence: Run the rule in warn mode across the whole estate and publish the violation count before enabling enforcement.
Enforce in the path that creates the thing, after you know what it will break.
Curated: · Written: · Reviewed:
QA-20The cluster and the repository disagree. Which is right and what do you do about it?(show answer)
The judgement in GitOps and drift is what stays visible and overridable, not how much gets hidden.
Under GitOps the repository is the declared intent and the cluster is the observation. A difference is information — someone changed something out of band, or reconciliation is failing — and neither silently overwriting nor ignoring it is a policy.
Concretely, detect and report drift explicitly, decide per resource class whether reconciliation reverts or alerts, and keep an audit trail of out-of-band changes so an emergency fix is recorded rather than erased.
The reason for that specificity is a failure I have seen: Automatic reconciliation reverted an emergency manual scale-up 90 seconds after an on-call engineer applied it during an incident; the service returned to the failing capacity and the outage extended by 25 minutes.
Response to an emergency out-of-band change.
| Policy | Incident outcome | Change recorded |
|---|---|---|
| auto-revert | +25 min outage | no |
| alert only | resolved | yes |
| revert with hold window | resolved | yes |
I would not consider it settled without evidence: Test the incident case explicitly: apply an out-of-band change and confirm the system's response is the one you intended.
Drift is a signal, and reverting it silently deletes the signal.
Curated: · Written: · Reviewed:
QA-21How should a change move from staging to production, and what makes an environment promotion meaningless?(show answer)
Where platform teams lose the interview on promotion between environments is mandating what has not yet earned adoption.
Promotion is only evidence if the artifact is the same and the environments differ in ways you have enumerated. Rebuilding per environment tests a different artifact, and an environment that differs in unknown ways tests nothing you can rely on.
Concretely, build once and promote the same immutable artifact by digest, keep configuration external and diffable, and maintain an explicit inventory of the intended differences between environments so an unintended one is visible.
The reason for that specificity is a failure I have seen: A pipeline rebuilt per environment; a transitive dependency published a new patch between staging and production builds, and the artifact that passed staging was not the one that failed in production.
What promotion proves.
| Pipeline | Same digest | Staging evidence applies | Escaped defects/quarter |
|---|---|---|---|
| rebuild per env | no | no | 7 |
| build once, promote | yes | yes | 1 |
| build once, config drift | yes | partly | 3 |
I would not consider it settled without evidence: Assert that the digest deployed to production is the digest that passed staging, and fail the promotion when it is not.
Promotion means the same artifact, or it means nothing.
Curated: · Written: · Reviewed:
QA-22You run a canary at 5 percent for 10 minutes. What makes that useful rather than ceremonial?(show answer)
I would answer progressive delivery and the abort condition by separating the risk control that is genuinely required from the preference that is not.
A canary is only useful if it can fail. That requires enough traffic to detect the effect you care about within the window, and a predeclared abort condition that something automated will act on.
Concretely, compute the traffic needed for the smallest effect worth catching, set the window from that rather than from convenience, define the abort metric and threshold before the rollout, and wire the rollback to fire without a human decision.
The reason for that specificity is a failure I have seen: A 5 percent canary over 10 minutes carried 340 requests against a baseline error rate of 0.4 percent; the smallest error-rate change it could distinguish was about 3 percentage points, so a regression that doubled errors passed every time.
Detectable effect against canary size.
| Share | Window | Requests | Smallest detectable |
|---|---|---|---|
| 5% | 10 min | 340 | ~3.0 pp |
| 5% | 60 min | 2040 | ~1.2 pp |
| 25% | 60 min | 10200 | ~0.5 pp |
I would not consider it settled without evidence: State the detectable effect size for the chosen traffic and window, and lengthen the window when it is larger than the regression you care about.
A canary that cannot fail is a delay, not a test.
Curated: · Written: · Reviewed:
QA-23How do you know that the container running in production came from the commit you think it did?(show answer)
The engineering content of build reproducibility and supply chain is the migration path and the support policy, not the templating language.
Without a signed link from source to artifact, the connection between a commit and a running image is an assumption maintained by convention. Tags are mutable and a rebuild from the same tag is not the same image.
Concretely, pin by digest rather than tag, generate provenance attesting to the source commit and build environment, sign it in the build system, and verify the signature at admission so an unattested image cannot run.
The reason for that specificity is a failure I have seen: A rollback deployed a moving tag that had been overwritten by a later build; the running image differed from the commit the team believed they had rolled back to, and the incident review reasoned from the wrong source for 2 hours.
What each reference guarantees.
| Reference | Immutable | Links to commit |
|---|---|---|
| :latest | no | no |
| :v1.4.2 | no | no |
| @sha256:... | yes | with provenance |
I would not consider it settled without evidence: Verify at admission that the running digest has provenance naming the expected source commit, and reject when it does not.
A tag is a label someone can move; a digest is the artifact.
Curated: · Written: · Reviewed:
QA-24Should teams share a cluster or get their own?(show answer)
Before rolling multi-tenancy isolation choices out I would write down what would tell me it had made things worse.
The question is what isolation the workloads require — security boundary, noisy-neighbour resistance, blast radius, and independent upgrade — and each answer has a different cost. A shared cluster is cheaper and couples upgrade cycles; a cluster per team is the reverse.
Concretely, decide per requirement rather than as one choice: namespace isolation with quotas and network policy for most, node pools for noisy or licensed workloads, separate clusters where a genuine security boundary or an independent upgrade cadence is required.
The reason for that specificity is a failure I have seen: Forty teams shared one cluster; a control-plane upgrade needed a coordinated window across all of them, took 3 attempts across 6 weeks, and one team's admission webhook blocked the rollout for everyone.
What each boundary gives you.
| Boundary | Security | Noisy neighbour | Independent upgrade | Cost |
|---|---|---|---|---|
| namespace | partial | with quotas | no | low |
| node pool | partial | yes | no | medium |
| cluster | yes | yes | yes | high |
I would not consider it settled without evidence: Write down which isolation property each tenant actually needs and check the chosen model provides it, rather than choosing on cost alone.
Isolation is several properties, and one boundary rarely provides them all.
Curated: · Written: · Reviewed:
QA-25A team's service is slow and their CPU usage graph looks fine. What do you check on a shared cluster?(show answer)
The first thing I would establish about resource requests, limits, and the noisy neighbour is which developer decision it is meant to remove.
Requests determine scheduling and the share a workload is guaranteed; limits determine throttling. A container under its limit can still be throttled against its request share, and average CPU usage will not show it.
Concretely, look at throttling counters rather than utilisation, set requests from observed usage percentiles rather than from a guess, and treat a large gap between request and limit as a scheduling risk since the scheduler placed the pod on the request.
The reason for that specificity is a failure I have seen: A service requested 100 millicores and limited at 2 cores; it ran at 900 millicores under load on a full node and was throttled 38 percent of periods while its average utilisation graph read 45 percent.
Same service, three views.
| Metric | Reading | Suggests |
|---|---|---|
| mean CPU | 45% | healthy |
| p95 CPU | 900m | near limit |
| throttled periods | 38% | starved |
I would not consider it settled without evidence: Report throttled period share per container alongside utilisation, and alert on throttling rather than on CPU.
Utilisation says how much it used; throttling says how much it was refused.
Curated: · Written: · Reviewed:
QA-26A service autoscales on CPU and still queues requests under load. What signal would you use instead?(show answer)
I would start autoscaling signals from the journey a developer actually walks, not from the architecture diagram.
CPU is a proxy for saturation that holds only when the work is CPU-bound. A service that spends its time waiting on a downstream call saturates on concurrency long before CPU moves, so scaling on CPU responds after the queue has already formed.
Concretely, scale on the resource that actually saturates — in-flight requests, queue depth, or a latency objective — keep CPU as a secondary signal, and set the scale-up response faster than the scale-down so a burst does not oscillate the replica count.
The reason for that specificity is a failure I have seen: An IO-bound service held 40 percent CPU while p99 latency rose from 120 ms to 4.2 s; the CPU-based scaler added no replicas for 9 minutes because its threshold was never crossed.
Which signal moves first under load.
| Load | CPU | In-flight | p99 |
|---|---|---|---|
| 50% | 32% | 40 | 118 ms |
| 90% | 38% | 190 | 940 ms |
| 110% | 40% | 610 | 4.2 s |
I would not consider it settled without evidence: Load-test to saturation and identify which signal moves first, then scale on that one.
Scale on what runs out, not on what is easy to measure.
Curated: · Written: · Reviewed:
QA-27How does a secret get from the vault to a running container, and where do most designs leak?(show answer)
This is an area where a high adoption number and a good handling of secrets distribution are not the same thing.
The hard part is not storage but the bootstrap: the workload needs some credential to fetch its secrets, and that credential is the real root of trust. A long-lived token in an environment variable moves the problem rather than solving it.
Concretely, derive workload identity from the platform — a projected service-account token or instance identity exchanged for a short-lived credential — inject at runtime rather than baking into images, keep secrets out of environment variables where the runtime dumps them into crash reports, and rotate on a schedule you have tested.
The reason for that specificity is a failure I have seen: A static vault token was mounted as an environment variable with a one-year expiry; it appeared in a crash dump uploaded to a third-party error tracker, and rotating it required a coordinated restart of 74 services.
Blast radius of the bootstrap credential.
| Bootstrap | Lifetime | Scope | Services to rotate |
|---|---|---|---|
| static token | 1 year | all secrets | 74 |
| workload identity | 10 min | one service | 0 |
I would not consider it settled without evidence: Trace one secret end to end and name the credential at each hop, then check the lifetime and blast radius of the longest-lived one.
Every secrets system has a bootstrap credential, and that is the one that matters.
Curated: · Written: · Reviewed:
QA-28Your secrets rotate every 90 days automatically. What still worries you?(show answer)
My answer to secret rotation that has been tested begins with who carries the work afterwards, because an abstraction is an obligation.
Rotation is only safe if every consumer picks up the new value without a restart, or if the restart is part of the tested procedure. Automation that rotates the stored value while consumers hold the old one in memory converts a routine event into an outage.
Concretely, support two valid credentials during an overlap window, make consumers refresh on a timer or on authentication failure, and exercise the rotation in a lower environment on the real schedule rather than assuming it works.
The reason for that specificity is a failure I have seen: An automatic rotation succeeded and 3 of 12 consumers cached the credential at process start; they failed authentication 40 minutes later when their connection pools recycled, well after the rotation had been marked successful.
Consumer behaviour at rotation.
| Consumer refresh | Overlap window | Failures |
|---|---|---|
| cached at start | none | 3 of 12 |
| cached at start | 24 h | 3 of 12, delayed |
| refresh on 401 | 24 h | 0 |
I would not consider it settled without evidence: Rotate in a lower environment and confirm every consumer recovers without a restart, per consumer rather than in aggregate.
Rotation you have not exercised is a scheduled outage.
Curated: · Written: · Reviewed:
QA-29Why prefer workload identity to a shared service account with a long-lived key?(show answer)
I would treat workload identity over shared credentials as a product decision with users who can route around me.
A shared long-lived key gives you no attribution and no containment: every workload holding it is indistinguishable in the audit log, and a leak from any one of them compromises all. Workload identity binds the credential to the running workload and makes both attribution and revocation per-workload.
Concretely, federate the platform's own workload identity to the cloud provider so each workload receives short-lived credentials scoped to what it needs, and make the shared key path unavailable rather than merely discouraged.
The reason for that specificity is a failure I have seen: A shared key appeared in a public repository; the audit log showed 400,000 calls with no way to tell which of 60 workloads made any of them, so the response was to rotate everything and restart the entire estate.
Incident response by credential model.
| Model | Attribution | Revocation scope | Restart |
|---|---|---|---|
| shared key | none | everything | 60 services |
| workload identity | per workload | one workload | 1 service |
I would not consider it settled without evidence: Take a single audit-log entry and confirm it identifies the specific workload, not a shared account.
A credential shared by sixty workloads identifies none of them.
Curated: · Written: · Reviewed:
QA-30The platform issues every service the same broad role because narrowing it is slow. What would you change?(show answer)
The useful question for least privilege in platform-issued roles is what happens on the day it breaks and the developer has to look underneath.
A permission set granted because it is convenient is a permission set nobody can reason about later. The cost of narrowing it rises with time, because by then something depends on each permission for reasons nobody recorded.
Concretely, generate the role from the declared needs in the service's own manifest, start from deny and add on evidence, and use access logs to find and remove permissions that have never been exercised, with a grace period rather than an immediate cut.
The reason for that specificity is a failure I have seen: A shared role accumulated 40 permissions over 2 years; an access-log analysis showed that only 18 of the 40 had ever been exercised, leaving 22 unused, and one of the unused ones allowed deleting production storage buckets.
Granted against exercised over 90 days.
| Permission class | Granted | Used | Destructive |
|---|---|---|---|
| read | 14 | 12 | no |
| write | 18 | 5 | some |
| delete | 8 | 1 | yes |
I would not consider it settled without evidence: Report permissions granted against permissions exercised over a representative window, and drive removal from that gap.
A permission nobody has used is a permission nobody will miss and an attacker might.
Curated: · Written: · Reviewed:
QA-31Should the platform default to open or closed pod-to-pod networking?(show answer)
I would settle network policy defaults by watching a team use it rather than by asking whether they liked it.
A default-open network means every service is reachable from every other, so a single compromised workload has lateral access to the whole estate. Default-closed is safe but breaks anything undeclared, so the migration matters as much as the policy.
Concretely, run in observation mode first to learn the actual traffic graph, generate initial policies from what is observed rather than from architecture diagrams, then flip to default-deny per namespace with the observed policies in place.
The reason for that specificity is a failure I have seen: A default-deny policy was applied from a diagram; 12 undocumented dependencies broke, including a metrics sidecar that every team's dashboards relied on, and the change was reverted within an hour.
Documented against observed edges.
| Source | Documented edges | Observed edges | Undocumented |
|---|---|---|---|
| diagram | 34 | — | — |
| observation | — | 46 | 12 |
I would not consider it settled without evidence: Compare the observed traffic graph against the documented one before enforcing, and enumerate every edge the policy would cut.
Default-deny is right and the traffic graph is not the diagram.
Curated: · Written: · Reviewed:
QA-32Every service emits metrics and developers still cannot answer why their request was slow. What is missing?(show answer)
The judgement in platform observability the developer can use is what stays visible and overridable, not how much gets hidden.
Aggregated service metrics answer questions about a service; they cannot follow one request across the services that handled it. Without correlated traces, a cross-service latency question requires joining several dashboards by eye and by timestamp.
Concretely, propagate a trace context through every platform-provided hop — ingress, service mesh, message queue, job runner — so the trace is complete by default rather than only where a team instrumented it, and sample in a way that keeps the slow requests.
The reason for that specificity is a failure I have seen: Traces broke at the message queue because the platform's own consumer did not propagate context; a latency investigation crossing that queue took 2 days and concluded with the wrong service.
Trace completeness by hop.
| Hop | Context propagated | Investigation time |
|---|---|---|
| ingress to service | yes | minutes |
| service to service | yes | minutes |
| across queue | no | 2 days |
I would not consider it settled without evidence: Send a request through every platform-provided hop and confirm the trace is unbroken end to end.
A trace that stops at your component stops being a trace.
Curated: · Written: · Reviewed:
QA-33Log spend has tripled in a year. How do you reduce it without losing the ability to debug?(show answer)
Where platform teams lose the interview on log cost and retention policy is mandating what has not yet earned adoption.
Log value falls sharply with age while cost accrues linearly, and most volume comes from a small number of high-rate, low-value lines. Cutting uniformly loses the rare lines that matter; cutting by value keeps them.
Concretely, attribute cost by service and by log line so the top contributors are visible, drop or sample the high-rate low-value lines at the source, keep short hot retention with longer cheap archival, and give teams the bill for their own volume.
The reason for that specificity is a failure I have seen: A uniform 30-day cut saved 40 percent and removed the only record of a quarterly batch failure, whose investigation window was 45 days; the next occurrence had no history to compare against.
Volume against investigative value.
| Line class | Share of volume | Used in incidents |
|---|---|---|
| per-request access | 71% | 4% |
| health check | 12% | 0% |
| error and warn | 3% | 88% |
I would not consider it settled without evidence: Rank log lines by volume and by how often they appear in incident investigations, and cut from the bottom of that ranking.
Cut by value per byte, not by age alone.
Curated: · Written: · Reviewed:
QA-34Cloud spend is rising and no team believes it is theirs. How do you fix the conversation?(show answer)
I would answer cost attribution to teams by separating the risk control that is genuinely required from the preference that is not.
Spend that is not attributable is spend nobody can act on. Attribution has to be enforced at resource creation, because retroactive tagging never reaches the resources that matter and shared costs stay unallocated by default.
Concretely, require ownership tags at creation through the platform's own provisioning path, allocate shared costs by a stated and defensible rule rather than leaving them unallocated, and show each team its own trend rather than the total.
The reason for that specificity is a failure I have seen: Forty-one percent of spend sat in an unallocated bucket because tagging was advisory; the largest single line in it was an abandoned test cluster running for 7 months at about 2,400 dollars a month.
Where the spend sits.
| Bucket | Share | Actionable |
|---|---|---|
| attributed to teams | 59% | yes |
| shared, allocated | 0% | no |
| unallocated | 41% | no |
I would not consider it settled without evidence: Report the unallocated share weekly and treat it as a platform defect rather than a finance one.
Unattributed cost is cost with no owner and no downward pressure.
Curated: · Written: · Reviewed:
QA-35The platform offers on-demand preview environments and cost has quadrupled. What went wrong?(show answer)
The engineering content of environment sprawl is the migration path and the support policy, not the templating language.
An environment that is easy to create and has no lifecycle accumulates. The cost of the feature is not the environment, it is the absence of an expiry, because nobody deletes something that is not visibly theirs.
Concretely, give every ephemeral environment a default expiry with an explicit extension, tie its life to the pull request that created it, and report per-team environment count and age so sprawl is visible before the invoice.
The reason for that specificity is a failure I have seen: Preview environments had no expiry; 340 were running against 41 open pull requests, and the oldest had been alive for 5 months after its branch was deleted.
Environments against open pull requests.
| Policy | Live environments | Open PRs | Ratio |
|---|---|---|---|
| no expiry | 340 | 41 | 8.3 |
| 72 h expiry | 52 | 41 | 1.3 |
I would not consider it settled without evidence: Report live environment count against open pull requests, and alarm on the ratio rather than on cost.
Anything easy to create needs a default expiry.
Curated: · Written: · Reviewed:
QA-36A cluster autoscaler is enabled and pods still sit pending for minutes. Why?(show answer)
Before rolling capacity headroom and cluster autoscaling out I would write down what would tell me it had made things worse.
Cluster autoscaling adds nodes after a pod is already unschedulable, and node provisioning takes minutes. The delay is structural, so a workload that cannot wait needs headroom rather than faster scaling.
Concretely, keep enough spare capacity for the largest expected burst, use low-priority placeholder workloads that are evicted to make room instantly, and size headroom from the observed burst distribution rather than from a round number.
The reason for that specificity is a failure I have seen: A traffic burst needed 14 pods; the autoscaler provisioned nodes in 4.5 minutes and the queue backed up for the whole interval, dropping about 12,000 requests before capacity arrived.
Burst response by strategy.
| Strategy | Capacity available in | Requests dropped |
|---|---|---|
| autoscale only | 4.5 min | 12000 |
| placeholder pods | 12 s | 40 |
| static headroom | 0 s | 0 |
I would not consider it settled without evidence: Measure node provisioning time in your own environment and compare it against how long the workload can queue.
Autoscaling reacts, and headroom is what covers the reaction time.
Curated: · Written: · Reviewed:
QA-37The platform is degraded and forty teams cannot deploy. What is your first move?(show answer)
The first thing I would establish about platform incident response is which developer decision it is meant to remove.
A platform incident is a multiplied incident: the affected population is every team that depends on the interface, and most of them will discover it by failing. Communication is part of mitigation because it prevents forty parallel investigations.
Concretely, publish the incident before it is understood, name the affected capability in the developer's terms rather than yours, give an explicit workaround or say there is none, and update on a stated cadence so teams stop checking.
The reason for that specificity is a failure I have seen: A registry outage went unannounced for 35 minutes; 14 teams independently investigated their own pipelines and 3 rolled back unrelated changes trying to fix a failure that was not theirs.
Cost of the silent interval.
| Interval | Teams investigating | Unrelated rollbacks |
|---|---|---|
| 0-10 min | 3 | 0 |
| 10-35 min | 14 | 3 |
| after notice | 0 | 0 |
I would not consider it settled without evidence: Measure time from detection to first published notice, and treat it as a platform reliability metric.
Silence during a platform incident multiplies the investigation.
Curated: · Written: · Reviewed:
QA-38What SLO does a platform owe its internal users, and how is it different from a product SLO?(show answer)
I would start platform service level objectives from the journey a developer actually walks, not from the architecture diagram.
A platform's users build on its availability, so a platform SLO has to be at least as strong as what it lets teams promise downstream. An internal SLO with no consequence is a statement of intent rather than an objective.
Concretely, define objectives on the developer-visible operations — can I deploy, can I provision, can I read logs — rather than on component uptime, publish the error budget, and tie budget exhaustion to a concrete change in platform behaviour such as pausing platform-side feature work.
The reason for that specificity is a failure I have seen: A platform reported 99.95 percent component uptime while the deploy path failed 4 percent of attempts, because retries eventually succeeded and the SLO measured components rather than the operation the developer cared about.
Component uptime against operation success.
| Measure | Value | Developer experience |
|---|---|---|
| component uptime | 99.95% | looks healthy |
| deploy first-attempt | 96.0% | 1 in 25 fails |
| deploy eventual | 99.9% | after retries |
I would not consider it settled without evidence: Measure the developer-visible operation end to end, counting a first-attempt failure as a failure.
Measure the operation the developer performs, not the parts you own.
Curated: · Written: · Reviewed:
QA-39A product team pages the platform for every alert they cannot immediately explain. How would you fix that?(show answer)
This is an area where a high adoption number and a good handling of on-call boundaries between platform and product are not the same thing.
An escalation boundary that is not written down defaults to whoever answers. The fix is a stated division of responsibility plus the diagnostic surface that lets a product team answer the questions on their side of it.
Concretely, publish what the platform owns and what the team owns, give the team the tools to distinguish the two themselves, and treat a recurring escalation class as a missing diagnostic rather than as a training problem.
The reason for that specificity is a failure I have seen: Sixty percent of platform pages were product-code issues; the platform on-call spent a median 25 minutes per page proving the platform was healthy, and the recurring cause was that teams could not see their own resource throttling.
Pages by actual cause.
| Cause | Share | Median time to prove |
|---|---|---|
| platform | 40% | 30 min |
| product code | 47% | 25 min |
| product config | 13% | 20 min |
I would not consider it settled without evidence: Classify pages by where the cause turned out to be, and build for the largest misrouted class.
A recurring wrong page is a missing diagnostic, not a misbehaving team.
Curated: · Written: · Reviewed:
QA-40Two pipelines applied the same infrastructure module at once. What protects you?(show answer)
My answer to infrastructure as code state begins with who carries the work afterwards, because an abstraction is an obligation.
Declarative infrastructure tools keep state, and concurrent applies against one state file produce a result that matches neither run. Locking is what makes the operation safe, and it has to be enforced by the state backend rather than by convention.
Concretely, use a backend with locking, keep state per environment and per component so the blast radius of a corrupted state is bounded, and make manual applies from a workstation impossible rather than discouraged.
The reason for that specificity is a failure I have seen: Two concurrent applies against one state file left 3 resources tracked in state but absent in the cloud and 2 present but untracked; reconciling took a day of manual import.
Result of a concurrent apply.
| Category | Resources | Reconciliation |
|---|---|---|
| in state, absent | 3 | manual removal |
| present, untracked | 2 | manual import |
| consistent | 41 | none |
I would not consider it settled without evidence: Confirm the backend takes an exclusive lock by starting a second apply during a first and observing it block.
State without a lock is state two runs can disagree about.
Curated: · Written: · Reviewed:
QA-41A shared infrastructure module has 30 input variables. What does that tell you?(show answer)
I would treat module design in infrastructure as code as a product decision with users who can route around me.
A module whose inputs mirror the underlying resource has not encapsulated a decision, it has added a layer. Its value comes from the choices it makes on the caller's behalf, and each input handed back is a choice it declined to make.
Concretely, express the module in terms of the intent it serves, make the safe configuration the default, and where a caller genuinely needs an unusual value, prefer a second module for that case over a flag that changes behaviour for everyone.
The reason for that specificity is a failure I have seen: A 30-input module was called 22 ways across the estate; a security change had to be verified against every call site individually, and 2 of them silently disabled it through an input combination nobody had anticipated.
Call-site variety by module design.
| Design | Inputs | Distinct configurations | Call sites |
|---|---|---|---|
| resource mirror | 30 | 22 | 24 |
| intent-shaped | 4 | 3 | 24 |
I would not consider it settled without evidence: Count distinct call configurations across the estate; a number close to the number of callers means the module is not encapsulating anything.
Every input is a decision the module refused to make.
Curated: · Written: · Reviewed:
QA-42How do you find infrastructure that exists but is not in code?(show answer)
The useful question for drift between infrastructure code and reality is what happens on the day it breaks and the developer has to look underneath.
Declarative code describes what it manages and is silent about everything else. Resources created by hand during an incident, or by a tool nobody inventoried, are invisible to a plan that only compares against its own state.
Concretely, reconcile the cloud inventory against the union of all state files on a schedule, report unmanaged resources with their creation identity and date, and either import or remove them rather than leaving a third category.
The reason for that specificity is a failure I have seen: An account held 61 unmanaged resources including 4 security groups with open inbound rules created during an incident 14 months earlier; every plan ran clean because none of them were in state.
Inventory reconciliation.
| Category | Count | Risk |
|---|---|---|
| managed | 412 | tracked |
| unmanaged | 61 | unreviewed |
| unmanaged, network | 4 | open inbound |
I would not consider it settled without evidence: Run a scheduled inventory reconciliation and report the unmanaged count as a tracked number.
A clean plan says nothing about what the code does not know exists.
Curated: · Written: · Reviewed:
QA-43How does the platform deploy itself, and what makes that different from deploying a product service?(show answer)
I would settle the platform's own deployment by watching a team use it rather than by asking whether they liked it.
A platform that deploys itself with itself has a circular dependency: a bad platform release can remove the mechanism needed to fix it. The recovery path has to exist outside the thing being recovered.
Concretely, keep a documented and exercised break-glass path that does not depend on the platform's own control plane, stage platform releases behind the product estate, and roll platform changes out progressively across tenants rather than globally.
The reason for that specificity is a failure I have seen: A broken platform release removed the ability to deploy, including the ability to deploy the platform; recovery took 3 hours and required an engineer with cluster credentials that only two people held.
Recovery path availability.
| Path | Depends on platform | People able | Time |
|---|---|---|---|
| normal deploy | yes | all | n/a in outage |
| break-glass, untested | no | 2 | 3 h |
| break-glass, drilled | no | 9 | 25 min |
I would not consider it settled without evidence: Exercise the break-glass path on a schedule with someone who does not normally use it, and time it.
The thing that fixes the platform cannot depend on the platform.
Curated: · Written: · Reviewed:
QA-44A service catalogue is 8 months old and half its entries are wrong. What would you change?(show answer)
The judgement in backstage-style service catalogues is what stays visible and overridable, not how much gets hidden.
A catalogue maintained by hand decays at the rate the estate changes, and a catalogue nobody trusts is worse than none because it produces confident wrong answers. Accuracy comes from deriving entries from systems of record rather than from asking people to update them.
Concretely, derive ownership from the repository and deployment metadata, mark every field with its source and freshness, and fail loudly on an entry whose source has gone stale rather than continuing to display it.
The reason for that specificity is a failure I have seen: An incident escalated to an owner listed in the catalogue who had left 5 months earlier; the repository's own metadata had been current the whole time and was not used.
Field accuracy by source.
| Field | Source | Accuracy |
|---|---|---|
| owner | manual | 54% |
| owner | repo metadata | 97% |
| runtime version | derived | 99% |
I would not consider it settled without evidence: Sample entries against the systems of record and report accuracy as a percentage, per field.
A catalogue is only as fresh as the source each field is derived from.
Curated: · Written: · Reviewed:
QA-45Where should platform documentation live and what makes it decay?(show answer)
Where platform teams lose the interview on documentation as part of the interface is mandating what has not yet earned adoption.
Documentation decays when it is separated from the thing it describes, because the change that invalidates it happens somewhere the writer is not looking. Keeping it in the same repository and testing its examples ties its life to the interface's.
Concretely, keep docs beside the code, execute the examples in the test suite so a broken example fails the build, and generate reference material from the schema rather than transcribing it.
The reason for that specificity is a failure I have seen: A quickstart's example manifest referenced an API version removed 4 months earlier; 30 percent of new-team onboarding tickets that quarter were people following it and failing.
Onboarding tickets by cause.
| Cause | Share |
|---|---|
| stale example | 30% |
| missing permission | 22% |
| genuine question | 48% |
I would not consider it settled without evidence: Execute every documented example in CI and fail on a broken one.
An example nobody runs is a claim nobody checks.
Curated: · Written: · Reviewed:
QA-46How long should it take a new service to reach production, and how would you know?(show answer)
I would answer onboarding time as a platform metric by separating the risk control that is genuinely required from the preference that is not.
Time from repository creation to first production deploy is the single measure that exercises most of the platform at once. It is hard to game because it is an outcome the developer feels, and it decomposes into steps you can attack.
Concretely, instrument each step of the journey, report the distribution rather than the best case, and attack the largest step rather than the most annoying one. Watch the tail, because the median hides the teams that got stuck.
The reason for that specificity is a failure I have seen: A platform advertised a 30-minute onboarding based on a demo; the measured median was 2.5 days and the p90 was 9 days, dominated by a manual access-approval step nobody had counted as part of the journey.
Where the onboarding time goes.
| Step | Median | p90 |
|---|---|---|
| scaffold | 8 min | 20 min |
| access approval | 1.8 d | 7 d |
| first deploy | 40 min | 3 h |
I would not consider it settled without evidence: Instrument the journey end to end including waits for approval, and report median and p90.
The waiting is part of the journey even when it is not your system.
Curated: · Written: · Reviewed:
QA-47Every service was created from the same template. Two years later, how similar are they?(show answer)
The engineering content of template drift across the estate is the migration path and the support policy, not the templating language.
A template is a copy at a moment in time. Without a mechanism that propagates change, the estate diverges from the day it is created, and the platform's mental model of "how our services work" stops matching any of them.
Concretely, distribute shared behaviour as a dependency that can be updated rather than as copied files where possible, and where copying is unavoidable, provide an automated update and measure the share of the estate on each version.
The reason for that specificity is a failure I have seen: A change to the health-check contract was applied to the template; 6 months later 38 of 71 services still had the old behaviour because nothing propagated it, and the platform's readiness logic had to support both indefinitely.
Estate by template version after 6 months.
| Version | Services | Behaviour |
|---|---|---|
| v3 current | 33 | new contract |
| v2 | 26 | old contract |
| v1 | 12 | old contract |
I would not consider it settled without evidence: Report the share of the estate on each template version, and treat a wide spread as the platform's own technical debt.
A template is copied once and diverges thereafter.
Curated: · Written: · Reviewed:
QA-48CI takes 22 minutes and developers batch their changes because of it. Where do you look first?(show answer)
Before rolling build cache and pipeline latency out I would write down what would tell me it had made things worse.
Pipeline latency changes developer behaviour: past a threshold people batch, and batching increases change size and blast radius. The fix is usually caching and parallelism rather than more compute, because most pipelines repeat work they already did.
Concretely, measure where the time goes per stage, cache dependency resolution and build outputs with a key that actually changes only when inputs do, parallelise independent stages, and run the slow full suite after merge rather than on every push where that is safe.
The reason for that specificity is a failure I have seen: A 22-minute pipeline spent 14 minutes resolving dependencies because the cache key included a timestamp and never hit; fixing the key took the pipeline to 7 minutes and merge frequency roughly doubled within a month.
Pipeline time by stage.
| Stage | Before | Cache hit | After |
|---|---|---|---|
| dependencies | 14 min | 0% | 40 s |
| compile | 5 min | 60% | 4 min |
| test | 3 min | n/a | 2.5 min |
I would not consider it settled without evidence: Report cache hit rate per stage alongside stage duration, since a cache that never hits looks the same as no cache.
A cache key that always changes is not a cache.
Curated: · Written: · Reviewed:
QA-49Teams retry CI until it passes. Is that their problem or yours?(show answer)
The first thing I would establish about flaky tests as a platform problem is which developer decision it is meant to remove.
A flake rate that makes retrying rational destroys the signal for everyone, and it is usually caused by shared infrastructure — contended runners, shared test databases, timeouts sized for an idle system — which is the platform's to fix.
Concretely, measure flake rate per test and per runner class, separate infrastructure-caused failures from test-caused ones, and quarantine a genuinely flaky test rather than letting it train everyone to retry.
The reason for that specificity is a failure I have seen: A 9 percent job failure rate was 7 points infrastructure contention and 2 points real; developers retried by reflex and a genuine regression was retried through 4 times before someone read the output.
Failure causes over one week.
| Cause | Share of failures | Owner |
|---|---|---|
| runner contention | 61% | platform |
| shared test DB | 17% | platform |
| real regression | 22% | team |
I would not consider it settled without evidence: Classify failures by cause and report the infrastructure-caused share as a platform metric.
When retrying is rational, the signal is already gone.
Curated: · Written: · Reviewed:
QA-50The container registry is 40 TB and growing. What is your retention policy?(show answer)
I would start artifact registry hygiene from the journey a developer actually walks, not from the architecture diagram.
Registries grow without bound because deletion is risky: an image that looks unused may be the one a rollback needs. Retention has to be driven by what is deployed and what could be rolled back to, not by age alone.
Concretely, retain everything currently deployed anywhere, plus a stated number of prior versions per service, plus anything referenced by a release tag, and expire the rest. Verify against live deployments before deleting rather than against a tag pattern.
The reason for that specificity is a failure I have seen: An age-based cleanup deleted images older than 90 days and removed the last-known-good image for a service that deployed quarterly; the next rollback had nothing to roll back to.
What each policy would delete.
| Policy | Deleted | Rollback targets lost |
|---|---|---|
| older than 90 d | 61% | 4 |
| keep 10 per service | 58% | 0 |
| keep 10 + deployed | 55% | 0 |
I would not consider it settled without evidence: Check candidate deletions against the set of images currently running and recently deployed, and refuse any that appear.
Retention follows what is deployable, not what is recent.
Curated: · Written: · Reviewed:
QA-51A provisioning API is called twice because a client retried. What should happen?(show answer)
This is an area where a high adoption number and a good handling of idempotency in platform APIs are not the same thing.
Any API a client can retry will be called twice, and a network timeout gives the client no way to know whether the first call succeeded. Idempotency is what makes retrying safe, and it has to be provided by the server because the client cannot infer it.
Concretely, accept a client-supplied idempotency key, store the outcome against it, and return the original result on a repeat rather than performing the work again. Scope the key's lifetime to longer than the client's retry window.
The reason for that specificity is a failure I have seen: A provisioning call timed out at the client after the server had succeeded; the retry created a second database instance, and the duplicate was found 3 weeks later on the invoice.
Retry outcomes by design.
| Design | Resources after 3 retries | Response |
|---|---|---|
| no key | 3 | 3 different |
| key, no store | 3 | 3 different |
| key with stored outcome | 1 | identical |
I would not consider it settled without evidence: Replay the same request with the same key and assert that exactly one resource exists and the same response is returned.
A client that times out cannot tell success from failure, so the server must.
Curated: · Written: · Reviewed:
QA-52Provisioning takes 4 minutes. How should the API express that?(show answer)
My answer to asynchronous provisioning and status begins with who carries the work afterwards, because an abstraction is an obligation.
An operation that outlasts a reasonable request timeout cannot be modelled as a synchronous call. Holding the connection open converts a slow operation into a failed one whenever anything in the path has a shorter timeout than the work.
Concretely, return an operation resource immediately, let the client poll or subscribe, make the operation carry enough status to distinguish queued from failed, and keep it retrievable after completion so a client that lost the connection can still learn the outcome.
The reason for that specificity is a failure I have seen: A synchronous provisioning call was cut by a 60-second load-balancer timeout; the work completed server-side and the client treated it as a failure and retried, producing duplicates the API had no way to detect.
Timeouts along the path.
| Hop | Timeout | Operation p99 |
|---|---|---|
| client | 120 s | 240 s |
| load balancer | 60 s | 240 s |
| server | 300 s | 240 s |
I would not consider it settled without evidence: Check the shortest timeout in the request path against the operation's p99 duration before choosing a synchronous design.
The shortest timeout in the path defines what synchronous means.
Curated: · Written: · Reviewed:
QA-53An inventory endpoint returns everything and started timing out at 8,000 resources. What is the fix?(show answer)
I would treat pagination in platform inventory APIs as a product decision with users who can route around me.
An unpaginated collection endpoint is a latency and memory problem that scales with the estate, and offset pagination over a changing collection skips and repeats items as things are inserted and deleted between pages.
Concretely, use cursor pagination keyed on a stable sort, cap page size at the server, and make the cursor opaque so its encoding can change. Where callers need everything, offer an export rather than letting them walk the whole collection page by page.
The reason for that specificity is a failure I have seen: A nightly reconciliation walked an offset-paginated inventory while resources were being created; it processed 40 duplicates and missed 31 resources, and the missing ones were silently treated as deleted.
Walking 8,000 resources under concurrent writes.
| Scheme | Duplicates | Missed |
|---|---|---|
| offset | 40 | 31 |
| cursor on stable key | 0 | 0 |
I would not consider it settled without evidence: Page through the collection while inserting and deleting concurrently and assert that every stable item appears exactly once.
Offset pagination assumes the collection stands still.
Curated: · Written: · Reviewed:
QA-54Should an internal platform API rate-limit its own colleagues?(show answer)
The useful question for rate limits on internal APIs is what happens on the day it breaks and the developer has to look underneath.
Internal callers are not more careful than external ones; they are less visible. A retry loop in one team's pipeline can saturate a shared control plane, and without limits the failure mode is that everyone's provisioning stops.
Concretely, limit per caller identity rather than per source address, return a retry-after the client library actually honours, and make the limit generous enough not to shape normal work while bounding the worst loop. Publish limits so a team can design against them.
The reason for that specificity is a failure I have seen: A misconfigured retry loop issued about 40,000 requests an hour against a control plane sized for 2,000; provisioning failed for every team for 90 minutes and the cause was one pipeline.
One caller against shared capacity.
| Caller behaviour | Requests/h | Capacity | Others affected |
|---|---|---|---|
| normal | 200 | 2000 | no |
| retry loop, no limit | 40000 | 2000 | all |
| retry loop, limited | 2000 cap | 2000 | no |
I would not consider it settled without evidence: Load-test the control plane to saturation and set the limit from measured capacity rather than from a round number.
A shared control plane is shared until one caller takes it.
Curated: · Written: · Reviewed:
QA-55The platform notifies teams by webhook and some events are missed. What did the design get wrong?(show answer)
I would settle webhooks from the platform to teams by watching a team use it rather than by asking whether they liked it.
A webhook is an at-least-once delivery attempt to an endpoint you do not control. Treating it as reliable delivery of an ordered stream means every consumer outage becomes lost state that nobody can reconstruct.
Concretely, retry with backoff and a dead-letter, sign the payload, include an event id and a sequence so consumers can detect gaps and duplicates, and provide a catch-up query so a consumer that was down can reconcile rather than asking you to replay.
The reason for that specificity is a failure I have seen: A team's endpoint was down for 40 minutes during a deploy; 61 events were dropped after 3 retries and the team discovered the gap 2 weeks later when their internal inventory disagreed with reality.
Consumer down for 40 minutes.
| Design | Events lost | Recovery |
|---|---|---|
| 3 retries only | 61 | none |
| retries + dead letter | 0 | manual |
| + catch-up query | 0 | automatic |
I would not consider it settled without evidence: Take a consumer offline for longer than the retry window and confirm it can reconstruct the missed state without the platform replaying it.
Design for the consumer being down, because it will be.
Curated: · Written: · Reviewed:
QA-56Where would you add a cache in a provisioning control plane, and where would you refuse to?(show answer)
The judgement in caching in a platform control plane is what stays visible and overridable, not how much gets hidden.
Caching authorisation decisions or resource state trades correctness for latency in exactly the places where correctness matters. A cached permission means a revoked access still works, and a cached resource state means a decision made against something that no longer exists.
Concretely, cache what is expensive and slow-changing — schemas, region metadata, template contents — and read through for authorisation and current state. Where a permission cache is unavoidable, keep the TTL inside the revocation SLA and state that SLA explicitly.
The reason for that specificity is a failure I have seen: A 15-minute permission cache meant a revoked engineer could still provision for a quarter of an hour after offboarding, which was discovered during an access review rather than by design.
What is safe to cache.
| Data | Change rate | Cacheable | TTL |
|---|---|---|---|
| region metadata | months | yes | 24 h |
| template contents | days | yes | 10 min |
| permissions | seconds | with care | < SLA |
I would not consider it settled without evidence: State the revocation SLA and confirm no cache TTL in the authorisation path exceeds it.
A cached permission is a permission that outlives its revocation.
Curated: · Written: · Reviewed:
QA-57Should a provisioning request go straight to the cloud API or through a queue?(show answer)
Where platform teams lose the interview on queue-based decoupling in platform workflows is mandating what has not yet earned adoption.
A queue decouples the request rate from the downstream capacity and survives a downstream outage, at the cost of an unbounded backlog and a delay between request and effect. The right choice depends on whether the caller needs a synchronous answer and how the downstream fails.
Concretely, queue where the downstream is rate-limited or flaky and the operation is already asynchronous, bound the queue and shed rather than growing without limit, and expose queue depth and age to the caller so a backlog is visible rather than experienced as slowness.
The reason for that specificity is a failure I have seen: An unbounded queue absorbed a cloud API outage and accumulated 14,000 provisioning requests; when the API recovered it processed them all, hit the account rate limit, and extended the outage by 2 hours.
Recovery after a 3-hour downstream outage.
| Drain policy | Backlog | Recovery time |
|---|---|---|
| unbounded, full speed | 14000 | +2 h |
| unbounded, rate-matched | 14000 | 90 min |
| bounded at 2000 | 2000 | 20 min |
I would not consider it settled without evidence: Test the recovery path after a downstream outage, not just the outage itself, and confirm the drain rate respects downstream limits.
The queue survives the outage and the drain can extend it.
Curated: · Written: · Reviewed:
QA-58How should the platform handle a schema migration during a rolling deployment?(show answer)
I would answer database migrations in a platform-managed pipeline by separating the risk control that is genuinely required from the preference that is not.
During a rolling deployment two versions of the application run against one database, so any migration that is not compatible with both will break whichever version it did not anticipate. The constraint is the overlap, not the migration itself.
Concretely, split incompatible changes into expand and contract steps separated by a release: add the new shape, write both, migrate readers, then remove the old shape once no running version depends on it. Make the platform's pipeline able to express that ordering rather than running migrations as one pre-deploy step.
The reason for that specificity is a failure I have seen: A column rename ran as a single pre-deploy migration; the previous version, still serving half the traffic, failed on every write for the 6 minutes the rollout took.
Rename as one step against expand and contract.
| Approach | Releases | Failing window |
|---|---|---|
| rename in place | 1 | 6 min |
| expand, migrate, contract | 3 | 0 |
I would not consider it settled without evidence: Run the old and new versions against the migrated schema together in a lower environment before promoting.
A rolling deploy means both versions must work on one schema.
Curated: · Written: · Reviewed:
QA-59Services autoscale and the shared database hits its connection limit. What is the platform's part of the fix?(show answer)
The engineering content of connection pooling behind a platform abstraction is the migration path and the support policy, not the templating language.
Connection count scales with replica count times pool size, which the platform controls through autoscaling while the database limit is fixed. A per-service pool sized sensibly in isolation becomes an estate-wide problem once replicas multiply.
Concretely, compute the estate-wide worst case from maximum replicas and default pool size, put a pooler in front of the database, and make the default pool size a platform decision informed by that arithmetic rather than a per-service default copied from a tutorial.
The reason for that specificity is a failure I have seen: A default pool of 20 across 12 services scaling to 30 replicas implied a worst case of 7,200 connections against a limit of 500; the limit was reached during a routine traffic peak and unrelated services could not connect.
Worst-case connections.
| Config | Replicas | Pool | Total | Limit |
|---|---|---|---|---|
| default | 30 x 12 | 20 | 7200 | 500 |
| reduced pool | 30 x 12 | 5 | 1800 | 500 |
| with pooler | 30 x 12 | 5 | 120 | 500 |
I would not consider it settled without evidence: Compute maximum replicas times pool size across the estate and compare against the database limit before enabling aggressive autoscaling.
The database sees the whole estate, not one service's sensible default.
Curated: · Written: · Reviewed:
QA-60Should the platform own the feature-flag system, and what is the risk if it does?(show answer)
Before rolling platform-managed feature flags out I would write down what would tell me it had made things worse.
A flag system becomes a control plane for behaviour across the estate, which means its availability requirement is at least as strong as the services that read it, and a bad flag change can affect everything at once.
Concretely, make the client fail static to the last known value rather than to the default, cache locally so an evaluation service outage does not become an estate outage, keep an audit trail of who changed which flag, and treat a global flag change with the same rollout discipline as a deployment.
The reason for that specificity is a failure I have seen: The flag service became unreachable and clients fell back to their compiled defaults; 9 services reverted to behaviour that had been off for 8 months, including one that re-enabled a deprecated payment path.
Client behaviour when the flag service is down.
| Fallback | Services changed behaviour |
|---|---|
| compiled default | 9 |
| last known value | 0 |
| last known + alarm | 0 |
I would not consider it settled without evidence: Take the evaluation service offline in a lower environment and confirm every client keeps its last known values.
Falling back to the default is falling back to a version nobody is running.
Curated: · Written: · Reviewed:
QA-61The platform runs in one region and the business wants a second. What is the hard part?(show answer)
The first thing I would establish about cross-region platform failover is which developer decision it is meant to remove.
The hard part is state and the decision to fail over, not compute. Stateless components duplicate easily; the control plane's database, the artifact store, and the secret material each need a replication and consistency answer, and an untested failover is an assumption.
Concretely, enumerate every piece of state with its replication mode and lag, decide whether failover is manual or automatic and who has authority, and exercise the failover on a schedule with production-like data rather than documenting it.
The reason for that specificity is a failure I have seen: A documented failover had never been run; when it was needed, the secondary's control-plane database was 40 minutes behind and the artifact store had not been replicated at all, so nothing could be deployed after failover.
State inventory at failover.
| State | Replication | Lag | Usable |
|---|---|---|---|
| control plane DB | async | 40 min | partly |
| artifact store | none | — | no |
| secrets | synced | 0 | yes |
I would not consider it settled without evidence: Run a real failover on a schedule and record recovery point and recovery time actually achieved.
A failover that has not been run is a document.
Curated: · Written: · Reviewed:
QA-62How do you decide whether the platform builds a capability or buys it?(show answer)
I would start choosing between build and buy for a capability from the journey a developer actually walks, not from the architecture diagram.
The decision is per capability rather than per platform, and the cost of buying is not the licence but the integration, the migration, and the loss of control over the roadmap. The cost of building is not the build but the maintenance across the years the interface lives.
Concretely, compare total cost over the expected life including maintenance and migration, keep the internal product ownership and the developer-facing interface even when the implementation is bought, and prefer buying where the capability is not a differentiator and the data is portable.
The reason for that specificity is a failure I have seen: A bought capability was exposed to developers with the vendor's own interface; when the contract ended 3 years later, 61 services had vendor-specific configuration and the migration took 8 months.
Migration cost by interface ownership.
| Interface | Services touched | Migration |
|---|---|---|
| vendor's own | 61 | 8 months |
| platform's own | 0 | 6 weeks |
I would not consider it settled without evidence: Ask what a migration away would cost before signing, and design the developer-facing interface so that answer stays small.
Buy the implementation, own the interface.
Curated: · Written: · Reviewed:
QA-63You need to retire a capability that 30 teams use. How do you run it?(show answer)
This is an area where a high adoption number and a good handling of deprecating a platform capability are not the same thing.
Deprecation is a migration project you are running on other people's schedules. Announcing an end date without providing the replacement and the path to it moves the work to teams who did not plan for it and will not prioritise it.
Concretely, ship the replacement first, provide automated migration for the mechanical part, instrument remaining usage per team, contact the tail individually, and set the removal date from observed usage rather than from the original announcement.
The reason for that specificity is a failure I have seen: A capability was announced for removal with 90 days' notice and no migration tooling; at the deadline 19 of 30 teams were still on it, the date slipped 3 times, and the announcement lost credibility for the next deprecation.
Teams remaining at the announced date.
| Support offered | Teams migrated | Date held |
|---|---|---|
| announcement only | 11 of 30 | no |
| + migration guide | 19 of 30 | no |
| + automated migration | 29 of 30 | yes |
I would not consider it settled without evidence: Track remaining usage per team weekly and only set a firm date when the tail is small enough to contact individually.
A date without a path is a date that will slip.
Curated: · Written: · Reviewed:
QA-64Every team asks for something different. How do you decide what the platform builds next?(show answer)
My answer to platform request intake begins with who carries the work afterwards, because an abstraction is an obligation.
Requests arrive as solutions, and building the requested solution for the loudest team produces a platform that is a pile of features. The useful unit is the underlying problem and how many teams share it.
Concretely, record the problem behind each request rather than the request, cluster them, size each cluster by the teams affected and the cost they carry, and publish the decision including what you declined and why so teams stop re-asking.
The reason for that specificity is a failure I have seen: A platform built 6 requested features for 6 teams in a quarter; a later analysis found 4 of them were the same underlying problem — no way to run a one-off job — solved 4 incompatible ways.
Requests against underlying problems.
| Quarter | Requests built | Distinct problems | Duplicated effort |
|---|---|---|---|
| as requested | 6 | 3 | 3 |
| clustered first | 3 | 3 | 0 |
I would not consider it settled without evidence: Cluster the recorded problems quarterly and check whether recent work addressed clusters or individual requests.
Build for the problem behind the request, or you build it several times.
Curated: · Written: · Reviewed:
QA-65Developers say the platform works in CI and not on their machine. Whose problem is that?(show answer)
I would treat supporting local development as a product decision with users who can route around me.
The developer's inner loop is part of the platform's surface whether or not the platform team claims it. A platform that only works in its own pipeline pushes every developer to invent a local approximation, and those approximations diverge from production in ways nobody tracks.
Concretely, provide a supported local path that uses the same configuration and artifacts as CI, state explicitly where it differs from production, and measure inner-loop time because that is what the developer experiences most often.
The reason for that specificity is a failure I have seen: With no supported local path, 5 teams each built their own docker-compose approximation; 3 of them diverged on the database version and a bug reproducing only on one version took 4 days to attribute.
Local setups in the estate.
| Approach | Distinct setups | Divergences | Inner loop |
|---|---|---|---|
| unsupported | 5 | 3 | 40 s-4 min |
| supported path | 1 | 0 | 25 s |
I would not consider it settled without evidence: Measure the inner-loop cycle time and the number of distinct local setups in use across the estate.
If you do not supply the inner loop, everyone builds their own.
Curated: · Written: · Reviewed:
QA-66Half your team's week goes to manual requests. How do you get it back?(show answer)
The useful question for reducing platform toil is what happens on the day it breaks and the developer has to look underneath.
Manual work that recurs is a missing interface. Automating the most frequent request first is usually right, but the most valuable target is the one whose manual handling carries the most risk of being done wrong.
Concretely, categorise the queue by frequency and by consequence of error, automate the intersection first, and make the automated path the only path so the manual one stops accumulating exceptions.
The reason for that specificity is a failure I have seen: A team automated the most frequent request — log access, 40 a month, low risk — and left production database restores manual at 3 a month; a restore was performed against the wrong instance and cost 4 hours of data.
Queue by frequency and consequence.
| Request | Per month | Consequence of error | Priority |
|---|---|---|---|
| log access | 40 | low | 2nd |
| DB restore | 3 | severe | 1st |
| quota bump | 22 | low | 3rd |
I would not consider it settled without evidence: Rank the queue by frequency times consequence rather than by frequency alone.
Automate what is dangerous done by hand, not only what is common.
Curated: · Written: · Reviewed:
QA-67Should platform engineers be embedded in product teams or kept as a separate team?(show answer)
I would settle platform team topology by watching a team use it rather than by asking whether they liked it.
A separate team builds reusable capability and risks losing touch with the work; embedding keeps context and produces bespoke solutions that do not generalise. The useful pattern is a platform team with a deliberate mechanism for context rather than a choice between the two.
Concretely, keep the platform team as the owner of the interfaces, and run rotations or enabling engagements where platform engineers work inside a product team for a bounded period with a stated goal of bringing findings back.
The reason for that specificity is a failure I have seen: A fully embedded model produced 4 deployment systems in 18 months; a fully separate model produced one system whose 3 highest-priority features were not what any team had asked for.
Outcome by model over 18 months.
| Model | Systems built | Reused across teams |
|---|---|---|
| fully embedded | 4 | 0 |
| fully separate | 1 | partly |
| central + rotation | 1 | yes |
I would not consider it settled without evidence: Check whether platform priorities are traceable to observed product-team work, and whether shared capability actually generalises across teams.
Own the interface centrally and get the context deliberately.
Curated: · Written: · Reviewed:
QA-68An auditor asks how you know every service encrypts data at rest. What answer holds up?(show answer)
The judgement in compliance controls in the paved road is what stays visible and overridable, not how much gets hidden.
A control asserted in a policy document is a claim about intent. A control implemented in the path that creates the resource, with evidence emitted automatically, is a claim about the estate that can be checked.
Concretely, implement the control in the provisioning path so non-compliance is impossible rather than detectable, emit evidence as a byproduct of normal operation, and continuously scan for resources created outside that path.
The reason for that specificity is a failure I have seen: A policy required encryption at rest; an audit found 4 of 210 storage buckets unencrypted, all created directly through the cloud console during incidents, and the policy document had no way to have caught them.
Where the exceptions came from.
| Creation route | Buckets | Unencrypted |
|---|---|---|
| paved road | 206 | 0 |
| console, direct | 4 | 4 |
I would not consider it settled without evidence: Scan the actual estate for the control rather than sampling the paved road, and report the count created outside it.
Evidence about the estate, not about the path most of it took.
Curated: · Written: · Reviewed:
QA-69What makes a platform audit log actually useful during an investigation?(show answer)
Where platform teams lose the interview on audit logging that survives review is mandating what has not yet earned adoption.
An audit log is only useful if it answers who did what to which resource and when, in a form that cannot be edited by the party being audited. Logs that record the platform's actions without the initiating human identity answer none of that.
Concretely, carry the initiating identity through every hop including automation, write to append-only storage outside the control of the systems being logged, and retain for the period a real investigation needs rather than the period logs are cheap.
The reason for that specificity is a failure I have seen: Every production change appeared in the audit log as the platform's own service account; attributing a specific deletion to a person required correlating 3 systems by timestamp and took 2 days.
Attribution by log design.
| Log records | Names person | Time to attribute |
|---|---|---|
| service account only | no | 2 days |
| + initiating identity | yes | 2 min |
I would not consider it settled without evidence: Take a real change and confirm the audit log names the initiating person without needing a second system.
An audit log that names your own automation names nobody.
Curated: · Written: · Reviewed:
QA-70The cloud provider is degraded and your platform is affected. What do you do beyond waiting?(show answer)
I would answer handling a shared incident with a cloud provider by separating the risk control that is genuinely required from the preference that is not.
A provider incident is still your incident from the developer's point of view, and the only useful posture is to know which of your capabilities depend on the degraded service and what each can do without it.
Concretely, maintain a dependency map from platform capability to provider service, publish which capabilities are affected in your own terms, and have a stated degraded mode per capability rather than a single up-or-down status.
The reason for that specificity is a failure I have seen: A regional object-storage degradation broke artifact pulls; the platform status page said operational because its own health checks did not exercise storage, and 40 teams reported the same failure independently.
What the health check exercised.
| Check | Exercises storage | Detected outage |
|---|---|---|
| control plane ping | no | no |
| end-to-end artifact pull | yes | yes |
I would not consider it settled without evidence: Make health checks exercise the provider dependencies the capability actually uses, end to end.
Your status page should reflect what developers can do, not what you host.
Curated: · Written: · Reviewed:
QA-71Which platform capabilities must keep working when everything else is broken?(show answer)
The engineering content of graceful degradation of platform features is the migration path and the support policy, not the templating language.
Not all platform capabilities are equal in an incident. The ability to deploy a fix, roll back, and read logs is what teams need during an outage; provisioning and cataloguing can wait. Designing for that ordering is different from designing for uniform availability.
Concretely, rank capabilities by what an incident requires, keep the incident-critical path on the fewest dependencies, and test it with the non-critical components deliberately failed.
The reason for that specificity is a failure I have seen: A shared control-plane database outage removed provisioning and, because rollback read its target version from the same database, removed rollback too; teams could not revert the change that had triggered the incident.
Capability survival under control-plane loss.
| Capability | Depends on shared DB | Survives | Teams blocked |
|---|---|---|---|
| provisioning | yes | no | 40 |
| rollback, coupled | yes | no | 40 |
| rollback, decoupled | no | yes | 0 |
| log read | no | yes | 0 |
I would not consider it settled without evidence: Fail each shared dependency in turn and confirm rollback and log access survive.
Rollback must not depend on what fails during a rollout.
Curated: · Written: · Reviewed:
QA-72How do you test a platform whose behaviour is mostly about other people's workloads?(show answer)
Before rolling testing the platform itself out I would write down what would tell me it had made things worse.
Unit tests cover the platform's own logic and say nothing about whether a real workload can be onboarded, deployed, and recovered. That requires exercising the developer journey against real infrastructure, which is slower and is the only test that covers the product.
Concretely, keep a continuously running synthetic service that goes through the full journey on every platform release, assert on the developer-visible outcome rather than internal state, and treat a failure of that journey as a release blocker.
The reason for that specificity is a failure I have seen: A release passed 900 unit tests and broke onboarding for new services, because the failure was in the interaction between the scaffolding tool and a permission change; no test exercised both.
What each layer caught.
| Layer | Tests | Caught this defect |
|---|---|---|
| unit | 900 | no |
| component | 120 | no |
| synthetic journey | 1 | yes |
I would not consider it settled without evidence: Run the full onboarding and deploy journey on every release candidate, against real infrastructure.
The platform's test is a developer's journey, not its own functions.
Curated: · Written: · Reviewed:
QA-73Name something the platform should deliberately leave exposed.(show answer)
The first thing I would establish about choosing what not to abstract is which developer decision it is meant to remove.
Anything a developer must reason about during an incident is a poor candidate for abstraction, because that is exactly when the abstraction's cost is highest and the platform's help is slowest. Domain-specific behaviour is another: it varies per team and generalising it produces a configuration language nobody wants.
Concretely, abstract the undifferentiated and the risky-by-default; leave exposed what varies genuinely between teams and what a developer needs during an incident. Revisit as the estate grows, since what varied across 5 teams may be uniform across 50.
The reason for that specificity is a failure I have seen: A platform abstracted request routing including retry and timeout policy; every team needed different values, the configuration surface grew to 22 fields, and the abstraction became a worse version of the library it hid.
Override rate by abstracted concern.
| Concern | Teams overriding | Verdict |
|---|---|---|
| TLS configuration | 0 of 24 | abstract |
| log shipping | 1 of 24 | abstract |
| retry policy | 21 of 24 | expose |
I would not consider it settled without evidence: Count how many teams override a given default; a high override rate means the thing should not have been abstracted.
A default everyone overrides was not a default.
Curated: · Written: · Reviewed:
QA-74You are replacing the deployment system for 120 services. How do you sequence it?(show answer)
I would start migrating an estate to a new platform from the journey a developer actually walks, not from the architecture diagram.
A migration of this size fails on the tail rather than the bulk. The first services are volunteers with simple needs; the last are the ones with the requirement nobody designed for, and the project's duration is set by them.
Concretely, find the hard cases first and design for them before the easy migrations build momentum you will have to unwind, run both systems in parallel with a stated end date driven by observed migration, and give the tail dedicated help rather than reminders.
The reason for that specificity is a failure I have seen: A migration moved 100 of 120 services in 4 months and took 14 more months for the last 20, because 3 of them needed a capability the new system had never been designed to support.
Migration progress against effort.
| Phase | Services | Elapsed |
|---|---|---|
| volunteers | 40 | 1 month |
| bulk | 60 | 3 months |
| tail | 20 | 14 months |
I would not consider it settled without evidence: Survey the estate for unusual requirements before starting and design for the hardest three.
The last twenty services set the schedule.
Curated: · Written: · Reviewed:
QA-75What is the cost of running the old and new platforms in parallel, and how do you keep it bounded?(show answer)
This is an area where a high adoption number and a good handling of running two platforms during a migration are not the same thing.
Parallel operation doubles the surface that must be patched, monitored, and supported, and it splits the platform team's attention at exactly the moment the new system needs it. The cost is real and rises with the length of the overlap.
Concretely, freeze feature work on the old system to maintenance only, publish the overlap end date and drive it from migration progress, and refuse new services on the old system from day one so the tail does not grow while you shrink it.
The reason for that specificity is a failure I have seen: The old system kept accepting new services during a migration; 14 were created on it after the migration started, extending the tail by 5 months and requiring their own migration afterwards.
New services during migration.
| Policy | New on old system | Tail extension |
|---|---|---|
| open | 14 | 5 months |
| closed to new | 0 | none |
I would not consider it settled without evidence: Report new-service creation per system weekly and confirm the old one is at zero.
A migration whose source keeps growing does not converge.
Curated: · Written: · Reviewed:
QA-76Where is Python the right choice for platform tooling and where does it hurt?(show answer)
My answer to writing platform tooling in Python begins with who carries the work afterwards, because an abstraction is an obligation.
Python is excellent for glue, one-off automation, and anything whose users are the platform team. It hurts where the tool ships to every developer's machine, because the dependency and interpreter version become a support burden you did not budget for.
Concretely, keep internal automation in whatever the team writes fastest, and distribute developer-facing tools as a single self-contained binary so the platform is not debugging virtual environments across the estate.
The reason for that specificity is a failure I have seen: A Python CLI distributed by pip drew 40 percent of a quarter's platform support tickets, almost all of them dependency conflicts with the developer's own project environment.
Support load by distribution method.
| Distribution | Install tickets/quarter | Versions in use |
|---|---|---|
| pip into project env | 61 | 9 |
| pipx isolated | 12 | 4 |
| single binary | 1 | 2 |
I would not consider it settled without evidence: Count support tickets attributable to the tool's installation rather than to its behaviour.
A tool everyone installs is a distribution problem before it is a language choice.
Curated: · Written: · Reviewed:
QA-77A reconciliation script processes 4,000 resources serially and takes 40 minutes. How would you speed it up?(show answer)
I would treat concurrency in platform automation as a product decision with users who can route around me.
The work is IO-bound on API calls, so concurrency helps and parallelism across cores does not. The limit is the downstream API's rate limit, which means unbounded concurrency converts a slow script into a throttled one and possibly an outage for other callers.
Concretely, use bounded concurrency sized from the downstream rate limit, honour retry-after rather than retrying immediately, and make the work resumable so a failure partway does not restart from the beginning.
The reason for that specificity is a failure I have seen: A script switched to unbounded concurrency across 4,000 resources; it saturated the account rate limit, was throttled to a crawl, and blocked two other teams' provisioning for 25 minutes.
Runtime and collateral effect.
| Concurrency | Runtime | Other callers throttled |
|---|---|---|
| 1 | 40 min | no |
| 20 | 3 min | no |
| unbounded | 22 min | yes |
I would not consider it settled without evidence: Measure the downstream rate limit and set the concurrency bound from it, then confirm other callers are unaffected while the job runs.
Concurrency is bounded by the downstream, not by your machine.
Curated: · Written: · Reviewed:
QA-78Every platform client retries on failure. What turns that into an outage?(show answer)
The useful question for retry and backoff in platform clients is what happens on the day it breaks and the developer has to look underneath.
Synchronised retries amplify a partial failure into a total one: when a downstream slows, every client retries at once, adding load exactly when there is least capacity. Backoff without jitter keeps the retries synchronised.
Concretely, use exponential backoff with full jitter, cap total attempts and total elapsed time, and add a circuit breaker so a persistently failing dependency is not retried at all until a probe succeeds.
The reason for that specificity is a failure I have seen: A 2-second downstream blip caused 400 clients to retry in lockstep at 1, 2, and 4 seconds; the retry load kept the downstream saturated for 6 minutes after the original cause had cleared.
Recovery after a 2-second blip.
| Retry policy | Peak retry load | Recovery |
|---|---|---|
| fixed interval | 400/s | 6 min |
| exponential | 400/s spikes | 3 min |
| exponential + jitter | 40/s | 12 s |
I would not consider it settled without evidence: Load-test a downstream blip with the real client population and confirm the retry load decays rather than sustaining.
Retries without jitter turn a blip into an outage.
Curated: · Written: · Reviewed:
QA-79A developer's manifest is accepted and the service misbehaves at runtime. What should have happened?(show answer)
I would settle structured configuration and validation by watching a team use it rather than by asking whether they liked it.
Configuration accepted now and interpreted later moves the error from the developer's terminal to production. Every field a platform accepts is a promise to interpret it, and an unknown field silently ignored is the most common way that promise is broken.
Concretely, validate against a schema at submission, reject unknown fields rather than ignoring them, check semantic constraints as well as types, and give the error with the field path and an example rather than a parser message.
The reason for that specificity is a failure I have seen: A misspelled field was silently ignored; the service ran with the default replica count of 1 rather than the intended 6, and the misconfiguration was found during a traffic peak.
Where the error surfaces.
| Validation | Error appears | Cost |
|---|---|---|
| none | traffic peak | outage |
| type only | deploy | minutes |
| schema + unknown-field reject | submit | seconds |
I would not consider it settled without evidence: Submit a manifest with an unknown field and confirm it is rejected at submission, not at runtime.
An ignored field is a silent default in production.
Curated: · Written: · Reviewed:
QA-80Reconciling 50,000 cloud resources against 50,000 state entries takes 20 minutes of CPU. What is wrong?(show answer)
The judgement in hash maps in inventory reconciliation is what stays visible and overridable, not how much gets hidden.
Matching two collections by scanning one for each element of the other is quadratic; at 50,000 each that is 2.5 billion comparisons. Indexing one side by key makes it linear and the difference is minutes against seconds.
Concretely, build a dictionary keyed on the stable identifier for one side, then walk the other once, collecting present-in-both, missing, and unmanaged in a single pass. Choose a key that is genuinely stable, since names change and identifiers usually do not.
The reason for that specificity is a failure I have seen: A nested-loop reconciliation over 50,000 resources ran 20 minutes and was scheduled hourly; it consumed a full core continuously and still could not be run more often as the estate grew.
Runtime against estate size.
| Resources | Nested loop | Indexed |
|---|---|---|
| 12500 | 75 s | 0.6 s |
| 25000 | 5 min | 1.2 s |
| 50000 | 20 min | 2.4 s |
I would not consider it settled without evidence: Measure runtime against estate size at two points; a fourfold rise for a doubled estate is the quadratic signature.
Index one side and walk the other once.
Curated: · Written: · Reviewed:
QA-81Infrastructure components must be created in dependency order and the list is maintained by hand. What would you do instead?(show answer)
Where platform teams lose the interview on topological ordering of platform dependencies is mandating what has not yet earned adoption.
A hand-maintained order encodes a graph as a list, and it silently becomes wrong whenever a dependency is added. Deriving the order from declared dependencies makes the graph the source of truth and turns a cycle into an error rather than a deadlock.
Concretely, declare dependencies per component, sort topologically to derive the order, detect and report cycles by name, and parallelise independent components since the sort also tells you which are independent.
The reason for that specificity is a failure I have seen: A hand-ordered list of 34 components missed a new dependency; the create ran in the wrong order, failed at component 19, and left 18 resources provisioned with no rollback path.
Ordering approach and outcome.
| Approach | Components | Cycle detected | Parallel groups |
|---|---|---|---|
| hand-ordered list | 34 | no | 1 |
| topological sort | 34 | yes | 7 |
I would not consider it settled without evidence: Assert that the declared graph is acyclic and that the derived order satisfies every declared edge, in CI.
Derive the order from the graph, or maintain a graph as a list.
Curated: · Written: · Reviewed:
QA-82A regression appeared somewhere in 200 deployments. How do you find it?(show answer)
I would answer binary search in a platform rollback by separating the risk control that is genuinely required from the preference that is not.
When the property being tested is monotonic — everything before the bad change is good and everything after is bad — bisection finds the change in about log2(n) steps rather than n. At 200 deployments that is 8 tests instead of up to 200.
Concretely, confirm monotonicity first, since a flaky or intermittent regression breaks the assumption and bisection will converge on the wrong change. Automate the test so each step is cheap, and record the tested revisions so the search is reproducible.
The reason for that specificity is a failure I have seen: A bisection over an intermittent regression converged on an unrelated change; the team reverted it, the symptom persisted, and the real cause was found 3 days later by reading the diff.
Steps to locate the change.
| Deployments | Linear scan | Bisection |
|---|---|---|
| 50 | up to 50 | 6 |
| 200 | up to 200 | 8 |
| 1000 | up to 1000 | 10 |
I would not consider it settled without evidence: Run the test several times at a known-good and known-bad revision to establish it is deterministic before bisecting.
Bisection assumes a single flip, so check the assumption first.
Curated: · Written: · Reviewed:
QA-83Token bucket or fixed window for a platform API limit?(show answer)
The engineering content of rate limiting algorithm choice is the migration path and the support policy, not the templating language.
A fixed window allows twice the intended rate across a window boundary, because a caller can spend a full window's allowance at its end and another at the start of the next. A token bucket smooths that while still permitting a bounded burst.
Concretely, use a token bucket or sliding window where the burst matters, size the bucket from the burst you intend to permit rather than from the sustained rate, and return the remaining allowance so a well-behaved client can pace itself.
The reason for that specificity is a failure I have seen: A 1,000-per-minute fixed window let a caller issue 2,000 requests in 4 seconds across a boundary; the control plane saturated even though the limit was, on paper, never exceeded.
Burst permitted at the window boundary.
| Algorithm | Sustained | Worst 4-second burst |
|---|---|---|
| fixed window 1000/min | 1000/min | 2000 |
| sliding window | 1000/min | 1000 |
| token bucket, 100 burst | 1000/min | 1100 |
I would not consider it settled without evidence: Test the boundary case deliberately, issuing the full allowance either side of a window edge.
A fixed window permits double its rate at the seam.
Curated: · Written: · Reviewed:
QA-84A cache tier is resized and hit rate collapses. What would consistent hashing have changed?(show answer)
Before rolling consistent hashing for platform routing out I would write down what would tell me it had made things worse.
Modulo-based assignment remaps almost every key when the node count changes, so a resize invalidates nearly the whole cache. Consistent hashing remaps roughly the share belonging to the changed nodes.
Concretely, use consistent hashing with virtual nodes so the load spreads evenly, and where a resize is unavoidable, warm the new topology before shifting traffic rather than accepting the miss storm.
The reason for that specificity is a failure I have seen: Growing a cache tier from 8 to 10 nodes with modulo assignment remapped about 80 percent of keys; hit rate fell from 94 to 19 percent and the origin took 5 times its normal load for 20 minutes.
Keys remapped growing 8 nodes to 10.
| Scheme | Keys remapped | Hit rate after |
|---|---|---|
| modulo | ~80% | 19% |
| consistent hashing | ~20% | 76% |
I would not consider it settled without evidence: Compute the fraction of keys that remap under the planned change before making it.
Modulo assignment makes every resize a cache flush.
Curated: · Written: · Reviewed:
QA-85A worker pool falls behind and memory grows until it is killed. What is missing?(show answer)
The first thing I would establish about queues and backpressure in platform workers is which developer decision it is meant to remove.
A producer that cannot be slowed will fill any buffer. Without backpressure the queue is bounded only by memory, so the failure mode is an out-of-memory kill that loses everything queued rather than a graceful slowdown.
Concretely, bound the queue and make the producer block or shed when it is full, prefer shedding with a clear error over silent dropping, and expose queue depth and oldest-item age so the condition is visible before the kill.
The reason for that specificity is a failure I have seen: An unbounded in-memory queue grew to 2.1 GB during a downstream slowdown; the worker was killed and 40,000 queued items were lost with no record of what they were.
Behaviour under sustained overload.
| Queue | Peak memory | Items lost |
|---|---|---|
| unbounded | 2.1 GB | 40000 |
| bounded, shed | 40 MB | shed, counted |
| bounded, block | 40 MB | 0 |
I would not consider it settled without evidence: Run the producer faster than the consumer for a sustained period and confirm the system sheds rather than growing.
Unbounded queues fail by losing everything at once.
Curated: · Written: · Reviewed:
QA-86Two replicas of a controller both act on the same resource. What went wrong?(show answer)
I would start leader election in platform controllers from the journey a developer actually walks, not from the architecture diagram.
Running multiple replicas for availability requires exactly one to be active, and a lease is what provides that. A lease held without checking its validity before each write allows a partitioned former leader to keep acting.
Concretely, acquire a lease with a bounded term, renew well inside it, and verify leadership before any write rather than only at acquisition. Size the term against the write's duration so a slow write cannot outlive the lease it was authorised under.
The reason for that specificity is a failure I have seen: A controller replica lost network connectivity, kept processing from its in-memory queue for 40 seconds after its 15-second lease expired, and duplicated 3 cloud resources that the new leader had already created.
Behaviour of a partitioned leader.
| Check | Writes after lease expiry | Duplicates |
|---|---|---|
| at acquisition only | 40 s worth | 3 |
| before each write | 0 | 0 |
I would not consider it settled without evidence: Partition the leader and confirm it stops writing before the lease expires, not after it notices.
Leadership is checked at the write, not at startup.
Curated: · Written: · Reviewed:
QA-87You create a resource and immediately read it back and it is not there. Is that a bug?(show answer)
This is an area where a high adoption number and a good handling of eventual consistency in cloud APIs are not the same thing.
Many cloud control planes are eventually consistent: a successful create means the intent was accepted, not that every read replica can see it. Code that creates and immediately reads will fail intermittently, and the rate depends on load.
Concretely, use the create response rather than re-reading, poll with backoff where you must read back, treat not-found after a successful create as a retryable condition for a bounded period, and never take absence as evidence to create again.
The reason for that specificity is a failure I have seen: A provisioning workflow created a resource, read back, saw not-found, and created again; under load the read lag rose and the workflow produced duplicates on about 1 in 30 runs.
Read-after-write lag under load.
| Load | p50 lag | p99 lag | Duplicate rate |
|---|---|---|---|
| light | 40 ms | 300 ms | 0 |
| heavy | 210 ms | 2.4 s | 1 in 30 |
I would not consider it settled without evidence: Measure the read-after-write lag distribution under load and size the polling window from its tail.
A successful create is a promise, not a visible fact.
Curated: · Written: · Reviewed:
QA-88A developer's request takes 30 seconds to fail through four platform hops. What is misconfigured?(show answer)
My answer to timeouts across a platform call chain begins with who carries the work afterwards, because an abstraction is an obligation.
Timeouts have to decrease along a call chain. When an inner hop's timeout exceeds the outer one's remaining budget, the outer gives up while the inner keeps working, wasting capacity on a result nobody will read.
Concretely, propagate a deadline rather than a per-hop timeout, so each hop knows how much budget remains and can refuse work it cannot finish, and make the outermost timeout the one the developer experiences.
The reason for that specificity is a failure I have seen: Four hops each configured with a 30-second timeout meant a failing innermost call held resources at all four levels for the full 30 seconds, and the retry at each level multiplied the load 16-fold.
Budget along the chain.
| Hop | Fixed timeout | With deadline |
|---|---|---|
| ingress | 30 s | 30 s |
| control plane | 30 s | 26 s |
| provisioner | 30 s | 20 s |
| cloud call | 30 s | 14 s |
I would not consider it settled without evidence: Trace one failing request and confirm each hop gives up before the one above it.
Propagate the deadline so each hop knows what is left.
Curated: · Written: · Reviewed:
QA-89Before changing a shared module, how do you find out what it affects?(show answer)
I would treat graph traversal for blast radius as a product decision with users who can route around me.
Dependency impact is a reachability question on a directed graph, and the answer is everything reachable from the changed node rather than its direct dependents. Stopping at one level understates the blast radius, usually by a lot.
Concretely, build the graph from declared dependencies and lockfiles rather than from documentation, traverse transitively while guarding against cycles, and report the affected set with the path so an owner can see why they are included.
The reason for that specificity is a failure I have seen: A change was reviewed against 4 direct dependents; the transitive set was 38 services, and 2 of them broke on a behaviour the direct dependents did not exercise.
Affected set by traversal depth.
| Depth | Services | Broke |
|---|---|---|
| direct | 4 | 0 |
| 2 levels | 17 | 1 |
| transitive | 38 | 2 |
I would not consider it settled without evidence: Report the transitive affected set with paths before approving a change to a shared component.
Blast radius is reachability, not adjacency.
Curated: · Written: · Reviewed:
QA-90A dashboard sorts 200,000 rows in the browser and freezes. What is the right split?(show answer)
The useful question for sorting and paging platform dashboards is what happens on the day it breaks and the developer has to look underneath.
Sorting and filtering belong where the data and the index are. Shipping the full set to the client to sort it moves both the bandwidth and the CPU to the least capable part of the system.
Concretely, sort and filter server-side against an index, return a page, and keep client-side sorting only for a page already in hand. Where the client genuinely needs everything, that is an export rather than a dashboard.
The reason for that specificity is a failure I have seen: A dashboard fetched 200,000 rows as 140 MB of JSON and sorted in JavaScript; first paint took 45 seconds and the tab used 2 GB before becoming responsive.
Dashboard load at 200,000 rows.
| Design | Payload | Time to interactive |
|---|---|---|
| client sort | 140 MB | 45 s |
| server sort, 50/page | 40 kB | 0.4 s |
I would not consider it settled without evidence: Measure payload size and time to interactive at the largest realistic dataset, not at a development sample.
Sort where the index is.
Curated: · Written: · Reviewed:
QA-91Should the platform patch running hosts or replace them?(show answer)
I would settle immutable infrastructure versus in-place patching by watching a team use it rather than by asking whether they liked it.
In-place patching produces hosts whose state is the sum of their history, so two hosts built from the same image diverge over time and a failure on one may not reproduce on another. Replacement makes the image the only state and the divergence impossible.
Concretely, build a new image, roll it out progressively, and replace rather than patch. Keep in-place change for genuine emergencies and treat any such change as a defect in the image pipeline to be fixed in the next build.
The reason for that specificity is a failure I have seen: Two years of in-place patching left 3 of 40 hosts with a different library version; a bug reproduced only on those 3 and took a week to attribute because the hosts were nominally identical.
Divergence across 40 nominally identical hosts.
| Model | Distinct package sets | Reproducibility |
|---|---|---|
| in-place patching | 6 | poor |
| image replacement | 1 | exact |
I would not consider it settled without evidence: Compare package inventories across hosts of the same role and report any divergence as a defect.
A host you patch is a host with a history nobody recorded.
Curated: · Written: · Reviewed:
QA-92Who is responsible for a critical CVE in the base image 200 services use?(show answer)
The judgement in base image maintenance is what stays visible and overridable, not how much gets hidden.
A shared base image concentrates both the benefit and the obligation: one fix protects everyone, and until services rebuild they all carry the vulnerability. Publishing a fixed image is not the same as the estate being fixed.
Concretely, publish the patched image, then drive and measure rebuild across the estate, automating the rebuild where the platform controls the pipeline. Report estate coverage rather than image availability.
The reason for that specificity is a failure I have seen: A patched base image was published and announced; 6 weeks later 84 of 200 services had not rebuilt because nothing triggered them to, and the platform's report said the CVE was remediated.
Remediation by measure.
| Measure | Week 1 | Week 6 |
|---|---|---|
| patched image published | yes | yes |
| services rebuilt | 41 of 200 | 116 of 200 |
| with automated rebuild | 180 of 200 | 200 of 200 |
I would not consider it settled without evidence: Report the share of running workloads on the patched image, not the existence of the patched image.
Publishing the fix is the start of remediation, not the end.
Curated: · Written: · Reviewed:
QA-93A scanner reports 4,000 findings across the estate. How do you make that actionable?(show answer)
Where platform teams lose the interview on dependency and supply-chain scanning noise is mandating what has not yet earned adoption.
A finding count is not a risk measure. Most findings are in dependencies that are not reachable at runtime or not exposed to untrusted input, and a list that does not distinguish those trains everyone to ignore all of them.
Concretely, filter by reachability and exploitability, prioritise by whether the affected code path is reachable from an untrusted input, deduplicate across services sharing a base image, and give teams a small ranked list rather than the raw output.
The reason for that specificity is a failure I have seen: A 4,000-finding report was distributed weekly; teams stopped reading it, and a genuinely exploitable finding in a public-facing service sat unactioned for 5 weeks among the noise.
Filtering 4,000 raw findings.
| Filter | Remaining | Actioned |
|---|---|---|
| raw | 4000 | 2% |
| deduplicated | 610 | 9% |
| reachable only | 44 | 84% |
I would not consider it settled without evidence: Measure the fraction of reported findings that teams action, and treat a low rate as a signal-quality problem rather than a compliance one.
A list nobody reads protects nothing.
Curated: · Written: · Reviewed:
QA-94How do you tell forty teams about a platform change without being ignored?(show answer)
I would answer platform changelog and communication by separating the risk control that is genuinely required from the preference that is not.
Announcement volume trains the audience. A channel that carries every change at the same urgency gets filtered out, and the one message that required action goes with it.
Concretely, separate changes that require action from those that do not, route action-required messages to the specific teams affected rather than broadcasting, and include what breaks, by when, and the exact command to fix it.
The reason for that specificity is a failure I have seen: A breaking change was announced in a channel carrying 30 routine notices a week; 22 of 40 teams missed it and their pipelines failed on the cutover morning.
Action rate by routing.
| Routing | Teams affected | Acted before deadline |
|---|---|---|
| broadcast channel | 40 | 18 |
| targeted, action-required | 40 | 37 |
I would not consider it settled without evidence: Measure the fraction of affected teams that acted before the deadline, per announcement, and adjust routing rather than repeating the broadcast.
Broadcast everything and nothing is read.
Curated: · Written: · Reviewed:
QA-95A senior team asks for direct production cluster access to debug. What is your answer?(show answer)
The engineering content of handling a request the platform should refuse is the migration path and the support policy, not the templating language.
Refusing without an alternative loses the argument and the relationship, because the underlying need — seeing what is happening in production — is legitimate. The right response is to satisfy the need through a path that keeps the audit trail and the blast radius.
Concretely, offer time-bounded, audited, read-mostly access through a break-glass path, find out which diagnostic gap drove the request, and close it so the next team does not need to ask.
The reason for that specificity is a failure I have seen: A blanket refusal led the team to obtain credentials from a colleague; the access was unaudited, indefinite, and discovered 4 months later during an access review.
Outcome by response.
| Response | Audited | Duration | Gap closed |
|---|---|---|---|
| refuse | no | 4 months | no |
| grant standing access | yes | indefinite | no |
| break-glass + fix gap | yes | 2 h | yes |
I would not consider it settled without evidence: Track break-glass requests by cause and confirm the recurring causes are being closed rather than repeatedly granted.
A refused legitimate need routes around you.
Curated: · Written: · Reviewed:
QA-96What makes a platform postmortem useful to teams who were not involved?(show answer)
Before rolling reading a platform postmortem out I would write down what would tell me it had made things worse.
A platform incident affects every dependent team, so the postmortem's audience is the estate rather than the platform team. It has to say what developers experienced, what they should have done, and what the platform changed so they do not need to do it next time.
Concretely, write the developer-visible symptom first, name the detection gap separately from the cause, state the action items with owners and dates, and publish where the affected teams already read rather than in the platform's own repository.
The reason for that specificity is a failure I have seen: A postmortem described a control-plane database failover in internal terms; 5 teams that had been unable to deploy for 90 minutes could not tell from it whether their own retries had been the right response.
What the postmortem answered.
| Question | Internal write-up | Estate write-up |
|---|---|---|
| what broke | yes | yes |
| what developers saw | no | yes |
| what to do next time | no | yes |
I would not consider it settled without evidence: Have someone from an affected team read the draft and confirm they can answer what they should do next time.
Write it for the people who felt it, not the people who fixed it.
Curated: · Written: · Reviewed:
QA-97Product leadership wants a feature and the platform needs to pay down its own debt. How do you argue it?(show answer)
The first thing I would establish about prioritising platform work against product pressure is which developer decision it is meant to remove.
Platform debt is invisible to the people who fund it because it shows up as everyone else's slowness rather than as an outage. Making it visible means expressing it in the currency the audience uses, which is delivery time and incident cost rather than architecture.
Concretely, quantify the debt as time it costs the estate — support hours, waiting, failed deployments — and present the trade as a choice between two measurable outcomes rather than as a principled objection.
The reason for that specificity is a failure I have seen: A platform team argued for a rewrite on architectural grounds for 3 quarters and lost each time; when they re-framed it as 340 engineer-hours a quarter of estate-wide waiting, it was funded in one cycle.
Same proposal, two framings.
| Framing | Quarters requested | Funded |
|---|---|---|
| architectural | 3 | no |
| 340 h/quarter of waiting | 1 | yes |
I would not consider it settled without evidence: Measure the estate-wide cost of the debt before asking for the time to fix it.
Debt funded in the currency the funder counts.
Curated: · Written: · Reviewed:
QA-98When should a platform team stop building a capability?(show answer)
I would start knowing when the platform is done enough from the journey a developer actually walks, not from the architecture diagram.
A capability is done when the marginal improvement no longer changes what teams can do, and continuing past that point spends the platform's credibility along with its time. The signal is in the outcome measures flattening while effort continues.
Concretely, set a target outcome before building, watch it rather than the feature list, and move on when it flattens. Where it does not move at all, the honest conclusion is that the capability was not the constraint.
The reason for that specificity is a failure I have seen: A team spent 2 further quarters improving a scaffolding tool after time-to-first-deploy had already flattened at 40 minutes; the actual constraint had moved to access approval, which nobody was working on.
Effort against outcome.
| Quarter | Effort | Time to first deploy |
|---|---|---|
| Q1 | high | 2.5 d to 4 h |
| Q2 | high | 4 h to 45 min |
| Q3 | high | 45 min to 41 min |
| Q4 | high | 41 min to 40 min |
I would not consider it settled without evidence: Track the outcome the capability was meant to move and stop when it flattens, then re-measure where the constraint went.
Stop when the number stops moving and find where it moved to.
Curated: · Written: · Reviewed:
QA-99A candidate proposes a full internal developer platform for a 12-engineer company. What would you probe?(show answer)
This is an area where a high adoption number and a good handling of platform interviews and the scale question are not the same thing.
Platform investment is justified by repeated cost across teams, and at small scale that repetition does not exist. The right answer at 12 engineers is usually conventions, a template, and a managed service, not a platform team.
Concretely, ask what work is being repeated and by how many people, size the investment against that, and prefer the smallest thing that removes the cost. Revisit as the organisation grows, since the answer genuinely changes.
The reason for that specificity is a failure I have seen: A 12-engineer company built an internal platform; 2 of 12 engineers maintained it, product delivery slowed measurably, and the platform was retired in favour of a managed service 14 months later.
Right answer by organisation size.
| Engineers | Teams | Right answer |
|---|---|---|
| 12 | 2 | conventions + managed service |
| 60 | 8 | templates + a part-time owner |
| 300 | 40 | a platform team |
I would not consider it settled without evidence: Count the teams sharing the problem before proposing a shared solution.
A platform needs enough teams to be worth having.
Curated: · Written: · Reviewed:
QA-100An interviewer asks about a platform decision you got wrong. What makes a good answer?(show answer)
My answer to explaining a platform decision that was wrong begins with who carries the work afterwards, because an abstraction is an obligation.
A useful answer names the decision, the evidence available at the time, the signal that should have changed your mind sooner, and what you now do differently. Anything that ends at "we learned a lot" describes an outcome rather than a change in practice.
Concretely, say what you would have measured earlier and what threshold would have triggered a reversal, because that is the part that transfers to the next decision.
The reason for that specificity is a failure I have seen: A team continued a bespoke orchestration layer for two quarters past the point where its adoption had flattened at 20 percent, because no threshold had been agreed at which they would stop.
Adoption against the reversal decision.
| Quarter | Adoption | Threshold agreed | Action |
|---|---|---|---|
| Q1 | 12% | none | continue |
| Q2 | 19% | none | continue |
| Q3 | 20% | none | continue, plateau begins |
| Q4 | 20% | none | continue |
| Q5 | 20% | 40% | would have stopped |
I would not consider it settled without evidence: Set a reversal threshold at the start of any significant platform investment and write it down.
Decide in advance what would tell you to stop.
Curated: · Written: · Reviewed:
