Overview
Curated: · Written: · Reviewed:
A platform API is a product contract for developer intent
This is the topic senior and staff candidates most often answer as a design essay: they list good principles but never land the judgement calls interviewers are actually probing for. The question behind almost every platform-API interview prompt is: what do you expose, what do you hide, and how do you keep that contract stable for consumers you can't see? A weak answer describes a REST facade over provider SDKs. A strong answer defends a boundary — where it sits, why, and what happens when it's wrong.
The core skill: expose vs. hide
A platform API is the stable contract through which developers, portals, CLIs, automation and other systems discover and operate platform capabilities. It should express domain intent — "a production database with this owner, data class and availability target" — instead of mirroring provider calls or exposing an orchestration graph. The portal is one client. If essential behavior exists only in the UI or a ticket workflow, the platform has no coherent API product.
The judgement call interviewers probe: which details are accidental, and which are decisions the user must own? Hide accidental implementation details, vendor-specific wiring and sequencing the platform owns. Keep intent visible when it affects cost, durability, security, reliability, locality or workload responsibility. Both failure modes lock teams in:
- An API that accepts every underlying provider field has not reduced cognitive load — teams now debug your abstraction and the provider's.
- An API that hides every meaningful tradeoff forces teams into undocumented workarounds, usually raw credentials.
How to recognise a leaky abstraction before it locks teams in: users start asking for "just one more provider field"; support tickets reference internal queue names; teams script around the portal because the API can't express what they need. Each of those is a signal the boundary is in the wrong place, not that the API needs more fields.
Model a small set of cohesive resources with stable identities and ownership. Separate desired configuration from observed status where the domain is declarative; use explicit operations for genuinely imperative or one-shot work rather than pretending everything is desired state. Define the full lifecycle — create, get, update, import, transfer, deprecate, delete. Avoid giant options maps, weakly typed blobs and universal resource endpoints; they move compatibility and validation problems to runtime, which is exactly where you can't fix them.
Choosing the interface medium
Interviewers often ask "REST or gRPC or CRDs?" — the weak answer picks one. The honest answer: an internal platform usually ends up with several consistent views of one abstraction, because different consumers need different interaction models:
- REST/HTTP+JSON for the portal, ad-hoc curl, and broad language coverage.
- gRPC for internal service-to-service paths where typed schemas and streaming matter.
- Kubernetes CRDs + Operators when the platform's control plane is Kubernetes and consumers already live in GitOps tooling — the CRD is the API, and reconciliation is the semantics.
- Terraform providers because teams manage infrastructure as code; a provider that wraps your API keeps one source of truth.
- CLI and generated SDKs as projections of the same schema, not hand-maintained parallels.
The rule that matters: transport parity does not require identical interaction. The portal can guide discovery; the CLI supports scripts; declarative workflows reconcile state. All must share server-side authentication, authorization, validation, policy and state semantics. Never rely on hidden client validation, and never give a UI privileged backdoors other authorized clients can't audit. If the CLI can do something the API can't, you have two products.
Abstracting over heterogeneous backends
When the platform fronts multiple clouds, clusters or IaaS providers, the trap is the lowest common denominator: a field set so diluted it describes no provider accurately. The opposite trap is leaking provider state into domain resources. The defensible middle:
- Define an adapter boundary so domain resources never expose tool-specific state. Specify adapter capabilities, error mapping, idempotency, ownership and observability; certify implementations with contract tests.
- When providers genuinely diverge — different durability models, consistency guarantees, networking semantics, lifecycle behavior — surface that as capability discovery ("this backend supports PITR, that one supports read replicas") or as distinct resource classes, not as a misleading shared field that means different things per provider.
- Do not pretend two providers offer identical semantics. A
highAvailability: truethat means synchronous replication on one cloud and best-effort failover on another is a contract violation waiting for an incident.
Escape hatches must be intentional: bounded extension points, composition, annotations or an off-road resource class with explicit support boundaries — not raw credentials or arbitrary execution. Validate and namespace extensions, prevent them overriding mandatory controls, make them visible in inventory and policy. Then analyze usage: recurring escape-hatch patterns tell you whether to pave a missing capability or preserve a separate path.
Long-running work, idempotency and partial failure
Provisioning takes minutes to hours, so the operation is part of the contract. Interviewers probe whether you treat async work as a first-class resource or an afterthought:
- Return durable operation identity, state, metadata and eventual result or structured error. Validate everything rejectable before starting.
- Expose queued/running/succeeded/failed/cancelled, progress or stage where meaningful, timestamps, actor, target, retryability and safe diagnostics.
- Cancellation stops future work but cannot promise atomic reversal of completed provider effects. Say so in the contract. Define whether concurrent operations queue, conflict or supersede.
Design idempotency explicitly. User-specified resource identity makes create retryable; request IDs deduplicate operations for a documented retention window. The test interviewers reach for: a retry after timeout must not create duplicate infrastructure. Bind request identity to caller, method and normalized intent; return a compatible prior outcome or current resource state; reject reuse for different input. And state the limit plainly — idempotency is not exactly-once execution; it's deduplicated effects with observable state.
Partial-failure recovery is where designs get exposed. If the database was created but the network policy wasn't, what does get return? Decide whether operations are resumable, whether status reports partial provider state, and who reconciles drift. Eventual consistency between desired and observed status is fine — undocumented inconsistency is not.
Errors, authorization and the resource model
Errors are part of the contract. Return a stable machine-readable reason and category plus a concise actionable message, safe context, remediation and a documentation link where useful. Distinguish invalid input, failed precondition, conflict, quota, unavailable dependency and internal error — clients need to know whether to fix, wait or retry. Check authorization before revealing resource existence (a 403 vs 404 policy is a contract decision, not an accident). Never require clients to parse human prose, leak provider secrets or dump raw internal failures.
Authorization follows the resource model. Scope every resource to a tenant, owner and environment; validate references and actions server-side; separate get/list, plan, apply, admin, override and delete. Filter collections so metadata doesn't leak. The platform's provider credentials do not authorize the caller. Preserve actor and delegation through orchestration and audit every privileged transition.
Compatibility for known-but-numerous consumers
Here's the nuance that separates staff answers from senior ones: an internal API is held to a different standard than a public one. You know who your consumers are and can reach them — but there are hundreds of teams, some with automation you'll never see. You get less freedom than "we'll just migrate them," and more than "never change anything."
- Compatibility includes source, wire and semantics. Additive fields can still break clients through new required behavior, enum values, changed defaults, pagination or stricter limits.
- Do not rename fields or resource identities in place, silently change meanings, or remove behavior because the portal no longer uses it — the CLI consumer you forgot about is the one that breaks.
- Publish stability tiers, supported versions, deprecation policy and migration paths. For internal APIs, deprecation can carry dates and owner outreach, not just a header.
- Test old clients against new servers, new clients against supported servers, persisted resources and operational workflows. Compatibility fixtures and adapter conformance suites are the mechanism.
Operating the API as a production dependency
The API and its adapters are privileged production dependencies. Version schemas and generated clients, publish changelogs, rate limits and SLOs, isolate tenants, protect credentials, bound payloads and queues, and design backpressure. Observe request rate, latency, error reason, operation age, stale status, provider health and abandonment — without high-cardinality or sensitive labels. One property worth stating in an interview: existing workloads should remain safe if the platform API cannot accept new work. Rejection must degrade gracefully, not cascade.
What interviewers probe, and what a weak answer sounds like
- "How do you decide what to expose?" Weak: "we abstract everything behind a clean interface." Strong: a named criterion (user-owned tradeoffs stay visible; platform-owned sequencing stays hidden) plus an example of each.
- "What happens when two providers diverge?" Weak: "we take the common subset." Strong: capability discovery or distinct resource classes, with the LCD trap named.
- "A client times out mid-provision and retries." Weak: "we make it idempotent." Strong: identity binding, retention window, what the retry returns, and the exactly-once caveat.
- "How do you version an internal API?" Weak: "semver." Strong: semantic vs. wire vs. behavioral compatibility, and how internal reachability changes deprecation.
- "The portal can do something the API can't." Weak: "that's a UI feature." Strong: that's a missing product capability and a contract gap.
Likely follow-ups: how you'd test the contract (validation, authorization, retry and request-ID reuse, timeout-after-side-effect, concurrent update, stale version, partial provider failure, cancellation, pagination, deletion and name reuse); how you'd migrate an existing portal-only platform to an API-first one; and where you'd deliberately not abstract.
The best abstraction is not the one with the fewest fields. It is the smallest stable contract that preserves the decisions users must own while letting the platform safely evolve everything else.
