Skip to content
Tech Interview Prep home
Technical interview guide

Self-Service Infrastructure

Letting teams provision what they need through an API or portal, with policy enforced automatically instead of manually.

Read
40 min
Practice MCQs
25
Interview QA
25
Edition
v3
Editorial status
Reviewed
Relevant for
Platform Engineer

Scope: AWS and Microsoft platform engineering guidance and HashiCorp Terraform/HCP Terraform documentation current 2026-08-31; confirm product-tier and version-specific policy features for the target deployment.

Overview

Curated: · Written: · Reviewed:

Self-service infrastructure is governed delegated authority

Self-service infrastructure lets an authorized team request, change, inspect and retire supported infrastructure through an API, CLI, portal or version-controlled declaration without a routine manual ticket. It is not unrestricted cloud access and it is not a form backed by a hidden human queue. The platform converts a small amount of workload intent into repeatable provider actions while enforcing organizational boundaries and returning durable status and evidence.

Start from a specific journey and eligible cohort: for example, a development database, ephemeral environment, message queue or production service account. Discover current waits, handoffs, errors, risks and lifecycle gaps. Define supported regions, sizes, environments, data classes, availability, retention and extension points. A universal “create any cloud resource” abstraction exposes provider complexity and concentrates privilege without delivering a coherent product.

Collect the minimum intent that cannot be safely inferred: owner, workload/system, environment, data classification, availability target, capacity bounds, connectivity, retention, cost center and expiry. Provide reviewed defaults and show their consequences. Hide accidental details, but expose cost, destructive changes, data durability and operational responsibilities. Preview or plan should show intended creates, updates, replacements and deletes before high-impact apply.

Authenticate the human or workload through the organizational identity system, authorize every request server-side against team, environment and resource class, and separate permission to plan from permission to apply or override policy. The orchestrator uses short-lived, scoped credentials for the target tenant; users do not inherit its broad service identity. Validate cross-resource references to prevent confused-deputy escalation and isolate state, runners, secrets, networks and provider accounts where risk warrants.

Use version-controlled reusable modules and declarative desired state when they fit. Pin provider and module dependencies, review changes, record provenance and generate a durable execution plan. Maintain a single authoritative binding between desired objects and real resources. Remote state and workflow metadata are sensitive production data: encrypt them, restrict and audit access, back them up and use locking or serialized reconciliation. Force-unlocking or editing state requires evidence that no writer remains and a recovery plan.

Policy belongs at several layers. Validate request shape and intrinsic relationships early; scan configuration and dependencies; evaluate the concrete plan for security, privacy, region, tagging and cost; verify effective deployed state and drift. Assign each policy an owner and choose prevent, warn, detect or approve based on risk. Return the rule, rationale, affected object and remediation. Overrides are privileged, scoped, justified, audited, time-bound and revisited—not a generic bypass button.

Provisioning is asynchronous and partially failing. Give each request immutable identity, actor, desired version and idempotency key; expose queued, planning, awaiting approval, applying, succeeded, failed and cancelled states with safe logs and resulting resources. Serialize operations that share state, make provider actions safely retryable and reconcile after ambiguous timeouts. Cancellation cannot promise rollback of already-completed effects; the system must observe and reconcile partial state.

Manage quotas and cost without turning every request into approval. Enforce team/environment capacity and rate limits, show estimated cost and major drivers, attribute actual cost through validated immutable ownership metadata, alert on anomalies and expire ephemeral resources. A budget is not an availability policy: production capacity changes must preserve SLO and data constraints. Provide a documented escalation for justified bursts and measure quota false positives and abandonment.

Drift policy must be explicit. Detect differences between declared, state and provider reality. For platform-authoritative fields, reconcile or propose correction; for emergency changes, preserve evidence and import or revert through an authorized workflow; for externally owned fields, observe rather than fight. Avoid endless oscillation between controllers. A portal inventory is not automatically the source of truth; identify authoritative systems for ownership, desired state, billing and runtime facts.

Deletion is part of self-service. Show dependencies, data and recovery consequences; distinguish stop, detach, archive and destroy. Require stronger authorization for irreversible or production deletion, use retention and legal-hold rules, revoke credentials and network access, delete or transfer dependents, verify provider completion and retain audit evidence. Name reuse must not cause a new request to adopt old resources accidentally.

The platform itself is a privileged production dependency. Protect templates, plugins, policies, runners, state and credentials; sign/version releases, canary changes and separate duties. Define SLOs for plan and apply, queue latency and correctness. Design degraded modes: existing workloads should remain safe if new provisioning is unavailable. Back up state and configuration, rehearse provider outage, state loss, stuck lock, policy-service failure, credential compromise and regional recovery.

Measure outcomes from a baseline: request completion time and success, environment wait, failed/retried runs, policy findings prevented early, drift age, deletion completion, idle cost, support tickets, eligible adoption, abandonment and user comprehension. Counts of portal clicks or resources created can reward waste. Segment by environment, resource class and team maturity, and use qualitative research to find hidden manual work.

Self-service is successful when teams gain faster safe autonomy and the organization retains truthful control, inventory and recovery. The central interview insight is that removing a human approver does not remove governance; it moves governance into product design, authorization, policy, evidence, bounded automation and lifecycle operations.