Overview
Curated: · Written: · Reviewed:
Incident management restores control; postmortems convert experience into safer systems
An incident is an event that threatens a service, users, data, security, legal obligations or business operations enough to require coordinated response. Declare early when scope or resolution is uncertain: a formal structure can shrink as evidence improves, while delayed declaration leaves parallel responders, contradictory changes and stakeholder silence. Severity is a current coordination and impact signal, not a verdict about competence or a permanent label; reclassify it as evidence changes.
Preparation happens before the page. Define declaration and escalation criteria, roles, authority, communication channels, durable working log, status templates, evidence handling, vendor/legal/privacy routes and accessible runbooks. Exercise the process through drills and prior-incident replays. Avoid relying exclusively on the same production system being repaired for coordination. Maintain contact and access paths for control-plane, identity, network and observability failures.
The incident commander owns coordination, priorities, roles, decision cadence and incident state—not every keyboard action. An operations lead directs technical investigation and mitigation. A communications lead gives accurate, audience-appropriate updates and protects responders from repeated inquiries. Scribes/timekeepers preserve decisions, hypotheses, observations and changes. Security, privacy, legal, support, vendor and executive roles join by consequence. One person may initially hold several roles, but delegate as complexity grows and hand off explicitly with current state, risks and next decisions.
Mitigate user harm before pursuing a complete causal explanation. Prefer bounded, reversible actions such as rollback, feature disablement, traffic shift, isolation, capacity relief or safe degraded mode. State the hypothesis, expected effect, owner, time and rollback before each material change; change one major variable at a time when feasible. Preserve volatile and durable evidence without delaying urgent safety actions. A mitigation is not proof of root cause, and service recovery is not incident closure.
Maintain one authoritative incident record: declaration/time basis, severity, affected journeys/cohorts, known/unknown, customer/data/security impact, owners/roles, timeline, hypotheses, commands/changes, decisions, communications, evidence references and next update. Separate facts from inference. Record detection time, impact start/end as later reconstructed, acknowledgement, mitigation and recovery; do not fabricate precision from incomplete clocks or telemetry. Sensitive logs, samples, credentials, customer data and security indicators require access, retention and chain-of-custody controls.
Communication must be timely, consistent and honest about uncertainty. Internal responder updates need technical state, active work, blockers and requests. Executives/support need scope, business risk, decisions and next update. External status should describe experienced symptoms, affected functions, mitigations and recovery without speculation, blame, exploitable detail or unsupported deadlines. Regulatory, contractual, law-enforcement and data-subject notification decisions belong to authorized legal/privacy/security processes and documented clocks.
Recovery requires more than a green graph. Validate critical user journeys and data integrity, stop unsafe retries, reconcile partial or duplicated effects, watch fallback and capacity headroom, and define an observation period. Remove emergency access and temporary controls through owned follow-up. Capture remaining unknowns, exposed populations and delayed work. Hand back from incident command to normal ownership explicitly, with monitoring and rollback criteria.
A postmortem is a reviewed evidence record of impact, detection, response, causal and contributing conditions, and owned improvement work. Predeclare triggers such as material user impact, data loss, security/privacy event, lengthy recovery, manual intervention, monitoring failure or high learning value. Blameless means assuming people acted reasonably with their information, incentives, tools and constraints; it does not mean consequence-free, vague, or unwilling to examine decisions. Look beyond a single “root cause” to defenses, latent conditions, coupling, change, detection, decision support and recovery factors.
Build the timeline from chat, alerts, logs, traces, deploys, tickets and interviews, reconciling clocks and confidence. Quantify affected users/requests/records/value, duration, regions, objectives and support burden; label estimates and unknowns. Explain why the system allowed the condition, why safeguards did not prevent or detect it, why mitigation took its observed time, and what worked well. Distinguish causal evidence from correlation and counterfactual claims.
Action items must change recurrence likelihood, blast radius, detection, mitigation or organizational readiness. Each needs an owner, priority justified by risk, due date, completion evidence and verification. “Be careful,” “retrain everyone,” “add monitoring” or “rewrite it” are not sufficient without the specific control, failure mode and proof. Track items with normal engineering work, escalate overdue high-risk work, verify through tests/drills and measure recurrence. A published document with abandoned actions has not completed the learning loop.
Review and share postmortems with audiences who can use the learning, while redacting or restricting personal, customer, security and legally sensitive material. Store searchable structured metadata to identify recurring contributing factors across incidents, but protect access and avoid ranking individuals or teams by incident count. Review the incident process itself: declaration delay, role clarity, change collisions, communication load, evidence gaps, responder fatigue and handoff quality.
Test the complete system. Run tabletop and production-safe drills for partial outage, data corruption, credential compromise, regional/control-plane failure, observability loss, vendor outage and simultaneous incidents. Verify access, role assignment, communications, evidence preservation, containment, rollback, recovery validation, notification routes and postmortem action closure. The goal is not a perfect document or a lower mean time metric at any cost; it is reduced user harm, trustworthy recovery and demonstrably stronger future response.
Worked example: checkout SEV1, clocks versus actions
Users report empty carts at 14:00. The error-rate graph is already red. Nobody is paged until 14:22. Incident declared 14:31. Rollback completes 14:48. Checkout returns 200 the whole time.
| clock | wall | minutes users are down | what the postmortem action must change |
|---|---|---|---|
| user reports | 14:00 | 0 (start) | not an action — that is detection failure |
| page | 14:22 | 22 | alert on checkout error rate, not HTTP 200 |
| declare | 14:31 | 31 | on-call knows who is commander without a Slack hunt |
| rollback | 14:48 | 48 | one-command revert, already rehearsed |
"Be more careful" does not move any of those four numbers. An owned alert with a 2-minute burn, a named commander, and a tested rollback does. That table is the interview.
