Overview
Curated: · Written: · Reviewed:
Behavioral interviews and the STAR method
Behavioral interviews look like conversation and are scored like an assessment. This guide treats STAR as a way to put past work onto a 1–5 rubric: the same questions, the same order, a Situation short enough to leave time for Action, and a Result with a number. It assumes you have actually shipped or failed at something — the bank of stories is the product — and it uses OPM structure, EEOC job-relatedness, Amazon's published principles, and SRE postmortem culture as the sources, not interview-coach slogans.
structured interviewing versus a chat
A structured interview asks every candidate the same job-related questions, in the same order, and scores answers on the same anchored scale.
OPM and Campion treat structure as the method: predetermined questions, consistent probes, and a 1–5 rubric. Schmidt and Hunter's meta-analysis puts structured interviews among the stronger predictors of job performance; an unstructured conversation is closer to an unvalidated chat.
Same prompt, two scoring regimes, two different hiring errors.
| regime | questions | scoring | typical failure |
|---|---|---|---|
| unstructured | whatever comes up | gut, 30 min | hire the fluent candidate |
| structured | 4 past-behavior + 1 hypothetical, same order | 1–5 anchors | miss the quiet owner of a 40% latency cut |
| mixed | same questions, no anchors | 'strong communicator' | 0.2 inter-rater kappa |
Interview trap. Believing an unstructured conversation predicts performance better because it feels like real work rewards fluency, which is the bias structure exists to remove.
Engineering practice. Treat the loop as an assessment: write the questions from the job analysis, train interviewers on the anchors, and refuse 'just talking' as the scoring method.
the STAR sequence
STAR is a scoring scaffold: Situation, Task, Action, Result, in that order, each with a job-related fact.
DDI's behavioral interview scores past behavior. Situation bounds the scene (team, system, constraint). Task is the decision you owned. Action is what you did. Result is the measured change. Skipping a letter leaves the interviewer nothing to put on the 1–5 scale.
Four sentences that fill a 1–5 anchor instead of a 90-second ramble.
| letter | 18-word draft | missing this letter costs |
|---|---|---|
| S | Checkout p99 was 1.8 s after a Black Friday 4× traffic spike. | 'big outage somewhere' |
| T | I owned cutting p99 under 400 ms without adding cache nodes. | 'the team decided' |
| A | I moved session reads to Redis with a 5-minute TTL and a stampede lock. | 'we used Redis' |
| R | p99 fell to 320 ms; checkout conversion rose 1.4 points the next week. | 'it went well' |
Interview trap. Treating STAR as optional window dressing leaves an anecdote with no Task and no Result, which gives the rater nothing to place on the 1-5 scale.
Engineering practice. Draft every story as four labeled sentences before the interview, then speak them without reading the labels.
Situation as scope not autobiography
Situation is a one-sentence bound on time, system, and constraint — not a career recap.
Interviewers need enough context to judge the Action. OPM's guide says behavioral questions should elicit a situation or task, the candidate's actions, and their impact; it does not prescribe a universal answer duration. In a four-minute practice answer, a Situation that burns 90 seconds leaves less room for the decision and result being scored.
Budget 20 seconds for S, not 40% of a 4-minute answer.
| part | budget | words | example |
|---|---|---|---|
| S | 20 s | ~40 | Q3 2025; payments API; 12% error budget burned in 6 days |
| T | 20 s | ~40 | I owned stopping the burn without a freeze |
| A | 120 s | ~240 | the three steps you took |
| R | 40 s | ~80 | error rate 2.1% → 0.3%; freeze avoided |
Interview trap. Opening with a five-minute origin story so the interviewer 'gets the culture' spends the Action and Result budget on context nobody scores.
Engineering practice. Cap Situation at two clauses: when, what broke, and the constraint. Move history into a follow-up if they ask.
Task as the decision you owned
Task names the decision, constraint, and owner — you — not the team's mission statement.
Structured scoring looks for job-related competencies on this candidate. 'We needed to scale' is a Situation leftover. 'I had to pick between a 2-week shard and a 2-day read replica, with a Friday launch' is a Task the 1–5 scale can mark.
The Task line is the one that would appear on an on-call ticket.
| weak Task | owned Task |
|---|---|
| Help the team ship search. | Choose whether to freeze ranking or ship a 48-hour BM25 fallback before Tuesday's demo. |
| Improve reliability. | Cut 5xx on /checkout from 1.8% to under 0.5% in 14 days without extra headcount. |
| Be a good teammate. | Decide whether to revert my migration after it added 90 ms p95 for EU users. |
Interview trap. Offering the company OKR as the Task describes a successful quarter rather than the decision this candidate owned, which is the thing the scale marks.
Engineering practice. Write Task as 'I had to X under Y by Z'. If you cannot name X, Y, and Z, you do not have a story yet.
Action as first-person verbs
Action is what you personally did, in verbs a skeptic can check, not what 'we' shipped.
Amazon's leadership loop and OPM's past-behavior probes both ask who moved the work. 'We migrated' hides whether you wrote the dual-write, reviewed the RFC, or watched Slack. First-person verbs (wrote, measured, reverted, paged, negotiated) are the evidence.
Count I-verbs. Zero is a group status update; three is an Action.
| sentence | I-verbs | scores as |
|---|---|---|
| We improved the pipeline. | 0 | Situation leftover |
| I added a unique constraint, I dual-wrote for 7 days, I cut over at 02:00. | 3 | Action |
| I asked SRE to 'look at it'. | 1, weak | delegation without a decision |
Interview trap. Saying 'we' throughout to sound collaborative leaves the panel unable to tell whether you wrote the dual-write, reviewed the RFC, or watched Slack.
Engineering practice. If a sentence has no I-verb, rewrite it or move it to Situation. Keep one sentence of who else was in the room.
Result as measured change
Result is a before/after number, a decision reversed, or a failure contained — not 'stakeholders were happy'.
The 1–5 anchor needs a criterion. Latency 1.8 s → 320 ms, error budget 12% → 2%, hire/no-hire reversed after a 3-page postmortem: those are Results. Mood and applause are not on the scale.
Same Action, two Results; only one survives a follow-up.
| Result | follow-up it survives |
|---|---|
| People liked the dashboard. | none; 'liked' has no unit |
| On-call pages for checkout dropped from 14/week to 3/week over 4 weeks. | 'what was the remaining 3?' |
| We saved money. | none |
| Reserved instances cut compute from $41k/month to $27k/month at the same QPS. | 'what is the 12-month commit risk?' |
Interview trap. Offering the team's good feeling as the Result gives the anchor no criterion, because mood has no unit and survives no follow-up.
Engineering practice. Carry two figures in every story: the starting measurement and the ending one, plus the window (7 days, 1 quarter).
the missing Result follow-up
If you omit Result, a trained interviewer will ask for it; inventing one on the spot is worse than saying you did not measure.
OPM's probes include 'what happened' and 'how did you know'. A candidate who never instrumented the change should say so and name the proxy they used (ticket count, error rate, a 20-user diary study). Fabricating a 40% improvement is an integrity fail, not a STAR fail.
Honesty with a proxy beats a round 40%.
| you say | interviewer hears |
|---|---|
| It improved things a lot, maybe 40%. | unverified |
| I did not have p99. I had 18 support tickets/week; after the fix, 5 in the next 14 days. | measured proxy, n=18→5 |
| Leadership said it was fine. | no Result |
Interview trap. Inventing a round percentage once the number is gone turns a measurement gap into an integrity problem, which costs more than the missing Result did.
Engineering practice. Keep a one-page story bank with the actual dashboards or commit SHAs. If there is no number, say the proxy and the sample size.
we versus I ownership
Credit the team in Situation; keep Action and Result attached to what you decided or built.
Panels score this person. A staff loop still wants the candidate's leverage: the RFC you wrote, the rollback you called, the 2-page decision you got VP sign-off on. Naming collaborators is required; hiding behind them is the tell.
Staff-level still needs an I-verb; the scope of the I-verb is what changed.
| level | Situation | Action I-verb |
|---|---|---|
| mid | 4-person squad, 1 service | I shipped the retry budget |
| senior | 3 squads, 2 regions | I set the dual-write plan and the abort metric |
| staff | 11 teams, a Friday freeze | I wrote the 3-page RFC and I called the rollback at 01:12 |
Interview trap. Avoiding 'I' entirely to sound humble hides the leverage being scored, because it is this candidate the panel has to rate.
Engineering practice. One sentence of 'with N people' in Situation, then I-verbs. If you truly only assisted, pick a different story.
conflict without villains
A conflict story scores the disagreement, the evidence, and the decision — not a cartoon of a foolish coworker.
Amazon's Earn Trust / Have Backbone pair and OPM interpersonal competencies look for disagreement under a shared goal. Name the trade-off (latency vs. completeness, Friday ship vs. 200 ms p95). Do not diagnose the other person's character.
Score the trade-off, not the personality.
| move | text | scores |
|---|---|---|
| villain | PM was reckless and would not listen. | 1 on Earn Trust |
| trade-off | PM needed Tuesday's demo; I owned a 200 ms p95 SLO. | 3, needs Result |
| settled | We shipped a feature flag to 5% for 48 h; p95 stayed 140 ms, then 100%. | 4–5 |
Interview trap. Casting the other person as obviously wrong and yourself as the one who overrode them scores against Earn Trust, because the trade-off was the thing to be assessed.
Engineering practice. State their constraint in one sentence they would accept, then your constraint, then the evidence that settled it.
failure without redemption theater
A failure story needs a real miss, a cause you owned, and a change to the system — not a humble-brag that ends in a promotion.
Google SRE's postmortem culture is the production template: timeline, contributing factors, what you will change, and no blame theater. Interview scoring is the same. 'I worked too hard' is not a miss. 'I shipped without a unique index and duplicated 12k orders' is.
Postmortem shape in four lines.
| field | example |
|---|---|
| miss | Duplicate charges: 12,041 extra rows in 47 minutes |
| cause I owned | I skipped the unique (user_id, idempotency_key) because the load test was 'green' |
| 15-minute now | Stop writers, unique index CONCURRENTLY, replay from outbox |
| guardrail | CI check that payments tables have an idempotency unique key |
Interview trap. Substituting a fake weakness that secretly proves dedication supplies no real miss, no cause you owned, and no change to the system.
Engineering practice. Pick a miss with a number, say what you would do in the first 15 minutes today, and name the guardrail you added.
leadership without a title
Leadership stories show you moved other people's work with a decision, a written plan, or a rollback call — not that you had the word Lead in your title.
Amazon's Bias for Action and Dive Deep are title-agnostic. OPM rates the behavior. A mid-level engineer who froze a bad migration at 01:12 and wrote the RFC is leading; a manager who 'aligned stakeholders' with no decision is not.
Title is not the competency.
| title | behavior | leadership? |
|---|---|---|
| Staff, no reports | I published a 3-page RFC; 4 teams changed their freeze date. | yes |
| Eng Manager | I forwarded the Slack thread. | no |
| Senior IC | I called rollback; 6 services reverted in 11 minutes. | yes |
Interview trap. Waiting for a manager title before telling a leadership story discards the decisions, documents, and rollback calls that the competency actually rates.
Engineering practice. Use a story where people changed what they were doing because of a document, a metric, or a call you made.
disagreement with a manager
Disagree-and-commit stories need the data you brought, the decision of record, and whether you executed the call you lost.
Have Backbone; Disagree and Commit is explicit: escalate with evidence, then execute the decision. Interviewers listen for whether you sandbagged after losing. Result is either the better metric you won, or the committed execution after you lost — both are valid.
Two passing endings; one failing one.
| path | next morning | score |
|---|---|---|
| you won | I shipped the slower, safer cutover; p95 held 180 ms. | pass |
| you lost, committed | I executed their Friday ship and I owned the 2-hour rollback playbook. | pass |
| you lost, sandbagged | I delayed the PR until the date slipped. | fail |
Interview trap. Answering that you never disagree with a manager leaves nothing to score, because the competency is escalating with evidence and then executing the decision.
Engineering practice. Name the artifact (graph, RFC, customer ticket IDs). Name the forum. Name what you did the next morning either way.
saying no to scope
A 'no' story scores the trade-off you named, the number that justified it, and the smaller yes you offered.
Scope cuts are product and engineering competency. 'We cannot do everything' is not a story. 'I cut the PDF export to keep the 200 ms p95, and I scheduled export for Q4 with a 3-day spike estimate' is.
No without a substitute is a stall; no with a 3-day spike is a decision.
| no | substitute | number |
|---|---|---|
| We cannot do PDF. | none | fail |
| Not in this sprint. | 3-day spike week of 12 Oct; p95 stays 200 ms. | pass |
| Engineering is blocked. | none | fail |
Interview trap. Treating 'no' as negative and accepting all scope hides the trade-off, the number that justified it, and the smaller yes that was the actual decision.
Engineering practice. Carry the cost of the cut (days, SLO, dollars) and the substitute. Offer a date, not a vibe.
delivering bad news
Bad-news stories score earliness, the fact you brought, and the option set — not optimism.
SRE and incident command both require early, specific status. 'We might slip' on Thursday after a Monday miss is late. 'As of 11:00 the dual-write is 62% caught up; I need a 24-hour slip or a feature flag' is the Action.
Hours late is the miss, not the tone of the email.
| clock | you knew | you said | miss |
|---|---|---|---|
| Mon 11:00 | dual-write 62% | two options, 24 h or flag | none |
| Thu 16:00 | same fact since Monday | 'looking good overall' | 3 days of silence |
| Fri 09:00 | launch is dead | surprise Slack | integrity |
Interview trap. Holding bad news until the fix is complete converts a 24-hour slip into days of silence, which is the miss the story gets scored on.
Engineering practice. Timestamp the moment you knew. Say what you knew. Give two options with costs. Do not wait for a perfect plan.
influence without authority
Influence stories show a person you did not manage changing a plan because of your evidence or artifact.
No-title leadership is a written artifact plus a metric. A Slack opinion is not influence. A 2-page design with a 12-hour canary plan that another team's EM adopted is.
Track the other team's diff, not your charisma.
| you did | they changed | influence? |
|---|---|---|
| 14 Slack messages | nothing | no |
| RFC + 12 h canary graph | they delayed freeze 48 h | yes |
| Complained in standup | they assigned you a ticket | no |
Interview trap. Equating influence with being liked, and dismissing artifacts as bureaucracy, leaves no evidence that anyone outside your reporting line changed course.
Engineering practice. Name who changed course, what they changed, and which graph or doc they cited.
mentoring versus doing the work
A mentoring story shows the other person's independent result after you stepped back — not that you pair-programmed the whole change.
The competency is developing others. If you still owned every commit, it is an Action story about you. Result is 'they shipped the next retry budget without me, p95 210 ms' or a promotion packet you wrote.
Step-back is the Action; their metric is the Result.
| Action | Result | mentoring? |
|---|---|---|
| I wrote the migration. | it shipped | no |
| I sat in their design review and asked for the abort metric. | they shipped v2 alone; abort fired once, no customer impact | yes |
| I rewrote their PR at 01:00. | green CI | no |
Interview trap. Doing the hard part so the other person cannot fail turns a mentoring story into an Action story about you, with no independent result to point at.
Engineering practice. Pick a story that ends with their commit, their page, or their design review — after a specific coaching move.
incident ownership and blamelessness
Incident stories separate command (you drove mitigation) from blame (you do not score a person as the root cause).
SRE postmortems list contributing factors in systems: missing unique keys, a 0-second retry, a pager that went to a laptop on mute. Your Action is the mitigation timeline. Your Result is TTI, TTM, and the guardrail. 'Bob is careless' is not a factor.
Blameless still has owners of follow-ups.
| heading | example |
|---|---|
| detection | 00:07 pager, error budget 8% → 40% in 11 min |
| mitigation | I flipped the flag at 00:14; 5xx 6.2% → 0.4% |
| factor | no unique idempotency key; retries at 0 ms |
| follow-up | I own the unique index by Fri; SRE owns retry jitter |
Interview trap. Naming the teammate who caused the incident to demonstrate standards replaces the contributing factors with a person, which is what postmortem culture exists to prevent.
Engineering practice. Use the postmortem headings: detection, mitigation, contributing factors, follow-up owners and dates.
trade-off stories with numbers
A trade-off story names two real options, a number for each, and the option you picked — not 'it depends' as the ending.
Interviewers use behavioral questions to hear judgment. 'It depends' is the start of Task, not Result. Consistency vs. latency: 50 ms extra vs. stale reads for 800 ms. You picked one, you measured, you reserved the right to reverse.
A 2×2 you can say in 40 seconds.
| option | user-visible cost | risk | pick? |
|---|---|---|---|
| sync write to primary | +50 ms p95 | none on stale | yes, checkout |
| async replica | 0 ms extra | 800 ms stale window | no for payments |
| flip if | fraud review can tolerate 800 ms stale | — | then replica |
Interview trap. Listing every possible trade-off and stopping leaves 'it depends' as the ending, which is the start of Task rather than a Result.
Engineering practice. Put two columns on the whiteboard (or in your draft): option, cost, risk, what you chose, what would flip it.
time-boxing a four-minute answer
Use the interview's stated format; four minutes is a rehearsal target, not a universal rule.
OPM describes four to six competencies, not minutes per answer. In a 45-minute practice plan, reserve 10 minutes for opening and closing and 10 for probes: 25 / 5 = 5 minutes per prompt. A nine-minute Situation consumes nearly twice that allowance before Result.
A 45-minute loop with 5 questions cannot absorb a 12-minute story.
| item | minutes | leftover for probes |
|---|---|---|
| 5 × 4 min answers | 20 | 25 |
| 1 × 12 min + 4 × 4 | 28 | 17, last question dropped |
| 5 × 2 min, no Result | 10 | 35 of unscored chat |
Interview trap. Treating length as a seniority signal spends a 45-minute loop on three questions, so the panel never hears the Results it is required to rate.
Engineering practice. Rehearse with a 4-minute timer. If you hit 3:00 without a Result sentence, skip to the number.
the follow-up probe
Probes are part of the structured method: they gather missing STAR letters, they are not a chance to restart the story.
Campion's structural elements include consistent probes. 'What did you do next?' and 'How did you know it worked?' are expected. Restating Situation wastes them. Answer the missing letter in one sentence.
Probe → letter → one sentence.
| probe | letter | one sentence |
|---|---|---|
| Who else was involved? | S | 3 backend, 1 SRE; I owned the dual-write. |
| What did you do next? | A | I added the unique index CONCURRENTLY, then I cut the old path. |
| How did you know? | R | Duplicate inserts went from 12,041 to 0 over the next 47 minutes. |
Interview trap. Reading a probe as rejection and switching to a more impressive story mid-answer abandons the letter they asked for and restarts the Situation.
Engineering practice. Map common probes to letters: 'who else' → Situation, 'what did you do' → Action, 'how measured' → Result. Answer that letter only.
mapping a story to a leadership principle
When a company publishes principles, pick the story that supplies evidence for that principle — do not recast every anecdote as Customer Obsession.
Amazon's loop assigns principles to interviewers. Dive Deep wants the metric you pulled. Bias for Action wants the 01:12 call. Inventing a mapping ('this is also Hire and Develop') dilutes the score. One story, one principle, maybe a second if asked.
Primary mapping; stretching loses the bar.
| story | primary | stretch that fails |
|---|---|---|
| 01:12 rollback | Bias for Action | also 'Hire and Develop the Best' |
| unique index after 12k dupes | Insist on the Highest Standards | also 'Think Big' |
| 3-page RFC adopted by 4 teams | Earn Trust | also 'Frugality' with no cost number |
Interview trap. Naming every principle in every answer dilutes the one this interviewer was assigned to score, because each is listening for specific evidence.
Engineering practice. Build a story bank with a primary principle column. If they ask a different principle, pick a different story rather than stretching.
rehearsed scripts versus reusable facts
Rehearse facts, figures, and the four STAR sentences — not a memorized paragraph that collapses on the first probe.
Structured interviews change the probe. A scripted 400-word block cannot answer 'what would you do differently'. A fact card (numbers, names of systems, the decision) can be reassembled. Over-rehearsal is audible: identical cadence, no pause when asked a new letter.
Fact card vs. script under a probe.
| store | 'what would you do differently?' |
|---|---|
| 400-word script | repeats the script |
| fact card: 12,041 dupes, missing unique key, 47 min | I would have required the unique key in the RFC checklist; I added that check the next day. |
Interview trap. Memorizing a 400-word block word-for-word produces an answer that cannot absorb 'what would you do differently', and the identical cadence is audible.
Engineering practice. Drill with a friend who is allowed only to ask for a missing letter. If you cannot answer off-script, you memorized prose, not the event.
protected-topic and pre-offer questions
U.S. federal rules distinguish pre-offer disability inquiries, family questions, and work-authorization questions; they are not one legal category.
EEOC generally bars pre-offer disability questions and says non-job-related marital or children questions may evidence discrimination. DOJ says employers may ask whether an applicant has the legal right to work in the United States and needs sponsorship. Jurisdictions differ; this is not legal advice.
Different legal categories call for different bounded responses.
| question | bounded response |
|---|---|
| Are you planning children? | I will stay on the role: here is a 14-day on-call rotation I ran. |
| What medications are you on? | That is medical; here is how I handled a 47-minute incident instead. |
| Are you authorized to work in the U.S., and will you need sponsorship? | Answer the authorization and sponsorship questions accurately. |
Interview trap. Treating every sensitive question as either harmless small talk or automatically illegal erases the pre-offer, job-relatedness, jurisdiction, and work-authorization distinctions that determine the right response.
Engineering practice. Prepare a calm redirect for non-job-related personal or medical questions, answer a properly framed work-authorization or sponsorship question, and seek qualified local advice when the distinction affects a real hiring decision.
calibrating seniority in the story
The same STAR letters scale by blast radius: mid owns a service, senior owns a cross-team cutover, staff owns a freeze and a written standard.
Panels hire for a level. A mid-level Action that is 'I refactored a file' under-indexes a staff loop. A staff Action that is 'I wrote a unique index' without changing other teams' plans under-indexes too. Match the story's radius to the level on the req.
Radius, not vocabulary, is the level signal.
| level | radius | Result that fits |
|---|---|---|
| mid | 1 service | p95 400 ms → 180 ms |
| senior | 3 services, 1 dependency | dual-write 7 days, then cut |
| staff | org freeze, 11 teams | RFC + abort metric adopted as the standard |
Interview trap. Telling the most technically intricate bug regardless of the level on the req mismatches the blast radius the panel is hiring for.
Engineering practice. Keep three radii in the bank. Pick the one that matches the job. Do not inflate; do not hide a staff-sized call in a mid loop if that is what they asked.
collecting a story bank
A story bank is a dated set of 8–12 real events with numbers, not 50 slogans; you draw from it under the principle they asked.
You cannot invent a 12,041-row incident in the lobby. Keep a private list: date, system, S/T/A/R one-liners, principle tags, a link to the postmortem or dashboard. Refresh quarterly. Two stories per common prompt (conflict, failure, no, incident, influence) covers a 5-question loop with a spare.
Eight rows cover a 5-question loop; one hero project does not.
| tag | stories in bank | if they ask twice |
|---|---|---|
| conflict | 2 | second is the manager-disagree |
| failure | 2 | one prod, one process |
| no-to-scope | 1 | — |
| incident | 2 | command vs. contributing factor |
| influence | 1 | RFC adopted |
| hero-project-only | 1 stretched 5 ways | probes break it |
Interview trap. Stretching one hero project across every prompt breaks on the first probe, because the bank needs separate dated events with their own numbers.
Engineering practice. Write the bank the week after the incident, while the numbers still exist. Never wait until the night before the onsite.
