Skip to content
Tech Interview Prep home
Technical interview guide

LLM Fundamentals

The core mechanics of how LLMs work and the practical parameters engineers actually tune: attention, tokenization, sampling, structured output, and cost/latency trade-offs.

Read
45 min
Practice MCQs
25
Interview QA
25
Edition
v4
Editorial status
Reviewed

Scope: Transformer, BPE, SentencePiece, GPT-3, scaling laws, Chinchilla, nucleus sampling, RoPE, FlashAttention, GQA, PagedAttention, speculative decoding, GPTQ, Model Cards, and NIST GenAI references reviewed 2026-09-06.

Overview

Curated: · Written: · Reviewed:

Understand the stack behind every generated token

Review status: rewritten after review — coverage, worked artefacts and interview framing were below bar. Quality score: pending re-review.

An interviewer asks "how does an LLM work?" and most candidates recite a vocabulary list: tokens, attention, temperature. The answer that gets staff-level marks treats the model as one probabilistic component inside a deterministic product boundary, and can trace a request from text to tokens to logits to a sampled token to a validated, billed API response. This guide walks that path and flags, at each step, what interviewers probe and what a weak answer sounds like.

The mental model: autoregressive next-token prediction

A decoder-only LLM is a function that takes a sequence of token IDs and returns a probability distribution over the vocabulary for the next token. Generation is a loop: sample a token, append it, re-run, repeat until a stop condition. Everything else in the product — chat templates, tools, retrieval, caching, safety filters — is scaffolding around that loop.

The model has no notion of "truth" at inference time. It has learned statistical structure from pretraining and been shaped by instruction and preference tuning. It is not a database, not a proof engine, not an authorization service, and its fluent explanations of its own outputs are generated text, not a faithful log of internal computation.

Interviewer probe: "Is the model reasoning or pattern-matching?" There is no clean answer, and saying either with confidence is a weak answer. A strong answer separates what you can claim (the output distribution is conditioned on the full context; chain-of-thought-style outputs measurably improve some task accuracy) from what you can't (that internal states implement a specific algorithm).

Tokenization: the economic and behavioural boundary layer

Text becomes integer IDs from a fixed vocabulary (typically 50k–250k entries) via subword methods like BPE or SentencePiece. Frequent strings get short tokens; rare strings shatter into many. This is where cost, latency and context limits are actually denominated — in tokens, not words.

A rough anchor for English prose: ~0.75 words per token, so 4,000 words ≈ 5,300 tokens. But the ratio collapses on other inputs. Worked comparison, same 30-character string:

inputapprox. tokens (BPE-style, English-centric vocab)why
"The quick brown fox jumps"6common words, one token each
"def quicksort(arr): return arr"~10keywords are single tokens, punctuation splits
"快速排序算法"~6–10CJK characters often one token each or byte-fallback
"🚀🎉✨"~9–12emoji fall back to byte tokens, 3–4 each

The exact numbers vary by tokenizer; the point a candidate can defend is the shape: code and non-English text cost 1.5–4× more tokens per unit of meaning, and adversarial strings (zero-width joiners, homoglyphs) can inflate or fragment in ways that break naive length checks.

Decision criteria: always count with the exact tokenizer and chat template for the selected model. How you do that depends on the provider: Anthropic's Messages API exposes a count_tokens endpoint, while OpenAI's API does not expose a token-counting endpoint — you count client-side with tiktoken against the model's encoding. Either way, test multilingual and adversarial inputs rather than estimating from English word counts. Preserve special-token IDs; leaking them into user text can alter or break the chat template.

Weak answer: "a token is about four characters." Strong answer: names the tokenizer, the failure modes on code/CJK/emoji, and that the tokenizer — not the model — sets your unit of cost.

Attention and the transformer block: what it costs and why

Each token ID maps to an embedding vector. Each transformer block computes self-attention: every position emits a query, key and value; each query is compared against all keys (dot product, scaled), the scores become weights via softmax, and the weighted sum of values is the position's new representation. Multiple heads run this in parallel over lower-dimensional slices, letting different heads track different relationships. Feed-forward networks, residual connections and normalization wrap around this. Attention weights are internal routing values — useful for analysis, not a causal explanation of the output.

Why it's O(n²) in sequence length: every one of n positions attends to every other position, so the score matrix is n×n. Concretely: doubling a 4k-token context to 8k makes the attention score computation 4× larger. This quadratic term is a large part of why long context is expensive and why prefill latency grows super-linearly with input length.

Causal masking: during training and decoding, position i may only attend to positions ≤ i. This is what makes one forward pass usable for next-token training at every position simultaneously, and what stops the decoder from reading the future at inference.

Positional information: attention is permutation-invariant over content — swap two tokens and the value mixing changes but nothing enforces order. Positional mechanisms (learned absolute embeddings, sinusoidal encodings, or rotary embeddings/RoPE, which rotates query/key pairs by position-dependent angles) inject order. RoPE's advantage is that relative displacement is expressed directly in the q·k dot product, which is why most current open models use it and why context-extension work (e.g., adjusting RoPE scaling) is common.

What the KV cache changes: decoding without a cache recomputes keys and values for the whole prefix every step — O(n²) work per token generated. The cache stores each position's keys and values once, so each new token only computes its own q/k/v and attends against the cached prefix: O(n) work per step. The trade is memory. Cache size for one sequence:

KV bytes = 2 (K and V) × n_layers × n_kv_heads × head_dim × seq_len × bytes_per_element

Worked figure, GPT-3-class geometry (96 layers, 96 heads, head_dim 128, fp16 = 2 bytes), full multi-head cache, 4k-token context:

2 × 96 × 96 × 128 × 4096 × 2 = 19.3 GB per sequence

That is why serving stacks use grouped-query attention (fewer KV heads — e.g., 8 KV heads instead of 96 cuts the above to ~1.6 GB), quantized caches, and PagedAttention-style paging to reduce fragmentation. If an interviewer asks "why is long context expensive," the memory arithmetic above is the answer that separates levels.

Pretraining, tuning, and choosing the right layer to change behavior

Pretraining minimizes next-token prediction loss over a large corpus; scale, data mixture, deduplication and training-token count set the capability ceiling. Instruction tuning and preference tuning (e.g., RLHF-style alignment) shape how the base model follows requests — same weights' knowledge, different behavior.

When you need to change what the system does, pick the least complex layer that satisfies the requirement:

layerchangeslatency costreversibilityuse when
promptingcurrent context onlyper-request tokensinstantbehavior shaping, format, examples
retrievalexternal evidence in contextindex + added tokensredeploy indexfacts that change, need citations
toolsauthoritative computation/actionsround-tripredeploy toolarithmetic, lookups, side effects
fine-tuningmodel weightstraining + eval cyclere-versionstable style/format a prompt can't reach

The rule underneath the table: keep anything that changes facts or permissions outside the weights. Weights are frozen at inference; the world isn't.

Weak answer: "we'll fine-tune it to be accurate." Strong answer: asks whether the failure is knowledge (retrieval), formatting (prompt or schema-constrained decoding), or behavior (tuning), and names the eval that would tell them apart.

Sampling and generation control

Each decoding step produces logits over the vocabulary. Controls reshape that distribution before sampling:

  • Greedy / T=0: argmax. Deterministic per model revision, but not a mathematical truth setting — it picks the most probable token, which can still be wrong.
  • Temperature: divides logits before softmax. T<1 sharpens (low-probability tokens vanish), T>1 flattens (tail tokens become reachable). T=1 samples the raw distribution.
  • Top-k: keep only the k highest-probability tokens. Fixed count, so the effective support varies with distribution shape.
  • Top-p (nucleus): keep the smallest set whose cumulative probability ≥ p. Variable count — a confident distribution keeps 1–2 tokens, a flat one keeps hundreds. Generally preferable to top-k for the same reason.
  • Frequency/presence penalties, stop sequences, max tokens: reshape repetition, terminate generation, and cap output length.

Worked trace — same model, same prompt ("What is 2 + 2?"), three samples each, hypothetical but representative:

decodings1s2s3what changed
greedy (T=0)444nothing — argmax
T=0.7, top-p=0.944"four"wording varies, arithmetic doesn't
T=1.545a jokethe distribution widened; facts didn't move

Temperature widens the next-token set; it does not make content more true. Decision rule: pin decoding (T=0 or low, top-p low) for extraction and classification; sample only where you will validate or where diversity is the point.

Seeds and reproducibility: a seed fixes the pseudo-random sampling stream, nothing else. Provider-side batching, floating-point non-associativity, routing, model revision changes, tool results and retrieval updates all still vary outputs. Deterministic product guarantees come from schema validators, executable checks, authoritative data and idempotent state machines — not from a seed. When testing stochastic behavior, use repeated trials and compare distributions; when testing exact contracts, make downstream acceptance deterministic.

Interviewer probe: "your eval score dropped 2% after a redeploy — how do you know it's the model?" Weak answer: re-run and compare averages. Strong answer: pin the model revision, fix decoding, replay a fixed input set, and separate sampling noise (repeated trials) from real drift (pinned-config deltas).

Context window: finite, and not uniformly used

The window counts everything: system/developer/user messages, tool definitions and results, retrieved passages, examples, hidden template tokens, and generated output, per the provider contract. A 128k window does not mean 128k of equally-attended evidence — long-context performance degrades with position and distractor density; important evidence can be truncated, diluted or ignored (the "lost in the middle" family of effects).

The window is also the memory bound for agents: an agent that accumulates tool results and observations will hit it, so the design question is what to summarize, what to drop, and what to re-retrieve. Budget explicitly:

budget = window − system/tools − reserved output − safety margin
retrieval + history must fit inside budget, ranked by relevance, with citations kept

Treat context content as untrusted data unless the application explicitly grants it authority — a retrieved web page saying "ignore previous instructions" is data, not instructions, and your architecture should enforce that distinction.

Structured output and tool calling

Schema-constrained decoding (JSON mode, grammar-constrained sampling) forces the output to parse — the sampler is restricted to tokens that keep the string valid against the schema. What it guarantees is syntax. What it never guarantees is semantics: a schema can force "total": 4 to be a valid integer while the correct total is 5.

Tool calls are represented in the message stream: the model emits a structured request (name + arguments) instead of, or alongside, text; your runtime executes it, appends the result as a tool-role message, and re-invokes the model. The model proposes; your code disposes.

Failure modes and the standard defenses:

failuredefense
invalid JSON (unconstrained mode)schema-constrained decoding, or parse-retry with the error fed back (cap retries, e.g., 2)
schema-valid but wrong valueserver-side validation: enums, ranges, cross-field invariants
fabricated tool argumentsvalidate against the tool's real signature and the user's permissions
duplicate side effects on retryidempotency keys on material operations
unverified postconditioncheck the authoritative state after the call, don't trust the model's narration

Minimal server-side check, Python 3.12, illustrative:

from pydantic import BaseModel, field_validator

class Refund(BaseModel):
    order_id: str
    amount_usd: float
    reason: str

    @field_validator("amount_usd")
    @classmethod
    def positive(cls, v: float) -> float:
        if v <= 0:
            raise ValueError("amount must be positive")
        return v

# After parsing: still enforce authorization and the authoritative
# order total server-side — the model never decides permissions or money.

Bind identity and permissions outside the model, require confirmation for material effects, and verify postconditions against the authoritative system. Weak answer: "we use JSON mode so it's reliable." Strong answer: names what JSON mode does and does not guarantee, and where the semantic checks live.

Hallucination: why, and what actually reduces it

Hallucination is confident-looking output unsupported by truth or evidence. It is not a bug in the sense of a broken component — next-token prediction rewards plausible continuations, and plausibility is only correlated with truth. Tuning reduces it; nothing in the architecture eliminates it.

Reduction, in order of leverage:

  1. Scope the task — closed questions with defined output beat open-ended generation.
  2. Ground in current retrieval — the model can't know post-cutoff facts; retrieval supplies them, with citations.
  3. Constrain decoding — schemas, enums and constrained decoding remove whole classes of fabrication (the model cannot invent a status code outside your enum).
  4. Use authoritative tools — arithmetic, lookups and side effects go to code, not generation.
  5. Design abstention — explicitly permit "I don't know / no supporting passage found," and make the no-answer path a first-class, tested behavior.
  6. Check claims, not documents — claim-level citation verification and calibrated human review on the residual.

Detection is empirical: run grounded-vs-ungrounded eval sets, measure unsupported-claim rates per slice, and monitor refusals and truncation in production. Model confidence language ("I'm certain that…") is not calibrated probability — never treat it as a signal.

Talking to a customer honestly: name the mechanism (plausible continuation, not verified retrieval), state what your system does about it (grounding, constrained output, abstention, human review on high-stakes paths), and never place access control, legal approval, financial arithmetic or safety-critical verification solely in free-form generation. A weak answer promises accuracy; a strong answer prices the residual risk and shows where it's caught.

Serving: where latency and cost actually come from

Two phases per request. Prefill processes all input tokens in parallel — its cost drives time-to-first-token (TTFT), and it grows with input length. Decoding emits tokens one at a time, each step attending over the KV cache — inter-token latency and total duration live here, along with contention from other tenants.

Hypothetical but representative figures for a mid-size hosted model: TTFT ~200–500 ms for a 1k-token prompt; decode ~20–50 ms/token, so a 500-token response adds 10–25 s of decode time. The levers follow directly: shrink input (retrieval over stuffing), cap max tokens, cache (prompt caching reuses prefill for shared prefixes), and stream so users see the first token at TTFT rather than at the end.

Serving optimizations you should be able to name with their trade-offs:

  • Continuous batching — inserts new requests into a running batch instead of waiting for the whole batch to finish; improves utilization, adds scheduling variance to tail latency.
  • PagedAttention — pages the KV cache like virtual memory; cuts fragmentation, adds bookkeeping.
  • Grouped-query attention — fewer KV heads, smaller cache, negligible quality loss in practice; this is the arithmetic from the attention section.
  • FlashAttention — exact attention computed in tiled fashion with far less memory traffic; same math, faster kernel.
  • Quantization (e.g., 8-bit or 4-bit weights) — less memory and bandwidth, potential accuracy and hardware-support caveats; benchmark on your tasks.
  • Speculative decoding — a small draft model proposes several tokens, the target verifies them in one pass; preserves the target's distribution under its assumptions, trades compute for decode latency.

Each needs quality, compatibility, latency-percentile and failure tests on the exact packaged runtime — a benchmark on a different quantization or kernel than you ship is not evidence.

Model selection, observability, and the product boundary

Model choice is a system decision, not a leaderboard lookup. Compare on representative and adversarial tasks from your domain: capability, context behavior at your real document lengths, tool/schema reliability, supported languages, latency and throughput percentiles, availability, privacy and retention terms, deployment region, safety controls, version lifecycle, and cost per successful task (a cheaper model that fails 20% of requests is often the expensive one). Route by task: a small model for classification, a large one for hard synthesis. Pin revisions where the provider allows it, and keep a tested fallback whose behavior you have actually validated — not one that is silently different.

Production requires an immutable trace per request: application version, model revision, prompt and template, decoding parameters, retrieval index version, tool versions, policy version, output validation results, latency, token counts, cost, and final state — with data minimization and access controls on the trace itself. Monitor errors, truncation, invalid schemas, refusals, safety events and user-outcome metrics by slice. Roll out with shadow traffic or sticky canaries and atomic rollback.

The abstraction to carry into an interview: not "ask a smart model," but a versioned probabilistic component inside a deterministic, observable product boundary — every place the model is uncertain, something deterministic outside it validates, decides, or catches the failure.

Likely follow-ups and the answers that hold up

  • "Why is long context expensive?" — O(n²) attention scores in prefill plus linear-in-n KV cache memory per sequence; give the 19.3 GB arithmetic.
  • "Your JSON outputs keep failing validation — walk me through fixing it." — constrained decoding first, then server-side semantic validation, then retry-with-error with a cap; distinguish syntax failures from semantic ones.
  • "How do you make the model deterministic?" — you don't; you make the contract deterministic: pinned revision, pinned decoding, validators, idempotent state machines.
  • "When would you fine-tune vs. prompt vs. retrieve?" — the layer table: facts outside weights, behavior in prompts, stable format in tuning; name the eval that decides.
  • "A customer asks if the product hallucinates." — yes, by mechanism; here is the grounding, abstention and review path, and here is the residual risk on their use case.

A weak answer recites definitions. A strong answer traces one request end to end, states the numbers, and knows exactly which component owns each guarantee.