Skip to content
Tech Interview Prep home
Technical interview guide

LLM Fundamentals

The core mechanics of how LLMs work and the practical parameters engineers actually tune: attention, tokenization, sampling, structured output, and cost/latency trade-offs.

Read
45 min
Practice MCQs
25
Interview QA
25
Edition
v4
Editorial status
Reviewed

Scope: Transformer, BPE, SentencePiece, GPT-3, scaling laws, Chinchilla, nucleus sampling, RoPE, FlashAttention, GQA, PagedAttention, speculative decoding, GPTQ, Model Cards, and NIST GenAI references reviewed 2026-09-06.

Interview QA

Treat each question like a live interview question: answer out loud first (structure, assumptions, tradeoffs), then open the model answer to spot gaps and rehearse a tighter follow-up.

Curated: · Written: · Reviewed:

QA-1

Explain an autoregressive LLM request end to end.

QA-2

Design token-budget handling for multilingual input.

QA-3

Compare content embeddings and positional information.

QA-4

Explain self-attention and causal masking to an interviewer.

QA-5

Choose an intervention for a new LLM product requirement.

QA-6

Compare greedy, temperature, top-k, and top-p decoding.

QA-7

Prevent truncation from corrupting downstream workflows.

QA-8

Design reproducible tests for a stochastic model API.

QA-9

Build a safe structured-output boundary.

QA-10

Select decoding controls for a constrained extraction task.

QA-11

Design a secure tool-calling execution loop.

QA-12

Design hallucination controls for a high-stakes assistant.

QA-13

Engineer context for a long-document workflow.

QA-14

Reason about KV-cache capacity in an LLM service.

QA-15

Evaluate adopting an optimized attention kernel.

QA-16

Compare multi-head, multi-query, and grouped-query attention.

QA-17

Validate a quantized model for production.

QA-18

Assess speculative decoding for an API workload.

QA-19

How do Time to First Token (TTFT) and Inter-Token Latency (ITL) trade off against each other in LLM serving systems?

QA-20

Operate continuous batching under variable sequence lengths.

QA-21

How do Chinchilla and Kaplan scaling laws differ, and how do they determine the optimal balance between parameters and training tokens?

QA-22

Design privacy-aware LLM request telemetry.

QA-23

Design safe LLM provider failover.

QA-24

Specify an LLM request trace and dashboards.

QA-25

Review an LLM-backed system end to end.