Skip to content
Tech Interview Prep home
Technical interview guide

Fine-Tuning & Adaptation

Adapting a pretrained LLM's behavior via further training, and how that differs from prompting or RAG.

Read
45 min
Practice MCQs
25
Interview QA
25
Edition
v4
Editorial status
Reviewed

Scope: Transformer, InstructGPT, LoRA, QLoRA, DPO, LIMA, Model Cards, Datasheets, Hugging Face PEFT/Transformers/TRL, NIST AI RMF, and OWASP supply-chain references reviewed 2026-09-06.

Overview

Curated: · Written: · Reviewed:

Train only after defining the behavior and its evidence

Fine-tuning updates model behavior using examples or preference data. It can improve task format, domain terminology, instruction following, classification, extraction, style or policy behavior. It is usually a poor mechanism for rapidly changing factual knowledge, exact authorization, guaranteed truth or live computation. Compare prompt and workflow changes, retrieval, tools, deterministic code and a stronger base model before accepting training cost and lifecycle complexity.

Write a behavior contract first. Define intended tasks and users, output schema and language, acceptable variation, refusal and escalation, latency/cost, safety and privacy boundaries, and concrete success/failure metrics. Establish a frozen base-model baseline with the production prompt and decoding configuration. Without a baseline, training can look impressive while a prompt fix would have been cheaper and safer.

Dataset governance determines the product. Record source, owner, consent/license, collection purpose, language/domain, timestamp, sensitivity, annotator process, transformations, quality checks, permitted use, retention/deletion, and lineage to each training example. Remove secrets, unnecessary personal data, copyrighted or restricted content under policy, prompt injections, malware and poisoned examples. Deduplicate exact and semantic copies. A dataset row is not safe merely because it came from production logs.

Define the supervision target precisely. Instruction tuning examples should match the production conversation or completion template, roles, tool schemas, separators, special tokens, maximum length and loss mask. Decide whether loss applies to assistant output only or other tokens. Preserve system-policy boundaries rather than training user text as privileged instruction. Weighting and sampling determine which languages, tasks, safety cases and edge conditions dominate optimization.

Splits must prevent leakage. Group related documents, users, templates, conversations, generated variants and near duplicates before train/validation/test division. Use time- or source-based holdouts where deployment faces future or unseen distributions. Keep the final test set untouched by hyperparameter and prompt selection. Scan base-model contamination where feasible and report uncertainty. A random row split can put paraphrases of the same answer in every partition.

Supervised fine-tuning minimizes token-level loss on demonstrations. Low training loss does not establish useful behavior. Monitor validation loss plus task metrics, exact schema validity, factuality, refusal, safety, calibration, diversity and general capability. Learning rate, schedule, warmup, batch and accumulation, sequence length, optimizer, precision, gradient clipping, epochs and packing affect stability. Token-weighted averages can hide poor performance on short or minority examples.

Parameter-efficient fine-tuning updates a small adapter rather than all base weights. LoRA injects trainable low-rank updates into selected modules; rank, alpha, dropout and target modules affect capacity. QLoRA keeps a quantized base for memory efficiency while training adapters, but quantization, compute dtype, kernels and merging/loading remain compatibility concerns. Smaller trainable state reduces resource and storage needs; it does not remove data, evaluation or security requirements.

Full fine-tuning offers capacity but costs more memory, compute and artifact storage and can alter broad capabilities. Continual or narrow tuning can cause catastrophic forgetting or behavior drift. Mix replay/general examples where justified, regularize, reduce learning rate/epochs, or use adapters for separable tasks, then measure retained capabilities. Multiple adapters create routing, composition, tenant-isolation and base-version dependencies.

Preference optimization uses comparisons rather than one target answer. Define prompt distribution, chosen/rejected labeling guidance, tie/uncertainty handling, rater quality and conflicts. DPO optimizes preference likelihood relative to a reference policy; its beta and data quality influence movement. RLHF adds reward model and policy optimization complexity. Preference wins can encode verbosity, sycophancy, demographic bias or evaluator shortcuts, so human evaluation must inspect why one answer wins.

Training is a reproducible artifact pipeline. Pin base model and revision, tokenizer/chat template, dataset snapshot and preprocessing code, seeds, framework/dependencies, hardware, distributed strategy, precision, hyperparameters and checkpoint selection rule. Track gradients/loss/resource and data counts, validate checkpoint integrity, and preserve resumable optimizer state according to policy. Sign and scan model/adapters, serialize safely, and never load untrusted executable model code into privileged environments.

Privacy and security need explicit tests. Models may memorize rare strings or personal data; training services and experiment logs can leak examples. Apply minimization, access, encryption, retention, deletion and approved privacy mechanisms. Test canaries and extraction attacks without publishing sensitive targets. Guard against data poisoning, backdoors, malicious adapters/base models, compromised dependencies and weight substitution with provenance, review, checksums/signatures and behavioral red teams.

Evaluate the exact deployed stack: base plus adapter/merged weights, tokenizer/template, system prompt, tools/RAG, decoding and safety filters. Use deterministic tests for schema/tool arguments, calibrated human review for open-ended quality, and adversarial suites for jailbreaks, prompt injection, privacy, bias and harmful content. Report confidence intervals and slices by task, domain, language, length, risk and data provenance. Compare win/loss/tie to baseline and require non-regression on general and safety capabilities.

Deployment must be reversible. Check runtime architecture, quantization, context size, special-token IDs, adapter/base compatibility, memory and latency before rollout. Shadow or canary with privacy controls, pin model version per request, prevent mixed caches, monitor output validity, task success, refusal, safety, drift, latency/cost and user corrections, and retain instant rollback. A merged adapter may be operationally simpler but loses separable switching; document exact lineage either way.

Close the feedback loop carefully. User ratings are selection-biased and can reward fluent wrong answers. Sample failures with consent and minimization, classify root cause as prompt, retrieval, tool, product, data or model, and add targeted examples only when training is the right lever. Version new data, keep holdouts independent, re-run all gates, and preserve the ability to delete or supersede harmful examples. Fine-tuning is an ongoing model-and-data product, not a one-time training command.

Worked example: loss 0.31, schema 61%, canary extracted

Support-bot LoRA. Gate: JSON tool args ≥ 90% and the planted canary SSN 078-05-1120 must not appear.

checkpointtrain losstool-JSON passcanary in 50 greedy samples
base—88%0
epoch 30.9094%0
epoch 120.3161%7

The lowest-loss checkpoint fails both gates. Pick epoch 3 or stop. That table is the interview.