Skip to content
Tech Interview Prep home
Technical interview guide

LLM Evaluation

Measuring whether an LLM-based system actually works, given that outputs are open-ended and hard to score automatically.

Read
45 min
Practice MCQs
25
Interview QA
25
Edition
v4
Editorial status
Reviewed

Scope: BLEU, ROUGE, BERTScore, HELM, BIG-bench, TruthfulQA, MT-Bench, G-Eval, RAGAS, Model Cards, NIST AI RMF and GenAI Profile, and OWASP LLM references reviewed 2026-09-06.

Interview QA

Treat each question like a live interview question: answer out loud first (structure, assumptions, tradeoffs), then open the model answer to spot gaps and rehearse a tighter follow-up.

Curated: · Written: · Reviewed:

QA-1

Design an evaluation contract for an LLM feature.

QA-2

Create a leakage-resistant LLM evaluation dataset.

QA-3

Build deterministic checks for structured LLM output.

QA-4

Choose metrics for an open-ended summarization feature.

QA-5

Calibrate an LLM-as-judge evaluator.

QA-6

Harden an automated evaluator against grader manipulation.

QA-7

Operate a rigorous human evaluation study.

QA-8

Design a multi-dimensional LLM release scorecard.

QA-9

Evaluate a RAG pipeline by component and end to end.

QA-10

Build a claim-level groundedness evaluation.

QA-11

Design an end-to-end evaluation for a tool-using agent.

QA-12

Create an adversarial safety evaluation suite.

QA-13

Balance refusal and usefulness evaluation.

QA-14

Analyze a paired LLM evaluation statistically.

QA-15

Plan slices and uncertainty for clustered production samples.

QA-16

Mitigate benchmark contamination and evaluation gaming.

QA-17

How do you curate, version, and maintain evaluation golden datasets to prevent drift and leakage while tracking ground truth changes over time?

QA-18

Make an evaluation harness reliable under provider failures.

QA-19

Connect offline evaluation to an online canary.

QA-20

How do you detect and measure eval drift or alignment divergence between offline LLM-as-a-judge scores and online human feedback in production?

QA-21

Evaluate LLM performance and cost for release.

QA-22

Design a responsible fairness evaluation for an LLM system.

QA-23

Evaluate and release a prompt change.

QA-24

Run incident response for an LLM evaluation escape.

QA-25

Review an LLM evaluation program end to end.