Technical interview guide
LLM Evaluation
Measuring whether an LLM-based system actually works, given that outputs are open-ended and hard to score automatically.
- Read
- 45 min
- Practice MCQs
- 25
- Interview QA
- 25
- Edition
- v4
- Editorial status
- Reviewed
Scope: BLEU, ROUGE, BERTScore, HELM, BIG-bench, TruthfulQA, MT-Bench, G-Eval, RAGAS, Model Cards, NIST AI RMF and GenAI Profile, and OWASP LLM references reviewed 2026-09-06.
Curated: · Written: · Reviewed:
