Skip to content
Tech Interview Prep home
Technical interview guide

LLM Observability & Monitoring

Monitoring, logging, and tracing LLM and agent systems in production — the specific signals and tooling that apply on top of general observability practice.

Read
45 min
Practice MCQs
25
Interview QA
25
Edition
v4
Editorial status
Reviewed

Scope: OpenTelemetry specification and semantic conventions reviewed at 1.44.0 on 2026-09-04, W3C Trace Context 2021 Recommendation, OpenMetrics 1.0, Google SRE monitoring, RAGAS, Lost in the Middle, Model Cards, NIST GenAI, and OWASP Prompt Injection/Sensitive Disclosure references.

Overview

Curated: · Written: · Reviewed:

Observe the product outcome and the exact path that produced it

The mental model: a 200 is not a success

LLM observability is ordinary distributed-system telemetry plus behavior evidence for probabilistic components. The reason it deserves its own treatment: the two failure classes split differently than in a deterministic service.

  • Operational failures — latency, timeouts, rate limits, provider outages — look like classic telemetry and classic golden signals catch them.
  • Semantic failures — the wrong answer, an ungrounded claim, a tool called with plausible-but-wrong arguments — return HTTP 200 with a well-formed stream and a confident tone. Nothing in the transport layer knows. Some failures are only visible in the output text.

So the mental model is two planes. The operational plane answers "is the service available, fast, and within budget?" The behavior plane answers "are the outputs correct, grounded, and useful?" A dashboard full of tokens and model latency cannot prove task success; a dashboard full of judge scores cannot tell you the provider is throwing 429s. You need both, and you need the join between them: which exact prompt, model revision, retrieval corpus, and tool version produced this output.

The decisions this system exists to serve: is the service safe and within budget; which change caused a regression; which users or tasks are affected; roll back or escalate? Define service-level indicators from user-visible outcomes (verified task completion, schema-valid finals, refund succeeded), then add component signals that explain them. Never let component vanity metrics substitute for the outcome.

What interviewers probe here: they'll ask "how do you know your LLM feature is working in production?" A weak answer stops at latency and token cost. A strong answer names a user-visible outcome metric, admits it's hard to measure directly, and describes the sampling/judge/implicit-signal machinery that approximates it.

What to capture on a single LLM call

Every model call should be attributable to an exact invocation. The minimum set:

categoryfieldswhy
identityrequest ID, trace/span ID, user/tenant, feature flag, experiment cohortslicing and blast-radius later
modelprovider, model name, resolved version/snapshot, deployment region"gpt-4o" is not a version; providers snapshot and deprecate
inputsprompt template ID + version, rendered prompt hash (not raw text by default), few-shot set IDreproduction without copying proprietary text into every event
decodingtemperature, top_p, max tokens, stop sequences, seed if supportedsame prompt, different decoding, different behavior
outputsresponse, finish reason, token counts in/out/cached/reasoning, cost, schema validitythe outcome and its price
timingqueue time, time to first token (TTFT), inter-token latency, total durationstreaming UX depends on TTFT, not total latency
safetyfilter verdicts, refusals, truncationa truncated "successful" call is a failed task

Token categories depend on provider contracts — cached and reasoning tokens are billed differently per vendor — so normalize cautiously, retain raw provider fields, and reconcile against invoices. Optimize cost per successful task, not cost per call: a cheaper model that doubles retries can cost more per completed refund.

A concrete span, OpenTelemetry-style (Python, OTel SDK conventions):

from opentelemetry.trace import Status, StatusCode

with tracer.start_as_current_span("llm.chat") as span:
    span.set_attribute("gen_ai.operation.name", "chat")
    span.set_attribute("gen_ai.request.model", "claude-sonnet-4-20250514")
    span.set_attribute("gen_ai.request.temperature", 0.2)
    span.set_attribute("llm.prompt_template.id", "refund-assistant")
    span.set_attribute("llm.prompt_template.version", "v19")
    span.set_attribute("llm.user_tier", "free")   # bounded label, not raw user ID
    t0 = time.monotonic()
    resp = client.messages.create(model=..., messages=rendered)
    ttft = time.monotonic() - t0   # with streaming: time to first chunk
    span.set_attribute("gen_ai.usage.input_tokens", resp.usage.input_tokens)
    span.set_attribute("gen_ai.usage.output_tokens", resp.usage.output_tokens)
    span.set_attribute("llm.finish_reason", resp.stop_reason)
    span.set_status(Status(StatusCode.OK))

The point of the snippet is the shape: bounded, versioned, joinable attributes — not a log line with the whole prompt dumped in it.

Tracing multi-step and agent runs

One flat log line per call is insufficient because it can't answer "which step of which task went wrong?" An agent run is a tree: a run/trace ID stitches the whole task; nested spans carry the causal path through the agent loop, each LLM call, each tool invocation, retrieval, and handoffs.

trace 7f3a  task: "process refund request #8841"
└─ span agent.loop (attempt 1)
   ├─ span llm.chat (plan)            820ms  412 in / 96 out
   ├─ span tool.call lookup_order     140ms  ok
   ├─ span llm.chat (decide)         1100ms  1,830 in / 210 out
   ├─ span tool.call issue_refund     310ms  ok   amount=$42.00
   └─ span llm.chat (summarize)       640ms  finish_reason=end_turn
   task.result  verified_refund=true  total 3.0s  $0.011

The span durations sum to 820 + 140 + 1,100 + 310 + 640 = 3,010 ms, matching the 3.0s task total (the spans run sequentially, so the sum is the wall-clock total).

Per tool span: name, arguments (redacted), result, latency, error. Per agent span: graph/node/attempt, routing and delegation, budgets, retries, loops, cancellation, human decisions, and the final authoritative state. Count workflows, not only calls — 10,000 model calls could be 9,000 tasks or 400 stuck ones.

Agent-specific failure modes to detect explicitly:

  • No-progress cycles — the loop calls the same tool with the same arguments three times. Span-level repetition detection catches this; per-call metrics never will.
  • Duplicate effects — the same idempotency key hitting the payment tool twice. Record idempotency keys on side-effecting spans.
  • Budget exhaustion — token or step budget hit before an answer. This is a task failure even though every span returned 200.
  • Orphan nodes and ambiguous commits — two nodes both wrote state; which one is authoritative?

Generated hidden reasoning (chain-of-thought) is neither necessary nor safe as routine telemetry. Store the checkable decisions and evidence instead.

RAG instrumentation: retrieval or synthesis?

The RAG question an operator actually asks is "was the bad answer caused by bad retrieval or bad generation?" You cannot answer it unless the trace records both sides and the join between them.

On the retrieval side: original and derived query IDs, authenticated filters and policy version, candidate source/version IDs with component scores, reranker output, embedding model and index version, selected chunks, context order and token budget, citations, cache hit, and freshness. On the generation side: which chunk was actually used in the prompt, and the groundedness verdict on the output.

span rag.retrieve   query="return policy for opened items?"
                  candidates: [doc#41 score 0.83, doc#17 score 0.79, doc#93 score 0.71]
                  reranker kept: [doc#41, doc#17]   index=v6  embed=bge-large-v1.5
span rag.generate  context=doc#41,doc#17  cited=doc#41  groundedness=0.91 (judge)

If doc#41 was the right document and the answer still contradicts it, the failure is synthesis. If the top candidate was doc#93 (a 2021 policy) and freshness lag shows the 2024 rewrite hadn't been indexed yet, the failure is the source-to-index pipeline. Same bad user answer, different fix — and only the trace distinguishes them.

Monitor: no-result rate, permission denials, source-to-index and deletion lag, retrieval latency, and sampled relevance/faithfulness. Avoid logging full retrieved text by default; store canonical protected record IDs so the context can be reconstructed under access control.

Measuring quality in production

Behavior metrics come from versioned evaluators and audited samples, because you cannot infer "was this answer correct" from latency histograms.

Sampled explicit evaluation. LLM-as-judge scoring for groundedness, relevance, correctness, and safety — sampled, not on every request, because a judge call doubles your model bill and adds its own error rate. Rules of thumb that survive scrutiny:

  • Calibrate the judge against expert-labeled samples before trusting it, and report the agreement rate. An uncalibrated judge is a random number generator with a dashboard.
  • Version the evaluator and preserve its outputs per version; a score drop might mean the judge changed, not the product.
  • Protect evaluators from prompt injection in the content they score.
  • Report coverage and uncertainty: "groundedness 0.87 ± 0.03 on a 2% sample (n=1,140)" is defensible; "groundedness is high" is not.

Implicit signals — free, biased, continuous: task success (tool completed, order placed), retry/regenerate rate, user edits of the output, thumbs up/down, abandonment mid-task, customer-support escalations. A spike in "regenerate" clicks after a prompt deploy is a regression signal hours before any human reviews a sample.

Human review queues — the gold standard, expensive, so targeted: oversample low-confidence judge scores, refusals, safety flags, and new task types.

Do not turn free-form model output into an unvalidated low-cardinality metric. "Top error themes" extracted by a model into a fixed taxonomy is fine if the taxonomy is versioned and the extraction is itself evaluated.

Sampling, cardinality, and the telemetry system itself

Sampling follows risk. Head sampling (decide at request start) is cheap but can't know the final outcome. Tail sampling decides after completion, so it can retain all errors, slow traces, policy events, and rare workflows — at the cost of collector complexity and buffering delay. Always retain required audit events through a separate governed path, and never write sampling rules that erase successful denominators or minority cohorts, or your quality metrics become lies. Record the sampling decision and probability on each trace so analysis can reweight.

Cardinality is a production risk. User IDs, request IDs, prompts, and arbitrary error text do not belong in metric labels. A provider outage that puts raw exception strings into a label can create millions of series and take the monitoring system down during the exact incident you needed it. Use bounded dimensions — model, operation, status, task class, deployment cohort — and connect sampled exemplars to protected traces. Cardinality budgets, drop policies, and collector backpressure protect the system; lost telemetry must itself be counted, because absence of events is not proof of health.

Telemetry is a sensitive data system. Prompts, outputs, retrieved content, tool arguments, and feedback carry personal data, secrets, and proprietary instructions. Define purpose and data classification, minimize and redact before export, tokenize identities, isolate tenants, encrypt, restrict and audit reads, set retention/deletion, and prevent unauthorized training reuse. Redaction itself needs tests — a regex that misses a base64-encoded API key is a breach with a dashboard. Never assume provider or exporter defaults are safe.

Drift, alerts, and incidents

Drift can hit the input distribution, the retrieval corpus, the model/provider (silent snapshot updates), prompt effectiveness, tools, outcomes, or the evaluators themselves. Monitor feature and task mixes, length/language/source slices, score distributions, and failure taxonomies — but require outcome evidence before calling every statistical change harmful. Frozen sentinels (a fixed eval set run on a schedule) catch model-side drift; production samples catch prompt- and data-side drift. Use change-point alerts with minimum-traffic thresholds and burn-rate logic, and record deployments for correlation.

Alerts must be actionable and symptom-oriented. Page for sustained user-impacting availability, critical safety/security events, corrupted authorization, or duplicated external effects. Ticket (don't page) for slower quality, cost, or drift trends. Every alert carries affected scope, current versions, dashboard and trace links, runbook, rollback, and owner. Test alerts with synthetic probes and failure injection; audit silence caused by broken instrumentation.

During an incident: preserve privacy-safe evidence, contain traffic or capabilities, pin or roll back the affected stack, find the earliest divergent span and affected cohort, and classify the failure — product, provider, prompt, retrieval, agent, tool, policy, or telemetry. Add regression coverage and monitoring after root cause. Telemetry changes are releases too: version schemas and semantic conventions, support dual emission, validate collectors, reconcile counts, canary, and roll back without silently losing critical evidence.

Worked example: TTFT is green while refunds fail

Canary prompt v19. Streaming looks healthy. Schema-valid JSON for the refund tool drops.

dashboardTTFT p50schema-valid finalsverified refunds/hrpage?
tokens + TTFT only180 msnot plottedunknownno
finish_reason + cost per successful task180 ms41%12 vs SLO 40yes, quality burn
raw prompt text as a metric label———monitoring dies (cardinality)

Diagnosis path: the outcome metric fired → tail-sampled traces of failed tasks → span diff shows v19 renders a longer system prompt, output hits max_tokens, finish_reason=max_tokens, JSON truncates before the tool argument closes. Fix: raise max_tokens or trim the template. A stream that opened is not a completed task. Interview the denominator, then the labels.

Misconceptions, trade-offs, and what changes at scale

Common misconceptions:

  • "Observability = logging the prompts." Logging everything is the opposite of observability: it's a privacy incident and an unqueryable firehose. Structure and sample instead.
  • "The judge score is ground truth." It's an estimate with its own error rate and version drift. Calibrate it, report uncertainty.
  • "Low latency means good product." TTFT p50 of 180 ms coexists with a 41% task success rate, as above.
  • "Metrics, logs, and traces are interchangeable." They aren't: metrics summarize numeric state over time, logs record discrete structured events, traces connect one request's causal path. Choose by information shape.

When not to build this: a prototype with five users needs a log of failures and a weekly manual review, not tail sampling and judge pipelines. The machinery above earns its cost when volume makes manual review impossible or when a regression can harm real users before anyone notices.

What changes at scale: sampling becomes mandatory (full-fidelity everything is unaffordable); judge cost becomes a budget line (2% of 10M requests/day is 200k judge calls/day); cardinality discipline becomes survival; multi-provider routing means the model-version attribute does the heavy lifting for attribution; and cross-tenant isolation moves from nice-to-have to compliance requirement.

Likely follow-ups interviewers ask: How do you detect a provider's silent model update? (Frozen sentinels + version-pinned requests + score change-point alerts.) How do you evaluate the evaluator? (Expert-labeled calibration sets, agreement rates per slice, versioned judges.) How do you alert on quality without paging on noise? (Burn-rate on outcome SLOs, ticket on drift trends.) How do you sample without biasing your denominators? (Record sampling probability, keep audit paths unsampled, reweight analysis.)

The complete loop: requirement → instrumentation → telemetry → evaluation → alert → diagnosis → change → canary → rollback or widening → incident → regression. A useful observability system reduces time to detect, scope, and safely correct real user harm while respecting the data it collects.