Skip to content
Tech Interview Prep home
Technical interview guide

Monitoring, Logging & Tracing

The three pillars of observability, and what question each one is actually good at answering.

Read
30 min
Practice MCQs
25
Interview QA
25
Edition
v4
Editorial status
Reviewed

Scope: OpenTelemetry, W3C Trace Context, Prometheus, and Google SRE current 2026-08-31.

Overview

Curated: · Written: · Reviewed:

Observability is the ability to answer operational questions from trustworthy evidence

Metrics, logs and traces are complementary telemetry signals, not three boxes that automatically make a system observable. Observability begins with user journeys, service objectives, failure modes and decisions. Instrumentation should let an operator detect user harm, locate it across boundaries, test hypotheses, reconstruct relevant events and verify recovery. A large telemetry volume without semantic consistency, coverage, ownership or actionable workflows is expensive noise.

Metrics are aggregatable measurements over time. Counters represent cumulative events and can reset; gauges represent current values that can rise or fall; histograms preserve distributions through buckets or native structures. Metrics are efficient for rates, ratios, dashboards, objectives and alerting, but aggregation removes individual-event detail. Avoid averages for skewed latency and never average precomputed percentiles across instances. Labels make dimensions queryable but every unique label set creates a time series: user IDs, request IDs, unbounded URLs and arbitrary error text can cause cardinality and cost explosions.

Logs are discrete records of events. Prefer structured fields with stable schema, severity, time, service/resource identity, operation, outcome, error classification and trace/span correlation where available. A log line is not automatically true, complete or unique: retries duplicate events, clocks skew, buffers drop records and multiple services describe different perspectives. Do not use logs as an unaudited transactional source of truth. Define retention, indexing and sampling by investigation need, and minimize secrets, tokens, payloads, personal data and sensitive business values.

A trace represents causal work across a request or workflow as spans with parent/child or link relationships, timing, attributes, events and status. Context must be injected at the sender and extracted by the receiver across HTTP, messaging and asynchronous boundaries. Broken propagation creates fragmented traces; an incorrect parent creates false causality. W3C traceparent/tracestate supports interoperability, but incoming headers are untrusted input and propagated context can expose internal information. Validate formats, bound sizes, define trust boundaries and avoid personal data in trace context or baggage.

Correlation increases signal value. A trace ID and span ID on logs can move from a slow span to detailed events; exemplars can connect a metric observation to a representative trace. Correlation IDs are not authorization, proof of identity or globally collision-free business keys. Normalize resource and semantic attributes so queries mean the same thing across languages and teams. Version material schema changes and measure instrumentation coverage.

Sampling is a cost and evidence tradeoff. Head sampling decides early and cannot know the final outcome; uniform low-rate sampling can miss rare failures. Tail sampling can retain slow or erroneous traces after observing more of the trace, but needs buffering, consistent routing, memory, decision latency and explicit incomplete-trace behavior. Sampling probability must be known for statistical estimates. Keep critical audit/security records in their appropriate durable system rather than assuming a sampled trace is evidence of every action.

Design from questions. User-symptom metrics answer whether and how broadly something is wrong. Traces show which distributed path and span contributed. Logs reveal detailed local events and state. Profiles, events, topology and change/deployment data may be equally necessary; the “three pillars” metaphor is not a completeness guarantee. Use high-cardinality event stores for investigative dimensions that would be unsafe as metric labels.

Alert on actionable user symptoms or imminent objective risk, then attach dashboards, traces, relevant logs, recent changes, ownership and a runbook. Component metrics support diagnosis and capacity planning but should not page merely because they moved. Test alerts against incidents and benign changes. Monitor telemetry pipelines themselves: collection lag, dropped spans/logs, scrape gaps, cardinality, queue pressure, export failures, storage/query latency and cost. Missing telemetry is unknown, not healthy.

Instrumentation must preserve application behavior. Telemetry exporters require bounded queues, timeouts, retry budgets and backpressure policy; they must not turn an observability backend outage into a production outage. Avoid synchronous blocking in critical paths. Resource limits may deliberately drop lower-value telemetry, but the loss must be measured and visible. Collector and backend tiers need capacity, tenant isolation, quotas, encryption, least privilege and recovery testing.

Security and privacy span the full lifecycle: collection, propagation, transport, processing, storage, search, export and deletion. Authenticate producers, authorize queries, isolate tenants, encrypt data, redact before export, audit sensitive access and enforce retention. Attackers can forge trace headers, inject log text, cause high-cardinality dimensions or use telemetry as an exfiltration channel. Treat dashboards and alert payloads as disclosure surfaces.

Validate observability with known operations and failures. Assert expected counters, histogram buckets, log fields, trace topology and correlations; test retries, async queues, broken propagation, sampling, clock skew, exporter outage and backend saturation. During incidents, compare telemetry with independent user and business evidence. A green dashboard or complete-looking trace proves only what the instrumented, retained and queryable evidence covers.

The pillars are signals, not a strategy

Describing observability as metrics, logs and traces answers what data exists, not what questions the data can answer, and the interview usually turns on the second. A useful framing is to start from the questions an operator must answer during an incident — is the user affected, which dependency changed, is this release or this tenant, has it happened before — and then say which signal answers each and what would make it unable to. That framing exposes the gaps that a pillar-by-pillar inventory hides, such as high-cardinality dimensions dropped at ingest, traces sampled away exactly when the system is degraded, or logs whose retention is shorter than the time it takes to notice the problem.

Cost is part of the design rather than a constraint imposed afterwards. Cardinality drives metric cost superlinearly, full-fidelity tracing is usually unaffordable at high request rates, and log volume grows with both traffic and verbosity. The engineering response is to decide deliberately: aggregate where the question is about rates, sample where the question is about typical behaviour, and retain full fidelity only where the question is about individual failures — with tail-based or error-biased sampling so the interesting traces are the ones kept.

The last point is trust. A signal that can be silently wrong is worse than one that is absent, because it is acted upon. Instrumentation needs its own freshness and completeness checks — pipeline lag, drop counts, agent health — and dashboards should make missing data visibly different from healthy data rather than rendering an empty graph that reads as zero.

Worked example: 10,000 checkouts, 40 failures, 1% head sample

Same hour: 10,000 checkout requests, 40 failures (0.4%). On-call opens traces, not the error-rate graph.

policytraces storedexpected failure tracescan they see the 40
1% head sampling~1000.4often zero — incident looks like a green trace explorer
tail: keep status=error, 1% of successes~40 + 10040yes, plus a success baseline
metrics only0 traces0rate is visible; the failing dependency span is not

A complete-looking sampled trace proves the sampled path. It does not prove the 40 failures were retained. That table is the interview.