Skip to content
Tech Interview Prep home
Technical interview guide

Neural Network Fundamentals

Forward pass, backpropagation, and the activation functions that make deep networks work.

Read
27 min
Practice MCQs
25
Interview QA
25
Edition
v3
Editorial status
Reviewed

Scope: Google ML Crash Course neural-network guidance, TensorFlow Keras guides, and PyTorch stable reproducibility guidance accessed 2026-08-31..

Overview

Curated: · Written: · Reviewed:

Key takeaways

  • A dense layer computes an affine transformation; nonlinear activations between layers let stacked networks represent nonlinear functions. Stacking affine layers without nonlinearities collapses to one affine map.
  • Output activation and loss must match the task and target representation: logits plus a numerically stable cross-entropy implementation are common for classification; regression outputs and losses depend on target support and decision cost.
  • Backpropagation applies the chain rule over the computation graph. Gradients stop where the graph is cut — detached tensors, frozen layers, stop_gradient. An optimizer uses those gradients; it does not determine whether the data split, objective, or labels are valid.
  • Initialization, scale, activation, depth, normalization, optimizer, learning-rate schedule, batch size, and precision interact. Diagnose training through curves, gradient/activation statistics, finite checks, and controlled baselines.
  • Training behavior and inference behavior can differ for dropout and batch normalization. Checkpoint/export must preserve the architecture, learned and non-trainable state, preprocessing contract, and serving signature.
  • Neural networks are not automatically superior. Compare with simple baselines using deployment-like splits, calibration and cohort metrics, latency, memory, reproducibility, and rollback evidence.

1. Forward computation and architecture

A unit commonly computes z = w·x + b and then applies an activation. A dense layer applies that operation across output units. Layer dimensions must compose: a batch has an explicit batch axis, each weight matrix maps input width to output width, and broadcasting of biases must be intentional. Interviewers frequently probe shape bookkeeping here — expect "given a batch of shape (B, 784) and a hidden layer of 256 units, what are the weight and bias shapes?" (W: 784×256, b: 256, output: (B, 256)). A weak answer hand-waves "the framework figures it out"; a strong one traces the shapes layer by layer and can spot a transposed weight or a silently broadcast bias.

Parameters are learned state; hyperparameters include width, depth, activation, optimizer settings, and training budget.

Without nonlinear activations, multiple dense layers compose into one affine transformation and cannot create nonlinear decision boundaries. ReLU returns max(0,x), sigmoid maps to (0,1), and tanh to (-1,1). ReLU is simple and often avoids sigmoid/tanh saturation-driven vanishing gradients, but units can become permanently inactive (dead ReLU) when their pre-activation stays negative across the data — a unit whose weights drift so the input never clears zero contributes a zero gradient to everything upstream of it. Activation choice depends on architecture and output semantics rather than a universal ranking.

Sequential models suit a plain single-input/single-output stack. Graph/functional models represent branches, skip connections, shared layers, multiple inputs, and multiple outputs. Parameter sharing changes the learning assumption; skip (residual) connections give gradients a short path to early layers and are a standard mitigation for vanishing gradients in deep stacks. Every tensor shape, mask, trainable state, and input name belongs to the versioned model contract.

2. Outputs, losses, and probability semantics

Binary classification commonly uses one logit with binary cross-entropy; multiclass single-label classification uses one logit per class and softmax cross-entropy; multilabel classification uses independent logits and binary losses. The pairing is not arbitrary: softmax cross-entropy consumes raw logits because the log-softmax and the loss combine into a numerically stable form (the "log-sum-exp trick"), and applying softmax yourself first throws away that stability and can produce log(0). The same logic pairs sigmoid with binary cross-entropy — the loss is written against logits so the sigmoid and log combine without overflow. A classic follow-up: "why did my loss go to NaN when I applied softmax then cross-entropy?" The answer is exactly this double-activation plus explicit-log mistake. One-hot versus integer labels must match the selected loss API.

Regression may use an unrestricted linear output, a positive-domain transform, or bounded output only when domain semantics justify it. MAE, MSE, Huber, and quantile losses impose different error costs. The training loss need not be the only evaluation metric, but it must be coherent with the task. Class/sample weights alter the objective and may change calibration; thresholds are chosen on validation evidence, not assumed from 0.5. A weak answer names a loss without saying what error cost it encodes; a strong one can say when MSE is the wrong choice (heavy-tailed residuals, where a few outliers dominate the gradient).

3. Backpropagation and optimization

Backpropagation records or reconstructs the computation graph, starts with the loss derivative, and repeatedly applies the chain rule from outputs to earlier parameters. Gradients accumulate contributions along paths; shared parameters receive all relevant contributions. Gradients also stop where the graph is cut: a detached tensor (.detach() in PyTorch, stop_gradient in JAX) or a frozen layer (requires_grad=False) contributes nothing upstream. Interviewers use this to test whether you understand autodiff as graph traversal — "what happens to the gradient at a detached branch?" The correct answer is that the branch's outputs are treated as constants; the weak answer is "it still learns somehow."

Gradient checking on tiny smooth examples can compare autodiff with finite differences, but nonsmooth points (ReLU kinks) and floating-point tolerance require care.

Gradient descent moves parameters opposite an estimate of the loss gradient. Mini-batches trade noisy frequent updates against throughput and memory. Momentum accumulates direction; adaptive optimizers rescale updates using gradient history. Learning rate often dominates success: too high can diverge or oscillate, too low can appear stuck. Schedules, warmup, clipping, weight decay, and mixed precision interact and require logged, controlled experiments rather than blind defaults.

4. Stable training and generalization

All-zero or identical initialization makes hidden units learn identical features — every unit in a layer receives the same gradient and stays identical forever. Naive random init (say, standard normal at every layer) breaks deep training differently: activations grow or shrink multiplicatively through the stack, so the forward signal saturates or vanishes before reaching the output. Variance-scaled initialization fixes this by keeping activation variance roughly constant across layers. Xavier/Glorot initialization is tuned for symmetric, zero-centered activations like tanh; He initialization scales for ReLU-family activations, whose negative half outputs zero and so halve the variance per unit. Using Xavier with deep ReLU stacks is a common silent failure — the expected follow-up is "which do you use with ReLU and why," and "He, because ReLU zeroes half the pre-activations" is the answer that survives.

Vanishing and exploding gradients have readable symptoms. Vanishing: loss decreases then plateaus early, per-layer gradient norms shrink geometrically with depth (early layers' norms orders of magnitude below late layers'), sigmoid/tanh activations saturating near 0/1 or ±1. Exploding: loss spikes to NaN or huge values, gradient norms growing with depth. Standard mitigations, roughly in order of how often they're cited: ReLU-family activations, residual connections, normalization layers, gradient clipping (caps update magnitude but masks the underlying cause — say this out loud), and sensible initialization. Batch normalization maintains learned affine parameters and running statistics, with different training and inference behavior; very small or shifted batches make statistics problematic. Layer normalization uses within-example features and has different semantics. Dropout masks activations only during training under normal inference.

Inspect per-layer activation and gradient distributions, parameter/update norms, finite values, dead-unit rates, and loss by batch. Early stopping, weight penalties, data augmentation, smaller architecture, and more representative data are different levers; fit and select them inside valid development evidence.

5. Evaluation, reproducibility, and serving

Split by time, entity, site, or other deployment unit before learned preprocessing. Compare against linear/tree/rule and dummy baselines. Report the declared primary metric and guardrails, uncertainty, calibration, confusion/residual evidence, relevant cohorts, train-versus-validation curves, and operational cost. A lower training loss is not a release criterion when held-out or product outcomes regress — interviewers probe this as "your offline metric improved but the A/B test didn't; what do you check first?" Distribution shift, leakage in the offline split, and metric-proxy mismatch are the answers that show judgment.

Seeds reduce some variation but do not guarantee identical results across framework releases, devices, kernels, or distributed execution. Record data, code, architecture, initialization, optimizer/schedule, precision, batch order, seeds, libraries, hardware, and determinism settings. Repeat important experiments and report dispersion; deterministic kernels can cost throughput.

A resumable training checkpoint needs architecture/configuration, trainable and non-trainable weights, optimizer/scheduler and step state, plus data position and randomness where exact continuation matters. A serving export may intentionally contain only forward inference. Validate reload and export on golden fixtures, input signatures, preprocessing, training/inference mode, numerical tolerance, latency, concurrency, and failure behavior. Canary, monitor model/data versions and mature outcomes, and retain rollback.

Worked example: two affine layers with and without ReLU

Hidden pre-activation z₁ = 2x + 1. Second layer z₂ = 3·a(z₁) − 4. If a is identity, the stack collapses to z₂ = 6x − 1. If a is ReLU, the output is stuck at −4 whenever z₁ < 0 (that is, x < −0.5).

xz₁ = 2x+1affine stack 6x−1ReLU hidden 3·max(0,z₁)−4
-2-3-13-4
-1-1-7-4
-0.50-4-4
01-1-1
1355
251111

At x = -2 the two models disagree by 9. At x = 1 they agree because the hidden unit is in its linear region. That single table is why a nonlinearity is not optional decoration: without it the extra layer is algebra, not capacity.