New batches starting this week Β· Limited seats

Deep Learning Interview Questions and Answers 2026 (65 Questions)

65 deep learning interview questions with clear answers, from activations and backpropagation to PyTorch training loops, GPU memory, distributed training, generative models and real debugging scenarios.

Deep learning interview questions 2026: 65 questions on backpropagation, optimisers, regularisation, CNNs, transformers and PyTorch
Last updated Β· 48 min read Β· 10,490 words

Deep learning interview questions in 2026 test whether you can train a neural network that actually converges, explain why it works, and debug it when it doesn't, not whether you can recite the definition of a neuron. This guide collects 65 high-value questions with model answers covering fundamentals, optimisation, regularisation, CNNs, RNNs and transformers, transfer and self-supervised learning, PyTorch, mixed precision, distributed training, generative models and real debugging scenarios.

How to use this guide

This page sits between two sibling guides. Our machine learning interview questions cover classical ML, metrics, validation and production ML systems, and touch neural networks only briefly. Our LLM interview questions go deep into transformer internals, tokenization, RLHF and LLM serving. Here the focus is the layer in between: how neural networks are built and trained, and how a deep learning engineer reasons about training behaviour. What interviewers commonly probe at each level:

  • Freshers: activation functions, backpropagation, the purpose of dropout and normalisation, what a convolution computes, and a correct PyTorch training loop.
  • Mid-level engineers: initialisation, AdamW versus Adam, warmup, batch norm pitfalls, transfer learning strategy, mixed precision and GPU memory arithmetic.
  • Senior engineers: distributed training choices, self-supervised pretraining, generative model trade-offs and structured diagnosis of training failures.

Answer each question aloud first, then compare. Where code appears, it is short and illustrative, not a full program.

Neural network fundamentals

1. What is a perceptron, and why can't a single one solve XOR?

Answer: A perceptron computes a weighted sum of its inputs plus a bias and passes it through a step function: output 1 if wΒ·x + b > 0, else 0. Geometrically it draws one straight line (a hyperplane) through the input space, so it can only separate classes that are linearly separable. XOR's positive points, (0,1) and (1,0), sit diagonally opposite each other, and no single line separates them from (0,0) and (1,1). Adding one hidden layer with a non-linear activation fixes this: the hidden units carve the space into regions that the output unit can then combine. That observation is the core argument for multi-layer networks.

2. Why do neural networks need non-linear activation functions?

Answer: Without non-linearity, stacking layers achieves nothing: two linear layers W2(W1x) collapse into one linear layer (W2W1)x, so a hundred-layer network is still a linear model. Non-linear activations between layers let the network compose simple functions into very complex ones. The universal approximation theorem says even one wide hidden layer with a suitable non-linearity can approximate any continuous function on a bounded domain, but in practice depth is far more parameter-efficient than width, because deep networks reuse intermediate features hierarchically (edges, then textures, then parts, then objects).

3. Compare sigmoid, tanh, ReLU, Leaky ReLU and GELU. Which would you use where?

Answer:

ActivationRangeStrengthWeaknessTypical use
Sigmoid(0, 1)Reads as a probabilitySaturates; max gradient 0.25; not zero-centredBinary output, gates in LSTMs
Tanh(-1, 1)Zero-centredStill saturates at extremesRNN hidden states
ReLU[0, ∞)Cheap; no saturation for positive inputs"Dead" units stuck at zeroDefault for CNNs and MLPs
Leaky ReLU(-∞, ∞)Small negative slope keeps gradient aliveExtra hyperparameterWhen dead ReLUs appear
GELU / SiLUSmooth, small negative dipSmooth gradients, strong empirical resultsSlightly costlierTransformers, modern CNNs

Hidden layers: ReLU for a plain baseline, GELU or SiLU for transformer-style models. Output layer depends on the task, not fashion: none (raw logits) for classification with a cross-entropy loss, sigmoid for independent multi-label outputs, identity for regression.

4. What is a "dying ReLU", and how do you detect and fix it?

Answer: A ReLU unit dies when its pre-activation is negative for every input in the data, so it outputs zero and receives zero gradient, and therefore never recovers. It usually follows a large update, often from too high a learning rate, that pushes the bias strongly negative. Detect it by logging the fraction of activations that are exactly zero per layer across a batch; a layer where most units are always zero is a red flag. Fixes: lower the learning rate or add warmup, use He initialisation, switch to Leaky ReLU, GELU or SiLU, and add normalisation before the activation.

5. Why do classifiers output logits, and why should softmax and cross-entropy be computed together?

Answer: Logits are unnormalised scores. Softmax turns them into probabilities, and cross-entropy penalises the negative log probability of the correct class. Computing them separately is numerically fragile: softmax of large logits can overflow, and the log of a tiny probability underflows to negative infinity. Fused implementations use the log-sum-exp trick (subtract the maximum logit first), which is stable. In PyTorch, nn.CrossEntropyLoss expects raw logits and applies log-softmax internally; feeding it softmax outputs is a classic bug that still trains, but slowly and badly. A neat consequence of the pairing: the gradient of the loss with respect to the logits is simply softmax(z) - y.

6. Explain backpropagation as a computational graph. What does it actually compute?

Answer: The forward pass builds a directed graph of operations from inputs and parameters to a scalar loss, storing intermediate values. Backpropagation walks that graph in reverse, applying the chain rule at each node: each operation receives the gradient of the loss with respect to its output and multiplies by its local derivative to get the gradient with respect to its inputs. Gradients flowing into a node from several paths are summed. This reverse-mode automatic differentiation yields every parameter's gradient for roughly the cost of a couple of forward passes, because there is one scalar output and millions of inputs.

Interview tip: Mention the memory cost: forward activations are kept until backward uses them, which is why activation checkpointing exists (Q47). The ML guide covers the chain rule derivation itself.

7. What causes vanishing and exploding gradients in deep networks, and which design choices address them?

Answer: The gradient reaching an early layer is a product of many per-layer Jacobians. If their typical scale is below one, the product shrinks exponentially with depth (vanishing); above one, it grows exponentially (exploding). Saturating activations such as sigmoid make vanishing worse, and recurrent networks, which multiply by the same weight matrix at every time step, suffer both. Modern deep learning is largely a set of answers to this one problem:

  • Initialisation that keeps activation variance stable across layers (Q8).
  • Non-saturating activations such as ReLU and GELU.
  • Normalisation layers that re-centre and re-scale activations (Q19).
  • Residual connections, which give gradients an identity path around each block (Q24).
  • Gating in LSTMs and GRUs (Q27).
  • Gradient clipping as a safety net for exploding gradients (Q14).

Initialisation and optimisation

8. Why not initialise all weights to zero? Explain Xavier and He initialisation.

Answer: If every weight in a layer starts equal, every neuron computes the same output and receives the same gradient, so they stay identical forever. Random initialisation breaks this symmetry. The scale matters too: too small and signals shrink to nothing through depth, too large and they explode. Xavier (Glorot) initialisation sets the weight variance to about 2 / (fan_in + fan_out), keeping variance roughly constant for tanh or sigmoid layers. He (Kaiming) initialisation uses 2 / fan_in, which compensates for ReLU zeroing half its inputs. Biases are usually initialised to zero. PyTorch layers come with sensible default initialisations, but if you change the activation or build an unusually deep network, set them deliberately.

9. How does SGD with momentum differ from plain SGD, and what does Nesterov momentum add?

Answer: Plain SGD steps along the current mini-batch gradient, which is noisy and zig-zags in narrow ravines of the loss surface. Momentum keeps an exponentially decaying average of past gradients (a velocity) and steps along that, so consistent directions accelerate and oscillating ones cancel out. A momentum coefficient around 0.9 is the common default. Nesterov momentum evaluates the gradient at the look-ahead position where the velocity is about to carry the parameters, which corrects overshoot a little earlier. Well-tuned SGD with momentum remains competitive for many vision models, at the cost of more learning-rate tuning.

10. What is the difference between Adam and AdamW, and why does it matter?

Answer: Adam keeps per-parameter running averages of the gradient (first moment) and squared gradient (second moment), with bias correction, and divides each step by the square root of the second moment, giving every parameter its own effective learning rate. The original way to add weight decay to Adam was to add an L2 term to the loss. The problem is that this L2 gradient then gets divided by the same adaptive denominator, so parameters with large historical gradients are barely regularised. AdamW decouples weight decay: it shrinks the weights directly, separate from the adaptive gradient step. That makes weight decay behave as intended and makes its value easier to tune independently of the learning rate. AdamW is the default for transformers and most modern fine-tuning.

11. Which learning-rate schedules are common, and how do you choose one?

Answer: The learning rate is usually the most important hyperparameter, and a fixed one is rarely optimal: you want large steps early to make progress and small steps late to settle into a good minimum. Common schedules:

  • Step decay: divide the rate by a factor at fixed epochs; simple and still used in vision.
  • Cosine decay: smoothly decreases the rate along a half cosine to near zero; a widely used default for pretraining and fine-tuning.
  • One-cycle: rises then falls within one run; useful for fast training on a fixed budget.
  • Reduce-on-plateau: lowers the rate when validation loss stalls; convenient when you don't know the run length.

Cosine, linear and one-cycle need the total step count up front; when fine-tuning, start from the architecture's published recipe.

12. What is learning-rate warmup, and why do transformers in particular need it?

Answer: Warmup starts the learning rate near zero and raises it linearly to the target over the first few hundred or thousand steps. At the start of training, Adam's second-moment estimates are based on very few samples and can be unreliable, and randomly initialised networks can produce large, poorly conditioned gradients. Full-size steps at that moment can push the model into a bad region or produce loss spikes and NaNs. Transformers, especially deep post-norm ones, are notably sensitive to this. Warmup lets the optimiser statistics and the network settle before large steps begin. It is cheap insurance and appears in almost every transformer training recipe.

  lr
   |      ________
   |     /        ``--..
   |    /               `-.
   |   /                   `.
   |__/______________________`__ steps
     warmup     cosine decay

13. How are batch size and learning rate related?

Answer: A larger batch gives a less noisy gradient estimate, so you can usually take larger steps. The linear scaling heuristic says that when you multiply the batch size by k, multiply the SGD learning rate by roughly k, with warmup; for Adam-style optimisers a square-root scaling is often a better starting point. The heuristic breaks down at very large batches, where returns diminish and generalisation can suffer because the regularising noise of small batches is lost. In practice, GPU memory sets the batch size and the learning rate is re-tuned to match.

14. What is gradient clipping, and when should you use it?

Answer: Gradient clipping caps the size of gradients before the optimiser step. Clipping by global norm, the common choice, rescales the whole gradient vector so its L2 norm does not exceed a threshold (values around 1.0 are typical for transformers), which preserves its direction. Clipping by value caps each element independently and can distort direction. Use it for RNNs, transformers and any run that shows occasional loss spikes. Log the pre-clip gradient norm: if clipping fires on almost every step, the threshold is too low or something else is wrong, such as the learning rate.

loss.backward()
torch.nn.utils.clip_grad_norm_(
    model.parameters(), max_norm=1.0)
optimizer.step()

Regularisation and normalisation

15. How does dropout behave differently in training and inference?

Answer: During training, dropout zeroes each activation with probability p and scales the survivors by 1 / (1 - p) (inverted dropout), so the expected activation is unchanged. This stops units from co-adapting and acts like training an ensemble of thinned networks. At inference, dropout is switched off and, because of the scaling during training, no further adjustment is needed. In PyTorch this switch is controlled by model.train() and model.eval(); forgetting eval() gives noisy, non-deterministic predictions. Dropout is common in fully connected heads and in transformer fine-tuning; large pretraining runs often use little or none because the data itself prevents overfitting.

16. Beyond dropout, which regularisation techniques do deep learning engineers actually use?

Answer: Weight decay (via AdamW) is almost universal. Data augmentation is often the strongest regulariser for images and audio. Label smoothing replaces one-hot targets with slightly softened ones, discouraging over-confident logits and often improving calibration. Mixup and CutMix train on blended images and labels. Stochastic depth randomly skips residual blocks during training. Early stopping halts when validation loss stops improving. Finally, more and cleaner data beats all of these. Match technique to failure: an underfitting model needs less regularisation, not more.

17. How do you design a data augmentation pipeline, and how can augmentation go wrong?

Answer: Augmentation should produce inputs the model could plausibly see in production while keeping the label valid. For natural images: random crops, flips, colour jitter and small rotations. For audio: noise, time shifting and masking parts of the spectrogram. For text: back-translation or synonym replacement, used carefully. It goes wrong when a transform changes the meaning. Flipping a digit "6" vertically, horizontally flipping text in a document image, or rotating a chest X-ray so the heart appears on the wrong side all create mislabelled data. Apply augmentation only to the training set, keep validation preprocessing deterministic, and check augmented samples visually before a long run. The computer vision interview questions go further into image-specific augmentation.

18. How do you implement early stopping correctly?

Answer: Evaluate on a validation set at fixed intervals, track the best validation metric, save a checkpoint whenever it improves, and stop when it has not improved for a patience window of several evaluations. At the end, restore the best checkpoint rather than using the final weights. Three common mistakes: choosing the stopping point using the test set (which leaks the test set into model selection), using too small a patience with a noisy metric, and stopping on validation loss when the business metric is accuracy or recall, which can peak at a different point. With cosine or one-cycle schedules, many teams instead train for a fixed budget and keep the best checkpoint.

19. Compare batch normalisation and layer normalisation. Why do transformers use layer norm?

Answer: Both standardise activations and then apply a learned scale and shift; they differ in which axis they compute statistics over.

AspectBatch normLayer norm
Statistics overThe batch (and spatial positions) for each channelAll features of a single example
Train vs evalBatch statistics in training; running averages at inferenceIdentical in both modes
Small batchesNoisy, unreliable statisticsUnaffected
Variable-length sequencesAwkward (padding pollutes statistics)Natural fit
Typical homeCNNsTransformers, RNNs

Transformers use layer norm (or its simpler variant RMSNorm, which drops mean-centring) because sequences vary in length, batches are often small per device, and autoregressive inference processes one example at a time, where batch statistics make no sense. Group norm is the usual fix for CNNs with tiny batches.

20. What is the difference between pre-norm and post-norm residual blocks?

Answer: In the original transformer (post-norm), normalisation is applied after adding the residual: x = LN(x + f(x)). In pre-norm, it is applied to the branch input: x = x + f(LN(x)). Pre-norm leaves a clean identity path from the input to the output of the whole stack, so gradients flow more easily and deep models train more stably, often with less dependence on warmup. Most modern large transformers use pre-norm, usually with a final normalisation before the output head. The LLM guide covers how this fits into a full decoder block.

Convolutional neural networks

21. What does a convolutional layer compute, and why is it so parameter-efficient?

Answer: A convolutional layer slides small learned filters (for example 3Γ—3 across all input channels) over the input and computes a dot product at each position, producing one output feature map per filter. Two properties make it efficient. Local connectivity: each output depends only on a small neighbourhood, matching the fact that nearby pixels are related. Weight sharing: the same filter is applied everywhere, so a feature detector learned in one place works in all places, giving translation equivariance. A 3Γ—3 convolution from 64 to 128 channels has 3Γ—3Γ—64Γ—128 weights plus 128 biases, roughly 74,000 parameters, regardless of image size; a fully connected layer between the same feature maps of even a small image would need many millions.

22. How do you calculate a convolution's output size? What do padding and stride do?

Answer: For input size n, kernel k, padding p and stride s, the output size per spatial dimension is floor((n + 2p - k) / s) + 1. Padding adds a border (usually zeros) so edges are processed and, with p = (k - 1) / 2 and stride 1, the size is preserved ("same" padding). Stride is the step between filter positions; stride 2 roughly halves each dimension and is a common learned alternative to pooling. Dilation spaces out the kernel taps to enlarge the receptive field without extra parameters. Example: a 224Γ—224 input with a 7Γ—7 kernel, stride 2 and padding 3 gives floor((224 + 6 - 7) / 2) + 1 = 112.

23. What is the receptive field, and why does it matter? Where does pooling fit?

Answer: The receptive field of a unit is the region of the original input that can influence it. Stacking two 3Γ—3 convolutions gives a 5Γ—5 receptive field and three give 7Γ—7, with fewer parameters and more non-linearity than one large kernel, which is why stacks of small kernels became standard. Downsampling (strided convolution or pooling) multiplies the growth, so deeper layers see large regions and can recognise whole objects. Max pooling keeps the strongest response in each window, adding a little translation invariance; global average pooling at the end collapses each feature map to one number, replacing large fully connected heads. If the receptive field is smaller than the objects you need to recognise, accuracy suffers no matter how much you train.

24. How do residual (skip) connections work, and why did ResNets allow much deeper networks?

Answer: Before ResNets, simply adding layers to a plain CNN eventually made even training error worse, which is an optimisation problem, not overfitting. A residual block computes y = x + F(x), so the layers learn a residual correction instead of a full mapping. If the best thing a block can do is nothing, it only has to push F towards zero, which is easy. In backpropagation, the addition passes the gradient straight through to earlier layers, so even very deep stacks receive a useful signal. When shapes differ (more channels or lower resolution), the skip path uses a 1Γ—1 convolution with stride to match. The same idea sits inside every transformer block.

x ---------------------------+
|                            |
+-> conv-BN-ReLU-conv-BN -->(+)--> ReLU --> y

25. What are 1Γ—1 convolutions and depthwise separable convolutions used for?

Answer: A 1Γ—1 convolution mixes information across channels at each position without looking at neighbours. It is used to change channel count cheaply, for example the bottleneck blocks in deeper ResNets, which reduce channels, apply a 3Γ—3, then expand again. A depthwise separable convolution factorises a standard convolution into a depthwise step (one spatial filter per input channel) followed by a 1Γ—1 pointwise step that mixes channels. This cuts computation and parameters substantially for a small accuracy cost, which is why it powers mobile-oriented architectures. It matters whenever you deploy to phones, cameras or edge devices; see on-device and edge AI for the deployment side.

RNNs, LSTMs and transformers

26. How does an RNN process a sequence, and what is backpropagation through time?

Answer: An RNN keeps a hidden state and updates it one step at a time: h_t = tanh(W_x x_t + W_h h_{t-1} + b), using the same weights at every step. Training unrolls the network across the sequence into one long feedforward graph and backpropagates through it; this is backpropagation through time (BPTT). Because the gradient to early steps multiplies by W_h repeatedly, it vanishes or explodes over long sequences, so plain RNNs struggle to learn dependencies more than a few dozen steps back. Truncated BPTT limits how far back gradients flow, saving memory at the cost of the longest dependencies.

27. How do LSTM gates solve the vanishing gradient problem, and how does a GRU differ?

Answer: An LSTM adds a separate cell state that runs along the sequence and is changed only by gated, mostly additive updates. The forget gate decides how much of the old cell state to keep, the input gate decides how much new candidate information to write, and the output gate decides how much of the cell to expose as the hidden state. Because the cell is updated additively and the forget gate can stay near one, gradients can flow across many steps without shrinking. A GRU merges the forget and input gates into one update gate, adds a reset gate, and has no separate cell state. It has fewer parameters and often performs comparably; choose by validation results rather than doctrine.

28. Why did transformers largely replace RNNs for language and many sequence tasks?

Answer: Three reasons. Parallelism: an RNN must process step t before t+1, so training cannot be parallelised across the sequence; self-attention processes all positions at once, which suits GPUs and made training on far larger datasets practical. Path length: in an RNN, information from 500 steps back must survive 500 state updates; in attention, any two positions connect directly in one layer. Scaling: transformers kept improving as data and parameters grew, and the same architecture transferred across text, images, audio and code. The trade-offs: attention cost grows quadratically with sequence length, and autoregressive inference needs a growing KV cache, whereas an RNN carries a fixed-size state. RNNs and small 1D CNNs remain reasonable for small, low-latency sensor or time-series tasks.

29. Explain attention conceptually. What problem was it originally invented to solve?

Answer: Attention was first added to RNN encoder-decoder translation models, which had to squeeze a whole source sentence into one fixed-size vector before decoding. Attention let the decoder, at each output step, look back over all encoder states and take a weighted average, with weights based on relevance to what it is currently generating. The transformer kept that idea and dropped the recurrence: every token produces a query, key and value; query-key similarity, passed through softmax, decides how much each token draws from every other token's value. Several heads run in parallel to capture different relationships, and positional information is added because attention by itself ignores order. For the maths, scaling, multi-head variants and positional encoding schemes, see the LLM interview guide, which covers these in depth.

30. How are transformers applied beyond text, for example to images?

Answer: A transformer only needs a sequence of vectors. A Vision Transformer (ViT) cuts an image into fixed-size patches, flattens and linearly projects each patch into an embedding, adds position embeddings, and feeds the sequence to a standard transformer encoder, often with a special classification token. ViTs have weaker built-in assumptions than CNNs (no locality or weight sharing by design), so they typically need more data or strong pretraining to beat CNNs, but they scale well and pretrained ViTs transfer strongly. The computer vision guide compares CNN and ViT choices for detection and segmentation.

Embeddings, transfer and self-supervised learning

31. What is representation learning, and what makes an embedding "good"?

Answer: Representation learning means the network learns its own features rather than relying on hand-engineered ones. The layers before the task head form a representation (an embedding) in which similar inputs land near each other. A good embedding makes downstream tasks easy: a simple linear classifier or nearest-neighbour search on top of it works well (the "linear probe" test). It should ignore nuisance variation such as lighting or phrasing. An embedding layer in PyTorch (nn.Embedding) is just a learned lookup table from IDs to vectors. For how embeddings are used in search and RAG, see embeddings explained.

32. What are the main transfer learning strategies, and how do you choose between them?

Answer: Transfer learning reuses a model pretrained on a large dataset as the starting point for your task. The options form a spectrum:

  • Feature extraction: freeze the backbone, train only a new head. Fast, needs little data, hard to overfit.
  • Partial fine-tuning: unfreeze the top layers, which hold the most task-specific features, and keep early generic layers frozen.
  • Full fine-tuning: update everything, usually with a lower learning rate for the backbone than the head (discriminative learning rates).
  • Parameter-efficient fine-tuning: train small adapters such as LoRA while the backbone stays frozen; common for large models.

Little data and a similar domain: freeze more. Plenty of data or a distant domain (satellite, medical, industrial imagery): fine-tune more. A common recipe is to train the head first with the backbone frozen, then unfreeze and fine-tune everything at a small learning rate, so a random head doesn't send noisy gradients into good pretrained weights.

33. What are the common pitfalls when fine-tuning a pretrained model?

Answer: Using the pretraining learning rate instead of a much smaller one, which wipes out useful features. Mismatched preprocessing: the pretrained model expects specific input sizes, normalisation statistics or tokenisation, and silently performs worse otherwise. Forgetting to replace or resize the head for your number of classes. Batch norm statistics updating from small, unrepresentative fine-tuning batches; sometimes it is better to keep BN layers in eval mode or frozen. Training too long on a small dataset and overfitting. Finally, for large language models specifically, there is catastrophic forgetting of general skills; the fine-tuning LLMs guide covers that case.

34. What is self-supervised learning, and how do contrastive and masked approaches differ?

Answer: Self-supervised learning creates the training signal from the data itself, so a model can learn from huge unlabelled collections and then be fine-tuned with few labels. Two families dominate:

  • Contrastive and joint-embedding methods: create two augmented views of the same input and train the encoder so their embeddings are close while embeddings of different inputs are far apart. Some variants avoid explicit negatives using techniques such as a momentum teacher network and stop-gradient, but must prevent collapse, where every input maps to the same vector.
  • Masked prediction: hide part of the input and predict it. Masked language modelling hides tokens; masked autoencoders hide image patches and reconstruct them. Next-token prediction in LLMs is also self-supervised.

Contrastive methods depend heavily on the augmentations, which define what the model learns to ignore. Masked methods need fewer augmentation choices and scale well with model size.

35. How does a contrastive loss work, and what does the temperature parameter do?

Answer: In an InfoNCE-style loss, each example's positive pair must be picked out from a set of negatives (often the other examples in the batch). Similarities, usually cosine similarities of normalised embeddings, are divided by a temperature and passed through softmax cross-entropy, where the "correct class" is the positive. Temperature controls how sharply the loss focuses on the hardest negatives: a low temperature magnifies small differences and concentrates on the most similar negatives; a high one spreads the penalty evenly. Because negatives come from the batch, larger batches usually help. The same loss underlies image-text models and retrieval embedding models.

PyTorch essentials

36. What should every PyTorch engineer know about tensors: shape, dtype, device, views and contiguity?

Answer: A tensor has a shape, a dtype (float32, bfloat16, int64 and so on) and a device (CPU or a specific GPU). Operations require compatible devices and usually compatible dtypes; mixing them causes the most common runtime errors. Many operations such as view, transpose, permute and slicing return views that share memory with the original, so modifying one modifies the other. Transposing makes a tensor non-contiguous in memory, and view then fails; reshape copies when needed, or call .contiguous() first. Broadcasting automatically expands dimensions of size one, which is convenient but hides bugs: subtracting a (N,) tensor from an (N, 1) tensor silently produces an (N, N) matrix.

x = torch.randn(32, 3, 224, 224, device="cuda")
y = x.permute(0, 2, 3, 1)      # view, not contiguous
z = y.reshape(32, -1)          # copies if needed
labels = labels.to(x.device)   # match devices

37. How does autograd work in PyTorch, and why do you call zero_grad()?

Answer: Tensors with requires_grad=True (all nn.Parameters by default) are tracked: each operation records a node in a graph built on the fly during the forward pass. loss.backward() traverses that graph and accumulates gradients into each leaf tensor's .grad, then frees the graph. Accumulation, not overwriting, is deliberate: it allows gradient accumulation across micro-batches (Q47) and summing losses from multiple passes. The consequence is that you must reset gradients every step, with optimizer.zero_grad(), or gradients from previous steps are added to the new ones. Recent PyTorch versions set gradients to None by default when zeroing, which saves memory. Use .detach() to cut a tensor out of the graph, and .item() to pull a scalar out for logging so you don't keep the graph alive.

38. How do you write an nn.Module properly? What is the difference between a parameter and a buffer?

Answer: Subclass nn.Module, create layers in __init__, define computation in forward, and call the module as model(x) rather than model.forward(x) so hooks run. Submodules and nn.Parameters assigned as attributes are registered automatically and appear in model.parameters() and the state dict. Layers stored in a plain Python list are not registered, so the optimiser never sees them; use nn.ModuleList or nn.ModuleDict. A buffer, created with register_buffer, is state that moves with the model across devices and is saved in the state dict but is not trained, for example batch norm running statistics or a fixed mask.

class Classifier(nn.Module):
    def __init__(self, d_in, n_cls):
        super().__init__()
        self.blocks = nn.ModuleList(
            [nn.Linear(d_in, d_in) for _ in range(2)])
        self.head = nn.Linear(d_in, n_cls)

    def forward(self, x):
        for blk in self.blocks:
            x = F.gelu(blk(x))
        return self.head(x)    # raw logits

39. Write a minimal, correct PyTorch training loop and explain each step.

Answer: The order matters: put the model in training mode, move each batch to the device, clear old gradients, run the forward pass, compute the loss, backpropagate, optionally clip, step the optimiser, then step a per-iteration scheduler. Validation runs in eval mode without gradient tracking.

for epoch in range(epochs):
    model.train()
    for xb, yb in train_loader:
        xb, yb = xb.to(dev), yb.to(dev)
        optimizer.zero_grad()
        loss = loss_fn(model(xb), yb)
        loss.backward()
        optimizer.step()
        scheduler.step()

    model.eval()
    correct = 0
    with torch.inference_mode():
        for xb, yb in val_loader:
            logits = model(xb.to(dev))
            pred = logits.argmax(dim=1).cpu()
            correct += (pred == yb).sum().item()

Interview tip: Interviewers look for the details: zero_grad before backward, model.eval() for validation, no gradient tracking during evaluation, .item() for metrics, and knowing whether the scheduler steps per batch or per epoch for the schedule you chose.

40. How do Dataset and DataLoader work, and how do you stop data loading becoming the bottleneck?

Answer: A map-style Dataset implements __len__ and __getitem__ to return one sample; an iterable-style dataset streams samples, which suits very large or remote data. DataLoader batches samples, shuffles (training only), and uses a collate_fn to combine samples, for example padding variable-length sequences. To keep the GPU fed: use num_workers greater than zero so loading and augmentation run in parallel processes, pin_memory=True with non_blocking=True transfers for faster host-to-GPU copies, persistent_workers=True to avoid restarting workers every epoch, and store data in formats that read quickly (pre-decoded or sharded files rather than millions of tiny files on network storage).

loader = DataLoader(
    train_ds, batch_size=64, shuffle=True,
    num_workers=8, pin_memory=True,
    persistent_workers=True, drop_last=True)

41. What is the difference between model.eval(), torch.no_grad() and torch.inference_mode()?

Answer: They solve different problems and you usually need both kinds. model.eval() changes layer behaviour: dropout turns off and batch norm uses its running statistics instead of batch statistics. It does not disable gradient tracking. torch.no_grad() disables gradient tracking, saving memory and compute, but does not change layer behaviour. torch.inference_mode() is a stricter, slightly faster version of no_grad for pure inference; tensors created inside it cannot later be used in autograd. So for validation and serving: model.eval() plus inference_mode() or no_grad(). And remember to call model.train() again before the next training epoch.

42. How do you save and load models safely in PyTorch?

Answer: Save the state_dict (a dictionary of tensors), not the whole pickled model object, which ties the file to your exact code layout. For resumable training, save a checkpoint with model, optimiser, scheduler and gradient scaler states, the epoch or step, and RNG states. When loading, the security point matters: torch.load uses Python pickle, and unpickling an untrusted file can execute arbitrary code. Since PyTorch 2.6 the default is weights_only=True, which restricts loading to tensors and a safe set of types; keep it that way and never set it to False for files you didn't produce. For sharing weights, the safetensors format stores raw tensors with no executable code. Load to CPU first with map_location when moving between machines.

torch.save({"model": model.state_dict(),
            "optim": optimizer.state_dict(),
            "step": step}, "ckpt.pt")

ckpt = torch.load("ckpt.pt", map_location="cpu",
                  weights_only=True)
model.load_state_dict(ckpt["model"])

Production consideration: Treat model files like code artefacts in your supply chain: record where they came from, scan or verify them, and store hashes. The MLOps interview questions cover model registries and versioning.

43. How do you make PyTorch training reproducible, and what are the limits?

Answer: Seed Python, NumPy and PyTorch (CPU and CUDA), seed DataLoader workers and pass a seeded generator for shuffling, and record library versions, hardware and the exact config. For bit-exact results on GPU, enable deterministic algorithms (torch.use_deterministic_algorithms(True)), disable cuDNN benchmark autotuning, and set any environment variables the docs require for deterministic cuBLAS behaviour. Limits: determinism can cost speed, some operations have no deterministic implementation, and results can still differ across GPU models, driver or library versions, and different numbers of GPUs. In practice, aim for statistical reproducibility: several seeds landing in a tight range.

44. What does torch.compile do, and when would you use it?

Answer: torch.compile(model) captures the model's Python code into graphs and compiles them into optimised kernels, fusing operations to cut memory traffic and Python overhead, while the code stays ordinary eager PyTorch. It often speeds up training and inference, especially for models with many small element-wise operations. Costs: the first iterations are slow while compiling, dynamic shapes or data-dependent Python control flow can cause graph breaks and recompilation, and debugging is harder. Use it once the model is correct in eager mode, and benchmark on your real shapes.

Mixed precision, GPU memory and distributed training

45. How does mixed-precision training work? Compare FP16 and BF16.

Answer: Mixed precision runs most operations, particularly matrix multiplies and convolutions, in a 16-bit format to use tensor cores and halve activation memory, while keeping master weights and numerically sensitive operations (reductions, softmax, loss) in FP32. In PyTorch, torch.autocast picks the precision per operation. FP16 has a narrow exponent range, so small gradients underflow to zero; a gradient scaler multiplies the loss before backward and unscales before the step, skipping steps that produce infinities. BF16 keeps FP32's exponent range with less mantissa precision, so it rarely needs loss scaling and is the usual choice on GPUs that support it. If you clip gradients with FP16, unscale first.

scaler = torch.amp.GradScaler("cuda")
with torch.autocast("cuda", dtype=torch.float16):
    loss = loss_fn(model(xb), yb)
scaler.scale(loss).backward()
scaler.unscale_(optimizer)      # before clipping
torch.nn.utils.clip_grad_norm_(model.parameters(), 1.0)
scaler.step(optimizer)
scaler.update()

46. What consumes GPU memory during training, and how do you estimate it?

Answer: Four main buckets: weights; gradients (same size as weights); optimiser state (Adam keeps two extra values per parameter, usually in FP32); and activations saved for backward, which grow with batch size, sequence length or image resolution, and depth. Plus framework workspace and fragmentation. As a rough rule for Adam with mixed precision, the weight-related part alone is in the order of 16 bytes per parameter (FP32 master weights, FP32 moments, 16-bit weights and gradients), before any activations, which is why fine-tuning needs far more memory than inference. A 1-billion-parameter model therefore needs roughly 16 GB just for that state. Measure rather than guess: torch.cuda.max_memory_allocated() after a few steps, and the memory profiler for detail. See GPU basics for AI engineers for how GPU specifications map to these numbers.

47. You hit CUDA out-of-memory. What are your options, in order?

Answer: Work from cheapest to most invasive:

  1. Check for leaks first: accumulating tensors that hold the graph, evaluation running without no_grad, or caching outputs on the GPU.
  2. Turn on mixed precision (BF16 or FP16).
  3. Reduce the batch size and use gradient accumulation to keep the same effective batch: run several micro-batches, divide each loss by the number of steps, and call optimizer.step() once.
  4. Use activation checkpointing, which discards some activations and recomputes them during backward, trading compute for memory.
  5. Reduce sequence length or image resolution if the task allows.
  6. Use memory-efficient optimisers or parameter-efficient fine-tuning so most weights need no gradients or optimiser state.
  7. Shard state across GPUs with FSDP or ZeRO-style approaches (Q49).

Remember that batch norm statistics depend on the micro-batch size, so accumulation is not fully equivalent for BN models.

48. What is the difference between DataParallel and DistributedDataParallel?

Answer: Both are data parallelism: each GPU holds a full model copy and processes a different slice of the batch. DataParallel is single-process and multi-threaded: one GPU scatters inputs, gathers outputs and holds the master copy, so it suffers from Python's global interpreter lock and an overloaded primary GPU. DistributedDataParallel (DDP) runs one process per GPU; each computes its own gradients, and they are averaged with an all-reduce operation that overlaps with the backward pass. DDP is faster, scales across machines and is the recommended approach. Practicalities: launch with torchrun, use a DistributedSampler so each process sees different data (and call set_epoch each epoch so shuffling changes), save checkpoints from rank 0 only, and remember that the effective batch size is the per-GPU batch times the number of GPUs.

49. When do you need more than data parallelism? Explain FSDP/ZeRO, tensor and pipeline parallelism.

Answer: Plain data parallelism requires each GPU to hold the full model, gradients and optimiser state. When that doesn't fit:

  • Sharded data parallelism (FSDP, ZeRO): shard parameters, gradients and optimiser state across GPUs and gather each layer's full weights only when needed. Often the first step beyond DDP because the model code barely changes.
  • Tensor parallelism: split individual large matrix multiplies across GPUs; it needs very fast interconnect, so it is usually kept within one server.
  • Pipeline parallelism: place different layers on different GPUs and stream micro-batches through them, accepting some idle "bubble" time.

Very large models combine all three. All of them trade communication for memory, so interconnect bandwidth becomes a design constraint. For most enterprise fine-tuning, DDP or FSDP on a single multi-GPU server is enough.

Generative models

50. How does a variational autoencoder (VAE) work, and why is it "variational"?

Answer: A plain autoencoder's latent space has gaps, so random codes decode to garbage. A VAE's encoder outputs the parameters of a distribution (a mean and variance) over the latent code rather than a single point. Training maximises a lower bound on the data likelihood (the ELBO): a reconstruction term plus a KL-divergence term that pulls each encoded distribution towards a simple prior, usually a standard normal. This makes the latent space smooth enough to sample from. The reparameterisation trick, writing the sample as mu + sigma * eps with random eps, makes the sampling step differentiable. VAEs train stably but samples tend to be blurrier than GANs or diffusion; their encoders are widely used to compress images into latent spaces where diffusion models then operate.

51. How do GANs work, and why are they hard to train?

Answer: A generator maps random noise to samples; a discriminator tries to tell real samples from generated ones. They train adversarially: the discriminator improves at detection, and the generator improves at fooling it. When it works, samples are sharp and generation is a single fast forward pass. Training is a two-player game with no single loss to minimise, so it can oscillate or diverge. Mode collapse is the classic failure: the generator produces a narrow set of outputs that fool the discriminator and ignores the rest of the data distribution. If the discriminator becomes too strong, the generator's gradients vanish. Remedies include alternative losses, gradient penalties and spectral normalisation. For many image generation tasks, diffusion models have displaced GANs because they train more reliably and cover the data distribution better, while GANs remain useful where single-step speed matters.

52. Explain diffusion models conceptually, and compare them with VAEs and GANs.

Answer: A diffusion model defines a forward process that gradually adds Gaussian noise to data over many steps until only noise remains, then trains a network (often a U-Net or transformer) to reverse one step: given a noisy input and the noise level, predict the noise that was added (or an equivalent target). This simple regression objective is why training is stable. Generation starts from noise and repeatedly denoises. Conditioning, such as a text prompt embedding injected through cross-attention, steers the output; classifier-free guidance strengthens the conditioning. Latent diffusion runs the process in a VAE's compressed latent space for efficiency. The cost is inference speed: many denoising steps, which is reduced by better samplers and distillation into few-step models.

ModelTraining stabilitySample qualitySampling speedDiversity
VAEStableOften blurryFast (one pass)Good
GANFragileSharpFast (one pass)Risk of mode collapse
DiffusionStableHighSlow (many steps) unless distilledStrong

Production consideration: Enterprises deploying image generation also need provenance labelling; see AI content provenance and watermarking.

If you want to go from answering these questions to training, evaluating and deploying models on real cloud infrastructure, the APEX AI, ML, Cloud and Cyber Security program covers the ML and deep learning foundations alongside cloud and security.

Real-world scenario questions

53. Scenario: your classifier's loss sits at about ln(number of classes) from the first step and never moves. What is happening?

Answer: A loss equal to ln(C) means the model predicts a uniform distribution over C classes, so it is learning nothing at all. This points to a pipeline bug rather than a tuning problem.

What I would check:

  1. Whether gradients reach the parameters: are any .grad values non-zero, and are the parameters actually passed to the optimiser (unregistered layers in a Python list, a frozen backbone with no head trainable)?
  2. Whether optimizer.step() and zero_grad() are called in the right order.
  3. Whether the learning rate is zero or absurdly small, for example a scheduler warmup bug.
  4. Whether labels match inputs: shuffling inputs and labels separately destroys the signal.
  5. Whether the output layer or loss is wrong, such as applying softmax before CrossEntropyLoss or a final ReLU clamping logits.
  6. Whether the model can overfit a single small batch. If it cannot memorise ten examples, the bug is in code, not capacity.

Production consideration: Build the "overfit one batch" test into your training template so it runs before every long job on shared GPUs.

54. Scenario: training runs fine for hours, then the loss suddenly becomes NaN. How do you debug it?

Answer: A late NaN usually means numerical instability or a bad data sample, not a fundamentally wrong model.

What I would check:

  1. Logs for loss spikes and rising gradient norm in the steps before the NaN; if the norm explodes, lower the learning rate, add or tighten clipping, or lengthen warmup.
  2. The batch at failure: corrupted inputs, NaN or infinite values in features, empty sequences or division by zero in a custom normalisation.
  3. Unsafe operations: log(0), sqrt of negatives, division without an epsilon, or a hand-written softmax instead of the stable built-in.
  4. FP16 overflow: whether the gradient scaler is in use, whether BF16 is an option, and which layer first produces infinities (forward hooks or anomaly detection on a short reproduction run, since it is slow).

Production consideration: Checkpoint regularly, have the training loop stop or skip a step on non-finite loss instead of corrupting weights, and resume from the last good checkpoint with a lower learning rate.

55. Scenario: training accuracy reaches near-perfect but validation accuracy stalls well below. What do you do?

Answer: This is classic overfitting, but confirm the validation set is sound before adding regularisation.

What I would check:

  1. That validation data comes from the same distribution as production and has no labelling inconsistencies; inspect the errors.
  2. Train and validation curves over time: when did the gap open, and did validation loss start rising (keep the best checkpoint)?
  3. More effective data: stronger and appropriate augmentation, more labelled data, or self-supervised pretraining on unlabelled data.
  4. Regularisation: weight decay, dropout in the head, label smoothing, stochastic depth.
  5. Capacity: a pretrained backbone with fewer trainable layers instead of a large model from scratch.

Production consideration: On small datasets, report results across several seeds and splits before claiming a fix worked.

56. Scenario: a model scores well in validation during training, but after loading the checkpoint for inference its accuracy drops sharply. Why?

Answer: The model is fine; the inference path is different from the validation path. This is one of the most common real-world deep learning bugs.

What I would check:

  1. Whether the serving code calls model.eval(); without it, dropout is active and batch norm uses the statistics of each (often tiny) inference batch.
  2. Preprocessing parity: resize method, crop, colour channel order, normalisation mean and standard deviation, tokeniser version and maximum length.
  3. Whether the right weights loaded: load_state_dict with strict=False hiding missing keys, or loading the last checkpoint rather than the best.
  4. Class index mapping: label order in training versus the mapping used in the API.
  5. Precision changes, such as exporting to a quantised or different runtime without re-validating.

Production consideration: Package preprocessing with the model and run a golden-set test in CI that compares serving outputs to training-time predictions on fixed examples.

57. Scenario: you fine-tune a pretrained CNN with batch size 4 because the images are high resolution, and results are unstable. What do you change?

Answer: With four images per batch, batch norm statistics are very noisy, and that noise differs between training and inference. Fix the normalisation before touching other hyperparameters.

What I would check:

  1. Freezing batch norm layers (keep them in eval mode, using pretrained running statistics) during fine-tuning.
  2. Replacing batch norm with group norm if training from scratch.
  3. Synchronised batch norm across GPUs if training with DDP, to compute statistics over the global batch.
  4. Mixed precision and activation checkpointing to fit a larger batch.
  5. Training on random crops at lower resolution, then fine-tuning briefly at full resolution.

Production consideration: Document which normalisation mode the shipped model uses; it matters when someone fine-tunes it again.

58. Scenario: you move from one GPU to eight with DDP, but training is barely faster. What is going on?

Answer: Something other than GPU compute is the bottleneck, or the communication cost dominates.

What I would check:

  1. GPU utilisation per device: low utilisation usually means the data pipeline cannot keep up; profile the DataLoader, increase workers, move decoding off the critical path, use faster storage.
  2. Whether the per-GPU batch is so small that all-reduce communication outweighs compute; increase it if memory allows.
  3. Interconnect: GPUs communicating over slow links or across nodes with limited network bandwidth.
  4. Synchronisation points: frequent .item() or .cpu() calls, logging every step, evaluation running on every rank.
  5. Whether the measurement is fair: eight GPUs process eight times the samples per step, so compare samples per second or time to a target metric, not step time.

Production consideration: On rented cloud GPUs, idle time is paid time; a short profiling run before a multi-day job is cheap. GPU basics for AI engineers covers interconnect and utilisation monitoring.

59. Scenario: a retailer in Hyderabad wants a product image classifier for several thousand categories, with only a few dozen labelled images for many of them. How would you approach it?

Answer: Lean on pretrained representations and treat the long tail as a retrieval problem rather than training a large classifier from scratch.

What I would check:

  1. A strong pretrained image encoder (or an image-text model) as the backbone, evaluated first with a linear probe and nearest-neighbour lookup on frozen embeddings.
  2. For tail categories, embedding similarity against labelled reference images, which also lets the catalogue team add new categories without retraining.
  3. Fine-tuning the backbone on head categories with class-balanced sampling or loss reweighting.
  4. Augmentation that reflects real photos: lighting, backgrounds, partial occlusion, phone camera quality.
  5. Per-category metrics, not just overall accuracy, which hides tail failures.

Production consideration: Monitor new and seasonal products, which shift the distribution constantly, and schedule periodic refreshes of the reference embeddings. Our computer vision interview questions go deeper into this kind of design.

60. Scenario: a bank wants to detect unusual sequences of card transactions. Would you use an LSTM, a transformer or something else?

Answer: Start from the constraints, not the architecture. Strict latency and explainability needs often make gradient-boosted trees on aggregated features the baseline to beat. Sequence models help when the order and timing of events carry signal those aggregates miss.

What I would check:

  1. Latency budget per authorisation and whether state can be updated incrementally: an RNN or GRU carries a compact state per card, which is cheap to update on each transaction.
  2. History length needed: if short windows suffice, a small transformer over the last N transactions is accurate and parallel to train.
  3. Label delay and scarcity: fraud labels arrive late and are rare, so consider self-supervised pretraining on transaction sequences, then fine-tuning.
  4. Explainability requirements from risk and compliance teams; plan reason codes from the start.
  5. Offline evaluation with a time-based split that respects how the model will be used.

Production consideration: Run any deep model in shadow mode next to the existing system before it affects decisions. The machine learning interview guide covers fraud system design and imbalanced metrics.

61. Scenario: fine-tuning a large pretrained model for one epoch on a single GPU runs out of memory immediately. What do you do?

Answer: Estimate first: full fine-tuning with Adam in mixed precision needs roughly eight times the memory of the 16-bit weights alone, before activations, so the plan often has to change rather than just the batch size.

What I would check:

  1. Whether full fine-tuning is needed at all; parameter-efficient fine-tuning (LoRA or adapters) removes most gradient and optimiser memory.
  2. BF16 mixed precision and activation checkpointing.
  3. Batch size 1 with gradient accumulation, and shorter sequences or smaller crops.
  4. Quantised base weights with adapters (QLoRA-style) if it still doesn't fit, or a smaller distilled model if it meets the accuracy target.
  5. FSDP across several GPUs if full fine-tuning is genuinely required.

Production consideration: Record peak memory for each configuration so the next GPU request is arithmetic, not guesswork.

62. Scenario: a hospital's chest X-ray model performs well on the internal test set but poorly on images from a newly connected hospital. What happened?

Answer: Distribution shift from a different acquisition setup, and possibly shortcut learning: deep networks readily latch on to scanner-specific artefacts, text markers or positioning that correlate with labels in the training data.

What I would check:

  1. Image characteristics at the new site: scanner manufacturer, resolution, bit depth, windowing, preprocessing, patient positioning and population.
  2. Saliency or attribution maps to see whether the model looks at anatomy or at markers and borders.
  3. Preprocessing parity, including how raw medical images are converted and normalised.
  4. Performance by subgroup and by site, not one overall number.
  5. Remedies: harmonised preprocessing, augmentation that simulates scanner variation, training on data from multiple sites, and site-held-out validation as the standard test.

Production consideration: In clinical settings, a model update needs clinical validation and governance sign-off; the system should keep a clinician in the loop and log inputs for audit.

63. Scenario: two teammates train "the same" model and get noticeably different validation scores. How do you resolve the debate?

Answer: First find out whether the difference is real or noise.

What I would check:

  1. Diff the configs exactly: learning rate, schedule, batch size, number of GPUs (which changes the effective batch), precision, augmentation and data version.
  2. Library and driver versions, and whether cuDNN autotuning or non-deterministic kernels are enabled.
  3. Run three to five seeds of each configuration; if the ranges overlap, the difference is seed variance.
  4. Whether both evaluate on the same validation split with the same preprocessing and checkpoint selection rule (best versus last).

Production consideration: Use an experiment tracker that records config, code commit, data version and environment for every run, so these debates take minutes instead of days.

64. Scenario: an image classifier is accurate but too slow and large for a factory's edge camera device. How do you shrink it?

Answer: Combine architecture, compression and runtime choices, measuring accuracy and latency on the target hardware after each step.

What I would check:

  1. A smaller efficient architecture (depthwise separable convolutions, lower input resolution) as the student.
  2. Knowledge distillation from the accurate teacher to the small student.
  3. Post-training quantisation to 8-bit integers with a representative calibration set, or quantisation-aware training if accuracy drops too far.
  4. Structured pruning of channels, which yields real speedups, unlike unstructured sparsity on most hardware.
  5. Export to the device's optimised runtime and benchmark on the device itself, not a workstation GPU.
  6. Accuracy on the hard cases specifically, such as rare defect types, not only the average.

Production consideration: Plan for model updates on fleets of devices: versioning, staged rollout and a fallback. On-device and edge AI covers the operational side.

65. Scenario: an interviewer asks you to walk through how you would take a deep learning model from idea to production at an enterprise. How do you structure it?

Answer: Structure it as a sequence of decisions, each with a check, and keep returning to the business metric.

Problem + metric --> Data + labels --> Baseline
       |                                  |
       v                                  v
  Pretrained model --> Fine-tune --> Evaluate
                                         |
  Monitor <-- Deploy <-- Optimise <------+

What I would check:

  1. The decision the model supports, the metric that reflects it, and the cost of each error type.
  2. Data sourcing, labelling quality, consent and privacy, and a split strategy that mirrors production (time or site based).
  3. A simple baseline, then a pretrained model with transfer learning before anything custom.
  4. Training hygiene: tracked experiments, several seeds, the right checkpoint selection, failure analysis by slice.
  5. Optimisation for the serving target: precision, batch size, export format, latency and cost per prediction.
  6. Deployment with preprocessing packaged alongside the model, shadow or canary rollout, monitoring of input drift and prediction distributions, and a retraining and rollback plan.

Production consideration: Connecting a model to real systems, security controls and monitoring inside a customer's environment is the work of a Forward Deployed Engineer, and the FDE PRO program focuses on exactly that integration and deployment layer.

Key takeaways

  • Most deep learning techniques exist to keep gradients healthy: initialisation, non-saturating activations, normalisation, residual connections, gating and clipping.
  • AdamW with warmup and a decaying schedule is the default starting point for transformers; know why decoupled weight decay matters.
  • Batch norm and dropout behave differently in train and eval mode, and that difference causes many real production bugs.
  • Transformers replaced RNNs mainly for parallelism, short paths between positions and scaling, at the cost of quadratic attention and a growing KV cache.
  • Transfer learning and self-supervised pretraining are usually the first answer when labels are scarce.
  • GPU memory is weights, gradients, optimiser state and activations; mixed precision, accumulation, checkpointing and sharding each attack a different part.
  • Debug training systematically: overfit one batch, check gradients and data, then tune.

Interview preparation checklist

  • Write a PyTorch training and validation loop from memory, including eval mode, no-grad evaluation and checkpointing.
  • Compute convolution output shapes and parameter counts by hand for a small CNN.
  • Explain backpropagation on a two-layer network and why gradients vanish in deep plain networks.
  • Be able to compare SGD with momentum, Adam and AdamW, and sketch warmup plus cosine decay.
  • Explain batch norm versus layer norm and the small-batch problem.
  • Fine-tune a pretrained image or text model end to end and report results over several seeds.
  • Train with mixed precision and measure peak GPU memory before and after.
  • Run a two-GPU DDP job, or at least know the launcher, sampler and rank-0 checkpointing details.
  • Prepare one story about debugging a training failure: the symptom, the hypotheses and the fix.
  • Explain VAEs, GANs and diffusion in two minutes each, with one trade-off for each.
  • Revise the neighbouring guides: Python for AI interview questions for the language side and Hugging Face interview questions for working with pretrained models.

FAQ

What skills are required for a deep learning engineer interview?

You need solid Python and PyTorch, the maths behind gradients and optimisation at an intuitive level, working knowledge of CNNs, sequence models and transformers, and practical experience training, debugging and evaluating models. For production-facing roles, GPU memory, mixed precision, distributed training and deployment matter too.

Is PyTorch or TensorFlow more important for deep learning interviews in 2026?

PyTorch is the more common framework in research and in many industry teams, and most recent open-source models ship PyTorch code first. Some organisations still run TensorFlow or Keras in production, so check the job description, but concepts transfer between frameworks.

How much maths do I need for deep learning interviews?

You should be comfortable with matrix multiplication, derivatives and the chain rule, basic probability, and why softmax with cross-entropy works. Most interviews test whether you can reason with these ideas rather than derive long proofs.

How is a deep learning interview different from a machine learning interview?

Machine learning interviews cover classical algorithms, feature engineering, metrics and validation. Deep learning interviews go further into network architectures, optimisation behaviour, training at scale on GPUs and debugging training runs.

Do I need a GPU to prepare for deep learning interviews?

Not for most preparation. Small models train on a laptop CPU, and free or low-cost notebook GPUs and short cloud GPU sessions are enough for fine-tuning experiments and mixed-precision practice.

Which deep learning topics are most commonly asked in 2026?

Commonly asked topics include backpropagation and vanishing gradients, AdamW and learning-rate schedules, batch versus layer normalisation, residual connections, why transformers replaced RNNs, transfer learning, mixed precision, GPU memory, DDP and debugging scenarios such as NaN losses.

What projects should I build before a deep learning interview?

Build one fine-tuned vision model and one fine-tuned text model with proper evaluation, a small model trained from scratch to show you understand the training loop, and one project that is served behind an API with monitoring. Document the failures you debugged along the way.

Is deep learning a good career path for Indian engineers?

Deep learning skills are used across GCCs, product companies and services firms in Hyderabad, Bengaluru and other cities, in roles from ML engineering to AI platform and applied research. Pairing model skills with cloud and production engineering keeps your options broad.

How long does it take to prepare for a deep learning interview?

It depends on your background. An engineer who already knows Python and classical ML can usually cover the core topics and one or two projects in a few focused weeks; a fresher should plan for longer and spend most of the time building and debugging models.

Ready to turn deep learning knowledge into production-grade AI systems? Cloudsoft's APEX program for AI, ML, cloud and cyber security builds these foundations with hands-on labs, in classroom sessions in Ameerpet or live online. Call +91 96660 19191 to book a free demo.

Share𝕏infβœ‰
EnrollWhatsAppCall us