New batches starting this week Β· Limited seats

Recommendation Systems Interview Questions and Answers 2026 (55 Questions)

55 recommender system interview questions with model answers, from matrix factorisation and two-tower retrieval to learning-to-rank, position bias, bandits, LLMs and eleven Indian product design scenarios.

Recommendation systems interview questions 2026: 55 questions on collaborative filtering, two-tower retrieval, ranking, cold start and evaluation
Last updated Β· 44 min read Β· 9,670 words

Recommendation system interview questions in 2026 test whether you can design the whole loop, not just name an algorithm: how interaction data becomes labels, how a two-tower model and an ANN index pull candidates from millions of items, how a ranker orders them, and how you evaluate, explore and serve it all within a latency budget. This guide collects 55 high-value recommender system interview questions with model answers, from collaborative filtering and matrix factorisation to learning-to-rank, position bias, bandits, LLMs in recommendations, DPDP consent and eleven system design scenarios set in Indian products.

How to use this guide

Recommender interviews mix ML theory, ranking system design and product judgement. What interviewers commonly probe at each level:

  • Freshers and junior engineers: collaborative versus content-based filtering, implicit versus explicit feedback, matrix factorisation, precision@k and NDCG, and why cold start is hard.
  • Mid-level ML engineers: the candidate generation, ranking and re-ranking funnel, two-tower retrieval, negative sampling, feature engineering without leakage, learning-to-rank losses and offline evaluation design.
  • Senior and staff roles: position bias, feedback loops, exploration, off-policy evaluation, real-time features, latency budgets, fairness to sellers or creators, privacy, and how to debug a launch where the offline metric went up and the business metric went down.

Questions are numbered continuously. Practise scenarios with a fixed structure: objective, label, funnel, metrics, failure modes. If you need classic ML revision first, the machine learning interview questions guide covers validation, metrics and model selection.

Fundamentals

1. What problem does a recommender system solve, and how is it different from search?

Answer: A recommender predicts which items a user is likely to value next and orders a small slate of them, usually without an explicit query. Search starts from a stated intent (the query) and its main job is relevance to that intent; recommendation has to infer intent from history, context and similar users. That changes the problem in three ways. First, the label is indirect: you learn from clicks, watches and purchases rather than judged relevance. Second, the system shapes its own training data, because users can only react to what was shown. Third, success is measured over time, not on one result page.

2. Compare collaborative filtering, content-based filtering and hybrid approaches.

Answer: Collaborative filtering (CF) uses only the interaction matrix: users who behaved alike in the past will like similar items. It discovers non-obvious relationships (people who buy a trekking pole also buy a rain cover) and needs no metadata, but it cannot score items or users with no interactions and it drifts towards popular items. Content-based filtering represents items by attributes or embeddings of text and images and recommends items similar to what the user engaged with. It handles new items and is easy to explain, but it tends to recommend more of the same and depends on metadata quality. Hybrids combine both, either by blending scores, by switching (content for new items, CF otherwise) or, most commonly today, by feeding interaction-based and content-based features into a single model such as a two-tower retriever and a feature-rich ranker.

ApproachUsesStrengthWeakness
CollaborativeUser-item interactionsFinds hidden affinitiesCold start, popularity bias
Content-basedItem attributes, text, imagesNew items, explainableNarrow, metadata-dependent
HybridBoth, in one model or a blendCoverage plus accuracyMore features and pipelines

3. What is the difference between explicit and implicit feedback, and how do you turn implicit signals into training labels?

Answer: Explicit feedback is a stated preference: ratings, thumbs up, "not interested". It is clear but sparse and biased towards users who bother to rate. Implicit feedback is observed behaviour: impressions, clicks, dwell time, add-to-cart, purchase, watch completion, skips. It is abundant but noisy, and crucially it has no reliable negatives: an item the user never clicked may simply never have been seen. To build labels, start from impressions (what was actually shown), mark positives by the action that most closely reflects value (a purchase or a long watch beats a click), and treat shown-but-ignored items as weak negatives. Grade the label where useful (view, cart, purchase), weight by confidence, and filter accidental signals such as clicks followed by an immediate bounce. Explicit negatives such as "not interested" or returns deserve strong negative weight.

4. Explain matrix factorisation for recommendations. How do ALS and SGD differ?

Answer: Matrix factorisation approximates the sparse user-item matrix R as the product of two low-rank matrices: each user u gets a vector p_u and each item i a vector q_i, and the predicted affinity is their dot product, often plus user and item biases. Training minimises squared error on observed entries plus L2 regularisation on the vectors. SGD updates both vectors per observed interaction and is easy to extend. Alternating least squares fixes item vectors, solves each user vector in closed form, then swaps; the independent solves parallelise well. ALS is also the natural choice for implicit feedback, where the weighted formulation treats every unobserved cell as a zero with low confidence and observed interactions as ones with confidence that grows with the interaction count.

Interview tip: Mention that the learned vectors are embeddings. That is the bridge from classic MF to two-tower models, which replace the free per-ID vectors with neural towers over features.

5. How does item-to-item collaborative filtering work, and why is it still useful?

Answer: Item-item CF computes similarity between items from the users who interacted with both, for example cosine similarity between item columns of the interaction matrix, or a normalised co-occurrence count within sessions. To recommend, take the user's recent items, look up each one's nearest neighbours and aggregate scores. It remains useful because item relationships are more stable than user tastes, so the similarity table can be precomputed offline and served by key lookup; it updates the user's recommendations immediately as they browse; and it explains itself ("because you viewed X"). Normalise co-occurrence by item frequency, or popular items dominate every list.

6. What baselines would you build before any deep learning model?

Answer: Build a ladder of cheap baselines so every complex model has to earn its place. Global popularity over a recent window; popularity within segments such as city, language or category; recency of the user's own history (re-show recently viewed, which is hard to beat on some surfaces); item-item co-occurrence; and a matrix factorisation or simple ranker on a handful of features. Evaluate all of them on the same temporal split with the same metrics.

7. Why are sparsity and the long tail central problems in recommendations?

Answer: Most users interact with a tiny fraction of the catalogue, so the matrix is overwhelmingly empty, and interactions concentrate on a small head of popular items. Models trained on this data learn the head well and the tail badly, so they recommend popular items to everyone, which further starves the tail of data. Remedies include using content features so tail items share statistical strength with similar items, negative sampling corrections that stop the model penalising popular items merely for being frequent, popularity-aware regularisation or re-ranking, exploration that gives tail items measured exposure, and metrics such as catalogue coverage alongside accuracy.

8. How do factorisation machines and deep models extend matrix factorisation with side features?

Answer: Plain MF knows only IDs. Factorisation machines give every feature (user ID, item ID, category, city, device, hour) its own embedding and model pairwise interactions as dot products between those embeddings, so the model can generalise to combinations it never saw directly, such as a new item in a familiar category. Deep recommender architectures build on the same idea: embed sparse categorical features, combine them with dense features, and learn interactions with explicit cross layers, attention or MLPs. The trade-off is more data, more tuning and more serving cost; for many ranking problems gradient-boosted trees over well-engineered features remain a strong competitor.

Architecture: retrieval, ranking and re-ranking

9. Explain the multi-stage recommendation architecture.

Answer: No single model can score millions of items for every request within a latency budget, so production systems use a funnel where each stage is more expensive per item and sees fewer items.

request (user, context)
     |
     v
candidate generation   millions -> thousands
 - two-tower + ANN, item-item, trending,
   recently viewed, editorial pools
     |
     v
ranking                thousands -> hundreds
 - rich features, multi-task model
     |
     v
re-ranking             hundreds -> a slate
 - diversity, rules, dedup, stock, fairness
     |
     v
serve, log impressions, positions, outcomes

Candidate generation optimises recall: do not miss good items, cheaply. Ranking optimises precision at the top with features too expensive to compute for the whole catalogue, including user-item cross features. Re-ranking optimises the slate as a whole and applies policy.

10. How does a two-tower retrieval model work and how is it trained?

Answer: A two-tower model has a query tower that encodes the user and context (ID embedding, recent items, demographics where permitted, time, device) and an item tower that encodes the item (ID embedding, category, price band, text and image embeddings). Each produces a vector of the same size, and the score is their dot product or cosine similarity. Because the towers never see each other's features until the final dot product, item vectors can be precomputed and indexed, and only the query vector is computed at request time. Training typically treats it as classification over items: for each positive pair, the other items in the batch act as negatives (in-batch softmax), sometimes plus sampled or hard negatives. In-batch negatives over-represent popular items, so a log-frequency correction (subtracting the log of each item's sampling probability from its logit) is commonly applied to avoid unfairly penalising popular items.

11. How is approximate nearest neighbour search used in retrieval, and what operational issues come with it?

Answer: After training, every item vector from the item tower is loaded into an ANN index (HNSW graphs, IVF partitions, often with quantisation), and at request time the query vector retrieves the top few hundred or thousand items by similarity in milliseconds. Operational issues: item vectors and the query tower must come from the same model version, so index rebuilds and model deploys must be coordinated (blue-green indexes work well); new items need vectors quickly, so the item tower must run in the ingestion pipeline; filters such as "in stock in this pincode" or "Telugu only" interact badly with pure post-filtering, so partitioned indexes or filtered search are used; and ANN recall should be measured against exact search on a sample. For index internals and parameter tuning, see the vector database interview questions.

12. What negative sampling strategies exist, and what can go wrong?

Answer: Random negatives (uniform over the catalogue) are cheap but too easy; the model learns only to separate obvious mismatches. In-batch negatives reuse other positives in the mini-batch, which is efficient but popularity-skewed, hence the frequency correction. Hard negatives are items that score highly but were not engaged with, such as items shown and skipped, or neighbours retrieved by the previous model; they sharpen discrimination at the top. Mixed strategies combine easy and hard. What goes wrong: false negatives (an item the user would have liked but never saw, especially in hard-negative mining from unseen items), over-hard negatives that destabilise training, and a negative distribution that does not resemble what the ranker will actually face. Evaluate with a retrieval recall metric, not training loss.

13. Why use several candidate sources, and how do you merge them?

Answer: Each source captures a different reason a user might want an item: personalised embeddings capture long-term taste, item-item captures current session intent, trending captures what is happening now, and editorial or business pools carry new launches or campaigns. Relying on one source makes the system brittle and narrow. Merge by union with per-source quotas and deduplication, then let the ranker score everything on one scale with the source as a feature. Quotas protect fresh sources from being crowded out; source-level logging shows which sources earn their latency.

14. What features would you engineer for a ranking model, and how do you avoid leakage?

Answer: Group features into user (tenure, activity level, preferred categories, price sensitivity), item (category, price, age, quality signals, historical click-through and conversion rates), context (time of day, device, location, surface, session depth), and user-item cross features (has the user viewed this item or brand, similarity between item embedding and user's recent items, time since last interaction with the category). Counters over multiple windows (last hour, day, month) capture both trends and stable preference. Leakage is the classic failure: every feature must be computed as of the moment of the impression, never using data from after it. A feature store with point-in-time joins, or logging features at serving time and training on those logs, prevents both leakage and training-serving skew.

15. Which ranking models are used, and why are rankers often multi-task?

Answer: Common choices are gradient-boosted trees over engineered features (strong, fast, interpretable feature importance) and deep rankers that learn from sparse IDs, sequences and dense features together. Rankers are usually multi-task because a single click objective is a poor proxy for value: the model predicts several outcomes such as click, add-to-cart, purchase, long watch, share or hide, using shared lower layers and separate heads. A final score combines them, for example a weighted sum of predicted probabilities or an expected-value formula, and the weights are business decisions tuned through online experiments.

16. What does a re-ranking stage do, and why keep it separate from the ranker?

Answer: The ranker scores items independently; the re-ranker builds a good slate. It enforces diversity (for example maximal marginal relevance, which trades each item's score against its similarity to items already chosen), removes duplicates and near-duplicates, applies availability and eligibility rules (stock, serviceable pincode, age restrictions), caps repeated sellers or brands, inserts sponsored or editorial slots under policy, and can apply exposure constraints for fairness. Version and log every rule so undocumented rules do not quietly override the model.

Learning to rank

17. Explain pointwise, pairwise and listwise learning to rank.

Answer: Pointwise methods treat each user-item pair independently and predict a label or probability, using regression or log loss. They are simple and produce calibrated scores, but ignore that only relative order matters. Pairwise methods learn from pairs within the same request or query: the item the user engaged with should score above the one they skipped, with losses such as a logistic loss on the score difference. They match the ranking goal better and are less sensitive to label scale. Listwise methods optimise a loss defined over the whole list, such as a softmax cross-entropy over items in the request or surrogates of NDCG, so they focus effort on the top positions. In practice, pointwise is common in recommendation rankers because predicted probabilities are needed downstream; pairwise and listwise losses help when the order at the top matters most and labels are graded.

FamilyUnit of trainingGood for
PointwiseOne itemCalibrated probabilities, multi-task value formulas
PairwiseTwo items, same requestRelative preference, noisy label scales
ListwiseWhole listTop-heavy metrics such as NDCG

Interview tip: The neighbouring search relevance interview questions go deeper on LTR for query-driven ranking.

18. What is the intuition behind LambdaRank and LambdaMART?

Answer: NDCG is not differentiable because it depends on sorted positions. LambdaRank sidesteps this by defining gradients directly: for each pair of items with different relevance, it uses a pairwise gradient scaled by how much NDCG would change if the two items swapped positions. Swaps near the top of the list change NDCG a lot, so those pairs get larger gradients, and the model focuses its capacity on getting the top right. LambdaMART applies these lambda gradients inside gradient-boosted trees, which made it a long-standing strong baseline for learning to rank and is available in common GBDT libraries as a ranking objective.

19. When does score calibration matter in a recommender?

Answer: If scores are only used to sort, calibration matters little. It matters when scores are combined or compared on an absolute scale: multi-task value formulas that add predicted click and purchase probabilities with business weights, blending organic and sponsored items by expected value, thresholds such as "only send a push notification if predicted open probability is high enough", and comparisons across surfaces or model versions. Negative downsampling during training shifts predicted probabilities and must be corrected. Check calibration with reliability plots per segment, and fix with post-hoc methods such as isotonic or Platt scaling on recent data. Pairwise and listwise models produce scores that are not probabilities, which is one reason pointwise heads are common.

Session and sequence models

20. What are session-based recommendations, and what approaches work?

Answer: Session-based recommendation predicts the next item from the current session's sequence of actions, often for anonymous or logged-out users with no long-term history. Approaches, from simple to complex: co-visitation counts (items viewed together within sessions), item-to-item similarity aggregated over the session with recency weighting, Markov-style next-item transitions, recurrent models that read the click sequence, and transformer models with self-attention over the sequence. Simple co-visitation and nearest-neighbour session methods are cheap and surprisingly competitive, so benchmark them before a neural model.

21. How do transformer-based sequential recommenders work conceptually?

Answer: They treat a user's interaction history as a sequence of item tokens, like words in a sentence. Each item is embedded, positional or time information is added, and self-attention layers learn which past items matter for predicting the next one. Two common training styles: causal (unidirectional) models predict the next item from all previous items, matching how the model is used at serving time; masked (bidirectional) models hide random items in the sequence and predict them from both sides, which gives more training signal per sequence. The final hidden state acts as a user representation that can feed a retrieval index or a ranker. Practical concerns: sequence length truncation for long histories, incorporating action types (view versus purchase) and time gaps, popularity bias in the output softmax over a huge catalogue (handled with sampled softmax), and serving cost when sequences change with every click.

22. How do you combine long-term preferences with short-term session intent?

Answer: Represent both and let the model weigh them. Long-term preference comes from a profile embedding or aggregated features computed in batch (favourite categories, price band, brands). Short-term intent comes from the current session's items, encoded with a sequence model or a recency-weighted average of item embeddings. Combine them as separate inputs to the query tower and ranker, possibly with a learned gate so a strong session signal (five views of running shoes in two minutes) overrides the long-term profile, while a weak session falls back to it. Also retrieve from both: one candidate source driven by the session, one by the profile.

LLMs in recommendations

23. Where do LLMs realistically fit in a recommender system in 2026?

Answer: Mostly around the core ranking loop rather than replacing it. Large language models are expensive per call and slow relative to a ranker scoring hundreds of items in milliseconds, so the common, defensible uses are: item understanding (extracting attributes, normalising messy seller titles, tagging content, generating text for content embeddings), better cold-start representations for new items and users, natural-language explanations of recommendations, conversational recommendation where the user states needs in their own words, and offline work such as synthetic evaluation queries or labelling. LLM-derived embeddings and features then feed conventional retrievers and rankers. Fully generative recommenders are an active area (Q25); treat claims that LLMs have replaced ranking pipelines with caution.

24. How would you generate "why you are seeing this" explanations, and what is the main risk?

Answer: The explanation must be grounded in the signals that actually produced the recommendation, otherwise it is a plausible story rather than an explanation. Start with structured reasons the system can prove: the candidate source (similar to an item you viewed), the top contributing features, or the session item that triggered retrieval. Templates turn those into text reliably. An LLM can make the wording more natural or personalised, but it must be constrained to the provided reasons and item facts, with checks that it does not invent attributes, prices or claims. The main risk is unfaithful explanations: fluent text that does not reflect the model's real reasoning, which misleads users and, in regulated or sensitive contexts, can become a compliance problem. See LLM hallucinations explained for why grounding matters.

25. What is generative retrieval for recommendations, described carefully?

Answer: In generative retrieval, instead of embedding the user and searching an index, a sequence model directly generates the identifier of the next item. Because raw item IDs carry no meaning and catalogues are huge, research approaches assign each item a short "semantic ID": a sequence of discrete codes produced by quantising its content embedding (for example with residual quantisation), so similar items share prefixes. A transformer is trained on users' histories expressed as these code sequences and decodes the next item's codes autoregressively, often with beam search to produce several candidates. Open issues: decoding can produce codes that map to no item (constrained decoding addresses this), the code assignment must be updated as the catalogue changes, latency and cost of autoregressive decoding, and controllability. It is an evolving technique, so present it as an option to evaluate against a two-tower baseline, not as settled practice.

26. What are the risks of using LLMs inside a recommender?

Answer: Hallucinated items or attributes when an LLM is asked to recommend from memory instead of from the catalogue (always retrieve real items and let the LLM only choose among or describe them). Prompt injection through item content, because seller-written titles, descriptions and reviews are untrusted text that an LLM will read. Popularity and cultural bias inherited from pretraining, which can push well-known brands or English-language content. Cost and latency at feed scale, which usually pushes LLM work offline or to a small number of requests. Privacy: sending user histories to an external model API needs a lawful basis, data minimisation and contractual controls. Mitigate with catalogue-grounded generation, offline precomputation, caching, guardrails and the same online experimentation as any model change.

Cold start

27. How do you handle new-user cold start?

Answer: Use whatever is known at the first request, then learn fast. At first touch you typically have context: location (city or pincode), device, language setting, referral source or campaign, time of day and the landing page. Serve segment popularity and trending items for that context, with diversity across categories so early clicks are informative. Short, skippable onboarding (languages, interests) helps. Then update in-session: item-item and session models react to the first few clicks immediately, long before any batch retraining. Bandit-style exploration across categories (Q37) reduces the time to identify interests.

28. How do you handle new-item cold start?

Answer: A new item has no interactions, so ID-based methods cannot score it. Build its representation from content: text and image embeddings, category, brand, price and seller signals, fed through an item tower that does not depend on the ID alone, so it gets a meaningful vector at ingestion and lands in the ANN index within minutes. Give new items controlled exposure through an exploration budget or a fresh-items candidate source with a quota, and update their statistics quickly using streaming counters with smoothing toward category priors.

Evaluation: offline and online

29. Define precision@k, recall@k and hit rate, with an example.

Answer: For one user, precision@k is the number of relevant items in the top k divided by k. Recall@k is the number of relevant items in the top k divided by the total number of relevant items for that user. Hit rate@k is 1 if at least one relevant item appears in the top k, else 0. Average each across users. Example: a user later bought 4 items; your top 10 contains 2 of them. Precision@10 is 2/10 = 0.2, recall@10 is 2/4 = 0.5, and hit rate@10 is 1. Recall@k is the natural metric for candidate generation (did the right items survive the funnel?), while precision-style and position-aware metrics suit the final slate. Define relevance explicitly (purchase, long watch) and evaluate on a held-out future period.

30. Explain NDCG and MAP with a small calculation.

Answer: NDCG rewards relevant items more when they appear higher. DCG@k sums each item's relevance divided by log2(position + 1); NDCG divides by the DCG of the ideal ordering so it falls between 0 and 1. Example with binary relevance for a top 3 of [relevant, not, relevant]: DCG = 1/log22 + 0 + 1/log24 = 1 + 0.5 = 1.5. The ideal order [relevant, relevant, not] gives 1 + 1/log23, about 1.63, so NDCG@3 is about 0.92. Graded relevance (purchase 3, cart 2, click 1) fits naturally into NDCG. MAP computes, for each user, the average of precision at each rank where a relevant item appears, then averages across users. For the same list: precision at rank 1 is 1, at rank 3 is 2/3, so average precision is about 0.83 if those are the user's only relevant items. MAP is binary and also rewards early hits; MRR uses only the first relevant item.

31. Which beyond-accuracy metrics matter: coverage, diversity, novelty and serendipity?

Answer: Catalogue coverage is the share of items (or sellers, creators) that appear in anyone's recommendations over a period; low coverage means the long tail is invisible. Intra-list diversity measures how dissimilar items within one slate are, typically one minus the average pairwise similarity of their embeddings or categories. Novelty measures how unpopular (and therefore less likely already known) recommended items are, often via the negative log of item popularity. Serendipity tries to capture relevant items that the user would not have found by themselves, usually relevant items outside their usual categories. Track them as guardrails next to NDCG so the model cannot win by showing everyone the same popular items.

32. How do you design an offline evaluation split without leakage?

Answer: Use time. Train on interactions up to a cut-off date, validate on the following period and test on the period after that, mirroring how the model will be used. Random splits leak the future: the model learns from a user's later behaviour to predict earlier behaviour, and item popularity statistics include the test period. Leave-last-out (hold out each user's most recent interaction) is common in research but still needs a global time cut so features do not see the future. Recompute all features as of the cut-off. Evaluate against the full candidate set where feasible; metrics computed by ranking the true item against a small random sample of negatives can disagree with full-ranking metrics and can even change which model looks better. Report results by segment (new versus returning users, head versus tail items, language).

33. Why do offline improvements often fail to show up online?

Answer: Offline logs contain outcomes only for items the old system showed, at the positions it showed them, so a new model that surfaces different items is evaluated on data biased toward the old model. Other causes: the offline label (click) differs from the online goal (purchases, retention); position and presentation effects are absent offline; training-serving skew in features; latency increases that hurt engagement; interactions with re-ranking rules that override the new ranking; and offline gains concentrated in segments that matter little for business. Narrow the gap with counterfactual evaluation (Q39), randomised traffic, consistent feature logging, and tracking whether past offline gains predicted online wins.

34. How would you design an A/B test for a new ranking model?

Answer: Randomise by user (not by request) so each person sees a consistent experience and outcomes are independent across units. Choose one primary metric tied to the objective, such as purchases or completed watch time per user, plus guardrails: latency, revenue, returns or cancellations, complaint and hide rates, coverage and diversity, and metrics for sensitive segments. Estimate sample size and duration from the metric's variance and the minimum effect you care about, run full weekly cycles, and avoid stopping when the result first looks significant. Watch for novelty effects and marketplace interference (treatment users buying stock control users would have bought). Interleaving two rankers in one list is far more sensitive for quick screening, but it measures preference, not business impact, so screen with interleaving and decide with A/B.

Recommenders sit at the intersection of ML, data engineering and production systems. If you want structured, hands-on practice across that stack, Cloudsoft's APEX AI, ML, Cloud and Cyber Security program covers machine learning, cloud deployment and secure AI engineering with live labs, in Ameerpet or online.

Position bias, feedback loops and exploration

35. What is position bias and how do you correct for it?

Answer: Users click higher positions more regardless of relevance, so click logs confound "this item is good" with "this item was at the top". A model trained naively learns to reproduce the old ranking. Corrections: include position as a feature during training and set it to a fixed value at serving, so the model learns relevance separately from position; train a separate shallow position tower whose output is added to the relevance logit in training and dropped at serving; weight clicks by the inverse of the estimated probability that the position was examined (inverse propensity scoring); and estimate those examination probabilities from small randomised experiments, such as swapping adjacent results, or from click models.

36. What are feedback loops in recommender systems, and how do you mitigate them?

Answer: The model decides what users see, users can only engage with what they see, and that engagement trains the next model. Over time this amplifies whatever the system already favours: popular items get more exposure and more data, niche items starve, and user interests appear narrower than they are. Effects include popularity bias, homogenisation across users, filter bubbles, and degraded ability to learn about new items. Mitigations: dedicated exploration traffic and exploration slots, logging propensities so training can be debiased, popularity-corrected training objectives, diversity and coverage constraints in re-ranking, holdout groups that receive randomised or non-personalised recommendations to measure the loop's effect, and periodic audits of exposure concentration.

Interview tip: Interviewers like hearing that you would keep a small, permanent randomised slice of traffic. It costs a little engagement but gives unbiased data for evaluation and training.

37. Explain the explore-exploit trade-off and common bandit algorithms.

Answer: Exploitation shows the items the model currently scores highest; exploration shows items with uncertain value to learn about them. Pure exploitation never learns about new items or changing tastes; too much exploration wastes user attention. Multi-armed bandit algorithms formalise this. Epsilon-greedy shows the top item most of the time and a random one with a small probability; simple but explores blindly. Upper confidence bound methods pick the item with the highest optimistic estimate (mean plus an uncertainty bonus that shrinks with more data), so under-explored items get tried. Thompson sampling keeps a posterior distribution over each item's reward, samples from each and picks the highest sample, so exploration naturally concentrates on items that might plausibly be the top choice. Thompson sampling and UCB generally waste less traffic than epsilon-greedy. For the broader theory, see the reinforcement learning interview questions.

38. How are contextual bandits used in recommendations?

Answer: Contextual bandits choose an action given features of the user and situation, learning which item or strategy works for which context, while still exploring. Linear contextual bandits such as LinUCB model the reward as a linear function of context features per arm and add a confidence bonus; neural variants use a model's uncertainty estimate, for example by Thompson sampling over an approximate posterior. Typical uses are not the whole feed but bounded decisions: which banner or hero slot to show, which new items to try in a fresh-items slot, which of several ranking strategies or notification templates to use, or cold-start category exploration. They suit problems with immediate rewards; long-term outcomes such as retention are harder to optimise and evaluate safely.

39. What is off-policy evaluation and why does it need logged propensities?

Answer: Off-policy evaluation estimates how a new recommendation policy would have performed using logs collected under the old policy, without running an experiment. Inverse propensity scoring reweights each logged reward by the ratio of the new policy's probability of taking that action to the old policy's probability; it is unbiased if the old policy gave every relevant action a non-zero chance, but variance explodes when propensities are tiny. Self-normalised IPS and weight clipping trade a little bias for much lower variance. Doubly robust estimators combine a reward model with IPS, staying accurate if either is good. All of this requires logging, at serving time, the probability with which each shown item was chosen, which deterministic rankers do not have, so systems that want off-policy evaluation deliberately add stochastic exploration and log propensities. For slates, estimates are noisy, so use them to shortlist candidates for online tests.

Real-time features, freshness and serving

40. How do you build real-time features and keep recommendations fresh?

Answer: Freshness has three layers. Event freshness: stream user actions (views, carts, skips) through an event bus into a stream processor that updates session features and windowed counters (item clicks in the last hour, a user's category views today) in an online feature store within seconds. Model freshness: retrain or fine-tune incrementally on recent data on a schedule that matches how fast the domain changes, and refresh item embeddings for new items continuously. Catalogue freshness: propagate price, stock and eligibility changes to retrieval filters quickly, so you never recommend items that sold out an hour ago. The hardest part is consistency: the same feature must have the same definition offline and online, or you get training-serving skew. Logging the exact features used at serving time and training on those logs is the most reliable fix. The data engineering interview questions cover streaming pipelines in more depth.

41. How would you meet a tight latency budget for a recommendation API, and what are the fallbacks?

Answer: Start from the end-to-end budget for the page and allocate it per stage, for example (illustrative) around 10 ms for feature fetch, 15 ms for parallel candidate retrieval, 30 ms for ranking and 5 ms for re-ranking, leaving headroom for network and tail latency. Techniques: precompute what does not depend on the request (item embeddings, batch user profiles, even full candidate lists for less active users), fetch candidate sources in parallel with per-source timeouts, batch feature lookups, cap the number of items ranked, use a lighter pre-ranker before the heavy model, distil or quantise the ranker, run on hardware that suits the model, and cache results per user for short periods where freshness allows. Design degraded modes in tiers: if the ranker times out, return candidates ordered by retrieval score; if retrieval fails, return segment popularity; if everything fails, serve a static, cached trending list. Monitor p95 and p99 and how often each fallback fires.

Production consideration: For GPU-served deep rankers, the LLM inference and serving interview questions cover batching and hardware trade-offs that apply here too.

Fairness, filter bubbles and privacy

42. What does fairness mean in a recommender, and how do you measure it?

Answer: Recommenders are multi-sided, so fairness has several meanings. User-side fairness asks whether recommendation quality differs across user groups, for example lower relevance for users of regional languages, older devices or smaller towns. Provider-side fairness asks whether sellers, creators, employers or job seekers receive exposure proportionate to their relevance or quality, rather than exposure driven by early popularity. Measure quality metrics per user segment, exposure distribution and concentration across providers, and exposure relative to relevance for comparable providers. Interventions include re-ranking with exposure constraints, exploration budgets for new providers, debiased training, and removing proxies for protected attributes where outcomes affect livelihoods, such as jobs or credit. See AI bias and fairness testing for test methods.

43. What are filter bubbles, and what can engineers actually do about them?

Answer: A filter bubble is the narrowing of what a user sees because the system keeps showing more of what they already engaged with, which can also reinforce one-sided views in news and social content. Engineering levers: diversity and novelty constraints in re-ranking, explicit exploration of adjacent interests, separating short-term engagement from long-term satisfaction in objectives (for example by including survey or "not interested" signals and return visits), reducing the weight of low-effort engagement such as outrage clicks, and user controls such as "show less of this", topic follow and unfollow, and a non-personalised option. Test diversity changes against long-term retention, not only same-session clicks.

44. How does India's DPDP Act affect consent for personalisation?

Answer: Under the Digital Personal Data Protection Act, 2023 and the DPDP Rules, 2025 (notified in November 2025, with most business obligations applying from May 2027), processing personal data for personalisation generally needs consent that is free, specific, informed, unambiguous and given by a clear affirmative action, after a notice describing the data and purpose. For engineers that means: a consent record per user and purpose (personalised recommendations as a distinct purpose from, say, order fulfilment); pipelines that check consent before using behaviour data for training or serving; a non-personalised fallback for users who decline; withdrawal that is as easy as giving consent, with events that stop personalisation promptly and trigger erasure of data no longer needed; and data minimisation in features. For children (under 18), verifiable parental consent is required and tracking, behavioural monitoring and targeted advertising directed at children are not permitted, subject to exemptions in the Rules. This is an engineering summary, not legal advice: confirm with your privacy team and the official text. More detail is in DPDP Act for AI applications.

Interview tip: Mention that per-user embeddings and features should be removed promptly on withdrawal, and that regular retraining windows exclude withdrawn users.

System design and debugging scenarios

45. Scenario: design the home feed for an Indian e-commerce marketplace.

Answer: Clarify the objective first: completed orders or gross merchandise value per active user, with returns, cancellations and seller health as guardrails, not raw clicks. The home feed is a mix of widgets (continue shopping, recommended for you, deals, categories), so design both the per-widget recommendation and the widget ordering. Retrieval combines a two-tower model over user history and product content, item-item co-views from recent sessions, trending by city and category, and campaign pools during sales. A multi-task ranker predicts click, add-to-cart and purchase, with features such as price band affinity, delivery promise to the user's pincode, seller rating and return rate. Re-ranking enforces serviceability and stock, caps repeated brands, diversifies categories and places sponsored items under policy. The AI in retail and e-commerce article covers the wider use cases around this feed.

What I would check:

  1. Is the label purchase or a long-funnel signal, and how are returns and cancellations folded back into training?
  2. Are items unavailable at the user's pincode filtered before ranking, not after the slate is built?
  3. How does the feed behave for new users arriving from an ad campaign, and for users in tier-2 and tier-3 cities with sparse data?
  4. Does widget ordering have its own model or rules, and is it tested separately?

Production consideration: Sponsored placements and organic ranking need a clear, logged policy for how they mix, both for user trust and for debugging. Keep a small non-personalised holdout to measure the feed's true incremental value.

46. Scenario: design job recommendations for freshers on a hiring platform.

Answer: This is a two-sided, capacity-constrained problem. Freshers have thin profiles (degree, branch, college, city, skills listed, perhaps projects and certifications) and little history, while each job has limited openings and an expiry date. The objective is not clicks but good matches: applications that lead to shortlists or interviews, measured for both sides. Embed resumes and job descriptions, normalise skills to a taxonomy, and parse eligibility (batch year, degree, location) as hard filters. Retrieval combines semantic match, skill overlap and recruiter-side signals; ranking predicts apply and shortlist probabilities. Re-ranking must avoid congestion, where every fresher is shown the same few popular jobs and those recruiters drown in applications while others get none, so apply exposure caps per job and boost under-applied relevant jobs. Expired or filled jobs must disappear quickly.

What I would check:

  1. Are eligibility rules (graduation year, degree, location) enforced as filters before the model, so freshers do not waste applications?
  2. Is there a trust and safety layer for fraudulent postings, such as jobs asking for fees, before they reach recommendations?
  3. Do features like college name act as proxies that unfairly reduce exposure for candidates from smaller colleges or towns?
  4. Are recruiter responses (shortlist, reject, no response) flowing back as labels, and how is the delay handled?

Production consideration: Employment is a high-impact domain. Keep explanations simple and honest ("matches your Python and SQL skills"), audit outcomes across segments regularly and give candidates control over the preferences that drive recommendations.

47. Scenario: design course recommendations for an ed-tech platform.

Answer: Course purchases are rare, deliberate and goal-driven, so interaction data is sparse and the long purchase cycle makes clicks a weak label. Frame recommendations around goals: a learner preparing for a cloud job, a student revising for an exam, a working professional reskilling. Use stated goals from onboarding, skill assessments, current enrolment progress and content-based similarity between course syllabi as core signals. Surfaces include next module (progress-driven), next course on a learning path (prerequisite-aware) and discovery. The objective should include completion and learning outcomes, not just enrolment, so a course that learners abandon should not be pushed because it converts well. Re-ranking enforces prerequisites and level fit.

What I would check:

  1. Are any users under 18? If so, personalisation based on behavioural monitoring needs verifiable parental consent and may be restricted under the DPDP Rules, so design a compliant non-personalised or goal-based mode.
  2. Which label are we optimising: enrolment, completion, assessment improvement or renewal?
  3. Do we have a prerequisite graph, and who maintains it?
  4. How are free courses and paid courses balanced so the system does not optimise purely for revenue at the learner's expense?

48. Scenario: design a news recommender where freshness dominates.

Answer: News items lose value within hours, new items arrive constantly, and every article is effectively cold start. Represent articles by content: text embeddings of title and body, entities, topic, language, source and publication time, computed at ingestion so new articles are retrievable within minutes. Cluster articles into stories so the feed shows one version of a breaking event, not ten near-duplicates. Rank with recency decay features, real-time popularity counters (clicks in the last few minutes, normalised by impressions), user topic and source affinities, and session context. Editorial overrides, sensitive-event policy, source credibility and viewpoint diversity belong in re-ranking.

What I would check:

  1. What is the end-to-end delay from publication to first possible recommendation?
  2. Are popularity counters normalised by exposure, so an article shown in a top slot is not credited for position alone?
  3. How are misinformation, sensitive content and low-credibility sources handled before ranking?
  4. Is engagement weighted by dwell time so clickbait headlines are not rewarded?

49. Scenario: design recommendations for an Indian-language content platform (video, audio or short-form).

Answer: Language is the first-class dimension. Infer language preference from explicit settings, consumption history and device locale, and allow multiple languages per user (many users consume Hindi and Telugu, or Tamil and English). Content metadata is often code-mixed (Hinglish, transliterated Telugu in Latin script) or missing, so use multilingual text embeddings that handle transliteration, plus audio and visual embeddings where titles are poor. Creator-side fairness matters because new regional creators need exposure to grow.

What I would check:

  1. Is recommendation quality (NDCG, completion rate) reported per language and region, not only overall?
  2. How well do embeddings handle code-mixed and transliterated titles? Test on labelled samples per language.
  3. Are users ever shown content in languages they do not understand, and how quickly does the system correct after a skip?
  4. Is there a moderation layer that works in each supported language?

Production consideration: Speech and text models vary in quality across Indian languages; validate them per language before trusting their features. The NLP interview questions cover multilingual and code-mixed text handling.

50. Scenario: after launching a new ranker, click-through rate went up but orders and revenue went down. What happened?

Answer: The model is likely optimising a proxy that diverged from value. Common causes: the objective weights click too heavily, so the ranker promotes attractive but low-intent items (cheap accessories, clickbait images, items out of the user's price range); the model surfaces items that are unavailable or slow to deliver to the user's location; or curiosity clicks from a novelty effect.

What I would check:

  1. Click-to-order conversion and add-to-cart rate per item category and price band, treatment versus control.
  2. Share of recommended items that are out of stock or not serviceable for the user.
  3. Whether the multi-task score weights changed, and whether purchase predictions are still calibrated.
  4. Whether the effect is concentrated in one segment, surface or device.

Production consideration: Make the business metric the primary experiment metric and CTR a diagnostic, and add guardrails that block a launch automatically when conversion or revenue drops.

51. Scenario: your new model improves offline NDCG clearly, but the A/B test is flat. How do you investigate?

Answer: Treat it as a question about the offline metric's validity, not just the model. The offline set only contains outcomes for items the old system showed, so the new model may be better at reproducing the old ranking rather than finding new good items.

What I would check:

  1. How much of the final slate actually changed between treatment and control after re-ranking.
  2. Offline gains by segment and position versus where online traffic and value concentrate.
  3. Feature parity: are online features computed the same way as the offline ones the model was evaluated on?
  4. Latency differences between arms, since a slower page can cancel a better ranking.

Production consideration: Keep a record of every experiment's offline and online deltas. Over time this tells you which offline metrics predict online success for your product.

52. Scenario: the model was excellent in validation but recommendations look poor right after deployment. What is your diagnosis process?

Answer: Suspect training-serving skew before suspecting the model. Typical causes: a feature computed differently online (different time windows, time zones, default values for missing data, or a stale feature store table), feature leakage that inflated validation scores, a mismatch between the item index and the query tower model versions, or a preprocessing step such as vocabulary mapping that changed between training and serving so IDs map to the wrong embeddings.

What I would check:

  1. Compare logged online feature values with offline values for the same users and items at the same timestamps.
  2. Check the share of requests with missing or default features, and the share of IDs mapped to the unknown token.
  3. Confirm the ANN index and the query tower were built from the same model version.
  4. Check whether fallbacks are firing because of timeouts, so users are actually seeing popularity lists.

Production consideration: Shadow-deploy new models (score live traffic without serving) and compare score distributions and feature values before any user sees results. The MLOps interview questions cover shadow deployment and monitoring.

53. Scenario: new sellers on a marketplace complain that their products never appear in recommendations.

Answer: This is new-item cold start combined with a feedback loop. New listings have no engagement, so ID-based retrieval cannot find them and the ranker scores them low on missing engagement features; never shown means never learned about. Fix the representation and the exposure. Ensure the item tower uses content features so new products get meaningful embeddings at listing time; add a fresh-items candidate source with a quota; give new items a measured exploration budget (for example Thompson sampling with a category-level prior) and graduate them based on observed performance; and adjust features so missing history is distinguishable from bad history.

What I would check:

  1. Time from listing to the item's first appearance in the ANN index.
  2. Exposure distribution for items by age and by seller tenure, before and after changes.
  3. Whether listing quality (images, titles, attributes) is the real problem, which an LLM-based attribute extraction step could help with.
  4. Whether exploration is hurting user metrics, and how the budget is tuned.

54. Scenario: during a festival sale, recommendation API p99 latency breaches the SLO and pages time out. What do you do?

Answer: Stabilise first, then fix. Immediate actions: switch to a degraded mode that ranks fewer candidates or skips the heavy ranker, serve cached recommendations for active users, and enable the precomputed trending fallback for the overflow.

What I would check:

  1. Per-stage latency traces to locate the slow stage and whether it is CPU, memory, network or a dependency.
  2. Hot keys in the feature store, such as counters for a handful of deal items read by every request.
  3. Autoscaling lag versus the traffic ramp, and whether capacity was pre-provisioned for the sale.
  4. Timeouts and retries that amplify load instead of shedding it.

Production consideration: Before known peaks, load-test the full pipeline with realistic traffic, pre-warm caches and indexes, and decide in advance which features and stages can be shed. Treat the degraded mode as a tested feature, not an emergency hack.

55. Scenario: users complain that recommendations are repetitive and keep showing items they already bought.

Answer: The system lacks purchase-aware and repetition-aware logic. Some categories are repeat-purchase (groceries, diapers, mobile recharges), while others are one-off (a refrigerator, a laptop), and after buying a one-off item the user needs complements, not substitutes. Model repurchase cycles per category, suppress recently purchased one-off items and their close substitutes, promote complements (a phone case after a phone), and add frequency caps so the same item is not shown in every session after it has been ignored several times.

What I would check:

  1. Share of slates containing already-purchased items, split by repeat-purchase and one-off categories.
  2. How many times an item is shown to a user without engagement before it is suppressed.
  3. Whether order and return events reach the serving system quickly enough to suppress items.
  4. Intra-list and across-session diversity trends for affected users.

Production consideration: Add impression-level fatigue features (times shown recently, times ignored) to the ranker. They are cheap and often fix repetition more cleanly than hand-written rules.

Key takeaways

  • Start from the business objective and a carefully defined label; clicks are a diagnostic, not the goal.
  • Know the funnel: multi-source candidate generation for recall, a feature-rich multi-task ranker for precision, and re-ranking for slate quality and policy.
  • Two-tower models and ANN indexes make retrieval scale; frequency-corrected negatives and version-aligned indexes make them work.
  • Evaluate offline with time-based splits and position-aware metrics, track coverage and diversity, and decide with user-randomised A/B tests.
  • Logs are biased by what was shown and where: correct for position bias, keep exploration traffic and log propensities.
  • LLMs add the most value around the core loop today: item understanding, cold start, grounded explanations and conversational entry points.
  • Consent, fairness to providers and filter bubbles are design requirements, especially in jobs, education, news and children's products.

Interview preparation checklist

  • Implement implicit-feedback matrix factorisation and item-item CF on a public dataset, with a time-based split.
  • Compute precision@k, recall@k, MAP and NDCG by hand for a small list, then in code.
  • Build a small two-tower model with in-batch negatives and serve it through an ANN index; measure retrieval recall.
  • Draw the candidate generation, ranking and re-ranking funnel from memory with metrics for each stage.
  • Prepare one clear explanation each of position bias correction, Thompson sampling and inverse propensity scoring.
  • Rehearse two system design scenarios end to end, including cold start, latency budget, fallbacks and A/B design.
  • Prepare a debugging story: an offline-online gap or training-serving skew, and how you found it.
  • Revise neighbouring topics: AI system design interview questions for broader design rounds and the vector search questions for retrieval internals.

FAQ

What skills are required for a recommendation systems engineer role?

You need solid Python and SQL, machine learning fundamentals, ranking metrics, experimentation and A/B testing, and an understanding of data pipelines and model serving. Senior roles add system design, latency engineering, bias correction and the product judgement to choose the right objective.

How should I prepare for recommender system interview questions?

Build one small end-to-end recommender on a public dataset with a time-based split, practise explaining the retrieval, ranking and re-ranking funnel aloud, and rehearse two or three system design scenarios. Prepare stories about labels, leakage and offline-online gaps, because interviewers probe real experience.

Are recommendation system interviews mostly theory or system design?

Most loops include both. Expect fundamentals such as matrix factorisation and NDCG early, then a design round on a feed, search or marketplace problem where you must cover objectives, data, models, metrics, serving and failure modes.

Do I need deep learning to work on recommender systems?

You should understand embeddings, two-tower models and the basics of sequence models. Many production rankers still use gradient-boosted trees, so strong feature engineering and evaluation skills matter as much as deep learning expertise.

What is asked in a collaborative filtering interview?

Expect the difference between user-based and item-based methods, matrix factorisation with ALS or SGD, handling implicit feedback, cold start and popularity bias. Interviewers often ask you to compute a similarity or a metric by hand and explain its limitations.

What is a two-tower model interview question usually testing?

It tests whether you understand why separate user and item encoders allow precomputed item vectors and ANN retrieval, how in-batch negatives and frequency correction work, and why a ranker with cross features is still needed afterwards.

Can freshers get recommendation systems roles?

Freshers usually enter through ML engineer, data scientist or data engineer roles on personalisation teams. Two or three well-documented projects with proper evaluation, and clear explanations of the fundamentals, help more than listing many algorithms.

How are LLMs changing recommendation systems?

LLMs are mainly used for item understanding, cold-start representations, explanations and conversational recommendation, while efficient retrievers and rankers still handle most scoring. Generative retrieval is an active area, so evaluate it against strong baselines rather than assuming it wins.

Is recommendation systems a good career path for Indian engineers?

Personalisation is central to e-commerce, media, ed-tech, fintech and job platforms, and many such teams work from Hyderabad and Bengaluru, including in global capability centres. The skills also transfer to search, ads and ranking problems.

Ready to build ML and ranking systems end to end, from data and models to cloud deployment and monitoring? Explore Cloudsoft's APEX program for machine learning, AI, cloud and security skills. If you want to take AI systems into real customer environments, the AI Forward Deployed Engineer FDE PRO program focuses on enterprise integration and deployment, with placement support until you're placed. Classroom training is in Ameerpet, Hyderabad, or live online; call +91 96660 19191 for a free demo.

Share𝕏infβœ‰
EnrollWhatsAppCall us