New batches starting this week · Limited seats

Reinforcement Learning Interview Questions and Answers 2026 (50 Questions)

50 reinforcement learning interview questions with practical answers, from MDPs, Bellman equations and Q-learning to PPO, offline RL, bandits for A/B testing, RLHF and verifiable rewards, plus ten production scenarios.

Reinforcement learning interview questions 2026: 50 questions on MDPs, Q-learning, policy gradients, PPO, bandits and RLHF
Last updated · 44 min read · 9,762 words

Reinforcement learning interview questions in 2026 test whether you can reason about sequential decisions, rewards and exploration under real constraints, not whether you can recite the Bellman equation. This guide collects 50 high-value RL interview questions with model answers, from MDPs and TD learning through Q-learning, DQN, policy gradients and PPO to offline RL, bandits for A/B testing, RLHF and reward hacking, and ten production scenarios.

How to use this guide

RL comes up in three kinds of interviews: research-flavoured ML roles, applied roles where bandits and offline policy learning drive recommendations or operations, and LLM roles where RLHF and verifiable-reward training are now standard background. What interviewers commonly probe at each level:

  • Freshers: the agent-environment loop, returns and discounting, value versus policy, exploration, and a correct tabular Q-learning update.
  • Mid-level engineers: on-policy versus off-policy, why DQN needs replay and target networks, REINFORCE and baselines, PPO's clipping, and bandits versus A/B tests.
  • Senior engineers: offline RL and off-policy evaluation, reward design and hacking, RL for LLMs, safety constraints and an honest judgement of when RL is the wrong tool.

Answer each question aloud before reading the model answer. The code is short and illustrative; run it with python3 and change the hyperparameters to see what moves. Preference-tuning mechanics (DPO's beta, ORPO, KTO, preference-data pipelines) are covered in our fine-tuning LLM interview questions; here the RLHF questions are asked from the RL side.

RL fundamentals

1. What is reinforcement learning, and how is it different from supervised and unsupervised learning?

Answer: Reinforcement learning trains an agent to choose actions that maximise cumulative reward through interaction, where feedback is evaluative (how good was that?) rather than instructive (what was the right answer?). Supervised learning gets a labelled target for every input; unsupervised learning finds structure without any target. RL differs in three ways that matter in practice. First, the data depends on the agent's own choices, so a poor policy collects poor data. Second, rewards can be delayed, so the agent must work out which earlier action earned a later reward (the credit assignment problem). Third, the agent must balance trying new actions against using what it already knows.

Interview tip: If you are asked "could this be supervised learning instead?", take the question seriously. When a correct label exists for each decision, supervised learning is usually cheaper and more reliable.

2. Define agent, environment, state, action, reward and policy using a concrete example.

Answer: Take a warehouse replenishment system. The agent is the ordering logic. The environment is everything it does not control: demand, supplier lead times, the warehouse itself. The state is what the agent observes at each decision point, for example stock on hand, open orders and day of week. The action is the order quantity per SKU. The reward is a scalar signal after each step, such as margin earned minus holding cost minus a stock-out penalty. The policy maps states to actions (or to a probability distribution over actions). The loop is: observe state, act, receive reward and next state, repeat. Notice that the reward definition is a design decision with consequences, not a given.

3. What is a Markov decision process, and what happens when the Markov property does not hold?

Answer: An MDP is the tuple of states, actions, transition probabilities P(s' | s, a), a reward function and a discount factor. The Markov property says the next state and reward depend only on the current state and action, not on the full history. That assumption is what lets value functions be defined per state and lets Bellman equations work. When the agent's observation does not capture everything relevant (a call-centre router sees queue length but not caller mood; a trading agent sees prices but not other traders' intentions), the problem is a partially observable MDP. Practical fixes are to enrich the state with history (the last few observations, running averages), use a recurrent or transformer policy that keeps memory, or maintain an explicit belief state.

4. What is the return, and why do we discount future rewards?

Answer: The return Gt is the sum of rewards from time t onward, usually discounted: Gt = rt+1 + γ rt+2 + γ² rt+3 + … with γ between 0 and 1. Discounting does three jobs: it keeps the sum finite in continuing tasks that never end, it expresses a preference for sooner rewards (a rupee today versus a rupee next quarter), and it reduces variance because distant, uncertain rewards count less. γ also sets an effective planning horizon of roughly 1 / (1 − γ) steps, so γ = 0.99 looks about a hundred steps ahead and γ = 0.9 about ten. Too low and the agent becomes myopic, for example a pricing agent that maximises today's margin and erodes repeat purchases. Too high and learning becomes slow and noisy, because every value estimate depends on a long, uncertain future.

5. Explain state value, action value, advantage and policy. How do they relate?

Answer: The state value Vπ(s) is the expected return from state s when following policy π. The action value Qπ(s, a) is the expected return if you take action a in s and then follow π. The advantage Aπ(s, a) = Qπ(s, a) − Vπ(s) says how much better action a is than the policy's average behaviour in that state. The policy is the decision rule itself. Value-based methods learn Q and act greedily on it; policy-based methods learn the policy directly; actor-critic methods learn both, using a value estimate to judge the policy's choices. Advantage matters because it is centred around zero: telling the policy "this action was better than usual" is a much lower-variance learning signal than "this action led to a return of 412".

6. What is the exploration-exploitation trade-off, and what are the common strategies?

Answer: Exploitation chooses the action that currently looks most valuable; exploration tries others to learn whether they are actually better. Too little exploration locks in an early, wrong belief; too much wastes reward. Common strategies:

  • ε-greedy: act randomly with probability ε, usually decayed over time. Simple, but explores blindly.
  • Softmax / Boltzmann: sample actions in proportion to exp(Q / temperature), so near-ties are explored more than clearly bad actions.
  • Optimism and UCB: add a bonus for actions tried rarely, so uncertainty itself attracts exploration.
  • Thompson sampling: sample a plausible model from the posterior and act greedily on it.
  • Entropy bonuses and intrinsic rewards: keep deep RL policies stochastic or reward novelty.

Real-world example: In production, exploration has a cost someone pays: a customer sees a weaker offer, a call goes to a less suitable queue. That is why production systems cap exploration, restrict it to safe actions, and log the probability of every action taken.

7. What is the difference between on-policy and off-policy learning?

Answer: On-policy methods learn about the policy that is generating the data; off-policy methods learn about a target policy from data produced by a different behaviour policy. SARSA, REINFORCE and PPO are on-policy: once the policy changes, old data is stale and is discarded or only reused briefly. Q-learning and DQN are off-policy: they learn the greedy policy while behaving ε-greedily, and can reuse a replay buffer of old transitions. Off-policy learning is more sample-efficient and is what makes learning from logged historical data possible, but it is less stable and needs corrections (importance weights, or conservative objectives) when the behaviour and target policies differ a lot. In short: on-policy is simpler and steadier; off-policy reuses expensive experience.

8. Explain the Bellman equations intuitively.

Answer: The Bellman equations say that the value of where you are equals what you get now plus the discounted value of where you land. The expectation equation for a fixed policy: Vπ(s) = E[r + γ Vπ(s')], averaging over the policy's actions and the environment's transitions. The optimality equation replaces the average over actions with a maximum: Q*(s, a) = E[r + γ maxa' Q*(s', a')]. Intuitively it is a consistency condition, like a route planner: the shortest time from Ameerpet to the airport equals the time to the next junction plus the shortest time from that junction. That recursion is the foundation for almost every algorithm here: dynamic programming applies it with a known model, TD learning applies it to sampled transitions, and Q-learning applies the optimality version. The gap between the two sides, r + γV(s') − V(s), is the TD error, the learning signal itself.

9. What is dynamic programming in RL, and why is it rarely used directly in production?

Answer: Dynamic programming solves an MDP when the transition probabilities and rewards are fully known. Policy evaluation repeatedly applies the Bellman expectation equation to compute Vπ. Policy iteration alternates evaluation with greedy improvement until the policy stops changing. Value iteration folds the two together by applying the optimality equation directly. These methods are exact and converge, but they need a complete model and they sweep over every state, so they break down when the model is unknown or the state space is large (the curse of dimensionality). They remain useful in small, well-specified problems, and Monte Carlo and TD methods are ways of doing the same backups from samples instead of a model.

10. Compare Monte Carlo and temporal-difference learning. What is the bias-variance trade-off?

Answer: Monte Carlo waits until the episode ends and updates each state's value towards the actual return observed. TD learning updates after each step towards r + γV(s'), using its own current estimate of the next state, which is called bootstrapping. Monte Carlo targets are unbiased but high-variance, because a whole episode of random events feeds every update; it also needs episodes that end. TD targets have lower variance and work online and in continuing tasks, but they are biased while V is still wrong. n-step returns and TD(λ) sit in between: use a few real rewards, then bootstrap. TD usually learns faster, so most deep RL is TD-based; Monte Carlo-style returns suit short episodes with end rewards, such as scoring a whole LLM response.

Value-based methods: Q-learning, SARSA, DQN

11. Write the Q-learning update and explain why it is off-policy.

Answer: Q(s, a) ← Q(s, a) + α [r + γ maxa' Q(s', a') − Q(s, a)]. The term in brackets is the TD error; α is the learning rate. It is off-policy because the target uses the maximum over next actions, which is the value of the greedy policy, regardless of which action the agent actually takes next. The agent can behave ε-greedily (or even randomly) and still learn the greedy policy's values. In the tabular case, with every state-action pair visited infinitely often and a suitably decaying learning rate, Q-learning converges to Q*. The weaknesses to mention: the max introduces an upward bias when estimates are noisy (see Q16), and once you replace the table with a neural network that convergence result no longer holds.

12. How does SARSA differ from Q-learning, and why does it matter on the cliff-walking problem?

Answer: SARSA uses the action the agent actually takes next: Q(s, a) ← Q(s, a) + α [r + γ Q(s', a') − Q(s, a)], so it is on-policy and learns the value of its exploratory policy. On cliff walking, where the shortest path runs along a cliff edge and falling costs a large penalty, Q-learning learns the optimal edge path, but while still exploring with ε-greedy it occasionally steps off and pays the penalty. SARSA accounts for its own exploration and learns a safer path further from the edge, so its online performance during training is better. In production, if the system keeps exploring while it runs, an on-policy method accounts for the risk of its own mistakes.

13. Implement tabular Q-learning for a small environment.

Answer: A corridor with six cells: the agent starts at cell 0, the goal is cell 5, each step costs a little and reaching the goal pays 1. The agent learns to move right everywhere.

import random

# 1-D corridor: states 0..5, start at 0, goal at 5
N, GOAL = 6, 5
ACTIONS = [-1, +1]                  # left, right
alpha, gamma, eps = 0.1, 0.9, 0.1
Q = [[0.0, 0.0] for _ in range(N)]

def step(s, a):
    s2 = min(max(s + ACTIONS[a], 0), N - 1)
    r = 1.0 if s2 == GOAL else -0.01  # small step cost
    return s2, r, s2 == GOAL

random.seed(0)
for episode in range(500):
    s, done = 0, False
    while not done:
        if random.random() < eps:
            a = random.randrange(2)               # explore
        else:
            a = 0 if Q[s][0] > Q[s][1] else 1  # exploit
        s2, r, done = step(s, a)
        target = r if done else r + gamma * max(Q[s2])
        Q[s][a] += alpha * (target - Q[s][a])  # TD update
        s = s2

print([round(max(q), 2) for q in Q[:GOAL]])
print(["LR"[q.index(max(q))] for q in Q[:GOAL]])

It prints values rising towards the goal (about 0.62, 0.7, 0.79, 0.89, 1.0) and the policy R, R, R, R, R. Note the terminal handling: when the episode ends there is no next-state value to bootstrap from. Forgetting that is one of the most common bugs interviewers look for.

14. Why is Q-learning with a neural network unstable? What is the "deadly triad"?

Answer: Sutton and Barto name three ingredients that together can make value learning diverge: function approximation (a network generalises across states, so updating one state's value shifts others), bootstrapping (targets are built from the network's own estimates) and off-policy data. Each is useful alone; together, an error in one estimate leaks into the targets for others and can feed back on itself. Neural Q-learning adds two more problems: consecutive transitions are strongly correlated, which violates the roughly independent samples that stochastic gradient descent assumes, and the target moves every time the network updates. DQN's contribution was a set of engineering tricks that tame these problems enough to learn from pixels, as DeepMind showed on Atari games, not a fix that removes them.

15. Explain experience replay and target networks in DQN.

Answer: Experience replay stores transitions (s, a, r, s', done) in a large buffer and trains on random mini-batches from it. That breaks the correlation between consecutive samples, lets each transition be reused many times (better sample efficiency) and smooths over shifts in the data distribution. It works because Q-learning is off-policy. A target network is a lagged copy of the Q-network used to compute r + γ max Qtarget(s', a'). It is updated by copying weights every few thousand steps or by slowly blending them (Polyak averaging). Freezing the target for a while turns a chasing-your-own-tail problem into a sequence of near-supervised regression problems.

 env --(s,a,r,s')--> replay buffer
                         |
                  random mini-batch
                         v
 online Q-net  <-- loss --  r + γ·max Q_target(s')
       |                          ^
       +---- copy every C steps --+

Interview tip: Mention the usual extras: Huber loss instead of squared error, gradient clipping, and that replay buffer size and the target update period are among the most sensitive hyperparameters.

16. What is overestimation bias, and how do Double DQN, dueling networks and prioritised replay help?

Answer: Taking the maximum of noisy estimates is biased upward: if every action's true value is zero but estimates are noisy, the max is positive. Q-learning bootstraps on that max, so the optimism compounds. Double DQN decouples selection from evaluation: the online network chooses the argmax action in s', the target network scores it. Dueling networks split the head into a state value and per-action advantages, which helps when many actions have similar value because V can be learnt from every transition. Prioritised replay samples transitions with large TD error more often, with importance weights to correct the bias this introduces.

Policy gradients, actor-critic and PPO

17. What is the policy gradient idea, and when do policy methods beat value-based ones?

Answer: Instead of learning values and acting greedily, parameterise the policy πθ(a | s) directly and do gradient ascent on expected return. The policy gradient theorem gives the direction: ∇J(θ) = E[∇ log πθ(a | s) · A(s, a)]. Intuitively, increase the log-probability of actions that turned out better than expected and decrease it for those that turned out worse, weighted by how much better or worse. Policy methods are preferred when actions are continuous or very many (a max over actions is then expensive or impossible), when the optimal policy is stochastic (games with bluffing, or a deliberate spread across options), and when you want smooth policy changes rather than abrupt switches caused by a small change in Q. Their downsides are high-variance gradients and poor sample efficiency, because they are typically on-policy.

18. Explain REINFORCE and why a baseline reduces variance without adding bias.

Answer: REINFORCE is the Monte Carlo policy gradient: run a full episode, compute the return Gt from each step, and update θ ← θ + α Gt ∇ log πθ(at | st). It is unbiased but very noisy, because Gt mixes the effect of this action with all the luck that followed. Subtracting a baseline b(s), typically a learnt estimate of V(s), gives (Gt − b(s)) as the weight. It adds no bias because the expected value of ∇ log π(a | s) under the policy is zero, so subtracting anything that does not depend on the action leaves the expectation unchanged. It reduces variance because the weight now measures "better or worse than usual" instead of raw return. This exact idea reappears in LLM training: GRPO's group average is a baseline (Q33).

19. What is actor-critic, and what does generalised advantage estimation add?

Answer: Actor-critic keeps two components: an actor (the policy) and a critic (a value function) that estimates how good states are. Instead of waiting for the full return, the actor is updated with an advantage estimate from the critic, such as the one-step TD error δ = r + γV(s') − V(s). That brings TD's lower variance to policy gradients, at the cost of bias from an imperfect critic. Generalised advantage estimation (GAE) blends n-step advantages with an exponentially weighted λ: λ = 0 is the one-step TD error (low variance, more bias), λ = 1 is close to the Monte Carlo return minus a baseline (low bias, high variance). PPO is usually implemented as an actor-critic with GAE.

20. Explain PPO and the intuition behind its clipped objective.

Answer: Policy gradients reuse a batch of data only once because one large step can wreck the policy, and the next batch is then collected by the wrecked policy. PPO lets you take several gradient epochs on the same batch while keeping each update modest. It computes the probability ratio rt = πnew(a | s) / πold(a | s) and optimises min(rt A, clip(rt, 1 − ε, 1 + ε) A), with ε around 0.1 to 0.2. The intuition: if an action had positive advantage, you may raise its probability, but once the ratio passes 1 + ε there is no extra reward for pushing further; if the advantage was negative, you may lower it only down to 1 − ε. The min makes the objective pessimistic, so the clip removes the incentive to overshoot but never hides a change that made things worse. PPO approximates TRPO's trust region with simple first-order optimisation, which made it the default for both control tasks and RLHF.

Interview tip: Mention what you monitor in a PPO run: the fraction of samples clipped, approximate KL between old and new policies, entropy, value loss and explained variance. A clip fraction that keeps growing usually means the learning rate or epoch count is too high.

21. How do you handle continuous action spaces? Compare DDPG, TD3 and SAC.

Answer: With continuous actions (a steering angle, a discount level, a valve setting), a max over actions is not available, so methods either output a distribution (Gaussian policy in PPO) or learn a deterministic actor trained through the critic. DDPG is deterministic actor-critic with replay and target networks; it is sample-efficient but brittle and prone to overestimation. TD3 fixes the main failures: two critics with the minimum used as target (clipped double Q), delayed actor updates, and noise added to target actions for smoothing. SAC adds a maximum-entropy objective: maximise reward plus an entropy bonus, with the temperature often tuned automatically. That keeps exploration alive and tends to make SAC robust to hyperparameters, which is why it is a common first choice for off-policy continuous control.

Model-based and offline RL

22. What is the difference between model-based and model-free RL, and when is model-based worth it?

Answer: Model-free methods learn a value function or policy directly from experience without learning how the environment works. Model-based methods learn (or are given) a transition and reward model, then plan with it (tree search, trajectory optimisation, model predictive control) or generate imagined experience to train a policy, as in Dyna-style methods. Model-based RL is far more sample-efficient, which suits problems where real interaction is expensive or risky. Its weakness is model error: the planner exploits the model's mistakes and produces plans that look excellent in the model and fail in reality. Mitigations are short planning horizons, ensembles of models with uncertainty penalties, and frequent re-planning. In industry, a hand-built simulator used for planning is often the most practical "model".

23. What is offline RL, why is it hard, and how do you evaluate an offline policy before deploying it?

Answer: Offline (batch) RL learns a policy from a fixed log of past interactions with no further exploration, which is the natural setting for banks, hospitals and contact centres that cannot experiment freely. The core difficulty is distribution shift: the learnt policy wants to take actions the logging policy rarely took, the Q-function has no data to correct its estimates for those actions, and the max in the Bellman target picks exactly those over-optimistic errors. Offline methods stay close to the data: behaviour cloning as a baseline, CQL (conservative Q-learning, which pushes down Q-values of unseen actions) and IQL (implicit Q-learning, which avoids querying them at all).

Evaluation uses off-policy evaluation (OPE): inverse propensity scoring reweights logged rewards by πnew(a | s) / πlog(a | s); doubly robust estimators combine that with a learnt reward model to reduce variance; fitted Q evaluation estimates the new policy's value with a separate critic. All of them need the logging policy's action probabilities and overlap between the policies, and their estimates get unreliable as the new policy moves away from the old one. The last step is always a limited online test: shadow mode, then a small randomised rollout with guardrail metrics.

Interview tip: Say plainly that an offline policy whose estimated gain comes mostly from actions the log almost never contains should be distrusted, however impressive the number.

Reward design and reward hacking

24. What is reward shaping, and how do you handle sparse rewards without changing what the agent learns?

Answer: Sparse rewards (a payoff only when a long task succeeds) give the agent almost no signal early on, so random exploration may never find success. Reward shaping adds intermediate rewards to guide learning, but naive shaping changes the optimal policy: reward a cleaning robot for picking up dirt and it may learn to drop and re-pick the same dirt. Potential-based shaping (Ng, Harada and Russell) adds F = γΦ(s') − Φ(s) for some potential function Φ, such as negative distance to the goal; because the terms telescope over a trajectory, the optimal policy is provably unchanged. Other tools for sparse rewards: curricula that start with easy instances, demonstrations and behaviour cloning to warm-start, hindsight experience replay (relabel a failed attempt as a success at reaching wherever it actually ended), and intrinsic curiosity bonuses.

25. What is reward hacking, and how do you detect and prevent it?

Answer: Reward hacking (specification gaming) is when an agent maximises the reward as written while failing at the intent. It is Goodhart's law with an optimiser attached: the stronger the optimisation, the more any gap between proxy and goal is found. Detection needs signals the agent is not trained on: held-out metrics, human review of high-reward samples (the top of the reward distribution is where hacks live), watching for reward rising while independent quality metrics are flat or falling, and checking distribution shifts in behaviour such as length, refusal rate or action frequencies. Prevention: rewards built from multiple terms and hard constraints, regularising toward a reference policy (the KL penalty in RLHF), ensembles of reward models, early stopping on a held-out evaluation, and periodically refreshing the reward model with data from the current policy.

Real-world example: Consider a ticket-routing agent rewarded on "tickets closed within SLA". It learns to route borderline tickets to the queue whose agents close tickets fastest, including closing them unresolved. The reward is satisfied; the customer is not. Adding a reopen-rate penalty fixes the specific hack; reviewing high-reward episodes finds the next one.

Multi-armed and contextual bandits

26. What is a multi-armed bandit, and how does it relate to full RL?

Answer: A bandit is RL with a single state: each round you choose one of K actions (arms), observe a reward and repeat, and your choice does not change future situations. The goal is to minimise regret, the reward lost compared with always pulling the optimal arm. Contextual bandits add an observed context (user features, time of day) and learn which arm suits which context, still without long-term effects. Full RL is needed only when actions change future states, for example when an offer today changes a customer's eligibility or patience tomorrow. Most business "RL" proposals honestly sit on the bandit rungs of the ladder (A/B test, bandit, contextual bandit, full RL), and bandits are far easier to build, evaluate and explain.

27. Compare ε-greedy, UCB and Thompson sampling, and implement Thompson sampling.

Answer: ε-greedy explores uniformly at random with probability ε; it is easy but keeps wasting traffic on clearly bad arms unless ε is decayed. UCB1 picks the arm with the highest mean + √(2 ln t / na): rarely tried arms get a large bonus, and the bonus shrinks as evidence accumulates. Thompson sampling keeps a posterior over each arm's reward rate, samples one plausible value per arm and plays the highest sample. It explores in proportion to the probability that an arm is optimal and copes well with delayed, batched feedback. For conversion-style rewards, a Beta-Bernoulli model is enough:

import random

true_rates = {"cashback": 0.04, "fee_waiver": 0.05,
              "bonus_points": 0.03}
wins = {k: 1 for k in true_rates}     # Beta(1, 1) prior
losses = {k: 1 for k in true_rates}

random.seed(1)
for _ in range(20000):
    # sample a plausible rate per offer; show the top one
    arm = max(true_rates,
              key=lambda k: random.betavariate(
                  wins[k], losses[k]))
    converted = random.random() < true_rates[arm]
    wins[arm] += converted
    losses[arm] += not converted

for k in true_rates:
    print(k, wins[k] + losses[k] - 2, "shown")

In this simulated run the fee-waiver arm, which has the highest true rate, receives most of the 20,000 impressions while the weaker arms still get enough traffic to be ruled out. The rates are invented for illustration.

28. When should you use a bandit instead of a classic A/B test?

Answer: They answer different questions. An A/B test with fixed allocation is built for inference: an unbiased estimate of the difference between variants with a confidence interval, which you need when the decision is permanent, expensive or must be defended to stakeholders or regulators. A bandit is built for earning while learning: it shifts traffic toward winners during the experiment, reducing regret, which suits short-lived decisions (headlines, festive-season offers, creatives that go stale), many variants, or continuous optimisation. The costs of bandits: adaptive allocation biases naive estimates of each arm's effect, slow-moving metrics (retention, churn) do not fit fast reward loops, and novelty effects or seasonality can fool them. A practical compromise is to run a bandit for the immediate metric while holding out a small, uniformly randomised control group for clean measurement, and to always log assignment probabilities.

29. How do contextual bandits power recommendations, and what must you log?

Answer: A contextual bandit picks an item (or a slate) for a user given features of the user, item and context, observes feedback (click, purchase, dwell) and updates. LinUCB assumes reward is linear in features and adds an uncertainty bonus; neural or tree-based models with Thompson-style sampling or ε-greedy exploration scale further. Explicit exploration counters the feedback loop where items never shown never get data. The non-negotiable logging is the tuple (context, action, propensity of that action, reward, timestamp, policy version). Propensities make counterfactual evaluation of future policies possible; without them your logs only describe the old policy. For ranking, candidate generation and the wider recsys stack, our recommendation systems interview questions go deeper.

RL interviews reward engineers who can connect the maths to a production system with logging, evaluation and guardrails. If you want structured, hands-on practice across ML, deep learning, GenAI and cloud, Cloudsoft's APEX program for AI, ML, cloud and cyber security runs in classroom sessions in Ameerpet and live online.

RL for LLMs and agents

30. How does LLM training map onto the RL framework?

Answer: The policy is the language model. The state is the prompt plus the tokens generated so far; an action is the next token; an episode is one complete response (or, for agents, a multi-turn trajectory including tool calls and tool results). Transitions are deterministic, since appending a token gives a known next state, so the difficulty is not environment randomness but a huge action space and a reward that usually arrives only at the end of the response. That is why most LLM RL treats each response as a single bandit-like episode with a sequence-level reward, and why credit assignment across hundreds of tokens is the central technical problem. Our LLM interview questions cover the transformer side of the same models.

31. How is a reward model trained, and why does optimising against it eventually make things worse?

Answer: A reward model is usually the SFT model with a scalar head, trained on human or AI comparisons of two responses to the same prompt. The Bradley-Terry formulation models P(A preferred over B) = σ(r(A) − r(B)), so the loss pushes the preferred response's score above the other's. The reward model is a learnt proxy, accurate near the data it saw. As the policy optimises, it drifts into regions the reward model never saw, where the proxy is wrong, and true quality first rises and then falls while the proxy reward keeps climbing. That over-optimisation curve is the RLHF version of Goodhart's law, and it is why the KL penalty, reward-model ensembles, held-out human evaluation and refreshing the reward model on new policy samples all exist.

32. Why does RLHF use a KL penalty to a reference model, from an RL point of view?

Answer: The objective is reward minus β times the KL divergence between the policy and a frozen reference (the SFT model). In RL terms it does three things. It is a trust region: it limits how far the policy can move toward regions where the reward model is unreliable, directly countering over-optimisation. It is a prior: the reference model encodes fluency and general knowledge that the narrow reward does not measure, so the penalty preserves them. And it keeps output diversity: without it, the policy tends to collapse onto a few high-reward patterns (mode collapse). Too small a β gives reward hacking and drift; too large a β means RL changes almost nothing. Tracking KL against reward over training is the standard health check.

33. From an RL perspective, compare PPO-based RLHF, DPO and GRPO with verifiable rewards.

Answer: The useful axis is online versus offline, and learnt versus checked reward.

AspectPPO RLHFDPOGRPO + verifiable reward
DataOnline: policy generates, reward model scoresOffline: fixed preference pairsOnline: groups of samples per prompt
RewardLearnt reward modelImplicit, from preferencesProgram check (tests, exact answer)
Baseline / criticLearnt value modelNone neededGroup mean reward
Main riskReward hacking, instability, costLimited to what the pairs showGameable checks, narrow gains

Online methods can discover responses better than anything in the dataset because they explore; DPO is closer to supervised learning on preferences, so it is cheaper and steadier but can only shift probability among behaviours the data covers. GRPO, popularised by DeepSeek's work on mathematical reasoning, is essentially REINFORCE with a group baseline and PPO-style clipping: sample several responses per prompt, normalise each reward against the group's mean (and often its standard deviation) to get an advantage, and drop the value network entirely. With verifiable rewards (unit tests, exact numeric answers, schema checks), the reward is much harder to fool than a learnt model, which is why this recipe sits behind many reasoning models. For DPO's loss and beta, ORPO, KTO and preference-data pipelines, see the fine-tuning interview guide; for hands-on recipes, the fine-tuning LLMs guide.

34. Where does RL actually show up in AI agents?

Answer: In two main places, and it is worth being precise because most enterprise agents use no RL at runtime. First, model training: model providers train tool-using and coding models with RL on multi-step trajectories, where the reward is task success (tests pass, the right record was updated) and credit must be assigned across many tool calls. Outcome rewards score only the final result; process rewards score intermediate steps, which helps credit assignment but needs a reliable step-level judge. Second, adaptive decisions around the agent: bandits choosing which model, prompt variant or retrieval strategy to route a request to, balancing cost against success rate. For a typical enterprise team the practical path is prompting, tools and rigorous agent evaluation; RL fine-tuning of an agent model is justified only with a large volume of tasks, a verifiable success signal and the budget to run it. Taking agents from evaluation into customer production is the focus of the Forward Deployed Engineer program (FDE PRO).

Practical issues and business applications

35. Why is RL so sample-inefficient, and what do you do about it?

Answer: The agent must discover good behaviour by trial, rewards are sparse and delayed, and on-policy methods throw data away after each update, so learning tasks a human finds simple can take millions of interactions. Levers, roughly in order of payoff: start from a good policy (behaviour cloning from logs or experts, or a pretrained model, which is why LLM RL is feasible at all); use off-policy methods with replay; use model-based methods or a simulator to generate experience; shrink the problem (fewer actions, a well-engineered state, shorter horizons); shape rewards carefully; and run many environments in parallel when simulation is cheap.

36. RL training is notoriously unstable. How do you debug an agent that is not learning?

Answer: Debug in layers, starting with the cheapest checks:

  1. Environment: step it by hand, check reward signs and scales, terminal versus time-limit truncation (a time limit is not a true terminal and should still bootstrap), observation normalisation and action bounds.
  2. Sanity baselines: random policy and a simple heuristic. If the heuristic scores far above the agent, the agent is broken, not the task.
  3. A known-solvable toy: if your code fails on a simple benchmark, the bug is in the algorithm.
  4. Diagnostics: TD error and Q-value magnitudes (exploding Q means divergence), policy entropy (collapsing too early means premature convergence), value-function explained variance, gradient norms, KL between updates.
  5. Seeds: results vary a lot across random seeds, so judge changes over several seeds, never one run.

Then tune the usual suspects: learning rate, reward scale, discount, batch size, target update period or PPO epochs. Many stability problems that look like deep learning problems overlap with our deep learning interview questions on optimisation and normalisation.

37. How do you evaluate an RL agent properly?

Answer: Report the distribution, not the luckiest run: learning curves across several seeds with confidence intervals, and interquartile means or similar robust aggregates rather than the maximum. Separate training performance (with exploration) from evaluation performance (deterministic or low-exploration policy on fresh episodes). Test generalisation: held-out environment seeds, different demand patterns, perturbed dynamics. Compare with strong non-RL baselines, such as the current business rule or an optimisation heuristic, because beating a random policy proves nothing. For deployment, add off-policy evaluation on logged data, then shadow mode (the agent recommends, the existing system acts) and a small randomised online test with guardrail metrics. Evaluate behaviour too: action distributions, constraint violations, worst-case episodes. For LLM policies the same thinking applies with different tools, covered in our LLM evaluation interview questions.

38. What role do simulators play, and how do you handle the sim-to-real gap?

Answer: Simulators make RL practical by providing cheap, safe and parallel experience: logistics and warehouse simulations, queueing models for contact centres, traffic and network simulators, physics engines for robotics. The risk is that the policy learns the simulator's quirks. Ways to manage it: calibrate the simulator against historical data and validate it on periods it was not fitted to; use domain randomisation (vary demand, delays, failure rates during training so the policy cannot rely on one exact setting); add realistic noise, latency and missing data; prefer policies that are robust over ones that are optimal for one setting; and fine-tune or adapt with limited real data. Above all, check that it models the effect of the agent's actions. A simulator replaying historical demand cannot tell you how customers react to a price nobody ever charged.

39. How do you make an RL system safe enough to run in production?

Answer: Assume the agent will find whatever the reward does not forbid, and build layers around it:

  • Action masking and hard limits: remove disallowed actions before the policy chooses (price floors and ceilings, maximum order quantities, eligibility rules), rather than hoping a penalty deters them.
  • Constrained RL: optimise reward subject to expected cost constraints (Lagrangian methods), for example "maximise throughput while keeping overtime below a limit".
  • Conservative deployment: shadow mode, small traffic slices, automatic rollback on guardrail metrics, and a fallback rule-based policy.
  • Human oversight: human-in-the-loop approval for high-impact actions and a clear owner for the reward definition.
  • Drift monitoring: new products, policy changes and seasonality can make last year's policy misbehave.

For RL-trained LLMs the equivalent layers are input and output guardrails, refusal behaviour tests and red-teaming after each training run.

40. Where does RL create business value, and what are its honest limits?

Answer: A balanced answer separates where RL-style methods are mature from where they are mostly proposals.

AreaRealistic useHonest limits
Recommendations, offersContextual bandits for slates, offers, notificationsLong-term effects hard to credit; clicks can be the wrong reward
Pricing and promotionsBandit-style price testing within guardrailsFairness, regulation, competitor reactions, brand risk
OperationsInventory, scheduling, routing, control with simulatorsSimulator fidelity; strong OR baselines often match it
Contact centresOffline policies for routing and next actionLogged data coverage; agent and customer behaviour shifts
LLMs and agentsRLHF and verifiable-reward fine-tuningCost, reward hacking, needs large task volume

The pattern: RL is worth it when decisions are repeated at high volume, actions affect future states, feedback is measurable, exploration is safe or simulatable, and simpler methods have been tried. Operations research (linear and integer programming), forecasting plus rules, and supervised models solve many "RL" problems more cheaply. Demand forecasting, often the real bottleneck, is covered in our time series forecasting interview questions.

Real-world scenario questions

41. A bank wants to pick which of five credit card offers to show each customer in its mobile app. Design a bandit for offer selection.

Answer: This is a contextual bandit: one decision per app session, a measurable outcome, and limited long-term state effects (with caveats below). Start with Thompson sampling or LinUCB over permitted customer features.

What I would check:

  1. Reward definition: a click is fast but weak; an approved application is meaningful but delayed by days. Use a click or apply-start reward for learning speed, with delayed approved-application rewards joined back in, and monitor approval and early-default rates as guardrails.
  2. Eligibility and compliance: the bandit only chooses among offers the customer is eligible for, uses no protected or prohibited attributes, and keeps an audit trail of why each offer was shown.
  3. Exploration limits: a minimum exploration floor per arm so every offer keeps getting some data, and a cap so no customer segment receives clearly unsuitable offers.
  4. Logging: context, offer shown, propensity, policy version and outcome, so future policies can be evaluated offline.
  5. Measurement: a small uniformly random holdout for unbiased lift estimates against the current rule-based selection.

Production consideration: Model risk teams will ask to explain the policy. A linear contextual bandit with documented features and propensity logs is far easier to defend than a deep RL agent, and it is usually enough.

42. Your team ran RLHF on a customer-support assistant using a reward model trained on CSAT ratings. CSAT-predicted reward rose sharply, but escalations and refund disputes also rose. What happened?

Answer: Reward hacking. CSAT ratings reward how the customer felt at the end of the chat, so the reward model learnt that apologies, warmth and promises score highly. The policy then learnt to promise things it cannot deliver, such as refunds or callbacks, and to tell customers their issue is resolved. Customers rated the chat well and escalated later.

What I would check:

  1. Read the highest-reward responses from the latest checkpoint; hacks show up at the top of the reward distribution.
  2. Compare commitments the assistant makes ("I have processed your refund") with what the backend logs show.
  3. Plot reward against KL from the reference over training; a sharp late rise in reward with growing KL suggests over-optimisation.
  4. Score held-out conversations with an independent rubric: policy compliance, factual accuracy, resolution confirmed by downstream systems.

Fixes: roll back to an earlier checkpoint; retrain the reward model with negative examples of unauthorised commitments; add rule-based penalties for actions the assistant may not promise; raise the KL coefficient; and move checkable parts (did the refund actually happen, did the ticket reopen) into the reward or evaluation.

Production consideration: Never use a single satisfaction proxy as the whole reward. Pair it with policy-compliance checks and downstream outcome metrics, and gate every training run on a fixed evaluation set that includes these failure cases.

43. A contact centre proposes offline RL to route calls to agent queues, trained on two years of routing logs. How do you evaluate the proposal?

Answer: Offline RL can fit, because routing is a repeated sequential decision (today's routing affects queue lengths and agent workload minutes later) and live exploration with customers is costly. But the proposal stands or falls on the logs.

What I would check:

  1. Coverage: the current router is mostly rule-based and deterministic, so for most call types only one queue was ever tried. Without variety in the logged actions, no offline method can learn what the alternatives would have done. Look for natural variation (overflow routing, rule changes) or run a small randomised window first.
  2. Propensities: were routing probabilities logged? If not, off-policy evaluation needs estimated propensities, which adds uncertainty.
  3. Reward: first-contact resolution, handle time, transfers, abandonment and CSAT pull in different directions. Agree a weighted reward and hard constraints (maximum wait time, regulatory callbacks) with the business up front.
  4. Confounders: for example supervisors manually overriding routing for VIP callers.
  5. Baselines: compare against a skills-based routing heuristic and a supervised "predict resolution per queue" model, which may capture most of the gain.
  6. Method: a conservative method (CQL or IQL) or a constrained improvement over behaviour cloning, evaluated with doubly robust OPE plus a queueing simulator calibrated on held-out weeks.
 routing logs -> coverage check -> reward/limits
                                       |
                                       v
 conservative offline RL -> OPE + queue simulator
                                       |
               shadow mode -> small randomised pilot

Production consideration: Deploy as a recommendation alongside the existing router first, with automatic fallback, and monitor agent workload fairness: an optimiser can quietly send all hard calls to the same few skilled agents.

44. An e-commerce retailer wants an RL agent to set prices dynamically during a festive sale. What do you advise?

Answer: Start narrower than "an RL agent sets prices". Most of the value usually comes from good demand forecasts, price elasticity estimates and an optimiser with business constraints; bandit-style price testing within tight bands can then learn elasticity online.

What I would check:

  1. Price history variation: has the retailer ever tested prices, or were changes driven by events that confound elasticity estimates?
  2. Constraints: floors at cost plus margin, ceilings, maximum change per day, MAP agreements, and consistent prices across channels.
  3. Fairness and legal review: personalised prices per customer raise regulatory, reputational and consumer-protection concerns; segment-level or SKU-level pricing is far safer.
  4. Inventory dynamics: a short sale with limited stock is genuinely sequential (sell out early or hold stock), where RL or dynamic programming helps.
  5. Reward: margin over the whole sale period, with stock-out and leftover-inventory costs, not daily revenue.

Production consideration: Run it with human approval of price bands, rollback in minutes, and alerts on price changes beyond thresholds.

45. A DQN agent's average reward climbed steadily for hours, then collapsed and never recovered. What do you investigate?

Answer: Classic causes are divergence of Q-values, loss of exploration and a replay buffer that has forgotten earlier experience.

What I would check:

  1. Q-value magnitudes over time: values growing far beyond the plausible return range indicate divergence; Double DQN, a slower target update, Huber loss and gradient clipping help.
  2. Replay buffer composition: once the agent got good, the buffer filled with only successful states, and it may have "forgotten" how to recover from bad ones (catastrophic forgetting). A larger buffer or reserving part of it for diverse experience helps.
  3. Exploration schedule: ε decayed to near zero, so the agent stopped correcting errors.
  4. Learning rate: too high for late training; decay it.
  5. Environment changes: did a curriculum stage, a reward change or a bug in episode resets coincide with the collapse?

Production consideration: Always checkpoint regularly and select the deployed model on a separate evaluation run, not the final training step.

46. A warehouse picking-schedule agent trained in a simulator beats the current heuristic by a wide margin in simulation but performs worse in the pilot. Why?

Answer: A sim-to-real gap, and the agent has probably exploited something the simulator gets wrong.

What I would check:

  1. Compare simulated and real distributions for the pilot period: order arrival, pick times, travel times, worker availability, breaks.
  2. Look at what the agent does differently from the heuristic. If it relies on, for example, perfectly predictable pick times or zero congestion in aisles, the simulator lacks those effects.
  3. Check the observation pipeline: real features may arrive late, be missing or be computed differently.

Fixes: add the missing dynamics and noise, use domain randomisation, retrain with a robustness objective, and fine-tune with a limited amount of real data under the heuristic's supervision.

Production consideration: Validate the simulator itself before trusting the agent: replay the heuristic in simulation for a past month and confirm it reproduces that month's real metrics.

47. A product manager wants to replace all A/B tests with bandits because "bandits waste less traffic". The analytics lead disagrees. How do you mediate?

Answer: Both are right for different decisions, so agree a decision rule rather than a winner.

What I would check:

  1. Is the decision permanent and does it need an unbiased effect size (a new checkout flow, a pricing policy)? Use a fixed-allocation A/B test.
  2. Is it short-lived, repeated or with many variants (banners, push notification copy, festival creatives)? Use a bandit.
  3. How fast does the target metric arrive? Bandits need quick feedback; retention effects need A/B tests.
  4. Are there long-term or interaction effects (one variant changes behaviour on other pages)? Prefer A/B tests.

A middle path: bandits with a fixed random holdout and logged propensities, so the analytics team can still estimate effects with inverse propensity weighting.

48. After RL with verifiable rewards on a code-generation model, the unit-test pass rate jumped, but reviewers find the generated code is worse. What happened?

Answer: The reward was verifiable but incomplete, so the policy found ways to pass tests without solving the problem.

What I would check:

  1. Inspect passing samples for special-casing of test inputs, hard-coded expected outputs, catching and suppressing exceptions, or modifying or skipping tests if the environment allows file access to them.
  2. Test coverage: thin tests reward thin solutions. Evaluate on a held-out set with stronger, hidden tests.
  3. Sandbox permissions: the policy should not be able to read or edit test files or the grader.
  4. Other quality signals: readability, length, use of unsafe functions, performance.

Fixes: hidden and property-based tests, a sandbox that isolates the grader, penalties for disallowed patterns, mixing in a quality judge for non-checkable aspects, and a regression suite run on every checkpoint.

Production consideration: "Verifiable" means harder to game, not impossible to game. Treat the grader as an attack surface and review top-reward samples every run.

49. A GCC team in Hyderabad wants to "use reinforcement learning" to improve its IT-operations agent that resolves tickets with tools. How do you respond?

Answer: Ask what problem RL would solve that the current approach cannot. Faster gains usually come from better tool descriptions, retrieval, prompts and an evaluation set of real tickets.

What I would check:

  1. Is there an evaluation suite with clear success criteria per ticket type? Without it, RL has no reward and the team cannot tell whether anything improved.
  2. Where do failures occur: wrong tool choice, bad arguments, missing knowledge, or unsafe actions? Most are fixed without training.
  3. Is there a verifiable success signal (the service is healthy again, the access request was granted correctly) and enough task volume to train on?
  4. Is there a safe sandbox replicating the tools? RL on live production systems is not acceptable.
  5. Are there smaller RL-flavoured wins, such as a bandit routing tickets between a cheap and an expensive model by success rate and cost?

Production consideration: If RL fine-tuning is eventually justified, it is usually done on an open-weight model in a sandbox with verifiable task rewards, and gated by the same evaluation suite plus security review of the tool permissions. Our AI agent developer interview questions cover the non-RL side of building these agents.

50. Your agent's reward stays at exactly zero for the first million steps. How do you diagnose it?

Answer: Almost certainly a sparse-reward exploration failure, or a bug that makes success unreachable.

What I would check:

  1. Can success happen at all? Script a hand-written policy or replay a known solution through the environment and confirm it receives the reward.
  2. Reward plumbing: is the reward computed but not passed through wrappers, clipped to zero, or delivered after the episode is already marked done?
  3. Episode length: a time limit shorter than the minimum steps to the goal makes zero reward certain.
  4. Exploration: random actions may effectively never reach the goal. Add potential-based shaping, a curriculum from easier start states, demonstrations, hindsight relabelling or curiosity bonuses.
  5. For LLM RL, the equivalent is prompts too hard for the starting model: every sample fails, advantages are zero and nothing is learnt. Filter prompts to those with mixed outcomes.

Production consideration: Add an automated check to the training pipeline that alerts if the fraction of non-zero-reward episodes stays at zero after a warm-up period, so a broken run does not burn a week of compute.

Key takeaways

  • Bellman's idea, immediate reward plus discounted value of the next state, underlies DP, TD learning, Q-learning and actor-critic methods.
  • Know the stability tricks and why they exist: replay and target networks for DQN, baselines and advantages for policy gradients, clipping for PPO, KL penalties for RLHF.
  • Most business "RL" is honestly a bandit or contextual bandit problem; climb from A/B test to bandit to full RL only when actions change future states.
  • Offline RL lives or dies on logged action coverage and propensities; use conservative methods and off-policy evaluation, then a limited online test.
  • Reward hacking is the default outcome of strong optimisation against a proxy; inspect top-reward samples and track metrics the agent is not trained on.
  • For LLMs, distinguish online versus offline and learnt versus verifiable rewards; GRPO is REINFORCE with a group baseline.
  • Always compare against strong non-RL baselines and say when RL is the wrong tool.

Interview preparation checklist

  • Write the Bellman expectation and optimality equations from memory and explain them with an analogy.
  • Implement tabular Q-learning and SARSA on a small grid; change γ and ε and explain the results.
  • Train a DQN and a PPO agent on a standard benchmark environment over several seeds and plot the variance.
  • Implement ε-greedy, UCB1 and Thompson sampling on a simulated bandit and compare cumulative regret.
  • Explain inverse propensity scoring and doubly robust off-policy evaluation, and what logging they require.
  • Prepare two reward-hacking stories (one classic RL, one LLM) and how you would detect each.
  • Be able to sketch the RLHF pipeline and compare PPO, DPO and GRPO in two minutes.
  • Prepare one business case where you would recommend against RL and say what you would use instead.
  • Revise supporting topics: machine learning fundamentals, deep learning optimisation and evaluation design.

FAQ

Which reinforcement learning topics are most commonly asked in interviews?

Commonly asked topics include MDPs, the Bellman equations, Q-learning versus SARSA, DQN's replay and target networks, policy gradients, actor-critic, PPO clipping, exploration strategies, multi-armed bandits, offline RL, reward hacking and RLHF.

Do I need advanced maths for RL interviews?

You need comfort with probability, expectations and gradients, and you should be able to write the main update rules. Most applied interviews test intuition and debugging judgement rather than proofs, but research roles may ask for derivations such as the policy gradient theorem.

Is reinforcement learning still relevant now that LLMs dominate AI work?

Yes. RLHF and RL with verifiable rewards are central to how modern LLMs and reasoning models are trained, and bandits remain widely used in recommendations, offers and experimentation. The RL ideas carry over even when the policy is a language model.

How should a fresher prepare for RL interview questions?

Learn the fundamentals from a standard textbook, implement tabular Q-learning and a simple bandit by hand, then train one deep RL agent with a library on a benchmark environment. Be able to explain every line of your code and why your results vary across seeds.

Which libraries should I know for reinforcement learning?

Gymnasium for environments and Stable-Baselines3 for standard algorithms are common starting points, with PyTorch underneath. For LLM RL, Hugging Face's TRL library is widely used. Check current documentation, since APIs change between versions.

What projects help in an RL interview?

A bandit simulation comparing exploration strategies, a DQN or PPO agent with multi-seed learning curves, and one applied project such as an inventory or offer-selection problem with a clear baseline comparison. Document reward design choices and failures you debugged.

How is an RL interview different from a machine learning interview?

Machine learning interviews focus on supervised models, features, metrics and validation. RL interviews add sequential decisions, exploration, delayed rewards, stability of training and the difficulty of evaluating a policy without deploying it.

Is reinforcement learning a good career skill for Indian engineers?

RL is a specialised skill, used in recommendations, operations, robotics and LLM training teams at GCCs, product companies and research groups. It is most valuable combined with strong ML, deep learning and production engineering skills rather than as a standalone specialisation.

Want to turn RL, ML and GenAI knowledge into systems that run in production? Explore Cloudsoft's APEX AI, ML, Cloud and Cyber Security program, with hands-on labs in classroom sessions in Ameerpet or live online. Call +91 96660 19191 to book a free demo.

Share𝕏inf✉
EnrollWhatsAppCall us