New batches starting this week Β· Limited seats

A/B Testing Interview Questions and Answers 2026 (50 Questions)

50 A/B testing and experimentation interview questions with model answers, from hypotheses and statsmodels power calculations to SRM, CUPED, switchbacks, Bayesian methods, LLM feature tests and ten Indian product scenarios.

A/B testing interview questions 2026: 50 questions on power and MDE, p-values, peeking, CUPED, sample ratio mismatch and testing AI features
Last updated Β· 43 min read Β· 9,419 words

A/B testing interview questions in 2026 check whether you can run an experiment that leads to a correct decision, not whether you can recite the definition of a p-value. Interviewers want you to pick the right randomisation unit and metrics, size the test honestly, spot a broken experiment before anyone trusts it, and explain the result to a product manager who wants to ship tomorrow. This guide has 50 high-value experimentation interview questions with model answers. It runs from hypotheses and power calculations to CUPED, sample ratio mismatch, interference, Bayesian methods, bandits, testing LLM features online, and ten scenarios set in Indian products.

How to use this guide

Experimentation comes up in data science, product analytics, ML engineering and product management loops. What interviewers commonly probe at each level:

  • Freshers and analysts: hypotheses, control and treatment, p-values and confidence intervals stated correctly, Type I and Type II errors, and simple sample-size reasoning.
  • Mid-level data scientists and product analysts: metric design, power and MDE, peeking, multiple testing, ratio metrics, SRM checks and variance reduction.
  • Senior and staff roles: interference and marketplace designs, heterogeneous effects, long-term effects, platform design, testing AI features with noisy quality metrics, and making the call when the data is ambiguous.

Questions are numbered continuously. For each scenario, practise the same order out loud: is the experiment valid, what does the data say, what would I decide, and what would I check next. If you need broader statistics, SQL and modelling revision first, the data science interview questions guide covers the groundwork. This page goes much deeper on experimentation.

Fundamentals

1. What is an A/B test, and why is it better than comparing before and after a launch?

Answer: An A/B test (an online controlled experiment) randomly assigns units such as users to a control experience and one or more treatments, runs them at the same time, and compares outcomes. Randomisation makes the groups alike on average in everything, both what you measure and what you do not, so a difference larger than chance can be attributed to the change. A before-and-after comparison mixes the change with everything else that happened at the same time: a sale, a marketing push, an app store feature, a festival, a competitor outage. Running the arms in parallel cancels those time effects because both arms experience them.

Interview tip: Use the word "causal" and explain why. Interviewers listen for the link between randomisation and comparable groups.

2. How do you write a good experiment hypothesis?

Answer: A good hypothesis names the change, the mechanism, the metric and the direction, and it is written before launch. For example: "Showing delivery-date estimates on the product listing page will reduce uncertainty, which will increase checkout conversion per visitor, without increasing returns." It should also state the decision rule: what you will do if the primary metric moves up, down or not at all. Writing the decision down first stops the team from turning any result into a story after the fact. The mechanism tells you which diagnostic metrics to watch (did users actually see the date?) and keeps a surprising result from being over-read.

3. How do you choose the randomisation unit?

Answer: Randomise at the level where the treatment is experienced and where units are reasonably independent. Common choices are user (logged-in ID), device or cookie (for logged-out traffic), session, page view, account or organisation (B2B), and cluster (city, store, seller). User-level randomisation gives a consistent experience and supports retention metrics. Page-view randomisation gives more units and more power, but a user who sees both versions can get confused, and their outcomes are correlated. The analysis unit should match the randomisation unit. If you randomise by user but analyse per session, sessions from the same user are not independent, and naive standard errors will be too small. When users interact, as in marketplaces or social products, move up to clusters (see Q25).

Real-world example: A B2B SaaS tool used by finance teams at Indian GCCs should randomise by company account. Colleagues share workflows and would notice two different invoice screens.

4. What are primary, guardrail, diagnostic and north star metrics?

Answer:

TypePurposeExample for a checkout change
North starThe long-term company metric that experiments should serve over time; often too slow or noisy to move in one testWeekly active buyers, or long-run revenue per customer
Primary (decision) metricThe one metric, chosen up front, that decides the launch; sensitive and aligned with the north starOrders per visitor
GuardrailMust not get worse beyond a set tolerance; protects users and the businessPayment failure rate, page latency, refunds, support contacts, crash rate
DiagnosticExplains the mechanism and confirms the treatment was deliveredClicks on the new button, exposure counts

Some teams combine several metrics into an overall evaluation criterion (OEC), a single weighted score agreed in advance. The aim is the same: agree what "good" means before you see the data.

5. What does a p-value actually mean?

Answer: A p-value is the probability of seeing a difference at least as extreme as the one observed, assuming the null hypothesis is true and the test's assumptions hold. A p-value of 0.03 means that if the change truly did nothing, data this extreme would turn up about 3 times in 100 repeats. It is not the probability that the null is true, not the probability that the treatment works, and not a measure of how big or important the effect is. A tiny, useless effect can give a very small p-value with enough traffic. A large, valuable effect can give p = 0.2 in an underpowered test.

Interview tip: Interviewers often ask you to explain this to a product manager. Try: "If this change did nothing, a gap this big would be unusual, so we have evidence it did something. How big, and whether it is worth it, comes from the confidence interval."

6. How do you interpret a 95 per cent confidence interval for a treatment effect?

Answer: The interval comes from a procedure that, over many repeated experiments, would contain the true effect 95 times out of 100. For a single interval, the practical reading is "the effects compatible with our data at this confidence level". An interval of +0.1 to +0.9 percentage points on conversion says the effect is likely positive but could be small. An interval of βˆ’0.2 to +1.5 says you cannot rule out zero or a small harm. Strictly, it is wrong to say "there is a 95 per cent probability the true value lies in this particular interval". That is a Bayesian credible-interval statement. Intervals are more useful than p-values for decisions because they show size and uncertainty together. Compare them with the smallest effect worth shipping, not just with zero.

7. What are Type I and Type II errors, and how do they relate to alpha and power?

Answer: A Type I error (false positive) is declaring an effect when there is none. Its rate is controlled by alpha, commonly 0.05. A Type II error (false negative) is missing a real effect of a given size. Its rate is beta, and power is 1 βˆ’ beta, commonly targeted at 0.8. For a fixed sample, lowering alpha raises beta. The only ways to get both down are more data, lower variance, or a larger true effect. Product teams often care more about one side. A risky payments change may want a strict alpha and a strong guardrail. A cheap copy test may accept more false positives in exchange for speed.

8. What is an A/A test, and why run one?

Answer: An A/A test splits traffic into two groups that get the identical experience. It checks the plumbing. The split ratio should match the design (no SRM), metrics should show no systematic difference, and across many A/A runs or simulated re-splits, roughly alpha of them should come out "significant". If many more do, your variance estimate is wrong. Common causes are analysing at the wrong unit, heavy-tailed metrics or correlated observations. Platforms often run continuous A/A tests, or replay historical data through random splits, to validate every new metric before teams rely on it.

9. Which statistical test would you use for a conversion metric versus a revenue metric?

Answer: For a binary per-user conversion, a two-proportion z-test (or chi-square test) is standard at the sample sizes online tests use. For a continuous per-user metric such as revenue or minutes watched, Welch's t-test on per-user values works well at large samples because of the central limit theorem, even though the raw data is skewed. Very heavy tails still need care: use pre-registered winsorisation (capping extreme values) or a bootstrap. For ratio metrics where the denominator is not the randomisation unit, such as clicks per page view, use the delta method (Q20). Non-parametric tests like Mann-Whitney test a different hypothesis (about distribution ranks, not means), so do not swap them in when the business cares about average revenue.

Power, sample size and MDE

10. What is the minimum detectable effect (MDE), and how do you choose it?

Answer: The MDE is the smallest true effect your test can reliably detect at the chosen alpha and power. It is a design input, not a prediction of what will happen. Choose it from the business side: the smallest lift that would justify the cost of building, maintaining and rolling out the change. Then check whether the traffic you have can detect it in a reasonable time. Sample size scales with the inverse square of the MDE, so halving the MDE needs roughly four times the users. If the honest MDE needs more traffic than you have, change the design: use a more sensitive metric, apply variance reduction, run longer, or accept that the test can only catch large effects.

11. How do you calculate sample size for a conversion-rate test? Show it in Python.

Answer: You need the baseline rate, the MDE, alpha, power and the split ratio. Here is an illustrative example with made-up numbers: baseline checkout conversion of 0.040, and we want to detect a lift to 0.044 (a 10 per cent relative lift) with a two-sided alpha of 0.05 and power of 0.8.

import math
from statsmodels.stats.power import NormalIndPower
from statsmodels.stats.proportion import (
    proportion_effectsize)

p_control, p_treat = 0.040, 0.044
h = proportion_effectsize(p_treat, p_control)
n = NormalIndPower().solve_power(
    effect_size=h, alpha=0.05, power=0.8,
    ratio=1.0, alternative="two-sided")
print(round(h, 4), math.ceil(n))
# 0.0199 39455   (users per arm)

h_small = proportion_effectsize(0.042, 0.040)
n_small = NormalIndPower().solve_power(
    effect_size=h_small, alpha=0.05, power=0.8)
print(math.ceil(n_small))
# 154283   (about 4x for half the MDE)

Running this with statsmodels gives about 39,455 users per arm for the 0.4-point lift and about 154,283 per arm for a 0.2-point lift. That shows the inverse-square rule. proportion_effectsize returns Cohen's h, an arcsine-transformed difference. Other calculators that use a pooled-variance formula give slightly different but similar answers. Then divide the total by eligible daily traffic to get the duration, and round up to whole weeks.

Interview tip: Also do the reverse. "We only have 20,000 users per arm; what can we detect?" Calling solve_power(nobs1=20000, alpha=0.05, power=0.8) returns h of about 0.028. From a 0.040 baseline, that is a lift to roughly 0.0457, about a 14 per cent relative lift. Telling a team what their test cannot detect is a senior skill.

12. How does metric variance affect sample size, and how do you size a revenue test?

Answer: For a continuous metric, the standardised effect is the absolute effect divided by the standard deviation. Revenue per user is usually very skewed: most users spend nothing and a few spend a lot. The standard deviation is therefore large relative to the mean, and tests need far more users than conversion tests. An illustrative example: mean revenue per user of β‚Ή500, standard deviation β‚Ή1,500, and we want to detect a β‚Ή10 change.

import math
from statsmodels.stats.power import TTestIndPower
d = 10 / 1500
n = TTestIndPower().solve_power(
    effect_size=d, alpha=0.05, power=0.8)
print(math.ceil(n))     # 353200 per arm

That is about 353,200 users per arm, nearly nine times the conversion example. Estimate the standard deviation from recent pre-period data for the same eligible population, not from all users. Capping extreme values and applying CUPED (Q22) both shrink it.

13. How long should an experiment run?

Answer: Run it at least long enough to reach the planned sample size. In practice, also cover at least one full weekly cycle, and preferably two, because weekday and weekend users behave differently. Starting on a Monday and stopping on Thursday over-represents weekday users. Account for ramp-up time, and for the fact that daily unique users do not add up linearly, since returning users are counted once. Watch out for running too long as well. Cookie churn and logged-out users re-randomising can dilute the effect, and long tests delay learning. If there are novelty effects (Q23), you may need a longer window or a separate long-running holdout.

14. When would you use an unequal traffic split, and what does it cost?

Answer: Unequal splits, such as giving the treatment a smaller share, limit exposure for a risky change, protect revenue, or reflect limited capacity (only so many users can get the new GPU-backed feature). The cost is power. For a fixed total sample, a 50/50 split minimises variance of the difference, and a lopsided split needs more total traffic for the same MDE. In statsmodels, the ratio argument handles this. Common practice is to ramp gradually, for example a small exposure first to catch crashes and guardrail breaches, then move to the full planned split for the analysis period. Analyse only the period at the final, stable split, or use methods that handle the ramp properly. Pooling ramp days with different ratios can bias the result if behaviour changes over time (a Simpson's paradox risk).

Analysis and common pitfalls

15. What is peeking, and why does it inflate false positives?

Answer: Peeking means checking a fixed-horizon test repeatedly and stopping the first time p drops below 0.05. The fixed-horizon test assumes one look at a pre-planned sample. Each extra look gives random noise another chance to cross the threshold, so the real false-positive rate climbs well above alpha. With daily looks over a long test it can rise several-fold. The fix is not to ban dashboards. Either commit to analysing once at the planned sample size, with dashboards showing guardrails only, or use a method designed for repeated looks (Q16). Stopping early for harm on a guardrail is fine and expected. That is a safety decision, not a claim of a win.

16. What is sequential testing, and when would you use it?

Answer: Sequential methods control the false-positive rate even though you look at the data many times. Two families are common:

  • Group sequential designs: a pre-planned number of interim analyses, with alpha spent across them by a spending function. O'Brien-Fleming-style boundaries are very strict early and close to the usual threshold at the end. Pocock-style boundaries are equal at each look.
  • Always-valid inference: for example, the mixture sequential probability ratio test (mSPRT) and confidence sequences. These give p-values and intervals that stay valid at any stopping time, so you can monitor continuously.

The trade-off is power. For the same maximum sample size, a sequential test is somewhat less powerful at the final look than a fixed-horizon test. The benefit is that it can stop early when the effect is large. Use sequential testing when early stopping has real value (a harmful change, an expensive treatment) or when stakeholders will look anyway. Many experimentation platforms now offer it by default.

17. How do you handle multiple metrics, variants and segments without fooling yourself?

Answer: Every extra comparison is another chance of a false positive. If you check 20 independent metrics at alpha 0.05, you should expect about one false "win" by chance. Options:

  • Pre-register one primary metric. Treat secondary metrics and segments as exploratory unless corrected.
  • Family-wise error control: Bonferroni (divide alpha by the number of tests) is simple but conservative. Holm's step-down method is uniformly more powerful and still controls family-wise error.
  • False discovery rate: Benjamini-Hochberg controls the expected share of false discoveries among the effects you call significant. It suits metric scorecards with dozens of metrics.
  • For many variants against one control, Dunnett-style comparisons or a correction across arms.

Interview tip: Say which error you are controlling and why. Family-wise control suits a launch decision, and FDR suits screening many metrics for follow-up.

18. What is a sample ratio mismatch (SRM), and how do you test for it?

Answer: SRM means the observed split between arms differs from the designed split by more than chance allows. Test it with a chi-square goodness-of-fit test on the unit counts:

from scipy.stats import chisquare
obs = [50_000, 51_200]   # control, treatment
exp = [sum(obs) / 2] * 2
print(chisquare(obs, exp))
# statistic about 14.2, p about 0.00016

A gap of 1,200 users in about 101,000 looks small, but it is very unlikely under a true 50/50 split. With [50,000, 50,300] the p-value is about 0.34, which is unremarkable. Teams often alert at a strict threshold such as p < 0.001 because the test runs on every experiment. An SRM means assignment or logging is broken, so the arms are no longer comparable. Treat the result as invalid until you find the cause. Do not reweight it away. Scenario Q42 walks through the investigation.

19. What are the usual causes of SRM?

Answer: Common causes, roughly in the order I would check them:

  • Exposure logging differences: the treatment page is slower or crashes, so its exposure event fires less often. Or the treatment logs exposure at a different point in the flow.
  • Redirects: a treatment served by redirect loses users who bounce during the extra hop.
  • Bot and fraud filtering that interacts with the treatment, for example a new page that triggers a bot heuristic more often.
  • Triggering conditions that differ between arms, such as counting only users who saw a component that exists only in treatment.
  • Assignment bugs: hash collisions with another experiment, a changed salt mid-test, caching that serves one arm to the wrong users, or a ramp change that was not handled in the analysis.
  • Data pipeline issues: late-arriving events for one arm or a join that drops rows.

Slice the SRM by platform, browser, app version, country and day. A mismatch concentrated in one slice usually points straight at the cause.

20. What is a ratio metric, and why do you need the delta method?

Answer: A ratio metric divides two sums whose denominator is not the randomisation unit, for example click-through rate (clicks over page views) or revenue per session when you randomise by user. The page views or sessions from one user are correlated, so treating each page view as independent gives standard errors that are too small and too many false positives. The delta method uses a first-order Taylor approximation to get the variance of the ratio of means at the user level. For R = mean(Y)/mean(X), the variance is approximately [Var(Θ²) βˆ’ 2RΒ·Cov(Θ², XΜ„) + RΒ²Β·Var(XΜ„)] / mean(X)Β², where Y and X are per-user sums. It needs only per-user totals, so it is cheap to compute in SQL over a warehouse. The alternative is a user-level bootstrap (resample users, not page views), which is simpler to reason about but costs more at scale.

Real-world example: A news app randomises by user and reports "articles read per session". A naive session-level t-test shows significance. The delta-method interval at user level is wider and includes zero, because a few heavy readers dominate the session counts.

21. How do you deal with outliers and heavy-tailed metrics?

Answer: Decide the rule before the test and apply it equally to both arms. Common approaches are winsorising (capping values at a high percentile computed on the pooled data), using a capped version of the metric as the decision metric while still reporting the uncapped one, and removing known bot or test accounts using filters set before the experiment. Never drop outliers after looking at which arm they fall in. In B2B or marketplace data, a single large customer can swing revenue, so consider stratified randomisation or analysing with and without the largest accounts. If capping changes the conclusion, report both results and treat the outcome as fragile.

Advanced methods

22. What is CUPED, and how does it reduce variance?

Answer: CUPED (Controlled-experiment Using Pre-Experiment Data), popularised by Microsoft's experimentation team, adjusts each unit's outcome using a covariate measured before the experiment, usually the same metric in a pre-period. The adjusted metric is Yβ€² = Y βˆ’ ΞΈ(X βˆ’ mean(X)), with ΞΈ = Cov(Y, X) / Var(X). Because X is measured before assignment, it is independent of treatment, so the adjustment removes noise without biasing the effect. The variance shrinks by a factor of (1 βˆ’ ρ²), where ρ is the correlation between X and Y. In the revenue example from Q12, a correlation of 0.6 cuts the needed sample from about 353,200 to about 226,000 users per arm:

rho = 0.6
d_adj = 10 / (1500 * math.sqrt(1 - rho**2))
n = TTestIndPower().solve_power(
    effect_size=d_adj, alpha=0.05, power=0.8)
print(math.ceil(n))     # 226049 per arm

Users with no pre-period data (new users) get X set to a constant or a missing-data flag. CUPED is a special case of regression adjustment (ANCOVA). Extensions use several covariates or ML predictions of the outcome as the covariate.

Interview tip: Mention that CUPED helps most for metrics with stable user behaviour, like spend or engagement, and least for one-off events like first purchase.

23. What are novelty and primacy effects, and how do you detect them?

Answer: A novelty effect is a temporary lift because something is new and users explore it. A primacy effect (change aversion) is a temporary dip while existing users relearn a familiar interface. Either way, the effect in week one is not the long-run effect. To detect them:

  • Plot the daily treatment effect over time, by days since first exposure rather than calendar date. A decaying or recovering curve is the signal.
  • Compare new users (who have no old habit) with returning users.
  • Run longer, or keep a long-term holdout (Q29) for big changes.

Be careful reading a time trend. The population mix changes over a test, because heavy users arrive early and light users arrive later. So split the curve by cohort rather than trusting one line.

24. What is interference (a SUTVA violation), and where does it occur?

Answer: The stable unit treatment value assumption (SUTVA) says one unit's outcome depends only on its own assignment. Interference breaks this, and then the simple treatment-minus-control difference is biased. Typical cases:

  • Two-sided marketplaces: food delivery, ride-hailing or quick commerce. A treatment that makes treated customers order more uses up delivery partners that control customers also need. Control looks worse, so the test overstates the effect.
  • Social and communication products: a sharing feature in treatment changes what control users receive.
  • Shared budgets: in ads auctions, treatment and control campaigns compete for the same budget and inventory.
  • Recommender feedback loops: a treatment ranker changes which items get popular, which changes training data for the control model. The recommendation systems interview questions guide covers this loop in depth.

25. How do cluster-randomised and switchback designs work?

Answer: Both move the randomisation to a level where interference is mostly contained.

  • Cluster randomisation: assign whole groups, such as cities, delivery zones, stores, or graph clusters of connected users, to treatment or control. Spillover then happens mostly within a cluster. You have far fewer effective units, so power drops sharply. Analyse at the cluster level, or use cluster-robust standard errors, and consider pairing or stratifying similar clusters.
  • Switchback (time-split) designs: the whole market alternates between treatment and control in time windows, for example one-hour or half-day blocks, assigned randomly per region. This suits pricing, dispatch and matching algorithms where everyone in a market shares supply. Watch carryover between windows (orders placed under one policy finishing under the next). Use buffer periods, choose a window length longer than the carryover, and balance windows across hours of the day and days of the week.
Switchback, one city, randomised blocks
time:  09-12  12-15  15-18  18-21  21-24
city:   T      C      C      T      C
        |--buffer--| excluded from analysis

Real-world example: A food-delivery company testing a new dispatch algorithm in Hyderabad cannot randomise customers, because all orders share the same riders. A switchback across zones and time blocks, analysed per block, gives a much less biased estimate.

26. What are heterogeneous treatment effects, and how do you analyse them responsibly?

Answer: Heterogeneous treatment effects (HTE) mean the effect differs across users, for example positive for new users and negative for power users, or different on low-end Android devices. The conditional average treatment effect (CATE) is the effect for users with given characteristics. Methods include pre-registered segment analysis with interaction terms, meta-learners (S-, T- and X-learners), and causal forests. The traps are multiple testing (slice enough segments and something will look "significant") and segments defined by post-treatment behaviour, which breaks randomisation. Use only pre-treatment attributes, correct for multiple comparisons, and treat surprising segment findings as hypotheses for a follow-up test. HTE is valuable for deciding whether to launch to everyone, to a segment, or with personalisation. A positive average can hide harm to an important group.

27. Bayesian vs frequentist A/B testing: how do they differ, and which would you use?

Answer:

AspectFrequentistBayesian
Outputp-value, confidence intervalPosterior distribution, probability the treatment beats control, expected loss
InputsAlpha, power, sample sizePrior plus data; choice of prior matters at small samples
InterpretationError rates over repeated experimentsDirect probability statements about the effect, given the prior
Typical decision ruleReject null if p < alphaShip if expected loss falls below a threshold

Bayesian outputs are easier for stakeholders to read ("the probability this is better is high, and the expected loss if we are wrong is small"). Informative priors from past experiments can also shrink exaggerated estimates. But Bayesian testing is not a free fix for peeking. If you stop the first time "probability to beat control" crosses a fixed cut-off, the frequency of shipping null changes still rises with the number of looks. With large samples and weak priors, the two approaches usually agree. I would choose based on the organisation: consistent methods across teams, clear pre-registered decision rules, and good calibration matter more than the philosophy.

28. When would you use a multi-armed bandit instead of an A/B test?

Answer: A bandit (for example Thompson sampling or upper confidence bound) shifts traffic towards better-performing arms while the test runs. That minimises regret, meaning the value lost by showing worse options. A/B tests hold allocation fixed to get an unbiased, well-powered estimate of the effect. Use bandits when the goal is to optimise rather than learn: many short-lived options such as headlines, promotional banners or festival offers, where the reward is fast and you do not need a precise effect size. Use A/B tests when you need to understand the effect, check guardrails and long-term metrics, or make a lasting product decision. Bandits complicate inference, because adaptive allocation biases naive estimates, and they struggle with delayed rewards and non-stationary behaviour. Contextual bandits personalise the choice per user. The reinforcement learning interview questions guide covers the exploration theory behind them.

29. How do you measure long-term effects that a two-week test cannot capture?

Answer: Options include a long-term holdout (a small share of users kept on the old experience for months, which measures the combined effect of many launches), surrogate metrics (short-term metrics validated as predictors of long-term outcomes such as retention, combined in a surrogate index), and cohort follow-up of experiment participants after the test ends. Holdouts cost something, because those users miss improvements, and the holdout group must stay clean and not drift through cookie churn. Ads load and notification volume are classic cases: more ads or pushes lift short-term revenue or sessions but can erode long-term engagement. Only long-term measurement shows that trade-off.

30. What are triggering and dilution, and why do they matter?

Answer: If a change only affects users who reach a certain point, such as the payment page, analysing all assigned users dilutes the effect with people who never saw it, so power drops. Triggered analysis restricts the analysis to users who would have been exposed. Crucially, the trigger must be evaluated the same way in both arms: log "would have seen the new component" for control users too (a counterfactual trigger). Otherwise you compare different populations and can create SRM. Report the triggered effect, and also translate it to the overall effect (effect Γ— trigger rate) when estimating business impact, so a large triggered lift on a small population is not oversold.

If you want to build these skills hands-on, from statistics and SQL to ML models and cloud deployment, Cloudsoft's APEX AI, ML, Cloud and Cyber Security program uses labs and projects rather than slides.

A/B testing LLM and AI features

31. What is different about A/B testing an LLM-powered feature?

Answer: Several things change:

  • Outputs are non-deterministic and depend on prompts, model versions, retrieval indexes and tool behaviour. Pin all of these per arm and log versions with every exposure.
  • Quality is hard to measure directly. There is no simple "correct" label in production. Thumbs up or down is sparse and biased towards strong reactions.
  • Cost and latency vary per request, because of token counts, retries and tool calls, so they become first-class guardrails.
  • Safety and policy risks, such as hallucinated facts, unsafe content or data leakage, need their own guardrails and sometimes a kill switch.
  • Behaviour shifts over time as users learn to use the assistant, so novelty effects are common.

Offline evaluation comes first: a regression suite and evaluation set must pass before any online exposure. The online test then answers whether users get more value. The LLM evaluation guide covers the offline side.

32. How do you handle noisy quality metrics for AI features online?

Answer: Combine several signals, each with known weaknesses. These are behavioural outcomes (task completion, the user copying or accepting a suggestion, fewer follow-up rephrasings, escalation to a human agent, repeat usage), explicit feedback (thumbs and ratings, sparse and biased), and sampled quality scores from human review or an LLM judge. Choose a primary metric close to user value, such as successful resolutions per user, and treat the others as diagnostics. Quality scores from sampled review have extra variance because only a sample of sessions is scored. Account for that sampling in the standard error, and size the review sample using the same power logic as Q11. Define every metric at the randomisation unit (user), not per message, because messages within a conversation are highly correlated.

33. Can you use an LLM-as-judge metric in an online experiment? What are the risks?

Answer: Yes, as a scaled proxy for human review, with controls:

  • Calibrate the judge against human labels on a representative sample from your traffic, and measure agreement before using it in decisions.
  • Blind and symmetric: the judge must not know which arm produced the output, and both arms must be scored with the same judge prompt and model version, on the same sampling scheme.
  • Pin the judge version for the whole test. A judge model update mid-test shifts scores in a way that looks like a treatment effect over time.
  • Known biases: judges can favour longer answers, a particular style, or outputs from their own model family. If the treatment changes answer length, the judge may reward length, not quality. Check this with length-controlled comparisons.
  • Privacy: sending user conversations to a judge model is a data flow that needs the same consent, masking and retention controls as the feature itself.

Use judge scores as a guardrail or diagnostic, and keep a business or behavioural metric as the primary decision metric. The LLM evaluation interview questions guide goes deeper on judge design.

34. What cost and latency guardrails would you set for an LLM feature test?

Answer: Track cost per user and per successful task, not just per request, because a cheaper model that needs more retries may cost more per resolved task. Useful guardrails include input and output tokens per user, model and tool-call count per session, p95 and p99 end-to-end latency and time to first token, error and timeout rates, and fallback rate to a secondary model. Set explicit tolerances before launch, for example "cost per successful task must not rise beyond an agreed threshold", and alert automatically. Make sure costs are attributed to the right arm: logs from a shared gateway must carry the experiment ID. An LLM gateway that tags every call with arm and version makes this straightforward. A feature that lifts engagement while sharply raising cost per user may still be a no-ship. The decision memo should show both numbers side by side.

35. How would you A/B test a prompt or model change for an existing AI assistant?

Answer: First run offline evaluation on a fixed test set plus a replay of recent anonymised traffic, checking quality, safety and cost. Then shadow-run the new version on live traffic without showing users its output, to compare latency, cost and judge scores on real inputs. Only then randomise users for the online test, with a small initial exposure. Keep users in the same arm across conversations so their experience is consistent and carryover effects stay contained. Compare at user level on the pre-agreed primary metric, with guardrails on safety flags, escalations, cost and latency. Because model and prompt changes ship often, many teams keep a reusable "evaluation then shadow then experiment" pipeline rather than designing each test from scratch.

offline eval --> shadow on live --> small ramp
     |                |                 |
  regressions?   cost/latency?     guardrails?
     v                v                 v
              full A/B at user level
                       |
            decision memo + rollout

Experimentation platforms

36. What are the main components of an experimentation platform?

Answer: Whether bought, open source or built in-house, platforms tend to have the same building blocks:

  • Configuration and feature flags: define experiments, eligibility rules, arms, traffic allocation and ramp schedules, with an audit trail and a kill switch.
  • Assignment service: deterministic, fast bucketing (Q37), available on server, web and mobile SDKs, with consistent results across them.
  • Exposure logging: an event recorded when a unit actually experiences the arm, carrying experiment ID, arm and version. This usually flows through an event stream; the Kafka interview questions guide covers that layer.
  • Metric pipeline: joins exposures to outcome events in a warehouse, computes per-unit aggregates, and applies the delta method, CUPED and corrections. These are usually scheduled jobs, so orchestration with Airflow or similar is a common interview follow-up.
  • Scorecard and alerts: effects with intervals, automatic SRM checks, guardrail alerts and a decision log.
flags/config --> assignment SDK --> exposure events
                                         |
outcome events --> warehouse join <------+
                         |
            per-user metrics, CUPED, delta
                         |
          scorecard + SRM/guardrail alerts

37. How does deterministic assignment with hashing work?

Answer: Compute a hash of the unit ID combined with an experiment-specific salt, for example hash(salt + user_id), map it to one of many buckets (say 1,000 or 10,000), and assign bucket ranges to arms. The same user always gets the same arm without any database lookup, assignment is effectively random with respect to user attributes, and different salts make assignments independent across experiments. Pitfalls include reusing a salt (carryover from a previous test, since the same users stay grouped), changing allocation mid-test in a way that reshuffles existing users, ID inconsistency between logged-out and logged-in states, and weak hash functions or modulo bias. Validate with A/A tests and SRM monitoring.

38. How do you run many experiments at once without them interfering?

Answer: Most platforms allow overlapping experiments, with each user in many tests at once, because independent salts make assignments orthogonal. Interaction effects between unrelated tests are usually small. Where two changes clearly conflict, such as two tests that both change the same checkout button, put them in a mutually exclusive layer or domain so a user can only be in one of them. Teams monitor for interactions by checking a test's effect across the arms of another overlapping test. Exclusive layers cost traffic, so reserve them for genuine conflicts rather than putting every test in isolation.

Ethics and consent

39. What ethical questions should you ask before running an experiment?

Answer: Ask whether any arm could plausibly harm users: financial harm (pricing, credit or loan offers), wellbeing (manipulating emotional content, addictive notification patterns), access (withholding a safety or accessibility feature from control), or fairness (an arm that disadvantages a protected group). A widely discussed 2014 social-media study that changed the emotional tone of users' feeds without explicit consent is often cited as the case that made the industry treat experiment ethics seriously. Good practice includes a lightweight review process for sensitive tests, a clear statement of what users agreed to, minimising exposure to risky arms, stopping rules for harm, and never using dark patterns as a "variant". Testing whether a page loads faster is low risk. Testing different loan interest rates on different users needs legal, compliance and ethics review.

40. How does privacy regulation, such as India's DPDP Act, affect experimentation?

Answer: Experiments process personal data: identifiers, behaviour and sometimes conversation content. Under India's Digital Personal Data Protection Act and similar laws elsewhere, you need a lawful basis and a clear purpose for that processing, data minimisation, retention limits, and respect for withdrawn consent. Practically, that means using pseudonymous IDs in experiment tables, keeping exposure and outcome logs only as long as needed, respecting users who opt out of analytics or personalisation (and checking the opt-out does not cause SRM), and being careful when AI experiments send user text to third-party models or judges. Check current legal guidance with your privacy team rather than assuming. The DPDP Act for AI applications article summarises the obligations for AI products.

Real-world scenarios

41. A test showed a significant lift, but the follow-up test shows nothing. What happened?

Answer: The original result was probably a false positive, an overestimate, or specific to its context. Significant results from underpowered tests systematically overstate the effect (the "winner's curse", or Type M error), so a replication at the true, smaller effect often fails. Other common causes are peeking or stopping early, picking the one significant metric out of many, a novelty effect that had faded, or a different population or season in the second run.

What I would check:

  1. Whether the first test was pre-registered with a fixed sample size, or stopped when it looked good.
  2. Its power for a realistic effect size. If power was low, the significant estimate was likely inflated.
  3. How many metrics, segments and variants were examined, and whether corrections were applied.
  4. SRM, instrumentation changes and data quality in both tests.
  5. Differences in population, season, platform mix or traffic source between the runs.
  6. The effect over time in the first test, for novelty decay.

Production consideration: Combine both tests in a pre-planned meta-analysis rather than picking the one you like, and judge the pooled estimate against the MDE. Going forward, use realistic MDEs, sequential methods if people need to peek, and a culture where "flat" is an acceptable result. Twyman's law applies here: any figure that looks unusually interesting is probably wrong.

42. Your SRM check fires on a checkout experiment two days in. What do you do?

Answer: Stop trusting the results, tell stakeholders the test is under investigation, and find the cause before reading any metric. If there is any sign of user harm, such as crashes or payment failures, pause the treatment.

What I would check:

  1. Confirm the SRM is real: designed versus observed counts, the right unit, the right date range, and whether the allocation was changed during a ramp.
  2. Slice by platform, app version, browser, country and day to find where the mismatch sits.
  3. Compare the exposure logging point in each arm. Does the treatment log exposure later in the flow, after a slow load?
  4. Check crash rates, page load time and redirect drop-off in treatment.
  5. Check bot filtering and any eligibility or trigger logic that touches one arm differently.
  6. Check for conflicts with other experiments or a salt reused from an earlier test.

Production consideration: Once fixed, restart with fresh assignment (a new salt) rather than continuing, because the users already exposed are not a clean sample. Add an automated SRM gate so scorecards hide effects until the check passes.

43. Click-through on a recommendations widget rose significantly, but revenue did not move. How do you explain it?

Answer: Clicks are a means, not the goal. Several mechanisms can raise clicks without raising revenue. The widget may cannibalise clicks that would have happened elsewhere on the page. The new layout may attract curiosity clicks that do not convert. It may push cheaper items, so orders rise but basket value falls. Or the revenue test may simply be underpowered, since revenue has far higher variance than clicks (Q12), so a real but small revenue effect may be undetectable.

What I would check:

  1. The power and confidence interval for revenue. "No significant change" with a wide interval is not "no effect".
  2. Total clicks and conversions across the whole page, not just the widget, to check for cannibalisation.
  3. The funnel: click β†’ add to cart β†’ purchase, and average order value per arm.
  4. Return and cancellation rates, since revenue net of returns is what matters.
  5. Segments such as new versus returning users, and app versus web.

Production consideration: Agree up front whether the primary metric is revenue (or margin) per user, with CTR as a diagnostic. If revenue is too noisy, use CUPED and capping, or validate a surrogate metric such as add-to-cart per user against revenue on past experiments. The AI in retail and e-commerce article covers how such widgets connect to business outcomes.

44. How would you design an experiment for a new AI assistant inside a banking app?

Answer: Consider an illustrative bank adding a GenAI assistant that answers account and product questions using retrieval over policy documents. I would randomise at customer level, start with a small, eligible population (for example, customers who opted in to new features), and pre-register a primary metric close to value: self-service resolution rate, meaning sessions that end without a branch visit, call or human chat escalation within a few days. Guardrails would cover complaint and escalation rate, safety and compliance flags from sampled review, hallucinated policy answers (judge plus human audit), p95 latency, and cost per resolved query.

What I would check:

  1. Offline evaluation and red-teaming results before any exposure, including regulatory content checks.
  2. That the control arm still has the existing help journey, so we measure the incremental effect.
  3. Novelty: plot effects by days since first exposure, and run long enough to see repeat usage.
  4. Spillover: are treatment customers calling less, but control customers waiting longer in a shared call-centre queue? That is interference through a shared resource.
  5. Consent and data handling for conversation logs used in evaluation.

Production consideration: Keep a kill switch and a human-handoff path in treatment, and get compliance sign-off on the experiment design itself. Measure downstream outcomes, such as complaints filed later, through cohort follow-up. Taking assistants like this into a bank's real systems, with identity, security and audit controls, is the work Forward Deployed Engineers do, and Cloudsoft's FDE PRO course includes a Secure Banking AI Assistant project.

45. Marketing wants a pricing test that runs only during the Diwali week. What do you advise?

Answer: A Diwali-week test can answer "what works during Diwali" but not "what works in a normal week". Festival traffic differs in intent, basket size, gifting behaviour, device mix and new-user share, so results often do not generalise. The window is also short, may not cover a full weekly cycle, and often coincides with code freezes and heavy promotions.

What I would check:

  1. Whether the decision is festival-specific (a Diwali banner or offer). If it is, a short test is fine, but power must work within the window.
  2. Power at festival traffic levels and variance. Revenue variance is often higher during sales.
  3. Interaction with other running promotions and campaigns, and whether marketing will send traffic unevenly.
  4. Pricing ethics and consistency: different prices for different customers needs legal review and clear policy.
  5. Delayed outcomes such as returns and cancellations that arrive after the festival.

Production consideration: For festival-specific creative, a bandit may suit better than a fixed test (Q28). For lasting decisions, run during Diwali and repeat in a normal period, or keep a holdout running across both. Use CUPED with pre-festival data to gain power. Write down in the decision memo that the result is season-specific.

46. The test result is p = 0.07 and the product manager wants to ship. What do you say?

Answer: Move the conversation from the threshold to the decision. Show the confidence interval: perhaps βˆ’0.1 to +1.2 points on the primary metric, with guardrails flat. Then ask what being wrong would cost. If the change is cheap, reversible and has no guardrail harm, shipping a likely-positive change can be a reasonable business call, as long as it is recorded honestly as "not statistically significant at our pre-set threshold", not reported as a win. If the change is costly to maintain or risky, extend the test to the planned sample (if it ended early) or run a properly powered follow-up. What I would not do is re-slice the data until something crosses 0.05. A decision framework agreed in advance, such as "ship if the interval excludes harm and the point estimate is positive for low-cost changes", avoids this argument entirely.

47. You need to test a new feature in a B2B product with only 200 customer accounts. How do you get a usable answer?

Answer: With few units, a standard A/B test has very little power, and a few large accounts dominate. Use stratified or paired randomisation (pair accounts by size and industry, then randomise within pairs), CUPED with each account's pre-period usage, and a metric measured per account with many observations, such as weekly active seats share. Consider a within-account design where safe, for example a crossover where accounts receive the feature in different periods, accepting carryover risk. Report intervals rather than relying on significance alone, and combine the experiment with qualitative evidence from customer interviews. Be explicit that the test can only detect large effects. That is still useful for catching harm.

48. Overall the treatment wins, but it loses on both the app and web segments separately. How is that possible?

Answer: This is Simpson's paradox. It happens when the mix of segments differs between arms or across time, for example if treatment received more web traffic, which has higher conversion, during a ramp. Under clean randomisation with a stable split, segment shares should be similar in both arms, so this pattern is itself a red flag. Check for SRM within segments, ramp-period pooling with changing ratios, and a platform-specific bug that changes who gets logged. If the mix really is balanced and the paradox persists, recheck the calculation, because with balanced shares the overall effect is a weighted average of segment effects and cannot be outside their range.

49. A delivery app tests faster-delivery promises on customer-level randomisation and sees a big lift. Why might you distrust it?

Answer: Interference through shared supply. Treated customers see faster promises and order more. Riders and dark-store capacity are shared, so control customers may get slower actual deliveries and order less. The difference between arms exaggerates the true effect of a full rollout, and launching to everyone may produce no lift and worse delivery times.

What I would check:

  1. Control-arm delivery times and cancellations during the test compared with the pre-period.
  2. Rider utilisation and order backlog by zone.
  3. Whether the effect is larger in zones where supply was tight.

Production consideration: Re-run as a zone-level cluster test or a switchback (Q25), and add guardrails on actual delivery-time accuracy, since a promise the operations team cannot keep becomes a trust problem.

50. Your fraud team wants to A/B test a new transaction-risk model. What is special about this test?

Answer: Outcomes are delayed and adversarial. Chargebacks and confirmed fraud can arrive weeks later, fraudsters adapt to whichever arm is easier to beat, and a weaker arm causes real financial loss. Use a small treatment share at first, a long enough observation window to capture delayed labels, and guardrails on false declines, which hurt genuine customers and revenue, as well as on fraud losses. Randomise by customer or card, not by transaction, so fraudsters cannot simply retry until they land in the weaker arm. Often the safer first step is a shadow test, where the new model scores live traffic without taking actions, followed by champion-challenger routing of a small share. The fraud and anomaly detection interview questions guide covers the modelling side.

Key takeaways

  • Write the hypothesis, primary metric, guardrails, MDE and decision rule before launch, and randomise at the level where the treatment is experienced.
  • Sample size scales with the inverse square of the MDE. Run a power calculation, and tell stakeholders what the test cannot detect.
  • Explain p-values and confidence intervals correctly, and decide with intervals against a practical threshold, not with p < 0.05 alone.
  • Check validity first: SRM, A/A health, exposure logging and analysis at the right unit (delta method for ratio metrics).
  • Peeking, many metrics and segment-mining inflate false positives. Use sequential methods and multiple-testing corrections deliberately.
  • CUPED, triggering and capping buy power cheaply. Cluster and switchback designs handle interference at a cost in power.
  • For LLM features, combine behavioural metrics, calibrated and blinded judge scores, and hard cost, latency and safety guardrails.

Interview preparation checklist

  • Run a power calculation in statsmodels for a conversion and a revenue metric, and the reverse calculation (MDE from available traffic).
  • Simulate a hundred A/A tests in Python with daily peeking and count the false positives, then repeat with a single look.
  • Implement CUPED and the delta method on a synthetic dataset and compare the interval widths with naive methods.
  • Write SQL that computes per-user metrics from exposure and event tables; the SQL interview questions for data and AI guide is useful practice.
  • Prepare a one-page experiment design for a feature you know: hypothesis, unit, metrics, MDE, duration, guardrails, decision rule.
  • Rehearse explaining a p-value and a confidence interval to a non-technical stakeholder in under a minute.
  • Prepare stories: a broken experiment you caught, a flat result you defended, and a decision you made with ambiguous data.
  • Revise interference, switchbacks and bandits for marketplace and AI roles, and offline LLM evaluation for AI product roles.

FAQ

Which roles get A/B testing interview questions?

Data scientists, product analysts, ML engineers, growth engineers and product managers all commonly face experimentation questions. The depth varies: analysts focus on design and interpretation, data scientists add statistics and variance reduction, and platform engineers add assignment, logging and metric pipelines.

What statistics do I need for an experimentation interview?

You need hypothesis testing, p-values, confidence intervals, Type I and Type II errors, power and sample size, the central limit theorem, and the basics of regression. For senior roles, add sequential testing, multiple-testing corrections, the delta method, CUPED and causal inference ideas.

How should I prepare for A/B testing interview questions?

Practise designing experiments end to end for products you use, run power calculations and simulations in Python, and rehearse scenario answers in a fixed order: validity, result, decision, next checks. Interviewers value clear reasoning about trade-offs more than memorised formulas.

Do I need to know Bayesian A/B testing?

You should be able to explain how Bayesian results differ from frequentist ones, what a prior does, and why Bayesian methods do not remove the risks of frequent stopping. Deep Bayesian modelling is needed mainly for teams whose platforms use it.

Are A/B testing questions asked in product analytics interviews in India?

Yes, experimentation is a common topic in product analytics loops at Indian product companies and global capability centres in Hyderabad and Bengaluru, often combined with SQL and metric-design case questions.

What tools are used for running online experiments?

Teams use feature-flag and experimentation platforms, either commercial, open source or built in-house, alongside a data warehouse, SQL and Python for analysis. Interviewers usually care more about the concepts than about a specific tool.

How is A/B testing AI features different from testing UI changes?

AI features add non-deterministic outputs, noisy quality metrics, per-request cost and latency, and safety risks. You need offline evaluation before exposure, calibrated quality measures online, and explicit cost and safety guardrails alongside the usual product metrics.

Can freshers answer A/B testing interview questions well?

Yes. Freshers who can define a hypothesis, pick sensible metrics, run a basic power calculation and correctly explain a p-value and confidence interval stand out. A small simulation project on peeking or SRM shows real understanding.

Is experimentation a good career skill for data professionals?

Experimentation connects analysis to decisions, so it is valued across product, growth, marketing and AI teams. The skills transfer to causal inference, marketplace analytics and evaluating AI systems in production.

Want to go from statistics on paper to experiments that drive real product decisions? Explore Cloudsoft's APEX program for hands-on machine learning, AI, cloud and data skills. Classroom training is in Ameerpet, Hyderabad, or live online. Call +91 96660 19191 for a free demo.

Share𝕏infβœ‰
EnrollWhatsAppCall us