Time series interview questions in 2026 test whether you can produce a forecast that a planner, a finance team or a capacity engineer will actually act on: a defensible baseline, the right model family for the data, an honest backtest and an uncertainty range, not just a low error on one lucky split. This guide collects 55 high-value time series and forecasting interview questions with model answers, from trend, seasonality and stationarity through ETS, ARIMA, gradient boosting, deep learning and time-series foundation models, to hierarchical and intermittent demand, evaluation metrics, production operations and real-world scenarios. It is written for data scientists, ML engineers and analysts preparing for demand forecasting, supply chain, FinOps and capacity planning roles.
How to use this guide
General ML theory such as bias-variance and regularisation is covered in our machine learning interview questions, and statistics and experimentation basics are in the data science interview questions, so this page stays on forecasting. What interviewers commonly probe at each level:
- Freshers and analysts: components of a series, stationarity, ACF and PACF, naive baselines, why a random train/test split is wrong, and MAE versus MAPE.
- Mid-level data scientists: ETS and ARIMA intuition, lag features without leakage, rolling-origin backtesting, holiday and promotion regressors, and prediction intervals.
- Senior engineers and leads: global models over thousands of series, hierarchical reconciliation, intermittent demand, foundation models versus trained models, retraining and drift, and translating forecast error into inventory or capacity cost.
The Python snippets use pandas and statsmodels and were run on synthetic daily data; treat them as illustrations of the idea, not production code.
- Fundamentals: components, stationarity and baselines (Q1βQ10)
- Classical models: ETS, ARIMA and Prophet (Q11βQ17)
- Machine learning, deep learning and foundation models (Q18βQ26)
- Hierarchies, intermittent demand and external regressors (Q27βQ32)
- Evaluation, metrics and uncertainty (Q33βQ38)
- Anomaly detection and production forecasting (Q39βQ44)
- Real-world scenario questions (Q45βQ55)
- Key takeaways
- Interview preparation checklist
- FAQ
Fundamentals: components, stationarity and baselines
1. What makes time series forecasting different from ordinary regression?
Answer: Observations are ordered and dependent, so the past is the main source of information and the future is never available at training time. That one fact changes everything downstream. You cannot shuffle rows into a random train/test split, because the model would learn from the future. Errors are autocorrelated, so standard regression assumptions about independent residuals break. The target distribution itself shifts over time through trend, seasonality, structural breaks and policy changes. And the output is usually a path over a horizon (the next 28 days, the next 12 weeks) rather than a single value, with accuracy that degrades as the horizon grows.
Interview tip: Mention the decision the forecast feeds. A forecast for store replenishment, cash in ATMs or GPU purchase orders has different horizons, granularity and cost of error, and a strong candidate asks about that before naming a model.
2. What are the components of a time series?
Answer: The classic decomposition has four parts. Trend is the long-run direction (steady growth in an e-commerce category). Seasonality is a pattern that repeats at a fixed, known period: day-of-week, month-of-year, hour-of-day. Cycles are rises and falls without a fixed period, usually tied to business or economic conditions; they are often confused with seasonality, and the difference is that you cannot predict when a cycle will turn from the calendar alone. Noise (the remainder) is what is left after the systematic parts are removed. Real series often add level shifts (a new warehouse opens), calendar effects (moving festivals, month-end salary days) and outliers.
Real-world example: Daily UPI-style payment volumes for a bank show upward trend, weekly seasonality, a month-start spike when salaries land, and festival peaks whose Gregorian date moves each year. A model that only knows "weekly plus yearly" seasonality will misplace the festival peak.
3. When is seasonality additive and when is it multiplicative? How do transformations help?
Answer: Additive seasonality adds a roughly constant amount each season regardless of level (weekends add about the same number of units whether the store is small or large). Multiplicative seasonality scales with the level (weekends add a percentage, so the swing grows as sales grow). You spot multiplicative behaviour when the seasonal amplitude widens as the series rises. A log transform turns multiplicative structure into additive and also stabilises variance; Box-Cox generalises this with a tunable parameter. Remember to back-transform forecasts, and note that back-transforming the mean of a log forecast gives roughly the median, not the mean, unless you apply a bias adjustment.
4. What is stationarity and why do classical models care about it?
Answer: A (weakly) stationary series has a constant mean, constant variance and an autocorrelation structure that depends only on the lag, not on time. ARIMA-type models assume the series, after differencing, is stationary, because their parameters describe a stable relationship between the present and the past. If the mean drifts, a model fitted on the early period will systematically miss later. Trend and seasonality are the usual sources of non-stationarity; changing variance (volatility clustering) is another. Tree-based ML models do not formally require stationarity, but they cannot extrapolate beyond the target range seen in training, so detrending or predicting differences and ratios still matters in practice.
5. How do you test for stationarity, and what does differencing do?
Answer: Plot first: rolling mean and rolling standard deviation over time reveal most problems. Then use tests. The Augmented Dickey-Fuller (ADF) test has a null hypothesis of a unit root (non-stationary), so a small p-value supports stationarity. The KPSS test reverses the null (stationary), so using both catches ambiguous cases. First differencing (yt β ytβ1) removes a stochastic trend; seasonal differencing (yt β ytβ7 for daily data with weekly seasonality) removes a stable seasonal pattern. Over-differencing adds noise and a characteristic negative autocorrelation at lag 1, so difference only as much as needed.
from statsmodels.tsa.stattools import adfuller
p_raw = adfuller(y)[1]
p_diff = adfuller(y.diff().dropna())[1]
# synthetic trending series: p_raw high, p_diff near 0
6. How do you read ACF and PACF plots?
Answer: The autocorrelation function (ACF) shows the correlation between the series and its own lags. The partial autocorrelation function (PACF) shows the correlation at lag k after removing the effect of the shorter lags. On a stationary series, the textbook pattern is: an AR(p) process has a PACF that cuts off after lag p and an ACF that decays; an MA(q) process has an ACF that cuts off after lag q and a PACF that decays. Spikes at lags 7, 14, 21 in daily data point to weekly seasonality. A slowly decaying ACF on the raw series is a sign that it needs differencing. In practice, mixed ARMA processes rarely give clean cut-offs, so the plots narrow the search and information criteria finish it.
7. What baseline forecasts should you always build, and why?
Answer: Three cheap baselines set the bar every model must beat. The naive forecast repeats the last observed value; it is surprisingly strong for random-walk-like series such as prices. The seasonal naive forecast repeats the value from one season ago (same weekday last week, same month last year). A moving average (or the mean of the last k periods) smooths noise for flat series. If a complex model cannot beat seasonal naive in a proper backtest, either the series has little predictable structure or the model is leaking or overfitting. Baselines also anchor metrics such as MASE (Q36) and make results explainable to business users.
train, test = y[:-28], y[-28:]
fc = np.tile(train.iloc[-7:].to_numpy(), 4) # weekly
mae = np.mean(np.abs(test.to_numpy() - fc))
Interview tip: Saying "I start with seasonal naive and report how much each model improves on it" is one of the clearest signals of forecasting maturity you can give.
8. What are horizon, granularity and lead time, and how do they shape the design?
Answer: Horizon is how far ahead you forecast. Granularity is the time bucket (hourly, daily, weekly) and the entity level (SKU-store, category-region). Lead time is the gap between when the forecast is produced and when the decision locks: if suppliers need orders 14 days ahead, the forecast that matters is the one for days 15 to 21, made today, and your backtest must evaluate exactly that offset. Finer granularity is noisier; coarser granularity hides patterns the business needs. The right answer comes from the decision: replenishment wants SKU-store-day, budget planning wants category-month.
9. How do you handle missing values, irregular timestamps and resampling?
Answer: First distinguish "missing" from "zero". A missing day because the data feed failed is different from a day with zero sales, and treating one as the other corrupts intermittent demand models. Reindex to a complete regular calendar (asfreq or resample in pandas), then fill deliberately: forward fill for slowly changing states, interpolation for smooth sensor signals, explicit zeros only where zero is real, and a flag column so the model knows a value was imputed. For irregular event data (transactions, logs), aggregate to a regular bucket with the right function: sum for volumes, mean or last for levels, max for peak load. Also align time zones before aggregating; a UTC-versus-IST mismatch shifts daily boundaries by five and a half hours and smears the daily pattern.
10. What is the difference between a point forecast and a probabilistic forecast?
Answer: A point forecast is one number per period, usually an estimate of the mean or median. A probabilistic forecast describes the range of outcomes: quantiles (P10, P50, P90), prediction intervals or a full distribution. Most business decisions are asymmetric, so they need the distribution. Safety stock depends on how bad the high-demand tail is; GPU capacity depends on the probability of running out; staffing depends on the chance of a queue overflow. A team that publishes only point forecasts forces planners to invent their own buffers, usually inconsistently.
Classical models: ETS, ARIMA and Prophet
11. Explain simple exponential smoothing, Holt and Holt-Winters.
Answer: Simple exponential smoothing forecasts a flat level that is a weighted average of past observations, with weights decaying geometrically; the smoothing parameter alpha controls how quickly it reacts. Holt's linear method adds a trend component with its own smoothing parameter, and a damped trend variant flattens the trend over longer horizons, which is usually safer than extrapolating a straight line forever. Holt-Winters adds a seasonal component, additive or multiplicative. These methods are fast, need little data, handle thousands of series cheaply and are hard to beat on short, regular series.
from statsmodels.tsa.holtwinters import ExponentialSmoothing
ets = ExponentialSmoothing(train, trend="add",
seasonal="add", seasonal_periods=7).fit()
fc = ets.forecast(28)
12. What does the ETS framework add over plain exponential smoothing?
Answer: ETS (Error, Trend, Seasonal) recasts exponential smoothing as a family of state space models, each labelled by its error type (additive or multiplicative), trend type (none, additive, damped) and seasonal type (none, additive, multiplicative). Because each is a proper statistical model with a likelihood, you can estimate parameters by maximum likelihood, compare candidates with information criteria such as AIC, and generate prediction intervals from the model rather than guessing. Automatic ETS selection across that family is a strong default for a large portfolio of regular series. Some combinations, such as multiplicative error with additive seasonality on series near zero, are numerically unstable and usually excluded.
13. Explain ARIMA intuitively. What are p, d and q?
Answer: ARIMA models the differenced series as a combination of its own past values and past forecast errors. d is the number of differences needed to make the series stationary. p is the autoregressive order: how many past values feed the prediction ("today looks like yesterday and the day before"). q is the moving-average order: how many past shocks feed it ("yesterday's surprise still echoes today"). The intuition is that ETS describes the series by its level, trend and seasonal states, while ARIMA describes its autocorrelation structure. Some ETS and ARIMA models are mathematically equivalent; neither family is strictly more general.
14. How does SARIMA extend ARIMA, and how do you choose the orders?
Answer: SARIMA adds seasonal terms (P, D, Q) at the seasonal period m, written ARIMA(p,d,q)(P,D,Q)m. Seasonal differencing (D) removes a stable seasonal pattern; seasonal AR and MA terms capture correlation at lags m, 2m and so on. To choose orders: decide d and D from plots and unit-root tests, use ACF/PACF to shortlist p, q, P, Q, fit candidates and compare AIC or AICc on the same differenced data, then check that residuals look like white noise (Ljung-Box test, residual ACF). Automated stepwise search (auto-ARIMA style) does this at scale. Keep orders small; high-order models overfit and become unstable.
from statsmodels.tsa.statespace.sarimax import SARIMAX
m = SARIMAX(train, order=(1, 1, 1),
seasonal_order=(0, 1, 1, 7)).fit(disp=False)
pred = m.get_forecast(28)
band = pred.conf_int(alpha=0.2) # 80 per cent interval
Interview tip: SARIMA handles one seasonal period well. For hourly data with daily, weekly and yearly patterns, say you would switch to Fourier terms as regressors, MSTL decomposition or a model designed for multiple seasonalities (Q15).
15. How do you model multiple seasonalities?
Answer: Hourly electricity load or call-centre volume has daily, weekly and yearly cycles at once. Options: Fourier terms (pairs of sine and cosine features at each period) added as regressors to a regression or ARIMA-with-regressors model, where the number of pairs controls smoothness; MSTL, which extends STL decomposition to several seasonal periods, after which you forecast the seasonally adjusted series and add the seasonal parts back; TBATS-style models built for complex seasonality; or tree and neural models with hour-of-day, day-of-week and day-of-year features. STL itself (Seasonal-Trend decomposition using LOESS) is also a useful diagnostic: it shows how strong each component is before you pick a model.
16. What is Prophet, and when would you use or avoid it?
Answer: Prophet is an open-source forecasting library released by Meta (originally Facebook). It fits an additive model: a piecewise-linear or logistic-growth trend with automatically detected changepoints, Fourier-based yearly and weekly seasonality, a holiday component from a user-supplied calendar, and optional extra regressors. Its strengths are analyst-friendliness, robustness to missing data and outliers, and easy injection of holiday and event knowledge, which makes it a reasonable first model for business series with strong calendar effects. Weaknesses: it does not model autocorrelation in residuals the way ARIMA does, its trend extrapolation can be overconfident, and on many benchmarks well-tuned ETS or a global gradient-boosting model is as good or better. Treat it as one candidate in the backtest, not a default.
17. When do classical statistical models beat machine learning?
Answer: When series are short, few in number, regular and driven mainly by their own history. With two years of monthly data you have 24 points; a gradient-boosting model has almost nothing to learn from, while ETS needs only a handful of parameters. Classical models also win when you need cheap, fast, explainable forecasts for many independent series, or calibrated intervals with little engineering. ML tends to win when you have many related series (so a global model can share patterns), rich covariates (price, promotions, weather, events) and enough history to learn interactions. A combination of a statistical model and an ML model is often more robust than either alone.
Machine learning, deep learning and foundation models
18. How do you turn a forecasting problem into a supervised learning problem?
Answer: Each row becomes "the state of the world at forecast time" plus the target value at some future step. Features usually include lags of the target (tβ7, tβ14, tβ364), rolling statistics (mean, std, min, max over windows), calendar features (day of week, month, week of year, days to the next festival, month-end flag), known-future covariates (planned price, planned promotion, holiday flags) and static attributes (store size, category, region). The key rule is that every feature must be computable at the moment the forecast would be made in production.
df = y.to_frame("sales")
df["lag_7"] = df["sales"].shift(7)
df["roll_mean_28"] = (df["sales"].shift(1)
.rolling(28).mean())
df["dow"] = df.index.dayofweek
Note the shift(1) before rolling: without it, the rolling mean includes the value you are trying to predict.
19. What are the common leakage pitfalls in time series ML?
Answer: Leakage is the main reason a forecasting model looks excellent offline and disappoints in production. Common sources:
- Random K-fold splits that let the model train on days after the test day.
- Rolling windows that include the target (missing
shift), or centred windows. - Lags shorter than the horizon: a lag-1 feature is fine for a one-day-ahead model, but for a 14-day-ahead forecast the most recent value you will actually have is lag 14.
- Using actuals for covariates that are only forecasts in production, such as observed temperature or realised promotion uplift.
- Target encoding or scaling fitted on the full dataset, including the test period.
- Revised data: training on the cleaned, back-filled history that did not exist at decision time (late-arriving returns, restated sales).
Interview tip: Describe an "as-of" feature pipeline: every feature is built from data available at the forecast timestamp, and the backtest replays that timestamp.
20. What is the difference between local and global forecasting models?
Answer: A local model is fitted separately per series (one ETS per SKU). A global model is one model trained across many series at once, using series identifiers and attributes as features. Global models share patterns across series, so a new or sparse SKU benefits from the behaviour of similar ones; they also scale operationally (one model to train, monitor and deploy rather than ten thousand). The costs: series with very different scales must be normalised (divide by each series' recent mean, or predict ratios), and a global model can underfit an unusual series that a local model would capture. A common production design is a global gradient-boosting or neural model plus local statistical models as fallbacks and benchmarks.
21. How would you set up gradient boosting (LightGBM, XGBoost, CatBoost) for demand forecasting?
Answer: One global model over all series, with lag, rolling, calendar, price, promotion and static features. Practical choices that matter:
- Target scaling per series, or a log1p transform, so high-volume items do not dominate the loss.
- Loss function matched to the data: Tweedie or Poisson objectives suit non-negative, zero-heavy demand; quantile loss gives P10/P50/P90 directly.
- No extrapolation: trees predict within the range seen in training, so a growing series needs a detrended or ratio target.
- Time-based validation for early stopping and hyperparameters, never random folds.
- Feature importance sanity checks: if one lag dominates, check for leakage.
This setup has a strong record in public retail forecasting competitions, which is why interviewers expect you to know it well.
22. Compare recursive, direct and multi-output forecasting strategies.
Answer: Recursive: train a one-step model and feed its predictions back as lags to roll forward. Simple and one model, but errors compound over the horizon. Direct: train a separate model (or a model with a horizon feature) for each step ahead, using only lags available at that offset. No compounding, but more models and possibly inconsistent paths. Multi-output (MIMO): one model predicts the whole horizon vector at once, as many neural forecasters do, keeping steps consistent. A hybrid that is common in practice is direct models for horizon buckets (days 1β7, 8β14, 15β28), each with lags no shorter than the bucket's start.
23. How do RNNs, LSTMs and temporal convolutional networks apply to forecasting?
Answer: RNNs process a sequence step by step with a hidden state; LSTMs and GRUs add gates that help retain information over longer spans and ease vanishing gradients. In forecasting they are usually used as global models over many series, often in an encoder-decoder layout that reads the history and emits the horizon, sometimes outputting distribution parameters for probabilistic forecasts (the DeepAR approach from Amazon research is a well-known example). Temporal convolutional networks (TCNs) use causal, dilated 1-D convolutions: causal so no output sees the future, dilated so the receptive field grows exponentially with depth. TCNs train in parallel and are often faster than RNNs. Both need plenty of data, careful scaling and more tuning than tree models; on small tabular-style problems they frequently lose to gradient boosting. The deep learning interview questions cover the training mechanics behind these architectures.
24. How are transformers used for time series, and what are their limits?
Answer: Transformers use attention to relate every time step to every other, which can capture long-range dependencies and mix in covariates. Well-known designs include the Temporal Fusion Transformer, which combines recurrent layers, attention and variable-selection gates for interpretable multi-horizon forecasts with static, known-future and observed inputs, and patch-based models that split a series into fixed-length patches treated as tokens, which shortens sequences and tends to work better than one token per time step. Limits: attention cost grows with sequence length, the models need lots of data, and research has shown that simple linear models can match some transformer architectures on standard long-horizon benchmarks. In an interview, show that you would benchmark a transformer against seasonal naive, ETS and gradient boosting before trusting it.
25. What are time-series foundation models?
Answer: They are large models pretrained on very large and diverse collections of time series so they can forecast a new series zero-shot: you pass recent history (the context) and get a forecast, without training on your data. Two prominent open examples: Chronos from Amazon Science, a family that originally tokenised scaled and quantised values and trained language-model architectures on them, with later generations (Chronos-Bolt, Chronos-2) adding faster patch-based designs and, in Chronos-2, multivariate and covariate support; and TimesFM from Google Research, a decoder-only forecasting model, which is also exposed through BigQuery's AI.FORECAST function. Strengths: no per-dataset training, decent quality on cold-start and long-tail series, built-in probabilistic outputs. Weaknesses: limited context length, inference cost per series, weaker use of business covariates than a purpose-built model, and pretraining data that may not reflect your domain. Model versions change quickly, so check the current model cards before quoting capabilities.
26. How would you choose between a foundation model, gradient boosting and a statistical model?
Answer: Let the backtest decide, but frame the shortlist by situation:
| Situation | Usually strong candidates | Why |
|---|---|---|
| Few, short, regular series | ETS, ARIMA, seasonal naive | Few parameters, calibrated intervals |
| Thousands of related series with price, promo, event data | Global gradient boosting | Learns covariate interactions, shares patterns |
| Very large datasets, multi-horizon, many covariates | Global neural models (DeepAR-style, TFT) | Joint horizon, probabilistic outputs |
| New series, little history, quick POC | Foundation model zero-shot | No training, reasonable quality fast |
| Long-tail, intermittent items | Croston family, aggregation, global model | Handles zeros sensibly (Q29) |
Also weigh operating cost, latency, explainability for planners and the team's ability to maintain the model. An ensemble of two different families is a robust default when the backtest is close.
Hierarchies, intermittent demand and external regressors
27. What is hierarchical forecasting, and what are the basic approaches?
Answer: Business series form hierarchies (SKU β category β department β total; store β city β region β national) and often grouped structures that cross them (SKU by channel by region). Different teams consume different levels, and forecasts must add up, or the regional plan will not match the sum of store plans. Approaches: bottom-up forecasts the lowest level and sums (keeps detail but bottom series are noisy); top-down forecasts the total and splits it by historical or forecast proportions (stable totals, but loses local patterns); middle-out forecasts a middle level and does both. These are special cases of reconciliation (Q28).
28. What is forecast reconciliation, and why does MinT come up in interviews?
Answer: Reconciliation takes independent "base" forecasts at every level, which generally do not add up, and adjusts them so they are coherent with the hierarchy. Formally the reconciled forecasts are a projection of the base forecasts onto the space of coherent forecasts defined by the summing matrix. OLS reconciliation treats all levels equally; WLS weights by each series' error variance or structure; MinT (minimum trace) uses an estimate of the base forecast error covariance to minimise the total variance of reconciled errors. The practical point: reconciliation often improves accuracy at several levels at once because it pools information, and it gives the business a single consistent set of numbers. Be ready to mention that reconciliation can produce negative values for low-volume series, which need non-negativity handling.
Total
/ \
North South base forecasts at every
/ \ / \ node, then reconcile so
S1 S2 S3 S4 children sum to parents
29. What is intermittent demand, and how do you forecast it?
Answer: Intermittent demand has many zero periods with occasional non-zero sales: spare parts, slow-moving SKUs, speciality medicines. Standard models forecast a smeared small value every day that is never right, and percentage metrics break on zeros. Methods: Croston's method separately smooths the size of non-zero demands and the interval between them, forecasting their ratio as a demand rate; the SBA (Syntetos-Boylan approximation) corrects Croston's known bias; TSB smooths the probability of demand instead of the interval, which lets the forecast decay for items going obsolete; temporal aggregation (forecast weekly or monthly, then disaggregate) reduces the zero share; and global models with Tweedie or Poisson loss work well when many sparse series share structure. Classify items first (by average inter-demand interval and variability of demand size) so each class gets an appropriate method.
Interview tip: For intermittent items, the decision is usually a stock level, so evaluate with service-level or quantile metrics rather than per-period accuracy.
30. How do you model holidays and festivals such as Diwali, whose dates move every year?
Answer: Build an explicit event calendar rather than relying on yearly seasonality, because Diwali, Eid, Holi, Pongal and Sankranti follow lunar or lunisolar calendars and shift across Gregorian dates, and many festivals are regional (Onam matters in Kerala, Durga Puja in West Bengal, Bathukamma and Bonalu in Telangana). Encode each event as features: a flag, days before and after the event (demand ramps up in the weeks before Diwali and drops after), event-window indicators, and region-specific flags joined by store location. Also encode adjacent effects: long weekends, salary dates, and sale events tied to festivals. In Prophet this is the holidays table with lower and upper windows; in a gradient-boosting model they are just features. Keep the calendar as versioned reference data reviewed with the business each year.
31. How do you handle promotions, price changes and cannibalisation?
Answer: Promotions are known-future covariates when they are planned in advance, so include promotion type, discount depth, display and marketing flags, and price relative to the item's regular price. Watch for three effects: the uplift during the promotion, the pull-forward dip after it (customers stocked up), and cannibalisation or halo effects on related items (a discount on one biscuit brand lowers sales of another). Cross-item features such as "a substitute in the same category is on promotion" capture cannibalisation in a global model. The data trap is that historical promotions were not random: planners promoted items they expected to sell, so the model can confuse selection with effect. For pricing decisions, as opposed to forecasting under a given plan, you need causal methods or experiments, not just a forecaster.
32. How would you use weather or monsoon data as a regressor?
Answer: Weather drives demand for beverages, ice cream, umbrellas, air conditioners, electricity and some medicines; monsoon onset and rainfall also affect agricultural inputs and rural consumption. The catch is availability: in production you only have weather forecasts, not observed weather, and their skill falls with horizon. So train on what you will have (archived forecasts as of each date, if you can obtain them) or at least test how the model behaves when observed weather is replaced by forecasts. For long horizons, use climatology or seasonal outlook categories rather than daily values. Aggregate spatially to match the business unit (district rainfall for a district's stores). Always check, in a backtest, that the regressor improves accuracy at the horizons the business uses; a regressor that helps day 1 but adds noise at week 4 should be horizon-specific.
Evaluation, metrics and uncertainty
33. How do you cross-validate a forecasting model?
Answer: Use rolling-origin (time-based) backtesting: choose several forecast origins, train only on data before each origin, forecast the horizon after it, and average errors across origins. An expanding window keeps all past data; a sliding window keeps a fixed length, which suits series whose behaviour changes. Leave a gap equal to the lead time between training end and the evaluated window if decisions lock in advance. Include origins that cover hard periods (festive season, year-end) and report errors by horizon step, because day-1 and day-28 accuracy can differ a lot.
def backtest(y, horizon=28, folds=4):
errs = []
for k in range(folds, 0, -1):
cut = len(y) - k * horizon
tr = y.iloc[:cut]
te = y.iloc[cut:cut + horizon]
f = ExponentialSmoothing(
tr, trend="add", seasonal="add",
seasonal_periods=7).fit().forecast(horizon)
errs.append(np.mean(np.abs(
te.to_numpy() - f.to_numpy())))
return errs
origin 1: [train......][test]
origin 2: [train..........][test]
origin 3: [train..............][test]
time --------------------->
34. MAE or RMSE: which should you use?
Answer: MAE (mean absolute error) treats all errors linearly and is minimised by forecasting the median. RMSE (root mean squared error) penalises large errors heavily and is minimised by the mean. Choose based on the cost of error: if one huge miss (an empty shelf on Diwali, a capacity outage) is far worse than several small ones, RMSE reflects that better; if costs are roughly linear and the data has outliers, MAE is more robust. Both are scale-dependent, so they cannot be averaged across series of different sizes without scaling. Matching the training loss to the evaluation metric matters: training on squared error and evaluating on MAE on skewed, zero-heavy data penalises you unnecessarily.
35. What is wrong with MAPE?
Answer: MAPE (mean absolute percentage error) is popular because business users understand percentages, but it has serious problems. It is undefined when actuals are zero and explodes when actuals are small, so slow movers dominate the average. It is asymmetric: over-forecasting can produce errors far above 100 per cent while under-forecasting is capped at 100, so optimising MAPE pushes forecasts low, which is dangerous for inventory. And averaging per-series MAPE weights a tiny SKU the same as a high-revenue one. If the business insists on a percentage, use WAPE (sum of absolute errors divided by sum of actuals), which is stable and volume-weighted.
36. Explain sMAPE, MASE and WAPE.
Answer: sMAPE divides the absolute error by the average of the absolute actual and forecast, which bounds it, but it is still unstable near zero and, despite the name, not truly symmetric between over- and under-forecasting. MASE (mean absolute scaled error) divides the model's MAE by the in-sample MAE of a naive or seasonal naive forecast; a value below one means the model beats the baseline, and because it is scale-free it can be averaged across series. It is generally the most defensible choice for comparing models across a portfolio. WAPE (also called MAD/mean ratio) is total absolute error over total actuals, which business teams read easily and which weights by volume. A common reporting set: WAPE for executives, MASE for model selection, bias for planners, and quantile loss for probabilistic forecasts.
scale = np.mean(np.abs(train.diff(7).dropna()))
mase = mae / scale # below 1 beats seasonal naive
37. How do you produce and evaluate prediction intervals?
Answer: Sources of intervals: analytical intervals from state space models (ETS, ARIMA); quantile regression (a gradient-boosting model per quantile with pinball loss); distributional neural models that output distribution parameters; bootstrapped residuals or simulated sample paths; and conformal prediction, which calibrates intervals from backtest residuals and adapts well to any point model. Evaluate with coverage (does the 80 per cent interval contain the actual about 80 per cent of the time in the backtest?), interval width (narrower is better at the same coverage), and pinball (quantile) loss or the weighted interval score, which reward both. Model-based intervals usually understate uncertainty because they ignore model and parameter uncertainty, so check coverage empirically.
Interview tip: Sum of quantiles is not a quantile of the sum. If a planner needs the P90 for a whole region, aggregate sample paths or forecast at that level; do not add up store-level P90s.
38. What is forecast bias, and what is forecast value added?
Answer: Bias is systematic over- or under-forecasting, measured as mean error or total forecast divided by total actual. A model with decent accuracy but consistent under-forecasting creates stockouts week after week, so track bias separately from accuracy. Forecast value added (FVA) measures how much each step of the forecasting process improves on a baseline: does the statistical model beat seasonal naive, and do planner overrides improve on the statistical model? FVA analysis often shows that some manual adjustments add error, which gives you an objective basis for scenario 50.
Anomaly detection and production forecasting
39. How do you detect anomalies in time series?
Answer: The most robust pattern is forecast-based detection: predict the expected value with an interval, and flag observations far outside it. This handles seasonality naturally (a busy Monday morning is not an anomaly). Simpler methods include rolling z-scores, median absolute deviation, and STL decomposition followed by thresholds on the remainder. For multivariate signals (many metrics from one service), isolation forests, autoencoders or reconstruction-error models detect unusual combinations. Distinguish point anomalies (a spike), contextual anomalies (normal value at the wrong time) and collective anomalies (a sustained level shift). Alert design matters as much as the algorithm: require persistence over several points, group related alerts, and tune thresholds to an acceptable false-alarm rate agreed with the on-call team.
40. How should anomalies, stockouts and structural breaks in history be treated before training?
Answer: Not all unusual history should be removed. Data errors (duplicate loads, a unit change) should be corrected. Genuine one-off events (a pandemic lockdown, a site outage) should be flagged with indicator features or down-weighted, so the model does not learn them as normal. Stockouts deserve special care: sales on days with zero stock are censored demand, not true demand, so training on them teaches the model to under-forecast exactly the items that sell out. Use inventory data to flag stockout days and either exclude them, impute demand from similar periods, or model them explicitly. For structural breaks (a new pricing policy, a store refit), shorten the training window or add a regime feature.
41. How do you decide on retraining frequency?
Answer: Separate refitting from re-forecasting. Most systems re-forecast every cycle with the latest data (daily or weekly), which statistical models do naturally by updating state. Full retraining of a global ML model, including hyperparameter search, can happen less often: monthly, or triggered by monitoring. Decide using a backtest of the policy itself: simulate "retrain weekly" versus "retrain monthly" and compare accuracy and cost. Also retrain before known regime changes such as the festive season, and keep a champion-challenger setup so a new model only replaces the current one after it wins on recent backtests. The MLOps interview questions cover model registries and promotion pipelines in more depth.
42. What does drift look like in forecasting, and how do you monitor it?
Answer: Drift shows up as a change in the input data (new stores, new SKUs, a changed price distribution, a feed now reporting in a different unit), a change in the relationship between inputs and demand (promotions stop working the way they used to), or simply rising error. Monitor at several layers: data quality (missing series, late files, volume anomalies in the input itself), feature distributions, accuracy and bias by segment once actuals arrive, and interval coverage. Because actuals arrive with delay, use early signals such as data checks and forecast-versus-plan gaps to catch issues before the error report does. Segment monitoring matters: a stable overall WAPE can hide a collapse in one region.
43. How do you architect forecasting for many series at scale?
Answer: Design for batch throughput, isolation of failures and reproducibility.
sales, inventory, price, calendar
|
feature store (as-of joins)
|
segment series: regular /
intermittent / new / dead
|
+------+--------+----------+
| | |
global GBM Croston/TSB cold-start
| | |
+------+--------+----------+
|
reconcile + quantiles
|
validate -> publish -> monitor
Key ideas: one global model rather than thousands of local ones where possible; segmentation so each class of series gets a fitting method; parallel execution (Spark, Ray or container jobs) for local models; per-series fallbacks to seasonal naive when a model fails or outputs nonsense; validation gates (no negatives, no extreme jumps versus last cycle) before publishing; and versioned inputs and models so any published forecast can be reproduced. Teams on Databricks often run this with pandas UDFs over grouped series; the Databricks interview questions go into that pattern.
44. What managed cloud services exist for forecasting in 2026?
Answer: Describe them generally and check current documentation, because this area changes. On AWS, Amazon Forecast is no longer available to new customers (existing customers can continue to use it), and AWS points users towards forecasting in Amazon SageMaker Canvas; teams that want full control build their own models on SageMaker AI. On Google Cloud, BigQuery ML offers ARIMA_PLUS models trained in SQL and an AI.FORECAST function that uses the built-in TimesFM model without training. On Azure, Azure Machine Learning automated ML supports forecasting tasks with time-series-aware validation. Managed services are attractive for quick wins and small teams; custom pipelines win when you need specific features (event calendars, reconciliation, custom loss) or tight cost control at scale. The SageMaker interview questions cover the AWS ML platform side.
If you want to practise forecasting alongside the cloud, MLOps and AI engineering skills these roles expect, Cloudsoft's APEX AI, ML, Cloud and Cyber Security program runs as classroom training in Ameerpet or live online; you can book a free demo on +91 96660 19191.
Real-world scenario questions
45. A retailer's forecasts badly under-predicted the Diwali spike last year. How do you fix it for this year?
Answer: The likely causes are that the model treated the festival as yearly seasonality (but the date moved), had too few festive seasons in history to learn from, and trained on stockout-censored sales from the previous peak. I would add an explicit festive calendar with days-to-event features, model the pre-festival ramp and post-festival dip, correct censored demand, and evaluate on backtests whose origins sit just before past festive seasons. Because each festival happens once a year, pooling across stores and categories in a global model matters more than per-series fitting.
What I would check:
- Did the model have the festival date as a feature, or only day-of-year?
- How many days were out of stock during last year's peak, and was that demand lost from the training target?
- Were promotions and sale events during the season captured as planned covariates?
- Does the error appear across all regions, or mostly where regional festivals differ?
- Did planners override the forecast, and did those overrides help or hurt?
Production consideration: Publish P50 and P90 for the festive window so replenishment can buy to the higher quantile for high-margin, non-perishable items, and set up daily re-forecasting with sell-through as the season starts, because the first few days of actual demand are the most informative signal you will get. The AI in retail and e-commerce guide covers how forecasting fits with pricing and personalisation in retail.
46. A new product launches next month with no sales history. How do you forecast it?
Answer: This is the cold-start problem. Use information other than the product's own history: analogue products (similar category, price band, brand, launch season) and their launch curves; a global model that uses product attributes as features so it can predict for an unseen item; a foundation model once a few weeks of data exist; and judgemental inputs from category managers, captured in a structured way (planned distribution, marketing spend). Forecast the launch curve shape (ramp, peak, decay) rather than a flat level.
What I would check:
- Is this a true new product, or a replacement or variant of an existing one whose history can be mapped?
- How many stores will stock it, and when? Distribution drives early sales more than the product does.
- What did comparable launches look like in their first eight weeks?
- Is there pre-order, wishlist or search data that signals interest?
Production consideration: Use wide intervals at first, re-forecast frequently, and shift weight from the analogue model to the product's own data as history accumulates. Track cold-start accuracy as a separate segment so it does not get hidden in the portfolio average.
47. You need daily forecasts for 10,000 SKUs across 200 stores. How do you design it?
Answer: That is millions of SKU-store series, so per-series model tuning is out. I would segment the series (regular, intermittent, new, discontinued), train one or a few global gradient-boosting models for regular series with scaled targets and Tweedie loss, use Croston/TSB or aggregation for intermittent ones, apply cold-start logic for new items, and reconcile to store, category and region totals so every team sees consistent numbers.
What I would check:
- Which level actually drives decisions: SKU-store-day for replenishment, or SKU-DC-week for purchasing?
- What share of series is intermittent, and what metric will the business accept for them?
- How long does a full run take, and does it finish before the replenishment cut-off?
- Are inventory and stockout flags available to correct censored demand?
- What is the fallback when a series fails validation?
Production consideration: Run in parallel on a distributed engine, write forecasts to a versioned table, and publish only after validation gates pass. Report accuracy weighted by revenue or volume, not a simple average over series, otherwise thousands of slow movers dominate the headline number. The data engineering interview questions cover the pipeline and orchestration side.
48. Forecast GPU capacity for an internal AI platform for the next two quarters.
Answer: GPU demand is driven less by history than by plans: new model launches, training runs, inference traffic growth and team onboarding. I would split demand into inference (forecastable from request volume, tokens per request and throughput per GPU, with daily and weekly seasonality), scheduled training (a pipeline of known projects with estimated GPU-hours) and ad hoc experimentation (statistical, from history). Forecast each, combine as scenarios (base, high, low), and translate them into instances by GPU type, since one accelerator type cannot always substitute for another.
What I would check:
- What is the procurement or reservation lead time, and the commitment term? That fixes the horizon that matters.
- What utilisation do current GPUs actually achieve, and how much demand is queued or rejected (censored)?
- Which product launches are confirmed, and what traffic do they assume?
- How will model or serving optimisations (quantisation, batching) change GPUs needed per request?
Production consideration: Present capacity as a probability of shortfall at each commitment level, and combine reserved capacity for the base forecast with on-demand or spot capacity for the upper tail. Revisit monthly as plans firm up. See GPU basics for AI engineers for how throughput per GPU is estimated and cloud cost optimisation for AI for the commitment trade-offs.
49. Forecast accuracy dropped sharply this month. How do you investigate?
Answer: Treat it like an incident: confirm the drop is real, localise it, then find the cause. A data problem is far more common than a model problem.
What I would check:
- Is the actuals feed complete? A late or partial load makes forecasts look wrong.
- Which segments, regions, channels or horizons account for most of the extra error?
- Did anything change upstream: a new store format, a product hierarchy remap, a unit-of-measure change, a promotion calendar not loaded?
- Was a new model version or feature deployed?
- Is there a real-world shift (competitor opening, regulation, weather extreme)?
Production consideration: Keep the previous model version deployable for fast rollback, and keep per-segment accuracy dashboards so the next drop is localised in minutes. Document the cause and add a data check that would have caught it.
50. Planners keep overriding the forecasts. What do you do?
Answer: Overrides are information, not just friction. Planners often know things the model does not (a supplier issue, a local event), but they also add bias. I would measure forecast value added for overrides by type and planner, feed the useful knowledge back into the model as features (event flags, supply constraints), and make it easy to record the reason for each override.
What I would check:
- Do overrides improve or worsen accuracy in aggregate, and for which item types?
- Are overrides mostly upward (fear of stockouts), indicating that planners need the high quantile, not a different median?
- Do planners understand what the forecast includes, such as planned promotions?
Production consideration: Provide explanations (drivers such as "festival week, promotion, price cut") and intervals with each forecast. Trust grows when people can see why a number moved.
51. A hospital wants to forecast emergency department arrivals for staffing. What is your approach?
Answer: Consider a multi-speciality hospital planning nurse rosters. Arrivals are counts with strong hour-of-day and day-of-week patterns, plus seasonal illness waves, festival and holiday effects, and local events. I would forecast hourly arrivals with a model that handles multiple seasonalities (Fourier terms or MSTL with a count-appropriate model, or a global model across departments), produce quantiles, and feed the upper quantiles into a staffing model, because understaffing during a surge is far costlier than a slightly idle shift.
What I would check:
- Is the target arrivals, or patients by acuity level? Staffing needs differ by acuity.
- Are there disease-season signals (public health surveillance, school calendars) available in time?
- How far ahead are rosters fixed?
Production consideration: Patient data falls under privacy rules, so forecast on aggregated counts, keep identifiers out of the forecasting pipeline and review the data flow with the compliance team.
52. A bank asks you to forecast cash withdrawals for its ATM network. What matters?
Answer: Consider a bank with ATMs across cities and towns that pays for every cash-replenishment trip. The cost is asymmetric: running out of cash hurts customers and reputation; overloading ties up idle cash. Demand has month-start salary peaks, weekly patterns, festival peaks and location effects (near a market, a hospital, an industrial area). I would build a global model across ATMs with location attributes and calendar features, forecast quantiles for the period until the next possible refill, and let an optimisation layer choose refill amounts and routes.
What I would check:
- Are days when the ATM was out of service or out of cash flagged? Withdrawals on those days are censored.
- What is the replenishment schedule and cassette capacity? That defines the horizon and the useful quantile.
- Do nearby ATMs substitute for each other when one is down?
Production consideration: Evaluate the forecast through the business metric (cash-out events and idle cash under the resulting refill plan), not only WAPE.
53. Forecast intraday contact-centre call volumes for a GCC support team.
Answer: Consider a Hyderabad GCC running support for global customers across time zones. Volumes arrive in 15- or 30-minute intervals with daily and weekly seasonality, customer-time-zone effects, release-day spikes and outage-driven surges. I would forecast daily totals with a model that uses business calendars of the served regions (their holidays, not only Indian ones), then distribute across intervals with historical intraday profiles by weekday, and feed interval forecasts into a workforce model (Erlang-style queueing) to convert volumes into required agents.
What I would check:
- Are abandoned calls included? They are demand that was not served.
- Do product releases or incidents explain past spikes, and can release calendars be a known-future regressor?
- How are channels shifting (chat versus voice), which changes handle time?
Production consideration: Add an intraday re-forecast that updates the rest of the day from morning actuals, since the first hours carry strong signal about the day's level.
54. Anomaly alerts on payment transaction volume are too noisy. The on-call team ignores them. How do you fix it?
Answer: Noisy alerts usually come from static thresholds that ignore seasonality, too-sensitive cut-offs, and alerting on single points. I would move to forecast-based detection with seasonal expectations (hour of day, day of week, salary dates, festival days), alert only on persistent deviations, and route by severity.
What I would check:
- What fraction of last month's alerts led to action? Label them and use that as the evaluation set.
- Are known events (scheduled maintenance, sale days) suppressed or expected?
- Is the alert on the right signal, for example failure rate rather than raw volume?
Production consideration: Agree an acceptable alert budget with the on-call team, tune thresholds to it on labelled history, and review weekly. For connecting anomaly signals to operational workflows, see AI observability and the AIOps interview guide.
55. You have two weeks to prove forecasting value to a sceptical supply chain head. What do you deliver?
Answer: A narrow, honest pilot. Pick one category and region with clean data, establish the current process (often a spreadsheet moving average plus planner judgement) as the baseline, and run a rolling-origin backtest of seasonal naive, ETS and a global gradient-boosting model with calendar and promotion features. Then translate the accuracy difference into business terms with a simple inventory simulation: stockout days and excess stock under each forecast at the same service level, using clearly labelled assumptions.
What I would check:
- What metric does the supply chain head already trust, and what decision does it drive?
- Is the backtest replaying information exactly as it was available at each forecast date?
- What would it take to run this weekly in production (data feeds, owner, review)?
Production consideration: Present the limitations (segments where the model does not beat the baseline) alongside the wins. Credibility on the first pilot decides whether there is a second one, and taking a model from a backtest into the business's real decision loop is exactly the gap between a demo and an enterprise outcome.
Key takeaways
- Start every forecasting answer with the decision, horizon, granularity and lead time, then name a baseline.
- Seasonal naive, ETS and ARIMA remain essential; global gradient boosting is the workhorse for large portfolios with covariates.
- Leakage is the most common reason offline results fail in production; build features as of the forecast timestamp.
- Use rolling-origin backtests, report error by horizon, and prefer MASE and WAPE over raw MAPE.
- Publish quantiles, check coverage, and remember quantiles do not add across a hierarchy.
- Hierarchies need reconciliation, intermittent items need their own methods, and stockout days hide true demand.
- Foundation models such as Chronos and TimesFM are useful for zero-shot and cold-start cases, but benchmark them like any other model.
Interview preparation checklist
- Decompose a real public series with STL and explain each component out loud.
- Fit seasonal naive, ETS and SARIMA in statsmodels and compare them with a rolling-origin backtest.
- Build a lag-feature gradient-boosting model for many series and deliberately find and fix one leakage bug.
- Compute MAE, RMSE, MAPE, WAPE and MASE on the same forecasts and explain why they disagree.
- Produce P10/P50/P90 with quantile loss or conformal intervals and measure coverage.
- Sketch a hierarchy and explain bottom-up, top-down and MinT reconciliation.
- Prepare a festive-season story: calendar features, censored demand, intervals for replenishment.
- Read the current docs for one managed forecasting service and one foundation model.
- Brush up on SQL window functions for feature building with the SQL interview questions for data and AI.
- Rehearse a one-minute explanation of a forecast for a non-technical stakeholder.
FAQ
What skills are needed for a time series forecasting role?
You need Python with pandas, a grounding in statistics, classical models such as ETS and ARIMA, feature engineering for gradient boosting, time-based validation, forecast metrics, SQL, and the ability to explain forecasts and uncertainty to business teams.
Is ARIMA still asked in forecasting interviews in 2026?
Yes. Interviewers use ARIMA and exponential smoothing to test whether you understand stationarity, autocorrelation and baselines, even if the production system uses gradient boosting or neural models.
Do I need deep learning to get a demand forecasting job?
Not usually. Strong fundamentals, gradient boosting with good features and honest backtesting matter more for most roles. Knowing how LSTMs, transformers and foundation models work conceptually is enough unless the role is research-focused.
How should I prepare for time series interview questions as a fresher?
Learn the components of a series, stationarity, ACF and PACF, naive baselines and error metrics, then complete one end-to-end project with a public dataset that includes a rolling backtest and a written comparison of models.
What is the most common mistake candidates make in forecasting interviews?
Validating with a random train-test split or building features that leak future information. Interviewers also notice when candidates skip the baseline or report only one accuracy number without intervals.
Which Python libraries should I know for forecasting?
pandas and NumPy for data handling, statsmodels for ETS, ARIMA and statistical tests, scikit-learn and a gradient-boosting library such as LightGBM, and familiarity with at least one forecasting framework or foundation model library.
Are time-series foundation models replacing traditional forecasting?
Not replacing, but adding an option. They are useful for quick zero-shot forecasts and sparse or new series, while trained models with business covariates often remain stronger for core demand planning. Benchmark both on your own data.
Is time series forecasting a good career path?
It can be. Forecasting supports retail, supply chain, banking, energy, healthcare and cloud capacity planning, and the skills combine well with data engineering, MLOps and AI engineering roles in GCCs and product companies.
How is forecasting different in Indian business contexts?
Moving festivals, regional holidays, monsoon effects, month-start salary cycles and large sale events create strong calendar-driven demand changes, so explicit event calendars and region-level features are especially important.
Forecasting interviews reward people who can connect statistics, data pipelines and business decisions. To build those skills with hands-on labs in ML, cloud and MLOps, explore Cloudsoft's APEX program for AI, ML, cloud and security. If your goal is to take models like these into customer environments end to end, the AI Forward Deployed Engineer FDE PRO program covers integration, deployment and observability, with placement support until you're placed. Both run in Ameerpet or live online; call +91 96660 19191 for a free demo.



