New batches starting this week Β· Limited seats

Search Relevance Interview Questions and Answers 2026 (50 Questions)

50 search relevance and search engineering interview questions with accurate answers, from inverted indexes and BM25 to hybrid search, learning to rank, evaluation and production scenarios.

Search relevance interview questions 2026: 50 questions on inverted indexes, BM25, analysers, hybrid kNN search, learning to rank and metrics
Last updated Β· 41 min read Β· 9,001 words

Search relevance interview questions in 2026 test whether you can explain why a search engine returned these results in this order, and how you would make that order better without breaking anything else. This guide covers 50 high-value questions with model answers on inverted indexes, analysers, BM25, query design, relevance tuning, learning to rank, semantic and hybrid search in Elasticsearch and OpenSearch, autocomplete, facets, Indian-language search, evaluation and cluster operations, finishing with ten production scenarios.

How to use this guide

Search engineer interviews mix information retrieval theory with very practical cluster work. What interviewers commonly probe at each level:

  • Freshers and junior engineers: what an inverted index is, how text is analysed, the difference between text and keyword fields, and match versus term queries.
  • Mid-level engineers: BM25 parameters, bool query structure, synonyms, autocomplete, aggregations, and kNN or hybrid search with RRF.
  • Senior and architect roles: offline and online evaluation, learning to rank, sharding and caching, mapping governance, access control and structured diagnosis of relevance regressions.

This guide focuses on search engineering and relevance. For ANN index internals such as HNSW, IVF and quantisation, use the vector database interview questions; for retrieval inside RAG pipelines, read hybrid search and reranking for RAG. Elasticsearch and OpenSearch share most core concepts, so the answers name both and call out differences where they matter. Features evolve quickly in both projects, so check current documentation before quoting exact parameter names in an interview.

Fundamentals: indexes, analysis and BM25

1. What is an inverted index, and why does full-text search depend on it?

Answer: An inverted index maps each term to the list of documents that contain it (a postings list), usually with term frequencies and positions. Instead of scanning every document for the word "refund", the engine looks up "refund" once and gets the matching document IDs directly. Queries with several terms intersect or union postings lists, which is cheap because the lists are sorted and compressed. Positions make phrase and proximity queries possible, and per-field statistics (document count, average field length) feed the scoring function. In Lucene, which underlies both Elasticsearch and OpenSearch, the index is written as immutable segments that are periodically merged.

Interview tip: Contrast it with a B-tree index in a relational database: a B-tree finds rows by an exact or range value of a whole column, while an inverted index finds documents by the individual terms inside a field.

2. Walk through an analyser. What do character filters, tokenisers and token filters each do?

Answer: An analyser turns raw text into the terms stored in the inverted index, in three stages. Character filters work on the raw string: stripping HTML, mapping characters (for example turning "&" into "and"). The tokeniser splits the string into tokens: the standard tokeniser follows Unicode word-boundary rules, while whitespace, pattern, n-gram and ICU tokenisers behave differently. Token filters then transform the token stream: lowercasing, ASCII folding, stop words, stemming, synonyms, word delimiters. The critical rule is that the query is analysed too, and the query-side analysis must produce terms compatible with what was indexed. Most "why doesn't this match?" bugs are analysis mismatches, and the _analyze API is the first tool to reach for.

Real-world example: A product code like "WD-40X" is split by the standard tokeniser into "wd" and "40x". If users search "wd40x", nothing matches unless you add a word-delimiter filter with catenation or a separate keyword sub-field for codes.

3. What is the difference between stemming and lemmatisation, and what goes wrong with aggressive stemming?

Answer: Stemming chops words to a crude root using rules ("running" to "run", "policies" to "polici"); lemmatisation maps words to a dictionary form using vocabulary and morphology. Search engines usually stem because it is fast and needs no linguistic model. Aggressive algorithmic stemmers raise recall but conflate words with different meanings, for example reducing "university" and "universe" to the same stem, or "organ" and "organisation". Lighter stemmers (such as the "light" or "minimal" variants offered for many languages) conflate less. Two practical mitigations are keeping an unstemmed sub-field that gets a higher boost, so exact forms rank above stemmed matches, and using a keyword-marker or stemmer-override filter for known problem words.

4. Explain text versus keyword fields and why multi-fields are common.

Answer: A text field is analysed into terms for full-text search and is scored. A keyword field stores the whole value as one term, used for exact filters, sorting and aggregations; it is not tokenised (though it can have a normaliser for lowercasing). Multi-fields index the same source value several ways: title as text with an English analyser, title.exact with a minimal analyser, title.raw as keyword for sorting. Each sub-field can be queried and boosted independently. The cost is index size and indexing time, so add sub-fields for real query needs, not by habit.

5. Explain BM25 intuitively.

Answer: BM25 scores a document for a query by summing, over each query term, three ideas. Rarity (inverse document frequency): matching a rare term like "chargeback" says more than matching "account". Term frequency with saturation: more occurrences help, but the gain flattens quickly, so the tenth mention adds far less than the second. Length normalisation: a match in a short field is more meaningful than the same match in a very long one, because long fields match many terms by chance. BM25 is the default similarity in Lucene-based engines and remains a strong baseline, especially for exact terms, names and codes.

Interview tip: Say "IDF is computed per shard by default", which leads naturally to why scores can differ slightly between shards on small indexes (and dfs_query_then_fetch as the fix when it matters).

6. What do the BM25 parameters k1 and b control, and when would you change them?

Answer: k1 controls term-frequency saturation: a low value makes the score saturate almost immediately (presence matters, repetition barely does), a high value lets repeated terms keep adding score. b controls length normalisation from 0 (ignore length) to 1 (full normalisation). Lucene's defaults are k1 = 1.2 and b = 0.75. You might lower b for a field where length does not signal dilution, such as short product titles of similar length or legal clauses where long ones are not less relevant, and lower k1 where keyword stuffing is a risk, such as seller-written product descriptions. Change them per field through a custom similarity, and only with an evaluation set that shows the effect; most relevance gains come from analysis, field design and query structure, not from k1 and b.

7. How is near-real-time search achieved, and what is the refresh interval?

Answer: New documents are first written to an in-memory buffer and a transaction log. A refresh writes the buffer into a new searchable segment, which is when documents become visible to search; Elasticsearch and OpenSearch refresh every second by default (Elasticsearch skips refreshes on shards that have been search-idle). A flush makes segments durable on disk and trims the translog. Background merges combine small segments and physically remove deleted documents. Frequent refreshes create many small segments and merge pressure, so bulk loads often set a longer refresh interval (or disable it) and restore it afterwards.

8. What are doc values, and why does it matter for sorting and aggregations?

Answer: The inverted index answers "which documents contain this term?" Sorting and aggregations need the opposite: "what is the value of this field for this document?" Doc values are a columnar, on-disk structure built at index time for that access pattern, enabled by default for keyword, numeric, date and similar fields. Text fields do not have doc values; aggregating on them requires fielddata, which is built in heap memory and can exhaust it, which is why it is disabled by default. The usual fix is to aggregate on a keyword sub-field instead.

9. Briefly, what is the licensing history of Elasticsearch and OpenSearch?

Answer: Elasticsearch was Apache 2.0 licensed until early 2021, when Elastic moved new versions (from 7.11) to a choice of the Server Side Public License (SSPL) and the Elastic License 2.0, neither of which is an OSI-approved open source licence. In response, AWS forked the last Apache-licensed versions (Elasticsearch and Kibana 7.10.2) as OpenSearch and OpenSearch Dashboards, which remain Apache 2.0. In August 2024 Elastic announced it was adding AGPLv3, an OSI-approved licence, as a third option alongside SSPL and ELv2 for the free parts of the code, rolled out around the 8.16 release. In September 2024 the Linux Foundation launched the OpenSearch Software Foundation, and AWS moved the OpenSearch project to it, with governance by a technical steering committee. The two projects have diverged since the fork, so APIs, plugins and features are not interchangeable beyond the shared basics.

Interview tip: Keep this factual and neutral. The practical point for an enterprise is that licence terms, managed-service options and feature availability by subscription tier all belong in the platform decision, and legal review owns the final call.

Query types and field boosting

10. What is the difference between a term query and a match query?

Answer: A term query looks up the exact term you give it, without analysis. A match query analyses the input with the field's search analyser first, then builds a query (usually a boolean OR of the resulting terms, configurable with operator or minimum_should_match). Use term queries on keyword, numeric and date fields for exact filters; use match on text fields. The classic bug is a term query for "Laptop" against a text field: the index holds "laptop" (lowercased), so nothing matches.

11. How do phrase queries and slop work?

Answer: match_phrase requires the terms to appear in order and adjacent, using the positions stored in the index. slop allows a number of position moves, so "credit card limit" with slop 1 still matches "credit card spending limit": each extra word in between costs one move, and swapped words cost more. A common relevance pattern is to keep a broad match clause for recall and add a match_phrase clause with a slop in a should clause as a boost, so documents containing the phrase rank higher without excluding the rest. match_phrase_prefix exists for simple prefix matching but is not a good autocomplete engine at scale.

12. Explain the bool query clauses and the difference between query and filter context.

Answer: A bool query combines clauses: must (must match, contributes to score), should (optional, adds score; with no must or filter clause at least one should is required unless you set minimum_should_match), filter (must match, no scoring) and must_not (excluded, no scoring). Filter context answers yes or no, so it skips scoring and its results can be cached, which makes it faster. Put structured constraints such as category, in-stock status, tenant ID, date ranges and permissions in filter, and keep must and should for the parts that should influence ranking.

Interview tip: Interviewers often check whether you know that putting a price range in must instead of filter silently changes the scores.

13. When is fuzzy matching useful, and what are its risks?

Answer: Fuzzy queries match terms within a Levenshtein edit distance (at most two edits), which catches typos such as "samsnug". fuzziness: AUTO scales the allowed edits with term length (no edits for very short terms), which is the sensible default. Risks: short or common terms fuzz into unrelated words ("cat" to "car"), fuzzy matching ignores meaning, it costs more CPU than exact lookups, and fuzzy matches of brand names can surface competitors. Mitigations are a prefix_length so the first characters must match, applying fuzziness only as a fallback when the exact query returns few results, and scoring fuzzy matches below exact ones.

14. Explain multi_match types and how you would boost fields.

Answer: multi_match runs one query across several fields. best_fields (the default) takes the score of the best-matching field, with tie_breaker adding a fraction of the others; it suits "the whole query should match one field". most_fields sums scores across fields, useful when the same text is indexed with different analysers (stemmed, exact, shingles). cross_fields treats the fields as one big field, term by term, which suits data split across fields like first name and last name or brand and model. phrase and phrase_prefix run phrase matching per field. Field boosts use caret syntax, for example title^3, brand^2, description. Boosts are multipliers on scores whose scales differ by field, so tune them against judgments, not intuition, and avoid extreme values that make one field dominate.

15. How do you combine text relevance with business signals such as popularity, recency or margin?

Answer: Use function_score or script_score to modify the text score: a field_value_factor on log-scaled sales or ratings, decay functions (gauss, exp, linear) on date or geo distance for recency and proximity, and weights for in-stock or promoted items. In Elasticsearch, rank_feature fields and queries are an efficient way to add static signals such as popularity. Keep text relevance as the primary signal and the business boosts bounded, typically with log or saturation functions, or popular-but-irrelevant items will float to the top for every query. Document every boost, its owner and its reason, because untracked business boosts are a common cause of mysterious ranking.

Real-world example: Consider a retailer whose merchandising team adds a margin boost. Searches for "phone charger" start showing high-margin power banks first. A capped, log-scaled boost applied only within the top text-relevant candidates fixes it.

Relevance tuning and learning to rank

16. How do you approach relevance tuning systematically rather than by trial and error?

Answer: Treat it as an engineering loop with a fixed test set. First collect representative queries from logs: head, torso and long-tail, plus known problem queries and zero-result queries. Get graded judgments for the top results of each. Measure a baseline (NDCG@10, zero-result rate, and per-segment results). Then change one thing at a time: analysis, field structure, query template, boosts. Use the _explain API and the profile API to see why a document scored as it did. Re-run the whole set after every change, because a fix for one query frequently breaks others. Only changes that improve the offline metric without serious regressions go to an online A/B test.

query logs -> sample set -> judgments
      |                          |
      v                          v
 change one thing -> offline NDCG + diff
      |                          |
      +-- regress? <- no --- A/B test -> ship

Interview tip: Mention a "query diff" report: for each test query, which documents entered or left the top ten. Reviewers can read it, and it catches regressions a single average hides.

17. Index-time versus search-time synonyms: which do you prefer, and why?

Answer: Search-time synonyms are usually preferred. Changing index-time synonyms requires reindexing, they inflate the index, and they distort term statistics. Search-time synonyms can be updated without reindexing (Elasticsearch supports updateable synonym filters and a synonyms API for managed sets; OpenSearch supports updateable search analysers too; check the version you run). Use the synonym_graph filter at search time so multi-word synonyms such as "tv" and "television set" produce correct phrase structure. Design choices: equivalence ("mobile, cellphone, smartphone") versus one-way expansion ("iphone => iphone, smartphone", but not the reverse), and limiting synonyms to the fields where they help. Every synonym list needs an owner, a review process and regression tests, because synonyms are one of the most common sources of relevance regressions (see Q42).

18. What is learning to rank, and how does it fit into a search engine?

Answer: Learning to rank (LTR) trains a model to order results using many features, instead of hand-tuned boosts. The engine first retrieves candidates with a cheap query (BM25, hybrid), then a rescoring stage applies the model to the top N, for example the top few hundred. Models are usually gradient-boosted trees (LambdaMART-style objectives via XGBoost or LightGBM) trained with pairwise or listwise losses on graded judgments. Elasticsearch has a native learning-to-rank rescorer in recent versions, and OpenSearch has an LTR plugin; both log features and score with a deployed model. The candidate stage limits what LTR can fix: if a relevant document is not retrieved, the reranker never sees it.

19. What features would you use in an LTR model, and what are the common pitfalls?

Answer: Features fall into groups: query-document (BM25 per field, phrase match, vector similarity, exact SKU match), document (popularity, rating, freshness, stock, price), query (length, predicted category, is it a brand query) and context (device, region, where allowed). Pitfalls: feature leakage, such as features computed with data from after the judgment time; training on raw clicks without correcting position bias, so the model learns to reproduce the old ranking; features computed differently offline and online; and too few judged queries for the tail. Log features at query time from the production engine to keep training and serving consistent, and keep a hand-tuned baseline to fall back to.

20. What is query understanding, and why does it often beat ranking changes?

Answer: Query understanding interprets the query before retrieval: spelling correction, category or intent classification ("shoes under 2000" means a price filter), entity recognition (brand, size, colour), and query rewriting or relaxation. Many bad results come from treating a structured query as a bag of words: "red nike running shoes size 9" works far better as a filter on brand, colour and size plus a text match on "running shoes". Classifiers can be simple rules and dictionaries at first, then models trained on click data or an LLM used offline to label queries. Keep rewrites explainable and logged, so you can tell when a rewrite caused a bad result.

Interview tip: Tie this to NLP interview questions topics like named entity recognition and intent classification; search teams reuse those skills directly.

Semantic and hybrid search

21. How is kNN vector search supported in Elasticsearch and OpenSearch?

Answer: Elasticsearch uses the dense_vector field type with HNSW-based approximate search (including quantised variants), queried through a top-level knn section, a knn query or a kNN retriever, with k and num_candidates controlling how many neighbours are explored per shard. It also offers semantic_text, which calls an inference endpoint to generate embeddings automatically at index and query time. OpenSearch uses the k-NN plugin with the knn_vector field type, backed by the Faiss or Lucene engines (the older nmslib engine was deprecated and is not available for new indexes in 3.x), and the neural search plugin to run embedding models through ingest and search pipelines. Both support filters during kNN search. The details of HNSW tuning belong to the vector database interview guide; in a search interview, focus on how vectors combine with text, filters and ranking.

22. What is learned sparse retrieval, and how does it differ from dense vectors?

Answer: Learned sparse models output a weighted bag of terms per document and query, including expanded terms the text did not contain, so "laptop bag" might also get weight on "sleeve" and "backpack". The result is stored in an inverted index and matched like keywords, which makes it explainable and efficient, and good for domain vocabulary. Elastic's ELSER and OpenSearch's neural sparse search are examples. Dense vectors capture overall meaning in a fixed-length embedding and are stronger for paraphrase, but are less interpretable and weaker at exact identifiers. Many teams combine BM25, sparse and dense signals; the right mix is decided by evaluation, not by preference. For the embedding basics, see embeddings explained.

23. How would you implement hybrid search, and when would you use RRF versus score normalisation?

Answer: Run a lexical query and a vector query, then fuse. Reciprocal rank fusion (RRF) uses only ranks: each document scores the sum of 1 / (k + rank) across lists, with k commonly 60. It needs no score calibration, so it is a robust default. Elasticsearch exposes it as an rrf retriever (with rank_constant and rank_window_size), and OpenSearch added RRF to hybrid search in 2.19 through a score-ranker processor in a search pipeline. Normalised score combination (min-max or L2 normalisation, then weighted sum; OpenSearch's normalisation processor, Elasticsearch's linear-style combinations in recent versions) can do better when tuned, because it keeps score gaps, but weights drift when models or data change. Start with RRF, tune the per-list window size, and move to weighted fusion only when judgments show a gain. The fusion mechanics are covered in depth in our hybrid search and reranking guide.

24. Where do rerankers fit in a search stack, and what do they cost?

Answer: A cross-encoder reranker reads the query and each candidate together and outputs a relevance score, which is more accurate than comparing separate embeddings, but costs one model call per candidate. So it runs only on the top N candidates after retrieval and fusion, typically tens to a couple of hundred, as a rescoring stage. Costs are latency (often the largest single item in the request), GPU or API spend, and an extra dependency that needs timeouts and a fallback to the un-reranked order. In product search, a reranker still needs business signals afterwards (stock, price filters already applied, diversity), so it is one stage, not the final ranking. Measure the gain in NDCG against the latency added at p95, not just the average.

25. How do filters interact with kNN search, and why might filtered vector queries return too few results?

Answer: If the filter is applied after the approximate search (post-filtering), the engine finds k nearest neighbours and then drops those that fail the filter, so a selective filter can leave only a few results. Both engines support filtering during the kNN search, where the filter restricts candidates as the graph is explored, and they can fall back to exact search when the filter is very selective. In Elasticsearch, put the filter inside the kNN section rather than in an outer post-filter, and raise num_candidates if recall is low. In OpenSearch, the Faiss and Lucene engines support efficient filtering; behaviour depends on engine and version. Permission and tenant filters must always be pre-filters, never post-processing in the application.

26. When does semantic search make results worse?

Answer: Semantic search hurts on queries where the user means something exact: SKUs, part numbers, model numbers ("iPhone 15" versus "iPhone 16"), people's names, error codes, legal section numbers and short navigational queries. Embeddings place "iPhone 15 case" close to "iPhone 16 case", which is the wrong answer for a shopper. It also struggles with negation ("without sugar"), numeric constraints and very domain-specific jargon the embedding model never saw. The fix is not to drop semantic search but to route: detect identifier-like queries and give lexical matching priority, keep exact-match boosts in hybrid queries, and evaluate per query class.

Autocomplete, spelling and facets

27. How would you build autocomplete, and what are the options?

Answer: Common options are: an edge n-gram analysed field (index "lap", "lapt", "lapto", "laptop" so prefixes match normal queries, with a standard analyser at search time); the search_as_you_type field type, which builds prefix and shingle sub-fields for you; and the completion suggester, an in-memory finite-state structure that is very fast for prefix lookups with weights but less flexible (prefix only, limited filtering through contexts). Many production systems autocomplete over a separate index of popular past queries and category names rather than over product documents, ranked by frequency and success rate, with filtering of offensive or zero-result suggestions. Latency budgets are tight, so keep the suggestion index small and the query simple.

28. How do you implement spelling correction ("did you mean")?

Answer: Options range from engine features to dedicated services. The term and phrase suggesters propose corrections from terms in the index; the phrase suggester uses n-gram language-model style scoring and handles multi-word queries better. Fuzzy queries handle typos implicitly during retrieval. Stronger systems learn corrections from query logs: a user types "samsnug", gets nothing, then types "samsung" and clicks, so that pair becomes a correction candidate. Decide when to auto-correct (high confidence and the original query has zero or very few results) versus show a suggestion, and never auto-correct valid rare terms such as brand names, which needs a protected-word list.

29. How do facets work with aggregations, and what is post_filter for?

Answer: Facets are aggregations over the matched documents: terms aggregations for brand or category counts, range or histogram for price, nested aggregations for variant data. The subtlety is multi-select facets. If a user selects brand "Asus", the brand facet should still show counts for other brands, while other facets should reflect the selection. post_filter applies a filter to the hits after aggregations are computed, so aggregations see the unfiltered set. For full multi-select behaviour, use filter aggregations that apply every selected facet except the facet's own. Aggregate on keyword or numeric fields with doc values, never on analysed text.

30. Why can terms aggregation counts be approximate, and how do you deal with it?

Answer: A terms aggregation asks each shard for its top terms (controlled by shard_size, which defaults to more than size), then the coordinating node merges them. A term that is moderately common on every shard but never in any single shard's top list can be undercounted or missed, which the response reports through error bounds. Increase shard_size, use fewer shards for small indexes, or route related documents to the same shard. cardinality (distinct counts) is approximate by design, using HyperLogLog++ with a tunable precision threshold. For facets these approximations are normally acceptable; for financial reports they are not, and that data belongs in a database or a composite aggregation that pages through all buckets.

Multilingual and Indian-language search

31. How would you design an index for multilingual content?

Answer: Three common patterns. Field per language (title_en, title_hi), each with its own analyser, and a query that targets the user's language plus a fallback. Index per language, simpler analysers and statistics per language, but more indexes to manage. One field with a language-neutral analyser (ICU tokeniser and folding), simplest but no stemming. Detect the document language at ingest (an ingest pipeline with a language-identification model, or upstream), and detect query language cautiously because short queries are ambiguous. Multilingual embedding models help a lot for cross-lingual queries and are a strong reason to add a semantic leg to multilingual search.

32. What makes Indian-language search hard, and how would you handle transliteration and script variants?

Answer: Users mix scripts and languages: the same item is searched as "ΰ€ͺΰ€¨ΰ₯€ΰ€°", "paneer" and "panir"; Telugu users may type romanised Telugu; queries mix English and Hindi ("kurta for ladies cotton"). Problems include romanisation with no standard spelling, Unicode variants of the same visible text (nukta forms, precomposed versus decomposed characters, zero-width joiners), and limited stemmers for some languages. Practical measures: apply Unicode normalisation (NFC or NFKC, for example with the ICU normaliser) at both index and query time; strip zero-width characters where they do not change meaning; index a transliterated Latin sub-field for native-script content (the ICU plugin's transform filter can convert scripts, though output quality varies by language and should be tested) and a native-script sub-field for romanised content where you have a mapping; add phonetic or n-gram matching for romanised spelling variants; and build synonym lists from query logs ("dal", "daal", "dhal"). Elasticsearch has built-in analysers for some Indian languages such as Hindi and Bengali; check current documentation for others. Multilingual embeddings add cross-script recall but must be evaluated on your own judged queries in each language.

Real-world example: Consider a grocery app in Hyderabad where "atta", "aata" and "ΰ€†ΰ€Ÿΰ€Ύ" should all find wheat flour. Normalised native script, a transliterated sub-field and a short log-mined synonym list together fix most of it before any model is involved.

Relevance evaluation

33. How do you build relevance judgments, and which metrics do you use?

Answer: Judgments are graded labels for query-document pairs, for example 0 (irrelevant) to 3 (perfect). Sources: trained human judges with written guidelines, domain experts for specialised content, click-derived labels, and LLM judges calibrated against a human-labelled subset. Sample queries across head, torso and tail, and pool the top results from several systems so new rankers are not penalised for retrieving unjudged documents. Metrics: NDCG@k rewards putting highly relevant results near the top and handles graded labels; MRR suits queries with one right answer (navigational search); precision@k and recall@k for coverage; zero-result rate and query reformulation rate as health signals. Elasticsearch and OpenSearch both have rank evaluation tooling, but a small script over your judgment file works too. For RAG-specific retrieval metrics, see RAG evaluation metrics.

34. Why can't you use raw clicks as relevance labels? What are click models?

Answer: Clicks are biased. Position bias: users click the top results more because they are on top. Presentation bias: images, prices and badges attract clicks regardless of relevance. Trust bias and the fact that unseen results get no clicks at all. Training on raw clicks makes the system reinforce its current ranking. Click models estimate relevance by modelling examination: the cascade model assumes users scan top-down and stop at a satisfying result; position-based models factor click probability into examination probability times attractiveness; dynamic Bayesian network models add satisfaction after a click. Practical alternatives include inverse propensity weighting with propensities estimated by small randomisation experiments, and using downstream signals (add-to-cart, purchase, dwell time) as stronger labels.

35. How do you run online experiments for search, and what is interleaving?

Answer: In an A/B test, users are split between control and treatment rankers and you compare business and engagement metrics: conversion, revenue per search, click-through, zero-result rate, reformulation and abandonment. Search A/B tests need careful randomisation by user, not by request, enough traffic for the effect size, guardrail metrics such as latency, and attention to novelty effects. Interleaving merges results from both rankers into one list for the same user (team-draft interleaving is a common method) and credits each ranker for clicks on its results. It is much more sensitive, so it needs less traffic to detect a preference, but it tells you which ranker users prefer, not the business impact, so teams often use interleaving to screen and A/B tests to confirm.

Performance, security and operations

36. How do you decide the number of shards and replicas?

Answer: A shard is a Lucene index; primaries split data, replicas copy it for availability and read throughput. The primary count is fixed at index creation (changeable only through split, shrink or reindex), while replicas can change any time. Too many small shards waste heap and coordination overhead and slow searches, because each query fans out to every shard; too few large shards make recovery and rebalancing slow. Size from data volume and growth, aiming for shard sizes in the range your engine's current documentation recommends (Elastic's guidance commonly cites tens of gigabytes), use time-based indexes with rollover for logs, and at least one replica in production. For search-heavy workloads, add replicas to scale reads; for indexing-heavy loads, replicas add write cost.

37. What caches matter in Elasticsearch and OpenSearch, and how do refresh settings affect them?

Answer: The node query cache caches results of filter clauses per segment, which is why repeated filters (in-stock, category) get fast. The shard request cache caches whole responses for size: 0 requests such as aggregation-only dashboards, invalidated when the shard refreshes with changes. The operating system file system cache holds index files and matters most of all, which is why you leave a large share of RAM outside the JVM heap. Each refresh creates new segments, so frequent refreshes on a busy index reduce cache effectiveness. Use filters for cacheable constraints, avoid "now" in date filters without rounding (for example now/h), set a longer refresh interval where seconds of freshness do not matter, and watch cache hit and eviction rates.

38. What is a mapping explosion, and how do you prevent it?

Answer: With dynamic mapping, every new JSON key becomes a new field in the mapping. If documents carry user-defined keys, such as product attributes, log labels or customer metadata, field count grows without bound, cluster state gets large, and the master node and heap suffer. Defaults limit total fields per index (1,000 in both engines unless changed) and indexing fails when exceeded. Prevention: set dynamic to strict or false for controlled indexes, model arbitrary attributes as key-value pairs in a nested or flattened structure (Elasticsearch's flattened and OpenSearch's flat_object field types store an object as keywords without per-key mappings), use index templates and review mapping changes like schema migrations. Raising the field limit is a last resort, not a fix.

39. How do you enforce security and access filtering in search?

Answer: Layers: authentication and TLS everywhere; role-based access to indexes; document-level security (roles carrying a query that limits visible documents) and field-level security (hiding fields such as PAN numbers or Aadhaar-linked data); and audit logging. Both engines provide these through their security features, with availability depending on distribution and subscription tier. For enterprise search over documents from many systems, store access control lists on each document (allowed groups and users) at index time, and add a filter built from the caller's identity to every query on the server side, never trusting a client-supplied filter. Keep ACLs in sync with the source system on change, because stale ACLs leak data. Aggregations, autocomplete and spelling suggestions can leak titles too, so apply the same filters to them. For identity patterns when an AI agent calls search, see AI agent identity and access.

40. How do you change a mapping or analyser in production without downtime?

Answer: Most analysis and field-type changes require a reindex. Use index aliases: applications always query an alias such as products. Create products_v2 with the new mapping, reindex from the source of truth (or with the reindex API from v1), dual-write or replay changes made during the copy, validate document counts and run the relevance test set against v2, then atomically switch the alias and keep v1 for rollback. The same pattern supports relevance experiments, since two index versions can serve an A/B test. Operational basics to mention alongside: snapshots to object storage, rolling upgrades one node at a time, monitoring cluster health, heap, GC, thread pool rejections and disk watermarks.

Want to build search and retrieval systems end to end, from ingestion and analysers to semantic ranking and cloud deployment? Explore the APEX AI, ML, Cloud and Cyber Security program, with classroom sessions in Ameerpet or live online.

Real-world scenario questions

41. Analytics show a large share of searches return zero results. How do you reduce zero-results queries?

Answer: Classify the zero-result queries before fixing anything, because the causes need different fixes: typos, synonyms or vocabulary gaps ("sneakers" versus "sports shoes"), over-strict query logic, identifier formats, other languages or scripts, and genuinely unavailable items.

What I would check:

  1. Pull the top zero-result queries by frequency and a random tail sample, and label each by cause.
  2. Run the worst ones through _analyze and _validate/query?explain to spot analysis mismatches such as hyphenated codes or unnormalised Unicode.
  3. Check the query template: an operator: and or a high minimum_should_match across many terms often causes zero results for long queries.
  4. Check filters applied silently, such as a default region, stock or category filter.
  5. Add a relaxation chain: exact, then fewer required terms, then fuzzy, then semantic, stopping at the first level with results.
  6. Add log-mined synonyms and spelling corrections for recurring vocabulary gaps.
  7. For items you genuinely do not stock, show helpful alternatives or a category page instead of an empty page.

Production consideration: Track zero-result rate per segment (language, device, category) and alert on jumps, which often reveal an ingestion failure or a bad deployment rather than a relevance problem. A relaxed result set is not automatically good; judge relaxed results too.

42. After a merchandiser added a batch of synonyms, ranking got noticeably worse for many queries. What happened, and how do you fix it?

Answer: Synonyms change which documents match and how terms are weighted. Common causes: equivalence rules that should have been one-way ("apple" equated with "iphone" now returns iPhones for "apple juice"); overly broad terms ("case" equated with "cover" and "sleeve" across all fields); multi-word synonyms applied with a non-graph filter, which breaks phrase positions; index-time synonyms that skewed IDF; and synonyms applied to fields such as brand where they never belonged.

What I would check:

  1. Diff the synonym file versions and list the new rules.
  2. Run the relevance test set with and without the new rules and produce a per-query diff to find which rules cause the regressions.
  3. Use _analyze on affected queries to see the expanded token graph.
  4. Check whether the filter is synonym_graph at search time and which fields use that analyser.
  5. Use _explain on a wrongly promoted document to see which expanded term matched.
  6. Roll back the batch through the updateable synonym set while fixing; then reapply rules one by one with one-way expansions and narrower field scope.

Production consideration: Treat synonyms as code: version control, review, a regression run against judgments in CI before release, and an owner per rule. Down-weight synonym matches relative to original terms (for example by querying an exact sub-field with a higher boost) so expansions add recall without outranking literal matches.

43. The category facets on the search page have become slow, and aggregation latency dominates the request. How do you investigate?

Answer: Aggregation cost scales with the number of matching documents, the number of buckets and the data structures used. Find out which aggregation is slow and why before scaling hardware.

What I would check:

  1. Use the profile API to time each aggregation, and compare with a size: 0 request.
  2. Check for high-cardinality terms aggregations (seller ID, raw SKU) with large size or shard_size, and nested sub-aggregations multiplying buckets.
  3. Check for aggregations on text fields with fielddata, scripts computing values at query time, or nested aggregations over large variant arrays.
  4. Look at shard count: many small shards add fan-out and merge overhead.
  5. Consider eager global ordinals for keyword fields that are aggregated on every request (shifts cost to refresh time), or a longer refresh interval.
  6. Precompute: store category paths as keyword fields at ingest instead of scripting them; use the request cache for repeated dashboard queries.
  7. Reduce work in the product: show the top categories first and lazy-load the long tail of facets.

Production consideration: Set a latency budget per facet, monitor search thread pool queues and circuit-breaker trips, and test aggregations against a realistic broad query ("shirt") rather than narrow test queries that match few documents.

44. An e-commerce company wants to add semantic search to its existing keyword product search. How would you do it?

Answer: Add semantic retrieval as an extra candidate source, fuse it with the existing lexical results, and prove it with evaluation before rollout, rather than replacing BM25.

query -> understanding (spell, intent, filters)
   |                      |
   v                      v
 BM25 + exact SKU     kNN on product vectors
   |                      |
   +------ RRF fusion ----+
              |
   rerank / LTR + business rules
              |
           results

What I would check:

  1. Baseline: build a judged query set covering head, tail, natural-language ("gift for a 10 year old who likes science") and identifier queries.
  2. Choose what to embed: title, brand, category, key attributes, perhaps a cleaned description; noisy seller descriptions can hurt.
  3. Pick an embedding model by evaluating a few candidates on your own queries, including Indian-language and Hinglish queries if your users write them.
  4. Index vectors in the same engine (a vector field in Elasticsearch or OpenSearch) so filters for stock, price and category apply during kNN search.
  5. Fuse with RRF first; keep exact-match boosts so SKUs and model numbers stay precise; route identifier-like queries to lexical priority.
  6. Measure NDCG per query class, latency at p95 including embedding inference, and memory growth.
  7. A/B test on conversion and revenue per search with latency and zero-result guardrails.

Production consideration: Embedding generation becomes part of the catalogue pipeline, so new products need vectors before they appear, and a model change means re-embedding the whole catalogue through the alias-swap pattern. For the wider business view, see AI in retail and e-commerce; personalised ranking on top of search connects to the recommendation systems interview questions.

45. Users paging deep into results cause heap pressure and slow queries. What do you change?

Answer: With from plus size pagination, each shard must collect from + size hits and the coordinating node sorts all of them, so page 500 is far more expensive than page 1; engines cap this with max_result_window (10,000 by default). For user-facing search, cap pagination, because almost nobody needs page 500, and offer better filters instead. For deep pagination needs such as "export all", use search_after with a point-in-time (PIT) view in Elasticsearch or OpenSearch, which pages by sort values efficiently and consistently. The scroll API is the older approach for bulk retrieval and holds resources open. Also check for bots and scrapers driving the deep pages.

What I would check:

  1. Slow logs and request patterns for high from values and their sources.
  2. Whether exports go through the search API instead of a data pipeline from the source of truth.
  3. Heap usage, GC pauses and circuit breakers during those requests.

Production consideration: Rate-limit per client and keep exports off the latency-critical cluster where possible.

46. Search p99 latency spikes every night when the catalogue is bulk-reindexed. How do you fix it?

Answer: Indexing and search compete for CPU, disk I/O and caches. Each refresh during the bulk load invalidates caches and creates segments that must be merged.

What I would check:

  1. Whether the job rewrites unchanged documents; incremental updates cut most of the load.
  2. Bulk request sizes and concurrency; too many parallel bulk requests cause rejections and merge storms.
  3. Refresh interval during the load; a longer interval or a separate new index built offline helps.
  4. Building into a fresh index with replicas set to zero, then adding replicas, force-merging if appropriate, warming and swapping the alias.
  5. Separating indexing and search onto different nodes if the engine and deployment support it.

Production consideration: Watch merge throttling, disk watermarks and thread pool rejections; schedule heavy jobs with the business, since "night" in one region is peak in another.

47. A marketplace sees poor search for users in smaller towns who type in Hindi, Telugu or romanised forms. How do you approach it?

Answer: Start from the queries, not from the analysers. Segment logs by script (Devanagari, Telugu, Latin) and measure zero-result and reformulation rates per segment, then fix the largest gaps.

What I would check:

  1. Unicode normalisation consistency between catalogue and queries; mismatched forms look identical on screen but do not match.
  2. Whether product titles exist only in English; if so, native-script queries need transliteration to Latin or cross-lingual embeddings.
  3. Romanised spelling variation; add phonetic or n-gram matching and log-mined variant synonyms.
  4. Code-mixed queries; tokenise both scripts correctly and avoid an English stemmer mangling romanised Indian words.
  5. Voice search input, which may produce native script for words the catalogue holds in English.
  6. A judged query set per language, built with native speakers, before and after changes.

Production consideration: Keep per-language metrics on the dashboard permanently; improvements for English often regress other languages unnoticed. Voice queries add their own issues, covered in the voice AI interview questions.

48. In an internal enterprise search at a bank, an employee found a confidential HR document title in autocomplete. How do you respond?

Answer: Treat it as a data exposure incident first and a search bug second: contain, assess, fix the root cause, then prevent recurrence.

What I would check:

  1. Contain: disable the suggestion source or the affected index for suggestions immediately.
  2. Find the path: the main search probably applies an ACL filter, but the suggestion index, spelling suggester or aggregations may not.
  3. Check ACL sync: whether the document was correctly restricted in the source system and whether the indexed ACL was stale.
  4. Assess exposure using query and access logs: who could have seen it, and when.
  5. Involve security, HR and compliance as the bank's incident process requires.
  6. Fix: apply the same identity-derived filter to every search surface, or build suggestions only from content visible to everyone.
  7. Add automated tests that query each surface as a low-privilege user and assert restricted content never appears.

Production consideration: Permissions are part of relevance infrastructure. The same rule applies when search backs a RAG assistant; see RAG interview questions for retrieval access control in that context.

49. An A/B test shows a new ranker improves offline NDCG and click-through, but conversion drops. What do you conclude?

Answer: Offline metrics and clicks measure what they measure, not business value. Possible explanations: judgments rewarded topical relevance but ignored price, stock or delivery time; the new ranker surfaces attractive but out-of-stock or expensive items that get clicks but not purchases; position-biased click labels; a segment effect (better on mobile, worse on desktop); or a latency increase that hurts conversion.

What I would check:

  1. Experiment validity: sample ratio mismatch, randomisation unit, test duration and novelty effects.
  2. Segment results by device, category, query class and new versus returning users.
  3. Latency per arm, including p95 and p99.
  4. Stock and price distribution of top results in each arm.
  5. Judgment guidelines: do they reflect purchase intent for commercial queries?

Production consideration: Do not ship. Define success metrics and guardrails before the next test, and add purchase-linked features or labels to the ranker.

50. A new team asks you to choose between Elasticsearch, OpenSearch and PostgreSQL full-text search plus pgvector for a product search service. How do you decide?

Answer: Decide from requirements, operating model and licensing, not brand preference.

ConsiderationSearch engine (Elasticsearch / OpenSearch)PostgreSQL FTS + pgvector
Relevance featuresRich analysers, BM25, synonyms, suggesters, LTR, hybrid fusionBasic full-text ranking (not BM25 by default); hybrid needs custom SQL
Facets and aggregationsBuilt for interactive facets at scalePossible with SQL, slower on broad queries
OperationsA separate cluster to run, or a managed serviceReuses an existing database and skills
ConsistencyNear-real-time, separate from the system of recordTransactional with the source data
LicensingElasticsearch: SSPL, ELv2 or AGPL options; OpenSearch: Apache 2.0PostgreSQL licence; pgvector open source

What I would check:

  1. Catalogue size, query volume and the latency target.
  2. Whether rich facets, autocomplete, multilingual analysis and relevance tuning are core product requirements.
  3. Team skills and whether a managed service (Elastic Cloud, Amazon OpenSearch Service) is acceptable.
  4. Licence review with legal, and which features require a paid tier.
  5. A proof of concept on a judged query set, comparing relevance and latency.

Production consideration: Small catalogues and internal tools often start well on PostgreSQL; consumer product search with heavy faceting and tuning usually justifies a search engine. Either way, keep a reindex pipeline from the source of truth so you can change engines later.

Key takeaways

  • Most "missing result" bugs are analysis mismatches; the _analyze and _explain APIs are the first diagnostic tools.
  • BM25 is about rarity, saturating term frequency and length normalisation; field design and query structure usually matter more than k1 and b.
  • Put structured constraints in filter context, keep business boosts bounded, and document every boost and synonym rule.
  • Semantic search is an extra candidate source; fuse it with lexical results (RRF is a robust default) and protect exact identifiers.
  • Relevance work needs a judged query set, NDCG and per-query diffs offline, and A/B or interleaving tests online with business guardrails.
  • Indian-language search needs Unicode normalisation, transliteration handling and per-language evaluation.
  • Operations shape relevance: shard sizing, refresh intervals, mapping governance, alias-based reindexing and access filters on every search surface.

Interview preparation checklist

  • Index a small product catalogue in Elasticsearch or OpenSearch with custom analysers and multi-fields, and inspect tokens with _analyze.
  • Write bool queries using must, should, filter and must_not, and explain the scores with _explain.
  • Build search-time synonyms with synonym_graph and measure their effect on a judged query set.
  • Implement autocomplete two ways (edge n-grams and a completion suggester) and compare.
  • Add a vector field, run kNN with filters and a hybrid query with RRF, and compare NDCG@10 against BM25 alone.
  • Write a short script that computes NDCG@10 and MRR from a judgment file.
  • Practise the alias-swap reindex and explain shard, replica and refresh choices.
  • Prepare one Indian-language example: normalisation, transliteration and how you would evaluate it.
  • Rehearse scenario answers aloud: zero results, synonym regressions, slow aggregations, semantic rollout, access leaks.
  • Review neighbouring guides: knowledge graph interview questions and data engineering interview questions for ingestion pipelines.

FAQ

What skills does a search relevance engineer need?

Information retrieval basics (inverted indexes, BM25, analysis), hands-on Elasticsearch or OpenSearch, query design, evaluation with judgments and metrics such as NDCG, some machine learning for learning to rank and embeddings, and enough operations knowledge to keep a cluster fast and secure.

Should I learn Elasticsearch or OpenSearch for interviews?

Either works for the core concepts, because both are built on Lucene and share analysis, query and aggregation basics. Learn one well, then read the differences in vector search, hybrid search, security and licensing so you can discuss both.

Are BM25 interview questions still asked now that semantic search exists?

Yes. BM25 remains the default lexical scorer and a key part of hybrid search, and interviewers use it to check whether you understand term rarity, term-frequency saturation and length normalisation rather than treating search as a black box.

How should a fresher prepare for search engineer interview questions?

Build a small search project end to end: index a public product or article dataset, add analysers, synonyms, autocomplete and facets, create a small judged query set, and measure NDCG before and after each change. Be ready to explain every decision.

Do I need machine learning knowledge for search relevance roles?

For many roles, basic ML is enough: embeddings, rerankers and how learning to rank uses features and graded labels. Senior relevance and ranking roles go deeper into click models, ranking losses and experiment design.

What is the difference between search relevance and recommendation systems?

Search responds to an explicit query and must match its intent, while recommendations predict interest without a query. They share ranking models, feature pipelines and evaluation methods, so the skills transfer in both directions.

Is search engineering a good career for Indian engineers?

Search sits behind e-commerce, enterprise knowledge tools and RAG assistants, so the skills are relevant at product companies, GCCs and services firms in Hyderabad, Bengaluru and elsewhere. Strong profiles combine relevance depth with production engineering and evaluation discipline.

How long does it take to prepare for a search relevance interview?

It depends on your background. A backend developer can cover the core concepts and build a measured hands-on project in a few focused weeks; a fresher should allow longer and spend most of the time building, evaluating and debugging real queries.

If you want guided, hands-on practice with retrieval, semantic search, evaluation and cloud deployment, the Cloudsoft APEX program covers AI, ML, cloud and security engineering together. To apply search and retrieval inside customer-facing enterprise AI systems, look at the AI Forward Deployed Engineer FDE PRO program, which includes placement support until you're placed. Classroom sessions run in Ameerpet, Hyderabad, or live online; call +91 96660 19191 for a free demo.

Share𝕏infβœ‰
EnrollWhatsAppCall us