New batches starting this week Β· Limited seats

NLP Interview Questions and Answers 2026 (60 Questions)

60 natural language processing interview questions with practical answers, from preprocessing and embeddings to NER, evaluation metrics, Indian-language NLP and production scenarios.

NLP interview questions 2026: 60 questions on tokenisation, embeddings, NER and classification, encoders vs decoders, metrics and Indic NLP
Last updated Β· 43 min read Β· 9,518 words

NLP interview questions in 2026 test whether you can pick the right tool for a language problem: a TF-IDF baseline, a fine-tuned encoder, an LLM prompt or a hybrid of the three, and then prove with the right metric that it works on real, messy text. This guide collects 60 high-value natural language processing interview questions with model answers, from preprocessing and tokenisation through embeddings, NER, classic tasks and evaluation, to multilingual and Indic NLP and production scenarios. It is written for engineers preparing for NLP engineer, applied scientist and AI engineer roles.

How to use this guide

This page focuses on language tasks and how they are solved. Transformer internals, decoding and LLM serving are covered in our LLM interview questions, and training mechanics such as backpropagation, optimisers and RNNs are covered in the deep learning interview questions, so they are only referenced here. What interviewers commonly probe at each level:

  • Freshers: preprocessing, TF-IDF, word2vec intuition, BIO tagging, precision, recall and F1, and why subword tokenisation exists.
  • Mid-level engineers: fine-tuning encoders for classification and NER, label alignment, class imbalance, metric choice and the limits of BLEU and ROUGE.
  • Senior engineers and architects: LLM versus fine-tuned model trade-offs, multilingual and code-mixed text, weak supervision, distillation, latency budgets and drift in production.

Text preprocessing and tokenisation

1. How has a typical NLP pipeline changed from the classic era to 2026?

Answer: The classic pipeline was a chain of hand-built stages: cleaning, tokenising, lowercasing, stop-word removal, stemming, feature extraction (bag-of-words or TF-IDF), then a linear model, SVM or CRF per task. Today most systems start from a pretrained transformer: either a fine-tuned encoder for a narrow, high-volume task, or a prompted LLM for tasks that need generation, reasoning or fast iteration. Preprocessing has shrunk to normalisation and the model's own tokeniser, and the real work has moved to data quality, labelling, evaluation, latency and cost.

Classic methods survive as baselines, as retrieval features (BM25) and as fallbacks where compute is constrained.

2. Which text preprocessing steps still matter with transformer models, and which ones hurt?

Answer: With pretrained transformers, apply only the preprocessing the model saw during pretraining. Keep: Unicode normalisation, removing markup and boilerplate (HTML, email signatures, quoted reply chains), de-duplication, consistent handling of URLs, emails, phone numbers and account numbers, and language identification. Usually avoid lowercasing a cased model, stop-word removal ("not" carries meaning), stemming and stripping punctuation.

Noisy user text needs targeted handling rather than aggressive cleaning. Keep emojis (they carry sentiment), normalise elongations ("sooooo" to a bounded form) only if your test set shows it helps, and avoid automatic spelling correction on names, product codes and romanised Indian-language words.

3. What is Unicode normalisation, and why does it matter for Indian-language text?

Answer: Unicode allows the same visible text to be encoded as different code-point sequences. Normalisation forms (NFC, NFD, NFKC, NFKD) convert text into a canonical representation so that visually identical strings compare equal.

For Indic scripts this is not cosmetic. Some characters with a nukta can be written as a precomposed code point or as a base consonant plus a combining nukta. Zero-width joiner and non-joiner characters change how conjuncts and certain letter forms render in scripts such as Devanagari and Malayalam. If training data and production text are normalised differently, the tokeniser produces different tokens for the "same" word, exact-match lookups fail and metrics drop for no visible reason.

Interview tip: Mention that you normalise once, at ingestion, with the same form across training, indexing and inference, and that you test by round-tripping samples from every source system.

4. What is the difference between stemming and lemmatisation, and are they still used?

Answer: Stemming chops affixes with rules (the Porter stemmer turns "studies" into "studi"), which is fast but produces non-words and over- or under-merges. Lemmatisation maps a word to its dictionary form using vocabulary and part-of-speech information ("better" to "good", "running" to "run").

They are rarely used as input to transformers, because subword tokenisation and contextual representations handle morphology implicitly. They still appear in keyword search, analytics dashboards, rule-based systems, and classic baselines. For morphologically rich, agglutinative languages such as Telugu, Tamil, Kannada and Malayalam, rule-based stemmers are weak, and a morphological analyser or subword-based model is usually more reliable.

5. Trace how tokenisation evolved from words to subwords. What problem did each step solve?

Answer: Word-level tokenisation (split on whitespace and punctuation) is intuitive but gives a huge vocabulary, and any unseen word becomes an unknown token. Character-level tokenisation has no unknowns and a tiny vocabulary, but sequences become very long and the model must learn spelling before meaning.

Subword tokenisation is the compromise: frequent words stay whole, rare words split into reusable pieces ("unbelievably" into "un", "believ", "ably"). Every string can be represented, the vocabulary stays fixed, and morphology is partly captured. The main algorithms are BPE (merge the most frequent pairs), WordPiece (merges chosen by likelihood gain, used by BERT) and the Unigram language model (start large, prune pieces that least reduce likelihood). Merge mechanics are covered in the LLM interview guide; the practical consequences for cost and context are explained in tokens and context windows explained.

6. What is SentencePiece, and why is it popular for multilingual models?

Answer: SentencePiece is a tokeniser library that trains BPE or Unigram models directly on raw text, treating whitespace as an ordinary symbol (shown as "▁"). Because it needs no language-specific pre-tokenisation, it works the same way for English, Hindi, Telugu, Chinese or Thai, where word boundaries are absent or written differently.

For multilingual models, the important decisions are vocabulary size and how the training corpus is sampled across languages. If English dominates the tokeniser's training data, Indic scripts get fragmented into many small pieces. Multilingual models usually upsample low-resource languages during tokeniser training to give them fairer vocabulary coverage.

Classic representations: BoW, TF-IDF, n-grams

7. What is bag-of-words, and what are its limitations?

Answer: Bag-of-words represents a document as a vector of word counts (or binary presence) over a fixed vocabulary, ignoring order. Limitations: word order is lost ("dog bites man" equals "man bites dog"), negation and context are invisible, vectors are huge and sparse, synonyms are unrelated dimensions ("refund" and "reimbursement" share nothing), and any word outside the vocabulary is ignored.

8. How does TF-IDF work, and how does it relate to BM25 and dense retrieval?

Answer: TF-IDF weights each term by its frequency in the document (TF) multiplied by its inverse document frequency (IDF, typically the log of total documents divided by documents containing the term). Words that appear everywhere ("the", "please", "issue") get low weight; distinctive words get high weight.

BM25 is a retrieval scoring function in the same family. It adds TF saturation (the tenth occurrence of a word adds little) and document-length normalisation, which is why it remains the default lexical ranker in search engines. Sparse methods win on exact identifiers, rare terms, codes and names; dense embeddings win on paraphrase and meaning. That is why production retrieval often combines both, as explained in hybrid search and reranking.

9. What are n-grams, and when are character n-grams better than word n-grams?

Answer: An n-gram is a contiguous sequence of n tokens. Adding word bigrams to bag-of-words captures short phrases and some negation ("not good", "credit card").

Character n-grams (for example three to five characters, within or across word boundaries) are robust to typos, inflection and spelling variation, because "refund", "refnd" and "refunds" share most of their character n-grams. That makes them valuable for noisy chat text, romanised Hindi or Telugu where the same word is spelt many ways, and morphologically rich languages. To control memory, the hashing trick maps n-grams into a fixed number of buckets without storing a vocabulary, at the cost of occasional collisions.

10. Why build a TF-IDF plus logistic regression baseline before fine-tuning a transformer?

Answer: It takes minutes to train, gives you a real number to beat, and exposes data problems early. If the baseline is already close to your target, a transformer may not justify its serving cost. If the baseline is unexpectedly high, look for leakage, such as a template phrase or agent name that reveals the label.

Word embeddings to contextual embeddings

11. What is the distributional hypothesis, and how does word2vec use it?

Answer: The distributional hypothesis says words that appear in similar contexts have similar meanings. Word2vec learns a dense vector per word by training a shallow network on a prediction task over a sliding window. CBOW predicts the centre word from its surrounding words; skip-gram predicts surrounding words from the centre word. Skip-gram is generally better for rare words, while CBOW is faster.

A full softmax over the vocabulary is expensive, so word2vec uses negative sampling: for each true (word, context) pair it samples a few random "negative" words and trains a binary classifier to tell real pairs from fake ones. The learned vectors place semantically related words close together and capture some regular relationships, such as the well-known analogy patterns, although analogies are a fragile evaluation.

12. How does GloVe differ from word2vec?

Answer: Word2vec is predictive and learns from local context windows one pair at a time. GloVe is count-based: it first builds a global word-word co-occurrence matrix over the corpus, then learns vectors whose dot products approximate the log of co-occurrence counts, using a weighted least-squares objective that down-weights very rare and very frequent pairs. Both produce one static vector per word.

13. What does fastText add, and why is it useful for Indian languages?

Answer: FastText represents each word as the sum of vectors for its character n-grams (with boundary markers), plus the word itself. Two consequences follow. It can build a vector for any unseen word from its pieces, so typos and new inflections are not unknown. And words sharing morphology share parameters, which helps languages where one root produces many forms, such as Telugu, Tamil, Kannada, Malayalam and Marathi. FastText also ships a very fast linear text classifier and a widely used language-identification model, both useful as baselines or for routing on CPU.

14. Why are static embeddings not enough, and what are contextual embeddings?

Answer: A static embedding gives one vector per word regardless of context, so "bank" in "river bank" and "bank account" is the same point, a blend of its senses. Contextual embeddings compute a vector for each token occurrence as a function of the whole sentence. The same word now gets different vectors in different contexts, which captures polysemy, negation scope and syntactic role. The trade-off is compute: you run a deep network for every input instead of looking up a table.

15. How do you get good sentence embeddings, and how do you check they work for your domain?

Answer: A raw pretrained encoder is not trained to place similar sentences close together, so taking its [CLS] vector or averaging token vectors often gives mediocre similarity. Sentence-embedding models are further trained with contrastive objectives on paired data (paraphrases, question-answer pairs, query-passage pairs), so that related texts are close in cosine space. The concepts behind vector similarity are covered in embeddings explained.

Check fit with an extrinsic test, not leaderboard scores: build a few hundred labelled pairs or queries from your own data (including your languages and jargon) and measure recall at k, MRR or clustering purity. Also test failure modes specific to embeddings: negation ("card blocked" versus "card not blocked" are usually close), numbers and identifiers, and cross-lingual pairs if users mix languages.

Encoders, decoders and encoder-decoders

16. Which NLP tasks map naturally to encoder, decoder and encoder-decoder models?

Answer: Encoder-only models (the BERT family, multilingual encoders such as XLM-R and MuRIL) see the whole input bidirectionally and output a representation per token. They fit understanding tasks with a fixed output space: classification, NER, extractive QA, embeddings and reranking.

Encoder-decoder models (T5, mT5, BART, IndicBART-style models, dedicated translation models) encode the input then generate a new sequence while attending to it. They fit input-to-output transformations where the output is a different text: translation, summarisation, and structured generation from documents. Decoder-only LLMs generate left to right and handle almost any task through prompting, which makes them flexible but larger and costlier per call.

Interview tip: A crisp answer is: "fixed label space and high volume, start with an encoder; text-to-text with a well-defined source, consider an encoder-decoder; open-ended tasks or few labels, start with an LLM and distil later if volume justifies it."

17. How does masked language modelling differ from causal language modelling as a pretraining objective?

Answer: Masked language modelling (BERT) hides a fraction of input tokens (about 15 percent in the original BERT recipe) and trains the model to recover them using context on both sides. This produces strong bidirectional representations but no natural way to generate text. Causal language modelling (GPT-style) predicts each next token from the tokens before it, which directly trains generation. The objective explains why encoders fine-tune efficiently for classification while decoders are natural for prompting.

18. How do you fine-tune an encoder for sentence classification, token classification and extractive QA?

Answer: The pretrained encoder stays the same; only the head and labels change.

  • Sentence classification: pool the sequence (the [CLS] token or mean pooling), add a linear layer over the classes, train with cross-entropy (or binary cross-entropy per label for multi-label).
  • Token classification (NER, POS): a linear layer on every token's output predicts a tag such as B-PER or O; loss is computed per token, with padding and non-first subwords masked out.
  • Extractive QA: feed question and context as a pair; two linear outputs score each token as the answer start and end; the predicted answer is the highest-scoring valid span, or "no answer" if the null score wins.

19. Classic BERT-style encoders accept around 512 tokens. How do you handle longer documents?

Answer: Options, in rough order of simplicity: truncate if the signal is at the start (ticket subject and first paragraph often suffice); chunk with overlap, classify or tag each chunk, and aggregate (max, mean or attention over chunk vectors; for NER, merge entities across overlapping regions); select relevant passages first with a cheap retriever, then run the encoder on those; or use a newer long-context encoder (models such as ModernBERT support much longer inputs).

Sequence labelling and information extraction

20. What is BIO tagging, and how does BIOES differ?

Answer: BIO turns span extraction into per-token classification. B-X marks the beginning of an entity of type X, I-X marks a token inside it, and O marks tokens outside any entity. "Sreya Rao lives in Hyderabad" becomes B-PER I-PER O O B-LOC. BIOES (also called BILOU) adds E for the end token and S for single-token entities, giving the model more explicit boundary signals; it sometimes helps slightly, at the cost of more labels. Standard BIO cannot represent nested entities ("State Bank of India Hyderabad branch" containing an organisation and a location), which need span-based or layered approaches.

21. What does a CRF layer add on top of a token classifier?

Answer: A plain softmax head predicts each token's tag independently, so it can output invalid sequences such as O followed by I-LOC, or B-PER followed by I-ORG. A linear-chain conditional random field models the whole tag sequence: it combines per-token emission scores with learned transition scores between adjacent tags, and at inference the Viterbi algorithm finds the highest-scoring valid sequence.

CRFs were the standard for NER with hand-crafted features and later with BiLSTM encoders. With strong transformer encoders the gain is often small, but a CRF, or simply constrained decoding that forbids invalid transitions, still helps on small datasets and on long entities where boundary consistency matters.

22. How do you align word-level labels with subword tokens in NER?

Answer: Annotations are per word, but the model sees subwords: "Bathalapalli" may become several pieces. The common approach is to assign the word's label to its first subword and mark the remaining subwords with an ignore index (for example -100 in PyTorch losses) so they do not contribute to the loss; at inference, read predictions from first subwords only. You need the tokeniser's word-to-token mapping (offset mappings) to do this correctly, and to project predicted spans back to character offsets in the original text, which downstream systems need.

23. How is NER solved in 2026: fine-tuned encoder, LLM extraction or a hybrid?

Answer: It depends on volume, label stability and latency.

ApproachStrengthsWeaknesses
Fine-tuned encoder (token classification)Fast on CPU or small GPU, cheap per document, exact character offsets, stable outputsNeeds labelled data per entity type; adding a type means relabelling and retraining
LLM extraction with a JSON schemaWorks with few or no labels, handles new types by editing the prompt, can normalise valuesHigher cost and latency, may paraphrase or invent values, offsets must be recovered, output varies with prompt and model version
HybridLLM labels data or handles rare types; encoder serves the bulk; rules validate formatsMore moving parts to evaluate and monitor

For pattern-defined entities (PAN, Aadhaar-format numbers, IFSC codes, dates, amounts), regular expressions with checksums or validators beat any model and should run alongside it. A common production pattern is to bootstrap labels with an LLM, review a sample by hand, fine-tune an encoder for serving, and keep the LLM for low-confidence or long-tail cases.

24. What is relation extraction, and how does entity linking fit into an extraction pipeline?

Answer: Relation extraction identifies typed relationships between entities, for example (Policy-123, insured_party, Ravi Kumar) or (drug, treats, condition). LLMs with a schema of allowed relation types now handle many low-volume relation tasks well, but need validation against the schema and the text.

Entity linking maps an extracted mention to a canonical record: "SBI", "State Bank" and "State Bank of India" to one organisation ID, or a customer name to a CRM entry. Businesses rarely want raw strings; they want the right record. Linking typically combines candidate generation (alias tables, fuzzy matching, embeddings) with disambiguation using context.

Classic NLP tasks in 2026

25. For a new text classification task, how do you choose between a fine-tuned encoder, embeddings plus a classifier, and LLM prompting?

Answer: Decide on four factors: labelled data available, request volume, latency budget and how often the label set changes.

  • Zero- or few-shot LLM prompting: no or few labels, low to moderate volume, labels that change often, or when you need an explanation with the label. Fast to start, highest cost per item.
  • Frozen embeddings plus a light classifier (logistic regression or kNN): a few hundred labels, need for cheap retraining, many classes that are added over time. Good accuracy-to-effort ratio.
  • Fine-tuned encoder: thousands of labels, high volume, tight latency or on-premises constraints. Usually the most accurate per unit of compute for a stable taxonomy.

Many teams start with an LLM to launch and to generate training labels, then move the bulk of traffic to a fine-tuned or distilled small model once volume grows.

26. How do you handle multi-label classification, class imbalance and threshold choice?

Answer: Multi-class means exactly one label per item (softmax); multi-label means any subset (independent sigmoid per label, binary cross-entropy).

For imbalance, first fix evaluation (macro F1, per-class recall, precision-recall curves rather than accuracy). Then consider class-weighted loss, oversampling rare classes, targeted labelling of rare classes, or merging classes nobody actually routes differently. Thresholds should be tuned per class on a validation set against the business cost of each error type, not left at 0.5. A "low confidence" band that routes to a human or a larger model is often worth more than another point of F1.

27. Why is sentiment analysis harder than it looks, and what is aspect-based sentiment?

Answer: Document-level positive or negative labels hide most of what businesses need. A review saying "delivery was quick but the phone heats up" is mixed: positive on delivery, negative on product. Aspect-based sentiment analysis extracts aspects (delivery, battery, price, staff) and the sentiment towards each, either as a joint extraction task or as classification over a fixed aspect list.

Other difficulties: negation and contrast ("not bad at all"), sarcasm, comparative statements, domain shift ("unpredictable" is good for a film plot, bad for a car's brakes), and code-mixed or romanised text. Lexicon-based methods fail on most of these.

28. Compare extractive and abstractive summarisation. What is the main production risk?

Answer: Extractive summarisation selects existing sentences (by centrality, as in TextRank, or a trained sentence scorer). It cannot invent facts but can be choppy and miss information spread across sentences. Abstractive summarisation generates new text with an encoder-decoder or LLM, which is fluent and concise but can introduce unsupported facts, wrong numbers or wrong attributions.

The main production risk is faithfulness. Mitigations: constrain the format (fixed sections, bullet fields), require citations to source sentences or timestamps, run a faithfulness check (an NLI model or an LLM judge testing each summary claim against the source), and validate numbers, names and dates by string matching. ROUGE barely detects these errors, so evaluate with targeted human review as well (see Q35 and Q36).

29. What is the difference between extractive QA and generative QA, and where does RAG fit?

Answer: Extractive QA (SQuAD-style) returns a span from a given passage, using an encoder with start and end heads. Answers are verifiable and grounded, but the system must be given the right passage and cannot synthesise across sources. Generative QA produces a free-text answer, either from the model's own parameters (closed-book, risky for facts) or from retrieved passages (retrieval-augmented generation). RAG is the dominant enterprise pattern because it combines retrieval for grounding with generation for synthesis. For RAG-specific depth see the RAG interview questions.

30. How is machine translation done in 2026, and when would you use a dedicated MT model instead of an LLM?

Answer: Neural MT uses encoder-decoder transformers trained on parallel corpora; LLMs can also translate through prompting, often fluently and with good handling of context and tone. Choose a dedicated MT model when you need high throughput at low cost, self-hosting for data residency, predictable output (no added commentary), or strong coverage of specific language pairs. For Indian languages, open models such as AI4Bharat's IndicTrans2, which covers all 22 scheduled Indian languages, are an important option to evaluate.

Use an LLM when context, style or document-level coherence matters more than throughput, or when translation is combined with other steps such as summarising. In either case, enforce terminology with glossaries (via prompt instructions or constrained decoding), protect placeholders and entities, and evaluate per language with chrF or a learned metric plus human review by native speakers.

31. Compare LDA with embedding-based topic modelling.

Answer: LDA is a generative probabilistic model: each document is a mixture of topics, each topic a distribution over words, inferred from bag-of-words counts. It is interpretable and cheap, but needs careful preprocessing, a fixed number of topics, and struggles with short texts like tweets or ticket subjects. Embedding-based approaches (for example BERTopic) embed documents, reduce dimensions, cluster them with a density-based algorithm, and describe each cluster with class-based TF-IDF keywords. They handle short and multilingual text better and capture semantic similarity, at the cost of more compute and some instability across runs.

Evaluation metrics and test design

32. Explain precision, recall and F1, and when to use micro, macro or weighted averaging.

Answer: Precision is the share of predicted positives that are correct; recall is the share of actual positives that were found; F1 is their harmonic mean, which punishes imbalance between the two. For multiple classes: micro-averaging pools all decisions, so frequent classes dominate (for single-label multi-class it equals accuracy); macro-averaging computes F1 per class and averages equally, so rare classes count as much as common ones; weighted averaging weights each class by its support, which sits between the two.

Report macro F1 and per-class metrics when rare classes matter (fraud complaints, legal threats), micro or weighted when overall throughput matters, and always pair them with a confusion matrix.

33. How is NER evaluated, and why does token-level accuracy mislead?

Answer: Most tokens are O, so token-level accuracy looks high even for a useless model. The standard is entity-level evaluation (as in the CoNLL convention and the seqeval library): a predicted entity counts as correct only if both the span boundaries and the type exactly match a gold entity; precision, recall and F1 are computed over entities. Partial-match schemes give credit for overlapping spans or correct boundaries with wrong types, which helps diagnose errors.

34. How do exact match and F1 work in question answering?

Answer: After normalising both strings (lowercasing, removing punctuation, articles and extra whitespace), exact match is 1 if the prediction equals any gold answer, else 0. Token-level F1 treats prediction and gold answer as bags of tokens and computes overlap precision and recall, taking the maximum over gold answers. F1 gives partial credit ("Hyderabad, Telangana" versus "Hyderabad"). For long generative answers they break down, because correct answers can be phrased in many ways; there you need judged correctness and faithfulness, covered in the LLM evaluation interview questions.

35. How do BLEU, ROUGE and chrF work, and what are their limits?

Answer: BLEU, the classic MT metric, measures modified n-gram precision (typically up to four-grams) of a candidate against references, combined with a brevity penalty, and is designed for corpus-level use. ROUGE, common for summarisation, is recall-oriented: ROUGE-1 and ROUGE-2 count unigram and bigram overlap, ROUGE-L uses the longest common subsequence. chrF computes an F-score over character n-grams, which is more forgiving of inflection and segmentation differences, so it often correlates better with human judgement for morphologically rich languages such as many Indian languages.

Limits: all are surface-overlap metrics, so valid paraphrases score low, while fluent output with a wrong number or a negation flipped can score high. They depend heavily on reference quality and count, and scores are sensitive to tokenisation. Learned metrics such as COMET (for MT) and BERTScore use neural models to compare meaning and usually track human judgement better, but they inherit their models' biases and language coverage. Use overlap metrics for regression tracking, not as proof of quality.

36. When do you need human evaluation, and how do you make it reliable?

Answer: Whenever automatic metrics cannot capture what matters: faithfulness, usefulness, tone, terminology, cultural appropriateness, and any language where your metrics or judges are poorly validated. Make it reliable with written guidelines and examples per score, a defined rubric (for example adequacy, fluency and factual errors as separate scores, or pairwise preference), blind and randomised presentation so raters do not know which system produced what, multiple raters on an overlapping subset, and an agreement measure such as Cohen's kappa for two raters or Krippendorff's alpha for more.

37. How do you build a test set that predicts production performance?

Answer: Sample from real production traffic, not from the clean training data. Split by time (test newer than train) and by entity such as customer or template, and remove near-duplicates across splits. Stratify so rare but important classes have enough examples. Define slices that matter (language, script, channel, region, product line, document type) and report metrics per slice, because averages hide failures. Freeze a gold set for regression testing, keep a separate rolling set for drift, and never tune prompts or thresholds on the gold set.

Multilingual and Indic NLP

38. What makes NLP for Indian languages hard?

Answer: Several factors combine:

  • Many scripts: most Indian languages use Brahmic abugidas (Devanagari, Bengali, Gurmukhi, Gujarati, Odia, Telugu, Kannada, Tamil, Malayalam), and Urdu uses a Perso-Arabic script. Conjuncts, vowel signs and joiners create encoding and normalisation issues.
  • Rich morphology: Dravidian languages are agglutinative, so one long word can encode what English says in a phrase, which inflates vocabulary and hurts word-level methods.
  • Romanised writing: much informal text is typed in Latin script with no standard spelling.
  • Code-mixing: users switch between English and one or more Indian languages within a sentence.
  • Uneven resources: labelled data and evaluation sets are much thinner than for English, and quality varies widely between languages.
  • Tokenisation cost: tokenisers trained mostly on English split Indic text into many more tokens.

39. What is the difference between transliteration and translation, and how do you handle romanised text?

Answer: Translation changes the language while preserving meaning ("where is my refund" into Hindi). Transliteration changes the script while preserving the language and pronunciation ("kahan hai mera refund" written in Devanagari). Romanised Indian-language text is very common in chat, reviews and support tickets, and spelling is inconsistent ("kya", "kyaa", "kia").

Approaches: train or fine-tune directly on romanised text (character n-gram models and multilingual encoders that saw transliterated data, such as MuRIL, handle it reasonably); transliterate to native script with a model such as AI4Bharat's IndicXlit (trained using the Aksharantar transliteration dataset), then use native-script models; or let an LLM handle it, checking quality per language. Back-transliteration is ambiguous (one romanisation can map to several native spellings), so evaluate transliteration on your own text.

40. What is code-mixing, and why does it break standard NLP pipelines?

Answer: Code-mixing is switching languages within a sentence or conversation, such as Hinglish ("payment fail ho gaya, please check") or Telugu-English mixes, often in Roman script. It breaks pipelines that assume one language per document: sentence-level language ID returns a single wrong label, monolingual tokenisers and lexicons miss half the words, translation-first pipelines translate fragments badly, and sentiment cues sit in the non-English part ("bakwas service").

Practical handling: avoid hard routing by language; use multilingual models fine-tuned on in-domain code-mixed data; use character n-gram features in baselines; consider token-level language identification when downstream steps need it; and build test sets that reflect actual mixing ratios and scripts.

41. How do you measure tokeniser fertility, and why does it influence model choice?

Answer: Fertility is the average number of tokens a tokeniser produces per word (or per character) of a language. Measure it by tokenising a representative sample of your real text in each language and script, including romanised and code-mixed samples, and comparing tokens per word against English. High fertility means more cost per request on token-billed APIs, more of the context window used, slower generation and sometimes lower quality, because the model sees fragmented pieces.

Add fertility to the model-selection matrix alongside accuracy: models whose tokenisers saw more Indic data can be cheaper and better for the same task (see tokens and context windows).

42. What open resources exist for Indic NLP, and how do you evaluate whether to use them?

Answer: AI4Bharat, a research lab at IIT Madras, publishes many open resources: IndicBERT and IndicBART models, the IndicTrans2 translation models, IndicXlit and the Aksharantar dataset for transliteration, Naamapadam, a named entity dataset for Indian languages, with IndicNER models trained on it, and monolingual corpora such as IndicCorp and Sangraha. Government initiatives such as Bhashini, the National Language Translation Mission, provide language services and datasets. Multilingual models from global labs (XLM-R, mT5, many open-weight LLMs) also cover Indian languages to varying degrees, and Google's MuRIL targets Indian languages including transliterated text.

Evaluate each before adopting: licence and commercial-use terms (check current terms on each project's page), which languages and scripts are actually covered and how much data each had, the domain of the training data (news and Wikipedia differ greatly from support chat), and performance on your own slice-level test set. Resources change quickly, so check current releases rather than relying on older benchmarks.

43. What are cross-lingual transfer, translate-train and translate-test?

Answer: Cross-lingual transfer means training on labelled data in one language (usually English) with a multilingual model and applying it to other languages, relying on shared representations. It works surprisingly well for related tasks but degrades for distant languages, scripts underrepresented in pretraining, and culture-specific labels. Translate-train machine-translates the training data into the target language and trains on it; translate-test translates incoming text into English and uses an English model. In practice, the strongest results usually come from even a small set of native in-domain labels added to transferred or translated data, so budget for some native annotation.

Data labelling, weak supervision and bias

44. How do you design annotation guidelines and measure label quality?

Answer: Start with a pilot: two or three people label the same small batch, disagreements are discussed, and the guideline is rewritten with definitions, decision rules for borderline cases, positive and negative examples, and an explicit "unclear" option. Then measure inter-annotator agreement on an overlapping sample throughout the project (Cohen's kappa, Krippendorff's alpha, or entity-level F1 between annotators for NER). Agreement sets a practical ceiling: if humans agree only moderately, a model reported as far above that is probably overfitting to one annotator's habits.

45. What is weak supervision, and how do labelling functions, active learning and LLM labellers fit together?

Answer: Weak supervision trains models from noisy, programmatic labels instead of fully hand-labelled data. In the labelling-function approach (popularised by Snorkel), experts write many small heuristics (keyword rules, regexes, lookups, outputs of existing models) that vote or abstain on each example; a label model estimates each function's accuracy and correlations and combines votes into probabilistic labels; a discriminative model is then trained on those labels and generalises beyond the rules.

LLMs are now a powerful labelling function or primary labeller: prompt them with the guideline, keep their confidence or agreement across prompts, and audit samples against human labels. Active learning complements both: the model selects the examples it is most uncertain about, or that are most diverse, for human labelling, so annotation budget goes where it changes the model most. A practical loop is LLM pre-labels, human correction of uncertain items, retrain, repeat. Synthetic data has similar trade-offs, covered in synthetic data for AI testing.

46. Where does bias show up in NLP systems, and how do you test for it?

Answer: Bias enters through pretraining text (embeddings associating professions with gender, or names from certain communities with negative contexts), through labelling (annotators rating dialect or romanised text as more "toxic" or "rude"), and through coverage (poorer accuracy for under-represented languages, regions or writing styles). Effects include moderation systems that over-flag certain dialects, and support classifiers that route regional-language customers to slower queues.

Test with sliced evaluation (metrics per language, script, region, gender-coded names), counterfactual tests (swap names, gendered terms, cities or communities in otherwise identical text and check that outputs do not change), and targeted challenge sets reviewed by people from the affected groups. See AI bias and fairness testing for a fuller method.

Production NLP: latency, small models, distillation

47. How do you meet tight latency and cost budgets for an NLP model in production?

Answer: Start from the budget (for example, classification must finish inside a ticketing system's synchronous call) and work back. Levers, roughly in order of effort:

  • Pick a smaller model: compact encoders (DistilBERT-style, MiniLM-style, small multilingual encoders) are often close in accuracy to larger ones for narrow tasks.
  • Optimise the runtime: export to ONNX or a similar optimised runtime, apply INT8 quantisation, batch requests dynamically, and keep sequence lengths short with dynamic padding.
  • Avoid work: cache results for repeated texts, truncate to the useful part of the input, and use a cheap first-stage filter.
  • Cascade: a small model answers confident cases; uncertain cases go to a larger model or an LLM.

Measure p95 and p99 latency, not averages, under realistic load. For broader context on when small models suit enterprises, see small language models in the enterprise; for LLM-side tuning, LLM latency optimisation.

48. How do you distil an LLM into a small task model, and how do you route between them?

Answer: Knowledge distillation trains a small student to imitate a larger teacher. For NLP tasks the practical version is: run the LLM (with a carefully evaluated prompt) over a large set of unlabelled in-domain texts, keep its labels and, if available, probability-like scores; filter or human-review a sample; fine-tune a small encoder on these labels plus any gold labels; then evaluate the student against a human-labelled test set, not against the teacher's outputs. The method is explained in model distillation explained; check the teacher model's terms of use before training on its outputs.

text -> small model -> confidence >= t ? -> label
                          |
                          no
                          v
                    LLM (or human) -> label
                          |
                          v
              log for review and retraining

Production consideration: Tune the threshold t on validation data so that the routed share fits the budget and the accuracy target, and monitor that share: if it rises, the input distribution has probably shifted.

If you want to practise this end to end, from data labelling to evaluation and deployment on cloud, the APEX AI, ML, Cloud and Cyber Security program covers the ML and AI engineering foundations behind these answers.

Real-world scenario questions

49. Scenario: an IT services GCC in Hyderabad wants to auto-route support tickets written in English, Hindi, Telugu and romanised mixes into about forty queues. How do you design the classifier?

Answer: Treat it as a multilingual, possibly multi-label classification problem with a human fallback. Start by pulling historical tickets with their final resolver group (not the initial, often wrong, assignment) as weak labels. Clean them: strip signatures, quoted replies and auto-generated text, normalise Unicode and mask personal data. Build a character n-gram TF-IDF baseline, then fine-tune a multilingual encoder that has seen Indian languages and romanised text. Serve the encoder with a confidence threshold; low-confidence tickets go to a triage queue whose decisions become new training data.

ticket -> clean + mask PII -> encoder classifier
                                  |
             confident ----------+---------- unsure
                 |                               |
           auto-route                     triage desk
                                               |
                               label -> retrain

What I would check:

  1. How noisy the historical labels are, by having experts relabel a random sample.
  2. Per-language and per-script metrics, especially for romanised Telugu and code-mixed tickets.
  3. Which queues are confused with each other, and whether the business routes them differently at all.

Production consideration: Track reassignment rate and time to first response alongside macro F1, and keep taxonomy changes cheap with periodic retrains.

50. Scenario: an insurer wants to extract names, policy numbers, dates and amounts from scanned claim forms, some handwritten. How do you approach NER on these documents?

Answer: This is a document-understanding problem first and a text NER problem second. Text NER on OCR output loses layout, which carries much of the meaning: a value is usually identified by the label printed next to it. Use a document pipeline: image cleanup, OCR or a document AI service that returns words with bounding boxes and confidence, then either key-value extraction by layout, a layout-aware model that uses text and position together, or a vision-capable LLM prompted with a schema. Validate every field: policy-number format and check digits, date parsing, amount reconciliation against totals, and lookup against the policy database. See document parsing for the parsing options and the multimodal AI interview questions for vision-language models.

What I would check:

  1. OCR quality per form template, scan quality and handwriting share, measured separately from extraction quality.
  2. Field-level accuracy after validation and normalisation, which is what the claims system consumes.

Production consideration: Report straight-through processing rate and field error rate per template, and keep an audit trail linking every extracted value to its image region.

51. Scenario: a retailer wants sentiment and aspect analysis on product reviews that mix Hindi and English, often in Roman script. What do you build?

Answer: Define the output first with the business: which aspects (delivery, packaging, quality, price, returns, seller behaviour) and what they will do with the results. Then sample real reviews, including the code-mixed and romanised ones, and have fluent annotators label aspect-level sentiment with clear rules for sarcasm and mixed opinions. Compare three candidates on that test set: an LLM with a JSON schema for aspects and polarity, a multilingual encoder fine-tuned on the labels (or on LLM-generated labels reviewed by humans), and a character n-gram baseline.

What I would check:

  1. Performance by script and mixing level, not only overall.
  2. Disagreement between star rating and text sentiment, which often reveals sarcasm or aspect-level nuance.
  3. Stability of LLM labels across repeated runs and prompt variants.

Production consideration: Dashboards should show aspect trends with sample reviews attached, so category managers can verify the model's reading rather than trusting a score.

52. Scenario: leadership asks whether to classify a high volume of daily messages with an LLM API or a fine-tuned BERT-style model. How do you make the cost decision?

Answer: Compare total cost of ownership at the required quality, using your own measurements. For the LLM option: messages per day multiplied by average input and output tokens per message (measured on real text, which matters for Indian languages because of fertility) multiplied by per-token prices, plus prompt-engineering and evaluation effort. For the fine-tuned option: labelling cost, training runs, serving infrastructure sized for peak load (often CPU or a small GPU), and engineering time for retraining and monitoring. Use placeholder values first, then replace them with measured ones. Typically the LLM is cheaper at low volume or when labels change often, and the small model wins at sustained high volume with a stable taxonomy, but the crossover point depends entirely on your inputs.

What I would check:

  1. Accuracy of both options on the same human-labelled test set, per class and per language.
  2. Latency at p95 against the integration requirement.
  3. How often the label set changes, and the cost of each retrain.

Production consideration: The answer is often "both": launch with the LLM, use it to generate training labels, distil to a small model, and keep the LLM for low-confidence cases.

53. Scenario: a support classifier that launched with good accuracy has slowly degraded over six months. What happened, and what do you do?

Answer: Most likely drift: new products, new error messages, a new channel such as WhatsApp with shorter and more informal text, reorganised queues, or seasonal issues.

What I would check:

  1. Input statistics over time: length, language and script mix, unknown or rare-token rates, channel distribution.
  2. Prediction distribution and confidence over time, and the share routed to humans.
  3. A fresh labelled sample from recent traffic, to measure real accuracy instead of guessing.
  4. Whether label definitions changed (queues merged or split) without the model being retrained.

Production consideration: Put drift monitors, a rolling labelled sample and a scheduled retrain in place, and version models with the data they were trained on. Operational practices like these are covered in the MLOps interview questions.

54. Scenario: a hospital wants to de-identify discharge summaries before using them for analytics. How do you build the PII detection?

Answer: De-identification is NER with an asymmetric cost: a missed name or phone number is a privacy incident, while an over-redacted word only costs some analytic value. So optimise for recall. Combine pattern detectors (phone numbers, email addresses, ID-number formats, dates, medical record numbers) with an NER model fine-tuned on annotated clinical notes for names, addresses and facilities, and add lookup lists of staff and local place names. Replace detected spans with consistent typed placeholders or realistic surrogates, so analytics can still distinguish "patient" from "doctor".

What I would check:

  1. Recall per PII type on a held-out set annotated by two people, with residual misses reviewed manually.
  2. Performance on different note templates, departments and hospitals.
  3. Legal basis and retention under the DPDP Act, with the privacy team; see DPDP Act for AI applications.

Production consideration: Keep processing inside the approved environment, log only placeholders, and re-audit samples periodically because templates change.

55. Scenario: an insurer wants policy documents and customer letters translated into Hindi, Telugu and Tamil. How do you ensure quality?

Answer: Legal and financial text needs accuracy and consistent terminology more than fluency. Build a bilingual glossary of insurance terms with the business and language experts, decide which terms stay in English, and enforce it through the MT system or the LLM prompt. Protect placeholders, amounts, dates and clause numbers so they pass through unchanged, and verify them automatically after translation. Evaluate candidate systems (a dedicated Indic MT model, a commercial MT service, an LLM) on a test set of real documents per language using chrF or COMET plus native-speaker review for adequacy and terminology. For customer-facing legal text, keep human post-editing and sign-off in the workflow.

What I would check:

  1. Glossary compliance rate and placeholder integrity, which can be checked automatically.

Production consideration: Store source, machine output and post-edited versions; the post-edits become training and evaluation data for the next iteration.

56. Scenario: call summaries generated for a bank's contact centre sometimes state commitments the agent never made. How do you fix this?

Answer: This is a faithfulness failure, possibly compounded by speech-to-text errors. First separate the stages: check the transcript against audio for a sample to see whether errors start in ASR (which is weaker on code-mixed speech) or in summarisation. For summarisation, switch to a structured format (customer issue, actions taken, commitments with timestamps), require every commitment to cite the transcript turn it came from, and add a verification step that checks each claimed commitment against the cited turn and drops or flags unsupported ones. For the transcription side, see the voice AI interview questions, and for a worked design, the call centre QA AI project.

What I would check:

  1. The share of unsupported claims on a labelled sample before and after the change.
  2. Whether downstream systems act on summaries automatically, which raises the required precision.

Production consideration: For regulated commitments, keep the summary advisory and link to the source turns, so a supervisor can verify in seconds.

57. Scenario: a bank's chatbot uses an LLM to classify user messages into several hundred intents, and it confuses similar intents and costs too much. What do you change?

Answer: Putting hundreds of intent descriptions into every prompt makes requests long and expensive and forces fine distinctions in one step. Restructure: use an embedding-based retriever or a fine-tuned encoder to shortlist the top few candidate intents, then let the LLM (or a cross-encoder) choose among those with their full descriptions and examples. Alternatively organise intents hierarchically (domain, then intent). Often a fine-tuned encoder over the full intent set, trained on historical utterances, handles most traffic alone, with the LLM only for low-confidence cases.

What I would check:

  1. The confusion matrix to find intent pairs that are genuinely ambiguous even to humans.
  2. Recall of the shortlist stage: if the right intent is not in the top candidates, the second stage cannot recover.
  3. Token count per request before and after.

Production consideration: Version the intent catalogue with its examples and re-run the evaluation whenever intents are added.

58. Scenario: a government-facing service receives the same complaint in many languages and wants duplicates grouped. How do you design it?

Answer: Use a multilingual sentence-embedding model so that semantically similar complaints land close together regardless of language or script, store embeddings in a vector index, and for each new complaint retrieve nearest neighbours within a time window and location filter. A cross-encoder or LLM can confirm whether top candidates describe the same incident. Cluster confirmed duplicates and attach them to a master record. Romanised text may need transliteration or a model trained on it, because some multilingual embedding models handle native script much better than Roman script.

What I would check:

  1. Cross-lingual retrieval quality on labelled duplicate pairs, per language pair and script.
  2. False merges of distinct incidents with similar wording (two different broken streetlights on the same road), which location and time filters should prevent.

Production consideration: Keep merges reversible and show staff why two complaints were grouped.

59. Scenario: a content moderation model on a community platform flags posts in one regional dialect far more often than others. How do you investigate and fix it?

Answer: This is a likely bias in data or labels rather than a true difference in behaviour. Measure first: build a sample of posts in that dialect and others, have fluent reviewers label them under the same guideline, and compare false-positive rates per group. Common causes are annotators unfamiliar with the dialect labelling ordinary slang as abusive, training data in which that dialect appears mainly in toxic examples, and lexicon features matching words that are offensive in one variety but neutral in another.

What I would check:

  1. False-positive and false-negative rates per dialect and script on a fairly labelled sample.
  2. Counterfactual tests: the same meaning expressed in different dialects.
  3. Which features or tokens drive the flags, using attribution or error analysis.

Production consideration: Fix the data and guidelines, add dialect-specific evaluation slices to release gates, and route borderline cases to reviewers fluent in the dialect. See the responsible AI interview questions for governance angles.

60. Scenario: an interviewer asks you to walk through how you would take an NLP feature from business request to production. How do you structure the answer?

Answer: Use a sequence that shows judgement at every step:

  1. Problem: what decision the output drives, who uses it, and the cost of each error type.
  2. Data: sources, languages and scripts, volume, privacy constraints, and existing labels and their noise.
  3. Test set and metric: a production-like, sliced test set and a metric tied to the business cost.
  4. Baseline: TF-IDF or a rules plus regex system, plus a zero-shot LLM, to set the bar quickly.
  5. Model choice: fine-tuned encoder, LLM, or hybrid, decided by volume, latency, label stability and cost.
  6. Integration: APIs, confidence thresholds, human review paths, fallbacks.
  7. Deployment and monitoring: latency, drift, routed-to-human share, per-slice metrics.
  8. Improvement loop: corrections become labels; scheduled retrains or prompt updates with regression tests.

Production consideration: Close with the business outcome, such as faster routing or fewer manual reviews, not with the model's F1. Delivering systems like this inside customer environments is the core of the Forward Deployed Engineer role, covered in the FDE engineer interview questions.

Key takeaways

  • Modern NLP is mostly about choosing between a classic baseline, a fine-tuned encoder, an LLM and a hybrid, and justifying the choice with data, latency and cost.
  • Subword tokenisation solved out-of-vocabulary problems, but tokeniser fertility still drives cost and quality for Indian languages.
  • Static embeddings gave way to contextual ones; for retrieval and similarity, use models trained for sentence embeddings and test them on your own data.
  • Sequence labelling depends on correct BIO tagging, subword alignment and entity-level evaluation; validators and rules still beat models for formatted fields.
  • Overlap metrics such as BLEU and ROUGE track regressions but cannot prove quality; use chrF or learned metrics, sliced evaluation and well-run human review.
  • Code-mixed and romanised text needs in-domain data, fluent annotators and per-script test slices, not translation-first shortcuts.
  • In production, cascades and distillation let a small model serve most traffic while an LLM handles the hard cases.

Interview preparation checklist

  • Train a TF-IDF plus logistic regression baseline and a fine-tuned multilingual encoder on the same dataset, and explain the difference.
  • Fine-tune a token-classification NER model, handle subword alignment, and evaluate with entity-level F1.
  • Measure tokeniser fertility for English, one Indian language and a romanised sample across two or three models.
  • Build an LLM zero-shot classifier, then distil it into a small model and compare accuracy, latency and cost.
  • Compute BLEU, chrF and ROUGE on a few outputs and find one example where the metric disagrees with your own judgement.
  • Write a one-page annotation guideline and measure agreement with a friend on fifty examples.
  • Review neighbouring guides: Hugging Face interview questions for tooling and LLM inference and serving interview questions for deployment.

FAQ

What skills does an NLP engineer need in 2026?

Python, text processing, a solid grasp of tokenisation, embeddings and transformers, fine-tuning encoders, prompting and evaluating LLMs, metric design, data labelling, and enough MLOps and cloud knowledge to deploy and monitor models. Multilingual experience is valuable for Indian products.

Is classic NLP still worth learning when LLMs can do most tasks?

Yes. Classic methods give cheap baselines, power lexical search, help debug data problems and still serve high-volume tasks where cost and latency matter. Interviewers often use them to check that you understand what the models are doing.

How should a fresher prepare for NLP interview questions?

Learn the fundamentals in this guide, then build two or three small projects: a text classifier with a baseline and a fine-tuned model, an NER model with proper evaluation, and an LLM-based extraction task with validation. Be ready to explain every metric you report.

Do I need to know deep learning maths for NLP interviews?

You should understand embeddings, attention at a conceptual level, cross-entropy loss and how fine-tuning works. Applied roles rarely ask for derivations, while research roles may go deeper into training objectives and architectures.

Which NLP topics are commonly asked in interviews?

Commonly asked topics include tokenisation, TF-IDF, word2vec and contextual embeddings, encoder versus decoder models, NER and BIO tagging, precision, recall and F1, BLEU and ROUGE limits, LLM versus fine-tuned model trade-offs, and scenario questions on drift, multilingual text and cost.

Is Indian-language NLP experience useful for my career?

It can be. Many products built in India serve users in several Indian languages, and engineers who understand scripts, transliteration and code-mixed text can solve problems that general-purpose pipelines handle poorly.

Which tools should I practise with for NLP interviews?

Practise with Python, scikit-learn for baselines, PyTorch, the Hugging Face transformers and datasets libraries, a sentence-embedding library, an LLM API, and an evaluation library for metrics such as seqeval and sacreBLEU.

How long does it take to prepare for an NLP engineer interview?

It depends on your background. A developer with some ML experience can usually cover the core topics and one project in a few focused weeks; a fresher should plan for longer and spend most of the time building and evaluating models rather than reading.

Ready to turn these concepts into hands-on skill with real data, cloud deployment and evaluation? Explore Cloudsoft's APEX program for AI, ML, cloud and cyber security. If you want to deliver AI systems inside enterprise customer environments, the AI Forward Deployed Engineer FDE PRO program adds enterprise projects and placement support until you're placed. Both run as classroom sessions in Ameerpet or live online; call +91 96660 19191 for a free demo.

Share𝕏infβœ‰
EnrollWhatsAppCall us