Computer vision interview questions in 2026 test whether you can turn pixels into reliable decisions: pick the right task formulation, choose between CNNs, vision transformers and vision-language models, measure performance honestly, and keep a model working when the camera, lighting or product line changes. Interviewers for CV engineer, applied scientist and edge AI roles want to see that you understand detection, segmentation and OCR beyond library calls, and that you can ship a model to a factory floor, a phone or a fleet of CCTV cameras. This guide covers 55 high-value questions with model answers, from convolution intuition and mAP to SAM-style foundation models, quantisation, video pipelines, DPDP-aware privacy design and eleven production scenarios.
How to use this guide
- Freshers and graduates: first rounds lean on the fundamentals and core tasks sections. Interviewers check that you can explain IoU, NMS or augmentation in your own words and draw them on a whiteboard.
- Working ML or CV engineers: expect architecture, data and evaluation questions with "what broke when you tried that?" follow-ups. Have one project you can walk through end to end: data, labels, model, metric, deployment.
- Senior candidates: deployment, system design and scenarios carry most weight. A strong answer covers the camera and environment, the label policy, the operating threshold, the latency budget and how you will detect drift.
- Neural network training basics (backpropagation, optimisers, normalisation) are covered in the companion deep learning interview questions guide, so this page stays focused on vision.
Contents
- Image fundamentals (Q1β10)
- Core vision tasks: detection, segmentation, OCR, tracking (Q11β24)
- Architectures and foundation models (Q25β33)
- Data, labelling and evaluation (Q34β38)
- Deployment, edge and video pipelines (Q39β44)
- Real-world scenarios (Q45β55)
- Key takeaways
- Interview preparation checklist
- FAQ
Image fundamentals
1. How is a digital image represented for a vision model?
Answer: An image is a grid of pixels stored as a tensor of shape height Γ width Γ channels (or channels Γ height Γ width in most PyTorch code). A colour image usually has three channels with integer values from 0 to 255 per channel at 8-bit depth; grayscale has one. Before a model sees it, you typically convert to floating point, scale to [0, 1], and normalise each channel with the mean and standard deviation the backbone was pretrained with.
Interview tip: Mention the pretrained normalisation statistics. Feeding unnormalised inputs into a pretrained backbone is one of the most common silent bugs in vision code.
2. What colour spaces should a CV engineer know, and when do you leave RGB?
Answer: RGB is the default for deep models. HSV or HSL separates hue from brightness, which helps classical colour thresholding (for example, detecting a red warning light under varying illumination). YCbCr and YUV separate luminance from chroma and are what video codecs and many camera pipelines use natively. Lab is designed to be closer to perceptual uniformity and is used in colour-difference measurement, for instance checking paint or textile shade consistency. Grayscale is enough for many document and inspection tasks and cuts compute.
A classic trap: OpenCV loads images as BGR, while most deep learning libraries and pretrained weights expect RGB. A model trained on RGB and served on BGR frames still runs, just with worse accuracy, which makes the bug hard to spot.
3. Explain convolution intuitively. Why does it work so well for images?
Answer: A convolution slides a small learned filter (say 3Γ3) across the image and computes a weighted sum at each position, producing a feature map that lights up wherever the local pattern matches the filter. Early layers learn edges and colour blobs, middle layers learn textures and parts, and deeper layers respond to object-level patterns.
It suits images because of two built-in assumptions. Locality: nearby pixels are related, so small filters are enough. Weight sharing: the same pattern can appear anywhere, so one filter is reused across all positions. That makes convolution translation-equivariant (shift the input, the feature map shifts too) and drastically reduces parameters compared with a fully connected layer over every pixel.
Interview tip: Be able to compute output size: (W - K + 2P) / S + 1 for input width W, kernel K, padding P and stride S.
4. What are receptive field, stride, padding and pooling?
Answer: The receptive field of a unit is the region of the input image that can influence it; it grows as you stack layers, downsample or use dilated convolutions. Stride is how far the filter moves each step; stride 2 halves spatial resolution. Padding adds border pixels so the output keeps its size and edge pixels are not under-represented. Pooling (max or average) downsamples feature maps, adding some local translation invariance and cutting compute; global average pooling at the end turns a feature map into a vector for classification.
5. Why not just use a fully connected network on raw pixels?
Answer: A fully connected layer on even a modest image needs an enormous number of weights, ignores spatial structure and must learn the same pattern separately at every location. It overfits easily and generalises poorly to shifted or resized objects. Convolutions (and, differently, transformers over patches) build in structure that makes learning from realistic dataset sizes feasible.
6. Which data augmentations do you use, and when can augmentation hurt?
Answer: Common augmentations are geometric (flips, rotations, scale and crop, perspective), photometric (brightness, contrast, colour jitter, blur, noise, JPEG compression) and mixing methods (Mixup, CutMix, Mosaic for detection). The goal is to simulate variation the model will meet in production that the training set under-represents.
Augmentation hurts when it breaks the label. Horizontal flips are wrong for text, for left/right body parts in medical images and for traffic signs with direction. Colour jitter is wrong when colour is the defect (a discoloured weld, a ripeness grade). Heavy rotation is wrong when objects always appear upright on a conveyor. For detection and segmentation, every geometric transform must also be applied to boxes, masks and keypoints.
Real-world example: Consider a dairy plant inspecting seal colour on cup lids. A team copied a generic augmentation recipe with strong hue jitter, and the model learned to ignore exactly the colour shift that indicates a bad seal. Removing hue jitter and adding realistic glare augmentation fixed it.
7. How do classification, detection, segmentation and keypoint estimation differ?
Answer: They differ in output granularity and labelling cost.
| Task | Output | Typical label | Example use |
|---|---|---|---|
| Classification | One or more labels per image | Image-level tag | Is this X-ray normal? Which product category? |
| Object detection | Boxes + class + score | Bounding boxes | Count vehicles, find helmets on workers |
| Semantic segmentation | Class per pixel | Pixel masks | Road vs footpath, crop vs weed |
| Instance segmentation | Mask per object | Per-object masks | Separate touching tablets on a tray |
| Keypoints / pose | Coordinates of landmarks | Points per object | Worker posture, gauge needle position |
Pick the cheapest formulation that answers the business question. If the plant only needs "reject or pass", image classification or anomaly detection may be enough; if it needs to measure defect size, you need segmentation.
8. What is transfer learning in vision, and when do you freeze vs fine-tune?
Answer: Transfer learning starts from a backbone pretrained on a large, diverse dataset and adapts it to your task. With little data and a domain close to natural photos, freeze the backbone and train only a new head, or use frozen features with a linear probe. With more data or a distant domain (thermal, microscopy, satellite, X-ray), fine-tune more layers, often with a lower learning rate for the backbone than the head, and unfreeze progressively.
9. How should you resize and preprocess images with different aspect ratios?
Answer: Options are direct resize (distorts shape), centre crop (may cut off the object), or letterboxing (resize keeping aspect ratio, pad the rest). Detection models commonly use letterboxing so geometry is preserved, and boxes must be mapped back through the same scale and offset at inference time. For high-resolution inputs with small targets (PCB inspection, satellite imagery), tiling the image into overlapping crops often beats shrinking it.
The rule that matters most: preprocessing at inference must match training exactly, including interpolation method, colour order and normalisation. Put preprocessing in shared, versioned code rather than reimplementing it in the serving layer.
10. What is IoU and why does it appear everywhere in detection?
Answer: Intersection over Union is the area of overlap between two regions divided by the area of their union. It ranges from 0 (no overlap) to 1 (identical). Detection uses it to decide whether a predicted box matches a ground-truth box (for example IoU at or above 0.5 counts as a hit), inside non-maximum suppression to decide which boxes are duplicates, in anchor or label assignment during training, and as the basis of box-regression losses such as GIoU, DIoU and CIoU, which stay informative even when boxes do not overlap. Segmentation uses the same idea per class as mean IoU.
Interview tip: Write the code for box IoU from corner coordinates, including the clamp to zero when boxes do not intersect. It is a frequent live-coding warm-up.
Core vision tasks: detection, segmentation, OCR, tracking
11. Compare two-stage and one-stage object detectors.
Answer: Two-stage detectors (the R-CNN family, such as Faster R-CNN) first propose candidate regions, then classify and refine each proposal. They tend to be strong on accuracy and small objects but have more moving parts and higher latency. One-stage detectors (SSD, RetinaNet, the YOLO family, FCOS) predict classes and boxes densely in a single pass over the feature map, which is faster and simpler to deploy, at the cost of handling a huge imbalance between background and object locations (addressed with focal loss and better label assignment).
In 2026 the gap is narrow for many workloads, and transformer-based detectors add a third family (Q28). The practical choice usually comes from the latency budget, target hardware support and the tooling your team can maintain, not from leaderboard positions.
12. What are anchors, and why have many detectors moved to anchor-free designs?
Answer: Anchors are predefined reference boxes of several sizes and aspect ratios placed at each feature-map location; the network predicts offsets from each anchor and whether it contains an object. They make regression easier but add hyperparameters (sizes, ratios, matching thresholds) that must be tuned to the dataset, and they produce many redundant candidates.
Anchor-free detectors predict objects directly from points: for example, the distance from a location to the four box edges (FCOS style) or object centres plus size (CenterNet style). They are simpler to configure and generalise better to unusual shapes such as long thin cracks or cables. If an anchor-based model misses very elongated objects, check whether any anchor ever matched them during training.
13. How does non-maximum suppression work, and what are its weaknesses?
Answer: A detector outputs many overlapping boxes for the same object. NMS sorts boxes by confidence, keeps the highest, removes all others of the same class whose IoU with it exceeds a threshold, and repeats. It is usually run per class (class-aware) so a person and a bag overlapping are not merged.
Weaknesses: in crowded scenes (people in a queue, stacked boxes in a warehouse) genuine neighbours get suppressed; the IoU threshold is a hand-tuned trade-off; and NMS is sequential, which can be awkward to accelerate on some edge hardware. Alternatives include Soft-NMS (decays scores instead of deleting), weighted box fusion for ensembles, and set-prediction detectors such as DETR that are trained to output one box per object and need no NMS.
14. Explain mAP for object detection.
Answer: For each class, sort all predictions by confidence and mark each as a true positive if it matches an unmatched ground-truth box at the chosen IoU threshold, otherwise a false positive. Sweep the confidence threshold to trace a precision-recall curve; Average Precision (AP) is the area under it (with interpolation details that differ between benchmarks). mAP is the mean of AP over classes. COCO-style evaluation also averages over IoU thresholds from 0.5 to 0.95, which rewards tight localisation, and reports AP for small, medium and large objects separately.
What interviewers want you to add: mAP summarises ranking quality across all thresholds, but production runs at one threshold. You still need precision and recall at the operating point, per-class breakdowns and cost-weighted errors. A model with higher mAP can be worse for the business if it gains on easy classes and loses on the one rare defect that matters.
15. How do you improve detection of small objects?
Answer: Small objects lose information through downsampling and get few positive training samples. Levers, roughly in order of payoff:
- Increase input resolution or tile large images into overlapping crops (slicing-aided inference), then merge detections.
- Use feature pyramids so high-resolution feature maps contribute to predictions (Q22).
- Tune label assignment or anchors so small objects receive enough positive matches.
- Augment with copy-paste of small objects and scale jitter; avoid crops that cut small objects in half.
- Check annotation quality: small boxes are where labellers are least consistent, and a few pixels of box error is a large IoU penalty.
Also check the camera: sometimes the cheapest fix is a better lens or moving the camera closer, which no model change can match.
16. What is the difference between semantic, instance and panoptic segmentation?
Answer: Semantic segmentation assigns a class to every pixel but does not separate individual objects; three touching cars become one "car" region. Instance segmentation produces a separate mask for each countable object (Mask R-CNN is the classic design: a detector plus a mask head). Panoptic segmentation unifies both: every pixel gets a class, and pixels of countable "things" also get an instance ID, while amorphous "stuff" such as sky or road gets only a class. Panoptic quality (PQ) combines segmentation quality with recognition quality.
17. Which losses and metrics do you use for segmentation?
Answer: Pixel-wise cross-entropy is the default loss. When the foreground is tiny (a hairline crack covering a minute fraction of pixels), cross-entropy is dominated by background, so teams add Dice loss or focal loss, or combine them. Boundary-aware losses help when edge accuracy matters, for example measuring defect length.
Metrics: mean IoU and Dice per class are standard; report them per class rather than only averaged. For inspection, pixel metrics are often less relevant than object-level ones (did we find each defect, and is its measured size within tolerance?). For thin structures, boundary F-score or skeleton-based metrics reflect quality better than area overlap.
18. How does keypoint or pose estimation work?
Answer: Most models predict a heatmap per keypoint, where the peak marks the likely location, sometimes with an offset map for sub-pixel accuracy. Top-down approaches detect each person or object first, then estimate keypoints inside each crop: accurate, but cost grows with the number of objects. Bottom-up approaches detect all keypoints in the image and group them into individuals (for example with part affinity fields): cost is more constant in crowded scenes. Evaluation commonly uses Object Keypoint Similarity (OKS), which scales error tolerance by object size and per-keypoint difficulty.
Real-world example: Keypoints are not only for humans. Reading an analogue pressure gauge can be framed as detecting the needle tip, the needle pivot and the scale end points, then computing the angle, which is far more robust than classifying the reading directly.
19. Walk through a modern OCR pipeline.
Answer: A classic pipeline has four stages: (1) preprocessing such as deskewing, perspective correction for phone photos, denoising and binarisation for poor scans; (2) text detection, finding word or line regions, often with segmentation-style detectors that handle curved or rotated text; (3) text recognition, turning each crop into characters, commonly a CNN or transformer encoder with a CTC or attention-based decoder; (4) layout and post-processing, ordering lines into reading order, detecting tables and key-value pairs, and applying lexicons or validation rules (dates, PAN or GSTIN formats, totals that add up).
End-to-end models and vision LLMs can now read documents directly (Q32), but the staged pipeline remains easier to debug, cheaper at volume and gives character-level confidence and coordinates. For how parsed output feeds retrieval, see document parsing for RAG.
Interview tip: Distinguish character error rate (CER) and word error rate (WER) from field-level accuracy. Business users care whether the invoice total is right, not the average character accuracy across the page.
20. How does multi-object tracking work in video?
Answer: Most production trackers use tracking-by-detection: run a detector on each frame, predict where existing tracks will move (often with a Kalman filter), then associate new detections to tracks using a cost built from IoU, motion and optionally appearance embeddings (a re-identification model), solved with the Hungarian algorithm or greedy matching. Unmatched detections start tentative tracks; tracks without matches for some frames are terminated. SORT, DeepSORT and ByteTrack are well-known designs; ByteTrack's idea of also associating low-confidence detections helps with occlusion.
Metrics include MOTA, IDF1 and HOTA, plus the raw count of ID switches. In business terms, ID switches are what break people-counting, dwell-time and line-crossing analytics.
21. How do you handle class imbalance in vision?
Answer: Imbalance appears at two levels. Between classes (a rare defect type, a rare vehicle category) and between foreground and background inside detection. Tools:
- Data: collect more of the rare class deliberately, oversample it or use repeat-factor sampling, copy-paste rare objects into backgrounds, or generate synthetic examples (Q36).
- Loss: class-weighted cross-entropy, focal loss (down-weights easy negatives), or logit adjustment for long-tailed distributions.
- Decision: per-class thresholds chosen on validation data against the actual cost of each error type.
- Formulation: when defects are extremely rare and varied, train an anomaly detector on good parts only, and use a supervised model for known defect types.
Always evaluate per class. An overall accuracy number on an imbalanced inspection set says almost nothing.
22. What is a feature pyramid network and why do detectors use one?
Answer: A CNN backbone produces feature maps at decreasing resolution: early maps are high resolution but semantically weak; late maps are semantically strong but coarse. A Feature Pyramid Network adds a top-down pathway that upsamples deep features and merges them with lateral connections from earlier layers, so every scale has strong semantics. Detectors then predict small objects from high-resolution levels and large objects from coarse levels. The trade-off is extra compute at high-resolution levels, which matters on edge devices.
23. How do you write annotation guidelines and measure label quality?
Answer: Labels are the specification of your model, so treat the guideline like a contract. It should define each class with positive and negative visual examples, say how to handle occlusion, truncation, reflections and ambiguous cases, specify box tightness or mask boundary rules, and include an "uncertain" option rather than forcing guesses.
Measure quality with inter-annotator agreement on a shared overlap set (IoU between annotators' boxes, Cohen's kappa for classes), a gold set seeded into labelling queues, and periodic expert review of disagreements.
Interview tip: Say that you would review disagreements with the domain expert before training anything. In inspection, two senior quality engineers often disagree on borderline defects, and that disagreement sets the ceiling for any model.
24. How do you evaluate an image classifier beyond accuracy?
Answer: Look at the confusion matrix to see which classes are confused, per-class precision and recall, and precision-recall curves for each class that matters. For multi-label problems, use per-label metrics and mean AP. Check calibration (do 0.9-confidence predictions turn out right about nine times in ten?) with reliability diagrams or expected calibration error; temperature scaling on a held-out set often fixes overconfidence. Finally, slice results by acquisition conditions: camera, site, shift, lighting, device model. Slices are where production surprises hide.
Architectures and foundation models
25. Describe the main CNN families conceptually.
Answer: Interviewers want the idea each family introduced, not layer counts:
| Family | Key idea | Why it matters |
|---|---|---|
| VGG | Deep stacks of small 3Γ3 convolutions | Showed depth with simple blocks works; heavy compute |
| Inception | Parallel filters of different sizes per block | Multi-scale features at controlled cost |
| ResNet | Residual (skip) connections | Made very deep networks trainable; still a common baseline |
| DenseNet | Each layer receives all earlier feature maps | Feature reuse, parameter efficiency |
| MobileNet / ShuffleNet | Depthwise separable convolutions, cheap channel mixing | Designed for phones and edge devices |
| EfficientNet | Compound scaling of depth, width and resolution | Principled way to scale a model family |
| ConvNeXt | CNN modernised with transformer-era design choices | Shows CNNs remain competitive with good training recipes |
26. How does a vision transformer (ViT) work?
Answer: A ViT splits the image into fixed-size patches (for example 16Γ16 pixels), flattens and linearly projects each patch into an embedding, adds position embeddings so the model knows where each patch came from, and feeds the sequence through standard transformer encoder layers. A special classification token, or a pooled output, feeds the final head. Self-attention lets every patch attend to every other patch from the first layer, so global context is available immediately rather than built up through many convolutions.
Cost grows quadratically with the number of patches, so high-resolution inputs get expensive. Hierarchical designs such as Swin Transformer use local windowed attention and patch merging to produce multi-scale feature maps, which makes transformers practical backbones for detection and segmentation. For attention mechanics, the deep learning guide goes deeper.
27. ViT or CNN: how do you choose?
Answer: CNNs carry strong inductive biases (locality, translation equivariance), so they learn well from smaller datasets and run efficiently on most edge accelerators. ViTs have weaker built-in biases, so they need large-scale pretraining to shine, but then transfer very well, handle global context naturally and are the backbone of most modern vision-language and foundation models.
Practical decision factors: how much labelled data and what pretrained checkpoints exist for your domain; target hardware (some NPUs and older accelerators handle convolutions far better than attention); latency at your input resolution; and whether you want to plug into a vision-language ecosystem. Hybrid models (convolutional stem plus transformer blocks) are common compromises. The honest interview answer is "benchmark both on my data and my hardware".
28. How do transformer-based detectors such as DETR differ from classic detectors?
Answer: DETR treats detection as set prediction. A backbone extracts features, a transformer encoder-decoder processes them, and a fixed number of learned "object queries" each output one box and class (or "no object"). During training, the Hungarian algorithm finds the one-to-one matching between predictions and ground truth, and the loss is computed on that matching. Because each object is matched to exactly one query, there are no anchors and no NMS.
The original design converged slowly and struggled with small objects; later variants (deformable attention, denoising training, real-time DETR-style models) fixed much of that. Interviewers often ask why removing NMS matters: it simplifies deployment, removes a hand-tuned threshold and behaves better in crowded scenes.
29. What is CLIP-style contrastive vision-language training, and what does zero-shot classification mean?
Answer: A CLIP-style model has an image encoder and a text encoder trained on large collections of image-caption pairs. The contrastive objective pulls the embeddings of matching image-text pairs together and pushes non-matching pairs apart within a batch. The result is a shared embedding space where an image and a description of it land close together.
Zero-shot classification follows directly: write a text prompt per class ("a photo of a forklift", "a photo of a pallet jack"), embed them, and assign the image to the closest text embedding, with no task-specific training. The same embeddings power text-to-image search over a photo archive and near-duplicate detection.
Limits: zero-shot accuracy drops on specialised domains (defect types, medical images, regional product packaging) that are rare in web data, results depend on prompt wording, and the models inherit social biases from web data. Linear probes or light fine-tuning on a few labelled examples usually help. See embeddings explained for the retrieval side.
30. What is open-vocabulary detection, and when is it useful?
Answer: A classic detector only knows the fixed classes it was trained on. Open-vocabulary detectors combine a detection architecture with vision-language alignment, so you can ask for objects by text prompt ("hard hat", "spilled liquid") at inference time, including categories not in the detector's labelled training set. Grounding models go further and localise the object referred to by a phrase.
They are excellent for prototyping, for long-tail categories and for auto-labelling: run an open-vocabulary detector to propose boxes, have humans correct them, then train a small closed-set detector that is cheaper and more predictable in production. Treat their raw outputs as noisy for safety-critical decisions and check current model licences before commercial use.
31. What are promptable segmentation foundation models such as SAM, and how do teams use them?
Answer: Segment Anything (SAM) and similar models are trained on very large mask datasets to segment whatever a prompt points at: a click, a box or (in some variants and pipelines) a text phrase. They have a heavy image encoder that runs once per image and a light prompt decoder that responds quickly, which makes interactive use practical.
The class-agnostic nature is the key point: SAM tells you where an object's boundary is, not what it is. Common enterprise uses:
- Labelling acceleration: annotators click once instead of drawing polygons, then correct.
- Pipelines: a detector or open-vocabulary model provides boxes, SAM refines them into masks.
- Measurement: segmenting a part to measure area or dimensions after a detector has found it.
For edge inference, the full encoder is usually too heavy; lightweight distilled variants exist, or you use SAM offline to build training masks for a small task-specific model (see model distillation explained).
32. When would you use a vision LLM instead of a classic OCR or detection pipeline?
Answer: Vision-capable LLMs accept images alongside text and can describe scenes, answer questions about charts, read handwriting, extract fields from documents into structured JSON and reason across a page. They shine when inputs are highly varied (hundreds of invoice layouts), when the task needs reasoning ("is this claim photo consistent with the description?") and when you need a working system before labelled data exists.
Prefer a classic pipeline when you need high throughput at low cost, strict latency, pixel-accurate coordinates, on-device execution, deterministic behaviour or easy auditability. Vision LLMs can hallucinate plausible field values, so production designs add schema validation, cross-checks (line items sum to the total), confidence routing and human review for low-confidence cases. A common hybrid: OCR provides text with coordinates, the LLM interprets layout and meaning, and rules validate. The enterprise picture is in multimodal AI in the enterprise.
33. What is self-supervised pretraining in vision?
Answer: Self-supervised learning builds representations from unlabelled images by inventing a task from the data itself. Contrastive and self-distillation methods (SimCLR, MoCo, DINO-style) make embeddings of two augmented views of the same image agree. Masked image modelling (MAE-style) hides most patches and trains the model to reconstruct them. The resulting backbones give strong features for downstream tasks with few labels.
This matters in industry because unlabelled images are cheap (a factory camera produces thousands per hour) while expert labels are expensive. Pretraining or continuing pretraining on in-domain unlabelled images before fine-tuning on a small labelled set is a strong strategy for unusual domains.
Data, labelling and evaluation
34. What is domain shift in vision, and how do you detect it?
Answer: Domain shift means production images differ from training images in ways that hurt the model. Typical causes: a new camera or lens, changed lighting or a seasonal sun angle, a new product variant, dust on the lens, different compression settings on a new video recorder, or a new site with different backgrounds. Label shift (the defect mix changes) is a related problem.
Detection approaches: monitor input statistics (brightness, contrast, sharpness, colour histograms), embedding-distribution distance between recent frames and a training reference, prediction-distribution changes (sudden rise in "no detection" or one class), confidence distributions, and, most reliably, a small stream of human-reviewed samples that gives real accuracy over time. Mitigations range from augmentation and calibration of the camera setup to fine-tuning on a few hundred labelled images from the new domain.
35. How do you split vision data to avoid leakage?
Answer: Random image-level splits leak badly in vision. Consecutive video frames, multiple photos of the same part, the same patient's scans or burst shots are near duplicates; if they land in both train and test, test scores look excellent and mean little. Split by group instead: by video, by part serial number, by patient, by store, by production day or by camera. Deduplicate with perceptual hashes or embedding similarity before splitting. For deployment to new sites, hold out entire sites to estimate how the model generalises to locations it has never seen.
Interview tip: If you mention one thing about CV evaluation, mention grouped splits. It separates practitioners from tutorial-followers.
36. When does synthetic data help in computer vision?
Answer: Synthetic images come from 3D rendering (CAD models of parts, simulated warehouses), compositing (pasting objects onto real backgrounds), procedural defect generation, or generative image models. They help most when real examples are rare or dangerous to collect (rare defects, accidents), when labels must be pixel-perfect (depth, segmentation), and when you need controlled variation (every lighting angle).
The risk is the sim-to-real gap: the model learns rendering artefacts. Mitigations are domain randomisation (vary textures, lighting and backgrounds widely so reality looks like one more variation), mixing synthetic with real data, and always validating on real images only. Synthetic data should never be part of the test set you report to stakeholders. For testing uses, see synthetic data for AI testing.
37. How would you set up an active learning loop for labelling?
Answer: Train an initial model on a small labelled set, run it over the unlabelled pool, and send the most informative images to annotators: low-confidence predictions, disagreement between ensemble members or augmented views, images far from the training distribution in embedding space, and a random sample to stay unbiased. Pre-label with the model so annotators correct rather than draw from scratch. Retrain, re-score and repeat. In production, route images the model is unsure about to human review and feed those decisions back as labels, which is a human-in-the-loop design.
Watch for two biases: pure uncertainty sampling oversamples confusing junk (blurred frames), and pre-labelling makes annotators accept model mistakes. Audit a sample of accepted pre-labels.
38. How do you choose the operating threshold and report results to the business?
Answer: Start with the cost of each error. In defect inspection, a missed defect (escape) may reach a customer; a false reject costs a re-inspection or scrapped good part. Estimate those costs with the quality team, then pick the threshold on a validation set that minimises expected cost or meets a hard constraint such as "recall on critical defects at least the agreed target". Validate the chosen threshold on a separate test set.
Report in business units: escapes per thousand parts, false rejects per shift, review workload per day, and how these vary by product variant and line. Show the precision-recall trade-off as a decision the business makes, not a number the ML team picks alone.
Deployment, edge and video pipelines
39. How do you deploy a vision model on edge or on-device hardware?
Answer: Export the trained model to a portable format (commonly ONNX), then compile or convert it for the target runtime: NVIDIA TensorRT for Jetson-class devices and GPUs, OpenVINO for Intel CPUs and accelerators, LiteRT (previously TensorFlow Lite) for Android and microcontroller-class devices, Core ML for Apple devices, or vendor SDKs for specific NPUs. Check that every operator is supported on the accelerator; unsupported ops fall back to the CPU and can wreck latency.
Then profile the whole pipeline on the real device, not just the model: camera capture, decode, resize, inference, post-processing and output. Plan for thermal throttling, power limits, offline operation and secure over-the-air model updates with rollback. The broader trade-offs are in on-device and edge AI.
40. Explain quantisation for vision models. PTQ or QAT?
Answer: Quantisation represents weights and activations with fewer bits, typically INT8 instead of FP32 or FP16, which shrinks the model and speeds inference on hardware with integer units. Post-training quantisation (PTQ) converts a trained model using a calibration set of representative images to choose value ranges; it is quick and often enough for CNN classifiers and detectors. Quantisation-aware training (QAT) simulates quantisation during fine-tuning so the model learns to tolerate it; use it when PTQ loses too much accuracy.
Practical points: the calibration set must reflect production conditions (night frames, glare, every product variant), per-channel weight quantisation usually beats per-tensor, some layers (the first and last, attention softmax, detection heads) may need to stay at higher precision, and you must re-check per-class metrics after quantising, not only the average (see Q54).
41. How do you budget latency in a real-time video pipeline?
Answer: Start from the requirement: a robot arm that must reject a part on a conveyor moving past a camera has a hard deadline; a dashboard counting footfall can tolerate seconds. Then break end-to-end latency into stages and measure each:
camera -> decode -> preprocess -> infer -> postprocess
-> track/logic -> action (PLC, alert, DB write)
Common optimisations: hardware video decode, preprocessing on the GPU, batching frames across cameras, lower input resolution where accuracy allows, running the detector every Nth frame and letting the tracker interpolate between, region-of-interest cropping, and pipelining stages so capture, inference and post-processing overlap.
42. Design a video analytics system for a few hundred CCTV cameras.
Answer: Consider a retailer that wants queue-length alerts and occupancy counts across many stores. A typical architecture:
[Cameras] --RTSP--> [Edge box per store]
| decode (HW)
| detect + track
| count / zone logic
v
events + metrics (no raw video)
|
v
[Cloud] queue -> stream proc -> dashboards
| alerts
v
sampled frames (blurred) -> review/labels
Key decisions: process at the edge so raw video never leaves the store (bandwidth and privacy); send only events and aggregates; keep a model registry and roll out new versions to a few stores first; monitor per-camera health (offline, frozen frame, lens obstructed, sudden drop in detections); and collect a small, privacy-protected sample for ongoing evaluation. Camera placement and angle determine more of the accuracy than the model choice, so involve the store operations team early.
43. How do you monitor a CV model in production?
Answer: Monitor four layers. System: latency percentiles, throughput, dropped frames, device temperature, GPU or NPU utilisation. Input: image quality statistics, blur and brightness, embedding drift against a reference set, camera uptime. Output: detection counts per class per camera, confidence distributions, rate of "no object" outputs, reject rate on the line. Outcome: accuracy on a regularly labelled sample, downstream confirmations (rejected parts confirmed defective by a human inspector), and customer complaints or escapes.
Alert on sudden changes per camera rather than global averages, because one dirty lens hides inside a fleet average. Pipeline and registry practices overlap with MLOps interview questions.
44. What do you design for privacy when a system sees faces, CCTV footage or identity documents in India?
Answer: Face images, CCTV footage of identifiable people and scanned identity documents are personal data, and India's Digital Personal Data Protection Act, 2023 and its Rules apply to how they are processed (check current obligations and timelines with your legal team). Engineering controls that support compliance:
- Purpose limitation and minimisation: if the use case is counting, never store identities; process at the edge and keep only aggregates.
- Anonymise early: blur or mask faces and number plates before footage is stored or sent for labelling.
- Retention: short, defined retention for raw video with automatic deletion; separate access for incident review.
- Notice and consent: clear notices where cameras operate; explicit consent for face recognition use cases such as attendance, plus a non-biometric alternative.
- Access and audit: role-based access to footage, logging of who viewed what, encryption at rest and in transit.
- Labelling vendors: contracts, masked data and restricted environments for third-party annotators.
Also test face-related models for performance differences across skin tones, ages and genders before deployment. The DPDP Act for AI applications guide covers the legal side in more depth.
Real-world scenarios
45. A defect detection model works well on Line 1 but fails on a newly commissioned Line 2. What do you do?
Answer: Treat it as domain shift until proven otherwise, and find out what is physically different before touching the model. The fastest wins usually come from aligning the imaging setup, then from a small amount of in-domain fine-tuning.
What I would check:
- Side-by-side images from both lines: camera model, lens, distance, angle, exposure, lighting type and position, background, conveyor speed and motion blur.
- Whether Line 2 produces product variants or materials that Line 1 never saw.
- The preprocessing path: resolution, colour order, bit depth and crop region configured for the new camera.
- Error patterns on a labelled sample from Line 2: are failures false positives on new textures, or missed defects?
- Embedding-space comparison to quantify how far Line 2 images sit from training data.
- After fixing setup differences, fine-tune with a few hundred labelled Line 2 images, keeping Line 1 data in the mix to avoid forgetting.
Production consideration: Write an imaging specification (camera, lens, lighting, mounting, calibration target) that every new line must meet, and run a short qualification phase in which the model runs in shadow mode alongside human inspection before it is allowed to reject parts. The manufacturing context is covered in AI in manufacturing.
46. Operators complain the defect detector raises too many false positives and are starting to ignore it. How do you fix it?
Answer: Alert fatigue destroys the value of an inspection system faster than missed defects, so this is urgent. Start by categorising the false positives rather than raising the threshold blindly, because a higher threshold also lets real defects through.
What I would check:
- Collect several hundred recent false positives and cluster them: dust, water droplets, reflections, label edges, acceptable cosmetic variation, new product variant.
- Compare against the label guideline: are some "false positives" actually defects the guideline never defined, or acceptable marks the model was never shown as negatives?
- Add hard negatives from each cluster to training; consider an explicit "acceptable variation" class.
- Fix physical causes: air knives for dust and water, polarising filters or diffuse lighting for glare.
- Recalibrate per-class thresholds against agreed escape and false-reject costs, and consider requiring the defect to appear in consecutive frames or multiple camera views.
- Add a second-stage classifier on detected crops if one class of confusions dominates.
Production consideration: Give operators a one-tap "not a defect" button whose feedback flows into the labelling queue, and report false rejects per shift on the line dashboard so improvement is visible to the people who lost trust.
47. You must extract fields from scanned documents in Telugu, Hindi and English, sometimes mixed on one page. How do you approach OCR?
Answer: Indian-language documents add script diversity, conjunct characters and diacritics, mixed scripts within a line, variable scan quality and handwriting in filled forms. Plan a pipeline that detects script per region, uses recognisers trained on Indic scripts, and validates extracted fields, with human review for low-confidence cases.
What I would check:
- Build a representative test set from real documents per language, form type and scan source, with field-level ground truth. Without it, every tool comparison is guesswork.
- Evaluate candidate OCR engines and vision LLMs on that set for each script, measuring CER and field accuracy separately for printed and handwritten text.
- Add script identification at line or word level so each region goes to the right recogniser.
- Check Unicode handling end to end: normalisation, correct rendering of conjuncts and matras, and fonts in the review UI. Many "OCR errors" are encoding bugs.
- Use field-specific validation: numerals (Indian languages have their own digit forms), dates, names transliterated against a master record, amounts in words versus figures.
- If accuracy is short on a script, fine-tune the recogniser with real crops plus synthetic text lines rendered in multiple fonts and degradations.
Production consideration: Store both the original script and a transliterated form, keep coordinates for each extracted field so reviewers can see the source, and route by confidence. The downstream document workflow is similar to the enterprise document intelligence project.
48. Your detector reported excellent validation mAP, but accuracy collapsed in the first week of production. What went wrong?
Answer: The most common cause is that validation was not representative: leakage from near-duplicate frames, or a validation set drawn from the same cameras and days as training. The second is a pipeline mismatch between training and serving.
What I would check:
- How the split was made: by image or by video, day, camera and site? Search for near duplicates across splits with embeddings.
- Serving preprocessing: colour order, resize method, normalisation, letterbox mapping of boxes back to the original frame.
- The threshold used in production versus the one assumed during validation; mAP hides operating-point behaviour.
- Production image samples: compression artefacts from the video recorder, lower resolution substreams, night infrared mode.
- Class mix in production versus validation.
Production consideration: Rebuild the test set with grouped splits and a held-out site, add a golden-image regression test that runs the exact serving container on known images, and release future models through shadow mode first.
49. A CCTV-based safety model (helmet and vest detection) fails at night and during monsoon rain. How do you improve it?
Answer: Night and rain change the input physics: cameras switch to infrared (grayscale, different textures), rain adds streaks and droplets on the lens, and low light means noise and motion blur. The training set probably under-represents these conditions.
What I would check:
- Slice metrics by time of day, IR mode and weather to confirm where errors concentrate.
- Collect and label night and rain footage specifically; one monsoon season of samples is worth more than heavy augmentation.
- Add augmentations that resemble real conditions: grayscale conversion, sensor noise, motion blur, synthetic rain streaks, glare from floodlights.
- Consider a separate model or fine-tuned head for IR mode, selected by camera mode metadata.
- Work with facilities on lighting and lens hoods; a camera with water on the lens needs maintenance, not a new model.
Production consideration: Monitor per-camera confidence and detection rates by hour so you catch degradation each season, and avoid automatic disciplinary actions from low-light detections; route them for human review.
50. The model meets accuracy targets but runs too slowly on the chosen edge device. How do you cut latency without losing accuracy where it matters?
Answer: Profile first; the model is often not the only bottleneck. Then apply optimisations in order of effort and risk.
What I would check:
- Per-stage timing on the device: decode, preprocessing (CPU resizing is a frequent culprit), inference, NMS and post-processing, I/O.
- Whether all layers run on the accelerator or some fall back to CPU because of unsupported operators.
- Compile with the vendor runtime and FP16; then try INT8 PTQ with a representative calibration set.
- Reduce input resolution or crop to the region of interest; test the accuracy impact per class.
- Switch to a smaller backbone or distil the current model into one (teacher-student training).
- Run inference on every Nth frame with tracking between, if the use case allows.
Production consideration: Test under sustained load at realistic ambient temperature; edge devices in a hot factory throttle, and a benchmark that passes in an air-conditioned lab may fail on the line.
51. In a crowded retail store, people-counting is off because the tracker keeps switching IDs. How do you fix it?
Answer: ID switches in crowds come from occlusion, missed detections and weak association cues. Fix detection first, then association, then the camera geometry.
What I would check:
- Detector recall on occluded people; crowded scenes may need training data with heavy overlap and a higher NMS IoU threshold or an NMS-free detector.
- Association parameters: IoU versus appearance weight, track buffer length, and whether low-confidence detections are used for association (ByteTrack-style).
- Add or improve a re-identification embedding for appearance matching; validate it on footage from these stores.
- Frame rate: low frame rates make motion prediction unreliable.
- Camera angle: an overhead or steeply angled camera at the entrance reduces occlusion dramatically.
Production consideration: For counting, line-crossing logic with hysteresis is more robust than relying on perfect long tracks. Appearance embeddings of shoppers are personal data, so keep them on the edge device, in memory only, and never link them to identities.
52. A retailer adds new products every week and wants shelf images recognised at SKU level. Retraining a classifier weekly is not sustainable. What do you design?
Answer: Turn it into detection plus retrieval. A class-agnostic product detector finds each item on the shelf; an embedding model maps each crop into a vector; recognition is nearest-neighbour search against a gallery of reference images per SKU. Adding a product means adding its reference images to the gallery, with no retraining.
What I would check:
- Reference image quality: multiple angles per SKU, ideally including shelf-like photos, not just studio packshots.
- Fine-grained confusions: pack sizes and flavour variants that look almost identical. Add text cues via OCR on the label, or price-tag reading, to separate them.
- An "unknown product" threshold on similarity so new or unregistered items are flagged rather than misnamed.
- Periodic fine-tuning of the embedding model with metric learning on hard pairs, which improves all SKUs at once.
Production consideration: Store embeddings in a vector index with SKU metadata and version the gallery alongside the model; a planogram compliance report is only as good as the gallery it was produced from. The indexing side overlaps with vector database interview questions.
53. An insurer wants to assess vehicle damage from customer-uploaded photos. What risks and design choices would you raise?
Answer: Consider an insurer that wants faster motor claims. The vision task (localise damaged parts, classify damage type and severity) is only part of the system; photo quality, fraud and fairness to customers matter as much.
What I would check:
- Capture guidance in the app: required angles, distance, lighting, and real-time quality checks for blur and framing before upload.
- Part segmentation plus damage segmentation, so severity maps to a repair-or-replace decision per part.
- Fraud signals: reused photos (perceptual hash against previous claims), EXIF and location inconsistencies, edited or AI-generated images, photos of a screen.
- Performance across vehicle colours, models common in Indian markets, and night or rain photos.
- Where a vision LLM is used to summarise damage, cross-check its statements against the segmentation output.
Production consideration: Use the model to fast-track simple, low-value claims and to assist surveyors, with a human making adverse decisions. Log model outputs with each claim for audit, and explain to customers what evidence the decision used.
54. After INT8 quantisation, overall accuracy barely changed, but recall on one small defect class dropped sharply. What happened and how do you fix it?
Answer: Aggregate metrics hid a class-specific regression. Small, low-contrast defects often depend on subtle activation differences that coarse quantisation ranges wipe out, especially if the calibration set contained few examples of that defect.
What I would check:
- Calibration set composition: include enough images with the rare defect and the conditions it appears in.
- Calibration method: try different range estimation (percentile or entropy-based instead of min-max) and per-channel weight quantisation.
- Layer sensitivity analysis: quantise layer by layer to find which layers cause the drop, and keep those at FP16 (mixed precision).
- Re-tune the confidence threshold for that class after quantisation, since score distributions shift.
- If PTQ cannot recover it, run quantisation-aware fine-tuning.
Production consideration: Make per-class metrics on the quantised artefact a release gate. The model that ships is the quantised one, so that is the one you must evaluate.
55. An HR head asks for face-recognition attendance at office entrances. How do you respond as the CV engineer?
Answer: Clarify the actual goal (accurate attendance, stopping proxy punching, contactless entry) before building anything, then present options with their privacy implications. Face templates are personal data of employees, and the design must satisfy the DPDP Act, company policy and employee trust, not just accuracy.
What I would check:
- Whether a less intrusive option meets the goal: badge plus PIN, mobile app check-in with geofencing, or existing access-control logs.
- If face recognition is chosen: explicit, informed consent with a genuine non-biometric alternative for those who decline.
- On-device matching where the terminal stores only templates, not raw photos; encryption; deletion when an employee leaves.
- Liveness detection against photo and screen spoofing.
- Accuracy across skin tones, ages, genders, spectacles, beards, head coverings and masks, tested on a consenting internal sample, plus an easy manual fallback when matching fails. The bias and fairness testing guide covers test design.
- Purpose limitation: attendance data must not be reused for surveillance or performance monitoring without fresh notice and consent.
Production consideration: Document the decision, data flows and retention in a privacy assessment reviewed by legal and the data protection contact, and give employees a clear way to see and withdraw their data. Saying "here is a safer way to meet the goal" is a sign of seniority, not of being difficult.
Want to go beyond interview answers and build vision and multimodal systems that run in production? Cloudsoft's APEX AI, ML, Cloud and Cyber Security program covers machine learning, deep learning, cloud deployment and security in one track, classroom in Ameerpet or live online.
Key takeaways
- Pick the cheapest task formulation (classification, detection, segmentation, anomaly detection, retrieval) that answers the business question.
- Know IoU, NMS, anchors versus anchor-free and mAP well enough to code and critique them; mAP is not your operating point.
- Grouped splits (by video, part, site, day) and per-class, per-slice metrics are the backbone of honest CV evaluation.
- CNNs, ViTs, CLIP-style models, SAM-style segmenters and vision LLMs each have a place; combine them in pipelines rather than treating one as the answer.
- Most production failures are imaging and data problems: cameras, lighting, label guidelines and domain shift.
- Edge deployment is a pipeline problem: profile every stage, quantise with representative calibration data and re-check per-class results.
- Faces, CCTV and identity documents are personal data; design for minimisation, edge processing and consent from day one.
Interview preparation checklist
- Code box IoU, NMS and a simple AP calculation from scratch in Python and NumPy.
- Fine-tune a pretrained classifier and a detector on a small custom dataset; record what augmentations helped and which hurt.
- Build a segmentation demo that uses a promptable model for labelling, then trains a smaller task model.
- Run a CLIP-style zero-shot classifier and compare it with a linear probe on a few labelled examples.
- Export one model to ONNX, run it with an edge runtime, quantise it to INT8 and compare latency and per-class metrics.
- Build an OCR pipeline on a few Indian-language documents and measure field-level accuracy, not just CER.
- Write a short video pipeline: decode, detect, track and count line crossings; measure p95 latency.
- Prepare one project story covering data, labels, splits, metric, threshold choice, deployment and monitoring.
- Be ready to discuss privacy controls for faces and CCTV under the DPDP Act in plain language.
- Revise backpropagation, normalisation and attention basics with the Python for AI interview questions and deep learning guides.
FAQ
What skills are required for a computer vision engineer role?
You need Python, PyTorch or a similar framework, OpenCV, a solid grasp of CNNs and vision transformers, detection and segmentation methods, evaluation metrics, data labelling practices, and increasingly model export, quantisation and edge or cloud deployment.
How should I prepare for a computer vision interview?
Revise image fundamentals and core tasks, practise coding IoU and NMS, build two or three end-to-end projects with honest evaluation, and prepare to explain one deployment in detail, including latency, thresholds and monitoring.
Are computer vision interview questions mostly theory or coding?
Most interviews mix both: a coding round with array or geometry problems such as IoU, a concepts round on architectures and metrics, and a design or scenario round about data, deployment and failures in production.
Do I need to know vision transformers and foundation models for CV interviews in 2026?
Yes, at a conceptual level. Expect questions on how ViTs process patches, how CLIP-style models enable zero-shot tasks, how SAM-style models help labelling, and when a vision LLM is or is not the right tool.
Is classical computer vision with OpenCV still asked?
Often, especially for industrial and edge roles. Thresholding, morphology, contours, camera calibration and colour spaces still solve many inspection problems cheaply and are useful preprocessing around deep models.
Can freshers get computer vision jobs?
Yes, when they show strong fundamentals and projects that go beyond tutorials, such as a custom-labelled dataset, grouped evaluation splits, a quantised model running on a small device, and a clear write-up of what failed.
Which industries hire computer vision engineers in India?
Manufacturing and quality inspection, retail analytics, insurance and banking document processing, healthcare imaging, agriculture, logistics, smart infrastructure, and global capability centres in cities such as Hyderabad and Bengaluru that build vision products.
Is computer vision a good career with the rise of multimodal LLMs?
Multimodal models have widened the field rather than replaced it. Teams still need engineers who understand imaging, data, evaluation, edge deployment and privacy, and who can decide when a specialised model beats a general one.
What projects should I put on my resume for a CV engineer interview?
Choose projects with business framing: a defect detector with threshold analysis, an OCR pipeline for Indian-language documents, a people-counting video pipeline on an edge device, or a document understanding system that combines OCR with a vision LLM.
If you want guided, hands-on preparation across ML, deep learning, cloud and security, explore the APEX program at Cloudsoft. Engineers who want to take vision and multimodal systems into customer environments end to end can also look at the AI Forward Deployed Engineer course (FDE PRO), which adds enterprise integration, security and deployment with placement support until you're placed. Classroom training is in Ameerpet, Hyderabad, or live online; call +91 96660 19191 for a free demo.



