New batches starting this week Β· Limited seats

Kubernetes for AI Interview Questions and Answers 2026 (55 Questions)

55 Kubernetes for AI interview questions with senior-level answers, from GPU device plugins and DRA to KServe, vLLM, Kueue, autoscaling and real GPU incident scenarios.

Kubernetes for AI interview questions 2026: 55 questions on GPU scheduling, autoscaling, KServe and vLLM, Kueue and Ray, model weights and cost
Last updated Β· 40 min read Β· 8,823 words

Kubernetes for AI interview questions test whether you can run GPU-hungry, slow-starting, long-streaming workloads on a platform that was designed for small stateless web services. Interviewers for AI platform, MLOps and LLM-serving roles want to hear how you schedule and share GPUs, autoscale on the right signals, get multi-gigabyte model weights into pods quickly, and keep inference, training and tenants from hurting each other. This guide collects 55 high-value questions, from device plugins and Dynamic Resource Allocation to KServe, vLLM, Kueue and real incident scenarios, with answers written the way a senior platform engineer would explain them.

How to use this guide

  • Freshers and juniors: master the fundamentals (Q1 to Q10). Interviewers check that you know how a pod asks for a GPU, why GPU nodes are tainted and what a startup probe is for.
  • DevOps and cloud engineers moving into AI: focus on GPU scheduling, autoscaling and serving (Q11 to Q30). The usual gap is knowing why CPU-based autoscaling and round-robin load balancing fail for LLMs.
  • Senior and platform roles: expect training orchestration, multi-tenancy, cost and the scenario round (Q31 to Q55), where you are judged on how you diagnose, not on definitions.
  • If plain Kubernetes is rusty, warm up with the general Kubernetes interview questions first.

Contents

Fundamentals

1. When does it make sense to run AI workloads on Kubernetes instead of a managed model endpoint?

Answer: Kubernetes earns its place when you run several long-lived services around the model, self-host open-weight models on GPUs, or the customer already mandates a hardened cluster standard. If the application only calls a hosted model API such as Amazon Bedrock, Azure OpenAI or Gemini, the model never touches your cluster; Kubernetes runs the API, agent orchestrator, ingestion workers and schedulers. Self-hosting adds GPU capacity planning, driver management, autoscaling and weight distribution, so it should be justified by data residency, cost at steady high volume, latency, or the need for a specific fine-tuned model. Managed inference endpoints are simpler when they meet the requirement.

Interview tip: Say "it depends" and then name the deciding factors. The Kubernetes guide for AI applications walks through this split, and self-hosting LLMs covers the build-versus-buy decision.

2. How does Kubernetes know a node has GPUs, and how does a pod request one?

Answer: The kubelet does not discover GPUs by itself. A vendor device plugin (for NVIDIA, the NVIDIA device plugin) runs on each GPU node, registers with the kubelet and advertises GPUs as an extended resource such as nvidia.com/gpu. The scheduler then treats it like any countable resource. A pod asks for GPUs in resources.limits, for example nvidia.com/gpu: 1. When the container starts, the device plugin tells the runtime which device files and libraries to inject. The node still needs a working driver and the NVIDIA container toolkit, which is why most teams install the stack through the GPU Operator (Q4). Newer clusters can also use Dynamic Resource Allocation (Q11).

3. Why can't a pod request half a GPU with the device plugin?

Answer: Extended resources are integers and cannot be overcommitted. You specify them in limits (if you set requests too, they must equal the limits), and a device-plugin GPU is allocated exclusively to one container; GPUs are not shared between containers by default. Fractional use has to come from somewhere else: MIG partitions a supported GPU into hardware-isolated instances that are advertised as separate resources, time-slicing advertises several "replicas" of one physical GPU, and DRA lets drivers describe shareable or partitionable devices. Each offers a different level of isolation, covered in Q13.

4. What does the NVIDIA GPU Operator install, and why use it?

Answer: The GPU Operator automates the software stack a GPU node needs, so GPU nodes can be managed like any other node pool. It typically deploys the GPU driver (as a container, unless the node image already has one), the NVIDIA container toolkit, the device plugin, GPU Feature Discovery (which labels nodes with GPU model, memory and capabilities), the DCGM exporter for metrics, and the MIG manager. It can also deploy NVIDIA's DRA driver as an alternative to the device plugin. The benefit is consistent, versioned driver and plugin rollout across nodes; the risk is that a driver upgrade becomes a cluster-wide change that needs a canary node pool and a rollback plan (see Q49). Some managed offerings preinstall drivers, in which case you enable only the operator components you need.

5. How do you keep ordinary workloads off expensive GPU nodes?

Answer: Put GPUs in a dedicated node pool, taint those nodes (for example nvidia.com/gpu=present:NoSchedule) and give only model-serving and training pods a matching toleration. A toleration only permits scheduling; it does not attract the pod. So also add a node selector or node affinity on a GPU label (pool name, or a GPU Feature Discovery label for the GPU model) so model pods land on the right hardware. Without the taint, a burst of web pods can fill GPU nodes' CPU and memory and leave GPUs stranded; without the affinity, a model pod might land on the wrong GPU type. Many managed node pools can apply the taint for you.

Interview tip: Interviewers like hearing that taints repel and affinity attracts, and that you need both.

6. Which Kubernetes workload types do you use for AI systems?

Answer:

WorkloadKubernetes objectWhy
Online inference server, API, orchestratorDeployment (or a KServe resource)Stateless replicas, rolling updates, HPA
Multi-node inference (one model across nodes)LeaderWorkerSetTreats a group of pods as one replica
Fine-tuning, batch embedding, evaluation runJob, Indexed Job, JobSet, or a Trainer/Ray CRDRuns to completion, retries, parallelism
Nightly re-index, scheduled evalCronJobTime-based trigger
In-cluster vector databaseStatefulSetStable identity and persistent volumes
Node-level agents (exporters, drivers)DaemonSetOne per node

7. Why do model servers need a startup probe, and what should readiness check?

Answer: Loading a large model can take minutes: download or mount weights, load them into GPU memory, warm up kernels and allocate the KV cache. A liveness probe that starts too early kills the pod mid-load, which produces an endless restart loop. A startup probe with a generous failure threshold holds off liveness and readiness until the server reports it is up. Readiness should mean "can serve a request now", typically the server's health endpoint after the model is loaded, so traffic never reaches a loading replica. Liveness should be cheap and detect a truly hung process, not a busy one; an overloaded but healthy server should shed load, not be restarted.

8. How is LLM inference traffic different from typical web traffic on a cluster?

Answer: Requests are long (seconds to minutes), streamed token by token, and vary enormously in cost depending on prompt and output length. Each replica is expensive, so you run few of them and keep them busy. Capacity is bounded by GPU memory for the KV cache, not CPU. This breaks several defaults: idle and request timeouts on ingress and load balancers cut streams, round-robin balancing piles long requests onto one replica, CPU-based autoscaling never fires, and rolling updates drop in-flight generations unless termination grace periods are long enough. The LLM latency guide covers the metrics (time to first token, inter-token latency) that matter here.

9. How should pods authenticate to cloud model APIs and object storage?

Answer: Through workload identity, not static keys in Secrets. A Kubernetes service account is mapped to a cloud identity: IAM roles for service accounts or EKS Pod Identity on EKS, Workload Identity Federation for GKE, and Microsoft Entra Workload ID on AKS. The pod receives short-lived credentials automatically, scoped to exactly what that workload needs, such as read access to one model bucket or invoke permission on specific Bedrock models. Give each workload its own service account so the ingestion worker cannot invoke models and the model server cannot write to the document store. Disable automatic token mounting for pods that never call the Kubernetes API.

10. How do you stop one team from consuming all the GPUs in a shared cluster?

Answer: At minimum, a namespace per team with a ResourceQuota that caps the GPU extended resource (for example requests.nvidia.com/gpu), plus LimitRanges for CPU and memory defaults. Quotas stop runaway usage but are rigid: idle quota in one team cannot help another. For training and batch work, Kueue (Q32) adds queueing, fair sharing and borrowing between teams. For serving, combine quotas with PriorityClasses so production inference outranks experiments. Quotas on DRA-based devices are expressed differently, so check how your version accounts for ResourceClaims.

GPU scheduling and sharing

11. What is Dynamic Resource Allocation (DRA), and what is its status?

Answer: DRA is a Kubernetes API for requesting devices by their attributes rather than as an opaque count. Its core APIs in the resource.k8s.io group graduated to general availability in Kubernetes 1.34 and are enabled by default. A DRA driver publishes ResourceSlices describing the devices on each node; administrators or drivers define DeviceClasses; workloads create ResourceClaims (or a ResourceClaimTemplate for per-pod claims) that select devices with expressions such as GPU model or memory size; and the scheduler allocates a matching device before binding the pod. Claims can be shared between containers or pods, which makes sharing explicit. Several advanced capabilities (such as prioritised alternatives and partitionable devices) have their own, separate feature states, so check them for your Kubernetes version.

Driver ---publishes---> ResourceSlice (devices, attrs)
Admin  ---defines-----> DeviceClass
Pod    ---references--> ResourceClaim (select by attrs)
Scheduler: match claim to slice, allocate, bind pod
Kubelet: driver prepares device for the container

12. Device plugin or DRA: how do you choose today?

Answer: The device plugin is simple, universally supported and fine when workloads ask for "N whole GPUs of this pool". DRA is worth it when you need attribute-based selection (a specific GPU model or memory size without a node pool per type), structured sharing of one device between containers, partitioning, or allocating related devices together. The practical constraints are driver maturity and ecosystem support: your GPU vendor's DRA driver, your managed Kubernetes version, your autoscaler and your queueing system must all understand ResourceClaims. A common path is to keep the device plugin for existing pools and pilot DRA on a new pool. Avoid running both against the same GPUs on one node.

13. Compare MIG, time-slicing and MPS for sharing a GPU.

Answer:

ApproachIsolationGood forWatch out for
MIG (Multi-Instance GPU)Hardware partitions with dedicated memory and compute, fault isolationSeveral small inference models or tenants on one large GPUOnly on supported data-centre GPUs; fixed partition profiles; reconfiguring needs the GPU drained
Time-slicingNone for memory or faults; processes interleaveDev, notebooks, light or bursty workloadsOne pod can exhaust memory and crash neighbours; no assured performance
MPS (Multi-Process Service)Concurrent kernels; limited memory and compute caps, weaker fault isolation than MIGMany small processes that each underuse a GPUShared failure domain; check operator support for your setup

With the GPU Operator's time-slicing, nodes advertise more nvidia.com/gpu units than physical GPUs (optionally renamed to nvidia.com/gpu.shared), so a pod asking for "one GPU" may get a slice. Make that visible in naming so no one mistakes a slice for dedicated capacity. MIG and time-slicing can be combined.

14. How do you target a specific GPU model, such as one with more memory?

Answer: With the device plugin, use node labels: either your own node pool label or the labels GPU Feature Discovery adds (GPU product, memory, count, MIG capability), combined with node affinity. Prefer "required" affinity for hard memory needs (a 70B-class model simply will not fit on a small card) and "preferred" affinity when several types would work. With DRA, put the selection in the ResourceClaim's device selector instead of on the node. Either way, size from first principles: weight memory at your chosen precision, plus KV cache for your target concurrency and context length, plus runtime overhead. GPU basics for AI engineers explains that sizing arithmetic.

15. How do you run a model that needs more than one GPU, or more than one node?

Answer: If it fits on one node, request several GPUs in one pod and let the inference server use tensor parallelism across them; keep all GPUs on the same node for fast interconnect. If it needs several nodes, a single Deployment replica is the wrong abstraction because one "replica" is now a group of pods that must start, fail and scale together. LeaderWorkerSet models exactly that: a leader pod plus workers form one unit, created and restarted as a group. Multi-node serving also needs a fast network between nodes and topology-aware placement, so check whether your cloud offers placement groups or compact placement for GPU nodes. Only go multi-node when quantisation or a larger single node cannot meet the requirement; it multiplies the failure surface.

16. How do PriorityClasses and preemption apply to GPU workloads?

Answer: PriorityClasses let the scheduler evict lower-priority pods to place higher-priority ones when the cluster is full. On a GPU pool a sensible order is production inference, then scheduled production batch, then experiments and notebooks. Two cautions. First, preempting a training job wastes all work since its last checkpoint, so preemptible jobs must checkpoint frequently (Q35). Second, preemption of a single pod in a distributed job can stall the remaining workers while still holding their GPUs, so batch preemption is better handled at the job level by Kueue than pod by pod. Use PodDisruptionBudgets on inference so voluntary disruptions such as node drains never take all replicas at once.

17. Why are AI container images so large, and why does it matter on Kubernetes?

Answer: CUDA libraries, deep-learning frameworks and inference engines can push images to many gigabytes before any weights are added, and baking weights into the image makes it far larger still. Large images slow every new node and every rollout, because pulls happen before the container can start, and they waste node disk. Keep weights out of the application image (Q36), use slim runtime base images with only the CUDA components you need, pin by digest, and order layers so dependencies change rarely. Platform-side options include pre-pulling images onto GPU nodes with a DaemonSet, node images with common layers cached, and lazy-loading or streaming image features where your provider supports them. The Docker for AI applications guide covers image hygiene in more depth.

Autoscaling AI workloads

18. Why is CPU utilisation the wrong autoscaling signal for LLM serving, and what do you use instead?

Answer: The work happens on the GPU, so the server's CPU stays modest while requests queue. GPU utilisation is also misleading because a busy-looking GPU can still have spare batching capacity. Scale on signals that reflect queueing and saturation: requests waiting in the inference server's queue (vLLM exposes vllm:num_requests_waiting), number of running requests per replica, KV-cache usage, or a latency target such as time to first token. Expose these through Prometheus and feed them to the Horizontal Pod Autoscaler via a custom or external metrics adapter, or use KEDA (Q19). Set a target per replica derived from load testing, for example the queue depth at which time to first token starts breaching its objective.

19. What is KEDA, and how does it relate to the HPA?

Answer: KEDA (Kubernetes Event-driven Autoscaling) is a CNCF project that scales workloads from external event sources. You declare a ScaledObject (or ScaledJob) with triggers such as a Prometheus query, a queue length (SQS, Kafka, RabbitMQ) or a cron schedule. KEDA creates and manages an HPA for the 1-to-N range and handles activation itself, which is how it scales a workload to zero and back. For AI that fits two patterns: inference servers scaled on a Prometheus query of queue depth, and batch workers (embedding, document parsing, evaluation) scaled on queue backlog, including ScaledJobs that launch one Job per batch of work.

20. Should GPU inference scale to zero?

Answer: Only where users can tolerate the cold start. Scaling to zero saves the most money on expensive GPUs, but the first request after idle may wait for a node to be provisioned, the image to pull and weights to load, which can take minutes. It suits internal tools used in office hours, dev and test endpoints, and batch endpoints. It does not suit customer-facing chat with latency objectives; keep a warm minimum there and scale on schedule ahead of known peaks. The core HPA does not scale to zero by default (it has sat behind an alpha feature gate), so teams use KEDA, or Knative through KServe's serverless mode. If you do scale to zero, return a clear "warming up" response or queue requests rather than letting clients time out.

21. Cluster Autoscaler or Karpenter for GPU nodes?

Answer: Both add nodes when pods are unschedulable and remove underused ones. Cluster Autoscaler scales predefined node groups, so you create a group per GPU type and it picks among them; behaviour is predictable and it works across many providers. Karpenter provisions nodes directly from flexible NodePool constraints (instance families, GPU types, capacity type such as spot or on-demand, zones), chooses a fitting instance for the pending pods, and consolidates underused nodes. It is common on EKS and also underpins AKS node auto-provisioning. For GPUs, whichever you use: restrict GPU NodePools to GPU workloads with taints, set limits so a bug cannot provision unbounded GPUs, tune consolidation so it does not evict long-running inference or training, and remember that cloud GPU capacity can be unavailable in a zone regardless of the autoscaler.

22. Walk through what happens when a traffic spike hits a self-hosted LLM service.

Answer: There are several layers, each with its own delay:

Spike -> queue depth rises on running replicas
  -> HPA/KEDA raises replica count (seconds)
  -> new pods Pending: no free GPU
  -> node autoscaler requests GPU node (minutes)
  -> node joins, driver + device plugin ready
  -> image pull (minutes if large, uncached)
  -> weights load into GPU memory
  -> startup probe passes, pod Ready
  -> router sends traffic to new replica

The total can easily be several minutes, so autoscaling alone cannot absorb a sudden spike. Combine it with headroom (a warm minimum, or low-priority placeholder pods that reserve GPU nodes and get preempted when real replicas need them), scheduled scaling before known peaks, faster weight loading (Q36), and admission control at the gateway so excess requests get a fast, honest "busy" response rather than timing out.

23. Can you use spot or preemptible GPUs for AI workloads?

Answer: Yes for fault-tolerant work: batch inference, embedding backfills, evaluation runs and training that checkpoints frequently. Spot capacity can be reclaimed with short notice, so handle the termination signal, checkpoint or drain, and let the job controller or Kueue requeue the work. For online inference, spot is usable only as extra capacity on top of an on-demand baseline that can carry the minimum load, with replicas spread across capacity types and zones. GPU spot availability is often scarce and uneven, so let the node autoscaler fall back across several instance types and keep an on-demand fallback for anything with a deadline.

Model serving on Kubernetes

24. What does a production-ready vLLM deployment on Kubernetes need?

Answer: Beyond the container and a GPU request: a pinned image digest and model revision; weights from a cache rather than a public hub at start-up (Q36); startup, readiness and liveness probes against the server's health endpoint; enough shared memory for multi-GPU tensor parallelism (an in-memory emptyDir mounted at /dev/shm); explicit settings for maximum context length and GPU memory fraction so the KV cache is sized deliberately; Prometheus scraping of the server's metrics; an autoscaler on queue depth; a long terminationGracePeriodSeconds with a pre-stop step so in-flight streams finish during rollouts; a PodDisruptionBudget; and a service account with no cloud permissions beyond reading its model. Put authentication, rate limiting and quotas in a gateway in front of it, because the inference server's OpenAI-compatible API is not a security boundary.

Interview tip: The shared-memory detail and the termination grace period are the kind of specifics that separate people who have run it from people who have read about it. The LLM inference serving interview questions go deeper on batching, KV cache and quantisation.

25. What does KServe add over a plain Deployment?

Answer: KServe, now a CNCF incubating project, gives you a model-serving custom resource so teams declare what to serve rather than hand-writing Deployments, Services, autoscalers and routes. The InferenceService bundles the model storage location, serving runtime, resources, autoscaling, and canary traffic splitting across predictive and generative runtimes. For LLMs it supports runtimes such as vLLM with OpenAI-compatible endpoints, and a dedicated LLMInferenceService resource for generative workloads. It offers a serverless mode (Knative-based, including scale to zero) and a standard mode on plain Deployments, plus a model cache feature that stores weights on node-local storage to speed up starts. The trade-off is another control plane to install, upgrade and debug, and a moving API surface, so pin versions and check the current docs.

26. What is the Gateway API Inference Extension, and what problem does it solve?

Answer: It is a Kubernetes SIG project that turns Gateway API gateways into inference-aware routers. Ordinary load balancing assumes requests cost about the same; LLM requests do not, and a replica that already holds a matching prompt prefix in its KV cache can answer faster. The extension adds an InferencePool resource (a set of model-server pods, selected by labels, used as a route backend) and an Endpoint Picker that chooses a pod per request using live model-server metrics such as queue length and KV-cache usage, and can account for prefixes and LoRA adapters. InferencePool reached v1 with the project's v1.0 release; the companion InferenceObjective resource, which carries request priority, is still alpha. Several gateways implement it, including Istio, NGINX Gateway Fabric, agentgateway and GKE's managed Inference Gateway. Check the conformance and implementations list for your gateway before relying on a feature.

27. What is llm-d, and when would you consider it?

Answer: llm-d is an open-source, Kubernetes-native stack for distributed LLM inference, now a CNCF Sandbox project. It builds on inference engines such as vLLM and on the Gateway API Inference Extension, and packages tested "well-lit path" recipes for patterns that are hard to assemble yourself: prefix-cache-aware routing, prefill/decode disaggregation (running the compute-heavy prompt-processing phase and the memory-bound token-generation phase on separate pools), and wide expert parallelism for large mixture-of-experts models. Consider it when you serve large models at a scale where routing and disaggregation measurably improve latency or cost. For one or two modest models, a well-configured vLLM Deployment or KServe resource is simpler. It moves quickly, so pin releases and validate with your own traffic.

28. What breaks when you stream LLM responses through ingress, and how do you fix it?

Answer: Streaming uses server-sent events or chunked HTTP (sometimes gRPC or WebSockets). Common breakages: proxy or load-balancer idle timeouts cut long generations; response buffering in the ingress controller holds tokens until the response completes, so users see nothing and then everything; HTTP/2 or gRPC settings differ between hops; and rolling updates or node drains kill open streams. Fixes: raise read and idle timeouts on every hop (CDN, cloud load balancer, ingress or gateway, service mesh), disable proxy buffering for streaming routes, send periodic keep-alive comments if a hop enforces idle limits, set a termination grace period longer than your longest expected generation, and make clients resume or retry cleanly. Test with a long generation end to end, not just a health check.

29. How do you roll out a new model version safely on Kubernetes?

Answer: Treat the model version as a deployable artifact, declared in git. Run the offline evaluation gate first, then deploy the new version alongside the old one: shadow traffic for comparison, then a small canary weight via KServe traffic splitting, Gateway API weighted routes or an Argo Rollouts analysis step. Watch quality signals (evaluation scores on sampled traffic, user feedback, refusal and error rates) as well as latency and GPU memory, because a new model can be "healthy" and still worse. Capacity matters: two versions side by side need GPUs for both, so plan headroom or use a smaller canary pool. Rollback is a git revert of the version pin. The MLOps and LLMOps interview questions cover evaluation gates in detail.

30. Where do you run embedding and reranking models?

Answer: These models are small compared with generative LLMs, and their traffic is often bursty (ingestion backfills) or latency-critical (query-time embedding and reranking in RAG). Options: a hosted embedding API, which is simplest; small CPU replicas for lighter models; or a GPU slice (MIG instance or time-sliced GPU) shared by several small models. Keep ingestion-time and query-time embedding on separate deployments so a backfill cannot starve live queries, scale the ingestion side on queue backlog with KEDA, and pin the exact embedding model version, because changing it means re-embedding the index.

Want to practise these patterns on real clusters rather than slides? Cloudsoft's Kubernetes training in Hyderabad covers scheduling, autoscaling, Helm, GitOps and troubleshooting hands-on, in the Ameerpet classroom or live online; call +91 96660 19191 for a free demo.

Training and batch workloads

31. How do Jobs, Indexed Jobs and JobSet fit AI batch work?

Answer: A Job runs pods to completion with retries; that suits a single-node fine-tune or an evaluation run. An Indexed Job gives each pod a stable completion index, ideal for sharded work such as embedding a corpus split into N shards. Pod failure policies let you fail fast on non-retriable errors (a bad config) while retrying infrastructure failures, and per-index backoff limits stop one bad shard from failing everything. JobSet groups several Jobs into one unit, for example a driver plus workers in distributed training, with shared lifecycle, failure and restart policies. Higher-level tools such as Kubeflow Trainer use JobSet underneath.

32. What is Kueue, and why do GPU clusters need it?

Answer: Kueue is a Kubernetes-native job queueing system. The default scheduler places pods as soon as resources exist, which on a busy GPU cluster leads to partially started distributed jobs holding GPUs while waiting for the rest, and to first-come-first-served unfairness between teams. Kueue admits whole workloads against quota: ClusterQueues hold quota per ResourceFlavor (for example, a GPU type or spot versus on-demand), teams submit to namespaced LocalQueues, ClusterQueues in a cohort can borrow unused quota, and preemption reclaims it by priority. Admission checks can wait for node provisioning before admitting, topology-aware scheduling keeps a job's pods close together, and MultiKueue dispatches jobs across clusters. It integrates with Jobs, JobSet, Kubeflow training jobs, RayJobs and more.

33. What is Kubeflow Trainer, and how does it differ from the older Training Operator?

Answer: Kubeflow Trainer (the v2 successor to Training Operator v1) is a Kubernetes-native project for distributed training and LLM fine-tuning across frameworks such as PyTorch, DeepSpeed, JAX and Hugging Face. Instead of a separate CRD per framework (PyTorchJob, TFJob and so on in v1), it uses a TrainJob that references a reusable TrainingRuntime or ClusterTrainingRuntime, so platform teams define the runtime once and data scientists submit jobs with less infrastructure detail, often through the Kubeflow Python SDK. It builds on JobSet and integrates with Kueue for queueing. If you inherit v1 jobs, plan a migration using the project's guide.

34. When would you use Ray on Kubernetes (KubeRay)?

Answer: When the workload is written for Ray: distributed Python data processing, batch inference over large datasets, hyperparameter search, reinforcement learning, or Ray Serve pipelines that compose several models. KubeRay provides three resources: RayCluster (a head pod and worker groups, optionally autoscaled by Ray's autoscaler), RayJob (creates a cluster, runs a job, optionally tears it down) and RayService (long-running Ray Serve apps with upgrades). Use separate worker groups for CPU and GPU so data loading does not occupy GPU nodes, and put RayJobs under Kueue on shared clusters. If the team does not already use Ray, plain Jobs or Trainer may be simpler.

35. A fine-tuning job runs for hours on preemptible capacity. How do you make it resilient?

Answer: Checkpoint model and optimiser state to durable storage at a fixed interval and on the termination signal; make the training script resume from the latest checkpoint automatically; and let the Job, JobSet or Trainer restart policy plus Kueue requeue the workload. Store checkpoints in object storage, not on the node, and keep a retention policy because checkpoints are large. Set the pod's termination grace period long enough to write a final checkpoint. Track progress (step, loss) in your experiment tracker so a resumed run is visibly the same run. For the fine-tuning side itself, see the fine-tuning LLM interview questions.

Storage, observability, security, tenancy and cost

36. What are the options for getting model weights into pods quickly?

Answer:

OptionProsCons
Download from object storage at start (init container)Simple, versionedSlow cold start on every new pod; egress and throttling
Shared PVC (ReadWriteMany file system)Download once, many pods readRead throughput can bottleneck many simultaneous starts
Node-local cache (local SSD, pre-warmed by a DaemonSet or KServe's model cache)Fast loads after first useCache management, disk per node
Weights as an OCI artifact via the image volume typeSame registry, signing and caching as images; read-only mountNeeds a recent Kubernetes and runtime support
Baked into the application imageOne artifactHuge images, slow pulls, rebuild per model change

The Kubernetes image volume, which mounts an OCI image or artifact read-only into a pod, is stable as of Kubernetes 1.36. Whatever the method, use the safetensors format, pin the revision and checksum, and never pull from a public hub in production.

37. What does the DCGM exporter give you, and how do you read GPU utilisation correctly?

Answer: The NVIDIA DCGM exporter (usually deployed by the GPU Operator) exposes GPU telemetry from NVIDIA's Data Center GPU Manager in Prometheus format, with pod and namespace labels so you can attribute GPUs to workloads. Useful series include GPU utilisation (DCGM_FI_DEV_GPU_UTIL), framebuffer memory used, temperature, power, XID error codes and, where supported, profiling metrics such as graphics engine activity. Be careful with "GPU utilisation": it measures the fraction of time any kernel was running, not how much of the GPU's compute was used, so a lightly batched server can show high utilisation while wasting capacity. Read it alongside engine-activity profiling metrics, memory use and the serving engine's own throughput and batch size.

38. What would you put on a dashboard for a self-hosted LLM service on Kubernetes?

Answer: Four layers. User experience: time to first token, inter-token latency, end-to-end latency percentiles, error and timeout rates. Serving engine: running and waiting requests, KV-cache usage, preemptions or recomputations, tokens per second in and out, batch size. GPU: utilisation, engine activity, memory, XID errors, power and thermal throttling from DCGM. Kubernetes: replica count versus desired, pending pods, node provisioning time, pod restarts and OOM kills, and image-pull and start-up durations. Add cost per thousand tokens per tenant. Alert on objectives (time to first token, error rate), not on raw utilisation. See AI observability for tracing the application layer above this.

39. How do you apply network policies to model-serving and AI workloads?

Answer: Start from default-deny ingress and egress in the namespace, then allow only what each workload needs: the gateway may reach the model server port; the orchestrator may reach the model service, vector database and approved tool endpoints; the model server needs no internet at all if weights come from an internal cache or registry. Route any external egress (to hosted model APIs or SaaS tools) through an egress gateway or proxy with an allow-list and logging, which also gives you a choke point for data-leak controls. Remember that network policy needs a CNI that enforces it, DNS egress must be allowed explicitly, and the metrics scraper needs access to metrics ports.

40. How do you manage secrets such as model-hub tokens and API keys?

Answer: Prefer no secret at all: workload identity (Q9) for cloud APIs and storage. Where a secret is unavoidable, such as a third-party API key or a model-hub token for the internal mirror job, keep it in a cloud secrets manager or Vault and sync it with the External Secrets Operator or mount it with the Secrets Store CSI driver, rather than committing it to git. Enable encryption at rest for Kubernetes Secrets, restrict who can read them with RBAC, mount them only into the pods that need them, and rotate. Never pass keys as plain environment variables in manifests or log them in request traces.

41. How do you secure the AI supply chain on a cluster?

Answer: Treat models like code. Mirror approved models into an internal registry or bucket after scanning, licence review and evaluation; store them in safetensors rather than pickle-based formats that can execute code on load; record provenance (source, revision, checksum, who approved). Sign images and, where you package weights as OCI artifacts, sign those too, then enforce verification with an admission policy (Kyverno, Gatekeeper or a cloud equivalent) so only signed artifacts from approved registries run. Scan images for CVEs, pin by digest, and run pods as non-root with read-only root file systems and no privilege escalation; the GPU driver containers are the main legitimate exception to privileged access. The DevSecOps for enterprise AI guide covers the pipeline side.

42. How do you design multi-tenancy for an internal AI platform?

Answer: Decide the isolation level per tenant type. For internal teams with similar trust, soft multi-tenancy works: a namespace per team, RBAC, ResourceQuotas, network policies, a LocalQueue per team mapped into Kueue ClusterQueues with fair sharing and borrowing, and separate service accounts with scoped cloud identities. For tenants that must not share hardware (regulated data, external customers), use dedicated node pools or MIG instances, or separate clusters. Shared model endpoints serving many teams need tenant identity carried in each request, per-tenant rate limits and token quotas at the gateway, and per-tenant cost attribution. The multi-tenant AI SaaS guide covers the application side of the same problem.

43. What are the main levers for controlling GPU cost on Kubernetes?

Answer: In rough order of impact: right-size the model (a smaller or quantised model that passes evaluation can need far fewer or smaller GPUs); keep GPUs busy with continuous batching and by packing small models with MIG or sharing; scale on real demand with sensible minimums and schedule-based scale-down outside office hours; use spot for batch and training; make idle visible with per-namespace GPU-hour reports from DCGM labels; set node-autoscaler limits; and kill forgotten notebooks and dev endpoints automatically. Also compare against hosted APIs honestly, including the engineering time for operating the GPU platform. The cloud cost optimisation for AI article and the FinOps interview questions expand on this.

44. How do EKS, AKS and GKE differ for AI workloads?

Answer: The Kubernetes concepts are the same; the differences are in GPU node provisioning, driver management, identity and add-ons. All three offer GPU node pools or node groups, GPU-ready node images, workload identity integration, and integration with their own model services. EKS users commonly pair accelerated node images with Karpenter and EKS Pod Identity. GKE offers GPU node pools with managed driver installation, GPU support in Autopilot, and a managed implementation of the Inference Gateway. AKS offers GPU node pools, Karpenter-based node auto-provisioning and the KAITO add-on for deploying open models. Feature names and availability change often and vary by region, so describe the pattern in interviews and check current documentation before committing to a design. GPU quota and regional capacity are usually the real constraint on any of them.

Real-world scenarios

45. Your new model-serving pods are stuck in Pending waiting for GPUs. How do you debug it?

Answer: Pending means the scheduler could not place the pod, or the autoscaler has not provided a node. The pod's events usually say why: insufficient nvidia.com/gpu, untolerated taint, node affinity mismatch, or an unbound volume.

What I would check:

  1. kubectl describe pod events: the exact scheduling failure message.
  2. Whether GPU nodes advertise the resource at all (kubectl describe node, allocatable nvidia.com/gpu). Zero usually means device plugin or driver pods are failing.
  3. Tolerations and affinity versus the actual taints and labels on GPU nodes; a typo in the label value is common.
  4. ResourceQuota in the namespace and, if Kueue is used, whether the workload is admitted or suspended in a queue.
  5. Node autoscaler logs or events: whether it tried to scale, hit its max, or got an "insufficient capacity" or quota error from the cloud.
  6. Fragmentation: enough free GPUs in total but not enough on any one node for a multi-GPU pod.

Production consideration: Alert on GPU pods pending beyond a few minutes, keep cloud GPU quota requests ahead of demand, and allow the autoscaler several GPU instance types and zones.

46. A new inference replica takes many minutes to become ready, so autoscaling is too slow. How do you cut cold-start latency?

Answer: Break the start-up into phases and time each: node provisioning, image pull, weight download, weight load into GPU memory, warm-up. Fix the biggest first.

What I would check:

  1. Node provisioning: keep warm capacity or placeholder pods so new replicas land on existing nodes.
  2. Image pull: slim the image, remove weights from it, pre-pull on GPU nodes, use the provider's image streaming if available.
  3. Weight download: switch to a node-local cache, shared volume or image volume; ensure the bucket is in the same region.
  4. Weight load: use safetensors, check storage throughput, and consider a quantised variant if it passes evaluation.
  5. Engine warm-up: review options such as graph capture that trade start time for runtime speed.
  6. Probes: confirm the startup probe is not adding delay with long periods.

Production consideration: Even after optimisation, cold start stays far slower than for web pods, so pair it with predictive or scheduled scaling and a warm minimum for user-facing endpoints.

47. Finance asks why your GPU nodes look underutilised. How do you investigate and improve it?

Answer: Underutilisation is usually an allocation problem, a batching problem, or both.

What I would check:

  1. Allocated versus used: DCGM metrics by pod and namespace. Are GPUs allocated to idle pods (forgotten notebooks, dev endpoints, overprovisioned minimum replicas)?
  2. Stranded GPUs: nodes where CPU or memory is exhausted by non-GPU pods, or GPUs left unused because pods request more than they need.
  3. Serving efficiency: batch size, running requests and KV-cache usage per replica. Low concurrency means the engine is not batching; consolidate replicas or route more traffic per replica.
  4. Model size: a small model on a large GPU is a candidate for MIG or sharing.
  5. Training jobs: data-loader bottlenecks show up as low engine activity with high CPU or I/O wait.

Production consideration: Publish a weekly GPU-hours-allocated versus GPU-hours-used report per team; visibility changes behaviour faster than policy does.

48. One tenant's heavy prompts are slowing everyone else on a shared inference endpoint. What do you do?

Answer: This is a noisy-neighbour problem: long prompts and long outputs from one tenant fill the batch and KV cache, so other tenants' time to first token rises.

What I would check:

  1. Per-tenant request rate, prompt tokens and output tokens at the gateway, to confirm who and what.
  2. Queue depth and KV-cache usage on the replicas during the slowdown.
  3. Whether limits exist: per-tenant rate limits, maximum tokens per request, concurrency caps.
  4. Routing: is traffic balanced by load, or round-robin? Inference-aware routing reduces hot spots.

Production consideration: Enforce per-tenant token quotas and concurrency limits at the gateway, use request priority (the Inference Extension's objectives or the engine's priority scheduling where available) so interactive traffic beats batch, and move heavy or batch tenants to a separate pool. A LLM gateway is the natural place for these controls.

49. After a GPU driver upgrade, several GPU nodes show zero allocatable GPUs. How do you respond?

Answer: Stop the rollout first, then diagnose. Zero allocatable GPUs means the device plugin is not registering devices, almost always because the driver or toolkit failed on those nodes.

What I would check:

  1. GPU Operator pod status on affected nodes: driver, toolkit, validator and device-plugin pods, and their logs.
  2. Compatibility between the driver version, the node kernel version and the CUDA version the serving images need.
  3. Node events and kernel logs for XID errors or module load failures.
  4. Whether the managed node image already ships a driver that conflicts with the operator-installed one.

Production consideration: Upgrade drivers on a canary node pool first, with workloads exercised end to end, and keep the previous driver version pinned for rollback. Cordon and drain nodes one at a time with PodDisruptionBudgets so serving capacity never drops below the minimum.

50. After raising the maximum context length, model pods start failing with out-of-memory errors at peak. What happened?

Answer: Longer context means more KV cache per sequence. If the engine's memory settings were tuned for the old limit, peak concurrency with long prompts can exhaust GPU memory, or the engine now admits far fewer concurrent sequences, which raises latency.

What I would check:

  1. Whether the errors are GPU memory (engine logs) or container memory (Kubernetes OOMKilled, which points at host RAM limits).
  2. The engine's GPU memory fraction, maximum sequence and batch settings versus the new context length.
  3. Prompt length distribution: did a new feature start stuffing large documents into context?
  4. KV-cache usage and preemption metrics over time.

Production consideration: Load-test context-length changes before release, cap context per route rather than globally, and consider KV-cache quantisation or a larger GPU only after trimming unnecessary context.

51. Users report that long answers stop at around the same point every time. Short answers are fine. How do you debug it?

Answer: A consistent cut-off on long streams almost always means a timeout on one hop of the path, not the model.

What I would check:

  1. Time-to-cut-off: if it is a round number of seconds, look for a matching idle or request timeout.
  2. Each hop in order: CDN or WAF, cloud load balancer, ingress or gateway, service mesh sidecar, application server.
  3. Proxy buffering settings on the streaming route.
  4. Whether the cut coincides with pod terminations during rollouts or scale-down.
  5. The model's own maximum output tokens setting, to rule out a genuine length stop.

Production consideration: Document the timeout budget for the whole path and add a synthetic check that streams a long response through production ingress.

52. A distributed training job keeps getting some workers scheduled while others stay Pending, and the cluster deadlocks. Why, and how do you fix it?

Answer: Without gang scheduling, the default scheduler places each pod independently. Two jobs that each need eight GPUs on a cluster with eight free GPUs can each get four and wait forever for the rest, holding GPUs that neither can use.

What I would check:

  1. Which jobs hold GPUs with incomplete worker sets.
  2. Whether the jobs go through Kueue or a gang-capable scheduler at all.
  3. Whether the node autoscaler can satisfy the full request, or quota caps it.

Production consideration: Admit whole workloads through Kueue, which admits a job only when its full quota is available and can evict and requeue workloads whose pods do not all become ready in time, or use a gang-scheduling plugin such as Volcano. Combine with admission checks that wait for nodes to be provisioned.

53. Monthly GPU spend in the dev cluster has jumped with no new projects. What do you do?

Answer: Find the allocations, then put guardrails in place so it does not recur.

What I would check:

  1. GPU pods by namespace, owner label and age; long-running notebooks and endpoints with no traffic.
  2. Node autoscaler history: new GPU instance types, nodes not scaling down because of pods without disruption permission.
  3. Minimum replica settings and scale-to-zero configuration on dev endpoints.
  4. Jobs that failed and retried in loops, keeping GPUs busy.

Production consideration: Enforce GPU quotas per dev namespace, require owner and expiry labels through admission policy, scale dev endpoints to zero after idle, shut down dev GPU pools outside working hours, and send per-team GPU-hour reports.

54. A bank wants a private LLM on its own cluster with no public internet egress. How do you design the model supply and serving path?

Answer: Consider a bank whose GCC platform team in Hyderabad runs the cluster. Separate the model supply chain from runtime. A controlled mirror job, in a separate environment, pulls approved model revisions, scans and evaluates them, records provenance and pushes them to the bank's internal registry or object store. The production cluster never reaches public hubs.

Public hub -> mirror job (scan, licence, eval, sign)
  -> internal registry / bucket (pinned revision)
     -> GPU node cache or image volume
        -> model server (no egress)
           <- internal gateway (authN, quotas, logs)
              <- banking apps

What I would check:

  1. Default-deny egress in the model namespace, with private endpoints for registry and storage.
  2. Admission policy that only allows signed images and artifacts from the internal registry.
  3. Workload identity with read-only access to the model bucket.
  4. Audit logging of every prompt and response at the gateway, with retention agreed with compliance and personal data handled under the DPDP Act.

Production consideration: Plan for model updates as change-managed releases, and keep capacity for running the old and new version side by side during evaluation. See open-weight LLMs for enterprise for the licence and risk review.

55. Design a shared Kubernetes platform for several product teams who want to serve and fine-tune open-weight models.

Answer: Split serving and training into separate GPU pools with different policies, give teams a self-service interface, and centralise the parts that must be consistent: model supply, gateway, identity, observability and cost.

Teams -> Git (model + runtime values) -> Argo CD
               |
   +-----------+-------------------------+
   | Gateway + Inference Extension       |
   |  (auth, quotas, priority, routing)  |
   +-----------+-------------------------+
               |
   Serving pool (on-demand GPUs, MIG for small models)
     vLLM / KServe per model, autoscaled on queue depth
   Training pool (spot + on-demand)
     Kueue queues per team -> TrainJob / RayJob
   Shared: model registry, weight cache, DCGM + Prometheus,
           per-team cost reports, admission policies

What I would check:

  1. Tenancy model per team (namespace and quotas versus dedicated pools) and who approves new models.
  2. Kueue ClusterQueues, cohorts and borrowing rules for training, with priority for production retraining.
  3. Serving objectives per model and the autoscaling and warm-minimum settings that meet them.
  4. Upgrade paths for Kubernetes, drivers, operators and serving engines, with canary pools.

Production consideration: Start with the smallest platform that meets current needs, such as one serving runtime, one queueing system and one gateway, and add components like disaggregated serving only when metrics show the need. Every operator you add is something your team must upgrade at two in the morning.

Key takeaways

  • GPUs reach pods through device plugins today and DRA increasingly; DRA's core APIs have been GA since Kubernetes 1.34.
  • Taint GPU nodes and add affinity; tolerations alone do not place pods.
  • Autoscale LLM serving on queue depth, KV-cache usage or latency, never CPU; plan for multi-minute cold starts.
  • Weight distribution (caching, image volumes, slim images) is often the biggest lever on start-up time.
  • LLM traffic needs inference-aware routing, long timeouts and streaming-safe rollouts.
  • Use Kueue for whole-job admission and fair sharing of GPUs between teams.
  • Read "GPU utilisation" carefully, and make allocated versus used GPU hours visible per team.

Interview preparation checklist

  • Deploy a small open model with vLLM on a GPU node (cloud or local) and explain every field in your manifest.
  • Write a taint, toleration and node affinity set for a GPU pool from memory.
  • Read the DRA concept page and be able to explain ResourceSlice, DeviceClass and ResourceClaim.
  • Configure an HPA or KEDA ScaledObject on a Prometheus metric and test it under load.
  • Time a cold start phase by phase and try one optimisation.
  • Run a Kueue ClusterQueue and LocalQueue with two namespaces and observe borrowing.
  • Install the DCGM exporter and build a dashboard of GPU memory, utilisation and engine activity.
  • Rehearse three scenarios aloud: pending GPU pods, slow cold start, noisy neighbour.
  • Revise the advanced scenario-based Kubernetes questions for the general troubleshooting round.

FAQ

What skills are required for a Kubernetes for AI role?

Solid core Kubernetes (scheduling, networking, storage, RBAC), GPU basics, one inference engine such as vLLM, autoscaling with custom metrics, Prometheus monitoring, Helm or GitOps, and enough ML knowledge to understand model size, batching and evaluation.

Do I need to know CUDA programming for these interviews?

No. Platform roles expect you to understand drivers, CUDA version compatibility, GPU memory and utilisation metrics, not to write kernels.

Can a DevOps engineer move into AI platform engineering?

Yes. Containers, Kubernetes, CI/CD, IaC and observability transfer directly. The gap is GPU scheduling, inference serving and ML workload patterns, which you can close with a few hands-on projects.

Which tools should I learn first?

Start with Kubernetes fundamentals, Helm, Prometheus and Grafana, then the NVIDIA GPU Operator, vLLM, KEDA and Kueue. Add KServe or KubeRay if the roles you target use them.

Do I need access to expensive GPUs to practise?

Not for most of it. Scheduling, taints, quotas, autoscaling and Kueue can be practised on CPU clusters. Rent a small cloud GPU for short sessions to practise the GPU Operator, vLLM and DCGM metrics, and shut it down afterwards.

Are cloud-specific questions common?

Often. Expect questions on the cloud the employer uses, such as GPU node pools, workload identity and autoscaling on EKS, AKS or GKE, but the underlying Kubernetes concepts carry across all three.

How should a fresher prepare for Kubernetes for AI interviews?

Learn core Kubernetes first, then deploy one small model end to end with probes, autoscaling and monitoring, and be ready to explain every decision. A working, well-explained project matters more than a long tool list.

Is Kubernetes for AI a good career direction in India?

It is a natural growth path for DevOps, SRE and cloud engineers, as GCCs and product companies in Hyderabad and Bengaluru build internal AI platforms. Strong fundamentals plus hands-on GPU and serving experience make you relevant to both platform and MLOps roles.

If you want structured, hands-on preparation, Cloudsoft's Kubernetes course in Ameerpet builds the cluster, scheduling, autoscaling and troubleshooting skills these questions test, and the APEX AI, ML, Cloud and Cyber Security program extends them into AI and MLOps. Both run in the Ameerpet classroom or live online; call +91 96660 19191 to book a free demo.

Share𝕏infβœ‰
EnrollWhatsAppCall us