These GKE interview questions cover what Google Kubernetes Engine interviews for cloud, DevOps, SRE and platform roles test in 2026: how Autopilot and Standard differ, how GKE upgrades, networks, secures and scales clusters, and how you debug a cluster on Google Cloud when something breaks. Interviewers already assume you know Kubernetes; on a GKE round they want to hear what Google manages, what you still own, and which GKE feature you would reach for when a pod cannot pull an image, a node pool runs out of IP addresses or an upgrade breaks a release. The 55 questions below run from fundamentals through release channels, VPC-native networking, Workload Identity Federation for GKE, autoscaling, cost and AI workloads to twelve real-world scenarios.
How to use this guide:
- Freshers and juniors are usually asked about Autopilot vs Standard, node pools, regional clusters, Artifact Registry and how to deploy an app.
- Mid-level DevOps and cloud engineers get release channels, VPC-native IP planning, load balancing, Workload Identity Federation and autoscaling.
- Senior and platform roles are pushed on upgrade strategy, private control plane access, fleets, policy, cost and multi-cluster design.
- This guide focuses on what is specific to GKE. For generic Kubernetes (Pods, Deployments, probes, RBAC basics) use our Kubernetes interview questions guide, and for deeper troubleshooting drills see the advanced scenario-based GKE interview questions.
Contents
- GKE fundamentals (Q1โQ9)
- Release channels, maintenance and upgrades (Q10โQ14)
- GKE networking (Q15โQ23)
- Identity and security (Q24โQ30)
- Autoscaling and cost (Q31โQ35)
- Delivery, storage and observability (Q36โQ40)
- AI on GKE (Q41โQ43)
- Real-world GKE scenarios (Q44โQ55)
- Key takeaways
- Interview preparation checklist
- FAQ
- Related Cloudsoft resources
GKE fundamentals
1. What is Google Kubernetes Engine, and what does Google manage versus what you manage?
Answer: GKE is Google Cloud's managed Kubernetes service. Google runs and patches the control plane (API server, scheduler, controllers and etcd), upgrades it, and integrates the cluster with Google Cloud networking, load balancing, IAM, logging and monitoring. What you own depends on the mode. In Standard you manage node pools, machine types, node capacity and much of the node configuration. In Autopilot Google also manages nodes, and you mostly own workload specs, resource requests, identity, network policy and your application.
In both modes you still own the things that cause most outages: container images, resource requests, PodDisruptionBudgets, probes, IAM bindings, network design and upgrade readiness. A good answer frames GKE as a shared-responsibility model, not "Google handles everything".
Interview tip: Name two things you still own in Autopilot (for example, IP planning and IAM). It shows you have run it, not just read the marketing page.
2. What is the difference between GKE Autopilot and GKE Standard, and which should you choose?
Answer: Autopilot is a mode where GKE provisions and manages the nodes based on your Pod specs, enforces a hardened security baseline and bills mainly on the resources your Pods request. Standard gives you node pools you configure yourself and bills for the nodes (VMs) whether or not they are full. Google's guidance is to use Autopilot for most workloads unless you need privileges or node configuration that Autopilot does not allow.
| Aspect | Autopilot | Standard |
|---|---|---|
| Nodes | Managed by GKE | You create and size node pools |
| Billing | Mostly per Pod resource requests | Per node VM |
| Security baseline | Enforced (no privileged Pods by default, Workload Identity Federation on) | You configure it |
| Flexibility | Constrained Pod specs | Full node and kernel-level control |
| Good fit | Most web, API and batch workloads | Custom agents, special kernels, tight bin-packing control |
Real-world example: A retailer's platform team in Hyderabad might run customer-facing APIs on Autopilot to cut node operations, and keep one Standard cluster for a security vendor's DaemonSet that needs host access.
3. What are compute classes, and can you run Autopilot workloads in a Standard cluster?
Answer: Yes. GKE's documentation states that Autopilot workloads can run in Autopilot clusters or in Standard clusters, using ComputeClasses. A ComputeClass is a custom resource that describes the node characteristics a workload wants (machine families, Spot or on-demand, accelerators, fallback priorities). A Pod selects one with a nodeSelector on cloud.google.com/compute-class, and GKE provisions matching nodes.
This blurs the old "pick a mode forever" decision. A Standard cluster can host most workloads in Autopilot-style managed capacity while keeping hand-built node pools for the few that need them. ComputeClasses are also how you express priorities such as "try Spot first, fall back to on-demand".
Interview tip: Feature availability differs by GKE version, so say "check the current docs for your version" rather than quoting a version number from memory.
4. What is the difference between a zonal and a regional GKE cluster?
Answer: A zonal cluster has a single control plane replica in one zone, so a control plane upgrade or zone problem makes the Kubernetes API unavailable for a while (running workloads keep serving). A regional cluster replicates the control plane across three zones in a region and, by default, spreads nodes across those zones. Autopilot clusters are always regional.
Regional clusters give you an available API during control plane upgrades and survive a zone failure, at the cost of more nodes (node counts apply per zone in Standard) and cross-zone traffic. For production, regional is the default answer; zonal fits dev, test and cost-sensitive batch.
Interview tip: Mention that in a regional Standard cluster, a node pool with "3 nodes" means three per zone, nine in total. Many candidates get this wrong.
5. What is a node pool, and when do you create more than one?
Answer: A node pool is a group of nodes in a Standard cluster that share a configuration: machine type, node image, disk, labels, taints, Spot or on-demand, accelerators and the service account. You create separate pools when workloads have different needs, for example a general pool, a high-memory pool for caches, a GPU pool with a taint so only ML Pods land there, and a Spot pool for batch.
The default node image is Container-Optimized OS with containerd, a locked-down image maintained by Google; Ubuntu images exist for cases that need extra kernel modules or tooling. Separate pools also let you upgrade or roll back one class of nodes at a time.
6. What do node auto-repair and node auto-upgrade do?
Answer: Auto-repair watches node health (for example a node reporting NotReady or running out of boot disk for a sustained period) and recreates the node when it stays unhealthy. Auto-upgrade keeps node versions in line with the control plane and the cluster's release channel. Both are on by default for clusters in a release channel, and Autopilot manages them for you.
Repairs and upgrades both drain nodes, so they rely on the same things: replicas spread across nodes and zones, PodDisruptionBudgets that allow some disruption, and graceful shutdown that completes within the termination period.
7. What happened to GKE Enterprise, and what is a fleet?
Answer: A fleet is a logical group of clusters (GKE and attached clusters) that you manage together: shared identity, multi-cluster networking, configuration and policy. Fleet features used to be sold as GKE Enterprise (earlier Anthos). Google's documentation now says GKE is a single offering without different editions or tiers: the features that were part of GKE Enterprise became part of standard GKE, or are products you add on, and clusters no longer have a tier (changes announced in September 2025).
In interviews, say "previously GKE Enterprise, now fleet management in GKE" and note that some capabilities, such as Cloud Service Mesh and multi-cluster ingress features, are still billed separately. Check current pricing before you promise a client anything is free.
8. How is GKE priced at a high level?
Answer: There is a per-cluster management fee, with a monthly free-tier credit that covers one zonal or Autopilot cluster per billing account. On top of that, Standard bills the Compute Engine VMs in your node pools, while Autopilot bills mostly by the CPU, memory and ephemeral storage your Pods request, with node-based billing for some compute classes and specific hardware. Extended support on the Extended channel costs extra. Load balancers, persistent disks, network egress, logging volume and Cloud NAT are billed separately.
Interview tip: Do not quote prices from memory. Say which line items drive cost and where you would check (the GKE pricing page and billing export with GKE cost allocation).
9. How do IAM and Kubernetes RBAC work together in GKE, and which service accounts are involved?
Answer: Google Cloud IAM controls who can reach the cluster and act on Google Cloud resources; Kubernetes RBAC controls what an authenticated user can do inside the cluster. A user needs IAM permission to get cluster credentials (for example container.clusters.get) and then RBAC (or broad IAM roles such as Kubernetes Engine Developer) to act on objects. You can bind RBAC roles to Google Groups so access follows your directory.
Three identities show up in interviews:
- Node service account: the IAM service account attached to the node VMs. It pulls images and writes logs and metrics. Use a dedicated, minimally privileged one instead of the Compute Engine default.
- Kubernetes ServiceAccount (KSA): the in-cluster identity of a Pod.
- IAM service account or principal: what a Pod becomes when it calls Google Cloud APIs through Workload Identity Federation for GKE (Q24).
Release channels, maintenance and upgrades
10. What are GKE release channels, and how do you choose one?
Answer: A release channel decides which GKE versions a cluster gets and how quickly it is auto-upgraded. There are four:
- Rapid: new minor versions shortly after upstream GA. Use it for pre-production testing.
- Regular (the default): a balance between new features and stability. Recommended for most users.
- Stable: versions arrive after longer validation. For production that values stability over features.
- Extended: lets a cluster stay on a minor version for longer (up to about 24 months in total, including an extended support period that is charged), with security patches. For regulated teams that can upgrade minor versions only rarely.
The old "no channel" (static) option is deprecated; Google's docs give a removal date in 2027. A common pattern is dev on Rapid, staging on Regular and production on Regular or Stable, so issues show up in lower environments first.
11. How do maintenance windows and maintenance exclusions work?
Answer: A maintenance window is a recurring period when GKE may run automatic maintenance such as control plane and node upgrades; work that overruns is paused and resumed in the next window. GKE requires a minimum amount of maintenance availability (at least 48 hours in any 92-day window). Maintenance exclusions block upgrades for a period, with three scopes:
- No upgrades: blocks minor and patch upgrades, up to 90 days.
- No minor upgrades: patches still apply; can run up to the version's end of support.
- No minor or node upgrades: only control plane patches; nodes are left alone, which avoids rescheduling Pods.
Scoped exclusions need a release channel. Long "no upgrades" exclusions are risky because you miss security patches.
Real-world example: Consider a bank in its quarter-end freeze. A "no minor or node upgrades" exclusion keeps control plane security patches flowing while ensuring no node drains during the batch window.
12. How does a GKE upgrade work end to end, and how do you prepare for one?
Answer: GKE upgrades the control plane first (regional clusters stay available; zonal clusters lose the API briefly), then node pools, within the supported version skew. Auto-upgrades follow the release channel, maintenance windows and exclusions; you can also upgrade manually ahead of schedule.
Preparation is mostly about Kubernetes API removals and workload resilience. GKE surfaces deprecation insights and recommendations when it detects calls to APIs that are removed in a later version, and can pause automatic minor upgrades for clusters still using them. Before a minor upgrade I check those insights, upgrade Helm charts and operators that ship old CRD versions, run the target version in a Rapid-channel staging cluster, and confirm PDBs and probes are correct.
dev (Rapid) -> staging (Regular) -> prod (Regular/Stable)
check deprecation insights -> upgrade control plane
-> upgrade one node pool -> watch SLOs -> continue
13. Compare surge upgrades and blue-green node pool upgrades.
Answer: Surge upgrade is the default for Standard node pools. It adds up to maxSurge new nodes and takes up to maxUnavailable old nodes out at a time (default maxSurge=1, maxUnavailable=0), upgrading node by node. It is cheap and quick but has no fast rollback.
Blue-green upgrade creates a full set of new ("green") nodes, cordons the old ("blue") pool, drains it in batches, then soaks (one hour by default, configurable up to seven days) before deleting blue. You can cancel, resume or roll back during the phases before deletion. It needs temporary capacity for two pools, so check quota.
| Need | Choose |
|---|---|
| Low cost, tolerant stateless apps | Surge, maybe higher maxSurge for speed |
| Fast rollback, critical workloads | Blue-green with a soak period |
| Long-running batch Pods that must not be evicted | Autoscaled blue-green (Preview at time of writing) |
14. How do PodDisruptionBudgets and graceful termination behave during GKE node upgrades?
Answer: When GKE drains a node it respects PodDisruptionBudgets, but only for a limited time (up to about an hour for surge upgrades) before it evicts anyway, so a PDB cannot block an upgrade forever. Pods also get their terminationGracePeriodSeconds, which GKE caps during node drains. The practical result: a PDB with minAvailable equal to replicas, or a single-replica Deployment, turns an upgrade into a forced eviction and an outage.
Interview tip: Say what you do instead: at least two replicas spread across zones with topology spread constraints, a PDB that allows one disruption, readiness probes, and a preStop hook so load balancers stop sending traffic before the container exits.
GKE networking
15. What is a VPC-native cluster, and how do alias IP ranges work?
Answer: In a VPC-native cluster, Pods get IP addresses from a secondary range on the cluster's subnet, using alias IP ranges attached to each node's network interface. Pod IPs are therefore real VPC addresses: routable across the VPC and peered or connected networks, visible to firewall rules, and usable directly as load balancer backends. Services get ClusterIPs from a Service range (newer clusters can use a Google-managed Service range so you don't have to plan one).
VPC-native is the default and is required for features such as container-native load balancing, Private Service Connect and Dataplane V2. The older routes-based mode used custom static routes per node and is legacy.
16. How do you plan Pod IP ranges for a GKE cluster?
Answer: By default, each node reserves a /24 alias range (256 addresses) from the Pod range to support up to 110 Pods per node; GKE reserves about twice the max Pods so IPs aren't reused too quickly. Pod range size therefore limits node count, not Pod count. With a /24 per node, a /20 Pod range supports only 16 nodes, whatever the Pods actually use.
Levers: lower max-pods-per-node (32 Pods needs a /26 per node, so the same range holds four times as many nodes), size the Pod range for future node counts, and add more ranges later with discontiguous multi-Pod CIDR. When RFC 1918 space is tight in a large enterprise network, GKE supports non-RFC 1918 ranges, including privately used public ranges, for Pods.
Pod range /20 = 4096 IPs
max-pods 110 -> /24 per node -> 16 nodes
max-pods 32 -> /26 per node -> 64 nodes
17. What is GKE Dataplane V2?
Answer: Dataplane V2 is GKE's eBPF-based dataplane, built on Cilium. It implements Kubernetes Services in eBPF instead of kube-proxy and iptables, enforces NetworkPolicy natively without a separate Calico add-on, and offers network policy logging. It is on by default in new Autopilot clusters and can be chosen for new Standard clusters; it cannot be switched on for an existing cluster.
Benefits are better scale for Services and endpoints and visibility into allowed and denied flows. The limits to mention: you cannot run your own custom eBPF programs that conflict with it, and some third-party eBPF tools may not be compatible.
Interview tip: When asked "how do you know if a NetworkPolicy dropped this connection?", the GKE-specific answer is network policy logging with Dataplane V2.
18. How do you expose an application on GKE, and what is container-native load balancing?
Answer: A Service of type LoadBalancer creates a Google Cloud passthrough Network Load Balancer (external, or internal with the internal annotation). An Ingress, or better a Gateway, creates an Application Load Balancer for HTTP(S), with Google-managed certificates, Cloud Armor and Cloud CDN available through policies.
Container-native load balancing uses network endpoint groups (NEGs) so the load balancer sends traffic straight to Pod IPs instead of node ports and an extra kube-proxy hop. It gives more even load, accurate health checks against Pods and fewer surprises during rollouts. It is the default for Ingress on VPC-native clusters.
19. How does Gateway API work on GKE?
Answer: GKE ships a managed Gateway controller with GatewayClasses that map to Google Cloud load balancers, for example gke-l7-global-external-managed (global external Application Load Balancer), gke-l7-regional-external-managed and gke-l7-rilb (internal Application Load Balancer). The platform team owns the Gateway (listeners, certificates, IP), and app teams attach HTTPRoutes from their own namespaces.
Compared with Ingress, Gateway API has built-in header matching, traffic weighting for canaries and a role-based model instead of many controller-specific annotations. Policies such as health checks and backend security attach through GKE policy resources. The generic Gateway API concepts are in our Kubernetes guide; on GKE, know the class names and which load balancer each creates.
20. What are multi-cluster Gateways, and how do they relate to fleets?
Answer: Multi-cluster Gateways use GatewayClasses ending in -mc (for example gke-l7-global-external-managed-mc). One Gateway, defined in a config cluster of the fleet, programs a single load balancer that sends traffic to backends in several clusters and regions. Services are exported with multi-cluster Services (ServiceExport/ServiceImport) so routes can target them across clusters.
Use cases are multi-region active-active, proximity-based routing and blue-green cluster migration (shift weight from an old cluster to a new one). Mention that multi-cluster features have their own pricing and quotas.
21. What is a private GKE cluster today, and how do you control access to the control plane?
Answer: The modern way to think about it is two separate settings. Private nodes have only internal IPs, so workers cannot be reached from the internet. Control plane access is configured through endpoints:
- DNS-based endpoint (recommended in current docs): every cluster gets a unique FQDN for its control plane, and access is decided by IAM and authentication, not IP addresses. It removes the need for a bastion or proxy to reach the API from other networks.
- IP-based endpoints: an external and/or internal endpoint. The most locked-down IP option is internal endpoint only. Protect IP endpoints with authorized networks (an IP allowlist).
The older "private cluster" setting combined these and needed VPC peering for the control plane; you will still see the term in interviews, so map it to private nodes plus restricted endpoints.
22. How do Pods on private nodes reach the internet and Google APIs, and how does Shared VPC fit in?
Answer: Private nodes have no external IP, so outbound internet traffic (third-party APIs, public registries) goes through Cloud NAT on the subnet's region. Google APIs, including Artifact Registry, are reached through Private Google Access on the subnet, which Cloud NAT enables automatically. Firewall rules and, for stricter setups, VPC Service Controls limit what can leave.
In Shared VPC, a host project owns the network and subnets (including the Pod and Service secondary ranges), and service projects run the clusters. The GKE service agent of the service project needs the right network roles in the host project. This is the standard enterprise layout: a central network team owns IP space, application teams own clusters.
23. What is Cloud Service Mesh on GKE, and when would you use it?
Answer: Cloud Service Mesh is Google's managed, Istio-based service mesh (previously Anthos Service Mesh). It adds mutual TLS between services, identity-based authorization, traffic splitting, retries and timeouts, and service-level telemetry in Cloud Monitoring, with a Google-managed control plane.
Use it when you need zero-trust service-to-service authentication, fine-grained traffic control across many teams, or consistent telemetry, typically in a fleet. Skip it for a few services where Gateway API routing and NetworkPolicy are enough. A mesh adds latency, resources, upgrade work and cost, so the honest answer includes the trade-off.
Identity and security
24. What is Workload Identity Federation for GKE, and how does it work?
Answer: Workload Identity Federation for GKE (previously just "Workload Identity") lets Pods call Google Cloud APIs as a Kubernetes ServiceAccount, without exported service account keys. When enabled, the cluster joins the project's workload identity pool PROJECT_ID.svc.id.goog. The GKE metadata server, running on each node, intercepts calls to the metadata endpoint and exchanges the Pod's Kubernetes token for a federated access token through the Security Token Service.
There are two ways to grant access:
- Direct IAM: grant roles straight to a principal identifier such as
principal://iam.googleapis.com/projects/PROJECT_NUMBER/locations/global/workloadIdentityPools/PROJECT_ID.svc.id.goog/subject/ns/NAMESPACE/sa/KSA. You can also target all Pods in a namespace or cluster. - Impersonation: annotate the KSA with
iam.gke.io/gcp-service-accountand grantroles/iam.workloadIdentityUseron the IAM service account. This is needed for the few APIs that don't accept federated principals.
Autopilot enables it by default; in Standard, node pools must use the GKE metadata server.
25. How do you harden a GKE cluster?
Answer: My checklist, roughly in priority order:
- Private nodes, with control plane access through the DNS-based endpoint and IAM, or IP endpoints limited by authorized networks.
- A dedicated least-privilege node service account; never the default Compute Engine account with an Editor role.
- Workload Identity Federation for GKE for every Pod that calls Google APIs; no JSON keys in Secrets.
- Shielded GKE nodes (secure boot, integrity monitoring) and Container-Optimized OS.
- Release channel with auto-upgrades so patches land.
- Pod Security Admission at baseline or restricted, NetworkPolicies with default deny, Binary Authorization for images.
- Application-layer Secrets encryption with Cloud KMS, and Secret Manager for sensitive values.
- Security posture dashboard, audit logs and alerts.
Autopilot gives you much of this by default, which is a fair argument for it in regulated environments.
26. What is Binary Authorization, and how do you use it on GKE?
Answer: Binary Authorization is a deploy-time control. When a Pod is created, GKE checks the image against a policy and blocks it if the policy is not met, for example "only images from our Artifact Registry repositories" or "only images with an attestation from the CI attestor proving they were built and scanned by our pipeline". Dry-run mode logs violations without blocking, and break-glass lets an authorised engineer deploy in an emergency, with an audit trail.
Continuous validation with check-based platform policies can also monitor running Pods for drift from policy (in Preview at the time of writing). Roll it out with dry-run first, fix violations, then enforce, starting with production namespaces.
27. What does the GKE security posture dashboard provide?
Answer: The security posture dashboard shows security concerns across clusters. Workload configuration auditing scans Kubernetes objects for risky settings such as privileged containers or missing security contexts, is on by default for new clusters, and has no extra charge. Findings carry severities and can go to Security Command Center and Cloud Logging.
Know that the vulnerability-scanning parts have changed: Google's docs mark container OS vulnerability scanning and Advanced Vulnerability Insights as deprecated with shutdown dates, so image vulnerability scanning is better done through Artifact Analysis on Artifact Registry and in your CI pipeline. Check the current deprecation notes before you design around a specific feature.
28. How should you manage secrets for workloads on GKE?
Answer: Kubernetes Secrets are only base64-encoded in the API, so on GKE I layer three controls. First, enable application-layer Secrets encryption so etcd data is encrypted with a key you control in Cloud KMS. Second, keep high-value secrets (database passwords, API keys) in Secret Manager and mount them through the Secret Manager add-on for GKE (CSI driver) or fetch them in code, authenticating with Workload Identity Federation. Third, limit RBAC get/list on Secrets, because anyone who can list Secrets in a namespace can read them.
Secret Manager gives versioning, rotation, IAM per secret and audit logs, which interviews for bank or insurer projects usually want to hear.
29. What are Config Sync and Policy Controller?
Answer: Config Sync is the GitOps service included with GKE. It continuously syncs Kubernetes configuration from a Git repository, OCI image or Helm chart to one cluster or a whole fleet, and reverts drift. Policy Controller is an admission-time policy engine built on Open Policy Agent Gatekeeper, with a library of constraint templates and bundles (for example, Pod Security or CIS-style checks) and audit of existing resources.
Both used to sit under Anthos Config Management and then GKE Enterprise; they are now standalone GKE features in the fleet. Teams that already use Argo CD may keep it; the GKE-specific advantage of Config Sync is fleet-wide rollout without running your own GitOps controllers. Our GitOps interview questions cover the generic patterns.
30. How do you design multi-tenancy on GKE?
Answer: Choose an isolation level per risk. Soft multi-tenancy on one cluster uses a namespace per team, RBAC bound to Google Groups, ResourceQuotas and LimitRanges, default-deny NetworkPolicies, a dedicated KSA per app with Workload Identity Federation, and Policy Controller guardrails. Fleet team scopes and fleet namespaces help manage this across clusters. For stronger isolation, add separate node pools with taints for sensitive tenants, GKE Sandbox (gVisor) for untrusted code, or separate clusters and projects.
Real-world example: Consider a GCC in Bengaluru hosting 30 internal teams. Internal tools share a regional cluster with namespace isolation; the payments team gets its own cluster in its own project because auditors want a clear boundary for IAM, logs and network.
Autoscaling and cost
31. How does the cluster autoscaler work on GKE?
Answer: The cluster autoscaler adds nodes to a node pool when Pods are unschedulable because of resource requests, and removes nodes whose Pods can be packed elsewhere. It works from requests, not live CPU usage, which is why accurate requests matter. You set minimum and maximum per pool (per zone or total), and it respects PDBs, node selectors, affinity and taints when deciding.
GKE adds autoscaling profiles: balanced (default) and optimize-utilization, which scales down more aggressively and favours packing, good for batch and cost-sensitive clusters. Pods annotated as not safe to evict, or using local storage, can block scale-down; check those first when nodes don't shrink.
32. What is node auto-provisioning, and how does it relate to ComputeClasses?
Answer: Node auto-provisioning, now documented as node pool auto-creation, lets GKE create and delete whole node pools that fit pending workloads, instead of only resizing pools you created. If a Pod asks for a GPU type or a large memory shape that no pool offers, GKE creates a suitable pool within resource limits you set for the cluster.
Google now recommends enabling auto-creation per workload through ComputeClasses rather than cluster-wide; on recent versions you can do that without turning on cluster-level auto-provisioning. It is not on by default in Standard. Autopilot does the equivalent for you.
33. How do HPA, VPA and multidimensional Pod autoscaling work on GKE?
Answer: The HorizontalPodAutoscaler changes replica counts from CPU, memory or custom and external metrics; on GKE, custom metrics usually come from Cloud Monitoring or Managed Service for Prometheus through an adapter. The VerticalPodAutoscaler is offered by GKE as a managed feature (enabled per cluster in Standard, on in Autopilot) and recommends or applies CPU and memory requests. Multidimensional Pod autoscaling (MultidimPodAutoscaler) combines horizontal scaling on CPU with vertical scaling on memory for the same workload.
The rule to state: do not let HPA and VPA both act on the same metric for the same workload. A common pattern is VPA in recommendation mode to size requests, and HPA on CPU or a request-rate metric for scale-out.
34. How do you use Spot VMs safely on GKE?
Answer: Spot VMs are much cheaper but can be reclaimed at any time. On a preemption notice, GKE gives Pods a short graceful termination window (30 seconds by default, extendable on recent versions). Spot nodes are labelled cloud.google.com/gke-spot=true; add a taint to Spot pools so only workloads that tolerate interruption land there. In Autopilot you request Spot through a ComputeClass or node selector.
Good fits: stateless workers behind queues, CI runners, batch, rendering, and fault-tolerant training with checkpoints. Bad fits: single-replica services, databases and anything that can't finish shutdown in the termination window. Keep an on-demand fallback (a ComputeClass priority list does this well).
35. What levers do you use to reduce GKE cost?
Answer: Start with measurement: billing export with GKE cost allocation, so cost is visible per namespace and label. Then, in rough order of impact:
- Right-size requests using VPA recommendations; over-requested Pods waste money in both modes.
- Pick the right mode: Autopilot removes idle node capacity; Standard can be cheaper for densely packed, steady workloads.
- Autoscale everything: HPA, cluster autoscaler with optimize-utilization for batch, scale dev to zero out of hours.
- Spot VMs for interruptible work; committed use discounts for the steady baseline.
- Watch the non-compute line items: logging volume, cross-zone and egress traffic, idle load balancers, orphaned disks and Cloud NAT.
Our FinOps interview questions cover the organisational side.
Want to practise these on real clusters rather than slides? Cloudsoft's GKE and Kubernetes training in Hyderabad covers Autopilot and Standard, VPC-native networking, Workload Identity Federation, upgrades and troubleshooting labs, in our Ameerpet classroom or live online. Call +91 96660 19191 for a free demo.
Delivery, storage and observability
36. How do you build a CI/CD pipeline to GKE?
Answer: A typical Google Cloud pipeline builds the image (Cloud Build, GitHub Actions or GitLab CI), pushes it to Artifact Registry, lets Artifact Analysis scan it, optionally signs an attestation for Binary Authorization, then deploys with Cloud Deploy (promotion through dev, staging and prod with approvals), Helm, or GitOps through Config Sync or Argo CD. CI should authenticate to Google Cloud with Workload Identity Federation from the CI provider (for example, GitHub's OIDC token), not a stored key.
Helm still packages applications; GitOps decides how releases reach clusters. Deploy by image digest, not a mutable tag, so what you tested is what runs. For the CI side, see our GitHub Actions interview questions.
37. What storage options does GKE offer for stateful workloads?
Answer: Through CSI drivers that GKE manages: Persistent Disk and Hyperdisk for block storage (zonal, or regional Persistent Disk replicated across two zones for zone-failure tolerance), Filestore for managed NFS shared across Pods, and Cloud Storage FUSE to mount buckets, which is common for ML datasets and model weights. StorageClasses decide disk type, and WaitForFirstConsumer binding keeps the disk in the same zone as the Pod.
The GKE-specific trap is zonality: a zonal disk can only attach in its zone, so a StatefulSet Pod rescheduled into another zone stays Pending. For production databases, many teams still prefer Cloud SQL, AlloyDB or Spanner over running databases themselves.
38. What is Backup for GKE, and how do you plan disaster recovery?
Answer: Backup for GKE is a managed service that backs up Kubernetes resource configuration and, optionally, Persistent Disk volume snapshots for chosen namespaces or applications. A BackupPlan defines source cluster, scope, schedule and retention; a RestorePlan defines the target cluster, conflict handling and transformation rules. You can restore into a different cluster, including another region or project, for DR or migration.
DR on GKE is more than backups: define RPO and RTO, keep cluster creation in Terraform, keep app config in Git, decide warm standby versus restore-on-demand, and test the restore. Cross-region backup storage adds network transfer charges.
39. How do logging and monitoring work on GKE?
Answer: GKE sends system logs and workload logs (container stdout and stderr) to Cloud Logging, and system metrics to Cloud Monitoring (Cloud Logging and Cloud Monitoring were previously branded Stackdriver). You choose which components to collect at cluster level. Cloud Monitoring has built-in GKE dashboards for clusters, nodes, workloads and Pods; Cloud Audit Logs record who changed what in the control plane and IAM.
In production, write structured JSON logs so fields are queryable, use log-based metrics and alerting policies on SLOs, and control cost with exclusion filters and sinks for noisy logs. When someone says "logs are missing", check the cluster's logging configuration, the node service account's logging permission and whether exclusions are dropping them.
40. What is Google Cloud Managed Service for Prometheus?
Answer: It is a managed Prometheus-compatible metrics service built on Monarch, the same datastore Google uses for its own monitoring. With managed collection, an operator runs collectors in the cluster, and you describe scrape targets with PodMonitoring (namespaced) and ClusterPodMonitoring custom resources. You query with PromQL from Cloud Monitoring, Grafana or the Prometheus API, with long retention and no Prometheus servers to scale.
It is Google's recommended approach for Kubernetes metrics. Trade-offs: billing is per sample ingested, so high-cardinality labels and short scrape intervals cost money. Self-deployed collection is also an option if you need to keep existing Prometheus configs. Our Prometheus and Grafana interview questions go deeper into PromQL and alerting.
AI on GKE
41. How do you run GPU and TPU workloads on GKE?
Answer: In Standard you create GPU node pools (GKE can install NVIDIA drivers for you) or TPU node pools and TPU slices; Pods request nvidia.com/gpu or TPU resources and land on tainted accelerator nodes. In Autopilot you request accelerators through node selectors or ComputeClasses and GKE provisions the hardware. Options such as GPU time-sharing and multi-instance GPUs let small inference workloads share a card.
Accelerator capacity is the hard part. Dynamic Workload Scheduler's flex-start mode provisions GPUs or TPUs when capacity is available, for up to seven days, which suits fine-tuning and batch inference; queued provisioning (through a ProvisioningRequest) gets all nodes for a distributed job at once. Reservations cover steady serving. The broader patterns are in our Kubernetes for AI interview questions.
42. What is GKE Inference Gateway?
Answer: GKE Inference Gateway extends the GKE Gateway with model-aware routing for generative AI serving. It implements the open-source Gateway API Inference Extension. An InferencePool groups model-server Pods that share a base model, accelerator and server; an InferenceObjective sets serving properties such as priority, so latency-critical and batch traffic can share a pool.
Instead of round-robin, it picks a replica using signals from the model servers: KV-cache utilisation, queue depth, prefix-cache matches and active LoRA adapters. That matters because LLM requests vary hugely in cost, and plain load balancing piles long requests on one replica. Some features (such as predicted-latency routing) are Preview, so check the docs.
43. A team wants to deploy an AI assistant on GKE that calls Gemini models. What does the GKE side look like?
Answer: The application (for example, a FastAPI service with a retrieval layer) runs as an ordinary Deployment behind a Gateway. It calls Gemini through Google's managed AI platform (Vertex AI, now branded Gemini Enterprise Agent Platform) using Workload Identity Federation for GKE, so the Pod's KSA gets only the AI user role it needs and no keys. Egress to Google APIs uses Private Google Access, with VPC Service Controls around the data if it is regulated. HPA scales on request rate or latency, not CPU, because the heavy work happens in the model API.
If the team self-hosts an open-weight model instead, add a GPU node pool or ComputeClass, a model server, Inference Gateway, and Cloud Storage FUSE or a cache for weights. Observability should include token counts, latency and errors per route. For the platform side, see our GCP AI interview questions and Google Cloud for AI engineers.
Real-world GKE scenarios
44. New Pods show ImagePullBackOff when pulling from Artifact Registry. How do you fix it?
Answer: On GKE, kubelet pulls images as the node's service account, not the Pod's Workload Identity. So most failures come down to that account's permissions, the node's access scopes or the network path to Artifact Registry.
What I would check:
kubectl describe podevents: is it403 Forbidden,not foundor a timeout?- Typo in the image path, region, project, repository or tag; confirm the digest exists.
- Node service account has
roles/artifactregistry.readeron the repository, including when the registry is in another project. - Node access scopes include storage read-only or cloud-platform (a common gap on custom node pools).
- Private nodes: Private Google Access enabled on the subnet, or Cloud NAT; firewall or VPC Service Controls not blocking it.
- If an imagePullSecret is used, the identity in that Secret has read access.
Production consideration: Give each cluster a dedicated node service account with reader access only to the repositories it needs, and test pulls as part of node pool creation in Terraform.
45. A Pod using Workload Identity Federation gets "permission denied" calling Cloud Storage. How do you troubleshoot it?
Answer: Work out which identity the Pod is actually using, then whether that identity has the right role on the right resource.
What I would check:
- Cluster has Workload Identity Federation enabled and the node pool uses the GKE metadata server (Standard).
- Pod spec uses the intended KSA (
serviceAccountName), notdefault. - From inside the Pod, ask the metadata server which identity it returns, and compare with what you expect.
- Direct IAM: the principal identifier has the right project number, namespace and KSA name, and the role is granted on the bucket or project.
- Impersonation: KSA annotation is correct and
roles/iam.workloadIdentityUseris granted on the IAM service account toPROJECT_ID.svc.id.goog[NAMESPACE/KSA]. - Code isn't using a mounted key file or
GOOGLE_APPLICATION_CREDENTIALSthat overrides metadata credentials; IAM changes can take a few minutes to apply.
Production consideration: Manage KSA and IAM bindings together in Terraform so namespaces and principals don't drift apart, and use Policy Troubleshooter to explain a denied request.
46. A VPC-native cluster can't add nodes and reports that the Pod IP range is exhausted. What do you do?
Answer: Pod IPs are allocated per node, so the cluster has run out of per-node blocks even if Pods are few. Fix the immediate problem, then the plan.
What I would check:
- Pod range size versus
max-pods-per-node; work out how many nodes the range really supports (Q16). - Short term: add another Pod range with discontiguous multi-Pod CIDR and create new node pools that use it.
- Create node pools with a lower max Pods per node if actual density is low, then move workloads and delete old pools.
- Check the subnet's primary range for node IPs and the Service range too; any of them can be the limit.
- If RFC 1918 space is scarce, agree non-RFC 1918 or privately used public ranges with the network team.
Production consideration: Put IP planning in the cluster design review: expected max nodes per cluster, Pod density and growth, owned by the network team in a Shared VPC with documented ranges.
47. After an automatic GKE upgrade, a workload stopped working. How do you respond?
Answer: Stabilise first, then find which layer changed: control plane version, node version or node image, or an API the workload relied on.
What I would check:
- Cluster operations and notifications: what was upgraded and when; compare with the time the errors started.
- If it was a node pool upgrade still in progress with blue-green, roll back; with surge, create a pool at the previous version (if still allowed) and move workloads there.
- Removed API versions or changed defaults: failing controllers, Helm releases, admission webhooks and CRDs that needed new versions.
- Node-level changes: kernel, containerd or cgroup behaviour affecting agents, privileged DaemonSets or JVM memory settings.
- Release notes for the version and known issues.
- Set a maintenance exclusion to stop further upgrades while you fix it.
Production consideration: Control plane downgrades are very limited, so the real fix is upstream: Rapid or Regular staging clusters, deprecation insights in the upgrade checklist, blue-green for critical pools and maintenance windows during staffed hours.
48. An Autopilot cluster rejects or changes your Pod spec. Why, and how do you handle it?
Answer: Autopilot enforces a security baseline and resource rules through admission. It rejects privileged containers, host namespaces, hostPath writes (read-only hostPath only under /var/log), some Linux capabilities and node selectors on labels it doesn't allow. It also changes resource requests: it applies defaults when requests are missing and adjusts values outside allowed minimums or CPU-to-memory ratios, which can change your bill.
What I would check:
- The exact admission error from
kubectl applyor events; it names the rule that was violated. - Whether the workload actually needs the privilege, or whether a chart default (for example, privileged true) can be turned off.
- For NET_ADMIN, the cluster-level setting that allows it; for monitoring or security agents, whether the vendor is an allowlisted Autopilot partner.
- Requests after admission (
kubectl get pod -o yaml) against what you set, and the compute class for special hardware. - If none of these fit, run the workload on a Standard node pool or a Standard cluster.
Production consideration: Validate Helm charts against Autopilot in CI before go-live, and keep requests explicit so the bill isn't decided by defaults.
49. The GKE bill rose sharply this month. How do you investigate?
Answer: Break the bill down by SKU and label before changing anything, because the cause is often not compute.
What I would check:
- Billing export grouped by service and SKU: GKE compute, Autopilot Pod resources, logging, network egress, load balancers, disks.
- GKE cost allocation by namespace and label to find the team or workload.
- Autoscaling: an HPA with a wrong target, a max replica cap removed, or a cluster autoscaler unable to scale down because of PDBs or local storage.
- Requests inflated by a new chart version, or Autopilot adjusting requests upward.
- Logging: a debug log level left on in production can multiply ingestion.
- Cross-zone or cross-region traffic from a new service, or a new GPU node pool left running.
Production consideration: Set budgets and alerts per project, label every workload with team and cost centre, and review the top movers monthly, with a FinOps owner.
50. Engineers cannot run kubectl against a cluster with private nodes and no public endpoint. How do you give them access safely?
Answer: Prefer the DNS-based endpoint, so access depends on IAM rather than network location; otherwise provide a controlled path into the VPC.
What I would check:
- Which endpoints are enabled: DNS-based, internal IP, external IP, and authorized networks.
- With the DNS-based endpoint: users have the IAM permission for it and fetch credentials for that endpoint (for example,
get-credentialswith the DNS endpoint option). - With an internal IP endpoint only: connect through Cloud VPN or Interconnect from the office network, or an IAP-protected bastion in the VPC, and add those ranges to authorized networks.
- For fleet clusters, the Connect gateway as another IAM-controlled path.
- Once connected, RBAC: users are in the right Google Groups bound to roles.
Production consideration: CI and GitOps should not need human-style access. Pull-based GitOps inside the cluster, or a CI identity limited to deployment, reduces how many people need the API at all.
51. Pods stay Pending and the cluster autoscaler is not adding nodes. What could be wrong?
Answer: The autoscaler only adds a node if a node from some pool would make the Pod schedulable, and only if it is allowed to create it.
What I would check:
- Pod events and the cluster autoscaler visibility events in Cloud Logging ("no scale-up" reasons).
- Node pool maximums reached, or autoscaling not enabled on the pool.
- Requests larger than any machine type in any pool, or a nodeSelector, affinity or toleration no pool satisfies; consider node pool auto-creation or a ComputeClass.
- Compute Engine quota (CPUs, GPUs, IP addresses) or zone capacity shortages, common for GPUs.
- Pod IP range exhaustion (Q46) or a PVC bound to a zone where no pool can scale.
Production consideration: Alert on Pods Pending for more than a few minutes and on autoscaler errors, and keep quota headroom for surge and blue-green upgrades.
52. A batch platform on Spot nodes loses many Pods at once and jobs restart from zero. How do you fix the design?
Answer: Spot reclamation is expected; jobs restarting from zero is a design issue.
What I would check:
- Whether jobs checkpoint progress to Cloud Storage or a database and resume from the checkpoint.
- Whether shutdown handling (SIGTERM) completes within the Spot termination window.
- Diversity: several machine families and zones in the ComputeClass so one reclaim doesn't hit everything.
- An on-demand fallback for jobs with deadlines, and Job
backoffLimitand Pod failure policies that don't count preemptions as application failures. - For accelerator jobs, flex-start provisioning instead of Spot when interruptions are too costly.
Production consideration: Track cost saved against work lost. If restarts cost more compute than Spot saves, move the long jobs to on-demand or flex-start.
53. The load balancer created by a Gateway or Ingress returns 502 errors while the Pods look healthy. What do you check?
Answer: With container-native load balancing, the Google Cloud load balancer health-checks Pods directly, so a mismatch between its health check and your app is the usual cause.
What I would check:
- Backend service health in the console or
gcloud: are NEG endpoints unhealthy? - Health check path, port and protocol versus what the app serves; set them with a health check policy rather than relying on defaults.
- Firewall rules that allow Google's health-check ranges to reach Pod IPs.
- During rollouts: readiness probes, a preStop delay and enough termination grace so the load balancer drains endpoints before Pods exit.
- Backend timeout versus slow endpoints, and HTTP/2 or TLS settings to the backend.
Production consideration: Add a synthetic check through the public endpoint and alert on load balancer 5xx rates, not just Pod health.
54. Design disaster recovery for a critical GKE application across two regions.
Answer: Start from RPO and RTO, then choose active-active or active-passive. Active-active runs regional clusters in two regions in one fleet, behind a multi-cluster Gateway, with a data layer that replicates across regions (for example Spanner, or a database with cross-region replicas). Active-passive keeps a smaller warm cluster or rebuilds from Terraform and restores with Backup for GKE.
Global external ALB (multi-cluster Gateway)
/ \
Region A regional cluster Region B regional cluster
(Config Sync from Git) (Config Sync from Git)
\ /
Cross-region data layer + Backup for GKE
Production consideration: The data layer and DNS or traffic shifting decide real RTO, not the clusters. Run game days where you actually drain one region, and keep images in a multi-region Artifact Registry repository.
55. Design a production GKE platform for a bank's customer-facing services.
Answer: Consider a bank moving its mobile banking APIs to Google Cloud. I would propose:
- Projects and network: Shared VPC with planned Pod and Service ranges; separate service projects per environment; VPC Service Controls around data services.
- Clusters: regional clusters with private nodes, DNS-based control plane endpoint with IAM, Regular or Stable channel, maintenance windows in staffed hours and exclusions for freeze periods. Autopilot unless a vendor agent forces Standard.
- Identity: Workload Identity Federation per app, RBAC through Google Groups, no service account keys.
- Supply chain: Artifact Registry with scanning, signed attestations, Binary Authorization enforced in production.
- Traffic: Gateway API with the global external Application Load Balancer, Cloud Armor, managed certificates; Cloud Service Mesh if mTLS between services is required.
- Policy and delivery: Config Sync and Policy Controller from Git across the fleet; Cloud Deploy or Argo CD with approvals.
- Operations: Managed Service for Prometheus, SLO-based alerts, audit logs to a locked logging project, Backup for GKE and a tested DR runbook.
Production consideration: Explain the order of rollout (network and identity first, then supply chain, then workloads) and who owns what between the platform, security and app teams. Interviewers score the trade-offs, not the length of the list.
Key takeaways
- GKE interviews assume Kubernetes knowledge and test the Google Cloud layer: modes, upgrades, networking, identity and cost.
- Autopilot is Google's recommended default; ComputeClasses let Standard clusters run Autopilot-style workloads too.
- Upgrades are controlled by release channels, maintenance windows, exclusions and the node upgrade strategy; PDBs only delay drains.
- VPC-native IP planning (per-node ranges, max Pods) is the networking topic that breaks real clusters.
- Use Workload Identity Federation for GKE for Pod access to Google APIs, but remember image pulls use the node service account.
- Control plane access is now about the DNS-based endpoint and IAM; "private cluster" maps to private nodes plus restricted endpoints.
- Fleet features formerly sold as GKE Enterprise are part of GKE now; some add-ons are still billed separately.
Interview preparation checklist
- Create one Autopilot and one Standard cluster; deploy the same app and compare billing, admission errors and node handling.
- Build a VPC-native cluster in Terraform with custom Pod and Service ranges; work out how many nodes it supports.
- Set up Workload Identity Federation for a Pod reading a Cloud Storage bucket, both with direct IAM and with impersonation.
- Push an image to Artifact Registry and deliberately break the node service account's permission to see the error.
- Expose an app with Gateway API and a health check policy; run a canary with weighted HTTPRoutes.
- Configure a release channel, a maintenance window and an exclusion; run a blue-green node pool upgrade and a rollback.
- Add HPA, VPA in recommendation mode and a Spot node pool with a taint; watch the cluster autoscaler events.
- Enable Managed Service for Prometheus with a PodMonitoring resource and build one SLO alert.
- Prepare two stories: an incident you debugged and a cost or reliability improvement, with the commands you ran.
FAQ
Is GKE a good skill for a cloud or DevOps career in 2026?
Yes, for teams on Google Cloud. GKE builds on Kubernetes, which transfers to EKS and AKS, and adds Google Cloud networking, IAM and operations skills that platform and SRE roles use every day.
Should I learn Kubernetes before GKE?
Yes. Learn Pods, Deployments, Services, probes, requests and RBAC first, then learn what GKE adds: Autopilot, release channels, VPC-native networking, Workload Identity Federation and Google Cloud load balancing.
Is Autopilot or Standard asked more in GKE interviews?
Both come up. Expect to explain the differences, when Autopilot rejects a Pod spec, how billing differs and when you would still pick Standard node pools.
Which Google Cloud certification relates to GKE?
The Associate Cloud Engineer and Professional Cloud DevOps Engineer exams both include GKE topics. A certification helps with shortlisting, but interviews still test hands-on troubleshooting.
How much hands-on GKE practice do I need before interviews?
Enough to have built, upgraded, scaled and broken at least one cluster yourself. Interviewers quickly spot answers learned only from documentation.
Do I need Terraform for GKE roles?
Usually yes. Most teams create clusters, node pools, networks and IAM bindings with Terraform, so be ready to read and explain a GKE module.
How do I keep my GKE knowledge current?
Read the GKE release notes and deprecation pages regularly. Features and names change often, for example Workload Identity Federation for GKE, DNS-based endpoints and the end of GKE Enterprise as a separate edition.
What GKE projects should I show in a portfolio?
Show a Terraform-built regional cluster with private nodes, Workload Identity Federation, Gateway API, GitOps deployments, autoscaling and dashboards, plus a short write-up of an upgrade or incident you handled.
If you want guided, lab-based preparation for GKE roles, Cloudsoft's GKE Kubernetes course covers cluster design, networking, security, upgrades and real troubleshooting, with mock interviews, in Ameerpet or live online. For a wider path across AI, ML, cloud and security, see the APEX AI, ML, Cloud and Cyber Security program. Call +91 96660 19191 for a free demo.



