Kubernetes and GKE Interview Questions

Kubernetes questions test whether you've actually operated clusters or just read about them. These are the K8s and GKE questions interviewers ask — from core concepts to the 3 AM troubleshooting scenarios.

Kubernetes Core

1. Walk me through what happens when you run kubectl apply -f deployment.yaml

Answer: the interview classic — they want the control-plane flow: kubectl sends the manifest to the API server (after auth/authz); the object is persisted to etcd; the controller manager notices the new Deployment and creates a ReplicaSet; the scheduler assigns pods to nodes; each node's kubelet pulls images and starts containers via the container runtime. Service networking is handled separately by the cluster's networking components when a Service exists. Say "desired state vs actual state — controllers continuously reconcile the difference" — that's the sentence that scores.

2. Pod, Deployment, StatefulSet, DaemonSet — when do you use each?

Answer:

  • Pod — the smallest unit; one or more tightly-coupled containers sharing network and storage. You almost never create pods directly.
  • Deployment — stateless apps. Manages ReplicaSets, rolling updates, rollbacks. Your default for APIs and workers.
  • StatefulSet — stateful apps needing stable identity and storage: databases, Kafka, ZooKeeper. Stable pod names, ordered rollout, per-pod PersistentVolumeClaims.
  • DaemonSet — one pod per node: log collectors, monitoring agents, CNI plugins.

The trap: "Can I run a database in a Deployment?" — a StatefulSet gives stateful workloads stable identity and storage semantics, but it doesn't make a database highly available by itself. The experienced answer: StatefulSet for the Kubernetes primitives, plus database-level replication, failover, backups, and recovery for actual HA. Know both halves.

3. How does service discovery work? ClusterIP vs NodePort vs LoadBalancer vs Ingress

Answer: a Service gives pods a stable virtual IP + DNS name (my-svc.my-ns.svc.cluster.local) in front of ephemeral pod IPs.

  • ClusterIP (default) — cluster-internal virtual IP.
  • NodePort — exposes a port on nodes; also used as a building block in some architectures.
  • LoadBalancer — asks the platform/cloud integration for a load balancer per service. Simple but expensive (one LB per service).
  • Ingress — HTTP(S) routing using the Ingress API and a controller: one LB, path/host rules fan out to many services.
  • Gateway API — the newer, more expressive traffic-routing model; Google describes Gateway as an evolution of Ingress capabilities.

GKE specifics interviewers probe: GKE Ingress provisions a Google Cloud external HTTPS load balancer. Know both Ingress and Gateway API — for new designs, be ready to discuss Gateway API.

4. ConfigMaps vs Secrets — and what's wrong with Secrets?

Answer: ConfigMaps hold non-sensitive config; Secrets hold sensitive data — but a Kubernetes Secret is not made secure merely because its value is base64-encoded (base64 is encoding, not encryption). Protect Secret data with proper RBAC and encryption-at-rest controls — and note GKE configuration can differ from vanilla Kubernetes defaults here. The stronger pattern on GKE: Secret Manager with Workload Identity Federation, so applications fetch secrets at runtime without static credentials (the external-secrets operator is the portable equivalent).

5. Requests vs limits — and what happens when you get them wrong?

Answer: requests = the amount Kubernetes uses for scheduling and resource-allocation decisions (bin-packing pods onto nodes; also influences QoS); limits = the maximum resource usage enforced or constrained at runtime, depending on resource type. Get them wrong and you meet the failure modes interviewers love:

  • Memory limit exceeded → OOMKilled (exit 137).
  • CPU limit exceeded → throttling (not killed — just slow, which confuses everyone).
  • Requests far above actual usage → nodes look "full" while half-idle (low bin-packing efficiency).
  • No limits on a noisy workload → it can consume excess resources and create contention with neighboring pods.

Follow-up: "How do you right-size them?" — Vertical Pod Autoscaler in recommendation mode, or GKE's cost-optimization dashboards.

Scaling & Updates

6. HPA, VPA, and cluster autoscaler — which scales what?

Answer: the three autoscalers form a stack, and mixing them up is the classic wrong answer:

  • Horizontal Pod Autoscaler (HPA) — scales pod count on CPU/memory/custom metrics.
  • Vertical Pod Autoscaler (VPA) — scales pod size (requests/limits). Be careful combining HPA and VPA when both react to the same resource signals — their feedback loops can interact; common designs separate what each autoscaler controls.
  • Cluster autoscaler — scales node count when pods can't be scheduled.

GKE also has node auto-provisioning (creates new node pools automatically). The interview one-liner: "HPA scales pods out, VPA scales pods up, cluster autoscaler scales nodes."

7. How do rolling updates work, and how do you do zero-downtime deploys?

Answer: a Deployment rolling update replaces pods gradually, controlled by maxSurge (how many extra pods) and maxUnavailable (how many can be down). Zero-downtime rolling deploys need: multiple replicas where required, readiness/startup probes (don't send traffic until the app is ready), appropriate maxUnavailable/maxSurge, graceful termination (handle SIGTERM — terminationGracePeriodSeconds), application/backward compatibility across versions, and adequate capacity. PodDisruptionBudgets are a separate mechanism — they limit voluntary disruptions (node drains, maintenance), not rollout mechanics. The trap: "Deploys cause 503s — why?" — start with readiness behavior, endpoint removal/termination races, rollout capacity, and application errors; 503s have many parents.

GKE Specifics

8. GKE Standard vs Autopilot?

Answer: Autopilot — Google manages node infrastructure, scaling, and security configuration; you focus on workloads. Standard — you manage node pools (machine types, autoscaling, upgrades) with deeper control over nodes and cluster infrastructure. Don't treat them as two sealed worlds: Autopilot billing can be pod-based or node-based depending on workload and configuration, Autopilot supports many DaemonSet use cases, and privileged workloads can be admitted under controlled policies — discuss workload-specific restrictions rather than memorizing "Autopilot cannot X." The interview follow-up: "What's the cost gotcha with Autopilot?" — with pod-based billing you pay for requests, so over-requesting burns money; right-size with VPA recommendations.

9. What is Workload Identity Federation for GKE, and why is it better than service account keys?

Answer: Workload Identity Federation for GKE gives workloads short-lived federated identities for accessing Google Cloud APIs — no service-account key files distributed anywhere. You can grant IAM roles directly to Kubernetes workload identities; service-account impersonation remains available for APIs and use cases that require it. It's better than JSON key files because keys leak, expire, and get committed to git — the three horsemen of credential incidents. The interview one-liner: "No static keys to rotate, no static keys to leak."

10. Private clusters, and how do you operate them?

Answer: a private GKE cluster has nodes with internal IPs rather than external IPs. Frame it as two independent questions. Can the nodes reach what they need? — Private Google Access and/or Cloud NAT for egress and API access, plus appropriate private connectivity. Can administrators reach the Kubernetes API endpoint? — configure the control-plane endpoint and access path according to the organization's network model (authorized networks, VPN, or a jump host/IAP-bastion architecture). The interview trap: "kubectl stopped working after we made the cluster private" — node privacy and control-plane reachability are separate dimensions; making nodes private doesn't automatically cut off your admin path, and fixing it means addressing the endpoint's access configuration.

Troubleshooting

11. Pod stuck in CrashLoopBackOff — your debug sequence?

Answer: interviewers want an ordered runbook, not random commands:

  1. kubectl describe pod — events tell you the story (failed probes? OOMKilled?).
  2. kubectl logs --previous — the crashed container's last words.
  3. Check the usual suspects: bad image tag (ImagePullBackOff is different — that's a registry problem), missing env/config, failing readiness/liveness probes, memory limit too low (exit 137).
  4. kubectl get events --sort-by=.lastTimestamp for the namespace-wide view.

The reasoning matters more than the order: first determine whether the failure is scheduling, image startup, container execution, probes, resources, or application behavior — then use events/describe and current/previous logs accordingly. logs --previous is especially valuable for containers that keep restarting.

12. OOMKilled vs Evicted — what's the difference?

Answer: OOMKilled tells you a process/container was killed for memory pressure — check container state, limits, and actual usage (it's often the memory limit, but kernel and container memory behavior can be more nuanced). Evicted means the kubelet removed the pod because the node hit an eviction condition — memory, ephemeral-storage/disk, or other resource pressure — starting with BestEffort pods, then Burstable, respecting QoS. So: OOMKilled → investigate the container's memory; Evicted → investigate node pressure, pod requests/usage, QoS, and competing workloads. A poorly configured pod can contribute to node pressure, and node configuration can contribute to OOM situations — don't split blame too neatly. The follow-up: "How do you prevent evictions?" — set realistic requests, size and autoscale nodes appropriately, monitor node memory and ephemeral storage, control noisy workloads, and understand QoS and eviction priority. (Note: PodDisruptionBudgets protect against voluntary disruptions — they don't prevent kubelet node-pressure eviction.)

13. A pod is stuck in Pending — walk me through why

Answer: Pending means it couldn't get scheduled or start — the complement to CrashLoopBackOff (started, then repeatedly failed) and ImagePullBackOff (couldn't retrieve the image). Reason through the scheduler's checklist:

  • Requests don't fit — no node has enough allocatable CPU/memory; check kubectl describe pod events.
  • Node selector / affinity / tolerations — the pod demands nodes that don't exist or won't accept it (taints without matching tolerations).
  • PVC binding — a PersistentVolumeClaim waiting for a volume (wrong storage class, no provisioner, zone mismatch).
  • Quota — namespace ResourceQuota exhausted.
  • Node availability — all nodes cordoned, or the cluster autoscaler hasn't caught up yet.

The three-way distinction to say out loud: Pending → couldn't get scheduled or start. CrashLoopBackOff → started, then repeatedly failed. ImagePullBackOff → couldn't retrieve or start the image. That framing is extremely interview-useful.

In this series

  1. Top 30 GCP Interview Questions and Answers (2026 Edition) — start here.
  2. GCP Interview Questions: Compute, Storage, and Networking Fundamentals — pick the right compute, storage, and network.
  3. GCP Data Engineering Interview Questions: BigQuery, Dataflow, Pub/Sub — analytics at scale.
  4. Kubernetes and GKE Interview Questions (this post).
  5. GCP IAM, Security, and SRE Interview Questions — identity, safety, and staying alive.

Related: System Design Interviews: A Practical Primer — the 4-step framework these cloud questions plug into.

Comments

Popular posts from this blog

JSP Servlet Interview Questions For Freshers Series 1

Java Banking Finance Services and Insurance (BFSI) domain interview questions

Java program to check even or odd number