Top 20 GCP Interview Questions and Answers (2026 Edition)

Preparing for a GCP interview? These are the 20 questions that come up most often — each with a crisp, interview-ready answer. Every answer links to a full deep-dive guide in this series when you want the follow-up traps and trade-offs.

How to use this: read the question, say your answer out loud, then check. Interviewers score decision rules and failure modes, not definitions — "when would you pick X, and what breaks?"

Compute, Storage & Networking

1. GCE vs GKE vs Cloud Run — when do you use each?

Answer: GCE when you need VM-level control (custom kernels, licensed software, lift-and-shift). GKE when you need Kubernetes orchestration at scale (autoscaling, rolling updates, custom operators). Cloud Run for serverless containers with scale-to-zero — the default to evaluate first for stateless services. Cloud Run functions (formerly Cloud Functions) for single-purpose event handlers. Deep dive →

2. What are Spot VMs, and when are they a trap?

Answer: spare GCP capacity at up to ~91% discount, reclaimable with 30 seconds' notice. Perfect for fault-tolerant batch work (rendering, CI, genomics) designed with checkpointing and idempotent work units. A trap for anything that can't tolerate interruption — no single-replica production services, no un-checkpointed long jobs. Deep dive →

3. How do you pick a Cloud Storage class?

Answer: one decision — how often do you read the data? Standard (hot), Nearline (monthly, 30-day minimum), Coldline (quarterly, 90-day minimum), Archive (yearly, 365-day minimum). The trap is early-deletion fees: delete a Nearline object on day 10 and you still pay for 30 days. Use Object Lifecycle Management or Autoclass for automatic transitions. Deep dive →

4. Which GCP database for which workload?

Answer: the decision rule is access pattern + consistency + scale + geography. Demanding PostgreSQL OLTP → AlloyDB; standard OLTP → Cloud SQL; global horizontally-scalable relational → Spanner; massive low-latency key-value → Bigtable; documents with realtime sync → Firestore; analytics over terabytes → BigQuery. Deep dive →

5. How do you choose a load balancer on GCP?

Answer: first Layer 7 vs Layer 4: HTTP/HTTPS → Application Load Balancer (URL maps, Cloud CDN, Cloud Armor attach here); TCP/UDP → Network Load Balancer. Then scope: external vs internal, global/cross-region vs regional. They're software-defined — they scale automatically, no pre-warming. Deep dive →

Data Engineering

6. How do you keep a BigQuery bill from exploding?

Answer: layers: dry runs and maximum bytes billed before execution; quotas and capacity reservations for governance; budgets and alerts for visibility (alerts don't stop a running query); partition by date and cluster by filter columns; avoid SELECT * on wide tables. Deep dive →

7. Batch vs streaming — how do you decide?

Answer: latency requirement vs cost and complexity. Streaming (Pub/Sub → Dataflow → BigQuery) gives seconds-fresh data but costs more and forces you to handle late/out-of-order data. The senior move: ask how real-time — sub-second, one minute, or fifteen? Most "real-time" needs turn out to be micro-batch. Deep dive →

8. How do you handle duplicates and poison messages in Pub/Sub?

Answer: Pub/Sub is at-least-once by default, so consumers must be idempotent — dedupe on a stable event ID. Configure a dead-letter topic so poison messages are quarantined after N attempts instead of redelivered forever, and monitor oldest unacked message age to catch stuck subscriptions. Deep dive →

9. What are slots in BigQuery, and when do you buy reservations?

Answer: a slot is a unit of BigQuery compute. On-demand (pay per TB scanned) suits spiky workloads; capacity reservations (BigQuery editions) suit steady, predictable workloads needing controlled capacity and cost. One-liner: "on-demand for spiky, reservations for steady." Deep dive →

10. Explain windows, watermarks, and triggers in streaming.

Answer: Window — which events belong together (fixed, sliding, session). Watermark — how far through event time the system believes it has progressed. Trigger — when results are emitted. Allowed lateness — how long late data is still accepted. Streaming is about deciding what to do with late data. Deep dive →

Kubernetes & GKE

11. Walk me through what happens when you run kubectl apply.

Answer: kubectl → API server (auth/authz) → persisted to etcd → controller manager creates a ReplicaSet → scheduler assigns pods to nodes → kubelet pulls images and starts containers. The scoring sentence: "desired state vs actual state — controllers continuously reconcile the difference." Deep dive →

12. Requests vs limits — what happens when you get them wrong?

Answer: requests drive scheduling decisions; limits cap runtime usage. Memory limit exceeded → OOMKilled; CPU limit exceeded → throttling (slow, not killed); oversized requests → nodes look full while half-idle; no limits → noisy workloads starve neighbors. Right-size with VPA recommendations. Deep dive →

13. HPA vs VPA vs cluster autoscaler — which scales what?

Answer: HPA scales pod count; VPA scales pod size (requests/limits); cluster autoscaler scales node count. One-liner: "HPA scales pods out, VPA scales pods up, cluster autoscaler scales nodes." Be careful combining HPA and VPA on the same signals — their feedback loops interact. Deep dive →

14. GKE Standard vs Autopilot — how do you choose?

Answer: Autopilot when you want Google to manage node infrastructure, scaling, and security configuration; Standard when you need deeper control over nodes and cluster infrastructure. Don't memorize "Autopilot cannot X" — discuss workload-specific restrictions, since the platform keeps getting more flexible. Deep dive →

15. A pod is stuck in CrashLoopBackOff — your debug sequence?

Answer: an ordered runbook: kubectl describe pod (events tell the story), kubectl logs --previous (the crashed container's last words), then check the usual suspects — bad image, missing config, failing probes, memory limit too low. Know the trio: Pending → couldn't schedule; CrashLoopBackOff → started then failed; ImagePullBackOff → couldn't pull the image. Deep dive →

IAM, Security & SRE

16. How does GCP IAM work, and what is least privilege in practice?

Answer: who (principals) + what (roles) + where (resource); allow bindings inherit downward, but deny policies can override grants. Least privilege in practice: predefined roles instead of Editor, the narrowest scope matching the operational model, IAM Recommender for over-privileged accounts, org policies blocking risky configurations. Deep dive →

17. How do workloads authenticate to GCP without service-account keys?

Answer: Workload Identity Federation for GKE gives pods short-lived federated identities — IAM roles can be granted directly to Kubernetes identities, no key files. Same pattern for external workloads (GitHub Actions, AWS, on-prem). One-liner: "No static keys to rotate, no static keys to leak." Deep dive →

18. Define SLI, SLO, SLA — and what is an error budget?

Answer: SLI = the measurement (e.g., proportion of requests succeeding under 300 ms); SLO = the target on the SLI (99.9% over 30 days); SLA = the contractual commitment with consequences. The error budget is the gap between 100% and the SLO — teams use it to balance release velocity against reliability work. Deep dive →

19. Your error budget is burning — walk me through the incident.

Answer: detect (burn-rate alert), triage (declare incident, assign roles), mitigate first (roll back, fail over, shed load — restore the SLI before root-causing), resolve, then a blameless postmortem with action items and owners. The scoring sentence: "Triage enough to choose a safe mitigation, then restore service before deep root-cause analysis." Deep dive →

20. How do you ship safely on GCP?

Answer: rolling (gradual, default), blue/green (two environments, instant rollback), canary (small traffic %, judge against production signals, then ramp). On GCP, Cloud Deploy orchestrates this with canary phases; automated progressive delivery uses monitoring/SLO checks as promotion gates. Deep dive →

In this series

  1. Top 20 GCP Interview Questions and Answers (2026 Edition) (this post) — start here.
  2. GCP Interview Questions: Compute, Storage, and Networking Fundamentals — pick the right compute, storage, and network.
  3. GCP Data Engineering Interview Questions: BigQuery, Dataflow, Pub/Sub — analytics at scale.
  4. Kubernetes and GKE Interview Questions — from kubectl to 3 AM debugging.
  5. GCP IAM, Security, and SRE Interview Questions — identity, safety, and staying alive.

Related: System Design Interviews: A Practical Primer — the 4-step framework these cloud questions plug into.

Comments

Popular posts from this blog

Java Banking Finance Services and Insurance (BFSI) domain interview questions

JSP Servlet Interview Questions For Freshers Series 1

Java program to check even or odd number