GCP IAM, Security, and SRE Interview Questions

The last mile of cloud interviews: identity, security, and keeping production alive. These are the GCP IAM, security, and SRE questions that separate "used the console" from "ran production."

IAM

1. How does GCP IAM actually work — principals, roles, policies?

Answer: the model has three parts: who (principals: Google accounts, service accounts, groups, domains), what (roles: collections of permissions), and where (the resource — org, folder, project, or individual resource). You attach a policy (a list of role bindings) to a resource, and allow bindings inherit downward — org → folder → project → resource. One modern nuance: deny policies also inherit downward and can block a permission even when a role grants it — effective access is not simply the union of all grants. Three role types: basic (Owner/Editor/Viewer — broad, avoid), predefined (job-function roles like roles/bigquery.dataViewer — prefer these), custom (your own permission sets). The interview one-liner: "Grant at the narrowest practical scope that matches the operational model, with predefined roles, following least privilege."

2. What is least privilege, and how do you practice it on GCP?

Answer: least privilege = every identity gets only the permissions it needs, nothing more. In practice: use predefined roles instead of Editor; scope bindings to match the operational model — sometimes a folder-level grant to the right group is cleaner and safer than hundreds of individual bindings; use IAM Recommender to find over-privileged accounts; enforce organization policies that restrict risky configurations, such as external IP usage or service-account key creation (iam.disableServiceAccountKeyCreation — enforced by default for newer organizations); and regularly audit with Policy Analyzer. The trap: "The app needs BigQuery access — just give it Editor, right?" — the answer they want is a firm no, then a scoped example: for a read-querying workload, that might mean roles/bigquery.jobUser on the project plus dataset-level dataViewer — but the right roles depend on whether the workload reads, writes, runs jobs, or uses the Storage API.

3. Service accounts — when do you use them, and how do you avoid key sprawl?

Answer: service accounts are identities for workloads (apps, VMs, pipelines), not humans. Rules: use distinct workload identities when workloads have different permission boundaries — don't make unrelated applications share one broad service account; grant minimal roles; and avoid downloading JSON keys — prefer Workload Identity Federation for GKE, attached service accounts (GCE), or Workload Identity Federation for external workloads (GitHub Actions, AWS, on-prem), which uses short-lived federated credentials instead of keys. Access can be granted directly to the federated principal or through service-account impersonation, depending on the API and use case. Key sprawl happens when teams download keys "temporarily"; the org policy iam.disableServiceAccountKeyCreation kills the habit at the source.

4. How do you give a developer temporary elevated access?

Answer: the patterns interviewers want, in order: Privileged Access Manager (PAM) for governed just-in-time elevation with approvals, justifications, and audit — the cleanest primary answer for humans; IAM Conditions for time-bound bindings ("this role expires Friday"); and service-account impersonation where operationally appropriate (grant roles/iam.serviceAccountTokenCreator so they act as a privileged SA briefly). What you never do: permanently grant Owner "because it's urgent."

Security

5. Defense in depth on GCP — name the layers

Answer: the answer is a layered list, outside-in:

  • Edge: Cloud Armor (WAF/DDoS) on the external LB, Cloud CDN.
  • Network: VPC firewall rules, Private Google Access, no public IPs, VPC Service Controls — a service perimeter around supported Google Cloud services that reduces exfiltration risk even when an identity holds valid IAM credentials.
  • Identity: least-privilege IAM, Workload Identity Federation, org policies.
  • Data: encryption by default; Cloud KMS for customer-managed keys (CMEK); Secret Manager for secrets — don't hardcode secret values into YAML, source, images, or repositories; retrieve them through a managed secret path at runtime.
  • Supply chain: Artifact Registry vulnerability scanning, Binary Authorization — enforce deployment policy, such as requiring trusted attestations before an image can deploy.

The follow-up: "What's VPC Service Controls, really?" — a service perimeter around supported Google Cloud services that reduces exfiltration risk even when an identity holds valid credentials (e.g., it makes it much harder for a compromised SA to copy data to an external project).

6. Where do secrets belong — and where do they always end up instead?

Answer: secrets belong in Secret Manager (versioned, audited, IAM-controlled, rotation support), consumed at runtime via the API or Workload Identity Federation. Where they end up instead: hardcoded in source, committed to git, pasted in Slack, baked into container images. The interview story they want: "We found a key in a repo — now what?" — rotate immediately, revoke the old key, check audit logs for misuse, then add Secret Manager + a pre-commit hook (like gitleaks) so it can't recur.

SRE & Operations

7. SLI, SLO, SLA — define them without hand-waving

Answer: SLI (indicator) = the measurement — e.g., the measured proportion of requests succeeding under 300 ms. SLO (objective) = the target on the SLI — e.g., 99.9% over 30 days. SLA (agreement) = the contractual commitment to the customer, with consequences — usually looser than the SLO. The gap between 100% and the SLO is the error budget — and the SRE insight interviewers test: teams often use an error-budget policy to slow or pause risky launches when reliability burns beyond the agreed budget, and prioritize reliability work instead. That's the whole philosophy in one sentence.

8. How do you monitor on GCP — and what do you alert on?

Answer: Cloud Monitoring (metrics, dashboards, alerting), Cloud Logging (centralized logs with Log Router sinks to BigQuery/GCS/Pub/Sub), Cloud Trace (distributed tracing), Cloud Profiler (production profiling). The alerting wisdom they test: page primarily on user-visible symptoms and SLO impact — "checkout success rate dropped" (SLI burn), not "CPU is 80%". Infrastructure signals (disk saturation, queue exhaustion, certificate expiry) belong as diagnostic alerts when they're actionable and genuinely worth proactive paging. Every alert needs a runbook link, and alerts should be actionable — if nobody would do anything different at 3 AM, it's not a page, it's a ticket.

9. Your error budget is burning — walk me through the incident

Answer: they want the incident-response shape: detect (SLO burn-rate alert fires — multiwindow alerts catch both fast and slow burns), triage (declare incident, assign roles: commander, comms, ops), mitigate first (roll back, fail over, shed load — restore the SLI before root-causing), resolve, then blameless postmortem (timeline, contributing factors, action items with owners). The sentence that scores: "Triage enough to choose a safe mitigation, then restore service before deep root-cause analysis — the SLI comes first, the root cause comes second."

DevOps & Cost

10. How do you ship safely — deployment strategies on GCP?

Answer: common deployment strategies:

  • Rolling (default) — replace instances/pods gradually. Cheap, but a bad build still reaches all users eventually.
  • Blue/green — two full environments, flip traffic. Instant rollback, double cost during switch.
  • Canary — route a small % of traffic to the new version, watch SLIs, then ramp. Useful when you want to expose a new version gradually and judge it against production signals before broader rollout.

On GCP: Cloud Deploy orchestrates this (with canary rollout phases and automation), Cloud Build builds and tests, Artifact Registry stores images. The follow-up: "How do you automate the canary decision?" — automated progressive delivery can use monitoring and SLO checks as promotion gates; if the health criteria fail, stop or roll back the rollout. The gates need real verification logic behind them — the deploy tool doesn't invent the health criteria.

11. How do you keep the GCP bill under control?

Answer: the cost playbook interviewers expect:

  • Committed use discounts — 1- or 3-year commitments that can materially reduce eligible steady-state compute cost; the exact discount depends on resource and commitment type, so don't memorize a single number.
  • Spot VMs — up to ~91% off for fault-tolerant batch.
  • Right-sizing — Recommender + GKE cost dashboards; kill idle IPs, disks, and LBs (idle forwarding rules still bill).
  • BigQuery — partitioning, clustering, no SELECT *, reservations for steady load.
  • Budgets + alerts — billing budgets with alert thresholds for visibility. Note: normal budget alerts don't cap spending automatically. You can attach automation to budget notifications for selected cases, but shutdown actions need careful safeguards — disabling billing can take services down and risk resource loss.
  • Labels — label everything by team/env so you can attribute spend (the unglamorous answer that's always correct).

The trap: "We got a surprise $10k bill — what happened?" — start with the billing breakdown by service, project, SKU, and label, find the delta, then investigate the responsible resource. Common things to check first: sudden BigQuery scan volume, egress, idle or forgotten accelerators and VMs, storage growth, and network/NAT costs. Prevention is budgets + labels + org policies.

In this series

  1. Top 20 GCP Interview Questions and Answers (2026 Edition) — start here.
  2. GCP Interview Questions: Compute, Storage, and Networking Fundamentals — pick the right compute, storage, and network.
  3. GCP Data Engineering Interview Questions: BigQuery, Dataflow, Pub/Sub — analytics at scale.
  4. Kubernetes and GKE Interview Questions — from kubectl to 3 AM debugging.
  5. GCP IAM, Security, and SRE Interview Questions (this post).

Related: System Design Interviews: A Practical Primer — the 4-step framework these cloud questions plug into.

Comments

Popular posts from this blog

Java Banking Finance Services and Insurance (BFSI) domain interview questions

JSP Servlet Interview Questions For Freshers Series 1

Java program to check even or odd number