GCP Data Engineering Interview Questions: BigQuery, Dataflow, Pub/Sub

Three services show up repeatedly in GCP data-engineering interviews: BigQuery for analytics, Pub/Sub for messaging, and Dataflow for managed data processing. These are the questions interviewers actually ask — each with the crisp answer and the cost or correctness trap underneath. (Real interviews can also cover Cloud Storage, Dataproc, Composer, IAM, SQL/data modeling, and system design — the framework below works for all of them.)

The interview framework: for every architecture question, answer in this order — requirement → service choice → guarantee → failure mode → cost. Example: "We need events visible within one minute; at-least-once input is acceptable but duplicates aren't. I'd use Pub/Sub → Dataflow → BigQuery, dedupe on a stable event ID, monitor subscriber and Dataflow lag, DLQ malformed events, partition BigQuery by event date — and then validate whether continuous streaming cost is justified by the one-minute SLO." That beats "use Pub/Sub, Dataflow and BigQuery" every time.

BigQuery

1. How does BigQuery pricing work, and how do you keep a query bill from exploding?

Answer: BigQuery charges two ways: on-demand ($ per TB scanned) and capacity-based pricing (BigQuery editions with reserved slots for predictable workloads). Cost control is a favorite interview topic — and it has layers:

  • Before execution: dry runs (--dry_run) to preview bytes scanned; maximum bytes billed per query as a hard guardrail.
  • Governance: custom quotas and capacity reservations for predictable workloads.
  • Visibility: budgets and alerts — note these tell you spending crossed a threshold; they don't stop a running query. Don't confuse an alert with a kill switch.
  • Table design: partition when queries naturally filter on a coarse column such as date (partition pruning); cluster when queries repeatedly filter or aggregate on selected columns. They complement each other, but you don't need both by default — a table can benefit from clustering alone, and small tables may need neither.
  • Avoid SELECT * on large production tables when you only need a subset of columns — columnar storage means you pay per column scanned.
  • Materialize common aggregations instead of recomputing them.

The trap: "Partitioning vs clustering — which first?" — it's not a universal order. Partition for coarse pruning (dates), cluster for repeated fine-grained filtering. Know both, apply per workload.

2. Streaming inserts vs batch loads in BigQuery?

Answer: batch loads (load jobs from GCS) are free and preferred for bulk historical data. For streaming, know the two APIs separately: the legacy tabledata.insertAll streaming inserts (per-row cost, best-effort dedup within a window — duplicates are normal) and the newer Storage Write API for high-throughput ingestion. The Write API's committed streams can support exactly-once writes when the client supplies and manages offsets correctly — and exactly-once writes still isn't exactly-once business processing. Follow-up: "Your dashboard shows duplicate rows — why?" — streaming duplicates; dedupe with a MERGE on a stable event ID or design idempotent downstream logic.

3. What are slots, and when do you buy them?

Answer: a slot is a unit of BigQuery compute. On-demand gives you a shared pool with burst limits; reservations (BigQuery editions) organize capacity that can be assigned across projects, folders, and job types. Reservations give you controlled, predictable capacity rather than relying entirely on on-demand compute — though workloads can still queue depending on capacity, concurrency, and workload behavior. Buy them when: you have predictable heavy workloads, you need predictable capacity and cost, or workload analysis shows reservations beat on-demand for your pattern. The interview one-liner: "On-demand for spiky, reservations for steady."

4. How do you design an incremental, idempotent BigQuery pipeline?

Answer: the pattern interviewers want: land raw events in a staging table, dedupe by stable event ID, then MERGE into the target table. The pipeline is idempotent — reprocessing the same input produces the same output. The killer follow-up: "The Dataflow job restarts and sends yesterday's records again — how do you prevent duplicates?" — the MERGE keys on the event ID, so replays are absorbed. The lesson to say out loud: delivery guarantee ≠ business correctness — at-least-once transport plus idempotent writes equals correct results. (BigQuery ML — training models with SQL where the data lives — is worth knowing for ML-adjacent roles: use it for fast SQL-native baselines, Vertex AI when you need custom training, deployment, and MLOps.)

Pub/Sub

5. Pub/Sub concepts: topics, subscriptions, push vs pull

Answer: publishers send to a topic; subscriptions receive from it (many subscriptions per topic = fan-out). Pull = your subscriber polls; push = Pub/Sub POSTs to your HTTPS endpoint. Use push for Cloud Run/functions endpoints; pull for workers that want flow control (Dataflow, GCE consumers). Key guarantees: at-least-once delivery by default — duplicates happen, so consumers must be idempotent. (Pub/Sub also supports exactly-once delivery for supported pull-subscription scenarios, with specific constraints — know it exists, design for at-least-once.) Ordering is opt-in per ordering key: messages sharing a key are ordered, so a hot ordering key serializes onto one worker and becomes a throughput bottleneck — the classic trade-off question.

6. How do you handle a poison message / failing subscriber?

Answer: configure a dead-letter topic — after N delivery attempts, the message is forwarded there instead of redelivered forever. Set a sensible ack deadline — subscriber libraries can extend leases, but chronically slow processing usually signals an architecture problem (bigger workers, more parallelism, or smaller messages), not a deadline to lengthen. Use retry policies with backoff, and monitor oldest unacked message age — the metric that tells you a subscription is stuck. The interview follow-up: "Messages are piling up — what do you check?" — subscriber throughput vs publish rate, ack deadlines expiring, and whether ordering keys are serializing everything onto one worker.

7. Pub/Sub vs Kafka — when does the interviewer want which answer?

Answer: Pub/Sub when you want deeply managed, GCP-native messaging with minimal infrastructure operations — no brokers or partitions to manage, global by default, with retention and replay via seeking to timestamps/snapshots. Kafka when its durable-log model, consumer-offset semantics, ecosystem (Kafka Streams, Connect), portability, or your existing Kafka investment is central to the architecture. Don't make replay the differentiator — both can replay; the difference is operational model and ecosystem. Discuss transactions and exactly-once carefully per product rather than from memory. Interviewers probe for trade-off thinking, not loyalty.

Dataflow & Processing

8. Dataflow and Apache Beam — what's the mental model?

Answer: Beam is the programming model (PCollection, PTransform, pipelines); Dataflow is GCP's managed runner for Beam pipelines. Beam gives you one programming model for bounded and unbounded data — although streaming pipelines still require streaming-specific design (event time, state, checkpointing, sink choices). The four streaming concepts to recite cold:

  • Window: which events belong together? (fixed, sliding, session)
  • Watermark: how far through event time does the system believe it has progressed?
  • Trigger: when should results be emitted?
  • Allowed lateness: how long do we keep accepting late data into a window?

One-liner: "Beam is the SDK, Dataflow is the managed runner, and streaming is about deciding what to do with late data."

9. Batch vs streaming — how do you decide?

Answer: the decision is latency requirement vs cost/complexity: streaming (Pub/Sub → Dataflow → BigQuery) gives seconds-fresh data but costs more and forces you to handle late/out-of-order data; batch (scheduled Dataflow/Dataproc → BigQuery) is cheaper and simpler for hourly/daily freshness. The interview trap: "We need real-time dashboards" — ask how real-time: sub-second, under a minute, five minutes, or fifteen? Most "real-time" requirements turn out to be micro-batch, which is dramatically simpler to build and operate. A candidate who asks for the latency SLO sounds senior; one who immediately draws Pub/Sub → Dataflow does not.

10. Dataproc vs Dataflow?

Answer: Dataproc = managed Spark/Hadoop clusters — use when you have existing Spark code or need the Hadoop ecosystem. Dataflow = managed Beam — evaluate it for new pipelines, especially streaming-centric ones (new Spark workloads can still be entirely appropriate). The migration interview question: "We have Spark jobs — move to Dataflow?" — only if the rewrite pays off; otherwise Dataproc Serverless removes cluster management while keeping Spark.

11. How do you orchestrate it all?

Answer: Cloud Composer (managed Airflow) when you need Airflow semantics and ecosystem — operators, sensors, DAG scheduling, backfills, or an existing Airflow estate; Workflows for simple serverless orchestration (call services in sequence, no workers to manage); Cloud Scheduler for cron-style triggers. The follow-up: "Composer vs Workflows?" — Composer when you need Airflow's ecosystem (operators, backfills, rich UI); Workflows when the orchestration is just "call A, then B, then C."

12. Sketch a real-time analytics pipeline on GCP

Answer: a common GCP streaming analytics design: events → Pub/Sub topic → Dataflow streaming pipeline (window, dedupe, enrich) → BigQuery (partitioned, clustered) → Looker Studio / Looker dashboard. Dead-letter topic for bad events, Cloud Monitoring alerts on unacked age and Dataflow lag, budget alerts on BigQuery. Say the dead-letter and the alerting parts — that's what separates a textbook answer from an experienced one.

In this series

  1. Top 20 GCP Interview Questions and Answers (2026 Edition) — start here.
  2. GCP Interview Questions: Compute, Storage, and Networking Fundamentals — pick the right compute, storage, and network.
  3. GCP Data Engineering Interview Questions: BigQuery, Dataflow, Pub/Sub (this post).
  4. Kubernetes and GKE Interview Questions — from kubectl to 3 AM debugging.
  5. GCP IAM, Security, and SRE Interview Questions — identity, safety, and staying alive.

Related: System Design Interviews: A Practical Primer — the 4-step framework these cloud questions plug into.

Comments

Popular posts from this blog

Java Banking Finance Services and Insurance (BFSI) domain interview questions

JSP Servlet Interview Questions For Freshers Series 1

Java program to check even or odd number