Data Engineering Interview Guide 2025-2026: PySpark, Kafka, and Airflow
Data engineering is one of the hottest career paths in India right now — and the interviews reflect that. Companies building modern data platforms probe your understanding of distributed processing, streaming architectures, pipeline reliability, and cloud data services depth. This guide covers everything you'll face in a senior data engineering interview: PySpark internals, Kafka design, Airflow patterns, data modelling, and cloud data platform questions — with the scenario-based answers that impress senior interviewers.
What Senior Data Engineering Interviews Really Test
Entry-level data engineering interviews test PySpark syntax and basic SQL. Senior interviews are fundamentally different — they test your system design instincts for data pipelines, your understanding of failure modes and recovery, your knowledge of distributed system tradeoffs (throughput vs latency, exactly-once vs at-least-once), and your ability to diagnose and fix data quality issues.
Companies like Flipkart, Amazon (data engineering roles), Juspay, PhonePe, ThoughtWorks, and Accenture's data practices run extensive data engineering interviews that cover the full modern data stack. This guide prepares you for all of it.
Data Engineering Interview Preparation Plan: 5 Weeks
5-Week Data Engineering Interview Roadmap
Round 1: PySpark and the Spark Execution Model
PySpark is the most tested area in data engineering interviews. You need to go beyond syntax to the execution model.
Spark Architecture: What Interviewers Probe
- Explain the difference between a Spark driver and executors. What happens if the driver dies vs an executor dies?
- What is the difference between a transformation and an action in Spark? Give 5 examples of each.
- Explain the Spark DAG (Directed Acyclic Graph). How does Spark use it for optimization?
- What triggers a shuffle? Why is shuffle expensive and how do you minimize it?
- What is the difference between repartition() and coalesce()? When would you use each?
PySpark Performance Optimization (Critical)
- Broadcast joins: When to use broadcast hints for small dimension tables. What threshold triggers auto-broadcast?
- Data skew: How to detect partitions with far more data than others. Solutions: salting, custom partitioner, AQE (Adaptive Query Execution).
- Caching strategy: When to use persist() vs cache(). Storage levels (MEMORY_ONLY, MEMORY_AND_DISK). When NOT to cache.
- File formats: Why Parquet over CSV for analytics. Columnar storage, predicate pushdown, column pruning.
Round 2: Apache Kafka Architecture and Design
Kafka is tested both for message queue internals and streaming pipeline design. Know both.
Kafka Internals (Frequently Asked)
- Explain Kafka's architecture: brokers, topics, partitions, consumer groups, and the role of ZooKeeper / KRaft in Kafka 3+.
- What is the difference between at-least-once, at-most-once, and exactly-once delivery semantics in Kafka?
- How does consumer group rebalancing work? What is the impact of rebalancing on a streaming pipeline?
- What is log compaction? When would you enable it on a Kafka topic?
- How do you handle message ordering in Kafka? What are the guarantees within a partition vs across partitions?
Kafka Design Question: Real-Time Order Events
"Design a Kafka-based pipeline to process 1 million order events per minute for real-time analytics and fraud detection."
- Producers publish to orders topic with order_id as partition key (ensures events for same order go to same partition)
- Multiple partitions (e.g., 50) for parallelism, replication factor 3 for durability
- Two consumer groups: one for real-time fraud detection (Spark Structured Streaming), one for analytics aggregation
- Exactly-once with idempotent producers + transactional consumers
- Kafka Connect sink to write to data warehouse (Snowflake/BigQuery) and OLTP store for fraud alerts
Round 3: Airflow Pipeline Design and Best Practices
Airflow questions in data engineering interviews test both your ability to design reliable pipelines and your understanding of Airflow's execution model.
Critical Airflow Concepts
- DAG design: Keep DAGs idempotent — running the same DAG run twice should produce the same result. Use execution_date, not hardcoded timestamps.
- Sensors vs Operators: Use sensors for external dependency detection (S3KeySensor, ExternalTaskSensor). Understand poke vs reschedule mode and their resource implications.
- XComs: For passing small data between tasks. Why XComs are not for large data — overhead on the metadata database.
- Task dependencies: set_upstream, set_downstream, BranchPythonOperator for conditional pipelines, ShortCircuitOperator.
- Backfill: What is it? When is it needed? The importance of idempotent tasks for safe backfill.
Round 4: Data Modelling for Data Warehouses
Data modelling is tested at every senior level. Know the tradeoffs, not just the definitions.
- Star schema: Central fact table with denormalized dimensions. Fast analytical queries, simple joins. Ideal for BI/reporting.
- Snowflake schema: Normalized dimensions. More tables, more joins. Saves storage, can hurt query performance.
- Slowly Changing Dimensions (SCD): Type 1 (overwrite), Type 2 (add row with effective dates), Type 3 (add column for previous value). Know when each is appropriate and how Delta Lake simplifies SCD Type 2.
- Data Vault: Hub (business key), Link (relationships), Satellite (descriptive attributes). For highly volatile schemas in enterprise data warehouses.
Round 5: Cloud Data Platforms
- Databricks: Delta Lake format (ACID transactions, time travel, schema enforcement), Unity Catalog, auto-scaling clusters, Delta Live Tables for declarative pipelines.
- AWS Glue: Serverless ETL, Glue Data Catalog as central metadata store, DynamicFrames vs DataFrames, job bookmarks for incremental processing.
- BigQuery: Serverless, columnar, partitioned and clustered tables, cost optimization with partition expiry and slot reservations.
- Snowflake: Virtual warehouses, time travel, zero-copy cloning, data sharing, stream and task for CDC patterns.
Salary Expectations for Data Engineers in India (2025-2026)
- 2–4 years, PySpark + SQL: ₹10L–₹22L in service companies, ₹14L–₹26L in product companies
- 4–7 years, Kafka + Airflow + cloud: ₹24L–₹45L for senior data engineering
- 7+ years, data platform architecture: ₹48L–₹80L at product companies and consulting firms
- Remote US roles: $110K–$175K for Senior Data Engineer / Data Platform Engineer
Get Expert Data Engineering Interview Coaching
RVK Tech's data engineering specialists coach on PySpark, Kafka, Airflow, and cloud data platforms. We run live mock interviews with real questions from Amazon, Flipkart, PhonePe, and consulting firms. Get the edge that lands offers.