Spark Pills
A series for data engineers who can run PySpark jobs but want to understand what happens underneath. Each pill starts from a real question and each has a quiz to check what stuck.
- 00
Spark Pill 0: Why Spark isn't Pandas
Have you ever wondered why you can't just load 80GB into Pandas? And if Spark only has 48GB of RAM total, why doesn't it crash too? This is question a lot us had as Junior Data Engineer. In this pill post We are going to see What Spark is about
- 01
Spark Pill 1: What The Hell Is MapReduce and Why Does Spark Still Use It?
Have you ever wondered what actually happens when Spark moves data between machines? Let's open the MapReduce box and see who the mappers and reducers really are.
- 02
Spark Pill 2: How Does the Spark DAG really work? And Why it doesn't execute line by line
Have you ever wondered why Spark waits until you call .show() before reading a single byte? Let's look at lazy evaluation, the DAG, and the Catalyst optimizer.
- 03
Spark Pill 3: Shuffle: Moving data across the Network
Have you wondered what happens physically when Spark redistributes data across machines? Let's trace the bytes through the sort-based shuffle, from serialization to TCP transfer.
- 04
Spark Pill 4: The Computer Science behind Broadcast or Sort-Merge Joins
Have you ever wondered why one join finishes in 30 seconds and another identical-looking join takes 90 minutes? The answer lies in which join strategy Spark picks, and whether you can eliminate the shuffle entirely.
- 05
Spark Pill 5: Data Skew: The Eternal Enemy of Spark
Have you ever wondered why your Spark job uses 10 executors but one task takes 50x longer than the rest? The culprit is Data Skew, and it can turn a 5-minute job into a 2-hour timeout.
- 06
Spark Pill 6: Cache, Persist, or Checkpoint? The Lineage Trap
Have you ever wondered why your Spark job spends 66 minutes planning and only 26 seconds executing? The answer involves two layers most engineers never separate: the Driver's planning and the Executors' execution.
- 07
Spark Pill 7: Learn to Read the Spark UI or You're Flying Blind
Have you ever wondered what all those numbers in the Spark UI actually mean? Jobs, stages, tasks, shuffle bytes, spill, GC time. Let's decode them so you can diagnose problems in minutes instead of guessing for hours.
- 08
Spark Pill 8: Iceberg Is Not Free: What Nobody Tells You About the Migration
Have you ever wondered why your Iceberg table gets slower over time even though the data hasn't grown? Or why a simple DELETE takes longer than rewriting the entire table? Let's look at the pitfalls that only show up in production.
- 09
Spark Pill 9: Why Lambda Doesn't Work for Streaming: The State Problem
Have you ever wondered why people keep saying 'just use Lambda' for real-time processing, and then teams end up rewriting everything in Flink? The answer comes down to one word: state.