← Back to Articles
  1. 00

    Spark Pill 0: Why Spark isn't Pandas

    Have you ever wondered why you can't just load 80GB into Pandas? And if Spark only has 48GB of RAM total, why doesn't it crash too? This is question a lot us had as Junior Data Engineer. In this pill post We are going to see What Spark is about

    SparkPandasDistributed ComputingMapReduce
  2. 01

    Spark Pill 1: What The Hell Is MapReduce and Why Does Spark Still Use It?

    Have you ever wondered what actually happens when Spark moves data between machines? Let's open the MapReduce box and see who the mappers and reducers really are.

    SparkMapReduceShuffleRDD
  3. 02

    Spark Pill 2: How Does the Spark DAG really work? And Why it doesn't execute line by line

    Have you ever wondered why Spark waits until you call .show() before reading a single byte? Let's look at lazy evaluation, the DAG, and the Catalyst optimizer.

    SparkLazy EvaluationDAGCatalyst
  4. 03

    Spark Pill 3: Shuffle: Moving data across the Network

    Have you wondered what happens physically when Spark redistributes data across machines? Let's trace the bytes through the sort-based shuffle, from serialization to TCP transfer.

    SparkShuffleNetwork I/ODisk I/O
  5. 04

    Spark Pill 4: The Computer Science behind Broadcast or Sort-Merge Joins

    Have you ever wondered why one join finishes in 30 seconds and another identical-looking join takes 90 minutes? The answer lies in which join strategy Spark picks, and whether you can eliminate the shuffle entirely.

    SparkJoinsBroadcast JoinSort-Merge Join
  6. 05

    Spark Pill 5: Data Skew: The Eternal Enemy of Spark

    Have you ever wondered why your Spark job uses 10 executors but one task takes 50x longer than the rest? The culprit is Data Skew, and it can turn a 5-minute job into a 2-hour timeout.

    SparkData SkewPerformanceSalting
  7. 06

    Spark Pill 6: Cache, Persist, or Checkpoint? The Lineage Trap

    Have you ever wondered why your Spark job spends 66 minutes planning and only 26 seconds executing? The answer involves two layers most engineers never separate: the Driver's planning and the Executors' execution.

    SparkCachePersistCheckpoint
  8. 07

    Spark Pill 7: Learn to Read the Spark UI or You're Flying Blind

    Have you ever wondered what all those numbers in the Spark UI actually mean? Jobs, stages, tasks, shuffle bytes, spill, GC time. Let's decode them so you can diagnose problems in minutes instead of guessing for hours.

    SparkSpark UIDebuggingPerformance
  9. 08

    Spark Pill 8: Iceberg Is Not Free: What Nobody Tells You About the Migration

    Have you ever wondered why your Iceberg table gets slower over time even though the data hasn't grown? Or why a simple DELETE takes longer than rewriting the entire table? Let's look at the pitfalls that only show up in production.

    SparkApache IcebergParquetData Lake
  10. 09

    Spark Pill 9: Why Lambda Doesn't Work for Streaming: The State Problem

    Have you ever wondered why people keep saying 'just use Lambda' for real-time processing, and then teams end up rewriting everything in Flink? The answer comes down to one word: state.

    SparkStreamingApache FlinkAWS Lambda