Back to Talks
Rediscovering single-node processing: When does it make sense to move from Spark to Polars?

Rediscovering single-node processing: When does it make sense to move from Spark to Polars?

Jonas Böer

Date
Thursday, April 16, 2026
Time
3:45 PM - 4:15 PM
Room
Helium [3rd Floor]
Talk PyData: Data Handling & Data Engineering
Transcription

Apache Spark is the industry standard for big data processing, rightfully so. But for many data processing applications, a more light-weight solution will work just as well, avoiding Spark's compute and configuration overhead. Polars offers such a solution, with a fast single-node processing engine and a syntax that will pose no problems for experienced Spark developers. I will give a short comparison of Spark and Polars, where they have similarities and differences and show an implementation of a typical ETL and Feature Engineering task in both. I will compare the deployment, performance and cost of the two and, while giving my opinion on the topic, hope to enable you to also make an informed decision on when you want to use Polars and when to use Spark.