Airflow and Spark are one of the most common combinations in modern data platforms, but the orchestration decisions made on the Airflow side often determine whether Spark jobs run fast and cheap or slow and expensive. In this episode, Kenten is joined by Meni Shmueli, Co-Founder and CEO at DataFlint, to dig into how Airflow and Spark fit together, where teams go wrong, and how AI is changing the way they reason about cost and performance.
Key Takeaways:
- 00:00 Introduction.
- 01:40 Meni's background as a data engineer and what led him to start DataFlint, a production observability agent for Apache Spark.
- 02:53 Why the Airflow plus Spark combination is so common, and how each tool plays to its strengths.
- 04:23 How orchestration decisions in Airflow directly impact Spark performance and cost.
- 04:44 A customer story where parallelizing Airflow tasks made Spark jobs slower, less stable, and more expensive, and the opposite case where sequential runs left compute on the table.
- 07:02 The number one mistake teams make benchmarking pipelines: only looking at the Airflow side and ignoring underlying Spark cost and resource usage.
- 08:13 What DataFlint's Airflow and Astro integration gives teams, and how it brings production context into AI agents.
- 10:46 How AI is changing pipeline optimization, including holistic scheduling across hundreds of pipelines and connecting context from Spark, FinOps, and cloud.
- 12:21 A customer migration from Databricks to EMR that cut workflow costs by 80%, with examples of up to 100x optimizations.
- 14:23 Where the Airflow and Spark story could be better, including Spark Declarative Pipelines and tighter feedback between the two projects.
Resources Mentioned:
Thanks for listening to "The Data Flowcast: Mastering Apache Airflow® for Data Engineering and AI." If you enjoyed this episode, please leave a 5-star review to help get the word out about the show. And be sure to subscribe so you never miss any of the insightful conversations.
