Simplify, streamline, and scale your data operations with data pipelines built on Apache Airflow
Data Pipelines with Apache Airflow has empowered thousands of data engineers to build more successful data platforms. This new second edition has been fully revised for Airflow 3 with coverage of all the latest features of Apache Airflow, including the Taskflow API, deferrable operators, and Large Language Model integration. Filled with real-world scenarios and examples, you'll be carefully guided from Airflow novice to expert.
In
Data Pipelines with Apache Airflow, Second Edition you'll learn how to:
- Master the core concepts of Airflow architecture and workflow design
- Schedule data pipelines using the Dataset API and time tables, including complex irregular schedules
- Develop custom Airflow components for your specific needs
- Implement comprehensive testing strategies for your pipelines
- Apply industry best practices for building and maintaining Airflow workflows
- Deploy and operate Airflow in production environments
- Orchestrate workflows in container-native environments
- Build and deploy Machine Learning and Generative AI models using Airflow
Using real-world scenarios and examples,
Data Pipelines with Apache Airflow, Second Edition teaches you how to simplify and automate data pipelines, reduce operational overhead, and smoothly integrate all the technologies in your stack. Part reference and part tutorial, each technique is illustrated with engaging hands-on examples, from training machine learning models for generative AI to optimizing delivery routes.