Apache Airflow
Data · Automate
Schedule dependent data tasks and practise retries, backfills and failure recovery.
Airflow/dbt/Spark basics: extract, transform, load into a warehouse with tests and lineage.
Explore 3 tools for this projectShare of job postings in India, per role, that name this capability.
Needs first: Query and model data with SQL, Write production-quality Python for AI work
Start with one tool for each part of your project. You don’t need to learn them all.
3 tools to explore
Data · Automate
Schedule dependent data tasks and practise retries, backfills and failure recovery.
Data · Test
Turn SQL transformations into documented, tested data models.
Data
Query datasets in a cloud warehouse and inspect query cost and performance.
Use the NYC TLC trip-record parquet files as a stand-in daily feed: split one month into day-sized partitions and land them one at a time so the pipeline sees a real arriving feed. Build an Airflow DAG that loads each day incrementally into local Postgres, transforms it with dbt staging and mart models, and runs dbt tests for nulls, uniqueness and referential integrity. Then break it on purpose — feed a day with corrupted rows and confirm the run fails loudly instead of quietly writing bad marts. Finish by backfilling a day you deliberately skipped.
Built a daily Airflow + dbt pipeline with incremental idempotent loads and four data tests — demonstrated it catching a corrupted day's feed before it reached the marts, and backfilling a missed date cleanly.