Skip to main content

Command Palette

Search for a command to run...

Building a Simple ETL Pipeline with Docker, MinIO, and Airflow

Updated
•3 min read•View as Markdown

1-How my idea started.

Initially, I wanted to practice working with AWS S3, but I worried about unexpected cloud costs if I couldn't control the billing properly. I previously created a free tier account, but unfortunately it got locked ambiguously. Since credit cards can only be used for one-time free credits, I needed another solution. Luckily, I discovered MinIO through this guide - it's very easy to set up using Docker Compose, and all the steps are clear to follow.

Read the full guide on Medium by Abhishek Jain

In the past, whenever I wanted to build a project using Docker, it would take me several days to resolve library and version issues. And I wasn't always lucky, sometimes my efforts failed completely, which made me want to give up many times. But today, with the help of AI, fixing these configuration issues is much easier. I just let the AI read the error messages and wait for it to help me fix them until everything works. :D

After setting these things, my components look like this:

MinIO holding 3 Parquet files migrated from CSV:

Connected to PostgreSQL using DBeaver to verify data ingestion into the RDBMS:

2- Then airflow was added.

To demonstrate an ETL process, I decided to add Airflow by including the Airflow component in my docker-compose.yml file (I will provide the full file at the end of this post). I also added the DAGs and the etl-manager files, all written in Python with suport from Copilot. Now, whenever I append new data, I don't have to start from the main function manually, I can trigger it right from Airflow with a manual start or set it on a schedule.

The generated code met about 90,99% of my requirements, but I had to tweak a few quirks along the way. Don't trust AI-generated code completely.

DAG graph displayed on the Airflow web server (commonly accessible on port 8080):

3- Adding More Dynamic Input.

CSV files are good for a demo, but commonly our input data will come from dynamic sources like a live database or an API. For the scope of my personal storage project, I decided to add an API data source using a free API from Stack Overflow - Question's data. I wrote another DAG file to handle this case.

Fetching data from the API:

Triggering the run through the Airflow:

4 - Add a very simple dashboard.

I chose Metabase to demonstrate dashboard connectivity and adapt end-to-end functionality, just loading the data from Stack Exchange. Maybe in the future, I will design a proper data model and metrics to make the dashboard more meaningful.

Metabase chart looks like it will be more than a static table once we can build a data model with several tables.

I tried some SQL aggregate functions and created this simple chart.

Now I'm thinking about what component to add next...

1.8K views