Building a Simple ETL Pipeline with Docker, MinIO, and Airflow
1-How my idea started.
Initially, I wanted to practice working with AWS S3, but I worried about unexpected cloud costs if I couldn't control the billing properly. I previously created a free tier account, but unfortunately it got locked ambiguously. Since credit cards can only be used for one-time free credits, I needed another solution. Luckily, I discovered MinIO through this guide - it's very easy to set up using Docker Compose, and all the steps are clear to follow.
Read the full guide on Medium by Abhishek Jain
In the past, whenever I wanted to build a project using Docker, it would take me several days to resolve library and version issues. And I wasn't always lucky, sometimes my efforts failed completely, which made me want to give up many times. But today, with the help of AI, fixing these configuration issues is much easier. I just let the AI read the error messages and wait for it to help me fix them until everything works. :D
After setting these things, my components look like this:
MinIO holding 3 Parquet files migrated from CSV:
Connected to PostgreSQL using DBeaver to verify data ingestion into the RDBMS:
2- Then airflow was added.
To demonstrate an ETL process, I decided to add Airflow by including the Airflow component in my docker-compose.yml file (I will provide the full file at the end of this post). I also added the DAGs and the etl-manager files, all written in Python with suport from Copilot. Now, whenever I append new data, I don't have to start from the main function manually, I can trigger it right from Airflow with a manual start or set it on a schedule.
The generated code met about 90,99% of my requirements, but I had to tweak a few quirks along the way. Don't trust AI-generated code completely.
DAG graph displayed on the Airflow web server (commonly accessible on port 8080):
3- Adding More Dynamic Input.
CSV files are good for a demo, but commonly our input data will come from dynamic sources like a live database or an API. For the scope of my personal storage project, I decided to add an API data source using a free API from Stack Overflow - Question's data. I wrote another DAG file to handle this case.
Fetching data from the API:
Triggering the run through the Airflow:
4 - Add a very simple dashboard.
I chose Metabase to demonstrate dashboard connectivity and adapt end-to-end functionality, just loading the data from Stack Exchange. Maybe in the future, I will design a proper data model and metrics to make the dashboard more meaningful.
Metabase chart looks like it will be more than a static table once we can build a data model with several tables.
I tried some SQL aggregate functions and created this simple chart.
Now I'm thinking about what component to add next...