Apache Airflow Core Repository
Open Source Contributor (Workflow Orchestration)
Contributed code modifications to the official core repository of Apache Airflow, standardizing provider testing infrastructure and CI/CD pipeline stability.
Data Engineer & Apache Airflow Contributor specializing in building scalable batch & real-time streaming pipelines, distributed processing architectures, and robust cloud data platforms.
I specialize in designing, scaling, and maintaining high-throughput data architectures and distributed processing platforms.
My focus lies in building resilient, automated data workflows. As an open-source contributor to Apache Airflow, I standardise provider testing frameworks (Snowflake, Teradata). I construct real-time streaming pipelines using Apache Kafka and Spark Structured Streaming, manage distributed compute clusters with Spark and Hadoop, and architect modern Data Lakehouses with Apache Iceberg.
Currently pursuing my M.Sc. in Data Science at IIIT Lucknow (2025–2027), I bridge the gap between raw unstructured data streams and production-ready analytical engines using Terraform, Docker, and AWS.
Open Source Contributor (Workflow Orchestration)
Contributed code modifications to the official core repository of Apache Airflow, standardizing provider testing infrastructure and CI/CD pipeline stability.
Kafka & Spark Structured Streaming
A production-grade streaming architecture processing live NYC taxi events at 100+ events/sec, calculating geospatial surge pricing, windowed aggregations, and dual storage persistence.
Hadoop Big Data Pipeline
An end-to-end distributed data pipeline performing ETL and trend analytics on high-volume website clickstream logs using Apache Flume, Pig, and Hive.
I am open to new data engineering opportunities, project collaborations, and discussions on distributed architectures.