Available for Data Engineering opportunities

Biplov Singh

Data Engineer & Apache Airflow Contributor specializing in building scalable batch & real-time streaming pipelines, distributed processing architectures, and robust cloud data platforms.

Airflow Open Source Contributor
100+ ev/s Real-Time Streaming Stack
IIIT Lucknow M.Sc. Data Science ('27)
01. About

Engineering Background

I specialize in designing, scaling, and maintaining high-throughput data architectures and distributed processing platforms.

My focus lies in building resilient, automated data workflows. As an open-source contributor to Apache Airflow, I standardise provider testing frameworks (Snowflake, Teradata). I construct real-time streaming pipelines using Apache Kafka and Spark Structured Streaming, manage distributed compute clusters with Spark and Hadoop, and architect modern Data Lakehouses with Apache Iceberg.

Currently pursuing my M.Sc. in Data Science at IIIT Lucknow (2025–2027), I bridge the gap between raw unstructured data streams and production-ready analytical engines using Terraform, Docker, and AWS.

Open Source Apache Airflow Contributor Refactored Snowflake & Teradata provider test suites
Education M.Sc. Data Science @ IIIT Lucknow Specialization: Distributed Big Data Systems
Location Jamshedpur, Jharkhand, India
Availability Open for Roles & Consultations
02. Stack

Technologies & Toolkit

Core Languages

Python SQL (PostgreSQL, MySQL) Bash / Shell

Orchestration & Streaming

Apache Airflow Apache Spark / PySpark Apache Kafka Hadoop (Hive, Pig, Flume)

Databases & Storage

Apache Iceberg / Parquet PostgreSQL & PostGIS Redis In-Memory Cache Great Expectations

Infrastructure & DevOps

Docker & Compose Terraform (IaC) / AWS Linux Administration Git / CI/CD (pytest)
03. Work

Open Source & Projects

OPEN SOURCE // 01 Apache Airflow

Apache Airflow Core Repository

Open Source Contributor (Workflow Orchestration)

Contributed code modifications to the official core repository of Apache Airflow, standardizing provider testing infrastructure and CI/CD pipeline stability.

Refactored provider test suites for Snowflake (#68898) and Teradata (#68190) modules.
Eliminated legacy unittest dependencies in favor of modern pytest conventions.
Validated pipeline stability across 80+ continuous integration checks on multi-platform runners.
PROJECT // 02 Streaming

Real-Time Ride Analytics Pipeline

Kafka & Spark Structured Streaming

A production-grade streaming architecture processing live NYC taxi events at 100+ events/sec, calculating geospatial surge pricing, windowed aggregations, and dual storage persistence.

Ingests high-frequency ride events using Apache Kafka brokers.
Applies 5-minute tumbling windows & 2-minute watermarks in Spark Structured Streaming.
Integrates PostGIS spatial geofencing for dynamic surge multipliers.
Dual storage sinks: Redis (hot caching) and S3 Apache Iceberg (lakehouse cold storage).
PROJECT // 03 Batch ETL

Website Clickstream Trend Analysis

Hadoop Big Data Pipeline

An end-to-end distributed data pipeline performing ETL and trend analytics on high-volume website clickstream logs using Apache Flume, Pig, and Hive.

Ingests clickstream log data continuously via Apache Flume agents.
Cleans raw logs and handles corrupted records using Apache Pig scripts.
Runs analytical queries and aggregates user session trends in Apache Hive.
Fully containerized distributed system environment managed via Docker.
04. Contact

Get In Touch

I am open to new data engineering opportunities, project collaborations, and discussions on distributed architectures.