Available for New Roles

Hi, I'm Sourabh Savre

I build robust 

Hands-on Data Engineer with experience building scalable ETL pipelines, orchestrating workflows with Airflow/Delta Live Tables, and implementing high-efficiency Medallion architectures in cloud environments.

About Me

I am a results-oriented Data Engineer who loves turning massive volumes of raw, unstructured data into clean, highly-optimized pipelines. My engineering journey is centered around designing automated architectures that handle high-velocity datasets without breaking.

Having built end-to-end data systems at scale, I specialize in leveraging PySpark, Databricks, and modern cloud infrastructures to deliver analytics-ready datasets. I approach data engineering not just as a support system, but as a critical driver for business analytics and machine learning applications.

Data Engineering Philosophy

"A great data pipeline is invisible—it must be modular, idempotent, and self-healing. I build with the conviction that data quality is absolute, latency should be minimized proactively, and infrastructure must scale dynamically with business needs."

Robust Pipelines

Fault-tolerant ETL workflows built using Airflow and Spark Streaming.

Optimized Compute

Performance tuning via partition pruning, broadcasting, and Z-Ordering.

Modern Storage

Delta Lake architectures enforcing ACID transactions and schema compliance.

Data Quality

SCD Type 2 schemas and comprehensive validation tests.

Technical Skills

Core technologies and architectures I use to design, build, and optimize scalable data systems.

Languages

PythonSQLScala (Basics)

Big Data

PySparkApache SparkDelta LakeDatabricks

Cloud Platforms

AWS (S3, Glue, EMR)Azure (ADF, Synapse)GCP (BigQuery, Dataflow)

Data Engineering

ETL PipelinesData ModelingData WarehousingApache Airflow

DevOps & Tools

DockerKubernetesGitGitHubCI/CD

Architecture

Medallion ArchitectureStar SchemaSnowflake SchemaSCD Type 2

Work Experience

Data Engineer

Active Role

Smallest.ai

Sep 2025 – PresentRemote / Bengaluru, India
PySparkApache SparkDatabricksDelta LakeDelta Live Tables (DLT)Apache Airflow
  • Designed and developed scalable ETL pipelines using PySpark and Apache Spark to process large-scale datasets efficiently.
  • Built and maintained data pipelines on Databricks leveraging Delta Lake for reliable, ACID-compliant data storage and processing.
  • Implemented Medallion Architecture (Bronze-Silver-Gold layers) to ensure clean, structured, analytics-ready data.
  • Orchestrated data workflows using Delta Live Tables (DLT) and Apache Airflow for automated, fault-tolerant pipeline execution.
  • Collaborated with data scientists and analysts to deliver high-quality datasets for ML model training and business reporting.

Featured Projects

Case studies of pipelines and systems built with an emphasis on scale, performance, and clean design.

Real-Time Data Pipeline

Jan 2025 – Apr 2025
PySparkDatabricksDelta LakeAWS S3Apache Airflow
Problem / Context

High ingestion latency and lack of transactional guarantees for event-driven streaming datasets, causing delays in analytical reports.

What was built

Built a multi-tier streaming pipeline ingesting raw events into a Bronze layer, applying cleanups/validations in Silver, and aggregating business metrics into a Gold layer.

Quantified Impact

Significantly reduced latency by adopting incremental loads and applying Z-ORDER cluster-key optimization on Delta tables for faster analytical queries.

View RepositoryCase Study

Cloud Data Warehouse Migration

Aug 2024 – Dec 2024
AWS GlueRedshiftPythonSQLPySpark
Problem / Context

Slow query processing times and costly maintenance of legacy, on-premise database warehouse systems with complex schema transformations.

What was built

Migrated the database to AWS Redshift, redesigned schemas as Star Schemas, and implemented SCD Type 2 patterns to track historical changes and automate deduplication.

Quantified Impact

Reduced query execution times by 40% through partitioning strategies, distribution keys, and broadcast join optimizations in Spark workloads.

View RepositoryCase Study

Crop Price Predictor

Mar 2024 – Jun 2024
PythonPandasScikit-learnNumPyEDA
Problem / Context

Inability for market players to predict crop price fluctuations due to scattered historic files and high seasonal price variance.

What was built

Conducted extensive EDA and feature engineering using Pandas and NumPy, then trained, cross-validated, and optimized multiple machine learning regression models.

Quantified Impact

Delivered a high-precision forecasting tool, choosing the best-performing regression algorithm to accurately predict future price trends.

View RepositoryCase Study

Architecture Showcase

Interactive representation of the Medallion Architecture (Bronze-Silver-Gold) pipelines I build for data governance and query performance.

Click on any stage above to inspect schema formats, processing techniques, and table configurations.

Licenses & Certifications

AWS Certified Data Engineer - Associate
Completed

AWS Certified Data Engineer - Associate

Amazon Web Services (AWS)

Date:Jun 2025
Databricks Certified Associate Developer for Apache Spark
Completed

Databricks Certified Associate Developer for Apache Spark

Databricks

Date:Dec 2024
Python for Data Engineering
Completed

Python for Data Engineering

Coursera

Date:Jul 2024

Education

B.Tech, Computer Science and Engineering

Indore Institute of Science and Technology

2023 – 2027Indore, India
Key Focus & Coursework
Data Structures & Algorithms
Database Management Systems (DBMS)
Object-Oriented Programming (OOP)
Operating Systems & Networks

Get in Touch

Have a question or want to discuss scaling a data pipeline? Drop me a message below or contact me directly.

Notice

If the automated form API fails due to system blocks, the submission defaults back to opening your local mail client.