Parth Mundhwa
Ouvert aux possibilitésBig Data Engineer | Data Engineer | Analytics Engineer
London, ON
À propos
Data Engineer with 5+ years of experience building and optimizing batch and real-time pipelines using PySpark, Kafka, and SQL across AWS and Azure environments. Processes large-scale datasets (10M+ records daily), improving performance, data quality, and delivery for analytics and machine learning. Experienced in ETL/ELT development, Airflow orchestration, and dbt-based transformations. Supports real-time analytics, reporting, and scalable data platforms with a focus on reliability, efficiency, and production-ready systems.
Compétences
- CI/CD
- Docker
- Git
- Kubernetes
- MySQL
- PostgreSQL
- Python
- Terraform
Expérience
Data Engineer
Softchoice
mars 2025 to Aujourd’hui
Ontario, Canada
• Built batch and streaming pipelines using PySpark and Kafka, processing multi-source datasets, improving data availability by 30 percent, and enabling near real-time reporting across client-facing analytics platforms. • Designed ELT workflows using SQL and dbt, transforming raw data into curated datasets, reducing data preparation time by 35 percent and improving reporting consistency across business teams. • Improved Spark job performance by optimizing partitioning, caching, and execution plans, reducing processing time by 25 percent and lowering cloud compute costs across distributed workloads. • Developed curated data models supporting analytics and BI use cases, improving dashboard performance by 20 percent and enabling faster access to business-critical insights. • Implemented validation rules, anomaly checks, and monitoring, reducing data inconsistencies by 40 percent and improving trust in datasets used for reporting and analytics. • Orchestrated workflows using Airflow with scheduling, retries, and alerting, improving pipeline reliability and ensuring consistent execution across batch and streaming data pipelines.
Data Engineer
OmniMD
mai 2023 to déc. 2023
Gujarat, India
• Processed over 1M healthcare records daily using PySpark pipelines, reducing reporting delays by 35 percent and improving data availability across clinical and operational analytics systems. • Reduced pipeline execution time by 30 percent by converting ETL workflows into SQL-driven ELT pipelines, improving transformation efficiency and data delivery timelines. • Improved machine learning model readiness by 25 percent by delivering standardized, feature-ready datasets, reducing preprocessing effort and enabling faster experimentation cycles. • Decreased pipeline failures by 40 percent through schema validation, null checks, and data quality enforcement, ensuring accurate and reliable datasets across production environments. • Standardized ingestion processes across multiple healthcare data sources, reducing manual effort and improving data consistency across pipelines. • Built structured datasets optimized for reporting, improving query performance and enabling faster access for analytics and business intelligence teams.
Big Data Engineer - II
Rysun Labs
janv. 2019 to déc. 2022
Gujarat, India
• Processed 10M+ records per day using distributed Spark pipelines, improving throughput and reducing duplicate data by 90 percent, increasing accuracy across analytics and reporting systems. • Reduced data latency by 15 percent by implementing Kafka-based streaming pipelines, enabling faster processing of real-time data for dashboards and monitoring applications. • Achieved sub-second query response times by designing a distributed data platform, improving data retrieval speed for large-scale analytics workloads. • Increased data availability by 40 percent by integrating over 600 data sources, enabling centralized access and improving reporting capabilities across multiple business domains. • Improved pipeline uptime to 99.5 percent by implementing Airflow-based scheduling, monitoring, and retry mechanisms, ensuring consistent data delivery across batch and streaming systems. • Reduced processing time by optimizing Spark jobs through partitioning and tuning, improving efficiency across high-volume data workloads. • Built reusable pipeline components, reducing development effort for new workflows and improving consistency across engineering processes. • Collaborated with analysts and engineering teams to deliver scalable data solutions aligned with reporting and operational requirements.
Formation
Fanshawe College
Postgraduate Degree, Artificial Intelligence and Machine Learning
Canada
2024 to 2025
Sal College Of Engineering (113)
Bachelor's degree, Computer Engineering
India
2019 to 2022