JK

Jamel Khan

Ouvert aux possibilités

Senior Data Engineer | Cloud Data Platform Architect | Real-Time & AI-Driven Systems

Toronto, ON

Optable
Not stated

À propos

Senior Data Engineer with 8+ years of experience architecting and scaling cloud-native data platforms and lakehouse solutions that support multi-terabyte data processing, advanced analytics, and real-time intelligence. Skilled in building high-performance ETL/ELT pipelines with dbt, Airflow, Fivetran, and PySpark, and in designing streaming architectures using Kafka and Flink across AWS, Azure, and GCP. Experienced in modernizing data ecosystems, enabling machine learning workflows, and delivering secure, compliant solutions aligned with HIPAA, FHIR, SOC 2, and GDPR standards. Hands-on with feature stores, real-time inference pipelines, and vector-based architectures (Pinecone, FAISS) supporting AI-driven applications. Recognized for optimizing performance, reliability, and cost efficiency while leading engineering initiatives and mentoring teams to deliver scalable, high-impact data systems across healthcare, financial services, and enterprise domains.

Compétences

  • Cross-Functional Collaboration
  • Docker
  • Java
  • Power BI
  • Python
  • API REST
  • Tableau
  • Terraform

Expérience

  1. Lead Data Engineer

    Optable

    janv. 2023 to Aujourd’hui

    • Led the design and operation of large-scale, cloud-native data platforms and lakehouse architectures on Azure and AWS, processing over 10M+ healthcare records daily while maintaining full compliance with HIPAA, FHIR, HL7, SOC 2, and GDPR standards. • Engineered high-throughput real-time data pipelines using Kafka, NiFi, and Flink, significantly improving data freshness and reducing latency by 50%, enabling near real-time healthcare analytics and insights. • Developed scalable ETL/ELT frameworks leveraging dbt, Airflow, and Fivetran to support machine learning feature stores, AI-driven pipelines, operational reporting, and self-service analytics, reducing pipeline execution time by 45%. • Implemented comprehensive data governance, security, and observability frameworks, including IAM, encryption, tokenization, Unity Catalog, and data monitoring tools, ensuring audit readiness and enterprise-grade compliance. • Collaborated with cross-functional stakeholders to design domain-driven data pipelines, standardize schemas, and enforce data contracts, improving data consistency by 40% and enhancing platform reliability. • Led and mentored a team of 5 engineers, driving best practices in cloud optimization, observability, and MLOps, and fostering a high-performance, compliance-focused engineering culture.

  2. Senior Data Engineer

    Qohash

    juin 2019 to déc. 2022

    • Designed and scaled high-volume ETL/ELT pipelines using Python, Spark, dbt, and SQL to process multi-terabyte datasets, increasing data throughput by 60% and enabling reliable, high-performance analytics. • Built resilient real-time data ingestion frameworks leveraging Kafka, Flume, and Debezium to support event-driven architectures, reducing data availability lag from hours to near real-time. • Developed optimized data warehouse models (star schema, OLAP) to support BI and reporting workloads, improving query performance by up to 70%. • Collaborated closely with data science teams to operationalize machine learning pipelines, including feature engineering, real-time inference, and AI-driven workflows for fraud detection and predictive analytics. • Automated infrastructure and deployment processes using Terraform, GitHub Actions, Docker, and Kubernetes, accelerating delivery cycles by 40% while maintaining secure and compliant environments. • Played a key role in data platform strategy, contributing to architecture decisions, governance standards, and cost optimization initiatives to support rapidly growing data volumes and long-term scalability.

  3. ETL & BI Engineer

    BlueDot

    août 2017 to mai 2019

    • Built and managed automated ETL pipelines using Python, SQL, and Airflow to integrate data from multiple sources, reducing manual effort by 30% and improving data availability. • Designed and enhanced data warehouse schemas (star models) to support analytics workloads, improving query efficiency and reducing reporting latency by 50%. • Partnered with business stakeholders to develop interactive dashboards and reports, enabling data-driven insights and improving cross-functional visibility. • Implemented robust data quality and validation frameworks, ensuring consistency, accuracy, and reliability across enterprise reporting systems. • Supported pipeline orchestration, scheduling, and documentation efforts while collaborating with senior engineers, gaining hands-on experience in cloud-based data platforms and large-scale data operations.

Formation

  1. Not stated

    Bachelor's, Computer Science