Senior Databricks Data Engineer
Offre en anglaisDesign and maintain enterprise-scale batch and streaming pipelines using a Medallion Architecture on a Databricks Lakehouse platform. Optimize Spark performance and implement data governance and security via Unity Catalog.
- Sur place
- Brampton, ON
- Publié 13 août 2026
- Postuler avant le 12 sept. 2026
- 1 poste
D’autres postes auxquels postuler directement
Des possibilités semblables publiées par des employeurs qui recrutent sur Jobs.ca, sans formulaire externe.
Forgeahead Solutions Corporation
Technical Lead and Senior Software Engineer
- Sur place
BC Public Schools
Manager, People, Performance and Culture
- Sur place
Alcohol and Gaming Commission of Ontario (AGCO)
Responsable de la gestion de l’information
- Sur place
Résumé du poste
Position Overview We are looking for a Senior, Super Hands-On Databricks Data Engineer who lives and breathes code, query optimization, and modern data architecture. In this role, you won't just design architectures on whiteboards—you will write production PySpark/SQL, optimize Databricks clusters, build streaming and batch pipelines, and enforce data governance. You will own end-to-end pipeline execution from raw ingestion to curated Gold layer models, playing a lead role in modernizing our Lakehouse platform. Key Responsibilities 1. Hands-On Pipeline Development & Lakehouse Architecture Design, build, and maintain enterprise-scale batch and real-time streaming pipelines using PySpark, SQL, Delta Live Tables (DLT), and Auto Loader. Implement and refine Medallion Architecture (Bronze Silver Gold) to support downstream BI, reporting, and Machine Learning workloads. Enforce schema evolution, ACID transactions, and data compaction using Delta Lake core constructs. 2. Performance Tuning & Optimization (Deep Tech) Diagnose and resolve Spark performance bottlenecks: data skew, OOM errors, excessive shufflings, and memory spills. Optimize queries using Liquid Clustering, Z-Ordering, Data Partitioning, AQE (Adaptive Query Execution), and Photon engine tuning. Benchmark and optimize Databricks compute workloads to minimize DBU (Databricks Unit) consumption and cloud costs (FinOps). 3. Governance, Security & Quality Implement end-to-end data governance, fine-grained access control (row/column-level security), and lineage tracking using Unity Catalog. Automate automated data quality validation checks and alert mechanisms across the pipeline life cycle. 4. Operations, CI/CD & DevOps Automate pipeline orchestration using Databricks Asset Bundles (DABs) or Databricks Workflows / Apache Airflow. Build CI/CD pipelines (GitHub Actions, Azure DevOps, or GitLab) for automated testing, deployment, and code promotions. Requirements Required Skills & Qualifications Must-Haves Experience: 8+ years in Data Engineering, with 4+ years of intensive, hands-on production experience on Databricks. Programming Mastery: Fluent in PySpark, Advanced SQL, and Python. Databricks Ecosystem: Deep experience with Delta Lake, Unity Catalog, Delta Live Tables (DLT), Auto Loader, and Databricks Workflows. Cloud Infrastructure: Strong hands-on experience in at least one primary cloud provider (AWS, Azure, or GCP) integration with Databricks (S3/ADLS Gen2, IAM, Key Vaults/Secret Manager). Data Modeling: Solid understanding of dimensional modeling (Kimball), One Big Table (OBT) strategies, and data vault patterns. CI/CD & Software Engineering: Proficient in Git workflows, unit testing PySpark code (pytest), and deployment automation. Preferred / Nice-to-Haves Certifications: Databricks Certified Data Engineer Professional. Streaming: Hands-on with Apache Kafka, Event Hubs, or Kinesis integration via Structured Streaming. GenAI / ML Ops: Familiarity with MLflow, Feature Store, or Vector Search within Databricks. Infrastructure as Code (IaC): Experience using Terraform to provision Databricks workspaces and storage resources. Performance Indicators (How success is measured) Pipeline Reliability: Maintaining strict SLA thresholds on critical Gold-layer models. Cost Efficiency: Measurable reduction in DBU costs through effective compute profiling and tuning. Code Quality: High test coverage and zero-downtime CI/CD deployments.
Ce que vous ferez
Design and maintain enterprise-scale batch and streaming pipelines using a Medallion Architecture on a Databricks Lakehouse platform. Optimize Spark performance and implement data governance and security via Unity Catalog.
Exigences
Requires over 8 years of data engineering experience, with at least 4 years of intensive production experience in Databricks. Mastery of PySpark, SQL, and Python is essential, along with proficiency in CI/CD and cloud infrastructure.
Compétences indiquées
- CI/CDSouhaitée
- PythonSouhaitée
Autres compétences pertinentes
Relevées dans la description du poste. Confirmez les exigences importantes ci-dessus.
- PySpark
- Advanced SQL
- Python
- Databricks
- Delta Lake
- Unity Catalog
- Delta Live Tables
- Auto Loader
- Medallion Architecture
- Spark Performance Tuning
- CI/CD
- Data Modeling
- Cloud Infrastructure
- Apache Airflow
- GitHub Actions
- Azure DevOps
Domaines d’emploi
- Data & Analytics
- Technology
- Software
- Engineering
- Transportation
Renseignements supplémentaires
- Expérience minimale
- 10+ ans
- Postuler avant le
- 12 sept. 2026
- Langue de l’offre
- anglais
- Heures de travail
- 40 heures par semaine
- Niveau d’expérience
- Mid-Senior level
- Mode de candidature
- La candidature directe est offerte