Spark Scala Developer / Big data engineer
Offre en anglaisDesign and develop large-scale data processing pipelines using Apache Spark, Scala, and PySpark. Build and optimize batch and real-time workflows utilizing Kafka and the Hadoop ecosystem to support business requirements.
- Sur place
- Mississauga, ON
- Publié 11 août 2026
- Postuler avant le 10 sept. 2026
- 1 poste
D’autres postes auxquels postuler directement
Des possibilités semblables publiées par des employeurs qui recrutent sur Jobs.ca, sans formulaire externe.
Résumé du poste
Hi, Hope you are doing good! We have a Full-time opportunity for you as Spark Scala Developer @ Mississauga, ON Role: Spark Scala Developer Locations: Mississauga, ON Type of Hiring: FTE Required Qualifications: At least 4 years of Information Technology experience 4+ years of experience in Big Data technologies. Strong expertise in: Apache Spark (Core, SQL, Data Frames, RDDs) Scala programming PySpark Hands-on experience with: Kafka (real-time streaming) Hadoop ecosystem (HDFS, Hive, Impala) NoSQL Databases (HBase, MongoDB, Couchbase) Strong understanding of distributed computing concepts and data processing frameworks. Experience in building ETL/data pipelines for large-scale datasets. Proficiency in SQL and data modeling. Preferred Qualifications: Hands-on experience with data lakes, data warehouses, and scalable ETL pipeline design, including batch and real-time processing architecture. Strong understanding and practical exposure to Agile software development methodologies (Scrum) and SDLC practices. Proven experience in Banking domain, supporting use cases such as fraud detection, risk analytics, regulatory reporting, and customer insights. Excellent analytical, problem-solving, and communication skills, with the ability to translate business requirements into scalable technical solutions. Demonstrated ability to work effectively in cross-functional, multi-stakeholder environments, collaborating with Business, Data Engineering, and Architecture teams. Experience with real-time data streaming frameworks such as Kafka and Spark Streaming for low-latency processing. Understanding data modeling concepts (dimensional modeling, snowflake schemas) to support analytics workloads. Experience and desire to work in a global delivery environment. Key Responsibilities: Design and develop large-scale data processing pipelines using Apache Spark (Scala & PySpark) Build and optimize batch and real-time data processing workflows using Spark, Kafka, and Hadoop ecosystem Develop Spark applications using RDDs, DataFrames, and Spark SQL for complex transformations Develop and optimize PySpark applications leveraging joins, Spark DAG execution flow, stage optimization, transformation techniques, and streaming with dynamic allocation and failover handling. Implement streaming pipelines using Kafka and Spark Streaming / Structured Streaming Develop and maintain HDFS, Hive, NoSql and Impala-based data lake solutions Convert existing SQL/Hive workloads into optimized Spark jobs for improved performance Work with ETL pipelines to ingest, cleanse, transform, and process large datasets Optimize performance through partitioning, caching, serialization, and tuning techniques Handle data formats such as Parquet, ORC, Avro, JSON Integrate multiple data sources including streaming systems, flat files RDBMS, and APIs Collaborate with cross-functional teams to understand business requirements and translate them into scalable technical solutions Ensure data quality, reliability, and performance monitoring across pipelines Participate in code reviews, design discussions, and best practices implementation Key Skills: Distributed Data Processing. Spark Optimization & Performance Tuning. Real-time Data Streaming. Data Modeling & ETL Design. Problem-solving and Analytical Thinking. Strong Communication & Stakeholder Management. Nice to Have: Exposure to Machine Learning pipelines or MLOps workflows. Experience with Databricks platform. Experience with AWS/GCP.
Ce que vous ferez
Design and develop large-scale data processing pipelines using Apache Spark, Scala, and PySpark. Build and optimize batch and real-time workflows utilizing Kafka and the Hadoop ecosystem to support business requirements.
Exigences
Requires at least 4 years of IT experience with a strong focus on Big Data technologies and distributed computing. Proficiency in Scala, PySpark, and NoSQL databases is essential, with banking domain experience preferred.
Compétences indiquées
- SQLSouhaitée
- AgileSouhaitée
Autres compétences pertinentes
Relevées dans la description du poste. Confirmez les exigences importantes ci-dessus.
- Apache Spark
- Scala
- PySpark
- Kafka
- Hadoop
- HDFS
- Hive
- Impala
- NoSQL
- ETL Design
- Data Modeling
- Distributed Computing
- SQL
- Spark Streaming
- Performance Tuning
- Agile
Domaines d’emploi
- Data & Analytics
- Software
- Technology
- Engineering
- Consulting
Renseignements supplémentaires
- Expérience minimale
- 5+ ans
- Postuler avant le
- 10 sept. 2026
- Langue de l’offre
- anglais
- Heures de travail
- 40 heures par semaine
- Niveau d’expérience
- Mid-Senior level
- Mode de candidature
- La candidature directe est offerte