Senior Software Developer, ML Ops
Offre en anglaisDesign and develop a self-service machine learning platform to enable the efficient deployment and operation of AI solutions. Build framework-agnostic services for model training, inference, and monitoring while establishing technical roadmaps for ML infrastructure.
- Télétravail
- Canada
- Publié 7 août 2026
- 1 poste
Résumé du poste
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Software Developer, ML Ops based in Canada. This role offers the opportunity to build the foundations powering advanced machine learning and AI-driven products at scale. You will design and develop reliable ML infrastructure that enables data scientists and engineers to bring models from experimentation into production. The position sits at the intersection of software engineering, machine learning operations, and large-scale platform development. You will collaborate with technical teams across the organization to solve complex infrastructure challenges and improve ML capabilities. This is a high-impact opportunity for an experienced engineer who enjoys autonomy, technical ownership, and building scalable systems. You will contribute to the evolution of modern ML platforms using cutting-edge technologies and industry-leading engineering practices. The ideal candidate is passionate about building robust systems that enable innovation across multiple product areas. Accountabilities As a Senior Software Developer focused on ML Operations, you will help build and evolve a self-service machine learning platform that enables teams to develop, deploy, and operate ML and AI solutions efficiently. You will work closely with engineers, data scientists, and infrastructure teams to design scalable systems, define technical direction, and deliver reliable platform capabilities. Design, build, and maintain framework-agnostic services supporting machine learning model training, inference, evaluation, and monitoring. Develop scalable ML infrastructure that enables data scientists to move models from experimentation into production reliably. Establish technical roadmaps and architecture decisions for machine learning platform capabilities. Build high-performance systems and contribute to ML infrastructure running on technologies such as Kubernetes and Ray. Collaborate with product, engineering, infrastructure, and data teams to solve complex technical challenges. Take ownership of projects from concept through implementation, deployment, and long-term maintenance. Improve platform reliability through strong engineering practices, observability, performance optimization, and operational excellence. Contribute to the evolution of ML tooling, workflows, and internal developer experiences. Support the adoption of modern AI and machine learning technologies across the organization. Requirements The ideal candidate is an experienced software engineer with a strong background in machine learning infrastructure, scalable systems, and production-grade software development. You should be comfortable owning complex technical projects, collaborating across teams, and building systems designed for long-term impact. 5+ years of experience developing and implementing production-ready software services. 5+ years of software engineering experience focused on ML infrastructure, ML tooling, or machine learning engineering. 3+ years of professional experience with Python. Proven experience designing and building highly scalable, performant, and observable systems. Strong understanding of software architecture, distributed systems, and production engineering practices. Experience owning projects end-to-end and maintaining complex systems over time. Ability to work independently, make technical decisions, and define solutions for ambiguous challenges. Strong collaboration and communication skills with both technical and non-technical stakeholders. Experience with machine learning ecosystems and tools such as PyTorch, Transformers, Ray, feature stores, vector databases, Triton, CUDA, DVC, or vLLM is a plus. Experience with Kubernetes, FastAPI, MLflow, cloud-based ML infrastructure, or data platforms is an advantage. Benefits Competitive base salary range of approximately CA$152,900 - CA$189,000 depending on experience, skills, and location. Equity compensation opportunities as part of total compensation. Comprehensive health benefits and life insurance coverage. Employer-matched group savings program. 20 vacation days, wellness days, and flexible sick and mental health days. Ability to work outside Canada for up to 90 days per year. Remote-friendly work environment with opportunities for collaboration across North America. Access to employee resource groups supporting inclusion and community. Opportunity to work with talented teams building innovative technology products. Supportive culture focused on curiosity, learning, and continuous improvement. How Jobgether Works We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team. We appreciate your interest and wish you the best! Why Apply Through Jobgether? Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time. We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
Ce que vous ferez
Design and develop a self-service machine learning platform to enable the efficient deployment and operation of AI solutions. Build framework-agnostic services for model training, inference, and monitoring while establishing technical roadmaps for ML infrastructure.
Exigences
Requires over 5 years of experience in production-ready software and ML infrastructure, with at least 3 years of professional Python experience. Candidates must demonstrate expertise in building scalable, observable distributed systems and owning complex technical projects end-to-end.
Avantages
• Equity compensation • Comprehensive health benefits • Life insurance coverage • Employer-matched group savings program • 20 vacation days • Wellness days • Flexible sick and mental health days • Ability to work outside Canada for up to 90 days per year • Remote-friendly work environment • Access to employee resource groups
Compétences indiquées
- KubernetesSouhaitée
- PythonSouhaitée
Autres compétences pertinentes
Relevées dans la description du poste. Confirmez les exigences importantes ci-dessus.
- ML Ops
- Python
- Kubernetes
- Distributed Systems
- Software Architecture
- Ray
- PyTorch
- FastAPI
- MLflow
- Model Monitoring
- Production Engineering
- Scalable Systems
- Infrastructure Design
- Vector Databases
- CUDA
- vLLM
Domaines d’emploi
- Software
- Technology
- Engineering
- Data & Analytics
- Science & Research
Renseignements supplémentaires
- Expérience minimale
- 5+ ans
- Langue de l’offre
- anglais
- Heures de travail
- 40 heures par semaine
- Niveau d’expérience
- Mid-Senior level
- Mode de candidature
- La candidature directe est offerte