MLOps / ML Platform Engineer
Offre en anglaisDesign and operate scalable ML infrastructure, including data pipelines, training workflows, and production model serving. Ensure system reliability, observability, and security while collaborating with research scientists to streamline the deployment process.
- Télétravail
- Canada, United States
- Publié 12 août 2026
- 1 poste
D’autres postes auxquels postuler directement
Des possibilités semblables publiées par des employeurs qui recrutent sur Jobs.ca, sans formulaire externe.
Résumé du poste
Responsibilities Design and operate ML infrastructure: Manage data, training, serving, and inference systems for high-throughput model workflows. Build scalable pipelines: Implement reproducible training and evaluation pipelines with versioning, scheduling, and artifact tracking. Optimize compute and cost: Tune GPU and CPU workloads, manage clusters, and drive efficiency via rightsizing, spot scheduling, and caching. Serve models in production: Operate APIs for low-latency inference with autoscaling, blue-green or canary rollouts, and rollback safety. Ensure reliability and observability: Define and own SLOs; instrument pipelines and services to track latency, cost, drift, and data quality. Secure and automate: Manage IAM, secrets, and container security; automate deployment pipelines via CI/CD and infrastructure as code. Collaborate cross-functionally: Partner with research scientists and AI engineers to deliver models from experiment to production with minimal friction. Document and enable: Build templates, runbooks, and internal tooling that make ML workflows repeatable, safe, and fast. Qualifications 4+ years of experience in ML platform, DevOps, or infrastructure engineering. Deep knowledge of Kubernetes, CI/CD, containers, and cloud infrastructure (AWS, GCP, or Azure). Hands-on experience managing GPU clusters and training/inference pipelines. Familiarity with data orchestration and storage formats (Delta, Parquet, Polars, Spark). Proven ability to ship and operate production ML systems with SLOs. Strong Python skills and comfort with infrastructure as code and automation. Experience with observability and cost optimization at scale. Nice to Have Experience with real-time or low-latency model serving (REST, gRPC). Exposure to model registry and promotion workflows. Familiarity with data quality, lineage, and curation pipelines. Background in sports analytics or other high-volume data domains. Experience integrating LLM workflows or evaluation pipelines. Benefits Competitive Salary and Bonus Plan Comprehensive health insurance plan Retirement savings plan (401k) with company match Remote working environment A flexible, unlimited time off policy Generous paid holiday schedule - 13 in total including Monday after the Super Bowl SumerSports is committed to fair and equitable compensation practices. Actual compensation packages are based on several factors that are unique to each candidate, including but not limited to skill set, depth of experience, certifications, and specific work location. This may be different in other locations due to differences in the cost of labor. The total compensation package for this position may also include annual performance bonus, benefits and/or other applicable incentive compensation plans.
Ce que vous ferez
Design and operate scalable ML infrastructure, including data pipelines, training workflows, and production model serving. Ensure system reliability, observability, and security while collaborating with research scientists to streamline the deployment process.
Exigences
Requires 4+ years of experience in ML platform, DevOps, or infrastructure engineering with deep knowledge of Kubernetes and cloud environments. Candidates must possess strong Python skills and hands-on experience managing GPU clusters and production ML systems.
Avantages
• Competitive Salary • Bonus Plan • Comprehensive Health Insurance • Retirement Savings Plan • 401k With Company Match • Remote Working Environment • Unlimited Time Off • Paid Holiday Schedule
Compétences indiquées
- KubernetesSouhaitée
- CI/CDSouhaitée
- Amazon Web ServicesSouhaitée
- Google CloudSouhaitée
- Microsoft AzureSouhaitée
- PythonSouhaitée
Autres compétences pertinentes
Relevées dans la description du poste. Confirmez les exigences importantes ci-dessus.
- MLOps
- Kubernetes
- CI/CD
- Cloud Infrastructure
- AWS
- GCP
- Azure
- Python
- GPU Clusters
- Infrastructure As Code
- Observability
- Data Orchestration
- Spark
- Model Serving
- Automation
- System Reliability
- gRPC
- Pipelines
- Sports Analytics
- Container Security
- Apache Parquet
- Workflow Management
- Infrastructure as Code (IaC)
- Research
- Application Programming Interface (API)
- Artificial Intelligence
- Amazon Web Services
- Microsoft Azure
- Data Quality
- DevOps
- Experimentation
- Scalability
- Python (Programming Language)
- Machine Learning
- Tooling
- Software Versioning
- Scheduling
- Templates
- Autoscaling
- Low Latency
- Machine Learning Infrastructure
Domaines d’emploi
- Technology
- Engineering
- Data & Analytics
- Software
- Sports & Recreation
- Platform Engineer
- Software and Applications Developers and Analysts Not Elsewhere Classified
- Validation Engineers
- Industrial Engineers
Renseignements supplémentaires
- Expérience minimale
- 4+ ans
- Langue de l’offre
- anglais
- Heures de travail
- 40 heures par semaine