Retour à la recherche
Logo de Nexxa.ai
Nexxa.aiSource d’offres vérifiée

Staff DevOps Engineer

Offre en anglais
  • Toronto, Ontario, Canada
  • Télétravail
  • Publié 27 août 2026
  • 1 poste

Ouvre un site externe

Connectez-vous pour enregistrer ce poste
Type d’emploi
Temps plein
Niveau d’expérience
Expérimenté · 6+ ans
Langue de l’offre
anglais
Heures de travail
40 heures par semaine
Exigences de lieu
Country, Canada

Résumé du poste

You will own and evolve core infrastructure, including compute, networking, and deployment systems to support AI and industrial workloads. Additionally, you will partner with engineering teams to define observability practices, reliability standards, and infrastructure roadmaps.

Détails du poste

Nexxa is building the best AI systems for heavy industries — enabling machines, systems, and operations to think, decide, and act autonomously across manufacturing, large-scale infrastructure, logistics, and legacy environments. Our mission is to translate deep technical breakthroughs into operational reality, solving some of the hardest systems-level problems in industry. About the Role We're looking for a Senior/Staff DevOps Engineer who has spent the last several years building and operating the infrastructure that lets AI and industrial systems run reliably at scale. You understand what it takes to keep production ML and data workloads fast, observable, and resilient — from GPU-backed training and inference clusters to the pipelines that connect them to real-world industrial environments. This role is ideal for candidates who want deep infrastructure ownership at a company where uptime, latency, and reliability directly affect physical operations — not just software. You'll partner closely with AI, data, and product engineering teams to make sure the systems they build can actually run in production, safely and at scale. What You'll Do Own and evolve Nexxa's core infrastructure — compute, networking, storage, and deployment systems — end-to-end Design and operate CI/CD pipelines that support fast, safe iteration across AI, data, and product engineering teams Build and maintain infrastructure-as-code (e.g., Terraform, Pulumi) for reproducible, auditable environments across cloud and on-prem/edge deployments Architect and manage Kubernetes-based platforms for training, inference, and application workloads, including GPU scheduling and autoscaling Partner with data and AI teams to support the infrastructure behind: Data warehouses and lakehouse architectures (e.g., Snowflake, BigQuery, Redshift, Databricks) Feature stores, embedding indices, and retrieval pipelines Model training, evaluation, and serving infrastructure Define and drive observability practices — metrics, logging, tracing, and alerting — across distributed systems Establish and enforce reliability practices: SLOs/SLIs, incident response, postmortems, and on-call rotations Design for security and compliance across cloud infrastructure, secrets management, and access control, particularly relevant to industrial and legacy-environment integrations Make pragmatic tradeoffs across cost, latency, reliability, and developer velocity Collaborate with engineering leadership to define infrastructure roadmap and platform strategy Mentor engineers on infrastructure best practices and raise the bar for operational excellence across the org Required Qualifications 6+ years of experience in DevOps, Site Reliability Engineering, Platform Engineering, or infrastructure-focused software engineering roles Deep hands-on experience with: Cloud platforms (AWS, GCP, or Azure) at production scale Kubernetes in production, including GPU workload scheduling Infrastructure-as-code tooling (Terraform, Pulumi, or equivalent) CI/CD systems (e.g., GitHub Actions, GitLab CI, CircleCI, Jenkins, ArgoCD) Strong track record designing and operating observability stacks (e.g., Prometheus, Grafana, Datadog, OpenTelemetry) Experience supporting ML/AI infrastructure — training clusters, model serving, data pipelines — a strong plus Excellent scripting/programming skills (Python, Go, or Bash) for automation and tooling Proven ability to independently scope and lead infrastructure projects from design through production rollout Strong incident management instincts — you can lead through an outage calmly and drive toward root cause Preferred Qualifications Experience operating infrastructure that bridges cloud and edge/on-prem environments, especially in industrial or manufacturing contexts Familiarity with data warehouse/lakehouse platforms (Snowflake, BigQuery, Redshift, Databricks) Experience with service mesh, zero-trust networking, or compliance frameworks relevant to industrial/critical infrastructure (e.g., SOC 2, IEC 62443) History of building internal developer platforms or self-service infrastructure tooling Experience scaling infrastructure teams or setting technical direction at a Staff level What Success Looks Like You can own ambiguous, high-stakes infrastructure problems end-to-end Systems you build stay reliable as usage and scale grow — you design for the next order of magnitude, not just today You bring strong technical judgment on tradeoffs between reliability, cost, and speed You raise the bar for operational rigor and engineering discipline across the team You help define what's next for the platform, not just execute what's known Why Join Nexxa.ai? Innovative Environment: Play a critical role in transforming heavy industries through groundbreaking AI and automation technologies Collaborative Culture: Be part of a team that values innovation, discipline, and continuous improvement Professional Growth: Benefit from significant opportunities for career development and advancement Competitive Compensation: Enjoy a comprehensive salary and equity package reflective of your expertise and contributions If you're passionate about building the infrastructure that powers advanced AI solutions in the real world, we'd love to connect.

Ce que vous ferez

You will own and evolve core infrastructure, including compute, networking, and deployment systems to support AI and industrial workloads. Additionally, you will partner with engineering teams to define observability practices, reliability standards, and infrastructure roadmaps.

Exigences

Candidates must have 6+ years of experience in DevOps, SRE, or platform engineering with deep expertise in cloud platforms and Kubernetes. Strong proficiency in infrastructure-as-code, CI/CD pipelines, and scripting languages like Python or Go is essential.

Avantages

• Competitive compensation • Equity package • Professional growth opportunities

Compétences indiquées

  • Microsoft Azure · Souhaitée
  • Kubernetes · Souhaitée
  • CI/CD · Souhaitée
  • Go · Souhaitée
  • Amazon Web Services · Souhaitée
  • Google Cloud · Souhaitée
  • Terraform · Souhaitée
  • Python · Souhaitée

Autres compétences pertinentes

Relevées dans la description du poste. Confirmez les exigences importantes ci-dessus.

  • DevOps
  • Site Reliability Engineering
  • Kubernetes
  • Terraform
  • Pulumi
  • AWS
  • GCP
  • Azure
  • Python
  • Go
  • Bash
  • CI/CD
  • Observability
  • Prometheus
  • Grafana
  • Datadog
  • Argo CD
  • Pipelines
  • OpenTelemetry
  • Go (Programming Language)
  • AI/ML Inference
  • Snowflake (Data Warehouse)
  • Resilience
  • Machine Learning Model Training
  • Infrastructure as Code (IaC)
  • Access Controls
  • Artificial Intelligence
  • Amazon Web Services
  • Automation
  • Microsoft Azure
  • Bash (Scripting Language)
  • Google BigQuery
  • Management
  • Cloud Infrastructure
  • Continuous Improvement Process
  • Course Evaluations
  • Data Warehousing
  • Incident Response
  • End Systems
  • Github
  • Leadership
  • Incident Management
  • Innovation
  • Python (Programming Language)
  • Machine Learning
  • Operational Excellence
  • Operations
  • Product Family Engineering
  • Product Engineering
  • Prometheus (Software)

Domaines d’emploi

  • Technology
  • Software
  • Engineering
  • Data & Analytics
  • Manufacturing
  • Staff DevOps Engineer
  • DevOps Engineer
  • Software and Applications Developers and Analysts Not Elsewhere Classified
  • Software Developers

D’autres postes auxquels postuler directement

Des possibilités semblables publiées par des employeurs qui recrutent sur Jobs.ca, sans formulaire externe.

Voir tous les postes à candidature simplifiée