Retour à la recherche
VA
Veeda AISource d’offres vérifiée

Senior AI Infrastructure Engineer — HPC & Compute Clusters

Offre en anglais

Design, deploy, and maintain large-scale bare-metal GPU clusters using Slurm and Kubernetes for distributed training. Optimize high-speed interconnects and distributed storage to ensure maximum GPU utilization and reliability.

  • Hybride
  • Toronto, ON
  • Publié 21 juill. 2026
  • 1 poste

Résumé du poste

About US Veeda AI is building the next generation of multimodal foundation world models for Physical AI. We're a small, fast-moving team of engineers and researchers from leading AI labs, tackling some of the most challenging problems at the intersection of AI, robotics, and embodied intelligence. If you're excited about pushing the boundaries of what's possible with Physical AI, you'll have the opportunity to make an outsized impact from day one. Responsibilities GPU Cluster Orchestration: Design, deploy, and maintain large-scale bare-metal GPU clusters running Slurm and Kubernetes to power multi-node distributed training runs. Workload Management & Scheduler Tuning: Configure, tune, and optimize Slurm schedulers, dynamic job queuing, priority topologies, and autoscaling for optimal GPU utilization and job throughput. Kubernetes Integration: Manage K8s clusters and hybrid Slurm-K8s environments (e.g., KubeRay, MPI Operator, Slurm on K8s) for serving, data processing pipelines, and interactive research environments. Storage & Network Performance: Optimize high-speed interconnects (InfiniBand/RoCE, NCCL, NVLink) and high-throughput distributed storage (Lustre, WEKA, Ceph) to keep tens of thousands of GPUs continuously saturated. Reliability & Monitoring: Build observability pipelines (Prometheus, Grafana, DCGM) to proactively detect hardware degradation, GPU silent errors, network flaps, and node failures before they impact training jobs. Requirements You have a Bachelor's degree in Computer Science, Computer Engineering, or equivalent hands-on experience in high-performance computing (HPC) or infrastructure engineering. You have deep hands-on expertise administering Linux-based HPC clusters running Slurm or Kubernetes (job scheduling, cgroups, fair-share policies, topology configuration). You possess strong troubleshooting skills in low-level Linux networking, kernel parameters, hardware diagnostics, and storage systems (e.g., NFS, NVMe-oF, or distributed filesystems). You are proficient in automation and infrastructure-as-code tools (e.g., Ansible, Terraform, Helm, Python/Bash scripting). Background supporting distributed deep learning frameworks (DeepSpeed, Ray). Nice to Have Experience managing high-density GPU infrastructure (NVIDIA H100/B200/B300 systems, DGX/HGX architectures, DCGM, NCCL tuning). Experience with high-speed network fabrics including InfiniBand (Subnet Manager, SMDB) and RoCE (v2).

Ce que vous ferez

Design, deploy, and maintain large-scale bare-metal GPU clusters using Slurm and Kubernetes for distributed training. Optimize high-speed interconnects and distributed storage to ensure maximum GPU utilization and reliability.

Exigences

Requires a Bachelor's degree in Computer Science or equivalent experience with deep expertise in Linux-based HPC clusters and infrastructure-as-code tools. Proficiency in low-level networking, hardware diagnostics, and distributed deep learning frameworks is essential.

Autres compétences pertinentes

Relevées dans la description du poste. Confirmez les exigences importantes ci-dessus.

  • GPU Cluster Orchestration
  • Slurm
  • Kubernetes
  • InfiniBand
  • RoCE
  • NCCL
  • NVLink
  • Lustre
  • WEKA
  • Ceph
  • Prometheus
  • Grafana
  • Ansible
  • Terraform
  • Python
  • Bash
  • Ceph (Software)
  • Pipelines
  • Distributed Machine Learning
  • Observability
  • Infrastructure as Code (IaC)
  • Research
  • Artificial Intelligence
  • Automation
  • Bash (Scripting Language)
  • Management
  • Computer Science
  • Data Processing
  • Computer Engineering
  • Linux
  • File Systems
  • Distributed Data Store
  • Python (Programming Language)
  • Message Passing Interface
  • Network File Systems
  • Network Performance Management
  • Prometheus (Software)
  • Robotics
  • Topology
  • Troubleshooting (Problem Solving)
  • Weka
  • Job Scheduling (Inventory Management)
  • High Performance Computing
  • Autoscaling
  • Deep Learning
  • Bare Metal
  • Slurm (Batch Scheduling Software)

Domaines d’emploi

  • Technology
  • Engineering
  • Software
  • Data & Analytics
  • Science & Research
  • Infrastructure Engineer
  • Artificial Intelligence Engineer (General)
  • Software Developers

Renseignements supplémentaires

Formation minimale
Baccalauréat
Expérience minimale
5+ ans
Langue de l’offre
anglais
Heures de travail
40 heures par semaine
Exigences de lieu
Country, Switzerland, Canada