Retour à la recherche
VA
Veeda AISource d’offres vérifiée

Senior Machine Learning Infrastructure Engineer (Precision, Diagnostics & Hardware)

Offre en anglais

Design and optimize high-throughput distributed training systems for multi-modal foundation models on large GPU clusters. Focus on resolving numerical instability, diagnosing hardware faults, and improving overall FLOPS utilization.

  • Hybride
  • Toronto, ON
  • Publié 21 juill. 2026
  • 1 poste

Résumé du poste

About Us Veeda AI is building the next generation of multimodal foundation world models for Physical AI. We're a small, fast-moving team of engineers and researchers from leading AI labs, tackling some of the most challenging problems at the intersection of AI, robotics, and embodied intelligence. If you're excited about pushing the boundaries of what's possible with Physical AI, you'll have the opportunity to make an outsized impact from day one. Responsibilities Distributed Training Systems & Scalability: Design, optimize, and maintain high-throughput distributed training systems across large-scale GPU clusters for multi-modal foundation models. Precision & Numerical Stability: Debug, diagnose, and resolve subtle numerical instability issues (underflow/overflow, loss spikes, gradient explosion, and mixed-precision divergence) in FP16, BF16, FP8, and custom quantization schemes. Fault Diagnostics & Recovery: Build advanced fault-detection mechanisms and automated diagnostics to rapidly pinpoint and isolate silent data corruption (SDC), hardware hang/deadlock, memory leaks, and "card-freeze" issues during large training runs. Performance Profiling & Optimization: Profile distributed communication bottlenecks, memory usage, and kernel execution to improve overall FLOPS utilization across multi-node, multi-GPU training jobs. Developer Tooling & Infrastructure: Develop resilient checkpointing systems, rapid fault-recovery pipelines, and execution telemetry to keep researcher productivity high and hardware downtime minimal. Requirements You have a Bachelor's degree or equivalent hands-on experience in Computer Science, Computer Engineering, or a related technical field. You have deep hands-on experience with deep learning training frameworks (e.g., PyTorch) and distributed training paradigms (FSDP, Megatron-LM, DeepSpeed, Tensor Parallelism, Pipeline Parallelism). You have proven experience in numerical precision analysis, low-precision training (BF16/FP8), and debugging complex loss divergence/stability issues in massive training runs. You have strong root-cause analysis skills for hardware/software interaction bugs, including stuck CUDA kernels, NCCL timeouts, GPU hardware faults, and silent training corruptions. You have strong programming skills in Python and C++/CUDA, with a deep understanding of low-level GPU architectures and memory hierarchies. Nice to Have You have experience running or porting large-scale training workloads on AMD GPUs (ROCm platform) or Google TPUs (JAX/XLA stack). You have contributed to low-level training infrastructure, custom CUDA/Triton kernels, or distributed training open-source projects. You have built resilient fault-tolerant training frameworks with dynamic node re-queueing and rapid checkpointing/saving mechanisms.

Ce que vous ferez

Design and optimize high-throughput distributed training systems for multi-modal foundation models on large GPU clusters. Focus on resolving numerical instability, diagnosing hardware faults, and improving overall FLOPS utilization.

Exigences

Requires a degree in Computer Science or equivalent experience with deep expertise in PyTorch and distributed training paradigms. Must be proficient in Python, C++/CUDA, and low-level GPU memory hierarchies.

Autres compétences pertinentes

Relevées dans la description du poste. Confirmez les exigences importantes ci-dessus.

  • Distributed Training
  • PyTorch
  • CUDA
  • C++
  • Python
  • FSDP
  • Megatron-LM
  • DeepSpeed
  • Numerical Precision Analysis
  • GPU Architecture
  • Performance Profiling
  • Fault Diagnostics
  • Triton
  • Mixed-Precision Training
  • NCCL
  • Tensor Parallelism
  • Scalability Design
  • Pipelines
  • Distributed Machine Learning
  • Resilience
  • Memory Leaks
  • Research
  • Artificial Intelligence
  • Automation
  • C++ (Programming Language)
  • Communication
  • Computer Science
  • Nvidia CUDA
  • Computer Engineering
  • Data Corruption
  • Debugging
  • Fault Tolerance
  • Python (Programming Language)
  • Telemetry
  • Robotics
  • Tooling
  • Root Cause Analysis
  • Deep Learning
  • Quantization
  • PyTorch (Machine Learning Library)

Domaines d’emploi

  • Technology
  • Software
  • Engineering
  • Data & Analytics
  • Science & Research
  • Infrastructure Engineer
  • Machine Learning Engineer
  • Software Developers
  • Computer and Information Research Scientists

Renseignements supplémentaires

Formation minimale
Baccalauréat
Expérience minimale
5+ ans
Langue de l’offre
anglais
Heures de travail
40 heures par semaine
Exigences de lieu
Country, Switzerland, Canada