Retour à la recherche
SS
Sapphire Stream TechnologySource d’offres vérifiée

AI Performance Engineer – GPU & ROCm

Offre en anglais

Optimize and deploy AI and machine learning workloads using ROCm to improve speed, memory efficiency, and scalability. Profile GPU applications to identify bottlenecks and develop performance benchmarks for various AI models.

  • Sur place
  • Toronto, ON
  • Publié 30 juill. 2026
  • Postuler avant le 29 août 2026
  • 1 poste

Résumé du poste

Job Summary We are looking for a GPU AI Performance Engineer to optimize and deploy AI and machine-learning workloads using ROCm. You will improve the speed, memory efficiency, and scalability of AI training and inference workloads. You will also work with GPU software, AI frameworks, profiling tools, and cloud or containerized environments. Key Responsibilities Optimize AI models for GPU performance. Improve training speed, inference latency, memory usage, and multi-GPU scalability. Work with large language models, vision models, multimodal AI, and generative AI. Develop and troubleshoot applications using ROCm and HIP. Work with PyTorch, TensorFlow, ONNX Runtime, vLLM, SGLang, or MIGraphX. Profile GPU applications and identify performance bottlenecks. Apply mixed precision, quantization, kernel tuning, operator fusion, and memory optimization. Support AI deployment on Linux, cloud, edge, and containerized platforms. Develop performance benchmarks and automated validation tests. Collaborate with hardware, compiler, runtime, framework, and infrastructure teams. Required Qualifications Bachelor’s or master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field. Strong programming skills in C++ and Python. Experience with Linux and GPU computing. Hands-on experience with ROCm, HIP, CUDA, or similar GPU technologies. Understanding of GPU architecture, parallel programming, and memory management. Experience with PyTorch, TensorFlow, ONNX Runtime, or similar AI frameworks. Experience profiling and optimizing GPU applications. Knowledge of AI model training, inference, and performance optimization. Preferred Qualifications Experience optimizing large language models or generative AI applications. Experience with vLLM, SGLang, or distributed inference. Familiarity with LLVM, MLIR, or compiler technologies. Experience migrating workloads between CUDA and HIP. Knowledge of Kubernetes, containers, and distributed AI systems. Contributions to ROCm or other open-source GPU projects.

Ce que vous ferez

Optimize and deploy AI and machine learning workloads using ROCm to improve speed, memory efficiency, and scalability. Profile GPU applications to identify bottlenecks and develop performance benchmarks for various AI models.

Exigences

Requires a degree in Computer Science or a related field with strong proficiency in C++, Python, and GPU computing. Candidates must have hands-on experience with ROCm, HIP, or CUDA and familiarity with major AI frameworks.

Autres compétences pertinentes

Relevées dans la description du poste. Confirmez les exigences importantes ci-dessus.

  • ROCm
  • HIP
  • GPU Optimization
  • C++
  • Python
  • PyTorch
  • TensorFlow
  • ONNX Runtime
  • vLLM
  • SGLang
  • MIGraphX
  • CUDA
  • Parallel Programming
  • Mixed Precision
  • Quantization
  • Kernel Tuning

Domaines d’emploi

  • Engineering
  • Technology
  • Software
  • Data & Analytics
  • Manufacturing

Renseignements supplémentaires

Formation minimale
Baccalauréat
Expérience minimale
5+ ans
Postuler avant le
29 août 2026
Langue de l’offre
anglais
Heures de travail
40 heures par semaine
Niveau d’expérience
Mid-Senior level
Mode de candidature
La candidature directe est offerte