DL Performance Software Engineer - LLM Inference
Sur place · Toronto
Architect and implement high-performance inference software and optimize GPU kernels to serve large-scale AI models efficiently. Collaborate across research and engineering teams to integrate state-of-the-art techniques into production-grade open-source software like vLLM.