Technical Lead, AI Infrastructure Optimization
Offre en anglaisLead the design of systems that collect telemetry and improve performance of inference workloads. Collaborate closely with other senior engineers and product managers to execute complex software development initiatives.
- Télétravail
- Canada
- Publié 6 juill. 2026
- 1 poste
Résumé du poste
About Kubex Kubex is building the future of autonomous, AI-driven infrastructure optimization. Our platform enables intelligent, policy-driven optimization across Kubernetes, cloud, and GPU-backed environments improving performance, reducing cost, and eliminating waste for some of the world’s most sophisticated technology organizations. As AI workloads increasingly run on Kubernetes, especially for inference at scale, Kubex is expanding its GPU support to provide more advanced optimization and automation aligned with the unique challenges of GPU-accelerated infrastructure. We combine deep systems expertise, advanced analytics, and patented optimization technology to help customers run AI workloads efficiently and reliably in real-world production environments. Role Overview Kubex is seeking a Technical Lead for AI Infrastructure Optimization to own the technical direction and hands-on development of our AI infrastructure optimization capabilities. This is a senior, hands-on technical leadership role within the office of the CTO. You will own the design and evolution of Kubex’s GPU observability and performance optimization solution. This role carries broad technical ownership and organizational influence. This role is ideal for someone who combines deep, practical experience with GPU infrastructure and Kubernetes with the ability to reason about system-level trade-offs, optimization strategies, & real-world customer environments, and who remains excited to write and ship production code Key Responsibilities Lead the design of systems that collect telemetry and improve performance of inference workloads Contribute directly to production code, remaining deeply hands-on in the design, implementation, and evolution of core platform components. Collaborate closely with other senior engineers, product managers and engineering leadership to coordinate and execute complex software development initiatives. Prototype, validate, and productionize new technical approaches related to AI workload observability and performance optimization. Identify opportunities to extend Kubex’s value beyond inference workloads, including potential future optimizations for training or hybrid workloads. Required Qualifications 5+ years of professional software engineering experience, including significant experience building complex, production systems. Hands-on experience with GPU-accelerated infrastructure, particularly inference on NVIDIA-based environments. Experience using observability tools within Kubernetes for performance optimization (ex: Prometheus, OpenTelemetry). Practical experience with CUDA, GPU telemetry, and performance considerations for AI workloads. Strong coding skills and a demonstrated commitment to remaining hands-on with production code. Preferred Qualifications Strong knowledge of Kubernetes, including how GPU-backed workloads are scheduled, scaled, and operated in real-world clusters Exposure to optimization-based systems, scheduling, bin-packing, or resource allocation problems Why Join Kubex? Play a key role in shaping the future of AI infrastructure optimization. Work on technically challenging problems at the intersection of Kubernetes, GPUs, and AI workloads. Collaborate with a highly experienced, deeply technical team. Influence product direction, architecture, and external technical positioning. Flexible, remote-first culture focused on impact and innovation. Competitive compensation, equity, and benefits
Ce que vous ferez
Lead the design of systems that collect telemetry and improve performance of inference workloads. Collaborate closely with other senior engineers and product managers to execute complex software development initiatives.
Exigences
Candidates should have 5+ years of professional software engineering experience, particularly with GPU-accelerated infrastructure and Kubernetes. Strong coding skills and hands-on experience with performance optimization tools are essential.
Avantages
• Competitive Compensation • Equity • Benefits
Autres compétences pertinentes
Relevées dans la description du poste. Confirmez les exigences importantes ci-dessus.
- AI Workloads
- Kubernetes
- GPU Infrastructure
- Performance Optimization
- Telemetry
- CUDA
- Observability Tools
- Software Development
- Production Code
- System-Level Trade-Offs
- Optimization Strategies
- Collaboration
- Prototyping
- Validation
- Technical Leadership
- Resource Allocation
Domaines d’emploi
- Technology
- Software
- Engineering
- Data & Analytics
- Consulting
Renseignements supplémentaires
- Expérience minimale
- 5+ ans
- Langue de l’offre
- anglais
- Heures de travail
- 40 heures par semaine
- Niveau d’expérience
- Mid-Senior level
- Mode de candidature
- La candidature directe est offerte