Retour à la recherche
K
KubexSource d’offres vérifiée

Technical Lead, AI Infrastructure Optimization

Offre en anglais

Lead the design of systems that collect telemetry and improve performance of inference workloads. Collaborate closely with other senior engineers and product managers to execute complex software development initiatives.

  • Télétravail
  • Canada
  • Publié 6 juill. 2026
  • 1 poste

Résumé du poste

About Kubex Kubex is building the future of autonomous, AI-driven infrastructure optimization. Our platform enables intelligent, policy-driven optimization across Kubernetes, cloud, and GPU-backed environments improving performance, reducing cost, and eliminating waste for some of the world’s most sophisticated technology organizations. As AI workloads increasingly run on Kubernetes, especially for inference at scale, Kubex is expanding its GPU support to provide more advanced optimization and automation aligned with the unique challenges of GPU-accelerated infrastructure. We combine deep systems expertise, advanced analytics, and patented optimization technology to help customers run AI workloads efficiently and reliably in real-world production environments. Role Overview Kubex is seeking a Technical Lead for AI Infrastructure Optimization to own the technical direction and hands-on development of our AI infrastructure optimization capabilities. This is a senior, hands-on technical leadership role within the office of the CTO. You will own the design and evolution of Kubex’s GPU observability and performance optimization solution. This role carries broad technical ownership and organizational influence. This role is ideal for someone who combines deep, practical experience with GPU infrastructure and Kubernetes with the ability to reason about system-level trade-offs, optimization strategies, & real-world customer environments, and who remains excited to write and ship production code Key Responsibilities Lead the design of systems that collect telemetry and improve performance of inference workloads Contribute directly to production code, remaining deeply hands-on in the design, implementation, and evolution of core platform components. Collaborate closely with other senior engineers, product managers and engineering leadership to coordinate and execute complex software development initiatives. Prototype, validate, and productionize new technical approaches related to AI workload observability and performance optimization. Identify opportunities to extend Kubex’s value beyond inference workloads, including potential future optimizations for training or hybrid workloads. Required Qualifications 5+ years of professional software engineering experience, including significant experience building complex, production systems. Hands-on experience with GPU-accelerated infrastructure, particularly inference on NVIDIA-based environments. Experience using observability tools within Kubernetes for performance optimization (ex: Prometheus, OpenTelemetry). Practical experience with CUDA, GPU telemetry, and performance considerations for AI workloads. Strong coding skills and a demonstrated commitment to remaining hands-on with production code. Preferred Qualifications Strong knowledge of Kubernetes, including how GPU-backed workloads are scheduled, scaled, and operated in real-world clusters Exposure to optimization-based systems, scheduling, bin-packing, or resource allocation problems Why Join Kubex? Play a key role in shaping the future of AI infrastructure optimization. Work on technically challenging problems at the intersection of Kubernetes, GPUs, and AI workloads. Collaborate with a highly experienced, deeply technical team. Influence product direction, architecture, and external technical positioning. Flexible, remote-first culture focused on impact and innovation. Competitive compensation, equity, and benefits

Ce que vous ferez

Lead the design of systems that collect telemetry and improve performance of inference workloads. Collaborate closely with other senior engineers and product managers to execute complex software development initiatives.

Exigences

Candidates should have 5+ years of professional software engineering experience, particularly with GPU-accelerated infrastructure and Kubernetes. Strong coding skills and hands-on experience with performance optimization tools are essential.

Avantages

• Competitive Compensation • Equity • Benefits

Autres compétences pertinentes

Relevées dans la description du poste. Confirmez les exigences importantes ci-dessus.

  • AI Workloads
  • Kubernetes
  • GPU Infrastructure
  • Performance Optimization
  • Telemetry
  • CUDA
  • Observability Tools
  • Software Development
  • Production Code
  • System-Level Trade-Offs
  • Optimization Strategies
  • Collaboration
  • Prototyping
  • Validation
  • Technical Leadership
  • Resource Allocation

Domaines d’emploi

  • Technology
  • Software
  • Engineering
  • Data & Analytics
  • Consulting

Renseignements supplémentaires

Expérience minimale
5+ ans
Langue de l’offre
anglais
Heures de travail
40 heures par semaine
Niveau d’expérience
Mid-Senior level
Mode de candidature
La candidature directe est offerte