Retour à la recherche
K
KubexSource d’offres vérifiée

Staff Software Engineer - Kubernetes Observability

Offre en anglais

Lead the architectural design and technical direction of Kubernetes observability capabilities to support infrastructure optimization. Contribute directly to production code while collaborating with senior leadership to prototype and productionize new technical approaches for AI workloads.

  • Télétravail
  • Canada
  • Publié 23 juill. 2026
  • 1 poste

Résumé du poste

About Kubex Kubex is building the future of autonomous, AI-driven infrastructure optimization. Our platform enables intelligent, policy-driven optimization across Kubernetes, cloud, and GPU-backed environments improving performance, reducing cost, and eliminating waste for some of the world’s most sophisticated technology organizations. As AI workloads increasingly run on Kubernetes, especially for inference at scale, Kubex is expanding its GPU support to provide more advanced optimization and automation aligned with the unique challenges of GPU-accelerated infrastructure. We combine deep systems expertise, advanced analytics, and patented optimization technology to help customers run AI workloads efficiently and reliably in real-world production environments. Role Overview Kubex is seeking a Technical Lead, Kubernetes Observability to own the architecture, technical direction and development of our Kubernetes observability capabilities. This is a senior technical leadership role for someone who can define how telemetry is collected, enriched, processed, and used to support infrastructure optimization across complex customer environments. You will lead the design and evolution of systems that capture workload behavior, resource usage, performance, and operational context across Kubernetes clusters. This includes making key architectural decisions, establishing engineering patterns, prototyping new approaches, and contributing directly to the most critical and technically challenging parts of the implementation. This role combines high-level ownership with strong hands-on engineering. You will use modern AI-assisted development workflows to accelerate implementation, testing, investigation, and iteration, while remaining accountable for system design, technical decisions, code quality, and production reliability. The ideal candidate has deep experience building observability or telemetry systems for Kubernetes and is comfortable working across metrics, events, logs, traces, workload metadata, and distributed data collection. Experience with GPU or AI workloads is valuable but not required. You will help build the observability foundation that supports Kubex’s current Kubernetes optimization capabilities and its continued expansion into GPU and AI infrastructure. Key Responsibilities Lead the design of systems that collect telemetry and improve performance of inference workloads. Contribute directly to production code, remaining deeply hands-on in the design, implementation, and evolution of core platform components. Collaborate closely with other senior engineers, product managers and engineering leadership to coordinate and execute complex software development initiatives. Prototype, validate, and productionize new technical approaches related to AI workload observability and performance optimization. Identify opportunities to extend Kubex’s value beyond inference workloads, including potential future optimizations for training or hybrid workloads. Required Qualifications 7+ years of professional software engineering experience, including significant experience building production software on Kubernetes. Strong experience designing or building observability and telemetry solutions using technologies such as Prometheus, OpenTelemetry, or similar platforms. Deep understanding of Kubernetes workloads, resource management, API interactions, and the operational challenges of running software across diverse customer clusters. Strong coding skills, preferably in Go, with experience building scalable, reliable, and testable distributed systems. Demonstrated ability to own technical architecture while remaining hands-on with implementation, prototyping, debugging, and production delivery. Preferred Qualifications Experience building collectors, exporters, agents, or telemetry-processing pipelines for Kubernetes environments. Knowledge of performance analysis, resource optimization, scheduling, or infrastructure efficiency. Exposure to GPU-backed infrastructure, AI inference workloads, or GPU observability technologies. Why Join Kubex? Play a key role in shaping the future of AI infrastructure optimization. Work on technically challenging problems at the intersection of Kubernetes, GPUs, and AI workloads. Collaborate with a highly experienced, deeply technical team. Influence product direction, architecture, and external technical positioning. Flexible, remote-first culture focused on impact and innovation. Competitive compensation, equity, and benefits.

Ce que vous ferez

Lead the architectural design and technical direction of Kubernetes observability capabilities to support infrastructure optimization. Contribute directly to production code while collaborating with senior leadership to prototype and productionize new technical approaches for AI workloads.

Exigences

Requires over 7 years of professional software engineering experience with a focus on building production software on Kubernetes. Must have strong expertise in observability tools like Prometheus or OpenTelemetry and proficiency in Go for building scalable distributed systems.

Avantages

• Equity • Competitive compensation • Benefits

Autres compétences pertinentes

Relevées dans la description du poste. Confirmez les exigences importantes ci-dessus.

  • Kubernetes
  • Observability
  • Telemetry
  • Go
  • Prometheus
  • OpenTelemetry
  • Distributed Systems
  • Cloud Infrastructure
  • GPU Optimization
  • AI Infrastructure
  • System Architecture
  • Software Engineering
  • Performance Optimization
  • Resource Management
  • API Interactions
  • Debugging

Domaines d’emploi

  • Software
  • Engineering
  • Technology
  • Data & Analytics

Renseignements supplémentaires

Expérience minimale
5+ ans
Langue de l’offre
anglais
Heures de travail
40 heures par semaine
Niveau d’expérience
Mid-Senior level
Mode de candidature
La candidature directe est offerte