Platform Observability Engineer
Offre en anglaisABOUT THE ROLE We are seeking an experienced Observability Engineer to join our Enterprise Kubernetes Platform team at a leading financial services organization. You’ll own the complete observability stack across 50+ production Kubernetes clusters, providing metrics, logging, tracing, and alerting capabilities that ensure exceptional reliability and performance for mission-critical applications. This role combines deep technical expertise in modern observability tools with emerging AI/ML capabilities to build intelligent monitoring solutions, predictive alerting, and self-healing infrastruct…
- Sur place
- ONTARIO
- Publié 13 juill. 2026
- Postuler avant le 12 août 2026
- 1 poste
Résumé du poste
ABOUT THE ROLE We are seeking an experienced Observability Engineer to join our Enterprise Kubernetes Platform team at a leading financial services organization. You’ll own the complete observability stack across 50+ production Kubernetes clusters, providing metrics, logging, tracing, and alerting capabilities that ensure exceptional reliability and performance for mission-critical applications. This role combines deep technical expertise in modern observability tools with emerging AI/ML capabilities to build intelligent monitoring solutions, predictive alerting, and self-healing infrastructure. WHAT YOULL DO • Design, deploy, and maintain enterprise-scale observability infrastructure including Prometheus, Grafana, Thanos, Loki, and modern collection agents • Manage observability deployments using GitOps principles and infrastructures code • Implement long-term metrics storage solutions with cloud object storage • Maintain and upgrade observability components across development, QA, UAT, production, and DR environments • Configure distributed observability architecture spanning multiple datacenters and cloud providers METRICS & MONITORING • Design and implement Prometheus monitoring strategies for Kubernetes infrastructure and containerized applications • Create Service Monitors, Pod Monitors for automated metrics collection • Develop rules for intelligent alerting with minimal false positives • Configure multi-cluster metrics federation and aggregation • Optimize metrics cardinality, storage deficiency, and query performance.
Ce que vous ferez
The Platform Observability Engineer will design, deploy, and maintain the observability infrastructure for over 50 production Kubernetes clusters. This includes implementing monitoring strategies, configuring alerting systems, and optimizing metrics collection.
Exigences
Candidates should have deep technical expertise in modern observability tools and experience with AI/ML capabilities. Familiarity with GitOps principles and infrastructure as code is also essential.
Autres compétences pertinentes
Relevées dans la description du poste. Confirmez les exigences importantes ci-dessus.
- Observability
- Kubernetes
- Prometheus
- Grafana
- Thanos
- Loki
- GitOps
- Infrastructure as Code
- Metrics Storage
- Alerting
- Monitoring
- Cloud
- AI
- ML
- Self-Healing Infrastructure
- Metrics Federation
Renseignements supplémentaires
- Expérience minimale
- 5+ ans
- Postuler avant le
- 12 août 2026