Monitoring & Observability Architect
Offre en anglaisDesign and operate scalable observability platforms providing metrics, logs, traces, and alerts for cloud-native systems. This includes building pipelines, defining SLOs/SLIs, and collaborating with SRE and platform teams to produce runbooks.
- Télétravail
- Canada, United States
- Publié 30 juill. 2026
- 1 poste
Résumé du poste
Job Role: Monitoring & Observability Architect (with focus on Splunk) Location: Canada/USA- Remote Role overview: Design, implement, and operate scalable observability platforms to provide metrics, logs, traces, and alerts for cloud-native, distributed systems. Responsibilities: · Build and maintain Observability pipelines and dashboards (metrics, logs, traces). · Define and track SLOs, SLIs, and SLAs; create actionable alerts. · Instrument apps/infrastructure with OpenTelemetry and vendor SDKs. · Integrate and operate tools (Splunk, Prometheus/Grafana, ELK/EFK, Datadog, New Relic) · Automate deployments with CI/CD and IaC (Terraform/CloudFormation). · Collaborate with dev, SRE, and platform teams; produce runbooks and post-incident reports. Required: · Hands-on with Splunk at least two monitoring tools (Splunk, Prometheus & Grafana, ELK/EFK, Datadog, New Relic). · Strong scripting/programming (Python, Go, or Bash). · Deep knowledge of Kubernetes, Docker, microservices, and distributed systems. · Experience defining SLOs/SLIs/SLAs. · Familiar with CI/CD and IaC; experience on AWS.
Ce que vous ferez
Design and operate scalable observability platforms providing metrics, logs, traces, and alerts for cloud-native systems. This includes building pipelines, defining SLOs/SLIs, and collaborating with SRE and platform teams to produce runbooks.
Exigences
Requires hands-on experience with Splunk and other monitoring tools, along with proficiency in scripting languages like Python or Go. Candidates must have deep knowledge of Kubernetes, AWS, and Infrastructure as Code (IaC).
Autres compétences pertinentes
Relevées dans la description du poste. Confirmez les exigences importantes ci-dessus.
- Splunk
- Prometheus
- Grafana
- ELK/EFK
- Datadog
- New Relic
- OpenTelemetry
- Python
- Go
- Bash
- Kubernetes
- Docker
- Terraform
- CloudFormation
- AWS
- CI/CD
Domaines d’emploi
- Technology
- Software
- Data & Analytics
- Engineering
- Consulting
Renseignements supplémentaires
- Expérience minimale
- 5+ ans
- Langue de l’offre
- anglais
- Heures de travail
- 40 heures par semaine
- Niveau d’expérience
- Mid-Senior level
- Mode de candidature
- La candidature directe est offerte