Site Reliability Engineering Manager
- Toronto, ON
- Hybride
- Publié 18 sept. 2026
- 1 poste
Ouvre un site externe
- Type d’emploi
- Contrat
- Niveau d’expérience
- Expérimenté · 6+ ans
- Postuler avant le
- 15 oct. 2026
- Langue de l’offre
- anglais
- Heures de travail
- 40 heures par semaine
- Niveau d’expérience
- Mid-Senior level
- Mode de candidature
- La candidature directe est offerte
Résumé du poste
Responsible for day-to-day reliability and the implementation of Service Level Availability (SLAs) and reliability metrics. Manages high-availability production environments, incident response, and service restoration using AI-powered engineering tools.
Détails du poste
Experienced Application Manager (Site Reliability Engineer) who is responsible for day-to-day reliability Implement and operationalize Service Level Availability (SLAs) and related reliability metrics 6+ years of experience in SRE, Production Engineering, Platform Engineering, Application Support, or Technology Service Delivery. Strong technical knowledge of cloud platforms, enterprise systems, and application, data, and platform architectures. Capital Markets Total Fund Market Investment Proven experience managing high-availability production environments, incident response, troubleshooting, and root cause analysis. AI-powered engineering tools, including IDE assistants and MCP-enabled integrations, to support troubleshooting and service restoration. Experience with observability, automation, Python, and PowerShell. Experience working with APIs, messaging/queueing platforms, distributed systems, CI/CD, Infrastructure as Code, DevOps, and DataOps. Working knowledge of Agile, Waterfall, DevOps, ITIL, and COBIT. Experience with Jira, Confluence, and Git. Excellent communication and collaboration skills, with the ability to partner effectively with architects, business analysts, DBAs, QA, and cross-functional technology teams. Contract/Hybrid Toronto
Ce que vous ferez
Responsible for day-to-day reliability and the implementation of Service Level Availability (SLAs) and reliability metrics. Manages high-availability production environments, incident response, and service restoration using AI-powered engineering tools.
Exigences
Requires 6+ years of experience in SRE or related platform engineering roles with strong knowledge of cloud platforms and distributed systems. Proficiency in Python, PowerShell, and various DevOps/DataOps methodologies is essential.
Compétences indiquées
- CI/CD · Souhaitée
- Root Cause Analysis · Souhaitée
- Agile · Souhaitée
- Python · Souhaitée
Autres compétences pertinentes
Relevées dans la description du poste. Confirmez les exigences importantes ci-dessus.
- SRE
- Production Engineering
- Cloud Platforms
- Incident Response
- Root Cause Analysis
- Observability
- Automation
- Python
- PowerShell
- APIs
- CI/CD
- Infrastructure as Code
- DevOps
- DataOps
- Agile
- ITIL
Domaines d’emploi
- Technology
- Software
- Engineering
- Management & Leadership
- Finance & Accounting
D’autres postes auxquels postuler directement
Des possibilités semblables publiées par des employeurs qui recrutent sur Jobs.ca, sans formulaire externe.
Bédard Ressources Humaines
ITAD Services Representative #1265
CommanditéEmployeur directCandidature simplifiée- Sur place
- Mississauga, ON
- Publié 16 sept. 2026
Sustainable Projects Group
Sustainability consultant/ Team Lead, Energy Consulting NOC 41400
CommanditéEmployeur directCandidature simplifiée- Hybride
- vancouver v5l 4s1, BC
- Publié 17 sept. 2026
Desjardins
Senior Litigation Advisor -
CommanditéEmployeur directCandidature simplifiée- Hybride
- Calgary, AB
- Publié 9 sept. 2026