Retour à la recherche
Logo de Bevertec
BevertecSource d’offres vérifiée

Site Reliability Engineering Manager

Offre en anglais
  • Toronto, ON
  • Hybride
  • Publié 18 sept. 2026
  • 1 poste

Ouvre un site externe

Connectez-vous pour enregistrer ce poste
Type d’emploi
Contrat
Niveau d’expérience
Expérimenté · 6+ ans
Postuler avant le
15 oct. 2026
Langue de l’offre
anglais
Heures de travail
40 heures par semaine
Niveau d’expérience
Mid-Senior level
Mode de candidature
La candidature directe est offerte

Résumé du poste

Responsible for day-to-day reliability and the implementation of Service Level Availability (SLAs) and reliability metrics. Manages high-availability production environments, incident response, and service restoration using AI-powered engineering tools.

Détails du poste

Experienced Application Manager (Site Reliability Engineer) who is responsible for day-to-day reliability Implement and operationalize Service Level Availability (SLAs) and related reliability metrics 6+ years of experience in SRE, Production Engineering, Platform Engineering, Application Support, or Technology Service Delivery. Strong technical knowledge of cloud platforms, enterprise systems, and application, data, and platform architectures. Capital Markets Total Fund Market Investment Proven experience managing high-availability production environments, incident response, troubleshooting, and root cause analysis. AI-powered engineering tools, including IDE assistants and MCP-enabled integrations, to support troubleshooting and service restoration. Experience with observability, automation, Python, and PowerShell. Experience working with APIs, messaging/queueing platforms, distributed systems, CI/CD, Infrastructure as Code, DevOps, and DataOps. Working knowledge of Agile, Waterfall, DevOps, ITIL, and COBIT. Experience with Jira, Confluence, and Git. Excellent communication and collaboration skills, with the ability to partner effectively with architects, business analysts, DBAs, QA, and cross-functional technology teams. Contract/Hybrid Toronto

Ce que vous ferez

Responsible for day-to-day reliability and the implementation of Service Level Availability (SLAs) and reliability metrics. Manages high-availability production environments, incident response, and service restoration using AI-powered engineering tools.

Exigences

Requires 6+ years of experience in SRE or related platform engineering roles with strong knowledge of cloud platforms and distributed systems. Proficiency in Python, PowerShell, and various DevOps/DataOps methodologies is essential.

Compétences indiquées

  • CI/CD · Souhaitée
  • Root Cause Analysis · Souhaitée
  • Agile · Souhaitée
  • Python · Souhaitée

Autres compétences pertinentes

Relevées dans la description du poste. Confirmez les exigences importantes ci-dessus.

  • SRE
  • Production Engineering
  • Cloud Platforms
  • Incident Response
  • Root Cause Analysis
  • Observability
  • Automation
  • Python
  • PowerShell
  • APIs
  • CI/CD
  • Infrastructure as Code
  • DevOps
  • DataOps
  • Agile
  • ITIL

Domaines d’emploi

  • Technology
  • Software
  • Engineering
  • Management & Leadership
  • Finance & Accounting

D’autres postes auxquels postuler directement

Des possibilités semblables publiées par des employeurs qui recrutent sur Jobs.ca, sans formulaire externe.

Voir tous les postes à candidature simplifiée