Retour à la recherche
Logo de Direct IT Recruiting Inc.
Direct IT Recruiting Inc.Source d’offres vérifiée

Site Reliability Engineer (1890)

Offre en anglais
  • Toronto, ON
  • Hybride
  • Publié 19 sept. 2026
  • 1 poste

85 $–94 $ / heure

Ouvre un site externe

Connectez-vous pour enregistrer ce poste
Type d’emploi
Contrat
Niveau d’expérience
Expérimenté · 5+ ans
Postuler avant le
16 oct. 2026
Langue de l’offre
anglais
Heures de travail
40 heures par semaine
Présence au bureau
3 jours par semaine
Niveau d’expérience
Mid-Senior level
Mode de candidature
La candidature directe est offerte

Résumé du poste

The role focuses on developing automation solutions, enhancing CI/CD pipelines, and improving the observability and resilience of critical financial platforms. The engineer will also participate in incident response, disaster recovery planning, and cross-team collaboration to ensure operational readiness.

Détails du poste

NOTE: Hybrid 3 days a week in Toronto Office Type: 9 month contract, 8 hours/day, 40 hours/week Rate: $85 to $94/hr SKILLS: Python, GitHub Actions, CI/CD, automation, Site Reliability Engineering, Dynatrace, Argo CD, PowerShell, observability Industry: Capital Markets and Investment Management DESCRIPTION We are seeking an experienced Site Reliability Engineer to support applications, data pipelines, platforms, and services across Capital Markets, Investment Analytics, and Total Fund Management. This is a hands on individual contributor role focused on Python development, operational automation, CI/CD, observability, disaster recovery readiness, and platform modernization. The successful candidate will improve health checks, deployment pipelines, telemetry, diagnostics, and operational processes while addressing a significant backlog of reliability initiatives. The environment is highly Python focused and includes GitHub Actions, Argo CD, Dynatrace, Elastic and ELK, PowerShell, Ansible, Bash, and cron jobs. Incident volume is relatively low, allowing the successful candidate to focus primarily on proactive engineering and sustainable automation. Capital markets experience is beneficial but not required. Strong Python, automation, CI/CD, and observability experience are the primary priorities. RESPONSIBILITIES • Build Automation and Scripting Solutions: Develop maintainable Python scripts, tools, and automation solutions that reduce manual effort, improve consistency, and strengthen operational efficiency. • Improve CI/CD and Deployment Processes: Enhance build, deployment, release, rollback, and production validation processes using GitHub Actions, Argo CD, and related technologies. • Advance Site Reliability Engineering Practices: Improve the availability, performance, resilience, supportability, and operational sustainability of applications, data pipelines, platforms, and services. • Enhance Observability and Telemetry: Strengthen application health checks, logging, monitoring, metrics, alerts, dashboards, diagnostics, and telemetry. Support the adoption of Dynatrace and the transition from the current Elastic and ELK logging environment. • Strengthen Incident Response and Root Cause Analysis: Investigate production issues using logs, metrics, traces, and diagnostic tools. Identify root causes and implement corrective and preventative actions. • Improve Detection and Recovery: Develop monitoring, diagnostics, automation, and operational runbooks that reduce Mean Time to Detect and Mean Time to Restore. • Improve Operational Readiness: Conduct production readiness reviews and ensure monitoring, runbooks, resilience testing, documentation, release processes, and support arrangements are in place before production deployment. • Support Disaster Recovery and Resilience: Contribute to disaster recovery plans, recovery procedures, resilience initiatives, and operational testing. • Support Production Systems: Participate in incident response, coordinate technical activities, communicate service impacts, and escalate major or cross service incidents through the appropriate channels. • Partner Across Technology Teams: Collaborate with Product Engineering, Data Solutions, Data Platform, Architecture, Security, Technology Services, database, and quality assurance teams to identify and resolve reliability, resilience, and supportability risks. REQUIREMENTS • 6 or more years of experience in Site Reliability Engineering, production engineering, platform engineering, application support, or technology service delivery. • Strong hands on Python development and scripting experience, including sound coding, testing, documentation, and source control practices. • Experience designing automation solutions and independently selecting the appropriate scripting language, platform, or tool to solve operational problems. • Experience developing, maintaining, or improving CI/CD pipelines, preferably using GitHub Actions and Argo CD. • Experience with observability, application monitoring, centralized logging, telemetry, alerting, and production diagnostics. Dynatrace experience is strongly preferred. • Experience with Elastic, ELK, PowerShell, Ansible, Bash, cron jobs, or similar scripting and automation technologies is beneficial. • Hands on experience operating and improving highly available applications, data pipelines, platforms, APIs, message queues, or distributed services. • Experience with incident response, root cause analysis, disaster recovery, resilience testing, capacity planning, and production readiness. • Experience with cloud environments, infrastructure as code, DevOps, DataOps, automated testing, release automation, and production validation practices. • Experience using Git, JIRA, Confluence, and related engineering collaboration tools. • Strong problem solving skills with the ability to understand an unfamiliar problem space, establish an effective approach, and independently drive work through completion. • Strong communication and collaboration skills, including the ability to explain technical risks, incidents, and service impacts in clear business language. • Experience working with developers, delivery leads, architects, business analysts, database administrators, quality assurance professionals, security teams, and infrastructure teams. • Knowledge of Agile, Waterfall, DevOps, ITIL, or COBIT practices. • Practical experience using approved AI assisted engineering tools to troubleshoot applications, databases, infrastructure, code, logs, and telemetry is an asset. • Capital markets, investment analytics, investment management, or financial services experience is beneficial but not required. Note: As part of our hiring process, we use AI based systems to support initial applicant screening. TO APPLY: https://directitrecruiting.com/job/site-reliability-engineer-1890/

Ce que vous ferez

The role focuses on developing automation solutions, enhancing CI/CD pipelines, and improving the observability and resilience of critical financial platforms. The engineer will also participate in incident response, disaster recovery planning, and cross-team collaboration to ensure operational readiness.

Exigences

Candidates must have 6 or more years of experience in Site Reliability Engineering or related fields with strong hands-on Python development skills. Proficiency in CI/CD tools, observability platforms, and operational automation is essential for this role.

Compétences indiquées

  • CI/CD · Souhaitée
  • Python · Souhaitée

Autres compétences pertinentes

Relevées dans la description du poste. Confirmez les exigences importantes ci-dessus.

  • Python
  • Site Reliability Engineering
  • CI/CD
  • Automation
  • GitHub Actions
  • Argo CD
  • Dynatrace
  • Observability
  • PowerShell
  • Elastic
  • ELK
  • Ansible
  • Bash
  • Disaster Recovery
  • Telemetry
  • Infrastructure as Code

Domaines d’emploi

  • Technology
  • Software
  • Data & Analytics
  • Finance & Accounting
  • Engineering

D’autres postes auxquels postuler directement

Des possibilités semblables publiées par des employeurs qui recrutent sur Jobs.ca, sans formulaire externe.

Voir tous les postes à candidature simplifiée