Retour à la recherche
S
SITASource d’offres vérifiée

Site Reliability Engineer/ Expert/ Specialist | Ingénieur(e) Site Reliability (SRE) | HYBRID role

Offre en anglais

Ensure the reliability and performance of production systems through deep technical analysis and proactive engineering. Lead incident investigations, perform root cause analysis, and automate operational tasks to improve service resilience.

  • Hybride
  • Montréal, QC
  • Publié 24 août 2026
  • Postuler avant le 23 sept. 2026
  • 1 poste

D’autres postes auxquels postuler directement

Des possibilités semblables publiées par des employeurs qui recrutent sur Jobs.ca, sans formulaire externe.

Résumé du poste

Overview Welcome to SITA! At SITA, we keep airports moving, airlines flying smoothly, and borders open. Our technology and communication innovations power the success of the global air travel industry. You’ll find us in 95% of international airports, working closely with over 2,500 transportation and government clients. Each partnership brings unique challenges, and we thrive on delivering fresh solutions and cutting-edge tech to keep operations running like clockwork. We don’t just move the world forward—we’re proud to be recognized as a Great Place to Work® by 79% of our employees and certified in most of our growing locations. Here, we feel empowered, supported, and inspired to grow. Are you ready to love your job? The adventure begins right here, with you, at SITA. About The Role & Team We are seeking a hands-on Site Reliability Engineer (SRE) with strong expertise across application support, Kubernetes environments, and CI/CD pipelines. This role is responsible for ensuring the reliability, performance, and observability of production systems through deep technical analysis and proactive engineering. The successful candidate will lead incident and problem investigations, identify root causes, and drive permanent resolutions in close collaboration with Development, Product, and Operations teams. This is an engineering-focused role requiring strong troubleshooting skills, production support experience, and a commitment to continuous service improvement. What You Will Do Analyze production incidents using logs, metrics, and traces to identify impacted application code and execution paths. Diagnose system issues and determine whether the root cause is related to application code, configuration, Kubernetes, or infrastructure. Troubleshoot Kubernetes workloads, including runtime behavior, networking, health probes, and failure scenarios. Serve as the technical escalation point during critical incidents, providing clear and timely guidance. Lead root cause analysis (RCA) and drive permanent corrective actions to improve reliability. Enhance observability and alerting to improve issue detection and resolution. Partner with Development and Platform teams to resolve systemic issues and strengthen service reliability. Automate repetitive operational tasks and promote engineering best practices. Assess the impact of deployments on production environments through CI/CD pipeline expertise. Monitor system performance and reliability, identifying opportunities to improve resilience and reduce outages. Qualifications ABOUT YOUR SKILLS 5+ years’ experience in SRE, DevOps, or Production Engineering supporting high-availability systems. Strong expertise in root cause analysis (RCA) and permanent issue resolution. Hands-on troubleshooting of production environments using logs, metrics, and traces. Strong knowledge of .NET and/or Java applications in production. Hands-on experience with Kubernetes and containerized workloads. Experience with monitoring and observability tools for distributed systems. Familiarity with CI/CD pipelines and deployment tools (e.g., Azure DevOps, Jenkins, GitHub Actions). Scripting and automation skills using Python, Bash, or similar. Strong analytical, communication, and cross-functional collaboration skills. Bachelor’s degree in Computer Science, Engineering, or related field, or equivalent experience. What We Offer We value diversity, operating in 200 countries and spanning 60 languages and cultures. Our inclusive offices are comfortable and fun, with the flexibility to work from home. Join our team and step closer to your best life. 🏡 Flex Week: Hybrid (2 days from home and 3 days in the Montreal office. 🌎 Flex Location: Take up to 30 days a year to work from any location in the world. 🌿Employee Wellbeing: We’ve got you covered with our Employee Assistance Program (EAP), for you and your dependents 24/7, 365 days/year. We also offer Champion Health a personalized platform that supports a range of well-being needs. 🚀Professional Development: Level up your skills with our training platforms, including LinkedIn Learning! 🙌🏽 Competitive Benefits: Competitive benefits that make sense with both your local market and employment status. SITA is an Employment Equity Employer and values a diverse workforce. In support of our Employment Equity Program, women, Aboriginal people, members of visible minorities, and/or persons with disabilities are encouraged to apply and self-identify in the application process.

Ce que vous ferez

Ensure the reliability and performance of production systems through deep technical analysis and proactive engineering. Lead incident investigations, perform root cause analysis, and automate operational tasks to improve service resilience.

Exigences

Requires 5+ years of experience in SRE or DevOps with strong expertise in Kubernetes, .NET/Java, and CI/CD tools. A Bachelor's degree in Computer Science or a related field is required or equivalent experience.

Avantages

• Hybrid Work • Work From Anywhere (up to 30 days/year) • Employee Assistance Program (EAP) • Champion Health Wellbeing Platform • Professional Development Training • LinkedIn Learning • Competitive Benefits

Compétences indiquées

  • KubernetesSouhaitée
  • .NETSouhaitée
  • JavaSouhaitée
  • PythonSouhaitée

Autres compétences pertinentes

Relevées dans la description du poste. Confirmez les exigences importantes ci-dessus.

  • Site Reliability Engineering
  • Kubernetes
  • CI/CD Pipelines
  • Root Cause Analysis
  • Observability
  • Production Support
  • Dotnet
  • Java
  • Python
  • Bash
  • Azure DevOps
  • Jenkins
  • GitHub Actions
  • Containerization
  • Troubleshooting
  • Automation

Domaines d’emploi

  • Technology
  • Software
  • Engineering
  • Consulting
  • Transportation

Renseignements supplémentaires

Formation minimale
Baccalauréat
Expérience minimale
5+ ans
Postuler avant le
23 sept. 2026
Langue de l’offre
anglais
Heures de travail
40 heures par semaine
Présence au bureau
3 jours par semaine
Niveau d’expérience
Not Applicable