Retour à la recherche
A
AptonetSource d’offres vérifiée

Application Support Engineer

Offre en anglais

Provide hands-on production support for cloud-native, Kubernetes-based applications in a large-scale retail environment. Responsibilities include monitoring application health, triaging production incidents, and improving observability solutions to ensure system reliability.

  • Hybride
  • Toronto, ON
  • Publié 18 août 2026
  • Postuler avant le 17 sept. 2026
  • 1 poste

D’autres postes auxquels postuler directement

Des possibilités semblables publiées par des employeurs qui recrutent sur Jobs.ca, sans formulaire externe.

Résumé du poste

Senior Application Support / Systems Engineer Job Title: Senior Associate Technology L2 Client: Leading Retail / eCommerce Company Location: Toronto, Ontario Work Arrangement: Hybrid – 2–3 days per week onsite Contract Duration: 6 months initially Target Start Date: Immediately Rate: $105–$115 CAD/hour Position Summary We are seeking a Senior Application Support / Systems Engineer to provide hands-on production support for cloud-native, Kubernetes-based applications within a large-scale retail environment. The successful candidate will be responsible for maintaining application health and reliability, configuring monitoring and alerting solutions, responding to and triaging production incidents, and troubleshooting performance, latency, and availability issues. This role requires a strong combination of application support, cloud engineering, observability, incident management, CI/CD, and production troubleshooting experience. The ideal candidate will have hands-on expertise with Google Cloud Platform (GCP), Google Kubernetes Engine (GKE), Kubernetes, Jenkins, and modern observability tools. Key Responsibilities Monitor application health, infrastructure signals, service availability, performance, and production workloads across cloud-based environments. Configure, maintain, and continuously improve observability solutions, dashboards, alerts, metrics, logs, and distributed tracing. Utilize monitoring and observability platforms such as Grafana, Prometheus, Datadog, Dynatrace, Splunk, or similar technologies. Support Kubernetes and GKE environments, including troubleshooting workloads, deployments, services, networking, and resource-related issues. Troubleshoot production incidents involving application failures, latency, performance degradation, integration issues, and service availability. Participate in and drive incident triage, escalation, root-cause analysis, and resolution activities. Collaborate with development, DevOps, infrastructure, and client-facing teams to resolve complex production issues. Support production deployments and releases, validating application health before, during, and after deployment activities. Work with Jenkins and CI/CD pipelines to support reliable software delivery and investigate deployment or pipeline failures. Analyze logs, metrics, traces, application behavior, and system dependencies to identify and resolve issues within distributed environments. Identify recurring operational problems and recommend automation, monitoring, configuration, or process improvements to increase system reliability. Contribute to operational documentation, troubleshooting procedures, knowledge sharing, and continuous improvement initiatives. Maintain clear and timely communication with technical stakeholders throughout production incidents and resolution efforts. Required Qualifications & Technical Skills Strong professional experience in Senior Application Support, Production Support, Systems Engineering, Site Reliability Engineering, or a related discipline. Strong hands-on experience supporting production applications and distributed systems. Extensive experience with Google Cloud Platform (GCP) – mandatory. Hands-on experience with Google Kubernetes Engine (GKE) and Kubernetes-based production environments. Strong understanding of observability and monitoring concepts, including: Metrics Logs Alerting Dashboards Distributed tracing Hands-on experience with one or more observability platforms such as Grafana, Prometheus, Datadog, Dynatrace, Splunk, or similar. Experience with Jenkins and CI/CD processes. Demonstrated experience supporting production deployments and troubleshooting deployment-related issues. Strong experience with incident management and production troubleshooting. Ability to diagnose application performance, latency, availability, and reliability issues across distributed environments. Strong analytical and problem-solving skills. Excellent communication and collaboration skills. Ability to work effectively in a fast-paced production environment where application availability and reliability are critical. Preferred Qualifications Experience supporting large-scale retail or eCommerce platforms. Experience working with microservices architectures and cloud-native applications. Experience with multiple observability platforms and the ability to determine the appropriate monitoring strategy for different applications and services. Experience with distributed tracing and troubleshooting complex service-to-service dependencies. Experience working in highly available, business-critical production environments. Familiarity with operational automation and practices designed to improve reliability and incident response. Experience working collaboratively with software development, DevOps, infrastructure, and platform engineering teams. Core Technical Skills Required: GCP GKE Kubernetes Jenkins Production Deployments Application Support Production Troubleshooting Incident Management Observability / Monitoring Distributed Systems Preferred: Grafana Prometheus Datadog Dynatrace Splunk Distributed Tracing Microservices CI/CD Cloud-Native Applications Retail / eCommerce Platforms Work Arrangement & Candidate Requirements This is a hybrid position based in Toronto, Ontario, with an expectation of working onsite at the office 2–3 days per week. Occasional local visits to the client site may also be required. Candidates who are unable to complete the interview process onsite with Publicis Sapient must be prepared to retrieve their company laptop from the Publicis Sapient Toronto office on their start date. Interview Process The expected interview process consists of: 1–2 internal interview rounds with Publicis Sapient. Final client interview with the end client. What Success Looks Like The successful candidate will help ensure stable, observable, and reliable production operations by quickly identifying issues, coordinating effective incident resolution, maintaining strong visibility into application health, and supporting safe and reliable production deployments. The ideal engineer will bring a proactive production-support mindset, strong GCP and Kubernetes expertise, and the ability to use observability data to quickly identify root causes. They will also contribute to continuous improvement by identifying recurring issues and implementing solutions that improve application reliability, performance, and operational efficiency.

Ce que vous ferez

Provide hands-on production support for cloud-native, Kubernetes-based applications in a large-scale retail environment. Responsibilities include monitoring application health, triaging production incidents, and improving observability solutions to ensure system reliability.

Exigences

Requires extensive experience with GCP, GKE, and Kubernetes, along with proficiency in observability tools and Jenkins CI/CD pipelines. Candidates must have a strong background in incident management and troubleshooting distributed systems.

Compétences indiquées

  • KubernetesSouhaitée
  • CI/CDSouhaitée

Autres compétences pertinentes

Relevées dans la description du poste. Confirmez les exigences importantes ci-dessus.

  • Google Cloud Platform
  • Google Kubernetes Engine
  • Kubernetes
  • Jenkins
  • Production Support
  • Incident Management
  • Observability
  • Monitoring
  • Distributed Systems
  • CI/CD
  • Grafana
  • Prometheus
  • Datadog
  • Dynatrace
  • Splunk
  • Distributed Tracing

Domaines d’emploi

  • Technology
  • Software
  • Engineering
  • Retail
  • Consulting

Renseignements supplémentaires

Expérience minimale
5+ ans
Postuler avant le
17 sept. 2026
Langue de l’offre
anglais
Heures de travail
40 heures par semaine
Présence au bureau
3 jours par semaine
Niveau d’expérience
Mid-Senior level
Mode de candidature
La candidature directe est offerte