SR

Shivani Rathnam

Ouvert aux possibilités

Cloud DevOps / Site Reliability Engineer

Toronto, ON

St. Clair College

À propos

Cloud DevOps / Site Reliability Engineer with 5+ years of experience supporting highly available cloud-native production environments across AWS, Azure, and Google Cloud Platform (GCP). Experienced in Kubernetes orchestration, Infrastructure as Code using Terraform, CI/CD automation, Linux administration, observability, incident response, and cloud operations. Proven expertise in improving platform reliability, automating operational workflows using Python, Bash, YAML, and JSON, implementing AI-assisted monitoring solutions, and supporting globally distributed production systems through 24×7 follow-the-sun operations.

Compétences

  • Amazon Web Services
  • Docker
  • Google Cloud
  • Kubernetes
  • Linux
  • Microsoft Azure
  • MongoDB
  • MySQL
  • Python
  • Redis
  • Terraform
  • Windows

Expérience

  1. Site Reliability Engineer

    Netskope

    déc. 2023 to août 2026

    • Engineered and operated highly available Kubernetes platforms across AWS and GCP, improving scalability, resilience, and operational consistency across production environments. • Owned Kubernetes lifecycle activities across EKS and GKE, including cluster upgrades, node maintenance, workload migration, capacity validation, and post-upgrade health verification. • Build and enhance Jenkins and GitLab CI/CD workflows for Kubernetes deployments, automating release steps and reducing manual deployment effort. • Developed Python and Bash automation for deployment validation, health checks, infrastructure operations, and recurring production workflows. • Designed observability capabilities using Prometheus, Grafana, Elasticsearch, Kibana, Sumo Logic, and CloudWatch for application, Kubernetes, and infrastructure monitoring. • Engineered alerting strategies around latency, resource utilization, Kafka lag, service availability, and infrastructure health to improve signal quality and operational response. • Drive capacity and performance improvements across Kubernetes clusters by analyzing CPU, memory, workload placement, scaling behavior, and node utilization. • Implemented reliability and automation improvements for distributed services including Kafka, Redis, and ClickHouse, covering deployment, scaling, recovery, and production health validation. • Standardized Kubernetes deployment patterns using Helm charts, manifests, configuration templates, and reusable deployment workflows across production environments. • Improved workload reliability by tuning resource requests and limits, pod placement, scaling behavior, and node utilization across Kubernetes clusters. • Built operational tooling and validation checks to identify deployment failures, unhealthy workloads, infrastructure degradation, and capacity risks before they impact production. • Automated repetitive platform maintenance activities to reduce manual intervention and improve consistency across environments. • Implemented monitoring and health-validation workflows for distributed data services, improving visibility into service performance, infrastructure utilization, and workload behavior. • Contributed to platform reliability improvements by analyzing recurring failure patterns and implementing automation, configuration, and infrastructure changes to prevent repeat incidents.

  2. Azure DevOps Engineer

    Optum

    juill. 2020 to déc. 2021

    • Managed Azure Kubernetes Service (AKS) environments supporting containerized microservices. • Automated deployments using Azure DevOps, YAML pipelines, Docker, Helm, and CI/CD best practices. • Worked with Terraform and Infrastructure as Code. • Developed Python and Bash automation scripts. • Implemented monitoring using Azure Monitor, Azure Log Analytics, Grafana, and Status Page. • Implemented container security scanning and DevSecOps practices. • Managed Kubernetes charts using Helm. Created reproducible builds of the Kubernetes applications, managed Kubernetes manifest files and managed releases of Helm packages. • Wrote Bash and Python scripts for automation tasks such as service restarts, log cleanup, and deployment checks. • Worked closely with cross-functional teams to align infrastructure capabilities with application requirements and mentored team members in using Azure tools effectively.

  3. DevOps Engineer

    Envytee Info solutions

    juill. 2019 to juill. 2020

    • Supported AWS cloud infrastructure using EC2, IAM, S3, VPC, and CloudFormation. • Built and maintained Infrastructure as Code using Terraform to provision and manage cloud resources consistently across environments. Developed Jenkins CI/CD pipelines for automated deployments. • Supported Kubernetes deployments and Helm releases. • Configured and maintained AWS CloudWatch monitoring, log collection, and alerting for cloud infrastructure. • Supported Docker-based application deployments and collaborated on Kubernetes environment management. • Participated in code reviews for Terraform and deployment scripts to maintain Infrastructure as Code standards.

Formation

  1. St. Clair College

    Postgraduate Diploma, Data Analytics

    Windsor, ON, Canada

    2022 to 2023

  2. Jawaharlal Nehru Technological University (JNTUH)

    Bachelor of Technology, Information Technology

    India

    2015 to 2019