Hardik Gandhi
Ouvert aux possibilitésSite Reliability Engineer|DevOps Engineer
Ottawa, ON
À propos
Site Reliability Engineer with 3+ years of experience building highly available, scalable, and secure cloud platforms across fintech and SaaS environments using AWS, Kubernetes, Terraform, Docker, Linux, and Infrastructure as Code. Proven expertise in SRE, DevOps, AI-assisted operations, CI/CD automation, observability, incident response, cloud reliability, and distributed systems, delivering resilient production platforms that support millions of users and high-volume transactions. Experienced in improving system reliability, automating infrastructure, optimizing cloud performance, strengthening security, and enabling faster software delivery through GitOps, DevSecOps, and modern platform engineering practices. Strong collaborator who partners with engineering, product, and security teams to solve complex operational challenges, reduce production risk, and build reliable, cost-efficient, and scalable cloud infrastructure that accelerates business growth.
Compétences
- Agile
- Docker
- Git
- Kubernetes
- Linux
- Python
- Redis
- Scrum
- Terraform
Expérience
Site Reliability Engineer
Stripe
avr. 2025 to Aujourd’hui
Canada
• Built AI-assisted incident response workflows with Python, OpenAI APIs, Kubernetes, AWS, and Grafana that automate log correlation, runbook suggestions, and root cause analysis, cutting average resolution time by 42 minutes across 180+ production incidents a year for global platform engineering teams. • Built highly available Kubernetes infrastructure on AWS EKS using Terraform, Helm, ArgoCD, and GitHub Actions to standardize multi-region deployments, supporting 520+ microservices that process nearly 24 million payment transactions daily on Stripe's global merchant platform. • Strengthened platform observability with Prometheus, Grafana, OpenTelemetry, Loki, and CloudWatch dashboards to set up SLI/SLO monitoring, cutting critical alert noise by roughly 1,400 alerts a month and giving SRE and engineering teams clearer operational visibility. • Automated infrastructure provisioning with Terraform modules, AWS CloudFormation, IAM, and policy-as-code, eliminating repetitive manual work and cutting environment provisioning time from 5 hours to under 20 minutes for development and platform teams. • Improved production reliability by designing progressive delivery pipelines with Argo Rollouts, automated canary deployments, health validation, and rollback automation, letting engineering teams safely ship 350+ production releases a month with no customer-facing disruptions. • Optimized Kubernetes resource utilization by tuning Horizontal Pod Autoscaler, Cluster Autoscaler, Redis caching, and container resource allocation, lowering compute consumption by nearly 1,900 vCPUs a month while maintaining 99% availability during global payment traffic peaks.
DevOps Engineer
Zoho
juill. 2021 to juill. 2023
India
• Designed and built enterprise CI/CD pipelines with Jenkins, GitLab CI, Maven, Docker, Kubernetes, and Ansible to automate application delivery, cutting release cycles from 6 hours to about 30 minutes across 85+ production applications for multiple Zoho SaaS products. • Provisioned secure AWS infrastructure with Terraform, Infrastructure as Code, and reusable deployment templates to standardize cloud environments, supporting 260+ EC2 instances, load balancers, and production databases across engineering teams. • Built centralized monitoring and alerting platforms with Prometheus, Grafana, ELK Stack, and CloudWatch, improving infrastructure visibility and reducing high-priority production incidents by 110 a year through proactive health monitoring and alert tuning. • Automated server configuration management, patching, deployments, and operational workflows using Ansible, Python, and Bash, eliminating repetitive manual tasks and improving deployment consistency for DevOps and release engineering teams. • Modernized container orchestration by deploying Kubernetes clusters with Helm, rolling deployments, readiness probes, and autoscaling, supporting 240+ containerized microservices with reliable scalability across production environments. • Integrated SonarQube, Trivy, dependency scanning, and secret detection into CI/CD pipelines, catching and fixing 950+ security and code quality issues before production deployment, improving release quality for development teams.
Formation
Carleton University
Master of Engineering, Computer Engineering
Ottawa, ON
2023 to 2025
Charotar University of Science and Technology.
Bachelor of Technology, Information Technology
Anand, Gujarat, India
2019 to 2023
Permis et certifications
Azure Administrator Associate (AZ-104)
AWS Certified Cloud Practitioner
Oracle Cloud AI Foundations Associate