AI Infrastructure Engineer
Offre en anglaisYou will design, build, and maintain secure, scalable cloud infrastructure to support real-time AI services and customer-facing applications. Additionally, you will improve service reliability through observability, failure testing, and the development of automated deployment systems.
- Sur place
- Toronto, ON
- Publié 12 août 2026
- 1 poste
D’autres postes auxquels postuler directement
Des possibilités semblables publiées par des employeurs qui recrutent sur Jobs.ca, sans formulaire externe.
Forgeahead Solutions Corporation
Technical Lead and Senior Software Engineer
- Sur place
Government of Ontario
Coordonnateur des services de santé; coordonnatrice des services de santé
- Sur place
Bédard Ressources Humaines
Conseiller sénior en ressources humaines
- Sur place
Résumé du poste
Palona’s AI agents operate continuously in production, handle real-time guest interactions, integrate with restaurant systems, and face sharp traffic peaks. Infrastructure is therefore part of the product: latency, reliability, deployment safety, observability, security, and cost directly shape the guest and operator experience. We are looking for an Infrastructure Engineer who combines cloud and reliability depth with strong software engineering judgment. You will build and operate the platform beneath Palona’s AI products, improve how engineers ship, and turn production signals into durable system improvements. This is not a ticket-driven IT or operations role. You will write production code, design systems, automate repetitive work, and own outcomes across the full service lifecycle. Our current environment includes Python services, Docker, AWS and selected Azure services, ECS and Lambda workloads, API Gateway, load balancers, relational data systems, OpenTofu/Terraform, Datadog, and CI/CD automation. We value the ability to learn and make sound tradeoffs more than exact tool-for-tool matching. What you will own: Design, build, and evolve secure, scalable cloud infrastructure for real-time AI services and customer-facing applications. Improve service reliability through clear SLOs, actionable observability, capacity planning, failure testing, and pragmatic incident prevention. Build deployment and release systems that make production changes fast, repeatable, auditable, and safe. Own infrastructure as code, environment consistency, and reusable platform patterns across development, staging, and production. Partner with product and AI engineers on architecture, performance, data flows, and operational readiness for new capabilities. Diagnose complex distributed-system failures across application, network, database, model-provider, and third-party integration boundaries. Reduce infrastructure and model-serving cost without compromising customer experience or engineering velocity. Strengthen secrets management, access controls, backup and recovery, vulnerability management, and other practical security foundations. Build internal tooling and paved paths that let engineers ship and operate services with less manual work. Participate in incident response and turn incidents into better systems, automation, documentation, and engineering judgment. 3+ years industrial experience in relevant technical domain. Strong software engineering fundamentals and experience building or operating production distributed systems. Hands-on experience with a major cloud platform; AWS experience is especially relevant. Experience with containers, infrastructure as code, CI/CD, monitoring, alerting, and production debugging. Ability to write reliable automation and services in Python or another modern programming language. Sound judgment around availability, latency, scalability, security, and cost tradeoffs. A track record of taking ambiguous operational problems from diagnosis through durable resolution. Clear communication during architecture reviews, launches, and incidents. AI-native working habits and curiosity about the operational behavior of LLM- and agent-powered systems. Competitive Salary and Stock Option Plan. Medical, dental, vision, retirement, leave, and disability benefits as applicable. Family Leave Short Term & Long Term Disability Paid time off and company holidays. Learning and development support.
Ce que vous ferez
You will design, build, and maintain secure, scalable cloud infrastructure to support real-time AI services and customer-facing applications. Additionally, you will improve service reliability through observability, failure testing, and the development of automated deployment systems.
Exigences
Candidates must have at least 3 years of industrial experience in distributed systems and cloud platforms, specifically AWS. Strong software engineering fundamentals, proficiency in Python, and experience with infrastructure as code and containerization are required.
Avantages
• Competitive Salary • Stock Option Plan • Medical Insurance • Dental Insurance • Vision Insurance • Retirement Benefits • Family Leave • Short Term Disability • Long Term Disability • Paid Time Off • Company Holidays • Learning and Development Support
Compétences indiquées
- Microsoft AzureSouhaitée
- CI/CDSouhaitée
- DockerSouhaitée
- Amazon Web ServicesSouhaitée
- TerraformSouhaitée
- PythonSouhaitée
Autres compétences pertinentes
Relevées dans la description du poste. Confirmez les exigences importantes ci-dessus.
- Python
- AWS
- Docker
- Infrastructure as Code
- Terraform
- OpenTofu
- CI/CD
- Datadog
- Distributed Systems
- Cloud Infrastructure
- Observability
- System Reliability
- API Gateway
- Azure
- Security Foundations
- Automation
- Inventory Staging
- Curiosity
- Time Off Management
- Systems Automation
- Infrastructure as Code (IaC)
- AI Agents
- Access Controls
- Artificial Intelligence
- Amazon Web Services
- Microsoft Azure
- Management
- IT Capacity Management
- Customer Service
- Communication
- Relational Databases
- Debugging
- Incident Response
- Scalability
- Python (Programming Language)
- Operations
- Restaurant Operation
- Software Engineering
- Tooling
- Vulnerability Management
- Production Code
- Docker (Software)
Domaines d’emploi
- Technology
- Software
- Engineering
- Data & Analytics
- Infrastructure Engineer
- Artificial Intelligence Engineer (General)
- Software Developers
Renseignements supplémentaires
- Expérience minimale
- 2+ ans
- Langue de l’offre
- anglais
- Heures de travail
- 40 heures par semaine