Senior Cloud / Platform Architect
- Calgary, AB
- Sur place
- Publié 10 sept. 2026
- 1 poste
Ouvre un site externe
- Type d’emploi
- Temps plein
- Niveau d’expérience
- Chef d’équipe · 10+ ans
- Postuler avant le
- 10 oct. 2026
- Langue de l’offre
- anglais
- Heures de travail
- 40 heures par semaine
- Niveau d’expérience
- Not Applicable
- Mode de candidature
- La candidature directe est offerte
Résumé du poste
Transition a working prototype of the TIP OpenWiFi cloud SDK and multi-tenant layer into a supportable production service. Lead the design and implementation of device identity renewal, runtime architecture, and site gateway redundancy models.
Détails du poste
This is not a greenfield project. We already run a working platform: the TIP OpenWiFi (uCentral) cloud SDK plus an in-house multi-tenant layer above it (Organisation -> Site -> Device model, canonical device catalogue, Keycloak/OIDC role-based access control, audit trail) and a React operator console. Real access points, switches and site gateways are provisioned end to end in our lab today. You are joining to take that from a working prototype to a supportable production service, and to extend it onto hardware and customers it does not yet support. What the platform does not yet have is an architecture that survives being handed to a customer. Three problems are open and they are yours: device identity has no working renewal path, so certificates expire and devices cannot recover without someone physically reaching them; there is no proven failover story for the equipment we put on a customer site; and the scale ceiling of the control plane is estimated rather than measured. This is a hands-on architecture role. You are expected to write the decision down, then build the thing you decided. Key Requirements: Identity and certificate lifecycle. Device identity is mutual TLS against our own certificate authority. Certificates are currently issued with a short lifetime, the device agent has no in-band renewal command, and a device that expires while offline cannot re-enrol on its own. Own the design and the implementation of a renewal and recovery path that does not depend on physical access, plus expiry as a monitored fleet-wide signal. Runtime architecture. Device connections are raw TLS over TCP, not HTTP. They terminate on a Kubernetes LoadBalancer service with local external traffic policy to preserve source addresses. An HTTP ingress controller in this path is incorrect and breaks mutual TLS. Own the service topology, storage placement across external Ceph and node-local NVMe, and the failure domains. Site gateway architecture. Customer sites terminate on a small x86 appliance running our own OpenWrt-based image: wide area network, guest network, RADIUS and policy enforcement. Own the redundancy model - a failover pair is the target - and the remote recovery path for an appliance nobody can reach. Delivery and environment promotion. Helm chart management for the upstream deployment charts plus our own, delivered through GitOps with a real promotion path between environments, and a documented delta against upstream that stays rebaseable. Observability that pages. Service level objectives, and alerting that reaches a human before a customer does. Dashboards nobody is paged from do not count. Scale envelope. Establish a defensible concurrent-device ceiling by measurement, name the first bottleneck, and re-measure it as the platform changes. Required experience: Production on-premises Kubernetes built and operated with kubeadm or equivalent - not managed cloud only. Certificate authority operations at a practical level: issuing policy and lifetimes, renewal, revocation, and what happens to a device that misses its window. HashiCorp Vault or comparable. Bare-metal load balancing with MetalLB or Cilium, including BGP or layer 2 modes, and source address preservation for non-HTTP services. Ceph as a consumer: block storage CSI configuration, storage class tuning, and diagnosing storage-induced application timeouts. Helm at an authoring level, and GitOps-based delivery with environment promotion. Prometheus, Grafana and Loki, with alerting tied to service level objectives rather than to raw thresholds. The ability to write an architecture decision down and defend it in review. We will ask to read something you wrote. Valuable but not required TIP OpenWiFi deployment charts, or Kafka operations. High availability at the network edge: VRRP or keepalived, multi-WAN, and remote recovery of unattended appliances. Capacity planning for telemetry-heavy workloads. Experience operating a service that other people's hardware depends on. What success looks like in the first 90 days A device certificate renewal path that is implemented, not just documented, including a stated recovery route for a device that expires while offline. The southbound service path running in production with source address preservation verified end to end. A measured storage latency baseline for the relational tier, with a go or no-go call on placement. Alerting that has caught at least one real incident before a user reported it. What We Offer At GuestTek, you won’t just have a job, you will have the opportunity to build your career while working with innovative technology and a global team. Competitive compensation and comprehensive benefits Opportunities for career growth and professional development Exposure to innovative technology, AI, cybersecurity, and global projects Collaborative and supportive work environment Opportunities to work with teams and customers around the world Challenging projects that make a real impact A culture that values innovation, teamwork, and employee contributions
Ce que vous ferez
Transition a working prototype of the TIP OpenWiFi cloud SDK and multi-tenant layer into a supportable production service. Lead the design and implementation of device identity renewal, runtime architecture, and site gateway redundancy models.
Exigences
Requires deep expertise in on-premises Kubernetes, certificate lifecycle management, and bare-metal load balancing. Candidates must be proficient with Helm, GitOps, and observability tools like Prometheus and Grafana.
Avantages
• Competitive compensation • Comprehensive benefits • Opportunities for career growth • Professional development
Compétences indiquées
- Kubernetes · Souhaitée
Autres compétences pertinentes
Relevées dans la description du poste. Confirmez les exigences importantes ci-dessus.
- Kubernetes
- Certificate Authority
- HashiCorp Vault
- MetalLB
- Cilium
- Ceph
- Helm
- GitOps
- Prometheus
- Grafana
- Loki
- BGP
- mTLS
- OpenWrt
- Kafka
- Capacity Planning
Domaines d’emploi
- Technology
- Software
- Engineering
- Consulting
- Data & Analytics
D’autres postes auxquels postuler directement
Des possibilités semblables publiées par des employeurs qui recrutent sur Jobs.ca, sans formulaire externe.
Desjardins
Senior Litigation Advisor -
CommanditéEmployeur directCandidature simplifiée- Hybride
- Calgary, AB
- Publié 9 sept. 2026
BC Public Schools
Manager, Financial Planning and Analysis
CommanditéEmployeur directCandidature simplifiée- Sur place
- Victoria, BC
- Publié 18 sept. 2026
Bédard Ressources Humaines
ITAD Services Representative #1265
CommanditéEmployeur directCandidature simplifiée- Sur place
- Mississauga, ON
- Publié 16 sept. 2026