Retour à la recherche
P
PolarGridSource d’offres vérifiée

GPU Infrastructure Deployment Lead

Offre en anglais

Lead the design, deployment, and rollout of GPU infrastructure clusters from initial requirements to production handoff. Coordinate with colocation providers, hardware vendors, and customers to ensure correct physical and network configurations.

  • Télétravail
  • Canada
  • Publié 10 août 2026
  • 1 poste

Résumé du poste

Role outcome Own the successful design, deployment, and rollout of GPU infrastructure for customers. You are responsible for getting GPU clusters from requirements and proposed architecture through colo coordination, installation, networking, validation, and production handoff. This is not primarily a data-center technician role. You will often coordinate or approve work performed by customers, colocation providers, hardware vendors, and remote-hands teams. For important or complex deployments, you may travel onsite to oversee or participate in rack-and-stack and commissioning. What you own Translate customer workloads and requirements into appropriate GPU cluster configurations. Review and approve server, networking, power, rack, cabling, and storage designs before deployment. Help customers determine what infrastructure they actually need. Coordinate deployment requirements with colocation facilities, hardware vendors, networking providers, and remote-hands teams. Create deployment plans, rack diagrams, cabling plans, BOMs, and implementation checklists. Ensure power, cooling, networking, optics, switching, and physical rack requirements are correct before equipment arrives. Oversee rack-and-stack, cabling, firmware configuration, burn-in, networking, and cluster bring-up. Troubleshoot hardware and infrastructure problems during deployment. Validate clusters before they are handed over for production workloads. Maintain deployment standards and improve the deployment process as the company scales. Act as the technical owner for infrastructure questions during customer deployments. Typical projects Deploy a new 8–64 GPU customer cluster in a third-party colo. Review whether a proposed network topology is appropriate for a multi-node inference or training cluster. Work with a customer to determine GPU count, server configuration, networking, power, and rack requirements. Coordinate a deployment between the server OEM, colo, network provider, and customer engineering team. Diagnose why a newly installed GPU node or fabric is not performing as expected. Travel onsite for a high-value deployment where hands-on oversight is warranted. Ideal background Strong candidates will have experience with several of: GPU servers such as NVIDIA HGX/DGX, GB200/GB300, B200/B300, H100/H200, or similar systems. Data-center server deployment and commissioning. High-speed networking: Ethernet and/or InfiniBand. Switch configuration, optics, transceivers, cabling, NICs, and network topology. Linux systems administration and server hardware troubleshooting. Cluster architecture for AI training or inference. Power, cooling, rack density, and colo deployment constraints. Working directly with customers, vendors, and data-center operators. Managing technical deployments involving multiple external parties. Success looks like Clusters arrive with the correct configuration and infrastructure requirements. Deployment problems are caught before hardware reaches the data center. Customers know exactly what they need to provide. Colo and vendor teams have clear instructions. New clusters move from delivery to production quickly. Deployments do not require senior engineering leadership to constantly intervene. The company develops a repeatable deployment playbook rather than treating every cluster as a one-off project.

Ce que vous ferez

Lead the design, deployment, and rollout of GPU infrastructure clusters from initial requirements to production handoff. Coordinate with colocation providers, hardware vendors, and customers to ensure correct physical and network configurations.

Exigences

Requires strong experience with NVIDIA GPU servers (HGX/DGX), high-speed networking, and data center server commissioning. Candidates must be proficient in Linux systems administration and capable of managing complex technical deployments with external parties.

Autres compétences pertinentes

Relevées dans la description du poste. Confirmez les exigences importantes ci-dessus.

  • GPU Infrastructure Deployment
  • Data Center Commissioning
  • High-Speed Networking
  • InfiniBand
  • Ethernet
  • Linux Systems Administration
  • Cluster Architecture
  • Hardware Troubleshooting
  • Colocation Coordination
  • Rack and Stack
  • Network Topology
  • BOM Creation
  • Firmware Configuration
  • Power and Cooling Management
  • Vendor Management
  • Customer Technical Ownership

Domaines d’emploi

  • Technology
  • Engineering
  • Data & Analytics
  • Management & Leadership
  • Logistics

Renseignements supplémentaires

Expérience minimale
5+ ans
Langue de l’offre
anglais
Heures de travail
40 heures par semaine
Exigences de lieu
Country, Canada
Mode de candidature
La candidature directe est offerte