Retour à la recherche
BA
Boson AISource d’offres vérifiée

Datacenter Technician

Offre en anglais

You will be responsible for installing, racking, and maintaining physical server and network infrastructure in the datacenter. This includes diagnosing hardware failures, performing preventive maintenance, and ensuring all assets are accurately documented and operational.

  • Sur place
  • Barrie, ON
  • Publié 28 août 2026
  • 1 poste

D’autres postes auxquels postuler directement

Des possibilités semblables publiées par des employeurs qui recrutent sur Jobs.ca, sans formulaire externe.

Résumé du poste

Boson AI is an early-stage startup building large language tools for everyone to use. Our founders (Alex Smola, Mu Li), and a team of Deep Learning, Optimization, NLP, AutoML and Statistics scientists and engineers are working on high quality generative AI models for language and beyond. You keep the machines running We are looking for a Datacenter Hardware Technician to keep the physical infrastructure behind our AI research running in Barrie, ON. You will work hands-on with the latest NVIDIA GPUs, thousands of disks, terabit networking and hundreds of Supermicro servers — racking them, repairing them, keeping firmware current, and making sure every machine that should be online is online. Our researchers train models around the clock, so the difference between a good day and a bad one is often a technician who noticed a marginal cable, logged a serial number correctly, or caught a failing drive before it took a training job down with it. This is careful, methodical work, and we treat it that way. This role is for our Barrie datacenter. You must live within 50 km of the site, such that you can be onsite on short notice. Workload can be bursty, i.e. periods of smooth sailing mixed with periods of intense work during hardware failures, upgrade and maintenance cycles. \n A day in the life Install, rack, cable and commission new Supermicro servers, storage and network equipment. Diagnose and repair hardware failures — replace DIMMs, drives, power supplies, fans, GPUs, cables and mainboards — and drive RMA cases with vendors through to resolution. Install and update drivers, BIOS and firmware across servers, NICs, HBAs and switches, keeping fleet versions consistent and documented. Test and install network connections, including structured cabling, optics and link validation, and troubleshoot physical-layer faults. Perform preventive maintenance: inspections, cable management, airflow and filter checks, and spare-parts inventory. Keep accurate records of every asset, serial number, part swap and rack location. Respond to hardware failure alerts, and escalate to the SRE team when a fault is not purely physical. Follow runbooks and ESD and safety procedures precisely — and improve them where they are unclear. \n $50,000 - $100,000 a year Minimum Qualifications 2+ years of hands-on experience with server hardware in a datacenter, colocation, IT operations or equivalent environment. Practical experience replacing and troubleshooting server components: drives, memory, power supplies, fans, GPUs and mainboards. Experience installing drivers, BIOS and firmware updates on server hardware. Experience testing and installing network connections, including structured cabling. Comfortable on the Linux command line — navigating the filesystem, reading logs, running diagnostics. Familiar with remote access and operational practices: SSH, VPN, and out-of-band management (IPMI, BMC, Redfish). Exceptionally organized and meticulous. Accurate records, careful labelling and consistent procedure matter more here than raw speed. Able to work onsite in Barrie, ON, living within 50 km of the site. Comfortable with the physical demands of the role: lifting and racking equipment, working in cold aisles, and occasional after-hours or on-call response. Clear written communication — your notes are what the next person relies on. Preferred Qualifications Direct experience with Supermicro servers, chassis and IPMI tooling. Experience with GPU servers (NVIDIA H100, A100 or similar) and their power and cooling requirements. Experience with high-speed networking: 100Gb+ Ethernet, InfiniBand, optics and DAC cabling. Familiarity with DCIM or asset-management tools such as NetBox. Basic scripting in Bash or Python to automate repetitive checks. Experience with power distribution units, structured cabling standards, and rack and power design. Experience handling RMA workflows with hardware vendors. Relevant certifications (CompTIA Server+, A+, Network+) or an equivalent hands-on track record. \n If you take pride in a tidy rack, a clean cable run, and a fleet where every machine is accounted for, we'd love to hear from you.

Ce que vous ferez

You will be responsible for installing, racking, and maintaining physical server and network infrastructure in the datacenter. This includes diagnosing hardware failures, performing preventive maintenance, and ensuring all assets are accurately documented and operational.

Exigences

Candidates must have at least 2 years of hands-on experience with server hardware and be comfortable working on the Linux command line. You must reside within 50 km of the Barrie, ON facility and be capable of performing physical tasks like lifting equipment.

Compétences indiquées

  • Network TroubleshootingSouhaitée
  • Preventive MaintenanceSouhaitée
  • Asset ManagementSouhaitée
  • PythonSouhaitée

Autres compétences pertinentes

Relevées dans la description du poste. Confirmez les exigences importantes ci-dessus.

  • Server hardware maintenance
  • Linux command line
  • Hardware troubleshooting
  • Structured cabling
  • Firmware updates
  • BIOS configuration
  • GPU maintenance
  • Network troubleshooting
  • IPMI
  • BMC
  • Redfish
  • Asset management
  • Bash
  • Python
  • RMA management
  • Preventive maintenance
  • Automated Machine Learning
  • Power Distribution Units
  • Generative Artificial Intelligence
  • Workflow Management
  • Musician and Artist Management
  • AI Research
  • Apache Airflow
  • Safety Procedures
  • Research
  • Artificial Intelligence
  • Asset Management
  • Bash (Scripting Language)
  • BIOS
  • Management
  • CompTIA A+
  • CompTIA Network+
  • File Systems
  • Network Interface Controllers
  • Firmware
  • Firmware Updates
  • InfiniBand
  • Information Technology Operations
  • Networking Hardware
  • Inventory Control Systems
  • Virtual Private Networks (VPN)
  • Python (Programming Language)
  • Physical Layers
  • Linux Commands
  • Preventive Maintenance
  • Network Connections
  • Remote Access Systems
  • CompTIA Server+
  • Statistics
  • Structured Cabling

Domaines d’emploi

  • Technology
  • Engineering
  • Data & Analytics
  • Software
  • Data Center Technician
  • Hardware Quality Assurance / Test Engineer
  • Electronics Engineers
  • Computer Hardware Engineers

Renseignements supplémentaires

Formation minimale
Diplôme professionnel
Expérience minimale
2+ ans
Langue de l’offre
anglais
Heures de travail
40 heures par semaine