Member of Technical Staff - ML Data
- Toronto, ON
- Hybride
- Publié 7 sept. 2026
- 1 poste
Ouvre un site externe
- Type d’emploi
- Temps plein
- Niveau d’expérience
- Intermédiaire · 2+ ans
- Formation minimale
- Baccalauréat
- Langue de l’offre
- anglais
- Heures de travail
- 40 heures par semaine
- Exigences de lieu
- Country, Switzerland, United States, Canada
Résumé du poste
You will build and maintain large-scale multimodal data pipelines for video, lidar, and robot trajectories using distributed computing frameworks. Additionally, you will be responsible for data curation, auto-labeling, and ensuring strict provenance and licensing compliance for all training datasets.
Détails du poste
Member of Technical Staff - ML Data About Us Veeda AI is building the next generation of multimodal foundation world models for Physical AI. We're a small, fast-moving team of engineers and researchers from leading AI labs, tackling some of the most challenging problems at the intersection of AI, robotics, and embodied intelligence. If you're excited about pushing the boundaries of what's possible with Physical AI, you'll have the opportunity to make an outsized impact from day one. Responsibilities ML Methods for Data: Design, develop, and validate ML methods for data selection, enrichment, annotation, and quality assessment. End-to-End Ownership: Own the full lifecycle, from problem formulation through deployment and continuous improvement. Large-Scale Application: Build reliable workflows for processing large-scale datasets. Evaluation & Experimentation: Measure data quality and assess its impact on model performance. Iterative Improvement: Use failure analysis and feedback to improve data-processing methods. Annotation: Produce and evaluate labels such as captions, camera poses, depth maps, and segmentation masks. Research Collaboration: Partner with researchers and engineers to develop effective data solutions. Requirements Master’s or Ph.D. in Computer Science, Engineering, or a related technical field, or equivalent hands-on experience. Demonstrated ability to develop original ML methods, evidenced by peer-reviewed publications or substantial research contributions with rigorous experimental validation. Experience owning the full lifecycle of an ML method: designing and implementing the approach, applying it to large-scale data, evaluating results, and improving it through successive iterations. Strong Python and PyTorch skills, with experience training, adapting, and evaluating machine learning models. Ability to design controlled experiments, establish meaningful metrics, and analyze errors to guide improvements. Strong software engineering skills, with an emphasis on reproducibility, reliability, and maintainable code. Nice to Have Publications in computer vision, robotics, or related machine learning fields, with a substantial personal contribution. Experience scaling ML inference and data processing with distributed computing tools. Experience optimizing large-scale inference or data-processing workflows. Experience building and operating distributed data pipelines over hundreds of terabytes, with a strong focus on idempotency, backfills, and schema evolution.
Ce que vous ferez
You will build and maintain large-scale multimodal data pipelines for video, lidar, and robot trajectories using distributed computing frameworks. Additionally, you will be responsible for data curation, auto-labeling, and ensuring strict provenance and licensing compliance for all training datasets.
Exigences
Candidates must have a bachelor's degree in Computer Science or equivalent experience in large-scale data engineering. You should possess strong Python skills and proven experience operating distributed data pipelines over hundreds of terabytes.
Compétences indiquées
- Apprentissage automatique · Souhaitée
- Python · Souhaitée
Autres compétences pertinentes
Relevées dans la description du poste. Confirmez les exigences importantes ci-dessus.
- Data Engineering
- Python
- Ray Data
- Daft
- Spark
- GPU Decoding
- Video Processing
- Distributed Systems
- Machine Learning
- Data Curation
- Computer Vision
- Robotics
- Columnar Storage
- Data Pipelines
- Schema Evolution
- Provenance Tracking
- Pipelines
- Cardiac Ablation
- Research
- Artificial Intelligence
- FishEye (Software)
- Computer Science
- Computer Engineering
- Distributed Data Store
- FFmpeg
- Packaging And Labeling
- Governance
- Python (Programming Language)
- Open Source Technology
- Robot Operating Systems
- Safety Assurance
- Statistics
- Tooling
- Curation
- Object Storage
- Feed Forward
- Light Detection And Ranging (LiDAR)
- Honesty
Domaines d’emploi
- Technology
- Data & Analytics
- Software
- Engineering
- Science & Research
- Member of Technical Staff
- Artificial Intelligence Engineer (General)
- Software Developers
D’autres postes auxquels postuler directement
Des possibilités semblables publiées par des employeurs qui recrutent sur Jobs.ca, sans formulaire externe.
Bédard Ressources Humaines
ITAD Services Representative #1265
CommanditéEmployeur directCandidature simplifiée- Sur place
- Mississauga, ON
- Publié 16 sept. 2026
Sustainable Projects Group
Sustainability consultant/ Team Lead, Energy Consulting NOC 41400
CommanditéEmployeur directCandidature simplifiée- Hybride
- vancouver v5l 4s1, BC
- Publié 17 sept. 2026
Desjardins
Senior Litigation Advisor -
CommanditéEmployeur directCandidature simplifiée- Hybride
- Calgary, AB
- Publié 9 sept. 2026