Research Scientist, Data
Offre en anglaisOwn the evaluation and data strategy across the training stack to identify capability gaps and improve AI models. Translate complex scientific workflows into rigorous benchmarks and build robust pipelines to ingest and transform large-scale heterogeneous datasets.
- Sur place
- Montréal, QC
- Publié 25 juill. 2026
- 1 poste
D’autres postes auxquels postuler directement
Des possibilités semblables publiées par des employeurs qui recrutent sur Jobs.ca, sans formulaire externe.
BC Public Schools
Manager, People, Performance and Culture
- Sur place
Alcohol and Gaming Commission of Ontario (AGCO)
Responsable de la gestion de l’information
- Sur place
Dalfen Ltée
Technicien Comptable
- Sur place
Résumé du poste
About Periodic Labs The most important scientific discoveries of our time won’t happen in a traditional lab. We’re an AI and physical sciences company building state-of-the-art models to accelerate breakthroughs across materials, energy, and beyond. Backed by world-class investors and growing rapidly, we operate at the pace the frontier requires. Our team brings deep expertise, genuine ownership, and an insatiable drive to push the boundaries of what’s scientifically possible. About the Role You will work on the most important aspect of Scientific AI creation: evaluations and data. This means constructing cutting-edge evaluations based on advanced scientific use cases, sourcing and procuring external datasets, integrating internally generated experimental data into the training stack, constructing training environments for RL. You’ll ensure that the team always has the right assets, in the right shape, to evaluate and improve AI models. You will work with computational and experimental scientists to translate complex scientific workflows into rigorous evaluations and agentic benchmarks, and partner with pretraining, midtraining, and reinforcement learning researchers to identify the data models needed, then build the datasets, environments, and pipelines to deliver it. Your goal will be to create a tight feedback loop between scientific use cases, model evaluation, and training data. What You’ll Do Own the evaluation and data strategy across the training stack, identifying capability gaps and shaping the roadmap with leads of physical science and AI research Work with domain experts to translate advanced scientific workflows into rigorous evals, benchmarks, and RL environments Source, evaluate, and procure external datasets across chemistry, physics, materials science, mathematics, simulations, and laboratory instrumentation Build robust pipelines to ingest, clean, and transform for training large-scale datasets from heterogeneous sources Build tooling and analysis workflows that help researchers inspect data, understand model failures, and determine which evaluations or datasets to develop next You Will Thrive in This Role If You Have Designed evaluations, benchmarks, or RL environments for language models, agents, or scientific AI systems Built large-scale data pipelines for LLM pretraining, midtraining, post-training, or evaluation Strong judgment about dataset and evaluation quality, including scientific relevance, coverage, provenance, licensing, and contamination risks Strong software and data engineering skills, including familiarity with data processing at scale, dataset versioning, lineage tracking A research-oriented mindset: you form hypotheses about data, run controlled experiments, measure model outcomes, and iterate with rigor Research experience in areas such as materials science, solid state chemistry, chemistry, computational physics, semiconductors Mechanics Minimum education: Bachelor’s degree or similar experience Location: Menlo Park, CA or Montreal, Canada. (Soon: San Francisco, too) Compensation: $250,000-350,000 + equity Visa sponsorship: Yes, we sponsor visas.
Ce que vous ferez
Own the evaluation and data strategy across the training stack to identify capability gaps and improve AI models. Translate complex scientific workflows into rigorous benchmarks and build robust pipelines to ingest and transform large-scale heterogeneous datasets.
Exigences
Requires a background in designing evaluations or RL environments for language models and strong software engineering skills for large-scale data processing. Candidates should possess a research-oriented mindset and experience with scientific datasets in chemistry, physics, or materials science.
Avantages
• Equity • Visa sponsorship
Autres compétences pertinentes
Relevées dans la description du poste. Confirmez les exigences importantes ci-dessus.
- Evaluation Design
- Benchmark Construction
- RL Environments
- Data Pipeline Engineering
- LLM Pretraining
- Dataset Versioning
- Lineage Tracking
- Scientific AI
- Data Cleaning
- Data Transformation
- Experimental Design
- Software Engineering
- Materials Science
- Pipelines
- Physical Science
- Language Models
- Workflow Management
- AI Research
- Laboratory Instrumentation
- Data Strategy
- Data Version Control (DVC)
- Research
- Artificial Intelligence
- Data Analysis
- Chemistry
- Data Processing
- Controlled Experiments
- Data Engineering
- Data Modeling
- Experimental Data
- Mathematics
- Mechanics
- Physics
- Tooling
- Reinforcement Learning
- Data Pipelines
- Artificial Intelligence Infrastructure
Domaines d’emploi
- Science & Research
- Data & Analytics
- Technology
- Software
- Engineering
- Research Scientist
- Physical and Engineering Science Technicians Not Elsewhere Classified
- Chemists
Renseignements supplémentaires
- Formation minimale
- Baccalauréat
- Expérience minimale
- 5+ ans
- Langue de l’offre
- anglais
- Heures de travail
- 40 heures par semaine
- Autorisation de travail
- Parrainage de visa mentionné