Retour à la recherche
C
Copoly.aiSource d’offres vérifiée

Junior AI/ML Data Engineer (Protein Structures)

Offre en anglais

Build and scale multimodal protein data pipelines to train large-scale generative machine learning models for drug and target discovery. Collaborate with multidisciplinary teams to integrate biological databases and implement novel data featurizations.

  • Sur place
  • Canada
  • Publié 5 août 2026
  • 1 poste

Résumé du poste

At Copoly.ai, we are a dynamic biotech and AI company driving innovation by working on our own proprietary products and developing specialized solutions for our clients in pharma, biotech, and beyond. We are transforming the future of early cancer detection through AI-powered diagnostic solutions. Our flagship product, OncoSage, leverages RNA sequencing and proprietary machine learning algorithms to deliver accurate, blood-based cancer detection. We are committed to advancing the field of oncology through cutting-edge technology, improving patient outcomes, and detecting cancer at its earliest stages. Join us in our mission to make revolutionary strides in healthcare technology. About the Job We are looking for a highly motivated and skilled AI/ML Scientist to join our project team at Copoly. This role is dedicated to building and scaling multimodal protein data pipelines for training state-of-the-art machine learning models that accelerate drug and target discovery efforts. The project spans large-scale generative models, multi-modal reasoning, and functional therapeutic design, with a strong emphasis on scientific discovery and flexible drug discovery workflows. The successful candidate will work in a multidisciplinary environment alongside AI scientists, AI engineers, and computational and wet-lab biologists to advance the frontier of AI and its impact on healthcare outcomes. Key Responsibilities Dataset Design: Build and scale protein data pipelines for training large-scale generative models; inform strategies for data acquisition efforts Multi-modal Data Integration: Integrate and organize metadata from multiple external and internal biological databases, including automated updates and extensibility to new sources Data-enabled Innovation: Work directly with researchers to conceive and implement novel data featurizations and augmentations that enable architectural innovation Quantify Impact: Define data splits and establish quantitative performance metrics to facilitate model benchmarking and hypothesis testing. Scientific Collaboration: Participate in technical discussions, code reviews, and cross-functional planning with ML scientists, engineers, and biologists. Educational Background Education: B.S., Master’s or PhD (Computational) Biology, Biophysics, Chemistry, Computer/Data Science, or a related technical field. Experience: 2+ years of relevant experience in industry, academia, or other technical environments Technical Skills Data Engineering: Strong foundations in data structures, algorithms, and software engineering principles. Biological Data: Deep familiarity with molecular structure databases and file formats (e.g. mmCIF). Programming: High-level proficiency in Python and experience with common packages for manipulating biological data (e.g. Biotite, Biopython). Modeling & Infrastructure: Experience constructing datasets for large-scale ML model training, as well as modern best practices for data splitting and quantitative model performance metrics. Preferred experience: experimental structural biology (X-ray, Cryo-EM, NMR); protein biophysics or macromolecular modeling (e.g. Rosetta, MD); familiarity with antibodies or other pharma-relevant therapeutic modalities, other biological data sources, modern ML libraries (e.g. PyTorch), or relational databases (e.g. Postgres). Soft Skills & Professional Attributes Problem Solving: Proven ability to take full ownership of technical challenges, incorporate stakeholder feedback, and proactively drive solutions end-to-end. Communication: Strong technical communication skills with the ability to articulate complex concepts to both technical and non-technical audiences. Research mindset: You are excited to participate in scientific discussions and motivated to ensure that training data are accurate, rigorous, and information-rich.

Ce que vous ferez

Build and scale multimodal protein data pipelines to train large-scale generative machine learning models for drug and target discovery. Collaborate with multidisciplinary teams to integrate biological databases and implement novel data featurizations.

Exigences

Requires a degree in Computational Biology, Computer Science, or a related field with at least 2 years of relevant experience. Proficiency in Python and familiarity with molecular structure databases and ML infrastructure is essential.

Compétences indiquées

  • PostgreSQLSouhaitée
  • PythonSouhaitée

Autres compétences pertinentes

Relevées dans la description du poste. Confirmez les exigences importantes ci-dessus.

  • Data Engineering
  • Python
  • Protein Data Pipelines
  • Biotite
  • Biopython
  • mmCIF
  • PyTorch
  • Postgres
  • Molecular Structure Databases
  • Generative Models
  • Data Featurization
  • Model Benchmarking
  • X-ray Crystallography
  • Cryo-EM
  • NMR
  • Macromolecular Modeling

Domaines d’emploi

  • Science & Research
  • Data & Analytics
  • Technology
  • Healthcare
  • Engineering

Renseignements supplémentaires

Formation minimale
Baccalauréat
Expérience minimale
2+ ans
Langue de l’offre
anglais
Heures de travail
40 heures par semaine
Niveau d’expérience
Entry level