Retour à la recherche
Logo de Firecrawl
FirecrawlSource d’offres vérifiée

Machine Learning Engineer

Offre en anglais
  • Toronto, ON
  • Hybride
  • Publié 14 sept. 2026
  • 1 poste

210 000 $ US–240 000 $ US / année

Ouvre un site externe

Connectez-vous pour enregistrer ce poste
Type d’emploi
Temps plein
Niveau d’expérience
Intermédiaire · 3+ ans
Langue de l’offre
anglais
Heures de travail
40 heures par semaine

Résumé du poste

You will build and optimize machine learning models for search ranking, relevance, and LLM-driven features within the Firecrawl platform. Additionally, you will design experimentation frameworks, conduct A/B testing, and manage the end-to-end lifecycle of models in production.

Détails du poste

Machine Learning Engineer You'll build the ML behind Firecrawl: the models and the systems that serve them. That starts with search: training and shipping the ranking and relevance models for one of our fastest-growing products, then extending that work across extraction quality and LLM-driven features. You'll also own how we measure: A/B testing launches and building the experimentation frameworks the whole team ships against. If you ship models into production, whether your title says ML engineer or data scientist, this is for you. Salary Range: $250,000–$290,000 USD/year (SF) / $210,000–$224,000 CAD/year (Toronto) Equity Range: Competitive equity. Details shared during the process. Location: San Francisco, CA (SF HQ) or Toronto, ON (Toronto Hub). On-site, five days a week. Job Type: Full-Time Experience: 3+ years building ML or data-heavy systems in production Work Authorization: Must be authorized to work in the United States or Canada. We're not able to sponsor US visas right now. For Canada, we'll consider sponsorship on a case-by-case basis through our Toronto Hub. About Firecrawl Firecrawl is the easiest way to turn the web into data AI agents can use. One API call converts any URL into clean, LLM-ready markdown or structured data. It's the boring-hard problem everyone building with LLMs eventually hits, solved. We hit 8 figures in ARR in year one and more than doubled it in year two. We have 180k+ GitHub stars, putting us in the top 50 repositories of all time, and developers, agents, and category-defining AI companies build on us every day. Growth like this is rare, and we're just getting started. We're a small team punching far above our weight, working out of SF HQ and our new Toronto Hub. Everyone here owns a real piece of the product and company, end to end, and runs it themselves. No hiding behind process or headcount. This is a place for people who want to work at the frontier: an AI company building the infrastructure other AI companies run on, not one bolting AI onto an existing product. We move fast, go deep, and are building the tools superintelligence will rely on to gather data from the web. What You'll Do Improve ranking and relevance for Firecrawl Search, from feature engineering to model training to production Build and tune models for learning-to-rank, query understanding, and LLM-driven retrieval Extend ML across Firecrawl's products: extraction quality, content classification, and evaluation of LLM-driven features Mine query logs and behavioral data at scale to find where our products win and where they fail Build the data pipelines that turn web-scale crawl and query data into training data and features Work hands-on with platform, search, and cloud DevOps engineers to get models running fast and cheap in production Design our testing strategy: the A/B testing frameworks and offline evaluation the team ships against Partner on product launches across Firecrawl: define success metrics, run the experiments, and make the ship/no-ship call on evidence Report on how releases perform post-launch and turn the findings into the next iteration What We're Looking For You've shipped ML models into production systems and owned them after launch: deploying, monitoring, and retraining them, not handing them off You have real ranking or relevance-modeling experience: learning-to-rank, recommendations, or search quality You're comfortable in large, data-heavy systems: query logs, pipelines, and datasets that don't fit in memory You write production-quality code (Python at minimum) and can work inside a real backend codebase You're rigorous about measurement. You've designed and analyzed A/B tests and know when a lift is real You can communicate results clearly to the team: what shipped, what moved, and what to do next Nice to Have MLOps experience: MLflow, experiment tracking, model registries, or feature stores. Kubernetes is a plus Experience building or standardizing an experimentation framework at a previous company Experience with embedding models, vector retrieval, or LLM-based relevance evaluation Experience evaluating LLM outputs at scale: quality scoring, structured-extraction accuracy, or agent behavior Spark or similar large-scale data processing experience What We're NOT Looking For A pure statistician or analyst who needs an engineering team to productionize their work Someone who wants to specialize narrowly and hand off everything else Someone who optimizes for process over shipping A Note On Pace We operate at an absurd level of urgency because the window for what we're building won't stay open forever. If that excites you, keep reading. If it doesn't, no hard feelings, but this role probably isn't for you. Benefits & Perks Available to all employees Salary that makes sense: $250,000–$290,000 USD/year (SF) / $210,000–$224,000 CAD/year (Toronto), based on impact, not tenure Own a piece: Gain competitive equity in what you're helping build Generous PTO: 15 days mandatory, anything after 24 days, just ask (holidays excluded). Take the time you need to recharge Parental leave: 12 weeks fully paid, for all parents Wellness stipend: $100 USD/month for the gym, therapy, massages, or whatever keeps you human Learning & Development: Expense up to $1,000 USD/year toward anything that helps you grow professionally Team offsites: A change of scenery, minus the trust falls Sabbatical: 3 paid months off after 4 years, do something fun and new Available to US-based full-time employees Full coverage, no red tape: Medical, dental, and vision (100% for employees, 50% for partner and kids). No weird loopholes, just care that works Life & Disability insurance: Employer-paid basic life and AD&D, short-term disability, and long-term disability. Coverage for life's curveballs Virtual care and a health guide: Teladoc for the couch doctor visit, plus Rightway to answer coverage questions and fight billing errors for you Mental health: Talkspace, therapy and psychiatry on your schedule Fertility and family building: Carrot, covering you and your partner EAP: Free confidential counseling, legal and financial consults, and online will prep through Guardian 401(k) plan: Retirement might be a ways off, but future-you will thank you Pre-tax benefits: HSA, FSA, and commuter benefits to help your wallet out a bit Supplemental options: Extra life and AD&D, accident, critical illness, hospital indemnity, plus pet, legal, and identity protection through MetLife Available to Canada-based full-time employees Full coverage, no red tape: Extended health, dental, and vision through Manulife (Diamond, the top tier), 100% employer-paid for you, your partner, and your kids Life & Disability insurance: Employer-paid life, AD&D, short-term disability, and long-term disability. Coverage for life's curveballs Virtual care: Dialogue Premium, so you can see a doctor or nurse from your couch, any hour Mental health: Talkspace Elite, therapy and psychiatry on your schedule Fertility and family building: Carrot, covering you and your partner Retirement: Group RRSP through Wealthsimple, so future-you can thank you Available to SF-based employees SF HQ perks: Snacks, drinks, team lunches, intense ping pong, and peak startup energy E-Bike transportation: A loaner electric bike to get you around the city, on us Available to Toronto-based employees Toronto Hub perks: Snacks, drinks, team lunches, glass-walled views down University Avenue, and a home base steps from Union Station Transit, covered: A PRESTO card loaded for GO Transit, subway, and streetcar, plus station parking if you drive to the train. Winter-proof, on us Interview Process Application Review: Send us your work and a quick note on why this excites you. Show us what you've built: ranking models, search systems, experimentation frameworks, pipelines that fed production models. We care about what you've shipped, not where you went to school. Intro Chat (~25 min): A quick conversation to get to know each other before we go deep. We'll talk about what you've been working on, what drew you to Firecrawl, and what you're looking for in your next role. Time for your questions too. Technical Chat (~45 min): We'll dig into a real problem from our world (improving ranking quality with noisy relevance signals, designing the A/B test for a product launch, building features from query logs) and talk through how you'd approach it. Come ready to think out loud. We care how you reason, not whether you memorized the answer. Founder Chat (~25 min): Culture, pace, ownership, and how you like to work. Time for your questions too. Paid Work Trial (~40 Hours): Work with the team on a real, scoped piece of the product, paid at a contractor rate. It's the truest signal for both sides. You see what building at Firecrawl actually feels like, and we see how you ship. Remote-friendly, and we'll flex around your current commitments. Decision: We move fast after the trial. If you want your models ranking results for the whole web, and to see the impact in production the same week, you should join us. 👉 Apply now.

Ce que vous ferez

You will build and optimize machine learning models for search ranking, relevance, and LLM-driven features within the Firecrawl platform. Additionally, you will design experimentation frameworks, conduct A/B testing, and manage the end-to-end lifecycle of models in production.

Exigences

Candidates must have 3+ years of experience building production-grade machine learning or data-heavy systems. Proficiency in Python and experience with ranking, relevance modeling, or large-scale data processing is required.

Avantages

• Health insurance • Dental insurance • Vision insurance • Equity • Paid time off • Parental leave • Wellness stipend • Learning and development stipend • Life insurance • Disability insurance • 401(k) plan • Pet insurance • Commuter benefits • Telehealth

Compétences indiquées

  • Analyse de données · Souhaitée
  • Apprentissage automatique · Souhaitée
  • Python · Souhaitée

Autres compétences pertinentes

Relevées dans la description du poste. Confirmez les exigences importantes ci-dessus.

  • Machine learning
  • Python
  • Ranking models
  • Relevance modeling
  • A/B testing
  • Data pipelines
  • Feature engineering
  • LLM
  • Search quality
  • Recommendation systems
  • Data analysis
  • Production deployment
  • Monitoring
  • Model training
  • Query understanding
  • Pipelines
  • MLOps (Machine Learning Operations)
  • MLflow
  • Query Understanding
  • Cloud DevOps
  • Schema Markup
  • Machine Learning Model Monitoring And Evaluation
  • AI Agents
  • A/B Testing
  • Application Programming Interface (API)
  • Artificial Intelligence
  • Data Processing
  • Github
  • Python (Programming Language)
  • Machine Learning
  • Markdown
  • Quality Scoring
  • Telehealth
  • Punching (Metal Forming)
  • Feature Engineering
  • Data Science
  • Kubernetes
  • Codebase
  • Data Pipelines

Domaines d’emploi

  • Technology
  • Software
  • Data & Analytics
  • Engineering
  • Machine Learning Engineer
  • Software Developers
  • Computer and Information Research Scientists

D’autres postes auxquels postuler directement

Des possibilités semblables publiées par des employeurs qui recrutent sur Jobs.ca, sans formulaire externe.

Voir tous les postes à candidature simplifiée