Protege Logo

Protege

Senior Machine Learning Researcher / Principal Scientist

Posted 7 Days Ago
Remote
Hiring Remotely in USA
Senior level
Remote
Hiring Remotely in USA
Senior level
Lead the evaluation and optimization of large-scale datasets for training AI models, enhancing data quality and performance through statistical and ML methods.
The summary above was generated by AI

Company Overview:

We are building Protege to solve the biggest unmet need in AI — getting access to the right training data. The process today is time intensive, incredibly expensive, and often ends in failure. The Protege platform facilitates the secure, efficient, and privacy-centric exchange of AI training data.

Solving AI's data problem is a generational opportunity. The company that succeeds will be one of the largest in AI — and in tech.

Role Overview

Data is the foundation of AI performance, and we believe model quality starts with data quality. You’ll be at the heart of shaping how we curate, assess, and prepare the training data that powers real-world AI systems.

We’re seeking a Senior Member of the Core Data Team/ Principal Scientist to lead the evaluation and optimization of large-scale datasets used to train state-of-the-art AI models. In this role, you’ll help define what "high-quality data" means in practice, using statistical, computational, and ML-driven methods to ensure our data is diverse, representative, and high-impact. You’ll work closely with research and engineering teams to improve model performance through better data. This is an ideal role for someone with a PhD in machine learning, CS, or a related applied field who is passionate about the role of data in AI training and excited to advance Protege’s mission to become the ubiquitous platform for AI training data.

Key Responsibilities

  • Design and apply statistical and machine learning methods to curate, filter, and enrich large-scale unstructured datasets

  • Develop frameworks to assess data diversity, duplication, and informativeness. Design statistical approaches to de-risk training datasets.

  • Collaborate with model training teams to identify data bottlenecks and optimize dataset performance. Emphasis on ability to collaborate with large foundational models and smaller startups.

  • Provide leadership on data quality strategy and shape internal best practices

  • Evaluate external datasets for integration, focusing on scalability, quality, and relevance to model performance. Help build data scorecards.

  • Contribute to research and development of tools that automate data preprocessing and validation

About You

  • PhD or equivalent Master's Degree + 4+ years industry experience in machine learning, economics, mathematics, engineering, computer science, statistics, or a related quantitative field

  • Strong understanding of AI model training pipelines, including pre-processing and evaluation

  • Experience working with large, unstructured datasets, especially text

  • Background in statistical analysis, bias detection, and data validation

  • Able to identify high-impact problems and drive independent solutions

Bonus if you have these attributes

  • Experience with synthetic data generation or augmentation strategies

  • Publications or open-source contributions in data-centric AI or related areas

  • Experience developing evaluation frameworks or performance metrics for training data

  • Cross-functional collaboration with product, infrastructure, or partnership teams

Top Skills

Machine Learning
Statistical Analysis

Similar Jobs

A Minute Ago
In-Office or Remote
Palo Alto, CA, USA
150K-287K Annually
Senior level
150K-287K Annually
Senior level
Aerospace • Artificial Intelligence • Computer Vision • Software • Analytics • Defense • Big Data Analytics
Manage financial reporting processes, improve data visibility, ensure compliance, and implement best practices across Accounting, Finance, and FP&A.
Top Skills: Accounting SoftwareFinancial Information SystemsMicrosoft Office Suite
A Minute Ago
In-Office or Remote
Palo Alto, CA, USA
161K-307K Annually
Expert/Leader
161K-307K Annually
Expert/Leader
Aerospace • Artificial Intelligence • Computer Vision • Software • Analytics • Defense • Big Data Analytics
The Director of Mission Architecture at Maxar Space leads the development of satellite system solutions, manages technical requirements, and oversees system architecture, while ensuring compliance with intelligence and defense standards.
Top Skills: MS Office
A Minute Ago
In-Office or Remote
Palo Alto, CA, USA
92K-176K Annually
Senior level
92K-176K Annually
Senior level
Aerospace • Artificial Intelligence • Computer Vision • Software • Analytics • Defense • Big Data Analytics
The Lead Contracts Specialist will manage government and commercial contracts, ensuring compliance, negotiation, and monitoring of contract performance while collaborating with internal and external stakeholders.

What you need to know about the Los Angeles Tech Scene

Los Angeles is a global leader in entertainment, so it’s no surprise that many of the biggest players in streaming, digital media and game development call the city home. But the city boasts plenty of non-entertainment innovation as well, with tech companies spanning verticals like AI, fintech, e-commerce and biotech. With major universities like Caltech, UCLA, USC and the nearby UC Irvine, the city has a steady supply of top-flight tech and engineering talent — not counting the graduates flocking to Los Angeles from across the world to enjoy its beaches, culture and year-round temperate climate.

Key Facts About Los Angeles Tech

  • Number of Tech Workers: 375,800; 5.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Snap, Netflix, SpaceX, Disney, Google
  • Key Industries: Artificial intelligence, adtech, media, software, game development
  • Funding Landscape: $11.6 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Strong Ventures, Fifth Wall, Upfront Ventures, Mucker Capital, Kittyhawk Ventures
  • Research Centers and Universities: California Institute of Technology, UCLA, University of Southern California, UC Irvine, Pepperdine, California Institute for Immunology and Immunotherapy, Center for Quantum Science and Engineering

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account