The most important scientific discoveries of our time won’t happen in a traditional lab. We’re an AI and physical sciences company building state-of-the-art models to accelerate breakthroughs across materials, energy, and beyond. Backed by world-class investors and growing rapidly, we operate at the pace the frontier requires. Our team brings deep expertise, genuine ownership, and an insatiable drive to push the boundaries of what’s scientifically possible.
About the RoleYou will work on the most important aspect of Scientific AI creation: evaluations and data. This means constructing cutting-edge evaluations based on advanced scientific use cases, sourcing and procuring external datasets, integrating internally generated experimental data into the training stack, constructing training environments for RL. You’ll ensure that the team always has the right assets, in the right shape, to evaluate and improve AI models.
You will work with computational and experimental scientists to translate complex scientific workflows into rigorous evaluations and agentic benchmarks, and partner with pretraining, midtraining, and reinforcement learning researchers to identify the data models needed, then build the datasets, environments, and pipelines to deliver it. Your goal will be to create a tight feedback loop between scientific use cases, model evaluation, and training data.
Own the evaluation and data strategy across the training stack, identifying capability gaps and shaping the roadmap with leads of physical science and AI research
Work with domain experts to translate advanced scientific workflows into rigorous evals, benchmarks, and RL environments
Source, evaluate, and procure external datasets across chemistry, physics, materials science, mathematics, simulations, and laboratory instrumentation
Build robust pipelines to ingest, clean, and transform for training large-scale datasets from heterogeneous sources
Build tooling and analysis workflows that help researchers inspect data, understand model failures, and determine which evaluations or datasets to develop next
Designed evaluations, benchmarks, or RL environments for language models, agents, or scientific AI systems
Built large-scale data pipelines for LLM pretraining, midtraining, post-training, or evaluation
Strong judgment about dataset and evaluation quality, including scientific relevance, coverage, provenance, licensing, and contamination risks
Strong software and data engineering skills, including familiarity with data processing at scale, dataset versioning, lineage tracking
A research-oriented mindset: you form hypotheses about data, run controlled experiments, measure model outcomes, and iterate with rigor
Minimum education: Bachelor’s degree or similar experience
Location: Menlo Park, CA or Montreal, Canada. (Soon: San Francisco, too)
Compensation: $250,000-350,000 + equity
Visa sponsorship: Yes, we sponsor visas.
Similar Jobs
What you need to know about the Los Angeles Tech Scene
Key Facts About Los Angeles Tech
- Number of Tech Workers: 375,800; 5.5% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Snap, Netflix, SpaceX, Disney, Google
- Key Industries: Artificial intelligence, adtech, media, software, game development
- Funding Landscape: $11.6 billion in venture capital funding in 2024 (Pitchbook)
- Notable Investors: Strong Ventures, Fifth Wall, Upfront Ventures, Mucker Capital, Kittyhawk Ventures
- Research Centers and Universities: California Institute of Technology, UCLA, University of Southern California, UC Irvine, Pepperdine, California Institute for Immunology and Immunotherapy, Center for Quantum Science and Engineering



