Luma AI Logo

Luma AI

Research Scientist / Engineer – Reinforcement Learning Infrastructure

Reposted 28 Days Ago
Remote or Hybrid
Hiring Remotely in CA, USA
188K-395K Annually
Senior level
Remote or Hybrid
Hiring Remotely in CA, USA
188K-395K Annually
Senior level
Design, build, and operate large-scale RL post-training systems for multimodal foundation models: distributed training, high-throughput rollout generation, environment and reward infrastructure, evaluation and debugging tooling, and efficiency/stability improvements. Collaborate with researchers to productionize RL methods and run stable RL training across thousands of GPUs.
The summary above was generated by AI
About Luma AI
Luma's mission is to build multimodal AI to expand human imagination and capabilities. We believe that multimodality is critical for intelligence. To go beyond language models and build more aware, capable and useful systems, the next step function change will come from vision. So, we are working on training and scaling up multimodal foundation models for systems that can see and understand, show and explain, and eventually interact with our world to effect change.

About the Role
Reinforcement learning is how our foundation models go from capable to useful — learning to reason, use tools, and act over long horizons. The RL Infrastructure team builds the systems that make this possible at scale: high-throughput distributed training that couples policy optimization with large fleets of inference workers, environments that expose models to realistic multi-step tasks, and the reward, verification, and evaluation systems that turn model behavior into learning signal.

Unlike pretraining, RL at scale is a full-loop systems problem — training, rollout generation, environment execution, and reward computation all run concurrently across thousands of GPUs and must stay fast, stable, and correct together. We are looking for engineers and scientists who have lived this problem: people who have post-trained LLMs with RL, built environments and verifiers from scratch, and debugged what happens when an asynchronous rollout pipeline meets a frontier-scale training run. You will work alongside our research team to design and operate the RL stack for our largest multimodal models.

Responsibilities
  • Design, build, and scale distributed RL post-training systems for large multimodal models — orchestrating trainer, rollout, environment, and reward workloads across thousands of GPUs
  • Build and optimize high-throughput rollout generation, including efficient integration of inference engines (e.g. vLLM, SGLang) into the training loop, weight synchronization, and asynchronous / off-policy training schemes
  • Design and implement RL environments for agentic and multi-step tasks — sandboxed code execution, tool use, computer use, and multimodal interaction — that are reproducible, hermetic, and scalable to millions of episodes
  • Build reward infrastructure: verifiable / programmatic rewards, reward model serving, LLM-as-judge pipelines, and defenses against reward hacking
  • Develop the evaluation, monitoring, and debugging tooling needed to keep large RL runs stable, diagnose convergence and throughput regressions, and understand model behavior mid-run
  • Advance RL training efficiency and stability: sequence packing for long multi-turn trajectories, KV cache reuse across rollouts, curriculum and task sampling, and resource scheduling across heterogeneous training/inference workloads
  • Collaborate closely with researchers to turn new post-training ideas (RLVR, agentic RL, long-horizon credit assignment, self-improvement loops) into production-quality training runs

Experience
  • Hands-on experience post-training LLMs with reinforcement learning (e.g. PPO / GRPO-family methods, RLHF, RLVR / RL from verifiable rewards) at meaningful scale
  • Extensive experience with distributed PyTorch training and parallelization strategies (FSDP, Tensor / Pipeline / Expert Parallel) for foundation models
  • Experience building RL environments, reward functions, verifiers, or evaluation harnesses for LLM agents — including sandboxed execution and multi-turn tool use
  • Deep familiarity with RL post-training frameworks and their systems tradeoffs (e.g. veRL, OpenRLHF, TRL, Ray-based orchestration) and inference engines used for rollouts (vLLM, SGLang)
  • Strong understanding of GPU clusters, networking, and communication libraries (NCCL, MPI), and how they behave under mixed training + inference workloads
  • (Preferred) Experience running RL training across >100 GPUs, including asynchronous or disaggregated trainer/rollout architectures
  • (Preferred) Experience with containerization and orchestration (Kubernetes, Ray) for large environment fleets and sandboxed workloads
  • (Preferred) Research contributions in RL for LLMs — reasoning, agents, reward modeling, or long-horizon tasks — or open-source contributions to RL training frameworks
Compensation
The base pay range for this role is $187,500 – $395,000 per year.
About Luma

Luma’s mission is to build unified general intelligence that can generate, understand, and operate in the physical world.

We believe that multimodality is critical for intelligence. To go beyond language models and build more aware, capable and useful systems, the next step function change will come from vision. So, we are working on training and scaling up multimodal foundation models for systems that can see and understand, show and explain, and eventually interact with our world to effect change.

Similar Jobs

Yesterday
Easy Apply
Remote or Hybrid
Easy Apply
143K-185K Annually
Senior level
143K-185K Annually
Senior level
Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
Design, develop, debug, and test cloud-connected firmware features (C/C++ and Go). Support and scale IoT device fleet, build observability and rollout metrics, triage customer/QA issues, and collaborate cross-functionally to deliver production-ready firmware.
Top Skills: BuildkiteC++CanCan-UtilsDatabricksGoGraphQLLinuxLinux DriversLteSocUsbWifi
Yesterday
Remote or Hybrid
CA, USA
136K-245K Annually
Senior level
136K-245K Annually
Senior level
eCommerce • Fintech • Hardware • Payments • Software • Financial Services
Develop and manage Square’s always-on B2B content engine across social, thought leadership, customer stories, partner content, and industry narratives. Build content calendars, formats, briefs, campaigns, and AI-assisted workflows; coordinate cross-functional teams and external partners; oversee production through launch; and analyze performance to optimize messaging, formats, and distribution.
Top Skills: Ai ToolsLinkedInSocial AnalyticsSocial Media Platforms
Yesterday
In-Office or Remote
CA, USA
136K-245K Annually
Senior level
136K-245K Annually
Senior level
Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
Develop and manage Square’s always-on B2B content engine across social, thought leadership, customer stories, partner content, and industry narratives. Build content calendars, formats, briefs, campaigns, and AI-assisted workflows; collaborate with marketing, sales, creative, communications, and external partners; manage production through launch; and use performance insights to optimize content and distribution.
Top Skills: Ai ToolsLinkedInSocial Media Platforms

What you need to know about the Los Angeles Tech Scene

Los Angeles is a global leader in entertainment, so it’s no surprise that many of the biggest players in streaming, digital media and game development call the city home. But the city boasts plenty of non-entertainment innovation as well, with tech companies spanning verticals like AI, fintech, e-commerce and biotech. With major universities like Caltech, UCLA, USC and the nearby UC Irvine, the city has a steady supply of top-flight tech and engineering talent — not counting the graduates flocking to Los Angeles from across the world to enjoy its beaches, culture and year-round temperate climate.

Key Facts About Los Angeles Tech

  • Number of Tech Workers: 375,800; 5.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Snap, Netflix, SpaceX, Disney, Google
  • Key Industries: Artificial intelligence, adtech, media, software, game development
  • Funding Landscape: $11.6 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Strong Ventures, Fifth Wall, Upfront Ventures, Mucker Capital, Kittyhawk Ventures
  • Research Centers and Universities: California Institute of Technology, UCLA, University of Southern California, UC Irvine, Pepperdine, California Institute for Immunology and Immunotherapy, Center for Quantum Science and Engineering

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account