Runpod Logo

Runpod

Senior ML Systems Engineer, Inference

Posted An Hour Ago
Remote
Hiring Remotely in USA
150K-220K Annually
Senior level
Remote
Hiring Remotely in USA
150K-220K Annually
Senior level
Lead end-to-end LLM inference performance across serving systems, models, GPU hardware, and workloads. Define rigorous benchmarks, diagnose bottlenecks from scheduling through kernels and interconnects, optimize single- and multi-node deployments, and deliver production-ready runtimes and configurations. Collaborate with product and infrastructure teams, evaluate inference technologies, and implement runtime fixes. The role requires strong Python engineering, production experience with serving engines such as vLLM or SGLang, GPU profiling, and knowledge of modern inference optimization techniques.
The summary above was generated by AI

Runpod is the AI Developer Cloud. More than one million developers, from indie researchers to teams running frontier models in production, use Runpod to experiment, train, fine-tune, deploy, and scale AI on one platform. The platform has processed more than 20 billion inference requests. We closed a $100M Series A in June 2026. We're at an inflection point for AI infrastructure, and we're building the platform the next generation of developers will depend on.

 

We're a small, remote-first team. We take ownership seriously, move fast, and ship work that more than a million developers rely on every day. We're looking for people who care deeply, build with urgency, and want to matter at scale.

 

Learn more in our CEO's funding announcement: https://www.runpod.io/blog/one-million-developers.

We're looking for a ML Systems Engineer, Inference. We want Runpod to be the best place in the world to run LLM inference, meaning the fastest and the most cost-efficient. You'll lead that effort. You'll own LLM serving performance end to end. That means measuring it, understanding it, and improving it across models, hardware generations, and workloads. The work you ship will show up directly in the latency and cost our customers experience. This is a hands-on engineering role for someone who likes finding the real bottleneck and fixing it, then turning that fix into something that runs reliably in production.

Responsibilities
  • Define how we measure inference performance, including throughput, time to first token, inter-token latency, and cost per token, and build the tooling that makes those measurements rigorous and repeatable.

  • Profile and diagnose performance problems across the serving stack, from scheduling and memory management down to kernels and interconnect.

  • Improve serving efficiency for large, state-of-the-art models on single-node and multi-node GPU deployments.

  • Turn what you learn into production-ready runtimes, configurations, and defaults that customers benefit from automatically.

  • Work closely with product and infrastructure teams to shape how inference is offered on Runpod.

  • Keep up with the fast-moving inference ecosystem, including the open-source community, and decide what's worth adopting, what's worth building, and what's worth contributing back.

  • Trace bottlenecks in the serving engine/runtime and implement fixes when configuration tuning is not enough.

Requirements
  • 5+ years of professional system engineering experience.

  • Deep, hands-on experience with vLLM, SGLang (or a comparable serving engine) in production or at serious benchmark scale.

  • Strong software engineering skills in Python. You're comfortable working in large, performance-critical codebases.

  • A solid understanding of what drives LLM inference performance: batching, memory, parallelism, and the trade-offs between latency and throughput.

  • Experience with modern inference optimization techniques such as quantization, speculative decoding, or distributed serving.

  • Rigor in benchmarking and performance analysis, plus comfort with GPU profiling tools.

  • The ability to explain your results clearly in writing and turn them into decisions.

Preferred
  • Experience writing or tuning GPU kernels in CUDA or Triton.

  • Contributions to inference or ML systems projects.

  • Experience with multi-node GPU systems and high-speed networking.

  • Experience at a company where inference cost and latency were core business metrics.

What You’ll Receive:

  • The competitive base pay for this position ranges from ($150,000 - $220,000). This salary range may be inclusive of several career levels at Runpod and will be narrowed during the interview process based on a number of factors, including the candidate’s experience, qualifications, and location

  • Meaningful equity in a fast-growing company- everyone on the team receives stock options — your impact drives our growth, and you share in the upside.

  • Generous medical, dental & vision plans

  • Flexible PTO- take the time you need to recharge

  • Most roles are remote work first with an inclusive, collaborative teams utilizing slack as the main form of internal communication

  • Join a passionate team on the cutting edge of AI infrastructure — where culture, learning, and ownership are at the heart of how we scale.

  • $1,200 Home Office & Equipment Stipend- We set you up for success from day one with gear and support to create your ideal workspace

Runpod is committed to maintaining a workplace free from discrimination and upholding the principles of equality and respect for all individuals. We believe that diversity in all its forms enhances our team. As an equal opportunity employer, Runpod is committed to creating an inclusive workforce at every level. We evaluate qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, marital status, protected veteran status, disability status, or any other characteristic protected by law. We welcome every qualified candidate eligible to work in the United States; however, we are currently unable to sponsor employment visas.

Similar Jobs at Runpod

Yesterday
Remote
USA
160K-280K Annually
Senior level
160K-280K Annually
Senior level
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Own a portfolio of fewer than 20 strategic accounts, driving retention, expansion, renewals, and multi-year growth. Build executive relationships, lead strategic reviews, negotiate complex commercial terms, forecast revenue, and maintain account health metrics. Partner with Forward Deployed Engineers to connect technical outcomes with commercial growth, influence internal product priorities, and introduce new capabilities. The role requires periodic travel to strategic customers and industry events.
Top Skills: Ai InfrastructureCloud ComputingDeveloper ToolingGpu CloudHubspotSlack
8 Days Ago
Remote
USA
200K-280K Annually
Senior level
200K-280K Annually
Senior level
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Lead and scale Runpod’s analytics and business intelligence function, managing analysts and BI practitioners while setting the analytics roadmap. Own core company metrics, consumption and marketplace analytics, self-service reporting, experimentation practices, and executive and board-level insights. Partner with Product, Finance, Engineering, Security, Supply, and Go-to-Market leaders to improve data quality, modeling, decision-making, and strategic outcomes.
Top Skills: DagsterDbtModeSnowflakeSoc 2SQL
16 Days Ago
Remote
USA
180K-260K Annually
Senior level
180K-260K Annually
Senior level
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Design, scale, and operate Runpod’s distributed storage platform across network volumes, local NVMe, and S3-compatible object storage. Tune storage and high-speed network performance, lead capacity expansions and migrations, write production automation and control-plane code, extend APIs, and build observability, dashboards, SLOs, and alerts. The role also participates in on-call operations, incident response, infrastructure-as-code practices, and long-term architecture decisions for petabyte-scale AI infrastructure.
Top Skills: APIsCephCi/CdCsiDatadogEthernetGoGpfsGpudirect StorageGrafanaInfinibandInfrastructure As CodeIscsiKubernetesLinuxLustreMinioMoosefsNfsNvmeNvme-OfPrometheusPythonRdmaRoceRustS3SmbSpectrum ScaleStatefulsetsVastWekafsZfs

What you need to know about the Los Angeles Tech Scene

Los Angeles is a global leader in entertainment, so it’s no surprise that many of the biggest players in streaming, digital media and game development call the city home. But the city boasts plenty of non-entertainment innovation as well, with tech companies spanning verticals like AI, fintech, e-commerce and biotech. With major universities like Caltech, UCLA, USC and the nearby UC Irvine, the city has a steady supply of top-flight tech and engineering talent — not counting the graduates flocking to Los Angeles from across the world to enjoy its beaches, culture and year-round temperate climate.

Key Facts About Los Angeles Tech

  • Number of Tech Workers: 375,800; 5.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Snap, Netflix, SpaceX, Disney, Google
  • Key Industries: Artificial intelligence, adtech, media, software, game development
  • Funding Landscape: $11.6 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Strong Ventures, Fifth Wall, Upfront Ventures, Mucker Capital, Kittyhawk Ventures
  • Research Centers and Universities: California Institute of Technology, UCLA, University of Southern California, UC Irvine, Pepperdine, California Institute for Immunology and Immunotherapy, Center for Quantum Science and Engineering

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account