NVIDIA Logo

NVIDIA

Distinguished Engineer, Scaled Out Inferencing

Posted One Month Ago
Be an Early Applicant
In-Office or Remote
Hiring Remotely in CA, USA
320K-489K Annually
Expert/Leader
In-Office or Remote
Hiring Remotely in CA, USA
320K-489K Annually
Expert/Leader
Lead architecture and global strategy for large-scale, low-latency distributed AI inference. Design high-throughput model serving, hardware-software co-optimization, orchestration for deployment/versioning/auto-scaling across cloud and datacenters, collaborate with customers and open-source ecosystems, and drive full software and system lifecycle for production reliability and performance.
The summary above was generated by AI

NVIDIA is leading the industry in delivering accelerated computing in cloud and enterprise environments. We’re a team of innovative engineers dedicated to solving some of the world’s biggest challenges, constantly driving advancements, and impacting millions of lives worldwide!
As a technology leader at NVIDIA, you will lead the development of our global strategy for scaled-out AI inferencing. You will architect the high-throughput, low-latency distributed pipelines and model serving strategies required for massive scale and production reliability. You will define and drive the technical roadmap for full-lifecycle, from deployment and versioning to automated scaling, across enterprise and cloud environments. Working with NVIDIA leadership, you will establish the systems and orchestration layers that enable the world’s most advanced AI models to run with peak efficiency on our accelerated computing hardware.

What You’ll Be Doing:

  • Various Architectural Work: Architect distributed pipelines, define and drive the technical implementation of high-throughput, low-latency, distributed inference systems to support massive-scale AI workloads.

  • Collaborate on Cross Domain Disciplines: Hardware-software co-optimization, drive performance tuning at the kernel and driver level, optimizing GPU resource management and hardware acceleration for production-grade model serving.

  • Collaborate on Open Source and Ecosystem Projects: Guide and influence open source projects Dynamo, TensorRT-LLM, and ecosystem projects (vLLM, SGLang, Linux, Kubernetes, Ray) to bring state of the art inferencing on NVIDIA accelerated hardware. 

  • Accelerate Integration: Orchestrate model lifecycles, lead the strategy for full-lifecycle model management, including automated deployment, versioning, and intelligent scaling across varied cloud and datacenter environments.

  • Engage Stakeholders: Collaborate with customers, infrastructure providers, and partners to ensure NVIDIA’s solutions set the industry standard for performance and availability. 

  • Full Software and System Lifecycle: From ideation to architecture, design, development, deployment, operations, and full lifecycle management, lead all technical aspects of planning and continuous evolution of a large technical scope.

What We Need to See:

  • 16+ overall years in technical roles with a recent long-term focus on AI infrastructure and more recent direct experience in large-scale inference orchestration. Proven track record building secure, highly available, and durable production distributed systems.

  • 7-10+ years of leadership experience

  • BS/MS or higher or equivalent experience in systems / software engineering, or related engineering fields 

  • Deep Technical Expertise: Proficiency in GPU architecture, hardware acceleration, and low-level performance tuning (CUDA, kernels) alongside cloud-native architectures for multi-tenant model serving.

  • Proven success delivering high-impact technically complex solutions that achieve high levels of transparency into resource utilization, performance, and operational insights.

  • Technical Leadership: Develop and advance consensus and organizational alignment across technical leadership and the highest level of senior corporate leadership. Ability to synthesize cross-functional needs into architecture and design while guiding internal execution across diverse teams. 

  • Communication and Teamwork: Strong collaboration and influence skills, capable of leading engineering engagement, communicating with peers, partners, and working with high performance and accelerated computing customers.

Ways to Stand Out from the Crowd:

  • Application of Artificial Intelligence: Real world experience building the systems to support AI/ML workloads. 

  • Industry Expertise: Direct experience in designing, developing, delivering and operating secure, highly available, scaled out systems in enterprise and cloud environments.

  • Engineering Enablement: Demonstrated history of creating scalable processes and extensible systems that facilitate cross-functional collaboration and operations at scale.

  • Open Source Collaboration: Familiarity with open source ecosystems and projects (e.g. Dynamo, TensorRT-LLM, vLLM, SGLang, Ray). Ability to collaborate and influence in open source project governance to represent NVIDIA, customers, and partners interests in technical alignment and direction.

NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. We have some of the most forward-thinking and hardworking people on the planet working for us. If you're creative, passionate and self-motivated, we want to hear from you!
With competitive salaries and a generous benefits package (www.nvidiabenefits.com), we are widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us and, due to outstanding growth, our best-in-class engineering teams are rapidly growing. If you're a creative and autonomous engineer with a real passion for technology, we want to hear from you!

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 320,000 USD - 488,750 USD.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until August 15, 2026.

This posting is for an existing vacancy. 

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Similar Jobs

7 Minutes Ago
Easy Apply
Remote or Hybrid
USA
Easy Apply
172K-245K Annually
Entry level
172K-245K Annually
Entry level
Cloud • Information Technology • Security • Software • Cybersecurity
Leads Data Security product growth across the AMS region by driving sales strategy, supporting technical and commercial opportunities, enabling field teams, and translating customer insights into product strategy and roadmap improvements. The role covers competitive analysis, pricing, pipeline and churn analysis, customer escalations, deployment support, and data-driven business updates. It focuses on AI-powered data security products, including DLP, DSPM, and AI security solutions.
Top Skills: Ai Agent SecurityAi/MlAispmAWSAzureCasbCloud DlpDnsDspmEmail DlpEndpoint DlpFirewallsGCPGenerative AiHttp/SInline DlpProxiesSAMLScimSecure BrowserSsl/TlsSspmSwgTcpUdpVpnZtna
11 Minutes Ago
Remote
USA
32-40 Hourly
Mid level
32-40 Hourly
Mid level
eCommerce • Retail
Creates performance-driven creative for paid social, YouTube, demand generation, email, and CTV campaigns. Develops video and image advertising assets, collaborates with growth and marketing teams, manages creative workflows, iterates rapidly based on testing, monitors industry trends, and maintains brand consistency while optimizing creative performance in a fast-paced e-commerce environment.
Top Skills: Adobe Creative SuiteAdobe FireflyAfter EffectsCapcutCtvFigmaGoogle AdsIllustratorMeta Ads ManagerPhotoshopPremiereTiktok AdsYoutube
12 Minutes Ago
Remote
USA
144K-280K Annually
Senior level
144K-280K Annually
Senior level
Consumer Web • Healthtech • Professional Services • Social Impact • Software
Owns Headway’s nationwide clinical risk program for behavioral health providers. Responsibilities include safety event intake, triage, investigation, severity and contributing-factor analysis, corrective-action tracking, dashboard ownership, leadership and board reporting, policy and training development, just-culture promotion, and regulatory or accreditation support. The role partners with clinical operations, legal, compliance, product, and data teams to convert event learnings into system improvements and reduce repeat harm.

What you need to know about the Los Angeles Tech Scene

Los Angeles is a global leader in entertainment, so it’s no surprise that many of the biggest players in streaming, digital media and game development call the city home. But the city boasts plenty of non-entertainment innovation as well, with tech companies spanning verticals like AI, fintech, e-commerce and biotech. With major universities like Caltech, UCLA, USC and the nearby UC Irvine, the city has a steady supply of top-flight tech and engineering talent — not counting the graduates flocking to Los Angeles from across the world to enjoy its beaches, culture and year-round temperate climate.

Key Facts About Los Angeles Tech

  • Number of Tech Workers: 375,800; 5.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Snap, Netflix, SpaceX, Disney, Google
  • Key Industries: Artificial intelligence, adtech, media, software, game development
  • Funding Landscape: $11.6 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Strong Ventures, Fifth Wall, Upfront Ventures, Mucker Capital, Kittyhawk Ventures
  • Research Centers and Universities: California Institute of Technology, UCLA, University of Southern California, UC Irvine, Pepperdine, California Institute for Immunology and Immunotherapy, Center for Quantum Science and Engineering

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account