NVIDIA Logo

NVIDIA

Senior Software Engineer, AI Frameworks

Posted 18 Hours Ago
In-Office or Remote
Hiring Remotely in CA, USA
152K-288K Annually
Senior level
In-Office or Remote
Hiring Remotely in CA, USA
152K-288K Annually
Senior level
Develop production-grade integrations for NVIDIA Grove across AI frameworks including Dynamo, llm-d, Ray, and PyTorch. Build adapters, plugins, operators, runtime components, reference workflows, and developer tooling. Optimize distributed multi-node and multi-GPU training and inference, improve Kubernetes observability, define APIs, and troubleshoot containers, networking, scheduling, CUDA, and runtime issues. Collaborate with engineering teams and open-source communities to upstream changes and maintain compatibility across evolving frameworks.
The summary above was generated by AI

We are seeking a Senior Software Engineer to drive integration of the NVIDIA Grove project within Dynamo and across a set of leading open-source AI frameworks. In this role, you will develop production-grade software enabling Grove capabilities to be adopted, scaled, and operated smoothly. In this role, you will build production-grade software that enables seamless adoption, scaling, and operation of Grove capabilities across environments such as Dynamo, llm-d, Ray, PyTorch, and other emerging frameworks in the AI ecosystem. You will collaborate across engineering teams and the open-source community to deliver robust integrations, reference implementations, and developer-focused tooling.

What you'll be doing:

  • Design and implement end-to-end integrations of Grove with open-source AI frameworks (e.g., Dynamo, llm-d, Ray, PyTorch, and related ecosystem projects).

  • Build and maintain adapters, plugins, operators, and/or runtime components that enable Grove features to work smoothly across training and inference stacks.

  • Partner with framework owners to upstream changes, contribute patches, and ensure long-term maintainability of integrations.

  • Develop reference workflows, sample apps, and best-practice guides that accelerate adoption by users and partners.

  • Optimize performance, scalability, and reliability for distributed training/inference, including multi-node and multi-GPU environments.

  • Improve observability and operational readiness (metrics, logging, tracing, debugging tools) for Kubernetes-based deployments.

  • Participate in technical design reviews, define APIs/contracts, and ensure compatibility across versions of frameworks and dependencies.

  • Diagnose complex issues spanning containers, networking, scheduling, CUDA/GPU utilization, and framework runtime behavior.

What we need to see:

  • BS/MS/PhD in Computer Science, Electrical Engineering, or related field (or equivalent experience)

  • 5+ years of proven experience in related field

  • Hands-on experience integrating with at least one major AI framework/runtime (e.g., PyTorch, Ray, Triton Inference Server ecosystem, distributed runtimes, model serving stacks).

  • Solid understanding of AI workloads: model development basics, training vs. inference tradeoffs, and performance considerations (throughput/latency, batching, memory).

  • Experience with distributed systems concepts (RPC, scheduling, fault tolerance, resource management).

  • Practical Kubernetes experience: deploying and operating services/jobs, Helm/Kustomize, operators/controllers (nice to have), and debugging clusters.

  • Familiarity with containers and cloud-native tooling (Docker, container registries, CI/CD pipelines).

  • Strong software engineering experience in Go, C++ and/or Python, with a track record of shipping reliable systems.

  • Strong interpersonal skills and ability to collaborate across teams and with open-source communities.

  • Exceptional collaboration, communication, and documentation habits.

Ways to stand out from the crowd:

  • Open-source contributions to Dynamo, PyTorch, Ray, llm-d, Kubernetes ecosystem, or related ML infrastructure projects.

  • Experience with large-scale model serving, distributed inference, or multi-tenant AI platforms.

  • Experience building SDKs/APIs or developer tooling that improves integration usability.

  • Knowledge of GPU performance profiling and optimization (Nsight tools or similar), and/or kernel-level performance tuning.

  • Experience with reproducibility, packaging, versioning, and compatibility testing across fast-moving dependencies.

NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most experienced and hard-working people in the world working for us. Are you creative and autonomous? Do you love a challenge? If so, we want to hear from you

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until September 24, 2026.

This posting is for an existing vacancy. 

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Similar Jobs

5 Days Ago
In-Office or Remote
United States
143K-331K Annually
Senior level
143K-331K Annually
Senior level
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Build and operate production-quality AI framework, runtime, benchmarking, performance, automation, observability, and developer-tooling systems. Optimize large language model training and inference across GPUs and Microsoft hardware, diagnose cross-stack reliability and performance issues, and improve model onboarding, hardware utilization, Azure efficiency, and deployment speed. Senior engineers own major components and projects; Principal engineers define technical direction, architecture, and multi-release strategy while leading cross-team initiatives and mentoring technical leaders.
Top Skills: Amd GpusAzureC++CudaNvidia GpusOnnx RuntimePythonPyTorchRocmTensorFlowTriton
6 Days Ago
Remote
United States
120K-261K Annually
Senior level
120K-261K Annually
Senior level
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Designs and operates production-quality AI framework, runtime, benchmarking, performance, automation, and observability components. Benchmarks, profiles, debugs, and optimizes large language model training and inference across GPUs and Microsoft hardware. The role drives projects from definition through deployment, investigates cross-stack reliability issues, partners with research, infrastructure, model, and hardware teams, contributes to technical standards, and mentors engineers.
Top Skills: Amd GpusC++CudaMicrosoft SiliconNvidia GpusPythonSglangTritonVllm
2 Days Ago
Remote
United States
120K-304K Annually
Senior level
120K-304K Annually
Senior level
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Develop and operate production AI framework components, runtimes, benchmarking systems, performance tools, and integrations. Benchmark and optimize large language model training and inference across GPUs and Microsoft hardware, build automation and observability, investigate cross-stack performance and reliability issues, and deliver scalable platform improvements. Principal engineers additionally set technical direction, lead multi-team initiatives, define architectures, and mentor technical leaders.
Top Skills: Ai FrameworksAi RuntimesAzureC++CudaGpusOnnx RuntimePythonPyTorchRocmTensorFlowTriton

What you need to know about the Los Angeles Tech Scene

Los Angeles is a global leader in entertainment, so it’s no surprise that many of the biggest players in streaming, digital media and game development call the city home. But the city boasts plenty of non-entertainment innovation as well, with tech companies spanning verticals like AI, fintech, e-commerce and biotech. With major universities like Caltech, UCLA, USC and the nearby UC Irvine, the city has a steady supply of top-flight tech and engineering talent — not counting the graduates flocking to Los Angeles from across the world to enjoy its beaches, culture and year-round temperate climate.

Key Facts About Los Angeles Tech

  • Number of Tech Workers: 375,800; 5.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Snap, Netflix, SpaceX, Disney, Google
  • Key Industries: Artificial intelligence, adtech, media, software, game development
  • Funding Landscape: $11.6 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Strong Ventures, Fifth Wall, Upfront Ventures, Mucker Capital, Kittyhawk Ventures
  • Research Centers and Universities: California Institute of Technology, UCLA, University of Southern California, UC Irvine, Pepperdine, California Institute for Immunology and Immunotherapy, Center for Quantum Science and Engineering

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account