NVIDIA Logo

NVIDIA

Senior Deep Learning Software Engineer, Inference

Posted 22 Days Ago
Be an Early Applicant
Remote
4 Locations
148K-288K
Senior level
Remote
4 Locations
148K-288K
Senior level
The role involves optimizing and analyzing the performance of deep learning models, collaborating on software solutions, and contributing to NVIDIA's inference frameworks.
The summary above was generated by AI

We are now looking for a Senior Deep Learning Software Engineer, Inference! NVIDIA is seeking an experienced Deep Learning Engineer focused on analyzing and improving performance of DL inference! NVIDIA is rapidly growing our research and development for Deep Learning Inference and is seeking excellent Software Engineers at all levels of expertise to join our team. Companies around the world are using NVIDIA GPUs to power a revolution in deep learning, enabling breakthroughs in areas like LLM, Generative AI, Recommenders and Vision that has put DL into every software solution. Join the team that builds the software to enable the performance optimization, deployment and serving of these DL solutions. We specialize in developing GPU-accelerated Deep learning software like vLLM and SGLang, DL benchmarking software and performant solutions to deploy and serve these models.

Collaborate with the deep learning community to implement the latest algorithms for public release in vLLM and SGLang and DL benchmarks. Identify performance opportunities and optimize SoTA DL models across the spectrum of NVIDIA accelerators, from datacenter GPUs to edge SoCs. Implement optimizations using vLLM and SGLang, its open source tools like Polygraphy, vLLM and SGLang plugins, Triton and CUDA kernels. Work and collaborate with a diverse set of teams involving performance modeling, performance analysis, kernel development and inference software development.

What you'll be doing:

  • Performance optimization, analysis, and tuning of DL models in various domains like LLM, Recommender, GNN, Generative AI.

  • Scale performance of DL models across different architectures and types of NVIDIA accelerators.

  • Contribute features and code to NVIDIA’s inference benchmarking frameworks, vLLM and SGLang, Triton and LLM software solutions.

  • Work with cross-collaborative teams across generative AI, automotive, image understanding, and speech understanding to develop innovative solutions.

What we need to see:

  • Masters or PhD or equivalent experience in relevant field (Computer Engineering, Computer Science, EECS, AI).

  • At least 5 years of relevant software development experience.

  • You'll need excellent C/C++ programming and software design skills. SW Agile skills are helpful and Python experience is a plus.

  • Prior experience with training, deploying or optimizing the inference of DL models in production is a plus.

  • Prior background with performance modelling, profiling, debug, and code optimization or architectural knowledge of CPU and GPU is a plus.

  • GPU programming experience (CUDA or OpenCL) is a plus.

GPU deep learning has provided the foundation for machines to learn, perceive, reason and solve problems posed using human language. The GPU started out as the engine for simulating human imagination, conjuring up the amazing virtual worlds of video games and Hollywood films. Now, NVIDIA's GPU runs deep learning algorithms, simulating human intelligence, and acts as the brain of computers, robots and self-driving cars that can perceive and understand the world. Just as human imagination and intelligence are linked, computer graphics and artificial intelligence come together in our architecture. Two modes of the human brain, two modes of the GPU. This may explain why NVIDIA GPUs are used broadly for deep learning, and NVIDIA is increasingly known as “the AI computing company.” Come, join our DL Architecture team, where you can help build the real-time, cost-effective computing platform driving our success in this exciting and quickly growing field.

#LI-Hybrid 

The base salary range is 148,000 USD - 287,500 USD. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions.

You will also be eligible for equity and benefits. NVIDIA accepts applications on an ongoing basis.

NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Top Skills

C++
Cuda
Dl Benchmarks
Dl Models
Gpu
Python
Tensorrt
Triton

Similar Jobs

4 Minutes Ago
Remote
Hybrid
2 Locations
Senior level
Senior level
Big Data • Cloud • Productivity • Software • Database • Analytics • Automation
Lead engineering teams and support the development of SaaS platforms, focusing on team growth, technical decision making, and alignment with company strategy.
Top Skills: SaaS
4 Minutes Ago
Remote
US
125K-200K Annually
Expert/Leader
125K-200K Annually
Expert/Leader
Cloud • Fintech • Food • Information Technology • Software • Hospitality
The Lead Salesforce Developer will oversee Salesforce development tasks, mentor team members, ensure code quality, and actively participate in deployment processes.
Top Skills: ApexGitLightningRestSalesforceSoapVf
4 Minutes Ago
Easy Apply
Remote
Hybrid
United States
Easy Apply
157K-253K Annually
Senior level
157K-253K Annually
Senior level
Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
The Staff Software Engineer will lead security initiatives in IAM, developing tools to enhance security, compliance and promote best practices across the organization.
Top Skills: AWSAzureLinuxUnix

What you need to know about the Los Angeles Tech Scene

Los Angeles is a global leader in entertainment, so it’s no surprise that many of the biggest players in streaming, digital media and game development call the city home. But the city boasts plenty of non-entertainment innovation as well, with tech companies spanning verticals like AI, fintech, e-commerce and biotech. With major universities like Caltech, UCLA, USC and the nearby UC Irvine, the city has a steady supply of top-flight tech and engineering talent — not counting the graduates flocking to Los Angeles from across the world to enjoy its beaches, culture and year-round temperate climate.

Key Facts About Los Angeles Tech

  • Number of Tech Workers: 375,800; 5.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Snap, Netflix, SpaceX, Disney, Google
  • Key Industries: Artificial intelligence, adtech, media, software, game development
  • Funding Landscape: $11.6 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Strong Ventures, Fifth Wall, Upfront Ventures, Mucker Capital, Kittyhawk Ventures
  • Research Centers and Universities: California Institute of Technology, UCLA, University of Southern California, UC Irvine, Pepperdine, California Institute for Immunology and Immunotherapy, Center for Quantum Science and Engineering

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account