NVIDIA

Site Reliability Engineer, HPC and LSF

Posted Yesterday

Be an Early Applicant

In-Office or Remote

5 Locations

120K-236K Annually

Junior

In-Office or Remote

5 Locations

120K-236K Annually

Junior

As a Site Reliability Engineer, you'll troubleshoot HPC environments, enhance automation, ensure system reliability, and collaborate to improve chip development processes.

The summary above was generated by AI

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.

As a member of the Hardware Infrastructure Farm team, you will provide leadership in the design and implementation of ground breaking compute clusters that powers all silicon development across NVIDIA. We seek an expert to build and operate these clusters at high reliability, efficiency, and performance and drive foundational improvements and automation to improve engineer's productivity. As a Site Reliability Engineer, you are responsible for the big picture of how our systems relate to each other, we use a breadth of tools and approaches to tackle a broad spectrum of problems. Practices such as limiting time spent on reactive operational work, blameless postmortems and proactive identification of potential outages factor into iterative improvement that is key to both product quality and interesting dynamic day-to-day work. SRE's culture of diversity, intellectual curiosity, problem solving and openness is important to our success. Our organization brings together people with a wide variety of backgrounds, experiences and perspectives. We encourage them to collaborate, think big and take risks in a blame-free environment. We promote self-direction to work on meaningful projects, while we also strive to build an environment that provides the support and mentorship needed to learn and grow

What you’ll be doing:

Troubleshoot incoming support requests in a large-scale HPC environment.
Contribute enhancements to existing deployment automation, configuration management, observability, and operational monitoring and day to day operation through automation.
Ensure compute servers are running correct Operating System and configuration.
Troubleshoot Complex Issues: Perform comprehensive troubleshooting from bare metal to application level, ensuring system reliability and efficiency.
Collaborate with specialist teams to drive issues to closure.
Collaborate with domain experts to improve how our chip development process utilizes our infrastructure.
Directly contribute to the overall quality and improve time to market for our next generation chips.

What we need to see:

Proficient in administering Centos/RHEL Linux distributions.
Understating of container technologies like Docker.
Proficiency in Python and UNIX scripting languages such as bash.
Excellent problem-solving skills, with the ability to analyze complex systems, identify bottlenecks, and implement scalable solutions.
Excellent communication and teamwork skills, with the ability to work effectively with diverse teams and individuals.
BS in Computer Science, similar degree (or equivalent experience) with 2+yrs of relevant post degree experience.
Solid understanding of cluster configuration managements tools such as Ansible.

Ways to stand out from the crowd:

Understanding of key Linux technologies such as NFS, automounter, LDAP, DNS, and TCP/IP networking in Red Hat Linux distribution flavors.
Familiarity with job scheduler administration (e.g. IBM Spectrum LSF or SLURM) and experience building/ operating large scale compute infrastructure.
Knowledge of the FlexLM license management system.
Proficiency in Perl for maintaining legacy automation scripts.
Familiarity with High-Speed Networking (InfiniBand, RDMA, RoCE etc.) and fast, distributed storage systems (Lustre, GPFS, etc.)

#LI-Hybrid

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 120,000 USD - 189,750 USD for Level 2, and 148,000 USD - 235,750 USD for Level 3.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until November 9, 2025.NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Top Skills

Centos,Rhel,Docker,Python,Bash,Ansible

Similar Jobs

Motive

Program Manager

17 Minutes Ago

Easy Apply

Remote

United States

Easy Apply

76K-116K Annually

Mid level

76K-116K Annually

Mid level

Artificial Intelligence • Fintech • Hardware • Information Technology • Sales • Software • Transportation

Manage Motive's global gifting and direct mail programs, partnering with cross-functional teams to enhance customer engagement and pipeline conversion.

Top Skills: Postal.IoReachdeskSalesforceSendoso

Luxury Presence

Staff Product Designer

17 Minutes Ago

Easy Apply

Remote or Hybrid

USA

Easy Apply

185K-230K Annually

Senior level

185K-230K Annually

Senior level

Marketing Tech • Real Estate • Software • PropTech • SEO

Lead design initiatives from discovery to launch, collaborating with product and engineering teams to shape product strategy and enhance user experiences. Mentor designers and improve usability through iterative design.

Top Skills: Figma

Cohere Health

Senior Manager, Actuarial

18 Minutes Ago

Easy Apply

Remote

United States

Easy Apply

150K-195K Annually

Senior level

150K-195K Annually

Senior level

Healthtech • Software

Lead the Customer-Facing Actuarial Team, providing actuarial support and insights to customers, managing a team of actuaries, and collaborating across departments to enhance financial impact evaluations.

Top Skills: ExcelPythonRSQL

What you need to know about the Los Angeles Tech Scene

Los Angeles is a global leader in entertainment, so it’s no surprise that many of the biggest players in streaming, digital media and game development call the city home. But the city boasts plenty of non-entertainment innovation as well, with tech companies spanning verticals like AI, fintech, e-commerce and biotech. With major universities like Caltech, UCLA, USC and the nearby UC Irvine, the city has a steady supply of top-flight tech and engineering talent — not counting the graduates flocking to Los Angeles from across the world to enjoy its beaches, culture and year-round temperate climate.

Key Facts About Los Angeles Tech

Number of Tech Workers: 375,800; 5.5% of overall workforce (2024 CompTIA survey)
Major Tech Employers: Snap, Netflix, SpaceX, Disney, Google
Key Industries: Artificial intelligence, adtech, media, software, game development
Funding Landscape: $11.6 billion in venture capital funding in 2024 (Pitchbook)
Notable Investors: Strong Ventures, Fifth Wall, Upfront Ventures, Mucker Capital, Kittyhawk Ventures
Research Centers and Universities: California Institute of Technology, UCLA, University of Southern California, UC Irvine, Pepperdine, California Institute for Immunology and Immunotherapy, Center for Quantum Science and Engineering