NVIDIA Logo

NVIDIA

Senior AI Tools Engineer, SRE Operations - GeForce NOW

Posted 2 Days Ago
In-Office or Remote
Hiring Remotely in CA, USA
144K-230K Annually
Senior level
In-Office or Remote
Hiring Remotely in CA, USA
144K-230K Annually
Senior level
Design, build, and deploy AI/ML tools and LLM/agent-based systems for GeForce NOW SRE. Transform production signals, metrics, and logs into actionable intelligence, automate root-cause analysis, manage large-scale data pipelines, and operate ML tooling on Kubernetes and AWS with monitoring and visualization (Grafana).
The summary above was generated by AI

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.

We are seeking a passionate AI Tools Engineer to join the Site Reliability Engineering (SRE) Data Team. Applicants with SRE or equivalent experience are encouraged.

What you will be doing:

You will build and deploy sophisticated AI-powered tools and products. These tools support the operation and optimization of a critical production global Geforce Now service. This role is critical for transforming extensive production data streams—such as signals, metrics, and logs—into actionable intelligence. The intelligence automates root cause analysis for incidents and predicts future service trends and patterns.

  • Build and implement robust AI/ML tools capable of analyzing production data to identify root causes for complex incidents and identify future operational trends.

  • Lead the development of brand-new LLM- and Agent-based systems to improve operational efficiency.

  • Establish and maintain excellent data management practices, including building pipelines to transform and handle large-scale data sources vital for model development.

  • Take charge of and enhance LLM-based pipelines while integrating a strong grasp of LLM progress into product development.

  • Act as a resident authority on AI Frameworks, recommending the best platforms, toolsets, and architectural approaches to ensure the long-term technical sustainability of the product.

What we need to see:

  • B.S. in Computer Science, Statistics, or Engineering (or equivalent experience), and 5+ years of experience.

  • Strong proficiency in Python; familiarity with Go or other systems languages is a plus.

  • Practical experience building, optimizing, and deploying AI tools.

  • Strong knowledge of the AI space and current developments, including understanding how LLM-based platforms are built, optimized, and which platforms work best.

  • Hands-on experience with container orchestration (Kubernetes) and cloud environments (AWS cloud).

  • Active engagement with developments in the AI field and the ability to distinguish meaningful advances from noise when making technical decisions.

  • Expertise in automation and handling large-scale data pipelines.

  • Experience applying monitoring and visualization tools, such as Grafana, to interact with data.

  • Excellent ability to handle data sources and pipelines to transform and manage data.

Ways to stand out in a crowd:

  • Understanding of SRE principles and experience managing production environments.

  • Strong in LLM improvement pipelines as well as a strong grasp of recent developments in LLM training.

  • Someone with excellent knowledge of LLMs and AI Models who can reason and recommend an approach that sustains the team and product long term. This person helps prevent grave mistakes by avoiding the wrong platform choice.

  • Understanding of SRE concepts and managing production environments as well as experience with Kubernetes, AWS, and other cloud technologies.

  • Proficiency in automation.

With a competitive salary package and benefits, NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. Are you a creative and autonomous AI Tools Engineer who loves challenges? Do you have a genuine passion for advancing the state of Site Reliability Engineering across a variety of industries? If so, we want to hear from you.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 144,000 USD - 230,000 USD.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until August 8, 2026.

This posting is for an existing vacancy. 

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Similar Jobs

5 Minutes Ago
Easy Apply
Remote or Hybrid
United States
Easy Apply
160K-195K Annually
Senior level
160K-195K Annually
Senior level
Fintech • Financial Services
Own and evolve the application security program: embed secure SDLC practices, partner with engineering on design and code reviews, manage AppSec tooling (SAST/DAST/ASM/WAF/mobile), harden AWS deployments, integrate security into CI/CD, and lead vulnerability investigation and remediation efforts.
Top Skills: AppdomeAsmAWSCi/Cd PipelinesCloudflare WafCryptographic Key ManagementDastEcsGithub Advanced SecurityGoHadrianIamInvictiMobile Application Security ToolsPythonReact NativeRuby On RailsSastScaSecret ScanningSsl Certificates
8 Minutes Ago
Easy Apply
Remote or Hybrid
United States
Easy Apply
151K-297K Annually
Expert/Leader
151K-297K Annually
Expert/Leader
Big Data • Cloud • Software • Database
Design and implement a distributed query optimization system for MongoDB. Research query systems, architect features, write and review production-quality C++ code, debug large codebases, coordinate cross-team integrations, and mentor engineers while shaping long-term roadmap and performance improvements.
Top Skills: AWSC++CompilersDistributed SystemsGoogle Cloud PlatformAzureMongoDBMongodb AtlasMongodb Query Language (Mql)Query Optimization
An Hour Ago
Remote
USA
224K-336K Annually
Senior level
224K-336K Annually
Senior level
Consumer Web • Healthtech • Professional Services • Social Impact • Software
Lead design and implementation of Headway's insurance systems (eligibility, pricing, matching, claims). Build data-intensive, ML-powered platforms and automation, mentor engineers, write high-impact code, and improve accuracy and reliability across state and payer complexity to scale insurance-covered therapy.
Top Skills: AWSClaudeCursorDatadogEddyFastapiGeminiKafkaMachine LearningNext.JsOcrPagerdutyPostgresPython 3ReactRedisRemixSentrySnowflakeSqlalchemyTemporalTypescript

What you need to know about the Los Angeles Tech Scene

Los Angeles is a global leader in entertainment, so it’s no surprise that many of the biggest players in streaming, digital media and game development call the city home. But the city boasts plenty of non-entertainment innovation as well, with tech companies spanning verticals like AI, fintech, e-commerce and biotech. With major universities like Caltech, UCLA, USC and the nearby UC Irvine, the city has a steady supply of top-flight tech and engineering talent — not counting the graduates flocking to Los Angeles from across the world to enjoy its beaches, culture and year-round temperate climate.

Key Facts About Los Angeles Tech

  • Number of Tech Workers: 375,800; 5.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Snap, Netflix, SpaceX, Disney, Google
  • Key Industries: Artificial intelligence, adtech, media, software, game development
  • Funding Landscape: $11.6 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Strong Ventures, Fifth Wall, Upfront Ventures, Mucker Capital, Kittyhawk Ventures
  • Research Centers and Universities: California Institute of Technology, UCLA, University of Southern California, UC Irvine, Pepperdine, California Institute for Immunology and Immunotherapy, Center for Quantum Science and Engineering

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account