DDN Storage Logo

DDN Storage

Senior/Staff AI Engineer

Posted Yesterday
Remote or Hybrid
Hiring Remotely in California, USA
Senior level
Remote or Hybrid
Hiring Remotely in California, USA
Senior level
Build and optimize production LLM serving and inference systems, improving GPU and CPU performance, memory usage, KV caching, storage, throughput, and latency. Design scalable infrastructure for RAG and retrieval-heavy workloads while solving distributed systems challenges across compute, memory, and storage. The role requires hands-on ownership of AI infrastructure and the ability to balance architecture decisions with implementation in performance-critical environments.
The summary above was generated by AI

What you’ll do
  • Build and optimize LLM serving and inference systems for production environments

  • Improve performance across GPU and CPU pathways

  • Work on KV cache, memory, storage, and throughput bottlenecks

  • Design and scale systems that support RAG and retrieval-heavy AI workloads

  • Contribute to infrastructure where storage architecture and systems efficiency materially affect AI performance

  • Solve engineering problems at the intersection of AI, high-performance systems, and distributed infrastructure

What we’re looking for
  • An engineer who has spent meaningful time building or optimizing production AI systems, not just experimenting with models

  • Someone who understands how inference performance is shaped by the interaction between compute, memory, storage, and serving architecture

  • Deep hands-on experience working close to the systems layer — for example, improving how workloads run across GPU and CPU resources, reducing bottlenecks, or tuning infrastructure for better throughput and latency

  • Evidence of real ownership in areas like model serving, retrieval, caching, storage, or distributed performance, rather than purely application-layer AI work

  • The ability to move comfortably between architecture decisions and hands-on implementation, especially in environments where efficiency and scale matter

  • A background that suggests you can operate in technically demanding environments, whether that comes from AI infrastructure, high-performance systems, storage platforms, or adjacent distributed systems work

  • PhD preferred, but far less important than having built serious systems in the real world

Why this role is compelling
  • This is not a “prompt engineering” job.

  • This is not an “AI wrapper” job.

  • This is not a generic backend role with AI sprinkled on top.

  • This is a chance to work on the infrastructure that determines whether modern AI systems are fast, scalable, efficient, and commercially viable.

  • If you want to work on the real mechanics of AI performance — serving, retrieval, compute efficiency, memory behavior, storage architecture, and inference at scale — this is where that work happens.

Who will love this role
  • Engineers who enjoy deep systems problems

  • Builders who care about performance, scale, and architecture

  • People who want to work where AI meets infrastructure

  • Candidates who would rather solve hard technical bottlenecks than ship surface-level AI features

Who should not apply

This role is not for:

  • Purely academic researchers without meaningful production ownership

  • Generic software engineers without clear AI systems or inference depth

  • Candidates focused mainly on prompt engineering or lightweight application integrations

  • MLOps generalists who have not worked deeply on serving, storage, or performance-critical AI systems

HQ

DDN Storage California, USA Office

9351 Deering Avenue, CA, United States, 91311

Similar Jobs

2 Days Ago
Remote
United States
242K-288K Annually
Senior level
242K-288K Annually
Senior level
Healthtech • Social Impact • Software • Telehealth
Develop and operate production ML and AI systems powering patient experiences such as matching, ranking, recommendations, onboarding, and personalization. Build reusable ML infrastructure, pipelines, evaluation, observability, and model-serving capabilities. Provide technical leadership across the ML lifecycle, guide architecture, evaluate custom versus foundation-model approaches, and partner with ML and Patient Engineering teams to deliver reliable, scalable AI products.
Top Skills: Artificial IntelligenceExperimentationFeature EngineeringFoundation ModelsGenerative AiMachine LearningMatching SystemsMl InfrastructureMl PipelinesModel EvaluationModel ServingModel TrainingObservabilityPersonalization SystemsPythonRanking SystemsRecommendation SystemsSearch Systems
23 Days Ago
Remote or Hybrid
191K-334K Annually
Senior level
191K-334K Annually
Senior level
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Own technical relationships with strategic enterprise customers throughout Moveworks implementations. Design AI agent solutions, enterprise integrations, workflows, and automations using APIs, iPaaS tools, cloud functions, and scripting. Advise customers on agentic AI roadmaps, coordinate cross-functional delivery, develop reusable capabilities, and mentor engineers. The role combines solution architecture, hands-on implementation, technical consulting, executive communication, and up to 25% travel.
Top Skills: Aws LambdaAzure FunctionsContext EngineeringGoJavaScriptJIRALinuxLlmsMoveworks PlatformOktaPrompt EngineeringPythonRest ApisServicenowServicenow Flow DesignerWindowsWorkatoWorkday
Yesterday
Remote or Hybrid
California, USA
Senior level
Senior level
Artificial Intelligence • Machine Learning • Software • Analytics
Leads hands-on design, development, and optimization of AI data movement and distributed storage systems. Responsibilities include integrating NVIDIA NIXL and DDN Infinia with GPU inference platforms, optimizing GPU-to-storage I/O using GPUDirect Storage, RDMA, and NVMe-over-Fabrics, developing KV cache and multi-tier storage strategies, benchmarking production systems, resolving performance bottlenecks, influencing distributed inference architecture, and mentoring engineers.
Top Skills: CC++Ddn InfiniaDistributed StorageGpu ComputingHpcInfinibandKv Cache ManagementLinuxLlm InferenceNvidia Gpudirect StorageNvidia NixlNvmeNvme-Over-FabricsObject StoragePythonPyTorchRdmaRetrieval-Augmented GenerationScalable File SystemsSsdTensorFlowVector Databases

What you need to know about the Los Angeles Tech Scene

Los Angeles is a global leader in entertainment, so it’s no surprise that many of the biggest players in streaming, digital media and game development call the city home. But the city boasts plenty of non-entertainment innovation as well, with tech companies spanning verticals like AI, fintech, e-commerce and biotech. With major universities like Caltech, UCLA, USC and the nearby UC Irvine, the city has a steady supply of top-flight tech and engineering talent — not counting the graduates flocking to Los Angeles from across the world to enjoy its beaches, culture and year-round temperate climate.

Key Facts About Los Angeles Tech

  • Number of Tech Workers: 375,800; 5.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Snap, Netflix, SpaceX, Disney, Google
  • Key Industries: Artificial intelligence, adtech, media, software, game development
  • Funding Landscape: $11.6 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Strong Ventures, Fifth Wall, Upfront Ventures, Mucker Capital, Kittyhawk Ventures
  • Research Centers and Universities: California Institute of Technology, UCLA, University of Southern California, UC Irvine, Pepperdine, California Institute for Immunology and Immunotherapy, Center for Quantum Science and Engineering

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account