Veeda AI Logo

Veeda AI

Senior Machine Learning Infrastructure Engineer (Precision, Diagnostics & Hardware)

Posted Yesterday
Remote or Hybrid
Hiring Remotely in California, USA
Senior level
Remote or Hybrid
Hiring Remotely in California, USA
Senior level
Design, optimize, and maintain large-scale distributed multi-GPU training systems. Diagnose numerical precision and stability issues (FP16/BF16/FP8), detect and recover from hardware/software faults, profile performance bottlenecks, and build developer tooling for resilient checkpointing and rapid fault recovery to maximize researcher productivity and hardware utilization.
The summary above was generated by AI
About Us

Veeda AI is building the next generation of multimodal foundation world models for Physical AI. We're a small, fast-moving team of engineers and researchers from leading AI labs, tackling some of the most challenging problems at the intersection of AI, robotics, and embodied intelligence. If you're excited about pushing the boundaries of what's possible with Physical AI, you'll have the opportunity to make an outsized impact from day one.

Responsibilities
  • Distributed Training Systems & Scalability: Design, optimize, and maintain high-throughput distributed training systems across large-scale GPU clusters for multi-modal foundation models.

  • Precision & Numerical Stability: Debug, diagnose, and resolve subtle numerical instability issues (underflow/overflow, loss spikes, gradient explosion, and mixed-precision divergence) in FP16, BF16, FP8, and custom quantization schemes.

  • Fault Diagnostics & Recovery: Build advanced fault-detection mechanisms and automated diagnostics to rapidly pinpoint and isolate silent data corruption (SDC), hardware hang/deadlock, memory leaks, and "card-freeze" issues during large training runs.

  • Performance Profiling & Optimization: Profile distributed communication bottlenecks, memory usage, and kernel execution to improve overall FLOPS utilization across multi-node, multi-GPU training jobs.

  • Developer Tooling & Infrastructure: Develop resilient checkpointing systems, rapid fault-recovery pipelines, and execution telemetry to keep researcher productivity high and hardware downtime minimal.

Requirements
  • You have a Bachelor's degree or equivalent hands-on experience in Computer Science, Computer Engineering, or a related technical field.

  • You have deep hands-on experience with deep learning training frameworks (e.g., PyTorch) and distributed training paradigms (FSDP, Megatron-LM, DeepSpeed, Tensor Parallelism, Pipeline Parallelism).

  • You have proven experience in numerical precision analysis, low-precision training (BF16/FP8), and debugging complex loss divergence/stability issues in massive training runs.

  • You have strong root-cause analysis skills for hardware/software interaction bugs, including stuck CUDA kernels, NCCL timeouts, GPU hardware faults, and silent training corruptions.

  • You have strong programming skills in Python and C++/CUDA, with a deep understanding of low-level GPU architectures and memory hierarchies.

Nice to Have
  • You have experience running or porting large-scale training workloads on AMD GPUs (ROCm platform) or Google TPUs (JAX/XLA stack).

  • You have contributed to low-level training infrastructure, custom CUDA/Triton kernels, or distributed training open-source projects.

  • You have built resilient fault-tolerant training frameworks with dynamic node re-queueing and rapid checkpointing/saving mechanisms.

Similar Jobs

7 Hours Ago
In-Office or Remote
Site of Old Bullion, NV, USA
177K-294K Annually
Senior level
177K-294K Annually
Senior level
Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
Lead HEOR and evidence strategy for obesity and internal medicine early-pipeline assets. Design and oversee RWD studies, economic models, and patient-reported outcomes; influence trial design; manage evidence generation, publications, vendor relationships, and cross-functional stakeholder engagement to secure reimbursement and patient access.
Top Skills: ChatgptMicrosoft Copilot
7 Hours Ago
In-Office or Remote
Site of Old Bullion, NV, USA
177K-294K Annually
Senior level
177K-294K Annually
Senior level
Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
Lead global HEOR and evidence-generation strategy for obesity assets. Design and execute RWE, economic models, and PRO strategies; create launch dossiers; manage cross-functional teams, vendors, budgets, and external partnerships to support reimbursement and patient access.
Top Skills: ChatgptMicrosoft Copilot
12 Hours Ago
Remote
Senior level
Senior level
Artificial Intelligence • Big Data • Cloud • Information Technology • Software • Cybersecurity • Data Privacy
Own technical relationship for Swiss enterprise and mid-market accounts: run discovery, demos, and technical validations (POC), architect Rubrik solutions across on‑prem, cloud, SaaS, identity and AI data workflows, qualify opportunities, present to technical and CxO stakeholders, and partner with account teams to close deals.
Top Skills: Ai-Driven Data WorkflowsBackup And Disaster RecoveryCloudCyber ResilienceData ProtectionData ResilienceIdentity SecurityObservabilityRemediationRubrikRubrik Agent CloudRubrik Security CloudSaaS

What you need to know about the Los Angeles Tech Scene

Los Angeles is a global leader in entertainment, so it’s no surprise that many of the biggest players in streaming, digital media and game development call the city home. But the city boasts plenty of non-entertainment innovation as well, with tech companies spanning verticals like AI, fintech, e-commerce and biotech. With major universities like Caltech, UCLA, USC and the nearby UC Irvine, the city has a steady supply of top-flight tech and engineering talent — not counting the graduates flocking to Los Angeles from across the world to enjoy its beaches, culture and year-round temperate climate.

Key Facts About Los Angeles Tech

  • Number of Tech Workers: 375,800; 5.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Snap, Netflix, SpaceX, Disney, Google
  • Key Industries: Artificial intelligence, adtech, media, software, game development
  • Funding Landscape: $11.6 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Strong Ventures, Fifth Wall, Upfront Ventures, Mucker Capital, Kittyhawk Ventures
  • Research Centers and Universities: California Institute of Technology, UCLA, University of Southern California, UC Irvine, Pepperdine, California Institute for Immunology and Immunotherapy, Center for Quantum Science and Engineering

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account