The role focuses on optimizing AI models for efficiency, involving GPU/CPU code profiling, high-performance programming, and developing performance tools.
About Luma AI
About the Role
Responsibilities
Experience
Luma's mission is to build multimodal AI to expand human imagination and capabilities. We believe that multimodality is critical for intelligence. To go beyond language models and build more aware, capable and useful systems, the next step function change will come from vision. So we are working on training and scaling up multimodal foundation models for systems that can see and understand, show and explain, and eventually interact with our world to effect change.
The Performance Optimization team at Luma is dedicated to maximizing the efficiency and performance of our AI models. Working closely with both research and engineering teams, this group ensures that our cutting-edge multimodal models can be trained efficiently and deployed at scale while maintaining the highest quality standards.
- Profile and optimize GPU/CPU/Accelerator code for maximum utilization and minimal latency
- Write high-performance PyTorch, Triton, CUDA, deferring to custom PyTorch operations if necessary
- Develop fused kernels and leverage tensor cores and modern hardware features for optimal hardware utilization on different hardware platforms
- Optimize model architectures and implementations for distributed multi-node production deployment
- Build performance monitoring and analysis tools and automation
- Research and implement cutting-edge optimization techniques for transformer model
- Expert-level proficiency in Triton/CUDA programming and GPU optimization
- Strong PyTorch skills
- Experience with PyTorch kernel development and custom operations
- Proficiency with profiling tools (NVIDIA Nsight, torch profiler, custom tooling)
- Deep understanding of transformer architectures and attention mechanisms
- (Preferred) Experience with compilers/exporters such as torch.compile, TensorRT, ONNX, XLA
- (Preferred) Experience optimizing inference workloads for latency and throughput
- (Preferred) Experience with Triton compiler and kernel fusion techniques
- (Preferred) Knowledge of warp-level intrinsics and advanced CUDA optimization
Your applications are reviewed by real people.
CompensationThe base pay range for this role is $187,500 – $395,000 per year.
About LumaLuma’s mission is to build unified general intelligence that can generate, understand, and operate in the physical world.
We believe that multimodality is critical for intelligence. To go beyond language models and build more aware, capable and useful systems, the next step function change will come from vision. So, we are working on training and scaling up multimodal foundation models for systems that can see and understand, show and explain, and eventually interact with our world to effect change.
Similar Jobs
Cloud • Security • Software • Cybersecurity • Automation
Design, maintain, and extend abuse detection and prevention systems for GitLab's SaaS platform. Serve as a core maintainer of a Ruby on Rails monolith, build anomaly detection and agentic AI mitigations, automate abuse response, collaborate with peer engineering and security teams, and produce runbooks and operational documentation.
Top Skills:
Agentic AiAWSGitlabGoogle Cloud Platform (Gcp)Large Language Models (Llms)RubyRuby On RailsSaaS
eCommerce • Fintech • Hardware • Payments • Software • Financial Services
Lead analytics and measurement for a Sales subdomain: monitor KPIs, investigate performance issues, design experiments, build data foundations and self-serve dashboards, use agentic AI to automate workflows, and communicate results to senior sales, finance, and operations stakeholders.
Top Skills:
A/B TestingAgentic AiAirflowCausal TestingDatabricksDbtLookerMatched-Market TestingNumpyOmniPandasPythonSalesforceSnowflakeSQL
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Detect, contain, and remediate security incidents across Windows, macOS, and Linux. Perform malware analysis and forensic investigations, develop detection and remediation processes, produce customer-facing reports and recommendations, and contribute public thought leadership. Work a hybrid on-site 4x10 schedule with weekend coverage; US work authorization required.
Top Skills:
.NetAi TechnologiesCC#Computer ForensicsDynamic AnalysisIncident ResponseLinuxmacOSMalware AnalysisNetwork Analysis ToolsNetwork ForensicsPerlPythonRuby On RailsStatic AnalysisVisual BasicWindows
What you need to know about the Los Angeles Tech Scene
Los Angeles is a global leader in entertainment, so it’s no surprise that many of the biggest players in streaming, digital media and game development call the city home. But the city boasts plenty of non-entertainment innovation as well, with tech companies spanning verticals like AI, fintech, e-commerce and biotech. With major universities like Caltech, UCLA, USC and the nearby UC Irvine, the city has a steady supply of top-flight tech and engineering talent — not counting the graduates flocking to Los Angeles from across the world to enjoy its beaches, culture and year-round temperate climate.
Key Facts About Los Angeles Tech
- Number of Tech Workers: 375,800; 5.5% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Snap, Netflix, SpaceX, Disney, Google
- Key Industries: Artificial intelligence, adtech, media, software, game development
- Funding Landscape: $11.6 billion in venture capital funding in 2024 (Pitchbook)
- Notable Investors: Strong Ventures, Fifth Wall, Upfront Ventures, Mucker Capital, Kittyhawk Ventures
- Research Centers and Universities: California Institute of Technology, UCLA, University of Southern California, UC Irvine, Pepperdine, California Institute for Immunology and Immunotherapy, Center for Quantum Science and Engineering


.png)
