Optimize large-model inference at scale: improve throughput, reduce latency and cost per token, build benchmarking harnesses, tune parallelism and quantization strategies, implement load-balancing in routing, and evaluate custom kernels and emerging inference hardware.
Dragonfly is a crypto-native Venture Capital and research firm with $3.6B+ in assets under management and 160+ portfolio companies. Our Talent team connects people with roles across our portfolio, opening the door to opportunities through our Talent Network.
This is an application to join our talent network. This is not a listing for an internal role at Dragonfly.
We're actively sourcing for a Senior Inference Optimization Engineer for one of our portfolio companies building privacy-first consumer AI infrastructure. You'll be on the bleeding edge of LLM inference performance, pushing throughput, driving down latency, and optimizing cost per token at significant scale.
Location: Remote, USA (open to excellent candidates outside the USA)
What We’re Looking For:
- 5+ years in performance optimization or HPC with deep GPU architecture and parallel programming knowledge
- Hands-on experience with at least one production LLM inference engine (vLLM, SGLang) running at high volume
- Demonstrated experience with LLM inference optimization: continuous batching, PagedAttention, KV cache management, speculative decoding, quantization, CUDA graphs, torch.compile
- Experience with distributed inference strategies: tensor parallelism, pipeline parallelism, MoE parallelism in multi-GPU and multi-node environments
- GPU profiling fluency: Nsight Systems, Nsight Compute, PyTorch Profiler
- Proficiency in Python, Rust, or Go. C++/CUDA a strong plus
- Bonus: custom Triton kernels, diffusion/image model inference optimization, open-source inference framework contributions
About the role:
- Stand up and optimize GPU infrastructure including B300 nodes in owned data centers
- Drive down TTFT and TPOT, push throughput, and improve cost per token for LLM inference workloads
- Build reproducible benchmarking harnesses across inference engines to identify optimal engine, quantization scheme, and parallelism strategy per workload and GPU SKU
- Optimize multivariate inference load-balancing algorithms within the inference routing system
- Evaluate emerging inference optimization techniques including custom CUDA/Triton kernels, novel attention variants, new quantization schemes, and compilation stack improvements
- Evaluate emerging inference hardware (FPGAs, ASICs, custom silicon) for viability in the stack
Even if you don't match every point above but are an engineer passionate about AI and/or crypto, we encourage you to apply. There may be other opportunities that fit your skill set.
Process:
- We'll review your application and assess fit for this role.
- If there's a match, we'll facilitate a warm introduction to the team.
- If the timing isn't right, we'll keep you in mind for future opportunities across the portfolio.
Compensation
- Early Career: $110,000–$150,000
- Mid: $150,000–$200,000
- Senior: $180,000–$250,000
- Staff/Leadership: $230,000–$330,000
The compensation range reflects variation by seniority, location, and hiring company. Final compensation is confirmed with the specific hiring company.
Submit your information below, and we’ll reach out if there’s a potential fit.
Similar Jobs
Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Leads the strategic evolution of Mastercard’s Cards Lab platform by defining content and intelligence strategy, governance, roadmap priorities, and growth opportunities. Manages content providers, subject-matter experts, research partners, and internal stakeholders to ensure accurate, refreshed, high-quality materials. Develops strategic recommendations, presentations, and business cases for senior leaders, identifies content gaps, and mentors junior team members. The role requires strong product strategy, analytical, communication, and stakeholder-management skills, plus knowledge of payment products and card value propositions.
Top Skills:
ExcelMicrosoft Powerpoint
Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
Leads health technology assessment, value, and evidence strategy for Pfizer’s genitourinary oncology portfolio. Manages HEOR, real-world evidence, economic models, global value dossiers, registries, and evidence dissemination to support reimbursement and patient access. Partners with global, regional, country, and cross-functional oncology teams, oversees vendors and project teams, and communicates findings through publications and conferences.
Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Participate in Mastercard’s 18-month Launch Graduate Program as an Associate Consultant. Work in small teams supporting major brands with strategic business challenges, using technology, data analysis, and marketing expertise. Own client deliverables, build external relationships, evaluate campaign performance, and develop solutions across industries including financial services, retail, restaurants, and hospitality. The program provides structured learning, mentorship, client exposure, and specialized marketing consulting experience.
Top Skills:
ExcelSASSQLTableau
What you need to know about the Los Angeles Tech Scene
Los Angeles is a global leader in entertainment, so it’s no surprise that many of the biggest players in streaming, digital media and game development call the city home. But the city boasts plenty of non-entertainment innovation as well, with tech companies spanning verticals like AI, fintech, e-commerce and biotech. With major universities like Caltech, UCLA, USC and the nearby UC Irvine, the city has a steady supply of top-flight tech and engineering talent — not counting the graduates flocking to Los Angeles from across the world to enjoy its beaches, culture and year-round temperate climate.
Key Facts About Los Angeles Tech
- Number of Tech Workers: 375,800; 5.5% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Snap, Netflix, SpaceX, Disney, Google
- Key Industries: Artificial intelligence, adtech, media, software, game development
- Funding Landscape: $11.6 billion in venture capital funding in 2024 (Pitchbook)
- Notable Investors: Strong Ventures, Fifth Wall, Upfront Ventures, Mucker Capital, Kittyhawk Ventures
- Research Centers and Universities: California Institute of Technology, UCLA, University of Southern California, UC Irvine, Pepperdine, California Institute for Immunology and Immunotherapy, Center for Quantum Science and Engineering


