ChipAgents Logo

ChipAgents

ML Systems Engineer

Posted Yesterday
In-Office
Santa Barbara, CA, USA
150K-350K Annually
Mid level
In-Office
Santa Barbara, CA, USA
150K-350K Annually
Mid level
Optimize and deploy LLM inference systems across multi-node clusters to maximize throughput and minimize latency. Implement and benchmark inference optimizations, profile GPU and memory bottlenecks, build evaluation and benchmarking frameworks, and collaborate with research scientists to integrate new model architectures and techniques into production.
The summary above was generated by AI
About ChipAgents

ChipAgents is redefining the future of chip design and verification with agentic AI workflows. Our platform leverages cutting-edge generative AI to assist engineers in RTL design, simulation, and verification, dramatically accelerating chip development. Founded by experts in AI and semiconductor engineering, we partner with top semiconductor firms, cloud providers, and innovative startups to build intelligent AI agents. The company is a Series A company backed by tier-1 VC firms. ChipAgents is deployed in production to companies that have shipped 16B chips.

Position Overview

We are seeking an ML Systems Engineer to optimize the performance and efficiency of large language model inference powering our agentic AI platform. This is a technical role focused on low-level systems optimization. You will implement performance optimizations, build evaluation harnesses, and architect multi-node clusters for training and inference that push the limits of LLM throughput and latency. Your work will directly impact the responsiveness and cost-efficiency of AI agents used by leading semiconductor companies to design chips.

Key Responsibilities
  • Design, deploy, and optimize LLM inference systems across multi-node clusters, maximizing throughput and minimizing latency for production workloads.

  • Implement and benchmark concrete inference optimizations.

  • Profile and analyze inference bottlenecks at the systems level—from GPU kernel execution to memory bandwidth constraints.

  • Build robust evaluation harnesses and benchmarking frameworks that measure accuracy, throughput, latency, and resource utilization across various parallelism strategies.

  • Collaborate with research scientists to integrate new model architectures and optimizations into production inference infrastructure.

  • Investigate and apply emerging techniques from research papers and open-source projects to continuously improve inference performance.

Qualifications
  • B.S., M.S., or PhD in Computer Science, Electrical Engineering, or related field (or equivalent experience).

  • Experience with large-scale ML systems, GPU computing, or high-performance inference optimization.

  • Strong proficiency in Python and C++/CUDA; hands-on experience with SGLang, vLLM, PyTorch, or similar inference frameworks.

  • Deep understanding of GPU architecture, memory hierarchies, and parallel computing paradigms.

  • Experience deploying and optimizing LLMs in production: model serving, batching strategies, distributed inference, or quantization.

  • Strong systems-level debugging and profiling skills; comfort working at multiple layers of the stack from CUDA kernels to application logic.

  • Familiarity with distributed computing frameworks (Ray, multi-node training/inference) is a plus.

  • Self-directed problem solver who is interested in working on ambitious optimization challenges.

Why Join Us
  • Work on cutting-edge LLM inference optimization problems with real-world production impact.

  • Access to substantial GPU compute resources for experimentation and benchmarking.

  • Collaborate with a world-class team spanning AI research, systems engineering, and EDA.

  • Shape the performance characteristics of AI systems used by leading semiconductor companies.

What we offer
  • $150K/yr – $350K/yr + Offers Equity. We are open to discuss above-scale compensation with exceptional candidates on a case-by-case basis.

  • Unlimited PTO and full benefits (medical, vision, dental, 401k).

  • Two engineering-centric offices with free parking, private gym, and free lunch, drinks and snacks.

 
HQ

ChipAgents Santa Barbara, California, USA Office

Santa Barbara, CA, United States

Similar Jobs

2 Days Ago
Hybrid
140K-217K Annually
Mid level
140K-217K Annually
Mid level
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Design, train, evaluate, and productionize LLMs and agentic AI systems. Build scalable compound AI architectures, improve latency/reliability, ensure responsible AI, and ship production-quality code and datasets.
Top Skills: DpoGoGraph Of ThoughtsHybrid Vector DatabasesLlmsmacOSMultimodal Foundation ModelsPythonRlaifRlhfTree Of Thoughts
2 Days Ago
Hybrid
161K-274K Annually
Senior level
161K-274K Annually
Senior level
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Design, build, and optimize scalable ML infrastructure for training, evaluating, and serving large language models. Automate ML workflows, improve LLM latency and production pipelines, and collaborate cross-functionally to deploy and monitor models at scale.
Top Skills: C++Distributed TrainingETLGoHuggingfaceInference PipelinesLlmsMicroservicesModel EvaluationModel MonitoringPythonPyTorchTensorrt-LlmVllm
2 Days Ago
Hybrid
Senior level
Senior level
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Develop and productionize NLU and agentic AI components using LLMs and multimodal models. Improve model quality, latency, reliability, evaluation, fine-tuning (RLHF/RLAIF/DPO), agent reasoning, multilingual and multimodal capabilities, and responsible AI infrastructure. Collaborate cross-functionally to ship product improvements and maintain high data quality and evaluation standards.
Top Skills: Abstractive SummarizationActive LearningDpoFew-Shot LearningGoGraph Of ThoughtsLarge Language Models (Llms)Multimodal Foundation ModelsPythonRlaifRlhfTree Of ThoughtsVector Databases

What you need to know about the Los Angeles Tech Scene

Los Angeles is a global leader in entertainment, so it’s no surprise that many of the biggest players in streaming, digital media and game development call the city home. But the city boasts plenty of non-entertainment innovation as well, with tech companies spanning verticals like AI, fintech, e-commerce and biotech. With major universities like Caltech, UCLA, USC and the nearby UC Irvine, the city has a steady supply of top-flight tech and engineering talent — not counting the graduates flocking to Los Angeles from across the world to enjoy its beaches, culture and year-round temperate climate.

Key Facts About Los Angeles Tech

  • Number of Tech Workers: 375,800; 5.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Snap, Netflix, SpaceX, Disney, Google
  • Key Industries: Artificial intelligence, adtech, media, software, game development
  • Funding Landscape: $11.6 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Strong Ventures, Fifth Wall, Upfront Ventures, Mucker Capital, Kittyhawk Ventures
  • Research Centers and Universities: California Institute of Technology, UCLA, University of Southern California, UC Irvine, Pepperdine, California Institute for Immunology and Immunotherapy, Center for Quantum Science and Engineering

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account