Positron AI Jobs

Software Engineer, Senior

Positron AI

Software Engineer, Senior

Reposted 9 Hours Ago

Remote

Hiring Remotely in United States

150K-250K Annually

Senior level

Remote

Hiring Remotely in United States

150K-250K Annually

Senior level

The Senior Software Engineer will develop high-performance software for executing open-source LLMs on custom hardware, focusing on optimizations and efficient libraries primarily in C++.

The summary above was generated by AI

About Positron AI

Positron AI specializes in developing custom hardware systems to accelerate AI inference. These inference systems offer significant performance and efficiency gains over traditional GPU-based systems, delivering advantages in both performance per dollar and performance per watt. Positron exists to create the world's best AI inference systems.

Role Overview

Senior Software Engineer – Machine Learning Systems & High-Performance LLM Inference

We are seeking a Senior Software Engineer to contribute to the development of high-performance software that powers execution of open-source large language models (LLMs) on our custom appliance. This appliance leverages a combination of FPGAs and x86 CPUs to accelerate transformer-based models. The software stack is written primarily in modern C++ (C++17/20) and heavily relies on templates, SIMD optimizations, and efficient parallel computing techniques.

Key Responsibilities

Design and implement high-performance inference software for LLMs on custom hardware.
Develop and optimize C++-based libraries that efficiently utilize SIMD instructions, threading, and memory hierarchy.
Work closely with FPGA and systems engineers to ensure efficient data movement and computational offloading between x86 CPUs and FPGAs.
Optimize model execution via low-level optimizations, including vectorization, cache efficiency, and hardware-aware scheduling.
Contribute to performance profiling tools and methodologies to analyze execution bottlenecks at the instruction and data flow levels.
Apply NUMA-aware memory management techniques to optimize memory access patterns for large-scale inference workloads.
Implement ML system-level optimizations such as token streaming, KV cache optimizations, and efficient batching for transformer execution.
Collaborate with ML researchers and software engineers to integrate model quantization techniques, sparsity optimizations, and mixed-precision execution.
Ensure all code contributions include unit, performance, acceptance, and regression tests as part of a continuous integration-based development process.

Required Qualifications

7+ years of professional experience in C++ software development, with a focus on performance-critical applications.
Strong understanding of C++ templates and modern memory management.
Hands-on experience with SIMD programming (AVX-512, SSE, or equivalent) and intrinsics-based vectorization.
Experience in high-performance computing (HPC), numerical computing, or ML inference optimization.
Experience with ML model execution optimizations, including efficient tensor computations and memory access patterns.
Knowledge of multi-threading, NUMA architectures, and low-level CPU optimization.
Proficiency with systems-level software development, profiling tools (perfetto, VTune, Valgrind), and benchmarking.
Experience working with hardware accelerators (FPGAs, GPUs, or custom ASICs) and designing efficient software-hardware interfaces.

Preferred Qualifications

Familiarity with LLVM/Clang or GCC compiler optimizations.
Experience in LLM quantization, sparsity optimizations, and mixed-precision computation.
Knowledge of distributed inference techniques and networking optimizations.
Understanding of graph partitioning and execution scheduling for large-scale ML models.

Leveling & Scope

While this role is currently posted at a specific level, we are a growth-oriented organization and are open to hiring at a more senior level for the right candidate. Please note that this job description serves as a focused but generalized overview of the role; specific responsibilities and impact expectations will be tailored to the experience and seniority of the final hire.

Why Join Us?

Work on a cutting-edge ML inference platform that redefines performance and efficiency for LLMs.
Tackle challenging low-level performance engineering problems in AI and HPC.
Collaborate with a team of hardware, software, and ML experts building an industry-first product.
Opportunity to contribute to and shape the future of open-source AI inference software.

Compensation and Benefits

The base salary range for this role is $150,000 – $250,000.

Please note that the figures provided represent the base salary range only and do not include other elements of our total compensation package, equity, or comprehensive benefits.

At Positron AI, we value the unique expertise each candidate brings. While the range above reflects our typical expectation for the position, we reserve the flexibility to exceed this range for candidates whose specialized skills, significant experience, or unique qualifications fall outside the standard scope of the role. Final offers are determined based on a variety of factors, including internal equity, and individual impact.

Equal Opportunity Employer. If you’re excited about the role but don’t meet every bullet, we’d still love to hear from you.

Similar Jobs

ServiceNow

Staff Software Engineer

2 Hours Ago

Remote or Hybrid

191K-334K Annually

Senior level

191K-334K Annually

Senior level

Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation

Lead performance, scalability, and reliability efforts across the platform. Define benchmarks and SLAs; run load, stress, and scalability tests; optimize services, databases, and infrastructure; build observability; lead RCA and long-term improvements; mentor engineers and influence scalable architecture.

Top Skills: AWSGo

ServiceNow

Staff Software Engineer

2 Hours Ago

Remote or Hybrid

172K-301K Annually

Senior level

172K-301K Annually

Senior level

Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation

Lead architecture and roadmap for a multi-agent AI platform, design agent orchestration, context/memory management, and productionize LLM capabilities at enterprise scale. Provide technical leadership, establish test/evaluation frameworks for non-deterministic AI systems, mentor senior engineers, and align cross-functional strategy to deploy secure, highly available agent-driven features.

Top Skills: AutogenAWSAzureDistributed SystemsGCPGoHigh-Throughput Api DesignIamJavaLangchainLlmsMicroservicesModel Context Protocol (Mcp)Python

Upstart

Senior Software Engineer

Yesterday

Easy Apply

Remote

United States

Easy Apply

167K-231K Annually

Senior level

167K-231K Annually

Senior level

Artificial Intelligence • Fintech • Machine Learning • Social Impact • Software

Design, build, and maintain the White Label platform's scalable APIs, services, and UIs. Partner with product and business stakeholders to deliver self-service investor tools, optimize workflows, ensure security, performance, and availability, and participate in code reviews, testing, and deployments.

Top Skills: AWSAzureGCPKafkaKotlinMicroservicesNext.JsPostgresPythonReactRuby On RailsSparkSQLVercel

What you need to know about the Los Angeles Tech Scene

Los Angeles is a global leader in entertainment, so it’s no surprise that many of the biggest players in streaming, digital media and game development call the city home. But the city boasts plenty of non-entertainment innovation as well, with tech companies spanning verticals like AI, fintech, e-commerce and biotech. With major universities like Caltech, UCLA, USC and the nearby UC Irvine, the city has a steady supply of top-flight tech and engineering talent — not counting the graduates flocking to Los Angeles from across the world to enjoy its beaches, culture and year-round temperate climate.

Key Facts About Los Angeles Tech

Number of Tech Workers: 375,800; 5.5% of overall workforce (2024 CompTIA survey)
Major Tech Employers: Snap, Netflix, SpaceX, Disney, Google
Key Industries: Artificial intelligence, adtech, media, software, game development
Funding Landscape: $11.6 billion in venture capital funding in 2024 (Pitchbook)
Notable Investors: Strong Ventures, Fifth Wall, Upfront Ventures, Mucker Capital, Kittyhawk Ventures
Research Centers and Universities: California Institute of Technology, UCLA, University of Southern California, UC Irvine, Pepperdine, California Institute for Immunology and Immunotherapy, Center for Quantum Science and Engineering