Graphcore Logo

Graphcore

Staff AI Performance Engineer

Posted 49 Minutes Ago
Be an Early Applicant
Hybrid
Milpitas, CA
Expert/Leader
Hybrid
Milpitas, CA
Expert/Leader
Leads performance optimization for large-scale AI training and inference workloads across distributed systems. Builds benchmarks, profiles workloads, identifies compute, memory, networking, and software bottlenecks, and develops C++ and Python tools. Coordinates complex performance improvements across hardware and software teams, validates results with reliable data, and guides engineering decisions that improve efficiency, scalability, and reliability at data center scale.
The summary above was generated by AI
About the Job

Lead performance work that turns large-scale AI workloads into faster, more efficient systems.

As a Staff AI Performance Engineer, you will analyze and optimize AI training and inference workloads across distributed systems. Your work will connect compute, memory, networking and software behavior.

You will lead complex performance workstreams across hardware and software. You will find bottlenecks, validate improvements and turn performance evidence into practical engineering decisions.

You will build benchmarks, profile demanding workloads and develop C++ and Python tools for analysis and optimization. Your recommendations will help teams improve efficiency, scalability and reliability.

This role gives you broad technical scope across AI infrastructure at data center scale. It is based in Milpitas, California.

The team and culture

The System Engineering Performance team architects, evaluates and optimizes high-performance infrastructure for large-scale data center deployments. The team works across the computing stack to understand real system behavior.

Work moves through evidence, ownership and clear technical judgment. You will lead investigations, coordinate across teams and validate impact with reliable performance data.

Decisions are shaped through benchmarks, models, simulations and practical engineering tradeoffs. The team values engineers who think big, act fast, take responsibility, speak up and lead beyond their own area.

What we’re looking for

· Strong experience profiling and optimizing AI, machine learning or high-performance computing workloads.
· Experience with distributed systems and communication libraries such as MPI, NCCL, UCX or libfabric.
· Strong C++ and Python skills, including reliable tools or performance-sensitive software.
· Deep understanding of compute, memory and communication behavior in large-scale systems.
· Ability to lead technically complex work and coordinate improvements across teams.
· Familiarity with MLPerf, accelerated architectures, ML frameworks or high-performance interconnects.

While we have outlined a set of requirements, we value transferable skills and diverse experiences. We also welcome engineers returning to the profession after a career break, including through returnship routes.

Benefits

Graphcore offers competitive compensation and benefits designed to support employees' health, financial well-being, work-life needs, and professional growth. Benefits and programs for eligible U.S. employees may include: 

• Medical, dental, and vision coverage, with options that may extend to eligible dependents.
• Mental health, wellness, and employee assistance resources.
• Retirement savings benefits and company contributions where applicable.
• Paid vacation, sick time, company holidays, and parental or family leave in accordance with applicable plans and policies.
• Life insurance and short-term or long-term disability coverage.
• Flexible working hours and hybrid working arrangements where compatible with the role and team requirements.
• Professional-development resources, learning programs, office amenities, and team-led activities. 

Benefits vary by work location, employment status, scheduled hours, and plan eligibility and are subject to the terms of the applicable plans and company policies. This overview is not a contract or guarantee of benefits. 

Graphcore is an equal opportunity employer. We consider qualified applicants without regard to race, color, religion, creed, sex, pregnancy, sexual orientation, gender identity or expression, national origin, ancestry, age, disability, genetic information, veteran status, or any other status protected by applicable law. 

Graphcore is committed to an inclusive and accessible hiring process. If you need a reasonable accommodation to participate in the application or interview process, please let the recruiting team know. 

Personal information submitted during the recruiting process will be handled in accordance with Graphcore's applicable candidate privacy notices.

Join the Team at Graphcore

Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry.

As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone.

Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore brings together deep expertise to solve complex problems and deliver meaningful progress in AI compute.

Ready to lead performance improvements across AI systems at data center scale? Apply now to join Graphcore in Milpitas.

Similar Jobs at Graphcore

Yesterday
Hybrid
Expert/Leader
Expert/Leader
Artificial Intelligence • Semiconductor
Leads the operating model and new-product-introduction framework for AI data center technologies. Defines deployment, service, diagnostics, telemetry, readiness, and lifecycle workflows across compute, networking, storage, power, cooling, and site operations. Coordinates multidisciplinary teams, conducts operational readiness reviews, establishes quality metrics, resolves complex platform issues, guides automation, and communicates risks and recommendations to engineering and operations leaders.
Top Skills: Ai ComputingAutomated DiagnosticsAutomationCompute InfrastructureData Center InfrastructureDirect Liquid CoolingFirmwareHigh-Performance ComputingLinuxMemoryNetworkingOperating SystemsPower DistributionRack-As-System ArchitecturesScriptingStorageTelemetry Platforms
2 Days Ago
Hybrid
Senior level
Senior level
Artificial Intelligence • Semiconductor
Defines end-to-end storage architecture and technology roadmaps for AI servers and data centers. Leads design of local, disaggregated, and distributed storage; optimizes Linux storage performance, NVMe lifecycle management, telemetry, and AI workloads. Diagnoses complex hardware and software issues across kernels, PCIe, networks, firmware, and storage platforms. Provides technical leadership across engineering, automation, supply chain, vendors, and product teams while evaluating emerging interconnect and memory-tiering technologies.
Top Skills: BashBlktraceCompute Express Link (Cxl)DaosDdnExt4FioGpu Direct StorageIostatLinuxLustreNvmeNvme Over FabricsPcie Gen5Pcie Gen6PerfPythonRest ApisRocev2Storage Performance Development Kit (Spdk)TcpVast DataWekaXfsZfs
2 Days Ago
Hybrid
Expert/Leader
Expert/Leader
Artificial Intelligence • Semiconductor
Lead end-to-end hardware and firmware security architecture for Arm-based AI server platforms. Define trusted computing, secure boot, attestation, cryptographic key management, lifecycle controls, interface security, and threat mitigations across silicon, firmware, virtualization, and system software. Translate security goals into engineering requirements, guide cross-functional teams, evaluate emerging threats, and protect AI models, data, workloads, and intellectual property from remote, physical, side-channel, and supply-chain attacks.
Top Skills: ArmArm Trusted FirmwareArmv9Confidential Compute ArchitectureContainersCryptographyKvmLinuxNvmeOpenbmcPcie Gen5Pcie Gen6Realm Management ExtensionRemote AttestationSecure BootTrusted Platform ModuleTrustzoneUefi

What you need to know about the Los Angeles Tech Scene

Los Angeles is a global leader in entertainment, so it’s no surprise that many of the biggest players in streaming, digital media and game development call the city home. But the city boasts plenty of non-entertainment innovation as well, with tech companies spanning verticals like AI, fintech, e-commerce and biotech. With major universities like Caltech, UCLA, USC and the nearby UC Irvine, the city has a steady supply of top-flight tech and engineering talent — not counting the graduates flocking to Los Angeles from across the world to enjoy its beaches, culture and year-round temperate climate.

Key Facts About Los Angeles Tech

  • Number of Tech Workers: 375,800; 5.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Snap, Netflix, SpaceX, Disney, Google
  • Key Industries: Artificial intelligence, adtech, media, software, game development
  • Funding Landscape: $11.6 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Strong Ventures, Fifth Wall, Upfront Ventures, Mucker Capital, Kittyhawk Ventures
  • Research Centers and Universities: California Institute of Technology, UCLA, University of Southern California, UC Irvine, Pepperdine, California Institute for Immunology and Immunotherapy, Center for Quantum Science and Engineering

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account