Graphcore Logo

Graphcore

Staff Engineering Operations Technical Program Manager

Posted 2 Hours Ago
Be an Early Applicant
Hybrid
Austin, TX
Senior level
Hybrid
Austin, TX
Senior level
Provides advanced operational and engineering support for Arm-based AI hardware platforms in lab and data center environments. Leads hardware bring-up, break-fix troubleshooting, validation, root cause analysis, and reliability improvements across server blades, motherboards, power systems, and rack-scale infrastructure. Develops troubleshooting documentation, collaborates with engineering and infrastructure teams, mentors junior staff, and supports on-call activities during critical hardware milestones.
The summary above was generated by AI
About us

Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry.

As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone.

Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation.


Job Summary

We are seeking a Staff Hardware Engineer to provide advanced operational, diagnostic, and engineering support for Graphcore’s Arm-based hardware platforms across lab and data center environments.

This role focuses on supporting hardware bring-up, validation, and troubleshooting of complex AI compute platforms, including server blades, racks, and rack-scale infrastructure. The successful candidate will collaborate closely with engineering, platform, and data center teams to ensure the reliability and performance of next-generation AI systems.


The Team

The Systems Engineering and Hardware Engineering teams are responsible for enabling the bring-up, validation, and operational reliability of Graphcore’s AI infrastructure platforms.

The team works closely with server engineering, firmware teams, platform architects, and data center operations to support the development, testing, and deployment of next-generation AI compute systems.

This collaborative environment enables rapid problem-solving and continuous improvement of Graphcore’s hardware platforms from early development through production deployment.


Responsibilities and Duties
  • Lead advanced break-fix troubleshooting for server blades, motherboards, power systems, and rack-scale infrastructure.
  • Support engineering bring-up activities, including component validation and firmware interaction testing.
  • Diagnose system-level failures involving thermal behavior, power anomalies, network configuration, and BIOS/BMC issues.
  • Collaborate with server engineering teams to perform root cause analysis and propose corrective actions or design improvements.
  • Support deployment and rollout of next-generation hardware platforms through structured validation and qualification cycles.
  • Interface with facilities and infrastructure teams to understand environmental factors impacting system reliability.
  • Develop and maintain standard operating procedures (SOPs), troubleshooting guides, and validation documentation.
  • Provide guidance and mentorship to junior technicians and engineers on troubleshooting methodologies and hardware diagnostics.
  • Participate in on-call rotations or off-hours support during critical engineering milestones or hardware bring-up phases.
Candidate ProfileEssential
  • Bachelor’s degree in Electrical Engineering, Computer Engineering, Computer Science, or related discipline.
  • 8 years of strong experience with server hardware architectures and board-level debugging.
  • Experience analyzing system logs, hardware telemetry, and power/thermal metrics to isolate hardware failures.
  • Hands-on experience with HPC systems, AI compute platforms, or rack-scale infrastructure.
  • Strong collaboration skills and ability to work effectively in fast-paced engineering environments.
  • Excellent written and verbal communication skills.
Desirable
  • Experience supporting prototype or pre-production hardware bring-up.
  • Familiarity with data center facilities, including liquid cooling and power distribution systems.
  • Experience using Python, Bash, or automation tools for hardware validation or troubleshooting.
  • Exposure to structured failure analysis and reliability engineering methodologies.

USA Benefits
In addition to a competitive salary, Graphcore offers flexible working and a comprehensive benefits package designed to support your health, wellbeing and financial future. Our benefits include medical, dental and vision coverage, Flexible Spending Accounts (FSAs), Health Savings Accounts (HSAs), disability and life insurance, a 401(k) retirement plan, commuter benefits, wellness services and an Employee Assistance Programme (EAP). We welcome people of different backgrounds and experiences; we're committed to building an inclusive work environment that makes Graphcore a great home for everyone. We offer an equal opportunity process and understand that there are visible and invisible differences in all of us. We can provide a flexible approach to interview and encourage you to chat to us if you require any reasonable adjustments.

Similar Jobs at Graphcore

Yesterday
Hybrid
Senior level
Senior level
Artificial Intelligence • Semiconductor
Provides advanced operational, diagnostic, and engineering support for Arm-based AI hardware platforms. Leads server, motherboard, power, thermal, networking, firmware, BIOS, and BMC troubleshooting; supports hardware bring-up, validation, qualification, deployment, and root-cause analysis. Develops troubleshooting documentation, collaborates with engineering and data center teams, mentors junior staff, and participates in on-call support during critical hardware milestones.
Top Skills: Ai Compute PlatformsArmBashBiosBmcFirmwareHpc SystemsLiquid CoolingMotherboard DebuggingPower Distribution SystemsPythonRack-Scale InfrastructureServer Hardware
Yesterday
Hybrid
Expert/Leader
Expert/Leader
Artificial Intelligence • Semiconductor
Develop and extend low-level diagnostic and stress-testing software for next-generation AI SoCs. Build configurable validation tools covering CPUs, accelerators, memory, PCIe, storage, firmware, BMC, and operating-system interactions. Investigate intermittent hardware failures, silent data corruption, performance issues, and reliability problems across server and rack-scale platforms. Collaborate with hardware, firmware, software, and validation teams to improve observability, fault isolation, silicon bring-up, and platform quality.
Top Skills: Ai AcceleratorsArmBmcCC++DdrEmbedded SystemsHbmLinuxNumaPciePythonRasServer PlatformsSoc Validation
Senior level
Artificial Intelligence • Semiconductor
Develop system-level diagnostics and stress-testing software for next-generation AI SoCs. Build reusable validation utilities, automate testing across server and rack-scale environments, expose hardware failures and silent data corruption, and improve fault isolation and observability. Collaborate with hardware, firmware, software, and validation teams while debugging interactions across CPUs, memory, PCIe, storage, networking, operating systems, and drivers.
Top Skills: Ai AcceleratorsArmBmcCC++DdrHbmLinuxNumaPciePythonRasSocs

What you need to know about the Los Angeles Tech Scene

Los Angeles is a global leader in entertainment, so it’s no surprise that many of the biggest players in streaming, digital media and game development call the city home. But the city boasts plenty of non-entertainment innovation as well, with tech companies spanning verticals like AI, fintech, e-commerce and biotech. With major universities like Caltech, UCLA, USC and the nearby UC Irvine, the city has a steady supply of top-flight tech and engineering talent — not counting the graduates flocking to Los Angeles from across the world to enjoy its beaches, culture and year-round temperate climate.

Key Facts About Los Angeles Tech

  • Number of Tech Workers: 375,800; 5.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Snap, Netflix, SpaceX, Disney, Google
  • Key Industries: Artificial intelligence, adtech, media, software, game development
  • Funding Landscape: $11.6 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Strong Ventures, Fifth Wall, Upfront Ventures, Mucker Capital, Kittyhawk Ventures
  • Research Centers and Universities: California Institute of Technology, UCLA, University of Southern California, UC Irvine, Pepperdine, California Institute for Immunology and Immunotherapy, Center for Quantum Science and Engineering

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account