Graphcore Logo

Graphcore

Principal Datacenter Technologist

Posted Yesterday
Be an Early Applicant
Hybrid
Milpitas, CA
Expert/Leader
Hybrid
Milpitas, CA
Expert/Leader
Leads the operating model and new-product-introduction framework for AI data center technologies. Defines deployment, service, diagnostics, telemetry, readiness, and lifecycle workflows across compute, networking, storage, power, cooling, and site operations. Coordinates multidisciplinary teams, conducts operational readiness reviews, establishes quality metrics, resolves complex platform issues, guides automation, and communicates risks and recommendations to engineering and operations leaders.
The summary above was generated by AI
About Graphcore

Graphcore is a leading innovator in artificial intelligence computing. We develop hardware, software, and data center infrastructure that provide the specialized processing and systems capabilities needed to advance AI while improving the efficiency required for broad adoption.

As part of SoftBank Group, Graphcore works alongside companies developing advanced technologies. Our teams bring together AI researchers, silicon designers, hardware and software engineers, and systems architects to solve complex technical problems across the computing stack.

The Opportunity

As Principal Datacenter Technologist, you will turn new platform architectures into repeatable operating models for Graphcore's AI data centers. You will lead the technical definition of how Graphcore introduces, deploys, services, and sustains new compute, networking, storage, memory, power, and cooling technologies at scale.

Working within Advanced Architecture, you will connect platform design with data center engineering and site reliability operations. You will define new-product-introduction workflows, readiness criteria, diagnostics, and cross-functional handoffs so that new systems can move from architecture through deployment with clear ownership, supportability, and operational feedback.

What You Will Do
  • Own the technical operating model and new-product-introduction framework for new hardware and infrastructure technologies entering Graphcore data centers.
  • Define end-to-end workflows for installation, configuration, validation, deployment, service, repair, upgrade, and sustained operation across the product lifecycle.
  • Translate architecture requirements into data center readiness criteria covering software, firmware, compute, networking, storage, memory, rack integration, power, cooling, space, and site operations.
  • Create clear runbooks, interface definitions, ownership models, acceptance criteria, and escalation paths that align architecture, engineering, data center, and site reliability teams.
  • Lead operational readiness reviews and identify gaps in tooling, diagnostics, telemetry, serviceability, spares, documentation, and training before deployment.
  • Serve as a senior escalation point for complex platform issues, ensuring that teams collect the right logs, telemetry, failure evidence, and environmental data to support root-cause analysis and disposition.
  • Define and guide the deployment of automated diagnostic and telemetry systems that identify systemic hardware, software, quality, handling, and site-integration issues across the fleet.
  • Establish quality and operational metrics for new technologies, analyze field trends, and feed findings back into architecture, design, supplier, and deployment decisions.
  • Coordinate cross-functional resolution of issues spanning hardware, firmware, software, networking, storage, facilities, and site operations, with clear owners, decisions, and follow-through.
  • Provide technical leadership, review implementation plans, mentor engineers, and communicate readiness, risks, tradeoffs, and recommendations to engineering and operations leaders.
What You Will Bring
  • A bachelor's degree or equivalent experience in information technology, computer science, computer engineering, electrical engineering, or a related field, or equivalent practical experience.
  • Extensive experience in data center infrastructure build, infrastructure operations, platform deployment, or related technical fields, including leadership at Principal, Staff, or Lead Engineer scope.  
  • Demonstrated experience defining processes, workflows, and readiness criteria for new-product introduction in large-scale data center or infrastructure environments.
  • Broad systems knowledge across compute, networking, storage, memory, firmware, operating systems, racks, power distribution, cooling, and data center operations.
  • Hands-on Linux experience, including scripting or automation used to collect diagnostics, telemetry, configuration, and health information.
  • Experience with failure analysis and validation of AI compute, server, networking, or other complex data center equipment.
  • Ability to define telemetry, diagnostic coverage, quality metrics, and fleet-level signals that distinguish product issues from deployment, handling, or environmental problems.
  • Proven ability to lead multidisciplinary technical work across architecture, hardware, software, facilities, data center engineering, and site reliability teams without relying on direct authority.
  • Clear written and verbal communication skills, including the ability to produce precise technical workflows and explain operational risks and decisions to engineering leaders.
Preferred Qualifications

These qualifications are helpful, not required. We encourage you to apply even if you do not meet every preferred qualification.

  • Experience leading data center engineering, platform introduction, or infrastructure readiness initiatives.
  • Experience operating or supporting high-density AI or high-performance computing systems at fleet scale.
  • Experience with telemetry platforms, automated diagnostics, serviceability tooling, and quality metrics for compute and networking equipment.
  • Experience with rack-as-a-system architectures, direct liquid cooling, high-power racks, and the operational controls required for safe deployment and service.
  • Experience working with suppliers, field-service teams, manufacturing, or logistics partners on equipment lifecycle, quality, repair, or failure-disposition processes.
  • United States Benefits Overview

    Graphcore offers compensation and benefits designed to support employees' health, financial well-being, work-life needs, and professional growth. Benefits and programs for eligible U.S. employees may include:

    • Medical, dental, and vision coverage, with options that may extend to eligible dependents.
    • Mental health, wellness, and employee assistance resources.
    • Retirement savings benefits and company contributions where applicable.
    • Paid vacation, sick time, company holidays, and parental or family leave in accordance with applicable plans and policies.
    • Life insurance and short-term or long-term disability coverage.
    • Flexible working hours and hybrid working arrangements where compatible with the role and team requirements.
    • Professional development resources, learning programs, office amenities, and team-led activities.

    Benefits vary by work location, employment status, scheduled hours, and plan eligibility and are subject to the terms of the applicable plans and company policies. This overview is not a contract or guarantee of benefits.

    Equal Opportunity and Accommodations

    Graphcore is an equal opportunity employer. We consider qualified applicants without regard to race, color, religion, creed, sex, pregnancy, sexual orientation, gender identity or expression, national origin, ancestry, age, disability, genetic information, veteran status, or any other status protected by applicable law.

    Graphcore is committed to an inclusive and accessible hiring process. If you need a reasonable accommodation to participate in the application or interview process, please let the recruiting team know.

    Candidate Privacy

    Personal information submitted during the recruiting process will be handled in accordance with Graphcore's applicable candidate privacy notices.

Similar Jobs at Graphcore

An Hour Ago
Hybrid
Expert/Leader
Expert/Leader
Artificial Intelligence • Semiconductor
Leads performance optimization for large-scale AI training and inference workloads across distributed systems. Builds benchmarks, profiles workloads, identifies compute, memory, networking, and software bottlenecks, and develops C++ and Python tools. Coordinates complex performance improvements across hardware and software teams, validates results with reliable data, and guides engineering decisions that improve efficiency, scalability, and reliability at data center scale.
Top Skills: C++Distributed SystemsHigh-Performance InterconnectsLibfabricMachine Learning FrameworksMlperfMpiNcclPythonUcx
2 Days Ago
Hybrid
Senior level
Senior level
Artificial Intelligence • Semiconductor
Defines end-to-end storage architecture and technology roadmaps for AI servers and data centers. Leads design of local, disaggregated, and distributed storage; optimizes Linux storage performance, NVMe lifecycle management, telemetry, and AI workloads. Diagnoses complex hardware and software issues across kernels, PCIe, networks, firmware, and storage platforms. Provides technical leadership across engineering, automation, supply chain, vendors, and product teams while evaluating emerging interconnect and memory-tiering technologies.
Top Skills: BashBlktraceCompute Express Link (Cxl)DaosDdnExt4FioGpu Direct StorageIostatLinuxLustreNvmeNvme Over FabricsPcie Gen5Pcie Gen6PerfPythonRest ApisRocev2Storage Performance Development Kit (Spdk)TcpVast DataWekaXfsZfs
2 Days Ago
Hybrid
Expert/Leader
Expert/Leader
Artificial Intelligence • Semiconductor
Lead end-to-end hardware and firmware security architecture for Arm-based AI server platforms. Define trusted computing, secure boot, attestation, cryptographic key management, lifecycle controls, interface security, and threat mitigations across silicon, firmware, virtualization, and system software. Translate security goals into engineering requirements, guide cross-functional teams, evaluate emerging threats, and protect AI models, data, workloads, and intellectual property from remote, physical, side-channel, and supply-chain attacks.
Top Skills: ArmArm Trusted FirmwareArmv9Confidential Compute ArchitectureContainersCryptographyKvmLinuxNvmeOpenbmcPcie Gen5Pcie Gen6Realm Management ExtensionRemote AttestationSecure BootTrusted Platform ModuleTrustzoneUefi

What you need to know about the Los Angeles Tech Scene

Los Angeles is a global leader in entertainment, so it’s no surprise that many of the biggest players in streaming, digital media and game development call the city home. But the city boasts plenty of non-entertainment innovation as well, with tech companies spanning verticals like AI, fintech, e-commerce and biotech. With major universities like Caltech, UCLA, USC and the nearby UC Irvine, the city has a steady supply of top-flight tech and engineering talent — not counting the graduates flocking to Los Angeles from across the world to enjoy its beaches, culture and year-round temperate climate.

Key Facts About Los Angeles Tech

  • Number of Tech Workers: 375,800; 5.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Snap, Netflix, SpaceX, Disney, Google
  • Key Industries: Artificial intelligence, adtech, media, software, game development
  • Funding Landscape: $11.6 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Strong Ventures, Fifth Wall, Upfront Ventures, Mucker Capital, Kittyhawk Ventures
  • Research Centers and Universities: California Institute of Technology, UCLA, University of Southern California, UC Irvine, Pepperdine, California Institute for Immunology and Immunotherapy, Center for Quantum Science and Engineering

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account