Hyperbolic Logo

Hyperbolic

Forward Deployed Infrastructure Engineer - Eastern US

Reposted 7 Days Ago
Remote
Hiring Remotely in USA
Mid level
Remote
Hiring Remotely in USA
Mid level
The role involves benchmarking infrastructure performance, designing tests, debugging customer trials, and maintaining benchmarking infrastructure while ensuring clear documentation and communication.
The summary above was generated by AI
Who We Are

Hyperbolic Labs is on a mission to democratize AI by breaking down the barriers to computing power with our Open-Access AI Cloud. By making better use of idle computing resources across the globe, we offer an innovative GPU marketplace and AI inference service that promise affordability and accessibility for all. As pioneers at the intersection of AI and open-source technology, we believe in an open future where AI innovation is limited only by imagination, not by access to resources. We're looking for forward-thinking individuals who share our passion for making AI universally accessible, secure, and affordable. Join us in building a platform that empowers innovators everywhere to turn their visionary AI projects into reality.

NOTE: This role is focused on EASTERN TIME ZONE

About the Role

Our reserved customers run large multinode GPU clusters, and when those clusters misbehave the problem is rarely simple. You are the engineer embedded with those customers: you stand their cluster up, you hand it over, and you stay with it.

You do not own tickets. You own environments. Technical Support Engineers own the ticket lifecycle and pull you in when an issue needs real depth: multinode collective performance, hardware faults, fabric problems, or a provider who needs to be told what is wrong with their hardware.

One thing we will be straight about, because it shapes the job. We aggregate capacity from suppliers rather than owning most of the hardware ourselves. That means a real part of this role is technical liaison work: proving where a fault actually lives, taking it to the provider with evidence, and coordinating the fix on the customer's behalf. The engineers who enjoy this role are the ones who find that interesting rather than frustrating.

Who You Are

  • Cluster stand-up and handoff. Build, validate, and benchmark new customer clusters, then hand them over with documentation the customer's own engineers can work from.

  • Deep escalations. Multinode and NCCL performance debugging, GPU and hardware faults (XID and ECC errors, lspci, dmesg), driver and fabric issues, container and scheduler problems.

  • Provider escalation and coordination. Maintenance windows, RMAs, hung nodes, and disputed fault attribution. You bring the evidence that makes the provider act, and you keep the customer informed while it happens.

  • Embedded ownership of named accounts. You are the engineer your customers know by name. You learn their workload, not just their infrastructure, and you tell them what to change before they hit the wall.

  • Proactive monitoring. Own the monitoring and alerting we put in front of customer clusters (Grafana, Prometheus) so we find faults before the customer reports them.

  • Tooling and pushing work down. Automate the repeat work and turn your own escalations into runbooks the L1 tier can run. Anything you fix three times should stop reaching you.

  • Deep Linux experience and total comfort in the CLI, including in someone else's broken environment.

  • Hands-on multinode GPU experience: NCCL, InfiniBand or RoCE, collective performance debugging, topology and placement.

  • Hardware fault triage on GPU nodes: XID and ECC errors, dmesg, lspci, nvidia-smi, thermal and power faults.

  • Experience provisioning and operating GPU clusters with Kubernetes, Slurm, or both.

  • Genuinely comfortable customer-facing, including delivering bad news and saying “this is ours” or “this is the provider’s” with confidence.

  • Sound judgment on when to keep digging and when to escalate.

Preferred Qualifications

  • Experience with parallel filesystems (Weka, Lustre, GPFS) and high-performance storage.

  • Grafana, Prometheus, or similar observability stacks in production.

  • Prior forward deployed engineer, solutions architect, or technical account manager experience.

  • Infrastructure as code (Terraform, Ansible) and CI for cluster provisioning.

  • Exposure to model training or inference workloads from the practitioner side.

Hyperbolic is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.

Similar Jobs

An Hour Ago
In-Office or Remote
IN, USA
171K-171K Annually
Entry level
171K-171K Annually
Entry level
Artificial Intelligence • Consumer Web • Digital Media • Information Technology • Social Impact • Software
Lead Circle’s creative direction and narrative across product launches, campaigns, advertising, landing pages, partnerships, video, events, and brand experiences. Partner closely with founders, design, product marketing, paid media, video, and events teams to create distinctive, high-performing work. Develop creative concepts, briefs, scripts, messaging, and campaign direction while raising creative standards through influence, collaboration, and deep understanding of the product, customers, and market.
An Hour Ago
In-Office or Remote
IN, USA
145K-160K Annually
Entry level
145K-160K Annually
Entry level
Artificial Intelligence • Consumer Web • Digital Media • Information Technology • Social Impact • Software
Build Circle’s social and media function by executing platform-native content strategies, producing and editing short-form video, managing content calendars and capture workflows, and partnering with leadership and cross-functional teams. Attend events to capture content, support long-form production, monitor social performance indicators, and recommend improvements to formats and workflows. The role requires strong platform fluency, independent video production, creator-economy awareness, operational judgment, and excellent English communication.
Top Skills: Adobe PremiereInstagramLinkedInTiktokYoutube
An Hour Ago
In-Office or Remote
IN, USA
140K-170K Annually
Senior level
140K-170K Annually
Senior level
Artificial Intelligence • Consumer Web • Digital Media • Information Technology • Social Impact • Software
Own Discover’s end-to-end product design for Circle’s two-sided marketplace. Set design direction, identify opportunities, design AI-driven discovery experiences, and create trust-building experiences for creators and consumers. Prototype rapidly with AI-assisted tools, collaborate closely with engineers through launch, understand marketplace dynamics, and mentor other designers. The role requires strong product ownership, systems thinking, comfort with ambiguity, and experience designing consumer marketplace products at scale.
Top Skills: AIClaude CodeCursorFigmaLovableV0

What you need to know about the Los Angeles Tech Scene

Los Angeles is a global leader in entertainment, so it’s no surprise that many of the biggest players in streaming, digital media and game development call the city home. But the city boasts plenty of non-entertainment innovation as well, with tech companies spanning verticals like AI, fintech, e-commerce and biotech. With major universities like Caltech, UCLA, USC and the nearby UC Irvine, the city has a steady supply of top-flight tech and engineering talent — not counting the graduates flocking to Los Angeles from across the world to enjoy its beaches, culture and year-round temperate climate.

Key Facts About Los Angeles Tech

  • Number of Tech Workers: 375,800; 5.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Snap, Netflix, SpaceX, Disney, Google
  • Key Industries: Artificial intelligence, adtech, media, software, game development
  • Funding Landscape: $11.6 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Strong Ventures, Fifth Wall, Upfront Ventures, Mucker Capital, Kittyhawk Ventures
  • Research Centers and Universities: California Institute of Technology, UCLA, University of Southern California, UC Irvine, Pepperdine, California Institute for Immunology and Immunotherapy, Center for Quantum Science and Engineering

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account