Mirantis Logo

Mirantis

Technical Product Manager, Observability – remote in the US

Posted 7 Days Ago
Remote
Hiring Remotely in USA
Senior level
Remote
Hiring Remotely in USA
Senior level
Own the observability vision, roadmap, backlog, and priorities for k0rdent AI’s GPU infrastructure control plane. Define integrations across GPU compute, networking, storage, DPUs, schedulers, inference systems, and databases using OpenTelemetry and Prometheus ecosystems. Translate customer and partner needs into product requirements, guide engineering trade-offs, track emerging standards, and support product marketing, field teams, customers, analysts, and ecosystem partners.
The summary above was generated by AI
Company Description

About Mirantis

Mirantis, an IREN company, is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. By combining open source innovation with deep expertise in Kubernetes orchestration, Mirantis empowers platform engineering teams to deliver composable, production-ready developer platforms across any environment—on-premises, in the cloud, at the edge, or in sovereign data centers. As enterprises navigate the growing complexity of AI-driven workloads, Mirantis delivers the automation, GPU orchestration, and policy-driven control needed to manage infrastructure with confidence and agility. Committed to open standards and freedom from lock-in, Mirantis ensures that customers retain full control of their infrastructure strategy.  https://www.mirantis.com/

Job Description

Job Summary

Mirantis is looking for a Technical Product Manager to own observability for k0rdent AI, our control plane for GPU infrastructure and distributed AI workloads. In this role, you will define the observability strategy, roadmap, and feature priorities that determine how operators gain visibility into the health, performance, and resource utilization of GPU clusters running large-scale training and inference. You will shape how k0rdent AI handles everything from GPU-level metrics and distributed tracing across AI workloads, to multi-tenant log aggregation and intelligent alerting — powered by the OpenTelemetry ecosystem, and Prometheus-compatible metrics pipelines.

The ideal candidate brings strong technical fluency in observability tooling and the AI infrastructure stack. You will work directly with engineering to shape requirements, with marketing to define positioning, and with customers to help ensure their success.

Responsibilities

  • Own the vision, roadmap, and priorities for k0rdent AI observability across the full stack: GPU compute, east-west fabric (InfiniBand, RoCE), high-performance storage, DPU/SmartNIC telemetry, workload schedulers, inference serving, and data services

  • Translate requirements from NeoClouds, GPU clouds, telcos, sovereign clouds, and enterprise platform teams into clear product direction; partner with engineering to define requirements and evaluate trade-offs

  • Manage the observability backlog using feedback from production deployments and design partners to refine priorities

  • Track and shape our response to emerging observability standards and technologies, including OpenTelemetry (OTel), DCGM GPU metrics, InfiniBand/RoCE fabric counters, storage platform telemetry APIs, and AI workload profiling

  • Define integration strategies for vendor telemetry sources across the ecosystem — NVIDIA compute and BlueField DPUs, storage platforms (VAST, Weka, DDN), workload managers (SLURM), inference stacks, and vector and relational databases — into a unified, operator-facing observability plane

  • Partner with product marketing and field teams on positioning, technical briefs, and reference architectures; represent Mirantis with customers, analysts, and ecosystem partners

Qualifications

 

  • 5+ years in product management or a senior technical role owning an observability product or operating large-scale monitoring infrastructure

  • Working knowledge of Prometheus, OpenTelemetry, distributed tracing (Jaeger, Tempo), and log aggregation (Loki, Elasticsearch/OpenSearch)

  • Fluency in Kubernetes observability, cloud-native monitoring, or metrics and alerting pipeline architecture

  • Ability to work directly with engineering on technical trade-offs and with field teams in competitive GPU cloud and NeoCloud deals

Strongly Preferred:

  • Exposure to GPU observability, including DCGM metrics, AI workload profiling and performance analysis

  • Familiarity with east-west fabric telemetry - InfiniBand counters, RoCEv2 congestion metrics (ECN, PFC, DCQCN), or switch-level fabric health

  • Experience with high-performance storage telemetry from platforms such as VAST Data, Weka, or DDN, including IOPS, latency, and throughput instrumentation at scale

  • Familiarity with NVIDIA BlueField DPU telemetry, SR-IOV, or offload pipeline observability

  • Exposure to workload-level visibility for SLURM job scheduling, inference serving stacks (vLLM, Triton, TensorRT-LLM), or data service telemetry from vector databases (Milvus, Qdrant) and relational databases in AI pipelines

 

Additional Information

Why you’ll love Mirantis

  • Build the observability foundation for the AI cloud era, working directly with leading GPU cloud operators, NeoClouds, sovereign clouds, and AI-first enterprises

  • Collaborate with a world-class, distributed team committed to openness and technical excellence

  • Shape the product narrative and influence go-to-market success

 

It is understood that Mirantis, Inc. may use automated decision-making technology (ADMT) for specific employment-related decisions. Opting out of ADMT use is requested for decisions about evaluation and review connected with the specific employment decision for the position applied for. You also have the right to appeal any decisions made by ADMT by sending your request to [email protected]

By submitting your resume, you consent to the processing and storage of your personal data in accordance with applicable data protection laws, for the purposes of considering your application for current and future job opportunities.

#remote

We are a Leader for Container Management in G2 (#2 after AWS)!

Similar Jobs

Senior level
Artificial Intelligence • Big Data • Healthtech • Software • Biotech
Leads end-to-end IVD, companion diagnostic, clinical trial assay, tumor profiling, and SaMD development programs. Aligns internal cross-functional teams and biopharma partners across strategy, validation, clinical operations, regulatory submission, manufacturing, and commercialization. Owns governance, milestones, budgets, resources, risks, and partner relationships while delivering programs through regulatory approval or pivotal clinical trials.
Top Skills: Analytical ValidationAssay DevelopmentBioinformaticsBlaClinical ValidationCompanion Diagnostics (Cdx)Design ControlsEu IvdrFda Regulatory RequirementsIndIvdMaaNdaPhase I/Ii/IiiSoftware As A Medical Device (Samd)
44 Minutes Ago
Remote or Hybrid
USA
170K-260K Annually
Expert/Leader
170K-260K Annually
Expert/Leader
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Lead product marketing strategy and go-to-market execution for Falcon Exposure Management. Develop positioning, messaging, launches, thought leadership, web strategy, and field enablement while partnering with product, sales, marketing, and executive teams. Translate cybersecurity capabilities into customer value and business outcomes, guide market education, analyze buyers and competitors, and strengthen CrowdStrike’s category leadership. The role requires deep exposure management expertise, strong cybersecurity knowledge, exceptional communication skills, and 12–15+ years of relevant experience.
Top Skills: AIB2B SaasCloud SecurityCybersecurityEnterprise SoftwareExternal Attack Surface ManagementFalcon Exposure ManagementIdentity SecuritySecurity OperationsVulnerability Management
44 Minutes Ago
Remote or Hybrid
USA
120K-180K Annually
Senior level
120K-180K Annually
Senior level
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Lead program management for CrowdStrike’s Office of Adversary Strategy, coordinating complex cross-functional projects from scoping through delivery. Responsibilities include managing schedules, dependencies, risks, metrics, communications, change management, process improvements, and escalations. The role partners with technical teams and leadership, applies Agile project management practices, coaches teams through obstacles, and uses AI technologies to improve decision-making, workflows, and business outcomes.
Top Skills: AgileArtificial IntelligenceChatgptClaudeExcelGeminiGenerative AiJIRAKanbanScaled Agile FrameworkScrumSystems Development Lifecycle

What you need to know about the Los Angeles Tech Scene

Los Angeles is a global leader in entertainment, so it’s no surprise that many of the biggest players in streaming, digital media and game development call the city home. But the city boasts plenty of non-entertainment innovation as well, with tech companies spanning verticals like AI, fintech, e-commerce and biotech. With major universities like Caltech, UCLA, USC and the nearby UC Irvine, the city has a steady supply of top-flight tech and engineering talent — not counting the graduates flocking to Los Angeles from across the world to enjoy its beaches, culture and year-round temperate climate.

Key Facts About Los Angeles Tech

  • Number of Tech Workers: 375,800; 5.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Snap, Netflix, SpaceX, Disney, Google
  • Key Industries: Artificial intelligence, adtech, media, software, game development
  • Funding Landscape: $11.6 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Strong Ventures, Fifth Wall, Upfront Ventures, Mucker Capital, Kittyhawk Ventures
  • Research Centers and Universities: California Institute of Technology, UCLA, University of Southern California, UC Irvine, Pepperdine, California Institute for Immunology and Immunotherapy, Center for Quantum Science and Engineering

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account