Mirantis Logo

Mirantis

Senior HPC Networking Engineer

Posted 2 Days Ago
Be an Early Applicant
Remote
Hiring Remotely in USA
Senior level
Remote
Hiring Remotely in USA
Senior level
Design, deploy, and maintain high-performance HPC network fabrics (InfiniBand/RoCE), troubleshoot complex latency/throughput issues, manage Fortinet security solutions, tune and monitor network performance, collaborate with compute/storage teams, document architectures, participate in on-call rotations, and support large-scale GPU/AI clusters and incident triage.
The summary above was generated by AI
Company Description

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. By combining open source innovation with deep expertise in Kubernetes orchestration, Mirantis empowers platform engineering teams to deliver composable, production-ready developer platforms across any environment—on-premises, in the cloud, at the edge, or in sovereign data centers. As enterprises navigate the growing complexity of AI-driven workloads, Mirantis delivers the automation, GPU orchestration, and policy-driven control needed to manage infrastructure with confidence and agility. Committed to open standards and freedom from lock-in, Mirantis ensures that customers retain full control of their infrastructure strategy.  https://www.mirantis.com/

Job Description

Location: US
Employment Type: Full-time

Role Overview:
We are seeking a highly skilled Senior HPC Networking Engineer to design, deploy, manage, and troubleshoot high-performance networking environments. The ideal candidate will have deep expertise in InfiniBand technologies, strong general networking knowledge, and hands-on experience with Fortinet solutions. You will play a critical role in ensuring the performance, reliability, and scalability of HPC infrastructure. 

Key Responsibilities:

  • Design, deploy, and maintain high-performance network infrastructures for HPC environments, with a strong focus on InfiniBand fabrics.

  • Troubleshoot complex network issues across InfiniBand and Ethernet environments, ensuring minimal downtime and optimal performance.

  • Manage and optimize InfiniBand components, including switches, HCAs, subnet managers, and fabric configurations.

  • Perform performance tuning, monitoring, and capacity planning for HPC networking systems.

  • Implement and maintain network security using Fortinet solutions (FortiGate, FortiManager, FortiAnalyzer).

  • Diagnose and resolve issues related to routing, switching, latency, and throughput across hybrid network environments.

  • Collaborate with compute, storage, and platform teams to support HPC workloads and cluster operations.

  • Develop and maintain documentation for network architecture, configurations, and operational procedures.

  • Participate in on-call rotations and provide escalation support for critical incidents.

  • Lead or contribute to network upgrades, migrations, and new deployments.

  • You will actively troubleshoot and resolve daily customer incident tickets to keep massive GPU training runs moving.

Build, operate, and scale next-generation GPU infrastructure:

  • You will be hands-on on the front lines driving daily triage, incident resolution, and SLA management for the world’s most advanced NVIDIA clusters, InfiniBand/RoCE fabrics, and AI workloads.

Build the playbook, then grow into the platform:

  • Designed for engineers energized by standing up new operations from the ground up, this role offers a direct trajectory from operationalizing bare-metal clusters to driving platform and AI capabilities as we scale.

Qualifications

Required:

  • 5+ years of experience in network engineering, with a focus on HPC or data center environments.

  • Strong hands-on experience with InfiniBand technologies (e.g., Mellanox/NVIDIA).

  • Solid understanding of networking fundamentals: TCP/IP, routing protocols (BGP, OSPF), VLANs, QoS, and network design.

  • Proven experience deploying and troubleshooting Fortinet solutions (FortiGate, FortiManager, VPNs, firewall policies).

  • Experience with network performance analysis and troubleshooting tools.

  • Familiarity with Linux systems and scripting for automation (e.g., Bash, Python).

  • Strong analytical and problem-solving skills.

Preferred:

  • Experience with large-scale HPC clusters or AI/ML infrastructure.

  • Knowledge of RDMA, MPI, and low-latency networking concepts.

  • Certifications such as FCSS/FCNSP (Fortinet), CCNP/CCIE, or equivalent.

  • Experience with automation and Infrastructure as Code tools (e.g., Ansible, Terraform).

Soft Skills:

  • Strong communication and collaboration skills.

  • Ability to work independently and handle complex technical challenges.

  • Detail-oriented with a proactive approach to problem-solving.

What We Offer:

  • Opportunity to work on cutting-edge HPC infrastructure.

  • Collaborative and innovative work environment.

  • Competitive salary and benefits package.

 

Additional Information

What does Mirantis offer you?

  • Work with an established Silicon Valley leader in the cloud infrastructure industry;
  • Work with exceptionally passionate, talented and engaging colleagues, helping Fortune 500 and Global 2000 customers implement next-generation cloud technologies;
  • Be a part of cutting-edge, open-source innovation;
  • Thrive in the high-energy environment of a young company where openness, collaboration, risk-taking, and continuous growth are valued;
  • Professional development and training;
  • Attend conferences and working groups;
  • Company outings, happy hours, hackathons, and tech talks;
  • Receive a competitive compensation package with a strong benefits plan.

We are a Leader for Container Management in G2 (#2 after AWS)!

Similar Jobs

20 Minutes Ago
Remote or Hybrid
United States
98K-165K Annually
Senior level
98K-165K Annually
Senior level
Artificial Intelligence • Cloud • Sales • Security • Software • Cybersecurity • Data Privacy
Lead end-to-end execution of a 12+ initiative product security program, managing cross-team dependencies, vendor procurement, BAU security operations, and executive reporting. Maintain roadmap and tracking in Jira/Confluence, drive procurement/POCs, produce executive-ready metrics, and establish program cadence to ensure timely remediation and predictable delivery.
Top Skills: Ai And Machine LearningCi/CdCloud-Native SecurityConfluenceDevsecopsIdentity And Access Management (Iam)JIRAScmVulnerability Management
21 Minutes Ago
Remote
USA
200K-250K Annually
Senior level
200K-250K Annually
Senior level
Artificial Intelligence • Machine Learning • Natural Language Processing • Software • Conversational AI
Build and scale Deepgram's federal partner network by owning strategy and execution for SI/prime and hyperscaler alliances. Drive sell-to/sell-through/sell-with motions, create joint solutions, enable partners at scale, instrument partner-sourced/influenced pipeline, and deploy AI-native automations to operate at scale. Represent Deepgram externally and coordinate cross-functionally to accelerate partner-driven revenue.
Top Skills: Agentic AiAgentsAws Connect GovcloudCio-Sp3CRMDeepgramDod Impact LevelFedrampGenesys GovernmentGovcloudGsa MasGwacsLlmsNasa SewpNiceRetrieval Augmented Generation (Rag)Sam.GovServicenow FederalSlackSpeech-To-Text (Stt)Text-To-Speech (Tts)Usaspending
24 Minutes Ago
Remote
United States
182K-203K Annually
Expert/Leader
182K-203K Annually
Expert/Leader
Healthtech • Social Impact • Software • Telehealth
Lead and scale organic growth across SEO, AEO, and CRO. Own multi-year strategy, experimentation, site architecture, and measurement. Partner with Engineering, Product, and Design to improve discoverability and conversions, ensure HIPAA/WCAG compliance, manage P&L, recruit and mentor teams, and drive revenue through AI-native discovery and conversion optimization.
Top Skills: AeoAIAPIsAutomationCmsContent Delivery Networks (Cdn)Conversion Rate Optimization (Cro)Core Web VitalsCro ToolingEdge RenderingHipaaJavaScriptLlmsProgrammatic ContentSchema MarkupSeoSeo ToolingServer-Side RenderingStructured DataWcag

What you need to know about the Los Angeles Tech Scene

Los Angeles is a global leader in entertainment, so it’s no surprise that many of the biggest players in streaming, digital media and game development call the city home. But the city boasts plenty of non-entertainment innovation as well, with tech companies spanning verticals like AI, fintech, e-commerce and biotech. With major universities like Caltech, UCLA, USC and the nearby UC Irvine, the city has a steady supply of top-flight tech and engineering talent — not counting the graduates flocking to Los Angeles from across the world to enjoy its beaches, culture and year-round temperate climate.

Key Facts About Los Angeles Tech

  • Number of Tech Workers: 375,800; 5.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Snap, Netflix, SpaceX, Disney, Google
  • Key Industries: Artificial intelligence, adtech, media, software, game development
  • Funding Landscape: $11.6 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Strong Ventures, Fifth Wall, Upfront Ventures, Mucker Capital, Kittyhawk Ventures
  • Research Centers and Universities: California Institute of Technology, UCLA, University of Southern California, UC Irvine, Pepperdine, California Institute for Immunology and Immunotherapy, Center for Quantum Science and Engineering

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account