Mirantis Logo

Mirantis

Senior Software Systems Engineer (Storage) - remote in the US

Posted Yesterday
Remote
Hiring Remotely in USA
Senior level
Remote
Hiring Remotely in USA
Senior level
Deploy and operate high-performance NFS storage for Kubernetes, bare-metal, GPU, and AI workloads. Integrate storage through CSI, tune Linux, NFS, networking, and kernel parameters, and manage capacity, quotas, snapshots, and lifecycle. Automate provisioning with Terraform/OpenTofu and GitOps, build observability, and troubleshoot performance and reliability issues across hybrid, edge, and air-gapped environments. The role also supports k0s, Cluster API, K0rdent, Harbor, and secure disconnected deployments.
The summary above was generated by AI
Company Description

About Mirantis

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. By combining open source innovation with deep expertise in Kubernetes orchestration, Mirantis empowers platform engineering teams to deliver composable, production-ready developer platforms across any environment—on-premises, in the cloud, at the edge, or in sovereign data centers. As enterprises navigate the growing complexity of AI-driven workloads, Mirantis delivers the automation, GPU orchestration, and policy-driven control needed to manage infrastructure with confidence and agility. Committed to open standards and freedom from lock-in, Mirantis ensures that customers retain full control of their infrastructure strategy.

Job Description

Overview

Deploy, integrate, and operate high-performance storage for GPU-accelerated compute and AI platforms. You will own the storage layer where Kubernetes meets bare metal — standing up NFS-based high-performance storage, wiring it into clusters via CSI, and tuning it to keep data flowing to GPU workloads at scale. Work spans hybrid, edge, and air-gapped deployments built on the Mirantis K0rdent stack.

About the Role

We are looking for a senior DevOps engineer who treats storage as infrastructure to be automated, observed, and tuned — not hand-managed. The right candidate is fluent in Kubernetes storage, deeply versed in Linux storage and networking fundamentals down to the kernel and NFS-client layer, and knows how to make high-performance NAS actually perform under demanding workloads. You should reach for infrastructure-as-code and GitOps by default, be self-directed in diagnosing performance and reliability issues end to end, set operational standards for others to follow, and communicate clearly across teams. Bare-metal hardware experience is a strong plus, but deep Linux storage knowledge is essential.

Responsibilities
1. Storage Integration & Operation

  • Integrate NFS-based high-performance storage (e.g., VAST, Dell PowerScale) into Kubernetes clusters via CSI, storage classes, and persistent volumes.
  • Tune the NFS data path — mount options, nconnect/RDMA, Linux client, and network settings — for high-throughput, low-latency GPU/AI workloads.
  • Deploy and operate storage services and operators; manage capacity, quotas, snapshots, and lifecycle.

2. Linux Platform & System Integration

  • Configure and optimize Linux systems for storage workloads, including driver setup, file system layout, network tuning, and kernel parameter optimization.
  • Deliver storage integration for k0s-based Kubernetes via Cluster API (CAPI) and K0rdent management/child cluster topologies.
  • Operate storage in fully disconnected (air-gapped) environments, including local artifact/mirror connectivity (Harbor) and PKI/TLS considerations.

3. Automation & Observability

  • Automate storage provisioning and configuration with infrastructure-as-code (Terraform/OpenTofu) and GitOps pipelines (ArgoCD or Flux).
  • Build monitoring, alerting, and observability for storage performance, capacity, and health.
  • Diagnose and resolve performance, reliability, and scaling issues across the storage stack.

Qualifications

Required qualifications:

  • 7+ years of experience in SRE or infrastructure operations

  • 5+ years of building/operating distributed production Storage systems at scale

  • Hands-on with High Performance Storage solutions (VAST, Weka, DDN, PowerScale)

  • Linux and Kubernetes storage fundamentals (NFS, CSI)

 

Preferred:

  • Bare-metal experience: hands-on experience with bare-metal host provisioning, raw disk/hardware layout, and physical server storage configurations.

  • Hands-on experience with VAST and/or Dell PowerScale.

  • Experience with GPUDirect Storage and RDMA/RoCE data paths

  • Experience with the Mirantis K0rdent stack (K0rdent Enterprise, K0rdent AI, k0s, MKE) and Cluster API.

  • Familiarity with other storage backends (Ceph, object/S3) and CSI driver operations.

  • Proven experience in sovereign or high-security air-gapped environments.

 

Additional Information

What does Mirantis offer you?

- Work with an established Silicon Valley leader in the cloud infrastructure industry;
- Work with exceptionally passionate, talented and engaging colleagues, helping Fortune 500 and Global 2000 customers implement next-generation cloud technologies;
- Be a part of cutting-edge, open-source innovation;
- Thrive in the high-energy environment of a young company where openness, collaboration, risk-taking, and continuous growth are valued;
- Professional development and training;
- Attend conferences and working groups;
- Company outings, happy hours, hackathons, and tech talks;
- Receive a competitive compensation package with a strong benefits plan.

It is understood that Mirantis, Inc. may use automated decision-making technology (ADMT) for specific employment-related decisions. Opting out of ADMT use is requested for decisions about evaluation and review connected with the specific employment decision for the position applied for. You also have the right to appeal any decisions made by ADMT by sending your request to [email protected]

By submitting your resume, you consent to the processing and storage of your personal data in accordance with applicable data protection laws, for the purposes of considering your application for current and future job opportunities.

We are a Leader for Container Management in G2 (#2 after AWS)!

Similar Jobs

An Hour Ago
In-Office or Remote
Mid level
Mid level
Aerospace • Information Technology • Software • Cybersecurity • Design • Defense • Manufacturing
Operates and maintains unmanned aircraft systems in remote and austere environments. Plans missions, briefs customers, monitors weather and airspace requirements, delivers flight and payload data, analyzes mission results, and completes reports. Performs system testing, troubleshooting, maintenance, inventory control, and customer support while coordinating with subject matter experts. Requires frequent deployment, physical outdoor work, an active Secret DoW clearance, and an FAA Class 2 medical certificate.
Top Skills: Faa Class 2 Medical CertificateIntegratorNetworkingNotamsRq-21ScaneagleSpinsUnmanned Aircraft Systems (Uas)
An Hour Ago
In-Office or Remote
Mid level
Mid level
Aerospace • Information Technology • Software • Cybersecurity • Design • Defense • Manufacturing
Operate and maintain unmanned aircraft systems during customer missions, including preflight planning, mission execution, full-motion video and payload data delivery, post-mission reporting, troubleshooting, maintenance, inventory control, and customer coordination. The role requires frequent deployment to remote or austere environments for four-to-six-month assignments, extensive travel, physical work outdoors, and an active Secret Department of War clearance.
Top Skills: Faa Class 2 Medical CertificateIntegratorNotamsRq-21ScaneagleSpinsUas
2 Hours Ago
Remote or Hybrid
United States
85K-143K Annually
Expert/Leader
85K-143K Annually
Expert/Leader
Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Oversee operational risk management for MetLife’s Group Benefits business by implementing the non-financial risk framework, advising business leaders, monitoring risk events and indicators, managing remediation, strengthening controls, and partnering with technology, operations, compliance, and control functions. The role also develops risk mitigation strategies, supports governance and risk appetite compliance, and improves risk practices through training, stakeholder engagement, analytics, and emerging technologies.
Top Skills: AIAutomationDashboardsRisk Analytics

What you need to know about the Los Angeles Tech Scene

Los Angeles is a global leader in entertainment, so it’s no surprise that many of the biggest players in streaming, digital media and game development call the city home. But the city boasts plenty of non-entertainment innovation as well, with tech companies spanning verticals like AI, fintech, e-commerce and biotech. With major universities like Caltech, UCLA, USC and the nearby UC Irvine, the city has a steady supply of top-flight tech and engineering talent — not counting the graduates flocking to Los Angeles from across the world to enjoy its beaches, culture and year-round temperate climate.

Key Facts About Los Angeles Tech

  • Number of Tech Workers: 375,800; 5.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Snap, Netflix, SpaceX, Disney, Google
  • Key Industries: Artificial intelligence, adtech, media, software, game development
  • Funding Landscape: $11.6 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Strong Ventures, Fifth Wall, Upfront Ventures, Mucker Capital, Kittyhawk Ventures
  • Research Centers and Universities: California Institute of Technology, UCLA, University of Southern California, UC Irvine, Pepperdine, California Institute for Immunology and Immunotherapy, Center for Quantum Science and Engineering

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account