VAST Data Logo

VAST Data

Senior Solutions Engineer, AI Infrastructure

Posted 20 Days Ago
Remote or Hybrid
Hiring Remotely in United States
Senior level
Remote or Hybrid
Hiring Remotely in United States
Senior level
The Senior Solutions Engineer will design and implement infrastructure for AI and HPC workloads, engage with customers, and lead technical discovery and architecture design.
The summary above was generated by AI
Description

We're looking for a deeply technical Solutions Architect to help customers design, evaluate, and deploy infrastructure for large-scale AI, HPC, analytics, and data-intensive workloads.

This is a customer-facing technical role for someone who has lived inside production infrastructure. You may have been a platform engineer, infrastructure engineer, SRE, MLOps engineer, AI infrastructure engineer, storage engineer, cloud engineer, or HPC systems engineer. What matters most is that you have built, operated, or architected real systems, and can bring that credibility into customer conversations.

Our customers are building infrastructure at serious scale: GPU clusters, high-performance storage systems, Kubernetes platforms, distributed training environments, inference platforms, data pipelines, lakehouses, and large enterprise systems. You'll help them reason about architectures involving 10,000+ GPUs, 100PB+ of storage, high-performance networking, distributed filesystems, orchestration layers, and demanding production workloads.

You'll own technical discovery, architecture design, PoC planning, competitive positioning, and customer technical strategy. You'll work from the first whiteboard session through evaluation, deployment planning, and production success. You'll also partner closely with product and engineering teams to bring field feedback into the roadmap.

We're looking for someone who can go deep technically, communicate clearly, operate without a rigid playbook, and translate complex infrastructure into customer outcomes.

Responsibilities

  • Lead technical discovery with customers across infrastructure, platform, ML, data, and executive stakeholders.
  • Design architectures for large-scale AI, HPC, analytics, and enterprise data workloads.
  • Help customers evaluate infrastructure involving GPUs, storage, networking, orchestration, and data movement.
  • Translate complex technical requirements into clear solution designs, reference architectures, and deployment guidance.
  • Debug customer issues across Linux, storage, networking, Kubernetes, schedulers, GPUs, and application workloads.
  • Build technical assets, runbooks, and field guidance for repeatable customer engagements.
  • Partner with product and engineering to communicate customer requirements, gaps, and roadmap opportunities.
  • Help customers move from architecture design to production deployment.
Requirements
  • 8 to 12+ years of technical experience, with significant hands-on infrastructure experience.
  • Experience building, operating, or architecting production platform infrastructure.
  • Strong understanding of Linux kernel implementation details, distributed systems including PAXOS and raft, storage implementations details like NAND or write amplification, networking store/forward, load balancing designs, and production operations.
  • Experience with one or more of: GPU infrastructure, large scale HPC systems, Kubernetes platforms from scratch, MLOps, storage systems, cloud infrastructure, data platforms, or large-scale enterprise infrastructure.
  • Ability to communicate credibly with engineers, architects, technical executives, and business stakeholders.
  • Strong discovery, problem-solving, and systems debugging skills.
  • Comfort operating in ambiguous, fast-moving environments.
  • Interest in customer-facing technical work, solution design, and business outcomes.

Preferred Experience

  • Experience with large-scale GPU clusters, distributed training, inference infrastructure, or AI platforms.
  • Experience with petabyte-scale storage or high-performance data systems.
  • Experience with Kubernetes, Slurm, Ray, Spark, or other orchestration / scheduling systems.
  • Domain Expertise with one or more of these - Lustre, Ceph, Weka, BeeGFS, GPFS, VAST, object storage, or distributed filesystems.
  • Experience with large-scale InfiniBand, RoCE, RDMA, high-performance Ethernet, or NVIDIA/Mellanox networking.
  • Direct Experience with CUDA, NCCL, DCGM, GPUDirect, checkpointing, dataset staging, or model-serving infrastructure.
  • Experience across multiple industries or customer environments.

Similar Jobs

An Hour Ago
In-Office or Remote
143K-258K Annually
Senior level
143K-258K Annually
Senior level
Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
Lead data strategy and analytics for Compliance, focusing on Financial Crimes/AML. Drive ETL, source validation, and analyses to quantify compliance risk, support regulatory engagements, and influence product and compliance roadmaps. Communicate insights through visualization and cross-functional collaboration while leading multiple high-impact workstreams.
Top Skills: ETLLookerPrefectPythonRSQLTableau
3 Hours Ago
In-Office or Remote
113K-193K Annually
Senior level
113K-193K Annually
Senior level
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Lead product strategy and execution for large, data-driven enterprise platforms. Own roadmap, requirements, and delivery across cross-functional teams; translate business needs into measurable outcomes. Partner with engineering and data science to scale AI/ML capabilities, ensure responsible implementation, and mentor product teams in a regulated healthcare payer environment.
Top Skills: AgileAi/MlData PlatformsData ScienceDistributed SystemsEnterprise Ai ToolsGenerative AiModern Application Architectures
3 Hours Ago
In-Office or Remote
113K-193K Annually
Senior level
113K-193K Annually
Senior level
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Manage strategic PBM relationships with health plan clients: lead contract renewals, drive retention and profitability, present performance reviews, implement benefit designs, supervise client implementations and teams, identify cost-savings and upsell opportunities, and maintain client communications and compliance.
Top Skills: ExcelMicrosoft PowerpointMicrosoft WordNavigatorRxclaimTracker

What you need to know about the Los Angeles Tech Scene

Los Angeles is a global leader in entertainment, so it’s no surprise that many of the biggest players in streaming, digital media and game development call the city home. But the city boasts plenty of non-entertainment innovation as well, with tech companies spanning verticals like AI, fintech, e-commerce and biotech. With major universities like Caltech, UCLA, USC and the nearby UC Irvine, the city has a steady supply of top-flight tech and engineering talent — not counting the graduates flocking to Los Angeles from across the world to enjoy its beaches, culture and year-round temperate climate.

Key Facts About Los Angeles Tech

  • Number of Tech Workers: 375,800; 5.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Snap, Netflix, SpaceX, Disney, Google
  • Key Industries: Artificial intelligence, adtech, media, software, game development
  • Funding Landscape: $11.6 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Strong Ventures, Fifth Wall, Upfront Ventures, Mucker Capital, Kittyhawk Ventures
  • Research Centers and Universities: California Institute of Technology, UCLA, University of Southern California, UC Irvine, Pepperdine, California Institute for Immunology and Immunotherapy, Center for Quantum Science and Engineering

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account