Vast.ai Logo

Vast.ai

Technical Support Engineer II (Linux)

Posted 2 Days Ago
Be an Early Applicant
In-Office
Los Angeles, CA, USA
90K-130K Annually
Entry level
In-Office
Los Angeles, CA, USA
90K-130K Annually
Entry level
Provides escalated technical support for AI infrastructure across Linux, Docker, NVIDIA GPUs, CUDA, networking, storage, virtualization, and host configurations. Diagnoses complex workload and performance issues, supports supplier onboarding, develops Python and Bash automation, and maintains runbooks and knowledge articles. Collaborates with engineering on systemic platform problems, assists TensorFlow and PyTorch users, and provides L1 overflow coverage through an on-call rotation.
The summary above was generated by AI

About Us

Vast.ai's cloud powers AI projects and businesses all over the world. We are democratizing and decentralizing AI computing — reshaping our future for the benefit of humanity. Our mission is to organize, optimize, and orient the world's computation.

We value elegance, ownership, integrity, and continuous learning. You'll have the opportunity to dive into state-of-the-art AI systems while collaborating with a globally distributed team.

About the Role

This is a technical support role focused on escalated infrastructure issues that go beyond frontline triage. You'll be the engineering resource our L1 support team leans on when tickets get complex: diagnosing and resolving issues across the full stack — hardware/BIOS/firmware, networking, Ubuntu, Docker, NVIDIA CUDA/GPU, and virtualization (KVM).

You'll handle higher-complexity issues, own escalation resolution end-to-end, and contribute to internal documentation and runbooks. The best engineers in this role don't just resolve tickets — they build the tooling and runbooks that eliminate recurring ones. You'll collaborate directly with the engineering team and host support team on systemic issues.

Strong technical depth and support experience are the primary requirements. You should be comfortable working autonomously across Ubuntu environments, diagnosing container and GPU issues, and communicating findings clearly to both technical and non-technical audiences.

Vast.ai users or hosts strongly preferred.

This role is full-time and onsite in our office in Westwood (LA)
Schedule: Sunday - Thursday.

Key Responsibilities

  • Handle escalated support tickets, including GPU workload failures, container issues, networking problems, account infrastructure, and host-side configuration

  • Diagnose and resolve issues across Docker, NVIDIA CUDA/GPU drivers, and virtualization environments (KVM)

  • Troubleshoot network-layer issues: VLAN, DNS, DHCP, VPN, NAT, firewall rules, and connectivity failures on host machines

  • Investigate performance issues on GPU utilization, container resource constraints, thermal throttling, driver conflicts, disk I/O bottlenecks

  • Advise suppliers (hosts) on installation best practices — hardware setup, driver configuration, BIOS/firmware settings, and network configuration for optimal performance

  • Provide managed support for supplier onboarding and ongoing machine management, acting as a technical resource through installation, configuration, and post-setup troubleshooting

  • Write and maintain internal runbooks, escalation guides, and knowledge base articles to reduce repeat escalations

  • Build diagnostic and automation tooling in Python and Bash to reduce manual triage overhead

  • Collaborate with the engineering team and infrastructure support team to flag and document systemic or recurring platform issues

  • Assist clients and infrastructure suppliers working with AI frameworks (TensorFlow, PyTorch) and GPU-accelerated workloads

  • Provide coverage for L1 support team overflow during peak periods or incidents, per a defined on-call rotation

You Are

  • Fluent in Linux — you navigate systems, read logs, and solve problems from the command line without hesitation

  • Methodical and thorough: you gather data, dig into root causes, and don't settle for surface-level fixes

  • A self-starter who can manage a queue of complex tickets with minimal supervision

  • Adaptable to a defined on-call rotation which may include weekend coverage

  • A clear written communicator: able to explain technical findings and write useful internal documentation

  • Genuinely curious about AI infrastructure, GPU computing, and distributed systems

Must-Haves

  • Solid Linux SysOps experience: Ubuntu Server, RHEL/CentOS, Debian; comfortable with systems, networking, storage, and permissions

  • Proficiency with Docker: container debugging, Docker Compose, image management, cgroup resource limits, Docker storage/filesystem management

  • Experience with virtualization: Proxmox VE, VMware, or similar hypervisors; provisioning and troubleshooting VMs

  • Networking fundamentals: VLAN, DNS, DHCP, NAT, VPN, firewall rules, and general L2/L3 troubleshooting

  • Hands-on experience with NVIDIA GPU drivers, CUDA, and GPU workload troubleshooting (essential)

  • Scripting in Python and Bash for automation and diagnostic tooling

  • Strong English written communication: clear, professional, and technically precise

  • Experience providing technical support in a customer-facing or internal helpdesk context

  • Ability to prioritize across a concurrent queue of escalated tickets, triaging by severity and customer impact, balancing reactive resolution against proactive documentation and tooling work, and making clear judgment calls on when to escalate versus own resolution end-to-end

Nice-to-Haves

  • Familiarity with AI/ML frameworks (TensorFlow, PyTorch) and running GPU-accelerated containers

  • Monitoring and observability experience (Prometheus, Grafana)

  • Relevant certifications: RHCSA, CompTIA Linux+, or similar

  • Knowledge of the Vast.ai platform as a client or infrastructure supplier

Annual Salary Range

$90,000 – $150,000 + equity + benefits

Vast.ai is hiring across all experience levels with compensation commensurate with background, experience and potential.

Benefits
  • Comprehensive health, dental, vision, and life insurance

  • 401(k) with company match

  • Meaningful early-stage equity

  • Onsite meals, snacks, and close collaboration with founders/tech leaders

  • Ambitious, fast-paced startup culture where initiative is rewarded

 

HQ

Vast.ai Los Angeles, California, USA Office

Los Angeles, CA, United States

Similar Jobs

35 Minutes Ago
Remote or Hybrid
USA
Senior level
Senior level
Machine Learning • Payments • Security • Software • Financial Services
Owns the vision, customer focus, and product backlog for a near-real-time data product. Prioritizes work based on business value, leads backlog grooming, communicates product direction, and partners with Scrum Masters and development teams to ensure delivery aligns with client requirements and business objectives.
Top Skills: Agile DevelopmentData VisualizationScrumUx Design
An Hour Ago
Remote or Hybrid
CA, USA
149K-248K Annually
Senior level
149K-248K Annually
Senior level
eCommerce • Fintech • Hardware • Payments • Software • Financial Services
Leads Square’s US Business Development Representative organization through managers and frontline teams. Owns pipeline generation strategy, operating cadence, performance management, forecasting, capacity planning, outbound execution, and sales technology improvements. Coaches managers, develops BDR talent, hires and retains staff, and partners with Sales, Marketing, Revenue Operations, Enablement, Analytics, and Strategy to improve funnel conversion and revenue contribution. Uses data, automation, AI, and prospecting tools to improve productivity and scale successful programs.
Top Skills: Artificial IntelligenceAutomationCRMData EnrichmentGongLookerOutreachSales Engagement ToolsSalesforceSalesloft
4 Hours Ago
Hybrid
37K-66K Hourly
Senior level
37K-66K Hourly
Senior level
Fintech • Financial Services
Manage and grow relationships with affluent customers, acquire new clients, provide multi-product financial and credit guidance, coordinate referrals across Wealth/Home Lending/Business Banking, handle account openings and service requests, and maintain compliance and required licensing.

What you need to know about the Los Angeles Tech Scene

Los Angeles is a global leader in entertainment, so it’s no surprise that many of the biggest players in streaming, digital media and game development call the city home. But the city boasts plenty of non-entertainment innovation as well, with tech companies spanning verticals like AI, fintech, e-commerce and biotech. With major universities like Caltech, UCLA, USC and the nearby UC Irvine, the city has a steady supply of top-flight tech and engineering talent — not counting the graduates flocking to Los Angeles from across the world to enjoy its beaches, culture and year-round temperate climate.

Key Facts About Los Angeles Tech

  • Number of Tech Workers: 375,800; 5.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Snap, Netflix, SpaceX, Disney, Google
  • Key Industries: Artificial intelligence, adtech, media, software, game development
  • Funding Landscape: $11.6 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Strong Ventures, Fifth Wall, Upfront Ventures, Mucker Capital, Kittyhawk Ventures
  • Research Centers and Universities: California Institute of Technology, UCLA, University of Southern California, UC Irvine, Pepperdine, California Institute for Immunology and Immunotherapy, Center for Quantum Science and Engineering

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account