Vast.ai

HQ
Los Angeles
41 Total Employees
Year Founded: 2018

Jobs at Vast.ai

Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.

Recently posted jobs

12 Days AgoSaved
In-Office
Los Angeles, CA, USA
Artificial Intelligence • On-Demand • Software
Own the infrastructure, security, scalability, observability, and reliability roadmap for a decentralized GPU cloud marketplace. Translate non-functional requirements into engineering specifications, prioritize initiatives, and drive delivery. Manage platform metrics including uptime, fleet reliability, latency, and cost-to-serve. Partner across engineering and leadership to improve distributed systems, compliance, abuse prevention, and infrastructure tooling at significant scale.
12 Days AgoSaved
In-Office
Los Angeles, CA, USA
Artificial Intelligence • On-Demand • Software
Execute manual, exploratory, regression, and automated testing for web applications and backend services. Develop test plans and cases, validate requirements, investigate failures, identify root causes, and file actionable defects. Collaborate with Product and Engineering on acceptance criteria, quality standards, and release readiness in a Linux-first environment.
12 Days AgoSaved
In-Office
Los Angeles, CA, USA
Artificial Intelligence • On-Demand • Software
Troubleshoot complex Linux, NVIDIA GPU, CUDA, Docker, virtualization, networking, hardware, and AI workload issues. Own escalated support tickets, assist clients and infrastructure suppliers, identify root causes, and collaborate with engineering teams. Build Python and Bash diagnostic tooling, create runbooks, improve monitoring and troubleshooting processes, and support supplier onboarding and machine management. The role is fully or mostly on-site in Los Angeles, with schedules covering Monday-Friday or Sunday-Thursday.
12 Days AgoSaved
In-Office
Los Angeles, CA, USA
Artificial Intelligence • On-Demand • Software
Design and implement secure architectures for a multi-tenant GPU cloud platform. Conduct threat modeling, security assessments, code reviews, penetration testing, and incident response. Manage SIEM, WAF, and EDR tools; develop security improvements with engineering teams; support compliance with SOC 2, ISO 27001, GDPR, NIST, and CIS Controls; maintain security documentation; and train development teams on secure practices.
12 Days AgoSaved
In-Office
Los Angeles, CA, USA
Artificial Intelligence • On-Demand • Software
Lead and scale Vast.ai’s two-sided GPU marketplace sales organization. Coach existing enterprise demand and strategic supply representatives, hire and onboard additional sellers, expand outreach, manage pipeline forecasting and CRM discipline, develop the sales technology stack, design compensation plans, set quotas, and manage strategic accounts. Partner with Marketing, Product, and Engineering to align go-to-market strategy and support aggressive revenue growth.
12 Days AgoSaved
In-Office
2 Locations
Artificial Intelligence • On-Demand • Software
Deploys, fine-tunes, and serves open-source AI models on rented GPUs using tools such as vLLM and SGLang. Builds public example repositories, Docker templates, benchmarks, and live endpoints, then teaches through technical writing, videos, talks, livestreams, and community support. Reports product friction to engineering and product teams, represents Vast at hackathons and conferences, and serves as a highly credible technical voice across developer communities.
12 Days AgoSaved
In-Office
Los Angeles, CA, USA
Artificial Intelligence • On-Demand • Software
Provides escalated technical support for AI infrastructure across Linux, Docker, NVIDIA GPUs, CUDA, networking, storage, virtualization, and host configurations. Diagnoses complex workload and performance issues, supports supplier onboarding, develops Python and Bash automation, and maintains runbooks and knowledge articles. Collaborates with engineering on systemic platform problems, assists TensorFlow and PyTorch users, and provides L1 overflow coverage through an on-call rotation.
12 Days AgoSaved
In-Office
2 Locations
Artificial Intelligence • On-Demand • Software
Design and scale infrastructure powering a global GPU marketplace. Responsibilities include GPU provisioning, workload scheduling, orchestration, provider onboarding, usage tracking, billing, pricing logic, resource management, and marketplace APIs. The engineer will optimize performance, reliability, security, fault tolerance, and multi-tenant cloud systems while collaborating with product and infrastructure teams. The role requires strong Python and C++ development, distributed systems expertise, and experience building large-scale compute platforms.
12 Days AgoSaved
In-Office
Los Angeles, CA, USA
Artificial Intelligence • On-Demand • Software
Lead accounting and finance operations for a high-volume AI-compute marketplace. Responsibilities include cash-to-accrual conversion, GAAP reporting, monthly close, audit readiness, marketplace revenue recognition, payout and cash controls, tax compliance, finance systems implementation, forecasting inputs, and cross-functional financial reporting. The Controller will build scalable, automated processes and eventually hire and lead a finance team.
12 Days AgoSaved
In-Office
2 Locations
Artificial Intelligence • On-Demand • Software
Develop GPU kernels, parallel libraries, tensor libraries, and auto-optimization tools to improve AI inference performance. Research and evaluate state-of-the-art models, frameworks, tools, and system architectures. Contribute to market-based resource management systems while collaborating with technical leadership and engineering teams. The role requires deep GPU architecture knowledge, neural network performance expertise, systems engineering skills, and strong research capabilities.
12 Days AgoSaved
In-Office
Los Angeles, CA, USA
Artificial Intelligence • On-Demand • Software
Investigate and resolve escalated infrastructure issues across Linux, hardware, networking, Docker, GPUs, CUDA, and virtualization environments. Manage supplier onboarding, troubleshoot GPU workloads and performance problems, support L1 overflow, create diagnostic automation with Python and Bash, and maintain runbooks and escalation documentation. Collaborate with engineering and support teams to address systemic infrastructure issues and communicate solutions clearly.
12 Days AgoSaved
In-Office
2 Locations
Artificial Intelligence • On-Demand • Software
Develop and extend a GPU cloud daemon that orchestrates hosts across a large fleet. Improve infrastructure performance, reliability, monitoring, telemetry, and containerization. Design market-based resource management systems, harden code and infrastructure to zero-trust standards, and benchmark and eliminate bottlenecks across hypervisor, container, and network layers.
12 Days AgoSaved
In-Office
2 Locations
Artificial Intelligence • On-Demand • Software
Design and optimize GPU kernels and tensor libraries for scalable AI inference. Apply HPC and parallel-computing techniques, evaluate emerging GPU architectures and resource-management approaches, and improve GPU infrastructure efficiency. The role requires advanced C++ development, parallel programming expertise, systems optimization, and performance tooling experience. This is a full-time, on-site position in San Francisco or Los Angeles.