Lambda Logo

Lambda

Technical Account Manager

Posted Yesterday
Be an Early Applicant
Remote or Hybrid
Hiring Remotely in San Jose, CA
231K-342K Annually
Senior level
Remote or Hybrid
Hiring Remotely in San Jose, CA
231K-342K Annually
Senior level
Own the technical health of post-sales AI cloud accounts, including architecture, GPU workload performance, POCs, SLA validation, reliability, escalations, and customer advocacy. Build dashboards, telemetry, health scoring, and early-warning tools while guiding training, fine-tuning, and inference workloads. Lead technical incidents, root-cause analyses, feature requests, and expansion recommendations. The role requires strong customer-facing infrastructure expertise across compute, networking, storage, and Slurm or Kubernetes, plus executive communication and comfort creating processes in ambiguous environments.
The summary above was generated by AI

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.

If you'd like to build the world's best AI cloud, join us.

*Note: This position requires presence in our San Francisco, San Jose, or Bellevue office location 4 days per week; Lambda’s designated work from home day is currently Tuesday.


The Technical Account Manager owns the technical health of the post-sales relationship for Lambda’s public cloud accounts, spanning AI-native startups, Enterprises, and Fortune 500 companies. Where the Customer Success Manager owns the commercial health of an account, you own its technical health: the customer’s workloads run well, the architecture is right, the SLA story is defensible, and technical risk is found and retired before it threatens revenue. Solutions Engineering carries the account through pre-sales and hypercare; at handoff, you take ownership of the technical relationship for the life of the contract.

This is a hands-on role, not a coordination role. You will understand what customers are actually building (training runs, fine-tuning pipelines, inference services) deeply enough to lead joint POC sessions, design and defend architectures, validate SLA events at the root-cause level, and build the tooling that makes account health measurable. You will be the customer’s most credible technical advocate inside Lambda and Lambda’s most trusted technical voice inside the account.

What You’ll Do

  • Own the technical health of your accounts. Take the technical handoff from Solutions Engineering at the end of hypercare and own the account’s technical outcomes through steady state, expansion, and renewal. Know the state of every cluster and workload you are accountable for, and keep your commercial counterparts ahead of technical risk.

  • Understand customer AI use cases end to end. Map what it means for each customer to train, fine-tune, and serve models on Lambda: frameworks, schedulers, parallelism strategy, data paths, and performance baselines. Build the customer user journey and convert it into value-add opportunities across the platform, documentation, and escalation routing.

  • Lead joint customer POC sessions. Define success criteria with the customer before a node is provisioned: acceptance thresholds, benchmarks, timelines. Own the execution plan, coordinate capacity and provisioning, run or oversee the tests, and drive the POC to a clear verdict: win the workload, close the gap through product, or qualify out.

  • Lead customer architecture designs. Produce and defend reference architectures spanning compute, networking, storage, connectivity, and scheduler integration (Slurm, Kubernetes). Make support boundaries explicit: what is managed and what is not. Own the design as it evolves after handoff, pulling in engineering domain experts with specific, well-framed questions.

  • Own SLA and reliability engineering. Build and own the canonical methodology for uptime, downtime, and credit calculation. Validate breach events at the technical level, down to the specific Ethernet or InfiniBand failure, and arm CSMs and leadership with defensible numbers. Partner with product to standardize SLA language and structure across 1CC, on-demand, and reserved offerings.

  • Build the tooling that makes accounts measurable. Own the technical data surfaces for customer health end to end: dashboards, telemetry and uptime history views, health scoring, and churn early-warning signals. Scope, build, and drive adoption. Replace “escalate to engineering to answer a basic question” with self-serve data for the whole GTM org.

  • Direct technical escalations and incidents. Serve as the technical lead during high-severity events on your accounts: drive root cause, hold the quality bar on RCAs, coordinate engineering, support, and vendors (NVIDIA, storage, networking), and give account teams a technically accurate narrative. Run proactive maintenance and known-issue communication so customers hear about problems from Lambda first.

  • Drive the technical voice of the customer. Run a structured feature-request pipeline into product with committed triage timelines. Audit the platform hands-on by provisioning as a customer and testing known friction points. Lead product-led POCs (for example, NVIDIA NIM) that open new value for customers.

  • Know the market technology landscape. Track GPU roadmaps, competing clouds and neoclouds, and the evolving training and inference stacks. Brief customers on what is coming and internal teams on where Lambda stands, and let that context shape architecture and expansion recommendations.

You

  • 5+ years in technical account management, solutions engineering or architecture, ML engineering, technical program or product management, or infrastructure engineering with significant customer-facing scope, in cloud, HPC, or AI infrastructure.

  • Hands-on fluency with GPU infrastructure: able to provision, benchmark, and debug across compute, networking (InfiniBand, Ethernet), storage, and schedulers (Slurm, Kubernetes), and to read results critically.

  • Working command of AI/ML workloads (training, fine-tuning, inference) sufficient to map a customer’s stack, identify constraints, and lead technical conversations with their ML and infrastructure engineers.

  • Track record leading structured technical engagements: POCs with defined success criteria, architecture designs, benchmark programs, or high-severity escalations.

  • A builder’s toolkit: scripting, SQL, and dashboarding, with a history of turning operational data into tools other people depend on.

  • Executive-grade communication of deeply technical content, in writing and in the room.

  • Comfort with ambiguity and a track record of building methodology where none exists.

Nice to Have

  • Experience at an AI cloud, neocloud, or hyperscaler serving large-scale GPU or HPC customers.

  • Applied LLM experience (fine-tuning, RAG systems, or inference serving) that mirrors the workloads Lambda customers run.

  • Depth in the NVIDIA ecosystem: NIM and NeMo, Base Command / BCM, the CUDA stack, DGX-class systems.

  • Familiarity with SLA structures, service credits, enterprise contract mechanics, and retention metrics (NRR/GRR).

  • Product management or TPM background with experience converting customer evidence into roadmap decisions.

Salary Range Information

The annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.

About Lambda

  • Founded in 2012, with 500+ employees, and growing fast

  • Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove

  • We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG

  • Our values are publicly available: https://lambda.ai/careers

  • We offer generous cash & equity compensation

  • Health, dental, and vision coverage for you and your dependents

  • Wellness and commuter stipends for select roles

  • 401k Plan with 2% company match (USA employees)

  • Flexible paid time off plan that we all actually use

Equal Opportunity Employer

Lambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.

Similar Jobs

23 Days Ago
Remote
United States
245K-245K Annually
Senior level
245K-245K Annually
Senior level
Security • Cybersecurity
Lead and grow a Technical Account Management team to drive Dragos Platform adoption in industrial security environments. Serve as escalation point, partner with engineering/product to prioritize customer needs, define KPIs, operationalize use cases, and improve processes and tooling to scale customer success. Advocate AI-driven improvements and collaborate with sales, solutions architects, and intel teams to prove platform value.
Top Skills: AIDragos PlatformElasticSIEMSplunk
Yesterday
Remote or Hybrid
231K-342K Annually
Senior level
231K-342K Annually
Senior level
Artificial Intelligence • Cloud • Machine Learning • Infrastructure as a Service (IaaS)
Own the technical health of post-sales AI cloud accounts, including customer workloads, architectures, POCs, SLAs, reliability, tooling, escalations, and product feedback. Design GPU infrastructure solutions across compute, networking, storage, and schedulers; lead benchmarks and incident response; build account-health dashboards; and advise customers and internal teams on AI infrastructure trends, workload performance, and expansion opportunities.
Top Skills: Base Command/BcmEthernetGpu InfrastructureInfinibandKubernetesLlm Fine-TuningNvidia CudaNvidia DgxNvidia NemoNvidia NimRag SystemsSlurmSQL
2 Days Ago
Remote
United States
139K-288K Annually
Senior level
139K-288K Annually
Senior level
Information Technology
Serve as a trusted post-sales technical advisor to strategic enterprise customers. Drive Docker adoption, operational maturity, value realization, risk mitigation, and renewals while managing technical relationships. Partner with Solutions Engineering, Product, Engineering, Support, and Sales; provide guidance on containers, Kubernetes, CI/CD, infrastructure as code, and developer productivity. Maintain customer health documentation, lead reviews and enablement programs, develop scalable playbooks, represent customer feedback internally, and mentor other TAMs.
Top Skills: Ai/Ml InfrastructureCi/CdCloud-Native DevelopmentDockerDocker DesktopDocker HubDocker ScoutInfrastructure As Code (Iac)Kubernetes

What you need to know about the Los Angeles Tech Scene

Los Angeles is a global leader in entertainment, so it’s no surprise that many of the biggest players in streaming, digital media and game development call the city home. But the city boasts plenty of non-entertainment innovation as well, with tech companies spanning verticals like AI, fintech, e-commerce and biotech. With major universities like Caltech, UCLA, USC and the nearby UC Irvine, the city has a steady supply of top-flight tech and engineering talent — not counting the graduates flocking to Los Angeles from across the world to enjoy its beaches, culture and year-round temperate climate.

Key Facts About Los Angeles Tech

  • Number of Tech Workers: 375,800; 5.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Snap, Netflix, SpaceX, Disney, Google
  • Key Industries: Artificial intelligence, adtech, media, software, game development
  • Funding Landscape: $11.6 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Strong Ventures, Fifth Wall, Upfront Ventures, Mucker Capital, Kittyhawk Ventures
  • Research Centers and Universities: California Institute of Technology, UCLA, University of Southern California, UC Irvine, Pepperdine, California Institute for Immunology and Immunotherapy, Center for Quantum Science and Engineering

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account