TensorWave Logo

TensorWave

Senior Network Engineer

Reposted 10 Days Ago
Remote
Hiring Remotely in USA
Senior level
Remote
Hiring Remotely in USA
Senior level
This role involves developing and managing a networking infrastructure for AI cloud services, integrating new technologies, ensuring network reliability, and collaborating with IT and AI teams.
The summary above was generated by AI

Our mission at Tensorwave Cloud is to build seamless, secure, reliable, and resilient AI infrastructure at scale, eliminating barriers and challenging the status quo to empower builders and support AI innovation.

About the role

We are seeking a Senior Network Engineer focused on implementing and operating large-scale, Arista-based RoCEv2 data center networks powering next generation AI and ML infrastructure.

You’ll work hand-in-hand with our network architect to design the infrastructure that keeps over 8,000 GPUs burring, and play a critical role in the implementation and maintenance of our next generation systems, with cluster sizes reaching over 100,000 GPUs.

You’ll work hands-on with high-speed optics, switching, and routing in production clusters and implement modern automation and tooling critical to how the network is deployed, validated, and operated.

Responsibilities

  • Design, deploy, and operate large-scale RoCEv2 data center networks supporting AI and ML clusters from thousands to 100,000+ GPUs

  • Own congestion management and performance tuning across RDMA fabrics, including PFC, ECN, and DCQCN, in production environments

  • Implement and maintain automation, validation, and observability tooling using Python, Ansible, Terraform, and modern DevOps workflows

  • Ensure high availability and reliability across multi-tenant environments by leading operational excellence, incident response, and continuous improvement

Required Experience

  • Bachelor of Science in Computer Science, Computer Engineering, or a related technical field, or equivalent practical experience

  • Deep experience with RDMA and RoCEv2 in large-scale production data centers supporting AI or HPC workloads

  • Strong Arista expertise, including EOS, hardware platforms, and operating high-speed Ethernet fabrics

  • Proven knowledge of congestion management and performance tuning using PFC, ECN, and DCQCN

  • Hands-on experience with high-speed optics and cabling including 400G, 800G, and AEC, AOC, DAC, and structured cabling in dense environments

  • Automation and operations mindset, with experience using Python, Ansible, Terraform, Git, and observability tooling in always-on production systems

What We Bring

  • Mission driven company

  • Competitive Salary

  • Stock Options

  • 100% paid Medical, Dental, and Vision insurance

  • Flexible PTO

  • Paid Holidays

  • 401(k)

  • Parental Leave

  • Flexible Spending Account

  • Short Term Disability Insurance

  • Life and Voluntary Supplemental Insurance

  • Mental Health Benefits through Spring Health

We’re looking for resilient, adaptable people to join our team, people who believe in the mission and think at massive scale. The solutions that worked on a handful of devices will not work at Exascale. Be prepared to be pushed daily, to learn a lot, and literally build the future.

Tensorwave is an equal opportunity employer, committed to fostering an inclusive and supportive workplace. All qualified applicants and candidates will receive consideration for employment without regard to race, color, religion, sex, disability, age, national origin, or veteran status.

Top Skills

Amd Gpu
Bgp
Ethernet Protocols
Nvidia Gpu
Rocev2

Similar Jobs

14 Days Ago
Remote
United States
203K-274K Annually
Expert/Leader
203K-274K Annually
Expert/Leader
Artificial Intelligence • Cloud • Consumer Web • Productivity • Software • App development • Data Privacy
Responsible for managing Dropbox's cloud networking infrastructure and ensuring reliability, scalability, and security across multi-cloud environments while collaborating with various teams to drive improvements in automation and observability.
Top Skills: AnsibleAWSAzureDatadogGCPOciPythonTerraform
14 Days Ago
Remote or Hybrid
San Diego, CA, USA
125K-213K Annually
Senior level
125K-213K Annually
Senior level
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Operate and maintain cloud network infrastructure, troubleshoot network issues, analyze metrics, and automate deployments in a hybrid cloud environment, collaborating with teams to enhance operational reliability.
Top Skills: AnsibleAWSAzureBashDockerF5GCPGitlabGrafanaKubernetesNginxPrometheusPythonSplunk
14 Days Ago
Remote or Hybrid
Santa Clara, CA, USA
125K-213K Annually
Senior level
125K-213K Annually
Senior level
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
The role involves maintaining cloud network infrastructure, troubleshooting issues, managing hybrid environments, and driving automation for rapid deployment.
Top Skills: AnsibleAWSAzureBashCiscoDockerF5GCPGitlab Ci/CdGrafanaJuniperKubernetesLinuxNginxPalo AltoPrometheusPythonRadwareSplunkTerraform

What you need to know about the Los Angeles Tech Scene

Los Angeles is a global leader in entertainment, so it’s no surprise that many of the biggest players in streaming, digital media and game development call the city home. But the city boasts plenty of non-entertainment innovation as well, with tech companies spanning verticals like AI, fintech, e-commerce and biotech. With major universities like Caltech, UCLA, USC and the nearby UC Irvine, the city has a steady supply of top-flight tech and engineering talent — not counting the graduates flocking to Los Angeles from across the world to enjoy its beaches, culture and year-round temperate climate.

Key Facts About Los Angeles Tech

  • Number of Tech Workers: 375,800; 5.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Snap, Netflix, SpaceX, Disney, Google
  • Key Industries: Artificial intelligence, adtech, media, software, game development
  • Funding Landscape: $11.6 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Strong Ventures, Fifth Wall, Upfront Ventures, Mucker Capital, Kittyhawk Ventures
  • Research Centers and Universities: California Institute of Technology, UCLA, University of Southern California, UC Irvine, Pepperdine, California Institute for Immunology and Immunotherapy, Center for Quantum Science and Engineering

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account