Together AI Logo

Together AI

Senior Network Operations Engineer

Reposted 8 Days Ago
In-Office
San Francisco, CA
160K-230K Annually
Mid level
In-Office
San Francisco, CA
160K-230K Annually
Mid level
The Senior Network Operations Engineer handles incident response for network issues, collaborates with teams, troubleshoots, and improves network monitoring and documentation.
The summary above was generated by AI

As a Senior Network Operations Engineer at Together AI, you are our front-line responder for break/fix incidents—owning alert triage, collaborating with SRE and MLOps teams, and driving rapid resolution to keep our global network and platform running smoothly. You combine strong operational discipline with hands-on troubleshooting and a bias for automation. Beyond traditional networking, you’ll work hands-on with Kubernetes and Slurm to diagnose issues that span infrastructure, container networking, and HPC job fabrics.

You’re fluent in routing/switching and network security fundamentals, comfortable on Linux, and thrive in fast-moving environments where clear communication and crisp execution matter. You’ll improve monitoring, runbooks, and recovery playbooks to reduce MTTA/MTTR and prevent repeat incidents.

Outstanding problem-solving abilities and a solid understanding of fundamental network theory are also critical to your success.

Responsibilities

  • Serve as first responder for network alerts and incidents: assess impact, prioritize, mitigate, and escalate as needed to SRE/MLOps/Network Engineering.
  • Own end-to-end incident lifecycle: detection, triage, containment, remediation, comms, and post-incident reviews with clear timelines and action items.
  • Monitor network health and capacity across routing/switching, firewalls, and data center fabrics; tune alert thresholds and dashboards to reduce noise.
  • Troubleshoot L2–L4 issues (ARP, VLAN/VXLAN/EVPN, routing protocols, ACLs/NAT, DNS, TLS termination, QoS) using packet capture and flow/telemetry tools.
  • Execute standard changes (MOPs) and emergency changes with rigorous change control and validation; document outcomes and update runbooks.
  • Operate multi-cluster add-ons (e.g., MetalLB/Traefik/NGINX), observe health via Prometheus/Grafana/Loki, and tune alerts to reduce noise.
  • Debug CNI/data plane (e.g., VXLAN/EVPN, iptables/nftables, network policies), kube-proxy/iptables mode, CoreDNS, Services (ClusterIP/NodePort/LoadBalancer), and Ingress/EGRESS.
  • Maintain accurate network documentation: diagrams, inventories, IPAM, device configs, and topology state.
  • Improve operational excellence: automate repetitive tasks, enhance self-service tooling, and contribute to SLOs, error budgets, and reliability roadmaps.
  • Participate in a shared on-call rotation providing 24×7 coverage for critical services.

Requirements

  • 3+ years in a NOC/Network Operations or Network Support role for large-scale data center or service provider-style environments (hybrid/on-prem + cloud).
  • Solid understanding of TCP/IP and core protocols: BGP, OSPF/IS-IS, VLAN, VXLAN, EVPN, ACLs/NAT, DHCP, DNS, and QoS.
  • Proficiency with troubleshooting tools: Wireshark/tcpdump, mtr/traceroute, nmap, curl, iperf; comfortable on Linux for diagnostics and log analysis.
  • Experience operating multi-vendor networks (e.g., Arista, Cisco, Juniper, NVIDIA/Mellanox) and load balancers/firewalls.
  • Familiarity with AWS/GCP/Azure networking concepts (VPC/VNet, IGW/NATGW, peering, PrivateLink, routing, security groups).
  • Strong scripting/automation fundamentals (e.g., Bash/Python), and comfort with Git-based workflows for config versioning and change reviews.
  • Clear, concise communicator—able to write incident timelines, RCAs, and user-facing updates under time pressure.

Preferred

  • Knowledge of RoCE and Infiniband protocols a plus
  • Hands-on Kubernetes troubleshooting experience: CNI fundamentals (policies, encapsulation), Services/Ingress, DNS (CoreDNS), kube-proxy, and container runtime basics a huge plus
  • Understanding of AI training workloads and the demands they exert on networks a plus

About Together AI

Together AI is a research-driven artificial intelligence company. We believe open and transparent AI systems will drive innovation and create the best outcomes for society, and together we are on a mission to significantly lower the cost of modern AI systems by co-designing software, hardware, algorithms, and models. We have contributed to leading open-source research, models, and datasets to advance the frontier of AI, and our team has been behind technological advancement such as FlashAttention, Hyena, FlexGen, and RedPajama. We invite you to join a passionate group of researchers and engineers in our journey in building the next generation AI infrastructure.

Compensation

We offer competitive compensation, startup equity, health insurance and other competitive benefits. The US base salary range for this full-time position is: $160,000 - $230,000 + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge.

Equal Opportunity

Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.

Please see our privacy policy at https://www.together.ai/privacy  

Top Skills

Arista
AWS
Azure
Bash
Cisco
Curl
GCP
Iperf
Juniper
Kubernetes
Linux
Mtr
Nmap
Nvidia/Mellanox
Python
Slurm
Tcpdump
Traceroute
Wireshark

Similar Jobs

10 Days Ago
Remote or Hybrid
Santa Clara, CA, USA
125K-213K Annually
Senior level
125K-213K Annually
Senior level
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
As a Senior Network Operations Engineer, you'll manage and optimize cloud network infrastructure, troubleshoot issues, and enhance automation for reliable application delivery.
Top Skills: AnsibleAWSAzureBashBgpCiscoDnsDockerF5GCPGitlab Ci/CdGrafanaJuniperKubernetesLinuxNginxPalo AltoPrometheusPythonRadwareSplunkTcp/IpTerraformTlsVpns
3 Hours Ago
In-Office
Los Angeles, CA, USA
104K-218K Annually
Senior level
104K-218K Annually
Senior level
Information Technology • Consulting • Defense
Lead enterprise-level network support and modernization for U.S. Air Force infrastructure, troubleshooting network issues, and overseeing complex engineering efforts.
Top Skills: BgpCisco RoutersDmvpnEigrpExcelFirewallsMs Office Suite (WordNetwork Automation TechnologiesPowerPointSwitchesVisio)
6 Days Ago
In-Office or Remote
2 Locations
136K-265K Annually
Senior level
136K-265K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
As a Senior Network Operations Engineer, you'll manage and maintain NVIDIA's cloud and datacenter networks, responding to incidents, improving operations, and collaborating with teams and vendors.
Top Skills: AristaAWSAzureBgpDnsEvpnFortinetGCPGrafanaGreIpsecIs-IsJuniperMacsecMplsNautobotNetboxOciOspfPanoptesPrometheusPythonQosShellTcp/IpUnix/LinuxVxlan

What you need to know about the Los Angeles Tech Scene

Los Angeles is a global leader in entertainment, so it’s no surprise that many of the biggest players in streaming, digital media and game development call the city home. But the city boasts plenty of non-entertainment innovation as well, with tech companies spanning verticals like AI, fintech, e-commerce and biotech. With major universities like Caltech, UCLA, USC and the nearby UC Irvine, the city has a steady supply of top-flight tech and engineering talent — not counting the graduates flocking to Los Angeles from across the world to enjoy its beaches, culture and year-round temperate climate.

Key Facts About Los Angeles Tech

  • Number of Tech Workers: 375,800; 5.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Snap, Netflix, SpaceX, Disney, Google
  • Key Industries: Artificial intelligence, adtech, media, software, game development
  • Funding Landscape: $11.6 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Strong Ventures, Fifth Wall, Upfront Ventures, Mucker Capital, Kittyhawk Ventures
  • Research Centers and Universities: California Institute of Technology, UCLA, University of Southern California, UC Irvine, Pepperdine, California Institute for Immunology and Immunotherapy, Center for Quantum Science and Engineering

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account