Top Tech Jobs & Startup Jobs in Los Angeles, CA

2 Days AgoSaved
Remote
USA
Senior level
Senior level
Software
Design, deploy, and maintain high-performance HPC network fabrics (InfiniBand/RoCE), troubleshoot complex latency/throughput issues, manage Fortinet security solutions, tune and monitor network performance, collaborate with compute/storage teams, document architectures, participate in on-call rotations, and support large-scale GPU/AI clusters and incident triage.
Top Skills: AnsibleBashBgpFortianalyzerFortigateFortimanagerHcasInfinibandLinuxMellanox/NvidiaMpiNvidia Gpu ClustersOspfPythonQosRdmaRoceSubnet ManagerTcp/IpTerraformVlanVpn
7 Days AgoSaved
Remote
USA
Senior level
Senior level
Software
Design and implement control-plane services, CSI drivers, and tooling to provision and operate high-performance storage for Kubernetes-based GPU/AI workloads. Integrate enterprise and scale-out storage (e.g., PowerScale, VAST), deliver IaC and GitOps pipelines, build observability for data-path performance, and support hybrid, bare-metal, and air-gapped deployments.
Top Skills: ArgocdCluster Api (Capi)CsiDell PowerscaleFluxGitopsGoHarborK0RdentK0SKubernetesOpentofuPersistent VolumesPki/TlsStorage ClassesTerraformVast
Mid level
Software
Design and build infrastructure services and APIs that manage bare-metal GPU servers and multi-tenant Kubernetes clusters. Implement provisioning workflows, reconciliation loops, and consoles to visualize inventory and provisioning state while ensuring reliability, idempotency, and durable long-running operation handling.
Top Skills: Api GatewayArgocdBmcCluster ApiDpuFluxGitopsGoGpuGrpcIpxeK0RdentKubernetesMetal3NicOpentofuPxeRedfishRestTemporalTerraform
7 Days AgoSaved
Remote
USA
Senior level
Senior level
Software
Operate and automate high-performance NFS/NAS storage for GPU-accelerated Kubernetes clusters. Integrate storage via CSI, tune NFS/RDMA data paths, provision bare-metal hosts, support air-gapped deployments, automate with Terraform/OpenTofu and GitOps (ArgoCD/Flux), build observability, and diagnose performance and reliability at scale.
Top Skills: ArgocdBashCephCluster ApiCsiDell PowerscaleFluxGitopsGoGpudirect StorageHarborK0RdentK0SKernelKubernetesLinuxMkeNasNconnectNfsOpentofuPersistent VolumesPki/TlsPythonRdmaRoceS3Storage ClassesTerraformVast
9 Days AgoSaved
Remote
USA
Senior level
Senior level
Software
Own the networking vision and roadmap for k0rdent AI, defining underlay/overlay fabrics, RDMA, DNS/IPAM, network automation, and DPU/SmartNIC integration. Translate customer and engineering requirements into priorities, track emerging interconnect standards, and support positioning, field engagements, and reference architectures.
Top Skills: Amd PensandoBgpDcqcnDnsDpuEcnEvpnFabric TelemetryGpudirect RdmaInfinibandIpamKubernetes NetworkingLinux NetworkingNcclNetwork AutomationNvidiaOverlay NetworkingOvnOvsPfcRcclRdmaRocev2Scale-Up EthernetSdnSmartnicSr-IovUalinkUltra Ethernet (Uec)VrfVxlan
9 Days AgoSaved
Remote
USA
Senior level
Senior level
Software
Design, deploy, manage, and troubleshoot high-performance InfiniBand and Ethernet networks for HPC. Perform performance tuning, capacity planning, and monitoring. Implement Fortinet security, resolve routing/switching/latency issues, collaborate with compute and storage teams, document architectures, and participate in on-call escalation and upgrades.
Top Skills: BashBgpFortianalyzerFortigateFortimanagerInfiniband (Mellanox/Nvidia)LinuxOspfPythonQosTcp/IpVlans
Senior level
Software
Lead operations for large-scale AI infrastructure: manage NVIDIA GPU and high-performance networking environments, troubleshoot Linux/Kubernetes/storage/hardware issues, drive incident response and root cause analysis, improve observability and automation, mentor engineers, and collaborate with engineering and datacenter teams to ensure platform reliability and scalability.
Top Skills: ElkGrafanaInfinibandInfrastructure-As-CodeK0RdentKubernetesLinuxNvidia GpusNvidia UfmOpentelemetryPrometheus
Mid level
Software
Develop positioning and messaging for Mirantis AI infrastructure products; translate technical capabilities into audience-specific content; own competitive intelligence; create technical assets (white papers, briefs, reference architectures, sales enablement); support GTM, launches, analyst/media briefings, and industry representation.
Top Skills: AnyscaleAWSAzureCloud-NativeCncfContainer OrchestrationGCPGpu Cluster OrchestrationGpu SchedulingHashicorpHugging FaceKubeflowKubernetesLlm InferenceMlflowMlopsMlops PipelinesModalNvidia Base CommandRancherRayRed Hat OpenshiftSuseVmware Tanzu
Mid level
Software
Operate and support production AI infrastructure across global datacenters, focusing on high-performance NVIDIA GPU platforms, Kubernetes, and high-speed networking. Monitor and troubleshoot infrastructure, network, hardware, and platform incidents; participate in incident response and root cause analysis; improve observability, automation, and operational runbooks.
Top Skills: KubernetesLinux
New

Track Smarter, Apply Better.

Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.

Use For Free
Application Tracker Preview
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account