Maximum of 25 job preferences reached.
Top SRE Engineer Jobs in Los Angeles, CA
Cloud • Software
The Site Reliability Engineer will ensure reliable cloud operations by applying Python for infrastructure automation, managing OpenStack and Kubernetes, and practicing devsecops in a fast-paced environment.
Top Skills:
KubernetesLinuxOpenstackPython
Cloud • Security • Software • Cybersecurity
As a Site Reliability Engineer II, you'll automate tasks, monitor AI workloads, enhance dashboards, support CI/CD processes, and collaborate with engineering teams on complex issues while participating in on-call rotations.
Top Skills:
GoGrafanaKubernetesLinuxPrometheusPythonSaltstackTerraform
Software • Cybersecurity
This role involves managing Kubernetes clusters, cloud infrastructure, and CI/CD pipelines. The engineer will enhance system reliability and efficiency while troubleshooting production issues.
Top Skills:
AlertmanagerAWSAzureBashCi/CdDockerElastic StackElasticsearchGCPGoGrafanaHelmKafkaKubernetesLokiMongoDBOciPrometheusPythonRedisSparkTerraform
Software • Web3
Lead reliability practices across teams: embed early in projects, define SLIs/SLOs, build multi-cloud paved roads with Terraform, run on-call, drive org-wide incident maturity and tooling.
Top Skills:
AWSAzureGCPRuby On RailsTerraformTypescriptWebcontainers
Aerospace • Hardware • Software • Defense • Manufacturing
As a Site Reliability Engineer, you'll ensure robotics system reliability, build telemetry integration, and develop tools for diagnostics and automation, collaborating with engineering teams for enhanced production reliability.
Top Skills:
C++DatadogGoKubernetesOpentelemetryPrometheusPythonRos2TelegrafTypescript
Events
The Site Reliability Engineer II designs and maintains scalable systems, focusing on automation, monitoring, incident response, and collaboration with developers to enhance operational practices and efficiency.
Top Skills:
BashCloud Service OperationsContainersContinuous DeliveryContinuous IntegrationGoInfrastructure As CodeOrchestration PlatformsPython
Fitness • Healthtech • Software
Own reliability and security of CI/CD and production services: define SLI/SLOs, lead incident response and postmortems, build observability (Datadog), operate Kubernetes and IaC (Terraform), harden pipelines with SAST/DAST/SCA and policy-as-code, and coach teams on reliability and operational best practices.
Top Skills:
AWSCi/CdConftestDastDatadogDockerGithub ActionsGoInfrastructure As CodeKubernetesKyvernoOpa/RegoPythonSastScaTerraformTypescript
Automotive
Leads SRE engineering leaders and engineers while defining enterprise observability, reliability, and platform strategy across GCP, on-premise, manufacturing, distribution, and campus environments. Oversees vendor-agnostic tooling, OpenTelemetry integrations, CI/CD observability, SRE maturity models, and Agentic AI initiatives. Drives adoption of SRE practices, develops technical roadmaps, partners with senior leadership and operational teams, and maintains hands-on architectural and technical credibility.
Top Skills:
Agentic AiAWSAzureCi/CdDatadogDynatraceGCPNew RelicOpentelemetryOtel Genai Semantic ConventionsSource Control PlatformsSplunkTerraform
Other
Design, build, and maintain highly available cloud-native systems. Improve reliability through automation, CI/CD, Kubernetes, observability, and incident management. Collaborate with developers, security, and product teams to define SLOs, implement self-healing, debug production issues, and ensure secure deployments.
Top Skills:
AWSAzure Cloud ServicesDatadogGCPGithub ActionsGitlab CiGoInfrastructure As CodeKubernetesOpsgeniePagerdutyPythonRubySite Reliability Engineering Foundation
Artificial Intelligence • Cloud • Information Technology • Software
Design and operate large-scale GPU infrastructure for distributed AI training, ensuring reliability, performance, and efficient customer partnerships.
Top Skills:
AnsibleCudaDeepspeedFsdpGpuHelmInfinibandKubernetesLinuxMegatronNcclNvidia A100Nvidia B200Nvidia H100NvlinkPyTorchRoceTerraform
Cloud • Security • Software • Cybersecurity
Design, build, and operate scalable infrastructure and CI/CD/IaC systems. Implement observability (monitoring, logging, alerting), automate reliability improvements, mentor engineers, collaborate on incident response, and participate in on-call rotations to maintain Akamai Cloud services.
Top Skills:
AlertingAnsibleBashChefCi/CdGithub ActionsGitlab Ci/CdGoInfrastructure As CodeJenkinsLoggingMonitoringPuppetPythonSaltstackTelemetryTerraform
Cloud • Security • Software • Cybersecurity
Design, develop, test, and operate scalable infrastructure and services for Akamai Cloud. Implement and manage Infrastructure-as-Code (Terraform and similar tools), CI/CD, and observability. Automate reliability improvements, mentor engineers, collaborate on incident response and root-cause remediation, and participate in on-call rotations.
Top Skills:
Alerting)AnsibleChefCi/CdInfrastructure As CodeLinuxLoggingObservability (MonitoringPuppetSaltstackTerraform
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
Artificial Intelligence • Information Technology • Consulting
Build and operate Nebius's network infrastructure: define SLIs/SLOs, improve site and inter-site reliability, lead incident response and postmortems, develop observability and alerting, automate change workflows, and collaborate with network and platform teams to embed operability.
Top Skills:
Ci/CdContainer PlatformsGoInfrastructure As CodeLinuxPython
Aerospace • Other
The Site Reliability Engineer, GNC at SpaceX oversees mission-critical GNC products, operates servers, maintains HPC clusters, and enhances services and infrastructure to support space operations.
Top Skills:
AnsibleBazelDockerGradleKubernetesLinuxMakeNpmPipPuppetPythonTerraformVagrant
Aerospace • Other
Manage and design server, HPC, storage, and networking infrastructure (including InfiniBand) to support propulsion engineering workflows. Integrate and optimize engineering applications (ANSYS, StarCCM+), automate deployments, troubleshoot performance bottlenecks, and coordinate with IT, facilities, and engineering teams to scale compute resources for rocket engine development.
Top Skills:
AnsibleAnsysBashDockerEnterprise NetworkingHpcInfinibandKubernetesLinuxPuppetPythonStarccm+VirtualizationWindows Server
Fintech • Financial Services
Lead Site Reliability Engineer responsible for ensuring system reliability, scalability, and performance. Develop automated deployment strategies, maintain monitoring/observability, define SLIs/SLOs, collaborate with cross-functional teams, drive reliability best practices, participate in on-call incident response, and improve delivery through automation and training.
Top Skills:
AWSAzureBashGCPPowershellPython
Software
The role involves managing compute infrastructure for decentralized applications, requiring critical thinking, documentation skills, and experience in Kubernetes and blockchain management.
Top Skills:
BlockchainGitopsInfrastructure-As-CodeKubernetesProgramming Languages
Artificial Intelligence • Information Technology • Software • Database
As a Site Reliability Engineer, you will design, implement, and maintain scalable infrastructure, ensure system reliability, automate processes, and collaborate with engineering teams.
Top Skills:
DockerElk StackGoGrafanaJavaKubernetesNode.jsPrometheusPulumiPythonRubyTerraform
Reposted 20 Days AgoSaved
Other • Social Impact
As a Senior Site Reliability Engineer, you will design, develop, and maintain reliable infrastructure for Wikimedia's API services, ensuring performance and availability while driving reliability engineering practices and improving developer experience.
Top Skills:
AnsibleArgocdAWSAzureGCPGitlabGoKubernetesOpentelemetryPrometheusPythonTerraform
Angel or VC Firm • Blockchain • Fintech • Cryptocurrency
Apply to join Galaxy Ventures' invite-only Talent Network for DevOps, SRE, QA, and Security professionals. Upon acceptance, your profile may be discreetly shared with portfolio companies for relevant roles, and you'll receive invitations to exclusive networking events. Participation is confidential and non-binding.
Aerospace • Other
Design, operate, and scale on-premise infrastructure for the Starshield satellite constellation. Build automation for Kubernetes cluster deployment and management, operate core infrastructure (databases, monitoring, distributed storage), collaborate with software teams, troubleshoot across the stack, improve service lifecycle, and ensure high availability through monitoring and performance improvements.
Top Skills:
AnsibleBashC++GoKubernetesLinuxOci ContainersPythonTerraform
Aerospace • Other
Design, deploy, and automate on‑prem and cloud compute infrastructure; manage core infrastructure (databases, monitoring, storage); collaborate with software teams to build scalable, operable systems; improve service lifecycle from design through deployment, operation, and refinement.
Top Skills:
AnsibleBashBazelDatabasesKubernetesLinuxMakeMakefilesMonitoringPythonStorageTcp/IpTerraform
Healthtech • Software
Design, automate, and maintain scalable infrastructure and SRE tooling. Manage Kubernetes clusters, CI/CD, monitoring, and incident response. Improve processes, reduce toil via automation, and collaborate with engineering and data teams to support domestic and international workloads.
Top Skills:
AWSAzureContainerdDnsDockerFirewallsGCPGoGrpcHelmKubernetesLinuxLoad BalancingPrometheusPythonRoutingShell ScriptingTcp/IpUdp
Artificial Intelligence • Insurance • Software • Automation
The Staff Site Reliability Engineer will build and scale infrastructure for Assured's platform, automate delivery, enhance observability, and lead mentoring initiatives.
Top Skills:
AWSKubernetesPostgresTerraform
Healthtech • Social Impact • Software
Own the operational lifecycle of cloud-native data infrastructure: design and automate reliable deployments, observability, incident response, SLIs/SLOs, autoscaling and IaC, and improve platform efficiency and data freshness across GKE and Cloud Run.
Top Skills:
BashBigQueryCloud BuildCloud MonitoringCloud RunDatadogDockerGCPGithub ActionsGkeGoGrafanaJIRAKubernetesPrometheusPulumiPythonSentrySlackSnykSonarqubeTerraform
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Los Angeles, CA Companies Hiring SRE Engineers
See AllPopular Los Angeles, CA Engineering Job Searches
Engineering Jobs in Los Angeles, CA
Software Engineer Jobs in Los Angeles, CA
Android Developer Jobs in Los Angeles, CA
C# Jobs in Los Angeles, CA
C++ Jobs in Los Angeles, CA
DevOps Jobs in Los Angeles, CA
Front End Developer Jobs in Los Angeles, CA
Golang Jobs in Los Angeles, CA
Hardware Engineer Jobs in Los Angeles, CA
iOS Developer Jobs in Los Angeles, CA
Java Developer Jobs in Los Angeles, CA
Javascript Jobs in Los Angeles, CA
Linux Jobs in Los Angeles, CA
Engineering Manager Jobs in Los Angeles, CA
.NET Developer Jobs in Los Angeles, CA
PHP Developer Jobs in Los Angeles, CA
Python Jobs in Los Angeles, CA
QA Jobs in Los Angeles, CA
Ruby Jobs in Los Angeles, CA
Salesforce Developer Jobs in Los Angeles, CA
Scala Jobs in Los Angeles, CA
Application Engineer Jobs in Los Angeles, CA
Associate Software Engineer Jobs in Los Angeles, CA
Automation Engineer Jobs in Los Angeles, CA
AWS Engineer Jobs in Los Angeles, CA
Backend Engineer Jobs in Los Angeles, CA
Cloud Engineer Jobs in Los Angeles, CA
Controls Engineer Jobs in Los Angeles, CA
CTO Jobs in Los Angeles, CA
Design Engineer Jobs in Los Angeles, CA
DevOps Engineer Jobs in Los Angeles, CA
Director of Engineering Jobs in Los Angeles, CA
Electrical Engineering Jobs in Los Angeles, CA
Embedded Software Engineer Jobs in Los Angeles, CA
Field Engineer Jobs in Los Angeles, CA
Firmware Engineer Jobs in Los Angeles, CA
Full-Stack Engineer Jobs in Los Angeles, CA
Game Engineer Jobs in Los Angeles, CA
Industrial Engineer Jobs in Los Angeles, CA
Infrastructure Engineer Jobs in Los Angeles, CA
Manufacturing Engineer Jobs in Los Angeles, CA
Mechanical Design Engineer Jobs in Los Angeles, CA
Mechanical Engineering Jobs in Los Angeles, CA
Mechatronics Engineering Jobs in Los Angeles, CA
Network Engineer Jobs in Los Angeles, CA
Platform Engineer Jobs in Los Angeles, CA
Principal Engineer Jobs in Los Angeles, CA
Principal Software Engineer Jobs in Los Angeles, CA
Process Engineer Jobs in Los Angeles, CA
Product Engineer Jobs in Los Angeles, CA
Project Engineer Jobs in Los Angeles, CA
QA Engineer Jobs in Los Angeles, CA
Reliability Engineer Jobs in Los Angeles, CA
Robotics Engineer Jobs in Los Angeles, CA
Software Architect Jobs in Los Angeles, CA
Software Engineering Manager Jobs in Los Angeles, CA
Solutions Engineer Jobs in Los Angeles, CA
SRE Engineer Jobs in Los Angeles, CA
Staff Engineer Jobs in Los Angeles, CA
Staff Software Engineer Jobs in Los Angeles, CA
Structural Engineer Jobs in Los Angeles, CA
Systems Engineer Jobs in Los Angeles, CA
VP of Engineering Jobs in Los Angeles, CA
Web Developer Jobs in Los Angeles, CA
All Filters
Total selected ()
No Results
No Results




























