Maximum of 25 job preferences reached.
Top SRE Engineer Jobs in Los Angeles, CA
Edtech
Lead infrastructure modernization and platform reliability across multiple cloud providers. Design infrastructure as code, operate Kubernetes and Linux environments, improve CI/CD and deployment tooling, establish SLI/SLO practices, strengthen observability, lead incident response, manage cloud costs, and partner on security and compliance. Provide technical leadership through architecture guidance, mentorship, engineering standards, and roadmap development while participating in on-call support.
Top Skills:
AWSCi/CdGCPJenkinsKubernetesLinuxPythonRubyRuby On RailsSoc 2SpinnakerTerraform
Cloud • Information Technology • Business Intelligence • Consulting
Design, build, and operate cloud infrastructure and SRE capabilities for an enterprise AI platform. Responsibilities include infrastructure-as-code, landing zones, networking, Kubernetes, CI/CD, observability, incident response, SLOs, production readiness, automation, cost optimization, and support for hybrid, edge, air-gapped, and customer-controlled environments. The role is remote, client-facing, and requires strong collaboration and reliability ownership.
Top Skills:
AlertingAzureAzure ArcAzure DevopsBicepCi/CdDashboardsDockerGithub ActionsGpu WorkloadsInfrastructure As CodeKubernetesLogsMetricsObservabilitySlosTerraformTraces
Cloud • Security • Software • Cybersecurity
The Site Reliability Engineer II ensures the reliability, availability, performance, and security of critical cloud systems and services. Responsibilities include developing automation for provisioning and configuration management, maintaining monitoring and alerting, optimizing infrastructure performance, supporting high availability, and enabling continuous integration and delivery. The role collaborates with security teams and drives operational improvements across cloud and network infrastructure.
Top Skills:
AnsibleAWSAzureChefContinuous DeliveryContinuous IntegrationDnsElk StackGCPGoGrafanaHTTPKubernetesLinuxPrometheusPuppetPythonShellTcp/IpUnix
Healthtech • Insurance
Lead cloud, DevOps, and SRE architecture efforts to scale CareSource's digital platform. Partner with Cloud, DevOps, Security, and SRE teams to design infrastructure, build Terraform templates, enhance CI/CD pipelines, implement monitoring/alerting, and improve reliability, scalability, and incident response for enterprise-scale digital products.
Top Skills:
Azure CloudDockerDynatraceGithub ActionsKubernetesSplunkTerraform Enterprise
Aerospace • Other
Design, upgrade, and operate large-scale distributed systems for Starlink. Improve sharding, geo-redundancy, multi-region deployment, monitoring, and performance. Manage petabyte-scale bare-metal clusters and collaborate across teams through the full software lifecycle.
Top Skills:
Apache KafkaSparkC#FlinkGoHbaseHdfsIstioJavaKubernetesLinuxPythonScala
Reposted One Month AgoSaved
Easy Apply
Easy Apply
Cloud • Information Technology • Security • Software • Cybersecurity
This internship role focuses on SRE skills, requiring collaboration and problem-solving in dynamic environments for Zscaler's Zero Trust Exchange team.
Top Skills:
AnsibleAws EcsKubernetesLinuxPythonTerraform
Cloud • Security • Software • Cybersecurity
Build and maintain reliable, scalable cloud compute platforms across distributed services. Troubleshoot Linux, networking, and production issues; develop automation and AI-assisted tooling; improve monitoring, alerting, SLIs, and SLOs; conduct incident response and root cause analysis; and partner with engineering teams on system design, deployment safety, and operational readiness.
Top Skills:
AnsibleDnsDockerElkGoGrafanaKubernetesLinuxLokiNomadOpensearchPodmanPrometheusPythonSaltTcp/IpTerraform
Artificial Intelligence • Healthtech • Software • Telehealth
Designs, deploys, and maintains resilient AWS and Kubernetes infrastructure. Builds automation, GitHub Actions components, internal AI-assisted operational tools, and observability systems. Leads incident response, postmortems, and SLO/SLI management while ensuring HIPAA compliance and high availability. Collaborates across teams on architecture reviews, risk reduction, clinical safety, and reliability best practices.
Top Skills:
Ai-Assisted OperationsAmazon Ec2Amazon EksAmazon RdsAmazon S3AWSBashDatadogGithub ActionsGoHelmKubernetesPythonTerraform
Artificial Intelligence • Healthtech • Machine Learning • Software
Define reliability strategy and technical roadmaps; design scalable cloud infrastructure; improve observability, incident response, disaster recovery, automation, and deployment systems. Partner across engineering, security, data, AI, and product teams to strengthen resilience, compliance, and operational excellence. Lead architecture reviews, resolve complex production issues, establish SRE practices, and mentor engineers while reducing operational toil and improving platform scalability.
Top Skills:
AutomationCi/CdCloud InfrastructureContainer OrchestrationContainersDisaster RecoveryDistributed SystemsInfrastructure As CodeMulti-Region ArchitectureNetworkingObservability
Gaming
Own and operate large-scale infrastructure for sports betting and media platforms across cloud and production environments. Lead infrastructure migrations, build Kubernetes platform tooling and CI/CD automation, improve observability and alerting, support development teams, and participate in incident response. The role requires strong distributed-systems expertise, production troubleshooting, cross-team project leadership, technical communication, and mentoring.
Top Skills:
ArgocdAWSBashCephCiliumDatadogGCPGithub ActionsGoHelmIstioKubernetesLinuxPgbouncerPostgresPythonTalos OsTerraform
Digital Media
Build, maintain, and operate Ookla’s globally distributed infrastructure platform at massive scale. Responsibilities include managing cloud instances, containers, serverless applications, databases, streaming systems, and big-data tooling; supporting 24/7 production operations and on-call rotations; implementing security programs; improving deployment pipelines, monitoring, observability, and reliability; and guiding software and data engineering teams on operational best practices and troubleshooting.
Top Skills:
Amazon AuroraAmazon RdsAnsibleSparkAWSChefCloudFormationDockerDynamoDBGitGitGoIds/IpsJavaKafkaKinesisKubernetesLinuxMongoDBMySQLPHPPostgresPythonRubySQLTerraformTypescript
Aerospace • Manufacturing
Build and lead a centralized observability platform for satellite, ground-station, and distributed network systems. Responsibilities include scaling metrics, logging, and tracing infrastructure; defining SLOs, SLIs, and error budgets; enabling application instrumentation; automating deployments with Terraform and ArgoCD; monitoring Kubernetes, GCP, and AWS environments; and developing incident response, alerting, and reliability practices. The role includes on-call responsibilities and requires an active Top Secret/SCI clearance.
Top Skills:
ArgocdAWSC++ElkGitlab CiGoGoogle Cloud PlatformGrafanaHoneycombIstioJaegerJavaKubernetesLinkerdLokiOpentelemetryPrometheusPythonTempoTerraform
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
Cloud • Security • Software
Design, deploy, and maintain resilient cloud infrastructure for Ping Identity’s mission-critical services. Build and optimize automated CI/CD pipelines, support cloud security and observability, evaluate technologies, participate in planning and on-call rotations, and help improve engineering practices. Collaborate across development and operations teams while sharing expertise and supporting distributed production systems.
Top Skills:
Ci/CdCloud PlatformsDistributed SystemsDockerGitGoIdentity And Access ManagementKubernetesNetworking
Defense • Manufacturing
Design, deploy, and maintain cloud and on-prem infrastructure with IaC; scale and operate mission-critical services; automate tooling and self-service platforms; collaborate with engineering teams to ensure high availability, reliability, security, and resiliency for vehicle and non-vehicle software systems.
Top Skills:
AWSAzureContainer OrchestrationGCPInfrastructure-As-Code (Iac)KubernetesLinuxNetworking
Aerospace • Other
Build, operate, and scale mission-critical application infrastructure and tooling for vehicle and satellite software delivery. Manage infrastructure as code, improve observability, collaborate with engineers, participate in on-call rotation, perform incident response and postmortems, and provide end-user support to reduce build and test times.
Top Skills:
AnsibleBazelBuckC#C++ClickhouseDockerJavaScriptKubernetesKvmLinuxMakeMySQLPostgresPuppetPythonQemuTerraformVsphere
Aerospace • Other
Build, operate, and scale mission-critical application platforms to accelerate vehicle software delivery. Manage infrastructure as code, improve observability, collaborate with developers, run on-call rotations, conduct blameless postmortems, and reduce performance bottlenecks to support Falcon, Starship, Dragon, and Starlink software lifecycles.
Top Skills:
AnsibleBazelBuckC#C++ClickhouseDockerJavaScriptKubernetesKvmLinuxMakeMySQLPostgresPuppetPythonQemuTerraformVsphere
Healthtech • Software
Provides technical leadership for site reliability and operational development. Designs automated, repeatable cloud and infrastructure solutions; monitors service objectives and business metrics; improves application resiliency, performance, efficiency, and cost. Builds monitoring, deployment, testing, and vulnerability-response automation while supporting mission-critical production systems. Collaborates with software, security, product, and business teams, conducts technical training and resilience exercises, and promotes sound development, change-management, and operational practices.
Top Skills:
Amazon EcsAnsibleAzureCC++ChefDockerGoJavaKubernetesLinuxPerlPuppetPythonRubyTerraformWindows
Artificial Intelligence • Software
Architects and owns highly available infrastructure and Kubernetes-based platforms supporting autonomous systems. Builds Golang backend services, platform tooling, observability systems, dashboards, alerts, and log aggregation. Partners with product teams to launch services, performs performance analysis, manages cloud upgrades, and participates in incident response and postmortems. Collaborates on cloud security risk assessments, intrusion detection, threat-feed systems, risk mitigation, and SaaS payment processes. Provides architectural leadership and mentorship across engineering teams.
Top Skills:
ArgocdArgocd Image UpdaterArtifactoryAWSGithub ActionsGoJavaScriptKubernetesPythonRustTerraform
Aerospace • Hardware • Software • Biotech • Pharmaceutical • Manufacturing
Lead design, build, and operate mission-critical infrastructure across cloud, on-prem, and spacecraft contexts. Implement IaC, CI/CD, observability, and scalable Kubernetes-based systems; respond to incidents, perform root cause analysis, optimize performance, and collaborate with software and hardware teams. Participate in on-call rotations and occasional travel.
Top Skills:
AnsibleArgocdAzureBashCi/CdContainerdDatabasesDockerFirewallsGitopsGpu WorkloadsGrafanaHpcInfluxdbKubernetesLinuxPowershellPrometheusPythonSaltSlurmSubnetsTerraformVpcVpns
Information Technology • Security • Cybersecurity
Own production operations and highly available cloud infrastructure for regulated government environments. Lead incident response, root-cause analysis, disaster recovery testing, compliance operationalization, audit readiness, vulnerability management, and continuous monitoring. Build automation, secure CI/CD pipelines, infrastructure-as-code, observability, and compliance tooling across Kubernetes, Linux, containers, and cloud platforms. Partner with security, compliance, and engineering teams to improve reliability, deployment safety, and regulatory sustainment.
Top Skills:
Aws GovcloudBashCi/CdDod Il4Dod Il5FedrampGitopsGoGrafanaKubernetesLinuxNist 800-53PrometheusPythonStigTerraformUnixZero Trust
Aerospace • Manufacturing
As a Site Reliability Engineer, you'll build and manage observability platforms for satellite communications, define SLOs/SLIs, and collaborate on incident response and deployment automation.
Top Skills:
ArgocdAWSElkGCPGoGrafanaIstioJaegerKubernetesLinkerdLokiOpentelemetryPrometheusPythonTempoTerraform
Cybersecurity
Assist with monitoring system performance, uptime, and reliability; support incident response, troubleshooting, and root cause analysis; monitor SRE alerts in Slack and report issues; learn cloud platforms; and collaborate with development and DevOps teams.
Top Skills:
AWSAzureDevOpsDevsecopsGCPGrafanaKubernetesPrometheusSlack
Artificial Intelligence • Machine Learning • Security • Software
The Senior Staff Site Reliability Engineer will be responsible for ensuring system reliability, debugging issues, mentoring the engineering team, and maintaining infrastructure and CI/CD pipelines.
Top Skills:
AWSDatadogDockerGithub ActionsGrafanaHelmKotlinKubernetesPostgresPrometheusPythonRustTerraformTerragruntTypescript
Artificial Intelligence • Fintech • Machine Learning • Natural Language Processing • Business Intelligence
Lead architecture and implementation of reliability platforms and SRE practices for a production SaaS. Build self-service reliability tooling, drive AIOps automation, advance observability (monitoring, tracing, profiling), lead incident response and postmortems, mentor engineers, and embed production readiness across teams to achieve 99.99% uptime.
Top Skills:
AWSAzureContinuous ProfilingDatadogDnsElkGCPGoGrafanaHttp/SKubernetesLoad BalancingOpentelemetryPrometheusPythonTcp/Ip
Artificial Intelligence • Fintech • Machine Learning • Software • App development • Conversational AI • Generative AI
Own and improve production infrastructure reliability, deployments, Infrastructure-as-Code, Kubernetes environments, automation, CI/CD, monitoring, alerting, and observability. Investigate incidents, optimize system performance, maintain documentation and runbooks, and support DNS, WAF, CDN, and caching infrastructure. The role requires strong Linux administration, Bash scripting, networking, Git, and containerization skills, with independent ownership and collaboration across development and operations teams.
Top Skills:
AkamaiAmqpAnsibleAWSBashCdnCloudflareDnsDockerGCPGitGitlab CiGrafanaHttp/HttpsKubernetesLinuxPodmanPrometheusPythonRabbitMQTerraformVictoriametricsWafZabbix
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Los Angeles, CA Companies Hiring SRE Engineers
See AllPopular Los Angeles, CA Engineering Job Searches
Engineering Jobs in Los Angeles, CA
Software Engineer Jobs in Los Angeles, CA
Android Developer Jobs in Los Angeles, CA
C# Jobs in Los Angeles, CA
C++ Jobs in Los Angeles, CA
DevOps Jobs in Los Angeles, CA
Front End Developer Jobs in Los Angeles, CA
Golang Jobs in Los Angeles, CA
Hardware Engineer Jobs in Los Angeles, CA
iOS Developer Jobs in Los Angeles, CA
Java Developer Jobs in Los Angeles, CA
Javascript Jobs in Los Angeles, CA
Linux Jobs in Los Angeles, CA
Engineering Manager Jobs in Los Angeles, CA
.NET Developer Jobs in Los Angeles, CA
PHP Developer Jobs in Los Angeles, CA
Python Jobs in Los Angeles, CA
QA Jobs in Los Angeles, CA
Ruby Jobs in Los Angeles, CA
Salesforce Developer Jobs in Los Angeles, CA
Scala Jobs in Los Angeles, CA
Application Engineer Jobs in Los Angeles, CA
Associate Software Engineer Jobs in Los Angeles, CA
Automation Engineer Jobs in Los Angeles, CA
AWS Engineer Jobs in Los Angeles, CA
Backend Engineer Jobs in Los Angeles, CA
Cloud Engineer Jobs in Los Angeles, CA
Controls Engineer Jobs in Los Angeles, CA
CTO Jobs in Los Angeles, CA
Design Engineer Jobs in Los Angeles, CA
DevOps Engineer Jobs in Los Angeles, CA
Director of Engineering Jobs in Los Angeles, CA
Electrical Engineering Jobs in Los Angeles, CA
Embedded Software Engineer Jobs in Los Angeles, CA
Field Engineer Jobs in Los Angeles, CA
Firmware Engineer Jobs in Los Angeles, CA
Full-Stack Engineer Jobs in Los Angeles, CA
Game Engineer Jobs in Los Angeles, CA
Industrial Engineer Jobs in Los Angeles, CA
Infrastructure Engineer Jobs in Los Angeles, CA
Manufacturing Engineer Jobs in Los Angeles, CA
Mechanical Design Engineer Jobs in Los Angeles, CA
Mechanical Engineering Jobs in Los Angeles, CA
Mechatronics Engineering Jobs in Los Angeles, CA
Network Engineer Jobs in Los Angeles, CA
Platform Engineer Jobs in Los Angeles, CA
Principal Engineer Jobs in Los Angeles, CA
Principal Software Engineer Jobs in Los Angeles, CA
Process Engineer Jobs in Los Angeles, CA
Product Engineer Jobs in Los Angeles, CA
Project Engineer Jobs in Los Angeles, CA
QA Engineer Jobs in Los Angeles, CA
Reliability Engineer Jobs in Los Angeles, CA
Robotics Engineer Jobs in Los Angeles, CA
Software Architect Jobs in Los Angeles, CA
Software Engineering Manager Jobs in Los Angeles, CA
Solutions Engineer Jobs in Los Angeles, CA
SRE Engineer Jobs in Los Angeles, CA
Staff Engineer Jobs in Los Angeles, CA
Staff Software Engineer Jobs in Los Angeles, CA
Structural Engineer Jobs in Los Angeles, CA
Systems Engineer Jobs in Los Angeles, CA
VP of Engineering Jobs in Los Angeles, CA
Web Developer Jobs in Los Angeles, CA
All Filters
Total selected ()
No Results
No Results

















.png)













