Top SRE Engineer Jobs in Los Angeles, CA

5 Days AgoSaved
Easy Apply
Remote
Los Angeles, CA
Easy Apply
191K-226K Annually
Senior level
191K-226K Annually
Senior level
Big Data • Healthtech • HR Tech • Machine Learning • Software • Telehealth • Big Data Analytics
Own the reliability, performance, resilience, observability, and security of AWS and Kubernetes infrastructure supporting products and AI/ML workloads. Define SLOs, lead incident response and root-cause analysis, build Terraform automation, optimize cloud costs, reduce operational toil, and establish deployment standards that help engineers ship reliably. Participate in on-call rotations and maintain HIPAA-compliant infrastructure.
Top Skills: AWSClaudeDatadogGitlabGoHipaaIstioKubernetesNatsPostgresPythonSoc 2TerraformTypescript
Reposted 15 Days AgoSaved
Remote
Los Angeles, CA
180K-220K Annually
Senior level
180K-220K Annually
Senior level
Software • Defense
Work as an SRE embedded with product teams to improve reliability by fixing application code (primarily TypeScript), building observability (Prometheus, Loki, Grafana, Alloy), defining SLIs/SLOs, leading incident response and postmortems, automating toil, and supporting deployments across on‑prem DoD and AWS environments.
Top Skills: AlloyAWSBashContainersDockerGithub ActionsGitlab Ci/CdGoGrafanaJenkinsKubectlKubernetesLokiNode.jsPrometheusPythonTypescript
Reposted 16 Days AgoSaved
Easy Apply
Remote or Hybrid
Los Angeles, CA
Easy Apply
127K-249K Annually
Senior level
127K-249K Annually
Senior level
Big Data • Cloud • Software • Database
As a Senior Site Reliability Engineer, you'll design and build complex systems, support Atlas platform operations, automate processes, and ensure high availability of services.
Top Skills: AWSAzureDnsGCPGoHTTPLinuxPythonRubyTls
Reposted 27 Days AgoSaved
Easy Apply
Hybrid
Los Angeles, CA
Easy Apply
170K-190K Annually
Senior level
170K-190K Annually
Senior level
AdTech • Big Data • Cloud • Marketing Tech • Software • Analytics
Lead SRE efforts to improve security, reliability, cost efficiency, and observability. Build automation, CI/CD, agentic AI platforms (MCPs), and tooling for capacity planning, incident response, and self-service. Evangelize SecDevOps and zero-trust designs across product and platform teams.
Top Skills: Ai Agentic FrameworksArgocdAWSBashCi/CdDevsecopsDockerEksGoGrafanaKubernetesLinuxLokiMcpNew RelicPrometheusPythonSamTerraformZero Trust
Reposted 18 Days AgoSaved
Easy Apply
Remote or Hybrid
Los Angeles, CA
Easy Apply
127K-249K Annually
Senior level
127K-249K Annually
Senior level
Big Data • Cloud • Software • Database
Develop and maintain Kubernetes runtime environments, support developers, resolve critical issues, and participate in on-call rotations for production systems.
Top Skills: AWSAzureCert-ManagerCorednsCrdsCriCsiGatekeeperGCPGoHelmKubernetesKustomizeOperatorsPythonTerraform
Reposted 20 Days AgoSaved
In-Office or Remote
Los Angeles, CA
161K-284K Annually
Senior level
161K-284K Annually
Senior level
Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
The Senior Site Reliability Engineer will enhance reliability of Block's platform, improve incident response using AI tools, and coordinate incident management. Responsibilities include building reliable systems, standardizing tools, and leading high-severity incidents during on-call rotations.
Top Skills: Amazon Web ServicesDatadogDynamoDBGrpcHTTPIstioJavaJSONKotlinKubernetesLaunchdarklyMySQLProtocol BuffersTerraformVitess
Reposted 20 Days AgoSaved
Remote or Hybrid
Los Angeles, CA
Senior level
Senior level
Fintech • Software
Lead SRE efforts for DFIN SaaS: ensure availability, performance, scalability, and automation. Implement monitoring, CI/CD, IaC, container orchestration, AI-enhanced observability, incident response, RCA, and runbook automation while collaborating across engineering teams.
Top Skills: .NetAiopsAksAnsibleAppdynamicsAWSAzureAzure DevopsBashC#Ci/CdCloud Ai ServicesContainersCosmosDatadogDynatraceEksFirewallHarnessIdera Sql Diagnostic ManagerInfrastructure As Code (Iac)JavaJenkinsKubernetesLinuxLoad BalancingNew RelicPowershellPythonRedgate Sql MonitorSolarwinds Database Performance AnalyzerSQLTerraformWindows
2 Days AgoSaved
In-Office
Los Angeles, CA
181K-265K Annually
Senior level
181K-265K Annually
Senior level
Fintech • Professional Services • Software
Own platform reliability, performance, observability, and resilience for scalable backend systems. Mentor engineering teams on secure, performant coding practices; establish OpenTelemetry observability and service-level objectives; design and improve incident response processes; and implement automated testing. Collaborate with engineering leaders, product managers, and cross-functional teams on shared infrastructure and challenging technical solutions. The role also contributes to cloud infrastructure, CI/CD, tabletop exercises, and modern distributed systems across the organization.
Top Skills: AWSClaudeDebeziumDockerGitlab Ci/CdHelmJavaKafkaKubernetesMySQLOpentelemetryPostgresRestful ApisSnowflakeSpring BootSQLTemporalTerraform
2 Days AgoSaved
In-Office
Los Angeles, CA
181K-225K Annually
Senior level
181K-225K Annually
Senior level
Fintech • Professional Services • Software
Build and improve resilient, scalable backend systems as part of the Platform team. Responsibilities include implementing observability with OpenTelemetry, defining SLOs, owning incident response processes and follow-ups, mentoring teams on secure and reliable coding, and applying automated testing. The role involves working with Java, Spring Boot, microservices, distributed systems, Kubernetes, cloud infrastructure, and CI/CD while collaborating with engineering leadership and cross-functional teams.
Top Skills: AWSClaudeDebeziumDistributed SystemsDockerGitlab Ci/CdHelmJavaKafkaKubernetesMicroservicesMySQLOpentelemetryPostgresRestful ApisSnowflakeSpring BootSQLTemporalTerraform
Reposted 22 Days AgoSaved
Easy Apply
Remote
Los Angeles, CA
Easy Apply
100K-110K Annually
Mid level
100K-110K Annually
Mid level
Healthtech • Software
Operate and maintain AWS-hosted MERN applications and large-scale data workflows. Manage serverless and Spark-based pipelines, perform incident response and on-call duties, engineer automation to eliminate operational toil, ensure HIPAA/SOC2/HITRUST compliance, build observability and lead blameless post-mortems.
Top Skills: Amazon EcsAmazon EksAmazon EmrAthenaAws GlueAws LambdaAws SnsAws SqsCloudwatchEc2IamJavaScriptMernMySQLNode.jsOpentofuPysparkPythonRabbitMQTerraformTypescriptVpc
23 Days AgoSaved
Remote or Hybrid
Los Angeles, CA
140K-215K Annually
Senior level
140K-215K Annually
Senior level
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Senior SRE owning availability, automation, and observability for CI/CD platform services. Build and operate infrastructure, run on-call, lead incident response, mentor engineers, drive design/capacity planning, integrate AI-assisted workflows, and improve cross-team reliability.
Top Skills: Active DirectoryAnsibleApache AirflowSparkAWSAzureBashBazelBitbucketCassandraChefDatadogDnsFirewall RulesGCPGitGithub ActionsGitlabGitlab CiGoGrafanaHoneycombHumio/LogscaleJenkinsKafkaKubernetesLoad BalancersMongoDBMySQLNasNew RelicNfsObject StorageOpensearchOraclePostgresPowershellPrometheusPulsarPuppetPythonRabbitMQRedis/ValkeyRedpandaRoutingSaltSanSplunkTerraformVarnishVipsWindows Server
5 Days AgoSaved
In-Office
Los Angeles, CA
108K-162K Annually
Senior level
108K-162K Annually
Senior level
Fintech • Financial Services
Operates and improves high-traffic, business-critical cloud and network systems. Designs scalable network solutions, manages performance and troubleshooting, automates deployment and monitoring, coordinates router and switch installations, and tests redundancy, resilience, and failover. Partners with development teams to improve operability, supports project planning and vendor comparisons, provides training, and participates in on-call coverage.
Top Skills: Cloud ComputingIpRoutersSwitchesVoip
New

Track Smarter, Apply Better.

Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.

Use For Free
Application Tracker Preview
Reposted 24 Days AgoSaved
Easy Apply
Remote or Hybrid
Los Angeles, CA
Easy Apply
126K-248K Annually
Senior level
126K-248K Annually
Senior level
Big Data • Cloud • Software • Database
The Senior Site Reliability Engineer will develop and support distributed storage services, ensuring reliability and operational safety, with a focus on automation and efficiency.
Top Skills: AWSAzureDnsGoGoogle Cloud PlatformKubernetesLinuxPythonTcp/IpTls
7 Days AgoSaved
In-Office
Los Angeles, CA
155K-195K Annually
Senior level
155K-195K Annually
Senior level
Aerospace
Design, build, and operate highly available ground and site network infrastructure across commercial and secured environments. Architect Kubernetes platforms, enforce security and compliance controls, develop Terraform-based infrastructure as code, and establish monitoring, logging, alerting, CI/CD, and automation. Serve as a site reliability owner for mission-critical systems, driving uptime and incident response. Partner with security and mission operations teams, define architectural standards, and mentor engineers.
Top Skills: ArgocdCi/CdDockerGitopsGoGrafanaInfrastructure As CodeKubernetesMimirNetwork SegmentationNetworkingObservabilityPrometheusPythonRoutingSwitchingTerraformTerragrunt
Reposted 26 Days AgoSaved
Remote or Hybrid
Los Angeles, CA
175K-200K Annually
Senior level
175K-200K Annually
Senior level
eCommerce • Fintech • Payments • Software
The role involves ensuring software reliability and performance, managing incidents, developing infrastructure automation, and mentoring junior engineers within a platform team.
Top Skills: AWSCloudFormationDatadogKubernetesOpentelemetryRubyRuby On RailsTerraform
Reposted YesterdaySaved
Remote
Los Angeles, CA
Internship
Internship
Artificial Intelligence • Cloud • Information Technology • Consulting
Internship SRE role responsible for availability, performance, and scalability of an e-commerce supply-chain platform. Tasks include SLO/SLA definition, observability (Prometheus/Grafana/Loki/Tempo/OpenTelemetry), incident response, capacity planning, disaster recovery for PostgreSQL, infrastructure-as-code (Terraform), CI/CD automation, and operational reliability for AI agent services. Mentored by Head of Technology/CTO with potential conversion to full-time based on performance.
Top Skills: BashCi/CdDockerGrafanaLangchainLlmLokiMakefileNestjsOpentelemetryOracle CloudPgbackrestPostgresql 15PrometheusPythonRedisTempoTerraformTraefik
11 Days AgoSaved
In-Office
Los Angeles, CA
180K-200K Annually
Senior level
180K-200K Annually
Senior level
Digital Media • Software
Deploy, operate, validate, and improve Ateme video delivery platforms in customer environments, including critical live production systems. Responsibilities include Linux administration, networking, video technology integration, troubleshooting, performance investigations, documentation, customer training, technical collaboration with R&D, and ensuring service availability. The role supports live production events and participates in a 24/7 on-call rotation.
Top Skills: ArqAutomationAv1AvcCdnCloud InfrastructureCmafDistributed SystemsDrmHevcHlsJpeg-XsKubernetesLinuxMpeg-TsNetworkingObservabilityScte-104Scte-224Scte-30Scte-35Smpte 2110UnixVideo Encoding
4 Days AgoSaved
Remote
Los Angeles, CA
Senior level
Senior level
Edtech • Kids + Family • Sports
Audit infrastructure, deployment pipelines, monitoring, alerting, incident response, on-call practices, and internal tools. Produce actionable audit reports, implement code and configuration fixes, improve SLOs and reliability practices, advise on scalable architecture, and partner with engineers on implementation and handoff. The role is a fully remote, part-time consulting engagement with potential for full-time conversion.
Top Skills: Ai Coding ToolsAWSCi/CdDatadogGrafanaPrometheus
4 Days AgoSaved
Remote
Los Angeles, CA
Senior level
Senior level
Cloud • Information Technology • Cybersecurity • Infrastructure as a Service (IaaS)
Owns reliability, observability, and incident response for a GPUaaS platform. Defines SLOs, builds monitoring and alerting systems, leads major incidents and post-incident reviews, automates operational processes, maintains runbooks, manages on-call operations, coordinates with engineering teams, drives chaos testing, reports SLA performance, and mentors junior engineers.
Top Skills: DatadogGoGpuaasGrafanaGremlinHpcKubernetesLitmusOpentelemetryPrometheusPython
Reposted 13 Days AgoSaved
In-Office or Remote
Los Angeles, CA
164K-270K Annually
Mid level
164K-270K Annually
Mid level
Aerospace • Hardware • Software • Defense • Manufacturing
Build scalable automated solutions for device fleet management, own and optimize MDM platforms, write OS-level scripts for self-healing, gather telemetry to prevent end-user disruption, translate compliance (CMMC) into code-managed baselines, and create dashboards and alerts measuring end-user SLOs.
Top Skills: AnsibleBashChefFleet DmIntuneJAMFOsqueryPowershellPulumiPuppetPythonSaltTerraformWorkspace One
Reposted 5 Days AgoSaved
Remote or Hybrid
Los Angeles, CA
138K-221K Annually
Senior level
138K-221K Annually
Senior level
Artificial Intelligence • Cloud • Fintech • Machine Learning • Mobile • Software
Lead design, development, deployment, and scaling of cloud infrastructure and SRE tooling. Build automation, CI/CD, observability, capacity planning, and reliability improvements; collaborate with product teams to define non-functional requirements and resolve production issues.
Top Skills: .NetApi GatewayAWSAzureC#Data LakehouseDatabricks DeltaDatadogElasticsearchElkEvent HubsFunctions/ServerlessGitGrafanaJavaJenkinsKafkaKibanaKubernetesLogstashPowershellSnowflakeSqsTeamcityVisual Basic
5 Days AgoSaved
Remote
Los Angeles, CA
Senior level
Senior level
Healthtech • Software
Manage and optimize a multi-account AWS environment supporting production healthcare applications and analytics platforms. Responsibilities include AWS infrastructure administration, CI/CD and infrastructure automation, observability, incident response, on-call support, disaster recovery validation, HIPAA/HiTrust compliance, IAM and security controls, and support for containerized Java and Python applications. The role also contributes to Kubernetes and EKS modernization initiatives and partners with developers to resolve complex production issues.
Top Skills: Amazon EksApi GatewayArgocdAuroraAWSAws CloudformationAws CodepipelineAws Security HubBashCloudfrontCloudwatchCortex CloudDatadogDockerEc2EcsFargateGitGuarddutyHelmIamJavaJenkinsKubernetesLambdaLinuxPrismaPythonRdsS3Spring BootUbuntuZabbix
Reposted 5 Days AgoSaved
Remote
Los Angeles, CA
Senior level
Senior level
Software
Owns reliability, observability, performance, and security for a multi-region SaaS platform. Responsibilities include managing Datadog, implementing APM and tracing, defining SLOs, developing automation, expanding infrastructure as code and CI/CD, automating operational workflows, maintaining security controls, participating in incident response, documenting procedures, and mentoring engineers.
Top Skills: ApmAzure DevopsAzure Kubernetes ServiceAzure SqlBashBicepCi/CdCosmos DbDatadogDistributed TracingHelmInfrastructure As CodeKey VaultKubernetesKustomizeManaged IdentitiesAzureMicrosoft Entra IdPowershellPythonRedisService BusTerraform
Reposted 14 Days AgoSaved
In-Office
Los Angeles, CA
72K-119K Annually
Junior
72K-119K Annually
Junior
Information Technology • Legal Tech • Analytics
Support and automate cloud infrastructure (primarily Azure) for insurance applications. Implement IaC, assist deployments, troubleshoot incidents, improve security controls, and maintain monitoring and documentation while collaborating with engineers and non-technical stakeholders.
Top Skills: Arm TemplatesAWSAzureBashCi/CdContainerizationDnsGCPGitGitMonitoring And Logging ToolsPowershellPythonServerless FunctionsTerraformVnets
Reposted One Month AgoSaved
Remote
Los Angeles, CA
150K-200K Annually
Senior level
150K-200K Annually
Senior level
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Ensure stability and resilience of Runpod's distributed AI platform by defining SLIs/SLOs, leading incident response, building observability and reliability tooling, automating operational workflows, and partnering with engineering teams to reduce toil and improve production readiness.
Top Skills: BashCi/CdContainerized Production SystemsGoGpu Observability ToolingGrafanaInfrastructure As CodeLinuxPrometheusPython
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account