Maximum of 25 job preferences reached.
Top SRE Engineer Jobs in Los Angeles, CA
Big Data • Healthtech • HR Tech • Machine Learning • Software • Telehealth • Big Data Analytics
Own the reliability, performance, resilience, observability, and security of AWS and Kubernetes infrastructure supporting products and AI/ML workloads. Define SLOs, lead incident response and root-cause analysis, build Terraform automation, optimize cloud costs, reduce operational toil, and establish deployment standards that help engineers ship reliably. Participate in on-call rotations and maintain HIPAA-compliant infrastructure.
Top Skills:
AWSClaudeDatadogGitlabGoHipaaIstioKubernetesNatsPostgresPythonSoc 2TerraformTypescript
Reposted 15 Days AgoSaved
Software • Defense
Work as an SRE embedded with product teams to improve reliability by fixing application code (primarily TypeScript), building observability (Prometheus, Loki, Grafana, Alloy), defining SLIs/SLOs, leading incident response and postmortems, automating toil, and supporting deployments across on‑prem DoD and AWS environments.
Top Skills:
AlloyAWSBashContainersDockerGithub ActionsGitlab Ci/CdGoGrafanaJenkinsKubectlKubernetesLokiNode.jsPrometheusPythonTypescript
Reposted 16 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
As a Senior Site Reliability Engineer, you'll design and build complex systems, support Atlas platform operations, automate processes, and ensure high availability of services.
Top Skills:
AWSAzureDnsGCPGoHTTPLinuxPythonRubyTls
AdTech • Big Data • Cloud • Marketing Tech • Software • Analytics
Lead SRE efforts to improve security, reliability, cost efficiency, and observability. Build automation, CI/CD, agentic AI platforms (MCPs), and tooling for capacity planning, incident response, and self-service. Evangelize SecDevOps and zero-trust designs across product and platform teams.
Top Skills:
Ai Agentic FrameworksArgocdAWSBashCi/CdDevsecopsDockerEksGoGrafanaKubernetesLinuxLokiMcpNew RelicPrometheusPythonSamTerraformZero Trust
Reposted 18 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
Develop and maintain Kubernetes runtime environments, support developers, resolve critical issues, and participate in on-call rotations for production systems.
Top Skills:
AWSAzureCert-ManagerCorednsCrdsCriCsiGatekeeperGCPGoHelmKubernetesKustomizeOperatorsPythonTerraform
Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
The Senior Site Reliability Engineer will enhance reliability of Block's platform, improve incident response using AI tools, and coordinate incident management. Responsibilities include building reliable systems, standardizing tools, and leading high-severity incidents during on-call rotations.
Top Skills:
Amazon Web ServicesDatadogDynamoDBGrpcHTTPIstioJavaJSONKotlinKubernetesLaunchdarklyMySQLProtocol BuffersTerraformVitess
Fintech • Software
Lead SRE efforts for DFIN SaaS: ensure availability, performance, scalability, and automation. Implement monitoring, CI/CD, IaC, container orchestration, AI-enhanced observability, incident response, RCA, and runbook automation while collaborating across engineering teams.
Top Skills:
.NetAiopsAksAnsibleAppdynamicsAWSAzureAzure DevopsBashC#Ci/CdCloud Ai ServicesContainersCosmosDatadogDynatraceEksFirewallHarnessIdera Sql Diagnostic ManagerInfrastructure As Code (Iac)JavaJenkinsKubernetesLinuxLoad BalancingNew RelicPowershellPythonRedgate Sql MonitorSolarwinds Database Performance AnalyzerSQLTerraformWindows
Fintech • Professional Services • Software
Own platform reliability, performance, observability, and resilience for scalable backend systems. Mentor engineering teams on secure, performant coding practices; establish OpenTelemetry observability and service-level objectives; design and improve incident response processes; and implement automated testing. Collaborate with engineering leaders, product managers, and cross-functional teams on shared infrastructure and challenging technical solutions. The role also contributes to cloud infrastructure, CI/CD, tabletop exercises, and modern distributed systems across the organization.
Top Skills:
AWSClaudeDebeziumDockerGitlab Ci/CdHelmJavaKafkaKubernetesMySQLOpentelemetryPostgresRestful ApisSnowflakeSpring BootSQLTemporalTerraform
Fintech • Professional Services • Software
Build and improve resilient, scalable backend systems as part of the Platform team. Responsibilities include implementing observability with OpenTelemetry, defining SLOs, owning incident response processes and follow-ups, mentoring teams on secure and reliable coding, and applying automated testing. The role involves working with Java, Spring Boot, microservices, distributed systems, Kubernetes, cloud infrastructure, and CI/CD while collaborating with engineering leadership and cross-functional teams.
Top Skills:
AWSClaudeDebeziumDistributed SystemsDockerGitlab Ci/CdHelmJavaKafkaKubernetesMicroservicesMySQLOpentelemetryPostgresRestful ApisSnowflakeSpring BootSQLTemporalTerraform
Healthtech • Software
Operate and maintain AWS-hosted MERN applications and large-scale data workflows. Manage serverless and Spark-based pipelines, perform incident response and on-call duties, engineer automation to eliminate operational toil, ensure HIPAA/SOC2/HITRUST compliance, build observability and lead blameless post-mortems.
Top Skills:
Amazon EcsAmazon EksAmazon EmrAthenaAws GlueAws LambdaAws SnsAws SqsCloudwatchEc2IamJavaScriptMernMySQLNode.jsOpentofuPysparkPythonRabbitMQTerraformTypescriptVpc
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Senior SRE owning availability, automation, and observability for CI/CD platform services. Build and operate infrastructure, run on-call, lead incident response, mentor engineers, drive design/capacity planning, integrate AI-assisted workflows, and improve cross-team reliability.
Top Skills:
Active DirectoryAnsibleApache AirflowSparkAWSAzureBashBazelBitbucketCassandraChefDatadogDnsFirewall RulesGCPGitGithub ActionsGitlabGitlab CiGoGrafanaHoneycombHumio/LogscaleJenkinsKafkaKubernetesLoad BalancersMongoDBMySQLNasNew RelicNfsObject StorageOpensearchOraclePostgresPowershellPrometheusPulsarPuppetPythonRabbitMQRedis/ValkeyRedpandaRoutingSaltSanSplunkTerraformVarnishVipsWindows Server
Fintech • Financial Services
Operates and improves high-traffic, business-critical cloud and network systems. Designs scalable network solutions, manages performance and troubleshooting, automates deployment and monitoring, coordinates router and switch installations, and tests redundancy, resilience, and failover. Partners with development teams to improve operability, supports project planning and vendor comparisons, provides training, and participates in on-call coverage.
Top Skills:
Cloud ComputingIpRoutersSwitchesVoip
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
Reposted 24 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
The Senior Site Reliability Engineer will develop and support distributed storage services, ensuring reliability and operational safety, with a focus on automation and efficiency.
Top Skills:
AWSAzureDnsGoGoogle Cloud PlatformKubernetesLinuxPythonTcp/IpTls
Aerospace
Design, build, and operate highly available ground and site network infrastructure across commercial and secured environments. Architect Kubernetes platforms, enforce security and compliance controls, develop Terraform-based infrastructure as code, and establish monitoring, logging, alerting, CI/CD, and automation. Serve as a site reliability owner for mission-critical systems, driving uptime and incident response. Partner with security and mission operations teams, define architectural standards, and mentor engineers.
Top Skills:
ArgocdCi/CdDockerGitopsGoGrafanaInfrastructure As CodeKubernetesMimirNetwork SegmentationNetworkingObservabilityPrometheusPythonRoutingSwitchingTerraformTerragrunt
eCommerce • Fintech • Payments • Software
The role involves ensuring software reliability and performance, managing incidents, developing infrastructure automation, and mentoring junior engineers within a platform team.
Top Skills:
AWSCloudFormationDatadogKubernetesOpentelemetryRubyRuby On RailsTerraform
Artificial Intelligence • Cloud • Information Technology • Consulting
Internship SRE role responsible for availability, performance, and scalability of an e-commerce supply-chain platform. Tasks include SLO/SLA definition, observability (Prometheus/Grafana/Loki/Tempo/OpenTelemetry), incident response, capacity planning, disaster recovery for PostgreSQL, infrastructure-as-code (Terraform), CI/CD automation, and operational reliability for AI agent services. Mentored by Head of Technology/CTO with potential conversion to full-time based on performance.
Top Skills:
BashCi/CdDockerGrafanaLangchainLlmLokiMakefileNestjsOpentelemetryOracle CloudPgbackrestPostgresql 15PrometheusPythonRedisTempoTerraformTraefik
Digital Media • Software
Deploy, operate, validate, and improve Ateme video delivery platforms in customer environments, including critical live production systems. Responsibilities include Linux administration, networking, video technology integration, troubleshooting, performance investigations, documentation, customer training, technical collaboration with R&D, and ensuring service availability. The role supports live production events and participates in a 24/7 on-call rotation.
Top Skills:
ArqAutomationAv1AvcCdnCloud InfrastructureCmafDistributed SystemsDrmHevcHlsJpeg-XsKubernetesLinuxMpeg-TsNetworkingObservabilityScte-104Scte-224Scte-30Scte-35Smpte 2110UnixVideo Encoding
Edtech • Kids + Family • Sports
Audit infrastructure, deployment pipelines, monitoring, alerting, incident response, on-call practices, and internal tools. Produce actionable audit reports, implement code and configuration fixes, improve SLOs and reliability practices, advise on scalable architecture, and partner with engineers on implementation and handoff. The role is a fully remote, part-time consulting engagement with potential for full-time conversion.
Top Skills:
Ai Coding ToolsAWSCi/CdDatadogGrafanaPrometheus
Cloud • Information Technology • Cybersecurity • Infrastructure as a Service (IaaS)
Owns reliability, observability, and incident response for a GPUaaS platform. Defines SLOs, builds monitoring and alerting systems, leads major incidents and post-incident reviews, automates operational processes, maintains runbooks, manages on-call operations, coordinates with engineering teams, drives chaos testing, reports SLA performance, and mentors junior engineers.
Top Skills:
DatadogGoGpuaasGrafanaGremlinHpcKubernetesLitmusOpentelemetryPrometheusPython
Aerospace • Hardware • Software • Defense • Manufacturing
Build scalable automated solutions for device fleet management, own and optimize MDM platforms, write OS-level scripts for self-healing, gather telemetry to prevent end-user disruption, translate compliance (CMMC) into code-managed baselines, and create dashboards and alerts measuring end-user SLOs.
Top Skills:
AnsibleBashChefFleet DmIntuneJAMFOsqueryPowershellPulumiPuppetPythonSaltTerraformWorkspace One
Artificial Intelligence • Cloud • Fintech • Machine Learning • Mobile • Software
Lead design, development, deployment, and scaling of cloud infrastructure and SRE tooling. Build automation, CI/CD, observability, capacity planning, and reliability improvements; collaborate with product teams to define non-functional requirements and resolve production issues.
Top Skills:
.NetApi GatewayAWSAzureC#Data LakehouseDatabricks DeltaDatadogElasticsearchElkEvent HubsFunctions/ServerlessGitGrafanaJavaJenkinsKafkaKibanaKubernetesLogstashPowershellSnowflakeSqsTeamcityVisual Basic
Healthtech • Software
Manage and optimize a multi-account AWS environment supporting production healthcare applications and analytics platforms. Responsibilities include AWS infrastructure administration, CI/CD and infrastructure automation, observability, incident response, on-call support, disaster recovery validation, HIPAA/HiTrust compliance, IAM and security controls, and support for containerized Java and Python applications. The role also contributes to Kubernetes and EKS modernization initiatives and partners with developers to resolve complex production issues.
Top Skills:
Amazon EksApi GatewayArgocdAuroraAWSAws CloudformationAws CodepipelineAws Security HubBashCloudfrontCloudwatchCortex CloudDatadogDockerEc2EcsFargateGitGuarddutyHelmIamJavaJenkinsKubernetesLambdaLinuxPrismaPythonRdsS3Spring BootUbuntuZabbix
Software
Owns reliability, observability, performance, and security for a multi-region SaaS platform. Responsibilities include managing Datadog, implementing APM and tracing, defining SLOs, developing automation, expanding infrastructure as code and CI/CD, automating operational workflows, maintaining security controls, participating in incident response, documenting procedures, and mentoring engineers.
Top Skills:
ApmAzure DevopsAzure Kubernetes ServiceAzure SqlBashBicepCi/CdCosmos DbDatadogDistributed TracingHelmInfrastructure As CodeKey VaultKubernetesKustomizeManaged IdentitiesAzureMicrosoft Entra IdPowershellPythonRedisService BusTerraform
Information Technology • Legal Tech • Analytics
Support and automate cloud infrastructure (primarily Azure) for insurance applications. Implement IaC, assist deployments, troubleshoot incidents, improve security controls, and maintain monitoring and documentation while collaborating with engineers and non-technical stakeholders.
Top Skills:
Arm TemplatesAWSAzureBashCi/CdContainerizationDnsGCPGitGitMonitoring And Logging ToolsPowershellPythonServerless FunctionsTerraformVnets
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Ensure stability and resilience of Runpod's distributed AI platform by defining SLIs/SLOs, leading incident response, building observability and reliability tooling, automating operational workflows, and partnering with engineering teams to reduce toil and improve production readiness.
Top Skills:
BashCi/CdContainerized Production SystemsGoGpu Observability ToolingGrafanaInfrastructure As CodeLinuxPrometheusPython
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Los Angeles, CA Companies Hiring SRE Engineers
See AllPopular Los Angeles, CA Engineering Job Searches
Engineering Jobs in Los Angeles, CA
Software Engineer Jobs in Los Angeles, CA
Android Developer Jobs in Los Angeles, CA
C# Jobs in Los Angeles, CA
C++ Jobs in Los Angeles, CA
DevOps Jobs in Los Angeles, CA
Front End Developer Jobs in Los Angeles, CA
Golang Jobs in Los Angeles, CA
Hardware Engineer Jobs in Los Angeles, CA
iOS Developer Jobs in Los Angeles, CA
Java Developer Jobs in Los Angeles, CA
Javascript Jobs in Los Angeles, CA
Linux Jobs in Los Angeles, CA
Engineering Manager Jobs in Los Angeles, CA
.NET Developer Jobs in Los Angeles, CA
PHP Developer Jobs in Los Angeles, CA
Python Jobs in Los Angeles, CA
QA Jobs in Los Angeles, CA
Ruby Jobs in Los Angeles, CA
Salesforce Developer Jobs in Los Angeles, CA
Scala Jobs in Los Angeles, CA
Application Engineer Jobs in Los Angeles, CA
Associate Software Engineer Jobs in Los Angeles, CA
Automation Engineer Jobs in Los Angeles, CA
AWS Engineer Jobs in Los Angeles, CA
Backend Engineer Jobs in Los Angeles, CA
Cloud Engineer Jobs in Los Angeles, CA
Controls Engineer Jobs in Los Angeles, CA
CTO Jobs in Los Angeles, CA
Design Engineer Jobs in Los Angeles, CA
DevOps Engineer Jobs in Los Angeles, CA
Director of Engineering Jobs in Los Angeles, CA
Electrical Engineering Jobs in Los Angeles, CA
Embedded Software Engineer Jobs in Los Angeles, CA
Field Engineer Jobs in Los Angeles, CA
Firmware Engineer Jobs in Los Angeles, CA
Full-Stack Engineer Jobs in Los Angeles, CA
Game Engineer Jobs in Los Angeles, CA
Industrial Engineer Jobs in Los Angeles, CA
Infrastructure Engineer Jobs in Los Angeles, CA
Manufacturing Engineer Jobs in Los Angeles, CA
Mechanical Design Engineer Jobs in Los Angeles, CA
Mechanical Engineering Jobs in Los Angeles, CA
Mechatronics Engineering Jobs in Los Angeles, CA
Network Engineer Jobs in Los Angeles, CA
Platform Engineer Jobs in Los Angeles, CA
Principal Engineer Jobs in Los Angeles, CA
Principal Software Engineer Jobs in Los Angeles, CA
Process Engineer Jobs in Los Angeles, CA
Product Engineer Jobs in Los Angeles, CA
Project Engineer Jobs in Los Angeles, CA
QA Engineer Jobs in Los Angeles, CA
Reliability Engineer Jobs in Los Angeles, CA
Robotics Engineer Jobs in Los Angeles, CA
Software Architect Jobs in Los Angeles, CA
Software Engineering Manager Jobs in Los Angeles, CA
Solutions Engineer Jobs in Los Angeles, CA
SRE Engineer Jobs in Los Angeles, CA
Staff Engineer Jobs in Los Angeles, CA
Staff Software Engineer Jobs in Los Angeles, CA
Structural Engineer Jobs in Los Angeles, CA
Systems Engineer Jobs in Los Angeles, CA
VP of Engineering Jobs in Los Angeles, CA
Web Developer Jobs in Los Angeles, CA
All Filters
Total selected ()
No Results
No Results
.png)










.png)



















