Top Reliability Engineer Jobs in Los Angeles, CA

Reposted 24 Days AgoSaved
Easy Apply
Remote or Hybrid
Los Angeles, CA
Easy Apply
126K-248K Annually
Senior level
126K-248K Annually
Senior level
Big Data • Cloud • Software • Database
The Senior Site Reliability Engineer will develop and support distributed storage services, ensuring reliability and operational safety, with a focus on automation and efficiency.
Top Skills: AWSAzureDnsGoGoogle Cloud PlatformKubernetesLinuxPythonTcp/IpTls
Reposted 29 Days AgoSaved
In-Office
Los Angeles, CA
142K-190K Annually
Senior level
142K-190K Annually
Senior level
Digital Media • Gaming • News + Entertainment • Sports
Design, build, and operate highly available, observable cloud-native production systems. Improve reliability through automation and infrastructure-as-code, troubleshoot performance issues, collaborate with engineers for scalable designs, and ensure security best practices across cloud and data-center environments.
Top Skills: Active DirectoryAmazon LinuxAnsibleAWSAzureDatadogDockerGitlab Ci/CdGoGCPGrafanaJenkinsKubernetesLdapNew RelicPing IdentityPythonRancherRubySwiftTerraformWindows
Reposted 29 Days AgoSaved
In-Office
Los Angeles, CA
112K-159K Annually
Entry level
112K-159K Annually
Entry level
3D Printing • Aerospace • Hardware • Software • Manufacturing
Perform system-level reliability analyses (FTA, FMEA/FMECA, RBDs), reliability predictions, and develop/maintain spaceflight component reliability data. Support trade studies, identify critical components, assist root cause/failure investigations and risk characterization for human-rated space station systems.
Top Skills: Fault Tree AnalysisFmeaFmecaFracasFtaProbabilistic Risk AnalysisReliability Block DiagramsReliability Growth ModelingReliability PredictionsStress-Strength Interference AnalysisWeibull Analysis
7 Days AgoSaved
In-Office
Los Angeles, CA
155K-195K Annually
Senior level
155K-195K Annually
Senior level
Aerospace
Design, build, and operate highly available ground and site network infrastructure across commercial and secured environments. Architect Kubernetes platforms, enforce security and compliance controls, develop Terraform-based infrastructure as code, and establish monitoring, logging, alerting, CI/CD, and automation. Serve as a site reliability owner for mission-critical systems, driving uptime and incident response. Partner with security and mission operations teams, define architectural standards, and mentor engineers.
Top Skills: ArgocdCi/CdDockerGitopsGoGrafanaInfrastructure As CodeKubernetesMimirNetwork SegmentationNetworkingObservabilityPrometheusPythonRoutingSwitchingTerraformTerragrunt
Reposted 26 Days AgoSaved
Remote or Hybrid
Los Angeles, CA
175K-200K Annually
Senior level
175K-200K Annually
Senior level
eCommerce • Fintech • Payments • Software
The role involves ensuring software reliability and performance, managing incidents, developing infrastructure automation, and mentoring junior engineers within a platform team.
Top Skills: AWSCloudFormationDatadogKubernetesOpentelemetryRubyRuby On RailsTerraform
One Month AgoSaved
In-Office
Los Angeles, CA
100K-160K Annually
Junior
100K-160K Annually
Junior
Aerospace • Other
Ensure Falcon and Dragon flight hardware meets design criteria and reliability standards. Perform design reviews and test approvals, develop requirements and verification approaches, enable aircraft-like operations, drive design improvements with hardware teams, and build specialized design/analysis tools.
Top Skills: AbaqusAnsysExcelFeaMatlabNastranPythonVisual Basic
Reposted YesterdaySaved
Remote
Los Angeles, CA
Internship
Internship
Artificial Intelligence • Cloud • Information Technology • Consulting
Internship SRE role responsible for availability, performance, and scalability of an e-commerce supply-chain platform. Tasks include SLO/SLA definition, observability (Prometheus/Grafana/Loki/Tempo/OpenTelemetry), incident response, capacity planning, disaster recovery for PostgreSQL, infrastructure-as-code (Terraform), CI/CD automation, and operational reliability for AI agent services. Mentored by Head of Technology/CTO with potential conversion to full-time based on performance.
Top Skills: BashCi/CdDockerGrafanaLangchainLlmLokiMakefileNestjsOpentelemetryOracle CloudPgbackrestPostgresql 15PrometheusPythonRedisTempoTerraformTraefik
11 Days AgoSaved
In-Office
Los Angeles, CA
180K-200K Annually
Senior level
180K-200K Annually
Senior level
Digital Media • Software
Deploy, operate, validate, and improve Ateme video delivery platforms in customer environments, including critical live production systems. Responsibilities include Linux administration, networking, video technology integration, troubleshooting, performance investigations, documentation, customer training, technical collaboration with R&D, and ensuring service availability. The role supports live production events and participates in a 24/7 on-call rotation.
Top Skills: ArqAutomationAv1AvcCdnCloud InfrastructureCmafDistributed SystemsDrmHevcHlsJpeg-XsKubernetesLinuxMpeg-TsNetworkingObservabilityScte-104Scte-224Scte-30Scte-35Smpte 2110UnixVideo Encoding
Reposted One Month AgoSaved
In-Office
Los Angeles, CA
105K-150K Annually
Junior
105K-150K Annually
Junior
Aerospace • Other
Perform failure analysis and root cause investigations on satellite PCBAs; collaborate with design, manufacturing, test, and supply chain teams to improve reliability; run benchtop environmental tests; monitor yields and quality metrics; document findings and drive corrective actions to enhance product robustness and test coverage.
Top Skills: CC++CanDfxDigital MultimetersEthernetHaltHassI2CLeanOscilloscopesPcbasPower SuppliesPythonRf TestingRs422Six SigmaSoldering EquipmentSpcSpiThermal TestingTvacVibration Testing
Reposted 26 Days AgoSaved
Remote
Los Angeles, CA
145K-180K Annually
Senior level
145K-180K Annually
Senior level
Legal Tech • Software
Lead automation and optimization of Filevine's data platform: performance tune MSSQL/Postgres, optimize Snowflake, provision infrastructure with Terraform/AWS, run stateful containers on Kubernetes, integrate AI/LLM and MCP for operational automation, manage CI/CD, capacity planning, documentation, and serve in 24/7 on-call rotation.
Top Skills: AWSC#DapperDockerDynamoDBEntity FrameworkGitlabKubernetesLlmsMcp (Model Context Protocol)Microsoft Sql Server (Mssql)Octopus DeployOpensearchPostgresPowershellPythonRedisSnowflakeTerraform
4 Days AgoSaved
Remote
Los Angeles, CA
Senior level
Senior level
Edtech • Kids + Family • Sports
Audit infrastructure, deployment pipelines, monitoring, alerting, incident response, on-call practices, and internal tools. Produce actionable audit reports, implement code and configuration fixes, improve SLOs and reliability practices, advise on scalable architecture, and partner with engineers on implementation and handoff. The role is a fully remote, part-time consulting engagement with potential for full-time conversion.
Top Skills: Ai Coding ToolsAWSCi/CdDatadogGrafanaPrometheus
4 Days AgoSaved
Remote
Los Angeles, CA
Senior level
Senior level
Cloud • Information Technology • Cybersecurity • Infrastructure as a Service (IaaS)
Owns reliability, observability, and incident response for a GPUaaS platform. Defines SLOs, builds monitoring and alerting systems, leads major incidents and post-incident reviews, automates operational processes, maintains runbooks, manages on-call operations, coordinates with engineering teams, drives chaos testing, reports SLA performance, and mentors junior engineers.
Top Skills: DatadogGoGpuaasGrafanaGremlinHpcKubernetesLitmusOpentelemetryPrometheusPython
New

Track Smarter, Apply Better.

Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.

Use For Free
Application Tracker Preview
Reposted 13 Days AgoSaved
In-Office or Remote
Los Angeles, CA
164K-270K Annually
Mid level
164K-270K Annually
Mid level
Aerospace • Hardware • Software • Defense • Manufacturing
Build scalable automated solutions for device fleet management, own and optimize MDM platforms, write OS-level scripts for self-healing, gather telemetry to prevent end-user disruption, translate compliance (CMMC) into code-managed baselines, and create dashboards and alerts measuring end-user SLOs.
Top Skills: AnsibleBashChefFleet DmIntuneJAMFOsqueryPowershellPulumiPuppetPythonSaltTerraformWorkspace One
One Month AgoSaved
Easy Apply
Remote
Los Angeles, CA
Easy Apply
173K-255K Annually
Senior level
173K-255K Annually
Senior level
Big Data • Fintech • Mobile • Payments • Financial Services
Lead design and delivery of a reliability platform for production systems. Own quarterly goals, guide engineers through ambiguity, collaborate with PM/design/analytics, monitor and operate services (on-call), set quality and code standards, and mentor teammates. Integrate AI/LLM tooling to improve automation, debugging, and service health.
Top Skills: Ai FrameworksAWSKotlinKubernetesLlmsMySQLPythonReactVue
Reposted 5 Days AgoSaved
Remote or Hybrid
Los Angeles, CA
138K-221K Annually
Senior level
138K-221K Annually
Senior level
Artificial Intelligence • Cloud • Fintech • Machine Learning • Mobile • Software
Lead design, development, deployment, and scaling of cloud infrastructure and SRE tooling. Build automation, CI/CD, observability, capacity planning, and reliability improvements; collaborate with product teams to define non-functional requirements and resolve production issues.
Top Skills: .NetApi GatewayAWSAzureC#Data LakehouseDatabricks DeltaDatadogElasticsearchElkEvent HubsFunctions/ServerlessGitGrafanaJavaJenkinsKafkaKibanaKubernetesLogstashPowershellSnowflakeSqsTeamcityVisual Basic
5 Days AgoSaved
Remote
Los Angeles, CA
Senior level
Senior level
Healthtech • Software
Manage and optimize a multi-account AWS environment supporting production healthcare applications and analytics platforms. Responsibilities include AWS infrastructure administration, CI/CD and infrastructure automation, observability, incident response, on-call support, disaster recovery validation, HIPAA/HiTrust compliance, IAM and security controls, and support for containerized Java and Python applications. The role also contributes to Kubernetes and EKS modernization initiatives and partners with developers to resolve complex production issues.
Top Skills: Amazon EksApi GatewayArgocdAuroraAWSAws CloudformationAws CodepipelineAws Security HubBashCloudfrontCloudwatchCortex CloudDatadogDockerEc2EcsFargateGitGuarddutyHelmIamJavaJenkinsKubernetesLambdaLinuxPrismaPythonRdsS3Spring BootUbuntuZabbix
Reposted 5 Days AgoSaved
Remote
Los Angeles, CA
Senior level
Senior level
Software
Owns reliability, observability, performance, and security for a multi-region SaaS platform. Responsibilities include managing Datadog, implementing APM and tracing, defining SLOs, developing automation, expanding infrastructure as code and CI/CD, automating operational workflows, maintaining security controls, participating in incident response, documenting procedures, and mentoring engineers.
Top Skills: ApmAzure DevopsAzure Kubernetes ServiceAzure SqlBashBicepCi/CdCosmos DbDatadogDistributed TracingHelmInfrastructure As CodeKey VaultKubernetesKustomizeManaged IdentitiesAzureMicrosoft Entra IdPowershellPythonRedisService BusTerraform
Reposted 14 Days AgoSaved
In-Office
Los Angeles, CA
72K-119K Annually
Junior
72K-119K Annually
Junior
Information Technology • Legal Tech • Analytics
Support and automate cloud infrastructure (primarily Azure) for insurance applications. Implement IaC, assist deployments, troubleshoot incidents, improve security controls, and maintain monitoring and documentation while collaborating with engineers and non-technical stakeholders.
Top Skills: Arm TemplatesAWSAzureBashCi/CdContainerizationDnsGCPGitGitMonitoring And Logging ToolsPowershellPythonServerless FunctionsTerraformVnets
Reposted One Month AgoSaved
In-Office
Los Angeles, CA
138K-230K Annually
Senior level
138K-230K Annually
Senior level
3D Printing • Aerospace • Hardware • Software • Manufacturing
Lead build reliability for vehicle integration and test, partnering across engineering, manufacturing, quality, and supply chain to define readiness criteria, drive process maturity, investigate root causes, implement corrective actions, and mentor teams to improve manufacturability, verification, and flight readiness for human-rated space stations.
Top Skills: As9100Control PlansDesign For ManufacturabilityFmeaIso 9001RcaSix Sigma Black BeltSpc
Reposted One Month AgoSaved
In-Office
Los Angeles, CA
162K-265K Annually
Senior level
162K-265K Annually
Senior level
3D Printing • Aerospace • Hardware • Software • Manufacturing
Lead reliability integration for Haven-1 structures: drive FMEA and fault-tree analysis, define test plans and verification/validation, shepherd design through build and test, create standards and tools, manage change/configuration, and partner across engineering, risk, and manufacturing to ensure human-rated structural reliability.
Top Skills: ExcelMatlabPythonTableau
Reposted One Month AgoSaved
Remote
Los Angeles, CA
150K-200K Annually
Senior level
150K-200K Annually
Senior level
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Ensure stability and resilience of Runpod's distributed AI platform by defining SLIs/SLOs, leading incident response, building observability and reliability tooling, automating operational workflows, and partnering with engineering teams to reduce toil and improve production readiness.
Top Skills: BashCi/CdContainerized Production SystemsGoGpu Observability ToolingGrafanaInfrastructure As CodeLinuxPrometheusPython
5 Days AgoSaved
Remote
Los Angeles, CA
Senior level
Senior level
Edtech
Lead infrastructure modernization and platform reliability across multiple cloud providers. Design infrastructure as code, operate Kubernetes and Linux environments, improve CI/CD and deployment tooling, establish SLI/SLO practices, strengthen observability, lead incident response, manage cloud costs, and partner on security and compliance. Provide technical leadership through architecture guidance, mentorship, engineering standards, and roadmap development while participating in on-call support.
Top Skills: AWSCi/CdGCPJenkinsKubernetesLinuxPythonRubyRuby On RailsSoc 2SpinnakerTerraform
5 Days AgoSaved
Remote
Los Angeles, CA
115K-175K Annually
Entry level
115K-175K Annually
Entry level
Cloud • Information Technology • Business Intelligence • Consulting
Design, build, and operate cloud infrastructure and SRE capabilities for an enterprise AI platform. Responsibilities include infrastructure-as-code, landing zones, networking, Kubernetes, CI/CD, observability, incident response, SLOs, production readiness, automation, cost optimization, and support for hybrid, edge, air-gapped, and customer-controlled environments. The role is remote, client-facing, and requires strong collaboration and reliability ownership.
Top Skills: AlertingAzureAzure ArcAzure DevopsBicepCi/CdDashboardsDockerGithub ActionsGpu WorkloadsInfrastructure As CodeKubernetesLogsMetricsObservabilitySlosTerraformTraces
6 Days AgoSaved
In-Office or Remote
Los Angeles, CA
95K-171K Annually
Junior
95K-171K Annually
Junior
Cloud • Security • Software • Cybersecurity
The Site Reliability Engineer II ensures the reliability, availability, performance, and security of critical cloud systems and services. Responsibilities include developing automation for provisioning and configuration management, maintaining monitoring and alerting, optimizing infrastructure performance, supporting high availability, and enabling continuous integration and delivery. The role collaborates with security teams and drives operational improvements across cloud and network infrastructure.
Top Skills: AnsibleAWSAzureChefContinuous DeliveryContinuous IntegrationDnsElk StackGCPGoGrafanaHTTPKubernetesLinuxPrometheusPuppetPythonShellTcp/IpUnix
Reposted 6 Days AgoSaved
Remote
Los Angeles, CA
136K-237K Annually
Expert/Leader
136K-237K Annually
Expert/Leader
Healthtech • Insurance
Lead cloud, DevOps, and SRE architecture efforts to scale CareSource's digital platform. Partner with Cloud, DevOps, Security, and SRE teams to design infrastructure, build Terraform templates, enhance CI/CD pipelines, implement monitoring/alerting, and improve reliability, scalability, and incident response for enterprise-scale digital products.
Top Skills: Azure CloudDockerDynatraceGithub ActionsKubernetesSplunkTerraform Enterprise
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account