Maximum of 25 job preferences reached.
Top Reliability Engineer Jobs in Los Angeles, CA
Reposted 24 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
The Senior Site Reliability Engineer will develop and support distributed storage services, ensuring reliability and operational safety, with a focus on automation and efficiency.
Top Skills:
AWSAzureDnsGoGoogle Cloud PlatformKubernetesLinuxPythonTcp/IpTls
Digital Media • Gaming • News + Entertainment • Sports
Design, build, and operate highly available, observable cloud-native production systems. Improve reliability through automation and infrastructure-as-code, troubleshoot performance issues, collaborate with engineers for scalable designs, and ensure security best practices across cloud and data-center environments.
Top Skills:
Active DirectoryAmazon LinuxAnsibleAWSAzureDatadogDockerGitlab Ci/CdGoGCPGrafanaJenkinsKubernetesLdapNew RelicPing IdentityPythonRancherRubySwiftTerraformWindows
3D Printing • Aerospace • Hardware • Software • Manufacturing
Perform system-level reliability analyses (FTA, FMEA/FMECA, RBDs), reliability predictions, and develop/maintain spaceflight component reliability data. Support trade studies, identify critical components, assist root cause/failure investigations and risk characterization for human-rated space station systems.
Top Skills:
Fault Tree AnalysisFmeaFmecaFracasFtaProbabilistic Risk AnalysisReliability Block DiagramsReliability Growth ModelingReliability PredictionsStress-Strength Interference AnalysisWeibull Analysis
Aerospace
Design, build, and operate highly available ground and site network infrastructure across commercial and secured environments. Architect Kubernetes platforms, enforce security and compliance controls, develop Terraform-based infrastructure as code, and establish monitoring, logging, alerting, CI/CD, and automation. Serve as a site reliability owner for mission-critical systems, driving uptime and incident response. Partner with security and mission operations teams, define architectural standards, and mentor engineers.
Top Skills:
ArgocdCi/CdDockerGitopsGoGrafanaInfrastructure As CodeKubernetesMimirNetwork SegmentationNetworkingObservabilityPrometheusPythonRoutingSwitchingTerraformTerragrunt
eCommerce • Fintech • Payments • Software
The role involves ensuring software reliability and performance, managing incidents, developing infrastructure automation, and mentoring junior engineers within a platform team.
Top Skills:
AWSCloudFormationDatadogKubernetesOpentelemetryRubyRuby On RailsTerraform
Aerospace • Other
Ensure Falcon and Dragon flight hardware meets design criteria and reliability standards. Perform design reviews and test approvals, develop requirements and verification approaches, enable aircraft-like operations, drive design improvements with hardware teams, and build specialized design/analysis tools.
Top Skills:
AbaqusAnsysExcelFeaMatlabNastranPythonVisual Basic
Artificial Intelligence • Cloud • Information Technology • Consulting
Internship SRE role responsible for availability, performance, and scalability of an e-commerce supply-chain platform. Tasks include SLO/SLA definition, observability (Prometheus/Grafana/Loki/Tempo/OpenTelemetry), incident response, capacity planning, disaster recovery for PostgreSQL, infrastructure-as-code (Terraform), CI/CD automation, and operational reliability for AI agent services. Mentored by Head of Technology/CTO with potential conversion to full-time based on performance.
Top Skills:
BashCi/CdDockerGrafanaLangchainLlmLokiMakefileNestjsOpentelemetryOracle CloudPgbackrestPostgresql 15PrometheusPythonRedisTempoTerraformTraefik
Digital Media • Software
Deploy, operate, validate, and improve Ateme video delivery platforms in customer environments, including critical live production systems. Responsibilities include Linux administration, networking, video technology integration, troubleshooting, performance investigations, documentation, customer training, technical collaboration with R&D, and ensuring service availability. The role supports live production events and participates in a 24/7 on-call rotation.
Top Skills:
ArqAutomationAv1AvcCdnCloud InfrastructureCmafDistributed SystemsDrmHevcHlsJpeg-XsKubernetesLinuxMpeg-TsNetworkingObservabilityScte-104Scte-224Scte-30Scte-35Smpte 2110UnixVideo Encoding
Aerospace • Other
Perform failure analysis and root cause investigations on satellite PCBAs; collaborate with design, manufacturing, test, and supply chain teams to improve reliability; run benchtop environmental tests; monitor yields and quality metrics; document findings and drive corrective actions to enhance product robustness and test coverage.
Top Skills:
CC++CanDfxDigital MultimetersEthernetHaltHassI2CLeanOscilloscopesPcbasPower SuppliesPythonRf TestingRs422Six SigmaSoldering EquipmentSpcSpiThermal TestingTvacVibration Testing
Legal Tech • Software
Lead automation and optimization of Filevine's data platform: performance tune MSSQL/Postgres, optimize Snowflake, provision infrastructure with Terraform/AWS, run stateful containers on Kubernetes, integrate AI/LLM and MCP for operational automation, manage CI/CD, capacity planning, documentation, and serve in 24/7 on-call rotation.
Top Skills:
AWSC#DapperDockerDynamoDBEntity FrameworkGitlabKubernetesLlmsMcp (Model Context Protocol)Microsoft Sql Server (Mssql)Octopus DeployOpensearchPostgresPowershellPythonRedisSnowflakeTerraform
Edtech • Kids + Family • Sports
Audit infrastructure, deployment pipelines, monitoring, alerting, incident response, on-call practices, and internal tools. Produce actionable audit reports, implement code and configuration fixes, improve SLOs and reliability practices, advise on scalable architecture, and partner with engineers on implementation and handoff. The role is a fully remote, part-time consulting engagement with potential for full-time conversion.
Top Skills:
Ai Coding ToolsAWSCi/CdDatadogGrafanaPrometheus
Cloud • Information Technology • Cybersecurity • Infrastructure as a Service (IaaS)
Owns reliability, observability, and incident response for a GPUaaS platform. Defines SLOs, builds monitoring and alerting systems, leads major incidents and post-incident reviews, automates operational processes, maintains runbooks, manages on-call operations, coordinates with engineering teams, drives chaos testing, reports SLA performance, and mentors junior engineers.
Top Skills:
DatadogGoGpuaasGrafanaGremlinHpcKubernetesLitmusOpentelemetryPrometheusPython
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
Aerospace • Hardware • Software • Defense • Manufacturing
Build scalable automated solutions for device fleet management, own and optimize MDM platforms, write OS-level scripts for self-healing, gather telemetry to prevent end-user disruption, translate compliance (CMMC) into code-managed baselines, and create dashboards and alerts measuring end-user SLOs.
Top Skills:
AnsibleBashChefFleet DmIntuneJAMFOsqueryPowershellPulumiPuppetPythonSaltTerraformWorkspace One
One Month AgoSaved
Easy Apply
Easy Apply
Big Data • Fintech • Mobile • Payments • Financial Services
Lead design and delivery of a reliability platform for production systems. Own quarterly goals, guide engineers through ambiguity, collaborate with PM/design/analytics, monitor and operate services (on-call), set quality and code standards, and mentor teammates. Integrate AI/LLM tooling to improve automation, debugging, and service health.
Top Skills:
Ai FrameworksAWSKotlinKubernetesLlmsMySQLPythonReactVue
Artificial Intelligence • Cloud • Fintech • Machine Learning • Mobile • Software
Lead design, development, deployment, and scaling of cloud infrastructure and SRE tooling. Build automation, CI/CD, observability, capacity planning, and reliability improvements; collaborate with product teams to define non-functional requirements and resolve production issues.
Top Skills:
.NetApi GatewayAWSAzureC#Data LakehouseDatabricks DeltaDatadogElasticsearchElkEvent HubsFunctions/ServerlessGitGrafanaJavaJenkinsKafkaKibanaKubernetesLogstashPowershellSnowflakeSqsTeamcityVisual Basic
Healthtech • Software
Manage and optimize a multi-account AWS environment supporting production healthcare applications and analytics platforms. Responsibilities include AWS infrastructure administration, CI/CD and infrastructure automation, observability, incident response, on-call support, disaster recovery validation, HIPAA/HiTrust compliance, IAM and security controls, and support for containerized Java and Python applications. The role also contributes to Kubernetes and EKS modernization initiatives and partners with developers to resolve complex production issues.
Top Skills:
Amazon EksApi GatewayArgocdAuroraAWSAws CloudformationAws CodepipelineAws Security HubBashCloudfrontCloudwatchCortex CloudDatadogDockerEc2EcsFargateGitGuarddutyHelmIamJavaJenkinsKubernetesLambdaLinuxPrismaPythonRdsS3Spring BootUbuntuZabbix
Software
Owns reliability, observability, performance, and security for a multi-region SaaS platform. Responsibilities include managing Datadog, implementing APM and tracing, defining SLOs, developing automation, expanding infrastructure as code and CI/CD, automating operational workflows, maintaining security controls, participating in incident response, documenting procedures, and mentoring engineers.
Top Skills:
ApmAzure DevopsAzure Kubernetes ServiceAzure SqlBashBicepCi/CdCosmos DbDatadogDistributed TracingHelmInfrastructure As CodeKey VaultKubernetesKustomizeManaged IdentitiesAzureMicrosoft Entra IdPowershellPythonRedisService BusTerraform
Information Technology • Legal Tech • Analytics
Support and automate cloud infrastructure (primarily Azure) for insurance applications. Implement IaC, assist deployments, troubleshoot incidents, improve security controls, and maintain monitoring and documentation while collaborating with engineers and non-technical stakeholders.
Top Skills:
Arm TemplatesAWSAzureBashCi/CdContainerizationDnsGCPGitGitMonitoring And Logging ToolsPowershellPythonServerless FunctionsTerraformVnets
3D Printing • Aerospace • Hardware • Software • Manufacturing
Lead build reliability for vehicle integration and test, partnering across engineering, manufacturing, quality, and supply chain to define readiness criteria, drive process maturity, investigate root causes, implement corrective actions, and mentor teams to improve manufacturability, verification, and flight readiness for human-rated space stations.
Top Skills:
As9100Control PlansDesign For ManufacturabilityFmeaIso 9001RcaSix Sigma Black BeltSpc
3D Printing • Aerospace • Hardware • Software • Manufacturing
Lead reliability integration for Haven-1 structures: drive FMEA and fault-tree analysis, define test plans and verification/validation, shepherd design through build and test, create standards and tools, manage change/configuration, and partner across engineering, risk, and manufacturing to ensure human-rated structural reliability.
Top Skills:
ExcelMatlabPythonTableau
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Ensure stability and resilience of Runpod's distributed AI platform by defining SLIs/SLOs, leading incident response, building observability and reliability tooling, automating operational workflows, and partnering with engineering teams to reduce toil and improve production readiness.
Top Skills:
BashCi/CdContainerized Production SystemsGoGpu Observability ToolingGrafanaInfrastructure As CodeLinuxPrometheusPython
Edtech
Lead infrastructure modernization and platform reliability across multiple cloud providers. Design infrastructure as code, operate Kubernetes and Linux environments, improve CI/CD and deployment tooling, establish SLI/SLO practices, strengthen observability, lead incident response, manage cloud costs, and partner on security and compliance. Provide technical leadership through architecture guidance, mentorship, engineering standards, and roadmap development while participating in on-call support.
Top Skills:
AWSCi/CdGCPJenkinsKubernetesLinuxPythonRubyRuby On RailsSoc 2SpinnakerTerraform
Cloud • Information Technology • Business Intelligence • Consulting
Design, build, and operate cloud infrastructure and SRE capabilities for an enterprise AI platform. Responsibilities include infrastructure-as-code, landing zones, networking, Kubernetes, CI/CD, observability, incident response, SLOs, production readiness, automation, cost optimization, and support for hybrid, edge, air-gapped, and customer-controlled environments. The role is remote, client-facing, and requires strong collaboration and reliability ownership.
Top Skills:
AlertingAzureAzure ArcAzure DevopsBicepCi/CdDashboardsDockerGithub ActionsGpu WorkloadsInfrastructure As CodeKubernetesLogsMetricsObservabilitySlosTerraformTraces
Cloud • Security • Software • Cybersecurity
The Site Reliability Engineer II ensures the reliability, availability, performance, and security of critical cloud systems and services. Responsibilities include developing automation for provisioning and configuration management, maintaining monitoring and alerting, optimizing infrastructure performance, supporting high availability, and enabling continuous integration and delivery. The role collaborates with security teams and drives operational improvements across cloud and network infrastructure.
Top Skills:
AnsibleAWSAzureChefContinuous DeliveryContinuous IntegrationDnsElk StackGCPGoGrafanaHTTPKubernetesLinuxPrometheusPuppetPythonShellTcp/IpUnix
Healthtech • Insurance
Lead cloud, DevOps, and SRE architecture efforts to scale CareSource's digital platform. Partner with Cloud, DevOps, Security, and SRE teams to design infrastructure, build Terraform templates, enhance CI/CD pipelines, implement monitoring/alerting, and improve reliability, scalability, and incident response for enterprise-scale digital products.
Top Skills:
Azure CloudDockerDynatraceGithub ActionsKubernetesSplunkTerraform Enterprise
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Los Angeles, CA Companies Hiring Reliability Engineers
See AllPopular Los Angeles, CA Engineering Job Searches
Engineering Jobs in Los Angeles, CA
Software Engineer Jobs in Los Angeles, CA
Android Developer Jobs in Los Angeles, CA
C# Jobs in Los Angeles, CA
C++ Jobs in Los Angeles, CA
DevOps Jobs in Los Angeles, CA
Front End Developer Jobs in Los Angeles, CA
Golang Jobs in Los Angeles, CA
Hardware Engineer Jobs in Los Angeles, CA
iOS Developer Jobs in Los Angeles, CA
Java Developer Jobs in Los Angeles, CA
Javascript Jobs in Los Angeles, CA
Linux Jobs in Los Angeles, CA
Engineering Manager Jobs in Los Angeles, CA
.NET Developer Jobs in Los Angeles, CA
PHP Developer Jobs in Los Angeles, CA
Python Jobs in Los Angeles, CA
QA Jobs in Los Angeles, CA
Ruby Jobs in Los Angeles, CA
Salesforce Developer Jobs in Los Angeles, CA
Scala Jobs in Los Angeles, CA
Application Engineer Jobs in Los Angeles, CA
Associate Software Engineer Jobs in Los Angeles, CA
Automation Engineer Jobs in Los Angeles, CA
AWS Engineer Jobs in Los Angeles, CA
Backend Engineer Jobs in Los Angeles, CA
Cloud Engineer Jobs in Los Angeles, CA
Controls Engineer Jobs in Los Angeles, CA
CTO Jobs in Los Angeles, CA
Design Engineer Jobs in Los Angeles, CA
DevOps Engineer Jobs in Los Angeles, CA
Director of Engineering Jobs in Los Angeles, CA
Electrical Engineering Jobs in Los Angeles, CA
Embedded Software Engineer Jobs in Los Angeles, CA
Field Engineer Jobs in Los Angeles, CA
Firmware Engineer Jobs in Los Angeles, CA
Full-Stack Engineer Jobs in Los Angeles, CA
Game Engineer Jobs in Los Angeles, CA
Industrial Engineer Jobs in Los Angeles, CA
Infrastructure Engineer Jobs in Los Angeles, CA
Manufacturing Engineer Jobs in Los Angeles, CA
Mechanical Design Engineer Jobs in Los Angeles, CA
Mechanical Engineering Jobs in Los Angeles, CA
Mechatronics Engineering Jobs in Los Angeles, CA
Network Engineer Jobs in Los Angeles, CA
Platform Engineer Jobs in Los Angeles, CA
Principal Engineer Jobs in Los Angeles, CA
Principal Software Engineer Jobs in Los Angeles, CA
Process Engineer Jobs in Los Angeles, CA
Product Engineer Jobs in Los Angeles, CA
Project Engineer Jobs in Los Angeles, CA
QA Engineer Jobs in Los Angeles, CA
Reliability Engineer Jobs in Los Angeles, CA
Robotics Engineer Jobs in Los Angeles, CA
Software Architect Jobs in Los Angeles, CA
Software Engineering Manager Jobs in Los Angeles, CA
Solutions Engineer Jobs in Los Angeles, CA
SRE Engineer Jobs in Los Angeles, CA
Staff Engineer Jobs in Los Angeles, CA
Staff Software Engineer Jobs in Los Angeles, CA
Structural Engineer Jobs in Los Angeles, CA
Systems Engineer Jobs in Los Angeles, CA
VP of Engineering Jobs in Los Angeles, CA
Web Developer Jobs in Los Angeles, CA
All Filters
Total selected ()
No Results
No Results




.png)
























