Maximum of 25 job preferences reached.
Top Reliability Engineer Jobs in Los Angeles, CA
AdTech • Big Data • Cloud • Marketing Tech • Software • Analytics
Lead SRE efforts to improve security, reliability, cost efficiency, and observability. Build automation, CI/CD, agentic AI platforms (MCPs), and tooling for capacity planning, incident response, and self-service. Evangelize SecDevOps and zero-trust designs across product and platform teams.
Top Skills:
Ai Agentic FrameworksArgocdAWSBashCi/CdDevsecopsDockerEksGoGrafanaKubernetesLinuxLokiMcpNew RelicPrometheusPythonSamTerraformZero Trust
Aerospace • Software
Lead mission assurance and flight readiness activities across vehicle programs, assess and mitigate risks, drive root-cause investigations, advance configuration and change management, support certification and customer/government interfaces, and partner with multidisciplinary teams to improve vehicle reliability and operations.
Aerospace • Software
Lead standardization and implementation of design processes across structures, propulsion, avionics, composites, and mechanisms. Define margining, configuration management, release workflows (CAD→PLM→MRP/ERP), review templates, and design-for-reliability practices. Create tools for traceability and use build/flight data to influence requirements. Drive quality, containment, corrective actions, and cross-functional problem solving with build, quality, and flight reliability teams.
Top Skills:
As9100CadErpFmeaGd&TIso 9001MrpPlmWeibull Analysis
Aerospace • Software
Lead reliability for propulsion hardware across design, manufacturing, and flight. Investigate and eliminate build defects, develop inspections and acceptance criteria, perform development testing, provide design-for-reliability guidance, optimize MRP/ERP traceability, build quality metrics, train technicians, and drive corrective actions and supplier quality improvements.
Top Skills:
As9100ErpFmeaGd&TIso 9001MrpWeibull Analysis
Reposted 18 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
Develop and maintain Kubernetes runtime environments, support developers, resolve critical issues, and participate in on-call rotations for production systems.
Top Skills:
AWSAzureCert-ManagerCorednsCrdsCriCsiGatekeeperGCPGoHelmKubernetesKustomizeOperatorsPythonTerraform
Aerospace
Design, build, and maintain highly available global ground-station network infrastructure. Define WAN/LAN architecture, configure routers/switches/firewalls, implement IP addressing and VLANs, enforce network security (VPN, ACLs), perform vulnerability scanning and hardening, deploy monitoring/observability, coordinate telecom providers, document configurations/diagrams, automate health reporting, and lead cybersecurity compliance for the Quartz program.
Top Skills:
BgpC++DhcpDnsFips 140-2/140-3FirewallsGoIpv4Ipv6Is-IsJuniperMplsNetwork MonitoringNist Sp 800-SeriesNtpOspfOut-Of-Band ManagementPalo Alto NetworksPythonRoutersRustSwitchesTcp/IpVlanVpnVulnerability Scanning
Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
The Senior Site Reliability Engineer will enhance reliability of Block's platform, improve incident response using AI tools, and coordinate incident management. Responsibilities include building reliable systems, standardizing tools, and leading high-severity incidents during on-call rotations.
Top Skills:
Amazon Web ServicesDatadogDynamoDBGrpcHTTPIstioJavaJSONKotlinKubernetesLaunchdarklyMySQLProtocol BuffersTerraformVitess
Fintech • Software
Lead SRE efforts for DFIN SaaS: ensure availability, performance, scalability, and automation. Implement monitoring, CI/CD, IaC, container orchestration, AI-enhanced observability, incident response, RCA, and runbook automation while collaborating across engineering teams.
Top Skills:
.NetAiopsAksAnsibleAppdynamicsAWSAzureAzure DevopsBashC#Ci/CdCloud Ai ServicesContainersCosmosDatadogDynatraceEksFirewallHarnessIdera Sql Diagnostic ManagerInfrastructure As Code (Iac)JavaJenkinsKubernetesLinuxLoad BalancingNew RelicPowershellPythonRedgate Sql MonitorSolarwinds Database Performance AnalyzerSQLTerraformWindows
Fintech • Professional Services • Software
Own platform reliability, performance, observability, and resilience for scalable backend systems. Mentor engineering teams on secure, performant coding practices; establish OpenTelemetry observability and service-level objectives; design and improve incident response processes; and implement automated testing. Collaborate with engineering leaders, product managers, and cross-functional teams on shared infrastructure and challenging technical solutions. The role also contributes to cloud infrastructure, CI/CD, tabletop exercises, and modern distributed systems across the organization.
Top Skills:
AWSClaudeDebeziumDockerGitlab Ci/CdHelmJavaKafkaKubernetesMySQLOpentelemetryPostgresRestful ApisSnowflakeSpring BootSQLTemporalTerraform
Fintech • Professional Services • Software
Build and improve resilient, scalable backend systems as part of the Platform team. Responsibilities include implementing observability with OpenTelemetry, defining SLOs, owning incident response processes and follow-ups, mentoring teams on secure and reliable coding, and applying automated testing. The role involves working with Java, Spring Boot, microservices, distributed systems, Kubernetes, cloud infrastructure, and CI/CD while collaborating with engineering leadership and cross-functional teams.
Top Skills:
AWSClaudeDebeziumDistributed SystemsDockerGitlab Ci/CdHelmJavaKafkaKubernetesMicroservicesMySQLOpentelemetryPostgresRestful ApisSnowflakeSpring BootSQLTemporalTerraform
Aerospace • Other
Responsible for developing and enforcing manufacturing, assembly, welding, NDE, inspection, and tooling processes for valves and aerospace components. Identify reliability gaps, partner with design and manufacturing, monitor production ramp-up, investigate failures, implement corrective actions, and document inspection and quality procedures while working hands-on in office, factory, and lab environments.
Top Skills:
ApqpCnc MachiningControl PlansDesign Of Experiments (Doe)Electron Beam (Eb) WeldingFluid System Pressure TestingLeanMeasurement Systems Analysis (Msa)MetrologyNon-Destructive Evaluation (Nde)PfmeaStatistical Process Control (Spc)Tig WeldingTooling Design
Renewable Energy
Own reliability, performance, and scalability of Postgres and ClickHouse databases. Build scalable data pipelines, design analytical schemas and DBT models, migrate data to ClickHouse, implement data quality checks, eliminate duplicates, and manage database infrastructure via IaC.
Top Skills:
Aws CdkAws Step FunctionsCi/CdClickhouseDagsterDbtPostgresPulumiPythonSQL
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
Software
As a Senior DevOps / Platform Reliability Engineer, you will manage CI/CD pipelines, automate infrastructure, operate Kubernetes, and enhance observability while ensuring security and compliance for enterprise systems.
Top Skills:
Argo CdAurora MysqlAWSBashCloudFormationEksElasticacheGithub ActionsGrafanaKubernetesLinuxMskOpentelemetryPrometheusPythonS3Terraform
Aerospace • Other
Develop highly reliable software used across SpaceX for launch vehicle production, flight operations, and Starlink growth. Responsibilities include designing and building applications, prototyping solutions, owning software and product development, investigating user problems, and contributing to architecture, design, and code reviews. The role requires onsite work and may involve extended hours or weekends based on launch schedules.
Top Skills:
.NetAgentic AiAngularC#Context RetrievalContinuous DeliveryContinuous IntegrationDockerGoJavaJavaScriptLarge Language ModelsPostgresPythonReactScalaSQL ServerVersion Control
Defense • Manufacturing
The Senior Design Reliability Engineer will develop design review processes, ensure compliance with reliability standards, and manage project documentation for spacecraft design.
Top Skills:
Aerospace SystemsLaunch VehiclesProject ManagementSpacecraft
Defense • Manufacturing
The Principal Design Reliability Engineer ensures spacecraft designs meet rigorous reliability standards, overseeing design reviews, configuration management, and documentation for aerospace engineering. They lead cross-disciplinary teams, drive process improvements, and ensure smooth transitions to production while capturing lessons learned for continuous improvement.
Top Skills:
Aerospace EngineeringComputer EngineeringElectrical EngineeringMechanical Engineering
Aerospace • Other
The PCB Reliability Engineer will design PCB stackups, manage reliability campaigns, optimize product quality, and collaborate with engineering teams to improve manufacturability for SpaceX's satellites.
Top Skills:
Cross-Section)Failure Analysis Tools (X-RayIpc Manufacturing StandardsPcb DesignPcb Fabrication
Marketing Tech
Design, deploy, and maintain cloud systems and tools; containerize applications; build CI/CD pipelines; instrument observability; support dev and ops teams; participate in on-call rotation; provide scoping/LOE and debug developer code as needed.
Top Skills:
AWSAws LambdaDockerGithub ActionsGoGoogle BigqueryGCPGoogle Cloud FunctionsKubernetesLinuxPythonServerlessSQLTerraform
Marketing Tech
Design, build, and operate cloud infrastructure and tools: containerize applications, implement Terraform and CI/CD, support dev and ops, ensure observability, and participate in on-call rotation.
Top Skills:
AWSAws LambdaContainerizationDockerGithub ActionsGoGoogle BigqueryGCPGoogle Cloud FunctionsKubernetesLinuxPythonServerlessSQLTerraform
Healthtech • Software
Operate and maintain AWS-hosted MERN applications and large-scale data workflows. Manage serverless and Spark-based pipelines, perform incident response and on-call duties, engineer automation to eliminate operational toil, ensure HIPAA/SOC2/HITRUST compliance, build observability and lead blameless post-mortems.
Top Skills:
Amazon EcsAmazon EksAmazon EmrAthenaAws GlueAws LambdaAws SnsAws SqsCloudwatchEc2IamJavaScriptMernMySQLNode.jsOpentofuPysparkPythonRabbitMQTerraformTypescriptVpc
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Senior SRE owning availability, automation, and observability for CI/CD platform services. Build and operate infrastructure, run on-call, lead incident response, mentor engineers, drive design/capacity planning, integrate AI-assisted workflows, and improve cross-team reliability.
Top Skills:
Active DirectoryAnsibleApache AirflowSparkAWSAzureBashBazelBitbucketCassandraChefDatadogDnsFirewall RulesGCPGitGithub ActionsGitlabGitlab CiGoGrafanaHoneycombHumio/LogscaleJenkinsKafkaKubernetesLoad BalancersMongoDBMySQLNasNew RelicNfsObject StorageOpensearchOraclePostgresPowershellPrometheusPulsarPuppetPythonRabbitMQRedis/ValkeyRedpandaRoutingSaltSanSplunkTerraformVarnishVipsWindows Server
Fintech • Financial Services
Operates and improves high-traffic, business-critical cloud and network systems. Designs scalable network solutions, manages performance and troubleshooting, automates deployment and monitoring, coordinates router and switch installations, and tests redundancy, resilience, and failover. Partners with development teams to improve operability, supports project planning and vendor comparisons, provides training, and participates in on-call coverage.
Top Skills:
Cloud ComputingIpRoutersSwitchesVoip
Automotive • Information Technology • Other • Transportation • Energy
Perform RAM and FMECA/FMEA analyses, develop fault trees and reliability predictions, support maintainability and logistics analyses, produce reliability growth test plans, contribute to systems engineering documentation, advise design engineers on R&M shortfalls, and present results to management and clients.
Top Skills:
Fault Tree AnalysisFmeaFmecaIntegrated Logistics Support (Ils/Ilsa)Iso-9000Mil-Hdbk-217FRam ModellingRam SoftwareStatistical Methods
Artificial Intelligence • Cloud • Information Technology • Software
The Site Reliability Engineer will provision and manage Kubernetes clusters, build automation tools, debug customer issues, and improve infrastructure reliability.
Top Skills:
AnsibleBashDatadogGoGrafanaHelmKubernetesLokiPrometheusPythonTerraform
Artificial Intelligence • Insurance • Software • Automation
Lead design, automation, and optimization of database infrastructure (PostgreSQL/Aurora). Build monitoring, tuning, and scaling strategies, create automation tooling, drive performance and reliability initiatives, and expand into broader SRE responsibilities to improve availability and system health for a growing SaaS platform.
Top Skills:
Amazon AuroraCi/CdDockerJavaScriptKubernetesNode.jsPostgresPrismaRedshiftTerraformTerragruntTypescript
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Los Angeles, CA Companies Hiring Reliability Engineers
See AllPopular Los Angeles, CA Engineering Job Searches
Engineering Jobs in Los Angeles, CA
Software Engineer Jobs in Los Angeles, CA
Android Developer Jobs in Los Angeles, CA
C# Jobs in Los Angeles, CA
C++ Jobs in Los Angeles, CA
DevOps Jobs in Los Angeles, CA
Front End Developer Jobs in Los Angeles, CA
Golang Jobs in Los Angeles, CA
Hardware Engineer Jobs in Los Angeles, CA
iOS Developer Jobs in Los Angeles, CA
Java Developer Jobs in Los Angeles, CA
Javascript Jobs in Los Angeles, CA
Linux Jobs in Los Angeles, CA
Engineering Manager Jobs in Los Angeles, CA
.NET Developer Jobs in Los Angeles, CA
PHP Developer Jobs in Los Angeles, CA
Python Jobs in Los Angeles, CA
QA Jobs in Los Angeles, CA
Ruby Jobs in Los Angeles, CA
Salesforce Developer Jobs in Los Angeles, CA
Scala Jobs in Los Angeles, CA
Application Engineer Jobs in Los Angeles, CA
Associate Software Engineer Jobs in Los Angeles, CA
Automation Engineer Jobs in Los Angeles, CA
AWS Engineer Jobs in Los Angeles, CA
Backend Engineer Jobs in Los Angeles, CA
Cloud Engineer Jobs in Los Angeles, CA
Controls Engineer Jobs in Los Angeles, CA
CTO Jobs in Los Angeles, CA
Design Engineer Jobs in Los Angeles, CA
DevOps Engineer Jobs in Los Angeles, CA
Director of Engineering Jobs in Los Angeles, CA
Electrical Engineering Jobs in Los Angeles, CA
Embedded Software Engineer Jobs in Los Angeles, CA
Field Engineer Jobs in Los Angeles, CA
Firmware Engineer Jobs in Los Angeles, CA
Full-Stack Engineer Jobs in Los Angeles, CA
Game Engineer Jobs in Los Angeles, CA
Industrial Engineer Jobs in Los Angeles, CA
Infrastructure Engineer Jobs in Los Angeles, CA
Manufacturing Engineer Jobs in Los Angeles, CA
Mechanical Design Engineer Jobs in Los Angeles, CA
Mechanical Engineering Jobs in Los Angeles, CA
Mechatronics Engineering Jobs in Los Angeles, CA
Network Engineer Jobs in Los Angeles, CA
Platform Engineer Jobs in Los Angeles, CA
Principal Engineer Jobs in Los Angeles, CA
Principal Software Engineer Jobs in Los Angeles, CA
Process Engineer Jobs in Los Angeles, CA
Product Engineer Jobs in Los Angeles, CA
Project Engineer Jobs in Los Angeles, CA
QA Engineer Jobs in Los Angeles, CA
Reliability Engineer Jobs in Los Angeles, CA
Robotics Engineer Jobs in Los Angeles, CA
Software Architect Jobs in Los Angeles, CA
Software Engineering Manager Jobs in Los Angeles, CA
Solutions Engineer Jobs in Los Angeles, CA
SRE Engineer Jobs in Los Angeles, CA
Staff Engineer Jobs in Los Angeles, CA
Staff Software Engineer Jobs in Los Angeles, CA
Structural Engineer Jobs in Los Angeles, CA
Systems Engineer Jobs in Los Angeles, CA
VP of Engineering Jobs in Los Angeles, CA
Web Developer Jobs in Los Angeles, CA
All Filters
Total selected ()
No Results
No Results




























