Maximum of 25 job preferences reached.
Top Remote Site Reliability Engineer Jobs in Los Angeles, CA
Reposted 22 Hours AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
The Senior Site Reliability Engineer will develop and support distributed storage services, ensuring reliability and operational safety, with a focus on automation and efficiency.
Top Skills:
AWSAzureDnsGoGoogle Cloud PlatformKubernetesLinuxPythonTcp/IpTls
Artificial Intelligence • Machine Learning
Lead development of AI-assisted reliability tooling, own incident response end-to-end, improve observability and SLO/SLI frameworks, scale single-tenant SaaS operations, mentor engineers, and reduce recurring operational toil through engineering and automation.
Top Skills:
Cloud PlatformsGoKubernetesLinuxLlm/Ai ToolingLogs And TracingObservability ToolingPythonSlo/Sli Frameworks
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Ensure stability and resilience of Runpod's distributed AI platform by defining SLIs/SLOs, leading incident response, building observability and reliability tooling, automating operational workflows, and partnering with engineering teams to reduce toil and improve production readiness.
Top Skills:
BashCi/CdContainerized Production SystemsGoGpu Observability ToolingGrafanaInfrastructure As CodeLinuxPrometheusPython
Reposted 10 Days AgoSaved
Easy Apply
Easy Apply
Cloud • Information Technology • Security • Software • Cybersecurity
This internship role focuses on SRE skills, requiring collaboration and problem-solving in dynamic environments for Zscaler's Zero Trust Exchange team.
Top Skills:
AnsibleAws EcsKubernetesLinuxPythonTerraform
Reposted 12 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
The Senior Site Reliability Engineer will lead security design and implementation for cloud infrastructures, mentor teams, and automate security solutions.
Top Skills:
AnsibleAWSAzureCloud Security ToolsCloudFormationGCPGoTerraform
Healthtech • Software
Operate and maintain AWS-hosted MERN applications and large-scale data workflows. Manage serverless and Spark-based pipelines, perform incident response and on-call duties, engineer automation to eliminate operational toil, ensure HIPAA/SOC2/HITRUST compliance, build observability and lead blameless post-mortems.
Top Skills:
Amazon EcsAmazon EksAmazon EmrAthenaAws GlueAws LambdaAws SnsAws SqsCloudwatchEc2IamJavaScriptMernMySQLNode.jsOpentofuPysparkPythonRabbitMQTerraformTypescriptVpc
Reposted 17 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
As a Senior Site Reliability Engineer, you'll design and build complex systems, support Atlas platform operations, automate processes, and ensure high availability of services.
Top Skills:
AWSAzureDnsGCPGoHTTPLinuxPythonRubyTls
Aerospace • Hardware • Software • Defense • Manufacturing
As a Site Reliability Engineer, you'll ensure robotics system reliability, build telemetry integration, and develop tools for diagnostics and automation, collaborating with engineering teams for enhanced production reliability.
Top Skills:
C++DatadogGoKubernetesOpentelemetryPrometheusPythonRos2TelegrafTypescript
eCommerce • Fintech • Payments • Software
The role involves ensuring software reliability and performance, managing incidents, developing infrastructure automation, and mentoring junior engineers within a platform team.
Top Skills:
AWSCloudFormationDatadogKubernetesOpentelemetryRubyRuby On RailsTerraform
Social Media • Software
Design, implement, and operate infrastructure for a federated social network. Own reliability, availability, observability, incident response, deployments, capacity planning, and cost management. Build automation and tooling, scale bare-metal and cloud systems for millions of users, lead incident reviews, mentor engineers, and manage vendor relationships to ensure operational excellence.
Top Skills:
Bare-MetalCapacity PlanningCloud ServicesColocationDatabasesDeployment And Rollback SystemsGoIncident ResponseKubernetesLinuxMonitoringNetworkingObservability SystemsProduction AutomationStorage
Artificial Intelligence • Cloud • Information Technology • Consulting
Internship SRE role responsible for availability, performance, and scalability of an e-commerce supply-chain platform. Tasks include SLO/SLA definition, observability (Prometheus/Grafana/Loki/Tempo/OpenTelemetry), incident response, capacity planning, disaster recovery for PostgreSQL, infrastructure-as-code (Terraform), CI/CD automation, and operational reliability for AI agent services. Mentored by Head of Technology/CTO with potential conversion to full-time based on performance.
Top Skills:
BashCi/CdDockerGrafanaLangchainLlmLokiMakefileNestjsOpentelemetryOracle CloudPgbackrestPostgresql 15PrometheusPythonRedisTempoTerraformTraefik
Digital Media • Social Media • Software • Sports
Lead the technical architecture and execution of migration to AWS, drive developer enablement, and automate infrastructure using code-first principles.
Top Skills:
Aws EksDatadogGithub ActionsGoIstioK6KubernetesNode.jsTerraform
New
Cut your apply time in half.
Use ourAI Assistantto automatically fill your job applications.
Use For Free
Artificial Intelligence • Marketing Tech • Mobile • Software
Lead design and implementation of scalable, reliable platform systems; define SLIs/SLOs and observability; drive cross-team strategic initiatives; mentor engineers; own production standards, incident management, and cost/operational optimization to improve platform reliability and scalability.
Top Skills:
GoJavaPythonTypescript
Healthtech • Pharmaceutical • Manufacturing
Support operational deployments and maintenance for Core Speech production systems. Implement and maintain monitoring, alerting, performance reporting, and capacity planning. Participate in on-call rotation for off-hours reliability support. Drive improvements to system performance and architecture, and collaborate on projects while complying with corporate and quality standards.
Information Technology
Design, build, and operate a reliable, scalable developer platform and cloud infrastructure (GCP/GKE). Lead SRE practices: SLO/SLI, observability, incident response, on-call, automation, security-by-default, Terraform IaC, CI/CD, capacity planning, mentoring, and platform enablement across teams.
Top Skills:
Ci/CdConfluenceDatadogDockerGCPGitGithub ActionsGkeGoGrafanaGsm (Gcp Secrets Manager)HclHelmHoneycombJIRAKubernetesNew RelicNode.jsOpentelemetryPrometheusPythonSslTerraform
Cloud • Security • Software
Design, build, and operate cloud infrastructure and CI/CD pipelines for a large identity platform. Implement automated deployments, ensure resiliency, observability, security, and cost optimization. Collaborate across teams, perform technology evaluations, participate in on-call rotation, and mentor peers.
Top Skills:
AWSCi/CdDockerGCPGitGoKubernetes
Healthtech • Insurance
Lead cloud, DevOps, and SRE architecture efforts to scale CareSource's digital platform. Partner with Cloud, DevOps, Security, and SRE teams to design infrastructure, build Terraform templates, enhance CI/CD pipelines, implement monitoring/alerting, and improve reliability, scalability, and incident response for enterprise-scale digital products.
Top Skills:
Azure CloudDockerDynatraceGithub ActionsKubernetesSplunkTerraform Enterprise
Software • Financial Services
Ensure platform reliability, performance, and availability by implementing observability, automating infrastructure, participating in on-call rotations and post-mortems, partnering with Product and Engineering, designing scalable architectures, mentoring teammates, and integrating Dynatrace with Azure DevOps and Jira while supporting compliance (SOC/FedRAMP).
Top Skills:
.NetAksAlpineAnsibleAppinsightsArm TemplatesAWSAzure DevopsBashBicepC#ChefCloudFormationDatadogDebianDynatraceEksGCPGitGitGksGrafanaHelmJIRAKubernetesLog AnalyticsAzureNew RelicOnestream SoftwareOpenshiftPowershellPowershell DscPrometheusPuppetPythonRest ApisSQLTerraformUbuntu
3D Printing • Artificial Intelligence • Software • Design
Lead design and operation of scalable, multi-tenant spatial streaming platforms. Build Terraform-based cloud infrastructure, optimize CDN/content delivery, implement observability (SLI/SLO), run incident response/on-call, conduct post-mortems, enforce compliance and security practices, and mentor DevOps engineers to improve reliability and production readiness.
Top Skills:
Aws FargateCdnCoreweaveGrafanaKubernetesPrometheusTerraform
Gaming • Software
The Site Reliability Engineer will manage infrastructure stability and scalability, lead cloud migrations, and optimize performance across systems while mentoring team members.
Top Skills:
AnsibleAWSAzureBashChefCloudFormationDatadogDockerElk StackGCPGoGrafanaKubernetesPrometheusPuppetPythonTerraformUnix/Linux
Software
As a Site Reliability Engineer, you'll enhance system reliability, collaborate on production readiness, define SLIs/SLOs, and improve incident response.
Top Skills:
AWSDatadogGrafanaKubernetesOpentelemetryPrometheusTypescript
Information Technology • Cybersecurity
Lead global SRE team to design, implement, and operate scalable, highly available cloud infrastructure. Drive observability, automation (IaC/CI-CD), cost optimization (FinOps), security hygiene, and post-incident improvements while remaining hands-on and exploring AI tooling for reliability.
Top Skills:
Ai ToolingAlloyAWSAzureCi/CdCloudFormationContainer OrchestrationDatadogEdge ComputingFinopsGCPGrafanaGrafana IrnInfrastructure-As-CodeKubernetesLokiPrometheusPulumiServerlessSplunkSpot InstancesTerraform
Cloud • Security • Software • Cybersecurity
Ensure reliability, scalability, and usability of network infrastructure for Akamai Connected Cloud. Define requirements and SLOs, build automation and CI/CD pipelines, collaborate with dev/QA to improve code and stability, troubleshoot complex network issues (on-call), and mentor teammates while driving architectural standards.
Top Skills:
AnsibleArgocdBashBirdChefFrrGithub ActionsGoGobgpJenkinsLinux NetworkingPuppetPythonSalt Stack
Software
Lead SRE to define strategy and roadmap for reliability, scalability, observability, and automation across cloud and hybrid environments. Design and operate containerized production workloads, infrastructure-as-code, monitoring/alerting, incident management, and compliance for regulated domains. Mentor SREs, partner with security and product teams, manage cloud costs and capacity, and build a developer platform to improve delivery and production stability.
Top Skills:
AWSAws MarketplaceAzureAzure MarketplaceGCPGoogle Cloud MarketplaceGrafanaKubernetesPrometheusTerraform
Legal Tech • Software
Lead observability and incident management efforts: define SLIs/SLOs, build monitoring/alerting, dashboards, logging, and tracing. Drive incident response, postmortems, and reliability improvements to reduce MTTD/MTTR. Integrate observability into CI/CD, maintain AWS and Kubernetes infrastructure, automate operations, and mentor engineers on SRE best practices.
Top Skills:
AWSBashCi/CdDatadogDistributed TracingDynatraceGrafanaKubernetesNew RelicOpentelemetryPowershellPrometheusPython
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Popular Job Searches
All Filters
Total selected ()
No Results
No Results






.png)

























