Maximum of 25 job preferences reached.
Top Remote Site Reliability Engineer Jobs in Los Angeles, CA
Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Leads architecture, modernization, optimization, and reliability initiatives for mainframe CICS, MQ, and z/OS Connect environments. Provides technical direction across development and operations teams, establishes governance and change processes, tunes performance using telemetry, resolves incidents, and develops modernization roadmaps. Collaborates with stakeholders and enterprise architects to deliver secure, scalable, high-availability solutions while evaluating automation, cloud integration, and AI technologies.
Top Skills:
AnsibleCicsCobolDevOpsIbm MqIbm Z/OsOpenshiftPythonRed Hat Ansible Automation PlatformZ/Os ConnectZlinux
17 Days AgoSaved
Easy Apply
Easy Apply
Cloud • Security • Software • Cybersecurity • Automation
Provide technical direction for GitLab Dedicated, a managed single-tenant SaaS platform. Lead architecture and transformation across resilience, failover, tenant orchestration, change management, automation, and platform integrations. Identify systemic reliability and scalability risks, establish reusable platform patterns, strengthen service ownership, and guide cross-team technical decisions. Mentor senior engineers and advance engineering excellence across the organization.
Top Skills:
Cloud InfrastructureDevsecopsDistributed SystemsGoInfrastructure As CodeObservabilityPythonRuby
Fintech • Financial Services
Develop and maintain reliable, scalable, and high-performing systems through automated deployments, monitoring, observability, and incident response. The role partners with development, operations, and QA teams to improve software delivery, establish reliability processes, track KPIs and SLIs, optimize database performance, and manage messaging systems. Participation in on-call rotations and occasional travel for team meetings are required.
Top Skills:
AlertingAutomated DeploymentAWSAzureBashEvent Bus QueuesGCPGroovyLoggingMessaging QueuesMetricsMonitoringNoSQLObservabilityPowershellSQL
Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3 • Infrastructure as a Service (IaaS)
Manage AWS and GCP cloud environments, scale infrastructure globally, shape technical architecture, and build reliable CI/CD pipelines. Automate security and compliance controls, improve infrastructure performance, and develop systems interacting with smart contracts across multiple blockchains. The role requires Terraform, shell scripting, GitHub Actions, Docker, production cloud operations, and observability experience, with Kubernetes, networking, and fintech compliance knowledge preferred.
Top Skills:
AWSCi/CdDockerFirewallsGCPGithub ActionsGkeGoHelmInfrastructure As CodeKubernetesLoad BalancersMtlsNode.jsPciShellSsl/TlsTerraformTypescriptVpcZero Trust
Software
Own the operational health, observability, reliability, and failover of AI model providers and endpoints. Build SLO monitoring, canaries, quality regression detection, automated provider-operations tooling, and load-testing systems. Lead incident response, communicate with external providers, produce reliability scorecards, and drive postmortems and corrective actions. Support capacity planning and launch readiness for high-volume LLM inference traffic.
Top Skills:
ClickhouseCloudflare WorkersGCPPostgresPythonTypescriptVercel
eCommerce
Own the reliability, availability, security, and observability of Tradeweb’s global AWS platform. Responsibilities include infrastructure-as-code automation, monitoring, SLO development, incident triage and resolution, performance analysis, architecture collaboration, and regular on-call support. The role requires cloud-native engineering, scripting, networking, Linux/Unix, security, and reliability expertise while contributing to an Agile engineering organization.
Top Skills:
Amazon EksAmazon SmsAmazon SnsArgocdAWSAws LambdaGitsecopsKubernetesKustomizeLgtmLinuxPulumiPythonUnix
Artificial Intelligence • Cloud • Information Technology • Consulting
Internship SRE role responsible for availability, performance, and scalability of an e-commerce supply-chain platform. Tasks include SLO/SLA definition, observability (Prometheus/Grafana/Loki/Tempo/OpenTelemetry), incident response, capacity planning, disaster recovery for PostgreSQL, infrastructure-as-code (Terraform), CI/CD automation, and operational reliability for AI agent services. Mentored by Head of Technology/CTO with potential conversion to full-time based on performance.
Top Skills:
BashCi/CdDockerGrafanaLangchainLlmLokiMakefileNestjsOpentelemetryOracle CloudPgbackrestPostgresql 15PrometheusPythonRedisTempoTerraformTraefik
Big Data • Software
Owns production reliability for a B2B SaaS platform, including SLOs, error budgets, observability, alerting, incident response, recovery exercises, performance and capacity engineering, load testing, disaster recovery, and production readiness. The role writes automation and production code, leads incidents, coaches engineering teams, and improves reliability across Kubernetes, Azure, databases, messaging, and service infrastructure.
Top Skills:
.NetAksAsp.NetAWSAzureAzure Service BusC#Dora MetricsFluxGatlingGitopsGoGrafanaIncident.IoIstioJmeterK6KafkaKubernetesLokiMongoDBNist 800-171PagerdutyPrometheusPythonSentrySignalrSoc 2TempoTemporalTypescript
Cloud • Security • Software • Cybersecurity
Performs site reliability engineering for large-scale cloud infrastructure, focusing on application and network performance, reliability, security, scalability, and capacity. Deploys highly available systems, automates cloud service deployments, monitors and troubleshoots services, analyzes logs and events, maintains SLAs, and resolves infrastructure issues. The role requires expertise in microservices, container orchestration, cloud migration, deployment automation, and multiple DevOps technologies.
Top Skills:
AzureChefCloud ComputingContainer OrchestrationGitIntellij IdeaJenkinsKubernetesLinuxMicroservicesMySQLPostgresPythonTerraformVMware
Cloud • Security • Software • Cybersecurity
Analyzes and resolves availability and performance issues in large-scale production and lab environments. Develops automation, monitoring, alerting, log analysis, debugging tools, and systems programming solutions. Collaborates with software development and engineering teams on CI/CD, platform architecture, incident resolution, and operational best practices within an agile SDLC.
Top Skills:
AlertingAnsibleAutomation ScriptingBashContinuous DeliveryContinuous IntegrationDevOpsLinuxLog AnalysisMonitoringPowershellSystems Programming
Software
Own production reliability and infrastructure operations for the GrayKey cloud platform. Responsibilities include AWS monitoring, incident response, vulnerability remediation, networking, IAM and Okta administration, EKS and Argo CD operations, backups, recovery testing, Terraform changes, cost optimization, automation, documentation, and compliance support. The role requires independent work during Pacific and Mountain time-zone hours and involves less than 5% travel.
Top Skills:
Active DirectoryAmazon EksAnsibleArgo CdAWSAzure AdBashDatadogDnsDockerEc2Entra IdFortigateGitlab Ci/CdHelmIamKubernetesLambdaLinuxOidcOktaPostgresPythonRdsRedisSAMLTcp/IpTerraformTlsVpcVpn
Artificial Intelligence • Information Technology • Consulting
The Linux Systems Administrator will maintain and troubleshoot Linux systems, support network services, and work on systems integration while collaborating with infrastructure teams.
Top Skills:
DhcpDnsLinuxNtpPython
New
Cut your apply time in half.
Use ourAI Assistantto automatically fill your job applications.
Use For Free
Digital Media • Gaming • Information Technology • Software • Sports • Esports • Big Data Analytics
Lead reliability, scalability, and operational excellence of large-scale database platforms across cloud and on-prem. Build automation-first database infrastructure (Kubernetes operators, IaC, GitOps), drive monitoring/SLOs, incident leadership, performance and cost optimization, and partner with application teams on safe schema/migration practices. Mentor engineers and evaluate AI-assisted workflows to improve productivity and reliability.
Top Skills:
AerospikeArgocdAuroraClaudeCloud SqlCursorDatabase OperatorsEksFluxcdGithub CopilotGitopsGkeGoKubernetesMcpMongoDBMySQLPersistent VolumesPostgresPulumiPythonRedisScylladbStatefulsetsTerraform
Information Technology
Leads SRE and cloud operations for highly available, scalable platforms and AI-powered solutions. Designs AWS infrastructure, Kubernetes environments, CI/CD pipelines, Infrastructure as Code, observability, automation, and incident response processes. Develops AI-driven operational capabilities, supports MLOps and model lifecycle management, establishes reliability metrics and SLOs, and mentors engineering teams. Partners with software, data, machine learning, security, and product teams to improve platform performance, resilience, and operational excellence.
Top Skills:
Amazon SagemakerAWSAws BedrockAzure DevopsAzure OpenaiCi/CdCloudFormationCloudwatchDatadogDockerEc2EcsEksElk StackGithub ActionsGitlab Ci/CdGrafanaIamInfrastructure As CodeJenkinsKubernetesLambdaLangchainMlopsNvidia AiOpenai ApisPrometheusPythonRdsS3ShellSplunkTerraformVpc
One Month AgoSaved
Easy Apply
Easy Apply
Cloud • Security • Software • Cybersecurity • Automation
Build and operate reliable, scalable production infrastructure for GitLab’s user-facing services. Responsibilities include developing infrastructure automation and tooling, managing Kubernetes deployments, maintaining infrastructure as code, supporting CI/CD and GitOps, participating in on-call and incident response, improving observability and SLOs, troubleshooting production systems, and documenting operational practices. The role spans Intermediate through Senior Staff levels and requires strong software engineering, cloud, reliability, and asynchronous collaboration skills.
Top Skills:
AlertingAWSCi/CdGCPGitopsGoInfrastructure As CodeKubernetesLoggingMetricsRubySlisSlosTerraform
Social Media • Software
Design, implement, and operate infrastructure for high-scale production systems serving millions of users. Own reliability, observability, incident response, deployments, capacity planning, cost management, and operational excellence across bare-metal and cloud environments. Develop automation and performance tooling, improve production readiness, manage vendors, lead incident reviews, and mentor engineers on reliability and distributed-systems practices.
Top Skills:
Bare-Metal ServersCloud ServicesDatabasesDistributed SystemsGoKubernetesLinuxNetworkingObservability SystemsStorage Systems
Edtech • Professional Services
The Site Reliability Engineer will operate and improve critical, high-traffic systems across AWS and Azure. Responsibilities include infrastructure-as-code, automation, incident response, root-cause analysis, monitoring, alerting, disaster recovery, capacity planning, performance tuning, high availability, cloud security, CI/CD, and multi-cloud networking. The role also involves mentoring, documentation, regular DR testing, and on-call participation.
Top Skills:
ArmAWSAws CdkAws Direct ConnectAzureAzure DevopsAzure ExpressrouteBashBicepCi/CdCircleCICloudFormationDatadogEc2Github ActionsIamInfrastructure As CodeKubernetesOwaspPagerdutyPowershellPythonRdsS3SnykSonarqubeTerraformVpcVpn
Other
Build and operate AWS GovCloud environments within FedRAMP boundaries, including multi-account infrastructure, guardrails, logging, security tooling, and Terraform provisioning. Develop automated fleet release, patching, verification, CI/CD, compliance evidence, telemetry, SLO, and alerting systems. Support VM and appliance fleets at scale, participate in on-call, and engineer recurring incident causes away. The role also requires compliance-as-code, continuous monitoring, and test harnesses validating operational state.
Top Skills:
AlertingArtifact ManagementAws ConfigAws GovcloudCi/CdDashboardsFedrampFipsGrafanaKubernetesOpaOpentelemetryOscalSecurity ScanningSlosTelemetryTerraform
Cloud • Security • Software • Cybersecurity
Ensure reliability, scalability, and usability of network infrastructure for Akamai Connected Cloud. Define requirements and SLOs, build automation and CI/CD pipelines, collaborate with dev/QA to improve code and stability, troubleshoot complex network issues (on-call), and mentor teammates while driving architectural standards.
Top Skills:
AnsibleArgocdBashBirdChefFrrGithub ActionsGoGobgpJenkinsLinux NetworkingPuppetPythonSalt Stack
Artificial Intelligence • Information Technology • Consulting
Lead end-to-end deployment and operation of cloud infrastructure and Symmetri environments for US customers. Provision, configure, monitor, maintain, upgrade, and troubleshoot cloud, single-tenant, on-premises, and air-gapped environments. Manage Kubernetes, networking, IAM, infrastructure as code, resource planning, incident response, runbooks, and root-cause analysis. Represent the company in customer security reviews and help establish US deployment and operations standards and teams.
Top Skills:
AWSAws CdkAzureDnsGitIamKubectlKubernetesLoad BalancingOpentofuPulumiPythonTerraformTerragruntVnetsVpc
Aerospace • Information Technology • Security • Cybersecurity • Defense
Designs and develops Python-based AWS cloud solutions, reliability tools, automation, observability, and Terraform infrastructure. Builds CI/CD pipelines, defines SRE standards and SLOs, supports incident management, and improves platform reliability through monitoring, postmortems, testing, and proactive toil reduction. Collaborates with cross-functional Agile teams to deliver scalable cloud automation across the Federal Reserve System.
Top Skills:
Automated TestingAWSAws CanaryAws CloudwatchCi/CdCloudFormationEc2EventbridgeGoGrafanaIamInfrastructure As Code (Iac)LambdaPythonS3Step FunctionsTerraformVpc
Mobile
Serve as Vida’s first dedicated Site Reliability Engineer, modernizing Terraform, standardizing environments, improving CI/CD, scaling GCP and Kubernetes infrastructure, upgrading databases and runtimes, and strengthening monitoring, observability, security, and operational processes. The role includes building runbooks, on-call and escalation procedures, supporting enterprise launches, managing infrastructure costs, and evaluating multi-cluster Kubernetes architecture in a fully remote environment.
Top Skills:
AirflowCloud MonitoringCloud SqlDatadogDjangoFirestoreGCPGithub ActionsGkeIamKubernetesMySQLPostgresPythonRedisTerraform
Professional Services • Consulting
Own and improve CI/CD, developer tooling, agent harnesses, Kubernetes infrastructure, observability, cost management, and incident-response practices. Build platform capabilities that reduce engineering toil and improve delivery speed for conversational AI products. Participate in an on-call rotation for core infrastructure and help shape reliable, scalable systems across a fully remote engineering organization.
Top Skills:
AlertingCi/CdHelmIncident ManagementKubernetesLlmsLogsMetricsMonitoringNode.jsObservabilityPythonTerraformTracingTypescript
Cloud • Information Technology
Own the reliability, scalability, and administration of production Vitess/MySQL and Cassandra databases. Design highly available architectures, replication, backup, recovery, security, and disaster recovery strategies. Lead incident response, on-call operations, monitoring, SLO development, root cause analysis, and production readiness reviews. Automate database and infrastructure operations using scripting, Kubernetes, Terraform, Ansible, and Jenkins. Establish runbooks, escalation procedures, training materials, and mentor junior database SREs while partnering across engineering and infrastructure teams.
Top Skills:
AnsibleAWSAzureBashCassandraCatchpointDockerElkFirehydrantGCPGoGrafanaItilJenkinsKubectlKubernetesLinuxMySQLMysqlshNomadOssPrometheusPythonSQLTerraformVaultVitess
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Build and operate software automation for large-scale NVIDIA GPU infrastructure. Responsibilities include bare-metal provisioning, hardware validation, firmware and software upgrades, cluster lifecycle management, BMC and Redfish tooling, failure diagnosis, observability, incident response, and automated repair. The role supports NVL72 systems, BlueField DPUs, Linux, Kubernetes, and cloud or on-premises environments while coordinating across hardware, networking, platform, data center, and partner teams.
Top Skills:
Argo CdBluefield-3 DpusBmcDriver ManagementFirmware ManagementGitopsGoInfinibandKubernetesLinuxNetwork BootNvidia Nvl72NvlinkPythonRedfishSlosSpectrum-X
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Los Angeles, CA Companies Hiring Remote Site Reliability Engineers
See AllPopular Job Searches
All Filters
Total selected ()
No Results
No Results














_1.png)















