Top Remote Site Reliability Engineer Jobs in Los Angeles, CA

16 Days AgoSaved
Remote or Hybrid
United States
111K-180K Annually
Senior level
111K-180K Annually
Senior level
Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Leads architecture, modernization, optimization, and reliability initiatives for mainframe CICS, MQ, and z/OS Connect environments. Provides technical direction across development and operations teams, establishes governance and change processes, tunes performance using telemetry, resolves incidents, and develops modernization roadmaps. Collaborates with stakeholders and enterprise architects to deliver secure, scalable, high-availability solutions while evaluating automation, cloud integration, and AI technologies.
Top Skills: AnsibleCicsCobolDevOpsIbm MqIbm Z/OsOpenshiftPythonRed Hat Ansible Automation PlatformZ/Os ConnectZlinux
17 Days AgoSaved
Easy Apply
Remote
United States
Easy Apply
223K-380K Annually
Expert/Leader
223K-380K Annually
Expert/Leader
Cloud • Security • Software • Cybersecurity • Automation
Provide technical direction for GitLab Dedicated, a managed single-tenant SaaS platform. Lead architecture and transformation across resilience, failover, tenant orchestration, change management, automation, and platform integrations. Identify systemic reliability and scalability risks, establish reusable platform patterns, strengthen service ownership, and guide cross-team technical decisions. Mentor senior engineers and advance engineering excellence across the organization.
Top Skills: Cloud InfrastructureDevsecopsDistributed SystemsGoInfrastructure As CodeObservabilityPythonRuby
3 Days AgoSaved
In-Office or Remote
Los Angeles, CA, USA
Senior level
Senior level
Fintech • Financial Services
Develop and maintain reliable, scalable, and high-performing systems through automated deployments, monitoring, observability, and incident response. The role partners with development, operations, and QA teams to improve software delivery, establish reliability processes, track KPIs and SLIs, optimize database performance, and manage messaging systems. Participation in on-call rotations and occasional travel for team meetings are required.
Top Skills: AlertingAutomated DeploymentAWSAzureBashEvent Bus QueuesGCPGroovyLoggingMessaging QueuesMetricsMonitoringNoSQLObservabilityPowershellSQL
26 Days AgoSaved
Remote
USA
170K-225K Annually
Mid level
170K-225K Annually
Mid level
Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3 • Infrastructure as a Service (IaaS)
Manage AWS and GCP cloud environments, scale infrastructure globally, shape technical architecture, and build reliable CI/CD pipelines. Automate security and compliance controls, improve infrastructure performance, and develop systems interacting with smart contracts across multiple blockchains. The role requires Terraform, shell scripting, GitHub Actions, Docker, production cloud operations, and observability experience, with Kubernetes, networking, and fintech compliance knowledge preferred.
Top Skills: AWSCi/CdDockerFirewallsGCPGithub ActionsGkeGoHelmInfrastructure As CodeKubernetesLoad BalancersMtlsNode.jsPciShellSsl/TlsTerraformTypescriptVpcZero Trust
YesterdaySaved
Remote
US
Mid level
Mid level
Software
Own the operational health, observability, reliability, and failover of AI model providers and endpoints. Build SLO monitoring, canaries, quality regression detection, automated provider-operations tooling, and load-testing systems. Lead incident response, communicate with external providers, produce reliability scorecards, and drive postmortems and corrective actions. Support capacity planning and launch readiness for high-volume LLM inference traffic.
Top Skills: ClickhouseCloudflare WorkersGCPPostgresPythonTypescriptVercel
YesterdaySaved
Remote
United States
170K-210K Annually
Senior level
170K-210K Annually
Senior level
eCommerce
Own the reliability, availability, security, and observability of Tradeweb’s global AWS platform. Responsibilities include infrastructure-as-code automation, monitoring, SLO development, incident triage and resolution, performance analysis, architecture collaboration, and regular on-call support. The role requires cloud-native engineering, scripting, networking, Linux/Unix, security, and reliability expertise while contributing to an Agile engineering organization.
Top Skills: Amazon EksAmazon SmsAmazon SnsArgocdAWSAws LambdaGitsecopsKubernetesKustomizeLgtmLinuxPulumiPythonUnix
Reposted YesterdaySaved
Remote
United States
Internship
Internship
Artificial Intelligence • Cloud • Information Technology • Consulting
Internship SRE role responsible for availability, performance, and scalability of an e-commerce supply-chain platform. Tasks include SLO/SLA definition, observability (Prometheus/Grafana/Loki/Tempo/OpenTelemetry), incident response, capacity planning, disaster recovery for PostgreSQL, infrastructure-as-code (Terraform), CI/CD automation, and operational reliability for AI agent services. Mentored by Head of Technology/CTO with potential conversion to full-time based on performance.
Top Skills: BashCi/CdDockerGrafanaLangchainLlmLokiMakefileNestjsOpentelemetryOracle CloudPgbackrestPostgresql 15PrometheusPythonRedisTempoTerraformTraefik
Reposted YesterdaySaved
Remote
United States
Senior level
Senior level
Big Data • Software
Owns production reliability for a B2B SaaS platform, including SLOs, error budgets, observability, alerting, incident response, recovery exercises, performance and capacity engineering, load testing, disaster recovery, and production readiness. The role writes automation and production code, leads incidents, coaches engineering teams, and improves reliability across Kubernetes, Azure, databases, messaging, and service infrastructure.
Top Skills: .NetAksAsp.NetAWSAzureAzure Service BusC#Dora MetricsFluxGatlingGitopsGoGrafanaIncident.IoIstioJmeterK6KafkaKubernetesLokiMongoDBNist 800-171PagerdutyPrometheusPythonSentrySignalrSoc 2TempoTemporalTypescript
YesterdaySaved
In-Office or Remote
United States
138K-171K Annually
Mid level
138K-171K Annually
Mid level
Cloud • Security • Software • Cybersecurity
Performs site reliability engineering for large-scale cloud infrastructure, focusing on application and network performance, reliability, security, scalability, and capacity. Deploys highly available systems, automates cloud service deployments, monitors and troubleshoots services, analyzes logs and events, maintains SLAs, and resolves infrastructure issues. The role requires expertise in microservices, container orchestration, cloud migration, deployment automation, and multiple DevOps technologies.
Top Skills: AzureChefCloud ComputingContainer OrchestrationGitIntellij IdeaJenkinsKubernetesLinuxMicroservicesMySQLPostgresPythonTerraformVMware
YesterdaySaved
In-Office or Remote
United States
111K-171K Annually
Junior
111K-171K Annually
Junior
Cloud • Security • Software • Cybersecurity
Analyzes and resolves availability and performance issues in large-scale production and lab environments. Develops automation, monitoring, alerting, log analysis, debugging tools, and systems programming solutions. Collaborates with software development and engineering teams on CI/CD, platform architecture, incident resolution, and operational best practices within an agile SDLC.
Top Skills: AlertingAnsibleAutomation ScriptingBashContinuous DeliveryContinuous IntegrationDevOpsLinuxLog AnalysisMonitoringPowershellSystems Programming
YesterdaySaved
Remote
United States
90K-160K Annually
Mid level
90K-160K Annually
Mid level
Software
Own production reliability and infrastructure operations for the GrayKey cloud platform. Responsibilities include AWS monitoring, incident response, vulnerability remediation, networking, IAM and Okta administration, EKS and Argo CD operations, backups, recovery testing, Terraform changes, cost optimization, automation, documentation, and compliance support. The role requires independent work during Pacific and Mountain time-zone hours and involves less than 5% travel.
Top Skills: Active DirectoryAmazon EksAnsibleArgo CdAWSAzure AdBashDatadogDnsDockerEc2Entra IdFortigateGitlab Ci/CdHelmIamKubernetesLambdaLinuxOidcOktaPostgresPythonRdsRedisSAMLTcp/IpTerraformTlsVpcVpn
Reposted YesterdaySaved
Remote
United States
100K-140K Annually
Mid level
100K-140K Annually
Mid level
Artificial Intelligence • Information Technology • Consulting
The Linux Systems Administrator will maintain and troubleshoot Linux systems, support network services, and work on systems integration while collaborating with infrastructure teams.
Top Skills: DhcpDnsLinuxNtpPython
New

Cut your apply time in half.

Use ourAI Assistantto automatically fill your job applications.

Use For Free
Application Tracker Preview
Reposted One Month AgoSaved
Remote or Hybrid
United States
168K-210K Annually
Senior level
168K-210K Annually
Senior level
Digital Media • Gaming • Information Technology • Software • Sports • Esports • Big Data Analytics
Lead reliability, scalability, and operational excellence of large-scale database platforms across cloud and on-prem. Build automation-first database infrastructure (Kubernetes operators, IaC, GitOps), drive monitoring/SLOs, incident leadership, performance and cost optimization, and partner with application teams on safe schema/migration practices. Mentor engineers and evaluate AI-assisted workflows to improve productivity and reliability.
Top Skills: AerospikeArgocdAuroraClaudeCloud SqlCursorDatabase OperatorsEksFluxcdGithub CopilotGitopsGkeGoKubernetesMcpMongoDBMySQLPersistent VolumesPostgresPulumiPythonRedisScylladbStatefulsetsTerraform
2 Days AgoSaved
In-Office or Remote
United States
111K-145K Annually
Senior level
111K-145K Annually
Senior level
Information Technology
Leads SRE and cloud operations for highly available, scalable platforms and AI-powered solutions. Designs AWS infrastructure, Kubernetes environments, CI/CD pipelines, Infrastructure as Code, observability, automation, and incident response processes. Develops AI-driven operational capabilities, supports MLOps and model lifecycle management, establishes reliability metrics and SLOs, and mentors engineering teams. Partners with software, data, machine learning, security, and product teams to improve platform performance, resilience, and operational excellence.
Top Skills: Amazon SagemakerAWSAws BedrockAzure DevopsAzure OpenaiCi/CdCloudFormationCloudwatchDatadogDockerEc2EcsEksElk StackGithub ActionsGitlab Ci/CdGrafanaIamInfrastructure As CodeJenkinsKubernetesLambdaLangchainMlopsNvidia AiOpenai ApisPrometheusPythonRdsS3ShellSplunkTerraformVpc
One Month AgoSaved
Easy Apply
Remote
United States
Easy Apply
Senior level
Senior level
Cloud • Security • Software • Cybersecurity • Automation
Build and operate reliable, scalable production infrastructure for GitLab’s user-facing services. Responsibilities include developing infrastructure automation and tooling, managing Kubernetes deployments, maintaining infrastructure as code, supporting CI/CD and GitOps, participating in on-call and incident response, improving observability and SLOs, troubleshooting production systems, and documenting operational practices. The role spans Intermediate through Senior Staff levels and requires strong software engineering, cloud, reliability, and asynchronous collaboration skills.
Top Skills: AlertingAWSCi/CdGCPGitopsGoInfrastructure As CodeKubernetesLoggingMetricsRubySlisSlosTerraform
4 Days AgoSaved
Remote
USA
200K-270K Annually
Expert/Leader
200K-270K Annually
Expert/Leader
Social Media • Software
Design, implement, and operate infrastructure for high-scale production systems serving millions of users. Own reliability, observability, incident response, deployments, capacity planning, cost management, and operational excellence across bare-metal and cloud environments. Develop automation and performance tooling, improve production readiness, manage vendors, lead incident reviews, and mentor engineers on reliability and distributed-systems practices.
Top Skills: Bare-Metal ServersCloud ServicesDatabasesDistributed SystemsGoKubernetesLinuxNetworkingObservability SystemsStorage Systems
5 Days AgoSaved
Remote
United States
113K-140K Annually
Senior level
113K-140K Annually
Senior level
Edtech • Professional Services
The Site Reliability Engineer will operate and improve critical, high-traffic systems across AWS and Azure. Responsibilities include infrastructure-as-code, automation, incident response, root-cause analysis, monitoring, alerting, disaster recovery, capacity planning, performance tuning, high availability, cloud security, CI/CD, and multi-cloud networking. The role also involves mentoring, documentation, regular DR testing, and on-call participation.
Top Skills: ArmAWSAws CdkAws Direct ConnectAzureAzure DevopsAzure ExpressrouteBashBicepCi/CdCircleCICloudFormationDatadogEc2Github ActionsIamInfrastructure As CodeKubernetesOwaspPagerdutyPowershellPythonRdsS3SnykSonarqubeTerraformVpcVpn
5 Days AgoSaved
Remote
United States
Entry level
Entry level
Other
Build and operate AWS GovCloud environments within FedRAMP boundaries, including multi-account infrastructure, guardrails, logging, security tooling, and Terraform provisioning. Develop automated fleet release, patching, verification, CI/CD, compliance evidence, telemetry, SLO, and alerting systems. Support VM and appliance fleets at scale, participate in on-call, and engineer recurring incident causes away. The role also requires compliance-as-code, continuous monitoring, and test harnesses validating operational state.
Top Skills: AlertingArtifact ManagementAws ConfigAws GovcloudCi/CdDashboardsFedrampFipsGrafanaKubernetesOpaOpentelemetryOscalSecurity ScanningSlosTelemetryTerraform
Reposted 9 Days AgoSaved
In-Office or Remote
United States
146K-264K Annually
Senior level
146K-264K Annually
Senior level
Cloud • Security • Software • Cybersecurity
Ensure reliability, scalability, and usability of network infrastructure for Akamai Connected Cloud. Define requirements and SLOs, build automation and CI/CD pipelines, collaborate with dev/QA to improve code and stability, troubleshoot complex network issues (on-call), and mentor teammates while driving architectural standards.
Top Skills: AnsibleArgocdBashBirdChefFrrGithub ActionsGoGobgpJenkinsLinux NetworkingPuppetPythonSalt Stack
9 Days AgoSaved
Remote
United States
Senior level
Senior level
Artificial Intelligence • Information Technology • Consulting
Lead end-to-end deployment and operation of cloud infrastructure and Symmetri environments for US customers. Provision, configure, monitor, maintain, upgrade, and troubleshoot cloud, single-tenant, on-premises, and air-gapped environments. Manage Kubernetes, networking, IAM, infrastructure as code, resource planning, incident response, runbooks, and root-cause analysis. Represent the company in customer security reviews and help establish US deployment and operations standards and teams.
Top Skills: AWSAws CdkAzureDnsGitIamKubectlKubernetesLoad BalancingOpentofuPulumiPythonTerraformTerragruntVnetsVpc
9 Days AgoSaved
Remote
United States
104K-166K Annually
Senior level
104K-166K Annually
Senior level
Aerospace • Information Technology • Security • Cybersecurity • Defense
Designs and develops Python-based AWS cloud solutions, reliability tools, automation, observability, and Terraform infrastructure. Builds CI/CD pipelines, defines SRE standards and SLOs, supports incident management, and improves platform reliability through monitoring, postmortems, testing, and proactive toil reduction. Collaborates with cross-functional Agile teams to deliver scalable cloud automation across the Federal Reserve System.
Top Skills: Automated TestingAWSAws CanaryAws CloudwatchCi/CdCloudFormationEc2EventbridgeGoGrafanaIamInfrastructure As Code (Iac)LambdaPythonS3Step FunctionsTerraformVpc
10 Days AgoSaved
Remote
United States
175K-185K Annually
Senior level
175K-185K Annually
Senior level
Mobile
Serve as Vida’s first dedicated Site Reliability Engineer, modernizing Terraform, standardizing environments, improving CI/CD, scaling GCP and Kubernetes infrastructure, upgrading databases and runtimes, and strengthening monitoring, observability, security, and operational processes. The role includes building runbooks, on-call and escalation procedures, supporting enterprise launches, managing infrastructure costs, and evaluating multi-cluster Kubernetes architecture in a fully remote environment.
Top Skills: AirflowCloud MonitoringCloud SqlDatadogDjangoFirestoreGCPGithub ActionsGkeIamKubernetesMySQLPostgresPythonRedisTerraform
10 Days AgoSaved
Remote
United States
160K-210K Annually
Senior level
160K-210K Annually
Senior level
Professional Services • Consulting
Own and improve CI/CD, developer tooling, agent harnesses, Kubernetes infrastructure, observability, cost management, and incident-response practices. Build platform capabilities that reduce engineering toil and improve delivery speed for conversational AI products. Participate in an on-call rotation for core infrastructure and help shape reliable, scalable systems across a fully remote engineering organization.
Top Skills: AlertingCi/CdHelmIncident ManagementKubernetesLlmsLogsMetricsMonitoringNode.jsObservabilityPythonTerraformTracingTypescript
10 Days AgoSaved
Remote
United States
125K-150K Annually
Senior level
125K-150K Annually
Senior level
Cloud • Information Technology
Own the reliability, scalability, and administration of production Vitess/MySQL and Cassandra databases. Design highly available architectures, replication, backup, recovery, security, and disaster recovery strategies. Lead incident response, on-call operations, monitoring, SLO development, root cause analysis, and production readiness reviews. Automate database and infrastructure operations using scripting, Kubernetes, Terraform, Ansible, and Jenkins. Establish runbooks, escalation procedures, training materials, and mentor junior database SREs while partnering across engineering and infrastructure teams.
Top Skills: AnsibleAWSAzureBashCassandraCatchpointDockerElkFirehydrantGCPGoGrafanaItilJenkinsKubectlKubernetesLinuxMySQLMysqlshNomadOssPrometheusPythonSQLTerraformVaultVitess
10 Days AgoSaved
In-Office or Remote
CA, USA
152K-288K Annually
Senior level
152K-288K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Build and operate software automation for large-scale NVIDIA GPU infrastructure. Responsibilities include bare-metal provisioning, hardware validation, firmware and software upgrades, cluster lifecycle management, BMC and Redfish tooling, failure diagnosis, observability, incident response, and automated repair. The role supports NVL72 systems, BlueField DPUs, Linux, Kubernetes, and cloud or on-premises environments while coordinating across hardware, networking, platform, data center, and partner teams.
Top Skills: Argo CdBluefield-3 DpusBmcDriver ManagementFirmware ManagementGitopsGoInfinibandKubernetesLinuxNetwork BootNvidia Nvl72NvlinkPythonRedfishSlosSpectrum-X
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account