Top Remote Site Reliability Engineer Jobs in Los Angeles, CA

Reposted 4 Days AgoSaved
Remote
USA
235K-275K Annually
Expert/Leader
235K-275K Annually
Expert/Leader
Legal Tech • Software
Senior technical leader for SRE driving observability, platform infrastructure, SLIs/SLOs, incident response, automation, and self-service platform capabilities. Shapes reliability strategy, mentors engineers, and ensures production-scale operational excellence.
Top Skills: AiopsBashDatadogGoInfrastructure As CodeKubernetesNew RelicObservabilityPython
Reposted 14 Days AgoSaved
In-Office or Remote
2 Locations
164K-270K Annually
Mid level
164K-270K Annually
Mid level
Aerospace • Hardware • Software • Defense • Manufacturing
Build scalable automated solutions for device fleet management, own and optimize MDM platforms, write OS-level scripts for self-healing, gather telemetry to prevent end-user disruption, translate compliance (CMMC) into code-managed baselines, and create dashboards and alerts measuring end-user SLOs.
Top Skills: AnsibleBashChefFleet DmIntuneJAMFOsqueryPowershellPulumiPuppetPythonSaltTerraformWorkspace One
Reposted 10 Days AgoSaved
In-Office or Remote
CA, USA
161K-284K Annually
Senior level
161K-284K Annually
Senior level
Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
The Senior Site Reliability Engineer will enhance reliability of Block's platform, improve incident response using AI tools, and coordinate incident management. Responsibilities include building reliable systems, standardizing tools, and leading high-severity incidents during on-call rotations.
Top Skills: Amazon Web ServicesDatadogDynamoDBGrpcHTTPIstioJavaJSONKotlinKubernetesLaunchdarklyMySQLProtocol BuffersTerraformVitess
Reposted 6 Days AgoSaved
In-Office or Remote
United States
146K-264K Annually
Senior level
146K-264K Annually
Senior level
Cloud • Security • Software • Cybersecurity
Lead and mentor SRE teams; partner with engineering, operations and product; apply statistical analysis and networking expertise to diagnose performance and reliability issues; define and implement data feeds; influence technical decisions and investments; build tooling to automate analytical workflows and increase platform reliability.
Top Skills: CCloudDistributed SystemsDnsEdgeHTTPJavaPerlPythonRSQLTcpTls
Reposted 6 Days AgoSaved
Remote
USA
Expert/Leader
Expert/Leader
Artificial Intelligence • Software • Cybersecurity
Design, build, and maintain scalable, highly available cloud infrastructure for an AI-native cybersecurity platform. Automate deployments and incident response, optimize performance for AI workloads, manage IaC across cloud environments, lead incident management and post-mortems, and collaborate with engineering and security teams to embed reliability.
Top Skills: AWSAzureDatadogDistributed DatabasesEksElkEvent-Driven SystemsGCPGkeGrafanaKubernetesMicroservicesNetworkingPrometheusPulumiStorageTerraform
Reposted 6 Days AgoSaved
Remote
USA
110K-140K Annually
Senior level
110K-140K Annually
Senior level
Real Estate • Financial Services • PropTech
Support and optimize products migrated to AWS, implement cloud best practices, maintain operational coverage, enhance automation, observability, CI/CD/GitOps, and security. Collaborate with development and platform teams to scale, troubleshoot, and ensure reliable SaaS operations.
Top Skills: AmisArgocdAWSAws Elastic BeanstalkAws Transfer FamilyAzure DevopsBashCloudwatchCurlDockerEc2EksFluxcdGitGitopsHTTPIstioKubernetesLinkerdLoad BalancerPowershellPythonRdsSQLTerraformWget
Mid level
Artificial Intelligence • Hardware • Software • Semiconductor
Operate and scale production AI inference infrastructure, run releases and capacity changes, build self-service CD pipelines and automation, extend telemetry and observability, collaborate on SLOs, post-mortems, and capacity planning to reduce operational toil.
Top Skills: Argo CdBazelFluxGitopsGoGrafanaInfluxdbKubernetesPrometheusPython
Reposted 8 Days AgoSaved
Remote
US
101K-161K Annually
Senior level
101K-161K Annually
Senior level
Cloud • Software • Analytics
Join Arista Networks as a Site Reliability Engineer to manage CloudVision service reliability, scalability, and stability in a FedRAMP environment, focusing on areas like architecture, security, and performance optimization.
Top Skills: AnsibleBashGCPGkeGoKubernetesPulumiPython
9 Days AgoSaved
Remote
USA
140K-170K Annually
Mid level
140K-170K Annually
Mid level
Artificial Intelligence • Software • Generative AI • Automation
Operate and harden Blitzy's self-hosted, Kubernetes-based AI platform inside customer-controlled secure cloud environments. Own deployments, upgrades, capacity planning, observability, incident response, and customer-facing technical coordination while championing security and feeding operational learnings back into the product roadmap.
Top Skills: Alerting)BashCloud (Aws/Gcp/Azure)Container OrchestrationGoInfrastructure-As-CodeKubernetesMetricsObservability (LoggingPulumiPythonTerraformTracing
9 Days AgoSaved
Remote
United States
Senior level
Senior level
Other
Lead architecture and delivery of highly available, resilient cloud and on‑prem systems for the Password Safe platform. Own platform engineering, CI/CD pipelines, IaC/GitOps, release orchestration, observability (metrics/logs/traces), chaos engineering, SLO/SLI definition, and core services. Mentor engineers, define SRE strategy, and drive reliability, security, and automation improvements.
Top Skills: AnsibleApi GatewaysAWSAzureBlue Green DeploymentsC#CachesCanary DeploymentsChaos EngineeringCi/CdConfiguration ManagementDatadogDevsecopsDockerGitopsGoGrafana CloudJavaKubernetesLinuxOpentelemetryOpentofuSecrets ManagementService MeshTerraformWindows
Reposted 10 Days AgoSaved
In-Office or Remote
California, USA
Senior level
Senior level
Artificial Intelligence
The Deployment Engineer will build and operate AI inference clusters, ensure scalable deployments, optimize allocation, and maintain infrastructure. Responsibilities include software updates, telemetry development, and collaborative improvements with teams.
Top Skills: DockerGrafanaInfluxdbK8SLinuxPrometheusPython
11 Days AgoSaved
Remote
United States
152K-253K Annually
Senior level
152K-253K Annually
Senior level
Cloud • Security • Software • Cybersecurity
Build and operate the Veeam Data Cloud GOV environment: map systems, write runbooks, define SLIs/SLOs, run incident response, close observability gaps, automate deployments and support fleet management while working across security and compliance constraints.
Top Skills: Api ManagementApplication InsightsArgocdAws CloudformationAzureAzure Arm TemplatesAzure DevopsAzure FunctionsAzure GovernmentAzure MonitorAzure StorageBitbucketC#Cosmos DbDaggerElastic StackElkEntra IdFluxcdGitGithub ActionsGitlab CiGoGrafanaJavaJavaScriptKubernetesMicrosoft TfsOpentelemetryPrometheusPulumiServerless FrameworkTerraformTerragruntTypescript
New

Cut your apply time in half.

Use ourAI Assistantto automatically fill your job applications.

Use For Free
Application Tracker Preview
17 Days AgoSaved
Remote
United States
180K-220K Annually
Senior level
180K-220K Annually
Senior level
Software • Defense
Work as an SRE embedded with product teams to improve reliability by fixing application code (primarily TypeScript), building observability (Prometheus, Loki, Grafana, Alloy), defining SLIs/SLOs, leading incident response and postmortems, automating toil, and supporting deployments across on‑prem DoD and AWS environments.
Top Skills: AlloyAWSBashContainersDockerGithub ActionsGitlab Ci/CdGoGrafanaJenkinsKubectlKubernetesLokiNode.jsPrometheusPythonTypescript
Reposted 11 Days AgoSaved
Remote
US
Senior level
Senior level
Big Data • Healthtech • Information Technology • Analytics
Serve as an SRE-focused AI enablement partner: evaluate AI architectures, coach engineering teams on LLM integrations and agentic/RAG patterns, advise on observability, SLOs, governance, and operational readiness, and develop internal standards and documentation.
Top Skills: Agentic FrameworksAnthropic ClaudeAWSAzureAzure Ai FoundryCi/CdDatabricksDatadogDockerGrafanaKubernetesLangchainLlamaindexLlm ApiOpentelemetryRagSemantic Kernel
Reposted 12 Days AgoSaved
Remote
USA
Senior level
Senior level
Information Technology • Cryptocurrency
The Site Reliability Engineer will lead technical initiatives, architect solutions, troubleshoot issues, mentor team members, and improve observability practices.
Top Skills: ArgocdBashElk StackGCPGoGrafanaHelmKubernetesPrometheusPythonTerraform
Reposted 12 Days AgoSaved
Remote
United States
154K-231K Annually
Senior level
154K-231K Annually
Senior level
Information Technology • Marketing Tech • Social Media
Lead technical design and architecture of large-scale Ceph clusters, drive capacity planning and major upgrades, resolve complex cross-functional production incidents, establish automation and reliability standards, and mentor engineers on advanced Ceph operations to scale the global storage platform.
Top Skills: AnsibleBluestoreCephCephfsCinderCrushCsiDdnErasure CodingGoIsilonKubernetesManilaMdsMgrMonNetappNeutronNovaOpenstackOsdPlacement GroupsPurePythonRbdRbd MirroringRgwRgw MultisiteRookS3SaltstackSwiftTerraformVastWeka
Reposted 12 Days AgoSaved
Remote
United States
128K-192K Annually
Senior level
128K-192K Annually
Senior level
Information Technology • Marketing Tech • Social Media
Own and maintain reliability, performance, scalability, and capacity of massive Ceph storage clusters. Troubleshoot distributed storage issues, build automation (Python/Shell, SaltStack, Ansible), define observability (SLIs/SLOs, PromQL/LogQL, Grafana), and lead lifecycle initiatives like upgrades, expansions, and migrations.
Top Skills: AnsibleCeph (RadosCephfs)ChefGrafanaLinuxLogqlPromqlPuppetPythonRbdRgwSaltstackShell
Reposted 13 Days AgoSaved
Remote
United States
100K-140K Annually
Mid level
100K-140K Annually
Mid level
Artificial Intelligence • Information Technology • Consulting
The Linux Systems Administrator will maintain and troubleshoot Linux systems, support network services, and work on systems integration while collaborating with infrastructure teams.
Top Skills: DhcpDnsLinuxNtpPython
14 Days AgoSaved
In-Office or Remote
USA
Senior level
Senior level
Artificial Intelligence • Cloud • Information Technology • Software
Embed with customer teams running large-scale GPU training and inference to onboard, tune, debug, and improve reliability. Diagnose fabric, driver, scheduler, and application failures; profile performance; build automation, monitoring, and preflight checks; lead incident response; and convert field learnings into product improvements and reusable reference configurations.
Top Skills: AnsibleBashCgroupsContainer RuntimesCudaDcgmDevice PluginsFabric ManagerGoGpfsHelmInfinibandKubernetesKv CacheLustreNamespacesNcclNvidia DriversNvidia-SmiNvlinkPythonRoceSlurmSshTerraformTopology-Aware SchedulingVastWeka
14 Days AgoSaved
Remote
USA
75K-90K Annually
Mid level
75K-90K Annually
Mid level
3PL: Third Party Logistics
Own and improve uptime for backend services, APIs, workers, and ML pipelines; build monitoring, alerting, auto-remediation, optimize GCP infrastructure and Postgres performance, run on-call rotation and blameless postmortems.
Top Skills: Apollo ServerBashCi/CdCloud RunCloud SqlDatadogDockerExpressGCPGcsGrafanaNext.JsNode.jsPostgresPrometheusPythonReactTerraformTypescriptZabbix
14 Days AgoSaved
Remote
US
175K-185K Annually
Senior level
175K-185K Annually
Senior level
Software
Lead improvements in reliability, performance, scalability, capacity, and observability through automation and tooling. Partner with developers to design infrastructure and monitoring, define SLIs/SLOs, conduct load and performance testing, participate in incident response and root cause analysis, and manage monitoring services for production systems.
Top Skills: DockerGoGradleJavaKubernetesOpentelemetrySpring BootTerraform
15 Days AgoSaved
In-Office or Remote
11 Locations
140K-150K Annually
Mid level
140K-150K Annually
Mid level
Healthtech
Build, operate, and scale AWS cloud infrastructure and Kubernetes workloads using Terraform and Helm. Improve observability, define SLIs/SLOs, automate deployments and incident response, support on-call rotation, and implement security and compliance (HIPAA, SOC 2) best practices while partnering with product and engineering teams.
Top Skills: AWSCi/CdEvent SourcingHelmKubernetesLinuxMonitoring/Logging/TracingNetworkingTerraform
Reposted 20 Days AgoSaved
Easy Apply
Remote or Hybrid
United States
Easy Apply
127K-249K Annually
Senior level
127K-249K Annually
Senior level
Big Data • Cloud • Software • Database
Develop and maintain Kubernetes runtime environments, support developers, resolve critical issues, and participate in on-call rotations for production systems.
Top Skills: AWSAzureCert-ManagerCorednsCrdsCriCsiGatekeeperGCPGoHelmKubernetesKustomizeOperatorsPythonTerraform
16 Days AgoSaved
Remote
USA
140K-165K Annually
Senior level
140K-165K Annually
Senior level
Hardware • Machine Learning • Security • Software
Design and maintain developer experience tooling and CI/CD pipelines, manage self-service deployment tools (secrets, rollbacks), partner with cloud teams to troubleshoot production systems, participate in on-call rotation, and write production-grade automation and tests in Go and TypeScript to improve developer velocity and reliability.
Top Skills: AWSGithub ActionsGoHelmKubernetesTerraformTypescript
Reposted 16 Days AgoSaved
Remote
United States
115K-135K Annually
Mid level
115K-135K Annually
Mid level
Aerospace • Manufacturing
As a Site Reliability Engineer, you'll build and manage observability platforms for satellite communications, define SLOs/SLIs, and collaborate on incident response and deployment automation.
Top Skills: ArgocdAWSElkGCPGoGrafanaIstioJaegerKubernetesLinkerdLokiOpentelemetryPrometheusPythonTempoTerraform
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account