Core Scientific Logo

Core Scientific

Staff Site Reliability Engineer

Posted One Month Ago
Be an Early Applicant
Remote
Hiring Remotely in USA
Senior level
Remote
Hiring Remotely in USA
Senior level
Lead design, implementation, and operation of reliable hybrid cloud and on-prem infrastructure. Deliver complex initiatives, build immutable infrastructure with IaC, improve observability and incident response, mentor engineers, and ensure reliability, scalability, and compliance in change-controlled environments.
The summary above was generated by AI

Who We Are 

Core Scientific is a leading provider of infrastructure for high-performance compute in North America. Our mission is to accelerate digital innovation by scaling high-value compute rapidly, efficiently, and responsibly. We transform energy into high-value compute with unmatched efficiency at scale. Core Scientific is a publicly trade company (NASDAQ: CORZ). 


We power AI, HPC, and other next-generation data center workloads demanding exceptional computing power, in addition to our digital asset mining operations. Our footprint consists of 11 data center campuses across seven states, housing advanced infrastructure for our customers. 


What sets us apart? We have an entrepreneurial culture, a "can-do" and collaborative attitude, and we own and control our infrastructure. These strategic advantages enable us to maintain operational excellence, increase efficiency, and rapidly deploy cutting-edge innovations developed by our team of experts. 


Join us and accelerate your career alongside our groundbreaking journey. We seek smart, creative, and collaborative professionals who thrive in a fast-paced, result-driven environment. Ready to be part of something exceptional? Apply today and make an impact at Core Scientific.


The Job 
We are seeking a capable, motivated generalist who thrives in a change-controlled, compliant environment and enjoys working across hybrid cloud and on-premises systems. This role partners closely with application architecture and peer engineering teams while contributing hands-on across platform engineering, DevOps, and SRE. 
This position is expected to take ownership of complex technical initiatives and see them through to completion—balancing hands-on implementation with effective delegation and cross-team coordination. 
Responsibilities 

  • Lead end-to-end delivery of complex technical initiatives, from problem definition and design through implementation, rollout, and operation.

  • Own the design, implementation, and reliability of systems across hybrid cloud and on-premises environments.

  • Take accountability for technical outcomes, including system reliability, scalability, and performance in regulated, change-controlled environments.

  • Drive execution by coordinating work across engineers and teams, delegating effectively while remaining hands-on where needed.

  • Partner with application architecture and peer teams to shape system design and influence technical decisions.

  • Build, deploy, and operate infrastructure and applications using automation and infrastructure as code.

  • Implement secure, immutable infrastructure using modern tooling (e.g., Terraform, Kubernetes, Helm, Ansible).

  • Improve observability, monitoring, and incident response practices.

  • Establish and promote best practices for reliability, security, and operational excellence across teams.

  • Mentor engineers and contribute to raising the technical bar across the organization.

  • Foster open, respectful, and professional communication directly within the team as well as with co-workers/ teammates and leaders across the organization.

  • Performs other duties as assigned.

Qualifications 

  • Bachelor's degree in Computer Science or a related field, 7+ years of experience, or equivalent demonstrated impact in SRE, DevOps, or Infrastructure Engineering.

  • Broad technical experience across infrastructure and distributed systems, with the ability to design effective solutions, apply appropriate patterns, and anticipate scaling, reliability, and operational challenges.

  • Strong understanding of distributed systems behavior, including application runtime characteristics, service-to-service communication, networking, and failure modes in production environments.

  • Experience operating in regulated, compliant, or change-controlled environments.

  • Experience working in hybrid environments (AWS preferred; on-premises infrastructure required).

  • Strong experience with Infrastructure as Code, configuration management, and orchestration tools (Terraform, Helm, Kustomize, Ansible).

  • Experience with Kubernetes and virtualization technologies.

  • Experience with observability platforms (e.g., Datadog), including building monitoring and alerting integrations.

  • Experience with build and release systems (e.g., GitHub Actions, Makefiles, Python tooling).

Reports To
Site Reliability Engineering Manager
Location  
To be considered for the role you must reside near Miami, FL or Austin, TX.  
Travel 
Occasional travel may be required as needed.
Work Environment  
This job typically operates in a professional office environment and routinely utilizes standard equipment, including laptop computers and smartphones. This role may also travel to data center sites, and the work environment at a data center may contain loud noise, construction, and other operational elements. 
Physical Demands  
While performing the duties of this job, the employee is frequently required to sit, stand, walk, use hands, and lift up to 25 pounds.    
Position Type/ Expected Hours of Work  
This is a full-time position. General hours and days of work are Monday through Friday, 8:00 a.m. to 5:00 p.m. The employee is expected to be available generally around U.S. time zones and will be part of an on-call rotation. The current rotation is 1 week every 5 weeks. 
Supervisory Experience (Yes or No)  
No

Similar Jobs

An Hour Ago
Remote or Hybrid
Senior level
Senior level
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Staff Software Engineer responsible for operating production Kubernetes across hybrid and multi-cloud environments, building auto-remediation and observability systems, developing Infrastructure-as-Code and GitOps pipelines, defining SLOs and alerting, supporting incident response, and reducing operational toil. The role also includes mentoring SRE engineers, improving reliability practices, and contributing to cloud, data center, and on-call operations.
Top Skills: AksAWSAzureBashCi/CdCloudFormationEc2EksGCPGitopsGkeGoInfrastructure As CodeKubernetesLinuxMachine LearningPythonRdsService MeshTerraform
9 Days Ago
Easy Apply
Remote
United States
Easy Apply
Senior level
Senior level
Cloud • Security • Software • Cybersecurity • Automation
Build and operate reliable, scalable production infrastructure for GitLab’s user-facing services. Responsibilities include developing infrastructure automation and tooling, managing Kubernetes deployments, maintaining infrastructure as code, supporting CI/CD and GitOps, participating in on-call and incident response, improving observability and SLOs, troubleshooting production systems, and documenting operational practices. The role spans Intermediate through Senior Staff levels and requires strong software engineering, cloud, reliability, and asynchronous collaboration skills.
Top Skills: AlertingAWSCi/CdGCPGitopsGoInfrastructure As CodeKubernetesLoggingMetricsRubySlisSlosTerraform
3 Days Ago
Remote
USA
177K-240K Annually
Expert/Leader
177K-240K Annually
Expert/Leader
Information Technology • Security • Software • Cybersecurity
Lead reliability strategy across eight engineering groups by defining SLIs, SLOs, and error budgets; strengthening incident response, alerting, postmortems, change safety, and failure testing; and coaching teams to own reliability. The role remains hands-on through production investigations, tooling, dashboards, and reference implementations. It also leads AI adoption in incident management and observability while partnering with architecture, platform, and product teams to improve distributed-system resilience.
Top Skills: Ai ToolsAWSDatadogDynamoDBElasticsearchGoInfrastructure As CodeKafkaKubernetesObservability ToolingRedisTypescript

What you need to know about the Los Angeles Tech Scene

Los Angeles is a global leader in entertainment, so it’s no surprise that many of the biggest players in streaming, digital media and game development call the city home. But the city boasts plenty of non-entertainment innovation as well, with tech companies spanning verticals like AI, fintech, e-commerce and biotech. With major universities like Caltech, UCLA, USC and the nearby UC Irvine, the city has a steady supply of top-flight tech and engineering talent — not counting the graduates flocking to Los Angeles from across the world to enjoy its beaches, culture and year-round temperate climate.

Key Facts About Los Angeles Tech

  • Number of Tech Workers: 375,800; 5.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Snap, Netflix, SpaceX, Disney, Google
  • Key Industries: Artificial intelligence, adtech, media, software, game development
  • Funding Landscape: $11.6 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Strong Ventures, Fifth Wall, Upfront Ventures, Mucker Capital, Kittyhawk Ventures
  • Research Centers and Universities: California Institute of Technology, UCLA, University of Southern California, UC Irvine, Pepperdine, California Institute for Immunology and Immunotherapy, Center for Quantum Science and Engineering

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account