MeridianLink Logo

MeridianLink

Sr. Site Reliability Engineer

Posted 11 Days Ago
Remote
Hiring Remotely in US
104K-163K Annually
Senior level
Remote
Hiring Remotely in US
104K-163K Annually
Senior level
Lead reliability, observability, and resilience for cloud-based financial SaaS. Define SLOs/SLIs, design monitoring/tracing, own incident response and runbooks, build IaC and automation, implement AIOps, perform chaos and load testing, and write production-grade Python tooling while ensuring security and compliance.
The summary above was generated by AI

About the Role

We are seeking a Senior Site Reliability Engineer to join our cloud engineering team. You will own the reliability, scalability, and observability of our critical financial SaaS applications and infrastructure, working across cloud platforms to ensure our customers experience is seamless, secure, and performant services. This is a high-impact role for someone who is passionate about building resilient systems and preventing outages before they happen.

Key Responsibilities

  • Design, implement, and maintain Service Level Objectives (SLOs) and Service Level Indicators (SLIs) across all critical systems; ensure we meet or exceed targets consistently

  • Lead observability strategy by designing comprehensive monitoring, logging, and tracing architectures; select and deploy observability tools that provide deep visibility into system behavior

  • Build and own runbooks, incident response procedures, and post-incident review processes; mentor the team on incident management and blameless postmortems

  • Architect and deploy cloud infrastructure on AWS or Azure; implement infrastructure-as-code practices and ensure high availability, disaster recovery, and business continuity

  • Develop automation and AIOps capabilities to reduce toil, accelerate incident detection, and enable self-healing systems; implement intelligent alerting to minimize false positives

  • Drive reliability improvements through load testing, chaos engineering, and failure scenario analysis; identify and eliminate single points of failure

  • Partner with application and backend teams to design reliable systems from inception; conduct architecture reviews and reliability assessments

  • Write production-grade Python tooling for automation, metrics collection, alert management, and operational workflows

  • Champion security and compliance in infrastructure; implement defense-in-depth principles for a regulated fintech environment

Required Qualifications

  • 7+ years in Site Reliability Engineering, DevOps, platform engineering, or closely related roles with significant responsibility for production systems

  • Expert-level experience with Azure or AWS (or both); deep knowledge of compute, networking, storage, and managed services; experience managing infrastructure at scale

  • Demonstrated expertise in observability: designing and implementing monitoring, alerting, logging, and distributed tracing solutions; hands-on with observability platforms (e.g., Prometheus, Grafana, ELK, Datadog, New Relic, or similar)

  • Strong background in SLOs, SLIs, and SLAs; experience defining meaningful objectives and building systems to meet them; understanding of error budgets and their role in prioritization

  • Proven experience designing and troubleshooting highly available, resilient, and scalable systems; deep understanding of distributed systems concepts and failure modes

  • Proficiency in Python, PowerShell, bash, etc. scripting languages for production automation, tooling, and systems programming; ability to write clean, maintainable code for operational workflows

  • Hands-on experience with AIOps practices: event correlation, intelligent alerting, predictive analytics, and automated remediation; familiarity with AIOps platforms is a plus

  • Experience with infrastructure-as-code tools (e.g., Terraform, CloudFormation, Ansible); version control and CI/CD pipeline design

  • Track record of incident management and on-call ownership; comfort with incident response and the ability to remain calm under pressure

  • Excellent communication skills; ability to work cross-functionally and influence without authority; comfort mentoring junior engineers

Preferred Qualifications

  • Experience in the fintech, payments, banking, or other regulated industries; understanding of compliance requirements (SOC 2, PCI-DSS, etc.)

  • Experience with Kubernetes and container orchestration; deep knowledge of containerized application deployment and management

  • Proficiency with observability as code; experience building custom metrics, dashboards, and alerts programmatically

  • Background in chaos engineering or reliability testing; experience using tools like Gremlin or similar platforms

  • Contribution to open-source observability or infrastructure projects

  • Expertise in network security, application security, or infrastructure hardening

  • Experience with database optimization, query performance tuning, and backup/recovery strategies

HQ

MeridianLink Costa Mesa, California, USA Office

3560 Hyland Ave, Suite #200, Costa Mesa, CA, United States, 92626

Similar Jobs

2 Days Ago
Remote or Hybrid
USA
140K-215K Annually
Senior level
140K-215K Annually
Senior level
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Senior SRE owning availability, automation, and observability for CI/CD platform services. Build and operate infrastructure, run on-call, lead incident response, mentor engineers, drive design/capacity planning, integrate AI-assisted workflows, and improve cross-team reliability.
Top Skills: Active DirectoryAnsibleApache AirflowSparkAWSAzureBashBazelBitbucketCassandraChefDatadogDnsFirewall RulesGCPGitGithub ActionsGitlabGitlab CiGoGrafanaHoneycombHumio/LogscaleJenkinsKafkaKubernetesLoad BalancersMongoDBMySQLNasNew RelicNfsObject StorageOpensearchOraclePostgresPowershellPrometheusPulsarPuppetPythonRabbitMQRedis/ValkeyRedpandaRoutingSaltSanSplunkTerraformVarnishVipsWindows Server
3 Days Ago
In-Office or Remote
153K-205K Annually
Senior level
153K-205K Annually
Senior level
Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3
Build and operate scalable cloud-native microservices and Kubernetes infrastructure, improve CI/CD and developer workflows, engineer and run an autonomous coding-agent orchestration platform, integrate and operationalize AI services with guardrails, implement observability and incident response, and collaborate with product and engineering teams to ensure secure, reliable production systems and cost optimization.
Top Skills: Ai ApisAutonomous AgentsAWSCi/CdGCPGoJavaJavaScriptKubernetesMonitoringObservabilityPythonRestful ApisRustSdksSQLTypescriptWorkflow Orchestration
5 Days Ago
Remote or Hybrid
130K-160K Annually
Senior level
130K-160K Annually
Senior level
AdTech • Cloud • Digital Media • Information Technology • News + Entertainment • App development
Design, deploy, and maintain on-premises and cloud playout infrastructure for IP video distribution. Build automation, CI/CD pipelines, monitoring, and scalable fault-tolerant systems. Drive releases, troubleshoot broadcast incidents, mentor SREs, and provide 24/7 on-call support.
Top Skills: AnsibleAWSAzureBashBroadcast TechnologiesCi/CdContainerizationGCPIp VideoJavaScriptKubernetesLinuxPerlPythonRubyStreamingTerraform

What you need to know about the Los Angeles Tech Scene

Los Angeles is a global leader in entertainment, so it’s no surprise that many of the biggest players in streaming, digital media and game development call the city home. But the city boasts plenty of non-entertainment innovation as well, with tech companies spanning verticals like AI, fintech, e-commerce and biotech. With major universities like Caltech, UCLA, USC and the nearby UC Irvine, the city has a steady supply of top-flight tech and engineering talent — not counting the graduates flocking to Los Angeles from across the world to enjoy its beaches, culture and year-round temperate climate.

Key Facts About Los Angeles Tech

  • Number of Tech Workers: 375,800; 5.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Snap, Netflix, SpaceX, Disney, Google
  • Key Industries: Artificial intelligence, adtech, media, software, game development
  • Funding Landscape: $11.6 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Strong Ventures, Fifth Wall, Upfront Ventures, Mucker Capital, Kittyhawk Ventures
  • Research Centers and Universities: California Institute of Technology, UCLA, University of Southern California, UC Irvine, Pepperdine, California Institute for Immunology and Immunotherapy, Center for Quantum Science and Engineering

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account