Backblaze Logo

Backblaze

Site Reliability Engineer III (DBA)

Posted 6 Hours Ago
Be an Early Applicant
Remote
Hiring Remotely in United States
125K-150K Annually
Senior level
Remote
Hiring Remotely in United States
125K-150K Annually
Senior level
Own the reliability, scalability, and administration of production Vitess/MySQL and Cassandra databases. Design highly available architectures, replication, backup, recovery, security, and disaster recovery strategies. Lead incident response, on-call operations, monitoring, SLO development, root cause analysis, and production readiness reviews. Automate database and infrastructure operations using scripting, Kubernetes, Terraform, Ansible, and Jenkins. Establish runbooks, escalation procedures, training materials, and mentor junior database SREs while partnering across engineering and infrastructure teams.
The summary above was generated by AI

About Backblaze

Backblaze is the object storage leader in the open cloud movement, fueling customer success with cloud storage built purposefully to unlock budgets, unburden administrators, and unleash innovators. Together with our partners, we’re helping customers break free from the restrictive, overpriced legacy solutions that hold them back, and blaze forward with the full power of the open cloud in their hands.

Founded in 2007, we scaled the business with less than $3 million in outside funding until 2021, when we did a traditional IPO on the Nasdaq stock exchange. Today, Backblaze generates over $100m in revenue and is the leading specialized storage cloud - managing over three billion gigabytes of data storage for 500K+ customers in 175+ countries, including businesses, developers, IT professionals, and individuals.

But while there is a lot to celebrate in our past, there is almost as much opportunity ahead of us. We’re seeking a Site Reliability Engineer III (DBA) to join our team!

About the Role:

Individuals fulfilling this role will be responsible for ensuring the stability, scalability, and reliability of our production database systems, primarily Vitess (distributed MySQL) and Cassandra, alongside the rest of our production services and infrastructure. This role carries the same on-call, incident response, and service ownership expectations as other SRE IIIs, with database systems serving as the area of deepest technical ownership.

Because our SRE Database Engineering function is new, this role will also help establish its operational foundation by designing database architecture and developing the runbooks, escalation guidance, procedures, and training materials that our Level 1 and Level 2 SRE Database Engineers will use as they onboard. The ideal candidate will have strong experience with production database systems, Linux, automation, distributed systems, Kubernetes, observability, and incident response, with a proactive approach to reliability and operational excellence.

What You'll Do:
  • Database Architecture & Administration:
    • Design, deploy, and own highly available database architecture for Vitess (distributed MySQL) and Cassandra
    • Establish and document operational procedures, runbooks, and escalation guidance for Level 1 and Level 2 SRE Database Engineers
    • Optimize database performance through query tuning, indexing strategies, schema design, and capacity planning
    • Own database backup, recovery, replication, and disaster recovery strategies
    • Perform and validate disaster recovery testing and database recovery procedures
    • Drive database security, access control, patching, hardening, and compliance practices
    • Partner with the DBA and Data Infrastructure teams on resharding, capacity planning, replication, and architecture decisions for sharded MySQL environments
  • Service Reliability & Operations:
    • Support the availability and durability of critical services across production environments
    • Monitor service health using SLIs, SLOs, error budgets, monitoring, logging, and alerting platforms
    • Partner with service owners to define and improve SLIs, SLOs, error budget policies, and alerting
    • Participate in on-call rotations, incident response, root cause analysis, and post-incident reviews
    • Serve as an escalation point for complex database production incidents
    • Follow established ITIL/OSS processes including incident, change, problem, and capacity management
    • Take ownership of operational issues and drive projects from problem discovery through resolution
  • Automation & Tooling:
    • Develop automation for common operational and database administration tasks to reduce manual intervention and operational toil
    • Contribute to monitoring, logging, and alerting frameworks including Prometheus, Grafana, Catchpoint, and ELK
    • Help integrate operational runbooks and incident response workflows with FireHydrant
    • Work with CI/CD pipelines, configuration management, and infrastructure as code tools including Terraform, Ansible, and Jenkins
    • Develop scripts using Bash, Python, Go, or similar technologies to improve reliability and operational efficiency
    • Operate and troubleshoot containerized production environments using Kubernetes and Docker
    • Work within Kubernetes and Vitess environments using technologies such as kubectl, mysqlsh, and Vitess keyspaces
  • Project Management:
    • Lead Production Readiness Reviews (PRRs) for functionality being handed off from engineering partner teams
    • Support the operational readiness of new database-backed services before they enter production
    • Build training plans, onboarding materials, and technical documentation for new Level 1 and Level 2 SRE Database Engineers
    • Partner with Engineering, Product, Operations, and DBA/Data Infrastructure teams on reliability initiatives
    • Assist with capacity planning, disaster recovery exercises, database migrations, and infrastructure projects
    • Work with vendors and service providers to troubleshoot service issues and track SLA performance
    • Identify opportunities for automation and process efficiency
  • Incident response
    • Respond to and resolve production database, infrastructure, and service incidents
    • Troubleshoot and escalate database, Linux, networking, application, and infrastructure issues as needed
    • Participate in the on-call rotation and serve as an escalation point for database-related incidents
    • Lead or contribute to root cause analysis and post-incident reviews
    • Identify recurring issues and develop long-term corrective actions to improve reliability
  • What we value:
    • A proactive mindset with a can-do attitude
    • Someone who can work independently, take ownership, and drive complex technical problems through resolution
    • Someone who steps up, supports teammates, mentors others, and shares knowledge freely
    • Strong problem-solving skills and a willingness to learn new technologies
    • Curiosity, reliability, and a desire to improve the reliability and scalability of production systems
Required Qualifications:
  • 6–8 years of experience in site reliability engineering, systems engineering, infrastructure operations, database engineering, or similar roles, with meaningful experience supporting production database systems.
  • Deep hands-on experience with MySQL and distributed or sharded database systems.
  • Experience with Vitess in a production environment strongly preferred.
  • Experience administering and supporting NoSQL databases such as Cassandra.
  • Experience designing high-availability database architecture, replication topology, backup strategies, and disaster recovery processes.
  • Strong SQL skills, including query performance analysis, indexing, schema design, and troubleshooting.
  • Solid Linux systems administration and troubleshooting skills.
  • Experience with security-focused operations including patching, system hardening, access controls, and vulnerability remediation.
  • Strong understanding of service reliability concepts including monitoring, alerting, incident response, root cause analysis, SLIs, SLOs, and error budgets.
  • Experience working with containers and orchestration platforms including Kubernetes and Docker.
  • Comfortable operating in Kubernetes and Vitess environments using tools such as kubectl, mysqlsh, and Vitess keyspaces.
  • Experience with infrastructure and configuration management technologies including Terraform, Ansible, Jenkins, and HashiCorp products such as Vault and Nomad.
  • Proficiency in at least one scripting language such as Python, Bash, or Go.
  • Experience establishing operational procedures, runbooks, documentation, and escalation processes.
  • Experience mentoring, training, or helping onboard engineers into complex technical environments.
  • Experience in SaaS, cloud services, service provider, or large-scale distributed systems environments preferred.
  • Experience with AWS, GCP, Azure, or similar cloud platforms preferred.
  • Familiarity with ITIL/OSS practices and SLA/SLO management preferred.
  • Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent professional experience.

Backblaze Perks:

  • Healthcare for family, including dental and vision
  • Competitive compensation and 401K
  • RSU grants for full-time employees
  • ESPP program
  • Flexible vacation policy
  • Maternity & paternity leave
  • MacBook Pro to use for work, plus a generous stipend to personalize your workstation
  • Childcare bonus (human children only)
  • Fertility treatment and support
  • Learning & development program
  • Commuter benefits
  • Culture that supports a healthy work-life balance

To provide greater transparency to candidates, we share base pay ranges for all US-based job postings regardless of state. We set standard base pay ranges for all roles based on function, level, and country location, benchmarked against similar-stage growth companies. Final offer amounts are determined by multiple factors, including candidate location, skills, depth of work experience, and relevant licenses/credentials, and may vary from the amounts listed below.

The expected salary range for this role is - $125,000 - $150,000.

At Backblaze, we value being fair and good to our customers, partners, and employees. That’s why diversity, equity, and inclusion are at the core of our values. We are committed to fostering a workforce where all employees feel a sense of belonging regardless of race, ethnicity, nationality, gender, sexual orientation, age, religion, socio-economic status, ability, veteran status, and education. We believe that our dedication to cultivating a diverse workspace not only allows us to better serve our customers in over 175 countries but further reinforces our commitment to doing the right thing. We are proud to be an Equal Opportunity Employer.

To understand more about the data we collect and process as part of your application, please view our Backblaze Employee Privacy Notice.

Similar Jobs

26 Minutes Ago
Remote or Hybrid
Junior
Junior
Consumer Web • eCommerce • Information Technology • Retail • Software • Analytics • App development
Coordinate end-to-end installation projects for customers by tracking progress, scheduling within SLAs, documenting interactions, resolving work order issues, and communicating with customers, service providers, stores and vendors. Maintain compliance documentation, use systems (Installation Management System, myRedVest, Salesforce), deliver customer support via inbound/outbound calls, and adapt to process changes while meeting performance goals.
Top Skills: Installation Management SystemMyredvestSalesforce
31 Minutes Ago
Remote or Hybrid
140K-175K Annually
Entry level
140K-175K Annually
Entry level
Artificial Intelligence • Big Data • Cloud • Information Technology • Software • Big Data Analytics • Automation
Serve as the public face and voice of Bluebox, building developer trust and community presence from the ground up. Create technical content, livestreams, videos, workshops, conference talks, and meetups. Help users share product stories, collaborate with influencers and community partners, identify onboarding and documentation friction, and relay evidence-based feedback to Product and Engineering. Success is measured by sign-ups and meaningful product usage.
Top Skills: Ai Coding AgentsAWSBlueskyDiscordGCPHacker NewsRedditX
31 Minutes Ago
Remote or Hybrid
45K-85K Annually
Junior
45K-85K Annually
Junior
Artificial Intelligence • Fintech • Insurance • Marketing Tech • Software • Analytics
Handles inbound calls and warm leads, consults customers on insurance needs, recommends appropriate property and casualty coverages, and converts prospects into policyholders. The role includes paid training and licensing, customer communication, sales closing, and representing the Liberty Mutual brand. Employees work remotely on assigned evening and weekend schedules, maintain a dedicated workspace and high-speed wired internet, and must obtain a Property and Casualty insurance license after hire.
Top Skills: Cable/Fiber/Dsl InternetPcWired High-Speed Internet

What you need to know about the Los Angeles Tech Scene

Los Angeles is a global leader in entertainment, so it’s no surprise that many of the biggest players in streaming, digital media and game development call the city home. But the city boasts plenty of non-entertainment innovation as well, with tech companies spanning verticals like AI, fintech, e-commerce and biotech. With major universities like Caltech, UCLA, USC and the nearby UC Irvine, the city has a steady supply of top-flight tech and engineering talent — not counting the graduates flocking to Los Angeles from across the world to enjoy its beaches, culture and year-round temperate climate.

Key Facts About Los Angeles Tech

  • Number of Tech Workers: 375,800; 5.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Snap, Netflix, SpaceX, Disney, Google
  • Key Industries: Artificial intelligence, adtech, media, software, game development
  • Funding Landscape: $11.6 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Strong Ventures, Fifth Wall, Upfront Ventures, Mucker Capital, Kittyhawk Ventures
  • Research Centers and Universities: California Institute of Technology, UCLA, University of Southern California, UC Irvine, Pepperdine, California Institute for Immunology and Immunotherapy, Center for Quantum Science and Engineering

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account