GiveCampus Logo

GiveCampus

Site Reliability Engineer

Posted 5 Days Ago
Remote
Hiring Remotely in United States
Senior level
Remote
Hiring Remotely in United States
Senior level
Operate and improve AWS production infrastructure, infrastructure as code, Kubernetes workloads, observability, CI/CD, and incident response. The role investigates root causes, strengthens application resilience, automates operational tasks, supports database and performance reliability, maintains documentation, and participates in 24/7 on-call rotations. The engineer owns scoped reliability projects and collaborates with product engineering teams on resilient, secure, and compliant systems.
The summary above was generated by AI

GiveCampus is the world's leading fundraising platform for non-profit educational institutions. Trusted by millions of donors and 1,300+ colleges, universities, and K-12 schools, our mission is to help advance the quality, the affordability, and the accessibility of education. At our current pace, we will facilitate $100 billion in charitable giving over the next decade–enough money to send more than 1 million students to college, tuition-free.

GiveCampus is backed by leading investors including Y Combinator, but we’re also practitioners of Sustainable Growth: we’ve made the Inc. 5000 list of America's fastest-growing private companies each of the last five years and we’ve been profitable nine of the last 10. In 2025, we celebrated a $140 million growth investment that included a major liquidity event for GiveCampus employees–the second in less than three years. 

Our purpose-driven team of 130+ is located in 30+ states across the US: team members work from anywhere they choose. We have a beautiful 12,000sf office in Washington, DC that is available for people to use whenever they want, and we regularly organize team meet-ups, visit partner institutions, and host retreats in various locations. 

While we operate at meaningful scale, we’re still small relative to the commercial and social good opportunities in front of us. Every GiveCampus employee plays a meaningful role in shaping what comes next, and we're growing the team in support of our ambitious plans–including a $100 million investment in AI product development. If you believe in the transformative power of education and want to join a fast-growing, mission-driven company, you’ll fit right in.

Location: This is a remote-first role based in the U.S. While we embrace flexible, distributed work, we also value in-person connection. Team members are expected to attend multiple company-wide and team-specific onsites throughout the year.

About the role

GiveCampus is looking for a hands-on Site Reliability Engineer to help improve the reliability, performance, and operational maturity of our platform.

With our migration to AWS complete, this role will focus on operating and strengthening our production environment: improving observability, automating infrastructure and operational work, responding to incidents, and partnering with product engineers to build resilient systems.

You will own well-scoped reliability projects and contribute to larger cross-functional initiatives. You should be comfortable working independently on straightforward problems, asking for guidance when needed, explaining technical tradeoffs, and keeping teammates informed at important milestones.

What you'll do
  • Operate, maintain, and improve production infrastructure in AWS.
  • Build and maintain infrastructure as code using Terraform.
  • Support workloads running on Kubernetes and Amazon EKS.
  • Improve dashboards, alerts, and service-level indicators using New Relic or comparable observability platforms.
  • Investigate production issues, identify root causes, and implement durable fixes.
  • Participate in the shared 24/7 on-call rotation and contribute to effective incident response.
  • Participate in blameless postmortems and complete follow-up actions that reduce the likelihood or impact of repeat incidents.
  • Partner with product engineers to troubleshoot performance and reliability issues throughout the application stack.
  • Improve application resilience using established patterns such as timeouts, retries, queuing, backpressure, and idempotency.
  • Maintain and improve CI/CD pipelines and deployment workflows using tools such as GitHub Actions and CircleCI.
  • Automate repetitive operational tasks and identify opportunities to reduce engineering toil.
  • Create and maintain runbooks, system diagrams, troubleshooting guides, and production documentation.
  • Contribute to capacity planning, performance testing, database reliability, and production-readiness reviews.
  • Apply established security, access-control, logging, and compliance practices to infrastructure work.
  • Own small-to-medium reliability improvements from technical design and work breakdown through delivery.
  • Communicate progress, risks, tradeoffs, and blockers clearly while incorporating feedback from engineering partners.
What we're looking for
  • Approximately 5+ years of related experience in software engineering, infrastructure, systems engineering, SRE, Platform Engineering, DevOps, or equivalent practical experience.
  • Hands-on experience operating production workloads in AWS.
  • Experience building or maintaining infrastructure using Terraform or a similar infrastructure-as-code tool.
  • Experience with New Relic, Datadog, or another modern observability platform.
  • Experience troubleshooting production incidents and participating in an on-call rotation.
  • Experience building or maintaining CI/CD pipelines.
  • Software development or scripting experience, with the ability to read, debug, and make targeted changes to application or automation code.
  • Working knowledge of Linux, networking, distributed systems, and relational databases.
  • Ability to articulate root causes, explain technical tradeoffs, and translate findings into practical solutions.
  • Ability to manage a well-scoped project with general direction and provide timely updates at key milestones.
  • Strong written and verbal communication skills and a collaborative approach to working across engineering disciplines.
  • A habit of automating repetitive work and improving the reliability of the systems you support.
Bonus points
  • Experience with Ruby or Ruby on Rails.
  • PostgreSQL administration or performance-tuning experience.
  • Experience with Kubernetes and Amazon EKS.
  • Experience with Redis, OpenSearch, or Amazon RDS.
  • Experience operating enterprise SaaS products at scale.
  • Familiarity with SLOs, SLIs, error budgets, capacity modeling, or load testing.
  • Experience with payments, fintech, or other regulated systems.
  • Experience supporting SOC 2 or similar security and compliance programs.
Ready to apply?

Be sure to keep an eye on your spam and promotions boxes in case our emails end up there!

At GiveCampus, we value diversity and we pledge to foster an environment of support, inclusivity, and learning, both on the job and throughout the application process. In this spirit, we encourage candidates of all backgrounds to apply.

GiveCampus is an Equal Opportunity Employer. Applicants and employees are not discriminated against because of race, color, creed, sex, sexual orientation, gender identity or expression, age, religion, national origin, citizenship status, disability, ancestry, marital status, veteran status, medical condition or any protected category prohibited by local, state or federal laws.

If you feel like you don't meet all of the requirements for this role, please apply anyways. We know confidence gaps and imposter syndrome often get in the way of connecting with incredible people, and we don't want them to prevent us from meeting you.

Similar Jobs

2 Days Ago
In-Office or Remote
92K-164K Annually
Senior level
92K-164K Annually
Senior level
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Designs and operates secure, reliable Azure cloud platforms using Terraform, GitHub Actions, containers, and automation. Responsibilities include CI/CD, observability, incident response, platform security, vulnerability remediation, disaster recovery, infrastructure troubleshooting, and SRE practices. The role supports production workloads, improves reliability and delivery processes, participates in on-call activities, and mentors engineers while partnering across development, security, architecture, and operations teams.
Top Skills: BashCi/CdCloud SecurityDockerGitGithub ActionsGitopsInfrastructure As CodeKubernetesAzureObservabilityPowershellPythonTerraform
2 Days Ago
Remote or Hybrid
United States
111K-180K Annually
Senior level
111K-180K Annually
Senior level
Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Leads architecture, modernization, optimization, and reliability initiatives for mainframe CICS, MQ, and z/OS Connect environments. Provides technical direction across development and operations teams, establishes governance and change processes, tunes performance using telemetry, resolves incidents, and develops modernization roadmaps. Collaborates with stakeholders and enterprise architects to deliver secure, scalable, high-availability solutions while evaluating automation, cloud integration, and AI technologies.
Top Skills: AnsibleCicsCobolDevOpsIbm MqIbm Z/OsOpenshiftPythonRed Hat Ansible Automation PlatformZ/Os ConnectZlinux
2 Days Ago
Remote or Hybrid
United States
111K-180K Annually
Senior level
111K-180K Annually
Senior level
Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Leads the architecture, modernization, resilience, security, and performance optimization of enterprise mainframe environments. Responsibilities include z/OS performance tuning, WLM and RACF administration, business continuity planning, automation, technical governance, incident resolution, stakeholder collaboration, and guidance of cross-functional engineering and operations teams. The role also evaluates cloud, DevOps, AI, and hybrid IT technologies for mainframe transformation.
Top Skills: AnsibleCsmGlobal MirrorIbm Z/OsMetro MirrorOpenshiftPr/SmPythonRacfRed Hat Ansible For Ibm Z CollectionsRmfSmfWlmZlinux

What you need to know about the Los Angeles Tech Scene

Los Angeles is a global leader in entertainment, so it’s no surprise that many of the biggest players in streaming, digital media and game development call the city home. But the city boasts plenty of non-entertainment innovation as well, with tech companies spanning verticals like AI, fintech, e-commerce and biotech. With major universities like Caltech, UCLA, USC and the nearby UC Irvine, the city has a steady supply of top-flight tech and engineering talent — not counting the graduates flocking to Los Angeles from across the world to enjoy its beaches, culture and year-round temperate climate.

Key Facts About Los Angeles Tech

  • Number of Tech Workers: 375,800; 5.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Snap, Netflix, SpaceX, Disney, Google
  • Key Industries: Artificial intelligence, adtech, media, software, game development
  • Funding Landscape: $11.6 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Strong Ventures, Fifth Wall, Upfront Ventures, Mucker Capital, Kittyhawk Ventures
  • Research Centers and Universities: California Institute of Technology, UCLA, University of Southern California, UC Irvine, Pepperdine, California Institute for Immunology and Immunotherapy, Center for Quantum Science and Engineering

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account