Zefr Logo

Zefr

Site Reliability Engineer

Posted 3 Days Ago
Hybrid
Marina del Rey, CA, USA
100K-130K Annually
Junior
Hybrid
Marina del Rey, CA, USA
100K-130K Annually
Junior
Support and improve cloud infrastructure, Kubernetes-based microservices, CI/CD pipelines, infrastructure-as-code, observability, and production reliability. The role involves debugging application and infrastructure issues, participating in a 24/7 on-call rotation, maintaining documentation and runbooks, and collaborating with engineering and data science teams to deliver scalable, automated systems.
The summary above was generated by AI
What we do:

Zefr is the leader in AI-powered content classifications for brands and advertisers. Zefr’s platform is purpose built for multi-modal content understanding on open platforms like YouTube, TikTok, Meta and Snap, with pre-bid activation and verification solutions. Our products safeguard media and AI investments, while maximizing performance and efficacy on those channels.

Headquartered in Los Angeles with global offices across New York, Chicago, London, Toronto, Singapore, and more, Zefr is redefining what trust and transparency means for social media in the age of AI.

 
What you’ll do:

As a Site Reliability Engineer at Zefr, you’ll join our Engineering team and grow your skills in cloud infrastructure, CI/CD, Observability, and core SRE concepts, helping deliver high-quality, reliable, and scalable solutions. You’ll work closely with the rest of Zefr’s Engineering and Data Science teams, ensuring the infrastructure behind our services is robust, efficient, and scalable.

We are looking for a passionate engineer to come collaborate and grow alongside our experienced SRE team. You bring curiosity, a strong work ethic, and a passion for continuous improvement and automation.

  • Support and build systems and tools that enable other engineers to deploy and manage product features both quickly and safely.

  • Help deploy and support our multi-cloud, micro-service architecture, deployed via Github Actions, ArgoCD & Kubernetes.

  • Write and maintain Infrastructure as Code with Terraform and Terragrunt, contributing changes through pull requests and code review.

  • Build and improve CI/CD pipelines and release workflows.

  • Help maintain the health of production environments, including monitoring application performance and resource utilization.

  • Participate in 24/7 on-call rotation, responding to system performance issues and outages alongside senior teammates.

  • Debug issues at the application and infrastructure level.

  • Write and maintain clear documentation and runbooks.

  • Contribute to our DevOps culture and philosophy of continuous improvement.

Technology Stack at Zefr:
  • Cloud Providers: Google Cloud Platform (primary), Amazon Web Services

  • Infrastructure as Code (IaC): Terraform, Terragrunt

  • Containerization & Orchestration: Docker, Kubernetes (GKE)

  • CI/CD: GitHub Actions, Argo CD

  • Primary Language: Python

  • Observability: Prometheus, OpenTelemetry, Chronosphere, Pagerduty

  • Application Languages/Frameworks: Python, FastAPI, Flask, Node.js, React

  • Workflow Orchestration: Apache Airflow, Ray

  • Relational Databases: PostgreSQL

  • NoSQL Databases: DynamoDB

  • Search Databases: OpenSearch

  • Data Warehousing: Snowflake

 
What we’re looking for:
 
  • 1-3 years of experience supporting Cloud Infrastructure in a production environment using AWS and/or GCP

  • Hands-on experience with containers and Kubernetes

  • Competency in Python and shell scripting

  • Strong problem-solving skills

  • Strong written and verbal communication, organization, and documentation skills

 
Nice to have:
  • Familiarity with modern CI/CD pipelines and GitOps (Github Actions, GitLab, Argo CD)

  • Exposure to Monitoring and Observability tools (Prometheus, Grafana, Chronosphere, Datadog, OpenTelemetry)

  • Experience collaborating across departments to execute complex, high-impact projects.

 
Benefits (for US based employees):
  • Flexible PTO

  • Medical, dental, and vision insurance with FSA options

  • Company-paid life insurance

  • Paid parental leave

  • 401(k) with company match

  • Professional development opportunities

  • 13 paid holidays off

  • Summer Fridays (we leave early)

  • In-office and hybrid work options available

  • In-office lunches and lots of free food

  • Optional in-person and virtual events (we like to celebrate!)

 
 
Compensation (for US based employees):

The anticipated salary for this position is between $100,000 and $130,000. Within the range, individual pay is determined by factors such as job-related skills, experience, and relevant education or training. If your compensation expectations fall outside of this range, it may still be worth having a conversation.

 

Zefr is an equal opportunity employer that embraces diversity and inclusion in the workplace. We are committed to building a team that represents a variety of backgrounds, skills, and perspectives because we know this only makes us better. We strongly encourage women, persons of color, LGBTQIA+ individuals, persons with disabilities, members of ethnic minorities, foreign-born residents, and veterans to apply even if you do not meet 100% of the qualifications.

HQ

Zefr Venice, California, USA Office

Venice, CA, United States

Zefr Los Angeles, California, USA Office

Los Angeles, United States, 0

Zefr Marina del Rey, California, USA Office

Marina del Rey, CA, United States, 90066

Similar Jobs

4 Days Ago
Hybrid
176K-308K Annually
Mid level
176K-308K Annually
Mid level
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Build and operate highly available, cloud-native Voice AI systems, including deployments, observability, monitoring, disaster recovery, and continuous integration. Develop scalable software, integrate LLMs with voice and real-time communication platforms, automate testing and delivery, and collaborate with product owners. Mentor colleagues and promote knowledge sharing across telecommunications and AI engineering disciplines.
Top Skills: AnsibleClaude CodeConfiguration ManagementCursorGitlab CiGoHelmInfrastructure As CodeJavaKubernetesLlmsPrometheusPythonReal-Time Communication SystemsSoftware-Defined NetworkingSplunkVoice AiWindsurf
5 Days Ago
In-Office
147K-234K Annually
Senior level
147K-234K Annually
Senior level
Fintech • Information Technology • Payments • Sharing Economy • Financial Services • Cryptocurrency
Leads reliability, scalability, performance, and security for large-scale cloud systems. Designs AWS infrastructure with Terraform, automates CI/CD and operational workflows, establishes SLOs, manages incident response and disaster recovery, develops monitoring and observability solutions, and builds internal tools. Partners with engineering teams on architecture and reliability practices, conducts code reviews, mentors SREs, and supports compliance, vulnerability management, and security integration.
Top Skills: Agentic ApplicationsAmazon Api GatewayAmazon AuroraAmazon CloudfrontAmazon CloudwatchAmazon DynamodbAmazon EbsAmazon Ec2Amazon EcsAmazon EfsAmazon RdsAmazon Route 53Amazon S3Amazon VpcAWSAws FargateAws LambdaAws X-RayChaos EngineeringCi/CdDastDatadogDistributed SystemsDockerEvent-Driven SystemsGitlabGitopsGrafanaIamInfrastructure As CodeJavaKubernetesLlmsMicroservicesNew RelicNode.jsOwasp Top 10PythonSastServerless ArchitecturesSplunkTerraform
16 Days Ago
Remote or Hybrid
United States
111K-180K Annually
Senior level
111K-180K Annually
Senior level
Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Leads architecture, modernization, optimization, and reliability initiatives for mainframe CICS, MQ, and z/OS Connect environments. Provides technical direction across development and operations teams, establishes governance and change processes, tunes performance using telemetry, resolves incidents, and develops modernization roadmaps. Collaborates with stakeholders and enterprise architects to deliver secure, scalable, high-availability solutions while evaluating automation, cloud integration, and AI technologies.
Top Skills: AnsibleCicsCobolDevOpsIbm MqIbm Z/OsOpenshiftPythonRed Hat Ansible Automation PlatformZ/Os ConnectZlinux

What you need to know about the Los Angeles Tech Scene

Los Angeles is a global leader in entertainment, so it’s no surprise that many of the biggest players in streaming, digital media and game development call the city home. But the city boasts plenty of non-entertainment innovation as well, with tech companies spanning verticals like AI, fintech, e-commerce and biotech. With major universities like Caltech, UCLA, USC and the nearby UC Irvine, the city has a steady supply of top-flight tech and engineering talent — not counting the graduates flocking to Los Angeles from across the world to enjoy its beaches, culture and year-round temperate climate.

Key Facts About Los Angeles Tech

  • Number of Tech Workers: 375,800; 5.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Snap, Netflix, SpaceX, Disney, Google
  • Key Industries: Artificial intelligence, adtech, media, software, game development
  • Funding Landscape: $11.6 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Strong Ventures, Fifth Wall, Upfront Ventures, Mucker Capital, Kittyhawk Ventures
  • Research Centers and Universities: California Institute of Technology, UCLA, University of Southern California, UC Irvine, Pepperdine, California Institute for Immunology and Immunotherapy, Center for Quantum Science and Engineering

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account