Senior technical leader for SRE driving observability, platform infrastructure, SLIs/SLOs, incident response, automation, and self-service platform capabilities. Shapes reliability strategy, mentors engineers, and ensures production-scale operational excellence.
Filevine is a Legal AI company delivering Legal Operating Intelligence for the future of legal work. Grounded in a singular system of truth, Filevine brings together data, documents, workflows, and teams into one unified platform—where modern legal work happens with clarity and consistency.
Powered by LOIS, the Legal Operating Intelligence System, Filevine connects context across every matter to transform legal operations from reactive to proactive. LOIS reads, understands, and reasons across your data to surface insight, automate complexity, and give professionals the clarity and confidence to see more, know more, and do more. Fueled by a team of exceptional collaborators and innovators, Filevine’s rapid growth has earned AI awards and recognition from Deloitte and Inc. as one of the most innovative and fastest-growing technology companies in the country.
Role Summary
As a Staff Site Reliability Engineer at Filevine, you are the senior technical authority on the SRE team
and a strategic partner to engineering leadership. You don’t just maintain systems — you shape
engineering culture, define the technical standard for how Filevine runs in production, and bridge the
gap between high-level business goals and robust, internet-scale technical execution. You bring a
forward-looking perspective — actively shaping how AI and machine learning drive the future of
reliability practice.
You own the roadmap across two critical SRE domains — Observability & Alerting and Platform
Infrastructure — and are accountable for ensuring the team solves reliability problems permanently
rather than absorbing them as toil. You operate as the senior IC counterpart to the Engineering
Manager: technical correctness lives with you. You partner with the Reliability Architect and engineering
leadership on significant technical decisions, mentor engineers across experience levels, and influence
reliability strategy across the broader organization. Reliability at Filevine protects revenue. You are the
senior technical voice responsible for ensuring that uptime, incident response, and every production
change meet the operational standard the business demands.
This role does not participate in on-call rotation, but you are deeply invested in the engineers who do —
shaping the on-call strategy, tooling, and culture that make production support sustainable and
effective.
and a strategic partner to engineering leadership. You don’t just maintain systems — you shape
engineering culture, define the technical standard for how Filevine runs in production, and bridge the
gap between high-level business goals and robust, internet-scale technical execution. You bring a
forward-looking perspective — actively shaping how AI and machine learning drive the future of
reliability practice.
You own the roadmap across two critical SRE domains — Observability & Alerting and Platform
Infrastructure — and are accountable for ensuring the team solves reliability problems permanently
rather than absorbing them as toil. You operate as the senior IC counterpart to the Engineering
Manager: technical correctness lives with you. You partner with the Reliability Architect and engineering
leadership on significant technical decisions, mentor engineers across experience levels, and influence
reliability strategy across the broader organization. Reliability at Filevine protects revenue. You are the
senior technical voice responsible for ensuring that uptime, incident response, and every production
change meet the operational standard the business demands.
This role does not participate in on-call rotation, but you are deeply invested in the engineers who do —
shaping the on-call strategy, tooling, and culture that make production support sustainable and
effective.
Who You Are
The Technical Authority
• Master of the Craft: You bring deep expertise in distributed systems, cloud infrastructure,
observability, and reliability engineering. You raise the technical standard for every engineer
around you and thrive where the challenges are complex and the stakes are real.
• Technical Leader and Mentor: You are passionate about mentoring engineers and investing in
their growth. You influence technical direction and communicate production risk clearly across
engineering, product, and executive audiences.
• Forward-Thinking & AI/ML Fluent: You bring deep knowledge of AIOps and drive the use of
AI and machine learning in observability, anomaly detection, incident response, automated
remediation, and resource optimization.
• Production-Scale Problem Solver: You turn ambiguous, complex reliability challenges into
durable solutions for systems where availability, performance, and production changes carry
meaningful business impact.
• Software-Minded Builder: You use software, automation, Infrastructure as Code, and platform
capabilities to eliminate toil and make systems safer, more scalable, and easier to operate.
• Master of the Craft: You bring deep expertise in distributed systems, cloud infrastructure,
observability, and reliability engineering. You raise the technical standard for every engineer
around you and thrive where the challenges are complex and the stakes are real.
• Technical Leader and Mentor: You are passionate about mentoring engineers and investing in
their growth. You influence technical direction and communicate production risk clearly across
engineering, product, and executive audiences.
• Forward-Thinking & AI/ML Fluent: You bring deep knowledge of AIOps and drive the use of
AI and machine learning in observability, anomaly detection, incident response, automated
remediation, and resource optimization.
• Production-Scale Problem Solver: You turn ambiguous, complex reliability challenges into
durable solutions for systems where availability, performance, and production changes carry
meaningful business impact.
• Software-Minded Builder: You use software, automation, Infrastructure as Code, and platform
capabilities to eliminate toil and make systems safer, more scalable, and easier to operate.
What you will do
Define and execute the technical strategy for Observability & Alerting, Platform Infrastructure,
and operational excellence.
• Lead the evolution of reliable, scalable, secure, and efficient cloud platforms and distributed
systems.
• Champion SLIs, SLOs, error budgets, capacity planning, operational readiness, and automation
across the service lifecycle.
• Lead the organization through complex production incidents and turn post-incident learning into
permanent engineering improvements.
• Build self-service platform capabilities that reduce toil, improve engineering safety and velocity,
and make every team more capable of owning their own reliability.
• Mentor engineers and serve as a trusted technical authority for long-term reliability and platform
direction.
and operational excellence.
• Lead the evolution of reliable, scalable, secure, and efficient cloud platforms and distributed
systems.
• Champion SLIs, SLOs, error budgets, capacity planning, operational readiness, and automation
across the service lifecycle.
• Lead the organization through complex production incidents and turn post-incident learning into
permanent engineering improvements.
• Build self-service platform capabilities that reduce toil, improve engineering safety and velocity,
and make every team more capable of owning their own reliability.
• Mentor engineers and serve as a trusted technical authority for long-term reliability and platform
direction.
Qualifications
12+ years of experience in software engineering, infrastructure, platform engineering, or SRE,
including 6+ years in SRE and 3+ years leading complex, cross-functional technical initiatives
for distributed production systems.
• Expert-level depth in observability and platform infrastructure, with broad expertise in incident
response, capacity planning, automation, and reliability engineering.
• Advanced experience with a major container-orchestration platform, preferably Kubernetes, and
an observability platform such as New Relic, Datadog, or equivalent.
• Strong software-engineering ability in Python, Go, Bash, or another general-purpose language,
with experience building production tooling, automation, or platform capabilities.
• Proven ability to mentor engineers and communicate technical risk clearly to engineering,
product, and executive audiences.
• Experience in a regulated environment such as FedRAMP, CJIS, HIPAA, SOC 2, or PCI is
strongly preferred.
including 6+ years in SRE and 3+ years leading complex, cross-functional technical initiatives
for distributed production systems.
• Expert-level depth in observability and platform infrastructure, with broad expertise in incident
response, capacity planning, automation, and reliability engineering.
• Advanced experience with a major container-orchestration platform, preferably Kubernetes, and
an observability platform such as New Relic, Datadog, or equivalent.
• Strong software-engineering ability in Python, Go, Bash, or another general-purpose language,
with experience building production tooling, automation, or platform capabilities.
• Proven ability to mentor engineers and communicate technical risk clearly to engineering,
product, and executive audiences.
• Experience in a regulated environment such as FedRAMP, CJIS, HIPAA, SOC 2, or PCI is
strongly preferred.
Cool Company Benefits:
- A dynamic, rapidly growing company, focused on helping organizations thrive
- Medical, Dental, & Vision Insurance (for full-time employees)
- Competitive & Fair Pay
- Maternity & paternity leave (for full-time employees)
- Short & long-term disability
- Opportunity to learn from a dedicated leadership team
- Top-of-the-line company swag
Privacy Policy Notice
Filevine will handle your personal information according to what’s outlined in our Privacy Policy.
Communication about this opportunity, or any open role at Filevine, will only come from representatives with email addresses using "filevine.com". Other addresses reaching out are not affiliated with Filevine and should not be responded to.
Similar Jobs
Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3
Design, build, and operate scalable blockchain infrastructure and Kubernetes platforms. Implement IaC, CI/CD, AI-powered automation, monitoring, incident response, and reliability improvements. Mentor engineers, lead cross-functional initiatives, and support network launches, upgrades, and production troubleshooting in a follow-the-sun on-call rotation.
Top Skills:
Agentic AutomationArcBaseBlue-Green DeploymentCanary ReleasesChaos EngineeringCi/CdCloud-Native ToolingContainerizationControllersDnsEthereumGenerative AiGoHelmInfrastructure As CodeKubernetesLoad BalancersMcp ServersObservability ToolingOperatorsPulumiPythonRbacSolanaSQLTerraformVpc
Artificial Intelligence • Hardware • Software • Semiconductor
Lead automation and platform engineering to eliminate toil and deliver self-service GitOps-driven CD, capacity provisioning, and observability for large-scale inference clusters. Define SLOs/SLIs, mentor SREs, support incident escalation, implement reliability practices, and measure impact via deployment velocity, SLO compliance, MTTR, and adoption of self-service workflows.
Top Skills:
Argo CdBazelCapacity PlanningChaos EngineeringCi/CdGitopsLokiMimirPredictive AutoscalingPrometheusTempoWafer-Scale Engine
Other
Lead architecture and delivery of highly available, resilient cloud and on‑prem systems for the Password Safe platform. Own platform engineering, CI/CD pipelines, IaC/GitOps, release orchestration, observability (metrics/logs/traces), chaos engineering, SLO/SLI definition, and core services. Mentor engineers, define SRE strategy, and drive reliability, security, and automation improvements.
Top Skills:
AnsibleApi GatewaysAWSAzureBlue Green DeploymentsC#CachesCanary DeploymentsChaos EngineeringCi/CdConfiguration ManagementDatadogDevsecopsDockerGitopsGoGrafana CloudJavaKubernetesLinuxOpentelemetryOpentofuSecrets ManagementService MeshTerraformWindows
What you need to know about the Los Angeles Tech Scene
Los Angeles is a global leader in entertainment, so it’s no surprise that many of the biggest players in streaming, digital media and game development call the city home. But the city boasts plenty of non-entertainment innovation as well, with tech companies spanning verticals like AI, fintech, e-commerce and biotech. With major universities like Caltech, UCLA, USC and the nearby UC Irvine, the city has a steady supply of top-flight tech and engineering talent — not counting the graduates flocking to Los Angeles from across the world to enjoy its beaches, culture and year-round temperate climate.
Key Facts About Los Angeles Tech
- Number of Tech Workers: 375,800; 5.5% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Snap, Netflix, SpaceX, Disney, Google
- Key Industries: Artificial intelligence, adtech, media, software, game development
- Funding Landscape: $11.6 billion in venture capital funding in 2024 (Pitchbook)
- Notable Investors: Strong Ventures, Fifth Wall, Upfront Ventures, Mucker Capital, Kittyhawk Ventures
- Research Centers and Universities: California Institute of Technology, UCLA, University of Southern California, UC Irvine, Pepperdine, California Institute for Immunology and Immunotherapy, Center for Quantum Science and Engineering



_1.png)