Accelerant Logo

Accelerant

Senior SRE

Reposted 22 Days Ago
Remote
Hiring Remotely in United States
Senior level
Remote
Hiring Remotely in United States
Senior level
Lead reliability and observability for the financial data platform: define SLOs/SLIs, build metrics pipelines, extend instrumentation across Velocity, Redpanda, MuleSoft, Snowflake, Fabric, and AWS; implement incident management (Datadog -> Incident.io -> ServiceNow), scale automation and remediation, and design AI-assisted SRE agents using Cursor for triage and root-cause analysis.
The summary above was generated by AI

About Accelerant

Accelerant is a data-driven risk exchange connecting underwriters of specialty insurance risk with risk capital providers. Accelerant was founded in 2018 by a group of longtime insurance industry executives and technology experts who shared a vision of rebuilding the way risk is exchanged – so that it works better, for everyone. The Accelerant risk exchange does business across more than 20 different countries and 250 specialty products, and we are proud that our insurers have been awarded an AM Best A- (Excellent) rating. For more information, please visit www.accelerant.ai.

The Role

We're building the financial data platform at Accelerant — the premium, claims, and paid data products that underpin financial processing, reserving analysis, and the monthly close — and it needs to stay fast, resilient, and observable as we scale. You'll drive the reliability and observability strategy across the platform and the enterprise systems it depends on: Velocity, MuleSoft, D365, Snowflake, Fabric, and the streaming and integration layers that move data through it. You are a key decider about what gets measured, how we define reliability, and where engineering needs to invest to keep production healthy.

We need someone who can prove a repeatable, define-to-alert observability pipeline, harden it, and scale it into systems that have never had real SLOs — and build modern, AI-assisted operational tooling that lets a small team punch far above its weight.

This Is a High-Autonomy, High-Impact Role for Someone Who:

      Sees a recurring alert or a fragile deploy path and cannot leave it alone. Excels at shipping the right fix and the right automation, not the perfect one.

      Has run real production systems at scale — not just written runbooks for them.

      Has built with Datadog, OpenTelemetry, incident tooling, and AI coding assistants long enough to have strong opinions about what fits our needs.

      Can prototype an operational agent in Cursor and iterate as they go.

      Operates with autonomy, and can carry a technical discussion on system architecture, failure modes, and tradeoffs.

      Is genuinely curious about applying emerging AI to reliability and operations.

What You'll Do

Drive the reliability and observability initiative

      Own the reliability roadmap end to end. Prove a repeatable define → emit → ingest → dashboard → alert metric pipeline, set SLOs and error budgets, prioritize the work, and drive execution. You'll partner with engineering on what we monitor, how, and when — indexing on user impact over low-level infrastructure.

Harden the foundational platform

      Take the financial data platform from functional to enterprise-grade, with a focus on availability, performance, and recoverability. Strengthen deployment paths, straight-through processing, and failover so the monthly close runs faster and cleaner as legacy hops are retired.

Expand observability breadth and depth

      Extend instrumentation across the six target systems — Velocity, Red Panda, MuleSoft, Snowflake, Fabric, and AWS (with D365 ledger to follow) — proving both push (OpenTelemetry) and pull (agent) ingestion. Cover service health (latency, error rates, throughput) and business KPIs (match rate, reconciliation completeness, settlement correctness and latency).

Implement a scalable incident and review process

      Build the on-call, alerting, and blameless postmortem process that keeps reliability high as systems and the team grow. Route alerts Datadog → Incident.io with ServiceNow as the system of record, and set severity standards, escalation norms, and follow-up tracking that actually closes the loop.

Scale automation, auditability, and reduce toil

      Build the tooling that automates routine operations, self-heals common failures, and surfaces signal over noise. Establish data lineage and retention, and validate reliability at scale — 5,000+ transactions before go-live — through auto-remediation, capacity planning, and actionable dashboards.

Build specialized SRE agents using Cursor AI

      Design and ship AI agents for incident triage, log analysis, and root-cause investigation (to name a few). Use Cursor as your build environment. Treat the agents as products solving specific problems.

Host SRE agents on the AI fabric

      Partner with the AI platform team to deploy your agents on the org's AI fabric. Make them discoverable, governed, and reusable across functions.

What You'll Bring

Must-Haves

      Proven experience designing, operating, and scaling reliable production systems.

      Deep hands-on expertise with modern observability tooling — Datadog, Prometheus/Grafana, and OpenTelemetry — including both push and pull ingestion patterns.

      Strong background defining SLIs, SLOs, and error budgets — and translating them into business-level KPIs, not just infrastructure metrics.

      Experience operating data platforms (Snowflake, Fabric) and enterprise integration layers (MuleSoft) alongside enterprise SaaS such as D365 (F&O and/or Power Apps).

      Hands-on incident management experience with tools like Incident.io and ServiceNow, and a track record of running effective on-call and postmortem practices.

      Hands-on experience building with LLMs and AI coding assistants — Cursor in particular. Bonus if you've built and deployed agents.

      Ability to define reliability strategy, reliability targets, and operational metrics — and defend them to engineering leadership and the business.

      Strong communication skills — you can explain a root cause to a junior engineer and a reliability risk to a product lead.

      Demonstrated bias for action and ability to operate autonomously in ambiguous, fast-changing environments.

Nice-to-Haves

      Experience in insurance, fintech, or other regulated financial services industries.

      Familiarity with insurance and finance concepts (premium, claims, settlement, reserving, monthly close) or willingness to learn them deeply.

      Experience with streaming and event pipelines (Red Panda / Kafka) and data lineage, retention, and auditability requirements.

      Strong working knowledge of chaos engineering, performance and load testing, and capacity planning.

      Experience deploying AI agents on an internal AI platform or fabric (governance, eval harnesses, prompt/version management).

Similar Jobs

17 Days Ago
Remote or Hybrid
USA
140K-215K Annually
Senior level
140K-215K Annually
Senior level
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Senior SRE owning availability, automation, and observability for CI/CD platform services. Build and operate infrastructure, run on-call, lead incident response, mentor engineers, drive design/capacity planning, integrate AI-assisted workflows, and improve cross-team reliability.
Top Skills: Active DirectoryAnsibleApache AirflowSparkAWSAzureBashBazelBitbucketCassandraChefDatadogDnsFirewall RulesGCPGitGithub ActionsGitlabGitlab CiGoGrafanaHoneycombHumio/LogscaleJenkinsKafkaKubernetesLoad BalancersMongoDBMySQLNasNew RelicNfsObject StorageOpensearchOraclePostgresPowershellPrometheusPulsarPuppetPythonRabbitMQRedis/ValkeyRedpandaRoutingSaltSanSplunkTerraformVarnishVipsWindows Server
18 Days Ago
In-Office or Remote
153K-205K Annually
Senior level
153K-205K Annually
Senior level
Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3
Build and operate scalable cloud-native microservices and Kubernetes infrastructure, improve CI/CD and developer workflows, engineer and run an autonomous coding-agent orchestration platform, integrate and operationalize AI services with guardrails, implement observability and incident response, and collaborate with product and engineering teams to ensure secure, reliable production systems and cost optimization.
Top Skills: Ai ApisAutonomous AgentsAWSCi/CdGCPGoJavaJavaScriptKubernetesMonitoringObservabilityPythonRestful ApisRustSdksSQLTypescriptWorkflow Orchestration
20 Days Ago
Remote or Hybrid
130K-160K Annually
Senior level
130K-160K Annually
Senior level
AdTech • Cloud • Digital Media • Information Technology • News + Entertainment • App development
Design, deploy, and maintain on-premises and cloud playout infrastructure for IP video distribution. Build automation, CI/CD pipelines, monitoring, and scalable fault-tolerant systems. Drive releases, troubleshoot broadcast incidents, mentor SREs, and provide 24/7 on-call support.
Top Skills: AnsibleAWSAzureBashBroadcast TechnologiesCi/CdContainerizationGCPIp VideoJavaScriptKubernetesLinuxPerlPythonRubyStreamingTerraform

What you need to know about the Los Angeles Tech Scene

Los Angeles is a global leader in entertainment, so it’s no surprise that many of the biggest players in streaming, digital media and game development call the city home. But the city boasts plenty of non-entertainment innovation as well, with tech companies spanning verticals like AI, fintech, e-commerce and biotech. With major universities like Caltech, UCLA, USC and the nearby UC Irvine, the city has a steady supply of top-flight tech and engineering talent — not counting the graduates flocking to Los Angeles from across the world to enjoy its beaches, culture and year-round temperate climate.

Key Facts About Los Angeles Tech

  • Number of Tech Workers: 375,800; 5.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Snap, Netflix, SpaceX, Disney, Google
  • Key Industries: Artificial intelligence, adtech, media, software, game development
  • Funding Landscape: $11.6 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Strong Ventures, Fifth Wall, Upfront Ventures, Mucker Capital, Kittyhawk Ventures
  • Research Centers and Universities: California Institute of Technology, UCLA, University of Southern California, UC Irvine, Pepperdine, California Institute for Immunology and Immunotherapy, Center for Quantum Science and Engineering

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account