Syllo Logo

Syllo

Staff Software Engineer, Search & Retrieval Infrastructure

Posted 8 Days Ago
Remote
Hiring Remotely in USA
190K-230K Annually
Senior level
Remote
Hiring Remotely in USA
190K-230K Annually
Senior level
Own and scale the search, indexing, and data-scanning infrastructure for petabyte-scale data. Optimize hybrid lexical and vector retrieval for sub-second latency, design cost-effective hot/warm/cold data tiering and massive asynchronous scans, drive down latency and cost, and provide technical leadership for resilient, highly available indexing and retrieval systems.
The summary above was generated by AI

About Syllo 

Syllo is on a mission to transform litigation. Our product is a unified litigation platform that enables lawyers and paralegals to safely harness the power of language models and agentic AI throughout the litigation life cycle. Since going to market, we have gained a diverse group of enterprise customers, including some of the biggest law firms and corporations in the country, and we are quickly expanding. By reducing the expense of litigation industry-wide, we aim to improve access to high-quality representation and promote the alignment of legal outcomes with merit. 


About the Role 

We are seeking a Staff Software Engineer to take ownership of our advanced search, indexing, and data scanning infrastructure as we scale to the next echelon of data volume.

Our retrieval stack is robust and proven, but as our ingest sizes push into the multi-petabyte range, the complexity of balancing speed, availability, and cost increases exponentially. You will own this constellation of scaling challenges. Your mission is to continuously optimize and evolve our systems—ensuring our hot indexes maintain sub-second latency for interactive workflows, while simultaneously designing highly concurrent, cost-effective architectures for deep scanning and vectorizing massive volumes of cold-storage data. You will lead the design and implementation of sophisticated data tiering and retrieval strategies that keep our platform operating at peak performance without inflating cloud compute costs.


Responsibilities 

  • Scale the Retrieval Stack: Lead the optimization and architectural evolution of our existing hybrid search infrastructure, maximizing the throughput and efficiency of both lexical search (e.g., Elasticsearch, Lucene) and dense vector databases.
  • Advanced Data Tiering & Scanning: Design and implement intelligent, cost-effective tiering strategies across hot, warm, and cold data states. Evolve our distributed pipelines to efficiently execute asynchronous, massive-scale scans of petabytes of data in varying states of availability.
  • Relentless Optimization: Drive down latency and cost-to-serve. Deeply analyze system bottlenecks, tune indexing and querying algorithms, and optimize cloud infrastructure (compute, storage, and networking) for maximum efficiency at extreme scale.
  • Technical Leadership: Act as the domain expert and owner of the indexing and search ecosystem. Set the long-term technical vision for data storage and retrieval, guiding engineering teams on best practices for high-volume data modeling and performance tuning.
  • Resiliency at Scale: Ensure fault-tolerant, highly available operations during massive parallel ingest events and complex, concurrent querying across millions of documents.

Qualifications 

  • Extreme Scale Experience: 8+ years of software engineering experience, with a proven track record operating at the Staff/Principal level optimizing and scaling highly distributed, high-throughput systems to handle petabyte-level data.
  • Search & Vector Mastery: Deep, production-level expertise tuning and scaling Lucene-based search engines (Elasticsearch, Solr) and modern vector indexing infrastructure. You deeply understand index internals, chunking strategies, and embedding retrieval optimization.
  • Cost-Aware Architecture: A strong history of managing the compute vs. storage trade-off. You know how to design sophisticated cold-storage scanning solutions and hot-index architectures that are highly performant but fundamentally cost-effective.
  • Distributed Systems: Extensive experience managing complex data pipelines, high-throughput event streaming (Kafka, Kinesis), and distributed compute architectures handling billions of records.
  • Cloud Infrastructure: Expert command of cloud primitives (GCP preferred), Kubernetes, and infrastructure-as-code.
  • Languages: Expert-level proficiency in systems-level and backend languages (Go, Rust, Python, or Java/C++).

Salary Range ($190- $230K) plus health insurance and equity. 

United States - Remote Pay Range
$190,000$230,000 USD

Similar Jobs

12 Days Ago
Remote
US
190K-270K Annually
Senior level
190K-270K Annually
Senior level
Artificial Intelligence
Design and build scalable backend components and indexing pipelines for semantic and hybrid retrieval, build retrieval orchestration and knowledge-graph services, improve retrieval quality via evaluation and observability, design APIs, and optimize latency, throughput, cost, reliability, and security for large-scale AI inference and retrieval workloads.
Top Skills: C++ElasticEmbeddingsGoHybrid RetrievalJavaKnowledge GraphKubernetesLlmsObservability FrameworksOpensearchPineconePulumiPythonRagRustSemantic SearchTerraformVector Databases
An Hour Ago
Remote or Hybrid
44K-100K Annually
Junior
44K-100K Annually
Junior
Artificial Intelligence • Fintech • Insurance • Marketing Tech • Software • Analytics
Handle inbound/warm insurance sales leads, consult customers on Property & Casualty coverage, convert leads to policyholders, complete paid licensing and training, work assigned shifts in a remote home-office with provided equipment, and meet sales goals while maintaining strong customer service and integrity.
2 Hours Ago
In-Office or Remote
215K-358K Annually
Senior level
215K-358K Annually
Senior level
Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
Lead strategy and investment for Pfizer Oncology patient solutions platforms and co-pay/patient support programs. Transform programs from reactive to AI-enabled predictive support, set outcome-based KPIs, manage the US & Global Patient Experience team, own budgets and vendor selection, partner cross-functionally, and represent Pfizer externally on patient support innovation.
Top Skills: AIConversational AiPredictive Analytics

What you need to know about the Los Angeles Tech Scene

Los Angeles is a global leader in entertainment, so it’s no surprise that many of the biggest players in streaming, digital media and game development call the city home. But the city boasts plenty of non-entertainment innovation as well, with tech companies spanning verticals like AI, fintech, e-commerce and biotech. With major universities like Caltech, UCLA, USC and the nearby UC Irvine, the city has a steady supply of top-flight tech and engineering talent — not counting the graduates flocking to Los Angeles from across the world to enjoy its beaches, culture and year-round temperate climate.

Key Facts About Los Angeles Tech

  • Number of Tech Workers: 375,800; 5.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Snap, Netflix, SpaceX, Disney, Google
  • Key Industries: Artificial intelligence, adtech, media, software, game development
  • Funding Landscape: $11.6 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Strong Ventures, Fifth Wall, Upfront Ventures, Mucker Capital, Kittyhawk Ventures
  • Research Centers and Universities: California Institute of Technology, UCLA, University of Southern California, UC Irvine, Pepperdine, California Institute for Immunology and Immunotherapy, Center for Quantum Science and Engineering

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account