SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 190,000 - 230,000 / annual
Syllo is a unified litigation platform that leverages language models and agentic AI to transform how lawyers and paralegals work throughout the litigation lifecycle. The company serves enterprise customers including major law firms and corporations, and is rapidly scaling.
You will own the advanced search, indexing, and data scanning infrastructure as Syllo scales to multi-petabyte data volumes. The retrieval stack is proven but faces exponential complexity as ingest sizes grow—you will lead the optimization and architectural evolution to handle this scale.
Key responsibilities include:
• Scale the Retrieval Stack: Lead optimization and architectural evolution of the hybrid search infrastructure, maximizing throughput and efficiency of both lexical search (Elasticsearch, Lucene) and dense vector databases.
• Advanced Data Tiering & Scanning: Design and implement intelligent, cost-effective tiering strategies across hot, warm, and cold data states. Evolve distributed pipelines to efficiently execute asynchronous, massive-scale scans of petabytes of data.
• Relentless Optimization: Drive down latency and cost-to-serve by deeply analyzing system bottlenecks, tuning indexing and querying algorithms, and optimizing cloud infrastructure (compute, storage, networking) for maximum efficiency at extreme scale.
• Technical Leadership: Act as the domain expert and owner of the indexing and search ecosystem. Set long-term technical vision for data storage and retrieval, guiding engineering teams on best practices for high-volume data modeling and performance tuning.
• Resiliency at Scale: Ensure fault-tolerant, highly available operations during massive parallel ingest events and complex, concurrent querying across millions of documents.
Required qualifications: 8+ years of software engineering experience at Staff/Principal level optimizing and scaling highly distributed, high-throughput systems to petabyte-level data. Deep, production-level expertise tuning and scaling Lucene-based search engines (Elasticsearch, Solr) and modern vector indexing infrastructure. Strong history of managing compute vs. storage trade-offs and designing cost-effective cold-storage scanning solutions. Extensive experience with complex data pipelines, high-throughput event streaming (Kafka, Kinesis), and distributed compute architectures. Expert command of cloud primitives (GCP preferred), Kubernetes, and infrastructure-as-code. Expert-level proficiency in systems-level and backend languages (Go, Rust, Python, Java, or C++).