SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Solve Intelligence is a fast-growing AI startup transforming the intellectual property industry. The company has achieved 20-30% MoM revenue growth, serves 700+ global IP teams (including DLA Piper and major enterprises), and recently closed a $40M Series B (total funding $55M) from top-tier investors including Y Combinator, Microsoft, and Thomson Reuters. Users report 50-90% efficiency gains using the platform.
You will build the ingestion and search systems powering Solve Intelligence's AI products. The data sources span global patent literature, case law, technical standards, scientific databases, academic papers, and web content—encompassing structured records, documents, images, audio, and video.
Key responsibilities:
• Large-scale ingestion: Design and operate high-throughput, resumable pipelines for large datasets with efficient incremental updates, monitoring, and failure recovery.
• Document processing and data quality: Extract useful content from complex documents and formats; handle malformed records and schema changes; validate outputs while preserving structure and metadata.
• Search and serving: Build keyword, vector, and structured search systems; design schemas, indexes, and partitioning for fast queries over tens to hundreds of millions of records.
• Data linking: Connect patents, scientific records, and supporting documents across sources; preserve dates and versions; ensure results are traceable to original sources.
• Performance engineering: Profile parsing, ingestion, database builds, and queries at realistic scale; diagnose CPU, memory, and storage I/O bottlenecks; tune jobs and infrastructure for throughput, latency, and cost.
You will own systems end-to-end from source acquisition through query serving, working closely with AI researchers and product engineers with substantial autonomy to choose approaches and build systems.
The founding team includes a PhD AI researcher (CRO, ex-Huawei R&D), a PhD AI researcher and published academic (CEO, ex-Alan Turing Institute), and a systems engineer (CTO, ex-Qualcomm).
Requirements:
Must have:
- Strong Python and SQL with experience designing and operating production databases
- Solid experience building and operating production data pipelines over large, messy datasets
- Expertise running search systems over large document collections
- End-to-end ownership from raw data to user-facing functionality
- Good understanding of schema design, indexing, and query optimization
- Track record of diagnosing and fixing performance bottlenecks in live systems through profiling and measurement
Nice to have:
- Experience with PostgreSQL/pgvector, OpenSearch (or Elasticsearch), Spark/Delta Lake, AWS, NoSQL databases, or Rust/C++