SlipstreamJobsFresh Startup & VC-Backed Jobs

Machine Learning Engineer

Scarlet - London, United Kingdom - In-office - posted 2026-10-02

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Scarlet is a medical device certification company that uses AI agents and software to accelerate regulatory approval timelines for medical devices without compromising safety. The company combines clinical, technical, and regulatory expertise to help customers reduce certification timelines by a year or more and compress product-update cycles from months to weeks. The Applied Machine Learning team owns production systems and pursues new ideas from conception through deployment and iterative improvement. You will work with clinicians, assessors, and engineers to build reliable ML systems that help bring medical devices to market faster while maintaining safety standards. Key responsibilities include: - Building agentic document understanding systems that search, parse, and visually inspect technical files (scanned certificates, tables, architecture diagrams, evidence documents). The systems must identify relevant information and provide exact source attribution. - Developing custom agent harnesses for retrieving information across large, varied document collections while preserving source attribution, minimizing hallucination, and making deliberate trade-offs between accuracy, latency, and cost. Systems must be production-ready. - Defining success metrics with domain experts and building datasets, benchmarks, and evaluation frameworks that capture the nuanced medical certification domain. Measuring retrieval quality, citation correctness, expert agreement, and impact on assessor effort, assessment quality, and customer experience. - Applying AI alignment principles to build agents that respect the impartiality and objectivity required of a certification body, helping users understand evidence, recognize uncertainty, and retain responsibility for consequential decisions. You will own meaningful, ambiguous problems: working directly with domain experts to define success, deciding what to build and test, and taking ML systems through deployment to measurable production improvement. REQUIREMENTS: - 3+ years shipping software to production, with ownership of deployment, monitoring, and incident investigation. - Deep understanding of deep learning, agent tools, context management, and information retrieval, and how they interact. - Experience building datasets and establishing evaluation frameworks from first principles; skepticism of benchmarks that don't match reality. - Experience designing, deploying, and operating production systems; ability to choose infrastructure that fits the workload and explain trade-offs in complexity, reliability, security, and cost. - Pragmatic security judgment for production systems: reasoning about data access, permissions, untrusted inputs, and actions with real consequences; designing proportionate mitigations. - Ability to own ambiguous problems, working with domain experts to define success and following work through experiments to measurable improvement. - Insatiable curiosity about real-world problems and commitment to building solutions with clear economic and human value. PREFERRED QUALIFICATIONS: - Experience deploying agent systems with tool use, sensitive data, or consequential actions, including evaluating behavior and designing safeguards. - Experience building retrieval or document-understanding systems whose outputs must be validated against complex source evidence.

Similar roles