SlipstreamJobsFresh Startup & VC-Backed Jobs

Quality Systems Lead

Encord - London, United Kingdom - In-office - posted 2026-08-28

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Encord is hiring a Quality Systems Lead to own data quality measurement and assurance across its AI data platform. The role combines technical engineering with operational management of a distributed QA function. You will build automated dataset quality evaluation systems using model-assisted screening, LLM-as-judge approaches, agreement analysis at scale, and anomaly/drift detection across annotation output. Simultaneously, you will hire, train, calibrate, and manage a dedicated audit team of approximately ten specialists based in India. This team produces labelled ground truth that trains and validates the automated layer, with performance measured on inter-rater agreement and catch rate rather than volume. Key responsibilities include: establishing quality standards for each data type Encord delivers (written rubrics, golden sets, acceptance criteria); building scoring systems that rank annotator performance and inform routing and staffing decisions; setting certification pass thresholds; reporting quality KPIs to leadership and customers (accuracy, inter-annotator agreement, rework rate, cost of rework, coverage); managing remediation workflows with Project Management; owning unit economics of quality; and partnering with Product and Engineering to embed quality measurement into the Encord platform as native capability. You must balance engineering and operations: writing evaluation pipelines while running weekly calibration sessions with a distributed team. Your instinct on coverage problems is to automate rather than scale headcount. You are statistically literate in practical terms (sampling design, agreement statistics, metric validation), can hold calibrated standards across distributed teams, and are a strong writer capable of producing rubrics for a workforce you don't sit with. You hold quality standards under commercial pressure and bring evidence to make them stick with delivery teams and clients. Required: 4+ years owning both technical and operational outcomes in data, AI, or service delivery environments where quality was measured; hands-on Python and SQL; practical experience applying models to quality/evaluation problems (LLM-as-judge, model-assisted QA, automated evaluation, anomaly detection); sampling methodology and agreement statistics applied to production data; experience hiring, training, and managing teams (ideally audit/QA, ideally distributed); track record building quality frameworks with leadership and customer reporting. Bonuses: direct annotation/evaluation/model-training workflow experience; STEM degree or data science/research engineering background; multilingual delivery and linguistic quality assessment.

Similar roles