SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior/Staff AI Engineer, Quality & Evals (m/f/x)

Cortea - Berlin, Germany - In-office - posted 2026-09-09

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Cortea is a Berlin-based AI startup transforming audits with intelligent software and specialized AI agents. The company has secured over €15M in funding from top-tier VCs, has a working product, and is scaling with paying customers. You will build the evaluation and observability foundation for production-grade LLM agents used in complex audit workflows. This role sits at the intersection of backend engineering, data infrastructure, and AI quality systems. You will design and implement evaluation systems that power multimodal retrieval agents and continuously improve critical quality metrics across current and future pipelines. Key responsibilities include: - Building online and offline evaluation systems for LLM agents, including pipelines using golden datasets, ground-truth data, human review workflows, and experiment results - Creating automated quality gates to test changes to prompts, context, models, or agent logic before production deployment - Analyzing large volumes of agent traces and executions in analytical databases (BigQuery, ClickHouse) to identify failure modes, quality regressions, latency issues, reliability gaps, and cost optimization opportunities - Building reliable data retention and replay mechanisms for long-term analysis of production agent behavior - Managing observability tools for tracing, monitoring, debugging, and experiment management of audit agents - Partnering with backend engineers to improve speed and reliability of retrieval and reasoning agents This is not a traditional analytics, BI, or dashboarding role. You will write production code, design data architecture, work inside backend systems, and directly improve quality, cost, reliability, and performance of LLM-based agents. The company values first-principles thinking, speed, trust, and kindness, with a collaborative Berlin office environment. REQUIREMENTS: - Strong Python and/or backend engineering experience - Solid understanding of LLM and agent system evaluation methods (deterministic checks, ground truth, LLM-as-judge, human review, quality metrics) and judgment about when each approach is appropriate - Deployed and operated systems in the cloud, ideally on GCP - Hands-on experience building end-to-end retrieval or ML pipeline evaluation systems - Experience with LLM observability or experimentation tools (Braintrust, MLflow, Langfuse, Weights & Biases) - Comfortable working with analytical databases, data warehouses, columnar stores, and high-volume event/trace data - Understanding of system design, reliability, observability, monitoring, logging, debugging, and operational trade-offs - Senior-level engineering judgment: ability to make architectural decisions, communicate trade-offs, and build extensible systems - Comfort with ambiguity and first-principles reasoning NICE-TO-HAVES: - Experience designing data pipelines, ETL/ELT workflows, event-processing systems, or feedback loops - Infrastructure work around LLM-based products or agentic systems (optimizing LLM usage, context windows, reasoning tokens, model selection) - Working with production traces from complex distributed systems - Building internal platforms for engineers, domain experts, or operations teams - Workflow orchestration systems (Temporal or similar) - Familiarity with audit, finance, compliance, or high-accuracy domains - Early-stage startup or fast-moving engineering environment experience

Similar roles