SlipstreamJobsFresh Startup & VC-Backed Jobs

ML Engineering Lead (LLM Ops)

Neko Health - Stockholm, Sweden - In-office - posted 2026-10-02

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Neko Health is building a preventative healthcare platform that combines proprietary sensor technology with clinical care to deliver personalized health insights in a single non-invasive visit. The company operates within a regulated medical device quality management system and serves multiple clinical domains including dermatology and cardiology. You will own the operational lifecycle of Neko's LLM, GenAI, and RAG-based systems, building the production-grade platform that enables clinical ML and GenAI workflows to run reliably on proprietary sensor and device data. This role sits within the ML Engineering area and works closely with the existing MLOps team as an integrated extension rather than a siloed function. In your first 6-12 months, you will: - Implement MLflow Tracing observability across production LLM and agent pipelines, tracking prompts, tool calls, retrievals, latency, and cost - Build a comprehensive evaluation suite combining LLM judges, custom scorers, and human-feedback loops via review apps to replace ad hoc review processes - Ship at least one RAG or agentic pipeline to production with prompt and application versioning through MLflow Prompt Registry and Unity Catalog, enabling safe rollout, A/B testing, and rollback - Develop a documented framework for cost, latency, and GPU capacity trade-offs when choosing serving strategies: third-party APIs vs. Databricks External Models vs. self-hosted solutions - Integrate LLM Ops tightly with the existing MLOps team to ensure it operates as a natural platform extension while protecting reliability for clinical workflows You will work on systems that directly impact patient outcomes, operating within healthcare compliance requirements and across complex technical stacks spanning medical devices, firmware, and hardware. REQUIREMENTS: - Solid MLOps fundamentals across the full lifecycle (experiment tracking, training, monitoring) demonstrated through independent ownership of complex, production-grade work - Fluent in Python with deep understanding of core ML concepts and a track record of shipping end-to-end production ML systems and platformization initiatives - Practical, hands-on experience building LLM or GenAI applications: prompt engineering, RAG, agents or chains using frameworks such as LangChain, LangGraph, or comparable orchestration tools - Working knowledge of PyTorch, distributed systems, and ML orchestration - Conceptual understanding of LLM-specific MLOps trade-offs: fine-tuning vs. prompting vs. RAG, vector databases, embedding models, and human-feedback loops for non-deterministic outputs - Genuine, demonstrable motivation for LLM, GenAI, and RAG work specifically (not generic MLOps), and ability to navigate complex systems spanning medical domain, regulation, firmware, and hardware PREFERRED QUALIFICATIONS: - Experience with agentic or AI-assisted coding workflows while retaining full ownership and understanding of output - Kubernetes and Terraform for infrastructure as code or self-hosting fine-tuned or open-source models outside managed serving - Exposure to LLM evaluation and observability practices: tracing, LLM-as-judge, guardrails, and safety scorers; Databricks MLflow 3 for GenAI experience is a strong plus - Comfort navigating a fast-moving tools and platform ecosystem and distilling recommendations relevant to Neko's specific context

Similar roles