SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Machine Learning Engineer, Voice Agents - EMEA Remote

Hugging Face - Remote - Remote - posted 2026-09-02

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Hugging Face is seeking a Senior Machine Learning Engineer to own a substantial portion of their open voice-agent stack. The role centers on two key initiatives: speech-to-speech, an open-source library for realtime voice agents already powering the Reachy Mini fleet, and hf-voice, a new product enabling developers to build and deploy voice agents via their Hugging Face account. You will take architectural ownership of the speech-to-speech library, including pipeline design, latency optimization, and realtime loop reliability. You'll integrate new ASR, TTS, and end-to-end speech models as they emerge, maintain clean abstractions, review community contributions, and grow the contributor ecosystem. On the product side, you'll design the developer API and streaming protocol (session lifecycle, WebSockets/WebRTC, authentication, error semantics, versioning), build the serving infrastructure (realtime GPU inference, concurrency, autoscaling, observability, cost optimization), and collaborate with Hub and inference teams to ensure seamless integration. You'll drive the product from demo to production with load testing, SLOs, and graceful degradation strategies. You'll also work in the open: writing documentation, examples, and templates to help developers launch agents quickly, supporting existing deployments like Reachy Mini, and sharing work publicly through blog posts, demos, and conference talks (travel covered). Required: Senior-level ownership of substantial architecture; experience building developer-facing infrastructure at AI or developer-tools companies (inference APIs, agent infrastructure); substantial open-source Python library contributions with async and distributed systems expertise; shipped realtime systems (streaming, WebSockets, WebRTC, audio/video pipelines, live inference); production LLM/multimodal experience; clear async communication and public collaboration habits; motivation for voice and conversational AI. Bonus: contributions to voice-agent frameworks (speech-to-speech, pipecat, LiveKit Agents, Vocode, TEN); llama.cpp or inference runtime work; hands-on ASR/TTS/end-to-end speech model experience; GPU serving, quantization, or on-device inference; audio pipeline knowledge (VAD, echo cancellation, jitter buffers, barge-in, turn detection); embedded/robotics shipping; public track record.

About Hugging Face

AI / Data / Infrastructure — open-source AI model hub, libraries, datasets, and enterprise tooling.

Similar roles