SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Simpplr is seeking a Voice AI Engineer to design and build production-grade voice agents for frontline-heavy verticals including healthcare, manufacturing, warehousing, retail, and hospitality. The role focuses on employee support, procurement, collections, logistics, and ordering use cases.
You will lead the technical design of low-latency, real-time voice systems that combine automatic speech recognition (ASR), text-to-speech (TTS), large language models (LLMs), conversational AI, enterprise workflows, knowledge retrieval, compliance controls, and human handoff capabilities. This is a hands-on technical leadership position requiring someone who can architect voice AI from design through production deployment.
Key responsibilities include: designing and building real-time voice runtime for live conversations; optimizing streaming ASR, TTS, voice activity detection (VAD), endpointing, turn-taking, and barge-in functionality; building adaptive voice pipelines for high-noise frontline environments (60-112 dB) including server-side noise cancellation and echo suppression; architecting multi-provider speech routing across multilingual matrices with code-switching support (e.g., Spanglish, Hinglish); developing voice agents supporting multi-turn, multi-intent conversations with context switching and recovery; integrating with workflows, APIs, CRM, ITSM, knowledge bases, and enterprise systems; implementing secure identity verification, consent, privacy, audit, and compliance controls; building warm transfer and seamless human handoff with full conversation context; optimizing multilingual voice quality across accents and latency; establishing production observability across the full call path; evaluating and integrating leading speech, telephony, and AI technologies; defining architecture and engineering standards; and mentoring engineers on critical technical design reviews.
Minimum qualifications: 7+ years software engineering experience; strong background building production distributed or real-time systems; hands-on experience with conversational AI, voice AI, speech AI, or LLM-based agents; strong programming skills in Python, Java, Go, or equivalent; experience with APIs, streaming systems, asynchronous architectures, and cloud-native platforms; strong understanding of system design, scalability, reliability, and observability.
Preferred qualifications include hands-on experience with Deepgram (real-time ASR), LiveKit (WebRTC and voice-agent runtime), ElevenLabs (low-latency TTS), OpenAI/Azure/Google Speech technologies, WebRTC/SIP/RTP/WebSockets/Twilio, LLM agents and orchestration frameworks, enterprise system integration (Salesforce, ServiceNow, Jira, Zendesk, Workday), multilingual speech handling, and high-scale multi-tenant SaaS platform experience.