SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Innovaccer is seeking a Principal AI Researcher to lead the design, development, and deployment of production-grade AI systems that solve complex problems at scale. You will work across product, engineering, and data teams to build intelligent applications including LLM-based solutions, AI agents, retrieval-augmented generation (RAG), and intelligent automation workflows.
In this role, you will own the full AI development lifecycle—from experimentation and prototyping through production deployment and optimization. You'll take ideas from concept to prototype to production, building the model layer of real products while maintaining product-level accuracy standards. You'll write production Python and PyTorch code, design rigorous experiments with clear hypotheses and ablations, and build evaluation frameworks before models. You'll measure results skeptically and explain technical work to non-AI stakeholders including clinicians and operators.
Key responsibilities include: architecting complete model systems spanning small, medium, and large models; making critical decisions on data mixture, objective design, and training run management; leading large-scale distributed training efforts across hundreds of GPUs; building and owning data pipelines including deduplication, filtering, decontamination, and normalization; mentoring and setting technical direction for other engineers; and defining what in-house modeling means at Innovaccer.
Required qualifications: MS or PhD in Computer Science, Machine Learning, or related quantitative field (exceptional BS candidates with substantial research or open-source work considered). You must have first-author publications at top venues (NeurIPS, ICML, ICLR, ACL, EMNLP) or equivalent research depth. Hands-on experience fine-tuning open-weight models, understanding parameter-efficient vs. full fine-tuning approaches. Strong Python and PyTorch proficiency with familiarity with modern training/serving stacks (HuggingFace, FSDP, DeepSpeed, vLLM, SGLang). Multi-GPU and multi-node training experience at scale (tens of GPUs minimum, models in tens of billions of parameters or larger). Proven ability to ship models in production and understand failure modes. Experience leading post-training on models in hundreds of billions of parameters across hundreds of GPUs. Deep familiarity with large-scale distributed training failure modes and operational discipline. A public record of papers, model releases, systems, or training programs demonstrating impact. Ability to set two-to-four-quarter technical direction and lead team execution.
Innovaccer offers competitive benefits including 20 days PTO annually, industry-leading parental leave, rewards and recognition programs, comprehensive insurance (medical, dental, vision, disability, life), and optional legal aid and pet insurance.