SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Niural is an AI-native platform unifying payroll, compliance, HR, and financial operations across 150+ countries. The company is backed by Marathon, M13, and Inspired Capital.
You will join the team building EMMA, Niural's AI orchestration system. This is a production-focused role shipping agent workflows, retrieval pipelines, and document understanding systems that operate on payroll, tax, and compliance data. Unlike research or bolted-on chatbot roles, the systems you build take real actions on money movement and statutory filings where accuracy is non-negotiable.
Key responsibilities include designing and shipping production LLM features (agent workflows, tool calling, MCP servers, RAG pipelines); building retrieval systems over statutory guidance, tax publications, contracts, and handbooks with ownership of chunking, embedding selection, hybrid search, and reranking; building structured extraction pipelines for payroll and compliance documents with schema enforcement and validation; defining evaluation infrastructure with golden datasets, regression suites, and LLM-as-judge calibration; optimizing latency and cost through model routing, caching, and context budgeting; implementing safety controls (PII detection, prompt injection defense, grounded citations, confidence thresholds); instrumenting and monitoring AI systems in production; collaborating with payroll, tax, and compliance experts to translate regulatory requirements; and staying current with the fast-moving AI field with a bias toward measurable improvement.
Required qualifications: 3+ years professional software engineering with recent production LLM experience; strong Python and service-oriented codebases (APIs, queues, workers, observability); hands-on RAG experience (embeddings, vector stores, hybrid search, reranking); hands-on agent architecture experience (tool calling, multi-step planning, state/memory, error recovery); structured output experience (JSON schema, constrained decoding, validation); LLM evaluation and regression testing design; prompt and context engineering knowledge; LLM observability tooling experience; sound judgment on model vs. deterministic code in regulated domains; clear written communication.
Nice-to-have skills include AWS Bedrock, MCP server design, agent frameworks (LangGraph, LlamaIndex, Pydantic AI), model fine-tuning/distillation, vision-language models/OCR, synthetic data generation, knowledge graphs, React/TypeScript, fintech/payroll/tax/HR domain experience, and open-source contributions.