SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
PagerDuty is seeking a Senior AI/ML Engineer to design and build AI systems that run in production at scale. The role sits at the intersection of large-scale distributed systems and applied AI, focusing on powering Incident Management AI Agents, event intelligence, and LLM-powered capabilities across the platform.
You will own the full lifecycle of AI features—from problem framing through production deployment and monitoring. Key responsibilities include designing AI-powered features like LLM agents and retrieval systems that operate on high-volume, real-time event streams; architecting systems for agent and prompt orchestration, retrieval pipelines, tool/API integrations, and low-latency inference at scale; reasoning about consistency, throughput, fault tolerance, and cost under bursty, unpredictable load; and establishing evaluation, guardrail, observability, and improvement loops to keep AI features accurate and trustworthy over time.
You'll partner with platform, product, and applied-research teams to define success criteria and integrate AI cleanly into existing services. The role includes raising the bar through technical example, code reviews, and mentorship.
Required qualifications: 5+ years of software engineering with meaningful production experience in distributed systems (high-throughput services, streaming/event-driven architectures, or large-scale data platforms); hands-on experience building and shipping AI systems in production (LLM applications, agents, or retrieval systems, not just prototypes); strong programming fundamentals and ability to move between systems and AI/application code; solid grounding in applied AI fundamentals (prompting, retrieval, agent patterns, evaluation, and guardrailing); experience with cloud infrastructure (AWS, GCP, or Azure), containers, and Kubernetes; pragmatic, reliability-minded approach to systems design; and strong communication and collaboration skills.
Nice-to-have skills include LLMOps tooling experience, deep production experience with retrieval-augmented generation and multi-step agents, background in anomaly detection or observability, familiarity with LLM frameworks (LangChain, LlamaIndex), vector databases, and distributed tools (Kafka, Airflow, Spark), and open-source contributions.