SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Rubrik is hiring a Senior Machine Learning Engineer to join the SAGE (Semantic AI Governance Engine) team, building real-time monitoring, governance, and remediation systems for autonomous AI agents. SAGE uses small language models as judges to enforce governance policies on every agent action in production, operating at enterprise scale with sub-second latency.
In this role, you will own the full machine learning lifecycle across four key areas:
**Model Development & Training (25%):** Lead training, fine-tuning, and distillation of production small language models and classifiers. Own base-model selection, supervised fine-tuning, preference optimization (DPO/RLAIF), and distillation from frontier models. Design adversarial training pipelines with automated red-teams that feed directly into the next training cycle. Optimize the accuracy-latency-cost Pareto frontier using techniques like LoRA, quantization-aware training, and GRPO, validated against production traffic patterns.
**Model Serving & Inference Infrastructure (25%):** Design multi-stage inference pipelines handling both real-time enforcement (inline prompt/response/tool-call blocking) and high-throughput batch workloads processing billions of tokens daily. Optimize deployments through shared GPU pools, KV-cache routing, continuous batching, FP8/INT8 quantization, and speculative decoding to maintain sub-second P99 SLOs. Build model gateway infrastructure and request routing that doesn't become a latency bottleneck. Own canary, shadow, and A/B traffic validation for new model variants.
**Data & Evaluation Systems (20%):** Build automated data curation pipelines mining live customer environments (with privacy guarantees) for high-value training examples. Implement policy back-testing by replaying historical agent traffic against new model versions to catch regressions. Create online evaluation systems including shadow scoring, drift detection, calibration monitoring, and policy-coverage gap analysis. Generate and validate synthetic data using frontier teachers.
**Insights & Adaptive Improvement (15%):** Mine failure patterns and diagnose model weaknesses. Build context harnesses fusing data sensitivity, identity, and historical behavior into real-time decisions. Drive continuous model improvement based on production signals and customer feedback.
The team ships models to enterprise customers within weeks and is passionate about proving small, specialized models outperform frontier LLMs on AI safety and governance problems. You'll work end-to-end on models that enforce policies in the live request path and power Agent Rewind, Rubrik's capability to instantly undo destructive autonomous-agent actions.