SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Spectrum Labs (operating as Alice) is seeking a Senior AI Researcher to lead post-training evaluation, red-teaming, and reinforcement learning (RL) gym audits on open-weight models. You will establish rigorous benchmarking methodologies and evaluate large language models against complex security threats, particularly Indirect Prompt Injections (IPI).
Key responsibilities include executing post-training runs using RL algorithms like GRPO on open-weight generalist models against security-focused environments; analyzing loss curves and rollout traces to identify reward hacking, lazy policy convergence, and verifier flaws; auditing tasks and multi-turn environments (tool use, web navigation, computer use) for realism and threat model accuracy; generating comprehensive evaluation cards detailing performance uplift, failure modes, and task-level success rates; and integrating dockerized environments into internal training frameworks while optimizing concurrency and throughput.
Required qualifications include an M.S. or Ph.D. in Data Science, Machine Learning, Computer Science, or equivalent deep learning experience. You must have strong hands-on experience training large-scale models using RL algorithms on open-weight architectures, solid understanding of LLM vulnerabilities and red-teaming methodologies, proficiency in PyTorch and Docker containerization, and ability to analyze agent rollout traces and debug complex reward shaping issues.
Preferred experience includes work with standard RL gym formats (Harbor), evaluation of complex agentic workflows in tool-use or web-browser environments, and familiarity with evaluating open-weight models like Llama or Mistral against adversarial workloads.
Alice is a trust, safety, and security company providing end-to-end AI lifecycle coverage, from model hardening and pre-deployment red-teaming to runtime guardrails and drift detection. The company supports frontier model labs, enterprises, and UGC platforms, and is considered a global leader in online safety and AI security.