SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Reflection AI is a research lab building open foundational models to make intelligence accessible to everyone. As a Research Software Engineer on the Safety team, you will design and own the infrastructure for running sensitive model evaluations in high-consequence domains including CBRN (chemical, biological, radiological, nuclear), child safety, and other dangerous-capability areas. These evaluations directly inform release decisions for open models, so the systems you build must be secure, isolated, reproducible, and trustworthy under external scrutiny.
You will work at the intersection of platform engineering, security, and safety research. Your responsibilities include designing secure, sandboxed execution environments for model evaluation; building controlled data pipelines and storage for sensitive evaluation material with least-privilege access controls, RBAC, encryption, audit logging, and data-minimization safeguards; and partnering with safety researchers and domain experts to translate evaluation designs into reliable, reproducible, scalable systems.
Key technical deliverables include eval-orchestration tooling and harnesses for high-throughput evaluations in isolated environments; infrastructure for measuring AI capability uplift in high-consequence domains; guardrails, monitoring, and compartmentalization to keep sensitive work appropriately siloed; and production-quality Python systems for data processing and evaluation pipelines.
You should have strong software engineering skills in Python with a track record building reliable, scalable infrastructure or platform systems. Experience with sandboxed or security-sensitive execution environments (containerization, VM isolation, secure compute) is essential. You need solid grounding in security fundamentals: least privilege, need-to-know access, RBAC, secrets management, encryption, audit logging, compartmentalization, and defense-in-depth design. Experience building data pipelines with sensitive or restricted data is required. You should be comfortable owning entire problems end-to-end, including ambiguous cross-functional ones, and thrive in a fast-paced, high-agency startup environment.
Bonus qualifications include experience with evaluation, benchmarking, or experimentation infrastructure for ML systems; familiarity with LLMs, agents, or ML training/inference pipelines; knowledge of dangerous-capability or dual-use domains and their information-security considerations; and familiarity with compliance frameworks for sensitive data handling.