SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Handshake AI is a rapidly growing data company supporting frontier AI labs with human-generated training data. As an AI Model Policy Trainer on the Violence & Fiction team, you will evaluate AI model responses to user requests involving violence, weapons, threats, and dark fiction content. Your core responsibility is distinguishing legitimate depictions of violence (in novels, games, screenplays, history, journalism, self-defense contexts) from requests seeking real-world harm capability or expressing genuine intent to harm.
You will read user requests, model responses, and full conversation history to classify each case against policy standards and assess whether the model's response was appropriate. The work focuses on nuanced edge cases: a torture scene that could be either thriller fiction or an interrogation manual; a frustrated message about a boss versus an actual plan; a combat question from a novelist versus a real-world capability question. Success requires understanding how context, intent, and a single word can change the answer entirely.
Key responsibilities include: evaluating requests and responses within full conversation context; distinguishing fictional, educational, historical, and defensive violence from real-world harm-seeking; assessing whether model responses provide meaningful real-world capability regardless of framing; distinguishing anger/frustration/dark humor from credible threats; selecting defensible classifications for genuinely ambiguous cases with clear rationales; writing adversarial prompts to probe model boundaries; identifying policy gaps and edge cases; participating in calibration discussions; applying customer policies consistently; and maintaining accuracy during repeated, feedback-heavy evaluations.
This is not rote annotation work. You will balance policy text and intent with customer expectations, conversation context, precedent, and team calibration. Strong candidates bring deep expertise in at least one domain: fiction writing/game design/screenwriting (understanding story-shaped extraction attempts), real-world violence exposure (military, law enforcement, security, EMS, emergency medicine), or crisis work (crisis lines, counseling, threat assessment, domestic violence advocacy, school safety). The role values clear reasoning, the ability to hold strong opinions without attachment to being right, and the capacity to separate personal views from applied standards.