SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Handshake AI is a rapidly scaling data business supporting frontier AI labs with human-generated training data. As an AI Model Policy Trainer on the Violence & Fiction team, you will evaluate AI model responses to user requests involving violence, weapons, threats, and dark fiction. Your core responsibility is distinguishing legitimate depictions of violence (in novels, games, screenplays, history, journalism, self-defense contexts) from requests seeking real-world harm capability or expressing genuine intent to harm.
You will read user requests, model responses, and full conversation history to classify each case against policy standards and assess whether the model's response was appropriate. The work focuses on edge cases where context is everything: a torture scene that could be either thriller fiction or an interrogation manual; a frustrated message about a boss versus an actual plan; a novelist's combat question that also conveys real-world capability interest. You will explain your reasoning clearly enough to train models and inform policy refinement.
Key responsibilities include: evaluating requests within full conversation context; distinguishing fictional, educational, historical, and defensive violence from real-world harm-seeking; assessing whether a model response grants meaningful real-world capability regardless of framing; distinguishing anger and dark humor from credible threats; selecting defensible classifications in genuinely ambiguous cases with cited rationales; writing adversarial prompts to probe model boundaries; identifying policy gaps and edge cases; participating actively in team calibration discussions; applying customer policy consistently; and maintaining accuracy across repeated, feedback-heavy evaluations.
This is not rote annotation. You will balance policy text and intent with customer expectations, conversation context, precedent, and team calibration. Strong judgment, clear reasoning, and the ability to hold opinions without attachment to being right are essential. You should have deep expertise in at least one domain: fiction writing and dark storytelling, real-world violence and its consequences (military, law enforcement, security, emergency medicine), or crisis intervention and threat assessment (crisis lines, counseling, domestic violence advocacy, school safety, trust and safety).