SlipstreamJobsFresh Startup & VC-Backed Jobs

Research Engineer, Privacy and Anonymization

hud - San Francisco, CA, USA - Hybrid - posted 2026-09-16

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

HUD is building infrastructure for RL training data and evaluations for frontier AI agents, with a marketplace connecting data providers to frontier labs. The company has raised $16M from top VCs and is a YC W25 company. You will own the design and implementation of privacy and anonymization systems that enable sensitive, real-world data to be safely used for AI training. This is a full-stack role spanning detection, transformation, pipeline development, and evaluation. Key responsibilities: - Build systems to detect PII, quasi-identifiers, credentials, and sensitive information, designing transformations appropriate to data type and downstream use case - Develop and benchmark detection approaches combining rules, statistical models, classifiers, and LLM-based methods - Build production pipelines that anonymize raw data before downstream processing, training, evaluation, or synthetic data generation - Create evaluation frameworks measuring privacy risk and retained data utility, including recall-weighted metrics, leakage tests, and adversarial re-identification attempts - Design systems robust to new data sources, schema drift, unusual formats, and sensitive information in unexpected fields - Collaborate with engineering, research, operations, and customers to translate privacy requirements into technical policies and safeguards The team is ~25 people, mostly full-time in-person but with some remote flexibility. The team includes 4 International Olympiad medalists, serial AI startup founders, and researchers with publications at top venues (ICLR, NeurIPS). The company is scaling profitably with strong demand. Location: Offices in San Francisco or Singapore, with openness to remote candidates who can overlap 70-80% with either timezone. Visa sponsorship and relocation support provided for strong candidates. Requirements: - Strong Python proficiency and experience building reliable production data or ML systems - Experience with information extraction, named-entity recognition, classification, or related methods for detecting rare or sensitive content - Strong experimental instincts and ability to compare approaches across recall, precision, latency, cost, and downstream data utility - Understanding of redaction, masking, pseudonymization, anonymization, and synthetic data—and when each is appropriate - High attention to detail and ability to reason about subtle leakage paths, edge cases, and adversarial failure modes - Experience building data processing pipelines end-to-end without fully prescribed roadmaps Strong additional qualifications: - Hands-on experience with privacy-enhancing technologies (differential privacy, k-anonymity, secure aggregation, format-preserving encryption) - Experience with sensitive data in healthcare, finance, or security - Built low-latency or high-throughput ML inference and data-processing systems - Worked in unstructured problem spaces from early research through production deployment - Early-stage startup experience and strong cross-team communication skills The company prioritizes technical aptitude and learning potential over years of experience.

Similar roles