SlipstreamJobsFresh Startup & VC-Backed Jobs

Member of Technical Staff - Low Level & Kernels Capabilities

Preference Model - San Francisco, CA, USA - In-office - posted 2026-08-06

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Preference Model is building automated ML research engineering, focusing on reinforcement learning (RL) environments that enable frontier models to handle real-world complexity. The company was founded by engineers from Anthropic's data team who built infrastructure and datasets for Claude, and is now partnering with leading AI labs. You will join the Low Level / Kernels Capabilities team, which builds RL environments at the lowest layers of the stack—GPU and accelerator kernels, vector ISAs, codec and crypto primitives, FPGA work, and more. These are domains where frontier models are weakest and underrepresented in training data. Your responsibilities include: - Design and build low-level, kernel-focused RL environments targeting specified models and difficulty distributions - Evaluate which environments are worth building based on domain niche, hardware features, research motivation, and reference benchmarks (cuBLAS, FFTW, OpenSSL, etc.) - Build deterministic, ungameable correctness and performance scoring that prevents models from exploiting loopholes - Own environments end-to-end: domain selection, task design, scoring infrastructure, and hardening against reward hacking - Ensure robust sandboxing so models must actually solve the kernel problem rather than gaming timers or other shortcuts Required qualifications: - Strong low-level/systems engineering: fluent in C/C++/CUDA or equivalent kernel languages; comfortable with assembly - Production-grade Python across automation, deployment, data analysis, and plotting (not notebook-only) - Hardware-aware coding with deep understanding of memory hierarchy, occupancy, data movement, parallelism, and latency vs. throughput tradeoffs - Proven kernel development and optimization experience using profilers - Adversarial mindset: ability to turn fuzzy goals into robust, ungameable scoring systems - Hands-on experience with LLMs - Ownership and autonomy to build, debug, and ship end-to-end with minimal supervision Desirable experience includes: shipping kernels approaching SOTA, depth in niche hardware (FPGA/HLS, RISC-V Vector, DSPs, SIMD/AVX, TPUs), HPC/heterogeneous clusters, hardware design (RTL/HDL), compilers (MLIR/LLVM, Mojo, Triton, gem5), formal verification, reading and implementing performance papers, open-source contributions, competitive programming in low-level languages, or prior RL environment and evaluation infrastructure work. Compensation is competitive (>90th percentile cash and equity). The role offers ownership in a fast-moving startup, collaboration with top ML engineers, standard benefits (health, vision, dental, 401K match), onsite lunch daily, weekly snacks, and visa sponsorship/relocation support.

Similar roles