SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Transfyr is building physical AI systems for science—capturing real scientific work and converting it into high-fidelity, machine-readable records of experimental execution. The company addresses a critical gap: scientists lack the detailed feedback loops that athletes have, making it difficult to automate physical lab work, train new scientists, or transfer knowledge across labs. Transfyr's platform helps teams learn from failures, transfer expertise, and provide grounded data for AI and robotics. Backed by a $25M seed round with advisors including Chris Ré (Stanford), David Baker (Nobel laureate), and Kevin Weil (former CPO at OpenAI).
As a Systems ML Engineer, you will own performance optimization across the full ML lifecycle—from large-scale training through production inference. You'll work embedded with the research team to accelerate training and with perception/production systems to optimize inference across cloud and edge infrastructure in real lab environments.
Key responsibilities include: profiling and optimizing training and inference workloads using tools like Nsight and PyTorch Profiler; improving distributed training efficiency with frameworks like PyTorch Distributed; designing and maintaining high-performance GPU kernels in Triton or CUDA; building and optimizing data loading and inference pipelines for multimodal lab data (vision, audio, sensors); managing deployment across cloud and edge devices with limited compute and connectivity; debugging performance bottlenecks and resource issues; partnering with research and perception teams; and building monitoring, versioning, and rollback mechanisms for safe model updates.
You combine deep understanding of ML systems with hands-on performance engineering expertise. Ideal candidates have experience profiling and optimizing large-scale models for both training and inference, are comfortable writing custom GPU kernels, and can manage cloud and edge infrastructure in real-world lab settings. The role spans deep ML, performance engineering, and cloud/DevOps responsibilities. This is an in-person role in Cambridge, MA. The company is building a team across levels and welcomes both hands-on builders and senior engineers who enjoy shaping training infrastructure and technical direction.