SlipstreamJobsFresh Startup & VC-Backed Jobs

Member of the Technical Staff - Systems ML Engineer

Transfyr - Cambridge, MA, United States - In-office

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Transfyr is building physical AI systems for science, creating high-fidelity, machine-readable records of scientific execution to help teams learn from failures, transfer knowledge, and train the next generation of scientists. The company has raised a $25M seed round and is backed by advisors including Chris Ré (Stanford), David Baker (Nobel laureate), and Kevin Weil (former CPO at OpenAI). As a Systems ML Engineer, you will own performance optimization across Transfyr's ML stack, from large-scale training through production inference. You'll work embedded with the research team to make training faster and more efficient, and with perception and production systems to optimize model performance across cloud and edge infrastructure in real lab environments. Key responsibilities include: - Profile and optimize training and inference workloads using tools like Nsight and PyTorch Profiler, identifying bottlenecks in data loading, gradient computation, and communication - Implement optimizations such as kernel fusion, sharding, and tiling to improve step time - Improve distributed training efficiency using frameworks like PyTorch Distributed - Design and maintain high-performance GPU kernels in Triton or CUDA for performance-critical workloads - Build and optimize data loading pipelines for training and inference, handling multimodal lab data (vision, audio, sensor, metadata) - Manage deployment across cloud infrastructure and edge devices in active lab environments with limited compute and connectivity - Debug and resolve performance bottlenecks, resource issues, and failures across the training and deployment stack - Partner with research and perception teams on training efficiency and data pipeline reliability - Build monitoring, versioning, and rollback mechanisms into deployments to ensure safe model updates You combine deep understanding of ML systems with hands-on performance engineering skill. You have experience profiling and optimizing large-scale models for both training and inference, are comfortable writing custom GPU kernels when off-the-shelf operations aren't fast enough, and can manage cloud and edge infrastructure that holds up in real lab environments. You're high-agency, biased toward action, and successful in ambiguity—you identify what needs to be learned, build the right scaffolding, and push work forward without waiting for perfect datasets or well-posed problems.

Similar roles