SlipstreamJobsFresh Startup & VC-Backed Jobs

Systems ML Engineer (Member of the Technical Staff)

Transfyr - Cambridge, MA, United States - In-office

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Transfyr is building physical AI systems for science, creating infrastructure to capture, analyze, and operationalize the tacit knowledge embedded in real-world scientific work. The company records multimodal data from laboratory environments and transforms it into durable, actionable knowledge—building the world's largest commercial dataset on scientific execution. As a Systems ML Engineer (Member of the Technical Staff), you will own performance optimization across Transfyr's full ML stack, from large-scale training through production inference. This role combines deep ML systems expertise with hands-on performance engineering and cloud/edge infrastructure management. Key responsibilities include: - Profile and optimize training and inference workloads using tools like Nsight and PyTorch Profiler, identifying bottlenecks in data loading, gradient computation, and communication - Implement optimizations including kernel fusion, sharding, and tiling to improve step time - Optimize distributed training pipelines using PyTorch Distributed, working closely with the research team - Design and maintain high-performance GPU kernels in Triton or CUDA for performance-critical workloads - Build and optimize data loading pipelines for training and inference, handling multimodal lab data (vision, audio, sensors, metadata) - Manage deployment across cloud infrastructure and edge devices operating in active lab environments with constrained compute and connectivity - Debug and resolve performance bottlenecks, resource issues, and failures across the training and deployment stack - Partner with research and perception teams to improve training efficiency and data flow reliability - Build monitoring, versioning, and rollback mechanisms to ensure safe model updates in production - Optimize cloud spend and ensure security and compliance You are a high-agency engineer who identifies what needs to be learned, builds the right scaffolding, and pushes work forward. You prototype quickly, test assumptions against real data, and iterate based on failure. You thrive in ambiguity, making deliberate tradeoffs between model complexity, robustness, and operational cost. You communicate clearly about model performance and limitations across engineering, science, and operations teams, and care deeply about the mission. Required expertise: demonstrated proficiency in ML systems engineering, including optimizing and deploying large-scale models in production, debugging performance and stability issues, building infrastructure for reproducible and monitored ML deployments, and optimizing inference throughput and resource utilization across cloud and edge environments. Strong background in distributed training and serving frameworks.

Similar roles