SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Transfyr is building physical AI systems for science—infrastructure that captures, analyzes, and operationalizes the tacit knowledge embedded in real-world scientific work. The company records multimodal data from laboratory environments to make experimental execution legible, reproducible, and automatable.
As a Systems ML Engineer, you will own performance optimization across Transfyr's full ML stack, from large-scale training through production inference. You'll work embedded with research and perception teams to make training faster and more efficient, and ensure models run well at inference across both cloud and edge infrastructure deployed in active lab environments.
Key responsibilities include:
- Profile and optimize training and inference workloads using tools like Nsight and PyTorch Profiler, identifying bottlenecks in data loading, gradient computation, and communication.
- Implement optimizations such as kernel fusion, sharding, and tiling to improve step time and distributed training efficiency.
- Design and maintain high-performance GPU kernels in Triton or CUDA for performance-critical ML workloads.
- Build and optimize data loading pipelines for training and inference pipelines that reliably serve multimodal lab data (vision, audio, sensor, metadata).
- Manage deployment across cloud infrastructure and edge devices in resource-constrained lab environments with limited compute and connectivity.
- Debug and resolve performance bottlenecks, resource issues, and failures across the training and deployment stack.
- Partner with research and perception teams on training efficiency and data flow reliability.
- Build monitoring, versioning, and rollback mechanisms to ensure safe model updates in production.
You combine deep understanding of ML systems with hands-on performance engineering skill. You have experience profiling and optimizing large-scale models for both training and inference, are comfortable writing custom GPU kernels when off-the-shelf operations aren't fast enough, and can manage cloud and edge infrastructure in real-world environments. You're high-agency, biased toward action, and thrive in ambiguity. You communicate clearly about model performance and limitations across engineering, science, and operations teams.