SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Applied Intuition is seeking a Machine Learning Performance Engineer to optimize large-scale distributed training and inference workloads. The role focuses on making ML systems fast and cost-efficient in the datacenter, with emphasis on throughput, cluster goodput, and cost per unit of data processed.
Key responsibilities include:
- Profile and optimize distributed training end-to-end: data loading, preprocessing, augmentation, kernel execution, gradient communication, and checkpointing
- Optimize large-scale offline and batch inference over petabyte-scale sensor logs using batching strategies, quantization, graph optimization, and accelerator saturation techniques
- Establish roofline and performance models for workloads, quantify performance gaps, and prioritize optimization opportunities
- Improve multi-node scaling efficiency through sharding, parallelism strategies, collective communication, and memory-bandwidth optimization
- Drive cluster goodput by reducing GPU idle time from input pipeline stalls, storage/network I/O, scheduling gaps, and failure recovery
- Build benchmarking, observability, and regression-detection tooling to prevent performance degradation
- Collaborate across engineering teams to solve complex data and compute problems at scale
Required qualifications:
- Hands-on ML performance engineering experience with profiling, roofline analysis, and throughput optimization in production
- Experience with distributed multi-node training at scale (FSDP, DeepSpeed, Megatron, NCCL)
- Deep familiarity with GPU/accelerator performance concepts: memory bandwidth, kernel launch overhead, occupancy, quantization, collective communication
- Experience with high-throughput or batch inference systems (NVIDIA Triton, TensorRT, ONNX Runtime, Ray)
- Fluency in Python and proficiency in C++ or systems language
- Excellent debugging, analytical, and problem-solving skills
- Deep understanding of machine learning foundations
Nice-to-have skills include GPU kernel development (CUDA, Triton), profiling toolchains (Nsight, PyTorch Profiler), GPU scheduling/orchestration (Kubernetes, Slurm, Ray), fault tolerance and elastic training experience, and familiarity with autonomy/robotics data.
Applied Intuition is a $15B Series A company founded in 2017, powering physical AI for automotive, defense, trucking, construction, mining, and agriculture. The company has 18 of the top 20 global automakers and the US military as customers. Headquarters in Sunnyvale with global offices.