SlipstreamJobsFresh Startup & VC-Backed Jobs

Software Engineer, GPU Performance

HeyGen - Los Angeles, CA, USA - Hybrid - posted 2026-09-28

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

HeyGen is building AI applications including Avatar IV, Photo Avatar, Interactive Avatar, and Video Translation. We're seeking a Software Engineer focused on GPU performance to optimize the systems powering these experiences, making them faster and more cost-efficient. In this role, you will work across model execution and inference infrastructure, using profiling and measurement to improve latency, throughput, and GPU cost. You'll investigate GPU utilization, kernel execution, memory bandwidth, and CPU–GPU data movement using tools like NVIDIA Nsight Systems, Nsight Compute, and PyTorch Profiler. You'll identify bottlenecks across model execution, preprocessing, and inference serving, then measure the impact of optimizations. Key responsibilities include improving performance through batching, scheduling, memory management, and better GPU utilization; developing or integrating high-performance GPU kernels when existing implementations limit performance; building benchmarks and automated checks to catch performance regressions across representative video workloads; collaborating with AI researchers and infrastructure engineers to bring optimizations into production; and measuring the effect of changes on latency, throughput, cost, and output quality. This role is ideal for an engineer who enjoys understanding how software uses the hardware beneath it and can turn performance experiments into reliable production changes while communicating tradeoffs clearly. QUALIFICATIONS: - Experience optimizing GPU-based AI workloads or high-performance computing systems - Proficiency in Python and experience with PyTorch or similar machine learning framework - Strong curiosity about GPU hardware, including memory bandwidth, cache behavior, tensor cores, and CPU–GPU data movement - Experience using profiling tools to connect hardware behavior to application-level bottlenecks and validate improvements - Ability to turn performance experiments into reliable production changes and communicate tradeoffs clearly PREFERRED QUALIFICATIONS: - Experience with CUDA, Triton, or C++ GPU programming - Experience optimizing video, image, audio, diffusion, or Transformer models - Familiarity with multi-GPU inference, GPU interconnects, quantization, or large-scale model serving - Experience building performance benchmarks or regression testing infrastructure - Prior experience in a fast-paced technology environment

Similar roles