SlipstreamJobsFresh Startup & VC-Backed Jobs

GPU Kernel Engineer

Sciforium - San Francisco, CA, USA - In-office - posted 2026-08-26

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Sciforium is an AI infrastructure company building next-generation multimodal AI models and a proprietary high-efficiency serving platform. Backed by multi-million-dollar funding and direct AMD sponsorship, the company is scaling rapidly to develop the full stack powering frontier AI models and real-time applications. As a GPU Kernel Engineer, you will design and optimize custom GPU kernels that power large-scale AI systems. You'll work across the hardware–software stack, from low-level kernel development to integrating optimized operations into high-level ML frameworks used for training and inference at scale. Key responsibilities include: - Design, implement, and optimize custom GPU kernels using C++, PTX, CUDA, ROCm, Triton, and/or JAX Pallas - Profile and optimize end-to-end performance of ML operations, focusing on large-scale LLM training and inference - Integrate low-level GPU kernels into frameworks such as PyTorch, JAX, and custom internal runtimes - Develop performance models, identify bottlenecks, and deliver kernel-level improvements that accelerate AI workloads - Collaborate with ML researchers, distributed systems engineers, and model-serving teams to optimize compute performance across the stack - Work closely with hardware vendors (NVIDIA/AMD) and stay current on latest GPU architecture capabilities - Contribute to tooling, documentation, benchmarking suites, and testing frameworks Required qualifications: - 5+ years of industry or research experience in GPU kernel development or high-performance computing - Bachelor's, Master's, or PhD in Computer Science, Computer Engineering, Electrical Engineering, Applied Mathematics, or related field - Strong programming skills in C++, Python, and familiarity with ML frameworks - Deep expertise in CUDA/ROCm, GPU memory models, and performance optimization strategies - Hands-on experience with Triton and/or JAX Pallas for custom kernel development - Strong understanding of PTX, GPU ASM, and low-level GPU execution - Extensive experience writing and optimizing custom GPU kernels in C++ and PTX - Proven ability to integrate low-level kernels into PyTorch, JAX, or similar frameworks - Experience working with large-scale LLM workloads (training or inference) Nice-to-haves include AMD GPU and ROCm optimization experience, JAX FFI familiarity, efficient model serving frameworks (vLLM, TensorRT), TPU/XLA experience, and open-source contributions to ML systems, compilers, or GPU kernels.

Similar roles