SlipstreamJobsFresh Startup & VC-Backed Jobs

GPU Kernel Engineer

Sciforium - San Francisco, CA, USA - In-office - posted 2026-08-28

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Sciforium is an AI infrastructure company developing next-generation multimodal AI models and a proprietary high-efficiency serving platform. Backed by multi-million-dollar funding and direct sponsorship from AMD, the company is scaling rapidly to build the full stack powering frontier AI models and real-time applications. You will design and optimize custom GPU kernels that power next-generation large-scale AI systems, working across the hardware–software stack from low-level kernel development to integrating optimized operations into high-level ML frameworks used for large-scale training and inference. Key responsibilities include: - Design, implement, and optimize custom GPU kernels using C++, PTX, CUDA, ROCm, Triton, and/or JAX Pallas - Profile and optimize end-to-end performance of ML operations, focusing on large-scale LLM training and inference - Integrate low-level GPU kernels into frameworks such as PyTorch, JAX, and custom internal runtimes - Develop performance models, identify bottlenecks, and deliver kernel-level improvements that accelerate AI workloads - Collaborate with ML researchers, distributed systems engineers, and model-serving teams to optimize compute performance across the stack - Work closely with hardware vendors (NVIDIA/AMD) and stay current on latest GPU architecture capabilities - Contribute to tooling, documentation, benchmarking suites, and testing frameworks Required qualifications: - 5+ years of industry or research experience in GPU kernel development or high-performance computing - Bachelor's, Master's, or PhD in Computer Science, Computer Engineering, Electrical Engineering, Applied Mathematics, or related field - Strong programming skills in C++, Python, and familiarity with ML frameworks - Deep expertise in CUDA/ROCm, GPU memory models, and performance optimization strategies - Hands-on experience with Triton and/or JAX Pallas for custom kernel development - Strong understanding of PTX, GPU ASM, and low-level GPU execution - Extensive experience writing and optimizing custom GPU kernels in C++ and PTX - Proven ability to integrate low-level kernels into PyTorch, JAX, or similar frameworks - Experience working with large-scale LLM workloads (training or inference) Nice-to-haves include experience with AMD GPUs and ROCm optimization, JAX FFI, efficient model serving frameworks (vLLM, TensorRT), TPUs/XLA, and open-source ML systems contributions.

Similar roles