SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Sciforium is an AI infrastructure company building next-generation multimodal AI models and a proprietary high-efficiency serving platform. Backed by multi-million-dollar funding and direct AMD sponsorship, the company is scaling rapidly to develop the full stack powering frontier AI models and real-time applications.
As a GPU Kernel Engineer, you will design and optimize custom GPU kernels that power large-scale AI systems. You'll work across the hardware–software stack, from low-level kernel development to integrating optimized operations into high-level ML frameworks used for training and inference at scale.
Key responsibilities include:
- Design, implement, and optimize custom GPU kernels using C++, PTX, CUDA, ROCm, Triton, and/or JAX Pallas
- Profile and optimize end-to-end performance of ML operations, focusing on large-scale LLM training and inference
- Integrate low-level GPU kernels into frameworks such as PyTorch, JAX, and custom internal runtimes
- Develop performance models, identify bottlenecks, and deliver kernel-level improvements that accelerate AI workloads
- Collaborate with ML researchers, distributed systems engineers, and model-serving teams to optimize compute performance across the stack
- Work closely with hardware vendors (NVIDIA/AMD) and stay current on latest GPU architecture capabilities
- Contribute to tooling, documentation, benchmarking suites, and testing frameworks
Required qualifications:
- 5+ years of industry or research experience in GPU kernel development or high-performance computing
- Bachelor's, Master's, or PhD in Computer Science, Computer Engineering, Electrical Engineering, Applied Mathematics, or related field
- Strong programming skills in C++, Python, and familiarity with ML frameworks
- Deep expertise in CUDA/ROCm, GPU memory models, and performance optimization strategies
- Hands-on experience with Triton and/or JAX Pallas for custom kernel development
- Strong understanding of PTX, GPU ASM, and low-level GPU execution
- Extensive experience writing and optimizing custom GPU kernels in C++ and PTX
- Proven ability to integrate low-level kernels into PyTorch, JAX, or similar frameworks
- Experience working with large-scale LLM workloads (training or inference)
Nice-to-haves include AMD GPU and ROCm optimization experience, JAX FFI familiarity, efficient model serving frameworks (vLLM, TensorRT), TPU/XLA experience, and open-source contributions to ML systems, compilers, or GPU kernels.