SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Sciforium is an AI infrastructure company developing next-generation multimodal AI models and a proprietary high-efficiency serving platform. Backed by multi-million-dollar funding and direct sponsorship from AMD, the company is scaling rapidly to build the full stack powering frontier AI models and real-time applications.
You will design and optimize custom GPU kernels that power next-generation large-scale AI systems, working across the hardware–software stack from low-level kernel development to integrating optimized operations into high-level ML frameworks used for large-scale training and inference.
Key responsibilities include:
- Design, implement, and optimize custom GPU kernels using C++, PTX, CUDA, ROCm, Triton, and/or JAX Pallas
- Profile and optimize end-to-end performance of ML operations, focusing on large-scale LLM training and inference
- Integrate low-level GPU kernels into frameworks such as PyTorch, JAX, and custom internal runtimes
- Develop performance models, identify bottlenecks, and deliver kernel-level improvements that accelerate AI workloads
- Collaborate with ML researchers, distributed systems engineers, and model-serving teams to optimize compute performance across the stack
- Work closely with hardware vendors (NVIDIA/AMD) and stay current on latest GPU architecture capabilities
- Contribute to tooling, documentation, benchmarking suites, and testing frameworks
Required qualifications:
- 5+ years of industry or research experience in GPU kernel development or high-performance computing
- Bachelor's, Master's, or PhD in Computer Science, Computer Engineering, Electrical Engineering, Applied Mathematics, or related field
- Strong programming skills in C++, Python, and familiarity with ML frameworks
- Deep expertise in CUDA/ROCm, GPU memory models, and performance optimization strategies
- Hands-on experience with Triton and/or JAX Pallas for custom kernel development
- Strong understanding of PTX, GPU ASM, and low-level GPU execution
- Extensive experience writing and optimizing custom GPU kernels in C++ and PTX
- Proven ability to integrate low-level kernels into PyTorch, JAX, or similar frameworks
- Experience working with large-scale LLM workloads (training or inference)
Nice-to-haves include experience with AMD GPUs and ROCm optimization, JAX FFI, efficient model serving frameworks (vLLM, TensorRT), TPUs/XLA, and open-source ML systems contributions.