SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Sciforium is an AI infrastructure company developing next-generation multimodal AI models and a proprietary high-efficiency serving platform. The company is backed by multi-million-dollar funding and direct sponsorship from AMD, with hands-on support from AMD engineers.
As a Distributed Training and Inference Engineer, you will build, optimize, and maintain the critical software stack powering large-scale AI training and serving workloads. You will work across the entire ML infrastructure from low-level CUDA/ROCm runtimes to high-level frameworks like JAX and PyTorch.
Key responsibilities include:
- Maintaining and optimizing critical ML libraries and frameworks (JAX, PyTorch, CUDA, ROCm) across multiple environments and hardware configurations
- Building and continuously improving the entire ML software stack from ROCm/CUDA drivers to high-level JAX/PyTorch tooling
- Ensuring all model implementations are efficiently sharded, partitioned, and configured for large-scale distributed training and serving
- Integrating and validating modules for runtime correctness, memory efficiency, and scalability across multi-node GPU/accelerator clusters
- Conducting detailed profiling of compilation graphs, training workloads, and runtime execution to identify and eliminate bottlenecks
- Troubleshooting complex hardware-software interaction issues including vLLM compilation failures, CUDA memory leaks, distributed runtime failures, and kernel-level inconsistencies
- Collaborating with research, infrastructure, and kernel engineering teams to improve system throughput, stability, and developer experience
Ideal candidates have 5+ years of industry experience in ML systems or distributed training, a Bachelor's or Master's degree in Computer Science or related field, strong programming skills in Python and C++, deep understanding of profiling tools, expertise with modern ML frameworks, and hands-on experience with distributed training systems and CUDA/ROCm technologies. Nice-to-have qualifications include extensive XLA/JAX experience, familiarity with distributed serving frameworks, GPU kernel optimization background, and deep understanding of low-level C++ components in ML frameworks.