SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Arago is re-engineering computing from first principles by fusing optical and CMOS technologies to deliver order-of-magnitude performance gains for AI workloads. The company has built the fastest custom processor of its kind and is backed by leading deep-tech investors and industry luminaries including the CEO of Arm, the founder of macOS, an Nvidia Fellow, and Google's Head of Optics.
As an ML Systems Engineer focused on Inference Acceleration, you will optimize the execution and serving of modern AI models on Arago's custom accelerator. Your work spans kernels, model execution, multi-device distribution, runtime, and inference serving, directly shaping the software stack around Arago's hardware capabilities.
Key responsibilities include:
- Analyzing modern AI workloads to identify kernel-, runtime-, memory-, and system-level bottlenecks on the accelerator
- Developing and optimizing custom kernels, fused operators, and execution strategies to maximize device utilization
- Designing efficient mappings of models and operators across multiple Arago devices with optimized communication and synchronization
- Implementing inference-serving techniques such as continuous batching, paged KV caches, prefix/context caching, and prefill/decode scheduling strategies
- Building profiling, benchmarking, and performance-analysis infrastructure across kernels, full models, and serving workloads
- Collaborating closely with hardware, compiler, and runtime teams to co-design software abstractions and influence future hardware features
Required expertise includes strong experience in high-performance ML inference, GPU/accelerator programming, or ML systems engineering. You need deep understanding of computer architecture, accelerator execution models, memory hierarchies, and performance optimization. Hands-on experience with custom kernel development using CUDA, Triton, ROCm/HIP, or equivalent is essential. You should have practical experience with modern inference-serving systems like vLLM, SGLang, or TensorRT-LLM, including KV-cache management and attention mechanisms. Strong C++ and Python skills are required, with MLIR exposure a plus. Proficiency in English is necessary.
Arago operates with three core values: do great things, move with high velocity, and operate as one unit. The environment demands constant learning, ownership, and execution excellence, offering exceptional people the opportunity to do their life's work in deep-tech AI infrastructure.