SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Archetype AI is building the world's first physical AI platform to bring artificial intelligence into the real world. The company's foundation model, Newton, understands the physical world through objective sensor data and generates real-time insights into complex physical behaviors, from industrial machinery and systems to wearable devices and smart environments. Founded by a high-caliber team from Google and backed by a renowned Silicon Valley venture fund, Archetype AI is in Series A and rapidly advancing its technology.
In this role, you will own the serving path for Newton and related multimodal models. Much of the inference stack is Rust-native, featuring model nodes in the agent runtime built on Rust ML stacks (candle, Burn) with custom GPU kernels, plus the routing layer that streams real-time inference to GPU nodes. You will drive GPU utilization, numerical precision, and low-latency serving from the kernel up.
Key responsibilities include:
- Build and own model nodes in the Rust inference runtime: loading, warmup, batching, streaming, and GPU memory pools
- Optimize kernels and the GPU path: custom CUDA kernels, mixed precision, quantization, and parity against research
- Own inference routing and serving: streaming path API to GPU node, request batching, SLO-backed latency and cost optimization
- Productionize research checkpoints: export, compilation, quantization, parity evals, and rollout
- Build observability for inference needs: latency histograms, GPU metrics, OOM signatures, and replayable traces
You will work closely with researchers, translating architecture changes into serving work and ensuring production systems meet performance and reliability requirements.
REQUIREMENTS:
- 6+ years of software engineering experience, with several years in ML systems, inference, or high-performance GPU computing
- Proven track record owning a production serving path end-to-end (not just benchmarking models)
- Expert-level PyTorch proficiency with models shipped to production; ability to identify performance bottlenecks before profiling
- Strong CUDA or equivalent GPU depth: memory hierarchy, occupancy, Nsight or equivalent profiling tools
- Proficiency in Rust or C++ alongside Python; solid Linux performance and production ops experience; ready to work in Rust daily
- Ability to work effectively with researchers and translate technical requirements into implementation
NICE TO HAVE:
- Experience with Rust ML stacks (candle, Burn) or comparable GPU compute frameworks in Rust
- Custom kernels and compiler stacks: Triton, CUTLASS, TorchInductor, TensorRT
- Production quantization and mixed precision work with numerical-correctness validation
- Multimodal, video, embedding, or time-series serving experience (beyond decoder-only chat LLMs)
- High-performance serving stacks: vLLM, SGLang, TensorRT-LLM with continuous batching and paged KV cache