SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Pluralis Research is building Protocol Learning: a system for training and serving large language models in a fully decentralized way across consumer-grade devices connected via the internet. The company has already demonstrated feasibility with Agora, a permissionless run that pretrained an 8B model from scratch on consumer GPUs distributed globally, with no single participant holding full weights.
As a Research Engineer on the pre-training team, you will build the training infrastructure that scales Protocol Learning from 8B to frontier-scale models. This role bridges cutting-edge research and production systems engineering, tackling challenges that break traditional datacenter assumptions: communication-efficient training across heterogeneous hardware and networks, fault tolerance as nodes join and drop mid-run, and robustness to malicious participants.
Key responsibilities include:
- Implementing and optimizing distributed model-parallel training (data, pipeline, and tensor parallelism) for large models on heterogeneous GPUs under low-bandwidth, high-latency internet links
- Developing performance optimization techniques that reduce communication overhead while maintaining model convergence
- Building elasticity and fault tolerance mechanisms: robust checkpointing, state synchronization, and recovery as participants join and leave
- Creating run instrumentation and monitoring systems that track throughput, bottlenecks, and model quality across hundreds of devices
You should have hands-on experience training models across many devices in PyTorch using FSDP, DeepSpeed, Megatron, or custom implementations. Strong production-quality Python engineering, understanding of concurrency and failure handling, and evidence of shipped systems or serious projects are essential. Nice-to-have skills include experience with large language model training (Nemotron, Qwen, OLMo), P2P networking and NAT traversal, post-training and RL, inference systems, and work at proprietary or open-source AI labs.
The company is backed by Union Square Ventures and operates a remote-first culture with teams distributed across Australia and North America. Visa sponsorship and relocation support are available. The role offers significant equity alongside competitive base salary, and you'll be working on largely unsolved problems at the frontier of distributed AI training.