SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Pluralis Research is building Protocol Learning: a system for training and serving large language models in a fully decentralized way on consumer-grade devices connected via the internet. The company recently achieved a major milestone with Agora, a permissionless run that pretrained an 8B model from scratch on consumer GPUs distributed across the internet, with no single participant ever holding the full weights.
As a Research Engineer focused on Post-Training, you will own the end-to-end post-training stack for decentralized model training. This is a unique challenge: standard RL post-training assumes datacenter infrastructure with synchronous rollouts, fast interconnects, and trusted workers. Your stack must operate on consumer GPUs and Macs spread across the public internet, training models whose weights no participant fully holds, with rollouts arriving from a geo-distributed inference pipeline at high latencies.
Key responsibilities include:
- Building the complete RL training loop: rollout ingestion from the geo-distributed inference pipeline, reward computation, policy updates, and weight distribution back to the network
- Inventing novel algorithms adapted to asynchronous, high-latency, partially-trusted generation: staleness tolerance, off-policy corrections, and communication-efficient policy updates
- Shipping the first post-trained decentralized models, including building evaluation frameworks and taking releases from experimental run to public artifact
You should have hands-on experience running RL post-training on large language models (RLHF, RLVR, or reasoning-focused RL), with direct experience in the systems layer: rollout generation, async training loops, and weight synchronization. Strong production-quality Python and PyTorch skills are essential, including concurrency, failure handling, and profiling. Research ability demonstrated through publications or detailed unpublished work in RL post-training, asynchronous or distributed RL is valued. Mission alignment with Protocol Learning as a viable path for collective, trustless, and sovereign AI is important.
Nice-to-have skills include experience training over slow networks, familiarity with serving-engine internals (vLLM, SGLang), reward modeling, P2P networking and NAT traversal, and prior work at proprietary, open-weight, or open-source AI labs.
The company is backed by Union Square Ventures and other tier-1 investors, with a world-class team of ML researchers. The role is remote-first with teams distributed across Australia and North America. Visa sponsorship and relocation support to either Australia or the US are available. Professional-level English proficiency (written and spoken) is required.