SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Pluralis Research is building Protocol Learning: a system for training and serving large language models in a fully decentralized way across consumer-grade devices connected via the internet. The company recently achieved a major milestone with Agora, a permissionless run that pretrained an 8B model from scratch on consumer GPUs distributed globally, with no single participant holding full weights.
As a Machine Learning Engineer on the ML Training Platform team, you will architect, build, and scale the infrastructure that enables continuous experimentation and large-scale training on top of decentralized, non-colocated consumer nodes and cloud instances. This is a systems-heavy role focused on platform reliability and orchestration rather than model research.
Key responsibilities include:
- Multi-cloud infrastructure: Design resource management systems that provision and orchestrate compute across AWS, GCP, and Azure using infrastructure-as-code (Pulumi/Terraform). Handle dynamic scaling, state synchronization, and concurrent operations across hundreds of heterogeneous nodes.
- Distributed training and inference systems: Architect fault-tolerant infrastructure for distributed ML, including GPU cluster management, NVIDIA runtime optimization, S3 checkpointing, large-dataset management and streaming, health monitoring, and resilient retry strategies.
- Real-world networking: Build systems that simulate and handle real network conditions including bandwidth shaping, latency injection, and packet loss. Manage node churn and keep data flowing across workers with heterogeneous connectivity.
Required qualifications: Production experience with infrastructure-as-code (Pulumi/Terraform/CloudFormation) managing multi-cloud deployments, Docker/Kubernetes (EKS), GPU workloads, and heterogeneous clusters at scale. Deep understanding of distributed training workflows including checkpointing, data sharding, model versioning, and long-running job orchestration. Strong Python engineering with asyncio, concurrency, retry logic, cloud SDKs, and CLI tooling. Hands-on observability and SRE practice with Prometheus/Grafana, performance profiling, and incident response. Experience with P2P networking, NAT traversal, traffic shaping, and real bandwidth constraints. Proven track record owning systems in a startup with heavy service orchestration or at big-tech scale.
Nice-to-have: Experience with foundation model pre-training, post-training, or RL. Background at proprietary, open-weight, or open-source AI labs.
The company is backed by Union Square Ventures and other tier-1 investors. Offers equity-heavy compensation package with significant ownership for key technical contributors, remote-first culture with global distribution, optional visa sponsorship and relocation support to Australia or the US, and the opportunity to work on frontier problems with no published solutions.