Machine Learning Engineer - ML Training Platform
pluralis-research
USA or Australia
Posted Aug 31, 2026
- Full-time
- Remote
- Engineering
Job description
Pluralis Research works on Protocol Learning: training and serving large models in a fully decentralized way on small consumer-grade devices connected via the internet. Despite being dismissed as infeasible, we have made significant advances on this problem, most recently Agora, a permissionless run that pretrained an 8B model from scratch on consumer GPUs spread over the internet, with no single participant ever holding the full weights ([tech report](https://arxiv.org/abs/2607.13332)). While many of the core research problems have been solved, Protocol Learning unlocks a series of new challenges. For the mission in full, read [A Third Path: Protocol Learning](https://pluralis.ai/blog/a-third-path-protocol-learning/). Our training and inference doesn't happen in a datacenter. It happens on consumer nodes and cloud instances that are not co-located, connected by ordinary internet, joining and leaving mid-run. Your primary role is to architect, build, and scale the platform that keeps continuous experimentation and large-scale training running on top of that: infrastructure orchestration, distributed compute, and the services that tie them together. ## Key Responsibilities - **Multi-cloud infrastructure**: Design the resource management systems that provision and orchestrate compute across AWS, GCP, and Azure with infrastructure-as-code (Pulumi/Terraform). Handle dynamic scaling, state synchronization, and concurrent operations across hundreds of heterogeneous nodes. - **Distributed training and inference systems**: Architect fault-tolerant infrastructure for distributed ML. GPU clusters, NVIDIA runtime, S3 checkpointing, large-dataset management and streaming, health monitoring, and resilient retry strategies. - **Real-world networking**: Build the systems that simulate and handle real network conditions such as bandwidth shaping, latency injection, packet loss. Managing node churn and keeping data flowing across workers with heterogeneous connectivity. ## What We're Looking For - **Infrastructure and platform engineering (required)**: Production experience with infrastructure-as-code (Pulumi/Terraform/CloudFormation) managing multi-cloud deployments, Docker/Kubernetes (EKS), GPU workloads, and heterogeneous clusters at scale. - **Distributed systems and ML infrastructure**: You understand distributed training workflows: checkpointing, data sharding, model versioning, long-running job orchestration. - **Decentralized networking**: P2P, NAT traversal, traffic shaping, real bandwidth constraints. - **Systems programming and reliability**: Strong Python engineering (asyncio, concurrency, retry logic, cloud SDKs, CLI tooling) with hands-on observability and SRE practice; Prometheus/Grafana, performance profiling, incident response. - **Environment fit**: You've done this in a startup with heavy service orchestration, or at big-tech scale, and you can show which systems you owned. - **Mission alignment**: You believe Protocol Learning is the viable third path for collective, trustless, and sovereign AI. ## Nice to Have - Experience with foundation model pre-training, post-training, or RL. - Experience at proprietary, open-weight and open-source AI labs ## Compensation & Benefits - **Equity-Heavy Package**: We offer significant ownership for key technical contributors in addition to a high base salary. - **Remote-First Culture**: Flexible work environment with team members distributed globally. - **Visa Sponsorship**: Optional full visa sponsorship and relocation support to either Australia or the US. - **Open Problems**: Training and serving frontier models on hardware you don't control, over networks you don't own, mostly has no published answers yet. You'll write some of the first ones. ## FYI's - We work remotely across the world, with the main teams in Australia and North America. You'll need to be comfortable working across timezones. - Applicants must have professional-level English proficiency (written and spoken). - Recruiters: we aren't looking for agency support at this time. We'll reach out if we need help. *We are backed by* [*Union Square Ventures*](https://www.usv.com/) *and other tier-1 investors, and we are a world-class, deeply technical team of ML researchers. Pluralis is unapologetically ideological. We believe AI, and the world, end up on a better path if we succeed in implementing the protocol for intelligence. If this resonates, please apply.*