Reinforcement Learning Engineer - Manipulation
humanoid
UK, London
Posted Jul 16, 2026
- Full-time
- Engineering
Job description
Here at Humanoid, we believe in a future where robots amplify human potential. That’s why we’ve set out on a mission to build the world’s most capable, commercially-scalable, and safe humanoid robots. We’re bringing that mission to life with HMND‑01 Alpha - our rapidly developed humanoid platform now running in real industrial pilots - and we’re growing the team to take it even further. # About the Role We're hiring a **Reinforcement Learning Engineer** to join our Autonomy team based in London. In this role you will leverage reinforcement learning in both simulation and physical reality to build highly performant and robust manipulation policies. # What You'll Do - Train language-vision conditioned manipulation policies via reinforcement learning (RL) in simulation and in the real world. - Construct challenging and diverse suites of manipulation tasks in simulation. - Partner with teleoperations to collect trajectories in simulation for behavior cloning. - Partner with testing and operations to establish real-world RL training pipelines. - Experiment with various ways of bringing policies trained in simulation to the real world. # What We're Looking For - 3+ years building deep‑learning systems (industry or research) with shipped models or published artifacts to show for it. - Hands‑on with at least one of: LLMs, VLMs, or image/video generative models — architecture, training, and inference. - Experience solving real problems using reinforcement learning with deep neural networks in any domain. - Strong Python + PyTorch/JAX; you can profile, debug numerics, and write maintainable research code. - You are self-driven, pro-active, communicate efficiently, document experiments clearly and communicate trade‑offs crisply. # Nice to have - Experience with simulators for robotics (Isaac Sim, MuJoCo etc.) - Experience in RL for robotics. - Experience building infrastructure for large-scale RL (e.g. using ray). - Publications at ICLR/ICML/NeurIPS or equivalent open‑source contributions. - Familiarity with OpenVLA, Physical Intelligence (π) models, or similar open VLA frameworks. # What We Offer - Competitive equity: stock options with meaningful upside as we scale. - 30+ paid days off, including 23 days of annual leave, all UK bank holidays, and additional company closure days (including Christmas–New Year shutdown). - Private healthcare, including virtual and in-person care. - Pension scheme with 8% total contribution (5% employee, 3% employer) on full earnings. - Free daily breakfast, catered lunch, and snacks in-office. - Work at the frontier - collaborate daily with world-class engineers, researchers, and product experts building the next generation of AI and humanoid robotics. - Real ownership - direct access to founding leadership, meaningful input on product direction, and the ability to drive key initiatives from day one.