Deep Learning Engineer - World Models
humanoid
UK, London
Posted Jul 16, 2026
- Full-time
- Engineering
Job description
Here at Humanoid, we believe in a future where robots amplify human potential. That’s why we’ve set out on a mission to build the world’s most capable, commercially-scalable, and safe humanoid robots. We’re bringing that mission to life with HMND‑01 Alpha - our rapidly developed humanoid platform now running in real industrial pilots - and we’re growing the team to take it even further. # Our Mission At Humanoid we strive to create the world's leading, commercially scalable, safe, and advanced humanoid robots that seamlessly integrate into daily life and amplify human capacity. # About the Role As a Research Engineer on the World Models team, you will build action-conditioned generative models that predict how the world evolves around our robots — future video, proprioception, contacts, and outcomes — from past observations and actions. World models serve four purposes in our stack: a pretrained, physics-aware prior for our VLA policies; an engine for rare data collection, cross-platform transfer, and sim-to-real transfer; a testbed for policy evaluation and testing before hardware; and a future-prediction rollout engine that surfaces what our policies intend to do, for safety and planning. This is a hands-on individual contributor role: you will design architectures, run large training jobs, and validate your models against real fleet data from industrial deployments. # What You'll Do - Design and train multimodal world models — video, state, action, and language — using diffusion-based and transformer architectures. - Build action-conditioned video prediction and dynamics models that stay physically consistent over long horizons, including contact-rich manipulation, and serve as pretrained priors for VLA policies. - Develop learned-simulator evaluation: score candidate policies offline, predict real-world success rates before deployment, and roll out policy futures to expose intended behaviour for safety review and planning. - Generate synthetic rollouts and counterfactual experience — including rare events, cross-platform transfer, and sim-to-real transfer — to augment policy training, and measure their effect on downstream task performance. - Establish fidelity metrics and calibration protocols that quantify where the world model can be trusted and where it diverges from reality. - Build data pipelines that turn fleet telemetry, teleoperation logs, and internet-scale video into training corpora for world models. - Run scaling and ablation studies on architecture, data mixture, and context length; communicate findings crisply. - Collaborate with pretraining, RL, and manipulation teams to integrate world models into policy training and evaluation loops. # What We're Looking For - A track record of training large generative models — video, world, or multimodal — with shipped models or published artifacts to show for it. - Deep hands-on experience with modern generative architectures: diffusion models, autoregressive transformers, latent-variable models, or video prediction. - Experience with large-scale distributed training: streaming datasets, checkpointing and state management, debugging numerics and training instabilities. - Strong Python + PyTorch/JAX; you can profile kernels, optimize data loaders, and write maintainable research code. - Empirical rigor: you design careful evaluations, run honest baselines, and document experiments clearly. - Excitement about grounding generative models in physical reality rather than pixels alone. ## Nice to have - Experience with world models for robotics or autonomous driving (e.g., action-conditioned video models, learned simulators, model-based RL). - Familiarity with robotics simulators (Isaac Sim, MuJoCo) and sim-to-real considerations. - Experience using world models for policy evaluation or synthetic data generation at scale. - Publications at top-tier deep learning conferences (NeurIPS, ICML, ICLR, CoRL, CVPR) or equivalent open-source contributions. - Experience optimizing generative models for fast inference. # What We Offer - Competitive equity: stock options with meaningful upside as we scale. - 30+ paid days off, including 23 days of annual leave, all UK bank holidays, and additional company closure days (including Christmas–New Year shutdown). - Private healthcare, including virtual and in-person care. - Pension scheme with 8% total contribution (5% employee, 3% employer) on full earnings. - Free daily breakfast, catered lunch, and snacks in-office. - Work at the frontier - collaborate daily with world-class engineers, researchers, and product experts building the next generation of AI and humanoid robotics. - Real ownership - direct access to founding leadership, meaningful input on product direction, and the ability to drive key initiatives from day one.