Daxiao Robotics, in collaboration with Nanyang Technological University's S-Lab and other institutions, has unveiled and open-sourced the unified multimodal world model framework Puffin-World. This announcement was made on the afternoon of September 10.
As a breakthrough achievement targeting spatial intelligence and embodied intelligence, Puffin-World marks the first time that physics, geometry, and appearance have been defined as three native three-dimensional world states within a single unified multimodal framework. The system integrates gravity field perception, free-viewpoint spatial simulation, 3D world generation and reconstruction, along with closed-loop exploration, enabling it to predict the state of the world following a robot's actions.
In conjunction with the launch, the Puffin-World framework's code, models, and the Puffin-16M dataset have all been made open source. The dataset comprises 15 million vision-language-camera triplets and 1 million challenging rotational trajectories. Additionally, absolute camera pose annotations covering approximately 44.5 million images sourced from 28 public datasets have been made available, providing a reproducible and scalable technical foundation for robot simulation, synthetic data production, and embodied intelligence training.