During the 2026 World Robot Conference (WRC) on August 24, ZhiXiang Future unveiled its native full-modal interactive world model, HiDream-O1-World, marking its first-ever evaluation and immediate rise to the top of the Navi leaderboard on the authoritative WBench benchmark.
HiDream-O1-World supports multi-modal inputs including text, images, and interactive control, featuring three core capabilities: roaming, editing, and interaction. It enables the construction of diverse subjects and stylized scenes, accurately recreating real-world spaces while also generating anime-style or 3A game-rendered worlds. Its key breakthrough lies in long-term spatiotemporal and physical consistency, allowing the model to "remember" previously explored scene structures during extended interactions and maintain high physical plausibility in complex scenarios.
Mei Tao, founder and CEO of ZhiXiang Future, noted that AI large models have advanced rapidly in recent years, with three technological paths—large language models, multi-modal models, and embodied intelligence—now accelerating their convergence. He observed that 2026 is widely dubbed the "year of world models."
Citing data, he highlighted that top-tier large models scored an IQ of roughly 130 early in the year, approaching 140 now, but he stressed that "high IQ doesn't mean all-around capability"—a model's excellent test performance doesn't guarantee it can reliably execute long-term tasks in the real physical world.
Discussing the capability foundation of embodied intelligence, Mei Tao summarized the critical relationship: "The development of world models and embodied intelligence requires connecting data, world models, intelligent agents, and the embodied entity itself." Data builds cognition, world models interpret the state of the environment and deduce physical laws, agents make decisions and plan based on predictions, and the embodied entity executes real-world actions—with real environmental feedback looping back as new training data, forming a closed-loop flywheel of "data-cognition-action-feedback."
Mei Tao emphasized that the world model serves as the cognitive core of this flywheel; without high-quality world models, simulations can't approach real physics, and the data flywheel of embodied intelligence cannot truly spin.