Building a Native 4D World Model: The Path to General Physical Intelligence

Deep News
昨天

At the China International Fair for Trade in Services (CIFTIS) special event on Physical AI and Embodied Intelligent Robot Innovation Ecosystem, held on September 11, 2026, in Beijing, YingShen Intelligence (影身智能) founder and CEO Min Wei delivered a keynote address. He introduced his company's vision for constructing a native 4D world model designed to underpin universal physical intelligence.

Min Wei explained that the proposed model is fundamentally four-dimensional, combining three spatial dimensions with one temporal dimension. This aligns with first principles since our physical world is inherently 4D. For embodied AI, the essential foundation must be four-dimensional. Other sensory inputs like touch and hearing represent physical phenomena within the 4D spacetime continuum. By focusing on reconstructing, understanding, generating, and predicting this basic 4D framework, a base can be formed to carry broader physical laws.

Historically, major technology shifts have been driven by data. ImageNet pioneered the scaling of annotated images from ten thousand to one million, marking the true beginning of AI. Similarly, one-dimensional large language models and two-dimensional video generation models succeeded by leveraging vast datasets, with LLMs solving the unlabeled data problem and internet videos fuelling 2D models. However, entering the 4D era presents the challenge of insufficient 4D data. Robots today struggle to leave the marathon track and enter factories or households because they lack generalization capabilities. This shortfall stems directly from the data disparity: robots are trained on 1D language and 2D video but must operate in a 4D world. Humans achieve impressive generalization precisely because they learn from 4D data.

The industry currently possesses only about 1,500 hours of annotated 4D datasets. Simulation-based virtual data suffers from the sim-to-real gap, a persistent challenge. Rather than focusing solely on data collection, the priority must shift to data distillation. Collection is labour-intensive, as seen during the autonomous driving era, but embodied intelligence demands extracting high-dimensional insights from low-dimensional inputs. By placing a few cameras in a scene, data can be captured without invasive head-mounted devices, reducing costs below that of generative synthetic data. After multi-view alignment, physical information can be distilled via models. Just as China excels at refining rare earth elements, though they exist worldwide, the core competency lies in data distillation and reproduction. A complete 4D dataset functions as a physical experiment, capturing geometry, occlusion relationships, contact processes, and spatial structure far richer than pixel-based 2D data. While humans instinctively reconstruct video content mentally, models struggle, making 4D capture essential for proper understanding.

Min Wei showcased a light-field capture from Shougang Park, originally used in a Spring Festival performance with dancer Liu Haocun in the program "Meng Di," demonstrating 4D data's real-world application through digital avatars. Traditional capture systems require 80 cameras and massive computing power, limiting productivity like steelmaking in feudal societies. YingShen Intelligence has reduced the requirement to just six cameras while maintaining 4D data capture and distillation capabilities, transforming handcrafted capture into industrialized production. This technology is being demonstrated through a live 4D broadcast at Hall 12 of CIFTIS.

Addressing model architecture, Min Wei asserted that current open-source models remain immature precisely because they lack an embodied-specific foundation. Two existing routes dominate: VLA models built on large language models and WAM models based on video generation. Recent discussions around GPT-6 suggest its base capabilities may already overshadow dedicated VLA approaches. Similarly, video-generation-based world models gained traction this year. However, embodied AI requires a 4D foundation grounded in physical world understanding rather than video pixel generation. Our world comprises molecules and point clouds, not mere pixels. Experiments with models like Seedance 2.5 from ByteDance demonstrate failures in generating delicate tasks such as tying shoelaces, with correct endpoints but erroneous intermediate steps. Even the most advanced 2D video generation models cannot handle precise manipulation tasks due to fundamental misunderstanding of physical interactions. Academic research is pivoting toward 3D and 4D approaches, exemplified by Li Feifei's PointWorld using 4D point clouds as representational backbone, though it does not solve the mass 4D data acquisition problem.

YingShen Intelligence's methodology comprises four stages: reconstruction, understanding, generation, and prediction. Through 4D reconstruction, understanding becomes possible; understanding enables generation and fills unseen perspectives. Li Feifei's recent Atlas framework, predicting next viewpoints rather than tokens, aligns closely with this philosophy. Traditional video or language-based foundations merely mimic rather than understand, a key limitation. In-context learning has emerged as significant progress, converting comprehension into broader imitation. The company's goal is achieving deeper understanding with fewer resources. A closed loop spanning data, models, and robots enables multi-view capture, 4D reconstruction, physical reproduction, and real-world execution.

Live demonstrations included generating a 4D scene from a single image, multi-view input, and one sentence description. Another demo produced six viewpoint videos plus dense point clouds from text and image inputs. Compared to pure video generation, this approach maintains enhanced spatiotemporal consistency and reduces hallucination due to synchronized dense point cloud generation. Additional demos featured chopping chili peppers with millimetre-level precision and shoe-lacing scenarios, requiring only a sentence and image input. The company targets dual-flexible-object manipulation across industries, focusing initially on footwear manufacturing and digital entertainment. They have established a unified infrastructure supporting two application tracks driven by a single data flywheel.

Min Wei noted that the industry has rapidly converged toward 4D approaches since his company's founding two years ago, with World Labs' Atlas release and major domestic tech firms developing spatial video generation models. He proposed a two-dimensional framework: the horizontal axis represents progression from low to high dimensions (language to video to true 4D), while the vertical axis depicts scaling from demonstrations to universal real-world deployment. Four critical elements determine 4D embodied breakthroughs: data, models, real scenarios, and scaled delivery. Next stages involve scaling 4D collection, iterating native world models, and expanding beyond apparel and entertainment into broader applications. Native 4D models will become the next-generation foundation alongside large language and video generation models. Min Wei concluded by inviting all attendees to collaborate in building this physical-world base ecosystem from Shijingshan, open to sharing 4D capture, data distillation, and foundation model capabilities.

免责声明:投资有风险,本文并非投资建议,以上内容不应被视为任何金融产品的购买或出售要约、建议或邀请,作者或其他用户的任何相关讨论、评论或帖子也不应被视为此类内容。本文仅供一般参考,不考虑您的个人投资目标、财务状况或需求。TTM对信息的准确性和完整性不承担任何责任或保证,投资者应自行研究并在投资前寻求专业建议。

热议股票

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10