XPeng Fills a Key Gap in Its Physical AI System with New Visual Foundation

Deep News
07/22

On July 21, XPeng Inc. (NYSE: XPEV) announced the official launch of the TuringViT efficient visual encoder, designed for VLM (Visual Language Model) and VLA (Visual Language Action Model) applications. The company stated it has systematically overhauled the architecture design, data paradigm, and training process. This encoder is intended for use across three core business areas: XPeng's intelligent driving systems, intelligent cockpit, and the IRON humanoid robot.

In VLM/VLA systems, the visual encoder is often considered the "perception gateway" for physical AI, responsible for converting images or videos into visual features for subsequent processing by language and action modules.

As VLM/VLA technology moves towards real-world application, high-resolution images, multi-view inputs, and continuous video frames place greater demands on the efficiency and temporal processing capabilities of visual encoders.

XPeng believes that the industry's common approach of "reusing open-source general-purpose ViT" models struggles to simultaneously meet the performance, latency, and customization requirements of scenarios like intelligent driving and embodied robotics.

The design of TuringViT is summarized across three dimensions.

Regarding architecture, TuringViT is primarily based on Turing Linear Attention (TLA), while retaining a small number of standard multi-head attention mechanisms, forming a hybrid Turing Block architecture. This allows computational complexity to scale nearly linearly instead of quadratically under high-resolution inputs.

According to test data released by XPeng, at a resolution of 1536×1536, the TuringViT-18L achieves an inference throughput 3.04 times that of the Seed1.5-ViT model.

The encoder is offered in two versions: the TuringViT-18L, which contains 3 Turing Block groups and is geared towards deployment with dynamic resolutions, and the TuringViT-24L, which contains 4 groups and focuses on enhancing representational capability.

On the data front, TuringViT employs the VISTA-Curation multimodal data governance pipeline. This pipeline conducts multi-stage filtering, re-captioning, and scoring of image-text and video data to increase the supervisory information per sample. TuringViT was pre-trained using 0.85B image-text pairs, approximately 10% of the training data scale used for SigLIP2-L. On six zero-shot classification benchmarks, including ImageNet-1K, the TuringViT-24L achieved an average accuracy of 83.6%.

For training, TuringViT uses a four-stage progressive native dynamic resolution training scheme. This approach adapts to the input characteristics of downstream VLM/VLA applications from the pre-training stage itself, and, combined with 2D rotary position encoding, supports inputs of varying sizes and aspect ratios. XPeng states this design reduces reliance on the "fixed-resolution pre-training + post-hoc adaptation" pathway.

The company outlined three primary business application scenarios for TuringViT.

In the intelligent driving domain, XPeng identifies TuringViT as the core visual encoder for its second-generation VLA model. It is tasked with processing high-resolution images from multi-view surround cameras and multi-frame dynamic road scene inputs, providing visual tokens for the predictive world model.

For the intelligent cockpit, TuringViT's VLM-native characteristics are used to align visual features with the language model and support visual inputs with different aspect ratios and proportions.

Within the technological framework of the XPeng IRON humanoid robot, TuringViT is positioned as the foundational visual module, handling functions such as object recognition, spatial relationship understanding, operable area detection, and dynamic environment tracking.

During an earnings call on March 20, He Xiaopeng stated that the XPeng IRON humanoid robot is planned for mass production by the end of 2026, with a target monthly production capacity in the thousands by year-end, and will be commercially deployed first in XPeng stores.

XPeng indicated that the architectural design and training process of TuringViT are not dependent on specific hardware platforms, offering the industry a reproducible training path for large visual models.

From a technology roadmap perspective, TuringViT addresses a missing component in XPeng's physical AI technology stack: the visual encoder.

Previously, XPeng has disclosed several technologies, including the multi-view generative world model X-World, the world model inference acceleration engine X-Cache, the predictive world model X-Foresight, and the X-Mind technology framework. TuringViT now occupies the perception input layer of this physical AI technology system.

However, the journey from a technical release of a visual encoder to large-scale deployment in real-world scenarios still involves engineering adaptation, computing power deployment, and scenario validation.

Whether TuringViT can operate stably across the three distinct physical environments of vehicles, cockpits, and robots remains to be seen through continued observation. XPeng stated it will continue to expand its scale of high-quality image-text-video data and deepen technological exploration in temporal modeling and embodied vision.

免责声明:投资有风险,本文并非投资建议,以上内容不应被视为任何金融产品的购买或出售要约、建议或邀请,作者或其他用户的任何相关讨论、评论或帖子也不应被视为此类内容。本文仅供一般参考,不考虑您的个人投资目标、财务状况或需求。TTM对信息的准确性和完整性不承担任何责任或保证,投资者应自行研究并在投资前寻求专业建议。

热议股票

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10