Unisound (09678) Upgrades Voice Foundation Models U2-ASR and U2-TTS with Expanded Multilingual Capabilities

Stock News
07/28

Unisound (09678) has announced a comprehensive upgrade to its U2-ASR and U2-TTS voice models, strengthening the company's multimodal large model capabilities. This upgrade marks a significant expansion from covering hundreds of Chinese dialects to supporting a wide range of global languages, reinforcing the infrastructure for global voice interaction. The enhanced models are designed to provide efficient and low-barrier technical support for cross-border applications, embodied intelligence, and agent-based interactions.

In this upgrade, U2-ASR has added recognition capabilities for 13 new international languages, covering key overseas markets in Europe, Southeast Asia, the Middle East, and Latin America. Meanwhile, U2-TTS now supports speech synthesis for 8 additional Southeast Asian languages. As a result, the U2 voice large model now supports over 100 Chinese dialects and more than 15 international languages. This allows enterprises to process audio content in different languages by connecting to just one model and one set of APIs, significantly lowering the development, deployment, and maintenance threshold for multilingual voice services.

Under a unified evaluation benchmark, U2-ASR has demonstrated outstanding performance compared to industry-leading models, achieving an average character error rate (CER) of just 6.58% across 113 languages. In real-world business scenarios where language tags are not provided, the model maintains high accuracy through automatic language identification and closed-set routing, effectively avoiding language misidentification. Additionally, in the ChinaVoices Challenge 2026 Chinese multi-dialect speech recognition competition, U2-ASR achieved a top ranking with a CER of 9.235%, securing first place in the limited-data track and second place in the open-data track, further validating its advanced speech recognition technology.

U2-TTS has also demonstrated leading performance in subjective evaluations of intelligibility and naturalness compared to mainstream industry models. It utilizes a streaming neural network acoustic model to output 24kHz high-fidelity audio chunk by chunk in real time, significantly reducing initial packet latency and meeting the demands of strong real-time interactions such as live voice conversations. This upgrade enables seamless collaboration between U2-ASR and U2-TTS, providing enterprises with a complete multilingual voice interaction pipeline for building agent-based systems, positioning voice technology as a core interaction infrastructure for global business.

The upgraded U2-ASR and U2-TTS models are now fully available on the company's TokenHub large model service platform, with standard APIs open for access. The company will continue to adhere to its philosophy of "intelligence for good," leveraging technology to break down language barriers and promote the universal sharing of AI technology, ensuring that regions at different stages of development can equally enjoy the technological dividends of the AGI era.

免责声明:投资有风险,本文并非投资建议,以上内容不应被视为任何金融产品的购买或出售要约、建议或邀请,作者或其他用户的任何相关讨论、评论或帖子也不应被视为此类内容。本文仅供一般参考,不考虑您的个人投资目标、财务状况或需求。TTM对信息的准确性和完整性不承担任何责任或保证,投资者应自行研究并在投资前寻求专业建议。

热议股票

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10