Unisound (09678) Upgrades Voice Foundation Models U2-ASR and U2-TTS with Expanded Multilingual Capabilities

Stock News
07/28

Unisound (09678) has announced a comprehensive upgrade to its U2-ASR and U2-TTS voice models, strengthening the company's multimodal large model capabilities. This upgrade marks a significant expansion from covering hundreds of Chinese dialects to supporting a wide range of global languages, reinforcing the infrastructure for global voice interaction. The enhanced models are designed to provide efficient and low-barrier technical support for cross-border applications, embodied intelligence, and agent-based interactions.

In this upgrade, U2-ASR has added recognition capabilities for 13 new international languages, covering key overseas markets in Europe, Southeast Asia, the Middle East, and Latin America. Meanwhile, U2-TTS now supports speech synthesis for 8 additional Southeast Asian languages. As a result, the U2 voice large model now supports over 100 Chinese dialects and more than 15 international languages. This allows enterprises to process audio content in different languages by connecting to just one model and one set of APIs, significantly lowering the development, deployment, and maintenance threshold for multilingual voice services.

Under a unified evaluation benchmark, U2-ASR has demonstrated outstanding performance compared to industry-leading models, achieving an average character error rate (CER) of just 6.58% across 113 languages. In real-world business scenarios where language tags are not provided, the model maintains high accuracy through automatic language identification and closed-set routing, effectively avoiding language misidentification. Additionally, in the ChinaVoices Challenge 2026 Chinese multi-dialect speech recognition competition, U2-ASR achieved a top ranking with a CER of 9.235%, securing first place in the limited-data track and second place in the open-data track, further validating its advanced speech recognition technology.

U2-TTS has also demonstrated leading performance in subjective evaluations of intelligibility and naturalness compared to mainstream industry models. It utilizes a streaming neural network acoustic model to output 24kHz high-fidelity audio chunk by chunk in real time, significantly reducing initial packet latency and meeting the demands of strong real-time interactions such as live voice conversations. This upgrade enables seamless collaboration between U2-ASR and U2-TTS, providing enterprises with a complete multilingual voice interaction pipeline for building agent-based systems, positioning voice technology as a core interaction infrastructure for global business.

The upgraded U2-ASR and U2-TTS models are now fully available on the company's TokenHub large model service platform, with standard APIs open for access. The company will continue to adhere to its philosophy of "intelligence for good," leveraging technology to break down language barriers and promote the universal sharing of AI technology, ensuring that regions at different stages of development can equally enjoy the technological dividends of the AGI era.

免責聲明:投資有風險,本文並非投資建議,以上內容不應被視為任何金融產品的購買或出售要約、建議或邀請,作者或其他用戶的任何相關討論、評論或帖子也不應被視為此類內容。本文僅供一般參考,不考慮您的個人投資目標、財務狀況或需求。TTM對信息的準確性和完整性不承擔任何責任或保證,投資者應自行研究並在投資前尋求專業建議。

熱議股票

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10