阿里云下代语音模型发布,Qwen落地多款智能终端

新浪科技
Sep 22

  新浪科技讯 9月22日上午消息,阿里巴巴在云栖大会上宣布语音模型家族Qwen-Audio-3.1全新升级,包含语音转写(ASR)、语音合成(TTS)和实时语音交互(Realtime)三大系列。

  基于下代架构的音频理解模型Qwen-Audio-3.1-ASR-Next全新发布,它能够理解人物情绪、音乐、环境声与机械声,完成声音描述、事件定位、音频问答与推理,让“这段录音中人的情绪怎样”这样的问题能够被理解和准确回答。

  同时,音频创作模型Qwen-Audio-3.1-TTS-Next全新亮相,它采用全新架构,可根据一段脚本,直接生成融合对白和环境声的完整音频,极大提高有声书、影视剧和播客等专业创作的效率。

  Qwen-Audio-3.1-Realtime实时交互能力再提升,能够边听、边想、边说,并进一步增强多语言交互、角色扮演和共情能力。目前,该系列模型已应用于千问办公和Qoder,并接入QwenNote A2随身助理、QwenNote Eva桌面机器人和千问AI眼镜等硬件产品。

  此外,同声传译模型Qwen3.8-LiveTranslate首次现身,人类同传平均时延在4秒以上,但该模型可将时延压缩至不到2.5秒。据了解,新模型已接入千问AI眼镜、QwenNote A2、钉钉耳机等AI硬件。

阿里大模型加速走向终端。AI手机全栈解决方案Qwen Intelligence也将于大会期间发布,专为手机场景深度优化,为合作伙伴提供基于千问大模型的 Agent 技术平台,让手机能够更可靠地完成跨应用的复杂任务。

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Most Discussed

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10