模型发布 | 谷歌发布Gemini 3.8 Live双模型,强化实时语音AI

老虎资讯综合
Yesterday

谷歌周二推出两款新模型Gemini 3.8 Live和Gemini 3.8 Live Extended Thinking,重点提升语音对话中的推理能力。

谷歌表示,Gemini 3.8 Live主要面向需要大规模使用的场景,在对话能力、响应速度和视觉理解方面进行了优化。Extended Thinking版本则针对更复杂的任务,可以进行多步骤推理。

两款模型都支持实时处理视觉信息。Gemini 3.8 Live支持97种语言,并可以在对话过程中自动切换语言。用户提出请求后,模型还可以在后台调用工具和API,同时继续与用户交流,不必等任务完成后再继续对话。

对于需要更多计算的任务,Gemini 3.8 Live Extended Thinking可以一边进行推理,一边与用户交流。谷歌表示,模型在执行较复杂的后台任务时,会通过语音告知用户当前进展。

在Artificial Analysis的Speech to Speech Quality Index测试中,Gemini 3.8 Live Extended Thinking得分82.6分,排名第一;该模型在Big Bench Audio测试中的得分为97.7%。Gemini 3.8 Live在Speech Agent Arena中排名第二。

在功能方面,Gemini 3.8 Live可以近实时处理视觉输入,为语音交流提供更多上下文。该模型支持97种语言,并能够在对话过程中自动切换语言。

新模型还可以在后台执行工具和API调用,同时继续与用户对话。例如,用户提出请求后,模型可以先进行回应,在后台完成任务,而不必让用户等待整个过程结束。

对于需要更复杂推理的任务,Gemini 3.8 Live Extended Thinking可以在思考的同时继续进行语音交流。谷歌表示,该模型能够通过“让我查一下”等语音提示回应用户,并在执行多步骤后台任务时实时播报进展。

谷歌表示,两款模型已开始陆续开放。开发者可以通过Gemini API和Google AI Studio使用,企业客户可以在Gemini Enterprise中进行预览。此外,Gemini 3.8 Live已开始在Google Search Live中推出,Extended Thinking版本也开始向Gemini Live用户开放。

谷歌还表示,其AI产品生成的音频均加入SynthID水印,以便识别AI生成内容。

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Most Discussed

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10