DeepSeek下一波大模型被指即将问世:3万亿参数 性能超GPT-6 Astra

快科技
Sep 13

快科技9月13日消息,本周DeepSeek发布了V4.1 Flash,这个大模型架构变化之大,完全可以称得上是V5了,但DeepSeek很低调地没换大版本号。

Pro系列的新品确定了是V4.1 Pro,具体升级还没公布,如果是V4到V4.1这样的提升幅度,那V4.1 Pro性能上限也会很高,但考虑到之前V4 Pro-0813正式版的不及预期,大家期待值不要拉太高。

后续还有一波大模型,爆料者称它叫做DeepSeek Code 2.0,即将发布,时间窗口是9月份,预计性能会击败Mythos 5.1和GPT-6 Astra,这是当前最强大的两个前沿级大模型。

其他方面,DeepSeek Code 2.0预计会继续保持开源开放,设计重点是Computer Use,也就是电脑控制,AI替代你操作电脑完成很多工作,Astra目前很多让人惊讶的使用就来自这一方面。

这个Code 2.0拥有高达3万亿以上的参数量,显然这也是它提高性能的关键,这方面国内目前最强的是K3大模型的2.8万亿,千问的Qwen 3.8 Max是2.4万亿参数量,DeepSeek V4 Pro是1.6万亿参数。

从V4.1 Flash换用新结构使得参数量翻倍来看,这个Code 2.0从当前1.6万亿提升到3万亿参数的可信度是有的,想跟Mythos 5.1及Astra这样的大模型竞争就需要这个级别的参数量,不然性能上限是很难提升的。

这个爆料的可信度不好说,DeepSeek过去的23-24年倒是推出过三款名字带Code的大模型,不过那时候表现平平,这次要是重推偏向代码/逻辑的Code大模型,作为旗舰级产品倒也不失为一个方向,就让V4.1 Pro及Flash面向通用Agent任务即可。

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Most Discussed

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10