Google 称 AI 服务器记忆体成本破 75%,推软硬体双轨策略

链捕手
Sep 02

ChainCatcher 消息,SEMICON Taiwan 2026 记忆体高峰论坛 1 日登场,Alphabet 旗下 Google Cloud 供应链基础架构资深总监 Nikhil Cherian 指出,随着多模态与混合专家架构普及,AI 运算已从算力受限转向记忆体受限,高效能记忆体已占 AI 服务器硬体物料清单成本 75% 以上。面对容量、频宽与功耗瓶颈,Google 透过推论与训练硬体分流、无损量化软体算法等软硬体双轨策略,突破 AI 记忆体瓶颈。

Google 在硬体架构上采取分流策略,推出针对低延迟推论的 TPU 8i 以及专攻超大规模训练的 TPU 8t。TPU 8i 配备 288 GB 高频宽记忆体,芯片上 SRAM 容量扩增 3 倍至 384 MiB,将动态对话状态与键值快取置于芯片内部,实现零芯片外延迟。TPU 8t 透过 9600 颗芯片组成超大型运算丛集,HBM 共享池化达到 2 PB 规模,消除芯片外数据传输瓶颈,搭配 TPU Direct Storage 技术。

Google 开发出免训练的 TurboQuant 无损量化算法,将大模型的键值快取从 32 位元压缩至 3 位元,精确度零损失前提下缩小 6 倍记忆体占用,带来 8 倍注意力运算加速,导入旧世代 DRAM 整合技术延长零组件生命周期。

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Most Discussed

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10