郭明𫓹:英伟达已重启Rubin CPX项目 将于2027年第一季度开始生产

环球市场播报
Sep 01

  知名分析师郭明𫓹发文表示,就在市场已经开始相信Rubin CPX已被英伟达从产品路线图中剔除之际,其最新的产业调研显示,英伟达已重启该项目,预计将于2027年第一季度(1Q27)开始生产。与先前设计相比,重启后的Rubin CPX提供更强的预填充(prefill)性能,并对GPU规格与机架架构均作出重大改动,凸显出英伟达对预填充解决方案的高度重视。

  主要变动如下:

  1.Rubin CPX GPU规格 CPX提供接近Rubin的计算性能,并匹配Rubin每颗GPU最高2,300瓦的额定功耗上限。CPX改用168GB的HBM4(Rubin为288GB,此前CPX设计为128GB的GDDR7)。

  2.机架设计 新版CPX采用独立的MGXETL机架,而非此前设计中与Rubin共用机架。客户可根据需求选择64、128、192或256颗CPX GPU。在单个CPX机架内,每64颗CPX GPU构成一个机架模块(rack module),该模块由八个计算托盘(每个托盘8颗CPX GPU)和一个交换托盘组成。

  3.纵向扩展与横向扩展 NVLink仅用于每个托盘内8颗CPX GPU之间的纵向扩展(scale-up),每颗CPX的NVLink带宽为1–1.5TB/s(相比之下,每颗Rubin为3.6TB/s)。每个机架模块内各托盘之间的横向扩展(scale-out)通过Spectrum-6以太网运行,采用全铜L1链路。跨机架模块的横向扩展则由各模块的Spectrum-6交换机经OSFP光链路处理。

  4.工作原理 CPX必须与Vera Rubin NVL72搭配使用,英伟达建议CPX与Rubin GPU按1:1的比例配置。CPX负责预填充并构建KV缓存,随后通过以太网RDMA将KV缓存传输给Rubin进行解码(decode)。

  5.产品定位:长上下文预填充场景下的单位成本最佳性能 如今超过50%的AI推理工作负载来自对输入上下文的处理及相应KV缓存的构建。因此,CPX提供了一种更灵活、成本更低的预填充处理方式。每个八颗CPX的托盘拥有约1.34TB的HBM4,足以满足大多数长上下文预填充工作负载及相应的KV缓存需求。

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Most Discussed

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10