NVIDIA Considers Reducing Memory in Rubin Ultra to Address High-Bandwidth Memory Supply Constraints

Deep News
08/07

The artificial intelligence infrastructure boom is pushing supply chains to their limits.

NVIDIA is evaluating a significant adjustment to its next-generation AI GPU, the Rubin Ultra. The company is considering producing versions with lower memory configurations than originally planned, aiming to alleviate production pressure caused by a shortage of high-end High Bandwidth Memory (HBM).

This move indicates that even NVIDIA, which holds a dominant position in the GPU market, must balance product specifications with supply availability. The change reflects that HBM supply has become a critical bottleneck in the AI industry chain. It also suggests that AI server costs could rise further, and the pressure on data center construction will continue to be passed down the line.

On Thursday, NVIDIA shares closed down 0.1%, while SK Hynix shares fell 4.97%.

Addressing HBM Shortages with Multiple Lower-Memory Variants

Sources familiar with the matter said that over the past few weeks, NVIDIA has tested at least three different versions of the Rubin Ultra GPU. Some of these versions use less memory capacity than initially planned. A key reason for considering lower-memory versions is the potential inability to secure enough high-end HBM chips to support mass production of the original design.

The Rubin Ultra is a key component of NVIDIA's next-generation AI computing platform. It is positioned above the upcoming Rubin series and is seen as an important hardware platform for training ultra-large-scale AI models in the future. The original plan was to equip it with higher-capacity, higher-bandwidth HBM to improve model training and inference performance. However, with HBM supply remaining tight, NVIDIA is reassessing the product configuration. The goal is to ensure the product launches on schedule by adjusting memory specifications, rather than waiting for the supply chain to expand capacity.

Reduced Memory May Require Deploying More GPUs

For AI training, HBM not only determines a GPU's data throughput but also directly impacts the size of large models it can handle. If the Rubin Ultra ends up with a lower memory configuration, customers running large AI workloads, such as large language models, may need to deploy more GPUs to complete computing tasks that fewer chips could have handled originally.

While NVIDIA can partially compensate for the performance loss by increasing GPU compute power or optimizing interconnect bandwidth, overall system deployment costs and cluster complexity are likely to rise. For major cloud providers like Microsoft, Meta, Amazon, and Google, which are continuously increasing their AI capital expenditures, this implies further increases in future data center construction costs.

The AI Boom Exposes the Industry's Biggest Bottleneck

HBM has become one of the most critically scarce core components in the AI industry chain. In recent years, the rapid growth in demand for large model training has continuously pushed up GPU shipments. The HBM required for these GPUs is primarily supplied by a few companies, including Samsung Electronics, SK Hynix, and Micron. Due to the complex manufacturing process, high yield requirements, and long expansion cycles for HBM, supply growth has consistently struggled to keep pace with AI demand.

Industry experts widely expect HBM to remain in short supply for the next several years. This is a major reason why AI server prices remain high. NVIDIA's consideration of reducing the memory configuration for the Rubin Ultra reflects that even with the strongest bargaining power in the industry, its product planning is still constrained by HBM supply.

Cost Pressure Spreads Across the Tech Industry

The impact of rising HBM prices is not limited to AI servers. As key components like GPUs and HBM remain in short supply, hardware costs across the entire technology industry are being pushed up. Companies are increasing budgets for AI infrastructure, and more firms are being forced to raise capital expenditures to meet the demands of generative AI deployment.

At the same time, rising upstream chip costs are beginning to affect the consumer electronics sector. Some hardware manufacturers, including Apple, have already absorbed supply chain cost pressures by raising the prices of their end products. Market analysts believe that as long as AI computing demand continues to grow rapidly and the rate of new HBM capacity release remains limited, the competition for high-end memory resources will persist. HBM supply capacity will continue to be a key factor determining the pace of AI industry expansion.

免責聲明:投資有風險,本文並非投資建議,以上內容不應被視為任何金融產品的購買或出售要約、建議或邀請,作者或其他用戶的任何相關討論、評論或帖子也不應被視為此類內容。本文僅供一般參考,不考慮您的個人投資目標、財務狀況或需求。TTM對信息的準確性和完整性不承擔任何責任或保證,投資者應自行研究並在投資前尋求專業建議。

熱議股票

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10