How the AI Ecosystem Is Navigating Sky-High Memory Prices: Spec Downgrades, Workload Splitting, and CXL Memory Pooling

Deep News
1小时前

The memory shortage has become the most persistent structural bottleneck of this AI buildout cycle, and the pace of AI compute expansion will not wait for new wafer fabs to come online. Faced with this contradiction, the entire supply chain is exploring three workaround paths 鈥?reducing memory specifications, disaggregating inference workloads, and pooling memory through CXL technology 鈥?while spawning new investment opportunities along the way.

In its latest research report, Morgan Stanley noted that NVIDIA CEO Jensen Huang recently made clear that the industry needs entirely new thinking to address the memory bottleneck. This is not a pessimistic signal but a proactive anticipation of a shortage that will persist for years. The intensity of the memory crunch may fluctuate, but AI will essentially absorb all available supply for the foreseeable future.

On the investment side, Morgan Stanley named Astera Labs (ALAB) and Marvell (MRVL) as potential beneficiaries of CXL and scale-up directions, with Cerebras (CBRS) as the core beneficiary of inference workload disaggregation architectures, while maintaining overweight ratings on Micron (MU) and SanDisk (SNDK), on the logic that the memory shortage cycle is far from over.

Spec Reduction: A Stopgap That Shifts Pressure Rather Than Eliminating It

With both DRAM and NAND prices climbing, the most direct response for ecosystem players is to compress memory configurations per device or per server.

Morgan Stanley noted that NVIDIA is offering multiple specification versions of its Rubin platform to customers: LPDDR5 capacity per rack has been compressed from an originally planned 54TB to 28TB, with modules shrinking from 192GB SOCAMM2 to 96GB; on the HBM side, Rubin was originally planned to carry 288GB of HBM per GPU (8-high 12hi HBM4), but a 192GB version (8-high 8hi HBM4) is now expected. For Rubin Ultra, specifications have also been downgraded from 1TB HBM4e, with the actual comparable baseline potentially falling to the 192GB to 384GB range.

However, spec reduction is not a genuine solution. The report emphasized that when capacity at one memory layer is cut, data pressure inevitably shifts downward 鈥?the shrinkage of HBM and LPDDR is creating incremental demand for NAND and networking interconnects, meaning the memory bottleneck has shifted rather than disappeared.

Structural factors driving continuous expansion of AI memory demand span three dimensions: model sizes doubling every six months in the LLM era; context window lengths expanding roughly 5x per year; and rising inference concurrency requiring independent storage of KV caches for each session.

Morgan Stanley believes that the continued evolution of these three vectors means spec reduction can only be a temporary measure. Once supply eases, the industry will rapidly revert to higher specifications, and suppressed demand will be more easily absorbed.

Workload Disaggregation: Dedicated Hardware for Different Inference Stages

Inference consists of two fundamentally distinct stages: prefill, which processes input prompts in parallel and is compute-intensive; and decode, which generates output token by token and is more dependent on memory bandwidth and capacity. Separating the two and assigning each to hardware best suited to its characteristics is the second path to improving memory usage efficiency.

This trend accelerated markedly around the end of 2025. After NVIDIA acquired Groq, it integrated Groq's LPU (based on high-bandwidth on-chip SRAM) into the Vera Rubin platform, with Rubin GPUs handling prefill and decode requiring large KV cache capacity, Groq handling feed-forward and mixture-of-experts components, and NVIDIA Dynamo coordinating and transferring activation values. AWS subsequently announced a heterogeneous architecture with Trainium handling prefill and Cerebras handling decode, planned for launch on Amazon Bedrock in the first quarter of 2027; a similar Cerebras-AMD collaboration is expected to enter production in the fourth quarter of 2026. Matrix's Corsair has also entered mass production.

Morgan Stanley named Cerebras as the most direct pure-play beneficiary of workload disaggregation architectures. Cerebras's wafer-scale processor integrates large amounts of SRAM with compute units on the same silicon die, significantly reducing the frequency of data movement between processors and external memory, giving it a standout advantage in the decode stage. More importantly, workload disaggregation can materially improve the economics of Cerebras Cloud 鈥?according to disclosures from Cerebras and AMD, the joint system can boost throughput by up to 5x while maintaining Cerebras's inference speed, generating more revenue per deployment unit while simultaneously lowering cost per token.

CXL: A New Battleground from Memory Expansion to AI Inference

CXL (Compute Express Link) is a high-speed interconnect protocol based on the PCIe physical layer, whose core value lies in allowing multiple processors to access the same memory resource pool, fundamentally breaking the traditional architecture in which memory is tied to a specific CPU.

Its three main application forms are: memory expansion (adding capacity beyond the local memory channel limit for a single processor), memory sharing (multiple processors accessing the same data), and memory pooling (dynamically allocating memory across processors or servers to reduce idle waste).

Morgan Stanley noted that CXL previously served primarily general-purpose CPU computing scenarios, with limited penetration in AI due to latency disadvantages. But as inference scenarios see rapid expansion in capacity demand for long context, agentic workloads, and KV caches, large amounts of data do not need to reside in HBM at all times, giving CXL-attached DRAM a foothold as a lower-cost, higher-capacity memory tier.

On market size, Morgan Stanley raised its CXL addressable market forecast from the previous CPU-based estimate of over $4 billion to approximately $6 billion by 2030, with the increment mainly coming from new deployment scenarios on the AI server side. Astera Labs's Leo memory controller family (including custom designs targeting KV cache offload, expected to enter mass production in 2027) and Marvell's full product portfolio covering memory expansion (Structera X) and rack-level pooling (Structera S) are the two core names Morgan Stanley favors. Marvell management has previously positioned CXL as a revenue contributor of over $1 billion around 2028, and Morgan Stanley expects more details to be disclosed at the upcoming investor day.

Impact on Memory Stocks: Cycle Logic Anchored in Duration, Not Amplitude

Morgan Stanley acknowledged that from the perspective of maximizing near-term earnings, the aforementioned workaround strategies have some negative impact on DRAM makers 鈥?if the AI ecosystem were to grind to a complete halt due to the memory shortage, near-term pricing could be more favorable. But that scenario was never realistic, the report emphasized, and analysis of this memory cycle should be anchored primarily in duration rather than price amplitude.

The core argument: spec reduction stems from supply constraints rather than fading demand, and once supply improves, the industry will rapidly revert to higher specifications, with pent-up demand easily absorbing new capacity. Morgan Stanley also flagged a risk factor: if an AI slowdown stems not from declining demand but from infrastructure bottlenecks such as land, power, or fab capacity, such a pause could deliver a unique shock to the memory sector different from that felt by the compute sector.

Based on these judgments, Morgan Stanley maintained overweight ratings on MU and SNDK, believing the memory shortage cycle has not yet ended and that the allocation value of both remains intact.

免责声明:投资有风险,本文并非投资建议,以上内容不应被视为任何金融产品的购买或出售要约、建议或邀请,作者或其他用户的任何相关讨论、评论或帖子也不应被视为此类内容。本文仅供一般参考,不考虑您的个人投资目标、财务状况或需求。TTM对信息的准确性和完整性不承担任何责任或保证,投资者应自行研究并在投资前寻求专业建议。

热议股票

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10