Morgan Stanley: AI Memory Shortage to Last Several Years, Three Paths to Break the Bottleneck, Structural Opportunities in Storage and Heterogeneous Computing

Stock News
Yesterday

According to a semiconductor industry research report released by Morgan Stanley, the DRAM memory shortage will persist throughout the current AI industry cycle, and AI computing expansion will not wait for new wafer fabrication plants to be built and brought online.

The industry is bypassing the memory bottleneck through three major technological paths — hardware de-speccing, inference architecture disaggregation, and CXL memory pooling — to continue advancing computing infrastructure construction under supply constraints.

The firm remains bullish on storage leaders such as Micron Technology Inc (MU.US) and SanDisk Corp (SNDK.US), while also favoring incremental opportunities in the CXL interconnect and heterogeneous inference sectors, recommending Astera Labs Inc (ALAB.US), Marvell Technology Inc (MRVL.US), Cerebras Systems Inc (CBRS.US), and NVIDIA Corp (NVDA.US).

Morgan Stanley emphasizes that the current AI memory shortage is not a short-term disruption but a structural contradiction in which computing performance growth far outpaces memory supply growth.

Frontier large model scale doubles every 6 months, top model context windows are growing at an annual rate of 5-6 times, and combined with continuously rising inference concurrency demand, memory capacity and bandwidth requirements continue to surge.

Meanwhile, DRAM fabrication plant construction cycles span several years, supply elasticity is extremely low, and the shortage situation will persist for years.

NVIDIA CEO Jensen Huang has also publicly stated that the industry needs to shift its approach and address memory constraints through architectural innovation rather than simply waiting for capacity expansion.

Facing tight HBM and main memory supply, the most direct response for the industry is selectively reducing per-device memory specifications (de-speccing).

Taking NVIDIA's Rubin architecture as an example, per-rack LPDDR5 capacity was reduced from a planned 54TB to 28TB, and per-GPU HBM capacity was lowered from 288GB to 192GB, ensuring overall system shipment volumes by reducing memory stack heights.

Morgan Stanley points out that de-speccing does not eliminate the memory bottleneck but rather shifts pressure across memory tiers: after cutting high-speed local memory, non-high-frequency data such as KV cache migrates down to NAND storage, while cross-GPU data interaction increases, driving demand for network interconnect bandwidth — essentially supplementing scarce high-bandwidth memory with higher-speed network interconnects and lower-cost storage resources.

The second path to breaking the bottleneck is inference architecture disaggregation.

AI inference comprises two significantly different phases: Prefill and Decode. Prefill is primarily compute-intensive, while Decode is highly dependent on memory bandwidth.

Traditional architectures use the same accelerator for both task types, resulting in low resource utilization.

The industry is now accelerating toward heterogeneous inference, splitting the two phases onto different hardware: compute-intensive Prefill is handled by general-purpose GPUs, while memory-bandwidth-intensive Decode is processed by specialized accelerator chips equipped with large on-chip SRAM.

Typical examples include Cerebras's wafer-scale engine and the Groq LPU architecture acquired by NVIDIA, both of which can deliver memory bandwidth efficiency in the decode phase far exceeding traditional GPUs.

Morgan Stanley believes that heterogeneous inference will become an important evolutionary direction for AI infrastructure, and compute vendors specializing in the decode phase will gain a clear incremental market.

The third path is CXL technology restructuring the memory system.

CXL (Compute Express Link) enables memory expansion, sharing, and pooling through high-speed interconnects, decoupling memory from a single processor and becoming a core technology path for breaking through memory capacity constraints.

Morgan Stanley estimates that AI demand will drive the CXL and related memory add-on chip market to reach approximately $6 billion by 2030, far exceeding the traditional CPU memory expansion market size.

Its core value lies in building a tiered memory architecture: the highest-frequency data remains in HBM, moderately frequent data is placed in CXL memory pools, and low-frequency cold data sinks to NAND, using cost gradients to match data access frequencies.

Regarding investment themes, Morgan Stanley reiterates three directions: first, continue to overweight Micron and SanDisk, as de-speccing stems from supply shortages rather than weakening demand — AI memory demand is trending upward long-term, and the supply gap will continue to absorb capacity; second, position in CXL and scale-up interconnect leaders Astera Labs and Marvell; third, watch heterogeneous inference beneficiary Cerebras, as well as NVIDIA, which is completing its heterogeneous layout through the Groq acquisition.

Risk warnings: AI computing demand growth falling short of expectations; CXL technology deployment progress slower than anticipated; memory capacity expansion exceeding expectations.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Most Discussed

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10