Following an initial scare, Wall Street is actively discussing the potential impact of Kimi K3 on computing power requirements.
On July 16th, the Chinese large language model company Moonshot AI unveiled its next-generation model, Kimi K3, boasting a parameter count of 2.8 trillion, making it the world's largest open-source model by this measure. The model natively supports visual understanding and features a 1 million token context window. Within hours of its release, it topped the Arena leaderboard, a highly regarded AI coding tool evaluation list.
The model quickly garnered global attention. Late on July 19th, the Kimi team at Moonshot AI issued a statement titled "Regarding Computing Power Shortages and Temporary Pause on New Memberships," announcing an immediate halt to new consumer subscriptions to prioritize all available computing power for existing subscribers. The announcement stated, "Over the past 48 hours, user request volume has far exceeded our projections and is approaching the load limit of our existing cluster."
The emergence of K3 reminded US stock investors of the market shock caused by DeepSeek-R1. In early 2025, when DeepSeek's R1 model was released, concerns that AI might require less computing power than expected led to a nearly $600 billion single-day loss in market capitalization for AI chip leader NVIDIA Corp (NASDAQ: NVDA).
Consequently, US markets reacted similarly to K3. On July 17th local time, the US semiconductor sector saw a significant decline. Chip stocks broadly fell, with the Philadelphia Semiconductor Index dropping 1.63%, down 20.2% from its all-time high set on June 22nd.
However, after a weekend to digest the news, the market staged a strong rebound. On July 20th, the US semiconductor sector surged, with the Philadelphia Semiconductor Index jumping over 2%, Micron Technology, Inc. (NASDAQ: MU) gaining over 4%, and both Intel Corporation (NASDAQ: INTC) and Advanced Micro Devices, Inc. (NASDAQ: AMD) rising more than 3%.
Behind the 2.8 Trillion Parameters: Memory Pressure Remains
Beyond its massive parameter count, the deeper reason for K3's popularity lies in its architectural innovation and leap in "long-horizon tasks" capability. The model employs a Stable LatentMoE architecture, achieving precise dynamic sparse activation across 896 expert modules, improving overall scaling efficiency by 2.5 times.
Combined with its 1 million token context window and native visual capabilities, K3 can autonomously complete highly complex workflows requiring days of effort, such as long-range programming, intricate scientific calculations, and even automated video editing.
However, analysis points out that comparing K3 and R1 may overlook a crucial distinction. DeepSeek's breakthrough lay in reducing the training and operational costs of AI models. While Kimi K3 improves computational efficiency, it does so with a much larger model, placing higher demands on memory infrastructure. This is expected to continue driving development for chipmakers like SK Hynix Inc (KRX: 000660), NVIDIA Corp (NASDAQ: NVDA), and Taiwan Semiconductor Manufacturing Company Limited (NYSE: TSM).
Bloomberg analysis notes that Kimi K3, with its 2.8 trillion parameters, is the largest AI model from China to date. The model pushes sparsity—a metric for computational efficiency—to a new historical high. Higher sparsity means a lower proportion of the model's total parameters are activated when processing specific tasks.
Nevertheless, all 2.8 trillion parameters still need to be fully resident in memory. Even with compression using low-precision data formats, Kimi K3 occupies approximately 1.4TB of memory space. Therefore, deploying this model necessitates clusters of AI processors equipped with massive memory.
"Open-source models are typically more memory-hungry than closed-source ones," an analyst team at UBS Group AG (SWX: UBSG) stated in a recent report, "because they have longer context windows. Even after quantization, the absolute demand for KV cache continues to grow, making deployment more reliant on HBM and memory chips."
Wall Street's Analysis: The Jevons Paradox Reappears
In response to K3's impact, several Wall Street investment banks have issued reports, converging on a core view: K3 is not a signal of diminishing computing power demand but an accelerator for infrastructure build-out.
Citi semiconductor analyst Peter Lee directly referenced the classic economic concept of the "Jevons Paradox" in his report—where technological progress improves the efficiency of resource use but ultimately leads to an increase, not a decrease, in total consumption of that resource.
Peter Lee argues that while K3's inference efficiency is excellent, the pressure from KV cache on general-purpose memory (like server DDR5 and eSSD) actually increases with context length expansion. This will continue to boost the performance of industry giants like SK Hynix Inc (KRX: 000660), Micron Technology, Inc. (NASDAQ: MU), and Taiwan Semiconductor Manufacturing Company Limited (NYSE: TSM).
On the software and application side, analysis firm Stifel believes that lower token prices improve the unit economics for SaaS companies, lower the barrier to innovation, and also bring robust incremental demand for cybersecurity service providers—as the exponential expansion of large model deployment inevitably creates more attack surfaces requiring protection.
However, from a risk perspective, Bank of America Merrill Lynch cautions that if models become cheaper but usage does not expand significantly in tandem, the growth logic for computing power demand could weaken: "If efficiency gains outpace workload growth, we might see a pullback in infrastructure build-out."
Meanwhile, Chinese AI models are accelerating their global penetration. According to statistics from the open API gateway OpenRouter, the share of global developer traffic using Chinese AI models has grown from less than 2% a year ago to now exceeding 45%.
Additionally, on the afternoon of July 19th, the official account for Alibaba's Qwen large model announced that the preview version of Qwen3.8-Max had debuted on Alibaba's Token Plan, Qoder, and QoderWork platforms. Qwen3.8 will be released and open-sourced soon, with the new model reaching 2.4 trillion parameters, potentially making it the most powerful model besides Fable 5.
These releases indicate that competition among top-tier AI models is broadening beyond US frontier labs like OpenAI and Anthropic PBC. This could pressure the pricing power of US model providers while continuing to drive investment in AI infrastructure.
Bank of America Securities semiconductor analyst Vivek Arya believes that in the current competitive environment, the response from US frontier AI labs is "not less compute, but more." If Chinese open-source models continue to close the gap, OpenAI, Anthropic, and Google must maintain differentiation through larger-scale training, heavier inference, and faster iteration.
Nomura Securities Asia-Pacific technology team analyst Duan Bing emphasized that competition and innovation in the global large model market will only intensify: "As we get closer to AGI, the application of generative AI in consumer and enterprise segments will continue to expand. Both frontier AI labs and hyperscale cloud platform companies are likely to continue investing during this phase to maintain competitive positions—as the scaling laws persist, we interpret this competition as a positive for the AI infrastructure value chain."