Is Kimi K3 a Repeat of the 'DeepSeek Moment'? Wall Street Says It Actually Boosts Compute Demand

Deep News
Jul 21

The market is panicking over Kimi K3 as a potential "DeepSeek Moment 2.0," but this time around, Wall Street's assessment is markedly different.

Late on July 16th, Moonshot AI unveiled Kimi K3 in Shanghai. This open-source model with 2.8 trillion parameters scored 57 on the Artificial Analysis intelligence index, ranking it between third and fourth globally, on par with Anthropic's Claude Opus 4.8 and OpenAI's GPT-5.5. More crucially, on the Frontend Code Arena programming benchmark created by UC Berkeley, K3 topped the chart with a score of 1679, surpassing Claude Fable 5 and GPT-5.6 Sol, making it the first open-source model to outperform all leading closed-source overseas models on a major programming benchmark.

On July 17th, the U.S. semiconductor sector saw a notable decline. The market's knee-jerk reaction is understandable—the release of DeepSeek R1 in early 2025 triggered a sharp sell-off in compute stocks, based on the logic: if Chinese models are getting stronger, do U.S. AI companies still need to invest so heavily in compute? If Chinese models can approach frontier capabilities at a lower cost, will demand for NVIDIA products, HBM, servers, and networking equipment be reassessed?

However, according to the latest research reports from investment banks including UBS, Nomura, Bank of America Merrill Lynch, and Citi, Kimi K3 is not a terminator for compute demand but rather an accelerator.

Kimi K3 and DeepSeek R1 represent different types of impact. R1 primarily showcased "efficiency" to the market; K3 emphasizes "scale." Features like 2.8 trillion parameters, a 1M token context window, always-on inference, native multimodal capabilities, and a MoE architecture do not tell a light-asset story. They collectively increase pressure on inference, memory, networking, and storage.

How Powerful is Kimi K3?

Kimi K3 was released by Moonshot AI on July 16, 2026, with its full model weights scheduled for public release on July 27th. It is a 2.8 trillion parameter open-weight large language model, described by several institutions as the largest open-weight LLM currently available.

Its core configuration includes three key points: First, a 1M token context window, enabling the model to handle longer texts, larger codebases, and more complex corporate documents and research tasks. Second, always-on inference, designed not just for simple Q&A but for long-chain reasoning and Agent-based tasks. Third, native visual capabilities, allowing K3 to process not only text but also video, images, game development, front-end design, CAD, and other multimodal tasks.

Architecturally, K3 employs Kimi Delta Attention, Attention Residuals, and Stable LatentMoE. The MoE component activates 16 out of 896 experts per token. Moonshot AI states that overall scaling efficiency has improved by approximately 2.5 times compared to Kimi K2.

This explains why K3 is not a "cheaper K2." Price data compiled by Nomura shows K3 input costs $3 per million tokens, cached input costs $0.30, and output costs $15 per million tokens. According to Artificial Analysis, K3 costs about $0.94 per task. This price is lower than Claude Fable 5's ~$2.75 and Claude Opus 4.8's ~$1.80, close to GPT-5.6 Sol's $1.04, but significantly higher than GLM-5.2's $0.32–$0.47 and far above DeepSeek V4 Pro's $0.04.

Thus, K3's positioning is not as the lowest-cost option, but as a model that offers frontier-level capabilities at a more accessible price.

Investment Banks' Consensus: Demand is Not Weakened

Addressing market fears of a "DeepSeek Moment" impact, Duan Bing, an analyst with Nomura's Asia-Pacific Technology team, wrote in a report: "We believe competition and innovation in the global large model market will not cease. As we move closer to Artificial General Intelligence (AGI), the application of generative AI in both consumer and enterprise sectors will continue to expand. Frontier AI labs and hyperscale cloud platform companies are likely to continue investing during this phase to maintain their competitive positions—we interpret this competition, as scaling laws persist, as a positive for the AI infrastructure value chain."

Citi semiconductor analyst Peter Lee titled his July 19th report "Another Jevons Paradox." What is the Jevons Paradox? Simply put: increased efficiency of coal-powered steam engines led to greater coal consumption because more people could afford them and more use cases emerged. The same applies to AI models—when high-quality models become cheaper, developers and enterprises deploy more applications and process more tokens, ultimately driving up compute consumption.

Peter Lee believes that even if K3 sees widespread adoption, demand for general-purpose memory like server DDR5 and eSSD will still increase. The reason is that K3's inference efficiency is comparable to other frontier models, but its KV cache footprint expands with context length, increasing rather than decreasing memory pressure.

Bank of America Securities semiconductor analyst Vivek Arya was more direct in his July 17th report. He argues the response from U.S. frontier AI labs will be "not less compute, but more." If Chinese open-source models continue to close the gap, OpenAI, Anthropic, and Google must maintain differentiation through larger-scale training, heavier inference loads, and faster iteration. Arya also mentioned an often-overlooked context: media reports indicate Google's Gemini 3.5 Pro has been delayed by several months, with its programming performance failing to meet internal targets, making the defense of frontier leadership "increasingly difficult."

The UBS analyst team led by Timo Arcuri noted in a July 20th report that parallels between K3 and DeepSeek R1 do exist, but K3 is more about scale—being the world's largest open-source model with 2.8 trillion parameters and a 1M token context window. The analysts emphasized that open-source models typically consume more memory than closed-source frontier models due to longer context windows; even after quantization, KV cache demand continues to grow in absolute terms, making the deployment of open-source models more reliant on HBM and storage.

The Real Beneficiaries of This Competition

Storage: The most direct beneficiary. UBS calculations show the storage and memory sector's cumulative free cash flow (FCF) by 2028 is projected to reach about 30% of its market value, the highest proportion among all sub-sectors—with Micron Technology (MU) alone reaching 47%. Both Citi and Nomura maintain Buy ratings on Samsung Electronics, citing an extremely tight supply situation in the global memory market. Citi's Peter Lee points out that Kimi K3's inference-side memory demand is no less than other frontier models, and the expansion of KV cache volume will directly drive demand for server DDR5 and enterprise SSDs (eSSD). He particularly notes that large-scale deployment of Kimi K3 requires "super-node" cluster configurations with over 64 GPUs.

Compute Infrastructure: TSMC and NVIDIA stand to benefit first. Whether from the continued validity of scaling laws on the training side or increased token demand on the inference side, the outcome points to greater demand for advanced process chips. Nomura reaffirmed Buy ratings for TSMC, ASE Technology Holding, and MediaTek. NVIDIA has publicly stated that modern MoE model inference on its GB300 NVL72 platform offers up to 25x better performance-per-watt compared to the previous Hopper architecture. Models like K3 naturally benefit from NVIDIA's latest hardware.

Networking: The super-node trend creates structural opportunities. Kimi K3's need for super-node clusters, combined with China's domestic compute limitations due to high-end chip export controls—requiring reliance on super-node architectures to compensate for single-card performance gaps—drives demand for networking layer suppliers like optical modules and optical chips. Nomura is bullish on Zhongji Innolight and Suzhou Innolight.

Cloud Platforms: Benefiting from ecosystem aggregation effects. Cloud platforms hosting multiple frontier open-source models gain stronger pricing power, reducing dependence on any single closed-source model provider. Nomura sees Alibaba Group (BABA) as a core player in China's AI cloud ecosystem and is positive on data center operators like GDS Holdings and VNET Group, Inc. (VNET).

The Pace of Global Penetration for Chinese AI Models

This is perhaps the most underestimated data point in the entire narrative.

According to statistics from the open API gateway OpenRouter, the share of token usage by Chinese AI models in global developer traffic has grown from less than 2% a year ago to over 45% currently. Data from Bank of America Merrill Lynch corroborates the acceleration of overall AI penetration: approximately 55% of U.S. enterprises have now subscribed to AI models, platforms, or tools, with enterprise adoption rates reaching 42% for Anthropic and 40% for OpenAI. The top 1% of enterprise AI consumers are spending $4,833 per employee per month on AI.

The market is bifurcating. On one side are Chinese open-source models like DeepSeek and Kimi K3, covering the economy and mid-to-high-end value segments. On the other, top U.S. frontier models are focusing on more complex workloads (like scientific computing) to maintain technological and pricing premiums. Nomura's judgment is that leading large model players on both sides of the Pacific will benefit—provided they can each remain at the forefront of the technology curve.

What K3 Really Changes is the Competitive Pace

Following K3's release, the market's initial reaction was to compare it to DeepSeek R1. This comparison is useful, but one shouldn't stop at the question of "will AI hardware be sold off again?"

DeepSeek forced the market to re-evaluate training efficiency. K3 shows the market something else: open-source models can also push scale, long context, Agents, and multimodality into the frontier range.

This will pressure U.S. frontier labs to continue investing and allow Chinese models to further expand within the global developer ecosystem. Closed-source top-tier models retain their technology and price premiums, open-source models cover more price points and deployment scenarios, cloud vendors provide model distribution and enterprise implementation, and the hardware chain bears the pressure of training and inference.

In the short term, trading may experience volatility due to "DeepSeek memories." In the medium term, as long as token usage continues to grow and long-context and Agent applications keep spreading, compute, HBM, storage, networking, and IDC remain unavoidable cost items.

This is why multiple institutions have reached similar conclusions post-K3: stronger open-source models are not the end of AI infrastructure demand; instead, they may be the gateway to the next wave of demand diffusion.

However, Bank of America Merrill Lynch also clearly flagged a tail risk: "If the pace of efficiency gains outpaces the growth in workloads, we might see some pullback in infrastructure build-out." In other words, if models become increasingly cheaper but usage does not expand significantly in sync, the growth logic for compute demand could be undermined.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Most Discussed

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10