AI Cloud Competition Shifts to Unit Task Cost, Full-Stack and Scale Advantages to Define Long-Term Winners

Stock News
7 hours ago

A recent research report from Zheshang Securities Co.,Ltd. highlights that as large model inference scales continue to expand and technology iterates rapidly, the "unit task completion cost" is set to become the core competitive battleground in the AI cloud business. The firm identifies full-stack capabilities and scale advantages as the two primary pillars of competitiveness in this sector. They maintain a positive outlook on long-term investment opportunities driven by robust AI cloud demand, emphasizing the strategic importance of these two factors.

As the industry matures, the competitive focus in AI cloud will increasingly center on unit task cost efficiency. Both Hyperscaler and Neocloud providers are leveraging their distinct infrastructure strengths to continuously improve cluster utilization, reduce per-token costs, and amplify the value and profitability of their computing assets. The report outlines several key perspectives on how these dynamics will shape the market.

Core competitiveness in cloud business: Full-stack and scale advantages

Full-stack technology maximizes the efficiency of every unit of computing power, while scale advantages boost cluster utilization and minimize idle capacity. Large CSPs are expected to maintain industry leadership by capitalizing on horizontal scale benefits and vertical full-stack integration capabilities. Meanwhile, Neocloud providers can build differentiated competitive moats through specialized pure-AI cluster architectures and efficiency gains from homogenous traffic in third-party model inference.

Full-stack advantage part 1: Moving from local optimization to system-wide optimization

Insufficient GPU utilization in computing clusters often stems from a pronounced "short-board effect." By coordinating the entire pipeline for global optimization, companies can achieve results unattainable through localized tweaks, reducing losses at each stage and enhancing overall system efficiency. Examples such as Google's full-stack approach and the collaborative tuning between Nvidia and Nebius demonstrate the performance leaps possible through holistic optimization.

Full-stack advantage part 2: Delivering higher added value to customers

AI cloud services can be categorized into four tiers from bottom to top: physical infrastructure, managed infrastructure, managed inference services, and agent optimization layers. Full-stack capabilities drive AI infrastructure upgrades toward higher-value-added segments. This approach not only improves resource utilization through global technical optimization but also helps vendors diversify their customer base and reduce concentration risks. A full-stack comparison between Nebius and CoreWeave, both leading emerging AI cloud providers in the US and Europe with deep ties to Nvidia, reveals divergent paths. Both focus on AI computing workloads, but their trajectories differ significantly: CoreWeave primarily operates at the second tier (managed multi-tenant infrastructure), while Nebius sits at the third tier (managed inference with its TokenFactory), boasting lower customer concentration and a more complete full-stack technology suite.

Scale advantage part 1: Leveraging the law of large numbers to smooth traffic fluctuations and boost cluster utilization

As computing power and business traffic scale up, massive inference volumes can smooth out tidal fluctuations in workload, raising GPU cluster utilization and lowering marginal inference costs. Large-scale clusters do not eliminate traffic volatility but rather hedge against it through the law of large numbers. Techniques such as GPU sharding, continuous batching, Prefill/Decode decoupling, and resource reuse all rely on large-scale task volumes to fully unlock performance benefits—their impact is limited in small-scale traffic scenarios.

Scale advantage part 2: Declining total cost of ownership

First, larger vendors reduce hardware procurement costs through bargaining power and ODM direct-sourcing models. Second, they secure long-term power supplies, and as computing scale and utilization grow, fixed energy consumption for power delivery and idle cooling is amortized, pushing down PUE. Third, fixed costs such as team overhead are spread across a larger revenue base. From 2023 to 2025, Google Cloud's personnel costs as a percentage of revenue steadily declined, leading to improved profit margins.

Understanding the competitive strengths of Hyperscalers vs. Neocloud providers

1. Core strengths of CSPs: Horizontal scale advantages—by 2025, the top three CSPs each generate tens of billions in cloud revenue, with AWS and Azure reaching the hundred-billion-dollar mark, while Neocloud players like Nebius and CoreWeave report revenue of only $530 million and $5.13 billion, respectively. Vertical full-stack integration—major cloud vendors are actively developing in-house chips to bolster their complete-stack capabilities and fortify competitive barriers.

2. Differentiated strengths of Neocloud: Although CSPs dominate the AI cloud market share, Neocloud providers such as Nebius and CoreWeave achieved impressive growth in 2025, with revenue surges of 351% and 168%, respectively. This high growth stems from their agile organizational structures and the mutual checks and balances among various industry players, in a market context where computing supply still lags behind demand. Looking ahead, Neocloud providers can build lasting differentiation by combining the low total cost of ownership from pure-AI clusters with efficiency gains from homogenous third-party inference traffic.

Risk warnings

Key risks include slower-than-expected AI technology iteration, fluctuations in the supply-demand balance for computing power, intensifying industry competition, volatility in customer demand, and policy and compliance uncertainties.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Most Discussed

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10