During the 2026 Yunqi Conference on September 24, Huang Shan, Strategic Technology Director of Lenovo China's Infrastructure Business Group, stated that the benchmark for measuring AI computing power is transitioning from "the cost per Token" to "the value each Token generates." He noted that as AI integrates into production processes, enterprises are increasingly focusing on how many Tokens can be produced per unit of cost, and are further linking Token output to business outcomes.
According to Huang, the rapid adoption of AI Agents is reshaping the demand logic for AI computing power. On one hand, Agents are driving exponential growth in Token consumption. Previously, user queries followed by model responses resulted in linear Token usage; however, with AI now expected to execute multi-step tasks and self-verify like digital employees, Token consumption is no longer a simple multiple but escalates several-fold or even exponentially. On the other hand, the rise of Agent-driven scenarios is altering the current structure of AI computing power demand. Huang pointed out that the ratio of Token consumption between inference and training has already surpassed 6:4, with inference-side computing demand growing rapidly. While the industry previously focused on expanding training cluster scale, the Agent era now raises a new question for infrastructure development: how to continuously and efficiently complete large volumes of inference tasks.
In response to growing Token demand, the supply side is prioritizing efficient Token production. Huang used Lenovo's operations to break down the cost structure of Token generation, noting that infrastructure accounts for roughly one-quarter of costs, with electricity representing about 10%, while software-hardware collaboration influences over 50% of Token costs. This implies that Token production efficiency is determined not only by chip hardware, but also by computing, communication, memory access, storage, inference optimization, and overall system scheduling capabilities. Competition in AI computing power is therefore moving from single-point hardware strength to comprehensive system capabilities.
On the demand side, planning approaches are also evolving in parallel with supply-side costs. Huang believes that enterprise computing power construction should no longer be defined solely by IT departments, but rather driven by business needs. Companies must first clarify the specific problems AI aims to solve, embed AI capabilities into business workflows, and then reverse-engineer the required Token volume, service-level agreements, and computing power scale from those workflows. This approach ties infrastructure investment to actual business rhythms, avoiding the pitfall of blindly purchasing computing power before identifying use cases.
This fundamental shift is also influencing how enterprises deploy AI. Public clouds handle elasticity and scheduling, while private deployments address security and data sovereignty requirements. The dynamic combination of these two, forming a hybrid AI model, will become the inevitable choice for enterprise AI deployment in the future. Huang emphasized that infrastructure providers spanning chips, systems, hardware, and software stacks are unlikely to build a complete advantage through any single capability alone. The key lies in integrating these frontier technology elements into system-level capabilities tailored to real business needs.
"In the future, Lenovo hopes to continue using hybrid AI as the foundation, transforming computing power resources into sustainable and measurable intelligent capacity, empowering enterprises to move AI from technical pilots to large-scale production," Huang concluded.