Z.AI MaaS open platform has registered nearly 7 million API users, an increase of about 2 million compared to early July, including 23,000 enterprise clients. The developer product ZCode, which rivals OpenAI's Codex, surpassed one million users within its first month of launch.
This year, Z.AI's ARR has grown 15-fold. According to our information, the ARR growth rate was not impacted by the release of new models from Kimi and Alibaba. One investor suggested the ARR has reached $2 billion, but Z.AI officially denies this, with sources close to the company indicating the actual figure is likely higher. If calculated at the $2 billion benchmark, the average monthly spend per registered user on the Z.AI open platform would be approximately 160 RMB.
On July 21, market rumors circulated that Z.AI had built a 1 GW domestic AI computing infrastructure. According to our sources, over 50,000 domestic AI chips have been activated to meet growing inference demands. On July 31, Z.AI's long-limited Coding Plan became available for purchase after a significant price increase.
Computing power is constrained, but there is still significant room for Infra efficiency improvement
The rapid ARR growth is fueled by a surge in demand from coding scenarios and the strong pull of frontier models. For comparison, Z.AI's ARR was about $100 million before the GLM-5 model was released in early February 2026. After the GLM-5 launch, Z.AI's ARR doubled in a short period. According to Artificial Analysis, GLM-5 ranked fourth globally at its release, behind two of Claude's most advanced models at the time and OpenAI's GPT-5.2.
The model of driving revenue growth through the release of advanced models has proven successful in both China and the US. Anthropic's annualized revenue was $9 billion at the end of 2025, rapidly growing to $50 billion six months later (and exceeding $60 billion by July, according to SemiAnalysis' calculations). Beyond the strong demand from coding scenarios, Anthropic's continuous release of models surpassing OpenAI is considered a key factor. Z.AI is the first company to replicate this model in China, with similar stories subsequently playing out at Moonshot AI and Alibaba.
On July 16, Kimi released K3, causing its API to be overloaded and interrupted for nearly six hours the next day, with Kimi even suspending new user registrations. On August 3, Alibaba released Qwen3.8-Max with 2.4 trillion parameters, driving its Hong Kong-listed shares up 7% that day and bringing its market cap back to 2.4 trillion HKD, a rebound of over 40% from its June low.
However, before Anthropic successfully navigated the coding path, almost no one, including Anthropic itself, had fully anticipated the bottleneck in computing power supply. For Chinese large model companies, the computing power shortage is even more severe. Unable to rapidly expand physical hardware, engineers have shifted focus to improving the inference efficiency of existing models through algorithms, especially for long-context tasks—the core pain point of coding demand.
Z.AI was previously perceived as lacking experience in Infra, but with the first-mover advantage brought by GLM-5, it has rapidly filled this gap over the past six months. In July, Z.AI completed the acquisition of Zhongke Jiahe, with whom it had been co-located for several months. In a technical blog, Z.AI stated that an average Agent task requires stuffing over 70,000 tokens of context into the model, far exceeding daily chat content. By splitting KV caches, the research results from Z.AI and Zhongke Jiahe can boost the throughput of a single inference service by up to 132% as long-context tasks grow. In another solution called ZCube, a large model inference network architecture, Z.AI claims its research can increase the throughput of the entire production cluster by up to 15% while reducing the need for switches and optical modules in AI infrastructure by up to one-third.
In late May, Z.AI launched a high-speed version of its API, achieving 400 tokens per second, roughly 8 times faster than the standard API service. The company stated it "sets a new global speed record for API services from large model vendors." This achievement came from collaboration with the TileRT team, which is from the same group behind TileLang—the programming language from the widely circulated Liang Wenfeng speech, considered to have the potential to dismantle Nvidia's CUDA moat. Many researchers agree that there are numerous aspects of Infra worthy of optimization.
On June 9, the TileRT team published results from a collaboration with Xiaomi, pushing the speed of the MiMo-V2.5-Pro model, which has 1 trillion parameters, to 1000 tokens/s. With the explosion in coding demand, Z.AI's new model training is increasingly oriented towards long-context tasks, while also considering inference needs in the model architecture. In a blog post about the collaboration with Z.AI, TileRT listed the next phase as the recoupling of models, compilers, and hardware.
Several core architectural changes in version 5.2 are almost all designed around inference costs. The technical report disclosed includes IndexShare, which reduces the computational cost per token in million-token scenarios by 2.9 times by reusing the indexer of the sparse attention layer. It also includes an improved MTP draft layer, increasing the acceptance length by 20% and generating speed by about 20%. In the post-training phase, Z.AI re-adopted PPO (Proximal Policy Optimization), paired with an asynchronous method called SAO (Single-rollout Asynchronous Optimization), allowing each task to enter learning immediately upon completion. Where industry methods typically collapse around 160 steps, this method stably trained for 1,000 steps. The self-developed framework, slime, can merge a dozen expert models into one in about two days.
The technical report did not disclose the accompanying modifications at the inference engine level. The ultimate result of all these optimizations is a normalized chart: if the throughput of GLM-5.1 processing 32,000 tokens of context is taken as the baseline, GLM-5.2's throughput at a 200,000 token context length is 4.7 times, while GLM-5.1's is 2.77 times, a difference of nearly 70%.
This is also the common optimization direction for current model companies. Kimi's K3 is more radical in its model architecture, converting three-quarters of its attention layers to so-called "linear attention." Compared to Z.AI's method, this is a "lossy" change, but the benefits are immediate: cache size is significantly reduced, and throughput is correspondingly increased. One investor suggested that Infra layer optimization will be a key theme for AI investment in the next year or two. As model capabilities converge, the company that can run more tokens on the same cards will have a better story. This has led to rapid salary growth for Infra talent; a 1% efficiency improvement, when scaled to large computing deployments, results in astronomical cost savings.
API gross margins are continuously improving, making selling tokens a good business
A simple calculation can illustrate the potential for scaling efficiency in large models: Even the high-speed API's 400 tokens/s only uses about 40% of the theoretical hardware limit of an 8-card H200 server's memory bandwidth. One industry source estimated that each round of optimization for large model Infra can still yield 15% to 30% efficiency gains, and even more so for Z.AI, which was not originally known for its Infra optimization prowess. The commercial value of this is greater than the technical significance, making the sale of tokens a more attractive business: on the cost side, inference efficiency offers continuous room for optimization; on the revenue side, user demand is steadily growing. Simultaneously, coding projects are becoming larger and more complex, consuming an increasing number of tokens.
Z.AI recognized this as a good business earlier than others. It was the first company in China to champion "Model as a Service," and its MaaS open platform domain, bigmodel.cn, was registered before ChatGPT's launch in 2022. However, its misstep was not securing more computing power at a more favorable price earlier. Even without the most advanced models, merely having immediately usable computing power to sublease would be a good business, as Elon Musk has demonstrated. Currently, AI's penetration in productivity scenarios is not deep enough, making tokens a purely incremental market. According to our sources, when Kimi released K3 in mid-July, Z.AI's API user growth and ARR growth rate were largely unaffected. At this point, every model vendor's growth is not at the expense of another's decline, giving everyone the confidence to raise prices.
On July 31, Z.AI lifted purchase restrictions on the Coding Plan, introducing a points-based calculation system, but the most notable change was the price increase. The introductory price was 20 RMB per month, while the Lite version is now 118 RMB per month. The Coding Plan was originally an early strategy to attract users, and no one anticipated the explosion in token demand. Now, Z.AI must prioritize supply for its API, forcing price increases for individual users. Even so, based on a price ratio of approximately 7:2:1 for input, cache hit, and output, the blended price for GLM-5.2 is about 8.8 RMB per million tokens, significantly lower than comparable US models.
Different analytical institutions have estimated API gross margins for various model vendors. SemiAnalysis estimates Anthropic's API gross margin exceeds 80%. The Information's exclusive report indicated DeepSeek's API gross margin is around 70% to 80%. According to our sources, Z.AI's API gross margin under ideal conditions, running on its own computing infrastructure, can reach 50% to 60%. These gross margin calculations do not include the upfront investment in API services, which in the large model field is the major cost—massive computing procurement and large personnel expenses. In July, Z.AI announced a new round of Hong Kong-listed share placement, which occurred the day after cornerstone investors' lock-up period expired. Z.AI's share price rose instead of falling. According to a source close to Z.AI, the entire placement subscription was completed in an afternoon, ultimately raising over 30 billion HKD. According to Z.AI's announcement, 55% of the funds will be used for R&D, covering both computing power procurement and talent expansion.
Another prerequisite for maintaining static API gross margins is that the intelligence provided by model vendors is differentiated, with pricing power belonging to the leading model. However, many believe that open-sourcing directly impacts the pricing power of advanced models, thereby affecting vendor revenue. Consequently, many vendors are tightening their open-source policies. Z.AI still uses the most permissive MIT license for GLM-5.2, but more than one vendor is planning or implementing changing open-source authorizations into commercial contracts. Terms include releasing a "crippled" version as open source, with the full version available only via the official API; requiring companies with annual revenue exceeding $20 million to apply for separate authorization; and stipulating that third parties hosting the model cannot price below the official API, with some plans requiring a revenue share of up to 30%.
Others disagree. Open-weight models are free, theoretically deployable by anyone with a few servers, but third-party pricing cannot undercut the model vendor's margins. As mentioned, the model's structure, training process, and inference efficiency are increasingly intertwined. Third parties can only optimize outside the model, while the model vendor holds the deepest know-how on inference efficiency. Taking DeepSeek as an example, it is publicly known that many third-party deployments of DeepSeek models cannot match the officially claimed efficiency, and DeepSeek is already one of the most technically open model vendors.
A brief window for independent model vendors
The opportunity in large models appears to be returning to startups, a scenario that was hard to imagine just over six months ago. Kimi's valuation hovered around $3 billion, while the prospectuses of Z.AI and MiniMax struggled to articulate the rationality of their business models: one derived 85% of its revenue from project-based local deployment; the other relied on C-end subscriptions, with a loss of over 1 billion RMB on revenue of over 2 billion RMB. The C-end battlefield seemed already irrelevant to independent model vendors. ByteDance's Doubao attracted a national-level user base but was losing money daily in the tens of millions—a mobile internet tactic of sacrificing scale for free users, finding no profitable loop.
The window comes from coding. The coding wave has truly ignited the demand for per-token billing, fundamentally changing the calculation method. According to Z.AI's 2025 annual report, doubling the price led to an 8-fold increase in usage volume, and the price gap with overseas frontier models is gradually narrowing. Kimi is more aggressive than Z.AI, pricing its K3 model's API output at $15 per million tokens, the same as the standard pricing for Claude Sonnet 5. Another aspect of the window is the absence of major tech companies. For half a year, Baidu, Alibaba, and Tencent were unable to produce a first-tier coding model. In February, the narrative was still about the Spring Festival red envelope war, but Z.AI had decided to focus its research on coding as early as the spring of 2025. From this perspective, the window may even be widening. ByteDance has announced it will not engage in distillation, effectively forgoing the opportunity to catch up to frontier models in the short term.
But the window won't stay open forever. For independent model companies without historical baggage, the advantage lies in being faster to adopt new technologies. However, this is a race about speed, repeated over and over. A major tech company only needs to win once for the entire narrative to be rewritten. ByteDance has already produced SOTA results in video models, and their commitment to coding is now visibly firm. Furthermore, major tech companies have advantages in computing power procurement and labor costs. At the May earnings call, Alibaba's Wu Yongming stated that investment in computing centers would far exceed the previously committed 380 billion RMB over three years. Tencent's operating capital expenditure in the first quarter of 2026 grew 18% year-on-year and 84% quarter-on-quarter, "accelerating investment in server infrastructure." In contrast, independent model vendors remain constrained by computing power bottlenecks, sometimes even renting computing power from potential competitors to handle massive training and inference demands. To keep models in the first tier, such investment cannot stop for a single day.
More importantly, the coding business model itself is rapidly evolving. More companies are building so-called "internal routers," combining models of different price tiers and performance levels based on task complexity. Companies like Palantir are educating enterprise clients to use only the API and avoid the additional tools offered by model vendors—tools that are crucial for enhancing user stickiness. Palantir's logic is that models will eventually become commodities, while data and workflows are a company's most precious assets. Companies should not continuously feed valuable business data to model vendors, effectively funding an intelligent system that will eventually replace them. Therefore, the better approach is to keep data, business logic, and workflows in-house, with the model as a plug-and-play component. This is the logic of the B2B business: everything defers to efficiency. To maintain a premium in the enterprise service space, one must control technical standards and formats, control customer relationships, or provide a unique service. The business model of large model APIs appears to satisfy none of these conditions: the lead time for top models is shrinking, users migrate freely between platforms, and there are almost no barriers.
If the improvement of Scaling Law truly has no end, it implies the realization of Artificial Superintelligence (ASI). At that point, the discussion will no longer be about whose gross margin is a few points higher. But before ASI arrives, the only certainty seems to be the ever-faster refreshing of the list of top models, which appears to be the only coping strategy currently available for large model players. The window is still open, computing power remains scarce, and people still crave the emergence of new top-tier models—ideally ones that can match or surpass the leading US closed-source models. The next release is already on the schedule. Both Z.AI and DeepSeek are expected to release new models in August, intensifying the competition.