Z.AI Founder Examines Scaling Law Evolution: Trillion-Parameter Models Were an Industry-Wide Detour, Post-Training Now Key

Deep News
08/19

Tang Jie, founder of Z.AI, shared his latest perspective on Scaling Law in a post on platform X on the 19th, using GLM-5.3 as a "controlled variable experiment" to support his reasoning. The post directly addresses external questions about why the company did not train a completely new foundation model.

In his post, Tang emphasized that parameter count has never been the only metric for evaluating model capability; data scale, compute allocation, and inference deployment conditions are equally indispensable. He noted that GLM-5.3 shares the identical foundation, architecture, 753B total parameters, and 40B activated parameters with its predecessor GLM-5.2, with the sole difference being an additional month of long-horizon environmental training and reinforcement learning (RL) post-training. The resulting capability improvement, he said, is "far from trivial." On the AA evaluation index, GLM-5.3 scored 60 points, a 7-point increase from GLM-5.2's 53 points.

The core takeaway from this experiment: when a foundation model's parameter scale already holds sufficient knowledge, additional capability gains typically come from "turning other dials" — particularly the post-training phase — rather than continuing to stack more parameters. This assessment offers direct reference value for how the broader AI industry allocates training resources today.

Trillion Parameters: A Detour the Industry Once Took Together

Tang traced the evolution of Scaling Law in his post. In 2020, Kaplan and colleagues concluded that parameter growth should outpace data growth at a ratio of approximately 2.7:1. This finding spurred the concentrated emergence of ultra-large models like GPT-3, Gopher, and MT-NLG, with trillion-parameter scale briefly viewed as an inevitable milestone on the path of industry progress.

However, in 2022, Hoffmann and colleagues re-ran experiments across more than 400 models and found that compute-optimal allocation should approach roughly 20 tokens per parameter, with parameters and data growing at roughly the same rate as compute expands — rather than the gap continuously widening. This is the widely cited "Chinchilla Scaling Law."

Tang stated bluntly that early fitting errors were magnified as compute scale grew, leaving the largest models of that era precisely the ones with the most imbalanced compute allocation:

"Looking back, the trillion-parameter round was a period where the entire field walked a detour together, and then collectively turned back."

Inference Cost and MoE: The Optimal Solution Keeps Moving

Even after Chinchilla, the optimal solution did not remain fixed. Tang pointed out that Chinchilla's optimization targeted a model that is "trained once and evaluated once," whereas today's mainstream models may be called billions of times daily, making inference costs far exceed one-time training costs over the full lifecycle.

When inference cost enters the objective function, the optimal solution shifts toward "smaller models trained longer." He cited Llama-2-7B and Gemma-2-9B as examples, with roughly 290 and 889 tokens per parameter respectively — both deliberately "overtrained."

The rise of Mixture-of-Experts (MoE) architecture further complicates the picture. Tang distinguished two dimensions that must be discussed separately in the MoE context: total parameters determine how much knowledge, factual information, and long-tail data a model can store; while activated parameters and effective depth per forward pass determine long-range reasoning capability.

He cited a 2025 study by Roberts and colleagues finding that the optimal tokens-per-parameter ratio varies by task — memory-oriented tasks favor more parameters, while reasoning-oriented tasks favor more data. Furthermore, under a fixed tokens-per-parameter ratio, increasing total parameters degrades reasoning ability, while activating more experts steadily improves reasoning performance. Tang concluded that capabilities like vulnerability identification, which require twenty-step causal chain reasoning, do not reside in total parameter count.

GLM-5.3: A Deliberate Controlled Experiment

Based on the above analysis, Tang positioned GLM-5.3 as an intentional "controlled variable experiment." Compared to GLM-5.2, GLM-5.3 keeps the foundation, architecture, total parameters, and activated parameters completely unchanged, with the sole variable being an additional month of long-horizon environmental scaling and reinforcement learning post-training.

He acknowledged that post-training was chosen as the "dial" because it currently offers the largest room for improvement — not because other dimensions have been exhausted. "This does not mean other dials have been turned to their limit," he clarified. Foundation model scale, pre-training data, and compute investment per forward pass all remain under consideration, with potential future adjustments involving mid-training and pre-training phases.

This approach directly responds to earlier external skepticism about why Z.AI did not train a new foundation model. In Tang's framework, dials need not be adjusted simultaneously; the key is identifying "which one is most worth turning next." At present, GLM-5.3's coding capability has been rated by multiple developers as the current best among domestic models.

免責聲明:投資有風險,本文並非投資建議,以上內容不應被視為任何金融產品的購買或出售要約、建議或邀請,作者或其他用戶的任何相關討論、評論或帖子也不應被視為此類內容。本文僅供一般參考,不考慮您的個人投資目標、財務狀況或需求。TTM對信息的準確性和完整性不承擔任何責任或保證,投資者應自行研究並在投資前尋求專業建議。

熱議股票

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10