ByteDance Considers Building a Model with Over 5 Trillion Parameters, Zhang Yiming: No Distillation, Don't Be Distracted by Short-Term Coding Trends

Deep News
Aug 06

ByteDance is reportedly exploring the development of a model with more than 5 trillion parameters, surpassing the scale of Alibaba's Qwen 3.8-Max (2.4 trillion parameters) and Moonshot AI's K3 (2.8 trillion parameters). This would make it the largest known model in China in terms of parameter count. Generally, larger models tend to exhibit higher intelligence, though the plan is still in its early stages and does not guarantee a final release. The new model will be led by Xiang Liang, head of the Seed Foundation, in collaboration with Shen Ke, who oversees LLM pre-training data. To support this, Seed is reorganizing its structure, clarifying responsibilities, and reallocating resources. Both Xiang Liang and Shen Ke are core members of ByteDance's AI and large model R&D system, both originating from the company's "Search, Advertising, and Recommendation" framework. Xiang Liang, who joined ByteDance in 2016 after earning a bachelor's degree from USTC and a PhD from the Institute of Automation, Chinese Academy of Sciences, previously led the AML team before moving to lead the Doubao Large Model Foundation team. Shen Ke, a 2018 Tsinghua graduate, now primarily manages LLM pre-training data.

Since ByteDance entered the large model space, Seed has been viewed as one of China's most promising teams. Its success is largely credited with Doubao achieving over 200 million daily active users. However, questions have emerged this year. In mid-February, Seed 2.0, a key model launched after Wu Yonghui took over Seed, received a lukewarm market response. Around the same time, Zhipu's open-source GLM-5 was hailed as the first domestic model comparable to the Anthropic Opus series, and in July, Moonshot AI's Kimi K3 was deemed by multiple third-party evaluations to be nearing the performance of overseas closed-source flagship models. Multiple ByteDance sources noted that many Seed teams have been reflecting on these developments over the first half of the year. This is a key reason behind the discussion of training a model with over 5 trillion parameters, as the company considers bypassing incremental improvements on existing sizes to leapfrog competitors. "It's like a gamble," said a source close to Seed.

Two weeks ago, ByteDance founder Zhang Yiming and Seed head Wu Yonghui held a Seed all-hands meeting. Zhang, who has been less visible recently, expressed his support for Seed. He reassured the team that training large models is inherently difficult and urged them not to be overly anxious. He stated that the company could accept the model lagging behind the industry for a period, setting a goal of aiming for top-tier global intelligence. While he acknowledged that coding is a current key direction, warranting the integration of resources from Volcano Engine, Feishu, and Doubao, he cautioned against being solely driven by this short-term trend, noting larger opportunities exist beyond coding. He praised Seedance for its market uniqueness and technological edge, encouraging Seed to build models with distinctive features rather than just following leaders. Zhang also expressed his opposition to model distillation, arguing that while it improves short-term performance, it essentially replicates existing capabilities, limiting the potential for true breakthroughs. He urged Seed to build AGI barriers from a more fundamental level. Over the past few years, ByteDance has heavily invested in AI and plans to commit even more resources in the future.

Seed's performance over the past six months showed multimodal leadership but lagged in language models. In the past half-year, Seed's multimodal models have achieved significant success. The video generation model Seedance 2.0 and image generation model Seedream 5.0 Lite, launched in February, quickly garnered attention, surpassing other similar models in the industry. Seedance 2.0 has become a cornerstone of Volcano Engine's MaaS business revenue. In contrast, the language model Seed 2.0, released simultaneously, had a limited market impact. While it was initially praised internally for handling the surge in consumer demand during the Spring Festival, the competitive landscape shifted when Anthropic's Opus series, with its coding capabilities, rapidly captured the developer and B2B market. Chinese tech giants then realized they had missed the coding window. Competitors like GLM-5 and Kimi K2.5 now have a clear lead in coding skills. This lag in model capability has affected revenue. Volcano Engine, currently China's largest model API seller with about half the market share, had a 2025 revenue of around 15 billion yuan and an internal target of over 40 billion yuan for this year. However, concerns exist. Token consumption from Doubao shows that over half of the tokens consumed come from the Seedance and Seedream multimodal series, with the language model's share being low. Since the second quarter, growth in Seedance's token consumption and revenue has slowed due to the peaking of the short drama industry and broader economic changes. Doubao's token consumption was 120 trillion in March and 180 trillion in June, below the planned 250-300 trillion target. Meanwhile, Zhipu and Moonshot have each surpassed annual recurring revenue of $1 billion and $300 million, respectively, driven by improved coding capabilities.

To address the coding gap, Zhang Yiming personally invited DeepSeek core researcher Guo Daya to join Seed, focusing on specialized coding training. ByteDance has consolidated its coding-related R&D resources under Guo's leadership. Alibaba's Tongyi Qwen team has also invested more resources in coding training. While training relies on data, and large companies can leverage internal business lines for feedback, Seed has debated the feasibility of distillation, with ByteDance leadership expressing concerns about its long-term harm. At the Seed all-hands meeting, Zhang Yiming made clear that even if avoiding distillation means temporarily lagging behind domestic competitors, the company will not take this shortcut. At a company-wide meeting on August 6, Doubao head Zhao Qi stated that the market potential for large models is immense, so economic returns are not an immediate worry, but revenue remains important. ByteDance's confidence is underpinned by its ample computing resources, industry-leading infrastructure, and strong talent pool. The success of Seedance has shown that once model capabilities are leading, they can quickly become a commercial moat. A ByteDance source noted that Seedance proved users naturally gravitate to the best-performing products, reducing the need for anxiety over short-term ranking setbacks.

ByteDance continues to push boundaries with heavy investment. ByteDance has long championed a "heavy investment yields miracles" approach: entering a domain by launching multiple teams, allocating substantial resources, and advancing several projects simultaneously. In 2012, ByteDance launched 13 apps, with Joke and Toutiao eventually gaining massive traction. In 2015, it launched over 20 apps for overseas markets, and in 2016, it used a three-pronged strategy with Volcano, Douyin, and Xigua for short videos. This led to the post-Toutiao growth curve of Douyin and TikTok. This model was repeated in gaming and education. In AI, ByteDance has been the most committed, with the largest capital expenditure, the most GPUs, and the most top-tier talent, exploring every business direction. In 2025, the mainstream view on video generation was that scaling pre-training along the DiT architecture had limited upside, leading many teams to focus on post-training. Kuaishou's Kling, for example, attempted but didn't fully commit to a larger model. Seedance took the opposite approach, investing heavily in pre-training. After nearly a year of optimization, Seedance 2.0 became the first video model to fully adopt the MoE architecture, with 200 billion parameters. Launched in February 2026, it was quickly recognized as the world's most powerful video model. Today, scaling model size is industry consensus: Moonshot launched the 2.8-trillion-parameter Kimi K3 in July; xAI expanded Grok from 500 billion to 1.5 trillion parameters, planning a next-generation model with 6 trillion; and OpenAI and Anthropic are also reportedly exploring larger models. ByteDance, lagging in language models, is betting on a similar approach to catch up. This is challenging, as video models require fewer resources, like lower demands on inference chips and cheaper research talent, with Seedance's core algorithm team having only about ten members. A larger language model, however, requires innovations in algorithms, handling of different-scale system engineering, more diverse data, and deep cross-departmental collaboration. Changes are underway. ByteDance leadership is pushing to eliminate internal competition within Seed, consolidate resources, and clearly define roles. Seed is breaking down departmental barriers to foster collaboration. A former Seed employee noted, "ByteDance's culture may not be ideal for innovation, but it excels at problem-solving." Seedance solved the first problem, and ByteDance now aims to apply the same approach to the next.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Most Discussed

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10