AIToday
Large Language ModelsAI Business & IndustryTHE DECODERPublished: Aug 7, 2026, 22:02 JST3 min read

Bytedance trains 10-trillion-parameter AI model, largest in China

Bytedance trains 10-trillion-parameter AI model, largest in China

Key takeaway

  • Bytedance is developing an AI model with up to ten trillion parameters, making it China's largest model and comparable in scale to Anthropic's Mythos 5.

  • The model is currently in pretraining, which typically lasts three to six months.

  • Bytedance founder Zhang Yiming has directed the company to pursue world-leading capabilities, and the company has independently trained the model without relying on distillation from competitors' outputs.

3 Key Points

  1. What happened

    Bytedance is training an AI model with up to ten trillion parameters, according to the Financial Times. That size is three times larger than Moonshot's Kimi K3, the current largest Chinese model, and places Bytedance in the same range as Anthropic's Mythos 5, estimated at around eight trillion parameters.

  2. Why it matters

    The model is still in pretraining, a phase that typically takes three to six months. Bytedance founder Zhang Yiming has told the company's 2,000-person Seed team to aim for world-leading model capabilities over the long term, signaling the company's ambition to compete at the frontier of AI development.

  3. What to watch

    One insider noted that Bytedance has avoided distillation—training on outputs from other companies' models—for over a year, suggesting the company is pursuing an independent path. xAI is also training Grok variants with six and ten trillion parameters on its Colossus 2 cluster, according to Elon Musk.

In Depth

Read the full story

According to the Financial Times, Bytedance is training an AI model with up to ten trillion parameters—a scale that would make it the largest model developed in China to date. The size is three times that of Moonshot's Kimi K3, which currently holds the title of largest Chinese model. Industry estimates place Anthropic's Mythos 5 at around eight trillion parameters, putting Bytedance's model in a comparable ballpark, though Anthropic has not publicly disclosed its own parameter counts.

Three insiders told the Financial Times that the model is in pretraining, a development phase that typically lasts three to six months. During pretraining, raw data is used to teach the model fundamental language patterns and knowledge. While the number of parameters determines how much information a model can store, actual performance also depends on the quality of the data and the training methods employed. One source revealed that Bytedance has deliberately avoided distillation—a technique where models are trained on outputs from other companies' systems—for over a year, suggesting an independent approach to model development.

Bytedance founder Zhang Yiming has internally directed the company's 2,000-person Seed team to pursue world-leading model capabilities as a long-term goal, underscoring the organization's ambitions in AI. Separately, xAI is also training Grok variants with six and ten trillion parameters using its Colossus 2 cluster, according to Elon Musk, indicating that multiple companies are pursuing models of similar scale.

Context & Analysis

Bytedance's ten-trillion-parameter model represents a significant escalation in China's AI development. The scale places it directly in competition with frontier systems from Western companies like Anthropic, narrowing what was once a clear gap between Chinese and Western AI capabilities. The model's current pretraining phase suggests Bytedance is still months away from releasing or fully evaluating it, meaning the competitive landscape could shift substantially once it emerges.

A notable aspect of Bytedance's approach is its avoidance of distillation for over a year. This suggests the company is building its model from foundational training rather than copying or refining outputs from competitors, which could indicate both confidence in its data and training infrastructure and a strategic choice to develop proprietary techniques. The involvement of founder Zhang Yiming in setting the company's AI ambitions underscores Bytedance's commitment to achieving world-leading capabilities, not merely matching existing systems.

FAQ

How long will the model take to finish training?
The model is in pretraining, a phase that typically takes three to six months.
How does Bytedance's model compare in size to other leading models?
Bytedance's model with up to ten trillion parameters is three times the size of Moonshot's Kimi K3, the current largest Chinese model, and is in the same range as Anthropic's Mythos 5, estimated at around eight trillion parameters.
Is Bytedance using outputs from other companies' models to train it?
No; one insider said Bytedance has avoided distillation—training on outputs from other companies' models—for over a year.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleCTO Circle: 350 engineering leaders share playbook for AI-native organizations

The AI news that matters, in one minute each morning.

Sign up free