AIToday
Large Language ModelsAI Business & IndustryTHE DECODERPublished: Aug 7, 2026, 22:02 JST

Bytedance trains 10-trillion-parameter AI model, largest in China

Bytedance trains 10-trillion-parameter AI model, largest in China

3 Key Points

  1. What happened

    Bytedance is training an AI model with up to ten trillion parameters, according to the Financial Times. That size is three times larger than Moonshot's Kimi K3, the current largest Chinese model, and places Bytedance in the same range as Anthropic's Mythos 5, estimated at around eight trillion parameters.

  2. Why it matters

    The model is still in pretraining, a phase that typically takes three to six months. Bytedance founder Zhang Yiming has told the company's 2,000-person Seed team to aim for world-leading model capabilities over the long term, signaling the company's ambition to compete at the frontier of AI development.

  3. What to watch

    One insider noted that Bytedance has avoided distillation—training on outputs from other companies' models—for over a year, suggesting the company is pursuing an independent path. xAI is also training Grok variants with six and ten trillion parameters on its Colossus 2 cluster, according to Elon Musk.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

Bytedance's ten-trillion-parameter model represents a significant escalation in China's AI development. The scale places it directly in competition with frontier systems from Western companies like Anthropic, narrowing what was once a clear gap between Chinese and Western AI capabilities. The model's current pretraining phase suggests Bytedance is still months away from releasing or fully evaluating it, meaning the competitive landscape could shift substantially once it emerges.

A notable aspect of Bytedance's approach is its avoidance of distillation for over a year. This suggests the company is building its model from foundational training rather than copying or refining outputs from competitors, which could indicate both confidence in its data and training infrastructure and a strategic choice to develop proprietary techniques. The involvement of founder Zhang Yiming in setting the company's AI ambitions underscores Bytedance's commitment to achieving world-leading capabilities, not merely matching existing systems.

FAQ
How long will the model take to finish training?
The model is in pretraining, a phase that typically takes three to six months.
How does Bytedance's model compare in size to other leading models?
Bytedance's model with up to ten trillion parameters is three times the size of Moonshot's Kimi K3, the current largest Chinese model, and is in the same range as Anthropic's Mythos 5, estimated at around eight trillion parameters.
Is Bytedance using outputs from other companies' models to train it?
No; one insider said Bytedance has avoided distillation—training on outputs from other companies' models—for over a year.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Agentic AI pushes identity to front of data security, Oracle saysSiliconANGLE AI · 54m ago
  • Momentic launches Mo, an AI agent that tests apps without scriptsSiliconANGLE AI · 54m ago
  • Anthropic debuts Claude Sonnet 5.5, 30% fasterSiliconANGLE AI · 54m ago

AI-summarized, only the topics you pick: one digest a day via Email, LINE, or Slack.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleCTO Circle: 350 engineering leaders share playbook for AI-native organizations