
Bytedance is developing an AI model with up to ten trillion parameters, making it China's largest model and comparable in scale to Anthropic's Mythos 5.
The model is currently in pretraining, which typically lasts three to six months.
Bytedance founder Zhang Yiming has directed the company to pursue world-leading capabilities, and the company has independently trained the model without relying on distillation from competitors' outputs.
What happened
Bytedance is training an AI model with up to ten trillion parameters, according to the Financial Times. That size is three times larger than Moonshot's Kimi K3, the current largest Chinese model, and places Bytedance in the same range as Anthropic's Mythos 5, estimated at around eight trillion parameters.
Why it matters
The model is still in pretraining, a phase that typically takes three to six months. Bytedance founder Zhang Yiming has told the company's 2,000-person Seed team to aim for world-leading model capabilities over the long term, signaling the company's ambition to compete at the frontier of AI development.
What to watch
One insider noted that Bytedance has avoided distillation—training on outputs from other companies' models—for over a year, suggesting the company is pursuing an independent path. xAI is also training Grok variants with six and ten trillion parameters on its Colossus 2 cluster, according to Elon Musk.
According to the Financial Times, Bytedance is training an AI model with up to ten trillion parameters—a scale that would make it the largest model developed in China to date. The size is three times that of Moonshot's Kimi K3, which currently holds the title of largest Chinese model. Industry estimates place Anthropic's Mythos 5 at around eight trillion parameters, putting Bytedance's model in a comparable ballpark, though Anthropic has not publicly disclosed its own parameter counts.
Three insiders told the Financial Times that the model is in pretraining, a development phase that typically lasts three to six months. During pretraining, raw data is used to teach the model fundamental language patterns and knowledge. While the number of parameters determines how much information a model can store, actual performance also depends on the quality of the data and the training methods employed. One source revealed that Bytedance has deliberately avoided distillation—a technique where models are trained on outputs from other companies' systems—for over a year, suggesting an independent approach to model development.
Bytedance founder Zhang Yiming has internally directed the company's 2,000-person Seed team to pursue world-leading model capabilities as a long-term goal, underscoring the organization's ambitions in AI. Separately, xAI is also training Grok variants with six and ten trillion parameters using its Colossus 2 cluster, according to Elon Musk, indicating that multiple companies are pursuing models of similar scale.
Bytedance's ten-trillion-parameter model represents a significant escalation in China's AI development. The scale places it directly in competition with frontier systems from Western companies like Anthropic, narrowing what was once a clear gap between Chinese and Western AI capabilities. The model's current pretraining phase suggests Bytedance is still months away from releasing or fully evaluating it, meaning the competitive landscape could shift substantially once it emerges.
A notable aspect of Bytedance's approach is its avoidance of distillation for over a year. This suggests the company is building its model from foundational training rather than copying or refining outputs from competitors, which could indicate both confidence in its data and training infrastructure and a strategic choice to develop proprietary techniques. The involvement of founder Zhang Yiming in setting the company's AI ambitions underscores Bytedance's commitment to achieving world-leading capabilities, not merely matching existing systems.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
Amazon and Google are intensifying competitive efforts against The Trade Desk (TTD), a major digital advertisi…

QumulusAI announced a GPU-as-a-Service agreement with DRW, a global trading firm, to supply a dedicated NVIDIA…

OpenAI introduced Premium Seats for ChatGPT Business, priced at $125 per user per month ($100 with annual bill…

Computer scientists at University of Tübingen, Max Planck Institute, MATS Research, and Snyk discovered a meth…

Anthropic pledged to embed machine-readable watermarks in Claude-generated text and digitally signed provenanc…

OpenAI announced that its unreleased model Astra had produced solutions to 10 long-standing mathematics proble…

The AI news that matters, in one minute each morning.
Sign up free