ByteDance is training an AI model with up to 10 quintillion parameters, competing with Mythos 5 from Anthropic
According to sources familiar with the matter, ByteDance is training an AI model with up to 10 quintillion parameters — three times as many as Kimi K3 from Moonshot. It is in an early stage of pretraining; the estimated size of Mythos 5 from Anthropic is around 8 quintillion parameters.
According to three sources familiar with the matter, ByteDance is training an AI model with up to 10 quintillion parameters. The model is currently in an early stage of pretraining, which typically lasts three to six months before fine-tuning and a possible release. According to the sources, the final size of the model will only be determined at a later stage, so its release is not certain.
According to the sources, this would be a model three times larger than Kimi K3 from Moonshot, the largest Chinese model released to date. Anthropic does not disclose the size of its models, but industry estimates put Mythos 5 at around 8 quintillion parameters and Fable 5 at around 5 quintillion. The number of parameters determines the basic capacity of a model to store information, but its actual capabilities also depend on data quality and training methods.
The effort by ByteDance shows the ambition of Chinese labs to not only catch up with but also surpass American competitors at the most advanced level of AI. Models from Moonshot and Alibaba have achieved strong benchmark results in recent weeks and trail only Fable 5 in some areas. Mythos 5, the most advanced model from Anthropic, is available only to approved organizations following a temporary ban in June over safety concerns. According to industry sources, several Chinese labs are training models the size of Fable 5, with ByteDance placing the greatest emphasis among them on achieving the largest possible size.
Why it matters
Model size alone does not guarantee better performance, because data quality and training methods also matter — the number of quintillions of parameters therefore cannot be taken as a direct measure of capabilities. Nevertheless, the report shows that the gap between leading models from the US and China is narrowing at the benchmark level, which is relevant to companies considering AI models from different providers. In addition, the model from ByteDance is still only in an early stage of training, and neither its release nor its final size is certain.
Relevant practical impact
What this means
For a business
The narrowing gap between Chinese and American models on benchmarks may expand the range of AI models with comparable performance in the future, which is relevant to companies planning to select a provider of AI technologies.
StrategyCheck the original
Event sources
only one source so far · 1 publisher, 1 independent. We count feeds from the same owner only once.