ByteDance discusses training a model with over 5 trillion parameters, scale approaching twice that of Kimi K3
2026-08-07 13:21:12
According to CoinMeta, ByteDance is discussing training a large language model with over 5 trillion parameters, in an attempt to catch up with leading competitors both domestically and internationally by expanding the scale of the model. If the plan is implemented, this model will surpass Alibaba’s 2.4 trillion parameters and MoZhiDanMian’s 2.8 trillion parameters, becoming the largest model known in China in terms of parameter size. However, the project is still in the early stages of discussion, and it is not yet determined whether the model will actually be trained and released. The new model is planned to be led by Xiang Liang, the person in charge of Seed Foundation, in collaboration with Shen Ke, who is responsible for the pre-training data of large language models. ByteDance has already begun pre-training a large model with a maximum of 10 trillion parameters, which, if trained to its full capacity, would exceed Kimi K3 by more than three times, making it the largest model known among Chinese teams so far. This model is currently in the pre-training phase, which usually takes 3 to 6 months, and further post-training will be required before it can be successfully released. Recently, Zhang Yiming has explicitly opposed distilling competitor models within Seed, urging the team to accept a short-term lag and rely on their own research to strive for a place in the world’s top tier.
Bullish 0
Bearish 0
Source:Internet
This content is for market information only and does not constitute investment advice.