星球日报|Aug 19, 2026 11:35
[Zhiyuan Founder Tang Jie: Scaling Large Models Is Not Just About Adding Parameters, Future Competition Will Focus on Post-Training and Inference Capabilities]
Odaily Planet Daily News - Zhiyuan founder Tang Jie published an article on the X platform discussing thoughts on the Scaling Law for large models. He stated that the current understanding of improving model capabilities in the AI industry is shifting from 'expanding parameter size' to multidimensional scaling. Parameter count is no longer the sole metric for evaluating model capabilities; it must be assessed comprehensively alongside data scale, allocation of computational resources, and the actual operational scenarios of the model.
Tang Jie pointed out that early research once drove the industry to rapidly expand model parameter sizes. Kaplan et al.'s 2020 study suggested that the growth rate of model parameters should exceed the growth rate of data, which spurred the development of ultra-large-scale models like GPT-3, Gopher, and MT-NLG. However, Hoffmann et al.'s 2022 analysis of hundreds of models revealed that the optimal computational strategy is closer to 'approximately 20 training tokens per parameter,' suggesting that model parameters and data scale should grow synchronously. The previous pursuit of trillion-parameter models was, in fact, a collective 'deviation' experienced by the industry.
As model application scenarios evolve, inference costs are gradually becoming a significant component of lifecycle costs. Optimization goals have shifted from merely reducing training costs to enhancing long-term operational efficiency, leading to increased attention on the 'small model + more thorough training' approach. Tang Jie noted that sparse mixture-of-experts (MoE) architectures have further altered the logic of Scaling. In MoE models, the total parameter count determines how much knowledge the model can store, while activated parameters and effective depth more significantly influence the model's ability to perform complex reasoning tasks. For tasks like vulnerability detection that require long-chain reasoning, the capability does not stem from simply memorizing more information but rather from maintaining coherence in multi-step reasoning processes.
Recent research indicates that there is no universal answer to the optimal 'token-to-parameter ratio': memory-intensive tasks require more parameters, while reasoning-intensive tasks rely more on data and computational depth. Under fixed training data scales, blindly increasing total parameters may even weaken reasoning capabilities, whereas increasing the number of activated experts is more conducive to improving model performance.
Regarding Zhiyuan's latest model developments, Tang Jie revealed that GLM-5.3 represents an experiment targeting the direction of Scaling. This model shares the same foundational model, architecture, and total parameter and activated parameter scales as GLM-5.2 but underwent post-training optimization through one month of large-scale, long-cycle environment training and reinforcement learning (RL). The performance improvement did not stem from an increase in parameters but rather from extensions during the post-training phase.
He concluded that competition in large models has shifted from simply comparing parameter sizes to exploring 'multidimensional Scaling.' Future advancements in model capabilities will rely more on training strategies, inference depth, and continuous optimization of post-training capabilities.
Share To
HotFlash
APP
X
Telegram
CopyLink