PANews丨APP全面升级|Aug 20, 2026 06:22
Zhiyuan AI Founder Tang Jie: Trillion-parameter models are a detour for the industry; post-training is the next key step.
On August 19, Tsinghua professor and Zhiyuan AI founder Tang Jie posted that parameter scale is no longer the sole metric for evaluating a model's capability.
He pointed out that Kaplan and others' 2020 research led the industry to scale parameters faster than data at a 2.7:1 ratio, giving rise to models like GPT-3 and Gopher. However, the 2022 Chinchilla paper re-ran experiments and found the optimal ratio should be 20 tokens per parameter. "Looking back, the trillion-parameter route was a detour the entire field collectively took and then retraced."
Tang Jie used GLM-5.3 as a comparative experiment: with the same base, architecture, and 753B parameters as GLM-5.2, it achieved significant capability improvements solely through one month of large-scale, long-cycle environment training and reinforcement learning.
He concluded: "Scaling isn't just one knob. This time, we turned the post-training knob because it still has the most room for improvement. Next time, it might be mid-training or pre-training.
Share To
HotFlash
APP
X
Telegram
CopyLink