PANews丨APP全面升级
PANews丨APP全面升级|Aug 20, 2026 06:22
Zhiyuan AI Founder Tang Jie: Trillion-parameter models are a detour for the industry; post-training is the next key step. On August 19, Tsinghua professor and Zhiyuan AI founder Tang Jie posted that parameter scale is no longer the sole metric for evaluating a model's capability. He pointed out that Kaplan and others' 2020 research led the industry to scale parameters faster than data at a 2.7:1 ratio, giving rise to models like GPT-3 and Gopher. However, the 2022 Chinchilla paper re-ran experiments and found the optimal ratio should be 20 tokens per parameter. "Looking back, the trillion-parameter route was a detour the entire field collectively took and then retraced." Tang Jie used GLM-5.3 as a comparative experiment: with the same base, architecture, and 753B parameters as GLM-5.2, it achieved significant capability improvements solely through one month of large-scale, long-cycle environment training and reinforcement learning. He concluded: "Scaling isn't just one knob. This time, we turned the post-training knob because it still has the most room for improvement. Next time, it might be mid-training or pre-training.
Share To

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads