深潮TechFlow
深潮TechFlow|Aug 26, 2026 14:27
[Zhipu Launches and Open-Sources 'Niulai' Model GLM-5.3-Flash, Enhancing Multimodal Capabilities and Inference Cost Efficiency] According to Deep Tide TechFlow, on August 26, Z.ai released GLM-5.3-Flash, calling it the first native multimodal model in the GLM-5 series. The model features 320 billion total parameters and 18 billion active parameters, surpassing GLM-5.2 in multiple coding, agent, and vision benchmarks, and approaching Claude Opus 4.8 in certain coding tasks. GLM-5.3-Flash adopts a hybrid architecture combining sparse attention and linear attention, and introduces mechanisms such as mHC and IndexPool to reduce long-context inference costs and KV cache overhead. Z.ai stated that the model has been deployed on a large-scale service using China's AI chip clusters, achieving a threefold improvement in end-to-end service performance compared to the initial baseline. The model weights have been made publicly available on Hugging Face.
+2
Mentioned
Share To

Timeline

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads