律动BlockBeats
律动BlockBeats|Aug 26, 2026 15:29
[NVIDIA's Moat Takes Another Hit: 'Niulai' Model GLM-5.3 Flash Processes 23 Trillion Tokens, All Powered by Domestic Chips] Beating AI Newsflash: Following the release of GLM-5.3 Flash, Zhipu has confirmed a previously undisclosed detail: during the anonymous testing phase of Ox Alpha, all inference computing power was provided by Chinese domestic AI chips. Within just six full calendar days of Ox Alpha going live on OpenRouter, it processed 23.2 trillion tokens—more than double that of DeepSeek-V4-Flash during the same period. OpenCode even previously claimed that the backend has the capacity to provide a daily free quota of 100 trillion tokens. Zhipu stated that they have optimized end-to-end inference performance on the same Chinese domestic hardware to three times its original efficiency. Now, hardware efficiency and per-token costs are approaching those of mainstream NVIDIA GPUs. GPAmericaemiAnalysis specifically highlighted this point. Previously, when the public saw the capability to supply 100 trillion tokens per day, there was speculation that such scale could only be supported by top-tier labs and NVIDIA GPUs. However, this time, the entire process was powered exclusively by Chinese domestic chips. [Original Article Link]
+3
Mentioned
Share To

Timeline

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads