律动BlockBeats|Aug 26, 2026 17:15
[NVIDIA: Qwen3.8-Flash-Next Achieves Over 16,000 Tokens/Second Single-Card Throughput on GB300 NVL72]
Beating AI Newsflash: NVIDIA has published an article stating that Alibaba's latest preview model, Qwen3.8-Flash-Next, is now supported on the NVIDIA GB300 NVL72 platform. The model features a total parameter scale of 176 billion, with approximately 6 billion parameters activated per token. It natively supports a 262,000-token context, which can be extended to 1 million tokens via YaRN, and is primarily designed for long-context agent applications such as intelligent programming, document processing, and tool invocation.
NVIDIA noted that Qwen3.8-Flash-Next adopts a hybrid architecture combining Gated DeltaNet (GDN) and Qwen Sparse Attention (QSA) to reduce computational and KV cache overhead in long-context scenarios. Testing shows that on the GB300 NVL72, the model achieves a single GPU throughput of over 16,000 tokens per second and a single-user throughput of over 200 tokens per second. It also supports inference frameworks such as SGLang, vLLM, and TensorRT-LLM.
Share To
Timeline
HotFlash
APP
X
Telegram
CopyLink