律动BlockBeats
律动BlockBeats|Aug 11, 2026 08:45
[Ant Group Open-Sources Ling-3.0-tiny: 7.9 Billion Parameters, Runs at 90 Tokens/s Locally on M4 Pro] According to monitoring by Beating, Ant Group's Ling-3.0-tiny weights have been officially released, offering BF16, FP8, and INT4 versions. The Hugging Face page lists the model under an MIT license. The model has a total of 7.9 billion parameters, with only 1.3 billion activated per token, primarily designed for local deployment and agent scenarios. Ling-3.0-tiny features 128 routing experts, with each token selecting only 8 of them, plus 1 shared expert used by all tokens. The attention mechanism adopts a 3:1 KDA and MLA hybrid architecture, meaning every 3 layers of Kimi Delta Attention are followed by 1 layer of MLA. The official statement claims the FP8 version achieves approximately 86–90 tokens/s on an M4 Pro MacBook and around 100–105 tokens/s on a DGX Spark. [Original Link]
+4
Mentioned
Share To

Timeline

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads