律动BlockBeats|8月 05, 2026 07:36
[Ant Ling-3.0-flash Officially Open-Sourced, FP8 Version Only 128GB]
According to monitoring by Beating, Ant Bailing (inclusionAI) has officially released the weights for Ling-3.0-flash, offering both the BF16 original version and the FP8 quantized version. Both versions are licensed under MIT and are now available on Hugging Face and ModelScope, deployable via SGLang or vLLM. The BF16 version weighs approximately 255GB, while the FP8 version is about 128GB, nearly halving the size. FP8 uses lower precision to store parameters, reducing storage and VRAM requirements. In the four tests listed by the official team, the maximum score difference between FP8 and BF16 is 1.57 points. Ling-3.0-flash has a total of 124 billion parameters, with only 5.1 billion activated per generation. The model supports a context length of up to 256,000 tokens and is primarily designed for agent tasks such as programming, search, deep research, and tool invocation. Official evaluations show that it matches or surpasses the trillion-parameter predecessor model Ring-2.6-1T on most benchmarks. [Original Link]
Share To
Timeline
HotFlash
APP
X
Telegram
CopyLink