律动BlockBeats
律动BlockBeats|Jul 27, 2026 16:21
Why did Kimi K3 approach Fable 5? The Dark Side of the Moon simultaneously rewrites attention, residuals, and MoE According to Beating monitoring, the dark side of the moon has released the Kimi K3 technical report. K3 has a total of 2.8 trillion parameters, with each token activating 104 billion parameters. Its base architecture has undergone a comprehensive upgrade compared to K2. The increase in parameters is only a part of it. According to the Dark Side of the Moon, the new architecture, data, and training methods have increased the overall scaling efficiency of K3 by about 2.5 times compared to K2. According to its Scaling Law curve, the training computation required to achieve the same validation loss is approximately 40% of K2. K3 first rewrote attention. It uses KDA to process long sequences and adds another layer of global MLA every three layers. KDA compresses the previous content into a fixed size state, reducing the computational and caching pressure of millions of Token contexts; MLA is responsible for preserving global information exchange. It also includes Attention Residents. The traditional model can only continuously stack the information of all previous layers, while K3 allows each layer to directly select the required content from earlier network blocks. After the model deepens, early information is not easily diluted. MoE has also been redesigned. K3 has 896 routing experts, with 16 activated per token, which is twice as many as K2. The routing expert will first calculate in the compressed space and then return to the model backbone. The new activation function and load balancing algorithm are used to prevent the training of super large models from losing control. What truly enhances the Agent's capabilities is post training. Train the general, agent, and code models separately on the dark side of the moon, then train three different levels of thinking intensity for each direction, and finally merge the nine experts. The training task can contain up to thousands of tool calls and continuously retain files, applications, and virtual machine states. K3 has approached or surpassed top closed source models such as Fable 5 and GPT-5.6 Sol in multiple code, search, and tool call tests. Its leap is not solely based on heap parameters, but has pushed the limits of larger bases, new architectures, and million token agent reinforcement learning. [Original link]
+3
Mentioned
Share To

Timeline

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads