Original author: Li Jia
Original source: Wall Street Insights
The large model Kimi K3 from the Dark Side of the Moon adopts a linear attention mechanism, raising concerns in the market about a potential weakening of the demand for Nvidia, HBM, and network devices.
However, the semiconductor research organization SemiAnalysis recently provided a completely different assessment: The massive parameter scale and inference architecture requirements of K3 will not weaken the demand for high-end AI hardware; rather, they may further strengthen the demand for Nvidia's high-end GPUs, HBM, and high-speed interconnect devices.
SemiAnalysis pointed out that K3 has a parameter scale exceeding 2.8 trillion, with model weight capacity exceeding 1.5TB of HBM. Even in relatively limited user concurrency scenarios, KV caching still requires a large amount of unloading to CPU DDR5 memory and NVMe storage, meaning there will not be significant surplus in HBM space.
More importantly, the Dark Side of the Moon previously revealed that efficient inference deployment for K3 requires a large-scale expansion domain architecture made up of at least 64 chips. This hardware requirement is highly compatible with the design direction of Nvidia's GB200/GB300 NVL72 and similar rack-level AI systems.
SemiAnalysis believes that the market's previous understanding of linear attention as "weakening GPU demand" is flawed. The real impact may be exactly the opposite: more efficient model architectures reduce AI inference costs, which will drive more applications to materialize, thereby stimulating the long-term demand for GPUs, HBM, DRAM, and network infrastructure.

Unafraid of linear attention iteration, demand for Nvidia chips remains robust
Market concerns mainly stem from the Kimi Delta Attention (KDA) mechanism used in Kimi K3.
Compared to traditional Transformer attention mechanisms, KDA can significantly reduce the data transmission requirements of KV caching, with potential reductions of up to about 10 times the network bandwidth pressure. This change has led some investors to recall the concerns about AI hardware demand following the release of DeepSeek R1, believing that improvements in model efficiency may decrease reliance on high-end computing hardware.
However, SemiAnalysis believes that this judgment overlooks another core demand in large model inference—the computational and interconnection pressure caused by parameter scale.
K3 has over 2.8 trillion parameters, and its model weights themselves require a large-scale distributed computing system for deployment. Meanwhile, K3 adopts the WideEP (Wide Expert Parallelism) optimization strategy, distributing 896 expert modules across multiple GPUs, so that a single GPU only needs to carry a part of the expert weights, thereby improving computational utilization.
However, WideEP also presents new challenges: the frequent data exchange among experts requires more powerful network interconnect capabilities. SemiAnalysis points out that the copper backplane interconnection architecture used in GB200/GB300 NVL72 provides an intra-rack bandwidth 18 times that of traditional DGX B200 systems, making it very suitable for such large-scale expert parallel inference tasks.
In other words, the KV caching communication demand saved by KDA may be partially offset by the weight exchange requirements introduced by WideEP, and the overall pressure on AI infrastructure has not significantly decreased.


The 64-chip expansion domain is not exclusively beneficial to Nvidia
However, there are differing opinions in the market. Insider GDP (@bookwormengr) pointed out that the "64-chip expansion domain" mentioned by the Dark Side of the Moon does not necessarily imply Nvidia's NVL72 solution. Huawei's Ascend 950 SuperPod also uses a 64-chip configuration and features a unified bus (UB) memory expansion capability similar to NVLink.
From an architectural capability perspective, the Ascend 950 SuperPod can support expansion to 1024 NPUs across 16 racks, making it competitive in meeting large-scale model inference demands as well. Therefore, the hardware demand increase brought by K3 does not mean that Nvidia will be the only beneficiary, but the trend of strengthening demand for high-end AI interconnect systems remains clear.

Jevons Paradox: AI efficiency improvements may drive hardware demand growth
SemiAnalysis further invokes Jevons Paradox to explain trends in AI infrastructure.
The theory suggests that when a technology improves resource utilization efficiency and reduces unit costs, demand often does not decrease; instead, it may increase due to an expanded application range. Applied to the AI field, after the linear attention mechanism reduces inference costs, it may drive more companies to deploy AI applications, further expanding the global AI inference scale, ultimately boosting the demand for GPUs, HBM, DRAM, and high-speed network devices.
However, GDP remains cautious about this. He acknowledges the long-term logic of Jevons Paradox but points out that KDA has practical significance for state storage optimization in long-context tasks. Even with huge model weight scales, due to the use of 4-bit quantization and highly sparse design, the actual memory pressure may be lower than market intuition.
He believes that the truly noteworthy variable is whether the AI companies with the largest global inference demand—OpenAI, Anthropic, and Google DeepMind—have already adopted or will in the future adopt linear attention schemes similar to KDA, DeepSeek CSA/HCA, etc. If leading AI companies massively adopt such architectures, the demand for memory and interconnection resources for long-context inference may significantly decrease, which would become an important factor influencing the future structure of AI hardware demand.
Market focus shifts: Can AI demand growth offset architectural efficiency improvements?
Overall, the emergence of Kimi K3 does not simply point to a decline in AI hardware demand. For Nvidia, the real competitive focus may shift from "single-card performance" to "system-level capabilities"—including high-speed interconnect, rack-level expansion, and large-scale inference optimization.
As model scales continue to expand, even with ongoing optimizations in attention mechanisms, AI infrastructure still faces pressure from parameter scale, expert parallelism, data exchange, and enhancements in inference throughput.
The core issue that the market needs to focus on in the future is not whether linear attention reduces the consumption of individual resources, but whether the new demand arising from the expansion of AI application scales can continuously exceed the resource savings brought by architectural efficiency improvements.
免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。
