Written by: Techub News整理
Introduction
At the AMD Advancing AI 2026 Annual Conference held in San Francisco, AMD Chairman and CEO Lisa Su delivered a significant keynote address. As a leader in the semiconductor industry, Su's speech not only summarized AMD's AI strategy over the past year but also outlined the forward-looking layout of the AI computing landscape for the coming years. In a presentation lasting over two hours, she systematically elaborated on the paradigm shift of AI from training to inference, from cloud to edge, and alongside numerous top AI companies and partners, unveiled a series of new products and technology roadmaps that are set to reshape the industry competition landscape.
Summary
- The demand for AI inference computing has surpassed training for the first time, with an estimated 60% of global AI computing power expected to be used for inference by 2026, with Agentic AI as the core driving force.
- Significantly raised market predictions: the AI accelerator market size is expected to reach $1.4 trillion by 2030. The server CPU market will grow by over 50% due to Agentic AI, surpassing $20 billion.
- Launched the new AI rack system Helios, integrating MI455 GPU, Venice CPU, and Pensando DPU, claiming to be "the highest performing AI rack in the world," which is now in mass production.
- Introduced the Venice CPU family based on the Zen6 core, optimized for Agentic AI workloads, offering different product lines from high frequency to high density, which has been fully produced.
- Deepened ecosystem cooperation: Anthropic will deploy up to 2 gigawatts of Helios systems; OpenAI confirmed expanding AMD's infrastructure deployment to 6 gigawatts; Meta emphasized deep collaboration with AMD on CPU/GPU co-design.
- Launched the ROCm.ai platform at the software level, utilizing AI-assisted GPU programming to lower development thresholds; partnered with Hugging Face, Cisco, and others to promote AI deployment and management in enterprises and at the edge.
- Expanded AI boundaries: introduced personal AI computing platforms like Ryzen AI Halo, and the Kore.ai robot development platform, laying out physical AI and edge intelligence.
AI Paradigm Shift: From Training to Inference, Agent-Driven Trillion-Dollar Market
Lisa Su opened by pointing out that AI is the most important technology of the past 50 years, and its development speed far exceeds imagination. She revealed through data the fundamental transformation occurring in the industry: currently, the number of AI tokens consumed globally each month has reached 35 quadrillion, growing 160 times in two years. More crucially, the focus of AI computing is rapidly shifting from model training to model inference.
“We expect that by 2026, the computing power used to run (infer) AI models will exceed that used for training for the first time,” Su stated, “This year, about 60% of global AI computing capacity will be used for inference.” The core driving force behind this shift is billions of users utilizing AI daily, along with the rise of Agentic AI. Agents are no longer simple chatbots that answer questions but are "virtual employees" capable of understanding objectives, invoking tools, accessing data, and continuously working until problems are solved. This working model has generated exponential demand for computing, requiring not only a large number of GPUs for inference but also massive CPUs to coordinate every step of the operations.
Based on this insight, AMD has significantly revised its market forecast. Last year, AMD predicted that the AI accelerator market would reach $500 billion by 2028, but this figure is now considered "too conservative." Su announced that the AI accelerator market size is expected to reach approximately $1.4 trillion by 2030, close to the current size of the entire semiconductor market. Among these, GPUs will occupy the vast majority of market share due to their programmability, which allows them to adapt to rapidly evolving algorithms and workloads.
Meanwhile, Agentic AI has opened a new growth curve for the server CPU market. Su revealed that based on communications with hyperscale customers, AMD expects the server CPU market to grow over 50% by 2030, exceeding $20 billion. This growth stems from the expansion of agents from millions to billions, demanding a large CPU infrastructure to support them.
Su emphasized that AMD's strategic vision is not limited to data centers. As AI integrates into everyday life, computing must happen where the work is generated— in the cloud, edge devices, and autonomous machines. Therefore, AMD is committed to building “ubiquitous AI” computational foundations, expecting its high-performance and adaptive computing products market to grow at approximately 40% compound annual growth rate in the coming years, approaching $2 trillion by 2030.
Hardware Innovation: Helios AI Rack and Venice CPU Family Unveiled
In the face of enormous market opportunities and complex workloads, Su revealed AMD’s answer: Helios—the industry’s highest performance AI rack system. Helios is designed for large-scale training and running cutting-edge AI models, and its core concept is to design the entire rack as a complete system.
The core components of Helios include:
- MI455 Accelerator: using TSMC’s 2nm/3nm process, it has 320 billion transistors, integrates 12 computing and I/O chiplets, equipped with 432 GB HBM4 memory. Su calls it "the industry's highest performance AI accelerator."
- Venice CPU: based on the new generation Zen6 core, used to drive GPUs and coordinate workloads.
- Pensando DPU: provides leading programmability and lateral scaling bandwidth for network connections.
Su showcased the Helios system, which integrates 75 GPUs through high-speed interconnections within a single rack. The data she released showed that compared to competitors, Helios can offer 15% higher computing performance, 50% more HBM4 memory capacity and bandwidth, and 50% more lateral scaling bandwidth. Under fixed rack power consumption, Helios leads competitors by 10-15% in average performance on mainstream inference workloads, while also having 30% higher processing capability per dollar token.
She excitedly announced that Helios has entered comprehensive mass production, with shipments expected to begin at the end of the third quarter, ramping up in the fourth quarter and the second half of the year, with "extremely strong" customer demand.
On the CPU front, Su emphasized the new Epic Venice CPU Family tailored for the "Agentic Era." She noted that Agentic AI is causing a differentiation in server workloads: GPU servers need high-frequency cores to rapidly feed the GPUs; Agentic sandbox servers require high-density cores to run thousands of agents simultaneously; traditional general-purpose servers seek energy efficiency.
The Venice family based on the Zen6 core uses chiplet design, flexibly responding to different needs:
- Venice HF: High-frequency version for AI host nodes, driving GPUs in Helios.
- Venice 256 cores: High-density version specifically designed for Agentic sandboxes, offering the industry’s highest compute density.
- Venice 128 cores: Optimized version aimed at enterprise and general servers for the best cost-performance ratio.
In terms of performance, Venice shows performance improvements of up to 1.8 times over the previous Turin generation. In Agentic sandbox scenarios, Venice’s agent processing capability per watt is 2.8 times leading ARM processors; at the rack level, performance per watt leads by 3.3 times. Su also announced that Venice has been fully produced, with the strongest customer demand for the Epic series ever, and all major server OEMs and cloud providers will begin rolling out in the fourth quarter.
Ecological Prosperity: Building an Open Future with Top AI Companies
Su's speech was not a solo performance; she invited several heavyweight partners to the stage to collectively validate the strength of the AMD platform and the value of the open ecosystem.
Anthropic co-founder and Chief Computing Officer Tom Brown took the stage to announce the deployment of up to 2 gigawatts of Helios systems. He shared a fun fact: an engineer at Anthropic once tried to deploy the MI355 system independently, successfully improving the model's performance curve with the help of the Claude agent over the weekend, demonstrating the ease of use of AMD’s open platform. The two parties will also collaborate to utilize Claude to optimize AMD's chip design and software kernel development.
OpenAI infrastructure head Sachin Katti then followed, reaffirming the depth of cooperation with AMD. He confirmed that OpenAI is advancing towards the previously promised 6 gigawatts of AMD infrastructure deployment and is one of the first to receive the MI455 racks. The engineers from both sides are working together to optimize the software stack so that GPT-level workloads can run on Helios. Katti specifically emphasized the importance of recursively improving the system itself using AI, and how AMD's open software ecosystem enables such innovations to benefit the entire industry.
Meta infrastructure head Santosh Janardhan stressed the critical importance of "co-design" in the AI era. He pointed out that AI is not just a competition for GPUs; CPUs have become at least equally important or even more so. Meta is transitioning from optimizing individual servers to designing the entire data center as an integrated system, which requires early and deep collaboration with partners like AMD. He revealed that the collaboration with AMD has spanned multiple generations of CPUs (Milan, Bergamo, Turin, Venice) and GPUs (MI300, MI350, MI450), and will continue to co-design systems aimed at 2027 and 2028.
Additionally, AMD announced a collaboration with AI chip company Cerebras to combine Cerebras’ wafer-scale engine with the Helios rack, creating a solution for ultra-low-latency inference scenarios. Cerebras CEO Andrew Feldman stated that this solution can maintain high speeds while providing five times the throughput, expected to be launched later this year.
Software and Edge: Making AI Accessible and Ubiquitous
Hardware is foundational, but software is key. AMD Senior Vice President Vamsi Boppana took the stage to introduce the latest developments in the ROCm software stack. He announced the launch of ROCm.ai, an agent AI platform aimed at using AI-assisted GPU programming to automate kernel generation, code optimization, and debugging, significantly lowering the development threshold, allowing developers to efficiently utilize AMD hardware without extensive ROCm expertise.
At the same time, AMD is pushing AI capabilities from the cloud to the edge and personal devices. Su and another executive showcased personal AI computing platforms like Ryzen AI Halo, supporting models with up to 300 billion parameters to run locally, providing developers with a seamless path from local experimentation to large-scale deployment. AMD announced an expanded collaboration with Hugging Face to provide optimized models and tools for Ryzen AI Halo, and each Halo device will come with one year of Hugging Face Pro services.
To manage the large-scale deployment of edge AI, AMD partnered with Cisco to launch integrated solutions. Cisco Chief Product Officer G2 Patel elucidated the concerns of enterprise CIOs regarding cost and security and demonstrated the Cisco cloud control platform integrated with AMD Halo devices, capable of unified management of AI inference infrastructure from edge to cloud, monitoring agent behavior and token costs, and implementing security policies.
Finally, AMD has expanded the boundaries of AI into the physical world by releasing the Kore.ai system-level module and robot development platform for the robotics field. This platform integrates CPUs, GPUs, NPUs, and unified memory, aiming to provide a real-time perception, inference, and action control brain for autonomous robots. Su stated that in third-party tests, the Kore platform outperforms competitors in real-time performance, concurrent agent numbers, and CPU capacity.
Future Blueprint: Continuous Iteration to Lead the Next Decade of AI
At the closing of her speech, Su looked ahead to AMD's future technology roadmap, showcasing its long-term commitments:
- CPU: Launching the Florence Epic family based on the Zen7 core in 2028; planning to launch the Ravenna family based on the Zen8 in 2030.
- GPU: Maintaining a generational iteration rhythm every year. The MI500 will use the next-generation HBM and copper/optical interconnect, expected to bring the largest intergenerational performance leap in Instinct history. The MI600 (based on the CDNA Next architecture) is in development, targeted for 2028.
- Systems: Customers can expect AMD to release new generations of Helios rack systems every year, continually improving performance, capacity, and energy efficiency.
Su concluded by stating that AI is moving from a stage of discussing possibilities to one of generating real, significant impacts. AMD is committed to building the technologies, roadmaps, and partnerships necessary to provide all the tools and capabilities needed for AI to create meaningful real-world impacts in every corner from data centers to personal devices to the physical world. “I have never been as confident as I am today that high-performance computing can make the world a better place,” she concluded, highlighting AMD's ambition to lead the next era of AI computing.
免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。