OpenAI Releases Test Results for Its First Custom Inference Chip Jalapeño, Leading in Efficiency and Latency Performance

深潮TechFlow
深潮TechFlow|8月 25, 2026 14:05
Deep潮 TechFlow reports that on August 25, OpenAI unveiled preliminary test results for its first custom inference chip, Jalapeño. Data shows that the chip achieved a 1.5 to 1.9 times peak performance per watt improvement compared to benchmark systems on models such as GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T, while reducing end-to-end latency by 1.7 to 3.6 times. Under high-interaction workloads, performance improvements reached 2.1 to 4.1 times. OpenAI stated that Jalapeño leverages a collaborative design across chip, memory, network, software, and rack-level systems to minimize data movement and communication latency, balancing high throughput with low latency. The company plans to deploy Jalapeño into its computing infrastructure by the end of this year and revealed that the second-generation chip is in advanced development, with the third-generation chip already in the planning stages.
+5
Mentioned
Share To

Timeline

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads