詹姆斯叉 | JamesX|8月 26, 2026 00:33
OpenAI isn’t building a chip.
It’s reclaiming the profit margin on every token from Nvidia.
OpenAI just announced the first round of testing for its in-house inference chip, Jalapeño:
- AI workload per watt increased by 1.5–1.9x
- End-to-end latency reduced by 1.7–3.6x
- Internal deployment planned by the end of this year
This doesn’t mean Nvidia will be replaced immediately.
Jalapeño is currently focused on inference, not training; and its initial deployment scale won’t be anywhere near Nvidia’s global supply chain.
But the AI business model is already shifting:
In the past, model companies bought compute power from chip companies and sold it to users per token.
In the future, leading model companies will control:
Models + inference software + chips + networks + racks.
Every time they lower the cost per token, they can choose to:
- Lower API prices;
- Increase free usage limits;
- Enable agents to run more steps;
- Or turn the savings into profit.
Five years from now, the biggest competition in the AI industry might not just be about “whose model is the smartest.”
It’ll be about:
Who can produce a useful token at the lowest cost.
Share To
Timeline
HotFlash
APP
X
Telegram
CopyLink