OpenAI launched three weeks ago with an 80% price reduction, and the Chinese model triggered the fiercest "token price war" in AI history.

CN
6 hours ago
The AI industry has officially entered a new phase of "good enough, whoever is cheaper will do."

Author: Claude, Deep Tide TechFlow

Deep Tide Guide: OpenAI has drastically slashed the price of GPT-5.6 Luna by 80% to just $0.2 per million input tokens, Terra has decreased by 20%, forced to significantly adjust prices just three weeks after launching. The driving force behind this is the strong penetration of Chinese open-source models into the U.S. enterprise market: the Dark Side of the Moon Kimi K3 has been open-sourced with 2.8 trillion parameters, and DeepSeek V4 Pro only costs 4 cents per task. The token usage share of Chinese models on OpenRouter has surged from 4.5% a year ago to a peak of 46%. Anthropic and Google are following up with price cuts or launching cost-effective new products. The AI industry has officially entered a new phase of "good enough, whoever is cheaper will do."

On July 30, OpenAI announced significant price reductions for two models in the GPT-5.6 series, just three weeks after its launch. CEO Sam Altman posted on the X platform saying "major price cuts today," adding that "we want to offer the best price/intelligence ratio at every level."

This is not a typical product iteration price reduction. The GPT-5.6 series was officially launched on July 9, and just three weeks later, they made a substantial cut on the mid to low-end models. The scale and speed of this move have never occurred in OpenAI's history. The context is clear: the share of American enterprise token usage for Chinese open-source models on the OpenRouter platform has soared from 4.5% in the first half of 2025 to a peak of 46% by mid-2026. OpenAI can no longer maintain its prices.

Luna cut by 80%, Terra cut by 20%, flagship Sol unchanged

The input price of GPT-5.6 Luna (the fastest and lightest version) has dropped from $1 per million tokens to $0.2, and the output has fallen from $6 to $1.2, a total reduction of 80%. Terra (mid-range) input fell from $2.5 to $2, and output from $15 to $12, a reduction of 20%. The flagship Sol remains unchanged at $5/$30.

At the same time, OpenAI has added a Fast mode for Sol, increasing the speed by 2.5 times with doubled prices (input $10/output $60), replacing the previous Priority Processing. The Auto-review feature in the ChatGPT app and Codex CLI has also upgraded from GPT-5.4 to Luna, which OpenAI claims will reduce the review cost of the agent workflow by about 10 times.

The strategy is clear: maintain the flagship premium for Sol while aggressively cutting prices on mid to low-end products to intercept customer churn. While OpenAI keeps the cutting-edge high-end intelligence intact, it treats the mid-market share as a line of defense.

Chinese models invading the U.S.: 46% token share

The direct catalyst for this price reduction is that the penetration speed of Chinese models in the U.S. enterprise market has exceeded everyone's expectations.

According to a CNBC report on July 7, since February 8, 2026, the usage share of Chinese models among U.S. enterprise tokens on OpenRouter has remained above 30% weekly, peaking at 46%. DeepSeek alone accounts for 17.6% of the total token volume on the OpenRouter platform, approximately 51.3 trillion tokens weekly, making it the largest single supplier on the platform, ahead of all U.S. labs. Alibaba's Tongyi Qianwen ranks second among Chinese manufacturers with a 13.9% share.

The price difference is the core driving force. DeepSeek V4 Flash charges only $0.14 for every million input tokens, while OpenAI's GPT-5.5 charges $5, making the difference as high as 35 times. The AI startup Lindy has migrated 100% of its traffic from Claude to DeepSeek, expected to save millions of dollars. Airbnb has confirmed it is using Tongyi Qianwen in production, which has even led to inquiries from the U.S. Congress.

Data from OpenRouter shows that the substantial token demand from U.S. enterprises is becoming a commodity, routed to the "good enough" cheapest models.

Kimi K3: 2.8 trillion parameters open-sourced, leading front-end coding capability

Kimi K3 from Moonshot AI is the most impactful variable in this price war.

K3 launched its API on July 16 and officially opened the weights on July 27. It boasts a total parameter count of 2.8 trillion, using an MoE architecture (activating 16 out of 896 experts per token, with 10.4 billion active parameters), and a context window of 1 million tokens. This is currently the largest open-source model, nearly double the previous record holder, DeepSeek V4 Pro, which had 1.6 trillion parameters.

In terms of performance, K3 ranks first in the LMArena front-end coding arena, surpassing Anthropic's flagship Claude Fable 5. On the single-task cost index from Artificial Analysis, K3 costs $0.94 to complete one task, while DeepSeek V4 Pro only costs 4 cents, OpenAI's flagship Sol costs $1.04, and Anthropic's Claude Opus 4.8 costs $1.80.

K3's API pricing is $3 per million input tokens and $15 per output, but the input price drops to $0.30 after cache hits. Moonshot AI reports that the cache hit rate exceeds 90% in programming workloads, meaning K3's effective input cost is just $0.30 per million tokens.

Due to strong demand, Moonshot AI has had to temporarily limit new subscriptions and API access because of computational constraints. The company's daily revenue has grown approximately sixfold since K3's release and is seeking a $50 billion valuation, preparing for a Hong Kong IPO.

DeepSeek V4 Pro: 4 cents to complete one task

DeepSeek is another disruptive force in this price war.

The API pricing for V4 Pro is $0.435 per million input tokens and $0.87 per output. This price comes from the 75% promotional discount in late May being converted into a permanent price, with the previous price being $1.74 for input and $3.48 for output. V4 Flash is even cheaper: $0.14 for input and $0.28 for output. Yahoo FinanceQuasa

To intuitively grasp this gap: the output price of GPT-5.6 Sol is $30 per million tokens, while DeepSeek V4 Pro is $0.87. With the same budget, the amount that can be processed using DeepSeek is 34 times that of OpenAI's flagship.

This explains why the total token throughput on OpenRouter has grown from about 50 trillion tokens weekly in April 2025 to over 200 trillion by April 2026, with the explosive growth of the platform primarily driven by the low prices brought by Chinese models. Towards AI

Full-scale melee: Anthropic, Google, and Microsoft in sync

The price cuts by OpenAI are not an isolated incident. In the past two weeks, nearly all top labs have made intensive adjustments in pricing and products.

On July 24, Anthropic released Claude Opus 5, priced at $5 per million input tokens and $25 for output, matching the previous generation Opus 4.8, which is only half of Fable 5's price ($10/$50). Anthropic positions Opus 5 as a daily use, not painful to the wallet, model, with the core selling point being the performance close to flagship Fable 5 at half the cost.

This month, Google launched three new Gemini Flash models, claiming that its top Flash model has lower single-task costs than Kimi K3 and other Chinese models. Microsoft is promoting its cyber security-specific model MAI-Cyber-1-Flash, with AI lead Mustafa Suleyman stating it achieves world-leading performance at "50% of the cost."

All players' actions point to one conclusion: enterprise customers are becoming increasingly sensitive to AI spending, demanding clearer returns on investment and starting to turn to cheaper alternatives.

The essence of the price war: AI tokens are being commoditized

The underlying logic of this round of price cuts is that AI inference tokens are transitioning from "scarce cutting-edge capabilities" to "interchangeable computing units."

OpenAI's pricing strategy for GPT-5.6 (Sol/Terra/Luna) itself illustrates the issue: layering the sale of model capabilities, maintaining flagship pricing while aggressively cutting mid and low-end prices is a mature industry strategy, not a pricing approach during a technological monopoly. OpenAI has directly inserted Luna's pricing into the price range of DeepSeek and Google’s low-end models, while Sol's output price of $30 remains high.

OpenAI attributes the price cuts to increased efficiency. According to TechTimes, GPT-5.6 Sol has autonomously rewritten its production GPU kernels and speculative decoding models in the Codex environment, marking it as the first known AI model to "optimize its own reasoning costs." OpenAI claims it is channeling these efficiency gains into pricing.

But the market's interpretation is more straightforward: Gartner analyst Arun Chandrasekaran believes this is a critical moment for buyers—negotiating prices and business terms with cutting-edge AI labs has been notoriously difficult.

The rise of Chinese open-source models has not only changed prices but has also altered the entire market structure. The so-called era of unrestricted AI spending is giving way to a more pragmatic era. This is good news for developers and enterprises. For the leading labs that rely on high prices to sustain research and development efforts, the compression of profit margins has only just begun.

免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。

Share To
APP

X

Telegram

Facebook

Reddit

CopyLink