OpenAI and Anthropic, prominent U.S. artificial intelligence laboratories, have initiated price reductions for their mid-tier AI models, leading to a decrease in consumer costs by approximately 25% since mid-July. This adjustment reflects a response to a market where customers are becoming more sensitive to the rising expenses associated with AI model usage. Simultaneously, Chinese AI developers, including Moonshot and DeepSeek, are expanding their presence in U.S. and European markets by providing models that are both more economical and increasingly comparable in capability to their Western counterparts.
OpenAI, backed by Microsoft, recently implemented an 80% price cut for its GPT-5.6 Luna model, reducing input tokens from $1 to $0.20 per million and output tokens from $6 to $1.20 per million. The company also lowered the price of its GPT-5.6 Terra model by 20%. Anthropic, supported by Amazon and Alphabet, introduced its Claude Opus 5 model at half the price of its predecessor, Fable 5, setting its input tokens at $5 per million and output tokens at $25 per million. Anthropic also canceled a planned price increase for its Sonnet 5 model, which was scheduled for September.
Despite these price reductions on mid-tier models, both OpenAI and Anthropic are either maintaining or increasing the prices of their top-tier offerings. Industry analysis suggests that a more capable, albeit more expensive, model might complete tasks with fewer tokens, potentially leading to a lower overall cost for users.
The underlying cost of generating AI tokens is also decreasing due to more efficient computing systems, contributing to the downward pressure on token prices. Data from OpenRouter, a service that provides access to various AI models, indicates that after OpenAI's price cuts, the effective price of using Luna fell approximately tenfold, while consumption surged about 14-fold. Similarly, Terra's effective price decreased roughly threefold, and its usage increased about fivefold. This surge in usage outpaced the price reductions, leading to an estimated 34% increase in revenue for OpenAI from Luna and a 45% increase from Terra, according to TD Cowen analysts.
Chinese AI firms have been actively competing on price. Baidu's ERNIE Bot offers a professional plan for approximately $8.20 per month, with a discounted rate for auto-renewal. Baidu also provides API access for its ERNIE 4.5 model, with input prices starting at RMB 0.004 per thousand tokens (approximately $0.00055 USD) and output prices at RMB 0.016 per thousand tokens (approximately $0.0022 USD). Alibaba Cloud's current flagship, Qwen3.8 Max, released on August 3, 2026, is priced at $2.00 per million input tokens and $6.00 per million output tokens. Its previous flagship, Qwen3.7 Max, is available at $1.25 and $3.75 per million tokens. The cheapest Qwen model, Qwen3.7 Flash, starts at $0.03 per million input tokens and $0.13 per million output tokens for requests under 32,000 tokens. Tencent's Hunyuan models offer competitive pricing, with the TurboS variant providing 256,000 context at $0.11 per million input tokens and $0.28 per million output tokens. Moonshot AI's Kimi K3 flagship model costs $3 per million input tokens and $15 per million output tokens, with cached input at $0.30 per million.
The competitive pricing landscape extends beyond per-token costs. Many providers, including Google, OpenAI, and Anthropic, offer a 50% discount for batch API processing, which allows for asynchronous job completion. This discount can be combined with prompt-cache pricing, further reducing costs for repetitive inputs. For example, a batched Opus 5 cache hit from Anthropic could cost $0.25 per million tokens.
The increased competition and falling prices are enabling businesses to apply AI to a wider range of tasks that were previously too expensive, leading to expanded usage across various applications.
