In June of 2025 I predicted that Chinese open weight models would set the LLM pricing floor. The piece pointed at DeepSeek-V3 and DeepSeek-R1 turning up in the cheap tiers of tools like Windsurf and said “these will likely become the new baseline for price-sensitive usage.” I wrote it a month after a $622 Anthropic bill (up from around $50 the month before) and after burning a $20 Zed plan’s entire monthly allocation in a few hours. My reasoning was this: the frontier providers were selling inference below cost, the subsidy would end, and the cheap models were waiting underneath.
I gave it two time windows. The tighter one sat with the claim that “Chinese competition will eventually drive costs down significantly,” following “a period—probably 6-12 months—where usage-based pricing hits hard.” That one ran out around June. The looser one sat with a broader claim, that Chinese pressure would “eventually drive down pricing across the board,” with the caveat that “eventually” might be 12-18 months out. That one runs to December.
The first window failed outright. The second is failing in a direction I didn’t anticipate: not that the floor held, but that the floor is about to go up.
DeepSeek’s own API pricing documentation now carries this: “We plan to raise the overall pricing for DeepSeek API services in the near future, with a significant increase expected.” No percentage. No effective date. No list of affected models. The South China Morning Post, which reported the notice, has DeepSeek telling developers that specifics are “to be advised later” and that “users should plan their usage accordingly.” SCMP attributes the pressure to demand for DeepSeek-V4-Flash-0731, a 284-billion-parameter model released on July 31, a week before the notice went up.
The company I expected to set the floor has signaled it will raise it. The labs I expected to raise prices spent the last four months cutting prices or raising limits.
TL;DR
In June 2025 I gave Chinese models two windows to set the price floor: 6-12 months, and 12-18 months. The first ran out around June 2026. The second runs to December, and DeepSeek is now warning of a “significant” price increase with no figure and no date attached, confirmed on its own pricing docs, not just in press coverage.
Over the same stretch the frontier moved the other way. Anthropic doubled Claude Code’s 5-hour limits and dropped peak-hour throttling (May 6). OpenAI split its $200 Pro tier into $100 and $200 tiers running identical models, differing only in usage volume (April 9). Google cut its top Gemini subscription from $249.99 to $199.99 (May 19).
On July 30 OpenAI cut GPT-5.6 Luna by 80%, to $0.20 input and $1.20 output per million tokens. Multiply DeepSeek V4-Flash by ten and it lands at $1.40/$2.80: 3.6x to 7x under Opus 5 and Fable 5 on input where today it is 36x to 71x under them, and more expensive than Luna.
OpenAI’s stated reason for the cut was efficiency: 20% lower serving costs and 15%+ better token-generation efficiency. That covers a 20% cut. Terra got 20%. Luna got 80%.
The rationing I predicted did happen, one layer down. Cursor converted flat billing to metered credits in June 2025. GitHub Copilot moved to token-metered AI Credits on June 1, 2026, at unchanged sticker prices.
What DeepSeek Disclosed
V4-Flash lists at $0.14 per million input tokens and $0.28 per million output, with cache hits at $0.0028. V4-Pro lists at $0.435 and $0.87. Those are the pre-hike numbers. DeepSeek disclosed neither the size of the increase nor when it lands. Any number you see attached to this story is somebody’s guess.
A “significant” increase off a base that low still leaves room. Double V4-Flash and it sits at $0.28/$0.56, cheaper than most of what the frontier sold a year ago. The magnitude we don’t know yet. The direction we do, and that is the part I called backwards.
The Frontier Went The Other Way
Anthropic doubled Claude Code’s 5-hour rate limits on May 6 for Pro, Max, Team, and seat-based Enterprise plans, and removed the peak-hours limit reduction for Pro and Max. Max 5x is still $100 a month. Max 20x is still $200. Both buy more than they did in the spring.
On April 9 of this year, OpenAI split its $200 Pro tier into $100 and $200 tiers running identical models and differing only in usage volume. That’s a second fixed-price plan at Claude Max’s number. Neither tier is uncapped. OpenAI doesn’t publish the Pro-model allowance on either, and running it out soft-degrades you to a smaller model rather than cutting you off.
Google restructured AI Ultra at I/O on May 19. The single $249.99 tier became $99.99 at 5x Pro limits and $199.99 at 20x, with daily prompt caps replaced by compute-weighted 5-hour refresh and weekly metering. The top tier alone dropped $50.
Who GPT-5.6 Luna Is For
Then came July 30. OpenAI cut GPT-5.6 Terra from $2.50/$15 to $2/$12, a 20% reduction. It cut GPT-5.6 Luna from $1/$6 to $0.20/$1.20, an 80% reduction. The stated reason, from OpenAI’s own account: “20% lower serving costs from production GPU kernel improvements. 15%+ better token-generation efficiency from improved speculative decoding.”
That explains Terra. Twenty percent cheaper to serve, twenty percent off the price, a straight pass-through. It doesn’t explain Luna. The same kernels and the same speculative decoding produced a cut four times larger on the cheaper model. Efficiency gains don’t sort themselves by tier. Pricing decisions do.
Max plans holding steady is not by itself evidence that providers depend on volume. Falling serving costs cover that on their own. A lab building a tier at the bottom of the market and then cutting it 80% in a single move is harder to explain that way, because a $0.20 input tier is not where a company with a $5/$25 flagship model makes its money. It’s where a company goes when it wants a number that reads lower than DeepSeek’s.
Two months ago I wrote that dealers give the first one away for a reason, and that the Max plan was my first bag. The habit runs in both directions. A dealer who keeps cutting the price of the first bag is a dealer who can’t afford to have you buy it somewhere else.
Run my June 2025 claim forward against a tenfold hike. DeepSeek disclosed nothing supporting that multiple or any other, so it’s a stress test, not a forecast. V4-Flash at ten times list: $1.40 input, $2.80 output. Against the models people think of when they say “frontier”, a multiple survives. Claude Opus 5 at $5/$25 is 3.6x the input price and 8.9x the output. Claude Fable 5 at $10/$50 is 7.1x and 17.9x. GPT-5.5 at $5/$30 is 3.6x and 10.7x. Those are reduced from much bigger numbers. At today’s list price Opus 5 costs 36x what V4-Flash does on input and 89x on output. A tenfold hike takes about ninety percent of that away. At 36x, price picks the model. At 3.6x, the model does.
Further down the market it reverses. Against Luna at $0.20/$1.20, a tenfold-hiked V4-Flash costs seven times more on input and more than twice as much on output. Against Gemini 3.6 Flash at $1.50/$7.50 it’s roughly a wash on input. Even Gemini 3.1 Pro at $2/$12 for prompts at or under 200k tokens (above that it goes to $4/$18) sits within 1.4x on input (4.3x on output). Run the stress test and the frontier ends up priced below the discounter, not for the flagship but for the tier built for exactly the work I said DeepSeek would take.
The Case Against
Three things complicate this reading.
DeepSeek’s prices track chip supply, not just strategy. The company cited “constraints in high-end compute capacity” when it priced V4-Pro in April 2026, a constraint widely linked to US export controls on Nvidia H20 chips, though DeepSeek didn’t specify. In late May it made a 75% V4-Pro price cut permanent without saying why, on timing that lines up with easing Huawei Ascend 950 supply. The August notice gives no reason at all, though SCMP’s reporting points at demand for a model released the week before. I can’t tell a capacity bottleneck from a strategic retreat from outside the company.
The rationing happened at the consumer level. Cursor replaced flat request-based Pro billing with usage-based credits pegged to API cost in June 2025, then ran a refund program from June 16 to July 4 after users got painful bills. GitHub Copilot moved from flat Premium Request Units to token-metered AI Credits on June 1, 2026, at unchanged sticker prices, with the code-review multiplier rising to 13x for legacy annual-plan holders. Anthropic added weekly caps on top of its 5-hour caps in August 2025, citing round-the-clock usage and account resale, and estimated at announcement that they would “apply to less than 5% of subscribers based on current usage.” Model-lab sticker prices held. The metering moved to heroin layer.
The efficiency explanation may be the whole explanation. Serving costs have fallen steeply, and Anthropic’s gross margin has reportedly climbed off a deeply negative base, roughly -94% in 2024. The Information reported in January 2026 that Anthropic had trimmed its own projection to about 40% for 2025, still a steep climb off that base. If cost per token falls faster than price per token, generous Max plans need no dependency on volume to explain them. My reading rests on the 20/80 split between Terra and Luna, and that’s one data point on one day from one vendor.
What Would Prove Me Wrong Again
The June 2025 piece got the $200 line right, which was the easy part. What I got wrong was which end of the market was fragile. Any of these would tell me I’m getting it wrong a second time.
DeepSeek publishes numbers and the increase is under 2x on V4-Flash, leaving it under $0.30 input. That would be a capacity adjustment — the floor is where I thought it was.
OpenAI raises Luna back toward $1 within two quarters. If that happens, July 30 was pass-through of a real engineering result, and I read a price war into it.
Anthropic or Google restores throttling, or cuts 5-hour allowances at unchanged prices, before the end of 2026. That would mean the generosity was a promotional window, and June 2025 was early rather than wrong.
The tool layer keeps metering while the model labs keep expanding: the pricing pressure was sitting at the tool layer the whole time, and both of my previous pieces were aimed at the wrong tier.
DeepSeek’s hike lands and its API volume holds anyway — which would mean price wasn’t what price-sensitive buyers were choosing on.
If you’re budgeting against a Chinese-model price floor for 2027, price in the possibility that the floor is now a frontier lab’s customer-acquisition tier.
Bob Matsuoka is CTO of Duetto, a hospitality revenue-management platform, and writes about AI-augmented engineering practice.
Related reading:
The Other Shoe Will Drop — The June 2025 prediction this piece corrects.
The Other Shoe Has Dropped — Why per-token price cuts stopped reaching the invoice.
AI Power Ranking — Tool comparisons and benchmarks for AI practitioners.
LinkedIn Newsletter — Strategic AI insights for CTOs and engineering leaders.






