GPT-5.6 Luna dropped 80% this week — $1.00/$6.00 per million input/output tokens down to $0.20/$1.20. Terra fell 20%, to $2/$12. Sol picked up a Fast mode: 2.5× the throughput for 2× the price, a genuine tradeoff rather than a freebie. Artificial Analysis published a chart the same week putting Luna at the top of a cost-versus-intelligence frontier.
Look at that chart for longer than a headline requires and it says something different than "OpenAI won." The connected frontier line runs entirely through OpenAI's own model family. Every competitor — GLM-5.2, Gemini 3.6 Flash, Claude Opus 5, Claude Sonnet 5, DeepSeek V4 — sits off the line as an isolated point, several of them at worse scores for higher prices. That's not necessarily unfair. It's also precisely the chart OpenAI's own team would draw. Treat it the way this blog treats every vendor's self-graded benchmark: real, but not neutral.
Which raises the actual question a price cut like this is supposed to answer: cheaper than whom, exactly, and for how long does that matter?
A market that's already been priced doesn't need repricing
Lay out what's currently on offer, per million tokens, and the shape of the market is hard to miss:
| Model | Price ($/M tokens) | Tier |
|---|---|---|
| GPT-5.6 Luna | 1.20 out | Fast, cheap |
| Gemini 3.5 Flash-Lite | 2.50 out | Fast, cheap |
| DeepSeek V4 | $0.87 out (reported) | Fast, cheap |
| GLM-5.2 | $4.40 out (reported) | Mid |
| GPT-5.6 Terra | 12.00 out | Mid |
| Gemini 3.1 Pro | 12.00 out | Mid |
| Claude Sonnet 5 | 10.00 out | Mid-high |
| Grok 4.5 | 6.00 out | Mid |
| Gemini 3.6 Flash | 7.50 out | Mid |
| GPT-5.6 Sol | 30.00 out | Flagship |
| Claude Opus 5 | 25.00 out | Flagship |
| Claude Fable 5 | 50.00 out | Frontier |
| Kimi K3 (hosted) | $15/M output | Frontier, open weights |
The bottom tier isn't a toy anymore — Luna's new price puts it in the same conversation as models that were flagship-grade eighteen months ago. Sol's Fast mode isn't a new model at all, just a new dial on an old one, which is probably the more honest preview of where this goes: not new tiers, more knobs on existing ones. And Kimi K3 sits entirely outside this table, because the moment you self-host 2.8T open weights, "price per token" stops being the number that matters and your own infrastructure bill takes over.
Here's the puzzle that table doesn't explain on its own: this information has been public, comparable, and screenshottable for over a year. If buyers actually shopped this market the way the chart assumes — scanning a price-performance frontier and routing to whoever's winning it this week — the field would have consolidated by now around whichever provider is cheapest at a given quality bar. It hasn't. Enterprises that started on one provider are, overwhelmingly, still on that provider, discount or no discount, chart or no chart.
Buyers aren't shopping, because leaving costs more than staying does
A 2026 Parallels survey of enterprise leaders makes the gap explicit: 94% say they're worried about AI vendor lock-in, and 89% believe they could switch providers within a month if they wanted to. Of the ones who actually tried, 58% say it either failed outright or took far more effort than they'd budgeted for. Seventy-four percent say losing their current vendor would disrupt daily operations or stop the business from functioning.
That's not a perception problem correcting itself over time — it's a structural one, and it's mostly invisible until someone tries to leave. Two models exposing an identical, OpenAI-compatible chat endpoint can still disagree on tool-calling format, refusal thresholds, and context handling, so a prompt tuned to one model's quirks quietly degrades on another's, and the failure shows up in production rather than in testing. Fine-tunes and eval suites are built against one model's tokenizer and failure modes and don't transfer — they get rebuilt, not redeployed. Enterprise pricing tiers often carry $10,000–$50,000 annual volume commitments, so switching means forfeiting money already spent. And the team that would benefit from a cheaper model is rarely the team that owns the integration and would have to redo the work.
Put those together and a price cut is answering a question almost nobody is actively asking. The buyer who'd benefit from Luna's new price isn't running a live comparison against Gemini Flash-Lite; they're not running any comparison, because the cost of finding out isn't the API bill, it's the rebuild.
What actually moves someone
If price rarely triggers a switch, something else has to — and the pattern across real cases is that it's involuntary. OpenRouter's routing layer exists largely because providers deprecate models out from under their customers: more than 70 pulled or restricted across major providers in the last few years, each one forcing a migration on the provider's calendar, not the customer's. Outages do more damage to loyalty than pricing ever does — routing infrastructure defaults to deprioritizing any provider with a recent failure, because a slow demo in front of a client costs more than a good rate ever saved. And nobody re-architects a production pipeline because a benchmark ticked up two points; they do it when a task the old model flatly couldn't do — long context, native audio, some real qualitative jump — becomes possible and stops being defensible to ignore.
The clearest version of an involuntary switch already happened this year, not hypothetically. Claude Fable 5 was pulled from availability entirely under a US export-control directive on June 12, restored nineteen days later. Nobody running production infrastructure on Fable 5 got a comparison chart or a grace period. They got a Friday.
Read against that, GPT-5.6's price cut stops looking like an opening bid in a negotiation nobody's having. It's insurance against the one scenario where the negotiation does happen — the rare day a customer is forced to look, whether by an outage, a deprecation notice, or a capability gap that finally became too expensive to work around. On that day, being the cheapest reasonable option is worth more than being the cheapest option was worth on any of the preceding 364.
But the API is only one of three places a person can "leave a provider" from — and it's the one where the question matters least.
One floor up: where switching is a feature, not a failure
At the raw API level, swapping really is close to trivial — compatible endpoints, a string change in a router. Everything above about enterprise stickiness is organizational cost wrapped around a technically easy operation. One floor up, in the tools developers actually work in, even that framing dissolves: a developer in Cursor doesn't migrate between providers at all. The model is a dropdown. Changing it isn't a project; it's a preference, sometimes set per task — this model for refactors, that one for tests.
Harnesses like Cursor exist precisely to make the model fungible: same keybindings, same context engine, same workflow, any brain. At this layer the switching cost has been engineered to zero on purpose — and loyalty doesn't attach to the model, because nothing accumulates in the model. It accumulates in the tool: muscle memory, project context, configuration, trust. Ask a Cursor user to drop Claude for a week and they'll shrug. Ask them to drop Cursor and you're back to the 58%-harder-than-expected story — the lock didn't disappear, it moved into the harness.
The labs noticed, which is why every one of them now ships a harness of its own — Claude Code, Codex, Gemini CLI. When the API is a commodity and the model is a dropdown, the only place left to rebuild attachment is the tool itself. Claude Code isn't a side feature; it's the structural answer to Cursor having made models interchangeable — re-bundling the model with the workflow, so that leaving the model once again means leaving something you'd miss.
And one floor down: the app, where nobody compares anything
Then there's the person all of this is ultimately for — not the developer, the ordinary subscriber with one chat app on their phone and €20 a month leaving their account. They will never see the pricing table above, never run an eval, never hear about this week's cut. Their version of the question isn't "is Luna cheaper per token." It's "would I leave this app?" — and almost nothing gradual will ever make them.
What holds them isn't the model's benchmark position; it's what the app has accumulated around them. Months of history. Memory — the assistant that already knows your job, your projects, your kids' names, your writing tics is the only one holding an asset that grows, and there is no import button for it anywhere: you can export your history, but no rival will ingest it into its own memory. The subscription is already paid, and it costs the same twenty everywhere — which is not an accident. None of the labs wants the consumer decision to be about price, because on price they're identical; the real competition is over who holds your context.
What actually moves a subscriber is the enterprise story translated into consumer terms: not comparison, but events. A capability moment that goes viral and is visibly missing from their app — an image style everyone is posting, a research feature everyone is sharing. A product that only exists on the other side — which is quietly how Claude Code moved a generation of developers into Anthropic's consumer ecosystem: the terminal tool came first, the app subscription followed it home. And above everything, defaults: Gemini arrives preinstalled in Android and Workspace, Copilot inside Windows and Office. For most people the best model is the one already on the home screen, and distribution beats benchmarks by a margin no price cut will ever touch.
Three markets wearing one name
So "leaving a provider" is three different acts in three different markets. At the API, switching is technically easy and organizationally rare — it happens involuntarily, on someone else's Friday. In the harness, switching is so easy it stops being an event at all — the model competes per task, and the lock has migrated into the tool. In the app, switching is rarest of all, and moves only on moments, products, and defaults — never on tokens per dollar.
This week's price cut lives entirely in the first market. In the second, it changes which dropdown entry gets picked this afternoon. In the third, the subscriber will never hear about it — which is exactly why the labs that want the subscriber are spending on memory, on default placement, and on tools like Claude Code, rather than on the number the headline was about.
Links
- Announcement: OpenAI — Advancing the price-performance frontier with GPT-5.6
- Coverage: VentureBeat — AI price wars · Axios
- Switching-cost data: Parallels 2026 enterprise vendor lock-in survey
- Routing mechanics: OpenRouter — how model routing works
- Related: Claude Fable 5's forced 19-day suspension · pricing history across 234 releases in the arcade