The Price Elasticity of Intelligence
Thesis
OpenAI CFO Sarah Friar said on September 7, 2026 that an 80% price reduction for GPT-5.6 Luna helped drive roughly 10x higher usage.
If usage is treated as a proxy for billable token volume, the simple math is striking: 0.2 × 10 = 2.0. Intelligence became 80% cheaper, yet the implied gross inference-spend proxy doubled.
That gives us a sharper question than "How cheap will models get?"
How elastic does demand for intelligence become as its price approaches zero?
The Price Elasticity of Intelligence
OpenAI cut the price of GPT-5.6 Luna by 80%. Usage rose roughly tenfold.¹
The first number looks like margin pressure.
The second may tell a different story.
If usage scales roughly with billable token volume, the arithmetic is simple:
Price: 1.0 → 0.2
Usage: 1.0 → 10.0
Spend proxy: 1.0 → 2.0
An 80% reduction in unit price can therefore coexist with roughly twice the gross inference spend.
That is not reported OpenAI revenue. It is an indexed calculation. Customer contracts, cached tokens, input-output mix, workloads and other pricing effects can change the actual economics. But the direction is important. Lower prices do not necessarily shrink the market. They can expand the amount of intelligence customers choose to consume.
OpenAI itself describes the mechanism directly. When useful intelligence becomes cheaper, more work becomes economical to perform. Luna now costs $0.20 per million input tokens and $1.20 per million output tokens, compared with prices five times higher before the July reduction.²
The architecture is:
price ↓ → viable workloads ↑ → usage ↑ → inference demand ↑ → compute demand ↑
The model market is usually discussed as if falling token prices must commoditize the model layer. That is only half the equation. The other half is demand elasticity.
Using the reported 80% price reduction and 10x usage increase as a simple illustration, the absolute implied log elasticity is about 1.43.
ln(10) ÷ ln(5) ≈ 1.43
An absolute elasticity above one means quantity expands proportionally faster than price falls.
A model provider can cut the price of intelligence while expanding the amount of intelligence customers buy quickly enough that aggregate spending still rises.
There is an important counterpoint. An NBER working paper by Mert Demirer, Andrey Fradkin, Nadav Tadelis and Sida Peng, using API data from OpenRouter and Microsoft Azure, estimates preliminary short-run price elasticities only slightly above one. The authors argue that this leaves relatively limited room for strong Jevons effects across the market as a whole.³
So Luna's 10x usage increase should not be treated as proof that the entire AI market has a 1.43 elasticity. Nor can the full usage increase necessarily be causally attributed to price. Product improvements, distribution, workload migration and customer adoption may also have contributed.
The model economy may be entering a phase where price per token becomes less informative than workload created per dollar of price reduction.
Suppose inference prices fall another 80%.
If customers simply run the same workloads more cheaply, model economics compress.
But if cheaper intelligence causes companies to automate tasks they previously would not automate, run agents longer, increase parallelism, analyze larger datasets and embed inference into more workflows, the addressable quantity of machine cognition expands.
Then efficiency does not merely reduce cost. It creates demand.
This is the logic behind Jevons' paradox. Improvements in resource efficiency can lower effective prices enough that total consumption rises instead of falls. In AI, the scarce resource is not only tokens. It is the compute, power and infrastructure required to produce them.
That gives us a second architecture:
model efficiency ↑ → inference price ↓ → workloads ↑ → tokens ↑ → compute utilization ↑ → infrastructure demand ↑
The competitive question therefore changes.
It is not simply:
Who can sell intelligence cheapest?
It becomes:
Who can turn falling intelligence prices into the largest expansion of useful workloads while preserving acceptable unit economics?
If demand remains elastic, the model price war may not destroy the inference market.
It may make the market much larger.
And the next binding constraint moves one layer down the stack.
From the price of intelligence to the infrastructure required to supply all of it.

References:
1. Reuters, September 9, 2026. Sarah Friar said OpenAI's 80% Luna price reduction helped drive roughly a tenfold increase in usage. OpenAI offers AI for chip design, Reuters
2. OpenAI, July 2026. GPT-5.6 Luna pricing was reduced by 80% to $0.20 per million input tokens and $1.20 per million output tokens. Advancing the price-performance frontier with GPT-5.6
3. Demirer, M., Fradkin, A., Tadelis, N. & Peng, S. The Emerging Market for Intelligence: Pricing, Supply, and Demand for LLMs. NBER Working Paper 34608, December 2025. NBER Working Paper 34608
4. OpenAI. Current GPT-5.6 Luna model documentation and API pricing. GPT-5.6 Luna API documentation
5. OpenAI. Economic framing on cheaper useful intelligence expanding economically viable work. Building abundant intelligence



Comments