The Jevons Trap
- Hurratul Maleka Taj
- 3 days ago
- 11 min read
Updated: 1 day ago
Why the AI Labs Win Volume and Lose the Margin.
AI demand will compound for a decade. The surplus won't reach the labs producing it.
The AI industry is having the wrong argument confidently.
The argument on the table is whether demand for intelligence keeps compounding as the technology gets cheaper and better. It does, and it will. That much is settled, and the people who settled it are right. My problem isn't with the answer. It's that this is the wrong question to feel certain about, because it only tells you how big the prize gets. It says nothing about who walks away holding it.
Here's the question I'd rather have certainty on, and it's an older and harder one. When a resource gets radically cheaper and everyone starts using enormous amounts of it, who actually captures the money? The historical record is blunt about this, and it is not kind to the people everyone is currently betting on. Cycle after cycle, the party that produced the newly abundant resource captured the least of the value it unleashed. The surplus flowed straight past them to whoever was sitting closest to demand.
That's the whole thesis. The Jevons Trap. The labs are cheering the exact mechanism that turns them into a commodity. The demand explosion they're right about is the same force that routes the margin away from them, and "we won the demand argument" and "we lost the margin war" are not opposite outcomes here. In a market where your product becomes interchangeable, they're the same sentence.
Here's the argument, layer by layer.
1. The demand side is settled. It's worth restating properly anyway.
Jevons' Paradox, named by William Stanley Jevons in 1865, is the observation that making a resource more efficient to use tends to increase total consumption of it rather than shrink it. Efficiency drops the effective cost of whatever work the resource does; demand for that work turns out to be wildly elastic; and the flood of new usage overwhelms the per-unit saving many times over. Cheaper doesn't mean less. It means vastly more.
The sharpest recent version of this argument for AI came from Tomasz Tunguz: https://www.linkedin.com/pulse/racing-sustain-jevons-paradox-tomasz-tunguz-g9d3c/
He has put the current pricing moment squarely in Jevons terms. His point: segmentation, a tiered market running from an expensive frontier down to near-free value models, is what keeps the paradox alive even while headline prices climb. I think he's right about the demand side, and the segmentation read is the correct one. So I'm going to take it as given and build from there.
But there's a measurement trap sitting inside the demand story, and most of the analysis walks straight into it, and clearing it is what opens the door to the question that actually matters. The supply side, though, is already splitting: open models now hold roughly a third of all tokens, so the cheap value floor this whole argument turns on isn't a forecast, it's already here.

Figure 1: Open vs. Closed Source Model
2. Nominal price is the wrong ruler
Jevons was never about the sticker price of coal. It was about the effective cost of the service coal delivered, once you run it through efficiency. That isn't a pedantic distinction. It's the entire game.
Nominal price is the number on the rate card: dollars per million tokens. Effective cost is what you actually pay to finish a real piece of work, and it quietly absorbs three things the rate card never shows you, how many tokens the task burns, how efficiently the model chews through them, and how many attempts it takes before the answer is usable. A model can double its price per token and still lower your cost-per-finished-task, provided it got more than twice as good at finishing the task.
Which is why rising rate cards don't threaten Jevons on their own. The frontier sticker can climb while the real cost of getting work done keeps sliding, pushed down by better tokenization, longer and cheaper context, and the value tier falling toward the floor. The paradox lives on the effective-cost curve. That curve still points down.
So when the charts show flagship prices repricing upward, the honest reading isn't "Jevons is breaking." It's narrower and stranger: the sticker and the effective cost have come apart, and only one of them was ever relevant to the paradox.
3. The price rise is margin, not cost
If the effective cost of work is falling, then why are the labs raising nominal prices at all? Here's where you have to be careful.
The obvious answer is "their costs went up." And to be fair, capex genuinely is going up, setting records every quarter, and it'll keep climbing as models scale. Concede that plainly. But capex is a fixed cost. It is not the marginal cost of serving a token, and those are very different animals. The marginal cost of the next token is close to nothing, mostly electricity. You build the datacenter once, and you're on the hook for it whether you serve a single token or a trillion.
So the move the labs are making is really a decision about how fast to earn that fixed investment back. They can stretch recovery across years of compounding volume, or they can lift price now and recover faster at fatter margin. Picking the second is a pricing-power call, not a cost pass-through. It's the pivot the vendors are themselves signaling out loud, cut when cash is plentiful and share is the prize, raise when cash is tight and margin is the prize.
That reframes segmentation entirely. Tiering isn't the hero rescuing Jevons. It's the price-discrimination surface that lets a lab recover capex and skim margin at the top while the value tier quietly keeps demand compounding at the bottom. You can watch how strained the top has gotten: the premium frontier now runs something like 13x the cost of the value tier for maybe 20% more measured intelligence, while the floor has collapsed to roughly three cents per million tokens. That's not a demand story. It's a margin structure.
And that's the hinge. Once you read the price rise as margin extraction rather than cost inflation, exactly one question is left standing, and almost nobody is asking it: who is actually positioned to extract that margin?
4. Jevons grows the pie. It has nothing to say about who eats it.
Jevons is a theory about how much gets consumed. It tells you consumption of a cheapening resource goes up. That's it. It has nothing to say, by construction, about value capture, about which layer of the chain gets to keep the surplus all that new consumption throws off.
For that question, there's a better framework: Ben Thompson's Aggregation Theory: https://stratechery.com/2015/aggregation-theory/

Figure 2: How value moves when distribution goes free
Once distribution and transaction costs fall to roughly zero, value stops pooling around whoever controls supply and starts pooling around whoever owns the demand relationship and can turn suppliers into interchangeable parts. You win not by making the thing but by owning the customer and commoditizing everyone who makes the thing. Google did it to publishers. Amazon did it to sellers, then to book publishers specifically. Netflix did it to studios. Airbnb did it to hotels. Each time, the maker of the underlying good got demoted to a swappable input, and the aggregator that owned demand kept the surplus.
Now apply it to AI, and use Tomasz's segmentation premise as the input:
- Once models are tiered and substitutable across price bands, they are, by definition, modularized suppliers. Interchangeability is the one thing Aggregation Theory needs to function, and segmentation is precisely what manufactures it.
- The layer sitting between the buyer and those interchangeable models, deciding which one handles each request, is the router. As Tunguz himself notes, the router is becoming the strategic layer between buyer and model. In Aggregation terms, that's the layer that owns demand.
- From there the framework doesn't guess, it predicts a direction: the router keeps the surplus, and the models get commoditized underneath it.

Figure 3: The AI Value Chain, Aggregated
5. It has happened almost every time. The producer rarely keeps it.
This isn't a hunch. It's the base rate.
Look at the last five times a market got aggregated. Search, e-commerce, streaming, ride-hail, lodging: in all five, the producer supplied the value and someone closer to demand captured it.

Figure 4: The producer rarely keeps the surplus
None of these suppliers were weak businesses. Hotels, studios, major brands, these were formidable incumbents with real assets and real leverage. They got commoditized anyway, because the instant their output became discoverable and swappable through a layer that owned the customer, pricing power slid one notch closer to demand and stayed there. The producer kept the operational headache. The aggregator kept the margin.
Now, the important caveat. It isn't a universal law. Producers who refused to become interchangeable, or who built a direct line to their own demand, held onto their pricing power. Nvidia sits upstream of the whole AI boom and captures enormous margin precisely because it is not substitutable. Luxury brands starved Amazon of their catalog rather than be modularized by it. Disney and HBO clawed back from Netflix by owning the content and, more importantly, the subscriber. So the rule isn't "the producer always loses." It's sharper than that: the producer loses when its output becomes interchangeable and someone else owns the customer. Escape either condition and you escape the trap. Hold that thought, because it's the entire back half of this piece.
The catch for the labs is that they walk into this pattern as the most capital-intensive suppliers it has ever seen. That doesn't spare them. It raises the stakes, because gigantic fixed costs sitting on top of an increasingly commoditized output is the exact profile that gets caught in the squeeze: forced to keep spending just to hold the frontier, unable to price for that spend because a cheaper tier is always one routing decision away.
And in AI, that value tier is already the volume. On OpenRouter, one open supplier, DeepSeek, has served more tokens than the next several providers combined, more than seventeen times Google's open-model volume over the same year.

Figure 5: Total token volume by model author
6. The Jevons Trap, said plainly
The labs are right that demand compounds. They're right that segmentation sustains it. But segmentation is the thing turning them into interchangeable suppliers, and interchangeable suppliers sitting behind a demand aggregator don't get to keep the surplus their own abundance creates. The mechanism they're celebrating, cheap intelligence pulling explosive usage across every tier, is the same mechanism quietly moving the margin up to the layer above them. They are winning the demand argument in the specific way that loses them the margin war.
Which reframes something we're all already watching. Why are the major labs sprinting to field an entry in every single tier, premium, mid-market, value, when the value-tier economics are genuinely miserable? Read through the trap, it stops looking like expansion and starts looking like defense. Each lab is trying to become its own router, to swallow the aggregation layer before an external one forms and disintermediates it. The three-tier scramble is a fight not to get cut out of your own distribution. Filling all three tiers is just what it costs to avoid being modularized by somebody else's router.
That's a very different story from the victory-lap version. It's the story of producers running hard to dodge the fate producers almost always meet in this pattern.
This isn't only theory. It's already visible in the data. In a 100-trillion-token study of real usage across models, a16z and OpenRouter found the market splitting exactly along these lines: closed frontier models cluster in the high-cost, high-usage corner while open models occupy the low-cost, high-volume floor, and the demand curve between them is nearly flat, a 10% price cut moves usage by well under 1%. Open-weight models already carry roughly a third of all tokens, and a single value-tier supplier, DeepSeek, served more tokens than the next several providers combined, more than seventeen times Google's open-model volume. The commoditization the trap depends on isn't a forecast here. It's a measured fact. What the same study is careful to say, and so am I, is that this hasn't yet compressed the labs' pricing power: premium models still command a premium, and the study argues the market hasn't fully commoditized. That gap, commoditized tokens but intact margins, is the exact tension this piece is about. The question isn't whether segmentation is real. The data settles that. The question is who ends up holding the margin once the router sits between every buyer and every model.

Figure 6: Log cost vs. log usage by category
7. Where this breaks, which is also where the escape is
Here's where the Jevons Trap actually breaks. It has three escape routes.
If frontier capability stays genuinely non-substitutable, the trap never closes.
Aggregation needs interchangeable suppliers to work. If one lab holds a lead big enough that swapping its model out is something the user immediately feels, that lab isn't a modularized supplier and no router can commoditize it. So the whole thing hinges on one empirical question, does frontier intelligence converge toward interchangeability, or stay differentiated? Tomasz's segmentation thesis is, underneath, a bet on convergence. The labs' capability arms race is a bet against it. Watch that variable before any other.
If a lab owns demand directly, it stops being the supplier and becomes the aggregator.
The way out of getting modularized is to be the one doing the modularizing. A lab with a genuinely dominant consumer surface, hundreds of millions of people showing up directly, every day, isn't sitting behind a router. It is the router. That's why the consumer app layer matters far more than a pricing chart makes it look. It's the whole difference between being Airbnb and being a hotel.
If routers never achieve winner-take-all scale, the surplus just scatters.
Aggregation only captures value when the aggregator gets real scale economics on the demand side. If routing stays cheap and undifferentiated, a dozen interchangeable routers, no lock-in, then no aggregator ever forms and the surplus disperses instead of pooling anywhere. The genuine bull case for the labs is a world where routers stay weak.
Notice none of these are comfort for the labs. They're a task list. Stay non-substitutable, or own demand outright, or make sure no router ever reaches aggregation scale. Every one of them is hard, and right now the market structure is drifting against all three at once.
8. So what do you do with it
If you're a founder: the durable margin in AI probably won't sit where the capital is. It'll sit at the layer that owns the demand relationship and gets to treat models as swappable inputs. The orchestration and routing layer isn't plumbing, it's the aggregation surface, and it's underpriced right now precisely because everyone is transfixed by the models. Own one defensible point on the demand relationship that a capital-intensive lab can't economically match, and you're building the aggregator, not the supplier.
If you're an investor: stop leading with "how good is the model." Lead with "who does this business modularize, and who could modularize it?" A model lab with no demand surface of its own is a supplier in a pattern that has rarely been kind to suppliers. A router that reaches real demand-side scale is the aggregator in a pattern that has rarely failed to reward them.
If you're a lab: the tell is already in your own behavior. The rush to fill every tier and to build direct consumer surfaces is exactly right. It's an attempt to escape by becoming the aggregator instead of staying the supplier. Whether any of the labs pulls it off is the real contest of the next five years, and it gets decided on the demand relationship, not the benchmark.
Jevons will hold. The pie grows for a decade. The only question that ends up mattering is who's standing closest to demand when it does, and history is unusually clear that it won't be whoever produced the resource.
That's the trap. The labs are right about nearly everything, except the part about who wins.
To sum up, Jevons tells you the pie grows; Aggregation Theory tells you who eats it; and in a segmented, router-mediated market, it isn't the labs. The demand win and the margin loss are the same event.
Sources:
Jevons' Paradox. William Stanley Jevons, The Coal Question, London: Macmillan, 1865.
Tomasz Tunguz, "Racing to Sustain Jevon's Paradox," 2026. https://www.linkedin.com/pulse/racing-sustain-jevons-paradox-tomasz-tunguz-g9d3c/
Aggregation Theory. Ben Thompson, "Aggregation Theory," Stratechery, July 2015. https://stratechery.com/2015/aggregation-theory/
Conservation of Attractive Profits. Clayton M. Christensen and Michael E. Raynor, The Innovator's Solution, Harvard Business School Press, 2003 (the framework Thompson builds on).
Empirical usage data. "State of AI: An Empirical 100 Trillion Token Study," a16z and OpenRouter, December 2025. https://openrouter.ai/state-of-ai (platform data, November 2024 to November 2025)
Model pricing and intelligence figures. Artificial Analysis Intelligence Index v4.1, accessed August 4, 2026, as compiled by Theory Ventures.
Tomasz Tunguz, "AI Model Inflation," Theory Ventures. https://tomtunguz.com/ai-model-inflation/
Figures:
Figure 1, "Open vs. closed source models." a16z and OpenRouter, State of AI: An Empirical 100 Trillion Token Study, December 2025.
Figure 2, "How value moves when distribution goes free." Diagram created by me, after Ben Thompson's Aggregation Theory (2015), building on Christensen's Conservation of Attractive Profits.
Figure 3, "The AI value chain, aggregated." Diagram created by me. Pricing and intelligence figures from Artificial Analysis Intelligence Index v4.1, per Theory Ventures, accessed August 4, 2026.
Figure 4, "The producer rarely keeps the surplus." Diagram created by me. Historical cases after Thompson, Aggregation Theory (2015); exception rows my own analysis.
Figure 5, "Total token volume by model author." a16z and OpenRouter, State of AI: An Empirical 100 Trillion Token Study, December 2025 (data November 2024 to November 2025, on OpenRouter).
Figure 6, "Log cost vs. log usage by category." a16z and OpenRouter, State of AI: An Empirical 100 Trillion Token Study, December 2025.



Comments