Deep Dive — Tokens halved in price. The bill to make them didn't

Share
Deep Dive — Tokens halved in price. The bill to make them didn't

Something unprecedented happened to the price of machine intelligence in August, and the industry's favorite growth story did not survive it. Silicon Data's LLM Token Spend Index — which tracks what buyers actually pay per million tokens across the market — fell 29% in a single month to $0.97, the first print under a dollar, and roughly half its May 28 peak of $2.07. Three months earlier, the consensus was that smarter models earn more pricing power. Now the flagship models are being discounted twice in a month, and the growth pillar underneath every unlisted lab's valuation is deflating in public.

Here is the part the headlines miss. At the same moment the output price collapsed, the input price went the other way. All of 2027 DRAM and HBM capacity is sold out. Chinese accelerator prices rose 20% to 50% in two months. An RTX 5090 that costs $1,999 at MSRP was trading at $5,997 on September 13 after touching $7,369 the day before. Memory is not a side story — it is the binding constraint on serving a token, and a single HBM4 stack now sells for around $392.

That is the trade nobody is pricing: the AI economy's selling price is falling faster than its cost of goods. The whole bull case depends on the quantity of tokens demanded eventually overwhelming that scissors. The August data is the first hard evidence that it might not do so fast enough.

Close-up of cooling fans in a server room, showcasing technology and efficiency.

What actually got cheaper

An index of average paid price can fall for two reasons, and both are happening at once. Either list prices drop, or buyers migrate to cheaper models so that the mix shifts downward. The current print is both, which is why it moved so violently.

The migration is visible in the routing data. JPMorgan's analysts read OpenRouter's August numbers and found routed token volume up roughly 47% month over month while dollar spend on the platform rose just 7%. The single largest consumer of tokens that month was DeepSeek V4 Flash-0731, at about 50.7 trillion tokens routed through it. Half of the platform's top ten models by volume were cheap Flash variants. When agents and developers route almost half their new volume to the discount tier, the average paid price falls even if nobody's list price moves.

The list-price cuts are real too, and they are being made through three distinct levers rather than one.

The first is headline price. OpenAI cut GPT-5.6 Luna permanently by 80% on July 30 and took 20% off Terra in the same release, then roughly a month later cut input prices on its own flagship GPT-5.6 Sol by 20% and its output prices by a third. DeepSeek, after raising prices as much as 12x in mid-August, reversed course on September 10 — retiring V4 Pro entirely and rerouting those requests onto V4.1 Flash at Flash pricing.

The second lever is the cleverest, and it is not a discount in any traditional sense: cache pricing. Anthropic's Claude Fable 5.1 left its headline input and output rates alone and cut cached reads 75%, from $1 to $0.25 per million tokens. The engineering reason is straightforward. Most of the compute in a long request is spent in the prefill pass — reading the whole prompt, running attention across all of it, and writing a KV cache into memory. A cache hit skips that work almost entirely, so the marginal cost of re-reading context the provider already computed is near zero. Undercutting cache prices is therefore nearly free margin-wise for the seller while directly subsidizing the fastest-growing category of buyer: agents, which may re-read the same long context dozens of times per task. DeepSeek took the same route from the hardware side, compressing KV cache footprint through architecture, attention mechanism and cache precision choices so that storing context costs less in the first place.

The third lever is model substitution, which is where the "Flash" phenomenon matters. Zhipu's GLM-5.3-Flash and Alibaba's Qwen3.8-Flash shipped the same day; the Qwen variant reportedly trained for about one-ninth the cost of its flagship. These are not the best models anyone has. They are good enough for the task at hand at a fraction of the price, and for most agent workloads "good enough" is the entire requirement.

One vendor's approach deserves separate mention because it points the other way: OpenAI has also been shortening output, so that a task consumes fewer tokens even after a rate increase. That is a genuine efficiency gain wearing the costume of a price rise, and users of GPT-6 Astra noticed their quotas draining faster than before the optimization. Cost-per-outcome and cost-per-token are no longer the same number.

The input side is not cooperating

Cutting price while your bill rises is a decision, not a strategy, and this is where the quarter's data gets awkward.

The memory shortage is the mechanism. AI demand has booked the entire 2027 output of Samsung, SK hynix and Micron through long-term agreements, some running five years, and HBM is the component that sets how fast a model can serve tokens. That scarcity is now reaching all the way up the stack: Nvidia has reportedly told top customers to expect increases around 15% from early 2027, the RTX PRO 6000 Blackwell doubled to $16,000, and Amazon raised Echo, Kindle and Fire TV prices by as much as 60% in August on memory and storage costs. Consumer devices, mid-tier laptops and frontier APIs are all bidding for the same wafers.

So the two curves point in opposite directions: a million tokens for $0.97, and the machines that produce them costing more every month.

The most informative refusal came on Monday morning from the least sentimental sources. Samsung Electronics and SK hynix both declined a request from Korea Electric Power Corp. to prepay roughly 25 trillion won — about $18.7 billion — in electricity, worth around five years of bills at last year's levels. KEPCO had asked Samsung for 20 trillion won and SK hynix for 5 trillion won to front-load financing for the semiconductor clusters under construction at Yongin and in the Honam region, and both companies completed internal reviews and said no. Their reasoning, as relayed by Reuters' sources, was that committing that much cash is hard to defend while long-term AI silicon demand is genuinely unknown. KEPCO's own position explains the ask: it ended June with 210.7 trillion won of debt and roughly 11.5 billion won a day in interest payments.

The two companies that sell the memory every AI accelerator needs looked at a five-year electricity bill for capacity they themselves are building and declined to underwrite it. Meanwhile SoftBank — OpenAI's most exposed backer — closed a bridge facility sized at $11.87 billion against a $10 billion target from around 20 lenders. Oversubscribed, and two years long. Capex on that scale is normally financed against a decade of assumed demand.

Put the three numbers in one line and the shape of the market's worry becomes legible: the token price is falling, the hardware price is rising, and the money on offer runs to roughly one-fifth of the horizon the buildout assumes.

Where the squeeze lands

For a listed chipmaker, deflation in the application layer is someone else's problem — indeed the memory suppliers have been the boom's best-paid passengers, with Samsung expecting 90 trillion to 110 trillion won available for shareholder returns this year. For a model lab, it is the balance sheet.

Both American frontier companies are private, both have told investors that token consumption growth is the core of the story, and both have pricing power they no longer demonstrably hold. DeepSeek's August reversal is the tell that the price war is not a subsidy war: it raised prices as much as 12x because the cost curve demanded it, then cut them back within 24 days because the market would not pay. Every discount now has to respect a cost line, which is exactly why the cache and Flash plays replaced blunt rate cuts — you can only discount what is genuinely cheaper to serve.

There is one disclosed figure that shows what the labs are absorbing rather than passing through. OpenAI has said it dedicates 20% of the compute it would otherwise spend on inference to monitoring reasoning traces in real time — a safety overhead carried on its own account, described in The safety compute tax: OpenAI's 20% overhead and the price of alignment. That is cost inflation from a direction nobody forecast, and it does not show up in a token price.

China, which expects to burn 100 quadrillion tokens this year, has already built a financial layer on the assumption that token consumption is a stable measure of economic activity. Beijing's E-Town district pays up to 50% of an enterprise's token spend under its Token Ten Measures, four ten-thousand-card token factories are planned by 2030, and Bank of China's Haizhu branch has begun writing three-year unsecured loans against a company's token bill — monthly consumption, compute contracts, API call logs — instead of real estate, in the invention we covered as Bank of China's first 'token loan' ties credit to AI usage. That is the first financial instrument priced directly off model usage, and the dollar falling under $1 is its first stress test: a borrower's collateral value is now partly a function of unit prices that moved 29% in a month.

What the skeptics say

The strongest counter-argument is that "intelligence got 50% cheaper" is a category error. The index measures a blended average, and a large part of the fall is composition — buyers moving to smaller models, not paying less for the same capability. Compare like with like and flagship rates have barely moved; Claude Fable 5.1 still costs what it cost, and the cuts landed on cache reads and mid-tier SKUs. On this reading there is no collapse, just a market discovering that most tasks do not need the best model available, which is a maturation, not a crisis.

The second objection attacks the Jevons framing itself. The 19th-century observation was that better steam engines increased total coal consumption, and the AI industry borrowed it as: cheaper tokens, more tokens, bigger revenue. The August numbers look like the paradox failing — volume plus 47%, spend plus 7%. But the demand curve has a visible reason to lag. The heavy consumers are agents and enterprise workflows, and enterprises spent the summer installing brakes rather than accelerators: Uber's CTO described per-engineer API costs of $500 to $2,000 a month and burned its full-year AI tools budget by April, after which Meta and Amazon both imposed monthly token caps on staff. Uber's COO Andrew Macdonald went further on the record, saying the link between tokens spent and features shipped is hard to establish. That is not absent demand. That is demand that has not yet been proven worth paying for, and it is the exact gap Goldman's One-Delta head Rich Privorotsky described when he said the industry has moved from asking whether the technology is feasible to asking whether the cost can be sustained.

The steelman of the bull case: deflation at the token layer is what makes the application layer viable. An industry being financed at frontier-lab scale cannot sell intelligence at premium rates when a single agent task consumes millions of tokens, and every serious agent deployment this summer hit the same wall — models cheap enough to demo and too expensive to run at scale. Falling token prices may be the precondition for the demand that eventually justifies the capex.

The honest position is that both things are true in the same quarter, and the market has started marking the difference. Global AI-linked stocks sold off on Monday after Anthropic's Dario Amodei called for pacing the frontier, with SK hynix down more than 6%, SoftBank 10% in Tokyo, ASML off more than 4%, Intel nearly 6% in U.S. premarket. RBC Brewin Dolphin's Zoe Gillespie named the mechanism: the equity rally is built on AI growth and productivity gains, and if that derails, "we may see this destabilize." The pacing debate gave investors a reason to sell. The token index gave them a number to sell against.

What to watch, in order

The September token-spend print, once DeepSeek's rerouting and Anthropic's cache cut land in the data. Holding under $1 with volume still climbing would confirm this is a structural repricing, not an August blip. Morgan Stanley's view — reported by Chinese financial media this month — is that the war stays orderly rather than becoming a race to the bottom; the September number is how you test that.

Whether anyone raises list prices again. DeepSeek did in August and reversed in 24 days. A second attempt, by anyone, would be the clearest available signal that unit economics have turned genuinely negative, because a lab only re-tests a price ceiling once it believes it can hold.

The memory pass-through. Watch whether Nvidia's reported mid-teens 2027 increases stick, whether HBM4 stack prices hold around $392, and whether the RTX 5090 stays at triple MSRP — consumer GPU pricing is the most transparent read on where the shortage ends and speculation begins. For the mechanics, our plain-English version is in AI 101 — What is HBM (high-bandwidth memory)?.

The power queue. KEPCO's Yongin schedule now has to advance without $18.7 billion of prepaid bills, and the same constraint repeats in every grid region negotiating a data-center connection.

And the first token-loan problem. Guangdong's token-backed book is small — around ¥28 million of approved lines in public reports — but it is the first credit product whose collateral is a usage bill. The first case of inflated or resold token consumption, or the first borrower whose collateral value halves because a price war moved 29% in a month, will teach everyone how liquid intelligence actually is as an asset.

What we now know is the price of a token. What nobody has priced is whether the cost side ever stops rising.

If a million tokens cost 97 cents and the GPU that serves them doubled in price, is AI getting cheaper or just getting subsidized differently? Tell us in the comments.

Sources: Huxiu / Lightcone Intelligence on the token spend index · CNBC — AI stocks slide after major CEOs urge slowdown · Reuters — Samsung, SK hynix reject KEPCO power prepayment · Techmeme — Silicon Data LLM Token Expenditure Index at 97 cents · videocardprices.com RTX 5090 tracker · Japan Times — SoftBank's upsized $11.9B loan