The average price of an LLM token just fell below $1
Two charts the industry doesn't usually read together — one on API prices, one on GPU shelves — both say the same thing: the AI cost curve is bending, and not in the direction the growth story needs.
The market price of a million LLM tokens has fallen below one dollar. Silicon Data's LLM Token Spend Index, which tracks the average paid price per million tokens across the market, dropped 29% in August to $0.97 — the first print under the dollar mark, and roughly a halving from the $2-plus level it held in May. The collapse is both a real price war and a mix shift: Goldman's One-Delta head Rich Privorotsky frames the industry as moving from "is the technology feasible" to "can the cost be sustained." The evidence is in the routing. DeepSeek cut its Flash series on September 10 (24 days after its own August 17 hike of up to 12x) and retired V4 Pro outright, rerouting those requests to V4.1 Flash at Flash prices; Claude Fable 5.1 dropped cache reads 75%, from $1 to $0.25 per million tokens; OpenAI has cut GPT-5.6 Luna 80% and then repriced its own flagship twice in a month. Per JPMorgan's read of OpenRouter, August routed volume grew about 47% month over month while dollar spend rose just 7% — and half of the platform's top ten models by tokens were cheap Flash variants. That's the Jevons paradox stalling: usage explodes, revenue doesn't follow. We've tracked the demand side of this story all summer — China expects to burn 100 quadrillion tokens this year and Zhipu's API revenue jumps 27x — but price deflation outrunning volume growth is a new number, and Morgan Stanley's take is that the war will stay orderly rather than turn into a race to the bottom. The uncomfortable implication: token economics, the pillar of every unlisted lab's valuation, is deflating in public.
The RTX 5090 is nearly gone from shelves, and local-AI builders are watching the price double. r/LocalLLaMA reporters say Micro Center is sold out down to leftover $14,000 RTX PRO 6000 cards, and the trackers agree: videocardprices.com has the 5090 at $5,997 as of September 13 — up 17.6% in under a month, 200% over the $1,999 MSRP, and it touched $7,369 on September 12. The driver is the same memory squeeze behind the data-center buildout: GDDR7 cost inflation plus resale demand from buyers modding or hoarding VRAM-class cards, including unverified listings for 96 GB versions that commenters in the thread call scams. For anyone running inference at home, the math now favors the API — which, given the block above, is cheaper than ever. We covered the API-price floor here, and the 33B model that matched DeepSeek V4 Pro on free tokens is the other reason a $6,000 GPU is a harder sell this week.
What to watch: whether September's token-spend print holds under $1 after DeepSeek's and Anthropic's cache cuts land in the index.
Would you pay $6,000 for a 5090 today, or just rent the tokens? Tell us in the comments.
Sources: Huxiu (Lightcone Intelligence) on the token spend index · Appinn on DeepSeek's Flash price cut · Tencent News on Morgan Stanley's price-war read · videocardprices.com RTX 5090 tracker · r/LocalLLaMA 5090 stock thread