Deep Dive — Nobody can price what an agent buys

Share
Deep Dive — Nobody can price what an agent buys

The Financial Times built an interactive piece this week around a suspicion enterprise buyers have nursed all year: the token is losing its claim to be the unit at which intelligence is sold. "Powerful AI tools burn through budgets," its framing reads, "prompting a rethink of how the technology is priced ahead of frontier lab IPOs." Two numbers inside it set the stakes. SemiAnalysis estimates that almost 90 percent of Anthropic's business now comes from agentic AI — coding tools and autonomous workflows, not chat. And according to people familiar with the company's IPO prospectus, almost a quarter of Anthropic's 2025 revenue came from just two clients. The most talked-about AI company of the moment is heading toward the largest IPO in history with nearly all of its revenue attached to a billing unit that neither it nor its customers can predict.

That is the story under the story. The last two years of AI economics were built on a simple promise: pay for what you use, by the token, the way you pay for electricity. Agents broke the promise from the demand side — they consume without a human in the loop — and the supply side answered by unmooring prices even further. What is left is an industry about to sell equity against revenue streams it meters, discounts and forecasts with tools it increasingly does not trust.

The meter everyone inherited

Token billing worked while the customer was a person. A chatbot answer had a readable size, a human checked it, and the bill scaled with use in a way finance departments could model. The unit even has an academic champion: Yale's Aleh Tsyvinski, who analyzed usage by people and companies consuming more than 380 trillion tokens across 400 different models, compares the token to the kilowatt-hour — "the basic building block of AI," as he put it to the Financial Times.

The economics underneath the meter moved fast anyway. Context windows went from GPT-3.5's 4,096 tokens in 2022 to the million-token standard of current models, and Epoch AI's research puts the price of a given level of performance falling roughly thirteen-fold a year since 2023. Both facts point in opposite directions from what buyers actually experienced, because usage grew faster than unit prices fell. Then, earlier this year, the leading labs made a structural change that barely got headlines: enterprise customers were moved off fixed subscriptions and onto uncapped usage-based pricing. The Financial Times reads that shift as a transfer of risk — longer, open-ended workflows now bill to the customer, and Uber, Meta, Cisco and Walmart all responded by installing spending caps.

What the agents actually burn

The cap story is where the abstraction turns into line items. Uber's chief technology officer, Praveen Neppalli Naga, told the Financial Times that the company had exhausted its full-year AI spending barely four months into 2026: "We budget based on what we know, but things are changing fast in this space. Every week there's a tectonic shift." Uber's answer was engineering, not negotiation — token budgets of roughly $1,500 a month per AI coding tool, model-switching through OpenRouter to route work to cheaper options, and cached context. Usage still grew ninefold since February, a figure the company gave investors in August; by the CTO's account spending has since stabilised, with the budget discipline treated as an engineering problem rather than a procurement one.

Not every blowout ends that neatly. One Amazon project ran 860 percent over budget after incomplete tasks quietly kept working in the background for five months, costing more than $1 million, according to people familiar with the matter cited by the Financial Times. Ramp's analysis of 70,000 businesses found average monthly AI spend per employee among the heaviest-using percent of companies more than doubling, to $7,200. And Amazon's internal culture has already produced the era's first accounting hallucination: a bout of "tokenmaxxing," in which staff faked activity to climb usage leaderboards.

The underlying reason bills resist forecasting is now documented. OpenAI's enterprise data, as reported by the Financial Times, shows that as of June its Codex agent was used by just 17 percent of customers yet generated 63 percent of all output tokens across ChatGPT and Codex combined. A study by researchers from Stanford, MIT, Google DeepMind and Microsoft found coding agents consumed about a thousand times as many tokens as simpler coding interactions — and that across repeated attempts at the same task, the most expensive run typically burned twice as many tokens as the cheapest, in some cases thirty times as many. The models themselves are bad estimates of their own cost, the authors note; agents, in the words of the University of Texas at Austin's Jiaxin Pei, one of the paper's authors, are "not trained to be budget-aware." Goldman Sachs forecast in May that agentic demand will drive a twenty-four-fold increase in global token consumption by 2030.

We have been charting this shift from the supply side since August, when AI agents now burn more tokens than humans on OpenRouter. The pricing story is the same shift seen from the invoice.

Close-up of professionals reviewing financial graphs at a business meeting.

The labs' receipts ride on this meter

Here is why a billing curiosity became an IPO question. If almost 90 percent of Anthropic's business is agentic — SemiAnalysis's estimate, relayed by the Financial Times — then nearly all of its revenue comes from the workload with the worst forecasting record. The growth behind it is real and spectacular: Anthropic's annualized revenue went from $9 billion at the end of 2025 to $65 billion in July, with backers forecasting more than $120 billion by year-end, figures the Financial Times has previously reported and we tracked in Anthropic hits $65B run rate, up $18B in two months. But run rates are extrapolations of metered sales, and the prospectus already concedes how narrow the base is: two clients carried almost a quarter of 2025 revenue, a concentration we put in context alongside the company's fixed commitments in Anthropic promised $518 billion. Most of it is owed either way.

Joey Brookhart, the SemiAnalysis analyst who covers the labs, put the forecasting problem in the bluntest terms the Financial Times printed: "At this point, if you are looking at 2027, no one knows what they're making." His named risk is a release — a strong new model from Google or Meta that forces prices down — which would prompt "a much different conversation about these AI labs and the IPOs very quickly." Meanwhile the price moves keep undermining the meter's stability: Anthropic's Claude Opus 5.5 launched at 40 percent less than its predecessor, and OpenAI's new Sol and Luna versions came in about 50 percent cheaper than the iterations they replaced. Cheaper units, unpredictable quantities, a customer base installing caps — the three variables a revenue forecast needs are all moving at once.

The bid for a new unit

The labs' own executives now say the unit is wrong. "Pricing in tokens doesn't make any sense," OpenAI president Greg Brockman said at the September launch of the Astra model. "Price per task is what matters. Can you get the thing done at an appropriate price and appropriate speed?" OpenAI is testing exactly that, following Salesforce and Sierra into outcome-shaped billing, and the Financial Times reports the direction of travel among researchers as well: NYU Tandon's Quanyan Zhu argues current usage pricing judges employees by time in the office rather than results, and expects more contracts to move to completed tasks or agreed outcomes. OpenAI's chief economist Ronnie Chatterji offered the cleanest formulation of the discontent: "This idea that all tokens are the same is just an artefact of us being able to count them."

There is a deeper version of that argument in the data. The same FT analysis notes that Codex's 17-percent-customer, 63-percent-of-tokens split means tokens measure where compute goes, not where value lands — a lawyer's finished contract review and a model's meandering path to it are billed in the same currency even though only one of them is what anyone bought. Anthropic's near-total bet on agentic work makes this more than philosophy: its revenue recognition depends on whether customers come to see a completed task as worth paying for, or as an unpredictable cost to be capped.

What the skeptics say

The bull case has not gone away, and it rests on volume. Proponents argue that falling prices plus faster-growing usage equals rising revenue — the Jevons pattern the industry has leaned on all along — and that cheap tokens are the precondition for the agent demand that justifies the data centers. The caching data supports part of it: as we noted when agents overtook humans on OpenRouter, roughly 70 percent of agent token consumption comes from cached prompts billed at deep discounts, so real dollars climb far more slowly than raw volume. Uber's stabilized spend is the corporate version of the same hope: once you treat budgets as an engineering problem, they become tractable.

Skeptics of the per-task answer counter that every new unit becomes the new target. Amazon's tokenmaxxing episode shows what happens when a metric pays; "complete the task" invites corner-cutting on quality in ways token counts at least made visible. And the transition itself is a hazard: two frontier labs changing billing models mid-cycle, both privately traded in intent if not yet in fact, makes comparability between quarters — and between the labs — harder exactly when investors will be judging them on it.

What to watch

  • Whether per-task pricing ships before the roadshows. Anthropic's S-1 must go public at least 15 days before its roadshow, expected after the November midterms. If the filing describes revenue in tokens while the CEO talks about tasks, the mismatch is the story.
  • Whether enterprise caps loosen or spread. Uber's stabilization is one data point; Ramp's 70,000-business dataset is the place to watch whether the heaviest percent of spenders keeps doubling or plateaus — and whether caps appear at companies that haven't hit a blowout yet.
  • Whether the disclosed concentration widens. Two clients at a quarter of revenue is a 2025 figure; any update in the public filing tells you whether the agentic boom diversified the base or concentrated it further.
  • Whether a hyperscaler release blinks the market. Brookhart's trigger — a Google or Meta model strong enough to reset prices — now has a visible precedent in Gemini 4 Argon's introductory pricing, which lands below its frontier rivals.

The token will not disappear; it is too useful as an accounting unit, and someone always has to count something. But the industry is moving toward a split the FT piece makes plain: meters for the machines that consume, tasks and outcomes for the customers who pay. Frontier labs used to sell intelligence by a unit nobody questioned. They are now preparing to sell stock while questioning it themselves.

If the token is the wrong unit for agents, would you rather pay per task, per outcome, or per token you can at least count? Tell us in the comments.

Sources: Financial Times — AI agents are rewriting the economics of computing · SemiAnalysis · Reuters — Anthropic's IPO prospectus shows surging costs and client concentration · Stanford/MIT/DeepMind/Microsoft study on coding agents' token consumption (arXiv) · Ars Technica — New Anthropic, OpenAI models make the same promise: a little more for a lot less money