AT&T swings 40% of its AI calls to open models to cut costs

Share
AT&T swings 40% of its AI calls to open models to cut costs

AT&T is quietly remaking how 100,000 employees use AI, and the lesson for every company staring at a frontier-lab invoice is that open-source models are now the default, not the fallback. VP Mark Austin told reporters the telecom plans to keep its spend on Anthropic and OpenAI's closed models flat over the coming years, and the lever is routing more internal traffic to open and open-weight models like NVIDIA's Nemotron, Meta's Llama, and Google's Gemma.

The numbers are the story. Around 40 percent of employee AI calls now run on open models, a share AT&T aims to push to 60–70 percent over the next few years. Its Ask AT&T system alone processes roughly 45 billion tokens a day — code generation, HR lookup, sales call summaries, customer-support retrieval — an engine that, if fully hosted on closed frontier models, Austin estimates would cost over $100 million a year at published pricing. Austin says the open models he's tested match older Anthropic and OpenAI releases "or do better," and he puts the gap between open weights and the frontier at 6–10 months and shrinking.

Cost routing is doing the heavy lifting. AT&T deploys model-routing that assigns each request to the cheapest model that can do the job — summaries go to open models, hard codegen still gets the top closed models — and after plugging in the LiteLLM routing layer, it cut the cost of advanced coding tasks by up to 56 percent while output quality fell just two percent. The company also runs some open-model workloads on its own data centers with NVIDIA and AMD silicon rather than renting cloud GPUs. It's separately evaluating DeepSeek and Moonshot as they weigh China's open-source models, but hasn't put either into production yet.

This matters well beyond one carrier. It's a public, dollar-signed vote that availability of cheap open models is eroding the pricing power OpenAI and Anthropic bank on right as they slug it out in a very public price war — OpenAI just cut GPT-5.6 Sol API pricing by over 20 percent. We flagged that fight this morning in OpenAI's price war gambit is a warning shot at Anthropic's IPO. AT&T's routing math shows why the pressure is structural, not just competitive: when your biggest customers can swap a modular open model in for a frontier API and lose two percent quality, the closed-model premium has to keep earning its keep.

If you're an AI buying team reading this: have you benchmarked your own token mix against open models yet? Tell us in the comments.

Read more

Lambda raises up to $4B from Blackstone ahead of its IPO

Lambda raises up to $4B from Blackstone ahead of its IPO

The neocloud money is consolidating fast, and today's inbox shows both ends of the market: a heavyweight pre-IPO round on one side, and a Google open model you can run on a phone on the other. Lambda is raising up to $4 billion led by Blackstone and Coatue at a $14.5 billion pre-money valuation — its last private round before a planned IPO. The Wall Street Journal reported the scoop from a letter to limited partners, and Reuters independently confirmed the headline terms: the round is led by t

South Korea bets $3.49B on its own frontier AI model

South Korea bets $3.49B on its own frontier AI model

Sovereign-model money is getting serious, and the hardware money is following it. Today's inbox: Korea's nine-figure upgrade to its homegrown model push, a physics-simulation startup priced like a chip designer, and Google turning a geospatial model loose on public health. South Korea is putting 4.7 trillion won — about $3.49 billion — of state equity behind a homegrown frontier AI model. The Ministry of Science and ICT confirmed the figure as part of its proposed 2027 budget, split into two t

Mistral's Le Chonk puts Europe's sovereignty bet on a download date

Mistral's Le Chonk puts Europe's sovereignty bet on a download date

Mistral's biggest model ever is real, benchmarked and for sale today — but the thing that makes it matter to Europe's sovereignty argument, the weights, is still three weeks out. The preview settles who built it; the release will settle whether it counts. What Mistral actually shipped Mistral opened a public preview of Mistral Large 4 — unofficially ML4, officially le Chonk — a 1 trillion-parameter mixture-of-experts model with 49 billion active parameters and native multimodal input. The p

Mistral unveils Le Chonk: a 1T-parameter open-weights model

Mistral unveils Le Chonk: a 1T-parameter open-weights model

The biggest open-weight release outside China lands in public preview today, and the country that spent the week promising its own frontier model just put a price on the ambition. Mistral has opened a public preview of Mistral Large 4 — codenamed "le Chonk" — a 1 trillion-parameter mixture-of-experts model with 49 billion active parameters, natively multimodal, which the company calls its largest and most capable model to date. The preview API is live today on Mistral Studio at $1.36 per milli