Today in AI — September 15, 2026

Share
Today in AI — September 15, 2026

Tuesday's through-line was where the model lives. Shanghai AI Lab put a 744-billion-parameter agent on Hugging Face with no announcement, MediaTek and Honor put one inside the phone's silicon and operating system, and Google handed its own engineers a rival's coder. The infrastructure bill for all of it also came due: gas forecasts, a stacked data center, and a cloud region AWS says it cannot rebuild.

Models & Research

  • Shanghai AI Laboratory posted a 744-billion-parameter agentic model with no announcement, no paper and an MIT licence. Atria Dawn Preview — built on the 744B GLM-5.2 foundation — ships with a 1,048,576-token context and 8 of 256 routed experts active per token, and its model card claims five benchmark leads: BrowseComp at 92.5 against GPT-5.6 Sol's 92.2, CyberGym at 86.5, DeepSearchQA at 96.0. The card also shows its losses, which is what makes it readable: SWE-bench Pro at 59.6 against Claude Opus 5's 74.7. Running it yourself means downloading roughly 756 GB in FP8 or 1.5 TB in BF16, and no independent evaluation exists yet.
  • Voice got cheaper and more Chinese on the same day. Google launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, putting the extended model at 82.6% on Artificial Analysis' speech-to-speech quality index against GPT-Live-1 Astra's 81.5% — while pricing an hour of input audio at $0.84, roughly a sixth of GPT-Live's $5.83. Hours earlier, Shanghai's StepFun had shipped five StepAudio 3 models at once and claimed first place globally on full-duplex conversational dynamics at 98.9 and a 1.7% word error rate on transcription. A one-point benchmark lead is inside the noise; a sixth of the price is a commoditization move.
  • Meta FAIR published the result that breaks distillation's ceiling: shrink the vocabulary to 256 bytes. In a paper written with a University of Washington team, 1.28-billion-parameter byte-level students distilled from token teachers start behind and climb a steeper curve, predicted to pass their token counterparts by up to 4% average accuracy at the limit while matching them on one-sixth of the data. The practical warning is aimed at anyone running evals: identical validation loss maps to different benchmark scores across tokenizers, so the comparison needs a downstream calibration step.
  • Two unpublished NeurIPS submissions, six days each and thousands of dollars of compute, and both papers came back rejected — by their own authors. The agents completed all of the engineering without human help and still made no substantial progress on the research questions, with five recurring failure modes: poor judgment about what counts as publishable, uncreative responses when the design hit trouble, ineffective backtracking, no awareness of their own budget, and instruction drift. The method is the more durable contribution — "shadow evaluations" hand an agent a real paper's central question and let the people who wrote that paper grade the result.
  • The models that don't chat had a day: Prior Labs shipped TabPFN-3.5 to the top of the tabular arenas, and TypeSafe came out of stealth with $40 million for a model that refuses to write prose. TabPFN-3.5 lands about 150 Elo points ahead of the previous leader on BeyondArena, adds a thinking variant worth roughly 44 more Elo on TabArena, and eats raw text and datetime columns with no preprocessing — it went into SAP AI Core on day one. TypeSafe's Jev returns typed decisions with confidence scores at $0.042 per million input tokens and output tokens free, though the company's own site concedes its headline 193x speed claim sits at the high end of what real workloads should expect.

Industry

  • Google now gives every engineer access to Anthropic's Claude Opus 5 — inside Antigravity, on a quota. A spokesperson called Gemini "primary and foundational" and third-party models supplementary, which restates the concession rather than softening it: after spending billions pushing its workforce onto its own model, the best coder in the building is a rival's. The move lands in the same week Palantir, Nvidia and Booz Allen were reported restricting Claude over data-retention fears.
  • Meta turned its custom silicon roadmap into deliveries: MTIA 450, codenamed Arke, goes into data centers in the first half of 2027, with MTIA 500 (Astrid) reaching production sites by the end of the year. Twelve MTIA 450 chips arrived from TSMC on September 1 measuring within 2% to 3% of Meta's simulations, and the company has committed to more than a gigawatt of them over twelve months. The design target is inference economics, not training bragging rights — memory bandwidth doubled to 18.4 TB/s and hardware attention acceleration, which is what answering for billions of daily users actually costs.
  • The agent moved into the hardware twice in one day. MediaTek launched the Dimensity 9600 Pro, a TSMC 2nm chip with dual NPUs, an "Agentic AI Engine" and compression that runs a 30B mixture-of-experts model entirely on-device at 40% lower always-on power. Hours later Honor shipped MagicOS 11 with what it calls the first commercially deployed system-level Agent Harness: roughly 700 callable system tools, more than 100 steps per instruction, and Honor's own claim of closing about nine of ten complex multi-app tasks. Google and Anthropic are selling agents by the token while the phone vendors sell them by the handset.
  • Physical AI collected the day's biggest cheques. CADDi raised $114 million at a $1.2 billion valuation for models that read 2D drawings and 3D CAD files general-purpose LLMs cannot, with more than half of Japan's 100 largest manufacturers on the platform. Shield AI was reported in talks for a round valuing it at $20 billion or more, up from $12.7 billion in March, and Agility unveiled Digit 5 — a humanoid designed to work beside people with no safety fencing, which it says has drawn more than $300 million in multi-year orders.
  • Universal Music Group sued DistroKid for deceptive trade practices and copyright infringement, and the filing's numbers do the arguing. One account, Lofi Chill, released 4,562 tracks in a twelve-month window; technical analysis in the 52-page complaint found more than 97% of Chill Flow Radio's catalogue and more than 98% of Mellow Vibes Radio's to be raw Suno output. The sharpest allegation is what the filing calls ISRC theft — uploading under another recording's code to collect its streams. The majors spent the year licensing the generators; the unlicensed pipe turned out to be the distributor.
  • Salesforce used Dreamforce to make two versions of the same argument: stop renting reasoning, and stop needing a user interface. Koa, built with Nvidia on the open-weight Nemotron base, is meant to handle Agentforce's multi-step thinking in-house — post-trained on simulated enterprise workflows with no customer data, and pilot-only until winter. AIforce pushes Salesforce's data, permissions and business logic out of its own UI and into Claude, Slack and whatever surface an employee already has open, with a claimed zero data retention. Both are the enterprise stack declining to be a tenant of the frontier labs.
  • One IT consultancy's Cursor renewal went from about $200,000 for 800 licences to a $1.5 million ask — 7.5 times — and the unbundling that followed is the story. The firm negotiated down to $250,000, Sanofi's chief digital officer said he may not renew at all, and Anthropic-backed Sapiom cut roughly 25% by swapping Claude Code for an open-source harness. "All the value is accruing to the harness," one consultancy founder told The Information — which is the same finding Mozilla's open-model report made from a different direction.
  • The build-out's physical limits got three new numbers. BloombergNEF now expects US data centers to add 15 billion cubic feet per day of gas demand by 2035 — more than Germany and Japan burn today combined, and more than double its own December forecast — while Huawei says its stacked three-floor data center design cuts electrical installation from six months to three. The hardest one came from AWS, which told customers it cannot restore its Bahrain region or one of its UAE availability zones six months after drone strikes, conceding the damage "exceeded what our regional and multi-AZ services are designed to withstand."

Policy

  • FTC Chairman Andrew Ferguson said "everyone should be deeply suspicious" of labs that ask for regulation and an antitrust exemption in the same breath. Speaking at Georgetown, he called that combination a request for "barriers to entry that will insulate their incumbency from challenge," without naming Anthropic chief executive Dario Amodei — the author of the proposal. Hours later, OpenAI told a FRONTIER Act sponsor it supports the bill's embedded-evaluator provision, moving the company from voluntary evaluators to mandatory ones while the preemption clause that would override state AI law stays the real fight.
  • The Commerce Department ordered Kalshi to pull its AI-compute forward curve on national security grounds, and pushed the CFTC to freeze approvals of new compute contracts for 60 days. The product aggregated users' bets on Nvidia rental prices into a public curve of where compute costs were heading; many of the underlying markets still trade. CME's planned October 5 launch of H100 and B200 rental-index futures now sits inside that window — a commodity market that only works when its prices flatter the incumbent trade is not a commodity market.
  • The referee question got three more answers, none of them from a government. Elon Musk proposed letting the top labs plus "three or four of the leading Chinese companies" run test harnesses on each other's unreleased models, framing it as competitors grading homework where labs currently grade their own; no lab has agreed. A DeepMind safety researcher who left in July posted that he "earnestly believe[s]" AI could kill us all, and Britain's government said ministers must "heed the warnings" while OpenAI urged London to write mandatory independent testing into law. The only version of the idea with a sponsor in Congress remains Amodei's embedded evaluators.
  • China wrote on-device models into its industrial plan. The 15th Five-Year Plan for electronic information manufacturing names on-device large models, compute chips, memory chips and operating systems as one compatibility problem to solve together, orders adaptation work between AI chips and models, and starts drafting a unified specification for high-speed interconnect buses under the name CLink. It is a procurement mandate rather than a research target, which explains why the open-weight 4B-to-30B race keeps accelerating there.

Tools

  • NVIDIA's research arm open-sourced an agent harness that discovered four of its own efficiency mechanisms through automated search. SoL-Pi reports 45% to 49% fewer tokens and roughly a third lower cost than the stock Pi agent while keeping about 94% of its average score, with all four mechanisms — bundling an edit with its validation command, archiving oversized tool output, verifying summaries against the original log, and compacting context only when projected savings beat the rewrite cost — opt-in and off by default. Every number is NVIDIA's own, no independent reproduction exists, and the preprint is still pending.
  • Agent infrastructure got a governance plane and a database on the same day. StackGen launched an Autonomous Operations Factory that runs infrastructure, DevOps, observability and reliability agents under one world model, one harness and one audit trail, with an Autonomy Index that scores how much authority each stage actually holds — the number an enterprise needs before it can defend a rollout to a board. KeewanoDB raised $12 million for an event store built for agents rather than SQL writers, keeping events ordered around an entity and running filtering in-database so the model receives a narrow window instead of the whole history; the company says that cuts token use by around 84%.
  • IBM Research has a fix for the agent that impresses in a demo and then takes a different path in production, and the diagnosis is more useful than the fix. On AppWorld, a ReAct agent running GPT-4.1 succeeded on 77.4% of runs but succeeded on all five repetitions for only 53.0% of tasks — a consistency gap that widens to roughly 30 points on the hard tier and is not sampling noise, since the runs were at temperature zero. Feeding the flagged decision steps back as guidelines cut the gap from 24.4 points to 12.0 and raised average accuracy rather than trading it away; the analyzer now ships in IBM's open-source toolkit.

What to watch: whether the Gulf's compute build-out gets repriced now that a hyperscaler has publicly written off a region, whether the FRONTIER Act reaches markup before the fall recess, and whether any lab says yes to Musk's peer-testing proposal.

The model stopped being the product today — a phone chip runs a 30B model, Google's engineers code on Anthropic's, and a 744B agent shipped with no launch post. Which layer do you think is actually defensible? Tell us in the comments.

Sources: Atria Dawn Preview (Hugging Face) · 观点网 Guandian — Shanghai AI Lab open-sources Atria Dawn Preview · Google DeepMind — Gemini 3.8 Live · arXiv — StepAudio 3 Realtime technical report · arXiv — Breaking the Token Ceiling · AILog — byte-model distillation · arXiv — Can AI agents conduct open-ended AI research? · Prior Labs — TabPFN-3.5 technical report · TypeSafe AI — Jev and System One models · Business Insider — Google lets all engineers use Claude · Bloomberg — Meta's in-house AI chips · Meta AI Engineering — MTIA · MediaTek press release — Dimensity 9600 Pro · Zhidx 智东西 — MagicOS 11 agent harness · Fortune — CADDi's $114M Series D · The Information — Shield AI valuation talks · The Robot Report — Agility's Digit 5 · Music Business Worldwide — UMG v. DistroKid · UMG Recordings v. DistroKid docket · TechCrunch — Salesforce and Nvidia's Koa · Salesforce — AIforce announcement · The Information — Cursor customers fight price hikes · The Next Web — the coding fight moved to the harness · BloombergNEF — US data center gas demand outlook · CNBC — AWS cannot restore Bahrain and a UAE zone · OSChina — Huawei's 3D data center · Reuters — FTC chair suspicious of AI antitrust exemptions · POLITICO — OpenAI backs third-party safety assessments · Semafor — Commerce ordered Kalshi to pull the compute curve · CFTC Release 9286-26 · CNBC — Musk urges labs to test each other's models · Bloomberg — DeepMind researcher's exit post · City AM — UK opens the door to AI laws · Economic Observer via Shanghai Securities News — 15th Five-Year Plan · NVlabs SoL-Pi (GitHub) · SoL-Pi project page · StackGen press release via Business Wire · Keewano · IBM Research — consistency gap analysis · ALTK-Evolve (GitHub)