Today in AI — September 21, 2026
Monday's through-line was accountability: who answers when a model misbehaves, and who pays when it works. Washington drew that line at management, two labs reportedly negotiated to attack each other's models, China's flagships kept winning the volume race while one of them took a shelf at Amazon, and two compression papers argued that smaller can now mean better rather than merely cheaper.
Models & Research
- ByteDance Seed and Tsinghua AIR open-sourced DAPO, a reinforcement-learning system whose full recipe — algorithm, dataset and training infrastructure — ships alongside the weights. The Decoupled Clip and Dynamic Sampling Policy Optimization method reaches 50 points on AIME 2024 from a Qwen2.5-32B base, beating DeepSeek-R1-Zero-Qwen-32B while using roughly half the training steps, and the team's training logs show why: response length, reward and entropy climb together in a way that keeps long reasoning runs stable instead of collapsing. The weights, the datasets and two reproduction scripts are public, which is the actual contribution — most labs publish a result and keep the pipeline.
- A 4-bit model that outperforms its own full-precision checkpoint is the claim in Quantization-Aware Healing, and Multiverse Computing has the receipts — with one acknowledged gap. The method distils the compressed student from the original pre-compression teacher rather than fine-tuning it against hard labels, and applied to a GPT-OSS 120B cut to 60B and re-quantized to 4-bit, the result wins on 7 of 9 benchmarks against that model's own bfloat16 version, gaining 7.4 points on long-context reasoning and 5.6 on AIME 2025. It also converges far faster and does not decay: QAH peaks in about 100 steps against 700 for quantization-aware training, which sheds nearly 19 points by step 1,200. The caveat is in the comments — a reviewer pointed out the two checkpoints received unequal training, and the authors conceded the table does not isolate "distilling under quantization" from "distilling longer."
- The same lab published a second compression result that reframes pruning as a physics problem: choose which blocks to delete by brute-forcing an Ising Hamiltonian on one GPU. Multiverse's constrained binary optimization searches up to tens of billions of configurations, and on Llama-3.3-70B-Instruct it pulls decisively ahead of block-influence baselines at aggressive compression — almost 23 MMLU points better with 40 of 80 blocks removed and no retraining. The formulation is architecture-agnostic, so it transferred unchanged to NVIDIA's Nemotron-3-Nano-30B hybrid, where removing two or three MoE layers or two attention layers still beat the baseline on AIME25 and GPQA. Redundancy in hybrid stacks is real but unevenly distributed, and the paper's useful admission is that the best configuration is often not the ground state.
- China's weekly model call volume ran ahead of the United States for a twenty-first straight week, at roughly four and a half times the American total, with DeepSeek V4.1-Flash taking the top single model slot. The week's tally put Chinese models at 67.46 trillion tokens, up 10.28%, against 14.21 trillion for US models, down 34.7%, and four of the global top five were Chinese. Volume is not capability — it is a price and distribution story, and it is the reason this chart has mattered all year, as we noted when China's LLM call volume topped the US for 15 straight weeks. The number to watch next is whether the gap survives DeepSeek's planned API price increase.
Industry
- Moonshot's Kimi K3 went generally available on Amazon Bedrock, and with it the first Chinese model distributed through a global cloud on a revenue-share basis. Enterprise developers can now call the open-weight model directly through Bedrock, and Moonshot confirmed the arrangement both outlets had been reporting as a rumour: cloud providers list the model and split usage revenue with the lab. Kimi K3 is the open flagship the company puts at 2.8 trillion parameters with a one-million-token context window. The commercial detail matters more than the model — it is the difference between being downloaded and being sold, and it is the same arithmetic that has Western startups dropping frontier-lab rentals for their own models.
- Anthropic has reportedly pushed its planned IPO from October to November so it can present a stronger third quarter. The expectation attached to the listing is a valuation near $2 trillion, and the reported drags are infrastructure cost — including roughly $1.25 billion a month under one compute deal — alongside unresolved security questions after a month of hacking incidents. Nothing here is confirmed by the company, and a one-month slip inside an IPO process is routine; what is worth noting is that the lab calling publicly for a slower industry is also the one racing a calendar.
- Kairos Power will get up to $100 million from Samsung C&T for the reactors it is building on Google's behalf, $70 million of it as equity and the rest as in-kind engineering. Samsung C&T has built or helped build about a dozen nuclear reactors worldwide, which is the point: Kairos needs a five-year path to first power and then needs to move faster than that on the reactors behind it. Hermes 1, the low-power demonstrator at Oak Ridge, and Hermes 2, the first commercial-scale unit, together count as the first 50 megawatts of the Google agreement. A fluoride salt-cooled design has been discussed for years without being built at commercial scale, which is the honest risk in the plan.
- NVIDIA opened DSX Ready, a qualification program that lets AI-factory builders pick power and cooling products against its reference design instead of guessing. The program launches with two categories: battery energy storage, with qualified offerings from Hitachi Energy, LG Energy Solution and Tesla, and cooling distribution units from LG Electronics, LiquidStack and Vertiv. NVIDIA is explicit that passing qualification does not replace site-level engineering or imply site-level stability — a fair hedge, given that optimising one part of a factory tends to move the bottleneck rather than remove it. Certification programs are how infrastructure markets consolidate, and the first two categories are the two that most often decide whether a site gets built.
- Industry data puts global humanoid robot sales at roughly 7,000 units last year, with many bought for research rather than productive work. That is the number that should sit next to every humanoid announcement of the past twelve months: unit counts in the thousands, not the tens of thousands, and a meaningful share of them functioning as lab equipment. Bank of America projects shipments reaching 90,000 this year, which is the forecast to hold the sector to. The gap between demonstration clips and deployed hours is still the story.
- Apple's $250 million Siri settlement is now open for claims from iPhone owners. Eligible users can file through the settlement site with a device serial number before December 21, 2026, for an estimated $25 per device — a figure that can rise toward $95 if claims come in low. It is a modest payout attached to a large claim about assistant features that were advertised and not delivered, and it closes one chapter of the consumer-AI marketing problem without answering it.
Policy
- Treasury Secretary Scott Bessent said the US proposed an AI incident-notification mechanism to China, and that both sides agreed to set up an AI dialogue ahead of the September 24 Trump–Xi summit. A hotline for AI incidents is the concrete version of the coordination both governments have talked about all year, and it lands with the same administration refusing to hand the industry liability relief: Bessent said the Trump administration will not provide liability exemptions to AI leaders. In the same stretch he said the Hugging Face hacking incident is "the responsibility of the OpenAI management, not a bunch of agents" — a sentence that places accountability with executives rather than models, and quietly rules out the defence the sector has been building.
- OpenAI was negotiating a legally binding deal with Anthropic for the two companies to stress-test each other's models, according to a report on discussions that had not previously been disclosed. The talks predate the recent run of security incidents, and they would have been a genuine departure: adversarial testing conducted by a competitor rather than by the lab itself or a hired auditor. The report does not say the deal was signed, and neither company has confirmed it. If it holds, it is a stronger form of oversight than the self-reporting frameworks both labs published this month — because the tester has no interest in a clean result.
- Prime Minister Mark Carney and President Emmanuel Macron urged the leading AI powers to cooperate on regulation and make the technology "safe and effective." Speaking from the French territory of Saint-Pierre-et-Miquelon, Carney grouped four countries at the frontier and called for international coordination on emerging technology. It is a statement of intent without a mechanism, and it arrives in a week when Washington and Beijing were negotiating one of their own — the pattern worth watching is whether the middle-power push becomes a third channel or just commentary on the two that matter.
Tools
- AWS released Strands Harness, an Apache-licensed general-purpose agent that runs locally or in any cloud, not only its own. The pitch is that local agent prototypes built on Claude Code or Codex stop working the moment a team tries to scale them, and a preassembled harness with read, write, shell and web-search tools removes that translation step. AWS says its own benchmarking puts the agent 26% more efficient than other frameworks on the same underlying model, and in one test using Anthropic's Fable 5 it cost 77% less than Claude Code on identical tasks while scoring higher on Terminal Bench 2.1. Treat vendor benchmarks as vendor benchmarks, but the context-management design — offloading tool results to files and caching reused request parts — is the part that generalises.
- ByteDance launched Dramagic, a platform built to take a short drama from script to screen across one pipeline. Access is by request through BytePlus, and the framing is industrial rather than creative: a single workflow covering the stages that Chinese short-drama studios currently stitch together from separate tools. Given how much of the genre's content is already AI-generated, the interesting question is not whether the tooling works but whether the format's economics survive everyone having it.
What to watch: whether the OpenAI–Anthropic stress-testing agreement ever gets signed or quietly dies, whether an incident hotline between Washington and Beijing outlives the summit it was announced ahead of, and whether the next 4-bit compression result arrives with the control arm a reviewer already asked for.
Companies keep volunteering to police their own models while a competitor's audit sits unsigned on the table. If labs won't test each other, who should — regulators, customers, or nobody? Tell us in the comments.
Sources: DAPO (GitHub) · DAPO paper · DAPO-Qwen-32B weights · Multiverse Computing — Quantization-Aware Healing · QAH paper · Multiverse Computing — Pruning LLMs Like a Physicist · Block removal paper · NBD — China's weekly call volume · Eastmoney — weekly token ranking · AWS — Kimi K3 on Amazon Bedrock · Jiemian — Moonshot's revenue-share deal · ITHome — Bedrock adds Kimi K3 · WSJ — Anthropic shifts planned IPO to November · The Information — Anthropic's IPO waiting game · TechCrunch — Kairos Power gets up to $100M from Samsung group · NVIDIA — DSX Ready qualification program · Reuters via Techmeme — humanoid sales hit 7,000 · The Express Tribune — humanoid robot sales tally · The Verge — Apple's $250 million Siri settlement · Financial Times via Techmeme — US proposes AI incident notification mechanism · Bloomberg via Techmeme — Bessent on the Hugging Face hack · The Information — OpenAI and Anthropic neared stress-testing deal · Techmeme — OpenAI–Anthropic stress-test talks · Anadolu Agency — Carney urges AI powers to cooperate · The Guardian — Macron and Carney announce closer ties · SiliconANGLE — AWS debuts Strands Harness · The Decoder — ByteDance launches Dramagic