Today in AI — September 26, 2026

Share
Today in AI — September 26, 2026

Saturday's through-line was who sets the terms: the two governments agreed to call the technology "super intelligence" and open a channel for incidents, Chinese open-weight models crossed into the majority of tokens on the developer gateways that measure such things, and the most interesting engineering of the day happened in the layers nobody markets — the agent harness, the kernel, the SSD.

Models & Research

  • Nvidia's SoL-Pi cuts coding-agent token use by roughly half by having a research agent rewrite the harness instead of the model. The paper treats the control layer between model and environment — tool use, context management, verification, abort logic — as the optimization surface: across 535 executable environments it explored 152 directions over 3,000-plus runs and 60,000-plus agent-environment interactions, landing on four mechanisms including Action Fusion (merging an edit and its test run into one model call) and ObservationPack (archiving long tool outputs and re-sending a summary). Reported savings are 50% against Codex and 54.3% against Claude Code on EdgeBench at roughly unchanged capability, with the held-out 40 tasks walled off from the search so the harness can't overfit to what it was tuned on.
  • A vLLM-alumni startup got 57% more throughput out of 16 Google TPUs than out of 16 Nvidia GB200s running the same model and the same inference engine. Inferact benchmarked Kimi K3 on 16 TPU v7 Ironwood parts against 16 GB200s at 709 versus 452 tokens per second, with the gap widening at small batch sizes and on Qwen 3.8 27B (1,515 versus 695 tokens per second on four chips); accuracy was unchanged on GPQA-Diamond and GSM8K. The trick is a "megakernel" written in Pallas that fuses all 92 MoE layers of the forward pass into one program, removing the bandwidth idle time between kernel launches and prefetching the next layer's weights while the current one computes — it compiled in under 90 seconds instead of XLA's 30-plus minutes, and the code is open.
  • Epoch AI's furniture-assembly test is the cleanest progress chart of the week: the best model went from 28% to 80% in ten months. The benchmark photographs three IKEA builds mid-assembly with deliberate errors, hands the model the manual, and asks it to say whether the build is right and what went wrong. OpenAI's GPT-6 Astra scores 80% at about three minutes per photo, against 70% for Claude Fable 5.1 and 61% for Opus 5; open-weight models still trail the leaders by roughly seven months, and the runtime is the honest caveat — too slow to help you actually assemble anything.
  • A five-experiment study with 3,132 participants found that simply having an AI answer available nearly erases people's willingness to say "I don't know." Participants were asked about fine visual details from films, questions the model (Step 3.5 Flash) got almost always wrong, so the shift can't be read as sensible delegation: withholding answers fell from 36% and 44% in the no-AI control groups to 6% and 3% with AI available, confidence rose from 29.6 to 75.9 on a 100-point scale, and accuracy fell from 27.6% to 10%. Cash incentives for correct answers reduced how often people consulted the model but did not fix the pattern.
  • A three-month-old company put a general-purpose robot model on top of RoboDojo, and its pitch is that AI is doing the research. Simate's Simate-beta leads the board with an average score of 33.95 and a 27.96% success rate, a result the team says came without any benchmark-specific tuning, from a group that previously pushed a single-stage end-to-end driving model to FSD-comparable, production-deployed performance. Its stack is explicitly AI-for-Physical-AI — an AutoResearch engine that runs the experiments, an AI-native infra layer, and a research interface designed to be machine-readable — with MIT, Caltech, Tsinghua and Peking University researchers already in the platform's beta.

Industry

  • Washington and Beijing agreed to call the technology "super intelligence" and to open a standing channel to talk about it. The White House says the new U.S.-China Super Intelligence Dialogue will meet on "risks and benefits related to SI," with the next session before November, and the two sides also set up an AI incident hotline that Axios likened to the Cold War-era red telephone; Chinese foreign-ministry spokespeople said Beijing respects the American phrasing while retaining its own. Chinese reports of the same summit package add a tariff-reduction arrangement worth about $30 billion — the AI channel is the part that will still matter in six months, because it is the first named mechanism either government has attached to the technology itself.
  • Chinese open-weight models went from a fringe share to the majority of tokens on two of the gateways Western companies use to buy model access. On OpenRouter, Chinese models accounted for 57% to 67% of tokens in the week of September 14, up from 6% to 13% in February; on Vercel the share reached 55% in August from 11% in January. Price is doing the work — a model that clears the quality bar for coding and agentic tasks at a fraction of frontier rates is an easy procurement decision — and two House committees are now investigating the adoption on security and influence grounds.
  • Quantum computing venture funding has already passed $4 billion this year, roughly matching all of 2025, which itself nearly matched the previous four years combined. The PitchBook figures land in a year when the AI trade taught investors that compute-adjacent infrastructure can re-rate violently, and quantum's capital cycle is now running on the same logic: fund the hardware layer before the software demand is provable.
  • Bloomberg's read on the recent departures from Google DeepMind is that a cohort of senior researchers is leaving to build alternatives to large language models outright. Departures from a lab that shaped the transformer era are the clearest signal yet that the next architecture argument is being funded privately rather than fought inside the incumbents — and it is the mirror image of the commercial story running all year, where the LLM stack keeps consolidating around fewer, bigger providers.
  • Russia has stepped up targeted strikes on Ukrainian data centers, knocking out internet access for roughly 100,000 Kyiv residents across Wednesday and Thursday. Civilian connectivity has been a front in this war since the first year, but data centers are now the target class — the same facilities that host cloud services, government systems and, increasingly, the AI workloads Ukraine sells abroad.
  • Walmart's incoming CEO says the company will not use its AI shopping assistant or its electronic shelf labels to set prices based on who is standing in the aisle. John Furner's commitment is narrower than the technology: identity-based pricing is technically trivial for a store that knows what you bought, where you are, and what you are looking at, and the pledge is worth exactly as much as the audit trail behind it. Retail's version of the AI question is not whether the model can discriminate, but whether the company can prove it didn't.

Policy

  • A bipartisan group of US lawmakers has introduced a bill to bar the federal government from putting Chinese optical transceivers into sensitive systems. The component is the least glamorous choke point in the AI buildout — the optics that move data between accelerators and across data centers — and the bill treats it as a supply-chain security question rather than a trade one. Expect the same framing to spread to the rest of the AI data-center bill of materials, because the argument that works on transceivers works on everything downstream of the chip.
  • Joe Lonsdale, the Palantir and 8VC cofounder who is also an Anthropic investor, says AI companies warn about existential risk to move public policy their way. The critique is not new — it is the sharpened version of a fight the labs have had with their own investors all year — but Lonsdale is arguing it from inside the money, which makes it harder for the industry to file under advocacy from outside. It also lands the same month a federal court upheld the government's power to blacklist a lab on national-security grounds, where the safety vocabulary was load-bearing.
  • Ukraine's former digital-transformation minister is pitching a privately funded robot army. Mykhailo Fedorov told IT Arena in Lviv that drones now account for more than 95% of target engagements, that Ukraine has over 700 drone makers, and that his programme will invest in defense-tech firms and test working systems with the military at scale — "robots should fight, not people." The humanoid hero shots in his promotional trailer are not battle-ready and may not make military sense, but the direction of travel is the story: autonomous-weapons policy is being debated at the UN while the deployment happens in the private sector, on a battlefield, without a treaty.

Tools

  • A 744-billion-parameter model now runs on a laptop with 25 GB of RAM, because a single developer moved the cold weights to the SSD and streamed them back on demand. Colibrì is a pure-C inference engine with no framework dependencies that keeps the dense part of a GLM-5.2 mixture-of-experts in memory at int4 (about 9.9 GB) and leaves the ~370 GB of routed experts on an NVMe drive behind an LRU cache, with an expert-prefetch heuristic that predicts the next layer's routing with about 71.6% accuracy. It has passed 32,000 GitHub stars, covers nine model families including 2.8-trillion-parameter Kimi K3, and hits roughly 1.8 tokens per second on a 128 GB CPU desktop and 5.8–6.8 on six RTX 5090s — slow, but frontier weights on hardware you already own.
  • KoboldCpp shipped a built-in agent harness with a 2,000-token system prompt, including all nine of its tools. It is a deliberate lightweight replacement for Codex, Claude Code and OpenCode: one checkbox or a launch flag, support for AGENTS.md, context compaction, MCP servers loaded from a mcp.json, and three approval modes for tool calls. The interesting number is the prompt budget — most harnesses spend tens of thousands of tokens on scaffolding before the user's first word, and for a local model on 12 GB of VRAM that overhead is the difference between usable and not.

What to watch: whether the SI Dialogue produces anything before its November meeting, whether Washington's Chinese-model inquiry turns into procurement rules, and whether anyone publishes a harness-optimization result on unfamiliar tasks rather than a held-out slice of the same distribution.

Two of today's best stories were about the same thing from opposite ends: Nvidia automated the search for a leaner agent harness, and a lone developer put a 744B model on an SSD. Both are bets that the next big win is in the plumbing, not the parameters. Where do you think the remaining headroom is — kernels, control layers, or the model itself? Tell us in the comments.

Sources: Nvidia — SoL-Pi harness optimization (arXiv) · THE DECODER — Nvidia's SoL-Pi cuts coding agent token usage nearly in half · Inferact — 700 TPS on Kimi K3: a case for TPU megakernels · QbitAI — Google TPUs run Kimi 57% faster than Nvidia GPUs · Epoch AI — Can AI spot mistakes in IKEA assembly? · THE DECODER — GPT-6 Astra on the furniture assembly benchmark · THE DECODER — AI access and the unwillingness to say "I don't know" · QbitAI — Three-month-old Simate tops RoboDojo with Simate-beta · Axios — US and China agree to a "super intelligence" dialogue · Lianhe Zaobao — 白宫称美中同意设立AI事件沟通渠道 · RFI — 中美达成涉及300亿美元关税减让安排,将启动人工智能对话 · CNBC — Chinese AI models surge in global popularity · Financial Times via Techmeme — Quantum VC funding passes $4B · Bloomberg via Techmeme — DeepMind researchers exit to build LLM alternatives · Financial Times via Techmeme — Russia strikes Ukrainian data centers · Financial Times via Techmeme — Walmart's pricing pledge on AI shelf labels · Reuters via Techmeme — Bill to bar Chinese optical transceivers from sensitive systems · Reuters via Techmeme — Lonsdale says AI firms use existential-risk talk to sway policy · THE DECODER — Fedorov pitches a private-sector robot army · Colibrì (GitHub) · QbitAI — A laptop streams a 744B GLM off its SSD · KoboldCpp Agent (LocalLLaMA)