Today in AI — September 27, 2026
Sunday's roundup runs heavier on rules than on models: Brussels' enforcement gap, a state attorney general asking Congress for a law, dark-web resellers undercutting the labs' own pricing, and Beijing selling the token rather than the model. The engineering that did land was about spending fewer tokens, not more.
Models & Research
- Fireworks Research shipped Ember-1, a version of Kimi K3 trained to stop over-thinking rather than to think harder. The company says K3 spends the majority of its generated tokens — sometimes more than 90% — on internal reasoning, and that in multi-turn agent work every turn re-reads and re-bills the reasoning of the turns before it. Ember-1 cuts reasoning by 35% to 50% at comparable accuracy, holding 92.2% on SWE-bench Verified against K3-max's 93.2% while beating it on Terminal Bench 2.1 (82.0% versus 80.9%) and DeepSWE 1.1 (75.2% versus 66.4%); two customer A/B tests on production coding traffic showed roughly 35% fewer tokens per task. It ships as a two-week research preview alongside K3 rather than a replacement, which makes the pricing question the real test — Fireworks is selling against its own base model.
- Princeton's plasma lab handed part of a tokamak's control loop to an AI framework that makes decisions every 20 milliseconds. PACMAN combines several machine-learning models into a repeating loop: read temperature, density and magnetic signals, check them for errors, then adjust heating and shaping before the plasma has time to misbehave. In one DIII-D experiment a model predicted a tearing-mode instability about 200 milliseconds before it would have formed and changed the plasma to prevent it rather than suppress it, and the framework coordinated all six of the machine's gyrotrons at once — work a focused human operator can measure in seconds. The write-up is a few weeks old but circulated hard again today, and the honest reading is that fusion needs fast controllers more than it needs a bigger plasma model.
- A llama.cpp contributor got prompt-lookup drafting up to 42x faster while using up to 2.6x less memory, by fixing data structures rather than algorithms. Hayder Tirmazi found the n-gram caches were copying inner maps on every drafting step (a bug more than an optimisation), swapped the outer hash map for a denser one, replaced most inner maps with sorted vectors searched by a branchless binary search, and started using a read-only map built on binary fuse filters — which alone cut static-cache load time from 3.76 seconds to 0.23 on a 541 MB corpus. Daniel Lemire then sent a pull request that checks the acceptance thresholds before scoring any candidates, taking the total to as much as 140x. Prompt lookup is a special case of speculative decoding with an n-gram model as the guesser, so this is the cheap half of the speedup stack getting genuinely cheap.
Industry
- Mark Gurman's read is that Meta's VR Glasses are the product Apple's Vision Pro should have been, and that Apple's Vision team is on "life support." Apple is reportedly still working on a lighter Vision Pro successor and experimenting with other headset designs — one AppleInsider piece predicts smart glasses in 2027 and a refreshed headset in 2028 — while moving Vision Pro leadership around underneath. The interesting part is the framing rather than the hardware: Apple spent a decade on the optics and Meta shipped a cheaper device that people might actually wear, which is the same trade-off Apple keeps failing on the consumer side of the AI glasses race.
- The Wall Street Journal profiled Jaan Tallinn, the Estonian investor who led Anthropic's $124 million Series A and has given roughly $170 million to safety work. Tallinn has been arguing about AI risk for over a decade, which makes him the rare funder who was early to both the returns and the warning — and the profile lands in a week when lab leaders' doomsday language is being read as either prophecy or negotiating position. The money is the subtext: the same person financed the company and the field that keeps telling the company to slow down.
- China's digital-trade expo gave tokens their own floor, and the framing is that compute has become a tradeable commodity. The fifth Global Digital Trade Expo in Hangzhou ran September 23–27 with the first Token-themed AI zone, walking visitors from compute clusters and chips up through model APIs to end applications; Zhejiang stood up a provincial Token Operations Center and a unified platform on September 24, and vendors pitched things like a "security token factory" that turns security data and tooling into metered token assets. Beijing has published a 2026–2028 action plan for what it calls the token economy, and Shantou has followed with subsidies and token vouchers. This is the most concrete version of the Chinese position on the token-share fight: if the West buys Chinese models by the token, the state wants that transaction denominated and taxed in China.
Policy
- The New York Times reports that AI's pace has outrun every regulator trying to govern it, with EU AI Act enforcement lagging badly behind the law's own calendar. The piece opens on European Commission president Ursula von der Leyen convening senior officials at a 19th-century palace outside Brussels — a meeting that reads less like enforcement and more like triage — and describes regulators caught between wanting to harness the technology and wanting to constrain it. The gap is structural rather than political: the rules were written for a slower product cycle, and every capability release since has added work for offices that have not grown.
- Google's threat intelligence team found dark-web marketplaces reselling access to Anthropic, Google and OpenAI models at discounts of up to 97%, a practice researchers are calling "LLM-jacking." The economics are the story — stolen API keys and hijacked accounts turn enterprise model access into a commodity that undercuts list pricing, and the victim is whoever pays the inference bill. It also puts a number on the security assumption underlying every agent deployment: if access can be resold anonymously at a 97% discount, access control is a billing problem, not a boundary.
- New Mexico's attorney general asked Congress for federal AI regulation, citing the attempted OpenAI-linked intrusion into a state university's systems. Raúl Torrez's office is arguing from the receiving end of the agent incidents the labs are still counting — a public university has neither the visibility nor the leverage to audit a frontier lab's model behaviour, and the state's answer is a federal standard rather than fifty state ones. It is the same request the bundle of state attorneys general made earlier this month, now with a local victim attached.
- Microsoft AI chief Mustafa Suleyman told Bloomberg that testing models ten times larger than today's may require removing guardrails, and that the industry needs a cross-industry safety body. His argument is that you cannot validate a system's limits while the safety layer constrains the tests — which is a defensible research position and also exactly the reasoning OpenAI's incident reports keep running into. The call for a shared body is the more consequential half: the labs have spent the year publishing per-company incident write-ups and no shared standard for what gets reported.
- Hong Kong's de facto central bank told the city's banks that AI has moved from advising to executing, and that governance has to catch up. Deputy chief executive Arthur Yuen, speaking at the Hong Kong Bankers Summit, said the question for boards is no longer whether AI is used but how it is governed, and raised technical resilience and the ability of bank systems to withstand increasingly complex AI-driven activity; the HKMA said practice guidance is coming. Banks are the right place for this test — they already have the model-risk frameworks, the audit trails and a supervisor willing to name the requirement.
Tools
- Imp is a full port of DSPy to the BEAM, so Elixir and Erlang teams get declarative, self-improving language-model programs without leaving the runtime. You describe what each model step takes and returns, choose how it reasons, and hand an optimiser examples of what good output looks like. It matters because the DSPy pattern is becoming the default way to build pipelines you can actually evaluate, and until now adopting it meant running Python beside your Erlang services.
- A sampler called Engram runs a tiny AI model on-device to mangle audio and hallucinate sounds that were never recorded. Thoughtful Things' Kickstarter instrument is deliberately not a push-button song generator — it is pitched as a field recorder for latent space, in the lineage of circuit bending, with firmware the company plans to open so owners can swap in their own models. The model runs locally with no internet connection and was trained only on openly licensed audio, and the campaign prices a limited first run at $675 against an expected retail of $850 to $900. The interesting bet is the inverse of the cloud-music play: keep the model small and weird, and sell the person a physical thing.
Which of today's moves actually binds anything — a trade-show floor, a state attorney general's letter, or a supervisor's guidance to banks? Tell us in the comments.
Sources: Fireworks Research — Introducing Ember-1 · Specialized Intelligence Index — Bedside Bench Pareto frontier · PPPL — PACMAN framework makes key decisions in milliseconds · ScienceDaily — AI can now control fusion plasma faster than humans can react · Hayder Tirmazi — 42x faster prompt lookup drafting in llama.cpp · Hacker News discussion — faster prompt lookup drafting in llama.cpp · Bloomberg — Meta's VR Glasses are what Apple's Vision Pro should have been · 9to5Mac — Apple working on a new, lighter Vision Pro · Wall Street Journal — Jaan Tallinn and the safety bet behind Anthropic · Xinhua — What to see in the digital trade expo's first token-themed zone · 21st Century Business Herald — The fifth Global Digital Trade Expo puts AI and tokens at the centre · New York Times — How AI's acceleration created a global policy vacuum · Financial Times — Hackers steal AI access to power a new cybercrime wave · Moneycontrol — Hackers are stealing AI access to power a new wave of cybercrime · Source New Mexico — Attorney General Torrez urges Congress for AI regulation after attempted OpenAI hack on UNM · Techmeme — Q&A with Microsoft's Mustafa Suleyman on AI safety incidents · Hong Kong Monetary Authority — Arthur Yuen's remarks at the HKMA Bankers Summit · Imp (GitHub) · The Verge — Engram is a sampler that turns broken AI hallucinations into music · Engram — Kickstarter campaign