Today in AI — September 23, 2026

Share
Today in AI — September 23, 2026

Wednesday's through-line was the leash. Altman and Amodei took the UN Security Council's stage to ask for international standards while Washington's own alert plan left technical experts off the list, and the sharpest evidence of why anyone is arguing arrived from Anthropic's own red team: models that can put a stranger within a kilometre of their home from a photograph, and fly a simulated drone into a moving car.

Models & Research

  • Anthropic's Frontier Red Team measured how much of a spy's and a weapons engineer's job current models can already do, and the answer is a meaningful fraction. On geolocating outdoor photos from a single static frame, Mythos Preview landed a median 37 km from the true location across 6,000 images, beating the strongest human baseline the researchers could find — 151 km, the median for the top 0.01% of competitive GeoGuessr players — with Mythos 5 at 47 km, Opus 5 at 181 km and Kimi K3, the newest open-weights model tested, at 385 km. Given anonymous posts and a sandboxed search tool, the frontier models placed a user's home to a median of about 20–22 km, and at least one model got 8% of 1,697 users within a kilometre. The team's own framing is the part to hold onto: these are floors, not ceilings, and the open-weights gap "should not be mistaken for safety."
  • The same evaluation suite turned models loose on drone flight software, and Opus 5 flew it best. In a simulated terminal-guidance task — a quadcopter with a 640×480 forward camera, no GPS and 90 seconds to hit a car — Opus 5 struck 80% of its launches against a parked, high-contrast vehicle and 47% against a moving one, against 15% and 1.6% for Kimi K3 and 5% and 0% for Sonnet 5; across all nine settings Opus 5 hit on 20% of 540 launches. The behaviours that separated it are unglamorous engineering habits: smaller code edits per iteration, proportional navigation and a target-tracking Kalman filter early, and — uniquely — writing its own physics model of the drone to test the controller before spending a launch. Camouflage, evasive targets and heavy wind gusting remained essentially unsolved by every model, which is where the honest caveats live.
  • OpenAI formed an independent mathematics advisory group after its internal model resolved more than 100 open problems, and the mathematicians' reception is the story to watch. OpenAI says the group is meant to be a bridge between the company and the mathematics community — helping assess new results and improve how they are disseminated — and that the model, which began training on August 28, has now cleared more than 100 long-standing problems across most areas of mathematics on top of the Navier–Stokes Millennium Prize claim that a mathematician already cried foul over. Researchers who spoke to NPR said they learn little from a solved problem without the reasoning that produced it, which is the real test of whether an advisory panel changes anything: publication practice, not headcount.
  • Google shipped two speech models that let you invent a voice from a text description instead of picking one off a shelf. Gemini 3.8 Flash TTS takes a prompt that can define a voice's role, accent and vocal traits across more than 100 languages, sits alongside a library of over 2,000 presets, and will clone a voice from a 30-second sample — provided the speaker records a consent statement that has to match the sample. Flash-Lite TTS is the cheap tier aimed at dubbing and voice agents, both models take per-line stage directions and scripted laughter or sighs, two-voice dialogue comes out of a single script, and every clip carries an inaudible SynthID watermark. One reviewer's test runs still picked up a background whine and a voice that shifted near the end of a clip, so treat the demo-quality claims as provisional.
  • Qualcomm's new flagship phone chips are the first 2nm Snapdragon parts, and the pitch is agents rather than benchmarks. At Snapdragon Summit the company launched two sixth-generation Snapdragon 8 platforms together for the first time, led by an eight-core CPU clocked up to 5 GHz — a first for a mobile part — with a new shared-cache design aimed squarely at agents that migrate threads and call tools across cores. The NPU adds an Element Accelerator for Transformer workloads with a 32k-token context window and 50% more shared memory, and Qualcomm says a 30B-parameter mixture-of-experts model (StepFun's StepEdgeOmni) now runs fully on-device at over 50% lower memory cost, with UFS 5.0 streaming weights from flash. High-bandwidth compute, its data-centre interconnect idea, is coming to phones as a co-processor, and Xiaomi is first to ship.

Industry

  • Enveda raised $311 million at a $2 billion valuation to keep mining plants and microbes for drugs with AI. The Series E was led by Catalio Capital Management with Iconiq participating, and doubles the valuation the company reached twelve months ago. Founded in 2019 by a former Recursion Pharmaceuticals employee, Enveda has several candidates in patients, including one for severe skin conditions and one aimed at holding weight off after GLP-1 treatment stops — a reminder that the AI-drug-discovery story is now measured in clinical trials rather than in pipeline announcements.
  • Bessemer raised $5.75 billion across two funds and pointed essentially all of it at AI. The firm split the raise into $1.75 billion for seed and early-stage cheques and $4 billion for growth, and says it has put $3 billion into AI-related companies and backed more than 260 AI-native startups since 2022, with Anthropic, Cognition, Perplexity and Waymo already on the sheet. The tell is in the reasoning: partner Byron Deeter says AI-native companies scale faster than any category the firm has backed, and that companies staying private longer is what forces ever-larger funds. War chests are a lagging indicator of a thesis, not evidence for it.
  • Meta's Muse is reportedly using human contractors to stand in for the AI on phone calls during testing, which is a strange look for an agent that just topped the App Store. Cailian Press reported the beta mechanism and the story travelled fast in Chinese tech media, arriving the same week Amazon blocked Muse from buying anything on its site — the standoff that has become the template for retailers guarding the customer relationship — and as shares of brokerages and travel sites fell on the read-through that an agent which books and buys for you is a threat to whoever currently owns the transaction. Meta has not commented on the contractor report, and the company's developer conference lands with Muse momentum doing most of the talking.
  • Flock Safety, the licence-plate camera network, is weighing a potential sale and has held preliminary talks with outside advisers. The report arrives after a rough stretch for the company: it offered buyouts to its 1,500 staff as cities dropped its cameras, following a run of disclosures about how the system is actually used. A surveillance vendor shopping itself is also a verdict on the procurement politics — the customers leaving are the ones who were supposed to be the growth story.

Policy

  • Sam Altman and Dario Amodei both addressed the UN Security Council, asking for international coordination on AI safety days after President Trump told the same body there would be no global rules. Altman's prepared remarks named two ways the technology goes badly — losing control of the future to systems moving faster than institutions can follow, and concentrating power in too few hands — and proposed complementary national and international frontier standards, incident-reporting protocols so failures are learned from before they become catastrophes, and secure channels between governments, infrastructure operators and technical experts. He also said OpenAI has slowed down unilaterally before and will again. Amodei's message was the same in substance: coordinate, and slow the pace. Neither lab has explained how standards get enforced if the largest government in the room has already declined to join.
  • Washington is touting a bilateral AI safety alert mechanism that leaves technical experts off the distribution list, and Beijing has not answered. The proposal would have the two governments warn each other about AI incidents ahead of Thursday's Trump–Xi meeting, but routes the alerts through official channels rather than the researchers and labs who would recognise one. The silence from China is doing most of the work in the story so far, and it lands as OpenAI argues publicly that the United States should lead on writing the international rules and that it expects cooperation with China to be part of the summit's agenda.
  • A US representative introduced legislation to shut down the southern border surveillance tower program, days after an investigation quantified the deaths inside its coverage. Delia Ramirez, an Illinois Democrat on the Homeland Security Committee, said the towers "just don't work" and cited the finding that nearly 1,100 people died within range of the billion-dollar system between 2015 and 2026, per MIT Technology Review's "Dying on Camera" investigation, which also showed the Department of Homeland Security does not track the program's outcomes. The bill, developed with immigrant-rights and Latinx political organisations, is framed as part of a broader plan to replace the department rather than as a fix to it. We looked at the audit gap when the investigation landed.
  • Illinois Governor JB Pritzker is assembling a state AI cabinet to assess AI threats, joining the run of state-level moves that Washington has not made. The panel is being stood up amid calls for tighter regulation and follows Seattle's ban on algorithmic grocery pricing and New York's frontier-developer registry in the same month. State governments are now the most active regulator of AI in the United States by default, which is a strange place for a technology argument this large to be settled.

Tools

  • ChatGPT's voice mode became an agent on mobile, which is the version of this feature people will actually use. Plus and Pro subscribers can now use the Work tab on their phone to draft documents and emails, summarise Slack messages, build sites and presentations, and drive the cloud browser, while Free and Go users get plugins and connected apps; conversations switch between text and voice and can be resumed on desktop. OpenAI says Voice can now be powered by GPT-6 Astra, Sol and Luna, which turns the phone into a front end for the same work surface the desktop app has been pushing all year.
  • Cisco Talos documented Windows malware that outsources its tactical decisions to a vote among four commercial models, and released the toolkit it used to find it. The 16.4 MB Go binary, which Talos calls CLOSEDQUORUM, has no command-and-control server: every five to 15 minutes it sends the host's state to Gemini, DeepSeek, Qwen and Mistral under a system prompt describing the model as an "advanced malware strategist", tallies the verdicts with DeepSeek breaking ties, and executes one of four modules for credential theft, code injection, persistence or lateral movement. Nothing suggests it has been used against anyone yet, and Talos says six samples trace a roughly week-long build chain back to a criminal-forum account posting about carding since 2025. The research toolkit, CAIRN, works from metadata alone — 24 filters looking for the fingerprints that model integration leaves behind — so nothing has to be detonated.
  • Anthropic made claude.ai about three times faster in a two-week sprint, with Claude finding and shipping most of the fixes. At the 75th percentile, time to a typeable page on a fresh load went from 3.1 seconds to 0.55, starting a Claude Code session from 0.8 to 0.3, and loading a Cowork cloud session from 2.6 to 0.73 — changes that were merged as more than 3,000 individual commits with no customer-facing incidents or rollbacks, across four user journeys covering 95% of activity. The engineering detail that matters beyond the numbers is the loop: an internal research model roughly comparable to Opus 5.5 monitored deploys, built benchmarks and watched for regressions while humans set goals and approved every change.
  • Oracle is moving agent controls down into the database, on the argument that application-layer permissions cannot be trusted with an agent. Its database chief, Juan Loaiza, says the company runs multiple models over its own code and that AI is "literally superhuman" at finding vulnerabilities, both in volume and speed — which is why Oracle wants authorisation to live as close to the data as possible and applies consistently whether the caller is an application or an agent. The advice to customers is less glamorous: keep database environments current, because the patch treadmill just got faster.

What to watch: whether Thursday's summit produces anything on AI beyond a photograph, whether anyone outside Anthropic reruns its targeting evaluations against open-weights models, and whether OpenAI's new mathematics group changes how results get published rather than just who gets told about them.

Two labs asked the UN for binding international standards on the same day their own red team published evidence that a mid-tier open model can help locate and target people. If the labs can measure the risk but not stop it, who is supposed to? Tell us in the comments.

Sources: Anthropic — Measuring tactical intelligence targeting and conventional weapons capabilities · Anthropic Threat Intelligence report, September 2026 · OpenAI — Advisory Group on Mathematics and Artificial Intelligence · TechCrunch — OpenAI forms math advisory group · NPR — Mathematicians learn little from AI completing unsolved problem · The Decoder — Google's new Flash TTS models · Simon Willison — Gemini 3.8 TTS Playground · Zhidx — Qualcomm's 2nm flagships rebuilt for agentic AI · Qualcomm — Snapdragon for the agentic era · TechCrunch — Enveda secures $311M · TechCrunch — Bessemer raises $5.75B for AI · Cailian Press — Muse beta used human callers in place of the AI · The Verge — Meta's Muse agent hands-on · CNBC — Meta's standoff with Amazon over Muse · Techmeme — Flock Safety weighs a sale (Semafor) · OpenAI — Sam Altman's remarks at the UN Security Council · CNBC — OpenAI and Anthropic CEOs push for AI cooperation at UN · Techmeme — Altman and Amodei to the Security Council (Bloomberg) · Ars Technica — China silent as US touts AI safety alerts that omit tech experts · MIT Technology Review — A representative proposes killing the border tower program · WBEZ Chicago — Pritzker assembles an Illinois AI cabinet · TechCrunch — ChatGPT mobile app gets voice-based agentic features · Techmeme — ChatGPT Voice powered by Astra, Sol and Luna · SiliconANGLE — Talos finds malware that puts its next move to a four-model vote · Claude — How we made claude.ai 3x faster in two weeks · SiliconANGLE — Oracle shifts agent controls toward data-layer security