Today in AI — October 10, 2026

A day where the boring machinery took center stage: proof checkers, order books, warehouses and the humble text message — plus fresh arguments about who gets to declare any of it safe.
Models & Research
- Mathematicians get a reliability primer for the Lean Theorem Prover. A guest post by Thomas Hales on Terence Tao's blog walks through what Lean actually guarantees — and devotes itself to the "Summer of Soundness Bugs," the string of kernel bugs found this summer that let false proofs through, including one that produced an illicit "disproof" of the Collatz conjecture and another that faked a proof of the Kepler conjecture. The good news is the telling part: frontier models in the hands of security researchers found the bugs — Dan Selsam's AI work with the Lean team turned up further kernel issues until it could find no more — and all were repaired with mathlib re-verified. The post's bottom line holds regardless: trust a Lean proof only after the kernel checks it and a human audits that the statement says what we think it says.
- Jane Street tested whether diffusion models can generate market data — and mostly can't. The quant firm's engineers trained autoregressive diffusion on four years of US equities order-book events, only to find the fundamental problem is that market data is neither continuous nor discrete: it has the discrete action space of token picks and the continuous drift of prices, and fully continuous diffusion smoothed away exactly the jaggedness that makes real book events. Their fixes were partial — DDPM diverged at high noise while flow matching held up — and the working design ended up hybrid, sampling the event kind from one head while a diffusion head generates the continuous features. It's a clean negative result for anyone assuming image-generation recipes transfer to finance.
- An open-source tactile sensor grows into a real ecosystem. Columbia University's Li Yunzhu team gave an IROS 2026 update on FlexiTac, a flexible piezoresistive tactile sensor whose full bill of materials and assembly instructions are public: build it in about three minutes, bend or cut it to fit a gripper or a robot's skin, and it plugs into the LeRobot learning framework. Eighteen months after the open source release the stack has spread across academia and industry, a Purdue collaboration on dexterous multi-finger gripping is a best-paper nominee at this year's conference, and Analog Devices is building a commercial version targeting roughly five times a human fingertip's tactile resolution. Tactile data has lacked the standardization vision has — this is the closest thing to a common sensor lab.
Industry
- AI is putting the law firm's billable hour on trial. Almost 90% of legal professionals in the UK and Ireland now use AI, according to Clio's 2026 legal insights report, and among firms using it, nearly 80% say they handle more work without adding resources. Deloitte's survey puts a number on the shift: the share of legal work charged by the hour is expected to fall from 72% to 44% over the next two to three years. UK magic-circle firms are already living it — A&O Shearman built AI agents for loan-agreement review with Harvey, Slaughter and May rolled it out firm-wide — while the harder question is training, since the document review that once made junior lawyers is exactly the work disappearing.
- Walmart's warehouse robots keep tripping over the real world. The Wall Street Journal's look at Walmart's push to automate its roughly 200 US warehouses finds setbacks rooted in the unglamorous stuff — cardboard boxes, turkeys, dog food — pushing the effort toward what participants call "peak complexity." Walmart owns 12.6% of lead partner Symbotic, so every slip in the rollout is also an investor problem for a company whose stock is priced on exactly this promise. Automation at this scale is an integration problem, not a robotics demo.
- Synopsys opens the door to Chinese AI labs. The chip-design software leader says it is exploring partnerships with Chinese AI labs to build AI-powered chip design tools for the Chinese market, per Nikkei Asia — a notable positioning move from the company whose software sits under most of the world's advanced chip design. If EDA vendors start leaning on frontier models for layout and verification, the tooling layer of the semiconductor industry gets an AI arms race of its own.
- Prime Agent rewrote itself in Rust — with 2,000 agents doing the work. Prime Intellect says its coding agent orchestrated a swarm of over 2,000 agents across more than 10,000 sandboxes, burning over 200 billion tokens of GLM-5.3 compute over two weeks to port its own codebase from TypeScript while chasing feature parity. The motivation is mundane and telling: Prime Agent has been downloaded over 300,000 times since August and processed over eight trillion tokens, so trimming memory, startup time and per-process runtime overhead now pays directly. Self-improvement, it turns out, is also cost engineering.
- A fuel-cell startup bets data centers can't wait for the grid. Petra Power, a 15-person company founded in 2017, sells solid oxide fuel cells — ceramic devices that turn fuel into electricity without combustion — pitched at data centers and defense vehicles that need power without the wait or the footprint of traditional generation. It has taken nearly $9 million in Defense Department contracts and is talking to neoclouds and infrastructure providers, with first deployments targeted for 2028 and full-scale production for 2029; on the vehicle side, nothing has deployed on a live vehicle yet.
Policy
- Inside the Slack channel where Washington meets AI vendors. CBS News profiles a 1,700-member Slack run by Medicare agency CMS, where executives from Microsoft, OpenAI and other companies have spent the better part of a year quietly helping shape policy on AI applications and access to medical records. It is a layer of influence that never shows up in a hearing: vendors and officials negotiating the plumbing of health AI in a group chat while Congress argues about frontier models.
- Who gets to say an AI system is safe? A SiliconANGLE column argues the industry has a "control gap" between what AI can do and what the evidence supports trusting it to do — a great demo proves nothing about staying inside a deployment's limits. The prescription: independent experts examine the evidence labs produce, and an authority with teeth can require remediation, restrict an activity or stop it, with the clarity of knowing what must be demonstrated beforehand — banking- and power-plant-style obligations rather than voluntary promises. Appian CEO Matt Calkins, the interview subject, wants the same standard applied to the agents running inside ordinary businesses.
Tools
- The AI agent you text instead of open. A new survey counts roughly two dozen agents that live in your messages: Caddy sits in iMessage and RCS to turn scattered emails and texts into calendar actions, the open-source Comma works across your browser, computer and files while staying reachable over Signal, Telegram and WeChat, and Instinct — fresh off a $1 billion round at a $10 billion valuation — is the segment's best-known name. No app to download, no new inbox: you text them like a person and they carry tasks to completion. Messaging is quietly becoming the interface layer for agents.
- REA hands your coding agent a disassembler. The project gives agent harnesses the tools to take a program apart and explain what it does — pitched at everything from understanding unfamiliar code to hunting hidden functionality — and connects to Claude Code, Codex, Cursor, Gemini CLI and most other coding agents rather than binding to one vendor. It drew a sizable Hacker News crowd this week, with commenters already pointing at application-security reviews as the obvious second act.
- The AI bird feeder that lost to a pigeon. A Verge review of the Kiwibit Bird Feeder 2 Pro found the 4K camera, species identification and migration facts genuinely fun — until a wood pigeon parked itself in front of the lens, ignored the device's "Nuisance Animal Alarm," and ate for 15 minutes straight, followed by mice underneath. The reviewer gave up and put the feeder in a closet: delightful until it wasn't, which is a fair epitaph for a lot of AI hardware.
Is the billable hour the first professional norm AI actually breaks? Tell us in the comments.




