GPT-6 Astra drove a real car through a cone course — its rivals didn't

Share
GPT-6 Astra drove a real car through a cone course — its rivals didn't

Three stories today about machines taking the controls: a language model steering a real Corolla, Beijing counting the switches in its own data centers, and an editor you brief in plain English instead of scrubbing timelines.

A group of independent researchers put frontier models in the driver's seat of a real Toyota Corolla — and GPT-6 Astra is the only one that finished the course. DrivingBench, a project by Aditya Ramabadran, Simon Mahns and Tobias Gessler, gave each model control of the car's steering, accelerator and brakes through tools built on comma.ai's openpilot stack, with the model reading the cameras and telemetry and issuing steering and speed commands on a cone course in an empty parking lot. A human sat in the driver's seat ready to brake, speeds stayed low, and each model got up to three attempts inside one continuous chat. Astra's first attempt ended in a collision at 49% of the course; on the second, it completed all 134.7 meters in 5 minutes 22 seconds, burning 246.6 million tokens and $7.74 including its post-run reflection. Claude Fable 5.1 climbed from 9% to 45% across three attempts without finishing; Grok 4.6 managed 11% and GPT-5.6 Sol 6%, barely moving between tries.

The honest framing is on the site itself: the authors say on Hacker News that the models "clearly aren't good enough to drive on an actual road yet" — this is a low-speed parking-lot test with a human safety net, and commenters note a person could walk the course in about fifteen seconds. But the benchmark captures something real: closed-loop control of physical hardware, where a bad decision is a crash, not a wrong token. The improvement across attempts is the interesting signal — Astra went 49% to 100% and Fable 9% to 45% within a single conversation, which is in-context learning applied to torque, not just trivia.

One commenter's detail is the best AI-safety anecdote of the week: Astra reportedly refused to drive at first because it recognized it was controlling a real car, and only complied once the tooling was renamed a "sandbox." Evaluation traces and the harness are public, so the leaderboard can be checked rather than believed. We've tracked Astra's agent record all year — GPT-6 Astra turns $15,515 running a vending business — and this is the first one with a chassis. It also lands days after NHTSA's probe into Comma is a test for open-source driving: the software layer this benchmark builds on is now under federal investigation for how it behaves on real roads.


China's state-assets regulator is counting how many Broadcom switches sit inside state-backed data centers — a survey that reads as the opening move against the last American monopoly in the AI stack. The Financial Times reports that SASAC, which oversees state-owned companies, has spent recent weeks surveying how much of their networking gear comes from Broadcom, whose switches move data inside AI clusters the world over. According to the FT's sources, the surveys found Broadcom equipment making up as much as 90% of networking hardware at some state-owned firms, and regulators are examining whether the company used that position. It pairs directly with this morning's Nscale filing — The chip ban's biggest hole is a rental agreement — export controls squeeze the top of China's AI hardware from one side while procurement independence sweeps the bottom from the other. Switches carry no export-control halo: swapping them is a purchasing decision, not an engineering moonshot.


YouTube's next AI feature edits video the way you'd brief an intern: conversationally. Announced at the Made On YouTube event, the tool takes natural-language instructions in a chat window — pick the best take of each outfit, add music that fits the dance, cut the dead air — and assembles the edit across long-form video and Shorts, with text overlays and pause removal in the same pass. The pitch, demoed by creator Happy Kelli: a one-hour shoot can eat ten hours in the edit, and the AI acts as an "infinitely patient helper" you can overrule by jumping back into the manual timeline whenever you want control. It ships in early 2027 in Shorts and the Create app. Coming a day after YouTube's new AI agent will run your channel's back catalog, the pattern is unmistakable — the agent took the audience, the editor gets the workflow, and both ends of a creator's job are being rebuilt around conversation.

What to watch: whether anyone replicates DrivingBench's leaderboard with the public harness — and whether SASAC's survey ends in domestic-switch procurement quotas.

If a language model can finish a cone course in a Corolla, what's the first real-world job you'd hand one with wheels? Tell us in the comments.

Sources: DrivingBench · Hacker News discussion · DrivingBench harness (GitHub) · Financial Times · Reuters · TechCrunch · Made On YouTube (Google)