Stanford's kitchen robot demo has no paper, no code, and two tasks
The weekend wire produced one genuinely new robot system, one number that decides the next two years of AI spending, and one executive saying out loud what every lab does quietly.
Stanford and Caltech researchers have deleted the trained control layer between a language model and a humanoid, and the demo works — there is just no paper, no code and no numbers behind it. HomeBody puts GPT-6 Astra in charge of a Unitree G1: the model explores an unfamiliar kitchen, uses Astra as a Real2Sim agent to rebuild the room as a digital twin in Nvidia's Isaac Sim, logs objects and their locations in spatial memory, then composes a five-skill library — pick, place, open drawer, pick from drawer, navigate — to tidy up and to fetch medicine from a drawer it can no longer see. The architectural claim is the interesting part. Instead of the usual chain (a reasoning VLM feeding a learned vision-language-action policy feeding a whole-body controller), HomeBody has the frontier model call skills directly, with execution feedback letting it retry when a grasp or a transition fails. That is the bet several labs are now making — and if it holds, the expensive part of robotics stops being policy training.
The evidence is thinner than the coverage implies. The project page's Paper link is a disabled pill reading "Coming soon"; the GitHub repository currently contains only website assets, with a README that says "Code coming soon"; the page shows two tasks and marks the rest "coming soon"; and the demonstration clips are sped up, at 1.25× on the individual skills and faster elsewhere. The stated limits are refreshingly honest — Astra's reasoning latency introduces pauses between skills, finger servos overheat in extended operation, and the Real2Sim step carries API cost — but nothing on the page measures success. The question that goes unanswered is refusal. The same frontier model completed 60 of 100 genuinely dangerous robot tasks in the RoboHarm benchmark we covered this month — The most capable robot refused 2 of its 100 harmful orders — and HomeBody's videos show a robot deciding for itself what to throw away in a room containing knives, toasters and bleach. Treat this as a capability demo, not a result.
Goldman Sachs has put a break-even figure on the AI buildout: the hyperscalers need roughly $300 billion a year in AI revenue just to stop losing money on what they are building. The number comes from Goldman strategist Ryan Hammond, whose team also expects hyperscaler AI capex to reach $1.2 trillion in 2027 — above Wall Street's $1.1 trillion consensus — with spending now exceeding what those companies generate from ongoing operations, which is why debt financing is climbing. About $1 trillion a year in end-user AI spending is what Goldman says would produce solid returns.
What needs correcting is the version of this circulating in Chinese media over the weekend: "AI revenue must reach $636 billion to cover capital expenditure." That figure is real, and it is mislabeled. It is the top row of Goldman's own return-sensitivity table, where $308 billion corresponds to a 0% return, $417 billion to 10%, $526 billion to 15% and $636 billion to 20%. Break-even is the $308 billion row, not the $636 billion one — the two get conflated whenever the table is screenshotted without its labels. We have spent the month writing around the buildout's physical limits — Deep Dive: Power, not GPUs, now sets the pace of AI — and this is the same story told in dollars: the constraint is not the capex, it is the revenue that has to show up behind it.
A senior OpenAI researcher says 80 to 90 percent of the company's research now targets GPT-7, GPT-8 and beyond, and calls the within-generation updates everyone else treats as launches "extremely shortsighted." Boris Power, OpenAI's head of applied research, said it at the Fellows Forum in Menlo Park on September 23, in a session carried on the conference's public Day 1 stream. His framing is that incremental releases like 5.1 to 5.2 come from specialized training data and are deliberate short-term bets — they help the company iterate faster today without being the long-term strategy, and after each generational jump the team has to relearn where the quick wins are. He also puts the bottleneck for assistants in onboarding rather than model quality, on the grounds that most ChatGPT users do not know what the product can already do.
The caveat matters: those percentages arrive as one outlet's paraphrase of a livestream, not a transcript or a filing, so read them as a signal about internal priorities rather than a measured statistic. Even at half the stated figure, it is a useful corrective to the assumption that labs are still polishing the current generation.
What to watch: whether the HomeBody repository ever ships code, and whether anyone at OpenAI puts the 80-to-90 percent figure in writing.
If a frontier model can be handed a kitchen and told to clean it, who is liable when it decides the bleach is rubbish? Tell us in the comments.
Sources: HomeBody (Stanford Movement Lab) · HomeBody (GitHub) · The Decoder — HomeBody · Robocurve — RoboHarm · Yahoo Finance — Goldman Sachs · Crypto Briefing — Goldman's AI capex phases · The Decoder — Goldman Sachs · Sina — Goldman's ROIC table · The Decoder — OpenAI targets GPT-7 · Fellows Forum 2026, Day 1