Figure's Helix 2.5 tidied 30 homes it had never entered

Share
Figure's Helix 2.5 tidied 30 homes it had never entered

Figure sent a humanoid into 30 Bay Area houses it had never trained on and had it tidy rooms, fold towels and make beds with no fine-tuning — 56% success, six times the 9% a from-scratch policy managed. China's provinces keep building token marketplaces, and CapCut put an agent inside the editing timeline.

Figure's Helix 2.5 is the first real evidence that whole-body robot behavior learned from human video transfers to homes the robot has never seen. The evaluation ran across 30 Bay Area homes with a single fixed checkpoint, no weights adapted to the evaluation homes or objects, and no data collected in any of them. Each task — tidying 13 to 15 scattered toys, folding towels to a corner-alignment grade, making a bed to a fixed rubric — was scored pass or fail with no partial credit, and any rollout needing human intervention for safety counted as a failure.

The controlled comparison is the number that matters. Figure trained two policies on identical task-specification data, one from random weights and one initialized from the Index-pretrained model, holding architecture, optimization, hyperparameters and evaluation fixed. Random init succeeded on 9% of zero-shot trials; Index pretraining on 56%. Because pretraining was the only variable, that gap is the cleanest measure yet of how much of a robot's competence comes from watching humans rather than from the site it will work in. Figure adds that Helix 2.5 matched a Helix 02 policy trained directly in its evaluation environment while using half as much adaptation data.

Figure also claims the property that made language models forecastable now applies to physical intelligence. It trained four models on nested subsets of Index spanning an eightfold increase in pretraining data, and held-out action-prediction loss fell smoothly enough that the team forecast the largest run's loss to four decimal places — a forecasting error of 0.54% of the total variation observed across that range. Index now generates roughly 35 minutes of new human experience every second, and Figure has committed $3.5 billion of compute to Helix training. No single evaluation task makes up more than 1.90% of the pretraining data, which is the breadth argument in one number.

The caveats are real and Figure states them plainly: the evaluation is internal with no independent replication, 56% still means the robot fails roughly 44% of the time, and the link that would turn "deployable home robot" into a computable resource problem — correlation between next-action prediction loss and actual task success — has not been measured. We looked at the other half of this problem, robots learning the physics of their own actions, in Deep Dive — Robots are learning what their actions do, not what to do.


China's provinces are turning token purchases into retail infrastructure: Hubei's state big-data group launched a token trading center on September 20, listing more than 200 model services from 326 ecosystem partners, and estimates it cuts enterprise AI costs by 15% to 30%. The pitch is one gateway and one bill — a unified protocol, single sign-on, key escrow, content filtering and full-chain audit, with every model's usage metered into a single token account and a dashboard showing which department spent what. It joins Inner Mongolia's green-compute token platform and Shaanxi's Silk Road token exchange in a national push to make tokens a traded commodity, which is a step past API resale: whoever runs the exchange sets the floor and owns the meter. We've tracked the demand side of this — Zhipu's API revenue jumps 27x as China's AI labs prove the token economy.


CapCut put an agent inside the timeline. At its September 20 launch, ByteDance's editing app introduced CapCut Hub — a workspace with an infinite canvas and multi-track editor that pulls generated images, video and audio in from ByteDance's other AI tools — plus a desktop assistant that takes over asset sorting, spoken-word trimming, subtitle correction and packaging, with custom Skills for creators who already have a workflow. The mobile version ships as Xiaoying, which doubles as a marketing assistant for merchants. The tell is the monetization: a new AI Ultra tier folds AI credits and professional editing perks into one subscription, alongside roughly 150 million yuan a year in creator incentives. CapCut says 67.68 million users exported a first project in the past year and 1.66 million small merchants used it — the installed base, not the model, is what makes an editor agent a distribution event.

What to watch: whether anyone outside Figure replicates the 30-home test, and whether the next doubling of Index data moves task success the way it moves prediction loss.

Would you let a robot that fails 44% of the time into your house while the industry works that out, or wait for the correlation to be measured? Tell us in the comments.

Sources: Figure — Helix 2.5: Zero-Shot 30-Home Generalization · Figure — Introducing Index · Unite.AI · TechTimes · Hubei Daily (湖北日报) · SASAC — Hubei Big Data Group's AI supply model · China Economic Weekly (经济网) — CapCut's new AI capabilities · Sohu — CapCut AI launch · 10jqka — CapCut launches AI Ultra subscription