Skild's S1 learns brand-new robot tasks from a single video — no retraining

Share
Skild's S1 learns brand-new robot tasks from a single video — no retraining

Robotics has its "just prompt it" moment — or at least the strongest claim yet. Skild AI, the Pittsburgh startup behind the $14 billion-valued Skild Brain, says its flagship foundation model S1 can execute tasks it has never seen in training after watching one video of a human doing them — with no fine-tuning and no post-training.

The company laid out the evidence in a technical blog post this week that is worth reading as a benchmark claim, not just a product pitch. Shown a single egocentric video of plant potting, S1 went from demonstration to autonomous execution on hardware in 11 minutes; it also handled unseen pancake flipping, pour-over coffee brewing, and kit assembly across ten-minute, multi-step horizons. On Skild's internal scaling study, the gap that matters is on out-of-distribution tasks: at 100k hours of pre-training data, an in-context policy hit a 66% per-step success rate where a conventional language-prompted VLA managed only 9%. The mechanism is the interesting part: instead of compressing instructions into language tokens, the demonstration sits in the model's context window like a prompt for GPT-3, so the weights never change — the task specification does.

The honest caveats are in Skild's own numbers. A 66% step-success rate still means roughly one intervention in three, which is why the study grades steps with humans recovering failures mid-rollout, and why Crypto Briefing's read of the results — 60–80% completion within hours of first data collection — frames this as impressive-but-not-production. Seen-task performance also favors classic VLAs at small data scale; in-context prompting only wins once pre-training gets big. But if the claim holds under outside scrutiny, the economics of robot deployment change shape: teaching a new task becomes minutes of demonstration rather than days of teleoperation, which is precisely the bottleneck keeping general-purpose robots out of warehouses. It lands at a moment when competitors are converging on similar ideas, and Skild has the war chest — $1.7 billion raised, NVIDIA and Amazon on the cap table — plus an NVIDIA-built training stack behind it.

We covered the previous state of the art in January — GEN-1.5 teaches robots new tasks from a single demo — and S1's bet is that scale plus in-context learning turns that party trick into a platform.

What to watch: independent replications of the 10-minute-horizon results, and the promised follow-up posts detailing how S1 was actually trained.

If your robot could learn one new skill by watching a single video, what would you teach it first? Tell us in the comments.

Sources: Skild AI · Crypto Briefing · Dealroom · Techmeme

Read more