Skild's S1 teaches robots a 10-minute task from one video

Share
Skild's S1 teaches robots a 10-minute task from one video

Robots normally need thousands of teleoperated practice runs to learn a single job. Today two stories pointed the other way — a foundation model that learns by watching once, and a European startup betting €28 million that the transformer is not the end of the story.

Skild AI's new S1 foundation model picks up a task the way a person does — by watching it done once. Show S1 a single egocentric video of a human demonstration and the robot reproduces the task on real hardware within minutes: in the company's own launch timeline, a plant-potting demonstration recorded at 9:22 PM had S1 executing autonomously by 9:27 PM. The tasks it was never trained on are the striking part — potting a plant, flipping a pancake, making pour-over coffee, assembling a kit — runs up to ten minutes long and dozens of manipulation steps, composed from primitives the model absorbed during pre-training. NVIDIA's blog today frames S1 as part of its Physical AI stack, with the model trained on NVIDIA infrastructure.

Why it matters: the standard robotics recipe is collect teleop data, fine-tune per task, redeploy. Skild's own comparisons argue that is the slow path — a conventional language-prompted VLA policy degraded up to three times as much as S1 when scenes shifted, and the team post-trained a VLA with between one and 2,000 demonstrations to try to match what S1 got from a single video. The take: language was always an awkward interface for robot hands; demonstration is the native prompt. If the company's in-context learning scaling laws hold — every additional hour of data buys more robust primitives — the robotics data argument flips from "collect more teleop" to "let the model watch more humans." We covered why robots need to learn what their actions cause, not just what to do — Deep Dive — Robots are learning what their actions do, not what to do. S1 is already running with commercial partners; the test now is whether third parties can rerun the results.


Paris-based Arlequin AI raised a €28 million Series A to build models on topological neural networks — a deliberate step away from the transformer. The round was co-led by Redalpine and OTB Ventures with Bpifrance's Defense Innovation Fund, and the company says European investors exclusively backed it. Where mainstream models treat input as sequences, Arlequin's architecture learns from how data points connect — multi-path relationships across documents, transactions, video and operational logs — and traces outcomes back to root causes. The target use cases are where that audit trail matters: counterterrorism analysis over billions of data points from seized devices, fraud and money laundering detection, cybersecurity. The pitch that should raise eyebrows in 2026: the architecture is designed to need significantly less compute than mainstream models, with research collaborations spanning INRIA, CNRS, the Max Planck Institute, Oxford, Cornell, Princeton and UC Santa Barbara.

The round is small by this week's standards, but the interesting claim is not the money — it is that relationship-structured learning can be cheaper and more auditable than pure scale, in markets where governments pay for explainability.

What to watch: whether Skild publishes S1 benchmarks third parties can rerun, and whether a topological-network model shows up in a real defense or fraud-detection deployment rather than a deck.

Would you trust a robot that learned a job from one video? Tell us in the comments.

Sources: Skild AI · NVIDIA Blog · Arlequin AI · SiliconANGLE · Tech.eu · OTB Ventures