
Skild's S1 robot model learns new jobs from a single video.
It scored 66% with demonstration prompting versus 9% with language.
Matching that performance via training would need about 380 demonstrations.
What happened
Skild AI says its S1 robot foundation model can watch a human demonstrate an unfamiliar job once, then attempt it on a robot without updating model weights. Examples include potting a plant and flipping pancakes, with tasks lasting up to ten minutes.
Why it matters
Skild estimates that matching S1's single-video performance through additional training would take about 380 teleoperated demonstrations, equivalent to 50–100 hours. This suggests teams could spend less time assembling an initial task dataset and begin evaluating robots sooner, though actual deployment savings remain unmeasured.
What to watch
In Skild's controlled comparison at 100,000 hours of pretraining, demonstration prompting scored 66%, compared with 9% for language prompting. However, the score averages cumulative success across task steps and includes human recovery interventions, leaving unresolved how often a robot completes an entire job without assistance.
Ask the AI about this article →
Summaries like this, in your inbox every morning.
Skild introduced S1 on August 18, with the announcement appearing in The Rundown's September 3 newsletter. The core idea is to apply a heavily pretrained foundation model (100,000 hours of pretraining) to new extended tasks by supplying a video as an input prompt, rather than updating weights. This positions S1 as a test of rapid adaptation for multi-step jobs that can take up to ten minutes.
The company's controlled comparison gives scale to the potential savings: one video prompt delivered performance that would otherwise require roughly 380 teleoperated demonstrations (50–100 hours). However, Skild's metric averages success across steps and includes human recovery interventions, meaning full-job completion rates without assistance remain unclear. A practical workflow may involve demonstrating a task, assessing performance, then collecting targeted additional data where reliability falls short.
Generalist offers a related example, reporting 59% average success across ten simple, short tasks from one 3–12 second demonstration without weight updates, and 83% after ten gradient steps with five minutes of task data. Skild's larger demonstration collection also exceeded its single-video score. This suggests hybrid approaches — initial video prompting followed by focused training — could balance speed and reliability, though how well these methods generalize across new environments and hardware remains an open question.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
In January, Ukraine's defense ministry said it would share millions of data points from tens of thousands of d…

Caterpillar has entered a collaboration with FieldAI to advance physical AI, autonomy, and robotics on jobsite…

Aptiv announced support for NVIDIA Jetson Orin Nano 2, extending its edge AI collaboration for drones, robotic…

OpenAI CEO Sam Altman said in an August 26 TIME interview that OpenAI would “definitely” make humanoid robots

Lockheed Martin and its partners demonstrated an AI-enabled 5G system designed to protect against drones

Yokohama Rubber has introduced a generative AI system that uses RAG (a technique that lets an AI draw on a com…
