
The humanoid robotics industry has raised billions in funding over the past 18 months, but much of it is quietly paying human workers to remotely operate robots — a method that has become the dominant way to train physical AI systems. The problem is that teleoperation datasets are over 100,000 times smaller than those used for language and vision models, and because the real world constantly changes (doors vary, shelves move, new package types appear), every variation requires a new human demonstration, meaning the data problem grows faster than any workforce can solve. If robots truly require a permanent stream of human input to function, they are not replacing human labor — they are just a labor system.
Summaries like this, in your inbox every morning.
Sign up free →What happened
The humanoid robotics industry has raised billions in the past 18 months, but the majority is funding human workers to operate robots remotely — a method called teleoperation — rather than building truly autonomous systems. Teleoperation datasets are over 100,000 times smaller than those used to train today's language and vision models, and the gap cannot close by hiring more operators because the real world constantly changes, requiring new demonstrations faster than humans can provide them.
Why it matters
The original pitch for humanoid robots is to replace human labor due to aging populations and labor shortages. But if robots require a permanent stream of human demonstrations to function, they are essentially just a labor system rather than an autonomy solution. This means companies may be building infrastructure that deepens dependency on human workers indefinitely rather than resolving it.
What to watch
Reinforcement learning in simulation offers a potential alternative path. Unlike teleoperation, RL systems learn through trial and error across millions of iterations without human operators in the loop, and simulation allows that training to scale directly with compute — more GPUs enable faster iteration — rather than being limited by human labor availability.
Nikita Rudin, co-founder and CEO of Flexion (a company that recently raised $50 million(約80億円) to build AI systems for humanoid robots), argues that the humanoid robotics industry is caught in a data trap disguised as progress. In the past 18 months, humanoid robotics companies have raised billions of dollars, but the majority of that capital is quietly funding human workers to remotely operate robots — a practice known as teleoperation. This has become the dominant method for training physical AI systems, attracting serious capital and earning enthusiastic coverage as evidence of industry progress. However, Rudin contends that the assumption underlying this approach — that enough demonstrations will eventually produce robots capable of generalizing across real environments — deserves far more scrutiny.
The core limitation is scale. Language models trained on text can draw from decades of existing writing, articles, and books accumulated before the AI era. Robots have no such archive. Every demonstration must be generated by a human, meaning the data can only grow as fast as human labor allows. More critically, teleoperation datasets are over 100,000 times smaller than the datasets used to train today's language and vision models. That gap cannot be closed by hiring more operators because the real world is not static. A shelf moves, a door handle is slightly different, a new package type shows up on an assembly line — each variation requires a new demonstration. The problem grows faster than the workforce can expand to meet it. Data quality adds another layer of difficulty: operators cannot feel what they are touching or judge depth reliably, so they move slowly and overcorrect, forcing robots to learn from footage of someone struggling with a controller. That is what the robot ends up practicing.
The industry's response has been to recruit more workers, predominantly in lower-wage economies, to film household tasks, operate robots remotely, or move through facilities wearing camera rigs. A commercial ecosystem has emerged, with startups across China, India, Europe, and the U.S. selling teleoperation data the same way companies once sold labeled text for language models. The irony is sharp: the original pitch for humanoid robots is that they will fill jobs humans will no longer be able to do due to demographic shifts, labor shortages, and aging populations. But if what the industry is actually building requires a permanent stream of human demonstrations to function, those humans might as well do the task directly. A system that cannot handle anything new without fresh human input is essentially just a labor system.
The standard defense is that teleoperation is a bridge — a temporary method while better robot training techniques catch up. For narrow, repetitive tasks in controlled environments, that argument holds some water. However, much of the industry is building infrastructure to generate demonstrations indefinitely, with no clear account of how or when that dependency ends. The field is tracking metrics that feel like progress — demonstrations collected, hours of footage logged, tasks completed in controlled settings — but none of these metrics reveal whether a robot can handle something it has never seen before in an environment not specifically set up for it. Building more teleoperation infrastructure deepens the dependency on humans rather than resolving it.
Rudin points to reinforcement learning as the path forward. When language model researchers trained early systems on massive amounts of text, they produced systems that could imitate the style of Shakespeare but did not quite make sense. The breakthrough came through reinforcement learning in synthetic environments, which produced systems capable of reasoning, coding, and following complex instructions. The robotics industry is largely stuck in that early imitation stage. Scaling teleoperation data is equivalent to scaling pre-training on 100,000 times less data — you get robots that somewhat move their arms, sometimes grab something, sometimes do not, and can vaguely imitate what a human operator showed them but cannot reason through a situation they have not seen before. There are intermediate approaches — egocentric video capture and devices like the Universal Manipulation Interface (UMI), which lets operators demonstrate tasks more naturally by wearing a handheld gripper rather than controlling a robot remotely — that reduce operator burden and produce somewhat more natural motion data. These can be useful stepping stones for narrow, well-defined tasks, but they do not resolve the industry's full dependency on humans. Reinforcement learning changes this. Rather than imitating a human operator, a system trained with RL figures things out through trial and error, attempting a task, failing, adjusting, and trying again across millions of iterations without a human in the loop. Simulation follows naturally: running millions of RL iterations in the real world destroys hardware and takes years, but in simulation you can reset instantly, run in parallel, and generate variation at a scale no human workforce could match. Unlike teleoperation, RL in simulation scales directly with compute — more GPUs mean more environments, more variation, and faster iteration.
The robotics industry faces a structural challenge that most of its funding strategies do not address. While billions have flowed into humanoid robotics companies over the past 18 months, a majority of that capital is being spent on hiring human workers to remotely operate robots — a data generation method that the article describes as a permanent labor solution masquerading as a technical bridge. The underlying assumption is that enough demonstrations will eventually produce robots capable of generalizing to new environments, but the numbers do not support this path: teleoperation datasets are over 100,000 times smaller than those used to train today's language and vision models, a gap that cannot be closed by hiring more operators.
The core issue is that the real world is not static. A shelf placement changes, a door handle varies, a new package type arrives on an assembly line — each variation requires a fresh human demonstration. This means the data generation problem grows faster than any workforce can expand to meet it, creating a structural wall that more labor cannot overcome. Data quality compounds the problem: operators working through controllers cannot feel what they are touching or judge depth reliably, so they move slowly and overcorrect, teaching robots to practice struggling with a controller rather than performing tasks efficiently.
The alternative path the article identifies is reinforcement learning in simulation. Rather than collecting human demonstrations, RL systems learn through trial and error across millions of iterations in synthetic environments, with no human in the loop. This approach scales directly with computational resources — adding more GPUs enables faster iteration and greater environmental variation — whereas teleoperation remains bound by human labor. The article argues that for the industry to move toward genuine autonomy, it must measure whether human dependency is actually decreasing over time; without such metrics, teleoperation stops being a temporary bridge and becomes the permanent method.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
No comments yet. Be the first to share your thoughts!
Log in to join the discussion





Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.
Get Started FreeFree · takes 30 seconds · unsubscribe anytime
1 minute a day. The AI essentials.
200+ sources · Email / LINE / Slack