AIToday
Large Language ModelsRoboticsAI Business & IndustryIEEE Spectrum RoboticsPublished: Jul 15, 2026, 01:00 JST4 min read

Chinese robot startup X Square bets on unified stack for general-purpose machines

Chinese robot startup X Square bets on unified stack for general-purpose machines

Key takeaway

  • X Square Robot, a Chinese embodied-AI startup, has presented an integrated stack for general-purpose robotics built on three layers: data collection, a world model called WALL-WM, and an action model called Wall-OSS-0.5. The company argues that data quality and scalable infrastructure, rather than model size, are the real limiting factors for robots that can generalize across tasks.

  • By collecting demonstrations via inexpensive wearable rigs rather than expensive teleoperation, then validating them through physical playback on real robots, X Square reports achieving comparable performance to all-robot datasets at roughly a 20-fold lower collection cost.

  • With code now open-sourced and valuation above US $2.9 billion(約4600億円), the field will soon test whether these principles scale beyond the company's own benchmarks.

3 Key Points

  1. What happened

    X Square Robot, a Chinese embodied-AI company, has unveiled an integrated foundation stack for robotics spanning data collection, a world model (WALL-WM), and an action model (Wall-OSS-0.5). The company is releasing the code openly and argues this layered approach—organized around physical events rather than fixed time slices—solves robotics' core problem: how to build capability that transfers across tasks and machines. The company's valuation has climbed above 20 billion yuan (about US $2.9 billion(約4600億円)).

  2. Why it matters

    Robotics has long assembled separate perception, planning, and control modules that don't generalize, unlike large language models which scale predictably. X Square's bet is that data quality and infrastructure, not model size, are the real bottleneck for general-purpose robots. The company reports reaching performance comparable to an all-robot dataset at roughly a 20-fold lower cost by pretraining on robot-free human demonstrations captured via wearable rigs, then anchoring to real-robot data. If validated independently, this cost reduction could reshape robot training economics.

  3. What to watch

    The company's strongest results are currently measured on its own robots and benchmarks. With code now released, the broader robotics community will test and reproduce these capabilities across different robots, tasks, and settings—a crucial step to confirm whether the stack's principles hold beyond X Square's controlled environment.

Ask the AI about this article →

Context & Analysis

The robotics field has long struggled with a fundamental problem that large language models solved: how to build systems that generalize. Where LLMs benefit from pretraining on broad data and then fine-tuning for specific tasks, robotics has historically cobbled together separate perception, planning, and control modules that rarely transfer knowledge across different robots or tasks. X Square Robot's explicit wager is that the solution lies in an integrated stack where data infrastructure, world modeling, and action generation are tightly coupled from the ground up.

The company's diagnosis—that data quality, not model size, is the real bottleneck—reflects a pragmatic shift in how the embodied-AI field is thinking about scaling. By using inexpensive wearable rigs to capture human demonstrations, then validating them through physical playback on real hardware, X Square claims to have cracked a cost problem that has long made robot learning expensive. The reported 20-fold reduction in data collection cost, if confirmed independently, would reshape the economics of robot training. Equally important is the architectural choice to organize the world model around semantic events—coherent behaviors like grasping or placing—rather than fixed time windows. This design respects what large video models already know while still yielding executable motion, a middle ground between pure prediction and pure control.

The fact that X Square's valuation has exceeded US $2.9 billion(約4600億円), and that the company is releasing code openly, signals confidence among investors that data infrastructure and foundation models will be long-term differentiators in embodied AI. However, much of the current evidence comes from the company's own robots and benchmarks. The next phase—independent testing and reproduction across a wider range of robots, tasks, and settings—will determine whether these principles truly generalize or remain specific to X Square's carefully controlled environment.

FAQ

How does X Square reduce the cost of training robot data?
The company collects demonstrations using a wearable rig with dual grippers instead of teleoperating robots, which is much cheaper. It then pretrains on this large volume of robot-free human data to build general representations, and adds only a small amount of real-robot data as an anchor to the specific machine's dynamics. The company reports this hybrid approach reaches performance comparable to an all-robot dataset at roughly a 20-fold lower cost of collection.
What makes X Square's world model (WALL-WM) different from standard approaches?
Most action models predict fixed-length chunks of motion, which segments behavior into windows determined by elapsed time. WALL-WM instead treats an action-grounded semantic event—such as reaching, grasping, or placing—as its unit, allowing variable-length segments suited to reasoning over long horizons while still producing the steady output a controller needs for real-time execution.
Why is the company's physical playback quality control step important?
Rather than accepting recorded trajectories as valid data, X Square physically replays a sample of them on a real robot and only counts those that actually complete the task. This ensures the dataset reflects what truly works; for example, a gripper that closes too early may look like a grasp in the recording but has actually failed the task, so it is not included.
IEEE Spectrum RoboticsRead Original Article

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Visko raises $10M, launches live AI video model OrbisSiliconANGLE AI · 33m ago
  • Runway unveils Solaris, an AI that generates app interfaces in real timeTHE DECODER · 33m ago
  • Google AI Search flags Facebook users as dangerTHE DECODER · 33m ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleDelaware proposes AIC legal entity for autonomous AI agents in regulatory sandbox