
RLWRLD last week presented RLDX-1, designed to tackle complex tasks in real-world industry using robotic hands. The model integrates a scalable data-collection pipeline, versatile architecture design, robust training methodologies, and optimized deployment strategies, and is deployable across single-arm, dual-arm, and humanoid embodiments.
RLDX-1 uses a Multi-Stream Action Transformer (MSAT) architecture where each sensory modality—torque (a high-rate continuous stream), video (sparse high-dimensional frames), and memory (stateful)—gets its own dedicated processing stream. The streams communicate through joint self-attention without being forced into a shared representation prematurely. A robot-specialized vision language model fine-tuned on robot visual question and answering targets spatial reasoning, task understanding, and action grounding.
On RoboCasa, the fine-tuned vision model achieved +3.42 percentage points over the vanilla model. On conveyor-belt pick-and-place tasks, the Motion Module achieved +37.5 percentage points over GR00T N1.6 and π₀.₅. The Cognition Interface achieved +35 percentage points inference speedup (16.3→22.1 Hz).
Ask the AI about this article →
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Geek+ reported interim results for the six months ended 30 June 2026

GE Aerospace's Lynn, Massachusetts facility is integrating robots and AI into its manufacturing operations, re…

Suntory Logistics and Toyota Industries announced on September 1 that they will begin a demonstration test in…

AWS Japan held a成果発表会 on August 31, 2026, where 13 projects from 14 companies presented results from its physi…

Google DeepMind unveiled Gemini Robotics 2, the latest vision-language-action (VLA) model, with whole-body con…

At RoboBusiness 2026, held Oct
