AIToday
Large Language ModelsRoboticsTHE DECODERPublished: Oct 6, 2026, 04:00 JST

Reka AI's Rho-1 omni-model handles text, images, video and robot control

Reka AI's Rho-1 omni-model handles text, images, video and robot control

3 Key Points

  1. What happened

    Reka AI released a research preview of Rho-1, a 19-billion-parameter omni-model that processes and generates text, images, video and robot control actions in one neural network, trained on 320 H100 GPUs over about three months.

  2. Why it matters

    Unlike most AI systems that route tasks to specialized models with tool calls or external models, Rho-1 runs all modalities as tokens in one shared context window, which appears to simplify multimodal AI and enable real-time video with on-the-fly instructions.

  3. What to watch

    The approach hinges on whether one model can match specialized systems, but it is a research preview so production readiness is unproven; watch for updates on the inverse dynamics model that pulls control signals from ordinary internet videos.

WHO IT HITSThis matters to AI researchers and robotics teams exploring unified models, who may see new ways to handle multimodal tasks without specialized tools.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

Reka AI has a track record in multimodal AI: in April 2024 it shipped Reka Core, a multimodal language model that competed with GPT-4, Claude 3, and Gemini Ultra on benchmarks. The new Rho-1 release fits a broader push in AI research toward so-called world models, which aim to understand and simulate environments. Unlike most systems that route tasks to specialized models, Rho-1 runs all modalities as tokens in one shared context window with no tool calls or external models, and it generates continuous video in real time while responding to new instructions without restarting. The same weights that predict camera images also drive robot movements, with an inverse dynamics model extracting control signals from ordinary internet videos to work around scarce robot training data.

The outcome hinges on whether this unified approach can match or exceed specialized models in practice, and on how the research community responds to the trade-offs of a single model versus modular systems. For now, Rho-1 is a research preview, so its path to production and real-world deployment is likely still being explored.

FAQ
What is Rho-1?
Rho-1 is a 19-billion-parameter omni-model from Reka AI that processes and generates text, images, video and robot control actions in a single neural network.
How was Rho-1 trained?
Rho-1 trained on 320 H100 GPUs over about three months, and Reka AI built an inverse dynamics model to pull control signals from ordinary internet videos due to scarce robot training data.

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleFCC robot curbs push AI onto the machine