AI RACE— The AI Race
New Models

Reka AI Unveils Rho-1, a 19-Billion-Parameter Omni-Model Combining Video and Robot Control

Reka AI has released a research preview of Rho-1, a unified 19-billion-parameter model capable of processing text, images, video, and direct robotic actions in a single neural network.

10/06/2026, 01:26
Reka AI ra mắt Rho-1: Mô hình omni 19 tỷ tham số kết hợp tạo video và điều khiển robot

Multimodal AI startup Reka AI has unveiled a research preview of Rho-1, a 19-billion-parameter omni-model capable of processing and generating text, images, video, and robotic actions within a single unified neural network.

The architecture departs from common industry designs that route requests across multiple specialized systems. Instead, Rho-1 treats all inputs and outputs as tokens within a single shared context window, operating without external models or tool calls.

Direct Integration of Vision and Robot Movement

Rho-1 is designed to support real-time interaction. The model can generate continuous video in real time while adapting to new instructions on the fly without needing to reset or restart its context.

Crucially, the exact same neural network weights responsible for predicting camera imagery also calculate and drive physical robot movements. By coupling perception and action into one model, Rho-1 connects visual understanding directly to physical manipulation.

Overcoming Data Bottlenecks with Inverse Dynamics

A long-standing obstacle in training embodied AI systems is the shortage of real-world robotic control datasets. To bypass this limitation, Reka AI engineered an inverse dynamics model capable of extracting underlying control signals directly from standard internet video footage.

This approach allowed the team to train Rho-1 using widely available web video rather than relying solely on dedicated robotic teleoperation logs. The training process required 320 Nvidia H100 GPUs running over approximately three months.

Advancing Toward World Models

The release highlights an ongoing transition in AI research toward "world models"—systems capable of simulating environments, understanding physical dynamics, and taking physical action based on real-time sensory inputs.

Reka AI previously made headlines in April 2024 with the launch of Reka Core, a multimodal language model designed to compete with frontier systems including OpenAI's GPT-4, Anthropic's Claude 3, and Google's Gemini Ultra. With Rho-1, the company is extending its core multimodal technology deeper into physical robotics and generative visual environments.

◗ Sources

The Decoder10/06

Related stories