Reka has released a research preview of Rho-1, a 19B parameter model trained from scratch that processes text, images, video, and robot actions within a single neural network. This architecture aims to replace traditional multimodal pipelines by eliminating handoffs between specialist models, allowing all modalities to be handled as tokens in one context window.

  • The model uses two expert streams (understanding and generation) that share attention and a single KV cache in every transformer block.
  • Discrete tokens handle text and reasoning, while continuous tokens manage image latents, video frames, and robot actions.
  • A distilled variant reduces denoising steps from 99 to 8, generating a 5.3-second clip in about one second with minimal quality loss.
  • The base model generates video at 0.79x real-time speed, with the first clip appearing in roughly 6 seconds.
  • For robotics, Rho-1 emits continuous action tokens and pairs with an Inverse Dynamics Model to infer control signals from raw video.

By collapsing the multimodal stack into a single model, Reka claims this approach reduces latency and allows for real-time steering and continuous rollouts without external tool calls or secondary models.