started · updated
HiDream.ai launches HiDream-O1-World interactive world model
HiDream.ai has launched HiDream-O1-World, a native omni-modal interactive world model capable of generating explorable 3D environments from text, images, or interactive controls. Built on the company’s proprietary Unified Transformer (UiT) architecture, the model aims to solve long-standing challenges in spatiotemporal and physical consistency.
Key technical breakthroughs include the ability to maintain stable scene geometry during camera movements and ensure that physical interactions, such as collisions and gravity, follow real-world causal logic. The model utilizes a memory mechanism to encode 3D spatial relationships and Test-Time Training (TTT) to maintain alignment with 3D constraints during inference.
In its debut on the WBench benchmark—a system developed by Meituan’s LongCat team and Fudan University—HiDream-O1-World topped the Navi sub-leaderboard with an average score of 80.9. It also ranked first in the Physical dimension and the Consistency dimension, outperforming several mainstream models.