started · updated
Ricoh develops physical AI to learn robot actions from video
Ricoh has developed a new physical AI technology capable of learning action representations from video data of humans or robots. This technology utilizes a Latent Action Model (LAM) to identify common behaviors, such as “picking up an object” or “moving forward,” from visual changes without requiring explicit labels or dependence on a specific robot's physical structure.
By using LAM to extract universal action concepts, Ricoh aims to overcome the high cost and data scarcity associated with traditional Vision-Language-Action (VLA) models. Currently, different robot configurations—such as humanoids or multi-joint arms—require individual, massive datasets for tuning. Ricoh’s approach allows for the reuse of human movement data to train robots, reducing the need for manual data collection from physical machines.
The development was supported by Amazon Web Services (AWS) Japan through its Physical AI Development Support Program, utilizing services like Amazon EC2 and Amazon S3 for large-scale data processing. Ricoh has successfully verified the technology through pick-and-place simulations and real-world testing using Sim2Real (Simulation-to-Real) techniques. The company plans to further evaluate applications in offices, factories, logistics, and maintenance sectors.