started · updated
Generative AI advancements include real-time video editing and lightweight agents
Several new developments in generative AI and agent technology have been highlighted. JD.com’s research division, Joy Future Academy, has introduced ‘JoyAI-Video-Edit’, a real-time video editing model. Unlike traditional offline models that process entire videos at once, this 16B parameter autoregressive diffusion model processes video in small chunks. This allows for continuous editing in streaming environments without significant memory increases or quality degradation. It can process 720p resolution at approximately 30FPS using a single Nvidia B200 GPU.
In the field of AI agents, Alibaba’s research team announced ‘LongHorizon-Harness’, a system designed to prevent the accumulation of errors during long-term tasks. This method addresses the issue where agents lose sight of their original objectives as their processing history grows.
Other notable releases include a lightweight version of the ‘MiniMax-H3’ video generation AI for ComfyUI, which is optimized for lower VRAM, and Liquid AI’s ‘LFM2.5-2.6B’, a lightweight AI agent capable of running on smartphones with performance comparable to models four times its size. Additionally, the Rust-based library ‘anydoc’ was noted for its ability to rapidly convert various document formats like Word and PDF into Markdown.