started · updated
Artificial Intelligence evolves toward Large Multimodal Models
Artificial intelligence is evolving beyond Large Language Models (LLMs) toward Large Multimodal Models (LMMs), which integrate various data types such as text, images, audio, and video. Unlike standard LLMs, LMMs can relate different types of information within a single task, such as analyzing a photographed document or describing visual content.
Technical reports for models like Qwen3-VL and Qwen3-Omni illustrate this shift. Qwen3-VL incorporates visual information across different layers and uses timestamps for video processing, while Qwen3-Omni processes text, images, audio, and video to generate text and voice.
Parallel to multimodality, the field is also advancing through the convergence of supervised learning, reinforcement learning (RL), and evolutionary algorithms. These methodologies address complex computational problems that exceed natural language inference, utilizing techniques like Monte Carlo Tree Search (MCTS) and genetic algorithms to optimize decision-making and multi-objective problems.