Ant Group launches Ling‑3.0‑Flash AI model delivering high performance with low parameter count
Ant Group announced the release of Ling‑3.0‑Flash, a next‑generation native hybrid‑reasoning foundational model aimed at production‑grade AI agent workflows. The model contains 124 billion total parameters but activates only 5.1 billion parameters per token, achieving performance comparable to larger models while using a fraction of the parameter scale.
Ling‑3.0‑Flash employs a hybrid‑linear attention architecture that alternates Kimi Delta Attention (KDA) and MLA layers in a 5:1 ratio, delivering a 256 K token context window that can scale to 1 million tokens. Optimisations include a compressed expert activation ratio (1/64) and enhanced long‑context efficiency, enabling rapid response, cost‑efficiency, and stable execution for tasks such as coding, task decomposition, and multi‑source research. The model has been trained on over 10 000 interactive environments and features self‑correction and long‑horizon planning mechanisms.
Entities: Ant Group · Hangzhou · Ling‑3.0‑Flash