What changed
2026-08-03 01:44 UTC → 2026-08-04 10:26 UTC ·
added
removed
DeepSeek continued its DeepSeek’s rapid rollout of cost‑efficient, open‑weight models. On 1 models continued through July 2026 it announced 2026. After the July 1 launch of V4 Flash, Flash – a 284‑billion‑parameter LLM that runs on high‑end laptops via a Mixture‑of‑Experts design that activates only about 13 billion parameters per token. The model, released under an MIT licence, targets devices such as the MacBook Pro M3 Max and Nvidia DGX Spark, and is accompanied by a developer guide for the V3 API model with a 64 K‑token 1‑million‑token context window and competitive per‑token pricing. A leaked three‑hour investor call pricing of 1 yuan in, 2 yuan out – the firm announced a full‑version V4 on 27 July 2026 reiterated DeepSeek’s strategy: prioritising artificial‑general‑intelligence research and open‑source releases over short‑term profit, while acknowledging that 21, claiming 40‑50× lower compute power remains the chief bottleneck. The call confirmed a recent external funding round that valued the company at roughly $52 billion cost than leading U.S. systems and outlined plans pricing up to build larger, domestically sourced compute clusters. On 31 July 2026 DeepSeek made 70 times cheaper than rivals. Alibaba followed with the public V4 Flash‑0731 model available. Early benchmarks show it surpasses open‑source rivals Qwen 3.8 max, positioning it as a lower‑price, cloud‑integrated alternative amid fierce competition from other Chinese models such as GLM‑5.2 and approaches the performance Moonshot’s Kimi K3. Parallel to these releases, a wave of premium proprietary systems, while per‑token costs technical guides made a variety of Chinese open‑source LLMs runnable on consumer hardware. DeepSeek‑V4 (7 B, 8K context) and Qwen 3.6‑27B were cut to 1 yuan packaged in GGUF format for inputs PCs with 12 GB VRAM; Qwen 3.5‑9B‑AWQ, Gemma‑4‑12B‑it, Gemma‑4‑26B‑A4B‑it, Qwen‑3.5‑2B, GLM‑5‑FP8, Qwen 3.5‑9B‑MLX‑8bit, and 2 yuan for outputs. Qwen 3.5‑9B‑NVFP4 each received step‑by‑step installers, quantisation (AWQ‑INT4, FP8, NVFP4) and hardware‑auto‑tuning scripts targeting CPUs, GPUs and even Windows‑only deployments. A financing round of about 500 billion yuan lifted July 17 playbook showed how Alibaba’s 397‑billion‑parameter Qwen 3.5‑MoE could be accelerated on Google’s Ironwood TPU v7x, achieving 3‑fold decode throughput gains. DeepSeek’s CEO Liang Wenfeng reiterated the firm’s valuation to over 3.5 trillion yuan compute gap with U.S. firms – roughly 20,000 H100‑class GPUs versus the 50,000‑plus needed for parity – but pledged continued expansion of domestic clusters and prompted open‑source releases as a doubling of staff across all departments, reinforcing its position in the AI‑as‑a‑service market. strategic moat. The combined focus on ultra‑low‑cost pricing, open licensing and easy local deployment underscores a broader Chinese strategy to challenge Western incumbents by making high‑quality LLMs accessible on modest hardware.