< Back to situations

We’ll email you as it develops, and you can follow the whole thread from day one.

[SITUATION] · [ACTIVE]

11 clusters · 36 sources · 33 days · First seen · Last updated

Categories: TECHNOLOGY

Chinese AI firms push cost‑focused open‑source LLMs

Entities: DeepSeek · OpenAI · China · V4-Flash · DeepSeek V4 Flash

Overview

DeepSeek’s rapid rollout of cost‑efficient, open‑weight models continued through July 2026. After the July 1 launch of V4 Flash – a 284‑billion‑parameter Mixture‑of‑Experts model with a 1‑million‑token context window and per‑token pricing of 1 yuan in, 2 yuan out – the firm announced a full‑version V4 on July 21, claiming 40‑50× lower compute cost than leading U.S. systems and pricing up to 70 times cheaper than rivals. Alibaba followed with the open‑source Qwen 3.8 max, positioning it as a lower‑price, cloud‑integrated alternative amid fierce competition from other Chinese models such as Moonshot’s Kimi K3.

Parallel to these releases, a wave of technical guides made a variety of Chinese open‑source LLMs runnable on consumer hardware. DeepSeek‑V4 (7 B, 8K context) and Qwen 3.6‑27B were packaged in GGUF format for PCs with 12 GB VRAM; Qwen 3.5‑9B‑AWQ, Gemma‑4‑12B‑it, Gemma‑4‑26B‑A4B‑it, Qwen‑3.5‑2B, GLM‑5‑FP8, Qwen 3.5‑9B‑MLX‑8bit, and Qwen 3.5‑9B‑NVFP4 each received step‑by‑step installers, quantisation (AWQ‑INT4, FP8, NVFP4) and hardware‑auto‑tuning scripts targeting CPUs, GPUs and even Windows‑only deployments. A July 17 playbook showed how Alibaba’s 397‑billion‑parameter Qwen 3.5‑MoE could be accelerated on Google’s Ironwood TPU v7x, achieving 3‑fold decode throughput gains.

DeepSeek’s CEO Liang Wenfeng reiterated the compute gap with U.S. firms – roughly 20,000 H100‑class GPUs versus the 50,000‑plus needed for parity – but pledged continued expansion of domestic clusters and open‑source releases as a strategic moat. The combined focus on ultra‑low‑cost pricing, open licensing and easy local deployment underscores a broader Chinese strategy to challenge Western incumbents by making high‑quality LLMs accessible on modest hardware.

Timeline

  1. 2 days ago

    [TECHNOLOGY] 10 sources
    DeepSeek launches V4-Flash, the cheapest AI model

    DeepSeek’s V4-Flash, released 31 July, costs $0.14/$0.28 per M tokens (≈3¢ per test), over 100× cheaper than Anthropic’s Claude Fable 5, and scores 50/100, matching Google’s Gemini 3.6 Flash while trailing top‑

  2. 8 days ago

    [TECHNOLOGY] 3 sources
    DeepSeek CEO Liang Wenfeng outlines AI strategy amid China-US competition

    DeepSeek CEO Liang Wenfeng said the firm values AGI and open‑source over profit, sees compute power as the main AI race bottleneck with the U.S., and plans its own large‑scale clusters after a $52 bn valuation.

  3. 12 days ago

    [TECHNOLOGY] 10 sources
    DeepSeek CEO cites compute gap, pledges open‑source AI models

    DeepSeek CEO says the lab lacks compute power versus U.S. rivals, needs tens of thousands of GPUs, but will expand capacity and keep top models open‑source while prioritising AGI over profit.

  4. 14 days ago

    [TECHNOLOGY] 4 sources
    DeepSeek V4 and Alibaba Qwen 3.8 Launch Boost Chinese AI Race

    DeepSeek V4 and Alibaba Qwen 3.8 debut, offering low‑cost, high‑performance AI models that intensify China’s competition with US giants and the Kimi K3.

  5. 17 days ago

    [TECHNOLOGY] 2 sources
    Qwen3.5-9B NVFP4 Language Model Available for Windows Installation

    The Qwen3.5‑9B NVFP4 language model, a 9 B‑parameter AI system, is now available for Windows, offering fast, low‑memory inference with required specs of i5/Ryzen 5 CPU, 32 GB RAM, 80 GB NVMe SSD, and CUDA 8.0+.

  6. 18 days ago

    [TECHNOLOGY] 2 sources
    Qwen AI models receive new local deployment guide and Google TPU optimization playbook

    A guide details local deployment of Qwen3‑VL‑Reranker‑8B, while Google publishes a playbook boosting Alibaba’s Qwen 3.5‑397B on Ironwood TPUs with up to 4.7× speed gains.

  7. 19 days ago

    [TECHNOLOGY] 2 sources
    Local Deployment Guides for GLM-5-FP8 and Qwen3.5‑9B‑MLX‑8bit AI Models

    Guides detail offline deployment of GLM-5-FP8 (176 B, FP8) and Qwen3.5‑9B‑MLX‑8bit (9 B, 8‑bit) models on consumer hardware, outlining hardware needs and setup scripts.

  8. 22 days ago

    [TECHNOLOGY] 2 sources
    Gemma‑4‑26B and Qwen‑3.5‑2B Enable High‑Performance Local AI Deployment

    Gemma‑4‑26B (4‑bit AWQ) and Qwen‑3.5‑2B (2 B parameters) can be installed locally with scripts that auto‑configure hardware, enabling fast, low‑resource AI deployment.

  9. 24 days ago

    [TECHNOLOGY] 2 sources
    New Qwen3.5‑9B‑AWQ and Gemma‑4‑12B‑it models enable fast local AI deployment

    Qwen3.5‑9B‑AWQ and Gemma‑4‑12B‑it language models launch with offline deployment tools for fast, consumer‑grade AI inference and multilingual capabilities.

  10. 24 days ago

    [TECHNOLOGY] 4 sources
    Deepseek‑V4 and Qwen3.6‑27B AI models now runnable on consumer PCs

    Open‑source AI models Deepseek‑V4 (7 B) and Qwen3.6‑27B (27 B) can now be installed on typical PCs using GGUF packages, with step‑by‑step guides outlining required hardware and automatic setup.

  11. about 1 month ago

    [TECHNOLOGY] 2 sources
    DeepSeek releases 284‑billion‑parameter V4 model for local laptop use

    DeepSeek launched a 284B‑parameter V4 model that runs on high‑end laptops using MoE and quantization, and issued a V3 API guide offering OpenAI‑compatible, low‑cost access for developers.

Sources

36kr.com · arquitectosdevalencia.es · bolnews.com · borncity.com · deccanchronicle.com · dev.to · fourweekmba.com · gamemag.it · games.yahoo.com.tw · geeky-gadgets.com · huguettetiegna.fr · igeekphone.com · independent.co.uk · infa.lt · isd.sorbonneonu.fr · isds.co.il · it-boltwise.de · ithome.com · jcnc.org · kvia.com · m.olhardigital.uol.com.br · mathrubhumi.com · memeburn.com · mononews.gr · nuevospapeles.com · radiolideranca103.com.br · revistaplus.com.py · sapo.pt · sitepoint.com · sspai.com · techjuice.pk · thisissoundcheck.co.uk · tmtpost.com · unwire.hk · wccftech.com · wtvbam.com

This summary has been updated 2 times: see revision history