# Chinese AI firms push cost‑focused open‑source LLMs

> Live situation record from CLSTR: https://clstr.news/situations/consumerpc-deployment-of-opensource-llms
> Updated: 2026-08-02T14:52:59.000Z. Sources: 36. Developments: 11.

DeepSeek’s rapid rollout of cost‑efficient, open‑weight models continued through July 2026. After the July 1 launch of V4 Flash – a 284‑billion‑parameter Mixture‑of‑Experts model with a 1‑million‑token context window and per‑token pricing of 1 yuan in, 2 yuan out – the firm announced a full‑version V4 on July 21, claiming 40‑50× lower compute cost than leading U.S. systems and pricing up to 70 times cheaper than rivals. Alibaba followed with the open‑source Qwen 3.8 max, positioning it as a lower‑price, cloud‑integrated alternative amid fierce competition from other Chinese models such as Moonshot’s Kimi K3.

Parallel to these releases, a wave of technical guides made a variety of Chinese open‑source LLMs runnable on consumer hardware. DeepSeek‑V4 (7 B, 8K context) and Qwen 3.6‑27B were packaged in GGUF format for PCs with 12 GB VRAM; Qwen 3.5‑9B‑AWQ, Gemma‑4‑12B‑it, Gemma‑4‑26B‑A4B‑it, Qwen‑3.5‑2B, GLM‑5‑FP8, Qwen 3.5‑9B‑MLX‑8bit, and Qwen 3.5‑9B‑NVFP4 each received step‑by‑step installers, quantisation (AWQ‑INT4, FP8, NVFP4) and hardware‑auto‑tuning scripts targeting CPUs, GPUs and even Windows‑only deployments. A July 17 playbook showed how Alibaba’s 397‑billion‑parameter Qwen 3.5‑MoE could be accelerated on Google’s Ironwood TPU v7x, achieving 3‑fold decode throughput gains.

DeepSeek’s CEO Liang Wenfeng reiterated the compute gap with U.S. firms – roughly 20,000 H100‑class GPUs versus the 50,000‑plus needed for parity – but pledged continued expansion of domestic clusters and open‑source releases as a strategic moat. The combined focus on ultra‑low‑cost pricing, open licensing and easy local deployment underscores a broader Chinese strategy to challenge Western incumbents by making high‑quality LLMs accessible on modest hardware.

## Timeline

### 2026-08-02: DeepSeek launches V4-Flash, the cheapest AI model

DeepSeek’s V4-Flash, released 31 July, costs $0.14/$0.28 per M tokens (≈3¢ per test), over 100× cheaper than Anthropic’s Claude Fable 5, and scores 50/100, matching Google’s Gemini 3.6 Flash while trailing top‑

10 sources. https://clstr.news/cluster/deepseek-launches-v4-flash-ai-model-with-strong-performancetocost-ratio

### 2026-07-27: DeepSeek CEO Liang Wenfeng outlines AI strategy amid China-US competition

DeepSeek CEO Liang Wenfeng said the firm values AGI and open‑source over profit, sees compute power as the main AI race bottleneck with the U.S., and plans its own large‑scale clusters after a $52 bn valuation.

3 sources. https://clstr.news/cluster/deepseek-ceo-liang-wenfeng-outlines-ai-strategy-amid-china-us-competition

### 2026-07-23: DeepSeek CEO cites compute gap, pledges open‑source AI models

DeepSeek CEO says the lab lacks compute power versus U.S. rivals, needs tens of thousands of GPUs, but will expand capacity and keep top models open‑source while prioritising AGI over profit.

10 sources. https://clstr.news/cluster/palebluedot-ai-secures-255-million-credit-facility-to-expand-agentic-ai-infrastructure

### 2026-07-21: DeepSeek V4 and Alibaba Qwen 3.8 Launch Boost Chinese AI Race

DeepSeek V4 and Alibaba Qwen 3.8 debut, offering low‑cost, high‑performance AI models that intensify China’s competition with US giants and the Kimi K3.

4 sources. https://clstr.news/cluster/deepseek-v4-and-alibaba-qwen-38-launch-boost-chinese-ai-race

### 2026-07-18: Qwen3.5-9B NVFP4 Language Model Available for Windows Installation

The Qwen3.5‑9B NVFP4 language model, a 9 B‑parameter AI system, is now available for Windows, offering fast, low‑memory inference with required specs of i5/Ryzen 5 CPU, 32 GB RAM, 80 GB NVMe SSD, and CUDA 8.0+.

2 sources. https://clstr.news/cluster/qwen35-9b-nvfp4-language-model-available-for-windows-installation

### 2026-07-17: Qwen AI models receive new local deployment guide and Google TPU optimization playbook

A guide details local deployment of Qwen3‑VL‑Reranker‑8B, while Google publishes a playbook boosting Alibaba’s Qwen 3.5‑397B on Ironwood TPUs with up to 4.7× speed gains.

2 sources. https://clstr.news/cluster/qwen-ai-models-receive-new-local-deployment-guide-and-google-tpu-optimization-playbook

### 2026-07-16: Local Deployment Guides for GLM-5-FP8 and Qwen3.5‑9B‑MLX‑8bit AI Models

Guides detail offline deployment of GLM-5-FP8 (176 B, FP8) and Qwen3.5‑9B‑MLX‑8bit (9 B, 8‑bit) models on consumer hardware, outlining hardware needs and setup scripts.

2 sources. https://clstr.news/cluster/local-deployment-guides-for-glm-5-fp8-and-qwen359bmlx8bit-ai-models

### 2026-07-13: Gemma‑4‑26B and Qwen‑3.5‑2B Enable High‑Performance Local AI Deployment

Gemma‑4‑26B (4‑bit AWQ) and Qwen‑3.5‑2B (2 B parameters) can be installed locally with scripts that auto‑configure hardware, enabling fast, low‑resource AI deployment.

2 sources. https://clstr.news/cluster/gemma426b-and-qwen352b-enable-highperformance-local-ai-deployment

### 2026-07-11: New Qwen3.5‑9B‑AWQ and Gemma‑4‑12B‑it models enable fast local AI deployment

Qwen3.5‑9B‑AWQ and Gemma‑4‑12B‑it language models launch with offline deployment tools for fast, consumer‑grade AI inference and multilingual capabilities.

2 sources. https://clstr.news/cluster/new-qwen359bawq-and-gemma412bit-models-enable-fast-local-ai-deployment

### 2026-07-11: Deepseek‑V4 and Qwen3.6‑27B AI models now runnable on consumer PCs

Open‑source AI models Deepseek‑V4 (7 B) and Qwen3.6‑27B (27 B) can now be installed on typical PCs using GGUF packages, with step‑by‑step guides outlining required hardware and automatic setup.

4 sources. https://clstr.news/cluster/deepseekv4-and-qwen3627b-ai-models-now-runnable-on-consumer-pcs

### 2026-07-01: DeepSeek releases 284‑billion‑parameter V4 model for local laptop use

DeepSeek launched a 284B‑parameter V4 model that runs on high‑end laptops using MoE and quantization, and issued a V3 API guide offering OpenAI‑compatible, low‑cost access for developers.

2 sources. https://clstr.news/cluster/deepseek-releases-284billionparameter-v4-model-for-local-laptop-use

---
Cite as: Chinese AI firms push cost‑focused open‑source LLMs. CLSTR, https://clstr.news/situations/consumerpc-deployment-of-opensource-llms
