< Back to all clusters
[TECHNOLOGY] · Brazil, Israel · 2 sources

New Qwen3.5‑9B‑AWQ and Gemma‑4‑12B‑it models enable fast local AI deployment

Two large language models have been released with detailed guides for offline, low‑resource deployment. The Qwen3.5‑9B‑AWQ model, a 9‑billion‑parameter system using activation‑aware quantisation, runs efficiently on consumer‑grade CPUs and GPUs, supports an 8K‑token context window and excels at code generation, dialogue and multilingual factual Q&A.

The Gemma‑4‑12B‑it model, with 12 billion parameters, offers rapid inference, a 2 K‑token context length and strong multilingual performance, achieving 85 % accuracy on reading‑comprehension benchmarks and a 78 % pass rate on code‑generation tests. Both models provide step‑by‑step installers, hardware‑auto‑tuning and no‑code deployment options for Windows PCs and Linux systems, targeting developers who need high‑accuracy AI without specialised server hardware.