< Back to all clusters
[TECHNOLOGY] · France, United States, China · 2 sources

Qwen AI models receive new local deployment guide and Google TPU optimization playbook

A detailed walkthrough shows how to run the Qwen3‑VL‑Reranker‑8B vision‑language re‑ranking model on a local machine. The guide lists hardware recommendations—including a GPU with at least 16 GB video memory, 150 GB storage and sufficient RAM—highlights the model’s 8 billion parameters, cross‑modal attention, and low‑latency inference speed of about 200 tokens per second.

Google’s developer blog released an engineering playbook for accelerating Alibaba’s open‑weight Qwen 3.5‑397B mixture‑of‑experts model on its Ironwood TPU v7x. The optimization achieved roughly 3.1× higher decode‑heavy throughput and 4.7× higher pre‑fill throughput, reaching about 82 % of the TPU’s theoretical roofline. The work incorporates hybrid sharding, custom JAX/Pallas kernels, SparseCore MoE routing and FP8 precision, and integrates with the open‑source serving frameworks vLLM and SGLang.