started · updated
AMD enables local execution of Qwen 3.8 27B AI model
AMD has announced Day 0 support for the Qwen 3.8 27B AI model, allowing developers and enthusiasts to run the 27-billion-parameter model locally on AMD hardware without relying on cloud infrastructure.
The model, part of Alibaba Cloud’s Qwen family, requires approximately 24GB of VRAM to operate effectively. Preliminary testing using the Vulkan backend via llama.cpp shows performance speeds of up to 24.5 tokens per second on AMD Ryzen AI Max+ 395 processors and up to 51.8 tokens per second on AMD Radeon AI PRO R9700 GPUs.
Users can access the model through a graphical interface via LM Studio for easier deployment. Additionally, developers can utilize AMD’s Lemonade platform, which serves as a hardware-aware inference layer to optimize load distribution across CPUs, GPUs, and NPUs.
Entities
AMD · Alibaba Cloud · LM Studio · Qwen