< Back to all clusters
[TECHNOLOGY] · United States · 2 sources

Apple evaluates PrismML's 27B‑parameter AI compression for iPhone

PrismML, a Caltech spin‑out backed by Khosla Ventures, says it has compressed Alibaba’s 27‑billion‑parameter Qwen 3.6 language model from roughly 54 GB to under 4 GB and run it on an iPhone 17 Pro with all parameters active. The startup emerged from stealth in March 2026 with a $16.25 million seed round and claims its Bonsai compression technique reduces memory footprint up to 14× and speeds inference up to 8× while keeping every parameter intact.

Apple has reportedly been in talks with PrismML about using the technology to extend its on‑device AI capabilities beyond the current AFM 3 Core Advanced model, which tops out at 20 billion parameters and relies on a sparse architecture. Running a full‑capacity large language model locally would improve privacy, eliminate the need for cloud‑based inference, and lower energy consumption, though the performance claims have not yet been independently verified.