Apple Evaluates PrismML's AI Model Compression for On‑Device iPhone Use
PrismML, a Caltech spin‑out, has demonstrated a technique that compresses large language models by more than 90 % using extreme 1‑bit and ternary quantization. The startup reduced Alibaba’s open‑source Qwen 3.6 27‑billion‑parameter model from roughly 54 GB to under 4 GB, allowing it to run on an iPhone 17 Pro at about 11–13 tokens per second.
Apple is in early‑stage talks with PrismML and is evaluating the technology’s speed, energy efficiency and on‑device performance. CEO Babak Hassibi told CNBC, “They’re really evaluating our technology right now.” Analysts note the breakthrough could let Apple keep more AI processing on devices, potentially speeding up Siri and improving user privacy by reducing reliance on cloud servers.
PrismML has released the compressed models under an Apache 2.0 license with custom kernels for Apple’s Metal framework. The company claims the 1‑bit version uses 10–15× less memory, runs 6–8× faster and consumes 3–6× less energy than the full‑precision model, while retaining roughly 90‑95 % of benchmark scores.