How to Run an 80B Qwen Model in 4.3GB of RAM: The Edge AI Revolution Explained

TL;DR AI
2 min readKey summary
A viral Hacker News report says an 80B Qwen model can run locally in just 4.3GB of RAM at about 4 tokens per second.
The same coverage claims a 35B model also runs on a base iPhone 18 Pro, highlighting fast progress in on-device LLM inference.
The gains are attributed to aggressive compression methods like mixed-precision quantization, pruning, codebook weight sharing, and low-rank factorization.
Apple Silicon’s unified memory and sparse-operation hardware are described as key enablers for running large models on consumer devices.
If accurate, the trend could make laptops and phones much more practical for private, low-latency AI without cloud GPUs.
