Google's Gemma 4 AI can run on smartphones, no Internet required

Key summary
Google released Gemma 4, an open-weight model family built on Gemini 3 and offered under Apache 2.0.
Gemma 4 has four variants; the largest are a 26B Mixture-of-Experts and a 31B Dense model (31B prioritizes raw quality and can be fine-tuned).
Those large models require an 80GB NVIDIA H100 to run unquantized in bfloat16; the 26B activates 3.8 billion parameters during inference.
Smaller Effective 2B and Effective 4B variants can run fully offline on phones and consumer devices like Raspberry Pi and Jetson Nano.
Gemma 4 is claimed to be faster and more capable than Gemma 3 for local hardware; Arena AI ranks the 31B at #3 and the 26B at #6. Google released only model parameters, not the full training pipeline, so it does not meet OSI open-source criteria.



