Google launches Gemma 4 with a broad licensing model

Key summary
Google DeepMind released Gemma 4, a family of four open-weight AI models offered under the Apache 2.0 license.
The models are built to run locally across devices—from small edge endpoints to workstations—with E2B and E4B running fully offline on phones, Raspberry Pi, and Nvidia Jetson Orin Nano with near-zero latency.
The lineup includes Effective 2B (E2B), Effective 4B (E4B), a 26B Mixture of Experts (MoE), and a 31B Dense model.
On the Arena AI text leaderboard the 31B Dense ranks third and the 26B MoE ranks sixth; Google says both beat models up to twenty times larger in parameter count.
Gemma was first donated to open-source in February 2024; the series has 400M+ downloads and over 100,000 community variants; models are available now via Google AI Studio, Kaggle, Ollama, and Hugging Face and include day-one compatibility with vLLM, llama.cpp, Ollama, NVIDIA NIM, and LM Studio.



