Google releases DiffusionGemma and bets that parallel text generation can upend the economics of local AI
Google DeepMind released DiffusionGemma on June 10, a 26-billion-parameter open model that generates text through diffusion rather than token-by-token prediction, reaching over 1,000 tokens per second on a single H100. The model ships under Apache 2.0 with immediate support in vLLM, Hugging Face Transformers, and Unsloth, and runs entirely on local hardware with no cloud dependency. Google flags it as experimental, but its day-one infrastructure support and NVIDIA hardware optimizations signal a