Google launches DiffusionGemma with 4x faster text generation
Google launched DiffusionGemma, an experimental open model that uses text diffusion to generate blocks of text in parallel, claiming up to 4x faster text generation than traditional token-by-token language models. The 26B-parameter Mixture-of-Experts model activates just 3.8B parameters during inference and can run within 18GB VRAM when quantized, making it relevant for both consumer and enterprise GPU workflows. Google says the model exceeds 1,000 tokens per second on NVIDIA H100 hardware and more than 700 tokens per second on RTX 5090 systems. The release is aimed at researchers and developers working on speed-critical local applications, though Google notes output quality is lower than its standard Gemma 4 models. The article is mildly positive for Google’s AI positioning and underscores NVIDIA’s importance as the hardware platform enabling these gains.