< Back to all clusters
[TECHNOLOGY] · 3 sources

started · updated

Google DeepMind advances Gemma 4 with DiffusionGemma and streamlined architecture

Google DeepMind has introduced technical advancements through its Gemma 4 model family, focusing on architectural efficiency and new diffusion capabilities.

A new model, DiffusionGemma, was developed by converting the existing Gemma-4-26B-A4B into a diffusion model using less than ten percent of the original training token budget. This model utilizes a two-stage training process involving reconstruction of noisy text blocks followed by a combined phase of Reinforcement Learning and Sampler Distillation (SD·RL). DiffusionGemma allows for bidirectional thinking, enabling the model to self-correct during the generation process, which has improved reasoning benchmarks by an average of ten points and significantly enhanced performance in tasks like Sudoku solving.

Additionally, the Gemma 4 12B model has demonstrated improved performance through aggressive architectural streamlining. By removing a dedicated 305-million-parameter audio encoder and replacing a 550-million-parameter vision encoder with a smaller 35-million-parameter matrix, the model achieved a 0.063 word error rate in English transcription. Other optimizations include a 37.5% reduction in the global KV cache and improved long-context retrieval capabilities, which increased from 8.6 to 79.5 at a 128k context window.

Entities

DiffusionGemma · Gemma 4 · Google · Google DeepMind