Skip to content
context New models

The Google Research team introduced the experimental DiffusionGemma model based on discrete diffusion

only one source so far

DiffusionGemma, a model created by fine-tuning Gemma 4, generates text through discrete diffusion in parallel blocks of 256 tokens and, according to the authors, achieves about 1 500 tokens per second on a single NVIDIA H100 GPU.

The Google Research team published a technical report on DiffusionGemma, an experimental language model with open weights that uses discrete diffusion instead of sequential token-by-token generation. The model iteratively refines blocks of 256 tokens in parallel, which, according to the authors, avoids the sequential limitation of conventional autoregressive models.

DiffusionGemma was created by fine-tuning Gemma 4, a model with a mixture-of-experts architecture that has 3.8 billion active and 25.2 billion total parameters. Training took place in two stages: first, supervised fine-tuning taught the model bidirectional denoising, then a combination of reinforcement learning and sampler distillation improved both generation quality and inference efficiency. According to the authors, the entire process used less than 10 % of the training token budget of the original autoregressive model.

The model generates around 20 tokens per forward pass on average and achieves approximately 1 500 output tokens per second on a single NVIDIA H100 GPU - according to the authors, substantially faster than autoregressive models even when using state-of-the-art speculative decoding. DiffusionGemma retains support for thinking mode, multimodal inputs and long contexts from the base model. Although the model was fine-tuned for diffusion, it remains capable of autoregressive generation with only a small drop in performance, which the authors describe as a possible path to hybrid diffusion-AR decoding.

What changed

Why it matters

For developers working with large language models, the work demonstrates an alternative decoding architecture that could significantly shorten text generation response times without having to train a model from scratch - fine-tuning an existing model is sufficient. For companies running LLM inference at scale, faster generation (up to approximately 1 500 tokens per second on a single GPU) could mean lower compute costs and shorter latency if the approach proves suitable for production use.

Two audiences, two different impacts

What this means

01

For individuals

Developers and researchers experimenting with open models gain evidence that diffusion decoding can offer significantly faster text generation than conventional autoregressive models through fine-tuning an existing model, rather than training from scratch.

What to do Monitor whether and when the weights for DiffusionGemma are released so you can test generation speed yourself.
More practical updates →
02

For a business

Companies running LLM inference at high volumes may infer from the results that diffusion decoding could reduce compute costs and generation latency if this approach becomes more widely adopted in production deployments.

Development
What to decide Monitor the development of diffusion-based inference as a potential way to reduce costs and latency when deploying LLMs at scale.
More business impacts →
DiffusionGemma discrete diffusion generation efficiency Gemma inference open-weight

Check the original

Event sources

only one source so far · 1 publisher, 1 independent. We count feeds from the same owner only once.

1
arXiv cs.CL (Computation and Language / NLP) research source · first detected DiffusionGemma Technical Report