The Google Research team introduced the experimental DiffusionGemma model based on discrete diffusion
DiffusionGemma, a model created by fine-tuning Gemma 4, generates text through discrete diffusion in parallel blocks of 256 tokens and, according to the authors, achieves about 1 500 tokens per second on a single NVIDIA H100 GPU.
The Google Research team published a technical report on DiffusionGemma, an experimental language model with open weights that uses discrete diffusion instead of sequential token-by-token generation. The model iteratively refines blocks of 256 tokens in parallel, which, according to the authors, avoids the sequential limitation of conventional autoregressive models.
DiffusionGemma was created by fine-tuning Gemma 4, a model with a mixture-of-experts architecture that has 3.8 billion active and 25.2 billion total parameters. Training took place in two stages: first, supervised fine-tuning taught the model bidirectional denoising, then a combination of reinforcement learning and sampler distillation improved both generation quality and inference efficiency. According to the authors, the entire process used less than 10 % of the training token budget of the original autoregressive model.
The model generates around 20 tokens per forward pass on average and achieves approximately 1 500 output tokens per second on a single NVIDIA H100 GPU - according to the authors, substantially faster than autoregressive models even when using state-of-the-art speculative decoding. DiffusionGemma retains support for thinking mode, multimodal inputs and long contexts from the base model. Although the model was fine-tuned for diffusion, it remains capable of autoregressive generation with only a small drop in performance, which the authors describe as a possible path to hybrid diffusion-AR decoding.
Why it matters
For developers working with large language models, the work demonstrates an alternative decoding architecture that could significantly shorten text generation response times without having to train a model from scratch - fine-tuning an existing model is sufficient. For companies running LLM inference at scale, faster generation (up to approximately 1 500 tokens per second on a single GPU) could mean lower compute costs and shorter latency if the approach proves suitable for production use.
Two audiences, two different impacts
What this means
For individuals
Developers and researchers experimenting with open models gain evidence that diffusion decoding can offer significantly faster text generation than conventional autoregressive models through fine-tuning an existing model, rather than training from scratch.
For a business
Companies running LLM inference at high volumes may infer from the results that diffusion decoding could reduce compute costs and generation latency if this approach becomes more widely adopted in production deployments.
DevelopmentCheck the original
Event sources
only one source so far · 1 publisher, 1 independent. We count feeds from the same owner only once.