Image
Gemma 4 up to 3x faster: Google introduces multi-token prediction for its open-source models
New speculative decoding technique lets Gemma 4 models generate text up to three times faster without any loss of accuracy. Thanks to the Apache 2.0 open-source license, developers can try it out instantly in Ollama, vLLM, or directly on mobile.