Skip to main content

gemma4:e4b

Google

https://ai.google.dev/gemma

About the model

Gemma 4 is Google's latest open-source model family. The e4b (efficient 4-bit) variant is optimized for fast inference on a 16 GB VRAM consumer GPU. Excellent performance-to-size ratio.

Capabilities

✅ Text 💻 Code

Technical specifications

Type ollama
Parameters 26B
Quantization Q4_K_M
Context window 32768

Test hardware

CPUAMD Ryzen
GPUNVIDIA RTX 5060 Ti 16GB
RAM32 GB DDR5
OSUbuntu 24.04 LTS

Test results

Test Run Tokens/s TTFT (ms) Duration (s) Tokens GPU VRAM Temp Quality Date
? #1 83.99 75 188.3 3380 4646 MB 69 °C - 25.07.2026
? #1 84.75 78 241.9 1650 4646 MB 64 °C - 25.07.2026
? #1 84.43 148 99.4 2425 4646 MB 67 °C - 25.07.2026
? #1 84.28 240 119.9 3546 4646 MB 68 °C - 25.07.2026
X

Don't miss out!

Subscribe for the latest news and updates.