About the model
Gemma 4 is Google's latest open-source model family. The e4b (efficient 4-bit) variant is optimized for fast inference on a 16 GB VRAM consumer GPU. Excellent performance-to-size ratio.
Capabilities
✅ Text 💻 Code
Technical specifications
| Type | ollama |
|---|---|
| Parameters | 26B |
| Quantization | Q4_K_M |
| Context window | 32768 |
Test hardware
| CPU | AMD Ryzen |
|---|---|
| GPU | NVIDIA RTX 5060 Ti 16GB |
| RAM | 32 GB DDR5 |
| OS | Ubuntu 24.04 LTS |
Test results
| Test | Run | Tokens/s | TTFT (ms) | Duration (s) | Tokens | GPU VRAM | Temp | Quality | Date |
|---|---|---|---|---|---|---|---|---|---|
| ? | #1 | 83.99 | 75 | 188.3 | 3380 | 4646 MB | 69 °C | - | 25.07.2026 |
| ? | #1 | 84.75 | 78 | 241.9 | 1650 | 4646 MB | 64 °C | - | 25.07.2026 |
| ? | #1 | 84.43 | 148 | 99.4 | 2425 | 4646 MB | 67 °C | - | 25.07.2026 |
| ? | #1 | 84.28 | 240 | 119.9 | 3546 | 4646 MB | 68 °C | - | 25.07.2026 |