Open-weight local model
Google Gemma 3 12B IT
Balanced multimodal model. Text estimates exclude image encoder and vision activation peaks.
- Parameters
- 12B
- Recommended quantization
- Q4_K_M
- Weight estimate
- 8.2GB
- Maximum context
- 128K
Quantization memory at the default context
| Quantization | Weights | KV cache | Total VRAM |
|---|---|---|---|
| FP16 | 24.50GB | 1.61GB | 28.56GB |
| Q8_0 | 13.50GB | 1.61GB | 16.46GB |
| Q6_K | 10.24GB | 1.61GB | 12.87GB |
| Q5_K_M | 8.66GB | 1.61GB | 11.14GB |
| Q4_K_M | 8.20GB | 1.61GB | 10.63GB |
| Q3_K_M | 6.80GB | 1.61GB | 9.21GB |
GPUs to check for this model
- Can RTX 4070 Ti SUPER 16GB run it?
- Can Apple MacBook Pro M3 Pro (14-core GPU) 18GB run it?
- Can Apple MacBook Pro M3 Pro (18-core GPU) 18GB run it?
- Can RTX 5080 16GB run it?
- Can RTX 5070 Ti 16GB run it?
- Can RTX 5070 Ti SUPER 16GB run it?
- Can RTX 5080 Mobile 16GB run it?
- Can RTX 6070 16GB run it?
- Can RTX 4080 SUPER 16GB run it?