Open-weight local model
Mistral Small 3.1 24B Instruct
High-quality local multimodal model. Text estimates exclude vision overhead; 24GB is the practical Q4 target.
- Parameters
- 24B
- Recommended quantization
- Q4_K_M
- Weight estimate
- 15GB
- Maximum context
- 128K
Quantization memory at the default context
| Quantization | Weights | KV cache | Total VRAM |
|---|---|---|---|
| FP16 | 48.50GB | 1.34GB | 54.69GB |
| Q8_0 | 26.50GB | 1.34GB | 30.49GB |
| Q6_K | 20.48GB | 1.34GB | 23.86GB |
| Q5_K_M | 17.32GB | 1.34GB | 20.40GB |
| Q4_K_M | 15.00GB | 1.34GB | 17.84GB |
| Q3_K_M | 12.40GB | 1.34GB | 14.98GB |
| Q2_K | 7.88GB | 1.34GB | 10.02GB |