Open-weight local model
DeepSeek R1 Distill Qwen 1.5B
Small reasoning checkpoint; long reasoning responses can reduce perceived chat speed.
- Parameters
- 1.777B
- Recommended quantization
- Q4_K_M
- Weight estimate
- 1.3GB
- Maximum context
- 128K
Quantization memory at the default context
| Quantization | Weights | KV cache | Total VRAM |
|---|---|---|---|
| FP16 | 3.70GB | 0.23GB | 4.73GB |
| Q8_0 | 2.10GB | 0.23GB | 3.13GB |
| Q6_K | 1.52GB | 0.23GB | 2.55GB |
| Q5_K_M | 1.28GB | 0.23GB | 2.32GB |
| Q4_K_M | 1.30GB | 0.23GB | 2.33GB |
| Q3_K_M | 1.10GB | 0.23GB | 2.13GB |