CanRunAI

Open-weight local model

DeepSeek R1 Distill Qwen 1.5B

Small reasoning checkpoint; long reasoning responses can reduce perceived chat speed.

Parameters
1.777B
Recommended quantization
Q4_K_M
Weight estimate
1.3GB
Maximum context
128K

Official model card

Quantization memory at the default context

QuantizationWeightsKV cacheTotal VRAM
FP163.70GB0.23GB4.73GB
Q8_02.10GB0.23GB3.13GB
Q6_K1.52GB0.23GB2.55GB
Q5_K_M1.28GB0.23GB2.32GB
Q4_K_M1.30GB0.23GB2.33GB
Q3_K_M1.10GB0.23GB2.13GB

GPUs to check for this model