CanRunAI

Open-weight local model

DeepSeek R1 8B

Llama-based reasoning distill; all displayed speed figures are calculated rather than benchmarked.

Parameters
8.03B
Recommended quantization
Q4_K_M
Weight estimate
5.4GB
Maximum context
128K

Official model card

Quantization memory at the default context

QuantizationWeightsKV cacheTotal VRAM
FP1616.40GB1.07GB19.11GB
Q8_09.10GB1.07GB11.08GB
Q6_K6.85GB1.07GB8.72GB
Q5_K_M5.80GB1.07GB7.67GB
Q4_K_M5.40GB1.07GB7.27GB
Q3_K_M4.50GB1.07GB6.37GB
Q2_K2.63GB1.07GB4.51GB

GPUs to check for this model