CanRunAI

Open-weight local model

Qwen 3 32B

Large dense reasoning model; Q4 is tight on 24GB once KV cache and runtime overhead are included.

Parameters
32.762B
Recommended quantization
Q4_K_M
Weight estimate
20.4GB
Maximum context
40K

Official model card

Quantization memory at the default context

QuantizationWeightsKV cacheTotal VRAM
FP1666.00GB2.15GB74.75GB
Q8_036.00GB2.15GB41.75GB
Q6_K27.95GB2.15GB32.89GB
Q5_K_M23.65GB2.15GB28.16GB
Q4_K_M20.40GB2.15GB24.59GB
Q3_K_M16.90GB2.15GB20.74GB
Q2_K10.75GB2.15GB13.97GB

GPUs to check for this model