CanRunAI

Open-weight local model

Google Gemma 3 27B IT

Near the limit of a 24GB card at 4-bit. Text estimates exclude vision overhead.

Parameters
27B
Recommended quantization
Q4_K_M
Weight estimate
17.2GB
Maximum context
128K

Official model card

Quantization memory at the default context

QuantizationWeightsKV cacheTotal VRAM
FP1654.50GB4.16GB64.11GB
Q8_029.80GB4.16GB36.94GB
Q6_K23.03GB4.16GB29.50GB
Q5_K_M19.49GB4.16GB25.60GB
Q4_K_M17.20GB4.16GB23.08GB
Q3_K_M14.30GB4.16GB19.89GB
Q2_K8.86GB4.16GB13.91GB

GPUs to check for this model