Open-weight local model
Qwen 2.5 Coder 7B Instruct
Popular code-specialized model that is practical on an 8GB GPU at 4-bit with modest context.
- Parameters
- 7.616B
- Recommended quantization
- Q4_K_M
- Weight estimate
- 5.1GB
- Maximum context
- 32K
Quantization memory at the default context
| Quantization | Weights | KV cache | Total VRAM |
|---|---|---|---|
| FP16 | 15.60GB | 0.47GB | 17.63GB |
| Q8_0 | 8.70GB | 0.47GB | 10.04GB |
| Q6_K | 6.50GB | 0.47GB | 7.77GB |
| Q5_K_M | 5.50GB | 0.47GB | 6.77GB |
| Q4_K_M | 5.10GB | 0.47GB | 6.37GB |
| Q3_K_M | 4.30GB | 0.47GB | 5.57GB |
GPUs to check for this model
- Can RTX 4060 8GB run it?
- Can Nvidia RTX A5000-8Q 8GB run it?
- Can RTX 3060 Ti GDDR6X 8GB run it?
- Can RTX 3070 Ti 8GB run it?
- Can RTX 3070 Ti 8 GB GA102 run it?
- Can Intel Arc A580 8GB run it?
- Can Intel Arc A750 8GB run it?
- Can Nvidia Quadro RTX 4000 Max-Q 8GB run it?
- Can Nvidia Quadro RTX 4000 Mobile 8GB run it?