Lesson 3: The GGUF Alphabet: K-Quants Decoded
Lesson 3: The GGUF Alphabet: K-Quants Decoded
GGUF is the container format of the llama.cpp universe: Ollama, LM Studio, koboldcpp, Jan, GPT4All, and every llama.cpp fork. It bundles weights, tokenizer, and metadata in one file โ that's the .gguf extension. Within GGUF, the quant name tells you everything. Here is the full ladder, with real file sizes from the unsloth Qwen3.8-27B repo:
| Quant name | ~bpw | Qwen3.8-27B size | What it is | Use when |
|---|---|---|---|---|
Q2_K | 2.9 | ~9.8 GB (UD-Q2_K_XL) | 2-bit K-quant | Desperate fit |
Q3_K_S / M / XL | 3.4-3.9 | ~13.2 GB (XL) | 3-bit K-quants | Fitting on 16 GB, quality drops |
Q4_0 (legacy) | 4.6 | 16.06 GB | Old plain 4-bit | Avoid if K-quants exist |
Q4_K_S | 4.6 | 15.36 GB | 4-bit, small tier | Tight 16 GB fit |
Q4_K_M | 4.9 | 16.46 GB | 4-bit, medium โ the default | Best all-round choice |
Q4_K_XL | 5.2 | 17.56 GB | 4-bit, large tier | 4-bit with extra attention precision |
Q5_K_S / M / XL | 5.5-6.2 | 18.7-20.9 GB | 5-bit K-quants | 24 GB GPUs, code/math work |
Q6_K | 6.6 | 21.98 GB | 6-bit, no tiers originally | 32 GB GPUs, near-lossless feel |
Q6_K_M / L / XL | 7-7.6 | 23.1-25.3 GB | New 6-bit tiers (2026) | Big-VRAM quality seekers |
Q8_0 | 8.5 | 29.05 GB | 8-bit integer | Essentially lossless; 32 GB+ |
Q8_K_L / XL | 8.5-9.5 | 28.1-31.5 GB | New 8-bit K-quants | Same, slightly better |
Q4_K_M = Q4 (4-bit level) + K (k-means mixed-precision scheme) + M (medium size tier). Q5_K_S = 5-bit, K-scheme, small tier. Q8_0 has no K โ it's the older-but-still-great plain 8-bit.
Which GGUF should you actually download?
- Default:
Q4_K_M. Best quality-per-gigabyte; the community consensus pick. - Headroom on a bigger GPU:
Q5_K_MโQ6_KโQ8_0as VRAM allows. Code and math reward the extra bits more than chat does. - 24 GB card with a 27B: Q4_K_M fits comfortably; Q5_K_M fits with a shorter context.
- 16 GB card with a 27B: Q4_K_S or Q3_K_XL or an IQ quant (next lesson).
- Never: legacy
Q4_0/Q4_1/Q5_0/Q5_1when a K-quant of the same bit level exists โ K-quants are strictly better at similar size. (Q4_0 survives only because it's trivial to compute and some tiny tools still use it.)
UD- prefix: Unsloth Dynamic
Repos by unsloth prefix their files with UD- (e.g. Qwen3.8-27B-UD-Q4_K_M.gguf). UD (Unsloth Dynamic 3.0) re-quantizes additional tensors to higher precision and is the current leader in accuracy-per-byte for GGUF โ unsloth claims more than 10% better accuracy at the same size versus plain quants. It's a quant method, not a new format: same .gguf, same runtimes, just better calibration. When a repo offers both plain and UD files at the same bpw, take the UD one.
UD-Q4_K_M (16.5 GB). 32 GB โ UD-Q5_K_M or UD-Q6_K. 48 GB โ Q8_0 or UD-Q8_K_L. 16 GB โ UD-Q4_K_S or the IQ quants in the next lesson.
🧠 Knowledge Check
1. What does the 'M' in Q4_K_M stand for?
2. Qwen3.8-27B at Q4_K_M is roughly:
3. Why prefer K-quants over legacy Q4_0 at the same bit level?