Lesson 3: The GGUF Alphabet: K-Quants Decoded

Lesson 3: The GGUF Alphabet: K-Quants Decoded

GGUF is the container format of the llama.cpp universe: Ollama, LM Studio, koboldcpp, Jan, GPT4All, and every llama.cpp fork. It bundles weights, tokenizer, and metadata in one file โ€” that's the .gguf extension. Within GGUF, the quant name tells you everything. Here is the full ladder, with real file sizes from the unsloth Qwen3.8-27B repo:

Quant name~bpwQwen3.8-27B sizeWhat it isUse when
Q2_K2.9~9.8 GB (UD-Q2_K_XL)2-bit K-quantDesperate fit
Q3_K_S / M / XL3.4-3.9~13.2 GB (XL)3-bit K-quantsFitting on 16 GB, quality drops
Q4_0 (legacy)4.616.06 GBOld plain 4-bitAvoid if K-quants exist
Q4_K_S4.615.36 GB4-bit, small tierTight 16 GB fit
Q4_K_M4.916.46 GB4-bit, medium โ€” the defaultBest all-round choice
Q4_K_XL5.217.56 GB4-bit, large tier4-bit with extra attention precision
Q5_K_S / M / XL5.5-6.218.7-20.9 GB5-bit K-quants24 GB GPUs, code/math work
Q6_K6.621.98 GB6-bit, no tiers originally32 GB GPUs, near-lossless feel
Q6_K_M / L / XL7-7.623.1-25.3 GBNew 6-bit tiers (2026)Big-VRAM quality seekers
Q8_08.529.05 GB8-bit integerEssentially lossless; 32 GB+
Q8_K_L / XL8.5-9.528.1-31.5 GBNew 8-bit K-quantsSame, slightly better
Reading the name: Q4_K_M = Q4 (4-bit level) + K (k-means mixed-precision scheme) + M (medium size tier). Q5_K_S = 5-bit, K-scheme, small tier. Q8_0 has no K โ€” it's the older-but-still-great plain 8-bit.

Which GGUF should you actually download?

  • Default: Q4_K_M. Best quality-per-gigabyte; the community consensus pick.
  • Headroom on a bigger GPU: Q5_K_M โ†’ Q6_K โ†’ Q8_0 as VRAM allows. Code and math reward the extra bits more than chat does.
  • 24 GB card with a 27B: Q4_K_M fits comfortably; Q5_K_M fits with a shorter context.
  • 16 GB card with a 27B: Q4_K_S or Q3_K_XL or an IQ quant (next lesson).
  • Never: legacy Q4_0/Q4_1/Q5_0/Q5_1 when a K-quant of the same bit level exists โ€” K-quants are strictly better at similar size. (Q4_0 survives only because it's trivial to compute and some tiny tools still use it.)

UD- prefix: Unsloth Dynamic

Repos by unsloth prefix their files with UD- (e.g. Qwen3.8-27B-UD-Q4_K_M.gguf). UD (Unsloth Dynamic 3.0) re-quantizes additional tensors to higher precision and is the current leader in accuracy-per-byte for GGUF โ€” unsloth claims more than 10% better accuracy at the same size versus plain quants. It's a quant method, not a new format: same .gguf, same runtimes, just better calibration. When a repo offers both plain and UD files at the same bpw, take the UD one.

Practical pick for Qwen3.8-27B: 24 GB GPU โ†’ UD-Q4_K_M (16.5 GB). 32 GB โ†’ UD-Q5_K_M or UD-Q6_K. 48 GB โ†’ Q8_0 or UD-Q8_K_L. 16 GB โ†’ UD-Q4_K_S or the IQ quants in the next lesson.

🧠 Knowledge Check

1. What does the 'M' in Q4_K_M stand for?

2. Qwen3.8-27B at Q4_K_M is roughly:

3. Why prefer K-quants over legacy Q4_0 at the same bit level?

Further Reading