Lesson 4: IQ Quants, imatrix, and Squeezing Below 4 Bits
Lesson 4: IQ Quants, imatrix, and Squeezing Below 4 Bits
Sometimes the model you want simply doesn't fit at 4 bits. Qwen3.8-27B at Q4_K_M is 16.5 GB โ too big for a 16 GB card once context is added. That's what the IQ family is for: importance-matrix-aware quants that wring more quality out of every bit, at the cost of a little extra dequantization work at inference (invisible on GPU, slightly slower on CPU).
| Quant | ~bpw | Qwen3.8-27B size | Fits in | Character |
|---|---|---|---|---|
IQ1_S / IQ1_M | 1.6-1.8 | 6.2 / 6.7 GB | 8 GB | Novelty โ heavy quality loss |
IQ2_XXS / IQ2_S | 2.1-2.5 | 7.3 / 8.4 GB | 8-12 GB | Usable if the model is big enough |
IQ3_XXS / IQ3_S | 3.1-3.4 | 10.9 / 12.0 GB | 12-16 GB | The 16 GB sweet spot for 27B |
IQ4_XS / IQ4_NL | 4.2-4.5 | 14.3 GB | 16 GB | Near-Q4_K_S quality, smaller |
The XXS / XS / S suffixes are size tiers within each bit level (XXS = extra-extra-small). IQ quants are the reason a 27B can run on a 16 GB card โ at IQ3_XXS (10.9 GB) you keep several GB free for context.
imatrix: the secret sauce
imatrix (importance matrix) is calibration data produced by running text through the fp16 model before quantizing. Quantizing with an imatrix measurably improves quality โ 10-20% better perplexity at Q4, and it's basically mandatory for Q3 and below. It shows up in filenames two ways:
-imatrixsuffix:Qwen3.8-27B-Q4_K_M-imatrix.gguf- Some repos ship the matrix as a separate
imatrix_unsloth.gguf-style file and bake it into all their quants.
The new experimental names: i1-, KS, KT
The llama.cpp world moves fast โ you'll see bleeding-edge repos like Qwen3.8-27B-i1-IQ4_KS_KT-GGUF. i1 is an experimental quant family (iterative/importance-based variants) that the community is currently benchmarking. General policy: unknown prefixes are a yellow flag. Check the model card, look for benchmarks or discussion, and if it's unproven, stick to Q4_K_M / IQ quants / UD quants โ they're the boring, reliable choices.
bpw in filenames
EXL2/EXL3-era files name themselves by exact bits per weight and even target size: qwen3.8-27b-abliterated-3.69bpw-12GB-MTP.gguf means ~3.69 bits per weight, a ~12 GB file, with a multi-token-prediction module. Sanity check: GB โ bpw ร params / 8. For 27B: 3.69 ร 27 / 8 = 12.5 GB โ. That one-line math decodes half the weird filenames on the Hub.
🧠 Knowledge Check
1. IQ quants (IQ2_XXS, IQ3_XS, IQ4_XS) are:
2. What does imatrix improve?
3. A 27B model at 3.69 bpw should be roughly: