-
Lesson 1: Why Quantization Exists
Precision, bit math, and why a 54 GB model can run on a 24 GB GPU.
-
Lesson 2: How Quantization Actually Works
Bits, groups, scales, bpw, and why some weights matter more than others.
-
Lesson 3: The GGUF Alphabet: K-Quants Decoded
Q4_K_M, Q5_K_S, Q6_K, Q8_0, XL tiers and the legacy Q4_0 โ with real Qwen3.8-27B file sizes.
-
Lesson 4: IQ Quants, imatrix, and Squeezing Below 4 Bits
IQ1/IQ2/IQ3/IQ4, importance matrices, and when going under 4 bits is a good idea.
-
Lesson 5: Beyond GGUF: The Format Zoo
AWQ, GPTQ, EXL2/EXL3, MLX, bitsandbytes NF4, FP8, NVFP4, QAT, AQLM, HQQ, EETQ, Marlin.
-
Lesson 6: The Fine-Tune Zoo: Non-Quant Suffixes
instruct vs base, thinking, uncensored/abliterated, flash, vision, MTP, MoE, dates, merges and more.
-
Lesson 7: Decoding Any Hugging Face Repo
Filename anatomy, what to verify, trusted publishers, and red flags.
-
Lesson 8: Choosing for Your Machine
The decision framework: hardware โ budget โ model size โ quant โ ecosystem โ test.
-
Grand Quiz
16-question comprehensive quiz across all lessons.