A Practical Guide to Compressing Large Language Models Without Losing Their Intelligence
Introduction: Why Compress the Mind?
- What You Will Learn
- How This Book Is Organized
- Prerequisites
Chapter 1: The Compression Imperative
- The Size Explosion: From GPT-2 to Today
- Why Raw Parameters Are a Bottleneck
- What Quantization Actually Does
- A Brief History of Model Compression
- What This Book Covers (and Does Not)
Chapter 2: Foundations of Numerical Precision in Deep Learning
- Floating-Point Arithmetic: FP32, FP16, BF16
- Integer Representations: INT8, INT4, INT2
- Mixed Precision and the NF4 Format
- How GPUs and CPUs Handle Different Datatypes
- The Information Theory of Weight Distributions
Chapter 3: Post-Training Quantization vs. Quantization-Aware Training
- Post-Training Quantization (PTQ): The Quick Path
- Quantization-Aware Training (QAT): The Careful Path
- Weight-Only Quantization: The LLM Sweet Spot
- Activation Quantization and the Outlier Problem
- Hybrid Approaches and Per-Token Strategies
Chapter 4: Calibration – Teaching Precision to Models
- The Role of Calibration Data
- Min-Max vs. Percentile Clipping
- Moving Average and Histogram-Based Methods
- Optimal Perceptual Quantization (OPQ)
- How Much Calibration Data Do You Really Need?
Chapter 5: GPTQ – Greedy One-Shot Quantization
- The Hessian Approximation Idea
- Layer-by-Layer Greedy Optimization
- The GPTQ Algorithm Step by Step
- AutoGPTQ: The Practical Implementation
- Strengths, Limitations, and Typical Results
- Mathematical Derivation: Why the Hessian Works
- Complete Production GPTQ Quantization Script
- Production Readiness Checklist for GPTQ
Chapter 6: AWQ – Activation-Aware Weight Quantization
- The Activation Magnitude Insight
- Weight Scaling Before Quantization
- The AWQ Algorithm: Smoothing and Rescaling
- AWQ vs. GPTQ: A Head-to-Head Comparison
- Practical Usage with AutoAWQ
- Complete Production AWQ Quantization Script
- Production Readiness Checklist for AWQ
- Case Study: Deploying a 70B Model on a Single A100
Chapter 7: GGUF and the llama.cpp Ecosystem
- The GGML Legacy and GGUF’s Design
- K-Quants: Q4_0, Q4_K_S, Q5_K_M, Q8_0
- How llama.cpp Runs Quantized Models on CPU
- Performance on CPUs vs. GPUs with Metal/Vulkan
- Community Toolchains and Model Hubs
- Importance Matrix (imatrix) Quantization: A Deep Dive
- Complete GGUF Conversion Pipeline
Chapter 8: BitsAndBytes and the NF4 Revolution
- The bitsandbytes Library Architecture
- NormalFloat4: Why It Beats Plain INT4
- QLoRA: Fine-Tuning in 4 Bits
- 8-Bit Adam and Optimizer Quantization
- Practical Usage with the Transformers Library
- Complete Production QLoRA Fine-Tuning Script
- Production Readiness Checklist for QLoRA
Chapter 9: Unsloth – Speed Through Quantized Fine-Tuning
- The Fine-Tuning Bottleneck
- Unsloth’s Architecture and Optimizations
- Patched Transformers: How the Speedup Works
- Benchmarks: Unsloth vs. Standard QLoRA
- Practical Usage and Limitations
Chapter 10: Advanced Frameworks – Bartowski, ByteShape, and Apex
- Bartowski’s Quantization Pipeline
- ByteShape: Understanding This Approach
- NVIDIA Apex: From Mixed-Precision Training to FP8
- Other Notable Methods: SmoothQuant, ZeroQuant, SPA
- The Fragmentation Problem in Toolchains
Chapter 11: Deployment Scenarios and Hardware Considerations
- Local Inference on Consumer GPUs
- CPU-Only Deployment and Edge Devices
- High-Throughput Server Serving
- Mobile and On-Device LLMs
- Case Study: Deploying a Customer Support Chatbot on a Single GPU
- The Memory Bandwidth Bottleneck
Chapter 12: Benchmarking, Best Practices, and Decision Framework
- A Reproducible Benchmarking Protocol
- How to Measure Quantization Quality
- Perplexity, Accuracy, and Latency Benchmarks
- Common Pitfalls and Debugging Tips
- The Quantization Decision Framework
- Future Directions: What’s Next in Model Compression
