Leanpub Header

Skip to main content

The Quantization Black Book

The Quantization Black Book
This book is 100% completeLast updated on 2026-08-27

The Quantization Black Book

Minimum price

$28.05

$28.05

You pay

Author earns

$

Also available for 1 book credit with a Reader Membership

PDF
112
Pages
About

About

About the Book

The definitive engineering reference for 4-bit quantization of large language models — from 70B all the way to 400B parameters. Written for the people who actually have to fold these models onto real hardware and keep them fast, accurate, and deployable.

This is not an overview. It is a deep technical field manual covering the full quantization stack: the math, the algorithms, the hardware setup, the calibration process, the runtime kernels, and the memory tricks that make the difference between a model that runs and a model that doesn't.

What's inside

  • Technical scope and core engineering objectives of quantization for 70B-400B class models
  • Mathematical foundation of Floating Point (FP32, FP16, BF16) vs Integer (INT8, INT4) representations
  • Quantization-aware outlier analysis for 70B+ parameter models and why they break naive schemes
  • High-performance computing environment setup, including multi-node H100 and A100 configurations
  • Linux kernel parameter tuning and NVMe scratch space for massive weight-sharding during quantization
  • Round-to-Nearest (RTN) quantization deep dive and its limitations at scale
  • GPTQ (Optimal Brain Quantization) theory and Hessian-based error compensation, implemented
  • Activation-aware Weight Quantization (AWQ) and preserving salient weights
  • GGUF and K-Quants mechanics for cross-platform deployment
  • EXL2 format deconstructed, with variable bitrate strategies for Llama-3 class models
  • QuIP# and the application of randomized Hadamard transforms to minimize quantization error
  • HQQ (Half-Quadratic Quantization) for fast, calibration-free 4-bit compression
  • Calibration dataset selection (The Pile, C4) for accurate weight distribution recovery
  • Layer-wise error tracking pipelines using KL-Divergence and Mean Squared Error
  • Offloading strategies for quantizing 400B models on limited VRAM via CPU-RAM swapping
  • Breaking the Memory Wall: optimizing CUDA kernels for 4-bit dequantization at runtime
  • QLoRA (Quantized Low-Rank Adaptation) for fine-tuning 4-bit artifacts
  • Double Quantization techniques to shrink the footprint of quantization constants themselves
  • KV-Cache quantization to 4-bit and 8-bit for long-context inference
  • And more — densely technical, cover to cover

Who this is for

ML engineers deploying local LLMs at scale, research engineers at AI startups, contractors compressing models for clients, and serious practitioners who are tired of scattered arXiv papers and want the full quantization stack in one place, with the math and the code.

Format

Delivered as PDF and TXT. Read the PDF cover to cover, or feed the TXT straight into your own local RAG, grep it, index it, pipe it into your local LLM. The book about quantization, ready to be consumed by a quantized model. + EXTRAS
PAGES 112

A note on responsibility

This material is provided for educational and research purposes. You are responsible for complying with the licenses of the models and tools you deploy, and with the laws of your jurisdiction. Build carefully and build for the right reasons.

Author

About the Author

Krzysztof Rybiński

I’m an independent AI systems developer, programmer, and technical researcher focused on the engineering behind modern computing systems.

My work spans artificial intelligence, machine learning, GPU computing, CUDA, quantum computing, automation, operating systems, cybersecurity, developer tooling, and high-performance software. I’m particularly interested in what happens beneath the abstractions: how systems actually execute, communicate, scale, fail, and can be engineered more efficiently.

I build and investigate practical systems rather than focusing exclusively on theory. This includes working with Python, C/C++, CUDA, Linux, AI infrastructure, local and distributed AI systems, GPU acceleration, quantum programming, automation frameworks, system administration, and security research.

My technical publishing is an extension of that work. I use Leanpub to document engineering knowledge, experiments, architectures, implementation techniques, and research-oriented material that can be useful to developers, engineers, researchers, and technically advanced readers.

I’m especially interested in emerging technologies where multiple disciplines meet — AI and systems engineering, GPU computing and machine learning, quantum computing and software development, automation and infrastructure, and cybersecurity and system architecture.

The goal is not simply to explain how a technology works, but to understand how to build with it, work around its limitations, examine its internals, and turn complex concepts into working systems.

I continuously develop, test, document, and refine these ideas through practical projects and technical research. My publications reflect that process: detailed engineering references created for people who want to go beyond the surface level and understand the technology they are working with.

https://businessofmachines.blogspot.com/

https://learn.microsoft.com/en-us/users/machinadeusex/

https://g.dev/machinadeusex

https://github.com/porucznikswext-source

Get the free Community Edition

You can get the free Community Edition in PDF or EPUB just by sharing your name and email address with the author.

 

Also by the Author

Also by the Author

The Leanpub 60 Day 100% Happiness Guarantee

Within 60 days of purchase you can get a 100% refund on any Leanpub purchase, in two clicks.

See full terms...

Earn $8 on a $10 Purchase, and $16 on a $20 Purchase

We pay 80% royalties on purchases of $7.99 or more, and 80% royalties minus a 50 cent flat fee on purchases between $0.99 and $7.98. You earn $8 on a $10 sale, and $16 on a $20 sale. So, if we sell 5000 non-refunded copies of your book for $20, you'll earn $80,000.

(Yes, some authors have already earned much more than that on Leanpub.)

In fact, authors have earned over $15 million writing, publishing and selling on Leanpub.

Learn more about writing on Leanpub

Free Updates. DRM Free.

If you buy a Leanpub book, you get free updates for as long as the author updates the book! Many authors use Leanpub to publish their books in-progress, while they are writing them. All readers get free updates, regardless of when they bought the book or how much they paid (including free).

Most Leanpub books are available in PDF (for computers) and EPUB (for phones, tablets and Kindle). The formats that a book includes are shown at the top right corner of this page.

Finally, Leanpub books don't have any DRM copy-protection nonsense, so you can easily read them on any supported device.

Learn more about Leanpub's ebook formats and where to read them

Write and Publish on Leanpub

You can use Leanpub to easily write, publish and sell in-progress and completed ebooks and online courses!

Leanpub is a powerful platform for serious authors, combining a simple, elegant writing and publishing workflow with a store focused on selling in-progress ebooks.

Leanpub is a magical typewriter for authors: just write in plain text, and to publish your ebook, just click a button. (Or, if you are producing your ebook your own way, you can even upload your own PDF and/or EPUB files and then publish with one click!) It really is that easy.

Learn more about writing on Leanpub