Leanpub Header

Skip to main content

vLLM and the Engineering of Fast LLM Inference

A Systems Guide to High-Performance Large Language Model Serving

vLLM and the Engineering of Fast LLM Inference
This book is 100% completeLast updated on 2026-09-07

Building fast LLM inference is about far more than turning the right knobs. This book takes you inside vLLM and the systems behind it, showing how memory, scheduling, GPUs and distributed serving shape real-world performance. Learn how to measure what matters, find bottlenecks and build inference infrastructure that scales without wasting money.

Minimum price

$19.00

$29.00

You pay

Author earns

$

Also available for 1 book credit with a Reader Membership

PDF
EPUB
WEB
APP
168
Pages
About

About

About the Book

This book teaches experienced engineers how to build, measure, and optimize production LLM inference systems. Using vLLM as its primary case study, it walks you from the first principles of transformer inference and GPU architecture through vLLM's internal design, advanced performance techniques, distributed deployment, and operational reliability. You will learn not just how to configure vLLM, but how to reason about memory bottlenecks, scheduling trade-offs, parallelism strategies, and cost-performance optimization so you can architect inference infrastructure that scales. The material is grounded in primary sources, source code, and measured benchmarks; version-dependent details are flagged explicitly so the underlying engineering principles endure beyond any single release.

Author

About the Author

Steve Publications

Steve is a technology professional with more than 20 years of experience in software development, server infrastructure, cybersecurity, vulnerability research and reverse engineering. Throughout his career, he has designed, secured, analyzed and tested complex software and infrastructure, with a particular focus on understanding how systems fail and how they can be made more secure.

Outside of work, Steve enjoys sharing knowledge with the technology community. He collaborates with researchers, industry experts and technology professionals to write practical books covering software development, cybersecurity, cloud computing, networking, DevOps, artificial intelligence and enterprise technologies. His books focus on practical learning through clear explanations, real-world examples and hands-on exercises. With more than two decades of industry experience, his goal is to help IT professionals, students and technology enthusiasts build useful skills and stay current in a rapidly changing industry.

We believe readers deserve to know how our books are created. Most of our authors are not native English speakers, so we use AI to help translate, proofread manuscripts, fix grammar, improve sentence structure and make technical explanations easier to read. AI is used as an editing tool only. It does not replace the research, technical knowledge or hands-on experience behind our books. Some of our authors also prefer to remain anonymous for privacy or professional reasons. In those cases, we publish their work under a different name. The author's name may be different, but the quality of the content and our review process remain the same.

Every book is written, reviewed and maintained by experienced technology professionals, with contributions from our private technical community of more than 420 engineers and researchers. We spend far more time validating technical accuracy and keeping our content up to date than generating text. We are always interested in working with experienced professionals who have deep expertise in a particular technology or domain. If you would like to publish a book with us or help review an existing manuscript, we'd love to hear from you. Send us a message describing your area of expertise. We are especially interested in niche technologies, specialized skills and emerging topics that are underrepresented in existing technical literature.

If you look through the contents of our books, you'll see practical examples, detailed explanations and material that is regularly updated. Our goal is to publish books that professionals can actually rely on, not low-effort AI-generated content. If you ever feel that one of our books does not meet that standard, Leanpub offers a 60-day money-back guarantee. Feel free to request a refund if you are not satisfied with your purchase.

Contents

Table of Contents

A Systems Guide to High-Performance Large Language Model Serving

Introduction: The LLM Inference Problem

  1. What Makes LLM Inference Hard
  2. A Day in the Life of an Inference Request
  3. Why vLLM Is Our Lens
  4. How to Use This Book

Chapter 1: Transformer Inference from First Principles

  1. The Transformer Forward Pass
  2. Self-Attention Mechanics and Complexity
  3. Prefill Versus Decode: Two Different Problems
  4. Why Attention Dominates Decode Performance
  5. The Role of Positional Embeddings
  6. FlashAttention and Modern Attention Kernels

Chapter 2: GPU Architecture for LLM Inference

  1. GPU Compute Units: SMs, CUDA Cores, and Tensor Cores
  2. The GPU Memory Hierarchy: DRAM, SRAM, Registers
  3. Memory Bandwidth as the Primary Bottleneck
  4. Arithmetic Intensity and Roofline Analysis
  5. PCIe, NVLink, and Interconnect Topology
  6. CUDA Streams and Concurrency Primitives

Chapter 3: Measuring What Matters: Inference Metrics and Workloads

  1. Latency Metrics: TTFT, ITL, and End-to-End Delay
  2. Throughput Metrics: Tokens per Second and Requests per Second
  3. GPU Utilization and Efficiency
  4. Workload Profiles: Interactive, Batch, and Hybrid
  5. Designing Reproducible Benchmarks
  6. Cost Models: Tokens, GPUs, and Dollars

Chapter 4: The KV-Cache and Memory Management

  1. Why the KV-Cache Exists
  2. KV-Cache Memory: How Much and Why It Matters
  3. Fragmentation and the Allocation Problem
  4. PagedAttention: The Core Innovation
  5. Block Tables and Virtual Memory for KV-Cache
  6. Eviction Policies and Cache Thrashing

Chapter 5: Batching and Scheduling Fundamentals

  1. Static Batching and Its Limitations
  2. Continuous Batching: Keeping the GPU Busy
  3. Request Scheduling Policies
  4. Preemption and Speculative Scheduling
  5. The Impact of Sequence Length Distribution
  6. Admission Control and Concurrency Limits

Chapter 6: Inside vLLM: Architecture and Execution Model

  1. The vLLM Process Model and Workers
  2. Engine, Scheduler, and Executor Roles
  3. The Request Lifecycle in vLLM
  4. Model Runner and GPU Worker Implementation
  5. Event Loop and CPU-GPU Coordination
  6. How vLLM Handles the API Server

Chapter 7: vLLM’s PagedAttention Implementation

  1. Virtual Block Allocation
  2. Block Table Construction and Maintenance
  3. The Attention Kernel with Paged KV-Cache
  4. Physical Memory Management
  5. Handling Variable Sequence Lengths
  6. Interaction with the Scheduler

Chapter 8: Continuous Batching in vLLM

  1. The Scheduler Loop
  2. Running, Waiting, and Swapped States
  3. Preemption Strategies
  4. Chunked Prefill
  5. Integration with PagedAttention
  6. Tuning Scheduler Parameters

Chapter 9: vLLM Configuration and Deployment Basics

  1. Installation: Dependencies, CUDA, and Compatibility
  2. Launching Your First vLLM Server
  3. OpenAI-Compatible API Configuration
  4. Docker and Container Deployments
  5. Model Selection and Loading Options
  6. Basic Configuration Parameters

Chapter 10: Quantization for Inference

  1. Why Quantize for Inference
  2. Weight-Only Quantization: AWQ, GPTQ, INT8, FP8
  3. Activation Quantization and Perplexity Impact
  4. vLLM’s Quantization Backends
  5. Choosing the Right Precision for Your Workload
  6. Quantized Model Serving Examples

Chapter 11: Tensor Parallelism and Multi-GPU Scaling

  1. The Limits of Single-GPU Inference
  2. How Tensor Parallelism Works
  3. vLLM’s Tensor Parallel Implementation
  4. Communication Overhead and NVLink
  5. Performance Scaling Curves
  6. Configuration for Multi-GPU Deployments

Chapter 12: Pipeline Parallelism and Expert Parallelism

  1. Pipeline Parallelism Basics
  2. Pipeline Bubble and Scheduling Challenges
  3. vLLM’s Pipeline Parallelism Support
  4. Mixture of Experts and Expert Parallelism
  5. DeepSeek and Qwen MoE Models in vLLM
  6. Hybrid Parallelism Strategies

Chapter 13: Distributed Inference Across Nodes

  1. Multi-Node Communication with Ray
  2. Network Topology and Bandwidth Planning
  3. Fault Tolerance Across Nodes
  4. Load Balancing Strategies
  5. Kubernetes Deployments with vLLM
  6. Operational Complexity of Distributed Serving

Chapter 14: Caching, Prefix Caching, and Context Optimization

  1. Prefix Caching in vLLM
  2. Hash-Based Cache Lookup
  3. Cache Validity and Model Versioning
  4. System Prompts and Repeated Contexts
  5. Multi-Query and Grouped-Query Attention
  6. Long-Context Optimization Techniques

Chapter 15: Speculative Decoding

  1. The Speculative Decoding Algorithm
  2. Draft Model Selection
  3. Token Tree Speculation
  4. vLLM’s Speculative Decoding Implementation
  5. Measuring Acceptance Rates and Speedup
  6. When Speculative Decoding Helps and Hurts

Chapter 16: CUDA Graphs and Kernel Optimization

  1. CUDA Kernel Launch Overhead
  2. CUDA Graphs in vLLM
  3. Attention Backend Selection
  4. FlashAttention Integration
  5. Profiling Kernel Performance
  6. When Low-Level Optimization Matters

Chapter 17: Production Serving Engineering

  1. Health Checks and Readiness Probes
  2. Autoscaling Strategies
  3. Rate Limiting and Admission Control
  4. Observability: Metrics, Logs, and Tracing
  5. Authentication and API Security
  6. Rolling Upgrades and Model Versioning

Chapter 18: Performance Engineering and Tuning

  1. Establishing Performance Baselines
  2. Profiling CPU and GPU with vLLM
  3. Diagnosing Bottlenecks
  4. Tuning for Interactive Workloads
  5. Tuning for Batch Workloads
  6. Configuration Decision Trees

Chapter 19: Troubleshooting and Failure Modes

  1. Out-of-Memory Failures
  2. CUDA and Driver Incompatibilities
  3. Scheduling and Throughput Issues
  4. GPU Utilization Problems
  5. Distributed System Failures
  6. Version and API Compatibility Issues

Chapter 20: The Ecosystem: vLLM Versus Alternatives

  1. Hugging Face Text Generation Inference
  2. TensorRT-LLM: NVIDIA’s Approach
  3. SGLang and Structured Generation
  4. llama.cpp and CPU/Edge Inference
  5. Feature Matrix and Workload Mapping

Chapter 21: The Economics and Future of LLM Inference

  1. Hardware Economics: GPUs and Inference Cost
  2. Capacity Planning Methodology
  3. Emerging Techniques and Trends
  4. What Will Remain Stable
  5. Architectural Recommendations

Conclusion: Engineering Judgment in Inference

  1. The Inference Engineering Mindset
  2. Key Principles That Endure
  3. Building Versus Buying
  4. A Final Architecture Recommendation

Back Matter

  1. Glossary of Key Terms
  2. Quick Reference: Key vLLM Parameters
  3. Performance Tuning Checklist
  4. References

Get the free sample chapters

Click the buttons to get the free sample in PDF or EPUB, or read the sample online here

The Leanpub 60 Day 100% Happiness Guarantee

Within 60 days of purchase you can get a 100% refund on any Leanpub purchase, in two clicks.

See full terms...

Earn $8 on a $10 Purchase, and $16 on a $20 Purchase

We pay 80% royalties on purchases of $7.99 or more, and 80% royalties minus a 50 cent flat fee on purchases between $0.99 and $7.98. You earn $8 on a $10 sale, and $16 on a $20 sale. So, if we sell 5000 non-refunded copies of your book for $20, you'll earn $80,000.

(Yes, some authors have already earned much more than that on Leanpub.)

In fact, authors have earned over $15 million writing, publishing and selling on Leanpub.

Learn more about writing on Leanpub

Free Updates. DRM Free.

If you buy a Leanpub book, you get free updates for as long as the author updates the book! Many authors use Leanpub to publish their books in-progress, while they are writing them. All readers get free updates, regardless of when they bought the book or how much they paid (including free).

Most Leanpub books are available in PDF (for computers) and EPUB (for phones, tablets and Kindle). The formats that a book includes are shown at the top right corner of this page.

Finally, Leanpub books don't have any DRM copy-protection nonsense, so you can easily read them on any supported device.

Learn more about Leanpub's ebook formats and where to read them

Write and Publish on Leanpub

You can use Leanpub to easily write, publish and sell in-progress and completed ebooks and online courses!

Leanpub is a powerful platform for serious authors, combining a simple, elegant writing and publishing workflow with a store focused on selling in-progress ebooks.

Leanpub is a magical typewriter for authors: just write in plain text, and to publish your ebook, just click a button. (Or, if you are producing your ebook your own way, you can even upload your own PDF and/or EPUB files and then publish with one click!) It really is that easy.

Learn more about writing on Leanpub