A Systems Guide to High-Performance Large Language Model Serving
Introduction: The LLM Inference Problem
- What Makes LLM Inference Hard
- A Day in the Life of an Inference Request
- Why vLLM Is Our Lens
- How to Use This Book
Chapter 1: Transformer Inference from First Principles
- The Transformer Forward Pass
- Self-Attention Mechanics and Complexity
- Prefill Versus Decode: Two Different Problems
- Why Attention Dominates Decode Performance
- The Role of Positional Embeddings
- FlashAttention and Modern Attention Kernels
Chapter 2: GPU Architecture for LLM Inference
- GPU Compute Units: SMs, CUDA Cores, and Tensor Cores
- The GPU Memory Hierarchy: DRAM, SRAM, Registers
- Memory Bandwidth as the Primary Bottleneck
- Arithmetic Intensity and Roofline Analysis
- PCIe, NVLink, and Interconnect Topology
- CUDA Streams and Concurrency Primitives
Chapter 3: Measuring What Matters: Inference Metrics and Workloads
- Latency Metrics: TTFT, ITL, and End-to-End Delay
- Throughput Metrics: Tokens per Second and Requests per Second
- GPU Utilization and Efficiency
- Workload Profiles: Interactive, Batch, and Hybrid
- Designing Reproducible Benchmarks
- Cost Models: Tokens, GPUs, and Dollars
Chapter 4: The KV-Cache and Memory Management
- Why the KV-Cache Exists
- KV-Cache Memory: How Much and Why It Matters
- Fragmentation and the Allocation Problem
- PagedAttention: The Core Innovation
- Block Tables and Virtual Memory for KV-Cache
- Eviction Policies and Cache Thrashing
Chapter 5: Batching and Scheduling Fundamentals
- Static Batching and Its Limitations
- Continuous Batching: Keeping the GPU Busy
- Request Scheduling Policies
- Preemption and Speculative Scheduling
- The Impact of Sequence Length Distribution
- Admission Control and Concurrency Limits
Chapter 6: Inside vLLM: Architecture and Execution Model
- The vLLM Process Model and Workers
- Engine, Scheduler, and Executor Roles
- The Request Lifecycle in vLLM
- Model Runner and GPU Worker Implementation
- Event Loop and CPU-GPU Coordination
- How vLLM Handles the API Server
Chapter 7: vLLM’s PagedAttention Implementation
- Virtual Block Allocation
- Block Table Construction and Maintenance
- The Attention Kernel with Paged KV-Cache
- Physical Memory Management
- Handling Variable Sequence Lengths
- Interaction with the Scheduler
Chapter 8: Continuous Batching in vLLM
- The Scheduler Loop
- Running, Waiting, and Swapped States
- Preemption Strategies
- Chunked Prefill
- Integration with PagedAttention
- Tuning Scheduler Parameters
Chapter 9: vLLM Configuration and Deployment Basics
- Installation: Dependencies, CUDA, and Compatibility
- Launching Your First vLLM Server
- OpenAI-Compatible API Configuration
- Docker and Container Deployments
- Model Selection and Loading Options
- Basic Configuration Parameters
Chapter 10: Quantization for Inference
- Why Quantize for Inference
- Weight-Only Quantization: AWQ, GPTQ, INT8, FP8
- Activation Quantization and Perplexity Impact
- vLLM’s Quantization Backends
- Choosing the Right Precision for Your Workload
- Quantized Model Serving Examples
Chapter 11: Tensor Parallelism and Multi-GPU Scaling
- The Limits of Single-GPU Inference
- How Tensor Parallelism Works
- vLLM’s Tensor Parallel Implementation
- Communication Overhead and NVLink
- Performance Scaling Curves
- Configuration for Multi-GPU Deployments
Chapter 12: Pipeline Parallelism and Expert Parallelism
- Pipeline Parallelism Basics
- Pipeline Bubble and Scheduling Challenges
- vLLM’s Pipeline Parallelism Support
- Mixture of Experts and Expert Parallelism
- DeepSeek and Qwen MoE Models in vLLM
- Hybrid Parallelism Strategies
Chapter 13: Distributed Inference Across Nodes
- Multi-Node Communication with Ray
- Network Topology and Bandwidth Planning
- Fault Tolerance Across Nodes
- Load Balancing Strategies
- Kubernetes Deployments with vLLM
- Operational Complexity of Distributed Serving
Chapter 14: Caching, Prefix Caching, and Context Optimization
- Prefix Caching in vLLM
- Hash-Based Cache Lookup
- Cache Validity and Model Versioning
- System Prompts and Repeated Contexts
- Multi-Query and Grouped-Query Attention
- Long-Context Optimization Techniques
Chapter 15: Speculative Decoding
- The Speculative Decoding Algorithm
- Draft Model Selection
- Token Tree Speculation
- vLLM’s Speculative Decoding Implementation
- Measuring Acceptance Rates and Speedup
- When Speculative Decoding Helps and Hurts
Chapter 16: CUDA Graphs and Kernel Optimization
- CUDA Kernel Launch Overhead
- CUDA Graphs in vLLM
- Attention Backend Selection
- FlashAttention Integration
- Profiling Kernel Performance
- When Low-Level Optimization Matters
Chapter 17: Production Serving Engineering
- Health Checks and Readiness Probes
- Autoscaling Strategies
- Rate Limiting and Admission Control
- Observability: Metrics, Logs, and Tracing
- Authentication and API Security
- Rolling Upgrades and Model Versioning
Chapter 18: Performance Engineering and Tuning
- Establishing Performance Baselines
- Profiling CPU and GPU with vLLM
- Diagnosing Bottlenecks
- Tuning for Interactive Workloads
- Tuning for Batch Workloads
- Configuration Decision Trees
Chapter 19: Troubleshooting and Failure Modes
- Out-of-Memory Failures
- CUDA and Driver Incompatibilities
- Scheduling and Throughput Issues
- GPU Utilization Problems
- Distributed System Failures
- Version and API Compatibility Issues
Chapter 20: The Ecosystem: vLLM Versus Alternatives
- Hugging Face Text Generation Inference
- TensorRT-LLM: NVIDIA’s Approach
- SGLang and Structured Generation
- llama.cpp and CPU/Edge Inference
- Feature Matrix and Workload Mapping
Chapter 21: The Economics and Future of LLM Inference
- Hardware Economics: GPUs and Inference Cost
- Capacity Planning Methodology
- Emerging Techniques and Trends
- What Will Remain Stable
- Architectural Recommendations
Conclusion: Engineering Judgment in Inference
- The Inference Engineering Mindset
- Key Principles That Endure
- Building Versus Buying
- A Final Architecture Recommendation
Back Matter
- Glossary of Key Terms
- Quick Reference: Key vLLM Parameters
- Performance Tuning Checklist
- References