From First Principles to Production Inference
Chapter 1: What Problem Are We Solving?
- The Sequence Modeling Problem
- What Came Before: RNNs, LSTMs, and Their Limits
- Why Parallelism Matters
- The Attention Insight
- A Map of This Book
Chapter 2: Tensors, Shapes, and Matrix Multiplication
- Tensors as n-Dimensional Arrays
- Element-Wise Operations
- Matrix Multiplication and Broadcasting
- Batch Dimensions and the Last-Dimension Convention
- A Running Example of Shape Tracking
Chapter 3: Neural Networks From First Principles
- Linear Layers and Affine Transformations
- Activation Functions: Why Non-Linearity Is Necessary
- Loss Functions: Cross-Entropy and What It Measures
- Backpropagation: The Chain Rule in Practice
- Optimizers: Gradient Descent and Its Variants
Chapter 4: From Text to Numbers: Tokenization and Embeddings
- Character-Level versus Token-Level Representations
- Byte-Pair Encoding and Subword Tokenization
- Building a Vocabulary and Handling Unknown Tokens
- The Embedding Layer as a Lookup Table
- Embeddings as Learned Vector Representations
Chapter 5: The Attention Mechanism
- What Is Attention, Intuitively?
- Query, Key, and Value: The Retrieval Analogy
- Dot-Product Similarity and the Scaling Problem
- Softmax as a Probability Distribution Over Positions
- Computing Attention: A Step-by-Step Numerical Example
Chapter 6: Multi-Head Attention
- Why One Head Is Not Enough
- Projecting Into Subspaces
- Concatenating and Mixing Head Outputs
- The Tensor Algebra of Multi-Head Attention
- What Different Heads Learn
Chapter 7: Positional Information
- The Permutation Invariance Problem
- Sinusoidal Positional Encodings: Derivation and Properties
- Learned Positional Embeddings
- Relative Position and Distance-Aware Attention
- Masking: Why and When to Zero Out Positions
Chapter 8: Feed-Forward Networks
- The Position-Wise Feed-Forward Network
- Why We Need a Separate Transformation After Attention
- Activation Functions: ReLU, GELU, and SwiGLU
- Width, Depth, and Computational Trade-Offs
- The Complete Transformer Block
Chapter 9: Implementing a Transformer in NumPy
- Designing the Data Structures
- The Embedding and Positional Encoding Layer
- Implementing Scaled Dot-Product Attention From Scratch
- Building Multi-Head Attention
- The Complete Encoder Stack
- Testing and Verification
Chapter 10: Building a Decoder and Encoder-Decoder Architecture
- Causal Masking and Autoregressive Generation
- The Decoder Self-Attention Layer
- Cross-Attention: Attending to the Encoder
- Output Projection and Logits
- Generating Text One Token at a Time
Chapter 11: Automatic Differentiation and Backpropagation Through a Transformer
- Gradients of Matrix Multiplication
- Gradients Through Softmax and Attention
- Gradients of the Feed-Forward Network
- The Residual Connection and Gradient Flow
- Verification Against Numerical Gradients
Chapter 12: Training a Transformer From Scratch With PyTorch
- Why PyTorch: Autograd and Dynamic Graphs
- Preparing a Tiny Training Dataset
- The Complete Training Loop
- Monitoring Loss and Generating Samples
- Saving, Loading, and Resuming Checkpoints
Chapter 13: GPT-Style Decoder-Only Models
- Why Decoder-Only Works for Language Generation
- Architectural Choices in the GPT Family
- Layer Normalization Placement and RMSNorm
- RoPE: Rotary Positional Embeddings
- SwiGLU and the Modern Feed-Forward Design
Chapter 14: BERT and Encoder-Only Models
- Bidirectional Attention and Pretraining Objectives
- Masked Language Modeling in Detail
- Next Sentence Prediction
- Fine-Tuning for Downstream Tasks
- The Encoder-Only Limitations
Chapter 15: Encoder-Decoder and Instruction-Tuned Models
- The T5 Architecture and Text-to-Text Pretraining
- Prefix-LM and Hybrid Designs
- Instruction Tuning and Alignment
- Parameter-Efficient Fine-Tuning: LoRA and Friends
- Architectural Trade-Offs in Practice
Chapter 16: Scaling Techniques and Mixture of Experts
- The Scaling Laws: Data, Parameters, and Compute
- Sparse Expert Routing
- Mixture-of-Experts Training Dynamics
- Load Balancing and Expert Capacity
- MoE in Production Models
Chapter 17: Long-Context Techniques
- The O(n²) Attention Bottleneck
- Grouped-Query and Multi-Query Attention
- Sliding Window Attention
- Memory Compression and Token Pruning
- Extending Context Through Data and Training
Chapter 18: Initialization, Optimization, and Scheduling
- Initialization Strategies and Variance Control
- AdamW and Why It Works for Transformers
- Learning Rate Warmup and Cosine Decay
- Gradient Clipping and Stability
- Hyperparameter Recipes That Work
Chapter 19: Distributed Training: Parallelism Strategies
- Data Parallelism and Gradient Synchronization
- Model Parallelism: Tensor and Pipeline
- ZeRO and Optimizer State Sharding
- Communication Overhead and Topology
- Real-World Training Runs
Chapter 20: Memory and Precision
- The Memory Budget of a Transformer Training Run
- Mixed Precision: FP16, BF16, and Loss Scaling
- Gradient Accumulation for Effective Batching
- Activation Checkpointing
- Memory Profiling and Debugging
Chapter 21: Inference Optimization
- Pre-Fill Versus Decode: Two Different Regimes
- KV Caching: Why and How
- Continuous Batching and Scheduling
- Speculative Decoding
- Speculative Decoding Correctness and Sampling
Chapter 22: Quantization and Efficient Formats
- Why Quantization Works for Neural Networks
- Post-Training Quantization: INT8 and Below
- Quantization-Aware Training
- Grouped Quantization and Outlier Handling
- GGUF, GGML, and the Open-Source Ecosystem
Chapter 23: FlashAttention and GPU Hardware
- GPU Memory Hierarchy and Tensor Cores
- The Attention Kernel: Why It Is Slow
- FlashAttention: IO-Aware Recomputation
- FlashAttention-2 and Beyond
- Measuring and Profiling GPU Utilization
Chapter 24: Serving Architectures
- Throughput Versus Latency Trade-Offs
- Model Serving Frameworks
- Multi-GPU Inference and Expert Parallelism
- Caching, Load Balancing, and Autoscaling
- Cost Modeling and Efficiency
Chapter 25: What Transformers Can and Cannot Do
- What Attention Can Represent
- Transformers as Approximate Algorithm Implementers
- Inductive Biases: What the Architecture Assumes
- Positional Encoding Limits and Extrapolation
- Failure Modes and Systematic Errors
Chapter 26: The History and Future of Attention
- The Pre-Transformer Landscape: 2014-2017
- “Attention Is All You Need”: Context and Impact
- The Architecture Arms Race: 2018-Present
- Post-Transformer Alternatives
- Open Questions and Where to Look Next
Conclusion
- The Transformer in One Page
- What You Can Do Now
- The Next Decade of Sequence Modeling