A Practical Guide to Training Your Own Transformer-Based AI in Python
Introduction: Why Build from Scratch?
- The Black Box Problem
- What You Will Build
- Prerequisites and How to Use This Book
- A Note on Hardware Requirements
Chapter 1: Tokens, Vocabularies, and Tokenization
- From Text to Numbers: The Tokenization Pipeline
- Character-Level vs Word-Level vs Subword Tokenization
- Building a Byte-Pair Encoding (BPE) Tokenizer
- Vocabulary Size, Special Tokens, and Edge Cases
- Walkthrough: Training and Analyzing a BPE Tokenizer on Shakespeare
Chapter 2: Embeddings: Turning Tokens into Vectors
- The Embedding Layer as a Lookup Table
- Learning vs Pre-trained Embeddings
- Vector Space Geometry and Similarity
- Implementing Embedding Layers in PyTorch
- Walkthrough: Visualizing an Embedding Space
Chapter 3: The Attention Mechanism
- What Is Attention and Why It Matters
- Scaled Dot-Product Attention from First Principles
- Multi-Head Attention: Parallelizing Understanding
- Causal Masking for Decoder-Only Models
- Walkthrough: Tracing Attention Through a Sentence
Chapter 4: Positional Encoding and Sequence Structure
- The Permutation Invariance Problem
- Sinusoidal Absolute Positional Encodings
- Learned Position Embeddings
- Rotary Positional Embeddings (RoPE)
- Exercise: Implement Three Position Encoding Schemes
Chapter 5: Building the Decoder-Only Transformer Architecture
- The Decoder-Only Design Decision
- Feed-Forward Networks and MLP Blocks
- Residual Connections and Layer Normalization
- Assembling the Full Transformer Block
- The Complete Decoder-Only Model
- Exercise: Build a 3-Layer Decoder from Scratch
Chapter 6: Mixture of Experts Architectures
- Why Mixture of Experts?
- How MoE Layers Work
- Switch Transformers and Single-Expert Routing
- Implementing an MoE Transformer Block
- Training Considerations for MoE Models
- Notable MoE Models
- Exercise: Train a Tiny MoE Model
Chapter 7: Data Preparation for Language Model Training
- The Data Landscape: Where Does Training Data Come From?
- Cleaning and Filtering Pipeline Design
- Deduplication Strategies
- Dataset Mixing and Domain Balancing
- Synthetic Data Generation Strategies
- Exercise: Build a Mini C4 Dataset
Chapter 8: The Training Loop: Loss, Optimizers, and Gradient Flow
- Cross-Entropy Loss and Next-Token Prediction
- The AdamW Optimizer and Why It Works
- Learning Rate Schedules: Warmup, Cosine Decay, and Beyond
- Gradient Clipping and Training Stability
- The Complete Training Loop
- Exercise: Train on a Tiny Dataset and Monitor Loss Curves
Chapter 9: Scaling Laws and Compute-Optimal Training
- The Empirical Laws of Language Model Scaling
- Chinchilla and Compute-Optimal Training
- Applying Scaling Laws to Your Training Run
- Understanding the Limits of Scaling
- Exercise: Analyze Scaling on Your Hardware
Chapter 10: Memory-Efficient Training Patterns
- The Memory Wall: Why Models Don’t Fit in GPU RAM
- Mixed-Precision Training with BF16/FP16
- Gradient Accumulation for Effective Batch Sizes
- Activation Checkpointing and Recomputation
- Exercise: Train a 10x Larger Model on the Same Hardware
Chapter 11: Distributed Training and Parallelism Strategies
- Data Parallelism and Distributed Data Parallel (DDP)
- Tensor Parallelism for Massive Models
- Pipeline Parallelism: Splitting the Forward Pass
- Fully Sharded Data Parallel (FSDP)
- Exercise: Multi-GPU Training Setup
Chapter 12: Checkpointing, Experiment Tracking, and Reproducibility
- Checkpointing Strategies and Recovery
- Experiment Tracking: Metrics, Configs, and Artifacts
- Reproducibility: Seeds, Determinism, and Hardware Variability
- Logging Design for Long Training Runs
- Exercise: Set Up a Production-Style Training Dashboard
Chapter 13: Fine-Tuning: LoRA, QLoRA, and Instruction Tuning
- The Fine-Tuning Landscape: Full vs Parameter-Efficient
- Low-Rank Adaptation (LoRA) from First Principles
- Quantized LoRA (QLoRA) for Memory-Constrained Fine-Tuning
- Instruction Tuning Dataset Design
- Exercise: Fine-Tune a Model on a Custom Task
Chapter 14: Alignment: RLHF and Beyond
- Why Raw Models Need Alignment
- The RLHF Pipeline: Reward Models and PPO
- Direct Preference Optimization (DPO) as a Simpler Alternative
- Constitutional AI and Rule-Based Alignment
- Walkthrough: Training a Model with Direct Preference Optimization
Chapter 15: Evaluation: Metrics, Benchmarks, and Red Teaming
- Perplexity as a Training Metric vs Real-World Performance
- Standard Benchmark Suites (MMLU, GSM8K, HumanEval)
- Qualitative Evaluation and LLM-as-Judge
- Safety Testing and Red Teaming Methodologies
- Exercise: Build an Evaluation Harness
Chapter 16: Deployment: Inference, Quantization, and Serving
- Inference Optimization: KV Cache and Speculative Decoding
- Quantization Strategies: FP16 -> INT8 -> INT4
- Building a Serving API with FastAPI
- Containerization and Production Deployment Patterns
- Exercise: Deploy Your Model Behind a Live API
Chapter 17: Advanced Inference: Speculative Decoding and Beyond
- The Inference Bottleneck
- How Speculative Decoding Works
- Practical Speedups and Real-World Performance
- N-Gram Speculation
- Exercise: Implement and Benchmark Speculative Decoding
Chapter 18: Context Window Extension and Long-Sequence Techniques
- The Context Window Problem
- Positional Encoding Interpolation
- YaRN: Yet Another RoPE Extension
- Attention Sparsity for Long Sequences
- Practical Recommendations for Long Context
Chapter 19: Retrieval-Augmented Generation (RAG)
- Why RAG? The Knowledge Problem
- The RAG Pipeline
- Building a Simple RAG System
- RAG vs Fine-Tuning
Chapter 20: Model Compression Beyond Quantization
- The Compression Landscape
- Pruning: Removing Unnecessary Weights
- Knowledge Distillation: Teaching Smaller Models
- Combining Compression Techniques
Capstone Project: From Raw Data to Live Inference
- Project Setup and Architecture Overview
- Data Preparation: Curating a 100M-Token Dataset
- Training Run: Configuration, Execution, and Monitoring
- Evaluation: Benchmarking Against Baselines
- Deployment: Containerized API with Health Checks
- Post-Mortem: What Went Well and What Would Be Different
Conclusion: The Road Ahead
- What You Have Accomplished
- Where to Go From Here: Scaling Up
- The State of Language Models in 2025
- A Final Word on Building vs Using
